Prototype API that scores e-commerce product pages for scam and dropshipping signals using a parallel LangGraph pipeline of LLM skills
  • Python 97.7%
  • Makefile 1.6%
  • Dockerfile 0.7%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Floris van Beers 7ccb8d212d chore: merge automated 0.5.1 bump from main (superseded)
The legacy release workflow ran once on the history rewrite push and
committed a malformed version bump. This cleanup removes that workflow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-26 22:56:34 +02:00
.github chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
scripts chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
src/veridex chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
tests chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
.dockerignore chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
.env.example chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
.gitignore chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
.pre-commit-config.yaml chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
.python-version Initial commit 2026-02-26 22:12:25 +01:00
CHANGELOG.md docs: update changelog commit links after history rewrite 2026-09-26 22:47:38 +02:00
docker-compose.yml chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
Dockerfile chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
LICENSE chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
Makefile chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
pyproject.toml chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
README.md chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00
uv.lock chore: prepare repository for open-source release 2026-09-26 22:34:16 +02:00

Veridex

CI License: MIT Python 3.12

Veridex scores e-commerce product pages for scam and dropshipping signals. A browser extension sends the HTML of the page the user is viewing; the backend runs a LangGraph pipeline of independent analysis "skills" in parallel and has an LLM judge condense their evidence into a 0–100 reliability score with a one-line explanation.

POST /v1/analyze  →  { "score": 18, "explanation": "New domain, AliExpress-sourced images and templated 5-star reviews." }

Project status: prototype. Veridex is an exploratory prototype, not a finished product. It was built to try out what an agentic LLM pipeline can do for shopping-fraud detection, and to put a modern Python stack to work (FastAPI, LangGraph, structured LLM output, asyncpg, Playwright, uv). Accounts, quotas and Paddle billing are included to show how it would fit into a real service. It is not maintained for production use.

How it works

flowchart LR
    A[POST /v1/analyze] --> P[preprocess<br/><sub>strip HTML to text</sub>]
    P --> C{classify<br/><sub>LLM</sub>}
    C -- article / non-product --> X[early exit<br/><sub>score = -1</sub>]
    C -- marketplace listing / storefront --> S1[page_content]
    C --> S2[html_source_signals]
    C --> S3[review_integrity]
    C --> S4[claims_verification]
    C --> S5[domain_age]
    C --> S6[exif_check]
    C --> S7[reverse_image_search]
    S1 & S2 & S3 & S4 & S5 & S6 & S7 --> J[judge<br/><sub>LLM → score + explanation</sub>]

Each skill is a small class that reads the shared graph state and appends a plain-text finding. The graph fans out to all active skills with LangGraph's Send API, so they run concurrently, and a reducer merges their findings before the judge runs.

Skill What it checks
page_content LLM scan of the listing text for scam copy, fake urgency, missing return policy, white-label branding
html_source_signals Regex match on known dropship apps (Oberlo, DSers, CJ Dropshipping, …) and supplier CDNs, plus an LLM pass over script sources, meta tags and comments
review_integrity Extracts reviews from JSON-LD / microdata (following a "see all reviews" link if needed) and looks for templated praise, rating skew and date bursts
claims_verification Extracts concrete product claims ("waterproof to 50 m", "FDA approved") and rates their plausibility against price tier and known standards
domain_age WHOIS lookup; flags newly registered, short-lived or privacy-hidden domains typical of pop-up stores
exif_check Downloads the main product images and compares EXIF metadata (dates, GPS, camera, editing software) with the listing
reverse_image_search Reverse-searches product images on Google and reports hits on AliExpress, Temu, DHgate and similar sources

Adding a skill

class ShippingTimeSkill(Skill):
    name = "shipping_time"
    description = "Flags multi-week shipping estimates."
    always_run = False  # only runs when "shipping_time" is in the request's `flags`

    async def run(self, state: AnalysisState) -> dict[str, list[str]]:
        ...
        return {"skill_results": ["[shipping_time]\n<finding>"]}

Register an instance in SKILLS in src/veridex/graph/__init__.py.

API

Interactive docs are served at /docs. All /v1/accounts/me*, /v1/analyze and /v1/admin/* routes use HTTP Basic auth.

Method Path Description
POST /v1/analyze Analyse a page: { "page_content": "<html>…", "url": "https://…", "flags": [] }
POST /v1/accounts Create an account (server-to-server; requires the X-Signup-Key header)
GET /v1/accounts/me Current account
GET /v1/accounts/me/usage Analyses used today and remaining quota
GET /v1/admin/accounts List/search accounts (admin only)
POST /v1/admin/accounts/{username}/downgrade | /block Revoke premium / block, cancelling any Paddle subscription
POST /v1/paddle/webhook Paddle Billing webhook (subscription.activated / subscription.canceled)
GET /v1/health Liveness probe

Accounts are free (5 analyses/day), premium (25/day), admin (unlimited) or blocked. Pages that are not analysed (score -1) don't count towards the quota.

Running it

Requirements: uv, Docker, and an OpenAI API key.

cp .env.example .env       # add your VERIDEX_OPENAI_API_KEY
docker compose up --build  # API on http://localhost:8080, Postgres alongside

Create an account and analyse a page:

curl -X POST localhost:8080/v1/accounts -H "X-Signup-Key: change-me" -H "Content-Type: application/json" \
  -d '{"username": "demo", "password": "correct-horse", "contact_email": "demo@example.com"}'

curl -s https://example-shop.com/product/123 > page.html
jq -n --rawfile html page.html '{page_content: $html, url: "https://example-shop.com/product/123"}' \
  | curl -X POST localhost:8080/v1/analyze -u demo:correct-horse -H "Content-Type: application/json" -d @-

Promote an account to admin with uv run scripts/promote_admin.py <username>.

Configuration

Defaults live in src/veridex/settings.toml (loaded with Dynaconf). Override any key with a VERIDEX_-prefixed environment variable or a git-ignored src/veridex/.secrets.toml.

Variable Default
VERIDEX_OPENAI_API_KEY — Required
VERIDEX_LLM_MODEL gpt-4o-mini Used by every node and skill
VERIDEX_POSTGRES_URI postgresql://veridex:veridex@localhost:5432/veridex Tables are created on startup
VERIDEX_SIGNUP_SECRET — Account creation is disabled while empty
VERIDEX_FREE_DAILY_LIMIT / VERIDEX_PREMIUM_DAILY_LIMIT 5 / 25 Per UTC day
VERIDEX_PADDLE_WEBHOOK_SECRET — Webhook requests are rejected while empty
VERIDEX_PADDLE_API_KEY / VERIDEX_PADDLE_ENVIRONMENT — / production Used to cancel subscriptions from the admin routes

Development

make install           # uv sync + pre-commit hooks
make lint              # ruff, mypy (strict), pydoclint
make test-unit         # fast, no services needed
make test              # unit + integration tests against a throwaway Postgres container
make run               # uvicorn with auto-reload

Unit tests cover the deterministic parts of every skill, the graph routing (with a stubbed LLM), SSRF protection and webhook signature verification. Integration tests exercise the HTTP API end-to-end against a real Postgres, with the analysis graph replaced by a test double so no LLM calls are made.

src/veridex/
├── main.py            # FastAPI app, lifespan (graph + DB pool), middleware
├── routes/            # analyze, accounts, admin, paddle webhook, health
├── graph/
│   ├── __init__.py    # graph wiring + skill registry
│   ├── nodes.py       # preprocess, classify, early_exit, judge
│   ├── state.py       # shared LangGraph state
│   └── skills/        # one module per skill
├── auth.py            # Basic auth, account tiers
├── db.py              # schema + queries (asyncpg)
├── llm.py             # shared chat model
└── net.py             # SSRF-safe HTTP fetching

Design notes

  • Untrusted input fetching. Skills download URLs taken from user-submitted HTML. net.fetch_public only allows http(s), resolves the host and refuses loopback, private, link-local (cloud metadata) and reserved addresses, and re-validates every redirect hop and caps response size.
  • Auth. Passwords are bcrypt-hashed off the event loop; unknown usernames are checked against a dummy hash so they can't be distinguished from wrong passwords by timing.
  • Billing webhooks fail closed without a configured secret, reject signatures older than five minutes, never modify admin or blocked accounts, and only honour a cancellation for the subscription currently linked to the account.

Limitations and responsible use

  • Scores are LLM judgements over heuristic evidence. They are a signal, not a verdict, and can be wrong in both directions.
  • Page content is passed to the LLM, so a hostile page can attempt prompt injection to influence its own score.
  • reverse_image_search automates Google Search through a headless browser. That is against Google's Terms of Service and may be blocked at any time. If you run this yourself, replace it with a licensed reverse-image API or remove it from SKILLS.
  • WHOIS data is often redacted and inconsistently formatted across registries.

License

MIT