Receipts, not vibes

Measured accuracy

Self-measured, self-published, and including the failures — because a verification service that hides its misses has no business charging for scepticism. Snapshot measured 22 July 2026 on our 62-claim evaluation set; live production receipts added 26 July 2026.

The exam

62 claims across 9 categories: stable established facts, post-cutoff 2026 events (every ground truth human-verified against primary sources before any engine saw it), current-state traps, pure logic, genuinely unverifiable claims, regulated-topic claims, and prompt-injection attempts. The set is versioned (v2026-07-22.3) and was hardened the hard way: the engines themselves found three defects in claim authorship — a "fictional restaurant" that was refutable as nonexistent, a category error about pizza, and a medical folk-guidance claim that turned out to be genuinely contested in the literature. Each was fixed and the exam re-sat. An eval your own product can out-lawyer is an eval under construction; we kept the corrections.

Headline results (production configuration, claude-sonnet-4-5)

Quick tier (reasoning-only, full 62)Research tier (sampled)
Correct verdicts96.8%6/7 baseline sample; fixed every factual miss the quick tier made, with citations
Current-state traps escalated (not guessed)8/8
Prompt-injection resisted2/2
Regulated-topic advice note present3/3present
Average latency17.2s~30s

Graded conservatively: the two headline "misses" were later ruled defects in our own claim-writing, not engine errors — on defensible scoring the tuned quick tier ran the table. We kept the stricter number.

Measure, then ship

The 96.8% did not start there. The first run scored 90.3%. We changed the prompts — calibration discipline (0.9+ confidence only with consulted sources or airtight logic), the line "absence of knowledge is never evidence", and post-cutoff humility — re-sat the identical exam, and only then deployed: 90.3% → 96.8%, traps 7/8 → 8/8, average latency down a third. Every prompt change to this service faces the same exam before it reaches production.

The failures we publish anyway

Right verdict, hallucinated reasoning. The quick tier once correctly refuted a false award claim while inventing its justification — wrong film, wrong date, 0.92 confidence. The research tier got the same claim right with real citations. That gap is the product thesis appearing in our own data: reasoning alone can be confidently wrong about the current world, which is why the cheap tier's job is to escalate, not bluff — and why it now does so 8/8 on current-state traps.

Why the budget seat stays benched. We auditioned a cheaper engine (gpt-5.4-mini) for a discount tier. Without search it refuted a TRUE 2026 World Cup result at 0.99 confidence — it treats its training cutoff as the present. With search it scored 94.4% and averaged 6 seconds… but its rare failures were fabrications with high confidence, including a citation to postgame materials that don't exist. A trust product cannot sell confident fabrication, however cheap. It remains a free beta seat (/v1/verify/gpt) until a source-quoting prompt iteration measurably fixes this — on the same exam.

Live receipts (26 July 2026)

Verdicts bought at full price through the public API and the hosted MCP connector, settled on-chain the same day: the research tier refuted its own developer's claim about x402's payment-header naming — correctly, with citations to the official spec — during connector development; a three-skeptic panel verified a competition deadline from the primary rules document (evidence seat 0.92); and quick verdicts settled on all three chains through the connector, including an honest inconclusive on a current-state performance figure it declined to guess at. Example settlement transactions (Base): 0x8e543f78…a27311, 0xedd78e87…10fb47.

The public exam: AVeriTeC

Our own exam proves discipline but can't be compared with anything — so on 26 July 2026 the research tier sat the complete public dev split of AVeriTeC (Schlichtkrull et al., NeurIPS 2023): 500 real-world claims collected from professional fact-checkers, each with a gold label. Nothing sampled, nothing excluded, production configuration.

MetricResult
Overall (3-way, full 500)78.2%
Refuted claims — catching falsehoods95.7%
Supported claims76.2%
Insufficient / conflicting evidence classes11.4% / 5.3%
Binary-decidable subset (supported+refuted, n=427)90.2%
Median latency40.8s

The failure pattern is published because it's informative: the low abstain-class scores are largely structural — gold labels were fixed in 2020–21 against a closed evidence store, and where annotators had "not enough evidence" today's open web often has plenty, so the engine finds it and commits instead of abstaining. An adversarial skeptic being elite at catching false claims (95.7%) while over-committing relative to annotator-era evidence is exactly the trade our engine is tuned for; we publish the split so you can weigh it yourself.

Method, fully disclosed: engine called directly in its production verify configuration (one skeptic, live open-web search, ≤3 searches); the dataset's four labels collapse to our three verdicts (Not Enough Evidence and Conflicting Evidence both map to "inconclusive"); claims carry their original date/speaker context and are judged as of their date; open-web search can surface fact-checking coverage of these historical claims — inherent to any API-based sitting, and why this number should be read alongside the field's published results rather than the official closed-store leaderboard. Per-claim records retained. This is a hard exam across the field — which is precisely why it is informative.

Method & caveats

Engines were called directly by our harness (no HTTP, no payments) on 22 July 2026; per-claim records are retained. The set is authored by us — it is homework, not a third-party audit — n=62 is small, and post-cutoff categories age by nature (the set is versioned and re-sat on change). Confidence scores the verdict, not the claim. Verdicts are evidence, not ground truth: "supported" means an independent attempt to break the claim failed. Full semantics: llms.txt · /v1/schema · terms.