Skip to content
Rex St. John
toollive

Verafy Swarm Explorer

Put a claim to hundreds of models at once and show where they disagree, rather than returning one confident answer from one of them. The product the company was built around.

Next.js / OpenRouter / Claude API

Demo

Problem

Ask a model whether something is true and you get one confident sentence. Ask a different model and you may get the opposite, equally confidently, and nothing in either answer tells you that you are in contested territory.

A single verdict hides the one thing you most need when the question actually matters: whether the models agree.

Solution

Put the claim to a panel — over 250 models reachable through the explorer — and show the breakdown rather than the average. Each model's verdict, each confidence, and the spread between them, on one screen.

Analysis

The screenshot is the whole argument. The Churchill claim comes back No at 93% overall, which reads as settled until you look at the row beneath it: Meta's model is at 0% and Qwen at 95%. A headline number would have buried that, and the disagreement is more useful than the verdict.

The cost is that it asks more of the reader than a green tick does. That is a real product trade-off, and it is the reason the FactCheck extension exists — same argument, delivered where people actually meet claims.

Result

The flagship product of Verafy, which raised $17M and built a developer community over 4,000 before the crypto market turned.

The company did not survive. The idea did: most of what I have built since is a descendant of it.

Screens

  • The Swarm Explorer fact-check interface, showing three past queries each with a verdict, a confidence score, and a per-model breakdown where individual models disagree
    The product argument in one screen. Churchill comes back No at 93% — but Meta's model is at 0% and Qwen at 95%. The spread is the information.
  • The Verafy Swarm Explorer landing page, showing the product on desktop and mobile beside copy about comparing responses across hundreds of models
    The pitch, which is the same as the architecture: compare across models rather than trusting one.