Why we built it
We measured the models, then built the check.
Long before Lenz, we spent years following the skeptic movement — the podcasts, the books, the arguments — on one practical question: how do you know what’s true and what’s not?
It has only got harder to answer: more of what people read is written to persuade rather than to inform. And when AI started writing the documents, answers and reports people act on, that question became an operational one.
So we measured the frontier models against recent, real-world claims, chosen to keep their verdicts out of the models’ training data as far as possible. On 63% of those claims the models did not agree on the answer, and they reported high confidence while disagreeing. Read the study →
We built on the research literature on multi-agent adversarial debate, evidence-based fact verification and multi-model judging panels, and turned it into a product. A full Lenz verification runs multiple models from multiple vendors across five stages, keeps both sides of the argument in the record, and returns the full trace with the verdict. The API and the SDKs put that process inside any workflow, and the MCP server puts it inside Claude and ChatGPT with no code at all.
Our goal is to check more and more AI output before anyone acts on it, so that fewer false claims get out there. A truer world is a simpler and kinder one to live in.