AI Super Simplified
Edition 307

This AI Refuses to Answer Unless It Can Prove It. Even Its $27M Makers Won’t Say “Hallucination-Proof.” | Edition 307

Edition 307 — Pramaana Labs raised $27M to prove AI answers in Lean. The honest part is the phrase it refuses to use.

By Jerry Croteau

If you saw this story at all, you probably saw it under a headline close to the one we had queued up: a startup is using math to make AI hallucination-proof.

So we went looking for who actually said that. It was not the startup.

Pramaana Labs raised $27 million in June, led by Khosla Ventures, to put a mathematical proof engine underneath AI answers in tax, law, and medicine. The money is real, the technique is real, and the people behind it are not lightweight. But the phrase “hallucination-proof” appears nowhere in the company’s press release, nowhere on its website, and in none of its founders’ quotes.

What appears instead is a sentence we find considerably more interesting: “It will refuse to answer before it proves.”

That is a smaller promise. It is also a much better one, and the gap between those two sentences is the whole story.

What actually happened

On June 17, Pramaana Labs announced a $27 million seed round led by Khosla Ventures, with Accel, Boldcap, Nexus Venture Partners, Premji Invest and Unbound also participating. The company was founded in 2025 and is headquartered in Palo Alto.

The three co-founders are IIT Madras alumni with relevant scar tissue rather than generic AI resumes:

  • Ranjan Rajagopalan (CEO) led Google Maps Moderation — keeping a planet-scale live database accurate.
  • Krishnan Raghavan spent three years at Glean building the first version of Glean Assistant and, in the company’s own words, fighting hallucinations, until he realised that solving them is a research problem, not a product problem.
  • Sanjay Ganapathy was a Staff Research Engineer at Google DeepMind and a core contributor to the Gemini models, where he built the tool-use system.

The backer list is the part that should make you take this seriously. Early backers include Pushmeet Kohli, VP at Google DeepMind, and Sriram Rajamani, Corporate VP at Microsoft CoreAI — both genuine formal-verification researchers, not celebrity angels. The tax work is advised by Danny Werfel, the former IRS Commissioner, and built with researchers from Yale Law School and Stanford.

This is not a thin wrapper with a good deck.

What the machine actually does

Today’s AI produces an answer by predicting what a good answer looks like. It is astonishingly effective and it comes with one permanent catch, which Pramaana’s own website states more bluntly than any critic would: “AI learned to sound right before it learned to be right.”

Pramaana’s approach adds a step that has nothing to do with prediction. It runs in three parts:

  1. Encode the rules. Domain experts — tax attorneys, clinicians, legal scholars — translate the actual rules of a field into Lean, a formal proof language. Lean does not do vibes. A statement is accepted only when every logical step is checked back to the axioms.
  2. Translate the question. When you ask something, the system converts your question into a formal statement in that same language.
  3. Prove it, or refuse. A proof engine searches for a chain of reasoning connecting the encoded rules to the answer. If it finds one, you get a machine-checkable proof. If it does not, the system tells you which rule breaks and why — instead of producing a confident paragraph.

The third step is the genuinely unusual product decision. Nearly every AI product ever shipped is optimised to always return something. This one is built to say I cannot prove that and stop.

The one strong claim they do make — read it carefully

We went through Pramaana’s materials specifically hunting for an absolute claim. There is exactly one, in the press release:

“It has never produced a confidently wrong verified answer.”

Read that again, slowly, because every word is load-bearing.

It is not never wrong. It is not never hallucinates. It is a claim about the answers that came out with a proof attached — a narrower set than “everything the product says.” Questions it refused, questions it formalised incorrectly, and questions outside the encoded rules are all sitting outside that sentence.

And it is a company self-report. There is no benchmark number, no error bar, no independent auditor, and no published evaluation behind it. That is not an accusation of dishonesty — it is a six-week-old funding announcement, and this is what funding announcements look like. But “we have never seen it happen” and “it cannot happen” are different statements, and only the second one would justify the word hallucination-proof.

The company does not use that word. We are not going to hand it to them.

Why we are this careful: it has been promised before

This is not a hypothetical worry. The exact promise has already been made in one of the exact domains Pramaana is targeting, and it has already been measured.

In 2024, a Stanford team — Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher Manning and Daniel Ho — ran the first preregistered evaluation of the big commercial legal-research AI tools. They noted what the vendors were claiming at the time: Casetext said its method would eliminate hallucinations, Thomson Reuters said avoid, and LexisNexis advertised “hallucination-free” legal citations.

The measured result: Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinated between 17% and 33% of the time - Lexis+ AI and Ask Practical Law AI at 17%, Westlaw AI-Assisted Research at 33%. The paper’s verdict on the marketing: the providers’ claims are overstated.

Worth being fair in both directions here. Only Lexis+ AI was clearly better than raw GPT-4, which hallucinated 43% of the time. Westlaw AI-Assisted Research still hallucinated 33%, and the paper found that Westlaw and Ask Practical Law AI answered fewer queries than GPT-4 without the answers they did give being significantly more trustworthy. The failure was not that the technology did nothing. The failure was the absolute word.

The background problem, meanwhile, keeps growing. Damien Charlotin, a researcher at HEC Paris, maintains a public database of court decisions where a judge found that someone filed AI-hallucinated material. When we checked it on 6 August 2026, it listed 1,847 cases — and it only counts decisions where a court explicitly found or clearly implied hallucinated content, not every accusation.

So the demand for what Pramaana is selling is completely real. That is precisely why the language around it needs to stay honest.

The actual weak point, and it is not the prover

Here is the part most coverage skipped, and it is the part that decides whether any of this works.

Lean is not the risk. Lean is about as trustworthy as software gets — if it certifies a proof, the proof holds. The risk lives one step earlier, in what researchers call autoformalization: turning messy human rules into formal logic.

Somebody has to translate a tax code, a clinical guideline, or a statute into Lean. If that translation is subtly wrong — a missed exception, a definition that varies by jurisdiction, a threshold off by one — then the proof engine will faithfully, rigorously, prove the wrong thing. You get a certificate of correctness for a rule that was never the real rule. And it arrives looking more authoritative than an ordinary AI answer, not less.

This is a documented, active research problem, not a theoretical quibble. Recent work on machine-generated Lean finds a wide gap between output that compiles and output that actually means the right thing — in the FormalMATH evaluation (arXiv:2505.02735), about 61% of syntactically valid Lean statements were filtered out as semantically misaligned. Formal-methods practitioners have a long-standing name for the underlying issue: a proof is only ever correct relative to its specification.

To Pramaana’s considerable credit, they are not hiding this. Their own engineering blog published a post on 14 July titled “How Lean Handles Ambiguity, Vagueness, and Gaps in Law”, subtitled Three Formalization Blockers and How to Deal with Them. A company selling a miracle does not publish a post about the three things that block the miracle.

That is the strongest signal in this entire story — stronger than the $27M, stronger than the advisor list. They are writing publicly about their own hard part.

The demo that actually landed

One result from their blog is worth knowing about, because it shows the approach doing something no chatbot can do.

Pramaana formalised Indian income tax law in Lean and, in the process, found a bug in the law itself — a marginal-relief flaw where earning more money could reduce take-home pay. They then proved a fix.

Sit with the direction of that. The system was not caught making an error. The system caught an error in the rules — the kind of edge case that exists because no human ever checked every interaction between every threshold. That is a category of work ordinary AI cannot do at all, because ordinary AI has no notion of a rule being internally inconsistent. It only knows what text tends to follow what text.

The company also reports that its systems formalised and proved all six International Mathematical Olympiad 2026 problems in Lean 4 in under seven hours for less than $150. That is their own claim on their own blog rather than an independent evaluation — but it is a checkable kind of claim, which is more than most AI marketing offers.

Run claims through a proof-carrying pipeline and watch it land on PROVED, REFUSED — and the case that should worry you, PROVED AND WRONG. Nothing you enter leaves the page. · Open full-screen ↗
The claimWho said itWhere it stands
AI made hallucination-proof by mathHeadlines and aggregatorsNot a claim Pramaana makes anywhere in its own materials
It will refuse to answer before it provesPramaana press releaseThe company actual design claim, and a modest testable one
Never produced a confidently wrong verified answerPramaana press releaseCompany self-report. No benchmark, no number, no third party
Hallucination-free legal citationsLexisNexis marketing, 2024Stanford measured those tools at 17 to 33% hallucination
Four claims about AI correctness. Only one of them came with an independent measurement.

What this means for you

You cannot buy this yet. There is no consumer product, no signup, no free tier — it is a six-week-old seed round aimed at tax firms, hospitals and compliance teams. So the useful takeaway is not a tool. It is a test you can start applying immediately.

Split your AI questions into two piles.

Pile one: questions with rules underneath them. Tax treatment, benefits eligibility, contract terms, dosage limits, filing deadlines, refund policies. These have a written rule that decides the answer. For anything in this pile, the right follow-up is always the same: “Which specific rule produces that answer, and quote it.” If the model cannot name the rule, you have received a plausible sentence, not an answer. This is the pile Pramaana is going after, and it is the pile where being wrong costs you money.

Pile two: questions with judgment underneath them. Is this a good strategy, is this writing any good, what should I do about a difficult colleague. No proof engine will ever help here, because there is no rule to check against. Fluent, confident AI is genuinely useful in this pile — just never mistake the confidence for verification.

The practical habit, starting today: when an AI answer would cost you real money or real risk if it were wrong, make it show you the rule. Not a citation — citations get fabricated, which is exactly what those 1,847 court cases are about. The rule, quoted, so you can go read it yourself.

And when a company tells you its AI cannot be wrong, notice that the most credible company in this space is the one carefully declining to say it.