"Do not guess": a single sentence cuts hallucinated fields from 71% to 20% across 16 frontier models
🔎 A six-word sentence against thousands of false data points
If you build data extraction pipelines with LLMs, you've probably already lived this moment: the agent returns an impeccable table, complete, every field filled in… and half the values exist nowhere on the source page. The model "completed" the gaps out of pure statistical reflex.
This is exactly the problem the "Do not guess" benchmark tackles, published by Earn an Honest Dollar, with additional analysis from RuntimeWire on September 27, 2026. The finding is disarmingly simple: adding a single sentence to the prompt — "Use null for any field whose value is not on the page. Do not guess." — drops invented fields from 70.7% to 20.2% on average.
Even more striking: the 16 frontier models tested, without exception, hallucinate more without this sentence. From GPT-5.5 to Gemini 3.1 Pro to Claude Opus 4.7, none escapes the fabrication reflex. The flaw is therefore not one of a particular model — it's structural, and it's fixable in one line.
But let's be honest: 20% of fields still being invented is huge for a production pipeline. And as the benchmark's protocol shows, the most effective traps are not empty pages, but pages stuffed with plausible bait. That's the whole story of this benchmark.
Key Takeaways
- One sentence is enough: adding "Use null for any field whose value is not on the page. Do not guess." reduces hallucinated fields from 405 out of 573 (70.7%) to 116 out of 574 (20.2%) on average across 16 frontier models.
- All models hallucinate without the sentence: none of the 16 tested proved an exception, suggesting a structural training bias rather than an isolated flaw.
- The traps are frighteningly effective: a former price "Was $493.00" is taken as the current price by all 16 models without the instruction; only one resists with it.
- The magic sentence has a cost: the arXiv study 2601.02023 shows it can reduce recall on certain tasks — a measurable "safety tax," especially late in the context.
- Downstream validation remains essential: downstream checkers catch 23 to 38 invented values out of 49, but systematically miss "near-miss" errors.
- A simple fetch + GPT-6 Luna costs $0.0049 per run and hallucinates 5 fields out of 36, versus 24 for a commercial extraction service without the instruction.
Recommended tools
To put these lessons into practice on your own pipelines:
| Tool | Main use | Price (September 2026) | Ideal for |
|---|---|---|---|
| OpenRouter | Multi-model access to test your prompts across several LLMs | Pay-as-you-go billing, free models available | Comparing "do not guess" behavior across models |
| Hermes Agent | Configure multiple providers and models in a single agent | Open source | Industrializing extraction pipelines with per-model instructions |
| Simple fetch + LLM (homegrown scraping) | Extraction without a specialized service | ~$0.005 per run (benchmark, September 2026) | Small volumes, full control over the prompt |
Note the most counterintuitive finding of the benchmark: the best-performing solution on hallucinated fields is not a sophisticated extraction service, but a simple HTTP fetch followed by a well-prompted LLM call. We come back to this below.
What does this benchmark actually measure?
Direct answer: it measures how many fields a model invents when the requested information doesn't exist on the page.
The protocol, detailed on the benchmark's page, is based on 42 pairs of synthetic web pages covering 7 page types. In each pair: one page contains the answer to the extraction question, the other doesn't — but both contain the same "bait" (decoy), a plausible, visible, tempting piece of data that is wrong for the requested task.
The model must extract specific fields. The score counts fields that were filled in even though they didn't exist on the page. Two conditions are tested: with and without the sentence "Use null for any field whose value is not on the page. Do not guess."
Key figures
| Condition | Invented fields | Rate |
|---|---|---|
| Without instruction | 405 / 573 missing fields | 70.7% |
| With instruction | 116 / 574 missing fields | 20.2% |
The nearly identical denominators (573 vs 574) allow a direct comparison. The drop is more than 50 percentage points — for six words added to the prompt.
The limitations, stated by the authors of the analysis
The analysis by RuntimeWire provides the necessary nuances: these are synthetic pages with artificially designed traps, and each contestant was evaluated only once. The denominators count responses, not measurements on real websites.
In other words: the clear instruction helps massively on this type of test, without any absolute guarantee on the real web, where ambiguity and deception take far more varied forms. It's a benchmark, not a universal law. But it's the first to properly quantify this reflex across so many frontier models.
Why does such a small sentence have such a big effect?
Direct answer: because an LLM is optimized for completion, not for saying "I don't know". The instruction short-circuits that default reflex.
A language model trained on data scraped from the web has learned, billions of times over, to fill in missing fields. A human scraper who encounters an empty field on a product listing often has an "intuition" of what that field should contain — and alignment training has probably reinforced this "helpful" behavior. Without instructions to the contrary, filling in is the default option. Saying "null" requires an explicit decision.
The sentence "Do not guess" turns that implicit decision into an explicit rule. It's a textbook case of what we already described in our prompt engineering guide: targeted negative instructions focused on a specific behavior outperform vague directives like "be precise and don't hallucinate".
A universal flaw, not a model flaw
The fact that all 16 tested models hallucinate more without the sentence is perhaps the most important takeaway. It's not "Gemini 3.1 Pro is reliable but GPT-5.5 isn't": it's "all frontier models share the same bias, and it can be corrected on the prompt side, not by choosing a different model".
For your pipelines, this means that switching from one model to another doesn't solve this problem. Changing the prompt does. This is consistent with our article on prompt debugging: before blaming the model, you need to audit what you're asking it — or rather, what you're forgetting to ask it.
The baits: where the benchmark gets really interesting
Direct answer: the benchmark's traps are designed to look like legitimate answers, and the models fall for them 16 out of 16 times.
Three examples from the protocol are worth a closer look, because they look exactly like what you'll find on real e-commerce or editorial pages:
- The old price: a page displays a struck-through "Was $493.00", followed by the current price. Without the instruction, all 16 models report 493 as the current price. With it, only one resists. The "Was X" formatting is so common that the model has learned to ignore it visually — but its tokenizer doesn't.
- The adjacent fact-checker: a fact-check label is placed near an author field, prompting the model to merge the two entities.
- The confusable date: an update date, placed in the same spot where a publication date usually sits, ends up being extracted as the publication date.
These traps work because they exploit not the model's ignorance, but its competence. It recognizes seemingly legitimate patterns and applies them to the wrong field. This is exactly the kind of error that spot-check human validation fails to catch, because the invented value is plausible.
The humiliation of the commercial extraction service
Direct answer: Firecrawl, a dedicated extraction service, invented 24 out of 36 fields — worse than 13 of the 16 models with the simple instruction.
The detail is delicious: Firecrawl copied the bait every single time, without exception. Meanwhile, a handcrafted setup — simple page fetch + GPT-6 Luna with the instruction — only invented 5 out of 36 fields, for a total run cost of $0.0049.
What should we take away from this? Two things, both embarrassing for the ecosystem:
- "Turnkey" extraction services don't have a monopoly on reliability. A home-made, transparent pipeline, where you control the prompt, can outperform a commercial product. This is one more argument for mastering your own stack, and for configuring your own models and providers rather than delegating blindly.
- The opacity of the agent-to-agent marketplace. RuntimeWire reminds us that this benchmark illustrates the central question of the Earn an Honest Dollar marketplace: can an agent buyer know when a service is bluffing? When a provider claims to extract "reliably," it doesn't tell you whether its pipeline contains the equivalent of "Do not guess" — nor whether it has ever been tested.
If you want to reduce the costs of your experiments, also note that you can test several models for free via OpenRouter or Groq before switching to production — a model's anti-hallucination behavior can be verified for a few cents.
The flip side: the "safety tax"
Direct answer: reducing fabrication doesn't mean winning everything — anti-hallucination instructions can cause recall to plummet on information that is nonetheless present.
This is what the study arXiv 2601.02023, "Not All Needles Are Found" documents, extending the needle-in-a-haystack benchmark to Gemini-2.5-flash, ChatGPT-5-mini, Claude-4.5-haiku and Deepseek-v3.2-chat. The finding: instructions like "Don't Make It Up" or "Do not guess" make some models too conservative.
The numbers speak for themselves:
| Condition | Literal extraction (ChatGPT-5-mini) |
|---|---|
| Standard prompt | 96.4% |
| Anti-hallucination prompt | 90.3% |
| Anti-hallucination + context at max capacity (272k tokens) | 72% |
| Facts concentrated in certain distributions | up to 0% |
This "safety tax" worsens toward the end of the context: the longer the context, the more the cautious model refuses to output information that is clearly present. On certain distributions of concentrated facts, extraction literally collapses to zero.
In practice: the magic phrase reduces fabrications, but it can make you lose real data. If your pipeline needs to maximize recall (for example in competitive intelligence where every missed field counts), you need to measure this trade-off on your own data, not assume it.
20% residual errors: why downstream validation remains mandatory
Direct answer: because a prompt phrase doesn't detect hallucinations, it reduces their frequency. The remaining 20% require a downstream safety net.
The benchmark tested two downstream verification strategies, with double-edged results:
- GPT-6 Luna catches 38 invented values out of 49, without rejecting any of the 47 correct values.
- Jev 1.13 catches 23 out of 49, with no false rejections among 48 correct values.
The zero false positive rate is remarkable: the verifiers don't wrongly reject correct data. But they all miss "near-miss" errors — a cooking time presented as a prep time, a price slightly close to the true value. The most dangerous hallucinations are precisely those that look the most like the truth.
Hence an architecture I recommend for any serious extraction pipeline:
- Prompt with explicit instruction ("Do not guess" or equivalent), possibly tuned to your recall tolerance.
- Structured output with a confidence or justification field: ask the model to cite the source passage for each field. A field without a citation is a suspect.
- Downstream verification by a second model, with a separate prompt.
- Statistical hallucination detection — our article on the phi-first method that detects hallucinations in a single token shows that you can spot a model inventing information before its response is even finished, without a costly second call.
The combination of "hardened prompt + token-level detection + semantic validation" costs a fraction of a second and a few cents per page. It turns a 20% error rate into something publishable.
What This Changes for Your Pipelines, in Concrete Terms
Direct answer: add the sentence, test it on your real data, measure the fabrication/recall trade-off, and keep downstream validation.
A five-step plan, no detours:
- Audit your current prompts. If your extraction instructions don't contain an explicit directive about missing fields, you're probably paying a 30 to 70% price in invented fields without knowing it. The exact sentence from the benchmark is a good starting point.
- Test on pairs of pages with and without the information. It's the only way to measure your own fabrication rate. A test on your typical pages (products, articles, internal records) takes an afternoon.
- Measure the safety tax on your data. Check that recall on fields that are actually present doesn't collapse, especially if your pages are long. Compare before/after the instruction.
- Choose your models with full knowledge of the facts. All frontier models share the bias, but their sensitivity to the safety tax differs. Use multi-model access — free to start — to refine.
- Stack the safeguards. Prompt, token-level detection, validation by a second model. No single link is enough on its own; the chain, however, holds.
And if you follow security news, note that defense in depth is also the logic behind systems like Palo Alto Networks' Unit 42 multi-model harness (September 22), which combines Claude Mythos 5, GPT-5.6-Cyber, and open models to validate one another — another echo of the same lesson picked up by AI Weekly: one model alone gets it wrong; several, cross-checked, correct one another.
❌ Common Mistakes
Mistake 1: believing a more expensive model hallucinates less
All 16 frontier models tested hallucinate more without the instruction, including the most expensive and best-ranked ones. The problem is structural, not a tier effect. Solution: harden the prompt first, change models second.
Mistake 2: treating the absence of data as data
The "Was $493.00" trap works because the model extracts a value present on the page, but attributes it to the wrong field. Your extraction schemas should distinguish between "field absent" (null) and "field filled with source justification". A field without justification is suspect by default.
Mistake 3: applying "Do not guess" everywhere without measuring
The safety tax documented by the arXiv study can cause you to lose real data, especially in long contexts. Solution: measure both fabrication AND recall on your own pages before generalizing the instruction.
Mistake 4: relying on an extraction service for reliability
Firecrawl invented 24 out of 36 fields by systematically copying the bait — worse than most raw models with a single sentence of prompt. Solution: ask your providers about their anti-hallucination methodology, and test them on trap pages like those in the benchmark.
Mistake 5: confusing downstream validation with an absolute guarantee
GPT-6 Luna catches 38 out of 49 errors, but misses the near-misses. A verifier that never rejects correct values (zero false positives) is excellent — but it is not infallible. Keep human sampling on critical fields.
❓ Frequently asked questions
What is the exact sentence to copy?
« Use null for any field whose value is not on the page. Do not guess. » That's the instruction tested by the benchmark. You can adapt it to your language and schema, but keep the dual structure: an explicit rule (null if absent) followed by a prohibition (do not guess).
Why 573 and 574 fields, and not the same number?
The two conditions test nearly identical but not strictly identical datasets; the one-field difference doesn't affect comparability. RuntimeWire notes that these denominators count responses on synthetic pages, not measurements on the real web.
Does this result hold for real web pages?
Partially. The benchmark uses synthetic pages with designed traps, and a single run per contestant. The instruction helps massively in this controlled setting, but the real web is more ambiguous and more deceptive. Validate it on your own data.
Should I stop using extraction services?
No, but test them. The Firecrawl result (24 invented fields out of 36) shows that a commercial product can be less reliable than a fetch + LLM well-prompted at $0.0049 per run. Compare on your use cases before committing.
How to detect hallucinations without a second call?
The phi-first method involves analyzing the probability distribution of the first generated token for each field: a hesitant model before inventing leaves detectable signatures. It's detailed in our dedicated article, and it's compatible with a high-throughput pipeline.
Does this work on open and free models?
The benchmark covers proprietary frontier models, but the mechanism — an explicit instruction countering a training reflex — should apply to open models. Since the bias is universal across the 16 tested, the instruction is probably beneficial everywhere; measure it yourself on the free models available via aggregators.
✅ Conclusion
A six-word sentence drops hallucinated fields from 71% to 20% across 16 frontier models — but the remaining 20%, and the safety tax it can induce, prove that no prompt can replace downstream validation. Test the "Do not guess" instruction on your next extraction pipeline, and read our method for detecting hallucinations in a single token to complete your safeguards.