📑 Table of contents

27.5% of the web is already AI-generated: the preprint that quantifies the poisoning of training data

Deep Tech 🟢 Beginner ⏱️ 10 min read 📅 2026-10-02

27.5% of the web is already AI-generated: the preprint that quantifies the poisoning of training data

🔎 The web is silently tipping over

An arXiv preprint dated September 30, 2026 has just put a number on the table: 27.5% of tokens in high-quality web crawls (FineWeb, June 2026) are AI-generated. Two months later, in August 2026, the proportion climbs to 31.1%. For reference, it was only 10.1% in June 2024 and 0.1% in June 2021. The acceleration is exponential.

The researchers trained 800 language models (from 19.9 million to 973 million parameters), systematically varying the ratio of AI tokens injected into the training data. The verdict: a slight benefit for data-starved models, then a clear reversal into negative territory once a satiation threshold is crossed. Training on unfiltered web data at the August 2026 rate already costs 1.6x more compute than a purely human subset. By 2028, the bill climbs to 3.0x.

This is no longer speculation. It's a two-term scaling law — benefit, then harm — measured experimentally. And it changes the game for anyone training, fine-tuning, or deploying models today.


Key takeaways

  • 27.5% of high-quality web tokens (June 2026) are AI — 31.1% in August 2026, versus 10.1% in June 2024 (source: arXiv 2609.40295).
  • Inverted U-shaped scaling law: AI tokens help "data-starved" models, then degrade performance beyond a satiation point.
  • Compute cost: +60% to train on today's unfiltered web, +200% projected for 2028.
  • Complicit quality filters: FineWeb and DCLM favor AI text instead of filtering it.
  • Disclosed conflict of interest: 4 of the 7 authors are from Pangram Labs, the vendor of the detector used for labeling.

Tool Main use Price (October 2026) Ideal for
Pangram Detector Large-scale AI detection Quote-based Data curation teams, LLM labs
FineWeb Quality-filtered web corpus Free (open) Pre-training, research
DCLM Baseline Data filtering pipeline Free (open) Building custom datasets
Hostinger AI-ready website hosting €2.99/month (Oct 2026, check hostinger.fr) Deploying lightweight LLM apps

The measurement: how we get to 27.5%

A sample of 310,000 documents over 5 years

The study covers 310,000 documents sampled from January 2021 to August 2026, run through the Pangram 4 classifier (AUROC 0.9916, false positives 0.0041%, false negatives 0.3396%). The source corpus: the FineWeb crawls, already filtered for quality — which makes the figure all the more worrying. This isn't the web's sewers. It's the "good" web.

The explosion by format and topic

Contamination is not uniform. According to the Pangram 4 technical report (arXiv 2607.27183), more than a third of new internet content would be AI. Long social media posts: 26%. News articles: 9%. ML conference reviews: 21%. Some technical topics (programming, SEO, marketing) saturate faster than others.

Why quality filters let it through

This is the most counterintuitive point. FineWeb and DCLM favor AI text. Their heuristics (length, coherence, absence of noise, clean HTML structure) preferentially select LLM output — which is by construction fluent, structured, and error-free. Raw human text, by contrast, is chaotic. "Quality" filters become magnets for synthetic content.


The experiment: 800 models, a two-term scaling law

The protocol: variable AI ratio, constant compute

The authors trained 800 models (19.9M to 973M parameters) by injecting "wild" text (WildAI, 83 billion tokens labeled by topic, format, and source) in controlled proportions. Human token budget fixed, identical compute. The only variable: the percentage of AI tokens.

Result: marginal benefit, then collapse

For models starved of human data (the "data-starved" regime), adding 10-20% AI tokens slightly improves perplexity. Synthetic text adds lexical diversity and fills coverage gaps. But as soon as you go past the satiation point — which varies depending on model size and the quality of available human text — the curve reverses. Every additional AI token degrades performance.

The two-term scaling law

The researchers formalize it: Performance = f(human_tokens) + g(AI_tokens, human_tokens). The second term is positive, then negative. This is not a simple power law. It's a non-monotonic function. Repeating human text (data augmentation) systematically beats adding AI text, even high-quality AI text.

Compute cost: 1.6x today, 3x in 2028

At the August 2026 contamination rate (31.1%), matching the performance of a model trained on purely human data requires 1.6x the compute (at the standard ~20 tokens/parameter ratio). Extrapolating the observed exponential growth (0.1% → 10.1% → 31.1% over 5 years), the factor rises to 3.0x by 2028. That means tripling the GPU budget for the same result.


The Conflict of Interest: Pangram as Judge and Jury

4 out of 7 authors affiliated with Pangram Labs

Four of the seven signatories of the preprint (arXiv 2609.40295) are employees of Pangram Labs, the company that sells the detector used to produce the study's "AI vs human" labels. The Pangram 4 detector (arXiv 2607.27183) shows impressive metrics (AUROC 0.9916), but the vendor has a direct interest in the problem appearing massive and measurable.

The detector as a product

Pangram sells the detection API. The more the "AI on the web problem" is quantified, the more valuable the solution (their classifier) becomes. The study publishes the WildAI corpus (83B tokens) and the code — which is commendable — but the entire measurement chain depends on a commercial black box. No independent third-party detector has validated the 27.5% / 31.1% figures.

False positives / false negatives: the impact on the numbers

With a false negative rate of 0.3396%, the detector slightly underestimates the AI share. But the false positive rate (0.0041%), though minuscule, applied to billions of human tokens, can create noise. Most importantly: the calibration of the decision threshold (binary AI/human) is not discussed in the preprint. A more conservative threshold would cause the percentage to drop.

Key takeaways

The figures are probably in the right order of magnitude — the trajectory 0.1% → 10% → 30% over 5 years is consistent with massive LLM adoption. But the decimal-point precision (27.5% vs 28.1%) should be taken with a grain of salt. The authors' transparency about their affiliation is present in the paper, but absent from most media coverage.


Practical Consequences for Training and Deployment

Data Curation Becomes Critical

If quality filters (FineWeb, DCLM) select AI text, the answer is not "more quality filtering". It's provenance filtering + AI detection + adaptive thresholds. Pre-training pipelines must integrate a detection step before quality scoring. Additional cost: not negligible, but lower than the 1.6x compute.

"Human Data" Becomes a Strategic Asset

Players who own human corpora that are verified, dated, pre-2023 (old Common Crawl, Wikipedia snapshots, books, pre-Copilot open source code) hold a resource that appreciates in value. The scarcity of clean human data becomes the #1 bottleneck, ahead of compute.

Fine-tuning and RAG: Same Problem

This isn't just a pre-training story. Fine-tuning on synthetic datasets (generated by GPT-5.5, Claude Opus 4.7, Gemini 3 Pro) reproduces the same phenomenon: initial benefit, then saturation. For RAG, indexing web content "freshly" generated by AI injects noise into the knowledge base. See our guide on how to train your AI avatar with your own data for safeguards.

Search and SEO: The Loop Closes

Google Search is now 100% generated by AI (AIO, AI Overviews). Synthetic responses are fed back into the web, recrawled, re-detected as AI, and the cycle continues. The article Google Search is now 100% generated by AI: the end of "ten blue links" and the earthquake for the entire web details this shift. For publishers, a website without AI is already obsolete in 2026 — but AI must be mastered, not merely endured.


Ecosystem responses: detectors, watermarking, payment

Detectors: an arms race

Pangram, GPTZero, Originality.ai, Turnitin — the detection market is exploding. But no detector is robust against adversarial attacks (paraphrasing, high temperature, fine-tuning on human style). Detection remains a probabilistic signal, not a binary truth.

Watermarking: an unfulfilled promise

Google (SynthID), OpenAI, and Meta have announced statistical watermarking. As of October 2026, no mass deployment is in effect on consumer-facing models. Open-weights models (Llama, Mistral, DeepSeek, Qwen) don't have it. Watermarking only covers proprietary APIs — a minority of the generated volume.

Cloudflare and the monetization of scraping

Cloudflare now offers to charge AI for scraping the web. Bots have become the majority of internet traffic. This doesn't solve the contamination of already-crawled data, but it changes the economics of future collection.

NeurIPS 2026: 28 submissions rejected for AI generation

The NeurIPS 2026 conference rejected 28 submissions detected as AI-generated. Academia is starting to crack down — but the detectors used (often Pangram or GPTZero) inherit the same limitations.


❌ Common Mistakes

Mistake 1: Believing that "more data = better" without distinguishing provenance

Injecting raw Common Crawl into your training pipeline today degrades the model. The AI/human ratio has flipped. Solution: audit your corpus with a detector (even an imperfect one), set a maximum threshold (e.g., <5% AI tokens), and supplement with verified human sources.

Mistake 2: Trusting standard quality filters (FineWeb, DCLM) to do the cleaning

These filters amplify the AI share. They select for fluency, and AI produces fluency at industrial scale. Solution: add an AI detection step before quality scoring, or use "frozen" pre-2023 corpora as an anchor.

Mistake 3: Ignoring the Pangram conflict of interest in your monitoring

Citing "27.5% of the web is AI" without mentioning that the figure comes from a commercial detector sold by 4 of the 7 authors is press-release regurgitation, not journalism. Solution: always cross-check with an independent estimate, or present the range along with the methodological caveat.

Mistake 4: Thinking watermarking will solve the problem

There is no universal watermarking deployed as of October 2026. Open-weights models (which generate a massive share of the volume) don't have it. Solution: don't wait for a magic technical fix—build your own data traceability chain.


❓ Frequently Asked Questions

Does the 27.5% figure apply to the entire internet or only "quality" web content?

Only the FineWeb crawls (quality-filtered). The raw web (unfiltered Common Crawl) probably has a higher rate, but the study doesn't quantify it. The 27.5% is therefore a lower bound on the web usable for training.

Why does repeating human text beat adding AI text?

Human text, even repeated, retains the natural distribution of language (idiosyncrasies, consistent errors, personal style). AI text, even high-quality, introduces a synthetic distribution that biases the model toward the "LLM average" — loss of diversity, latent mode collapse.

Is the satiation point the same for all models?

No. It depends on: (1) model size (larger = later satiation), (2) the amount of human data available (less human data = earlier satiation), (3) the relative quality of AI vs. human text. The study maps this surface across 800 configurations.

Should we stop using web crawls for training?

Not necessarily. But it must be actively curated. Labs that continue using raw Common Crawl without AI detection are already paying the 1.6x compute tax. By 2028, it will be 3x. Curation becomes profitable well before that.

Are current models (GPT-5.5, Gemini 3 Pro, Claude Opus 4.7) affected?

Yes. They were trained on data cut off before the 2024-2026 explosion (cutoff generally late 2023 / early 2024). The next training cycles (GPT-6, Gemini 4, Claude 5) will have to tackle this problem head-on. That's why the preprint matters now.


✅ Conclusion

The web has tipped. Nearly a third of "quality" tokens are now synthetic, and the filters we use to clean the data worsen the problem. The scaling law measured across 800 models is unequivocal: beyond a certain threshold, every AI token added costs performance. The compute needed to compensate doubles every 18 months.

The exact figure (27.5% vs 31.1%) deserves caution — the detector is commercial, and the authors are judge and jury in their own case. But the trajectory is undeniable. Whoever lacks a strategy for curating human-verified data for their next training runs will pay the exponential tax. Scarcity is no longer compute. It's humans.