ScholarCatalyst: the CMU/Stanford benchmark that measures an AI's ability to find the papers that inspire real research
🔎 Your research agents find relevant papers. Not necessarily the right ones.
AI research agents have proliferated in 2026. They read, synthesize, write literature reviews, propose hypotheses. But ask the uncomfortable question: how do we know they're finding the right papers — the ones that actually move a project forward? Until now, the answer rested on three outdated proxies: relevance, QA, citations.
On October 2, 2026, a team from Stanford, CMU, UW, SNU, AI2, and MIT — signed, among others, by Chelsea Finn, Omar Khattab, Yejin Choi, Pang Wei Koh, and Graham Neubig — published ScholarCatalyst (arXiv 2610.02202, 57 pages). The first benchmark that measures "scientific inspiration" as judged by the authors themselves. Not citations. Not relevance. Inspiration.
The finding is damning: agentic search performs no better than embedding-based retrieval (0.42 vs 0.48 Recall@20). The missing link in research agents is neither the LLM nor the harness. It's the evaluation signal.
Key Takeaways
- 184 lead authors of 207 recent computer science papers self-labeled the prior publications that advanced — or could have advanced — their work, with detailed justifications.
- 894 queries (207 CoreQ + 687 SubQ) over a corpus of 191,000 papers, temporally restricted to rule out any contamination.
- Agentic search loses to embeddings: 0.42 vs 0.48 Recall@20, even though it uses the same retriever as a tool.
- An agent based on Claude Fable 5.1, possibly exposed to the papers during training, plateaus at 0.51 R@20: it still misses nearly one in two inspiration papers.
- Scientific retrievers (SPECTER2, OpenScholar) trail general-purpose dense retrievers by 16 to 27 points, with the latter placing 39% of inspiration papers in the top 20 (CoreQ) and 51% (SubQ).
- The first benchmark built on author-judged inspiration, with pre-discovery queries and a temporally restricted corpus.
- Automated pipeline: a single arXiv ID is all it takes to generate new data that can be annotated by the authors. The benchmark grows with the literature.
Recommended Tools
| Tool | Main use | Price (October 2026) | Ideal for |
|---|---|---|---|
| ScholarCatalyst | Evaluate a retriever or an agent on scientific inspiration | Free (open benchmark) | Teams building research agents |
| OpenScholar (AI2) | Scientific QA with sources | Free (open source) | Document baseline — bearing in mind it loses 16 to 27 points on inspiration |
| SPECTER2 (AI2) | Scientific paper embeddings | Free (open source) | Semantic search over an internal corpus |
| Hostinger | Host a monitoring portal or an AI analysis blog | From ~€3/month (October 2026, check hostinger.com) | Publish your monitoring content and evaluation results |
The SPECTER2 and OpenScholar figures come from the evaluation published in ScholarCatalyst. If you take away only one thing from this table: test your embeddings first, the rest afterwards.
What is ScholarCatalyst, and why citations were no longer enough
ScholarCatalyst is a scientific retrieval benchmark whose ground truth comes from the authors of the papers themselves — not from citations, not from external annotators.
The protocol fits in one sentence: ask 184 lead authors of 207 recent computer science papers which earlier papers advanced their project — or could have — and why. Each labeling comes with a detailed justification. The question asked is not "is this paper relevant?" but "without it, would your project have been the same?".
From these annotations come 894 queries:
- 207 CoreQ: the central question of the research project;
- 687 SubQ: the decomposed sub-questions, which broaden the evaluation surface.
The corpus contains 191,000 papers, and this is the detail that changes everything: it is temporally restricted. For each query, the corpus ends before the start of the source project — and before the knowledge cutoff of models predating the source papers. A model cannot "remember" the answer: it has to find it.
Why inspiration rather than relevance? Because the promise being sold for the past two years is the "AI co-scientist": a machine that finds the missing idea. That promise had never been tested on a realistic task — finding, before the discovery, the work that makes it possible.
The problem with citations
Counting citations means measuring a paper's visibility, not its real influence. A citation can be a courtesy, a methodological obligation, or self-promotion. And above all: it takes years for one to exist.
ScholarCatalyst reverses the logic. The benchmark doesn't ask AIs to retrieve what was cited, but to retrieve what should have been found. It's a counterfactual task — and that's precisely what makes it close to real research work.
A benchmark that grows with the literature
One last technical point, and not the least: the pipeline is automated. An arXiv ID is all it takes to produce data that can be annotated by the authors. ScholarCatalyst is therefore not a frozen snapshot: it can grow at the pace of the literature, with queries that are always pre-discovery and a corpus that is always restricted.
The numbers: agentic search loses to embeddings
Agentic search doesn't outperform embedding-based retrieval: 0.42 vs 0.48 Recall@20. And the agent was using the very same retriever as a tool.
In other words, all the machinery — query decomposition, iterations, multi-step reasoning — doesn't add a single piece of information the retriever didn't already have. Worse: it degrades the result.
| Approach | Recall@20 | Key takeaway |
|---|---|---|
| Embedding-based retrieval (general-purpose dense) | 0.48 | The baseline to beat |
| Agentic search (same retriever as a tool) | 0.42 | Orchestration degrades |
| Agent based on Claude Fable 5.1 | 0.51 | Even with possible contamination, the gap remains narrow |
An agent that chains thirty LLM calls to lose six points of Recall isn't an innovation. It's a bill.
Claude Fable 5.1 at 0.51: the contamination question
The benchmark's best-performing agent relies on Claude Fable 5.1 — a model that may have seen the papers during training, the authors note. Even in this favorable case, it reaches 0.51 R@20: it still misses nearly one in two inspiration papers in the top 20.
This is the number to throw at anyone selling you an "autonomous research agent". On core questions (CoreQ), general-purpose dense retrievers place only 39% of inspiration papers in the top 20 — 51% on sub-questions (SubQ). Six key papers out of ten are missing.
This isn't an isolated case
The result converges with SciNet (arXiv 2601.03260), a relation-aware benchmark published in January 2026: 269 million papers, 8,940 tasks, and agent precision often below 20% on relational tasks. Two independent benchmarks, two methodologies, same diagnosis.
My reading: the bottleneck is informational, not procedural. An agent can reason all it wants — if it doesn't have the right document in its context, it's reasoning in a vacuum.
Why "scientific" retrievers lose 16 to 27 points
Science-specialized retrievers lose 16 to 27 points to generalist dense retrievers on inspiration queries.
The generalists place 39% of inspiration papers in the top 20 (CoreQ) and 51% (SubQ), while SPECTER2 and OpenScholar — designed precisely for the scientific literature — do markedly worse.
The most plausible hypothesis: these retrievers are trained on citation signals. They therefore learn citation biases — popularity, recency, community effects — and not inspiration. Yet a query like "what would have unlocked this project?" has no statistical citation signature. The specialist is trained on the wrong target.
The lesson goes beyond the scientific case: domain specialization is not an evaluation strategy. On inspiration queries, the quality of the training signal matters more than the domain vocabulary.
The Missing Link in Research Agents
ScholarCatalyst fills the void between "finding relevant papers" and "finding papers useful for a discovery". It's exactly the hole nobody was measuring.
Existing benchmarks — SciFact, DORIS-MAE, ScholarQABench, LitSearch, MIR — evaluate claim verification, literature QA, or retrieval relevance. None, the authors point out in the paper, relies on inspiration judged by the authors themselves, with pre-discovery queries and a temporally restricted corpus. We were testing reading, not impact.
ScholarCatalyst also fits into a 2026 wave of benchmarks that measure real dynamics rather than superficial scores: OmniGameArena evaluates the learning dynamics of VLM agents in games running on Unreal Engine 5, FutureSim has agents replay three months of real events, and Homebody shows a Stanford humanoid piloted without a learned policy — 7/100 on the desks benchmark, and that raw figure is precisely part of the value. ScholarCatalyst applies the same standard to science: judge real impact, not claimed performance.
Why is this the missing link? Because without this signal, teams optimize what is measurable — relevance, latency, cost — not what matters. You end up with agents that find 50 moderately relevant papers instead of the 3 that would have changed everything.
Three lessons if you're building a research agent
The paper takes an evening to read and changes three design decisions. Here are the lessons, in the order I would prioritize them.
1. Benchmark the retriever alone, before the harness
0.42 versus 0.48: if your agentic pipeline doesn't improve the Recall@20 of your base retriever, it adds nothing — it costs. Measure the embeddings alone first, then only add orchestration if it beats that floor. It's the same finding as in our analysis on the harness layer, now infrastructure in its own right: orchestration is a product, not a retrieval fix.
2. Track contamination systematically
An agent that could have ingested the corpus during pre-training isn't being evaluated: it's remembering. Restrict your evaluation sets temporally, as ScholarCatalyst does with its 191,000 papers cut off before the start of each project. Any metric obtained on a corpus posterior to the model's cutoff should be treated as suspect.
3. Have usefulness judged by humans who know
Relevance can be measured automatically; usefulness cannot. The 184 authors of ScholarCatalyst justified every labeling — that's what gives the benchmark its value. In your context, the equivalent: have your results judged by the people who actually run the projects. Agent memory, meanwhile, is becoming a product in its own right — see Hindsight — and it helps accumulate context across sessions. But no memory creates the signal: it takes human judgment on what was actually useful.
Note in passing: whatever reasoning engine you choose — Gemini 3.1 Pro, GPT-5.5, Claude Opus 4.7 — the bottleneck identified here sits upstream of the LLM. Switching models doesn't fix a retriever that misses.
The limitations you need to know about
ScholarCatalyst is the most rigorous benchmark of its kind, but three limitations deserve to be laid out plainly.
The judgment remains human and retrospective. 184 authors with detailed justifications is solid — but inspiration is judged after the fact. Pre-discovery queries and the restricted corpus limit information leakage, not the subjectivity of the judgment.
Computer science only, for now. The 207 source papers come from computer science. Generalization to biology, medicine, or physics remains to be demonstrated. The automated pipeline (an arXiv ID is all it takes) is the authors' answer — to be confirmed in practice.
The counterfactual may dilute the signal. The 687 SubQ concern papers that "could have" advanced the project. That's useful — research also advances through missed papers — but it's less clear-cut than the "did advance" of the CoreQ.
None of these limitations invalidate the benchmark. They define its scope: computer science, inspiration, author judgment.
❌ Common Mistakes
Mistake 1: Treating citations as a measure of impact
A citation measures visibility, not influence: courtesy citations, self-citations, methodological obligations. The solution: in your internal evaluations, have end users judge real usefulness — or rely on a benchmark based on author judgments, like ScholarCatalyst.
Mistake 2: Stacking agentic on top of a mediocre retriever
0.42 vs 0.48: iterative reasoning can't make up for a retriever that misses the right papers. The solution: benchmark the retriever on its own (Recall@20), and only add orchestration if it beats that score. Otherwise, you haven't built an agent — you've built a cost.
Mistake 3: Evaluating on a corpus the model may have seen
A flattering score on a corpus published after the knowledge cutoff proves nothing: the model may have memorized it. The solution: restrict your evaluation sets temporally, to before the cutoff of each model tested — the method behind ScholarCatalyst's 191,000 papers.
❓ Frequently Asked Questions
What is ScholarCatalyst?
A scientific retrieval benchmark published on October 2, 2026 (arXiv 2610.02202, 57 pages) by teams from Stanford, CMU, UW, SNU, AI2, and MIT. 184 authors of 207 papers label the prior publications that have advanced or could have advanced their research: 894 queries, a temporally restricted corpus of 191,000 papers.
Why does agentic search lose to embeddings?
Because the agent uses the same retriever as its tool: multi-step reasoning adds no information the retriever doesn't already have. The result: 0.42 vs. 0.48 Recall@20. Even the agent based on Claude Fable 5.1, advantaged by possible exposure during training, caps out at 0.51.
Which models and retrievers were tested?
The paper compares embedding-based retrieval (0.48 R@20), agentic search with the same retriever as a tool (0.42), and an agent based on Claude Fable 5.1 (0.51), as well as the scientific retrievers SPECTER2 and OpenScholar, which lose 16 to 27 points against generalist dense retrievers. The full details fit within the paper's 57 pages.
How does it differ from SciNet?
SciNet (arXiv 2601.03260) evaluates agents on 269 million papers and 8,940 relational tasks, with accuracy often below 20%. ScholarCatalyst measures something else: inspiration as judged by the authors, across 894 queries and 191,000 papers. The two complement each other more than they replace each other.
How can I use ScholarCatalyst for my project?
The pipeline is automated: a single arXiv ID is enough to produce data that can be annotated by the authors. You can evaluate your retriever or your agent on the 894 existing queries, or contribute new tasks as the literature advances. Starting point: the paper.
✅ Conclusion
ScholarCatalyst doesn't claim that research agents are bad: it proves, with the numbers to back it up, that they've been evaluated on the wrong question all along — and that on the right one, even the best systems miss one inspiration paper out of every two. Read the full paper, then benchmark your retriever. And if you document your evaluations online, Hostinger runs your research monitoring portal for a few euros a month (October 2026, check hostinger.com).