Sakana AI's MLR Review System Detects 73.43% of Core-Claim Errors in Benchmark
Sakana AI's Beyond Imitation paper introduces a contradiction benchmark and a multi-layered review system that caught 73.43% of core-claim errors with four reviews, versus 14.81% for the best baseline, but exact matches on retracted papers were 16.11%.
The Contradiction Benchmark inserts contradictions into real papers and checks whether reviewers catch them. The research team collected 257 CC-licensed papers from ACL, AISTATS, CVPR and ICML 2025, plus NeurIPS 2024. Gemini 2.5 Pro builds a knowledge graph of each paper's claims, evidence and methods. The distance of a node from a main claim sets severity: distance 0 hits a core claim, while larger distances hit details. GPT-4.1 then rewrites one node per distance into a contradiction, producing 1,164 data points. An o3 judge scores each review 10 times. On clean papers, the judge reached 99.9% accuracy. It showed 86.8% sensitivity on manually confirmed catches, so the reported scores may be conservative.
MLR is an agentic review system that reads a paper before critiquing it. It uses three agents on off-the-shelf Claude models. The Appendix Agent, based on Claude Haiku 3.5, summarizes experiments and implementation details from the appendix. The optional Literature Review Agent, based on Claude Sonnet 4, uses web search to place the paper in prior work. The Review Agent, also based on Claude Sonnet 4, runs a three-pass prompt chain inspired by Keshav's Three-Pass Approach. Pass 1 writes a high-level outline. Pass 2 reads in detail and flags weaknesses, assumptions and gaps. Pass 3 merges all agent outputs into Strengths, Weaknesses, Questions, Recommendation, Score and a To-Do list. The PDF is passed directly, so figures and equations survive. MLR reads up to 10 pages of main text.
MLR led every baseline on the Contradiction Benchmark. With four reviews, it caught 73.43% of distance-0 contradictions and 40.95% of contradictions overall. The best baseline, AgentReview, caught 14.81% at distance 0. A single MLR review still caught 60.79%. An ablation separates model from design. Swapping GPT-4.1 for Claude Sonnet 4 inside LLM-Review lifted distance-0 detection from 14.56% to 35.40%. MLR's design added about 25 more points on a single review. Accuracy falls as node distance grows, which supports the severity scoring. On real retracted papers from WithdrarXiv-Check, a set of 211 papers, gains shrink. MLR scored 26.07% on similar matches and 16.11% on exact matches. The strongest baselines scored 18.48% and 9.00%.
On agreement with human reviewers, scores mostly align. On ICLR 2025 submissions, MLR's predicted scores reached a Pearson correlation of 0.586 with human scores. The human-to-human reference was 0.742. On ICML 2025, the AI Reviewer edged it, 0.439 versus 0.429. On focus, the systems differ. MLR stresses validity and experiments, while humans weigh clarity and novelty more. The authors frame this as a complementary perspective, not a replacement. The paper also reports that MLR still falls for hidden prompt injection.
MLR costs about $0.47 per review, excluding the optional literature agent. It uses 189,062 input tokens, about half of the AI Reviewer's 403,654. A single-prompt variant cut cost by about two-thirds. Its detection dropped about 3.5 points on a subset of the benchmark.
In the paper's comparison, MLR used Claude Sonnet 4 and Claude Haiku 3.5, LLM-Review used GPT-4.1, AI Reviewer used o4-mini, and AgentReview used GPT-4o. MLR's design is three agents with a three-pass chain; LLM-Review uses a single prompt with text truncated; AI Reviewer uses a five-review ensemble, meta-review and reflection; AgentReview uses reviewer, author and area chair roles. On the full Contradiction Benchmark, MLR scored 40.95% with four reviews, against 6.39% for LLM-Review, 6.50% for AI Reviewer and 5.95% for AgentReview. On core-claim errors at distance 0, MLR scored 73.43%, against 14.56%, 11.17% and 14.81%. On WithdrarXiv-Check, MLR scored 26.07% on similar and 16.11% on exact matches, against 5.21% and 2.37% for LLM-Review, 13.74% and 9.00% for AI Reviewer, and 18.48% and 5.69% for AgentReview. On ICLR 2025 Pearson correlation with human scores, MLR scored 0.586, LLM-Review -0.013, AI Reviewer 0.538 and AgentReview 0.195. Input tokens per review were 189,062 for MLR, 6,517 for LLM-Review, 403,654 for AI Reviewer and 310,964 for AgentReview. Costs per review were about $0.47, $0.01, $0.49 and $0.81. MLR's code is available on request; the other systems have open code.
Editor's Summary
Sakana AI's Beyond Imitation paper presents a Contradiction Benchmark and Multi-Layered Review system that detected 73.43% of core-claim errors with four reviews, far exceeding the best baseline, but exact matches on retracted papers remained 16.11%. The system runs on off-the-shelf API models for about $0.47 per review, aligns moderately with human scores and still fails on hidden prompt injection. The authors present it as complementary to human review rather than a replacement.