AI News Feed
Market watch
Research

Google Cloud AI Research Open-Sources RRSI to Regularize Agent Self-Improvement

MarkTechPost reports Google Cloud AI Research and university partners have released RRSI, an Apache 2.0 framework that lets an LLM agent edit its own prompts, tools, memory and control flow while model weights stay frozen. The project reports held-out gains and lower token use than unregularized evolution.

The code is released under Apache 2.0 and requires Python 3.10 or later, the report said. It accepts any LiteLLM model string, and its defaults assume Claude Opus 4.8 on Vertex AI. RRSI is described as deployable as a research framework.

The problem RRSI targets is overfitting in harness evolution loops. Such loops propose edits, score them on a fixed evolve set and keep the winner, but the same tasks are reused every round, so the loop can memorize them. The RRSI research names three failure modes: benchmark-specific fitting, noise chasing and complexity accumulation. Each widens the gap between evolve-set scores and real transfer, the report said.

RRSI keeps every harness component editable and regularizes how the search moves. On the proposal side, an annealed edit budget uses a cosine schedule so early rounds can bundle several edits while late rounds allow a single attributable change. Evidence-aware credit logs each candidate with its component, hypothesis, diff, score change and cost change, and the proposer reads this ledger so falsified ideas are not retried. Structured exploration shifts budget to components the run never touched when progress stalls inside the noise band.

On the selection side, a leakage critic rejects task names, entities, answers or benchmark-specific logic before any scoring. A noise-adjusted floor requires gains to clear the variance measured on the unchanged base harness. A cost rule requires extra inference tokens to be paid for by measured gain. Pruning makes components that stop producing gains into deletion targets. The research team frames these as analogies to classic regularizers: the edit budget maps to L0, pruning to Lasso, or L1, and the cost rule to Ridge, or L2.

Across eight benchmarks, the report said RRSI raised Terminal-Bench 2.1 on the evolve split from 74.2% to 80.2% and SWE-bench Verified, which was never used for selection, from 82.0% to 83.8%. Out of distribution, JobBench rose by 4.7 points, GDPval by 3.5 points and APEX-Agents by 3.7 points. EngDesign gained 4.9 points on the evolve split, and Frontier-Eng gained 4.3 Medal points. Harvey LAB improved by 1.1 points on the evolve split and 2.3 points on its held-out split. All six held-out splits improved.

With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7, and SWE-bench Verified rose from 76.8 to 79.0. The harness is also lighter. On the agentic workspace instance, RRSI uses 2.42 million policy tokens per trial, while unregularized evolution uses 3.80 million. The report noted a discrepancy in the savings figure: the abstract says 30% fewer tokens, while the project page says 36%.

Scores in Table 1 of the RRSI research paper compare RRSI with Meta-Harness, AHETT, HE and HarnessX under the same starting harness, policy, evolve split and candidate budget. All five freeze model weights, while only RRSI includes the cost rule and pruning, according to the RRSI research team. Meta-Harness leads the Harvey LAB evolve split at 93.0, followed by HarnessX at 91.8, HE at 91.1, AHETT at 90.7 and RRSI at 90.5. On the out-of-distribution average, with H0 at 39.7, RRSI scores 43.6, Meta-Harness 40.6, AHETT 39.2, HE 38.0 and HarnessX 39.7. That average is the mean of JobBench, GDPval and APEX-Agents, computed from Table 1. RRSI has the smallest evolve gain but the only out-of-distribution average more than 1 point above H0, the report said.