Research

Google Open-Sources RRSI, Agents That Rewrite Their Own Harness

Google Cloud AI Research open-sourced RRSI, a framework where frozen-weight agents rewrite their own harness under regularization — improving held-out benchmarks while cutting tokens 30%.

Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without OverfittingAI-generated
By Marcus Bennett5 min read

Updated

Why it matters

  • RRSI, released with UNC-Chapel Hill, Stanford and WashU under Apache 2.0, lets frozen-weight LLM agents rewrite prompts, tools, memory, control flow and sub-agents.
  • All six held-out benchmark splits improved; OOD average rose from baseline 39.7 to 43.6, the only method more than a point above baseline among five compared.
  • RRSI used 2.42M policy tokens per trial versus 3.80M for unregularized evolution — a 30% reduction per the abstract, 36% per the project page.

Google Cloud AI Research has open-sourced RRSI (Regularized Recursive Self-Improvement), a framework that lets an LLM agent rewrite its own harness — prompts, tools, memory, control flow and sub-agents — while the model weights stay frozen. The release, built with collaborators at UNC-Chapel Hill, Stanford and Washington University in St. Louis, tackles the central failure mode of self-improving agents: gains that evaporate the moment the agent meets a benchmark it did not optimize against.

The code is available under Apache 2.0 on GitHub. It requires Python 3.10+, accepts any LiteLLM model string, and defaults to Claude Opus 4.8 on Vertex AI. The research paper is on arXiv (2609.24972), and the project page is live at regularized-rsi.com.

The stakes are straightforward. Harness evolution — loops that propose edits, score them on a fixed set of tasks and keep the winner — has become a popular way to squeeze more capability out of frozen models without retraining. But the same tasks are reused every round, so the loop can simply memorize them. The RRSI paper names three failure modes: benchmark-specific fitting, noise chasing and complexity accumulation. Each one widens the gap between evolve-set scores and real transfer. If agents that "improve themselves" only improve on their own homework, the technique is a benchmark trick, not an engineering advance.

How RRSI works

RRSI keeps every harness component editable. What it regularizes is the search itself, on both the proposal side and the selection side.

On the proposal side, three mechanisms constrain how edits are generated. An annealed edit budget follows a cosine schedule: early rounds can bundle several coordinated edits together to find a mechanism, while late rounds allow a single, attributable change. Evidence-aware credit logs every candidate with its component, hypothesis, diff, score change and cost change, so the proposer reads the full ledger and falsified ideas are not retried. Structured exploration shifts budget to untouched components when progress stalls inside the noise band.

On the selection side, a leakage critic rejects edits containing task names, entities, answers or benchmark-specific logic before any scoring happens. A noise-adjusted floor requires gains to clear the variance measured on the unchanged base harness. A cost rule demands that extra inference tokens be paid for by measured gain. And pruning turns components that stop producing gains into deletion targets.

The research team frames these mechanisms as analogies to classic regularizers from statistics: the edit budget maps to L0, pruning to Lasso (L1), and the cost rule to Ridge (L2).

Results across 8 benchmarks

The numbers span agentic coding, engineering design and professional-domain evaluation. On Terminal-Bench 2.1, the evolve split score rose from 74.2% to 80.2%. On SWE-bench Verified — never used for selection — the score rose from 82.0% to 83.8%.

Out-of-distribution results are the framework's core claim. JobBench improved by 4.7 points, GDPval by 3.5 points and APEX-Agents by 3.7 points. On the evolve splits, EngDesign gained 4.9 points and Frontier-Eng gained 4.3 Medal points. Harvey LAB, from the legal AI company Harvey, improved by 1.1 points on the evolve split and 2.3 points on its held-out split. All six held-out splits improved.

The gains are not tied to a single policy model. With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7, and SWE-bench Verified rose from 76.8 to 79.0.

The regularized harness is also cheaper to run. On the agentic workspace instance, RRSI used 2.42 million policy tokens per trial, against 3.80 million for unregularized evolution. The paper's abstract reports this as a 30% reduction; the project page says 36%.

RRSI versus the closest competitors

The paper's Table 1 compares RRSI against four related methods — Meta-Harness, AHE, TTHE and HarnessX — all sharing the same starting harness, policy, evolve split and candidate budget. All five freeze model weights. Per the RRSI research team, none of the competitors implements a cost rule or pruning.

The trade-off shows up clearly in the numbers. On the Harvey LAB evolve split, Meta-Harness leads with 93.0, ahead of HarnessX at 91.8, TTHE at 91.1, AHE at 90.7 and RRSI at 90.5. RRSI posts the smallest evolve gain of the group.

But on the out-of-distribution average — the mean of JobBench, GDPval and APEX-Agents, where the baseline harness H0 scores 39.7 — RRSI leads with 43.6. Meta-Harness reaches 40.6, HarnessX stays at 39.7, AHE drops to 39.2 and TTHE falls to 38.0. RRSI is the only method whose OOD average sits more than a point above the baseline.

That contrast is the paper's argument in miniature: unregularized evolution buys bigger scores on the tasks it trains on and gives back most of the value everywhere else. Regularization trades a smaller evolve-set gain for transfer that actually holds.

Why it matters

The context matters for anyone building or buying agentic systems. Self-improvement loops that overfit their evolve sets produce leaderboards that look good and deployments that disappoint. RRSI's contribution is a set of cheap, explicit acceptance gates — leakage screening, noise floors, cost accounting, pruning — that any team running harness search could adopt independently of the rest of the framework. The arithmetic is favorable: fewer tokens per trial plus better OOD transfer means the regularization pays for itself rather than costing margin.

The open questions are equally concrete. The delta on SWE-bench Verified is 1.8 points on an already-strong 82.0% baseline, so the ceiling on saturated benchmarks may be low. And the competitor comparison comes from the RRSI team's own paper, under an evaluation setup they chose.

The framework is deployable now as research infrastructure: Apache 2.0 code, LiteLLM compatibility and Vertex AI defaults mean a competent team can point it at its own benchmarks this week. Whether the regularization holds on private, messier task distributions than the eight public benchmarks in the paper is the test that will decide if regularized self-improvement becomes standard practice in agent engineering or stays a well-gated research result.

Original: regularized-rsi.com

Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

Senior reporter covering consumer brands and retail at AI In Context.

124 articles

Related articles

  1. Google Releases DiffusionGemma, a 26B Model That Generates Text Four Times Faster
  2. Agentic AI Is Driving a CPU Comeback — and a Shortage
  3. Google's Decoupled DiLoCo Trains LLMs Across Data Centers 20x Faster
  4. OpenAI Launches Model Distillation Suite in Its API
  5. Google DeepMind Joins DOE's Genesis Mission to Bring AI to 17 National Labs

« Previous article