Google researchers propose RRSI to stop self-improving AI agents from memorizing tests

Google researchers propose RRSI to curb test memorization in self-improving AI agents, improving unseen benchmark scores while reducing runtime token use.

AI News

Google researchers have proposed a method for improving AI agents without allowing them to overfit to the tasks used to optimize them. Called Regularized Recursive Self-Improvement of Agent Harnesses, or RRSI, the approach targets a problem that becomes more serious as systems automatically rewrite their own prompts, workflows, tools, and memory logic.

According to a research paper described by The Decoder, RRSI improved results on previously unseen benchmarks by as much as 4.7 points and used roughly 30% fewer runtime tokens than an unregularized optimization approach. The reported gains come from changing the agent’s harness around a frozen model, rather than updating the model’s weights.

Why agent harnesses become the target

Many production AI agents consist of a fixed language model surrounded by an agent harness: the prompts, tool calls, workflow rules, memory systems, recovery behavior, and output handling that determine how the model operates. The harness may decide whether an agent checks a file before editing it, retries after an error, or structures its final response.

Researchers from Google Cloud AI Research and university collaborators are studying ways to automate improvements to that layer. In a typical recursive self-improvement loop, a language model proposes changes to the harness, evaluates them against a set of tasks, and uses the results to generate further changes.

The risk is that repeated optimization on a small collection of tests can make an agent better at those exact tests without making it more capable in general. The system may learn benchmark-specific patterns, select changes that succeed by chance, or accumulate unnecessary complexity that raises the measured score while increasing cost and fragility.

That issue matters because AI agents are often evaluated on limited task suites, while their intended deployment environments are much less predictable. A harness that performs well on familiar workflows may fail when file structures, instructions, tools, or user goals change.

How RRSI limits overfitting

RRSI applies controls at two points in the optimization process. When generating candidate revisions, it limits how many independent edits can be bundled into one proposal. The allowed number of edits becomes smaller over time, moving the process from broad redesigns toward more targeted modifications.

The system also records earlier attempts, helping it avoid repeatedly exploring changes that have already failed. When progress stops, it directs experimentation toward portions of the harness that have not yet been examined.

A separate critic evaluates proposed changes before they become permanent. The critic rejects revisions that appear to hardcode task names, solutions, or other benchmark-specific behavior. RRSI also requires an observable performance benefit before accepting changes that increase computational cost, and removes components that no longer contribute.

Together, those rules are intended to favor smaller, explainable improvements that transfer to new tasks rather than aggressive changes that maximize a narrow test score. The method leaves the harness editable, but places limits on how quickly and freely it can rewrite itself.

What the reported results show

The researchers tested RRSI across eight benchmarks covering coding, office-oriented agent work, and engineering design. The underlying model, identified in the report as Claude Opus 4.8, remained frozen. The comparison included an unmodified baseline harness and four other optimization methods.

The results reported by the researchers show a tradeoff. RRSI produced gains of up to 14.1 points on tasks used during optimization, but the more significant result was its performance on five unseen benchmarks, where the maximum improvement reached 4.7 points. The RRSI harness did not fall below the baseline on any of those unseen evaluations, according to the paper as reported by The Decoder.

Other approaches reportedly performed strongly on their training tasks but transferred less effectively. Two methods fell below the baseline on new tasks. RRSI delivered the smallest training-set improvement among the tested variants, which the researchers interpret as evidence that it was sacrificing benchmark specialization for broader generalization.

The token result is also relevant to teams operating agents at scale. The optimized RRSI system used approximately 30% fewer tokens than the unregularized version and required fewer steps among the optimized harnesses. The original baseline remained more economical, however, so regularization did not make the entire system cheaper than every alternative.

These are research results, not independent production benchmarks. The performance figures are claims from the researchers’ evaluation, and the source evidence does not establish how the method would behave across larger task distributions, different models, or live enterprise workloads.

Implications for AI builders and buyers

For builders, the work suggests that improving an agent should involve more than maximizing a development benchmark. Evaluation sets need to include tasks that the optimization process never sees, otherwise a self-improving system can reward itself for learning the test rather than improving its underlying workflow.

RRSI’s design also points toward operational controls for recursive self-improvement. Teams could impose limits on the number of simultaneous changes, maintain a history of failed experiments, require cost justification for more expensive workflows, and block edits that appear tied to specific test cases. Those controls could make automated harness optimization easier to audit and roll back.

The approach may be especially relevant to enterprise AI, where token costs, predictable behavior, and reliability across varied internal processes matter as much as peak benchmark scores. A harness that generalizes across unfamiliar documents or procedures may be more useful than one that achieves a higher score on a fixed collection of demonstrations.

The work does not resolve the broader safety and governance questions around agents that alter their own operating logic. RRSI constrains harness changes, but the source evidence does not show whether its critics can reliably detect every form of hidden benchmark-specific behavior. Nor does it address systems in which model weights change during optimization.

The reported transfer across models is nevertheless notable. A coding harness discovered with Gemini 3.5 Flash reportedly raised the accuracy of the weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points without modifying the latter model. The finding suggests that some workflow improvements may be portable across model capabilities, although it comes from the same research report and requires broader validation.

What to watch next

The first signal will be independent replication of RRSI across additional models and task families. Results should be compared on held-out tasks that are hidden from both the harness optimizer and its critic.

Researchers and product teams should also test whether the method remains effective when tools, prompts, memory stores, and data distributions change after deployment. Cost measurements will matter as well: the reported 30% token reduction is relative to an unregularized optimized system, not necessarily to a carefully designed baseline.

Another open question is whether similar controls can govern agents that update model weights, rather than only the surrounding harness. The current study covers frozen models, leaving that more consequential form of self-improvement outside its scope.

Creati.ai perspective

RRSI addresses a practical weakness in the current agent race: teams can automate the search for better workflows faster than they can determine whether those workflows generalize. Its strongest contribution is therefore methodological. It treats unseen-task performance and computational cost as first-class constraints instead of relying on a single optimization score.

The research does not prove that recursive self-improvement is ready for unsupervised deployment. But it gives builders a clearer design principle: an agent should earn the right to become more complex by demonstrating reliable gains beyond the tests that produced the change.

Ads