Google DeepMind’s Dream-RSI lets AI agents improve by replaying past searches

Google DeepMind’s Dream-RSI lets AI agents reuse past search runs to tune exploration, reducing costly attempts without changing the underlying model.

AI News

Google DeepMind researchers have developed Dream-RSI, a method that allows AI agents to improve how they search for solutions by replaying earlier attempts instead of rerunning every experiment. The approach targets a central cost in autonomous problem-solving: deciding which possibilities to explore, which to abandon, and how much compute to spend on each path.

According to testing reported by The Decoder, Dream-RSI improved or matched existing results across programming, mathematical optimization, and GPU-kernel tasks. In some experiments, it reduced the number of attempts by more than half while leaving the underlying Gemini model unchanged. The work matters because it shifts self-improvement away from retraining a model and toward optimizing the process that directs its search.

How Dream-RSI reuses past attempts

AI agents working on difficult problems often follow an iterative loop. They generate a candidate solution, evaluate it, and use the result to decide what to try next. As the number of possible paths grows, exploration can consume large amounts of compute, particularly when each attempt requires code generation, execution, or another expensive evaluation.

A fixed search strategy may repeatedly pursue unproductive directions. An adaptive strategy can respond to results, but learning which strategy works may itself require many live trials. Dream-RSI addresses that trade-off by recording the agent’s previous attempts and their outcomes in a searchable history.

The method then tests alternative decisions against those stored results. Rather than generating and evaluating every candidate again, the system can simulate what might have happened if it had selected a different branch, stopped exploring a dead end earlier, or allocated more effort to a promising path. The researchers describe this retrospective process as “dreaming.”

The resulting strategy is used in a subsequent live search. After that run, the new search history can be analyzed again, creating a cycle in which the agent’s exploration policy improves over time. Dream-RSI changes the search strategy, not the weights or capabilities of the model producing the candidate solutions.

What the reported tests show

The Decoder reports that the researchers evaluated Dream-RSI with Gemini 3.1 Pro and Gemini 3.7 Flash on eight tasks spanning three areas. The comparisons used the same starting conditions but contrasted Dream-RSI with a baseline using a fixed search strategy.

One task involved writing a fast program for a statistical calculation used in genomics and finance. In the reported tests, Dream-RSI produced programs that ran faster than the established sklearn and glmnet libraries across six datasets. With Gemini 3.1 Pro, average runtime fell from 3,587 milliseconds to 2,931 milliseconds, while the number of attempts declined from 550 to 317.

The article also reports a comparison with SimpleTES, which required 51,200 runs on that task, versus 317 attempts for Dream-RSI. Additional tests involving mathematical optimization and GPU kernels showed either comparable or better results at lower search cost. On two GPU tasks, Dream-RSI matched performance while reducing the number of runs by as much as a factor of 2.43. On two others, it achieved up to 2.09 times the performance within the same compute budget.

These figures are research results reported through The Decoder, not independently verified production benchmarks. The available evidence does not establish how Dream-RSI performs across broader workloads, different models, or changing evaluation environments. It also does not show whether the method produces consistently better final solutions when the recorded search history is small or unrepresentative.

A follow-up analysis exposed another limitation. The researchers tested whether search histories could be condensed into explicit instructions telling the agent where to look. On one GPU task, that instruction-based version performed worse than the replay-based system. The researchers suggested that overly specific guidance can narrow exploration and prevent the agent from finding less obvious alternatives.

Why the distinction matters for builders

For teams building AI agents, Dream-RSI points to a potentially practical way to reduce inference and evaluation costs without immediately fine-tuning or retraining a foundation model. A coding agent, optimization system, or scientific discovery tool could preserve detailed traces of earlier work and use them to improve allocation of future search effort.

That could be useful in workflows where candidate solutions are expensive to test. GPU-kernel generation is one example: compiling and benchmarking each variant can take considerably more time than selecting among previously evaluated branches. Similar economics may apply to code optimization, design search, and automated experimentation.

The approach also highlights an important systems boundary. Improvements may come from better orchestration rather than a more capable base model. Teams could therefore evaluate search policies, replay systems, and compute allocation separately from model upgrades. In principle, this makes progress easier to measure: builders can compare the number of attempts, evaluation cost, and final quality under the same model.

There are operational risks. A search history can contain misleading evaluations, failures caused by temporary conditions, or gaps in the explored space. Replaying that history may make an agent more efficient at following a narrow map rather than better at solving the underlying problem. The instruction experiment reported by the researchers reinforces the need to preserve room for exploration instead of turning successful past behavior into rigid rules.

Dream-RSI sits within a wider research direction. Google DeepMind’s AlphaEvolve uses model-generated code and evolutionary selection to search for improved programs, while Dream-RSI operates one level above by tuning how that search is conducted. Other systems, including approaches that store failures as reusable instructions, make a similar bet on accumulated experience but may constrain exploration more directly.

What to watch next

The next important signal will be whether Dream-RSI is tested beyond the reported eight tasks and Gemini configurations. Independent evaluations should examine different foundation models, noisy or shifting benchmarks, and workloads where the search space changes between runs.

Researchers and product teams will also need clearer accounting of total cost. Fewer attempts do not automatically mean lower expense if recording, replaying, storing, and evaluating search histories adds substantial overhead. Reliability measurements will matter as well: a strategy that saves compute but misses rare high-quality solutions may not suit safety-critical or research applications.

Another open question is how Dream-RSI handles transfer. A strategy learned on one class of optimization problems may not be useful for another, and a history that improves code search may be a poor guide for mathematical reasoning. Evidence that the method can generalize without over-constraining exploration would determine whether it is a reusable platform technique or mainly a task-specific optimization.

Creati.ai perspective

Dream-RSI is notable less as a claim that agents can independently redesign themselves than as an example of where near-term efficiency gains may come from. The model remains fixed; the system becomes better at deciding how to spend its search budget. For AI builders, that is a more concrete and testable path than assuming every improvement requires a larger or newly trained model.

The central caveat is that replay is only as valuable as the experience being replayed. If search histories are narrow, biased, or costly to maintain, the agent may become efficient without becoming broadly capable. The strongest next step would be transparent, independent testing that measures quality, reliability, and total cost—not just the number of attempts saved.

Ads