Danijar Hafner is developing world-model agents for robots that can predict outcomes and adapt to unfamiliar homes, objects, and environments.

A new startup led by former Google DeepMind researcher Danijar Hafner is developing AI agents designed to anticipate what may happen in environments they have not seen before, according to a profile published by MIT Technology Review. The company remains in stealth mode, has not publicly disclosed its name or product, and is operating from a sparsely furnished San Francisco office containing humanoid robots.
The work extends Hafner’s long-running research into agents that learn inside simulated environments before acting in the real world. Its immediate significance is not a commercial launch, but a technical direction: using internal models of reality to let robots plan ahead instead of relying primarily on trial and error after deployment.
Hafner left Google DeepMind to form the startup in fall 2025. He has revealed few operational details, including the company’s name, funding, planned customers, or launch timetable. MIT Technology Review reported that the office contains humanoid robots sourced from China, suggesting that physical robotics is central to the venture’s current work, although the report does not establish a specific product or deployment target.
The technical foundation is model-based reinforcement learning. In this approach, an agent learns a model of its surroundings and uses that model to predict the consequences of possible actions. Rather than repeatedly testing every action in a physical environment, the agent can evaluate imagined outcomes and select a course of action based on those predictions.
For a robot operating in a home, this could mean adapting to a floor plan, furniture arrangement, or unexpected interruption that was absent from its training data. The concept is particularly relevant to service robots, where environments are variable and mistakes can damage property or create safety risks.
Hafner’s research record provides the clearest public evidence for the approach. His PlaNet system used a learned model to support forward planning. Later work on the Dreamer family of agents applied the same broad idea to increasingly difficult tasks in simulated environments.
According to MIT Technology Review, Dreamer 2 reached human-level performance on Atari 2600 games using a world model, while Dreamer 3 solved the Minecraft Diamond challenge by learning to mine the in-game resource. Dreamer 4 reportedly went further by learning from recorded gameplay data without directly interacting with the game during training.
These results matter because they show how planning can reduce the amount of direct interaction required during learning. However, success in games does not by itself demonstrate reliable performance in homes, factories, warehouses, or public spaces. Games provide defined rules and measurable objectives; the physical world is less predictable, harder to simulate accurately, and more consequential when an action fails.
Hafner’s DayDreamer project represents the bridge toward robotics. The project used the Dreamer approach to help robots operate in unfamiliar environments and respond to new experiences, including being pushed over. The startup appears to be pursuing a more ambitious version of that transition, with humanoid machines as the physical platform.
The strongest evidence available is Hafner’s published research history and the account of his current laboratory environment. The claims about PlaNet, Dreamer, and DayDreamer are reported by MIT Technology Review in the context of Hafner’s prior work. They describe research achievements and demonstrations, not independently verified commercial capabilities of the new startup.
There is also no public evidence in the supplied reporting of customer deployments, production contracts, revenue, safety certifications, or comparative testing against other robotics systems. Hafner’s former Google colleague Timothy Lillicrap praised him as an exceptional researcher, but that assessment is an individual endorsement rather than a performance benchmark.
The distinction is important for buyers and developers. A model that predicts game states or adapts to a controlled robot experiment still needs to prove that its predictions remain accurate when sensors are noisy, objects move unexpectedly, and the cost of a wrong decision is high. The startup’s stealth status makes those questions impossible to assess today.
If Hafner’s systems work outside tightly controlled demonstrations, they could address one of robotics’ central bottlenecks: the cost and difficulty of collecting enough real-world experience. Training robots through repeated physical failures is slow, expensive, and potentially unsafe. Planning inside a learned environment could reduce the number of physical trials needed before deployment.
For product teams, the practical test will be whether the agent can transfer knowledge between settings without extensive retraining. A warehouse operator, for example, would need a robot to handle changing layouts and unfamiliar packages. A household robot would need to distinguish between novel situations that are harmless and those that require stopping or asking for help.
That creates additional requirements beyond prediction. Enterprises will need visibility into how the world model was built, how uncertainty is represented, and when the robot falls back to a safe behavior. They will also need tools for evaluating rare events, updating the system after deployment, and separating genuine adaptation from unsafe improvisation.
The work may also intensify competition between two broad strategies in embodied AI: larger models trained on more demonstrations, and more structured agents that learn to simulate and plan. The two approaches are not mutually exclusive, but Hafner’s research puts forward a case that better internal prediction could be as important as more training data.
The first signal will be whether the startup identifies itself and explains its initial product or research objective. Investors, partners, and prospective customers will also look for evidence that the humanoid robots are being used for more than laboratory experimentation.
Technical disclosures will matter more than broad claims. Useful details would include the environments used for evaluation, the amount of real-world data required, how performance changes in unfamiliar settings, and the system’s behavior when its predictions are uncertain. Independent testing on manipulation, navigation, recovery from falls, and safety-critical interruptions would provide stronger evidence than game benchmarks alone.
Finally, the company’s choice of business model will clarify its ambition. A robotics platform, a licensed control system, and a vertically integrated robot product would each imply different requirements for hardware reliability, deployment support, and enterprise economics.
Hafner’s new venture is notable because it targets a specific weakness in current AI systems: the inability to act reliably when conditions differ from training. The use of world models and model-based reinforcement learning offers a plausible route to more deliberate behavior, but the gap between virtual achievement and dependable physical operation remains substantial.
For now, this is a research-led company to watch rather than a validated robotics supplier. The decisive question will be whether Hafner can turn predictive planning into measurable gains in real environments—fewer training trials, safer recovery from surprises, and consistent performance without custom engineering for every new location.