Reka AI’s Rho-1 combines text, images, video, and robot control in one 19B research model, signaling a push toward unified AI systems.

Reka AI has released a research preview of Rho-1, a 19-billion-parameter model designed to process and generate text, images, video, and robot-control actions within one neural network. The release matters because it takes aim at a common architecture for capabilities that are typically split across language models, vision systems, video generators, and robotics policies.
According to reporting by The Decoder, Rho-1 treats different media and actions as tokens in a shared context window rather than handing tasks to separate models through tool calls. Reka says the system can generate continuous video in real time and respond to new instructions while a scene is unfolding, without restarting the generation process.
The announcement places Reka AI among researchers exploring whether a single model can serve as a more general-purpose interface for perception, generation, and action. It is still a research preview, however, and the available evidence does not establish how Rho-1 performs against leading commercial models or real-world robotic systems.
Most production AI stacks divide multimodal work into components. A language model may interpret a user request, a vision model may inspect an image, a video model may generate frames, and a separate policy may translate observations into physical movements. These components can be connected, but each handoff adds engineering complexity, latency, and potential failure points.
Rho-1 is intended to avoid that arrangement. The model reportedly represents text, images, video, and robot actions in one shared sequence. In principle, the same context can contain an instruction, visual observations, generated frames, and action outputs, allowing the system to connect what it sees with what it should do.
The approach also gives Rho-1 a direct relationship between visual prediction and control. The Decoder reports that the same weights used to predict camera images also drive robot movements. That does not mean the model is ready for broad deployment in physical environments, but it reflects a research direction in which video prediction and action planning are trained together rather than treated as unrelated problems.
The model was trained on 320 H100 GPUs over approximately three months, according to The Decoder’s account of the release. Reka AI also developed an inverse dynamics model to address a central robotics problem: high-quality robot-control data is scarce compared with ordinary internet video.
Inverse dynamics generally attempts to infer the action that produced an observed change in a scene. In Reka’s reported approach, that capability is used to extract possible control signals from non-robot videos. This could expand the amount of training material available to a system that must learn how actions affect the world, although internet video does not provide the same reliability, embodiment, or sensor detail as demonstrations collected from a specific robot.
The reported compute requirement is notable because Rho-1 has 19 billion parameters, substantially smaller than many frontier systems commonly discussed in the market. Still, the 320 H100 GPUs used for three months represent a major training investment. The figure should therefore be read as a description of the research run, not as evidence that the model is inexpensive to train or operate for every deployment scenario.
The strongest details available in this report come from The Decoder’s coverage of Reka AI’s research preview. MarkTechPost separately described Rho-1 as a 19-billion-parameter omni-reasoning model that understands and generates video while producing robot actions, but the supplied MarkTechPost material does not include the full article or independent test results.
The available evidence does not provide benchmark scores, model-access terms, latency measurements, safety evaluations, or demonstrations that can be independently assessed here. Claims about real-time video generation, adaptive responses, and the relationship between visual prediction and robot movement should therefore be treated as company-reported or source-reported capabilities rather than settled performance findings.
That distinction is important for builders and buyers. A unified model may reduce the need to maintain multiple interfaces, but it can also make failures harder to isolate. A system that produces fluent language and convincing video may still make unsafe or physically incorrect decisions when asked to control equipment. Rho-1’s research-preview status leaves open questions about reliability, evaluation methodology, and access to the model and its weights.
Reka AI has previously worked on multimodal systems. The company released Reka Core in April 2024, positioning it as a competitor to GPT-4, Claude 3, and Gemini Ultra on benchmark evaluations. Rho-1 extends that direction beyond multimodal understanding toward video generation and embodied action.
For product teams, Rho-1’s main significance is architectural. If one model can maintain a shared representation of language, visual content, temporal changes, and actions, developers could build applications around a single inference surface rather than coordinating several specialized services. Potential use cases include interactive video assistants, simulation environments, visual agents, and robotics research tools.
The trade-off is concentration of risk. Specialized systems can be replaced or tuned independently, while a unified model may require retraining or broader evaluation when one modality underperforms. A model that is strong at text but weak at temporal reasoning, for example, could produce plausible explanations while failing to track a changing physical scene.
Enterprises considering such systems will also need to examine infrastructure and governance. Continuous video generation can create high bandwidth and storage demands. Robot-control outputs require hard safety boundaries, monitoring, and a fallback policy that can override the model. For regulated or safety-sensitive deployments, the ability to explain which visual evidence led to an action may matter as much as the model’s general benchmark performance.
The broader research debate concerns world models: systems intended to represent how environments change and how actions affect them. Rho-1 fits that agenda by combining visual prediction with action generation, but a unified token stream alone does not prove that the model has a dependable physical model of the world. Long-horizon planning, unusual conditions, and transfer between different robots remain difficult tests.
The first signal to watch is whether Reka AI publishes reproducible evaluations for each capability separately and in combination. Useful results would include video quality and temporal consistency, instruction-following under changing scenes, and robot-control success rates on clearly defined tasks.
Access will be another key indicator. If Rho-1 becomes available to outside researchers, independent testing could clarify whether its unified architecture delivers practical advantages over pipelines built from specialized models. Details on inference hardware, context length, licensing, and latency will determine whether the system is useful beyond demonstrations.
For robotics, the most important follow-up will be evidence from real hardware rather than internet-video or simulation results. Researchers will need to test safety, recovery from unexpected events, performance across embodiments, and the effect of noisy camera and sensor inputs.
Finally, the market will be watching whether Reka’s approach remains a research preview or becomes a product platform. The answer will show whether omni-models can move from an appealing architecture into dependable tools for AI agents, video applications, and physical automation.
Rho-1 is a meaningful research signal because it frames multimodality as a shared modeling problem that includes action, not only text, images, and video. That could simplify parts of the AI stack and make it easier to build systems that connect perception with response.
But the announcement is not yet proof that one model is better than a carefully engineered collection of specialists. Until independent evaluations and real-robot results are available, builders should treat Rho-1 as an architectural experiment with promising scope—not as a validated replacement for production multimodal or robotics systems.