Reka Releases Rho-1, a 19B Omni-Reasoning Model That Generates Video and Emits Robot Actions
Reka released Rho-1, a 19B omni-reasoning model that handles text, image, video and robot actions in one network.
Most multimodal systems today are pipelines: a central model plans, then hands jobs to specialists for images, video or detection. Each handoff adds latency, and each specialist sees only a narrow request. Rho-1 removes those handoffs by turning text, vision and robotic actions into tokens inside one context window. According to Reka, one unedited session shows the full loop: the model draws a lighthouse, boxes it, animates it, edits the video into a snowstorm, and explains the difference, all in five turns with no tool call and no second model.
Every input and output uses one of two native formats. Discrete tokens carry text, symbolic reasoning and high-level commands. Continuous tokens carry image latents, video frames, robot actions and proprioception. Each transformer block holds two expert weight streams: an understanding stream for language and visual parsing, and a generation stream that denoises latents into images and video. Both streams share attention and operate over the same KV cache. When a reply needs pixels, the understanding stream emits a discrete handoff token, and the generation stream renders from the full accumulated state. Training combines next-token prediction for discrete sequences with flow matching for continuous outputs.
Reka says the design has practical effects: bounding boxes come out as coordinate tokens rather than from a separate detector, and a video's first frame reuses the in-context image representation instead of a re-encoded copy.
The base model generates video at 0.79x real time at the median, with a watchable stream starting in roughly six seconds. Reka measured 7.0 seconds to a first clip, against what it presents as an illustrative 13.8 seconds for a multi-agent pipeline. A distilled variant cuts denoising from 99 steps to 8 with minimal quality loss reported, returning a 5.3-second clip in about a second. In Reka's internal tests it matched the fastest dedicated image models and was the quickest model tested to the first text token. These are vendor-run tests, not independent benchmarks.
Rho-1 streams continuously, clip after clip, with new instructions entering through the understanding stream and updating state mid-rollout. Reka demonstrates one opening forked into "bank left" and "bank right" continuations. For robotics, actions and future frames decode from the same latent state; a LIBERO simulation episode shows Rho-1 emitting 7 action channels. To scale past scarce teleoperation logs, Reka pairs Rho-1 with its Inverse Dynamics Model, which infers control signals from raw video.
Reka's own comparison places Rho-1 against ByteDance's BAGEL, BAAI's Emu3.5 and Google's Genie 3. Rho-1 is listed at 19B parameters, with text, image, video, actions and proprioception as both inputs and outputs, native video generation capped at 672×384, continuous rollouts with real-time steering, and native continuous action tokens. BAGEL is listed at 14B total parameters with 7B active under a mixture-of-transformers design, taking and producing text and images, without native video generation, real-time steering or robot actions, and with open weights under Apache 2.0. Emu3.5 is listed at 34B parameters, producing image frames rather than native clips, with embodied manipulation demos and open weights under Apache 2.0. Genie 3's parameter count is not disclosed; it is described as an interactive video world generating 720p at 24 frames per second, taking navigation actions rather than emitting them, and available through Project Genie for Google AI Ultra subscribers. Rho-1 is not available as open weights.
Reka said Rho-1 was trained on 320 H100 GPUs for three months. Access today is limited to a research preview requested through contact@reka.ai, with no public weights, API or pricing announced. The model's video output is capped at 672×384.
Editor's Summary
Reka has released a 19B omni-reasoning model, Rho-1, that handles text, images, video and robot actions within a single network using two expert streams over a shared KV cache. Reka reports a distilled variant producing 5.3-second clips in about a second, though the speed figures come from vendor-run tests. The model is limited to a research preview, with no open weights, API or pricing announced, and its video output is capped at 672×384.