MirroS Releases AgentGarten, a Real-Time Training Ground for Evolving AI Agents
MirroS released AgentGarten, pairing code-based physics with neural rendering so AI agents can act, observe and revise strategies.
The design addresses limits in two existing approaches. Code and game engine environments provide precise state and transparent rules, but their visuals are often rough white models and repeated textures; building hundreds of photorealistic scenes demands huge art and engineering resources. Video-generation world models produce high-definition images, but their physics remains implicit in neural network memory, making the state difficult to query or intervene in. AgentGarten separates the tasks: code handles physics, neural rendering handles appearance.
In a standard interaction loop, an agent issues an action. The code environment calculates collision and state updates, and exports a lightweight geometric sketch, such as depth or surface normals, from the agent's first-person camera. The neural renderer reads that sketch and previous visual memory to generate a photorealistic next frame. The QbitAI article says this keeps physical rules deterministic and controllable, and makes new worlds fast to build. The same visual interface can cover wing-suit flight, robotic arm operation, kitchen cooking and multi-car racing, shifting environment expansion from art labor to code-based procedural extension. The renderer is streaming: it generates a short chunk of video for each action, so the agent sees the result before deciding the next step.
The team recreated OpenAI's 2019 hide-and-seek experiment. Hiders and seekers compete one-on-one. The hider first blocks entrances; then the seeker enters. Agents write Python code to control movement and grasping, and decide only from generated first-person images, with no knowledge of spatial coordinates or opponent positions. Both sides start with a blank manual and maintain a strategy library of individual skills. Each round consists of ten games, followed by review, and the next round randomly selects some skills from the library. Hiders learned to move boards to build cover by round four. Seekers learned to use a ramp to climb a wall by round ten. The QbitAI article compares this with the 2019 study, where reinforcement learning from scratch required about 25 million episodes to discover cover-building and about 100 million episodes to learn ramp-climbing.
The rapid progress comes from a four-stage training loop. Agents receive a task file with goals and rules but no answer key. They act under strict step and time limits, with only camera images and no god's-eye coordinates, minimap or score before the end. After a round, they review and write down what they tried, what they saw, which judgments are uncertain and which hypotheses to test next. The manual is then frozen and passed to the next round as prior knowledge. The QbitAI article says the manuals resemble scientific inquiry: models distinguish observed facts from speculation and leave warnings. Hiders concluded that defense depends on blocking passages, not only sight lines. Seekers learned to test whether an object is truly grasped by pulling slightly and checking movement against room landmarks. A driver warned that a target disk sliding into a vehicle's blind spot does not mean the vehicle has arrived safely.
The same loop ran in four other worlds. In a companion-dog task, the agent had 60 seconds to pet and play ball to maintain the dog's favorability; its score rose from 13 in the first round to 19 in the fourth. In a narrow-bridge passing task, two cars had to negotiate and swap ends on a single-lane bridge; total time fell from 71 seconds to 41 seconds. In cooperative sheep herding, two dogs had to drive four sheep into a pen and keep them there for five seconds; after a first round that timed out with only three sheep, the next three rounds consistently penned all four. In a quarry loader task, a heavy loader had to push rocks, feed material and store it; in round one it barely pushed a rock, but by round four it completed the full flow 31 seconds before the 360-second limit.
To keep visuals synchronized with actions, MirroS started from a base omni-modal model and proposed an Adversarial Forcing training method with inference optimization. The renderer generates streamed chunks rather than a full video at once. An exact replay mechanism runs the first pass without gradients and the second pass with gradients in the same chunk-by-chunk rhythm to reproduce the trajectory and maintain temporal consistency over long interactions. A real-video discriminator provides adversarial signals to avoid degradation and grid-like artifacts and to keep generated results aligned with code-defined geometry. Custom Triton fused operators, CUDA Graph and a lightweight decoder with less than 10 milliseconds of latency support 480p resolution and more than 30 frames per second for closed-loop interaction.
MirroS describes the work as an exploration of Physical RSI, or recursive self-improvement in the physical world. Executable code allows training worlds to be procedurally generated; the learned renderer gives those worlds perceptible physical feedback; and the accumulated manuals let each round build on the last. The QbitAI article says language-model scaling is well underway, while scaling of intelligence in the physical world is only beginning. The team published a blog, technical report, code and project page with the release.