AI News Feed
Market watch
Research

Kaiming He Team Releases Multimodal Harness Framework VISTA

Kaiming He's team proposes VISTA, a visual-native harness for multimodal models to store and revisit raw frames during reasoning.

Traditional harnesses for long tasks record task state in text and code. In November 2025, Anthropic described such a harness for long coding tasks, using task lists, progress files, and code records to carry work across context windows. But visual environments require remembering object positions, orientations, and changes before and after each action; a first look does not reveal which detail will matter later. Text notes cannot fully capture this information, and early frames can drop out of context or be replaced by summaries. In ARC-AGI-3, an interactive visual reasoning benchmark that does not provide rules or goals in advance and requires agents to explore through observation and action, one existing approach converts frames into numeric text grids and has the model write programs to simulate rules and test actions. That becomes harder as environments grow more complex.

VISTA stores every frame returned by the environment outside the context, including intermediate animation frames, and retrieves them when the model needs them. In ARC-AGI-3, it can compare frames from different moments and zoom in on local regions to inspect orientation markers on small squares. Text notes record the model's current understanding, while raw frames preserve evidence for renewed judgment. The researchers call this explicit attention to interaction history: the model chooses which visual information re-enters context based on the question it is considering. The system must save raw frames and provide tools to retrieve and inspect them at any time.

VISTA has three core components: visual observation, lossless visual memory, and active visual inspection. Visual observation provides environment frames to the model. In ARC-AGI-3 experiments, VISTA upscales the official 64x64 frames to 512x512 PNG images while preserving color, appearance, and spatial relations. Lossless visual memory saves all frames returned after each action, including animation intermediate frames, indexed by turn number and frame number, so the model need not decide in advance what is worth remembering. Active visual inspection uses an inspect tool: the model can specify a historical turn and bring up the corresponding frame, crop or zoom into local regions to check details, or fetch multiple frames at once to compare changes before and after an action. The process is like giving the model a visual archive it can consult.

In VISTA's execution flow, each turn begins with the model observing the current frame and available actions. It combines accumulated experience, judges the current environment state, and proposes a hypothesis for the next action. If existing information is insufficient, the model can call tools to review historical frames, zoom into local regions, or read pixel information to seek evidence. After finding evidence, it predicts what changes the action may bring. After the action is executed, it compares the actual result with the prediction, checks whether its judgment was correct, and revises its understanding of game rules. The process repeats, letting the model explore and accumulate experience. To maintain coherence in long tasks, VISTA adds two text notes: GUIDE.md records reusable rules and experience across levels, and WORKING.md records the current level's state, progress, and plans. When the context approaches its limit, the model first writes a handoff summary, then enters a new context window to continue, with text notes, action history, and the visual archive retained.

VISTA saves visual history but does not put all images into the model context. In the full setup, after each action the framework by default shows only the last frame to the model. Intermediate frames are stored in the visual archive and retrieved on demand. This preserves complete visual information without letting large numbers of historical images occupy context space. VISTA also does not train a new model for this purpose. Environment understanding, action planning, and reasoning are still performed by an off-the-shelf multimodal model, while the harness executes tool calls, saves visual history, and provides evidence when needed. VISTA changes how models obtain, save, and use visual information, not the model itself.

In evaluations, with VISTA, Claude Opus 5.0 completed all 25 public games in ARC-AGI-3, reached a full relative human action efficiency score of 100, and used 57.4% fewer game actions than the first-time human baseline. GPT-5.6 Sol also completed all games, scoring 99. The team further tested browser games, mazes, and connect-the-dots tasks, and listed embodied tasks closer to the physical world as a future research direction. On GameWorld's 170 tasks, the same GPT-5.6 Sol model improved its success rate from 40.0% under the official base framework to 63.3%. On 10 games in AI GameStore, the composite score rose from 47.3 to 140.3, with the human median normalized to 100. On 39 maze and connect-the-dots questions from BabyVision, accuracy rose from 41.0% to 63.2%; these tasks contain only static images, and the model still improved its judgment by zooming into local regions and checking pixels. The same VISTA framework required only minor adaptation for different games and puzzles.

The paper's three co-first authors are Qiushi Han, Keya Hu, and Linlu Qiu. The other two authors are Cathy Wu and Kaiming He. Qiushi Han, also known as Josh Han, is a PhD student at the MIT Operations Research Center, advised by Cathy Wu. Keya Hu is a PhD student in MIT's Department of Electrical Engineering and Computer Science, co-advised by Kaiming He and Jacob Andreas. She graduated from Shanghai Jiao Tong University's ACM class, and her research interests center on the intersection of language and vision, aiming to build more data-efficient and generalizable agents. Linlu Qiu is a PhD student in MIT's Department of Electrical Engineering and Computer Science and the Computer Science and Artificial Intelligence Laboratory, advised by Yoon Kim and Jacob Andreas. Her research includes natural language processing and machine learning, and she has conducted research at Google Research and Meta FAIR. Cathy Wu is an associate professor in MIT's Department of Civil and Environmental Engineering and the Institute for Data, Systems, and Society, studying how to use machine learning and reinforcement learning to improve complex systems such as transportation. She received her bachelor's and engineer's degrees from MIT and her PhD from the University of California, Berkeley. Kaiming He is a tenured associate professor in MIT's Department of Electrical Engineering and Computer Science and a main author of ResNet, Mask R-CNN, and MAE. He received his bachelor's degree from Tsinghua University and his PhD from the Chinese University of Hong Kong, worked at Microsoft Research Asia and FAIR, and joined MIT in 2024. The paper is available at arXiv:2610.02200.