AI News Feed
Market watch
Research

Kaiming He Team's NAT-ARC Uses ImageNet Pretraining to Advance Pure-Vision ARC Solving

A team led by Kaiming He has proposed NAT-ARC, a pure-vision method that uses ImageNet MAE pretraining to solve ARC abstraction tasks. Its best single model reached 63.4% pass@2 on ARC-1, and an ensemble reached 70.2%.

ARC was introduced in 2019 by François Chollet, the creator of Keras. Each task gives two to five input-output examples on a two-dimensional grid of up to 30 by 30 cells with up to 10 colors. A solver must infer the hidden transformation rule from the examples and apply it to a new grid. Mainstream approaches have translated the grid into text or symbol sequences and handed them to a language model for reasoning.

A separate visual line of work has sought to process ARC grids as images directly. VARC, from the same MIT team, reframed ARC as a conditional image-to-image translation task and reached 54% with a 19M-parameter vision model. LoopViT added recurrent reasoning and reached 65.8% with 18M parameters. Loop-OWM used a video-pretrained model for few-shot learning and reached 68.5% with 10.6M parameters. These models are orders of magnitude smaller than LLM-based systems, but most visual ARC models still train from random initialization and do not benefit from large-scale pretraining, which limits scaling.

NAT-ARC inserts an ImageNet MAE pretraining step before the VARC visual pipeline. MAE, or Masked Autoencoder, was proposed by He at FAIR in 2022. It masks most of an image and asks the model to reconstruct the original from the remaining patches, a task that teaches shape, structure and spatial relationships. NAT-ARC uses MAE encoder weights pretrained on ImageNet, a dataset of about 1.3 million natural images including cats, dogs, flowers and plants, to initialize its visual encoder. The decoder starts from random initialization. The method uses a public MAE checkpoint without additional pretraining cost. Because the checkpoint was designed for larger ImageNet images while ARC grids are 64 by 64 pixels, NAT-ARC discards the original patch embedding and positional encoding, keeps the backbone weights and replaces positional encoding with 2D RoPE.

The pipeline has three steps: pretrain the encoder with MAE on ImageNet; offline-train the full encoder-decoder on the ARC training set; and at test time apply per-task LoRA fine-tuning. During fine-tuning, each task provides only two to five demonstration pairs, from which the model attempts to learn that task's rule.

Attention visualizations offer a clue to why natural-image pretraining helps abstract grid reasoning. Before any ARC training, a randomly initialized model spreads attention evenly, while an ImageNet MAE-pretrained model already separates foreground from background and focuses on meaningful pattern regions in the grid. The model has never seen an ARC grid, but its ImageNet-learned ability to distinguish objects from background transfers to ARC grids. A task-level breakdown identified 15 ARC tasks that improved markedly after pretraining. Manual annotation found that six involved match-and-copy and six involved connected-component reasoning. These correspond to visual priors from natural images: separating a cat from a grass background becomes recognizing a connected colored region in a grid.

On ARC-1, the best single NAT-ARC model used the huge scale with 0.6B parameters and scored 63.4±0.7% pass@2. For the ensemble, the researchers pooled models trained under three pretraining strategies, including no pretraining, ImageNet MAE pretraining and ARC-style grid MAE pretraining, and used majority voting. The ensemble reached 70.2±0.6% pass@2 with about 2B total parameters. As a reference, The ARChitects, fine-tuned specifically for ARC with an 8B language model, reached 71.6%. The visual route thus used about one-quarter of the parameters to approach a specialized LLM system.

The paper reports that pretraining addresses the visual route's scaling bottleneck. Without pretraining, performance improved from base to large but fell from large to huge, a sign that the small ARC dataset leads to overfitting as model capacity grows. With ImageNet pretraining, the scaling curve turned positive across base, large and huge, and convergence was faster at all three scales. The researchers also ran a visualization experiment: they trained an autoencoder to map ImageNet images into a 60 by 60, 10-color discrete grid representation, effectively translating a cat photo into an ARC-style grid. NAT-ARC then applied ARC transformations, rotating the grid 180 degrees and adding a blue border, before decoding it back to pixels. The cat was rotated and the border was added, which the paper presents as visual evidence that natural images and ARC grids share a representation space.

The first author, Xiaoman Delores Ding, is a member of MIT CSAIL in He's group and was also first author of VARC. Other authors include Keya Hu, one of He's first female students and a graduate of Shanghai Jiao Tong University, as well as Katelyn Gan and Victor Yin, MIT undergraduates and CSAIL student researchers. The corresponding author, Kaiming He, is the author of ResNet and MAE. VARC was accepted at CVPR 2026, and NAT-ARC is its direct follow-up. The paper is available at https://eccv.ecva.net/virtual/2026/poster/5527.