AI News Feed
Market watch
Research

Kaiming He Team's NAT-ARC Uses ImageNet Pretraining to Push Visual ARC Reasoning

A purely visual ARC method from Kaiming He's team, NAT-ARC uses ImageNet MAE pretraining and reaches 63.4% single-model and 70.2% ensemble pass@2 on ARC-1.

ARC was introduced by François Chollet in 2019. Each task gives a 2D grid up to 30×30 with up to 10 colors and two to five input-output examples. The solver must infer the hidden transformation and apply it to a new input grid. The tasks vary, data are scarce, and precise rule execution is required, making ARC a benchmark for abstract reasoning. Most leading systems are large language models that translate grids into text or symbols. A visual route has emerged, including VARC, which reframes ARC as conditional image-to-image translation and reaches 54 percent with a 19M-parameter model; LoopViT, which adds recurrent reasoning and reaches 65.8 percent with 18M parameters; and Loop-OWM, which uses a video-pretrained model for few-shot learning and reaches 68.5 percent with 10.6M parameters. These models are far smaller than LLM solutions but have begun to close the gap. Their structural weakness is that most start from random initialization and do not benefit from large-scale pretraining, unlike LLMs.

NAT-ARC inserts ImageNet MAE pretraining before the VARC visual pipeline. MAE, proposed by Kaiming He at FAIR in 2022, is a self-supervised method that masks much of an image and trains the model to reconstruct it, encouraging understanding of shape, structure, and spatial relationships. NAT-ARC initializes the visual encoder with MAE encoder weights trained on ImageNet, which contains about 1.3 million natural images, while the decoder starts from random initialization. The pipeline has three stages: first, pretrain the encoder with ImageNet MAE; second, train the full encoder-decoder offline on the ARC training set; third, at test time, run per-task LoRA fine-tuning using the two to five demonstration pairs for each task. The method uses a public MAE checkpoint and does not pay for additional pretraining. Because the checkpoint was designed for larger ImageNet images and ARC grids are 64×64 pixels, NAT-ARC discards the original patch embedding and positional encoding, keeps the backbone weights, and uses 2D RoPE for positional encoding.

The paper uses attention visualization to explain why natural-image pretraining helps. When models with no pretraining, ImageNet MAE pretraining, and ARC-style grid MAE pretraining are placed before the same ARC task before ARC training, the randomly initialized model has diffuse attention. The ImageNet-pretrained model already separates foreground from background and focuses on meaningful patterns in the grid. The researchers also selected 15 ARC tasks that improved significantly after pretraining. Manual annotation found that six are match-and-copy tasks, where the model must identify matching objects, colors, or patterns and copy the corresponding structure to the output, and six are connected-component reasoning tasks, where the model must identify, fill, or recolor spatially connected regions. These abilities correspond to visual priors from natural images.

On ARC-1, the best single NAT-ARC model uses the huge scale, has 0.6B parameters, and scores 63.4±0.7 percent pass@2. An ensemble that pools models trained with three pretraining strategies, no pretraining, ImageNet MAE pretraining, and ARC-style grid MAE pretraining, and uses majority voting reaches 70.2±0.6 percent pass@2, with about 2B total parameters. For comparison, The ARChitects, fine-tuned specifically for ARC, reaches 71.6 percent with an 8B language model. The visual route thus approaches a specialized LLM system with about one quarter of the parameters, according to the paper.

The paper also reports that pretraining changes scaling behavior. Without pretraining, performance improves from base to large but drops from large to huge, suggesting that the small ARC dataset causes larger models to overfit rather than generalize. With ImageNet pretraining, the scaling curve becomes positive across base, large, and huge scales, and larger models perform better. In offline training, ImageNet-pretrained models converge faster and reach higher final accuracy at all three scales.

In a visualization experiment, the researchers trained an autoencoder to map ImageNet images to a discrete 60×60, 10-color grid representation, effectively translating a cat image into an ARC-like grid. They then applied ARC transformations to that grid, rotating it 180 degrees and adding a blue border, and decoded it back to pixels. The cat was rotated and the border was added. The experiment provides visual evidence that natural images and ARC grids share a representation space, so transformations learned on ARC grids can apply to latent representations of natural images, and vice versa.

The first author is Xiaoman Delores Ding, a member of MIT CSAIL in Kaiming He's group and first author of VARC. Other authors include Keya Hu, one of Kaiming He's first female students and a Shanghai Jiao Tong University graduate, and Katelyn Gan and Victor Yin, both MIT undergraduates and student researchers at CSAIL. Kaiming He is the corresponding author, known for ResNet and MAE. VARC was accepted at CVPR 2026, and NAT-ARC is its direct follow-up. The team has now published two papers pursuing a purely visual route to ARC while LLM-based methods dominate the field.