JEPA-Anything Applies One World-Model Recipe Across Seven Domains
Researchers from PhAI Labs, CUHK, Fudan, Stanford, Oxford and Princeton have released JEPA-Anything, a domain-agnostic framework that applies one shared learning recipe to build world models across seven fields, according to MarkTechPost. Its Orthogonal Predictive Factorization splits a JEPA target into orthogonal factors and improves matched dynamics tasks, while planning gains remain environment-dependent.
The reported domains are vision, biology, clinical trajectories, control, molecular dynamics, physical fields and weather. A standard JEPA such as I-JEPA or V-JEPA 2 uses a context encoder, an EMA target encoder and one predictor, which outputs one monolithic target embedding. The researchers describe this as a capacity-allocation problem: high-variance structure dominates, and weaker modes receive conflicting gradients.
OPF splits the latent target of width d into K learned subspaces of width r, with d = K × r. Most experiments used K = 4. Each factor receives a dedicated predictor. The factor predictions are recombined through the Moore-Penrose pseudoinverse of the projector matrix, yielding one complete latent state for decoding, planning or rollout. Three regularizers keep factors useful: an orthogonality loss that keeps columns within each projector orthonormal and different projectors in non-overlapping subspaces; a factor-activity loss that applies a hinge on per-coordinate standard deviation so no factor goes dead; and an encoder-variance loss that sends a direct anti-collapse signal to the online encoder. The OPF loss is added to each domain's original training loss. Domain adapters handle tokenization and encoders, while the core library exposes the shared core as OrthogonalFactorProjection.
Orthogonality mattered for stable synthesis. On CITRIS Interventional Pong, a capacity-matched unconstrained multi-head model had a condition number of 438.52. The orthogonal version reached 1.00005, with cross-factor overlap near zero.
In terminal readout results, zero-shot PBMC clustering on single-cell data rose to 0.7752 in AvgBIO, compared with 0.7194 for Cell-JEPA. Norman perturbation Pearson rose from 0.787 to 0.814. For forecasting over 1,000 clinical events on UK Biobank data, mean PRAUC was 0.718, compared with 0.711 for the matched standard JEPA.
For latent world dynamics, single-intervention MSE on Interventional Pong fell 34.83 percent. Unseen combined interventions improved 12.90 percent, and 6-step free rollout improved 8.58 percent. JEPA-Anything improved reported metrics on all 10 matched dynamics tasks. The benchmarks include CausalWorld, DeepMind Control, PDEBench and WeatherBench2. On APEBench Burgers, 6-step rollout error dropped about 44.7 percent, improving in every seed. For 100-step molecular rollouts with a TrajCast-style backbone, it posted the lowest MAE and RMSD on water, quartz, paracetamol and benzene.
Planning results were mixed. With parameters matched within 0.3 percent, JEPA-Anything improved CEM return on Walker2d and HalfCheetah. Hopper favored standard JEPA.
In scientific analysis, factor analysis nominated IL-18 plus CD73 blockade as a cancer intervention. Wet-lab tests supported it in co-cultures, patient-derived organoids, tumor fragments and mice. Latent orbital modes also recovered Kepler's law with a fitted slope of −1.4991 against the theoretical −1.5.
Compared with other world models, JEPA-Anything uses factorized latent dynamics recombined via pseudoinverse, while V-JEPA 2 uses video JEPA plus action-conditioned V-JEPA 2-AC and a single latent target for video understanding and robot manipulation. DINO-WM builds a world model on pretrained visual features with a single latent target for PointMaze, PushT, Wall and deformables. DreamerV3 combines a world model with an actor-critic trained in imagination and uses categorical latent states across diverse RL domains with fixed hyperparameters. TD-MPC2 uses decoder-free latent dynamics plus MPC, a single latent state, and has been tested on 104 continuous-control tasks across four domains. JEPA-Anything reports per-domain research checkpoints on Hugging Face under an Apache-2.0 license, while V-JEPA 2 offers public checkpoints from 300M to 1B parameters under MIT, some files Apache-2.0; DINO-WM, DreamerV3 and TD-MPC2 list MIT licenses, and TD-MPC2 lists more than 300 checkpoints up to 317M. The checkpoint status for JEPA-Anything was taken from its Hugging Face card and verified Oct. 5, 2026.
The paper and GitHub repository are available. The core code is Apache-2.0, and per-domain research checkpoints are hosted on Hugging Face.