Researchers from PhAI Labs, CUHK, Fudan, Stanford, Oxford and Princeton have released JEPA-Anything, a domain-agnostic framework for building world models. Instead of designing a new predictive model for each field, it applies one shared learning recipe to very different systems. It extends joint-embedding predictive architectures (JEPAs) with a method called Orthogonal Predictive Factorization (OPF). The research team tested it across 7 domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields and weather.
What problem does JEPA-Anything solve?
A standard JEPA, such as I-JEPA or V-JEPA 2, uses a context encoder, an EMA target encoder and one predictor. The predictor outputs one monolithic target embedding. The research team call this a capacity-allocation problem: high-variance structure dominates, and weaker modes get conflicting gradients.
How does Orthogonal Predictive Factorization work?
OPF splits the latent target of width d into K learned subspaces of width r, with d = K × r. Most experiments use K = 4. Each factor gets a dedicated predictor. The factor predictions are then recombined through the Moore-Penrose pseudoinverse of the projector matrix. The result is 1 complete latent state for decoding, planning or rollout.
Three regularizers keep the factors useful:
- Orthogonality loss: keeps columns within each projector orthonormal and different projectors in non-overlapping subspaces.
- Factor-activity loss: a hinge on per-coordinate standard deviation so no factor goes dead.
- Encoder-variance loss: sends a direct anti-collapse signal to the online encoder.
The OPF loss is simply added to each domain’s original training loss. Domain adapters handle tokenization and encoders; the core library exposes the shared core as OrthogonalFactorProjection.
Orthogonality matters for stable synthesis. On CITRIS Interventional Pong, a capacity-matched unconstrained multi-head model had a condition number of 438.52. The orthogonal version reached 1.00005, with cross-factor overlap near zero.
‘;});}
else{x.rows.forEach(function(r){var mx=Math.max(r[1],r[2])*1.08;h+=’
‘;});
if(x.extra){h+=’
Other baselines: ‘+x.extra.map(function(e){return e[0]+’ ‘+e[1]}).join(‘, ‘)+’.
‘;}}
document.getElementById(‘bars’).innerHTML=h;
document.getElementById(‘dl’).innerHTML=’‘+x.d+’‘;
document.getElementById(‘rsrc’).textContent=(x.lo?’Lower is better. ‘:’Higher is better. ‘)+’Source: ‘+x.s;
setTimeout(function(){R.querySelectorAll(‘#bars .bf’).forEach(function(b){b.style.width=b.dataset.w+’%’})},40);setTimeout(rs,60);setTimeout(rs,900);}
drawRes(0);
window.addEventListener(‘resize’,function(){sz(cv);sz(cv3);rs();});window.addEventListener(‘load’,rs);setTimeout(rs,300);
})();


