Physical world modeling · Technical report

STRIKE: Learning Visual State Transitions for Physical World Modeling

Predict how an interaction changes the scene.
Use those visual states to generate the motion.

1Applied Intuition · 2University of Southern California · 3University of California, Berkeley

§Project Lead · †Corresponding author: wei.zhan@applied.co

arXiv (on release)

Overview

States before
trajectories.

Event-aligned supervision.
Timed visual states.
Conditional video dynamics.

Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that explicitly learns visual state transitions and uses the predicted states to condition dense video generation. At its core is an image-based transition model trained on paired visual states to predict the next scene configuration from the current image, a transition specification, and elapsed time. A pretrained vision-language model supplies timed transition specifications, and recursive application of the learned transition model produces a sequence of future visual states. A conditional video diffusion model then generates the full rollout from these states and their temporal locations. This separates supervised state-transition learning from conditional dynamics modeling. Experiments on Physics-IQ Verified, PhyGenBench, Pisa-Experiments, and RoboTwin2.0 show improvements over the corresponding video-backbone baselines in benchmark measures of physical realism and manipulation-video fidelity. These results support learned visual state transitions as an effective intermediate representation for physical world modeling.

Predict the state change, then the motion

01 / Framework

From events to states to motion

Two learned generative components, connected by a timed visual-state interface.

01

Plan the events

A pretrained vision-language model proposes what should change and when, from the initial image and task context.

qψ → transition + target time

02

Predict visual states

The learned transition model recursively predicts each future state from the current image, event specification and elapsed time.

qφ → sparse future states

03

Generate the dynamics

A separately trained video model synthesizes the full rollout, softly conditioned on those states and their temporal locations.

qθ → dense video rollout

Framework overview

02 / Qualitative results

Watch the interactions unfold

03 / Quantitative results

Benchmark evaluation

Physics-IQ Verified
Model# paramsEFLOPsScore ↑S-IoU ↑ST-IoU ↑WS-IoU ↑
Wan2.214B49.10034.250.325.229.9
Cosmos-Predict-2.52B2.56032.245.527.230.1
CogVideoX5B6.64030.540.430.324.1
Cosmos3-Nano16B11.80227.639.720.222.1
CoECT20B+5B29.60019.927.623.214.8
Wan2.2-5B5B2.73919.425.019.416.3
Open-Sora-v211B5.03519.027.021.210.1
LTX-Video13B1.15218.431.014.115.0
CausalMotion13B0.39512.719.015.08.1
Cosmos3-Super64B-42.7---
MiniMax-H333B-39.8---
Ours (CogVideoX-5B)20B+5B10.0241.946.753.536.0
Ours (Wan2.2-5B)20B+5B4.1939.745.253.035.4
PhyGenBench
Method# paramMechanics ↑Optics ↑Thermal ↑Material ↑Average ↑
VideoCrafter21.4B40.0058.0028.8934.1742.08
LaVie0.91B29.1750.6723.3332.5035.63
Open-Sora v211B55.0066.6747.7850.8356.25
LTX-Video13B45.8365.3343.3337.5049.37
DreamWorld1.3B54.1764.6751.1143.3354.17
Cosmos3-Nano16B57.5075.3348.8946.6758.75
Cosmos-Predict2.52B60.0068.0052.2243.3356.87
PhysVid1.7B44.1760.0038.8933.3345.42
CausalMotion13B73.3375.3365.5664.1770.21
Wan2.25B50.8365.3341.1140.0050.83
CogVideoX5B50.0067.3344.4451.6754.79
VideoREPA5B44.1767.3343.3344.1751.25
LaMo5B45.8365.3346.6744.1751.67
Ours (CogVideoX-5B)20B+5B69.1780.6781.1176.6776.91
Ours (Wan2.2-5B)20B+5B67.5079.3382.2274.1775.62
PISA Experiments
MethodL2 ↓CD ↓IoU ↑
CausalMotion0.12920.34350.0966
CogVideoX-5B0.11800.30170.1481
Cosmos-Predict2.50.13730.38710.1467
Cosmos3-Nano0.12500.34370.1361
LaMo-5B0.11130.28880.1590
LTX-Video0.11810.30710.1017
OpenSoraV20.12930.33100.0846
Ours (CogVideoX-5B)0.09490.24950.1753
Ours (Wan2.2-5B)0.10100.26370.1485
RoboTwin2 val500
MethodEWMScore ↑Trajectory ↑Interaction ↑Perspectivity ↑Instruction ↑Semantic ↑
OpenDW44.260.02200.25560.60960.20800.8685
Ctrl-World60.160.22170.52680.77760.49600.8782
CogVideoX-5B58.040.23010.54320.79520.51800.8934
Wan2.2-5B59.160.23720.50880.80400.45600.8866
Ours (CogVideoX-5B)60.510.33060.59000.81600.59160.8975
Ours (Wan2.2-5B)62.530.33680.60120.84480.60120.8943

04 / Ablation study

Learning the transition model matters

With the same input and planned transition targets, the learned transition model produces more faithful key states and stronger downstream videos than pretrained Qwen-Image-Edit without transition training.

05 / Citation

Cite this work

BibTeX citation coming soon.