Plan the events
A pretrained vision-language model proposes what should change and when, from the initial image and task context.
qψ → transition + target time
Physical world modeling · Technical report
Predict how an interaction changes the scene.
Use those visual states to generate the motion.
Overview
Event-aligned supervision.
Timed visual states.
Conditional video dynamics.
Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that explicitly learns visual state transitions and uses the predicted states to condition dense video generation. At its core is an image-based transition model trained on paired visual states to predict the next scene configuration from the current image, a transition specification, and elapsed time. A pretrained vision-language model supplies timed transition specifications, and recursive application of the learned transition model produces a sequence of future visual states. A conditional video diffusion model then generates the full rollout from these states and their temporal locations. This separates supervised state-transition learning from conditional dynamics modeling. Experiments on Physics-IQ Verified, PhyGenBench, Pisa-Experiments, and RoboTwin2.0 show improvements over the corresponding video-backbone baselines in benchmark measures of physical realism and manipulation-video fidelity. These results support learned visual state transitions as an effective intermediate representation for physical world modeling.
01 / Framework
Two learned generative components, connected by a timed visual-state interface.
A pretrained vision-language model proposes what should change and when, from the initial image and task context.
qψ → transition + target time
The learned transition model recursively predicts each future state from the current image, event specification and elapsed time.
qφ → sparse future states
A separately trained video model synthesizes the full rollout, softly conditioned on those states and their temporal locations.
qθ → dense video rollout
02 / Qualitative results
03 / Quantitative results
| Model | # params | EFLOPs | Score ↑ | S-IoU ↑ | ST-IoU ↑ | WS-IoU ↑ |
|---|---|---|---|---|---|---|
| Wan2.2 | 14B | 49.100 | 34.2 | 50.3 | 25.2 | 29.9 |
| Cosmos-Predict-2.5 | 2B | 2.560 | 32.2 | 45.5 | 27.2 | 30.1 |
| CogVideoX | 5B | 6.640 | 30.5 | 40.4 | 30.3 | 24.1 |
| Cosmos3-Nano | 16B | 11.802 | 27.6 | 39.7 | 20.2 | 22.1 |
| CoECT | 20B+5B | 29.600 | 19.9 | 27.6 | 23.2 | 14.8 |
| Wan2.2-5B | 5B | 2.739 | 19.4 | 25.0 | 19.4 | 16.3 |
| Open-Sora-v2 | 11B | 5.035 | 19.0 | 27.0 | 21.2 | 10.1 |
| LTX-Video | 13B | 1.152 | 18.4 | 31.0 | 14.1 | 15.0 |
| CausalMotion | 13B | 0.395 | 12.7 | 19.0 | 15.0 | 8.1 |
| Cosmos3-Super | 64B | - | 42.7 | - | - | - |
| MiniMax-H3 | 33B | - | 39.8 | - | - | - |
| Ours (CogVideoX-5B) | 20B+5B | 10.02 | 41.9 | 46.7 | 53.5 | 36.0 |
| Ours (Wan2.2-5B) | 20B+5B | 4.19 | 39.7 | 45.2 | 53.0 | 35.4 |
| Method | # param | Mechanics ↑ | Optics ↑ | Thermal ↑ | Material ↑ | Average ↑ |
|---|---|---|---|---|---|---|
| VideoCrafter2 | 1.4B | 40.00 | 58.00 | 28.89 | 34.17 | 42.08 |
| LaVie | 0.91B | 29.17 | 50.67 | 23.33 | 32.50 | 35.63 |
| Open-Sora v2 | 11B | 55.00 | 66.67 | 47.78 | 50.83 | 56.25 |
| LTX-Video | 13B | 45.83 | 65.33 | 43.33 | 37.50 | 49.37 |
| DreamWorld | 1.3B | 54.17 | 64.67 | 51.11 | 43.33 | 54.17 |
| Cosmos3-Nano | 16B | 57.50 | 75.33 | 48.89 | 46.67 | 58.75 |
| Cosmos-Predict2.5 | 2B | 60.00 | 68.00 | 52.22 | 43.33 | 56.87 |
| PhysVid | 1.7B | 44.17 | 60.00 | 38.89 | 33.33 | 45.42 |
| CausalMotion | 13B | 73.33 | 75.33 | 65.56 | 64.17 | 70.21 |
| Wan2.2 | 5B | 50.83 | 65.33 | 41.11 | 40.00 | 50.83 |
| CogVideoX | 5B | 50.00 | 67.33 | 44.44 | 51.67 | 54.79 |
| VideoREPA | 5B | 44.17 | 67.33 | 43.33 | 44.17 | 51.25 |
| LaMo | 5B | 45.83 | 65.33 | 46.67 | 44.17 | 51.67 |
| Ours (CogVideoX-5B) | 20B+5B | 69.17 | 80.67 | 81.11 | 76.67 | 76.91 |
| Ours (Wan2.2-5B) | 20B+5B | 67.50 | 79.33 | 82.22 | 74.17 | 75.62 |
| Method | L2 ↓ | CD ↓ | IoU ↑ |
|---|---|---|---|
| CausalMotion | 0.1292 | 0.3435 | 0.0966 |
| CogVideoX-5B | 0.1180 | 0.3017 | 0.1481 |
| Cosmos-Predict2.5 | 0.1373 | 0.3871 | 0.1467 |
| Cosmos3-Nano | 0.1250 | 0.3437 | 0.1361 |
| LaMo-5B | 0.1113 | 0.2888 | 0.1590 |
| LTX-Video | 0.1181 | 0.3071 | 0.1017 |
| OpenSoraV2 | 0.1293 | 0.3310 | 0.0846 |
| Ours (CogVideoX-5B) | 0.0949 | 0.2495 | 0.1753 |
| Ours (Wan2.2-5B) | 0.1010 | 0.2637 | 0.1485 |
| Method | EWMScore ↑ | Trajectory ↑ | Interaction ↑ | Perspectivity ↑ | Instruction ↑ | Semantic ↑ |
|---|---|---|---|---|---|---|
| OpenDW | 44.26 | 0.0220 | 0.2556 | 0.6096 | 0.2080 | 0.8685 |
| Ctrl-World | 60.16 | 0.2217 | 0.5268 | 0.7776 | 0.4960 | 0.8782 |
| CogVideoX-5B | 58.04 | 0.2301 | 0.5432 | 0.7952 | 0.5180 | 0.8934 |
| Wan2.2-5B | 59.16 | 0.2372 | 0.5088 | 0.8040 | 0.4560 | 0.8866 |
| Ours (CogVideoX-5B) | 60.51 | 0.3306 | 0.5900 | 0.8160 | 0.5916 | 0.8975 |
| Ours (Wan2.2-5B) | 62.53 | 0.3368 | 0.6012 | 0.8448 | 0.6012 | 0.8943 |
04 / Ablation study
With the same input and planned transition targets, the learned transition model produces more faithful key states and stronger downstream videos than pretrained Qwen-Image-Edit without transition training.
Static locked-off single-shot with fixed frame throughout filmed with constant framerate in real-time. The scene shows a realistic scientific demonstration. The scene only contains the described setup and actions. A row of colorful wooden blocks has been lined up on a wooden table next to a black platform. The wooden stick attached to the black platform rotates clockwise and hits the first block.
State 1Input first frame.
State 2purple wooden block tips left after the stick impact and is leaning left in contact with the blue block.
State 3wooden-block domino row finishes toppling from right to left and is collapsed leftward and resting on the wooden table.
Static locked-off single-shot with fixed frame throughout filmed with constant framerate in real-time. The scene shows a realistic scientific demonstration. The scene only contains the described setup and actions. A piece of clear glass resting on the edge of a light-colored wooden table against a plain white wall. The blue tennis ball rolls from the wooden table over the glass.
State 1Input first frame.
State 2blue tennis ball rolls onto the glass and is fully supported by the glass surface.
State 3blue tennis ball continues off the visible glass and is out of frame to the right.
05 / Citation
BibTeX citation coming soon.