NeurIPS 2026 · Under Review

Phys4D: Injecting Physical Structure into Video Diffusion via Lifted RGB-D-Motion 4D Interfaces

A structure-focused, appearance-constrained physics injection framework that transfers simulation-derived physical structure into pretrained video diffusion models — while preserving their real-video generative prior.

Anonymous Authors

Phys4D teaser: simulation teaches physics, not appearance
Phys4D: simulation teaches physics, not appearance. Phys4D uses simulation-derived depth, motion, masks, and trajectories to inject physical structure into pretrained video diffusion models through a 4D world interface, while blocking simulator RGB as the dominant appearance target. This improves observable physics and 4D geometry–motion consistency while preserving the real-video generative prior.

Abstract

Recent video diffusion models generate increasingly realistic videos, yet visual plausibility alone does not imply physical understanding. Physical laws govern evolving 3D scene states, while videos observe only their 2D projections. We introduce Phys4D, a structure-focused, appearance-constrained physics injection framework that transfers simulation-derived physical structure into pretrained video diffusion models while preserving their real-video generative prior. Phys4D operationalizes physics-aware video generation through a prior-preserving RGB-D-motion interface. Our key principle is simulation teaches physics, not appearance: simulation provides dense structural supervision, including depth, motion, and masks, while adaptation remains anchored to the pretrained real-video prior rather than drifting toward simulator-specific visual statistics. To this end, Phys4D exposes geometry and motion with lightweight prediction heads, injects local physical structure through gated geometry-motion supervision, and aligns sequence-level evolution using a simulation-grounded trajectory objective. Supported by a scalable coupled-physics simulation system, Phys4D improves observable physical behavior, geometry-motion consistency, and trajectory-level evolution across multiple video generation backbones, while maintaining visual realism and fidelity. We release the code, datasets, introduction slides, and video demos on this project page.

Key Principle
“Simulation teaches physics, not appearance.”

Simulated videos are procedurally rendered and visually mismatched with internet-scale real video. Directly fitting a generator to them transfers simulator-specific visual bias and degrades the pretrained prior. Phys4D therefore uses simulation selectively: prioritize physical structure while constraining appearance drift.

Key Contributions

1

Structure-transfer formulation. We formulate physical knowledge injection for video generation as a structure-transfer problem, where simulation provides geometry, motion, and trajectory supervision while the pretrained real-video generative prior is preserved.

2

The Phys4D framework. A structure-focused physics injection framework that constructs a prior-preserving RGB-D-motion interface, injects local physical structure through gated geometry-motion supervision, and aligns global scene evolution through a unified trajectory-level policy objective.

3

Scalable coupled-physics simulator. We build a unified simulation system spanning primitive and coupled physical regimes, and show that Phys4D improves fine-grained physical behavior, geometry-motion consistency, and trajectory-level evolution across multiple video backbones.

4

Open release. To facilitate reproducibility and broader use, we release our code, simulation datasets, introduction slides, and video demos on this project page.

Method Overview

Phys4D injects simulation-derived physical structure into a pretrained video diffusion model through a prior-preserving RGB-D-motion interface, in three stages.

Overview of the Phys4D three-stage training pipeline
Overview of the Phys4D training pipeline. Stage I freezes the video DiT and trains depth/motion heads to build an RGB-D-motion interface. Stage II injects local physical structure through mask-gated high-noise adapters with simulation depth, motion, and warp supervision. Stage III aligns generated 4D point trajectories with simulator trajectories using a 4D Chamfer reward.
Stage I

Prior-Preserving 4D Interface Construction

Two lightweight auxiliary heads are attached to the pretrained DiT backbone: a motion head predicting inter-frame optical flow and a depth head predicting per-frame depth. The heads are supervised by pseudo-labels from off-the-shelf estimators on curated internet videos and self-generated videos.

  • No simulation data is used
  • Backbone frozen throughout
  • Only heads trained — the RGB generator is unaltered
Stage II

Mask-Gated Local Physics Injection

Simulation supervision is restricted along three axes — when, where, and what to learn. A LoRA-style residual physics adapter is the only trainable pathway through which simulation modifies the generator; a noise gate activates it at high-noise structure-forming steps, and a dynamic-region gate localizes it to regions where physical changes occur.

  • When: high-noise denoising steps
  • Where: dynamic regions (mask-gated)
  • What: geometry & motion, not simulator RGB
Stage III

Simulation-Grounded 4D Alignment via Direct Policy Optimization

Adjacent-frame losses cannot tell whether the completed evolution follows the intended physical trajectory. Stage III lifts generated RGB-D-motion into 4D point trajectories and aligns them with simulator trajectories from paired rollouts, using a symmetric 4D Chamfer reward under a direct policy optimization objective.

  • Offline & simulation-paired training scenes only
  • Heads and gates frozen to avoid reward hacking
  • No simulator input needed at inference
Mask-gated physics adapter
h = h + α(σ) · ĜA(h)

α(σ) is the noise gate (when), Ĝ the learned dynamic-region gate (where), and A the lightweight physics adapter — the only pathway simulation may use.

4D point distance
d4D(p,q) = ‖xpxq22 + λtp − τq|2

A unified metric over lifted spatiotemporal points, combining 3D position and normalized time so that trajectories are compared jointly in space and time.

Trajectory alignment reward
R4D(V;s) = −CD4D(Pdyngen, Pdynsim)

Symmetric 4D Chamfer distance between generated and simulated dynamic point sets from the same paired rollout — a simulation-grounded alignment reward, not a universal physical-law reward.

Simulation as a Scalable Physical Teacher

Real videos provide rich appearance priors but rarely expose dense 4D physical states. We use simulation not as a target visual domain, but as an executable physical teacher that converts solvers and constraints into video-compatible structural supervision.

Simulation as a scalable physical teacher
Left: Asynchronous rollouts expand 200 scenes into 250K environments and 1.25M videos. Top right: The simulator covers primitive physical regimes. Middle right: These regimes are composed into coupled physical rollouts. Bottom right: Each rollout exports RGB, depth, motion, masks, and 4D point trajectories.
200
base scenes
250K
environments
1.25M
annotated videos
3.5K
hours of rollouts
15TB+
multi-modal annotations

Primitive Physical Regimes

Rigid Bodies Articulated Structures Garments Fluids Thermodynamics Deformables Inflatables Ropes Granular Materials

Coupled Physical Rollouts

Rigid × Granular Fluid × Inflatable Rigid × Deformable Garment × Articulated Deformable × Garment

Unified 4D Supervision Exported per Rollout

RGB frames Depth Scene flow / motion Dynamic masks Camera poses 4D point trajectories

Built on Isaac Sim and GarmentLab, integrating rigid-body dynamics, PBD, FEM, and task-specific solvers into a unified scene-level framework. Domain randomization over mass, friction, restitution, stiffness, gravity, and scale exposes regimes governed by shared physical laws while preventing memorization of fixed motion templates.

Video Demos

Physics-IQ scenarios under the matched switch-frame I2V protocol. Each clip shows the text prompt, the WAN2.2-5B baseline, and the same backbone with Phys4D.

Left: input prompt Middle: WAN2.2-5B (baseline) Right: WAN2.2-5B + Phys4D (ours)
Solid Mechanics

Ball placed on a rotating platform

A grabber arm lowers a tennis ball onto cardstock propped on a clockwise-rotating platform. Phys4D keeps contact and support consistent as the platform turns.

Solid Mechanics

Ball rolling out of a pipe

An orange ball rolls out of a black pipe across a coffee table. The baseline distorts object scale and count; Phys4D maintains correct object identity and a coherent rolling trajectory.

Solid Mechanics

Knife slicing a tangerine

A knife slices through a halved tangerine on a glass cutting board — a contact and separation event requiring consistent local geometry through the cut.

Fluid Dynamics

Beverage dispenser pouring into a glass

Grapefruit juice pours from a glass dispenser into a glass holding water. Phys4D produces flow consistent with gravity and container geometry rather than emitting from the wrong side.

Optics

Spotlight shadow on a rotating turntable

A blue block on a clockwise turntable casts a long shadow on the wall behind it — shadow geometry must track the block’s rotating 3D pose.

Fluid Dynamics

Paper towel released onto liquid

A grabber releases a paper towel onto a shallow dish of blue liquid — a coupled deformable–fluid interaction with wicking and sagging.

Thermodynamics

Folded paper burning

A folded paper burns on a glass cutting board while white smoke rises — a thermodynamic state change with directional buoyant motion.

Solid Mechanics

Tennis ball rolling past a lampshade

A grey tennis ball exits a black tube and rolls rightward across a wood surface, requiring a stable rolling trajectory and consistent depth ordering with the lampshade.

Solid Mechanics

Kettlebell lowered onto pillows

A 30 lb kettlebell and a sheet of green paper are lowered onto two pillows — mass-dependent deformation that appearance-driven models tend to under-express.

Qualitative comparison on Physics-IQ scenarios
Qualitative comparison on Physics-IQ scenarios. WAN2.2-5B baseline (middle) vs. WAN2.2-5B + Phys4D (right) across three physical interaction types. Phys4D produces more consistent object geometry, physically plausible motion, and stable temporal dynamics, compared to the baseline, which exhibits shape distortion and incoherent physical behavior.
Qualitative visualization of 4D interface-level diagnostics
Qualitative visualization of 4D interface-level diagnostics. From left to right: ground-truth 4D point cloud; generated 4D point clouds at 1/4, 1/2, and 3/4 (novel time) of the sequence; and the final frame. This visualization illustrates whether the lifted RGB-D-motion representation maintains coherent geometry and object motion over the generated sequence.

Results

We evaluate Phys4D from four complementary angles: observable video-level physics, 4D interface consistency, preservation of the real-video prior, and which design choices matter.

Does Phys4D improve observable physical plausibility in generated videos? Phys4D improves observable physical plausibility across three backbones and three independent physics diagnostics. Physics-IQ is reported under the matched switch-frame I2V protocol.

BackboneParamsMethod Physics-IQ MSE ↓Physics-IQ Score ↑ VBench2.0-Physics ↑PhyGenBench ↑
CogVideoX5BBase 0.013 ±0.00218.8 ±1.4 0.4849 ±0.0140.45 ±0.02
CogVideoX5B+ Phys4D 0.008 ±0.00132.8 ±1.9 0.5954 ±0.0180.58 ±0.03
WAN2.25BBase 0.016 ±0.00216.8 ±1.3 0.4928 ±0.0150.45 ±0.02
WAN2.25B+ Phys4D 0.013 ±0.00130.9 ±1.8 0.6247 ±0.0170.57 ±0.03
Open-Sora V1.21.1BBase 0.021 ±0.00314.5 ±1.2 0.4623 ±0.0160.44 ±0.02
Open-Sora V1.21.1B+ Phys4D 0.014 ±0.00224.7 ±1.6 0.5899 ±0.0190.57 ±0.03

Consistent gains across all three backbones suggest that simulation-derived physical structure transfers to generated videos beyond simulator-specific metrics.

Does Phys4D improve 4D world-level geometry–motion–trajectory consistency? Held-out simulated scenes with ground-truth depth, flow, and object trajectories, covering per-frame geometry, local geometry–motion consistency, global trajectory evolution, and novel-time continuity.

BackboneMethod AbsRel ↓Warp L1 ↓Flow EPE ↓ 4D Chamfer ↓Mean Drift ↓Fail Rate ↓ Novel Depth ↓Novel Warp ↓
WAN2.2Base + OTF0.39290.79901.25160.50580.536912.38%0.58411.1076
CogVideoXBase + OTF0.34830.70541.23430.49230.523911.34%0.54321.1254
Open-Sora V1.2Base + OTF0.42160.74361.27180.52860.563713.54%0.65341.1573
Phys4D + WAN2.2RGB + OTF0.33470.66981.08240.48730.512610.91%0.55271.0954
Phys4D + WAN2.2Ours 0.23860.45580.4478 0.41590.43128.36% 0.47461.0289

OTF denotes an off-the-shelf diagnostic pipeline that derives depth and motion cues from pretrained estimators (Depth Anything V2 and SEA-RAFT). Phys4D is evaluated through its own learned RGB-D-motion interface. All methods use the same held-out simulation scenes and camera settings.

Does Phys4D preserve the pretrained real-video generative prior? Evaluated with VBench 1.0 under the same prompts, sampling settings, and evaluation pipeline. Visual Avg. is the mean of Aesthetic Quality, Imaging Quality, and Color.

BackboneParamsMethod Aesthetic Quality ↑Imaging Quality ↑Color ↑ Visual Avg. ↑Appearance Style ↑Human Action ↑
CogVideoX5BBase0.61240.69310.90180.73580.21070.7452
CogVideoX5B+ Phys4D0.60970.68860.89950.73260.21840.7728
Open-Sora V1.21.1BBase0.59480.67150.88420.71680.19720.6872
Open-Sora V1.21.1B+ Phys4D0.59170.67820.88210.71730.20110.7023
WAN2.25BBase0.62860.71640.91120.75210.21390.7433
WAN2.25B+ Phys4D0.63510.71190.90870.75190.21620.7712

Phys4D keeps Visual Avg. nearly unchanged across all three backbones while consistently maintaining or improving Appearance Style and Human Action — open-domain visual quality is preserved, and simulation transfers mainly as physical structure rather than simulator appearance.

Which design choices are responsible for physical knowledge injection? All scores are Physics-IQ on WAN2.2-5B under the same matched switch-frame I2V protocol. Rows are grouped by ablation type and should be compared within each group.

Training stages

Base16.8
Stage III only18.4
+ Stage II27.8
+ Stage II + III30.9

Takeaway: local SFT initializes global trajectory alignment — Stage III is most effective when it starts from Stage II.

Auxiliary signals

no depth / flow21.4
depth only24.2
flow only24.8
depth + flow27.8

Takeaway: gains are not explained by auxiliary heads alone — removing depth and flow still improves over the base model.

Noise interval

low-noise22.9
full interval24.6
high-noise27.8

Takeaway: physics is best injected at structure-forming timesteps; high-noise adaptation consistently beats full-interval or low-noise updates.

Loss components

w/o ℒdepth23.8
w/o ℒmotion24.6
w/o ℒwarp25.9
w/o ℒrgb_warp29.4
full30.9

Takeaway: geometry, motion, and warp coupling are complementary — Phys4D relies on coupled geometry–motion learning rather than a single loss term.