
10 hours agoResearch
A driving world model that only predicts the parts that move

World action models — the systems that predict what a scene will look like a few frames after an agent acts — almost all work in 2D video, which means redrawing every pixel of every frame, including the buildings that have not moved since the drive began. Yueen Ma and colleagues at the Chinese University of Hong Kong, Fudan, and the Shanghai Academy of AI for Science argue that this is a strange way to spend a prediction budget, and propose keeping the background.
Why it matters: A video world model treats the future as an image-generation problem. It has no explicit notion that the parked car on the left is an object, that the object has a pose, or that the terrace behind it was fully observed ninety frames ago and has not changed since. Every frame is regenerated from scratch, and the model spends most of its capacity reproducing content it already had.












