SplatsThe evolution of media, in brief
RSS

One handheld video is now enough to rebuild a person in 4D

A bass player stands at the centre of a ring of generated camera views, with a phone at lower left labelled casually captured monocular video and the surrounding frames labelled generated multiview videos

Figure: Jin et al., Zhejiang University / Robbyant / Ant Group / HKUST · Research

Wen Jiang

Wen Jiang

Aug 21, 2026, 9:00 AM ET-Research

Yudong Jin, Tao Xie and colleagues at Zhejiang University's State Key Lab of CAD&CG, with Robbyant, Ant Group and HKUST, have built 4DAnyone, which takes a casually captured monocular video with unknown camera intrinsics and poses, generates the tens of multiview-consistent videos a 4D gaussian splatting reconstruction needs, and lifts them into a model you can orbit.

Why it matters: Free-viewpoint video of people has been a rig problem. DNA-Rendering, the benchmark this paper reports on, was captured with 48 synchronised cameras. That is the standard setup, and it is the reason volumetric human capture stayed inside studios that could afford one.

The alternative has been to generate the missing views with a video diffusion model, which produces footage that looks right and reconstructs badly. The gap between plausible and reconstruction-grade is where this work lives.

The bottleneck they named: The failure has a specific shape, and naming it is most of the contribution. A diffusion transformer can only hold so many target views in one forward pass. Past that, the views get split into groups, and two things break at once: conditioning each group on every previously generated view grows as O(N), which dilutes appearance guidance, and disjoint groups cannot see each other at all, which lets global structure drift apart.

Reference Context Packing compresses the growing set of reference views into a fixed-length mixed-resolution context, taking that first cost to O(1). Target Context Routing handles the second by rotating which views share a group as denoising proceeds — cycling the four-view groups at high noise, when global structure is being decided, then holding them fixed at low noise so neighbouring views can settle detail together.

By the numbers:

  • 24.33 dB PSNR for generated-video consistency on DNA-Rendering, against 21.47 for a fine-tuned ReCamMaster given the same modules and data, 21.25 for MV-Performer and 13.56 for TrajectoryCrafter.
  • 24.15 dB on the 4DGS reconstruction itself, against 20.55 for the strongest baseline — the consistency gains carry through the lift rather than washing out.
  • 23.28 dB on DyMVHumans, which is out of distribution for every method tested.
  • The ablation moves 21.09 → 22.63 dB, with roughly equal credit to each of the two mechanisms and a sliding group schedule beating random and strided.
  • MVGameHuman, the dataset they built for this in an in-house game engine: 38,000 videos, 24 cameras, 318 actors at 2560×1440.

Yes, but: The training bill is not casual even if the capture is. The model was trained on 128 H20-3E GPUs against a purpose-built game-engine dataset plus light-stage and in-the-wild footage. What moved to the phone is inference, not the work behind it.

The evaluation is also narrow: 10 DNA-Rendering sequences and 3 from DyMVHumans, thirteen in total. And the numbers thin out at the edge of the distribution — generated-video reconstruction on DyMVHumans lands at 21.03 dB, several dB below the in-distribution figures, which is the honest measure of how far in-the-wild generalisation actually reaches today.

The big picture: The teaser figure is a person filmed on a phone in one room, surrounded by a ring of generated viewpoints that never existed. The back of her jacket is a hallucination consistent with the front, and the paper says so plainly.

Which is the interesting property and the uncomfortable one. A capture rig records what was there. A system like this decides what was probably there, at a fidelity good enough to reconstruct and re-render, from footage anyone can shoot without the subject noticing the camera was doing anything unusual.

Go deeper:

  • 4DAnyone: Create Anyone in 4D from a Casual Monocular Video on arXiv
  • Project page, video results and source code
⟵ Back to the brief

More stories

Two rows of photoacoustic reconstructions of a branching vascular phantom, shown for SlingBAG, for PAGS, and as the ground-truth digital phantom. The SlingBAG panels carry a mottled noise floor around the vessels; the PAGS panels are cleaner, with the vessel network closer to the crisp white tracery of the phantom

Splatting, but the light is sound and the camera is a transducer

5 mins ago

A schematic of a scene divided into a wireframe grid of cells against black. One cell is outlined in yellow and holds a sharp green cylinder; a blurred blue slab sits behind it and a red slab in front, standing in for the frozen regions flattened into single background and foreground images

Their trick makes VRAM independent of scene size. The test scenes were too small to show it.

6 hours ago

A fairground drop-tower ride rendered twice: on the left from a degraded reconstruction, where the tower and foliage dissolve into white streaks and smears, and on the right after refinement, sharp and photographic against a clear sky

Give it scattered keypoints and it matches a full splat reconstruction

3 days ago

Six holographic reconstructions of laboratory equipment photographed against black through a HoloLens: a Bunsen burner and a rack of capped test tubes above, and below them a shredded, torn reconstruction of a mortar and pestle beside two further mortar-and-pestle models whose pestles are visibly deformed

PSNR said Gaussian splatting won. Seventeen people said it didn't.

3 days ago

splats

Short daily briefs on the evolution of media — gaussian splats, volumetric video, dome theaters, headsets, and the research underneath.

Newsroom

  • Latest
  • All stories
  • RSS feed
© 2026 Splats · Terms · Privacy