SplatsThe evolution of media, in brief
RSS

Humanoid robots keep falling over inside gaussian splats

Two reconstructed interiors shown as a photographic render beside three surface-normal visualisations, the first a noisy scribble of colour and the last resolving into clean flat walls and furniture

Figure: Pham et al., VinMotion / University of Southern California · Research

Priya Raghunathan

Priya Raghunathan

Aug 17, 2026, 7:30 AM ET-Research

Quan-Dung Pham and colleagues at VinMotion and the University of Southern California have built HumanoidVLN, a physics-grounded simulator whose environments are drawn partly from gaussian splatting reconstructions of real interiors — and whose headline result is how badly current navigation models cope with having legs.

Why it matters: Vision-language navigation benchmarks have mostly assumed a wheeled robot gliding along a floor plan. A bipedal robot has to stay upright, its camera pitches and rolls with every step, and no two humanoid platforms are shaped alike.

The test environments are real interiors, captured as gaussian splats and filtered down to those with more than 100 square metres of navigable floor. The capture is not the demonstration — it is the ground the robot has to walk on, and it has to hold up under a physics engine rather than a camera.

The result: Across four models and four robots, the best combination — JanusVLN — reaches a mean success rate of 43.55%. That is the ceiling, not the average.

The fall rates are more striking than the success rates. Unitree's H1 goes over in 27% of episodes under the steadiest model and 71% under the least steady, while the other three platforms stay between roughly 2.7% and 10%. Same instructions, same scenes — the body decides the outcome.

Zoom in:

  • Built on NVIDIA Isaac Sim, with four humanoids spanning 10–12 lower-body degrees of freedom and heights from 1.17m to 1.80m.
  • Control is a reinforcement-learning locomotion policy under interchangeable PD or MPC path trackers, so the controller can be varied independently of the navigation model.
  • 933 collision-aware reference episodes across 87 scenes and 17 indoor classes, each with one precise instruction plus formal, natural and casual paraphrases.
  • A 20-episode sim-to-real pilot on a real Unitree G1 found navigation error in simulation correlated with reality at r=0.935, mean absolute difference 0.68m.

The geometry problem: Buried in an ablation is a finding worth pulling out: vanilla gsplat produces noisy, discontinuous surface normals, particularly around textureless walls, furniture boundaries and thin structures — which makes the resulting collision meshes unusable for physics.

Their pipeline adds depth-normal consistency and unbiased depth rendering to get surfaces coherent enough to fuse into a mesh a simulator will accept. Pretty renders were never the bottleneck; walkable geometry was.

Yes, but: The sim-to-real pilot covers two scenes and one model checkpoint. The authors are explicit that it is episode-level evidence, not a claim that reconstructed scenes generalise.

They also name human verification as a scaling bottleneck — every instruction in the benchmark passed through a human, which is what makes 933 episodes a respectable number rather than a small one.

Go deeper:

  • HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation on arXiv
  • Project page
⟵ Back to the brief

More stories

Two rows of photoacoustic reconstructions of a branching vascular phantom, shown for SlingBAG, for PAGS, and as the ground-truth digital phantom. The SlingBAG panels carry a mottled noise floor around the vessels; the PAGS panels are cleaner, with the vessel network closer to the crisp white tracery of the phantom

Splatting, but the light is sound and the camera is a transducer

5 mins ago

A schematic of a scene divided into a wireframe grid of cells against black. One cell is outlined in yellow and holds a sharp green cylinder; a blurred blue slab sits behind it and a red slab in front, standing in for the frozen regions flattened into single background and foreground images

Their trick makes VRAM independent of scene size. The test scenes were too small to show it.

6 hours ago

A fairground drop-tower ride rendered twice: on the left from a degraded reconstruction, where the tower and foliage dissolve into white streaks and smears, and on the right after refinement, sharp and photographic against a clear sky

Give it scattered keypoints and it matches a full splat reconstruction

3 days ago

Six holographic reconstructions of laboratory equipment photographed against black through a HoloLens: a Bunsen burner and a rack of capped test tubes above, and below them a shredded, torn reconstruction of a mortar and pestle beside two further mortar-and-pestle models whose pestles are visibly deformed

PSNR said Gaussian splatting won. Seventeen people said it didn't.

3 days ago

splats

Short daily briefs on the evolution of media — gaussian splats, volumetric video, dome theaters, headsets, and the research underneath.

Newsroom

  • Latest
  • All stories
  • RSS feed
© 2026 Splats · Terms · Privacy