SplatsThe evolution of media, in brief
RSS

Radiance fields stop being pictures and start being places

The same orchard branch twice: on the left the radiance field rendered photorealistically, on the right the semantic field, with apples picked out in red and foliage in green

Figure: Heider et al. · Research

Priya Raghunathan

Priya Raghunathan

Aug 15, 2026, 7:00 AM ET-Research

Nico Heider, Michał Jan Włodarczyk and colleagues argue that Semantic Radiance Fields — captures that carry per-class meaning alongside colour and geometry — can serve as simulators for training embodied agents, closing the gap between synthetic worlds and real ones.

Why it matters: Simulators for robots come in two flavours, and both are compromised. Synthetic environments know exactly what every object is, because someone authored them, but they don't look like the world. Reconstructions of real places look right and know nothing — a splat of your kitchen has no idea which blob is a kettle.

The proposal is to make one representation do both jobs: keep the photometric fidelity of a capture, and attach to every gaussian a label saying what it is part of. A capture stops being footage and becomes a place an agent can be tested in.

How it works: Segmentations from pretrained 2D vision models are lifted into the 3D field, so geometry, appearance, and per-class identity are jointly encoded in one representation reconstructed from ordinary posed RGB captures.

Each class gets its own independent binary head rather than a single softmax across classes, so a point can belong to more than one thing at once — a design choice that keeps geometry from collapsing onto class boundaries.

Zoom in:

  • The field exposes three operations an agent can call: render a posed RGB image from a camera pose, query the per-class probability at any 3D point, and query occupancy for collision detection.
  • That trio covers viewpoint sampling, object localisation, collision checking, and per-pixel ground truth from a single grounded representation.
  • Semantic ground truth comes free — it is baked into the capture rather than hand-labelled per scene.
  • The worked example is an orchard: a robot reaching for an apple, with the radiance field supplying rendering, semantics, and occupancy to a physics engine.

Yes, but: This is a position paper with an outlined application, not a benchmarked system. The apple-reaching simulator is described as an example rather than evaluated against existing robot simulators.

The semantics are only as good as the 2D model that produced them — the field inherits whatever the segmentation backbone gets wrong, now baked into three dimensions.

The big picture: Captured media is quietly becoming infrastructure. The same orchard capture is a picture, a map, a collision mesh, and a labelled dataset, depending on which query you send it.

That is a more interesting destination for this technology than better-looking flythroughs.

Go deeper:

  • Semantic Radiance Fields as Simulators for Spatial Reasoning in Real-World Scenes on arXiv
⟵ Back to the brief

More stories

The SuperSplat 3.3.0 editor with a Gaussian splat capture of a street cafe loaded — furled yellow umbrellas over metal tables, parked cars and a tree-lined street behind. The scene manager and transform panels sit at the left, the tool strip along the bottom, and the status bar reads two million splats

SuperSplat rewrote itself on WebGPU and deleted the fallback

Today

Two rows of photoacoustic reconstructions of a branching vascular phantom, shown for SlingBAG, for PAGS, and as the ground-truth digital phantom. The SlingBAG panels carry a mottled noise floor around the vessels; the PAGS panels are cleaner, with the vessel network closer to the crisp white tracery of the phantom

Splatting, but the light is sound and the camera is a transducer

Sep 1, 2026

A schematic of a scene divided into a wireframe grid of cells against black. One cell is outlined in yellow and holds a sharp green cylinder; a blurred blue slab sits behind it and a red slab in front, standing in for the frozen regions flattened into single background and foreground images

Their trick makes VRAM independent of scene size. The test scenes were too small to show it.

Aug 31, 2026

A fairground drop-tower ride rendered twice: on the left from a degraded reconstruction, where the tower and foliage dissolve into white streaks and smears, and on the right after refinement, sharp and photographic against a clear sky

Give it scattered keypoints and it matches a full splat reconstruction

Aug 29, 2026

splats

Short daily briefs on the evolution of media — gaussian splats, volumetric video, dome theaters, headsets, and the research underneath.

Newsroom

  • Latest
  • All stories
  • RSS feed
© 2026 Splats · Terms · Privacy