SplatsThe evolution of media, in brief
RSS

Pull an object out of a splat you didn't capture

A query photo of a red handbag among clutter, beside a baseline extraction where the bag comes out surrounded by smeared background, beside Seed2GS's clean isolated bag on black

Figure: Ding et al., Chinese Academy of Sciences · Research

Yusuf Demirci

Yusuf Demirci

Aug 16, 2026, 7:20 AM ET-Research

Zongjian Ding and colleagues at the Chinese Academy of Sciences, HKUST, Zhejiang University and the Beijing Institute of Technology report the highest published LERF-MASK accuracy for object extraction from a pre-built splat scene — with the scene frozen and no access to the cameras that built it.

Why it matters: Most 3D editing workflows receive a finished splat, not a capture session. The source images and reconstruction cameras are somebody else's, from months ago, and were never shipped with the asset.

Nearly every existing extraction method assumes otherwise — it wants the original cameras back, or it wants to train a per-scene representation for tens of minutes before you can ask it a question.

The insight: The authors separate two things that earlier methods tangle together: identity — which object do you mean — and coverage — where does it extend in 3D.

Identity is fixed once, from a single reliable reference mask chosen among open-vocabulary candidates. Coverage is then built up by lifting that seed and orbiting virtual viewpoints around it, with tracking carrying the seed forward instead of re-detecting the object in every view.

By the numbers:

  • 92.1% mean IoU on LERF-MASK, 3.7 points above the strongest scene-trained baseline and 7.6 above the closest camera-free one.
  • 9.3 seconds of measured compute-only latency, against tens of minutes for scene-trained methods.
  • 95.7% mIoU on 3D-OVS.
  • With one fixed reference view per scene the full pipeline still holds 91.1% mIoU.
  • Swapping the predicted seed for a ground-truth mask gains only 0.72 points — the seed selection is close to the ceiling.

The tell: In the comparison figure the baselines don't fail by missing the object — they smear. Pull out a red bag and you get the bag plus a haze of background gaussians dragged along with it.

That difference matters more than the mIoU gap for anyone actually editing: a clean extraction is an asset, a smeared one is a cleanup job.

What's next: The scene stays frozen throughout — masks supervise a single temporary foreground value per gaussian rather than modifying the representation, which is what makes the query cheap and repeatable.

Nine seconds is the number to watch. It is the difference between an offline batch process and something that can sit behind a click in an editor.

Go deeper:

  • Seed2GS: Camera-Free, Training-Free Object Extraction from 3D Gaussian Scenes on arXiv
⟵ Back to the brief

More stories

The SuperSplat 3.3.0 editor with a Gaussian splat capture of a street cafe loaded — furled yellow umbrellas over metal tables, parked cars and a tree-lined street behind. The scene manager and transform panels sit at the left, the tool strip along the bottom, and the status bar reads two million splats

SuperSplat rewrote itself on WebGPU and deleted the fallback

Today

Two rows of photoacoustic reconstructions of a branching vascular phantom, shown for SlingBAG, for PAGS, and as the ground-truth digital phantom. The SlingBAG panels carry a mottled noise floor around the vessels; the PAGS panels are cleaner, with the vessel network closer to the crisp white tracery of the phantom

Splatting, but the light is sound and the camera is a transducer

Sep 1, 2026

A schematic of a scene divided into a wireframe grid of cells against black. One cell is outlined in yellow and holds a sharp green cylinder; a blurred blue slab sits behind it and a red slab in front, standing in for the frozen regions flattened into single background and foreground images

Their trick makes VRAM independent of scene size. The test scenes were too small to show it.

Aug 31, 2026

A fairground drop-tower ride rendered twice: on the left from a degraded reconstruction, where the tower and foliage dissolve into white streaks and smears, and on the right after refinement, sharp and photographic against a clear sky

Give it scattered keypoints and it matches a full splat reconstruction

Aug 29, 2026

splats

Short daily briefs on the evolution of media — gaussian splats, volumetric video, dome theaters, headsets, and the research underneath.

Newsroom

  • Latest
  • All stories
  • RSS feed
© 2026 Splats · Terms · Privacy