SplatsThe evolution of media, in brief
RSS

Your phone records three viewpoints per shot. The pipeline throws two away.

Three columns comparing a reconstruction of a hand moving across a carpet — ground truth, a monocular reconstruction in which the hand dissolves into a vertical smear, and an iPhone multi-camera reconstruction in which the closed fist is legible — each shown as a wide view above a zoomed crop

Figure: Li et al., Cornell / Berkeley / Adobe / Georgia Tech (CC BY 4.0) · Research

Wen Jiang

Wen Jiang

Aug 26, 2026, 9:15 PM ET-Research

The three rear cameras on an iPhone see the same scene from viewpoints about five degrees apart, simultaneously, every time the shutter fires. An Apple Vision Pro's stereo pair sits roughly fifteen degrees apart. A Lytro Illum records a thirteen-by-thirteen grid of them at once. Shamus Li and colleagues point out that essentially every reconstruction pipeline takes one of those streams and discards the rest, then spends considerable effort inventing the parallax it just threw out.

Why it matters: The received wisdom is that consumer camera baselines are too small to matter — five degrees of separation between phone lenses is nothing next to walking around an object. So the field standardised on a moving monocular camera, and when the camera cannot move enough, on learned priors that hallucinate the missing views.

That reasoning holds only when the camera can move. It fails completely for a single exposure, and it fails for anything in motion: a stationary monocular camera watching a moving hand has no angular diversity at all, and no amount of prior fixes a measurement that was never taken.

The paper separates two regimes that usually get conflated. Sensor-limited multi-view is one sensor trading spatial resolution for angular resolution — a light field camera, where more viewpoints means fewer pixels each. Exposure-limited multi-view is several sensors on one device capturing the same instant, where the extra views are free and the only question is whether the baseline is wide enough to be worth anything.

By the numbers:

  • On real single-exposure static scenes with a Lytro Illum: monocular 3DGS scores 26.24 dB and monocular SparseGS 27.75 dB. Feeding the same reconstruction all 81 sub-views takes plain 3DGS to 30.50 dB — past the specialist sparse-view method by 2.75 dB.
  • With an Apple Vision Pro's stereo pair, single exposure: 17.57 dB monocular against 20.89 dB using both cameras, and 21.57 dB with FSGS on top.
  • Depth is where the gain is starkest. From one exposure, mean absolute relative depth error falls from 0.0939 monocular to 0.0526 with stereo, and RMSE from 0.4588 to 0.3038 — roughly half the error, from hardware that was already in the device.
  • On synthetic scenes at a ten-degree baseline, one exposure: 16.25 dB monocular, 24.78 dB with a light field camera, 26.39 dB with the multiplexed prototype the authors built.
  • The dataset is the contribution that will outlast the analysis: static and dynamic real captures of each scene through an iPhone 15 Pro, an Apple Vision Pro and a Lytro Illum, which is three quite different angular sampling patterns of the same content.

Yes, but: The result does not generalise to the case most people are actually in, and the authors say so rather than letting the reader find out. On casual handheld video — someone walking around a scene with a phone — the extra cameras buy almost nothing: 21.75 dB monocular against 22.12 dB using all three iPhone lenses, and 27.89 against 28.44 for stereo.

Worse for the thesis, adding a diffusion prior to the monocular stream beats the multi-view capture outright. Monocular with Difix3D+ scores 23.67 dB where the iPhone's three cameras with the same prior score 23.25. The paper's own figure caption explains it plainly: camera motion already supplies angular diversity, so either added views or the prior can repair the ghosting, and you do not need both.

So the honest summary is narrower than the title suggests. Multi-camera capture matters when monocular capture is most constrained — one shot, or a fixed camera on a moving subject — and stops mattering once you are walking around. The authors state exactly this in their closing paragraph, which is more than many papers manage about their own headline claim.

There is also a measurement caveat worth reading the tables carefully for: single-view and three-view methods are scored against different camera sets, so the columns within Table 1 are not all directly comparable to each other.

The big picture: The interesting comparison is not multi-view against monocular; it is multi-view against the enormous research effort spent compensating for monocular. Sparse-view methods, regularisers and diffusion priors all exist because angular coverage is scarce. Here is a paper observing that on a large fraction of consumer hardware it is not scarce, merely discarded at the file format.

For anyone capturing dynamic scenes the practical instruction is unambiguous, because it is the one regime where the alternative does not exist. A fixed monocular camera cannot recover motion and geometry simultaneously — the hand smears — and the second lens is not an improvement but a precondition.

Go deeper:

  • Sparse Light Field Sampling Improves Casual 3D and 4D Reconstruction on arXiv
  • Project page
⟵ Back to the brief

More stories

Two rows of photoacoustic reconstructions of a branching vascular phantom, shown for SlingBAG, for PAGS, and as the ground-truth digital phantom. The SlingBAG panels carry a mottled noise floor around the vessels; the PAGS panels are cleaner, with the vessel network closer to the crisp white tracery of the phantom

Splatting, but the light is sound and the camera is a transducer

6 mins ago

A schematic of a scene divided into a wireframe grid of cells against black. One cell is outlined in yellow and holds a sharp green cylinder; a blurred blue slab sits behind it and a red slab in front, standing in for the frozen regions flattened into single background and foreground images

Their trick makes VRAM independent of scene size. The test scenes were too small to show it.

6 hours ago

A fairground drop-tower ride rendered twice: on the left from a degraded reconstruction, where the tower and foliage dissolve into white streaks and smears, and on the right after refinement, sharp and photographic against a clear sky

Give it scattered keypoints and it matches a full splat reconstruction

3 days ago

Six holographic reconstructions of laboratory equipment photographed against black through a HoloLens: a Bunsen burner and a rack of capped test tubes above, and below them a shredded, torn reconstruction of a mortar and pestle beside two further mortar-and-pestle models whose pestles are visibly deformed

PSNR said Gaussian splatting won. Seventeen people said it didn't.

3 days ago

splats

Short daily briefs on the evolution of media — gaussian splats, volumetric video, dome theaters, headsets, and the research underneath.

Newsroom

  • Latest
  • All stories
  • RSS feed
© 2026 Splats · Terms · Privacy