SplatsThe evolution of media, in brief
RSS

Your phone records three viewpoints per shot. The pipeline throws two away.

Three columns comparing a reconstruction of a hand moving across a carpet — ground truth, a monocular reconstruction in which the hand dissolves into a vertical smear, and an iPhone multi-camera reconstruction in which the closed fist is legible — each shown as a wide view above a zoomed crop

Figure: Li et al., Cornell / Berkeley / Adobe / Georgia Tech (CC BY 4.0) · Research

Wen Jiang

Wen Jiang

Aug 26, 2026, 9:15 PM ET-Research

The three rear cameras on an iPhone see the same scene from viewpoints about five degrees apart, simultaneously, every time the shutter fires. An Apple Vision Pro's stereo pair sits roughly fifteen degrees apart. A Lytro Illum records a thirteen-by-thirteen grid of them at once. Shamus Li and colleagues point out that essentially every reconstruction pipeline takes one of those streams and discards the rest, then spends considerable effort inventing the parallax it just threw out.

Why it matters: The received wisdom is that consumer camera baselines are too small to matter — five degrees of separation between phone lenses is nothing next to walking around an object. So the field standardised on a moving monocular camera, and when the camera cannot move enough, on learned priors that hallucinate the missing views.

That reasoning holds only when the camera can move. It fails completely for a single exposure, and it fails for anything in motion: a stationary monocular camera watching a moving hand has no angular diversity at all, and no amount of prior fixes a measurement that was never taken.

The paper separates two regimes that usually get conflated. Sensor-limited multi-view is one sensor trading spatial resolution for angular resolution — a light field camera, where more viewpoints means fewer pixels each. Exposure-limited multi-view is several sensors on one device capturing the same instant, where the extra views are free and the only question is whether the baseline is wide enough to be worth anything.

By the numbers:

  • On real single-exposure static scenes with a Lytro Illum: monocular 3DGS scores 26.24 dB and monocular SparseGS 27.75 dB. Feeding the same reconstruction all 81 sub-views takes plain 3DGS to 30.50 dB — past the specialist sparse-view method by 2.75 dB.
  • With an Apple Vision Pro's stereo pair, single exposure: 17.57 dB monocular against 20.89 dB using both cameras, and 21.57 dB with FSGS on top.
  • Depth is where the gain is starkest. From one exposure, mean absolute relative depth error falls from 0.0939 monocular to 0.0526 with stereo, and RMSE from 0.4588 to 0.3038 — roughly half the error, from hardware that was already in the device.
  • On synthetic scenes at a ten-degree baseline, one exposure: 16.25 dB monocular, 24.78 dB with a light field camera, 26.39 dB with the multiplexed prototype the authors built.
  • The dataset is the contribution that will outlast the analysis: static and dynamic real captures of each scene through an iPhone 15 Pro, an Apple Vision Pro and a Lytro Illum, which is three quite different angular sampling patterns of the same content.

Yes, but: The result does not generalise to the case most people are actually in, and the authors say so rather than letting the reader find out. On casual handheld video — someone walking around a scene with a phone — the extra cameras buy almost nothing: 21.75 dB monocular against 22.12 dB using all three iPhone lenses, and 27.89 against 28.44 for stereo.

Worse for the thesis, adding a diffusion prior to the monocular stream beats the multi-view capture outright. Monocular with Difix3D+ scores 23.67 dB where the iPhone's three cameras with the same prior score 23.25. The paper's own figure caption explains it plainly: camera motion already supplies angular diversity, so either added views or the prior can repair the ghosting, and you do not need both.

So the honest summary is narrower than the title suggests. Multi-camera capture matters when monocular capture is most constrained — one shot, or a fixed camera on a moving subject — and stops mattering once you are walking around. The authors state exactly this in their closing paragraph, which is more than many papers manage about their own headline claim.

There is also a measurement caveat worth reading the tables carefully for: single-view and three-view methods are scored against different camera sets, so the columns within Table 1 are not all directly comparable to each other.

The big picture: The interesting comparison is not multi-view against monocular; it is multi-view against the enormous research effort spent compensating for monocular. Sparse-view methods, regularisers and diffusion priors all exist because angular coverage is scarce. Here is a paper observing that on a large fraction of consumer hardware it is not scarce, merely discarded at the file format.

For anyone capturing dynamic scenes the practical instruction is unambiguous, because it is the one regime where the alternative does not exist. A fixed monocular camera cannot recover motion and geometry simultaneously — the hand smears — and the second lens is not an improvement but a precondition.

Go deeper:

  • Sparse Light Field Sampling Improves Casual 3D and 4D Reconstruction on arXiv
  • Project page
⟵ Back to the brief

More stories

A figure summarising the study: five systems and four water regimes across the top, a row of murky underwater renderings of a submerged structure, and beneath them two Gaussian point clouds of a sunken car — one coherent and car-shaped labelled 3DGS, one scattered and diffuse labelled SeaSplat, each captioned with its PSNR and chamfer error

The water got murkier and the scores went up

Today

The Tanks and Temples Truck scene — a pale blue vintage flatbed pickup parked on a pavement — rendered sharply inside the Splat.js browser interface, with a readout showing 579,748 splats and a Train button in the toolbar

Arrival.Space gave away the browser version of what it sells

6 hours ago

A soft, hazy Gaussian splat render of San Francisco seen from Telegraph Hill, with Coit Tower rising in the centre of the frame and the downtown skyline behind it

The framework under deck.gl just shipped a Gaussian splat renderer

10 hours ago

Two rows of underwater reef photographs, each showing a raw teal-cast frame, a SeaSplat restoration and NemoSplat's restoration, in which the water's colour cast lifts and pink and orange coral becomes visible

Where the light fails, optical flow still beats the foundation models

13 hours ago

splats

Short daily briefs on the evolution of media — gaussian splats, volumetric video, dome theaters, headsets, and the research underneath.

Newsroom

  • Latest
  • All stories
  • RSS feed
© 2026 Splats · Terms · Privacy