Splatting, but the light is sound and the camera is a transducer

Figure: Ge et al., Shanghai Jiao Tong University · Research

Aug 31, 2026, 9:30 PM ETResearch
Photoacoustic tomography fires a laser pulse into tissue, waits for the absorbed light to heat and expand it, and listens to the ultrasound that comes back. Reconstructing where the sound came from requires knowing how fast it travelled — and tissue is not uniform, so the assumption of a single sound speed smears the result. A group at Shanghai Jiao Tong and collaborators have attacked that with machinery borrowed wholesale from Gaussian splatting, in which the camera is a transducer and the spherical harmonics encode acoustic propagation instead of colour.
Why it matters: The speed of sound in tissue varies by scene, and getting it wrong changes acoustic time-of-flight, which defocuses everything. The two existing answers are both awkward: calibrate the acoustic properties in advance, or optimise a dense physical model of the medium, which is expensive and scales badly in three dimensions.
PAGS declines to recover the medium at all. It keeps the initial pressure field as sparse Gaussian sources — the direct analogue of splats — and replaces the medium model with a compact field of what the authors call anisotropic path-averaged sound speed, parameterised by spherical harmonic probes. For a given source and a given transducer direction, that field returns the one number the reconstruction actually needs: the effective speed along that path.
This is a genuinely economical piece of reasoning. The full medium is a hard, high-dimensional inverse problem. The arrival time is not. PAGS solves only for what changes the answer, and the spherical harmonics — which in ordinary splatting encode how colour varies with viewing angle — here encode how sound speed varies with direction.
By the numbers:
- On the simulated dataset PAGS reaches 30.3 dB PSNR, against 29.0 for SlingBAG, the closest Gaussian-based prior method, with RMSE improving from 0.0354 to 0.0305.
- The structural similarity gain is the striking one: 0.537 against 0.331, a jump of 0.206 — considerably larger than the PSNR movement suggests.
- Conventional universal back-projection assuming uniform sound speed manages 21.2 dB and 0.230 SSIM.
- The simulated acquisition uses 4,600 point transducers on a spherical cap of 12.5 mm radius, each recording 1,024 samples at 50 MHz, against a ground-truth volume at 0.1 mm voxels, with a 1,550 m/s ellipsoidal inclusion in a 1,450 m/s background.
- The representation runs to 10⁵–10⁶ Gaussian sources and about 10,000 spherical-harmonic probes carrying nine weights each, optimised for 200 to 500 iterations on an A100.
Yes, but: The headline comparison needs stating precisely. PAGS scores 30.3 dB against 28.2 for a back-projection baseline given the true dual sound speeds, and it is tempting to read that as blind beating oracle. It is not quite that, because two things differ at once: the reconstruction operator and the sound-speed knowledge. What it shows is that a Gaussian representation with a learned propagation field beats back-projection with correct physics — which is a claim about the representation as much as about the blind estimation.
The path-averaged field is also, deliberately, not a medium model. It returns an effective speed per source-transducer path and nothing more, so it will not hand you a sound-speed map of the tissue for any other purpose. That is the trade that makes it cheap, and it does mean the method solves imaging rather than characterisation.
Validation is one simulated dataset and one physical phantom. For a technique aimed at deep-tissue imaging, that is an early-stage evidence base, and the paper presents it as such.
The big picture: The interesting thing here is not the imaging result but the portability. Gaussian splatting arrived as a way to render photographs from novel viewpoints, and the parts being reused in a medical scanner are the parts that were never really about light: an explicit sparse primitive set, an analytic differentiable forward projection, adaptive density control, and spherical harmonics as a cheap basis for something that varies with direction.
Swap the forward model for acoustics and the same optimisation loop closes. That suggests the representation is more general than the rendering problem it was invented for, and that other sensing modalities with a differentiable forward model — anything where you know how a measurement is formed and want to invert it — are candidates for the same treatment.



