Talk at Télécom Paris: From Reconstruction to Reasoning, Structured 3D Representations

Slides for my talk at Télécom Paris (Institut Polytechnique de Paris) on 21 September 2026. [slides, HTML] [slides, PDF] [appendix, PDF]

The HTML version includes the video demonstrations. Use the arrow keys to advance, O for an overview of all slides, and S for the speaker view with notes.

Abstract

How can visual observations become representations that we can understand, manipulate and improve through feedback? A photorealistic reconstruction is not enough on its own: each operation we want to perform needs a handle in the representation, and each claim about a method needs a measurement that can fail. The talk follows this question through three parts.

Reliable reconstruction. FourieRF (3DV 2025) controls the spatial frequencies a tensor radiance field may learn, so that a few views constrain the geometry rather than only the images. MILo (ACM TOG, SIGGRAPH Asia 2025) brings an explicit mesh into Gaussian Splatting optimization and produces compact, usable surfaces. From Blobs to Spokes (ECCV 2026) gives each Gaussian an oriented surface meaning, recovers structures as thin as bicycle spokes, and shows that the standard surface benchmark rewards dense meshes, proposing two corrected protocols.

Semantic structure. ZeroKey (ICCV 2025) names 3D points through language and a point-level multimodal model, without 3D training. PatchAlign3D (CVPR 2026) learns language-aligned local features directly on point clouds, in a single feed-forward pass.

Feedback and reasoning. Is It a Good Match? (under review) asks whether a vision-language model can judge a proposed correspondence and, through that judgment alone, adapt a visual model that then runs without it. 3DHarnessBench (preprint, with Télécom Paris) measures what agents can reconstruct as executable Blender programs when the evidence they may access changes.

The closing part outlines the research programme From Video to Animatable and Efficient 3D Worlds: learn representations and operations together, from images and video, with checks we can trust, starting from verified trajectories for compact models that reconstruct and edit articulated objects.