EDBT 2026 Demo / reviewers in the wild / expert
Tanner Schmidt
dblp:164/8399
· DBLP profile ↗
8ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-5708-1257ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
3D vision · 68% Segmentation and scene understanding · 11% Efficient and distributed learning · 11% | |
| Computer graphics and multimedia
3 papers |
Computer animation and physical simulation · 50% Geometric modeling and processing · 38% Rendering · 13% |
Topics — the 20 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
neural radiance field |
1.1 | 2 | 2022 | Neural 3D Video Synthesis from Multi-view Video · CVPR 2022 STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.9 | 1 | 2025 | Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation · CVPR 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation · CVPR 2025 |
Computer vision › 3D vision › neural radiance field
dynamic neural radiance field |
0.6 | 1 | 2022 | Neural 3D Video Synthesis from Multi-view Video · CVPR 2022 |
Computer vision › 3D vision › novel view synthesis
multi-view video generation |
0.6 | 1 | 2022 | Neural 3D Video Synthesis from Multi-view Video · CVPR 2022 |
Computer vision › 3D vision
novel view synthesis |
0.6 | 1 | 2022 | Neural 3D Video Synthesis from Multi-view Video · CVPR 2022 |
Computer animation and physical simulation › motion synthesis
motion interpolation |
0.6 | 1 | 2022 | Neural 3D Video Synthesis from Multi-view Video · CVPR 2022 |
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction |
0.5 | 1 | 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021 |
Computer vision › 3D vision › object pose estimation
object pose tracking |
0.5 | 1 | 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021 |
Computer vision › 3D vision › motion estimation
rigid motion estimation |
0.5 | 1 | 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021 |
Computer vision › 3D vision
3d reconstruction |
0.4 | 1 | 2020 | Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction · ECCV (29) 2020 |
Computer vision › 3D vision › pose estimation
shape and pose estimation |
0.4 | 1 | 2020 | FroDO: From Detections to 3D Objects · CVPR 2020 |
Computer vision › 3D vision › 3d shape representation › implicit surface representation
signed distance function |
0.4 | 1 | 2020 | Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction · ECCV (29) 2020 |
Geometric modeling and processing
3d reconstruction |
0.4 | 1 | 2020 | FroDO: From Detections to 3D Objects · CVPR 2020 |
Computer vision › Video understanding and tracking › object tracking
articulated object tracking |
0.2 | 1 | 2015 | Depth-based tracking with physical constraints for robot manipulation · ICRA 2015 |
Robotics › Robot manipulation
grasping |
0.2 | 1 | 2015 | Depth-based tracking with physical constraints for robot manipulation · ICRA 2015 |
Robotics › Robot manipulation › grasping
grasp planning |
0.2 | 1 | 2015 | Depth-based tracking with physical constraints for robot manipulation · ICRA 2015 |
Computer vision › Video understanding and tracking
object tracking |
0.2 | 1 | 2015 | Depth-based tracking with physical constraints for robot manipulation · ICRA 2015 |
Rendering
novel view synthesis |
0.1 | 1 | 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021 |
Human-robot interaction
shared control |
0.1 | 1 | 2015 | Depth-based tracking with physical constraints for robot manipulation · ICRA 2015 |
Methods — techniques the papers use, named apart from their topics
time-conditioned neural radiance field · 1.1ray importance sampling · 1.1hierarchical training · 1.1volume rendering · 1.0self-supervised learning · 1.0joint optimization · 1.0variable-resolution patch tokenization · 0.9multi-view optimization · 0.9encoder network · 0.9deep signed distance function · 0.9torque sensing · 0.2physical constraint modeling · 0.2depth-based tracking · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Segment This Thing: Foveated Tokenization for Efficient Point-Prompted SegmentationabstractThis paper presents Segment This Thing (STT), a new efficient image segmentation model designed to produce a single segment given a single point prompt. Instead of following prior work and increasing efficiency by decreasing model size, we gain efficiency by foveating input images. Given an image and a point prompt, we extract a crop centered on the prompt and apply a novel variable-resolution patch tokenization in which patches are downsampled at a rate that increases with increased distance from the prompt. This approach yields far fewer image tokens than uniform patch tokenization. As a result we can drastically reduce the computational cost of segmentation without reducing model size. Furthermore, the foveation focuses the model on the region of interest, a potentially useful inductive bias. We show that our Segment This Thing model is more efficient than prior work while remaining competitive on segmentation benchmarks. It can easily run at interactive frame rates on consumer hardware and is thus a promising tool for augmented reality or robotics applications. Tanner Schmidt, Richard A. Newcombe |
CVPR | 1 |
| 2022 | Neural 3D Video Synthesis from Multi-view VideoabstractWe propose a novel approach for 3D video synthesis that is able to represent multi-view video recordings of a dynamic real-world scene in a compact, yet expressive representation that enables high-quality view synthesis and motion interpolation. Our approach takes the high quality and compactness of static neural radiance fields in a new direction: to a model-free, dynamic setting. At the core of our approach is a novel time-conditioned neural radiance field that represents scene dynamics using a set of compact latent codes. We are able to significantly boost the training speed and perceptual quality of the generated imagery by a novel hierarchical training scheme in combination with ray importance sampling. Our learned representation is highly compact and able to represent a 10 second 30 FPS multi-view video recording by 18 cameras with a model size of only 28MB. We demonstrate that our method can render high-fidelity wide-angle novel views at over 1K resolution, even for complex and dynamic scenes. We perform an extensive qualitative and quantitative evaluation that shows that our approach outperforms the state of the art. Project website: https://neural-3d-video.github.io/. Tianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green, Christoph Lassner, Changil Kim 0001, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard A. Newcombe, Zhaoyang Lv |
CVPR | 7 |
| 2021 | STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural RenderingabstractWe present STaR, a novel method that performs Self-supervised Tracking and Reconstruction of dynamic scenes with rigid motion from multi-view RGB videos without any manual annotation. Recent work has shown that neural networks are surprisingly effective at the task of compressing many views of a scene into a learned function which maps from a viewing ray to an observed radiance value via volume rendering. Unfortunately, these methods lose all their predictive power once any object in the scene has moved. In this work, we explicitly model rigid motion of objects in the context of neural representations of radiance fields. We show that without any additional human specified supervision, we can reconstruct a dynamic scene with a single rigid object in motion by simultaneously decomposing it into its two constituent parts and encoding each with its own neural representation. We achieve this by jointly optimizing the parameters of two neural radiance fields and a set of rigid poses which align the two fields at each frame. On both synthetic and real world datasets, we demonstrate that our method can render photorealistic novel views, where novelty is measured on both spatial and temporal axes. Our factored representation furthermore enables animation of unseen object motion. Zhaoyang Lv, Tanner Schmidt, Steven Lovegrove |
CVPR | 3 |
| 2020 | FroDO: From Detections to 3D ObjectsabstractObject-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction of object instances from RGB video that infers their location, pose and shape in a coarse to fine manner. Key to FroDO is to embed object shapes in a novel learnt shape space that allows seamless switching between sparse point cloud and dense DeepSDF decoding. Given an input sequence of localized RGB frames, FroDO first aggregates 2D detections to instantiate a 3D bounding box per object. A shape code is regressed using an encoder network before optimizing shape and pose further under the learnt shape priors using sparse or dense shape representations. The optimization uses multi-view geometric, photometric and silhouette losses. We evaluate on real-world datasets, including Pix3D, Redwood-OS, and ScanNet, for single-view, multi-view, and multi-object reconstruction. Martin Rünz, Kejie Li, Meng Tang 0001, Lingni Ma, Chen Kong, Tanner Schmidt, Ian D. Reid 0001, Lourdes Agapito, Julian Straub, Steven Lovegrove, Richard A. Newcombe |
CVPR | 6 |
| 2020 | Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction
Rohan Chabra, Jan Eric Lenssen, Eddy Ilg, Tanner Schmidt, Julian Straub, Steven Lovegrove, Richard A. Newcombe |
ECCV (29) | 4 |
| 2017 | Dynamic High Resolution Deformable Articulated TrackingabstractThe last several years have seen significant progress in using depth cameras for tracking articulated objects such as human bodies, hands, and robotic manipulators. Most approaches focus on tracking skeletal parameters of a fixed shape model, which makes them insufficient for applications that require accurate estimates of deformable object surfaces. To overcome this limitation, we present a 3D model-based tracking system for articulated deformable objects. Our system is able to track human body pose and high resolution surface contours in real time using a commodity depth sensor and GPU hardware. We implement this as a joint optimization over a skeleton to account for changes in pose, and over the vertices of a high resolution mesh to track the subject's shape. Through experimental results we show that we are able to capture dynamic sub-centimeter surface detail such as folds and wrinkles in clothing. We also show that this shape estimation aids kinematic pose estimation by providing a more accurate target to match against the point cloud. The end result is highly accurate spatiotemporal and semantic information which is well suited for physical human robot interaction as well as virtual and augmented reality systems. Aaron Walsman, Weilin Wan 0001, Tanner Schmidt, Dieter Fox |
3DV | 3 |
| 2017 | Self-directed Lifelong Learning for Robot Vision
Tanner Schmidt, Dieter Fox |
ISRR | 1 |
| 2015 | Depth-based tracking with physical constraints for robot manipulationabstractThis work integrates visual and physical constraints to perform real-time depth-only tracking of articulated objects, with a focus on tracking a robot's manipulators and manipulation targets in realistic scenarios. As such, we extend DART, an existing visual articulated object tracker, to additionally avoid interpenetration of multiple interacting objects, and to make use of contact information collected via torque sensors or touch sensors. To achieve greater stability, the tracker uses a switching model to detect when an object is stationary relative to the table or relative to the palm and then uses information from multiple frames to converge to an accurate and stable estimate. Deviation from stable states is detected in order to remain robust to failed grasps and dropped objects. The tracker is integrated into a shared autonomy system in which it provides state estimates used by a grasp planner and the controller of two anthropomorphic hands. We demonstrate the advantages and performance of the tracking system in simulation and on a real robot. Qualitative results are also provided for a number of challenging manipulations that are made possible by the speed, accuracy, and stability of the tracking system. Tanner Schmidt, Katharina Hertkorn, Richard A. Newcombe, Zoltan-Csaba Marton, Michael Suppa, Dieter Fox |
ICRA | 1 |