Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tanner Schmidt

dblp:164/8399 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-5708-1257ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 68% Segmentation and scene understanding · 11% Efficient and distributed learning · 11%
Computer graphics and multimedia
3 papers
Computer animation and physical simulation · 50% Geometric modeling and processing · 38% Rendering · 13%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
neural radiance field
1.122022
Neural 3D Video Synthesis from Multi-view Video · CVPR 2022
STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021
Computer vision › Segmentation and scene understanding
image segmentation
0.912025
Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation · CVPR 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation · CVPR 2025
Computer vision › 3D vision › neural radiance field
dynamic neural radiance field
0.612022
Neural 3D Video Synthesis from Multi-view Video · CVPR 2022
Computer vision › 3D vision › novel view synthesis
multi-view video generation
0.612022
Neural 3D Video Synthesis from Multi-view Video · CVPR 2022
Computer vision › 3D vision
novel view synthesis
0.612022
Neural 3D Video Synthesis from Multi-view Video · CVPR 2022
Computer animation and physical simulation › motion synthesis
motion interpolation
0.612022
Neural 3D Video Synthesis from Multi-view Video · CVPR 2022
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.512021
STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021
Computer vision › 3D vision › object pose estimation
object pose tracking
0.512021
STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021
Computer vision › 3D vision › motion estimation
rigid motion estimation
0.512021
STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021
Computer vision › 3D vision
3d reconstruction
0.412020
Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction · ECCV (29) 2020
Computer vision › 3D vision › pose estimation
shape and pose estimation
0.412020
FroDO: From Detections to 3D Objects · CVPR 2020
Computer vision › 3D vision › 3d shape representation › implicit surface representation
signed distance function
0.412020
Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction · ECCV (29) 2020
Geometric modeling and processing
3d reconstruction
0.412020
FroDO: From Detections to 3D Objects · CVPR 2020
Computer vision › Video understanding and tracking › object tracking
articulated object tracking
0.212015
Depth-based tracking with physical constraints for robot manipulation · ICRA 2015
Robotics › Robot manipulation
grasping
0.212015
Depth-based tracking with physical constraints for robot manipulation · ICRA 2015
Robotics › Robot manipulation › grasping
grasp planning
0.212015
Depth-based tracking with physical constraints for robot manipulation · ICRA 2015
Computer vision › Video understanding and tracking
object tracking
0.212015
Depth-based tracking with physical constraints for robot manipulation · ICRA 2015
Rendering
novel view synthesis
0.112021
STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering · CVPR 2021
Human-robot interaction
shared control
0.112015
Depth-based tracking with physical constraints for robot manipulation · ICRA 2015

Methods — techniques the papers use, named apart from their topics

time-conditioned neural radiance field · 1.1ray importance sampling · 1.1hierarchical training · 1.1volume rendering · 1.0self-supervised learning · 1.0joint optimization · 1.0variable-resolution patch tokenization · 0.9multi-view optimization · 0.9encoder network · 0.9deep signed distance function · 0.9torque sensing · 0.2physical constraint modeling · 0.2depth-based tracking · 0.2
YearPublicationVenuePosition
2025 Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation
abstract
This paper presents Segment This Thing (STT), a new efficient image segmentation model designed to produce a single segment given a single point prompt. Instead of following prior work and increasing efficiency by decreasing model size, we gain efficiency by foveating input images. Given an image and a point prompt, we extract a crop centered on the prompt and apply a novel variable-resolution patch tokenization in which patches are downsampled at a rate that increases with increased distance from the prompt. This approach yields far fewer image tokens than uniform patch tokenization. As a result we can drastically reduce the computational cost of segmentation without reducing model size. Furthermore, the foveation focuses the model on the region of interest, a potentially useful inductive bias. We show that our Segment This Thing model is more efficient than prior work while remaining competitive on segmentation benchmarks. It can easily run at interactive frame rates on consumer hardware and is thus a promising tool for augmented reality or robotics applications.
Tanner Schmidt, Richard A. Newcombe
CVPR1
2022 Neural 3D Video Synthesis from Multi-view Video
abstract
We propose a novel approach for 3D video synthesis that is able to represent multi-view video recordings of a dynamic real-world scene in a compact, yet expressive representation that enables high-quality view synthesis and motion interpolation. Our approach takes the high quality and compactness of static neural radiance fields in a new direction: to a model-free, dynamic setting. At the core of our approach is a novel time-conditioned neural radiance field that represents scene dynamics using a set of compact latent codes. We are able to significantly boost the training speed and perceptual quality of the generated imagery by a novel hierarchical training scheme in combination with ray importance sampling. Our learned representation is highly compact and able to represent a 10 second 30 FPS multi-view video recording by 18 cameras with a model size of only 28MB. We demonstrate that our method can render high-fidelity wide-angle novel views at over 1K resolution, even for complex and dynamic scenes. We perform an extensive qualitative and quantitative evaluation that shows that our approach outperforms the state of the art. Project website: https://neural-3d-video.github.io/.
Tianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green, Christoph Lassner, Changil Kim 0001, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard A. Newcombe, Zhaoyang Lv
CVPR7
2021 STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering
abstract
We present STaR, a novel method that performs Self-supervised Tracking and Reconstruction of dynamic scenes with rigid motion from multi-view RGB videos without any manual annotation. Recent work has shown that neural networks are surprisingly effective at the task of compressing many views of a scene into a learned function which maps from a viewing ray to an observed radiance value via volume rendering. Unfortunately, these methods lose all their predictive power once any object in the scene has moved. In this work, we explicitly model rigid motion of objects in the context of neural representations of radiance fields. We show that without any additional human specified supervision, we can reconstruct a dynamic scene with a single rigid object in motion by simultaneously decomposing it into its two constituent parts and encoding each with its own neural representation. We achieve this by jointly optimizing the parameters of two neural radiance fields and a set of rigid poses which align the two fields at each frame. On both synthetic and real world datasets, we demonstrate that our method can render photorealistic novel views, where novelty is measured on both spatial and temporal axes. Our factored representation furthermore enables animation of unseen object motion.
Zhaoyang Lv, Tanner Schmidt, Steven Lovegrove
CVPR3
2020 FroDO: From Detections to 3D Objects
abstract
Object-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction of object instances from RGB video that infers their location, pose and shape in a coarse to fine manner. Key to FroDO is to embed object shapes in a novel learnt shape space that allows seamless switching between sparse point cloud and dense DeepSDF decoding. Given an input sequence of localized RGB frames, FroDO first aggregates 2D detections to instantiate a 3D bounding box per object. A shape code is regressed using an encoder network before optimizing shape and pose further under the learnt shape priors using sparse or dense shape representations. The optimization uses multi-view geometric, photometric and silhouette losses. We evaluate on real-world datasets, including Pix3D, Redwood-OS, and ScanNet, for single-view, multi-view, and multi-object reconstruction.
Martin Rünz, Kejie Li, Meng Tang 0001, Lingni Ma, Chen Kong, Tanner Schmidt, Ian D. Reid 0001, Lourdes Agapito, Julian Straub, Steven Lovegrove, Richard A. Newcombe
CVPR6
2020 Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction
Rohan Chabra, Jan Eric Lenssen, Eddy Ilg, Tanner Schmidt, Julian Straub, Steven Lovegrove, Richard A. Newcombe
ECCV (29)4
2017 Dynamic High Resolution Deformable Articulated Tracking
abstract
The last several years have seen significant progress in using depth cameras for tracking articulated objects such as human bodies, hands, and robotic manipulators. Most approaches focus on tracking skeletal parameters of a fixed shape model, which makes them insufficient for applications that require accurate estimates of deformable object surfaces. To overcome this limitation, we present a 3D model-based tracking system for articulated deformable objects. Our system is able to track human body pose and high resolution surface contours in real time using a commodity depth sensor and GPU hardware. We implement this as a joint optimization over a skeleton to account for changes in pose, and over the vertices of a high resolution mesh to track the subject's shape. Through experimental results we show that we are able to capture dynamic sub-centimeter surface detail such as folds and wrinkles in clothing. We also show that this shape estimation aids kinematic pose estimation by providing a more accurate target to match against the point cloud. The end result is highly accurate spatiotemporal and semantic information which is well suited for physical human robot interaction as well as virtual and augmented reality systems.
Aaron Walsman, Weilin Wan 0001, Tanner Schmidt, Dieter Fox
3DV3
2017 Self-directed Lifelong Learning for Robot Vision
Tanner Schmidt, Dieter Fox
ISRR1
2015 Depth-based tracking with physical constraints for robot manipulation
abstract
This work integrates visual and physical constraints to perform real-time depth-only tracking of articulated objects, with a focus on tracking a robot's manipulators and manipulation targets in realistic scenarios. As such, we extend DART, an existing visual articulated object tracker, to additionally avoid interpenetration of multiple interacting objects, and to make use of contact information collected via torque sensors or touch sensors. To achieve greater stability, the tracker uses a switching model to detect when an object is stationary relative to the table or relative to the palm and then uses information from multiple frames to converge to an accurate and stable estimate. Deviation from stable states is detected in order to remain robust to failed grasps and dropped objects. The tracker is integrated into a shared autonomy system in which it provides state estimates used by a grasp planner and the controller of two anthropomorphic hands. We demonstrate the advantages and performance of the tracking system in simulation and on a real robot. Qualitative results are also provided for a number of challenging manipulations that are made possible by the speed, accuracy, and stability of the tracking system.
Tanner Schmidt, Katharina Hertkorn, Richard A. Newcombe, Zoltan-Csaba Marton, Michael Suppa, Dieter Fox
ICRA1