David Joseph Tan

dblp:119/1285 · DBLP profile ↗
← Back
27ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-7835-6280ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3
YearPublicationVenuePosition
2025 Towards Real-Time Open-Vocabulary Video Instance Segmentation
Bin Yan 0004, Martin Sundermeyer, David Joseph Tan, Huchuan Lu, Federico Tombari
WACV3
2024 RaNeuS: Ray-adaptive Neural Surface Reconstruction
abstract
Our objective is to leverage a differentiable radiance field e.g. NeRF to reconstruct detailed 3D surfaces in addition to producing the standard novel view renderings. There have been related methods that perform such tasks, usually by utilizing a signed distance field (SDF). However, the state-of-the-art approaches still fail to correctly reconstruct the small-scale details, such as the leaves, ropes, and textile surfaces. Considering that different methods formulate and optimize the projection from SDF to radiance field with a globally constant Eikonal regularization, we improve with a ray-wise weighting factor to prioritize the rendering and zero-crossing surface fitting on top of establishing a perfect SDF. We propose to adaptively adjust the regularization on the signed distance field so that unsatisfying rendering rays won’t enforce strong Eikonal regularization which is ineffective, and allow the gradients from regions with well-learned radiance to effectively back-propagated to the SDF. Consequently, balancing the two objectives in order to generate accurate and detailed surfaces. Additionally, concerning whether there is a geometric bias between the zero-crossing surface in SDF and rendering points in the radiance field, the projection becomes adjustable as well depending on different 3D locations during optimization. Our proposed RaNeuS1are extensively evaluated on both synthetic and real datasets, achieving state-of-the-art results on both novel view synthesis and geometric reconstruction.1Codes are released at https://github.com/wangyida/ra-neus.
Yida Wang 0001, David Joseph Tan, Nassir Navab, Federico Tombari
3DV2
2024 SemiVL: Semi-Supervised Semantic Segmentation with Vision-Language Guidance
Lukas Hoyer, David Joseph Tan, Muhammad Ferjad Naeem, Luc Van Gool, Federico Tombari
ECCV (39)2
2024 Self-Supervised Latent Space Optimization With Nebula Variational Coding
abstract
Deep learning approaches process data in a layer-by-layer way with intermediate (or latent) features. We aim at designing a general solution to optimize the latent manifolds to improve the performance on classification, segmentation, completion and/or reconstruction through probabilistic models. This paper proposes a variational inference model which leads to a clustered embedding. We introduce additional variables in the latent space, called nebula anchors, that guide the latent variables to form clusters during training. To prevent the anchors from clustering among themselves, we employ the variational constraint that enforces the latent features within an anchor to form a Gaussian distribution, resulting in a generative model we refer as Nebula Variational Coding (NVC). Since each latent feature can be labeled with the closest anchor, we also propose to apply metric learning in a self-supervised way to make the separation between clusters more explicit. As a consequence, the latent variables of our variational coder form clusters which adapt to the generated semantic of the training data, e.g., the categorical labels of each sample. We demonstrate experimentally that it can be used within different architectures designed to solve different problems including text sequence, images, 3D point clouds and volumetric data, validating the advantage of our proposed method.
Yida Wang 0001, David Joseph Tan, Nassir Navab, Federico Tombari
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Shape, Pose, and Appearance from a Single Image via Bootstrapped Radiance Field Inversion
abstract
Neural Radiance Fields (NeRF) coupled with CANs represent a promising direction in the area of 3D reconstruction from a single view, owing to their ability to efficiently model arbitrary topologies. Recent work in this area, however, has mostly focused on synthetic datasets where exact ground-truth poses are known, and has overlooked pose estimation, which is important for certain down-stream applications such as augmented reality (AR) and robotics. We introduce a principled end-to-end reconstructionframeworkfor natural images, where accurate ground-truth poses are not available. Our approach recovers an SDF-parameterized 3D shape, pose, and appearance from a single image of an object, without exploiting multiple views during training. More specifically, we leverage an unconditional 3D-aware generator, to which we apply a hybrid inversion scheme where a model produces a first guess of the solution which is then refined via optimization. Our frame-work can de-render an image in as few as 10 steps, enabling its use in practical scenarios. We demonstrate state-of-the-art results on a variety of real and synthetic benchmarks.
Dario Pavllo, David Joseph Tan, Marie-Julie Rakotosaona, Federico Tombari
CVPR2
2022 Learning Local Displacements for Point Cloud Completion
abstract
We propose a novel approach aimed at object and semantic scene completion from a partial scan represented as a 3D point cloud. Our architecture relies on three novel layers that are used successively within an encoder-decoder structure and specifically developed for the task at hand. The first one carries out feature extraction by matching the point features to a set of pre-trained local descriptors. Then, to avoid losing individual descriptors as part of standard operations such as max-pooling, we propose an alternative neighbor-pooling operation that relies on adopting the feature vectors with the highest activations. Finally, upsampling in the decoder modifies our feature extraction in order to increase the output dimension. While this model is already able to achieve competitive results with the state of the art, we further propose a way to increase the versatility of our approach to process point clouds. To this aim, we introduce a second model that assembles our layers within a transformer architecture. We evaluate both architectures on object and indoor scene completion tasks, achieving state-of-the-art performance.
Yida Wang 0001, David Joseph Tan, Nassir Navab, Federico Tombari
CVPR2
2022 SoftPool++: An Encoder-Decoder Network for Point Cloud Completion
abstract
Abstract We propose a novel convolutional operator for the task of point cloud completion. One striking characteristic of our approach is that, conversely to related work it does not require any max-pooling or voxelization operation. Instead, the proposed operator used to learn the point cloud embedding in the encoder extracts permutation-invariant features from the point cloud via a soft-pooling of feature activations, which are able to preserve fine-grained geometric details. These features are then passed on to a decoder architecture. Due to the compression in the encoder, a typical limitation of this type of architectures is that they tend to lose parts of the input shape structure. We propose to overcome this limitation by using skip connections specifically devised for point clouds, where links between corresponding layers in the encoder and the decoder are established. As part of these connections, we introduce a transformation matrix that projects the features from the encoder to the decoder and vice-versa. The quantitative and qualitative results on the task of object completion from partial scans on the ShapeNet dataset show that incorporating our approach achieves state-of-the-art performance in shape completion both at low and high resolutions.
Yida Wang 0001, David Joseph Tan, Nassir Navab, Federico Tombari
Int. J. Comput. Vis.2
2020 A Divide et Impera Approach for 3D Shape Reconstruction from Multiple Views
abstract
Estimating the 3D shape of an object from a single or multiple images has gained popularity thanks to the recent breakthroughs powered by deep learning. Most approaches regress the full object shape in a canonical pose, possibly extrapolating the occluded parts based on the learned priors. However, their viewpoint invariant technique often discards the unique structures visible from the input images. In contrast, this paper proposes to rely on viewpoint variant reconstructions by merging the visible information from the given views. Our approach is divided into three steps. Starting from the sparse views of the object, we first align them into a common coordinate system by estimating the relative pose between all the pairs. Then, inspired by the traditional voxel carving, we generate an occupancy grid of the object taken from the silhouette on the images and their relative poses. Finally, we refine the initial reconstruction to build a clean 3D model which preserves the details from each viewpoint. To validate the proposed method, we perform a comprehensive evaluation on the ShapeNet reference benchmark in terms of relative pose estimation and 3D shape reconstruction.
Riccardo Spezialetti, David Joseph Tan, Alessio Tonioni, Keisuke Tateno, Federico Tombari
3DV2
2020 SoftPoolNet: Shape Descriptor for Point Cloud Completion and Classification
Yida Wang 0001, David Joseph Tan, Nassir Navab, Federico Tombari
ECCV (3)2
2019 ForkNet: Multi-Branch Volumetric Semantic Completion From a Single Depth Image
abstract
We propose a novel model for 3D semantic completion from a single depth image, based on a single encoder and three separate generators used to reconstruct different geometric and semantic representations of the original and completed scene, all sharing the same latent space. To transfer information between the geometric and semantic branches of the network, we introduce paths between them concatenating features at corresponding network layers. Motivated by the limited amount of training samples from real scenes, an interesting attribute of our architecture is the capacity to supplement the existing dataset by generating a new training dataset with high quality, realistic scenes that even includes occlusion and real noise. We build the new dataset by sampling the features directly from latent space which generates a pair of partial volumetric surface and completed volumetric semantic surface. Moreover, we utilize multiple discriminators to increase the accuracy and realism of the reconstructions. We demonstrate the benefits of our approach on standard benchmarks for the two most common completion tasks: semantic 3D scene completion and 3D object completion.
Yida Wang 0001, David Joseph Tan, Nassir Navab, Federico Tombari
ICCV2
2018 Adversarial Semantic Scene Completion from a Single Depth Image
abstract
We propose a method to reconstruct, complete and semantically label a 3D scene from a single input depth image. We improve the accuracy of the regressed semantic 3D maps by a novel architecture based on adversarial learning. In particular, we suggest using multiple adversarial loss terms that not only enforce realistic outputs with respect to the ground truth, but also an effective embedding of the internal features. This is done by correlating the latent features of the encoder working on partial 2.5D data with the latent features extracted from a variational 3D auto-encoder trained to reconstruct the complete semantic scene. In addition, differently from other approaches that operate entirely through 3D convolutions, at test time we retain the original 2.5D structure of the input during downsampling to improve the effectiveness of the internal representation of our model. We test our approach on the main benchmark datasets for semantic scene completion to qualitatively and quantitatively assess the effectiveness of our proposal.
Yida Wang 0001, David Joseph Tan, Nassir Navab, Federico Tombari
3DV2
2018 Human Motion Analysis with Deep Metric Learning
Huseyin Coskun, David Joseph Tan, Sailesh Conjeti, Nassir Navab, Federico Tombari
ECCV (14)2
2018 Real-Time Accurate 3D Head Tracking and Pose Estimation with Consumer RGB-D Cameras
David Joseph Tan, Federico Tombari, Nassir Navab
Int. J. Comput. Vis.1
2017 One For All: Adaptive Learning-based Temporal Tracker for 3D Head Shape Models
David Joseph Tan, Federico Tombari, Nassir Navab
BMVC1
2017 Looking Beyond the Simple Scenarios: Combining Learners and Optimizers in 3D Temporal Tracking
abstract
3D object temporal trackers estimate the 3D rotation and 3D translation of a rigid object by propagating the transformation from one frame to the next. To confront this task, algorithms either learn the transformation between two consecutive frames or optimize an energy function to align the object to the scene. The motivation behind our approach stems from a consideration on the nature of learners and optimizers. Throughout the evaluation of different types of objects and working conditions, we observe their complementary nature - on one hand, learners are more robust when undergoing challenging scenarios, while optimizers are prone to tracking failures due to the entrapment at local minima; on the other, optimizers can converge to a better accuracy and minimize jitter. Therefore, we propose to bridge the gap between learners and optimizers to attain a robust and accurate RGB-D temporal tracker that runs at approximately 2 ms per frame using one CPU core. Our work is highly suitable for Augmented Reality (AR), Mixed Reality (MR) and Virtual Reality (VR) applications due to its robustness, accuracy, efficiency and low latency. Aiming at stepping beyond the simple scenarios used by current systems, often constrained by having a single object in the absence of clutter, averting to touch the object to prevent close-range partial occlusion or selecting brightly colored objects to easily segment them individually, we demonstrate the capacity to handle challenging cases under clutter, partial occlusion and varying lighting conditions.
David Joseph Tan, Nassir Navab, Federico Tombari
IEEE Trans. Vis. Comput. Graph.1
2016 Fits Like a Glove: Rapid and Reliable Hand Shape Personalization
abstract
We present a fast, practical method for personalizing a hand shape basis to an individual user's detailed hand shape using only a small set of depth images. To achieve this, we minimize an energy based on a sum of render-and-compare cost functions called the golden energy. However, this energy is only piecewise continuous, due to pixels crossing occlusion boundaries, and is therefore not obviously amenable to efficient gradient-based optimization. A key insight is that the energy is the combination of a smooth low-frequency function with a high-frequency, low-amplitude, piecewisecontinuous function. A central finite difference approximation with a suitable step size can therefore jump over the discontinuities to obtain a good approximation to the energy's low-frequency behavior, allowing efficient gradient-based optimization. Experimental results quantitatively demonstrate for the first time that detailed personalized models improve the accuracy of hand tracking and achieve competitive results in both tracking and model registration.
David Joseph Tan, Thomas J. Cashman 0001, Jonathan Taylor 0001, Andrew W. Fitzgibbon, Daniel Tarlow, Sameh Khamis, Shahram Izadi, Jamie Shotton
CVPR1
2016 Real-Time Online Adaption for Robust Instrument Tracking and Pose Estimation
Nicola Rieke, David Joseph Tan, Federico Tombari, Josué Page Vizcaíno, Chiara Amat di San Filippo, Abouzar Eslami, Nassir Navab
MICCAI (1)2
2016 Real-time localization of articulated surgical instruments in retinal microsurgery
Nicola Rieke, David Joseph Tan, Chiara Amat di San Filippo, Federico Tombari, Mohamed Alsheakhali, Vasileios Belagiannis, Abouzar Eslami, Nassir Navab
Medical Image Anal.2
2015 A Combined Generalized and Subject-Specific 3D Head Pose Estimation
abstract
We propose a real-time method for 3D head pose estimation from RGB-D sequences. Our algorithm relies on a Random Forest framework that is able to regress the head pose at every frame in a temporal tracking manner. Such framework is learned once from a generic dataset of 3D head models and refined online to adapt the forest to the specific characteristics of each subject. Through the qualitative experiments under different conditions, it demonstrates remarkable properties in terms of robustness to occlusions, computational efficiency and capacity of handling a variety of challenging head poses. In addition, it also outperforms the state of the art on the reference benchmark dataset with regards to the accuracy of the estimated head poses.
David Joseph Tan, Federico Tombari, Nassir Navab
3DV1
2015 A Versatile Learning-Based 3D Temporal Tracker: Scalable, Robust, Online
abstract
This paper proposes a temporal tracking algorithm based on Random Forest that uses depth images to estimate and track the 3D pose of a rigid object in real-time. Compared to the state of the art aimed at the same goal, our algorithm holds important attributes such as high robustness against holes and occlusion, low computational cost of both learning and tracking stages, and low memory consumption. These are obtained (a) by a novel formulation of the learning strategy, based on a dense sampling of the camera viewpoints and learning independent trees from a single image for each camera view, as well as, (b) by an insightful occlusion handling strategy that enforces the forest to recognize the object's local and global structures. Due to these attributes, we report state-of-the-art tracking accuracy on benchmark datasets, and accomplish remarkable scalability with the number of targets, being able to simultaneously track the pose of over a hundred objects at 30~fps with an off-the-shelf CPU. In addition, the fast learning time enables us to extend our algorithm as a robust online tracker for model-free 3D objects under different viewpoints and appearance changes as demonstrated by the experiments.
David Joseph Tan, Federico Tombari, Slobodan Ilic, Nassir Navab
ICCV1
2015 Surgical Tool Tracking and Pose Estimation in Retinal Microsurgery
Nicola Rieke, David Joseph Tan, Mohamed Alsheakhali, Federico Tombari, Chiara Amat di San Filippo, Vasileios Belagiannis, Abouzar Eslami, Nassir Navab
MICCAI (1)2
2015 Efficient Learning of Linear Predictors for Template Tracking
Stefan Holzer, Slobodan Ilic, David Joseph Tan, Marc Pollefeys, Nassir Navab
Int. J. Comput. Vis.3
2014 Deformable Template Tracking in 1ms
David Joseph Tan, Stefan Holzer, Nassir Navab, Slobodan Ilic
BMVC1
2014 Multi-forest Tracker: A Chameleon in Tracking
abstract
In this paper, we address the problem of object tracking in intensity images and depth data. We propose a generic framework that can be used either for tracking 2D templates in intensity images or for tracking 3D objects in depth images. To overcome problems like partial occlusions, strong illumination changes and motion blur, that notoriously make energy minimization-based tracking methods get trapped in a local minimum, we propose a learning-based method that is robust to all these problems. We use random forests to learn the relation between the parameters that defines the object's motion, and the changes they induce on the image intensities or the point cloud of the template. It follows that, to track the template when it moves, we use the changes on the image intensities or point cloud to predict the parameters of this motion. Our algorithm has an extremely fast tracking performance running at less than 2 ms per frame, and is robust to partial occlusions. Moreover, it demonstrates robustness to strong illumination changes when tracking templates using intensity images, and robustness in tracking 3D objects from arbitrary viewpoints even in the presence of motion blur that causes missing or erroneous data in depth images. Extensive experimental evaluation and comparison to the related approaches strongly demonstrates the benefits of our method.
David Joseph Tan, Slobodan Ilic
CVPR1
2013 Multi-task Forest for Human Pose Estimation in Depth Images
abstract
In this paper, we address the problem of human body pose estimation from depth data. Previous works Based on random forests relied either on a classification strategy to infer the different body parts or on a regression approach to predict directly the joint positions. To permit the inference of very generic poses, those approaches did not consider additional information during the learning phase, e.g. the performed activity. In the present work, we introduce a novel approach to integrate additional information at training time that actually improves the pose prediction during the testing. Our main contribution is a multi-task forest that aims at solving a joint regression-classification task: each foreground pixel from a depth image is associated to its relative displacements to the 3D joint positions as well as the activity class. Integrating activity information in the objective function during forest training permits a better partitioning of the 3D pose space that leads to a better modelling of the posterior. Thereby, our approach provides an improved pose prediction, and as a by-product, can give an estimate of the performed activity. We performed experiments on a dataset performed by 10 people associated with the ground truth body poses from a motion capture system. To demonstrate the benefits of our approach, poses are divided into 10 different activities for the training phase. Results on this dataset show that our multi-task forest provides improved human pose estimation compared to a pure regression forest approach.
Joé Lallemand, Olivier Pauly, Loren Arthur Schwarz, David Joseph Tan, Slobodan Ilic
3DV4
2012 Efficient Learning of Linear Predictors Using Dimensionality Reduction
Stefan Holzer, Slobodan Ilic, David Joseph Tan, Nassir Navab
ACCV (3)3
2012 Online Learning of Linear Predictors for Real-Time Tracking
Stefan Holzer, Marc Pollefeys, Slobodan Ilic, David Joseph Tan, Nassir Navab
ECCV (1)4