EDBT 2026 Demo / reviewers in the wild / expert
Vincent Lepetit
dblp:80/5556
· DBLP profile ↗
178ranked-venue papers
9as first author
38since 2021 · last 2026
0000-0001-9985-4433ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 142 · 7 first-author · 30 since 2021Artificial intelligence and machine learning · 129 · 7 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 17 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is Clustering Enough for LiDAR Instance Segmentation? A State-of-the-Art Training-Free Baseline
Corentin Sautier, Gilles Puy, Alexandre Boulch, Renaud Marlet, Vincent Lepetit |
3DV | 5 |
| 2026 | BOP-Distrib: Revisiting 6D Pose Estimation Benchmarks for Better Evaluation under Visual Ambiguitiesabstract6D pose estimation aims at determining the object pose that best explains the camera observation. The unique solution for non-ambiguous objects can turn into a multimodal pose distribution for symmetrical objects or when occlusions of symmetry-breaking elements happen, depending on the viewpoint. Currently, 6D pose estimation methods are benchmarked on datasets that consider, for their ground truth annotations, visual ambiguities as only related to global object symmetries, whereas they should be defined per-image to account for the camera viewpoint. We thus first propose an automatic method to re-annotate those datasets with a 6D pose distribution specific to each image, taking into account the object surface visibility in the image to correctly determine the visual ambiguities. Second, given this improved ground truth, we re-evaluate the state-of-the-art single pose methods and show that this greatly modifies the ranking of these methods. Third, as some recent works focus on estimating the complete set of solutions, we derive a precision/recall formulation to evaluate them against our image-wise distribution ground truth, making it the first benchmark for pose distribution methods on real images. Boris Meden, Asma Brazi, Fabrice Mayran de Chamisso, Steve Bourgeois, Vincent Lepetit |
WACV | 5 |
| 2025 | UNIT: Unsupervised Online Instance Segmentation Through TimeabstractOnline object segmentation and tracking in Lidar point clouds enables autonomous agents to understand their surroundings and make safe decisions. Unfortunately, manual annotations for these tasks are prohibitively costly. We tackle this problem with the task of class-agnostic unsupervised online instance segmentation and tracking. To that end, we leverage an instance segmentation backbone and propose a new training recipe that enables the online tracking of objects. Our network is trained on pseudo-labels, eliminating the need for manual annotations. We conduct an evaluation using metrics adapted for temporal instance segmentation. Computing these metrics requires temporally-consistent instance labels. When unavailable, we construct these labels using the available 3D bounding boxes and semantic labels in the dataset. We compare our method against strong baselines and demonstrate its superiority across two different outdoor Lidar datasets. Project page: csautier.github.io/unit Corentin Sautier, Gilles Puy, Alexandre Boulch, Renaud Marlet, Vincent Lepetit |
3DV | 5 |
| 2025 | NextBestPath: Efficient 3D Mapping of Unseen EnvironmentsabstractThis work addresses the problem of active 3D mapping, where an agent must find an efficient trajectory to exhaustively reconstruct a new scene.
Previous approaches mainly predict the next best view near the agent's location, which is prone to getting stuck in local areas. Additionally, existing indoor datasets are insufficient due to limited geometric complexity and inaccurate ground truth meshes.
To overcome these limitations, we introduce a novel dataset AiMDoom with a map generator for the Doom video game, enabling to better benchmark active 3D mapping in diverse indoor environments.
Moreover, we propose a new method we call next-best-path (NBP), which predicts long-term goals rather than focusing solely on short-sighted views.
The model jointly predicts accumulated surface coverage gains for long-term goals and obstacle maps, allowing it to efficiently plan optimal paths with a unified model.
By leveraging online data collection, data augmentation and curriculum learning, NBP significantly outperforms state-of-the-art methods on both the existing MP3D dataset and our AiMDoom dataset, achieving more efficient mapping in indoor environments of varying complexity. Antoine Guédon, Clémentin Boittiaux, Shizhe Chen, Vincent Lepetit |
ICLR | 5 |
| 2025 | Alligat0R: Pre-Training through Covisibility Segmentation for Relative Camera Pose RegressionabstractPre-training techniques have greatly advanced computer vision, with CroCo’s cross-view completion approach yielding impressive results in tasks like 3D reconstruction and pose regression. However, cross-view completion is ill-posed in non-covisible regions, limiting its effectiveness. We introduce Alligat0R, a novel pre-training approach that replaces cross-view learning with a covisibility segmentation task. Our method predicts whether each pixel in one image is covisible in the second image, occluded, or outside the field of view, making the pre-training effective in both covisible and non-covisible regions, and provides interpretable predictions. To support this, we present Cub3, a large-scale dataset with 5M image pairs and dense covisibility annotations derived from the nuScenes and ScanNet datasets. Cub3 includes diverse scenarios with varying degrees of overlap. The experiments show that our novel pre-training method Alligat0R significantly outperforms CroCo in relative pose regression. Alligat0R and Cub3 will be made publicly available. Thibaut Loiseau, Guillaume Bourmaud, Vincent Lepetit |
NeurIPS | 3 |
| 2025 | GuideFlow3D: Optimization-Guided Rectified Flow For Appearance TransferabstractTransferring appearance to 3D assets using different representations of the appearance object - such as images or text - has garnered interest due to its wide range of applications in industries like gaming, augmented reality, and digital content creation. However, state-of-the-art methods still fail when the geometry between the input and appearance objects is significantly different. A straightforward approach is to directly apply a 3D generative model, but we show that this ultimately fails to produce appealing results. Instead, we propose a principled approach inspired by universal guidance. Given a pretrained rectified flow model conditioned on image or text, our training-free method interacts with the sampling process by periodically adding guidance. This guidance can be modeled as a differentiable loss function, and we experiment with two different types of guidance including part-aware losses for appearance and self-similarity. Our experiments show that our approach successfully transfers texture and geometric details to the input 3D asset, outperforming baselines both qualitatively and quantitatively. We also show that traditional metrics are not suitable for evaluating the task due to their inability of focusing on local details and comparing dissimilar inputs, in absence of ground truth data. We thus evaluate appearance transfer quality with a GPT-based system objectively ranking outputs, ensuring robust and human-like assessment, as further confirmed by our user study. Beyond showcased scenarios, our method is general and could be extended to different types of diffusion models and guidance functions. Project Page: *[https://sayands.github.io/guideflow3d](https://sayands.github.io/guideflow3d)* Sayan Deb Sarkar, Sinisa Stekovic, Vincent Lepetit, Iro Armeni |
NeurIPS | 3 |
| 2025 | MCTS With Refinement for Proposals Selection Games in Scene UnderstandingabstractWe propose a novel method applicable in many scene understanding problems that adapts the Monte Carlo Tree Search (MCTS) algorithm, originally designed to learn to play games of high-state complexity. From a generated pool of proposals, our method jointly selects and optimizes proposals that minimize the objective term. In our first application for floor plan reconstruction from point clouds, our method selects and refines the room proposals, modelled as 2D polygons, by optimizing on an objective function combining the fitness as predicted by a deep network and regularizing terms on the room shapes. We also introduce a novel differentiable method for rendering the polygonal shapes of these proposals. Our evaluations on the recent and challenging Structured3D and Floor-SP datasets show significant improvements over the state-of-the-art both in speed and quality of reconstructions, without imposing hard constraints nor assumptions on the floor plan configurations. In our second application, we extend our approach to reconstruct general 3D room layouts from a color image and obtain accurate room layouts. We also show that our differentiable renderer can easily be extended for rendering 3D planar polygons and polygon embeddings. Our method shows high performance on the Matterport3D-Layout dataset, without introducing hard constraints on room layout configurations. Sinisa Stekovic, Mahdi Rad, Friedrich Fraundorfer, Vincent Lepetit |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | HOC-Search: Efficient CAD Model and Pose Retrieval From RGB-D ScansabstractWe present an automated and efficient approach for retrieving high-quality CAD models of objects and their poses in a scene captured by a moving RGB-D camera. We first investigate various objective functions to measure similarity between a candidate CAD object model and the available data, and the best objective function appears to be a ”render-and-compare” method comparing depth and mask rendering. We thus introduce a fast-search method that approximates an exhaustive search based on this objective function for simultaneously retrieving the object category, a CAD model, and the pose of an object given an approximate 3D bounding box. This method involves a search tree that organizes the CAD models and object properties including object category and pose for fast retrieval and an algorithm inspired by Monte Carlo Tree Search, that efficiently searches this tree. We show that this method retrieves CAD models that fit the real objects very well, with a speed-up factor of 10$\times$ to 120$\times$ compared to exhaustive search. We used our method to automatically retrieve CAD models for objects in the ScanNet dataset; our annotations are available at https://github.com/stefanainetter/SCANnotateDataset. Stefan Ainetter, Sinisa Stekovic, Friedrich Fraundorfer, Vincent Lepetit |
3DV | 4 |
| 2024 | BEVContrast: Self-Supervision in BEV Space for Automotive Lidar Point CloudsabstractWe present a surprisingly simple and efficient method for self-supervision of 3D backbone on automotive Lidar point clouds. We design a contrastive loss between features of Lidar scans captured in the same scene. Several such approaches have been proposed in the literature from PointConstrast [40], which uses a contrast at the level of points, to the state-of-the-art TARL [30], which uses a contrast at the level of segments, roughly corresponding to objects. While the former enjoys a great simplicity of implementation, it is surpassed by the latter, which however requires a costly pre-processing. In BEVContrast, we define our contrast at the level of 2D cells in the Bird’s Eye View plane. Resulting cell-level representations offer a good trade-off between the point-level representations exploited in PointContrast and segment-level representations exploited in TARL: we retain the simplicity of PointContrast (cell representations are cheap to compute) while surpassing the performance of TARL in downstream semantic segmentation. The code is available at github.com/valeoai/BEVContrast Corentin Sautier, Gilles Puy, Alexandre Boulch, Renaud Marlet, Vincent Lepetit |
3DV | 5 |
| 2024 | SuGaR: Surface-Aligned Gaussian Splatting for Efficient 3D Mesh Reconstruction and High-Quality Mesh RenderingabstractWe propose a method to allow precise and extremely fast mesh extraction from 3D Gaussian Splatting [15]. Gaussian Splatting has recently become very popular as it yields realistic rendering while being significantly faster to train than NeRFs. It is however challenging to extract a mesh from the millions of tiny 3D Gaussians as these Gaussians tend to be unorganized after optimization and no method has been proposed so far. Our first key contribution is a regularization term that encourages the Gaussians to align well with the surface of the scene. We then introduce a method that exploits this alignment to extract a mesh from the Gaussians using Poisson reconstruction, which is fast, scalable, and preserves details, in contrast to the Marching Cubes algorithm usually applied to extract meshes from Neural SDFs. Finally, we introduce an optional refinement strategy that binds Gaussians to the surface of the mesh, and jointly optimizes these Gaussians and the mesh through Gaussian splatting rendering. This enables easy editing, sculpting, animating, and relighting of the Gaussians by manipulating the mesh instead of the Gaussians themselves. Retrieving such an editable mesh for realistic rendering is done within minutes with our method, compared to hours with the state-of-the-art method on SDFs, while providing a better rendering quality. Antoine Guédon, Vincent Lepetit |
CVPR | 2 |
| 2024 | NOPE: Novel Object Pose Estimation from a Single ImageabstractThe practicality of 3D object pose estimation remains limited for many applications due to the need for prior knowledge of a 3D model and a training period for new objects. To address this limitation, we propose an approach that takes a single image of a new object as input and pre-dicts the relative pose of this object in new images without prior knowledge of the object's 3D model and without re-quiring training time for new objects and categories. We achieve this by training a model to directly predict discrim-inative embeddings for viewpoints surrounding the object. This prediction is done using a simple U-Net architecture with attention and conditioned on the desired pose, which yields extremely fast inference. We compare our approach to state-of-the-art methods and show it outperforms them both in terms of accuracy and robustness. Van Nguyen Nguyen, Thibault Groueix, Georgy Ponimatkin, Yinlin Hu, Renaud Marlet, Mathieu Salzmann, Vincent Lepetit |
CVPR | 7 |
| 2024 | GigaPose: Fast and Robust Novel Object Pose Estimation via One CorrespondenceabstractWe present GigaPose, afast, robust, and accurate method for CAD-based novel object pose estimation in RGB images. GigaPose first leverages discriminative “templates ”, ren-dered images of the CAD models, to recover the out-of-plane rotation and then uses patch correspondences to estimate the four remaining parameters. Our approach samples tem-plates in only a two-degrees-of-freedom space instead of the usual three and matches the input image to the templates using fast nearest-neighbor search in feature space, results in a speedup factor of 35x compared to the state of the art. More-over, GigaPose is significantly more robust to segmentation errors. Our extensive evaluation on the seven core datasets of the BOP challenge demonstrates that it achieves state-of-the-art accuracy and can be seamlessly integrated with existing refinement methods. Additionally, we show the potential of GigaPose with 3D models predicted by recent work on 3D reconstruction from a single image, relaxing the need for CAD models and making 6D pose object estimation much more convenient. Our source code and trained models are publicly available at https://github.conllnv-nguyenlgigaPose. Van Nguyen Nguyen, Thibault Groueix, Mathieu Salzmann, Vincent Lepetit |
CVPR | 4 |
| 2024 | Gaussian Frosting: Editable Complex Radiance Fields with Real-Time Rendering
Antoine Guédon, Vincent Lepetit |
ECCV (65) | 2 |
| 2024 | Correspondences of the Third Kind: Camera Pose Estimation from Object Reflection
Kohei Yamashita 0001, Vincent Lepetit, Ko Nishino |
ECCV (65) | 2 |
| 2024 | Two Projections Suffice for Cerebral Vascular Reconstruction
Alexandre Cafaro, Reuben Dorent, Nazim Haouchine, Vincent Lepetit, Nikos Paragios, William M. Wells III, Sarah F. Frisken |
MICCAI (7) | 4 |
| 2023 | MACARONS: Mapping and Coverage Anticipation with RGB Online Self-SupervisionabstractWe introduce a method that simultaneously learns to explore new large environments and to reconstruct them in 3D from color images only. This is closely related to the Next Best View problem (NBV), where one has to identify where to move the camera next to improve the coverage of an unknown scene. However, most of the current NBV methods rely on depth sensors, need 3D supervision and/or do not scale to large scenes. Our method requires only a color camera and no 3D supervision. It simultaneously learns in a self-supervised fashion to predict a “volume occupancy field” from color images and, from this field, to predict the NBV. Thanks to this approach, our method performs well on new scenes as it is not biased towards any training 3D data. We demonstrate this on a recent dataset made of various 3D scenes and show it performs even better than recent methods requiring a depth sensor, which is not a realistic assumption for outdoor scenes captured with a flying drone. Antoine Guédon, Tom Monnier, Pascal Monasse, Vincent Lepetit |
CVPR | 4 |
| 2023 | In-Hand 3D Object Scanning from an RGB SequenceabstractWe propose a method for in-hand 3D scanning of an unknown object with a monocular camera. Our method relies on a neural implicit surface representation that captures both the geometry and the appearance of the object, however, by contrast with most NeRF-based methods, we do not assume that the camera-object relative poses are known. Instead, we simultaneously optimize both the object shape and the pose trajectory. As direct optimization over all shape and pose parameters is prone to fail without coarse-level initialization, we propose an incremental approach that starts by splitting the sequence into carefully selected overlapping segments within which the optimization is likely to succeed. We reconstruct the object shape and track its poses independently within each segment, then merge all the segments before performing a global optimization. We show that our method is able to reconstruct the shape and color of both textured and challenging texture-less objects, outperforms classical methods that rely only on appearance features, and that its performance is close to recent methods that assume known camera poses. Shreyas Hampali, Tomas Hodan, Luan Tran, Lingni Ma, Cem Keskin, Vincent Lepetit |
CVPR | 6 |
| 2023 | You Never Get a Second Chance To Make a Good First Impression: Seeding Active Learning for 3D Semantic SegmentationabstractWe propose SeedAL, a method to seed active learning for efficient annotation of 3D point clouds for semantic segmentation. Active Learning (AL) iteratively selects relevant data fractions to annotate within a given budget, but requires a first fraction of the dataset (a ’seed’) to be already annotated to estimate the benefit of annotating other data fractions. We first show that the choice of the seed can significantly affect the performance of many AL methods. We then propose a method for automatically constructing a seed that will ensure good performance for AL. Assuming that images of the point clouds are available, which is common, our method relies on powerful unsupervised image features to measure the diversity of the point clouds. It selects the point clouds for the seed by optimizing the diversity under an annotation budget, which can be done by solving a linear optimization problem. Our experiments demonstrate the effectiveness of our approach compared to random seeding and existing methods on both the S3DIS and SemanticKitti datasets. Code is available at https://github.com/nerminsamet/seedal. Nermin Samet, Oriane Siméoni, Gilles Puy, Georgy Ponimatkin, Renaud Marlet, Vincent Lepetit |
ICCV | 6 |
| 2023 | X2Vision: 3D CT Reconstruction from Biplanar X-Rays with Deep Structure Prior
Alexandre Cafaro, Quentin Spinat, Amaury Leroy, Pauline Maury, Alexandre Munoz, Guillaume Beldjoudi, Charlotte Robert, Eric Deutsch, Vincent Grégoire, Vincent Lepetit, Nikos Paragios |
MICCAI (10) | 10 |
| 2023 | StructuRegNet: Structure-Guided Multimodal 2D-3D Registration
Amaury Leroy, Alexandre Cafaro, Grégoire Gessain, Anne Champagnac, Vincent Grégoire, Eric Deutsch, Vincent Lepetit, Nikos Paragios |
MICCAI (10) | 7 |
| 2023 | Automatically Annotating Indoor Images with CAD Models via RGB-D ScansabstractWe present an automatic method for annotating images of indoor scenes with the CAD models of the objects by relying on RGB-D scans. Through a visual evaluation by 3D experts, we show that our method retrieves annotations that are at least as accurate as manual annotations, and can thus be used as ground truth without the burden of manually annotating 3D data. We do this using an analysis-by-synthesis approach, which compares renderings of the CAD models with the captured scene. We introduce a ’cloning procedure’ that identifies objects that have the same geometry, to annotate these objects with the same CAD models. This allows us to obtain complete annotations for the Scan-Net dataset and the recent ARKitScenes dataset. We will release these annotations publicly, as we believe they will be very useful for the computer vision community. Stefan Ainetter, Sinisa Stekovic, Friedrich Fraundorfer, Vincent Lepetit |
WACV | 4 |
| 2023 | Back to MLP: A Simple Baseline for Human Motion PredictionabstractThis paper tackles the problem of human motion prediction, consisting in forecasting future body poses from historically observed sequences. State-of-the-art approaches provide good results, however, they rely on deep learning architectures of arbitrary complexity, such as Recurrent Neural Networks (RNN), Transformers or Graph Convolutional Networks (GCN), typically requiring multiple training stages and more than 2 million parameters. In this paper, we show that, after combining with a series of standard practices, such as applying Discrete Cosine Transform (DCT), predicting residual displacement of joints and optimizing velocity as an auxiliary loss, a light-weight network based on multi-layer perceptrons (MLPs) with only 0.14 million parameters can surpass the state-of-the-art performance. An exhaustive evaluation on the Human3.6M, AMASS, and 3DPW datasets shows that our method, named siMLpe, consistently outperforms all other approaches. We hope that our simple method could serve as a strong baseline for the community and allow re-thinking of the human motion prediction problem. The code is publicly available at https://github.com/dulucas/siMLPe. Wen Guo 0004, Yuming Du, Xi Shen 0001, Vincent Lepetit, Xavier Alameda-Pineda, Francesc Moreno-Noguer |
WACV | 4 |
| 2023 | A Simple and Powerful Global Optimization for Unsupervised Video Object SegmentationabstractWe propose a simple, yet powerful approach for unsupervised object segmentation in videos. We introduce an objective function whose minimum represents the mask of the main salient object over the input sequence. It only relies on independent image features and optical flows, which can be obtained using off-the-shelf self-supervised methods. It scales with the length of the sequence with no need for superpixels or sparsification, and it generalizes to different datasets without any specific training. This objective function can actually be derived from a form of spectral clustering applied to the entire video. Our method achieves on-par performance with the state of the art on standard bench-marks (DAVIS2016, SegTrack-v2, FBMS59), while being conceptually and practically much simpler. Georgy Ponimatkin, Nermin Samet, Yang Xiao 0009, Yuming Du, Renaud Marlet, Vincent Lepetit |
WACV | 6 |
| 2023 | Few-Shot Object Detection and Viewpoint Estimation for Objects in the WildabstractDetecting objects and estimating their viewpoints in images are key tasks of 3D scene understanding. Recent approaches have achieved excellent results on very large benchmarks for object detection and viewpoint estimation. However, performances are still lagging behind for novel object categories with few samples. In this paper, we tackle the problems of few-shot object detection and few-shot viewpoint estimation. We demonstrate on both tasks the benefits of guiding the network prediction with class-representative features extracted from data in different modalities: image patches for object detection, and aligned 3D models for viewpoint estimation. Despite its simplicity, our method outperforms state-of-the-art methods by a large margin on a range of datasets, including PASCAL and COCO for few-shot object detection, and Pascal3D+ and ObjectNet3D for few-shot viewpoint estimation. Furthermore, when the 3D model is not available, we introduce a simple category-agnostic viewpoint estimation method by exploiting geometrical similarities and consistent pose labeling across different classes. While it moderately reduces performance, this approach still obtains better results than previous methods in this setting. Last, for the first time, we tackle the combination of both few-shot tasks, on three challenging benchmarks for viewpoint estimation in the wild, ObjectNet3D, Pascal3D+ and Pix3D, showing very promising results. Yang Xiao 0009, Vincent Lepetit, Renaud Marlet |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | PIZZA: A Powerful Image-only Zero-Shot Zero-CAD Approach to 6 DoF TrackingabstractEstimating the relative pose of a new object without prior knowledge is a hard problem, while it is an ability very much needed in robotics and Augmented Reality. We present a method for tracking the 6D motion of objects in RGB video sequences when neither the training images nor the 3D geometry of the objects are available. In contrast to previous works, our method can therefore consider unknown objects in open world instantly, without requiring any prior information or a specific training phase. We consider two architectures, one based on two frames, and the other relying on a Transformer Encoder, which can exploit an arbitrary number of past frames. We train our architectures using only synthetic renderings with domain randomization. Our results on challenging datasets are on par with previous works that require much more information (training images of the target objects, 3D models, and/or depth data). Our source code is available at https://github.com/nv-nguyen/pizza. Van Nguyen Nguyen, Yuming Du, Yang Xiao 0007, Michaël Ramamonjisoa, Vincent Lepetit |
3DV | 5 |
| 2022 | Keypoint Transformer: Solving Joint Identification in Challenging Hands and Object Interactions for Accurate 3D Pose EstimationabstractWe propose a robust and accurate method for estimating the 3D poses of two hands in close interaction from a single color image. This is a very challenging problem, as large occlusions and many confusions between the joints may happen. State-of-the-art methods solve this problem by regressing a heatmap for each joint, which requires solving two problems simultaneously: localizing the joints and recognizing them. In this work, we propose to separate these tasks by relying on a CNN to first localize joints as 2D keypoints, and on self-attention between the CNN features at these keypoints to associate them with the corresponding hand joint. The resulting architecture, which we call “Keypoint Transformer”, is highly efficient as it achieves state-of-the-art performance with roughly half the number of model parameters on the InterHand2.6M dataset. We also show it can be easily extended to estimate the 3D pose of an object manipulated by one or two hands with high performance. Moreover, we created a new dataset of more than 75,000 images of two hands manipulating an object fully annotated in 3D and will make it publicly available. Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, Vincent Lepetit |
CVPR | 4 |
| 2022 | Templates for 3D Object Pose Estimation Revisited: Generalization to New Objects and Robustness to OcclusionsabstractWe present a method that can recognize new objects and estimate their 3D pose in RGB images even under partial occlusions. Our method requires neither a training phase on these objects nor real images depicting them, only their CAD models. It relies on a small set of training objects to learn local object representations, which allow us to locally match the input image to a set of “templates”, rendered images of the CAD models for the new objects. In contrast with the state-of-the-art methods, the new objects on which our method is applied can be very different from the training objects. As a result, we are the first to show generalization without retraining on the LINEMOD and Occlusion-LINEMOD datasets. Our analysis of the failure modes of previous template-based approaches further confirms the benefits of local features for template matching. We outperform the state-of-the-art template matching methods on the LINEMOD, Occlusion-LINEMOD and T-LESS datasets. Our source code and data are publicly available at https://github.com/nv-nguyen/template-pose. Van Nguyen Nguyen, Yinlin Hu, Yang Xiao 0009, Mathieu Salzmann, Vincent Lepetit |
CVPR | 5 |
| 2022 | MonteBoxFinder: Detecting and Filtering Primitives to Fit a Noisy Point Cloud
Michaël Ramamonjisoa, Sinisa Stekovic, Vincent Lepetit |
ECCV (28) | 3 |
| 2022 | Visual Correspondence Hallucination
Hugo Germain, Vincent Lepetit, Guillaume Bourmaud |
ICLR | 2 |
| 2022 | Multi-Finger Grasping Like HumansabstractRobots with multi-fingered grippers could perform advanced manipulation tasks for us if we were able to properly specify to them what to do. In this study, we take a step in that direction by making a robot grasp an object like a grasping demonstration performed by a human. We propose a novel optimization-based approach for transferring human grasp demonstrations to any multi-fingered grippers, which produces robotic grasps that mimic the human hand orientation and the contact area with the object, while alleviating interpenetration. Extensive experiments with the Allegro and BarrettHand grippers show that our method leads to grasps more similar to the human demonstration than existing approaches, without requiring any gripper-specific tuning. We confirm these findings through a user study and validate the applicability of our approach on a real robot. Yuming Du, Philippe Weinzaepfel, Vincent Lepetit, Romain Brégier |
IROS | 3 |
| 2022 | SCONE: Surface Coverage Optimization in Unknown Environments by Volumetric IntegrationabstractNext Best View computation (NBV) is a long-standing problem in robotics, and consists in identifying the next most informative sensor position(s) for reconstructing a 3D object or scene efficiently and accurately. Like most current methods, we consider NBV prediction from a depth sensor like Lidar systems. Learning-based methods relying on a volumetric representation of the scene are suitable for path planning, but have lower accuracy than methods using a surface-based representation. However, the latter do not scale well with the size of the scene and constrain the camera to a small number of poses. To obtain the advantages of both representations, we show that we can maximize surface metrics by Monte Carlo integration over a volumetric representation. In particular, we propose an approach, SCONE, that relies on two neural modules: The first module predicts occupancy probability in the entire volume of the scene. Given any new camera pose, the second module samples points in the scene based on their occupancy probability and leverages a self-attention mechanism to predict the visibility of the samples. Finally, we integrate the visibility to evaluate the gain in surface coverage for the new camera pose. NBV is selected as the pose that maximizes the gain in total surface coverage. Our method scales to large scenes and handles free camera motion: It takes as input an arbitrarily large point cloud gathered by a depth sensor as well as camera poses to predict NBV. We demonstrate our approach on a novel dataset made of large and complex 3D scenes. Antoine Guédon, Pascal Monasse, Vincent Lepetit |
NeurIPS | 3 |
| 2022 | Simultaneous completion and spatiotemporal grouping of corrupted motion tracksabstractAbstract Given an unordered list of 2D or 3D point trajectories corrupted by noise and partial observations, in this paper we introduce a framework to simultaneously recover the incomplete motion tracks and group the points into spatially and temporally coherent clusters. This advances existing work, which only addresses partial problems and without considering a unified and unsupervised solution. We cast this problem as a matrix completion one, in which point tracks are arranged into a matrix with the missing entries set as zeros. In order to perform the double clustering, the measurement matrix is assumed to be drawn from a dual union of spatiotemporal subspaces. The bases and the dimensionality for these subspaces, the affinity matrices used to encode the temporal and spatial clusters to which each point belongs, and the non-visible tracks, are then jointly estimated via augmented Lagrange multipliers in polynomial time. A thorough evaluation on incomplete motion tracks for multiple-object typologies shows that the accuracy of the matrix we recover compares favorably to that obtained with existing low-rank matrix completion methods, specially under noisy measurements. In addition, besides recovering the incomplete tracks, the point trajectories are directly grouped into different object instances, and a number of semantically meaningful temporal primitive actions are automatically discovered. Antonio Agudo, Vincent Lepetit, Francesc Moreno-Noguer |
Vis. Comput. | 2 |
| 2021 | Neural Reprojection Error: Merging Feature Learning and Camera Pose EstimationabstractAbsolute camera pose estimation is usually addressed by sequentially solving two distinct subproblems: First a feature matching problem that seeks to establish putative 2D-3D correspondences, and then a Perspective-n-Point problem that minimizes, w.r.t. the camera pose, the sum of so-called Reprojection Errors (RE). We argue that generating putative 2D-3D correspondences 1) leads to an important loss of information that needs to be compensated as far as possible, within RE, through the choice of a robust loss and the tuning of its hyperparameters and 2) may lead to an RE that conveys erroneous data to the pose estimator. In this paper, we introduce the Neural Reprojection Error (NRE) as a substitute for RE. NRE allows to rethink the camera pose estimation problem by merging it with the feature learning problem, hence leveraging richer information than 2D-3D correspondences and eliminating the need for choosing a robust loss and its hyperparameters. Thus NRE can be used as training loss to learn image descriptors tailored for pose estimation. We also propose a coarse-to-fine optimization method able to very efficiently minimize a sum of NRE terms w.r.t. the camera pose. We experimentally demonstrate that NRE is a good substitute for RE as it significantly improves both the robustness and the accuracy of the camera pose estimate while being computationally and memory highly efficient. From a broader point of view, we believe this new way of merging deep learning and 3D geometry may be useful in other computer vision applications. Source code and model weights will be made available at hugogermain.com/nre. Hugo Germain, Vincent Lepetit, Guillaume Bourmaud |
CVPR | 2 |
| 2021 | Monte Carlo Scene Search for 3D Scene UnderstandingabstractWe explore how a general AI algorithm can be used for 3D scene understanding to reduce the need for training data. More exactly, we propose a modification of the Monte Carlo Tree Search (MCTS) algorithm to retrieve objects and room layouts from noisy RGB-D scans. While MCTS was developed as a game-playing algorithm, we show it can also be used for complex perception problems. Our adapted MCTS algorithm has few easy-to-tune hyperparameters and can optimise general losses. We use it to optimise the posterior probability of objects and room layout hypotheses given the RGB-D data. This results in an analysis-by-synthesis approach that explores the solution space by rendering the current solution and comparing it to the RGB-D observations. To perform this exploration even more efficiently, we propose simple changes to the standard MCTS’ tree construction and exploration policy. We demonstrate our approach on the ScanNet dataset. Our method often retrieves configurations that are better than some manual annotations, especially on layouts. Shreyas Hampali, Sinisa Stekovic, Sayan Deb Sarkar, Chetan Srinivasa Kumar, Friedrich Fraundorfer, Vincent Lepetit |
CVPR | 6 |
| 2021 | Single Image Depth Prediction With Wavelet DecompositionabstractWe present a novel method for predicting accurate depths from monocular images with high efficiency. This optimal efficiency is achieved by exploiting wavelet decomposition, which is integrated in a fully differentiable encoder-decoder architecture. We demonstrate that we can reconstruct high-fidelity depth maps by predicting sparse wavelet coefficients.In contrast with previous works, we show that wavelet coefficients can be learned without direct supervision on coefficients. Instead we supervise only the final depth image that is reconstructed through the inverse wavelet transform. We additionally show that wavelet coefficients can be learned in fully self-supervised scenarios, without access to ground-truth depth. Finally, we apply our method to different state-of-the-art monocular depth estimation models, in each case giving similar or better results compared to the original model, while requiring less than half the multiply-adds in the decoder network. Michaël Ramamonjisoa, Michael Firman, Jamie Watson, Vincent Lepetit, Daniyar Turmukhambetov |
CVPR | 4 |
| 2021 | Back to the Feature: Learning Robust Camera Localization From Pixels To PoseabstractCamera pose estimation in known scenes is a 3D geometry task recently tackled by multiple learning algorithms. Many regress precise geometric quantities, like poses or 3D points, from an input image. This either fails to generalize to new viewpoints or ties the model parameters to a specific scene. In this paper, we go Back to the Feature: we argue that deep networks should focus on learning robust and invariant visual features, while the geometric estimation should be left to principled algorithms. We introduce PixLoc, a scene-agnostic neural network that estimates an accurate 6-DoF pose from an image and a 3D model. Our approach is based on the direct alignment of multiscale deep features, casting camera localization as metric learning. PixLoc learns strong data priors by end-to-end training from pixels to pose and exhibits exceptional generalization to new scenes by separating model parameters and scene geometry. The system can localize in large environments given coarse pose priors but also improve the accuracy of sparse feature matching by jointly refining keypoints and poses with little overhead. The code will be publicly available at github.com/cvg/pixloc. Paul-Edouard Sarlin, Ajaykumar Unagar, Måns Larsson, Hugo Germain, Carl Toft, Viktor Larsson, Marc Pollefeys, Vincent Lepetit, Lars Hammarstrand, Fredrik Kahl, Torsten Sattler |
CVPR | 8 |
| 2021 | Learning to Better Segment Objects from Unseen Classes with Unlabeled VideosabstractThe ability to localize and segment objects from unseen classes would open the door to new applications, such as autonomous object learning in active vision. Nonetheless, improving the performance on unseen classes requires additional training data, while manually annotating the objects of the unseen classes can be labor-extensive and expensive. In this paper, we explore the use of unlabeled video sequences to automatically generate training data for objects of unseen classes. It is in principle possible to apply existing video segmentation methods to unlabeled videos and automatically obtain object masks, which can then be used as a training set even for classes with no manual labels available. However, our experiments show that these methods do not perform well enough for this purpose. We therefore introduce a Bayesian method that is specifically designed to automatically create such a training set: Our method starts from a set of object proposals and relies on (non-realistic) analysis-by-synthesis to select the correct ones by performing an efficient optimization over all the frames simultaneously. Through extensive experiments, we show that our method can generate a high-quality training set which significantly boosts the performance of segmenting objects of unseen classes. We thus believe that our method could open the door for open-world instance segmentation by exploiting abundant Internet videos. Yuming Du, Yang Xiao 0007, Vincent Lepetit |
ICCV | 3 |
| 2021 | MonteFloor: Extending MCTS for Reconstructing Accurate Large-Scale Floor PlansabstractWe propose a novel method for reconstructing floor plans from noisy 3D point clouds. Our main contribution is a principled approach that relies on the Monte Carlo Tree Search (MCTS) algorithm to maximize a suitable objective function efficiently despite the complexity of the problem. Like previous work, we first project the input point cloud to a top view to create a density map and extract room proposals from it. Our method selects and optimizes the polygonal shapes of these room proposals jointly to fit the density map and outputs an accurate vectorized floor map even for large complex scenes. To do this, we adapt MCTS, an algorithm originally designed to learn to play games, to select the room proposals by maximizing an objective function combining the fitness with the density map as predicted by a deep network and regularizing terms on the room shapes. We also introduce a refinement step to MCTS that adjusts the shape of the room proposals. For this step, we propose a novel differentiable method for rendering the polygonal shapes of these proposals. We evaluate our method on the recent and challenging Structured3D and Floor-SP datasets and show a significant improvement over the state-of-the-art, without imposing any hard constraints nor assumptions on the floor plan configurations. Sinisa Stekovic, Mahdi Rad, Friedrich Fraundorfer, Vincent Lepetit |
ICCV | 4 |
| 2020 | 3D Object Detection and Pose Estimation of Unseen Objects in Color Images with Local Surface Embeddings
Giorgia Pitteri, Aurélie Bugeau, Slobodan Ilic, Vincent Lepetit |
ACCV (1) | 4 |
| 2020 | HOnnotate: A Method for 3D Annotation of Hand and Object PosesabstractWe propose a method for annotating images of a hand manipulating an object with the 3D poses of both the hand and the object, together with a dataset created using this method. Our motivation is the current lack of annotated real images for this problem, as estimating the 3D poses is challenging, mostly because of the mutual occlusions between the hand and the object. To tackle this challenge, we capture sequences with one or several RGB-D cameras and jointly optimize the 3D hand and object poses over all the frames simultaneously. This method allows us to automatically annotate each frame with accurate estimates of the poses, despite large mutual occlusions. With this method, we created HO-3D, the first markerless dataset of color images with 3D annotations for both the hand and object. This dataset is currently made of 77,558 frames, 68 sequences, 10 persons, and 10 objects. Using our dataset, we develop a single RGB image-based method to predict the hand pose when interacting with objects under severe occlusions and show it generalizes to objects not seen in the dataset. Shreyas Hampali, Mahdi Rad, Markus Oberweger, Vincent Lepetit |
CVPR | 4 |
| 2020 | Predicting Sharp and Accurate Occlusion Boundaries in Monocular Depth Estimation Using Displacement FieldsabstractCurrent methods for depth map prediction from monocular images tend to predict smooth, poorly localized contours for the occlusion boundaries in the input image. This is unfortunate as occlusion boundaries are important cues to recognize objects, and as we show, may lead to a way to discover new objects from scene reconstruction. To improve predicted depth maps, recent methods rely on various forms of filtering or predict an additive residual depth map to refine a first estimate. We instead learn to predict, given a depth map predicted by some reconstruction method, a 2D displacement field able to re-sample pixels around the occlusion boundaries into sharper reconstructions. Our method can be applied to the output of any depth estimation method, in an end-to-end trainable fashion. For evaluation, we manually annotated the occlusion boundaries in all the images in the test split of popular NYUv2-Depth dataset. We show that our approach improves the localization of occlusion boundaries for all state-of-the-art monocular depth estimation methods that we could evaluate, without degrading the depth accuracy for the rest of the images. Michaël Ramamonjisoa, Yuming Du, Vincent Lepetit |
CVPR | 3 |
| 2020 | Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation Under Hand-Object Interaction
Anil Armagan, Guillermo Garcia-Hernando, Seungryul Baek, Shreyas Hampali, Mahdi Rad, Shipeng Xie, Mingxiu Chen, Boshen Zhang, Fu Xiong, Yang Xiao 0007, Zhiguo Cao 0001, Junsong Yuan 0001, Pengfei Ren 0001, Weiting Huang, Haifeng Sun 0001, Marek Hrúz, Jakub Kanis, Zdenek Krnoul, Qingfu Wan, Shile Li, Linlin Yang 0001, Dongheui Lee, Angela Yao, Weiguo Zhou, Sijia Mei, Adrian Spurr, Umar Iqbal 0001, Pavlo Molchanov 0001, Philippe Weinzaepfel, Romain Brégier, Grégory Rogez, Vincent Lepetit, Tae-Kyun Kim 0001 |
ECCV (23) | 34 |
| 2020 | S2DNet: Learning Image Features for Accurate Sparse-to-Dense Matching
Hugo Germain, Guillaume Bourmaud, Vincent Lepetit |
ECCV (3) | 3 |
| 2020 | Geometric Correspondence Fields: Learned Differentiable Rendering for 3D Pose Refinement in the Wild
Alexander Grabner, Yaming Wang, Peizhao Zhang, Peihong Guo, Tong Xiao 0003, Peter Vajda, Peter M. Roth, Vincent Lepetit |
ECCV (16) | 8 |
| 2020 | General 3D Room Layout from a Single View by Render-and-Compare
Sinisa Stekovic, Shreyas Hampali, Mahdi Rad, Sayan Deb Sarkar, Friedrich Fraundorfer, Vincent Lepetit |
ECCV (16) | 6 |
| 2020 | Smart Hypothesis Generation for Efficient and Robust Room Layout EstimationabstractWe propose a novel method to efficiently estimate the spatial layout of a room from a single monocular RGB image. As existing approaches based on low-level feature extraction, followed by a vanishing point estimation are very slow and often unreliable in realistic scenarios, we build on semantic segmentation of the input image. To obtain better segmentations, we introduce a robust, accurate and very efficient hypothesize-and-test scheme. The key idea is to use three segmentation hypotheses, each based on a different number of visible walls. For each hypothesis, we predict the image locations of the room corners and select the hypothesis for which the layout estimated from the room corners is consistent with the segmentation. We demonstrate the efficiency and robustness of our method on three challenging benchmark datasets, where we significantly outperform the state-of-the-art. Martin Hirzer, Peter M. Roth, Vincent Lepetit |
WACV | 3 |
| 2020 | Casting Geometric Constraints in Semantic Segmentation as Semi-Supervised LearningabstractWe propose a simple yet effective method to learn to segment new indoor scenes from video frames: State-of- the-art methods trained on one dataset, even as large as the SUNRGB-D dataset, can perform poorly when applied to images that are not part of the dataset, because of the dataset bias, a common phenomenon in computer vision. To make semantic segmentation more useful in practice, one can exploit geometric constraints. Our main contribution is to show that these constraints can be cast conveniently as semi-supervised terms, which enforce the fact that the same class should be predicted for the projections of the same 3D location in different images. This is interesting as we can exploit general existing techniques de- veloped for semi-supervised learning to efficiently incorporate the constraints. We show that this approach can efficiently and accurately learn to segment target sequences of ScanNet and our own target sequences using only annotations from SUNRGB-D, and geometric relations between the video frames of target sequences. Sinisa Stekovic, Friedrich Fraundorfer, Vincent Lepetit |
WACV | 3 |
| 2020 | ALCN: Adaptive Local Contrast Normalization
Mahdi Rad, Peter M. Roth, Vincent Lepetit |
Comput. Vis. Image Underst. | 3 |
| 2020 | Generalized Feedback Loop for Joint Hand-Object Pose EstimationabstractWe propose an approach to estimating the 3D pose of a hand, possibly handling an object, given a depth image. We show that we can correct the mistakes made by a Convolutional Neural Network trained to predict an estimate of the 3D pose by using a feedback loop. The components of this feedback loop are also Deep Networks, optimized using training data. This approach can be generalized to a hand interacting with an object. Therefore, we jointly estimate the 3D pose of the hand and the 3D pose of the object. Our approach performs en-par with state-of-the-art methods for 3D hand pose estimation, and outperforms state-of-the-art methods for joint hand-object pose estimation when using depth images only. Also, our approach is efficient as our implementation runs in real-time on a single GPU. Markus Oberweger, Paul Wohlhart, Vincent Lepetit |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Sparse-to-Dense Hypercolumn Matching for Long-Term Visual LocalizationabstractWe propose a novel approach to feature point matching, suitable for robust and accurate outdoor visual localization in long-term scenarios. Given a query image, we first match it against a database of registered reference images, using recent retrieval techniques. This gives us a first estimate of the camera pose. To refine this estimate, like previous approaches, we match 2D points across the query image and the retrieved reference image. This step, however, is prone to fail as it is still very difficult to detect and match sparse feature points across images captured in potentially very different conditions. Our key contribution is to show that we need to extract sparse feature points only in the retrieved reference image: We then search for the corresponding 2D locations in the query image exhaustively. This search can be performed efficiently using convolutional operations, and robustly by using hypercolumn descriptors, i.e. image features computed for retrieval. We refer to this method as 'Sparse-to-Dense Hypercolumn Matching'. Because we know the 3D locations of the sparse feature points in the reference images thanks to an offline reconstruction stage, it is then possible to accurately estimate the camera pose from these matches. Our experiments show that this method allows us to outperform the state-of-the-art on several challenging outdoor datasets. Hugo Germain, Guillaume Bourmaud, Vincent Lepetit |
3DV | 3 |
| 2019 | Location Field Descriptors: Single Image 3D Model Retrieval in the WildabstractWe present Location Field Descriptors, a novel approach for single image 3D model retrieval in the wild. In contrast to previous methods that directly map 3D models and RGB images to an embedding space, we establish a common low-level representation in the form of location fields from which we compute pose invariant 3D shape descriptors. Location fields encode correspondences between 2D pixels and 3D surface coordinates and, thus, explicitly capture 3D shape and 3D pose information without appearance variations which are irrelevant for the task. This early fusion of 3D models and RGB images results in three main advantages: First, the bottleneck location field prediction acts as a regularizer during training. Second, major parts of the system benefit from training on a virtually infinite amount of synthetic data. Finally, the predicted location fields are visually interpretable and unblackbox the system. We evaluate our proposed approach on three challenging real-world datasets (Pix3D, Comp, and Stanford) with different object categories and significantly outperform the state-of-the-art by up to 20% absolute in multiple 3D retrieval metrics. Alexander Grabner, Peter M. Roth, Vincent Lepetit |
3DV | 3 |
| 2019 | On Object Symmetries and 6D Pose Estimation from ImagesabstractObjects with symmetries are common in our daily life and in industrial contexts, but are often ignored in the recent literature on 6D pose estimation from images. In this paper, we study in an analytical way the link between the symmetries of a 3D object and its appearance in images. We explain why symmetrical objects can be a challenge when training machine learning algorithms that aim at estimating their 6D pose from images. We propose an efficient and simple solution that relies on the normalization of the pose rotation. Our approach is general and can be used with any 6D pose estimation algorithm. Moreover, our method is also beneficial for objects that are 'almost symmetrical', i.e. objects for which only a detail breaks the symmetry. We validate our approach within a Faster-RCNN framework on a synthetic dataset made with objects from the T-Less dataset, which exhibit various types of symmetries, as well as real sequences from T-Less. Giorgia Pitteri, Michaël Ramamonjisoa, Slobodan Ilic, Vincent Lepetit |
3DV | 4 |
| 2019 | Speed Invariant Time Surface for Learning to Detect Corner Points With Event-Based CamerasabstractWe propose a learning approach to corner detection for event-based cameras that is stable even under fast and abrupt motions. Event-based cameras offer high temporal resolution, power efficiency, and high dynamic range. However, the properties of event-based data are very different compared to standard intensity images, and simple extensions of corner detection methods designed for these images do not perform well on event-based data. We first introduce an efficient way to compute a time surface that is invariant to the speed of the objects. We then show that we can train a Random Forest to recognize events generated by a moving corner from our time surface. Random Forests are also extremely efficient, and therefore a good choice to deal with the high capture frequency of event-based cameras-our implementation processes up to 1.6Mev/s on a single CPU. Thanks to our time surface formulation and this learning approach, our method is significantly more robust to abrupt changes of direction of the corners compared to previous ones. Our method also naturally assigns a confidence score for the corners, which can be useful for postprocessing. Moreover, we introduce a high-resolution dataset suitable for quantitative evaluation and comparison of corner detection methods for event-based cameras. We call our approach SILC, for Speed Invariant Learned Corners, and compare it to the state-of-the-art with extensive experiments, showing better performance. Jacques Manderscheid, Amos Sironi, Nicolas Bourdis, Davide Migliore, Vincent Lepetit |
CVPR | 5 |
| 2019 | GP2C: Geometric Projection Parameter Consensus for Joint 3D Pose and Focal Length Estimation in the WildabstractWe present a joint 3D pose and focal length estimation approach for object categories in the wild. In contrast to previous methods that predict 3D poses independently of the focal length or assume a constant focal length, we explicitly estimate and integrate the focal length into the 3D pose estimation. For this purpose, we combine deep learning techniques and geometric algorithms in a two-stage approach: First, we estimate an initial focal length and establish 2D-3D correspondences from a single RGB image using a deep network. Second, we recover 3D poses and refine the focal length by minimizing the reprojection error of the predicted correspondences. In this way, we exploit the geometric prior given by the focal length for 3D pose estimation. This results in two advantages: First, we achieve significantly improved 3D translation and 3D pose accuracy compared to existing methods. Second, our approach finds a geometric consensus between the individual projection parameters, which is required for precise 2D-3D alignment. We evaluate our proposed approach on three challenging real-world datasets (Pix3D, Comp, and Stanford) with different object categories and significantly outperform the state-of-the-art by up to 20% absolute in multiple different metrics. Alexander Grabner, Peter M. Roth, Vincent Lepetit |
ICCV | 3 |
| 2019 | AssemblyNet: A Novel Deep Decision-Making Process for Whole Brain MRI Segmentation
Pierrick Coupé, Boris Mansencal, Michaël Clément, Rémi Giraud, Baudouin Denis de Senneville, Vinh-Thong Ta 0002, Vincent Lepetit, José V. Manjón |
MICCAI (3) | 7 |
| 2018 | Domain Transfer for 3D Pose Estimation from Color Images Without Manual Annotations
Mahdi Rad, Markus Oberweger, Vincent Lepetit |
ACCV (5) | 3 |
| 2018 | 3D Pose Estimation and 3D Model Retrieval for Objects in the WildabstractWe propose a scalable, efficient and accurate approach to retrieve 3D models for objects in the wild. Our contribution is twofold. We first present a 3D pose estimation approach for object categories which significantly outperforms the state-of-the-art on Pascal3D+. Second, we use the estimated pose as a prior to retrieve 3D models which accurately represent the geometry of objects in RGB images. For this purpose, we render depth images from 3D models under our predicted pose and match learned image descriptors of RGB images against those of rendered depth images using a CNN-based multi-view metric learning approach. In this way, we are the first to report quantitative results for 3D model retrieval on Pascal3D+, where our method chooses the same models as human annotators for 50% of the validation images on average. In addition, we show that our method, which was trained purely on Pascal3D+, retrieves rich and accurate 3D models from ShapeNet given RGB images of objects in the wild. Alexander Grabner, Peter M. Roth, Vincent Lepetit |
CVPR | 3 |
| 2018 | Geometry-Aware Network for Non-Rigid Shape Prediction From a Single ViewabstractWe propose a method for predicting the 3D shape of a deformable surface from a single view. By contrast with previous approaches, we do not need a pre-registered template of the surface, and our method is robust to the lack of texture and partial occlusions. At the core of our approach is a geometry-aware deep architecture that tackles the problem as usually done in analytic solutions: first perform 2D detection of the mesh and then estimate a 3D shape that is geometrically consistent with the image. We train this architecture in an end-to-end manner using a large dataset of synthetic renderings of shapes under different levels of deformation, material properties, textures and lighting conditions. We evaluate our approach on a test split of this dataset and available real benchmarks, consistently improving state-of-the-art solutions with a significantly lower computational time. Albert Pumarola, Antonio Agudo, Lorenzo Porzi, Alberto Sanfeliu, Vincent Lepetit, Francesc Moreno-Noguer |
CVPR | 5 |
| 2018 | Feature Mapping for Learning Fast and Accurate 3D Pose Inference From Synthetic ImagesabstractWe propose a simple and efficient method for exploiting synthetic images when training a Deep Network to predict a 3D pose from an image. The ability of using synthetic images for training a Deep Network is extremely valuable as it is easy to create a virtually infinite training set made of such images, while capturing and annotating real images can be very cumbersome. However, synthetic images do not resemble real images exactly, and using them for training can result in suboptimal performance. It was recently shown that for exemplar-based approaches, it is possible to learn a mapping from the exemplar representations of real images to the exemplar representations of synthetic images. In this paper, we show that this approach is more general, and that a network can also be applied after the mapping to infer a 3D pose: At run-time, given a real image of the target object, we first compute the features for the image, map them to the feature space of synthetic images, and finally use the resulting features as input to another network which predicts the 3D pose. Since this network can be trained very effectively by using synthetic images, it performs very well in practice, and inference is faster and more accurate than with an exemplar-based approach. We demonstrate our approach on the LINEMOD dataset for 3D object pose estimation from color images, and the NYU dataset for 3D hand pose estimation from depth maps. We show that it allows us to outperform the state-of-the-art on both datasets. Mahdi Rad, Markus Oberweger, Vincent Lepetit |
CVPR | 3 |
| 2018 | Learning to Find Good CorrespondencesabstractWe develop a deep architecture to learn to find good correspondences for wide-baseline stereo. Given a set of putative sparse matches and the camera intrinsics, we train our network in an end-to-end fashion to label the correspondences as inliers or outliers, while simultaneously using them to recover the relative pose, as encoded by the essential matrix. Our architecture is based on a multi-layer perceptron operating on pixel coordinates rather than directly on the image, and is thus simple and small. We introduce a novel normalization technique, called Context Normalization, which allows us to process each data point separately while embedding global information in it, and also makes the network invariant to the order of the correspondences. Our experiments on multiple challenging datasets demonstrate that our method is able to drastically improve the state of the art with little training data. Kwang Moo Yi, Eduard Trulls, Yuki Ono, Vincent Lepetit, Mathieu Salzmann, Pascal Fua |
CVPR | 4 |
| 2018 | Making Deep Heatmaps Robust to Partial Occlusions for 3D Object Pose Estimation
Markus Oberweger, Mahdi Rad, Vincent Lepetit |
ECCV (15) | 3 |
| 2018 | Efficient Physics-Based Implementation for Realistic Hand-Object Interaction in Virtual RealityabstractWe propose an efficient physics-based method for dexterous `real hand' - `virtual object' interaction in Virtual Reality environments. Our method is based on the Coulomb friction model, and we show how to efficiently implement it in a commodity VR engine for realtime performance. This model enables very convincing simulations of many types of actions such as pushing, pulling, grasping, or even dexterous manipulations such as spinning objects between fingers without restrictions on the objects' shapes or hand poses. Because it is an analytic model, we do not require any prerecorded data, in contrast to previous methods. For the evaluation of our method, we conduction a pilot study that shows that our method is perceived more realistic and natural, and allows for more diverse interactions. Further, we evaluate the computational complexity of our method to show real-time performance in VR environments. Markus Höll, Markus Oberweger, Clemens Arth, Vincent Lepetit |
VR | 4 |
| 2018 | Editorial for ACCV'16 award papers
Shang-Hong Lai, Vincent Lepetit, Ko Nishino, Yoichi Sato 0001 |
Comput. Vis. Image Underst. | 2 |
| 2018 | Learning Latent Representations of 3D Human Pose with Deep Neural Networks
Isinsu Katircioglu, Bugra Tekin, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
Int. J. Comput. Vis. | 4 |
| 2018 | Robust 3D Object Tracking from Monocular Images Using Stable PartsabstractWe present an algorithm for estimating the pose of a rigid object in real-time under challenging conditions. Our method effectively handles poorly textured objects in cluttered, changing environments, even when their appearance is corrupted by large occlusions, and it relies on grayscale images to handle metallic environments on which depth cameras would fail. As a result, our method is suitable for practical Augmented Reality applications including industrial environments. At the core of our approach is a novel representation for the 3D pose of object parts: We predict the 3D pose of each part in the form of the 2D projections of a few control points. The advantages of this representation is three-fold: We can predict the 3D pose of the object even when only one part is visible; when several parts are visible, we can easily combine them to compute a better pose of the object; the 3D pose we obtain is usually very accurate, even when only few parts are visible. We show how to use this representation in a robust 3D tracking framework. In addition to extensive comparisons with the state-of-the-art, we demonstrate our method on a practical Augmented Reality application for maintenance assistance in the ATLAS particle detector at CERN. Alberto Crivellaro, Mahdi Rad, Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2017 | Accurate Camera Registration in Urban Environments Using High-Level Feature Matching
Anil Armagan, Martin Hirzer, Peter M. Roth, Vincent Lepetit |
BMVC | 4 |
| 2017 | Efficient 3D Tracking in Urban Environments with Semantic Segmentation
Martin Hirzer, Peter M. Roth, Vincent Lepetit |
BMVC | 3 |
| 2017 | Deep Learning for 3D Localization
Vincent Lepetit |
BMVC | 1 |
| 2017 | Adaptive Local Contrast Normalization for Robust Object Detection and Pose Estimation
Mahdi Rad, Vincent Lepetit, Peter M. Roth |
BMVC | 2 |
| 2017 | Learning to Align Semantic Segmentation and 2.5D Maps for GeolocalizationabstractWe present an efficient method for geolocalization in urban environments starting from a coarse estimate of the location provided by a GPS and using a simple untextured 2.5D model of the surrounding buildings. Our key contribution is a novel efficient and robust method to optimize the pose: We train a Deep Network to predict the best direction to improve a pose estimate, given a semantic segmentation of the input image and a rendering of the buildings from this estimate. We then iteratively apply this CNN until converging to a good pose. This approach avoids the use of reference images of the surroundings, which are difficult to acquire and match, while 2.5D models are broadly available. We can therefore apply it to places unseen during training. Anil Armagan, Martin Hirzer, Peter M. Roth, Vincent Lepetit |
CVPR | 4 |
| 2017 | BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using DepthabstractWe introduce a novel method for 3D object detection and pose estimation from color images only. We first use segmentation to detect the objects of interest in 2D even in presence of partial occlusions and cluttered background. By contrast with recent patch-based methods, we rely on a “holistic” approach: We apply to the detected objects a Convolutional Neural Network (CNN) trained to predict their 3D poses in the form of 2D projections of the corners of their 3D bounding boxes. This, however, is not sufficient for handling objects from the recent T-LESS dataset: These objects exhibit an axis of rotational symmetry, and the similarity of two images of such an object under two different poses makes training the CNN challenging. We solve this problem by restricting the range of poses used for training, and by introducing a classifier to identify the range of a pose at run-time before estimating it. We also use an optional additional step that refines the predicted poses. We improve the state-of-the-art on the LINEMOD dataset from 73.7% [2] to 89.3% of correctly registered RGB frames. We are also the first to report results on the Occlusion dataset [1] using color images only. We obtain 54% of frames passing the Pose 6D criterion on average on several sequences of the T-LESS dataset, compared to the 67% of the state-of-the-art [10] on the same sequences which uses both color and depth. The full approach is also scalable, as a single network can be trained for multiple objects simultaneously. Mahdi Rad, Vincent Lepetit |
ICCV | 2 |
| 2017 | Learning Lightprobes for Mixed Reality IlluminationabstractThis paper presents the first photometric registration pipeline for Mixed Reality based on high quality illumination estimation using convolutional neural networks (CNNs). For easy adaptation and deployment of the system, we train the CNNs using purely synthetic images and apply them to real image data. To keep the pipeline accurate and efficient, we propose to fuse the light estimation results from multiple CNN instances and show an approach for caching estimates over time. For optimal performance, we furthermore explore multiple strategies for the CNN training. Experimental results show that the proposed method yields highly accurate estimates for photo-realistic augmentations. David Mandl, Kwang Moo Yi, Peter Mohr, Peter M. Roth, Pascal Fua, Vincent Lepetit, Dieter Schmalstieg, Denis Kalkofen |
ISMAR | 6 |
| 2017 | Pose-specific non-linear mappings in feature space towards multiview facial expression recognition
Mahdi Jampour, Vincent Lepetit, Thomas Mauthner, Horst Bischof |
Image Vis. Comput. | 2 |
| 2017 | Detecting Flying Objects Using a Single Moving CameraabstractWe propose an approach for detecting flying objects such as Unmanned Aerial Vehicles (UAVs) and aircrafts when they occupy a small portion of the field of view, possibly moving against complex backgrounds, and are filmed by a camera that itself moves. We argue that solving such a difficult problem requires combining both appearance and motion cues. To this end we propose a regression-based approach for object-centric motion stabilization of image patches that allows us to achieve effective classification on spatio-temporal image cubes and outperform state-of-the-art techniques. As this problem has not yet been extensively studied, no test datasets are publicly available. We therefore built our own, both for UAVs and aircrafts, and will make them publicly available so they can be used to benchmark future flying object detection and collision avoidance algorithms. Artem Rozantsev, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Structured Prediction of 3D Human Pose with Deep Neural Networks
Bugra Tekin, Isinsu Katircioglu, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
BMVC | 4 |
| 2016 | Efficiently Creating 3D Training Data for Fine Hand Pose EstimationabstractWhile many recent hand pose estimation methods critically rely on a training set of labelled frames, the creation of such a dataset is a challenging task that has been overlooked so far. As a result, existing datasets are limited to a few sequences and individuals, with limited accuracy, and this prevents these methods from delivering their full potential. We propose a semi-automated method for efficiently and accurately labeling each frame of a hand depth video with the corresponding 3D locations of the joints: The user is asked to provide only an estimate of the 2D reprojections of the visible joints in some reference frames, which are automatically selected to minimize the labeling work by efficiently optimizing a sub-modular loss function. We then exploit spatial, temporal, and appearance constraints to retrieve the full 3D poses of the hand over the complete sequence. We show that this data can be used to train a recent state-of-the-art hand pose estimation method, leading to increased accuracy. Markus Oberweger, Gernot Riegler, Paul Wohlhart, Vincent Lepetit |
CVPR | 4 |
| 2016 | Direct Prediction of 3D Body Poses from Motion Compensated SequencesabstractWe propose an efficient approach to exploiting motion information from consecutive frames of a video sequence to recover the 3D pose of people. Previous approaches typically compute candidate poses in individual frames and then link them in a post-processing step to resolve ambiguities. By contrast, we directly regress from a spatio-temporal volume of bounding boxes to a 3D pose in the central frame. We further show that, for this approach to achieve its full potential, it is essential to compensate for the motion in consecutive frames so that the subject remains centered. This then allows us to effectively overcome ambiguities and improve upon the state-of-the-art by a large margin on the Human3.6m, HumanEva, and KTH Multiview Football 3D human pose estimation benchmarks. Bugra Tekin, Artem Rozantsev, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2016 | Learning to Assign Orientations to Feature PointsabstractWe show how to train a Convolutional Neural Network to assign a canonical orientation to feature points given an image patch centered on the feature point. Our method improves feature point matching upon the state-of-the art and can be used in conjunction with any existing rotation sensitive descriptors. To avoid the tedious and almost impossible task of finding a target orientation to learn, we propose to use Siamese networks which implicitly find the optimal orientations during training. We also propose a new type of activation function for Neural Networks that generalizes the popular ReLU, maxout, and PReLU activation functions. This novel activation performs better for our task. We validate the effectiveness of our method extensively with four existing datasets, including two non-planar datasets, as well as our own dataset. We show that we outperform the state-of-the-art without the need of retraining for each dataset. Kwang Moo Yi, Yannick Verdie, Pascal Fua, Vincent Lepetit |
CVPR | 4 |
| 2016 | Going Further with Point Pair Features
Stefan Hinterstoißer, Vincent Lepetit, Naresh Rajkumar, Kurt Konolige |
ECCV (3) | 2 |
| 2016 | LIFT: Learned Invariant Feature Transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, Pascal Fua |
ECCV (6) | 3 |
| 2016 | Vision-based Unmanned Aerial Vehicle detection and tracking for sense and avoid systemsabstractWe propose an approach for on-line detection of small Unmanned Aerial Vehicles (UAVs) and estimation of their relative positions and velocities in the 3D environment from a single moving camera in the context of sense and avoid systems. This problem is challenging both from a detection point of view, as there are no markers on the targets available, and from a tracking perspective, due to misdetection and false positives. Furthermore, the methods need to be computationally light, despite the complexity of computer vision algorithms, to be used on UAVs with limited payload. To address these issues we propose a multi-staged framework that incorporates fast object detection using an AdaBoost-based approach, coupled with an on-line visual-based tracking algorithm and a recent sensor fusion and state estimation method. Our framework allows for achieving real-time performance with accurate object detection and tracking without any need of markers and customized, high-performing hardware resources. Krishna Raj Sapkota, Steven Roelofsen, Artem Rozantsev, Vincent Lepetit, Denis Gillet, Pascal Fua, Alcherio Martinoli |
IROS | 4 |
| 2016 | Automated Age Estimation from Hand MRI Volumes Using Deep Learning
Darko Stern, Christian Payer, Vincent Lepetit, Martin Urschler |
MICCAI (2) | 3 |
| 2016 | Multiscale Centerline DetectionabstractFinding the centerline and estimating the radius of linear structures is a critical first step in many applications, ranging from road delineation in 2D aerial images to modeling blood vessels, lung bronchi, and dendritic arbors in 3D biomedical image stacks. Existing techniques rely either on filters designed to respond to ideal cylindrical structures or on classification techniques. The former tend to become unreliable when the linear structures are very irregular while the latter often has difficulties distinguishing centerline locations from neighboring ones, thus losing accuracy. We solve this problem by reformulating centerline detection in terms of a regression problem. We first train regressors to return the distances to the closest centerline in scale-space, and we apply them to the input images or volumes. The centerlines and the corresponding scale then correspond to the regressors local maxima, which can be easily identified. We show that our method outperforms state-of-the-art techniques for various 2D and 3D datasets. Moreover, our approach is very generic and also performs well on contour detection. We show an improvement above recent contour detection algorithms on the BSDS500 dataset. Amos Sironi, Engin Türetken, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Leaf Segmentation under Loosely Controlled ConditionsabstractExtracting accurately the shape of a leaf is a crucial step in image-based plant identification systems. The partial or total absence of textures on leaf surface and the high color variability of leaves belonging to same species make shape as the main recognition element. For such reason, leaf segmentation plays a decisive role in the leaf recognition process. Even though many general segmentation methods have been proposed in the last decades, leaf segmentation presents specific challenges. In particular, a pixel-level precision is required in order to highlight fine scale boundary structures and discriminate similar global shapes. Moreover, even if the input image can typically be taken in controlled conditions, where the leaf is the only visible object over a white background, the user taking the picture is not necessarily an expert and the conditions are often not ideal: the leaf exhibits specular reflections, casts shadows, the background is never exactly white and is usually non-uniform, and the image can be blurry. Recently, a solution for the problem at hand has been proposed in [1] where leaf segmentation is carried out by estimating the probability distribution of foreground and background pixels. However, several drawbacks appear in this formulation due to challenging leaves like pine needles, false positives detection related to shadows and false negatives detection related to specularities. Prior distributions and post-processing operations are employed to tackle such problem, with the risk of hurting the final leaf shape. In this paper we introduce a new solution by training a pixel-wise classifier [3] that learns filter responses associated to background and foreground regions in images of leaves. Our classifier is trained by selecting positive (leaf) and negative (non-leaf) feature samples that lie on the neighboorhod of the leaf boundary thus focusing learning only on sensitive pixels. Such classifier is then applied to each pixel location of a given unknown test image. This provides a score map that we then threshold using two different thresholds to detect pixels that belong to foreground and background with a high level of probability. With these pixels at hand we initialize an EM algorithm with a good initial estimate of foreground and background cluster parameters in the saturation-value color space, differently from [1] which has to initialize the EM segmentation with the same values for all the images. The other difference with [1] is that we can consider as unlabeled data only the pixels that are in the neighborhood of the detected leaf boundary. This allows to keep focusing on segmenting correctly the pixels around the leaf boundary, and in practice it is enough to get a good segmentation of the other pixels, which are easier to classify. For evaluation we use the Leafsnap Field image dataset publicly available online [1] where different leaves of different species are acquired against solid background and variable light conditions thus simulating typical images that a user could provide for plant recognition. To train our pixel-wise classifier we randomly select one image for each species and we manually produce segmentation and thicker contours to discriminate between positive and negative training samples placed in the neighborhood of boundary. Since segmentation ground truth is not available and its manual production for thousands of images would require an inestimable amount of time, we considered a subset of the original Field dataset. Our testing set is made of 300 images: 150 images for which the EM approach of [1] performs already well thus producing faithful segmentation in accordance with the leaf shape plus 150 more challenging images for which EM partially or totally fails. The general behavior of different methods can be qualitatively appreciated looking at Fig. 1 where results returned by Leafsnap, Leafsnap without post-processing (marked with *), GrabCut and our method are reported. As the reader can see comparing ground truth details with real segmentations, it is confirmed that post-processing hurts quality of (a) Leaf image (b) Ground truth Simone Buoncompagni, Dario Maio, Vincent Lepetit |
BMVC | 3 |
| 2015 | Hashmod: A Hashing Method for Scalable 3D Object DetectionabstractWe present a scalable method for detecting objects and estimating their 3D poses in RGB-D data. To this end, we rely on an efficient representation of object views and employ hashing techniques to match these views against the input frame in a scalable way. While a similar approach already exists for 2D detection, we show how to extend it to estimate the 3D pose of the detected objects. In particular, we explore different hashing strategies and identify the one which is more suitable to our problem. We show empirically that the complexity of our method is sublinear with the number of objects and we enable detection and pose estimation of many 3D objects with high accuracy while outperforming the state-of-the-art in terms of runtime. Wadim Kehl, Federico Tombari, Nassir Navab, Slobodan Ilic, Vincent Lepetit |
BMVC | 5 |
| 2015 | Flying objects detection from a single moving cameraabstractWe propose an approach to detect flying objects such as UAVs and aircrafts when they occupy a small portion of the field of view, possibly moving against complex backgrounds, and are filmed by a camera that itself moves. Artem Rozantsev, Vincent Lepetit, Pascal Fua |
CVPR | 2 |
| 2015 | TILDE: A Temporally Invariant Learned DEtectorabstractWe introduce a learning-based approach to detect repeatable keypoints under drastic imaging changes of weather and lighting conditions to which state-of-the-art keypoint detectors are surprisingly sensitive. We first identify good keypoint candidates in multiple training images taken from the same viewpoint. We then train a regressor to predict a score map whose maxima are those points so that they can be found by simple non-maximum suppression. As there are no standard datasets to test the influence of these kinds of changes, we created our own, which we will make publicly available. We will show that our method significantly outperforms the state-of-the-art methods in such challenging conditions, while still achieving state-of-the-art performance on untrained standard datasets. Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
CVPR | 4 |
| 2015 | Learning descriptors for object recognition and 3D pose estimationabstractDetecting poorly textured objects and estimating their 3D pose reliably is still a very challenging problem. We introduce a simple but powerful approach to computing descriptors for object views that efficiently capture both the object identity and 3D pose. By contrast with previous manifold-based approaches, we can rely on the Euclidean distance to evaluate the similarity between descriptors, and therefore use scalable Nearest Neighbor search methods to efficiently handle a large number of objects under a large range of poses. To achieve this, we train a Convolutional Neural Network to compute these descriptors by enforcing simple similarity and dissimilarity constraints between the descriptors. We show that our constraints nicely untangle the images from different objects and different views into clusters that are not only well-separated but also structured as the corresponding sets of poses: The Euclidean distance between descriptors is large when the descriptors are from different objects, and directly related to the distance between the poses when the descriptors are from the same object. These important properties allow us to outperform state-of-the-art object views representations on challenging RGB and RGB-D data. Paul Wohlhart, Vincent Lepetit |
CVPR | 2 |
| 2015 | A Novel Representation of Parts for Accurate 3D Object Detection and Tracking in Monocular ImagesabstractWe present a method that estimates in real-time and under challenging conditions the 3D pose of a known object. Our method relies only on grayscale images since depth cameras fail on metallic objects, it can handle poorly textured objects, and cluttered, changing environments, the pose it predicts degrades gracefully in presence of large occlusions. As a result, by contrast with the state-of-the-art, our method is suitable for practical Augmented Reality applications even in industrial environments. To be robust to occlusions, we first learn to detect some parts of the target object. Our key idea is to then predict the 3D pose of each part in the form of the 2D projections of a few control points. The advantages of this representation is three-fold: We can predict the 3D pose of the object even when only one part is visible, when several parts are visible, we can combine them easily to compute a better pose of the object, the 3D pose we obtain is usually very accurate, even when only few parts are visible. Alberto Crivellaro, Mahdi Rad, Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
ICCV | 6 |
| 2015 | Training a Feedback Loop for Hand Pose EstimationabstractWe propose an entirely data-driven approach to estimating the 3D pose of a hand given a depth image. We show that we can correct the mistakes made by a Convolutional Neural Network trained to predict an estimate of the 3D pose by using a feedback loop. The components of this feedback loop are also Deep Networks, optimized using training data. They remove the need for fitting a 3D model to the input data, which requires both a carefully designed fitting function and algorithm. We show that our approach outperforms state-of-the-art methods, and is efficient as our implementation runs at over 400 fps on a single GPU. Markus Oberweger, Paul Wohlhart, Vincent Lepetit |
ICCV | 3 |
| 2015 | Projection onto the Manifold of Elongated Structures for Accurate ExtractionabstractDetection of elongated structures in 2D images and 3D image stacks is a critical prerequisite in many applications and Machine Learning-based approaches have recently been shown to deliver superior performance. However, these methods essentially classify individual locations and do not explicitly model the strong relationship that exists between neighboring ones. As a result, isolated erroneous responses, discontinuities, and topological errors are present in the resulting score maps. We solve this problem by projecting patches of the score map to their nearest neighbors in a set of ground truth training patches. Our algorithm induces global spatial consistency on the classifier score map and returns results that are provably geometrically consistent. We apply our algorithm to challenging datasets in four different domains and show that it compares favorably to state-of-the-art methods. Amos Sironi, Vincent Lepetit, Pascal Fua |
ICCV | 2 |
| 2015 | An Efficient Minimal Solution for Multi-camera MotionabstractSummary form only given. We propose an efficient method for estimating the motion of a multi-camera rig from a minimal set of feature correspondences. Existing methods for solving the multi-camera relative pose problem require extra correspondences, are slow to compute, and/or produce a multitude of solutions. Our solution uses a first-order approximation to relative pose in order to simplify the problem and produce an accurate estimate quickly. The solver is applicable to sequential multi-camera motion estimation and is fast enough for real-time implementation in a random sampling framework. Our experiments show that our approach is both stable and efficient on challenging test sequences. Jonathan Ventura, Clemens Arth, Vincent Lepetit |
ICCV | 3 |
| 2015 | You Should Use Regression to Detect Cells
Philipp Kainz, Martin Urschler, Samuel Schulter, Paul Wohlhart, Vincent Lepetit |
MICCAI (3) | 5 |
| 2015 | Introduction to the CVIU special issue on "Parts and Attributes: Mid-level representation for object recognition, scene classification and object detection"
Trevor Darrell, Vittorio Ferrari, Frédéric Jurie, Vincent Lepetit |
Comput. Vis. Image Underst. | 4 |
| 2015 | On rendering synthetic images for training an object detector
Artem Rozantsev, Vincent Lepetit, Pascal Fua |
Comput. Vis. Image Underst. | 2 |
| 2015 | Learning Separable FiltersabstractLearning filters to produce sparse image representations in terms of over-complete dictionaries has emerged as a powerful way to create image features for many different purposes. Unfortunately, these filters are usually both numerous and non-separable, making their use computationally expensive. In this paper, we show that such filters can be computed as linear combinations of a smaller number of separable ones, thus greatly reducing the computational complexity at no cost in terms of performance. This makes filter learning approaches practical even for large images or 3D volumes, and we show that we significantly outperform state-of-the-art methods on the curvilinear structure extraction task, in terms of both accuracy and speed. Moreover, our approach is general and can be used on generic convolutional filter banks to reduce the complexity of the feature extraction step. Amos Sironi, Bugra Tekin, Roberto Rigamonti, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Learning Image Descriptors with BoostingabstractWe propose a novel and general framework to learn compact but highly discriminative floating-point and binary local feature descriptors. By leveraging the boosting-trick we first show how to efficiently train a compact floating-point descriptor that is very robust to illumination and viewpoint changes. We then present the main contribution of this paper-a binary extension of the framework that demonstrates the real advantage of our approach and allows us to compress the descriptor even further. Each bit of the resulting binary descriptor, which we call BinBoost, is computed with a boosted binary hash function, and we show how to efficiently optimize the hash functions so that they are complementary, which is key to compactness and robustness. As we do not put any constraints on the weak learner configuration underlying each hash function, our general framework allows us to optimize the sampling patterns of recently proposed hand-crafted descriptors and significantly improve their performance. Moreover, our boosting scheme can easily adapt to new applications and generalize to other types of image data, such as faces, while providing state-of-the-art results at a fraction of the matching time and memory footprint. Tomasz Trzcinski, C. Mario Christoudias, Vincent Lepetit |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Instant Outdoor Localization and SLAM Initialization from 2.5D MapsabstractWe present a method for large-scale geo-localization and global tracking of mobile devices in urban outdoor environments. In contrast to existing methods, we instantaneously initialize and globally register a SLAM map by localizing the first keyframe with respect to widely available untextured 2.5D maps. Given a single image frame and a coarse sensor pose prior, our localization method estimates the absolute camera orientation from straight line segments and the translation by aligning the city map model with a semantic segmentation of the image. We use the resulting 6DOF pose, together with information inferred from the city map model, to reliably initialize and extend a 3D SLAM map in a global coordinate system, applying a model-supported SLAM mapping approach. We show the robustness and accuracy of our localization approach on a challenging dataset, and demonstrate unconstrained global SLAM mapping and tracking of arbitrary camera motion on several sequences. Clemens Arth, Christian Pirchheim, Jonathan Ventura, Dieter Schmalstieg, Vincent Lepetit |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2014 | Robust 3D Tracking with Descriptor FieldsabstractWe introduce a method that can register challenging images from specular and poorly textured 3D environments, on which previous approaches fail. We assume that a small set of reference images of the environment and a partial 3D model are available. Like previous approaches, we register the input images by aligning them with one of the reference images using the 3D information. However, these approaches typically rely on the pixel intensities for the alignment, which is prone to fail in presence of specularities or in absence of texture. Our main contribution is an efficient novel local descriptor that we use to describe each image location. We show that we can rely on this descriptor in place of the intensities to significantly improve the alignment robustness at a minor increase of the computational cost, and we analyze the reasons behind the success of our descriptor. Alberto Crivellaro, Vincent Lepetit |
CVPR | 2 |
| 2014 | Multiscale Centerline Detection by Learning a Scale-Space Distance TransformabstractWe propose a robust and accurate method to extract the centerlines and scale of tubular structures in 2D images and 3D volumes. Existing techniques rely either on filters designed to respond to ideal cylindrical structures, which lose accuracy when the linear structures become very irregular, or on classification, which is inaccurate because locations on centerlines and locations immediately next to them are extremely difficult to distinguish. We solve this problem by reformulating centerline detection in terms of a regression problem. We first train regressors to return the distances to the closest centerline in scale-space, and we apply them to the input images or volumes. The centerlines and the corresponding scale then correspond to the regressors local maxima, which can be easily identified. We show that our method outperforms state-of-the-art techniques for various 2D and 3D datasets. Amos Sironi, Vincent Lepetit, Pascal Fua |
CVPR | 2 |
| 2014 | Tracking texture-less, shiny objects with descriptor fieldsabstractOur demo demonstrates the method we published at CVPR this year for tracking specular and poorly textured objects, and lets the visitors experiment with it and with their own patterns. Our approach only requires a standard monocular camera (no need for a depth sensor), and can be easily integrated within existing systems to improve their robustness and accuracy. Alberto Crivellaro, Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
ISMAR | 5 |
| 2014 | On the relevance of sparsity for image classification
Roberto Rigamonti, Vincent Lepetit, Germán González, Engin Türetken, Fethallah Benmansour, Matthew A. Brown, Pascal Fua |
Comput. Vis. Image Underst. | 2 |
| 2014 | Real-time landing place assessment in man-made environments
Xiaolu Sun, C. Mario Christoudias, Vincent Lepetit, Pascal Fua |
Mach. Vis. Appl. | 3 |
| 2013 | Learning Separable FiltersabstractLearning filters to produce sparse image representations in terms of over complete dictionaries has emerged as a powerful way to create image features for many different purposes. Unfortunately, these filters are usually both numerous and non-separable, making their use computationally expensive. In this paper, we show that such filters can be computed as linear combinations of a smaller number of separable ones, thus greatly reducing the computational complexity at no cost in terms of performance. This makes filter learning approaches practical even for large images or 3D volumes, and we show that we significantly outperform state-of-the-art methods on the linear structure extraction task, in terms of both accuracy and speed. Moreover, our approach is general and can be used on generic filter banks to reduce the complexity of the convolutions. Roberto Rigamonti, Amos Sironi, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2013 | Boosting Binary Keypoint DescriptorsabstractBinary key point descriptors provide an efficient alternative to their floating-point competitors as they enable faster processing while requiring less memory. In this paper, we propose a novel framework to learn an extremely compact binary descriptor we call Bin Boost that is very robust to illumination and viewpoint changes. Each bit of our descriptor is computed with a boosted binary hash function, and we show how to efficiently optimize the different hash functions so that they complement each other, which is key to compactness and robustness. The hash functions rely on weak learners that are applied directly to the image patches, which frees us from any intermediate representation and lets us automatically learn the image gradient pooling configuration of the final descriptor. Our resulting descriptor significantly outperforms the state-of-the-art binary descriptors and performs similarly to the best floating-point descriptors at a fraction of the matching time and memory footprint. Tomasz Trzcinski, C. Mario Christoudias, Pascal Fua, Vincent Lepetit |
CVPR | 4 |
| 2013 | Supervised Feature Learning for Curvilinear Structure Segmentation
Carlos J. Becker, Roberto Rigamonti, Vincent Lepetit, Pascal Fua |
MICCAI (1) | 3 |
| 2012 | Model Based Training, Detection and Pose Estimation of Texture-Less 3D Objects in Heavily Cluttered Scenes
Stefan Hinterstoißer, Vincent Lepetit, Slobodan Ilic, Stefan Holzer, Gary R. Bradski, Kurt Konolige, Nassir Navab |
ACCV (1) | 2 |
| 2012 | Efficient Discriminative Projections for Compact Binary Descriptors
Tomasz Trzcinski, Vincent Lepetit |
ECCV (1) | 2 |
| 2012 | Accurate and Efficient Linear Structure Segmentation by Leveraging Ad Hoc Features with Learned Filters
Roberto Rigamonti, Vincent Lepetit |
MICCAI (1) | 2 |
| 2012 | Learning Image Descriptors with the Boosting-TrickabstractIn this paper we apply boosting to learn complex non-linear local visual feature representations, drawing inspiration from its successful application to visual object detection. The main goal of local feature descriptors is to distinctively represent a salient image region while remaining invariant to viewpoint and illumination changes. This representation can be improved using machine learning, however, past approaches have been mostly limited to learning linear feature mappings in either the original input or a kernelized input feature space. While kernelized methods have proven somewhat effective for learning non-linear local feature descriptors, they rely heavily on the choice of an appropriate kernel function whose selection is often difficult and non-intuitive. We propose to use the boosting-trick to obtain a non-linear mapping of the input to a high-dimensional feature space. The non-linear feature mapping obtained with the boosting-trick is highly intuitive. We employ gradient-based weak learners resulting in a learned descriptor that closely resembles the well-known SIFT. As demonstrated in our experiments, the resulting descriptor can be learned directly from intensity patches achieving state-of-the-art performance. Tomasz Trzcinski, C. Mario Christoudias, Vincent Lepetit, Pascal Fua |
NIPS | 3 |
| 2012 | Real-time interactive modeling and scalable multiple object tracking for AR
Kiyoung Kim, Vincent Lepetit, Woontack Woo |
Comput. Graph. | 2 |
| 2012 | BRIEF: Computing a Local Binary Descriptor Very FastabstractBinary descriptors are becoming increasingly popular as a means to compare feature points very fast while requiring comparatively small amounts of memory. The typical approach to creating them is to first compute floating-point ones, using an algorithm such as SIFT, and then to binarize them. In this paper, we show that we can directly compute a binary descriptor, which we call BRIEF, on the basis of simple intensity difference tests. As a result, BRIEF is very fast both to build and to match. We compare it against SURF and SIFT on standard benchmarks and show that it yields comparable recognition accuracy, while running in an almost vanishing fraction of the time required by either. Michael Calonder, Vincent Lepetit, Mustafa Özuysal, Tomasz Trzcinski, Christoph Strecha, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Gradient Response Maps for Real-Time Detection of Textureless ObjectsabstractWe present a method for real-time 3D object instance detection that does not require a time-consuming training stage, and can handle untextured objects. At its core, our approach is a novel image representation for template matching designed to be robust to small image transformations. This robustness is based on spread image gradient orientations and allows us to test only a small subset of all possible pixel locations when parsing the image, and to represent a 3D object with a limited set of templates. In addition, we demonstrate that if a dense depth sensor is available we can extend our approach for an even better performance also taking 3D surface normal orientations into account. We show how to take advantage of the architecture of modern computers to build an efficient but very discriminant representation of the input images that can be used to consider thousands of templates in real time. We demonstrate in many experiments on real data that our method is much faster and more robust with respect to background clutter than current state-of-the-art methods. Stefan Hinterstoißer, Cedric Cagniart, Slobodan Ilic, Peter F. Sturm, Nassir Navab, Pascal Fua, Vincent Lepetit |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2012 | Thick boundaries in binary space and their influence on nearest-neighbor search
Tomasz Trzcinski, Vincent Lepetit, Pascal Fua |
Pattern Recognit. Lett. | 2 |
| 2012 | 3-D Head Tracking via Invariant Keypoint LearningabstractKeypoint matching is a standard tool to solve the correspondence problem in vision applications. However, in 3-D face tracking, this approach is often deficient because the human face complexities, together with its rich viewpoint, nonrigid expression, and lighting variations in typical applications, can cause many variations impossible to handle by existing keypoint detectors and descriptors. In this paper, we propose a new approach to tailor keypoint matching to track the 3-D pose of the user head in a video stream. The core idea is to learn keypoints that are explicitly invariant to these challenging transformations. First, we select keypoints that are stable under randomly drawn small viewpoints, nonrigid deformations, and illumination changes. Then, we treat keypoint descriptor learning at different large angles as an incremental scheme to learn discriminative descriptors. At matching time, to reduce the ratio of outlier correspondences, we use second-order color information to prune keypoints unlikely to lie on the face. Moreover, we integrate optical flow correspondences in an adaptive way to remove motion jitter efficiently. Extensive experiments show that the proposed approach can lead to fast, robust, and accurate 3-D head tracking results even under very challenging scenarios. Franck Davoine, Vincent Lepetit, Christophe Chaillou, Chunhong Pan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Handling Motion-Blur in 3D Tracking and Rendering for Augmented RealityabstractThe contribution of this paper is two-fold. First, we show how to extend the ESM algorithm to handle motion blur in 3D object tracking. ESM is a powerful algorithm for template matching-based tracking, but it can fail under motion blur. We introduce an image formation model that explicitly consider the possibility of blur, and shows its results in a generalization of the original ESM algorithm. This allows to converge faster, more accurately and more robustly even under large amount of blur. Our second contribution is an efficient method for rendering the virtual objects under the estimated motion blur. It renders two images of the object under 3D perspective, and warps them to create many intermediate images. By fusing these images we obtain a final image for the virtual objects blurred consistently with the captured image. Because warping is much faster than 3D rendering, we can create realistically blurred images at a very low computational cost. Vincent Lepetit, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2011 | Are sparse representations really relevant for image classification?abstractRecent years have seen an increasing interest in sparse representations for image classification and object recognition, probably motivated by evidence from the analysis of the primate visual cortex. It is still unclear, however, whether or not sparsity helps classification. In this paper we evaluate its impact on the recognition rate using a shallow modular architecture, adopting both standard filter banks and filter banks learned in an unsupervised way. In our experiments on the CIFAR-10 and on the Caltech-101 datasets, enforcing sparsity constraints actually does not improve recognition performance. This has an important practical impact in image descriptor design, as enforcing these constraints can have a heavy computational cost. Roberto Rigamonti, Matthew A. Brown, Vincent Lepetit |
CVPR | 3 |
| 2011 | Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenesabstractWe present a method for detecting 3D objects using multi-modalities. While it is generic, we demonstrate it on the combination of an image and a dense depth map which give complementary object information. It works in real-time, under heavy clutter, does not require a time consuming training stage, and can handle untextured objects. It is based on an efficient representation of templates that capture the different modalities, and we show in many experiments on commodity hardware that our approach significantly outperforms state-of-the-art methods on single modalities. Stefan Hinterstoißer, Stefan Holzer, Cedric Cagniart, Slobodan Ilic, Kurt Konolige, Nassir Navab, Vincent Lepetit |
ICCV | 7 |
| 2011 | Texture-less object tracking with online training using an RGB-D cameraabstractWe propose a texture-less object detection and 3D tracking method which automatically extracts on the fly the information it needs from color images and the corresponding depth maps. While texture-less 3D tracking is not new, it requires a prior CAD model, and real-time methods for detection still have to be developed for robust tracking. To detect the target, we propose to rely on a fast template-based method, which provides an initial estimate of its 3D pose, and we refine this estimate using the depth and image contours information. We automatically extract a 3D model for the target from the depth information. To this end, we developed methods to enhance the depth map and to stabilize the 3D pose estimation. We demonstrate our method on challenging sequences exhibiting partial occlusions and fast motions. Vincent Lepetit, Woontack Woo |
ISMAR | 2 |
| 2011 | General chairs
Martin Wiedmer, Vincent Lepetit |
ISMAR | 2 |
| 2011 | Learning Real-Time Perspective Patch Rectification
Stefan Hinterstoißer, Vincent Lepetit, Selim Benhimane, Pascal Fua, Nassir Navab |
Int. J. Comput. Vis. | 2 |
| 2011 | Video-Based In Situ Tagging on Mobile PhonesabstractWe propose a novel way to augment a real-world scene with minimal user intervention on a mobile phone; the user only has to point the phone camera to the desired location of the augmentation. Our method is valid for horizontal or vertical surfaces only, but this is not a restriction in practice in manmade environments, and it avoids going through any reconstruction of the 3-D scene, which is still a delicate process on a resource-limited system like a mobile phone. Our approach is inspired by recent work on perspective patch recognition, but we adapt it for better performances on mobile phones. We reduce user interaction with real scenes by exploiting the phone accelerometers to relax the need for fronto-parallel views. As a result, we can learn a planar target in situ from arbitrary viewpoints and augment it with virtual objects in real-time on a mobile phone. Wonwoo Lee, Vincent Lepetit, Woontack Woo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Extended Keyframe Detection with Stable Tracking for Multiple 3D Object TrackingabstractWe present a method that is able to track several 3D objects simultaneously, robustly, and accurately in real time. While many applications need to consider more than one object in practice, the existing methods for single object tracking do not scale well with the number of objects, and a proper way to deal with several objects is required. Our method combines object detection and tracking: frame-to-frame tracking is less computationally demanding but is prone to fail, while detection is more robust but slower. We show how to combine them to take the advantages of the two approaches and demonstrate our method on several real sequences. Vincent Lepetit, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2010 | Pareto-optimal dictionaries for signaturesabstractWe present an effective method to optimize over the parameters of an image patch descriptor to obtain one that is computationally more efficient while maintaining a high recognition rate. We formulate the optimization problem in a multi-objective manner, which balances two conflicting goals while removing the need for traditional weighting coefficients. To this end we introduce the Pareto efficiency criterion, which helps finding solutions that increase one objective without decreasing the other. Despite the vast size of the search space, we show how a state-of-the-art Genetic Algorithm can be tailored to find good solutions. Not only does the resulting descriptor perform better than state-of-the-art ones, but our approach is of broader significance as optimization problems with balanced goals are often encountered in Computer Vision. Michael Calonder, Vincent Lepetit, Pascal Fua |
CVPR | 2 |
| 2010 | Dominant orientation templates for real-time detection of texture-less objectsabstractWe present a method for real-time 3D object detection that does not require a time consuming training stage, and can handle untextured objects. At its core, is a novel template representation that is designed to be robust to small image transformations. This robustness based on dominant gradient orientations lets us test only a small subset of all possible pixel locations when parsing the image, and to represent a 3D object with a limited set of templates. We show that together with a binary representation that makes evaluation very fast and a branch-and-bound approach to efficiently scan the image, it can detect untextured objects in complex situations and provide their 3D pose in real-time. Stefan Hinterstoißer, Vincent Lepetit, Slobodan Ilic, Pascal Fua, Nassir Navab |
CVPR | 2 |
| 2010 | BRIEF: Binary Robust Independent Elementary Features
Michael Calonder, Vincent Lepetit, Christoph Strecha, Pascal Fua |
ECCV (4) | 2 |
| 2010 | Combining Geometric and Appearance Priors for Robust Homography Estimation
Eduard Serradell, Mustafa Özuysal, Vincent Lepetit, Pascal Fua, Francesc Moreno-Noguer |
ECCV (3) | 3 |
| 2010 | Keyframe-based modeling and tracking of multiple 3D objectsabstractWe propose a real-time solution for modeling and tracking multiple 3D objects in unknown environments. Our contribution is two-fold: First, we show how to scale with the number of objects. This is done by combining recent techniques for image retrieval and online Structure from Motion, which can be run in parallel. As a result, tracking 40 objects in 3D can be done within 6 to 25 milliseconds per frame, even under difficult conditions for tracking. Second, we propose a method to let the user add new objects very quickly. The user simply has to select in an image a 2D region lying on the object. A 3D primitive is then fitted to the features within this region, and adjusted to create the object 3D model. In practice, this procedure takes less than a minute. Kiyoung Kim, Vincent Lepetit, Woontack Woo |
ISMAR | 2 |
| 2010 | Point-and-shoot for ubiquitous tagging on mobile phonesabstractWe propose a novel way to augment a real scene with minimalist user intervention on a mobile phone: The user only has to point the phone camera to the desired location of the augmentation. Our method is valid for vertical or horizontal surfaces only, but this is not a restriction in practice in man-made environments, and avoids to go through any reconstruction of the 3D scene, which is still a delicate process. Our approach is inspired by recent work on perspective patch recognition and we show how to modify it for better performances on mobile phones and how to exploit the phone accelerometers to relax the need for fronto-parallel views. In addition, our implementation allows to share the augmentations and the required data over peer-to-peer communication to build a shared AR space on mobile phones. Wonwoo Lee, Vincent Lepetit, Woontack Woo |
ISMAR | 3 |
| 2010 | Augmented reality for board gamesabstractWe introduce a new type of Augmented Reality games: By using a simple webcam and Computer Vision techniques, we turn a standard real game board pawns into an AR game. We use these objects as a tangible interface, and augment them with visual effects. The game logic can be performed automatically by the computer. This results in a better immersion compared to the original board game alone and provides a different experience than a video game. We demonstrate our approach on Monopoly™, but it is very generic and could easily be adapted to any other board game. Eray Molla, Vincent Lepetit |
ISMAR | 2 |
| 2010 | A Fully Automated Approach to Segmentation of Irregularly Shaped Cellular Structures in EM Images
Aurélien Lucchi, Kevin Smith 0001, Radhakrishna Achanta, Vincent Lepetit, Pascal Fua |
MICCAI (2) | 4 |
| 2010 | From Canonical Poses to 3D Motion Capture Using a Single CameraabstractWe combine detection and tracking techniques to achieve robust 3D motion recovery of people seen from arbitrary viewpoints by a single and potentially moving camera. We rely on detecting key postures, which can be done reliably, using a motion model to infer 3D poses between consecutive detections, and finally refining them over the whole sequence using a generative model. We demonstrate our approach in the cases of golf motions filmed using a static camera and walking motions acquired using a potentially moving one. We will show that our approach, although monocular, is both metrically accurate because it integrates information over many frames and robust because it can recover from a few misdetections. Andrea Fossati, Miodrag Dimitrijevic, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | Fast Keypoint Recognition Using Random FernsabstractWhile feature point recognition is a key component of modern approaches to object detection, existing approaches require computationally expensive patch preprocessing to handle perspective distortion. In this paper, we show that formulating the problem in a naive Bayesian classification framework makes such preprocessing unnecessary and produces an algorithm that is simple, efficient, and robust. Furthermore, it scales well as the number of classes grows. To recognize the patches surrounding keypoints, our classifier uses hundreds of simple binary features and models class posterior probabilities. We make the problem computationally tractable by assuming independence between arbitrary sets of features. Even though this is not strictly true, we demonstrate that our classifier nevertheless performs remarkably well on image data sets containing very significant perspective changes. Mustafa Özuysal, Michael Calonder, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | DAISY: An Efficient Dense Descriptor Applied to Wide-Baseline StereoabstractIn this paper, we introduce a local image descriptor, DAISY, which is very efficient to compute densely. We also present an EM-based algorithm to compute dense depth and occlusion maps from wide-baseline image pairs using this descriptor. This yields much better results in wide-baseline situations than the pixel and correlation-based algorithms that are commonly used in narrow-baseline stereo. Also, using a descriptor makes our algorithm robust against many photometric and geometric transformations. Our descriptor is inspired from earlier ones such as SIFT and GLOH but can be computed much faster for our purposes. Unlike SURF, which can also be computed efficiently at every pixel, it does not introduce artifacts that degrade the matching performance when used densely. It is important to note that our approach is the first algorithm that attempts to estimate dense depth maps from wide-baseline image pairs, and we show that it is a good one at that with many experiments for depth estimation accuracy, occlusion detection, and comparing it against other descriptors on laser-scanned ground truth scenes. We also tested our approach on a variety of indoor and outdoor scenes with different photometric and geometric transformations and our experiments support our claim to being robust against these. Engin Tola, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Scalable real-time planar targets tracking for digilog books
Kiyoung Kim, Vincent Lepetit, Woontack Woo |
Vis. Comput. | 2 |
| 2009 | Appearance-based keypoint clusteringabstractWe present an algorithm for clustering sets of detected interest points into groups that correspond to visually distinct structure. Through the use of a suitable colour and texture representation, our clustering method is able to identify keypoints that belong to separate objects or background regions. These clusters are then used to constrain the matching of keypoints over pairs of images, resulting in greatly improved matching under difficult conditions. We present a thorough evaluation of each component of the algorithm, and show its usefulness on difficult matching problems. Francisco J. Estrada, Pascal Fua, Vincent Lepetit, Sabine Süsstrunk |
CVPR | 3 |
| 2009 | Real-time learning of accurate patch rectificationabstractRecent work showed that learning-based patch rectification methods are both faster and more reliable than affine region methods. Unfortunately, their performance improvements are founded in a computationally expensive offline learning stage, which is not possible for applications such as SLAM. In this paper we propose an approach whose training stage is fast enough to be performed at run-time without the loss of accuracy or robustness. To this end, we developed a very fast method to compute the mean appearances of the feature points over sets of small variations that span the range of possible camera viewpoints. Then, by simply matching incoming feature points against these mean appearances, we get a coarse estimate of the viewpoint that is refined afterwards. Because there is no need to compute descriptors for the input image, the method is very fast at run-time. We demonstrate our approach on tracking-by-detection for SLAM, real-time object detection and pose estimation applications. Stefan Hinterstoißer, Oliver Kutter, Nassir Navab, Pascal Fua, Vincent Lepetit |
CVPR | 5 |
| 2009 | Capturing 3D stretchable surfaces from single images in closed formabstractWe present a closed form solution to the problem of recovering the 3D shape of a nonrigid potentially stretchable surface from 3D-to-2D correspondences. In other words, we can reconstruct a surface from a single image without a priori knowledge of its deformations in that image. State of the art solutions to nonrigid 3D shape recovery rely on the fact that distances between neighboring surface points must be preserved and are therefore limited to inelastic surfaces. Here, we show that replacing the inextensibility constraints by shading ones removes this limitation while still allowing 3D reconstruction in closed-form. We demonstrate our method and compare it to an earlier one using both synthetic and real data. Francesc Moreno-Noguer, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2009 | Pose estimation for category specific multiview object localizationabstractWe propose an approach to overcome the two main challenges of 3D multiview object detection and localization: The variation of object features due to changes in the viewpoint and the variation in the size and aspect ratio of the object. Our approach proceeds in three steps. Given an initial bounding box of fixed size, we first refine its aspect ratio and size. We can then predict the viewing angle, under the hypothesis that the bounding box actually contains an object instance. Finally, a classifier tuned to this particular viewpoint checks the existence of an instance. As a result, we can find the object instances and estimate their poses, without having to search over all window sizes and potential orientations. We train and evaluate our method on a new object database specifically tailored for this task, containing real-world objects imaged over a wide range of smoothly varying viewpoints and significant lighting changes. We show that the successive estimations of the bounding box and the viewpoint lead to better localization results. Mustafa Özuysal, Vincent Lepetit, Pascal Fua |
CVPR | 2 |
| 2009 | Compact signatures for high-speed interest point description and matchingabstractProminent feature point descriptors such as SIFT and SURF allow reliable real-time matching but at a computational cost that limits the number of points that can be handled on PCs, and even more on less powerful mobile devices. A recently proposed technique that relies on statistical classification to compute signatures has the potential to be much faster but at the cost of using very large amounts of memory, which makes it impractical for implementation on low-memory devices. In this paper, we show that we can exploit the sparseness of these signatures to compact them, speed up the computation, and drastically reduce memory usage. We base our approach on Compressive Sensing theory. We also highlight its effectiveness by incorporating it into two very different SLAM packages and demonstrating substantial performance increases. Michael Calonder, Vincent Lepetit, Pascal Fua, Kurt Konolige, James Bowman, Patrick Mihelich |
ICCV | 2 |
| 2009 | Fast Ray features for learning irregular shapesabstractWe introduce a new class of image features, the Ray feature set, that consider image characteristics at distant contour points, capturing information which is difficult to represent with standard feature sets. This property allows Ray features to efficiently and robustly recognize deformable or irregular shapes, such as cells in microscopic imagery. Experiments show Ray features clearly outperform other powerful features including Haar-like features and Histograms of Oriented Gradients when applied to detecting irregularly shaped neuron nuclei and mitochondria. Ray features can also provide important complementary information to Haar features for other tasks such as face detection, reducing the number of weak learners and computational cost. Ray features can be efficiently precomputed to reduce cost, just as precomputing integral images reduces the overall cost of Haar features. While Rays are slightly more expensive to precompute, their computational cost is less than that of Haar features for scanning an AdaBoost-based detector window across an image at run-time. Kevin Smith 0001, Alan Carleton, Vincent Lepetit |
ICCV | 3 |
| 2009 | ESM-Blur: Handling & rendering blur in 3D tracking and augmentationabstractThe contribution of this paper is two-fold. First, we show how to extend the ESM algorithm to handle motion blur in 3D object tracking. ESM is a powerful algorithm for template matching-based tracking, but it can fail under motion blur. We introduce an image formation model that explicitly considers the possibility of blur, and show it results in a generalization of the original ESM algorithm. This allows to converge faster, more accurately and more robustly even under large amount of blur. Our second contribution is an efficient method for rendering the virtual objects under the estimated motion blur. It renders two images of the object under 3D perspective, and warps them to create many intermediate images. By fusing these images we obtain a final image for the virtual objects blurred consistently with the captured image. Because warping is much faster that 3D rendering, we can create realistically blurred images at a very low computational cost. Vincent Lepetit, Woontack Woo |
ISMAR | 2 |
| 2009 | EPnP: An Accurate O(n) Solution to the PnP Problem
Vincent Lepetit, Francesc Moreno-Noguer, Pascal Fua |
Int. J. Comput. Vis. | 1 |
| 2008 | Simultaneous Recognition and Homography Extraction of Local Patches with a Simple Linear ClassifierabstractWe show that the simultaneous estimation of keypoint identities and poses is more reliable than the two separate steps undertaken by previous approaches. A simple linear classifier coupled with linear predictors trained during a learning phase appears to be sufficient for this task. The retrieved poses are subpixel accurate due to the linear predictors. We demonstrate the advantages of our approach on real-time 3D object detection and tracking applications. Thanks to the high accuracy, one single keypoint is often enough to precisely estimate the object pose. As a result, we can deal in real-time with objects that are significantly less textured than the ones required by state-of-the-art methods. 1 Stefan Hinterstoißer, Selim Benhimane, Vincent Lepetit, Pascal Fua, Nassir Navab |
BMVC | 3 |
| 2008 | Online learning of patch perspective rectification for efficient object detectionabstractFor a large class of applications, there is time to train the system. In this paper, we propose a learning-based approach to patch perspective rectification, and show that it is both faster and more reliable than state-of-the-art ad hoc affine region detection methods. Our method performs in three steps. First, a classifier provides for every keypoint not only its identity, but also a first estimate of its transformation. This estimate allows carrying out, in the second step, an accurate perspective rectification using linear predictors. We show that both the classifier and the linear predictors can be trained online, which makes the approach convenient. The last step is a fast verification - made possible by the accurate perspective rectification - of the patch identity and its sub-pixel precision position estimation. We test our approach on real-time 3D object detection and tracking applications. We show that we can use the estimated perspective rectifications to determine the object pose and as a result, we need much fewer correspondences to obtain a precise pose estimation. Stefan Hinterstoißer, Selim Benhimane, Nassir Navab, Pascal Fua, Vincent Lepetit |
CVPR | 5 |
| 2008 | 3D pose refinement from reflectionsabstractWe demonstrate how to exploit reflections for accurate registration of shiny objects: The lighting environment can be retrieved from the reflections under a distant illumination assumption. Since it remains unchanged when the camera or the object of interest moves, this provides powerful additional constraints that can be incorporated into standard pose estimation algorithms. The key idea and main contribution of the paper is therefore to show that the registration should also be performed in the lighting environment space, instead of in the image space only. This lets us recover very accurate pose estimates because the specularities are very sensitive to pose changes. An interesting side result is an accurate estimate of the lighting environment. Furthermore, since the mapping from lighting environment to specularities has no analytical expression for objects represented as 3D meshes, and is not 1-to-1, registering lighting environments is far from trivial. However we propose a general and effective solution. Our approach is demonstrated on both synthetic and real images. Pascal Lagger, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2008 | General constraints for batch Multiple-Target Tracking applied to large-scale videomicroscopyabstractWhile there is a large class of Multiple-Target Tracking (MTT) problems for which batch processing is possible and desirable, batch MTT remains relatively unexplored in comparison to sequential approaches. In this paper, we give a principled probabilistic formalization of batch MTT in which we introduce two new, very general constraints that considerably help us in reaching the correct solution. First, we exploit the correlation between the appearance of a target and its motion. Second, entrances and departures of targets are encouraged to occur at the boundaries of the scene. We show how to implement these constraints in a formal and efficient manner. Our approach is applied to challenging 3-D biomedical imaging data where the number of targets is unknown and may vary, and numerous challenging tracking events occur. We demonstrate the ability of our model to simultaneously track the nuclei of over one hundred migrating neuron precursor cells in image stack series collected from a 2-photon microscope. Kevin Smith 0001, Alan Carleton, Vincent Lepetit |
CVPR | 3 |
| 2008 | A fast local descriptor for dense matchingabstractWe introduce a novel local image descriptor designed for dense wide-baseline matching purposes. We feed our descriptors to a graph-cuts based dense depth map estimation algorithm and this yields better wide-baseline performance than the commonly used correlation windows for which the size is hard to tune. As a result, unlike competing techniques that require many high-resolution images to produce good reconstructions, our descriptor can compute them from pairs of low-quality images such as the ones captured by video streams. Our descriptor is inspired from earlier ones such as SIFT and GLOH but can be computed much faster for our purposes. Unlike SURF which can also be computed efficiently at every pixel, it does not introduce artifacts that degrade the matching performance. Our approach was tested with ground truth laser scanned depth maps as well as on a wide variety of image pairs of different resolutions and we show that good reconstructions are achieved even with only two low quality images. Engin Tola, Vincent Lepetit, Pascal Fua |
CVPR | 2 |
| 2008 | Keypoint Signatures for Fast Learning and Recognition
Michael Calonder, Vincent Lepetit, Pascal Fua |
ECCV (1) | 2 |
| 2008 | Pose Priors for Simultaneously Solving Alignment and Correspondence
Francesc Moreno-Noguer, Vincent Lepetit, Pascal Fua |
ECCV (2) | 2 |
| 2008 | Closed-Form Solution to Non-rigid 3D Surface Registration
Mathieu Salzmann, Francesc Moreno-Noguer, Vincent Lepetit, Pascal Fua |
ECCV (4) | 3 |
| 2008 | Multiple 3D Object tracking for augmented realityabstractWe present a method that is able to track several 3D objects simultaneously, robustly, and accurately in real-time. While many applications need to consider more than one object in practice, the existing methods for single object tracking do not scale well with the number of objects, and a proper way to deal with several objects is required. Our method combines object detection and tracking: Frame-to-frame tracking is less computationally demanding but is prone to fail, while detection is more robust but slower. We show how to combine them to take the advantages of the two approaches, and demonstrate our method on several real sequences. Vincent Lepetit, Woontack Woo |
ISMAR | 2 |
| 2008 | The haunted bookabstractThis paper describes an artwork that relies on recent computer vision and augmented reality techniques to animate the illustrations of a poetry book. Because we donpsilat need markers, we can achieve seamless integration of real and virtual elements to create the desired atmosphere. The visualization is done on a computer screen to avoid cumbersome head-mounted displays. The camera is hidden into a desk lamp for easing even more the spectator immersion. Camille Scherrer, Julien Pilet, Pascal Fua, Vincent Lepetit |
ISMAR | 4 |
| 2008 | Fast Non-Rigid Surface Detection, Registration and Realistic Augmentation
Julien Pilet, Vincent Lepetit, Pascal Fua |
Int. J. Comput. Vis. | 2 |
| 2007 | Linear and Quadratic Subsets for Template-Based TrackingabstractWe propose a method that dramatically improves the performance of template-based matching in terms of size of convergence region and computation time. This is done by selecting a subset of the template that verifies the assumption (made during optimization) of linearity or quadraticity with respect to the motion parameters. We call these subsets linear or quadratic subsets. While subset selection approaches have already been proposed, they generally do not attempt to provide linear or quadratic subsets and rely on heuristics such as textured-ness. Because a naive search for the optimal subset would result in a combinatorial explosion for large templates, we propose a simple algorithm that does not aim for the optimal subset but provides a very good linear or quadratic subset at low cost, even for large templates. Simulation results and experiments with real sequences show the superiority of the proposed method compared to existing subset selection approaches. Selim Benhimane, Alexander Ladikos, Vincent Lepetit, Nassir Navab |
CVPR | 3 |
| 2007 | Bridging the Gap between Detection and Tracking for 3D Monocular Video-Based Motion CaptureabstractWe combine detection and tracking techniques to achieve robust 3-D motion recovery of people seen from arbitrary viewpoints by a single and potentially moving camera. We rely on detecting key postures, which can be done reliably, using a motion model to infer 3-D poses between consecutive detections, and finally refining them over the whole sequence using a generative model. We demonstrate our approach in the case of people walking against cluttered backgrounds and filmed using a moving camera, which precludes the use of simple background subtraction techniques. In this case, the easy-to-detect posture is the one that occurs at the end of each step when people have their legs furthest apart. Andrea Fossati, Miodrag Dimitrijevic, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2007 | Fast Keypoint Recognition in Ten Lines of CodeabstractWhile feature point recognition is a key component of modern approaches to object detection, existing approaches require computationally expensive patch preprocessing to handle perspective distortion. In this paper, we show that formulating the problem in a Naive Bayesian classification framework makes such preprocessing unnecessary and produces an algorithm that is simple, efficient, and robust. Furthermore, it scales well to handle large number of classes. To recognize the patches surrounding keypoints, our classifier uses hundreds of simple binary features and models class posterior probabilities. We make the problem computationally tractable by assuming independence between arbitrary sets of features. Even though this is not strictly true, we demonstrate that our classifier nevertheless performs remarkably well on image datasets containing very significant perspective changes. Mustafa Özuysal, Pascal Fua, Vincent Lepetit |
CVPR | 3 |
| 2007 | Deformable Surface Tracking AmbiguitiesabstractWe study from a theoretical standpoint the ambiguities that occur when tracking a generic deformable surface under monocular perspective projection given 3D to 2D correspondences. We show that, additionally to the known scale ambiguity, a set of potential ambiguities can be clearly identified. From this, we deduce a minimal set of constraints required to disambiguate the problem and incorporate them into a working algorithm that runs on real noisy data. Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
CVPR | 2 |
| 2007 | Accurate Non-Iterative O(n) Solution to the PnP ProblemabstractWe propose a non-iterative solution to the PnP problem-the estimation of the pose of a calibrated camera from n 3D-to-2D point correspondences—whose computational complexity grows linearly with 𝑛5) or even 𝑂(𝑛8), without being more accurate. Our method is applicable for all 𝑛≥4 and handles properly both planar and non-planar configurations. Our central idea is to express the 𝑛 3D points as a weighted sum of four virtual control points. The problem then reduces to estimating the coordinates of these control points in the camera referential, which can be done in 𝑂(𝑛) time by expressing these coordinates as weighted sum of the eigenvectors of a 12 × 12 matrix and solving a small constant number of quadratic equations to pick the right weights. The advantages of our method are demonstrated by thorough testing on both synthetic and real-data. Francesc Moreno-Noguer, Vincent Lepetit, Pascal Fua |
ICCV | 2 |
| 2007 | Retexturing in the Presence of Complex Illumination and OcclusionsabstractWe present a nonrigid registration technique that achieves spatial, photometric, and visibility accuracy. It lets us photo-realistically augment 3D deformable surfaces under complex illumination conditions and in spite of severe occlusions. There are many approaches that address some of these issues but very few that simultaneously handle all of them as we do. We use triangulated meshes to model the geometry and introduce explicit visibility maps as well as separate illumination parameters for each mesh vertex. We cast our registration problem in an expectation maximization framework that allows robust and fully automated operation. It provides explicit illumination and occlusion models that can be used for rendering purposes. Julien Pilet, Vincent Lepetit, Pascal Fua |
ISMAR | 2 |
| 2006 | Feature Harvesting for Tracking-by-Detection
Mustafa Özuysal, Vincent Lepetit, François Fleuret, Pascal Fua |
ECCV (3) | 2 |
| 2006 | An all-in-one solution to geometric and photometric calibrationabstractWe propose a fully automated approach to calibrating multiple cameras whose fields of view may not all overlap. Our technique only requires waving an arbitrary textured planar pattern in front of the cameras, which is the only manual intervention that is required. The pattern is then automatically detected in the frames where it is visible and used to simultaneously recover geometric and photometric camera calibration parameters. In other words, even a novice user can use our system to extract all the information required to add virtual 3D objects into the scene and light them convincingly. This makes it ideal for Augmented Reality applications and we distribute the code under a GPL license. Julien Pilet, Andreas Geiger 0001, Pascal Lagger, Vincent Lepetit, Pascal Fua |
ISMAR | 4 |
| 2006 | Human body pose detection using Bayesian spatio-temporal templates
Miodrag Dimitrijevic, Vincent Lepetit, Pascal Fua |
Comput. Vis. Image Underst. | 2 |
| 2006 | Keypoint Recognition Using Randomized TreesabstractIn many 3D object-detection and pose-estimation problems, runtime performance is of critical importance. However, there usually is time to train the system, which we will show to be very useful. Assuming that several registered images of the target object are available, we developed a keypoint-based approach that is effective in this context by formulating wide-baseline matching of keypoints extracted from the input images to those found in the model images as a classification problem. This shifts much of the computational burden to a training phase, without sacrificing recognition performance. As a result, the resulting algorithm is robust, accurate, and fast-enough for frame-rate performance. This reduction in runtime computational complexity is our first contribution. Our second contribution is to show that, in this context, a simple and fast keypoint detector suffices to support detection and tracking even under large perspective and scale variations. While earlier methods require a detector that can be expected to produce very repeatable results, in general, which usually is very time-consuming, we simply find the most repeatable object keypoints for the specific target object during the training phase. We have incorporated these ideas into a real-time system that detects planar, nonplanar, and deformable objects. It then estimates the pose of the rigid ones and the deformations of the others. Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Randomized Trees for Real-Time Keypoint RecognitionabstractIn earlier work, we proposed treating wide baseline matching of feature points as a classification problem, in which each class corresponds to the set of all possible views of such a point. We used a K-mean plus Nearest Neighbor classifier to validate our approach, mostly because it was simple to implement. It has proved effective but still too slow for real-time use. In this paper, we advocate instead the use of randomized trees as the classification technique. It is both fast enough for real-time performance and more robust. It also gives us a principled way not only to match keypoints but to select during a training phase those that are the most recognizable ones. This results in a real-time system able to detect and position in 3D planar, non-planar, and even deformable objects. It is robust to illuminations changes, scale changes and occlusions. Vincent Lepetit, Pascal Lagger, Pascal Fua |
CVPR (2) | 1 |
| 2005 | Real-Time Non-Rigid Surface DetectionabstractWe present a real-time method for detecting deformable surfaces, with no need whatsoever for a priori pose knowledge. Our method starts from a set of wide baseline point matches between an undeformed image of the object and the image in which it is to be detected. The matches are used not only to detect but also to compute a precise mapping from one to the other. The algorithm is robust to large deformations, lighting changes, motion blur, and occlusions. It runs at 10 frames per second on a 2.8 GHz PC and we are not aware of any other published technique that produces similar results. Combining deformable meshes with a well designed robust estimator is key to dealing with the large number of parameters involved in modeling deformable surfaces and rejecting erroneous matches for error rates of up to 95%, which is considerably more than what is required in practice. Julien Pilet, Vincent Lepetit, Pascal Fua |
CVPR (1) | 2 |
| 2005 | Augmenting Deformable Objects in Real-TimeabstractWe present a real-time system that can draw virtual patterns or images on deforming real objects by estimating both the deformations and the shading parameters. We show that this is what is required to render the virtual elements so that they blend convincingly with the surrounding real textures. The whole process of uncompressing the video stream, measuring the deformations, estimating the lighting parameters, and realistically augmenting the input image takes about 100 ms on a 2.8 GHz PC. It is fully automated and does not require any manual initialization or engineering of the scene. It is also robust to large deformations, lighting changes, motion blur, specularities, and occlusions. It can therefore be demonstrated live on a simple laptop. Julien Pilet, Vincent Lepetit, Pascal Fua |
ISMAR | 2 |
| 2004 | Markov-based Silhouette Extraction for Three--Dimensional Body Tracking in Presence of Cluttered BackgroundabstractWe propose a novel method to detect human body contours in presence of clutter and complex texture. Contours are extracted using a novel Markovbased approach which learns a texture along a given scanline in order to detect texture crossings. In contrast to conventional silhouette detection algorithms based on gradient, our texture boundary detection method allows extraction of silhouettes of textured and non-textured objects under difficult conditions such as having a cluttered/moving background. We demonstrate on demanding examples of monocular body tracking that our proposed method yields better results than gradient-based techniques. 1. Ali Shahrokni, Vincent Lepetit, Tom Drummond, Pascal Fua |
BMVC | 2 |
| 2004 | Point Matching as a Classification Problem for Fast and Robust Object Pose Estimation
Vincent Lepetit, Julien Pilet, Pascal Fua |
CVPR (2) | 1 |
| 2004 | ombining Edge and Texture Information for Real-Time Accurate 3D Camera TrackingabstractWe present an effective way to combine the information provided by edges and by feature points for the purpose of robust real-time 3-D tracking. This lets our tracker handle both textured and untextured objects. As it can exploit more of the image information, it is more stable and less prone to drift that purely edge or feature-based ones. We start with a feature-point based tracker we developed in earlier work and integrate the ability to take edge-information into account. Achieving optimal performance in the presence of cluttered or textured backgrounds, however, is far from trivial because of the many spurious edges that bedevil typical edge-detectors. We overcome this difficulty by proposing a method for handling multiple hypotheses for potential edge-locations that is similar in speed to approaches that consider only single hypotheses and therefore much faster than conventional multiple-hypothesis ones. This results in a real-time 3-D tracking algorithm that exploits both texture and edge information without being sensitive to misleading background information and that does not drift over time. Luca Vacchetti, Vincent Lepetit, Pascal Fua |
ISMAR | 2 |
| 2004 | Stable Real-Time 3D Tracking Using Online and Offline InformationabstractWe propose an efficient real-time solution for tracking rigid objects in 3D using a single camera that can handle large camera displacements, drastic aspect changes, and partial occlusions. While commercial products are already available for offline camera registration, robust online tracking remains an open issue because many real-time algorithms described in the literature still lack robustness and are prone to drift and jitter. To address these problems, we have formulated the tracking problem in terms of local bundle adjustment and have developed a method for establishing image correspondences that can equally well handle short and wide-baseline matching. We then can merge the information from preceding frames with that provided by a very limited number of keyframes created during a training stage, which results in a real-time tracker that does not jitter or drift and can deal with significant aspect changes. Luca Vacchetti, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Visual Golf Club Tracking for Enhanced Swing AnalysisabstractCVLAB Nicolas Gehrig, Vincent Lepetit, Pascal Fua |
BMVC | 2 |
| 2003 | Robust Data Association For Online ApplicationsabstractWe present a method for performing data association that handles complex motion models while increasing the robustness of tracking and being suitable for real-time applications. Instead of using motion model in standard recursive fashion, we robustly fit it over multiple frames simultaneously. This allows us to naturally handle arbitrarily complex motion models, to automate the initialization and to deal with occlusion and false alarms. This is effective even if the motion model is not entirely accurate and if there are frequent false-negatives and false-positives. Our algorithm is easy to implement and we show its performances on two real examples of complex motion tracking. Vincent Lepetit, Ali Shahrokni, Pascal Fua |
CVPR (1) | 1 |
| 2003 | Fusing Online and Offline Information for Stable 3D Tracking in Real-TimeabstractWe propose an efficient online real-time solution for single-camera 3D tracking of rigid objects that can handle large camera displacements, drastic aspect changes, and partial occlusions. While the offline camera registration problem can be considered as essentially solved, robust online tracking remains an open issue because many real-time algorithms described in the literature still lack robustness and are prone to drift and jitter. To solve these problems, we have developed a robust approach to 3D feature matching that can handle wide-baseline matching: our method merges the information from preceding frames in traditional recursive tracking fashion with that provided by a very limited number of keyframes created during an offline stage. This combination results in a system that does not suffer from the above difficulties and can deal with drastic aspect changes. We use augmented reality applications to demonstrate its behavior because they are particularly demanding in terms of tracking performance. Luca Vacchetti, Vincent Lepetit, Pascal Fua |
CVPR (2) | 2 |
| 2003 | Fully Automated and Stable Registration for Augmented Reality ApplicationsabstractWe present a fully automated approach to camera registration for augmented reality systems. It relies on purely passive vision techniques to solve the initialization and real-time tracking problems, given a rough CAD model of parts of the real scene. It does not require a controlled environment, for example placing markers. It handles arbitrarily complex models, occlusions, large camera displacements and drastic aspect changes. This is made possible by two major contributions: the first one is a fast recognition method that detects the known part of the scene, registers the camera with respect to it, and initializes a real-time tracker, which is the second contribution. Our tracker eliminates drift and jitter by merging the information from preceding frames in a traditional recursive tracking fashion with that of a very limited number of key-frames created off-line. In the rare instances where it fails, for example because of large occlusion, it detects the failure and reinvokes the initialization procedure. We present experimental results on several different kinds of objects and scenes. Vincent Lepetit, Luca Vacchetti, Daniel Thalmann, Pascal Fua |
ISMAR | 1 |
| 2003 | Real-Time Augmented FaceabstractThis real-time augmented reality demonstration relies on our tracking algorithm described in V. Lepetit et al (2003). This algorithm considers natural feature points, and then does not require engineering of the environment. It merges the information from preceding frames in traditional recursive tracking fashion with that provided by a very limited number of reference frames. This combination results in a system that does not suffer from jitter and drift, and can deal with drastic changes. The tracker recovers the full 3D pose of the tracked object, allowing insertion of 3D virtual objects for augmented reality applications. Vincent Lepetit, Luca Vacchetti, Daniel Thalmann, Pascal Fua |
ISMAR | 1 |
| 2002 | Polyhedral Object Detection and Pose Estimation for Augmented Reality ApplicationsabstractIn augmented reality applications, tracking and registration of both cameras and objects is required because, to combine real and rendered scenes, we must project synthetic models at the right location in real images. Although much work has been done to track objects of interest, initialization of theses trackers often remains manual. Our work aims at automating this step by integrating object recognition and tracking into an AR system. Our emphasis is on the initialization phase of the tracking. We address all the three major aspects of the problem of model-to-image registration: feature detection, correspondence and pose estimation. We have developed a novel approach based on facet detection that greatly reduces the number of possible feature correspondences making it possible to directly compute the transformation which best maps 3-D object to the image plane. We will argue that this approach offers a one-fold speed-up over existing methods. Results of our AR system which integrates initialization and tracking are shown. Our method takes about 5 seconds on our example images. Ali Shahrokni, Luca Vacchetti, Vincent Lepetit, Pascal Fua |
CA | 3 |
| 2000 | A Semi-Automatic Method for Resolving Occlusion in Augmented RealityabstractRealistic merging of virtual and real objects requires that the augmented patterns be correctly occluded by foreground objects. In this paper we propose a semi-automatic method for resolving occlusion in augmented reality which makes use of key-views. Once the user has outlined the occluding objects in the key-views, our system detects automatically these occluding objects in the intermediate views. A region of interest that contains the occluding objects is first computed from the outlined silhouettes. One of the main contribution of this paper is that this region takes into account the uncertainty on the computed interframe motion. Then a deformable region-based approach is used to recover the actual occluding boundary within the region of interest from this prediction. Vincent Lepetit, Marie-Odile Berger |
CVPR | 1 |