VLDB 2026 Research / reviewers in the wild / expert
Pascal Fua
dblp:f/PFua
· DBLP profile ↗
374ranked-venue papers
25as first author
69since 2021 · last 2025
0000-0002-6702-9970ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 304 · 24 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 251 · 13 first-author · 43 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 10 since 2021Human-computer interaction and ubiquitous computing · 9Systems, architecture and hardware · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected TexturesabstractReal-time free-view human rendering from sparse-view RGB inputs is a challenging task due to the sensor scarcity and the tight time budget. To ensure efficiency, recent methods leverage 2D CNNs operating in texture space to learn rendering primitives. However, they either jointly learn geometry and appearance, or completely ignore sparse image information for geometry estimation, significantly harming visual quality and robustness to unseen body poses. To address these issues, we present Double Unprojected Textures, which at the core disentangles coarse geometric deformation estimation from appearance synthesis, enabling robust and photorealistic 4K rendering in real-time. Specifically, we first introduce a novel image-conditioned template deformation network, which estimates the coarse deformation of the human template from a first unprojected texture. This updated geometry is then used to apply a second and more accurate texture unprojection. The resulting texture map has fewer artifacts and better alignment with input views, which benefits our learning of finer-level geometry and appearance represented by Gaussian splats. We validate the effectiveness and efficiency of the proposed method in quantitative and qualitative experiments, which significantly surpasses other state-of-the-art methods. Guoxing Sun 0001, Rishabh Dabral, Heming Zhu, Pascal Fua, Christian Theobalt, Marc Habermann |
CVPR | 4 |
| 2025 | Counting Stacked Objects
Corentin Dumery, Noa Etté, Aoxiang Fan, Hieu Le 0001, Pascal Fua |
ICCV | 7 |
| 2025 | A View-Consistent Sampling Method for Regularized Training of Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) has emerged as a compelling framework for scene representation and 3D recovery. To improve its performance on real-world data, depth regularizations have proven to be the most effective ones. However, depth estimation models not only require expensive 3D supervision in training, but also suffer from generalization issues. As a result, the depth estimations can be erroneous in practice, especially for outdoor unbounded scenes. In this paper, we propose to employ view-consistent distributions instead of fixed depth value estimations to regularize NeRF training. Specifically, the distribution is computed by utilizing both low-level color features and high-level distilled features from foundation models at the projected 2D pixel-locations from per-ray sampled 3D points. By sampling from the view-consistency distributions, an implicit regularization is imposed on the training of NeRF. We also utilize a depth-pushing loss that works in conjunction with the sampling technique to jointly provide effective regularizations for eliminating the failure modes. Extensive experiments conducted on various scenes from public datasets demonstrate that our proposed method can generate significantly better novel view synthesis results than state-of-the-art NeRF variants as well as different depth regularization methods. Aoxiang Fan, Corentin Dumery, Nicolas Talabot, Pascal Fua |
ICCV | 4 |
| 2025 | LeFusion: Controllable Pathology Synthesis via Lesion-Focused Diffusion ModelsabstractPatient data from real-world clinical practice often suffers from data scarcity and long-tail imbalances, leading to biased outcomes or algorithmic unfairness. This study addresses these challenges by generating lesion-containing image-segmentation pairs from lesion-free images. Previous efforts in medical imaging synthesis have struggled with separating lesion information from background, resulting in low-quality backgrounds and limited control over the synthetic output. Inspired by diffusion-based image inpainting, we propose LeFusion, a lesion-focused diffusion model. By redesigning the diffusion learning objectives to focus on lesion areas, we simplify the learning process and improve control over the output while preserving high-fidelity backgrounds by integrating forward-diffused background contexts into the reverse diffusion process. Additionally, we tackle two major challenges in lesion texture synthesis: 1) multi-peak and 2) multi-class lesions. We introduce two effective strategies: histogram-based texture control and multi-channel decomposition, enabling the controlled generation of high-quality lesions in difficult scenarios. Furthermore, we incorporate lesion mask diffusion, allowing control over lesion size, location, and boundary, thus increasing lesion diversity. Validated on 3D cardiac lesion MRI and lung nodule CT datasets, LeFusion-generated data significantly improves the performance of state-of-the-art segmentation models, including nnUNet and SwinUNETR. Yuhe Liu, Jiancheng Yang, Shouhong Wan, Pascal Fua |
ICLR | 7 |
| 2025 | IT3: Idempotent Test-Time TrainingabstractDeep learning models often struggle when deployed in real-world settings due to distribution shifts between training and test data. While existing approaches like domain adaptation and test-time training (TTT) offer partial solutions, they typically require additional data or domain-specific auxiliary tasks. We present Idempotent Test-Time Training (IT3), a novel approach that enables on-the-fly adaptation to distribution shifts using only the current test instance, without any auxiliary task design. Our key insight is that enforcing idempotence---where repeated applications of a function yield the same result---can effectively replace domain-specific auxiliary tasks used in previous TTT methods. We theoretically connect idempotence to prediction confidence and demonstrate that minimizing the distance between successive applications of our model during inference leads to improved out-of-distribution performance. Extensive experiments across diverse domains (including image classification, aerodynamics prediction, and aerial segmentation) and architectures (MLPs, CNNs, GNNs) show that IT3 consistently outperforms existing approaches while being simpler and more widely applicable. Our results suggest that idempotence provides a universal principle for test-time adaptation that generalizes across domains and architectures. Nikita Durasov, Assaf Shocher, Doruk Öner, Gal Chechik, Alexei A. Efros, Pascal Fua |
ICML | 6 |
| 2025 | Pairwise-Constrained Implicit Functions for 3D Human Heart Modeling
Hieu Le 0001, Nicolas Talabot, Jiancheng Yang, Pascal Fua |
MICCAI (16) | 5 |
| 2025 | DiffAtlas: GenAI-Fying Atlas Segmentation via Image-Mask Diffusion
Yuhe Liu, Jiancheng Yang, Weidong Guo, Pascal Fua |
MICCAI (16) | 6 |
| 2025 | High Resolution UDF Meshing via Iterative NetworksabstractUnsigned Distance Fields (UDFs) are a natural implicit representation for open surfaces but, unlike Signed Distance Fields (SDFs), are challenging to triangulate into explicit meshes. This is especially true at high resolutions where neural UDFs exhibit higher noise levels, which makes it hard to capture fine details.
Most current techniques perform within single voxels without reference to their neighborhood, resulting in missing surface and holes where the UDF is ambiguous or noisy. We show that this can be remedied by performing several passes and by reasoning on previously extracted surface elements to incorporate neighborhood information. Our key contribution is an iterative neural network that does this and progressively improves surface recovery within each voxel by spatially propagating information from increasingly distant neighbors. Unlike single-pass methods, our approach integrates newly detected surfaces, distance values, and gradients across multiple iterations, effectively correcting errors and stabilizing extraction in challenging regions. Experiments on diverse 3D models demonstrate that our method produces significantly more accurate and complete meshes than existing approaches, particularly for complex geometries, enabling UDF surface extraction at higher resolutions where traditional methods fail. Federico Stella, Nicolas Talabot, Hieu Le 0001, Pascal Fua |
NeurIPS | 4 |
| 2025 | Efficient anatomical labeling of pulmonary tree structures via deep point-graph representation-based implicit fieldsabstractPulmonary diseases rank prominently among the principal causes of death worldwide. Curing them will require, among other things, a better understanding of the complex 3D tree-shaped structures within the pulmonary system, such as airways, arteries, and veins. Traditional approaches using high-resolution image stacks and standard CNNs on dense voxel grids face challenges in computational efficiency, limited resolution, local context, and inadequate preservation of shape topology. Our method addresses these issues by shifting from dense voxel to sparse point representation, offering better memory efficiency and global context utilization. However, the inherent sparsity in point representation can lead to a loss of crucial connectivity in tree-shaped structures. To mitigate this, we introduce graph learning on skeletonized structures, incorporating differentiable feature fusion for improved topology and long-distance context capture. Furthermore, we employ an implicit function for efficient conversion of sparse representations into dense reconstructions end-to-end. The proposed method not only delivers state-of-the-art performance in labeling accuracy, both overall and at key locations, but also enables efficient inference and the generation of closed surface shapes. Addressing data scarcity in this field, we have also curated a comprehensive dataset to validate our approach. Data and code are available at https://github.com/M3DV/pulmonary-tree-labeling. Kangxian Xie, Jiancheng Yang, Donglai Wei 0001, Ziqiao Weng, Pascal Fua |
Medical Image Anal. | 5 |
| 2025 | Vision-based power line cables and pylons detection for low flying aircraftabstractAbstract Power lines are dangerous for low-flying aircraft, especially in low-visibility conditions. Thus, a vision-based system able to analyze the aircraft’s surroundings and to provide the pilots with a “second pair of eyes” can contribute to enhancing their safety. To this end, we develop a deep learning approach to jointly detect power line cables and pylons from images captured at distances of several hundred meters by aircraft-mounted cameras. In doing so, we combine a modern convolutional architecture with transfer learning and a loss function adapted to curvilinear structure delineation. We use a single network for both detection tasks and demonstrate its performance on two benchmarking datasets. We have also integrated it within an onboard system and run it inflight. We show with our experiments that it outperforms the prior distant cable detection method by Stambler et al. (in: International Conference on Robotics and Automation, 2019) on both datasets, while also successfully detecting pylons, given their annotations are available for the data. Jakub Gwizdala, Doruk Öner, Soumava Kumar Roy, Mian Akbar Shah, Ad Eberhard, Ivan Egorov, Philipp Krüsi, Grigory Yakushev, Pascal Fua |
Mach. Vis. Appl. | 9 |
| 2025 | Temporally-Consistent Surface Reconstruction Using Metrically-Consistent AtlasesabstractWe propose a method for unsupervised reconstruction of a temporally-consistent sequence of surfaces from a sequence of time-evolving point clouds. It yields dense and semantically meaningful correspondences between frames. We represent the reconstructed surfaces as atlases computed by a neural network, which enables us to establish correspondences between frames. The key to making these correspondences semantically meaningful is to guarantee that the metric tensors computed at corresponding points are as similar as possible. We have devised an optimization strategy that makes our method robust to noise and global motions, without a priori correspondences or pre-alignment steps. As a result, our approach outperforms state-of-the-art ones on several challenging datasets. Jan Bednarík, Noam Aigerman, Vladimir G. Kim, Siddhartha Chaudhuri, Shaifali Parashar, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Unsupervised 3D Keypoint Discovery with Multi-View GeometryabstractAnalyzing and training 3D body posture models depend heavily on the availability of joint labels that are commonly acquired through laborious manual annotation of body joints or via marker-based joint localization using carefully curated markers and capturing systems. However, such annotations are not always available, especially for people performing unusual activities. In this paper, we propose an algorithm that learns to discover 3D keypoints on human bodies from multiple-view images without any supervision or labels other than the constraints multiple-view geometry provides. To ensure that the discovered 3D keypoints are meaningful, they are re-projected to each view to estimate the person’s mask that the model itself has initially estimated without supervision. Our approach discovers more interpretable and accurate 3D keypoints compared to other state-of-the-art unsupervised approaches on Human3.6M and MPI-INF-3DHP benchmark datasets. Sina Honari, Chen Zhao 0025, Mathieu Salzmann, Pascal Fua |
3DV | 4 |
| 2024 | Occlusion Resilient 3D Human Pose EstimationabstractOcclusions remain one of the key challenges in 3D body pose estimation from single-camera video sequences. Temporal consistency has been extensively used to mitigate their impact but the existing algorithms in the literature do not explicitly model them.Here, we apply this by representing the deforming body as a spatio-temporal graph. We then introduce a refinement network that performs graph convolutions over this graph to output 3D poses. To ensure robustness to occlusions, we train this network with a set of binary masks that we use to disable some of the edges as in drop-out techniques.In effect, we simulate the fact that some joints can be hidden for periods of time and train the network to be immune to that. We demonstrate the effectiveness of this approach compared to state-of-the-art techniques that infer poses from single-camera sequences. Soumava Kumar Roy, Ilia Badanin, Sina Honari, Pascal Fua |
3DV | 4 |
| 2024 | AttEntropy: On the Generalization Ability of Supervised Semantic Segmentation Transformers to New Objects in New Domains
Krzysztof Lis, Matthias Rottmann, Annika Mütze, Sina Honari, Pascal Fua, Mathieu Salzmann |
BMVC | 5 |
| 2024 | CLOAF: CoLlisiOn-Aware Human FlowabstractEven the best current algorithms for estimating body 3D shape and pose yield results that include body self- intersections. In this paper, we present CLOAF, which exploits the diffeomorphic nature of Ordinary Differential Equations to eliminate such self-intersections while still im- posing body shape constraints. We show that, unlike earlier approaches to addressing this issue, ours completely elim- inates the self-intersections without compromising the ac- curacy of the reconstructions. Being differentiable, CLOAF can be used to fine-tune pose and shape estimation base- lines to improve their overall performance and eliminate self-intersections in their predictions. Furthermore, we demonstrate how our CLOAF strategy can be applied to practically any motion field induced by the user. CLOAF also makes it possible to edit motion to interact with the environment without worrying about potential collision or loss of body-shape prior. Andrey Davydov, Martin Engilberge, Mathieu Salzmann, Pascal Fua |
CVPR | 4 |
| 2024 | Garment Recovery with Shape and Deformation PriorsabstractWhile modeling people wearing tight-fitting clothing has made great strides in recent years, loose-fitting clothing remains a challenge. We propose a method that delivers realistic garment models from real-world images, regardless of garment shape or deformation. To this end, we introduce a fitting approach that utilizes shape and deformation priors learned from synthetic data to accurately capture garment shapes and deformations, including large ones. Not only does our approach recover the garment geometry accurately, it also yields models that can be directly used by downstream applications such as animation and simulation. Corentin Dumery, Benoît Guillard, Pascal Fua |
CVPR | 4 |
| 2024 | Neural Surface Detection for Unsigned Distance Fields
Federico Stella, Nicolas Talabot, Hieu Le 0001, Pascal Fua |
ECCV (58) | 4 |
| 2024 | MetaCap: Meta-learning Priors from Multi-view Imagery for Sparse-View Human Performance Capture and Rendering
Guoxing Sun 0001, Rishabh Dabral, Pascal Fua, Christian Theobalt, Marc Habermann |
ECCV (46) | 3 |
| 2024 | Enabling Uncertainty Estimation in Iterative Neural NetworksabstractTurning pass-through network architectures into iterative ones, which use their own output as input, is a well-known approach for boosting performance. In this paper, we argue that such architectures offer an additional benefit: The convergence rate of their successive outputs is highly correlated with the accuracy of the value to which they converge. Thus, we can use the convergence rate as a useful proxy for uncertainty. This results in an approach to uncertainty estimation that provides state-of-the-art estimates at a much lower computational cost than techniques like Ensembles, and without requiring any modifications to the original iterative model. We demonstrate its practical value by embedding it in two application domains: road detection in aerial images and the estimation of aerodynamic properties of 2D and 3D shapes. Nikita Durasov, Doruk Öner, Jonathan Donier, Hieu Le 0001, Pascal Fua |
ICML | 5 |
| 2024 | Generating Anatomically Accurate Heart Structures via Neural Implicit Fields
Jiancheng Yang, Ekaterina Sedykh, Jason Ken Adhinarta, Hieu Le 0001, Pascal Fua |
MICCAI (1) | 5 |
| 2024 | Reconstruction of Manipulated Garment with Guided Deformation PriorabstractModeling the shape of garments has received much attention, but most existing approaches assume the garments to be worn by someone, which constrains the range of shapes they can assume. In this work, we address shape recovery when garments are being manipulated instead of worn, which gives rise to an even larger range of possible shapes. To this end, we leverage the implicit sewing patterns (ISP) model for garment modeling and extend it by adding a diffusion-based deformation prior to represent these shapes. To recover 3D garment shapes from incomplete 3D point clouds acquired when the garment is folded, we map the points to UV space, in which our priors are learned, to produce partial UV maps, and then fit the priors to recover complete UV maps and 2D to 3D mappings. Experimental results demonstrate the superior reconstruction accuracy of our method compared to previous ones, especially when dealing with large non-rigid deformations arising from the manipulations. Corentin Dumery, Zhantao Deng, Pascal Fua |
NeurIPS | 4 |
| 2024 | DeepMesh: Differentiable Iso-Surface ExtractionabstractGeometric Deep Learning has recently made striking progress with the advent of continuous deep implicit fields. They allow for detailed modeling of watertight surfaces of arbitrary topology while not relying on a 3D euclidean grid, resulting in a learnable parameterization that is unlimited in resolution. Unfortunately, these methods are often unsuitable for applications that require an explicit mesh-based surface representation because converting an implicit field to such a representation relies on the Marching Cubes algorithm, which cannot be differentiated with respect to the underlying implicit field. In this work, we remove this limitation and introduce a differentiable way to produce explicit surface mesh representations from Deep Implicit Fields. Our key insight is that by reasoning on how implicit field perturbations impact local surface geometry, one can ultimately differentiate the 3D location of surface samples with respect to the underlying deep implicit field. We exploit this to define DeepMesh - an end-to-end differentiable mesh representation that can vary its topology. We validate our theoretical insight through several applications: Single view 3D Reconstruction via Differentiable Rendering, Physically-Driven Shape Optimization, Full Scene 3D Reconstruction from Scans and End-to-End Training. In all cases our end-to-end differentiable parameterization gives us an edge over state-of-the-art algorithms. Benoît Guillard, Edoardo Remelli, Artem Lukoianov, Pierre Yvernay, Stephan R. Richter, Timur M. Bagautdinov, Pierre Baqué, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Detecting Road Obstacles by Erasing ThemabstractVehicles can encounter a myriad of obstacles on the road, and it is impossible to record them all beforehand to train a detector. Instead, we select image patches and inpaint them with the surrounding road texture, which tends to remove obstacles from those patches. We then use a network trained to recognize discrepancies between the original patch and the inpainted one, which signals an erased obstacle. Krzysztof Lis, Sina Honari, Pascal Fua, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | A Closed-Form, Pairwise Solution to Local Non-Rigid Structure-From-MotionabstractA recent trend in Non-Rigid Structure-from-Motion (NRSfM) is to express local, differential constraints between pairs of images, from which the surface normal at any point can be obtained by solving a system of polynomial equations. While this approach is more successful than its counterparts relying on global constraints, the resulting methods face two main problems: First, most of the equation systems they formulate are of high degree and must be solved using computationally expensive polynomial solvers. Some methods use polynomial reduction strategies to simplify the system, but this adds some phantom solutions. In any event, an additional mechanism is employed to pick the best solution, which adds to the computation without any guarantees on the reliability of the solution. Second, these methods formulate constraints between a pair of images. Even if there is enough motion between them, they may suffer from local degeneracies that make the resulting estimates unreliable without any warning mechanism. %Unfortunately, these systems are of high degree with up to five real solutions. Hence, a computationally expensive strategy is required to select a unique solution. Furthermore, they suffer from degeneracies that make the resulting estimates unreliable, without any mechanism to identify this situation. In this paper, we solve these problems for isometric/conformal NRSfM. We show that, under widely applicable assumptions, we can derive a new system of equations in terms of the surface normals, whose two solutions can be obtained in closed-form and can easily be disambiguated locally. Our formalism also allows us to assess how reliable the estimated local normals are and to discard them if they are not. Our experiments show that our reconstructions, obtained from two or more views, are significantly more accurate than those of state-of-the-art methods, while also being faster. %In this paper, we show that, under widely applicable assumptions, we can derive a new system of equations in terms of the surface normals, whose two solutions can be obtained in closed-form and can easily be disambiguated locally. Our formalism also allows us to assess how reliable the estimated local normals are and to discard them if they are not. Our experiments show that our reconstructions, obtained from two or more views, are significantly more accurate than those of state-of-the-art methods, while also being faster. Shaifali Parashar, Yuxuan Long, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | DrapeNet: Garment Generation and Self-Supervised DrapingabstractRecent approaches to drape garments quickly over arbitrary human bodies leverage self-supervision to eliminate the need for large training sets. However, they are designed to train one network per clothing item, which severely limits their generalization abilities. In our work, we rely on self-supervision to train a single network to drape multiple garments. This is achieved by predicting a 3D deformation field conditioned on the latent codes of a generative network, which models garments as unsigned distance fields. Our pipeline can generate and drape previously unseen garments of any topology, whose shape can be edited by manipulating their latent codes. Being fully differentiable, our formulation makes it possible to recover accurate 3D models of garments from partial observations - images or 3D scans - via gradient descent. Our code is publicly available at https://github.com/liren2515/DrapeNet Luca De Luigi, Benoît Guillard, Mathieu Salzmann, Pascal Fua |
CVPR | 5 |
| 2023 | LightDepth: Single-View Depth Self-Supervision from Illumination DeclineabstractSingle-view depth estimation can be remarkably effective if there is enough ground-truth depth data for supervised training. However, there are scenarios, especially in medicine in the case of endoscopies, where such data cannot be obtained. In such cases, multi-view self-supervision and synthetic-to-real transfer serve as alternative approaches, however, with a considerable performance reduction in comparison to supervised case. Instead, we propose a single-view self-supervised method that achieves a performance similar to the supervised case. In some medical devices, such as endoscopes, the camera and light sources are co-located at a small distance from the target surfaces. Thus, we can exploit that, for any given albedo and surface orientation, pixel brightness is inversely proportional to the square of the distance to the surface, providing a strong single-view self-supervisory signal. In our experiments, our self-supervised models deliver accuracies comparable to those of fully supervised ones, while being applicable without depth ground-truth data. Javier Rodriguez Puigvert, Victor M. Batlle, J. M. M. Montiel, Ruben Martinez-Cantin, Pascal Fua, Juan D. Tardós, Javier Civera 0001 |
ICCV | 5 |
| 2023 | GECCO: Geometrically-Conditioned Point Diffusion ModelsabstractDiffusion models generating images conditionally on text, such as Dall-E 2 [51] and Stable Diffusion [53], have recently made a splash far beyond the computer vision community. Here, we tackle the related problem of generating point clouds, both unconditionally, and conditionally with images. For the latter, we introduce a novel geometrically-motivated conditioning scheme based on projecting sparse image features into the point cloud and attaching them to each individual point, at every step in the denoising process. This approach improves geometric consistency and yields greater fidelity than current methods relying on unstructured, global latent codes. Additionally, we show how to apply recent continuous-time diffusion schemes [59], [21]. Our method performs on par or above the state of art on conditional and unconditional experiments on synthetic data, while being faster, lighter, and delivering tractable likelihoods. We show it can also scale to diverse indoors scenes. Michal J. Tyszkiewicz, Pascal Fua, Eduard Trulls |
ICCV | 2 |
| 2023 | Tracking Adaptation to Improve SuperPoint for 3D Reconstruction in Endoscopy
Oscar León Barbed, J. M. M. Montiel, Pascal Fua, Ana Cristina Murillo |
MICCAI (1) | 3 |
| 2023 | LightNeuS: Neural Surface Reconstruction in Endoscopy Using Illumination Decline
Victor M. Batlle, J. M. M. Montiel, Pascal Fua, Juan D. Tardós |
MICCAI (10) | 3 |
| 2023 | ISP: Multi-Layered Garment Draping with Implicit Sewing PatternsabstractMany approaches to draping individual garments on human body models are realistic, fast, and yield outputs that are differentiable with respect to the body shape on which they are draped. However, they are either unable to handle multi-layered clothing, which is prevalent in everyday dress, or restricted to bodies in T-pose. In this paper, we introduce a parametric garment representation model that addresses these limitations. As in models used by clothing designers, each garment consists of individual 2D panels. Their 2D shape is defined by a Signed Distance Function and 3D shape by a 2D to 3D mapping. The 2D parameterization enables easy detection of potential collisions and the 3D parameterization handles complex shapes effectively. We show that this combination is faster and yields higher quality reconstructions than purely implicit surface representations, and makes the recovery of layered garments from images possible thanks to its differentiability. Furthermore, it supports rapid editing of garment shapes and texture by modifying individual 2D panels. Benoît Guillard, Pascal Fua |
NeurIPS | 3 |
| 2023 | Multi-view Tracking Using Weakly Supervised Human Motion PredictionabstractMulti-view approaches to people-tracking have the potential to better handle occlusions than single-view ones in crowded scenes. They often rely on the tracking-by-detection paradigm, which involves detecting people first and then connecting the detections. In this paper, we argue that an even more effective approach is to predict people motion over time and infer people’s presence in individual frames from these. This enables to enforce consistency both over time and across views of a single temporal frame. We validate our approach on the PETS2009 and WILDTRACK datasets and demonstrate that it outperforms state-of-the-art methods. Martin Engilberge, Weizhe Liu, Pascal Fua |
WACV | 3 |
| 2023 | Two-level Data Augmentation for Calibrated Multi-view DetectionabstractData augmentation has proven its usefulness to improve model generalization and performance. While it is commonly applied in computer vision application when it comes to multi-view systems, it is rarely used. Indeed geometric data augmentation can break the alignment among views. This is problematic since multi-view data tend to be scarce and it is expensive to annotate.In this work we propose to solve this issue by introducing a new multi-view data augmentation pipeline that preserves alignment among views. Additionally to traditional augmentation of the input image we also propose a second level of augmentation applied directly at the scene level. When combined with our simple multi-view detection model, our two-level augmentation pipeline outperforms all existing baselines by a significant margin on the two main multi-view multi-person detection datasets WILD-TRACK and MultiviewX. Martin Engilberge, Haixin Shi, Zhiye Wang, Pascal Fua |
WACV | 4 |
| 2023 | State of the Art in Dense Monocular Non-Rigid 3D ReconstructionabstractAbstract 3D reconstruction of deformable (ornon‐rigid) scenes from a set of monocular 2D image observations is a long‐standing and actively researched area of computer vision and graphics. It is an ill‐posed inverse problem, since—without additional prior assumptions—it permits infinitely many solutions leading to accurate projection to the input 2D images. Non‐rigid reconstruction is a foundational building block for downstream applications like robotics, AR/VR, or visual content creation. The key advantage of using monocular cameras is their omnipresence and availability to the end users as well as their ease of use compared to more sophisticated camera set‐ups such as stereo or multi‐view systems. This survey focuses on state‐of‐the‐art methods for dense non‐rigid 3D reconstruction of various deformable objects and composite scenes from monocular videos or sets of monocular views. It reviews the fundamentals of 3D reconstruction and deformation modeling from 2D image observations. We then start from general methods—that handle arbitrary scenes and make only a few prior assumptions—and proceed towards techniques making stronger assumptions about the observed objects and types of deformations (e.g. human faces, bodies, hands, and animals). A significant part of this STAR is also devoted to classification and a high‐level comparison of the methods, as well as an overview of the datasets for training and evaluation of the discussed techniques. We conclude by discussing open challenges in the field and the social aspects associated with the usage of the reviewed methods. Edith Tretschk, Navami Kairanda, Mallikarjun B. R. 0001, Rishabh Dabral, Adam Kortylewski, Bernhard Egger 0001, Marc Habermann, Pascal Fua, Christian Theobalt, Vladislav Golyanik |
Comput. Graph. Forum | 8 |
| 2023 | Overcoming the Domain Gap in Neural Action RepresentationsabstractAbstract Relating behavior to brain activity in animals is a fundamental goal in neuroscience, with practical applications in building robust brain-machine interfaces. However, the domain gap between individuals is a major issue that prevents the training of general models that work on unlabeled subjects. Since 3D pose data can now be reliably extracted from multi-view video sequences without manual intervention, we propose to use it to guide the encoding of neural action representations together with a set of neural and behavioral augmentations exploiting the properties of microscopy imaging. To test our method, we collect a large dataset that features flies and their neural activity. To reduce the domain gap, during training, we mix features of neural and behavioral data across flies that seem to be performing similar actions. To show our method can generalize further neural modalities and other downstream tasks, we test our method on a human neural Electrocorticography dataset, and another RGB video data of human activities from different viewpoints. We believe our work will enable more robust neural decoding algorithms to be used in future brain-machine interfaces. Semih Günel, Florian Aymanns, Sina Honari, Pavan Ramdya, Pascal Fua |
Int. J. Comput. Vis. | 5 |
| 2023 | Temporal Representation Learning on Monocular Videos for 3D Human Pose EstimationabstractIn this article we propose an unsupervised feature extraction method to capture temporal information on monocular videos, where we detect and encode subject of interest in each frame and leverage contrastive self-supervised (CSS) learning to extract rich latent vectors. Instead of simply treating the latent features of nearby frames as positive pairs and those of temporally-distant ones as negative pairs as in other CSS approaches, we explicitly disentangle each latent vector into a time-variant component and a time-invariant one. We then show that applying contrastive loss only to the time-variant features and encouraging a gradual transition on them between nearby and away frames while also reconstructing the input, extract rich temporal features, well-suited for human pose estimation. Our approach reduces error by about 50% compared to the standard CSS strategies, outperforms other unsupervised single-view methods and matches the performance of multi-view techniques. When 2D pose is available, our approach can extract even richer latent features and improve the 3D pose estimation accuracy, outperforming other state-of-the-art weakly supervised methods. Sina Honari, Victor Constantin, Helge Rhodin, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Persistent Homology With Improved Locality Information for More Effective DelineationabstractPersistent Homology (PH) has been successfully used to train networks to detect curvilinear structures and to improve the topological quality of their results. However, existing methods are very global and ignore the location of topological features. In this paper, we remedy this by introducing a new filtration function that fuses two earlier approaches: thresholding-based filtration, previously used to train deep networks to segment medical images, and filtration with height functions, typically used to compare 2D and 3D shapes. We experimentally demonstrate that deep networks trained using our PH-based loss function yield reconstructions of road networks and neuronal processes that reflect ground-truth connectivity better than networks trained with existing loss functions based on PH. Doruk Öner, Adélie Garin, Mateusz Kozinski, Kathryn Hess, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Long Term Motion Prediction Using KeyposesabstractLong term human motion prediction is essential in safety-critical applications such as human-robot interaction and autonomous driving. In this paper we show that to achieve long term forecasting, predicting human pose at every time instant is unnecessary. Instead, it is more effective to predict a few keyposes and approximate intermediate ones by interpolating the keyposes. We demonstrate that our approach enables us to predict realistic motions for up to 5 seconds in the future, which is far longer than the typical 1 second encountered in the literature. Furthermore, because we model future keyposes probabilistically, we can generate multiple plausible future motions by sampling at inference time. Over this extended time period, our predictions are more realistic, more diverse and better preserve the motion dynamics than those state-of-the-art methods yield. Sena Kiciroglu, Wei Wang 0108, Mathieu Salzmann, Pascal Fua |
3DV | 4 |
| 2022 | On Triangulation as a Form of Self-Supervision for 3D Human Pose EstimationabstractSupervised approaches to 3D pose estimation from single images are remarkably effective when labeled data is abundant. However, as the acquisition of ground-truth 3D labels is labor intensive and time consuming, recent attention has shifted towards semi- and weakly-supervised learning. Generating an effective form of supervision with little annotations still poses major challenge in crowded scenes. In this paper we propose to impose multi-view geometrical constraints by means of a weighted differentiable tri-angulation and use it as a form of self-supervision when no labels are available. We therefore train a 2D pose estimator in such a way that its predictions correspond to the re-projection of the triangulated 3D pose and train an auxiliary network on them to produce the final 3D poses. We complement the triangulation with a weighting mechanism that alleviates the impact of noisy predictions caused by self-occlusion or occlusion from other subjects. We demonstrate the effectiveness of our semi-supervised approach on Human3.6M and MPI-INF-3DHP datasets, as well as on a new multi-view multi-person dataset that features occlusion. Soumava Kumar Roy, Leonardo Citraro, Sina Honari, Pascal Fua |
3DV | 4 |
| 2022 | HybridSDF: Combining Deep Implicit Shapes and Geometric Primitives for 3D Shape Representation and ManipulationabstractDeep implicit surfaces excel at modeling generic shapes but do not always capture the regularities present in manufactured objects, which is something simple geometric primitives are particularly good at. In this paper, we propose a representation combining latent and explicit parameters that can be decoded into a set of deep implicit and geometric shapes that are consistent with each other. As a result, we can effectively model both complex and highly regular shapes that coexist in manufactured objects. This enables our approach to manipulate 3D shapes in an efficient and precise manner. Subeesh Vasu, Nicolas Talabot, Artem Lukoianov, Pierre Baqué, Jonathan Donier, Pascal Fua |
3DV | 6 |
| 2022 | DIG: Draping Implicit Garment over the Human Body
Benoît Guillard, Edoardo Remelli, Pascal Fua |
ACCV (1) | 4 |
| 2022 | 3D Pose Based Feedback for Physical Exercises
Sena Kiciroglu, Hugues Vinzant, Isinsu Katircioglu, Mathieu Salzmann, Pascal Fua |
ACCV (4) | 7 |
| 2022 | Adversarial Parametric Pose PriorabstractThe Skinned Multi-Person Linear (SMPL) model represents human bodies by mapping pose and shape parameters to body meshes. However, not all pose and shape parameter values yield physically-plausible or even realistic body meshes. In other words, SMPL is under-constrained and may yield invalid results. We propose learning a prior that restricts the SMPL parameters to values that produce realistic poses via adversarial training. We show that our learned prior covers the diversity of the real-data distribution, facilitates optimization for 3D reconstruction from 2D keypoints, and yields better pose estimates when used for regression from images. For all these tasks, it outperforms the state-of-the-art VAE-based approach to constraining the SMPL parameters. The code will be made available at https://github.com/cvlab-epfl/adv_param_pose_prior. Andrey Davydov, Anastasia Remizova, Victor Constantin, Sina Honari, Mathieu Salzmann, Pascal Fua |
CVPR | 6 |
| 2022 | Leveraging Self-Supervision for Cross-Domain Crowd CountingabstractState-of-the-art methods for counting people in crowded scenes rely on deep networks to estimate crowd density. While effective, these data-driven approaches rely on large amount of data annotation to achieve good performance, which stops these models from being deployed in emergencies during which data annotation is either too costly or cannot be obtained fast enough. One popular solution is to use synthetic data for training. Unfortunately, due to domain shift, the resulting models generalize poorly on real imagery. We remedy this shortcoming by training with both synthetic images, along with their associated labels, and unlabeled real images. To this end, we force our network to learn perspective-aware features by training it to recognize upside-down real images from regular ones and incorporate into it the ability to predict its own uncertainty so that it can generate useful pseudo labels for fine-tuning purposes. This yields an algorithm that consistently outperforms state-of-the-art cross-domain crowd counting ones without any extra computation at inference time. Code is publicly available at https://github.com/weizheliu/Cross-Domain-Crowd-Counting. Weizhe Liu, Nikita Durasov, Pascal Fua |
CVPR | 3 |
| 2022 | Learning to Align Sequential Actions in the WildabstractState-of-the-art methods for self-supervised sequential action alignment rely on deep networks that find correspondences across videos in time. They either learn frame-to-frame mapping across sequences, which does not leverage temporal information, or assume monotonic alignment between each video pair, which ignores variations in the order of actions. As such, these methods are not able to deal with common real-world scenarios that involve background frames or videos that contain non-monotonic sequence of actions. In this paper, we propose an approach to align sequential actions in the wild that involve diverse temporal variations. To this end, we propose an approach to enforce temporal priors on the optimal transport matrix, which leverages temporal consistency, while allowing for variations in the order of actions. Our model accounts for both monotonic and non-monotonic sequences and handles background frames that should not be aligned. We demonstrate that our approach consistently outperforms the state-of-the-art in self-supervised sequential action representation learning on four different benchmark datasets. Code is publicly available at https://github.com/weizheliu/VAVA. Weizhe Liu, Bugra Tekin, Huseyin Coskun, Vibhav Vineet, Pascal Fua, Marc Pollefeys |
CVPR | 5 |
| 2022 | ImplicitAtlas: Learning Deformable Shape Templates in Medical ImagingabstractDeep implicit shape models have become popular in the computer vision community at large but less so for biomed-ical applications. This is in part because large training databases do not exist and in part because biomedical an-notations are often noisy. In this paper, we show that by introducing templates within the deep learning pipeline we can overcome these problems. The proposed framework, named ImplicitAtlas, represents a shape as a deformation field from a learned template field, where multiple templates could be integrated to improve the shape representation ca-pacity at negligible computational cost. Extensive experi-ments on three medical shape datasets prove the superiority over current implicit representation methods. Jiancheng Yang, Udaranga Wickramasinghe, Bingbing Ni, Pascal Fua |
CVPR | 4 |
| 2022 | MeshUDF: Fast and Differentiable Meshing of Unsigned Distance Field Networks
Benoît Guillard, Federico Stella, Pascal Fua |
ECCV (3) | 3 |
| 2022 | Perspective Flow Aggregation for Data-Limited 6D Object Pose Estimation
Yinlin Hu, Pascal Fua, Mathieu Salzmann |
ECCV (2) | 2 |
| 2022 | Learning to Simulate Realistic LiDARsabstractSimulating realistic sensors is a challenging part in data generation for autonomous systems, often involving carefully handcrafted sensor design, scene properties, and physics modeling. To alleviate this, we introduce a pipeline for data-driven simulation of a realistic LiDAR sensor. We propose a model that learns a mapping between RGB images and corresponding LiDAR features such as raydrop or perpoint intensities directly from real datasets. We show that our model can learn to encode realistic effects such as dropped points on transparent surfaces or high intensity returns on reflective materials. When applied to naively raycasted point clouds provided by off-the-shelf simulator software, our model enhances the data by predicting intensities and removing points based on the scene's appearance to match a real LiDAR sensor. We use our technique to learn models of two distinct LiDAR sensors and use them to improve simulated LiDAR data accordingly. Through a sample task of vehicle segmentation, we show that enhancing simulated point clouds with our technique improves downstream task performance. Benoît Guillard, Sai Vemprala, Jayesh K. Gupta, Ondrej Miksik, Vibhav Vineet, Pascal Fua, Ashish Kapoor |
IROS | 6 |
| 2022 | Enforcing Connectivity of 3D Linear Structures Using Their 2D Projections
Doruk Öner, Hussein Osman, Mateusz Kozinski, Pascal Fua |
MICCAI (5) | 4 |
| 2022 | Weakly Supervised Volumetric Image Segmentation with Deformed Templates
Udaranga Wickramasinghe, Patrick M. Jensen, Mian Shah, Jiancheng Yang, Pascal Fua |
MICCAI (5) | 5 |
| 2022 | Neural Annotation Refinement: Development of a New 3D Dataset for Adrenal Gland Analysis
Jiancheng Yang, Udaranga Wickramasinghe, Qikui Zhu, Bingbing Ni, Pascal Fua |
MICCAI (4) | 6 |
| 2022 | GarNet++: Improving Fast and Accurate Static 3D Cloth Draping by Curvature LossabstractIn this paper, we tackle the problem of static 3D cloth draping on virtual human bodies. We introduce a two-stream deep network model that produces a visually plausible draping of a template cloth on virtual 3D bodies by extracting features from both the body and garment shapes. Our network learns to mimic a physics-based simulation (PBS) method while requiring two orders of magnitude less computation time. To train the network, we introduce loss terms inspired by PBS to produce plausible results and make the model collision-aware. To increase the details of the draped garment, we introduce two loss functions that penalize the difference between the curvature of the predicted cloth and PBS. Particularly, we study the impact of mean curvature normal and a novel detail-preserving loss both qualitatively and quantitatively. Our new curvature loss computes the local covariance matrices of the 3D points, and compares the Rayleigh quotients of the prediction and PBS. This leads to more details while performing favorably or comparably against the loss that considers mean curvature normal vectors in the 3D triangulated meshes. We validate our framework on four garment types for various body shapes and poses. Finally, we achieve superior performance against a recently proposed data-driven method. Erhan Gundogdu, Victor Constantin, Shaifali Parashar, Amrollah Seifoddini, Minh Dang, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Self-Supervised Human Detection and Segmentation via Background InpaintingabstractWhile supervised object detection and segmentation methods achieve impressive accuracy, they generalize poorly to images whose appearance significantly differs from the data they have been trained on. To address this when annotating data is prohibitively expensive, we introduce a self-supervised detection and segmentation approach that can work with single images captured by a potentially moving camera. At the heart of our approach lies the observation that object segmentation and background reconstruction are linked tasks, and that, for structured scenes, background regions can be re-synthesized from their surroundings, whereas regions depicting the moving object cannot. We encode this intuition into a self-supervised loss function that we exploit to train a proposal-based segmentation network. To account for the discrete nature of the proposals, we develop a Monte Carlo-based training strategy that allows the algorithm to explore the large space of object proposals. We apply our method to human detection and segmentation in images that visually depart from those of standard benchmarks and outperform existing self-supervised methods. Isinsu Katircioglu, Helge Rhodin, Victor Constantin, Jörg Spörri, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Counting People by Estimating People FlowsabstractModern methods for counting people in crowded scenes rely on deep networks to estimate people densities in individual images. As such, only very few take advantage of temporal consistency in video sequences, and those that do only impose weak smoothness constraints across consecutive frames. In this paper, we advocate estimating people flows across image locations between consecutive images and inferring the people densities from these flows instead of directly regressing them. This enables us to impose much stronger constraints encoding the conservation of the number of people. As a result, it significantly boosts performance without requiring a more complex architecture. Furthermore, it allows us to exploit the correlation between people flow and optical flow to further improve the results. We also show that leveraging people conservation constraints in both a spatial and temporal manner makes it possible to train a deep crowd counting model in an active learning setting with much fewer annotations. This significantly reduces the annotation cost while still leading to similar performance to the full supervision case. Weizhe Liu, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Promoting Connectivity of Network-Like Structures by Enforcing Region SeparationabstractWe propose a novel, connectivity-oriented loss function for training deep convolutional networks to reconstruct network-like structures, like roads and irrigation canals, from aerial images. The main idea behind our loss is to express the connectivity of roads, or canals, in terms of disconnections that they create between background regions of the image. In simple terms, a gap in the predicted road causes two background regions, that lie on the opposite sides of a ground truth road, to touch in prediction. Our loss function is designed to prevent such unwanted connections between background regions, and therefore close the gaps in predicted roads. It also prevents predicting false positive roads and canals by penalizing unwarranted disconnections of background regions. In order to capture even short, dead-ending road segments, we evaluate the loss in small image crops. We show, in experiments on two standard road benchmarks and a new data set of irrigation canals, that convnets trained with our loss function recover road connectivity so well that it suffices to skeletonize their output to produce state of the art maps. A distinct advantage of our approach is that the loss can be plugged in to any existing training setup without further modifications. Doruk Öner, Mateusz Kozinski, Leonardo Citraro, Nathan C. Dadap, Alexandra Georges Konings, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Robust Differentiable SVDabstractEigendecomposition of symmetric matrices is at the heart of many computer vision algorithms. However, the derivatives of the eigenvectors tend to be numerically unstable, whether using the SVD to compute them analytically or using the Power Iteration (PI) method to approximate them. This instability arises in the presence of eigenvalues that are close to each other. This makes integrating eigendecomposition into deep networks difficult and often results in poor convergence, particularly when dealing with large matrices. While this can be mitigated by partitioning the data into small arbitrary groups, doing so has no theoretical basis and makes it impossible to exploit the full power of eigendecomposition. In previous work, we mitigated this using SVD during the forward pass and PI to compute the gradients during the backward pass. However, the iterative deflation procedure required to compute multiple eigenvectors using PI tends to accumulate errors and yield inaccurate gradients. Here, we show that the Taylor expansion of the SVD gradient is theoretically equivalent to the gradient obtained using PI without relying in practice on an iterative process and thus yields more accurate gradients. We demonstrate the benefits of this increased accuracy for image classification and style transfer. Wei Wang 0108, Zheng Dang, Yinlin Hu, Pascal Fua, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Adjusting the Ground Truth Annotations for Connectivity-Based Learning to DelineateabstractDeep learning-based approaches to delineating 3D structure depend on accurate annotations to train the networks. Yet in practice, people, no matter how conscientious, have trouble precisely delineating in 3D and on a large scale, in part because the data is often hard to interpret visually and in part because the 3D interfaces are awkward to use. In this paper, we introduce a method that explicitly accounts for annotation inaccuracies. To this end, we treat the annotations as active contour models that can deform themselves while preserving their topology. This enables us to jointly train the network and correct potential errors in the original annotations. The result is an approach that boosts performance of deep networks trained with potentially inaccurate annotations. Doruk Öner, Mateusz Kozinski, Leonardo Citraro, Pascal Fua |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Leveraging Spatial and Photometric Context for Calibrated Non-Lambertian Photometric StereoabstractThe problem of estimating a surface shape from its observed reflectance properties still remains a challenging task in computer vision. The presence of global illumination effects such as inter-reflections or cast shadows makes the task particularly difficult for non-convex real-world surfaces. State-of-the-art methods for calibrated photometric stereo address these issues using convolutional neural networks (CNNs) that primarily aim to capture either the spatial context among adjacent pixels or the photometric one formed by illuminating a sample from adjacent directions.In this paper, we bridge these two objectives and introduce an efficient fully-convolutional architecture that can leverage both spatial and photometric context simultaneously. In contrast to existing approaches that rely on standard 2D CNNs and regress directly to surface normals, we argue that using separable 4D convolutions and regressing to 2D Gaussian heat-maps severely reduces the size of the network and leads to more stable predictions.Our experimental results on a real-world photometric stereo benchmark show that the proposed approach outperforms the existing published methods in accuracy. The source code for our method is available at https://github.com/DawyD/UNet-PS-4D. David Honzátko, Engin Türetken, Pascal Fua, L. Andrea Dunbar |
3DV | 3 |
| 2021 | Masksembles for Uncertainty EstimationabstractDeep neural networks have amply demonstrated their prowess but estimating the reliability of their predictions remains challenging. Deep Ensembles are widely considered as being one of the best methods for generating uncertainty estimates but are very expensive to train and evaluate.MC-Dropout is another popular alternative, which is less expensive, but also less reliable. Our central intuition is that there is a continuous spectrum of ensemble-like models of which MC-Dropout and Deep Ensembles are extreme examples. The first one uses effectively infinite number of highly correlated models while the second one relies on a finite number of independent models.To combine the benefits of both, we introduce Masksembles. Instead of randomly dropping parts of the network as in MC-dropout, Masksemble relies on a fixed number of binary masks, which are parameterized in a way that allows to change correlations between individual models. Namely, by controlling the overlap between the masks and their size one can choose the optimal configuration for the task at hand. This leads to a simple and easy to implement method with performance on par with Ensembles at a fraction of the cost. We experimentally validate Masksembles on two widely used datasets, CIFAR10 and ImageNet. Nikita Durasov, Timur M. Bagautdinov, Pierre Baqué, Pascal Fua |
CVPR | 4 |
| 2021 | Wide-Depth-Range 6D Object Pose Estimation in Spaceabstract6D pose estimation in space poses unique challenges that are not commonly encountered in the terrestrial setting. One of the most striking differences is the lack of atmospheric scattering, allowing objects to be visible from a great distance while complicating illumination conditions. Currently available benchmark datasets do not place a sufficient emphasis on this aspect and mostly depict the target in close proximity.Prior work tackling pose estimation under large scale variations relies on a two-stage approach to first estimate scale, followed by pose estimation on a resized image patch. We instead propose a single-stage hierarchical end-to-end trainable network that is more robust to scale variations. We demonstrate that it outperforms existing approaches not only on images synthesized to resemble images taken in space but also on standard benchmarks. Yinlin Hu, Sébastien Speierer, Wenzel Jakob, Pascal Fua, Mathieu Salzmann |
CVPR | 4 |
| 2021 | Deep Active Surface ModelsabstractActive Surface Models have a long history of being useful to model complex 3D surfaces. But only Active Contours have been used in conjunction with deep networks, and then only to produce the data term as well as meta-parameter maps controlling them. In this paper, we advocate a much tighter integration. We introduce layers that implement them that can be integrated seamlessly into Graph Convolutional Networks to enforce sophisticated smoothness priors at an acceptable computational cost.We will show that the resulting Deep Active Surface Models outperform equivalent architectures that use traditional regularization loss terms to impose smoothness priors for 3D surface reconstruction from 2D images and for 3D volume segmentation. Udaranga Wickramasinghe, Pascal Fua, Graham Knott |
CVPR | 2 |
| 2021 | PCLs: Geometry-Aware Neural Reconstruction of 3D Pose With Perspective Crop LayersabstractLocal processing is an essential feature of CNNs and other neural network architectures—it is one of the reasons why they work so well on images where relevant information is, to a large extent, local. However, perspective effects stemming from the projection in a conventional camera vary for different global positions in the image. We introduce Perspective Crop Layers (PCLs)—a form of perspective crop of the region of interest based on the camera geometry— and show that accounting for the perspective consistently improves the accuracy of state-of-the-art 3D pose reconstruction methods. PCLs are modular neural network layers, which, when inserted into existing CNN and MLP architectures, deterministically remove the location-dependent perspective effects while leaving end-to-end training and the number of parameters of the underlying neural network unchanged. We demonstrate that PCL leads to improved 3D human pose reconstruction accuracy for CNN architectures that use cropping operations, such as spatial transformer networks (STN), and, somewhat surprisingly, MLPs used for 2D-to-3D key-point lifting. Our conclusion is that it is important to utilize camera calibration information when available, for classical and deep-learning-based computer vision alike. PCL offers an easy way to improve the accuracy of existing 3D reconstruction networks by making them geometry-aware. Our code is publicly available at github.com/yu-frank/PerspectiveCropLayers. Frank Yu, Mathieu Salzmann, Pascal Fua, Helge Rhodin |
CVPR | 3 |
| 2021 | Temporally-Coherent Surface Reconstruction via Metric-Consistent AtlasesabstractWe propose a method for the unsupervised reconstruction of a temporally-coherent sequence of surfaces from a sequence of time-evolving point clouds, yielding dense, semantically meaningful correspondences between all keyframes. We represent the reconstructed surface as an atlas, using a neural network. Using canonical correspondences defined via the atlas, we encourage the reconstruction to be as isometric as possible across frames, leading to semantically-meaningful reconstruction. Through experiments and comparisons, we empirically show that our method achieves results that exceed that state of the art in the accuracy of unsupervised correspondences and accuracy of surface reconstruction. Jan Bednarík, Vladimir G. Kim, Siddhartha Chaudhuri, Shaifali Parashar, Mathieu Salzmann, Pascal Fua, Noam Aigerman |
ICCV | 6 |
| 2021 | Sketch2Mesh: Reconstructing and Editing 3D Shapes from SketchesabstractReconstructing 3D shape from 2D sketches has long been an open problem because the sketches only provide very sparse and ambiguous information. In this paper, we use an encoder/decoder architecture for the sketch to mesh translation. When integrated into a user interface that provides camera parameters for the sketches, this enables us to leverage its latent parametrization to represent and refine a 3D mesh so that its projections match the external contours outlined in the sketch. We will show that this approach is easy to deploy, robust to style changes, and effective. Furthermore, it can be used for shape refinement given only single pen strokes.We compare our approach to state-of-the-art methods on sketches—both hand-drawn and synthesized—and demonstrate that we outperform them. Benoît Guillard, Edoardo Remelli, Pierre Yvernay, Pascal Fua |
ICCV | 4 |
| 2021 | Human Detection and Segmentation via Multi-view ConsensusabstractSelf-supervised detection and segmentation of foreground objects aims for accuracy without annotated training data. However, existing approaches predominantly rely on restrictive assumptions on appearance and motion.For scenes with dynamic activities and camera motion, we propose a multi-camera framework in which geometric constraints are embedded in the form of multi-view consistency during training via coarse 3D localization in a voxel grid and fine-grained offset regression. In this manner, we learn a joint distribution of proposals over multiple views. At inference time, our method operates on single RGB images. We outperform state-of-the-art techniques both on images that visually depart from those of standard benchmarks and on those of the classical Human3.6M dataset. Isinsu Katircioglu, Helge Rhodin, Jörg Spörri, Mathieu Salzmann, Pascal Fua |
ICCV | 5 |
| 2021 | Image Matching Across Wide Baselines: From Paper to Practice
Yuhe Jin, Dmytro Mishkin, Anastasiia Mishchuk, Jiri Matas, Pascal Fua, Kwang Moo Yi, Eduard Trulls |
Int. J. Comput. Vis. | 5 |
| 2021 | Defect segmentation for multi-illumination quality control systemsabstractAbstract Thanks to recent advancements in image processing and deep learning techniques, visual surface inspection in production lines has become an automated process as long as all the defects are visible in a single or a few images. However, it is often necessary to inspect parts under many different illumination conditions to capture all the defects. Training deep networks to perform this task requires large quantities of annotated data, which are rarely available and cumbersome to obtain. To alleviate this problem, we devised an original augmentation approach that, given a small image collection, generates rotated versions of the images while preserving illumination effects, something that random rotations cannot do. We introduce three real multi-illumination datasets, on which we demonstrate the effectiveness of our illumination preserving rotation approach. Training deep neural architectures with our approach delivers a performance increase of up to 51% in terms of AuPRC score over using standard rotations to perform data augmentation. David Honzátko, Engin Türetken, Siavash Arjomand Bigdeli, L. Andrea Dunbar, Pascal Fua |
Mach. Vis. Appl. | 5 |
| 2021 | Eigendecomposition-Free Training of Deep Networks for Linear Least-Square ProblemsabstractMany classical Computer Vision problems, such as essential matrix computation and pose estimation from 3D to 2D correspondences, can be tackled by solving a linear least-square problem, which can be done by finding the eigenvector corresponding to the smallest, or zero, eigenvalue of a matrix representing a linear system. Incorporating this in deep learning frameworks would allow us to explicitly encode known notions of geometry, instead of having the network implicitly learn them from data. However, performing eigendecomposition within a network requires the ability to differentiate this operation. While theoretically doable, this introduces numerical instability in the optimization process in practice. In this paper, we introduce an eigendecomposition-free approach to training a deep network whose loss depends on the eigenvector corresponding to a zero eigenvalue of a matrix predicted by the network. We demonstrate that our approach is much more robust than explicit differentiation of the eigendecomposition using two general tasks, outlier rejection and denoising, with several practical examples including wide-baseline stereo, the perspective-n-point problem, and ellipse fitting. Empirically, our method has better convergence properties and yields state-of-the-art results. Zheng Dang, Kwang Moo Yi, Yinlin Hu, Fei Wang 0008, Pascal Fua, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Matching Seqlets: An Unsupervised Approach for Locality Preserving Sequence MatchingabstractIn this paper, we propose a novel unsupervised approach for sequence matching by explicitly accounting for the locality properties in the sequences. In contrast to conventional approaches that rely on frame-to-frame matching, we conduct matching using sequencelet or seqlet, a sub-sequence wherein the frames share strong similarities and are thus grouped together. The optimal seqlets and matching between them are learned jointly, without any supervision from users. The learned seqlets preserve the locality information at the scale of interest and resolve the ambiguities during matching, which are omitted by frame-based matching methods. We show that our proposed approach outperforms the state-of-the-art ones on datasets of different domains including human actions, facial expressions, speech, and character strokes. Jiayan Qiu, Xinchao Wang, Pascal Fua, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Better Patch Stitching for Parametric Surface ReconstructionabstractRecently, parametric mappings have emerged as highly effective surface representations, yielding low reconstruction error. In particular, the latest works represent the target shape as an atlas of multiple mappings, which can closely encode object parts. Atlas representations, however, suffer from one major drawback: The individual mappings are not guaranteed to be consistent, which results in holes in the reconstructed shape or in jagged surface areas.We introduce an approach that explicitly encourages global consistency of the local mappings. To this end, we introduce two novel loss terms. The first term exploits the surface normals and requires that they remain locally consistent when estimated within and across the individual mappings. The second term further encourages better spatial configuration of the mappings by minimizing novel stitching error. We show on standard benchmarks that the use of normal consistency requirement outperforms the baselines quantitatively while enforcing better stitching leads to much better visual quality of the reconstructed objects as compared to the state-of-the-art. Zhantao Deng, Jan Bednarík, Mathieu Salzmann, Pascal Fua |
3DV | 4 |
| 2020 | Motion Prediction Using Temporal Inception Module
Tim Lebailly, Sena Kiciroglu, Mathieu Salzmann, Pascal Fua, Wei Wang 0108 |
ACCV (2) | 4 |
| 2020 | Shape Reconstruction by Learning Differentiable Surface RepresentationsabstractGenerative models that produce point clouds have emerged as a powerful tool to represent 3D surfaces, and the best current ones rely on learning an ensemble of parametric representations. Unfortunately, they offer no control over the deformations of the surface patches that form the ensemble and thus fail to prevent them from either overlapping or collapsing into single points or lines. As a consequence, computing shape properties such as surface normals and curvatures becomes difficult and unreliable. In this paper, we show that we can exploit the inherent differentiability of deep networks to leverage differential surface properties during training so as to prevent patch collapse and strongly reduce patch overlap. Furthermore, this lets us reliably compute quantities such as surface normals and curvatures. We will demonstrate on several tasks that this yields more accurate surface reconstructions than the state-of-the-art methods in terms of normals estimation and amount of collapsed and overlapped patches. Jan Bednarík, Shaifali Parashar, Erhan Gundogdu, Mathieu Salzmann, Pascal Fua |
CVPR | 5 |
| 2020 | Single-Stage 6D Object Pose EstimationabstractMost recent 6D pose estimation frameworks first rely on a deep network to establish correspondences between 3D object keypoints and 2D image locations and then use a variant of a RANSAC-based Perspective-n-Point (PnP) algorithm. This two-stage process, however, is suboptimal: First, it is not end-to-end trainable. Second, training the deep network relies on a surrogate loss that does not directly reflect the final 6D pose estimation task. In this work, we introduce a deep architecture that directly regresses 6D poses from correspondences. It takes as input a group of candidate correspondences for each 3D keypoint and accounts for the fact that the order of the correspondences within each group is irrelevant, while the order of the groups, that is, of the 3D keypoints, is fixed. Our architecture is generic and can thus be exploited in conjunction with existing correspondence-extraction networks so as to yield single-stage 6D pose estimation frameworks. Our experiments demonstrate that these single-stage frameworks consistently outperform their two-stage counterparts in terms of both accuracy and speed. Yinlin Hu, Pascal Fua, Wei Wang 0108, Mathieu Salzmann |
CVPR | 2 |
| 2020 | ActiveMoCap: Optimized Viewpoint Selection for Active Human Motion CaptureabstractThe accuracy of monocular 3D human pose estimation depends on the viewpoint from which the image is captured. While freely moving cameras, such as on drones, provide control over this viewpoint, automatically positioning them at the location which will yield the highest accuracy remains an open problem. This is the problem that we address in this paper. Specifically, given a short video sequence, we introduce an algorithm that predicts which viewpoints should be chosen to capture future frames so as to maximize 3D human pose estimation accuracy. The key idea underlying our approach is a method to estimate the uncertainty of the 3D body pose estimates. We integrate several sources of uncertainty, originating from deep learning based regressors and temporal smoothness. Our motion planner yields improved 3D body pose estimates and outperforms or matches existing ones that are based on person following and orbiting. Sena Kiciroglu, Helge Rhodin, Sudipta N. Sinha, Mathieu Salzmann, Pascal Fua |
CVPR | 5 |
| 2020 | Deformation-Aware Unpaired Image Translation for Pose Estimation on Laboratory AnimalsabstractOur goal is to capture the pose of real animals using synthetic training examples, without using any manual supervision. Our focus is on neuroscience model organisms, to be able to study how neural circuits orchestrate behaviour. Human pose estimation attains remarkable accuracy when trained on real or simulated datasets consisting of millions of frames. However, for many applications simulated models are unrealistic and real training datasets with comprehensive annotations do not exist. We address this problem with a new sim2real domain transfer method. Our key contribution is the explicit and independent modeling of appearance, shape and pose in an unpaired image translation framework. Our model lets us train a pose estimator on the target domain by transferring readily available body keypoint locations from the source domain to generated target images. We compare our approach with existing domain transfer methods and demonstrate improved pose estimation accuracy on Drosophila melanogaster (fruit fly), Caenorhabditis elegans (worm) and Danio rerio (zebrafish), without requiring any manual annotation on the target domain and despite using simplistic off-the-shelf animal characters for simulation, or simple geometric shapes as models. Our new datasets, code and trained models will be published to support future computer vision and neuroscientific studies. Semih Günel, Mirela Ostrek, Pavan Ramdya, Pascal Fua, Helge Rhodin |
CVPR | 5 |
| 2020 | Local Non-Rigid Structure-From-Motion From Diffeomorphic MappingsabstractWe propose a new formulation to the non-rigid structure-from-motion problem that only requires the deforming surface to meaning that its differential structure is preserved. This is a much weaker assumption than the traditional ones of isometry or conformality. We show that it is nevertheless sufficient to establish local correspondences between the surface in two different images and therefore to perform point-wise reconstruction using only up to first-order derivatives. We formulate differential constraints and solve them algebraically using the theory of resultants. We will demonstrate that our approach is more widely applicable, more stable in noisy and sparse imaging conditions and much faster than earlier ones, while delivering similar accuracy. The code is available at https//github.com/cvlab-epf1/diff-nrsfm/. Shaifali Parashar, Mathieu Salzmann, Pascal Fua |
CVPR | 3 |
| 2020 | Lightweight Multi-View 3D Pose Estimation Through Camera-Disentangled RepresentationabstractWe present a lightweight solution to recover 3D pose from multi-view images captured with spatially calibrated cameras. Building upon recent advances in interpretable representation learning, we exploit 3D geometry to fuse input images into a unified latent representation of pose, which is disentangled from camera view-points. This allows us to reason effectively about 3D pose across different views without using compute-intensive volumetric grids. Our architecture then conditions the learned representation on camera projection operators to produce accurate per-view 2d detections, that can be simply lifted to 3D via a differentiable Direct Linear Transform (DLT) layer. In order to do it efficiently, we propose a novel implementation of DLT that is orders of magnitude faster on GPU architectures than standard SVD-based triangulation methods. We evaluate our approach on two large-scale human pose datasets (H36M and Total Capture): our method outperforms or performs comparably to the state-of-the-art volumetric methods, while, unlike them, yielding real-time performance. Edoardo Remelli, Shangchen Han, Sina Honari, Pascal Fua, Robert Wang 0002 |
CVPR | 4 |
| 2020 | Towards Reliable Evaluation of Algorithms for Road Network Reconstruction from Aerial Images
Leonardo Citraro, Mateusz Kozinski, Pascal Fua |
ECCV (28) | 3 |
| 2020 | Estimating People Flows to Better Count Them in Crowded Scenes
Weizhe Liu, Mathieu Salzmann, Pascal Fua |
ECCV (15) | 3 |
| 2020 | TopoAL: An Adversarial Learning Approach for Topology-Aware Road Segmentation
Subeesh Vasu, Mateusz Kozinski, Leonardo Citraro, Pascal Fua |
ECCV (27) | 4 |
| 2020 | Domain Adaptive Multibranch Networks
Róger Bermúdez-Chacón, Mathieu Salzmann, Pascal Fua |
ICLR | 3 |
| 2020 | Voxel2Mesh: 3D Mesh Model Generation from Volumetric Data
Udaranga Wickramasinghe, Edoardo Remelli, Graham Knott, Pascal Fua |
MICCAI (4) | 4 |
| 2020 | UCLID-Net: Single View Reconstruction in Object SpaceabstractMost state-of-the-art deep geometric learning single-view reconstruction approaches rely on encoder-decoder architectures that output either shape parametrizations or implicit representations. However, these representations rarely preserve the Euclidean structure of the 3D space objects exist in. In this paper, we show that building a geometry preserving 3-dimensional latent space helps the network concurrently learn global shape regularities and local reasoning in the object coordinate space and, as a result, boosts performance. We demonstrate both on ShapeNet synthetic images, which are often used for benchmarking purposes, and on real-world images that our approach outperforms state-of-the-art ones. Furthermore, the single-view pipeline naturally extends to multi-view reconstruction, which we also show. Benoît Guillard, Edoardo Remelli, Pascal Fua |
NeurIPS | 3 |
| 2020 | MeshSDF: Differentiable Iso-Surface ExtractionabstractGeometric Deep Learning has recently made striking progress with the advent of continuous Deep Implicit Fields. They allow for detailed modeling of watertight surfaces of arbitrary topology while not relying on a 3D Euclidean grid, resulting in a learnable parameterization that is not limited in resolution. Unfortunately, these methods are often not suitable for applications that require an explicit mesh-based surface representation because converting an implicit field to such a representation relies on the Marching Cubes algorithm, which cannot be differentiated with respect to the underlying implicit field. In this work, we remove this limitation and introduce a differentiable way to produce explicit surface mesh representations from Deep Signed Distance Functions. Our key insight is that by reasoning on how implicit field perturbations impact local surface geometry, one can ultimately differentiate the 3D location of surface samples with respect to the underlying deep implicit field. We exploit this to define MeshSDF, an end-to-end differentiable mesh representation which can vary its topology. We use two different applications to validate our theoretical insight: Single-View Reconstruction via Differentiable Rendering and Physically-Driven Shape Optimization. In both cases our differentiable parameterization gives us an edge over state-of-the-art algorithms. Edoardo Remelli, Artem Lukoianov, Stephan R. Richter, Benoît Guillard, Timur M. Bagautdinov, Pierre Baqué, Pascal Fua |
NeurIPS | 7 |
| 2020 | DISK: Learning local features with policy gradientabstractLocal feature frameworks are difficult to learn in an end-to-end fashion due to the discreteness inherent to the selection and matching of sparse keypoints. We introduce DISK (DIScrete Keypoints), a novel method that overcomes these obstacles by leveraging principles from Reinforcement Learning (RL), optimizing end-to-end for a high number of correct feature matches. Our simple yet expressive probabilistic model lets us keep the training and inference regimes close, while maintaining good enough convergence properties to reliably train from scratch. Our features can be extracted very densely while remaining discriminative, challenging commonly held assumptions about what constitutes a good keypoint, as showcased in Fig. 1, and deliver state-of-the-art results on three public benchmarks. Michal J. Tyszkiewicz, Pascal Fua, Eduard Trulls |
NeurIPS | 2 |
| 2020 | Tracing in 2D to reduce the annotation effort for 3D deep delineation of linear structures
Mateusz Kozinski, Agata Mosinska, Mathieu Salzmann, Pascal Fua |
Medical Image Anal. | 4 |
| 2020 | Real-time camera pose estimation for sports fields
Leonardo Citraro, Pablo Márquez-Neila, Stefano Savare, Vivek Jayaram, Charles Dubout, Félix Renaut, Andres Hasfura, Horesh Ben Shitrit, Pascal Fua |
Mach. Vis. Appl. | 9 |
| 2020 | Joint Segmentation and Path Classification of Curvilinear StructuresabstractDetection of curvilinear structures in images has long been of interest. One of the most challenging aspects of this problem is inferring the graph representation of the curvilinear network. Most existing delineation approaches first perform binary segmentation of the image and then refine it using either a set of hand-designed heuristics or a separate classifier that assigns likelihood to paths extracted from the pixel-wise prediction. In our work, we bridge the gap between segmentation and path classification by training a deep network that performs those two tasks simultaneously. We show that this approach is beneficial because it enforces consistency across the whole processing pipeline. We apply our approach on roads and neurons datasets. Agata Mosinska, Mateusz Kozinski, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Visual Correspondences for Unsupervised Domain Adaptation on Electron Microscopy ImagesabstractWe present an Unsupervised Domain Adaptation strategy to compensate for domain shifts on Electron Microscopy volumes. Our method aggregates visual correspondences-motifs that are visually similar across different acquisitions-to infer changes on the parameters of pretrained models, and enable them to operate on new data. In particular, we examine the annotations of an existing acquisition to determine pivot locations that characterize the reference segmentation, and use a patch matching algorithm to find their candidate visual correspondences in a new volume. We aggregate all the candidate correspondences by a voting scheme and we use them to construct a consensus heatmap: a map of how frequently locations on the new volume are matched to relevant locations from the original acquisition. This information allows us to perform model adaptations in two different ways: either by a) optimizing model parameters under a Multiple Instance Learning formulation, so that predictions between reference locations and their sets of correspondences agree, or by b) using high-scoring regions of the heatmap as soft labels to be incorporated in other domain adaptation pipelines, including deep learning ones. We show that these unsupervised techniques allow us to obtain high-quality segmentations on unannotated volumes, qualitatively consistent with results obtained under full supervision, for both mitochondria and synapses, with no need for new annotation effort. Róger Bermúdez-Chacón, Okan Altingövde, Carlos J. Becker, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Medical Imaging | 5 |
| 2020 | XNect: real-time multi-person 3D motion capture with a single RGB cameraabstractWe present a real-time approach for multi-person 3D motion capture at over 30 fps using a single RGB camera. It operates successfully in generic scenes which may contain occlusions by objects and by other people. Our method operates in subsequent stages. The first stage is a convolutional neural network (CNN) that estimates 2D and 3D pose features along with identity assignments for all visible joints of all individuals. We contribute a new architecture for this CNN, called SelecSLS Net , that uses novel selective long and short range skip connections to improve the information flow allowing for a drastically faster network without compromising accuracy. In the second stage, a fullyconnected neural network turns the possibly partial (on account of occlusion) 2D pose and 3D pose features for each subject into a complete 3D pose estimate per individual. The third stage applies space-time skeletal model fitting to the predicted 2D and 3D pose per subject to further reconcile the 2D and 3D pose, and enforce temporal coherence. Our method returns the full skeletal pose in joint angles for each subject. This is a further key distinction from previous work that do not produce joint angle results of a coherent skeleton in real time for multi-person scenes. The proposed system runs on consumer hardware at a previously unseen speed of more than 30 fps given 512x320 images as input while achieving state-of-the-art accuracy, which we will demonstrate on a range of challenging real-world scenes. Dushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller 0001, Weipeng Xu, Mohamed A. Elgharib, Pascal Fua, Hans-Peter Seidel, Helge Rhodin, Gerard Pons-Moll, Christian Theobalt |
ACM Trans. Graph. | 6 |
| 2019 | Motion Capture from Pan-Tilt Cameras with Unknown OrientationabstractIn sports, such as alpine skiing, coaches would like to know the speed and various biomechanical variables of their athletes and competitors. Existing methods use either body-worn sensors, which are cumbersome to setup, or manual image annotation, which is time consuming. We propose a method for estimating an athlete's global 3D position and articulated pose using multiple cameras. By contrast to classical markerless motion capture solutions, we allow cameras to rotate freely so that large capture volumes can be covered. In a first step, tight crops around the skier are predicted and fed to a 2D pose estimator network. The 3D pose is then reconstructed using a bundle adjustment method. Key to our solution is the rotation estimation of Pan-Tilt cameras in a joint optimization with the athlete pose and conditioning on relative background motion computed with feature tracking. Furthermore, we created a new alpine skiing dataset and annotated it with 2D pose labels, to overcome shortcomings of existing ones. Our method estimates accurate global 3D poses from images only and provides coaches with an automatic and fast tool for measuring and improving an athlete's performance. Roman Bachmann 0001, Jörg Spörri, Pascal Fua, Helge Rhodin |
3DV | 3 |
| 2019 | Segmentation-Driven 6D Object Pose EstimationabstractThe most recent trend in estimating the 6D pose of rigid objects has been to train deep networks to either directly regress the pose from the image or to predict the 2D locations of 3D keypoints, from which the pose can be obtained using a PnP algorithm. In both cases, the object is treated as a global entity, and a single pose estimate is computed. As a consequence, the resulting techniques can be vulnerable to large occlusions. In this paper, we introduce a segmentation-driven 6D pose estimation framework where each visible part of the objects contributes a local pose prediction in the form of 2D keypoint locations. We then use a predicted measure of confidence to combine these pose candidates into a robust set of 3D-to-2D correspondences, from which a reliable pose estimate can be obtained. We outperform the state-of-the-art on the challenging Occluded-LINEMOD and YCB-Video datasets, which is evidence that our approach deals well with multiple poorly-textured objects occluding each other. Furthermore, it relies on a simple enough architecture to achieve real-time performance. Yinlin Hu, Joachim Hugonot, Pascal Fua, Mathieu Salzmann |
CVPR | 3 |
| 2019 | Context-Aware Crowd CountingabstractState-of-the-art methods for counting people in crowded scenes rely on deep networks to estimate crowd density. They typically use the same filters over the whole image or over large image patches. Only then do they estimate local scale to compensate for perspective distortion. This is typically achieved by training an auxiliary classifier to select, for predefined image patches, the best kernel size among a limited set of choices. As such, these methods are not end-to-end trainable and restricted in the scope of context they can leverage. In this paper, we introduce an end-to-end trainable deep architecture that combines features obtained using multiple receptive field sizes and learns the importance of each such feature at each image location. In other words, our approach adaptively encodes the scale of the contextual information required to accurately predict crowd density. This yields an algorithm that outperforms state-of-the-art crowd counting methods, especially when perspective effects are strong. Weizhe Liu, Mathieu Salzmann, Pascal Fua |
CVPR | 3 |
| 2019 | Eliminating Exposure Bias and Metric Mismatch in Multiple Object TrackingabstractIdentity Switching remains one of the main difficulties Multiple Object Tracking (MOT) algorithms have to deal with. Many state-of-the-art approaches now use sequence models to solve this problem but their training can be affected by biases that decrease their efficiency. In this paper, we introduce a new training procedure that confronts the algorithm to its own mistakes while explicitly attempting to minimize the number of switches, which results in better training. We propose an iterative scheme of building a rich training set and using it to learn a scoring function that is an explicit proxy for the target tracking metric. Whether using only simple geometric features or more sophisticated ones that also take appearance into account, our approach outperforms the state-of-the-art on several MOT benchmarks. Andrii Maksai, Pascal Fua |
CVPR | 2 |
| 2019 | Neural Scene Decomposition for Multi-Person Motion CaptureabstractLearning general image representations has proven key to the success of many computer vision tasks. For example, many approaches to image understanding problems rely on deep networks that were initially trained on ImageNet, mostly because the learned features are a valuable starting point to learn from limited labeled data. However, when it comes to 3D motion capture of multiple people, these features are only of limited use. In this paper, we therefore propose an approach to learning features that are useful for this purpose. To this end, we introduce a self-supervised approach to learning what we call a neural scene decomposition (NSD) that can be exploited for 3D pose estimation. NSD comprises three layers of abstraction to represent human subjects: spatial layout in terms of bounding-boxes and relative depth; a 2D shape representation in terms of an instance segmentation mask; and subject-specific appearance and 3D pose information. By exploiting self-supervision coming from multiview data, our NSD model can be trained end-to-end without any 2D or 3D supervision. In contrast to previous approaches, it works for multiple persons and full-frame images. Because it encodes 3D geometry, NSD can then be effectively leveraged to train a 3D pose estimation network from small amounts of annotated data. Helge Rhodin, Victor Constantin, Isinsu Katircioglu, Mathieu Salzmann, Pascal Fua |
CVPR | 5 |
| 2019 | Recurrent U-Net for Resource-Constrained SegmentationabstractState-of-the-art segmentation methods rely on very deep networks that are not always easy to train without very large training datasets and tend to be relatively slow to run on standard GPUs. In this paper, we introduce a novel recurrent U-Net architecture that preserves the compactness of the original U-Net [33], while substantially increasing its performance to the point where it outperforms the state of the art on several benchmarks. We will demonstrate its effectiveness for several tasks, including hand segmentation, retina vessel segmentation, and road segmentation. We also introduce a large-scale dataset for hand segmentation. Wei Wang 0108, Kaicheng Yu, Joachim Hugonot, Pascal Fua, Mathieu Salzmann |
ICCV | 4 |
| 2019 | Gravity as a Reference for Estimating a Person's Height From VideoabstractEstimating the metric height of a person from monocular imagery without additional assumptions is ill-posed. Existing solutions either require manual calibration of ground plane and camera geometry, special cameras, or reference objects of known size. We focus on motion cues and exploit gravity on earth as an omnipresent reference 'object' to translate acceleration, and subsequently height, measured in image-pixels to values in meters. We require videos of motion as input, where gravity is the only external force. This limitation is different to those of existing solutions that recover a person's height and, therefore, our method opens up new application fields. We show theoretically and empirically that a simple motion trajectory analysis suffices to translate from pixel measurements to the person's metric height, reaching a MAE of up to 3.9 cm on jumping motions, and that this works without camera and ground plane calibration. Didier Bieler, Semih Günel, Pascal Fua, Helge Rhodin |
ICCV | 3 |
| 2019 | Beyond Cartesian Representations for Local DescriptorsabstractThe dominant approach for learning local patch descriptors relies on small image regions whose scale must be properly estimated a priori by a keypoint detector. In other words, if two patches are not in correspondence, their descriptors will not match. A strategy often used to alleviate this problem is to “pool” the pixel-wise features over log-polar regions, rather than regularly spaced ones. By contrast, we propose to extract the “support region” directly with a log-polar sampling scheme. We show that this provides us with a better representation by simultaneously oversampling the immediate neighbourhood of the point and undersampling regions far away from it. We demonstrate that this representation is particularly amenable to learning descriptors with deep networks. Our models can match descriptors across a much wider range of scales than was possible before, and also leverage much larger support regions without suffering from occlusions. We report state-of-the-art results on three different datasets. Patrick Ebel 0002, Eduard Trulls, Kwang Moo Yi, Pascal Fua, Anastasiia Mishchuk |
ICCV | 4 |
| 2019 | GarNet: A Two-Stream Network for Fast and Accurate 3D Cloth DrapingabstractWhile Physics-Based Simulation (PBS) can accurately drape a 3D garment on a 3D body, it remains too costly for real-time applications, such as virtual try-on. By contrast, inference in a deep network, requiring a single forward pass, is much faster. Taking advantage of this, we propose a novel architecture to fit a 3D garment template to a 3D body. Specifically, we build upon the recent progress in 3D point cloud processing with deep networks to extract garment features at varying levels of detail, including point-wise, patch-wise and global features. We fuse these features with those extracted in parallel from the 3D body, so as to model the cloth-body interactions. The resulting two-stream architecture, which we call as GarNet, is trained using a loss function inspired by physics-based modeling, and delivers visually plausible garment shapes whose 3D points are, on average, less than 1 cm away from those of a PBS method, while running 100 times faster. Moreover, the proposed method can model various garment types with different cutting patterns when parameters of those patterns are given as input to the network. Erhan Gundogdu, Victor Constantin, Amrollah Seifoddini, Minh Dang, Mathieu Salzmann, Pascal Fua |
ICCV | 6 |
| 2019 | Detecting the Unexpected via Image ResynthesisabstractClassical semantic segmentation methods, including the recent deep learning ones, assume that all classes observed at test time have been seen during training. In this paper, we tackle the more realistic scenario where unexpected objects of unknown classes can appear at test time. The main trends in this area either leverage the notion of prediction uncertainty to flag the regions with low confidence as unknown, or rely on autoencoders and highlight poorly-decoded regions. Having observed that, in both cases, the detected regions typically do not correspond to unexpected objects, in this paper, we introduce a drastically different strategy: It relies on the intuition that the network will produce spurious labels in regions depicting unexpected objects. Therefore, resynthesizing the image from the resulting semantic map will yield significant appearance differences with respect to the input image. In other words, we translate the problem of detecting unknown classes to one of identifying poorly-resynthesized image regions. We show that this outperforms both uncertainty- and autoencoder-based methods. Krzysztof Lis, Krishna K. Nakka, Pascal Fua, Mathieu Salzmann |
ICCV | 3 |
| 2019 | Geometric and Physical Constraints for Drone-Based Head Plane Crowd Density EstimationabstractState-of-the-art methods for counting people in crowded scenes rely on deep networks to estimate crowd density in the image plane. While useful for this purpose, this image-plane density has no immediate physical meaning because it is subject to perspective distortion. This is a concern in sequences acquired by drones because the viewpoint changes often. This distortion is usually handled implicitly by either learning scale-invariant features or estimating density in patches of different sizes, neither of which accounts for the fact that scale changes must be consistent over the whole scene. In this paper, we explicitly model the scale changes and reason in terms of people per square-meter. We show that feeding the perspective model to the network allows us to enforce global scale consistency and that this model can be obtained on the fly from the drone sensors. In addition, it also enables us to enforce physically-inspired temporal consistency constraints that do not have to be learned. This yields an algorithm that outperforms state-of-the-art methods in inferring crowd density from a moving drone camera especially when perspective effects are strong. Weizhe Liu, Krzysztof Lis, Mathieu Salzmann, Pascal Fua |
IROS | 4 |
| 2019 | Probabilistic Atlases to Enforce Topological Constraints
Udaranga Wickramasinghe, Graham Knott, Pascal Fua |
MICCAI (1) | 3 |
| 2019 | Backpropagation-Friendly EigendecompositionabstractEigendecomposition (ED) is widely used in deep networks. However, the backpropagation of its results tends to be numerically unstable, whether using ED directly or approximating it with the Power Iteration method, particularly when dealing with large matrices. While this can be mitigated by partitioning the data in small and arbitrary groups, doing so has no theoretical basis and makes its impossible to exploit the power of ED to the full. In this paper, we introduce a numerically stable and differentiable approach to leveraging eigenvectors in deep networks. It can handle large matrices without requiring to split them. We demonstrate the better robustness of our approach over standard ED and PI for ZCA whitening, an alternative to batch normalization, and for PCA denoising, which we introduce as a new normalization strategy for deep networks, aiming to further denoise the network's features. Wei Wang 0108, Zheng Dang, Yinlin Hu, Pascal Fua, Mathieu Salzmann |
NeurIPS | 4 |
| 2019 | Geometry in active learning for binary and multi-class image segmentation
Ksenia Konyushkova, Raphael Sznitman, Pascal Fua |
Comput. Vis. Image Underst. | 3 |
| 2019 | Beyond Sharing Weights for Deep Domain AdaptationabstractThe performance of a classifier trained on data coming from a specific domain typically degrades when applied to a related but different one. While annotating many samples from the new domain would address this issue, it is often too expensive or impractical. Domain Adaptation has therefore emerged as a solution to this problem; It leverages annotated data from a source domain, in which it is abundant, to train a classifier to operate in a target domain, in which it is either sparse or even lacking altogether. In this context, the recent trend consists of learning deep architectures whose weights are shared for both domains, which essentially amounts to learning domain invariant features. Here, we show that it is more effective to explicitly model the shift from one domain to the other. To this end, we introduce a two-stream architecture, where one operates in the source domain and the other in the target domain. In contrast to other approaches, the weights in corresponding layers are related but not shared. We demonstrate that this both yields higher accuracy than state-of-the-art methods on several object recognition and detection tasks and consistently outperforms networks with shared weights in both supervised and unsupervised settings. Artem Rozantsev, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | A Performance Evaluation of Local Features for Image-Based 3D ReconstructionabstractThis paper performs a comprehensive and comparative evaluation of the state-of-the-art local features for the task of image-based 3D reconstruction. The evaluated local features cover the recently developed ones by using powerful machine learning techniques and the elaborately designed handcrafted features. To obtain a comprehensive evaluation, we choose to include both float type features and binary ones. Meanwhile, two kinds of datasets have been used in this evaluation. One is a dataset of many different scene types with groundtruth 3D points, containing images of different scenes captured at fixed positions, for quantitative performance evaluation of different local features in the controlled image capturing situation. The other dataset contains Internet scale image sets of several landmarks with a lot of unrelated images, which is used for qualitative performance evaluation of different local features in the free image collection situation. Our experimental results show that binary features are competent to reconstruct scenes from controlled image sequences with only a fraction of processing time compared to using float type features. However, for the case of a large scale image set with many distracting images, float type features show a clear advantage over binary ones. Currently, the most traditional SIFT is very stable with regard to scene types in this specific task and produces very competitive reconstruction results among all the evaluated local features. Meanwhile, although the learned binary features are not as competitive as the handcrafted ones, learning float type features with CNN is promising but still requires much effort in the future. Bin Fan 0001, Qingqun Kong, Xinchao Wang, Zhiheng Wang 0001, Shiming Xiang, Chunhong Pan, Pascal Fua |
IEEE Trans. Image Process. | 7 |
| 2019 | Mo2Cap2: Real-time Mobile 3D Motion Capture with a Cap-mounted Fisheye CameraabstractWe propose the first real-time system for the egocentric estimation of 3D human body pose in a wide range of unconstrained everyday activities. This setting has a unique set of challenges, such as mobility of the hardware setup, and robustness to long capture sessions with fast recovery from tracking failures. We tackle these challenges based on a novel lightweight setup that converts a standard baseball cap to a device for high-quality pose estimation based on a single cap-mounted fisheye camera. From the captured egocentric live stream, our CNN based 3D pose estimation approach runs at 60 Hz on a consumer-level GPU. In addition to the lightweight hardware setup, our other main contributions are: 1) a large ground truth training corpus of top-down fisheye images and 2) a disentangled 3D pose estimation approach that takes the unique properties of the egocentric viewpoint into account. As shown by our evaluation, we achieve lower 3D joint error as well as better 2D overlay than the existing baselines. Weipeng Xu, Avishek Chatterjee, Michael Zollhöfer, Helge Rhodin, Pascal Fua, Hans-Peter Seidel, Christian Theobalt |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2018 | Learning to Reconstruct Texture-Less Deformable Surfaces from a Single ViewabstractRecent years have seen the development of mature solutions for reconstructing deformable surfaces from a single image, provided that they are relatively well-textured. By contrast, recovering the 3D shape of texture-less surfaces remains an open problem, and essentially relates to Shape-from-Shading. In this paper, we introduce a data-driven approach to this problem. We introduce a general framework that can predict diverse 3D representations, such as meshes, normals, and depth maps. Our experiments show that meshes are ill-suited to handle texture-less 3D reconstruction in our context. Furthermore, we demonstrate that our approach generalizes well to unseen objects, and that it yields higher-quality reconstructions than a state-of-the-art SfS technique, particularly in terms of normal estimates. Our reconstructions accurately model the fine details of the surfaces, such as the creases of a T-Shirt worn by a person. Jan Bednarík, Pascal Fua, Mathieu Salzmann |
3DV | 2 |
| 2018 | Every Smile Is Unique: Landmark-Guided Diverse Smile GenerationabstractEach smile is unique: one person surely smiles in different ways (e.g. closing/opening the eyes or mouth). Given one input image of a neutral face, can we generate multiple smile videos with distinctive characteristics? To tackle this one-to-many video generation problem, we propose a novel deep learning architecture named Conditional Multi-Mode Network (CMM-Net). To better encode the dynamics of facial expressions, CMM-Net explicitly exploits facial landmarks for generating smile sequences. Specifically, a variational auto-encoder is used to learn a facial landmark embedding. This single embedding is then exploited by a conditional recurrent network which generates a landmark embedding sequence conditioned on a specific expression (e.g. spontaneous smile). Next, the generated landmark embeddings are fed into a multi-mode recurrent landmark generator, producing a set of landmark sequences still associated to the given smile class but clearly distinct from each other. Finally, these landmark sequences are translated into face videos. Our experimental results demonstrate the effectiveness of our CMM-Net in generating realistic videos of multiple smile expressions. Wei Wang 0108, Xavier Alameda-Pineda, Dan Xu 0002, Pascal Fua, Elisa Ricci 0001, Nicu Sebe |
CVPR | 4 |
| 2018 | Modeling Facial Geometry Using Compositional VAEsabstractWe propose a method for learning non-linear face geometry representations using deep generative models. Our model is a variational autoencoder with multiple levels of hidden variables where lower layers capture global geometry and higher ones encode more local deformations. Based on that, we propose a new parameterization of facial geometry that naturally decomposes the structure of the human face into a set of semantically meaningful levels of detail. This parameterization enables us to do model fitting while capturing varying level of detail under different types of geometrical constraints. Timur M. Bagautdinov, Chenglei Wu, Jason M. Saragih, Pascal Fua, Yaser Sheikh |
CVPR | 4 |
| 2018 | WILDTRACK: A Multi-Camera HD Dataset for Dense Unscripted Pedestrian DetectionabstractPeople detection methods are highly sensitive to occlusions between pedestrians, which are extremely frequent in many situations where cameras have to be mounted at a limited height. The reduction of camera prices allows for the generalization of static multi-camera set-ups. Using joint visual information from multiple synchronized cameras gives the opportunity to improve detection performance. In this paper, we present a new large-scale and high-resolution dataset. It has been captured with seven static cameras in a public open area, and unscripted dense groups of pedestrians standing and walking. Together with the camera frames, we provide an accurate joint (extrinsic and intrinsic) calibration, as well as 7 series of 400 annotated frames for detection at a rate of 2 frames per second. This results in over 40 000 bounding boxes delimiting every person present in the area of interest, for a total of more than 300 individuals. We provide a series of benchmark results using baseline algorithms published over the recent months for multi-view detection with deep neural networks, and trajectory estimation using a non-Markovian model. Tatjana Chavdarova, Pierre Baqué, Stéphane Bouquet, Andrii Maksai, Cijo Jose, Timur M. Bagautdinov, Louis Lettry, Pascal Fua, Luc Van Gool, François Fleuret |
CVPR | 8 |
| 2018 | Beyond the Pixel-Wise Loss for Topology-Aware DelineationabstractDelineation of curvilinear structures is an important problem in Computer Vision with multiple practical applications. With the advent of Deep Learning, many current approaches on automatic delineation have focused on finding more powerful deep architectures, but have continued using the habitual pixel-wise losses such as binary cross-entropy. In this paper we claim that pixel-wise losses alone are unsuitable for this problem because of their inability to reflect the topological impact of mistakes in the final prediction. We propose a new loss term that is aware of the higher-order topological features of linear structures. We also exploit a refinement pipeline that iteratively applies the same model over the previous delineation to refine the predictions at each step, while keeping the number of parameters and the complexity of the model constant. When combined with the standard pixel-wise loss, both our new loss term and an iterative refinement boost the quality of the predicted delineations, in some cases almost doubling the accuracy as compared to the same classifier trained with the binary cross-entropy alone. We show that our approach outperforms state-of-the-art methods on a wide range of data, from microscopy to aerial images. Agata Mosinska, Pablo Márquez-Neila, Mateusz Kozinski, Pascal Fua |
CVPR | 4 |
| 2018 | Learning Monocular 3D Human Pose Estimation From Multi-View ImagesabstractAccurate 3D human pose estimation from single images is possible with sophisticated deep-net architectures that have been trained on very large datasets. However, this still leaves open the problem of capturing motions for which no such database exists. Manual annotation is tedious, slow, and error-prone. In this paper, we propose to replace most of the annotations by the use of multiple views, at training time only. Specifically, we train the system to predict the same pose in all views. Such a consistency constraint is necessary but not sufficient to predict accurate poses. We therefore complement it with a supervised loss aiming to predict the correct pose in a small set of labeled images, and with a regularization term that penalizes drift from initial predictions. Furthermore, we propose a method to estimate camera pose jointly with human pose, which lets us utilize multiview footage where calibration is difficult, e.g., for pan-tilt or moving handheld cameras. We demonstrate the effectiveness of our approach on established benchmarks, as well as on a new Ski dataset with rotating cameras and expert ski motion, for which annotations are truly hard to obtain. Helge Rhodin, Jörg Spörri, Isinsu Katircioglu, Victor Constantin, Frédéric Meyer, Erich Müller, Mathieu Salzmann, Pascal Fua |
CVPR | 8 |
| 2018 | Residual Parameter Transfer for Deep Domain AdaptationabstractThe goal of Deep Domain Adaptation is to make it possible to use Deep Nets trained in one domain where there is enough annotated training data in another where there is little or none. Most current approaches have focused on learning feature representations that are invariant to the changes that occur when going from one domain to the other, which means using the same network parameters in both domains. While some recent algorithms explicitly model the changes by adapting the network parameters, they either severely restrict the possible domain changes, or significantly increase the number of model parameters. By contrast, we introduce a network architecture that includes auxiliary residual networks, which we train to predict the parameters in the domain with little annotated data from those in the other one. This architecture enables us to flexibly preserve the similarities between domains where they exist and model the differences when necessary. We demonstrate that our approach yields higher accuracy than state-of-the-art methods without undue complexity. Artem Rozantsev, Mathieu Salzmann, Pascal Fua |
CVPR | 3 |
| 2018 | Real-Time Seamless Single Shot 6D Object Pose PredictionabstractWe propose a single-shot approach for simultaneously detecting an object in an RGB image and predicting its 6D pose without requiring multiple stages or having to examine multiple hypotheses. Unlike a recently proposed single-shot technique for this task [10] that only predicts an approximate 6D pose that must then be refined, ours is accurate enough not to require additional post-processing. As a result, it is much faster - 50 fps on a Titan X (Pascal) GPU - and more suitable for real-time processing. The key component of our method is a new CNN architecture inspired by [27, 28] that directly predicts the 2D image locations of the projected vertices of the object's 3D bounding box. The object's 6D pose is then estimated using a PnP algorithm. For single object and multiple object pose estimation on the LINEMOD and OCCLUSION datasets, our approach substantially outperforms other recent CNN-based approaches [10, 25] when they are all used without postprocessing. During post-processing, a pose refinement step can be used to boost the accuracy of these two methods, but at 10 fps or less, they are much slower than our method. Bugra Tekin, Sudipta N. Sinha, Pascal Fua |
CVPR | 3 |
| 2018 | Learning to Find Good CorrespondencesabstractWe develop a deep architecture to learn to find good correspondences for wide-baseline stereo. Given a set of putative sparse matches and the camera intrinsics, we train our network in an end-to-end fashion to label the correspondences as inliers or outliers, while simultaneously using them to recover the relative pose, as encoded by the essential matrix. Our architecture is based on a multi-layer perceptron operating on pixel coordinates rather than directly on the image, and is thus simple and small. We introduce a novel normalization technique, called Context Normalization, which allows us to process each data point separately while embedding global information in it, and also makes the network invariant to the order of the correspondences. Our experiments on multiple challenging datasets demonstrate that our method is able to drastically improve the state of the art with little training data. Kwang Moo Yi, Eduard Trulls, Yuki Ono, Vincent Lepetit, Mathieu Salzmann, Pascal Fua |
CVPR | 6 |
| 2018 | Eigendecomposition-Free Training of Deep Networks with Zero Eigenvalue-Based Losses
Zheng Dang, Kwang Moo Yi, Yinlin Hu, Fei Wang 0008, Pascal Fua, Mathieu Salzmann |
ECCV (5) | 5 |
| 2018 | Unsupervised Geometry-Aware Representation for 3D Human Pose EstimationabstractModern 3D human pose estimation techniques rely on deep networks, which require large amounts of training data. While weakly-supervised methods require less supervision, by utilizing 2D poses or multi-view imagery without annotations, they still need a sufficiently large set of samples with 3D annotations for learning to succeed. In this paper, we propose to overcome this problem by learning a geometry-aware body representation from multi-view images without annotations. To this end, we use an encoder-decoder that predicts an image from one viewpoint given an image from another viewpoint. Because this representation encodes 3D geometry, using it in a semi-supervised setting makes it easier to learn a mapping from it to 3D human pose. As evidenced by our experiments, our approach significantly outperforms fully-supervised methods given the same amount of labeled data, and improves over other semi-supervised methods while using as little as 1% of the labeled data. Helge Rhodin, Mathieu Salzmann, Pascal Fua |
ECCV (10) | 3 |
| 2018 | FishEyeRecNet: A Multi-context Collaborative Deep Network for Fisheye Image Rectification
Xiaoqing Yin, Xinchao Wang, Jun Yu 0002, Maojun Zhang, Pascal Fua, Dacheng Tao |
ECCV (10) | 5 |
| 2018 | Geodesic Convolutional Shape OptimizationabstractAerodynamic shape optimization has many industrial applications. Existing methods, however, are so computationally demanding that typical engineering practices are to either simply try a limited number of hand-designed shapes or restrict oneself to shapes that can be parameterized using only few degrees of freedom. In this work, we introduce a new way to optimize complex shapes fast and accurately. To this end, we train Geodesic Convolutional Neural Networks to emulate a fluidynamics simulator. The key to making this approach practical is remeshing the original shape using a poly-cube map, which makes it possible to perform the computations on GPUs instead of CPUs. The neural net is then used to formulate an objective function that is differentiable with respect to the shape parameters, which can then be optimized using a gradient-based technique. This outperforms state-of-the-art methods by 5 to 20% for standard problems and, even more importantly, our approach applies to cases that previous methods cannot handle. Pierre Baqué, Edoardo Remelli, François Fleuret, Pascal Fua |
ICML | 4 |
| 2018 | Learning to Segment 3D Linear Structures Using Only 2D Annotations
Mateusz Kozinski, Agata Mosinska, Mathieu Salzmann, Pascal Fua |
MICCAI (2) | 4 |
| 2018 | LF-Net: Learning Local Features from ImagesabstractWe present a novel deep architecture and a training strategy to learn a local feature pipeline from scratch, using collections of images without the need for human supervision. To do so we exploit depth and relative camera pose cues to create a virtual target that the network should achieve on one image, provided the outputs of the network for the other image. While this process is inherently non-differentiable, we show that we can optimize the network in a two-branch setup by confining it to one branch, while preserving differentiability in the other. We train our method on both indoor and outdoor datasets, with depth data from 3D sensors for the former, and depth estimates from an off-the-shelf Structure-from-Motion solution for the latter. Our models outperform the state of the art on sparse feature matching on both datasets, while running at 60+ fps for QVGA images. Yuki Ono, Eduard Trulls, Pascal Fua, Kwang Moo Yi |
NeurIPS | 3 |
| 2018 | Learning Latent Representations of 3D Human Pose with Deep Neural Networks
Isinsu Katircioglu, Bugra Tekin, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
Int. J. Comput. Vis. | 5 |
| 2018 | Robust 3D Object Tracking from Monocular Images Using Stable PartsabstractWe present an algorithm for estimating the pose of a rigid object in real-time under challenging conditions. Our method effectively handles poorly textured objects in cluttered, changing environments, even when their appearance is corrupted by large occlusions, and it relies on grayscale images to handle metallic environments on which depth cameras would fail. As a result, our method is suitable for practical Augmented Reality applications including industrial environments. At the core of our approach is a novel representation for the 3D pose of object parts: We predict the 3D pose of each part in the form of the 2D projections of a few control points. The advantages of this representation is three-fold: We can predict the 3D pose of the object even when only one part is visible; when several parts are visible, we can easily combine them to compute a better pose of the object; the 3D pose we obtain is usually very accurate, even when only few parts are visible. We show how to use this representation in a robust 3D tracking framework. In addition to extensive comparisons with the state-of-the-art, we demonstrate our method on a practical Augmented Reality application for maintenance assistance in the ATLAS particle detector at CERN. Alberto Crivellaro, Mahdi Rad, Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Reconstructing Evolving Tree Structures in Time Lapse Sequences by Enforcing Time-ConsistencyabstractWe propose a novel approach to reconstructing curvilinear tree structures evolving over time, such as road networks in 2D aerial images or neural structures in 3D microscopy stacks acquired in vivo. To enforce temporal consistency, we simultaneously process all images in a sequence, as opposed to reconstructing structures of interest in each image independently. We formulate the problem as a Quadratic Mixed Integer Program and demonstrate the additional robustness that comes from using all available visual clues at once, instead of working frame by frame. Furthermore, when the linear structures undergo local changes over time, our approach automatically detects them. Przemyslaw Glowacki, Miguel Amável Pinheiro, Agata Mosinska, Engin Türetken, Daniel Lebrecht, Raphael Sznitman, Anthony Holtmaat, Jan Kybic, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2017 | Monocular 3D Human Pose Estimation in the Wild Using Improved CNN SupervisionabstractWe propose a CNN-based approach for 3D human body pose estimation from single RGB images that addresses the issue of limited generalizability of models trained solely on the starkly limited publicly available 3D pose data. Using only the existing 3D pose data and 2D pose data, we show state-of-the-art performance on established benchmarks through transfer of learned features, while also generalizing to in-the-wild scenes. We further introduce a new training set for human body pose estimation from monocular images of real humans that has the ground truth captured with a multi-camera marker-less motion capture system. It complements existing corpora with greater diversity in pose, human appearance, clothing, occlusion, and viewpoints, and enables an increased scope of augmentation. We also contribute a new benchmark that covers outdoor and indoor scenes, and demonstrate that our 3D pose dataset shows better in-the-wild performance than existing annotated data, which is further improved in conjunction with transfer learning from 2D pose data. All in all, we argue that the use of transfer learning of representations in tandem with algorithmic and data contributions is crucial for general 3D body pose estimation. Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, Christian Theobalt |
3DV | 4 |
| 2017 | Social Scene Understanding: End-to-End Multi-person Action Localization and Collective Activity RecognitionabstractWe present a unified framework for understanding human social behaviors in raw image sequences. Our model jointly detects multiple individuals, infers their social actions, and estimates the collective actions with a single feed-forward pass through a neural network. We propose a single architecture that does not rely on external detection algorithms but rather is trained end-to-end to generate dense proposal maps that are refined via a novel inference scheme. The temporal consistency is handled via a person-level matching Recurrent Neural Network. The complete model takes as input a sequence of frames and outputs detections along with the estimates of individual actions and collective activities. We demonstrate state-of-the-art performance of our algorithm on multiple publicly available benchmarks. Timur M. Bagautdinov, Alexandre Alahi, François Fleuret, Pascal Fua, Silvio Savarese |
CVPR | 4 |
| 2017 | Multi-modal Mean-Fields via Cardinality-Based Clamping
Pierre Baqué, François Fleuret, Pascal Fua |
CVPR | 3 |
| 2017 | Flight Dynamics-Based Recovery of a UAV Trajectory Using Ground CamerasabstractWe propose a new method to estimate the 6-dof trajectory of a flying object such as a quadrotor UAV within a 3D airspace monitored using multiple fixed ground cameras. It is based on a new structure from motion formulation for the 3D reconstruction of a single moving point with known motion dynamics. Our main contribution is a new bundle adjustment procedure, which in addition to optimizing the camera poses, regularizes the point trajectory using a prior based on motion dynamics (or specifically flight dynamics). Furthermore, we can infer the underlying control input sent to the UAVs autopilot that determined its flight trajectory. Our method requires neither perfect single-view tracking nor appearance matching across views. For robustness, we allow the tracker to generate multiple detections per frame in each video. The true detections and the data association across videos is estimated using robust multi-view triangulation and subsequently refined in our bundle adjustment formulation. Quantitative evaluation on simulated data and experiments on real videos from indoor and outdoor scenes shows that our technique is superior to existing methods. Artem Rozantsev, Sudipta N. Sinha, Debadeepta Dey, Pascal Fua |
CVPR | 4 |
| 2017 | Deep Occlusion Reasoning for Multi-camera Multi-target DetectionabstractPeople detection in single 2D images has improved greatly in recent years. However, comparatively little of this progress has percolated into multi-camera multi-people tracking algorithms, whose performance still degrades severely when scenes become very crowded. In this work, we introduce a new architecture that combines Convolutional Neural Nets and Conditional Random Fields to explicitly model those ambiguities. One of its key ingredients are high-order CRF terms that model potential occlusions and give our approach its robustness even when many people are present. Our model is trained end-to-end and we show that it outperforms several state-of-the-art algorithms on challenging scenes. Pierre Baqué, François Fleuret, Pascal Fua |
ICCV | 3 |
| 2017 | Non-Markovian Globally Consistent Multi-object TrackingabstractMany state-of-the-art approaches to multi-object tracking rely on detecting them in each frame independently, grouping detections into short but reliable trajectory segments, and then further grouping them into full trajectories. This grouping typically relies on imposing local smoothness constraints but almost never on enforcing more global ones on the trajectories. In this paper, we propose a non-Markovian approach to imposing global consistency by using behavioral patterns to guide the tracking algorithm. When used in conjunction with state-of-the-art tracking algorithms, this further increases their already good performance on multiple challenging datasets. We show significant improvements both in supervised settings where ground truth is available and behavioral patterns can be learned from it, and in completely unsupervised settings. Andrii Maksai, Xinchao Wang, François Fleuret, Pascal Fua |
ICCV | 4 |
| 2017 | Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose EstimationabstractMost recent approaches to monocular 3D human pose estimation rely on Deep Learning. They typically involve regressing from an image to either 3D joint coordinates directly or 2D joint locations from which 3D coordinates are inferred. Both approaches have their strengths and weaknesses and we therefore propose a novel architecture designed to deliver the best of both worlds by performing both simultaneously and fusing the information along the way. At the heart of our framework is a trainable fusion scheme that learns how to fuse the information optimally instead of being hand-designed. This yields significant improvements upon the state-of-the-art on standard 3D human pose estimation benchmarks. Bugra Tekin, Pablo Márquez-Neila, Mathieu Salzmann, Pascal Fua |
ICCV | 4 |
| 2017 | Learning Lightprobes for Mixed Reality IlluminationabstractThis paper presents the first photometric registration pipeline for Mixed Reality based on high quality illumination estimation using convolutional neural networks (CNNs). For easy adaptation and deployment of the system, we train the CNNs using purely synthetic images and apply them to real image data. To keep the pipeline accurate and efficient, we propose to fuse the light estimation results from multiple CNN instances and show an approach for caching estimates over time. For optimal performance, we furthermore explore multiple strategies for the CNN training. Experimental results show that the proposed method yields highly accurate estimates for photo-realistic augmentations. David Mandl, Kwang Moo Yi, Peter Mohr, Peter M. Roth, Pascal Fua, Vincent Lepetit, Dieter Schmalstieg, Denis Kalkofen |
ISMAR | 5 |
| 2017 | Simultaneous Recognition and Pose Estimation of Instruments in Minimally Invasive Surgery
Thomas Kurmann, Pablo Márquez-Neila, Xiaofei Du 0001, Pascal Fua, Danail Stoyanov, Sebastian Wolf 0005, Raphael Sznitman |
MICCAI (2) | 4 |
| 2017 | Active Learning and Proofreading for Delineation of Curvilinear Structures
Agata Mosinska, Jakub Tarnawski, Pascal Fua |
MICCAI (2) | 3 |
| 2017 | Learning Active Learning from DataabstractIn this paper, we suggest a novel data-driven approach to active learning (AL). The key idea is to train a regressor that predicts the expected error reduction for a candidate sample in a particular learning state. By formulating the query selection procedure as a regression problem we are not restricted to working with existing AL heuristics; instead, we learn strategies based on experience from previous AL outcomes. We show that a strategy can be learnt either from simple synthetic 2D datasets or from a subset of domain-specific data. Our method yields strategies that work well on real data from a wide range of domains. Ksenia Konyushkova, Raphael Sznitman, Pascal Fua |
NIPS | 3 |
| 2017 | Geometric Graph Matching Using Monte Carlo Tree SearchabstractWe present an efficient matching method for generalized geometric graphs. Such graphs consist of vertices in space connected by curves and can represent many real world structures such as road networks in remote sensing, or vessel networks in medical imaging. Graph matching can be used for very fast and possibly multimodal registration of images of these structures. We formulate the matching problem as a single player game solved using Monte Carlo Tree Search, which automatically balances exploring new possible matches and extending existing matches. Our method can handle partial matches, topological differences, geometrical distortion, does not use appearance information and does not require an initial alignment. Moreover, our method is very efficient-it can match graphs with thousands of nodes, which is an order of magnitude better than the best competing method, and the matching only takes a few seconds. Miguel Amável Pinheiro, Jan Kybic, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Detecting Flying Objects Using a Single Moving CameraabstractWe propose an approach for detecting flying objects such as Unmanned Aerial Vehicles (UAVs) and aircrafts when they occupy a small portion of the field of view, possibly moving against complex backgrounds, and are filmed by a camera that itself moves. We argue that solving such a difficult problem requires combining both appearance and motion cues. To this end we propose a regression-based approach for object-centric motion stabilization of image patches that allows us to achieve effective classification on spatio-temporal image cubes and outperform state-of-the-art techniques. As this problem has not yet been extensively studied, no test datasets are publicly available. We therefore built our own, both for UAVs and aircrafts, and will make them publicly available so they can be used to benchmark future flying object detection and collision avoidance algorithms. Artem Rozantsev, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Network Flow Integer Programming to Track Elliptical Cells in Time-Lapse SequencesabstractWe propose a novel approach to automatically tracking elliptical cell populations in time-lapse image sequences. Given an initial segmentation, we account for partial occlusions and overlaps by generating an over-complete set of competing detection hypotheses. To this end, we fit ellipses to portions of the initial regions and build a hierarchy of ellipses, which are then treated as cell candidates. We then select temporally consistent ones by solving to optimality an integer program with only one type of flow variables. This eliminates the need for heuristics to handle missed detections due to partial occlusions and complex morphology. We demonstrate the effectiveness of our approach on a range of challenging sequences consisting of clumped cells and show that it outperforms state-of-the-art techniques. Engin Türetken, Xinchao Wang, Carlos J. Becker, Carsten Haubold, Pascal Fua |
IEEE Trans. Medical Imaging | 5 |
| 2016 | Structured Prediction of 3D Human Pose with Deep Neural Networks
Bugra Tekin, Isinsu Katircioglu, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
BMVC | 5 |
| 2016 | Learning to Match Aerial Images with Deep Attentive ArchitecturesabstractImage matching is a fundamental problem in Computer Vision. In the context of feature-based matching, SIFT and its variants have long excelled in a wide array of applications. However, for ultra-wide baselines, as in the case of aerial images captured under large camera rotations, the appearance variation goes beyond the reach of SIFT and RANSAC. In this paper we propose a data-driven, deep learning-based approach that sidesteps local correspondence by framing the problem as a classification task. Furthermore, we demonstrate that local correspondences can still be useful. To do so we incorporate an attention mechanism to produce a set of probable matches, which allows us to further increase performance. We train our models on a dataset of urban aerial imagery consisting of 'same' and 'different' pairs, collected for this purpose, and characterize the problem via a human study with annotations from Amazon Mechanical Turk. We demonstrate that our models outperform the state-of-the-art on ultra-wide baseline matching and approach human accuracy. Hani Altwaijry, Eduard Trulls, James Hays, Pascal Fua, Serge J. Belongie |
CVPR | 4 |
| 2016 | Principled Parallel Mean-Field Inference for Discrete Random FieldsabstractMean-field variational inference is one of the most popular approaches to inference in discrete random fields. Standard mean-field optimization is based on coordinate descent and in many situations can be impractical. Thus, in practice, various parallel techniques are used, which either rely on ad hoc smoothing with heuristically set parameters, or put strong constraints on the type of models. In this paper, we propose a novel proximal gradient-based approach to optimizing the variational objective. It is naturally parallelizable and easy to implement. We prove its convergence, and demonstrate that, in practice, it yields faster convergence and often finds better optima than more traditional mean-field optimization techniques. Moreover, our method is less sensitive to the choice of parameters. Pierre Baqué, Timur M. Bagautdinov, François Fleuret, Pascal Fua |
CVPR | 4 |
| 2016 | What Players do with the Ball: A Physically Constrained Interaction ModelingabstractTracking the ball is critical for video-based analysis of team sports. However, it is difficult, especially in low-resolution images, due to the small size of the ball, its speed that creates motion blur, and its often being occluded by players. In this paper, we propose a generic and principled approach to modeling the interaction between the ball and the players while also imposing appropriate physical constraints on the ball's trajectory. We show that our approach, formulated in terms of a Mixed Integer Program, is more robust and more accurate than several state-of-the-art approaches on real-life volleyball, basketball, and soccer sequences. Andrii Maksai, Xinchao Wang, Pascal Fua |
CVPR | 3 |
| 2016 | Active Learning for Delineation of Curvilinear StructuresabstractMany recent delineation techniques owe much of their increased effectiveness to path classification algorithms that make it possible to distinguish promising paths from others. The downside of this development is that they require annotated training data, which is tedious to produce. In this paper, we propose an Active Learning approach that considerably speeds up the annotation process. Unlike standard ones, it takes advantage of the specificities of the delineation problem. It operates on a graph and can reduce the training set size by up to 80% without compromising the reconstruction quality. We will show that our approach outperforms conventional ones on various biomedical and natural image datasets, thus showing that it is broadly applicable. Agata Mosinska, Raphael Sznitman, Przemyslaw Glowacki, Pascal Fua |
CVPR | 4 |
| 2016 | Direct Prediction of 3D Body Poses from Motion Compensated SequencesabstractWe propose an efficient approach to exploiting motion information from consecutive frames of a video sequence to recover the 3D pose of people. Previous approaches typically compute candidate poses in individual frames and then link them in a post-processing step to resolve ambiguities. By contrast, we directly regress from a spatio-temporal volume of bounding boxes to a 3D pose in the central frame. We further show that, for this approach to achieve its full potential, it is essential to compensate for the motion in consecutive frames so that the subject remains centered. This then allows us to effectively overcome ambiguities and improve upon the state-of-the-art by a large margin on the Human3.6m, HumanEva, and KTH Multiview Football 3D human pose estimation benchmarks. Bugra Tekin, Artem Rozantsev, Vincent Lepetit, Pascal Fua |
CVPR | 4 |
| 2016 | Learning to Assign Orientations to Feature PointsabstractWe show how to train a Convolutional Neural Network to assign a canonical orientation to feature points given an image patch centered on the feature point. Our method improves feature point matching upon the state-of-the art and can be used in conjunction with any existing rotation sensitive descriptors. To avoid the tedious and almost impossible task of finding a target orientation to learn, we propose to use Siamese networks which implicitly find the optimal orientations during training. We also propose a new type of activation function for Neural Networks that generalizes the popular ReLU, maxout, and PReLU activation functions. This novel activation performs better for our task. We validate the effectiveness of our method extensively with four existing datasets, including two non-planar datasets, as well as our own dataset. We show that we outperform the state-of-the-art without the need of retraining for each dataset. Kwang Moo Yi, Yannick Verdie, Pascal Fua, Vincent Lepetit |
CVPR | 3 |
| 2016 | LIFT: Learned Invariant Feature Transform
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, Pascal Fua |
ECCV (6) | 4 |
| 2016 | Vision-based Unmanned Aerial Vehicle detection and tracking for sense and avoid systemsabstractWe propose an approach for on-line detection of small Unmanned Aerial Vehicles (UAVs) and estimation of their relative positions and velocities in the 3D environment from a single moving camera in the context of sense and avoid systems. This problem is challenging both from a detection point of view, as there are no markers on the targets available, and from a tracking perspective, due to misdetection and false positives. Furthermore, the methods need to be computationally light, despite the complexity of computer vision algorithms, to be used on UAVs with limited payload. To address these issues we propose a multi-staged framework that incorporates fast object detection using an AdaBoost-based approach, coupled with an on-line visual-based tracking algorithm and a recent sensor fusion and state estimation method. Our framework allows for achieving real-time performance with accurate object detection and tracking without any need of markers and customized, high-performing hardware resources. Krishna Raj Sapkota, Steven Roelofsen, Artem Rozantsev, Vincent Lepetit, Denis Gillet, Pascal Fua, Alcherio Martinoli |
IROS | 6 |
| 2016 | Analyzing Volleyball Match Data from the 2014 World Championships Using Machine Learning TechniquesabstractThis paper proposes a relational-learning based approach for discovering strategies in volleyball matches based on optical tracking data. In contrast to most existing methods, our approach permits discovering patterns that account for both spatial (that is, partial configurations of the players on the court) and temporal (that is, the order of events and positions) aspects of the game. We analyze both the men's and women's final match from the 2014 FIVB Volleyball World Championships, and are able to identify several interesting and relevant strategies from the matches. Jan Van Haaren, Horesh Ben Shitrit, Jesse Davis, Pascal Fua |
KDD | 4 |
| 2016 | Scalable Unsupervised Domain Adaptation for Electron Microscopy
Róger Bermúdez-Chacón, Carlos J. Becker, Mathieu Salzmann, Pascal Fua |
MICCAI (2) | 4 |
| 2016 | Simultaneous segmentation and anatomical labeling of the cerebral vasculature
David Robben, Engin Türetken, Stefan Sunaert, Vincent Thijs, Guy Willms, Pascal Fua, Frederik Maes, Paul Suetens |
Medical Image Anal. | 6 |
| 2016 | Parsing human skeletons in an operating room
Vasileios Belagiannis, Xinchao Wang, Horesh Ben Shitrit, Kiyoshi Hashimoto, Ralf Stauder, Yoshimitsu Aoki, Michael Kranzfelder, Armin Schneider, Pascal Fua, Slobodan Ilic, Hubertus Feußner, Nassir Navab |
Mach. Vis. Appl. | 9 |
| 2016 | Template-Based Monocular 3D Shape Recovery Using Laplacian MeshesabstractWe show that by extending the Laplacian formalism, which was first introduced in the Graphics community to regularize 3D meshes, we can turn the monocular 3D shape reconstruction of a deformable surface given correspondences with a reference image into a much better-posed problem. This allows us to quickly and reliably eliminate outliers by simply solving a linear least squares problem. This yields an initial 3D shape estimate, which is not necessarily accurate, but whose 2D projections are. The initial shape is then refined by a constrained optimization problem to output the final surface reconstruction. Our approach allows us to reduce the dimensionality of the surface reconstruction problem without sacrificing accuracy, thus allowing for real-time implementations. Dat Tien Ngo, Jonas Östlund, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Multiscale Centerline DetectionabstractFinding the centerline and estimating the radius of linear structures is a critical first step in many applications, ranging from road delineation in 2D aerial images to modeling blood vessels, lung bronchi, and dendritic arbors in 3D biomedical image stacks. Existing techniques rely either on filters designed to respond to ideal cylindrical structures or on classification techniques. The former tend to become unreliable when the linear structures are very irregular while the latter often has difficulties distinguishing centerline locations from neighboring ones, thus losing accuracy. We solve this problem by reformulating centerline detection in terms of a regression problem. We first train regressors to return the distances to the closest centerline in scale-space, and we apply them to the input images or volumes. The centerlines and the corresponding scale then correspond to the regressors local maxima, which can be easily identified. We show that our method outperforms state-of-the-art techniques for various 2D and 3D datasets. Moreover, our approach is very generic and also performs well on contour detection. We show an improvement above recent contour detection algorithms on the BSDS500 dataset. Amos Sironi, Engin Türetken, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Reconstructing Curvilinear Networks Using Path Classifiers and Integer ProgrammingabstractWe propose a novel approach to automated delineation of curvilinear structures that form complex and potentially loopy networks. By representing the image data as a graph of potential paths, we first show how to weight these paths using discriminatively-trained classifiers that are both robust and generic enough to be applied to very different imaging modalities. We then present an Integer Programming approach to finding the optimal subset of paths, subject to structural and topological constraints that eliminate implausible solutions. Unlike earlier approaches that assume a tree topology for the networks, ours explicitly models the fact that the networks may contain loops, and can reconstruct both cyclic and acyclic ones. We demonstrate the effectiveness of our approach on a variety of challenging datasets including aerial images of road networks and micrographs of neural arbors, and show that it outperforms state-of-the-art techniques. Engin Türetken, Fethallah Benmansour, Bjoern Andres, Przemyslaw Glowacki, Hanspeter Pfister, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2016 | Tracking Interacting Objects Using Intertwined FlowsabstractIn this paper, we show that tracking different kinds of interacting objects can be formulated as a network-flow mixed integer program. This is made possible by tracking all objects simultaneously using intertwined flow variables and expressing the fact that one object can appear or disappear at locations where another is in terms of linear flow constraints. Our proposed method is able to track invisible objects whose only evidence is the presence of other objects that contain them. Furthermore, our tracklet-based implementation yields real-time tracking performance. We demonstrate the power of our approach on scenes involving cars and pedestrians, bags being carried and dropped by people, and balls being passed from one player to the next in team sports. In particular, we show that by estimating jointly and globally the trajectories of different types of objects, the presence of the ones which were not initially detected based solely on image evidence can be inferred from the detections of the others. Xinchao Wang, Engin Türetken, François Fleuret, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Probability occupancy maps for occluded depth imagesabstractWe propose a novel approach to computing the probabilities of presence of multiple and potentially occluding objects in a scene from a single depth map. To this end, we use a generative model that predicts the distribution of depth images that would be produced if the probabilities of presence were known and then to optimize them so that this distribution explains observed evidence as closely as possible. This allows us to exploit very effectively the available evidence and outperform state-of-the-art methods without requiring large amounts of data, or without using the RGB signal that modern RGB-D sensors also provide. Timur M. Bagautdinov, François Fleuret, Pascal Fua |
CVPR | 3 |
| 2015 | Flying objects detection from a single moving cameraabstractWe propose an approach to detect flying objects such as UAVs and aircrafts when they occupy a small portion of the field of view, possibly moving against complex backgrounds, and are filmed by a camera that itself moves. Artem Rozantsev, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2015 | TILDE: A Temporally Invariant Learned DEtectorabstractWe introduce a learning-based approach to detect repeatable keypoints under drastic imaging changes of weather and lighting conditions to which state-of-the-art keypoint detectors are surprisingly sensitive. We first identify good keypoint candidates in multiple training images taken from the same viewpoint. We then train a regressor to predict a score map whose maxima are those points so that they can be found by simple non-maximum suppression. As there are no standard datasets to test the influence of these kinds of changes, we created our own, which we will make publicly available. We will show that our method significantly outperforms the state-of-the-art methods in such challenging conditions, while still achieving state-of-the-art performance on untrained standard datasets. Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
CVPR | 3 |
| 2015 | A Novel Representation of Parts for Accurate 3D Object Detection and Tracking in Monocular ImagesabstractWe present a method that estimates in real-time and under challenging conditions the 3D pose of a known object. Our method relies only on grayscale images since depth cameras fail on metallic objects, it can handle poorly textured objects, and cluttered, changing environments, the pose it predicts degrades gracefully in presence of large occlusions. As a result, by contrast with the state-of-the-art, our method is suitable for practical Augmented Reality applications even in industrial environments. To be robust to occlusions, we first learn to detect some parts of the target object. Our key idea is to then predict the 3D pose of each part in the form of the 2D projections of a few control points. The advantages of this representation is three-fold: We can predict the 3D pose of the object even when only one part is visible, when several parts are visible, we can combine them easily to compute a better pose of the object, the 3D pose we obtain is usually very accurate, even when only few parts are visible. Alberto Crivellaro, Mahdi Rad, Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
ICCV | 5 |
| 2015 | Hot or Not: Exploring Correlations between Appearance and TemperatureabstractIn this paper we explore interactions between the appearance of an outdoor scene and the ambient temperature. By studying statistical correlations between image sequences from outdoor cameras and temperature measurements we identify two interesting interactions. First, semantically meaningful regions such as foliage and reflective oriented surfaces are often highly indicative of the temperature. Second, small camera motions are correlated with the temperature in some scenes. We propose simple scene-specific temperature prediction algorithms which can be used to turn a camera into a crude temperature sensor. We find that for this task, simple features such as local pixel intensities outperform sophisticated, global features such as from a semantically-trained convolutional neural network. Daniel Glasner, Pascal Fua, Todd E. Zickler, Lihi Zelnik-Manor |
ICCV | 2 |
| 2015 | Introducing Geometry in Active Learning for Image SegmentationabstractWe propose an Active Learning approach to training a segmentation classifier that exploits geometric priors to streamline the annotation process in 3D image volumes. To this end, we use these priors not only to select voxels most in need of annotation but to guarantee that they lie on 2D planar patch, which makes it much easier to annotate than if they were randomly distributed in the volume. A simplified version of this approach is effective in natural 2D images. We evaluated our approach on Electron Microscopy and Magnetic Resonance image volumes, as well as on natural images. Comparing our approach against several accepted baselines demonstrates a marked performance increase. Ksenia Konyushkova, Raphael Sznitman, Pascal Fua |
ICCV | 3 |
| 2015 | Dense Image Registration and Deformable Surface Reconstruction in Presence of Occlusions and Minimal TextureabstractDeformable surface tracking from monocular images is well-known to be under-constrained. Occlusions often make the task even more challenging, and can result in failure if the surface is not sufficiently textured. In this work, we explicitly address the problem of 3D reconstruction of poorly textured, occluded surfaces, proposing a framework based on a template-matching approach that scales dense robust features by a relevancy score. Our approach is extensively compared to current methods employing both local feature matching and dense template alignment. We test on standard datasets as well as on a new dataset (that will be made publicly available) of a sparsely textured, occluded surface. Our framework achieves state-of-the-art results for both well and poorly textured, occluded surfaces. Dat Tien Ngo, Sanghyuk Park, Anne Jorstad, Alberto Crivellaro, Chang Dong Yoo, Pascal Fua |
ICCV | 6 |
| 2015 | Discriminative Learning of Deep Convolutional Feature Point DescriptorsabstractDeep learning has revolutionalized image-level tasks such as classification, but patch-level tasks, such as correspondence, still rely on hand-crafted features, e.g. SIFT. In this paper we use Convolutional Neural Networks (CNNs) to learn discriminant patch representations and in particular train a Siamese network with pairs of (non-)corresponding patches. We deal with the large number of potential pairs with the combination of a stochastic sampling of the training set and an aggressive mining strategy biased towards patches that are hard to classify. By using the L2 distance during both training and testing we develop 128-D descriptors whose euclidean distances reflect patch similarity, and which can be used as a drop-in replacement for any task involving SIFT. We demonstrate consistent performance gains over the state of the art, and generalize well against scaling and rotation, perspective transformation, non-rigid deformation, and illumination changes. Our descriptors are efficient to compute and amenable to modern GPUs, and are publicly available. Edgar Simo-Serra, Eduard Trulls, Luis Ferraz, Iasonas Kokkinos, Pascal Fua, Francesc Moreno-Noguer |
ICCV | 5 |
| 2015 | Projection onto the Manifold of Elongated Structures for Accurate ExtractionabstractDetection of elongated structures in 2D images and 3D image stacks is a critical prerequisite in many applications and Machine Learning-based approaches have recently been shown to deliver superior performance. However, these methods essentially classify individual locations and do not explicitly model the strong relationship that exists between neighboring ones. As a result, isolated erroneous responses, discontinuities, and topological errors are present in the resulting score maps. We solve this problem by projecting patches of the score map to their nearest neighbors in a set of ground truth training patches. Our algorithm induces global spatial consistency on the classifier score map and returns results that are provably geometrically consistent. We apply our algorithm to challenging datasets in four different domains and show that it compares favorably to state-of-the-art methods. Amos Sironi, Vincent Lepetit, Pascal Fua |
ICCV | 3 |
| 2015 | Kullback-Leibler Proximal Variational InferenceabstractWe propose a new variational inference method based on the Kullback-Leibler (KL) proximal term. We make two contributions towards improving efficiency of variational inference. Firstly, we derive a KL proximal-point algorithm and show its equivalence to gradient descent with natural gradient in stochastic variational inference. Secondly, we use the proximal framework to derive efficient variational algorithms for non-conjugate models. We propose a splitting procedure to separate non-conjugate terms from conjugate ones. We then linearize the non-conjugate terms and show that the resulting subproblem admits a closed-form solution. Overall, our approach converts a non-conjugate model to subproblems that involve inference in well-known conjugate models. We apply our method to many models and derive generalizations for non-conjugate exponential family. Applications to real-world datasets show that our proposed algorithms are easy to implement, fast to converge, perform well, and reduce computations. Mohammad E. Khan, Pierre Baqué, François Fleuret, Pascal Fua |
NIPS | 4 |
| 2015 | On rendering synthetic images for training an object detector
Artem Rozantsev, Vincent Lepetit, Pascal Fua |
Comput. Vis. Image Underst. | 3 |
| 2015 | Editorial
Björn Stenger, Norimichi Ukita, Yoichi Sato 0001, Pascal Fua, David J. Fleet |
Comput. Vis. Image Underst. | 4 |
| 2015 | Non-Rigid Graph Registration Using Active Testing SearchabstractWe present a new approach for matching sets of branching curvilinear structures that form graphs embedded in R2 or R3 and may be subject to deformations. Unlike earlier methods, ours does not rely on local appearance similarity nor does require a good initial alignment. Furthermore, it can cope with non-linear deformations, topological differences, and partial graphs. To handle arbitrary non-linear deformations, we use Gaussian process regressions to represent the geometrical mapping relating the two graphs. In the absence of appearance information, we iteratively establish correspondences between points, update the mapping accordingly, and use it to estimate where to find the most likely correspondences that will be used in the next step. To make the computation tractable for large graphs, the set of new potential matches considered at each iteration is not selected at random as with many RANSAC-based algorithms. Instead, we introduce a so-called Active Testing Search strategy that performs a priority search to favor the most likely matches and speed-up the process. We demonstrate the effectiveness of our approach first on synthetic cases and then on angiography data, retinal fundus images, and microscopy image stacks acquired at very different resolutions. Eduard Serradell, Miguel Amável Pinheiro, Raphael Sznitman, Jan Kybic, Francesc Moreno-Noguer, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2015 | Learning Separable FiltersabstractLearning filters to produce sparse image representations in terms of over-complete dictionaries has emerged as a powerful way to create image features for many different purposes. Unfortunately, these filters are usually both numerous and non-separable, making their use computationally expensive. In this paper, we show that such filters can be computed as linear combinations of a smaller number of separable ones, thus greatly reducing the computational complexity at no cost in terms of performance. This makes filter learning approaches practical even for large images or 3D volumes, and we show that we significantly outperform state-of-the-art methods on the curvilinear structure extraction task, in terms of both accuracy and speed. Moreover, our approach is general and can be used on generic convolutional filter banks to reduce the complexity of the feature extraction step. Amos Sironi, Bugra Tekin, Roberto Rigamonti, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2015 | Domain Adaptation for Microscopy ImagingabstractElectron and light microscopy imaging can now deliver high-quality image stacks of neural structures. However, the amount of human annotation effort required to analyze them remains a major bottleneck. While machine learning algorithms can be used to help automate this process, they require training data, which is time-consuming to obtain manually, especially in image stacks. Furthermore, due to changing experimental conditions, successive stacks often exhibit differences that are severe enough to make it difficult to use a classifier trained for a specific one on another. This means that this tedious annotation process has to be repeated for each new stack. In this paper, we present a domain adaptation algorithm that addresses this issue by effectively leveraging labeled examples across different acquisitions and significantly reducing the annotation requirements. Our approach can handle complex, nonlinear image feature transformations and scales to large microscopy datasets that often involve high-dimensional feature spaces and large 3D data volumes. We evaluate our approach on four challenging electron and light microscopy applications that exhibit very different image modalities and where annotation is very costly. Across all applications we achieve a significant improvement over the state-of-the-art machine learning methods and demonstrate our ability to greatly reduce human annotation effort. Carlos J. Becker, C. Mario Christoudias, Pascal Fua |
IEEE Trans. Medical Imaging | 3 |
| 2015 | Learning Structured Models for Segmentation of 2-D and 3-D ImageryabstractEfficient and accurate segmentation of cellular structures in microscopic data is an essential task in medical imaging. Many state-of-the-art approaches to image segmentation use structured models whose parameters must be carefully chosen for optimal performance. A popular choice is to learn them using a large-margin framework and more specifically structured support vector machines (SSVM). Although SSVMs are appealing, they suffer from certain limitations. First, they are restricted in practice to linear kernels because the more powerful nonlinear kernels cause the learning to become prohibitively expensive. Second, they require iteratively finding the most violated constraints, which is often intractable for the loopy graphical models used in image segmentation. This requires approximation that can lead to reduced quality of learning. In this paper, we propose three novel techniques to overcome these limitations. We first introduce a method to "kernelize" the features so that a linear SSVM framework can leverage the power of nonlinear kernels without incurring much additional computational cost. Moreover, we employ a working set of constraints to increase the reliability of approximate subgradient methods and introduce a new way to select a suitable step size at each iteration. We demonstrate the strength of our approach on both 2-D and 3-D electron microscopic (EM) image data and show consistent performance improvement over state-of-the-art approaches. Aurélien Lucchi, Pablo Márquez-Neila, Carlos J. Becker, Yunpeng Li 0002, Kevin Smith 0001, Graham Knott, Pascal Fua |
IEEE Trans. Medical Imaging | 7 |
| 2015 | Live Texturing of Augmented Reality Characters from Colored DrawingsabstractColoring books capture the imagination of children and provide them with one of their earliest opportunities for creative expression. However, given the proliferation and popularity of digital devices, real-world activities like coloring can seem unexciting, and children become less engaged in them. Augmented reality holds unique potential to impact this situation by providing a bridge between real-world activities and digital enhancements. In this paper, we present an augmented reality coloring book App in which children color characters in a printed coloring book and inspect their work using a mobile device. The drawing is detected and tracked, and the video stream is augmented with an animated 3-D version of the character that is textured according to the child's coloring. This is possible thanks to several novel technical contributions. We present a texturing process that applies the captured texture from a 2-D colored drawing to both the visible and occluded regions of a 3-D character in real time. We develop a deformable surface tracking method designed for colored drawings that uses a new outlier rejection algorithm for real-time tracking and surface deformation recovery. We present a content creation pipeline to efficiently create the 2-D and 3-D content. And, finally, we validate our work with two user studies that examine the quality of our texturing algorithm and the overall App experience. Stéphane Magnenat, Dat Tien Ngo, Fabio Zünd, Mattia Ryffel, Gioacchino Noris, Gerhard Röthlin, Alessia Marra, Maurizio Nitti, Pascal Fua, Markus Gross 0001, Robert W. Sumner |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2014 | Reconstructing Evolving Tree Structures in Time Lapse SequencesabstractWe propose an approach to reconstructing tree structures that evolve over time in 2D images and 3D image stacks such as neuronal axons or plant branches. Instead of reconstructing structures in each image independently, we do so for all images simultaneously to take advantage of temporal-consistency constraints. We show that this problem can be formulated as a Quadratic Mixed Integer Program and solved efficiently. The outcome of our approach is a framework that provides substantial improvements in reconstructions over traditional single time-instance formulations. Furthermore, an added benefit of our approach is the ability to automatically detect places where significant changes have occurred over time, which is challenging when considering large amounts of data. Przemyslaw Glowacki, Miguel Amável Pinheiro, Engin Türetken, Raphael Sznitman, Daniel Lebrecht, Jan Kybic, Anthony Holtmaat, Pascal Fua |
CVPR | 8 |
| 2014 | Multiscale Centerline Detection by Learning a Scale-Space Distance TransformabstractWe propose a robust and accurate method to extract the centerlines and scale of tubular structures in 2D images and 3D volumes. Existing techniques rely either on filters designed to respond to ideal cylindrical structures, which lose accuracy when the linear structures become very irregular, or on classification, which is inaccurate because locations on centerlines and locations immediately next to them are extremely difficult to distinguish. We solve this problem by reformulating centerline detection in terms of a regression problem. We first train regressors to return the distances to the closest centerline in scale-space, and we apply them to the input images or volumes. The centerlines and the corresponding scale then correspond to the regressors local maxima, which can be easily identified. We show that our method outperforms state-of-the-art techniques for various 2D and 3D datasets. Amos Sironi, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2014 | Free-Shape Polygonal Object Localization
Xiaolu Sun, C. Mario Christoudias, Pascal Fua |
ECCV (6) | 3 |
| 2014 | Tracking Interacting Objects Optimally Using Integer Programming
Xinchao Wang, Engin Türetken, François Fleuret, Pascal Fua |
ECCV (1) | 4 |
| 2014 | Tracking texture-less, shiny objects with descriptor fieldsabstractOur demo demonstrates the method we published at CVPR this year for tracking specular and poorly textured objects, and lets the visitors experiment with it and with their own patterns. Our approach only requires a standard monocular camera (no need for a depth sensor), and can be easily integrated within existing systems to improve their robustness and accuracy. Alberto Crivellaro, Yannick Verdie, Kwang Moo Yi, Pascal Fua, Vincent Lepetit |
ISMAR | 4 |
| 2014 | Exploiting Enclosing Membranes and Contextual Cues for Mitochondria Segmentation
Aurélien Lucchi, Carlos J. Becker, Pablo Márquez-Neila, Pascal Fua |
MICCAI (1) | 4 |
| 2014 | Simultaneous Segmentation and Anatomical Labeling of the Cerebral Vasculature
David Robben, Engin Türetken, Stefan Sunaert, Vincent Thijs, Guy Willms, Pascal Fua, Frederik Maes, Paul Suetens |
MICCAI (1) | 6 |
| 2014 | Fast Part-Based Classification for Instrument Detection in Minimally Invasive Surgery
Raphael Sznitman, Carlos J. Becker, Pascal Fua |
MICCAI (2) | 3 |
| 2014 | On the relevance of sparsity for image classification
Roberto Rigamonti, Vincent Lepetit, Germán González, Engin Türetken, Fethallah Benmansour, Matthew A. Brown, Pascal Fua |
Comput. Vis. Image Underst. | 7 |
| 2014 | Take your eyes off the ball: Improving ball-tracking by focusing on team play
Xinchao Wang, Vitaly Ablavsky, Horesh Ben Shitrit, Pascal Fua |
Comput. Vis. Image Underst. | 4 |
| 2014 | Real-time landing place assessment in man-made environments
Xiaolu Sun, C. Mario Christoudias, Vincent Lepetit, Pascal Fua |
Mach. Vis. Appl. | 4 |
| 2014 | Multi-Commodity Network Flow for Tracking Multiple PeopleabstractIn this paper, we show that tracking multiple people whose paths may intersect can be formulated as a multi-commodity network flow problem. Our proposed framework is designed to exploit image appearance cues to prevent identity switches. Our method is effective even when such cues are only available at distant time intervals. This is unlike many current approaches that depend on appearance being exploitable from frame-to-frame. Furthermore, our algorithm lends itself to a real-time implementation. We validate our approach on three publicly available datasets that contain long and complex sequences, the APIDIS basketball dataset, the ISSIA soccer dataset, and the PETS'09 pedestrian dataset. We also demonstrate its performance on a newer basketball dataset that features complete world championship basketball matches. In all cases, our approach preserves identity better than state-of-the-art tracking algorithms. Horesh Ben Shitrit, Jérôme Berclaz, François Fleuret, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | Receptive Fields Selection for Binary Feature DescriptionabstractFeature description for local image patch is widely used in computer vision. While the conventional way to design local descriptor is based on expert experience and knowledge, learning-based methods for designing local descriptor become more and more popular because of their good performance and data-driven property. This paper proposes a novel data-driven method for designing binary feature descriptor, which we call receptive fields descriptor (RFD). Technically, RFD is constructed by thresholding responses of a set of receptive fields, which are selected from a large number of candidates according to their distinctiveness and correlations in a greedy way. Using two different kinds of receptive fields (namely rectangular pooling area and Gaussian pooling area) for selection, we obtain two binary descriptors RFDR and RFDG .accordingly. Image matching experiments on the well-known patch data set and Oxford data set demonstrate that RFD significantly outperforms the state-of-the-art binary descriptors, and is comparable with the best float-valued descriptors at a fraction of processing time. Finally, experiments on object recognition tasks confirm that both RFDR and RFDG successfully bridge the performance gap between binary descriptors and their floating-point competitors. Bin Fan 0001, Qingqun Kong, Tomasz Trzcinski, Zhiheng Wang 0001, Chunhong Pan, Pascal Fua |
IEEE Trans. Image Process. | 6 |
| 2013 | Learning for Structured Prediction Using Approximate Subgradient Descent with Working SetsabstractWe propose a working set based approximate sub gradient descent algorithm to minimize the margin-sensitive hinge loss arising from the soft constraints in max-margin learning frameworks, such as the structured SVM. We focus on the setting of general graphical models, such as loopy MRFs and CRFs commonly used in image segmentation, where exact inference is intractable and the most violated constraints can only be approximated, voiding the optimality guarantees of the structured SVM's cutting plane algorithm as well as reducing the robustness of existing sub gradient based methods. We show that the proposed method obtains better approximate sub gradients through the use of working sets, leading to improved convergence properties and increased reliability. Furthermore, our method allows new constraints to be randomly sampled instead of computed using the more expensive approximate inference techniques such as belief propagation and graph cuts, which can be used to reduce learning time at only a small cost of performance. We demonstrate the strength of our method empirically on the segmentation of a new publicly available electron microscopy dataset as well as the popular MSRC data set and show state-of-the-art results. Aurélien Lucchi, Yunpeng Li 0002, Pascal Fua |
CVPR | 3 |
| 2013 | Learning Separable FiltersabstractLearning filters to produce sparse image representations in terms of over complete dictionaries has emerged as a powerful way to create image features for many different purposes. Unfortunately, these filters are usually both numerous and non-separable, making their use computationally expensive. In this paper, we show that such filters can be computed as linear combinations of a smaller number of separable ones, thus greatly reducing the computational complexity at no cost in terms of performance. This makes filter learning approaches practical even for large images or 3D volumes, and we show that we significantly outperform state-of-the-art methods on the linear structure extraction task, in terms of both accuracy and speed. Moreover, our approach is general and can be used on generic filter banks to reduce the complexity of the convolutions. Roberto Rigamonti, Amos Sironi, Vincent Lepetit, Pascal Fua |
CVPR | 4 |
| 2013 | Fast Object Detection with Entropy-Driven EvaluationabstractCascade-style approaches to implementing ensemble classifiers can deliver significant speed-ups at test time. While highly effective, they remain challenging to tune and their overall performance depends on the availability of large validation sets to estimate rejection thresholds. These characteristics are often prohibitive and thus limit their applicability. We introduce an alternative approach to speeding-up classifier evaluation which overcomes these limitations. It involves maintaining a probability estimate of the class label at each intermediary response and stopping when the corresponding uncertainty becomes small enough. As a result, the evaluation terminates early based on the sequence of responses observed. Furthermore, it does so independently of the type of ensemble classifier used or the way it was trained. We show through extensive experimentation that our method provides 2 to 10 fold speed-ups, over existing state-of-the-art methods, at almost no loss in accuracy on a number of object classification tasks. Raphael Sznitman, Carlos J. Becker, François Fleuret, Pascal Fua |
CVPR | 4 |
| 2013 | Boosting Binary Keypoint DescriptorsabstractBinary key point descriptors provide an efficient alternative to their floating-point competitors as they enable faster processing while requiring less memory. In this paper, we propose a novel framework to learn an extremely compact binary descriptor we call Bin Boost that is very robust to illumination and viewpoint changes. Each bit of our descriptor is computed with a boosted binary hash function, and we show how to efficiently optimize the different hash functions so that they complement each other, which is key to compactness and robustness. The hash functions rely on weak learners that are applied directly to the image patches, which frees us from any intermediate representation and lets us automatically learn the image gradient pooling configuration of the final descriptor. Our resulting descriptor significantly outperforms the state-of-the-art binary descriptors and performs similarly to the best floating-point descriptors at a fraction of the matching time and memory footprint. Tomasz Trzcinski, C. Mario Christoudias, Pascal Fua, Vincent Lepetit |
CVPR | 3 |
| 2013 | Reconstructing Loopy Curvilinear Structures Using Integer ProgrammingabstractWe propose a novel approach to automated delineation of linear structures that form complex and potentially loopy networks. This is in contrast to earlier approaches that usually assume a tree topology for the networks. At the heart of our method is an Integer Programming formulation that allows us to find the global optimum of an objective function designed to allow cycles but penalize spurious junctions and early terminations. We demonstrate that it outperforms state-of-the-art techniques on a wide range of datasets. Engin Türetken, Fethallah Benmansour, Bjoern Andres, Hanspeter Pfister, Pascal Fua |
CVPR | 5 |
| 2013 | Detecting Irregular Curvilinear Structures in Gray Scale and Color Imagery Using Multi-directional Oriented FluxabstractWe propose a new approach to detecting irregular curvilinear structures in noisy image stacks. In contrast to earlier approaches that rely on circular models of the cross-sections, ours allows for the arbitrarily-shaped ones that are prevalent in biological imagery. This is achieved by maximizing the image gradient flux along multiple directions and radii, instead of only two with a unique radius as is usually done. This yields a more complex optimization problem for which we propose a computationally efficient solution. We demonstrate the effectiveness of our approach on a wide range of challenging gray scale and color datasets and show that it outperforms existing techniques, especially on very irregular structures. Engin Türetken, Carlos J. Becker, Przemyslaw Glowacki, Fethallah Benmansour, Pascal Fua |
ICCV | 5 |
| 2013 | An Optimal Policy for Target Localization with Application to Electron MicroscopyabstractThis paper considers the task of finding a target location by making a limited number of sequential observations. Each observation results from evaluating an imperfect classifier of a chosen cost and accuracy on an interval of chosen length and position. Within a Bayesian framework, we study the problem of minimizing an objective that combines the entropy of the posterior distribution with the cost of the questions asked. In this problem, we show that the one-step lookahead policy is Bayes-optimal for any arbitrary time horizon. Moreover, this one-step lookahead policy is easy to compute and implement. We then use this policy in the context of localizing mitochondria in electron microscope images, and experimentally show that significant speed ups in acquisition can be gained, while maintaining near equal image quality at target locations, when compared to current policies. Raphael Sznitman, Aurélien Lucchi, Peter I. Frazier, Bruno Jedynak, Pascal Fua |
ICML (1) | 5 |
| 2013 | Supervised Feature Learning for Curvilinear Structure Segmentation
Carlos J. Becker, Roberto Rigamonti, Vincent Lepetit, Pascal Fua |
MICCAI (1) | 4 |
| 2013 | Flash Scanning Electron Microscopy
Raphael Sznitman, Aurélien Lucchi, Marco Cantoni, Graham Knott, Pascal Fua |
MICCAI (3) | 5 |
| 2013 | Non-Linear Domain Adaptation with BoostingabstractA common assumption in machine vision is that the training and test samples are drawn from the same distribution. However, there are many problems when this assumption is grossly violated, as in bio-medical applications where different acquisitions can generate drastic variations in the appearance of the data due to changing experimental conditions. This problem is accentuated with 3D data, for which annotation is very time-consuming, limiting the amount of data that can be labeled in new acquisitions for training. In this paper we present a multi-task learning algorithm for domain adaptation based on boosting. Unlike previous approaches that learn task-specific decision boundaries, our method learns a single decision boundary in a shared feature space, common to all tasks. We use the boosting-trick to learn a non-linear mapping of the observations in each task, with no need for specific a-priori knowledge of its global analytical form. This yields a more parameter-free domain adaptation approach that successfully leverages learning on new tasks where labeled data is scarce. We evaluate our approach on two challenging bio-medical datasets and achieve a significant improvement over the state-of-the-art. Carlos J. Becker, C. Mario Christoudias, Pascal Fua |
NIPS | 3 |
| 2013 | TPAMI CVPR Special SectionabstractThe articles in this special issue include papers from the CVPR'11 conference which was held in Colorado Spring, CO, June 2011. Pedro F. Felzenszwalb, David A. Forsyth, Pascal Fua, Terrance E. Boult |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | Stochastic Exploration of Ambiguities for Nonrigid Shape RecoveryabstractRecovering the 3D shape of deformable surfaces from single images is known to be a highly ambiguous problem because many different shapes may have very similar projections. This is commonly addressed by restricting the set of possible shapes to linear combinations of deformation modes and by imposing additional geometric constraints. Unfortunately, because image measurements are noisy, such constraints do not always guarantee that the correct shape will be recovered. To overcome this limitation, we introduce a stochastic sampling approach to efficiently explore the set of solutions of an objective function based on point correspondences. This allows us to propose a small set of ambiguous candidate 3D shapes and then use additional image information to choose the best one. As a proof of concept, we use either motion or shading cues to this end and show that we can handle a complex objective function without having to solve a difficult nonlinear minimization problem. The advantages of our method are demonstrated on a variety of problems including both real and synthetic data. Francesc Moreno-Noguer, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Learning Context Cues for Synapse SegmentationabstractWe present a new approach for the automated segmentation of synapses in image stacks acquired by electron microscopy (EM) that relies on image features specifically designed to take spatial context into account. These features are used to train a classifier that can effectively learn cues such as the presence of a nearby post-synaptic region. As a result, our algorithm successfully distinguishes synapses from the numerous other organelles that appear within an EM volume, including those whose local textural properties are relatively similar. Furthermore, as a by-product of the segmentation, our method flawlessly determines synaptic orientation, a crucial element in the interpretation of brain circuits. We evaluate our approach on three different datasets, compare it against the state-of-the-art in synapse segmentation and demonstrate our ability to reliably collect shape, density, and orientation statistics over hundreds of synapses. Carlos J. Becker, Karim Ali 0002, Graham Knott, Pascal Fua |
IEEE Trans. Medical Imaging | 4 |
| 2012 | Robust non-rigid registration of 2D and 3D graphsabstractWe present a new approach to matching graphs embedded in ℝ2or ℝ3. Unlike earlier methods, our approach does not rely on the similarity of local appearance features, does not require an initial alignment, can handle partial matches, and can cope with non-linear deformations and topological differences. To handle arbitrary non-linear deformations, we represent them as Gaussian Processes. In the absence of appearance information, we iteratively establish correspondences between graph nodes, update the structure accordingly, and use the current mapping estimate to find the most likely correspondences that will be used in the next iteration. This makes the computation tractable. We demonstrate the effectiveness of our approach first on synthetic cases and then on angiography data, retinal fundus images, and microscopy image stacks acquired at very different resolutions. Eduard Serradell, Przemyslaw Glowacki, Jan Kybic, Francesc Moreno-Noguer, Pascal Fua |
CVPR | 5 |
| 2012 | Automated reconstruction of tree structures using path classifiers and Mixed Integer ProgrammingabstractAlthough tracing linear structures in 2D images and 3D image stacks has received much attention over the years, full automation remains elusive. In this paper, we formulate the delineation problem as one of solving a Quadratic Mixed Integer Program (Q-MIP) in a graph of potential paths, which can be done optimally up to a very small tolerance. We further propose a novel approach to weighting these paths, which results in a Q-MIP solution that accurately matches the ground truth. We demonstrate that our approach outperforms a state-of-the-art technique based on the k-Minimum Spanning Tree formulation on a 2D dataset of aerial images and a 3D dataset of confocal microscopy stacks. Engin Türetken, Fethallah Benmansour, Pascal Fua |
CVPR | 3 |
| 2012 | A constrained latent variable modelabstractLatent variable models provide valuable compact representations for learning and inference in many computer vision tasks. However, most existing models cannot directly encode prior knowledge about the specific problem at hand. In this paper, we introduce a constrained latent variable model whose generated output inherently accounts for such knowledge. To this end, we propose an approach that explicitly imposes equality and inequality constraints on the model's output during learning, thus avoiding the computational burden of having to account for these constraints at inference. Our learning mechanism can exploit non-linear kernels, while only involving sequential closed-form updates of the model parameters. We demonstrate the effectiveness of our constrained latent variable model on the problem of non-rigid 3D reconstruction from monocular images, and show that it yields qualitative and quantitative improvements over several baselines. Aydin Varol, Mathieu Salzmann, Pascal Fua, Raquel Urtasun |
CVPR | 3 |
| 2012 | Worldwide Pose Estimation Using 3D Point Clouds
Yunpeng Li 0002, Noah Snavely, Daniel P. Huttenlocher, Pascal Fua |
ECCV (1) | 4 |
| 2012 | Structured Image Segmentation Using Kernelized Features
Aurélien Lucchi, Yunpeng Li 0002, Kevin Smith 0001, Pascal Fua |
ECCV (2) | 4 |
| 2012 | Laplacian Meshes for Monocular 3D Shape Recovery
Jonas Östlund, Aydin Varol, Dat Tien Ngo, Pascal Fua |
ECCV (3) | 4 |
| 2012 | Learning Context Cues for Synapse Segmentation in EM Volumes
Carlos J. Becker, Karim Ali 0002, Graham Knott, Pascal Fua |
MICCAI (1) | 4 |
| 2012 | Data-Driven Visual Tracking in Retinal Microsurgery
Raphael Sznitman, Karim Ali 0002, Rogério Richa, Russell H. Taylor, Gregory D. Hager, Pascal Fua |
MICCAI (2) | 6 |
| 2012 | Efficient Scanning for EM Based Target Localization
Raphael Sznitman, Aurélien Lucchi, Natasa Pjescic-Emedji, Graham Knott, Pascal Fua |
MICCAI (3) | 5 |
| 2012 | Learning Image Descriptors with the Boosting-TrickabstractIn this paper we apply boosting to learn complex non-linear local visual feature representations, drawing inspiration from its successful application to visual object detection. The main goal of local feature descriptors is to distinctively represent a salient image region while remaining invariant to viewpoint and illumination changes. This representation can be improved using machine learning, however, past approaches have been mostly limited to learning linear feature mappings in either the original input or a kernelized input feature space. While kernelized methods have proven somewhat effective for learning non-linear local feature descriptors, they rely heavily on the choice of an appropriate kernel function whose selection is often difficult and non-intuitive. We propose to use the boosting-trick to obtain a non-linear mapping of the input to a high-dimensional feature space. The non-linear feature mapping obtained with the boosting-trick is highly intuitive. We employ gradient-based weak learners resulting in a learned descriptor that closely resembles the well-known SIFT. As demonstrated in our experiments, the resulting descriptor can be learned directly from intensity patches achieving state-of-the-art performance. Tomasz Trzcinski, C. Mario Christoudias, Vincent Lepetit, Pascal Fua |
NIPS | 4 |
| 2012 | Efficient large-scale multi-view stereo for ultra high-resolution image sets
Engin Tola, Christoph Strecha, Pascal Fua |
Mach. Vis. Appl. | 3 |
| 2012 | SLIC Superpixels Compared to State-of-the-Art Superpixel MethodsabstractComputer vision applications have come to rely increasingly on superpixels in recent years, but it is not always clear what constitutes a good superpixel algorithm. In an effort to understand the benefits and drawbacks of existing methods, we empirically compare five state-of-the-art superpixel algorithms for their ability to adhere to image boundaries, speed, memory efficiency, and their impact on segmentation performance. We then introduce a new superpixel algorithm, simple linear iterative clustering (SLIC), which adapts a k-means clustering approach to efficiently generate superpixels. Despite its simplicity, SLIC adheres to boundaries as well as or better than previous methods. At the same time, it is faster and more memory efficient, improves segmentation performance, and is straightforward to extend to supervoxel generation. Radhakrishna Achanta, Appu Shaji, Kevin Smith 0001, Aurélien Lucchi, Pascal Fua, Sabine Süsstrunk |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2012 | A Real-Time Deformable DetectorabstractWe propose a new learning strategy for object detection. The proposed scheme forgoes the need to train a collection of detectors dedicated to homogeneous families of poses, and instead learns a single classifier that has the inherent ability to deform based on the signal of interest. We train a detector with a standard AdaBoost procedure by using combinations of pose-indexed features and pose estimators. This allows the learning process to select and combine various estimates of the pose with features able to compensate for variations in pose without the need to label data for training or explore the pose space in testing. We validate our framework on three types of data: hand video sequences, aerial images of cars, and face images. We compare our method to a standard boosting framework, with access to the same ground truth, and show a reduction in the false alarm rate of up to an order of magnitude. Where possible, we compare our method to the state of the art, which requires pose annotations of the training data, and demonstrate comparable performance. Karim Ali 0002, François Fleuret, David Hasler, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | BRIEF: Computing a Local Binary Descriptor Very FastabstractBinary descriptors are becoming increasingly popular as a means to compare feature points very fast while requiring comparatively small amounts of memory. The typical approach to creating them is to first compute floating-point ones, using an algorithm such as SIFT, and then to binarize them. In this paper, we show that we can directly compute a binary descriptor, which we call BRIEF, on the basis of simple intensity difference tests. As a result, BRIEF is very fast both to build and to match. We compare it against SURF and SIFT on standard benchmarks and show that it yields comparable recognition accuracy, while running in an almost vanishing fraction of the time required by either. Michael Calonder, Vincent Lepetit, Mustafa Özuysal, Tomasz Trzcinski, Christoph Strecha, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2012 | Gradient Response Maps for Real-Time Detection of Textureless ObjectsabstractWe present a method for real-time 3D object instance detection that does not require a time-consuming training stage, and can handle untextured objects. At its core, our approach is a novel image representation for template matching designed to be robust to small image transformations. This robustness is based on spread image gradient orientations and allows us to test only a small subset of all possible pixel locations when parsing the image, and to represent a 3D object with a limited set of templates. In addition, we demonstrate that if a dense depth sensor is available we can extend our approach for an even better performance also taking 3D surface normal orientations into account. We show how to take advantage of the architecture of modern computers to build an efficient but very discriminant representation of the input images that can be used to consider thousands of templates in real time. We demonstrate in many experiments on real data that our method is much faster and more robust with respect to background clutter than current state-of-the-art methods. Stefan Hinterstoißer, Cedric Cagniart, Slobodan Ilic, Peter F. Sturm, Nassir Navab, Pascal Fua, Vincent Lepetit |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2012 | LDAHash: Improved Matching with Smaller DescriptorsabstractSIFT-like local feature descriptors are ubiquitously employed in computer vision applications such as content-based retrieval, video analysis, copy detection, object recognition, photo tourism, and 3D reconstruction. Feature descriptors can be designed to be invariant to certain classes of photometric and geometric transformations, in particular, affine and intensity scale transformations. However, real transformations that an image can undergo can only be approximately modeled in this way, and thus most descriptors are only approximately invariant in practice. Second, descriptors are usually high dimensional (e.g., SIFT is represented as a 128-dimensional vector). In large-scale retrieval and matching problems, this can pose challenges in storing and retrieving descriptor data. We map the descriptor vectors into the Hamming space in which the Hamming metric is used to compare the resulting representations. This way, we reduce the size of the descriptors by representing them as short binary strings and learn descriptor invariance from examples. We show extensive experimental validation, demonstrating the advantage of the proposed approach. Christoph Strecha, Alexander M. Bronstein, Michael M. Bronstein, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Monocular 3D Reconstruction of Locally Textured SurfacesabstractMost recent approaches to monocular nonrigid 3D shape recovery rely on exploiting point correspondences and work best when the whole surface is well textured. The alternative is to rely on either contours or shading information, which has only been demonstrated in very restrictive settings. Here, we propose a novel approach to monocular deformable shape recovery that can operate under complex lighting and handle partially textured surfaces. At the heart of our algorithm are a learned mapping from intensity patterns to the shape of local surface patches and a principled approach to piecing together the resulting local shape estimates. We validate our approach quantitatively and qualitatively using both synthetic and real data. Aydin Varol, Appu Shaji, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Thick boundaries in binary space and their influence on nearest-neighbor search
Tomasz Trzcinski, Vincent Lepetit, Pascal Fua |
Pattern Recognit. Lett. | 3 |
| 2012 | Supervoxel-Based Segmentation of Mitochondria in EM Image Stacks With Learned Shape FeaturesabstractIt is becoming increasingly clear that mitochondria play an important role in neural function. Recent studies show mitochondrial morphology to be crucial to cellular physiology and synaptic function and a link between mitochondrial defects and neuro-degenerative diseases is strongly suspected. Electron microscopy (EM), with its very high resolution in all three directions, is one of the key tools to look more closely into these issues but the huge amounts of data it produces make automated analysis necessary. State-of-the-art computer vision algorithms designed to operate on natural 2-D images tend to perform poorly when applied to EM data for a number of reasons. First, the sheer size of a typical EM volume renders most modern segmentation schemes intractable. Furthermore, most approaches ignore important shape cues, relying only on local statistics that easily become confused when confronted with noise and textures inherent in the data. Finally, the conventional assumption that strong image gradients always correspond to object boundaries is violated by the clutter of distracting membranes. In this work, we propose an automated graph partitioning scheme that addresses these issues. It reduces the computational complexity by operating on supervoxels instead of voxels, incorporates shape features capable of describing the 3-D shape of the target objects, and learns to recognize the distinctive appearance of true boundaries. Our experiments demonstrate that our approach is able to segment mitochondria at a performance level close to that of a human annotator, and outperforms a state-of-the-art 3-D segmentation technique. Aurélien Lucchi, Kevin Smith 0001, Radhakrishna Achanta, Graham Knott, Pascal Fua |
IEEE Trans. Medical Imaging | 5 |
| 2011 | Are spatial and global constraints really necessary for segmentation?abstractMany state-of-the-art segmentation algorithms rely on Markov or Conditional Random Field models designed to enforce spatial and global consistency constraints. This is often accomplished by introducing additional latent variables to the model, which can greatly increase its complexity. As a result, estimating the model parameters or computing the best maximum a posteriori (MAP) assignment becomes a computationally expensive task. In a series of experiments on the PASCAL and the MSRC datasets, we were unable to find evidence of a significant performance increase attributed to the introduction of such constraints. On the contrary, we found that similar levels of performance can be achieved using a much simpler design that essentially ignores these constraints. This more simple approach makes use of the same local and global features to leverage evidence from the image, but instead directly biases the preferences of individual pixels. While our investigation does not prove that spatial and consistency constraints are not useful in principle, it points to the conclusion that they should be validated in a larger context. Aurélien Lucchi, Yunpeng Li 0002, Xavier Boix, Kevin Smith 0001, Pascal Fua |
ICCV | 5 |
| 2011 | Conditional Random Fields for multi-camera object detectionabstractWe formulate a model for multi-class object detection in a multi-camera environment. From our knowledge, this is the first time that this problem is addressed taken into account different object classes simultaneously. Given several images of the scene taken from different angles, our system estimates the ground plane location of the objects from the output of several object detectors applied at each viewpoint. We cast the problem as an energy minimization modeled with a Conditional Random Field (CRF). Instead of predicting the presence of an object at each image location independently, we simultaneously predict the labeling of the entire scene. Our CRF is able to take into account occlusions between objects and contextual constraints among them. We propose an effective iterative strategy that renders tractable the underlying optimization problem, and learn the parameters of the model with the max-margin paradigm. We evaluate the performance of our model on several challenging multi-camera pedestrian detection datasets namely PETS 2009 [5] and EPFL terrace sequence [9]. We also introduce a new dataset in which multiple classes of objects appear simultaneously in the scene. It is here where we show that our method effectively handles occlusions in the multi-class case. Gemma Roig, Xavier Boix, Horesh Ben Shitrit, Pascal Fua |
ICCV | 4 |
| 2011 | Tracking multiple people under global appearance constraintsabstractIn this paper, we show that tracking multiple people whose paths may intersect can be formulated as a convex global optimization problem. Our proposed framework is designed to exploit image appearance cues to prevent identity switches. Our method is effective even when such cues are only available at distant time intervals. This is unlike many current approaches that depend on appearance being exploitable from frame to frame. We validate our approach on three multi-camera sport and pedestrian datasets that contain long and complex sequences. Our algorithm perseveres identities better than state-of-the-art algorithms while keeping similar MOTA scores. Horesh Ben Shitrit, Jérôme Berclaz, François Fleuret, Pascal Fua |
ICCV | 4 |
| 2011 | Long Term Real Trajectory Reuse through Region Goal Satisfaction
Junghyun Ahn, Stéphane Gobron, Quentin Silvestre, Horesh Ben Shitrit, Mirko Raca, Julien Pettré, Daniel Thalmann, Pascal Fua, Ronan Boulic |
MIG | 8 |
| 2011 | Learning Real-Time Perspective Patch Rectification
Stefan Hinterstoißer, Vincent Lepetit, Selim Benhimane, Pascal Fua, Nassir Navab |
Int. J. Comput. Vis. | 4 |
| 2011 | Real-time vehicle tracking for driving assistance
Andrea Fossati, Patrick Schönmann, Pascal Fua |
Mach. Vis. Appl. | 3 |
| 2011 | Multiple Object Tracking Using K-Shortest Paths OptimizationabstractMulti-object tracking can be achieved by detecting objects in individual frames and then linking detections across frames. Such an approach can be made very robust to the occasional detection failure: If an object is not detected in a frame but is in previous and following ones, a correct trajectory will nevertheless be produced. By contrast, a false-positive detection in a few frames will be ignored. However, when dealing with a multiple target problem, the linking step results in a difficult optimization problem in the space of all possible families of trajectories. This is usually dealt with by sampling or greedy search based on variants of Dynamic Programming which can easily miss the global optimum. In this paper, we show that reformulating that step as a constrained flow optimization results in a convex problem. We take advantage of its particular structure to solve it using the k-shortest paths algorithm, which is very fast. This new approach is far simpler formally and algorithmically than existing techniques and lets us demonstrate excellent performance in two very different contexts. Jérôme Berclaz, François Fleuret, Engin Türetken, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Linear Local Models for Monocular Reconstruction of Deformable SurfacesabstractRecovering the 3D shape of a nonrigid surface from a single viewpoint is known to be both ambiguous and challenging. Resolving the ambiguities typically requires prior knowledge about the most likely deformations that the surface may undergo. It often takes the form of a global deformation model that can be learned from training data. While effective, this approach suffers from the fact that a new model must be learned for each new surface, which means acquiring new training data, and may be impractical. In this paper, we replace the global models by linear local models for surface patches, which can be assembled to represent arbitrary surface shapes as long as they are made of the same material. Not only do they eliminate the need to retrain the model for different surface shapes, they also let us formulate 3D shape reconstruction from correspondences as either an algebraic problem that can be solved in closed form or a convex optimization problem whose solution can be found using standard numerical packages. We present quantitative results on synthetic data, as well as qualitative results on real images. Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Pareto-optimal dictionaries for signaturesabstractWe present an effective method to optimize over the parameters of an image patch descriptor to obtain one that is computationally more efficient while maintaining a high recognition rate. We formulate the optimization problem in a multi-objective manner, which balances two conflicting goals while removing the need for traditional weighting coefficients. To this end we introduce the Pareto efficiency criterion, which helps finding solutions that increase one objective without decreasing the other. Despite the vast size of the search space, we show how a state-of-the-art Genetic Algorithm can be tailored to find good solutions. Not only does the resulting descriptor perform better than state-of-the-art ones, but our approach is of broader significance as optimization problems with balanced goals are often encountered in Computer Vision. Michael Calonder, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2010 | Delineating trees in noisy 2D images and 3D image-stacksabstractWe present a novel approach to fully automated delineation of tree structures in noisy 2D images and 3D image stacks. Unlike earlier methods that rely mostly on local evidence, our method builds a set of candidate trees over many different subsets of points likely to belong to the final one and then chooses the best one according to a global objective function. Since we are not systematically trying to span all nodes, our algorithm is able to eliminate noise while retaining the right tree structure. Manually annotated dendrite micrographs and retinal scans are used to evaluate the performance of our method, which is shown to be able to reject noise while retaining the tree structure. Germán González, Engin Türetken, François Fleuret, Pascal Fua |
CVPR | 4 |
| 2010 | Dominant orientation templates for real-time detection of texture-less objectsabstractWe present a method for real-time 3D object detection that does not require a time consuming training stage, and can handle untextured objects. At its core, is a novel template representation that is designed to be robust to small image transformations. This robustness based on dominant gradient orientations lets us test only a small subset of all possible pixel locations when parsing the image, and to represent a 3D object with a limited set of templates. We show that together with a binary representation that makes evaluation very fast and a branch-and-bound approach to efficiently scan the image, it can detect untextured objects in complex situations and provide their 3D pose in real-time. Stefan Hinterstoißer, Vincent Lepetit, Slobodan Ilic, Pascal Fua, Nassir Navab |
CVPR | 4 |
| 2010 | Simultaneous pose, correspondence and non-rigid shapeabstractRecent works have shown that 3D shape of non-rigid surfaces can be accurately retrieved from a single image given a set of 3D-to-2D correspondences between that image and another one for which the shape is known. However, existing approaches assume that such correspondences can be readily established, which is not necessarily true when large deformations produce significant appearance changes between the input and the reference images. Furthermore, it is either assumed that the pose of the camera is known, or the estimated solution is pose-ambiguous. In this paper we relax all these assumptions and, given a set of 3D and 2D unmatched points, we present an approach to simultaneously solve their correspondences, compute the camera pose and retrieve the shape of the surface in the input image. This is achieved by introducing weak priors on the pose and shape that we model as Gaussian Mixtures. By combining them into a Kalman filter we can progressively reduce the number of 2D candidates that can be potentially matched to each 3D point, while pose and shape are refined. This lets us to perform a complete and efficient exploration of the solution space and retain the best solution. Jordi Sanchez-Riera, Jonas Östlund, Pascal Fua, Francesc Moreno-Noguer |
CVPR | 3 |
| 2010 | Simultaneous point matching and 3D deformable surface reconstructionabstractIt has been shown that the 3D shape of a deformable surface in an image can be recovered by establishing correspondences between that image and a reference one in which the shape is known. These matches can then be used to set-up a convex optimization problem in terms of the shape parameters, which is easily solved. However, in many cases, the correspondences are hard to establish reliably. In this paper, we show that we can solve simultaneously for both 3D shape and correspondences, thereby using 3D shape constraints to guide the image matching and increasing robustness, for example when the textures are repetitive. This involves solving a mixed integer quadratic problem. While optimizing this problem is NP-hard in general, we show that its solution can nevertheless be approximated effectively by a branch-and-bound algorithm. Appu Shaji, Aydin Varol, Lorenzo Torresani, Pascal Fua |
CVPR | 4 |
| 2010 | Dynamic and scalable large scale image reconstructionabstractRecent approaches to reconstructing city-sized areas from large image collections usually process them all at once and only produce disconnected descriptions of image subsets, which typically correspond to major landmarks. In contrast, we propose a framework that lets us take advantage of the available meta-data to build a single, consistent description from these potentially disconnected descriptions. Furthermore, this description can be incrementally updated and enriched as new images become available. We demonstrate the power of our approach by building large-scale reconstructions using images of Lausanne and Prague. Christoph Strecha, Timo Pylvänäinen, Pascal Fua |
CVPR | 3 |
| 2010 | BRIEF: Binary Robust Independent Elementary Features
Michael Calonder, Vincent Lepetit, Christoph Strecha, Pascal Fua |
ECCV (4) | 4 |
| 2010 | Exploring Ambiguities for Monocular Non-rigid Shape Estimation
Francesc Moreno-Noguer, Josep M. Porta, Pascal Fua |
ECCV (3) | 3 |
| 2010 | Combining Geometric and Appearance Priors for Robust Homography Estimation
Eduard Serradell, Mustafa Özuysal, Vincent Lepetit, Pascal Fua, Francesc Moreno-Noguer |
ECCV (3) | 4 |
| 2010 | Making Action Recognition Robust to Occlusions and Viewpoint Changes
Daniel Weinland, Mustafa Özuysal, Pascal Fua |
ECCV (3) | 3 |
| 2010 | A Fully Automated Approach to Segmentation of Irregularly Shaped Cellular Structures in EM Images
Aurélien Lucchi, Kevin Smith 0001, Radhakrishna Achanta, Vincent Lepetit, Pascal Fua |
MICCAI (2) | 5 |
| 2010 | Reconstructing Geometrically Consistent Tree Structures from Noisy Images
Engin Türetken, Christian Blum 0001, Germán González, Pascal Fua |
MICCAI (1) | 4 |
| 2010 | From Canonical Poses to 3D Motion Capture Using a Single CameraabstractWe combine detection and tracking techniques to achieve robust 3D motion recovery of people seen from arbitrary viewpoints by a single and potentially moving camera. We rely on detecting key postures, which can be done reliably, using a motion model to infer 3D poses between consecutive detections, and finally refining them over the whole sequence using a generative model. We demonstrate our approach in the cases of golf motions filmed using a static camera and walking motions acquired using a potentially moving one. We will show that our approach, although monocular, is both metrically accurate because it integrates information over many frames and robust because it can recover from a few misdetections. Andrea Fossati, Miodrag Dimitrijevic, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | Fast Keypoint Recognition Using Random FernsabstractWhile feature point recognition is a key component of modern approaches to object detection, existing approaches require computationally expensive patch preprocessing to handle perspective distortion. In this paper, we show that formulating the problem in a naive Bayesian classification framework makes such preprocessing unnecessary and produces an algorithm that is simple, efficient, and robust. Furthermore, it scales well as the number of classes grows. To recognize the patches surrounding keypoints, our classifier uses hundreds of simple binary features and models class posterior probabilities. We make the problem computationally tractable by assuming independence between arbitrary sets of features. Even though this is not strictly true, we demonstrate that our classifier nevertheless performs remarkably well on image data sets containing very significant perspective changes. Mustafa Özuysal, Michael Calonder, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | DAISY: An Efficient Dense Descriptor Applied to Wide-Baseline StereoabstractIn this paper, we introduce a local image descriptor, DAISY, which is very efficient to compute densely. We also present an EM-based algorithm to compute dense depth and occlusion maps from wide-baseline image pairs using this descriptor. This yields much better results in wide-baseline situations than the pixel and correlation-based algorithms that are commonly used in narrow-baseline stereo. Also, using a descriptor makes our algorithm robust against many photometric and geometric transformations. Our descriptor is inspired from earlier ones such as SIFT and GLOH but can be computed much faster for our purposes. Unlike SURF, which can also be computed efficiently at every pixel, it does not introduce artifacts that degrade the matching performance when used densely. It is important to note that our approach is the first algorithm that attempts to estimate dense depth maps from wide-baseline image pairs, and we show that it is a good one at that with many experiments for depth estimation accuracy, occlusion detection, and comparing it against other descriptors on laser-scanned ground truth scenes. We also tested our approach on a variety of indoor and outdoor scenes with different photometric and geometric transformations and our experiments support our claim to being robust against these. Engin Tola, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Appearance-based keypoint clusteringabstractWe present an algorithm for clustering sets of detected interest points into groups that correspond to visually distinct structure. Through the use of a suitable colour and texture representation, our clustering method is able to identify keypoints that belong to separate objects or background regions. These clusters are then used to constrain the matching of keypoints over pairs of images, resulting in greatly improved matching under difficult conditions. We present a thorough evaluation of each component of the algorithm, and show its usefulness on difficult matching problems. Francisco J. Estrada, Pascal Fua, Vincent Lepetit, Sabine Süsstrunk |
CVPR | 2 |
| 2009 | Observable subspaces for 3D human motion recoveryabstractThe articulated body models used to represent human motion typically have many degrees of freedom, usually expressed as joint angles that are highly correlated. The true range of motion can therefore be represented by latent variables that span a low-dimensional space. This has often been used to make motion tracking easier. However, learning the latent space in a problem- independent way makes it non trivial to initialize the tracking process by picking appropriate initial values for the latent variables, and thus for the pose. In this paper, we show that by directly using observable quantities as our latent variables, we eliminate this problem and achieve full automation given only modest amounts of training data. More specifically, we exploit the fact that the trajectory of a person's feet or hands strongly constrains body pose in motions such as skating, skiing, or golfing. These trajectories are easy to compute and to parameterize using a few variables. We treat these as our latent variables and learn a mapping between them and sequences of body poses. In this manner, by simply tracking the feet or the hands, we can reliably guess initial poses over whole sequences and, then, refine them. Andrea Fossati, Mathieu Salzmann, Pascal Fua |
CVPR | 3 |
| 2009 | Learning rotational features for filament detectionabstractState-of-the-art approaches for detecting filament-like structures in noisy images rely on filters optimized for signals of a particular shape, such as an ideal edge or ridge. While these approaches are optimal when the image conforms to these ideal shapes, their performance quickly degrades on many types of real data where the image deviates from the ideal model, and when noise processes violate a Gaussian assumption. In this paper, we show that by learning rotational features, we can outperform state-of-the-art filament detection techniques on many different kinds of imagery. More specifically, we demonstrate superior performance for the detection of blood vessel in retinal scans, neurons in brightfield microscopy imagery, and streets in satellite imagery. Germán González, François Fleuret, Pascal Fua |
CVPR | 3 |
| 2009 | Real-time learning of accurate patch rectificationabstractRecent work showed that learning-based patch rectification methods are both faster and more reliable than affine region methods. Unfortunately, their performance improvements are founded in a computationally expensive offline learning stage, which is not possible for applications such as SLAM. In this paper we propose an approach whose training stage is fast enough to be performed at run-time without the loss of accuracy or robustness. To this end, we developed a very fast method to compute the mean appearances of the feature points over sets of small variations that span the range of possible camera viewpoints. Then, by simply matching incoming feature points against these mean appearances, we get a coarse estimate of the viewpoint that is refined afterwards. Because there is no need to compute descriptors for the input image, the method is very fast at run-time. We demonstrate our approach on tracking-by-detection for SLAM, real-time object detection and pose estimation applications. Stefan Hinterstoißer, Oliver Kutter, Nassir Navab, Pascal Fua, Vincent Lepetit |
CVPR | 4 |
| 2009 | Capturing 3D stretchable surfaces from single images in closed formabstractWe present a closed form solution to the problem of recovering the 3D shape of a nonrigid potentially stretchable surface from 3D-to-2D correspondences. In other words, we can reconstruct a surface from a single image without a priori knowledge of its deformations in that image. State of the art solutions to nonrigid 3D shape recovery rely on the fact that distances between neighboring surface points must be preserved and are therefore limited to inelastic surfaces. Here, we show that replacing the inextensibility constraints by shading ones removes this limitation while still allowing 3D reconstruction in closed-form. We demonstrate our method and compare it to an earlier one using both synthetic and real data. Francesc Moreno-Noguer, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
CVPR | 4 |
| 2009 | Pose estimation for category specific multiview object localizationabstractWe propose an approach to overcome the two main challenges of 3D multiview object detection and localization: The variation of object features due to changes in the viewpoint and the variation in the size and aspect ratio of the object. Our approach proceeds in three steps. Given an initial bounding box of fixed size, we first refine its aspect ratio and size. We can then predict the viewing angle, under the hypothesis that the bounding box actually contains an object instance. Finally, a classifier tuned to this particular viewpoint checks the existence of an instance. As a result, we can find the object instances and estimate their poses, without having to search over all window sizes and potential orientations. We train and evaluate our method on a new object database specifically tailored for this task, containing real-world objects imaged over a wide range of smoothly varying viewpoints and significant lighting changes. We show that the successive estimations of the bounding box and the viewpoint lead to better localization results. Mustafa Özuysal, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2009 | Reconstructing sharply folding surfaces: A convex formulationabstractIn recent years, 3D deformable surface reconstruction from single images has attracted renewed interest. It has been shown that preventing the surface from either shrinking or stretching is an effective way to resolve the ambiguities inherent to this problem. However, while the geodesic distances on the surface may not change, the Euclidean ones decrease when folds appear. Therefore, when applied to discrete surface representations, such constant-distance constraints are only effective for smoothly deforming surfaces, and become inaccurate for more flexible ones that can exhibit sharp folds. In such cases, surface points must be allowed to come closer to each other. In this paper, we show that replacing the equality constraints of earlier approaches by inequality constraints that let the mesh representation of the surface shrink but not expand yields not only a more faithful representation, but also a convex formulation of the reconstruction problem. As a result, we can accurately reconstruct surfaces undergoing complex deformations that include sharp folds from individual images. Mathieu Salzmann, Pascal Fua |
CVPR | 2 |
| 2009 | Joint pose estimator and feature learning for object detectionabstractA new learning strategy for object detection is presented. The proposed scheme forgoes the need to train a collection of detectors dedicated to homogeneous families of poses, and instead learns a single classifier that has the inherent ability to deform based on the signal of interest. Specifically, we train a detector with a standard AdaBoost procedure by using combinations of pose-indexed features and pose estimators instead of the usual image features. This allows the learning process to select and combine various estimates of the pose with features able to implicitly compensate for variations in pose. We demonstrate that a detector built in such a manner provides noticeable gains on two hand video sequences and analyze the performance of our detector as these data sets are synthetically enriched in pose while not increased in size. Karim Ali 0002, François Fleuret, David Hasler, Pascal Fua |
ICCV | 4 |
| 2009 | Compact signatures for high-speed interest point description and matchingabstractProminent feature point descriptors such as SIFT and SURF allow reliable real-time matching but at a computational cost that limits the number of points that can be handled on PCs, and even more on less powerful mobile devices. A recently proposed technique that relies on statistical classification to compute signatures has the potential to be much faster but at the cost of using very large amounts of memory, which makes it impractical for implementation on low-memory devices. In this paper, we show that we can exploit the sparseness of these signatures to compact them, speed up the computation, and drastically reduce memory usage. We base our approach on Compressive Sensing theory. We also highlight its effectiveness by incorporating it into two very different SLAM packages and demonstrating substantial performance increases. Michael Calonder, Vincent Lepetit, Pascal Fua, Kurt Konolige, James Bowman, Patrick Mihelich |
ICCV | 3 |
| 2009 | Template-free monocular reconstruction of deformable surfacesabstractIt has recently been shown that deformable 3D surfaces could be recovered from single video streams. However, existing techniques either require a reference view in which the shape of the surface is known a priori, which often may not be available, or require tracking points over long sequences, which is hard to do. In this paper, we overcome these limitations. To this end, we establish correspondences between pairs of frames in which the shape is different and unknown. We then estimate homographies between corresponding local planar patches in both images. These yield approximate 3D reconstructions of points within each patch up to a scale factor. Since we consider overlapping patches, we can enforce them to be consistent over the whole surface. Finally, a local deformation model is used to fit a triangulated mesh to the 3D point cloud, which makes the reconstruction robust to both noise and outliers in the image data. Aydin Varol, Mathieu Salzmann, Engin Tola, Pascal Fua |
ICCV | 4 |
| 2009 | Steerable Features for Statistical 3D Dendrite Detection
Germán González, François Aguet, François Fleuret, Michael Unser, Pascal Fua |
MICCAI (1) | 5 |
| 2009 | EPnP: An Accurate O(n) Solution to the PnP Problem
Vincent Lepetit, Francesc Moreno-Noguer, Pascal Fua |
Int. J. Comput. Vis. | 3 |
| 2009 | Classification-Based Probabilistic Modeling of Texture Transition for Fast Line Search Tracking and DelineationabstractWe introduce a classification-based approach to finding occluding texture boundaries. The classifier is composed of a set of weak learners which operate on image intensity discriminative features which are defined on small patches and fast to compute. A database which is designed to simulate digitized occluding contours of textured objects in natural images is used to train the weak learners. The trained classifier score is then used to obtain a probabilistic model for the presence of texture transitions which can readily be used for line search texture boundary detection in the direction normal to an initial boundary estimate. This method is fast and therefore suitable for real-time and interactive applications. It works as a robust estimator which requires a ribbon like search region and can handle complex texture structures without requiring a large number of observations. We demonstrate results both in the context of interactive 2-D delineation and fast 3-D tracking and compare its performance with other existing methods for line search boundary detection. Ali Shahrokni, Tom Drummond, François Fleuret, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2008 | Simultaneous Recognition and Homography Extraction of Local Patches with a Simple Linear ClassifierabstractWe show that the simultaneous estimation of keypoint identities and poses is more reliable than the two separate steps undertaken by previous approaches. A simple linear classifier coupled with linear predictors trained during a learning phase appears to be sufficient for this task. The retrieved poses are subpixel accurate due to the linear predictors. We demonstrate the advantages of our approach on real-time 3D object detection and tracking applications. Thanks to the high accuracy, one single keypoint is often enough to precisely estimate the object pose. As a result, we can deal in real-time with objects that are significantly less textured than the ones required by state-of-the-art methods. 1 Stefan Hinterstoißer, Selim Benhimane, Vincent Lepetit, Pascal Fua, Nassir Navab |
BMVC | 4 |
| 2008 | Online learning of patch perspective rectification for efficient object detectionabstractFor a large class of applications, there is time to train the system. In this paper, we propose a learning-based approach to patch perspective rectification, and show that it is both faster and more reliable than state-of-the-art ad hoc affine region detection methods. Our method performs in three steps. First, a classifier provides for every keypoint not only its identity, but also a first estimate of its transformation. This estimate allows carrying out, in the second step, an accurate perspective rectification using linear predictors. We show that both the classifier and the linear predictors can be trained online, which makes the approach convenient. The last step is a fast verification - made possible by the accurate perspective rectification - of the patch identity and its sub-pixel precision position estimation. We test our approach on real-time 3D object detection and tracking applications. We show that we can use the estimated perspective rectifications to determine the object pose and as a result, we need much fewer correspondences to obtain a precise pose estimation. Stefan Hinterstoißer, Selim Benhimane, Nassir Navab, Pascal Fua, Vincent Lepetit |
CVPR | 4 |
| 2008 | 3D pose refinement from reflectionsabstractWe demonstrate how to exploit reflections for accurate registration of shiny objects: The lighting environment can be retrieved from the reflections under a distant illumination assumption. Since it remains unchanged when the camera or the object of interest moves, this provides powerful additional constraints that can be incorporated into standard pose estimation algorithms. The key idea and main contribution of the paper is therefore to show that the registration should also be performed in the lighting environment space, instead of in the image space only. This lets us recover very accurate pose estimates because the specularities are very sensitive to pose changes. An interesting side result is an accurate estimate of the lighting environment. Furthermore, since the mapping from lighting environment to specularities has no analytical expression for objects represented as 3D meshes, and is not 1-to-1, registering lighting environments is far from trivial. However we propose a general and effective solution. Our approach is demonstrated on both synthetic and real images. Pascal Lagger, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
CVPR | 4 |
| 2008 | Local deformation models for monocular 3D shape recoveryabstractWithout a deformation model, monocular 3D shape recovery of deformable surfaces is severely under-constrained. Even when the image information is rich enough, prior knowledge of the feasible deformations is required to overcome the ambiguities. This is further accentuated when such information is poor, which is a key issue that has not yet been addressed. In this paper, we propose an approach to learning shape priors to solve this problem. By contrast with typical statistical learning methods that build models for specific object shapes, we learn local deformation models, and combine them to reconstruct surfaces of arbitrary global shapes. Not only does this improve the generality of our deformation models, but it also facilitates learning since the space of local deformations is much smaller than that of global ones. While using a texture-based approach, we show that our models are effective to reconstruct from single videos poorly-textured surfaces of arbitrary shape, made of materials as different as cardboard, that deforms smoothly, and much lighter tissue paper whose deformations may be far more complex. Mathieu Salzmann, Raquel Urtasun, Pascal Fua |
CVPR | 3 |
| 2008 | On benchmarking camera calibration and multi-view stereo for high resolution imageryabstractIn this paper we want to start the discussion on whether image based 3-D modelling techniques can possibly be used to replace LIDAR systems for outdoor 3D data acquisition. Two main issues have to be addressed in this context: (i) camera calibration (internal and external) and (ii) dense multi-view stereo. To investigate both, we have acquired test data from outdoor scenes both with LIDAR and cameras. Using the LIDAR data as reference we estimated the ground-truth for several scenes. Evaluation sets are prepared to evaluate different aspects of 3D model building. These are: (i) pose estimation and multi-view stereo with known internal camera parameters; (ii) camera calibration and multi-view stereo with the raw images as the only input and (iii) multi-view stereo. Christoph Strecha, Wolfgang von Hansen, Luc Van Gool, Pascal Fua, Ulrich Thoennessen |
CVPR | 4 |
| 2008 | A fast local descriptor for dense matchingabstractWe introduce a novel local image descriptor designed for dense wide-baseline matching purposes. We feed our descriptors to a graph-cuts based dense depth map estimation algorithm and this yields better wide-baseline performance than the commonly used correlation windows for which the size is hard to tune. As a result, unlike competing techniques that require many high-resolution images to produce good reconstructions, our descriptor can compute them from pairs of low-quality images such as the ones captured by video streams. Our descriptor is inspired from earlier ones such as SIFT and GLOH but can be computed much faster for our purposes. Unlike SURF which can also be computed efficiently at every pixel, it does not introduce artifacts that degrade the matching performance. Our approach was tested with ground truth laser scanned depth maps as well as on a wide variety of image pairs of different resolutions and we show that good reconstructions are achieved even with only two low quality images. Engin Tola, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2008 | Multi-camera Tracking and Atypical Motion Detection with Behavioral Maps
Jérôme Berclaz, François Fleuret, Pascal Fua |
ECCV (3) | 3 |
| 2008 | Keypoint Signatures for Fast Learning and Recognition
Michael Calonder, Vincent Lepetit, Pascal Fua |
ECCV (1) | 3 |
| 2008 | Linking Pose and Motion
Andrea Fossati, Pascal Fua |
ECCV (4) | 2 |
| 2008 | Automated Delineation of Dendritic Networks in Noisy Image Stacks
Germán González, François Fleuret, Pascal Fua |
ECCV (4) | 3 |
| 2008 | Pose Priors for Simultaneously Solving Alignment and Correspondence
Francesc Moreno-Noguer, Vincent Lepetit, Pascal Fua |
ECCV (2) | 3 |
| 2008 | Making Background Subtraction Robust to Sudden Illumination Changes
Julien Pilet, Christoph Strecha, Pascal Fua |
ECCV (4) | 3 |
| 2008 | Closed-Form Solution to Non-rigid 3D Surface Registration
Mathieu Salzmann, Francesc Moreno-Noguer, Vincent Lepetit, Pascal Fua |
ECCV (4) | 4 |
| 2008 | The haunted bookabstractThis paper describes an artwork that relies on recent computer vision and augmented reality techniques to animate the illustrations of a poetry book. Because we donpsilat need markers, we can achieve seamless integration of real and virtual elements to create the desired atmosphere. The visualization is done on a computer screen to avoid cumbersome head-mounted displays. The camera is hidden into a desk lamp for easing even more the spectator immersion. Camille Scherrer, Julien Pilet, Pascal Fua, Vincent Lepetit |
ISMAR | 3 |
| 2008 | Retrieving multiple light sources in the presence of specular reflections and texture
Pascal Lagger, Pascal Fua |
Comput. Vis. Image Underst. | 2 |
| 2008 | Fast Non-Rigid Surface Detection, Registration and Realistic Augmentation
Julien Pilet, Vincent Lepetit, Pascal Fua |
Int. J. Comput. Vis. | 3 |
| 2008 | Multicamera People Tracking with a Probabilistic Occupancy MapabstractGiven two to four synchronized video streams taken at eye level and from different angles, we show that we can effectively combine a generative model with dynamic programming to accurately follow up to six individuals across thousands of frames in spite of significant occlusions and lighting changes. In addition, we also derive metrically accurate trajectories for each one of them. Our contribution is twofold. First, we demonstrate that our generative model can effectively handle occlusions in each time frame independently, even when the only data available comes from the output of a simple background subtraction algorithm and when the number of individuals is unknown a priori. Second, we show that multi-person tracking can be reliably achieved by processing individual trajectories separately over long sequences, provided that a reasonable heuristic is used to rank these individuals and avoid confusing them with one another. François Fleuret, Jérôme Berclaz, Richard Lengagne, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Robust Multi-View Change DetectionabstractWe present a multi-view change detection approach aimed at being robust \nwith respect to common “disturbance factors” yielding image changes in realworld \napplications. Disturbance factors causing “slow” or “fast-and-global” \nimage variations, such as light changes and dynamic adjustments of camera \nparameters (e.g. auto-exposure and auto-gain control), are dealt with by a \nproper single-view change detector run independently on each view. The \ncomputed change masks are then fused into a “synergy mask” defined into a \ncommon virtual top-view, so as to detect and filter-out “fast-and-local” image \nchanges due to physical points lying on the ground surface (e.g. shadows cast \nby moving objects and light spots hitting the ground surface). Alessandro Lanza, Luigi Di Stefano, Jérôme Berclaz, François Fleuret, Pascal Fua |
BMVC | 5 |
| 2007 | Bridging the Gap between Detection and Tracking for 3D Monocular Video-Based Motion CaptureabstractWe combine detection and tracking techniques to achieve robust 3-D motion recovery of people seen from arbitrary viewpoints by a single and potentially moving camera. We rely on detecting key postures, which can be done reliably, using a motion model to infer 3-D poses between consecutive detections, and finally refining them over the whole sequence using a generative model. We demonstrate our approach in the case of people walking against cluttered backgrounds and filmed using a moving camera, which precludes the use of simple background subtraction techniques. In this case, the easy-to-detect posture is the one that occurs at the end of each step when people have their legs furthest apart. Andrea Fossati, Miodrag Dimitrijevic, Vincent Lepetit, Pascal Fua |
CVPR | 4 |
| 2007 | Fast Keypoint Recognition in Ten Lines of CodeabstractWhile feature point recognition is a key component of modern approaches to object detection, existing approaches require computationally expensive patch preprocessing to handle perspective distortion. In this paper, we show that formulating the problem in a Naive Bayesian classification framework makes such preprocessing unnecessary and produces an algorithm that is simple, efficient, and robust. Furthermore, it scales well to handle large number of classes. To recognize the patches surrounding keypoints, our classifier uses hundreds of simple binary features and models class posterior probabilities. We make the problem computationally tractable by assuming independence between arbitrary sets of features. Even though this is not strictly true, we demonstrate that our classifier nevertheless performs remarkably well on image datasets containing very significant perspective changes. Mustafa Özuysal, Pascal Fua, Vincent Lepetit |
CVPR | 2 |
| 2007 | Deformable Surface Tracking AmbiguitiesabstractWe study from a theoretical standpoint the ambiguities that occur when tracking a generic deformable surface under monocular perspective projection given 3D to 2D correspondences. We show that, additionally to the known scale ambiguity, a set of potential ambiguities can be clearly identified. From this, we deduce a minimal set of constraints required to disambiguate the problem and incorporate them into a working algorithm that runs on real noisy data. Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
CVPR | 3 |
| 2007 | Non-Linear Beam Model for Tracking Large DeformationsabstractIn this paper we investigate physics-based plane beam model, frequently used in mechanical and civil engineering, to track large non-linear deformations in images. Such models do not only contribute to robust and precise tracking, in the presence of clutter and partial occlusions, but also allow to compute the forces that produce observed deformations. We verify the correctness of the recovered forces by using them in a simulation and compare the results to the original image displacements. We apply this method to track deformations of the pole vault, the rat whiskers and the car antenna. Slobodan Ilic, Pascal Fua |
ICCV | 2 |
| 2007 | Accurate Non-Iterative O(n) Solution to the PnP ProblemabstractWe propose a non-iterative solution to the PnP problem-the estimation of the pose of a calibrated camera from n 3D-to-2D point correspondences—whose computational complexity grows linearly with 𝑛5) or even 𝑂(𝑛8), without being more accurate. Our method is applicable for all 𝑛≥4 and handles properly both planar and non-planar configurations. Our central idea is to express the 𝑛 3D points as a weighted sum of four virtual control points. The problem then reduces to estimating the coordinates of these control points in the camera referential, which can be done in 𝑂(𝑛) time by expressing these coordinates as weighted sum of the eigenvectors of a 12 × 12 matrix and solving a small constant number of quadratic equations to pick the right weights. The advantages of our method are demonstrated by thorough testing on both synthetic and real-data. Francesc Moreno-Noguer, Vincent Lepetit, Pascal Fua |
ICCV | 3 |
| 2007 | Convex Optimization for Deformable Surface 3-D Trackingabstract3-D shape recovery of non-rigid surfaces from 3-D to 2-D correspondences is an under-constrained problem that requires prior knowledge of the possible deformations. State-of-the-art solutions involve enforcing smoothness constraints that limit their applicability and prevent the recovery of sharply folding and creasing surfaces. Here, we propose a method that does not require such smoothness constraints. Instead, we represent surfaces as triangulated meshes and, assuming the pose in the first frame to be known, disallow large changes of edge orientation between consecutive frames, which is a generally applicable constraint when tracking surfaces in a 25 frames- per-second video sequence. We will show that tracking under these constraints can be formulated as a Second Order Cone Programming feasibility problem. This yields a convex optimization problem with stable solutions for a wide range of surfaces with very different physical properties. Mathieu Salzmann, Richard I. Hartley, Pascal Fua |
ICCV | 3 |
| 2007 | Retexturing in the Presence of Complex Illumination and OcclusionsabstractWe present a nonrigid registration technique that achieves spatial, photometric, and visibility accuracy. It lets us photo-realistically augment 3D deformable surfaces under complex illumination conditions and in spite of severe occlusions. There are many approaches that address some of these issues but very few that simultaneously handle all of them as we do. We use triangulated meshes to model the geometry and introduce explicit visibility maps as well as separate illumination parameters for each mesh vertex. We cast our registration problem in an expectation maximization framework that allows robust and fully automated operation. It provides explicit illumination and occlusion models that can be used for rendering purposes. Julien Pilet, Vincent Lepetit, Pascal Fua |
ISMAR | 3 |
| 2007 | Implicit Meshes for Effective Silhouette Handling
Slobodan Ilic, Mathieu Salzmann, Pascal Fua |
Int. J. Comput. Vis. | 3 |
| 2007 | Surface Deformation Models for Nonrigid 3D Shape RecoveryabstractThree-dimensional detection and shape recovery of a nonrigid surface from video sequences require deformation models to effectively take advantage of potentially noisy image data. Here, we introduce an approach to creating such models for deformable 3D surfaces. We exploit the fact that the shape of an inextensible triangulated mesh can be parameterized in terms of a small subset of the angles between its facets. We use this set of angles to create a representative set of potential shapes, which we feed to a simple dimensionality reduction technique to produce low-dimensional 3D deformation models. We show that these models can be used to accurately model a wide range of deforming 3D surfaces from video sequences acquired under realistic conditions. Mathieu Salzmann, Julien Pilet, Slobodan Ilic, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2006 | Robust People Tracking with Global Trajectory OptimizationabstractGiven three or four synchronized videos taken at eye level and from different angles, we show that we can effectively use dynamic programming to accurately follow up to six individuals across thousands of frames in spite of significant occlusions. In addition, we also derive metrically accurate trajectories for each one of them. Our main contribution is to show that multi-person tracking can be reliably achieved by processing individual trajectories separately over long sequences, provided that a reasonable heuristic is used to rank these individuals and avoid confusing them with one another. In this way, we achieve robustness by finding optimal trajectories over many frames while avoiding the combinatorial explosion that would result from simultaneously dealing with all the individuals. Jérôme Berclaz, François Fleuret, Pascal Fua |
CVPR (1) | 3 |
| 2006 | 3D People Tracking with Gaussian Process Dynamical ModelsabstractWe advocate the use of Gaussian Process Dynamical Models (GPDMs) for learning human pose and motion priors for 3D people tracking. A GPDM provides a lowdimensional embedding of human motion data, with a density function that gives higher probability to poses and motions close to the training data. With Bayesian model averaging a GPDM can be learned from relatively small amounts of data, and it generalizes gracefully to motions outside the training set. Here we modify the GPDM to permit learning from motions with significant stylistic variation. The resulting priors are effective for tracking a range of human walking styles, despite weak and noisy image measurements and significant occlusions. Raquel Urtasun, David J. Fleet, Pascal Fua |
CVPR (1) | 3 |
| 2006 | Feature Harvesting for Tracking-by-Detection
Mustafa Özuysal, Vincent Lepetit, François Fleuret, Pascal Fua |
ECCV (3) | 4 |
| 2006 | An all-in-one solution to geometric and photometric calibrationabstractWe propose a fully automated approach to calibrating multiple cameras whose fields of view may not all overlap. Our technique only requires waving an arbitrary textured planar pattern in front of the cameras, which is the only manual intervention that is required. The pattern is then automatically detected in the frames where it is visible and used to simultaneously recover geometric and photometric camera calibration parameters. In other words, even a novice user can use our system to extract all the information required to add virtual 3D objects into the scene and light them convincingly. This makes it ideal for Augmented Reality applications and we distribute the code under a GPL license. Julien Pilet, Andreas Geiger 0001, Pascal Lagger, Vincent Lepetit, Pascal Fua |
ISMAR | 5 |
| 2006 | Human body pose detection using Bayesian spatio-temporal templates
Miodrag Dimitrijevic, Vincent Lepetit, Pascal Fua |
Comput. Vis. Image Underst. | 3 |
| 2006 | Modeling people: Vision-based understanding of a person's shape, appearance, movement, and behaviour
Adrian Hilton 0001, Pascal Fua, Rémi Ronfard |
Comput. Vis. Image Underst. | 2 |
| 2006 | Temporal motion models for monocular and multiview 3D human body tracking
Raquel Urtasun, David J. Fleet, Pascal Fua |
Comput. Vis. Image Underst. | 3 |
| 2006 | Implicit Meshes for Surface ReconstructionabstractDeformable 3D models can be represented either as traditional explicit surfaces, such as triangulated meshes, or as implicit surfaces. Explicit surfaces are widely accepted because they are simple to deform and render, but fitting them involves minimizing a nondifferentiable distance function. By contrast, implicit surfaces allow fitting by minimizing a differentiable algebraic distance, but are harder to meaningfully deform and render. Here, we propose a method that combines the strength of both approaches. It relies on a technique that can turn a completely arbitrary triangulated mesh, such as one taken from the Web, into an implicit surface that closely approximates it and can deform in tandem with it. This allows both automated algorithms to take advantage of the attractive properties of implicit surfaces for fitting purposes and people to use standard deformation tools they feel comfortable for interaction and animation purposes. We demonstrate the applicability of our technique to modeling the human upper-body, including face, neck, shoulders, and ears, from noisy stereo and silhouette data. Slobodan Ilic, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Keypoint Recognition Using Randomized TreesabstractIn many 3D object-detection and pose-estimation problems, runtime performance is of critical importance. However, there usually is time to train the system, which we will show to be very useful. Assuming that several registered images of the target object are available, we developed a keypoint-based approach that is effective in this context by formulating wide-baseline matching of keypoints extracted from the input images to those found in the model images as a classification problem. This shifts much of the computational burden to a training phase, without sacrificing recognition performance. As a result, the resulting algorithm is robust, accurate, and fast-enough for frame-rate performance. This reduction in runtime computational complexity is our first contribution. Our second contribution is to show that, in this context, a simple and fast keypoint detector suffices to support detection and tracking even under large perspective and scale variations. While earlier methods require a detector that can be expected to produce very repeatable results, in general, which usually is very time-consuming, we simply find the most repeatable object keypoints for the specific target object during the training phase. We have incorporated these ideas into a real-time system that detects planar, nonplanar, and deformable objects. It then estimates the pose of the rigid ones and the deformations of the others. Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2005 | Physically Valid Shape Parameterization for Monocular 3-D Deformable Surface TrackingabstractCVLAB Mathieu Salzmann, Slobodan Ilic, Pascal Fua |
BMVC | 3 |
| 2005 | Classifier-based Contour Tracking for Rigid and Deformable ObjectsabstractThis paper proposes a machine learning approach to the problem of modelbased contour tracking for rigid or deformable objects. The motion of the target is calculated by tracking its contours in a video sequence. We develop a probabilistic representation of contours that allows robust contour tracking in presence of texture and clutter. We use boosting to train a predictor of the conditional probability of texture transition, given the pixel intensities. The most likely connected contours are obtained by maximizing the posterior probability of object model parameters. The probabilistic formulation allows spatial connectivity of the contours to be formulated in a natural manner. For deformable objects, we use a Hidden Markov Model to calculate the joint law of the conditional probabilities of contour points while for rigid objects geometric properties of the model are used in a framework of random sample consensus algorithm to find the optimal model pose. We demonstrate that the proposed method is fast and robust for tracking deformable and rigid objects. We also compare our algorithm to several other contour tracking methods. 1 Ali Shahrokni, François Fleuret, Pascal Fua |
BMVC | 3 |
| 2005 | Implicit Surfaces Make for Better SilhouettesabstractThis paper advocates an implicit-surface representation of generic 3-D surfaces to take advantage of occluding edges in a very robust way. This lets us exploit silhouette constraints in uncontrolled environments that may involve occlusions and changing or cluttered backgrounds, which limit the applicability of most silhouette based methods. This desirable behavior is completely independent from the way the surface deformations are parametrized. To show this, we demonstrate our technique in three very different cases: modeling the deformations of a piece of paper represented by an ordinary triangulated mesh; tracking a person's shoulders whose deformations are expressed in terms of Dirichlet free form deformations; reconstructing the shape of a human face parametrized in terms of a principal component analysis model. Slobodan Ilic, Mathieu Salzmann, Pascal Fua |
CVPR (1) | 3 |
| 2005 | Randomized Trees for Real-Time Keypoint RecognitionabstractIn earlier work, we proposed treating wide baseline matching of feature points as a classification problem, in which each class corresponds to the set of all possible views of such a point. We used a K-mean plus Nearest Neighbor classifier to validate our approach, mostly because it was simple to implement. It has proved effective but still too slow for real-time use. In this paper, we advocate instead the use of randomized trees as the classification technique. It is both fast enough for real-time performance and more robust. It also gives us a principled way not only to match keypoints but to select during a training phase those that are the most recognizable ones. This results in a real-time system able to detect and position in 3D planar, non-planar, and even deformable objects. It is robust to illuminations changes, scale changes and occlusions. Vincent Lepetit, Pascal Lagger, Pascal Fua |
CVPR (2) | 3 |
| 2005 | Real-Time Non-Rigid Surface DetectionabstractWe present a real-time method for detecting deformable surfaces, with no need whatsoever for a priori pose knowledge. Our method starts from a set of wide baseline point matches between an undeformed image of the object and the image in which it is to be detected. The matches are used not only to detect but also to compute a precise mapping from one to the other. The algorithm is robust to large deformations, lighting changes, motion blur, and occlusions. It runs at 10 frames per second on a 2.8 GHz PC and we are not aware of any other published technique that produces similar results. Combining deformable meshes with a well designed robust estimator is key to dealing with the large number of parameters involved in modeling deformable surfaces and rejecting erroneous matches for error rates of up to 95%, which is considerably more than what is required in practice. Julien Pilet, Vincent Lepetit, Pascal Fua |
CVPR (1) | 3 |
| 2005 | Monocular 3-D Tracking of the Golf SwingabstractWe propose an approach to incorporating dynamic models into the human body tracking process that yields full 3D reconstructions from monocular sequences. We formulate the tracking problem in terms of minimizing a differentiable criterion whose differential structure is rich enough for successful optimization using a simple hill-climbing approach as opposed to a multihypotheses probabilistic one. In other words, we avoid the computational complexity of multihypotheses algorithms while obtaining excellent results under challenging conditions. To demonstrate this, we focus on monocular tracking of a golf swing from ordinary video. It involves both dealing with potentially very different swing styles, recovering arm motions that are perpendicular to the camera plane and handling strong self-occlusions. Raquel Urtasun, David J. Fleet, Pascal Fua |
CVPR (2) | 3 |
| 2005 | Monocular 3D Tracking of the Golf SwingabstractThis paper showcases our tracking algorithm described in (R. Urtasun et al., 2005). The proposed approach incorporates dynamic motion models into the human body tracking process yielding full 3D reconstruction from monocular sequences. The tracking is formulated in terms of minimizing a differentiable criterion whose differential structure is rich enough for successful optimization using a single hypothesis. In other words we avoid the computational complexity of multihypotheses algorithms while obtaining excellent results under challenging conditions. Raquel Urtasun, David J. Fleet, Pascal Fua |
CVPR (2) | 3 |
| 2005 | Fixed Point Probability Field for Complex Occlusion HandlingabstractIn this paper, we show that in a multi-camera context, we can effectively handle occlusions in real-time at each frame independently, even when the only available data comes from the binary output of a simple blob detector, and the number of present individuals is a priori unknown. We start from occupancy probability estimates in a top view and rely on a generative model to yield probability images to be compared with the actual input images. We then refine the estimates so that the probability images match the binary input images as well as possible. We demonstrate the quality of our results on several sequences involving complex occlusions. François Fleuret, Richard Lengagne, Pascal Fua |
ICCV | 3 |
| 2005 | Fast Texture-Based Tracking and Delineation Using Texture EntropyabstractWe propose a fast texture-segmentation approach to the problem of 2D and 3D model-based contour tracking, which is suitable for real-time or interactive applications. Our approach relies on detecting texture boundaries in the direction normal to the contour boundaries and on using a hidden Markov model to link these boundary points in the other direction. The probabilities that appear in this computation closely relate to texture entropy and Kullback-Leibler divergence, a property we use to compute and update dynamic texture models. We demonstrate results both in the context of interactive 2D delineation and fast 3D tracking Ali Shahrokni, Tom Drummond, Pascal Fua |
ICCV | 3 |
| 2005 | Priors for People Tracking from Small Training SetsabstractWe advocate the use of scaled Gaussian process latent variable models (SGPLVM) to learn prior models of 3D human pose for 3D people tracking. The SGPLVM simultaneously optimizes a low-dimensional embedding of the high-dimensional pose data and a density function that both gives higher probability to points close to training data and provides a nonlinear probabilistic mapping from the low-dimensional latent space to the full-dimensional pose space. The SGPLVM is a natural choice when only small amounts of training data are available. We demonstrate our approach with two distinct motions, golfing and walking. We show that the SGPLVM sufficiently constrains the problem such that tracking can be accomplished with straightforward deterministic optimization. Raquel Urtasun, David J. Fleet, Aaron Hertzmann, Pascal Fua |
ICCV | 4 |
| 2005 | Augmenting Deformable Objects in Real-TimeabstractWe present a real-time system that can draw virtual patterns or images on deforming real objects by estimating both the deformations and the shading parameters. We show that this is what is required to render the virtual elements so that they blend convincingly with the surrounding real textures. The whole process of uncompressing the video stream, measuring the deformations, estimating the lighting parameters, and realistically augmenting the input image takes about 100 ms on a 2.8 GHz PC. It is fully automated and does not require any manual initialization or engineering of the scene. It is also robust to large deformations, lighting changes, motion blur, specularities, and occlusions. It can therefore be demonstrated live on a simple laptop. Julien Pilet, Vincent Lepetit, Pascal Fua |
ISMAR | 3 |
| 2005 | Hierarchical implicit surface joint limits for human body tracking
Lorna Herda, Raquel Urtasun, Pascal Fua |
Comput. Vis. Image Underst. | 3 |
| 2004 | Human Shape and Motion from VideoabstractIn recent years, because cameras have become inexpensive and ever more prevalent, there has been increasing interest in modeling human shape and motion from image data. This type of modeling has many applications, such as electronic publishing, entertainment, sports medicine and athletic training. This, however, is an inherently difficult task, both because the body is very complex and because the data that can be extracted from images is often incomplete, noisy and ambiguous. EPFL’s Computer Vision Laboratory seeks to overcome these difficulties by using facial and body animation models, not only to represent the data, but also to guide the fitting process, thereby substantially improving performance. Start from sophisticated 3-D animation models, we reformulate them so that they can be used for data analysis in the three following research areas. 1 Augmented reality and 3-D tracking In augmented reality applications, tracking and registration of cameras and objects are Pascal Fua |
BMVC | 1 |
| 2004 | Markov-based Silhouette Extraction for Three--Dimensional Body Tracking in Presence of Cluttered BackgroundabstractWe propose a novel method to detect human body contours in presence of clutter and complex texture. Contours are extracted using a novel Markovbased approach which learns a texture along a given scanline in order to detect texture crossings. In contrast to conventional silhouette detection algorithms based on gradient, our texture boundary detection method allows extraction of silhouettes of textured and non-textured objects under difficult conditions such as having a cluttered/moving background. We demonstrate on demanding examples of monocular body tracking that our proposed method yields better results than gradient-based techniques. 1. Ali Shahrokni, Vincent Lepetit, Tom Drummond, Pascal Fua |
BMVC | 4 |
| 2004 | Accurate Face Models from Uncalibrated and Ill-Lit Video Sequences
Miodrag Dimitrijevic, Slobodan Ilic, Pascal Fua |
CVPR (2) | 3 |
| 2004 | Point Matching as a Classification Problem for Fast and Robust Object Pose Estimation
Vincent Lepetit, Julien Pilet, Pascal Fua |
CVPR (2) | 3 |
| 2004 | Hierarchical Implicit Surface Joint Limits to Constrain Video-Based Motion Capture
Lorna Herda, Raquel Urtasun, Pascal Fua |
ECCV (2) | 3 |
| 2004 | Texture Boundary Detection for Real-Time Tracking
Ali Shahrokni, Tom Drummond, Pascal Fua |
ECCV (2) | 3 |
| 2004 | 3D Human Body Tracking Using Deterministic Temporal Motion Models
Raquel Urtasun, Pascal Fua |
ECCV (3) | 2 |
| 2004 | ombining Edge and Texture Information for Real-Time Accurate 3D Camera TrackingabstractWe present an effective way to combine the information provided by edges and by feature points for the purpose of robust real-time 3-D tracking. This lets our tracker handle both textured and untextured objects. As it can exploit more of the image information, it is more stable and less prone to drift that purely edge or feature-based ones. We start with a feature-point based tracker we developed in earlier work and integrate the ability to take edge-information into account. Achieving optimal performance in the presence of cluttered or textured backgrounds, however, is far from trivial because of the many spurious edges that bedevil typical edge-detectors. We overcome this difficulty by proposing a method for handling multiple hypotheses for potential edge-locations that is similar in speed to approaches that consider only single hypotheses and therefore much faster than conventional multiple-hypothesis ones. This results in a real-time 3-D tracking algorithm that exploits both texture and edge information without being sensitive to misleading background information and that does not drift over time. Luca Vacchetti, Vincent Lepetit, Pascal Fua |
ISMAR | 3 |
| 2004 | Style-Based Motion SynthesisabstractAbstract Representing motions as linear sums of principal components has become a widely accepted animation technique. While powerful, the simplest version of this approach is not particularly well suited to modeling the specific style of an individual whose motion had not yet been recorded when building the database: it would take an expert to adjust the PCA weights to obtain a motion style that is indistinguishable from his. Consequently, when realism is required, the current practice is to perform a full motion capture session each time a new person must be considered. In this paper, we extend the PCA approach so that this requirement can be drastically reduced: for whole classes of cyclic and noncyclic motions such as walking, running or jumping, it is enough to observe the newcomer moving only once at a particular speed or jumping a particular distance using either an optical motion capture system or a simple pair of synchronized video cameras. This one observation is used to compute a set of principal component weights that best approximates the motion and to extrapolate in real‐time realistic animations of the same person walking or running at different speeds, and jumping a different distance. Raquel Urtasun, Pascal Glardon, Ronan Boulic, Daniel Thalmann, Pascal Fua |
Comput. Graph. Forum | 5 |
| 2004 | Stable Real-Time 3D Tracking Using Online and Offline InformationabstractWe propose an efficient real-time solution for tracking rigid objects in 3D using a single camera that can handle large camera displacements, drastic aspect changes, and partial occlusions. While commercial products are already available for offline camera registration, robust online tracking remains an open issue because many real-time algorithms described in the literature still lack robustness and are prone to drift and jitter. To address these problems, we have formulated the tracking problem in terms of local bundle adjustment and have developed a method for establishing image correspondences that can equally well handle short and wide-baseline matching. We then can merge the information from preceding frames with that provided by a very limited number of keyframes created during a training stage, which results in a real-time tracker that does not jitter or drift and can deal with significant aspect changes. Luca Vacchetti, Vincent Lepetit, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | Visual Golf Club Tracking for Enhanced Swing AnalysisabstractCVLAB Nicolas Gehrig, Vincent Lepetit, Pascal Fua |
BMVC | 3 |
| 2003 | Implicit Meshes for Modeling and ReconstructionabstractExplicit surfaces, such as triangulations or wireframe models, have been extensively used to represent the deformable 3D models that are used to fit 3D point and 2D silhouette data. The resulting approaches, however, suffer from the fact that fitting typically involves finding the facets that are closest to the 3D data points or most likely to be silhouette facets. This requires searching, which is slow, and dealing with the non-differentiability of the distance function. By contrast, implicit surface representations allow fitting without search, since one can simply evaluate a differentiable field function at every data point. However, implicit representations are not necessarily the most intuitive ones and users, such as graphics designers, tend to prefer explicit models. Slobodan Ilic, Pascal Fua |
CVPR (2) | 2 |
| 2003 | Robust Data Association For Online ApplicationsabstractWe present a method for performing data association that handles complex motion models while increasing the robustness of tracking and being suitable for real-time applications. Instead of using motion model in standard recursive fashion, we robustly fit it over multiple frames simultaneously. This allows us to naturally handle arbitrarily complex motion models, to automate the initialization and to deal with occlusion and false alarms. This is effective even if the motion model is not entirely accurate and if there are frequent false-negatives and false-positives. Our algorithm is easy to implement and we show its performances on two real examples of complex motion tracking. Vincent Lepetit, Ali Shahrokni, Pascal Fua |
CVPR (1) | 3 |
| 2003 | Fusing Online and Offline Information for Stable 3D Tracking in Real-TimeabstractWe propose an efficient online real-time solution for single-camera 3D tracking of rigid objects that can handle large camera displacements, drastic aspect changes, and partial occlusions. While the offline camera registration problem can be considered as essentially solved, robust online tracking remains an open issue because many real-time algorithms described in the literature still lack robustness and are prone to drift and jitter. To solve these problems, we have developed a robust approach to 3D feature matching that can handle wide-baseline matching: our method merges the information from preceding frames in traditional recursive tracking fashion with that provided by a very limited number of keyframes created during an offline stage. This combination results in a system that does not suffer from the above difficulties and can deal with drastic aspect changes. We use augmented reality applications to demonstrate its behavior because they are particularly demanding in terms of tracking performance. Luca Vacchetti, Vincent Lepetit, Pascal Fua |
CVPR (2) | 3 |
| 2003 | Robust tracking and segmentation of human motion in an image sequenceabstractWe present a method for improving robustness in feature-based tracking of human motion. Motion flows of features estimated by a standard tracker are modified to be coherent with neighboring ones. This coherence constraint is computed based on a smooth approximation to initial motion flows computed by the tracker. With these tracking results, we demonstrate motion segmentation of different body parts in an image sequence. Jose Juarez Gonzalez, Ik Soo Lim, Pascal Fua, Daniel Thalmann |
ICASSP (3) | 3 |
| 2003 | Fully Automated and Stable Registration for Augmented Reality ApplicationsabstractWe present a fully automated approach to camera registration for augmented reality systems. It relies on purely passive vision techniques to solve the initialization and real-time tracking problems, given a rough CAD model of parts of the real scene. It does not require a controlled environment, for example placing markers. It handles arbitrarily complex models, occlusions, large camera displacements and drastic aspect changes. This is made possible by two major contributions: the first one is a fast recognition method that detects the known part of the scene, registers the camera with respect to it, and initializes a real-time tracker, which is the second contribution. Our tracker eliminates drift and jitter by merging the information from preceding frames in a traditional recursive tracking fashion with that of a very limited number of key-frames created off-line. In the rare instances where it fails, for example because of large occlusion, it detects the failure and reinvokes the initialization procedure. We present experimental results on several different kinds of objects and scenes. Vincent Lepetit, Luca Vacchetti, Daniel Thalmann, Pascal Fua |
ISMAR | 4 |
| 2003 | Real-Time Augmented FaceabstractThis real-time augmented reality demonstration relies on our tracking algorithm described in V. Lepetit et al (2003). This algorithm considers natural feature points, and then does not require engineering of the environment. It merges the information from preceding frames in traditional recursive tracking fashion with that provided by a very limited number of reference frames. This combination results in a system that does not suffer from jitter and drift, and can deal with drastic changes. The tracker recovers the full 3D pose of the tracked object, allowing insertion of 3D virtual objects for augmented reality applications. Vincent Lepetit, Luca Vacchetti, Daniel Thalmann, Pascal Fua |
ISMAR | 4 |
| 2003 | Self-Consistency and MDL: A Paradigm for Evaluating Point-Correspondence Algorithms, and Its Application to Detecting Changes in Surface Elevation
Yvan G. Leclerc, Quang-Tuan Luong, Pascal Fua |
Int. J. Comput. Vis. | 3 |
| 2003 | Articulated Soft Objects for Multiview Shape and Motion CaptureabstractWe develop a framework for 3D shape and motion recovery of articulated deformable objects. We propose a formalism that incorporates the use of implicit surfaces into earlier robotics approaches that were designed to handle articulated structures. We demonstrate its effectiveness for human body modeling from synchronized video sequences. Our method is both robust and generic. It could easily be applied to other shape and motion recovery problems. Ralf Plänkers, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Polyhedral Object Detection and Pose Estimation for Augmented Reality ApplicationsabstractIn augmented reality applications, tracking and registration of both cameras and objects is required because, to combine real and rendered scenes, we must project synthetic models at the right location in real images. Although much work has been done to track objects of interest, initialization of theses trackers often remains manual. Our work aims at automating this step by integrating object recognition and tracking into an AR system. Our emphasis is on the initialization phase of the tracking. We address all the three major aspects of the problem of model-to-image registration: feature detection, correspondence and pose estimation. We have developed a novel approach based on facet detection that greatly reduces the number of possible feature correspondences making it possible to directly compute the transformation which best maps 3-D object to the image plane. We will argue that this approach offers a one-fold speed-up over existing methods. Results of our AR system which integrates initialization and tracking are shown. Our method takes about 5 seconds on our example images. Ali Shahrokni, Luca Vacchetti, Vincent Lepetit, Pascal Fua |
CA | 4 |
| 2002 | Using Dirichlet Free Form Deformation to Fit Deformable Models to Noisy 3-D Data
Slobodan Ilic, Pascal Fua |
ECCV (2) | 2 |
| 2002 | Recovery of Reflectances and Varying Illuminants from Multiple Views
Quang-Tuan Luong, Pascal Fua, Yvan G. Leclerc |
ECCV (3) | 2 |
| 2002 | Model-Based Silhouette Extraction for Accurate People Tracking
Ralf Plänkers, Pascal Fua |
ECCV (2) | 2 |
| 2002 | The Radiometry of Multiple ImagesabstractWe introduce a methodology for radiometric reconstruction (i.e. the simultaneous recovery of multiple illuminants and surface albedoes from multiple views), assuming that the geometry of the scene and of the cameras is known. We formulate a linear theory of multiple illuminants and show its similarity to the theory of geometric recovery of multiple views. Linear and nonlinear implementations are proposed, simulation results are discussed and, finally, results on real images are presented. Quang-Tuan Luong, Pascal Fua, Yvan G. Leclerc |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Incorporating Differential Constraints in a 3D Reconstruction Process. Application to Stereo
Richard Lengagne, Pascal Fua |
ICCV | 2 |
| 2001 | Articulated Soft Objects for Video-based Body Modeling
Ralf Plänkers, Pascal Fua |
ICCV | 2 |
| 2001 | Modeling People Toward Vision-Based Understanding of a Person's Shape, Appearance, and Movement
Adrian Hilton 0001, Pascal Fua |
Comput. Vis. Image Underst. | 2 |
| 2001 | Tracking and Modeling People in Video Sequences
Ralf Plänkers, Pascal Fua |
Comput. Vis. Image Underst. | 2 |
| 2000 | Skeleton-based Motion Capture for Robust Reconstruction of Human MotionabstractOptical motion capture provides an impressive ability to replicate gestures. However, even with a highly professional system there are many instances where crucial markers are occluded or when the algorithm confuses the trajectory of one marker with that of another. This requires much editing work on the part of the animator before the virtual characters are ready for their screen debuts. In this paper, we present an approach to increasing the robustness of a motion capture system by using a sophisticated anatomic human model. It includes a precise description of the skeleton's mobility and an approximated envelope. It allows us to accurately predict the 3-D location and visibility of markers, thus significantly increasing the robustness of marker tracking and assignment, and drastically reducing-or even eliminating-the need for human intervention during the 3D reconstruction process. Lorna Herda, Pascal Fua, Ralf Plänkers, Ronan Boulic, Daniel Thalmann |
CA | 2 |
| 2000 | Augmented Reality for Real and Virtual HumansabstractCurrent virtual reality technologies provide many ways to interact with virtual humans. Most of those techniques, however, are limited to synthetic elements and require cumbersome sensors. We have combined a real-time simulation and rendering platform with a real-time, non-invasive vision-based recognition system to investigate interactions in a mixed environment with real and synthetic elements. In this paper, we present the resulting system, the example of a checkers game between a real person and an autonomous virtual human to demonstrate its performance. Selim Balcisoy, Rémy Torre, Michal Ponder, Pascal Fua, Daniel Thalmann |
Computer Graphics International | 4 |
| 2000 | Detecting Changes in 3-D Shape Using Self-ConsistencyabstractA method for reliably detecting change in the 3-D shape of objects that are well-modeled as single-value functions Z=f(x,y) is presented. It uses an estimate of the accuracy of the 3-D models derived from a set of images taken simultaneously. This accuracy estimate is used to distinguish between significant and insignificant changes in 3-D models derived from different image sets. The accuracy of the 3-D model is estimated using a general methodology, called self-consistency, for estimating the accuracy of computer vision algorithms, which does not require prior establishment of "ground truth". A novel image-matching measure based on Minimum Description Length (MDL) theory allows us to estimate the accuracy of individual elements of the 3-D model. Experiments to demonstrate the utility of the procedure are presented. Yvan G. Leclerc, Quang-Tuan Luong, Pascal Fua, Koji Miyajima |
CVPR | 3 |
| 2000 | Variable Albedo Surface Reconstruction from Stereo and Shape from ShadingabstractWe present a multiview method for the computation of object shape and reflectance characteristics based on the integration of shape from shading (SFS) and stereo, for nonconstant albedo and non-uniformly Lambertian surfaces. First we perform stereo fitting on the input stereo pairs or image sequences. When the images are uncalibrated, we recover the camera parameters using bundle adjustment. Eased on the stereo result, we can automatically segment the albedo map (which is taken to be piece-wise constant) using a minimum description length (MDL) based metric, to identify areas suitable for SFS (typically smooth textureless areas) and to derive illumination information. The shape and the illumination parameter estimates are refined using a deformable model SFS algorithm, which iterates between computing shape and illumination parameters. Our method takes into account the viewing angle dependent for shortening and specularity effects, and compensates as much as possible by utilizing information from more than one images. We demonstrate that we can extend the applicability of SFS algorithms to real world situations when some of its traditional assumptions are violated. We demonstrate our method by applying it to face shape reconstruction. Experimental results indicate a significant improvement over SFS-only or stereo-only based reconstruction. Model accuracy and detail are improved, especially in areas of low texture detail. Albedo information is retrieved and can be used to accurately re-render the model under different illumination conditions. Dimitris Samaras, Dimitris N. Metaxas, Pascal Fua, Yvan G. Leclerc |
CVPR | 3 |
| 2000 | Measuring the Self-Consistency of Stereo Algorithms
Yvan G. Leclerc, Quang-Tuan Luong, Pascal Fua |
ECCV (1) | 3 |
| 2000 | A framework for rapid evaluation of prototypes with augmented realityabstractIn this paper we present a new framework in Augmented Reality context for rapid evaluation of prototypes before manufacture. The design of such prototypes is a time consuming process, leading t o the need of previous evaluation in realistic interactive environments. We have extended the definition of modelling object geometry with modelling object behaviour being able to evaluate them in a mixed environment. Such enhancements allow the development of tools and methods to test object behaviour, and perform interactions between real virtual humans and complex real and virtual objects.We propose a framework for testing the design of objects i an augmented reality context, where a virtual human is able to perform evaluation tests with an object composed of rea and virtual components. In this paper our framework is described and a case study is presented. Selim Balcisoy, Marcelo Kallmann, Pascal Fua, Daniel Thalmann |
VRST | 3 |
| 2000 | Regularized Bundle-Adjustment to Model Heads from Image Sequences without Calibration Data
Pascal Fua |
Int. J. Comput. Vis. | 1 |
| 2000 | 3D stereo reconstruction of human faces driven by differential constraints
Richard Lengagne, Pascal Fua, Olivier Monga |
Image Vis. Comput. | 2 |
| 1999 | From Synthesis to Analysis: Fitting Human Animation Models to Image DataabstractWe show that we can effectively fit complex animation models to noisy image data. Our approach is based on robust least squares adjustment and takes advantage of three complementary sources of information: stereo data, silhouette edges and 2D feature points. We take stereo to be our main information source and use the other two whenever available. In this way, complete head models-including ears and hair-can be acquired with a cheap and entirely passive sensor, such as an ordinary video camera. The motion parameters of limbs can be similarly captured. They can then be fed to existing animation software to produce synthetic sequences. Pascal Fua, Ralf Plänkers, Daniel Thalmann |
Computer Graphics International | 1 |
| 1999 | Using Model-Driven Bundle-Adjustment to Model Heads from Raw Video SequencesabstractWe show that we can effectively and automatically fit a complex facial animation model to uncalibrated image sequences. Our approach is based on model-driven bundle adjustment followed by least squares fitting. It takes advantage of three complementary sources of information: stereo data, silhouette edges and 2D feature points. In this way, complete head models can be acquired with a cheap and entirely passive sensor such as an ordinary video camera. They can then be fed to existing animation software to produce synthetic sequences. Pascal Fua |
ICCV | 1 |
| 1999 | Animated Heads from Ordinary Images: A Least-Squares Approach
Pascal Fua, C. Miccio |
Comput. Vis. Image Underst. | 1 |
| 1998 | From Regular Images to Animated Heads: A Least Squares Approach
Pascal Fua, C. Miccio |
ECCV (1) | 1 |
| 1998 | 3D Face Modeling from Stereo and Differential Constraints
Richard Lengagne, Pascal Fua, Olivier Monga |
FG | 2 |
| 1998 | Using differential constraints to generate a 3D face model from stereoabstractWe propose a way to incorporate a priori information in a 3D stereo reconstruction process from a pair of calibrated face images. A 3D mesh modeling the surface is iteratively deformed in order to minimize an energy function. Differential information about the object shape is used to generate an adaptive mesh that can fulfil the compacity and the accuracy requirements. Moreover in areas where the stereo information is not reliable enough to accurately recover the surface shape, because of inappropriate texture or bad lighting conditions, we incorporate geometric constraints related to the differential properties of the surface, that can be intuitive or refer to predefined geometric properties of the object to be reconstructed. They can be applied to scalar fields, such as curvature values, or structural features, such as crest lines. Therefore, we generate a 3D face model using computer vision techniques that is compact, accurate and consistent with the a priori knowledge about the underlying surface. Richard Lengagne, Pascal Fua, Olivier Monga |
ICPR | 2 |
| 1998 | Fast, Accurate and Consistent Modeling of Drainage and Surrounding Terrain
Pascal Fua |
Int. J. Comput. Vis. | 1 |
| 1997 | Model-Based approach to Accurate and Consistent 3-D Modeling of Drainage and Surrounding TerrainabstractWe propose an automated approach to modeling drainage channels-and more generally, linear features that lie on the terrain-from multiple images, which results not only in high-resolution accurate and consistent models of the features, but also of the surrounding terrain. In our specific case, we have chosen to exploit the fact that rivers flow downhill and lie at the bottom of local depressions in the terrain, valley floors tend to be "U" shaped, and the drainage pattern appears as a network of linear features that can be visually detected in single gray-level images. Different approaches have explored individual facets of this problem. Ours unifies these elements in a common framework. We accurately model terrain and features as 3-dimensional objects from several information sources that may be in error and inconsistent with one another. This approach allows us to generate models that are faithful to sensor data, internally consistent and consistent with physical constraints. Pascal Fua |
CVPR | 1 |
| 1997 | Using Differential Constraints to Reconstruct Complex Surfaces from StereoabstractStereo reconstruction algorithms often fail to properly deal with complex surfaces, because there is not enough image information. To overcome this problem, we propose to guide the reconstruction process using a priori information about the differential geometry of the object surfaces. We use both linear structures such as crest lines or scalar fields such as curvature values to generate a reconstruction of the surface which is consistent with the differential properties. This method improves the accuracy of the reconstruction around the discontinuities and increases the compactness of the surface representation. Richard Lengagne, Olivier Monga, Pascal Fua |
CVPR | 3 |
| 1997 | Imposing Hard Constraints on Deformable Models through Optimization in Orthogonal Subspaces
Pascal Fua, Christian Brechbühler |
Comput. Vis. Image Underst. | 1 |
| 1997 | Velcro Surfaces: Fast Initialization of Deformable Models
Walter M. Neuenschwander, Pascal Fua, Gábor Székely, Olaf Kübler |
Comput. Vis. Image Underst. | 2 |
| 1997 | From Multiple Stereo Views to Multiple 3-D Surfaces
Pascal Fua |
Int. J. Comput. Vis. | 1 |
| 1997 | Ziplock Snakes
Walter M. Neuenschwander, Pascal Fua, Lee Iverson, Gábor Székely, Olaf Kübler |
Int. J. Comput. Vis. | 2 |
| 1996 | Automatic Extraction of Generic House Roofs from High Resolution Aerial Imagery
Frank Bignone, Olof Henricsson, Pascal Fua, Markus A. Stricker |
ECCV (1) | 3 |
| 1996 | Imposing Hard Constraints on Soft Snakes
Pascal Fua, Christian Brechbühler |
ECCV (2) | 1 |
| 1996 | Using crest lines to guide surface reconstruction from stereoabstractWe propose an approach to interleave surface reconstruction from multiple images and feature extraction. It uses an object-centered representation that is optimized to conform to the surface shape. It also extracts typical features such as crest lines and uses them to guide the optimization and the surface reconstruction. We present results using aerial images and terrain data. Richard Lengagne, Olivier Monga, Pascal Fua |
ICIP (2) | 3 |
| 1996 | Using crest lines to guide surface reconstruction from stereoabstractWe propose an approach to interleave surface reconstruction from multiple images and feature extraction. It uses an object-centered representation that is optimized to conform to the surface shape. It also extracts typical features such as crest lines and uses them to guide the optimization. We present results using aerial images and terrain data. Richard Lengagne, Pascal Fua, Olivier Monga |
ICPR | 2 |
| 1996 | Taking Advantage of Image-Based and Geometry-Based Constraints to Recover 3-D Surfaces
Pascal Fua, Yvan G. Leclerc |
Comput. Vis. Image Underst. | 1 |
| 1995 | Reconstructing Complex Surfaces from Multiple Stereo ViewsabstractWe present a framework for 3D surface reconstruction that can be used to model fully 3 dimensional scenes from an arbitrary number of stereo views. Taken from vastly different viewpoints. This is a key step toward producing 3D world descriptions of complex scenes using stereo and is a very challenging problem: real world scenes tend to contain many 3D objects, they do not usually conform to the 2-1/2D assumption made by traditional algorithms, and one cannot take it for granted that the computed 3D points can easily be clustered into separate groups. By combining a particle based representation, robust fitting, and optimization of an image based objective function, we have been able to reconstruct surfaces without any a priori knowledge of their topology and in spite of the noisiness of the stereo data. Our current implementation goes through three steps-initializing a set of particles from the input 3D data, optimizing their location, and finally grouping them into global surfaces. Using several complex scenes containing multiple objects, we demonstrate its competence and ability to merge information and thus to go beyond what can be done with conventional stereo alone.> Pascal Fua |
ICCV | 1 |
| 1995 | Deformable Velcro(tm) SurfacesabstractWe present a new approach to segmentation of 3-D shapes that initializes and then optimizes a 3-D surface model given only the data and a very small number of 3-D seed points and corresponding surface normals. This is a valuable capability for medical, robotic and cartographic applications where such seed points can be naturally supplied. In effect, the surface model is clamped onto the object boundary in manner reminiscent of a Velcro being closed. We develop the method's mathematic framework and show preliminary results using volumetric medical data.> Walter M. Neuenschwander, Pascal Fua, Gábor Székely, Olaf Kübler |
ICCV | 2 |
| 1995 | Object-centered surface reconstruction: Combining multi-image stereo and shading
Pascal Fua, Yvan G. Leclerc |
Int. J. Comput. Vis. | 1 |
| 1994 | Registration without correspondencesabstractWe present a method for registering images of complex 3D surfaces that does not require explicit correspondences between features across the images. Our method relies on the use of a full 3D model of the surface to adjust the position and orientation of the camera by minimizing an objective function based on the projections of the images onto the model. This approach constrains the camera parameters strongly enough so that the models do not need, initially, to be accurate to yield good results. When registration has been achieved, the models can be refined and the fine details recovered. Our method is applicable to the calibration of stereo imagery, the precise registration of new images of a scene and the tracking of deformable objects. It can therefore lead to important applications in fields such as augmented reality in a medical context or data compression for transmission purposes. We demonstrate its applicability by using images of faces and of terrain.> Pascal Fua, Yvan G. Leclerc |
CVPR | 1 |
| 1994 | Initializing snakes [object delineation]abstractWe propose a snake-based approach that lets a user specify only the distant end points of the curve he wishes to delineate without having to supply an almost complete polygonal approximation. We achieve much better convergence properties than those of traditional snakes by using the image information around these end points to provide boundary conditions and by introducing an optimization schedule that allows the snake to take image information into account first only near its extremities and then, progressively, towards its center. These snakes could be used to alleviate the often repetitive task practitioners have to face when segmenting images by abolishing the need to sketch a feature of interest in its entirety, that is, to perform a painstaking, almost complete, manual segmentation.> Walter M. Neuenschwander, Pascal Fua, Gábor Székely, Olaf Kübler |
CVPR | 2 |
| 1994 | Using 3-Dimensional Meshes To Combine Image-Based and Geometry-Based Constraints
Pascal Fua, Yvan G. Leclerc |
ECCV (2) | 1 |
| 1994 | Using contextual information to set control parameters of a vision processabstractThis paper presents a novel approach to supervised learning of image-understanding tactics based on the use of contextual information. We use a database to store past experiences. From this database and the context elements computed from the task specifications and the input data, we determine whether or not an algorithm is applicable, and which parameters are suitable for it. The database is continuously updated with information of success or failure of the system. Stéphane Houzelle, Thomas M. Strat, Pascal Fua, Martin A. Fischler |
ICPR (1) | 3 |
| 1994 | Making snakes converge from minimal initializationabstractIn this paper, we present a new snake-based method to delineate contours where the user has to specify only the two distant end points without having to supply an almost complete polygonal approximation. We achieve much better convergence properties than those of traditional snakes by propagating the image information along the curve from both end points towards its center. We use our new method to outline curved object boundaries and for interactive road delineation. Walter M. Neuenschwander, Pascal Fua, Gábor Székely, Olaf Kübler |
ICPR (1) | 2 |
| 1992 | Segmenting Unstructured 3D Points into Surfaces
Pascal Fua, Peter T. Sander |
ECCV | 1 |
| 1992 | Autonomous planetary rover (VAP): on-board perception system concept and stereovision by correlation approachabstractThe authors present some of the work currently being conducted at CEA, CNRS, INRIA, ONERA, and CNES in the framework of the Autonomous Planetary Rover program (VAP). A system concept approach for the VAP application is first presented. Attention is then given to the VAP perception system: some results obtained with stereovision by a correlation algorithm on outdoor simulated Mars terrain scenes are shown.> L. Boissier, Bernard Hotz, Catherine Proy, Olivier D. Faugeras, Pascal Fua |
ICRA | 5 |
| 1992 | A parallel stereo algorithm that produces dense depth maps and preserves image features
Pascal Fua |
Mach. Vis. Appl. | 1 |
| 1991 | Combining Stereo and Monocular Information to Compute Dense Depth Maps that Preserve Depth Discontinuities
Pascal Fua |
IJCAI | 1 |
| 1991 | An optimization framework for feature extraction
Pascal Fua, Andrew J. Hanson |
Mach. Vis. Appl. | 1 |
| 1989 | Objective Functions for Feature Discrimination
Pascal Fua, Andrew J. Hanson |
IJCAI | 1 |
| 1989 | Model driven edge detection
Pascal Fua, Yvan G. Leclerc |
Mach. Vis. Appl. | 1 |
| 1987 | Using Generic Geometric Models for Intelligent Shape Extraction
Pascal Fua, Andrew J. Hanson |
AAAI | 1 |
| 1987 | Resegmentation using generic shape: Locating general cultural objects
Pascal Fua, Andrew J. Hanson |
Pattern Recognit. Lett. | 1 |
| 1986 | Using probability-density functions in the framework of evidential reasoning
Pascal Fua |
IPMU | 1 |