EDBT 2026 Demo / reviewers in the wild / expert
Mathieu Salzmann
dblp:18/4533
· DBLP profile ↗
225ranked-venue papers
13as first author
91since 2021 · last 2026
0000-0002-8347-8637ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 198 · 13 first-author · 81 since 2021Graphics, computer vision, multimedia, augmented reality and games · 156 · 10 first-author · 57 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FastPose-ViT: A Vision Transformer for Real-Time Spacecraft Pose EstimationabstractEstimating the 6-degrees-of-freedom (6DoF) pose of a spacecraft from a single image is critical for autonomous operations like in-orbit servicing and space debris removal. Existing state-of-the-art methods often rely on iterative Perspective-n-Point (PnP)-based algorithms, which are computationally intensive and ill-suited for real-time deployment on resource-constrained edge devices. To overcome these limitations, we propose FastPose-ViT, a Vision Transformer (ViT)-based architecture that directly regresses the 6DoF pose. Our approach processes cropped images from object bounding boxes and introduces a novel mathematical formalism to map these localized predictions back to the full-image scale. This formalism is derived from the principles of projective geometry and the concept of "apparent rotation", where the model predicts an apparent rotation matrix that is then corrected to find the true orientation. We demonstrate that our method outperforms other non-PnP strategies and achieves performance competitive with state-of-the-art PnP-based techniques on the SPEED dataset. Furthermore, we validate our model’s suitability for real-world space missions by quantizing it and deploying it on power-constrained edge hardware. On the NVIDIA Jetson Orin Nano, our end-to-end pipeline achieves a latency of 75 ms per frame under sequential execution, and a non-blocking throughput of up to 33 FPS when stages are scheduled concurrently. Pierre Ancey, Andrew Lawrence Price, Saqib Javed, Mathieu Salzmann |
WACV | 4 |
| 2026 | Pose without Guesses: Generalizable Object Pose Estimation From a Single Reference
Chen Zhao 0025, Tong Zhang 0023, Zheng Dang, Mathieu Salzmann |
Int. J. Comput. Vis. | 4 |
| 2025 | Free-Moving Object Reconstruction and Pose Estimation with Virtual CameraabstractWe propose an approach for reconstructing free-moving object from a monocular RGB video. Most existing methods either assume scene prior, hand pose prior, object category pose prior, or rely on local optimization with multiple sequence segments. We propose a method that allows free interaction with the object in front of a moving camera without relying on any prior, and optimizes the sequence globally without any segments. We progressively optimize the object shape and pose simultaneously based on an implicit neural representation. A key aspect of our method is a virtual camera system that reduces the search space of the optimization significantly. We evaluate our method on the standard HO3D dataset and a collection of egocentric RGB sequences captured with a head-mounted device. We demonstrate that our approach outperforms most methods significantly, and is on par with recent techniques that assume prior information. Haixin Shi, Yinlin Hu, Daniel Koguciuk, Juan-Ting Lin, Mathieu Salzmann, David Ferstl |
AAAI | 5 |
| 2025 | GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition ControlabstractWe present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth outputs for richer spatial understanding. We introduce autoregressive noise schedules to enable stable long-horizon generations. Our dataset is comprised of 4000+ hours of multimodal data across domains like autonomous driving, egocentric human activities, and drone flights. Pseudo-labels are used to get depth maps, ego-trajectories, and human poses. We use a comprehensive evaluation framework, including a new Control of Object Manipulation (COM) metric, to assess controllability. Experiments show GEM excels at generating diverse, controllable scenarios and temporal consistency over long generations. Code, models, and datasets are fully open-sourced1. Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, Pedro M. B. Rezende, Yasaman Haghighi, David Brüggemann, Isinsu Katircioglu, Xiaoran Chen, Marco Cannici, Elie Aljalbout, Botao Ye, Xi Wang 0021, Aram Davtyan, Mathieu Salzmann, Davide Scaramuzza 0001, Marc Pollefeys, Paolo Favaro, Alexandre Alahi |
CVPR | 16 |
| 2025 | MotionMap: Representing Multimodality in Human Pose ForecastingabstractHuman pose forecasting is inherently multimodal since multiple futures exist for an observed pose sequence. However, evaluating multimodality is challenging since the task is ill-posed. Therefore, we first propose an alternative paradigm to make the task well-posed. Next, while state-of-the-art methods predict multimodality, this requires oversampling a large volume of predictions. This raises key questions: (1) Can we capture multimodality by efficiently sampling fewer predictions? (2) Subsequently, which of the predicted futures is more likely for an observed pose sequence? We address these questions with MotionMap, a simple yet effective heatmap based representation for multimodality. We extend heatmaps to represent a spatial distribution over the space of all possible motions, where different local maxima correspond to different forecasts for a given observation. MotionMap can capture a variable number of modes per observation and provide confidence measures for different modes. Further, MotionMap allows us to introduce the notion of uncertainty and controllability over the forecasted pose sequence. Finally, MotionMap captures rare modes that are non-trivial to evaluate yet critical for safety. We support our claims through multiple qualitative and quantitative experiments using popular 3D human pose datasets: Human3.6M and AMASS, highlighting the strengths and limitations of our proposed method. https://vita-epfl.github.io/MotionMap Reyhaneh HosseiniNejad, Megh Shukla, Saeed Saadatnejad, Mathieu Salzmann, Alexandre Alahi |
CVPR | 4 |
| 2025 | Self-Ensembling Gaussian Splatting for Few-Shot Novel View Synthesisabstract3D Gaussian Splatting (3DGS) has demonstrated remarkable effectiveness in novel view synthesis (NVS). However, 3DGS tends to overfit when trained with sparse views, limiting its generalization to novel viewpoints. In this paper, we address this overfitting issue by introducing Self-Ensembling Gaussian Splatting (SE-GS). We achieve self-ensembling by incorporating an uncertainty-aware perturbation strategy during training. A $\mathbfΔ$-model and a $\mathbfΣ$-model are jointly trained on the available images. The $\mathbfΔ$-model is dynamically perturbed based on rendering uncertainty across training steps, generating diverse perturbed models with negligible computational overhead. Discrepancies between the $\mathbfΣ$-model and these perturbed models are minimized throughout training, forming a robust ensemble of 3DGS models. This ensemble, represented by the $\mathbfΣ$-model, is then used to generate novel-view images during inference. Experimental results on the LLFF, Mip-NeRF360, DTU, and MVImgNet datasets demonstrate that our approach enhances NVS quality under few-shot training conditions, outperforming existing state-of-the-art methods. The code is released at: https://sailor-z.github.io/projects/SEGS.html. Chen Zhao 0025, Xuan Wang 0009, Tong Zhang 0023, Saqib Javed, Mathieu Salzmann |
ICCV | 5 |
| 2025 | Enhancing Compositional Text-to-Image Generation with Reliable Random SeedsabstractText-to-image diffusion models have demonstrated remarkable capability in generating realistic images from arbitrary text prompts. However, they often produce inconsistent results for compositional prompts such as "two dogs" or "a penguin on the right of a bowl". Understanding these inconsistencies is crucial for reliable image generation. In this paper, we highlight the significant role of initial noise in these inconsistencies, where certain noise patterns are more reliable for compositional prompts than others. Our analyses reveal that different initial random seeds tend to guide the model to place objects in distinct image areas, potentially adhering to specific patterns of camera angles and image composition associated with the seed. To improve the model's compositional ability, we propose a method for mining these reliable cases, resulting in a curated training set of generated images without requiring any manual annotation.
By fine-tuning text-to-image models on these generated images, we significantly enhance their compositional capabilities. For numerical composition, we observe relative increases of 29.3\% and 19.5\% for Stable Diffusion and PixArt-$\alpha$, respectively. Spatial composition sees even larger gains, with 60.7\% for Stable Diffusion and 21.1\% for PixArt-$\alpha$. Shuangqi Li, Hieu Le 0001, Mathieu Salzmann |
ICLR | 4 |
| 2025 | Towards Self-Supervised Covariance Estimation in Deep Heteroscedastic RegressionabstractDeep heteroscedastic regression models the mean and covariance of the target distribution through neural networks. The challenge arises from heteroscedasticity, which implies that the covariance is sample dependent and is often unknown. Consequently, recent methods learn the covariance through unsupervised frameworks, which unfortunately yield a trade-off between computational complexity and accuracy. While this trade-off could be alleviated through supervision, obtaining labels for the covariance is non-trivial.
Here, we study self-supervised covariance estimation in deep heteroscedastic regression. We address two questions: (1) How should we supervise the covariance assuming ground truth is available? (2) How can we obtain pseudo labels in the absence of the ground-truth? We address (1) by analysing two popular measures: the KL Divergence and the 2-Wasserstein distance. Subsequently, we derive an upper bound on the 2-Wasserstein distance between normal distributions with non-commutative covariances that is stable to optimize. We address (2) through a simple neighborhood based heuristic algorithm which results in surprisingly effective pseudo labels for the covariance. Our experiments over a wide range of synthetic and real datasets demonstrate that the proposed 2-Wasserstein bound coupled with pseudo label annotations results in a computationally cheaper yet accurate deep heteroscedastic regression. Megh Shukla, Aziz Shameem, Mathieu Salzmann, Alexandre Alahi |
ICLR | 3 |
| 2025 | QT-DoG: Quantization-Aware Training for Domain GeneralizationabstractA key challenge in Domain Generalization (DG) is preventing overfitting to source domains, which can be mitigated by finding flatter minima in the loss landscape. In this work, we propose Quantization-aware Training for Domain Generalization (QT-DoG) and demonstrate that weight quantization effectively leads to flatter minima in the loss landscape, thereby enhancing domain generalization. Unlike traditional quantization methods focused on model compression, QT-DoG exploits quantization as an implicit regularizer by inducing noise in model weights, guiding the optimization process toward flatter minima that are less sensitive to perturbations and overfitting. We provide both an analytical perspective and empirical evidence demonstrating that quantization inherently encourages flatter minima, leading to better generalization across domains. Moreover, with the benefit of reducing the model size through quantization, we demonstrate that an ensemble of multiple quantized models further yields superior accuracy than the state-of-the-art DG approaches with no computational or memory overheads. Code is released at: https://saqibjaved1.github.io/QT_DoG/. Saqib Javed, Hieu Le 0001, Mathieu Salzmann |
ICML | 3 |
| 2025 | Demystifying Singular Defects in Large Language ModelsabstractLarge transformer models are known to produce high-norm tokens. In vision transformers (ViTs), such tokens have been mathematically modeled through the singular vectors of the linear approximations of layers. However, in large language models (LLMs), the underlying causes of high-norm tokens remain largely unexplored, and their different properties from those of ViTs require a new analysis framework. In this paper, we provide both theoretical insights and empirical validation across a range of recent models, leading to the following observations: i) The layer-wise singular direction predicts the abrupt explosion of token norms in LLMs. ii) The negative eigenvalues of a layer explain its sudden decay. iii) The computational pathways leading to high-norm tokens differ between initial and noninitial tokens. iv) High-norm tokens are triggered by the right leading singular vector of the matrix approximating the corresponding modules. We showcase two practical applications of these findings: the improvement of quantization schemes and the design of LLM signatures. Our findings not only advance the understanding of singular defects in LLMs but also open new avenues for their application. We expect that this work will stimulate further research into the internal mechanisms of LLMs. Code is released at https://github.com/haoqiwang/singular_defect. Tong Zhang 0023, Mathieu Salzmann |
ICML | 3 |
| 2025 | 6Img-to-3D: Few-Image Large-Scale Outdoor Novel View SynthesisabstractCurrent 3D reconstruction techniques struggle to infer unbounded scenes from a few images faithfully. Most existing methods have high computational demands, require detailed pose information, and cannot reconstruct occluded regions reliably. We introduce 6Img-to-3D, a novel transformer-based encoder-renderer method for single-shot image-to-3D reconstruction. Our method outputs a 3D-consistent param-eterized triplane from only six outward-facing input images for large-scale, unbounded outdoor driving scenarios. We take a step towards resolving existing shortcomings by combining contracted custom cross- and self-attention mechanisms for tri-plane parameterization, differentiable volume rendering, scene contraction, and image feature projection. We showcase on synthetic data that six surround-view vehicle images from a single timestamp are enough to reconstruct 360° scenes during inference time, taking 395 ms. Our method allows, for example, rendering third-person images and birds-eye views. Code, and more results are available at https: / /6Img-to-3D. GitHub. io/. Théo Gieruc, Marius Kästingschäfer, Sebastian Bernhard, Mathieu Salzmann |
IV | 4 |
| 2025 | Leveraging Gradient Information for Out-of-Domain Performance Estimations
Ekaterina Khramtsova, Mahsa Baktash, Guido Zuccon, Xi Wang 0021, Mathieu Salzmann |
ECML/PKDD (6) | 5 |
| 2025 | Temporally-Consistent Surface Reconstruction Using Metrically-Consistent AtlasesabstractWe propose a method for unsupervised reconstruction of a temporally-consistent sequence of surfaces from a sequence of time-evolving point clouds. It yields dense and semantically meaningful correspondences between frames. We represent the reconstructed surfaces as atlases computed by a neural network, which enables us to establish correspondences between frames. The key to making these correspondences semantically meaningful is to guarantee that the metric tensors computed at corresponding points are as similar as possible. We have devised an optimization strategy that makes our method robust to noise and global motions, without a priori correspondences or pre-alignment steps. As a result, our approach outperforms state-of-the-art ones on several challenging datasets. Jan Bednarík, Noam Aigerman, Vladimir G. Kim, Siddhartha Chaudhuri, Shaifali Parashar, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Unsupervised 3D Keypoint Discovery with Multi-View GeometryabstractAnalyzing and training 3D body posture models depend heavily on the availability of joint labels that are commonly acquired through laborious manual annotation of body joints or via marker-based joint localization using carefully curated markers and capturing systems. However, such annotations are not always available, especially for people performing unusual activities. In this paper, we propose an algorithm that learns to discover 3D keypoints on human bodies from multiple-view images without any supervision or labels other than the constraints multiple-view geometry provides. To ensure that the discovered 3D keypoints are meaningful, they are re-projected to each view to estimate the person’s mask that the model itself has initially estimated without supervision. Our approach discovers more interpretable and accurate 3D keypoints compared to other state-of-the-art unsupervised approaches on Human3.6M and MPI-INF-3DHP benchmark datasets. Sina Honari, Chen Zhao 0025, Mathieu Salzmann, Pascal Fua |
3DV | 3 |
| 2024 | LocPoseNet: Robust Location Prior for Unseen Object Pose EstimationabstractObject location prior is critical for the standard 6D object pose estimation setting. The prior can be used to initialize the 3D object translation and facilitate 3D object rotation estimation. Unfortunately, the object detectors that are used for this purpose do not generalize to unseen objects. Therefore, existing 6D pose estimation methods for unseen objects either assume the ground-truth object location to be known or yield inaccurate results when it is unavailable. In this paper, we address this problem by developing a method, LocPoseNet, able to robustly learn location prior for unseen objects. Our method builds upon a template matching strategy, where we propose to distribute the reference kernels and convolve them with a query to efficiently compute multi-scale correlations. We then introduce a novel translation estimator, which decouples scale-aware and scale-robust features to predict different object location parameters. Our method outperforms existing works by a large margin on LINEMOD and GenMOP. We further construct a challenging synthetic dataset, which allows us to highlight the better robustness of our method to various noise sources. Our project website is at: https://sailorz.github.io/projects/3DV2024_LocPoseNet.html. Chen Zhao 0025, Yinlin Hu, Mathieu Salzmann |
3DV | 3 |
| 2024 | AttEntropy: On the Generalization Ability of Supervised Semantic Segmentation Transformers to New Objects in New Domains
Krzysztof Lis, Matthias Rottmann, Annika Mütze, Sina Honari, Pascal Fua, Mathieu Salzmann |
BMVC | 6 |
| 2024 | CLOAF: CoLlisiOn-Aware Human FlowabstractEven the best current algorithms for estimating body 3D shape and pose yield results that include body self- intersections. In this paper, we present CLOAF, which exploits the diffeomorphic nature of Ordinary Differential Equations to eliminate such self-intersections while still im- posing body shape constraints. We show that, unlike earlier approaches to addressing this issue, ours completely elim- inates the self-intersections without compromising the ac- curacy of the reconstructions. Being differentiable, CLOAF can be used to fine-tune pose and shape estimation base- lines to improve their overall performance and eliminate self-intersections in their predictions. Furthermore, we demonstrate how our CLOAF strategy can be applied to practically any motion field induced by the user. CLOAF also makes it possible to edit motion to interact with the environment without worrying about potential collision or loss of body-shape prior. Andrey Davydov, Martin Engilberge, Mathieu Salzmann, Pascal Fua |
CVPR | 3 |
| 2024 | NOPE: Novel Object Pose Estimation from a Single ImageabstractThe practicality of 3D object pose estimation remains limited for many applications due to the need for prior knowledge of a 3D model and a training period for new objects. To address this limitation, we propose an approach that takes a single image of a new object as input and pre-dicts the relative pose of this object in new images without prior knowledge of the object's 3D model and without re-quiring training time for new objects and categories. We achieve this by training a model to directly predict discrim-inative embeddings for viewpoints surrounding the object. This prediction is done using a simple U-Net architecture with attention and conditioned on the desired pose, which yields extremely fast inference. We compare our approach to state-of-the-art methods and show it outperforms them both in terms of accuracy and robustness. Van Nguyen Nguyen, Thibault Groueix, Georgy Ponimatkin, Yinlin Hu, Renaud Marlet, Mathieu Salzmann, Vincent Lepetit |
CVPR | 6 |
| 2024 | GigaPose: Fast and Robust Novel Object Pose Estimation via One CorrespondenceabstractWe present GigaPose, afast, robust, and accurate method for CAD-based novel object pose estimation in RGB images. GigaPose first leverages discriminative “templates ”, ren-dered images of the CAD models, to recover the out-of-plane rotation and then uses patch correspondences to estimate the four remaining parameters. Our approach samples tem-plates in only a two-degrees-of-freedom space instead of the usual three and matches the input image to the templates using fast nearest-neighbor search in feature space, results in a speedup factor of 35x compared to the state of the art. More-over, GigaPose is significantly more robust to segmentation errors. Our extensive evaluation on the seven core datasets of the BOP challenge demonstrates that it achieves state-of-the-art accuracy and can be seamlessly integrated with existing refinement methods. Additionally, we show the potential of GigaPose with 3D models predicted by recent work on 3D reconstruction from a single image, relaxing the need for CAD models and making 6D pose object estimation much more convenient. Our source code and trained models are publicly available at https://github.conllnv-nguyenlgigaPose. Van Nguyen Nguyen, Thibault Groueix, Mathieu Salzmann, Vincent Lepetit |
CVPR | 3 |
| 2024 | HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance FieldsabstractHuman hands are highly articulated and versatile at handling objects. Jointly estimating the 3D poses of a hand and the object it manipulates from a monocular camera is challenging due to frequent occlusions. Thus, existing methods often rely on intermediate 3D shape representations to increase performance. These representations are typically explicit, such as 3D point clouds or meshes, and thus provide information in the direct surroundings of the intermediate hand pose estimate. To address this, we in-troduce HOISDF, a Signed Distance Field (SDF) guided hand-object pose estimation network, which jointly exploits hand and object SDFs to provide a global, implicit repre-sentation over the complete reconstruction volume. Specif-ically, the role of the SDFs is threefold: equip the visual encoder with implicit shape information, help to encode hand-object interactions, and guide the hand and object pose regression via SDF-based sampling and by augmenting the feature representations. We show that HOISDF achieves state-of-the-art results on hand-object pose esti-mation benchmarks (DexYCB and H03Dv2). Code is avail-able at https://github.com/amathislabIHOISDF. Haozhe Qi, Chen Zhao 0025, Mathieu Salzmann, Alexander Mathis |
CVPR | 3 |
| 2024 | Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning Through Object ExchangeabstractIn the realm of point cloud scene understanding, particularly in indoor scenes, objects are arranged following human habits, resulting in objects of certain semantics being closely positioned and displaying notable inter-object cor-relations. This can create a tendency for neural networks to exploit these strong dependencies, bypassing the individ-ual object patterns. To address this challenge, we introduce a novel self-supervised learning (SSL) strategy. Our approach leverages both object patterns and contextual cues to produce robust features. It begins with the formulation of an object-exchanging strategy, where pairs of objects with comparable sizes are exchanged across different scenes, effectively disentangling the strong contextual dependencies. Subsequently, we introduce a context-aware feature learning strategy, which encodes object patterns without relying on their specific context by aggregating object features across various scenes. Our extensive experiments demonstrate the superiority of our method over existing SSL techniques, further showing its better robustness to environmental changes. Moreover, we showcase the applicability of our approach by transferring pre-trained models to diverse point cloud datasets.11Our code is available at https:/lgithub.com/YanhaoWu/OESSL Yanhao Wu, Tong Zhang 0023, Wei Ke 0003, Congpei Qiu, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 6 |
| 2024 | DVMNet: Computing Relative Pose for Unseen Objects Beyond HypothesesabstractDetermining the relative pose of an object between two images is pivotal to the success of generalizable object pose estimation. Existing approaches typically approximate the continuous pose representation with a large number of discrete pose hypotheses, which incurs a computationally expensive process of scoring each hypothesis at test time. By contrast, we present a Deep Voxel Matching Network (DVMNet) that eliminates the need for pose hypotheses and computes the relative object pose in a single pass. To this end, we map the two input RGB images, reference and query, to their respective voxelized 3D representations. We then pass the resulting voxels through a pose estimation module, where the voxels are aligned and the pose is computed in an end-to-end fashion by solving a least-squares problem. To enhance robustness, we introduce a weighted closest voxel algorithm capable of mitigating the impact of noisy voxels. We conduct extensive experiments on the CO3D, LINEMOD, and Objaverse datasets, demonstrating that our method delivers more accurate relative pose estimates for novel objects at a lower computational cost compared to state-of-the-art methods. Our code is released at: https://github.com/sailor-z/DVMNet/. Chen Zhao 0025, Tong Zhang 0023, Zheng Dang, Mathieu Salzmann |
CVPR | 4 |
| 2024 | Data Augmentation via Latent Diffusion for Saliency Prediction
Bahar Aydemir, Deblina Bhattacharjee, Tong Zhang 0023, Mathieu Salzmann, Sabine Süsstrunk |
ECCV (78) | 4 |
| 2024 | Source-Free Domain-Invariant Performance Prediction
Ekaterina Khramtsova, Mahsa Baktash, Guido Zuccon, Xi Wang 0021, Mathieu Salzmann |
ECCV (80) | 5 |
| 2024 | SINDER: Repairing the Singular Defects of DINOv2
Tong Zhang 0023, Mathieu Salzmann |
ECCV (7) | 3 |
| 2024 | 3D Single-Object Tracking in Point Clouds with High Temporal Variation
Qiao Wu, Kun Sun 0002, Pei An, Mathieu Salzmann, Yanning Zhang 0001, Jiaqi Yang 0002 |
ECCV (7) | 4 |
| 2024 | Neural SDF Flow for 3D Reconstruction of Dynamic ScenesabstractIn this paper, we tackle the problem of 3D reconstruction of dynamic scenes from multi-view videos. Previous dynamic scene reconstruction works either attempt to model the motion of 3D points in space, which constrains them to handle a single articulated object or require depth maps as input. By contrast, we propose to directly estimate the change of Signed Distance Function (SDF), namely SDF
flow, of the dynamic scene. We show that the SDF flow captures the evolution of the scene surface. We further derive the mathematical relation between the SDF flow and the scene flow, which allows us to calculate the scene flow from the SDF flow analytically by solving linear equations. Our experiments on real-world multi-view video datasets show that our reconstructions are better than those of the state-of-the-art methods. Our code is available at https://github.com/wei-mao-2019/SDFFlow.git. Wei Mao 0001, Richard I. Hartley, Mathieu Salzmann, Miaomiao Liu 0001 |
ICLR | 3 |
| 2024 | Mind Your Augmentation: The Key to Decoupling Dense Self-Supervised LearningabstractDense Self-Supervised Learning (SSL) creates positive pairs by building positive paired regions or points, thereby aiming to preserve local features, for example of individual objects. However, existing approaches tend to couple objects by leaking information from the neighboring contextual regions when the pairs have a limited overlap. In this paper, we first quantitatively identify and confirm the existence of such a coupling phenomenon. We then address it by developing a remarkably simple yet highly effective solution comprising a novel augmentation method, Region Collaborative Cutout (RCC), and a corresponding decoupling branch. Importantly, our design is versatile and can be seamlessly integrated into existing SSL frameworks, whether based on Convolutional Neural Networks (CNNs) or Vision Transformers (ViTs). We conduct extensive experiments, incorporating our solution into two CNN-based and two ViT-based methods, with results confirming the effectiveness of our approach. Moreover, we provide empirical evidence that our method significantly contributes to the disentanglement of feature representations among objects, both in quantitative and qualitative terms. Congpei Qiu, Tong Zhang 0023, Yanhao Wu, Wei Ke 0003, Mathieu Salzmann, Sabine Süsstrunk |
ICLR | 5 |
| 2024 | 3D-Aware Hypothesis & Verification for Generalizable Relative Object Pose EstimationabstractPrior methods that tackle the problem of generalizable object pose estimation highly rely on having dense views of the unseen object. By contrast, we address the scenario where only a single reference view of the object is available. Our goal then is to estimate the relative object pose between this reference view and a query image that depicts the object in a different pose. In this scenario, robust generalization is imperative due to the presence of unseen objects during testing and the large-scale object pose variation between the reference and the query. To this end, we present a new hypothesis-and-verification framework, in which we generate and evaluate multiple pose hypotheses, ultimately selecting the most reliable one as the relative object pose. To measure reliability, we introduce a 3D-aware verification that explicitly applies 3D transformations to the 3D object representations learned from the two input images. Our comprehensive experiments on the Objaverse, LINEMOD, and CO3D datasets evidence the superior accuracy of our approach in relative pose estimation and its robustness in large-scale pose variations, when dealing with unseen objects. Chen Zhao 0025, Tong Zhang 0023, Mathieu Salzmann |
ICLR | 3 |
| 2024 | TIC-TAC: A Framework For Improved Covariance Estimation In Deep Heteroscedastic RegressionabstractDeep heteroscedastic regression involves jointly optimizing the mean and covariance of the predicted distribution using the negative log-likelihood. However, recent works show that this may result in sub-optimal convergence due to the challenges associated with covariance estimation. While the literature addresses this by proposing alternate formulations to mitigate the impact of the predicted covariance, we focus on improving the predicted covariance itself. We study two questions: (1) Does the predicted covariance truly capture the randomness of the predicted mean? (2) In the absence of supervision, how can we quantify the accuracy of covariance estimation? We address (1) with a Taylor Induced Covariance (TIC), which captures the randomness of the predicted mean by incorporating its gradient and curvature through the second order Taylor polynomial. Furthermore, we tackle (2) by introducing a Task Agnostic Correlations (TAC) metric, which combines the notion of correlations and absolute error to evaluate the covariance. We evaluate TIC-TAC across multiple experiments spanning synthetic and real-world datasets. Our results show that not only does TIC accurately learn the covariance, it additionally facilitates an improved convergence of the negative log-likelihood. Our code is available at https://github.com/vita-epfl/TIC-TAC Megh Shukla, Mathieu Salzmann, Alexandre Alahi |
ICML | 2 |
| 2024 | LiDAR-NeRF: Novel LiDAR View Synthesis via Neural Radiance Fields
Tang Tao, Longfei Gao, Guangrun Wang, Yixing Lao, Peng Chen 0054, Hengshuang Zhao, Dayang Hao, Xiaodan Liang, Mathieu Salzmann, Kaicheng Yu |
ACM Multimedia | 9 |
| 2024 | Generalize or Detect? Towards Robust Semantic Segmentation Under Multiple Distribution ShiftsabstractIn open-world scenarios, where both novel classes and domains may exist, an ideal segmentation model should detect anomaly classes for safety and generalize to new domains. However, existing methods often struggle to distinguish between domain-level and semantic-level distribution shifts, leading to poor OOD detection or domain generalization performance. In this work, we aim to equip the model to generalize effectively to covariate-shift regions while precisely identifying semantic-shift regions. To achieve this, we design a novel generative augmentation method to produce coherent images that incorporate both anomaly (or novel) objects and various covariate shifts at both image and object levels. Furthermore, we introduce a training strategy that recalibrates uncertainty specifically for semantic shifts and enhances the feature extractor to align features associated with domain shifts. We validate the effectiveness of our method across benchmarks featuring both semantic and domain shifts. Our method achieves state-of-the-art performance across all benchmarks for both OOD detection and domain generalization. Code is available at https://github.com/gaozhitong/MultiShiftSeg. Zhitong Gao, Mathieu Salzmann, Xuming He 0001 |
NeurIPS | 3 |
| 2024 | On the Impact of Hard Adversarial Instances on Overfitting in Adversarial TrainingabstractAdversarial training is a popular method to robustify models against adversarial attacks. However, it exhibits much more severe overfitting than training on clean inputs. In this work, we investigate this phenomenon from the perspective of training instances, i.e., training input-target pairs. Based on a quantitative metric measuring the relative difficulty of an instance in the training set, we analyze the model's behavior on training instances of different difficulty levels. This lets us demonstrate that the decay in generalization performance of adversarial training is a result of fitting hard adversarial instances. We theoretically verify our observations for both linear and general nonlinear models, proving that models trained on hard instances have worse generalization performance than ones trained on easy instances, and that this generalization gap increases with the size of the adversarial budget. Finally, we investigate solutions to mitigate adversarial overfitting in several scenarios, including fast adversarial training and fine-tuning a pretrained model with additional data. Our results demonstrate that using training data adaptively improves the model's robustness. Chen Liu 0027, Zhichao Huang 0002, Mathieu Salzmann, Tong Zhang 0023, Sabine Süsstrunk |
J. Mach. Learn. Res. | 3 |
| 2024 | Match Normalization: Learning-Based Point Cloud Registration for 6D Object Pose Estimation in the Real WorldabstractIn this work, we tackle the task of estimating the 6D pose of an object from point cloud data. While recent learning-based approaches have shown remarkable success on synthetic datasets, we have observed them to fail in the presence of real-world data. We investigate the root causes of these failures and identify two main challenges: The sensitivity of the widely-used SVD-based loss function to the range of rotation between the two point clouds, and the difference in feature distributions between the source and target point clouds. We address the first challenge by introducing a directly supervised loss function that does not utilize the SVD operation. To tackle the second, we introduce a new normalization strategy, Match Normalization. Our two contributions are general and can be applied to many existing learning-based 3D object registration frameworks, which we illustrate by implementing them in two of them, DCP and IDAM. Our experiments on the real-scene TUD-L Hodan et al. 2018, LINEMOD Hinterstoisser et al. 2012 and Occluded-LINEMOD Brachmann et al. 2014 datasets evidence the benefits of our strategies. They allow for the first-time learning-based 3D object registration methods to achieve meaningful results on real-world data. We therefore expect them to be key to the future developments of point cloud registration methods. Zheng Dang, Lizhou Wang, Yu Guo 0006, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Detecting Road Obstacles by Erasing ThemabstractVehicles can encounter a myriad of obstacles on the road, and it is impossible to record them all beforehand to train a detector. Instead, we select image patches and inpaint them with the surrounding road texture, which tends to remove obstacles from those patches. We then use a network trained to recognize discrepancies between the original patch and the inpainted one, which signals an erased obstacle. Krzysztof Lis, Sina Honari, Pascal Fua, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | A Closed-Form, Pairwise Solution to Local Non-Rigid Structure-From-MotionabstractA recent trend in Non-Rigid Structure-from-Motion (NRSfM) is to express local, differential constraints between pairs of images, from which the surface normal at any point can be obtained by solving a system of polynomial equations. While this approach is more successful than its counterparts relying on global constraints, the resulting methods face two main problems: First, most of the equation systems they formulate are of high degree and must be solved using computationally expensive polynomial solvers. Some methods use polynomial reduction strategies to simplify the system, but this adds some phantom solutions. In any event, an additional mechanism is employed to pick the best solution, which adds to the computation without any guarantees on the reliability of the solution. Second, these methods formulate constraints between a pair of images. Even if there is enough motion between them, they may suffer from local degeneracies that make the resulting estimates unreliable without any warning mechanism. %Unfortunately, these systems are of high degree with up to five real solutions. Hence, a computationally expensive strategy is required to select a unique solution. Furthermore, they suffer from degeneracies that make the resulting estimates unreliable, without any mechanism to identify this situation. In this paper, we solve these problems for isometric/conformal NRSfM. We show that, under widely applicable assumptions, we can derive a new system of equations in terms of the surface normals, whose two solutions can be obtained in closed-form and can easily be disambiguated locally. Our formalism also allows us to assess how reliable the estimated local normals are and to discard them if they are not. Our experiments show that our reconstructions, obtained from two or more views, are significantly more accurate than those of state-of-the-art methods, while also being faster. %In this paper, we show that, under widely applicable assumptions, we can derive a new system of equations in terms of the surface normals, whose two solutions can be obtained in closed-form and can easily be disambiguated locally. Our formalism also allows us to assess how reliable the estimated local normals are and to discard them if they are not. Our experiments show that our reconstructions, obtained from two or more views, are significantly more accurate than those of state-of-the-art methods, while also being faster. Shaifali Parashar, Yuxuan Long, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | TempSAL - Uncovering Temporal Information for Deep Saliency PredictionabstractDeep saliency prediction algorithms complement the object recognition features, they typically rely on additional information such as scene context, semantic relationships, gaze direction, and object dissimilarity. However, none of these models consider the temporal nature of gaze shifts during image observation. We introduce a novel saliency prediction model that learns to output saliency maps in sequential time intervals by exploiting human temporal attention patterns. Our approach locally modulates the saliency predictions by combining the learned temporal maps. Our experiments show that our method outperforms the state-of-the-art models, including a multi-duration saliency model, on the SALICON benchmark and CodeCharts1k dataset. Our code is publicly available on GitHub11https://ivrl.github.io/Tempsal/. Bahar Aydemir, Ludo Hoffstetter, Tong Zhang 0023, Mathieu Salzmann, Sabine Süsstrunk |
CVPR | 4 |
| 2023 | Knowledge Distillation for 6D Pose Estimation by Aligning Distributions of Local PredictionsabstractKnowledge distillation facilitates the training of a compact student network by using a deep teacher one. While this has achieved great success in many tasks, it remains completely unstudied for image-based 6D object pose estimation. In this work, we introduce the first knowledge distillation method driven by the 6D pose estimation task. To this end, we observe that most modern 6D pose estimation frameworks output local predictions, such as sparse 2D keypoints or dense representations, and that the compact student network typically struggles to predict such local quantities precisely. Therefore, instead of imposing prediction-to-prediction supervision from the teacher to the student, we propose to distill the teacher's distribution of local predictions into the student network, facilitating its training. Our experiments on several benchmarks show that our distillation method yields state-of-the-art results with different compact student models and for both keypoint-based and dense prediction-based architectures. Shuxuan Guo, Yinlin Hu, José M. Álvarez 0004, Mathieu Salzmann |
CVPR | 4 |
| 2023 | Rigidity-Aware Detection for 6D Object Pose EstimationabstractMost recent 6D object pose estimation methods first use object detection to obtain 2D bounding boxes before actually regressing the pose. However, the general object detection methods they use are ill-suited to handle cluttered scenes, thus producing poor initialization to the subsequent pose network. To address this, we propose a rigidity-aware detection method exploiting the fact that, in 6D pose estimation, the target objects are rigid. This lets us introduce an approach to sampling positive object regions from the entire visible object area during training, instead of naively drawing samples from the bounding box center where the object might be occluded. As such, every visible object part can contribute to the final bounding box prediction, yielding better detection robustness. Key to the success of our approach is a visibility map, which we propose to build using a minimum barrier distance between every pixel in the bounding box and the box boundary. Our results on seven challenging 6D pose estimation datasets evidence that our method outperforms general detection frameworks by a large margin. Furthermore, combined with a pose regression network, we obtain state-of-the-art pose estimation results on the challenging BOP benchmark. Yang Hai, Rui Song 0003, Jiaojiao Li 0001, Mathieu Salzmann, Yinlin Hu |
CVPR | 4 |
| 2023 | Robust Outlier Rejection for 3D Registration with Variational BayesabstractLearning-based outlier (mismatched correspondence) rejection for robust 3D registration generally formulates the outlier removal as an inlier/outlier classification problem. The core for this to be successful is to learn the discriminative inlier/outlier feature representations. In this paper, we develop a novel variational non-local network-based outlier rejection framework for robust alignment. By reformulating the non-local feature learning with variational Bayesian inference, the Bayesian-driven long-range dependencies can be modeled to aggregate discriminative geometric context information for inlier/outlier distinction. Specifically, to achieve such Bayesian-driven contextual dependencies, each query/key/value component in our nonlocal network predicts a prior feature distribution and a posterior one. Embedded with the inlier/outlier label, the posterior feature distribution is label-dependent and discriminative. Thus, pushing the prior to be close to the discriminative posterior in the training step enables the features sampled from this prior at test time to model highquality long-range dependencies. Notably, to achieve effective posterior feature guidance, a specific probabilistic graphical model is designed over our non-local model, which lets us derive a variational low bound as our optimization objective for model training. Finally, we propose a voting-based inlier searching strategy to cluster the high-quality hypothetical inliers for transformation estimation. Extensive experiments on 3DMatch, 3DLoMatch, and KITTI datasets verify the effectiveness of our method. Code is available at https://github.com/Jiang-HB/VBReg. Haobo Jiang, Zheng Dang, Zhen Wei 0001, Jin Xie 0001, Jian Yang 0003, Mathieu Salzmann |
CVPR | 6 |
| 2023 | DrapeNet: Garment Generation and Self-Supervised DrapingabstractRecent approaches to drape garments quickly over arbitrary human bodies leverage self-supervision to eliminate the need for large training sets. However, they are designed to train one network per clothing item, which severely limits their generalization abilities. In our work, we rely on self-supervision to train a single network to drape multiple garments. This is achieved by predicting a 3D deformation field conditioned on the latent codes of a generative network, which models garments as unsigned distance fields. Our pipeline can generate and drape previously unseen garments of any topology, whose shape can be edited by manipulating their latent codes. Being fully differentiable, our formulation makes it possible to recover accurate 3D models of garments from partial observations - images or 3D scans - via gradient descent. Our code is publicly available at https://github.com/liren2515/DrapeNet Luca De Luigi, Benoît Guillard, Mathieu Salzmann, Pascal Fua |
CVPR | 4 |
| 2023 | CLIP the Gap: A Single Domain Generalization Approach for Object DetectionabstractSingle Domain Generalization (SDG) tackles the problem of training a model on a single source domain so that it generalizes to any unseen target domain. While this has been well studied for image classification, the literature on SDG object detection remains almost non-existent. To address the challenges of simultaneously learning robust object localization and representation, we propose to leverage a pre-trained vision-language model to introduce semantic domain concepts via textual prompts. We achieve this via a semantic augmentation strategy acting on the features extracted by the detector backbone, as well as a text-based classification loss. Our experiments evidence the benefits of our approach, outperforming by 10% the only existing SDG object detection method, Single-DGOD [52], on their own diverse weather-driving benchmark. Vidit Vidit, Martin Engilberge, Mathieu Salzmann |
CVPR | 3 |
| 2023 | Learning Transformations to Reduce the Geometric Shift in Object DetectionabstractThe performance of modern object detectors drops when the test distribution differs from the training one. Most of the methods that address this focus on object appearance changes caused by, e.g., different illumination conditions, or gaps between synthetic and real images. Here, by contrast, we tackle geometric shifts emerging from variations in the image capture process, or due to the constraints of the environment causing differences in the apparent geometry of the content itself. We introduce a self-training approach that learns a set of geometric transformations to minimize these shifts without leveraging any labeled data in the new domain, nor any information about the cameras. We evaluate our method on two different shifts, i.e., a camera's field of view (FoV) change and a viewpoint change. Our results evidence that learning geometric transformations helps detectors to perform better in the target domains. Vidit Vidit, Martin Engilberge, Mathieu Salzmann |
CVPR | 3 |
| 2023 | Spatiotemporal Self-Supervised Learning for Point Clouds in the WildabstractSelf-supervised learning (SSL) has the potential to benefit many applications, particularly those where manually annotating data is cumbersome. One such situation is the semantic segmentation of point clouds. In this context, existing methods employ contrastive learning strategies and define positive pairs by performing various augmentation of point clusters in a single frame. As such, these methods do not exploit the temporal nature of LiDAR data. In this paper, we introduce an SSL strategy that leverages positive pairs in both the spatial and temporal domain. To this end, we design (i) a point-to-cluster learning strategy that aggregates spatial information to distinguish objects; and (ii) a cluster-to-cluster learning strategy based on unsupervised object tracking that exploits temporal correspondences. We demonstrate the benefits of our approach via extensive experiments performed by self-supervised training on two large-scale LiDAR datasets and transferring the resulting models to other point cloud segmentation benchmarks. Our results evidence that our method outperforms the state-of-the-art point cloud SSL methods.11Our code and pretrained models will be found at https://github.com/YanhaoWu/STSSL. Correspondence to Ke Wei. Yanhao Wu, Tong Zhang 0023, Wei Ke 0003, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 5 |
| 2023 | Vision Transformer Adapters for Generalizable Multitask LearningabstractWe introduce the first multitasking vision transformer adapters that learn generalizable task affinities which can be applied to novel tasks and domains. Integrated into an off-the-shelf vision transformer backbone, our adapters can simultaneously solve multiple dense vision tasks in a parameter-efficient manner, unlike existing multitasking transformers that are parametrically expensive. In contrast to concurrent methods, we do not require retraining or fine-tuning whenever a new task or domain is added. We introduce a task-adapted attention mechanism within our adapter framework that combines gradient-based task similarities with attention-based ones. The learned task affinities generalize to the following settings: zero-shot task transfer, unsupervised domain adaptation, and generalization without fine-tuning to novel domains. We demonstrate that our approach outperforms not only the existing convolutional neural network-based multitasking methods but also the vision transformer-based ones. Our project page is at https://ivrl.github.io/VTAGML. Deblina Bhattacharjee, Sabine Süsstrunk, Mathieu Salzmann |
ICCV | 3 |
| 2023 | AutoSynth: Learning to Generate 3D Training Data for Object Point Cloud RegistrationabstractIn the current deep learning paradigm, the amount and quality of training data are as critical as the network architecture and its training details. However, collecting, processing, and annotating real data at scale is difficult, expensive, and time-consuming, particularly for tasks such as 3D object registration. While synthetic datasets can be created, they require expertise to design and include a limited number of categories. In this paper, we introduce a new approach called AutoSynth, which automatically generates 3D training data for point cloud registration. Specifically, AutoSynth automatically curates an optimal dataset by exploring a search space encompassing millions of potential datasets with diverse 3D shapes at a low cost. To achieve this, we generate synthetic 3D datasets by assembling shape primitives, and develop a meta-learning strategy to search for the best training data for 3D registration on real point clouds. For this search to remain tractable, we replace the point cloud registration network with a much smaller surrogate network, leading to a 4056.43 times speedup. We demonstrate the generality of our approach by implementing it with two different point cloud registration networks, BPNet [13] and IDAM [34]. Our results on TUD-L [26], LINEMOD [23] and Occluded-LINEMOD [7] evidence that a neural network trained on our searched dataset yields consistently better performance than the same one trained on the widely used ModelNet40 dataset [65]. Zheng Dang, Mathieu Salzmann |
ICCV | 2 |
| 2023 | Center-Based Decoupled Point Cloud Registration for 6D Object Pose EstimationabstractIn this paper, we propose a novel center-based decoupled point cloud registration framework for robust 6D object pose estimation in real-world scenarios. Our method decouples the translation from the entire transformation by predicting the object center and estimating the rotation in a center-aware manner. This center offset-based translation estimation is correspondence-free, freeing us from the difficulty of constructing correspondences in challenging scenarios, thus improving robustness. To obtain reliable center predictions, we use a multi-view (bird’s eye view and front view) object shape description of the source-point features, with both views jointly voting for the object center. Additionally, we propose an effective shape embedding module to augment the source features, largely completing the missing shape information due to partial scanning, thus facilitating the center prediction. With the center-aligned source and model point clouds, the rotation predictor utilizes feature similarity to establish putative correspondences for SVD-based rotation estimation. In particular, we introduce a center-aware hybrid feature descriptor with a normal correction technique to extract discriminative, part-aware features for high-quality correspondence construction. Our experiments show that our method outperforms the state-of-the-art methods by a large margin on real-world datasets such as TUD-L, LINEMOD, and Occluded-LINEMOD. Code is available at https://github.com/JiangHB/CenterReg. Haobo Jiang, Zheng Dang, Shuo Gu, Jin Xie 0001, Mathieu Salzmann, Jian Yang 0003 |
ICCV | 5 |
| 2023 | Linear-Covariance Loss for End-to-End Learning of 6D Pose EstimationabstractMost modern image-based 6D object pose estimation methods learn to predict 2D-3D correspondences, from which the pose can be obtained using a PnP solver. Because of the non-differentiable nature of common PnP solvers, these methods are supervised via the individual correspondences. To address this, several methods have designed differentiable PnP strategies, thus imposing supervision on the pose obtained after the PnP step. Here, we argue that this conflicts with the averaging nature of the PnP problem, leading to gradients that may encourage the network to degrade the accuracy of individual correspondences. To address this, we derive a loss function that exploits the ground truth pose before solving the PnP problem. Specifically, we linearize the PnP solver around the ground-truth pose and compute the covariance of the resulting pose distribution. We then define our loss based on the diagonal covariance elements, which entails considering the final pose estimate yet not suffering from the PnP averaging issue. Our experiments show that our loss consistently improves the pose estimation accuracy for both dense and sparse correspondence based methods, achieving state-of-the-art results on both Linemod-Occluded and YCB-Video. Yinlin Hu, Mathieu Salzmann |
ICCV | 3 |
| 2023 | MixCycle: Mixup Assisted Semi-Supervised 3D Single Object Tracking with Cycle Consistencyabstract3D single object tracking (SOT) is an indispensable part of automated driving. Existing approaches rely heavily on large, densely labeled datasets. However, annotating point clouds is both costly and time-consuming. Inspired by the great success of cycle tracking in unsupervised 2D SOT, we introduce the first semi-supervised approach to 3D SOT. Specifically, we introduce two cycle-consistency strategies for supervision: 1) Self tracking cycles, which leverage labels to help the model converge better in the early stages of training; 2) forward-backward cycles, which strengthen the tracker’s robustness to motion variations and the template noise caused by the template update strategy. Furthermore, we propose a data augmentation strategy named SOTMixup to improve the tracker’s robustness to point cloud diversity. SOTMixup generates training samples by sampling points in two point clouds with a mixing rate and assigns a reasonable loss weight for training according to the mixing rate. The resulting MixCycle approach generalizes to appearance matching-based trackers. On the KITTI benchmark, based on the P2B tracker [16], MixCycle trained with 10% labels outperforms P2B trained with 100% labels, and achieves a 28.4% precision improvement when using 1% labels. Our code will be released at https://github.com/Mumuqiao/MixCycle. Qiao Wu, Jiaqi Yang 0002, Kun Sun 0002, Chu'ai Zhang, Yanning Zhang 0001, Mathieu Salzmann |
ICCV | 6 |
| 2023 | Towards Stable and Efficient Adversarial Training against l1 Bounded Adversarial AttacksabstractWe address the problem of stably and efficiently training a deep neural network robust to adversarial perturbations bounded by an $l_1$ norm. We demonstrate that achieving robustness against $l_1$-bounded perturbations is more challenging than in the $l_2$ or $l_\infty$ cases, because adversarial training against $l_1$-bounded perturbations is more likely to suffer from catastrophic overfitting and yield training instabilities. Our analysis links these issues to the coordinate descent strategy used in existing methods. We address this by introducing Fast-EG-$l_1$, an efficient adversarial training algorithm based on Euclidean geometry and free of coordinate descent. Fast-EG-$l_1$ comes with no additional memory costs and no extra hyper-parameters to tune. Our experimental results on various datasets demonstrate that Fast-EG-$l_1$ yields the best and most stable robustness against $l_1$-bounded adversarial attacks among the methods of comparable computational complexity. Code and the checkpoints are available at https://github.com/IVRL/FastAdvL. Yulun Jiang, Chen Liu 0027, Zhichao Huang 0002, Mathieu Salzmann, Sabine Süsstrunk |
ICML | 4 |
| 2023 | SE(3) Diffusion Model-based Point Cloud Registration for Robust 6D Object Pose EstimationabstractIn this paper, we introduce an SE(3) diffusion model-based point cloud registration framework for 6D object pose estimation in real-world scenarios. Our approach formulates the 3D registration task as a denoising diffusion process, which progressively refines the pose of the source point cloud to obtain a precise alignment with the model point cloud. Training our framework involves two operations: An SE(3) diffusion process and an SE(3) reverse process. The SE(3) diffusion process gradually perturbs the optimal rigid transformation of a pair of point clouds by continuously injecting noise (perturbation transformation). By contrast, the SE(3) reverse process focuses on learning a denoising network that refines the noisy transformation step-by-step, bringing it closer to the optimal transformation for accurate pose estimation. Unlike standard diffusion models used in linear Euclidean spaces, our diffusion model operates on the SE(3) manifold. This requires exploiting the linear Lie algebra $\mathfrak{se}(3)$ associated with SE(3) to constrain the transformation transitions during the diffusion and reverse processes. Additionally, to effectively train our denoising network, we derive a registration-specific variational lower bound as the optimization objective for model learning. Furthermore, we show that our denoising network can be constructed with a surrogate registration model, making our approach applicable to different deep registration networks. Extensive experiments demonstrate that our diffusion registration framework presents outstanding pose estimation performance on the real-world TUD-L, LINEMOD, and Occluded-LINEMOD datasets. Haobo Jiang, Mathieu Salzmann, Zheng Dang, Jin Xie 0001, Jian Yang 0003 |
NeurIPS | 2 |
| 2023 | Center-aware Adversarial Augmentation for Single Domain GeneralizationabstractDomain generalization (DG) aims to learn a model from multiple training (i.e., source) domains that can generalize well to the unseen test (i.e., target) data coming from a different distribution. Single domain generalization (Single-DG) has recently emerged to tackle a more challenging, yet realistic setting, where only one source domain is available at training time. The existing Single-DG approaches typically are based on data augmentation strategies and aim to expand the span of source data by augmenting out-of-domain samples. Generally speaking, they aim to generate hard examples to confuse the classifier. While this may make the classifier robust to small perturbation, the generated samples are typically not diverse enough to mimic a large domain shift, resulting in sub-optimal generalization performance. To alleviate this, we propose a center-aware adversarial augmentation technique that expands the source distribution by altering the source samples so as to push them away from the class centers via a novel angular center loss. We conduct extensive experiments to demonstrate the effectiveness of our approach on several benchmark datasets for Single-DG and show that our method outperforms the state-of-the-art in most cases. Mahsa Baktash, Zijian Wang 0009, Mathieu Salzmann |
WACV | 4 |
| 2023 | Temporal Representation Learning on Monocular Videos for 3D Human Pose EstimationabstractIn this article we propose an unsupervised feature extraction method to capture temporal information on monocular videos, where we detect and encode subject of interest in each frame and leverage contrastive self-supervised (CSS) learning to extract rich latent vectors. Instead of simply treating the latent features of nearby frames as positive pairs and those of temporally-distant ones as negative pairs as in other CSS approaches, we explicitly disentangle each latent vector into a time-variant component and a time-invariant one. We then show that applying contrastive loss only to the time-variant features and encouraging a gradual transition on them between nearby and away frames while also reconstructing the input, extract rich temporal features, well-suited for human pose estimation. Our approach reduces error by about 50% compared to the standard CSS strategies, outperforms other unsupervised single-view methods and matches the performance of multi-view techniques. When 2D pose is available, our approach can extract even richer latent features and improve the 3D pose estimation accuracy, outperforming other state-of-the-art weakly supervised methods. Sina Honari, Victor Constantin, Helge Rhodin, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Fast Adversarial Training With Adaptive Step SizeabstractWhile adversarial training and its variants have shown to be the most effective algorithms to defend against adversarial attacks, their extremely slow training process makes it hard to scale to large datasets like ImageNet. The key idea of recent works to accelerate adversarial training is to substitute multi-step attacks (e.g., PGD) with single-step attacks (e.g., FGSM). However, these single-step methods suffer from catastrophic overfitting, where the accuracy against PGD attack suddenly drops to nearly 0% during training, and the network totally loses its robustness. In this work, we study the phenomenon from the perspective of training instances. We show that catastrophic overfitting is instance-dependent, and fitting instances with larger input gradient norm is more likely to cause catastrophic overfitting. Based on our findings, we propose a simple but effective method, Adversarial Training with Adaptive Step size (ATAS). ATAS learns an instance-wise adaptive step size that is inversely proportional to its gradient norm. Our theoretical analysis shows that ATAS converges faster than the commonly adopted non-adaptive counterparts. Empirically, ATAS consistently mitigates catastrophic overfitting and achieves higher robust accuracy on CIFAR10, CIFAR100, and ImageNet when evaluated on various adversarial budgets. Our code is released at https://github.com/HuangZhiChao95/ATAS. Zhichao Huang 0002, Yanbo Fan, Chen Liu 0027, Yong Zhang 0034, Mathieu Salzmann, Sabine Süsstrunk, Jue Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | Training Provably Robust Models by Polyhedral Envelope RegularizationabstractTraining certifiable neural networks enables us to obtain models with robustness guarantees against adversarial attacks. In this work, we introduce a framework to obtain a provable adversarial-free region in the neighborhood of the input data by a polyhedral envelope, which yields more fine-grained certified robustness than existing methods. We further introduce polyhedral envelope regularization (PER) to encourage larger adversarial-free regions and thus improve the provable robustness of the models. We demonstrate the flexibility and effectiveness of our framework on standard benchmarks; it applies to networks of different architectures and with general activation functions. Compared with state of the art, PER has negligible computational overhead; it achieves better robustness guarantees and accuracy on the clean data in various settings. Chen Liu 0027, Mathieu Salzmann, Sabine Süsstrunk |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Long Term Motion Prediction Using KeyposesabstractLong term human motion prediction is essential in safety-critical applications such as human-robot interaction and autonomous driving. In this paper we show that to achieve long term forecasting, predicting human pose at every time instant is unnecessary. Instead, it is more effective to predict a few keyposes and approximate intermediate ones by interpolating the keyposes. We demonstrate that our approach enables us to predict realistic motions for up to 5 seconds in the future, which is far longer than the typical 1 second encountered in the literature. Furthermore, because we model future keyposes probabilistically, we can generate multiple plausible future motions by sampling at inference time. Over this extended time period, our predictions are more realistic, more diverse and better preserve the motion dynamics than those state-of-the-art methods yield. Sena Kiciroglu, Wei Wang 0108, Mathieu Salzmann, Pascal Fua |
3DV | 3 |
| 2022 | 3D Pose Based Feedback for Physical Exercises
Sena Kiciroglu, Hugues Vinzant, Isinsu Katircioglu, Mathieu Salzmann, Pascal Fua |
ACCV (4) | 6 |
| 2022 | Weakly-supervised Action Transition Learning for Stochastic Human Motion PredictionabstractWe introduce the task of action-driven stochastic human motion prediction, which aims to predict multiple plausible future motions given a sequence of action labels and a short motion history. This differs from existing works, which predict motions that either do not respect any specific action category, or follow a single action label. In particular, addressing this task requires tackling two challenges: The transitions between the different actions must be smooth; the length of the predicted motion depends on the action sequence and varies significantly across samples. As we cannot realistically expect training data to cover sufficiently diverse action transitions and motion lengths, we propose an effective training strategy consisting of combining multiple motions from different actions and introducing a weak form of supervision to encourage smooth transitions. We then design a VAE-based model conditioned on both the observed motion and the action label sequence, allowing us to generate multiple plausible future motions of varying length. We illustrate the generality of our approach by exploring its use with two different temporal encoding mod-els, namely RNNs and Transformers. Our approach out-performs baseline models constructed by adapting state-of-the-art single action-conditioned motion generation methods and stochastic human motion prediction approaches to our new task of action-driven stochastic motion prediction. Our code is available at https://github.com/wei-mao-2019/WAT. Wei Mao 0001, Miaomiao Liu 0001, Mathieu Salzmann |
CVPR | 3 |
| 2022 | MuIT: An End-to-End Multitask Learning TransformerabstractWe propose an end-to-end Multitask Learning Transformer framework, named MulT, to simultaneously learn multiple high-level vision tasks, including depth estimation, semantic segmentation, reshading, surface normal estimation, 2D keypoint detection, and edge detection. Based on the Swin transformer model, our framework encodes the input image into a shared representation and makes predictions for each vision task using task-specific transformer-based decoder heads. At the heart of our approach is a shared attention mechanism modeling the dependencies across the tasks. We evaluate our model on several multitask benchmarks, showing that our MulT framework outperforms both the state-of-the art multitask convolutional neural network models and all the respective single task transformer models. Our experiments further highlight the benefits of sharing attention across all the tasks, and demonstrate that our MulT model is robust and generalizes well to new domains. Our project website is at https://ivrl.github.io/MulT/. Deblina Bhattacharjee, Tong Zhang 0023, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 4 |
| 2022 | Adversarial Parametric Pose PriorabstractThe Skinned Multi-Person Linear (SMPL) model represents human bodies by mapping pose and shape parameters to body meshes. However, not all pose and shape parameter values yield physically-plausible or even realistic body meshes. In other words, SMPL is under-constrained and may yield invalid results. We propose learning a prior that restricts the SMPL parameters to values that produce realistic poses via adversarial training. We show that our learned prior covers the diversity of the real-data distribution, facilitates optimization for 3D reconstruction from 2D keypoints, and yields better pose estimates when used for regression from images. For all these tasks, it outperforms the state-of-the-art VAE-based approach to constraining the SMPL parameters. The code will be made available at https://github.com/cvlab-epfl/adv_param_pose_prior. Andrey Davydov, Anastasia Remizova, Victor Constantin, Sina Honari, Mathieu Salzmann, Pascal Fua |
CVPR | 5 |
| 2022 | Templates for 3D Object Pose Estimation Revisited: Generalization to New Objects and Robustness to OcclusionsabstractWe present a method that can recognize new objects and estimate their 3D pose in RGB images even under partial occlusions. Our method requires neither a training phase on these objects nor real images depicting them, only their CAD models. It relies on a small set of training objects to learn local object representations, which allow us to locally match the input image to a set of “templates”, rendered images of the CAD models for the new objects. In contrast with the state-of-the-art methods, the new objects on which our method is applied can be very different from the training objects. As a result, we are the first to show generalization without retraining on the LINEMOD and Occlusion-LINEMOD datasets. Our analysis of the failure modes of previous template-based approaches further confirms the benefits of local features for template matching. We outperform the state-of-the-art template matching methods on the LINEMOD, Occlusion-LINEMOD and T-LESS datasets. Our source code and data are publicly available at https://github.com/nv-nguyen/template-pose. Van Nguyen Nguyen, Yinlin Hu, Yang Xiao 0009, Mathieu Salzmann, Vincent Lepetit |
CVPR | 4 |
| 2022 | Leverage Your Local and Global Representations: A New Self-Supervised Learning StrategyabstractSelf-supervised learning (SSL) methods aim to learn view-invariant representations by maximizing the similar-ity between the features extracted from different crops of the same image regardless of cropping size and content. In essence, this strategy ignores the fact that two crops may truly contain different image information, e.g., background and small objects, and thus tends to restrain the diversity of the learned representations. In this work, we address this issue by introducing a new self-supervised learning strat-egy, LoGo, that explicitly reasons about Local and Global crops. To achieve view invariance, LoGo encourages similarity between global crops from the same image, as well as between a global and a local crop. However, to correctly encode the fact that the content of smaller crops may differ entirely, LoGo promotes two local crops to have dissimi-lar representations, while being close to global crops. Our LoGo strategy can easily be applied to existing SSL meth-ods. Our extensive experiments on a variety of datasets and using different self-supervised learning frameworks vali-date its superiority over existing approaches. Noticeably, we achieve better results than supervised models on trans-fer learning when using only 1/10 of the data.11Our code and pretrained models can be found at https://github.com/ztt1024/LoGo-SSL. Tong Zhang 0023, Congpei Qiu, Wei Ke 0003, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 5 |
| 2022 | Learning-Based Point Cloud Registration for 6D Object Pose Estimation in the Real World
Zheng Dang, Lizhou Wang, Yu Guo 0006, Mathieu Salzmann |
ECCV (1) | 4 |
| 2022 | Perspective Flow Aggregation for Data-Limited 6D Object Pose Estimation
Yinlin Hu, Pascal Fua, Mathieu Salzmann |
ECCV (2) | 3 |
| 2022 | Fusing Local Similarities for Retrieval-Based 3D Orientation Estimation of Unseen Objects
Chen Zhao 0025, Yinlin Hu, Mathieu Salzmann |
ECCV (1) | 3 |
| 2022 | Contrastive Class-aware Adaptation for Domain GeneralizationabstractDomain generalization (DG) tackles the problem of learning a model that generalizes to data drawn from a target domain that was unseen during training. A major trend in this area consists of learning a domain-invariant representation by minimizing the discrepancy across multiple source domains. This strategy, however, does not apply to the challenging yet realistic single-source scenario. In this paper, in contrast to existing methods that focus on domain discrepancy, we exploit the fact that discrepancies also arise across samples from the same class. We therefore develop a unified framework for both multisource and single-source DG that exploits contrastive learning to maximize the gap between samples from the same class, either from different domains or from the same one, while separating the samples from different classes. Our results on standard multisource and single-source DG benchmark datasets demonstrate the benefits of our method over the state-of-the-art ones in both settings. Mahsa Baktash, Mathieu Salzmann |
ICPR | 3 |
| 2022 | Contact-aware Human Motion ForecastingabstractIn this paper, we tackle the task of scene-aware 3D human motion forecasting, which consists of predicting future human poses given a 3D scene and a past human motion. A key challenge of this task is to ensure consistency between the human and the scene, accounting for human-scene interactions. Previous attempts to do so model such interactions only implicitly, and thus tend to produce artifacts such as ``ghost motion" because of the lack of explicit constraints between the local poses and the global motion. Here, by contrast, we propose to explicitly model the human-scene contacts. To this end, we introduce distance-based contact maps that capture the contact relationships between every joint and every 3D scene point at each time instant. We then develop a two-stage pipeline that first predicts the future contact maps from the past ones and the scene point cloud, and then forecasts the future human poses by conditioning them on the predicted contact maps. During training, we explicitly encourage consistency between the global motion and the local poses via a prior defined using the contact maps and future poses. Our approach outperforms the state-of-the-art human motion forecasting and human synthesis methods on both synthetic and real datasets. Our code is available at https://github.com/wei-mao-2019/ContAwareMotionPred. Wei Mao 0001, Miaomiao Liu 0001, Richard I. Hartley, Mathieu Salzmann |
NeurIPS | 4 |
| 2022 | Robust Binary Models by Pruning Randomly-initialized NetworksabstractRobustness to adversarial attacks was shown to require a larger model capacity, and thus a larger memory footprint. In this paper, we introduce an approach to obtain robust yet compact models by pruning randomly-initialized binary networks. Unlike adversarial training, which learns the model parameters, we initialize the model parameters as either +1 or −1, keep them fixed, and find a subnetwork structure that is robust to attacks. Our method confirms the Strong Lottery Ticket Hypothesis in the presence of adversarial attacks, and extends this to binary networks. Furthermore, it yields more compact networks with competitive performance than existing works by 1) adaptively pruning different network layers; 2) exploiting an effective binary initialization scheme; 3) incorporating a last batch normalization layer to improve training stability. Our experiments demonstrate that our approach not only always outperforms the state-of-the-art robust binary networks, but also can achieve accuracy better than full-precision ones on some datasets. Finally, we show the structured patterns of our pruned binary networks. Chen Liu 0027, Sabine Süsstrunk, Mathieu Salzmann |
NeurIPS | 4 |
| 2022 | Learning to Generate the Unknowns as a Remedy to the Open-Set Domain ShiftabstractIn many situations, the data one has access to at test time follows a different distribution from the training data. Over the years, this problem has been tackled by closed-set domain adaptation techniques. Recently, open-set domain adaptation has emerged to address the more realistic scenario where additional unknown classes are present in the target data. In this setting, existing techniques focus on the challenging task of isolating the unknown target samples, so as to avoid the negative transfer resulting from aligning the source feature distributions with the broader target one that encompasses the additional unknown classes. Here, we propose a simpler and more effective solution consisting of complementing the source data distribution and making it comparable to the target one by enabling the model to generate source samples corresponding to the unknown target classes. We formulate this as a general module that can be incorporated into any existing closed-set approach and show that this strategy allows us to outperform the state of the art on open-set domain adaptation benchmark datasets. Mahsa Baktash, Mathieu Salzmann |
WACV | 3 |
| 2022 | Estimating Image Depth in the Comics DomainabstractEstimating the depth of comics images is challenging as such images a) are monocular; b) lack ground-truth depth annotations; c) differ across different artistic styles; d) are sparse and noisy. We thus, use an off-the-shelf unsupervised image to image translation method to translate the comics images to natural ones and then use an attention-guided monocular depth estimator to predict their depth. This lets us leverage the depth annotations of existing natural images to train the depth estimator. Furthermore, our model learns to distinguish between text and images in the comics panels to reduce text-based artefacts in the depth estimates. Our method consistently outperforms the existing state-of-the-art approaches across all metrics on both the DCM and eBDtheque images. Finally, we introduce a dataset to evaluate depth prediction on comics. Deblina Bhattacharjee, Martin Nicolas Everaert, Mathieu Salzmann, Sabine Süsstrunk |
WACV | 3 |
| 2022 | Attention-based domain adaptation for single-stage detectorsabstractAbstract While domain adaptation has been used to improve the performance of object detectors when the training and test data follow different distributions, previous work has mostly focused on two-stage detectors. This is because their use of region proposals makes it possible to perform local adaptation, which has been shown to significantly improve the adaptation effectiveness. Here, by contrast, we target single-stage architectures, which are better suited to resource-constrained detection than two-stage ones but do not provide region proposals. To nonetheless benefit from the strength of local adaptation, we introduce an attention mechanism that lets us identify the important regions on which adaptation should focus. Our method gradually adapts the features from global, image level to local, instance level. Our approach is generic and can be integrated into any Single-Shot Detector. We demonstrate this on standard benchmark datasets by applying it to both the single-shot detector (SSD) and a recent variant of the You Only Look Once detector (YOLOv5). Furthermore, for equivalent single-stage architectures, our method outperforms the state-of-the-art domain adaptation techniques even though they were designed for specific detectors. Vidit Vidit, Mathieu Salzmann |
Mach. Vis. Appl. | 2 |
| 2022 | GarNet++: Improving Fast and Accurate Static 3D Cloth Draping by Curvature LossabstractIn this paper, we tackle the problem of static 3D cloth draping on virtual human bodies. We introduce a two-stream deep network model that produces a visually plausible draping of a template cloth on virtual 3D bodies by extracting features from both the body and garment shapes. Our network learns to mimic a physics-based simulation (PBS) method while requiring two orders of magnitude less computation time. To train the network, we introduce loss terms inspired by PBS to produce plausible results and make the model collision-aware. To increase the details of the draped garment, we introduce two loss functions that penalize the difference between the curvature of the predicted cloth and PBS. Particularly, we study the impact of mean curvature normal and a novel detail-preserving loss both qualitatively and quantitatively. Our new curvature loss computes the local covariance matrices of the 3D points, and compares the Rayleigh quotients of the prediction and PBS. This leads to more details while performing favorably or comparably against the loss that considers mean curvature normal vectors in the 3D triangulated meshes. We validate our framework on four garment types for various body shapes and poses. Finally, we achieve superior performance against a recently proposed data-driven method. Erhan Gundogdu, Victor Constantin, Shaifali Parashar, Amrollah Seifoddini, Minh Dang, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Self-Supervised Human Detection and Segmentation via Background InpaintingabstractWhile supervised object detection and segmentation methods achieve impressive accuracy, they generalize poorly to images whose appearance significantly differs from the data they have been trained on. To address this when annotating data is prohibitively expensive, we introduce a self-supervised detection and segmentation approach that can work with single images captured by a potentially moving camera. At the heart of our approach lies the observation that object segmentation and background reconstruction are linked tasks, and that, for structured scenes, background regions can be re-synthesized from their surroundings, whereas regions depicting the moving object cannot. We encode this intuition into a self-supervised loss function that we exploit to train a proposal-based segmentation network. To account for the discrete nature of the proposals, we develop a Monte Carlo-based training strategy that allows the algorithm to explore the large space of object proposals. We apply our method to human detection and segmentation in images that visually depart from those of standard benchmarks and outperform existing self-supervised methods. Isinsu Katircioglu, Helge Rhodin, Victor Constantin, Jörg Spörri, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Counting People by Estimating People FlowsabstractModern methods for counting people in crowded scenes rely on deep networks to estimate people densities in individual images. As such, only very few take advantage of temporal consistency in video sequences, and those that do only impose weak smoothness constraints across consecutive frames. In this paper, we advocate estimating people flows across image locations between consecutive images and inferring the people densities from these flows instead of directly regressing them. This enables us to impose much stronger constraints encoding the conservation of the number of people. As a result, it significantly boosts performance without requiring a more complex architecture. Furthermore, it allows us to exploit the correlation between people flow and optical flow to further improve the results. We also show that leveraging people conservation constraints in both a spatial and temporal manner makes it possible to train a deep crowd counting model in an active learning setting with much fewer annotations. This significantly reduces the annotation cost while still leading to similar performance to the full supervision case. Weizhe Liu, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Robust Differentiable SVDabstractEigendecomposition of symmetric matrices is at the heart of many computer vision algorithms. However, the derivatives of the eigenvectors tend to be numerically unstable, whether using the SVD to compute them analytically or using the Power Iteration (PI) method to approximate them. This instability arises in the presence of eigenvalues that are close to each other. This makes integrating eigendecomposition into deep networks difficult and often results in poor convergence, particularly when dealing with large matrices. While this can be mitigated by partitioning the data into small arbitrary groups, doing so has no theoretical basis and makes it impossible to exploit the full power of eigendecomposition. In previous work, we mitigated this using SVD during the forward pass and PI to compute the gradients during the backward pass. However, the iterative deflation procedure required to compute multiple eigenvectors using PI tends to accumulate errors and yield inaccurate gradients. Here, we show that the Taylor expansion of the SVD gradient is theoretically equivalent to the gradient obtained using PI without relying in practice on an iterative process and thus yields more accurate gradients. We demonstrate the benefits of this increased accuracy for image classification and style transfer. Wei Wang 0108, Zheng Dang, Yinlin Hu, Pascal Fua, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | An Analysis of Super-Net Heuristics in Weight-Sharing NASabstractWeight sharing promises to make neural architecture search (NAS) tractable even on commodity hardware. Existing methods in this space rely on a diverse set of heuristics to design and train the shared-weight backbone network, a.k.a. the super-net. Since heuristics substantially vary across different methods and have not been carefully studied, it is unclear to which extent they impact super-net training and hence the weight-sharing NAS algorithms. In this paper, we disentangle super-net training from the search algorithm, isolate 14 frequently-used training heuristics, and evaluate them over three benchmark search spaces. Our analysis uncovers that several commonly-used heuristics negatively impact the correlation between super-net and stand-alone performance, whereas simple, but often overlooked factors, such as proper hyper-parameter settings, are key to achieve strong performance. Equipped with this knowledge, we show that simple random search achieves competitive performance to complex state-of-the-art NAS algorithms when the super-net is properly trained. Kaicheng Yu, René Ranftl, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | SD-Pose: Semantic Decomposition for Cross-Domain 6D Object Pose EstimationabstractThe current leading 6D object pose estimation methods rely heavily on annotated real data, which is highly costly to acquire. To overcome this, many works have proposed to introduce computer-generated synthetic data. However, bridging the gap between the synthetic and real data remains a severe problem. Images depicting different levels of realism/semantics usually have different transferability between the synthetic and real domains. Inspired by this observation, we introduce an approach, SD-Pose, that explicitly decomposes the input image into multi-level semantic representations and then combines the merits of each representation to bridge the domain gap. Our comprehensive analyses and experiments show that our semantic decomposition strategy can fully utilize the different domain similarities of different representations, thus allowing us to outperform the state of the art on modern 6D object pose datasets without accessing any real data during training. Yinlin Hu, Mathieu Salzmann, Xiangyang Ji |
AAAI | 3 |
| 2021 | Wide-Depth-Range 6D Object Pose Estimation in Spaceabstract6D pose estimation in space poses unique challenges that are not commonly encountered in the terrestrial setting. One of the most striking differences is the lack of atmospheric scattering, allowing objects to be visible from a great distance while complicating illumination conditions. Currently available benchmark datasets do not place a sufficient emphasis on this aspect and mostly depict the target in close proximity.Prior work tackling pose estimation under large scale variations relies on a two-stage approach to first estimate scale, followed by pose estimation on a resized image patch. We instead propose a single-stage hierarchical end-to-end trainable network that is more robust to scale variations. We demonstrate that it outperforms existing approaches not only on images synthesized to resemble images taken in space but also on standard benchmarks. Yinlin Hu, Sébastien Speierer, Wenzel Jakob, Pascal Fua, Mathieu Salzmann |
CVPR | 5 |
| 2021 | Probabilistic Tracklet Scoring and Inpainting for Multiple Object TrackingabstractDespite the recent advances in multiple object tracking (MOT), achieved by joint detection and tracking, dealing with long occlusions remains a challenge. This is due to the fact that such techniques tend to ignore the long-term motion information. In this paper, we introduce a probabilistic autoregressive motion model to score tracklet proposals by directly measuring their likelihood. This is achieved by training our model to learn the underlying distribution of natural tracklets. As such, our model allows us not only to assign new detections to existing tracklets, but also to inpaint a tracklet when an object has been lost for a long time, e.g., due to occlusion, by sampling tracklets so as to fill the gap caused by misdetections. Our experiments demonstrate the superiority of our approach at tracking objects in challenging sequences; it outperforms the state of the art in most standard MOT metrics on multiple MOT benchmark datasets, including MOT16, MOT17, and MOT20. Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Seyed Hamid Rezatofighi, Mathieu Salzmann, Stephen Gould |
CVPR | 4 |
| 2021 | Landmark Regularization: Ranking Guided Super-Net Training in Neural Architecture SearchabstractWeight sharing has become a de facto standard in neural architecture search because it enables the search to be done on commodity hardware. However, recent works have empirically shown a ranking disorder between the performance of stand-alone architectures and that of the corresponding shared-weight networks. This violates the main assumption of weight-sharing NAS algorithms, thus limiting their effectiveness. We tackle this issue by proposing a regularization term that aims to maximize the correlation between the performance rankings of the shared-weight network and that of the standalone architectures using a small set of landmark architectures. We incorporate our regularization term into three different NAS algorithms and show that it consistently improves performance across algorithms, search-spaces, and tasks. Kaicheng Yu, René Ranftl, Mathieu Salzmann |
CVPR | 3 |
| 2021 | PCLs: Geometry-Aware Neural Reconstruction of 3D Pose With Perspective Crop LayersabstractLocal processing is an essential feature of CNNs and other neural network architectures—it is one of the reasons why they work so well on images where relevant information is, to a large extent, local. However, perspective effects stemming from the projection in a conventional camera vary for different global positions in the image. We introduce Perspective Crop Layers (PCLs)—a form of perspective crop of the region of interest based on the camera geometry— and show that accounting for the perspective consistently improves the accuracy of state-of-the-art 3D pose reconstruction methods. PCLs are modular neural network layers, which, when inserted into existing CNN and MLP architectures, deterministically remove the location-dependent perspective effects while leaving end-to-end training and the number of parameters of the underlying neural network unchanged. We demonstrate that PCL leads to improved 3D human pose reconstruction accuracy for CNN architectures that use cropping operations, such as spatial transformer networks (STN), and, somewhat surprisingly, MLPs used for 2D-to-3D key-point lifting. Our conclusion is that it is important to utilize camera calibration information when available, for classical and deep-learning-based computer vision alike. PCL offers an easy way to improve the accuracy of existing 3D reconstruction networks by making them geometry-aware. Our code is publicly available at github.com/yu-frank/PerspectiveCropLayers. Frank Yu, Mathieu Salzmann, Pascal Fua, Helge Rhodin |
CVPR | 2 |
| 2021 | Contextually Plausible and Diverse 3D Human Motion PredictionabstractWe tackle the task of diverse 3D human motion prediction, that is, forecasting multiple plausible future 3D poses given a sequence of observed 3D poses. In this context, a popular approach consists of using a Conditional Variational Autoencoder (CVAE). However, existing approaches that do so either fail to capture the diversity in human motion, or generate diverse but semantically implausible continuations of the observed motion. In this paper, we address both of these problems by developing a new variational framework that accounts for both diversity and context of the generated future motion. To this end, and in contrast to existing approaches, we condition the sampling of the latent variable that acts as source of diversity on the representation of the past observation, thus encouraging it to carry relevant information. Our experiments demonstrate that our approach yields motions not only of higher quality while retaining diversity, but also that preserve the contextual information contained in the observed motion. Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Lars Petersson, Stephen Gould, Mathieu Salzmann |
ICCV | 5 |
| 2021 | Temporally-Coherent Surface Reconstruction via Metric-Consistent AtlasesabstractWe propose a method for the unsupervised reconstruction of a temporally-coherent sequence of surfaces from a sequence of time-evolving point clouds, yielding dense, semantically meaningful correspondences between all keyframes. We represent the reconstructed surface as an atlas, using a neural network. Using canonical correspondences defined via the atlas, we encourage the reconstruction to be as isometric as possible across frames, leading to semantically-meaningful reconstruction. Through experiments and comparisons, we empirically show that our method achieves results that exceed that state of the art in the accuracy of unsupervised correspondences and accuracy of surface reconstruction. Jan Bednarík, Vladimir G. Kim, Siddhartha Chaudhuri, Shaifali Parashar, Mathieu Salzmann, Pascal Fua, Noam Aigerman |
ICCV | 5 |
| 2021 | Human Detection and Segmentation via Multi-view ConsensusabstractSelf-supervised detection and segmentation of foreground objects aims for accuracy without annotated training data. However, existing approaches predominantly rely on restrictive assumptions on appearance and motion.For scenes with dynamic activities and camera motion, we propose a multi-camera framework in which geometric constraints are embedded in the form of multi-view consistency during training via coarse 3D localization in a voxel grid and fine-grained offset regression. In this manner, we learn a joint distribution of proposals over multiple views. At inference time, our method operates on single RGB images. We outperform state-of-the-art techniques both on images that visually depart from those of standard benchmarks and on those of the classical Human3.6M dataset. Isinsu Katircioglu, Helge Rhodin, Jörg Spörri, Mathieu Salzmann, Pascal Fua |
ICCV | 4 |
| 2021 | Generating Smooth Pose Sequences for Diverse Human Motion PredictionabstractRecent progress in stochastic motion prediction, i.e., predicting multiple possible future human motions given a single past pose sequence, has led to producing truly diverse future motions and even providing control over the motion of some body parts. However, to achieve this, the state-of-the-art method requires learning several mappings for diversity and a dedicated model for controllable motion prediction. In this paper, we introduce a unified deep generative network for both diverse and controllable motion prediction. To this end, we leverage the intuition that realistic human motions consist of smooth sequences of valid poses, and that, given limited data, learning a pose prior is much more tractable than a motion one. We therefore design a generator that predicts the motion of different body parts sequentially, and introduce a normalizing flow based pose prior, together with a joint angle loss, to achieve motion realism. Our experiments on two standard benchmark datasets, Human3.6M and HumanEva-I, demonstrate that our approach outperforms the state-of-the-art baselines in terms of both sample diversity and accuracy. The code is available at https://github.com/wei-mao-2019/gsps Wei Mao 0001, Miaomiao Liu 0001, Mathieu Salzmann |
ICCV | 3 |
| 2021 | Progressive Correspondence Pruning by Consensus LearningabstractCorrespondence pruning aims to correctly remove false matches (outliers) from an initial set of putative correspondences. The pruning process is challenging since putative matches are typically extremely unbalanced, largely dominated by outliers, and the random distribution of such outliers further complicates the learning process for learning-based methods. To address this issue, we propose to progressively prune the correspondences via a local-to-global consensus learning procedure. We introduce a "pruning" block that lets us identify reliable candidates among the initial matches according to consensus scores estimated using local-to-global dynamic graphs. We then achieve progressive pruning by stacking multiple pruning blocks sequentially. Our method outperforms state-of-the-arts on robust line fitting, camera pose estimation and retrieval-based image localization benchmarks by significant margins and shows promising generalization ability to different datasets and detector/descriptor combinations. Chen Zhao 0025, Yixiao Ge, Feng Zhu 0006, Rui Zhao 0001, Hongsheng Li 0001, Mathieu Salzmann |
ICCV | 6 |
| 2021 | Distilling Image Classifiers in Object DetectorsabstractKnowledge distillation constitutes a simple yet effective way to improve the performance of a compact student network by exploiting the knowledge of a more powerful teacher. Nevertheless, the knowledge distillation literature remains limited to the scenario where the student and the teacher tackle the same task. Here, we investigate the problem of transferring knowledge not only across architectures but also across tasks. To this end, we study the case of object detection and, instead of following the standard detector-to-detector distillation approach, introduce a classifier-to-detector knowledge transfer framework. In particular, we propose strategies to exploit the classification teacher to improve both the detector's recognition accuracy and localization performance. Our experiments on several detectors with different backbones demonstrate the effectiveness of our approach, allowing us to outperform the state-of-the-art detector-to-detector distillation methods. Shuxuan Guo, José M. Álvarez 0004, Mathieu Salzmann |
NeurIPS | 3 |
| 2021 | Learning Transferable Adversarial PerturbationsabstractWhile effective, deep neural networks (DNNs) are vulnerable to adversarial attacks. In particular, recent work has shown that such attacks could be generated by another deep network, leading to significant speedups over optimization-based perturbations. However, the ability of such generative methods to generalize to different test-time situations has not been systematically studied. In this paper, we, therefore, investigate the transferability of generated perturbations when the conditions at inference time differ from the training ones in terms of the target architecture, target data, and target task. Specifically, we identify the mid-level features extracted by the intermediate layers of DNNs as common ground across different architectures, datasets, and tasks. This lets us introduce a loss function based on such mid-level features to learn an effective, transferable perturbation generator. Our experiments demonstrate that our approach outperforms the state-of-the-art universal and transferable attack strategies. Krishna K. Nakka, Mathieu Salzmann |
NeurIPS | 2 |
| 2021 | Multi-level Motion Attention for Human Motion Prediction
Wei Mao 0001, Miaomiao Liu 0001, Mathieu Salzmann, Hongdong Li |
Int. J. Comput. Vis. | 3 |
| 2021 | Eigendecomposition-Free Training of Deep Networks for Linear Least-Square ProblemsabstractMany classical Computer Vision problems, such as essential matrix computation and pose estimation from 3D to 2D correspondences, can be tackled by solving a linear least-square problem, which can be done by finding the eigenvector corresponding to the smallest, or zero, eigenvalue of a matrix representing a linear system. Incorporating this in deep learning frameworks would allow us to explicitly encode known notions of geometry, instead of having the network implicitly learn them from data. However, performing eigendecomposition within a network requires the ability to differentiate this operation. While theoretically doable, this introduces numerical instability in the optimization process in practice. In this paper, we introduce an eigendecomposition-free approach to training a deep network whose loss depends on the eigenvector corresponding to a zero eigenvalue of a matrix predicted by the network. We demonstrate that our approach is much more robust than explicit differentiation of the eigendecomposition using two general tasks, outlier rejection and denoising, with several practical examples including wide-baseline stereo, the perspective-n-point problem, and ellipse fitting. Empirically, our method has better convergence properties and yields state-of-the-art results. Zheng Dang, Kwang Moo Yi, Yinlin Hu, Fei Wang 0008, Pascal Fua, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Geometry-Aware Deep Recurrent Neural Networks for Hyperspectral Image ClassificationabstractVariants of deep networks have been widely used for hyperspectral image (HSI)-classification tasks. Among them, in recent years, recurrent neural networks (RNNs) have attracted considerable attention in the remote sensing community. However, complex geometries cannot be learned easily by the traditional recurrent units [e.g., long short-term memory (LSTM) and gated recurrent unit (GRU)]. In this article, we propose a geometry-aware deep recurrent neural network (Geo-DRNN) for HSI classification. We build this network upon two modules: a U-shaped network (U-Net) and RNNs. We first input the original HSI patches to the U-Net, which can be trained with very few images and obtain a preliminary classification result. We then add RNNs on the top of the U-Net so as to mimic the human brain to refine continuously the output-classification map. However, instead of using the traditional dot product in each gate of the RNNs, we introduce a Net-Gated GRU that increases the nonlinear representation power. Finally, we use a pretrained ResNet as a regularizer to improve further the ability of the proposed network to describe complex geometries. To this end, we construct a geometry-aware ResNet loss, which leverages the pretrained ResNet's knowledge about the different structures in the real world. Our experimental results on real HSIs and road topology images demonstrate that our approach outperforms the state-of-the-art classification methods and can learn complex geometries. Siyuan Hao, Wei Wang 0108, Mathieu Salzmann |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Better Patch Stitching for Parametric Surface ReconstructionabstractRecently, parametric mappings have emerged as highly effective surface representations, yielding low reconstruction error. In particular, the latest works represent the target shape as an atlas of multiple mappings, which can closely encode object parts. Atlas representations, however, suffer from one major drawback: The individual mappings are not guaranteed to be consistent, which results in holes in the reconstructed shape or in jagged surface areas.We introduce an approach that explicitly encourages global consistency of the local mappings. To this end, we introduce two novel loss terms. The first term exploits the surface normals and requires that they remain locally consistent when estimated within and across the individual mappings. The second term further encourages better spatial configuration of the mappings by minimizing novel stitching error. We show on standard benchmarks that the use of normal consistency requirement outperforms the baselines quantitatively while enforcing better stitching leads to much better visual quality of the reconstructed objects as compared to the state-of-the-art. Zhantao Deng, Jan Bednarík, Mathieu Salzmann, Pascal Fua |
3DV | 3 |
| 2020 | Motion Prediction Using Temporal Inception Module
Tim Lebailly, Sena Kiciroglu, Mathieu Salzmann, Pascal Fua, Wei Wang 0108 |
ACCV (2) | 3 |
| 2020 | Towards Robust Fine-Grained Recognition by Maximal Separation of Discriminative Features
Krishna K. Nakka, Mathieu Salzmann |
ACCV (6) | 2 |
| 2020 | A Stochastic Conditioning Scheme for Diverse Human Motion PredictionabstractHuman motion prediction, the task of predicting future 3D human poses given a sequence of observed ones, has been mostly treated as a deterministic problem. However, human motion is a stochastic process: Given an observed sequence of poses, multiple future motions are plausible. Existing approaches to modeling this stochasticity typically combine a random noise vector with information about the previous poses. This combination, however, is done in a deterministic manner, which gives the network the flexibility to learn to ignore the random noise. Alternatively, in this paper, we propose to stochastically combine the root of variations with previous pose information, so as to force the model to take the noise into account. We exploit this idea for motion prediction by incorporating it into a recurrent encoder-decoder network with a conditional variational autoencoder block that learns to exploit the perturbations. Our experiments on two large-scale motion prediction datasets demonstrate that our model yields high-quality pose sequences that are much more diverse than those from state-of-the-art stochastic motion prediction techniques. Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Lars Petersson, Stephen Gould |
CVPR | 3 |
| 2020 | Shape Reconstruction by Learning Differentiable Surface RepresentationsabstractGenerative models that produce point clouds have emerged as a powerful tool to represent 3D surfaces, and the best current ones rely on learning an ensemble of parametric representations. Unfortunately, they offer no control over the deformations of the surface patches that form the ensemble and thus fail to prevent them from either overlapping or collapsing into single points or lines. As a consequence, computing shape properties such as surface normals and curvatures becomes difficult and unreliable. In this paper, we show that we can exploit the inherent differentiability of deep networks to leverage differential surface properties during training so as to prevent patch collapse and strongly reduce patch overlap. Furthermore, this lets us reliably compute quantities such as surface normals and curvatures. We will demonstrate on several tasks that this yields more accurate surface reconstructions than the state-of-the-art methods in terms of normals estimation and amount of collapsed and overlapped patches. Jan Bednarík, Shaifali Parashar, Erhan Gundogdu, Mathieu Salzmann, Pascal Fua |
CVPR | 4 |
| 2020 | DUNIT: Detection-Based Unsupervised Image-to-Image TranslationabstractImage-to-image translation has made great strides in recent years, with current techniques being able to handle unpaired training images and to account for the multi-modality of the translation problem. Despite this, most methods treat the image as a whole, which makes the results they produce for content-rich scenes less realistic. In this paper, we introduce a Detection-based Unsupervised Image-to-image Translation (DUNIT) approach that explicitly accounts for the object instances in the translation process. To this end, we extract separate representations for the global image and for the instances, which we then fuse into a common representation from which we generate the translated image. This allows us to preserve the detailed content of object instances, while still modeling the fact that we aim to produce an image of a single consistent scene. We introduce an instance consistency loss to maintain the coherence between the detections. Furthermore, by incorporating a detector into our architecture, we can still exploit object instances at test time. As evidenced by our experiments, this allows us to outperform the state-of-the-art unsupervised image-to-image translation methods. Furthermore, our approach can also be used as an unsupervised domain adaptation strategy for object detection, and it also achieves state-of-the-art performance on this task. Deblina Bhattacharjee, Seungryong Kim, Guillaume Vizier, Mathieu Salzmann |
CVPR | 4 |
| 2020 | Single-Stage 6D Object Pose EstimationabstractMost recent 6D pose estimation frameworks first rely on a deep network to establish correspondences between 3D object keypoints and 2D image locations and then use a variant of a RANSAC-based Perspective-n-Point (PnP) algorithm. This two-stage process, however, is suboptimal: First, it is not end-to-end trainable. Second, training the deep network relies on a surrogate loss that does not directly reflect the final 6D pose estimation task. In this work, we introduce a deep architecture that directly regresses 6D poses from correspondences. It takes as input a group of candidate correspondences for each 3D keypoint and accounts for the fact that the order of the correspondences within each group is irrelevant, while the order of the groups, that is, of the 3D keypoints, is fixed. Our architecture is generic and can thus be exploited in conjunction with existing correspondence-extraction networks so as to yield single-stage 6D pose estimation frameworks. Our experiments demonstrate that these single-stage frameworks consistently outperform their two-stage counterparts in terms of both accuracy and speed. Yinlin Hu, Pascal Fua, Wei Wang 0108, Mathieu Salzmann |
CVPR | 4 |
| 2020 | ActiveMoCap: Optimized Viewpoint Selection for Active Human Motion CaptureabstractThe accuracy of monocular 3D human pose estimation depends on the viewpoint from which the image is captured. While freely moving cameras, such as on drones, provide control over this viewpoint, automatically positioning them at the location which will yield the highest accuracy remains an open problem. This is the problem that we address in this paper. Specifically, given a short video sequence, we introduce an algorithm that predicts which viewpoints should be chosen to capture future frames so as to maximize 3D human pose estimation accuracy. The key idea underlying our approach is a method to estimate the uncertainty of the 3D body pose estimates. We integrate several sources of uncertainty, originating from deep learning based regressors and temporal smoothness. Our motion planner yields improved 3D body pose estimates and outperforms or matches existing ones that are based on person following and orbiting. Sena Kiciroglu, Helge Rhodin, Sudipta N. Sinha, Mathieu Salzmann, Pascal Fua |
CVPR | 4 |
| 2020 | Local Non-Rigid Structure-From-Motion From Diffeomorphic MappingsabstractWe propose a new formulation to the non-rigid structure-from-motion problem that only requires the deforming surface to meaning that its differential structure is preserved. This is a much weaker assumption than the traditional ones of isometry or conformality. We show that it is nevertheless sufficient to establish local correspondences between the surface in two different images and therefore to perform point-wise reconstruction using only up to first-order derivatives. We formulate differential constraints and solve them algebraically using the theory of resultants. We will demonstrate that our approach is more widely applicable, more stable in noisy and sparse imaging conditions and much faster than earlier ones, while delivering similar accuracy. The code is available at https//github.com/cvlab-epf1/diff-nrsfm/. Shaifali Parashar, Mathieu Salzmann, Pascal Fua |
CVPR | 2 |
| 2020 | Volumetric Transformer Networks
Seungryong Kim, Sabine Süsstrunk, Mathieu Salzmann |
ECCV (28) | 3 |
| 2020 | Estimating People Flows to Better Count Them in Crowded Scenes
Weizhe Liu, Mathieu Salzmann, Pascal Fua |
ECCV (15) | 2 |
| 2020 | History Repeats Itself: Human Motion Prediction via Motion Attention
Wei Mao 0001, Miaomiao Liu 0001, Mathieu Salzmann |
ECCV (14) | 3 |
| 2020 | Indirect Local Attacks for Context-Aware Semantic Segmentation Networks
Krishna K. Nakka, Mathieu Salzmann |
ECCV (5) | 2 |
| 2020 | Domain Adaptive Multibranch Networks
Róger Bermúdez-Chacón, Mathieu Salzmann, Pascal Fua |
ICLR | 2 |
| 2020 | Evaluating The Search Phase of Neural Architecture Search
Kaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Cristian Musat, Mathieu Salzmann |
ICLR | 5 |
| 2020 | ExpandNets: Linear Over-parameterization to Train Compact Convolutional NetworksabstractWe introduce an approach to training a given compact network. To this end, we leverage over-parameterization, which typically improves both neural network optimization and generalization. Specifically, we propose to expand each linear layer of the compact network into multiple consecutive linear layers, without adding any nonlinearity. As such, the resulting expanded network, or ExpandNet, can be contracted back to the compact one algebraically at inference. In particular, we introduce two convolutional expansion strategies and demonstrate their benefits on several tasks, including image classification, object detection, and semantic segmentation. As evidenced by our experiments, our approach outperforms both training the compact network from scratch and performing knowledge distillation from a teacher. Furthermore, our linear over-parameterization empirically reduces gradient confusion during training and improves the network generalization. Shuxuan Guo, José M. Álvarez 0004, Mathieu Salzmann |
NeurIPS | 3 |
| 2020 | On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome ThemabstractWe analyze the influence of adversarial training on the loss landscape of machine learning models. To this end, we first provide analytical studies of the properties of adversarial loss functions under different adversarial budgets. We then demonstrate that the adversarial loss landscape is less favorable to optimization, due to increased curvature and more scattered gradients. Our conclusions are validated by numerical analyses, which show that training under large adversarial budgets impede the escape from suboptimal random initialization, cause non-vanishing gradients and make the models' minima found sharper. Based on these observations, we show that a periodic adversarial scheduling (PAS) strategy can effectively overcome these challenges, yielding better results than vanilla adversarial training while being much less sensitive to the choice of learning rate. Chen Liu 0027, Mathieu Salzmann, Tao Lin 0004, Ryota Tomioka, Sabine Süsstrunk |
NeurIPS | 2 |
| 2020 | Tracing in 2D to reduce the annotation effort for 3D deep delineation of linear structures
Mateusz Kozinski, Agata Mosinska, Mathieu Salzmann, Pascal Fua |
Medical Image Anal. | 3 |
| 2020 | Visual Correspondences for Unsupervised Domain Adaptation on Electron Microscopy ImagesabstractWe present an Unsupervised Domain Adaptation strategy to compensate for domain shifts on Electron Microscopy volumes. Our method aggregates visual correspondences-motifs that are visually similar across different acquisitions-to infer changes on the parameters of pretrained models, and enable them to operate on new data. In particular, we examine the annotations of an existing acquisition to determine pivot locations that characterize the reference segmentation, and use a patch matching algorithm to find their candidate visual correspondences in a new volume. We aggregate all the candidate correspondences by a voting scheme and we use them to construct a consensus heatmap: a map of how frequently locations on the new volume are matched to relevant locations from the original acquisition. This information allows us to perform model adaptations in two different ways: either by a) optimizing model parameters under a Multiple Instance Learning formulation, so that predictions between reference locations and their sets of correspondences agree, or by b) using high-scoring regions of the heatmap as soft labels to be incorporated in other domain adaptation pipelines, including deep learning ones. We show that these unsupervised techniques allow us to obtain high-quality segmentations on unannotated volumes, qualitatively consistent with results obtained under full supervision, for both mitochondria and synapses, with no need for new annotation effort. Róger Bermúdez-Chacón, Okan Altingövde, Carlos J. Becker, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Medical Imaging | 4 |
| 2019 | Segmentation-Driven 6D Object Pose EstimationabstractThe most recent trend in estimating the 6D pose of rigid objects has been to train deep networks to either directly regress the pose from the image or to predict the 2D locations of 3D keypoints, from which the pose can be obtained using a PnP algorithm. In both cases, the object is treated as a global entity, and a single pose estimate is computed. As a consequence, the resulting techniques can be vulnerable to large occlusions. In this paper, we introduce a segmentation-driven 6D pose estimation framework where each visible part of the objects contributes a local pose prediction in the form of 2D keypoint locations. We then use a predicted measure of confidence to combine these pose candidates into a robust set of 3D-to-2D correspondences, from which a reliable pose estimate can be obtained. We outperform the state-of-the-art on the challenging Occluded-LINEMOD and YCB-Video datasets, which is evidence that our approach deals well with multiple poorly-textured objects occluding each other. Furthermore, it relies on a simple enough architecture to achieve real-time performance. Yinlin Hu, Joachim Hugonot, Pascal Fua, Mathieu Salzmann |
CVPR | 4 |
| 2019 | Context-Aware Crowd CountingabstractState-of-the-art methods for counting people in crowded scenes rely on deep networks to estimate crowd density. They typically use the same filters over the whole image or over large image patches. Only then do they estimate local scale to compensate for perspective distortion. This is typically achieved by training an auxiliary classifier to select, for predefined image patches, the best kernel size among a limited set of choices. As such, these methods are not end-to-end trainable and restricted in the scope of context they can leverage. In this paper, we introduce an end-to-end trainable deep architecture that combines features obtained using multiple receptive field sizes and learns the importance of each such feature at each image location. In other words, our approach adaptively encodes the scale of the contextual information required to accurately predict crowd density. This yields an algorithm that outperforms state-of-the-art crowd counting methods, especially when perspective effects are strong. Weizhe Liu, Mathieu Salzmann, Pascal Fua |
CVPR | 2 |
| 2019 | Neural Scene Decomposition for Multi-Person Motion CaptureabstractLearning general image representations has proven key to the success of many computer vision tasks. For example, many approaches to image understanding problems rely on deep networks that were initially trained on ImageNet, mostly because the learned features are a valuable starting point to learn from limited labeled data. However, when it comes to 3D motion capture of multiple people, these features are only of limited use. In this paper, we therefore propose an approach to learning features that are useful for this purpose. To this end, we introduce a self-supervised approach to learning what we call a neural scene decomposition (NSD) that can be exploited for 3D pose estimation. NSD comprises three layers of abstraction to represent human subjects: spatial layout in terms of bounding-boxes and relative depth; a 2D shape representation in terms of an instance segmentation mask; and subject-specific appearance and 3D pose information. By exploiting self-supervision coming from multiview data, our NSD model can be trained end-to-end without any 2D or 3D supervision. In contrast to previous approaches, it works for multiple persons and full-frame images. Because it encodes 3D geometry, NSD can then be effectively leveraged to train a 3D pose estimation network from small amounts of annotated data. Helge Rhodin, Victor Constantin, Isinsu Katircioglu, Mathieu Salzmann, Pascal Fua |
CVPR | 4 |
| 2019 | Recurrent U-Net for Resource-Constrained SegmentationabstractState-of-the-art segmentation methods rely on very deep networks that are not always easy to train without very large training datasets and tend to be relatively slow to run on standard GPUs. In this paper, we introduce a novel recurrent U-Net architecture that preserves the compactness of the original U-Net [33], while substantially increasing its performance to the point where it outperforms the state of the art on several benchmarks. We will demonstrate its effectiveness for several tasks, including hand segmentation, retina vessel segmentation, and road segmentation. We also introduce a large-scale dataset for hand segmentation. Wei Wang 0108, Kaicheng Yu, Joachim Hugonot, Pascal Fua, Mathieu Salzmann |
ICCV | 5 |
| 2019 | GarNet: A Two-Stream Network for Fast and Accurate 3D Cloth DrapingabstractWhile Physics-Based Simulation (PBS) can accurately drape a 3D garment on a 3D body, it remains too costly for real-time applications, such as virtual try-on. By contrast, inference in a deep network, requiring a single forward pass, is much faster. Taking advantage of this, we propose a novel architecture to fit a 3D garment template to a 3D body. Specifically, we build upon the recent progress in 3D point cloud processing with deep networks to extract garment features at varying levels of detail, including point-wise, patch-wise and global features. We fuse these features with those extracted in parallel from the 3D body, so as to model the cloth-body interactions. The resulting two-stream architecture, which we call as GarNet, is trained using a loss function inspired by physics-based modeling, and delivers visually plausible garment shapes whose 3D points are, on average, less than 1 cm away from those of a PBS method, while running 100 times faster. Moreover, the proposed method can model various garment types with different cutting patterns when parameters of those patterns are given as input to the network. Erhan Gundogdu, Victor Constantin, Amrollah Seifoddini, Minh Dang, Mathieu Salzmann, Pascal Fua |
ICCV | 5 |
| 2019 | Detecting the Unexpected via Image ResynthesisabstractClassical semantic segmentation methods, including the recent deep learning ones, assume that all classes observed at test time have been seen during training. In this paper, we tackle the more realistic scenario where unexpected objects of unknown classes can appear at test time. The main trends in this area either leverage the notion of prediction uncertainty to flag the regions with low confidence as unknown, or rely on autoencoders and highlight poorly-decoded regions. Having observed that, in both cases, the detected regions typically do not correspond to unexpected objects, in this paper, we introduce a drastically different strategy: It relies on the intuition that the network will produce spurious labels in regions depicting unexpected objects. Therefore, resynthesizing the image from the resulting semantic map will yield significant appearance differences with respect to the input image. In other words, we translate the problem of detecting unknown classes to one of identifying poorly-resynthesized image regions. We show that this outperforms both uncertainty- and autoencoder-based methods. Krzysztof Lis, Krishna K. Nakka, Pascal Fua, Mathieu Salzmann |
ICCV | 4 |
| 2019 | Learning Trajectory Dependencies for Human Motion PredictionabstractHuman motion prediction, i.e., forecasting future body poses given observed pose sequence, has typically been tackled with recurrent neural networks (RNNs). However, as evidenced by prior work, the resulted RNN models suffer from prediction errors accumulation, leading to undesired discontinuities in motion prediction. In this paper, we propose a simple feed-forward deep network for motion prediction, which takes into account both temporal smoothness and spatial dependencies among human body joints. In this context, we then propose to encode temporal information by working in trajectory space, instead of the traditionally-used pose space. This alleviates us from manually defining the range of temporal dependencies (or temporal convolutional filter size, as done in previous work). Moreover, spatial dependency of human pose is encoded by treating a human pose as a generic graph (rather than a human skeletal kinematic tree) formed by links between every pair of body joints. Instead of using a pre-defined graph structure, we design a new graph convolutional network to learn graph connectivity automatically. This allows the network to capture long range dependencies beyond that of human kinematic tree. We evaluate our approach on several standard benchmark datasets for motion prediction, including Human3.6M, the CMU motion capture dataset and 3DPW. Our experiments clearly demonstrate that the proposed approach achieves state of the art performance, and is applicable to both angle-based and position-based pose representations. The code is available at https://github.com/wei-mao-2019/LearnTrajDep. Wei Mao 0001, Miaomiao Liu 0001, Mathieu Salzmann, Hongdong Li |
ICCV | 3 |
| 2019 | Field Typing for Improved Recognition on Heterogeneous Handwritten FormsabstractOffline handwriting recognition has undergone continuous progress over the past decades. However, existing methods are typically benchmarked on free-form text datasets that are biased towards good-quality images and handwriting styles, and homogeneous content. In this paper, we show that state-of-the-art algorithms, employing long short-term memory (LSTM) layers, do not readily generalize to real-world structured documents, such as forms, due to their highly heterogeneous and out-of-vocabulary content, and to the inherent ambiguities of this content. To address this, we propose to leverage the content type within an LSTM-based architecture. Furthermore, we introduce a procedure to generate synthetic data to train this architecture without requiring expensive manual annotations. We demonstrate the effectiveness of our approach at transcribing text on a challenging, real-world dataset of European Accident Statements. Ciprian Tomoiaga, Paul Feng, Mathieu Salzmann, Patrick Jayet |
ICDAR | 3 |
| 2019 | Learning Factorized Representations for Open-Set Domain Adaptation
Mahsa Baktash, Masoud Faraki, Tom Drummond, Mathieu Salzmann |
ICLR (Poster) | 4 |
| 2019 | Overcoming Multi-model ForgettingabstractWe identify a phenomenon, which we refer to as multi-model forgetting, that occurs when sequentially training multiple deep networks with partially-shared parameters; the performance of previously-trained models degrades as one optimizes a subsequent one, due to the overwriting of shared parameters. To overcome this, we introduce a statistically-justified weight plasticity loss that regularizes the learning of a model’s shared parameters according to their importance for the previous models, and demonstrate its effectiveness when training two models sequentially and for neural architecture search. Adding weight plasticity in neural architecture search preserves the best models to the end of the search and yields improved results in both natural language processing and computer vision tasks. Yassine Benyahia, Kaicheng Yu, Kamil Bennani-Smires, Martin Jaggi, Anthony C. Davison, Mathieu Salzmann, Claudiu Cristian Musat |
ICML | 6 |
| 2019 | Geometric and Physical Constraints for Drone-Based Head Plane Crowd Density EstimationabstractState-of-the-art methods for counting people in crowded scenes rely on deep networks to estimate crowd density in the image plane. While useful for this purpose, this image-plane density has no immediate physical meaning because it is subject to perspective distortion. This is a concern in sequences acquired by drones because the viewpoint changes often. This distortion is usually handled implicitly by either learning scale-invariant features or estimating density in patches of different sizes, neither of which accounts for the fact that scale changes must be consistent over the whole scene. In this paper, we explicitly model the scale changes and reason in terms of people per square-meter. We show that feeding the perspective model to the network allows us to enforce global scale consistency and that this model can be obtained on the fly from the drone sensors. In addition, it also enables us to enforce physically-inspired temporal consistency constraints that do not have to be learned. This yields an algorithm that outperforms state-of-the-art methods in inferring crowd density from a moving drone camera especially when perspective effects are strong. Weizhe Liu, Krzysztof Lis, Mathieu Salzmann, Pascal Fua |
IROS | 3 |
| 2019 | Backpropagation-Friendly EigendecompositionabstractEigendecomposition (ED) is widely used in deep networks. However, the backpropagation of its results tends to be numerically unstable, whether using ED directly or approximating it with the Power Iteration method, particularly when dealing with large matrices. While this can be mitigated by partitioning the data in small and arbitrary groups, doing so has no theoretical basis and makes its impossible to exploit the power of ED to the full. In this paper, we introduce a numerically stable and differentiable approach to leveraging eigenvectors in deep networks. It can handle large matrices without requiring to split them. We demonstrate the better robustness of our approach over standard ED and PI for ZCA whitening, an alternative to batch normalization, and for PCA denoising, which we introduce as a new normalization strategy for deep networks, aiming to further denoise the network's features. Wei Wang 0108, Zheng Dang, Yinlin Hu, Pascal Fua, Mathieu Salzmann |
NeurIPS | 5 |
| 2019 | Memory Efficient Max Flow for Multi-Label Submodular MRFsabstractMulti-label submodular Markov Random Fields (MRFs) have been shown to be solvable using max-flow based on an encoding of the labels proposed by Ishikawa, in which each variable$X_i$is represented by$\ell$nodes (where$\ell$is the number of labels) arranged in a column. However, this method in general requires$2\;\ell ^2$edges for each pair of neighbouring variables. This makes it inapplicable to realistic problems with many variables and labels, due to excessive memory requirement. In this paper, we introduce a variant of the max-flow algorithm that requires much less storage. Consequently, our algorithm makes it possible to optimally solve multi-label submodular problems involving large numbers of variables and labels on a standard computer. Thalaiyasingam Ajanthan, Richard I. Hartley, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Beyond Sharing Weights for Deep Domain AdaptationabstractThe performance of a classifier trained on data coming from a specific domain typically degrades when applied to a related but different one. While annotating many samples from the new domain would address this issue, it is often too expensive or impractical. Domain Adaptation has therefore emerged as a solution to this problem; It leverages annotated data from a source domain, in which it is abundant, to train a classifier to operate in a target domain, in which it is either sparse or even lacking altogether. In this context, the recent trend consists of learning deep architectures whose weights are shared for both domains, which essentially amounts to learning domain invariant features. Here, we show that it is more effective to explicitly model the shift from one domain to the other. To this end, we introduce a two-stream architecture, where one operates in the source domain and the other in the target domain. In contrast to other approaches, the weights in corresponding layers are related but not shared. We demonstrate that this both yields higher accuracy than state-of-the-art methods on several object recognition and detection tasks and consistently outperforms networks with shared weights in both supervised and unsupervised settings. Artem Rozantsev, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Efficient Relaxations for Dense CRFs with Sparse Higher-Order PotentialsabstractDense conditional random fields (CRFs) have become a popular framework for modeling several problems in computer vision such as stereo correspondence and multiclass semantic segmentation. By modeling long-range interactions, dense CRFs provide a labeling that captures finer detail than their sparse counterparts. Currently, the state-of-the-art algorithm performs mean-field inference using a filter-based method but fails to provide a strong theoretical guarantee on the quality of the solution. A question naturally arises as to whether it is possible to obtain a maximum a posteriori (MAP) estimate of a dense CRF using a principled method. Within this paper, we show that this is indeed possible. Specifically, we will show that, by using a filter-based method, continuous relaxations of the MAP problem can be optimized efficiently using state-of-the-art algorithms. Specifically, we will solve a quadratic programming relaxation using the Frank--Wolfe algorithm and a linear programming relaxation by developing a proximal minimization framework. By exploiting labeling consistency in the higher-order potentials and utilizing the filter-based method, we are able to formulate the above algorithms such that each iteration has a complexity linear in the number of classes and random variables. The presented algorithms can be applied to any labeling problem using a dense CRF with sparse higher-order potentials. In this paper, we use semantic segmentation as an example application as it demonstrates the ability of the algorithm to scale to dense CRFs with large dimensions. We perform experiments on the Pascal dataset to indicate that the presented algorithms are able to attain lower energies than the mean-field inference method. Thomas Joy, Alban Desmaison, Thalaiyasingam Ajanthan, Rudy Bunel, Mathieu Salzmann, Pushmeet Kohli, Philip Torr 0001, M. Pawan Kumar |
SIAM J. Imaging Sci. | 5 |
| 2018 | Learning to Reconstruct Texture-Less Deformable Surfaces from a Single ViewabstractRecent years have seen the development of mature solutions for reconstructing deformable surfaces from a single image, provided that they are relatively well-textured. By contrast, recovering the 3D shape of texture-less surfaces remains an open problem, and essentially relates to Shape-from-Shading. In this paper, we introduce a data-driven approach to this problem. We introduce a general framework that can predict diverse 3D representations, such as meshes, normals, and depth maps. Our experiments show that meshes are ill-suited to handle texture-less 3D reconstruction in our context. Furthermore, we demonstrate that our approach generalizes well to unseen objects, and that it yields higher-quality reconstructions than a state-of-the-art SfS technique, particularly in terms of normal estimates. Our reconstructions accurately model the fine details of the surfaces, such as the creases of a T-Shirt worn by a person. Jan Bednarík, Pascal Fua, Mathieu Salzmann |
3DV | 3 |
| 2018 | 3D Box Proposals From a Single Monocular Image of an Indoor SceneabstractModern object detection methods typically rely on bounding box proposals as input. While initially popularized in the 2D case, this idea has received increasing attention for 3D bounding boxes. Nevertheless, existing 3D box proposal techniques all assume having access to depth as input, which is unfortunately not always available in practice. In this paper, we therefore introduce an approach to generating 3D box proposals from a single monocular RGB image. To this end, we develop an integrated, fully differentiable framework that inherently predicts a depth map, extracts a 3D volumetric scene representation and generates 3D object proposals. At the core of our approach lies a novel residual, differentiable truncated signed distance function module, which, accounting for the relatively low accuracy of the predicted depth map, extracts a 3D volumetric representation of the scene. Our experiments on the standard NYUv2 dataset demonstrate that our framework lets us generate high-quality 3D box proposals and that it outperforms the two-stage technique consisting of successively performing state-of-the-art depth prediction and depth-based 3D proposal generation. Wei Zhuo 0004, Mathieu Salzmann, Xuming He 0001, Miaomiao Liu 0001 |
AAAI | 2 |
| 2018 | VIENA ^2 : A Driving Anticipation Dataset
Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Basura Fernando, Lars Petersson, Lars Andersson |
ACCV (1) | 3 |
| 2018 | Deep Attentional Structured Representation Learning for Visual Recognition
Krishna K. Nakka, Mathieu Salzmann |
BMVC | 2 |
| 2018 | Geometry-Aware Deep Network for Single-Image Novel View SynthesisabstractThis paper tackles the problem of novel view synthesis from a single image. In particular, we target real-world scenes with rich geometric structure, a challenging task due to the large appearance variations of such scenes and the lack of simple 3D models to represent them. Modern, learning-based approaches mostly focus on appearance to synthesize novel views and thus tend to generate predictions that are inconsistent with the underlying scene structure. By contrast, in this paper, we propose to exploit the 3D geometry of the scene to synthesize a novel view. Specifically, we approximate a real-world scene by a fixed number of planes, and learn to predict a set of homographies and their corresponding region masks to transform the input image into a novel view. To this end, we develop a new region-aware geometric transform network that performs these multiple tasks in a common framework. Our results on the outdoor KITTI and the indoor ScanNet datasets demonstrate the effectiveness of our network in generating high-quality synthetic views that respect the scene geometry, thus outperforming the state-of-the-art methods. Miaomiao Liu 0001, Xuming He 0001, Mathieu Salzmann |
CVPR | 3 |
| 2018 | Learning Monocular 3D Human Pose Estimation From Multi-View ImagesabstractAccurate 3D human pose estimation from single images is possible with sophisticated deep-net architectures that have been trained on very large datasets. However, this still leaves open the problem of capturing motions for which no such database exists. Manual annotation is tedious, slow, and error-prone. In this paper, we propose to replace most of the annotations by the use of multiple views, at training time only. Specifically, we train the system to predict the same pose in all views. Such a consistency constraint is necessary but not sufficient to predict accurate poses. We therefore complement it with a supervised loss aiming to predict the correct pose in a small set of labeled images, and with a regularization term that penalizes drift from initial predictions. Furthermore, we propose a method to estimate camera pose jointly with human pose, which lets us utilize multiview footage where calibration is difficult, e.g., for pan-tilt or moving handheld cameras. We demonstrate the effectiveness of our approach on established benchmarks, as well as on a new Ski dataset with rotating cameras and expert ski motion, for which annotations are truly hard to obtain. Helge Rhodin, Jörg Spörri, Isinsu Katircioglu, Victor Constantin, Frédéric Meyer, Erich Müller, Mathieu Salzmann, Pascal Fua |
CVPR | 7 |
| 2018 | Residual Parameter Transfer for Deep Domain AdaptationabstractThe goal of Deep Domain Adaptation is to make it possible to use Deep Nets trained in one domain where there is enough annotated training data in another where there is little or none. Most current approaches have focused on learning feature representations that are invariant to the changes that occur when going from one domain to the other, which means using the same network parameters in both domains. While some recent algorithms explicitly model the changes by adapting the network parameters, they either severely restrict the possible domain changes, or significantly increase the number of model parameters. By contrast, we introduce a network architecture that includes auxiliary residual networks, which we train to predict the parameters in the domain with little annotated data from those in the other one. This architecture enables us to flexibly preserve the similarities between domains where they exist and model the differences when necessary. We demonstrate that our approach yields higher accuracy than state-of-the-art methods without undue complexity. Artem Rozantsev, Mathieu Salzmann, Pascal Fua |
CVPR | 2 |
| 2018 | Learning to Find Good CorrespondencesabstractWe develop a deep architecture to learn to find good correspondences for wide-baseline stereo. Given a set of putative sparse matches and the camera intrinsics, we train our network in an end-to-end fashion to label the correspondences as inliers or outliers, while simultaneously using them to recover the relative pose, as encoded by the essential matrix. Our architecture is based on a multi-layer perceptron operating on pixel coordinates rather than directly on the image, and is thus simple and small. We introduce a novel normalization technique, called Context Normalization, which allows us to process each data point separately while embedding global information in it, and also makes the network invariant to the order of the correspondences. Our experiments on multiple challenging datasets demonstrate that our method is able to drastically improve the state of the art with little training data. Kwang Moo Yi, Eduard Trulls, Yuki Ono, Vincent Lepetit, Mathieu Salzmann, Pascal Fua |
CVPR | 5 |
| 2018 | Eigendecomposition-Free Training of Deep Networks with Zero Eigenvalue-Based Losses
Zheng Dang, Kwang Moo Yi, Yinlin Hu, Fei Wang 0008, Pascal Fua, Mathieu Salzmann |
ECCV (5) | 6 |
| 2018 | Unsupervised Geometry-Aware Representation for 3D Human Pose EstimationabstractModern 3D human pose estimation techniques rely on deep networks, which require large amounts of training data. While weakly-supervised methods require less supervision, by utilizing 2D poses or multi-view imagery without annotations, they still need a sufficiently large set of samples with 3D annotations for learning to succeed. In this paper, we propose to overcome this problem by learning a geometry-aware body representation from multi-view images without annotations. To this end, we use an encoder-decoder that predicts an image from one viewpoint given an image from another viewpoint. Because this representation encodes 3D geometry, using it in a semi-supervised setting makes it easier to learn a mapping from it to 3D human pose. As evidenced by our experiments, our approach significantly outperforms fully-supervised methods given the same amount of labeled data, and improves over other semi-supervised methods while using as little as 1% of the labeled data. Helge Rhodin, Mathieu Salzmann, Pascal Fua |
ECCV (10) | 2 |
| 2018 | Effective Use of Synthetic Data for Urban Scene Semantic Segmentation
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004 |
ECCV (2) | 3 |
| 2018 | Statistically-Motivated Second-Order Pooling
Kaicheng Yu, Mathieu Salzmann |
ECCV (7) | 2 |
| 2018 | Learning to Segment 3D Linear Structures Using Only 2D Annotations
Mateusz Kozinski, Agata Mosinska, Mathieu Salzmann, Pascal Fua |
MICCAI (2) | 3 |
| 2018 | Learning Latent Representations of 3D Human Pose with Deep Neural Networks
Isinsu Katircioglu, Bugra Tekin, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
Int. J. Comput. Vis. | 3 |
| 2018 | Dimensionality Reduction on SPD Manifolds: The Emergence of Geometry-Aware MethodsabstractRepresenting images and videos with Symmetric Positive Definite (SPD) matrices, and considering the Riemannian geometry of the resulting space, has been shown to yield high discriminative power in many visual recognition tasks. Unfortunately, computation on the Riemannian manifold of SPD matrices -especially of high-dimensional ones- comes at a high cost that limits the applicability of existing techniques. In this paper, we introduce algorithms able to handle high-dimensional SPD matrices by constructing a lower-dimensional SPD manifold. To this end, we propose to model the mapping from the high-dimensional SPD manifold to the low-dimensional one with an orthonormal projection. This lets us formulate dimensionality reduction as the problem of finding a projection that yields a low-dimensional manifold either with maximum discriminative power in the supervised scenario, or with maximum variance of the data in the unsupervised one. We show that learning can be expressed as an optimization problem on a Grassmann manifold and discuss fast solutions for special cases. Our evaluation on several classification tasks evidences that our approach leads to a significant accuracy gain over state-of-the-art methods. Mehrtash Harandi, Mathieu Salzmann, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Incorporating Network Built-in Priors in Weakly-Supervised Semantic SegmentationabstractPixel-level annotations are expensive and time consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recently, CNN-based methods have proposed to fine-tune pre-trained networks using image tags. Without additional information, this leads to poor localization accuracy. This problem, however, was alleviated by making use of objectness priors to generate foreground/background masks. Unfortunately these priors either require pixel-level annotations/bounding boxes, or still yield inaccurate object boundaries. Here, we propose a novel method to extract accurate masks from networks pre-trained for the task of object recognition, thus forgoing external objectness modules. We first show how foreground/background masks can be obtained from the activations of higher-level convolutional layers of a network. We then show how to obtain multi-class masks by the fusion of foreground/background ones with information extracted from a weakly-supervised localization network. Our experiments evidence that exploiting these masks in conjunction with a weakly-supervised training loss yields state-of-the-art tag-based weakly-supervised semantic segmentation results. Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004, Stephen Gould |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Efficient Model-Free Anthropometry from Depth DataabstractExisting depth-based approaches to predicting anthropometric measurements, such as body height, arm span and hip circumference, either directly compute the measurements on 3D point clouds, and thus are sensitive to noise, or fit a model to the observed depth values, which typically is time-consuming. In this paper, we rely on the intuition that, to predict a specific anthropometric measurement, one does not need to have detailed information about the entire body shape. We therefore introduce an approach to anthropometry based on a random regression forest trained from local depth cues. The local predictions are then accumulated into one global, image-level anthropometric measurement prediction. We introduce a forest refinement scheme, whose objective function directly relies on both the image-level prediction, as well as on the local predictions' reliability. The resulting approach has the advantage of being both computationally highly efficient and accurate. Thomas Probst, Andrea Fossati, Mathieu Salzmann, Luc Van Gool |
3DV | 3 |
| 2017 | Efficient Linear Programming for Dense CRFsabstractThe fully connected conditional random field (CRF) with Gaussian pairwise potentials has proven popular and effective for multi-class semantic segmentation. While the energy of a dense CRF can be minimized accurately using a linear programming (LP) relaxation, the state-of-the-art algorithm is too slow to be useful in practice. To alleviate this deficiency, we introduce an efficient LP minimization algorithm for dense CRFs. To this end, we develop a proximal minimization framework, where the dual of each proximal problem is optimized via block coordinate descent. We show that each block of variables can be efficiently optimized. Specifically, for one block, the problem decomposes into significantly smaller subproblems, each of which is defined over a single pixel. For the other block, the problem is optimized via conditional gradient descent. This has two advantages: 1) the conditional gradient can be computed in a time linear in the number of pixels and labels, and 2) the optimal step size can be computed analytically. Our experiments on standard datasets provide compelling evidence that our approach outperforms all existing baselines including the previous LP based approach for dense CRFs. Thalaiyasingam Ajanthan, Alban Desmaison, Rudy Bunel, Mathieu Salzmann, Philip Torr 0001, M. Pawan Kumar |
CVPR | 4 |
| 2017 | Boundary-Aware Instance SegmentationabstractWe address the problem of instance-level semantic segmentation, which aims at jointly detecting, segmenting and classifying every individual object in an image. In this context, existing methods typically propose candidate objects, usually as bounding boxes, and directly predict a binary mask within each such proposal. As a consequence, they cannot recover from errors in the object candidate generation process, such as too small or shifted boxes. In this paper, we introduce a novel object segment representation based on the distance transform of the object masks. We then design an object mask network (OMN) with a new residual-deconvolution architecture that infers such a representation and decodes it into the final binary object mask. This allows us to predict masks that go beyond the scope of the bounding boxes and are thus robust to inaccurate object candidates. We integrate our OMN into a Multitask Network Cascade framework, and learn the resulting boundary-aware instance segmentation (BAIS) network in an end-to-end manner. Our experiments on the PASCAL VOC 2012 and the Cityscapes datasets demonstrate the benefits of our approach, which outperforms the state-of-the-art in both object proposal generation and instance segmentation. Zeeshan Hayder, Xuming He 0001, Mathieu Salzmann |
CVPR | 3 |
| 2017 | Indoor Scene Parsing with Instance Segmentation, Semantic Labeling and Support Relationship InferenceabstractOver the years, indoor scene parsing has attracted a growing interest in the computer vision community. Existing methods have typically focused on diverse subtasks of this challenging problem. In particular, while some of them aim at segmenting the image into regions, such as object or surface instances, others aim at inferring the semantic labels of given regions, or their support relationships. These different tasks are typically treated as separate ones. However, they bear strong connections: good regions should respect the semantic labels, support can only be defined for meaningful regions, support relationships strongly depend on semantics. In this paper, we therefore introduce an approach to jointly segment the instances and infer their semantic labels and support relationships from a single input image. By exploiting a hierarchical segmentation, we formulate our problem as that of jointly finding the regions in the hierarchy that correspond to instances and estimating their class labels and pairwise support relationships. We express this via a Markov Random Field, which allows us to further encode links between the different types of variables. Inference in this model can be done exactly via integer linear programming, and we learn its parameters in a structural SVM framework. Our experiments on NYUv2 demonstrate the benefits of reasoning jointly about all these subtasks of indoor scene parsing. Wei Zhuo 0004, Mathieu Salzmann, Xuming He 0001, Miaomiao Liu 0001 |
CVPR | 2 |
| 2017 | Encouraging LSTMs to Anticipate Actions Very EarlyabstractIn contrast to the widely studied problem of recognizing an action given a complete sequence, action anticipation aims to identify the action from only partially available videos. As such, it is therefore key to the success of computer vision applications requiring to react as early as possible, such as autonomous navigation. In this paper, we propose a new action anticipation method that achieves high prediction accuracy even in the presence of a very small percentage of a video sequence. To this end, we develop a multi-stage LSTM architecture that leverages context-aware and action-aware features, and introduce a novel loss function that encourages the model to predict the correct class as early as possible. Our experiments on standard benchmark datasets evidence the benefits of our approach; We outperform the state-of-the-art action anticipation methods for early prediction by a relative increase in accuracy of 22.0% on JHMDB-21, 14.0% on UT-Interaction and 49.9% on UCF-101. Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Basura Fernando, Lars Petersson, Lars Andersson |
ICCV | 3 |
| 2017 | Bringing Background into the Foreground: Making All Classes Equal in Weakly-Supervised Video Semantic SegmentationabstractPixel-level annotations are expensive and time-consuming to obtain. Hence, weak supervision using only image tags could have a significant impact in semantic segmentation. Recent years have seen great progress in weakly-supervised semantic segmentation, whether from a single image or from videos. However, most existing methods are designed to handle a single background class. In practical applications, such as autonomous navigation, it is often crucial to reason about multiple background classes. In this paper, we introduce an approach to doing so by making use of classifier heatmaps. We then develop a two-stream deep architecture that jointly leverages appearance and motion, and design a loss based on our heatmaps to train it. Our experiments demonstrate the benefits of our classifier heatmaps and of our two-stream architecture on challenging urban scene datasets and on the YouTube-Objects benchmark, where we obtain state-of-the-art results. Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, José M. Álvarez 0004 |
ICCV | 3 |
| 2017 | Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose EstimationabstractMost recent approaches to monocular 3D human pose estimation rely on Deep Learning. They typically involve regressing from an image to either 3D joint coordinates directly or 2D joint locations from which 3D coordinates are inferred. Both approaches have their strengths and weaknesses and we therefore propose a novel architecture designed to deliver the best of both worlds by performing both simultaneously and fusing the information along the way. At the heart of our framework is a trainable fusion scheme that learns how to fuse the information optimally instead of being hand-designed. This yields significant improvements upon the state-of-the-art on standard 3D human pose estimation benchmarks. Bugra Tekin, Pablo Márquez-Neila, Mathieu Salzmann, Pascal Fua |
ICCV | 3 |
| 2017 | Joint Dimensionality Reduction and Metric Learning: A Geometric TakeabstractTo be tractable and robust to data noise, existing metric learning algorithms commonly rely on PCA as a pre-processing step. How can we know, however, that PCA, or any other specific dimensionality reduction technique, is the method of choice for the problem at hand? The answer is simple: We cannot! To address this issue, in this paper, we develop a Riemannian framework to jointly learn a mapping performing dimensionality reduction and a metric in the induced space. Our experiments evidence that, while we directly work on high-dimensional features, our approach yields competitive runtimes with and higher accuracy than state-of-the-art metric learning algorithms. Mehrtash Harandi, Mathieu Salzmann, Richard I. Hartley |
ICML | 2 |
| 2017 | Compression-aware Training of Deep NetworksabstractIn recent years, great progress has been made in a variety of application domains thanks to the development of increasingly deeper neural networks. Unfortunately, the huge number of units of these networks makes them expensive both computationally and memory-wise. To overcome this, exploiting the fact that deep networks are over-parametrized, several compression strategies have been proposed. These methods, however, typically start from a network that has been trained in a standard manner, without considering such a future compression. In this paper, we propose to explicitly account for compression in the training process. To this end, we introduce a regularizer that encourages the parameter matrix of each layer to have low rank during training. We show that accounting for compression during training allows us to learn much more compact, yet at least as effective, models than state-of-the-art compression techniques. José M. Álvarez 0004, Mathieu Salzmann |
NIPS | 2 |
| 2017 | Deep Subspace Clustering NetworksabstractWe present a novel deep neural network architecture for unsupervised subspace clustering. This architecture is built upon deep auto-encoders, which non-linearly map the input data into a latent space. Our key idea is to introduce a novel self-expressive layer between the encoder and the decoder to mimic the "self-expressiveness" property that has proven effective in traditional subspace clustering. Being differentiable, our new self-expressive layer provides a simple but effective way to learn pairwise affinities between all data points through a standard back-propagation procedure. Being nonlinear, our neural-network based method is able to cluster data points having complex (often nonlinear) structures. We further propose pre-training and fine-tuning strategies that let us effectively learn the parameters of our subspace clustering networks. Our experiments show that the proposed method significantly outperforms the state-of-the-art unsupervised subspace clustering methods. Pan Ji, Tong Zhang 0023, Hongdong Li, Mathieu Salzmann, Ian D. Reid 0001 |
NIPS | 4 |
| 2016 | Structured Prediction of 3D Human Pose with Deep Neural Networks
Bugra Tekin, Isinsu Katircioglu, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
BMVC | 3 |
| 2016 | Memory Efficient Max Flow for Multi-label Submodular MRFsabstractMulti-label submodular Markov Random Fields (MRFs) have been shown to be solvable using max-flow based on an encoding of the labels proposed by Ishikawa, in which each variable Xi is represented by l nodes (where l is the number of labels) arranged in a column. However, this method in general requires 2 l2 edges for each pair of neighbouring variables. This makes it inapplicable to realistic problems with many variables and labels, due to excessive memory requirement. In this paper, we introduce a variant of the max-flow algorithm that requires much less storage. Consequently, our algorithm makes it possible to optimally solve multi-label submodular problems involving large numbers of variables and labels on a standard computer. Thalaiyasingam Ajanthan, Richard I. Hartley, Mathieu Salzmann |
CVPR | 3 |
| 2016 | When VLAD Met HilbertabstractIn many challenging visual recognition tasks where training data is limited, Vectors of Locally Aggregated Descriptors (VLAD) have emerged as powerful image/video representations that compete with or outperform state-of the-art approaches. In this paper, we address two fundamental limitations of VLAD: its requirement for the local descriptors to have vector form and its restriction to linear classifiers due to its high-dimensionality. To this end, we introduce a kernelized version of VLAD. This not only lets us inherently exploit more sophisticated classification schemes, but also enables us to efficiently aggregate nonvector descriptors (e.g., manifold-valued data) in the VLAD framework. Furthermore, we propose an approximate formulation that allows us to accelerate the coding process while still benefiting from the properties of kernel VLAD. Our experiments demonstrate the effectiveness of our approach at handling manifold-valued data, such as covariance descriptors, on several classification tasks. Our results also evidence the benefits of our nonlinear VLAD descriptors against the linear ones in Euclidean space using several standard benchmark datasets. Mehrtash Harandi, Mathieu Salzmann, Fatih Porikli |
CVPR | 2 |
| 2016 | Learning to Co-Generate Object Proposals with a Deep Structured NetworkabstractGenerating object proposals has become a key component of modern object detection pipelines. However, most existing methods generate the object candidates independently of each other. In this paper, we present an approach to co-generating object proposals in multiple images, thus leveraging the collective power of multiple object candidates. In particular, we introduce a deep structured network that jointly predicts the objectness scores and the bounding box locations of multiple object candidates. Our deep structured network consists of a fully-connected Conditional Random Field built on top of a set of deep Convolutional Neural Networks, which learn features to model both the individual object candidates and the similarity between multiple candidates. To train our deep structured network, we develop an end-to-end learning algorithm that, by unrolling the CRF inference procedure, lets us backpropagate the loss gradient throughout the entire structured network. We demonstrate the effectiveness of our approach on two benchmark datasets, showing significant improvement over state-of-the-art object proposal algorithms. Zeeshan Hayder, Xuming He 0001, Mathieu Salzmann |
CVPR | 3 |
| 2016 | Robust Multi-Body Feature Tracker: A Segmentation-Free ApproachabstractFeature tracking is a fundamental problem in computer vision, with applications in many computer vision tasks, such as visual SLAM and action recognition. This paper introduces a novel multi-body feature tracker that exploits a multi-body rigidity assumption to improve tracking robustness under a general perspective camera model. A conventional approach to addressing this problem would consist of alternating between solving two subtasks: motion segmentation and feature tracking under rigidity constraints for each segment. This approach, however, requires knowing the number of motions, as well as assigning points to motion groups, which is typically sensitive to the motion estimates. By contrast, here, we introduce a segmentationfree solution to multi-body feature tracking that bypasses the motion assignment step and reduces to solving a series of subproblems with closed-form solutions. Our experiments demonstrate the benefits of our approach in terms of tracking accuracy and robustness to noise. Pan Ji, Hongdong Li, Mathieu Salzmann, Yiran Zhong |
CVPR | 3 |
| 2016 | Sample and Filter: Nonparametric Scene Parsing via Efficient FilteringabstractScene parsing has attracted a lot of attention in computer vision. While parametric models have proven effective for this task, they cannot easily incorporate new training data. By contrast, nonparametric approaches, which bypass any learning phase and directly transfer the labels from the training data to the query images, can readily exploit new labeled samples as they become available. Unfortunately, because of the computational cost of their label transfer procedures, state-of-the-art nonparametric methods typically filter out most training images to only keep a few relevant ones to label the query. As such, these methods throw away many images that still contain valuable information and generally obtain an unbalanced set of labeled samples. In this paper, we introduce a nonparametric approach to scene parsing that follows a sample-andfilter strategy. More specifically, we propose to sample labeled superpixels according to an image similarity score, which allows us to obtain a balanced set of samples. We then formulate label transfer as an efficient filtering procedure, which lets us exploit more labeled samples than existing techniques. Our experiments evidence the benefits of our approach over state-of-the-art nonparametric methods on two benchmark datasets. Mohammad Najafi, Sarah Taghavi Namin, Mathieu Salzmann, Lars Petersson |
CVPR | 3 |
| 2016 | Building Scene Models by Completing and Hallucinating Depth and Semantics
Miaomiao Liu 0001, Xuming He 0001, Mathieu Salzmann |
ECCV (6) | 3 |
| 2016 | Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Mathieu Salzmann, Lars Petersson, Stephen Gould, José M. Álvarez 0004 |
ECCV (8) | 3 |
| 2016 | Template-Free 3D Reconstruction of Poorly-Textured Nonrigid Surfaces
Xuan Wang 0009, Mathieu Salzmann, Fei Wang 0008, Jizhong Zhao |
ECCV (7) | 2 |
| 2016 | Scalable Unsupervised Domain Adaptation for Electron Microscopy
Róger Bermúdez-Chacón, Carlos J. Becker, Mathieu Salzmann, Pascal Fua |
MICCAI (2) | 3 |
| 2016 | Learning the Number of Neurons in Deep NetworksabstractNowadays, the number of layers and of neurons in each layer of a deep network are typically set manually. While very deep and wide networks have proven effective in general, they come at a high memory and computation cost, thus making them impractical for constrained platforms. These networks, however, are known to have many redundant parameters, and could thus, in principle, be replaced by more compact architectures. In this paper, we introduce an approach to automatically determining the number of neurons in each layer of a deep network during learning. To this end, we propose to make use of a group sparsity regularizer on the parameters of the network, where each group is defined to act on a single neuron. Starting from an overcomplete network, we show that our approach can reduce the number of parameters by up to 80\% while retaining or even improving the network accuracy. José M. Álvarez 0004, Mathieu Salzmann |
NIPS | 2 |
| 2016 | Efficient transductive semantic segmentationabstractSemantically describing the contents of images is one of the classical problems of computer vision. With huge numbers of images being made available daily, there is increasing interest in methods for semantic pixel labelling that exploit large image sets. Graph transduction provides a framework for the flexible inclusion of labeled data that can be exploited in the classification of unlabeled samples without requiring a trained classifier. Unfortunately, current approaches lack the scalability to tackle the joint segmentation of large image sets. Here we introduce an efficient flexible graph transduction approach to semantic segmentation that allows simple and efficient leveraging of large image sets without requiring separate computation of unary potentials, or a trained classifier. We demonstrate that this technique can handle far larger graphs than previous methods, and that results continue to improve as more labeled images are made available. Furthermore, we show that the method is able to benefit from dense or sparse unary labels when they are available. José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
WACV | 2 |
| 2016 | Semantic labeling for prosthetic vision
Lachlan Horne, José M. Álvarez 0004, Chris McCarthy, Mathieu Salzmann, Nick Barnes |
Comput. Vis. Image Underst. | 4 |
| 2016 | Distribution-Matching Embedding for Visual Domain AdaptationabstractDomain-invariant representations are key to addressing the domain shift problem where the training and test examples follow different distributions. Existing techniques that have attempted to match the distributions of the source and target domains typically compare these distributions in the original feature space. This space, however, may not be directly suitable for such a comparison, since some of the features may have been distorted by the domain shift, or may be domain specific. In this paper, we introduce a Distribution-Matching Embedding approach: An unsupervised domain adaptation method that overcomes this issue by mapping the data to a latent space where the distance between the empirical distributions of the source and target examples is minimized. In other words, we seek to extract the information that is invariant across the source and target data. In particular, we study two different distances to compare the source and target distributions: the Maximum Mean Discrepancy and the Hellinger distance. Furthermore, we show that our approach allows us to learn either a linear embedding, or a nonlinear one. We demonstrate the benefits of our approach on the tasks of visual object recognition, text categorization, and WiFi localization. Mahsa Baktash, Mehrtash Harandi, Mathieu Salzmann |
J. Mach. Learn. Res. | 3 |
| 2016 | Exploiting Large Image Sets for Road Scene ParsingabstractThere is an increasing interest in exploiting multiple images for scene understanding, with great progress in areas such as cosegmentation and video segmentation. Jointly analyzing the images in a large set offers the opportunity to exploit a greater source of information than when considering a single image on its own. However, this also yields challenges since, to effectively exploit all the available information, the resulting methods need to consider not just local connections, but efficiently analyze similarity between all pairs of pixels within and across all the images. In this paper, we propose to model an image set as a fully connected pairwise Conditional Random Field (CRF) defined over the image pixels, or superpixels, with Gaussian edge potentials. We show that this lets us co-label the images of a large set efficiently, thus yielding increased accuracy at no additional computational cost compared to sequential labeling of the images. Furthermore, we extend our framework to incorporate temporal dependence, thus effectively encompassing video segmentation as a special case of our approach, as well as to modeling label dependence over larger image regions. Our experimental evaluation demonstrates that our framework lets us handle over 10 000 images in a matter of seconds. José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2015 | Iteratively reweighted graph cut for multi-label MRFs with non-convex priorsabstractWhile widely acknowledged as highly effective in computer vision, multi-label MRFs with non-convex priors are difficult to optimize. To tackle this, we introduce an algorithm that iteratively approximates the original energy with an appropriately weighted surrogate energy that is easier to minimize. Our algorithm guarantees that the original energy decreases at each iteration. In particular, we consider the scenario where the global minimizer of the weighted surrogate energy can be obtained by a multi-label graph cut algorithm, and show that our algorithm then lets us handle of large variety of non-convex priors. We demonstrate the benefits of our method over state-of-the-art MRF energy minimization techniques on stereo and inpainting problems. Thalaiyasingam Ajanthan, Richard I. Hartley, Mathieu Salzmann, Hongdong Li |
CVPR | 3 |
| 2015 | Riemannian coding and dictionary learning: Kernels to the rescueabstractWhile sparse coding on non-flat Riemannian manifolds has recently become increasingly popular, existing solutions either are dedicated to specific manifolds, or rely on optimization problems that are difficult to solve, especially when it comes to dictionary learning. In this paper, we propose to make use of kernels to perform coding and dictionary learning on Riemannian manifolds. To this end, we introduce a general Riemannian coding framework with its kernel-based counterpart. This lets us (i) generalize beyond the special case of sparse coding; (ii) introduce efficient solutions to two coding schemes; (iii) learn the kernel parameters; (iv) perform unsupervised and supervised dictionary learning in a much simpler manner than previous Riemannian coding methods. We demonstrate the effectiveness of our approach on three different types of non-flat manifolds, and illustrate its generality by applying it to Euclidean spaces, which also are Riemannian manifolds. Mehrtash Harandi, Mathieu Salzmann |
CVPR | 2 |
| 2015 | Indoor scene structure analysis for single image depth estimationabstractWe tackle the problem of single image depth estimation, which, without additional knowledge, suffers from many ambiguities. Unlike previous approaches that only reason locally, we propose to exploit the global structure of the scene to estimate its depth. To this end, we introduce a hierarchical representation of the scene, which models local depth jointly with mid-level and global scene structures. We formulate single image depth estimation as inference in a graphical model whose edges let us encode the interactions within and across the different layers of our hierarchy. Our method therefore still produces detailed depth estimates, but also leverages higher-level information about the scene. We demonstrate the benefits of our approach over local depth estimation methods on standard indoor datasets. Wei Zhuo 0004, Mathieu Salzmann, Xuming He 0001, Miaomiao Liu 0001 |
CVPR | 2 |
| 2015 | Beyond Gauss: Image-Set Matching on the Riemannian Manifold of PDFsabstractState-of-the-art image-set matching techniques typically implicitly model each image-set with a Gaussian distribution. Here, we propose to go beyond these representations and model image-sets as probability distribution functions (PDFs) using kernel density estimators. To compare and match image-sets, we exploit Csiszar f-divergences, which bear strong connections to the geodesic distance defined on the space of PDFs, i.e., the statistical manifold. Furthermore, we introduce valid positive definite kernels on the statistical manifolds, which let us make use of more powerful classification schemes to match image-sets. Finally, we introduce a supervised dimensionality reduction technique that learns a latent space where f-divergences reflect the class labels of the data. Our experiments on diverse problems, such as video-based face recognition and dynamic texture classification, evidence the benefits of our approach over the state-of-the-art image-set matching methods. Mehrtash Harandi, Mathieu Salzmann, Mahsa Baktash |
ICCV | 2 |
| 2015 | Structural Kernel Learning for Large Scale Multiclass Object Co-detectionabstractExploiting contextual relationships across images has recently proven key to improve object detection. The resulting object co-detection algorithms, however, fail to exploit the correlations between multiple classes and, for scalability reasons are limited to modeling object instance similarity with relatively low-dimensional hand-crafted features. Here, we address the problem of multiclass object co-detection for large scale datasets. To this end, we formulate co-detection as the joint multiclass labeling of object candidates obtained in a class-independent manner. To exploit the correlations between objects, we build a fully-connected CRF on the candidates, which explicitly incorporates both geometric layout relations across object classes and similarity relations across multiple images. We then introduce a structural boosting algorithm that lets us exploits rich, high-dimensional deep network features to learn object similarity within our fully-connected CRF. Our experiments on PASCAL VOC 2007 and 2012 evidences the benefits of our approach over object detection with RCNN, single-image CRF methods and state-of-the-art co-detection algorithms. Zeeshan Hayder, Xuming He 0001, Mathieu Salzmann |
ICCV | 3 |
| 2015 | Shape Interaction Matrix Revisited and Robustified: Efficient Subspace Clustering with Corrupted and Incomplete DataabstractThe Shape Interaction Matrix (SIM) is one of the earliest approaches to performing subspace clustering (i.e., separating points drawn from a union of subspaces). In this paper, we revisit the SIM and reveal its connections to several recent subspace clustering methods. Our analysis lets us derive a simple, yet effective algorithm to robustify the SIM and make it applicable to realistic scenarios where the data is corrupted by noise. We justify our method by intuitive examples and the matrix perturbation theory. We then show how this approach can be extended to handle missing data, thus yielding an efficient and general subspace clustering algorithm. We demonstrate the benefits of our approach over state-of-the-art subspace clustering methods on several challenging motion segmentation and face clustering problems, where the data includes corruptions and missing measurements. Pan Ji, Mathieu Salzmann, Hongdong Li |
ICCV | 2 |
| 2015 | Cutting Edge: Soft Correspondences in Multimodal Scene ParsingabstractExploiting multiple modalities for semantic scene parsing has been shown to improve accuracy over the single modality scenario. Existing methods, however, assume that corresponding regions in two modalities have the same label. In this paper, we address the problem of data misalignment and label inconsistencies, e.g., due to moving objects, in semantic labeling, which violate the assumption of existing techniques. To this end, we formulate multimodal semantic labeling as inference in a CRF, and introduce latent nodes to explicitly model inconsistencies between two domains. These latent nodes allow us not only to leverage information from both domains to improve their labeling, but also to cut the edges between inconsistent regions. To eliminate the need for hand tuning the parameters of our model, we propose to learn intra-domain and inter-domain potential functions from training data. We demonstrate the benefits of our approach on two publicly available datasets containing 2D imagery and 3D point clouds. Thanks to our latent nodes and our learning strategy, our method outperforms the state-of-the-art in both cases. Sarah Taghavi Namin, Mohammad Najafi, Mathieu Salzmann, Lars Petersson |
ICCV | 3 |
| 2015 | Deformable 3D Fusion: From Partial Dynamic 3D Observations to Complete 4D ModelsabstractCapturing the 3D motion of dynamic, non-rigid objects has attracted significant attention in computer vision. Existing methods typically require either complete 3D volumetric observations, or a shape template. In this paper, we introduce a template-less 4D reconstruction method that incrementally fuses highly-incomplete 3D observations of a deforming object, and generates a complete, temporally-coherent shape representation of the object. To this end, we design an online algorithm that alternatively registers new observations to the current model estimate and updates the model. We demonstrate the effectiveness of our approach at reconstructing non-rigidly moving objects from highly-incomplete measurements on both sequences of partial 3D point clouds and Kinect videos. WeiPeng Xu, Mathieu Salzmann, Yongtian Wang, Yue Liu 0005 |
ICCV | 2 |
| 2015 | Efficient scene parsing by sampling unary potentials in a fully-connected CRFabstractEfficient, fully-connected CRF inference enables fast semantic labelling of images. However, this requires high-quality unary potentials to be computed, which is currently time-consuming. While some recent work attempts to address this issue by only computing a subset of unary potentials, a need remains for a simple, fast way to decide which unary potentials should be computed, without sacrificing accuracy. In particular, for embedded applications, a method which avoids time or memory-intensive operations is desired. In this paper, we introduce an approach to selecting good locations to compute unary potentials. We implement an efficient morphological approach to select a small proportion of pixel locations where unary potentials will be calculated. The speed of our labelling method allows us to directly search a large parameter space to optimize our method for a given task. We show that our method can achieve comparable accuracy to what can be achieved when all unary potentials are calculated, with significant time saving. Furthermore, we show that it is possible to tune our method to yield improved accuracy for certain classes of interest. We demonstrate this over multiple datasets representing challenging applications for our approach. Lachlan Horne, José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
Intelligent Vehicles Symposium | 3 |
| 2015 | A Linear Chain Markov Model for Detection and Localization of Cells in Early Stage Embryo DevelopmentabstractWe address the problem of detecting and localizing cells in time lapse microscopy images during early stage embryo development. Our approach is based on a linear chain Markov model that estimates the number and location of cells at each time step. The state space for each time step is derived from a randomized ellipse fitting algorithm that attempts to find individual cell candidates within the embryo. These cell candidates are combined into embryo hypotheses, and our algorithm finds the most likely sequence of hypotheses over all time steps. We restrict our attention to detect and localize up to four cells, which is sufficient for many important applications such as predicting blast cyst and can be used for assessing embryos in vitro fertilization procedures. We evaluate our method on twelve sequences of developing embryos and find that we can reliably detect and localize cells up to the four cell stage. Aisha Khan, Stephen Gould, Mathieu Salzmann |
WACV | 3 |
| 2015 | A Multi-modal Graphical Model for Scene AnalysisabstractIn this paper, we introduce a multi-modal graphical model to address the problems of semantic segmentation using 2D-3D data exhibiting extensive many-to-one correspondences. Existing methods often impose a hard correspondence between the 2D and 3D data, where the 2D and 3D corresponding regions are forced to receive identical labels. This results in performance degradation due to misalignments, 3D-2D projection errors and occlusions. We address this issue by defining a graph over the entire set of data that models soft correspondences between the two modalities. This graph encourages each region in a modality to leverage the information from its corresponding regions in the other modality to better estimate its class label. We evaluate our method on a publicly available dataset and beat the state-of-the-art. Additionally, to demonstrate the ability of our model to support multiple correspondences for objects in 3D and 2D domains, we introduce a new multi-modal dataset, which is composed of panoramic images and LIDAR data, and features a rich set of many-to-one correspondences. Sarah Taghavi Namin, Mohammad Najafi, Mathieu Salzmann, Lars Petersson |
WACV | 3 |
| 2015 | Kernel Methods on Riemannian Manifolds with Gaussian RBF KernelsabstractIn this paper, we develop an approach to exploiting kernel methods with manifold-valued data. In many computer vision problems, the data can be naturally represented as points on a Riemannian manifold. Due to the non-Euclidean geometry of Riemannian manifolds, usual Euclidean computer vision and machine learning algorithms yield inferior results on such data. In this paper, we define Gaussian radial basis function (RBF)-based positive definite kernels on manifolds that permit us to embed a given manifold with a corresponding metric in a high dimensional reproducing kernel Hilbert space. These kernels make it possible to utilize algorithms developed for linear spaces on nonlinear manifold-valued data. Since the Gaussian RBF defined with any given metric is not always positive definite, we present a unified framework for analyzing the positive definiteness of the Gaussian RBF on a generic metric space. We then use the proposed framework to identify positive definite kernels on two specific manifolds commonly encountered in computer vision: the Riemannian manifold of symmetric positive definite matrices and the Grassmann manifold, i.e., the Riemannian manifold of linear subspaces of a Euclidean space. We show that many popular algorithms designed for Euclidean spaces, such as support vector machines, discriminant analysis and principal component analysis can be generalized to Riemannian manifolds with the help of such positive definite Gaussian kernels. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Mirror Surface Reconstruction from a Single ImageabstractThis paper tackles the problem of reconstructing the shape of a smooth mirror surface from a single image. In particular, we consider the case where the camera is observing the reflection of a static reference target in the unknown mirror. We first study the reconstruction problem given dense correspondences between 3D points on the reference target and image locations. In such conditions, our differential geometry analysis provides a theoretical proof that the shape of the mirror surface can be recovered if the pose of the reference target is known. We then relax our assumptions by considering the case where only sparse correspondences are available. In this scenario, we formulate reconstruction as an optimization problem, which can be solved using a nonlinear least-squares method. We demonstrate the effectiveness of our method on both synthetic and real images. We then provide a theoretical analysis of the potential degenerate cases with and without prior knowledge of the pose of the reference target. Finally we show that our theory can be similarly applied to the reconstruction of the surface of transparent object. Miaomiao Liu 0001, Richard I. Hartley, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Domain Adaptation on the Statistical ManifoldabstractIn this paper, we tackle the problem of unsupervised domain adaptation for classification. In the unsupervised scenario where no labeled samples from the target domain are provided, a popular approach consists in transforming the data such that the source and target distributions become similar. To compare the two distributions, existing approaches make use of the Maximum Mean Discrepancy (MMD). However, this does not exploit the fact that probability distributions lie on a Riemannian manifold. Here, we propose to make better use of the structure of this manifold and rely on the distance on the manifold to compare the source and target distributions. In this framework, we introduce a sample selection method and a subspace-based method for unsupervised domain adaptation, and show that both these manifold-based techniques outperform the corresponding approaches based on the MMD. Furthermore, we show that our subspace-based approach yields state-of-the-art results on a standard object recognition benchmark. Mahsa Baktash, Mehrtash Harandi, Brian C. Lovell, Mathieu Salzmann |
CVPR | 4 |
| 2014 | Bregman Divergences for Infinite Dimensional Covariance MatricesabstractWe introduce an approach to computing and comparing Covariance Descriptors (CovDs) in infinite-dimensional spaces. CovDs have become increasingly popular to address classification problems in computer vision. While CovDs offer some robustness to measurement variations, they also throw away part of the information contained in the original data by only retaining the second-order statistics over the measurements. Here, we propose to overcome this limitation by first mapping the original data to a high-dimensional Hilbert space, and only then compute the CovDs. We show that several Bregman divergences can be computed between the resulting CovDs in Hilbert space via the use of kernels. We then exploit these divergences for classification purpose. Our experiments demonstrate the benefits of our approach on several tasks, such as material and texture recognition, person re-identification, and action recognition from motion capture data. Mehrtash Harandi, Mathieu Salzmann, Fatih Porikli |
CVPR | 2 |
| 2014 | Optimizing over Radial Kernels on Compact ManifoldsabstractWe tackle the problem of optimizing over all possible positive definite radial kernels on Riemannian manifolds for classification. Kernel methods on Riemannian manifolds have recently become increasingly popular in computer vision. However, the number of known positive definite kernels on manifolds remain very limited. Furthermore, most kernels typically depend on at least one parameter that needs to be tuned for the problem at hand. A poor choice of kernel, or of parameter value, may yield significant performance drop-off. Here, we show that positive definite radial kernels on the unit n-sphere, the Grassmann manifold and Kendall's shape manifold can be expressed in a simple form whose parameters can be automatically optimized within a support vector machine framework. We demonstrate the benefits of our kernel learning algorithm on object, face, action and shape recognition. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
CVPR | 3 |
| 2014 | Discrete-Continuous Depth Estimation from a Single ImageabstractIn this paper, we tackle the problem of estimating the depth of a scene from a single image. This is a challenging task, since a single image on its own does not provide any depth cue. To address this, we exploit the availability of a pool of images for which the depth is known. More specifically, we formulate monocular depth estimation as a discrete-continuous optimization problem, where the continuous variables encode the depth of the superpixels in the input image, and the discrete ones represent relationships between neighboring superpixels. The solution to this discrete-continuous optimization problem is then obtained by performing inference in a graphical model using particle belief propagation. The unary potentials in this graphical model are computed by making use of the images with known depth. We demonstrate the effectiveness of our model in both the indoor and outdoor scenarios. Our experimental evaluation shows that our depth estimates are more accurate than existing methods on standard datasets. Miaomiao Liu 0001, Mathieu Salzmann, Xuming He 0001 |
CVPR | 2 |
| 2014 | From Manifold to Manifold: Geometry-Aware Dimensionality Reduction for SPD Matrices
Mehrtash Harandi, Mathieu Salzmann, Richard I. Hartley |
ECCV (2) | 2 |
| 2014 | Expanding the Family of Grassmannian Kernels: An Embedding Perspective
Mehrtash Harandi, Mathieu Salzmann, Sadeep Jayasumana, Richard I. Hartley, Hongdong Li |
ECCV (7) | 2 |
| 2014 | Object Co-detection via Efficient Inference in a Fully-Connected CRF
Zeeshan Hayder, Mathieu Salzmann, Xuming He 0001 |
ECCV (3) | 2 |
| 2014 | Robust Motion Segmentation with Unknown Correspondences
Pan Ji, Hongdong Li, Mathieu Salzmann, Yuchao Dai |
ECCV (6) | 3 |
| 2014 | Non-associative Higher-Order Markov Networks for Point Cloud Classification
Mohammad Najafi, Sarah Taghavi Namin, Mathieu Salzmann, Lars Petersson |
ECCV (5) | 3 |
| 2014 | Nonrigid Surface Registration and Completion from RGBD Images
WeiPeng Xu, Mathieu Salzmann, Yongtian Wang, Yue Liu 0005 |
ECCV (2) | 2 |
| 2014 | Null space clustering with applications to motion segmentation and face clusteringabstractThe problems of motion segmentation and face clustering can be addressed in a framework of subspace clustering methods. In this paper, we tackle the more general problem of clustering data points lying in a union of low-dimensional linear(or affine) subspaces, which can be naturally applied in motion segmentation and face clustering. For data points drawn from linear (or affine) subspaces, we propose a novel algorithm called Null Space Clustering (NSC), utilizing the null space of the data matrix to construct the affinity matrix. To better deal with noise and outliers, it is converted to an equivalent problem with Frobenius norm minimization, which can be solved efficiently. We demonstrate that the proposed NSC leads to improved performance in terms of clustering accuracy and efficiency when compared to state-of-the-art algorithms on two well-known datasets, i.e., Hopkins 155 and Extended Yale B. Pan Ji, Yiran Zhong, Hongdong Li, Mathieu Salzmann |
ICIP | 4 |
| 2014 | Large-scale semantic co-labeling of image setsabstractAs evidenced by video segmentation and cosegmentation approaches, exploiting multiple images is key to the success of visual scene understanding. With the availability of increasingly large sets of images, there is a clear need for methods that can efficiently analyze the similarities and structure across huge numbers of image pixels. Furthermore, to make effective use of this data, these similarities should not just be considered locally between neighboring pixels, but between all pairs of pixels across all images. In this paper, we tackle this challenging scenario by introducing a semantic co-labeling approach that performs efficient inference in a fully-connected CRF defined over the pixels, or superpixels, of an image set. Our experimental evaluation demonstrates that our approach yields improved accuracy while coming at no additional computation cost compared to performing segmentation sequentially on individual images. Furthermore, our formulation lets us perform inference over ten thousand images in a matter of seconds. José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
WACV | 2 |
| 2014 | Data-driven road detectionabstractIn this paper, we tackle the problem of road detection from RGB images. In particular, we follow a data-driven approach to segmenting the road pixels in an image. To this end, we introduce two road detection methods: A top-down approach that builds an image-level road prior based on the traffic pattern observed in an input image, and a bottom-up technique that estimates the probability that an image superpixel belongs to the road surface in a nonparametric manner. Both our algorithms work on the principle of label transfer in the sense that the road prior is directly constructed from the ground-truth segmentations of training images. Our experimental evaluation on four different datasets shows that this approach outperforms existing top-down and bottom-up techniques, and is key to the robustness of road detection algorithms to the dataset bias. José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
WACV | 2 |
| 2014 | Efficient dense subspace clusteringabstractIn this paper, we tackle the problem of clustering data points drawn from a union of linear (or affine) subspaces. To this end, we introduce an efficient subspace clustering algorithm that estimates dense connections between the points lying in the same subspace. In particular, instead of following the standard compressive sensing approach, we formulate subspace clustering as a Frobenius norm minimization problem, which inherently yields denser con- nections between the data points. While in the noise-free case we rely on the self-expressiveness of the observations, in the presence of noise we simultaneously learn a clean dictionary to represent the data. Our formulation lets us address the subspace clustering problem efficiently. More specifically, the solution can be obtained in closed-form for outlier-free observations, and by performing a series of linear operations in the presence of outliers. Interestingly, we show that our Frobenius norm formulation shares the same solution as the popular nuclear norm minimization approach when the data is free of any noise, or, in the case of corrupted data, when a clean dictionary is learned. Our experimental evaluation on motion segmentation and face clustering demonstrates the benefits of our algorithm in terms of clustering accuracy and efficiency. Pan Ji, Mathieu Salzmann, Hongdong Li |
WACV | 2 |
| 2014 | Discriminative Non-Linear Stationary Subspace Analysis for Video ClassificationabstractLow-dimensional representations are key to the success of many video classification algorithms. However, the commonly-used dimensionality reduction techniques fail to account for the fact that only part of the signal is shared across all the videos in one class. As a consequence, the resulting representations contain instance-specific information, which introduces noise in the classification process. In this paper, we introduce non-linear stationary subspace analysis: a method that overcomes this issue by explicitly separating the stationary parts of the video signal (i.e., the parts shared across all videos in one class), from its non-stationary parts (i.e., the parts specific to individual videos). Our method also encourages the new representation to be discriminative, thus accounting for the underlying classification problem. We demonstrate the effectiveness of our approach on dynamic texture recognition, scene classification and action recognition. Mahsa Baktash, Mehrtash Harandi, Brian C. Lovell, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Kernel Methods on the Riemannian Manifold of Symmetric Positive Definite MatricesabstractSymmetric Positive Definite (SPD) matrices have become popular to encode image information. Accounting for the geometry of the Riemannian manifold of SPD matrices has proven key to the success of many algorithms. However, most existing methods only approximate the true shape of the manifold locally by its tangent plane. In this paper, inspired by kernel methods, we propose to map SPD matrices to a high dimensional Hilbert space where Euclidean geometry applies. To encode the geometry of the manifold in the mapping, we introduce a family of provably positive definite kernels on the Riemannian manifold of SPD matrices. These kernels are derived from the Gaussian kernel, but exploit different metrics on the manifold. This lets us extend kernel-based algorithms developed for Euclidean spaces, such as SVM and kernel PCA, to the Riemannian manifold of SPD matrices. We demonstrate the benefits of our approach on the problems of pedestrian detection, object categorization, texture analysis, 2D motion segmentation and Diffusion Tensor Imaging (DTI) segmentation. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
CVPR | 3 |
| 2013 | Mirror Surface Reconstruction from a Single ImageabstractThis paper tackles the problem of reconstructing the shape of a smooth mirror surface from a single image. In particular, we consider the case where the camera is ob-serving the reflection of a static reference target in the un-known mirror. We first study the reconstruction problem given dense correspondences between 3D points on the ref-erence target and image locations. In such conditions, our differential geometry analysis provides a theoretical proof that the shape of the mirror surface can be uniquely recov-ered if the pose of the reference target is known. We then relax our assumptions by considering the case where only sparse correspondences are available. In this scenario, we formulate reconstruction as an optimization problem, which can be solved using a nonlinear least-squares method. We demonstrate the effectiveness of our method on both syn-thetic and real images. 1. Miaomiao Liu 0001, Richard I. Hartley, Mathieu Salzmann |
CVPR | 3 |
| 2013 | Continuous Inference in Graphical Models with Polynomial EnergiesabstractIn this paper, we tackle the problem of performing inference in graphical models whose energy is a polynomial function of continuous variables. Our energy minimization method follows a dual decomposition approach, where the global problem is split into sub problems defined over the graph cliques. The optimal solution to these sub problems is obtained by making use of a polynomial system solver. Our algorithm inherits the convergence guarantees of dual decomposition. To speed up optimization, we also introduce a variant of this algorithm based on the augmented Lagrangian method. Our experiments illustrate the diversity of computer vision problems that can be expressed with polynomial energies, and demonstrate the benefits of our approach over existing continuous inference methods. Mathieu Salzmann |
CVPR | 1 |
| 2013 | Unsupervised Domain Adaptation by Domain Invariant ProjectionabstractDomain-invariant representations are key to addressing the domain shift problem where the training and test examples follow different distributions. Existing techniques that have attempted to match the distributions of the source and target domains typically compare these distributions in the original feature space. This space, however, may not be directly suitable for such a comparison, since some of the features may have been distorted by the domain shift, or may be domain specific. In this paper, we introduce a Domain Invariant Projection approach: An unsupervised domain adaptation method that overcomes this issue by extracting the information that is invariant across the source and target domains. More specifically, we learn a projection of the data to a low-dimensional latent space where the distance between the empirical distributions of the source and target examples is minimized. We demonstrate the effectiveness of our approach on the task of visual object recognition and show that it outperforms state-of-the-art methods on a standard domain adaptation benchmark dataset. Mahsa Baktash, Mehrtash Harandi, Brian C. Lovell, Mathieu Salzmann |
ICCV | 4 |
| 2013 | A Framework for Shape Analysis via Hilbert Space EmbeddingabstractWe propose a framework for 2D shape analysis using positive definite kernels defined on Kendall's shape manifold. Different representations of 2D shapes are known to generate different nonlinear spaces. Due to the nonlinearity of these spaces, most existing shape classification algorithms resort to nearest neighbor methods and to learning distances on shape spaces. Here, we propose to map shapes on Kendall's shape manifold to a high dimensional Hilbert space where Euclidean geometry applies. To this end, we introduce a kernel on this manifold that permits such a mapping, and prove its positive definiteness. This kernel lets us extend kernel-based algorithms developed for Euclidean spaces, such as SVM, MKL and kernel PCA, to the shape manifold. We demonstrate the benefits of our approach over the state-of-the-art methods on shape classification, clustering and retrieval. Sadeep Jayasumana, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
ICCV | 2 |
| 2013 | Real-time keystone correction for hand-held projectors with an RGBD cameraabstractThis paper introduces a novel and simple approach to realtime continuous keystone correction for hand-held projectors. An RGBD camera is attached to the projector to form a projector-RGBD-camera system. The system is first calibrated in an offline stage. At run-time, we then estimate the relative pose between the projector and the screen using the RGBD camera, which lets us correct the keystone distortion by warping the projected image accordingly. Experimental results show that our method outperforms existing techniques in terms of both accuracy and efficiency. WeiPeng Xu, Yongtian Wang, Yue Liu 0005, Dongdong Weng, Mengwen Tan, Mathieu Salzmann |
ICIP | 6 |
| 2013 | Non-Linear Stationary Subspace Analysis with Application to Video ClassificationabstractLow-dimensional representations are key to the success of many video classification algorithms. However, the commonly-used dimensionality reduction techniques fail to account for the fact that only part of the signal is shared across all the videos in one class. As a consequence, the resulting representations contain instance-specific information, which introduces noise in the classification process. In this paper, we introduce Non-Linear Stationary Subspace Analysis: A method that overcomes this issue by explicitly separating the stationary parts of the video signal (i.e., the parts shared across all videos in one class), from its non-stationary parts (i.e., specific to individual videos). We demonstrate the effectiveness of our approach on action recognition, dynamic texture classification and scene recognition. Mahsa Baktash, Mehrtash Harandi, Abbas Bigdeli, Brian C. Lovell, Mathieu Salzmann |
ICML (3) | 5 |
| 2013 | Learning appearance models for road detectionabstractWe introduce an approach to image-based road detection that exploits the availability of unannotated training images to learn an appearance model. Our approach allows us to remove the standard assumption that the lower part of the input image belongs to the road surface, which does not always hold and often yields strongly biased appearance models. Instead, we exploit this assumption in the training images, which yields a much more general appearance model. We then use the learned model to classify the pixels of an input image as road or background without requiring any assumptions about this image. Our experimental evaluation shows the benefits of our approach over existing methods in challenging real-world driving scenarios. José M. Álvarez 0004, Mathieu Salzmann, Nick Barnes |
Intelligent Vehicles Symposium | 2 |
| 2012 | A constrained latent variable modelabstractLatent variable models provide valuable compact representations for learning and inference in many computer vision tasks. However, most existing models cannot directly encode prior knowledge about the specific problem at hand. In this paper, we introduce a constrained latent variable model whose generated output inherently accounts for such knowledge. To this end, we propose an approach that explicitly imposes equality and inequality constraints on the model's output during learning, thus avoiding the computational burden of having to account for these constraints at inference. Our learning mechanism can exploit non-linear kernels, while only involving sequential closed-form updates of the model parameters. We demonstrate the effectiveness of our constrained latent variable model on the problem of non-rigid 3D reconstruction from monocular images, and show that it yields qualitative and quantitative improvements over several baselines. Aydin Varol, Mathieu Salzmann, Pascal Fua, Raquel Urtasun |
CVPR | 2 |
| 2012 | Beyond Feature Points: Structured Prediction for Monocular Non-rigid 3D Reconstruction
Mathieu Salzmann, Raquel Urtasun |
ECCV (4) | 1 |
| 2012 | Monocular 3D Reconstruction of Locally Textured SurfacesabstractMost recent approaches to monocular nonrigid 3D shape recovery rely on exploiting point correspondences and work best when the whole surface is well textured. The alternative is to rely on either contours or shading information, which has only been demonstrated in very restrictive settings. Here, we propose a novel approach to monocular deformable shape recovery that can operate under complex lighting and handle partially textured surfaces. At the heart of our algorithm are a learned mapping from intensity patterns to the shape of local surface patches and a principled approach to piecing together the resulting local shape estimates. We validate our approach quantitatively and qualitatively using both synthetic and real data. Aydin Varol, Appu Shaji, Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Learning cross-modality similarity for multinomial dataabstractMany applications involve multiple-modalities such as text and images that describe the problem of interest. In order to leverage the information present in all the modalities, one must model the relationships between them. While some techniques have been proposed to tackle this problem, they either are restricted to words describing visual objects only, or require full correspondences between the different modalities. As a consequence, they are unable to tackle more realistic scenarios where a narrative text is only loosely related to an image, and where only a few image-text pairs are available. In this paper, we propose a model that addresses both these challenges. Our model can be seen as a Markov random field of topic models, which connects the documents based on their similarity. As a consequence, the topics learned with our model are shared across connected documents, thus encoding the relations between different modalities. We demonstrate the effectiveness of our model for image retrieval from a loosely related text. Yangqing Jia, Mathieu Salzmann, Trevor Darrell |
ICCV | 2 |
| 2011 | Physically-based motion models for 3D tracking: A convex formulationabstractIn this paper, we propose a physically-based dynamical model for tracking. Our model relies on Newton's second law of motion, which governs any real-world dynamical system. As a consequence, it can be generally applied to very different tracking problems. Furthermore, since the equations describing Newton's second law are simple linear equalities, they can be incorporated in any tracking framework at very little cost. Leveraging this lets us introduce a convex formulation of 3D tracking from monocular images. We demonstrate the strengths of our approach on various types of motion, such as billiards and acrobatics. Mathieu Salzmann, Raquel Urtasun |
ICCV | 1 |
| 2011 | Linear Local Models for Monocular Reconstruction of Deformable SurfacesabstractRecovering the 3D shape of a nonrigid surface from a single viewpoint is known to be both ambiguous and challenging. Resolving the ambiguities typically requires prior knowledge about the most likely deformations that the surface may undergo. It often takes the form of a global deformation model that can be learned from training data. While effective, this approach suffers from the fact that a new model must be learned for each new surface, which means acquiring new training data, and may be impractical. In this paper, we replace the global models by linear local models for surface patches, which can be assembled to represent arbitrary surface shapes as long as they are made of the same material. Not only do they eliminate the need to retrain the model for different surface shapes, they also let us formulate 3D shape reconstruction from correspondences as either an algebraic problem that can be solved in closed form or a convex optimization problem whose solution can be found using standard numerical packages. We present quantitative results on synthetic data, as well as qualitative results on real images. Mathieu Salzmann, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Combining discriminative and generative methods for 3D deformable surface and articulated pose reconstructionabstractHistorically non-rigid shape recovery and articulated pose estimation have evolved as separate fields. Recent methods for non-rigid shape recovery have focused on improving the algorithmic formulation, but have only considered the case of reconstruction from point-to-point correspondences. In contrast, many techniques for pose estimation have followed a discriminative approach, which allows for the use of more general image cues. However, these techniques typically require large training sets and suffer from the fact that standard discriminative methods do not enforce constraints between output dimensions. In this paper, we combine ideas from both domains and propose a unified framework for articulated pose estimation and 3D surface reconstruction. We address some of the issues of discriminative methods by explicitly constraining their prediction. Furthermore, our formulation allows for the combination of generative and discriminative methods into a single, common framework. Mathieu Salzmann, Raquel Urtasun |
CVPR | 1 |
| 2010 | Learning to Recognize Objects from Unseen Modalities
C. Mario Christoudias, Raquel Urtasun, Mathieu Salzmann, Trevor Darrell |
ECCV (1) | 3 |
| 2010 | Factorized Latent Spaces with Structured SparsityabstractRecent approaches to multi-view learning have shown that factorizing the information into parts that are shared across all views and parts that are private to each view could effectively account for the dependencies and independencies between the different input modalities. Unfortunately, these approaches involve minimizing non-convex objective functions. In this paper, we propose an approach to learning such factorized representations inspired by sparse coding techniques. In particular, we show that structured sparsity allows us to address the multi-view learning problem by alternately solving two convex optimization problems. Furthermore, the resulting factorized latent spaces generalize over existing approaches in that they allow :having latent dimensions shared between any subset of the views instead of between all the views only. We show that our approach outperforms state-of-the-art methods on the task of human pose estimation. Yangqing Jia, Mathieu Salzmann, Trevor Darrell |
NIPS | 2 |
| 2010 | Implicitly Constrained Gaussian Process Regression for Monocular Non-Rigid Pose EstimationabstractEstimating 3D pose from monocular images is a highly ambiguous problem. Physical constraints can be exploited to restrict the space of feasible configurations. In this paper we propose an approach to constraining the prediction of a discriminative predictor. We first show that the mean prediction of a Gaussian process implicitly satisfies linear constraints if those constraints are satisfied by the training examples. We then show how, by performing a change of variables, a GP can be forced to satisfy quadratic constraints. As evidenced by the experiments, our method outperforms state-of-the-art approaches on the tasks of rigid and non-rigid pose estimation. Mathieu Salzmann, Raquel Urtasun |
NIPS | 1 |
| 2009 | Observable subspaces for 3D human motion recoveryabstractThe articulated body models used to represent human motion typically have many degrees of freedom, usually expressed as joint angles that are highly correlated. The true range of motion can therefore be represented by latent variables that span a low-dimensional space. This has often been used to make motion tracking easier. However, learning the latent space in a problem- independent way makes it non trivial to initialize the tracking process by picking appropriate initial values for the latent variables, and thus for the pose. In this paper, we show that by directly using observable quantities as our latent variables, we eliminate this problem and achieve full automation given only modest amounts of training data. More specifically, we exploit the fact that the trajectory of a person's feet or hands strongly constrains body pose in motions such as skating, skiing, or golfing. These trajectories are easy to compute and to parameterize using a few variables. We treat these as our latent variables and learn a mapping between them and sequences of body poses. In this manner, by simply tracking the feet or the hands, we can reliably guess initial poses over whole sequences and, then, refine them. Andrea Fossati, Mathieu Salzmann, Pascal Fua |
CVPR | 2 |
| 2009 | Capturing 3D stretchable surfaces from single images in closed formabstractWe present a closed form solution to the problem of recovering the 3D shape of a nonrigid potentially stretchable surface from 3D-to-2D correspondences. In other words, we can reconstruct a surface from a single image without a priori knowledge of its deformations in that image. State of the art solutions to nonrigid 3D shape recovery rely on the fact that distances between neighboring surface points must be preserved and are therefore limited to inelastic surfaces. Here, we show that replacing the inextensibility constraints by shading ones removes this limitation while still allowing 3D reconstruction in closed-form. We demonstrate our method and compare it to an earlier one using both synthetic and real data. Francesc Moreno-Noguer, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
CVPR | 2 |
| 2009 | Reconstructing sharply folding surfaces: A convex formulationabstractIn recent years, 3D deformable surface reconstruction from single images has attracted renewed interest. It has been shown that preventing the surface from either shrinking or stretching is an effective way to resolve the ambiguities inherent to this problem. However, while the geodesic distances on the surface may not change, the Euclidean ones decrease when folds appear. Therefore, when applied to discrete surface representations, such constant-distance constraints are only effective for smoothly deforming surfaces, and become inaccurate for more flexible ones that can exhibit sharp folds. In such cases, surface points must be allowed to come closer to each other. In this paper, we show that replacing the equality constraints of earlier approaches by inequality constraints that let the mesh representation of the surface shrink but not expand yields not only a more faithful representation, but also a convex formulation of the reconstruction problem. As a result, we can accurately reconstruct surfaces undergoing complex deformations that include sharp folds from individual images. Mathieu Salzmann, Pascal Fua |
CVPR | 1 |
| 2009 | Template-free monocular reconstruction of deformable surfacesabstractIt has recently been shown that deformable 3D surfaces could be recovered from single video streams. However, existing techniques either require a reference view in which the shape of the surface is known a priori, which often may not be available, or require tracking points over long sequences, which is hard to do. In this paper, we overcome these limitations. To this end, we establish correspondences between pairs of frames in which the shape is different and unknown. We then estimate homographies between corresponding local planar patches in both images. These yield approximate 3D reconstructions of points within each patch up to a scale factor. Since we consider overlapping patches, we can enforce them to be consistent over the whole surface. Finally, a local deformation model is used to fit a triangulated mesh to the 3D point cloud, which makes the reconstruction robust to both noise and outliers in the image data. Aydin Varol, Mathieu Salzmann, Engin Tola, Pascal Fua |
ICCV | 2 |
| 2008 | 3D pose refinement from reflectionsabstractWe demonstrate how to exploit reflections for accurate registration of shiny objects: The lighting environment can be retrieved from the reflections under a distant illumination assumption. Since it remains unchanged when the camera or the object of interest moves, this provides powerful additional constraints that can be incorporated into standard pose estimation algorithms. The key idea and main contribution of the paper is therefore to show that the registration should also be performed in the lighting environment space, instead of in the image space only. This lets us recover very accurate pose estimates because the specularities are very sensitive to pose changes. An interesting side result is an accurate estimate of the lighting environment. Furthermore, since the mapping from lighting environment to specularities has no analytical expression for objects represented as 3D meshes, and is not 1-to-1, registering lighting environments is far from trivial. However we propose a general and effective solution. Our approach is demonstrated on both synthetic and real images. Pascal Lagger, Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
CVPR | 2 |
| 2008 | Local deformation models for monocular 3D shape recoveryabstractWithout a deformation model, monocular 3D shape recovery of deformable surfaces is severely under-constrained. Even when the image information is rich enough, prior knowledge of the feasible deformations is required to overcome the ambiguities. This is further accentuated when such information is poor, which is a key issue that has not yet been addressed. In this paper, we propose an approach to learning shape priors to solve this problem. By contrast with typical statistical learning methods that build models for specific object shapes, we learn local deformation models, and combine them to reconstruct surfaces of arbitrary global shapes. Not only does this improve the generality of our deformation models, but it also facilitates learning since the space of local deformations is much smaller than that of global ones. While using a texture-based approach, we show that our models are effective to reconstruct from single videos poorly-textured surfaces of arbitrary shape, made of materials as different as cardboard, that deforms smoothly, and much lighter tissue paper whose deformations may be far more complex. Mathieu Salzmann, Raquel Urtasun, Pascal Fua |
CVPR | 1 |
| 2008 | Closed-Form Solution to Non-rigid 3D Surface Registration
Mathieu Salzmann, Francesc Moreno-Noguer, Vincent Lepetit, Pascal Fua |
ECCV (4) | 1 |
| 2007 | Deformable Surface Tracking AmbiguitiesabstractWe study from a theoretical standpoint the ambiguities that occur when tracking a generic deformable surface under monocular perspective projection given 3D to 2D correspondences. We show that, additionally to the known scale ambiguity, a set of potential ambiguities can be clearly identified. From this, we deduce a minimal set of constraints required to disambiguate the problem and incorporate them into a working algorithm that runs on real noisy data. Mathieu Salzmann, Vincent Lepetit, Pascal Fua |
CVPR | 1 |
| 2007 | Convex Optimization for Deformable Surface 3-D Trackingabstract3-D shape recovery of non-rigid surfaces from 3-D to 2-D correspondences is an under-constrained problem that requires prior knowledge of the possible deformations. State-of-the-art solutions involve enforcing smoothness constraints that limit their applicability and prevent the recovery of sharply folding and creasing surfaces. Here, we propose a method that does not require such smoothness constraints. Instead, we represent surfaces as triangulated meshes and, assuming the pose in the first frame to be known, disallow large changes of edge orientation between consecutive frames, which is a generally applicable constraint when tracking surfaces in a 25 frames- per-second video sequence. We will show that tracking under these constraints can be formulated as a Second Order Cone Programming feasibility problem. This yields a convex optimization problem with stable solutions for a wide range of surfaces with very different physical properties. Mathieu Salzmann, Richard I. Hartley, Pascal Fua |
ICCV | 1 |
| 2007 | Implicit Meshes for Effective Silhouette Handling
Slobodan Ilic, Mathieu Salzmann, Pascal Fua |
Int. J. Comput. Vis. | 2 |
| 2007 | Surface Deformation Models for Nonrigid 3D Shape RecoveryabstractThree-dimensional detection and shape recovery of a nonrigid surface from video sequences require deformation models to effectively take advantage of potentially noisy image data. Here, we introduce an approach to creating such models for deformable 3D surfaces. We exploit the fact that the shape of an inextensible triangulated mesh can be parameterized in terms of a small subset of the angles between its facets. We use this set of angles to create a representative set of potential shapes, which we feed to a simple dimensionality reduction technique to produce low-dimensional 3D deformation models. We show that these models can be used to accurately model a wide range of deforming 3D surfaces from video sequences acquired under realistic conditions. Mathieu Salzmann, Julien Pilet, Slobodan Ilic, Pascal Fua |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Physically Valid Shape Parameterization for Monocular 3-D Deformable Surface TrackingabstractCVLAB Mathieu Salzmann, Slobodan Ilic, Pascal Fua |
BMVC | 1 |
| 2005 | Implicit Surfaces Make for Better SilhouettesabstractThis paper advocates an implicit-surface representation of generic 3-D surfaces to take advantage of occluding edges in a very robust way. This lets us exploit silhouette constraints in uncontrolled environments that may involve occlusions and changing or cluttered backgrounds, which limit the applicability of most silhouette based methods. This desirable behavior is completely independent from the way the surface deformations are parametrized. To show this, we demonstrate our technique in three very different cases: modeling the deformations of a piece of paper represented by an ordinary triangulated mesh; tracking a person's shoulders whose deformations are expressed in terms of Dirichlet free form deformations; reconstructing the shape of a human face parametrized in terms of a principal component analysis model. Slobodan Ilic, Mathieu Salzmann, Pascal Fua |
CVPR (1) | 2 |