EDBT 2026 Demo / reviewers in the wild / expert
Patrick Pérez
dblp:71/1167
· DBLP profile ↗
208ranked-venue papers
8as first author
46since 2021 · last 2026
0000-0002-8124-1206ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 158 · 5 first-author · 23 since 2021Artificial intelligence and machine learning · 132 · 3 first-author · 40 since 2021Systems, architecture and hardware · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Theory of computation · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLIP's Visual Embedding Projector is a Few-shot CornucopiaabstractWe introduce ProLIP, a simple and architecture-agnostic method for adapting contrastively pretrained vision-language models, such as CLIP [36], to few-shot classification. ProLIP fine-tunes the vision encoder’s projection matrix with Frobenius norm regularization on its deviation from the pretrained weights. It achieves state-of-the-art performance on 11 few-shot classification benchmarks under both "few-shot validation" [23] and "validation-free" [42] settings. Moreover, by rethinking the non-linear CLIP-Adapter [13] through ProLIP’s lens, we design a Regularized Linear Adapter (RLA) that performs better, requires no hyperparameter tuning, is less sensitive to learning rate values, and offers an alternative to ProLIP in black-box scenarios where model weights are inaccessible. Beyond few-shot classification, ProLIP excels in cross-dataset transfer, domain generalization, base-to-new class generalization, and test-time adaptation—where it outperforms prompt tuning while being an order of magnitude faster to train. Code is available at https://github.com/astra-vision/ProLIP. Mohammad Fahes, Andrei Bursuc, Patrick Pérez, Raoul de Charette |
WACV | 4 |
| 2026 | Domain Adaptation with a Single Vision-Language Embedding
Mohammad Fahes, Andrei Bursuc, Patrick Pérez, Raoul de Charette |
Int. J. Comput. Vis. | 4 |
| 2025 | ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger BridgeabstractDiffusion models break down the challenging task of generating data from high-dimensional distributions into a series of easier denoising steps. Inspired by this paradigm, we propose a novel approach that extends the diffusion framework into modality space, decomposing the complex task of RGB image generation into simpler, interpretable stages. Our method, termed {\papernameAbbrev}, cascades modality-specific models, each responsible for generating an intermediate representation, such as contours, palettes, and detailed textures, ultimately culminating in a high-quality RGB image.
Instead of relying on the naive LDM concatenation conditioning mechanism to connect the different stages together, we employ Schr\"odinger Bridge to determine the optimal transport between different modalities.
Although employing a cascaded pipeline introduces more stages, which could lead to a more complex architecture, each stage is meticulously formulated for efficiency and accuracy, surpassing Stable-Diffusion (LDM) performance.
Modality composition not only enhances overall performance but enables emerging proprieties such as consistent editing, interaction capabilities, high-level interpretability, and faster convergence and sampling rate.
Extensive experiments on diverse datasets, including LSUN-Churches, ImageNet, CelebHQ, and LAION-Art, demonstrate the efficacy of our approach, consistently outperforming state-of-the-art methods.
For instance, {\papernameAbbrev} achieves notable efficiency, matching LDM performance on LSUN-Churches while operating 2$\times$ faster with a 3$\times$ smaller architecture.
The project website is available at:
\href{https://toddlerdiffusion.github.io/website/}{$https://toddlerdiffusion.github.io/website/$} Eslam Mohamed Bakr, Liangbing Zhao, Vincent Tao Hu, Matthieu Cord, Patrick Pérez |
ICLR | 5 |
| 2025 | Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
Oriane Siméoni, Eloi Zablocki, Spyros Gidaris, Gilles Puy, Patrick Pérez |
Int. J. Comput. Vis. | 5 |
| 2025 | Unsupervised Semantic Segmentation of Urban Scenes via Cross-Modal DistillationabstractAbstract Semantic image segmentation models typically require extensive pixel-wise annotations, which are costly to obtain and prone to biases. Our work investigates learning semantic segmentation in urban scenes without any manual annotation. We propose a novel method for learning pixel-wise semantic segmentation using raw, uncurated data from vehicle-mounted cameras and LiDAR sensors, thus eliminating the need for manual labeling. Our contributions are as follows. First, we develop a novel approach for cross-modal unsupervised learning of semantic segmentation by leveraging synchronized LiDAR and image data. A crucial element of our method is the integration of an object proposal module that examines the LiDAR point cloud to generate proposals for spatially consistent objects. Second, we demonstrate that these 3D object proposals can be aligned with corresponding images and effectively grouped into semantically meaningful pseudo-classes. Third, we introduce a cross-modal distillation technique that utilizes image data partially annotated with the learnt pseudo-classes to train a transformer-based model for semantic image segmentation. Fourth, we demonstrate further significant improvements of our approach by extending the proposed model using a teacher-student distillation with an exponential moving average and incorporating soft targets from the teacher. We show the generalization capabilities of our method by testing on four different testing datasets (Cityscapes, Dark Zurich, Nighttime Driving, and ACDC) without any fine-tuning. We present an in-depth experimental analysis of the proposed model including results when using another pre-training dataset, per-class and pixel accuracy results, confusion matrices, PCA visualization, k-NN evaluation, ablations of the number of clusters and LiDAR’s density, supervised finetuning as well as additional qualitative results and their analysis. Antonín Vobecký, David Hurych, Oriane Siméoni, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Josef Sivic |
Int. J. Comput. Vis. | 6 |
| 2025 | Manipulating Trajectory Prediction Models With BackdoorsabstractAutonomous vehicles depend on accurate trajectory prediction to navigate safely in complex traffic. Yet current models are vulnerable to stealthy backdoor attacks: an adversary embeds subtle, physically plausible triggers during training that remain latent until activated. To address this risk, we introduce a structured framework categorizing four trigger types—spatial, kinetic (braking), coordinated, and composite—and demonstrate on two benchmarks (nuScenes and Argoverse 2) and two state-of-the-art architectures (Autobot and Wayformer) that poisoning as little as 5% of training samples can reliably hijack future predictions. We further propose a real-time defense leveraging social attention: by encoding agent histories, computing cross-attention to the target vehicle, and filtering out agents with anomalously high weights, our method neutralizes backdoor triggers without degrading clean-data accuracy. Comprehensive experiments show our defense reduces attack success rates across diverse urban scenarios—intersections, roundabouts, multi-lane roads—highlighting both the severity of backdoor threats and a promising pathway to secure trajectory predictors in autonomous driving systems. Kaouther Messaoud, Kathrin Grosse, Mickaël Chen, Matthieu Cord, Patrick Pérez, Alexandre Alahi |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | PointBeV: A Sparse Approach to BeV PredictionsabstractBird's-eye View (BeV) representations have emerged as the de-facto shared space in driving applications, offering a unified space for sensor data fusion and supporting various downstream tasks. However, conventional models use grids with fixed resolution and range and face computational inefficiencies due to the uniform allocation of resources across all cells. To address this, we propose Point-BeV, a novel sparse BeV segmentation model operating on sparse BeV cells instead of dense grids. This approach offers precise control over memory usage, enabling the use of long temporal contexts and accommodating memory-constrained platforms. PointBeV employs an efficient two-pass strategy for training, enabling focused computation on regions of interest. At inference time, it can be used with various memory/performance trade-offs and flexibly adjusts to new specific use cases. PointBeV achieves state-of-the-art results on the nuScenes dataset for vehicle, pedes-trian, and lane segmentation, showcasing superior performance in static and temporal settings despite being trained solely with sparse signals. We release our code with two new efficient modules used in the architecture: Sparse Feature Pulling, designed for the effective extraction of features from images to BeV, and Submanifold Attention, which en-ables efficient temporal modeling. The code is available at https://github.com/valeoai/PointBeV. Loïck Chambon, Eloi Zablocki, Mickaël Chen, Florent Bartoccioni, Patrick Pérez, Matthieu Cord |
CVPR | 5 |
| 2024 | A Simple Recipe for Language-Guided Domain Generalized SegmentationabstractGeneralization to new domains not seen during training is one of the longstanding challenges in deploying neural networks in real-world applications. Existing generalization techniques either necessitate external images for augmentation, and/or aim at learning invariant representations by imposing various alignment constraints. Largescale pretraining has recently shown promising generalization capabilities, along with the potential of binding different modalities. For instance, the advent of vision-language models like CLIP has opened the doorway for vision models to exploit the textual modality. In this paper, we introduce a simple framework for generalizing semantic segmentation networks by employing language as the source of randomization. Our recipe comprises three key ingredients: (i) the preservation of the intrinsic CLIP robustness through mini-mal fine-tuning, (ii) language-driven local style augmentation, and (iii) randomization by locally mixing the source and augmented styles during training. Extensive experiments report state-of-the-art results on various generalization benchmarks. Code is accessible on the project page11https://astra-vision.github.io/FAMix. Mohammad Fahes, Andrei Bursuc, Patrick Pérez, Raoul de Charette |
CVPR | 4 |
| 2024 | Three Pillars Improving Vision Foundation Model Distillation for LidarabstractSelf-supervised image backbones can be used to address complex 2D tasks (e.g., semantic segmentation, object discovery) very efficiently and with little or no downstream supervision. Ideally, 3D backbones for lidar should be able to inherit these properties after distillation of these powerful 2D features. The most recent methods for image-to-lidar distillation on autonomous driving data show promising results, obtained thanks to distillation methods that keep improving. Yet, we still notice a large performance gap when measuring by linear probing the quality of distilled vs fully supervised features. In this work, instead of focusing only on the distillation method, we study the effect of three pillars for distillation: the 3D backbone, the pretrained 2D backbone, and the pretraining 2D+3D dataset. In particular, thanks to our scalable distillation method named ScaLR, we show that scaling the 2D and 3D backbones and pretraining on diverse datasets leads to a substantial improvement of the feature quality. This allows us to significantly reduce the gap between the quality of distilled and fully-supervised 3D features, and to improve the robustness of the pretrained backbones to domain gaps and perturbations. The code is available at https://github.com/valeoai/ScaLR. Gilles Puy, Spyros Gidaris, Alexandre Boulch, Oriane Siméoni, Corentin Sautier, Patrick Pérez, Andrei Bursuc, Renaud Marlet |
CVPR | 6 |
| 2024 | Reliability in Semantic Segmentation: Can We Use Synthetic Data?
Thibaut Loiseau, Mickaël Chen, Patrick Pérez, Matthieu Cord |
ECCV (23) | 4 |
| 2024 | CLIP-DINOiser: Teaching CLIP a Few DINO Tricks for Open-Vocabulary Semantic Segmentation
Monika Wysoczanska, Oriane Siméoni, Michaël Ramamonjisoa, Andrei Bursuc, Tomasz Trzcinski, Patrick Pérez |
ECCV (61) | 6 |
| 2024 | Winner-takes-all learners are geometry-aware conditional density estimatorsabstractWinner-takes-all training is a simple learning paradigm, which handles ambiguous tasks by predicting a set of plausible hypotheses. Recently, a connection was established between Winner-takes-all training and centroidal Voronoi tessellations, showing that, once trained, hypotheses should quantize optimally the shape of the conditional distribution to predict. However, the best use of these hypotheses for uncertainty quantification is still an open question. In this work, we show how to leverage the appealing geometric properties of the Winner-takes-all learners for conditional density estimation, without modifying its original training scheme. We theoretically establish the advantages of our novel estimator both in terms of quantization and density estimation, and we demonstrate its competitiveness on synthetic and real-world datasets, including audio data. Victor Letzelter, David Perera, Cédric Rommel, Mathieu Fontaine 0002, Slim Essid, Gaël Richard, Patrick Pérez |
ICML | 7 |
| 2024 | Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?abstractMotion forecasting is crucial in enabling autonomous vehicles to anticipate the future trajectories of surrounding agents. To do so, it requires solving mapping, detection, tracking, and then forecasting problems, in a multi-step pipeline. In this complex system, advances in conventional forecasting methods have been made using curated data, i.e., with the assumption of perfect maps, detection, and tracking. This paradigm, however, ignores any errors from upstream modules. Meanwhile, an emerging end-to-end paradigm, that tightly integrates the perception and forecasting architectures into joint training, promises to solve this issue. However, the evaluation protocols between the two methods were so far incompatible and their comparison was not possible. In fact, conventional forecasting methods are usually not trained nor tested in real-world pipelines (e.g., with upstream detection, tracking, and mapping modules). In this work, we aim to bring forecasting models closer to the real-world deployment. First, we propose a unified evaluation pipeline for forecasting methods with real-world perception inputs, allowing us to compare conventional and end-to-end methods for the first time. Second, our in-depth study uncovers a substantial performance gap when transitioning from curated to perception-based data. In particular, we show that this gap (1) stems not only from differences in precision but also from the nature of imperfect inputs provided by perception modules, and that (2) is not trivially reduced by simply finetuning on perception outputs. Based on extensive experiments, we provide recommendations for critical areas that require improvement and guidance towards more robust motion forecasting in the real world. The evaluation library for benchmarking models under standardized and practical conditions is provided: https://github.com/valeoai/MFEval. Loïck Chambon, Eloi Zablocki, Mickaël Chen, Alexandre Alahi, Matthieu Cord, Patrick Pérez |
ICRA | 7 |
| 2024 | ManiPose: Manifold-Constrained Multi-Hypothesis 3D Human Pose EstimationabstractWe propose ManiPose, a manifold-constrained multi-hypothesis model for human-pose 2D-to-3D lifting. We provide theoretical and empirical evidence that, due to the depth ambiguity inherent to monocular 3D human pose estimation, traditional regression models suffer from pose-topology consistency issues, which standard evaluation metrics (MPJPE, P-MPJPE and PCK) fail to assess. ManiPose addresses depth ambiguity by proposing multiple candidate 3D poses for each 2D input, each with its estimated plausibility. Unlike previous multi-hypothesis approaches, ManiPose forgoes generative models, greatly facilitating its training and usage. By constraining the outputs to lie on the human pose manifold, ManiPose guarantees the consistency of all hypothetical poses, in contrast to previous works. We showcase the performance of ManiPose on real-world datasets, where it outperforms state-of-the-art models in pose consistency by a large margin while being very competitive on the MPJPE metric. Cédric Rommel, Victor Letzelter, Nermin Samet, Renaud Marlet, Matthieu Cord, Patrick Pérez, Eduardo Valle |
NeurIPS | 6 |
| 2023 | Unsupervised Object Localization: Observing the Background to Discover ObjectsabstractRecent advances in self-supervised visual representation learning have paved the way for unsupervised methods tackling tasks such as object discovery and instance segmentation. However, discovering objects in an image with no supervision is a very hard task; what are the desired objects, when to separate them into parts, how many are there, and of what classes? The answers to these questions de-pend on the tasks and datasets of evaluation. In this work, we take a different approach and propose to look for the background instead. This way, the salient objects emerge as a by-product without any strong assumption on what an object should be. We propose FOUND, a simple model made of a single conv1 x 1 initialized with coarse background masks extracted from self-supervised patch-based representations. After fast training and refining these seed masks, the model reaches state-of-the-art results on unsupervised saliency detection and object discovery benchmarks. Moreover, we show that our approach yields good results in the unsupervised semantic segmentation retrieval task. The code to reproduce our results is available at https://github.com/valeoai/FOUND. Oriane Siméoni, Chloé Sekkat, Gilles Puy, Antonín Vobecký, Eloi Zablocki, Patrick Pérez |
CVPR | 6 |
| 2023 | OCTET: Object-aware Counterfactual ExplanationsabstractNowadays, deep vision models are being widely deployed in safety-critical applications, e.g., autonomous driving, and explainability of such models is becoming a pressing concern. Among explanation methods, counter-factual explanations aim to find minimal and interpretable changes to the input image that would also change the output of the model to be explained. Such explanations point end-users at the main factors that impact the decision of the model. However, previous methods struggle to explain decision models trained on images with many objects, e.g., urban scenes, which are more difficult to work with but also arguably more critical to explain. In this work, we propose to tackle this issue with an object-centric framework for counterfactual explanation generation. Our method, inspired by recent generative modeling works, encodes the query image into a latent space that is structured in a way to ease object-level manipulations. Doing so, it provides the end-user with control over which search directions (e.g., spatial displacement of objects, style modification, etc.) are to be explored during the counterfactual generation. We conduct a set of experiments on counterfactual explanation benchmarks for driving scenes, and we show that our method can be adapted beyond classification, e.g., to explain semantic segmentation models. To complete our analysis, we design and run a user study that measures the usefulness of counterfactual explanations in understanding a decision model. Code is available at https://github.com/valeoai/OCTET. Mehdi Zemni, Mickaël Chen, Eloi Zablocki, Hédi Ben-Younes, Patrick Pérez, Matthieu Cord |
CVPR | 5 |
| 2023 | PØDA: Prompt-driven Zero-shot Domain AdaptationabstractDomain adaptation has been vastly investigated in computer vision but still requires access to target images at train time, which might be intractable in some uncommon conditions. In this paper, we propose the task of ‘Prompt-driven Zero-shot Domain Adaptation’, where we adapt a model trained on a source domain using only a general description in natural language of the target domain, i.e., a prompt. First, we leverage a pretrained contrastive vision-language model (CLIP) to optimize affine transformations of source features, steering them towards the target text embedding while preserving their content and semantics. To achieve this, we propose Prompt-driven Instance Normalization (PIN). Second, we show that these prompt-driven augmentations can be used to perform zero-shot domain adaptation for semantic segmentation. Experiments demonstrate that our method significantly outperforms CLIP-based style transfer baselines on several datasets for the downstream task at hand, even surpassing one-shot unsupervised domain adaptation. A similar boost is observed on object detection and image classification. The code is available at https://github.com/astra-vision/PODA. Mohammad Fahes, Andrei Bursuc, Patrick Pérez, Raoul de Charette |
ICCV | 4 |
| 2023 | Self-supervised learning with rotation-invariant kernels
Léon Zheng, Gilles Puy, Elisa Riccietti, Patrick Pérez, Rémi Gribonval |
ICLR | 4 |
| 2023 | T-UDA: Temporal Unsupervised Domain Adaptation in Sequential Point CloudsabstractDeep perception models have to reliably cope with an open-world setting of domain shifts induced by different geographic regions, sensor properties, mounting positions, and several other reasons. Since covering all domains with annotated data is technically intractable due to the endless possible variations, researchers focus on unsupervised domain adaptation (UDA) methods that adapt models trained on one (source) domain with annotations available to another (target) domain for which only unannotated data are available. Current predominant methods either leverage semi-supervised approaches, e.g., teacher-student setup, or exploit privileged data, such as other sensor modalities or temporal data consistency. We introduce a novel domain adaptation method that leverages the best of both approaches. Our approach combines input data's temporal and cross-sensor geometric consistency with the mean teacher method. Dubbed T-UDA for “temporal UDA”, such a combination yields massive performance gains for the task of 3D semantic segmentation of driving scenes. Experiments are conducted on Waymo Open Dataset, nuScenes, and SemanticKITTI, for two popular 3D point cloud architectures, Cylinder3D and MinkowskiNet. Our codes are publicly available on https://github.com/ctu-vras/T-UDA. Awet Haileslassie Gebrehiwot, David Hurych, Karel Zimmermann, Patrick Pérez, Tomás Svoboda |
IROS | 4 |
| 2023 | Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysisabstractWe introduce Resilient Multiple Choice Learning (rMCL), an extension of the MCL approach for conditional distribution estimation in regression settings where multiple targets may be sampled for each training input.
Multiple Choice Learning is a simple framework to tackle multimodal density estimation, using the Winner-Takes-All (WTA) loss for a set of hypotheses. In regression settings, the existing MCL variants focus on merging the hypotheses, thereby eventually sacrificing the diversity of the predictions. In contrast, our method relies on a novel learned scoring scheme underpinned by a mathematical framework based on Voronoi tessellations of the output space, from which we can derive a probabilistic interpretation.
After empirically validating rMCL with experiments on synthetic data, we further assess its merits on the sound source localization problem, demonstrating its practical usefulness and the relevance of its interpretation. Victor Letzelter, Mathieu Fontaine 0002, Mickaël Chen, Patrick Pérez, Slim Essid, Gaël Richard |
NeurIPS | 4 |
| 2023 | POP-3D: Open-Vocabulary 3D Occupancy Prediction from ImagesabstractWe describe an approach to predict open-vocabulary 3D semantic voxel occupancy map from input 2D images with the objective of enabling 3D grounding, segmentation and retrieval of free-form language queries. This is a challenging problem because of the 2D-3D ambiguity and the open-vocabulary nature of the target tasks, where obtaining annotated training data in 3D is difficult. The contributions of this work are three-fold.
First, we design a new model architecture for open-vocabulary 3D semantic occupancy prediction. The architecture consists of a 2D-3D encoder together with occupancy prediction and 3D-language heads. The output is a dense voxel map of 3D grounded language embeddings enabling a range of open-vocabulary tasks.
Second, we develop a tri-modal self-supervised learning algorithm that leverages three modalities: (i) images, (ii) language and (iii) LiDAR point clouds, and enables training the proposed architecture using a strong pre-trained vision-language model without the need for any 3D manual language annotations.
Finally, we demonstrate quantitatively the strengths of the proposed model on several open-vocabulary tasks:
Zero-shot 3D semantic segmentation using existing datasets; 3D grounding and retrieval of free-form language queries, using a small dataset that we propose as an extension of nuScenes. You can find the project page here https://vobecant.github.io/POP3D. Antonín Vobecký, Oriane Siméoni, David Hurych, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Josef Sivic |
NeurIPS | 6 |
| 2023 | LiDARTouch: Monocular metric depth estimation with a few-beam LiDAR
Florent Bartoccioni, Eloi Zablocki, Patrick Pérez, Matthieu Cord, Karteek Alahari |
Comput. Vis. Image Underst. | 3 |
| 2023 | Cross-Modal Learning for Domain Adaptation in 3D Semantic SegmentationabstractDomain adaptation is an important task to enable learning when labels are scarce. While most works focus only on the image modality, there are many important multi-modal datasets. In order to leverage multi-modality for domain adaptation, we propose cross-modal learning, where we enforce consistency between the predictions of two modalities via mutual mimicking. We constrain our network to make correct predictions on labeled data and consistent predictions across modalities on unlabeled target-domain data. Experiments in unsupervised and semi-supervised domain adaptation settings prove the effectiveness of this novel domain adaptation strategy. Specifically, we evaluate on the task of 3D semantic segmentation from either the 2D image, the 3D point cloud or from both. We leverage recent driving datasets to produce a wide variety of domain adaptation scenarios including changes in scene layout, lighting, sensor setup and weather, as well as the synthetic-to-real setup. Our method significantly improves over previous uni-modal adaptation baselines on all adaption scenarios. Code will be made available upon publication. Maximilian Jaritz, Raoul de Charette, Émilie Wirbel, Patrick Pérez |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Decaf: Monocular Deformation Capture for Face and Hand InteractionsabstractExisting methods for 3D tracking from monocular RGB videos predominantly consider articulated and rigid objects ( e.g. , two hands or humans interacting with rigid environments). Modelling dense non-rigid object deformations in this setting ( e.g. when hands are interacting with a face), remained largely unaddressed so far, although such effects can improve the realism of the downstream applications such as AR/VR, 3D virtual avatar communications, and character animations. This is due to the severe ill-posedness of the monocular view setting and the associated challenges ( e.g. , in acquiring a dataset for training and evaluation or obtaining the reasonable non-uniform stiffness of the deformable object). While it is possible to naïvely track multiple non-rigid objects independently using 3D templates or parametric 3D models, such an approach would suffer from multiple artefacts in the resulting 3D estimates such as depth ambiguity, unnatural intra-object collisions and missing or implausible deformations. Hence, this paper introduces the first method that addresses the fundamental challenges depicted above and that allows tracking human hands interacting with human faces in 3D from single monocular RGB videos. We model hands as articulated objects inducing non-rigid face deformations during an active interaction. Our method relies on a new hand-face motion and interaction capture dataset with realistic face deformations acquired with a markerless multi-view camera system. As a pivotal step in its creation, we process the reconstructed raw 3D shapes with position-based dynamics and an approach for non-uniform stiffness estimation of the head tissues, which results in plausible annotations of the surface deformations, hand-face contact regions and head-hand positions. At the core of our neural approach are a variational auto-encoder supplying the hand-face depth prior and modules that guide the 3D tracking by estimating the contacts and the deformations. Our final 3D hand and face reconstructions are realistic and more plausible compared to several baselines applicable in our setting, both quantitatively and qualitatively. https://vcai.mpi-inf.mpg.de/projects/Decaf Soshi Shimada, Vladislav Golyanik, Patrick Pérez, Christian Theobalt |
ACM Trans. Graph. | 3 |
| 2022 | Raw High-Definition Radar for Multi-Task LearningabstractWith their robustness to adverse weather conditions and ability to measure speeds, radar sensors have been part of the automotive landscape for more than two decades. Recent progress toward High Definition (HD) Imaging radar has driven the angular resolution below the degree, thus approaching laser scanning performance. However, the amount of data a HD radar delivers and the computational cost to estimate the angular positions remain a challenge. In this paper, we propose a novel HD radar sensing model, FFT-RadNet, that eliminates the overhead of computing the range-azimuth-Doppler 3D tensor, learning instead to recover angles from a range-Doppler spectrum. FFT-RadNet is trained both to detect vehicles and to segment free driving space. On both tasks, it competes with the most recent radar-based models while requiring less compute and memory. Also, we collected and annotated 2-hour worth of raw data from synchronized automotive-grade sensors (camera, laser, HD radar) in various environments (city street, highway, countryside road). This unique dataset, nick-named RADIal for “Radar, LiDAR et al.”, is available at https://github.com/valeoai/RADIal. Julien Rebut, Arthur Ouaknine, Waqas Malik, Patrick Pérez |
CVPR | 4 |
| 2022 | STEEX: Steering Counterfactual Explanations with Semantics
Paul Jacob, Eloi Zablocki, Hédi Ben-Younes, Mickaël Chen, Patrick Pérez, Matthieu Cord |
ECCV (12) | 5 |
| 2022 | HULC: 3D HUman Motion Capture with Pose Manifold SampLing and Dense Contact Guidance
Soshi Shimada, Vladislav Golyanik, Zhi Li 0055, Patrick Pérez, Weipeng Xu, Christian Theobalt |
ECCV (22) | 4 |
| 2022 | Active Learning Strategies for Weakly-Supervised Object Detection
Huy V. Vo, Oriane Siméoni, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Jean Ponce |
ECCV (30) | 5 |
| 2022 | Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-Modal Distillation
Antonín Vobecký, David Hurych, Oriane Siméoni, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Josef Sivic |
ECCV (38) | 6 |
| 2022 | Diverse Probabilistic Trajectory Forecasting with Admissibility ConstraintsabstractPredicting multiple trajectories for road users is important for automated driving systems: ego-vehicle motion planning indeed requires a clear view of the possible motions of the surrounding agents. However, the generative models used for multiple-trajectory forecasting suffer from a lack of diversity in their proposals. To avoid this form of collapse, we propose a novel method for structured prediction of diverse trajectories. To this end, we complement an underlying pretrained generative model with a diversity component, based on a determinantal point process (DPP). We balance and structure this diversity with the inclusion of knowledge-based quality constraints, independent from the underlying generative model. We combine these two novel components with a gating operation, ensuring that the predictions are both diverse and within the drivable area. We demonstrate on the nuScenes driving dataset the relevance of our compound approach, which yields significant improvements in the diversity and the quality of the generated trajectories. Laura Calem, Hédi Ben-Younes, Patrick Pérez, Nicolas Thome |
ICPR | 3 |
| 2022 | Explainability of Deep Vision-Based Autonomous Driving Systems: Review and Challenges
Eloi Zablocki, Hédi Ben-Younes, Patrick Pérez, Matthieu Cord |
Int. J. Comput. Vis. | 3 |
| 2022 | Spherical perspective on learning with normalization layers
Simon Roburin, Yann de Mont-Marin, Andrei Bursuc, Renaud Marlet, Patrick Pérez, Mathieu Aubry |
Neurocomputing | 5 |
| 2022 | Confidence Estimation via Auxiliary ModelsabstractReliably quantifying the confidence of deep neural classifiers is a challenging yet fundamental requirement for deploying such models in safety-critical applications. In this paper, we introduce a novel target criterion for model confidence, namely the true class probability (TCP). We show that TCP offers better properties for confidence estimation than standard maximum class probability (MCP). Since the true class is by essence unknown at test time, we propose to learn TCP criterion from data with an auxiliary model, introducing a specific learning scheme adapted to this context. We evaluate our approach on the task of failure prediction and of self-training with pseudo-labels for domain adaptation, which both necessitate effective confidence estimates. Extensive experiments are conducted for validating the relevance of the proposed approach in each task. We study various network architectures and experiment with small and large datasets for image classification and semantic segmentation. In every tested benchmark, our approach outperforms strong baselines. Charles Corbière, Nicolas Thome, Antoine Saporta, Matthieu Cord, Patrick Pérez |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Driving behavior explanation with multi-level fusion
Hédi Ben-Younes, Eloi Zablocki, Patrick Pérez, Matthieu Cord |
Pattern Recognit. | 3 |
| 2022 | Deep Reinforcement Learning for Autonomous Driving: A SurveyabstractWith the development of deep representation learning, the domain of reinforcement learning (RL) has become a powerful learning framework now capable of learning complex policies in high dimensional environments. This review summarises deep reinforcement learning (DRL) algorithms and provides a taxonomy of automated driving tasks where (D)RL methods have been employed, while addressing key computational challenges in real world deployment of autonomous driving agents. It also delineates adjacent domains such as behavior cloning, imitation learning, inverse reinforcement learning that are related but are not classical RL algorithms. The role of simulators in training agents, methods to validate, test and robustify existing solutions in RL are discussed. Bangalore Ravi Kiran, Ibrahim Sobh, Victor Talpaert, Patrick Mannion, Ahmad A. Al Sallab, Senthil Kumar Yogamani, Patrick Pérez |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2022 | Detecting 32 Pedestrian Attributes for Autonomous VehiclesabstractPedestrians are arguably one of the most safety-critical road users to consider for autonomous vehicles in urban areas. In this paper, we address the problem of jointly detecting pedestrians and recognizing 32 pedestrian attributes from a single image. These encompass visual appearance and behavior, and also include the forecasting of road crossing, which is a main safety concern. For this, we introduce a Multi-Task Learning (MTL) model relying on a composite field framework, which achieves both goals in an efficient way. Each field spatially locates pedestrian instances and aggregates attribute predictions over them. This formulation naturally leverages spatial context, making it well suited to low resolution scenarios such as autonomous driving. By increasing the number of attributes jointly learned, we highlight an issue related to the scales of gradients, which arises in MTL with numerous tasks. We solve it by normalizing the gradients coming from different objective functions when they join at the fork in the network architecture during the backward pass, referred to as fork-normalization. Experimental validation is performed on JAAD, a dataset providing numerous attributes for pedestrian analysis from autonomous vehicles, and shows competitive detection and attribute recognition results, as well as a more stable MTL training. Taylor Mordan, Matthieu Cord, Patrick Pérez, Alexandre Alahi |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Artificial Dummies for Urban Dataset AugmentationabstractExisting datasets for training pedestrian detectors in images suffer from limited appearance and pose variation. The most challenging scenarios are rarely included because they are too difficult to capture due to safety reasons, or they are very unlikely to happen. The strict safety requirements in assisted and autonomous driving applications call for an extra high detection accuracy also in these rare situations. Having the ability to generate people images in arbitrary poses, with arbitrary appearances and embedded in different background scenes with varying illumination and weather conditions, is a crucial component for the development and testing of such applications. The contributions of this paper are three-fold. First, we describe an augmentation method for the controlled synthesis of urban scenes containing people, thus producing rare or never-seen situations. This is achieved with a data generator (called DummyNet) with disentangled control of the pose, the appearance, and the target background scene. Second, the proposed generator relies on novel network architecture and associated loss that takes into account the segmentation of the foreground person and its composition into the background scene. Finally, we demonstrate that the data generated by our DummyNet improve the performance of several existing person detectors across various datasets as well as in challenging situations, such as night-time conditions, where only a limited amount of training data is available. In the setup with only day-time data available, we improve the night-time detector by 17% log-average miss rate over the detector trained with the day-time data only. Antonín Vobecký, David Hurych, Michal Uricár, Patrick Pérez, Josef Sivic |
AAAI | 4 |
| 2021 | Localizing Objects with Self-supervised Transformers and no Labels
Oriane Siméoni, Gilles Puy, Huy V. Vo, Simon Roburin, Spyros Gidaris, Andrei Bursuc, Patrick Pérez, Renaud Marlet, Jean Ponce |
BMVC | 7 |
| 2021 | OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised LearningabstractLearning image representations without human supervision is an important and active research field. Several recent approaches have successfully leveraged the idea of making such a representation invariant under different types of perturbations, especially via contrastive-based instance discrimination training. Although effective visual representations should indeed exhibit such invariances, there are other important characteristics, such as encoding contextual reasoning skills, for which alternative reconstruction-based approaches might be better suited.With this in mind, we propose a teacher-student scheme to learn representations by training a convolutional net to reconstruct a bag-of-visual-words (BoW) representation of an image, given as input a perturbed version of that same image. Our strategy performs an online training of both the teacher network (whose role is to generate the BoW targets) and the student network (whose role is to learn representations), along with an online update of the visual-words vocabulary (used for the BoW targets). This idea effectively enables fully online BoW-guided unsupervised learning. Extensive experiments demonstrate the interest of our BoWbased strategy, which surpasses previous state-of-the-art methods (including contrastive-based ones) in several applications. For instance, in downstream tasks such Pascal object detection, Pascal classification and Places205 classification, our method improves over all prior unsupervised approaches, thus establishing new state-of-the-art results that are also significantly better even than those of supervised pre-training. We provide the implementation code at https://github.com/valeoai/obow. Spyros Gidaris, Andrei Bursuc, Gilles Puy, Nikos Komodakis, Matthieu Cord, Patrick Pérez |
CVPR | 6 |
| 2021 | Semantic Palette: Guiding Scene Generation With Class ProportionsabstractDespite the recent progress of generative adversarial networks (GANs) at synthesizing photo-realistic images, producing complex urban scenes remains a challenging problem. Previous works break down scene generation into two consecutive phases: unconditional semantic layout synthesis and image synthesis conditioned on layouts. In this work, we propose to condition layout generation as well for higher semantic control: given a vector of class proportions, we generate layouts with matching composition. To this end, we introduce a conditional framework with novel architecture designs and learning objectives, which effectively accommodates class proportions to guide the scene generation process. The proposed architecture also allows partial layout editing with interesting applications. Thanks to the semantic control, we can produce layouts close to the real distribution, helping enhance the whole scene generation process. On different metrics and urban scene benchmarks, our models outperform existing baselines. Moreover, we demonstrate the merit of our approach for data augmentation: semantic segmenters trained on real layout-image pairs along with additional ones generated by our approach outperform models only trained on real pairs. Guillaume Le Moing, Himalaya Jain, Patrick Pérez, Matthieu Cord |
CVPR | 4 |
| 2021 | Multi-View Radar Semantic SegmentationabstractUnderstanding the scene around the ego-vehicle is key to assisted and autonomous driving. Nowadays, this is mostly conducted using cameras and laser scanners, despite their reduced performance in adverse weather conditions. Automotive radars are low-cost active sensors that measure properties of surrounding objects, including their relative speed, and have the key advantage of not being impacted by rain, snow or fog. However, they are seldom used for scene understanding due to the size and complexity of radar raw data and the lack of annotated datasets. Fortunately, recent open-sourced datasets have opened up research on classification, object detection and semantic segmentation with raw radar signals using end-to-end trainable models. In this work, we propose several novel architectures, and their associated losses, which analyse multiple "views" of the range-angle-Doppler radar tensor to segment it semantically. Experiments conducted on the recent CARRADA dataset demonstrate that our best model outperforms alternative models, derived either from the semantic segmentation of natural images or from radar scene understanding, while requiring significantly fewer parameters. Both our code and trained models are available at https://github.com/valeoai/MVRSS. Arthur Ouaknine, Alasdair Newson, Patrick Pérez, Florence Tupin, Julien Rebut |
ICCV | 3 |
| 2021 | Multi-Target Adversarial Frameworks for Domain Adaptation in Semantic SegmentationabstractIn this work, we address the task of unsupervised domain adaptation (UDA) for semantic segmentation in presence of multiple target domains: The objective is to train a single model that can handle all these domains at test time. Such a multi-target adaptation is crucial for a variety of scenarios that real-world autonomous systems must handle. It is a challenging setup since one faces not only the domain gap between the labeled source set and the un-labeled target set, but also the distribution shifts existing within the latter among the different target domains. To this end, we introduce two adversarial frameworks: (i) multi-discriminator, which explicitly aligns each target domain to its counterparts, and (ii) multi-target knowledge transfer, which learns a target-agnostic model thanks to a multi-teacher/single-student distillation mechanism. The evaluation is done on four newly-proposed multi-target bench-marks for UDA in semantic segmentation. In all tested scenarios, our approaches consistently outperform baselines, setting competitive standards for the novel task. Antoine Saporta, Matthieu Cord, Patrick Pérez |
ICCV | 4 |
| 2021 | StyleLess layer: Improving robustness for real-world drivingabstractDeep Neural Networks (DNNs) are a critical component for self-driving vehicles. They achieve impressive performance by reaping information from high amounts of labeled data. Yet, the full complexity of the real world cannot be encapsulated in the training data, no matter how big the dataset, and DNNs can hardly generalize to unseen conditions. Robustness to various image corruptions, caused by changing weather conditions or sensor degradation and aging, is crucial for safety when such vehicles are deployed in the real world. We address this problem through a novel type of layer, dubbed StyleLess, which enables DNNs to learn robust and informative features that can cope with varying external conditions. We propose multiple variations of this layer that can be integrated in most of the architectures and trained jointly with the main task. We validate our contribution on typical autonomous-driving tasks (detection, semantic segmentation), showing that in most cases, this approach improves predictive performance on unseen conditions (fog, rain), while preserving performance on seen conditions and objects. Julien Rebut, Andrei Bursuc, Patrick Pérez |
IROS | 3 |
| 2021 | Large-Scale Unsupervised Object DiscoveryabstractExisting approaches to unsupervised object discovery (UOD) do not scale up to large datasets without approximations that compromise their performance. We propose a novel formulation of UOD as a ranking problem, amenable to the arsenal of distributed methods available for eigenvalue problems and link analysis. Through the use of self-supervised features, we also demonstrate the first effective fully unsupervised pipeline for UOD. Extensive experiments on COCO~\cite{Lin2014cocodataset} and OpenImages~\cite{openimages} show that, in the single-object discovery setting where a single prominent object is sought in each image, the proposed LOD (Large-scale Object Discovery) approach is on par with, or better than the state of the art for medium-scale datasets (up to 120K images), and over 37\% better than the only other algorithms capable of scaling up to 1.7M images. In the multi-object discovery setting where multiple objects are sought in each image, the proposed LOD is over 14\% better in average precision (AP) than all other methods for datasets ranging from 20K to 1.7M images. Using self-supervised features, we also show that the proposed method obtains state-of-the-art UOD performance on OpenImages. Huy V. Vo, Elena Sizikova, Cordelia Schmid, Patrick Pérez, Jean Ponce |
NeurIPS | 4 |
| 2021 | Handling new target classes in semantic segmentation with domain adaptation
Maxime Bucher, Matthieu Cord, Patrick Pérez |
Comput. Vis. Image Underst. | 4 |
| 2021 | Neural monocular 3D human motion capture with physical awarenessabstractWe present a new trainable system for physically plausible markerless 3D human motion capture, which achieves state-of-the-art results in a broad range of challenging scenarios. Unlike most neural methods for human motion capture, our approach, which we dub "physionical", is aware of physical and environmental constraints. It combines in a fully-differentiable way several key innovations, i.e. , 1) a proportional-derivative controller, with gains predicted by a neural network, that reduces delays even in the presence of fast motions, 2) an explicit rigid body dynamics model and 3) a novel optimisation layer that prevents physically implausible foot-floor penetration as a hard constraint. The inputs to our system are 2D joint keypoints, which are canonicalised in a novel way so as to reduce the dependency on intrinsic camera parameters---both at train and test time. This enables more accurate global translation estimation without generalisability loss. Our model can be finetuned only with 2D annotations when the 3D annotations are not available. It produces smooth and physically-principled 3D motions in an interactive frame rate in a wide variety of challenging scenes, including newly recorded ones. Its advantages are especially noticeable on in-the-wild sequences that significantly differ from common 3D pose estimation benchmarks such as Human 3.6M and MPI-INF-3DHP. Qualitative results are provided in the supplementary video. Soshi Shimada, Vladislav Golyanik, Weipeng Xu, Patrick Pérez, Christian Theobalt |
ACM Trans. Graph. | 4 |
| 2020 | The Missing Data Encoder: Cross-Channel Image Completion with Hide-and-Seek Adversarial NetworkabstractImage completion is the problem of generating whole images from fragments only. It encompasses inpainting (generating a patch given its surrounding), reverse inpainting/extrapolation (generating the periphery given the central patch) as well as colorization (generating one or several channels given other ones). In this paper, we employ a deep network to perform image completion, with adversarial training as well as perceptual and completion losses, and call it the “missing data encoder” (MDE). We consider several configurations based on how the seed fragments are chosen. We show that training MDE for “random extrapolation and colorization” (MDE-REC), i.e. using random channel-independent fragments, allows a better capture of the image semantics and geometry. MDE training makes use of a novel “hide-and-seek” adversarial loss, where the discriminator seeks the original non-masked regions, while the generator tries to hide them. We validate our models qualitatively and quantitatively on several datasets, showing their interest for image completion, representation learning as well as face occlusion handling. Arnaud Dapogny, Matthieu Cord, Patrick Pérez |
AAAI | 3 |
| 2020 | Learning Representations by Predicting Bags of Visual WordsabstractSelf-supervised representation learning targets to learn convnet-based image representations from unlabeled data. Inspired by the success of NLP methods in this area, in this work we propose a self-supervised approach based on spatially dense image descriptions that encode discrete visual concepts, here called visual words. To build such discrete representations, we quantize the feature maps of a first pre-trained self-supervised convnet, over a k-means based vocabulary. Then, as a self-supervised task, we train another convnet to predict the histogram of visual words of an image (i.e., its Bag-of-Words representation) given as input a perturbed version of that image. The proposed task forces the convnet to learn perturbation-invariant and context-aware image features, useful for downstream image understanding tasks. We extensively evaluate our method and demonstrate very strong empirical results, e.g., our pre-trained self-supervised representations transfer better on detection task and similarly on classification over classes "unseen'' during pre-training, when compared to the supervised case. This also shows that the process of image discretization into visual words can provide the basis for very powerful self-supervised approaches in the image domain, thus allowing further connections to be made to related methods from the NLP domain that have been extremely successful so far. Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, Matthieu Cord |
CVPR | 4 |
| 2020 | xMUDA: Cross-Modal Unsupervised Domain Adaptation for 3D Semantic SegmentationabstractUnsupervised Domain Adaptation (UDA) is crucial to tackle the lack of annotations in a new domain. There are many multi-modal datasets, but most UDA approaches are uni-modal. In this work, we explore how to learn from multi-modality and propose cross-modal UDA (xMUDA) where we assume the presence of 2D images and 3D point clouds for 3D semantic segmentation. This is challenging as the two input spaces are heterogeneous and can be impacted differently by domain shift. In xMUDA, modalities learn from each other through mutual mimicking, disentangled from the segmentation objective, to prevent the stronger modality from adopting false predictions from the weaker one. We evaluate on new UDA scenarios including day-to-night, country-to-country and dataset-to-dataset, leveraging recent autonomous driving datasets. xMUDA brings large improvements over uni-modal UDA on all tested scenarios, and is complementary to state-of-the-art UDA techniques. Code is available at https://github.com/valeoai/xmuda. Maximilian Jaritz, Raoul de Charette, Émilie Wirbel, Patrick Pérez |
CVPR | 5 |
| 2020 | StyleRig: Rigging StyleGAN for 3D Control Over Portrait ImagesabstractStyleGAN generates photorealistic portrait images of faces with eyes, teeth, hair and context (neck, shoulders, background), but lacks a rig-like control over semantic face parameters that are interpretable in 3D, such as face pose, expressions, and scene illumination. Three-dimensional morphable face models (3DMMs) on the other hand offer control over the semantic parameters, but lack photorealism when rendered and only model the face interior, not other parts of a portrait image (hair, mouth interior, background). We present the first method to provide a face rig-like control over a pretrained and fixed StyleGAN via a 3DMM. A new rigging network, is trained between the 3DMM's semantic parameters and StyleGAN's input. The network is trained in a self-supervised manner, without the need for manual annotations. At test time, our method generates portrait images with the photorealism of StyleGAN and provides explicit control over the 3D semantic parameters of the face. Ayush Tewari, Mohamed A. Elgharib, Gaurav Bharaj, Florian Bernard 0001, Hans-Peter Seidel, Patrick Pérez, Michael Zollhöfer, Christian Theobalt |
CVPR | 6 |
| 2020 | QuEST: Quantized Embedding Space for Transferring Knowledge
Himalaya Jain, Spyros Gidaris, Nikos Komodakis, Patrick Pérez, Matthieu Cord |
ECCV (21) | 4 |
| 2020 | Toward Unsupervised, Multi-object Discovery in Large-Scale Image Collections
Huy V. Vo, Patrick Pérez, Jean Ponce |
ECCV (23) | 2 |
| 2020 | This Dataset Does Not Exist: Training Models from Generated ImagesabstractCurrent generative networks are increasingly proficient in generating high-resolution realistic images. These generative networks, especially the conditional ones, can potentially become a great tool for providing new image datasets. This naturally brings the question: Can we train a classifier only on the generated data? This potential availability of nearly unlimited amounts of training data challenges standard practices for training machine learning models, which have been crafted across the years for limited and fixed size datasets. In this work we investigate this question and its related challenges. We identify ways to improve significantly the performance over naive training on randomly generated images with regular heuristics. We propose three standalone techniques that can be applied at different stages of the pipeline, i.e., data generation, training on generated data, and deploying on real data. We evaluate our proposed approaches on a subset of the ImageNet dataset and show encouraging results compared to classifiers trained on real images. Victor Besnier, Himalaya Jain, Andrei Bursuc, Matthieu Cord, Patrick Pérez |
ICASSP | 5 |
| 2020 | CARRADA Dataset: Camera and Automotive Radar with Range- Angle- Doppler AnnotationsabstractHigh quality perception is essential for autonomous driving (AD) systems. To reach the accuracy and robustness thatare required by such systems, several types of sensors must be combined. Currently, mostly cameras and laser scanners (lidar) are deployed to build a representation of the world around the vehicle. While radar sensors have been used fora long time in the automotive industry, they are still under-used for AD despite their appealing characteristics (notably, their ability to measure the relative speed of obstacles and to operate even in adverse weather conditions). To alarge extent, this situation is due to the relative lack of automotive datasets with real radar signals that are both raw and annotated. In this work, we introduce CARRADA, a dataset of synchronized camera and radar recordings with range-angle-Doppler annotations. We also present a semi-automatic annotation approach, which was used to annotate the dataset, and a radar semantic segmentation baseline, which we evaluate on several metrics. Both our code and dataset are available online. Arthur Ouaknine, Alasdair Newson, Julien Rebut, Florence Tupin, Patrick Pérez |
ICPR | 5 |
| 2020 | Adversarial frontier stitching for remote neural network watermarking
Erwan Le Merrer, Patrick Pérez, Gilles Trédan |
Neural Comput. Appl. | 2 |
| 2020 | ROAM: A Rich Object Appearance Model with Application to RotoscopingabstractRotoscoping, the detailed delineation of scene elements through a video shot, is a painstaking task of tremendous importance in professional post-production pipelines. While pixel-wise segmentation techniques can help for this task, professional rotoscoping tools rely on parametric curves that offer the artists a much better interactive control on the definition, editing and manipulation of the segments of interest. Sticking to this prevalent rotoscoping paradigm, we propose a novel framework to capture and track the visual aspect of an arbitrary object in a scene, given an initial closed outline of this object. This model combines a collection of local foreground/background appearance models spread along the outline, a global appearance model of the enclosed object and a set of distinctive foreground landmarks. The structure of this rich appearance model allows simple initialization, efficient iterative optimization with exact minimization at each step, and on-line adaptation in videos. We further extend this model by so-called trimaps which serve as an input to alpha-matting algorithms to allow truly seamless compositing. To this end, we leverage local classifiers attached to the roto-curves to define a confidence measure that is well-suited to define trimaps with adaptive band-widths. The resulting trimaps are parametric, temporally consistent and remain fully editable by the artist. We demonstrate qualitatively and quantitatively the merit of this framework through comparisons with tools based on either dynamic segmentation with a closed curve or pixel-wise binary labelling. Juan-Manuel Pérez-Rúa, Ondrej Miksik, Tomás Crivelli, Patrick Bouthemy, Philip Torr 0001, Patrick Pérez |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | High-Fidelity Monocular Face Reconstruction Based on an Unsupervised Model-Based Face AutoencoderabstractIn this work, we propose a novel model-based deep convolutional autoencoder that addresses the highly challenging problem of reconstructing a 3D human face from a single in-the-wild color image. To this end, we combine a convolutional encoder network with an expert-designed generative model that serves as decoder. The core innovation is the differentiable parametric decoder that encapsulates image formation analytically based on a generative model. Our decoder takes as input a code vector with exactly defined semantic meaning that encodes detailed face pose, shape, expression, skin reflectance, and scene illumination. Due to this new way of combining CNN-based with model-based face reconstruction, the CNN-based encoder learns to extract semantically meaningful parameters from a single monocular input image. For the first time, a CNN encoder and an expert-designed generative model can be trained end-to-end in an unsupervised manner, which renders training on very large (unlabeled) real world datasets feasible. The obtained reconstructions compare favorably to current state-of-the-art approaches in terms of quality and richness of representation. This work is an extended version of [1] , where we additionally present a stochastic vertex sampling technique for faster training of our networks, and moreover, we propose and evaluate analysis-by-synthesis and shape-from-shading refinement approaches to achieve a high-fidelity reconstruction. Ayush Tewari, Michael Zollhöfer, Florian Bernard 0001, Pablo Garrido 0001, Hyeongwoo Kim, Patrick Pérez, Christian Theobalt |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Weakly Supervised Representation Learning for Audio-Visual Scene AnalysisabstractAudio-visual (AV) representation learning is an important task from the perspective of designing machines with the ability to understand complex events. To this end, we propose a novel multimodal framework that instantiates multiple instance learning. Specifically, we develop methods that identify events and localize corresponding AV cues in unconstrained videos. Importantly, this is done using weak labels where only video-level event labels are known without any information about their location in time. We show that the learnt representations are useful for performing several tasks such as event/object classification, audio event detection, audio source separation and visual object localization. An important feature of our method is its capacity to learn from unsynchronized audio-visual events. We also demonstrate our framework's ability to separate out the audio source of interest through a novel use of nonnegative matrix factorization. State-of-the-art classification results, with a F1-score of 65.0, are achieved on DCASE 2017 smart cars challenge data with promising generalization to diverse object types such as musical instruments. Visualizations of localized visual regions and audio segments substantiate our system's efficacy, especially when dealing with noisy situations where modality-specific cues appear asynchronously. Sanjeel Parekh, Slim Essid, Alexey Ozerov, Ngoc Q. K. Duong, Patrick Pérez, Gaël Richard |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | PIE: portrait image embedding for semantic controlabstractEditing of portrait images is a very popular and important research topic with a large variety of applications. For ease of use, control should be provided via a semantically meaningful parameterization that is akin to computer animation controls. The vast majority of existing techniques do not provide such intuitive and fine-grained control, or only enable coarse editing of a single isolated control parameter. Very recently, high-quality semantically controlled editing has been demonstrated, however only on synthetically created StyleGAN images. We present the first approach for embedding real portrait images in the latent space of StyleGAN, which allows for intuitive editing of the head pose, facial expression, and scene illumination in the image. Semantic editing in parameter space is achieved based on StyleRig, a pretrained neural network that maps the control space of a 3D morphable face model to the latent space of the GAN. We design a novel hierarchical non-linear optimization problem to obtain the embedding. An identity preservation energy term allows spatially coherent edits while maintaining facial integrity. Our approach runs at interactive frame rates and thus allows the user to explore the space of possible edits. We evaluate our approach on a wide set of portrait photos, compare it to the current state of the art, and validate the effectiveness of its components in an ablation study. Ayush Tewari, Mohamed A. Elgharib, Mallikarjun B. R. 0001, Florian Bernard 0001, Hans-Peter Seidel, Patrick Pérez, Michael Zollhöfer, Christian Theobalt |
ACM Trans. Graph. | 6 |
| 2019 | SoDeep: A Sorting Deep Net to Learn Ranking Loss SurrogatesabstractSeveral tasks in machine learning are evaluated using non-differentiable metrics such as mean average precision or Spearman correlation. However, their non-differentiability prevents from using them as objective functions in a learning framework. Surrogate and relaxation methods exist but tend to be specific to a given metric. In the present work, we introduce a new method to learn approximations of such non-differentiable objective functions. Our approach is based on a deep architecture that approximates the sorting of arbitrary sets of scores. It is trained virtually for free using synthetic data. This sorting deep (SoDeep) net can then be combined in a plug-and-play manner with existing deep architectures. We demonstrate the interest of our approach in three different tasks that require ranking: Cross-modal text-image retrieval, multi-label image classification and visual memorability ranking. Our approach yields very competitive results on these three tasks, which validates the merit and the flexibility of SoDeep as a proxy for sorting operation in ranking-based losses. Martin Engilberge, Louis Chevallier, Patrick Pérez, Matthieu Cord |
CVPR | 3 |
| 2019 | A Flexible Convolutional Solver for Fast Style TransfersabstractWe propose a new flexible deep convolutional neural network (convnet) to perform fast neural style transfers. Our network is trained to solve approximately, but rapidly, the artistic style transfer problem of [Gatys et al.] for arbritary styles. While solutions already exist, our network is uniquely flexible by design: it can be manipulated at runtime to enforce new constraints on the final output. As examples, we show that it can be modified to perform tasks such as fast photorealistic style transfer, or fast video style transfer with short term consistency, with no retraining. This flexibility stems from the proposed architecture which is obtained by unrolling the gradient descent algorithm used in [Gatys et al.]. Regularisations added to [Gatys et al.] to solve a new task can be reported on-the-fly in our network, even after training. Gilles Puy, Patrick Pérez |
CVPR | 2 |
| 2019 | FML: Face Model Learning From VideosabstractMonocular image-based 3D reconstruction of faces is a long-standing problem in computer vision. Since image data is a 2D projection of a 3D face, the resulting depth ambiguity makes the problem ill-posed. Most existing methods rely on data-driven priors that are built from limited 3D face scans. In contrast, we propose multi-frame video-based self-supervised training of a deep network that (i) learns a face identity model both in shape and appearance while (ii) jointly learning to reconstruct 3D faces. Our face model is learned using only corpora of in-the-wild video clips collected from the Internet. This virtually endless source of training data enables learning of a highly general 3D face model. In order to achieve this, we propose a novel multi-frame consistency loss that ensures consistent shape and appearance across multiple frames of a subject's face, thus minimizing depth ambiguity. At test time we can use an arbitrary number of frames, so that we can perform both monocular as well as multi-frame reconstruction. Ayush Tewari, Florian Bernard 0001, Pablo Garrido 0001, Gaurav Bharaj, Mohamed A. Elgharib, Hans-Peter Seidel, Patrick Pérez, Michael Zollhöfer, Christian Theobalt |
CVPR | 7 |
| 2019 | Unsupervised Image Matching and Object Discovery as OptimizationabstractLearning with complete or partial supervision is power- ful but relies on ever-growing human annotation efforts. As a way to mitigate this serious problem, as well as to serve specific applications, unsupervised learning has emerged as an important field of research. In computer vision, unsu- pervised learning comes in various guises. We focus here on the unsupervised discovery and matching of object cate- gories among images in a collection, following the work of Cho et al. [12]. We show that the original approach can be reformulated and solved as a proper optimization problem. Experiments on several benchmarks establish the merit of our approach. Huy V. Vo, Francis R. Bach, Minsu Cho, Kai Han 0001, Yann LeCun, Patrick Pérez, Jean Ponce |
CVPR | 6 |
| 2019 | ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic SegmentationabstractSemantic segmentation is a key problem for many computer vision tasks. While approaches based on convolutional neural networks constantly break new records on different benchmarks, generalizing well to diverse testing environments remains a major challenge. In numerous real-world applications, there is indeed a large gap between data distributions in train and test domains, which results in severe performance loss at run-time. In this work, we address the task of unsupervised domain adaptation in semantic segmentation with losses based on the entropy of the pixel-wise predictions. To this end, we propose two novel, complementary methods using (i) entropy loss and (ii) adversarial loss respectively. We demonstrate state-of-the-art performance in semantic segmentation on two challenging “synthetic-2-real” set-ups and show that the approach can also be used for detection. Himalaya Jain, Maxime Bucher, Matthieu Cord, Patrick Pérez |
CVPR | 5 |
| 2019 | Boosting Few-Shot Visual Learning With Self-SupervisionabstractFew-shot learning and self-supervised learning address different facets of the same problem: how to train a model with little or no labeled data. Few-shot learning aims for optimization methods and models that can learn efficiently to recognize patterns in the low data regime. Self-supervised learning focuses instead on unlabeled data and looks into it for the supervisory signal to feed high capacity deep neural networks. In this work we exploit the complementarity of these two domains and propose an approach for improving few-shot learning through self-supervision. We use self-supervision as an auxiliary task in a few-shot learning pipeline, enabling feature extractors to learn richer and more transferable visual representations while still using few annotated samples. Through self-supervision, our approach can be naturally extended towards using diverse unlabeled data from other datasets in the few-shot setting. We report consistent improvements across an array of architectures, datasets and self-supervision techniques. We provide the implementation code at: https://github.com/valeoai/BF3S. Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, Matthieu Cord |
ICCV | 4 |
| 2019 | DADA: Depth-Aware Domain Adaptation in Semantic SegmentationabstractUnsupervised domain adaptation (UDA) is important for applications where large scale annotation of representative data is challenging. For semantic segmentation in particular, it helps deploy on real “target domain” data models that are trained on annotated images from a different “source domain”, notably a virtual environment. To this end, most previous works consider semantic segmentation as the only mode of supervision for source domain data, while ignoring other, possibly available, information like depth. In this work, we aim at exploiting at best such a privileged information while training the UDA model. We propose a unified depth-aware UDA framework that leverages in several complementary ways the knowledge of dense depth in the source domain. As a result, the performance of the trained semantic segmentation model on the target domain is boosted. Our novel approach indeed achieves state-of-the-art performance on different challenging synthetic-2-real benchmarks. Himalaya Jain, Maxime Bucher, Matthieu Cord, Patrick Pérez |
ICCV | 5 |
| 2019 | WoodScape: A Multi-Task, Multi-Camera Fisheye Dataset for Autonomous DrivingabstractFisheye cameras are commonly employed for obtaining a large field of view in surveillance, augmented reality and in particular automotive applications. In spite of their prevalence, there are few public datasets for detailed evaluation of computer vision algorithms on fisheye images. We release the first extensive fisheye automotive dataset, WoodScape, named after Robert Wood who invented the fisheye camera in 1906. WoodScape comprises of four surround view cameras and nine tasks including segmentation, depth estimation, 3D bounding box detection and soiling detection. Semantic annotation of 40 classes at the instance level is provided for over 10,000 images and annotation for other tasks are provided for over 100,000 images. With WoodScape, we would like to encourage the community to adapt computer vision models for fisheye camera instead of using naive rectification. Senthil Kumar Yogamani, Christian Witt, Hazem Rashed, Sanjaya Nayak, Saquib Mansoor, Padraig Varley, Xavier Perrotton, Derek O'Dea, Patrick Pérez, Ciarán Eising, Jonathan Horgan, Ganesh Sistu, Sumanth Chennupati, Michal Uricár, Stefan Milz, Martin Simon, Karl Amende |
ICCV | 9 |
| 2019 | Photo Style Transfer With Consistency LossesabstractWe address the problem of style transfer between two photos and propose a new way to preserve photorealism. Using the single pair of photos available as input, we train a pair of deep convolution networks (convnets), each of which transfers the style of one photo to the other. To enforce photorealism, we introduce a content preserving mechanism by combining a cycle-consistency loss with a self-consistency loss. Experimental results show that this method does not suffer from typical artifacts observed in methods working in the same settings [1], [2]. We then further analyze some properties of these trained convnets. First, we notice that they can be used to stylize other unseen images with same known style. Second, we show that retraining only a small subset of the network parameters can be sufficient to adapt these convnets to new styles. Gilles Puy, Patrick Pérez |
ICIP | 3 |
| 2019 | Zero-Shot Semantic SegmentationabstractSemantic segmentation models are limited in their ability to scale to large numbers of object classes. In this paper, we introduce the new task of zero-shot semantic segmentation: learning pixel-wise classifiers for never-seen object categories with zero training examples. To this end, we present a novel architecture, ZS3Net, combining a deep visual segmentation model with an approach to generate visual representations from semantic word embeddings. By this way, ZS3Net addresses pixel classification tasks where both seen and unseen categories are faced at test time (so called generalized zero-shot classification). Performance is further improved by a self-training step that relies on automatic pseudo-labeling of pixels from unseen classes. On the two standard segmentation datasets, Pascal-VOC and Pascal-Context, we propose zero-shot benchmarks and set competitive baselines. For complex scenes as ones in the Pascal-Context dataset, we extend our approach by using a graph-context encoding to fully leverage spatial context priors coming from class-wise segmentation maps. Maxime Bucher, Matthieu Cord, Patrick Pérez |
NeurIPS | 4 |
| 2019 | Addressing Failure Prediction by Learning Model ConfidenceabstractAssessing reliably the confidence of a deep neural net and predicting its failures is of primary importance for the practical deployment of these models. In this paper, we propose a new target criterion for model confidence, corresponding to the True Class Probability (TCP). We show how using the TCP is more suited than relying on the classic Maximum Class Probability (MCP). We provide in addition theoretical guarantees for TCP in the context of failure prediction. Since the true class is by essence unknown at test time, we propose to learn TCP criterion on the training set, introducing a specific learning scheme adapted to this context. Extensive experiments are conducted for validating the relevance of the proposed approach. We study various network architectures, small and large scale datasets for image classification and semantic segmentation. We show that our approach consistently outperforms several strong methods, from MCP to Bayesian uncertainty, as well as recent approaches specifically designed for failure prediction. Charles Corbière, Nicolas Thome, Avner Bar-Hen, Matthieu Cord, Patrick Pérez |
NeurIPS | 5 |
| 2018 | Finding Beans in Burgers: Deep Semantic-Visual Embedding With LocalizationabstractSeveral works have proposed to learn a two-path neural network that maps images and texts, respectively, to a same shared Euclidean space where geometry captures useful semantic relationships. Such a multi-modal embedding can be trained and used for various tasks, notably image captioning. In the present work, we introduce a new architecture of this type, with a visual path that leverages recent space-aware pooling mechanisms. Combined with a textual path which is jointly trained from scratch, our semantic-visual embedding offers a versatile model. Once trained under the supervision of captioned images, it yields new state-of-the-art performance on cross-modal retrieval. It also allows the localization of new concepts from the embedding space into any input image, delivering state-of-the-art result on the visual grounding of phrases. Martin Engilberge, Louis Chevallier, Patrick Pérez, Matthieu Cord |
CVPR | 3 |
| 2018 | Learning a Complete Image Indexing PipelineabstractTo work at scale, a complete image indexing system comprises two components: An inverted file index to restrict the actual search to only a subset that should contain most of the items relevant to the query; An approximate distance computation mechanism to rapidly scan these lists. While supervised deep learning has recently enabled improvements to the latter, the former continues to be based on unsupervised clustering in the literature. In this work, we propose a first system that learns both components within a unifying neural framework of structured binary encoding. Himalaya Jain, Joaquin Zepeda, Patrick Pérez, Rémi Gribonval |
CVPR | 3 |
| 2018 | Self-Supervised Multi-Level Face Model Learning for Monocular Reconstruction at Over 250 HzabstractThe reconstruction of dense 3D models of face geometry and appearance from a single image is highly challenging and ill-posed. To constrain the problem, many approaches rely on strong priors, such as parametric face models learned from limited 3D scan data. However, prior models restrict generalization of the true diversity in facial geometry, skin reflectance and illumination. To alleviate this problem, we present the first approach that jointly learns 1) a regressor for face shape, expression, reflectance and illumination on the basis of 2) a concurrently learned parametric face model. Our multi-level face model combines the advantage of 3D Morphable Models for regularization with the out-of-space generalization of a learned corrective space. We train end-to-end on in-the-wild images without dense annotations by fusing a convolutional encoder with a differentiable expert-designed renderer and a self-supervised training loss, both defined at multiple detail levels. Our approach compares favorably to the state-of-the-art in terms of reconstruction quality, better generalizes to real world faces, and runs at over 250 Hz. Ayush Tewari, Michael Zollhöfer, Pablo Garrido 0001, Florian Bernard 0001, Hyeongwoo Kim, Patrick Pérez, Christian Theobalt |
CVPR | 6 |
| 2018 | Audio Style Transferabstract“Style transfer” among images has recently emerged as a very active research topic, fuelled by the power of convolution neural networks (CNNs), and has become fast a very popular technology in social media. This paper investigates the analogous problem in the audio domain: How to transfer the style of a reference audio signal to a target audio content? We propose a flexible framework for the task, which uses a sound texture model to extract statistics characterizing the reference audio style, followed by an optimization-based audio texture synthesis to modify the target content. In contrast to mainstream optimization-based visual transfer method, the proposed process is initialized by the target content instead of random noise and the optimized loss is only about texture, not structure. These differences proved key for audio style transfer in our experiments. In order to extract features of interest, we investigate different architectures, whether pre-trained on other tasks, as done in image style transfer, or engineered based on the human auditory system. Experimental results on different types of audio signal confirm the potential of the proposed approach. Eric Grinstein, Ngoc Q. K. Duong, Alexey Ozerov, Patrick Pérez |
ICASSP | 4 |
| 2018 | Structural inpaintingabstractScene-agnostic visual inpainting remains very challenging despite progress in patch-based methods. Recently, Pathak et al. [26] have introduced convolutional "context encoders'' (CEs) for unsupervised feature learning through image completion tasks. With the additional help of adversarial training, CEs turned out to be a promising tool to complete complex structures in real inpainting problems. In the present paper we propose to push further this key ability by relying on perceptual reconstruction losses at training time. We show on a wide variety of visual scenes the merit of the approach forstructural inpainting, and confirm it through a user study. Combined with the optimization-based refinement of [32] with neural patches, our context encoder opens up new opportunities for prior-free visual inpainting. Huy V. Vo, Ngoc Q. K. Duong, Patrick Pérez |
ACM Multimedia | 3 |
| 2018 | State of the Art on Monocular 3D Face Reconstruction, Tracking, and ApplicationsabstractAbstract The computer graphics and vision communities have dedicated long standing efforts in building computerized tools for reconstructing, tracking, and analyzing human faces based on visual input. Over the past years rapid progress has been made, which led to novel and powerful algorithms that obtain impressive results even in the very challenging case of reconstruction from a single RGB or RGB‐D camera. The range of applications is vast and steadily growing as these technologies are further improving in speed, accuracy, and ease of use. Motivated by this rapid progress, this state‐of‐the‐art report summarizes recent trends in monocular facial performance capture and discusses its applications, which range from performance‐based animation to real‐time facial reenactment. We focus our discussion on methods where the central task is to recover and track a three dimensional model of the human face using optimization‐based reconstruction algorithms. We provide an in‐depth overview of the underlying concepts of real‐world image formation, and we discuss common assumptions and simplifications that make these algorithms practical. In addition, we extensively cover the priors that are used to better constrain the under‐constrained monocular reconstruction problem, and discuss the optimization techniques that are employed to recover dense, photo‐geometric 3D face models from monocular 2D data. Finally, we discuss a variety of use cases for the reviewed algorithms in the context of motion capture, facial animation, as well as image and video editing. Michael Zollhöfer, Justus Thies, Pablo Garrido 0001, Derek Bradley, Thabo Beeler, Patrick Pérez, Marc Stamminger, Matthias Nießner, Christian Theobalt |
Comput. Graph. Forum | 6 |
| 2018 | Deep video portraitsabstractWe present a novel approach that enables photo-realistic re-animation of portrait videos using only an input video. In contrast to existing approaches that are restricted to manipulations of facial expressions only, we are the first to transfer the full 3D head position, head rotation, face expression, eye gaze, and eye blinking from a source actor to a portrait video of a target actor. The core of our approach is a generative neural network with a novel space-time architecture. The network takes as input synthetic renderings of a parametric face model, based on which it predicts photo-realistic video frames for a given target actor. The realism in this rendering-to-video transfer is achieved by careful adversarial training, and as a result, we can create modified target videos that mimic the behavior of the synthetically-created input. In order to enable source-to-target video re-animation, we render a synthetic target video with the reconstructed head animation parameters from a source video, and feed it into the trained network - thus taking full control of the target. With the ability to freely recombine source and target parameters, we are able to demonstrate a large variety of video rewrite applications without explicitly modeling hair, body or background. For instance, we can reenact the full head using interactive user-controlled editing, and realize high-fidelity visual dubbing. To demonstrate the high quality of our output, we conduct an extensive series of experiments and evaluations, where for instance a user study shows that our video edits are hard to detect. Hyeongwoo Kim, Pablo Garrido 0001, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Nießner, Patrick Pérez, Christian Richardt, Michael Zollhöfer, Christian Theobalt |
ACM Trans. Graph. | 7 |
| 2017 | ROAM: A Rich Object Appearance Model with Application to RotoscopingabstractRotoscoping, the detailed delineation of scene elements through a video shot, is a painstaking task of tremendous importance in professional post-production pipelines. While pixel-wise segmentation techniques can help for this task, professional rotoscoping tools rely on parametric curves that offer the artists a much better interactive control on the definition, editing and manipulation of the segments of interest. Sticking to this prevalent rotoscoping paradigm, we propose a novel framework to capture and track the visual aspect of an arbitrary object in a scene, given a first closed outline of this object. This model combines a collection of local foreground/background appearance models spread along the outline, a global appearance model of the enclosed object and a set of distinctive foreground landmarks. The structure of this rich appearance model allows simple initialization, efficient iterative optimization with exact minimization at each step, and on-line adaptation in videos. We demonstrate qualitatively and quantitatively the merit of this framework through comparisons with tools based on either dynamic segmentation with a closed curve or pixel-wise binary labelling. Ondrej Miksik, Juan-Manuel Pérez-Rúa, Philip Torr 0001, Patrick Pérez |
CVPR | 4 |
| 2017 | Kernel Square-Loss Exemplar Machines for Image RetrievalabstractZepeda and Perez [41] have recently demonstrated the promise of the exemplar SVM (ESVM) as a feature encoder for image retrieval. This paper extends this approach in several directions: We first show that replacing the hinge loss by the square loss in the ESVM cost function significantly reduces encoding time with negligible effect on accuracy. We call this model square-loss exemplar machine, or SLEM. We then introduce a kernelized SLEM which can be implemented efficiently through low-rank matrix decomposition, and displays improved performance. Both SLEM variants exploit the fact that the negative examples are fixed, so most of the SLEM computational complexity is relegated to an offline process independent of the positive examples. Our experiments establish the performance and computational advantages of our approach using a large array of base features and standard image retrieval datasets. Rafael S. Rezende, Joaquin Zepeda, Jean Ponce, Francis R. Bach, Patrick Pérez |
CVPR | 5 |
| 2017 | Motion informed audio source separationabstractIn this paper we tackle the problem of single channel audio source separation driven by descriptors of the sounding object's motion. As opposed to previous approaches, motion is included as a soft-coupling constraint within the nonnegative matrix factorization framework. The proposed method is applied to a multimodal dataset of instruments in string quartet performance recordings where bow motion information is used for separation of string instruments. We show that the approach offers better source separation result than an audio-based baseline and the state-of-the-art multimodal-based approaches on these very challenging music mixtures. Sanjeel Parekh, Slim Essid, Alexey Ozerov, Ngoc Q. K. Duong, Patrick Pérez, Gaël Richard |
ICASSP | 5 |
| 2017 | Informed source separation via compressive graph signal samplingabstractWe propose a novel informed source separation method for audio object coding based on a recent sampling theory for smooth signals on graphs. Assuming that only one source is active at each time-frequency point, we compute an ideal map indicating which source is active at each time-frequency point at the encoder. This map is then sampled with a compressive graph signal sampling strategy that guarantees accurate and stable recovery at the decoder. The graph is built using feature vectors, computed using non-negative matrix factorization, that allows us to connect similar source activations in the time-frequency plane. We show that the proposed approach performs better than state-of-the-art methods at low bitrate. Gilles Puy, Alexey Ozerov, Ngoc Q. K. Duong, Patrick Pérez |
ICASSP | 4 |
| 2017 | SuBiC: A Supervised, Structured Binary Code for Image SearchabstractFor large-scale visual search, highly compressed yet meaningful representations of images are essential. Structured vector quantizers based on product quantization and its variants are usually employed to achieve such compression while minimizing the loss of accuracy. Yet, unlike binary hashing schemes, these unsupervised methods have not yet benefited from the supervision, end-to-end learning and novel architectures ushered in by the deep learning revolution. We hence propose herein a novel method to make deep convolutional neural networks produce supervised, compact, structured binary codes for visual search. Our method makes use of a novel block-softmax nonlinearity and of batch-based entropy losses that together induce structure in the learned encodings. We show that our method outperforms state-of-the-art compact representations based on deep hashing or structured quantization in single and cross-domain category retrieval, instance retrieval and classification. We make our code and models publicly available online. Himalaya Jain, Joaquin Zepeda, Patrick Pérez, Rémi Gribonval |
ICCV | 3 |
| 2017 | MoFA: Model-Based Deep Convolutional Face Autoencoder for Unsupervised Monocular ReconstructionabstractIn this work we propose a novel model-based deep convolutional autoencoder that addresses the highly challenging problem of reconstructing a 3D human face from a single in-the-wild color image. To this end, we combine a convolutional encoder network with an expert-designed generative model that serves as decoder. The core innovation is the differentiable parametric decoder that encapsulates image formation analytically based on a generative model. Our decoder takes as input a code vector with exactly defined semantic meaning that encodes detailed face pose, shape, expression, skin reflectance and scene illumination. Due to this new way of combining CNN-based with model-based face reconstruction, the CNN-based encoder learns to extract semantically meaningful parameters from a single monocular input image. For the first time, a CNN encoder and an expert-designed generative model can be trained end-to-end in an unsupervised manner, which renders training on very large (unlabeled) real world data feasible. The obtained reconstructions compare favorably to current state-of-the-art approaches in terms of quality and richness of representation. Ayush Tewari, Michael Zollhöfer, Hyeongwoo Kim, Pablo Garrido 0001, Florian Bernard 0001, Patrick Pérez, Christian Theobalt |
ICCV | 6 |
| 2016 | Cotemporal Multi-View Video SegmentationabstractWe address the problem of multi-view video segmentation of dynamic scenes in general and outdoor environments with possibly moving cameras. Multi-view methods for dynamic scenes usually rely on geometric calibration to impose spatial shape constraints between viewpoints. In this paper, we show that the calibration constraint can be relaxed while still getting competitive segmentation results using multi-view constraints. We introduce new multi-view cotemporality constraints through motion correlation cues, in addition to common appearance features used by co-segmentation methods to identify co-instances of objects. We also take advantage of learning based segmentation strategies by casting the problem as the selection of monocular proposals that satisfy multi-view constraints. This yields a fully automated method that can segment subjects of interest without any particular pre-processing stage. Results on several challenging outdoor datasets demonstrate the feasibility and robustness of our approach. Abdelaziz Djelouah, Jean-Sébastien Franco, Edmond Boyer, Patrick Pérez, George Drettakis |
3DV | 4 |
| 2016 | Maximum Margin Linear Classifiers in Unions of Subspaces
Xinrui Lyu, Joaquin Zepeda, Patrick Pérez |
BMVC | 3 |
| 2016 | Discovering motion hierarchies via tree-structured coding of trajectories
Juan-Manuel Pérez-Rúa, Tomás Crivelli, Patrick Pérez, Patrick Bouthemy |
BMVC | 3 |
| 2016 | Determining Occlusions from Space and Time Image ReconstructionsabstractThe problem of localizing occlusions between consecutive frames of a video is important but rarely tackled on its own. In most works, it is tightly interleaved with the computation of accurate optical flows, which leads to a delicate chicken-and-egg problem. With this in mind, we propose a novel approach to occlusion detection where visibility or not of a point in next frame is formulated in terms of visual reconstruction. The key issue is now to determine how well a pixel in the first image can be "reconstructed" from co-located colors in the next image. We first exploit this reasoning at the pixel level with a new detection criterion. Contrary to the ubiquitous displaced-framedifference and forward-backward flow vector matching, the proposed alternative does not critically depend on a precomputed, dense displacement field, while being shown to be more effective. We then leverage this local modeling within an energy-minimization framework that delivers occlusion maps. An easy-to-obtain collection of parametric motion models is exploited within the energy to provide the required level of motion information. Our approach outperforms state-of-the-art detection methods on the challenging MPI Sintel dataset. Juan-Manuel Pérez-Rúa, Tomás Crivelli, Patrick Bouthemy, Patrick Pérez |
CVPR | 4 |
| 2016 | Approximate Search with Quantized Sparse Representations
Himalaya Jain, Patrick Pérez, Rémi Gribonval, Joaquin Zepeda, Hervé Jégou |
ECCV (7) | 2 |
| 2016 | SPLeaP: Soft Pooling of Learned Parts for Image Classification
Praveen Kulkarni 0003, Frédéric Jurie, Joaquin Zepeda, Patrick Pérez, Louis Chevallier |
ECCV (8) | 4 |
| 2016 | Automatic allocation of NTF components for user-guided audio source separationabstractNonnegative matrix or tensor factorization is a very popular approach for audio source separation. One important problem in nonnegative tensor factorization (NTF) in the context of user-guided audio source separation is the necessity to manually assign the NTF components to audio sources in order to be able to enforce prior information on the sources during the estimation process. In this paper, two new approaches to NTF based source separation are proposed, which do not require any manual component assignment to the sources, but estimate the underlying assignment automatically. Both algorithms use the prior information on the source samples in the estimation process along with either a limit on the minimum number of components each source uses or with a restriction that each component is used by sparse number of sources. The proposed methods are shown to outperform the classic approach with a manual distribution of the components equally among the sources. Cagdas Bilen, Alexey Ozerov, Patrick Pérez |
ICASSP | 3 |
| 2016 | Sketching for large-scale learning of mixture modelsabstractLearning parameters from voluminous data can be prohibitive in terms of memory and computational requirements. We propose a "compressive learning" framework where we first sketch the data by computing random generalized moments of the underlying probability distribution, then estimate mixture model parameters from the sketch using an iterative algorithm analogous to greedy sparse signal recovery. We exemplify our framework with the sketched estimation of Gaussian Mixture Models (GMMs). We experimentally show that our approach yields results comparable to the classical Expectation-Maximization (EM) technique while requiring significantly less memory and fewer computations when the number of database elements is large. We report large-scale experiments in speaker verification, where our approach makes it possible to fully exploit a corpus of 1000 hours of speech signal to learn a universal background model at scales computationally inaccessible to EM. Nicolas Keriven, Anthony Bourrier, Rémi Gribonval, Patrick Pérez |
ICASSP | 4 |
| 2016 | Multichannel audio declippingabstractAudio declipping consists in recovering so-called clipped audio samples that are set to a maximum / minimum threshold. Many different approaches were proposed to solve this problem in case of singlechannel (mono) recordings. However, while most of audio recordings are multichannel nowadays, there is no method designed specifically for multichannel audio declipping, where the inter-channel correlations may be efficiently exploited for a better declipping result. In this work we propose for the first time such a multichannel audio declipping method. Our method is based on representing a multichannel audio recording as a convolutive mixture of several audio sources, and on modeling the source power spectrograms and mixing filters by nonnegative tensor factorization model and full-rank covariance matrices, respectively. A generalized expectation-maximization algorithm is proposed to estimate model parameters. It is shown experimentally that the proposed multichannel audio de-clipping algorithm outperforms in average and in most cases a state-of-the-art single-channel declipping algorithm applied to each channel independently. Alexey Ozerov, Cagdas Bilen, Patrick Pérez |
ICASSP | 3 |
| 2016 | Supervised learning of low-rank transforms for image retrievalabstractIn this paper we propose a new method to automatically select the rank of linear transforms during supervised learning. Our approach relies on a sparsity-enforcing element-wise soft-thresholding operation applied after the linear transform. This novel approach to supervised rank learning has the important advantage that it is very simple to implement and incurs no extra complexity relative to linear transform learning. Furthermore, we propose a simple Stochastic Gradient Descent (SGD) implementation suitable for large scale learning, where SGD solvers have established themselves as the default workhorse. We compare our method to various other metric learning techniques in the application of image retrieval. This is one of the remaining few areas where supervised learning of low-rank linear transforms has not been fully exploited. The main reason for this is the lack of adequate datasets that are large enough, and hence we further introduce a new dataset consisting of groups of matching images derived from Cable News Network (CNN) videos using geometric verification and manual selection to find matching frames with adequate variability. Cagdas Bilen, Joaquin Zepeda, Patrick Pérez |
ICIP | 3 |
| 2016 | Hierarchical motion decomposition for dynamic scene parsingabstractA number of applications in video analysis rely on a per-frame motion segmentation of the scene as key preprocessing step. Moreover, different settings in video production require extracting segmentation masks of multiple moving objects and object parts in a hierarchical fashion. In order to tackle this problem, we propose to analyze and exploit the compositional structure of scene motion to provide a segmentation which is not purely driven by local image information. Specifically, we leverage a hierarchical motion-based partition of the scene to capture a mid-level understanding of the dynamic video content. We present experimental results showing the strengths of this approach in comparison to current video segmentation approaches. Juan-Manuel Pérez-Rúa, Tomás Crivelli, Patrick Pérez, Patrick Bouthemy |
ICIP | 3 |
| 2016 | Visual object trapping
Tomás Crivelli, Patrick Pérez, Lionel Oisel |
Comput. Vis. Image Underst. | 2 |
| 2016 | Object-guided motion estimation
Juan-Manuel Pérez-Rúa, Tomás Crivelli, Patrick Pérez |
Comput. Vis. Image Underst. | 3 |
| 2016 | Reconstruction of Personalized 3D Face Rigs from Monocular VideoabstractWe present a novel approach for the automatic creation of a personalized high-quality 3D face rig of an actor from just monocular video data (e.g., vintage movies). Our rig is based on three distinct layers that allow us to model the actor’s facial shape as well as capture his person-specific expression characteristics at high fidelity, ranging from coarse-scale geometry to fine-scale static and transient detail on the scale of folds and wrinkles. At the heart of our approach is a parametric shape prior that encodes the plausible subspace of facial identity and expression variations. Based on this prior, a coarse-scale reconstruction is obtained by means of a novel variational fitting approach. We represent person-specific idiosyncrasies, which cannot be represented in the restricted shape and expression space, by learning a set of medium-scale corrective shapes. Fine-scale skin detail, such as wrinkles, are captured from video via shading-based refinement, and a generative detail formation model is learned. Both the medium- and fine-scale detail layers are coupled with the parametric prior by means of a novel sparse linear regression formulation. Once reconstructed, all layers of the face rig can be conveniently controlled by a low number of blendshape expression parameters, as widely used by animation artists. We show captured face rigs and their motions for several actors filmed in different monocular video formats, including legacy footage from YouTube, and demonstrate how they can be used for 3D animation and 2D video editing. Finally, we evaluate our approach qualitatively and quantitatively and compare to related state-of-the-art methods. Pablo Garrido 0001, Michael Zollhöfer, Dan Casas, Levi Valgaerts, Kiran Varanasi, Patrick Pérez, Christian Theobalt |
ACM Trans. Graph. | 6 |
| 2016 | Corrective 3D reconstruction of lips from monocular videoabstractIn facial animation, the accurate shape and motion of the lips of virtual humans is of paramount importance, since subtle nuances in mouth expression strongly influence the interpretation of speech and the conveyed emotion. Unfortunately, passive photometric reconstruction of expressive lip motions, such as a kiss or rolling lips, is fundamentally hard even with multi-view methods in controlled studios. To alleviate this problem, we present a novel approach for fully automatic reconstruction of detailed and expressive lip shapes along with the dense geometry of the entire face, from just monocular RGB video. To this end, we learn the difference between inaccurate lip shapes found by a state-of-the-art monocular facial performance capture approach, and the true 3D lip shapes reconstructed using a high-quality multi-view system in combination with applied lip tattoos that are easy to track. A robust gradient domain regressor is trained to infer accurate lip shapes from coarse monocular reconstructions, with the additional help of automatically extracted inner and outer 2D lip contours. We quantitatively and qualitatively show that our monocular approach reconstructs higher quality lip shapes, even for complex shapes like a kiss or lip rolling, than previous monocular approaches. Furthermore, we compare the performance of person-specific and multi-person generic regression strategies and show that our approach generalizes to new individuals and general scenes, enabling high-fidelity reconstruction even from commodity video footage. Pablo Garrido 0001, Michael Zollhöfer, Chenglei Wu, Derek Bradley, Patrick Pérez, Thabo Beeler, Christian Theobalt |
ACM Trans. Graph. | 5 |
| 2015 | Learning the Structure of Deep Architectures Using L1 RegularizationabstractInternational audience Praveen Kulkarni 0003, Joaquin Zepeda, Frédéric Jurie, Patrick Pérez, Louis Chevallier |
BMVC | 4 |
| 2015 | The Semantic Paintbrush: Interactive 3D Mapping and Recognition in Large Outdoor SpacesabstractWe present an augmented reality system for large scale 3D reconstruction and recognition in outdoor scenes. Unlike existing prior work, which tries to reconstruct scenes using active depth cameras, we use a purely passive stereo setup, allowing for outdoor use and extended sensing range. Our system not only produces a map of the 3D environment in real-time, it also allows the user to draw (or 'paint') with a laser pointer directly onto the reconstruction to segment the model into objects. Given these examples our system then learns to segment other parts of the 3D map during online acquisition. Unlike typical object recognition systems, ours therefore very much places the user 'in the loop' to segment particular objects of interest, rather than learning from predefined databases. The laser pointer additionally helps to 'clean up' the stereo reconstruction and final 3D map, interactively. Using our system, within minutes, a user can capture a full 3D map, segment it into objects of interest, and refine parts of the model during capture. We provide full technical details of our system to aid replication, as well as quantitative evaluation of system components. We demonstrate the possibility of using our system for helping the visually impaired navigate through spaces. Beyond this use, our system can be used for playing large-scale augmented reality games, shared online to augment streetview data, and used for more detailed car and person navigation. Ondrej Miksik, Vibhav Vineet, Morten Lidegaard, Ram Prasaath, Matthias Nießner, Stuart Golodetz, Stephen L. Hicks, Patrick Pérez, Shahram Izadi, Philip Torr 0001 |
CHI | 8 |
| 2015 | Exemplar SVMs as visual feature encodersabstractIn this work, we investigate the use of exemplar SVMs (linear SVMs trained with one positive example only and a vast collection of negative examples) as encoders that turn generic image features into new, task-tailored features. The proposed feature encoding leverages the ability of the exemplar-SVM (E-SVM) classifier to extract, from the original representation of the exemplar image, what is unique about it. While existing image description pipelines rely on the intuition of the designer to encode uniqueness into the feature encoding process, our proposed approach does it explicitly relative to a “universe” of features represented by the generic negatives. We show that such a post-processing enhances the performance of state-of-the art image retrieval methods based on aggregated image features, as well as the performance of nearest class mean and K-nearest neighbor image classification methods. We establish these advantages for several features, including “traditional” features as well as features derived from deep convolutional neural nets. As an additional contribution, we also propose a recursive extension of this E-SVM encoding scheme (RE-SVM) that provides further performance gains. Joaquin Zepeda, Patrick Pérez |
CVPR | 2 |
| 2015 | Hybrid multi-layer deep CNN/aggregator feature for image classificationabstractDeep Convolutional Neural Networks (DCNN) have established a remarkable performance benchmark in the field of image classification, displacing classical approaches based on hand-tailored aggregations of local descriptors. Yet DCNNs impose high computational burdens both at training and at testing time, and training them requires collecting and annotating large amounts of training data. Supervised adaptation methods have been proposed in the literature that partially re-learn a transferred DCNN structure from a new target dataset. Yet these require expensive bounding-box annotations and are still computationally expensive to learn. In this paper, we address these shortcomings of DCNN adaptation schemes by proposing a hybrid approach that combines conventional, unsupervised aggregators such as Bag-of-Words (BoW), with the DCNN pipeline by treating the output of intermediate layers as densely extracted local descriptors. We test a variant of our approach that uses only intermediate DCNN layers on the standard PASCAL VOC 2007 dataset and show performance significantly higher than the standard BoW model and comparable to Fisher vector aggregation but with a feature that is 150 times smaller. A second variant of our approach that includes the fully connected DCNN layers significantly outperforms Fisher vector schemes and performs comparably to DCNN approaches adapted to Pascal VOC 2007, yet at only a small fraction of the training and testing cost. Praveen Kulkarni 0003, Joaquin Zepeda, Frédéric Jurie, Patrick Pérez, Louis Chevallier |
ICASSP | 4 |
| 2015 | Background-foreground tracking for video object segmentationabstractWe present a method to segment objects of interest in video sequences by combining robust background and foreground point tracking with joint color and motion-based segmentation. Our approach is sequential in time, avoiding a global processing of the video, while being simple and generic. This makes the method attractive for online applications, including video editing or augmented reality, as it can be adapted for both automated and interactive work-flows. We present visual and quantitative experiments to compare with existing algorithms, showing promising results. Juan-Manuel Pérez-Rúa, Tomás Crivelli, Patrick Pérez |
ICIP | 3 |
| 2015 | Incremental dense semantic stereo fusion for large-scale semantic scene reconstructionabstractOur abilities in scene understanding, which allow us to perceive the 3D structure of our surroundings and intuitively recognise the objects we see, are things that we largely take for granted, but for robots, the task of understanding large scenes quickly remains extremely challenging. Recently, scene understanding approaches based on 3D reconstruction and semantic segmentation have become popular, but existing methods either do not scale, fail outdoors, provide only sparse reconstructions or are rather slow. In this paper, we build on a recent hash-based technique for large-scale fusion and an efficient mean-field inference algorithm for densely-connected CRFs to present what to our knowledge is the first system that can perform dense, large-scale, outdoor semantic reconstruction of a scene in (near) real time. We also present a `semantic fusion' approach that allows us to handle dynamic objects more effectively than previous approaches. We demonstrate the effectiveness of our approach on the KITTI dataset, and provide qualitative and quantitative results showing high-quality dense reconstruction and labelling of a number of scenes. Vibhav Vineet, Ondrej Miksik, Morten Lidegaard, Matthias Nießner, Stuart Golodetz, Victor Adrian Prisacariu, Olaf Kähler, David William Murray 0001, Shahram Izadi, Patrick Pérez, Philip Torr 0001 |
ICRA | 10 |
| 2015 | Incremental dense multi-modal 3D scene reconstructionabstractAquiring reliable depth maps is an essential prerequisite for accurate and incremental 3D reconstruction used in a variety of robotics applications. Depth maps produced by affordable Kinect-like cameras have become a de-facto standard for indoor reconstruction and the driving force behind the success of many algorithms. However, Kinect-like cameras are less effective outdoors where one should rely on other sensors. Often, we use a combination of a stereo camera and lidar, however, process the acquired data in independent pipelines which generally leads to sub-optimal performance since both sensors suffer from different drawbacks. In this paper, we propose a probabilistic model that efficiently exploits complementarity between different depth-sensing modalities for incremental dense scene reconstruction. Our model uses a piecewise planarity prior assumption which is common in both the indoor and outdoor scenes. We demonstrate the effectiveness of our approach on the KITTI dataset, and provide qualitative and quantitative results showing high-quality dense reconstruction of a number of scenes. Ondrej Miksik, Yousef Amar, Vibhav Vineet, Patrick Pérez, Philip Torr 0001 |
IROS | 4 |
| 2015 | VDub: Modifying Face Video of Actors for Plausible Visual Alignment to a Dubbed Audio TrackabstractAbstract In many countries, foreign movies and TV productions are dubbed, i.e., the original voice of an actor is replaced with a translation that is spoken by a dubbing actor in the country's own language. Dubbing is a complex process that requires specific translations and accurately timed recitations such that the new audio at least coarsely adheres to the mouth motion in the video. However, since the sequence of phonemes and visemes in the original and the dubbing language are different, the video‐to‐audio match is never perfect, which is a major source of visual discomfort. In this paper, we propose a system to alter the mouth motion of an actor in a video, so that it matches the new audio track. Our paper builds on high‐quality monocular capture of 3D facial performance, lighting and albedo of the dubbing and target actors, and uses audio analysis in combination with a space‐time retrieval method to synthesize a new photo‐realistically rendered and highly detailed 3D shape model of the mouth region to replace the target performance. We demonstrate plausible visual quality of our results compared to footage that has been professionally dubbed in the traditional way, both qualitatively and through a user study. Pablo Garrido 0001, Levi Valgaerts, H. Sarmadi, Ingmar Steiner, Kiran Varanasi, Patrick Pérez, Christian Theobalt |
Comput. Graph. Forum | 6 |
| 2015 | Sparse Multi-View Consistency for Object SegmentationabstractMultiple view segmentation consists in segmenting objects simultaneously in several views. A key issue in that respect and compared to monocular settings is to ensure propagation of segmentation information between views while minimizing complexity and computational cost. In this work, we first investigate the idea that examining measurements at the projections of a sparse set of 3D points is sufficient to achieve this goal. The proposed algorithm softly assigns each of these 3D samples to the scene background if it projects on the background region in at least one view, or to the foreground if it projects on foreground region in all views. Second, we show how other modalities such as depth may be seamlessly integrated in the model and benefit the segmentation. The paper exposes a detailed set of experiments used to validate the algorithm, showing results comparable with the state of art, with reduced computational complexity. We also discuss the use of different modalities for specific situations, such as dealing with a low number of viewpoints or a scene with color ambiguities between foreground and background. Abdelaziz Djelouah, Jean-Sébastien Franco, Edmond Boyer, François Le Clerc, Patrick Pérez |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2015 | Robust Optical Flow IntegrationabstractWe analyze the problem of how to correctly construct dense point trajectories from optical flow fields. First, we show that simple Euler integration is unavoidably inaccurate, no matter how good is the optical flow estimator. Then, an inverse integration scheme is analyzed which is more robust to bias and input noise and shows better stability properties. Our contribution is threefold: 1) a theoretical analysis that demonstrates why and in what sense inverse integration is more accurate; 2) a rich experimental validation both on synthetic and real (image) data; and 3) an algorithm for approximate online inverse integration. This new technique is precious whether one is trying to propagate information densely available on a reference frame to the other frames in the sequence or, conversely, to assign information densely over each frame by pulling it from the reference. Tomás Crivelli, Matthieu Fradet, Pierre-Henri Conze, Philippe Robert, Patrick Pérez |
IEEE Trans. Image Process. | 5 |
| 2014 | EPML: Expanded Parts Based Metric Learning for Occlusion Robust Face Verification
Gaurav Sharma 0004, Frédéric Jurie, Patrick Pérez |
ACCV (4) | 3 |
| 2014 | Distributed Non-convex ADMM-based inference in large-scale random fields
Ondrej Miksik, Vibhav Vineet, Patrick Pérez, Philip Torr 0001 |
BMVC | 3 |
| 2014 | Automatic Face ReenactmentabstractWe propose an image-based, facial reenactment system that replaces the face of an actor in an existing target video with the face of a user from a source video, while preserving the original target performance. Our system is fully automatic and does not require a database of source expressions. Instead, it is able to produce convincing reenactment results from a short source video captured with an off-the-shelf camera, such as a webcam, where the user performs arbitrary facial gestures. Our reenactment pipeline is conceived as part image retrieval and part face transfer: The image retrieval is based on temporal clustering of target frames and a novel image matching metric that combines appearance and motion to select candidate frames from the source video, while the face transfer uses a 2D warping strategy that preserves the user's identity. Our system excels in simplicity as it does not rely on a 3D face model, it is robust under head motion and does not require the source and target performance to be similar. We show convincing reenactment results for videos that we recorded ourselves and for low-quality footage taken from the Internet. Pablo Garrido 0001, Levi Valgaerts, Ole Rehmsen, Thorsten Thormählen, Patrick Pérez, Christian Theobalt |
CVPR | 5 |
| 2014 | Disparity-guided demosaicking of light field imagesabstractLight-field imaging has been recently introduced to mass market by the hand held plenoptic camera Lytro. Thanks to a microlens array placed between the main lens and the sensor, the captured data contains different views of the scene from different view points. This offers several post-capture applications, e.g., computationally changing the main lens focus. The raw data conversion in such cameras is however barely studied in the literature. The goal of this paper is to study the particularly overlooked problem of demosaicking the views for plenoptic cameras such as Lytro. We exploit the redundant sampling of scene content in the views, and show that disparities estimated from the mosaicked data can guide the demosaicking, resulting in minimum artifacts compared to the state of art methods. Besides, by properly addressing the view demultiplexing step, we take the first step towards light field super-resolution with negligible computational overload. Mozhdeh Seifi, Neus Sabater, Valter Drazic, Patrick Pérez |
ICIP | 4 |
| 2014 | Video Inpainting of Complex ScenesabstractWe propose an automatic video inpainting algorithm which relies on the optimization of a global, patch-based functional. Our algorithm is able to deal with a variety of challenging situations which naturally arise in video inpainting, such as the correct reconstruction of dynamic textures, multiple moving objects, and moving background. Furthermore, we achieve this in an order of magnitude less execution time with respect to the state-of-the-art. We are also able to achieve good quality results on high-definition videos. Finally, we provide specific algorithmic details to make implementation of our algorithm as easy as possible. The resulting algorithm requires no segmentation or manual input other than the definition of the inpainting mask and can deal with a wider variety of situations than is handled by previous work. Alasdair Newson, Andrés Almansa, Matthieu Fradet, Yann Gousseau, Patrick Pérez |
SIAM J. Imaging Sci. | 5 |
| 2014 | Robust Automatic Line Scratch Detection in FilmsabstractLine scratch detection in old films is a particularly challenging problem due to the variable spatiotemporal characteristics of this defect. Some of the main problems include sensitivity to noise and texture, and false detections due to thin vertical structures belonging to the scene. We propose a robust and automatic algorithm for frame-by-frame line scratch detection in old films, as well as a temporal algorithm for the filtering of false detections. In the frame-by-frame algorithm, we relax some of the hypotheses used in previous algorithms in order to detect a wider variety of scratches. This step's robustness and lack of external parameters is ensured by the combined use of an a contrario methodology and local statistical estimation. In this manner, over-detection in textured or cluttered areas is greatly reduced. The temporal filtering algorithm eliminates false detections due to thin vertical structures by exploiting the coherence of their motion with that of the underlying scene. Experiments demonstrate the ability of the resulting detection procedure to deal with difficult situations, in particular in the presence of noise, texture, and slanted or partial scratches. Comparisons show significant advantages over previous work. Alasdair Newson, Andrés Almansa, Yann Gousseau, Patrick Pérez |
IEEE Trans. Image Process. | 4 |
| 2014 | Fundamental Performance Limits for Ideal Decoders in High-Dimensional Linear Inverse ProblemsabstractThe primary challenge in linear inverse problems is to design stable and robust decoders to reconstruct high-dimensional vectors from a low-dimensional observation through a linear operator. Sparsity, low-rank, and related assumptions are typically exploited to design decoders, whose performance is then bounded based on some measure of deviation from the idealized model, typically using a norm. This paper focuses on characterizing the fundamental performance limits that can be expected from an ideal decoder given a general model, i.e., a general subset of simple vectors of interest. First, we extend the so-called notion of instance optimality of a decoder to settings where one only wishes to reconstruct some part of the original high-dimensional vector from a low-dimensional observation. This covers practical settings, such as medical imaging of a region of interest, or audio source separation, when one is only interested in estimating the contribution of a specific instrument to a musical recording. We define instance optimality relatively to a model much beyond the traditional framework of sparse recovery, and characterize the existence of an instance optimal decoder in terms of joint properties of the model and the considered linear operator. Noiseless and noise-robust settings are both considered. We show somewhat surprisingly that the existence of noise-aware instance optimal decoders for all noise levels implies the existence of a noise-blind decoder. A consequence of our results is that for models that are rich enough to contain an orthonormal basis, the existence of an ℓ2/ℓ2instance optimal decoder is only possible when the linear operator is not substantially dimension-reducing. This covers well-known cases (sparse vectors, low-rank matrices) as well as a number of seemingly new situations (structured sparsity and sparse inverse covariance matrices for instance). We exhibit an operator-dependent norm which, under a model-specific generalization of the restricted isometry property, always yields a feasible instance optimality property. This norm can be upper bounded by an atomic norm relative to the considered model. Anthony Bourrier, Mike E. Davies 0001, Tomer Peleg, Patrick Pérez, Rémi Gribonval |
IEEE Trans. Inf. Theory | 4 |
| 2013 | Compressive Gaussian Mixture estimationabstractWhen fitting a probability model to voluminous data, memory and computational time can become prohibitive. In this paper, we propose a framework aimed at fitting a mixture of isotropic Gaussians to data vectors by computing a low-dimensional sketch of the data. The sketch represents empirical moments of the underlying probability distribution. Deriving a reconstruction algorithm by analogy with compressive sensing, we experimentally show that it is possible to precisely estimate the mixture parameters provided that the sketch is large enough. Our algorithm provides good reconstruction and scales to higher dimensions than previous probability mixture estimation algorithms, while consuming less memory in the case of numerous data. It also provides a privacy-preserving data analysis tool, since the sketch doesn't disclose information about individual datum it is based on. Anthony Bourrier, Rémi Gribonval, Patrick Pérez |
ICASSP | 3 |
| 2013 | Multi-view Object Segmentation in Space and TimeabstractIn this paper, we address the problem of object segmentation in multiple views or videos when two or more viewpoints of the same scene are available. We propose a new approach that propagates segmentation coherence information in both space and time, hence allowing evidences in one image to be shared over the complete set. To this aim the segmentation is cast as a single efficient labeling problem over space and time with graph cuts. In contrast to most existing multi-view segmentation methods that rely on some form of dense reconstruction, ours only requires a sparse 3D sampling to propagate information between viewpoints. The approach is thoroughly evaluated on standard multi-view datasets, as well as on videos. With static views, results compete with state of the art methods but they are achieved with significantly fewer viewpoints. With multiple videos, we report results that demonstrate the benefit of segmentation propagation through temporal cues. Abdelaziz Djelouah, Jean-Sébastien Franco, Edmond Boyer, François Le Clerc, Patrick Pérez |
ICCV | 5 |
| 2013 | Temporal filtering of line scratch detections in degraded filmsabstractThe film defect known as the line scratch is difficult to restore automatically due to the large number of false alarms present in scratch detection algorithms. In this paper, an algorithm for dealing with these false alarms is proposed. Validating true scratches, which is the approach generally proposed in the literature, is a difficult task since scratch characteristics are hard to determine, making tracking these defects problematic. Instead, we eliminate false alarms by analysing their compatibility with a global motion estimation. We compare our algorithm with two other scratch detection methods from the literature. Experiments show that our algorithm outperforms these two, and that the proposed temporal filtering greatly improves precision while maintaining high recall. Alasdair Newson, Andrés Almansa, Yann Gousseau, Patrick Pérez |
ICIP | 4 |
| 2013 | On evaluating face tracks in moviesabstractAutomatic extraction of face tracks is a key component of systems that analyse people in audio-visual content such as TV programs and movies. Due to the lack of properly annotated content of this type, popular algorithms for extracting face tracks have not been fully assessed in the literature. We introduce and make publicly available a new dataset, based on the full annotation of a feature movie, to help fill this gap. We show in particular that, thanks to this dataset, state-of-art tracking metrics can now be exploited to evaluate face tracks used by, e.g., automatic character naming systems. We conduct such an evaluation on different variants of a novel system that we introduce as a generalization of existing ones. Alexey Ozerov, Jean-Ronan Vigouroux, Louis Chevallier, Patrick Pérez |
ICIP | 4 |
| 2013 | Revisiting the VLAD image representationabstractRecent works on image retrieval have proposed to index images by compact representations encoding powerful local descriptors, such as the closely related VLAD and Fisher vector. By combining such a representation with a suitable coding technique, it is possible to encode an image in a few dozen bytes while achieving excellent retrieval results. This paper revisits some assumptions proposed in this context regarding the handling of "visual burstiness", and shows that ad-hoc choices are implicitly done which are not desirable. Focusing on VLAD without loss of generality, we propose to modify several steps of the original design. Albeit simple, these modifications significantly improve VLAD and make it compare favorably against the state of the art. Jonathan Delhumeau, Philippe Henri Gosselin, Hervé Jégou, Patrick Pérez |
ACM Multimedia | 4 |
| 2013 | Correspondence Map-Aided Neighbor Embedding for Image Intra PredictionabstractThis paper describes new image prediction methods based on neighbor embedding (NE) techniques. Neighbor embedding methods are used here to approximate an input block (the block to be predicted) in the image as a linear combination of K nearest neighbors. However, in order for the decoder to proceed similarly, the K nearest neighbors are found by computing distances between the known pixels in a causal neighborhood (called template) of the input block and the co-located pixels in candidate patches taken from a causal window. Similarly, the weights used for the linear approximation are computed in order to best approximate the template pixels. Although efficient, these methods suffer from limitations when the template and the block to be predicted are not correlated, e.g., in non homogenous texture areas. To cope with these limitations, this paper introduces new image prediction methods based on NE techniques in which the K-NN search is done in two steps and aided, at the decoder, by a block correspondence map, hence the name map-aided neighbor embedding (MANE) method. Another optimized variant of this approach, called oMANE method, is also studied. In these methods, several alternatives have also been proposed for the K-NN search. The resulting prediction methods are shown to bring significant rate-distortion performance improvements when compared to H.264 Intra prediction modes (up to 44.75% rate saving at low bit rates). Safa Chérigui, Christine Guillemot, Dominique Thoreau, Philippe Guillotel, Patrick Pérez |
IEEE Trans. Image Process. | 5 |
| 2012 | Multi-step flow fusion: towards accurate and dense correspondences in long video shotsabstractInternational audience Tomás Crivelli, Pierre-Henri Conze, Philippe Robert, Matthieu Fradet, Patrick Pérez |
BMVC | 5 |
| 2012 | N-tuple Color Segmentation for Multi-view Silhouette Extraction
Abdelaziz Djelouah, Jean-Sébastien Franco, Edmond Boyer, François Le Clerc, Patrick Pérez |
ECCV (5) | 5 |
| 2012 | Hybrid template and block matching algorithm for image intra predictionabstractTemplate matching has been shown to outperform the H.264 prediction modes for Intra video coding thanks to better spatial prediction and no additional ancillary data to transmit. The method indeed works well when the template and the block to be predicted are highly correlated, e.g., in homogenous image areas, however, it obviously fails in areas with non homogeneous textures. This paper explores the idea of using a block-matching intra prediction algorithm which, thanks to a Rate-Distorsion (RD) based decision mechanism, will naturally be used in image areas when template matching (TM) fails. This new method offers a significant coding gain compared to H.264 Intra prediction modes and the template matching based prediction. Indeed, the TM-based algorithm and the proposed hybrid algorithm lead, with the Bjontergaard measure, to rate gains of up to respectively 38.02% and 48.38% at low bitrates when compared with H.264 Intra only. Safa Chérigui, Christine Guillemot, Dominique Thoreau, Philippe Guillotel, Patrick Pérez |
ICASSP | 5 |
| 2012 | Map-Aided Locally Linear Embedding methods for image predictionabstractImage prediction methods based on data dimensionality reduction techniques have been introduced in [1]. Although efficient, these methods suffer from limitations when the block to be predicted and its neighborhood (or template) are not correlated, e.g. in non homogenous texture areas. To cope with these limitations, this paper introduces new image prediction methods based on locally linear embedding (LLE) technique in which the required K-NN search is aided, at the decoder, by a block correspondence map, hence the name Map-Aided Locally Linear Embedding (MALLE) method. Another optimized variant of this approach, called oMALLE method, is also studied. The resulting prediction methods are shown to bring significant Rate-Distortion (RD) performance improvements when compared to H.264 Intra prediction modes (up to 40.78 % rate saving at low bit rates). Safa Chérigui, Christine Guillemot, Dominique Thoreau, Philippe Guillotel, Patrick Pérez |
ICIP | 5 |
| 2012 | From optical flow to dense long term correspondencesabstractDense point matching and tracking in image sequences is an open issue with implications in several domains, from content analysis to video editing. We observe that for long term dense point matching, some regions of the image are better matched by concatenation of consecutive motion vectors, while for others a direct long term matching is preferred. We propose a method to optimally estimate the correspondence of a point w.r.t. a reference image from a set of input motion estimations over different temporal intervals. Results on texture insertion by point tracking in the context of video editing are presented and compared with a state-of-the-art approach. Tomás Crivelli, Pierre-Henri Conze, Philippe Robert, Patrick Pérez |
ICIP | 4 |
| 2012 | A contrario shot detectionabstractThis paper presents a novel technique for video shot detection, including hard cuts and gradual transitions. The method builds on various similarity criteria, namely histogram difference, motion estimation and distribution of intensity difference. While these metrics are well known, the novelty of the paper resides in the probabilistic framework to assess the similarity between two images. The a contrario framework, introduced in [1, 2] enables to control explicitly the number of false alarms given a background noise model. Experiments have been conducted on the Trecvid 2007 database: for all transitions (cuts and gradual), a recall of 0.95 and a precision of 0.96 was obtained. Pierre Hellier, Vincent Demoulin, Lionel Oisel, Patrick Pérez |
ICIP | 4 |
| 2012 | Aggregating Local Image Descriptors into Compact CodesabstractThis paper addresses the problem of large-scale image search. Three constraints have to be taken into account: search accuracy, efficiency, and memory usage. We first present and evaluate different ways of aggregating local image descriptors into a vector and show that the Fisher kernel achieves better performance than the reference bag-of-visual words approach for any given vector dimension. We then jointly optimize dimensionality reduction and indexing in order to obtain a precise vector comparison as well as a compact representation. The evaluation shows that the image representation can be reduced to a few dozen bytes while preserving high accuracy. Searching a 100 million image data set takes about 250 ms on one processor core. Hervé Jégou, Florent Perronnin, Matthijs Douze, Jorge Sánchez 0002, Patrick Pérez, Cordelia Schmid |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2011 | Reconstructing an image from its local descriptorsabstractThis paper shows that an image can be approximately reconstructed based on the output of a blackbox local description software such as those classically used for image indexing. Our approach consists first in using an off-the-shelf image database to find patches that are visually similar to each region of interest of the unknown input image, according to associated local descriptors. These patches are then warped into input image domain according to interest region geometry and seamlessly stitched together. Final completion of still missing texture-free regions is obtained by smooth interpolation. As demonstrated in our experiments, visually meaningful reconstructions are obtained just based on image local descriptors like SIFT, provided the geometry of regions of interest is known. The reconstruction most often allows the clear interpretation of the semantic image content. As a result, this work raises critical issues of privacy and rights when local descriptors of photos or videos are given away for indexing and search purpose. Philippe Weinzaepfel, Hervé Jégou, Patrick Pérez |
CVPR | 3 |
| 2011 | Context-driven moving object detection in aerial scenes with user inputabstractAerial video sequences are a common source for applications such as intelligence, surveillance or search and rescue. Their off-line analysis however requires a certain level of assistance to reduce the expert's workload. This study focuses on detecting mobile vehicles in such sequences. The proposed approach exploits two types of contextual information: loose user input as tagged areas in a reference frame, and knowledge-based priors to describe specific constraints. Our main contribution is the design of a two-step general framework able to combine these two types of information. The first step is a pixelwise semantic classification labelling each sequence frame structure in vehicle, road and background; the classifier is based on local motion and appearance features and is organized as an iterative refining process. The second step exploits knowledge-based spatial reasoning to filter out false alarms. A quantitative evaluation on real video sequences demonstrates the usefulness of each level of contextual information. Christophe Guilmart, Stéphane Herbin, Patrick Pérez |
ICIP | 3 |
| 2011 | Joint pose estimation and action recognition in image graphsabstractHuman analysis in images and video is a hard problem due to the large variation in human pose, clothing, camera view-points, lighting and other factors. While the explicit modeling of this variability is difficult, the huge amount of available person images motivates for the implicit, data-driven approach to human analysis. In this work we aim to explore this approach using the large amount of images spanning a subspace of human appearance. We model this subspace by connecting images into a graph and propagating information through such a graph using a discriminatively-trained graphical model. We particularly address the problems of human pose estimation and action recognition and demonstrate how image graphs help solving these problems jointly. We report results on still images with human actions from the KTH dataset. Kumar Raja, Ivan Laptev, Patrick Pérez, Lionel Oisel |
ICIP | 3 |
| 2011 | Epitome-based image compression using translational sub-pel mappingabstractThis paper addresses the problem of epitome construction for image compression. An optimized epitome construction method is first described, where the epitome and the associated image reconstruction, are both successively performed at full pel and sub-pel accuracy. The resulting complete still image compression scheme is then discussed with details on some innovative tools. The PSNR-rate performance achieved with this epitome-based compression method is significantly higher than the one obtained with H.264 Intra and with state of the art epitome construction method. A bit-rate saving up to 16% comparatively to H.264 Intra is achieved. Safa Chérigui, Christine Guillemot, Dominique Thoreau, Philippe Guillotel, Patrick Pérez |
MMSP | 5 |
| 2011 | View-Independent Action Recognition from Temporal Self-SimilaritiesabstractThis paper addresses recognition of human actions under view changes. We explore self-similarities of action sequences over time and observe the striking stability of such measures across views. Building upon this key observation, we develop an action descriptor that captures the structure of temporal similarities and dissimilarities within an action sequence. Despite this temporal self-similarity descriptor not being strictly view-invariant, we provide intuition and experimental validation demonstrating its high stability under view changes. Self-similarity descriptors are also shown to be stable under performance variations within a class of actions when individual speed fluctuations are ignored. If required, such fluctuations between two different instances of the same action class can be explicitly recovered with dynamic time warping, as will be demonstrated, to achieve cross-view action synchronization. More central to the current work, temporal ordering of local self-similarity descriptors can simply be ignored within a bag-of-features type of approach. Sufficient action discrimination is still retained in this way to build a view-independent action recognition system. Interestingly, self-similarities computed from different image features possess similar properties and can be used in a complementary fashion. Our method is simple and requires neither structure recovery nor multiview correspondence estimation. Instead, it relies on weak geometric properties and combines them with machine learning for efficient cross-view action recognition. The method is validated on three public data sets. It has similar or superior performance compared to related methods and it performs well even in extreme conditions, such as when recognizing actions from top views while using side views only for training. Imran N. Junejo, Emilie Dexter, Ivan Laptev, Patrick Pérez |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Geometrically Guided Exemplar-Based InpaintingabstractExemplar-based methods have proven their efficiency for the reconstruction of missing parts in a digital image. Texture as well as local geometry are often very well restored by such methods. Some applications, however, require the ability to reconstruct nonlocal geometric features, e.g., long edges. In order to do so, we propose to first compute a geometric sketch, which is then interpolated and used as a guide for the global reconstruction. In comparison with other related approaches, the originality of our work relies on the following points: (1) The geometric sketch computation is parameter-free and based on level lines, which provides a complete, reliable, and stable representation of the image. (2) The completion of the geometric sketch is fully automatic. It is done using a new—and interesting on its own—geometric inpainting approach that interpolates level lines with Euler spirals. Euler spirals are natural curves for shape completion and have been used already for edge completion and inpainting. It is the first time, however, that these curves are used for completing the whole level lines structure. (3) The general reconstruction is performed using a guided version of a classical exemplar-based method. However, we do not constrain the exemplar-based reconstruction to strictly follow the geometric guide. We actually use a new metric between blocks that consists of the sum of the classical ${{\mathrm L}^2}$ metric between any two blocks of the general image plus an ${{\mathrm L}^2}$ metric between the corresponding blocks in the completed geometric image. This is equivalent to a Lagrangian relaxation of a strictly guided reconstruction. We discuss in the paper the details of the method and some related mathematical issues, and we illustrate its efficiency on several examples. Frédéric Cao, Yann Gousseau, Simon Masnou, Patrick Pérez |
SIAM J. Imaging Sci. | 4 |
| 2010 | Aggregating local descriptors into a compact image representationabstractWe address the problem of image search on a very large scale, where three constraints have to be considered jointly: the accuracy of the search, its efficiency, and the memory usage of the representation. We first propose a simple yet efficient way of aggregating local image descriptors into a vector of limited dimension, which can be viewed as a simplification of the Fisher kernel representation. We then show how to jointly optimize the dimension reduction and the indexing algorithm, so that it best preserves the quality of vector comparison. The evaluation shows that our approach significantly outperforms the state of the art: the search accuracy is comparable to the bag-of-features approach for an image representation that fits in 20 bytes. Searching a 10 million image dataset takes about 50ms. Hervé Jégou, Matthijs Douze, Cordelia Schmid, Patrick Pérez |
CVPR | 4 |
| 2010 | Compact Video Description for Copy Detection with Precise Temporal Alignment
Matthijs Douze, Hervé Jégou, Cordelia Schmid, Patrick Pérez |
ECCV (1) | 4 |
| 2010 | Stochastic Filtering of Level Sets for Curve TrackingabstractThis paper focuses on the tracking of free curves using non-linear stochastic filtering techniques. It relies on a particle filter which includes color measurements. The curve and its velocity are defined through two coupled implicit level set representations. The stochastic dynamics of the curve is expressed directly on the level set function associated to the curve representation and combines a velocity field captured from the additional second level set attached to the past curve's points location. The curve's dynamics combines a low-dimensional noise model and a data-driven local force. We demonstrate how this approach allows the tracking of highly and rapidly deforming objects, such as convective cells in infra-red satellite images, while providing a location-dependent assessment of the estimation confidence. Christophe Avenel, Étienne Mémin, Patrick Pérez |
ICPR | 3 |
| 2010 | Geodesic image and video editingabstractThis article presents a new, unified technique to perform general edge-sensitive editing operations on n-dimensional images and videos efficiently. The first contribution of the article is the introduction of a Generalized Geodesic Distance Transform (GGDT), based on soft masks. This provides a unified framework to address several edge-aware editing operations. Diverse tasks such as denoising and nonphotorealistic rendering are all dealt with fundamentally the same, fast algorithm. Second, a new Geodesic Symmetric Filter (GSF) is presented which imposes contrast-sensitive spatial smoothness into segmentation and segmentation-based editing tasks (cutout, object highlighting, colorization, panorama stitching). The effect of the filter is controlled by two intuitive, geometric parameters. In contrast to existing techniques, the GSF filter is applied to real-valued pixel likelihoods (soft masks), thanks to GGDTs and it can be used for both interactive and automatic editing. Complex object topologies are dealt with effortlessly. Finally, the parallelism of GGDTs enables us to exploit modern multicore CPU architectures as well as powerful new GPUs, thus providing great flexibility of implementation and deployment. Our technique operates on both images and videos, and generalizes naturally to n-dimensional data. The proposed algorithm is validated via quantitative and qualitative comparisons with existing, state-of-the-art approaches. Numerous results on a variety of image and video editing tasks further demonstrate the effectiveness of our method. Antonio Criminisi, Toby Sharp, Carsten Rother, Patrick Pérez |
ACM Trans. Graph. | 4 |
| 2009 | Multi-view Synchronization of Human Actions and Dynamic ScenesabstractThis paper deals with the temporal synchronization of image sequences. Two instances of this problem are considered: (a) synchronization of human actions and (b) synchronization of dynamic scenes with view changes. To address both tasks and to reliably handle large view variations, we use self-similarity matrices which remain stable across views. We propose time-adaptive descriptors that capture the structure of these matrices while being invariant to the impact of time warps between views. Synchronizing two sequences is then performed by aligning their temporal descriptors using the Dynamic Time Warping algorithm. We present quantitative comparison results between time-fixed and time-adaptive descriptors for image sequences with different frame rates. We also illustrate the performance of the approach on several challenging videos with large view variations, drastic independent camera motions and within-class variability of human actions. Emilie Dexter, Patrick Pérez, Ivan Laptev |
BMVC | 2 |
| 2009 | Detection and segmentation of moving objects in complex scenes
Aurélie Bugeau, Patrick Pérez |
Comput. Vis. Image Underst. | 2 |
| 2008 | Semi-automatic Motion Segmentation with Motion Layer Mosaics
Matthieu Fradet, Patrick Pérez, Philippe Robert |
ECCV (3) | 2 |
| 2008 | Cross-View Action Recognition from Temporal Self-similarities
Imran N. Junejo, Emilie Dexter, Ivan Laptev, Patrick Pérez |
ECCV (2) | 4 |
| 2008 | Time-sequential extraction of motion layersabstractA new time-sequential approach for motion layer extraction is presented. We assume that the scene can be described by a set of layers associated to affine motion models. In one or more key frames, the segmentation is obtained using a semi-automatic method. At a subsequent instant, the first step of the proposed algorithm is the prediction of the segmentation from one image to the next one, using motion models estimated for each layer. The second step is the refinement of the predicted motion boundaries by graph cut. Only the appearing areas and a strip around the predicted boundaries are questioned. A new rigidity constraint improves the temporal consistency of foreground rigid objects. Experimental results show that our sequential approach is at least as effective as more complex simultaneous approaches while being less computationally demanding. Matthieu Fradet, Patrick Pérez, Philippe Robert |
ICIP | 2 |
| 2007 | Joint Tracking and Segmentation of Objects Using Graph Cuts
Aurélie Bugeau, Patrick Pérez |
ACIVS | 2 |
| 2007 | Detection and segmentation of moving objects in highly dynamic scenesabstractDetecting and segmenting moving objects in dynamic scenes is a hard but essential task in a number of applications such as surveillance. Most existing methods only give good results in the case of persistent or slowly changing background, or if both the objects and the background are rigid. In this paper, we propose a new method for direct detection and segmentation of foreground moving objects in the absence of such constraints. First, groups of pixels having similar motion and photometric features are extracted. For this first step only a sub-grid of image pixels is used to reduce computational cost and improve robustness to noise. We introduce the use of p-value to validate optical flow estimates and of automatic bandwidth selection in the mean shift clustering algorithm. In a second stage, segmentation of the object associated to a given cluster is performed in a MAP/MRF framework. Our method is able to handle moving camera and several different motions in the background. Experiments on challenging sequences show the performance of the proposed method and its utility for video analysis in complex scenes. Aurélie Bugeau, Patrick Pérez |
CVPR | 2 |
| 2007 | Probabilistic Color and Adaptive Multi-Feature Tracking with Dynamically Switched Priority Between CuesabstractWe present a probabilistic multi-cue tracking approach constructed by employing a novel randomized template tracker and a constant color model based particle filter. Our approach is based on deriving simple binary confidence measures for each tracker which aid priority based switching between the two fundamental cues for state estimation. Thereby the state of the object is estimated from one of the two distributions associated to the cues at each tracking step. This switching also brings about interaction between the cues at irregular intervals in the form of cross sampling. Within this scheme, we tackle the important aspect of dynamic target model adaptation under randomized template tracking which, by construction, possesses the ability to adapt to changing object appearances. Further, to track the object through occlusions we interrupt sequential resampling and achieve relock using the color cue. In order to evaluate the efficacy of this scheme, we put it to test against several state of art trackers using the VIVID online evaluation program and make quantitative comparisons. Vijay Badrinarayanan, Patrick Pérez, François Le Clerc, Lionel Oisel |
ICCV | 2 |
| 2007 | Retrieving actions in moviesabstractWe address recognition and localization of human actions in realistic scenarios. In contrast to the previous work studying human actions in controlled settings, here we train and test algorithms on real movies with substantial variation of actions in terms of subject appearance, motion, surrounding scenes, viewing angles and spatio-temporal extents. We introduce a new annotated human action dataset and use it to evaluate several existing methods. We in particular focus on boosted space-time window classifiers and introduce "keyframe priming" that combines discriminative models of human motion and shape within an action. Keyframe priming is shown to significantly improve the performance of action detection. We present detection results for the action class "drinking" evaluated on two episodes of the movie "Coffee and Cigarettes". Ivan Laptev, Patrick Pérez |
ICCV | 2 |
| 2007 | On Uncertainties, Random Features and Object TrackingabstractAlgorithms for probabilistic visual tracking hypothesize a distribution of the target state (location, scale, etc.) at every tracking step with an associated information content or equivalently, an uncertainty. One measure of this uncertainty is the differential entropy. In this paper, we present a unified way to approximate the differential entropy of tracking distributions, which then makes it suitable, among other factors, for a qualitative assessment of both deterministic and sequential Monte Carlo simulation based tracking algorithms. We then illustrate the usefulness of this assessment measure via tracking an object by choosing a set of randomly picked features on it, each individually tracked, removed according to an uncertainty analysis and replaced randomly, without any aid of a feature selection algorithm as in current use. Vijay Badrinarayanan, Patrick Pérez, François Le Clerc, Lionel Oisel |
ICIP (5) | 2 |
| 2007 | Robust tracking with motion estimation and local Kernel-based color modeling
Venkatesh Babu Radhakrishnan, Patrick Pérez, Patrick Bouthemy |
Image Vis. Comput. | 2 |
| 2006 | Kernel-Based Robust Tracking for Objects Undergoing Occlusion
Venkatesh Babu Radhakrishnan, Patrick Pérez, Patrick Bouthemy |
ACCV (2) | 2 |
| 2006 | An adaptive mixture color model for robust visual trackingabstractGlobal color characterization is a very powerful tool to model in a simple yet discriminant way the visual appearance of complex objects. A fixed reference model of this type can be used within both deterministic and probabilistic sequential estimation frameworks to track targets that undergo drastic changes of detailed appearance. However, changes of illumination as well as occlusions require that reference model is updated while avoiding drift. Within the particle filtering framework, we propose to address this adaptation problem using a dynamic mixture of color models with two components which are respectively fixed and rapidly updated. The merit of this approach is demonstrated on tracking players in team sport videos. Antoine Lehuger, Patrick Lechat, Patrick Pérez |
ICIP | 3 |
| 2005 | Periodic Motion Detection and Segmentation via Approximate Sequence AlignmentabstractA method for detecting and segmenting periodic motion is presented. We exploit periodicity as a cue and detect periodic motion in complex scenes where common methods for motion segmentation are likely to fail. We note that periodic motion detection can be seen as an approximate case of sequence alignment where an image sequence is matched to itself over one or more periods of time. To use this observation, we first consider alignment of two video sequences obtained by independently moving cameras. Under assumption of constant translation, the fundamental matrices and the homographies are shown to be time-linear matrix functions. These dynamic quantities can be estimated by matching corresponding space-time points with similar local motion and shape. For periodic motion, we match corresponding points across periods and develop a RANSAC procedure to simultaneously estimate the period and the dynamic geometric transformations between periodic views. Using this method, we demonstrate detection and segmentation of human periodic motion in complex scenes with nonrigid backgrounds, moving camera and motion parallax. Ivan Laptev, Serge J. Belongie, Patrick Pérez, Josh Wills |
ICCV | 3 |
| 2005 | Robust tracking with motion estimation and kernel-based color modellingabstractVisual tracking is still a challenging problem in computer vision. The applications of visual tracking are far-reaching, ranging from surveillance and monitoring to smart rooms. In this work, we propose a new method to track arbitrary objects using both sum-of-squared differences (SSD) and color-based mean-shift (MS) trackers in the Kalman filter framework. The SSD and the MS trackers complement each other by overcoming their respective disadvantages. The rapid model change in SSD tracker is overcome by the MS tracker module, while the inability of MS tracker to handle large displacements and occlusions is circumvented by the SSD module. In addition, rapid scale changes of the object generated by camera ego-motion or zooming are measured by a global affine motion estimation. Finally, the global appearance model on which MS relies is updated, based on the Bhattacharyya distance between this target model and current candidate model. This permits to tackle global appearance changes of the object. The performance of the proposed tracker is better than the individual SSD and MS trackers. Venkatesh Babu Radhakrishnan, Patrick Pérez, Patrick Bouthemy |
ICIP (1) | 2 |
| 2005 | Bayesian visual tracking with existence processabstractMost object tracking approaches either assume that the number of objects is constant, or that information about object existence is provided by some external source. Here, we show how object existence can be rigorously integrated within the Bayesian single and multiple object tracking framework. We provide a general treatment that impacts as little as possible on existing tracking algorithms, so that software can be reused, and that allows implementation with Kalman filters, extended Kalman filters, particle filters, etc. We apply the proposed framework to colour-based tracking of multiple objects. Jaco Vermaak, Simon Maskell, Mark Briers, Patrick Pérez |
ICIP (1) | 4 |
| 2004 | Interactive Image Segmentation Using an Adaptive GMMRF Model
Andrew Blake 0001, Carsten Rother, Matthew A. Brown, Patrick Pérez, Philip Torr 0001 |
ECCV (1) | 4 |
| 2004 | Hierarchical Markovian segmentation of multispectral images for the reconstruction of water depth maps
Jean-Noël Provost, Christophe Collet 0001, Philippe Rostaing, Patrick Pérez, Patrick Bouthemy |
Comput. Vis. Image Underst. | 4 |
| 2004 | Data fusion for visual tracking with particlesabstractThe effectiveness of probabilistic tracking of objects in image sequences has been revolutionized by the development of particle filtering. Whereas Kalman filters are restricted to Gaussian distributions, particle filters can propagate more general distributions, albeit only approximately. This is of particular benefit in visual tracking because of the inherent ambiguity of the visual world that stems from its richness and complexity. One important advantage of the particle filtering framework is that it allows the information from different measurement sources to be fused in a principled manner. Although this fact has been acknowledged before, it has not been fully exploited within a visual tracking context. Here we introduce generic importance sampling mechanisms for data fusion and discuss them for fusing color with either stereo sound, for teleconferencing, or with motion, for surveillance with a still camera. We show how each of the three cues can be modeled by an appropriate data likelihood function, and how the intermittent cues (sound or motion) are best handled by generating proposal distributions from their likelihood functions. Finally, the effective fusion of the cues by particle filtering is demonstrated on real teleconference and surveillance data. Patrick Pérez, Jaco Vermaak, Andrew Blake 0001 |
Proc. IEEE | 1 |
| 2004 | Region filling and object removal by exemplar-based image inpaintingabstractA new algorithm is proposed for removing large objects from digital images. The challenge is to fill in the hole that is left behind in a visually plausible way. In the past, this problem has been addressed by two classes of algorithms: 1) "texture synthesis" algorithms for generating large image regions from sample textures and 2) "inpainting" techniques for filling in small image gaps. The former has been demonstrated for "textures"--repeating two-dimensional patterns with some stochasticity; the latter focus on linear "structures" which can be thought of as one-dimensional patterns, such as lines and object contours. This paper presents a novel and efficient algorithm that combines the advantages of these two approaches. We first note that exemplar-based texture synthesis contains the essential process required to replicate both texture and structure; the success of structure propagation, however, is highly dependent on the order in which the filling proceeds. We propose a best-first algorithm in which the confidence in the synthesized pixel values is propagated in a manner similar to the propagation of information in inpainting. The actual color values are computed using exemplar-based synthesis. In this paper, the simultaneous propagation of texture and structure information is achieved by a single, efficient algorithm. Computational efficiency is achieved by a block-based sampling process. A number of examples on real and synthetic images demonstrate the effectiveness of our algorithm in removing large occluding objects, as well as thin scratches. Robustness with respect to the shape of the manually selected target region is also demonstrated. Our results compare favorably to those obtained by existing techniques. Antonio Criminisi, Patrick Pérez, Kentaro Toyama |
IEEE Trans. Image Process. | 2 |
| 2003 | Object Removal by Exemplar-Based InpaintingabstractA new algorithm is proposed for removing large objects from digital images. The challenge is to fill in the hole that is left behind in a visually plausible way. In the past, this problem has been addressed by two classes of algorithms: (i) "texture synthesis" algorithms for generating large image regions from sample textures, and (ii) "inpainting" techniques for filling in small image gaps. The former work well for "textures" - repeating two dimensional patterns with some stochasticity; the latter focus on linear "structures" which can be thought of as one dimensional patterns, such as lines and object contours. This paper presents a novel and efficient algorithm that combines the advantages of these two approaches. We first note that exemplar-based texture synthesis contains the essential process required to replicate both texture and structure; the success of structure propagation, however, is highly dependent on the order in which the filling proceeds. We propose a best-first algorithm in which the confidence in the synthesized pixel values is propagated in a manner similar to the propagation of information in inpainting. The actual color values are computed using exemplar-based synthesis. Computational efficiency is achieved by a block-based sampling process. A number of examples on real and synthetic images demonstrate the effectiveness of our algorithm in removing large occluding objects as well as thin scratches. Robustness with respect to the shape of the manually selected target region is also demonstrated. Our results compare favorably to those obtained by existing techniques. Antonio Criminisi, Patrick Pérez, Kentaro Toyama |
CVPR (2) | 2 |
| 2003 | Variational Inference for Visual TrackingabstractThe likelihood models used in probabilistic visual tracking applications are often complex non-linear and/or non-Gaussian functions, leading to analytically intractable inference. Solutions then require numerical approximation techniques, of which the particle filter is a popular choice. Particle filters, however, degrade in performance as the dimensionality of the state space increases and the support of the likelihood decreases. As an alternative to particle filters this paper introduces a variational approximation to the tracking recursion. The variational inference is intractable in itself, and is combined with an efficient importance sampling procedure to obtain the required estimates. The algorithm is shown to compare favorably with particle filtering techniques on a synthetic example and two real tracking problems. The first involves the tracking of a designated object in a video sequence based on its color properties, whereas the second involves contour extraction in a single image. Jaco Vermaak, Neil D. Lawrence, Patrick Pérez |
CVPR (1) | 3 |
| 2003 | Constrained Subspace ModellingabstractWhen performing subspace modeling of data using principal component analysis (PCA) it may be desirable to constrain certain directions to be more meaningful in the context of the problem being investigated. This need arises due to the data often being approximately isotropic along the lesser principal components, making the choice of directions for these components more-or-less arbitrary. Furthermore, constraining may be imperative to ensure viable solutions in problems where the dimensionality of the data space is of the same order as the number of data points available. This paper adopts a Bayesian approach and augments the likelihood implied by probabilistic principal component analysis (PPCA) (Tipping and Bishop, 1999) with a prior designed to achieve the constraining effect. The subspace parameters are computed efficiently using the EM algorithm. The constrained modeling approach is illustrated on two pertinent problems, one from speech analysis, and one from computer vision. Jaco Vermaak, Patrick Pérez |
CVPR (2) | 2 |
| 2003 | Maintaining Multi-Modality through Mixture TrackingabstractIn recent years particle filters have become a tremendously popular tool to perform tracking for nonlinear and/or nonGaussian models. This is due to their simplicity, generality and success over a wide range of challenging applications. Particle filters, and Monte Carlo methods in general, are however poor at consistently maintaining the multimodality of the target distributions that may arise due to ambiguity or the presence of multiple objects. To address this shortcoming this paper proposes to model the target distribution as a nonparametric mixture model, and presents the general tracking recursion in this case. It is shown how a Monte Carlo implementation of the general recursion leads to a mixture of particle filters that interact only in the computation of the mixture weights, thus leading to an efficient numerical algorithm, where all the results pertaining to standard particle filters apply. The ability of the new method to maintain posterior multimodality is illustrated on a synthetic example and a real world tracking problem involving the tracking of football players in a video sequence. Jaco Vermaak, Arnaud Doucet, Patrick Pérez |
ICCV | 3 |
| 2003 | Poisson image editingabstractUsing generic interpolation machinery based on solving Poisson equations, a variety of novel tools are introduced for seamless editing of image regions. The first set of tools permits the seamless importation of both opaque and transparent source image regions into a destination region. The second set is based on similar mathematical ideas and allows the user to modify the appearance of the image seamlessly, within a selected region. These changes can be arranged to affect the texture, the illumination, and the color of objects lying in the region, or to make tileable a rectangular selection. Patrick Pérez, Michel Gangnet, Andrew Blake 0001 |
ACM Trans. Graph. | 1 |
| 2002 | Rapid Summarisation and Browsing of Video SequencesabstractThis paper presents a strategy for rapid summarisation and browsing of video sequences. The input video is first transformed into a sequence of representative feature vectors. Using this representation a utility function is designed that assigns high reward to subsequences of keyframes that are maximally distinct and individually carry the most information. For a specified level of detail and endpoints the keyframe sequence that maximises this utility function can be obtained by a non-iterative Dynamic Programming procedure, thus allowing the user to efficiently zoom in on any part or all of the video sequence. For the sake of compactness and clarity the working of the algorithm is illustrated on a television commercial. 1 Jaco Vermaak, Patrick Pérez, Michel Gangnet, Andrew Blake 0001 |
BMVC | 2 |
| 2002 | Dense Motion Analysis in Fluid Imagery
Thomas Corpetti, Étienne Mémin, Patrick Pérez |
ECCV (1) | 3 |
| 2002 | Color-Based Probabilistic Tracking
Patrick Pérez, Carine Hue, Jaco Vermaak, Michel Gangnet |
ECCV (1) | 1 |
| 2002 | Towards Improved Observation Models for Visual Tracking: Selective Adaptation
Jaco Vermaak, Patrick Pérez, Michel Gangnet, Andrew Blake 0001 |
ECCV (1) | 2 |
| 2002 | Hierarchical Estimation and Segmentation of Dense Motion Fields
Étienne Mémin, Patrick Pérez |
Int. J. Comput. Vis. | 2 |
| 2002 | Dense Estimation of Fluid FlowsabstractIn this paper, we address the problem of estimating and analyzing the motion of fluids in image sequences. Due to the great deal of spatial and temporal distortions that intensity patterns exhibit in images of fluids, the standard techniques from computer vision, originally designed for quasi-rigid motions with stable salient features, are not well adapted in this context. We thus investigate a dedicated minimization-based motion estimator. The cost function to be minimized includes a novel data term relying on an integrated version of the continuity equation of fluid mechanics, which is compatible with large displacements. This term is associated with an original second-order div-curl regularization which prevents the washing out of the salient vorticity and divergence structures. The performance of the resulting fluid flow estimator is demonstrated on meteorological satellite images. In addition, we show how the sequences of dense motion fields we estimate can be reliably used to reconstruct trajectories and to extract the regions of high vorticity and divergence. Thomas Corpetti, Étienne Mémin, Patrick Pérez |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2002 | Nonparametric motion characterization using causal probabilistic models for video indexing and retrievalabstractThis paper describes an original approach for content-based video indexing and retrieval. We aim at providing a global interpretation of the dynamic content of video shots without any prior motion segmentation and without any use of dense optic flow fields. To this end, we exploit the spatio-temporal distribution, within a shot, of appropriate local motion-related measurements derived from the spatio-temporal derivatives of the intensity function. These distributions are then represented by causal Gibbs models. To be independent of camera movement, the motion-related measurements are computed in the image sequence generated by compensating the estimated dominant image motion in the original sequence. The statistical modeling framework considered makes the exact computation of the conditional likelihood of a video shot belonging to a given motion or more generally to an activity class feasible. This property allows us to develop a general statistical framework for video indexing and retrieval with query-by-example. We build a hierarchical structure of the processed video database according to motion content similarity. This results in a binary tree where each node is associated to an estimated causal Gibbs model. We consider a similarity measure inspired from Kullback-Leibler divergence. Then, retrieval with query-by-example is performed through this binary tree using the maximum a posteriori (MAP) criterion. We have obtained promising results on a set of various real image sequences. Ronan Fablet, Patrick Bouthemy, Patrick Pérez |
IEEE Trans. Image Process. | 3 |
| 2001 | A New Algorithm for Super-Resolution from Image Sequences
Fabien Dekeyser, Patrick Bouthemy, Patrick Pérez |
CAIP | 3 |
| 2001 | JetStream: Probabilistic Contour Extraction with ParticlesabstractThe problem of extracting continuous structures from noisy or cluttered images is a difficult one. Successful extraction depends critically on the ability to balance prior constraints on continuity and smoothness against evidence garnered from image analysis. Exact, deterministic optimisation algorithms, based on discretized functionals, suffer from severe limitations on the form of prior constraint that can be imposed tractably. This paper proposes a sequential Monte-Carlo technique, termed JetStream, that enables constraints on curvature, corners, and contour parallelism. To be mobilized, all of which are infeasible under exact optimization. The power of JetStream is demonstrated in two contexts: (1) interactive cut-out in photo-editing applications, and (2) the recovery of roads in aerial photographs. Patrick Pérez, Andrew Blake 0001, Michel Gangnet |
ICCV | 1 |
| 2001 | Sequential Monte Carlo Fusion of Sound and Vision for Speaker Tracking
Jaco Vermaak, Michel Gangnet, Andrew Blake 0001, Patrick Pérez |
ICCV | 4 |
| 2000 | An Energy-Based Framework for Dense 3D Registration of Volumetric Brain ImagesabstractIn this paper we describe a new method for medical image registration. The registration is formulated as a minimization problem involving robust estimators. We propose an efficient hierarchical optimization framework which is both multiresolution and multigrid. An anatomical segmentation of the cortex is introduced in the adaptive partitioning of the volume on which the multigrid minimization is based. This allows to limit the estimation to the areas of interest, to accelerate the algorithm, and to refine the estimation in specified areas. Furthermore we introduce a methodology to constrain the registration with landmarks such as anatomical structures. The performances of this method are objectively evaluated on simulated data and its benefits are demonstrated on a large database of real acquisitions. Pierre Hellier, Christian Barillot, Étienne Mémin, Patrick Pérez |
CVPR | 4 |
| 2000 | Spot Satellite Data Analysis for Bathymetric MappingabstractThis paper presents the determination of bathymetric maps from the analysis of multispectral SPOT images. To this end, we have developed a multispectral segmentation method based on a hierarchical Markovian modeling including the unsupervised estimation of the model parameters. In each segmented region, an adaptive bathymetric inversion model is then applied in order to recover the water depth (mainly in coastal areas). Bathymetric estimation has been validated on real data, for which control points are available and correspond to bathymetric measures supplied by previous hydrographic campaigns. Christophe Collet 0001, Jean-Noël Provost, Philippe Rostaing, Patrick Pérez, Patrick Bouthemy |
ICIP | 4 |
| 2000 | Spatio-Temporal Wiener Filtering of Image Sequences Using a Parametric Motion ModelabstractThis paper deals with the use of a 2D parametric motion model in a spatio-temporal filtering scheme to reduce noise in image sequences. We estimate with a robust method an affine motion model accounting for the dominant image motion. Then, we cancel it before applying an adaptive spatiotemporal filter. We have compared the performance of several filtering techniques and evaluated the influence of the motion compensation step on this performance. Fabien Dekeyser, Patrick Bouthemy, Patrick Pérez |
ICIP | 3 |
| 2000 | Estimating Fluid Optical FlowabstractWe address the problem of fluid motion estimation in image sequences. For such motions, standard optical flow methods, based on intensity conservation and spatial coherence of motion field, are not suitable. This is due to the highly deformable nature of a fluid medium. For all applications where fluid motions are to be recovered from images, it is then important to have specific techniques. We investigate such dedicated models which include an original observation constraint, based on the continuity equation from fluid mechanics, and a new div-curl-type smoothness term. Our method is validated on synthetic and real meteorological images. Thomas Corpetti, Étienne Mémin, Patrick Pérez |
ICPR | 3 |
| 2000 | Super-Resolution from Noisy Image Sequences Exploiting a 2D Parametric Motion ModelabstractWe propose a low cost scheme for reconstructing high resolution images from noisy, and eventually blurred image sequences. The super-resolution is achieved by an iterative back projection method. To account for noise in image sequence, we first apply a spatio-temporal Wiener filter computed via a 3D DFT. In the filtering process, we need to compensate for apparent motion to ensure proper results. Furthermore, the knowledge of subpixel motion is necessary for super-resolution. In both cases, we exploit a parametric motion model to keep a good trade-off between accuracy and computation time. Fabien Dekeyser, Patrick Bouthemy, Patrick Pérez, Étienne Payot |
ICPR | 3 |
| 2000 | Markov Random Field and Fuzzy Logic Modeling in Sonar Imagery: Application to the Classification of Underwater Floor
Max Mignotte, Christophe Collet 0001, Patrick Pérez, Patrick Bouthemy |
Comput. Vis. Image Underst. | 3 |
| 2000 | Hybrid Genetic Optimization and Statistical Model-Based Approach for the Classification of Shadow Shapes in Sonar ImageryabstractWe present an original statistical classification method using a deformable template model to separate natural objects from man-made objects in an image provided by a high resolution sonar. A prior knowledge of the manufactured object shadow shape is captured by a prototype template, along with a set of admissible linear transformations, to take into account the shape variability. Then, the classification problem is defined as a two-step process: 1) the detection problem of a region of interest in the input image is stated as the minimization of a cost function; and 2) the value of this function at convergence allows one to determine whether the desired object is present or not in the sonar image. The energy minimization problem is tackled using relaxation techniques. In this context, we compare the results obtained with a deterministic relaxation technique and two stochastic relaxation methods: simulated annealing and a hybrid genetic algorithm. This latter method has been successfully tested on real and synthetic sonar images, yielding very promising results. Max Mignotte, Christophe Collet 0001, Patrick Pérez, Patrick Bouthemy |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2000 | Noniterative manipulation of discrete energy-based models for image analysis
Patrick Pérez, Annabelle Chardin, Jean-Marc Laferté |
Pattern Recognit. | 1 |
| 2000 | Discrete Markov image modeling and inference on the quadtreeabstractNoncasual Markov (or energy-based) models are widely used in early vision applications for the representation of images in high-dimensional inverse problems. Due to their noncausal nature, these models generally lead to iterative inference algorithms that are computationally demanding. In this paper, we consider a special class of nonlinear Markov models which allow one to circumvent this drawback. These models are defined as discrete Markov random fields (MRF) attached to the nodes of a quadtree. The quadtree induces causality properties which enable the design of exact, noniterative inference algorithms, similar to those used in the context of Markov chain models. We first introduce an extension of the Viterbi algorithm which enables exact maximum a posteriori (MAP) estimation on the quadtree. Two other algorithms, related to the MPM criterion and to Bouman and Shapiro's (1994) sequential-MAP (SMAP) estimator are derived on the same hierarchical structure. The estimation of the model hyper parameters is also addressed. Two expectation-maximization (EM)-type algorithms, allowing unsupervised inference with these models are defined. The practical relevance of the different models and inference algorithms is investigated in the context of image classification problem, on both synthetic and natural images. Jean-Marc Laferté, Patrick Pérez, Fabrice Heitz |
IEEE Trans. Image Process. | 2 |
| 2000 | Sonar image segmentation using an unsupervised hierarchical MRF modelabstractThis paper is concerned with hierarchical Markov random field (MRP) models and their application to sonar image segmentation. We present an original hierarchical segmentation procedure devoted to images given by a high-resolution sonar. The sonar image is segmented into two kinds of regions: shadow (corresponding to a lack of acoustic reverberation behind each object lying on the sea-bed) and sea-bottom reverberation. The proposed unsupervised scheme takes into account the variety of the laws in the distribution mixture of a sonar image, and it estimates both the parameters of noise distributions and the parameters of the Markovian prior. For the estimation step, we use an iterative technique which combines a maximum likelihood approach (for noise model parameters) with a least-squares method (for MRF-based prior). In order to model more precisely the local and global characteristics of image content at different scales, we introduce a hierarchical model involving a pyramidal label field. It combines coarse-to-fine causal interactions with a spatial neighborhood structure. This new method of segmentation, called the scale causal multigrid (SCM) algorithm, has been successfully applied to real sonar images and seems to be well suited to the segmentation of very noisy images. The experiments reported in this paper demonstrate that the discussed method performs better than other hierarchical schemes for sonar image segmentation. Max Mignotte, Christophe Collet 0001, Patrick Pérez, Patrick Bouthemy |
IEEE Trans. Image Process. | 3 |
| 1999 | Unsupervised Image Classification with a Hierarchical EM AlgorithmabstractThis work is undertaken in the context of hierarchical stochastic models for the resolution of discrete inverse problems from low level vision. Some of these models lie on the nodes of a quadtree which leads to non-iterative inference procedures. Nevertheless, if they circumvent the algorithmic drawbacks of grid-based models (computational load and/or great dependance on the initialization), they admit modeling shortcomings (cumbersome and somehow artificial). We investigate a new hierarchical stochastic model which benefits from both the spatial and hierarchical prior modeling. The independence graph is based on a tree which has been pollarded with nodes at the coarsest resolution exhibiting a grid-based interaction structure. For this class of model, we address the critical problem of parameter estimation. To this end, we derive an EM algorithm on the hybrid structure which mixes an exact EM algorithm on each subtree and a low cost Gibbs EM algorithm on the coarse spatial grid. Experiments on a synthetic image and multispectral satellite images are reported. Annabelle Chardin, Patrick Pérez |
ICCV | 2 |
| 1999 | Fluid Motion Recovery by Coupling Dense and Parametric Vector FieldsabstractAddresses the problem of estimating and analyzing the motion in image sequences that involve fluid phenomena. In this context, standard motion estimation techniques are not well-adapted, and more dedicated approaches have to be designed. We thus propose to estimate, in a joint and cooperative way, a dense motion field and a peculiar parametric representation of the flow. The parametric model is derived from an extension of the Rankine vortex model and includes a laminar flow field. Dense and parametric fields are estimated by minimizing a robust global objective function, thanks to a specific alternate scheme. The method has been validated on different kinds of meteorological image sequences. Étienne Mémin, Patrick Pérez |
ICCV | 2 |
| 1999 | Mode of Posterior Marginals with Hierarchical ModelsabstractThis work takes place in the context of hierarchical stochastic models for the resolution of discrete inverse problems from low level vision. We investigate a new hybrid hierarchical structure: a Markov random field attached to the nodes of a truncated tree. It thus combines causal hierarchical prior on trees with a non-causal spatial prior at the coarsest level. We address the problem of computing posterior marginals with such a prior structure. This is performed using non-iterative two-sweep marginalizations on trees, combined with a low cost Gibbs sampler in between the two sweeps. Posterior marginals are thus obtained in a semi-iterative way. They are then used to infer unknown variables according to the IMPM estimator. This is illustrated by experiments on synthetic data and on multispectral satellite images. Annabelle Chardin, Patrick Pérez |
ICIP (1) | 2 |
| 1999 | Dense/Parametric Estimation of Fluid FlowsabstractIn this work, we propose to cope in a joint and cooperative way with two types of problems that were addressed within distinct approaches so far: (a) the estimation of dense velocity (or displacement) fields of fluid flows, using a robust extension of the minimization-based approach proposed by Horn and Schunck; (b) the extraction of kinematic entities of particular interest from the fluid mechanics point of view: vortices, sinks, and sources. As concerns the second item, we extend Rankine vortex model, and thus design an original non-linear parametric model of the motion stemming from vortices, sinks and sources. We make interact this parametric model with the dense flow field within a global objective function, and we design a specific alternate scheme to perform the minimization of this function. Étienne Mémin, Patrick Pérez |
ICIP (3) | 2 |
| 1999 | A Hierarchical Unsupervised Multispectral Model to Segment Spot Images for Ocean CartographyabstractThis paper presents an unsupervised image segmentation method with applications to ocean cartography. By using SPOT satellite data, the aim is to improve the automatic production of bathymetric charts. Indeed, in coastal areas, satellite images provide the radiometry of the electromagnetic waves backscattered by the vegetation, the sea or the sea floor depending on the sea depth in littoral areas. The proposed segmentation method is based on a hierarchical Markovian model defined on a quad-tree, combining multispectral data. One of its interests is to take explicitly into account the correlation between the three spectral channels of the observation. Classification results obtained with synthetic and real images demonstrate the efficiency of the method. Jean-Noël Provost, Christophe Collet 0001, Patrick Pérez, Patrick Bouthemy |
ICIP (1) | 3 |
| 1999 | Medical Image Registration with Robust Multigrid Techniques
Pierre Hellier, Christian Barillot, Étienne Mémin, Patrick Pérez |
MICCAI | 4 |
| 1999 | Three-Class Markovian Segmentation of High-Resolution Sonar Images
Max Mignotte, Christophe Collet 0001, Patrick Pérez, Patrick Bouthemy |
Comput. Vis. Image Underst. | 3 |
| 1998 | Joint Estimation-Segmentation of Optic Flow
Étienne Mémin, Patrick Pérez |
ECCV (2) | 2 |
| 1998 | Statistical model and genetic optimization: application to pattern detection in sonar imagesabstractWe present a new classification method using a deformable template model to separate natural objects from man made objects in an image given by a high resolution sonar. A prior knowledge of the manufactured object shadow shape is described by a prototype template and a set of admissible linear transformations to take into account the shape variability. Then, the classification problem is defined as a two step process; firstly the detection problem of a region of interest in the input image is stated in a Bayesian framework and is posed as an equivalent energy minimization problem of an objective function: in this paper, this energy minimization problem is solved by using a hybrid genetic algorithm (GA). Secondly, the value of this function at convergence allows one to determine the presence of the desired object in the sonar image. This method has been successfully tested on real and synthetic sonar images, yielding very promising results. Max Mignotte, Christophe Collet 0001, Patrick Pérez, Patrick Bouthemy |
ICASSP | 3 |
| 1998 | A Multigrid Approach for Hierarchical Motion EstimationabstractThis paper focuses on the estimation of the apparent motion field between two consecutive frames in an image sequence. The approach developed here is a tradeoff between methods based on global parameterized flow models and local dense optic flow estimators. The method relies on an adaptive multigrid minimization approach. In addition to accelerated convergence toward good estimates, it allows to mix different parameterizations of the estimate relative to adaptive partitions of the image. The performances of the resulting algorithms are demonstrated in the difficult context of a non-convex energy. Experimental results on real world Meteosat sequences are presented. Étienne Mémin, Patrick Pérez |
ICCV | 2 |
| 1998 | Semi-Iterative Inference with Hierarchical ModelsabstractThis paper deals with hierarchical Markov random field models. We propose to introduce new hierarchical models based on a hybrid structure which combines a spatial grid of a reduced size at the coarsest level with sub-trees appended below it, down to the finest level. These models circumvent the algorithmic drawbacks of grid-based models (computational load and/or great dependance on the initialization) and the modeling drawbacks of tree-based approaches (cumbersome and somehow artificial structure). The hybrid structure leads to algorithms that mix a non-iterative inference on sub-trees with an iterative deterministic inference at the top of the structure. Experiments on a synthetic image demonstrate the gains provided in terms of both computational efficiency and the quality of results. Then experiments on real aerial images illustrate the ability of hybrid models to perform the multiresolution and multispectral image fusion. Annabelle Chardin, Patrick Pérez |
ICIP (1) | 2 |
| 1998 | Dense estimation and object-based segmentation of the optical flow with robust techniquesabstractIn this paper, we address the issue of recovering and segmenting the apparent velocity field in sequences of images. As for motion estimation, we minimize an objective function involving two robust terms. The first one cautiously captures the optical flow constraint, while the second (a priori) term incorporates a discontinuity-preserving smoothness constraint. To cope with the nonconvex minimization problem thus defined, we design an efficient deterministic multigrid procedure. It converges fast toward estimates of good quality, while revealing the large discontinuity structures of flow fields. We then propose an extension of the model by attaching to it a flexible object-based segmentation device based on deformable closed curves (different families of curve equipped with different kinds of prior can be easily supported). Experimental results on synthetic and natural sequences are presented, including an analysis of sensitivity to parameter tuning. Étienne Mémin, Patrick Pérez |
IEEE Trans. Image Process. | 2 |
| 1997 | Unsupervised Markovian segmentation of sonar imagesabstractThis work deals with unsupervised sonar image segmentation. We present a new estimation segmentation procedure using the an iterative method called iterative conditional estimation (ICE). This method takes into account the variety of the laws in the distribution mixture of a sonar image and the estimation of the parameters of the label field (modeled by a Markov random field (MRF)). For the estimation step we use a maximum likelihood estimation for the noise model parameters and the least square method proposed by Derin et al. (1987) to estimate the MRF prior model. Then, in order to obtain a good segmentation and to speed up the convergence rate, we use a multigrid strategy with the previously estimated parameters. This technique has been successfully applied to real sonar images and is compatible with an automatic treatment of massive amounts of data. Max Mignotte, Christophe Collet 0001, Patrick Pérez, Patrick Bouthemy |
ICASSP | 3 |
| 1997 | Generalized likelihood ratio-based face detection and extraction of mouth features
Charles Kervrann, Franck Davoine, Patrick Pérez, Robert Forchheimer, Claude Labit |
Pattern Recognit. Lett. | 3 |
| 1996 | Hierarchical MRF modeling for sonar picture segmentationabstractThis paper deals with sonar image segmentation based on a hierarchical Markovian modeling. The designed Markov random field (MRF) model takes into account both the phenomenon of speckle noise through Rayleigh's law, and notions of geometry related to the shape of object shadows. We adopt an 8-connexity neighbourhood in order to discriminate geometric and non-regular shadows. MRF are well adapted for this kind of segmentation where a priori knowledge about the shapes we are searching is available. Besides, the introduced hierarchical modeling allows us to successfully improve the sonar image segmentation while speeding up the iterative optimization scheme. Christophe Collet 0001, Pierre Thourel, Patrick Pérez, Patrick Bouthemy |
ICIP (3) | 3 |
| 1996 | Adaptive detection of moving objects using multiscale techniquesabstractIn this paper we address an important issue in motion analysis: the detection of moving objects. A statistical approach is adopted in order to formulate the problem. The inter-frame difference is modeled by a mixture of Laplacian distributions, and a Gibbs random field is used for describing the label set. A new method to determine the regularization parameter is proposed, based on a voting technique. Then two different multiscale algorithms are evaluated, and the labeling problem is solved using either ICM (iterated conditional modes) or HCF (highest confidence first) algorithms. Experimental results are provided using synthetic and real video sequences. Nikos Paragios, Patrick Pérez, Georgios Tziritas, Claude Labit, Patrick Bouthemy |
ICIP (1) | 2 |
| 1996 | Parallelized robust multiresolution motion estimationabstractMotion estimation is crucial in domains such as robot vision or image sequence coding. Fast or real time computing is often required but is problematic due to computational complexity. In this paper, an algorithm using multiresolution multigrid Markov random fields is parallelized. Such realizations have already been studied for fine parallelizations on massively parallel machines. The goal of this paper is to achieve a coarse parallelization on a network of workstations, using the Parallel Virtual Machine (PVM) software package. Patrick Piscaglia, Benoît Macq, Étienne Mémin, Patrick Pérez, Claude Labit |
ICIP (1) | 4 |
| 1996 | Statistical model-based estimation and tracking of non-rigid motionabstractWe describe a method for the temporal tracking of stochastic deformable models in image sequences. The object representation relies on a hierarchical statistical description of the deformations applied to a template. The optimal Bayesian estimate of deformations is obtained by maximizing nonlinear probability distributions using optimization techniques. The method may be sensitive to local maxima of the distributions and require an initial configuration close to the optimal solution. In our approach, the initialization is provided by a robust estimate of the rigid and statistically constrained nonrigid motions from the normal optical flow computed along the deformable contour. The approach is demonstrated on real-world sequences showing mouth movements and cardiac motions with missing data. Charles Kervrann, Fabrice Heitz, Patrick Pérez |
ICPR | 3 |
| 1996 | A multiresolution EM algorithm for unsupervised image classificationabstractUsing the causal Markov model defined on a quadtree we derive a multiresolution EM algorithm for unsupervised image classification. This algorithm is an efficient alternative to expensive or approximate EM algorithms associated with Markov random fields (MRFs). We show on synthetic and real images that our algorithm also provides good or even better results than those obtained by spatial MRF models. Jean-Marc Laferté, Fabrice Heitz, Patrick Pérez |
ICPR | 3 |
| 1996 | Robust discontinuity-preserving model for estimating optical flowabstractWe address the problem of incremental optical flow estimation. Following Black et al., we design a cost function whose data and prior terms both involve robust M-estimators. The non-convex minimization is lead with an efficient multigrid algorithm which converges fast toward estimates of good quality while providing, at very low cost, crude estimates revealing the large discontinuity structures of the flow field. Étienne Mémin, Patrick Pérez |
ICPR | 2 |
| 1996 | Restriction of a Markov random field on a graph and multiresolution statistical image modelingabstractThe association of statistical models and multiresolution data analysis in a consistent and tractable mathematical framework remains an intricate theoretical and practical issue. Several consistent approaches have been proposed previously to combine Markov random field (MRF) models and multiresolution algorithms in image analysis: renormalization group, subsampling of stochastic processes, MRFs defined on trees or pyramids, etc. For the simulation or a practical use of these models in statistical estimation, an important issue is the preservation of the local Markovian property of the representation at the different resolution levels. It is shown that this key problem may be studied by considering the restriction of a Markov random field (defined on some simple finite nondirected graph) to a part of its original site set. Several general properties of the restricted field are derived. The general form of the distribution of the restriction is given. "Locality" of the field is studied by exhibiting a neighborhood structure with respect to which the restricted field is an MRF. Sufficient conditions for the new neighborhood structure to be "minimal" are derived. Several consequences of these general results related to various "multiresolution" MRF-based modeling approaches in image analysis are presented. Patrick Pérez, Fabrice Heitz |
IEEE Trans. Inf. Theory | 1 |
| 1995 | Hierarchical Statistical Models for the Fusion of Multiresolution Image DataabstractThis paper presents a class of nonlinear hierarchical algorithms for the fusion of multiresolution image data in low-level vision. The approach combines nonlinear causal Markov models defined on hierarchical graph structures, with standard bayesian estimation theory. Two random processes defined on simple hierarchical graphs (quadtrees or "ternary graphs") are introduced to represent the multiresolution observations at hand and the hidden labels to be estimated. An optimal algorithm (inspired from the Viterbi algorithm) is developed to compute the bayesian estimates on the hierarchical graph structures. Estimates are obtained within two passes on the graph structure. This algorithm is non-iterative and yields a per pixel computational complexity which is independent of image size. This approach is compared to the multiscale algorithm proposed by (Bouman et al., 1994) for single-resolution image segmentation (that we have extended for multiresolution data fusion).> Jean-Marc Laferté, Fabrice Heitz, Patrick Pérez, Eric Fabre |
ICCV | 3 |
| 1994 | Global non-linear multigrid optimization for image analysis tasksabstractIn this paper, comprehensive non-linear optimization algorithms for the minimization of global objective functions in image analysis tasks, are investigated. We explore several versions of the approach corresponding to different exploration strategies of the multiresolution structure. The efficiency of this general approach to global optimization in image processing is demonstrated on a highly non-linear MRF motion estimation model. We show, that, associated with a deterministic relaxation algorithm, very significant gains on the convergence speed and on the quality of the final solutions may be obtained.> Jean-Marc Laferté, Patrick Pérez, Fabrice Heitz |
ICASSP (5) | 2 |
| 1994 | Motion Detection and Tracking using Deformable TemplatesabstractWe propose an object-based framework for detection and tracking of moving objects in a sequence of images. Two key ingredients of the approach are appropriate object models based on Grenander's (see General Pattern Theory, 1993) deformable templates and spatio-temporal data models. Detection and tracking problems are formulated as optimization problems. Detection employs a Metropolis-type procedure starting from a random initial configuration, while tracking involves a deterministic nonlinear Gauss-Seidel algorithm. We present experimental results with real data on a highway traffic sequence.> Patrick Pérez, Basilis Gidas |
ICIP (2) | 1 |
| 1992 | Multiscale Markov random fields and constrained relaxation in low level image analysisabstractThe authors investigate a new approach to multigrid image analysis based on Markov random field (MRF) models. The multigrid algorithms under consideration are based on constrained optimization schemes. The global optimization problem associated with MRF modeling is solved sequentially over particular subsets of the original configuration space. Those subsets consist of constrained configurations describing the desired resulting field at different scales. The constrained optimization can be implemented via a coarse-to-fine multigrid algorithm defined on a sequence of consistent multiscale MRF models. The proposed multiscale paradigm yields fast convergence toward high-quality estimates when compared to standard monoresolution or multigrid relaxation schemes.> Patrick Pérez, Fabrice Heitz |
ICASSP | 1 |