EDBT 2026 Demo / reviewers in the wild / expert
Matthieu Cord
dblp:68/3117
· DBLP profile ↗
179ranked-venue papers
6as first author
69since 2021 · last 2025
0000-0002-0627-5844ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 126 · 2 first-author · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 101 · 4 first-author · 30 since 2021Databases, data management, data science and information retrieval · 8Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Generalizable Trajectory Prediction using Dual-Level Representation Learning and Adaptive PromptingabstractExisting vehicle trajectory prediction models struggle with generalizability, prediction uncertainties, and handling complex interactions. It is often due to limitations like complex architectures customized for a specific dataset and inefficient multimodal handling. We propose Perceiver with Register queries (PerReg+), a novel trajectory prediction framework that introduces: (1) Dual-Level Representation Learning via Self-Distillation (SD) and Masked Reconstruction (MR), capturing global context and finegrained details. Additionally, our approach of reconstructing segment-level trajectories and lane segments from masked inputs with query drop, enables effective use of contextual information and improves generalization; (2) Enhanced Multimodality using register-based queries and pretraining, eliminating the need for clustering and suppression; and (3) Adaptive Prompt Tuning during fine-tuning, freezing the main architecture and optimizing a small number of prompts for efficient adaptation. PerReg+ sets a state-of-the-art performance on nuScenes [5], Argoverse 2 [42], and Waymo Open Motion Dataset (WOMD) [13]. Remarkably, our model reduces the error by 6.8% on smaller datasets, and multi-dataset training enhances generalization. In cross-domain tests, PerReg+ reduces B-FDE by 11.8% compared to its non-pretrained variant. Kaouther Messaoud, Matthieu Cord, Alexandre Alahi |
CVPR | 2 |
| 2025 | GaussRender: Learning 3D Occupancy with Gaussian RenderingabstractUnderstanding the 3D geometry and semantics of driving scenes is critical for safe autonomous driving. Recent advances in 3D occupancy prediction have improved scene representation but often suffer from spatial inconsistencies, leading to floating artifacts and poor surface localization. Existing voxel-wise losses (e.g., cross-entropy) fail to enforce geometric coherence. In this paper, we propose GaussRender, a module that improves 3D occupancy learning by enforcing projective consistency. Our key idea is to project both predicted and ground-truth 3D occupancy into 2D camera views, where we apply supervision. Our method penalizes 3D configurations that produce inconsistent 2D projections, thereby enforcing a more coherent 3D structure. To achieve this efficiently, we leverage differentiable rendering with Gaussian splatting. GaussRender seamlessly integrates with existing architectures while maintaining efficiency and requiring no inference-time modifications. Extensive evaluations on multiple benchmarks (SurroundOcc-nuScenes, Occ3D-nuScenes, SSCBench-KITTI360) demonstrate that GaussRender significantly improves geometric fidelity across various 3D occupancy models (TPVFormer, SurroundOcc, Symphonies), achieving state-of-the-art results, particularly on surface-sensitive metrics. The code is open-sourced at https://github.com/valeoai/GaussRender. Loïck Chambon, Eloi Zablocki, Alexandre Boulch, Mickaël Chen, Matthieu Cord |
ICCV | 5 |
| 2025 | Analyzing Fine-Tuning Representation Shift for Multimodal LLMs Steering
Pegah Khayatan, Mustafa Shukor, Jayneel Parekh, Arnaud Dapogny, Matthieu Cord |
ICCV | 5 |
| 2025 | Scaling Laws for Native Multimodal ModelsabstractBuilding general-purpose models that can effectively perceive the world through multimodal signals has been a long-standing goal. Current approaches involve integrating separately pre-trained components, such as connecting vision encoders to LLMs and continuing multimodal training. While such approaches exhibit remarkable sample efficiency, it remains an open question whether such late-fusion architectures are inherently superior. In this work, we revisit the architectural design of native multimodal models (NMMs)-those trained from the ground up on all modalities-and conduct an extensive scaling laws study, spanning 457 trained models with different architectures and training mixtures. Our investigation reveals no inherent advantage to late-fusion architectures over early-fusion ones, which do not rely on image encoders or tokenizers. On the contrary, early-fusion exhibits stronger performance at lower parameter counts, is more efficient to train, and is easier to deploy. Motivated by the strong performance of the early-fusion architectures, we show that incorporating Mixture of Experts (MoEs) allows models to learn modality-specific weights, significantly benefiting performance. Mustafa Shukor, Enrico Fini, Victor G. T. da Costa, Matthieu Cord, Joshua M. Susskind, Alaaeldin El-Nouby |
ICCV | 4 |
| 2025 | ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger BridgeabstractDiffusion models break down the challenging task of generating data from high-dimensional distributions into a series of easier denoising steps. Inspired by this paradigm, we propose a novel approach that extends the diffusion framework into modality space, decomposing the complex task of RGB image generation into simpler, interpretable stages. Our method, termed {\papernameAbbrev}, cascades modality-specific models, each responsible for generating an intermediate representation, such as contours, palettes, and detailed textures, ultimately culminating in a high-quality RGB image.
Instead of relying on the naive LDM concatenation conditioning mechanism to connect the different stages together, we employ Schr\"odinger Bridge to determine the optimal transport between different modalities.
Although employing a cascaded pipeline introduces more stages, which could lead to a more complex architecture, each stage is meticulously formulated for efficiency and accuracy, surpassing Stable-Diffusion (LDM) performance.
Modality composition not only enhances overall performance but enables emerging proprieties such as consistent editing, interaction capabilities, high-level interpretability, and faster convergence and sampling rate.
Extensive experiments on diverse datasets, including LSUN-Churches, ImageNet, CelebHQ, and LAION-Art, demonstrate the efficacy of our approach, consistently outperforming state-of-the-art methods.
For instance, {\papernameAbbrev} achieves notable efficiency, matching LDM performance on LSUN-Churches while operating 2$\times$ faster with a 3$\times$ smaller architecture.
The project website is available at:
\href{https://toddlerdiffusion.github.io/website/}{$https://toddlerdiffusion.github.io/website/$} Eslam Mohamed Bakr, Liangbing Zhao, Vincent Tao Hu, Matthieu Cord, Patrick Pérez |
ICLR | 4 |
| 2025 | Halton Scheduler for Masked Generative Image TransformerabstractMasked Generative Image Transformers (MaskGIT) have emerged as a scalable
and efficient image generation framework, able to deliver high-quality visuals with
low inference costs. However, MaskGIT’s token unmasking scheduler, an essential
component of the framework, has not received the attention it deserves. We analyze
the sampling objective in MaskGIT, based on the mutual information between
tokens, and elucidate its shortcomings. We then propose a new sampling strategy
based on our Halton scheduler instead of the original Confidence scheduler. More
precisely, our method selects the token’s position according to a quasi-random,
low-discrepancy Halton sequence. Intuitively, that method spreads the tokens
spatially, progressively covering the image uniformly at each step. Our analysis
shows that it allows reducing non-recoverable sampling errors, leading to simpler
hyper-parameters tuning and better quality images. Our scheduler does not require
retraining or noise injection and may serve as a simple drop-in replacement for
the original sampling strategy. Evaluation of both class-to-image synthesis on
ImageNet and text-to-image generation on the COCO dataset demonstrates that the
Halton scheduler outperforms the Confidence scheduler quantitatively by reducing
the FID and qualitatively by generating more diverse and more detailed images.
Our code is at https://github.com/valeoai/Halton-MaskGIT. Victor Besnier, Mickaël Chen, David Hurych, Eduardo Valle, Matthieu Cord |
ICLR | 5 |
| 2025 | LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Models for Referring Expression ComprehensionabstractVision Language Models (VLMs) have demonstrated remarkable capabilities in various open-vocabulary tasks, yet their zero-shot performance lags behind task-specific fine-tuned models, particularly in complex tasks like Referring Expression Comprehension (REC). Fine-tuning usually requires ‘white-box’ access to the model’s architecture and weights, which is not always feasible due to proprietary or privacy concerns. In this work, we propose LLM-wrapper, a method for ‘black-box’ adaptation of VLMs for the REC task using Large Language Models (LLMs). LLM-wrapper capitalizes on the reasoning abilities of LLMs, improved with a light fine-tuning, to select the most relevant bounding box to match the referring expression, from candidates generated by a zero-shot black-box VLM. Our approach offers several advantages: it enables the adaptation of closed-source models without needing access to their internal workings, it is versatile and works with any VLM, transfers to new VLMs, and it allows for the adaptation of an ensemble of VLMs. We evaluate LLM-wrapper on multiple datasets using different VLMs and LLMs, demonstrating significant performance improvements and highlighting the versatility of our method. While LLM-wrapper is not meant to directly compete with standard white-box fine-tuning, it offers a practical and effective alternative for black-box VLM adaptation. The code will be open-sourced. Amaia Cardiel, Eloi Zablocki, Elias Ramzi, Oriane Siméoni, Matthieu Cord |
ICLR | 5 |
| 2025 | Annealed Winner-Takes-All for Motion ForecastingabstractIn autonomous driving, motion prediction aims at forecasting the future trajectories of nearby agents, helping the ego vehicle to anticipate behaviors and drive safely. A key challenge is generating a diverse set of future predictions, commonly addressed using data-driven models with Multiple Choice Learning (MCL) architectures and Winner-Takes-All (WTA) training objectives. However, these methods face initialization sensitivity and training instabilities. Additionally, to compensate for limited performance, some approaches rely on training with a large set of hypotheses, requiring a post-selection step during inference to significantly reduce the number of predictions. To tackle these issues, we take inspiration from annealed MCL, a recently introduced technique that improves the convergence properties of MCL methods through an annealed Winner-Takes-All loss (aWTA). In this paper, we demonstrate how the aWTA loss can be integrated with state-of-the-art motion forecasting models to enhance their performance using only a minimal set of hypotheses, eliminating the need for the cumbersome post-selection step. Our approach can be easily incorporated into any trajectory prediction model normally trained using WTA and yields significant improvements. To facilitate the application of our approach to future motion forecasting models, the code is made publicly available: https://github.com/valeoai/MF_aWTA. Victor Letzelter, Mickaël Chen, Eloi Zablocki, Matthieu Cord |
ICRA | 5 |
| 2025 | FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion ModelsabstractFoundation models have exhibited unprecedented capabilities across various domains and tasks. Models like CLIP bridge cross-modal representations, while text-to-image diffusion models excel in realistic image generation. While the complexity of these models makes retraining infeasible, their superior performance has driven research to explore how to efficiently use them for downstream tasks. Our work explores how to leverage these models for dense visual prediction tasks, specifically image segmentation. To avoid the annotation cost or training large diffusion models, we constrain our method to be zero-shot and training-free. Our pipeline, dubbed FreeSeg-Diff, uses open-source foundation models to perform open-vocabulary segmentation as follows: (a) retrieving image caption (via BLIP-2) and visual features (via Stable Diffusion), (b) clustering and binarizing features to form class-agnostic object masks, (c) mapping these masks to textual classes using CLIP with open vocabulary support, and (d) refining coarse masks. FreeSeg-Diff surpasses many training-based methods on Pascal VOC and COCO datasets and delivers competitive results against recent weakly-supervised segmentation approaches. We provide experiments demonstrating the superiority of diffusion model features over other pre-trained models. Project page: https://bcorrad.github.io/freesegdiff/. Barbara Toniella Corradini, Mustafa Shukor, Paul Couairon, Guillaume Couairon, Franco Scarselli, Matthieu Cord |
IJCNN | 6 |
| 2025 | JAFAR: Jack up Any Feature at Any ResolutionabstractFoundation Vision Encoders have become indispensable across a wide range of dense vision tasks. However, their operation at low spatial feature resolutions necessitates subsequent feature decompression to enable full-resolution processing. To address this limitation, we introduce JAFAR, a lightweight and flexible feature upsampler designed to enhance the spatial resolution of visual features from any Foundation Vision Encoder to any target resolution. JAFAR features an attention-based upsampling module that aligns the spatial representations of high-resolution queries with semantically enriched low-resolution keys via Spatial Feature Transform modulation. Despite the absence of high-resolution feature ground truth; we find that learning at low upsampling ratios and resolutions generalizes surprisingly well to much higher scales. Extensive experiments demonstrate that JAFAR recovers intricate pixel-level details and consistently outperforms existing feature upsampling techniques across a diverse set of dense downstream applications. Paul Couairon, Loïck Chambon, Louis Serrano, Jean-Emmanuel Haugeard, Matthieu Cord, Nicolas Thome |
NeurIPS | 5 |
| 2025 | Learning to Steer: Input-dependent Steering for Multimodal LLMsabstractSteering has emerged as a practical approach to enable post-hoc guidance of LLMs towards enforcing a specific behavior.
However, it remains largely underexplored for multimodal LLMs (MLLMs); furthermore, existing steering techniques, such as \textit{mean} steering, rely on a single steering vector, applied independently of the input query. This paradigm faces limitations when the desired behavior is dependent on the example at hand. For example, a safe answer may consist in abstaining from answering when asked for an illegal activity, or may point to external resources or consultation with an expert when asked about medical advice. In this paper, we investigate a fine-grained steering that uses an input-specific linear shift. This shift is computed using contrastive input-specific prompting. However, the input-specific prompts required for this approach are not known at test time. Therefore, we propose to train a small auxiliary module to predict the input-specific steering vector. Our approach, dubbed as L2S (Learn-to-Steer), demonstrates that it reduces hallucinations and enforces safety in MLLMs, outperforming other static baselines. We will open-source our code. Jayneel Parekh, Pegah Khayatan, Mustafa Shukor, Arnaud Dapogny, Alasdair Newson, Matthieu Cord |
NeurIPS | 6 |
| 2025 | Manipulating Trajectory Prediction Models With BackdoorsabstractAutonomous vehicles depend on accurate trajectory prediction to navigate safely in complex traffic. Yet current models are vulnerable to stealthy backdoor attacks: an adversary embeds subtle, physically plausible triggers during training that remain latent until activated. To address this risk, we introduce a structured framework categorizing four trigger types—spatial, kinetic (braking), coordinated, and composite—and demonstrate on two benchmarks (nuScenes and Argoverse 2) and two state-of-the-art architectures (Autobot and Wayformer) that poisoning as little as 5% of training samples can reliably hijack future predictions. We further propose a real-time defense leveraging social attention: by encoding agent histories, computing cross-attention to the target vehicle, and filtering out agents with anomalously high weights, our method neutralizes backdoor triggers without degrading clean-data accuracy. Comprehensive experiments show our defense reduces attack success rates across diverse urban scenarios—intersections, roundabouts, multi-lane roads—highlighting both the severity of backdoor threats and a promising pathway to secure trajectory predictors in autonomous driving systems. Kaouther Messaoud, Kathrin Grosse, Mickaël Chen, Matthieu Cord, Patrick Pérez, Alexandre Alahi |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | PointBeV: A Sparse Approach to BeV PredictionsabstractBird's-eye View (BeV) representations have emerged as the de-facto shared space in driving applications, offering a unified space for sensor data fusion and supporting various downstream tasks. However, conventional models use grids with fixed resolution and range and face computational inefficiencies due to the uniform allocation of resources across all cells. To address this, we propose Point-BeV, a novel sparse BeV segmentation model operating on sparse BeV cells instead of dense grids. This approach offers precise control over memory usage, enabling the use of long temporal contexts and accommodating memory-constrained platforms. PointBeV employs an efficient two-pass strategy for training, enabling focused computation on regions of interest. At inference time, it can be used with various memory/performance trade-offs and flexibly adjusts to new specific use cases. PointBeV achieves state-of-the-art results on the nuScenes dataset for vehicle, pedes-trian, and lane segmentation, showcasing superior performance in static and temporal settings despite being trained solely with sparse signals. We release our code with two new efficient modules used in the architecture: Sparse Feature Pulling, designed for the effective extraction of features from images to BeV, and Submanifold Attention, which en-ables efficient temporal modeling. The code is available at https://github.com/valeoai/PointBeV. Loïck Chambon, Eloi Zablocki, Mickaël Chen, Florent Bartoccioni, Patrick Pérez, Matthieu Cord |
CVPR | 6 |
| 2024 | UniTraj: A Unified Framework for Scalable Vehicle Trajectory Prediction
Lan Feng, Mohammadhossein Bahari, Kaouther Messaoud, Eloi Zablocki, Matthieu Cord, Alexandre Alahi |
ECCV (12) | 5 |
| 2024 | Reliability in Semantic Segmentation: Can We Use Synthetic Data?
Thibaut Loiseau, Mickaël Chen, Patrick Pérez, Matthieu Cord |
ECCV (23) | 5 |
| 2024 | Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learningabstractFollowing the success of Large Language Models (LLMs), Large Multimodal Models (LMMs), such as the Flamingo model and its subsequent competitors, have started to emerge as natural steps towards generalist agents. However, interacting with recent LMMs reveals major limitations that are hardly captured by the current evaluation benchmarks. Indeed, task performances (e.g., VQA accuracy) alone do not provide enough clues to understand their real capabilities, limitations, and to which extent such models are aligned to human expectations. To refine our understanding of those flaws, we deviate from the current evaluation paradigm, and (1) evaluate 10 recent open-source LMMs from 3B up to 80B parameter scale, on 5 different axes; hallucinations, abstention, compositionality, explainability and instruction following. Our evaluation on these axes reveals major flaws in LMMs. While the current go-to solution to align these models is based on training, such as instruction tuning or RLHF, we rather (2) explore the training-free in-context learning (ICL) as a solution, and study how it affects these limitations. Based on our ICL study, (3) we push ICL further and propose new multimodal ICL variants such as; Multitask-ICL, Chain-of-Hindsight-ICL, and Self-Correcting-ICL. Our findings are as follows; (1) Despite their success, LMMs have flaws that remain unsolved with scaling alone. (2) The effect of ICL on LMMs flaws is nuanced; despite its effectiveness for improved explainability, answer abstention, ICL only slightly improves instruction following, does not improve compositional abilities, and actually even amplifies hallucinations. (3) The proposed ICL variants are promising as post-hoc approaches to efficiently tackle some of those flaws. The code is available here: https://github.com/mshukor/EvALign-ICL. Mustafa Shukor, Alexandre Ramé, Corentin Dancette, Matthieu Cord |
ICLR | 4 |
| 2024 | Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?abstractMotion forecasting is crucial in enabling autonomous vehicles to anticipate the future trajectories of surrounding agents. To do so, it requires solving mapping, detection, tracking, and then forecasting problems, in a multi-step pipeline. In this complex system, advances in conventional forecasting methods have been made using curated data, i.e., with the assumption of perfect maps, detection, and tracking. This paradigm, however, ignores any errors from upstream modules. Meanwhile, an emerging end-to-end paradigm, that tightly integrates the perception and forecasting architectures into joint training, promises to solve this issue. However, the evaluation protocols between the two methods were so far incompatible and their comparison was not possible. In fact, conventional forecasting methods are usually not trained nor tested in real-world pipelines (e.g., with upstream detection, tracking, and mapping modules). In this work, we aim to bring forecasting models closer to the real-world deployment. First, we propose a unified evaluation pipeline for forecasting methods with real-world perception inputs, allowing us to compare conventional and end-to-end methods for the first time. Second, our in-depth study uncovers a substantial performance gap when transitioning from curated to perception-based data. In particular, we show that this gap (1) stems not only from differences in precision but also from the nature of imperfect inputs provided by perception modules, and that (2) is not trivially reduced by simply finetuning on perception outputs. Based on extensive experiments, we provide recommendations for critical areas that require improvement and guidance towards more robust motion forecasting in the real world. The evaluation library for benchmarking models under standardized and practical conditions is provided: https://github.com/valeoai/MFEval. Loïck Chambon, Eloi Zablocki, Mickaël Chen, Alexandre Alahi, Matthieu Cord, Patrick Pérez |
ICRA | 6 |
| 2024 | DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized CutabstractFoundation models have emerged as powerful tools across various domains including language, vision, and multimodal tasks. While prior works have addressed unsupervised semantic segmentation, they significantly lag behind supervised models. In this paper, we use a diffusion UNet encoder as a foundation vision encoder and introduce DiffCut, an unsupervised zero-shot segmentation method that solely harnesses the output features from the final self-attention block. Through extensive experimentation, we demonstrate that using these diffusion features in a graph based segmentation algorithm, significantly outperforms previous state-of-the-art methods on zero-shot segmentation. Specifically, we leverage a recursive Normalized Cut algorithm that regulates the granularity of detected objects and produces well-defined segmentation maps that precisely capture intricate image details. Our work highlights the remarkably accurate semantic knowledge embedded within diffusion UNet encoders that could then serve as foundation vision encoders for downstream tasks. Paul Couairon, Mustafa Shukor, Jean-Emmanuel Haugeard, Matthieu Cord, Nicolas Thome |
NeurIPS | 4 |
| 2024 | What matters when building vision-language models?abstractThe growing interest in vision-language models (VLMs) has been driven by improvements in large language models and vision transformers. Despite the abundance of literature on this subject, we observe that critical decisions regarding the design of VLMs are often not justified. We argue that these unsupported decisions impede progress in the field by making it difficult to identify which choices improve model performance. To address this issue, we conduct extensive experiments around pre-trained models, architecture choice, data, and training methods. Our consolidation of findings includes the development of Idefics2, an efficient foundational VLM of 8 billion parameters. Idefics2 achieves state-of-the-art performance within its size category across various multimodal benchmarks, and is often on par with models four times its size. We release the model (base, instructed, and chat) along with the datasets created for its training. Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor Sanh |
NeurIPS | 3 |
| 2024 | A Concept-Based Explainability Framework for Large Multimodal ModelsabstractLarge multimodal models (LMMs) combine unimodal encoders and large language models (LLMs) to perform multimodal tasks. Despite recent advancements towards the interpretability of these models, understanding internal representations of LMMs remains largely a mystery. In this paper, we present a novel framework for the interpretation of LMMs. We propose a dictionary learning based approach, applied to the representation of tokens. The elements of the learned dictionary correspond to our proposed concepts. We show that these concepts are well semantically grounded in both vision and text. Thus we refer to these as ``multi-modal concepts''.
We qualitatively and quantitatively evaluate the results of the learnt concepts. We show that the extracted multimodal concepts are useful to interpret representations of test samples. Finally, we evaluate the disentanglement between different concepts and the quality of grounding concepts visually and textually. Our implementation is publicly available: https://github.com/mshukor/xl-vlms. Jayneel Parekh, Pegah Khayatan, Mustafa Shukor, Alasdair Newson, Matthieu Cord |
NeurIPS | 5 |
| 2024 | ManiPose: Manifold-Constrained Multi-Hypothesis 3D Human Pose EstimationabstractWe propose ManiPose, a manifold-constrained multi-hypothesis model for human-pose 2D-to-3D lifting. We provide theoretical and empirical evidence that, due to the depth ambiguity inherent to monocular 3D human pose estimation, traditional regression models suffer from pose-topology consistency issues, which standard evaluation metrics (MPJPE, P-MPJPE and PCK) fail to assess. ManiPose addresses depth ambiguity by proposing multiple candidate 3D poses for each 2D input, each with its estimated plausibility. Unlike previous multi-hypothesis approaches, ManiPose forgoes generative models, greatly facilitating its training and usage. By constraining the outputs to lie on the human pose manifold, ManiPose guarantees the consistency of all hypothetical poses, in contrast to previous works. We showcase the performance of ManiPose on real-world datasets, where it outperforms state-of-the-art models in pose consistency by a large margin while being very competitive on the MPJPE metric. Cédric Rommel, Victor Letzelter, Nermin Samet, Renaud Marlet, Matthieu Cord, Patrick Pérez, Eduardo Valle |
NeurIPS | 5 |
| 2024 | Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal InputsabstractLarge Language Models (LLMs) have demonstrated impressive performance on multimodal tasks, without any multimodal finetuning. They are the de facto building block for Large Multimodal Models (LMMs), yet, we still lack a proper understanding of their success. In this work, we expose frozen LLMs to image, video, audio and text inputs and analyse their internal representation with the attempt to understand their generalization beyond textual inputs. Our work provides the following **findings.** Perceptual tokens (1) are easily distinguishable from textual ones inside LLMs, with significantly different representations (e.g. live in different narrow cones), and complete translation to textual tokens does not exists. Yet, (2) both perceptual and textual tokens activate similar LLM weights. Despite their differences, (3) perceptual tokens are implicitly aligned to textual tokens inside LLMs, we call this the implicit multimodal alignment effect (IMA), and argue that this is linked to architectural design, helping LLMs to generalize. This provide more evidence to believe that the generalization of LLMs to multimodal inputs is mainly due to their architecture. These findings lead to several **implications.** This work provides several implications. (1) We find a positive correlation between the implicit alignment score and the task performance, suggesting that this could act as a proxy metric for model evaluation and selection. (2) A negative correlation exists regarding hallucinations (e.g. describing non-existing objects in images), revealing that this problem is mainly due to misalignment between the internal perceptual and textual representations. (3) Perceptual tokens change slightly throughout the model, thus, we propose different approaches to skip computations (e.g. in FFN layers), and significantly reduce the inference cost. (4) Due to the slowly changing embeddings across layers, and the high overlap between textual and multimodal activated weights, we compress LLMs by keeping only 1 subnetwork (called alpha-SubNet) that works well across a wide range of multimodal tasks. The code is available here: https://github.com/mshukor/ima-lmms. Mustafa Shukor, Matthieu Cord |
NeurIPS | 2 |
| 2024 | GradPaint: Gradient-guided inpainting with diffusion modelsabstractDenoising Diffusion Probabilistic Models (DDPMs) have recently achieved remarkable results in conditional and unconditional image generation. The pre-trained models can be adapted without further training to different downstream tasks, by guiding their iterative denoising process at inference time to satisfy additional constraints. For the specific task of image inpainting, the current guiding mechanism relies on copying-and-pasting the known regions from the input image at each denoising step. However, diffusion models are strongly conditioned by the initial random noise, and therefore struggle to harmonize predictions inside the inpainting mask with the real parts of the input image, often producing results with unnatural artifacts. Our method, dubbed GradPaint, steers the generation towards a globally coherent image. At each step in the denoising process, we leverage the model’s “denoised image estimation” by calculating a custom loss measuring its coherence with the masked input image. Our guiding mechanism uses the gradient obtained from backpropagating this loss through the diffusion model itself. GradPaint generalizes well to diffusion models trained on various datasets, improving upon current state-of-the-art supervised and unsupervised methods. Our code will be made available upon publication. Asya Grechka, Guillaume Couairon, Matthieu Cord |
Comput. Vis. Image Underst. | 3 |
| 2024 | Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval
Mustafa Shukor, Nicolas Thome, Matthieu Cord |
Comput. Vis. Image Underst. | 3 |
| 2024 | Semantic augmentation by mixing contents for semi-supervised learning
Rémy Sun, Clément Masson, Gilles Hénaff, Nicolas Thome, Matthieu Cord |
Pattern Recognit. | 5 |
| 2023 | CoMFormer: Continual Learning in Semantic and Panoptic SegmentationabstractContinual learning for segmentation has recently seen increasing interest. However, all previous works focus on narrow semantic segmentation and disregard panoptic segmentation, an important task with real-world impacts. In this paper, we present the first continual learning model capable of operating on both semantic and panoptic segmentation. Inspired by recent transformer approaches that consider segmentation as a mask-classification problem, we design CoMFormer. Our method carefully exploits the properties of transformer architectures to learn new classes over time. Specifically, we propose a novel adaptive distillation loss along with a mask-based pseudo-labeling technique to effectively prevent forgetting. To evaluate our approach, we introduce a novel continual panoptic segmentation benchmark on the challenging ADE20K dataset. Our CoMFormer outperforms all the existing baselines by forgetting less old classes but also learning more effectively new classes. In addition, we also report an extensive evaluation in the large-scale continual semantic segmentation scenario showing that CoMFormer also significantly outperforms state-of-the-art methods.11https://github.com/fcdl94/CoMFormer Fabio Cermelli, Matthieu Cord, Arthur Douillard |
CVPR | 2 |
| 2023 | Improving Selective Visual Question Answering by Learning from Your PeersabstractDespite advances in Visual Question Answering (VQA), the ability of models to assess their own correctness remains under-explored. Recent work has shown that VQA models, out-of-the-box, can have difficulties abstaining from answering when they are wrong. The option to abstain, also called Selective Prediction, is highly relevant when deploying systems to users who must trust the system's output (e.g., VQA assistants for users with visual impairments). For such scenarios, abstention can be especially important as users may provide out-of-distribution (OOD) or adversarial inputs that make incorrect answers more likely. In this work, we explore Selective VQA in both in-distribution (ID) and OOD scenarios, where models are presented with mixtures of ID and OOD data. The goal is to maximize the number of questions answered while minimizing the risk of error on those questions. We propose a simple yet effective Learning from Your Peers (LYP) approach for training multimodal selection functions for making abstention decisions. Our approach uses predictions from models trained on distinct subsets of the training data as targets for optimizing a Selective VQA model. It does not require additional manual labels or held-out data and provides a signal for identifying examples that are easy/difficult to generalize to. In our extensive evaluations, we show this benefits a number of models across different architectures and scales. Overall, for ID, we reach 32.92% in the selective prediction metric coverage at 1 % risk of error$(\mathcal{C} {@} 1\%)$which doubles the previous best coverage of 15.79% on this task. For mixed ID/OOD, using models' softmax confidences for abstention decisions performs very poorly, answering$\mathcal{C}$@1%. Corentin Dancette, Spencer Whitehead, Rishabh Maheshwary, Ramakrishna Vedantam, Stefan Scherer, Xinlei Chen, Matthieu Cord, Marcus Rohrbach |
CVPR | 7 |
| 2023 | Co-training 2L Submodels for Visual RecognitionabstractWe introduce submodel co-training, a regularization method related to co-training, self-distillation and stochastic depth. Given a neural network to be trained, for each sample we implicitly instantiate two altered networks, “submodels”, with stochastic depth: we activate only a subset of the layers. Each network serves as a soft teacher to the other, by providing a loss that complements the regular loss provided by the one-hot label. Our approach, dubbed “co-sub”, uses a single set of weights, and does not involve a pre-trained external model or temporal averaging. Experimentally, we show that submodel co-training is effective to train backbones for recognition tasks such as image classification and semantic segmentation. Our approach is compatible with multiple architectures, including RegNet, ViT, PiT, XCiT, Swin and ConvNext. Our training strategy improves their results in comparable settings. For instance, a ViT-B pretrained with cosub on ImageNet-21k obtains 87.4% top1 acc. @448 on ImageNet-val. Hugo Touvron, Matthieu Cord, Maxime Oquab, Piotr Bojanowski, Jakob Verbeek, Hervé Jégou |
CVPR | 2 |
| 2023 | OCTET: Object-aware Counterfactual ExplanationsabstractNowadays, deep vision models are being widely deployed in safety-critical applications, e.g., autonomous driving, and explainability of such models is becoming a pressing concern. Among explanation methods, counter-factual explanations aim to find minimal and interpretable changes to the input image that would also change the output of the model to be explained. Such explanations point end-users at the main factors that impact the decision of the model. However, previous methods struggle to explain decision models trained on images with many objects, e.g., urban scenes, which are more difficult to work with but also arguably more critical to explain. In this work, we propose to tackle this issue with an object-centric framework for counterfactual explanation generation. Our method, inspired by recent generative modeling works, encodes the query image into a latent space that is structured in a way to ease object-level manipulations. Doing so, it provides the end-user with control over which search directions (e.g., spatial displacement of objects, style modification, etc.) are to be explored during the counterfactual generation. We conduct a set of experiments on counterfactual explanation benchmarks for driving scenes, and we show that our method can be adapted beyond classification, e.g., to explain semantic segmentation models. To complete our analysis, we design and run a user study that measures the usefulness of counterfactual explanations in understanding a decision model. Code is available at https://github.com/valeoai/OCTET. Mehdi Zemni, Mickaël Chen, Eloi Zablocki, Hédi Ben-Younes, Patrick Pérez, Matthieu Cord |
CVPR | 6 |
| 2023 | Zero-shot spatial layout conditioning for text-to-image diffusion modelsabstractLarge-scale text-to-image diffusion models have significantly improved the state of the art in generative image modeling and allow for an intuitive and powerful user interface to drive the image generation process. Expressing spatial constraints, e.g. to position specific objects in particular locations, is cumbersome using text; and current text-based image generation models are not able to accurately follow such instructions. In this paper we consider image generation from text associated with segments on the image canvas, which combines an intuitive natural language interface with precise spatial control over the generated content. We propose ZestGuide, a "zero-shot" segmentation guidance approach that can be plugged into pre-trained text-to-image diffusion models, and does not require any additional training. It leverages implicit segmentation maps that can be extracted from cross-attention layers, and uses them to align the generation with input masks. Our experimental results combine high image quality with accurate alignment of generated content with input segmentations, and improve over prior work both quantitatively and qualitatively, including methods that require training on images with corresponding segmentations. Compared to Paint with Words, the previous state-of-the art in image generation with zero-shot segmentation conditioning, we improve by 5 to 10 mIoU points on the COCO dataset with similar FID scores. Guillaume Couairon, Marlène Careil, Matthieu Cord, Stéphane Lathuilière, Jakob Verbeek |
ICCV | 3 |
| 2023 | eP-ALM: Efficient Perceptual Augmentation of Language ModelsabstractLarge Language Models (LLMs) have so far impressed the world, with unprecedented capabilities that emerge in models at large scales. On the vision side, transformer models (i.e., ViT) are following the same trend, achieving the best performance on challenging benchmarks. With the abundance of such unimodal models, a natural question arises; do we need also to follow this trend to tackle multimodal tasks? In this work, we propose to rather direct effort to efficient adaptations of existing models, and propose to augment Language Models with perception. Existing approaches for adapting pretrained models for vision-language tasks still rely on several key components that hinder their efficiency. In particular, they still train a large number of parameters, rely on large multimodal pretraining, use encoders (e.g., CLIP) trained on huge image-text datasets, and add significant inference overhead. In addition, most of these approaches have focused on Zero-Shot and In Context Learning, with little to no effort on direct finetuning. We investigate the minimal computational effort needed to adapt unimodal models for multimodal tasks and propose a new challenging setup, alongside different approaches, that efficiently adapts unimodal pretrained models. We show that by freezing more than 99% of total parameters, training only one linear projection layer, and prepending only one trainable token, our approach (dubbed eP-ALM) significantly outperforms other baselines on VQA and Captioning across Image, Video, and Audio modalities, following the proposed setup. The code is available here: https://github.com/mshukor/eP-ALM. Mustafa Shukor, Corentin Dancette, Matthieu Cord |
ICCV | 3 |
| 2023 | DiffEdit: Diffusion-based semantic image editing with mask guidance
Guillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu Cord |
ICLR | 4 |
| 2023 | PowerQuant: Automorphism Search for Non-Uniform Quantization
Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
ICLR | 3 |
| 2023 | Model Ratatouille: Recycling Diverse Models for Out-of-Distribution GeneralizationabstractFoundation models are redefining how AI systems are built. Practitioners now follow a standard procedure to build their machine learning solutions: from a pre-trained foundation model, they fine-tune the weights on the target task of interest. So, the Internet is swarmed by a handful of foundation models fine-tuned on many diverse tasks: these individual fine-tunings exist in isolation without benefiting from each other. In our opinion, this is a missed opportunity, as these specialized models contain rich and diverse features. In this paper, we thus propose model ratatouille, a new strategy to recycle the multiple fine-tunings of the same foundation model on diverse auxiliary tasks. Specifically, we repurpose these auxiliary weights as initializations for multiple parallel fine-tunings on the target task; then, we average all fine-tuned weights to obtain the final model. This recycling strategy aims at maximizing the diversity in weights by leveraging the diversity in auxiliary tasks. Empirically, it improves the state of the art on the reference DomainBed benchmark for out-of-distribution generalization. Looking forward, this work contributes to the emerging paradigm of updatable machine learning where, akin to open-source software development, the community collaborates to reliably update machine learning models. Alexandre Ramé, Kartik Ahuja, Matthieu Cord, Léon Bottou, David Lopez-Paz |
ICML | 4 |
| 2023 | OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text DocumentsabstractLarge multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks. However, the datasets used to train these models have not been released, and the collection process has not been fully specified. We introduce the OBELICS dataset, an open web-scale filtered dataset of interleaved image-text documents comprising 141 million web pages extracted from Common Crawl, 353 million associated images, and 115 billion text tokens. We describe the dataset creation process, present comprehensive filtering rules, and provide an analysis of the dataset's content. To show the viability of OBELICS, we train on the dataset vision and language models of 9 and 80 billion parameters, IDEFICS-9B and IDEFICS, and obtain competitive performance on different multimodal benchmarks. We release our dataset, models and code. Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M. Rush, Douwe Kiela, Matthieu Cord, Victor Sanh |
NeurIPS | 11 |
| 2023 | Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsabstractFoundation models are first pre-trained on vast unsupervised datasets and then fine-tuned on labeled data. Reinforcement learning, notably from human feedback (RLHF), can further align the network with the intended usage. Yet the imperfections in the proxy reward may hinder the training and lead to suboptimal results; the diversity of objectives in real-world tasks and human opinions exacerbate the issue. This paper proposes embracing the heterogeneity of diverse rewards by following a multi-policy strategy. Rather than focusing on a single a priori reward, we aim for Pareto-optimal generalization across the entire space of preferences. To this end, we propose rewarded soup, first specializing multiple networks independently (one for each proxy reward) and then interpolating their weights linearly. This succeeds empirically because we show that the weights remain linearly connected when fine-tuned on diverse rewards from a shared pre-trained initialization. We demonstrate the effectiveness of our approach for text-to-text (summarization, Q&A, helpful assistant, review), text-image (image captioning, text-to-image generation, visual grounding), and control (locomotion) tasks. We hope to enhance the alignment of deep models, and how they interact with the world in all its diversity. Alexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya, Mustafa Shukor, Laure Soulier, Matthieu Cord |
NeurIPS | 7 |
| 2023 | REx: Data-Free Residual Quantization Error ExpansionabstractDeep neural networks (DNNs) are ubiquitous in computer vision and natural language processing, but suffer from high inference cost. This problem can be addressed by quantization, which consists in converting floating point operations into a lower bit-width format. With the growing concerns on privacy rights, we focus our efforts on data-free methods. However, such techniques suffer from their lack of adaptability to the target devices, as a hardware typically only supports specific bit widths. Thus, to adapt to a variety of devices, a quantization method shall be flexible enough to find good accuracy v.s. speed trade-offs for every bit width and target device. To achieve this, we propose REx, a quantization method that leverages residual error expansion, along with group sparsity.
We show experimentally that REx enables better trade-offs (in terms of accuracy given any target bit-width) on both convnets and transformers for computer vision, as well as NLP models. In particular, when applied to large language models, we show that REx elegantly solves the outlier problem that hinders state-of-the-art quantization methods.
In addition, REx is backed off by strong theoretical guarantees on the preservation of the predictive function of the original model. Lastly, we show that REx is agnostic to the quantization operator and can be used in combination with previous quantization work. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
NeurIPS | 3 |
| 2023 | SPIQ: Data-Free Per-Channel Static Input QuantizationabstractComputationally expensive neural networks are ubiquitous in computer vision and solutions for efficient inference have drawn a growing attention in the machine learning community. Examples of such solutions comprise quantization, i.e. converting the processing values (weights and inputs) from floating point into integers e.g. int8 or int4. Concurrently, the rise of privacy concerns motivated the study of less invasive acceleration methods, such as data-free quantization of pre-trained models weights and activations. Previous approaches either exploit statistical information to deduce scalar ranges and scaling factors for the activations in a static manner, or dynamically adapt this range on-the-fly for each input of each layer (also referred to as activations): the latter generally being more accurate at the expense of significantly slower inference. In this work, we argue that static input quantization can reach the accuracy levels of dynamic methods by means of a per-channel input quantization scheme that allows one to more finely preserve cross-channel dynamics. We show through a thorough empirical evaluation on multiple computer vision problems (e.g. ImageNet classification, Pascal VOC object detection as well as CityScapes semantic segmentation) that the proposed method, dubbed SPIQ, achieves accuracies rivalling dynamic approaches with static-level inference speed, significantly outperforming state-of-the-art quantization methods on every benchmark. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
WACV | 3 |
| 2023 | LiDARTouch: Monocular metric depth estimation with a few-beam LiDAR
Florent Bartoccioni, Eloi Zablocki, Patrick Pérez, Matthieu Cord, Karteek Alahari |
Comput. Vis. Image Underst. | 4 |
| 2023 | ResMLP: Feedforward Networks for Image Classification With Data-Efficient TrainingabstractWe present ResMLP, an architecture built entirely upon multi-layer perceptrons for image classification. It is a simple residual network that alternates (i) a linear layer in which image patches interact, independently and identically across channels, and (ii) a two-layer feed-forward network in which channels interact independently per patch. When trained with a modern training strategy using heavy data-augmentation and optionally distillation, it attains surprisingly good accuracy/complexity trade-offs on ImageNet. We also train ResMLP models in a self-supervised setup, to further remove priors from employing a labelled dataset. Finally, by adapting our model to machine translation we achieve surprisingly good results. We share pre-trained models and our code based on the Timm library. Hugo Touvron, Piotr Bojanowski, Mathilde Caron, Matthieu Cord, Alaaeldin El-Nouby, Edouard Grave, Gautier Izacard, Armand Joulin, Gabriel Synnaeve, Jakob Verbeek, Hervé Jégou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | RED++ : Data-Free Pruning of Deep Neural Networks via Input Splitting and Output MergingabstractPruning Deep Neural Networks (DNNs) is a prominent field of study in the goal of inference runtime acceleration. In this paper, we introduce a novel data-free pruning protocol RED++. Only requiring a trained neural network, and not specific to any particular DNN, we exploit an adaptive data-free scalar hashing which exhibits redundancies among neuron weight values. We study the theoretical and empirical guarantees on the preservation of the accuracy from the hashing as well as the expected pruning ratio resulting from the exploitation of said redundancies. We propose a novel data-free pruning technique of DNN layers which removes the input-wise redundant operations. This algorithm is straightforward, parallelizable and offers novel perspective on DNN pruning by shifting the burden of large computation to efficient memory access and allocation. We provide theoretical guarantees on RED++ performance and empirically demonstrate its superiority over other data-free pruning methods and its competitiveness with data-driven ones on ResNets, MobileNets, and EfficientNets. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Efficient Vision-Language Pretraining with Visual Concepts and Hierarchical Alignment
Mustafa Shukor, Guillaume Couairon, Matthieu Cord |
BMVC | 3 |
| 2022 | FlexIT: Towards Flexible Semantic Image TranslationabstractDeep generative models, like GANs, have considerably improved the state of the art in image synthesis, and are able to generate near photo-realistic images in structured domains such as human faces. Based on this success, recent work on image editing proceeds by projecting images to the GAN latent space and manipulating the latent vector. However, these approaches are limited in that only images from a narrow domain can be transformed, and with only a limited number of editing operations. We propose FlexIT, a novel method which can take any input image and a user-defined text instruction for editing. Our method achieves flexible and natural editing, pushing the limits of semantic image translation. First, FlexIT combines the input image and text into a single target point in the CLIP multimodal embedding space. Via the latent space of an autoencoder, we iteratively transform the input image toward the target point, ensuring coherence and quality with a variety of novel regularization terms. We propose an evaluation protocol for semantic image translation, and thoroughly evaluate our method on ImageNet. Code will be available at https://github.com/facebookresearch/SemanticImageTranslation/. Guillaume Couairon, Asya Grechka, Jakob Verbeek, Holger Schwenk, Matthieu Cord |
CVPR | 5 |
| 2022 | DyTox: Transformers for Continual Learning with DYnamic TOken eXpansionabstractDeep network architectures struggle to continually learn new tasks without forgetting the previous tasks. A recent trend indicates that dynamic architectures based on an ex-pansion of the parameters can reduce catastrophic forget-ting efficiently in continual learning. However, existing approaches often require a task identifier at test-time, need complex tuning to balance the growing number of parameters, and barely share any information across tasks. As a result, they struggle to scale to a large number of tasks without significant overhead. In this paper, we propose a transformer architecture based on a dedicated encoder/decoder framework. Critically, the encoder and decoder are shared among all tasks. Through a dynamic expansion of special tokens, we specialize each forward of our decoder network on a task distribution. Our strategy scales to a large number of tasks while having neg-ligible memory and time overheads due to strict control of the expansion of the parameters. Moreover, this efficient strategy doesn't need any hyperparameter tuning to control the network's expansion. Our model reaches excellent results on CIFAR100 and state-of-the-art performances on the large-scale ImageNet100 and ImageNet100 while having fewer parameters than concurrent dynamic frameworks.11Code is released at https://github.com/arthurdouillard/dytox. Arthur Douillard, Alexandre Ramé, Guillaume Couairon, Matthieu Cord |
CVPR | 4 |
| 2022 | STEEX: Steering Counterfactual Explanations with Semantics
Paul Jacob, Eloi Zablocki, Hédi Ben-Younes, Mickaël Chen, Patrick Pérez, Matthieu Cord |
ECCV (12) | 6 |
| 2022 | Three Things Everyone Should Know About Vision Transformers
Hugo Touvron, Matthieu Cord, Alaaeldin El-Nouby, Jakob Verbeek, Hervé Jégou |
ECCV (24) | 2 |
| 2022 | DeiT III: Revenge of the ViT
Hugo Touvron, Matthieu Cord, Hervé Jégou |
ECCV (24) | 2 |
| 2022 | Fishr: Invariant Gradient Variances for Out-of-Distribution GeneralizationabstractLearning robust models that generalize well under changes in the data distribution is critical for real-world applications. To this end, there has been a growing surge of interest to learn simultaneously from multiple training domains - while enforcing different types of invariance across those domains. Yet, all existing approaches fail to show systematic benefits under controlled evaluation protocols. In this paper, we introduce a new regularization - named Fishr - that enforces domain invariance in the space of the gradients of the loss: specifically, the domain-level variances of gradients are matched across training domains. Our approach is based on the close relations between the gradient covariance, the Fisher Information and the Hessian of the loss: in particular, we show that Fishr eventually aligns the domain-level loss landscapes locally around the final weights. Extensive experiments demonstrate the effectiveness of Fishr for out-of-distribution generalization. Notably, Fishr improves the state of the art on the DomainBed benchmark and performs consistently better than Empirical Risk Minimization. Our code is available at https://github.com/alexrame/fishr. Alexandre Ramé, Corentin Dancette, Matthieu Cord |
ICML | 3 |
| 2022 | Swapping Semantic Contents for Mixing ImagesabstractDeep architecture have proven capable of solving many tasks provided a sufficient amount of labeled data. In fact, the amount of available labeled data has become the principal bottleneck in low label settings such as Semi-Supervised Learning. Mixing Data Augmentations do not typically yield new labeled samples, as indiscriminately mixing contents creates between-class samples. In this work, we introduce the SciMix framework that can learn to replace the global semantic content from one sample. By teaching a StyleGan generator to embed a semantic style code into image backgrounds, we obtain new mixing scheme for data augmentation. We then demonstrate that SciMix yields novel mixed samples that inherit many characteristics from their non-semantic parents. Afterwards, we verify those samples can be used to improve the performance semi-supervised frameworks like Mean Teacher or Fixmatch, and even fully supervised learning on a small labeled dataset. Rémy Sun, Clément Masson, Gilles Hénaff, Nicolas Thome, Matthieu Cord |
ICPR | 5 |
| 2022 | Diverse Weight Averaging for Out-of-Distribution GeneralizationabstractStandard neural networks struggle to generalize under distribution shifts in computer vision. Fortunately, combining multiple networks can consistently improve out-of-distribution generalization. In particular, weight averaging (WA) strategies were shown to perform best on the competitive DomainBed benchmark; they directly average the weights of multiple networks despite their nonlinearities. In this paper, we propose Diverse Weight Averaging (DiWA), a new WA strategy whose main motivation is to increase the functional diversity across averaged models. To this end, DiWA averages weights obtained from several independent training runs: indeed, models obtained from different runs are more diverse than those collected along a single run thanks to differences in hyperparameters and training procedures. We motivate the need for diversity by a new bias-variance-covariance-locality decomposition of the expected error, exploiting similarities between WA and standard functional ensembling. Moreover, this decomposition highlights that WA succeeds when the variance term dominates, which we show occurs when the marginal distribution changes at test time. Experimentally, DiWA consistently improves the state of the art on DomainBed without inference overhead. Alexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy, Patrick Gallinari, Matthieu Cord |
NeurIPS | 6 |
| 2022 | SInGE: Sparsity via Integrated Gradients Estimation of Neuron RelevanceabstractThe leap in performance in state-of-the-art computer vision methods is attributed to the development of deep neural networks. However it often comes at a computational price which may hinder their deployment. To alleviate this limitation, structured pruning is a well known technique which consists in removing channels, neurons or filters, and is commonly applied in order to produce more compact models. In most cases, the computations to remove are selected based on a relative importance criterion. At the same time, the need for explainable predictive models has risen tremendously and motivated the development of robust attribution methods that highlight the relative importance of pixels of an input image or feature map. In this work, we discuss the limitations of existing pruning heuristics, among which magnitude and gradient-based methods. We draw inspiration from attribution methods to design a novel integrated gradient pruning criterion, in which the relevance of each neuron is defined as the integral of the gradient variation on a path towards this neuron removal. Furthermore, We propose an entwined DNN pruning and fine-tuning flowchart to better preserve DNN accuracy while removing parameters. We show through extensive validation on several datasets, architectures as well as pruning scenarios that the proposed method, dubbed SInGE, significantly outperforms existing state-of-the-art DNN pruning methods. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
NeurIPS | 3 |
| 2022 | Explainability of Deep Vision-Based Autonomous Driving Systems: Review and Challenges
Eloi Zablocki, Hédi Ben-Younes, Patrick Pérez, Matthieu Cord |
Int. J. Comput. Vis. | 4 |
| 2022 | Confidence Estimation via Auxiliary ModelsabstractReliably quantifying the confidence of deep neural classifiers is a challenging yet fundamental requirement for deploying such models in safety-critical applications. In this paper, we introduce a novel target criterion for model confidence, namely the true class probability (TCP). We show that TCP offers better properties for confidence estimation than standard maximum class probability (MCP). Since the true class is by essence unknown at test time, we propose to learn TCP criterion from data with an auxiliary model, introducing a specific learning scheme adapted to this context. We evaluate our approach on the task of failure prediction and of self-training with pseudo-labels for domain adaptation, which both necessitate effective confidence estimates. Extensive experiments are conducted for validating the relevance of the proposed approach in each task. We study various network architectures and experiment with small and large datasets for image classification and semantic segmentation. In every tested benchmark, our approach outperforms strong baselines. Charles Corbière, Nicolas Thome, Antoine Saporta, Matthieu Cord, Patrick Pérez |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Driving behavior explanation with multi-level fusion
Hédi Ben-Younes, Eloi Zablocki, Patrick Pérez, Matthieu Cord |
Pattern Recognit. | 4 |
| 2022 | Detecting 32 Pedestrian Attributes for Autonomous VehiclesabstractPedestrians are arguably one of the most safety-critical road users to consider for autonomous vehicles in urban areas. In this paper, we address the problem of jointly detecting pedestrians and recognizing 32 pedestrian attributes from a single image. These encompass visual appearance and behavior, and also include the forecasting of road crossing, which is a main safety concern. For this, we introduce a Multi-Task Learning (MTL) model relying on a composite field framework, which achieves both goals in an efficient way. Each field spatially locates pedestrian instances and aggregates attribute predictions over them. This formulation naturally leverages spatial context, making it well suited to low resolution scenarios such as autonomous driving. By increasing the number of attributes jointly learned, we highlight an issue related to the scales of gradients, which arises in MTL with numerous tasks. We solve it by normalizing the gradients coming from different objective functions when they join at the fork in the network architecture during the backward pass, referred to as fork-normalization. Experimental validation is performed on JAAD, a dataset providing numerous attributes for pedestrian analysis from autonomous vehicles, and shows competitive detection and attribute recognition results, as well as a more stable MTL training. Taylor Mordan, Matthieu Cord, Patrick Pérez, Alexandre Alahi |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | MAGECally invert images for realistic editing
Asya Grechka, Jean-François Goudou, Matthieu Cord |
BMVC | 3 |
| 2021 | PLOP: Learning Without Forgetting for Continual Semantic SegmentationabstractDeep learning approaches are nowadays ubiquitously used to tackle computer vision tasks such as semantic segmentation, requiring large datasets and substantial computational power. Continual learning for semantic segmentation (CSS) is an emerging trend that consists in updating an old model by sequentially adding new classes. However, continual learning methods are usually prone to catastrophic forgetting. This issue is further aggravated in CSS where, at each step, old classes from previous iterations are collapsed into the background. In this paper, we propose Local POD, a multi-scale pooling distillation scheme that preserves long- and short-range spatial relationships at feature level. Furthermore, we design an entropy-based pseudo-labelling of the background w.r.t. classes predicted by the old model to deal with background shift and avoid catastrophic forgetting of the old classes. Our approach, called PLOP, significantly outperforms state-of-the-art methods in existing CSS scenarios, as well as in newly proposed challenging benchmarks1. Arthur Douillard, Arnaud Dapogny, Matthieu Cord |
CVPR | 4 |
| 2021 | OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised LearningabstractLearning image representations without human supervision is an important and active research field. Several recent approaches have successfully leveraged the idea of making such a representation invariant under different types of perturbations, especially via contrastive-based instance discrimination training. Although effective visual representations should indeed exhibit such invariances, there are other important characteristics, such as encoding contextual reasoning skills, for which alternative reconstruction-based approaches might be better suited.With this in mind, we propose a teacher-student scheme to learn representations by training a convolutional net to reconstruct a bag-of-visual-words (BoW) representation of an image, given as input a perturbed version of that same image. Our strategy performs an online training of both the teacher network (whose role is to generate the BoW targets) and the student network (whose role is to learn representations), along with an online update of the visual-words vocabulary (used for the BoW targets). This idea effectively enables fully online BoW-guided unsupervised learning. Extensive experiments demonstrate the interest of our BoWbased strategy, which surpasses previous state-of-the-art methods (including contrastive-based ones) in several applications. For instance, in downstream tasks such Pascal object detection, Pascal classification and Places205 classification, our method improves over all prior unsupervised approaches, thus establishing new state-of-the-art results that are also significantly better even than those of supervised pre-training. We provide the implementation code at https://github.com/valeoai/obow. Spyros Gidaris, Andrei Bursuc, Gilles Puy, Nikos Komodakis, Matthieu Cord, Patrick Pérez |
CVPR | 5 |
| 2021 | Semantic Palette: Guiding Scene Generation With Class ProportionsabstractDespite the recent progress of generative adversarial networks (GANs) at synthesizing photo-realistic images, producing complex urban scenes remains a challenging problem. Previous works break down scene generation into two consecutive phases: unconditional semantic layout synthesis and image synthesis conditioned on layouts. In this work, we propose to condition layout generation as well for higher semantic control: given a vector of class proportions, we generate layouts with matching composition. To this end, we introduce a conditional framework with novel architecture designs and learning objectives, which effectively accommodates class proportions to guide the scene generation process. The proposed architecture also allows partial layout editing with interesting applications. Thanks to the semantic control, we can produce layouts close to the real distribution, helping enhance the whole scene generation process. On different metrics and urban scene benchmarks, our models outperform existing baselines. Moreover, we demonstrate the merit of our approach for data augmentation: semantic segmenters trained on real layout-image pairs along with additional ones generated by our approach outperform models only trained on real pairs. Guillaume Le Moing, Himalaya Jain, Patrick Pérez, Matthieu Cord |
CVPR | 5 |
| 2021 | Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question AnsweringabstractWe introduce an evaluation methodology for visual question answering (VQA) to better diagnose cases of shortcut learning. These cases happen when a model exploits spurious statistical regularities to produce correct answers but does not actually deploy the desired behavior. There is a need to identify possible shortcuts in a dataset and assess their use before deploying a model in the real world. The research community in VQA has focused exclusively on question-based shortcuts, where a model might, for example, answer "What is the color of the sky" with "blue" by relying mostly on the question-conditional training prior and give little weight to visual evidence. We go a step further and consider multimodal shortcuts that involve both questions and images. We first identify potential shortcuts in the popular VQA v2 training set by mining trivial predictive rules such as co-occurrences of words and visual elements. We then introduce VQA-CounterExamples (VQACE), an evaluation protocol based on our subset of CounterExamples i.e. image-question-answer triplets where our rules lead to incorrect answers. We use this new evaluation in a large-scale study of existing approaches for VQA. We demonstrate that even state-of-the-art models perform poorly and that existing techniques to reduce biases are largely ineffective in this context. Our findings suggest that past work on question-based biases in VQA has only addressed one facet of a complex issue. The code for our method is available at https://github.com/cdancette/detect-shortcuts Corentin Dancette, Rémi Cadène, Damien Teney, Matthieu Cord |
ICCV | 4 |
| 2021 | MixMo: Mixing Multiple Inputs for Multiple Outputs via Deep SubnetworksabstractRecent strategies achieved ensembling "for free" by fitting concurrently diverse subnetworks inside a single base network. The main idea during training is that each sub-network learns to classify only one of the multiple inputs simultaneously provided. However, the question of how to best mix these multiple inputs has not been studied so farIn this paper, we introduce MixMo, a new generalized framework for learning multi-input multi-output deep subnetworks. Our key motivation is to replace the suboptimal summing operation hidden in previous approaches by a more appropriate mixing mechanism. For that purpose, we draw inspiration from successful mixed sample data augmentations. We show that binary mixing in features - particularly with rectangular patches from CutMix - enhances results by making subnetworks stronger and more diverse.We improve state of the art for image classification on CIFAR-100 and Tiny ImageNet datasets. Our easy to implement models notably outperform data augmented deep ensembles, without the inference and memory overheads. As we operate in features and simply better leverage the expressiveness of large networks, we open a new line of research complementary to previous works. Alexandre Ramé, Rémy Sun, Matthieu Cord |
ICCV | 3 |
| 2021 | Multi-Target Adversarial Frameworks for Domain Adaptation in Semantic SegmentationabstractIn this work, we address the task of unsupervised domain adaptation (UDA) for semantic segmentation in presence of multiple target domains: The objective is to train a single model that can handle all these domains at test time. Such a multi-target adaptation is crucial for a variety of scenarios that real-world autonomous systems must handle. It is a challenging setup since one faces not only the domain gap between the labeled source set and the un-labeled target set, but also the distribution shifts existing within the latter among the different target domains. To this end, we introduce two adversarial frameworks: (i) multi-discriminator, which explicitly aligns each target domain to its counterparts, and (ii) multi-target knowledge transfer, which learns a target-agnostic model thanks to a multi-teacher/single-student distillation mechanism. The evaluation is done on four newly-proposed multi-target bench-marks for UDA in semantic segmentation. In all tested scenarios, our approaches consistently outperform baselines, setting competitive standards for the novel task. Antoine Saporta, Matthieu Cord, Patrick Pérez |
ICCV | 3 |
| 2021 | Going deeper with Image TransformersabstractTransformers have been recently adapted for large scale image classification, achieving high scores shaking up the long supremacy of convolutional neural networks. However the optimization of vision transformers has been little studied so far. In this work, we build and optimize deeper transformer networks for image classification. In particular, we investigate the interplay of architecture and optimization of such dedicated transformers. We make two architecture changes that significantly improve the accuracy of deep transformers. This leads us to produce models whose performance does not saturate early with more depth, for in-stance we obtain 86.5% top-1 accuracy on Imagenet when training with no external data, we thus attain the current sate of the art with less floating-point operations and parameters. Our best model establishes the new state of the art on Imagenet with Reassessed labels and Imagenet-V2 / match frequency, in the setting with no additional training data. We share our code and models1. Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, Hervé Jégou |
ICCV | 2 |
| 2021 | Grafit: Learning fine-grained image representations with coarse labelsabstractThis paper tackles the problem of learning a finer representation than the one provided by training labels. This enables fine-grained category retrieval of images in a collection annotated with coarse labels only.Our network is learned with a nearest-neighbor classifier objective, and an instance loss inspired by self-supervised learning. By jointly leveraging the coarse labels and the underlying fine-grained latent space, it significantly improves the accuracy of category-level retrieval methods.Our strategy outperforms all competing methods for retrieving or classifying images at a finer granularity than that available at train time. It also improves the accuracy for transfer learning tasks to fine-grained datasets. Hugo Touvron, Alexandre Sablayrolles, Matthijs Douze, Matthieu Cord, Hervé Jégou |
ICCV | 4 |
| 2021 | DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial Estimation
Alexandre Ramé, Matthieu Cord |
ICLR | 2 |
| 2021 | Training data-efficient image transformers & distillation through attentionabstractRecently, neural networks purely based on attention were shown to address image understanding tasks such as image classification. These high-performing vision transformers are pre-trained with hundreds of millions of images using a large infrastructure, thereby limiting their adoption. In this work, we produce competitive convolution-free transformers trained on ImageNet only using a single computer in less than 3 days. Our reference vision transformer (86M parameters) achieves top-1 accuracy of 83.1% (single-crop) on ImageNet with no external data. We also introduce a teacher-student strategy specific to transformers. It relies on a distillation token ensuring that the student learns from the teacher through attention, typically from a convnet teacher. The learned transformers are competitive (85.2% top-1 acc.) with the state of the art on ImageNet, and similarly when transferred to other tasks. We will share our code and models. Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, Hervé Jégou |
ICML | 2 |
| 2021 | Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity AnalysisabstractWe describe a novel attribution method which is grounded in Sensitivity Analysis and uses Sobol indices. Beyond modeling the individual contributions of image regions, Sobol indices provide an efficient way to capture higher-order interactions between image regions and their contributions to a neural network's prediction through the lens of variance.We describe an approach that makes the computation of these indices efficient for high-dimensional problems by using perturbation masks coupled with efficient estimators to handle the high dimensionality of images.Importantly, we show that the proposed method leads to favorable scores on standard benchmarks for vision (and language models) while drastically reducing the computing time compared to other black-box methods -- even surpassing the accuracy of state-of-the-art white-box methods which require access to internal representations. Our code is freely available:github.com/fel-thomas/Sobol-Attribution-Method. Thomas Fel, Rémi Cadène, Mathieu Chalvidal, Matthieu Cord, David Vigouroux, Thomas Serre |
NeurIPS | 4 |
| 2021 | RED : Looking for Redundancies for Data-FreeStructured Compression of Deep Neural NetworksabstractDeep Neural Networks (DNNs) are ubiquitous in today's computer vision landscape, despite involving considerable computational costs. The mainstream approaches for runtime acceleration consist in pruning connections (unstructured pruning) or, better, filters (structured pruning), both often requiring data to retrain the model. In this paper, we present RED, a data-free, unified approach to tackle structured pruning. First, we propose a novel adaptive hashing of the scalar DNN weight distribution densities to increase the number of identical neurons represented by their weight vectors. Second, we prune the network by merging redundant neurons based on their relative similarities, as defined by their distance. Third, we propose a novel uneven depthwise separation technique to further prune convolutional layers. We demonstrate through a large variety of benchmarks that RED largely outperforms other data-free pruning methods, often reaching performance similar to unconstrained, data-driven methods. Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
NeurIPS | 3 |
| 2021 | Handling new target classes in semantic segmentation with domain adaptation
Maxime Bucher, Matthieu Cord, Patrick Pérez |
Comput. Vis. Image Underst. | 3 |
| 2020 | The Missing Data Encoder: Cross-Channel Image Completion with Hide-and-Seek Adversarial NetworkabstractImage completion is the problem of generating whole images from fragments only. It encompasses inpainting (generating a patch given its surrounding), reverse inpainting/extrapolation (generating the periphery given the central patch) as well as colorization (generating one or several channels given other ones). In this paper, we employ a deep network to perform image completion, with adversarial training as well as perceptual and completion losses, and call it the “missing data encoder” (MDE). We consider several configurations based on how the seed fragments are chosen. We show that training MDE for “random extrapolation and colorization” (MDE-REC), i.e. using random channel-independent fragments, allows a better capture of the image semantics and geometry. MDE training makes use of a novel “hide-and-seek” adversarial loss, where the discriminator seeks the original non-masked regions, while the generator tries to hide them. We validate our models qualitatively and quantitatively on several datasets, showing their interest for image completion, representation learning as well as face occlusion handling. Arnaud Dapogny, Matthieu Cord, Patrick Pérez |
AAAI | 2 |
| 2020 | Learning Representations by Predicting Bags of Visual WordsabstractSelf-supervised representation learning targets to learn convnet-based image representations from unlabeled data. Inspired by the success of NLP methods in this area, in this work we propose a self-supervised approach based on spatially dense image descriptions that encode discrete visual concepts, here called visual words. To build such discrete representations, we quantize the feature maps of a first pre-trained self-supervised convnet, over a k-means based vocabulary. Then, as a self-supervised task, we train another convnet to predict the histogram of visual words of an image (i.e., its Bag-of-Words representation) given as input a perturbed version of that image. The proposed task forces the convnet to learn perturbation-invariant and context-aware image features, useful for downstream image understanding tasks. We extensively evaluate our method and demonstrate very strong empirical results, e.g., our pre-trained self-supervised representations transfer better on detection task and similarly on classification over classes "unseen'' during pre-training, when compared to the supervised case. This also shows that the process of image discretization into visual words can provide the basis for very powerful self-supervised approaches in the image domain, thus allowing further connections to be made to related methods from the NLP domain that have been extremely successful so far. Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, Matthieu Cord |
CVPR | 5 |
| 2020 | PODNet: Pooled Outputs Distillation for Small-Tasks Incremental Learning
Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert 0001, Eduardo Valle |
ECCV (20) | 2 |
| 2020 | QuEST: Quantized Embedding Space for Transferring Knowledge
Himalaya Jain, Spyros Gidaris, Nikos Komodakis, Patrick Pérez, Matthieu Cord |
ECCV (21) | 5 |
| 2020 | Deep Entwined Learning Head Pose and Face Alignment Inside an Attentional Cascade with Doubly-Conditional fusionabstractHead pose estimation and face alignment constitute a backbone preprocessing for many applications relying on face analysis. While both are closely related tasks, they are generally addressed separately, e.g. by deducing the head pose from the landmark locations. In this paper, we propose to entwine face alignment and head pose tasks inside an attentional cascade. This cascade uses a geometry transfer network for integrating heterogeneous annotations to enhance landmark localization accuracy. Furthermore, we propose a doubly-conditional fusion scheme to select relevant feature maps, and regions thereof, based on a current head pose and landmark localization estimate. We empirically show the benefit of entwining head pose and landmark localization objectives inside our architecture, and that the proposed AC-DC model enhances the state-of-the-art accuracy on multiple databases for both face alignment and head pose estimation tasks. Arnaud Dapogny, Kevin Bailly, Matthieu Cord |
FG | 3 |
| 2020 | This Dataset Does Not Exist: Training Models from Generated ImagesabstractCurrent generative networks are increasingly proficient in generating high-resolution realistic images. These generative networks, especially the conditional ones, can potentially become a great tool for providing new image datasets. This naturally brings the question: Can we train a classifier only on the generated data? This potential availability of nearly unlimited amounts of training data challenges standard practices for training machine learning models, which have been crafted across the years for limited and fixed size datasets. In this work we investigate this question and its related challenges. We identify ways to improve significantly the performance over naive training on randomly generated images with regular heuristics. We propose three standalone techniques that can be applied at different stages of the pipeline, i.e., data generation, training on generated data, and deploying on real data. We evaluate our proposed approaches on a subset of the ImageNet dataset and show encouraging results compared to classifiers trained on real images. Victor Besnier, Himalaya Jain, Andrei Bursuc, Matthieu Cord, Patrick Pérez |
ICASSP | 4 |
| 2020 | SEMEDA: Enhancing segmentation precision with semantic edge aware loss
Arnaud Dapogny, Matthieu Cord |
Pattern Recognit. | 3 |
| 2019 | BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship DetectionabstractMultimodal representation learning is gaining more and more interest within the deep learning community. While bilinear models provide an interesting framework to find subtle combination of modalities, their number of parameters grows quadratically with the input dimensions, making their practical implementation within classical deep learning pipelines challenging. In this paper, we introduce BLOCK, a new multimodal fusion based on the block-superdiagonal tensor decomposition. It leverages the notion of block-term ranks, which generalizes both concepts of rank and mode ranks for tensors, already used for multimodal fusion. It allows to define new ways for optimizing the tradeoff between the expressiveness and complexity of the fusion model, and is able to represent very fine interactions between modalities while maintaining powerful mono-modal representations. We demonstrate the practical interest of our fusion model by using BLOCK for two challenging tasks: Visual Question Answering (VQA) and Visual Relationship Detection (VRD), where we design end-to-end learnable architectures for representing relevant interactions between modalities. Through extensive experiments, we show that BLOCK compares favorably with respect to state-of-the-art multimodal fusion models for both VQA and VRD tasks. Our code is available at https://github.com/Cadene/block.bootstrap.pytorch. Hédi Ben-Younes, Rémi Cadène, Nicolas Thome, Matthieu Cord |
AAAI | 4 |
| 2019 | MUREL: Multimodal Relational Reasoning for Visual Question AnsweringabstractMultimodal attentional networks are currently state-of-the-art models for Visual Question Answering (VQA) tasks involving real images. Although attention allows to focus on the visual content relevant to the question, this simple mechanism is arguably insufficient to model complex reasoning features required for VQA or other high-level tasks. In this paper, we propose MuRel, a multimodal relational network which is learned end-to-end to reason over real images. Our first contribution is the introduction of the MuRel cell, an atomic reasoning primitive representing interactions between question and image regions by a rich vectorial representation, and modeling region relations with pairwise combinations. Secondly, we incorporate the cell into a full MuRel network, which progressively refines visual and question interactions, and can be leveraged to define visualization schemes finer than mere attention maps. We validate the relevance of our approach with various ablation studies, and show its superiority to attention-based methods on three datasets: VQA 2.0, VQA-CP v2 and TDIUC. Our final MuRel network is competitive to or outperforms state-of-the-art results in this challenging context. Our code is available: github.com/Cadene/murel.bootstrap.pytorch Rémi Cadène, Hédi Ben-Younes, Matthieu Cord, Nicolas Thome |
CVPR | 3 |
| 2019 | SoDeep: A Sorting Deep Net to Learn Ranking Loss SurrogatesabstractSeveral tasks in machine learning are evaluated using non-differentiable metrics such as mean average precision or Spearman correlation. However, their non-differentiability prevents from using them as objective functions in a learning framework. Surrogate and relaxation methods exist but tend to be specific to a given metric. In the present work, we introduce a new method to learn approximations of such non-differentiable objective functions. Our approach is based on a deep architecture that approximates the sorting of arbitrary sets of scores. It is trained virtually for free using synthetic data. This sorting deep (SoDeep) net can then be combined in a plug-and-play manner with existing deep architectures. We demonstrate the interest of our approach in three different tasks that require ranking: Cross-modal text-image retrieval, multi-label image classification and visual memorability ranking. Our approach yields very competitive results on these three tasks, which validates the merit and the flexibility of SoDeep as a proxy for sorting operation in ranking-based losses. Martin Engilberge, Louis Chevallier, Patrick Pérez, Matthieu Cord |
CVPR | 4 |
| 2019 | ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic SegmentationabstractSemantic segmentation is a key problem for many computer vision tasks. While approaches based on convolutional neural networks constantly break new records on different benchmarks, generalizing well to diverse testing environments remains a major challenge. In numerous real-world applications, there is indeed a large gap between data distributions in train and test domains, which results in severe performance loss at run-time. In this work, we address the task of unsupervised domain adaptation in semantic segmentation with losses based on the entropy of the pixel-wise predictions. To this end, we propose two novel, complementary methods using (i) entropy loss and (ii) adversarial loss respectively. We demonstrate state-of-the-art performance in semantic segmentation on two challenging “synthetic-2-real” set-ups and show that the approach can also be used for detection. Himalaya Jain, Maxime Bucher, Matthieu Cord, Patrick Pérez |
CVPR | 4 |
| 2019 | Exploring Complex Time-series Representations for Riemannian Machine Learning of Radar DataabstractClassification of radar observations with machine learning tools is of primary importance for the identification of non-cooperative radar targets such as drones. These observations are made of complex-valued time series which possess a strong underlying structure. These signals can be processed through a time-frequency analysis, through their self-correlation (or covariance) matrices or directly as the raw signal. All representations are linked but distinct and it is known that the input representation is critical for the success of any machine learning method. In this article, we explore these three possible input representation spaces with the help of two kinds of neural networks: a temporal fully convolutional network and a Riemannian network working direcly on the manifold of covariances matrices. We show that all the considered input representations are a particular case of a generic machine learning pipeline which goes from the raw complex data to the final classification stage through con-volutional layers and Riemannian layers. This pipeline can be learnt end-to-end and is shown experimentally to give the best classification accuracy together with the best robustness to lack of data. Daniel A. Brooks, Olivier Schwander, Frédéric Barbaresco, Jean-Yves Schneider, Matthieu Cord |
ICASSP | 5 |
| 2019 | DeCaFA: Deep Convolutional Cascade for Face Alignment in the WildabstractFace Alignment is an active computer vision domain, that consists in localizing a number of facial landmarks that vary across datasets. State-of-the-art face alignment methods either consist in end-to-end regression, or in refining the shape in a cascaded manner, starting from an initial guess. In this paper, we introduce an end-to-end deep convolutional cascade (DeCaFA) architecture for face alignment. Face Alignment is an active computer vision domain, that consists in localizing a number of facial landmarks that vary across datasets. State-of-the-art face alignment methods either consist in end-to-end regression, or in refining the shape in a cascaded manner, starting from an initial guess. In this paper, we introduce DeCaFA, an end-to-end deep convolutional cascade architecture for face alignment. DeCaFA uses fully-convolutional stages to keep full spatial resolution throughout the cascade. Between each cascade stage, DeCaFA uses multiple chained transfer layers with spatial softmax to produce landmark-wise attention maps for each of several landmark alignment tasks. Weighted intermediate supervision, as well as efficient feature fusion between the stages allow to learn to progressively refine the attention maps in an end-to-end manner. We show experimentally that DeCaFA significantly outperforms existing approaches on 300W, CelebA and WFLW databases. In addition, we show that DeCaFA can learn fine alignment with reasonable accuracy from very few images using coarsely annotated data. Arnaud Dapogny, Matthieu Cord, Kevin Bailly |
ICCV | 2 |
| 2019 | Boosting Few-Shot Visual Learning With Self-SupervisionabstractFew-shot learning and self-supervised learning address different facets of the same problem: how to train a model with little or no labeled data. Few-shot learning aims for optimization methods and models that can learn efficiently to recognize patterns in the low data regime. Self-supervised learning focuses instead on unlabeled data and looks into it for the supervisory signal to feed high capacity deep neural networks. In this work we exploit the complementarity of these two domains and propose an approach for improving few-shot learning through self-supervision. We use self-supervision as an auxiliary task in a few-shot learning pipeline, enabling feature extractors to learn richer and more transferable visual representations while still using few annotated samples. Through self-supervision, our approach can be naturally extended towards using diverse unlabeled data from other datasets in the few-shot setting. We report consistent improvements across an array of architectures, datasets and self-supervision techniques. We provide the implementation code at: https://github.com/valeoai/BF3S. Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, Matthieu Cord |
ICCV | 5 |
| 2019 | DiscoNet: Shapes Learning on Disconnected Manifolds for 3D EditingabstractEditing 3D models is a very challenging task, as it requires complex interactions with the 3D shape to reach the targeted design, while preserving the global consistency and plausibility of the shape. In this work, we present an intelligent and user-friendly 3D editing tool, where the edited model is constrained to lie onto a learned manifold of realistic shapes. Due to the topological variability of real 3D models, they often lie close to a disconnected manifold, which cannot be learned with a common learning algorithm. Therefore, our tool is based on a new deep learning model, DiscoNet, which extends 3D surface autoencoders in two ways. Firstly, our deep learning model uses several autoencoders to automatically learn each connected component of a disconnected manifold, without any supervision. Secondly, each autoencoder infers the output 3D surface by deforming a pre-learned 3D template specific to each connected component. Both advances translate into improved 3D synthesis, thus enhancing the quality of our 3D editing tool. Éloi Mehr, Ariane Jourdan, Nicolas Thome, Matthieu Cord, Vincent Guitteny |
ICCV | 4 |
| 2019 | DADA: Depth-Aware Domain Adaptation in Semantic SegmentationabstractUnsupervised domain adaptation (UDA) is important for applications where large scale annotation of representative data is challenging. For semantic segmentation in particular, it helps deploy on real “target domain” data models that are trained on annotated images from a different “source domain”, notably a virtual environment. To this end, most previous works consider semantic segmentation as the only mode of supervision for source domain data, while ignoring other, possibly available, information like depth. In this work, we aim at exploiting at best such a privileged information while training the UDA model. We propose a unified depth-aware UDA framework that leverages in several complementary ways the knowledge of dense depth in the source domain. As a result, the performance of the trained semantic segmentation model on the target domain is boosted. Our novel approach indeed achieves state-of-the-art performance on different challenging synthetic-2-real benchmarks. Himalaya Jain, Maxime Bucher, Matthieu Cord, Patrick Pérez |
ICCV | 4 |
| 2019 | Delving Deep into Interpreting Neural Nets with Piece-Wise Affine RepresentationabstractDeep convolutional neural networks (CNNs) are now ubiquitous in computer vision problems. However, these models usually describe very complicated functions of the input images. For a number of application, it is of utmost importance to be able to explain the decisions of a network, e.g. by highlighting the most relevant pixels in an image or a feature map w.r.t. a particular class. In this paper, we show that CNNs locally describe piece-wise affine functions of each pixel, whose coefficient and bias can be retrieved analytically. We apply our methodology on several popular CNNs and draw interesting conclusions on the relative contributions of pixels and biases for these networks. Antoine Saporta, Arnaud Dapogny, Matthieu Cord |
ICIP | 4 |
| 2019 | Reve: Regularizing Deep Learning with Variational Entropy BoundabstractStudies on generalization performance of machine learning algorithms under the scope of information theory suggest that compressed representations can guarantee good generalization, inspiring many compression-based regularization methods. In this paper, we introduce REVE, a new regularization scheme. Noting that compressing the representation can be sub-optimal, our first contribution is to identify a variable that is directly responsible for the final prediction. Our method aims at compressing the class conditioned entropy of this latter variable. Second, we introduce a variational upper bound on this conditional entropy term. Finally, we propose a scheme to instantiate a tractable loss that is integrated within the training procedure of the neural network and demonstrate its efficiency on different neural networks and datasets. Antoine Saporta, Michaël Blot, Matthieu Cord |
ICIP | 4 |
| 2019 | Riemannian batch normalization for SPD neural networksabstractCovariance matrices have attracted attention for machine learning applications due to their capacity to capture interesting structure in the data. The main challenge is that one needs to take into account the particular geometry of the Riemannian manifold of symmetric positive definite (SPD) matrices they belong to. In the con- text of deep networks, several architectures for these matrices have recently been proposed. In our article, we introduce a Riemannian batch normalization (batch- norm) algorithm, which generalizes the one used in Euclidean nets. This novel layer makes use of geometric operations on the manifold, notably the Riemannian barycenter, parallel transport and non-linear structured matrix transformations. We derive a new manifold-constrained gradient descent algorithm working in the space of SPD matrices, allowing to learn the batchnorm layer. We validate our proposed approach with experiments in three different contexts on diverse data types: a drone recognition dataset from radar observations, and on emotion and action recognition datasets from video and motion capture data. Experiments show that the Riemannian batchnorm systematically gives better classification performance compared with leading methods and a remarkable robustness to lack of data. Daniel A. Brooks, Olivier Schwander, Frédéric Barbaresco, Jean-Yves Schneider, Matthieu Cord |
NeurIPS | 5 |
| 2019 | Zero-Shot Semantic SegmentationabstractSemantic segmentation models are limited in their ability to scale to large numbers of object classes. In this paper, we introduce the new task of zero-shot semantic segmentation: learning pixel-wise classifiers for never-seen object categories with zero training examples. To this end, we present a novel architecture, ZS3Net, combining a deep visual segmentation model with an approach to generate visual representations from semantic word embeddings. By this way, ZS3Net addresses pixel classification tasks where both seen and unseen categories are faced at test time (so called generalized zero-shot classification). Performance is further improved by a self-training step that relies on automatic pseudo-labeling of pixels from unseen classes. On the two standard segmentation datasets, Pascal-VOC and Pascal-Context, we propose zero-shot benchmarks and set competitive baselines. For complex scenes as ones in the Pascal-Context dataset, we extend our approach by using a graph-context encoding to fully leverage spatial context priors coming from class-wise segmentation maps. Maxime Bucher, Matthieu Cord, Patrick Pérez |
NeurIPS | 3 |
| 2019 | RUBi: Reducing Unimodal Biases for Visual Question AnsweringabstractVisual Question Answering (VQA) is the task of answering questions about an image. Some VQA models often exploit unimodal biases to provide the correct answer without using the image information. As a result, they suffer from a huge drop in performance when evaluated on data outside their training set distribution. This critical issue makes them unsuitable for real-world settings. We propose RUBi, a new learning strategy to reduce biases in any VQA model. It reduces the importance of the most biased examples, i.e. examples that can be correctly classified without looking at the image. It implicitly forces the VQA model to use the two input modalities instead of relying on statistical regularities between the question and the answer. We leverage a question-only model that captures the language biases by identifying when these unwanted regularities are used. It prevents the base VQA model from learning them by influencing its predictions. This leads to dynamically adjusting the loss in order to compensate for biases. We validate our contributions by surpassing the current state-of-the-art results on VQA-CP v2. This dataset is specifically designed to assess the robustness of VQA models when exposed to different question biases at test time than what was seen during training. Rémi Cadène, Corentin Dancette, Hédi Ben-Younes, Matthieu Cord, Devi Parikh |
NeurIPS | 4 |
| 2019 | Addressing Failure Prediction by Learning Model ConfidenceabstractAssessing reliably the confidence of a deep neural net and predicting its failures is of primary importance for the practical deployment of these models. In this paper, we propose a new target criterion for model confidence, corresponding to the True Class Probability (TCP). We show how using the TCP is more suited than relying on the classic Maximum Class Probability (MCP). We provide in addition theoretical guarantees for TCP in the context of failure prediction. Since the true class is by essence unknown at test time, we propose to learn TCP criterion on the training set, introducing a specific learning scheme adapted to this context. Extensive experiments are conducted for validating the relevance of the proposed approach. We study various network architectures, small and large scale datasets for image classification and semantic segmentation. We show that our approach consistently outperforms several strong methods, from MCP to Bayesian uncertainty, as well as recent approaches specifically designed for failure prediction. Charles Corbière, Nicolas Thome, Avner Bar-Hen, Matthieu Cord, Patrick Pérez |
NeurIPS | 4 |
| 2019 | End-to-End Learning of Latent Deformable Part-Based Representations for Object Detection
Taylor Mordan, Nicolas Thome, Gilles Hénaff, Matthieu Cord |
Int. J. Comput. Vis. | 4 |
| 2019 | Distributed optimization for deep learning with gossip exchange
Michaël Blot, David Picard, Nicolas Thome, Matthieu Cord |
Neurocomputing | 4 |
| 2019 | Exploiting Negative Evidence for Deep Latent Structured ModelsabstractThe abundance of image-level labels and the lack of large scale detailed annotations (e.g. bounding boxes, segmentation masks) promotes the development of weakly supervised learning (WSL) models. In this work, we propose a novel framework for WSL of deep convolutional neural networks dedicated to learn localized features from global image-level annotations. The core of the approach is a new latent structured output model equipped with a pooling function which explicitly models negative evidence, e.g. a cow detector should strongly penalize the prediction of the bedroom class. We show that our model can be trained end-to-end for different visual recognition tasks: multi-class and multi-label classification, and also structured average precision (AP) ranking. Extensive experiments highlight the relevance of the proposed method: our model outperforms state-of-the art results on six datasets. We also show that our framework can be used to improve the performance of state-of-the-art deep models for large scale image classification on ImageNet. Finally, we evaluate our model for weakly supervised tasks: in particular, a direct adaptation for weakly supervised segmentation provides a very competitive model. Thibaut Durand, Nicolas Thome, Matthieu Cord |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Finding Beans in Burgers: Deep Semantic-Visual Embedding With LocalizationabstractSeveral works have proposed to learn a two-path neural network that maps images and texts, respectively, to a same shared Euclidean space where geometry captures useful semantic relationships. Such a multi-modal embedding can be trained and used for various tasks, notably image captioning. In the present work, we introduce a new architecture of this type, with a visual path that leverages recent space-aware pooling mechanisms. Combined with a textual path which is jointly trained from scratch, our semantic-visual embedding offers a versatile model. Once trained under the supervision of captioned images, it yields new state-of-the-art performance on cross-modal retrieval. It also allows the localization of new concepts from the embedding space into any input image, delivering state-of-the-art result on the visual grounding of phrases. Martin Engilberge, Louis Chevallier, Patrick Pérez, Matthieu Cord |
CVPR | 4 |
| 2018 | Manifold Learning in Quotient SpacesabstractWhen learning 3D shapes we are usually interested in their intrinsic geometry rather than in their orientation. To deal with the orientation variations the usual trick consists in augmenting the data to exhibit all possible variability, and thus let the model learn both the geometry as well as the rotations. In this paper we introduce a new auto-encoder model for encoding and synthesis of 3D shapes. To get rid of undesirable input variability our model learns a manifold in a quotient space of the input space. Typically, we propose to quotient the space of 3D models by the action of rotations. Thus, our quotient auto-encoder allows to directly learn in the space of interest, ignoring side information. This is reflected in better performances on reconstruction and interpolation tasks, as our experiments show that our model outperforms a vanilla auto-encoder on the well-known Shapenet dataset. Moreover, our model learns a rotation-invariant representation, leading to interesting results in shapes co-alignment. Finally, we extend our quotient auto-encoder to quotient by non-rigid transformations. Éloi Mehr, André Lieutier, Fernando Sanchez Bermudez, Vincent Guitteny, Nicolas Thome, Matthieu Cord |
CVPR | 6 |
| 2018 | HybridNet: Classification and Reconstruction Cooperation for Semi-supervised Learning
Thomas Robert 0001, Nicolas Thome, Matthieu Cord |
ECCV (7) | 3 |
| 2018 | Shade: Information-Based Regularization for Deep LearningabstractRegularization is a big issue for training deep neural networks. In this paper, we propose a new information-theory-based regularization scheme named SHADE for SHAnnon DEcay. The originality of the approach is to define a prior based on conditional entropy, which explicitly decouples the learning of invariant representations in the regularizer and the learning of correlations between inputs and labels in the data fitting term. Our second contribution is to derive a stochastic version of the regularizer compatible with deep learning, resulting in a tractable training scheme. We empirically validate the efficiency of our approach to improve classification performances compared to standard regularization schemes on several standard architectures. Michaël Blot, Thomas Robert 0001, Nicolas Thome, Matthieu Cord |
ICIP | 4 |
| 2018 | Revisiting Multi-Task Learning with ROCK: a Deep Residual Auxiliary Block for Visual DetectionabstractMulti-Task Learning (MTL) is appealing for deep learning regularization. In this paper, we tackle a specific MTL context denoted as primary MTL, where the ultimate goal is to improve the performance of a given primary task by leveraging several other auxiliary tasks. Our main methodological contribution is to introduce ROCK, a new generic multi-modal fusion block for deep learning tailored to the primary MTL context. ROCK architecture is based on a residual connection, which makes forward prediction explicitly impacted by the intermediate auxiliary representations. The auxiliary predictor's architecture is also specifically designed to our primary MTL context, by incorporating intensive pooling operators for maximizing complementarity of intermediate representations. Extensive experiments on NYUv2 dataset (object detection with scene classification, depth prediction, and surface normal estimation as auxiliary tasks) validate the relevance of the approach and its superiority to flat MTL approaches. Our method outperforms state-of-the-art object detection models on NYUv2 dataset by a large margin, and is also able to handle large-scale heterogeneous inputs (real and synthetic images) with missing annotation modalities. Taylor Mordan, Nicolas Thome, Gilles Hénaff, Matthieu Cord |
NeurIPS | 4 |
| 2018 | Cross-Modal Retrieval in the Cooking Context: Learning Semantic Text-Image EmbeddingsabstractDesigning powerful tools that support cooking activities has rapidly gained popularity due to the massive amounts of available data, as well as recent advances in machine learning that are capable of analyzing them. In this paper, we propose a cross-modal retrieval model aligning visual and textual data (like pictures of dishes and their recipes) in a shared representation space. We describe an effective learning scheme, capable of tackling large-scale problems, and validate it on the Recipe1M dataset containing nearly 1 million picture-recipe pairs. We show the effectiveness of our approach regarding previous state-of-the-art models and present qualitative results over computational cooking use cases. Micael Carvalho, Rémi Cadène, David Picard, Laure Soulier, Nicolas Thome, Matthieu Cord |
SIGIR | 6 |
| 2018 | Classifying low-resolution images by integrating privileged information in deep CNNs
Marion Chevalier, Nicolas Thome, Gilles Hénaff, Matthieu Cord |
Pattern Recognit. Lett. | 4 |
| 2018 | SyMIL: MinMax Latent SVM for Weakly Labeled DataabstractDesigning powerful models able to handle weakly labeled data are a crucial problem in machine learning. In this paper, we propose a new multiple instance learning (MIL) framework. Examples are represented as bags of instances, but we depart from standard MIL assumptions by introducing a symmetric strategy (SyMIL) that seeks discriminative instances in positive and negative bags. The idea is to use the instance the most distant from the hyper-plan to classify the bag. We provide a theoretical analysis featuring the generalization properties of our model. We derive a large margin formulation of our problem, which is cast as a difference of convex functions, and optimized using concave-convex procedure. We provide a primal version optimizing with stochastic subgradient descent and a dual version optimizing with one-slack cutting-plane. Successful experimental results are reported on standard MIL and weakly supervised object detection data sets: SyMIL significantly outperforms competitive methods (mi/MI/Latent-SVM), and gives very competitive performance compared to state-of-the-art works. We also analyze the selected instances of symmetric and asymmetric approaches on weakly supervised object detection and text classification tasks. Finally, we show complementarity of SyMIL with recent works on learning with label proportions on standard MIL data sets. Thibaut Durand, Nicolas Thome, Matthieu Cord |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Deformable Part-based Fully Convolutional Network for Object Detection
Taylor Mordan, Nicolas Thome, Gilles Hénaff, Matthieu Cord |
BMVC | 4 |
| 2017 | WILDCAT: Weakly Supervised Learning of Deep ConvNets for Image Classification, Pointwise Localization and SegmentationabstractThis paper introduces WILDCAT, a deep learning method which jointly aims at aligning image regions for gaining spatial invariance and learning strongly localized features. Our model is trained using only global image labels and is devoted to three main visual recognition tasks: image classification, weakly supervised object localization and semantic segmentation. WILDCAT extends state-of-the-art Convolutional Neural Networks at three main levels: the use of Fully Convolutional Networks for maintaining spatial resolution, the explicit design in the network of local features related to different class modalities, and a new way to pool these features to provide a global image prediction required for weakly supervised training. Extensive experiments show that our model significantly outperforms state-of-the-art methods. Thibaut Durand, Taylor Mordan, Nicolas Thome, Matthieu Cord |
CVPR | 4 |
| 2017 | MUTAN: Multimodal Tucker Fusion for Visual Question AnsweringabstractBilinear models provide an appealing framework for mixing and merging information in Visual Question Answering (VQA) tasks. They help to learn high level associations between question meaning and visual concepts in the image, but they suffer from huge dimensionality issues. We introduce MUTAN, a multimodal tensor-based Tucker decomposition to efficiently parametrize bilinear interactions between visual and textual representations. Additionally to the Tucker framework, we design a low-rank matrix-based decomposition to explicitly constrain the interaction rank. With MUTAN, we control the complexity of the merging scheme while keeping nice interpretable fusion relations. We show how the Tucker decomposition framework generalizes some of the latest VQA architectures, providing state-of-the-art results. Hédi Ben-Younes, Rémi Cadène, Matthieu Cord, Nicolas Thome |
ICCV | 3 |
| 2017 | Learning a Distance Metric from Relative Comparisons between Quadruplets of Images
Marc T. Law, Nicolas Thome, Matthieu Cord |
Int. J. Comput. Vis. | 3 |
| 2017 | Gaze latent support vector machine for image classification improved by weakly supervised region selection
Xin Wang 0053, Nicolas Thome, Matthieu Cord |
Pattern Recognit. | 3 |
| 2016 | WELDON: Weakly Supervised Learning of Deep Convolutional Neural NetworksabstractIn this paper, we introduce a novel framework for WEakly supervised Learning of Deep cOnvolutional neural Networks (WELDON). Our method is dedicated to automatically selecting relevant image regions from weak annotations, e.g. global image labels, and encompasses the following contributions. Firstly, WELDON leverages recent improvements on the Multiple Instance Learning paradigm, i.e. negative evidence scoring and top instance selection. Secondly, the deep CNN is trained to optimize Average Precision, and fine-tuned on the target dataset with efficient computations due to convolutional feature sharing. A thorough experimental validation shows that WELDON outperforms state-of-the-art results on six different datasets. Thibaut Durand, Nicolas Thome, Matthieu Cord |
CVPR | 3 |
| 2016 | Closed-Form Training of Mahalanobis Distance for Supervised ClusteringabstractClustering is the task of grouping a set of objects so that objects in the same cluster are more similar to each other than to those in other clusters. The crucial step in most clustering algorithms is to find an appropriate similarity metric, which is both challenging and problem-dependent. Supervised clustering approaches, which can exploit labeled clustered training data that share a common metric with the test set, have thus been proposed. Unfortunately, current metric learning approaches for supervised clustering do not scale to large or even medium-sized datasets. In this paper, we propose a new structured Mahalanobis Distance Metric Learning method for supervised clustering. We formulate our problem as an instance of large margin structured prediction and prove that it can be solved very efficiently in closed-form. The complexity of our method is (in most cases) linear in the size of the training dataset. We further reveal a striking similarity between our approach and multivariate linear regression. Experiments on both synthetic and real datasets confirm several orders of magnitude speedup while still achieving state-of-the-art performance. Marc T. Law, Yaoliang Yu, Matthieu Cord, Eric P. Xing |
CVPR | 3 |
| 2016 | Max-min convolutional neural networks for image classificationabstractConvolutional neural networks (CNN) are widely used in computer vision, especially in image classification. However, the way in which information and invariance properties are encoded through in deep CNN architectures is still an open question. In this paper, we propose to modify the standard convolutional block of CNN in order to transfer more information layer after layer while keeping some invariance within the network. Our main idea is to exploit both positive and negative high scores obtained in the convolution maps. This behavior is obtained by modifying the traditional activation function step before pooling. We are doubling the maps with specific activations functions, called MaxMin strategy, in order to achieve our pipeline. Extensive experiments on two classical datasets, MNIST and CIFAR-10, show that our deep MaxMin convolutional net outperforms standard CNN. Michaël Blot, Matthieu Cord, Nicolas Thome |
ICIP | 2 |
| 2016 | Deep Neural Networks Under StressabstractIn recent years, deep architectures have been used for transfer learning with state-of-the-art performance in many datasets. The properties of their features remain, however, largely unstudied under the transfer perspective. In this work, we present an extensive analysis of the resiliency of feature vectors extracted from deep models, with special focus on the trade-off between performance and compression rate. By introducing perturbations to image descriptions extracted from a deep convolutional neural network, we change their precision and number of dimensions, measuring how it affects the final score. We show that deep features are more robust to these disturbances when compared to classical approaches, achieving a compression rate of 98.4%, while losing only 0.88% of their original score for Pascal VOC 2007. Micael Carvalho, Matthieu Cord, Sandra Eliza Fontes de Avila, Nicolas Thome, Eduardo Valle |
ICIP | 2 |
| 2016 | Gaze latent support vector machine for image classificationabstractThis paper deals with image categorization from weak supervision, e.g. global image labels. We propose to improve the region selection performed in latent variable models such as Latent Support Vector Machine (LSVM) by leveraging human eye movement features collected from an eye-tracker device. We introduce a new model, Gaze Latent Support Vector Machine (G-LSVM), whose region selection during training is biased toward regions with a large gaze density ratio. On this purpose, the training objective is enriched with a gaze loss, from which we derive a convex upper bound, leading to a Concave-Convex Procedure (CCCP) optimization scheme. Experiments show that G-LSVM significantly outperforms LSVM in both object detection and action recognition problems on PASCAL VOC 2012. We also show that our G-LSVM is even slightly better than a model trained from bounding box annotations, while gaze labels are much cheaper to collect. Xin Wang 0053, Nicolas Thome, Matthieu Cord |
ICIP | 3 |
| 2015 | MANTRA: Minimum Maximum Latent Structural SVM for Image Classification and RankingabstractIn this work, we propose a novel Weakly Supervised Learning (WSL) framework dedicated to learn discriminative part detectors from images annotated with a global label. Our WSL method encompasses three main contributions. Firstly, we introduce a new structured output latent variable model, Minimum mAximum lateNt sTRucturAl SVM (MANTRA), which prediction relies on a pair of latent variables: h+(resp. h-) provides positive (resp. negative) evidence for a given output y. Secondly, we instantiate MANTRA for two different visual recognition tasks: multi-class classification and ranking. For ranking, we propose efficient solutions to exactly solve the inference and the loss-augmented problems. Finally, extensive experiments highlight the relevance of the proposed method: MANTRA outperforms state-of-the art results on five different datasets. Thibaut Durand, Nicolas Thome, Matthieu Cord |
ICCV | 3 |
| 2015 | Exemplar based metric learning for robust visual localizationabstractThis paper presents an exemplar based metric learning framework dedicated to robust visual localization in complex scenes, e.g. street images. The proposed framework learns off-line a specific (local) metric for each image of the database, so that the distance between a database image and a query image representing the same scene is smaller than the distance between the current image and other images of the database. To achieve this goal, we generate geometric and photometric transformations as proxies for query images. From the generated constraints, the learning problem is cast as a convex optimization problem over the cone of positive semi-definite matrices, which is efficiently solved using a projected gradient descent scheme. Successful experiments, conducted using a freely available geo-referenced image database, reveal that the proposed method significantly improves results over the metric in the input space, while being as efficient at test time. In addition, we show that the model learns discriminating features for the localization task, and is able to gain invariance to meaningful transformations. Cédric Le Barz, Nicolas Thome, Matthieu Cord, Stéphane Herbin, Martial Sanfourche |
ICIP | 3 |
| 2015 | LR-CNN for fine-grained classification with varying resolutionabstractIn this work, we present an extended study of image representations for fine-grained classification with respect to image resolution. Understudied in literature, this parameter yet presents many practical and theoretical interests, e.g. in embedded systems where restricted computational resources prevent treating high-resolution images. It is thus interesting to figure out which representation provides the best results in this particular context. On this purpose, we evaluate Fisher Vectors and deep representations on two significant finegrained oriented datasets: FGVC Aircraft [1] and PPMI [2]. We also introduce LR-CNN, a deep structure designed for classification of low-resolution images with strong semantic content. This net provides rich compact features and outperforms both pre-trained deep features and Fisher Vectors. Marion Chevalier, Nicolas Thome, Matthieu Cord, Jérôme Fournier, Gilles Hénaff, Elodie Dusch |
ICIP | 3 |
| 2014 | Fantope Regularization in Metric LearningabstractThis paper introduces a regularization method to explicitly control the rank of a learned symmetric positive semidefinite distance matrix in distance metric learning. To this end, we propose to incorporate in the objective function a linear regularization term that minimizes the k smallest eigenvalues of the distance matrix. It is equivalent to minimizing the trace of the product of the distance matrix with a matrix in the convex hull of rank-k projection matrices, called a Fantope. Based on this new regularization method, we derive an optimization scheme to efficiently learn the distance matrix. We demonstrate the effectiveness of the method on synthetic and challenging real datasets of face verification and image classification with relative attributes, on which our method outperforms state-of-the-art metric learning algorithms. Marc T. Law, Nicolas Thome, Matthieu Cord |
CVPR | 3 |
| 2014 | Semantic pooling for image categorization using multiple kernel learningabstractIn this paper, we propose a new method for taking into account the spatial information in image categorization. More specifically, we remove the loss of spatial information in Bag of Words related methods by computing the image signature over specific regions selected by object detectors. We propose to select the detectors using Multiple Kernel Learning techniques. We carry out experiments on the well known VOC 2007 dataset, and show our semantic pooling obtains promising results. Thibaut Durand, David Picard, Nicolas Thome, Matthieu Cord |
ICIP | 4 |
| 2014 | Incremental learning of latent structural SVM for weakly supervised image classificationabstractVisual learning with weak supervision is a promising research area, since it offers the possibility to build large image datasets at reasonable cost. In this paper, we address the problem of weakly supervised object detection, where the goal is to predict the label of the image using object position as latent variable. We propose a new method that builds upon the Latent Structural SVM (LSSVM) formalism. Specifically, we introduce an original coarse-to-fine approach that limits the evolution of the latent parameter subspace. This incremental strategy drives the learning towards better solutions, providing a model with increased predictive accuracy. In addition, this leads to a significant speed up during learning and inference compared to standard sliding window methods. Experiments carried out on Mammal dataset validate the good performances and fast training of the method compared to state-of-the-art works. Thibaut Durand, Nicolas Thome, Matthieu Cord, David Picard |
ICIP | 3 |
| 2014 | SnooperText: A text detection system for automatic indexing of urban scenes
Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J. Leite, Jorge Stolfi |
Comput. Vis. Image Underst. | 3 |
| 2014 | Model-Based Analysis-Synthesis for Realistic Tree Reconstruction and Growth SimulationabstractDue to complexity, vegetation analysis and reconstruction of remote sensing data are challenging problems. Using architectural tree models combined with model inputs estimated from aerial image analysis, this paper presents an analysis-synthesis approach for urban vegetation detection, modeling, and reconstruction. Tree species, height, and crown size information are extracted by aerial image analysis. These variables serve for model inversion to retrieve plant age, climatic growth conditions, and competition with neighbors. Functional-structural individual-based tree models are used to reconstruct and visualize virtual trees and their time evolutions realistically in a 3-D viewer rendering the models with geographical coordinates in the reconstructed scene. Our main contributions are: 1) a novel approach for generating plant models in 3-D reconstructed scenes based on the analysis of the geometric properties of the data, and 2) a modeling workflow for the reconstruction and growth simulation of vegetation in urban or natural environments. Corina Iovan, Paul-Henry Cournède, Thomas Guyard, Benoit Bayol, Didier Boldo, Matthieu Cord |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2014 | Learning Deep Hierarchical Visual Feature CodingabstractIn this paper, we propose a hybrid architecture that combines the image modeling strengths of the bag of words framework with the representational power and adaptability of learning deep architectures. Local gradient-based descriptors, such as SIFT, are encoded via a hierarchical coding scheme composed of spatial aggregating restricted Boltzmann machines (RBM). For each coding layer, we regularize the RBM by encouraging representations to fit both sparse and selective distributions. Supervised fine-tuning is used to enhance the quality of the visual representation for the categorization task. We performed a thorough experimental evaluation using three image categorization data sets. The hierarchical coding scheme achieved competitive categorization accuracies of 79.7% and 86.4% on the Caltech-101 and 15-Scenes data sets, respectively. The visual representations learned are compact and the model's inference is fast, as compared with sparse coding methods. The low-level representations of descriptors that were learned using this method result in generic features that we empirically found to be transferrable between different image data sets. Further analysis reveal the significance of supervised fine-tuning when the architecture has two layers of representations as opposed to a single layer. Hanlin Goh, Nicolas Thome, Matthieu Cord, Joo-Hwee Lim |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2013 | Dynamic Scene Classification: Learning Motion Descriptors with Slow Features AnalysisabstractIn this paper, we address the challenging problem of categorizing video sequences composed of dynamic natural scenes. Contrarily to previous methods that rely on handcrafted descriptors, we propose here to represent videos using unsupervised learning of motion features. Our method encompasses three main contributions: 1) Based on the Slow Feature Analysis principle, we introduce a learned local motion descriptor which represents the principal and more stable motion components of training videos. 2) We integrate our local motion feature into a global coding/pooling architecture in order to provide an effective signature for each video sequence. 3) We report state of the art classification performances on two challenging natural scenes data sets. In particular, an outstanding improvement of 11% in classification score is reached on a data set introduced in 2012. Christian Theriault, Nicolas Thome, Matthieu Cord |
CVPR | 3 |
| 2013 | Quadruplet-Wise Image Similarity LearningabstractThis paper introduces a novel similarity learning framework. Working with inequality constraints involving quadruplets of images, our approach aims at efficiently modeling similarity from rich or complex semantic label relationships. From these quadruplet-wise constraints, we propose a similarity learning framework relying on a convex optimization scheme. We then study how our metric learning scheme can exploit specific class relationships, such as class ranking (relative attributes), and class taxonomy. We show that classification using the learned metrics gets improved performance over state-of-the-art methods on several datasets. We also evaluate our approach in a new application to learn similarities between web page screenshots in a fully unsupervised way. Marc T. Law, Nicolas Thome, Matthieu Cord |
ICCV | 3 |
| 2013 | Image classification using object detectorsabstractImage categorization is one of the most competitive topic in computer vision and image processing. In this paper, we propose to use trained object and region detectors to represent the visual content of each image. Compared to similar methods found in the literature, our method encompasses two main areas of novelty: introducing a new spatial pooling formalism and designing a late fusion strategy for combining our representation with state-of-the art methods based on low-level descriptors, e.g. Fisher Vectors and BossaNova. Our experiments carried out in the challenging PASCAL VOC 2007 dataset reveal outstanding performances. When combined with low-level representations, we reach more than 67.6% in MAP, outperforming recently reported results in this dataset with a large margin. Thibaut Durand, Nicolas Thome, Matthieu Cord, Sandra Eliza Fontes de Avila |
ICIP | 3 |
| 2013 | Top-Down Regularization of Deep Belief NetworksabstractDesigning a principled and effective algorithm for learning deep architectures is a challenging problem. The current approach involves two training phases: a fully unsupervised learning followed by a strongly discriminative optimization. We suggest a deep learning strategy that bridges the gap between the two phases, resulting in a three-phase learning procedure. We propose to implement the scheme using a method to regularize deep belief networks with top-down information. The network is constructed from building blocks of restricted Boltzmann machines learned by combining bottom-up and top-down sampled signals. A global optimization procedure that merges samples from a forward bottom-up pass and a top-down pass is used. Experiments on the MNIST dataset show improvements over the existing algorithms for deep belief networks. Object recognition results on the Caltech-101 dataset also yield competitive results. Hanlin Goh, Nicolas Thome, Matthieu Cord, Joo-Hwee Lim |
NIPS | 3 |
| 2013 | Pooling in image representation: The visual codeword point of view
Sandra Eliza Fontes de Avila, Nicolas Thome, Matthieu Cord, Eduardo Valle, Arnaldo de Albuquerque Araújo |
Comput. Vis. Image Underst. | 3 |
| 2013 | JKernelMachines: a simple framework for kernel machine
David Picard, Nicolas Thome, Matthieu Cord |
J. Mach. Learn. Res. | 3 |
| 2013 | Text detection in street level images
Jonathan Fabrizio, Beatriz Marcotegui, Matthieu Cord |
Pattern Anal. Appl. | 3 |
| 2013 | T-HOG: An effective gradient-based descriptor for single line text regions
Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J. Leite, Jorge Stolfi |
Pattern Recognit. | 3 |
| 2013 | Extended Coding and Pooling in the HMAX ModelabstractThis paper presents an extension of the HMAX model, a neural network model for image classification. The HMAX model can be described as a four-level architecture, with the first level consisting of multiscale and multiorientation local filters. We introduce two main contributions to this model. First, we improve the way the local filters at the first level are integrated into more complex filters at the last level, providing a flexible description of object regions and combining local information of multiple scales and orientations. These new filters are discriminative and yet invariant, two key aspects of visual classification. We evaluate their discriminative power and their level of invariance to geometrical transformations on a synthetic image set. Second, we introduce a multiresolution spatial pooling. This pooling encodes both local and global spatial information to produce discriminative image signatures. Classification results are reported on three image data sets: Caltech101, Caltech256, and fifteen scenes. We show significant improvements over previous architectures using a similar framework. Christian Theriault, Nicolas Thome, Matthieu Cord |
IEEE Trans. Image Process. | 3 |
| 2012 | Structural and visual comparisons for web page archivingabstractIn this paper, we propose a Web page archiving system that combines state-of-the-art comparison methods based on the source codes of Web pages, with computer vision techniques. To detect whether successive versions of a Web page are similar or not, our system is based on: (1) a combination of structural and visual comparison methods embedded in a statistical discriminative model, (2) a visual similarity measure designed for Web pages that improves change detection, (3) a supervised feature selection method adapted to Web archiving. We train a Support Vector Machine model with vectors of similarity scores between successive versions of pages. The trained model then determines whether two versions, defined by their vector of similarity scores, are similar or not. Experiments on real archives validate our approach. Marc T. Law, Nicolas Thome, Stéphane Gançarski, Matthieu Cord |
ACM Symposium on Document Engineering | 4 |
| 2012 | Unsupervised and Supervised Visual Codes with Restricted Boltzmann Machines
Hanlin Goh, Nicolas Thome, Matthieu Cord, Joo-Hwee Lim |
ECCV (5) | 3 |
| 2012 | Learning geometric combinations of Gaussian kernels with alternating Quasi-Newton algorithm
David Picard, Nicolas Thome, Matthieu Cord, Alain Rakotomamonjy |
ESANN | 3 |
| 2012 | Contextual detection of drawn symbols in old mapsabstractIn this paper, we tackle the problem of detecting drawn symbols in old maps. We propose a novel approach that combines powerful low level descriptors to represent the local content of the objects, and contextual features to overcome the local analysis ambiguity. Our contribution is two-fold. Firstly, we propose a novel contextual feature adapted to our problem, where the context is integrated at two levels. In a close neighborhood, a local analysis is carried out to remove visual ambiguities between symbols. In a larger extent, co-occurrence statistics between classes are stored. Secondly, we propose an entire processing chain for learning and detection. The proposed method is evaluated on real french maps from the 18thcentury. The experiments show the efficiency of the detection system, and validate the relevance of the proposed contextual feature to improve detection performances. Jonathan Guyomard, Nicolas Thome, Matthieu Cord, Thierry Artières |
ICIP | 3 |
| 2012 | Classification of Urban Scenes from Geo-referenced Images in Urban Street-View ContextabstractThis paper addresses the challenging problem of scene classification in street-view georeferenced images of urban environments. More precisely, the goal of this task is semantic image classification, consisting in predicting in a given image, the presence or absence of a pre-defined class (e.g. shops, vegetation, etc.). The approach is based on the BOSSA representation, which enriches the Bag of Words (BoW) model, in conjunction with the Spatial Pyramid Matching scheme and kernel-based machine learning techniques. The proposed method handles problems that arise in large scale urban environments due to acquisition conditions (static and dynamic objects/pedestrians) combined with the continuous acquisition of data along the vehicle's direction, the varying light conditions and strong occlusions (due to the presence of trees, traffic signs, cars, etc.) giving rise to high intra-class variability. Experiments were conducted on a large dataset of high resolution images collected from two main avenues from the 12th district in Paris and the approach shows promising results. Corina Iovan, David Picard, Nicolas Thome, Matthieu Cord |
ICMLA (2) | 4 |
| 2012 | An application of swarm intelligence to distributed image retrieval
David Picard, Arnaud Revel, Matthieu Cord |
Inf. Sci. | 3 |
| 2012 | Locality-Sensitive Hashing for Chi2 DistanceabstractIn the past 10 years, new powerful algorithms based on efficient data structures have been proposed to solve the problem of Nearest Neighbors search (or Approximate Nearest Neighbors search). If the Euclidean Locality Sensitive Hashing algorithm, which provides approximate nearest neighbors in a euclidean space with sublinear complexity, is probably the most popular, the euclidean metric does not always provide as accurate and as relevant results when considering similarity measure as the Earth-Mover Distance and 2 distances. In this paper, we present a new LSH scheme adapted to 2 distance for approximate nearest neighbors search in high-dimensional spaces. We define the specific hashing functions, we prove their local-sensitivity, and compare, through experiments, our method with the Euclidean Locality Sensitive Hashing algorithm in the context of image retrieval on real image databases. The results prove the relevance of such a new LSH scheme either providing far better accuracy in the context of image retrieval than euclidean scheme for an equivalent speed, or providing an equivalent accuracy but with a high gain in terms of processing speed. David Gorisse, Matthieu Cord, Frédéric Precioso |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | BOSSA: Extended bow formalism for image classificationabstractIn image classification, the most powerful statistical learning approaches are based on the Bag-of-Words paradigm. In this article, we propose an extension of this formalism. Considering the Bag-of-Features, dictionary coding and pooling steps, we propose to focus on the pooling step. Instead of using the classical sum or max pooling strategies, we introduced a density function-based pooling strategy. This flexible formalism allows us to better represent the links between dictionary codewords and local descriptors in the resulting image signature. We evaluate our approach in two very challenging tasks of video and image classification, involving very high level semantic categories with large and nuanced visual diversity. Sandra Eliza Fontes de Avila, Nicolas Thome, Matthieu Cord, Eduardo Valle, Arnaldo de Albuquerque Araújo |
ICIP | 3 |
| 2011 | Learning invariant color features with sparse topographic restricted Boltzmann machinesabstractOur objective is to learn invariant color features directly from data via unsupervised learning. In this paper, we introduce a method to regularize restricted Boltzmann machines during training to obtain features that are sparse and topographically organized. Upon analysis, the features learned are Gabor-like and demonstrate a coding of orientation, spatial position, frequency and color that vary smoothly with the topography of the feature map. There is also differentiation between monochrome and color filters, with some exhibiting color-opponent properties. We also found that the learned representation is more invariant to affine image transformations and changes in illumination color. Hanlin Goh, Lukasz Kusmierz, Joo-Hwee Lim, Nicolas Thome, Matthieu Cord |
ICIP | 5 |
| 2011 | Snoopertrack: Text detection and tracking for outdoor videosabstractIn this work we introduced SnooperTrack, an algorithm for the automatic detection and tracking of text objects - such as store names, traffic signs, license plates, and advertisements - in videos of out door scenes. The purpose is to improve the performances of text detection process in still images by taking advantage of the temporal coherence in videos. We first propose an efficient tracking algorithm using particle filtering framework with original region descriptors. The second contribution is our strategy to merge tracked regions and new detections. We also propose an improved version of our previously published text detection algorithm in still images. Tests indicate that SnooperTrack is fast, robust, enable false positive suppression, and achieved great performances in complex videos of outdoor scenes. Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J. Leite, Jorge Stolfi |
ICIP | 3 |
| 2011 | Efficient Bag-of-Feature kernel representation for image similarity searchabstractAlthough “Bag-of-Features” image models have shown very good potential for object matching and image retrieval, such a complex data representation requires computationally expensive similarity measure evaluation. In this paper, we propose a framework unifying dictionary-based and kernel-based similarity functions that highlights the tradeoff between powerful data representation and eff cient similarity computation. On the basis of this formalism, we propose a new kernel-based similarity approach for Bag-of-Feature descriptions. We introduce a method for fast similarity search in large image databases. The conducted experiments prove that our approach is very competitive among State-of-the-art methods for similarity retrieval tasks. Frédéric Precioso, Matthieu Cord, David Gorisse, Nicolas Thome |
ICIP | 2 |
| 2011 | HMAX-S: Deep scale representation for biologically inspired image categorizationabstractThis paper presents an improvement on a biologically inspired network for image classification. Previous models have used a multi-scale and multi-orientation architecture to gain robustness to transformations and to extract complex visual features. Our contribution to this type of architecture resides in the building of complex visual features which are better tuned to images structures. We allow the network to build complex features with richer information in terms of the local scales of image structures. Our classification results show significant improvements over previous architectures using the same framework. Christian Theriault, Nicolas Thome, Matthieu Cord |
ICIP | 3 |
| 2011 | Spatio-Temporal Tube data representation and Kernel design for SVM-based video object retrieval system
Shuji Zhao, Frédéric Precioso, Matthieu Cord |
Multim. Tools Appl. | 3 |
| 2011 | SALSAS: Sub-linear active learning strategy with approximate k-NN search
David Gorisse, Matthieu Cord, Frédéric Precioso |
Pattern Recognit. | 2 |
| 2010 | Scalable active learning strategy for object category retrievalabstractSince the digital revolution, the volume of images to be processed has grown exponentially. Interactive search systems have to deal with these huge databases to remain effective. As the complexity of on-line learning methods is at least linear in the size of the database, scalability is the major problem for these methods. Fast retrieval systems, with index structures for fast navigation, have hence become like a Holy Grail. In this article, we propose a strategy to overcome this scalability limitation. Our technique exploits ultra fast retrieval methods as Locally Sensitive Hashing to speed up active learning system. Experiments on database of 180 K images are reported. The results show that our method is 45 times faster than state of the art approaches for similar accuracy. David Gorisse, Matthieu Cord, Frédéric Precioso |
ICIP | 2 |
| 2010 | Snoopertext: A multiresolution system for text detection in complex visual scenesabstractText detection in natural images remains a very challenging task. For instance, in an urban context, the detection is very difficult due to large variations in terms of shape, size, color, orientation, and the image may be blurred or have irregular illumination, etc. In this paper, we describe a robust and accurate multiresolution approach to detect and classify text regions in such scenarios. Based on generation/validation paradigm, we first segment images to detect character regions with a multiresolution algorithm able to manage large character size variations. The segmented regions are then filtered out using shape-based classification, and neighboring characters are merged to generate text hypotheses. A validation step computes a region signature based on texture analysis to reject false positives. We evaluate our algorithm in two challenging databases, achieving very good results. Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Jonathan Fabrizio, Beatriz Marcotegui |
ICIP | 3 |
| 2010 | An efficient system for combining complementary kernels in complex visual categorization tasksabstractRecently, increasing interest has been brought to improve image categorization performances by combining multiple descriptors. However, very few approaches have been proposed for combining features based on complementary aspects, and evaluating the performances in realistic databases. In this paper, we tackle the problem of combining different feature types (edge and color), and evaluate the performance gain in the very challenging VOC 2009 benchmark. Our contribution is three-fold. First, we propose new local color descriptors, unifying edge and color feature extraction into the “Bag Of Word” model. Second, we improve the Spatial Pyramid Matching (SPM) scheme for better incorporating spatial information into the similarity measurement. Last but not least, we propose a new combination strategy based on ℓ1Multiple Kernel Learning (MKL) that simultaneously learns individual kernel parameters and the kernel combination. Experiments prove the relevance of the proposed approach, which outperforms baseline combination methods while being computationally effective. David Picard, Nicolas Thome, Matthieu Cord |
ICIP | 3 |
| 2010 | STTK-based video object recognitionabstractIn this paper, we extend our video object recognition system to multiclass object recognition context, dealing with unbalanced data sets and comparing our resuls to state-of-the-art methods. Our approach is based on a Spatio-Temporal data representation, a dedicated kernel design and statistical learning techniques for object recognition. From video tracks made of segmented object regions in the successive frames, we extract sets of spatio-temporally coherent SIFT-based features, called Spatio-Temporal Tubes. To compare these complex tube objects, we integrate a Spatio-Temporal Tube Kernel (STTK) function into a multi-class classification framework with balancing process for unequal classes. Our approach is successfully evaluated on episodes from “Buffy, the Vampire Slayer” TV series which have been used in other works targeting same objectives. Our method proved to be more robust than dictionary based, facial feature based and key-frame based approaches. Our method is also tested on a small car database and preliminary results for car identification task illustrate its generalization potential. Shuji Zhao, Frédéric Precioso, Matthieu Cord |
ICIP | 3 |
| 2009 | Geometric consistency checking for local-descriptor based document retrievalabstractInternational audience Eduardo Valle, David Picard, Matthieu Cord |
ACM Symposium on Document Engineering | 3 |
| 2009 | Text segmentation in natural scenes using Toggle-MappingabstractWe offer, in this paper, a new method to segment text in natural scenes. This method is based on the use of a morphological operator: the Toggle Mapping. The efficiency of the method is illustrated and the method is compared, according to various criteria, with common methods issued from the state of the art. This comparison shows that our method gives better results and is faster than the state of the art methods. Our method reduces also the number of segmented regions. This can lead to time saving in a complete scheme (executing time of multiple processing steps usually depends on the number of regions) and proves that our algorithm is more relevant. Jonathan Fabrizio, Beatriz Marcotegui, Matthieu Cord |
ICIP | 3 |
| 2009 | Optimization on active learning strategy for object category retrievalabstractActive learning is a machine learning technique which has attracted a lot of research interest in the content-based image retrieval (CBIR) in recent years. To be effective, an active learning system must be fast and efficient using as few (relevance) feedback iterations as possible. Scalability is the major problem for such an on-line learning method, since the complexity of such methods on a database of size n is in the best case O(n * log(n)). In this article we propose a strategy to overcome this limitation. Our technique exploits ultra fast retrieval methods like Locality Sensitive Hashing (LSH), recently applied for unsupervised image retrieval. Combined with active selection, our method is able to achieve very fast active learning task in very large database. Experiments on VOC2006 database are reported, results are obtained four times faster while preserving the accuracy. David Gorisse, Matthieu Cord, Frédéric Precioso |
ICIP | 2 |
| 2009 | Spatio-Temporal Tube Kernel for actor retrievalabstractThis paper presents an actor video retrieval system based on face video-tubes extraction and representation with sets of temporally coherent features. Visual features, SIFT points, are tracked along a video shot, resulting in sets of feature point chains (spatio-temporal tubes). These tubes are then classified and retrieved using a kernel-based SVM learning framework for actor retrieval in a movie. In this paper, we present optimized feature tubes, we extend our feature representation with spatial location of SIFT points and we describe the new Spatio-Temporal Tube Kernel (STTK) of our content-based retrieval system. Our approach has been tested on a real movie and proved to be faster and more robust for actor retrieval task. Shuji Zhao, Frédéric Precioso, Matthieu Cord |
ICIP | 3 |
| 2008 | High-dimensional descriptor indexing for large multimedia databasesabstractIn this paper we address the subject of large multimedia database indexing for content-based retrieval. Eduardo Valle, Matthieu Cord, Sylvie Philipp-Foliguet |
CIKM | 2 |
| 2008 | Fast identification of visual documents using local descriptorsabstractIn this paper we introduce a system for the identification of visual documents. Since it stems from content-based document indexing and retrieval, our system does not need to rely on textual annotations, watermarks or other metadata, which can be missing or incorrect. Our retrieval system is based on local descriptors, which have been shown to provide accurate and robust description. Because of the high computational costs associated to the matching of local descriptors, we propose Projection KD-Forest: an indexing technique which allows efficient approximate k nearest neighbors search. Experiments demonstrate that the Projection KD-Forest allows the system to provide prompt results with negligible loss on accuracy. The Projection KD-Forest also compares well when contrasted to other strategies of k nearest neighbors search. Eduardo Valle, Matthieu Cord, Sylvie Philipp-Foliguet |
ACM Symposium on Document Engineering | 2 |
| 2008 | Long term learning for image retrieval over networksabstractIn this paper, we present a long term learning system for content based image retrieval over a network. Relevant feedback is used among different sessions to learn both the similarity function and the best routing for the searched category. Our system is based on mobile agents crawling the network in search of relevant images. An ant-behavior algorithm is used to learn the category dependent routing. With experiments on trecvid'05 key-frame dataset, we show that the smart association of category dependent routing and active learning leads to an improvement of the quality of the retrieval over time. David Picard, Arnaud Revel, Matthieu Cord |
ICIP | 3 |
| 2008 | Fast approximate kernel-based similarity search for image retrieval taskabstractIn content based image retrieval, the success of any distance-based indexing scheme depends critically on the quality of the chosen distance metric. We propose in this paper a kernel-based similarity approach working on sets of vectors to represent images. We introduce a method for fast approximate similarity search in large image databases with our kernel-based similarity metric. We evaluate our algorithm on image retrieval task and show it to be accurate and faster than linear scanning. David Gorisse, Matthieu Cord, Frédéric Precioso, Sylvie Philipp-Foliguet |
ICPR | 2 |
| 2008 | Combining visual dictionary, kernel-based similarity and learning strategy for image category retrieval
Philippe Henri Gosselin, Matthieu Cord, Sylvie Philipp-Foliguet |
Comput. Vis. Image Underst. | 2 |
| 2008 | Active Learning Methods for Interactive Image RetrievalabstractActive learning methods have been considered with increased interest in the statistical learning community. Initially developed within a classification framework, a lot of extensions are now being proposed to handle multimedia applications. This paper provides algorithms within a statistical framework to extend active learning for online content-based image retrieval (CBIR). The classification framework is presented with experiments to compare several powerful classification techniques in this information retrieval context. Focusing on interactive methods, active learning strategy is then described. The limitations of this approach for CBIR are emphasized before presenting our new active selection process RETIN. First, as any active method is sensitive to the boundary estimation between classes, the RETIN strategy carries out a boundary correction to make the retrieval process more robust. Second, the criterion of generalization error to optimize the active learning selection is modified to better represent the CBIR objective of database ranking. Third, a batch processing of images is proposed. Our strategy leads to a fast and efficient active learning scheme to retrieve sets of online images (query concept). Experiments on large databases show that the RETIN method performs well in comparison to several other active strategies. Philippe Henri Gosselin, Matthieu Cord |
IEEE Trans. Image Process. | 2 |
| 2008 | Image Retrieval Over Networks: Active Learning Using Ant AlgorithmabstractIn this article, we present a framework for distributed content based image retrieval with online learning based on ant-like mobile agents. Mobile agents crawl the network to find images matching a given example query. The images retrieved are shown to the user who labels them, following the classical relevant feedback scheme. The labels are used both to improve the similarity measure used for the retrieval and to learn paths leading to sites containing relevant images. The relevant paths are learned in an ethologically inspired way. We made experiments on the trecvid 2005 keyframe dataset showing that learning both the similarity function and the localization of the relevant images leads to a significant improvement. We also present an extension with the reuse of learned paths for later sessions leading to a further improvement. David Picard, Matthieu Cord, Arnaud Revel |
IEEE Trans. Multim. | 2 |
| 2007 | Matching Local Descriptors for Image Identification on Cultural DatabasesabstractIn this paper we present a new method for high- dimensional descriptor matching, based on the KD-tree, which is a classic method for nearest neighbours search. This new method, which we name 3-way tree, avoids the boundary effects that disrupt the KD-tree in higher dimensionalities, by the addition of redundant, overlapping sub-trees. That way, more precision is obtained for the same querying times. We evaluate our method in the context of image identification for cultural collections, a task which can greatly benefit from the use of high-dimensional local descriptors computed around Pol (Points of Interest). Eduardo Valle, Matthieu Cord, Sylvie Philipp-Foliguet |
ICDAR | 2 |
| 2007 | Kernels on Bags of Fuzzy Regions for Fast Object retrievalabstractWe propose in this paper a general kernel framework to deal with database object retrieval embedded in images with heterogeneous background. We use local features computed on fuzzy regions for image representation summarized in bags, and we propose original kernel functions to deal with sets of features and spatial constraints. Combined with SVMs classification and online learning scheme, the resulting algorithm satisfies the robustness requirements for representation and classification of objects. Experiments on a specific database having objects with heterogeneous backgrounds show the performance of our object retrieval technique. Philippe Henri Gosselin, Matthieu Cord, Sylvie Philipp-Foliguet |
ICIP (1) | 2 |
| 2007 | 3-Way-Trees: A Similarity Search Method for High-Dimensional Descriptor MatchingabstractIn this paper we look into the problem of high-dimensional local descriptor matching for image identification on cultural databases, presenting an important improvement over a classic method, the KD-tree. Our method, the 3-way tree, uses redundant, overlapping sub-trees, in order to avoid the boundary effects that disrupt the KD-tree in higher dimensionalities, achieving more precision for the same querying times. Eduardo Valle, Matthieu Cord, Sylvie Philipp-Foliguet |
ICIP (1) | 2 |
| 2007 | Stochastic exploration and active learning for image retrieval
Matthieu Cord, Philippe Henri Gosselin, Sylvie Philipp-Foliguet |
Image Vis. Comput. | 1 |
| 2006 | Image Retrieval using Long-Term Semantic LearningabstractThe automatic computation of features for content-based image retrieval still has difficulties to represent the concepts the user has in mind. Whenever an additional learning strategy (such as relevance feedback) can improve the results of the search, the system performances still depend on the representation of the image collection. We introduce in this paper a supervised optimization of a set of feature vectors. According to an incomplete set of partial labels, the method improves the representation of the image collection, even if the size, the number, and the structure of the concepts are unknown. Experiments have been carried out on a large general database in order to validate our approach. Matthieu Cord, Philippe Henri Gosselin |
ICIP | 1 |
| 2006 | Precision-Oriented Active Selection for Interactive Image RetrievalabstractActive learning methods have been considered with an increased interest in the content-based image retrieval (CBIR) community. These methods have been developed for classification problems, and do not deal with the particular characteristics of the CBIR. One of these characteristics is the criterion to optimize, for instance the error of generalization for classification, which is not the best adapted to CBIR context. We introduce in this paper an active selection which chooses the image the user should label such as the mean average precision is increased. The method is smartly combined with previous propositions, and leads to a fast and efficient active learning scheme. Experiments on a large database have been carried out in order to compare our approach to several other methods. Philippe Henri Gosselin, Matthieu Cord |
ICIP | 2 |
| 2006 | CBIR in Distributed Databases using a Multi-Agent SystemabstractInformation retrieval techniques have to face both the growing amount of data to be processed and the "natural"' distribution of these data over the network. Hence, we introduce in this paper a new architecture for image retrieval in distributed image databases, based on multi-agent, systems. Our system, inspired by "ant-agents", uses labels provided by the user for learning both the searched category of images and die path to the most relevant databases. We then show how effective can be our architecture on a generalist image database network. David Picard, Matthieu Cord, Arnaud Revel |
ICIP | 2 |
| 2006 | Performances of Mobile-Agents for Interactive Image RetrievalabstractIn this paper, we present a system for image retrieval over a network of computer based on "ant-like" mobile-agents. Image databases are hosted on the network, and the user wants to find all the images matching a specific concept (cars, flower, Italy, etc...). Usually, content based image retrieval systems (CBIR) do not consider the dispertion of the data among the network. We train a SVM classifier with examples annotated by the user and then launch mobile agents which explore the network in order to retrieve the most relevant images. Several interactive session (launching of agents then annotation of the results) are made to improve the classifier. Experiments are made both to see the influence of localization of the search concept on the quality of the learning, and to focus on the quality of the agent based solution compared to a centralizing system within a fixed amount of time for the interaction David Picard, Matthieu Cord |
Web Intelligence | 2 |
| 2006 | Feature-based approach to semi-supervised similarity learning
Philippe Henri Gosselin, Matthieu Cord |
Pattern Recognit. | 2 |
| 2005 | Semantic kernel learning for interactive image retrievalabstractContent-based image retrieval systems still have difficulties to bridge the semantic gap between the low-level representation of images and the high level concepts the user is looking for. Relevance feedback methods deal with this problem using labels provided by users, but only during the current retrieval session. In this paper, we introduce a semantic learning method to manage user labels in CBIR applications. Our approach uses a kernel matrix to represent semantic information in a statistical learning framework. The kernel matrix is updated according to labels provided by users after retrieval sessions. Experiments have been carried out on a large generalist database in order to validate our approach. Philippe Henri Gosselin, Matthieu Cord |
ICIP (1) | 2 |
| 2004 | Retin al: an active learning strategy for image category retrievalabstractActive learning methods have been considered with an increasing interest in the content-based image retrieval (CBIR) community. In this article, we propose an efficient method based on active learning strategy to retrieve large image categories. At each feedback step, the system optimizes the image set presented to the user in order to speed up the retrieval. Experimental tests on COREL photo database have been carried out. Philippe Henri Gosselin, Matthieu Cord |
ICIP | 2 |
| 2004 | Smooth Surface Reconstruction Using Tensor Fields as Structuring ElementsabstractAbstract We propose a new strategy to estimate surface normal information from highly noisy sparse data. Our approach is based on a tensor field morphologically adapted to infer normals. It acts as a three‐dimensional structuring element of smooth surfaces. Robust orientation inference for all input elements is performed by morphological operations using the tensor field. A general normal estimator is defined by combining the inferred normals, their confidences and the tensor field. This estimator can be used to directly reconstruct the surface or give input normals to other reconstruction methods. We present qualitative and quantitative results to show the behavior of the original methods and ours. A comparative discussion of these results shows the efficiency of our propositions. Marcelo Bernardes Vieira, Paulo P. Martins Jr., Arnaldo de Albuquerque Araújo, Matthieu Cord, Sylvie Philipp-Foliguet |
Comput. Graph. Forum | 4 |
| 2002 | Terrain surface modeling from altimetric dataabstractThe paper presents a method for terrain surface reconstruction from altimetric data. The approach is based on terrain modeling using a parametric surface. On the assumption that above-ground objects are outliers, M-estimators supply a theoretical framework to obtain a robust estimation of the ground surface parameters from a digital elevation model. Instead of the usual M-estimator functions, an asymmetrical weight function is introduced to improve the optimization process. We present results obtained with simulated and real data. Matthieu Cord, Thomas Belli |
ICIP (3) | 1 |
| 2002 | Long-term similarity learning in content-based image retrievalabstractThis paper presents a new learning technique for the similarity model refinement in CBIR systems. We propose a whole retrieval strategy based on a new relevance feedback scheme and on a long-term similarity learning algorithm which uses feedback information of previous sessions. We introduce this technique as the simple evolution of the short-term relevance feedback approach into a long-term similarity learning, without additional need of user interaction. Our algorithm is validated via a quality assessment realized on a heterogeneous database of 1,200 color images. Jérôme Fournier, Matthieu Cord |
ICIP (1) | 2 |
| 2001 | Back-propagation algorithm for relevance feedback in image retrievalabstractContent-based image retrieval (CBIR) usually relies on pre-attentive similarities. Results are often coarse because of the gap between the pre-attentive level and the semantic level of the user's request. The aim of relevance feedback is to refine results by taking user's expertise into account. This paper presents a new feedback architecture for CBIR. Images are compared through a weighted dissimilarity function which can be represented as a "network of dissimilarities". The weights are updated via an error backpropagation algorithm using the user's annotations of the successive set of result images. It allows an iterative refinement of the search through a simple interactive process (the user has just to specify if images are relevant or not). A quality assessment realized with three databases containing about 10,000 images shows the performance improvement after feedback. Jérôme Fournier, Matthieu Cord, Sylvie Philipp-Foliguet |
ICIP (1) | 2 |
| 2001 | Accurate Building Structure Recovery from High Resolution Aerial Imagery
Matthieu Cord, Michel Jordan, Jean Pierre Cocquerez |
Comput. Vis. Image Underst. | 1 |
| 2001 | RETIN: A Content-Based Image Indexing and Retrieval System
Jérôme Fournier, Matthieu Cord, Sylvie Philipp-Foliguet |
Pattern Anal. Appl. | 2 |
| 2001 | Three-dimensional building detection and modeling using a statistical approachabstractIn this paper, we address the problem of building reconstruction in high-resolution stereoscopic aerial imagery. We present a hierarchical strategy to detect and model buildings in urban sites, based on a global focusing process, followed by a local modeling. During the first step, we extract the building regions by exploiting to the full extent the depth information obtained with a new adaptive correlation stereo matching. In the modeling step, we propose a statistical approach, which is competitive to the sequential methods using segmentation and modeling. This parametric method is based on a multiplane model of the data, interpreted as a mixture model. From a Bayesian point of view the so-called augmentation of the model with indicator variables allows using stochastic algorithms to achieve both model parameter estimation and plane segmentation. We then report a Monte Carlo study of the performance of the stochastic algorithm on synthetic data, before displaying results on real data. Matthieu Cord, David Declercq |
IEEE Trans. Image Process. | 1 |
| 1999 | Bayesian Model Identification: Application to Building Reconstruction in Aerial Imagery
Matthieu Cord, David Declercq |
ICIP (3) | 1 |
| 1998 | Building Detection and Reconstruction from Mid- and High-Resolution Aerial Imagery
Nicolas Paparoditis, Matthieu Cord, Michel Jordan, Jean Pierre Cocquerez |
Comput. Vis. Image Underst. | 2 |