EDBT 2026 Demo / reviewers in the wild / expert
Nicolas Thome
dblp:66/4095
· DBLP profile ↗
86ranked-venue papers
7as first author
23since 2021 · last 2025
0000-0003-4871-3045ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 3 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ViLU: Learning Vision-Language Uncertainties for Failure PredictionabstractReliable Uncertainty Quantification (UQ) and failure prediction remain open challenges for Vision-Language Models (VLMs). We introduce ViLU, a new Vision-Language Uncertainty quantification framework that contextualizes uncertainty estimates by leveraging all task-relevant textual representations. ViLU constructs an uncertainty-aware multi-modal representation by integrating the visual embedding, the predicted textual embedding, and an image-conditioned textual representation via cross-attention. Unlike traditional UQ methods based on loss prediction, ViLU trains an uncertainty predictor as a binary classifier to distinguish correct from incorrect predictions using a weighted binary cross-entropy loss, making it loss-agnostic. In particular, our proposed approach is well-suited for post-hoc settings, where only vision and text embeddings are available without direct access to the model itself. Extensive experiments on diverse datasets show the significant gains of our method compared to state-of-the-art failure prediction methods. We apply our method to standard classification datasets, such as ImageNet-1k, as well as large-scale image-caption datasets like CC12M and LAION-400M. Ablation studies highlight the critical role of our architecture and training in achieving effective uncertainty quantification. Our code is publicly available and can be found here: https://github.com/ykrmm/ViLU. Marc Lafon, Yannis Karmim, Julio Silva-Rodríguez, Paul Couairon, Clément Rambour, Raphaël Fournier-S'niehotta, Ismail Ben Ayed, Jose Dolz, Nicolas Thome |
ICCV | 9 |
| 2025 | DIP: Unsupervised Dense In-Context Post-Training of Visual RepresentationsabstractWe introduce DIP, a novel unsupervised post-training method designed to enhance dense image representations in large-scale pretrained vision encoders for in-context scene understanding. Unlike prior approaches that rely on complex self-distillation architectures, our method trains the vision encoder using pseudo-tasks that explicitly simulate downstream in-context scenarios, inspired by meta-learning principles. To enable post-training on unlabeled data, we propose an automatic mechanism for generating in-context tasks that combines a pretrained diffusion model and the vision encoder itself. DIP is simple, unsupervised, and computationally efficient, requiring less than 9 hours on a single A100 GPU. By learning dense representations through pseudo in-context tasks, it achieves strong performance across a wide variety of downstream real-world in-context scene understanding tasks. It outperforms both the initial vision encoder and prior methods, offering a practical and effective solution for improving dense representations. Code available here: https://github.com/sirkosophia/DIP Sophia Sirko-Galouchenko, Spyros Gidaris, Antonín Vobecký, Andrei Bursuc, Nicolas Thome |
ICCV | 5 |
| 2025 | RT-HCP: Dealing with Inference Delays and Sample Efficiency to Learn Directly on Robotic PlatformsabstractLearning a controller directly on the robot requires extreme sample efficiency. Model-based reinforcement learning (RL) methods are the most sample efficient, but they often suffer from a too long inference time to meet the robot control frequency requirements. In this paper, we address the sample efficiency and inference time challenges with two contributions. First, we define a general framework to deal with inference delays where the slow inference robot controller provides a sequence of actions to feed the control-hungry robotic platform without execution gaps. Then, we compare several RL algorithms in the light of this framework and propose RT-HCP, an algorithm that offers an excellent trade-off between performance, sample efficiency and inference time. We validate the superiority of RT-HCP with experiments where we learn a controller directly on a simple but high frequency FURUTA pendulum platform. Code: github.com/elasriz/RTHCP Zakariae El Asri, Ibrahim Laiche, Clément Rambour, Olivier Sigaud, Nicolas Thome |
IROS | 5 |
| 2025 | JAFAR: Jack up Any Feature at Any ResolutionabstractFoundation Vision Encoders have become indispensable across a wide range of dense vision tasks. However, their operation at low spatial feature resolutions necessitates subsequent feature decompression to enable full-resolution processing. To address this limitation, we introduce JAFAR, a lightweight and flexible feature upsampler designed to enhance the spatial resolution of visual features from any Foundation Vision Encoder to any target resolution. JAFAR features an attention-based upsampling module that aligns the spatial representations of high-resolution queries with semantically enriched low-resolution keys via Spatial Feature Transform modulation. Despite the absence of high-resolution feature ground truth; we find that learning at low upsampling ratios and resolutions generalizes surprisingly well to much higher scales. Extensive experiments demonstrate that JAFAR recovers intricate pixel-level details and consistently outperforms existing feature upsampling techniques across a diverse set of dense downstream applications. Paul Couairon, Loïck Chambon, Louis Serrano, Jean-Emmanuel Haugeard, Matthieu Cord, Nicolas Thome |
NeurIPS | 6 |
| 2025 | CLIPTTA: Robust Contrastive Vision-Language Test-Time AdaptationabstractVision-language models (VLMs) like CLIP exhibit strong zero-shot capabilities but often fail to generalize under distribution shifts. Test-time adaptation (TTA) allows models to update at inference time without labeled data, typically via entropy minimization. However, this objective is fundamentally misaligned with the contrastive image-text training of VLMs, limiting adaptation performance and introducing failure modes such as pseudo-label drift and class collapse. We propose CLIPTTA, a new gradient-based TTA method for vision-language models that leverages a soft contrastive loss aligned with CLIP’s pre-training objective. We provide a theoretical analysis of CLIPTTA’s gradients, showing how its batch-aware design mitigates the risk of collapse. We further extend CLIPTTA to the open-set setting, where both in-distribution (ID) and out-of-distribution (OOD) samples are encountered, using an Outlier Contrastive Exposure (OCE) loss to improve OOD detection. Evaluated on 75 datasets spanning diverse distribution shifts, CLIPTTA consistently outperforms entropy-based objectives and is highly competitive with state-of-the-art TTA methods, outperforming them on a large number of datasets and exhibiting more stable performance across diverse shifts. Marc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers, Nicolas Thome |
NeurIPS | 5 |
| 2025 | Optimization of Rank Losses for Image RetrievalabstractIn image retrieval, standard evaluation metrics rely on score ranking, e.g. average precision (AP), recall at k (R@k), normalized discounted cumulative gain (NDCG). In this work, we introduce a general framework for robust and decomposable rank losses optimization. It addresses two major challenges for end-to-end training of deep neural networks with rank losses: non-differentiability and non-decomposability. First, we propose a general surrogate for ranking operator, SupRank, that is amenable to stochastic gradient descent. It provides an upperbound for rank losses and ensures robust training. Second, we use a simple yet effective loss function to reduce the decomposability gap between the averaged batch approximation of ranking losses and their values on the whole training set. We apply our framework to two standard metrics for image retrieval: AP and R@k. Additionally, we apply our framework to hierarchical image retrieval. We introduce an extension of AP, the hierarchical average precision $\mathcal {H}{\mathrm -AP}$H- AP , and optimize it as well as the NDCG. Finally, we create the first hierarchical landmarks retrieval dataset. We use a semi-automatic pipeline to create hierarchical labels, extending the large scale Google Landmarks v2 dataset. Elias Ramzi, Nicolas Audebert, Clément Rambour, André Araújo 0001, Xavier Bitot, Nicolas Thome |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | GalLoP: Learning Global and Local Prompts for Vision-Language Models
Marc Lafon, Elias Ramzi, Clément Rambour, Nicolas Audebert, Nicolas Thome |
ECCV (61) | 5 |
| 2024 | DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized CutabstractFoundation models have emerged as powerful tools across various domains including language, vision, and multimodal tasks. While prior works have addressed unsupervised semantic segmentation, they significantly lag behind supervised models. In this paper, we use a diffusion UNet encoder as a foundation vision encoder and introduce DiffCut, an unsupervised zero-shot segmentation method that solely harnesses the output features from the final self-attention block. Through extensive experimentation, we demonstrate that using these diffusion features in a graph based segmentation algorithm, significantly outperforms previous state-of-the-art methods on zero-shot segmentation. Specifically, we leverage a recursive Normalized Cut algorithm that regulates the granularity of detected objects and produces well-defined segmentation maps that precisely capture intricate image details. Our work highlights the remarkably accurate semantic knowledge embedded within diffusion UNet encoders that could then serve as foundation vision encoders for downstream tasks. Paul Couairon, Mustafa Shukor, Jean-Emmanuel Haugeard, Matthieu Cord, Nicolas Thome |
NeurIPS | 5 |
| 2024 | Supra-Laplacian Encoding for Transformer on Dynamic GraphsabstractFully connected Graph Transformers (GT) have rapidly become prominent in the static graph community as an alternative to Message-Passing models, which suffer from a lack of expressivity, oversquashing, and under-reaching.
However, in a dynamic context, by interconnecting all nodes at multiple snapshots with self-attention,GT loose both structural and temporal information. In this work, we introduce Supra-LAplacian encoding for spatio-temporal TransformErs (SLATE), a new spatio-temporal encoding to leverage the GT architecture while keeping spatio-temporal information.
Specifically, we transform Discrete Time Dynamic Graphs into multi-layer graphs and take advantage of the spectral properties of their associated supra-Laplacian matrix.
Our second contribution explicitly model nodes' pairwise relationships with a cross-attention mechanism, providing an accurate edge representation for dynamic link prediction.
SLATE outperforms numerous state-of-the-art methods based on Message-Passing Graph Neural Networks combined with recurrent models (e.g, LSTM), and Dynamic Graph Transformers,
on~9 datasets. Code is open-source and available at this link https://github.com/ykrmm/SLATE. Yannis Karmim, Marc Lafon, Raphaël Fournier-S'niehotta, Nicolas Thome |
NeurIPS | 4 |
| 2024 | MERLIN-Seg: Self-supervised despeckling for label-efficient semantic segmentationabstractRemote sensing satellites acquire a continuous stream of data on a daily basis. As most of those data are unlabeled, the development of algorithms requiring weak supervision is of paramount importance. In this paper, we show that the need for annotation for Synthetic Aperture Radar data can be reduced by coupling a despeckling task (self-supervised) and a segmentation task (supervised). The proposed self-supervised learning framework, called MERLIN-Seg, has been trained for building footprint extraction and achieves favorable performances even with 1% of annotated data. We show that conditioning the network on despeckling without labels is beneficial for supervised segmentation. Our experiments demonstrate that the joint training of the two tasks achieves better performances than a vanilla segmentation network in terms of IoU, F1 score, and accuracy on both simulated and real SAR images. Emanuele Dalsasso, Clément Rambour, Nicolas Trouvé, Nicolas Thome |
Comput. Vis. Image Underst. | 4 |
| 2024 | Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval
Mustafa Shukor, Nicolas Thome, Matthieu Cord |
Comput. Vis. Image Underst. | 2 |
| 2024 | Semantic augmentation by mixing contents for semi-supervised learning
Rémy Sun, Clément Masson, Gilles Hénaff, Nicolas Thome, Matthieu Cord |
Pattern Recognit. | 4 |
| 2023 | EAGLE: Large-scale Learning of Turbulent Fluid Dynamics with Mesh Transformers
Steeven Janny, Aurélien Béneteau, Madiha Nadri Wolf, Julie Digne, Nicolas Thome, Christian Wolf 0001 |
ICLR | 5 |
| 2023 | Hybrid Energy Based Model in the Feature Space for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection is a critical requirement for the deployment of deep neural networks. This paper introduces the HEAT model, a new post-hoc OOD detection method estimating the density of in-distribution (ID) samples using hybrid energy-based models (EBM) in the feature space of a pre-trained backbone. HEAT complements prior density estimators of the ID density, e.g. parametric models like the Gaussian Mixture Model (GMM), to provide an accurate yet robust density estimation. A second contribution is to leverage the EBM framework to provide a unified density estimation and to compose several energy terms. Extensive experiments demonstrate the significance of the two contributions. HEAT sets new state-of-the-art OOD detection results on the CIFAR-10 / CIFAR-100 benchmark as well as on the large-scale Imagenet benchmark. The code is available at: https://github.com/MarcLafon/heatood. Marc Lafon, Elias Ramzi, Clément Rambour, Nicolas Thome |
ICML | 4 |
| 2023 | Full Contextual Attention for Multi-resolution Transformers in Semantic SegmentationabstractTransformers have proved to be very effective for visual recognition tasks. In particular, vision transformers construct compressed global representations through self-attention and learnable class tokens. Multi-resolution transformers have shown recent successes in semantic segmentation but can only capture local interactions in high-resolution feature maps. This paper extends the notion of global tokens to build GLobal Attention Multi-resolution (GLAM) transformers. GLAM is a generic module that can be integrated into most existing transformer backbones. GLAM includes learnable global tokens, which unlike previous methods can model interactions between all image regions, and extracts powerful representations during training. Extensive experiments show that GLAM-Swin or GLAM-Swin-Unet exhibit substantially better performances than their vanilla counterparts on ADE20K and Cityscapes. Moreover, GLAM can be used to segment large 3D medical images, and GLAM-nnFormer achieves new state-of-the-art performance on the BCV dataset. Loic Themyr, Clément Rambour, Nicolas Thome, Toby Collins, Alexandre Hostettler |
WACV | 3 |
| 2023 | Deep Time Series Forecasting With Shape and Temporal CriteriaabstractThis paper addresses the problem of multi-step time series forecasting for non-stationary signals that can present sudden changes. Current state-of-the-art deep learning forecasting methods, often trained with variants of the MSE, lack the ability to provide sharp predictions in deterministic and probabilistic contexts. To handle these challenges, we propose to incorporate shape and temporal criteria in the training objective of deep models. We define shape and temporal similarities and dissimilarities, based on a smooth relaxation of Dynamic Time Warping (DTW) and Temporal Distortion Index (TDI), that enable to build differentiable loss functions and positive semi-definite (PSD) kernels. With these tools, we introduce DILATE (DIstortion Loss including shApe and TimE), a new objective for deterministic forecasting, that explicitly incorporates two terms supporting precise shape and temporal change detection. For probabilistic forecasting, we introduce STRIPE++ (Shape and Time diverRsIty in Probabilistic for Ecasting), a framework for providing a set of sharp and diverse forecasts, where the structured shape and time diversity is enforced with a determinantal point process (DPP) diversity loss. Extensive experiments and ablations studies on synthetic and real-world datasets confirm the benefits of leveraging shape and time features in time series forecasting. Vincent Le Guen, Nicolas Thome |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Complementing Brightness Constancy with Deep Networks for Optical Flow Prediction
Vincent Le Guen, Clément Rambour, Nicolas Thome |
ECCV (21) | 3 |
| 2022 | Hierarchical Average Precision Training for Pertinent Image Retrieval
Elias Ramzi, Nicolas Audebert, Nicolas Thome, Clément Rambour, Xavier Bitot |
ECCV (14) | 3 |
| 2022 | Diverse Probabilistic Trajectory Forecasting with Admissibility ConstraintsabstractPredicting multiple trajectories for road users is important for automated driving systems: ego-vehicle motion planning indeed requires a clear view of the possible motions of the surrounding agents. However, the generative models used for multiple-trajectory forecasting suffer from a lack of diversity in their proposals. To avoid this form of collapse, we propose a novel method for structured prediction of diverse trajectories. To this end, we complement an underlying pretrained generative model with a diversity component, based on a determinantal point process (DPP). We balance and structure this diversity with the inclusion of knowledge-based quality constraints, independent from the underlying generative model. We combine these two novel components with a gating operation, ensuring that the predictions are both diverse and within the drivable area. We demonstrate on the nuScenes driving dataset the relevance of our compound approach, which yields significant improvements in the diversity and the quality of the generated trajectories. Laura Calem, Hédi Ben-Younes, Patrick Pérez, Nicolas Thome |
ICPR | 4 |
| 2022 | Swapping Semantic Contents for Mixing ImagesabstractDeep architecture have proven capable of solving many tasks provided a sufficient amount of labeled data. In fact, the amount of available labeled data has become the principal bottleneck in low label settings such as Semi-Supervised Learning. Mixing Data Augmentations do not typically yield new labeled samples, as indiscriminately mixing contents creates between-class samples. In this work, we introduce the SciMix framework that can learn to replace the global semantic content from one sample. By teaching a StyleGan generator to embed a semantic style code into image backgrounds, we obtain new mixing scheme for data augmentation. We then demonstrate that SciMix yields novel mixed samples that inherit many characteristics from their non-semantic parents. Afterwards, we verify those samples can be used to improve the performance semi-supervised frameworks like Mean Teacher or Fixmatch, and even fully supervised learning on a small labeled dataset. Rémy Sun, Clément Masson, Gilles Hénaff, Nicolas Thome, Matthieu Cord |
ICPR | 4 |
| 2022 | Confidence Estimation via Auxiliary ModelsabstractReliably quantifying the confidence of deep neural classifiers is a challenging yet fundamental requirement for deploying such models in safety-critical applications. In this paper, we introduce a novel target criterion for model confidence, namely the true class probability (TCP). We show that TCP offers better properties for confidence estimation than standard maximum class probability (MCP). Since the true class is by essence unknown at test time, we propose to learn TCP criterion from data with an auxiliary model, introducing a specific learning scheme adapted to this context. We evaluate our approach on the task of failure prediction and of self-training with pseudo-labels for domain adaptation, which both necessitate effective confidence estimates. Extensive experiments are conducted for validating the relevance of the proposed approach in each task. We study various network architectures and experiment with small and large datasets for image classification and semantic segmentation. In every tested benchmark, our approach outperforms strong baselines. Charles Corbière, Nicolas Thome, Antoine Saporta, Matthieu Cord, Patrick Pérez |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Augmenting Physical Models with Deep Networks for Complex Dynamics Forecasting
Vincent Le Guen, Jérémie Donà, Emmanuel de Bézenac, Ibrahim Ayed, Nicolas Thome, Patrick Gallinari |
ICLR | 6 |
| 2021 | Robust and Decomposable Average Precision for Image RetrievalabstractIn image retrieval, standard evaluation metrics rely on score ranking, e.g. average precision (AP). In this paper, we introduce a method for robust and decomposable average precision (ROADMAP) addressing two major challenges for end-to-end training of deep neural networks with AP: non-differentiability and non-decomposability.Firstly, we propose a new differentiable approximation of the rank function, which provides an upper bound of the AP loss and ensures robust training. Secondly, we design a simple yet effective loss function to reduce the decomposability gap between the AP in the whole training set and its averaged batch approximation, for which we provide theoretical guarantees.Extensive experiments conducted on three image retrieval datasets show that ROADMAP outperforms several recent AP approximation methods and highlight the importance of our two contributions. Finally, using ROADMAP for training deep models yields very good performances, outperforming state-of-the-art results on the three datasets.Code and instructions to reproduce our results will be made publicly available at https://github.com/elias-ramzi/ROADMAP. Elias Ramzi, Nicolas Thome, Clément Rambour, Nicolas Audebert, Xavier Bitot |
NeurIPS | 2 |
| 2020 | Disentangling Physical Dynamics From Unknown Factors for Unsupervised Video PredictionabstractLeveraging physical knowledge described by partial differential equations (PDEs) is an appealing way to improve unsupervised video forecasting models. Since physics is too restrictive for describing the full visual content of generic video sequences, we introduce PhyDNet, a two-branch deep architecture, which explicitly disentangles PDE dynamics from unknown complementary information. A second contribution is to propose a new recurrent physical cell (PhyCell), inspired from data assimilation techniques, for performing PDE-constrained prediction in latent space. Extensive experiments conducted on four various datasets show the ability of PhyDNet to outperform state-of-the-art methods. Ablation studies also highlight the important gain brought out by both disentanglement and PDE-constrained prediction. Finally, we show that PhyDNet presents interesting features for dealing with missing data and long-term forecasting. Vincent Le Guen, Nicolas Thome |
CVPR | 2 |
| 2020 | Probabilistic Time Series Forecasting with Shape and Temporal DiversityabstractProbabilistic forecasting consists in predicting a distribution of possible future outcomes. In this paper, we address this problem for non-stationary time series, which is very challenging yet crucially important. We introduce the STRIPE model for representing structured diversity based on shape and time features, ensuring both probable predictions while being sharp and accurate. STRIPE is agnostic to the forecasting model, and we equip it with a diversification mechanism relying on determinantal point processes (DPP). We introduce two DPP kernels for modelling diverse trajectories in terms of shape and time, which are both differentiable and proved to be positive semi-definite. To have an explicit control on the diversity structure, we also design an iterative sampling mechanism to disentangle shape and time representations in the latent space. Experiments carried out on synthetic datasets show that STRIPE significantly outperforms baseline methods for representing diversity, while maintaining accuracy of the forecasting model. We also highlight the relevance of the iterative sampling scheme and the importance to use different criteria for measuring quality and diversity. Finally, experiments on real datasets illustrate that STRIPE is able to outperform state-of-the-art probabilistic forecasting approaches in the best sample prediction. Vincent Le Guen, Nicolas Thome |
NeurIPS | 2 |
| 2019 | BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship DetectionabstractMultimodal representation learning is gaining more and more interest within the deep learning community. While bilinear models provide an interesting framework to find subtle combination of modalities, their number of parameters grows quadratically with the input dimensions, making their practical implementation within classical deep learning pipelines challenging. In this paper, we introduce BLOCK, a new multimodal fusion based on the block-superdiagonal tensor decomposition. It leverages the notion of block-term ranks, which generalizes both concepts of rank and mode ranks for tensors, already used for multimodal fusion. It allows to define new ways for optimizing the tradeoff between the expressiveness and complexity of the fusion model, and is able to represent very fine interactions between modalities while maintaining powerful mono-modal representations. We demonstrate the practical interest of our fusion model by using BLOCK for two challenging tasks: Visual Question Answering (VQA) and Visual Relationship Detection (VRD), where we design end-to-end learnable architectures for representing relevant interactions between modalities. Through extensive experiments, we show that BLOCK compares favorably with respect to state-of-the-art multimodal fusion models for both VQA and VRD tasks. Our code is available at https://github.com/Cadene/block.bootstrap.pytorch. Hédi Ben-Younes, Rémi Cadène, Nicolas Thome, Matthieu Cord |
AAAI | 3 |
| 2019 | MUREL: Multimodal Relational Reasoning for Visual Question AnsweringabstractMultimodal attentional networks are currently state-of-the-art models for Visual Question Answering (VQA) tasks involving real images. Although attention allows to focus on the visual content relevant to the question, this simple mechanism is arguably insufficient to model complex reasoning features required for VQA or other high-level tasks. In this paper, we propose MuRel, a multimodal relational network which is learned end-to-end to reason over real images. Our first contribution is the introduction of the MuRel cell, an atomic reasoning primitive representing interactions between question and image regions by a rich vectorial representation, and modeling region relations with pairwise combinations. Secondly, we incorporate the cell into a full MuRel network, which progressively refines visual and question interactions, and can be leveraged to define visualization schemes finer than mere attention maps. We validate the relevance of our approach with various ablation studies, and show its superiority to attention-based methods on three datasets: VQA 2.0, VQA-CP v2 and TDIUC. Our final MuRel network is competitive to or outperforms state-of-the-art results in this challenging context. Our code is available: github.com/Cadene/murel.bootstrap.pytorch Rémi Cadène, Hédi Ben-Younes, Matthieu Cord, Nicolas Thome |
CVPR | 4 |
| 2019 | DiscoNet: Shapes Learning on Disconnected Manifolds for 3D EditingabstractEditing 3D models is a very challenging task, as it requires complex interactions with the 3D shape to reach the targeted design, while preserving the global consistency and plausibility of the shape. In this work, we present an intelligent and user-friendly 3D editing tool, where the edited model is constrained to lie onto a learned manifold of realistic shapes. Due to the topological variability of real 3D models, they often lie close to a disconnected manifold, which cannot be learned with a common learning algorithm. Therefore, our tool is based on a new deep learning model, DiscoNet, which extends 3D surface autoencoders in two ways. Firstly, our deep learning model uses several autoencoders to automatically learn each connected component of a disconnected manifold, without any supervision. Secondly, each autoencoder infers the output 3D surface by deforming a pre-learned 3D template specific to each connected component. Both advances translate into improved 3D synthesis, thus enhancing the quality of our 3D editing tool. Éloi Mehr, Ariane Jourdan, Nicolas Thome, Matthieu Cord, Vincent Guitteny |
ICCV | 3 |
| 2019 | Addressing Failure Prediction by Learning Model ConfidenceabstractAssessing reliably the confidence of a deep neural net and predicting its failures is of primary importance for the practical deployment of these models. In this paper, we propose a new target criterion for model confidence, corresponding to the True Class Probability (TCP). We show how using the TCP is more suited than relying on the classic Maximum Class Probability (MCP). We provide in addition theoretical guarantees for TCP in the context of failure prediction. Since the true class is by essence unknown at test time, we propose to learn TCP criterion on the training set, introducing a specific learning scheme adapted to this context. Extensive experiments are conducted for validating the relevance of the proposed approach. We study various network architectures, small and large scale datasets for image classification and semantic segmentation. We show that our approach consistently outperforms several strong methods, from MCP to Bayesian uncertainty, as well as recent approaches specifically designed for failure prediction. Charles Corbière, Nicolas Thome, Avner Bar-Hen, Matthieu Cord, Patrick Pérez |
NeurIPS | 2 |
| 2019 | Shape and Time Distortion Loss for Training Deep Time Series Forecasting ModelsabstractThis paper addresses the problem of time series forecasting for non-stationary signals and multiple future steps prediction. To handle this challenging task, we introduce DILATE (DIstortion Loss including shApe and TimE), a new objective function for training deep neural networks. DILATE aims at accurately predicting sudden changes, and explicitly incorporates two terms supporting precise shape and temporal change detection. We introduce a differentiable loss function suitable for training deep neural nets, and provide a custom back-prop implementation for speeding up optimization. We also introduce a variant of DILATE, which provides a smooth generalization of temporally-constrained Dynamic TimeWarping (DTW). Experiments carried out on various non-stationary datasets reveal the very good behaviour of DILATE compared to models trained with the standard Mean Squared Error (MSE) loss function, and also to DTW and variants. DILATE is also agnostic to the choice of the model, and we highlight its benefit for training fully connected networks as well as specialized recurrent architectures, showing its capacity to improve over state-of-the-art trajectory forecasting approaches. Vincent Le Guen, Nicolas Thome |
NeurIPS | 2 |
| 2019 | End-to-End Learning of Latent Deformable Part-Based Representations for Object Detection
Taylor Mordan, Nicolas Thome, Gilles Hénaff, Matthieu Cord |
Int. J. Comput. Vis. | 2 |
| 2019 | Distributed optimization for deep learning with gossip exchange
Michaël Blot, David Picard, Nicolas Thome, Matthieu Cord |
Neurocomputing | 3 |
| 2019 | Exploiting Negative Evidence for Deep Latent Structured ModelsabstractThe abundance of image-level labels and the lack of large scale detailed annotations (e.g. bounding boxes, segmentation masks) promotes the development of weakly supervised learning (WSL) models. In this work, we propose a novel framework for WSL of deep convolutional neural networks dedicated to learn localized features from global image-level annotations. The core of the approach is a new latent structured output model equipped with a pooling function which explicitly models negative evidence, e.g. a cow detector should strongly penalize the prediction of the bedroom class. We show that our model can be trained end-to-end for different visual recognition tasks: multi-class and multi-label classification, and also structured average precision (AP) ranking. Extensive experiments highlight the relevance of the proposed method: our model outperforms state-of-the art results on six datasets. We also show that our framework can be used to improve the performance of state-of-the-art deep models for large scale image classification on ImageNet. Finally, we evaluate our model for weakly supervised tasks: in particular, a direct adaptation for weakly supervised segmentation provides a very competitive model. Thibaut Durand, Nicolas Thome, Matthieu Cord |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Manifold Learning in Quotient SpacesabstractWhen learning 3D shapes we are usually interested in their intrinsic geometry rather than in their orientation. To deal with the orientation variations the usual trick consists in augmenting the data to exhibit all possible variability, and thus let the model learn both the geometry as well as the rotations. In this paper we introduce a new auto-encoder model for encoding and synthesis of 3D shapes. To get rid of undesirable input variability our model learns a manifold in a quotient space of the input space. Typically, we propose to quotient the space of 3D models by the action of rotations. Thus, our quotient auto-encoder allows to directly learn in the space of interest, ignoring side information. This is reflected in better performances on reconstruction and interpolation tasks, as our experiments show that our model outperforms a vanilla auto-encoder on the well-known Shapenet dataset. Moreover, our model learns a rotation-invariant representation, leading to interesting results in shapes co-alignment. Finally, we extend our quotient auto-encoder to quotient by non-rigid transformations. Éloi Mehr, André Lieutier, Fernando Sanchez Bermudez, Vincent Guitteny, Nicolas Thome, Matthieu Cord |
CVPR | 5 |
| 2018 | HybridNet: Classification and Reconstruction Cooperation for Semi-supervised Learning
Thomas Robert 0001, Nicolas Thome, Matthieu Cord |
ECCV (7) | 2 |
| 2018 | Shade: Information-Based Regularization for Deep LearningabstractRegularization is a big issue for training deep neural networks. In this paper, we propose a new information-theory-based regularization scheme named SHADE for SHAnnon DEcay. The originality of the approach is to define a prior based on conditional entropy, which explicitly decouples the learning of invariant representations in the regularizer and the learning of correlations between inputs and labels in the data fitting term. Our second contribution is to derive a stochastic version of the regularizer compatible with deep learning, resulting in a tractable training scheme. We empirically validate the efficiency of our approach to improve classification performances compared to standard regularization schemes on several standard architectures. Michaël Blot, Thomas Robert 0001, Nicolas Thome, Matthieu Cord |
ICIP | 3 |
| 2018 | Revisiting Multi-Task Learning with ROCK: a Deep Residual Auxiliary Block for Visual DetectionabstractMulti-Task Learning (MTL) is appealing for deep learning regularization. In this paper, we tackle a specific MTL context denoted as primary MTL, where the ultimate goal is to improve the performance of a given primary task by leveraging several other auxiliary tasks. Our main methodological contribution is to introduce ROCK, a new generic multi-modal fusion block for deep learning tailored to the primary MTL context. ROCK architecture is based on a residual connection, which makes forward prediction explicitly impacted by the intermediate auxiliary representations. The auxiliary predictor's architecture is also specifically designed to our primary MTL context, by incorporating intensive pooling operators for maximizing complementarity of intermediate representations. Extensive experiments on NYUv2 dataset (object detection with scene classification, depth prediction, and surface normal estimation as auxiliary tasks) validate the relevance of the approach and its superiority to flat MTL approaches. Our method outperforms state-of-the-art object detection models on NYUv2 dataset by a large margin, and is also able to handle large-scale heterogeneous inputs (real and synthetic images) with missing annotation modalities. Taylor Mordan, Nicolas Thome, Gilles Hénaff, Matthieu Cord |
NeurIPS | 2 |
| 2018 | Cross-Modal Retrieval in the Cooking Context: Learning Semantic Text-Image EmbeddingsabstractDesigning powerful tools that support cooking activities has rapidly gained popularity due to the massive amounts of available data, as well as recent advances in machine learning that are capable of analyzing them. In this paper, we propose a cross-modal retrieval model aligning visual and textual data (like pictures of dishes and their recipes) in a shared representation space. We describe an effective learning scheme, capable of tackling large-scale problems, and validate it on the Recipe1M dataset containing nearly 1 million picture-recipe pairs. We show the effectiveness of our approach regarding previous state-of-the-art models and present qualitative results over computational cooking use cases. Micael Carvalho, Rémi Cadène, David Picard, Laure Soulier, Nicolas Thome, Matthieu Cord |
SIGIR | 5 |
| 2018 | Classifying low-resolution images by integrating privileged information in deep CNNs
Marion Chevalier, Nicolas Thome, Gilles Hénaff, Matthieu Cord |
Pattern Recognit. Lett. | 2 |
| 2018 | SyMIL: MinMax Latent SVM for Weakly Labeled DataabstractDesigning powerful models able to handle weakly labeled data are a crucial problem in machine learning. In this paper, we propose a new multiple instance learning (MIL) framework. Examples are represented as bags of instances, but we depart from standard MIL assumptions by introducing a symmetric strategy (SyMIL) that seeks discriminative instances in positive and negative bags. The idea is to use the instance the most distant from the hyper-plan to classify the bag. We provide a theoretical analysis featuring the generalization properties of our model. We derive a large margin formulation of our problem, which is cast as a difference of convex functions, and optimized using concave-convex procedure. We provide a primal version optimizing with stochastic subgradient descent and a dual version optimizing with one-slack cutting-plane. Successful experimental results are reported on standard MIL and weakly supervised object detection data sets: SyMIL significantly outperforms competitive methods (mi/MI/Latent-SVM), and gives very competitive performance compared to state-of-the-art works. We also analyze the selected instances of symmetric and asymmetric approaches on weakly supervised object detection and text classification tasks. Finally, we show complementarity of SyMIL with recent works on learning with label proportions on standard MIL data sets. Thibaut Durand, Nicolas Thome, Matthieu Cord |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Deformable Part-based Fully Convolutional Network for Object Detection
Taylor Mordan, Nicolas Thome, Gilles Hénaff, Matthieu Cord |
BMVC | 2 |
| 2017 | WILDCAT: Weakly Supervised Learning of Deep ConvNets for Image Classification, Pointwise Localization and SegmentationabstractThis paper introduces WILDCAT, a deep learning method which jointly aims at aligning image regions for gaining spatial invariance and learning strongly localized features. Our model is trained using only global image labels and is devoted to three main visual recognition tasks: image classification, weakly supervised object localization and semantic segmentation. WILDCAT extends state-of-the-art Convolutional Neural Networks at three main levels: the use of Fully Convolutional Networks for maintaining spatial resolution, the explicit design in the network of local features related to different class modalities, and a new way to pool these features to provide a global image prediction required for weakly supervised training. Extensive experiments show that our model significantly outperforms state-of-the-art methods. Thibaut Durand, Taylor Mordan, Nicolas Thome, Matthieu Cord |
CVPR | 3 |
| 2017 | MUTAN: Multimodal Tucker Fusion for Visual Question AnsweringabstractBilinear models provide an appealing framework for mixing and merging information in Visual Question Answering (VQA) tasks. They help to learn high level associations between question meaning and visual concepts in the image, but they suffer from huge dimensionality issues. We introduce MUTAN, a multimodal tensor-based Tucker decomposition to efficiently parametrize bilinear interactions between visual and textual representations. Additionally to the Tucker framework, we design a low-rank matrix-based decomposition to explicitly constrain the interaction rank. With MUTAN, we control the complexity of the merging scheme while keeping nice interpretable fusion relations. We show how the Tucker decomposition framework generalizes some of the latest VQA architectures, providing state-of-the-art results. Hédi Ben-Younes, Rémi Cadène, Matthieu Cord, Nicolas Thome |
ICCV | 4 |
| 2017 | Learning a Distance Metric from Relative Comparisons between Quadruplets of Images
Marc T. Law, Nicolas Thome, Matthieu Cord |
Int. J. Comput. Vis. | 2 |
| 2017 | Gaze latent support vector machine for image classification improved by weakly supervised region selection
Xin Wang 0053, Nicolas Thome, Matthieu Cord |
Pattern Recognit. | 2 |
| 2016 | WELDON: Weakly Supervised Learning of Deep Convolutional Neural NetworksabstractIn this paper, we introduce a novel framework for WEakly supervised Learning of Deep cOnvolutional neural Networks (WELDON). Our method is dedicated to automatically selecting relevant image regions from weak annotations, e.g. global image labels, and encompasses the following contributions. Firstly, WELDON leverages recent improvements on the Multiple Instance Learning paradigm, i.e. negative evidence scoring and top instance selection. Secondly, the deep CNN is trained to optimize Average Precision, and fine-tuned on the target dataset with efficient computations due to convolutional feature sharing. A thorough experimental validation shows that WELDON outperforms state-of-the-art results on six different datasets. Thibaut Durand, Nicolas Thome, Matthieu Cord |
CVPR | 2 |
| 2016 | Max-min convolutional neural networks for image classificationabstractConvolutional neural networks (CNN) are widely used in computer vision, especially in image classification. However, the way in which information and invariance properties are encoded through in deep CNN architectures is still an open question. In this paper, we propose to modify the standard convolutional block of CNN in order to transfer more information layer after layer while keeping some invariance within the network. Our main idea is to exploit both positive and negative high scores obtained in the convolution maps. This behavior is obtained by modifying the traditional activation function step before pooling. We are doubling the maps with specific activations functions, called MaxMin strategy, in order to achieve our pipeline. Extensive experiments on two classical datasets, MNIST and CIFAR-10, show that our deep MaxMin convolutional net outperforms standard CNN. Michaël Blot, Matthieu Cord, Nicolas Thome |
ICIP | 3 |
| 2016 | Deep Neural Networks Under StressabstractIn recent years, deep architectures have been used for transfer learning with state-of-the-art performance in many datasets. The properties of their features remain, however, largely unstudied under the transfer perspective. In this work, we present an extensive analysis of the resiliency of feature vectors extracted from deep models, with special focus on the trade-off between performance and compression rate. By introducing perturbations to image descriptions extracted from a deep convolutional neural network, we change their precision and number of dimensions, measuring how it affects the final score. We show that deep features are more robust to these disturbances when compared to classical approaches, achieving a compression rate of 98.4%, while losing only 0.88% of their original score for Pascal VOC 2007. Micael Carvalho, Matthieu Cord, Sandra Eliza Fontes de Avila, Nicolas Thome, Eduardo Valle |
ICIP | 4 |
| 2016 | Gaze latent support vector machine for image classificationabstractThis paper deals with image categorization from weak supervision, e.g. global image labels. We propose to improve the region selection performed in latent variable models such as Latent Support Vector Machine (LSVM) by leveraging human eye movement features collected from an eye-tracker device. We introduce a new model, Gaze Latent Support Vector Machine (G-LSVM), whose region selection during training is biased toward regions with a large gaze density ratio. On this purpose, the training objective is enriched with a gaze loss, from which we derive a convex upper bound, leading to a Concave-Convex Procedure (CCCP) optimization scheme. Experiments show that G-LSVM significantly outperforms LSVM in both object detection and action recognition problems on PASCAL VOC 2012. We also show that our G-LSVM is even slightly better than a model trained from bounding box annotations, while gaze labels are much cheaper to collect. Xin Wang 0053, Nicolas Thome, Matthieu Cord |
ICIP | 2 |
| 2015 | MANTRA: Minimum Maximum Latent Structural SVM for Image Classification and RankingabstractIn this work, we propose a novel Weakly Supervised Learning (WSL) framework dedicated to learn discriminative part detectors from images annotated with a global label. Our WSL method encompasses three main contributions. Firstly, we introduce a new structured output latent variable model, Minimum mAximum lateNt sTRucturAl SVM (MANTRA), which prediction relies on a pair of latent variables: h+(resp. h-) provides positive (resp. negative) evidence for a given output y. Secondly, we instantiate MANTRA for two different visual recognition tasks: multi-class classification and ranking. For ranking, we propose efficient solutions to exactly solve the inference and the loss-augmented problems. Finally, extensive experiments highlight the relevance of the proposed method: MANTRA outperforms state-of-the art results on five different datasets. Thibaut Durand, Nicolas Thome, Matthieu Cord |
ICCV | 2 |
| 2015 | Exemplar based metric learning for robust visual localizationabstractThis paper presents an exemplar based metric learning framework dedicated to robust visual localization in complex scenes, e.g. street images. The proposed framework learns off-line a specific (local) metric for each image of the database, so that the distance between a database image and a query image representing the same scene is smaller than the distance between the current image and other images of the database. To achieve this goal, we generate geometric and photometric transformations as proxies for query images. From the generated constraints, the learning problem is cast as a convex optimization problem over the cone of positive semi-definite matrices, which is efficiently solved using a projected gradient descent scheme. Successful experiments, conducted using a freely available geo-referenced image database, reveal that the proposed method significantly improves results over the metric in the input space, while being as efficient at test time. In addition, we show that the model learns discriminating features for the localization task, and is able to gain invariance to meaningful transformations. Cédric Le Barz, Nicolas Thome, Matthieu Cord, Stéphane Herbin, Martial Sanfourche |
ICIP | 2 |
| 2015 | LR-CNN for fine-grained classification with varying resolutionabstractIn this work, we present an extended study of image representations for fine-grained classification with respect to image resolution. Understudied in literature, this parameter yet presents many practical and theoretical interests, e.g. in embedded systems where restricted computational resources prevent treating high-resolution images. It is thus interesting to figure out which representation provides the best results in this particular context. On this purpose, we evaluate Fisher Vectors and deep representations on two significant finegrained oriented datasets: FGVC Aircraft [1] and PPMI [2]. We also introduce LR-CNN, a deep structure designed for classification of low-resolution images with strong semantic content. This net provides rich compact features and outperforms both pre-trained deep features and Fisher Vectors. Marion Chevalier, Nicolas Thome, Matthieu Cord, Jérôme Fournier, Gilles Hénaff, Elodie Dusch |
ICIP | 2 |
| 2014 | Fantope Regularization in Metric LearningabstractThis paper introduces a regularization method to explicitly control the rank of a learned symmetric positive semidefinite distance matrix in distance metric learning. To this end, we propose to incorporate in the objective function a linear regularization term that minimizes the k smallest eigenvalues of the distance matrix. It is equivalent to minimizing the trace of the product of the distance matrix with a matrix in the convex hull of rank-k projection matrices, called a Fantope. Based on this new regularization method, we derive an optimization scheme to efficiently learn the distance matrix. We demonstrate the effectiveness of the method on synthetic and challenging real datasets of face verification and image classification with relative attributes, on which our method outperforms state-of-the-art metric learning algorithms. Marc T. Law, Nicolas Thome, Matthieu Cord |
CVPR | 2 |
| 2014 | Semantic pooling for image categorization using multiple kernel learningabstractIn this paper, we propose a new method for taking into account the spatial information in image categorization. More specifically, we remove the loss of spatial information in Bag of Words related methods by computing the image signature over specific regions selected by object detectors. We propose to select the detectors using Multiple Kernel Learning techniques. We carry out experiments on the well known VOC 2007 dataset, and show our semantic pooling obtains promising results. Thibaut Durand, David Picard, Nicolas Thome, Matthieu Cord |
ICIP | 3 |
| 2014 | Incremental learning of latent structural SVM for weakly supervised image classificationabstractVisual learning with weak supervision is a promising research area, since it offers the possibility to build large image datasets at reasonable cost. In this paper, we address the problem of weakly supervised object detection, where the goal is to predict the label of the image using object position as latent variable. We propose a new method that builds upon the Latent Structural SVM (LSSVM) formalism. Specifically, we introduce an original coarse-to-fine approach that limits the evolution of the latent parameter subspace. This incremental strategy drives the learning towards better solutions, providing a model with increased predictive accuracy. In addition, this leads to a significant speed up during learning and inference compared to standard sliding window methods. Experiments carried out on Mammal dataset validate the good performances and fast training of the method compared to state-of-the-art works. Thibaut Durand, Nicolas Thome, Matthieu Cord, David Picard |
ICIP | 2 |
| 2014 | SnooperText: A text detection system for automatic indexing of urban scenes
Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J. Leite, Jorge Stolfi |
Comput. Vis. Image Underst. | 2 |
| 2014 | Learning Deep Hierarchical Visual Feature CodingabstractIn this paper, we propose a hybrid architecture that combines the image modeling strengths of the bag of words framework with the representational power and adaptability of learning deep architectures. Local gradient-based descriptors, such as SIFT, are encoded via a hierarchical coding scheme composed of spatial aggregating restricted Boltzmann machines (RBM). For each coding layer, we regularize the RBM by encouraging representations to fit both sparse and selective distributions. Supervised fine-tuning is used to enhance the quality of the visual representation for the categorization task. We performed a thorough experimental evaluation using three image categorization data sets. The hierarchical coding scheme achieved competitive categorization accuracies of 79.7% and 86.4% on the Caltech-101 and 15-Scenes data sets, respectively. The visual representations learned are compact and the model's inference is fast, as compared with sparse coding methods. The low-level representations of descriptors that were learned using this method result in generic features that we empirically found to be transferrable between different image data sets. Further analysis reveal the significance of supervised fine-tuning when the architecture has two layers of representations as opposed to a single layer. Hanlin Goh, Nicolas Thome, Matthieu Cord, Joo-Hwee Lim |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Dynamic Scene Classification: Learning Motion Descriptors with Slow Features AnalysisabstractIn this paper, we address the challenging problem of categorizing video sequences composed of dynamic natural scenes. Contrarily to previous methods that rely on handcrafted descriptors, we propose here to represent videos using unsupervised learning of motion features. Our method encompasses three main contributions: 1) Based on the Slow Feature Analysis principle, we introduce a learned local motion descriptor which represents the principal and more stable motion components of training videos. 2) We integrate our local motion feature into a global coding/pooling architecture in order to provide an effective signature for each video sequence. 3) We report state of the art classification performances on two challenging natural scenes data sets. In particular, an outstanding improvement of 11% in classification score is reached on a data set introduced in 2012. Christian Theriault, Nicolas Thome, Matthieu Cord |
CVPR | 2 |
| 2013 | Quadruplet-Wise Image Similarity LearningabstractThis paper introduces a novel similarity learning framework. Working with inequality constraints involving quadruplets of images, our approach aims at efficiently modeling similarity from rich or complex semantic label relationships. From these quadruplet-wise constraints, we propose a similarity learning framework relying on a convex optimization scheme. We then study how our metric learning scheme can exploit specific class relationships, such as class ranking (relative attributes), and class taxonomy. We show that classification using the learned metrics gets improved performance over state-of-the-art methods on several datasets. We also evaluate our approach in a new application to learn similarities between web page screenshots in a fully unsupervised way. Marc T. Law, Nicolas Thome, Matthieu Cord |
ICCV | 2 |
| 2013 | Image classification using object detectorsabstractImage categorization is one of the most competitive topic in computer vision and image processing. In this paper, we propose to use trained object and region detectors to represent the visual content of each image. Compared to similar methods found in the literature, our method encompasses two main areas of novelty: introducing a new spatial pooling formalism and designing a late fusion strategy for combining our representation with state-of-the art methods based on low-level descriptors, e.g. Fisher Vectors and BossaNova. Our experiments carried out in the challenging PASCAL VOC 2007 dataset reveal outstanding performances. When combined with low-level representations, we reach more than 67.6% in MAP, outperforming recently reported results in this dataset with a large margin. Thibaut Durand, Nicolas Thome, Matthieu Cord, Sandra Eliza Fontes de Avila |
ICIP | 2 |
| 2013 | Top-Down Regularization of Deep Belief NetworksabstractDesigning a principled and effective algorithm for learning deep architectures is a challenging problem. The current approach involves two training phases: a fully unsupervised learning followed by a strongly discriminative optimization. We suggest a deep learning strategy that bridges the gap between the two phases, resulting in a three-phase learning procedure. We propose to implement the scheme using a method to regularize deep belief networks with top-down information. The network is constructed from building blocks of restricted Boltzmann machines learned by combining bottom-up and top-down sampled signals. A global optimization procedure that merges samples from a forward bottom-up pass and a top-down pass is used. Experiments on the MNIST dataset show improvements over the existing algorithms for deep belief networks. Object recognition results on the Caltech-101 dataset also yield competitive results. Hanlin Goh, Nicolas Thome, Matthieu Cord, Joo-Hwee Lim |
NIPS | 2 |
| 2013 | Pooling in image representation: The visual codeword point of view
Sandra Eliza Fontes de Avila, Nicolas Thome, Matthieu Cord, Eduardo Valle, Arnaldo de Albuquerque Araújo |
Comput. Vis. Image Underst. | 2 |
| 2013 | JKernelMachines: a simple framework for kernel machine
David Picard, Nicolas Thome, Matthieu Cord |
J. Mach. Learn. Res. | 2 |
| 2013 | T-HOG: An effective gradient-based descriptor for single line text regions
Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J. Leite, Jorge Stolfi |
Pattern Recognit. | 2 |
| 2013 | Extended Coding and Pooling in the HMAX ModelabstractThis paper presents an extension of the HMAX model, a neural network model for image classification. The HMAX model can be described as a four-level architecture, with the first level consisting of multiscale and multiorientation local filters. We introduce two main contributions to this model. First, we improve the way the local filters at the first level are integrated into more complex filters at the last level, providing a flexible description of object regions and combining local information of multiple scales and orientations. These new filters are discriminative and yet invariant, two key aspects of visual classification. We evaluate their discriminative power and their level of invariance to geometrical transformations on a synthetic image set. Second, we introduce a multiresolution spatial pooling. This pooling encodes both local and global spatial information to produce discriminative image signatures. Classification results are reported on three image data sets: Caltech101, Caltech256, and fifteen scenes. We show significant improvements over previous architectures using a similar framework. Christian Theriault, Nicolas Thome, Matthieu Cord |
IEEE Trans. Image Process. | 2 |
| 2012 | Structural and visual comparisons for web page archivingabstractIn this paper, we propose a Web page archiving system that combines state-of-the-art comparison methods based on the source codes of Web pages, with computer vision techniques. To detect whether successive versions of a Web page are similar or not, our system is based on: (1) a combination of structural and visual comparison methods embedded in a statistical discriminative model, (2) a visual similarity measure designed for Web pages that improves change detection, (3) a supervised feature selection method adapted to Web archiving. We train a Support Vector Machine model with vectors of similarity scores between successive versions of pages. The trained model then determines whether two versions, defined by their vector of similarity scores, are similar or not. Experiments on real archives validate our approach. Marc T. Law, Nicolas Thome, Stéphane Gançarski, Matthieu Cord |
ACM Symposium on Document Engineering | 2 |
| 2012 | Unsupervised and Supervised Visual Codes with Restricted Boltzmann Machines
Hanlin Goh, Nicolas Thome, Matthieu Cord, Joo-Hwee Lim |
ECCV (5) | 2 |
| 2012 | Learning geometric combinations of Gaussian kernels with alternating Quasi-Newton algorithm
David Picard, Nicolas Thome, Matthieu Cord, Alain Rakotomamonjy |
ESANN | 2 |
| 2012 | Contextual detection of drawn symbols in old mapsabstractIn this paper, we tackle the problem of detecting drawn symbols in old maps. We propose a novel approach that combines powerful low level descriptors to represent the local content of the objects, and contextual features to overcome the local analysis ambiguity. Our contribution is two-fold. Firstly, we propose a novel contextual feature adapted to our problem, where the context is integrated at two levels. In a close neighborhood, a local analysis is carried out to remove visual ambiguities between symbols. In a larger extent, co-occurrence statistics between classes are stored. Secondly, we propose an entire processing chain for learning and detection. The proposed method is evaluated on real french maps from the 18thcentury. The experiments show the efficiency of the detection system, and validate the relevance of the proposed contextual feature to improve detection performances. Jonathan Guyomard, Nicolas Thome, Matthieu Cord, Thierry Artières |
ICIP | 2 |
| 2012 | Classification of Urban Scenes from Geo-referenced Images in Urban Street-View ContextabstractThis paper addresses the challenging problem of scene classification in street-view georeferenced images of urban environments. More precisely, the goal of this task is semantic image classification, consisting in predicting in a given image, the presence or absence of a pre-defined class (e.g. shops, vegetation, etc.). The approach is based on the BOSSA representation, which enriches the Bag of Words (BoW) model, in conjunction with the Spatial Pyramid Matching scheme and kernel-based machine learning techniques. The proposed method handles problems that arise in large scale urban environments due to acquisition conditions (static and dynamic objects/pedestrians) combined with the continuous acquisition of data along the vehicle's direction, the varying light conditions and strong occlusions (due to the presence of trees, traffic signs, cars, etc.) giving rise to high intra-class variability. Experiments were conducted on a large dataset of high resolution images collected from two main avenues from the 12th district in Paris and the approach shows promising results. Corina Iovan, David Picard, Nicolas Thome, Matthieu Cord |
ICMLA (2) | 3 |
| 2011 | BOSSA: Extended bow formalism for image classificationabstractIn image classification, the most powerful statistical learning approaches are based on the Bag-of-Words paradigm. In this article, we propose an extension of this formalism. Considering the Bag-of-Features, dictionary coding and pooling steps, we propose to focus on the pooling step. Instead of using the classical sum or max pooling strategies, we introduced a density function-based pooling strategy. This flexible formalism allows us to better represent the links between dictionary codewords and local descriptors in the resulting image signature. We evaluate our approach in two very challenging tasks of video and image classification, involving very high level semantic categories with large and nuanced visual diversity. Sandra Eliza Fontes de Avila, Nicolas Thome, Matthieu Cord, Eduardo Valle, Arnaldo de Albuquerque Araújo |
ICIP | 2 |
| 2011 | Learning invariant color features with sparse topographic restricted Boltzmann machinesabstractOur objective is to learn invariant color features directly from data via unsupervised learning. In this paper, we introduce a method to regularize restricted Boltzmann machines during training to obtain features that are sparse and topographically organized. Upon analysis, the features learned are Gabor-like and demonstrate a coding of orientation, spatial position, frequency and color that vary smoothly with the topography of the feature map. There is also differentiation between monochrome and color filters, with some exhibiting color-opponent properties. We also found that the learned representation is more invariant to affine image transformations and changes in illumination color. Hanlin Goh, Lukasz Kusmierz, Joo-Hwee Lim, Nicolas Thome, Matthieu Cord |
ICIP | 4 |
| 2011 | Snoopertrack: Text detection and tracking for outdoor videosabstractIn this work we introduced SnooperTrack, an algorithm for the automatic detection and tracking of text objects - such as store names, traffic signs, license plates, and advertisements - in videos of out door scenes. The purpose is to improve the performances of text detection process in still images by taking advantage of the temporal coherence in videos. We first propose an efficient tracking algorithm using particle filtering framework with original region descriptors. The second contribution is our strategy to merge tracked regions and new detections. We also propose an improved version of our previously published text detection algorithm in still images. Tests indicate that SnooperTrack is fast, robust, enable false positive suppression, and achieved great performances in complex videos of outdoor scenes. Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J. Leite, Jorge Stolfi |
ICIP | 2 |
| 2011 | Efficient Bag-of-Feature kernel representation for image similarity searchabstractAlthough “Bag-of-Features” image models have shown very good potential for object matching and image retrieval, such a complex data representation requires computationally expensive similarity measure evaluation. In this paper, we propose a framework unifying dictionary-based and kernel-based similarity functions that highlights the tradeoff between powerful data representation and eff cient similarity computation. On the basis of this formalism, we propose a new kernel-based similarity approach for Bag-of-Feature descriptions. We introduce a method for fast similarity search in large image databases. The conducted experiments prove that our approach is very competitive among State-of-the-art methods for similarity retrieval tasks. Frédéric Precioso, Matthieu Cord, David Gorisse, Nicolas Thome |
ICIP | 4 |
| 2011 | HMAX-S: Deep scale representation for biologically inspired image categorizationabstractThis paper presents an improvement on a biologically inspired network for image classification. Previous models have used a multi-scale and multi-orientation architecture to gain robustness to transformations and to extract complex visual features. Our contribution to this type of architecture resides in the building of complex visual features which are better tuned to images structures. We allow the network to build complex features with richer information in terms of the local scales of image structures. Our classification results show significant improvements over previous architectures using the same framework. Christian Theriault, Nicolas Thome, Matthieu Cord |
ICIP | 2 |
| 2011 | A cognitive and video-based approach for multinational License Plate Recognition
Nicolas Thome, Antoine Vacavant, Lionel Robinault, Serge Miguet |
Mach. Vis. Appl. | 1 |
| 2010 | Fast People Counting Using Head Detection From Skeleton GraphabstractIn this paper, we present a new method for counting people. This method is based on the head detection after a segmentation of the human body by skeleton graph process. The skeleton silhouette is computed and decomposed into a set of segments corresponding to the head , torso and limbs. This structure captures the minimal information about the skeleton shape. No assumption is made about the viewpoint, this is done after the head pose process. Several results present the efficiency of the labelling process , particularly its structural properties for the detection of heads within a crowd. A proposed method are evaluated on the crowd counting task in the PETS 2010 dataset. Djamel Merad, Kheir-Eddine Aziz, Nicolas Thome |
AVSS | 3 |
| 2010 | Fast People Counting Using Head Detection from Skeleton GraphabstractIn this paper, we present a new method for counting people. This method is based on the head detection after a segmentation of the human body by skeleton graph process. The skeleton silhouette is computed and decomposed into a set of segments corresponding to the head, torso and limbs. This structure captures the minimal information about the skeleton shape. No assumption is made about the viewpoint, this is done after the head pose process. Several results present the efficiency of the labelling process , particularly its structural properties for the detection of heads within a crowd. A proposed method has been tested with an experiment of counting the number of pedestrians passing in a specific area. Djamel Merad, Kheir-Eddine Aziz, Nicolas Thome |
AVSS | 3 |
| 2010 | Snoopertext: A multiresolution system for text detection in complex visual scenesabstractText detection in natural images remains a very challenging task. For instance, in an urban context, the detection is very difficult due to large variations in terms of shape, size, color, orientation, and the image may be blurred or have irregular illumination, etc. In this paper, we describe a robust and accurate multiresolution approach to detect and classify text regions in such scenarios. Based on generation/validation paradigm, we first segment images to detect character regions with a multiresolution algorithm able to manage large character size variations. The segmented regions are then filtered out using shape-based classification, and neighboring characters are merged to generate text hypotheses. A validation step computes a region signature based on texture analysis to reject false positives. We evaluate our algorithm in two challenging databases, achieving very good results. Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Jonathan Fabrizio, Beatriz Marcotegui |
ICIP | 2 |
| 2010 | An efficient system for combining complementary kernels in complex visual categorization tasksabstractRecently, increasing interest has been brought to improve image categorization performances by combining multiple descriptors. However, very few approaches have been proposed for combining features based on complementary aspects, and evaluating the performances in realistic databases. In this paper, we tackle the problem of combining different feature types (edge and color), and evaluate the performance gain in the very challenging VOC 2009 benchmark. Our contribution is three-fold. First, we propose new local color descriptors, unifying edge and color feature extraction into the “Bag Of Word” model. Second, we improve the Spatial Pyramid Matching (SPM) scheme for better incorporating spatial information into the similarity measurement. Last but not least, we propose a new combination strategy based on ℓ1Multiple Kernel Learning (MKL) that simultaneously learns individual kernel parameters and the kernel combination. Experiments prove the relevance of the proposed approach, which outperforms baseline combination methods while being computationally effective. David Picard, Nicolas Thome, Matthieu Cord |
ICIP | 2 |
| 2008 | A bottom-up, view-point invariant human detectorabstractWe propose a bottom-up human detector that can deal with arbitrary poses and viewpoints. Heads, limbs and torsos are individually detected, and an efficient assembly strategy is used to perform the human detection and the part segmentation. Firstly, a topological model is used to represent the structure of the human body, and the topologically equivalent configurations are ranked with additional priors. Promising results prove the approach efficiency for detecting people in low-resolution and compressed images. Nicolas Thome, Sebastien Ambellouis |
ICPR | 1 |
| 2008 | Learning articulated appearance models for tracking humans: A spectral graph matching approach
Nicolas Thome, Djamel Merad, Serge Miguet |
Signal Process. Image Commun. | 1 |
| 2008 | A Real-Time, Multiview Fall Detection System: A LHMM-Based ApproachabstractAutomatic detection of a falling person in video sequences has interesting applications in video-surveillance and is an important part of future pervasive home monitoring systems. In this paper, we propose a multiview approach to achieve this goal, where motion is modeled using a layered hidden Markov model (LHMM). The posture classification is performed by a fusion unit, merging the decision provided by the independently processing cameras in a fuzzy logic context. In each view, the fall detection is optimized in a given plane by performing a metric image rectification, making it possible to extract simple and robust features, and being convenient for real-time purpose. A theoretical analysis of the chosen descriptor enables us to define the optimal camera placement for detecting people falling in unspecified situations, and we prove that two cameras are sufficient in practice. Regarding event detection, the LHMM offers a principle way for solving the inference problem. Moreover, the hierarchical architecture decouples the motion analysis into different temporal granularity levels, making the algorithm able to detect very sudden changes, and robust to low-level steps errors. Nicolas Thome, Serge Miguet, Sebastien Ambellouis |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Human Body Part Labeling and Tracking Using Graph Matching TheoryabstractProperly labeling human body parts in video sequences is essential for robust tracking and motion interpretation frameworks. We propose to perform this task by using Graph Matching. The silhouette skeleton is computed and decomposed into a set of segments corresponding to the different limbs. A Graph capturing the topology of the segments is generated and matched against a 3D model of the human skeleton. The limb identification is carried out for each node of the graph, potentially leading to the absence of correspondence. The method captures the minimal information about the skeleton shape. No assumption about the viewpoint, the human pose, the geometry or the appearance of the limbs is done during the matching process, making the approach applicable to every configuration. Some correspondences that might be ambiguous only relying on topology are enforced by tracking each graph node over time. Several results present the efficiency of the labeling, particularly its robustness to limb detection errors that are likely to occur in real situations because of occlusions or low level system failures. Finally the relevance of the labeling in an overall tracking system is described. Nicolas Thome, Djamel Merad, Serge Miguet |
AVSS | 1 |
| 2006 | A HHMM-Based Approach for Robust Fall DetectionabstractAutomatic detection of a falling person in video sequences is an important part of future pervasive home monitoring systems. We propose here a robust method to achieve this goal. Motion is modeled by a hierarchical hidden Markov model (HHMM) whose first layer states are related to the orientation of the tracked person. Finding a consistent way for robustly linking the observation vector to the human poses is the heart of our contribution. In that sense, we carefully study the relationship between angles in the 3D world and their projection onto the image plane. After performing an initial image metric rectification, we derive theoretical properties making it possible to bound the error angle introduced by the image formation process for a standing posture. This allows us to confidently identify other poses as "non-standing" ones, and thus to robustly analyze pose sequences against a given motion model. Several results illustrate the efficiency of the algorithm by pointing out its ability to accurately recognize a person falling down from another walking or sitting, as well as its capacity to run in an unspecified configuration Nicolas Thome, Serge Miguet |
ICARCV | 1 |
| 2005 | A robust appearance model for tracking human motionsabstractWe propose an original method for tracking people based on the construction of a 2-D human appearance model. The general framework, which is a region-based tracking approach, is applicable to any type of object. We show how to specialize the method for taking advantage of the structural properties of the human body. We segment its visible parts, construct and update the appearance model. This latter one provides a discriminative feature capturing both color and shape properties of the different limbs, making it possible to recognize people after they have temporarily disappeared. The method does not make use of skin color detection, which allows us to perform tracking under any viewpoint. The only assumption for the recognition is the approximate viewpoint correspondence during the matching process between the different models. Several results in complex situations prove the efficiency of the algorithm, which runs in near real time. Finally, the model provides an important clue for further human motion analysis process. Nicolas Thome, Serge Miguet |
AVSS | 1 |