VLDB 2026 Research / reviewers in the wild / expert
Hervé Le Borgne
dblp:60/6993
· DBLP profile ↗
52ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0003-0520-8436ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 26 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 10 · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image GenerationabstractState-of-the-art text-to-image models produce visually impressive results but often struggle with precise alignment to text prompts, leading to missing critical elements or unintended blending of distinct concepts. We propose a novel approach that learns a high-success-rate distribution conditioned on a target prompt, ensuring that generated images faithfully reflect the corresponding prompts. Our method explicitly models the signal component during the denoising process, offering fine-grained control that mitigates over-optimization and out-of-distribution artifacts. Moreover, our framework is training-free and seamlessly integrates with both existing diffusion and flow matching architectures. It also supports additional conditioning modalities -- such as bounding boxes -- for enhanced spatial alignment. Extensive experiments demonstrate that our approach outperforms current state-of-the-art methods. Paul Grimal, Michaël Soumm, Hervé Le Borgne, Olivier Ferret, Akihiro Sugimoto |
AAAI | 3 |
| 2026 | Interpretable Image Recognition with Variable Number of Prototypes
Katia Benali, Bianca Vieru-Dimulescu, Marin Ferecatu, Hervé Le Borgne |
ICPR (7) | 4 |
| 2026 | Unique Step Refinement for Transformer-Based Generative Models
Paul Grimal, Hervé Le Borgne, Olivier Ferret |
ICPR (5) | 2 |
| 2026 | MVAT: Multi-View Aware Teacher for Weakly Supervised 3D Object DetectionabstractAnnotating 3D data remains a costly bottleneck for 3D object detection, motivating the development of weakly supervised annotation methods that rely on more accessible 2D box annotations. However, relying solely on 2D boxes introduces projection ambiguities since a single 2D box can correspond to multiple valid 3D poses. Furthermore, partial object visibility under a single viewpoint setting makes accurate 3D box estimation difficult. We propose MVAT, a novel framework that leverages temporal multi-view present in sequential data to address these challenges. Our approach aggregates object-centric point clouds across time to build 3D object representations as dense and complete as possible. A Teacher-Student distillation paradigm is employed: The Teacher network learns from single viewpoints but targets are derived from temporally aggregated static objects. Then the Teacher generates high quality pseudo-labels that the Student learns to predict from a single viewpoint for both static and moving objects. The whole framework incorporates a multi-view 2D projection loss to enforce consistency between predicted 3D boxes and all available 2D annotations. Experiments on the nuScenes and Waymo Open datasets demonstrate that MVAT achieves state-of-the-art performance for weakly supervised 3D object detection, significantly narrowing the gap with fully supervised methods without requiring any 3D box annotations. % \footnote{Code available upon acceptance} Our code is available in our public repository (\href{https://github.com/CEA-LIST/MVAT}{code}). Saad Lahlali, Alexandre Fournier-Montgieux, Nicolas Granger 0001, Hervé Le Borgne, Quoc Cuong Pham |
WACV | 4 |
| 2025 | Cross-Modal Distillation for 2D/3D Multi-Object Discovery from 2D MotionabstractObject discovery, which refers to the task of localizing objects without human annotations, has gained significant attention in 2D image analysis. However, despite this growing interest, it remains under-explored in 3D data, where approaches rely exclusively on 3D motion, despite its several challenges. In this paper, we present a novel framework that leverages advances in 2D object discovery which are based on 2D motion to exploit the advantages of such motion cues being more flexible and generalizable and to bridge the gap between 2D and 3D modalities. Our primary contributions are twofold: (i) we introduce DIOD-3D, the first baseline for multi-object discovery in 3D data using 2D motion, incorporating scene completion as an auxiliary task to enable dense object localization from sparse input data; (ii) we develop xMOD, a cross-modal training framework that integrates 2D and 3D data while always using 2D motion cues. xMOD employs a teacher-student training paradigm across the two modalities to mitigate confirmation bias by leveraging the domain gap. During inference, the model supports both RGB-only and point cloud-only inputs. Additionally, we propose a late-fusion technique tailored to our pipeline that further enhances performance when both modalities are available at inference. We evaluate our approach extensively on synthetic (TRIP-PD) and challenging real-world datasets (KITTI and Waymo). Notably, our approach yields a substantial performance improvement compared with the 2D object discovery state-of-the-art on all datasets with gains ranging from +8.7 to +15.1 in F 1@50 score. The code is available at https://github.com/CEA-LIST/xMOD Saad Lahlali, Sandra Kara, Hejer Ammar, Florian Chabot, Nicolas Granger 0001, Hervé Le Borgne, Quoc Cuong Pham |
CVPR | 6 |
| 2025 | Entity-Aware Cross-Modal Pretraining for Knowledge-Based Visual Question Answering
Omar Adjali, Olivier Ferret, Sahar Ghannay, Hervé Le Borgne |
ECIR (3) | 4 |
| 2025 | CaMiT: A Time-Aware Car Model Dataset for Classification and GenerationabstractAI systems must adapt to the evolving visual landscape, especially in domains where object appearance shifts over time. While prior work on time-aware vision models has primarily addressed commonsense-level categories, we introduce Car Models in Time (CaMiT). This fine-grained dataset captures the temporal evolution of this representative subset of technological artifacts. CaMiT includes 787K labeled samples of 190 car models (2007–2023) and 5.1M unlabeled samples (2005–2023), supporting supervised and self-supervised learning. We show that static pretraining on in-domain data achieves competitive performance with large-scale generalist models, offering a more resource-efficient solution. However, accuracy degrades when testing a year's models backward and forward in time. To address this, we evaluate CaMiT in a time-incremental classification setting, a realistic continual learning scenario with emerging, evolving, and disappearing classes. We investigate two mitigation strategies: time-incremental pretraining, which updates the backbone model, and time-incremental classifier learning, which updates the final classification layer, with positive results in both cases. Finally, we introduce time-aware image generation by consistently using temporal metadata during training. Results indicate improved realism compared to standard generation. CaMiT provides a rich resource for exploring temporal adaptation in a fine-grained visual context for discriminative and generative AI systems. Frédéric Lin, Biruk Abere Ambaw, Adrian Popescu 0001, Hejer Ammar, Romaric Audigier, Hervé Le Borgne |
NeurIPS | 6 |
| 2025 | Fairer Analysis and Demographically Balanced Face Generation for Fairer Face VerificationabstractFace recognition and verification are two computer vision tasks whose performances have advanced with the introduction of deep representations. However, ethical, legal, and technical challenges due to the sensitive nature of face data and biases in real-world training datasets hinder their development. Generative AI addresses privacy by creating fictitious identities, but fairness problems remain. Using the existing DCFace SOTA framework, we introduce a new controlled generation pipeline that improves fairness. Through classical fairness metrics and a proposed indepth statistical analysis based on logit models and ANOVA, we show that our generation pipeline improves fairness more than other bias mitigation approaches while slightly improving raw performance. Alexandre Fournier-Montgieux, Michaël Soumm, Adrian Popescu 0001, Bertrand Luvison, Hervé Le Borgne |
WACV | 5 |
| 2025 | ALPI: Auto-Labeller with Proxy Injection for 3D Object Detection using 2D Labels Onlyabstract3D object detection plays a crucial role in various applications such as autonomous vehicles, robotics and augmented reality. However, training 3D detectors requires a costly precise annotation, which is a hindrance to scaling annotation to large datasets. To address this challenge, we propose a weakly supervised 3D annotator that relies solely on 2D bounding box annotations from images, along with size priors. One major problem is that supervising a 3D detection model using only 2D boxes is not reliable due to ambiguities between different 3D poses and their identical 2D projection. We introduce a simple yet effective and generic solution: we build 3D proxy objects with annotations by construction and add them to the training dataset. Our method requires only size priors to adapt to new classes. To better align 2D supervision with 3D detection, our method ensures depth invariance with a novel expression of the 2D losses. Finally, to detect more challenging instances, our annotator follows an offline pseudo-labelling scheme which gradually improves its 3D pseudo-labels. Extensive experiments on the KITTI dataset demonstrate that our method not only performs on-par or above previous works on the Car category, but also achieves performance close to fully supervised methods on more challenging classes. We further demonstrate the effectiveness and robustness of our method by being the first to experiment on the more challenging nuScenes dataset. We additionally propose a setting where weak labels are obtained from a 2D detector pre-trained on MS-COCO instead of human annotations. The code is available at https://github.com/CEA-LIST/ALPI Saad Lahlali, Nicolas Granger 0001, Hervé Le Borgne, Quoc Cuong Pham |
WACV | 3 |
| 2024 | Multi-Level Information Retrieval Augmented Generation for Knowledge-based Visual Question AnsweringabstractThe Knowledge-Aware Visual Question Answering about Entity task aims to disambiguate entities using textual and visual information, as well as knowledge.It usually relies on two independent steps, information retrieval then reading comprehension, that do not benefit each other.Retrieval Augmented Generation (RAG) offers a solution by using generated answers as feedback for retrieval training.RAG usually relies solely on pseudo-relevant passages retrieved from external knowledge bases which can lead to ineffective answer generation.In this work, we propose a multi-level information RAG approach that enhances answer generation through entity retrieval and query expansion.We formulate a joint-training RAG loss such that answer generation is conditioned on both entity and passage retrievals.We show through experiments new state-of-the-art performance on the VIQuAE KB-VQA benchmark and demonstrate that our approach can help retrieve more actual relevant knowledge to generate accurate answers. Omar Adjali, Olivier Ferret, Sahar Ghannay, Hervé Le Borgne |
EMNLP | 4 |
| 2024 | Semantic Generative Augmentations for Few-Shot CountingabstractWith the availability of powerful text-to-image diffusion models, recent works have explored the use of synthetic data to improve image classification performances. These works show that it can effectively augment or even replace real data. In this work, we investigate how synthetic data can benefit few-shot class-agnostic counting. This requires to generate images that correspond to a given input number of objects. However, text-to-image models struggle to grasp the notion of count. We propose to rely on a double conditioning of Stable Diffusion with both a prompt and a density map in order to augment a training dataset for few-shot counting. Due to the small dataset size, the fine-tuned model tends to generate images close to the training images. We propose to enhance the diversity of synthesized images by exchanging captions between images thus creating unseen configurations of object types and spatial layout. Our experiments show that our diversified generation strategy significantly improves the counting accuracy of two recent and performing few-shot counting models on FSC147 and CARPK. Perla Doubinsky, Nicolas Audebert, Michel Crucianu, Hervé Le Borgne |
WACV | 4 |
| 2024 | TIAM - A Metric for Evaluating Alignment in Text-to-Image GenerationabstractThe progress in the generation of synthetic images has made it crucial to assess their quality. While several metrics have been proposed to assess the rendering of images, it is crucial for Text-to-Image (T2I) models, which generate images based on a prompt, to consider additional aspects such as to which extent the generated image matches the important content of the prompt. Moreover, although the generated images usually result from a random starting point, the influence of this one is generally not considered. In this article, we propose a new metric based on prompt templates to study the alignment between the content specified in the prompt and the corresponding generated images. It allows us to better characterize the alignment in terms of the type of the specified objects, their number, and their color. We conducted a study on several recent T2I models about various aspects. An additional interesting result we obtained with our approach is that image quality can vary drastically depending on the noise used as a seed for the images. We also quantify the influence of the number of concepts in the prompt, their order as well as their (color) attributes. Finally, our method allows us to identify some seeds that produce better images than others, opening novel directions of research on this understudied topic. Paul Grimal, Hervé Le Borgne, Olivier Ferret, Julien Tourille |
WACV | 2 |
| 2023 | Wasserstein loss for Semantic Editing in the Latent Space of GANsabstractThe latent space of GANs contains rich semantics reflecting the training data. Different methods propose to learn edits in latent space corresponding to semantic attributes, thus allowing to modify generated images. Most supervised methods rely on the guidance of classifiers to produce such edits. However, classifiers can lead to out-of-distribution regions and be fooled by adversarial samples. We propose an alternative formulation based on the Wasserstein loss that avoids such problems, while maintaining performance on-par with classifier-based approaches. We demonstrate the effectiveness of our method on two datasets (digits and faces) using StyleGAN2. Code is available at: https://github.com/perladoubinsky/latent-wasserstein Perla Doubinsky, Nicolas Audebert, Michel Crucianu, Hervé Le Borgne |
CBMI | 4 |
| 2023 | Generalized Pseudo-Labeling in Consistency Regularization for Semi-Supervised LearningabstractSemi-Supervised Learning (SSL) reduces annotation cost by exploiting large amounts of unlabeled data. A popular idea in SSL image classification is Pseudo-Labeling (PL), where the predictions of a network are used in order to assign a label to an unlabeled image. However, this practice exposes learning to confirmation bias. In this paper we propose Generalized Pseudo-Labeling (GPL), a simple and generic way to exploit negative pseudo-labels in consistency regularization, entailing minimal additional computational overhead and hyperpameter fine-tuning. GPL makes learning more robust by using the information that an image does not belong to a certain class, which is more abundant and reliable. We showcase GPL in the context of FixMatch. In the benchmark using only 40 labels of the CIFAR-10 dataset, adding GPL on top of FixMatch improves the error rate from 7.93% to 6.58%, and on CIFAR-100 with 2500 labels, from 28.02% to 26.85%. Nikolaos Karaliolios, Florian Chabot, Camille Dupont, Hervé Le Borgne, Quoc Cuong Pham, Romaric Audigier |
ICIP | 4 |
| 2023 | Explicit Knowledge Integration for Knowledge-Aware Visual Question Answering about Named EntitiesabstractRecent years have shown unprecedented growth of interest in Vision-Language related tasks, with the need to address the inherent challenges of integrating linguistic and visual information to solve real-world applications. Such a typical task is Visual Question Answering (VQA), which aims to answer questions about visual content. The limitations of the VQA task in terms of question redundancy and poor linguistic variability encouraged researchers to propose Knowledge-aware Visual Question Answering tasks as a natural extension of VQA. In this paper, we tackle the KVQAE (Knowledge-based Visual Question Answering about named Entities) task, which proposes to answer questions about named entities defined in a knowledge base and grounded in visual content. In particular, besides the textual and visual information, we propose to leverage the structural information extracted from syntactic dependency trees and external knowledge graphs to help answer questions about a large spectrum of entities of various types. Thus, by combining contextual and graph-based representations using Graph Convolutional Networks (GCNs), we are able to learn meaningful embeddings for Information Retrieval tasks. Experiments on the ViQuAE public dataset show how our approach improves the state-of-the-art baselines while demonstrating the interest of injecting external knowledge to enhance multimodal information retrieval. Omar Adjali, Paul Grimal, Olivier Ferret, Sahar Ghannay, Hervé Le Borgne |
ICMR | 5 |
| 2023 | Learning semantic ambiguities for zero-shot learningabstractAbstract Zero-shot learning (ZSL) aims at recognizing classes for which no visual sample is available at training time. To address this issue, one can rely on a semantic description of each class. A typical ZSL model learns a mapping between the visual samples of seen classes and the corresponding semantic descriptions, in order to do the same on unseen classes at test time. State of the art approaches rely on generative models that synthesize visual features from the prototype of a class, such that a classifier can then be learned in a supervised manner. However, these approaches are usually biased towards seen classes whose visual instances are the only one that can be matched to a given class prototype. We propose a regularization method that can be applied to any conditional generative-based ZSL method, by leveraging only the semantic class prototypes. It learns to synthesize discriminative features for possible semantic description that are not available at training time, that is the unseen ones. The approach is evaluated for ZSL and GZSL on four datasets commonly used in the literature, either in inductive or transductive settings, with results on-par or above state of the art approaches. The code is available at https://github.com/hanouticelina/lsa-zsl . Celina Hanouti, Hervé Le Borgne |
Multim. Tools Appl. | 2 |
| 2022 | Self-Improving SLAM in Dynamic Environments: Learning When to Mask
Adrian Bojko, Romain Dupont, Mohamed Tamaazousti, Hervé Le Borgne |
BMVC | 4 |
| 2022 | ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named EntitiesabstractWhether to retrieve, answer, translate, or reason, multimodality opens up new challenges and perspectives. In this context, we are interested in answering questions about named entities grounded in a visual context using a Knowledge Base (KB). To benchmark this task, called KVQAE (Knowledge-based Visual Question Answering about named Entities), we provide ViQuAE, a dataset of 3.7K questions paired with images. This is the first KVQAE dataset to cover a wide range of entity types (e.g. persons, landmarks, and products). The dataset is annotated using a semi-automatic method. We also propose a KB composed of 1.5M Wikipedia articles paired with images. To set a baseline on the benchmark, we address KVQAE as a two-stage problem: Information Retrieval and Reading Comprehension, with both zero- and few-shot learning methods. The experiments empirically demonstrate the difficulty of the task, especially when questions are not about persons. This work paves the way for better multimodal entity representations and question answering. The dataset, KB, code, and semi-automatic annotation pipeline are freely available at https://github.com/PaulLerner/ViQuAE. Paul Lerner, Olivier Ferret, Camille Guinaudeau, Hervé Le Borgne, Romaric Besançon, José G. Moreno 0001, Jesús Lovón-Melgarejo |
SIGIR | 4 |
| 2022 | Multi-attribute balanced sampling for disentangled GAN controls
Perla Doubinsky, Nicolas Audebert, Michel Crucianu, Hervé Le Borgne |
Pattern Recognit. Lett. | 4 |
| 2020 | Webly Supervised Semantic Embeddings for Large Scale Zero-Shot Learning
Yannick Le Cacheux, Adrian Popescu 0001, Hervé Le Borgne |
ACCV (6) | 3 |
| 2020 | Multimodal Entity Linking for Tweets
Omar Adjali, Romaric Besançon, Olivier Ferret, Hervé Le Borgne, Brigitte Grau |
ECIR (1) | 4 |
| 2020 | Controlling generative models with continuous factors of variations
Antoine Plumerault, Hervé Le Borgne, Céline Hudelot |
ICLR | 2 |
| 2020 | Learning to Segment Dynamic Objects using SLAM OutliersabstractWe present a method to automatically learn to segment dynamic objects using SLAM outliers. It requires only one monocular sequence per dynamic object for training and consists in localizing dynamic objects using SLAM outliers, creating their masks, and using these masks to train a semantic segmentation network. We integrate the trained network in ORB-SLAM 2 and LDSO. At runtime we remove features on dynamic objects, making the SLAM unaffected by them. We also propose a new stereo dataset and new metrics to evaluate SLAM robustness. Our dataset includes consensus inversions, i.e., situations where the SLAM uses more features on dynamic objects that on the static background. Consensus inversions are challenging for SLAM as they may cause major SLAM failures. Our approach performs better than the State-of-the-Art on the TUM RGB-D dataset in monocular mode and on our dataset in both monocular and stereo modes. Adrian Bojko, Romain Dupont, Mohamed Tamaazousti, Hervé Le Borgne |
ICPR | 4 |
| 2020 | AVAE: Adversarial Variational Auto EncoderabstractAmong the wide variety of image generative models, two models stand out: Variational Auto Encoders (VAE) and Generative Adversarial Networks (GAN). GANs can produce realistic images, but they suffer from mode collapse and do not provide simple ways to get the latent representation of an image. On the other hand, VAEs do not have these problems, but they often generate images less realistic than GANs. In this article, we explain that this lack of realism is partially due to a common underestimation of the natural image manifold dimensionality. To solve this issue we introduce a new framework that combines VAE and GAN in a novel and complementary way to produce an auto-encoding model that keeps VAEs properties while generating images of GAN-quality. We evaluate our approach both qualitatively and quantitatively on five image datasets. Antoine Plumerault, Hervé Le Borgne, Céline Hudelot |
ICPR | 2 |
| 2020 | Building a Multimodal Entity Linking Dataset From TweetsabstractThe task of Entity linking, which aims at associating an entity mention with a unique entity in a knowledge base (KB), is useful for advanced Information Extraction tasks such as relation extraction or event detection. Most of the studies that address this problem rely only on textual documents while an increasing number of sources are multimedia, in particular in the context of social media where messages are often illustrated with images. In this article, we address the Multimodal Entity Linking (MEL) task, and more particularly the problem of its evaluation. To this end, we propose a novel method to quasi-automatically build annotated datasets to evaluate methods on the MEL task. The method collects text and images to jointly build a corpus of tweets with ambiguous mentions along with a Twitter KB defining the entities. We release a new annotated dataset of Twitter posts associated with images. We study the key characteristics of the proposed dataset and evaluate the performance of several MEL approaches on it. Omar Adjali, Romaric Besançon, Olivier Ferret, Hervé Le Borgne, Brigitte Grau |
LREC | 4 |
| 2020 | Learning More Universal Representations for Transfer-LearningabstractA representation is supposed universal if it encodes any element of the visual world (e.g., objects, scenes) in any configuration (e.g., scale, context). While not expecting pure universal representations, the goal in the literature is to improve the universality level, starting from a representation with a certain level. To improve that universality level, one can diversify the source-task, but it requires many additive annotated data that is costly in terms of manual work and possible expertise. We formalize such a diversification process then propose two methods to improve the universality of CNN representations that limit the need for additive annotated data. The first relies on human categorization knowledge and the second on re-training using fine-tuning. We propose a new aggregating metric to evaluate the universality in a transfer-learning scheme, that addresses more aspects than previous works. Based on it, we show the interest of our methods on 10 target-problems, relating to classification on a variety of visual domains. Youssef Tamaazousti, Hervé Le Borgne, Céline Hudelot, Mohamed El Amine Seddik, Mohamed Tamaazousti |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Modeling Inter and Intra-Class Relations in the Triplet Loss for Zero-Shot LearningabstractRecognizing visual unseen classes, i.e. for which no training data is available, is known as Zero Shot Learning (ZSL). Some of the best performing methods apply the triplet loss to seen classes to learn a mapping between visual representations of images and attribute vectors that constitute class prototypes. They nevertheless make several implicit assumptions that limit their performance on real use cases, particularly with fine-grained datasets comprising a large number of classes. We identify three of these assumptions and put forward corresponding novel contributions to address them. Our approach consists in taking into account both inter-class and intra-class relations, respectively by being more permissive with confusions between similar classes, and by penalizing visual samples which are atypical to their class. The approach is tested on four datasets, including the large-scale ImageNet, and exhibits performances significantly above recent methods, even generative methods based on more restrictive hypotheses. Yannick Le Cacheux, Hervé Le Borgne, Michel Crucianu |
ICCV | 2 |
| 2019 | From Classical to Generalized Zero-Shot Learning: A Simple Adaptation Process
Yannick Le Cacheux, Hervé Le Borgne, Michel Crucianu |
MMM (2) | 2 |
| 2018 | Learning Finer-class Networks for Universal Representations
Julien Girard, Youssef Tamaazousti, Hervé Le Borgne, Céline Hudelot |
BMVC | 3 |
| 2017 | Supervised Learning of Entity Disambiguation Models by Negative Sample Selection
Hani Daher, Romaric Besançon, Olivier Ferret, Hervé Le Borgne, Anne-Laure Daquo, Youssef Tamaazousti |
CICLing (1) | 4 |
| 2017 | MuCaLe-Net: Multi Categorical-Level Networks to Generate More Discriminating FeaturesabstractIn a transfer-learning scheme, the intermediate layers of a pre-trained CNN are employed as universal image representation to tackle many visual classification problems. The current trend to generate such representation is to learn a CNN on a large set of images labeled among the most specific categories. Such processes ignore potential relations between categories, as well as the categorical-levels used by humans to classify. In this paper, we propose Multi Categorical-Level Networks (MuCaLe-Net) that include human-categorization knowledge into the CNN learning process. A MuCaLe-Net separates generic categories from each other while it independently distinguishes specific ones. It thereby generates different features in the intermediate layers that are complementary when combined together. Advantageously, our method does not require additive data nor annotation to train the network. The extensive experiments over four publicly available benchmarks of image classification exhibit state-of-the-art performances. Youssef Tamaazousti, Hervé Le Borgne, Céline Hudelot |
CVPR | 2 |
| 2017 | AMECON: Abstract Meta-Concept Features for Text-IllustrationabstractCross-media retrieval is a problem of high interest that is at the frontier between computer vision and natural language processing. The state-of-the-art in the domain consists of learning a common space with regard to some constraints of correlation or similarity from two textual and visual modalities that are processed in parallel and possibly jointly. This paper proposes a different approach that considers the cross-modal problem as a supervised mapping of visual modalities to textual ones. Each modality is thus seen as a particular projection of an abstract meta-concept, each of its dimension subsuming several semantic concepts (``meta'' aspect) but may not correspond to an actual one (``abstract'' aspect). In practice, the textual modality is used to generate a multi-label representation, further used to map the visual modality through a simple shallow neural network. While being quite easy to implement, the experiments show that our approach significantly outperforms the state-of-the-art on Flickr-8K and Flickr-30K datasets for the text-illustration task. The source code is available at http://perso.ecp.fr/~tamaazouy/. Ines Chami, Youssef Tamaazousti, Hervé Le Borgne |
ICMR | 3 |
| 2017 | Vision-language integration using constrained local semantic features
Youssef Tamaazousti, Hervé Le Borgne, Adrian Popescu 0001, Etienne Gadeski, Alexandru-Lucian Gînsca, Céline Hudelot |
Comput. Vis. Image Underst. | 2 |
| 2017 | Harnessing noisy Web images for deep representation
Phong D. Vo, Alexandru-Lucian Gînsca, Hervé Le Borgne, Adrian Popescu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2017 | Fast and robust duplicate image detection on the web
Etienne Gadeski, Hervé Le Borgne, Adrian Popescu 0001 |
Multim. Tools Appl. | 2 |
| 2016 | Aggregating Image and Text Quantized Correlated ComponentsabstractCross-modal tasks occur naturally for multimedia content that can be described along two or more modalities like visual content and text. Such tasks require to "translate" information from one modality to another. Methods like kernelized canonical correlation analysis (KCCA) attempt to solve such tasks by finding aligned subspaces in the description spaces of different modalities. Since they favor correlations against modality-specific information, these methods have shown some success in both cross-modal and bi-modal tasks. However, we show that a direct use of the subspace alignment obtained by KCCA only leads to coarse translation abilities. To address this problem, we first put forward a new representation method that aggregates information provided by the projections of both modalities on their aligned subspaces. We further suggest a method relying on neighborhoods in these subspaces to complete uni-modal information. Our proposal exhibits state-of-the-art results for bi-modal classification on Pascal VOC07 and improves it by over 60% for cross-modal retrieval on FlickR 8K/30K. Thi Quynh Nhi Tran, Hervé Le Borgne, Michel Crucianu |
CVPR | 2 |
| 2016 | Constrained Local Enhancement of Semantic Features by Content-Based SparsityabstractSemantic features represent images by the outputs of a set of visual concept classifiers and have shown interesting performances in image classification and retrieval. All classifier outputs are usually exploited but it was recently shown that feature sparsification improves both performance and scalability. However, existing approaches consider a fixed sparsity level which disregards the actual content of individual images. In this paper, we propose a method to determine automatically a level of sparsity for the semantic features that is adapted to each image content. This method takes into account the amount of information contained by the image through a modeling of the semantic feature entropy and the confidence of individual dimensions of the feature. We also investigate the use of local regions of the image to further improve the quality of semantic features. Experimental validation is conducted on three benchmarks (Pascal VOC 2007, VOC 2012 and MIT Indoor) for image classification and two of them for image retrieval. Our method obtains competitive results on image classification and achieves state-of-the-art performances on image retrieval. Youssef Tamaazousti, Hervé Le Borgne, Adrian Popescu 0001 |
ICMR | 2 |
| 2016 | Diverse Concept-Level Features for Multi-Object ClassificationabstractWe consider the problem of image classification with semantic features that are built from a set of base classifier outputs, each reflecting visual concepts. However, existing approaches consider visual concepts independently from each other whereas they are often linked together. When those relations are considered, existing models strongly rely on image low-level features, yielding in irrelevant relations when the low-level representation fails. On the contrary, the approach we propose, uses existing human knowledge, the application context itself and the human categorization mechanism to reflect complex relations between concepts. By nesting this human knowledge and the application context in the concept detection and selection processes, our final semantic feature captures the most useful information for an effective categorization. Thus, it enables to give good representation, even if some important concepts failed to be recognized. Experimental validation is conducted on three publicly available benchmarks of multi-class object classification and leads to results that outperforms comparable approaches. Youssef Tamaazousti, Hervé Le Borgne, Céline Hudelot |
ICMR | 2 |
| 2015 | Combining Generic and Specific Information for Cross-modal RetrievalabstractCross-modal retrieval increasingly relies on joint statistical models built from large amounts of data represented according to several modalities. However, some information that is poorly represented by these models can be very significant for a retrieval task. We show that, by appropriately identifying and taking such information into account, the results of cross-modal retrieval can be strongly improved. We apply our model to three benchmarks for the text illustration task and find that the more data has misrepresented information, the more our model is comparatively effective. Thi Quynh Nhi Tran, Hervé Le Borgne, Michel Crucianu |
ICMR | 2 |
| 2015 | Large-Scale Image Mining with Flickr Groups
Alexandru-Lucian Gînsca, Adrian Popescu 0001, Hervé Le Borgne, Nicolas Ballas, Dinh-Phong Vo, Ioannis Kanellos |
MMM (1) | 3 |
| 2013 | Tag completion based on belief theory and neighbor votingabstractWe address the problem of tag completion for automatic image annotation. Our method consists in two main steps: creating a list of "candidate tags" from the visual neighbors of the untagged image then using them as pieces of evidence to be combined to provide the final list of predicted tags. Both steps introduce a scheme to tackle with imprecision and uncertainty. First, a bag-of-words (BOW) signature is generated for each neighbor using local soft coding. Second, a sum-pooling operation across the BOW of the k nearest neighbors provides the list of "candidate tags". Finally, we use neighbors as pieces of evidence to be combined according to the Dempster's rule to predict the more relevant tags. The method is evaluated in the context of image classification and that of tag suggestion. The database used for visual neighbors search contains 1.2 million images extracted from Flickr. Classification is evaluated on the well known Pascal VOC 2007 and MIR Flickr datasets, on which we obtain similar or better results than the state-of-the-art. For tag suggestion, we manually annotated 241 queries. As well, we obtain competitive results on this task. Amel Znaidia, Hervé Le Borgne, Céline Hudelot |
ICMR | 2 |
| 2012 | Locality-constrained and spatially regularized coding for scene categorizationabstractImproving coding and spatial pooling for bag-of-words based feature design have gained a lot of attention in recent works addressing object recognition and scene classification. Regarding the coding step in particular, properties such as sparsity, locality and saliency have been investigated. The main contribution of this work consists in taking into acount the local spatial context of an image into the usual coding strategies proposed in the state-of-the-art. For this purpose, given an imgae, dense local features are extracted and structured in a lattice. The latter is endowed with a neighborhood system and pairwise interactions. We propose a new objective function to encode local features, which preserves locality constraints both in the feature space and the spatial domain of the image. In addition, an appropriate efficient optimization algorithm is provided, inspired from the graph-cut framework. In conjunction with the maximum-pooling operation and the spatial pyramid matching, that reflects a global spatial layout, the proposed method improves the performances of several state-of-the-art coding schemes for scene classification on three publicly available benchmarks (UIUC 8-sport, Scene-15 and Caltech-101). Aymen Shabou, Hervé Le Borgne |
CVPR | 2 |
| 2012 | Bag-of-multimedia-words for image classification
Amel Znaidia, Aymen Shabou, Hervé Le Borgne, Céline Hudelot, Nikos Paragios |
ICPR | 3 |
| 2012 | Multimodal feature generation framework for semantic image classificationabstractThe automatic attribution of semantic labels to unlabeled or weakly labeled images has received considerable attention but, given the complexity of the problem, remains a hard research topic. Here we propose a unified classification framework which mixes textual and visual information in a seamless manner. Unlike most recent previous works, computer vision techniques are used as inspiration to process textual information. To do so, we consider two types of complementary tag similarities, respectively computed from a conceptual hierarchy and from data collected from a photo sharing platform. Visual content is processed using recent techniques for bag-of visual-words feature generation. A central contribution of our work is to infer the coding step of the general bag-of-word framework with such similarities and to aggregate these tag-codes by max-pooling to obtain a single representative vector (signature). Final image annotations are obtained via late fusion, where the three modalities (two text-based and one visual-based) are merged during the classification step. Experimental results on the Pascal VOC 2007 and MIR Flickr datasets show an improvement over the state-of-the-art methods, while significantly decreasing the computational complexity of the learning system. Amel Znaidia, Aymen Shabou, Adrian Popescu 0001, Hervé Le Borgne, Céline Hudelot |
ICMR | 4 |
| 2012 | Fast shared boosting for large-scale concept detection
Hervé Le Borgne, Nicolas Honnorat |
Multim. Tools Appl. | 1 |
| 2011 | Nonparametric Estimation of Fisher Vectors to Aggregate Image Descriptors
Hervé Le Borgne, Pablo Muñoz Fuentes |
ACIVS | 1 |
| 2007 | Learning Midlevel Image Features for Natural Scene and Texture ClassificationabstractThis paper deals with coding of natural scenes in order to extract semantic information. We present a new scheme to project natural scenes onto a basis in which each dimension encodes statistically independent information. Basis extraction is performed by independent component analysis (ICA) applied to image patches culled from natural scenes. The study of the resulting coding units (coding filters) extracted from well-chosen categories of images shows that they adapt and respond selectively to discriminant features in natural scenes. Given this basis, we define global and local image signatures relying on the maximal activity of filters on the input image. Locally, the construction of the signature takes into account the spatial distribution of the maximal responses within the image. We propose a criterion to reduce the size of the space of representation for faster computation. The proposed approach is tested in the context of texture classification (111 classes), as well as natural scenes classification (11 categories, 2037 images). Using a common protocol, the other commonly used descriptors have at most 47.7% accuracy on average while our method obtains performances of up to 63.8%. We show that this advantage does not depend on the size of the signature and demonstrate the efficiency of the proposed criterion to select ICA filters and reduce the dimension. Hervé Le Borgne, Anne Guérin-Dugué, Noel E. O'Connor |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Pre-Classification for Automatic Image OrientationabstractIn this paper, we propose a novel method for automatic orientation of digital images. The approach is based on exploiting the properties of local statistics of natural scenes. In this way, we address some of the difficulties encountered in previous works in this area. The main contribution of this paper is to introduce a pre-classification step into carefully defined categories in order to simplify subsequent orientation detection. The proposed algorithm was tested on 9068 images and compared to existing state of the art in the area. Results show a significant improvement over previous works Hervé Le Borgne, Noel E. O'Connor |
ICASSP (2) | 1 |
| 2006 | Fischlar-TRECVid-2004: combined text- and image-based searching of video archivesabstractThe Fischlar-TRECVid-2004 system was developed for Dublin City University's participation in the 2004 TRECVid video information retrieval benchmarking activity. The system allows search and retrieval of video shots from over 60 hours of content. The shot retrieval engine employed is based on a combination of query text matched against spoken dialogue combined with image-image matching where a still image (sourced externally), or a keyframe (from within the video archive itself), is matched against all keyframes in the video archive. Three separate text retrieval engines are employed for closed caption text, automatic speech recognition and video OCR. Visual shot matching is primarily based on MPEG-7 low-level descriptors. The system supports relevance feedback at the shot level enabling augmentation and refinement using relevant shots located by the user. Two variants of the system were developed, one that supports both text- and image-based searching and one that supports image only search. A user evaluation experiment compared the use of the two systems. Results show that while the system combining text- and image-based searching achieves greater retrieval effectiveness, users make more varied and extensive queries with the image only based searching version. Noel E. O'Connor, Hyowon Lee 0001, Alan F. Smeaton, Gareth J. F. Jones, Eddie Cooke, Hervé Le Borgne, Cathal Gurrin |
ISCAS | 6 |
| 2005 | Natural Scene Classification and Retrieval Using Ridgelet-Based Image Signatures
Hervé Le Borgne, Noel E. O'Connor |
ACIVS | 1 |
| 2005 | Fusing MPEG-7 Visual Descriptors for Image Classification
Evaggelos Spyrou, Hervé Le Borgne, Theofilos P. Mailis, Eddie Cooke, Yannis Avrithis, Noel E. O'Connor |
ICANN (2) | 2 |
| 2004 | Representation of images for classification with independent features
Hervé Le Borgne, Anne Guérin-Dugué, Anestis Antoniadis |
Pattern Recognit. Lett. | 1 |