EDBT 2026 Demo / reviewers in the wild / expert
Thibaut Durand
dblp:141/9848
· DBLP profile ↗
14ranked-venue papers
10as first author
2since 2021 · last 2025
0000-0003-3698-0920ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 8 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 23% Learning paradigms · 23% Generative modeling · 14% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 100% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
prompting |
0.9 | 1 | 2025 | LAST SToP for Modeling Asynchronous Time Series · ICML 2025 |
Natural language and speech › Language models and text generation
prompt tuning |
0.9 | 1 | 2025 | LAST SToP for Modeling Asynchronous Time Series · ICML 2025 |
Machine learning › Learning paradigms
multi-label classification |
0.8 | 2 | 2019 | Exploiting Negative Evidence for Deep Latent Structured Models · IEEE Trans. Pattern Anal. Mach. Intell. 2019 Learning a Deep ConvNet for Multi-Label Classification With Partial Labels · CVPR 2019 |
Machine learning › Generative modeling
variational autoencoder |
0.8 | 2 | 2019 | LayoutVAE: Stochastic Scene Layout Generation From a Label Set · ICCV 2019 A Variational Auto-Encoder Model for Stochastic Point Processes · CVPR 2019 |
Computer vision › Segmentation and scene understanding › semantic segmentation
weakly supervised semantic segmentation |
0.7 | 2 | 2019 | Exploiting Negative Evidence for Deep Latent Structured Models · IEEE Trans. Pattern Anal. Mach. Intell. 2019 WILDCAT: Weakly Supervised Learning of Deep ConvNets for Image Classification, Pointwise Localization and Segmentation · CVPR 2017 |
Machine learning › Learning paradigms
weakly supervised learning |
0.5 | 2 | 2016 | WELDON: Weakly Supervised Learning of Deep Convolutional Neural Networks · CVPR 2016 MANTRA: Minimum Maximum Latent Structural SVM for Image Classification and Ranking · ICCV 2015 |
Machine learning › Learning paradigms › weakly supervised learning
partial label learning |
0.4 | 1 | 2019 | Learning a Deep ConvNet for Multi-Label Classification With Partial Labels · CVPR 2019 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process |
0.4 | 1 | 2019 | A Variational Auto-Encoder Model for Stochastic Point Processes · CVPR 2019 |
Machine learning › Generative modeling › scene generation
scene layout generation |
0.4 | 1 | 2019 | LayoutVAE: Stochastic Scene Layout Generation From a Label Set · ICCV 2019 |
Computer vision › Image recognition and object detection › object recognition
weakly supervised object recognition |
0.4 | 1 | 2019 | Exploiting Negative Evidence for Deep Latent Structured Models · IEEE Trans. Pattern Anal. Mach. Intell. 2019 |
Visual content generation and editing
layout generation |
0.4 | 1 | 2019 | LayoutVAE: Stochastic Scene Layout Generation From a Label Set · ICCV 2019 |
Computer vision › Image recognition and object detection
image classification |
0.4 | 3 | 2017 | MANTRA: Minimum Maximum Latent Structural SVM for Image Classification and Ranking · ICCV 2015 WILDCAT: Weakly Supervised Learning of Deep ConvNets for Image Classification, Pointwise Localization and Segmentation · CVPR 2017 WELDON: Weakly Supervised Learning of Deep Convolutional Neural Networks · CVPR 2016 |
Computer vision › Image recognition and object detection › object localization
weakly supervised object localization |
0.3 | 1 | 2017 | WILDCAT: Weakly Supervised Learning of Deep ConvNets for Image Classification, Pointwise Localization and Segmentation · CVPR 2017 |
Machine learning › Learning paradigms
multiple instance learning |
0.2 | 1 | 2016 | WELDON: Weakly Supervised Learning of Deep Convolutional Neural Networks · CVPR 2016 |
Machine learning › Kernel, tree and ensemble methods › support vector machine
latent structural SVM |
0.2 | 1 | 2015 | MANTRA: Minimum Maximum Latent Structural SVM for Image Classification and Ranking · ICCV 2015 |
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification |
0.2 | 1 | 2015 | MANTRA: Minimum Maximum Latent Structural SVM for Image Classification and Ranking · ICCV 2015 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
user representation learning |
0.1 | 1 | 2020 | Learning User Representations for Open Vocabulary Image Hashtag Prediction · CVPR 2020 |
Machine learning › Efficient and distributed learning › data-efficient learning
label-efficient learning |
0.1 | 1 | 2019 | Learning a Deep ConvNet for Multi-Label Classification With Partial Labels · CVPR 2019 |
Natural language and speech › Language models and text generation › text generation
stochastic generation |
0.1 | 1 | 2019 | LayoutVAE: Stochastic Scene Layout Generation From a Label Set · ICCV 2019 |
Methods — techniques the papers use, named apart from their topics
prompt tuning · 0.9QLoRA · 0.9variational autoencoder · 0.8cross-modal embedding · 0.4variational inference · 0.4latent structured model · 0.4latent representation · 0.4curriculum learning · 0.4convolutional neural network · 0.4classification loss · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LAST SToP for Modeling Asynchronous Time SeriesabstractWe present a novel prompt design for Large Language Models (LLMs) tailored to Asynchronous Time Series. Unlike regular time series, which assume values at evenly spaced time points, asynchronous time series consist of timestamped events occurring at irregular intervals, each described in natural language. Our approach effectively utilizes the rich natural language of event descriptions, allowing LLMs to benefit from their broad world knowledge for reasoning across different domains and tasks. This allows us to extend the scope of asynchronous time series analysis beyond forecasting to include tasks like anomaly detection and data imputation. We further introduce Stochastic Soft Prompting, a novel prompt-tuning mechanism that significantly improves model performance, outperforming existing finetuning methods such as QLORA. Through extensive experiments on real-world datasets, we demonstrate that our approach achieves state-of-the-art performance across different tasks and datasets. Thibaut Durand, Graham W. Taylor, Lilian W. Bialokozowicz |
ICML | 2 |
| 2021 | Variational Selective Autoencoder: Learning from Partially-Observed Heterogeneous DataabstractLearning from heterogeneous data poses challenges such as combining data from various sources and of different types. Meanwhile, heterogeneous data are often associated with missingness in real-world applications due to heterogeneity and noise of input sources. In this work, we propose the variational selective autoencoder (VSAE), a general framework to learn representations from partially-observed heterogeneous data. VSAE learns the latent dependencies in heterogeneous data by modeling the joint distribution of observed data, unobserved data, and the imputation mask which represents how the data are missing. It results in a unified model for various downstream tasks including data generation and imputation. Evaluation on both low-dimensional and high-dimensional heterogeneous datasets for these two tasks shows improvement over state-of-the-art models. Hossein Hajimirsadeghi, Jiawei He 0001, Thibaut Durand, Greg Mori |
AISTATS | 4 |
| 2020 | Learning User Representations for Open Vocabulary Image Hashtag PredictionabstractIn this paper, we introduce an open vocabulary model for image hashtag prediction - the task of mapping an image to its accompanying hashtags. Recent work shows that to build an accurate hashtag prediction model, it is necessary to model the user because of the self-expression problem, in which similar image content may be labeled with different tags. To take into account the user behaviour, we propose a new model that extracts a representation of a user based on his/her image history. Our model allows to improve a user representation with new images or add a new user without retraining the model. Because new hashtags appear all the time on social networks, we design an open vocabulary model which can deal with new hashtags without retraining the model. Our model learns a cross-modal embedding between user conditional visual representations and hashtag word representations. Experiments on a subset of the YFCC100M dataset demonstrate the efficacy of our user representation in user conditional hashtag prediction and user retrieval. We further validate the open vocabulary prediction ability of our model. Thibaut Durand |
CVPR | 1 |
| 2019 | Learning a Deep ConvNet for Multi-Label Classification With Partial LabelsabstractDeep ConvNets have shown great performance for single-label image classification (e.g. ImageNet), but it is necessary to move beyond the single-label classification task because pictures of everyday life are inherently multi-label. Multi-label classification is a more difficult task than single-label classification because both the input images and output label spaces are more complex. Furthermore, collecting clean multi-label annotations is more difficult to scale-up than single-label annotations. To reduce the annotation cost, we propose to train a model with partial labels i.e. only some labels are known per image. We first empirically compare different labeling strategies to show the potential for using partial labels on multi-label datasets. Then to learn with partial labels, we introduce a new classification loss that exploits the proportion of known labels per example. Our approach allows the use of the same training settings as when learning with all the annotations. We further explore several curriculum learning based strategies to predict missing labels. Experiments are performed on three large-scale multi-label datasets: MS COCO, NUS-WIDE and Open Images. Thibaut Durand, Nazanin Mehrasa, Greg Mori |
CVPR | 1 |
| 2019 | A Variational Auto-Encoder Model for Stochastic Point ProcessesabstractWe propose a novel probabilistic generative model for action sequences. The model is termed the Action Point Process VAE (APP-VAE), a variational auto-encoder that can capture the distribution over the times and categories of action sequences. Modeling the variety of possible action sequences is a challenge, which we show can be addressed via the APP-VAE's use of latent representations and non-linear functions to parameterize distributions over which event is likely to occur next in a sequence and at what time. We empirically validate the efficacy of APP-VAE for modeling action sequences on the MultiTHUMOS and Breakfast datasets. Nazanin Mehrasa, Akash Abdu Jyothi, Thibaut Durand, Jiawei He 0001, Leonid Sigal, Greg Mori |
CVPR | 3 |
| 2019 | LayoutVAE: Stochastic Scene Layout Generation From a Label SetabstractRecently there is an increasing interest in scene generation within the research community. However, models used for generating scene layouts from textual description largely ignore plausible visual variations within the structure dictated by the text. We propose LayoutVAE, a variational autoencoder based framework for generating stochastic scene layouts. LayoutVAE is a versatile modeling framework that allows for generating full image layouts given a label set, or per label layouts for an existing image given a new label. In addition, it is also capable of detecting unusual layouts, potentially providing a way to evaluate layout generation problem. Extensive experiments on MNIST-Layouts and challenging COCO 2017 Panoptic dataset verifies the effectiveness of our proposed framework. Akash Abdu Jyothi, Thibaut Durand, Jiawei He 0001, Leonid Sigal, Greg Mori |
ICCV | 2 |
| 2019 | Exploiting Negative Evidence for Deep Latent Structured ModelsabstractThe abundance of image-level labels and the lack of large scale detailed annotations (e.g. bounding boxes, segmentation masks) promotes the development of weakly supervised learning (WSL) models. In this work, we propose a novel framework for WSL of deep convolutional neural networks dedicated to learn localized features from global image-level annotations. The core of the approach is a new latent structured output model equipped with a pooling function which explicitly models negative evidence, e.g. a cow detector should strongly penalize the prediction of the bedroom class. We show that our model can be trained end-to-end for different visual recognition tasks: multi-class and multi-label classification, and also structured average precision (AP) ranking. Extensive experiments highlight the relevance of the proposed method: our model outperforms state-of-the art results on six datasets. We also show that our framework can be used to improve the performance of state-of-the-art deep models for large scale image classification on ImageNet. Finally, we evaluate our model for weakly supervised tasks: in particular, a direct adaptation for weakly supervised segmentation provides a very competitive model. Thibaut Durand, Nicolas Thome, Matthieu Cord |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | SyMIL: MinMax Latent SVM for Weakly Labeled DataabstractDesigning powerful models able to handle weakly labeled data are a crucial problem in machine learning. In this paper, we propose a new multiple instance learning (MIL) framework. Examples are represented as bags of instances, but we depart from standard MIL assumptions by introducing a symmetric strategy (SyMIL) that seeks discriminative instances in positive and negative bags. The idea is to use the instance the most distant from the hyper-plan to classify the bag. We provide a theoretical analysis featuring the generalization properties of our model. We derive a large margin formulation of our problem, which is cast as a difference of convex functions, and optimized using concave-convex procedure. We provide a primal version optimizing with stochastic subgradient descent and a dual version optimizing with one-slack cutting-plane. Successful experimental results are reported on standard MIL and weakly supervised object detection data sets: SyMIL significantly outperforms competitive methods (mi/MI/Latent-SVM), and gives very competitive performance compared to state-of-the-art works. We also analyze the selected instances of symmetric and asymmetric approaches on weakly supervised object detection and text classification tasks. Finally, we show complementarity of SyMIL with recent works on learning with label proportions on standard MIL data sets. Thibaut Durand, Nicolas Thome, Matthieu Cord |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | WILDCAT: Weakly Supervised Learning of Deep ConvNets for Image Classification, Pointwise Localization and SegmentationabstractThis paper introduces WILDCAT, a deep learning method which jointly aims at aligning image regions for gaining spatial invariance and learning strongly localized features. Our model is trained using only global image labels and is devoted to three main visual recognition tasks: image classification, weakly supervised object localization and semantic segmentation. WILDCAT extends state-of-the-art Convolutional Neural Networks at three main levels: the use of Fully Convolutional Networks for maintaining spatial resolution, the explicit design in the network of local features related to different class modalities, and a new way to pool these features to provide a global image prediction required for weakly supervised training. Extensive experiments show that our model significantly outperforms state-of-the-art methods. Thibaut Durand, Taylor Mordan, Nicolas Thome, Matthieu Cord |
CVPR | 1 |
| 2016 | WELDON: Weakly Supervised Learning of Deep Convolutional Neural NetworksabstractIn this paper, we introduce a novel framework for WEakly supervised Learning of Deep cOnvolutional neural Networks (WELDON). Our method is dedicated to automatically selecting relevant image regions from weak annotations, e.g. global image labels, and encompasses the following contributions. Firstly, WELDON leverages recent improvements on the Multiple Instance Learning paradigm, i.e. negative evidence scoring and top instance selection. Secondly, the deep CNN is trained to optimize Average Precision, and fine-tuned on the target dataset with efficient computations due to convolutional feature sharing. A thorough experimental validation shows that WELDON outperforms state-of-the-art results on six different datasets. Thibaut Durand, Nicolas Thome, Matthieu Cord |
CVPR | 1 |
| 2015 | MANTRA: Minimum Maximum Latent Structural SVM for Image Classification and RankingabstractIn this work, we propose a novel Weakly Supervised Learning (WSL) framework dedicated to learn discriminative part detectors from images annotated with a global label. Our WSL method encompasses three main contributions. Firstly, we introduce a new structured output latent variable model, Minimum mAximum lateNt sTRucturAl SVM (MANTRA), which prediction relies on a pair of latent variables: h+(resp. h-) provides positive (resp. negative) evidence for a given output y. Secondly, we instantiate MANTRA for two different visual recognition tasks: multi-class classification and ranking. For ranking, we propose efficient solutions to exactly solve the inference and the loss-augmented problems. Finally, extensive experiments highlight the relevance of the proposed method: MANTRA outperforms state-of-the art results on five different datasets. Thibaut Durand, Nicolas Thome, Matthieu Cord |
ICCV | 1 |
| 2014 | Semantic pooling for image categorization using multiple kernel learningabstractIn this paper, we propose a new method for taking into account the spatial information in image categorization. More specifically, we remove the loss of spatial information in Bag of Words related methods by computing the image signature over specific regions selected by object detectors. We propose to select the detectors using Multiple Kernel Learning techniques. We carry out experiments on the well known VOC 2007 dataset, and show our semantic pooling obtains promising results. Thibaut Durand, David Picard, Nicolas Thome, Matthieu Cord |
ICIP | 1 |
| 2014 | Incremental learning of latent structural SVM for weakly supervised image classificationabstractVisual learning with weak supervision is a promising research area, since it offers the possibility to build large image datasets at reasonable cost. In this paper, we address the problem of weakly supervised object detection, where the goal is to predict the label of the image using object position as latent variable. We propose a new method that builds upon the Latent Structural SVM (LSSVM) formalism. Specifically, we introduce an original coarse-to-fine approach that limits the evolution of the latent parameter subspace. This incremental strategy drives the learning towards better solutions, providing a model with increased predictive accuracy. In addition, this leads to a significant speed up during learning and inference compared to standard sliding window methods. Experiments carried out on Mammal dataset validate the good performances and fast training of the method compared to state-of-the-art works. Thibaut Durand, Nicolas Thome, Matthieu Cord, David Picard |
ICIP | 1 |
| 2013 | Image classification using object detectorsabstractImage categorization is one of the most competitive topic in computer vision and image processing. In this paper, we propose to use trained object and region detectors to represent the visual content of each image. Compared to similar methods found in the literature, our method encompasses two main areas of novelty: introducing a new spatial pooling formalism and designing a late fusion strategy for combining our representation with state-of-the art methods based on low-level descriptors, e.g. Fisher Vectors and BossaNova. Our experiments carried out in the challenging PASCAL VOC 2007 dataset reveal outstanding performances. When combined with low-level representations, we reach more than 67.6% in MAP, outperforming recently reported results in this dataset with a large margin. Thibaut Durand, Nicolas Thome, Matthieu Cord, Sandra Eliza Fontes de Avila |
ICIP | 1 |