Dumitru Erhan

dblp:97/2914 · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0001-7650-1475ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
22 papers
Generative modeling · 22% Reinforcement learning · 15% Trustworthy machine learning · 15%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 76% Multimedia analysis and retrieval · 24%

Topics — the 30 heaviest of 42, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
video prediction
1.032019
High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks · NeurIPS 2019
Hierarchical Long-term Video Prediction without Supervision · ICML 2018
Stochastic Variational Video Prediction · ICLR (Poster) 2018
Machine learning › Reinforcement learning
model-based reinforcement learning
1.022022
Information Prioritization through Empowerment in Visual Model-based RL · ICLR 2022
Model Based Reinforcement Learning for Atari · ICLR 2020
Computer vision › Image recognition and object detection
object detection
0.842016
SSD: Single Shot MultiBox Detector · ECCV (1) 2016
Going deeper with convolutions · CVPR 2015
Scalable Object Detection Using Deep Neural Networks · CVPR 2014
Visual content generation and editing
video generation
0.822023
Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions · ICLR 2023
VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation · ICLR 2020
Machine learning › Trustworthy machine learning › interpretability
attribution methods
0.722019
A Benchmark for Interpretability Methods in Deep Neural Networks · NeurIPS 2019
Learning how to explain neural networks: PatternNet and PatternAttribution · ICLR (Poster) 2018
Machine learning › Trustworthy machine learning
interpretability
0.722019
A Benchmark for Interpretability Methods in Deep Neural Networks · NeurIPS 2019
Learning how to explain neural networks: PatternNet and PatternAttribution · ICLR (Poster) 2018
Machine learning › Generative modeling › video generation
text-to-video generation
0.712023
Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions · ICLR 2023
Visual content generation and editing › multimodal content generation
story visualization
0.712023
StoryBench: A Multifaceted Benchmark for Continuous Story Visualization · NeurIPS 2023
Visual content generation and editing › video generation
text-to-video generation
0.712023
StoryBench: A Multifaceted Benchmark for Continuous Story Visualization · NeurIPS 2023
Machine learning › Reinforcement learning › exploration › intrinsically motivated reinforcement learning
empowerment
0.612022
Information Prioritization through Empowerment in Visual Model-based RL · ICLR 2022
Machine learning › Reinforcement learning
exploration
0.612022
Information Prioritization through Empowerment in Visual Model-based RL · ICLR 2022
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.522017
Unsupervised Pixel-Level Domain Adaptation with Generative Adversarial Networks · CVPR 2017
Domain Separation Networks · NIPS 2016
Computer vision › Vision and language
image captioning
0.522017
Show and Tell: Lessons Learned from the 2015 MSCOCO Image Captioning Challenge · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Show and tell: A neural image caption generator · CVPR 2015
Computer vision › 3D vision
3d scene reconstruction
0.412020
SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous Driving · CVPR 2020
Machine learning › Generative modeling › image generation
GAN-based image generation
0.412020
SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous Driving · CVPR 2020
Machine learning › Generative modeling
image generation
0.412020
SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous Driving · CVPR 2020
Machine learning › Generative modeling
normalizing flow
0.412020
VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation · ICLR 2020
Robotics › Robot manipulation › robot simulation
sensor simulation
0.412020
SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous Driving · CVPR 2020
Machine learning › Trustworthy machine learning › interpretability › attribution methods
saliency methods
0.412019
A Benchmark for Interpretability Methods in Deep Neural Networks · NeurIPS 2019
Machine learning › Deep learning architectures and training › recurrent neural network
stochastic recurrent neural network
0.412019
High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks · NeurIPS 2019
Machine learning › Generative modeling
autoregressive model
0.312018
Hierarchical Long-term Video Prediction without Supervision · ICML 2018
Computer vision › Video understanding and tracking › video prediction
long-term video prediction
0.312018
Hierarchical Long-term Video Prediction without Supervision · ICML 2018
Machine learning › Trustworthy machine learning › interpretability
neural network interpretation
0.312018
Learning how to explain neural networks: PatternNet and PatternAttribution · ICLR (Poster) 2018
Computer vision › Video understanding and tracking › video prediction
stochastic video prediction
0.312018
Stochastic Variational Video Prediction · ICLR (Poster) 2018
Machine learning › Generative modeling
generative adversarial network
0.312017
Unsupervised Pixel-Level Domain Adaptation with Generative Adversarial Networks · CVPR 2017
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.312017
Unsupervised Pixel-Level Domain Adaptation with Generative Adversarial Networks · CVPR 2017
Machine learning › Transfer learning and domain adaptation › domain adaptation › visual domain adaptation
pixel-level domain adaptation
0.312017
Unsupervised Pixel-Level Domain Adaptation with Generative Adversarial Networks · CVPR 2017
Machine learning › Deep learning architectures and training
convolutional neural network
0.322016
Scalable Object Detection Using Deep Neural Networks · CVPR 2014
SSD: Single Shot MultiBox Detector · ECCV (1) 2016
Computer vision › Image recognition and object detection › object detection
one-stage object detection
0.212016
SSD: Single Shot MultiBox Detector · ECCV (1) 2016
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network architecture
0.212015
Going deeper with convolutions · CVPR 2015

Methods — techniques the papers use, named apart from their topics

transformer · 1.3human evaluation · 1.3autoregressive modeling · 1.3automatic metrics · 1.3conditional flow · 0.9convolutional neural network · 0.8mutual information maximization · 0.6deep recurrent architecture · 0.5surfel representation · 0.4GAN · 0.4
YearPublicationVenuePosition
2023 Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions
Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang 0010, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, Dumitru Erhan
ICLR9
2023 StoryBench: A Multifaceted Benchmark for Continuous Story Visualization
abstract
Generating video stories from text prompts is a complex task. In addition to having high visual quality, videos need to realistically adhere to a sequence of text prompts whilst being consistent throughout the frames. Creating a benchmark for video generation requires data annotated over time, which contrasts with the single caption used often in video datasets. To fill this gap, we collect comprehensive human annotations on three existing datasets, and introduce StoryBench: a new, challenging multi-task benchmark to reliably evaluate forthcoming text-to-video models. Our benchmark includes three video generation tasks of increasing difficulty: action execution, where the next action must be generated starting from a conditioning video; story continuation, where a sequence of actions must be executed starting from a conditioning video; and story generation, where a video must be generated from only text prompts. We evaluate small yet strong text-to-video baselines, and show the benefits of training on story-like data algorithmically generated from existing video captions. Finally, we establish guidelines for human evaluation of video stories, and reaffirm the need of better automatic metrics for video generation. StoryBench aims at encouraging future research efforts in this exciting new area.
Emanuele Bugliarello, Hernan Moraldo, Ruben Villegas, Mohammad Babaeizadeh, Mohammad Taghi Saffar, Han Zhang 0010, Dumitru Erhan, Vittorio Ferrari, Pieter-Jan Kindermans, Paul Voigtlaender
NeurIPS7
2022 Information Prioritization through Empowerment in Visual Model-based RL
Homanga Bharadhwaj, Mohammad Babaeizadeh, Dumitru Erhan, Sergey Levine
ICLR3
2020 SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous Driving
abstract
Autonomous driving system development is critically dependent on the ability to replay complex and diverse traffic scenarios in simulation. In such scenarios, the ability to accurately simulate the vehicle sensors such as cameras, lidar or radar is hugely helpful. However, current sensor simulators leverage gaming engines such as Unreal or Unity, requiring manual creation of environments, objects, and material properties. Such approaches have limited scalability and fail to produce realistic approximations of camera, lidar, and radar data without significant additional work. In this paper, we present a simple yet effective approach to generate realistic scenario sensor data, based only on a limited amount of lidar and camera data collected by an autonomous vehicle. Our approach uses texture-mapped surfels to efficiently reconstruct the scene from an initial vehicle pass or set of passes, preserving rich information about object 3D geometry and appearance, as well as the scene conditions. We then leverage a SurfelGAN network to reconstruct realistic camera images for novel positions and orientations of the self-driving vehicle and moving objects in the scene. We demonstrate our approach on the Waymo Open Dataset and show that it can synthesize realistic camera data for simulated scenarios. We also create a novel dataset that contains cases in which two self-driving vehicles observe the same scene at the same time. We use this dataset to provide additional evaluation and demonstrate the usefulness of our SurfelGAN model.
Zhenpei Yang, Yuning Chai, Dragomir Anguelov, Dumitru Erhan, Sean Rafferty, Henrik Kretzschmar
CVPR6
2020 Model Based Reinforcement Learning for Atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H. Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, Afroz Mohiuddin, Ryan Sepassi, George Tucker, Henryk Michalewski
ICLR7
2020 VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation
Manoj Kumar 0019, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn, Sergey Levine, Laurent Dinh, Durk Kingma
ICLR3
2019 A Benchmark for Interpretability Methods in Deep Neural Networks
abstract
We propose an empirical measure of the approximate accuracy of feature importance estimates in deep neural networks. Our results across several large-scale image classification datasets show that many popular interpretability methods produce estimates of feature importance that are not better than a random designation of feature importance. Only certain ensemble based approaches---VarGrad and SmoothGrad-Squared---outperform such a random assignment of importance. The manner of ensembling remains critical, we show that some approaches do no better then the underlying method but carry a far higher computational burden.
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, Been Kim
NeurIPS2
2019 High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
abstract
Predicting future video frames is extremely challenging, as there are many factors of variation that make up the dynamics of how frames change through time. Previously proposed solutions require complex inductive biases inside network architectures with highly specialized computation, including segmentation masks, optical flow, and foreground and background separation. In this work, we question if such handcrafted architectures are necessary and instead propose a different approach: finding minimal inductive bias for video prediction while maximizing network capacity. We investigate this question by performing the first large-scale empirical study and demonstrate state-of-the-art performance by learning large models on three different datasets: one for modeling object interactions, one for modeling human motion, and one for modeling car driving.
Ruben Villegas, Arkanath Pathak, Harini Kannan, Dumitru Erhan, Quoc V. Le, Honglak Lee
NeurIPS4
2018 Stochastic Variational Video Prediction
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H. Campbell, Sergey Levine
ICLR (Poster)3
2018 Learning how to explain neural networks: PatternNet and PatternAttribution
Pieter-Jan Kindermans, Kristof Schütt, Maximilian Alber, Klaus-Robert Müller, Dumitru Erhan, Been Kim, Sven Dähne
ICLR (Poster)5
2018 Hierarchical Long-term Video Prediction without Supervision
abstract
Much of recent research has been devoted to video prediction and generation, yet most of the previous works have demonstrated only limited success in generating videos on short-term horizons. The hierarchical video prediction method by Villegas et al. (2017) is an example of a state-of-the-art method for long-term video prediction, but their method is limited because it requires ground truth annotation of high-level structures (e.g., human joint landmarks) at training time. Our network encodes the input frame, predicts a high-level encoding into the future, and then a decoder with access to the first frame produces the predicted image from the predicted encoding. The decoder also produces a mask that outlines the predicted foreground object (e.g., person) as a by-product. Unlike Villegas et al. (2017), we develop a novel training method that jointly trains the encoder, the predictor, and the decoder together without highlevel supervision; we further improve upon this by using an adversarial loss in the feature space to train the predictor. Our method can predict about 20 seconds into the future and provides better results compared to Denton and Fergus (2018) and Finn et al. (2016) on the Human 3.6M dataset.
Nevan Wichers, Ruben Villegas, Dumitru Erhan, Honglak Lee
ICML3
2017 Unsupervised Pixel-Level Domain Adaptation with Generative Adversarial Networks
abstract
Collecting well-annotated image datasets to train modern machine learning algorithms is prohibitively expensive for many tasks. One appealing alternative is rendering synthetic data where ground-truth annotations are generated automatically. Unfortunately, models trained purely on rendered images fail to generalize to real images. To address this shortcoming, prior work introduced unsupervised domain adaptation algorithms that have tried to either map representations between the two domains, or learn to extract features that are domain-invariant. In this work, we approach the problem in a new light by learning in an unsupervised manner a transformation in the pixel space from one domain to the other. Our generative adversarial network (GAN)-based method adapts source-domain images to appear as if drawn from the target domain. Our approach not only produces plausible samples, but also outperforms the state-of-the-art on a number of unsupervised domain adaptation scenarios by large margins. Finally, we demonstrate that the adaptation process generalizes to object classes unseen during training.
Konstantinos Bousmalis, Nathan Silberman, David Dohan, Dumitru Erhan, Dilip Krishnan
CVPR4
2017 Show and Tell: Lessons Learned from the 2015 MSCOCO Image Captioning Challenge
abstract
Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In this paper, we present a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image. The model is trained to maximize the likelihood of the target description sentence given the training image. Experiments on several datasets show the accuracy of the model and the fluency of the language it learns solely from image descriptions. Our model is often quite accurate, which we verify both qualitatively and quantitatively. Finally, given the recent surge of interest in this task, a competition was organized in 2015 using the newly released COCO dataset. We describe and analyze the various improvements we applied to our own baseline and show the resulting performance in the competition, which we won ex-aequo with a team from Microsoft Research.
Oriol Vinyals, Alexander Toshev, Samy Bengio, Dumitru Erhan
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 SSD: Single Shot MultiBox Detector
Wei Liu 0015, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott E. Reed, Cheng-Yang Fu, Alexander C. Berg
ECCV (1)3
2016 Domain Separation Networks
abstract
The cost of large scale data collection and annotation often makes the application of machine learning algorithms to new tasks or datasets prohibitively expensive. One approach circumventing this cost is training models on synthetic data where annotations are provided automatically. Despite their appeal, such models often fail to generalize from synthetic to real images, necessitating domain adaptation algorithms to manipulate these models before they can be successfully applied. Existing approaches focus either on mapping representations from one domain to the other, or on learning to extract features that are invariant to the domain from which they were extracted. However, by focusing only on creating a mapping or shared representation between the two domains, they ignore the individual characteristics of each domain. We hypothesize that explicitly modeling what is unique to each domain can improve a model's ability to extract domain-invariant features. Inspired by work on private-shared component analysis, we explicitly learn to extract image representations that are partitioned into two subspaces: one component which is private to each domain and one which is shared across domains. Our model is trained to not only perform the task we care about in the source domain, but also to use the partitioned representation to reconstruct the images from both domains. Our novel architecture results in a model that outperforms the state-of-the-art on a range of unsupervised domain adaptation scenarios and additionally produces visualizations of the private and shared representations enabling interpretation of the domain adaptation process.
Konstantinos Bousmalis, George Trigeorgis, Nathan Silberman, Dilip Krishnan, Dumitru Erhan
NIPS5
2015 Going deeper with convolutions
abstract
We propose a deep convolutional neural network architecture codenamed Inception that achieves the new state of the art for classification and detection in the ImageNet Large-Scale Visual Recognition Challenge 2014 (ILSVRC14). The main hallmark of this architecture is the improved utilization of the computing resources inside the network. By a carefully crafted design, we increased the depth and width of the network while keeping the computational budget constant. To optimize quality, the architectural decisions were based on the Hebbian principle and the intuition of multi-scale processing. One particular incarnation used in our submission for ILSVRC14 is called GoogLeNet, a 22 layers deep network, the quality of which is assessed in the context of classification and detection.
Christian Szegedy, Wei Liu 0015, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, Andrew Rabinovich
CVPR7
2015 Show and tell: A neural image caption generator
abstract
Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing. In this paper, we present a generative model based on a deep recurrent architecture that combines recent advances in computer vision and machine translation and that can be used to generate natural sentences describing an image. The model is trained to maximize the likelihood of the target description sentence given the training image. Experiments on several datasets show the accuracy of the model and the fluency of the language it learns solely from image descriptions. Our model is often quite accurate, which we verify both qualitatively and quantitatively. For instance, while the current state-of-the-art BLEU-1 score (the higher the better) on the Pascal dataset is 25, our approach yields 59, to be compared to human performance around 69. We also show BLEU-1 score improvements on Flickr30k, from 56 to 66, and on SBU, from 19 to 28. Lastly, on the newly released COCO dataset, we achieve a BLEU-4 of 27.7, which is the current state-of-the-art.
Oriol Vinyals, Alexander Toshev, Samy Bengio, Dumitru Erhan
CVPR4
2015 Challenges in representation learning: A report on three machine learning contests
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron C. Courville, Mehdi Mirza, Benjamin Hamner, Will Cukierski, Yichuan Tang, Dave Thaler, Yingbo Zhou 0002, Chetan Ramaiah, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006, Dimitris Athanasakis, John Shawe-Taylor, Maxim Milakov, John Park, Radu Tudor Ionescu, Marius Popescu, Cristian Grozea, James Bergstra, Jingjing Xie, Lukasz Romaszko, Yoshua Bengio
Neural Networks2
2014 Scalable Object Detection Using Deep Neural Networks
abstract
Deep convolutional neural networks have recently achieved state-of-the-art performance on a number of image recognition benchmarks, including the ImageNet Large-Scale Visual Recognition Challenge (ILSVRC-2012). The winning model on the localization sub-task was a network that predicts a single bounding box and a confidence score for each object category in the image. Such a model captures the whole-image context around the objects but cannot handle multiple instances of the same object in the image without naively replicating the number of outputs for each instance. In this work, we propose a saliency-inspired neural network model for detection, which predicts a set of class-agnostic bounding boxes along with a single score for each box, corresponding to its likelihood of containing any object of interest. The model naturally handles a variable number of instances for each class and allows for cross-class generalization at the highest levels of the network. We are able to obtain competitive recognition performance on VOC2007 and ILSVRC2012, while using only the top few predicted locations in each image and a small number of neural network evaluations.
Dumitru Erhan, Christian Szegedy, Alexander Toshev, Dragomir Anguelov
CVPR1
2013 Challenges in Representation Learning: A Report on Three Machine Learning Contests
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron C. Courville, Mehdi Mirza, Benjamin Hamner, Will Cukierski, Yichuan Tang, Dave Thaler, Yingbo Zhou 0002, Chetan Ramaiah, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006, Dimitris Athanasakis, John Shawe-Taylor, Maxim Milakov, John Park, Radu Tudor Ionescu, Marius Popescu, Cristian Grozea, James Bergstra, Jingjing Xie, Lukasz Romaszko, Yoshua Bengio
ICONIP (3)2
2013 Deep Neural Networks for Object Detection
abstract
Deep Neural Networks (DNNs) have recently shown outstanding performance on the task of whole image classification. In this paper we go one step further and address the problem of object detection -- not only classifying but also precisely localizing objects of various classes using DNNs. We present a simple and yet powerful formulation of object detection as a regression to object masks. We define a multi-scale inference procedure which is able to produce a high-resolution object detection at a low cost by a few network applications. The approach achieves state-of-the-art performance on Pascal 2007 VOC.
Christian Szegedy, Alexander Toshev, Dumitru Erhan
NIPS3
2012 Sample-efficient Nonstationary Policy Evaluation for Contextual Bandits
Miroslav Dudík, Dumitru Erhan, John Langford 0001, Lihong Li 0001
UAI2
2010 Why Does Unsupervised Pre-training Help Deep Learning?
Dumitru Erhan, Yoshua Bengio, Aaron C. Courville, Pierre-Antoine Manzagol, Pascal Vincent, Samy Bengio
J. Mach. Learn. Res.1
2008 Zero-data Learning of New Tasks
Hugo Larochelle, Dumitru Erhan, Yoshua Bengio
AAAI2
2007 An empirical evaluation of deep architectures on problems with many factors of variation
abstract
Recently, several learning algorithms relying on models with deep architectures have been proposed. Though they have demonstrated impressive performance, to date, they have only been evaluated on relatively simple problems such as digit recognition in a controlled environment, for which many machine learning algorithms already report reasonable results. Here, we present a series of experiments which indicate that these models show promise in solving harder learning problems that exhibit many factors of variation. These models are compared with well-established algorithms such as Support Vector Machines and single hidden-layer feed-forward neural networks.
Hugo Larochelle, Dumitru Erhan, Aaron C. Courville, James Bergstra, Yoshua Bengio
ICML2
2006 Aggregate features and ADABOOSTfor music classification
James Bergstra, Norman Casagrande, Dumitru Erhan, Douglas Eck, Balázs Kégl
Mach. Learn.3