VLDB 2026 Research / reviewers in the wild / expert
Iulian Serban
dblp:164/5610 · also Iulian Vlad Serban
· DBLP profile ↗
17ranked-venue papers
7as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Question answering and dialogue systems · 40% Information extraction and text analysis · 16% Probabilistic and Bayesian machine learning · 14% | |
| Human-computer interaction and pervasive computing
2 papers |
Learning and educational technologies · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computing education · 100% |
Topics — the 16 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
question generation |
1.2 | 3 | 2024 | How Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational Quizzes · AAAI 2024 Generating Factoid Questions With Recurrent Neural Networks: The 30M Factoid Question-Answer Corpus · ACL (1) 2016 Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage Retrieval · EMNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation |
1.1 | 4 | 2017 | A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues · AAAI 2017 Multiresolution Recurrent Neural Networks: An Application to Dialogue Response Generation · AAAI 2017 How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis
discourse analysis |
0.5 | 1 | 2021 | Deep Discourse Analysis for Generating Personalized Feedback in Intelligent Tutor Systems · AAAI 2021 |
Natural language and speech › Information extraction and text analysis › text segmentation
discourse segmentation |
0.5 | 1 | 2021 | Deep Discourse Analysis for Generating Personalized Feedback in Intelligent Tutor Systems · AAAI 2021 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.5 | 1 | 2021 | Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage Retrieval · EMNLP (1) 2021 |
Learning and educational technologies
intelligent tutoring systems |
0.5 | 1 | 2021 | Deep Discourse Analysis for Generating Personalized Feedback in Intelligent Tutor Systems · AAAI 2021 |
Learning and educational technologies › educational feedback
personalized feedback |
0.5 | 1 | 2021 | Deep Discourse Analysis for Generating Personalized Feedback in Intelligent Tutor Systems · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems › dialogue evaluation
dialogue response evaluation |
0.3 | 1 | 2017 | Towards an Automatic Turing Test: Learning to Evaluate Dialogue Responses · ACL (1) 2017 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.3 | 1 | 2017 | A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues · AAAI 2017 |
Natural language and speech › Language models and text generation › large language model evaluation › automatic evaluation
learned evaluation metric |
0.3 | 1 | 2017 | Towards an Automatic Turing Test: Learning to Evaluate Dialogue Responses · ACL (1) 2017 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › amortized inference
neural variational inference |
0.3 | 1 | 2017 | Piecewise Latent Variables for Neural Variational Text Processing · EMNLP 2017 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2017 | Multiresolution Recurrent Neural Networks: An Application to Dialogue Response Generation · AAAI 2017 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.3 | 1 | 2017 | Piecewise Latent Variables for Neural Variational Text Processing · EMNLP 2017 |
Natural language and speech › Machine translation › machine translation evaluation
automatic evaluation metrics |
0.2 | 1 | 2016 | How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation · EMNLP 2016 |
Information retrieval › document retrieval
passage retrieval |
0.1 | 1 | 2021 | Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage Retrieval · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › text generation › neural text generation
recurrent neural network text generation |
0.1 | 1 | 2016 | Generating Factoid Questions With Recurrent Neural Networks: The 30M Factoid Question-Answer Corpus · ACL (1) 2016 |
Methods — techniques the papers use, named apart from their topics
large language model · 2.3self-training · 1.0relational graph · 1.0neural discourse segmentation · 1.0neural classifier · 1.0consistency filtering · 1.0stochastic latent variable · 0.3sequence-to-sequence · 0.3multiresolution recurrent neural network · 0.3hierarchical latent variable encoder-decoder · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | How Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational QuizzesabstractQuestion generation (QG) is a natural language processing task with an abundance of potential benefits and use cases in the educational domain. In order for this potential to be realized, QG systems must be designed and validated with pedagogical needs in mind. However, little research has assessed or designed QG approaches with the input of real teachers or students. This paper applies a large language model-based QG approach where questions are generated with learning goals derived from Bloom's taxonomy. The automatically generated questions are used in multiple experiments designed to assess how teachers use them in practice. The results demonstrate that teachers prefer to write quizzes with automatically generated questions, and that such quizzes have no loss in quality compared to handwritten versions. Further, several metrics indicate that automatically generated questions can even improve the quality of the quizzes created, showing the promise for large scale use of QG in the classroom setting. Sabina Elkins, Ekaterina Kochmar, Jackie Chi Kit Cheung, Iulian Serban |
AAAI | 4 |
| 2022 | Raising Student Completion Rates with Adaptive Curriculum and Contextual Bandits
Robert Belfer, Ekaterina Kochmar, Iulian Serban |
AIED (1) | 3 |
| 2021 | Deep Discourse Analysis for Generating Personalized Feedback in Intelligent Tutor SystemsabstractWe explore creating automated, personalized feedback in an intelligent tutoring system (ITS). Our goal is to pinpoint correct and incorrect concepts in student answers in order to achieve better student learning gains. Although automatic methods for providing personalized feedback exist, they do not explicitly inform students about which concepts in their answers are correct or incorrect. Our approach involves decomposing students answers using neural discourse segmentation and classification techniques. This decomposition yields a relational graph over all discourse units covered by the reference solutions and student answers. We use this inferred relational graph structure and a neural classifier to match student answers with reference solutions and generate personalized feedback. Although the process is completely automated and data-driven, the personalized feedback generated is highly contextual, domain-aware and effectively targets each student's misconceptions and knowledge gaps. We test our method in a dialogue-based ITS and demonstrate that our approach results in high-quality feedback and significantly improved student learning gains. Matt Grenander, Robert Belfer, Ekaterina Kochmar, Iulian Serban, François St-Hilaire, Jackie Chi Kit Cheung |
AAAI | 4 |
| 2021 | A Comparative Study of Learning Outcomes for Online Learning Platforms
François St-Hilaire, Nathan Burns, Robert Belfer, Muhammad Shayan, Ariella Smofsky, Dung Do Vu, Antoine Frau, Joseph Potochny, Farid Faraji, Vincent Pavero, Neroli Ko, Ansona Onyi Ching, Sabina Elkins, Anush Stepanyan, Adela Matajova, Laurent Charlin, Yoshua Bengio, Iulian Serban, Ekaterina Kochmar |
AIED (2) | 18 |
| 2021 | Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage RetrievalabstractIn this work, we introduce back-training, an alternative to self-training for unsupervised domain adaptation (UDA) from source to target domain.While self-training generates synthetic training data where natural inputs are aligned with noisy outputs, back-training results in natural outputs aligned with noisy inputs.This significantly reduces the gap between the target domain and synthetic data distribution, and reduces model overfitting to the source domain.We run UDA experiments on question generation and passage retrieval from the Natural Questions domain to machine learning and biomedical domains.We find that back-training vastly outperforms selftraining by a mean improvement of 7.8 BLEU-4 points on generation, and 17.6% top-20 retrieval accuracy across both domains.We further propose consistency filters to remove low-quality synthetic data before training.We also release a new domain-adaptation dataset-MLQuestions containing 35K unaligned questions, 50K unaligned passages, and 3K aligned question-passage pairs. Devang Kulshreshtha, Robert Belfer, Iulian Serban, Siva Reddy |
EMNLP (1) | 3 |
| 2020 | Automated Personalized Feedback Improves Learning Gains in An Intelligent Tutoring System
Ekaterina Kochmar, Dung Do Vu, Robert Belfer, Iulian Serban, Joelle Pineau |
AIED (2) | 5 |
| 2020 | A Large-Scale, Open-Domain, Mixed-Interface Dialogue-Based ITS for STEM
Iulian Serban, Ekaterina Kochmar, Dung Do Vu, Robert Belfer, Joelle Pineau, Aaron C. Courville, Laurent Charlin, Yoshua Bengio |
AIED (2) | 1 |
| 2020 | The Bottleneck Simulator: A Model-Based Deep Reinforcement Learning ApproachabstractDeep reinforcement learning has recently shown many impressive successes. However, one major obstacle towards applying such methods to real-world problems is their lack of data-efficiency. To this end, we propose the Bottleneck Simulator: a model-based reinforcement learning method which combines a learned, factorized transition model of the environment with rollout simulations to learn an effective policy from few examples. The learned transition model employs an abstract, discrete (bottleneck) state, which increases sample efficiency by reducing the number of model parameters and by exploiting structural properties of the environment. We provide a mathematical analysis of the Bottleneck Simulator in terms of fixed points of the learned policy, which reveals how performance is affected by four distinct sources of error: an error related to the abstract space structure, an error related to the transition model estimation variance, an error related to the transition model estimation bias, and an error related to the transition model class bias. Finally, we evaluate the Bottleneck Simulator on two natural language processing tasks: a text adventure game and a real-world, complex dialogue response selection task. On both tasks, the Bottleneck Simulator yields excellent performance beating competing approaches. Iulian Serban, Chinnadhurai Sankar, Joelle Pineau, Yoshua Bengio |
J. Artif. Intell. Res. | 1 |
| 2017 | Multiresolution Recurrent Neural Networks: An Application to Dialogue Response GenerationabstractWe introduce a new class of models called multiresolution recurrent neural networks, which explicitly model natural language generation at multiple levels of abstraction. The models extend the sequence-to-sequence framework to generate two parallel stochastic processes: a sequence of high-level coarse tokens, and a sequence of natural language words (e.g. sentences). The coarse sequences follow a latent stochastic process with a factorial representation, which helps the models generalize to new examples. The coarse sequences can also incorporate task-specific knowledge, when available. In our experiments, the coarse sequences are extracted using automatic procedures, which are designed to capture compositional structure and semantics. These procedures enable training the multiresolution recurrent neural networks by maximizing the exact joint log-likelihood over both sequences. We apply the models to dialogue response generation in the technical support domain and compare them with several competing models. The multiresolution recurrent neural networks outperform competing models by a substantial margin, achieving state-of-the-art results according to both a human evaluation study and automatic evaluation metrics. Furthermore, experiments show the proposed models generate more fluent, relevant and goal-oriented responses. Iulian Serban, Tim Klinger, Gerald Tesauro, Kartik Talamadupula, Bowen Zhou 0002, Yoshua Bengio, Aaron C. Courville |
AAAI | 1 |
| 2017 | A Hierarchical Latent Variable Encoder-Decoder Model for Generating DialoguesabstractSequential data often possesses hierarchical structures with complex dependencies between sub-sequences, such as found between the utterances in a dialogue. To model these dependencies in a generative framework, we propose a neural network-based generative architecture, with stochastic latent variables that span a variable number of time steps. We apply the proposed model to the task of dialogue response generation and compare it with other recent neural-network architectures. We evaluate the model performance through a human evaluation study. The experiments demonstrate that our model improves upon recently proposed models and that the latent variables facilitate both the generation of meaningful, long and diverse responses and maintaining dialogue state. Iulian Serban, Alessandro Sordoni, Ryan Lowe, Laurent Charlin, Joelle Pineau, Aaron C. Courville, Yoshua Bengio |
AAAI | 1 |
| 2017 | Towards an Automatic Turing Test: Learning to Evaluate Dialogue ResponsesabstractRyan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, Joelle Pineau. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Ryan Lowe, Michael Noseworthy, Iulian Serban, Nicolas Angelard-Gontier, Yoshua Bengio, Joelle Pineau |
ACL (1) | 3 |
| 2017 | Piecewise Latent Variables for Neural Variational Text ProcessingabstractAdvances in neural variational inference have facilitated the learning of powerful directed graphical models with continuous latent variables, such as variational autoencoders.The hope is that such models will learn to represent rich, multi-modal latent factors in real-world data, such as natural language text.However, current models often assume simplistic priors on the latent variables -such as the uni-modal Gaussian distributionwhich are incapable of representing complex latent factors efficiently.To overcome this restriction, we propose the simple, but highly flexible, piecewise constant distribution.This distribution has the capacity to represent an exponential number of modes of a latent target distribution, while remaining mathematically tractable.Our results demonstrate that incorporating this new latent distribution into different models yields substantial improvements in natural language processing tasks such as document modeling and natural language generation for dialogue. Iulian Serban, Alexander Ororbia, Joelle Pineau, Aaron C. Courville |
EMNLP | 1 |
| 2016 | Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network ModelsabstractWe investigate the task of building open domain, conversational dialogue systems based on large dialogue corpora using generative models. Generative models produce system responses that are autonomously generated word-by-word, opening up the possibility for realistic, flexible interactions. In support of this goal, we extend the recently proposed hierarchical recurrent encoder-decoder neural network to the dialogue domain, and demonstrate that this model is competitive with state-of-the-art neural language models and back-off n-gram models. We investigate the limitations of this and similar approaches, and show how its performance can be improved by bootstrapping the learning from a larger question-answer pair corpus and from pretrained word embeddings. Iulian Serban, Alessandro Sordoni, Yoshua Bengio, Aaron C. Courville, Joelle Pineau |
AAAI | 1 |
| 2016 | Generating Factoid Questions With Recurrent Neural Networks: The 30M Factoid Question-Answer CorpusabstractIulian Vlad Serban, Alberto García-Durán, Caglar Gulcehre, Sungjin Ahn, Sarath Chandar, Aaron Courville, Yoshua Bengio. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Iulian Serban, Alberto García-Durán, Caglar Gulcehre, Sungjin Ahn, Sarath Chandar, Aaron C. Courville, Yoshua Bengio |
ACL (1) | 1 |
| 2016 | How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response GenerationabstractWe investigate evaluation metrics for dialogue response generation systems where supervised labels, such as task completion, are not available.Recent works in response generation have adopted metrics from machine translation to compare a model's generated response to a single target response.We show that these metrics correlate very weakly with human judgements in the non-technical Twitter domain, and not at all in the technical Ubuntu domain.We provide quantitative and qualitative results highlighting specific weaknesses in existing metrics, and provide recommendations for future development of better automatic evaluation metrics for dialogue systems. Chia-Wei Liu, Ryan Lowe, Iulian Serban, Michael Noseworthy, Laurent Charlin, Joelle Pineau |
EMNLP | 3 |
| 2016 | On the Evaluation of Dialogue Systems with Next Utterance ClassificationabstractAn open challenge in constructing dialogue systems is developing methods for automatically learning dialogue strategies from large amounts of unlabelled data. Recent work has proposed Next-Utterance-Classification (NUC) as a surrogate task for building dialogue systems from text data. In this paper we investigate the performance of humans on this task to validate the relevance of NUC as a method of evaluation. Our results show three main findings: (1) humans are able to correctly classify responses at a rate much better than chance, thus confirming that the task is feasible, (2) human performance levels vary across task domains (we consider 3 datasets) and expertise levels (novice vs experts), thus showing that a range of performance is possible on this type of task, (3) automated dialogue systems built using state-of-the-art machine learning methods have similar performance to the human novices, but worse than the experts, thus confirming the utility of this class of tasks for driving further research in automated dialogue systems. Ryan Lowe, Iulian Serban, Michael Noseworthy, Laurent Charlin, Joelle Pineau |
SIGDIAL Conference | 2 |
| 2015 | The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue SystemsabstractThis paper introduces the Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words.This provides a unique resource for research into building dialogue managers based on neural language models that can make use of large amounts of unlabeled data.The dataset has both the multi-turn property of conversations in the Dialog State Tracking Challenge datasets, and the unstructured nature of interactions from microblog services such as Twitter.We also describe two neural learning architectures suitable for analyzing this dataset, and provide benchmark performance on the task of selecting the best next response. Ryan Lowe, Nissan Pow, Iulian Serban, Joelle Pineau |
SIGDIAL Conference | 3 |