VLDB 2026 Research / reviewers in the wild / expert
Emily Dinan
dblp:213/7927
· DBLP profile ↗
17ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0003-0624-6311ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Question answering and dialogue systems · 30% Language models and text generation · 30% Trustworthy machine learning · 23% | |
| Human-computer interaction and pervasive computing
1 paper |
Games and playful interaction · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Software testing · 100% |
Topics — the 22 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
open-domain dialogue |
1.0 | 2 | 2022 | SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems · ACL (1) 2022 The Dialogue Dodecathlon: Open-Domain Knowledge and Image Grounded Conversational Agents · ACL 2020 |
Machine learning › Trustworthy machine learning
fairness |
0.9 | 2 | 2020 | Multi-Dimensional Gender Bias Classification · EMNLP (1) 2020 Queens are Powerful too: Mitigating Gender Bias in Dialogue Generation · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | BTS: Harmonizing Specialized Experts into a Generalist LLM · EMNLP 2025 |
Machine learning › Efficient and distributed learning
model merging |
0.9 | 1 | 2025 | BTS: Harmonizing Specialized Experts into a Generalist LLM · EMNLP 2025 |
Machine learning › Trustworthy machine learning
safety evaluation |
0.6 | 1 | 2022 | SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems · ACL (1) 2022 |
Natural language and speech › Question answering and dialogue systems
dialogue generation |
0.4 | 1 | 2020 | Queens are Powerful too: Mitigating Gender Bias in Dialogue Generation · EMNLP (1) 2020 |
Machine learning › Trustworthy machine learning › fairness › bias mitigation
gender bias mitigation |
0.4 | 1 | 2020 | Queens are Powerful too: Mitigating Gender Bias in Dialogue Generation · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.4 | 1 | 2020 | Adversarial NLI: A New Benchmark for Natural Language Understanding · ACL 2020 |
Natural language and speech › Language models and text generation › text generation
neural text generation |
0.4 | 1 | 2020 | Neural Text Generation With Unlikelihood Training · ICLR 2020 |
Natural language and speech › Language models and text generation
unlikelihood training |
0.4 | 1 | 2020 | Neural Text Generation With Unlikelihood Training · ICLR 2020 |
Games and playful interaction › game AI
procedural content generation |
0.4 | 1 | 2020 | Generating Interactive Worlds with Text · AAAI 2020 |
Software testing › test infrastructure
benchmark construction |
0.4 | 1 | 2020 | Adversarial NLI: A New Benchmark for Natural Language Understanding · ACL 2020 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.4 | 1 | 2019 | Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Question answering and dialogue systems
dialogue |
0.4 | 1 | 2019 | Learning to Speak and Act in a Fantasy Text Adventure Game · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Question answering and dialogue systems › dialogue
dialogue safety |
0.4 | 1 | 2019 | Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack · EMNLP/IJCNLP (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph |
0.4 | 1 | 2019 | Wizard of Wikipedia: Knowledge-Powered Conversational Agents · ICLR (Poster) 2019 |
Natural language and speech › Question answering and dialogue systems
knowledge-grounded dialogue |
0.4 | 1 | 2019 | Wizard of Wikipedia: Knowledge-Powered Conversational Agents · ICLR (Poster) 2019 |
Natural language and speech › Question answering and dialogue systems
personalized dialogue |
0.3 | 1 | 2018 | Personalizing Dialogue Agents: I have a dog, do you have pets too? · ACL (1) 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.1 | 1 | 2020 | Generating Interactive Worlds with Text · AAAI 2020 |
Computer vision › Vision and language › multimodal dialogue
image-grounded dialogue |
0.1 | 1 | 2020 | The Dialogue Dodecathlon: Open-Domain Knowledge and Image Grounded Conversational Agents · ACL 2020 |
Natural language and speech › Information extraction and text analysis
text classification |
0.1 | 1 | 2020 | Multi-Dimensional Gender Bias Classification · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue policy learning |
0.1 | 1 | 2018 | Personalizing Dialogue Agents: I have a dog, do you have pets too? · ACL (1) 2018 |
Methods — techniques the papers use, named apart from their topics
neural network · 0.9compositional generation · 0.9adversarial data collection · 0.9reinforcement learning · 0.7targeted data collection · 0.4multi-task learning · 0.4counterfactual data augmentation · 0.4bias controlled training · 0.4bias classifiers · 0.4BERT · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BTS: Harmonizing Specialized Experts into a Generalist LLMabstractQizhen Zhang, Prajjwal Bhargava, Chloe Bi, Chris X. Cai, Jakob Nicolaus Foerster, Jeremy Fu, Punit Singh Koura, Ruan Silva, Sheng Shen, Emily Dinan, Suchin Gururangan, Mike Lewis. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Qizhen Zhang 0002, Prajjwal Bhargava, Chloe Bi, Chris X. Cai, Jakob N. Foerster, Jeremy Fu, Punit Singh Koura, Ruan Silva, Sheng Shen 0016, Emily Dinan, Suchin Gururangan, Mike Lewis |
EMNLP | 10 |
| 2024 | When Life Gives You Lemons, Make Cherryade: Converting Feedback from Bad Responses into Good LabelsabstractWeiyan Shi, Emily Dinan, Kurt Shuster, Jason Weston, Jing Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Weiyan Shi 0001, Emily Dinan, Kurt Shuster 0001, Jason Weston, Jing Xu 0014 |
NAACL-HLT | 2 |
| 2022 | SafetyKit: First Aid for Measuring Safety in Open-domain Conversational SystemsabstractEmily Dinan, Gavin Abercrombie, A. Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, Verena Rieser. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Emily Dinan, Gavin Abercrombie, A. Stevie Bergman, Shannon L. Spruit, Dirk Hovy, Y-Lan Boureau, Verena Rieser |
ACL (1) | 1 |
| 2022 | Guiding the Release of Safer E2E Conversational AI through Value Sensitive DesignabstractA. Stevie Bergman, Gavin Abercrombie, Shannon Spruit, Dirk Hovy, Emily Dinan, Y-Lan Boureau, Verena Rieser. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022. A. Stevie Bergman, Gavin Abercrombie, Shannon L. Spruit, Dirk Hovy, Emily Dinan, Y-Lan Boureau, Verena Rieser |
SIGDIAL | 5 |
| 2022 | Reducing Conversational Agents' Overconfidence through Linguistic CalibrationabstractAbstract While improving neural dialogue agents’ factual accuracy is the object of much research, another important aspect of communication, less studied in the setting of neural dialogue, is transparency about ignorance. In this work, we analyze to what extent state-of-the-art chit-chat models are linguistically calibrated in the sense that their verbalized expression of doubt (or confidence) matches the likelihood that the model’s responses are factually incorrect (or correct). We find that these models are poorly calibrated, yet we show that likelihood of correctness can accurately be predicted. By incorporating such metacognitive features into the training of a controllable generation model, we obtain a dialogue agent with greatly improved linguistic calibration. Sabrina J. Mielke, Arthur Szlam, Emily Dinan, Y-Lan Boureau |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Recipes for Building an Open-Domain ChatbotabstractStephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, Jason Weston. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Stephen Roller, Emily Dinan, Naman Goyal 0001, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu 0014, Myle Ott, Eric Michael Smith, Y-Lan Boureau, Jason Weston |
EACL | 2 |
| 2021 | Bot-Adversarial Dialogue for Safe Conversational AgentsabstractJing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, Emily Dinan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jing Xu 0014, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, Emily Dinan |
NAACL-HLT | 6 |
| 2020 | Generating Interactive Worlds with TextabstractProcedurally generating cohesive and interesting game environments is challenging and time-consuming. In order for the relationships between the game elements to be natural, common-sense has to be encoded into arrangement of the elements. In this work, we investigate a machine learning approach for world creation using content from the multi-player text adventure game environment LIGHT (Urbanek et al. 2019). We introduce neural network based models to compositionally arrange locations, characters, and objects into a coherent whole. In addition to creating worlds based on existing elements, our models can generate new game content. Humans can also leverage our models to interactively aid in worldbuilding. We show that the game environments created with our approach are cohesive, diverse, and preferred by human evaluators compared to other machine learning based world construction algorithms. Angela Fan, Jack Urbanek, Pratik Ringshia, Emily Dinan, Emma Qian, Siddharth Karamcheti, Shrimai Prabhumoye, Douwe Kiela, Tim Rocktäschel, Arthur Szlam, Jason Weston |
AAAI | 4 |
| 2020 | Adversarial NLI: A New Benchmark for Natural Language UnderstandingabstractWe introduce a new large-scale NLI benchmark dataset, collected via an iterative, adversarial human-and-model-in-the-loop procedure.We show that training models on this new dataset leads to state-of-the-art performance on a variety of popular NLI benchmarks, while posing a more difficult challenge with its new test set.Our analysis sheds light on the shortcomings of current state-of-theart models, and shows that non-expert annotators are successful at finding their weaknesses.The data collection method can be applied in a never-ending learning scenario, becoming a moving target for NLU, rather than a static benchmark that will quickly saturate. Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, Douwe Kiela |
ACL | 3 |
| 2020 | The Dialogue Dodecathlon: Open-Domain Knowledge and Image Grounded Conversational AgentsabstractWe introduce dodecaDialogue: a set of 12 tasks that measures if a conversational agent can communicate engagingly with personality and empathy, ask questions, answer questions by utilizing knowledge resources, discuss topics and situations, and perceive and converse about images.By multi-tasking on such a broad large-scale set of data, we hope to both move towards and measure progress in producing a single unified agent that can perceive, reason and converse with humans in an open-domain setting.We show that such multi-tasking improves over a BERT pretrained baseline, largely due to multi-tasking with very large dialogue datasets in a similar domain, and that the multi-tasking in general provides gains to both text and image-based tasks using several metrics in both the finetune and task transfer settings.We obtain stateof-the-art results on many of the tasks, providing a strong baseline for this challenge. Kurt Shuster 0001, Da Ju, Stephen Roller, Emily Dinan, Y-Lan Boureau, Jason Weston |
ACL | 4 |
| 2020 | Queens are Powerful too: Mitigating Gender Bias in Dialogue GenerationabstractModels often easily learn biases present in the training data, and their predictions directly reflect this bias.We analyze gender bias in dialogue data, and examine how this bias is actually amplified in subsequent generative chit-chat dialogue models.We measure gender bias in six existing dialogue datasets, and focus on the most biased one, the multiplayer text-based fantasy adventure dataset LIGHT (Urbanek et al., 2019), as a testbed for our bias mitigation techniques.The LIGHT dataset is highly imbalanced with respect to gender, containing predominantly male characters, likely because it is entirely collected by crowdworkers and reflects common biases that exist in fantasy or medieval settings.We consider three techniques to mitigate gender bias: counterfactual data augmentation, targeted data collection, and bias controlled training.We show that our proposed techniques mitigate gender bias in LIGHT by balancing the genderedness of generated dialogue utterances and are particularly effective in combination.We quantify performance using various evaluation methods-such as quantity of gendered words, a dialogue safety classifier, and human studies-all of which show that our models generate less gendered, but equally engaging chit-chat responses. Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, Jason Weston |
EMNLP (1) | 1 |
| 2020 | Multi-Dimensional Gender Bias ClassificationabstractMachine learning models are trained to find patterns in data.NLP models can inadvertently learn socially undesirable patterns when training on gender biased text.In this work, we propose a novel, general framework that decomposes gender bias in text along several pragmatic and semantic dimensions: bias from the gender of the person being spoken about, bias from the gender of the person being spoken to, and bias from the gender of the speaker.Using this fine-grained framework, we automatically annotate eight large scale datasets with gender information.In addition, we collect a new, crowdsourced evaluation benchmark.Distinguishing between gender bias along multiple dimensions enables us to train better and more fine-grained gender bias classifiers.We show our classifiers are valuable for a variety of applications, like controlling for gender bias in generative models, detecting gender bias in arbitrary text, and classifying text as offensive based on its genderedness. Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, Adina Williams |
EMNLP (1) | 1 |
| 2020 | Neural Text Generation With Unlikelihood Training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, Jason Weston |
ICLR | 4 |
| 2019 | Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human AttackabstractEmily Dinan, Samuel Humeau, Bharath Chintagunta, Jason Weston. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Emily Dinan, Samuel Humeau 0001, Bharath Chintagunta, Jason Weston |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Learning to Speak and Act in a Fantasy Text Adventure GameabstractJack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur Szlam, Jason Weston. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau 0001, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur Szlam, Jason Weston |
EMNLP/IJCNLP (1) | 6 |
| 2019 | Wizard of Wikipedia: Knowledge-Powered Conversational Agents
Emily Dinan, Stephen Roller, Kurt Shuster 0001, Angela Fan, Michael Auli, Jason Weston |
ICLR (Poster) | 1 |
| 2018 | Personalizing Dialogue Agents: I have a dog, do you have pets too?abstractSaizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, Jason Weston. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, Jason Weston |
ACL (1) | 2 |