VLDB 2026 Research / reviewers in the wild / expert
Behnam Hedayatnia
dblp:194/7461
· DBLP profile ↗
21ranked-venue papers
3as first author
12since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Overview of the Ninth Dialog System Technology Challenge: DSTC9abstractThis paper introduces the Ninth Dialog System Technology Challenge (DSTC-9). This edition of the DSTC focuses on applying end-to-end dialog technologies for four distinct tasks in dialog systems, namely, 1. Task-oriented dialog Modeling with Unstructured Knowledge Access, 2. Multi-domain task-oriented dialog, 3. Interactive evaluation of dialog and 4. Situated interactive multimodal dialog. This paper describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. R. Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro, Abhinav Rastogi, Yun-Nung Chen, Mihail Eric, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Yang Liu 0004, Chao-Wei Huang, Dilek Hakkani-Tür, Jinchao Li, Qi Zhu 0007, Lingxiao Luo, Lars Liden, Kaili Huang, Shahin Shayandeh, Runze Liang, Baolin Peng, Zheng Zhang 0020, Swadheen Shukla, Minlie Huang, Jianfeng Gao 0001, Shikib Mehri, Yulan Feng, Carla Gordon, Seyed Hossein Alavi, David R. Traum, Maxine Eskénazi, Ahmad Beirami, Eunjoon Cho, Paul A. Crook, Ankita De, Alborz Geramifard, Satwik Kottur, Seungwhan Moon, Shivani Poddar, Rajen Subba |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Towards Credible Human Evaluation of Open-Domain Dialog Systems Using Interactive SetupabstractEvaluating open-domain conversation models has been an open challenge due to the open-ended nature of conversations. In addition to static evaluations, recent work has started to explore a variety of per-turn and per-dialog interactive evaluation mechanisms and provide advice on the best setup. In this work, we adopt the interactive evaluation framework and further apply to multiple models with a focus on per-turn evaluation techniques. Apart from the widely used setting where participants select the best response among different candidates at each turn, one more novel per-turn evaluation setting is adopted, where participants can select all appropriate responses with different fallback strategies to continue the conversation when no response is selected. We evaluate these settings based on sensitivity and consistency using four GPT2-based models that differ in model sizes or fine-tuning data. To better generalize to any model groups with no prior assumptions on their rankings and control evaluation costs for all setups, we also propose a methodology to estimate the required sample size given a minimum performance gap of interest before running most experiments. Our comprehensive human evaluation results shed light on how to conduct credible human evaluations of open domain dialog systems using the interactive setup, and suggest additional future directions. Sijia Liu 0007, Patrick Lange, Behnam Hedayatnia, Alexandros Papangelis, Di Jin 0005, Andrew Wirth, Yang Liu 0004, Dilek Hakkani-Tür |
AAAI | 3 |
| 2023 | MERCY: Multiple Response Ranking Concurrently in Realistic Open-Domain Conversational SystemsabstractAutomatic Evaluation (AE) and Response Selection (RS) models assign quality scores to various candidate responses and rank them in conversational setups.Prior response ranking research compares various models' performance on synthetically generated test sets.In this work, we investigate the performance of model-based reference-free AE and RS models on our constructed response ranking datasets that mirror real-case scenarios of ranking candidates during inference time.Metrics' unsatisfying performance can be interpreted as their low generalizability over more pragmatic conversational domains such as human-chatbot dialogs.To alleviate this issue we propose a novel RS model called MERCY that simulates human behavior in selecting the best candidate by taking into account distinct candidates concurrently and learns to rank them.In addition, MERCY leverages natural language feedback as another component to help the ranking task by explaining why each candidate response is relevant/irrelevant to the dialog context.These feedbacks are generated by prompting large language models in a few-shot setup.Our experiments show the better performance of MERCY over baselines for the response ranking task in our curated realistic datasets. Sarik Ghazarian, Behnam Hedayatnia, Di Jin 0005, Sijia Liu 0007, Nanyun Peng 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 2 |
| 2023 | Investigating the Representation of Open Domain Dialogue Context for Transformer ModelsabstractVishakh Padmakumar, Behnam Hedayatnia, Di Jin, Patrick Lange, Seokhwan Kim, Nanyun Peng, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Vishakh Padmakumar, Behnam Hedayatnia, Di Jin 0005, Patrick Lange, Seokhwan Kim, Nanyun Peng 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 2 |
| 2023 | "What do others think?": Task-Oriented Conversational Modeling with Subjective KnowledgeabstractChao Zhao, Spandana Gella, Seokhwan Kim, Di Jin, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Spandana Gella, Seokhwan Kim, Di Jin 0005, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 7 |
| 2022 | Think Before You Speak: Explicitly Generating Implicit Commonsense Knowledge for Response GenerationabstractPei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren 0001, Yang Liu 0004, Dilek Hakkani-Tür |
ACL (1) | 3 |
| 2022 | A Systematic Evaluation of Response Selection for Open Domain DialogueabstractRecent progress on neural approaches for language processing has triggered a resurgence of interest on building intelligent open-domain chatbots.However, even the state-of-the-art neural chatbots cannot produce satisfying responses for every turn in a dialog.A practical solution is to generate multiple response candidates for the same context, and then perform response ranking/selection to determine which candidate is the best.Previous work in response selection typically trains response rankers using synthetic data that is formed from existing dialogs by using a ground truth response as the single appropriate response and constructing inappropriate responses via random selection or using adversarial methods.In this work, we curated a dataset where responses from multiple response generators produced for the same dialog context are manually annotated as appropriate (positive) and inappropriate (negative).We argue that such training data better matches the actual use case examples, enabling the models to learn to rank responses effectively.With this new dataset, we conduct a systematic evaluation of state-of-the-art methods for response selection, and demonstrate that both strategies of using multiple positive candidates and using manually verified hard negative candidates can bring in significant performance improvement in comparison to using the adversarial training data, e.g., increase of 3% and 13% in Recall@1 score, respectively. Behnam Hedayatnia, Di Jin 0005, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 1 |
| 2021 | "How Robust R U?": Evaluating Task-Oriented Dialogue Systems on Spoken ConversationsabstractMost prior work in dialogue modeling has been on written conversations mostly because of existing data sets. However, written dialogues are not sufficient to fully capture the nature of spoken conversations as well as the potential speech recognition errors in practical spoken dialogue systems. This work presents a new benchmark on spoken task-oriented conversations, which is intended to study multi-domain dialogue state tracking and knowledge-grounded dialogue modeling. We report that the existing state-of-the-art models trained on written conversations are not performing well on our spoken data, as expected. Furthermore, we observe improvements in task performances when leveraging$n$-best speech recognition hypotheses such as by combining predictions based on individual hypotheses. Our data set enables speech-based benchmarking of task-oriented dialogue systems. Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Dilek Hakkani-Tür |
ASRU | 6 |
| 2021 | Multi-Sentence Knowledge Selection in Open-Domain DialogueabstractMihail Eric, Nicole Chartier, Behnam Hedayatnia, Karthik Gopalakrishnan, Pankaj Rajan, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 14th International Conference on Natural Language Generation. 2021. Mihail Eric, Nicole Chartier, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Pankaj Rajan, Yang Liu 0004, Dilek Hakkani-Tür |
INLG | 3 |
| 2021 | Commonsense-Focused Dialogues for Response Generation: An Empirical StudyabstractPei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 3 |
| 2021 | Go Beyond Plain Fine-Tuning: Improving Pretrained Models for Social CommonsenseabstractPretrained language models have demonstrated outstanding performance in many NLP tasks recently. However, their social intelligence, which requires commonsense reasoning about the current situation and mental states of others, is still developing. Towards improving language models' social intelligence, in this study we focus on the Social IQA dataset, a task requiring social and emotional commonsense reasoning. Building on top of the pretrained RoBERTa and GPT2 models, we propose several architecture variations and extensions, as well as leveraging external commonsense corpora, to optimize the model for Social IQA. Our proposed system achieves competitive results as those top-ranking models on the leaderboard. This work demonstrates the strengths of pretrained language models, and provides viable ways to improve their performance for a particular task. Ting-Yun Chang, Yang Liu 0004, Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Dilek Hakkani-Tür |
SLT | 4 |
| 2020 | Policy-Driven Neural Response Generation for Knowledge-Grounded Dialog SystemsabstractOpen-domain dialog systems aim to generate relevant, informative and engaging responses.In this paper, we propose using a dialog policy to plan the content and style of target, opendomain responses in the form of an action plan, which includes knowledge sentences related to the dialog context, targeted dialog acts, topic information, etc.For training, the attributes within the action plan are obtained by automatically annotating the publicly released Topical-Chat dataset.We condition neural response generators on the action plan which is then realized as target utterances at the turn and sentence levels.We also investigate different dialog policy models to predict an action plan given the dialog context.Through automated and human evaluation, we measure the appropriateness of the generated responses and check if the generation models indeed learn to realize the given action plans.We demonstrate that a basic dialog policy that operates at the sentence level generates better responses in comparison to turn level generation as well as baseline models with no action plan.Additionally the basic dialog policy has the added benefit of controllability. Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Seokhwan Kim, Yang Liu 0004, Mihail Eric, Dilek Hakkani-Tür |
INLG | 1 |
| 2020 | Are Neural Open-Domain Dialog Systems Robust to Speech Recognition Errors in the Dialog History? An Empirical StudyabstractLarge end-to-end neural open-domain chatbots are becoming increasingly popular. However, research on building such chatbots has typically assumed that the user input is written in nature and it is not clear whether these chatbots would seamlessly integrate with automatic speech recognition (ASR) models to serve the speech modality. We aim to bring attention to this important question by empirically studying the effects of various types of synthetic and actual ASR hypotheses in the dialog history on TransferTransfo, a state-of-the-art Generative Pre-trained Transformer (GPT) based neural open-domain dialog system from the NeurIPS ConvAI2 challenge. We observe that TransferTransfo trained on written data is very sensitive to such hypotheses introduced to the dialog history during inference time. As a baseline mitigation strategy, we introduce synthetic ASR hypotheses to the dialog history during training and observe marginal improvements, demonstrating the need for further research into techniques to make end-to-end open-domain chatbots fully speech-robust. To the best of our knowledge, this is the first study to evaluate the effects of synthetic and actual ASR hypotheses on a state-of-the-art neural open-domain dialog system and we hope it promotes speech-robustness as an evaluation criterion in open-domain dialog. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Longshaokan Wang, Yang Liu 0004, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2020 | Beyond Domain APIs: Task-oriented Conversational Modeling with Unstructured Knowledge AccessabstractMost prior work on task-oriented dialogue systems are restricted to a limited coverage of domain APIs, while users oftentimes have domain related requests that are not covered by the APIs.In this paper, we propose to expand coverage of task-oriented dialogue systems by incorporating external unstructured knowledge sources.We define three sub-tasks: knowledge-seeking turn detection, knowledge selection, and knowledge-grounded response generation, which can be modeled individually or jointly.We introduce an augmented version of MultiWOZ 2.1, which includes new out-of-API-coverage turns and responses grounded on external knowledge sources.We present baselines for each sub-task using both conventional and neural approaches.Our experimental results demonstrate the need for further research in this direction to enable more informative conversational systems. Seokhwan Kim, Mihail Eric, Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Yang Liu 0004, Dilek Hakkani-Tür |
SIGdial | 4 |
| 2019 | Natural Language Generation at Scale: A Case Study for Open Domain Question AnsweringabstractAlessandra Cervone, Chandra Khatri, Rahul Goel, Behnam Hedayatnia, Anu Venkatesh, Dilek Hakkani-Tur, Raefer Gabriel. Proceedings of the 12th International Conference on Natural Language Generation. 2019. Alessandra Cervone, Chandra Khatri, Rahul Goel, Behnam Hedayatnia, Anu Venkatesh, Dilek Hakkani-Tür, Raefer Gabriel |
INLG | 4 |
| 2019 | Towards Coherent and Engaging Spoken Dialog Response Generation Using Automatic Conversation EvaluatorsabstractSanghyun Yi, Rahul Goel, Chandra Khatri, Alessandra Cervone, Tagyoung Chung, Behnam Hedayatnia, Anu Venkatesh, Raefer Gabriel, Dilek Hakkani-Tur. Proceedings of the 12th International Conference on Natural Language Generation. 2019. Sanghyun Yi, Rahul Goel, Chandra Khatri, Alessandra Cervone, Tagyoung Chung, Behnam Hedayatnia, Anu Venkatesh, Raefer Gabriel, Dilek Hakkani-Tür |
INLG | 6 |
| 2019 | Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations
Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Qinlang Chen 0001, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2018 | Contextual Language Model Adaptation for Conversational AgentsabstractStatistical language models (LM) play a key role in Automatic Speech Recognition (ASR) systems used by conversational agents. These ASR systems should provide a high accuracy under a variety of speaking styles, domains, vocabulary and argots. In this paper, we present a DNN-based method to adapt the LM to each user-agent interaction based on generalized contextual information, by predicting an optimal, context-dependent set of LM interpolation weights. We show that this framework for contextual adaptation provides accuracy improvements under different possible mixture LM partitions that are relevant for both (1) Goal-oriented conversational agents where it's natural to partition the data by the requested application and for (2) Non-goal oriented conversational agents where the data can be partitioned using topic labels that come from predictions of a topic classifier. We obtain a relative WER improvement of 3% with a 1-pass decoding strategy and 6% in a 2-pass decoding framework, over an unadapted model. We also show up to a 15% relative improvement in recognizing named entities which is of significant value for conversational ASR systems. Anirudh Raju, Behnam Hedayatnia, Linda Liu, Ankur Gandhe, Chandra Khatri, Angeliki Metallinou, Anu Venkatesh, Ariya Rastrow |
INTERSPEECH | 2 |
| 2018 | Contextual Topic Modeling For Dialog SystemsabstractAccurate prediction of conversation topics can be a valuable signal for creating coherent and engaging dialog systems. In this work, we focus on context-aware topic classification methods for identifying topics in free-form human-chatbot dialogs. We extend previous work on neural topic classification and unsupervised topic keyword detection by incorporating conversational context and dialog act features. On annotated data, we show that incorporating context and dialog acts leads to relative gains in topic classification accuracy by 35% and on unsupervised keyword detection recall by 11% for conversational interactions where topics frequently span multiple utterances. We show that topical metrics such as topical depth is highly correlated with dialog evaluation metrics such as coherence and engagement implying that conversational topic models can predict user satisfaction. Our work for detecting conversation topics and keywords can be used to guide chatbots towards coherent dialog. Chandra Khatri, Rahul Goel, Behnam Hedayatnia, Angeliki Metanillou, Anu Venkatesh, Raefer Gabriel, Arindam Mandal |
SLT | 3 |
| 2016 | Determining feature extractors for unsupervised learning on satellite imagesabstractAdvances in satellite imagery presents unprecedented opportunities for understanding natural and social phenomena at global and regional scales. Although the field of satellite remote sensing has evaluated imperative questions to human and environmental sustainability, scaling those techniques to very high spatial resolutions at regional scales remains a challenge. Satellite imagery is now more accessible with greater spatial, spectral and temporal resolution creating a data bottleneck in identifying the content of images. Because satellite images are unlabeled, unsupervised methods allow us to organize images into coherent groups or clusters. However, the performance of unsupervised methods, like all other machine learning methods, depends on features. Recent studies using features from pre-trained networks have shown promise for learning in new datasets. This suggests that features from pre-trained networks can be used for learning in temporally and spatially dynamic data sources such as satellite imagery. It is not clear, however, which features from which layer and network architecture should be used for learning new tasks. In this paper, we present an approach to evaluate the transferability of features from pre-trained Deep Convolutional Neural Networks for satellite imagery. We explore and evaluate different features and feature combinations extracted from various deep network architectures, and systematically evaluate over 2,000 network-layer combinations. In addition, we test the transferability of our engineered features and learned features from an unlabeled dataset to a different labeled dataset. Our feature engineering and learning are done on the unlabeled Draper Satellite Chronology dataset, and we test on the labeled UC Merced Land dataset to achieve near state-of-the-art classification results. These results suggest that even without any or minimal training, these networks can generalize well to other datasets. This method could be useful in the task of clustering unlabeled images and other unsupervised machine learning tasks. Behnam Hedayatnia, Mehrdad Yazdani, Mai H. Nguyen, Jessica Block, Ilkay Altintas |
IEEE BigData | 1 |