Shikib Mehri

dblp:212/0069 · DBLP profile ↗
← Back
20ranked-venue papers
12as first author
11since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 12 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Goal Alignment in LLM-Based User Simulators for Conversational AI
abstract
Abstract User simulators are essential to conversational AI, enabling scalable agent development and evaluation through simulated interactions. While current Large Language Models (LLMs) have advanced user simulation capabilities, we reveal that they struggle to consistently demonstrate goal-oriented behavior across multi-turn conversations, which is a critical limitation that compromises their reliability in downstream applications. We introduce User Goal State Tracking (UGST), a novel framework that tracks user goal progression throughout conversations. Leveraging UGST, we present a three-stage methodology for developing user simulators that can autonomously track goal progression and reason to generate goal-aligned responses. Moreover, we establish comprehensive evaluation metrics for measuring goal alignment in user simulators, and demonstrate that our approach yields substantial improvements across two benchmarks (MultiWOZ 2.4 and τ-Bench). Our contributions address a critical gap in conversational AI and establish UGST as an essential framework for developing goal-aligned user simulators. All code and data is released to facilitate future research 1.
Shuhaib Mehri, Xiaocheng Yang, Takyoung Kim, Gökhan Tür, Shikib Mehri, Dilek Hakkani-Tür
Trans. Assoc. Comput. Linguistics5
2025 Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
abstract
Abstract Large Language Models (LLMs) are often aligned using contrastive alignment objectives and preference pair datasets. The interaction between model, paired data, and objective makes alignment a complicated procedure, sometimes producing subpar results. We study this and find that (i) preference data gives a better learning signal when the underlying responses are contrastive, and (ii) alignment objectives lead to better performance when they specify more control over the model during training. Based on these insights, we introduce Contrastive Learning from AI Revisions (CLAIR), a data-creation method which leads to more contrastive preference pairs, and Anchored Preference Optimization (APO), a controllable and more stable alignment objective. We align Llama-3-8B-Instruct using various comparable datasets and alignment objectives and measure MixEval-Hard scores, which correlate highly with human judgments. The CLAIR preferences lead to the strongest performance out of all datasets, and APO consistently outperforms less controllable objectives. Our best model, trained on 32K CLAIR preferences with APO, improves Llama-3-8B-Instruct by 7.65%, closing the gap with GPT4-turbo by 45%. Our code and datasets are available.
Karel D'Oosterlinck, Winnie Xu, Chris Develder, Thomas Demeester, Amanpreet Singh, Christopher Potts, Douwe Kiela, Shikib Mehri
Trans. Assoc. Comput. Linguistics8
2024 Overview of the Ninth Dialog System Technology Challenge: DSTC9
abstract
This paper introduces the Ninth Dialog System Technology Challenge (DSTC-9). This edition of the DSTC focuses on applying end-to-end dialog technologies for four distinct tasks in dialog systems, namely, 1. Task-oriented dialog Modeling with Unstructured Knowledge Access, 2. Multi-domain task-oriented dialog, 3. Interactive evaluation of dialog and 4. Situated interactive multimodal dialog. This paper describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks.
R. Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro, Abhinav Rastogi, Yun-Nung Chen, Mihail Eric, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Yang Liu 0004, Chao-Wei Huang, Dilek Hakkani-Tür, Jinchao Li, Qi Zhu 0007, Lingxiao Luo, Lars Liden, Kaili Huang, Shahin Shayandeh, Runze Liang, Baolin Peng, Zheng Zhang 0020, Swadheen Shukla, Minlie Huang, Jianfeng Gao 0001, Shikib Mehri, Yulan Feng, Carla Gordon, Seyed Hossein Alavi, David R. Traum, Maxine Eskénazi, Ahmad Beirami, Eunjoon Cho, Paul A. Crook, Ankita De, Alborz Geramifard, Satwik Kottur, Seungwhan Moon, Shivani Poddar, Rajen Subba
IEEE ACM Trans. Audio Speech Lang. Process.24
2023 CESAR: Automatic Induction of Compositional Instructions for Multi-turn Dialogs
abstract
Instruction-based multitasking has played a critical role in the success of large language models (LLMs) in multi-turn dialog applications.While publicly-available LLMs have shown promising performance, when exposed to complex instructions with multiple constraints, they lag against state-of-the-art models like Chat-GPT.In this work, we hypothesize that the availability of large-scale complex demonstrations is crucial in bridging this gap.Focusing on dialog applications, we propose a novel framework, CESAR, that unifies a large number of dialog tasks in the same format and allows programmatic induction of complex instructions without any manual effort.We apply CESAR on InstructDial, a benchmark for instruction-based dialog tasks.We further enhance InstructDial with new datasets and tasks and utilize CESAR to induce complex tasks with compositional instructions.This results in a new benchmark called InstructDial++, which includes 63 datasets with 86 basic tasks and 68 composite tasks.Through rigorous experiments, we demonstrate the scalability of CESAR in providing rich instructions.Models trained on InstructDial++ can follow compositional prompts, such as prompts that ask for multiple stylistic constraints.
Taha Aksu, Devamanyu Hazarika, Shikib Mehri, Seokhwan Kim, Dilek Hakkani-Tür, Yang Liu 0004, Mahdi Namazifar
EMNLP3
2022 InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning
abstract
Instruction tuning is an emergent paradigm in NLP wherein natural language instructions are leveraged with language models to induce zeroshot performance on unseen tasks.Dialogue is an especially interesting area in which to explore instruction tuning because dialogue systems perform multiple tasks related to language (e.g., natural language understanding and generation, domain-specific interaction), yet instruction tuning has not been systematically explored for dialogue-related tasks.We introduce INSTRUCTDIAL, an instruction tuning framework for dialogue, which consists of a repository of 48 diverse dialogue tasks in a unified text-to-text format created from 59 openly available dialogue datasets.We explore crosstask generalization ability on models tuned on INSTRUCTDIAL across diverse dialogue tasks.Our analysis reveals that INSTRUCTDIAL enables good zero-shot performance on unseen datasets and tasks such as dialogue evaluation and intent detection, and even better performance in a few-shot setting.To ensure that models adhere to instructions, we introduce novel meta-tasks.We establish benchmark zero-shot and few-shot performance of models trained using the proposed framework on multiple dialogue tasks 1 .
Prakhar Gupta, Cathy Jiao, Shikib Mehri, Maxine Eskénazi, Jeffrey P. Bigham
EMNLP4
2022 Interactive Evaluation of Dialog Track at DSTC9
abstract
The ultimate goal of dialog research is to develop systems that can be effectively used in interactive settings by real users. To this end, we introduced the Interactive Evaluation of Dialog Track at the 9th Dialog System Technology Challenge. This track consisted of two sub-tasks. The first sub-task involved building knowledge-grounded response generation models. The second sub-task aimed to extend dialog models beyond static datasets by assessing them in an interactive setting with real users. Our track challenges participants to develop strong response generation models and explore strategies that extend them to back-and-forth interactions with real users. The progression from static corpora to interactive evaluation introduces unique challenges and facilitates a more thorough assessment of open-domain dialog systems. This paper provides an overview of the track, including the methodology and results. Furthermore, it provides insights into how to best evaluate open-domain dialog models.
Shikib Mehri, Yulan Feng, Carla Gordon, Seyed Hossein Alavi, David R. Traum, Maxine Eskénazi
LREC1
2022 The DialPort tools
abstract
The DialPort project (http://dialport. org/), funded by the
Jessica Huynh, Shikib Mehri, Cathy Jiao, Maxine Eskénazi
SIGDIAL2
2022 LAD: Language Models as Data for Zero-Shot Dialog
abstract
To facilitate zero-shot generalization in taskoriented dialog, this paper proposes Language Models as Data (LAD).LAD is a paradigm for creating diverse and accurate synthetic data which conveys the necessary structural constraints and can be used to train a downstream neural dialog model.LAD leverages GPT-3 to induce linguistic diversity.LAD achieves significant performance gains in zero-shot settings on intent prediction (+15%), slot filling (+31.4F-1) and next action prediction (+11 F-1).Furthermore, an interactive human evaluation shows that training with LAD is competitive with training on human dialogs.
Shikib Mehri, Yasemin Altun, Maxine Eskénazi
SIGDIAL1
2021 Example-Driven Intent Prediction with Observers
abstract
A key challenge of dialog systems research is to effectively and efficiently adapt to new domains.A scalable paradigm for adaptation necessitates the development of generalizable models that perform well in few-shot settings.In this paper, we focus on the intent classification problem which aims to identify user intents given utterances addressed to the dialog system.We propose two approaches for improving the generalizability of utterance classification models: (1) observers and (2) example-driven training.Prior work has shown that BERT-like models tend to attribute a significant amount of attention to the [CLS] token, which we hypothesize results in diluted representations.Observers are tokens that are not attended to, and are an alternative to the [CLS] token as a semantic representation of utterances.Example-driven training learns to classify utterances by comparing to examples, thereby using the underlying encoder as a sentence similarity model.These methods are complementary; improving the representation through observers allows the example-driven model to better measure sentence similarities.When combined, the proposed methods attain state-of-the-art results on three intent prediction datasets (BANKING77, CLINC150, HWU64) in both the full data and few-shot (10 examples per intent) settings.Furthermore, we demonstrate that the proposed approach can transfer to new intents and across datasets without any additional training.
Shikib Mehri, Mihail Eric
NAACL-HLT1
2021 GenSF: Simultaneous Adaptation of Generative Pre-trained Models and Slot Filling
abstract
In transfer learning, it is imperative to achieve strong alignment between a pre-trained model and a downstream task.Prior work has done this by proposing task-specific pre-training objectives, which sacrifices the inherent scalability of the transfer learning paradigm.We instead achieve strong alignment by simultaneously modifying both the pre-trained model and the formulation of the downstream task, which is more efficient and preserves the scalability of transfer learning.We present GENSF (Generative Slot Filling), which leverages a generative pre-trained open-domain dialog model for slot filling.GENSF (1) adapts the pre-trained model by incorporating inductive biases about the task and (2) adapts the downstream task by reformulating slot filling to better leverage the pre-trained model's capabilities.GENSF achieves state-of-the-art results on two slot filling datasets with strong gains in few-shot and zero-shot settings.We achieve a 9 F 1 score improvement in zeroshot slot filling.This highlights the value of strong alignment between the pre-trained model and the downstream task.
Shikib Mehri, Maxine Eskénazi
SIGDIAL1
2021 Schema-Guided Paradigm for Zero-Shot Dialog
abstract
Developing mechanisms that flexibly adapt dialog systems to unseen tasks and domains is a major challenge in dialog research.Neural models implicitly memorize task-specific dialog policies from the training data.We posit that this implicit memorization has precluded zero-shot transfer learning.To this end, we leverage the schema-guided paradigm, wherein the task-specific dialog policy is explicitly provided to the model.We introduce the Schema Attention Model (SAM) and improved schema representations for the STAR corpus.SAM obtains significant improvement in zero-shot settings, with a +22 F 1 score improvement over prior work.These results validate the feasibility of zero-shot generalizability in dialog.Ablation experiments are also presented to demonstrate the efficacy of SAM.
Shikib Mehri, Maxine Eskénazi
SIGDIAL1
2020 "None of the Above": Measure Uncertainty in Dialog Response Retrieval
abstract
This paper discusses the importance of uncovering uncertainty in end-to-end dialog tasks and presents our experimental results on uncertainty classification on the processed Ubuntu Dialog Corpus 1 .We show that instead of retraining models for this specific purpose, we can capture the original retrieval model's underlying confidence concerning the best prediction using trivial additional computation.
Yulan Feng, Shikib Mehri, Maxine Eskénazi
ACL2
2020 USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation
abstract
The lack of meaningful automatic evaluation metrics for dialog has impeded open-domain dialog research.Standard language generation metrics have been shown to be ineffective for evaluating dialog models.To this end, this paper presents USR, an UnSupervised and Reference-free evaluation metric for dialog.USR is a reference-free metric that trains unsupervised models to measure several desirable qualities of dialog.USR is shown to strongly correlate with human judgment on both Topical-Chat (turn-level: 0.42, systemlevel: 1.0) and PersonaChat (turn-level: 0.48 and system-level: 1.0).USR additionally produces interpretable measures for several desirable properties of dialog.
Shikib Mehri, Maxine Eskénazi
ACL1
2020 Unsupervised Evaluation of Interactive Dialog with DialoGPT
abstract
It is important to define meaningful and interpretable automatic evaluation metrics for open-domain dialog research.Standard language generation metrics have been shown to be ineffective for dialog.This paper introduces the FED metric (fine-grained evaluation of dialog), an automatic evaluation metric which uses DialoGPT, without any fine-tuning or supervision.It also introduces the FED dataset which is constructed by annotating a set of human-system and human-human conversations with eighteen fine-grained dialog qualities.The FED metric (1) does not rely on a ground-truth response, (2) does not require training data and (3) measures fine-grained dialog qualities at both the turn and whole dialog levels.FED attains moderate to strong correlation with human judgement at both levels.
Shikib Mehri, Maxine Eskénazi
SIGdial1
2019 Pretraining Methods for Dialog Context Representation Learning
abstract
This paper examines various unsupervised pretraining objectives for learning dialog context representations. Two novel methods of pretraining dialog context encoders are proposed, and a total of four methods are examined. Each pretraining objective is fine-tuned and evaluated on a set of downstream dialog tasks using the MultiWoz dataset and strong performance improvement is observed. Further evaluation shows that our pretraining objectives result in not only better performance, but also better convergence, models that are less data hungry and have better domain generalizability.
Shikib Mehri, Evgeniia Razumovskaia, Maxine Eskénazi
ACL (1)1
2019 Multi-Granularity Representations of Dialog
abstract
Shikib Mehri, Maxine Eskenazi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shikib Mehri, Maxine Eskénazi
EMNLP/IJCNLP (1)1
2019 Investigating Evaluation of Open-Domain Dialogue Systems With Human Generated Multiple References
abstract
The aim of this paper is to mitigate the shortcomings of automatic evaluation of open-domain dialog systems through multireference evaluation.Existing metrics have been shown to correlate poorly with human judgement, particularly in open-domain dialog.One alternative is to collect human annotations for evaluation, which can be expensive and time consuming.To demonstrate the effectiveness of multi-reference evaluation, we augment the test set of DailyDialog with multiple references.A series of experiments show that the use of multiple references results in improved correlation between several automatic metrics and human judgement for both the quality and the diversity of system output.
Prakhar Gupta, Shikib Mehri, Amy Pavel, Maxine Eskénazi, Jeffrey P. Bigham
SIGdial2
2019 Structured Fusion Networks for Dialog
abstract
Neural dialog models have exhibited strong performance, however their end-to-end nature lacks a representation of the explicit structure of dialog.This results in a loss of generalizability, controllability and a data-hungry nature.Conversely, more traditional dialog systems do have strong models of explicit structure.This paper introduces several approaches for explicitly incorporating structure into neural models of dialog.Structured Fusion Networks first learn neural dialog modules corresponding to the structured components of traditional dialog systems and then incorporate these modules in a higher-level generative model.Structured Fusion Networks obtain strong results on the MultiWOZ dataset, both with and without reinforcement learning.Structured Fusion Networks are shown to have several valuable properties, including better domain generalizability, improved performance in reduced data scenarios and robustness to divergence during reinforcement learning.
Shikib Mehri, Tejas Srinivasan, Maxine Eskénazi
SIGdial1
2018 Middle-Out Decoding
abstract
Despite being virtually ubiquitous, sequence-to-sequence models are challenged by their lack of diversity and inability to be externally controlled. In this paper, we speculate that a fundamental shortcoming of sequence generation models is that the decoding is done strictly from left-to-right, meaning that outputs values generated earlier have a profound effect on those generated later. To address this issue, we propose a novel middle-out decoder architecture that begins from an initial middle-word and simultaneously expands the sequence in both directions. To facilitate information flow and maintain consistent decoding, we introduce a dual self-attention mechanism that allows us to model complex dependencies between the outputs. We illustrate the performance of our model on the task of video captioning, as well as a synthetic sequence de-noising task. Our middle-out decoder achieves significant improvements on de-noising and competitive performance in the task of video captioning, while quantifiably improving the caption diversity. Furthermore, we perform a qualitative analysis that demonstrates our ability to effectively control the generation process of our decoder.
Shikib Mehri, Leonid Sigal
NeurIPS1
2017 Chat Disentanglement: Identifying Semantic Reply Relationships with Random Forests and Recurrent Neural Networks
abstract
Thread disentanglement is a precursor to any high-level analysis of multiparticipant chats. Existing research approaches the problem by calculating the likelihood of two messages belonging in the same thread. Our approach leverages a newly annotated dataset to identify reply relationships. Furthermore, we explore the usage of an RNN, along with large quantities of unlabeled data, to learn semantic relationships between messages. Our proposed pipeline, which utilizes a reply classifier and an RNN to generate a set of disentangled threads, is novel and performs well against previous work.
Shikib Mehri, Giuseppe Carenini
IJCNLP(1)1