VLDB 2026 Research / reviewers in the wild / expert
Dilek Hakkani-Tür
dblp:h/DilekZHakkaniTur · also Dilek Hakkani-Tur, Dilek Z. Hakkani-Tür, Dilek Zeynep Hakkani
· DBLP profile ↗
244ranked-venue papers
28as first author
62since 2021 · last 2026
0000-0001-5246-2117ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 165 · 17 first-author · 52 since 2021Graphics, computer vision, multimedia, augmented reality and games · 149 · 21 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Current Agents Fail to Leverage World Model as Tool for ForesightabstractCheng Qian, Emre Can Acikgoz, Bingxuan Li, Xiusi Chen, Yuji Zhang, Bingxiang He, Qinyu Luo, Gokhan Tur, Dilek Hakkani-Tür, Yunzhu Li, Heng Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Cheng Qian 0008, Emre Can Acikgoz, Xiusi Chen, Yuji Zhang 0002, Bingxiang He, Qinyu Luo, Gökhan Tür, Dilek Hakkani-Tür, Yunzhu Li, Heng Ji 0001 |
ACL (1) | 9 |
| 2026 | DialDefer: A Framework for Detecting and Mitigating LLM Dialogic DeferenceabstractParisa Rabbani, Priyam Sahoo, Ruben Mathew, Aishee Mondal, Harshita Ketharaman, Nimet Beyza Bozdag, Dilek Hakkani-Tür. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Parisa Rabbani, Priyam Sahoo, Ruben Mathew, Aishee Mondal, Harshita Ketharaman, Nimet Beyza Bozdag, Dilek Hakkani-Tür |
ACL (1) | 7 |
| 2026 | Do LLMs Encode Functional Importance of Reasoning Tokens ?abstractLarge language models solve complex tasks by generating long reasoning chains, achieving higher accuracy at the cost of increased computational cost and reduced ability to isolate functionally relevant reasoning.Prior work on compact reasoning shortens such chains through probabilistic sampling, heuristics, or supervision from frontier models, but offers limited insight into whether models internally encode token-level functional importance for answer generation.We address this gap diagnostically and propose greedy pruning, a likelihoodpreserving deletion procedure that iteratively removes reasoning tokens whose removal minimally degrades model likelihood under a specified objective, yielding length-controlled reasoning chains.We evaluate pruned reasoning in a distillation framework and show that students trained on pruned chains outperform a frontier-model-supervised compression baseline at matched reasoning lengths.Finally, our analysis reveals systematic pruning patterns and shows that attention scores can predict greedy pruning ranks, further suggesting that models encode a nontrivial functional importance structure over reasoning tokens. 1 Janvijay Singh, Dilek Hakkani-Tür |
ACL (1) | 2 |
| 2026 | Embodied Multi-Agent Coordination by Aligning World Models Through DialogueabstractEffective collaboration between embodied agents requires more than acting in a shared environment, it demands communication grounded in each agent’s evolving understanding of the world. When agents can only partially observe their surroundings, coordination without communication is provably hard, but communication can, in principle, bridge this gap by allowing agents to share observations and align their world models. In this work, we examine whether LLM-based embodied agents actually realize the ability to communicate. We extend PARTNR, a benchmark for collaborative household robotics, with a natural-language dialogue channel that enables two agents with partial observability to communicate during task execution. To evaluate whether dialogue leads to genuine world-model alignment rather than superficial coordination, we propose a framework for measuring world-model alignment defined over per-agent world graphs: observation convergence (do private world models align over time?), information novelty (do messages convey what the partner lacks?), and belief-sensitive messaging (do agents model what their partner knows?). Our experiments across three LLMs reveal that dialogue reduces action conflicts 40–83 percentage points but degrades task success relative to silent coordination. Using our metrics, we characterize the gap between superficial coordination and genuine world-model alignment, and identify where current models fall on this spectrum. Vardhan Dongre, Dilek Hakkani-Tür |
SIGDIAL | 2 |
| 2026 | Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent SystemsabstractLarge language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model’s opinion. While prior work has mostly studied this in single-agent settings, it remains underexplored in collaborative multi-agent systems. We ask whether awareness of other agents’ sycophancy levels influences discussion outcomes. To investigate this, we run controlled experiments with six open-source LLMs, providing agents with peer sycophancy rankings that estimate each peer’s tendency toward sycophancy. These rankings are based on scores calculated using various static (pre-discussion) and dynamic (online) strategies. We find that providing sycophancy priors reduces the influence of sycophancy-prone peers, mitigates error-cascades, and improves final discussion accuracy by an absolute 10.5%. Thus, this is a lightweight and efficient way to reduce model sycophancy during discussions and subsequently improve downstream accuracy. Vira Kasprova, Amruta Parulekar, Abdulrahman AlRabah, Krishna Agaram, Ritwik Garg, Sagar Jha, Nimet Beyza Bozdag, Dilek Hakkani-Tür |
SIGDIAL | 8 |
| 2026 | GBC: Gradient-Based Connections for Optimizing Multi-Agent SystemsabstractMulti-agent systems (MAS) built on large language models (LLMs) provide a promising framework for solving complex tasks through role specialization and structured interaction. However, their performance is often limited by miscoordination and, more fundamentally, the lack of fine-grained credit assignment across agents. Existing approaches typically rely on coarse-grained feedback, making it difficult to identify which agents or interaction steps are responsible for errors. We propose Gradient-Based Connections (GBC), an approach for fine-grained attribution and optimization of multi-agent systems. GBC models a MAS as a computational graph and introduces gradient-based connection weights to quantify the influence of each agent’s output on downstream agents at the token level. By constructing an attribution graph and propagating task-specific loss signals backward, our method enables precise identification of error sources and targeted prompt optimization. We further develop AgentChord, an efficient implementation that leverages prefix-based gradient computation. Experiments on MultiWOZ and τ-bench show that GBC improves multi-agent performance and outperforms strong single-agent and multi-agent baselines, and higher attribution quality is associated with greater optimization effectiveness. Code is available at: https://github.com/yxc-cyber/AgentChord. Xiaocheng Yang, Abdulrahman Alrabah, Dilek Hakkani-Tür, Gökhan Tür |
SIGDIAL | 3 |
| 2026 | Goal Alignment in LLM-Based User Simulators for Conversational AIabstractAbstract User simulators are essential to conversational AI, enabling scalable agent development and evaluation through simulated interactions. While current Large Language Models (LLMs) have advanced user simulation capabilities, we reveal that they struggle to consistently demonstrate goal-oriented behavior across multi-turn conversations, which is a critical limitation that compromises their reliability in downstream applications. We introduce User Goal State Tracking (UGST), a novel framework that tracks user goal progression throughout conversations. Leveraging UGST, we present a three-stage methodology for developing user simulators that can autonomously track goal progression and reason to generate goal-aligned responses. Moreover, we establish comprehensive evaluation metrics for measuring goal alignment in user simulators, and demonstrate that our approach yields substantial improvements across two benchmarks (MultiWOZ 2.4 and τ-Bench). Our contributions address a critical gap in conversational AI and establish UGST as an essential framework for developing goal-aligned user simulators. All code and data is released to facilitate future research 1. Shuhaib Mehri, Xiaocheng Yang, Takyoung Kim, Gökhan Tür, Shikib Mehri, Dilek Hakkani-Tür |
Trans. Assoc. Comput. Linguistics | 6 |
| 2025 | Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language ModelabstractEmre Can Acikgoz, Jeremiah Greer, Akul Datta, Ze Yang, William Zeng, Oussama Elachqar, Emmanouil Koukoumidis, Dilek Hakkani-Tür, Gokhan Tur. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Emre Can Acikgoz, Jeremiah Greer, Akul Datta, William Zeng, Oussama Elachqar, Emmanouil Koukoumidis, Dilek Hakkani-Tür, Gökhan Tür |
ACL (1) | 8 |
| 2025 | Know Your Mistakes: Towards Preventing Overreliance on Task-Oriented Conversational AI Through Accountability ModelingabstractRecent LLMs have enabled significant advancements for conversational agents.However, they are also well known to hallucinate, producing responses that seem plausible but are factually incorrect.On the other hand, users tend to over-rely on LLM-based AI agents, accepting AI's suggestion even when it is wrong.Adding positive friction, such as explanations or getting user confirmations, has been proposed as a mitigation in AI-supported decision-making systems.In this paper, we propose an accountability model for LLM-based task-oriented dialogue agents to address user overreliance via friction turns in cases of model uncertainty and errors associated with dialogue state tracking (DST).The accountability model is an augmented LLM with an additional accountability head that functions as a binary classifier to predict the relevant slots of the dialogue state mentioned in the conversation.We perform our experiments with multiple backbone LLMs on two established benchmarks (MultiWOZ and Snips).Our empirical findings demonstrate that the proposed approach not only enables reliable estimation of AI agent errors but also guides the decoder in generating more accurate actions.We observe around 3% absolute improvement in joint goal accuracy (JGA) of DST output by incorporating accountability heads into modern LLMs.Self-correcting the detected errors further increases the JGA from 67.13 to 70.51, achieving state-of-the-art DST performance.Finally, we show that error correction through user confirmations (friction turn) achieves a similar performance gain, highlighting its potential to reduce user overreliance.1 Suvodip Dey, Yi-Jyun Sun, Gökhan Tür, Dilek Hakkani-Tür |
ACL (1) | 4 |
| 2025 | Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic InteractionsabstractAs chatbots continue to evolve toward humanlike, real-world, interactions, multimodality remains an active area of research and exploration.So far, efforts to integrate multimodality into chatbots have primarily focused on image-centric tasks, such as visual dialogue and image-based instructions, placing emphasis on the "eyes" of human perception while neglecting the "ears", namely auditory aspects.Moreover, these studies often center around static interactions that focus on discussing the modality rather than naturally incorporating it into the conversation, which limits the richness of simultaneous, dynamic engagement.Furthermore, while multimodality has been explored in multi-party and multi-session conversations, task-specific constraints have hindered its seamless integration into dynamic, natural conversations.To address these challenges, this study aims to equip chatbots with "eyes and ears" capable of more immersive interactions with humans.As part of this effort, we introduce a new multimodal conversation dataset, Multimodal Multi-Session Multi-Party Conversation (M 3 C), and propose a novel multimodal conversation model featuring multimodal memory retrieval.Our model, trained on the M 3 C, demonstrates the ability to seamlessly engage in long-term conversations with multiple speakers in complex, real-world-like settings, effectively processing visual and auditory inputs to understand and respond appropriately.Human evaluations highlight the model's strong performance in maintaining coherent and dynamic interactions, demonstrating its potential for advanced multimodal conversational agents. 1 Jihyoung Jang, Minwook Bae, Dilek Hakkani-Tür, Hyounghun Kim |
ACL (1) | 4 |
| 2025 | Beliefs in Motion: Simulating Opinion Dynamics via LLM-Powered Community Reactions
Dachun Sun, Dilek Hakkani-Tür, Tarek F. Abdelzaher |
ASONAM (1) | 3 |
| 2025 | Aligning LLMs with Individual Preferences via InteractionabstractAs large language models (LLMs) demonstrate increasingly advanced capabilities, aligning their behaviors with human values and preferences becomes crucial for their wide adoption. While previous research focuses on general alignment to principles such as helpfulness, harmlessness, and honesty, the need to account for individual and diverse preferences has been largely overlooked, potentially undermining customized human experiences. To address this gap, we train LLMs that can “interact to align”, essentially cultivating the meta-skill of LLMs to implicitly infer the unspoken personalized preferences of the current user through multi-turn conversations, and then dynamically align their following behaviors and responses to these inferred preferences. Our approach involves establishing a diverse pool of 3,310 distinct user personas by initially creating seed examples, which are then expanded through iterative self-generation and filtering. Guided by distinct user personas, we leverage multi-LLM collaboration to develop a multi-turn preference dataset containing 3K+ multi-turn conversations in tree structures. Finally, we apply supervised fine-tuning and reinforcement learning to enhance LLMs using this dataset. For evaluation, we establish the ALOE (ALign with custOmized prEferences) benchmark, consisting of 100 carefully selected examples and well-designed metrics to measure the customized alignment performance during conversations. Experimental results demonstrate the effectiveness of our method in enabling dynamic, personalized alignment via interaction. The code and dataset will be made public. Shujin Wu, Yi R. Fung 0001, Cheng Qian 0008, Dilek Hakkani-Tür, Heng Ji 0001 |
COLING | 5 |
| 2025 | Spark: A System for Scientifically Creative Idea Generation
Aishik Sanyal, Samuel Schapiro, Sumuk Shashidhar, Royce Moon, Lav R. Varshney, Dilek Hakkani-Tür |
ICCC | 6 |
| 2025 | Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMsabstractChain-of-Thought (CoT) prompting enhances mathematical reasoning in large language models (LLMs) by enabling detailed step-by-step solutions. However, due to the verbosity of LLMs, the resulting reasoning chains can be long, making it harder to verify the reasoning steps and trace issues resulting from dependencies between the steps that may be farther away in the sequence of steps. Importantly, mathematical reasoning allows each step to be derived from a small set of premises, which are a subset of the preceding steps in the reasoning chain. In this paper, we present a framework that identifies the premises for each step, to improve the evaluation of reasoning. We restructure conventional linear reasoning chains into Premise Augmented Reasoning Chains (PARC) by introducing premise links, resulting in a directed acyclic graph where the nodes are the steps and the edges are the premise links. Through experiments with a PARC-based dataset that we built, namely (Premises and ERrors identification in LLMs), we demonstrate that LLMs can reliably identify premises within complex reasoning chains. In particular, even open-source LLMs achieve 90% recall in premise identification. We also show that PARC helps to identify errors in reasoning chains more reliably. The accuracy of error identification improves by 6% to 16% absolute when step-by-step verification is carried out in PARC under the premises. Our findings highlight the utility of premise-centric representations in addressing complex problem-solving tasks and open new avenues for improving the reliability of LLM-based reasoning evaluations. Sagnik Mukherjee, Abhinav Chinta, Takyoung Kim, Tarun Anoop Sharma, Dilek Hakkani-Tür |
ICML | 5 |
| 2025 | Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning DataabstractInfluence functions provide crucial insights into model training, but existing methods suffer from large computational costs and limited generalization. Particularly, recent works have proposed various metrics and algorithms to calculate the influence of data using language models, which do not scale well with large models and datasets. This is because of the expensive forward and backward passes required for computation, substantial memory requirements to store large models, and poor generalization of influence estimates to new data. In this paper, we explore the use of small neural networks -- which we refer to as the InfluenceNetwork -- to estimate influence values, achieving up to 99% cost reduction. Our evaluation demonstrates that influence values can be estimated with models just 0.0007% the size of full language models (we average across 1.5B-22B versions). We apply our algorithm of estimating influence values (called NN-CIFT: Neural Networks for effiCient Instruction Fine-Tuning) to the downstream task of subset selection for general instruction fine-tuning. In our study, we include four state-of-the-art influence functions and show no compromise in performance, despite large speedups, between NN-CIFT and the original influence functions. We provide an in-depth hyperparameter analyses of NN-CIFT. The code for our method can be found here: https://github.com/agarwalishika/NN-CIFT/tree/main. Ishika Agarwal, Dilek Hakkani-Tür |
NeurIPS | 2 |
| 2025 | MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided ConversationsabstractWe introduce MIRAGE, a new benchmark for multimodal expert-level reasoning and decision-making in consultative interaction settings. Designed for the domain of agriculture, MIRAGE captures the full complexity of expert consultations by combining natural user queries, expert-authored responses, and image-based context, offering a high-fidelity benchmark for evaluating models on grounded reasoning, clarification strategies, and long-form generation in a real-world, knowledge-intensive domain. Grounded in over 35,000 real user-expert interactions, and curated through a carefully designed multi-step pipeline, MIRAGE spans diverse crop health, pest diagnosis, and crop management scenarios. The benchmark includes more than 7,000 unique biological entities, covering plant species, pests, and diseases, making it one of the most taxonomically diverse benchmarks available for vision-language models in real-world expert-guided domains. Unlike existing benchmarks that rely on well-specified user inputs, MIRAGE features underspecified, context-rich scenarios, requiring models to infer latent knowledge gaps and either proactively guide the interaction or respond. Our benchmark comprises two core components. The Single-turn Challenge to reason over a single user turn and image set, identify relevant entities, infer causal explanations, and generate actionable recommendations; and a Multi-Turn challenge for dialogue state tracking, goal-driven generation, and expert-level conversational decision-making. We evaluate more than 20 closed and open-source frontier vision-language models (VLMs), using three reasoning language models as evaluators, highlighting the significant challenges posed by MIRAGE in both single-turn and multi-turn interaction settings. Even the advanced GPT4.1 and GPT4o models achieve 44.6% and 40.9% accuracy, respectively, indicating significant room for improvement. Vardhan Dongre, Chi Gui, Shubham Garg, Hooshang Nayyeri, Gökhan Tür, Dilek Hakkani-Tür, Vikram S. Adve |
NeurIPS | 6 |
| 2025 | Reinforcement Learning Finetunes Small Subnetworks in Large Language ModelsabstractReinforcement learning (RL) yields substantial improvements in large language models’ (LLMs) downstream task performance and alignment with human values. Surprisingly, such large gains result from updating only a small subnetwork comprising just 5%-30% of the parameters, with the rest effectively unchanged. We refer to this phenomenon as parameter update sparsity induced by RL. It is observed across all 7 widely-used RL algorithms (e.g., PPO, GRPO, DPO) and all 10 LLMs from different families in our experiments.
This sparsity is intrinsic and occurs without any explicit sparsity-promoting regularizations or architectural constraints. Finetuning the subnetwork alone recovers the test accuracy, and, remarkably, produces a model nearly identical to the one obtained via full finetuning.
The subnetworks from different random seeds, training data, and even RL algorithms show substantially greater overlap than expected by chance. Our analysis suggests that this sparsity is not due to updating only a subset of layers; instead, nearly all parameter matrices receive similarly sparse updates. Moreover, the updates to almost all parameter matrices are nearly full-rank,
suggesting RL updates a small subset of parameters that nevertheless span almost the full subspaces that the parameter matrices can represent. We conjecture that the this update sparsity can be primarily attributed to training on data that is near the policy distribution;
techniques that encourage the policy to remain close to the pretrained model, such as the KL regularization and gradient clipping, have limited impact. Sagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tür, Hao Peng 0009 |
NeurIPS | 3 |
| 2025 | ToolRL: Reward is All Tool Learning NeedsabstractCurrent Large Language Models (LLMs) often undergo supervised fine-tuning (SFT) to acquire tool use capabilities. However, SFT struggles to generalize to unfamiliar or complex tool use scenarios. Recent advancements in reinforcement learning (RL), particularly with R1-like models, have demonstrated promising reasoning and generalization abilities. Yet, reward design for tool use presents unique challenges: multiple tools may be invoked with diverse parameters, and coarse-grained reward signals, such as answer matching, fail to offer the finegrained feedback required for effective learning.
In this work, we present the first comprehensive study on reward design for tool selection and application tasks within the RL paradigm. We systematically explore a wide range of reward strategies, analyzing their types, scales, granularity, and temporal dynamics. Building on these insights, we propose a principled reward design tailored for tool use tasks and apply it to train LLMs using RL methods.
Empirical evaluations across diverse benchmarks demonstrate that our approach yields robust, scalable, and stable training, achieving a 17\% improvement over base models and a 15\% gain over SFT models. These results highlight the critical role of thoughtful reward design in enhancing the tool use capabilities and generalization performance of LLMs. All the codes are released to facilitate future research. Cheng Qian 0008, Emre Can Acikgoz, Hongru Wang 0011, Xiusi Chen, Dilek Hakkani-Tür, Gökhan Tür, Heng Ji 0001 |
NeurIPS | 6 |
| 2025 | TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level ComparisonsabstractTask-oriented dialogue (TOD) systems are experiencing a revolution driven by Large Language Models (LLMs), yet the evaluation methodologies for these systems remain insufficient for their growing sophistication. While traditional automatic metrics effectively assessed earlier modular systems, they focus solely on the dialogue level and cannot detect critical intermediate errors that can arise during user-agent interactions. In this paper, we introduce TD-EVAL (Turn and Dialogue-level Evaluation), a two-step evaluation framework that unifies fine-grained turn-level analysis with holistic dialogue-level comparisons. At turn-level, we assess each response along three TOD-specific dimensions: conversation cohesion, backend knowledge consistency, and policy compliance. Meanwhile, we design TOD Agent Arena that uses pairwise comparisons to provide a measure of dialogue-level quality. Through experiments on MultiWOZ 2.4 and Tau-Bench, we demonstrate that TD-EVAL effectively identifies the conversational errors that conventional metrics miss. Furthermore, TD-EVAL exhibits better alignment with human judgments than traditional and LLM-based metrics. These findings demonstrate that TD-EVAL introduces a new paradigm for TOD system evaluation, efficiently assessing both turn and system levels with an easily reproducible framework for future research. Emre Can Acikgoz, Carl Guo, Suvodip Dey, Akul Datta, Takyoung Kim, Gökhan Tür, Dilek Hakkani-Tür |
SIGDIAL | 7 |
| 2025 | DocCHA: Towards LLM-Augmented Interactive Online diagnosis SystemabstractDespite the impressive capabilities of Large Language Models (LLMs), existing Conversational Health Agents (CHAs) remain static and brittle, incapable of adaptive multi-turn reasoning, symptom clarification, or transparent decision-making. This hinders their real-world applicability in clinical diagnosis, where iterative and structured dialogue is essential. We propose DocCHA, a confidence-aware, modular framework that emulates clinical reasoning by decomposing the diagnostic process into three stages: (1) symptom elicitation, (2) history acquisition, and (3) causal graph construction. Each module uses interpretable confidence scores to guide adaptive questioning, prioritize informative clarifications, and refine weak reasoning links. Evaluated on two real-world Chinese consultation datasets (IMCS21, DX), DocCHA consistently outperforms strong prompting-based LLM baselines (GPT-3.5, GPT-4o, LLaMA-3), achieving up to 5.18% higher diagnostic accuracy and over 30% improvement in symptom recall, with only modest increase in dialogue turns. These results demonstrate DocCHA’s effectiveness in enabling structured, transparent, and efficient diagnostic conversations—paving the way for trustworthy LLM-powered clinical assistants in multilingual and resource-constrained settings. Dachun Sun, Yi R. Fung 0001, Dilek Hakkani-Tür, Tarek F. Abdelzaher |
SIGDIAL | 4 |
| 2024 | Unsupervised Human Preference LearningabstractLarge language models demonstrate impressive reasoning abilities but struggle to provide personalized content due to their lack of individual user preference information.Existing methods, such as in-context learning and parameter-efficient fine-tuning, fall short in capturing the complexity of human preferences, especially given the small, personal datasets individuals possess.In this paper, we propose a novel approach utilizing small parameter models as preference agents to generate natural language rules that guide a larger, pre-trained model, enabling efficient personalization.Our method involves a small, local "steering wheel" model that directs the outputs of a much larger foundation model, producing content tailored to an individual's preferences while leveraging the extensive knowledge and capabilities of the large model.Importantly, this personalization is achieved without the need to fine-tune the large model.Experimental results demonstrate that our technique significantly outperforms baseline personalization methods.By allowing foundation models to adapt to individual preferences in a data and compute-efficient manner, our approach paves the way for highly personalized language model applications. Sumuk Shashidhar, Abhinav Chinta, Vaibhav Sahai, Dilek Hakkani-Tür |
EMNLP | 4 |
| 2024 | Dialog Flow Induction for Constrainable LLM-Based ChatbotsabstractStuti Agrawal, Pranav Pillai, Nishi Uppuluri, Revanth Gangi Reddy, Sha Li, Gokhan Tur, Dilek Hakkani-Tur, Heng Ji. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024. Stuti Agrawal, Pranav Pillai, Nishi Uppuluri, Revanth Gangi Reddy, Gökhan Tür, Dilek Hakkani-Tür, Heng Ji 0001 |
SIGDIAL | 7 |
| 2024 | Large Language Models as User-Agents For Evaluating Task-Oriented-Dialogue SystemsabstractTraditionally, offline datasets have been used to evaluate task-oriented dialogue (TOD) models. These datasets lack context awareness, making them suboptimal benchmarks for conversational systems. In contrast, user-agents, which are context-aware, can simulate the variability and unpredictability of human conversations, making them better alternatives as evaluators. Prior research has utilized large language models (LLMs) to develop user-agents. Our work builds upon this by using LLMs to create user-agents for the evaluation of TOD systems. This involves prompting an LLM, using in-context examples as guidance, and tracking the user-goal state. Our evaluation of diversity and task completion metrics for the user-agents shows improved performance with the use of better prompts. Additionally, we propose methodologies for the automatic evaluation of TOD models within this dynamic framework. We make our code publicly available11https://github.com/TaahaKazi/user-agent Taaha Kazi, Ruiliang Lyu, Sizhe Zhou, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 4 |
| 2024 | Confidence Estimation For LLM-Based Dialogue State TrackingabstractEstimation of a model’s confidence on its outputs is critical for Conversational AI systems based on large language models (LLMs), especially for reducing hallucination and preventing over-reliance. In this work, we provide an exhaustive exploration of methods, including approaches proposed for open- and closed-weight LLMs, aimed at quantifying and leveraging model uncertainty to improve the reliability of LLM-generated responses, specifically focusing on dialogue state tracking (DST) in task-oriented dialogue systems (TODS). Regardless of the model type, well-calibrated confidence scores are essential to handle uncertainties, thereby improving model performance. We evaluate four methods for estimating confidence scores based on softmax, raw token scores, verbalized confidences, and a combination of these methods, using the area under the curve (AUC) metric to assess calibration, with higher AUC indicating better calibration. We also enhance these with a self-probing mechanism, proposed for closed models. Furthermore, we assess these methods using an open-weight model fine-tuned for the task of DST, achieving superior joint goal accuracy (JGA). Our findings also suggest that fine-tuning open-weight LLMs can result in enhanced AUC performance, indicating better confidence score calibration. Yi-Jyun Sun, Suvodip Dey, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 3 |
| 2024 | Overview of the Ninth Dialog System Technology Challenge: DSTC9abstractThis paper introduces the Ninth Dialog System Technology Challenge (DSTC-9). This edition of the DSTC focuses on applying end-to-end dialog technologies for four distinct tasks in dialog systems, namely, 1. Task-oriented dialog Modeling with Unstructured Knowledge Access, 2. Multi-domain task-oriented dialog, 3. Interactive evaluation of dialog and 4. Situated interactive multimodal dialog. This paper describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. R. Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro, Abhinav Rastogi, Yun-Nung Chen, Mihail Eric, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Yang Liu 0004, Chao-Wei Huang, Dilek Hakkani-Tür, Jinchao Li, Qi Zhu 0007, Lingxiao Luo, Lars Liden, Kaili Huang, Shahin Shayandeh, Runze Liang, Baolin Peng, Zheng Zhang 0020, Swadheen Shukla, Minlie Huang, Jianfeng Gao 0001, Shikib Mehri, Yulan Feng, Carla Gordon, Seyed Hossein Alavi, David R. Traum, Maxine Eskénazi, Ahmad Beirami, Eunjoon Cho, Paul A. Crook, Ankita De, Alborz Geramifard, Satwik Kottur, Seungwhan Moon, Shivani Poddar, Rajen Subba |
IEEE ACM Trans. Audio Speech Lang. Process. | 11 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 18 |
| 2023 | Towards Credible Human Evaluation of Open-Domain Dialog Systems Using Interactive SetupabstractEvaluating open-domain conversation models has been an open challenge due to the open-ended nature of conversations. In addition to static evaluations, recent work has started to explore a variety of per-turn and per-dialog interactive evaluation mechanisms and provide advice on the best setup. In this work, we adopt the interactive evaluation framework and further apply to multiple models with a focus on per-turn evaluation techniques. Apart from the widely used setting where participants select the best response among different candidates at each turn, one more novel per-turn evaluation setting is adopted, where participants can select all appropriate responses with different fallback strategies to continue the conversation when no response is selected. We evaluate these settings based on sensitivity and consistency using four GPT2-based models that differ in model sizes or fine-tuning data. To better generalize to any model groups with no prior assumptions on their rankings and control evaluation costs for all setups, we also propose a methodology to estimate the required sample size given a minimum performance gap of interest before running most experiments. Our comprehensive human evaluation results shed light on how to conduct credible human evaluations of open domain dialog systems using the interactive setup, and suggest additional future directions. Sijia Liu 0007, Patrick Lange, Behnam Hedayatnia, Alexandros Papangelis, Di Jin 0005, Andrew Wirth, Yang Liu 0004, Dilek Hakkani-Tür |
AAAI | 8 |
| 2023 | KILM: Knowledge Injection into Encoder-Decoder Language ModelsabstractYan Xu, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yan Xu 0012, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar, Yang Liu 0004, Dilek Hakkani-Tür |
ACL (1) | 6 |
| 2023 | Selective In-Context Data Augmentation for Intent Detection using Pointwise V-InformationabstractYen-Ting Lin, Alexandros Papangelis, Seokhwan Kim, Sungjin Lee, Devamanyu Hazarika, Mahdi Namazifar, Di Jin, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Alexandros Papangelis, Seokhwan Kim, Devamanyu Hazarika, Mahdi Namazifar, Di Jin 0005, Yang Liu 0004, Dilek Hakkani-Tür |
EACL | 9 |
| 2023 | CESAR: Automatic Induction of Compositional Instructions for Multi-turn DialogsabstractInstruction-based multitasking has played a critical role in the success of large language models (LLMs) in multi-turn dialog applications.While publicly-available LLMs have shown promising performance, when exposed to complex instructions with multiple constraints, they lag against state-of-the-art models like Chat-GPT.In this work, we hypothesize that the availability of large-scale complex demonstrations is crucial in bridging this gap.Focusing on dialog applications, we propose a novel framework, CESAR, that unifies a large number of dialog tasks in the same format and allows programmatic induction of complex instructions without any manual effort.We apply CESAR on InstructDial, a benchmark for instruction-based dialog tasks.We further enhance InstructDial with new datasets and tasks and utilize CESAR to induce complex tasks with compositional instructions.This results in a new benchmark called InstructDial++, which includes 63 datasets with 86 basic tasks and 68 composite tasks.Through rigorous experiments, we demonstrate the scalability of CESAR in providing rich instructions.Models trained on InstructDial++ can follow compositional prompts, such as prompts that ask for multiple stylistic constraints. Taha Aksu, Devamanyu Hazarika, Shikib Mehri, Seokhwan Kim, Dilek Hakkani-Tür, Yang Liu 0004, Mahdi Namazifar |
EMNLP | 5 |
| 2023 | Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied DialogueabstractEmbodied task completion is a challenge where an agent in a simulated environment must predict environment actions to complete tasks based on natural language instructions and egocentric visual observations.We propose a variant of this problem where the agent predicts actions at a higher level of abstraction called a plan, which helps make agent actions more interpretable and can be obtained from the appropriate prompting of large language models.We show that multimodal transformer models can outperform language-only models for this problem but fall significantly short of oracle plans.Since collecting human-human dialogues for embodied environments is expensive and time-consuming, we propose a method to synthetically generate such dialogues, which we then use as training data for plan prediction.We demonstrate that multimodal transformer models can attain strong zero-shot performance from our synthetic data, outperforming language-only models trained on humanhuman data. Aishwarya Padmakumar, Mert Inan, Spandana Gella, Patrick Lange, Dilek Hakkani-Tür |
EMNLP | 5 |
| 2023 | Identifying Entrainment in Task-Oriented ConversationsabstractHuman interlocutors adapt their behavior to each other in a conversation through entrainment. While entrainment has been found in long chit-chat conversations, much less research has been conducted on task-oriented dialogs. In this paper, we investigate short task-oriented Wizard-of-Oz conversations for acoustic-prosodic and lexical entrainment. We conduct significance tests that reveal changes in speech pitch and frequent words as important indicators of entrainment. Our findings will guide user-entraining dialog systems to improve the quality of conversations. Run Chen, Seokhwan Kim, Alexandros Papangelis, Julia Hirschberg, Yang Liu 0004, Dilek Hakkani-Tür |
ICASSP | 6 |
| 2023 | Role of Bias Terms in Dot-Product AttentionabstractDot-product attention is a core module in the present generation of neural network models, particularly transformers, and is being leveraged across numerous areas such as natural language processing and computer vision. This attention module is comprised of three linear transformations, namely query, key, and value linear transformations, each of which has a bias term. In this work, we study the role of these bias terms, and mathematically show that the bias term of the key linear transformation is redundant and could be omitted without any impact on the attention module. Moreover, we argue that the bias term of the value linear transformation has a more prominent role than that of the bias term of the query linear transformation. We empirically verify these findings through multiple experiments on language modeling, natural language understanding, and natural language generation tasks. Mahdi Namazifar, Devamanyu Hazarika, Dilek Hakkani-Tür |
ICASSP | 3 |
| 2023 | Conversational Text-to-SQL: An Odyssey into State-of-the-Art and Challenges AheadabstractConversational, multi-turn, text-to-SQL (CoSQL) tasks map natural language utterances in a dialogue to SQL queries. State-of-the-art (SOTA) systems use large, pre-trained and finetuned language models, such as the T5-family, in conjunction with constrained decoding. With multi-tasking (MT) over coherent tasks with discrete prompts during training, we improve over specialized text-to-SQL T5-family models. Based on Oracle analyses over n-best hypotheses, we apply a query plan model and a schema linking algorithm as rerankers. Combining MT and reranking, our results using T53B show absolute accuracy improvements of 1.0% in exact match and 3.4% in execution match over a SOTA baseline on CoSQL. While these gains consistently manifest at turn level, context dependent turns are considerably harder. We conduct studies to tease apart errors attributable to domain and compositional generalization, with the latter remaining a challenge for multi-turn conversations, especially in generating SQL with unseen parse trees. Sree Hari Krishnan Parthasarathi, Dilek Hakkani-Tür |
ICASSP | 3 |
| 2023 | Alexa Arena: A User-Centric Interactive Platform for Embodied AIabstractWe introduce Alexa Arena, a user-centric simulation platform to facilitate research in building assistive conversational embodied agents. Alexa Arena features multi-room layouts and an abundance of interactable objects. With user-friendly graphics and control mechanisms, the platform supports the development of gamified robotic tasks readily accessible to general human users, allowing high-efficiency data collection and EAI system evaluation. Along with the platform, we introduce a dialog-enabled task completion benchmark with online human evaluations. Qiaozi Gao, Govind Thattai, Suhaila M. Shakiah, Xiaofeng Gao 0002, Shreyas Pansare, Vasu Sharma, Gaurav S. Sukhatme, Hangjie Shi, Bofei Yang, Lucy Hu, Karthika Arumugam, Shui Hu, Matthew Wen, Dinakar Guthy, Shunan Chung, Rohan Khanna, Osman Ipek, Leslie Ball, Kate Bland, Heather Rocker, Michael Johnston, Reza Ghanadan, Dilek Hakkani-Tür, Premkumar Natarajan |
NeurIPS | 24 |
| 2023 | MERCY: Multiple Response Ranking Concurrently in Realistic Open-Domain Conversational SystemsabstractAutomatic Evaluation (AE) and Response Selection (RS) models assign quality scores to various candidate responses and rank them in conversational setups.Prior response ranking research compares various models' performance on synthetically generated test sets.In this work, we investigate the performance of model-based reference-free AE and RS models on our constructed response ranking datasets that mirror real-case scenarios of ranking candidates during inference time.Metrics' unsatisfying performance can be interpreted as their low generalizability over more pragmatic conversational domains such as human-chatbot dialogs.To alleviate this issue we propose a novel RS model called MERCY that simulates human behavior in selecting the best candidate by taking into account distinct candidates concurrently and learns to rank them.In addition, MERCY leverages natural language feedback as another component to help the ranking task by explaining why each candidate response is relevant/irrelevant to the dialog context.These feedbacks are generated by prompting large language models in a few-shot setup.Our experiments show the better performance of MERCY over baselines for the response ranking task in our curated realistic datasets. Sarik Ghazarian, Behnam Hedayatnia, Di Jin 0005, Sijia Liu 0007, Nanyun Peng 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 7 |
| 2023 | Investigating the Representation of Open Domain Dialogue Context for Transformer ModelsabstractVishakh Padmakumar, Behnam Hedayatnia, Di Jin, Patrick Lange, Seokhwan Kim, Nanyun Peng, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Vishakh Padmakumar, Behnam Hedayatnia, Di Jin 0005, Patrick Lange, Seokhwan Kim, Nanyun Peng 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 8 |
| 2023 | "What do others think?": Task-Oriented Conversational Modeling with Subjective KnowledgeabstractChao Zhao, Spandana Gella, Seokhwan Kim, Di Jin, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Spandana Gella, Seokhwan Kim, Di Jin 0005, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 10 |
| 2022 | Attention Biasing and Context Augmentation for Zero-Shot Control of Encoder-Decoder Transformers for Natural Language GenerationabstractControlling neural network-based models for natural language generation (NLG) to realize desirable attributes in the generated outputs has broad applications in numerous areas such as machine translation, document summarization, and dialog systems. Approaches that enable such control in a zero-shot manner would be of great importance as, among other reasons, they remove the need for additional annotated data and training. In this work, we propose novel approaches for controlling encoder-decoder transformer-based NLG models in zero shot. While zero-shot control has previously been observed in massive models (e.g., GPT3), our method enables such control for smaller models. This is done by applying two control knobs, attention biasing and context augmentation, to these models directly during decoding and without additional training or auxiliary models. These knobs control the generation process by directly manipulating trained NLG models (e.g., biasing cross-attention layers). We show that not only are these NLG models robust to such manipulations but also their behavior could be controlled without an impact on their generation performance. Devamanyu Hazarika, Mahdi Namazifar, Dilek Hakkani-Tür |
AAAI | 3 |
| 2022 | TEACh: Task-Driven Embodied Agents That ChatabstractRobots operating in human spaces must be able to engage in natural language interaction, both understanding and executing instructions, and using conversation to resolve ambiguity and correct mistakes. To study this, we introduce TEACh, a dataset of over 3,000 human-human, interactive dialogues to complete household tasks in simulation. A Commander with access to oracle information about a task communicates in natural language with a Follower. The Follower navigates through and interacts with the environment to complete tasks varying in complexity from "Make Coffee" to "Prepare Breakfast", asking questions and getting additional information from the Commander. We propose three benchmarks using TEACh to study embodied intelligence challenges, and we evaluate initial models' abilities in dialogue understanding, language grounding, and task execution. Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gökhan Tür, Dilek Hakkani-Tür |
AAAI | 9 |
| 2022 | Think Before You Speak: Explicitly Generating Implicit Commonsense Knowledge for Response GenerationabstractPei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren 0001, Yang Liu 0004, Dilek Hakkani-Tür |
ACL (1) | 8 |
| 2022 | ALFRED-L: Investigating the Role of Language for Action Learning in Interactive Visual EnvironmentsabstractArjun Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tur. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Arjun R. Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tür |
EMNLP | 7 |
| 2022 | Inducer-tuning: Connecting Prefix-tuning and Adapter-tuningabstractPrefix-tuning, or more generally continuous prompt tuning, has become an essential paradigm of parameter-efficient transfer learning.Using a large pre-trained language model (PLM), prefix-tuning can obtain strong performance by training only a small portion of parameters.In this paper, we propose to understand and further develop prefix-tuning through the kernel lens.Specifically, we make an analogy between prefixes and inducing variables in kernel methods and hypothesize that prefixes serving as inducing variables would improve their overall mechanism.From the kernel estimator perspective, we suggest a new variant of prefix-tuning-inducer-tuning, which shares the exact mechanism as prefix-tuning while leveraging the residual form found in adaptertuning.This mitigates the initialization issue in prefix-tuning.Through comprehensive empirical experiments on natural language understanding and generation tasks, we demonstrate that inducer-tuning can close the performance gap between prefix-tuning and fine-tuning. Yifan Chen 0004, Devamanyu Hazarika, Mahdi Namazifar, Yang Liu 0004, Di Jin 0005, Dilek Hakkani-Tür |
EMNLP | 6 |
| 2022 | Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding SystemsabstractWe present results from a large-scale experiment on pretraining encoders with non-embedding parameter counts ranging from 700M to 9.3B, their subsequent distillation into smaller models ranging from 17M-170M parameters, and their application to the Natural Language Understanding (NLU) component of a virtual assistant system. Though we train using 70% spoken-form data, our teacher models perform comparably to XLM-R and mT5 when evaluated on the written-form Cross-lingual Natural Language Inference (XNLI) corpus. We perform a second stage of pretraining on our teacher models using in-domain data from our system, improving error rates by 3.86% relative for intent classification and 7.01% relative for slot filling. We find that even a 170M-parameter model distilled from our Stage 2 teacher model has 2.88% better intent classification and 7.69% better slot filling error rates when compared to the 2.3B-parameter teacher trained only on public data (Stage 1), emphasizing the importance of in-domain data for pretraining. When evaluated offline using labeled NLU data, our 17M-parameter Stage 2 distilled model outperforms both XLM-R Base (85M params) and DistillBERT (42M params) by 4.23% to 6.14%, respectively. Finally, we present results from a full virtual assistant experimentation platform, where we find that models trained using our pretraining and distillation pipeline outperform models distilled from 85M-parameter teachers by 3.74%-4.91% on an automatic measurement of full-system user dissatisfaction. Jack FitzGerald, Shankar Ananthakrishnan, Konstantine Arkoudas, Davide Bernardi, Abhishek Bhagia, Claudio Delli Bovi, Jin Cao 0003, Rakesh Chada, Amit Chauhan, Luoxin Chen, Anurag Dwarakanath, Satyam Dwivedi, Turan Gojayev, Karthik Gopalakrishnan 0001, Thomas Gueudré, Dilek Hakkani-Tür, Wael Hamza, Jonathan J. Hüser, Kevin Martin Jose, Haidar Khan, Beiye Liu, Jianhua Lu, Alessandro Manzotti, Pradeep Natarajan, Karolina Owczarzak, Gokmen Oz, Enrico Palumbo, Charith Peris, Chandana Satya Prakash, Stephen Rawls, Andy Rosenbaum, Anjali Shenoy, Saleh Soltan, Mukund Sridhar, Lizhen Tan, Fabian Triefenbach, Pan Wei, Shuai Zheng 0004, Gökhan Tür, Premkumar Natarajan |
KDD | 16 |
| 2022 | Sketching as a Tool for Understanding and Accelerating Self-attention for Long SequencesabstractYifan Chen, Qi Zeng, Dilek Hakkani-Tur, Di Jin, Heng Ji, Yun Yang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yifan Chen 0004, Qi Zeng 0001, Dilek Hakkani-Tür, Di Jin 0005, Heng Ji 0001 |
NAACL-HLT | 3 |
| 2022 | Enhancing Knowledge Selection for Grounded Dialogues via Document Semantic GraphsabstractSha Li, Mahdi Namazifar, Di Jin, Mohit Bansal, Heng Ji, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Mahdi Namazifar, Di Jin 0005, Mohit Bansal, Heng Ji 0001, Yang Liu 0004, Dilek Hakkani-Tür |
NAACL-HLT | 7 |
| 2022 | Dialog Acts for Task Driven Embodied AgentsabstractEmbodied agents need to be able to interact in natural language -understanding task descriptions and asking appropriate follow up questions to obtain necessary information to be effective at successfully accomplishing tasks for a wide range of users.In this work, we propose a set of dialog acts for modelling such dialogs and annotate the TEACh dataset that includes over 3,000 situated, task oriented conversations (consisting of 39.5k utterances in total) with dialog acts.TEACh-DA is one of the first large scale dataset of dialog act annotations for embodied task completion.Furthermore, we demonstrate the use of this annotated dataset in training models for tagging the dialog acts of a given utterance, predicting the dialog act of the next response given a dialog history, and use the dialog acts to guide agent's non-dialog behaviour.In particular, our experiments on the TEACh Execution from Dialog History task where the model predicts the sequence of low level actions to be executed in the environment for embodied task completion, demonstrate that dialog acts can improve end task success rate by up to 2 points compared to the system without dialog acts. Spandana Gella, Aishwarya Padmakumar, Patrick Lange, Dilek Hakkani-Tür |
SIGDIAL | 4 |
| 2022 | A Systematic Evaluation of Response Selection for Open Domain DialogueabstractRecent progress on neural approaches for language processing has triggered a resurgence of interest on building intelligent open-domain chatbots.However, even the state-of-the-art neural chatbots cannot produce satisfying responses for every turn in a dialog.A practical solution is to generate multiple response candidates for the same context, and then perform response ranking/selection to determine which candidate is the best.Previous work in response selection typically trains response rankers using synthetic data that is formed from existing dialogs by using a ground truth response as the single appropriate response and constructing inappropriate responses via random selection or using adversarial methods.In this work, we curated a dataset where responses from multiple response generators produced for the same dialog context are manually annotated as appropriate (positive) and inappropriate (negative).We argue that such training data better matches the actual use case examples, enabling the models to learn to rank responses effectively.With this new dataset, we conduct a systematic evaluation of state-of-the-art methods for response selection, and demonstrate that both strategies of using multiple positive candidates and using manually verified hard negative candidates can bring in significant performance improvement in comparison to using the adversarial training data, e.g., increase of 3% and 13% in Recall@1 score, respectively. Behnam Hedayatnia, Di Jin 0005, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 4 |
| 2022 | Improving Bot Response Contradiction Detection via Utterance RewritingabstractThough chatbots based on large neural models can often produce fluent responses in open domain conversations, one salient error type is contradiction or inconsistency with the preceding conversation turns.Previous work has treated contradiction detection in bot responses as a task similar to natural language inference, e.g., detect the contradiction between a pair of bot utterances.However, utterances in conversations may contain co-references or ellipsis, and using these utterances as is may not always be sufficient for identifying contradictions.This work aims to improve the contradiction detection via rewriting all bot utterances to restore antecedents and ellipsis.We curated a new dataset for utterance rewriting and built a rewriting model on it.We empirically demonstrate that this model can produce satisfactory rewrites to make bot utterances more complete.Furthermore, using rewritten utterances improves contradiction detection performance significantly, e.g., the AUPR and joint accuracy scores (detecting contradiction along with evidence) increase by 6.5% and 4.5% (absolute increase), respectively. Di Jin 0005, Sijia Liu 0007, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 4 |
| 2022 | Knowledge-Grounded Conversational Data Augmentation with Generative Conversational NetworksabstractWhile rich, open-domain textual data are generally available and may include interesting phenomena (humor, sarcasm, empathy, etc.) most are designed for language processing tasks, and are usually in a non-conversational format.In this work, we take a step towards automatically generating conversational data using Generative Conversational Networks, aiming to benefit from the breadth of available language and knowledge data, and train open domain social conversational agents.We evaluate our approach on conversations with and without knowledge on the Topical Chat dataset using automatic metrics and human evaluators.Our results show that for conversations without knowledge grounding, GCN can generalize from the seed data, producing novel conversations that are less relevant but more engaging and for knowledge-grounded conversations, it can produce more knowledge-focused, fluent, and engaging conversations.Specifically, we show that for open-domain conversations with 10% of seed data, our approach performs close to the baseline that uses 100% of the data, while for knowledge-grounded conversations, it achieves the same using only 1% of the data, on human ratings of engagingness, fluency, and relevance. Alexandros Papangelis, Seokhwan Kim, Dilek Hakkani-Tür |
SIGDIAL | 4 |
| 2022 | N-Best Hypotheses Reranking for Text-to-SQL SystemsabstractText-to-SQL task maps natural language utterances to structured queries that can be issued to a database. State-of-the-art (SOTA) systems rely on finetuning large, pre-trained language models in conjunction with constrained decoding applying a SQL parser. On the well established Spider dataset, we begin with Oracle studies: specifically, choosing an Oracle hypothesis from a SOTA model's 10-best list, yields a 7.7% absolute improvement in both exact match (EM) and execution (EX) accuracy, showing significant potential improvements with reranking. Identifying coherence and correctness as reranking approaches, we design a model generating a query plan and propose a heuristic schema linking algorithm. Combining both approaches, with T5-Large, we obtain a consistent 1% improvement in EM accuracy, and a 2.5% improvement in EX, establishing a new SOTA for this task. Our comprehensive error studies on DEV data show the underlying difficulty in making progress on this task. Sree Hari Krishnan Parthasarathi, Dilek Hakkani-Tür |
SLT | 3 |
| 2022 | Revisiting the Boundary between ASR and NLU in the Age of Conversational Dialog SystemsabstractAbstract As more users across the world are interacting with dialog agents in their daily life, there is a need for better speech understanding that calls for renewed attention to the dynamics between research in automatic speech recognition (ASR) and natural language understanding (NLU). We briefly review these research areas and lay out the current relationship between them. In light of the observations we make in this article, we argue that (1) NLU should be cognizant of the presence of ASR models being used upstream in a dialog system’s pipeline, (2) ASR should be able to learn from errors found in NLU, (3) there is a need for end-to-end data sets that provide semantic annotations on spoken input, (4) there should be stronger collaboration between ASR and NLU research communities. Manaal Faruqui, Dilek Hakkani-Tür |
Comput. Linguistics | 2 |
| 2022 | Towards Textual Out-of-Domain Detection Without In-Domain LabelsabstractIn many real-world settings, machine learning models need to identify user inputs that are out-of-domain (OOD) so as to avoid performing wrong actions. This work focuses on a challenging case of OOD detection, where no labels for in-domain data are accessible (e.g., no intent labels for the intent classification task). To this end, we first evaluate different language model based approaches that predict likelihood for a sequence of tokens. Furthermore, we propose a novel representation learning based method by combining unsupervised clustering and contrastive learning so that better data representations for OOD detection can be learned. Through extensive experiments, we demonstrate that this method can significantly outperform likelihood-based methods and can be even competitive to the state-of-the-art supervised approaches with label information. Di Jin 0005, Shuyang Gao, Seokhwan Kim, Yang Liu 0004, Dilek Hakkani-Tür |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | "How Robust R U?": Evaluating Task-Oriented Dialogue Systems on Spoken ConversationsabstractMost prior work in dialogue modeling has been on written conversations mostly because of existing data sets. However, written dialogues are not sufficient to fully capture the nature of spoken conversations as well as the potential speech recognition errors in practical spoken dialogue systems. This work presents a new benchmark on spoken task-oriented conversations, which is intended to study multi-domain dialogue state tracking and knowledge-grounded dialogue modeling. We report that the existing state-of-the-art models trained on written conversations are not performing well on our spoken data, as expected. Furthermore, we observe improvements in task performances when leveraging$n$-best speech recognition hypotheses such as by combining predictions based on individual hypotheses. Our data set enables speech-based benchmarking of task-oriented dialogue systems. Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Dilek Hakkani-Tür |
ASRU | 7 |
| 2021 | Few Shot Dialogue State Tracking using Meta-learningabstractSaket Dingliwal, Shuyang Gao, Sanchit Agarwal, Chien-Wei Lin, Tagyoung Chung, Dilek Hakkani-Tur. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Saket Dingliwal, Shuyang Gao, Sanchit Agarwal, Chien-Wei Lin, Tagyoung Chung, Dilek Hakkani-Tür |
EACL | 6 |
| 2021 | Language Model is all You Need: Natural Language Understanding as Question AnsweringabstractDifferent flavors of transfer learning have shown tremendous impact in advancing research and applications of machine learning. In this work we study the use of a certain family of transfer learning, where the target domain is mapped to the source domain. Specifically we map Natural Language Understanding (NLU) problems to Question Answering (QA) problems and we show that in low data regimes this approach offers significant improvements compared to other approaches to NLU. Moreover, we show that these gains could be increased through sequential transfer learning across NLU problems from different domains. We show that our approach could reduce the amount of required data for the same performance by up to a factor of 10. Mahdi Namazifar, Alexandros Papangelis, Gökhan Tür, Dilek Hakkani-Tür |
ICASSP | 4 |
| 2021 | Multi-Sentence Knowledge Selection in Open-Domain DialogueabstractMihail Eric, Nicole Chartier, Behnam Hedayatnia, Karthik Gopalakrishnan, Pankaj Rajan, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 14th International Conference on Natural Language Generation. 2021. Mihail Eric, Nicole Chartier, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Pankaj Rajan, Yang Liu 0004, Dilek Hakkani-Tür |
INLG | 7 |
| 2021 | Correcting Automated and Manual Speech Transcription Errors Using Warped Language ModelsabstractMasked language models have revolutionized natural language processing systems in the past few years. A recently introduced generalization of masked language models called warped language models are trained to be more robust to the types of errors that appear in automatic or manual transcriptions of spoken language by exposing the language model to the same types of errors during training. In this work we propose a novel approach that takes advantage of the robustness of warped language models to transcription noise for correcting transcriptions of spoken language. We show that our proposed approach is able to achieve up to 10% reduction in word error rates of both automatic and manual transcriptions of spoken language. Mahdi Namazifar, John Malik, Li Erran Li, Gökhan Tür, Dilek Hakkani-Tür |
Interspeech | 5 |
| 2021 | Generative Conversational NetworksabstractAlexandros Papangelis, Karthik Gopalakrishnan, Aishwarya Padmakumar, Seokhwan Kim, Gokhan Tur, Dilek Hakkani-Tur. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Alexandros Papangelis, Karthik Gopalakrishnan 0001, Aishwarya Padmakumar, Seokhwan Kim, Gökhan Tür, Dilek Hakkani-Tür |
SIGDIAL | 6 |
| 2021 | Commonsense-Focused Dialogues for Response Generation: An Empirical StudyabstractPei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 8 |
| 2021 | Go Beyond Plain Fine-Tuning: Improving Pretrained Models for Social CommonsenseabstractPretrained language models have demonstrated outstanding performance in many NLP tasks recently. However, their social intelligence, which requires commonsense reasoning about the current situation and mental states of others, is still developing. Towards improving language models' social intelligence, in this study we focus on the Social IQA dataset, a task requiring social and emotional commonsense reasoning. Building on top of the pretrained RoBERTa and GPT2 models, we propose several architecture variations and extensions, as well as leveraging external commonsense corpora, to optimize the model for Social IQA. Our proposed system achieves competitive results as those top-ranking models on the leaderboard. This work demonstrates the strengths of pretrained language models, and provides viable ways to improve their performance for a particular task. Ting-Yun Chang, Yang Liu 0004, Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Dilek Hakkani-Tür |
SLT | 6 |
| 2021 | Warped Language Models for Noise Robust Language UnderstandingabstractMasked Language Models (MLM) are self-supervised neural networks trained to fill in the blanks in a given sentence with masked tokens. Despite the tremendous success of MLMs for various text based tasks, they are not robust for spoken language understanding, especially for spontaneous conversational speech recognition noise. In this work we introduce Warped Language Models (WLM) in which input sentences at training time go through the same modifications as in MLM, plus two additional modifications, namely inserting and dropping random tokens. These two modifications extend and contract the sentence in addition to the modifications in MLMs, hence the word "warped" in the name. The insertion and drop modification of the input text during training of WLM resemble the types of noise due to Automatic Speech Recognition (ASR) errors, and as a result WLMs are likely to be more robust to ASR noise. Through computational results we show that natural language understanding systems built on top of WLMs perform better compared to those built based on MLMs, especially in the presence of ASR errors. Mahdi Namazifar, Gökhan Tür, Dilek Hakkani-Tür |
SLT | 3 |
| 2020 | Just Ask: An Interactive Learning Framework for Vision and Language NavigationabstractIn the vision and language navigation task (Anderson et al. 2018), the agent may encounter ambiguous situations that are hard to interpret by just relying on visual information and natural language instructions. We propose an interactive learning framework to endow the agent with the ability to ask for users' help in such situations. As part of this framework, we investigate multiple learning approaches for the agent with different levels of complexity. The simplest model-confusion-based method lets the agent ask questions based on its confusion, relying on the predefined confidence threshold of a next action prediction model. To build on this confusion-based method, the agent is expected to demonstrate more sophisticated reasoning such that it discovers the timing and locations to interact with a human. We achieve this goal using reinforcement learning (RL) with a proposed reward shaping term, which enables the agent to ask questions only when necessary. The success rate can be boosted by at least 15% with only one question asked on average during the navigation. Furthermore, we show that the RL agent is capable of adjusting dynamically to noisy human responses. Finally, we design a continual learning strategy, which can be viewed as a data augmentation method, for the agent to improve further utilizing its interaction history with a human. We demonstrate the proposed strategy is substantially more realistic and data-efficient compared to previously proposed pre-exploration techniques. Ta-Chung Chi, Minmin Shen, Mihail Eric, Seokhwan Kim, Dilek Hakkani-Tür |
AAAI | 5 |
| 2020 | MMM: Multi-Stage Multi-Task Learning for Multi-Choice Reading ComprehensionabstractMachine Reading Comprehension (MRC) for question answering (QA), which aims to answer a question given the relevant context passages, is an important way to test the ability of intelligence systems to understand human language. Multiple-Choice QA (MCQA) is one of the most difficult tasks in MRC because it often requires more advanced reading comprehension skills such as logical reasoning, summarization, and arithmetic operations, compared to the extractive counterpart where answers are usually spans of text within given passages. Moreover, most existing MCQA datasets are small in size, making the task even harder. We introduce MMM, a Multi-stage Multi-task learning framework for Multi-choice reading comprehension. Our method involves two sequential stages: coarse-tuning stage using out-of-domain datasets and multi-task learning stage using a larger in-domain dataset to help model generalize better with limited data. Furthermore, we propose a novel multi-step attention network (MAN) as the top-level classifier for this task. We demonstrate MMM significantly advances the state-of-the-art on four representative MCQA datasets. Di Jin 0005, Shuyang Gao, Jiun-Yu Kao, Tagyoung Chung, Dilek Hakkani-Tür |
AAAI | 5 |
| 2020 | MA-DST: Multi-Attention-Based Scalable Dialog State TrackingabstractTask oriented dialog agents provide a natural language interface for users to complete their goal. Dialog State Tracking (DST), which is often a core component of these systems, tracks the system's understanding of the user's goal throughout the conversation. To enable accurate multi-domain DST, the model needs to encode dependencies between past utterances and slot semantics and understand the dialog context, including long-range cross-domain references. We introduce a novel architecture for this task to encode the conversation history and slot semantics more robustly by using attention mechanisms at multiple granularities. In particular, we use cross-attention to model relationships between the context and slots at different semantic levels and self-attention to resolve cross-domain coreferences. In addition, our proposed architecture does not rely on knowing the domain ontologies beforehand and can also be used in a zero-shot setting for new domains or unseen slot values. Our model improves the joint goal accuracy by 5% (absolute) in the full-data setting and by up to 2% (absolute) in the zero-shot setting over the present state-of-the-art on the MultiWoZ 2.1 dataset. Adarsh Kumar 0001, Peter Ku, Anuj Kumar Goyal, Angeliki Metallinou, Dilek Hakkani-Tür |
AAAI | 5 |
| 2020 | Schema-Guided Natural Language GenerationabstractYuheng Du, Shereen Oraby, Vittorio Perera, Minmin Shen, Anjali Narayan-Chen, Tagyoung Chung, Anushree Venkatesh, Dilek Hakkani-Tur. Proceedings of the 13th International Conference on Natural Language Generation. 2020. Yuheng Du, Shereen Oraby, Vittorio Perera, Minmin Shen, Anjali Narayan-Chen, Tagyoung Chung, Anu Venkatesh, Dilek Hakkani-Tür |
INLG | 8 |
| 2020 | Policy-Driven Neural Response Generation for Knowledge-Grounded Dialog SystemsabstractOpen-domain dialog systems aim to generate relevant, informative and engaging responses.In this paper, we propose using a dialog policy to plan the content and style of target, opendomain responses in the form of an action plan, which includes knowledge sentences related to the dialog context, targeted dialog acts, topic information, etc.For training, the attributes within the action plan are obtained by automatically annotating the publicly released Topical-Chat dataset.We condition neural response generators on the action plan which is then realized as target utterances at the turn and sentence levels.We also investigate different dialog policy models to predict an action plan given the dialog context.Through automated and human evaluation, we measure the appropriateness of the generated responses and check if the generation models indeed learn to realize the given action plans.We demonstrate that a basic dialog policy that operates at the sentence level generates better responses in comparison to turn level generation as well as baseline models with no action plan.Additionally the basic dialog policy has the added benefit of controllability. Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Seokhwan Kim, Yang Liu 0004, Mihail Eric, Dilek Hakkani-Tür |
INLG | 6 |
| 2020 | Are Neural Open-Domain Dialog Systems Robust to Speech Recognition Errors in the Dialog History? An Empirical StudyabstractLarge end-to-end neural open-domain chatbots are becoming increasingly popular. However, research on building such chatbots has typically assumed that the user input is written in nature and it is not clear whether these chatbots would seamlessly integrate with automatic speech recognition (ASR) models to serve the speech modality. We aim to bring attention to this important question by empirically studying the effects of various types of synthetic and actual ASR hypotheses in the dialog history on TransferTransfo, a state-of-the-art Generative Pre-trained Transformer (GPT) based neural open-domain dialog system from the NeurIPS ConvAI2 challenge. We observe that TransferTransfo trained on written data is very sensitive to such hypotheses introduced to the dialog history during inference time. As a baseline mitigation strategy, we introduce synthetic ASR hypotheses to the dialog history during training and observe marginal improvements, demonstrating the need for further research into techniques to make end-to-end open-domain chatbots fully speech-robust. To the best of our knowledge, this is the first study to evaluate the effects of synthetic and actual ASR hypotheses on a state-of-the-art neural open-domain dialog system and we hope it promotes speech-robustness as an evaluation criterion in open-domain dialog. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Longshaokan Wang, Yang Liu 0004, Dilek Hakkani-Tür |
INTERSPEECH | 5 |
| 2020 | MultiWOZ 2.1: A Consolidated Multi-Domain Dialogue Dataset with State Corrections and State Tracking BaselinesabstractMultiWOZ 2.0 (Budzianowski et al., 2018) is a recently released multi-domain dialogue dataset spanning 7 distinct domains and containing over 10,000 dialogues. Though immensely useful and one of the largest resources of its kind to-date, MultiWOZ 2.0 has a few shortcomings. Firstly, there are substantial noise in the dialogue state annotations and dialogue utterances which negatively impact the performance of state-tracking models. Secondly, follow-up work (Lee et al., 2019) has augmented the original dataset with user dialogue acts. This leads to multiple co-existent versions of the same dataset with minor modifications. In this work we tackle the aforementioned issues by introducing MultiWOZ 2.1. To fix the noisy state annotations, we use crowdsourced workers to re-annotate state and utterances based on the original utterances in the dataset. This correction process results in changes to over 32% of state annotations across 40% of the dialogue turns. In addition, we fix 146 dialogue utterances by canonicalizing slot values in the utterances to the values in the dataset ontology. To address the second problem, we combined the contributions of the follow-up works into MultiWOZ 2.1. Hence, our dataset also includes user dialogue acts as well as multiple slot descriptions per dialogue state slot. We then benchmark a number of state-of-the-art dialogue state tracking models on the MultiWOZ 2.1 dataset and show the joint state tracking performance on the corrected state annotations. We are publicly releasing MultiWOZ 2.1 to the community, hoping that this dataset resource will allow for more effective models across various dialogue subproblems to be built in the future. Mihail Eric, Rahul Goel, Shachi Paul, Abhishek Sethi, Sanchit Agarwal, Shuyang Gao, Adarsh Kumar 0001, Anuj Kumar Goyal, Peter Ku, Dilek Hakkani-Tür |
LREC | 10 |
| 2020 | Beyond Domain APIs: Task-oriented Conversational Modeling with Unstructured Knowledge AccessabstractMost prior work on task-oriented dialogue systems are restricted to a limited coverage of domain APIs, while users oftentimes have domain related requests that are not covered by the APIs.In this paper, we propose to expand coverage of task-oriented dialogue systems by incorporating external unstructured knowledge sources.We define three sub-tasks: knowledge-seeking turn detection, knowledge selection, and knowledge-grounded response generation, which can be modeled individually or jointly.We introduce an augmented version of MultiWOZ 2.1, which includes new out-of-API-coverage turns and responses grounded on external knowledge sources.We present baselines for each sub-task using both conventional and neural approaches.Our experimental results demonstrate the need for further research in this direction to enable more informative conversational systems. Seokhwan Kim, Mihail Eric, Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Yang Liu 0004, Dilek Hakkani-Tür |
SIGdial | 6 |
| 2020 | Learning from Mistakes: Combining Ontologies via Self-Training for Dialogue GenerationabstractNatural language generators (NLGs) for taskoriented dialogue typically take a meaning representation (MR) as input, and are trained endto-end with a corpus of MR/utterance pairs, where the MRs cover a specific set of dialogue acts and domain attributes.Creation of such datasets is labor intensive and time consuming.Therefore, dialogue systems for new domain ontologies would benefit from using data for pre-existing ontologies.Here we explore, for the first time, whether it is possible to train an NLG for a new larger ontology using existing training sets for the restaurant domain, where each set is based on a different ontology.We create a new, larger combined ontology, and then train an NLG to produce utterances covering it.For example, if one dataset has attributes for family friendly and rating information, and the other has attributes for decor and service, our aim is an NLG for the combined ontology that can produce utterances that realize values for family friendly, rating, decor and service.Initial experiments with a baseline neural sequence-to-sequence model show that this task is surprisingly challenging.We then develop a novel self-training method that identifies (errorful) model outputs, automatically constructs a corrected MR input to form a new (MR, utterance) training pair, and then repeatedly adds these new instances back into the training data.We then test the resulting model on a new test set.The result is a selftrained model whose performance is an absolute 75.4% improvement over the baseline model.We also report a human qualitative evaluation of the final model showing that it achieves high naturalness, semantic coherence and grammaticality. Lena Reed, Vrindavan Harrison, Shereen Oraby, Dilek Hakkani-Tür, Marilyn A. Walker |
SIGdial | 4 |
| 2020 | ConvERSe'20: The WSDM 2020 Workshop on Conversational Systems for E-Commerce Recommendations and SearchabstractConversational systems have improved dramatically recently, and are receiving increasing attention in academic literature. These systems are also becoming adapted in E-Commerce due to increased integration of E-Commerce search and recommendation source with virtual assistants such as Alexa, Siri, and Google assistant. However, significant research challenges remain spanning areas of dialogue systems, spoken natural language processing, human-computer interaction, and search and recommender systems, which all are exacerbated with demanding requirements of E-Commerce. Eugene Agichtein, Dilek Hakkani-Tür, Surya Kallumadi, Shervin Malmasi |
WSDM | 2 |
| 2019 | Robust Zero-Shot Cross-Domain Slot Filling with Example ValuesabstractTask-oriented dialog systems increasingly rely on deep learning-based slot filling models, usually needing extensive labeled training data for target domains.Often, however, little to no target domain training data may be available, or the training and target domain schemas may be misaligned, as is common for web forms on similar websites.Prior zero-shot slot filling models use slot descriptions to learn concepts, but are not robust to misaligned schemas.We propose utilizing both the slot description and a small number of examples of slot values, which may be easily available, to learn semantic representations of slots which are transferable across domains and robust to misaligned schemas.Our approach outperforms state-ofthe-art models on two multi-domain datasets, especially in the low-data setting. Darsh J. Shah, Amir A. Fayazi, Dilek Hakkani-Tür |
ACL (1) | 4 |
| 2019 | Learning to Navigate the Web
Izzeddin Gur, Ulrich Rückert 0003, Aleksandra Faust, Dilek Hakkani-Tür |
ICLR (Poster) | 4 |
| 2019 | Natural Language Generation at Scale: A Case Study for Open Domain Question AnsweringabstractAlessandra Cervone, Chandra Khatri, Rahul Goel, Behnam Hedayatnia, Anu Venkatesh, Dilek Hakkani-Tur, Raefer Gabriel. Proceedings of the 12th International Conference on Natural Language Generation. 2019. Alessandra Cervone, Chandra Khatri, Rahul Goel, Behnam Hedayatnia, Anu Venkatesh, Dilek Hakkani-Tür, Raefer Gabriel |
INLG | 6 |
| 2019 | Towards Coherent and Engaging Spoken Dialog Response Generation Using Automatic Conversation EvaluatorsabstractSanghyun Yi, Rahul Goel, Chandra Khatri, Alessandra Cervone, Tagyoung Chung, Behnam Hedayatnia, Anu Venkatesh, Raefer Gabriel, Dilek Hakkani-Tur. Proceedings of the 12th International Conference on Natural Language Generation. 2019. Sanghyun Yi, Rahul Goel, Chandra Khatri, Alessandra Cervone, Tagyoung Chung, Behnam Hedayatnia, Anu Venkatesh, Raefer Gabriel, Dilek Hakkani-Tür |
INLG | 9 |
| 2019 | HyST: A Hybrid Approach for Flexible and Accurate Dialogue State TrackingabstractRecent works on end-to-end trainable neural network based approaches have demonstrated state-of-the-art results on dialogue state tracking. The best performing approaches estimate a probability distribution over all possible slot values. However, these approaches do not scale for large value sets commonly present in real-life applications and are not ideal for tracking slot values that were not observed in the training set. To tackle these issues, candidate-generation-based approaches have been proposed. These approaches estimate a set of values that are possible at each turn based on the conversation history and/or language understanding outputs, and hence enable state tracking over unseen values and large value sets however, they fall short in terms of performance in comparison to the first group. In this work, we analyze the performance of these two alternative dialogue state tracking methods, and present a hybrid approach (HyST) which learns the appropriate method for each slot type. To demonstrate the effectiveness of HyST on a rich-set of slot types, we experiment with the recently released MultiWOZ-2.0 multi-domain, task-oriented dialogue-dataset. Our experiments show that HyST scales to multi-domain applications. Our best performing model results in a relative improvement of 24% and 10% over the previous SOTA and our best baseline respectively. Rahul Goel, Shachi Paul, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2019 | Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations
Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Qinlang Chen 0001, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, Dilek Hakkani-Tür |
INTERSPEECH | 8 |
| 2019 | Towards Universal Dialogue Act Tagging for Task-Oriented DialoguesabstractMachine learning approaches for building task-oriented dialogue systems require large conversational datasets with labels to train on. We are interested in building task-oriented dialogue systems from human-human conversations, which may be available in ample amounts in existing customer care center logs or can be collected from crowd workers. Annotating these datasets can be prohibitively expensive. Recently multiple annotated task-oriented human-machine dialogue datasets have been released, however their annotation schema varies across different collections, even for well-defined categories such as dialogue acts (DAs). We propose a Universal DA schema for task-oriented dialogues and align existing annotated datasets with our schema. Our aim is to train a Universal DA tagger (U-DAT) for task-oriented dialogues and use it for tagging human-human conversations. We investigate multiple datasets, propose manual and automated approaches for aligning the different schema, and present results on a target corpus of human-human dialogues. In unsupervised learning experiments we achieve an F1 score of 54.1% on system turns in human-human dialogues. In a semi-supervised setup, the F1 score increases to 57.7% which would otherwise require at least 1.7K manually annotated turns. For new domains, we show further improvements when unlabeled or labeled target domain data is available. Shachi Paul, Rahul Goel, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2019 | Learning Question-Guided Video Representation for Multi-Turn Video Question AnsweringabstractUnderstanding and conversing about dynamic scenes is one of the key capabilities of AI agents that navigate the environment and convey useful information to humans.Video question answering is a specific scenario of such AI-human interaction where an agent generates a natural language response to a question regarding the video of a dynamic scene.Incorporating features from multiple modalities, which often provide supplementary information, is one of the challenging aspects of video question answering.Furthermore, a question often concerns only a small segment of the video, hence encoding the entire video sequence using a recurrent neural network is not computationally efficient.Our proposed question-guided video representation module efficiently generates the token-level video summary guided by each word in the question.The learned representations are then fused with the question to generate the answer.Through empirical evaluation on the Audio Visual Scene-aware Dialog (AVSD) dataset (Alamri et al., 2019a), our proposed models in single-turn and multiturn question answering achieve state-of-theart performance on several automatic natural language generation evaluation metrics. Guan-Lin Chao, Abhinav Rastogi, Semih Yavuz, Dilek Hakkani-Tür, Jindong Chen, Ian Lane |
SIGdial | 4 |
| 2019 | Dialog State Tracking: A Neural Reading Comprehension ApproachabstractDialog state tracking is used to estimate the current belief state of a dialog given all the preceding conversation.Machine reading comprehension, on the other hand, focuses on building systems that read passages of text and answer questions that require some understanding of passages.We formulate dialog state tracking as a reading comprehension task to answer the question what is the state of the current dialog?after reading conversational context.In contrast to traditional state tracking methods where the dialog state is often predicted as a distribution over a closed set of all the possible slot values within an ontology, our method uses a simple attention-based neural network to point to the slot values within the conversation.Experiments on MultiWOZ-2.0 cross-domain dialog dataset show that our simple system can obtain similar accuracies compared to the previous more complex methods.By exploiting recent advances in contextual word embeddings, adding a model that explicitly tracks whether a slot value should be carried over to the next turn, and combining our method with a traditional joint state tracking method that relies on closed set vocabulary, we can obtain a joint-goal accuracy of 47.33% on the standard test split, exceeding current state-of-the-art by 11.75%**. Shuyang Gao, Abhishek Sethi, Sanchit Agarwal, Tagyoung Chung, Dilek Hakkani-Tür |
SIGdial | 5 |
| 2019 | DeepCopy: Grounded Response Generation with Hierarchical Pointer NetworksabstractRecent advances in neural sequence-tosequence models have led to promising results for several language generation-based tasks, including dialogue response generation, summarization, and machine translation.However, these models are known to have several problems, especially in the context of chit-chat based dialogue systems: they tend to generate short and dull responses that are often too generic.Furthermore, these models do not ground conversational responses on knowledge and facts, resulting in turns that are not accurate, informative and engaging for the users.In this paper, we propose and experiment with a series of response generation models that aim to serve in the general scenario where in addition to the dialogue context, relevant unstructured external knowledge in the form of text is also assumed to be available for models to harness.Our proposed approach extends pointer-generator networks (See et al., 2017) by allowing the decoder to hierarchically attend and copy from external knowledge in addition to the dialogue context.We empirically show the effectiveness of the proposed model compared to several baselines including (Ghazvininejad et al., 2018;Zhang et al., 2018) through both automatic evaluation metrics and human evaluation on CONVAI2 dataset. Semih Yavuz, Abhinav Rastogi, Guan-Lin Chao, Dilek Hakkani-Tür |
SIGdial | 4 |
| 2019 | Inaugural Editorial Innovations in an Era of Ubiquitous Audio, Speech, and Language Processing
Dilek Hakkani-Tür |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | (Almost) Zero-Shot Cross-Lingual Spoken Language UnderstandingabstractSpoken language understanding (SLU) is a component of goal-oriented dialogue systems that aims to interpret user's natural language queries in system's semantic representation format. While current state-of-the-art SLU approaches achieve high performance for English domains, the same is not true for other languages. Approaches in the literature for extending SLU models and grammars to new languages rely primarily on machine translation. This poses a challenge in scaling to new languages, as machine translation systems may not be reliable for several (especially low resource) languages. In this work, we examine different approaches to train a SLU component with little supervision for two new languages - Hindi and Turkish, and show that with only a few hundred labeled examples we can surpass the approaches proposed in the literature. Our experiments show that training a model bilingually (i.e., jointly with English), enables faster learning, in that the model requires fewer labeled instances in the target language to generalize. Qualitative analysis shows that rare slot types benefit the most from the bilingual training. Shyam Upadhyay, Manaal Faruqui, Gökhan Tür, Dilek Hakkani-Tür, Larry Heck |
ICASSP | 4 |
| 2018 | An Efficient Approach to Encoding Context for Spoken Language UnderstandingabstractIn task-oriented dialogue systems, spoken language understanding, or SLU, refers to the task of parsing natural language user utterances into semantic frames. Making use of context from prior dialogue history holds the key to more effective SLU. State of the art approaches to SLU use memory networks to encode context by processing multiple utterances from the dialogue at each turn, resulting in significant trade-offs between accuracy and computational efficiency. On the other hand, downstream components like the dialogue state tracker (DST) already keep track of the dialogue state, which can serve as a summary of the dialogue history. In this work, we propose an efficient approach to encoding context from prior utterances for SLU. More specifically, our architecture includes a separate recurrent neural network (RNN) based encoding module that accumulates dialogue context to guide the frame parsing sub-tasks and can be shared between SLU and DST. In our experiments, we demonstrate the effectiveness of our approach on dialogues from two domains. Abhinav Rastogi, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2018 | Deep Learning based Situated Goal-oriented Dialogue Systems
Dilek Hakkani-Tür |
INTERSPEECH | 1 |
| 2018 | Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue SystemsabstractBing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, Larry Heck. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Bing Liu 0024, Gökhan Tür, Dilek Hakkani-Tür, Pararth Shah, Larry Heck |
NAACL-HLT | 3 |
| 2018 | Multi-task Learning for Joint Language Understanding and Dialogue State TrackingabstractThis paper presents a novel approach for multi-task learning of language understanding (LU) and dialogue state tracking (DST) in task-oriented dialogue systems.Multi-task training enables the sharing of the neural network layers responsible for encoding the user utterance for both LU and DST and improves performance while reducing the number of network parameters.In our proposed framework, DST operates on a set of candidate values for each slot that has been mentioned so far.These candidate sets are generated using LU slot annotations for the current user utterance, dialogue acts corresponding to the preceding system utterance and the dialogue state estimated for the previous turn, enabling DST to handle slots with a large or unbounded set of possible values and deal with slot values not seen during training.Furthermore, to bridge the gap between training and inference, we investigate the use of scheduled sampling on LU output for the current user utterance as well as the DST output for the preceding turn. Abhinav Rastogi, Dilek Hakkani-Tür |
SIGDIAL Conference | 3 |
| 2018 | User Modeling for Task Oriented DialoguesabstractWe introduce end-to-end neural network based models for simulating users of task-oriented dialogue systems. User simulation in dialogue systems is crucial from two different perspectives: (i) automatic evaluation of different dialogue models, and (ii) training task-oriented dialogue systems. We design a hierarchical sequence-to-sequence model that first encodes the initial user goal and system turns into fixed length representations using Recurrent Neural Networks (RNN). It then encodes the dialogue history using another RNN layer. At each turn, user responses are decoded from the hidden representations of the dialogue level RNN. This hierarchical user simulator (HUS) approach allows the model to capture undiscovered parts of the user goal without the need of an explicit dialogue state tracking. We further develop several variants by utilizing a latent variable model to inject random variations into user responses to promote diversity in simulated user responses and a novel goal regularization mechanism to penalize divergence of user responses from the initial user goal. We evaluate the proposed models on movie ticket booking domain by systematically interacting each user simulator with various dialogue system policies trained with different objectives and users. Izzeddin Gur, Dilek Hakkani-Tür, Gökhan Tür, Pararth Shah |
SLT | 2 |
| 2018 | Resolving Referring Expressions in Images with Labeled ElementsabstractImages may have elements containing text and a bounding box associated with them, for example, text identified via optical character recognition on a computer screen image, or a natural image with labeled objects. We present an end-to-end trainable architecture to incorporate the information from these elements and the image to segment/identify the part of the image a natural language expression is referring to. We calculate an embedding for each element and then project it onto the corresponding location (i.e., the associated bounding box) of the image feature map. We show that this architecture gives an improvement in resolving referring expressions, over only using the image, and other methods that incorporate the element information. We demonstrate experimental results on the referring expression datasets based on COCO and on a webpage image referring expression dataset that we developed. Nevan Wichers, Dilek Hakkani-Tür, Jindong Chen |
SLT | 2 |
| 2017 | Scalable multi-domain dialogue state trackingabstractDialogue state tracking (DST) is a key component of task-oriented dialogue systems. DST estimates the user's goal at each user turn given the interaction until then. State of the art approaches for state tracking rely on deep learning methods, and represent dialogue state as a distribution over all possible slot values for each slot present in the ontology. Such a representation is not scalable when the set of possible values are unbounded (e.g., date, time or location) or dynamic (e.g., movies or usernames). Furthermore, training of such models requires labeled data, where each user turn is annotated with the dialogue state, which makes building models for new domains challenging. In this paper, we present a scalable multi-domain deep learning based approach for DST. We introduce a novel framework for state tracking which is independent of the slot value set, and represent the dialogue state as a distribution over a set of values of interest (candidate set) derived from the dialogue history or knowledge. Restricting these candidate sets to be bounded in size addresses the problem of slot-scalability. Furthermore, by leveraging the slot-independent architecture and transfer learning, we show that our proposed approach facilitates quick adaptation to new domains. Abhinav Rastogi, Dilek Hakkani-Tür, Larry Heck |
ASRU | 2 |
| 2017 | Learning concepts through conversations in spoken dialogue systemsabstractSpoken dialogue systems must be able to recover gracefully from unexpected user inputs. In many cases, these unexpected utterances may be within the scope of the system, but include previously unseen phrases that the system cannot interpret. In this work, we augment a spoken dialogue system with the ability to learn about new concepts by conversing with the user in natural language. We present a novel model that detects phrases corresponding to such concepts, using information from a neural slotfiller as well as syntactic cues. The system then prompts the user for a definition of the detected phrases, and uses these definitions to re-parse the original utterance. We demonstrate significant gains by learning from the user, compared to a baseline system. Robin Jia, Larry Heck, Dilek Hakkani-Tür, Georgi Nikolov |
ICASSP | 3 |
| 2017 | End-to-end joint learning of natural language understanding and dialogue managerabstractNatural language understanding and dialogue policy learning are both essential in conversational systems that predict the next system actions in response to a current user utterance. Conventional approaches aggregate separate models of natural language understanding (NLU) and system action prediction (SAP) as a pipeline that is sensitive to noisy outputs of error-prone NLU. To address the issues, we propose an end-to-end deep recurrent neural network with limited contextual dialogue memory by jointly training NLU and SAP on DSTC4 multi-domain human-human dialogues. Experiments show that our proposed model significantly outperforms the state-of-the-art pipeline models for both NLU and SAP, which indicates that our joint model is capable of mitigating the affects of noisy NLU outputs, and NLU model can be refined by error flows backpropagating from the extra supervised signals of system actions. Xuesong Yang, Yun-Nung Chen, Dilek Hakkani-Tür, Paul A. Crook, Xiujun Li, Jianfeng Gao 0001, Li Deng 0001 |
ICASSP | 3 |
| 2017 | Towards Zero-Shot Frame Semantic Parsing for Domain ScalingabstractState-of-the-art slot filling models for goal-oriented human/machine conversational language understanding systems rely on deep learning methods. While multi-task training of such models alleviates the need for large in-domain annotated datasets, bootstrapping a semantic parsing model for a new domain using only the semantic frame, such as the back-end API or knowledge graph schema, is still one of the holy grail tasks of language understanding for dialogue systems. This paper proposes a deep learning based approach that can utilize only the slot description in context without the need for any labeled or unlabeled in-domain examples, to quickly bootstrap a new domain. The main idea of this paper is to leverage the encoding of the slot names and descriptions within a multi-task deep learned slot filling model, to implicitly align slots across domains. The proposed approach is promising for solving the domain scaling problem and eliminating the need for any manually annotated data or explicit schema alignment. Furthermore, our experiments on multiple domains show that this approach results in significantly better slot-filling performance when compared to using only in-domain data, especially in the low data regime. Ankur Bapna, Gökhan Tür, Dilek Hakkani-Tür, Larry Heck |
INTERSPEECH | 3 |
| 2017 | To Plan or not to Plan? Discourse Planning in Slot-Value Informed Sequence to Sequence Models for Language Generation
Neha Nayak, Dilek Hakkani-Tür, Marilyn A. Walker, Larry Heck |
INTERSPEECH | 2 |
| 2017 | Sequential Dialogue Context Modeling for Spoken Language UnderstandingabstractSpoken Language Understanding (SLU) is a key component of goal oriented dialogue systems that would parse user utterances into semantic frame representations.Traditionally SLU does not utilize the dialogue history beyond the previous system turn and contextual ambiguities are resolved by the downstream components.In this paper, we explore novel approaches for modeling dialogue context in a recurrent neural network (RNN) based language understanding system.We propose the Sequential Dialogue Encoder Network, that allows encoding context from the dialogue history in chronological order.We compare the performance of our proposed architecture with two context models, one that uses just the previous turn context and another that encodes dialogue context in a memory network, but loses the order of utterances in the dialogue history.Experiments with a multi-domain dialogue dataset demonstrate that the proposed architecture results in reduced semantic frame error rates. Ankur Bapna, Gökhan Tür, Dilek Hakkani-Tür, Larry Heck |
SIGDIAL Conference | 3 |
| 2017 | Spoken language understanding and interaction: machine learning for human-like conversational systems
Milica Gasic, Dilek Hakkani-Tür, Asli Celikyilmaz |
Comput. Speech Lang. | 2 |
| 2016 | Zero-shot learning of intent embeddings for expansion by convolutional deep structured semantic modelsabstractThe recent surge of intelligent personal assistants motivates spoken language understanding of dialogue systems. However, the domain constraint along with the inflexible intent schema remains a big issue. This paper focuses on the task of intent expansion, which helps remove the domain limit and make an intent schema flexible. A con-volutional deep structured semantic model (CDSSM) is applied to jointly learn the representations for human intents and associated utterances. Then it can flexibly generate new intent embeddings without the need of training samples and model-retraining, which bridges the semantic relation between seen and unseen intents and further performs more robust results. Experiments show that CDSSM is capable of performing zero-shot learning effectively, e.g. generating embeddings of previously unseen intents, and therefore expand to new intents without re-training, and outperforms other semantic embeddings. The discussion and analysis of experiments provide a future direction for reducing human effort about annotating data and removing the domain constraint in spoken dialogue systems. Yun-Nung Chen, Dilek Hakkani-Tür, Xiaodong He 0001 |
ICASSP | 2 |
| 2016 | A New Pre-Training Method for Training Deep Learning Models with Application to Spoken Language UnderstandingabstractWe propose a simple and efficient approach for pre-training deep learning models with application to slot filling tasks in spoken language understanding. The proposed approach leverages unlabeled data to train the models and is generic enough to work with any deep learning model. In this study, we consider the CNN2CRF architecture that contains Convolutional Neural Network (CNN) with Conditional Random Fields (CRF) as top layer, since it has shown great potential for learning useful representations for supervised sequence learning tasks. The proposed pre-training approach with this architecture learns the feature representations from both labeled and unlabeled data at the CNN layer, covering features that would not be observed in limited labeled data. At the CRF layer, the unlabeled data uses predicted classes of words as latent sequence labels together with labeled sequences. Latent labeled sequences, in principle, has the regularization effect on the labeled sequences, yielding a better generalized model. This allows the network to learn representations that are useful for not only slot tagging using labeled data but also learning dependencies both within and between latent clusters of unseen words. The proposed pre-training method with the CRF2CNN architecture achieves significant gains with respect to the strongest semi-supervised baseline. Asli Celikyilmaz, Ruhi Sarikaya, Dilek Hakkani-Tür, Nikhil Ramesh, Gökhan Tür |
INTERSPEECH | 3 |
| 2016 | End-to-End Memory Networks with Knowledge Carryover for Multi-Turn Spoken Language UnderstandingabstractSpoken language understanding (SLU) is a core component of a spoken dialogue system. In the traditional architecture of dialogue systems, the SLU component treats each utterance independent of each other, and then the following components aggregate the multi-turn information in the separate phases. However, there are two challenges: 1) errors from previous turns may be propagated and then degrade the performance of the current turn; 2) knowledge mentioned in the long history may not be carried into the current turn. This paper addresses the above issues by proposing an architecture using end-to-end memory networks to model knowledge carryover in multi-turn conversations, where utterances encoded with intents and slots can be stored as embeddings in the memory and the decoding phase applies an attention model to leverage previously stored semantics for intent prediction and slot tagging simultaneously. The experiments on Microsoft Cortana conversational data show that the proposed memory network architecture can effectively extract salient semantics for modeling knowledge carryover in the multi-turn conversations and outperform the results using the state-of-the-art recurrent neural network framework (RNN) designed for single-turn SLU. Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür, Jianfeng Gao 0001, Li Deng 0001 |
INTERSPEECH | 2 |
| 2016 | Multi-Domain Joint Semantic Frame Parsing Using Bi-Directional RNN-LSTMabstractSequence-to-sequence deep learning has recently emerged as a new paradigm in supervised learning for spoken language understanding. However, most of the previous studies explored this framework for building single domain models for each task, such as slot filling or domain classification, comparing deep learning based approaches with conventional ones like conditional random fields. This paper proposes a holistic multi-domain, multi-task (i.e. slot filling, domain and intent detection) modeling approach to estimate complete semantic frames for all user utterances addressed to a conversational system, demonstrating the distinctive power of deep learning methods, namely bi-directional recurrent neural network (RNN) with long-short term memory (LSTM) cells (RNN-LSTM) to handle such complexity. The contributions of the presented work are three-fold: (i) we propose an RNN-LSTM architecture for joint modeling of slot filling, intent determination, and domain classification; (ii) we build a joint multi-domain model enabling multi-task deep learning where the data from each domain reinforces each other; (iii) we investigate alternative architectures for modeling lexical context in spoken language understanding. In addition to the simplicity of the single model framework, experimental results show the power of such an approach on Microsoft Cortana real user data over alternative methods based on single domain/task deep learning. Dilek Hakkani-Tür, Gökhan Tür, Asli Celikyilmaz, Yun-Nung Chen, Jianfeng Gao 0001, Li Deng 0001, Ye-Yi Wang |
INTERSPEECH | 1 |
| 2016 | AIMU: Actionable Items for Meeting Understanding
Yun-Nung Chen, Dilek Hakkani-Tür |
LREC | 2 |
| 2016 | Syntax or semantics? knowledge-guided joint semantic frame parsingabstractSpoken language understanding (SLU) is a core component of a spoken dialogue system, which involves intent prediction and slot filling and also called semantic frame parsing. Recently recurrent neural networks (RNN) obtained strong results on SLU due to their superior ability of preserving sequential information over time. Traditionally, the SLU component parses semantic frames for utterances considering their flat structures, as the underlying RNN structure is a linear chain. However, natural language exhibits linguistic properties that provide rich, structured information for better understanding. This paper proposes to apply knowledge-guided structural attention networks (K-SAN), which additionally incorporate non-flat network topologies guided by prior knowledge, to a language understanding task. The model can effectively figure out the salient substructures that are essential to parse the given utterance into its semantic frame with an attention mechanism, where two types of knowledge, syntax and semantics, are utilized. The experiments on the benchmark Air Travel Information System (ATIS) data and the conversational assistant Cortana data show that 1) the proposed K-SAN models with syntax or semantics outperform the state-of-the-art neural network based results, and 2) the improvement for joint semantic frame parsing is more significant, because the structured information provides rich cues for sentence-level understanding, where intent prediction and slot filling can be mutually improved. Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür, Asli Celikyilmaz, Jianfeng Gao 0001, Li Deng 0001 |
SLT | 2 |
| 2015 | A universal model for flexible item selection in conversational dialogsabstractHuman-computer interaction and statistical natural language understanding has changed with the addition of a visual display screen in modern mobile devices, as visual rendering is used to communicate the dialog system's response. Onscreen item identification and resolution when interpreting the user utterances is one critical problem to achieve the natural and accurate human-machine communication. This problem, also called Flexible Item Selection (FIS), has been posed as a classification task to correctly identify intended on-screen item(s) from user utterances. This paper presents a universal FIS model that can be applied to dialog systems developed in different languages. We design a set of input features for the FIS model that makes it largely language-independent. We demonstrate that a single universal FIS model can be used in place of language specific FIS models with no loss in accuracy. We also show that such a model can generalize well to new unseen languages with minimal loss in accuracy on held out languages including English, French, Spanish, Italian, German, and Chinese. Eliminating the need for building and maintaining a separate FIS model for each new language, the universal FIS model helps scaling an existing dialogue system to new languages faster at a lower development cost. Asli Celikyilmaz, Zhaleh Feizollahi, Dilek Hakkani-Tür, Ruhi Sarikaya |
ASRU | 3 |
| 2015 | Detecting actionable items in meetings by convolutional deep structured semantic modelsabstractThe recent success of voice interaction with smart devices (human-machine genre) and improvements in speech recognition for conversational speech show the possibility of conversation-related applications. This paper investigates the task of actionable item detection in meetings (human-human genre), where the intelligent assistant dynamically provides the participants access to information (e.g. scheduling a meeting, taking notes) without interrupting the meetings. A convolutional deep structured semantic model (CDSSM) is applied to learn the latent semantics for human actions and utterances from human-machine (source genre) and human-human (target) interactions. Furthermore, considering the mismatch between source and target genre and scarcity of annotated data sets for the target genre, we develop adaptation techniques that adjust the learned embeddings to better fit the target genre. Experiments show that CDSSM performs better for actionable item detection compared to baselines using lexical features (27.5% relative) and other semantic features (15.9% relative) when the source genre and target genre match with each other. When the target genre mismatches with the source genre, our proposed adaptation techniques further improve the performance. The discussion and analysis of the experiments provide a reasonable direction for such an actionable item detection task1. Yun-Nung Chen, Dilek Hakkani-Tür, Xiaodong He 0001 |
ASRU | 2 |
| 2015 | Investigation of ensemble models for sequence learningabstractWhile ensemble models have proven useful for sequence learning tasks there is relatively fewer work that provide insights into what makes them powerful. In this paper, we investigate the empirical behavior of the ensemble approaches on sequence modeling, specifically for the semantic tagging task. We explore this by comparing the performance of commonly used and easy to implement ensemble methods such as majority voting, linear combination and stacking to a learning based and rather complex ensemble method. Next, we ask the question: when models of different learning methods such as predictive and representation learning (e.g., deep learning) are aggregated, do we get performance gains over the individual baseline models. We explore these questions on a range of datasets on syntactic and semantic tagging tasks such as slot filling. Our findings show that a ranking based ensemble model outperforms all other well-known ensemble models. Asli Celikyilmaz, Dilek Hakkani-Tür |
ICASSP | 2 |
| 2015 | Probabilistic features for connecting eye gaze to spoken language understandingabstractMany users obtain content from a screen and want to make requests of a system based on items that they have seen. Eye-gaze information is a valuable signal in speech recognition and spoken-language understanding (SLU) because it provides context for a user's next utterance-what the user says next is probably conditioned on what they have seen. This paper investigates three types of features for connecting eye-gaze information to an SLU system: lexical, and two types of eye-gaze features. These features help us to understand which object (i.e. a link) that a user is referring to on a screen. We show a 17% absolute performance improvement in the referenced-object F-score by adding eye-gaze features to conventional methods based on a lexical comparison of the spoken utterance and the text on the screen. Anna Prokofieva, Malcolm Slaney, Dilek Hakkani-Tür |
ICASSP | 3 |
| 2015 | Clustering novel intents in a conversational interaction system with semantic parsingabstractSpoken language understanding (SLU) in today’s conversational systems focuses on recognizing a set of domains, intents, and associated arguments, that are determined by application developers. User requests that are not covered by these are usually directed to search engines, and may remain unhandled. We propose a method that aims to find common user intents amongst these uncovered, out-of-domain utterances, with the goal of supporting future phases of dialog system design. Our approach relies on finding common semantic patterns in uncovered user utterances using an Abstract Meaning Representation based semantic parser. We represent the corpus as a graph and find subgraphs that represent clusters, by pruning the corpus graph according to frequency and entropy. We employ crowd-workers to select and label the resulting clusters and compare resulting clusters with two baselines. Experimental analyses show that we obtain higher coverage and accuracy with the semantic parsing based clustering method. Furthermore, since the intents and candidate slots are already induced, these utterances can also be used in unsupervised SLU modeling. In intent classification experiments, we show that the statistical model trained using the clusters formed by this approach results in higher classification F-measure (showing about 25% relative improvement) in comparison to the alternatives. Dilek Hakkani-Tür, Yun-Cheng Ju, Geoffrey Zweig, Gökhan Tür |
INTERSPEECH | 1 |
| 2015 | Unsupervised relation detection using automatic alignment of query patterns extracted from knowledge graphs and query click logsabstractTraditional methods for building spoken language understanding systems require manual rules or annotated data, which are expensive. In this work, we present an unsupervised method for bootstrapping a relation classifier, which identifies the knowledge graph relations present in an input query. Unlike existing work, we utilize only one knowledge graph entity instead of two for mining relevant query patterns from query click logs. As a result, the mined patterns can be used to infer both explicit relations (where the objects of the relations are expressed in the queries) and implicit relations (where the objects of the relations are being asked about). Using only the mined queries, the final classifier achieves an F-measure of 55.5%, which is significantly higher than the previous unsupervised learning baselines. Panupong Pasupat, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2015 | Keynote: Graph-based Approaches for Spoken Language UnderstandingabstractFollowing an upsurge in mobile device usage and improvements in speech recognition performance, multiple virtual personal assistant systems have emerged, and have been widely adopted by users.While these assistants proved to be beneficial, their usage has been limited to certain scenarios and domains, with underlying language understanding models that have been finely tuned by their builders.Simultaneously, there have been several recent advancements in semantic web knowledge graphs especially used for basic question answering, efforts for integrating statistical information on these graphs, graph-based generic semantic representations and parsers, allproviding opportunities for open domain spoken language understanding.In this talk, I plan to summarize recent work in these areas, focusing on their connection, as a promise for wide coverage spoken language understanding in conversational systems, while at the same time investigating what is still lacking for natural human-machine interactions and related challenges. Dilek Hakkani-Tür |
SIGDIAL Conference | 1 |
| 2015 | Using Recurrent Neural Networks for Slot Filling in Spoken Language UnderstandingabstractSemantic slot filling is one of the most challenging problems in spoken language understanding (SLU). In this paper, we propose to use recurrent neural networks (RNNs) for this task, and present several novel architectures designed to efficiently model past and future temporal dependencies. Specifically, we implemented and compared several important RNN architectures, including Elman, Jordan, and hybrid variants. To facilitate reproducibility, we implemented these networks with the publicly available Theano neural network toolkit and completed experiments on the well-known airline travel information system (ATIS) benchmark. In addition, we compared the approaches on two custom SLU data sets from the entertainment and movies domains. Our results show that the RNN-based models outperform the conditional random field (CRF) baseline by 2% in absolute error reduction on the ATIS benchmark. We improve the state-of-the-art by 0.5% in the Entertainment domain, and 6.7% for the movies domain. Grégoire Mesnil, Yann N. Dauphin, Kaisheng Yao, Yoshua Bengio, Li Deng 0001, Dilek Hakkani-Tür, Xiaodong He 0001, Larry Heck, Gökhan Tür, Dong Yu 0001, Geoffrey Zweig |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2014 | Resolving Referring Expressions in Conversational Dialogs for Natural User InterfacesabstractUnlike traditional over-the-phone spoken dialog systems (SDSs), modern dialog systems tend to have visual rendering on the device screen as an additional modality to communicate the system's response to the user.Visual display of the system's response not only changes human behavior when interacting with devices, but also creates new research areas in SDSs.Onscreen item identification and resolution in utterances is one critical problem to achieve a natural and accurate humanmachine communication.We pose the problem as a classification task to correctly identify intended on-screen item(s) from user utterances.Using syntactic, semantic as well as context features from the display screen, our model can resolve different types of referring expressions with up to 90% accuracy.In the experiments we also show that the proposed model is robust to domain and screen layout changes. Asli Celikyilmaz, Zhaleh Feizollahi, Dilek Hakkani-Tür, Ruhi Sarikaya |
EMNLP | 3 |
| 2014 | Extending domain coverage of language understanding systems via intent transfer between domains using knowledge graphs and search query click logsabstractThis paper proposes a new technique to enable Natural Language Understanding (NLU) systems to handle user queries beyond their original semantic schemas defined by intents and slots. Knowledge graph and search query logs are used to extend NLU system's coverage by transferring intents from other domains to a given domain. The transferred intents as well as existing intents are then applied to a set of new slots that they are not trained with. The knowledge graph and search click logs are used to determine whether the new slots (i.e. entities) or their attributes in the graph can be used together with transfered intents without re-training the underlying NLU models with the expanded (i.e. with new intents and slots) schema. Experimental results show that the proposed technique can in fact be used in extending NLU system's domain coverage in fulfilling the user's request. Ali El-Kahky, Ruhi Sarikaya, Gökhan Tür, Dilek Hakkani-Tür, Larry Heck |
ICASSP | 5 |
| 2014 | A variational Bayesian model for user intent detectionabstractIntent detectors in state-of-the-art spoken language understanding systems are often trained with a small number of manually annotated examples collected from the application domain. Search query logs provide a large number of unlabeled queries that would be beneficial to improve such supervised classification. Furthermore, the contents of user queries as well as the clicked URLs provide information about user's intent. In this paper, we propose a variational Bayesian approach for modeling latent intents of user queries and clicked URLs when available. We use this model to enhance supervised intent classification of user queries from conversational interactions. Experiments were run with large volumes of search queries and show significant improvements over state-of-the-art systems. Yangfeng Ji, Dilek Hakkani-Tür, Asli Celikyilmaz, Larry Heck, Gökhan Tür |
ICASSP | 2 |
| 2014 | Automatic characterization of speaking styles in educational videosabstractRecent studies have shown the importance of using online videos along with textual material in educational instruction, especially for better content retention and improved concept understanding. A key question is how to select videos to maximize student engagement, particularly when there are multiple possible videos on the same topic. While there are many aspects that drive student engagement, in this paper we focus on presenter speaking styles in the video. We use crowd-sourcing to explore speaking style dimensions in online educational videos, and identify six broad dimensions: liveliness, speaking rate, pleasantness, clarity, formality and confidence. We then propose techniques based solely on acoustic features for automatically identifying a subset of the dimensions. Finally, we perform video re-ranking experiments to learn how users apply their speaking style preferences to augment textbook material. Our findings also indicate how certain dimensions are correlated with perceptions of general pleasantness of the voice. Soroosh Mariooryad, Anitha Kannan, Dilek Hakkani-Tür, Elizabeth Shriberg |
ICASSP | 3 |
| 2014 | Leveraging semantic web search and browse sessions for multi-turn spoken dialog systemsabstractTraining statistical dialog models in spoken dialog systems (SDS) requires large amounts of annotated data. The lack of scalable methods for data mining and annotation poses a significant hurdle for state-of-the-art statistical dialog managers. This paper presents an approach that directly leverage billions of web search and browse sessions to overcome this hurdle. The key insight is that task completion through web search and browse sessions is (a) predictable and (b) generalizes to spoken dialog task completion. The new method automatically mines behavioral search and browse patterns from web logs and translates them into spoken dialog models. We experiment with naturally occurring spoken dialogs and large scale web logs. Our session-based models outperform the state-of-the-art method for entity extraction task in SDS. We also achieve better performance for both entity and relation extraction on web search queries when compared with nontrivial baselines. Lu Wang 0008, Larry Heck, Dilek Hakkani-Tür |
ICASSP | 3 |
| 2014 | Eye Gaze for Spoken Language Understanding in Multi-modal Conversational InteractionsabstractWhen humans converse with each other, they naturally amalgamate information from multiple modalities (i.e., speech, gestures, speech prosody, facial expressions, and eye gaze). This paper focuses on eye gaze and its combination with speech. We develop a model that resolves references to visual (screen) elements in a conversational web browsing system. The system detects eye gaze, recognizes speech, and then interprets the user's browsing intent (e.g., click on a specific element) through a combination of spoken language understanding and eye gaze tracking. We experiment with multi-turn interactions collected in a wizard-of-Oz scenario where users are asked to perform several web-browsing tasks. We compare several gaze features and evaluate their effectiveness when combined with speech-based lexical features. The resulting multi-modal system not only increases user intent (turn) accuracy by 17%, but also resolves the referring expression ambiguity commonly observed in dialog systems with a 10% increase in F-measure. Dilek Hakkani-Tür, Malcolm Slaney, Asli Celikyilmaz, Larry Heck |
ICMI | 1 |
| 2014 | The Relation of Eye Gaze and Face Pose: Potential Impact on Speech RecognitionabstractWe are interested in using context to improve speech recognition and speech understanding. Knowing what the user is attending to visually helps us predict their utterances and thus makes speech recognition easier. Eye gaze is one way to access this signal, but is often unavailable (or expensive to gather) at longer distances. In this paper we look at joint eye-gaze and facial-pose information while users perform a speech reading task. We hypothesize, and verify experimentally, that the eyes lead, and then the face follows. Face pose might not be as fast, or as accurate a signal of visual attention as eye gaze, but based on experiments correlating eye gaze with speech recognition, we conclude that face pose provides useful information to bias a recognizer toward higher accuracy. Malcolm Slaney, Andreas Stolcke, Dilek Hakkani-Tür |
ICMI | 3 |
| 2014 | Rapidly building domain-specific entity-centric language models using semantic web knowledge sourcesabstractFor domain-specific speech recognition tasks, it is best if the statistical language model component is trained with text data that is content-wise and style-wise similar to the targeted domain for which the application is built. For state-of-the-art language modeling techniques that can be used in real-time within speech recognition engines during first-pass decoding (e.g., N-gram models), the above constraints have to be fulfilled in the training data. However collecting such data, even through crowd sourcing, is expensive and time consuming, and can still be not representative of how a much larger user population would interact with the recognition system. In this paper, we address this problem by employing several semantic web sources that already contain the domain-specific knowledge, such as query click logs and knowledge graphs. We build statistical language models that meet the requirements listed above for domain-specific recognition tasks where natural language is used and the user queries are about name entities in a specific domain. As a case study, in the movies domain where users’ voice queries are movie related, compared to a generic web language model, a language model trained with the above resources not only yields significant perplexity and word-errorrate improvements, but also presents an approach where such language models can be rapidly developed for other domains. Murat Akbacak, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 2 |
| 2014 | Probabilistic enrichment of knowledge graph entities for relation detection in conversational understandingabstractKnowledge encoded in semantic graphs such as Freebase has been shown to benefit semantic parsing and interpretation of natural language user utterances. In this paper, we propose new methods to assign weights to semantic graphs that reflect common usage types of the entities and their relations. Such statistical information can improve the disambiguation of entities in natural language utterances. Weights for entity types can be derived from the populated knowledge in the semantic graph, based on the frequency of occurrence of each type. They can also be learned from the usage frequencies in real world natural language text, such as related Wikipedia documents or user queries posed to a search engine. We compare the proposed methods with the unweighted version of the semantic knowledge graph for the relation detection task and show that all weighting methods result in better performance in comparison to using the unweighted version. Dilek Hakkani-Tür, Asli Celikyilmaz, Larry Heck, Gökhan Tür, Geoffrey Zweig |
INTERSPEECH | 1 |
| 2014 | Segmentation and disfluency removal for conversational speech translationabstractIn this paper we focus on the effect of on-line speech segmentation and disfluency removal methods on conversational speech translation. In a real-time conversational speech to speech translation system, on-line segmentation of speech is required to avoid latency beyond few seconds. While sentential unit segmentation and disfluency removal have been heavily studied mainly for off-line speech processing, to the best of our knowledge, the combined effect of these tasks on conversational speech translation has not been investigated. Furthermore, optimization of performance given maximum allowable system latency to enable a conversation is a newer problem for these tasks. We show that the conventional assumption of doing segmentation followed by disfluency removal is not the best practice. We propose a new approach to do simple-disfluency removal followed by segmentation and then by complex-disfluency removal. The proposed approach shows a significant gain on translation performance of up to 3 Bleu points with only 6 second latency to look ahead, using state-ofthe art machine translation and speech recognition systems. Index Terms: speech translation, disfluency removal, segmentation, sentence units, speech processing Hany Hassan, Lee Schwartz, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 3 |
| 2014 | Detecting out-of-domain utterances addressed to a virtual personal assistantabstractConversational understanding systems, especially virtual personal assistants (VPAs), perform “targeted” natural language understanding, assuming their users stay within the walled gardens of covered domains, and back-off to generic web search otherwise. However, users usually do not know the concept of domains and sometimes simply do not distinguish the system from simple voice search. Hence it becomes an important problem to identify these rejected out-of-domain utterances which are actually intended for the VPA. This paper presents a study tackling this new task, showing that how one utters a request is more important for this task than what is uttered, resembling addressee detection or dialog act tagging. To this end, syntactic and semantic parse “structure” features are extracted in addition to lexical features to train a binary SVM classifier using a large number of random web search queries and VPA utterances from multiple domains. We present controlled experiments leaving one domain out and check the precision of the model when combined with unseen queries. Our results indicate that such structured features result in higher precision especially when the test domain bears little resemblance to the existing domains. Gökhan Tür, Anoop Deoras, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2014 | Deriving local relational surface forms from dependency-based entity embeddings for unsupervised spoken language understandingabstractRecent works showed the trend of leveraging web-scaled structured semantic knowledge resources such as Freebase for open domain spoken language understanding (SLU). Knowledge graphs provide sufficient but ambiguous relations for the same entity, which can be used as statistical background knowledge to infer possible relations for interpretation of user utterances. This paper proposes an approach to capture the relational surface forms by mapping dependency-based contexts of entities from the text domain to the spoken domain. Relational surface forms are learned from dependency-based entity embeddings, which encode the contexts of entities from dependency trees in a deep learning model. The derived surface forms carry functional dependency to the entities and convey the explicit expression of relations. The experiments demonstrate the efficiency of leveraging derived relational surface forms as local cues together with prior background knowledge. Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 2 |
| 2014 | Entity ranking for descriptive queriesabstractWe investigate the problem of entity ranking towards descriptive queries, that aims to match entities referred in user queries to entities of a large knowledge base (KB). Entity ranking faces the primary challenge of the sparseness of entity related data, such as various ways of referring to an entity. The lack of sufficient variations of entity referring expressions in KB makes it difficult to find entities referred in user queries, especially when the queries are descriptive. We tackle this problem by enriching KB entries using web documents and query click logs. First, we propose a novel method of injecting textual information from web documents to the KB on a large scale. Since the number of web documents can be large, we propose to use keyword extraction and summarization techniques for compactly representing entity-related information. Second, we mine web search query logs to link entities to existing queries. Experiments show significant improvements after the KB enrichment, compared with two competitive baselines. We also achieve further improvements by combining the data from these two resources. Kai Hong, Pengjun Pei, Ye-Yi Wang, Dilek Hakkani-Tür |
SLT | 4 |
| 2014 | Personal knowledge graph population from user utterances in conversational understandingabstractKnowledge graphs provide a powerful representation of entities and the relationships between them, but automatically constructing such graphs from spoken language utterances presents the novelty and numerous challenges. In this paper, we introduce a statistical language understanding approach to automatically construct personal (user-centric) knowledge graphs in conversational dialogs. Such information has the potential to better understand the users' requests, fulfilling them, and enabling other technologies such as developing better inferences or proactive interactions. Knowledge encoded in semantic graphs such as Freebase has been shown to benefit semantic parsing and interpretation of natural language utterances. Hence, as a first step, we exploit the personal factual relation triples from Freebase to mine natural language snippets with a search engine, and the resulting snippets containing pairs of related entities to create the training data. This data is then used to build three key language understanding components: (1) Personal Assertion Classification identifies the user utterances that are relevant with personal facts, e.g., “my mother's name is Rosa”; (2) Relation Detection classifies the personal assertion utterance into one of the predefined relation classes, e.g., “parents”; and (3) Slot Filling labels the attributes or arguments of relations, e.g., “name(parents): Rosa”. Our experiments using the Microsoft conversational understanding system demonstrate the performance of this proposed approach on the population of personal knowledge graphs. Xiang Li 0066, Gökhan Tür, Dilek Hakkani-Tür, Qi Li 0014 |
SLT | 3 |
| 2014 | Distributed open-domain conversational understanding framework with domain independent extractorsabstractTraditional spoken dialog systems are usually based on a centralized architecture, in which the number of domains is predefined, and the provider is fixed for a given domain and intent. The spoken language understanding (SLU) component is responsible for detecting domain and intents, and filling domain-specific slots. It is expensive and time-consuming in this architecture to add new and/or competing domains, intents, or providers. The rapid growth of service providers in the mobile computing market calls for an extensible dialog system framework. This paper presents a distributed dialog infrastructure where each domain or provider is agnostic of others, and processes the user utterances independently using their own knowledge or models, so that a new domain and new provider can be easily incorporated in. In addition, to facilitate each service provider building their own SLU models or algorithms, we introduce a new component, extractors, to provide intermediate semantic annotations such as entity mention tags, which can be plugged in arbitrarily as well. Each service provider can then rapidly develop their SLU parser with minimum efforts by providing some example sentences with intents and slots if needed. Our preliminary experimental results demonstrate the power of this new framework compared to a centralized architecture. Qi Li 0014, Gökhan Tür, Dilek Hakkani-Tür, Xiang Li 0066, Tim Paek, Asela Gunawardana, Chris Quirk |
SLT | 3 |
| 2014 | Eye gaze for understanding conversational speechabstractEye gaze is a useful indication of attention and, as such, can be a valuable feature to improve spoken-language understanding in human-computer interaction. Based on the hypothesis that users look at a link before selecting it, we investigate the use of novel eye-gaze features to improve link click event prediction. Our data comprises users performing a variety of online tasks such as form filling and web browsing, and we show significant performance improvement by incorporating the use of gaze features. In addition, our analysis shows that there is much user-specific variation in gaze, so we are also looking to improve the modeling of gaze by user- and task-specific adaptation. Anna Prokofieva, Dilek Hakkani-Tür, Malcolm Slaney |
SLT | 2 |
| 2013 | Semi-Supervised Semantic Tagging of Conversational Understanding using Markov Topic Regression
Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür, Ruhi Sarikaya |
ACL (1) | 2 |
| 2013 | Easy contextual intent prediction and slot detectionabstractSpoken language understanding (SLU) is one of the main tasks of a dialog system, aiming to identify semantic components in user utterances. In this paper, we investigate the incorporation of context into the SLU tasks of intent prediction and slot detection. Using a corpus that contains session-level information, including the start and end of a session and the sequence of utterances within it, we experiment with the incorporation of information from previous intra-session utterances into the SLU tasks on a given utterance. For slot detection, we find that including features indicating the slots appearing in the previous utterances gives no significant increase in performance. In contrast, for intent prediction we find that a similar approach that incorporates the intent of the previous utterance as a feature yields relative error rate reductions of 6.7% on transcribed data and 8.7% on automatically-recognized data. We also find similar gains when treating intent prediction of utterance sequences as a sequential tagging problem via SVM-HMMs. Aditya Bhargava, Asli Celikyilmaz, Dilek Hakkani-Tür, Ruhi Sarikaya |
ICASSP | 3 |
| 2013 | Using a knowledge graph and query click logs for unsupervised learning of relation detectionabstractIn this paper, we introduce a novel statistical language understanding paradigm inspired by the emerging semantic web: Instead of building models for the target application, we propose relying on the semantic space already defined and populated in the knowledge graph for the target domain. As a first step towards this direction, we present unsupervised methods for training relation detection models exploiting the semantic knowledge graphs of the semantic web. The detected relations are used to mine natural language queries against a back-end knowledge base. For each relation, we leverage the complete set of entities that are connected to each other in the graph with the specific relation, and search these entity pairs on the web. We use the snippets that the search engine returns to create natural language examples that can be used as the training data for each relation. We further refine the annotations of these examples using the knowledge graph itself and iterate using a bootstrap approach. Furthermore, we explot the URLs returned for these pairs by the search engine to mine additional examples from the search engine query click logs. In our experiments, we show that, we can achieve relation detection models that perform about 60% macro F-measure on the relations that are in the knowledge graph without any manual labeling, resulting in a comparable performance with supervised training. Dilek Hakkani-Tür, Larry Heck, Gökhan Tür |
ICASSP | 1 |
| 2013 | Multi-style adaptive training for robust cross-lingual spoken language understandingabstractGiven the increasingly available machine translation (MT) services nowadays, one efficient strategy for cross-lingual spoken language understanding (SLU) is to first translate the input utterance from the second language into the primary language, and then call the primary language SLU system to decode the semantic knowledge. However, errors introduced in the MT process create a condition similar to the “mismatch” condition encountered in robust speech recognition. Such mismatch makes the performance of cross-lingual SLU far from acceptable. Motivated by successful solutions developed in robust speech recognition, we in this paper propose a multi-style adaptive training method to improve the robustness of the SLU system for cross-lingual SLU tasks. For evaluation, we created an English-Chinese bilingual ATIS database, and then carried out a series of experiments on that database to experimentally assess the proposed methods. Experimental results show that, without relying on any data in the second language, the proposed method significantly improves the performance on a cross-lingual SLU task while producing no degradation for input in the primary language. This greatly facilitates porting SLU to as many languages as there are MT systems without any human effort. We further study the robustness of this approach to another type of mismatch condition, caused by speech recognition errors, and demonstrate its success also. Xiaodong He 0001, Li Deng 0001, Dilek Hakkani-Tür, Gökhan Tür |
ICASSP | 3 |
| 2013 | Latent semantic modeling for slot filling in conversational understandingabstractIn this paper, we propose a new framework for semantic template filling in a conversational understanding (CU) system. Our method decomposes the task into two steps: latent n-gram clustering using a semi-supervised latent Dirichlet allocation (LDA) and sequence tagging for learning semantic structures in a CU system. Latent semantic modeling has been investigated to improve many natural language processing tasks such as syntactic parsing or topic tracking. However, due to several complexity problems caused by issues involving utterance length or dialog corpus size, it has not been analyzed directly for semantic parsing tasks. In this paper, we propose extending the LDA by introducing prior knowledge we obtain from semantic knowledge bases. Then, the topic posteriors obtained from the new LDA model are used as additional constraints to a sequence learning model for the semantic template filling task. The experimental results show significant performance gains on semantic slot filling models when features from latent semantic models are used in a conditional random field (CRF). Gökhan Tür, Asli Celikyilmaz, Dilek Hakkani-Tür |
ICASSP | 3 |
| 2013 | Understanding computer-directed utterances in multi-user dialog systemsabstractThis work aims to understand user requests when multiple users are interacting with each other and a spoken dialog system. More specifically, we explore the use of multi-human conversational context to improve domain detection in a human-computer interaction system. We investigate the different effects of human-directed context and computer-directed context, and compare the impact of using different context window sizes. Furthermore, we employ topic segmentation to chunk conversations for determining context boundaries. The experimental results show that the use of conversational context helps reduce domain detection error rate, especially in some specific domains. And though computer directed context is more reliable, the results show that the combination of both computer and human addressed utterances within a reasonable window size performs the best. Dilek Hakkani-Tür, Gökhan Tür |
ICASSP | 2 |
| 2013 | IsNL? a discriminative approach to detect natural language like queries for conversational understandingabstractWhile data-driven methods for spoken language understanding (SLU) provide state of the art performances and reduce maintenance and model adaptation costs compared to handcrafted parsers, the collection and annotation of domain-specific natural language utterances for training remains a time-consuming task. A recent line of research has focused on enriching the training data with in-domain utterances by mining search engine query logs to improve the SLU tasks. However genre mismatch is a big obstacle as search queries are typically keywords. In this paper, we present an efficient discriminative binary classification method that filters large collection of online web search queries only to select the natural language like queries. The training data used to build this classifier is mined from search query click logs, represented as a bipartite graph. Starting from queries which contain natural language salient phrases, random graph walk algorithms are employed to mine corresponding keyword queries. Then an active learning method is employed for quickly improving on top of this automatically mined data. The results show that our method is robust to noise in search queries by improving over a baseline model previously used for SLU data collection. We also show the effectiveness of detected natural language like queries in extrinsic evaluations on domain detection and slot filling tasks. Asli Celikyilmaz, Gökhan Tür, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2013 | A weakly-supervised approach for discovering new user intents from search query logsabstractState-of-the art spoken language understanding models that automatically capture user intents in human to machine dialogs are trained with manually annotated data, which is cumbersome and time-consuming to prepare. For bootstrapping the learning algorithm that detects relations in natural language queries to a conversational system, one can rely on publicly available knowledge graphs, such as Freebase, and mine corresponding data from the web. In this paper, we present an unsupervised approach to discover new user intents using a novel Bayesian hierarchical graphical model. Our model employs search query click logs to enrich the information extracted from bootstrapped models. We use the clicked URLs as implicit supervision and extend the knowledge graph based on the relational information discovered from this model. The posteriors from the graphical model relate the newly discovered intents with the search queries. These queries are then used as additional training examples to complement the bootstrapped relation detection models. The experimental results demonstrate the effectiveness of this approach, showing extended coverage to new intents without impacting the known intents. Index Terms: spoken language understanding, graphical models, search query click logs, intent discovery. Dilek Hakkani-Tür, Asli Celikyilmaz, Larry Heck, Gökhan Tür |
INTERSPEECH | 1 |
| 2013 | Leveraging knowledge graphs for web-scale unsupervised semantic parsingabstractThe past decade has seen the emergence of web-scale structured and linked semantic knowledge resources (e.g., Freebase, DB-Pedia). These semantic knowledge graphs provide a scalable “schema for the web”, representing a significant opportunity for the spoken language understanding (SLU) research community. This paper leverages these resources to bootstrap a web-scale semantic parser with no requirement for semantic schema de-sign, no data collection, and no manual annotations. Our ap-proach is based on an iterative graph crawl algorithm. From an initial seed node (entity-type), the method learns the related entity-types from the graph structure, and automatically anno-tates documents that can be linked to the node (e.g., Wikipedia articles, web search documents). Following the branches, the graph is crawled and the procedure is repeated. The resulting collection of annotated documents is used to bootstrap web-scale conditional random field (CRF) semantic parsers. Finally, we use a maximum-a-posteriori (MAP) unsupervised adapta-tion technique on sample data from a specific domain to refine the parsers. The scale of the unsupervised parsers is on the order of thousands of domains and entity-types, millions of entities, and hundreds of millions of relations. The precision-recall of the semantic parsers trained with our unsupervised method ap-proaches those trained with supervised annotations. Index Terms: semantic parsing, semantic web, semantic search, dialog, natural language understanding Larry Heck, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 2 |
| 2013 | Semantic parsing using word confusion networks with conditional random fieldsabstractA challenge in large vocabulary spoken language understand-ing (SLU) is robustness to automatic speech recognition (ASR) errors. The state of the art approaches for semantic parsing rely on using discriminative sequence classification methods, such as conditional random fields (CRFs). Most dialog systems em-ploy a cascaded approach where the best hypotheses from the ASR system are fed into the following SLU system. In our pre-vious work, we have proposed the use of lattices towards joint recognition and parsing. In this paper, extending this idea, we propose to exploit word confusion networks (WCNs), compiled from ASR lattices for both CRF modeling and decoding. WCNs provide a compact representation of multiple aligned ASR hy-potheses, without compromising recognition accuracy. For slot filling, we show significant semantic parsing performance im-provements using WCNs compared to ASR 1-best output, ap-proximating the oracle path performance. Index Terms: conditional random field, semantic parsing, word confusion network, natural language understanding Gökhan Tür, Anoop Deoras, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2013 | Joint Discriminative Decoding of Words and Semantic Tags for Spoken Language UnderstandingabstractMost Spoken Language Understanding (SLU) systems today employ a cascade approach, where the best hypothesis from Automatic Speech Recognizer (ASR) is fed into understanding modules such as slot sequence classifiers and intent detectors. The output of these modules is then further fed into downstream components such as interpreter and/or knowledge broker. These statistical models are usually trained individually to optimize the error rate of their respective output. In such approaches, errors from one module irreversibly propagates into other modules causing a serious degradation in the overall performance of the SLU system. Thus it is desirable to jointly optimize all the statistical models together. As a first step towards this, in this paper, we propose a joint decoding framework in which we predict the optimal word as well as slot sequence (semantic tag sequence) jointly given the input acoustic stream. Furthermore, the improved recognition output is then used for an utterance classification task, specifically, we focus on intent detection task. On a SLU task, we show 1.5% absolute reduction (7.6% relative reduction) in word error rate (WER) and 1.2% absolute improvement in F measure for slot prediction when compared to a very strong cascade baseline comprising of state-of-the-art large vocabulary ASR followed by conditional random field (CRF) based slot sequence tagger. Similarly, for intent detection, we show 1.2% absolute reduction (12% relative reduction) in classification error rate. Anoop Deoras, Gökhan Tür, Ruhi Sarikaya, Dilek Hakkani-Tür |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | A Joint Model for Discovery of Aspects in Utterances
Asli Celikyilmaz, Dilek Hakkani-Tür |
ACL (1) | 2 |
| 2012 | Translating natural language utterances to search queries for SLU domain detection using query click logsabstractLogs of user queries from a search engine (such as Bing or Google) together with the links clicked provide valuable implicit feedback to improve statistical spoken language understanding (SLU) models. However, the form of natural language utterances occurring in spoken interactions with a computer differs stylistically from that of keyword search queries. In this paper, we propose a machine translation approach to learn a mapping from natural language utterances to search queries. We train statistical translation models, using task and domain independent semantically equivalent natural language and keyword search query pairs mined from the search query click logs. We then extend our previous work on enriching the existing classification feature sets for input utterance domain detection with features computed using the click distribution over a set of clicked URLs from search engine query click logs of user utterances with automatically translated queries. This approach results in significant improvements for domain detection, especially when detecting the domains of user utterances that are formulated as natural language queries and effectively complements to the earlier work using syntactic transformations. Dilek Hakkani-Tür, Gökhan Tür, Rukmini Iyer, Larry Heck |
ICASSP | 1 |
| 2012 | Towards deeper understanding: Deep convex networks for semantic utterance classificationabstractFollowing the recent advances in deep learning techniques, in this paper, we present the application of special type of deep architecture - deep convex networks (DCNs) - for semantic utterance classification (SUC). DCNs are shown to have several advantages over deep belief networks (DBNs) including classification accuracy and training scalability. However, adoption of DCNs for SUC comes with non-trivial issues. Specifically, SUC has an extremely sparse input feature space encompassing a very large number of lexical and semantic features. This is about a few thousand times larger than the feature space for acoustic modeling, yet with a much smaller number of training samples. Experimental results we obtained on a domain classification task for spoken language understanding demonstrate the effectiveness of DCNs. The DCN-based method produces higher SUC accuracy than the Boosting-based discriminative classifier with word trigrams. Gökhan Tür, Li Deng 0001, Dilek Hakkani-Tür, Xiaodong He 0001 |
ICASSP | 3 |
| 2012 | Joint Decoding for Speech Recognition and Semantic TaggingabstractMost conversational understanding (CU) systems today employ a cascade approach, where the best hypothesis from automatic speech recognizer (ASR) is fed into spoken language understanding (SLU) module, whose best hypothesis is then fed into other systems such as interpreter or dialog manager. In such approaches, errors from one statistical module irreversibly propagates into another module causing a serious degradation in the overall performance of the conversational understanding system. Thus it is desirable to jointly optimize all the statistical modules together. As a first step towards this, in this paper, we propose a joint decoding framework in which we predict the optimal word as well as slot (semantic tag) sequence jointly given the input acoustic stream. On Microsoft’s CU system, we show 1.3 % absolute reduction in word error rate (WER) and 1.2% absolute improvement in F measure for slot prediction when compared to a very strong cascade baseline comprising of the state-of-the-art recognizer followed by a slot sequence tagger. Anoop Deoras, Ruhi Sarikaya, Gökhan Tür, Dilek Hakkani-Tür |
INTERSPEECH | 4 |
| 2012 | A Discriminative Classification-Based Approach to Information State Updates for a Multi-Domain Dialog SystemabstractWe propose a discriminative classification approach for updating the current information state of a multi-domain dialog system based on user responses. Our method uses a set of lexical and domain independent features to compare the spoken language understanding (SLU) output for the current user turn with the previous information state. We then update the information state accordingly, employing a discriminative machine learning approach. Using a data set collected from our conversational interaction system, we investigate the impact of features based on context dependent and context independent SLU tagging schemas. We show that the proposed approach outperforms two non-trivial baselines, one based on manually crafted rules and the other on classification with lexical features alone. Furthermore, such an approach allows the addition of new domains to the dialog manager in a seamless way. Dilek Hakkani-Tür, Gökhan Tür, Larry Heck, Ashley Fidler, Asli Celikyilmaz |
INTERSPEECH | 1 |
| 2012 | Learning When to Listen: Detecting System-Addressed Speech in Human-Human-Computer DialogabstractNew challenges arise for addressee detection when multiple people interact jointly with a spoken dialog system using unconstrained natural language. We study the problem of discriminating computer-directed from human-directed speech in a new corpus of human-human-computer (H-H-C) dialog, using lexical and prosodic features. The prosodic features use no word, context, or speaker information. Results with 19% WER speech recognition show improvements from lexical features (EER=23.1%) to prosodic features (EER=12.6%) to a combined model (EER=11.1%). Prosodic features also provide a 35% error reduction over a lexical model using true words (EER from 10.2% to 6.7%). Modeling energy contours with GMMs provides a particularly good prosodic model. While lexical models perform well for commands, they confuse free-form system-directed speech with human-human speech. Prosodic models dramatically reduce these confusions, implying that users change speaking style as they shift addressees (computer versus human) within a session. Overall results provide strong support for combining simple acoustic-prosodic models with lexical models to detect speaking style differences for this task. Elizabeth Shriberg, Andreas Stolcke, Dilek Hakkani-Tür, Larry Heck |
INTERSPEECH | 3 |
| 2012 | Exploiting the Semantic Web for Unsupervised Natural Language Semantic Parsing
Gökhan Tür, Minwoo Jeong, Ye-Yi Wang, Dilek Hakkani-Tür, Larry Heck |
INTERSPEECH | 4 |
| 2012 | Statistical semantic interpretation modeling for spoken language understanding with enriched semantic featuresabstractIn natural language human-machine statistical dialog systems, semantic interpretation is a key task typically performed following semantic parsing, and aims to extract canonical meaning representations of semantic components. In the literature, usually manually built rules are used for this task, even for implicitly mentioned non-named semantic components (like genre of a movie or price range of a restaurant). In this study, we present statistical methods for modeling interpretation, which can also benefit from semantic features extracted from large in-domain knowledge sources. We extract features from user utterances using a semantic parser and additional semantic features from textual sources (online reviews, synopses, etc.) using a novel tree clustering approach, to represent unstructured information that correspond to implicit semantic components related to targeted slots in the user's utterances. We evaluate our models on a virtual personal assistance system and demonstrate that our interpreter is effective in that it does not only improve the utterance interpretation in spoken dialog systems (reducing the interpretation error rate by 36% relative compared to a language model baseline), but also unveils hidden semantic units that are otherwise nearly impossible to extract from purely manual lexical features that are typically used in utterance interpretation. Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 2 |
| 2012 | Use of kernel deep convex networks and end-to-end learning for spoken language understandingabstractWe present our recent and ongoing work on applying deep learning techniques to spoken language understanding (SLU) problems. The previously developed deep convex network (DCN) is extended to its kernel version (K-DCN) where the number of hidden units in each DCN layer approaches infinity using the kernel trick. We report experimental results demonstrating dramatic error reduction achieved by the K-DCN over both the Boosting-based baseline and the DCN on a domain classification task of SLU, especially when a highly correlated set of features extracted from search query click logs are used. Not only can DCN and K-DCN be used as a domain or intent classifier for SLU, they can also be used as local, discriminative feature extractors for the slot filling task of SLU. The interface of K-DCN to slot filling systems via the softmax function is presented. Finally, we outline an end-to-end learning strategy for training the softmax parameters (and potentially all DCN and K-DCN parameters) where the learning objective can take any performance measure (e.g. the F-measure) for the full SLU system. Li Deng 0001, Gökhan Tür, Xiaodong He 0001, Dilek Hakkani-Tür |
SLT | 4 |
| 2012 | Exploiting the Semantic Web for unsupervised spoken language understandingabstractThis paper proposes an unsupervised training approach for SLU systems that leverages the structured semantic knowledge graphs of the emerging Semantic Web. The approach creates natural language surface forms of entity-relation-entity portions of knowledge graphs using a combination of web search retrieval and syntax-based dependency parsing. The new forms are used to train an SLU system in an unsupervised manner. This paper tests the approach on the problem of intent detection, and shows that the unsupervised training procedure matches the performance of supervised training over operating points important for commercial applications. Larry Heck, Dilek Hakkani-Tür |
SLT | 2 |
| 2011 | Discovery of Topically Coherent Sentences for Extractive Summarization
Asli Celikyilmaz, Dilek Hakkani-Tür |
ACL | 2 |
| 2011 | Exploiting distance based similarity in topic models for user intent detectionabstractOne of the main components of spoken language understanding is intent detection, which allows user goals to be identified. A challenging sub-task of intent detection is the identification of intent bearing phrases from a limited amount of training data, while maintaining the ability to generalize well. We present a new probabilistic topic model for jointly identifying semantic intents and common phrases in spoken language utterances. Our model jointly learns a set of intent dependent phrases and captures semantic intent clusters as distributions over these phrases based on a distance dependent sampling method. This sampling method uses proximity of words utterances when assigning words to latent topics. We evaluate our method on labeled utterances and present several examples of discovered semantic units. We demonstrate that our model outperforms standard topic models based on bag-of-words assumption. Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür, Ashley Fidler, Dustin Hillard |
ASRU | 2 |
| 2011 | Employing web search query click logs for multi-domain spoken language understandingabstractLogs of user queries from a search engine (such as Bing or Google) together with the links clicked provide valuable implicit feedback to improve statistical spoken language understanding (SLU) models. In this work, we propose to enrich the existing classification feature set for domain detection with features computed using the click distribution over a set of clicked URLs from search query click logs (QCLs) of user utterances. Since the form of natural language utterances differs stylistically from that of keyword search queries, to be able to match natural language utterances with related search queries, we perform a syntax-based transformation of the original utterances, after filtering out domain-independent salient phrases. This approach results in significant improvements for domain detection, especially when detecting the domains of web-related user utterances. Dilek Hakkani-Tür, Gökhan Tür, Larry Heck, Asli Celikyilmaz, Ashley Fidler, Dustin Hillard, Rukmini Iyer, Sarangarajan Parthasarathy |
ASRU | 1 |
| 2011 | Concept-based classification for multi-document summarizationabstractDocuments often contain inherently many concepts reflecting specific and generic aspects. To automatically generate a short summary text of documents on similar topics, it is imperative that we discover general aspects in documents be cause summaries usually contain general rather than specific concepts. This paper presents a semi-supervised extractive summarization model based upon latent concept classification that can differentiate between the two types of aspects as hidden concepts being mentioned in documents. A classifier is trained on hidden concepts discovered from documents and their corresponding human-generated summaries using a probabilistic Bayesian model: the summary-focused topic model. Experimental results based on ROUGE evaluations indicate that ranking sentences to be included in summary text based on the latent summary concept classification has improvements on the quality of the generated summaries. Asli Celikyilmaz, Dilek Hakkani-Tür |
ICASSP | 2 |
| 2011 | Exploiting query click logs for utterance domain detection in spoken language understandingabstractIn this paper, we describe methods to exploit search queries mined from search engine query logs to improve domain detection in spoken language understanding. We propose extending the label propagation algorithm, a graph-based semi-supervised learning approach, to incorporate noisy domain information estimated from search engine links the users click following their queries. The main contributions of our work are the use of search query logs for domain classification, integration of noisy supervision into the semi-supervised label propagation algorithm, and sampling of high-quality query click data by mining query logs and using classification confidence scores. We show that most semi-supervised learning methods we experimented with improve the performance of the supervised training, and the biggest improvement is achieved by label propagation that uses noisy supervision. We reduce the to error rate of domain detection by 20% relative, from 6.2% to 5.0%. Dilek Hakkani-Tür, Larry Heck, Gökhan Tür |
ICASSP | 1 |
| 2011 | Sentence simplification for spoken language understandingabstractIn this paper, we present a sentence simplification method and demonstrate its use to improve intent determination and slot filling tasks in spoken language understanding (SLU) systems. This research is motivated by the observation that, while current statistical SLU models usually perform accurately for simple, well-formed sentences, error rates increase for more complex, longer, more natural or spontaneous utterances. Furthermore, users familiar with web search usually formulate their information requests as a keyword search query, suggesting that frameworks which can handle both forms of inputs is required. We propose a dependency parsing-based sentence simplification approach that extracts a set of keywords from natural language sentences and uses those in addition to entire utterances for completing SLU tasks. We evaluated this approach using the well studied ATIS corpus with manual and automatic transcriptions and observed significant error reductions for both intent determination (30% relative) and slot filling (15% relative) tasks over the state-of the-art performances. Gökhan Tür, Dilek Hakkani-Tür, Larry Heck, Sarangarajan Parthasarathy |
ICASSP | 2 |
| 2011 | Approximate Inference for Domain Detection in Spoken Language UnderstandingabstractThis paper presents a semi-latent topic model for semantic domain detection in spoken language understanding systems. We use labeled utterance information to capture latent topics, which directly correspond to semantic domains. Additionally, we introduce an ’informative prior ’ for Bayesian inference that can simultaneously segment utterances of known domains into classes and divide them from out-of-domain utterances. We show that our model generalizes well on the task of classify-ing spoken language utterances and compare its results to those of an unsupervised topic model, which does not use labeled in-formation. Index Terms: spoken language understanding, generative mod-els, gibbs sampling. Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 2 |
| 2011 | Bootstrapping Domain Detection Using Query Click Logs for New DomainsabstractDomain detection in spoken dialog systems is usually treated as a multi-class, multi-label classification problem, and training of domain classifiers requires collection and manual annotation of example utterances. In order to extend a dialog system to new domains in a way that is seamless for users, domain detection should be able to handle utterances from the new domain as soon as it is introduced. In this work, we propose using web search query logs, which include queries entered by users and the links they subsequently click on, to bootstrap domain detection for new domains. While sampling user queries from the query click logs to train new domain classifiers, we introduce two types of measures based on the behavior of the users who entered a query and the form of the query. We show that both types of measures result in reductions in the error rate as compared to randomly sampling training queries. In controlled experiments over five domains, we achieve the best gain from the combination of the two types of sampling criteria. Dilek Hakkani-Tür, Gökhan Tür, Larry Heck, Elizabeth Shriberg |
INTERSPEECH | 1 |
| 2011 | Learning Weighted Entity Lists from Web Click Logs for Spoken Language UnderstandingabstractNamed entity lists provide important features for language understanding, but typical lists can contain many ambiguous or incorrect phrases. We present an approach for automatically learning weighted entity lists by mining user clicks from web search logs. The approach significantly outperforms multiple baseline approaches and the weighted lists improve spoken language understanding tasks such as domain detection and slot filling. Our methods are general and can be easily applied to large quantities of entities, across any number of lists. Index Terms: spoken language understanding, domain detection, slot filling, named entity lists, click logs Dustin Hillard, Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 3 |
| 2011 | Towards Unsupervised Spoken Language Understanding: Exploiting Query Click Logs for Slot FillingabstractIn this paper, we present a novel approach to exploit user queries mined from search engine query click logs to bootstrap or improve slot filling models for spoken language understanding. We propose extending the earlier gazetteer population techniques to mine unannotated training data for semantic parsing. The automatically annotated mined data can then be used to train slot specific parsing models. We show that this method can be used to bootstrap slot filling models and can be combined with any available annotated data to improve performance. Furthermore, this approach may eliminate the need for populating and maintaining in-domain gazetteers, in addition to providing complementary information if they are already available. Index Terms: spoken language understanding, slot filling, data mining, named entity extraction, unsupervised learning Gökhan Tür, Dilek Hakkani-Tür, Dustin Hillard, Asli Celikyilmaz |
INTERSPEECH | 2 |
| 2011 | Towards spoken clinical-question answering: evaluating and adapting automatic speech-recognition systems for spoken clinical questionsabstractOBJECTIVE: To evaluate existing automatic speech-recognition (ASR) systems to measure their performance in interpreting spoken clinical questions and to adapt one ASR system to improve its performance on this task. DESIGN AND MEASUREMENTS: The authors evaluated two well-known ASR systems on spoken clinical questions: Nuance Dragon (both generic and medical versions: Nuance Gen and Nuance Med) and the SRI Decipher (the generic version SRI Gen). The authors also explored language model adaptation using more than 4000 clinical questions to improve the SRI system's performance, and profile training to improve the performance of the Nuance Med system. The authors reported the results with the NIST standard word error rate (WER) and further analyzed error patterns at the semantic level. RESULTS: Nuance Gen and Med systems resulted in a WER of 68.1% and 67.4% respectively. The SRI Gen system performed better, attaining a WER of 41.5%. After domain adaptation with a language model, the performance of the SRI system improved 36% to a final WER of 26.7%. CONCLUSION: Without modification, two well-known ASR systems do not perform well in interpreting spoken clinical questions. With a simple domain adaptation, one of the ASR systems improved significantly on the clinical question task, indicating the importance of developing domain/genre-specific ASR systems. Gökhan Tür, Dilek Hakkani-Tür, Hong Yu 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2010 | A Hybrid Hierarchical Model for Multi-Document Summarization
Asli Celikyilmaz, Dilek Hakkani-Tür |
ACL | 2 |
| 2010 | Evaluation of semantic role labeling and dependency parsing of automatic speech recognition outputabstractSemantic role labeling (SRL) is an important module of spoken language understanding systems. This work extends the standard evaluation metrics for joint dependency parsing and SRL of text in order to be able to handle speech recognition output with word errors and sentence segmentation errors. We propose metrics based on word alignments and bags of relations, and compare their results on the output of several SRL systems on broadcast news and conversations of the OntoNotes corpus. We evaluate and analyze the relation between the performance of the subtasks that lead to SRL, including ASR, part-of-speech tagging or sentence segmentation. The tools are made available to the community. Benoît Favre, Bernd Bohnet, Dilek Hakkani-Tür |
ICASSP | 3 |
| 2010 | Summarization- and learning-based approaches to information distillationabstractInformation distillation is the task that aims to extract relevant passages of text from massive volumes of textual and audio sources, given a query. In this paper, we investigate two perspectives that use shallow language processing for answering open-ended distillation queries, such as “List me facts about [event]”. The first approach is a summarization-based approach that uses the unsupervised maximum marginal relevance (MMR) technique to successfully capture relevant but not redundant information. The second approach is based on supervised classification and trains support vector machines (SVMs) to discriminate relevant snippets from irrelevant snippets using a variety of features. Furthermore, we investigate the merit of using the ROUGE metric for its ability to evaluate redundancy alongside the conventionally used F-measure for evaluating distillation systems. Our experimental results with textual data indicate that SVM and MMR perform similarly in terms of ROUGE-2 scores while SVM is better than MMR in terms of F1 measure. Moreover, when speech recognizer output is used, SVM outperforms MMR in terms of both scores. Boriska Toth, Dilek Hakkani-Tür, Sibel Yaman |
ICASSP | 2 |
| 2010 | Extractive summarization using a latent variable model
Asli Celikyilmaz, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2010 | Speech-based automated cognitive status assessmentabstractVerbal interviews performed by trained clinicians are a common form of assessments to measure cognitive decline. The aim in this paper is to study the usability of automated methods for evaluating verbal cognitive status assessment tests for the elderly. If reliable, such methods for cognitive assessment can be used for frequent, non-intrusive, low-cost screenings and provide objective and longitudinal cognitive status monitoring data that can complement regular clinical visits and would be useful for early detection of conditions associated with language and communication impairments. This study focuses on two types of tests: a story-recall test, used for memory and language functioning assessment, and a picture description test, used to assess the information content in speech. A data collection was designed for this study involving recordings of about 100 people, mostly over 70 years old, performing these tests. The speech samples were manually transcribed and annotated with semantic units in order to obtain manual evaluation scores. We explore the use of automatic speech recognition and language processing methods to derive objective, automatically extracted metrics of cognitive status that are highly correlated with the manual scores. We use recall and precision based metrics based on semantic content units associated with the tests. Our experiments show high correlation between manually obtained scores and the automatic metrics obtained using either manual or automatic speech transcriptions. Index Terms: speech recognition, language processing, automated cognitive status assessment, elderly speech Dilek Hakkani-Tür, Dimitra Vergyri, Gökhan Tür |
INTERSPEECH | 1 |
| 2010 | Domain adaptation and compensation for emotion detectionabstractInspired by the recent improvements in domain adapta-tion and session variability compensation techniques used for speech and speaker processing, we study their effect for emo-tion prediction. More specifically, we investigated the use of publicly available out-of-domain data with emotion annotations for improving the performance of the in-domain model trained using 911 emergency-hotline calls. Following the emotion de-tection literature, we use prosodic (pitch, energy, and speaking rate) features as the inputs to a discriminative classifier. We performed segment-level n-fold cross validation emotion pre-diction experiments. Our results indicate significant improve-ment of performance for emotion prediction exploiting out-of-domain data. Index Terms: emotion detection, domain adaptation 1. Michelle Hewlett Sanchez, Gökhan Tür, Luciana Ferrer, Dilek Hakkani-Tür |
INTERSPEECH | 4 |
| 2010 | Social role discovery from spoken language using dynamic Bayesian networksabstractIn this paper, we focus on inferring social roles in con-versations using information extracted only from the speaking styles of the speakers. We model the turn-taking behavior of the speakers with dynamic Bayesian networks (DBNs), which provide the capability of naturally formulating the dependen-cies between random variables. More specifically, we first ex-plore the usefulness of a simple DBN, namely, a hidden Markov model (HMM), for this problem. As it turns out, the knowl-edge of the segments that belong to the same speaker can be augmented into this HMM structure, which results in a more sophisticated DBN. This information places a constraint on two subsequent speaker roles such that the current speaker role de-pends not only on the previous speaker’s role but also on that most recent role assigned to the same speaker. We conducted an experimental study to compare these two modeling approaches using broadcast shows. In our experiments, the approach with the constraint on same speaker segments assigned 89.5 % turns the correct role whereas the HMM-based approach assigned 79.2 % of turns their correct role. Index Terms: Social role discovery, speaker turn detection, spoken language understanding Sibel Yaman, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 2 |
| 2010 | Probabilistic model-based sentiment analysis of twitter messagesabstractWe present a machine learning approach to sentiment classification on twitter messages (tweets). We classify each tweet into two categories: polar and non-polar. Tweets with positive or negative sentiment are considered polar. They are considered non-polar otherwise. Sentiment analysis of tweets can potentially benefit different parties, such as consumers and marketing researchers, for obtaining opinions on different products and services. We present methods for text normalization of the noisy tweets and their classification with respect to the polarity. We experiment with a mixture model approach for generation of sentimental words, which are later used as indicator features of the classification model. Based on a gold standard manually annotated ensemble of tweets, with the new approach, we obtain F-scores that are relatively 10% better than a classification baseline that uses raw word n-gram features. Asli Celikyilmaz, Dilek Hakkani-Tür, Junlan Feng |
SLT | 2 |
| 2010 | What is left to be understood in ATIS?abstractOne of the main data resources used in many studies over the past two decades for spoken language understanding (SLU) research in spoken dialog systems is the airline travel information system (ATIS) corpus. Two primary tasks in SLU are intent determination (ID) and slot filling (SF). Recent studies reported error rates below 5% for both of these tasks employing discriminative machine learning techniques with the ATIS test set. While these low error rates may suggest that this task is close to being solved, further analysis reveals the continued utility of ATIS as a research corpus. In this paper, our goal is not experimenting with domain specific techniques or features which can help with the remaining SLU errors, but instead exploring methods to realize this utility via extensive error analysis. We conclude that even with such low error rates, ATIS test set still includes many unseen example categories and sequences, hence requires more data. Better yet, new annotated larger data sets from more complex tasks with realistic utterances can avoid over-tuning in terms of modeling and feature design. We believe that advancements in SLU can be achieved by having more naturally spoken data sets and employing more linguistically motivated features while preserving robustness due to speech recognition noise and variance due to natural language. Gökhan Tür, Dilek Hakkani-Tür, Larry Heck |
SLT | 2 |
| 2010 | Cascaded model adaptation for dialog act segmentation and tagging
Ümit Güz, Gökhan Tür, Dilek Hakkani-Tür, Sébastien Cuendet |
Comput. Speech Lang. | 3 |
| 2010 | Long story short - Global unsupervised models for keyphrase based meeting summarization
Korbinian Riedhammer, Benoît Favre, Dilek Hakkani-Tür |
Speech Commun. | 3 |
| 2010 | Multi-View Semi-Supervised Learning for Dialog Act Segmentation of SpeechabstractSentence segmentation of speech aims at determining sentence boundaries in a stream of words as output by the speech recognizer. Typically, statistical methods are used for sentence segmentation. However, they require significant amounts of labeled data, preparation of which is time-consuming, labor-intensive, and expensive. This work investigates the application of multi-view semi-supervised learning algorithms on the sentence boundary classification problem by using lexical and prosodic information. The aim is to find an effective semi-supervised machine learning strategy when only small sets of sentence boundary-labeled data are available. We especially focus on two semi-supervised learning approaches, namely, self-training and co-training. We also compare different example selection strategies for co-training, namely, agreement and disagreement. Furthermore, we propose another method, called self-combined, which is a combination of self-training and co-training. The experimental results obtained on the ICSI Meeting (MRDA) Corpus show that both multi-view methods outperform self-training, and the best results are obtained using co-training alone. This study shows that sentence segmentation is very appropriate for multi-view learning since the data sets can be represented by two disjoint and redundantly sufficient feature sets, namely, using lexical and prosodic information. Performance of the lexical and prosodic models is improved by 26% and 11% relative, respectively, when only a small set of manually labeled examples is used. When both information sources are combined, the semi-supervised learning methods improve the baseline F-Measure of 69.8% to 74.2%. Ümit Güz, Sébastien Cuendet, Dilek Hakkani-Tür, Gökhan Tür |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | The CALO Meeting Assistant SystemabstractThe CALO Meeting Assistant (MA) provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper presents the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, topic identification and segmentation, question-answer pair identification, action item recognition, decision extraction, and summarization. Gökhan Tür, Andreas Stolcke, L. Lynn Voss, Stanley Peters, Dilek Hakkani-Tür, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri |
IEEE Trans. Speech Audio Process. | 5 |
| 2009 | Who, What, When, Where, Why? Comparing Multiple Approaches to the Cross-Lingual 5W Task
Kristen Parton, Kathy McKeown, Bob Coyne, Mona T. Diab, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Heng Ji 0001, Wei-Yun Ma, Adam Meyers 0001, Sara Stolbach, Ang Sun, Gökhan Tür, Wei Xu 0004, Sibel Yaman |
ACL/IJCNLP | 6 |
| 2009 | Any questions? Automatic question detection in meetingsabstractIn this paper, we describe our efforts toward the automatic detection of English questions in meetings. We analyze the utility of various features for this task, originating from three distinct classes: lexico-syntactic, turn-related, and pitch-related. Of particular interest is the use of parse tree information in classification, an approach as yet unexplored. Results from experiments on the ICSI MRDA corpus demonstrate that lexico-syntactic features are most useful for this task, with turn-and pitch-related features providing complementary information in combination. In addition, experiments using reference parse trees on the broadcast conversation portion of the OntoNotes release 2.9 data set illustrate the potential of parse trees to outperform word lexical features. Kofi Boakye, Benoît Favre, Dilek Hakkani-Tür |
ASRU | 3 |
| 2009 | Integrating prosodic features in extractive meeting summarizationabstractSpeech contains additional information than text that can be valuable for automatic speech summarization. In this paper, we evaluate how to effectively use acoustic/prosodic features for extractive meeting summarization, and how to integrate prosodic features with lexical and structural information for further improvement. To properly represent prosodic features, we propose different normalization methods based on speaker, topic, or local context information. Our experimental results show that using only the prosodic features we achieve better performance than using the non-prosodic information on both the human transcripts and recognition output. In addition, a decision-level combination of the prosodic and non-prosodic features yields further gain, outperforming the individual models. Shasha Xie, Dilek Hakkani-Tür, Benoît Favre, Yang Liu 0004 |
ASRU | 2 |
| 2009 | Syntactically-informed models for comma predictionabstractProviding punctuation in speech transcripts not only improves readability, but it also helps downstream text processing such as information extraction or machine translation. In this paper, we improve by 7% the accuracy of comma prediction in English broadcast news by introducing syntactic features inspired by the role of commas as described in linguistics studies. We conduct an analysis of the impact of those features on other subsets of features (prosody, words...) when combined through CRFs. The syntactic cues can help characterizing large syntactic patterns such as appositions and lists which are not necessarily marked by prosody. Benoît Favre, Dilek Hakkani-Tür, Elizabeth Shriberg |
ICASSP | 2 |
| 2009 | A global optimization framework for meeting summarizationabstractWe introduce a model for extractive meeting summarization based on the hypothesis that utterances convey bits of information, or concepts. Using keyphrases as concepts weighted by frequency, and an integer linear program to determine the best set of utterances, that is, covering as many concepts as possible while satisfying a length constraint, we achieve ROUGE scores at least as good as a ROUGE-based oracle derived from human summaries. This brings us to a critical discussion of ROUGE and the future of extractive meeting summarization. Daniel Gillick, Korbinian Riedhammer, Benoît Favre, Dilek Hakkani-Tür |
ICASSP | 4 |
| 2009 | Towards automatic argument diagramming of multiparity meetingsabstractThis paper focuses on a lesser studied multiparty meetings processing task of argument diagramming. Argument diagramming aims at tagging the utterances and their relationships to represent the flow and structure of reasoning in conversations, especially in discussions and arguments. In this work, we tackle the problem of automatically assigning node types to user utterances using several lexical and prosodic features. We performed experiments using the AMI Meeting Corpus annotated according to the the Twente Argumentation Schema. Our results indicate that while lexical and prosodic features both provide orthogonal information for this task, using a cascaded approach, eliminating backchannel utterances improves the performance. With this final approach, when all features are used, we achieve about 9% relatively better error rates than a simpler classifier based on only lexical features. Dilek Hakkani-Tür |
ICASSP | 1 |
| 2009 | Phrase and word level strategies for detecting appositions in speechabstractAppositions are grammatical constructs in which two noun phrases are placed side-by-side, one modifying the other. Detecting them in speech can help extract semantic information useful, for instance, for co-reference resolution and question answering. We compare and combine three approaches: wordlevel and phrase-level classifiers, and a syntactic parser trained to generate appositions. On reference parses, the phrase-level classifier outperforms the other approaches while on automatic parses and ASR output, the combination of the appositiongenerating parser and the word-level classifier works best. An analysis of the system errors reveals that parsing accuracy and world knowledge are very important for this task. Benoît Favre, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2009 | Clusterrank: a graph based method for meeting summarizationabstractThis paper presents an unsupervised, graph based approach for extractive summarization of meetings. Graph based methods such as TextRank have been used for sentence extraction from news articles. These methods model text as a graph with sentences as nodes and edges based on word overlap. A sentence node is then ranked according to its similarity with other nodes. The spontaneous speech in meetings leads to incomplete, informed sentences with high redundancy and calls for additional measures to extract relevant sentences. We propose an extension of the TextRank algorithm that clusters the meeting utterances and uses these clusters to construct the graph. We evaluate this method on the AM I meeting corpus and show a significant improvement over TextRank and other baseline methods. Nikhil Garg 0005, Benoît Favre, Korbinian Riedhammer, Dilek Hakkani-Tür |
INTERSPEECH | 4 |
| 2009 | Leveraging sentence weights in a concept-based optimization framework for extractive meeting summarizationabstractInternational audience Shasha Xie, Benoît Favre, Dilek Hakkani-Tür, Yang Liu 0004 |
INTERSPEECH | 3 |
| 2009 | Combining semantic and syntactic information sources for 5-w question answeringabstractThis paper focuses on combining answers generated by a semantic parser that produces semantic role labels (SRLs) and those generated by syntactic parser that produces function tags for answering 5-W questions, i.e., who, what, when, where, and why. We take a probabilistic approach in which a system’s ability to correctly answer 5-W questions is measured with the likelihood that its answers are produced for the given word sequence. This is achieved by training statistical language models (LMs) that are used to predict whether the answers returned by semantic parse or those returned by the syntactic parser are more likely. We evaluated our approach using the OntoNotes dataset. Our experimental results indicate that the proposed LM-based combination strategy was able to improve the performance of the best individual system in terms of both F1 measure and accuracy. Furthermore, the error rates for each question type were also significantly reduced with the help of the proposed approach. Sibel Yaman, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 2 |
| 2009 | Classification-based strategies for combining multiple 5-w question answering systemsabstractWe describe and analyze inference strategies for combining outputs from multiple question answering systems each of which was developed independently. Specifically, we address the DARPA-funded GALE information distillation Year 3 task of finding answers to the 5-Wh questions (who, what, when, where, and why) for each given sentence. The approach we take revolves around determining the best system using discriminative learning. In particular, we train support vector machines with a set of novel features that encode systems’ capabilities of returning as many correct answers as possible. We analyze two combination strategies: one combines multiple systems at the granularity of sentences, and the other at the granularity of individual fields. Our experimental results indicate that the proposed features and combination strategies were able to improve the overall performance by 22% to 36% relative to a random selection, 16% to 35% relative to a majority voting scheme, and 15% to 23% relative to the best individual system. Index Terms: Question answering, Systems for spoken language understanding Sibel Yaman, Dilek Hakkani-Tür, Gökhan Tür, Ralph Grishman, Mary P. Harper, Kathy McKeown, Adam Meyers 0001, Kartavya Sharma |
INTERSPEECH | 2 |
| 2009 | IXIR: A statistical information distillation system
Michael Levit, Dilek Hakkani-Tür, Gökhan Tür, Daniel Gillick |
Comput. Speech Lang. | 2 |
| 2009 | Generative and Discriminative Methods Using Morphological Information for Sentence Segmentation of TurkishabstractThis paper presents novel methods for generative, discriminative, and hybrid sequence classification for segmentation of Turkish word sequences into sentences. In the literature, this task is generally solved using statistical models that take advantage of lexical information among others. However, Turkish has a productive morphology that generates a very large vocabulary, making the task much harder. In this paper, we introduce a new set of morphological features, extracted from words and their morphological analyses. We also extend the established method of hidden event language modeling (HELM) to factored hidden event language modeling (fHELM) to handle morphological information. In order to capture non-lexical information, we extract a set of prosodic features, which are mainly motivated from our previous work for other languages. We then employ discriminative classification techniques, boosting and conditional random fields (CRFs), combined with fHELM, for the task of Turkish sentence segmentation. Ümit Güz, Benoît Favre, Dilek Hakkani-Tür, Gökhan Tür |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Introduction to the Special Issue on Processing Morphologically Rich LanguagesabstractThe 12 papers in this special issue span a variety of speech and language processing applications highlighting the challenges and providing solutions for dealing with the morphological complexity of different languages. Ruhi Sarikaya, Katrin Kirchhoff, Tanja Schultz, Dilek Hakkani-Tür |
IEEE Trans. Speech Audio Process. | 4 |
| 2008 | Punctuating speech for information extractionabstractThis paper studies the effect of automatic sentence boundary detection and comma prediction on entity and relation extraction in speech. We show that punctuating the machine generated transcript according to maximum F-measure of period and comma annotation results in suboptimal information extraction. Precisely, period and comma decision thresholds can be chosen in order to improve the entity value score and the relation value score by 4% relative. Error analysis shows that preventing noun-phrase splitting by generating longer sentences and fewer commas can be harmful for IE performance. Indeed, it seems that missed punctuation allows syntactic parsers to merge noun-phrases and prevent the extraction of correct information. Benoît Favre, Ralph Grishman, Dustin Hillard, Heng Ji 0001, Dilek Hakkani-Tür, Mari Ostendorf |
ICASSP | 5 |
| 2008 | An iterative unsupervised learning method for information distillationabstractInformation distillation techniques are used to analyze and interpret large volumes of speech and text archives in multiple languages and produce structured information of interest to the user. In this work, we propose an iterative unsupervised sentence extraction method to answer open-ended natural language queries about an event. The approach consists of finding the subset of sentences that are very likely to be relevant or irrelevant for the query from candidate documents, and iteratively training a classification model using these examples. Our results indicate that performance of the system may be improved by around 30% relative in terms of F-measure, by using the proposed method. Kamand Kamangar, Dilek Hakkani-Tür, Gökhan Tür, Michael Levit |
ICASSP | 2 |
| 2008 | Name-aware speech recognition for interactive question answeringabstractIn this work we show how interactivity in a voice-enabled question answering application may improve speech recognition. We allow the user to provide a target named entity before asking the question. Then we build a named entity specific language model using the documents containing the named entity. The question-specific model is obtained by merging the named entity specific model with the model built on a set of questions. We present a set of experiments using the TREC question set on the AQUAINT corpus. The question-specific language model is compared with the baseline model built by merging a model of the AQUAINT corpus and past TREC questions. The question-specific model achieves 32.2% reduction in word error rate from the baseline using the questions where pronominal references are resolved. Svetlana Stoyanchev, Gökhan Tür, Dilek Hakkani-Tür |
ICASSP | 3 |
| 2008 | Unsupervised learning of edit parameters for matching name variantsabstractSince named entities are often written in different ways, question answering (QA) and other language processing tasks stand to benefit from entity matching. We address the problem of finding equivalent person names in unstructured text. Our approach is a generalization of spelling correction: We compare to candidate matches by applying a set of edits to an input name. We introduce a novel unsupervised method for learning spelling edit probabilities which improves overall F-Measure on our own name-matching task by 12%. Relevance is demonstrated by application to the GALE Distillation task. Index Terms: equivalent names, entity matching, unsupervised learning Daniel Gillick, Dilek Hakkani-Tür, Michael Levit |
INTERSPEECH | 2 |
| 2008 | Packing the meeting summarization knapsackabstractDespite considerable work in automatic meeting summarization over the last few years, comparing results remains difficult due to varied task conditions and evaluations. To address this issue, we present a method for determining the best possible extractive summary given an evaluation metric like ROUGE. Our oracle system is based on a knapsack-packing framework, and though NP-Hard, can be solved nearly optimally by a genetic algorithm. To frame new research results in a meaningful context, we suggest presenting our oracle results alongside two simple baselines. We show oracle and baseline results for a variety of evaluation scenarios that have recently appeared in this field. Korbinian Riedhammer, Daniel Gillick, Benoît Favre, Dilek Hakkani-Tür |
INTERSPEECH | 4 |
| 2008 | Cross-lingual sentence extraction for information distillationabstractInformation distillation aims to analyze and interpret large vol-umes of speech and text archives in multiple languages and pro-duce structured information of interest to the user. In this work, we investigate cross-lingual information distillation, where non-English (source language) documents are searched for user queries that are in English (target language). We propose to per-form distillation both on the original source language data and their English translations output by machine translation, and combine the two outputs. We experimentally show that com-bination approach results in 8 % to 16 % absolute (13 % to 31% relative) F-measure improvement over the previous work. Index Terms: information distillation, sentence extraction, cross-lingual processing, and classification model combination. Adish Kumar Singla, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2008 | Role recognition for meeting participants: an approach based on lexical information and social network analysisabstractThis paper presents experiments on the automatic recognition of roles in meetings. The proposed approach combines two sources of information: the lexical choices made by people playing different roles on one hand, and the Social Networks describing the interactions between the meeting participants on the other hand. Both sources lead to role recognition results significantly higher than chance when used separately, but the best results are obtained with their combination. Preliminary experiments obtained over a corpus of 138 meeting recordings (over 45 hours of material) show that around 70% of the time is labeled correctly in terms of role. Neha P. Garg, Sarah Favre, Hugues Salamin, Dilek Hakkani-Tür, Alessandro Vinciarelli |
ACM Multimedia | 4 |
| 2008 | Efficient sentence segmentation using syntactic featuresabstractTo enable downstream language processing,automatic speech recognition output must be segmented into its individual sentences. Previous sentence segmentation systems have typically been very local,using low-level prosodic and lexical features to independently decide whether or not to segment at each word boundary position. In this work,we leverage global syntactic information from a syntactic parser, which is better able to capture long distance dependencies. While some previous work has included syntactic features, ours is the first to do so in a tractable, lattice-based way, which is crucial for scaling up to long-sentence contexts. Specifically, an initial hypothesis lattice is constructed using local features. Candidate sentences are then assigned syntactic language model scores. These global syntactic scores are combined with local low-level scores in a log-linear model. The resulting system significantly outperforms the most popular long-span model for sentence segmentation (the hidden event language model) on both reference text and automatic speech recognizer output from news broadcasts. Benoît Favre, Dilek Hakkani-Tür, Slav Petrov, Daniel Klein 0001 |
SLT | 2 |
| 2008 | Efficient data selection for machine translationabstractPerformance of statistical machine translation (SMT) systems relies on the availability of a large parallel corpus which is used to estimate translation probabilities. However, the generation of such corpus is a long and expensive process. In this paper, we introduce two methods for efficient selection of training data to be translated by humans. Our methods are motivated by active learning and aim to choose new data that adds maximal information to the currently available data pool. The first method uses a measure of disagreement between multiple SMT systems, whereas the second uses a perplexity criterion. We performed experiments on Chinese-English data in multiple domains and test sets. Our results show that we can select only one-fifth of the additional training data and achieve similar or better translation performance, compared to that of using all available data. Arindam Mandal, Dimitra Vergyri, Wen Wang 0001, Jing Zheng 0001, Andreas Stolcke, Gökhan Tür, Dilek Hakkani-Tür, Necip Fazil Ayan |
SLT | 7 |
| 2008 | A keyphrase based approach to interactive meeting summarizationabstractRooted in multi-document summarization, maximum marginal relevance (MMR) is a widely used algorithm for meeting summarization (MS). A major problem in extractive MS using MMR is finding a proper query: the centroid based query which is commonly used in the absence of a manually specified query, can not significantly outperform a simple baseline system. We introduce a simple yet robust algorithm to automatically extract keyphrases (KP) from a meeting which can then be used as a query in the MMR algorithm. We show that the KP based system significantly outperforms both baseline and centroid based systems. As human refined KPs show even better summarization performance, we outline how to integrate the KP approach into a graphical user interface allowing interactive summarization to match the user's needs in terms of summary length and topic focus. Korbinian Riedhammer, Benoît Favre, Dilek Hakkani-Tür |
SLT | 3 |
| 2008 | The CALO meeting speech recognition and understanding systemabstractThe CALO Meeting Assistant provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper summarizes the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, question-answer pair identification, action item recognition, decision extraction, and summarization. Gökhan Tür, Andreas Stolcke, L. Lynn Voss, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Dilek Hakkani-Tür, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Stanley Peters, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri |
SLT | 11 |
| 2008 | Bootstrapping spoken dialogue systems by exploiting reusable librariesabstractAbstract Building natural language spoken dialogue systems requires large amounts of human transcribed and labeled speech utterances to reach useful operational service performances. Furthermore, the design of such complex systems consists of several manual steps. The User Experience (UE) expert analyzes and defines by hand the system core functionalities: the system semantic scope (call-types) and the dialogue manager strategy that will drive the human–machine interaction. This approach is extensive and error-prone since it involves several nontrivial design decisions that can be evaluated only after the actual system deployment. Moreover, scalability is compromised by time, costs, and the high level of UE know-how needed to reach a consistent design. We propose a novel approach for bootstrapping spoken dialogue systems based on the reuse of existing transcribed and labeled data, common reusable dialogue templates, generic language and understanding models, and a consistent design process. We demonstrate that our approach reduces design and development time while providing an effective system without any application-specific data. Giuseppe Di Fabbrizio, Gökhan Tür, Dilek Hakkani-Tür, Mazin Gilbert, Bernard Renger, David C. Gibbon, Zhu Liu 0001, Behzad Shahraray |
Nat. Lang. Eng. | 3 |
| 2007 | Integrating several annotation layers for statistical information distillationabstractWe present a sentence extraction algorithm for Information Distillation, a task where for a given templated query, relevant passages must be extracted from massive audio and textual document sources. For each sentence of the relevant documents (that are assumed to be known from the upstream stages) we employ statistical classification methods to estimate the extent of its relevance to the query, whereby two aspects of relevance are taken into account: the template (type) of the query and its slots (free-text descriptions of names, organizations, topic, events and so on, around which templates are centered). The idiosyncrasy of the presented method is in the choice of features used for classification. We extract our features from charts, compilations of elements from various annotation levels, such as word transcriptions, syntactic and semantic parses, and Information Extraction annotations. In our experiments we show that this integrated approach outperforms a purely lexical baseline by as much as 30% relative in terms of F-measure. We also investigate the algorithm's behavior under noisy conditions, by comparing its performance on ASR output and on corresponding manual transcriptions. Michael Levit, Dilek Hakkani-Tür, Gökhan Tür, Daniel Gillick |
ASRU | 2 |
| 2007 | Statistical Sentence Extraction for Information DistillationabstractInformation distillation aims to extract the most useful pieces of information related to a given query from massive, possibly multilingual, audio and textual document sources. One critical component in a distillation engine is detecting sentences to be extracted from each relevant document. In this paper, we present a statistical sentence extraction approach for distillation. Basically, we frame this tack as a classification problem, where each candidate sentence in documents is classified as a relevant to the query or not. These documents may be textual or audio format and in a number of languages. For audio documents, we use both manual and automatic transcriptions, for non-English documents, we use automatic translations. In this work, we use AdaBoost, a discriminative classification method with both lexical and semantic features. The results indicate 11%-13% relative improvement over a baseline keyword-spotting-based approach. We also show the robustness of our method on the audio subset of the document sources using manual and automatic transcriptions. Dilek Hakkani-Tür, Gökhan Tür |
ICASSP (4) | 1 |
| 2007 | Entropy Based Classifier Combination for Sentence SegmentationabstractWe describe recent extensions to our previous work, where we explored the use of individual classifiers, namely, boosting and maximum entropy models for sentence segmentation. In this paper we extend the set of classification methods with support vector machine (SVM). We propose a new dynamic entropy-based classifier combination approach to combine these classifiers, and compare it with the traditional classifier combination techniques, namely, voting, linear regression and logistic regression. Furthermore, we also investigate the combination of hidden event language models with the output of the proposed classifier combination, and the output of individual classifiers. Experimental studies conducted on the Mandarin TDT4 broadcast news database shows that the SVM classifier as an individual classifier improves over our previous best system. However, the proposed entropy-based classifier combination approach shows the best improvement in F-measure of 1% absolute, and the voting approach shows the best reduction in NIST error rate of 2.7% absolute when compared to the previous best system. Mathew Magimai-Doss, Dilek Hakkani-Tür, Özgür Çetin, Elizabeth Shriberg, James G. Fung, Nikki Mirghafori |
ICASSP (4) | 2 |
| 2007 | Cross-linguistic analysis of prosodic features for sentence segmentationabstractIn this paper, we perform a cross-linguistic study of prosodic features in sentence segmentation by using two different feature selection approaches: a forward search wrapper and feature filtering. Experiments in Arabic, English, and Mandarin show that prosodic features make significant contributions in all three languages. Feature selection results indicate that feature relevancy can vary greatly depending on the target language, and therefore the optimal feature subset varies considerably between languages. We observe patterns in the feature selection and the affinity of the different languages toward certain feature types, which gives us insight into future feature selection and feature design. Index Terms: prosodic features, cross lingual, feature selection, sentence segmentation James G. Fung, Dilek Hakkani-Tür, Mathew Magimai-Doss, Elizabeth Shriberg, Sébastien Cuendet, Nikki Mirghafori |
INTERSPEECH | 2 |
| 2007 | Co-training using prosodic and lexical information for sentence segmentationabstractWe investigate the application of the co-training learning algorithm on the sentence boundary classification problem by using lexical and prosodic information. Co-training is a semi-supervised machine learning algorithm that uses multiple weak classifiers with a relatively small amount of labeled data and incrementally uses unlabeled data. The assumption in co-training is that the classifiers can co-train each other, as one can label samples that are difficult for the other. The sentence segmentation problem is very appropriate for the co-training method since it satisfies the main requirements of the co-training algorithm: the dataset can be described by two disjoint and natural views that are redundantly sufficient. In our case, the feature sets are capturing lexical and prosodic information. The experimental results on the ICSI Meeting (MRDA) corpus show the effectiveness of the co-training algorithm for this task. Index Terms: co-training, sentence segmentation, prosody, self-training, Boosting Ümit Güz, Sébastien Cuendet, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 3 |
| 2007 | Exploiting information extraction annotations for document retrieval in distillation tasksabstractInformation distillation aims to extract relevant pieces of information related to a given query from massive, possibly multilingual, audio and textual document sources. In this paper, we present our approach for using information extraction annotations to augment document retrieval for distillation. We take advantage of the fact that some of the distillation queries can be associated with annotation elements introduced for the NIST Automatic Content Extraction (ACE) task. We experimentally show that using the ACE events to constrain the document set returned by an information retrieval engine significantly improves the precision at various recall rates for two different query templates. Index Terms: information distillation, information retrieval, information extraction, document retrieval Dilek Hakkani-Tür, Gökhan Tür, Michael Levit |
INTERSPEECH | 1 |
| 2007 | Improving speech translation with automatic boundary predictionabstractThis paper investigates the influence of automatic sentence boundary and sub-sentence punctuation prediction on machine translation (MT) of automatically recognized speech.We use prosodic and lexical cues to determine sentence boundaries, and successfully combine two complementary approaches to sentence boundary prediction.We also introduce a new feature for segmentation prediction that directly considers the assumptions of the phrase translation model.In addition, we show how automatically predicted commas can be used to constrain reordering in MT search.We evaluate the presented methods using a state-of-the-art phrase-based statistical MT system on two large vocabulary tasks.We find that careful optimization of the segmentation parameters directly for translation quality improves the translation results in comparison to independent optimization for segmentation quality of the predicted source language sentence boundaries. Evgeny Matusov, Dustin Hillard, Mathew Magimai-Doss, Dilek Hakkani-Tür, Mari Ostendorf, Hermann Ney |
INTERSPEECH | 4 |
| 2006 | Webtalk: Towards Automatically Building Spoken Dialog Systems Through MiningwebsitesabstractWeb Talk is a system for analyzing unstructured information from company websites to support automtic reation of spoken dialog applications. The goal is to completely automate the process of building, maintaining and deploying dialog applications by leveraging the wealth of information on the World Wide Web. Web Talk employs technologies in web mining, document understanding, question/answering, and speech and language processing. In this paper, we review extensions to these technologies to make them suitable for creating a Web Talk application. e present an evaluation study of a Web Talk spoken dialog system that has been instantiated on a telecom company website. Experiments with 30 different scenarios indicate promising results and provide evidence that such systems can potentially revolutionize the paradigm for creating and scaling spoken dialog services. Junlan Feng, Dilek Hakkani-Tür, Giuseppe Di Fabbrizio, Mazin Gilbert, Marc C. Beutnagel |
ICASSP (1) | 2 |
| 2006 | Bootstrapping Language Models for Spoken Dialog Systems From The World Wide WebabstractIn this paper, we describe our approach for bootstrapping statistical language models for spoken dialog systems using indomain Web data and utterances collected from previous applications. The approach is based on the idea of stitching conversational templates with the predicate and arguments extracted from the Web pages using semantic role labeling, to generate conversational style utterances. The conversational templates represent the task-independent portions of user utterances and can be built by hand, or learned from utterances collected from other domain applications. Experiments have shown that, stitching with both types of conversational templates have resulted in significantly better ASR word accuracy. Furthermore, the new language model bootstrapping approach can be combined with unsupervised and active learning to improve word accuracy even with very little in-domain transcribed data Dilek Hakkani-Tür, Mazin G. Rahim |
ICASSP (1) | 1 |
| 2006 | QASR: question answering using semantic roles for speech interfaceabstractIn this paper, we evaluate a semantic role labeling approach to the extraction of answers in the open domain question answering task. We show that this technique especially improves the system performance when answers are communicated to the user by voice. Semantic role labeling identifies predicates and semantic argument phrases in a sentence. With this information we are able to analyze and extract structure from both questions and candidate sentences, which helps us identify more relevant and precise answers in a long list of candidate sentences. When searching for an answer to a question, we match the missing argument in the question to the semantic parses of the candidate answers. This technique significantly improves the accuracy of the question answering system and results in more concise and grammatical answers, which is essential for enabling voice interfaces to question answering systems. In this paper we apply our approach to factoid questions containing predicates; however, this technique can be also useful in answering more complex questions. Index Terms: question answering, semantic roles Svetlana Stenchikova, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 2 |
| 2006 | The ICSI+ multilingual sentence segmentation systemabstractThe ICSI+ multilingual sentence segmentation with results for English and Mandarin broadcast news automatic speech recognizer transcriptions represents a joint effort involving ICSI, SRI, and UT Dallas. Our approach is based on using hidden event language models for exploiting lexical information, and maximum entropy and boosting classifiers for exploiting lexical, as well as prosodic, speaker change and syntactic information. We demonstrate that the proposed methodology including pitch- and energy-related prosodic features performs significantly better than a baseline system that uses words and simple pause features only. Furthermore, the obtained improvements are consistent across both languages, and no language-specific adaptation of the methodology is necessary. The best results were achieved by combining hidden event language models with a boosting-based classifier that to our knowledge has not previously been applied for this task. M. Zimmerman, Dilek Hakkani-Tür, James G. Fung, Nikki Mirghafori, Luke R. Gottlieb, Elizabeth Shriberg, Yang Liu 0004 |
INTERSPEECH | 2 |
| 2006 | Let's Discoh: Collecting an Annotated Open Corpuswith Dialogue Acts and Reward signals for Natural Language HelpdesksabstractWe motivate and explain the DlSCoH project, which uses a publicly deployed spoken dialogue system for conference services to collect a richly annotated corpus of mixed-initiative human- machine spoken dialogues. System users are able to call a phone number and learn about a conference, including paper submission, program, venue, accommodation options and costs, etc. The collected corpus is (1) usable for training, evaluating and comparing statistical models, (2) naturally spoken and task oriented, (3) extendible / generalizable, (4) collected using state-of-the-art research and commercial technology, (5) freely available to researchers. We explain the principles behind the dialogue context representations and reward signals collected by the system, as well as the overall system design, call types, and call flow. We also present results regarding the initial ASR models and spoken language understanding models. We expect the resulting corpora to be used in advanced dialogue research over the coming years. Giovanni Andreani, Giuseppe Di Fabbrizio, Mazin Gilbert, Daniel Gillick, Dilek Hakkani-Tür, Oliver Lemon |
SLT | 5 |
| 2006 | Model Adaptation for Sentence Segmentation from speechabstractThis paper analyzes various methods to adapt sentence segmentation models trained on conversational telephone speech (CTS) to meeting style conversations. The sentence segmentation model trained using a large amount of CTS data is used to improve the performance when various amounts of meeting data are available. We test the sentence segmentation performance on both reference and speech-to-text (STT) conditions on the ICSI MRDA meeting corpus using the switchboard CTS Corpus as the out-of-domain data. Results show that the sentence segmentation performance is significantly improved by the adapted classification model compared to the one obtained by using in-domain data only, independently of the amount of in-domain data used: 17.5% and 8.4% relative error reductions with only 1,000 and 3,000 in-domain sentences, respectively, and 3.7% relative error reduction with all in-domain data of 80,000 words. Sébastien Cuendet, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 2 |
| 2006 | Impact of Automatic Comma Prediction on POS/Name Tagging of speechabstractThis work looks at the impact of automatically predicted commas on part-of-speech (POS) and name tagging of speech recognition transcripts of Mandarin broadcast news. There is a significant gain in both POS and name tagging accuracy due to using automatically predicted commas over sentence boundary prediction alone. One difference between Mandarin and English is that there are two types of commas, and experiments here show that, while they can be reliably distinguished in automatic prediction, the distinction does not give a clear benefit for POS or name tagging. Dustin Hillard, Zhongqiang Huang, Heng Ji 0001, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Mari Ostendorf, Wen Wang 0001 |
SLT | 5 |
| 2006 | Model Adaptation for Dialog Act TaggingabstractIn this paper, we analyze the effect of model adaptation for dialog act tagging. The goal of adaptation is to improve the performance of the tagger using out-of-domain data or models. Dialog act tagging aims to provide a basis for further discourse analysis and understanding in conversational speech. In this study we used the ICSI meeting corpus with high-level meeting recognition dialog act (MRDA) tags, that is, question, statement, backchannel, disruptions, and floor grabbers/holders. We performed controlled adaptation experiments using the Switchboard (SWBD) corpus with SWBD-DAMSL tags as the out-of-domain corpus. Our results indicate that we can achieve significantly better dialog act tagging by automatically selecting a subset of the Switchboard corpus and combining the confidences obtained by both in-domain and out-of-domain models via logistic regression, especially when the in-domain data is limited. Gökhan Tür, Ümit Güz, Dilek Hakkani-Tür |
SLT | 3 |
| 2006 | Beyond ASR 1-best: Using word confusion networks in spoken language understanding
Dilek Hakkani-Tür, Frédéric Béchet, Giuseppe Riccardi, Gökhan Tür |
Comput. Speech Lang. | 1 |
| 2006 | Introduction to the Special Issue on Spoken Language Understanding in Conversational Systems
Srinivas Bangalore, Dilek Hakkani-Tür, Gökhan Tür |
Speech Commun. | 2 |
| 2006 | The AT&T spoken language understanding systemabstractSpoken language understanding (SLU) aims at extracting meaning from natural language speech. Over the past decade, a variety of practical goal-oriented spoken dialog systems have been built for limited domains. SLU in these systems ranges from understanding predetermined phrases through fixed grammars, extracting some predefined named entities, extracting users' intents for call classification, to combinations of users' intents and named entities. In this paper, we present the SLU system of VoiceTone/spl reg/ (a service provided by AT&T where AT&T develops, deploys and hosts spoken dialog applications for enterprise customers). The SLU system includes extracting both intents and the named entities from the users' utterances. For intent determination, we use statistical classifiers trained from labeled data, and for named entity extraction we use rule-based fixed grammars. The focus of our work is to exploit data and to use machine learning techniques to create scalable SLU systems which can be quickly deployed for new domains with minimal human intervention. These objectives are achieved by 1) using the predicate-argument representation of semantic content of an utterance; 2) extending statistical classifiers to seamlessly integrate hand crafted classification rules with the rules learned from data; and 3) developing an active learning framework to minimize the human labeling effort for quickly building the classifier models and adapting them to changes. We present an evaluation of this system using two deployed applications of VoiceTone/spl reg/. Narendra K. Gupta, Gökhan Tür, Dilek Hakkani-Tür, Srinivas Bangalore, Giuseppe Riccardi, Mazin Gilbert |
IEEE Trans. Speech Audio Process. | 3 |
| 2005 | The AT&T WATSON Speech RecognizerabstractThis paper describes the AT&T WATSON real-time speech recognizer, the product of several decades of research at AT&T. The recognizer handles a wide range of vocabulary sizes and is based on continuous-density hidden Markov models for acoustic modeling and finite state networks for language modeling. The recognition network is optimized for efficient search. We identify the algorithms used for high-accuracy, real-time and low-latency recognition. We present results for small and large vocabulary tasks taken from the AT&T VoiceTone/sup /spl reg// service, showing word accuracy improvement of about 5% absolute and real-time processing speed-up by a factor between 2 and 3. Vincent Goffin, Cyril Allauzen, Enrico Bocchieri, Dilek Hakkani-Tür, Andrej Ljolje, Sarangarajan Parthasarathy, Mazin G. Rahim, Giuseppe Riccardi, Murat Saraclar |
ICASSP (1) | 4 |
| 2005 | Error Prediction in Spoken Dialog: From Signal-to-Noise Ratio to Semantic Confidence ScoresabstractSpoken dialog systems aim to interpret the meanings of users' utterances and respond to them accordingly. The users' utterances are first recognized by an automatic speech recognizer (ASR) and the intents of the users are extracted by the spoken language understanding (SLU) unit. Both ASR and SLU are noisy and in general their noise statistics are not correlated. Our goal is to exploit the signal-to-noise information and ASR lattice-based and semantic confidence scores for SLU error prediction and prevention of these by rejecting erroneous utterances, or asking confirmation questions. In our experiments, we have shown up to 80% relative decrease in the error rate of the accepted utterances collected using the AT&T How May I Help You/spl trade/ spoken dialog system used for customer care. Dilek Hakkani-Tür, Gökhan Tür, Giuseppe Riccardi, Hong Kook Kim |
ICASSP (1) | 1 |
| 2005 | Automated wizard-of-oz for spoken dialogue systemsabstractDesigning and building natural language spoken dialogue systems require large amounts of speech utterances, which adequately represent the intended human-machine dialogues.For this purpose, typically, first a "Wizard-of-Oz" data collection is performed, and then the collected data is transcribed and labeled by expert labelers.Finally, the data is used to train both the speech recognizer and the spoken language understanding stochastic models.Data collection and labeling is an expensive and time consuming manual process.In this paper we propose a completely Automated Wizard, which is capable of recognizing and understanding application independent requests reusing the previously labeled and transcribed data from similar domains, and improving the informativeness of the collected data.We demonstrate that, in the context of automated call routing, compared to the existing data collection systems, the Automated Wizard better captures the user intentions and produces substantially shorter interactions resulting in a better user experience and a less intrusive approach. Giuseppe Di Fabbrizio, Gökhan Tür, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2005 | Using context to improve emotion detection in spoken dialog systemsabstractMost research that explores the emotional state of users of spoken dialog systems does not fully utilize the contextual nature that the dialog structure provides.This paper reports results of machine learning experiments designed to automatically classify the emotional state of user turns using a corpus of 5,690 dialogs collected with the "How May I Help You SM " spoken dialog system.We show that augmenting standard lexical and prosodic features with contextual features that exploit the structure of spoken dialog and track user state increases classification accuracy by 2.6%. Jackson Liscombe, Giuseppe Riccardi, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2005 | Combining active and semi-supervised learning for spoken language understanding
Gökhan Tür, Dilek Hakkani-Tür, Robert E. Schapire |
Speech Commun. | 2 |
| 2005 | Active learning: theory and applications to automatic speech recognitionabstractWe are interested in the problem of adaptive learning in the context of automatic speech recognition (ASR). In this paper, we propose an active learning algorithm for ASR. Automatic speech recognition systems are trained using human supervision to provide transcriptions of speech utterances. The goal of Active Learning is to minimize the human supervision for training acoustic and language models and to maximize the performance given the transcribed and untranscribed data. Active learning aims at reducing the number of training examples to be labeled by automatically processing the unlabeled examples, and then selecting the most informative ones with respect to a given cost function for a human to label. In this paper we describe how to estimate the confidence score for each utterance through an on-line algorithm using the lattice output of a speech recognizer. The utterance scores are filtered through the informativeness function and an optimal subset of training samples is selected. The active learning algorithm has been applied to both batch and on-line learning scheme and we have experimented with different selective sampling algorithms. Our experiments show that by using active learning the amount of labeled data needed for a given word accuracy can be reduced by more than 60% with respect to random sampling. Giuseppe Riccardi, Dilek Hakkani-Tür |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | Mining Spoken Dialogue Corpora for System Evaluation and Modelin
Frédéric Béchet, Giuseppe Riccardi, Dilek Hakkani-Tür |
EMNLP | 3 |
| 2004 | Unsupervised and active learning in automatic speech recognition for call classificationabstractA key challenge in rapidly building spoken natural language dialog applications is minimizing the manual effort required in transcribing and labeling speech data. This task is not only expensive but also time consuming. We present a novel approach that aims at reducing the amount of manually transcribed in-domain data required for building automatic speech recognition (ASR) models in spoken language dialog systems. Our method is based on mining relevant text from various conversational systems and Web sites. An iterative process is employed where the performance of the models can be improved through both unsupervised and active learning of the ASR models. We have evaluated the robustness of our approach on a call classification task that has been selected from AT&T VoiceTone/sup SM/ customer care. Our results indicate that with unsupervised learning it is possible to achieve a call classification performance that is only 1.5% lower than the upper bound set when using all available in-domain transcribed data. Dilek Hakkani-Tür, Gökhan Tür, Mazin G. Rahim, Giuseppe Riccardi |
ICASSP (1) | 1 |
| 2004 | Extending boosting for call classification using word confusion networksabstractWe are interested in the problem of robust understanding from noisy spontaneous speech input. In goal driven human-machine dialog, utterance classification is a key component of the understanding process to determine the intent of the speaker. We propose a novel algorithm for exploiting ASR word confidence scores for better classification of spoken utterances. Word confidence scores for automatic speech recognition (ASR) provide estimates for word error rates. While previous work has focused on straightforward combination of word confidence scores into Bayesian classifiers, we extend the mathematical formulation for boosting classifiers. This extension of the algorithm allows confidence scores to be exploited from a 1-best ASR output or from word confusion networks (WCNs). We present methods for on-line and off-line score combinations. The results we show are for a large database of utterances collected using the AT&T VoiceTone/sup SM/ spoken dialog system. Our experiments show between 5% and 10% reduction in error (1-precision) for a given recall using WCNs compared to ASR output. Gökhan Tür, Dilek Hakkani-Tür, Giuseppe Riccardi |
ICASSP (1) | 2 |
| 2004 | Detecting and extracting named entities from spontaneous speech in a mixed-initiative spoken dialogue context: How May I Help You?sm, tm
Frédéric Béchet, Allen L. Gorin, Jeremy H. Wright, Dilek Hakkani-Tür |
Speech Commun. | 4 |
| 2003 | A general algorithm for word graph matrix decompositionabstractIn automatic speech recognition, word graphs (lattices) are commonly used as an approximate representation of the complete word search space. Usually these word lattices are acyclic and have no a-priori structure. More recently a new class of normalized word lattices have been proposed. These word lattices (a.k.a. sausages) are very efficient (space) and they provide a normalization (chunking) of the lattice, by aligning words from all possible hypotheses. We propose a general framework for lattice chunking, the pivot algorithm. There are four important components of the pivot algorithm. First, the time information is not necessary but is beneficial for the overall performance. Second, the algorithm allows the definition of a predefined chunk structure of the final word lattice. Third, the algorithm operates on both weighted and unweighted lattices. Fourth, the labels on the graph are generic, and could be words as well as part of speech tags or parse tags. While the algorithm has applications to many tasks (e.g. parsing, named entity extraction) we present results on the performance of confidence scores for different large vocabulary speech recognition tasks. We compare the results of our algorithms against off-the-shelf methods and show significant improvements. Dilek Hakkani-Tür, Giuseppe Riccardi |
ICASSP (1) | 1 |
| 2003 | Active learning for spoken language understandingabstractWe describe active learning methods for reducing the labeling effort in a statistical call classification system. Active learning aims to minimize the number of labeled utterances by automatically selecting for labeling the utterances that are likely to be most informative. The first method, inspired by certainty-based active learning, selects the examples that the classifier is least confident about. The second method, inspired by committee-based active learning, selects the examples that multiple classifiers do not agree on. We have evaluated these active learning methods using a call classification system used for AT&T customer care. Our results indicate that it is possible to reduce human labeling effort at least by a factor of two. Gökhan Tür, Robert E. Schapire, Dilek Hakkani-Tür |
ICASSP (1) | 3 |
| 2003 | Multi-channel sentence classification for spoken dialogue language modelingabstractIn traditional language modeling word prediction is based on the local context (e.g. n-gram). In spoken dialog, language statistics are affected by the multidimensional structure of the human-machine interaction. In this paper we investigate the statistical dependencies of users’ responses with respect to the system’s and user’s channel. The system channel components are the prompts’ text, dialogue history, dialogue state. The user channel components are the Automatic Speech Recognition (ASR) transcriptions, the semantic classifier output and the sentence length. We describe an algorithm for language model rescoring using users’ response classification. The user’s response is first mapped into a multidimensional state and the state specific language model is applied for ASR rescoring. We present perplexity and ASR results on the How May I Help You ? sm 100K spoken dialogs. Frédéric Béchet, Giuseppe Riccardi, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2003 | Active and unsupervised learning for automatic speech recognitionabstractState-of-the-art speech recognition systems are trained using human transcriptions of speech utterances. In this paper, we describe a method to combine active and unsupervised learning for automatic speech recognition (ASR). The goal is to minimize the human supervision for training acoustic and language models and to maximize the performance given the transcribed and untranscribed data. Active learning aims at reducing the number of training examples to be labeled by automatically processing the unlabeled examples, and then selecting the most informative ones with respect to a given cost function. For unsupervised learning, we utilize the remaining untranscribed data by using their ASR output and word confidence scores. Our experiments show that the amount of labeled data needed for a given word accuracy can be reduced by 75% by combining active and unsupervised learning. Giuseppe Riccardi, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2003 | Exploiting unlabeled utterances for spoken language understandingabstractState of the art spoken language understanding systems are trained using labeled utterances, which is labor intensive and time consuming to prepare. In this paper, we propose methods for exploiting the unlabeled data in a statistical call classification system within a natural language dialog system. The basic assumption is that some amount of labeled data and relatively larger chunks of unlabeled data is available. The first method augments the training data by using the machine-labeled call-types for the unlabeled utterances. The second method, instead, augments the classification model trained using the human-labeled utterances with the machine-labeled ones in a weighted manner. We have evaluated these methods using a call classification system used for AT&T natural dialog customer care system. For call classification, we have used a boosting algorithm. Our results indicate that it is possible to obtain the same classification performance by using 30% less labeled data when the unlabeled data is utilized. This corresponds to a 1-1.5% absolute classification error rate reduction, using the same amount of labeled data. Gökhan Tür, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2003 | Active labeling for spoken language understandingabstractState-of-the-art spoken language understanding (SLU) systems are trained using human-labeled utterances, preparation of which is labor intensive and time consuming. Labeling is an error-prone process due to various reasons, such as labeler errors or imperfect description of classes. Thus, usually a second (or maybe more) pass(es) of labeling is required in order to check and fix the labeling errors and inconsistencies of the first (or earlier) pass(es). In this paper, we check the effect of labeling errors for statistical call classification and evaluate methods of finding and correcting these errors by checking minimum amount of data. We describe two alternative methods to speed up the labeling effort, one is based on the confidences obtained from a prior model and the other completely unsupervised. We call the labeling process employing one of these methods as active labeling. Active labeling aims to minimize the number of utterances to be checked again by automatically selecting the ones that are likely to be erroneous or inconsistent with the previously labeled examples. Although very same methods can be used as a postprocessing step to correct labeling errors, we only consider them as part of the labeling process. We have evaluated these active labeling methods using a call classification system used for AT&T natural dialog customer care system. Our results indicate that it is possible to find about 90% of the labeling errors or inconsistencies by checking just half the data. Gökhan Tür, Mazin G. Rahim, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2003 | A statistical information extraction system for TurkishabstractThis paper presents the results of a study on information extraction from unrestricted Turkish text using statistical language processing methods. In languages like English, there is a very small number of possible word forms with a given root word. However, languages like Turkish have very productive agglutinative morphology. Thus, it is an issue to build statistical models for specific tasks using the surface forms of the words, mainly because of the data sparseness problem. In order to alleviate this problem, we used additional syntactic information, i.e. the morphological structure of the words. We have successfully applied statistical methods using both the lexical and morphological information to sentence segmentation, topic segmentation, and name tagging tasks. For sentence segmentation, we have modeled the final inflectional groups of the words and combined it with the lexical model, and decreased the error rate to 4.34%, which is 21% better than the result obtained using only the surface forms of the words. For topic segmentation, stems of the words (especially nouns) have been found to be more effective than using the surface forms of the words and we have achieved 10.90% segmentation error rate on our test set according to the weighted TDT-2 segmentation cost metric. This is 32% better than the word-based baseline model. For name tagging, we used four different information sources to model names. Our first information source is based on the surface forms of the words. Then we combined the contextual cues with the lexical model, and obtained some improvement. After this, we modeled the morphological analyses of the words, and finally we modeled the tag sequence, and reached an F-Measure of 91.56%, according to the MUC evaluation criteria. Our results are important in the sense that, using linguistic information, i.e. morphological analyses of the words, and a corpus large enough to train a statistical model significantly improves these basic information extraction tasks for Turkish. Gökhan Tür, Dilek Hakkani-Tür, Kemal Oflazer |
Nat. Lang. Eng. | 2 |
| 2002 | Active learning for automatic speech recognitionabstractState-of-the-art speech recognition systems are trained using transcribed utterances, preparation of which is labor intensive and time-consuming. In this paper, we describe a new method for reducing the transcription effort for training in automatic speech recognition (ASR). Active learning aims at reducing the number of training examples to be labeled by automatically processing the unlabeled examples, and then selecting the most informative ones with respect to a given cost function for a human to label. We automatically estimate a confidence score for each word of the utterance, exploiting the lattice output of a speech recognizer, which was trained on a small set of transcribed data. We compute utterance confidence scores based on these word confidence scores, then selectively sample the utterances to be transcribed using the utterance confidence scores. In our experiments, we show that we reduce the amount of labeled data needed for a given word accuracy by 27%. 1. Dilek Hakkani-Tür, Giuseppe Riccardi, Allen L. Gorin |
ICASSP | 1 |
| 2002 | Named entity extraction from spontaneous speech in how may i help you?
Frédéric Béchet, Allen L. Gorin, Jeremy H. Wright, Dilek Hakkani-Tür |
INTERSPEECH | 4 |
| 2002 | Improving spoken language understanding using word confusion networksabstractA natural language spoken dialog system includes a large vocabulary automatic speech recognition (ASR) engine, whose output is used as the input of a spoken language understanding component. Two challenges in such a framework are that the ASR component is far from being perfect and the users can say the same thing in very different ways. So, it is very important to be tolerant to recognition errors and some amount of orthographic variability. In this paper, we present our work on developing new methods and investigating various ways of robust recognition and understanding of an utterance. To this end, we exploit word-level confusion networks (sausages), obtained from ASR word graphs (lattices) instead of the ASR 1-best hypothesis. Using sausages with an improved confidence model, we decreased the calltype classification error rate for AT&T's How May I Help You (HMIHY ) natural dialog system by 38%. Gökhan Tür, Jeremy H. Wright, Allen L. Gorin, Giuseppe Riccardi, Dilek Hakkani-Tür |
INTERSPEECH | 5 |
| 2001 | Integrating Prosodic and Lexical Cues for Automatic Topic SegmentationabstractWe present a probabilistic model that uses both prosodic and lexical cues for the automatic segmentation of speech into topically coherent units. We propose two methods for combining lexical and prosodic information using hidden Markov models and decision trees. Lexical information is obtained from a speech recognizer, and prosodic features are extracted automatically from speech waveforms. We evaluate our approach on the Broadcast News corpus, using the DARPA-TDT evaluation metrics. Results show that the prosodic model alone is competitive with word-based segmentation methods. Furthermore, we achieve a significant reduction in error by combining the prosodic and word-based knowledge sources. Gökhan Tür, Dilek Hakkani-Tür, Andreas Stolcke, Elizabeth Shriberg |
Comput. Linguistics | 2 |
| 2000 | Statistical Morphological Disambiguation for Agglutinative Languages
Dilek Hakkani-Tür, Kemal Oflazer, Gökhan Tür |
COLING | 1 |
| 2000 | Prosody-based automatic segmentation of speech into sentences and topics
Elizabeth Shriberg, Andreas Stolcke, Dilek Hakkani-Tür, Gökhan Tür |
Speech Commun. | 3 |
| 1999 | Combining words and prosody for information extraction from speechabstractThe design principles and collection procedures behind a speech synthesis corpus directly impact the performance of the resulting text-to-speech system. This paper describes the design and collection of the Victoria corpus, created to support speech synthesis research and development at Apple Computer. This corpus is composed of ve constituent parts, each designed to cover a speci c aspect of speech synthesis: polyphones, prosodic contexts, reiterant speech, function word sequences, and continuous speech. It was spoken in general U.S. English by one linguisticallytrained adult female. Portions of the corpus are being used in the statistical estimation of duration and pitch models for Apple's next-generation textto-speech system, MacinTalk 4. Dilek Hakkani-Tür, Gökhan Tür, Andreas Stolcke, Elizabeth Shriberg |
EUROSPEECH | 1 |
| 1999 | Modeling the prosody of hidden events for improved word recognitionabstractWe investigate a new approach for using speech prosody as a knowledge source for speech recognition. The idea is to penalize word hypotheses that are inconsistent with prosodic features such as duration and pitch. To model the interaction between words and prosody we modify the language model to represent hidden events such as sentence boundaries and various forms of disfluency, and combine with it decision trees that predict such events from prosodic features. N-best rescoring experiments on the Switchboard corpus show a small but consistent reduction of word error as a result of this modeling. We conclude with a preliminary analysis of the types of errors that are corrected by the prosodically informed model. 1. Andreas Stolcke, Elizabeth Shriberg, Dilek Hakkani-Tür, Gökhan Tür |
EUROSPEECH | 3 |
| 1998 | Automatic detection of sentence boundaries and disfluencies based on recognized wordsabstractWe study the problem of detecting linguistic events at interword boundaries, such as sentence boundaries and disfluency locations, in speech transcribed by an automatic recognizer. Recovering such events is crucial to facilitate speech understanding and other natural language processing tasks. Our approach is based on a combination of prosodic cues modeled by decision trees, and word-based event N-gram language models. Several model combination approaches are investigated. The techniques are evaluated on conversational speech from the Switchboard corpus. Model combination is shown to give a significant win over individual knowledge sources. 1. INTRODUCTION Current automatic speech recognition systems output a string of words. Most natural language understanding systems, however, require structural information such as punctuation, which is present in text but not overtly indicated in spoken language. Similarly, for speech understanding and information extraction, it is important to fi... Andreas Stolcke, Elizabeth Shriberg, Rebecca Bates 0001, Mari Ostendorf, Dilek Hakkani-Tür, Madelaine Plauché, Gökhan Tür |
ICSLP | 5 |
| 1998 | Tactical generation in a free constituent order language
Dilek Hakkani-Tür, Kemal Oflazer |
Nat. Lang. Eng. | 1 |
| 1996 | Tactical Generation in a Free Constituent Order LanguageabstractThis paper describes tactical generation in Turkish, a free con~stituent order language, in which the order of the constituents may change according to the information structure of the sentences to be generated.In the absence of any information regarding the information structure of a sentence (i.e., topic, focus, background, etc.), the constituents of the sentence obey a default order, but the order is almost freely changeable, depending on the constraints of the text flow or discourse.We have used a recursively structured finite state machine for handling the changes in constituent order, implemented as a right-linear grammar backbone.Our implementation environment is the GenKit system, developed at Carnegie Mellon University-Center for Machine Translation.Morphological realization has been implemented using an external morpholggical analysis/generation component which performs concrete morpheme selection and handles morphographemic processes. Dilek Hakkani-Tür, Kemal Oflazer, Ilyas Cicekli |
INLG (1) | 1 |