EDBT 2026 Demo / reviewers in the wild / expert
Gökhan Tür
dblp:24/4469 · also Gokhan Tur
· DBLP profile ↗
131ranked-venue papers
26as first author
19since 2021 · last 2026
0009-0008-7740-2557ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 95 · 22 first-author · 7 since 2021Artificial intelligence and machine learning · 81 · 13 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Current Agents Fail to Leverage World Model as Tool for ForesightabstractCheng Qian, Emre Can Acikgoz, Bingxuan Li, Xiusi Chen, Yuji Zhang, Bingxiang He, Qinyu Luo, Gokhan Tur, Dilek Hakkani-Tür, Yunzhu Li, Heng Ji. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Cheng Qian 0008, Emre Can Acikgoz, Xiusi Chen, Yuji Zhang 0002, Bingxiang He, Qinyu Luo, Gökhan Tür, Dilek Hakkani-Tür, Yunzhu Li, Heng Ji 0001 |
ACL (1) | 8 |
| 2026 | GBC: Gradient-Based Connections for Optimizing Multi-Agent SystemsabstractMulti-agent systems (MAS) built on large language models (LLMs) provide a promising framework for solving complex tasks through role specialization and structured interaction. However, their performance is often limited by miscoordination and, more fundamentally, the lack of fine-grained credit assignment across agents. Existing approaches typically rely on coarse-grained feedback, making it difficult to identify which agents or interaction steps are responsible for errors. We propose Gradient-Based Connections (GBC), an approach for fine-grained attribution and optimization of multi-agent systems. GBC models a MAS as a computational graph and introduces gradient-based connection weights to quantify the influence of each agent’s output on downstream agents at the token level. By constructing an attribution graph and propagating task-specific loss signals backward, our method enables precise identification of error sources and targeted prompt optimization. We further develop AgentChord, an efficient implementation that leverages prefix-based gradient computation. Experiments on MultiWOZ and τ-bench show that GBC improves multi-agent performance and outperforms strong single-agent and multi-agent baselines, and higher attribution quality is associated with greater optimization effectiveness. Code is available at: https://github.com/yxc-cyber/AgentChord. Xiaocheng Yang, Abdulrahman Alrabah, Dilek Hakkani-Tür, Gökhan Tür |
SIGDIAL | 4 |
| 2026 | Goal Alignment in LLM-Based User Simulators for Conversational AIabstractAbstract User simulators are essential to conversational AI, enabling scalable agent development and evaluation through simulated interactions. While current Large Language Models (LLMs) have advanced user simulation capabilities, we reveal that they struggle to consistently demonstrate goal-oriented behavior across multi-turn conversations, which is a critical limitation that compromises their reliability in downstream applications. We introduce User Goal State Tracking (UGST), a novel framework that tracks user goal progression throughout conversations. Leveraging UGST, we present a three-stage methodology for developing user simulators that can autonomously track goal progression and reason to generate goal-aligned responses. Moreover, we establish comprehensive evaluation metrics for measuring goal alignment in user simulators, and demonstrate that our approach yields substantial improvements across two benchmarks (MultiWOZ 2.4 and τ-Bench). Our contributions address a critical gap in conversational AI and establish UGST as an essential framework for developing goal-aligned user simulators. All code and data is released to facilitate future research 1. Shuhaib Mehri, Xiaocheng Yang, Takyoung Kim, Gökhan Tür, Shikib Mehri, Dilek Hakkani-Tür |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language ModelabstractEmre Can Acikgoz, Jeremiah Greer, Akul Datta, Ze Yang, William Zeng, Oussama Elachqar, Emmanouil Koukoumidis, Dilek Hakkani-Tür, Gokhan Tur. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Emre Can Acikgoz, Jeremiah Greer, Akul Datta, William Zeng, Oussama Elachqar, Emmanouil Koukoumidis, Dilek Hakkani-Tür, Gökhan Tür |
ACL (1) | 9 |
| 2025 | Know Your Mistakes: Towards Preventing Overreliance on Task-Oriented Conversational AI Through Accountability ModelingabstractRecent LLMs have enabled significant advancements for conversational agents.However, they are also well known to hallucinate, producing responses that seem plausible but are factually incorrect.On the other hand, users tend to over-rely on LLM-based AI agents, accepting AI's suggestion even when it is wrong.Adding positive friction, such as explanations or getting user confirmations, has been proposed as a mitigation in AI-supported decision-making systems.In this paper, we propose an accountability model for LLM-based task-oriented dialogue agents to address user overreliance via friction turns in cases of model uncertainty and errors associated with dialogue state tracking (DST).The accountability model is an augmented LLM with an additional accountability head that functions as a binary classifier to predict the relevant slots of the dialogue state mentioned in the conversation.We perform our experiments with multiple backbone LLMs on two established benchmarks (MultiWOZ and Snips).Our empirical findings demonstrate that the proposed approach not only enables reliable estimation of AI agent errors but also guides the decoder in generating more accurate actions.We observe around 3% absolute improvement in joint goal accuracy (JGA) of DST output by incorporating accountability heads into modern LLMs.Self-correcting the detected errors further increases the JGA from 67.13 to 70.51, achieving state-of-the-art DST performance.Finally, we show that error correction through user confirmations (friction turn) achieves a similar performance gain, highlighting its potential to reduce user overreliance.1 Suvodip Dey, Yi-Jyun Sun, Gökhan Tür, Dilek Hakkani-Tür |
ACL (1) | 3 |
| 2025 | MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided ConversationsabstractWe introduce MIRAGE, a new benchmark for multimodal expert-level reasoning and decision-making in consultative interaction settings. Designed for the domain of agriculture, MIRAGE captures the full complexity of expert consultations by combining natural user queries, expert-authored responses, and image-based context, offering a high-fidelity benchmark for evaluating models on grounded reasoning, clarification strategies, and long-form generation in a real-world, knowledge-intensive domain. Grounded in over 35,000 real user-expert interactions, and curated through a carefully designed multi-step pipeline, MIRAGE spans diverse crop health, pest diagnosis, and crop management scenarios. The benchmark includes more than 7,000 unique biological entities, covering plant species, pests, and diseases, making it one of the most taxonomically diverse benchmarks available for vision-language models in real-world expert-guided domains. Unlike existing benchmarks that rely on well-specified user inputs, MIRAGE features underspecified, context-rich scenarios, requiring models to infer latent knowledge gaps and either proactively guide the interaction or respond. Our benchmark comprises two core components. The Single-turn Challenge to reason over a single user turn and image set, identify relevant entities, infer causal explanations, and generate actionable recommendations; and a Multi-Turn challenge for dialogue state tracking, goal-driven generation, and expert-level conversational decision-making. We evaluate more than 20 closed and open-source frontier vision-language models (VLMs), using three reasoning language models as evaluators, highlighting the significant challenges posed by MIRAGE in both single-turn and multi-turn interaction settings. Even the advanced GPT4.1 and GPT4o models achieve 44.6% and 40.9% accuracy, respectively, indicating significant room for improvement. Vardhan Dongre, Chi Gui, Shubham Garg, Hooshang Nayyeri, Gökhan Tür, Dilek Hakkani-Tür, Vikram S. Adve |
NeurIPS | 5 |
| 2025 | ToolRL: Reward is All Tool Learning NeedsabstractCurrent Large Language Models (LLMs) often undergo supervised fine-tuning (SFT) to acquire tool use capabilities. However, SFT struggles to generalize to unfamiliar or complex tool use scenarios. Recent advancements in reinforcement learning (RL), particularly with R1-like models, have demonstrated promising reasoning and generalization abilities. Yet, reward design for tool use presents unique challenges: multiple tools may be invoked with diverse parameters, and coarse-grained reward signals, such as answer matching, fail to offer the finegrained feedback required for effective learning.
In this work, we present the first comprehensive study on reward design for tool selection and application tasks within the RL paradigm. We systematically explore a wide range of reward strategies, analyzing their types, scales, granularity, and temporal dynamics. Building on these insights, we propose a principled reward design tailored for tool use tasks and apply it to train LLMs using RL methods.
Empirical evaluations across diverse benchmarks demonstrate that our approach yields robust, scalable, and stable training, achieving a 17\% improvement over base models and a 15\% gain over SFT models. These results highlight the critical role of thoughtful reward design in enhancing the tool use capabilities and generalization performance of LLMs. All the codes are released to facilitate future research. Cheng Qian 0008, Emre Can Acikgoz, Hongru Wang 0011, Xiusi Chen, Dilek Hakkani-Tür, Gökhan Tür, Heng Ji 0001 |
NeurIPS | 7 |
| 2025 | TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level ComparisonsabstractTask-oriented dialogue (TOD) systems are experiencing a revolution driven by Large Language Models (LLMs), yet the evaluation methodologies for these systems remain insufficient for their growing sophistication. While traditional automatic metrics effectively assessed earlier modular systems, they focus solely on the dialogue level and cannot detect critical intermediate errors that can arise during user-agent interactions. In this paper, we introduce TD-EVAL (Turn and Dialogue-level Evaluation), a two-step evaluation framework that unifies fine-grained turn-level analysis with holistic dialogue-level comparisons. At turn-level, we assess each response along three TOD-specific dimensions: conversation cohesion, backend knowledge consistency, and policy compliance. Meanwhile, we design TOD Agent Arena that uses pairwise comparisons to provide a measure of dialogue-level quality. Through experiments on MultiWOZ 2.4 and Tau-Bench, we demonstrate that TD-EVAL effectively identifies the conversational errors that conventional metrics miss. Furthermore, TD-EVAL exhibits better alignment with human judgments than traditional and LLM-based metrics. These findings demonstrate that TD-EVAL introduces a new paradigm for TOD system evaluation, efficiently assessing both turn and system levels with an easily reproducible framework for future research. Emre Can Acikgoz, Carl Guo, Suvodip Dey, Akul Datta, Takyoung Kim, Gökhan Tür, Dilek Hakkani-Tür |
SIGDIAL | 6 |
| 2024 | Dialog Flow Induction for Constrainable LLM-Based ChatbotsabstractStuti Agrawal, Pranav Pillai, Nishi Uppuluri, Revanth Gangi Reddy, Sha Li, Gokhan Tur, Dilek Hakkani-Tur, Heng Ji. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024. Stuti Agrawal, Pranav Pillai, Nishi Uppuluri, Revanth Gangi Reddy, Gökhan Tür, Dilek Hakkani-Tür, Heng Ji 0001 |
SIGDIAL | 6 |
| 2024 | Large Language Models as User-Agents For Evaluating Task-Oriented-Dialogue SystemsabstractTraditionally, offline datasets have been used to evaluate task-oriented dialogue (TOD) models. These datasets lack context awareness, making them suboptimal benchmarks for conversational systems. In contrast, user-agents, which are context-aware, can simulate the variability and unpredictability of human conversations, making them better alternatives as evaluators. Prior research has utilized large language models (LLMs) to develop user-agents. Our work builds upon this by using LLMs to create user-agents for the evaluation of TOD systems. This involves prompting an LLM, using in-context examples as guidance, and tracking the user-goal state. Our evaluation of diversity and task completion metrics for the user-agents shows improved performance with the use of better prompts. Additionally, we propose methodologies for the automatic evaluation of TOD models within this dynamic framework. We make our code publicly available11https://github.com/TaahaKazi/user-agent Taaha Kazi, Ruiliang Lyu, Sizhe Zhou, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 5 |
| 2024 | Confidence Estimation For LLM-Based Dialogue State TrackingabstractEstimation of a model’s confidence on its outputs is critical for Conversational AI systems based on large language models (LLMs), especially for reducing hallucination and preventing over-reliance. In this work, we provide an exhaustive exploration of methods, including approaches proposed for open- and closed-weight LLMs, aimed at quantifying and leveraging model uncertainty to improve the reliability of LLM-generated responses, specifically focusing on dialogue state tracking (DST) in task-oriented dialogue systems (TODS). Regardless of the model type, well-calibrated confidence scores are essential to handle uncertainties, thereby improving model performance. We evaluate four methods for estimating confidence scores based on softmax, raw token scores, verbalized confidences, and a combination of these methods, using the area under the curve (AUC) metric to assess calibration, with higher AUC indicating better calibration. We also enhance these with a self-probing mechanism, proposed for closed models. Furthermore, we assess these methods using an open-weight model fine-tuned for the task of DST, achieving superior joint goal accuracy (JGA). Our findings also suggest that fine-tuning open-weight LLMs can result in enhanced AUC performance, indicating better confidence score calibration. Yi-Jyun Sun, Suvodip Dey, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 4 |
| 2023 | MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesabstractJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, Prem Natarajan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gökhan Tür, Premkumar Natarajan |
ACL (1) | 15 |
| 2022 | TEACh: Task-Driven Embodied Agents That ChatabstractRobots operating in human spaces must be able to engage in natural language interaction, both understanding and executing instructions, and using conversation to resolve ambiguity and correct mistakes. To study this, we introduce TEACh, a dataset of over 3,000 human-human, interactive dialogues to complete household tasks in simulation. A Commander with access to oracle information about a task communicates in natural language with a Follower. The Follower navigates through and interacts with the environment to complete tasks varying in complexity from "Make Coffee" to "Prepare Breakfast", asking questions and getting additional information from the Commander. We propose three benchmarks using TEACh to study embodied intelligence challenges, and we evaluate initial models' abilities in dialogue understanding, language grounding, and task execution. Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gökhan Tür, Dilek Hakkani-Tür |
AAAI | 8 |
| 2022 | Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding SystemsabstractWe present results from a large-scale experiment on pretraining encoders with non-embedding parameter counts ranging from 700M to 9.3B, their subsequent distillation into smaller models ranging from 17M-170M parameters, and their application to the Natural Language Understanding (NLU) component of a virtual assistant system. Though we train using 70% spoken-form data, our teacher models perform comparably to XLM-R and mT5 when evaluated on the written-form Cross-lingual Natural Language Inference (XNLI) corpus. We perform a second stage of pretraining on our teacher models using in-domain data from our system, improving error rates by 3.86% relative for intent classification and 7.01% relative for slot filling. We find that even a 170M-parameter model distilled from our Stage 2 teacher model has 2.88% better intent classification and 7.69% better slot filling error rates when compared to the 2.3B-parameter teacher trained only on public data (Stage 1), emphasizing the importance of in-domain data for pretraining. When evaluated offline using labeled NLU data, our 17M-parameter Stage 2 distilled model outperforms both XLM-R Base (85M params) and DistillBERT (42M params) by 4.23% to 6.14%, respectively. Finally, we present results from a full virtual assistant experimentation platform, where we find that models trained using our pretraining and distillation pipeline outperform models distilled from 85M-parameter teachers by 3.74%-4.91% on an automatic measurement of full-system user dissatisfaction. Jack FitzGerald, Shankar Ananthakrishnan, Konstantine Arkoudas, Davide Bernardi, Abhishek Bhagia, Claudio Delli Bovi, Jin Cao 0003, Rakesh Chada, Amit Chauhan, Luoxin Chen, Anurag Dwarakanath, Satyam Dwivedi, Turan Gojayev, Karthik Gopalakrishnan 0001, Thomas Gueudré, Dilek Hakkani-Tür, Wael Hamza, Jonathan J. Hüser, Kevin Martin Jose, Haidar Khan, Beiye Liu, Jianhua Lu, Alessandro Manzotti, Pradeep Natarajan, Karolina Owczarzak, Gokmen Oz, Enrico Palumbo, Charith Peris, Chandana Satya Prakash, Stephen Rawls, Andy Rosenbaum, Anjali Shenoy, Saleh Soltan, Mukund Sridhar, Lizhen Tan, Fabian Triefenbach, Pan Wei, Shuai Zheng 0004, Gökhan Tür, Premkumar Natarajan |
KDD | 40 |
| 2021 | Learning Better Visual Dialog Agents With Pretrained Visual-Linguistic RepresentationabstractGuessWhat?! is a visual dialog guessing game which incorporates a Questioner agent that generates a sequence of questions, while an Oracle agent answers the respective questions about a target object in an image. Based on this dialog history between the Questioner and the Oracle, a Guesser agent makes a final guess of the target object. While previous work has focused on dialogue policy optimization and visual-linguistic information fusion, most work learns the vision-linguistic encoding for the three agents solely on the GuessWhat?! dataset without shared and prior knowledge of vision-linguistic representation. To bridge these gaps, this paper proposes new Oracle, Guesser and Questioner models that take advantage of a pretrained vision-linguistic model, VilBERT. For Oracle model, we introduce a two-way background/target fusion mechanism to understand both intra and inter-object questions. For Guesser model, we introduce a state-estimator that best utilizes VilBERT’s strength in single-turn referring expression comprehension. For the Questioner, we share the state-estimator from pretrained Guesser with Questioner to guide the question generator. Experimental results show that our proposed models outperform state-of-the-art models significantly by 7%, 10%, 12% for Oracle, Guesser and End-to-End Questioner respectively. Tao Tu 0002, Qing Ping, Govindarajan Thattai, Gökhan Tür, Premkumar Natarajan |
CVPR | 4 |
| 2021 | Language Model is all You Need: Natural Language Understanding as Question AnsweringabstractDifferent flavors of transfer learning have shown tremendous impact in advancing research and applications of machine learning. In this work we study the use of a certain family of transfer learning, where the target domain is mapped to the source domain. Specifically we map Natural Language Understanding (NLU) problems to Question Answering (QA) problems and we show that in low data regimes this approach offers significant improvements compared to other approaches to NLU. Moreover, we show that these gains could be increased through sequential transfer learning across NLU problems from different domains. We show that our approach could reduce the amount of required data for the same performance by up to a factor of 10. Mahdi Namazifar, Alexandros Papangelis, Gökhan Tür, Dilek Hakkani-Tür |
ICASSP | 3 |
| 2021 | Correcting Automated and Manual Speech Transcription Errors Using Warped Language ModelsabstractMasked language models have revolutionized natural language processing systems in the past few years. A recently introduced generalization of masked language models called warped language models are trained to be more robust to the types of errors that appear in automatic or manual transcriptions of spoken language by exposing the language model to the same types of errors during training. In this work we propose a novel approach that takes advantage of the robustness of warped language models to transcription noise for correcting transcriptions of spoken language. We show that our proposed approach is able to achieve up to 10% reduction in word error rates of both automatic and manual transcriptions of spoken language. Mahdi Namazifar, John Malik, Li Erran Li, Gökhan Tür, Dilek Hakkani-Tür |
Interspeech | 4 |
| 2021 | Generative Conversational NetworksabstractAlexandros Papangelis, Karthik Gopalakrishnan, Aishwarya Padmakumar, Seokhwan Kim, Gokhan Tur, Dilek Hakkani-Tur. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Alexandros Papangelis, Karthik Gopalakrishnan 0001, Aishwarya Padmakumar, Seokhwan Kim, Gökhan Tür, Dilek Hakkani-Tür |
SIGDIAL | 5 |
| 2021 | Warped Language Models for Noise Robust Language UnderstandingabstractMasked Language Models (MLM) are self-supervised neural networks trained to fill in the blanks in a given sentence with masked tokens. Despite the tremendous success of MLMs for various text based tasks, they are not robust for spoken language understanding, especially for spontaneous conversational speech recognition noise. In this work we introduce Warped Language Models (WLM) in which input sentences at training time go through the same modifications as in MLM, plus two additional modifications, namely inserting and dropping random tokens. These two modifications extend and contract the sentence in addition to the modifications in MLMs, hence the word "warped" in the name. The insertion and drop modification of the input text during training of WLM resemble the types of noise due to Automatic Speech Recognition (ASR) errors, and as a result WLMs are likely to be more robust to ASR noise. Through computational results we show that natural language understanding systems built on top of WLMs perform better compared to those built based on MLMs, especially in the presence of ASR errors. Mahdi Namazifar, Gökhan Tür, Dilek Hakkani-Tür |
SLT | 2 |
| 2020 | Joint Contextual Modeling for ASR Correction and Language UnderstandingabstractThe quality of automatic speech recognition (ASR) is critical to Dialogue Systems as ASR errors propagate to and directly impact downstream tasks such as language understanding (LU). In this paper, we propose multi-task neural approaches to perform contextual language correction on ASR outputs jointly with LU to improve the performance of both tasks simultaneously. To measure the effectiveness of this approach we used a public benchmark, the 2nd Dialogue State Tracking (DSTC2) corpus. As a baseline approach, we trained task specific Statistical Language Models (SLM) and fine-tuned state-of-the-art Generative Pre-training (GPT) Language Model to re-rank the n-best ASR hypotheses, followed by a model to identify the dialog act and slots. i) We further trained ranker models using GPT and Hierarchical CNN-RNN models with discriminatory losses to detect the best output given n-best hypotheses. We extended these ranker models to first select the best ASR output and then identify the dialogue act and slots in an end to end fashion. ii) We also proposed a novel joint ASR error correction and LU model, a word confusion pointer network (WCN-Ptr) with multihead self attention on top, which consumes the word confusions populated from the n-best. We show that the error rates of off the shelf ASR and following LU systems can be reduced significantly by 14% relative with joint models trained using small amounts of in-domain data. Yue Weng, Sai Sumanth Miryala, Chandra Khatri, Huaixiu Zheng, Piero Molino, Mahdi Namazifar, Alexandros Papangelis, Hugh Williams, Franziska Bell, Gökhan Tür |
ICASSP | 11 |
| 2020 | Exploration Based Language Learning for Text-Based GamesabstractThis work presents an exploration and imitation-learning-based agent capable of state-of-the-art performance in playing text-based computer games. These games are of interest as they can be seen as a testbed for language understanding, problem-solving, and language generation by artificial agents. Moreover, they provide a learning setting in which these skills can be acquired through interactions with an environment rather than using fixed corpora. One aspect that makes these games particularly challenging for learning agents is the combinatorially large action space. Existing methods for solving text-based games are limited to games that are either very simple or have an action space restricted to a predetermined set of admissible actions. In this work, we propose to use the exploration approach of Go-Explore for solving text-based games. More specifically, in an initial exploration phase, we first extract trajectories with high rewards, after which we train a policy to solve the game by imitating these trajectories. Our experiments show that this approach outperforms existing solutions in solving text-based games, and it is more sample efficient in terms of the number of interactions with the environment. Moreover, we show that the learned policy can generalize better than existing solutions to unseen games without using any restriction on the action space. Andrea Madotto, Mahdi Namazifar, Joost Huizinga, Piero Molino, Adrien Ecoffet, Huaixiu Zheng, Alexandros Papangelis, Dian Yu 0002, Chandra Khatri, Gökhan Tür |
IJCAI | 10 |
| 2020 | Multi-Task Siamese Neural Network for Improving Replay Attack DetectionabstractAutomatic speaker verification systems are vulnerable to audio replay attacks which bypass security by replaying recordings of authorized speakers.Replay attack detection (RA) detection systems built upon Residual Neural Networks (ResNet)s have yielded astonishing results on the public benchmark ASVspoof 2019 Physical Access challenge.With most teams using fine-tuned feature extraction pipelines and model architectures, the generalizability of such systems remains questionable though.In this work, we analyse the effect of discriminative feature learning in a multi-task learning (MTL) setting can have on the generalizability and discriminability of RA detection systems.We use a popular ResNet architecture optimized by the cross-entropy criterion as our baseline and compare it to the same architecture optimized by MTL using Siamese Neural Networks (SNN).It can be shown that SNN outperform the baseline by relative 26.8% Equal Error Rate (EER).We further enhance the model's architecture and demonstrate that SNN with additional reconstruction loss yield another significant improvement of relative 13.8 % EER. Patrick von Platen, Gökhan Tür |
INTERSPEECH | 3 |
| 2019 | OCC: A Smart Reply System for Efficient In-App CommunicationsabstractSmart reply systems have been developed for various messaging platforms. In this paper, we introduce Uber's smart reply system: one-click-chat (OCC), which is a key enhanced feature on top of the Uber in-app chat system. It enables driver-partners to quickly respond to rider messages using smart replies. The smart replies are dynamically selected according to conversation content using machine learning algorithms. Our system consists of two major components: intent detection and reply retrieval, which are very different from standard smart reply systems where the task is to directly predict a reply. It is designed specifically for mobile applications with short and non-canonical messages. Reply retrieval utilizes pairings between intent and reply based on their popularity in chat messages as derived from historical data. For intent detection, a set of embedding and classification techniques are experimented with, and we choose to deploy a solution using unsupervised distributed embedding and nearest-neighbor classifier. It has the advantage of only requiring a small amount of labeled training data, simplicity in developing and deploying to production, and fast inference during serving and hence highly scalable. At the same time, it performs comparably with deep learning architectures such as word-level convolutional neural network. Overall, the system achieves a high accuracy of 76% on intent detection. Currently, the system is deployed in production for English-speaking countries and 71% of in-app communications between riders and driver-partners adopted the smart replies to speedup the communication process. Yue Weng, Huaixiu Zheng, Franziska Bell, Gökhan Tür |
KDD | 4 |
| 2019 | Collaborative Multi-Agent Dialogue Model Training Via Reinforcement LearningabstractWe present the first complete attempt at concurrently training conversational agents that communicate only via self-generated language.Using DSTC2 as seed data, we trained natural language understanding (NLU) and generation (NLG) networks for each agent and let the agents interact online.We model the interaction as a stochastic collaborative game where each agent (player) has a role ("assistant", "tourist", "eater", etc.) and their own objectives, and can only interact via natural language they generate.Each agent, therefore, needs to learn to operate optimally in an environment with multiple sources of uncertainty (its own NLU and NLG, the other agent's NLU, Policy, and NLG).In our evaluation, we show that the stochastic-game agents outperform deep learning based supervised baselines. Alexandros Papangelis, Yi-Chia Wang, Piero Molino, Gökhan Tür |
SIGdial | 4 |
| 2019 | Flexibly-Structured Model for Task-Oriented DialoguesabstractThis paper proposes a novel end-to-end architecture for task-oriented dialogue systems.It is based on a simple and practical yet very effective sequence-to-sequence approach, where language understanding and state tracking tasks are modeled jointly with a structured copy-augmented sequential decoder and a multi-label decoder for each slot.The policy engine and language generation tasks are modeled jointly following that.The copyaugmented sequential decoder deals with new or unknown values in the conversation, while the multi-label decoder combined with the sequential decoder ensures the explicit assignment of values to slots.On the generation part, slot binary classifiers are used to improve performance.This architecture is scalable to real-world scenarios and is shown through an empirical evaluation to achieve state-of-the-art performance on both the Cambridge Restaurant dataset and the Stanford in-car assistant dataset 1 . Lei Shu 0004, Piero Molino, Mahdi Namazifar, Hu Xu 0001, Bing Liu 0001, Huaixiu Zheng, Gökhan Tür |
SIGdial | 7 |
| 2018 | (Almost) Zero-Shot Cross-Lingual Spoken Language UnderstandingabstractSpoken language understanding (SLU) is a component of goal-oriented dialogue systems that aims to interpret user's natural language queries in system's semantic representation format. While current state-of-the-art SLU approaches achieve high performance for English domains, the same is not true for other languages. Approaches in the literature for extending SLU models and grammars to new languages rely primarily on machine translation. This poses a challenge in scaling to new languages, as machine translation systems may not be reliable for several (especially low resource) languages. In this work, we examine different approaches to train a SLU component with little supervision for two new languages - Hindi and Turkish, and show that with only a few hundred labeled examples we can surpass the approaches proposed in the literature. Our experiments show that training a model bilingually (i.e., jointly with English), enables faster learning, in that the model requires fewer labeled instances in the target language to generalize. Qualitative analysis shows that rare slot types benefit the most from the bilingual training. Shyam Upadhyay, Manaal Faruqui, Gökhan Tür, Dilek Hakkani-Tür, Larry Heck |
ICASSP | 3 |
| 2018 | Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue SystemsabstractBing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, Larry Heck. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Bing Liu 0024, Gökhan Tür, Dilek Hakkani-Tür, Pararth Shah, Larry Heck |
NAACL-HLT | 2 |
| 2018 | User Modeling for Task Oriented DialoguesabstractWe introduce end-to-end neural network based models for simulating users of task-oriented dialogue systems. User simulation in dialogue systems is crucial from two different perspectives: (i) automatic evaluation of different dialogue models, and (ii) training task-oriented dialogue systems. We design a hierarchical sequence-to-sequence model that first encodes the initial user goal and system turns into fixed length representations using Recurrent Neural Networks (RNN). It then encodes the dialogue history using another RNN layer. At each turn, user responses are decoded from the hidden representations of the dialogue level RNN. This hierarchical user simulator (HUS) approach allows the model to capture undiscovered parts of the user goal without the need of an explicit dialogue state tracking. We further develop several variants by utilizing a latent variable model to inject random variations into user responses to promote diversity in simulated user responses and a novel goal regularization mechanism to penalize divergence of user responses from the initial user goal. We evaluate the proposed models on movie ticket booking domain by systematically interacting each user simulator with various dialogue system policies trained with different objectives and users. Izzeddin Gur, Dilek Hakkani-Tür, Gökhan Tür, Pararth Shah |
SLT | 3 |
| 2017 | Towards Zero-Shot Frame Semantic Parsing for Domain ScalingabstractState-of-the-art slot filling models for goal-oriented human/machine conversational language understanding systems rely on deep learning methods. While multi-task training of such models alleviates the need for large in-domain annotated datasets, bootstrapping a semantic parsing model for a new domain using only the semantic frame, such as the back-end API or knowledge graph schema, is still one of the holy grail tasks of language understanding for dialogue systems. This paper proposes a deep learning based approach that can utilize only the slot description in context without the need for any labeled or unlabeled in-domain examples, to quickly bootstrap a new domain. The main idea of this paper is to leverage the encoding of the slot names and descriptions within a multi-task deep learned slot filling model, to implicitly align slots across domains. The proposed approach is promising for solving the domain scaling problem and eliminating the need for any manually annotated data or explicit schema alignment. Furthermore, our experiments on multiple domains show that this approach results in significantly better slot-filling performance when compared to using only in-domain data, especially in the low data regime. Ankur Bapna, Gökhan Tür, Dilek Hakkani-Tür, Larry Heck |
INTERSPEECH | 2 |
| 2017 | Sequential Dialogue Context Modeling for Spoken Language UnderstandingabstractSpoken Language Understanding (SLU) is a key component of goal oriented dialogue systems that would parse user utterances into semantic frame representations.Traditionally SLU does not utilize the dialogue history beyond the previous system turn and contextual ambiguities are resolved by the downstream components.In this paper, we explore novel approaches for modeling dialogue context in a recurrent neural network (RNN) based language understanding system.We propose the Sequential Dialogue Encoder Network, that allows encoding context from the dialogue history in chronological order.We compare the performance of our proposed architecture with two context models, one that uses just the previous turn context and another that encodes dialogue context in a memory network, but loses the order of utterances in the dialogue history.Experiments with a multi-domain dialogue dataset demonstrate that the proposed architecture results in reduced semantic frame error rates. Ankur Bapna, Gökhan Tür, Dilek Hakkani-Tür, Larry Heck |
SIGDIAL Conference | 2 |
| 2016 | A New Pre-Training Method for Training Deep Learning Models with Application to Spoken Language UnderstandingabstractWe propose a simple and efficient approach for pre-training deep learning models with application to slot filling tasks in spoken language understanding. The proposed approach leverages unlabeled data to train the models and is generic enough to work with any deep learning model. In this study, we consider the CNN2CRF architecture that contains Convolutional Neural Network (CNN) with Conditional Random Fields (CRF) as top layer, since it has shown great potential for learning useful representations for supervised sequence learning tasks. The proposed pre-training approach with this architecture learns the feature representations from both labeled and unlabeled data at the CNN layer, covering features that would not be observed in limited labeled data. At the CRF layer, the unlabeled data uses predicted classes of words as latent sequence labels together with labeled sequences. Latent labeled sequences, in principle, has the regularization effect on the labeled sequences, yielding a better generalized model. This allows the network to learn representations that are useful for not only slot tagging using labeled data but also learning dependencies both within and between latent clusters of unseen words. The proposed pre-training method with the CRF2CNN architecture achieves significant gains with respect to the strongest semi-supervised baseline. Asli Celikyilmaz, Ruhi Sarikaya, Dilek Hakkani-Tür, Nikhil Ramesh, Gökhan Tür |
INTERSPEECH | 6 |
| 2016 | End-to-End Memory Networks with Knowledge Carryover for Multi-Turn Spoken Language UnderstandingabstractSpoken language understanding (SLU) is a core component of a spoken dialogue system. In the traditional architecture of dialogue systems, the SLU component treats each utterance independent of each other, and then the following components aggregate the multi-turn information in the separate phases. However, there are two challenges: 1) errors from previous turns may be propagated and then degrade the performance of the current turn; 2) knowledge mentioned in the long history may not be carried into the current turn. This paper addresses the above issues by proposing an architecture using end-to-end memory networks to model knowledge carryover in multi-turn conversations, where utterances encoded with intents and slots can be stored as embeddings in the memory and the decoding phase applies an attention model to leverage previously stored semantics for intent prediction and slot tagging simultaneously. The experiments on Microsoft Cortana conversational data show that the proposed memory network architecture can effectively extract salient semantics for modeling knowledge carryover in the multi-turn conversations and outperform the results using the state-of-the-art recurrent neural network framework (RNN) designed for single-turn SLU. Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür, Jianfeng Gao 0001, Li Deng 0001 |
INTERSPEECH | 3 |
| 2016 | Multi-Domain Joint Semantic Frame Parsing Using Bi-Directional RNN-LSTMabstractSequence-to-sequence deep learning has recently emerged as a new paradigm in supervised learning for spoken language understanding. However, most of the previous studies explored this framework for building single domain models for each task, such as slot filling or domain classification, comparing deep learning based approaches with conventional ones like conditional random fields. This paper proposes a holistic multi-domain, multi-task (i.e. slot filling, domain and intent detection) modeling approach to estimate complete semantic frames for all user utterances addressed to a conversational system, demonstrating the distinctive power of deep learning methods, namely bi-directional recurrent neural network (RNN) with long-short term memory (LSTM) cells (RNN-LSTM) to handle such complexity. The contributions of the presented work are three-fold: (i) we propose an RNN-LSTM architecture for joint modeling of slot filling, intent determination, and domain classification; (ii) we build a joint multi-domain model enabling multi-task deep learning where the data from each domain reinforces each other; (iii) we investigate alternative architectures for modeling lexical context in spoken language understanding. In addition to the simplicity of the single model framework, experimental results show the power of such an approach on Microsoft Cortana real user data over alternative methods based on single domain/task deep learning. Dilek Hakkani-Tür, Gökhan Tür, Asli Celikyilmaz, Yun-Nung Chen, Jianfeng Gao 0001, Li Deng 0001, Ye-Yi Wang |
INTERSPEECH | 2 |
| 2016 | Syntax or semantics? knowledge-guided joint semantic frame parsingabstractSpoken language understanding (SLU) is a core component of a spoken dialogue system, which involves intent prediction and slot filling and also called semantic frame parsing. Recently recurrent neural networks (RNN) obtained strong results on SLU due to their superior ability of preserving sequential information over time. Traditionally, the SLU component parses semantic frames for utterances considering their flat structures, as the underlying RNN structure is a linear chain. However, natural language exhibits linguistic properties that provide rich, structured information for better understanding. This paper proposes to apply knowledge-guided structural attention networks (K-SAN), which additionally incorporate non-flat network topologies guided by prior knowledge, to a language understanding task. The model can effectively figure out the salient substructures that are essential to parse the given utterance into its semantic frame with an attention mechanism, where two types of knowledge, syntax and semantics, are utilized. The experiments on the benchmark Air Travel Information System (ATIS) data and the conversational assistant Cortana data show that 1) the proposed K-SAN models with syntax or semantics outperform the state-of-the-art neural network based results, and 2) the improvement for joint semantic frame parsing is more significant, because the structured information provides rich cues for sentence-level understanding, where intent prediction and slot filling can be mutually improved. Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür, Asli Celikyilmaz, Jianfeng Gao 0001, Li Deng 0001 |
SLT | 3 |
| 2016 | Intent detection using semantically enriched word embeddingsabstractState-of-the-art targeted language understanding systems rely on deep learning methods using 1-hot word vectors or off-the-shelf word embeddings. While word embeddings can be enriched with information from semantic lexicons (such as WordNet and PPDB) to improve their semantic representation, most previous research on word-embedding enriching has focused on improving intrinsic word-level tasks such as word analogy and antonym detection. In this work, we enrich word embeddings to force semantically similar or dissimilar words to be closer or farther away in the embedding space to improve the performance of an extrinsic task, namely, intent detection for spoken language understanding. We utilize several semantic lexicons, such as WordNet, PPDB, and Macmillan Dictionary to enrich the word embeddings and later use them as initial representation of words for intent detection. Thus, we enrich embeddings outside the neural network as opposed to learning the embeddings within the network, and, on top of the embeddings, build bidirectional LSTM for intent detection. Our experiments on ATIS and a real log dataset from Microsoft Cortana show that word embeddings enriched with semantic lexicons can improve intent detection. Joo-Kyung Kim, Gökhan Tür, Asli Celikyilmaz, Ye-Yi Wang |
SLT | 2 |
| 2015 | Clustering novel intents in a conversational interaction system with semantic parsingabstractSpoken language understanding (SLU) in today’s conversational systems focuses on recognizing a set of domains, intents, and associated arguments, that are determined by application developers. User requests that are not covered by these are usually directed to search engines, and may remain unhandled. We propose a method that aims to find common user intents amongst these uncovered, out-of-domain utterances, with the goal of supporting future phases of dialog system design. Our approach relies on finding common semantic patterns in uncovered user utterances using an Abstract Meaning Representation based semantic parser. We represent the corpus as a graph and find subgraphs that represent clusters, by pruning the corpus graph according to frequency and entropy. We employ crowd-workers to select and label the resulting clusters and compare resulting clusters with two baselines. Experimental analyses show that we obtain higher coverage and accuracy with the semantic parsing based clustering method. Furthermore, since the intents and candidate slots are already induced, these utterances can also be used in unsupervised SLU modeling. In intent classification experiments, we show that the statistical model trained using the clusters formed by this approach results in higher classification F-measure (showing about 25% relative improvement) in comparison to the alternatives. Dilek Hakkani-Tür, Yun-Cheng Ju, Geoffrey Zweig, Gökhan Tür |
INTERSPEECH | 4 |
| 2015 | Using Recurrent Neural Networks for Slot Filling in Spoken Language UnderstandingabstractSemantic slot filling is one of the most challenging problems in spoken language understanding (SLU). In this paper, we propose to use recurrent neural networks (RNNs) for this task, and present several novel architectures designed to efficiently model past and future temporal dependencies. Specifically, we implemented and compared several important RNN architectures, including Elman, Jordan, and hybrid variants. To facilitate reproducibility, we implemented these networks with the publicly available Theano neural network toolkit and completed experiments on the well-known airline travel information system (ATIS) benchmark. In addition, we compared the approaches on two custom SLU data sets from the entertainment and movies domains. Our results show that the RNN-based models outperform the conditional random field (CRF) baseline by 2% in absolute error reduction on the ATIS benchmark. We improve the state-of-the-art by 0.5% in the Entertainment domain, and 6.7% for the movies domain. Grégoire Mesnil, Yann N. Dauphin, Kaisheng Yao, Yoshua Bengio, Li Deng 0001, Dilek Hakkani-Tür, Xiaodong He 0001, Larry Heck, Gökhan Tür, Dong Yu 0001, Geoffrey Zweig |
IEEE ACM Trans. Audio Speech Lang. Process. | 9 |
| 2014 | Extending domain coverage of language understanding systems via intent transfer between domains using knowledge graphs and search query click logsabstractThis paper proposes a new technique to enable Natural Language Understanding (NLU) systems to handle user queries beyond their original semantic schemas defined by intents and slots. Knowledge graph and search query logs are used to extend NLU system's coverage by transferring intents from other domains to a given domain. The transferred intents as well as existing intents are then applied to a set of new slots that they are not trained with. The knowledge graph and search click logs are used to determine whether the new slots (i.e. entities) or their attributes in the graph can be used together with transfered intents without re-training the underlying NLU models with the expanded (i.e. with new intents and slots) schema. Experimental results show that the proposed technique can in fact be used in extending NLU system's domain coverage in fulfilling the user's request. Ali El-Kahky, Ruhi Sarikaya, Gökhan Tür, Dilek Hakkani-Tür, Larry Heck |
ICASSP | 4 |
| 2014 | A variational Bayesian model for user intent detectionabstractIntent detectors in state-of-the-art spoken language understanding systems are often trained with a small number of manually annotated examples collected from the application domain. Search query logs provide a large number of unlabeled queries that would be beneficial to improve such supervised classification. Furthermore, the contents of user queries as well as the clicked URLs provide information about user's intent. In this paper, we propose a variational Bayesian approach for modeling latent intents of user queries and clicked URLs when available. We use this model to enhance supervised intent classification of user queries from conversational interactions. Experiments were run with large volumes of search queries and show significant improvements over state-of-the-art systems. Yangfeng Ji, Dilek Hakkani-Tür, Asli Celikyilmaz, Larry Heck, Gökhan Tür |
ICASSP | 5 |
| 2014 | Rapidly building domain-specific entity-centric language models using semantic web knowledge sourcesabstractFor domain-specific speech recognition tasks, it is best if the statistical language model component is trained with text data that is content-wise and style-wise similar to the targeted domain for which the application is built. For state-of-the-art language modeling techniques that can be used in real-time within speech recognition engines during first-pass decoding (e.g., N-gram models), the above constraints have to be fulfilled in the training data. However collecting such data, even through crowd sourcing, is expensive and time consuming, and can still be not representative of how a much larger user population would interact with the recognition system. In this paper, we address this problem by employing several semantic web sources that already contain the domain-specific knowledge, such as query click logs and knowledge graphs. We build statistical language models that meet the requirements listed above for domain-specific recognition tasks where natural language is used and the user queries are about name entities in a specific domain. As a case study, in the movies domain where users’ voice queries are movie related, compared to a generic web language model, a language model trained with the above resources not only yields significant perplexity and word-errorrate improvements, but also presents an approach where such language models can be rapidly developed for other domains. Murat Akbacak, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 3 |
| 2014 | Probabilistic enrichment of knowledge graph entities for relation detection in conversational understandingabstractKnowledge encoded in semantic graphs such as Freebase has been shown to benefit semantic parsing and interpretation of natural language user utterances. In this paper, we propose new methods to assign weights to semantic graphs that reflect common usage types of the entities and their relations. Such statistical information can improve the disambiguation of entities in natural language utterances. Weights for entity types can be derived from the populated knowledge in the semantic graph, based on the frequency of occurrence of each type. They can also be learned from the usage frequencies in real world natural language text, such as related Wikipedia documents or user queries posed to a search engine. We compare the proposed methods with the unweighted version of the semantic knowledge graph for the relation detection task and show that all weighting methods result in better performance in comparison to using the unweighted version. Dilek Hakkani-Tür, Asli Celikyilmaz, Larry Heck, Gökhan Tür, Geoffrey Zweig |
INTERSPEECH | 4 |
| 2014 | Segmentation and disfluency removal for conversational speech translationabstractIn this paper we focus on the effect of on-line speech segmentation and disfluency removal methods on conversational speech translation. In a real-time conversational speech to speech translation system, on-line segmentation of speech is required to avoid latency beyond few seconds. While sentential unit segmentation and disfluency removal have been heavily studied mainly for off-line speech processing, to the best of our knowledge, the combined effect of these tasks on conversational speech translation has not been investigated. Furthermore, optimization of performance given maximum allowable system latency to enable a conversation is a newer problem for these tasks. We show that the conventional assumption of doing segmentation followed by disfluency removal is not the best practice. We propose a new approach to do simple-disfluency removal followed by segmentation and then by complex-disfluency removal. The proposed approach shows a significant gain on translation performance of up to 3 Bleu points with only 6 second latency to look ahead, using state-ofthe art machine translation and speech recognition systems. Index Terms: speech translation, disfluency removal, segmentation, sentence units, speech processing Hany Hassan, Lee Schwartz, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 4 |
| 2014 | Detecting out-of-domain utterances addressed to a virtual personal assistantabstractConversational understanding systems, especially virtual personal assistants (VPAs), perform “targeted” natural language understanding, assuming their users stay within the walled gardens of covered domains, and back-off to generic web search otherwise. However, users usually do not know the concept of domains and sometimes simply do not distinguish the system from simple voice search. Hence it becomes an important problem to identify these rejected out-of-domain utterances which are actually intended for the VPA. This paper presents a study tackling this new task, showing that how one utters a request is more important for this task than what is uttered, resembling addressee detection or dialog act tagging. To this end, syntactic and semantic parse “structure” features are extracted in addition to lexical features to train a binary SVM classifier using a large number of random web search queries and VPA utterances from multiple domains. We present controlled experiments leaving one domain out and check the precision of the model when combined with unseen queries. Our results indicate that such structured features result in higher precision especially when the test domain bears little resemblance to the existing domains. Gökhan Tür, Anoop Deoras, Dilek Hakkani-Tür |
INTERSPEECH | 1 |
| 2014 | Deriving local relational surface forms from dependency-based entity embeddings for unsupervised spoken language understandingabstractRecent works showed the trend of leveraging web-scaled structured semantic knowledge resources such as Freebase for open domain spoken language understanding (SLU). Knowledge graphs provide sufficient but ambiguous relations for the same entity, which can be used as statistical background knowledge to infer possible relations for interpretation of user utterances. This paper proposes an approach to capture the relational surface forms by mapping dependency-based contexts of entities from the text domain to the spoken domain. Relational surface forms are learned from dependency-based entity embeddings, which encode the contexts of entities from dependency trees in a deep learning model. The derived surface forms carry functional dependency to the entities and convey the explicit expression of relations. The experiments demonstrate the efficiency of leveraging derived relational surface forms as local cues together with prior background knowledge. Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 3 |
| 2014 | Joint semantic utterance classification and slot filling with recursive neural networksabstractIn recent years, continuous space models have proven to be highly effective at language processing tasks ranging from paraphrase detection to language modeling. These models are distinctive in their ability to achieve generalization through continuous space representations, and compositionality through arithmetic operations on those representations. Examples of such models include feed-forward and recurrent neural network language models. Recursive neural networks (RecNNs) extend this framework by providing an elegant mechanism for incorporating both discrete syntactic structure and continuous-space word and phrase representations into a powerful compositional model. In this paper, we show that RecNNs can be used to perform the core spoken language understanding (SLU) tasks in a spoken dialog system, more specifically domain and intent determination, concurrently with slot filling, in one jointly trained model. We find that a very simple RecNN model achieves competitive performance on the benchmark ATIS task, as well as on a Microsoft Cortana conversational understanding task. Zhaohan Guo, Gökhan Tür, Scott Yih, Geoffrey Zweig |
SLT | 2 |
| 2014 | Personal knowledge graph population from user utterances in conversational understandingabstractKnowledge graphs provide a powerful representation of entities and the relationships between them, but automatically constructing such graphs from spoken language utterances presents the novelty and numerous challenges. In this paper, we introduce a statistical language understanding approach to automatically construct personal (user-centric) knowledge graphs in conversational dialogs. Such information has the potential to better understand the users' requests, fulfilling them, and enabling other technologies such as developing better inferences or proactive interactions. Knowledge encoded in semantic graphs such as Freebase has been shown to benefit semantic parsing and interpretation of natural language utterances. Hence, as a first step, we exploit the personal factual relation triples from Freebase to mine natural language snippets with a search engine, and the resulting snippets containing pairs of related entities to create the training data. This data is then used to build three key language understanding components: (1) Personal Assertion Classification identifies the user utterances that are relevant with personal facts, e.g., “my mother's name is Rosa”; (2) Relation Detection classifies the personal assertion utterance into one of the predefined relation classes, e.g., “parents”; and (3) Slot Filling labels the attributes or arguments of relations, e.g., “name(parents): Rosa”. Our experiments using the Microsoft conversational understanding system demonstrate the performance of this proposed approach on the population of personal knowledge graphs. Xiang Li 0066, Gökhan Tür, Dilek Hakkani-Tür, Qi Li 0014 |
SLT | 2 |
| 2014 | Distributed open-domain conversational understanding framework with domain independent extractorsabstractTraditional spoken dialog systems are usually based on a centralized architecture, in which the number of domains is predefined, and the provider is fixed for a given domain and intent. The spoken language understanding (SLU) component is responsible for detecting domain and intents, and filling domain-specific slots. It is expensive and time-consuming in this architecture to add new and/or competing domains, intents, or providers. The rapid growth of service providers in the mobile computing market calls for an extensible dialog system framework. This paper presents a distributed dialog infrastructure where each domain or provider is agnostic of others, and processes the user utterances independently using their own knowledge or models, so that a new domain and new provider can be easily incorporated in. In addition, to facilitate each service provider building their own SLU models or algorithms, we introduce a new component, extractors, to provide intermediate semantic annotations such as entity mention tags, which can be plugged in arbitrarily as well. Each service provider can then rapidly develop their SLU parser with minimum efforts by providing some example sentences with intents and slots if needed. Our preliminary experimental results demonstrate the power of this new framework compared to a centralized architecture. Qi Li 0014, Gökhan Tür, Dilek Hakkani-Tür, Xiang Li 0066, Tim Paek, Asela Gunawardana, Chris Quirk |
SLT | 2 |
| 2013 | Semi-Supervised Semantic Tagging of Conversational Understanding using Markov Topic Regression
Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür, Ruhi Sarikaya |
ACL (1) | 3 |
| 2013 | Using a knowledge graph and query click logs for unsupervised learning of relation detectionabstractIn this paper, we introduce a novel statistical language understanding paradigm inspired by the emerging semantic web: Instead of building models for the target application, we propose relying on the semantic space already defined and populated in the knowledge graph for the target domain. As a first step towards this direction, we present unsupervised methods for training relation detection models exploiting the semantic knowledge graphs of the semantic web. The detected relations are used to mine natural language queries against a back-end knowledge base. For each relation, we leverage the complete set of entities that are connected to each other in the graph with the specific relation, and search these entity pairs on the web. We use the snippets that the search engine returns to create natural language examples that can be used as the training data for each relation. We further refine the annotations of these examples using the knowledge graph itself and iterate using a bootstrap approach. Furthermore, we explot the URLs returned for these pairs by the search engine to mine additional examples from the search engine query click logs. In our experiments, we show that, we can achieve relation detection models that perform about 60% macro F-measure on the relations that are in the knowledge graph without any manual labeling, resulting in a comparable performance with supervised training. Dilek Hakkani-Tür, Larry Heck, Gökhan Tür |
ICASSP | 3 |
| 2013 | Multi-style adaptive training for robust cross-lingual spoken language understandingabstractGiven the increasingly available machine translation (MT) services nowadays, one efficient strategy for cross-lingual spoken language understanding (SLU) is to first translate the input utterance from the second language into the primary language, and then call the primary language SLU system to decode the semantic knowledge. However, errors introduced in the MT process create a condition similar to the “mismatch” condition encountered in robust speech recognition. Such mismatch makes the performance of cross-lingual SLU far from acceptable. Motivated by successful solutions developed in robust speech recognition, we in this paper propose a multi-style adaptive training method to improve the robustness of the SLU system for cross-lingual SLU tasks. For evaluation, we created an English-Chinese bilingual ATIS database, and then carried out a series of experiments on that database to experimentally assess the proposed methods. Experimental results show that, without relying on any data in the second language, the proposed method significantly improves the performance on a cross-lingual SLU task while producing no degradation for input in the primary language. This greatly facilitates porting SLU to as many languages as there are MT systems without any human effort. We further study the robustness of this approach to another type of mismatch condition, caused by speech recognition errors, and demonstrate its success also. Xiaodong He 0001, Li Deng 0001, Dilek Hakkani-Tür, Gökhan Tür |
ICASSP | 4 |
| 2013 | Latent semantic modeling for slot filling in conversational understandingabstractIn this paper, we propose a new framework for semantic template filling in a conversational understanding (CU) system. Our method decomposes the task into two steps: latent n-gram clustering using a semi-supervised latent Dirichlet allocation (LDA) and sequence tagging for learning semantic structures in a CU system. Latent semantic modeling has been investigated to improve many natural language processing tasks such as syntactic parsing or topic tracking. However, due to several complexity problems caused by issues involving utterance length or dialog corpus size, it has not been analyzed directly for semantic parsing tasks. In this paper, we propose extending the LDA by introducing prior knowledge we obtain from semantic knowledge bases. Then, the topic posteriors obtained from the new LDA model are used as additional constraints to a sequence learning model for the semantic template filling task. The experimental results show significant performance gains on semantic slot filling models when features from latent semantic models are used in a conditional random field (CRF). Gökhan Tür, Asli Celikyilmaz, Dilek Hakkani-Tür |
ICASSP | 1 |
| 2013 | Understanding computer-directed utterances in multi-user dialog systemsabstractThis work aims to understand user requests when multiple users are interacting with each other and a spoken dialog system. More specifically, we explore the use of multi-human conversational context to improve domain detection in a human-computer interaction system. We investigate the different effects of human-directed context and computer-directed context, and compare the impact of using different context window sizes. Furthermore, we employ topic segmentation to chunk conversations for determining context boundaries. The experimental results show that the use of conversational context helps reduce domain detection error rate, especially in some specific domains. And though computer directed context is more reliable, the results show that the combination of both computer and human addressed utterances within a reasonable window size performs the best. Dilek Hakkani-Tür, Gökhan Tür |
ICASSP | 3 |
| 2013 | IsNL? a discriminative approach to detect natural language like queries for conversational understandingabstractWhile data-driven methods for spoken language understanding (SLU) provide state of the art performances and reduce maintenance and model adaptation costs compared to handcrafted parsers, the collection and annotation of domain-specific natural language utterances for training remains a time-consuming task. A recent line of research has focused on enriching the training data with in-domain utterances by mining search engine query logs to improve the SLU tasks. However genre mismatch is a big obstacle as search queries are typically keywords. In this paper, we present an efficient discriminative binary classification method that filters large collection of online web search queries only to select the natural language like queries. The training data used to build this classifier is mined from search query click logs, represented as a bipartite graph. Starting from queries which contain natural language salient phrases, random graph walk algorithms are employed to mine corresponding keyword queries. Then an active learning method is employed for quickly improving on top of this automatically mined data. The results show that our method is robust to noise in search queries by improving over a baseline model previously used for SLU data collection. We also show the effectiveness of detected natural language like queries in extrinsic evaluations on domain detection and slot filling tasks. Asli Celikyilmaz, Gökhan Tür, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2013 | A weakly-supervised approach for discovering new user intents from search query logsabstractState-of-the art spoken language understanding models that automatically capture user intents in human to machine dialogs are trained with manually annotated data, which is cumbersome and time-consuming to prepare. For bootstrapping the learning algorithm that detects relations in natural language queries to a conversational system, one can rely on publicly available knowledge graphs, such as Freebase, and mine corresponding data from the web. In this paper, we present an unsupervised approach to discover new user intents using a novel Bayesian hierarchical graphical model. Our model employs search query click logs to enrich the information extracted from bootstrapped models. We use the clicked URLs as implicit supervision and extend the knowledge graph based on the relational information discovered from this model. The posteriors from the graphical model relate the newly discovered intents with the search queries. These queries are then used as additional training examples to complement the bootstrapped relation detection models. The experimental results demonstrate the effectiveness of this approach, showing extended coverage to new intents without impacting the known intents. Index Terms: spoken language understanding, graphical models, search query click logs, intent discovery. Dilek Hakkani-Tür, Asli Celikyilmaz, Larry Heck, Gökhan Tür |
INTERSPEECH | 4 |
| 2013 | Leveraging knowledge graphs for web-scale unsupervised semantic parsingabstractThe past decade has seen the emergence of web-scale structured and linked semantic knowledge resources (e.g., Freebase, DB-Pedia). These semantic knowledge graphs provide a scalable “schema for the web”, representing a significant opportunity for the spoken language understanding (SLU) research community. This paper leverages these resources to bootstrap a web-scale semantic parser with no requirement for semantic schema de-sign, no data collection, and no manual annotations. Our ap-proach is based on an iterative graph crawl algorithm. From an initial seed node (entity-type), the method learns the related entity-types from the graph structure, and automatically anno-tates documents that can be linked to the node (e.g., Wikipedia articles, web search documents). Following the branches, the graph is crawled and the procedure is repeated. The resulting collection of annotated documents is used to bootstrap web-scale conditional random field (CRF) semantic parsers. Finally, we use a maximum-a-posteriori (MAP) unsupervised adapta-tion technique on sample data from a specific domain to refine the parsers. The scale of the unsupervised parsers is on the order of thousands of domains and entity-types, millions of entities, and hundreds of millions of relations. The precision-recall of the semantic parsers trained with our unsupervised method ap-proaches those trained with supervised annotations. Index Terms: semantic parsing, semantic web, semantic search, dialog, natural language understanding Larry Heck, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 3 |
| 2013 | Semantic parsing using word confusion networks with conditional random fieldsabstractA challenge in large vocabulary spoken language understand-ing (SLU) is robustness to automatic speech recognition (ASR) errors. The state of the art approaches for semantic parsing rely on using discriminative sequence classification methods, such as conditional random fields (CRFs). Most dialog systems em-ploy a cascaded approach where the best hypotheses from the ASR system are fed into the following SLU system. In our pre-vious work, we have proposed the use of lattices towards joint recognition and parsing. In this paper, extending this idea, we propose to exploit word confusion networks (WCNs), compiled from ASR lattices for both CRF modeling and decoding. WCNs provide a compact representation of multiple aligned ASR hy-potheses, without compromising recognition accuracy. For slot filling, we show significant semantic parsing performance im-provements using WCNs compared to ASR 1-best output, ap-proximating the oracle path performance. Index Terms: conditional random field, semantic parsing, word confusion network, natural language understanding Gökhan Tür, Anoop Deoras, Dilek Hakkani-Tür |
INTERSPEECH | 1 |
| 2013 | Joint Discriminative Decoding of Words and Semantic Tags for Spoken Language UnderstandingabstractMost Spoken Language Understanding (SLU) systems today employ a cascade approach, where the best hypothesis from Automatic Speech Recognizer (ASR) is fed into understanding modules such as slot sequence classifiers and intent detectors. The output of these modules is then further fed into downstream components such as interpreter and/or knowledge broker. These statistical models are usually trained individually to optimize the error rate of their respective output. In such approaches, errors from one module irreversibly propagates into other modules causing a serious degradation in the overall performance of the SLU system. Thus it is desirable to jointly optimize all the statistical models together. As a first step towards this, in this paper, we propose a joint decoding framework in which we predict the optimal word as well as slot sequence (semantic tag sequence) jointly given the input acoustic stream. Furthermore, the improved recognition output is then used for an utterance classification task, specifically, we focus on intent detection task. On a SLU task, we show 1.5% absolute reduction (7.6% relative reduction) in word error rate (WER) and 1.2% absolute improvement in F measure for slot prediction when compared to a very strong cascade baseline comprising of state-of-the-art large vocabulary ASR followed by conditional random field (CRF) based slot sequence tagger. Similarly, for intent detection, we show 1.2% absolute reduction (12% relative reduction) in classification error rate. Anoop Deoras, Gökhan Tür, Ruhi Sarikaya, Dilek Hakkani-Tür |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Translating natural language utterances to search queries for SLU domain detection using query click logsabstractLogs of user queries from a search engine (such as Bing or Google) together with the links clicked provide valuable implicit feedback to improve statistical spoken language understanding (SLU) models. However, the form of natural language utterances occurring in spoken interactions with a computer differs stylistically from that of keyword search queries. In this paper, we propose a machine translation approach to learn a mapping from natural language utterances to search queries. We train statistical translation models, using task and domain independent semantically equivalent natural language and keyword search query pairs mined from the search query click logs. We then extend our previous work on enriching the existing classification feature sets for input utterance domain detection with features computed using the click distribution over a set of clicked URLs from search engine query click logs of user utterances with automatically translated queries. This approach results in significant improvements for domain detection, especially when detecting the domains of user utterances that are formulated as natural language queries and effectively complements to the earlier work using syntactic transformations. Dilek Hakkani-Tür, Gökhan Tür, Rukmini Iyer, Larry Heck |
ICASSP | 2 |
| 2012 | Towards deeper understanding: Deep convex networks for semantic utterance classificationabstractFollowing the recent advances in deep learning techniques, in this paper, we present the application of special type of deep architecture - deep convex networks (DCNs) - for semantic utterance classification (SUC). DCNs are shown to have several advantages over deep belief networks (DBNs) including classification accuracy and training scalability. However, adoption of DCNs for SUC comes with non-trivial issues. Specifically, SUC has an extremely sparse input feature space encompassing a very large number of lexical and semantic features. This is about a few thousand times larger than the feature space for acoustic modeling, yet with a much smaller number of training samples. Experimental results we obtained on a domain classification task for spoken language understanding demonstrate the effectiveness of DCNs. The DCN-based method produces higher SUC accuracy than the Boosting-based discriminative classifier with word trigrams. Gökhan Tür, Li Deng 0001, Dilek Hakkani-Tür, Xiaodong He 0001 |
ICASSP | 1 |
| 2012 | Joint Decoding for Speech Recognition and Semantic TaggingabstractMost conversational understanding (CU) systems today employ a cascade approach, where the best hypothesis from automatic speech recognizer (ASR) is fed into spoken language understanding (SLU) module, whose best hypothesis is then fed into other systems such as interpreter or dialog manager. In such approaches, errors from one statistical module irreversibly propagates into another module causing a serious degradation in the overall performance of the conversational understanding system. Thus it is desirable to jointly optimize all the statistical modules together. As a first step towards this, in this paper, we propose a joint decoding framework in which we predict the optimal word as well as slot (semantic tag) sequence jointly given the input acoustic stream. On Microsoft’s CU system, we show 1.3 % absolute reduction in word error rate (WER) and 1.2% absolute improvement in F measure for slot prediction when compared to a very strong cascade baseline comprising of the state-of-the-art recognizer followed by a slot sequence tagger. Anoop Deoras, Ruhi Sarikaya, Gökhan Tür, Dilek Hakkani-Tür |
INTERSPEECH | 3 |
| 2012 | A Discriminative Classification-Based Approach to Information State Updates for a Multi-Domain Dialog SystemabstractWe propose a discriminative classification approach for updating the current information state of a multi-domain dialog system based on user responses. Our method uses a set of lexical and domain independent features to compare the spoken language understanding (SLU) output for the current user turn with the previous information state. We then update the information state accordingly, employing a discriminative machine learning approach. Using a data set collected from our conversational interaction system, we investigate the impact of features based on context dependent and context independent SLU tagging schemas. We show that the proposed approach outperforms two non-trivial baselines, one based on manually crafted rules and the other on classification with lexical features alone. Furthermore, such an approach allows the addition of new domains to the dialog manager in a seamless way. Dilek Hakkani-Tür, Gökhan Tür, Larry Heck, Ashley Fidler, Asli Celikyilmaz |
INTERSPEECH | 2 |
| 2012 | Exploiting the Semantic Web for Unsupervised Natural Language Semantic Parsing
Gökhan Tür, Minwoo Jeong, Ye-Yi Wang, Dilek Hakkani-Tür, Larry Heck |
INTERSPEECH | 1 |
| 2012 | Statistical semantic interpretation modeling for spoken language understanding with enriched semantic featuresabstractIn natural language human-machine statistical dialog systems, semantic interpretation is a key task typically performed following semantic parsing, and aims to extract canonical meaning representations of semantic components. In the literature, usually manually built rules are used for this task, even for implicitly mentioned non-named semantic components (like genre of a movie or price range of a restaurant). In this study, we present statistical methods for modeling interpretation, which can also benefit from semantic features extracted from large in-domain knowledge sources. We extract features from user utterances using a semantic parser and additional semantic features from textual sources (online reviews, synopses, etc.) using a novel tree clustering approach, to represent unstructured information that correspond to implicit semantic components related to targeted slots in the user's utterances. We evaluate our models on a virtual personal assistance system and demonstrate that our interpreter is effective in that it does not only improve the utterance interpretation in spoken dialog systems (reducing the interpretation error rate by 36% relative compared to a language model baseline), but also unveils hidden semantic units that are otherwise nearly impossible to extract from purely manual lexical features that are typically used in utterance interpretation. Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 3 |
| 2012 | Use of kernel deep convex networks and end-to-end learning for spoken language understandingabstractWe present our recent and ongoing work on applying deep learning techniques to spoken language understanding (SLU) problems. The previously developed deep convex network (DCN) is extended to its kernel version (K-DCN) where the number of hidden units in each DCN layer approaches infinity using the kernel trick. We report experimental results demonstrating dramatic error reduction achieved by the K-DCN over both the Boosting-based baseline and the DCN on a domain classification task of SLU, especially when a highly correlated set of features extracted from search query click logs are used. Not only can DCN and K-DCN be used as a domain or intent classifier for SLU, they can also be used as local, discriminative feature extractors for the slot filling task of SLU. The interface of K-DCN to slot filling systems via the softmax function is presented. Finally, we outline an end-to-end learning strategy for training the softmax parameters (and potentially all DCN and K-DCN parameters) where the learning objective can take any performance measure (e.g. the F-measure) for the full SLU system. Li Deng 0001, Gökhan Tür, Xiaodong He 0001, Dilek Hakkani-Tür |
SLT | 2 |
| 2011 | Exploiting distance based similarity in topic models for user intent detectionabstractOne of the main components of spoken language understanding is intent detection, which allows user goals to be identified. A challenging sub-task of intent detection is the identification of intent bearing phrases from a limited amount of training data, while maintaining the ability to generalize well. We present a new probabilistic topic model for jointly identifying semantic intents and common phrases in spoken language utterances. Our model jointly learns a set of intent dependent phrases and captures semantic intent clusters as distributions over these phrases based on a distance dependent sampling method. This sampling method uses proximity of words utterances when assigning words to latent topics. We evaluate our method on labeled utterances and present several examples of discovered semantic units. We demonstrate that our model outperforms standard topic models based on bag-of-words assumption. Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür, Ashley Fidler, Dustin Hillard |
ASRU | 3 |
| 2011 | Employing web search query click logs for multi-domain spoken language understandingabstractLogs of user queries from a search engine (such as Bing or Google) together with the links clicked provide valuable implicit feedback to improve statistical spoken language understanding (SLU) models. In this work, we propose to enrich the existing classification feature set for domain detection with features computed using the click distribution over a set of clicked URLs from search query click logs (QCLs) of user utterances. Since the form of natural language utterances differs stylistically from that of keyword search queries, to be able to match natural language utterances with related search queries, we perform a syntax-based transformation of the original utterances, after filtering out domain-independent salient phrases. This approach results in significant improvements for domain detection, especially when detecting the domains of web-related user utterances. Dilek Hakkani-Tür, Gökhan Tür, Larry Heck, Asli Celikyilmaz, Ashley Fidler, Dustin Hillard, Rukmini Iyer, Sarangarajan Parthasarathy |
ASRU | 2 |
| 2011 | Exploiting query click logs for utterance domain detection in spoken language understandingabstractIn this paper, we describe methods to exploit search queries mined from search engine query logs to improve domain detection in spoken language understanding. We propose extending the label propagation algorithm, a graph-based semi-supervised learning approach, to incorporate noisy domain information estimated from search engine links the users click following their queries. The main contributions of our work are the use of search query logs for domain classification, integration of noisy supervision into the semi-supervised label propagation algorithm, and sampling of high-quality query click data by mining query logs and using classification confidence scores. We show that most semi-supervised learning methods we experimented with improve the performance of the supervised training, and the biggest improvement is achieved by label propagation that uses noisy supervision. We reduce the to error rate of domain detection by 20% relative, from 6.2% to 5.0%. Dilek Hakkani-Tür, Larry Heck, Gökhan Tür |
ICASSP | 3 |
| 2011 | Sentence simplification for spoken language understandingabstractIn this paper, we present a sentence simplification method and demonstrate its use to improve intent determination and slot filling tasks in spoken language understanding (SLU) systems. This research is motivated by the observation that, while current statistical SLU models usually perform accurately for simple, well-formed sentences, error rates increase for more complex, longer, more natural or spontaneous utterances. Furthermore, users familiar with web search usually formulate their information requests as a keyword search query, suggesting that frameworks which can handle both forms of inputs is required. We propose a dependency parsing-based sentence simplification approach that extracts a set of keywords from natural language sentences and uses those in addition to entire utterances for completing SLU tasks. We evaluated this approach using the well studied ATIS corpus with manual and automatic transcriptions and observed significant error reductions for both intent determination (30% relative) and slot filling (15% relative) tasks over the state-of the-art performances. Gökhan Tür, Dilek Hakkani-Tür, Larry Heck, Sarangarajan Parthasarathy |
ICASSP | 1 |
| 2011 | Approximate Inference for Domain Detection in Spoken Language UnderstandingabstractThis paper presents a semi-latent topic model for semantic domain detection in spoken language understanding systems. We use labeled utterance information to capture latent topics, which directly correspond to semantic domains. Additionally, we introduce an ’informative prior ’ for Bayesian inference that can simultaneously segment utterances of known domains into classes and divide them from out-of-domain utterances. We show that our model generalizes well on the task of classify-ing spoken language utterances and compare its results to those of an unsupervised topic model, which does not use labeled in-formation. Index Terms: spoken language understanding, generative mod-els, gibbs sampling. Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 3 |
| 2011 | Bootstrapping Domain Detection Using Query Click Logs for New DomainsabstractDomain detection in spoken dialog systems is usually treated as a multi-class, multi-label classification problem, and training of domain classifiers requires collection and manual annotation of example utterances. In order to extend a dialog system to new domains in a way that is seamless for users, domain detection should be able to handle utterances from the new domain as soon as it is introduced. In this work, we propose using web search query logs, which include queries entered by users and the links they subsequently click on, to bootstrap domain detection for new domains. While sampling user queries from the query click logs to train new domain classifiers, we introduce two types of measures based on the behavior of the users who entered a query and the form of the query. We show that both types of measures result in reductions in the error rate as compared to randomly sampling training queries. In controlled experiments over five domains, we achieve the best gain from the combination of the two types of sampling criteria. Dilek Hakkani-Tür, Gökhan Tür, Larry Heck, Elizabeth Shriberg |
INTERSPEECH | 2 |
| 2011 | Learning Weighted Entity Lists from Web Click Logs for Spoken Language UnderstandingabstractNamed entity lists provide important features for language understanding, but typical lists can contain many ambiguous or incorrect phrases. We present an approach for automatically learning weighted entity lists by mining user clicks from web search logs. The approach significantly outperforms multiple baseline approaches and the weighted lists improve spoken language understanding tasks such as domain detection and slot filling. Our methods are general and can be easily applied to large quantities of entities, across any number of lists. Index Terms: spoken language understanding, domain detection, slot filling, named entity lists, click logs Dustin Hillard, Asli Celikyilmaz, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 4 |
| 2011 | Multi-Task Learning for Spoken Language Understanding with Shared SlotsabstractThis paper addresses the problem of learning multiple spoken language understanding (SLU) tasks that have overlapping sets of slots. In such a scenario, it is possible to achieve better slot filling performance by learning multiple tasks simultaneously, as opposed to learning them independently. We focus on pre-senting a number of simple multi-task learning algorithms for slot filling systems based on semi-Markov CRFs, assuming the knowledge of shared slots. Furthermore, we discuss an intra-domain clustering method that automatically discovers shared slots from training data. The effectiveness of our proposed ap-proaches is demonstrated in an SLU application that involves three different yet related tasks. Index Terms: spoken language understanding, slot filling, multi-task learning Xiao Li 0006, Ye-Yi Wang, Gökhan Tür |
INTERSPEECH | 3 |
| 2011 | Towards Unsupervised Spoken Language Understanding: Exploiting Query Click Logs for Slot FillingabstractIn this paper, we present a novel approach to exploit user queries mined from search engine query click logs to bootstrap or improve slot filling models for spoken language understanding. We propose extending the earlier gazetteer population techniques to mine unannotated training data for semantic parsing. The automatically annotated mined data can then be used to train slot specific parsing models. We show that this method can be used to bootstrap slot filling models and can be combined with any available annotated data to improve performance. Furthermore, this approach may eliminate the need for populating and maintaining in-domain gazetteers, in addition to providing complementary information if they are already available. Index Terms: spoken language understanding, slot filling, data mining, named entity extraction, unsupervised learning Gökhan Tür, Dilek Hakkani-Tür, Dustin Hillard, Asli Celikyilmaz |
INTERSPEECH | 1 |
| 2011 | Towards spoken clinical-question answering: evaluating and adapting automatic speech-recognition systems for spoken clinical questionsabstractOBJECTIVE: To evaluate existing automatic speech-recognition (ASR) systems to measure their performance in interpreting spoken clinical questions and to adapt one ASR system to improve its performance on this task. DESIGN AND MEASUREMENTS: The authors evaluated two well-known ASR systems on spoken clinical questions: Nuance Dragon (both generic and medical versions: Nuance Gen and Nuance Med) and the SRI Decipher (the generic version SRI Gen). The authors also explored language model adaptation using more than 4000 clinical questions to improve the SRI system's performance, and profile training to improve the performance of the Nuance Med system. The authors reported the results with the NIST standard word error rate (WER) and further analyzed error patterns at the semantic level. RESULTS: Nuance Gen and Med systems resulted in a WER of 68.1% and 67.4% respectively. The SRI Gen system performed better, attaining a WER of 41.5%. After domain adaptation with a language model, the performance of the SRI system improved 36% to a final WER of 26.7%. CONCLUSION: Without modification, two well-known ASR systems do not perform well in interpreting spoken clinical questions. With a simple domain adaptation, one of the ASR systems improved significantly on the clinical question task, indicating the importance of developing domain/genre-specific ASR systems. Gökhan Tür, Dilek Hakkani-Tür, Hong Yu 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2010 | Automatic disfluency removal for improving spoken language translationabstractStatistical machine translation (SMT) systems for spoken languages suffer from conversational speech phenomena, in particular, the presence of speech disfluencies. We examine the impact of disfluencies from broadcast conversation data on our hierarchical phrase-based SMT system and implement automatic disfluency removal approaches for cleansing the MT input. We evaluate the efficacy of proposed approaches and investigate the impact of disfluency removal on SMT performance across different disfluency types. We show that for translating Mandarin broadcast conversational transcripts into English, our automatic disfluency removal approaches could produce significant improvement in BLEU and TER. Wen Wang 0001, Gökhan Tür, Jing Zheng 0001, Necip Fazil Ayan |
ICASSP | 2 |
| 2010 | Speech-based automated cognitive status assessmentabstractVerbal interviews performed by trained clinicians are a common form of assessments to measure cognitive decline. The aim in this paper is to study the usability of automated methods for evaluating verbal cognitive status assessment tests for the elderly. If reliable, such methods for cognitive assessment can be used for frequent, non-intrusive, low-cost screenings and provide objective and longitudinal cognitive status monitoring data that can complement regular clinical visits and would be useful for early detection of conditions associated with language and communication impairments. This study focuses on two types of tests: a story-recall test, used for memory and language functioning assessment, and a picture description test, used to assess the information content in speech. A data collection was designed for this study involving recordings of about 100 people, mostly over 70 years old, performing these tests. The speech samples were manually transcribed and annotated with semantic units in order to obtain manual evaluation scores. We explore the use of automatic speech recognition and language processing methods to derive objective, automatically extracted metrics of cognitive status that are highly correlated with the manual scores. We use recall and precision based metrics based on semantic content units associated with the tests. Our experiments show high correlation between manually obtained scores and the automatic metrics obtained using either manual or automatic speech transcriptions. Index Terms: speech recognition, language processing, automated cognitive status assessment, elderly speech Dilek Hakkani-Tür, Dimitra Vergyri, Gökhan Tür |
INTERSPEECH | 3 |
| 2010 | Domain adaptation and compensation for emotion detectionabstractInspired by the recent improvements in domain adapta-tion and session variability compensation techniques used for speech and speaker processing, we study their effect for emo-tion prediction. More specifically, we investigated the use of publicly available out-of-domain data with emotion annotations for improving the performance of the in-domain model trained using 911 emergency-hotline calls. Following the emotion de-tection literature, we use prosodic (pitch, energy, and speaking rate) features as the inputs to a discriminative classifier. We performed segment-level n-fold cross validation emotion pre-diction experiments. Our results indicate significant improve-ment of performance for emotion prediction exploiting out-of-domain data. Index Terms: emotion detection, domain adaptation 1. Michelle Hewlett Sanchez, Gökhan Tür, Luciana Ferrer, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2010 | Social role discovery from spoken language using dynamic Bayesian networksabstractIn this paper, we focus on inferring social roles in con-versations using information extracted only from the speaking styles of the speakers. We model the turn-taking behavior of the speakers with dynamic Bayesian networks (DBNs), which provide the capability of naturally formulating the dependen-cies between random variables. More specifically, we first ex-plore the usefulness of a simple DBN, namely, a hidden Markov model (HMM), for this problem. As it turns out, the knowl-edge of the segments that belong to the same speaker can be augmented into this HMM structure, which results in a more sophisticated DBN. This information places a constraint on two subsequent speaker roles such that the current speaker role de-pends not only on the previous speaker’s role but also on that most recent role assigned to the same speaker. We conducted an experimental study to compare these two modeling approaches using broadcast shows. In our experiments, the approach with the constraint on same speaker segments assigned 89.5 % turns the correct role whereas the HMM-based approach assigned 79.2 % of turns their correct role. Index Terms: Social role discovery, speaker turn detection, spoken language understanding Sibel Yaman, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 3 |
| 2010 | What is left to be understood in ATIS?abstractOne of the main data resources used in many studies over the past two decades for spoken language understanding (SLU) research in spoken dialog systems is the airline travel information system (ATIS) corpus. Two primary tasks in SLU are intent determination (ID) and slot filling (SF). Recent studies reported error rates below 5% for both of these tasks employing discriminative machine learning techniques with the ATIS test set. While these low error rates may suggest that this task is close to being solved, further analysis reveals the continued utility of ATIS as a research corpus. In this paper, our goal is not experimenting with domain specific techniques or features which can help with the remaining SLU errors, but instead exploring methods to realize this utility via extensive error analysis. We conclude that even with such low error rates, ATIS test set still includes many unseen example categories and sequences, hence requires more data. Better yet, new annotated larger data sets from more complex tasks with realistic utterances can avoid over-tuning in terms of modeling and feature design. We believe that advancements in SLU can be achieved by having more naturally spoken data sets and employing more linguistically motivated features while preserving robustness due to speech recognition noise and variance due to natural language. Gökhan Tür, Dilek Hakkani-Tür, Larry Heck |
SLT | 1 |
| 2010 | Cascaded model adaptation for dialog act segmentation and tagging
Ümit Güz, Gökhan Tür, Dilek Hakkani-Tür, Sébastien Cuendet |
Comput. Speech Lang. | 2 |
| 2010 | Multi-View Semi-Supervised Learning for Dialog Act Segmentation of SpeechabstractSentence segmentation of speech aims at determining sentence boundaries in a stream of words as output by the speech recognizer. Typically, statistical methods are used for sentence segmentation. However, they require significant amounts of labeled data, preparation of which is time-consuming, labor-intensive, and expensive. This work investigates the application of multi-view semi-supervised learning algorithms on the sentence boundary classification problem by using lexical and prosodic information. The aim is to find an effective semi-supervised machine learning strategy when only small sets of sentence boundary-labeled data are available. We especially focus on two semi-supervised learning approaches, namely, self-training and co-training. We also compare different example selection strategies for co-training, namely, agreement and disagreement. Furthermore, we propose another method, called self-combined, which is a combination of self-training and co-training. The experimental results obtained on the ICSI Meeting (MRDA) Corpus show that both multi-view methods outperform self-training, and the best results are obtained using co-training alone. This study shows that sentence segmentation is very appropriate for multi-view learning since the data sets can be represented by two disjoint and redundantly sufficient feature sets, namely, using lexical and prosodic information. Performance of the lexical and prosodic models is improved by 26% and 11% relative, respectively, when only a small set of manually labeled examples is used. When both information sources are combined, the semi-supervised learning methods improve the baseline F-Measure of 69.8% to 74.2%. Ümit Güz, Sébastien Cuendet, Dilek Hakkani-Tür, Gökhan Tür |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | The CALO Meeting Assistant SystemabstractThe CALO Meeting Assistant (MA) provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper presents the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, topic identification and segmentation, question-answer pair identification, action item recognition, decision extraction, and summarization. Gökhan Tür, Andreas Stolcke, L. Lynn Voss, Stanley Peters, Dilek Hakkani-Tür, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Who, What, When, Where, Why? Comparing Multiple Approaches to the Cross-Lingual 5W Task
Kristen Parton, Kathy McKeown, Bob Coyne, Mona T. Diab, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Heng Ji 0001, Wei-Yun Ma, Adam Meyers 0001, Sara Stolbach, Ang Sun, Gökhan Tür, Wei Xu 0004, Sibel Yaman |
ACL/IJCNLP | 13 |
| 2009 | Co-adaptation: Adaptive co-training for semi-supervised learningabstractInspired by popular co-training and domain adaptation methods, we propose a co-adaptation algorithm. The goal is improving the performance of a dialog act segmentation model by exploiting the vast amount of unlabeled data. This task provides a nice framework for multiview learning, as it has been shown that lexical and prosodic features provide complementary information. Instead of simply adding machine-labeled data to the set of manually labeled data, co-adaptation technique adapts the existing models. While both co-training and domain adaptation techniques have been employed for dialog act segmentation, our experiments show that the proposed co-adaptation algorithm results in significantly better performance. Gökhan Tür |
ICASSP | 1 |
| 2009 | Exploiting user feedback for language model adaptation in meeting recognitionabstractWe investigate language model (LM) adaptation in a meeting recognition application, where the LM is adapted based on recognition output from relevant prior meetings and partial manual corrections. Unlike previous work, which has considered either completely unsupervised or supervised adaptation, we investigate a scenario where a human (e.g., a meeting participant) can correct some of the recognition mistakes. We find that recognition accuracy using the adapted LM can be enhanced substantially by partial correction. In particular, if all content words (about half of all recognition errors) are corrected, recognition improves to the same accuracy as if completely error-free (manually created) transcriptions had been used for adaptation. We also compare and combine a variety of adaptation methods, including linear interpolation, unigram marginal adaptation, and a discriminative method based on ldquopositiverdquo and ldquonegativerdquo N-grams. Dimitra Vergyri, Andreas Stolcke, Gökhan Tür |
ICASSP | 3 |
| 2009 | Combining semantic and syntactic information sources for 5-w question answeringabstractThis paper focuses on combining answers generated by a semantic parser that produces semantic role labels (SRLs) and those generated by syntactic parser that produces function tags for answering 5-W questions, i.e., who, what, when, where, and why. We take a probabilistic approach in which a system’s ability to correctly answer 5-W questions is measured with the likelihood that its answers are produced for the given word sequence. This is achieved by training statistical language models (LMs) that are used to predict whether the answers returned by semantic parse or those returned by the syntactic parser are more likely. We evaluated our approach using the OntoNotes dataset. Our experimental results indicate that the proposed LM-based combination strategy was able to improve the performance of the best individual system in terms of both F1 measure and accuracy. Furthermore, the error rates for each question type were also significantly reduced with the help of the proposed approach. Sibel Yaman, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 3 |
| 2009 | Classification-based strategies for combining multiple 5-w question answering systemsabstractWe describe and analyze inference strategies for combining outputs from multiple question answering systems each of which was developed independently. Specifically, we address the DARPA-funded GALE information distillation Year 3 task of finding answers to the 5-Wh questions (who, what, when, where, and why) for each given sentence. The approach we take revolves around determining the best system using discriminative learning. In particular, we train support vector machines with a set of novel features that encode systems’ capabilities of returning as many correct answers as possible. We analyze two combination strategies: one combines multiple systems at the granularity of sentences, and the other at the granularity of individual fields. Our experimental results indicate that the proposed features and combination strategies were able to improve the overall performance by 22% to 36% relative to a random selection, 16% to 35% relative to a majority voting scheme, and 15% to 23% relative to the best individual system. Index Terms: Question answering, Systems for spoken language understanding Sibel Yaman, Dilek Hakkani-Tür, Gökhan Tür, Ralph Grishman, Mary P. Harper, Kathy McKeown, Adam Meyers 0001, Kartavya Sharma |
INTERSPEECH | 3 |
| 2009 | IXIR: A statistical information distillation system
Michael Levit, Dilek Hakkani-Tür, Gökhan Tür, Daniel Gillick |
Comput. Speech Lang. | 3 |
| 2009 | Generative and Discriminative Methods Using Morphological Information for Sentence Segmentation of TurkishabstractThis paper presents novel methods for generative, discriminative, and hybrid sequence classification for segmentation of Turkish word sequences into sentences. In the literature, this task is generally solved using statistical models that take advantage of lexical information among others. However, Turkish has a productive morphology that generates a very large vocabulary, making the task much harder. In this paper, we introduce a new set of morphological features, extracted from words and their morphological analyses. We also extend the established method of hidden event language modeling (HELM) to factored hidden event language modeling (fHELM) to handle morphological information. In order to capture non-lexical information, we extract a set of prosodic features, which are mainly motivated from our previous work for other languages. We then employ discriminative classification techniques, boosting and conditional random fields (CRFs), combined with fHELM, for the task of Turkish sentence segmentation. Ümit Güz, Benoît Favre, Dilek Hakkani-Tür, Gökhan Tür |
IEEE Trans. Speech Audio Process. | 4 |
| 2008 | An iterative unsupervised learning method for information distillationabstractInformation distillation techniques are used to analyze and interpret large volumes of speech and text archives in multiple languages and produce structured information of interest to the user. In this work, we propose an iterative unsupervised sentence extraction method to answer open-ended natural language queries about an event. The approach consists of finding the subset of sentences that are very likely to be relevant or irrelevant for the query from candidate documents, and iteratively training a classification model using these examples. Our results indicate that performance of the system may be improved by around 30% relative in terms of F-measure, by using the proposed method. Kamand Kamangar, Dilek Hakkani-Tür, Gökhan Tür, Michael Levit |
ICASSP | 3 |
| 2008 | Extracting question/answer pairs in multi-party meetingsabstractUnderstanding multi-party meetings involves tasks such as dialog act segmentation and tagging, action item extraction, and summarization. In this paper we introduce a new task for multi-party meetings: extracting question/answer pairs. This is a practical application for further processing such as summarization. We propose a method based on discriminative classification of individual sentences as questions and answers via lexical, speaker, and dialog act tag information, followed by a contextual optimization via Markov models. Our results indicate that it is possible to outperform a non-trivial baseline using dialog act tag information. More specifically, our method achieves a 13% relative improvement over the baseline for the task of detecting answers in meetings. Andreas Kathol, Gökhan Tür |
ICASSP | 2 |
| 2008 | Name-aware speech recognition for interactive question answeringabstractIn this work we show how interactivity in a voice-enabled question answering application may improve speech recognition. We allow the user to provide a target named entity before asking the question. Then we build a named entity specific language model using the documents containing the named entity. The question-specific model is obtained by merging the named entity specific model with the model built on a set of questions. We present a set of experiments using the TREC question set on the AQUAINT corpus. The question-specific language model is compared with the baseline model built by merging a model of the AQUAINT corpus and past TREC questions. The question-specific model achieves 32.2% reduction in word error rate from the baseline using the questions where pronominal references are resolved. Svetlana Stoyanchev, Gökhan Tür, Dilek Hakkani-Tür |
ICASSP | 2 |
| 2008 | Exploiting dialogue act tagging and prosodic information for action item identificationabstractAn important task for multiparty meeting understanding is extracting action items. Action items are a set of tasks that are agreed on by the participants for execution after the meeting, with specific due dates and owners. Dialogue acts, the pragmatic function of an utterance, such as question or backchannel, have been reported to be useful for various dialogue understanding tasks. On the other hand, prosodic information, such as pitch, volume, and speech rate, has been reported to be useful for segmenting a dialogue into utterances or detecting questions. In this paper we investigate the use of dialogue act tagging to improve the identification of action item descriptions and prosodic information to improve action item agreements. Our results indicate that dialogue act tagging improves the identification of action item descriptions by 5% over lexical information, and prosodic information helps discriminating backchannels from agreements with 25% absolute improvement over a baseline. Gökhan Tür, Elizabeth Shriberg |
ICASSP | 2 |
| 2008 | Efficient data selection for machine translationabstractPerformance of statistical machine translation (SMT) systems relies on the availability of a large parallel corpus which is used to estimate translation probabilities. However, the generation of such corpus is a long and expensive process. In this paper, we introduce two methods for efficient selection of training data to be translated by humans. Our methods are motivated by active learning and aim to choose new data that adds maximal information to the currently available data pool. The first method uses a measure of disagreement between multiple SMT systems, whereas the second uses a perplexity criterion. We performed experiments on Chinese-English data in multiple domains and test sets. Our results show that we can select only one-fifth of the additional training data and achieve similar or better translation performance, compared to that of using all available data. Arindam Mandal, Dimitra Vergyri, Wen Wang 0001, Jing Zheng 0001, Andreas Stolcke, Gökhan Tür, Dilek Hakkani-Tür, Necip Fazil Ayan |
SLT | 6 |
| 2008 | The CALO meeting speech recognition and understanding systemabstractThe CALO Meeting Assistant provides for distributed meeting capture, annotation, automatic transcription and semantic analysis of multiparty meetings, and is part of the larger CALO personal assistant system. This paper summarizes the CALO-MA architecture and its speech recognition and understanding components, which include real-time and offline speech transcription, dialog act segmentation and tagging, question-answer pair identification, action item recognition, decision extraction, and summarization. Gökhan Tür, Andreas Stolcke, L. Lynn Voss, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael W. Frandsen, Clint Frederickson, Martin Graciarena, Dilek Hakkani-Tür, Donald Kintzing, Kyle Leveque, Shane Mason, John Niekrasz, Stanley Peters, Matthew Purver, Korbinian Riedhammer, Elizabeth Shriberg, Jing Tien, Dimitra Vergyri |
SLT | 1 |
| 2008 | Bootstrapping spoken dialogue systems by exploiting reusable librariesabstractAbstract Building natural language spoken dialogue systems requires large amounts of human transcribed and labeled speech utterances to reach useful operational service performances. Furthermore, the design of such complex systems consists of several manual steps. The User Experience (UE) expert analyzes and defines by hand the system core functionalities: the system semantic scope (call-types) and the dialogue manager strategy that will drive the human–machine interaction. This approach is extensive and error-prone since it involves several nontrivial design decisions that can be evaluated only after the actual system deployment. Moreover, scalability is compromised by time, costs, and the high level of UE know-how needed to reach a consistent design. We propose a novel approach for bootstrapping spoken dialogue systems based on the reuse of existing transcribed and labeled data, common reusable dialogue templates, generic language and understanding models, and a consistent design process. We demonstrate that our approach reduces design and development time while providing an effective system without any application-specific data. Giuseppe Di Fabbrizio, Gökhan Tür, Dilek Hakkani-Tür, Mazin Gilbert, Bernard Renger, David C. Gibbon, Zhu Liu 0001, Behzad Shahraray |
Nat. Lang. Eng. | 2 |
| 2007 | Integrating several annotation layers for statistical information distillationabstractWe present a sentence extraction algorithm for Information Distillation, a task where for a given templated query, relevant passages must be extracted from massive audio and textual document sources. For each sentence of the relevant documents (that are assumed to be known from the upstream stages) we employ statistical classification methods to estimate the extent of its relevance to the query, whereby two aspects of relevance are taken into account: the template (type) of the query and its slots (free-text descriptions of names, organizations, topic, events and so on, around which templates are centered). The idiosyncrasy of the presented method is in the choice of features used for classification. We extract our features from charts, compilations of elements from various annotation levels, such as word transcriptions, syntactic and semantic parses, and Information Extraction annotations. In our experiments we show that this integrated approach outperforms a purely lexical baseline by as much as 30% relative in terms of F-measure. We also investigate the algorithm's behavior under noisy conditions, by comparing its performance on ASR output and on corresponding manual transcriptions. Michael Levit, Dilek Hakkani-Tür, Gökhan Tür, Daniel Gillick |
ASRU | 3 |
| 2007 | Statistical Sentence Extraction for Information DistillationabstractInformation distillation aims to extract the most useful pieces of information related to a given query from massive, possibly multilingual, audio and textual document sources. One critical component in a distillation engine is detecting sentences to be extracted from each relevant document. In this paper, we present a statistical sentence extraction approach for distillation. Basically, we frame this tack as a classification problem, where each candidate sentence in documents is classified as a relevant to the query or not. These documents may be textual or audio format and in a number of languages. For audio documents, we use both manual and automatic transcriptions, for non-English documents, we use automatic translations. In this work, we use AdaBoost, a discriminative classification method with both lexical and semantic features. The results indicate 11%-13% relative improvement over a baseline keyword-spotting-based approach. We also show the robustness of our method on the audio subset of the document sources using manual and automatic transcriptions. Dilek Hakkani-Tür, Gökhan Tür |
ICASSP (4) | 2 |
| 2007 | Unsupervised Languagemodel Adaptation for Meeting RecognitionabstractWe present an application of unsupervised language model (ML) adaptation to meeting recognition, in a scenario where sequences of multiparty meetings on related topics are to be recognized, but no prior in-domain data for LM training is available. The recognizer LMs are adapted according to the recognition output on temporally preceding meetings, either in speaker-dependent or speaker-independent mode. Model adaptation is carried out by interpolating the n-gram probabilities of a large generic LM with those of a small LM estimated from adaptation data, and minimizing perplexity on the automatic transcripts of a separate meeting set, also previously recognized. The adapted LMs yield about 5.9% relative reduction in word error compared to the baseline. This improvement is about half of what can be achieved with supervised adaptation, i.e. using human-generated speech transcripts. Gökhan Tür, Andreas Stolcke |
ICASSP (4) | 1 |
| 2007 | Co-training using prosodic and lexical information for sentence segmentationabstractWe investigate the application of the co-training learning algorithm on the sentence boundary classification problem by using lexical and prosodic information. Co-training is a semi-supervised machine learning algorithm that uses multiple weak classifiers with a relatively small amount of labeled data and incrementally uses unlabeled data. The assumption in co-training is that the classifiers can co-train each other, as one can label samples that are difficult for the other. The sentence segmentation problem is very appropriate for the co-training method since it satisfies the main requirements of the co-training algorithm: the dataset can be described by two disjoint and natural views that are redundantly sufficient. In our case, the feature sets are capturing lexical and prosodic information. The experimental results on the ICSI Meeting (MRDA) corpus show the effectiveness of the co-training algorithm for this task. Index Terms: co-training, sentence segmentation, prosody, self-training, Boosting Ümit Güz, Sébastien Cuendet, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 4 |
| 2007 | Exploiting information extraction annotations for document retrieval in distillation tasksabstractInformation distillation aims to extract relevant pieces of information related to a given query from massive, possibly multilingual, audio and textual document sources. In this paper, we present our approach for using information extraction annotations to augment document retrieval for distillation. We take advantage of the fact that some of the distillation queries can be associated with annotation elements introduced for the NIST Automatic Content Extraction (ACE) task. We experimentally show that using the ACE events to constrain the document set returned by an information retrieval engine significantly improves the precision at various recall rates for two different query templates. Index Terms: information distillation, information retrieval, information extraction, document retrieval Dilek Hakkani-Tür, Gökhan Tür, Michael Levit |
INTERSPEECH | 2 |
| 2007 | Duration and pronunciation conditioned lexical modeling for speaker verificationabstractWe propose a method to improve speaker recognition lexical model performance using acoustic-prosodic information. More specifically, the lexical model is trained using durationand pronunciation-conditioned word N-grams, simultaneously modeling lexical information along with their acoustic and prosodic characteristics. Support vector machines are used for modeling and scoring, with N-gram frequency vectors serving as features. Experimental results using NIST Speaker Recognition Evaluation data sets show that this method outperforms the regular word N-gram-based lexical models. Furthermore, our approach gives additional information when combined with a high-accuracy acoustic speaker model. We believe that this is a promising step toward integrated speaker recognition models that combine multiple types of high-level features. Gökhan Tür, Elizabeth Shriberg, Andreas Stolcke, Sachin S. Kajarekar |
INTERSPEECH | 1 |
| 2007 | Extending boosting for large scale spoken language understanding
Gökhan Tür |
Mach. Learn. | 1 |
| 2006 | Multitask Learning for Spoken Language UnderstandingabstractIn this paper, we present a multitask learning (MTL) method for intent classification in goal oriented human-machine spoken dialog systems. MTL aims at training tasks in parallel while using a shared representation. What is learned for each task can help other tasks be learned better. Our goal is to automatically re-use the existing labeled data from various applications, which are similar but may have different intents or intent distributions, in order to improve the performance. For this purpose, we propose an automated intent mapping algorithm across applications. We also propose employing active learning to selectively sample the data to be re-used. Our results indicate that we can achieve significant improvements in intent classification performance especially when the labeled data size is limited Gökhan Tür |
ICASSP (1) | 1 |
| 2006 | QASR: question answering using semantic roles for speech interfaceabstractIn this paper, we evaluate a semantic role labeling approach to the extraction of answers in the open domain question answering task. We show that this technique especially improves the system performance when answers are communicated to the user by voice. Semantic role labeling identifies predicates and semantic argument phrases in a sentence. With this information we are able to analyze and extract structure from both questions and candidate sentences, which helps us identify more relevant and precise answers in a long list of candidate sentences. When searching for an answer to a question, we match the missing argument in the question to the semantic parses of the candidate answers. This technique significantly improves the accuracy of the question answering system and results in more concise and grammatical answers, which is essential for enabling voice interfaces to question answering systems. In this paper we apply our approach to factoid questions containing predicates; however, this technique can be also useful in answering more complex questions. Index Terms: question answering, semantic roles Svetlana Stenchikova, Dilek Hakkani-Tür, Gökhan Tür |
INTERSPEECH | 3 |
| 2006 | Model Adaptation for Sentence Segmentation from speechabstractThis paper analyzes various methods to adapt sentence segmentation models trained on conversational telephone speech (CTS) to meeting style conversations. The sentence segmentation model trained using a large amount of CTS data is used to improve the performance when various amounts of meeting data are available. We test the sentence segmentation performance on both reference and speech-to-text (STT) conditions on the ICSI MRDA meeting corpus using the switchboard CTS Corpus as the out-of-domain data. Results show that the sentence segmentation performance is significantly improved by the adapted classification model compared to the one obtained by using in-domain data only, independently of the amount of in-domain data used: 17.5% and 8.4% relative error reductions with only 1,000 and 3,000 in-domain sentences, respectively, and 3.7% relative error reduction with all in-domain data of 80,000 words. Sébastien Cuendet, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 3 |
| 2006 | Model Adaptation for Dialog Act TaggingabstractIn this paper, we analyze the effect of model adaptation for dialog act tagging. The goal of adaptation is to improve the performance of the tagger using out-of-domain data or models. Dialog act tagging aims to provide a basis for further discourse analysis and understanding in conversational speech. In this study we used the ICSI meeting corpus with high-level meeting recognition dialog act (MRDA) tags, that is, question, statement, backchannel, disruptions, and floor grabbers/holders. We performed controlled adaptation experiments using the Switchboard (SWBD) corpus with SWBD-DAMSL tags as the out-of-domain corpus. Our results indicate that we can achieve significantly better dialog act tagging by automatically selecting a subset of the Switchboard corpus and combining the confidences obtained by both in-domain and out-of-domain models via logistic regression, especially when the in-domain data is limited. Gökhan Tür, Ümit Güz, Dilek Hakkani-Tür |
SLT | 1 |
| 2006 | Beyond ASR 1-best: Using word confusion networks in spoken language understanding
Dilek Hakkani-Tür, Frédéric Béchet, Giuseppe Riccardi, Gökhan Tür |
Comput. Speech Lang. | 4 |
| 2006 | Introduction to the Special Issue on Spoken Language Understanding in Conversational Systems
Srinivas Bangalore, Dilek Hakkani-Tür, Gökhan Tür |
Speech Commun. | 3 |
| 2006 | The AT&T spoken language understanding systemabstractSpoken language understanding (SLU) aims at extracting meaning from natural language speech. Over the past decade, a variety of practical goal-oriented spoken dialog systems have been built for limited domains. SLU in these systems ranges from understanding predetermined phrases through fixed grammars, extracting some predefined named entities, extracting users' intents for call classification, to combinations of users' intents and named entities. In this paper, we present the SLU system of VoiceTone/spl reg/ (a service provided by AT&T where AT&T develops, deploys and hosts spoken dialog applications for enterprise customers). The SLU system includes extracting both intents and the named entities from the users' utterances. For intent determination, we use statistical classifiers trained from labeled data, and for named entity extraction we use rule-based fixed grammars. The focus of our work is to exploit data and to use machine learning techniques to create scalable SLU systems which can be quickly deployed for new domains with minimal human intervention. These objectives are achieved by 1) using the predicate-argument representation of semantic content of an utterance; 2) extending statistical classifiers to seamlessly integrate hand crafted classification rules with the rules learned from data; and 3) developing an active learning framework to minimize the human labeling effort for quickly building the classifier models and adapting them to changes. We present an evaluation of this system using two deployed applications of VoiceTone/spl reg/. Narendra K. Gupta, Gökhan Tür, Dilek Hakkani-Tür, Srinivas Bangalore, Giuseppe Riccardi, Mazin Gilbert |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Error Prediction in Spoken Dialog: From Signal-to-Noise Ratio to Semantic Confidence ScoresabstractSpoken dialog systems aim to interpret the meanings of users' utterances and respond to them accordingly. The users' utterances are first recognized by an automatic speech recognizer (ASR) and the intents of the users are extracted by the spoken language understanding (SLU) unit. Both ASR and SLU are noisy and in general their noise statistics are not correlated. Our goal is to exploit the signal-to-noise information and ASR lattice-based and semantic confidence scores for SLU error prediction and prevention of these by rejecting erroneous utterances, or asking confirmation questions. In our experiments, we have shown up to 80% relative decrease in the error rate of the accepted utterances collected using the AT&T How May I Help You/spl trade/ spoken dialog system used for customer care. Dilek Hakkani-Tür, Gökhan Tür, Giuseppe Riccardi, Hong Kook Kim |
ICASSP (1) | 2 |
| 2005 | Model Adaptation For Spoken Language UnderstandingabstractWe present a novel adaptation method for intent classification using boosting in a spoken language understanding system. The goal is adapting an existing model to a new target application, which is similar, but may have different intents or intent distributions. Adaptation can also be employed for a single application where the intent distribution varies with time. We assume the target application has a small amount of labeled data. We also propose employing active learning to sample selectively the data to label for adaptation. Our results indicate that we can achieve the same intent classification accuracy using less than half of the labeled data when there is not much training data available. Furthermore, combined with active learning, we see 18.6% relative reduction in classification error rate. Gökhan Tür |
ICASSP (1) | 1 |
| 2005 | Automated wizard-of-oz for spoken dialogue systemsabstractDesigning and building natural language spoken dialogue systems require large amounts of speech utterances, which adequately represent the intended human-machine dialogues.For this purpose, typically, first a "Wizard-of-Oz" data collection is performed, and then the collected data is transcribed and labeled by expert labelers.Finally, the data is used to train both the speech recognizer and the spoken language understanding stochastic models.Data collection and labeling is an expensive and time consuming manual process.In this paper we propose a completely Automated Wizard, which is capable of recognizing and understanding application independent requests reusing the previously labeled and transcribed data from similar domains, and improving the informativeness of the collected data.We demonstrate that, in the context of automated call routing, compared to the existing data collection systems, the Automated Wizard better captures the user intentions and produces substantially shorter interactions resulting in a better user experience and a less intrusive approach. Giuseppe Di Fabbrizio, Gökhan Tür, Dilek Hakkani-Tür |
INTERSPEECH | 2 |
| 2005 | Combining active and semi-supervised learning for spoken language understanding
Gökhan Tür, Dilek Hakkani-Tür, Robert E. Schapire |
Speech Commun. | 1 |
| 2004 | Unsupervised and active learning in automatic speech recognition for call classificationabstractA key challenge in rapidly building spoken natural language dialog applications is minimizing the manual effort required in transcribing and labeling speech data. This task is not only expensive but also time consuming. We present a novel approach that aims at reducing the amount of manually transcribed in-domain data required for building automatic speech recognition (ASR) models in spoken language dialog systems. Our method is based on mining relevant text from various conversational systems and Web sites. An iterative process is employed where the performance of the models can be improved through both unsupervised and active learning of the ASR models. We have evaluated the robustness of our approach on a call classification task that has been selected from AT&T VoiceTone/sup SM/ customer care. Our results indicate that with unsupervised learning it is possible to achieve a call classification performance that is only 1.5% lower than the upper bound set when using all available in-domain transcribed data. Dilek Hakkani-Tür, Gökhan Tür, Mazin G. Rahim, Giuseppe Riccardi |
ICASSP (1) | 2 |
| 2004 | Extending boosting for call classification using word confusion networksabstractWe are interested in the problem of robust understanding from noisy spontaneous speech input. In goal driven human-machine dialog, utterance classification is a key component of the understanding process to determine the intent of the speaker. We propose a novel algorithm for exploiting ASR word confidence scores for better classification of spoken utterances. Word confidence scores for automatic speech recognition (ASR) provide estimates for word error rates. While previous work has focused on straightforward combination of word confidence scores into Bayesian classifiers, we extend the mathematical formulation for boosting classifiers. This extension of the algorithm allows confidence scores to be exploited from a 1-best ASR output or from word confusion networks (WCNs). We present methods for on-line and off-line score combinations. The results we show are for a large database of utterances collected using the AT&T VoiceTone/sup SM/ spoken dialog system. Our experiments show between 5% and 10% reduction in error (1-precision) for a given recall using WCNs compared to ASR output. Gökhan Tür, Dilek Hakkani-Tür, Giuseppe Riccardi |
ICASSP (1) | 1 |
| 2004 | Cost-sensitive call classificationabstractWe present an efficient and effective method which extends the Boosting family of classifiers to allow the weighted classes. Typically classifiers do not treat individual classes separately. For most real world applications, this is not the case, not all classes have the same importance. The accuracy of a particular class can be more critical than others. In this paper we extend the mathematical formulation for Boosting to weigh the classes differently during training. We have evaluated this method for call classification in AT&T spoken language understanding system. Our results indicate significant improvements in the “important ” classes without a significant loss in the overall performance. 1. Gökhan Tür |
INTERSPEECH | 1 |
| 2003 | Optimizing SVMs for complex call classificationabstractLarge margin classifiers such as support vector machines (SVM) or Adaboost are obvious choices for natural language document or call routing. However, how to combine several binary classifiers to optimize the whole routing process and how this process scales when it involves many different decisions (or classes) is a complex problem that has only received partial answers. We propose a global optimization process based on an optimal channel communication model that allows a combination of possibly heterogeneous binary classifiers. As in Markov modeling, computational feasibility is achieved through simplifications and independence assumptions that are easy to interpret. Using this approach, we have managed to decrease the call-type classification error rate for AT&T's How May I Help You (HMIHY/sup (sm)/) natural dialog system by 50 %. Patrick Haffner, Gökhan Tür, Jeremy H. Wright |
ICASSP (1) | 2 |
| 2003 | Active learning for spoken language understandingabstractWe describe active learning methods for reducing the labeling effort in a statistical call classification system. Active learning aims to minimize the number of labeled utterances by automatically selecting for labeling the utterances that are likely to be most informative. The first method, inspired by certainty-based active learning, selects the examples that the classifier is least confident about. The second method, inspired by committee-based active learning, selects the examples that multiple classifiers do not agree on. We have evaluated these active learning methods using a call classification system used for AT&T customer care. Our results indicate that it is possible to reduce human labeling effort at least by a factor of two. Gökhan Tür, Robert E. Schapire, Dilek Hakkani-Tür |
ICASSP (1) | 1 |
| 2003 | Exploiting unlabeled utterances for spoken language understandingabstractState of the art spoken language understanding systems are trained using labeled utterances, which is labor intensive and time consuming to prepare. In this paper, we propose methods for exploiting the unlabeled data in a statistical call classification system within a natural language dialog system. The basic assumption is that some amount of labeled data and relatively larger chunks of unlabeled data is available. The first method augments the training data by using the machine-labeled call-types for the unlabeled utterances. The second method, instead, augments the classification model trained using the human-labeled utterances with the machine-labeled ones in a weighted manner. We have evaluated these methods using a call classification system used for AT&T natural dialog customer care system. For call classification, we have used a boosting algorithm. Our results indicate that it is possible to obtain the same classification performance by using 30% less labeled data when the unlabeled data is utilized. This corresponds to a 1-1.5% absolute classification error rate reduction, using the same amount of labeled data. Gökhan Tür, Dilek Hakkani-Tür |
INTERSPEECH | 1 |
| 2003 | Active labeling for spoken language understandingabstractState-of-the-art spoken language understanding (SLU) systems are trained using human-labeled utterances, preparation of which is labor intensive and time consuming. Labeling is an error-prone process due to various reasons, such as labeler errors or imperfect description of classes. Thus, usually a second (or maybe more) pass(es) of labeling is required in order to check and fix the labeling errors and inconsistencies of the first (or earlier) pass(es). In this paper, we check the effect of labeling errors for statistical call classification and evaluate methods of finding and correcting these errors by checking minimum amount of data. We describe two alternative methods to speed up the labeling effort, one is based on the confidences obtained from a prior model and the other completely unsupervised. We call the labeling process employing one of these methods as active labeling. Active labeling aims to minimize the number of utterances to be checked again by automatically selecting the ones that are likely to be erroneous or inconsistent with the previously labeled examples. Although very same methods can be used as a postprocessing step to correct labeling errors, we only consider them as part of the labeling process. We have evaluated these active labeling methods using a call classification system used for AT&T natural dialog customer care system. Our results indicate that it is possible to find about 90% of the labeling errors or inconsistencies by checking just half the data. Gökhan Tür, Mazin G. Rahim, Dilek Hakkani-Tür |
INTERSPEECH | 1 |
| 2003 | A statistical information extraction system for TurkishabstractThis paper presents the results of a study on information extraction from unrestricted Turkish text using statistical language processing methods. In languages like English, there is a very small number of possible word forms with a given root word. However, languages like Turkish have very productive agglutinative morphology. Thus, it is an issue to build statistical models for specific tasks using the surface forms of the words, mainly because of the data sparseness problem. In order to alleviate this problem, we used additional syntactic information, i.e. the morphological structure of the words. We have successfully applied statistical methods using both the lexical and morphological information to sentence segmentation, topic segmentation, and name tagging tasks. For sentence segmentation, we have modeled the final inflectional groups of the words and combined it with the lexical model, and decreased the error rate to 4.34%, which is 21% better than the result obtained using only the surface forms of the words. For topic segmentation, stems of the words (especially nouns) have been found to be more effective than using the surface forms of the words and we have achieved 10.90% segmentation error rate on our test set according to the weighted TDT-2 segmentation cost metric. This is 32% better than the word-based baseline model. For name tagging, we used four different information sources to model names. Our first information source is based on the surface forms of the words. Then we combined the contextual cues with the lexical model, and obtained some improvement. After this, we modeled the morphological analyses of the words, and finally we modeled the tag sequence, and reached an F-Measure of 91.56%, according to the MUC evaluation criteria. Our results are important in the sense that, using linguistic information, i.e. morphological analyses of the words, and a corpus large enough to train a statistical model significantly improves these basic information extraction tasks for Turkish. Gökhan Tür, Dilek Hakkani-Tür, Kemal Oflazer |
Nat. Lang. Eng. | 1 |
| 2002 | Improving spoken language understanding using word confusion networksabstractA natural language spoken dialog system includes a large vocabulary automatic speech recognition (ASR) engine, whose output is used as the input of a spoken language understanding component. Two challenges in such a framework are that the ASR component is far from being perfect and the users can say the same thing in very different ways. So, it is very important to be tolerant to recognition errors and some amount of orthographic variability. In this paper, we present our work on developing new methods and investigating various ways of robust recognition and understanding of an utterance. To this end, we exploit word-level confusion networks (sausages), obtained from ASR word graphs (lattices) instead of the ASR 1-best hypothesis. Using sausages with an improved confidence model, we decreased the calltype classification error rate for AT&T's How May I Help You (HMIHY ) natural dialog system by 38%. Gökhan Tür, Jeremy H. Wright, Allen L. Gorin, Giuseppe Riccardi, Dilek Hakkani-Tür |
INTERSPEECH | 1 |
| 2001 | Integrating Prosodic and Lexical Cues for Automatic Topic SegmentationabstractWe present a probabilistic model that uses both prosodic and lexical cues for the automatic segmentation of speech into topically coherent units. We propose two methods for combining lexical and prosodic information using hidden Markov models and decision trees. Lexical information is obtained from a speech recognizer, and prosodic features are extracted automatically from speech waveforms. We evaluate our approach on the Broadcast News corpus, using the DARPA-TDT evaluation metrics. Results show that the prosodic model alone is competitive with word-based segmentation methods. Furthermore, we achieve a significant reduction in error by combining the prosodic and word-based knowledge sources. Gökhan Tür, Dilek Hakkani-Tür, Andreas Stolcke, Elizabeth Shriberg |
Comput. Linguistics | 1 |
| 2000 | Statistical Morphological Disambiguation for Agglutinative Languages
Dilek Hakkani-Tür, Kemal Oflazer, Gökhan Tür |
COLING | 3 |
| 2000 | Prosody-based automatic segmentation of speech into sentences and topics
Elizabeth Shriberg, Andreas Stolcke, Dilek Hakkani-Tür, Gökhan Tür |
Speech Commun. | 4 |
| 1999 | Combining words and prosody for information extraction from speechabstractThe design principles and collection procedures behind a speech synthesis corpus directly impact the performance of the resulting text-to-speech system. This paper describes the design and collection of the Victoria corpus, created to support speech synthesis research and development at Apple Computer. This corpus is composed of ve constituent parts, each designed to cover a speci c aspect of speech synthesis: polyphones, prosodic contexts, reiterant speech, function word sequences, and continuous speech. It was spoken in general U.S. English by one linguisticallytrained adult female. Portions of the corpus are being used in the statistical estimation of duration and pitch models for Apple's next-generation textto-speech system, MacinTalk 4. Dilek Hakkani-Tür, Gökhan Tür, Andreas Stolcke, Elizabeth Shriberg |
EUROSPEECH | 2 |
| 1999 | Modeling the prosody of hidden events for improved word recognitionabstractWe investigate a new approach for using speech prosody as a knowledge source for speech recognition. The idea is to penalize word hypotheses that are inconsistent with prosodic features such as duration and pitch. To model the interaction between words and prosody we modify the language model to represent hidden events such as sentence boundaries and various forms of disfluency, and combine with it decision trees that predict such events from prosodic features. N-best rescoring experiments on the Switchboard corpus show a small but consistent reduction of word error as a result of this modeling. We conclude with a preliminary analysis of the types of errors that are corrected by the prosodically informed model. 1. Andreas Stolcke, Elizabeth Shriberg, Dilek Hakkani-Tür, Gökhan Tür |
EUROSPEECH | 4 |
| 1998 | Automatic detection of sentence boundaries and disfluencies based on recognized wordsabstractWe study the problem of detecting linguistic events at interword boundaries, such as sentence boundaries and disfluency locations, in speech transcribed by an automatic recognizer. Recovering such events is crucial to facilitate speech understanding and other natural language processing tasks. Our approach is based on a combination of prosodic cues modeled by decision trees, and word-based event N-gram language models. Several model combination approaches are investigated. The techniques are evaluated on conversational speech from the Switchboard corpus. Model combination is shown to give a significant win over individual knowledge sources. 1. INTRODUCTION Current automatic speech recognition systems output a string of words. Most natural language understanding systems, however, require structural information such as punctuation, which is present in text but not overtly indicated in spoken language. Similarly, for speech understanding and information extraction, it is important to fi... Andreas Stolcke, Elizabeth Shriberg, Rebecca Bates 0001, Mari Ostendorf, Dilek Hakkani-Tür, Madelaine Plauché, Gökhan Tür |
ICSLP | 7 |
| 1997 | Morphological Disambiguation by Voting ConstraintsabstractWe present a constraint-based morphological disambiguation system in which individual constraints vote on matching morphological parses, and disambiguation of all the tokens in a sentence is performed at the end by selecting parses that receive the highest votes. This constraint application paradigm makes the outcome of the disambiguation independent of the rule sequence, and hence relieves the rule developer from worrying about potentially conflicting rule sequencing. Our results for disambiguating Turkish indicate that using about 500 constraint rules and some additional simple statistics, we can attain a recall of 95--96% and a precision of 94--95% with about 1.01 parses per token. Our system is implemented in Prolog and we are currently investigating an efficient implementation based on finite state transducers. Kemal Oflazer, Gökhan Tür |
ACL | 2 |
| 1996 | Combining Hand-crafted Rules and Unsupervised Learning in Constraint-based Morphological Disambiguation
Kemal Oflazer, Gökhan Tür |
EMNLP | 2 |