EDBT 2026 Demo / reviewers in the wild / expert
Yun-Nung Chen
dblp:04/9878 · also Yun-Nung Vivian Chen
· DBLP profile ↗
95ranked-venue papers
26as first author
26since 2021 · last 2026
0000-0003-1777-3942ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 66 · 14 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 18 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Role-Playing Evaluation: Anonymous Benchmarking and A Systematic Study of Personality EffectsabstractLarge Language Models (LLMs) have shown remarkable potential in developing role-playing agents (RPAs). However, current evaluation frameworks rely heavily on well-known fictional characters, raising a critical concern: models may be leveraging their internal training memory of these characters rather than demonstrating role-playing capabilities. This reliance often leads to significant performance degradation when RPAs encounter unseen or out-of-distribution personas. To address this, we propose a more rigorous evaluation protocol designed to decouple role-playing proficiency from character recognition. Our experiments across multiple benchmarks demonstrate that anonymizing characters degrades performance, confirming that name exposure provides implicit cues that mask a model’s true capability. To mitigate this, we investigate diverse personality augmentation as a method to enhance role fidelity in anonymous settings. We systematically analyze the impact of various personality-description methods on agent behavior and consistency. Our results show that incorporating personality information consistently improves RPA performance. This work establishes a more equitable evaluation standard and validates a scalable, personality-enhanced framework for constructing robust RPAs. Ji-Lun Peng, Yun-Nung Chen |
SIGDIAL | 2 |
| 2026 | BEATS: Bootstrapping E-commerce Attribute Taxonomies for Search through Iterative Human-AI CollaborationabstractE-commerce platforms in emerging markets often operate with underdeveloped product catalogs that contain only category taxonomies but lack structured attribute schemas. This absence of fine-grained product attributes limits search capabilities---preventing faceted filtering, degrading query understanding, and weakening semantic representations used by search systems. We present BEATS, a human-in-the-loop LLM framework for bootstrapping product attribute taxonomies entirely from scratch. Our approach extends a multi-stage LLM generation pipeline with two critical production stages: (1) proactive quality checking by model developers to filter erroneous outputs, and (2) human annotation by domain-expert local staff to validate generated attributes. The framework operates iteratively---prompts at each generation stage are refined based on quality check observations and annotator feedback across successive rounds, progressively improving attribute quality. Once the attribute taxonomy is established, we employ LLMs to perform structured attribute tagging on individual product items, enriching their contextual representations. The enriched catalog directly benefits multiple components of the search system: enabling granular attribute-based filtering, providing structured features for ranking models, and improving semantic representations for dense retrieval. We validate the generated taxonomy by training dense retrieval models on attribute-enriched product data, demonstrating consistent improvements over baselines using original catalog information. Our system has been deployed at Rakuten Taiwan, enriching 9 major categories spanning 2,694 sub-categories with 67,277 generated attributes, and over 5.4 million products have been tagged with the generated attributes, with plans to enrich the entire product catalog. Yung-Yu Shih, Shang-Yu Su, Tzu-I Ho, Dongzhe Wang, Yun-Nung Chen |
SIGIR | 5 |
| 2025 | From Simulation to Strategy: Automating Personalized Interaction Planning for Conversational AgentsabstractAmid the rapid rise of agentic dialogue models, realistic user-simulator studies are essential for tuning effective conversation strategies. This work investigates a sales-oriented agent that adapts its dialogue based on user profiles spanning age, gender, and occupation. While age and gender influence overall performance, occupation produces the most pronounced differences in conversational intent. Leveraging this insight, we introduce a lightweight, occupation-conditioned strategy that guides the agent to prioritize intents aligned with user preferences, resulting in shorter and more successful dialogues. Our findings highlight the importance of rich simulator profiles and demonstrate how simple persona-informed strategies can enhance the effectiveness of sales-oriented dialogue systems.11Code & Prompts: https://github.com/MiuLab/PersonalizedAgent/ Wen-Yu Chang, Tzu-Hung Huang, Chih-Ho Chen, Yun-Nung Chen |
ASRU | 4 |
| 2025 | Company-Specific Knowledge Matters: Retrieval-Augmented Generation for Earnings Call Answer RehearsalabstractRetrieval-augmented generation (RAG) has long been used to guide generative models in producing more accurate answers, with most discussions focusing on reading comprehension-based question answering (QA). However, their role in real-world answer rehearsal scenarios remains underexplored. The rise of large language models (LLMs) presents new opportunities to develop systems that assist professionals, making this discussion both timely and essential. This paper explores how to better support corporate executives in answering questions from professional analysts during earnings calls. We compare the impact of two external knowledge sources: large-scale causal knowledge graphs (KGs) and historical Q&A records-retrieved either from a global pool or company-specific archives. Our findings suggest that a company's historical Q&A records are more influential than causal KGs in improving response quality. To the best of our knowledge, this is the first study to systematically compare and analyze different knowledge resources in answer rehearsal. Our findings show the potential of inspiring future research on the interplay between KG and historical QA pairs for answer rehearsal. Yung-Yu Shih, Yun-Nung Chen, Chung-Chi Chen 0001 |
CIKM | 2 |
| 2025 | Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future DirectionsabstractRecent advancements in large language models (LLMs) and AI systems have led to a paradigm shift in the design and optimization of complex AI workflows.By integrating multiple components, compound AI systems have become increasingly adept at performing sophisticated tasks.However, as these systems grow in complexity, new challenges arise in optimizing not only individual components but also their interactions.While traditional optimization methods such as supervised fine-tuning (SFT) and reinforcement learning (RL) remain foundational, the rise of natural language feedback introduces promising new approaches, especially for optimizing non-differentiable systems.This paper provides a systematic review of recent progress in optimizing compound AI systems, encompassing both numerical and languagebased techniques.We formalize the notion of compound AI system optimization, classify existing methods along several key dimensions, and highlight open research challenges and future directions in this rapidly evolving field. 1 Yu-Ang Lee, Guan-Ting Yi, Mei-Yi Liu, Jui-Chao Lu, Guan-Bo Yang, Yun-Nung Chen |
EMNLP | 6 |
| 2025 | Creativity in LLM-based Multi-Agent Systems: A SurveyabstractYi-Cheng Lin, Kang-Chieh Chen, Zhe-Yan Li, Tzu-Heng Wu, Tzu-Hsuan Wu, Kuan-Yu Chen, Hung-yi Lee, Yun-Nung Chen. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yi-Cheng Lin, Kang-Chieh Chen, Zhe-Yan Li, Tzu-Heng Wu, Tzu-Hsuan Wu, Kuan-Yu Chen 0005, Hung-yi Lee, Yun-Nung Chen |
EMNLP | 8 |
| 2025 | Using Language Models to Generate and Forget the Narrative Memories of an Assistive Robot
Angel F. Garcia Contreras, Wen-Yu Chang, Seiya Kawano, Yun-Nung Chen, Koichiro Yoshino |
MMM (5) | 4 |
| 2025 | Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token LearningabstractMaintaining consistent model performance across domains is a fundamental challenge in machine learning. While recent work has explored using LLM-generated data for fine-tuning, its impact on cross-domain generalization remains poorly understood. This paper presents a systematic analysis revealing that fine-tuning with LLM-generated data not only improves target task performance but also reduces non-target task degradation compared to fine-tuning with ground truth data. Through analyzing the data sequence in tasks of various domains, we demonstrate that this enhancement of non-target task robustness stems from the reduction of high perplexity tokens found in LLM-generated sequences. Following our findings, we showed that masking high perplexity tokens in ground truth training data achieves similar non-target task performance preservation, comparable to using LLM-generated data. Extensive experiments across different model families and scales, including Gemma 2 IT 2B, Llama 3 8B Instruct, and three additional models, agree with our findings. To the best of our knowledge, this is the first work to provide an empirical explanation based on token perplexity reduction to mitigate catastrophic forgetting in LLMs after fine-tuning, offering valuable insights for developing more robust fine-tuning strategies. Chao-Chung Wu, Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Shao-Hua Sun, Hung-yi Lee |
NeurIPS | 4 |
| 2024 | Efficient Unseen Language Adaptation for Multilingual Pre-Trained Language ModelsabstractMultilingual pre-trained language models (mPLMs) have demonstrated notable effectiveness in zero-shot cross-lingual transfer tasks.Specifically, they can be fine-tuned solely on tasks in the source language and subsequently applied to tasks in the target language.However, for low-resource languages unseen during pre-training, relying solely on zero-shot language transfer often yields sub-optimal results.One common strategy is to continue training PLMs using masked language modeling objectives on the target language.Nonetheless, this approach can be inefficient due to the need to adjust all parameters for language adaptation.In this paper, we propose a more efficient solution: soft-prompt tuning for language adaptation.Our experiments demonstrate that with carefully designed prompts, soft-prompt tuning enables mPLMs to achieve effective zero-shot cross-lingual transfer to downstream tasks in previously unseen languages.Notably, we found that prompt tuning outperforms continuously trained baselines on two text classification benchmarks, encompassing 20 low-resource languages while utilizing a mere 0.28% of the tuned parameters.These results underscore the superior adaptability of mPLMs to previously unseen languages afforded by soft-prompt tuning compared to traditional fine-tuning methods. 1 Po-Heng Chen, Yun-Nung Chen |
EMNLP | 2 |
| 2024 | PairDistill: Pairwise Relevance Distillation for Dense RetrievalabstractEffective information retrieval (IR) from vast datasets relies on advanced techniques to extract relevant information in response to queries.Recent advancements in dense retrieval have showcased remarkable efficacy compared to traditional sparse retrieval methods.To further enhance retrieval performance, knowledge distillation techniques, often leveraging robust crossencoder rerankers, have been extensively explored.However, existing approaches primarily distill knowledge from pointwise rerankers, which assign absolute relevance scores to documents, thus facing challenges related to inconsistent comparisons.This paper introduces Pairwise Relevance Distillation (PAIRDISTILL) to leverage pairwise reranking, offering finegrained distinctions between similarly relevant documents to enrich the training of dense retrieval models.Our experiments demonstrate that PAIRDISTILL outperforms existing methods, achieving new state-of-the-art results across multiple benchmarks.This highlights the potential of PAIRDISTILL in advancing dense retrieval techniques effectively. 1 Chao-Wei Huang, Yun-Nung Chen |
EMNLP | 2 |
| 2024 | DogeRM: Equipping Reward Models with Domain Knowledge through Model MergingabstractReinforcement learning from human feedback (RLHF) is a popular strategy for aligning large language models (LLMs) with desired behaviors.Reward modeling is a crucial step in RLHF.However, collecting paired preference data for training reward models is often costly and time-consuming, especially for domainspecific preferences requiring expert annotation.To address this challenge, we propose the Domain knowledge merged Reward Model (DogeRM), a novel framework that integrates domain-specific knowledge into a general reward model by model merging.The experiments demonstrate that DogeRM enhances performance across different benchmarks and provide a detailed analysis showcasing the effects of model merging, showing the great potential of facilitating model alignment.1 Tzu-Han Lin, Chen-An Li, Hung-yi Lee, Yun-Nung Chen |
EMNLP | 4 |
| 2024 | I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL GenerationabstractThis study explores the proactive ability of LLMs to seek user support.We propose metrics to evaluate the trade-off between performance improvements and user burden, and investigate whether LLMs can determine when to request help under varying information availability.Our experiments show that without external feedback, many LLMs struggle to recognize their need for user support.The findings highlight the importance of external signals and provide insights for future research on improving support-seeking strategies. Cheng-Kuang Wu, Zhi Rui Tam, Chao-Chung Wu, Chieh-Yen Lin, Hung-yi Lee, Yun-Nung Chen |
EMNLP | 6 |
| 2024 | StreamBench: Towards Benchmarking Continuous Improvement of Language AgentsabstractRecent works have shown that large language model (LLM) agents are able to improve themselves from experience, which is an important ability for continuous enhancement post-deployment. However, existing benchmarks primarily evaluate their innate capabilities and do not assess their ability to improve over time. To address this gap, we introduce StreamBench, a pioneering benchmark designed to evaluate the continuous improvement of LLM agents over an input-feedback sequence. StreamBench simulates an online learning environment where LLMs receive a continuous flow of feedback stream and iteratively enhance their performance. In addition, we propose several simple yet effective baselines for improving LLMs on StreamBench, and provide a comprehensive analysis to identify critical components that contribute to successful streaming strategies. Our work serves as a stepping stone towards developing effective online learning strategies for LLMs, paving the way for more adaptive AI systems in streaming scenarios. Cheng-Kuang Wu, Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Hung-yi Lee |
NeurIPS | 4 |
| 2024 | Large Language Model Based Generative Error Correction: A Challenge and Baselines For Speech Recognition, Speaker Tagging, and Emotion RecognitionabstractGiven recent advances in generative AI technology, a key question is how large language models (LLMs) can enhance acoustic modeling tasks using text decoding results from a frozen, pretrained automatic speech recognition (ASR) model. To explore new capabilities in language modeling for speech processing, we introduce the generative speech transcription error correction (GenSEC) challenge. This challenge comprises three post-ASR language modeling tasks: (i) post-ASR transcription correction, (ii) speaker tagging, and (iii) emotion recognition. These tasks aim to emulate future LLM-based agents handling voice-based interfaces while remaining accessible to a broad audience by utilizing open pretrained language models or agent-based APIs. We also discuss insights from baseline evaluations, as well as lessons learned for designing future evaluations. Chao-Han Huck Yang, Taejin Park, Yuan Gong 0001, Yuanchao Li, Zhehuai Chen, Chen Chen 0075, Kunal Dhawan, Piotr Zelasko, Chao Zhang 0031, Yun-Nung Chen, Yu Tsao 0001, Jagadeesh Balam, Boris Ginsburg, Sabato Marco Siniscalchi, Chng Eng Siong, Peter Bell 0001, Catherine Lai, Shinji Watanabe 0001, Andreas Stolcke |
SLT | 12 |
| 2024 | Overview of the Ninth Dialog System Technology Challenge: DSTC9abstractThis paper introduces the Ninth Dialog System Technology Challenge (DSTC-9). This edition of the DSTC focuses on applying end-to-end dialog technologies for four distinct tasks in dialog systems, namely, 1. Task-oriented dialog Modeling with Unstructured Knowledge Access, 2. Multi-domain task-oriented dialog, 3. Interactive evaluation of dialog and 4. Situated interactive multimodal dialog. This paper describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. R. Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro, Abhinav Rastogi, Yun-Nung Chen, Mihail Eric, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Yang Liu 0004, Chao-Wei Huang, Dilek Hakkani-Tür, Jinchao Li, Qi Zhu 0007, Lingxiao Luo, Lars Liden, Kaili Huang, Shahin Shayandeh, Runze Liang, Baolin Peng, Zheng Zhang 0020, Swadheen Shukla, Minlie Huang, Jianfeng Gao 0001, Shikib Mehri, Yulan Feng, Carla Gordon, Seyed Hossein Alavi, David R. Traum, Maxine Eskénazi, Ahmad Beirami, Eunjoon Cho, Paul A. Crook, Ankita De, Alborz Geramifard, Satwik Kottur, Seungwhan Moon, Shivani Poddar, Rajen Subba |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Joint Dual Learning With Mutual Information Maximization for Natural Language Understanding and Generation in DialoguesabstractModular conversational systems heavily rely on the performance of their natural language understanding (NLU) and natural language generation (NLG) components. NLU focuses on extracting core semantic concepts from input texts, while NLG constructs coherent sentences based on these extracted semantics. Inspired by information theory in digital communication, we introduce a one-way communication model that mirrors human conversations, comprising two distinct phases: (1) the conversion of thoughts into messages, similar to NLG, and (2) the comprehension of received messages, similar to NLU. This paper presents a novel algorithm that trains NLU and NLG collaboratively by concatenating their models and maximizing mutual information between inputs and outputs. This approach efficiently facilitates the transmission of semantics, leading to enhanced learning performance for both components. Our experimental results, based on three benchmark datasets, consistently demonstrate significant improvements for both NLU and NLG tasks, highlighting the practical promise of our proposed method. Shang-Yu Su, Yung-Sung Chung, Yun-Nung Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Self-ICL: Zero-Shot In-Context Learning with Self-Generated DemonstrationsabstractLarge language models (LLMs) have exhibited striking in-context learning (ICL) ability to adapt to target tasks with a few inputoutput demonstrations.For better ICL, different methods are proposed to select representative demonstrations from existing training corpora.However, such settings are not aligned with real-world practices, as end-users usually query LMs without access to demonstration pools.In this work, we introduce SELF-ICL-a simple framework which bootstraps LMs' intrinsic capabilities to perform zero-shot ICL.Given a test input, SELF-ICL first prompts the model to generate pseudoinputs.Next, the model predicts pseudo-labels for the pseudo-inputs via zero-shot prompting.Finally, we perform ICL for the test input with the pseudo-input-label pairs as demonstrations.Evaluation on 23 BIG-Bench Hard tasks shows SELF-ICL outperforms zero-shot baselines on both average accuracy and head-to-head comparison.Moreover, with zero-shot chain-ofthought, SELF-ICL achieves results comparable to using real demonstrations.Additionally, we conduct a range of analyses to validate SELF-ICL's effectiveness and provide insights for its behaviors under different settings.1 * Equal contribution. 1 https://github.com/ntunlplab/Self-ICLFollowing is an example instance for the task: Evaluate the result of a random Boolean expression.Please come up with 3 new, diverse, and creative instances for the task.Example instance: Q: not ( True ) and ( True ) is New instance 1: Q: ( False ) or ( False ) and ( True ) is New instance 2: Q: ( True ) and ( False ) or ( False ) is New instance 3: Q: not ( False ) and ( True ) or ( False ) Task description: Evaluate the result of a random Boolean expression.Q: ( False ) or ( False ) and ( True ) is A: False Q: ( True ) and ( False ) or ( False ) is A: False Q: not ( False ) and ( True ) or ( False ) A: True Q: not ( True ) and ( True ) is Weilin Chen 0005, Cheng-Kuang Wu, Yun-Nung Chen, Hsin-Hsi Chen |
EMNLP | 3 |
| 2023 | CONVERSER: Few-shot Conversational Dense Retrieval with Synthetic Data GenerationabstractConversational search provides a natural interface for information retrieval (IR).Recent approaches have demonstrated promising results in applying dense retrieval to conversational IR.However, training dense retrievers requires large amounts of in-domain paired data.This hinders the development of conversational dense retrievers, as abundant in-domain conversations are expensive to collect.In this paper, we propose CONVERSER, a framework for training conversational dense retrievers with at most 6 examples of in-domain dialogues.Specifically, we utilize the in-context learning capability of large language models to generate conversational queries given a passage in the retrieval corpus.Experimental results on conversational retrieval benchmarks OR-QuAC and TREC CAsT 19 show that the proposed CON-VERSER achieves comparable performance to fully-supervised models, demonstrating the effectiveness of our proposed framework in fewshot conversational dense retrieval.1 Chao-Wei Huang, Chen-Yu Hsu 0003, Tsu-Yuan Hsu, Chen-An Li, Yun-Nung Chen |
SIGDIAL | 5 |
| 2022 | SalesBot: Transitioning from Chit-Chat to Task-Oriented DialoguesabstractDialogue systems are usually categorized into two types, open-domain and task-oriented.The first one focuses on chatting with users and making them engage in the conversations, where selecting a proper topic to fit the dialogue context is essential for a successful dialogue.The other one focuses on a specific task instead of casual talks, e.g., finding a movie on Friday night, playing a song.These two directions have been studied separately due to their different purposes.However, how to smoothly transition from social chatting to task-oriented dialogues is important for triggering the business opportunities, and there is no any public data focusing on such scenarios.Hence, this paper focuses on investigating the conversations starting from open-domain social chatting and then gradually transitioning to taskoriented purposes, and releases a large-scale dataset with detailed annotations for encouraging this research direction.To achieve this goal, this paper proposes a framework to automatically generate many dialogues without human involvement, in which any powerful opendomain dialogue generation model can be easily leveraged.The human evaluation shows that our generated dialogue data has a natural flow at a reasonable quality, showing that our released data has a great potential of guiding future research directions and commercial activities.Furthermore, the released models allow researchers to automatically generate unlimited dialogues in the target scenarios, which can greatly benefit semi-supervised and unsupervised approaches. 1 Ssu Chiu, Yun-Nung Chen |
ACL (1) | 4 |
| 2022 | Contrastive Learning for Improving ASR Robustness in Spoken Language UnderstandingabstractSpoken language understanding (SLU) is an essential task for machines to understand human speech for better interactions. However, errors from the automatic speech recognizer (ASR) usually hurt the understanding performance. In reality, ASR systems may not be easy to adjust for the target scenarios. Therefore, this paper focuses on learning utterance representations that are robust to ASR errors using a contrastive objective, and further strengthens the generalization ability by combining supervised contrastive learning and self-distillation in model fine-tuning. Experiments on three benchmark datasets demonstrate the effectiveness of our proposed approach. Ya-Hsin Chang, Yun-Nung Chen |
INTERSPEECH | 2 |
| 2022 | Controllable User Dialogue Act Augmentation for Dialogue State TrackingabstractPrior work has demonstrated that data augmentation is useful for improving dialogue state tracking.However, there are many types of user utterances, while the prior method only considered the simplest one for augmentation, raising the concern about poor generalization capability.In order to better cover diverse dialogue acts and control the generation quality, this paper proposes controllable user dialogue act augmentation (CUDA-DST) to augment user utterances with diverse behaviors.With the augmented data, different state trackers gain improvement and show better robustness, achieving the state-of-the-art performance on MultiWOZ 2.1.1 Chun-Mao Lai, Ming-Hao Hsu, Chao-Wei Huang, Yun-Nung Chen |
SIGDIAL | 4 |
| 2022 | TREND: Trigger-Enhanced Relation-Extraction Network for DialoguesabstractThe goal of dialogue relation extraction (DRE) is to identify the relation between two entities in a given dialogue.During conversations, speakers may expose their relations to certain entities by explicit or implicit clues, such evidences called "triggers".However, trigger annotations may not be always available for the target data, so it is challenging to leverage such information for enhancing the performance.Therefore, this paper proposes to learn how to identify triggers from the data with trigger annotations and then transfers the trigger-finding capability to other datasets for better performance.The experiments show that the proposed approach is capable of improving relation extraction performance of unseen relations and also demonstrate the transferability of our proposed trigger-finding model across different domains and datasets.1 Po-Wei Lin, Shang-Yu Su, Yun-Nung Chen |
SIGDIAL | 3 |
| 2021 | Real-time Tropical Cyclone Intensity Estimation by Handling Temporally Heterogeneous Satellite DataabstractAnalyzing big geophysical observational data collected by multiple advanced sensors on various satellite platforms promotes our understanding of the geophysical system. For instance, convolutional neural networks (CNN) have achieved great success in estimating tropical cyclone (TC) intensity based on satellite data with fixed temporal frequency (e.g., ~3 h). However, to achieve more timely (under 30 min) and accurate TC intensity estimates, a deep learning model is demanded to handle temporally-heterogeneous satellite observations. Specifically, infrared (IR1) and water vapor (WV) images are available within every 15 minute period, while passive microwave rain rate (PMW) is available about every 3 hours. Meanwhile, the visible (VIS) channel is severely affected by noise and sunlight intensity, making it difficult to be utilized. Therefore, we propose a novel framework that combines generative adversarial network (GAN) with CNN. The model utilizes all data during the training phase including VIS and PMW information and eventually uses only the high-frequent IR1 and WV data for providing intensity estimates during the predicting phase. Experimental results demonstrate that the hybrid GAN-CNN framework achieves comparable precision to the state-of-the-art models, while possessing the capability of increasing the maximum estimation frequency from 3 hours to less than 15 minutes. Please visit https://github.com/BoyoChen/CNN-GAN-TC for codes and implementation details. Boyo Chen, Buo-Fu Chen, Yun-Nung Chen |
AAAI | 3 |
| 2021 | Efficient Multi-Task Auxiliary Learning: Selecting Auxiliary Data by Feature SimilarityabstractMulti-task auxiliary learning utilizes a set of relevant auxiliary tasks to improve the performance of a primary task.A common usage is to manually select multiple auxiliary tasks for multi-task learning on all data, which raises two issues: (1) selecting beneficial auxiliary tasks for a primary task is nontrivial; (2) when the auxiliary datasets are large, training on all data becomes time-expensive and impractical.Therefore, this paper focuses on addressing these problems and proposes a timeefficient sampling method to select the data that is most relevant to the primary task.The proposed method allows us to only train on the most beneficial sub-datasets from the auxiliary tasks, achieving efficient multi-task auxiliary learning.The experiments on three benchmark datasets (RTE, MRPC, STS-B) show that our method significantly outperforms random sampling and ST-DNN.Also, by applying our method, the model can surpass fully-trained MT-DNN on RTE, MRPC, STS-B, using only 50%, 66%, and 1% of data, respectively.1 Po-Nien Kung, Sheng-Siang Yin, Tse-Hsuan Yang, Yun-Nung Chen |
EMNLP (1) | 5 |
| 2021 | Modeling Diagnostic Label Correlation for Automatic ICD CodingabstractGiven the clinical notes written in electronic health records (EHRs), it is challenging to predict the diagnostic codes which is formulated as a multi-label classification task.The large set of labels, the hierarchical dependency, and the imbalanced data make this prediction task extremely hard.Most existing work built a binary prediction for each label independently, ignoring the dependencies between labels.To address this problem, we propose a two-stage framework to improve automatic ICD coding by capturing the label correlation.Specifically, we train a label set distribution estimator to rescore the probability of each label set candidate generated by a base predictor.This paper is the first attempt at learning the label set distribution as a reranking module for medical code prediction.In the experiments, our proposed framework is able to improve upon best-performing predictors on the benchmark MIMIC datasets.1 Shang-Chi Tsai, Chao-Wei Huang, Yun-Nung Chen |
NAACL-HLT | 3 |
| 2020 | An Empirical Study of Content Understanding in Conversational Question AnsweringabstractWith a lot of work about context-free question answering systems, there is an emerging trend of conversational question answering models in the natural language processing field. Thanks to the recently collected datasets, including QuAC and CoQA, there has been more work on conversational question answering, and recent work has achieved competitive performance on both datasets. However, to best of our knowledge, two important questions for conversational comprehension research have not been well studied: 1) How well can the benchmark dataset reflect models' content understanding? 2) Do the models well utilize the conversation content when answering questions? To investigate these questions, we design different training settings, testing settings, as well as an attack to verify the models' capability of content understanding on QuAC and CoQA. The experimental results indicate some potential hazards in the benchmark datasets, QuAC and CoQA, for conversational comprehension research. Our analysis also sheds light on both what models may learn and how datasets may bias the models. With deep investigation of the task, it is believed that this work can benefit the future progress of conversation comprehension. The source code is available at https://github.com/MiuLab/CQA-Study. Ting-Rui Chiang, Hao-Tong Ye, Yun-Nung Chen |
AAAI | 3 |
| 2020 | Learning Spoken Language Representations with Neural Lattice Language ModelingabstractPre-trained language models have achieved huge improvement on many NLP tasks.However, these methods are usually designed for written text, so they do not consider the properties of spoken language.Therefore, this paper aims at generalizing the idea of language model pre-training to lattices generated by recognition systems.We propose a framework that trains neural lattice language models to provide contextualized representations for spoken language understanding tasks.The proposed two-stage pre-training approach reduces the demands of speech data and has better efficiency.Experiments on intent detection and dialogue act recognition datasets demonstrate that our proposed method consistently outperforms strong baselines when evaluated on spoken inputs. 1 Chao-Wei Huang, Yun-Nung Chen |
ACL | 2 |
| 2020 | Towards Unsupervised Language Understanding and Generation by Joint Dual LearningabstractIn modular dialogue systems, natural language understanding (NLU) and natural language generation (NLG) are two critical components, where NLU extracts the semantics from the given texts and NLG is to construct corresponding natural language sentences based on the input semantic representations.However, the dual property between understanding and generation has been rarely explored.The prior work (Su et al., 2019) is the first attempt that utilized the duality between NLU and NLG to improve the performance via a dual supervised learning framework.However, the prior work still learned both components in a supervised manner; instead, this paper introduces a general learning framework to effectively exploit such duality, providing flexibility of incorporating both supervised and unsupervised learning algorithms to train language understanding and generation models in a joint fashion.The benchmark experiments demonstrate that the proposed approach is capable of boosting the performance of both NLU and NLG. 1 Shang-Yu Su, Chao-Wei Huang, Yun-Nung Chen |
ACL | 3 |
| 2020 | Lifelong Language Knowledge DistillationabstractIt is challenging to perform lifelong language learning (LLL) on a stream of different tasks without any performance degradation comparing to the multi-task counterparts.To address this issue, we present Lifelong Language Knowledge Distillation (L2KD), a simple but efficient method that can be easily applied to existing LLL architectures in order to mitigate the degradation.Specifically, when the LLL model is trained on a new task, we assign a teacher model to first learn the new task, and pass the knowledge to the LLL model via knowledge distillation.Therefore, the LLL model can better adapt to the new task while keeping the previously learned knowledge.Experiments show that the proposed L2KD consistently improves previous state-ofthe-art models, and the degradation comparing to multi-task models in LLL tasks is well mitigated for both sequence generation and text classification tasks. 1 Yung-Sung Chuang, Shang-Yu Su, Yun-Nung Chen |
EMNLP (1) | 3 |
| 2020 | What Do Position Embeddings Learn? An Empirical Study of Pre-Trained Language Model Positional EncodingabstractIn recent years, pre-trained Transformers have dominated the majority of NLP benchmark tasks.Many variants of pre-trained Transformers have kept breaking out, and most focus on designing different pre-training objectives or variants of self-attention.Embedding the position information in the self-attention mechanism is also an indispensable factor in Transformers however is often discussed at will.Therefore, this paper carries out an empirical study on position embeddings of mainstream pre-trained Transformers, which mainly focuses on two questions: 1) Do position embeddings really learn the meaning of positions?2) How do these different learned position embeddings affect Transformers for NLP tasks?This paper focuses on providing a new insight of pre-trained position embeddings through feature-level analysis and empirical experiments on most of iconic NLP tasks.It is believed that our experimental results can guide the future work to choose the suitable positional encoding function for specific tasks given the application property.1 Yu-An Wang, Yun-Nung Chen |
EMNLP (1) | 2 |
| 2020 | Learning Asr-Robust Contextualized Embeddings for Spoken Language UnderstandingabstractEmploying pre-trained language models (LM) to extract contextualized word representations has achieved state-of-the-art performance on various NLP tasks. However, applying this technique to noisy transcripts generated by automatic speech recognizer (ASR) is concerned. Therefore, this paper focuses on making contextualized representations more ASR-robust. We propose a novel confusion-aware fine-tuning method to mitigate the impact of ASR errors on pre-trained LMs. Specifically, we fine-tune LMs to produce similar representations for acoustically confusable words that are obtained from word confusion networks (WCNs) produced by ASR. Experiments on multiple benchmark datasets show that the proposed method significantly improves the performance of spoken language understanding when performing on ASR transcripts. Chao-Wei Huang, Yun-Nung Chen |
ICASSP | 2 |
| 2020 | TaylorGAN: Neighbor-Augmented Policy Update Towards Sample-Efficient Natural Language GenerationabstractScore function-based natural language generation (NLG) approaches such as REINFORCE, in general, suffer from low sample efficiency and training instability problems. This is mainly due to the non-differentiable nature of the discrete space sampling and thus these methods have to treat the discriminator as a black box and ignore the gradient information. To improve the sample efficiency and reduce the variance of REINFORCE, we propose a novel approach, TaylorGAN, which augments the gradient estimation by off-policy update and the first-order Taylor expansion. This approach enables us to train NLG models from scratch with smaller batch size --- without maximum likelihood pre-training, and outperforms existing GAN-based methods on multiple metrics of quality and diversity. Chun-Hsing Lin, Siang-Ruei Wu, Hung-yi Lee, Yun-Nung Chen |
NeurIPS | 4 |
| 2020 | Learning Multi-Level Information for Dialogue Response Selection by Highway Recurrent Transformer
Ting-Rui Chiang, Chao-Wei Huang, Shang-Yu Su, Yun-Nung Chen |
Comput. Speech Lang. | 4 |
| 2020 | RAP-Net: Recurrent Attention Pooling Networks for Dialogue Response Selection
Chao-Wei Huang, Ting-Rui Chiang, Shang-Yu Su, Yun-Nung Chen |
Comput. Speech Lang. | 4 |
| 2020 | Knowledge-Grounded Response Generation with Deep Attentional Latent-Variable Model
Hao-Tong Ye, Kai-Ling Lo, Shang-Yu Su, Yun-Nung Chen |
Comput. Speech Lang. | 4 |
| 2019 | Dual Supervised Learning for Natural Language Understanding and GenerationabstractNatural language understanding (NLU) and natural language generation (NLG) are both critical research topics in the NLP and dialogue fields.Natural language understanding is to extract the core semantic meaning from the given utterances, while natural language generation is opposite, of which the goal is to construct corresponding sentences based on the given semantics.However, such dual relationship has not been investigated in literature.This paper proposes a novel learning framework for natural language understanding and generation on top of dual supervised learning, providing a way to exploit the duality.The preliminary experiments show that the proposed approach boosts the performance for both tasks, demonstrating the effectiveness of the dual relationship.1 Shang-Yu Su, Chao-Wei Huang, Yun-Nung Chen |
ACL (1) | 3 |
| 2019 | Adapting Pretrained Transformer to Lattices for Spoken Language UnderstandingabstractLattices are compact representations that encode multiple hypotheses, such as speech recognition results or different word segmentations. It is shown that encoding lattices as opposed to 1-best results generated by automatic speech recognizer (ASR) boosts the performance of spoken language understanding (SLU). Recently, pre-trained language models with the transformer architecture have achieved the state-of-the-art results on natural language understanding, but their ability of encoding lattices has not been explored. Therefore, this paper aims at adapting pre-trained transformers to lattice inputs in order to perform understanding tasks specifically for spoken language. Our experiments on the benchmark ATIS dataset show that fine-tuning pre-trained transformers with lattice inputs yields clear improvement over fine-tuning with 1-best results. Further evaluation demonstrates the effectiveness of our methods under different acoustic conditions11The code is available at https://github.com/MiuLab/Lattice-SLU. Chao-Wei Huang, Yun-Nung Chen |
ASRU | 2 |
| 2019 | Dialogue Environments are Different from Games: Investigating Variants of Deep Q-Networks for Dialogue PolicyabstractThe dialogue manager is an important component in a task-oriented dialogue system, which focuses on deciding dialogue policy given the dialogue state in order to fulfill the user goal. Learning dialogue policy is usually framed as a reinforcement learning (RL) problem, where the objective is to maximize the reward indicating whether the conversation is successful and how efficient it is. However, even there are many variants of deep Q-networks (DQN) achieving better performance on game playing scenarios, no prior work analyzed the performance of dialogue policy learning using these improved versions. Considering that dialogue interactions differ a lot from game playing, this paper investigates variants of DQN models together with different exploration strategies in a benchmark experimental setup, and then we examine which RL methods are more suitable for task-completion dialogue policy learning1. Yu-An Wang, Yun-Nung Chen |
ASRU | 2 |
| 2019 | What Does This Word Mean? Explaining Contextualized Embeddings with Natural Language DefinitionabstractTing-Yun Chang, Yun-Nung Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ting-Yun Chang, Yun-Nung Chen |
EMNLP/IJCNLP (1) | 2 |
| 2019 | DyKgChat: Benchmarking Dialogue Generation Grounding on Dynamic Knowledge GraphsabstractYi-Lin Tuan, Yun-Nung Chen, Hung-yi Lee. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yi-Lin Tuan, Yun-Nung Chen, Hung-yi Lee |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Tree Transformer: Integrating Tree Structures into Self-AttentionabstractYaushian Wang, Hung-Yi Lee, Yun-Nung Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yau-Shian Wang, Hung-yi Lee, Yun-Nung Chen |
EMNLP/IJCNLP (1) | 3 |
| 2019 | QAInfomax: Learning Robust Question Answering System by Mutual Information MaximizationabstractYi-Ting Yeh, Yun-Nung Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yun-Nung Chen |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Mitigating the Impact of Speech Recognition Errors on Spoken Question Answering by Adversarial Domain AdaptationabstractSpoken question answering (SQA) is challenging due to complex reasoning on top of the spoken documents. The recent studies have also shown the catastrophic impact of automatic speech recognition (ASR) errors on SQA. Therefore, this work proposes to mitigate the ASR errors by aligning the mismatch between ASR hypotheses and their corresponding reference transcriptions. An adversarial model is applied to this domain adaptation task, which forces the model to learn domain-invariant features the QA model can effectively utilize in order to improve the SQA results. The experiments successfully demonstrate the effectiveness of our proposed model, and the results are better than the previous best model by 2% EM score. Chia-Hsuan Lee 0001, Yun-Nung Chen, Hung-yi Lee |
ICASSP | 2 |
| 2019 | Compound Variational Auto-encoderabstractAmortized variational inference (AVI) enables efficient training of deep generative models to scale to large datasets. The quality of the approximate inference is determined by various reasons, such as the ability of producing proper variational parameters for each datapoint in the recognition network and whether the variational distribution matches the true posterior, etc. This paper focuses on the inference sub-optimality of variational auto-encoders (VAE), where the goal is to reduce the difference caused by amortizing the variational distribution parameters over the entire training set instead of optimizing for each training example individually, which is also known as the amortization gap. This paper extends Bayesian inference in VAE from the latent level to both latent and weight levels by adopting Bayesian neural networks (BNN) in the encoder, so that each datapoint obtains its own distribution for better modeling. The hybrid design in the proposed compound VAE is empirically demonstrated to be capable of mitigating the amortization gap1. Shang-Yu Su, Shan-Wei Lin, Yun-Nung Chen |
ICASSP | 3 |
| 2019 | Dynamically Context-sensitive Time-decay Attention for Dialogue ModelingabstractSpoken language understanding (SLU) is an essential component in conversational systems. Considering that contexts provide informative cues for better understanding, history can be leveraged for contextual SLU. However, most prior work only paid attention to the related history utterances and ignored the temporal information. In dialogues, prior work considers an inflexible decaying time-aware attention to allow the model to pay more attention to the most recent utterances than the least recent ones. To improve the flexibility of function design, this paper allows the model to automatically learn a time-decay attention function where the attentional weights can be dynamically decided based on the content of each role's contexts, which effectively integrates both content-aware and time-aware perspectives and demonstrates remarkable flexibility to complex dialogue contexts. The experiments on the benchmark Dialogue State Tracking Challenge (DSTC4) dataset show that the proposed dynamically context-sensitive time-decay attention mechanisms significantly improve the state-of-the-art model for contextual understanding performance1. Shang-Yu Su, Pei-Chieh Yuan, Yun-Nung Chen |
ICASSP | 3 |
| 2019 | Modeling Melodic Feature Dependency with Modularized Variational Auto-encoderabstractAutomatic melody generation has been a long-time aspiration for both AI researchers and musicians. However, learning to generate euphonious melodies has turned out to be highly challenging. This paper introduces 1) a new variant of variational autoencoder (VAE), where the model structure is designed in a modularized manner in order to model polyphonic and dynamic music with domain knowledge, and 2) a hierarchical encoding/decoding strategy, which explicitly models the dependency between melodic features. The proposed framework is capable of generating distinct melodies that sounds natural, and the experiments for evaluating generated music clips show that the proposed model outperforms the baselines in human evaluation.1 Yu-An Wang, Tzu-Chuan Lin, Shang-Yu Su, Yun-Nung Chen |
ICASSP | 5 |
| 2018 | CLUSE: Cross-Lingual Unsupervised Sense EmbeddingsabstractThis paper proposes a modularized sense induction and representation learning model that jointly learns bilingual sense embeddings that align well in the vector space, where the crosslingual signal in the English-Chinese parallel corpus is exploited to capture the collocation and distributed characteristics in the language pair.The model is evaluated on the Stanford Contextual Word Similarity (SCWS) dataset to ensure the quality of monolingual sense embeddings.In addition, we introduce Bilingual Contextual Word Similarity (BCWS), a large and high-quality dataset for evaluating crosslingual sense embeddings, which is the first attempt of measuring whether the learned embeddings are indeed aligned well in the vector space.The proposed approach shows the superior quality of sense embeddings evaluated in both monolingual and bilingual spaces. 1 Ta-Chung Chi, Yun-Nung Chen |
EMNLP | 2 |
| 2018 | Discriminative Deep Dyna-Q: Robust Planning for Dialogue Policy LearningabstractThis paper presents a Discriminative Deep Dyna-Q (D3Q) approach to improving the effectiveness and robustness of Deep Dyna-Q (DDQ), a recently proposed framework that extends the Dyna-Q algorithm to integrate planning for task-completion dialogue policy learning.To obviate DDQ's high dependency on the quality of simulated experiences, we incorporate an RNN-based discriminator in D3Q to differentiate simulated experience from real user experience in order to control the quality of training data.Experiments show that D3Q significantly outperforms DDQ by controlling the quality of simulated experience used for planning.The effectiveness and robustness of D3Q is further demonstrated in a domain extension setting, where the agent's capability of adapting to a changing environment is tested. 1 Shang-Yu Su, Xiujun Li, Jianfeng Gao 0001, Jingjing Liu 0001, Yun-Nung Chen |
EMNLP | 5 |
| 2018 | Adversarial Advantage Actor-Critic Model for Task-Completion Dialogue Policy LearningabstractThis paper presents a new method - adversarial advantage actor-critic (Adversarial A2C), which significantly improves the efficiency of dialogue policy learning in task-completion dialogue systems. Inspired by generative adversarial networks (GAN), we train a discriminator to differentiate responses/actions generated by dialogue agents from responses/actions by experts. Then, we incorporate the discriminator as another critic into the advantage actor-critic (A2C) framework, to encourage the dialogue agent to explore state-action within the regions where the agent takes actions similar to those of the experts. Experimental results in a movie-ticket booking domain show that the proposed Adversarial A2C can accelerate policy exploration efficiently. Baolin Peng, Xiujun Li, Jianfeng Gao 0001, Jingjing Liu 0001, Yun-Nung Chen, Kam-Fai Wong |
ICASSP | 5 |
| 2018 | How Time Matters: Learning Time-Decay Attention for Contextual Spoken Language Understanding in DialoguesabstractShang-Yu Su, Pei-Chieh Yuan, Yun-Nung Chen. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Shang-Yu Su, Pei-Chieh Yuan, Yun-Nung Chen |
NAACL-HLT | 3 |
| 2018 | Abstractive Dialogue Summarization with Sentence-Gated Modeling Optimized by Dialogue ActsabstractNeural abstractive summarization has been increasingly studied, where the prior work mainly focused on summarizing single-speaker documents (news, scientific publications, etc). In dialogues, there are diverse interactive patterns between speakers, which are usually defined as dialogue acts. The interactive signals may provide informative cues for better summarizing dialogues. This paper proposes to explicitly leverage dialogue acts in a neural summarization model, where a sentence-gated mechanism is designed for modeling the relationships between dialogue acts and the summary. The experiments show that our proposed model significantly improves the abstractive summarization performance compared to the state-of-the-art baselines on the AMI meeting corpus, demonstrating the usefulness of the interactive signal provided by dialogue acts.1 Chih-Wen Goo, Yun-Nung Chen |
SLT | 2 |
| 2018 | Investigating Linguistic Pattern Ordering In Hierarchical Natural Language GenerationabstractNatural language generation (NLG) is a critical component in spoken dialogue system, which can be divided into two phases: (1) sentence planning: deciding the overall sentence structure, (2) surface realization: determining specific word forms and flattening the sentence structure into a string. With the rise of deep learning, most modern NLG models are based on a sequence-to-sequence (seq2seq) model, which basically contains an encoder-decoder structure; these NLG models generate sentences from scratch by jointly optimizing sentence planning and surface realization. However, such simple encoder-decoder architecture usually fail to generate complex and long sentences, because the decoder has difficulty learning all grammar and diction knowledge well. This paper introduces an NLG model with a hierarchical attentional decoder, where the hierarchy focuses on leveraging linguistic knowledge in a specific order. The experiments show that the proposed method significantly outperforms the traditional seq2seq model with a smaller model size, and the design of the hierarchical attentional decoder can be applied to various NLG systems. Furthermore, different generation strategies based on linguistic patterns are investigated and analyzed in order to guide future NLG research work1. Shang-Yu Su, Yun-Nung Chen |
SLT | 2 |
| 2017 | Towards End-to-End Reinforcement Learning of Dialogue Agents for Information AccessabstractBhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, Li Deng. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Bhuwan Dhingra, Lihong Li 0001, Xiujun Li, Jianfeng Gao 0001, Yun-Nung Chen, Faisal Ahmed 0001, Li Deng 0001 |
ACL (1) | 5 |
| 2017 | Dynamic time-aware attention to speaker roles and contexts for spoken language understandingabstractSpoken language understanding (SLU) is an essential component in conversational systems. Most SLU component treats each utterance independently, and then the following components aggregate the multi-turn information in the separate phases. In order to avoid error propagation and effectively utilize contexts, prior work leveraged history for contextual SLU. However, the previous model only paid attention to the content in history utterances without considering their temporal information and speaker roles. In the dialogues, the most recent utterances should be more important than the least recent ones. Furthermore, users usually pay attention to 1) self history for reasoning and 2) others utterances for listening, the speaker of the utterances may provides informative cues to help understanding. Therefore, this paper proposes an attention-based network that additionally leverages temporal information and speaker role for better SLU, where the attention to contexts and speaker roles can be automatically learned in an end-to-end manner. The experiments on the benchmark Dialogue State Tracking Challenge 4 (DSTC4) dataset show that the time-aware dynamic role attention networks significantly improve the understanding performance. Po-Chun Chen, Ta-Chung Chi, Shang-Yu Su, Yun-Nung Chen |
ASRU | 4 |
| 2017 | MUSE: Modularizing Unsupervised Sense EmbeddingsabstractThis paper proposes to address the word sense ambiguity issue in an unsupervised manner, where word sense representations are learned along a word sense selection mechanism given contexts.Prior work focused on designing a single model to deliver both mechanisms, and thus suffered from either coarse-grained representation learning or inefficient sense selection.The proposed modular approach, MUSE, implements flexible modules to optimize distinct mechanisms, achieving the first purely sense-level representation learning system with linear-time sense selection.We leverage reinforcement learning to enable joint training on the proposed modules, and introduce various exploration techniques on sense selection for better robustness.The experiments on benchmark data show that the proposed approach achieves the state-of-the-art performance on synonym selection as well as on contextual word similarities in terms of MaxSimC. Guang-He Lee, Yun-Nung Chen |
EMNLP | 2 |
| 2017 | End-to-end joint learning of natural language understanding and dialogue managerabstractNatural language understanding and dialogue policy learning are both essential in conversational systems that predict the next system actions in response to a current user utterance. Conventional approaches aggregate separate models of natural language understanding (NLU) and system action prediction (SAP) as a pipeline that is sensitive to noisy outputs of error-prone NLU. To address the issues, we propose an end-to-end deep recurrent neural network with limited contextual dialogue memory by jointly training NLU and SAP on DSTC4 multi-domain human-human dialogues. Experiments show that our proposed model significantly outperforms the state-of-the-art pipeline models for both NLU and SAP, which indicates that our joint model is capable of mitigating the affects of noisy NLU outputs, and NLU model can be refined by error flows backpropagating from the extra supervised signals of system actions. Xuesong Yang, Yun-Nung Chen, Dilek Hakkani-Tür, Paul A. Crook, Xiujun Li, Jianfeng Gao 0001, Li Deng 0001 |
ICASSP | 2 |
| 2017 | End-to-End Task-Completion Neural Dialogue SystemsabstractOne of the major drawbacks of modularized task-completion dialogue systems is that each module is trained individually, which presents several challenges. For example, downstream modules are affected by earlier modules, and the performance of the entire system is not robust to the accumulated errors. This paper presents a novel end-to-end learning framework for task-completion dialogue systems to tackle such issues. Our neural dialogue system can directly interact with a structured database to assist users in accessing information and accomplishing certain tasks. The reinforcement learning based dialogue manager offers robust capabilities to handle noises caused by other components of the dialogue system. Our experiments in a movie-ticket booking domain show that our end-to-end system not only outperforms modularized dialogue system baselines for both objective and subjective evaluation, but also is robust to noises as demonstrated by several systematic experiments with different error granularity and rates specific to the language understanding module. Xiujun Li, Yun-Nung Chen, Lihong Li 0001, Jianfeng Gao 0001, Asli Celikyilmaz |
IJCNLP(1) | 2 |
| 2017 | Order-Preserving Abstractive Summarization for Spoken Content Based on Connectionist Temporal ClassificationabstractConnectionist temporal classification (CTC) is a powerful approach for sequence-to-sequence learning, and has been popularly used in speech recognition.The central ideas of CTC include adding a label "blank" during training.With this mechanism, CTC eliminates the need of segment alignment, and hence has been applied to various sequence-to-sequence learning problems.In this work, we applied CTC to abstractive summarization for spoken content.The "blank" in this case implies the corresponding input data are less important or noisy; thus it can be ignored.This approach was shown to outperform the existing methods in term of ROUGE scores over Chinese Gigaword and MATBN corpora.This approach also has the nice property that the ordering of words or characters in the input documents can be better preserved in the generated summaries. Bo-Ru Lu, Frank Shyu, Yun-Nung Chen, Hung-yi Lee, Lin-Shan Lee |
INTERSPEECH | 3 |
| 2016 | Zero-shot learning of intent embeddings for expansion by convolutional deep structured semantic modelsabstractThe recent surge of intelligent personal assistants motivates spoken language understanding of dialogue systems. However, the domain constraint along with the inflexible intent schema remains a big issue. This paper focuses on the task of intent expansion, which helps remove the domain limit and make an intent schema flexible. A con-volutional deep structured semantic model (CDSSM) is applied to jointly learn the representations for human intents and associated utterances. Then it can flexibly generate new intent embeddings without the need of training samples and model-retraining, which bridges the semantic relation between seen and unseen intents and further performs more robust results. Experiments show that CDSSM is capable of performing zero-shot learning effectively, e.g. generating embeddings of previously unseen intents, and therefore expand to new intents without re-training, and outperforms other semantic embeddings. The discussion and analysis of experiments provide a future direction for reducing human effort about annotating data and removing the domain constraint in spoken dialogue systems. Yun-Nung Chen, Dilek Hakkani-Tür, Xiaodong He 0001 |
ICASSP | 1 |
| 2016 | Unsupervised user intent modeling by feature-enriched matrix factorizationabstractSpoken language interfaces are being incorporated into various devices such as smart phones and TVs. However, dialogue systems may fail to respond correctly when users' request functionality is not supported by currently installed apps. This paper proposes a feature-enriched matrix factorization (MF) approach to model open domain intents, which allows a system to dynamically add unexplored domains according to users' requests. First we leverage the structured knowledge from Wikipedia and Freebase to automatically acquire domain-related semantics to enrich features of input utterances, and then MF is applied to model automatically acquired knowledge, published app textual descriptions and users' spoken requests in a joint fashion; this generates latent feature vectors for utterances and user intents without need of prior annotations. Experiments show that the proposed MF models incorporated with rich features significantly improve intent prediction, achieving about 34% of mean average precision (MAP) for both ASR and manual transcripts. Yun-Nung Chen, Ming Sun 0001, Alexander I. Rudnicky, Anatole Gershman |
ICASSP | 1 |
| 2016 | End-to-End Memory Networks with Knowledge Carryover for Multi-Turn Spoken Language UnderstandingabstractSpoken language understanding (SLU) is a core component of a spoken dialogue system. In the traditional architecture of dialogue systems, the SLU component treats each utterance independent of each other, and then the following components aggregate the multi-turn information in the separate phases. However, there are two challenges: 1) errors from previous turns may be propagated and then degrade the performance of the current turn; 2) knowledge mentioned in the long history may not be carried into the current turn. This paper addresses the above issues by proposing an architecture using end-to-end memory networks to model knowledge carryover in multi-turn conversations, where utterances encoded with intents and slots can be stored as embeddings in the memory and the decoding phase applies an attention model to leverage previously stored semantics for intent prediction and slot tagging simultaneously. The experiments on Microsoft Cortana conversational data show that the proposed memory network architecture can effectively extract salient semantics for modeling knowledge carryover in the multi-turn conversations and outperform the results using the state-of-the-art recurrent neural network framework (RNN) designed for single-turn SLU. Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür, Jianfeng Gao 0001, Li Deng 0001 |
INTERSPEECH | 1 |
| 2016 | Multi-Domain Joint Semantic Frame Parsing Using Bi-Directional RNN-LSTMabstractSequence-to-sequence deep learning has recently emerged as a new paradigm in supervised learning for spoken language understanding. However, most of the previous studies explored this framework for building single domain models for each task, such as slot filling or domain classification, comparing deep learning based approaches with conventional ones like conditional random fields. This paper proposes a holistic multi-domain, multi-task (i.e. slot filling, domain and intent detection) modeling approach to estimate complete semantic frames for all user utterances addressed to a conversational system, demonstrating the distinctive power of deep learning methods, namely bi-directional recurrent neural network (RNN) with long-short term memory (LSTM) cells (RNN-LSTM) to handle such complexity. The contributions of the presented work are three-fold: (i) we propose an RNN-LSTM architecture for joint modeling of slot filling, intent determination, and domain classification; (ii) we build a joint multi-domain model enabling multi-task deep learning where the data from each domain reinforces each other; (iii) we investigate alternative architectures for modeling lexical context in spoken language understanding. In addition to the simplicity of the single model framework, experimental results show the power of such an approach on Microsoft Cortana real user data over alternative methods based on single domain/task deep learning. Dilek Hakkani-Tür, Gökhan Tür, Asli Celikyilmaz, Yun-Nung Chen, Jianfeng Gao 0001, Li Deng 0001, Ye-Yi Wang |
INTERSPEECH | 4 |
| 2016 | An Intelligent Assistant for High-Level Task UnderstandingabstractPeople are able to interact with domain-specific intelligent assistants (IAs) and get help with tasks. But sometimes user goals are complex and may require interactions with multiple applications. However current IAs are limited to specific applications and users have to directly manage execution spanning multiple applications in order to engage in more complex activities. An ideal personal agent would be able to learn, over time, about tasks that span different resources. This paper addresses the problem of cross-domain task assistance in the context of spoken dialogue systems. We propose approaches to discover users' high-level intentions and using this information to assist users in their task. We collected real-life smartphone usage data from 14 participants and investigated how to extract high-level intents from users' descriptions of their activities. Our experiments show that understanding high-level tasks allows the agent to actively suggest apps relevant to pursuing particular user goals and reduce the cost of users' self-management. Ming Sun 0001, Yun-Nung Chen, Alexander I. Rudnicky |
IUI | 2 |
| 2016 | AIMU: Actionable Items for Meeting Understanding
Yun-Nung Chen, Dilek Hakkani-Tür |
LREC | 1 |
| 2016 | AppDialogue: Multi-App Dialogues for Intelligent Assistants
Ming Sun 0001, Yun-Nung Chen, Zhenhao Hua, Yulian Tamres-Rudnicky, Arnab Dash, Alexander I. Rudnicky |
LREC | 2 |
| 2016 | Syntax or semantics? knowledge-guided joint semantic frame parsingabstractSpoken language understanding (SLU) is a core component of a spoken dialogue system, which involves intent prediction and slot filling and also called semantic frame parsing. Recently recurrent neural networks (RNN) obtained strong results on SLU due to their superior ability of preserving sequential information over time. Traditionally, the SLU component parses semantic frames for utterances considering their flat structures, as the underlying RNN structure is a linear chain. However, natural language exhibits linguistic properties that provide rich, structured information for better understanding. This paper proposes to apply knowledge-guided structural attention networks (K-SAN), which additionally incorporate non-flat network topologies guided by prior knowledge, to a language understanding task. The model can effectively figure out the salient substructures that are essential to parse the given utterance into its semantic frame with an attention mechanism, where two types of knowledge, syntax and semantics, are utilized. The experiments on the benchmark Air Travel Information System (ATIS) data and the conversational assistant Cortana data show that 1) the proposed K-SAN models with syntax or semantics outperform the state-of-the-art neural network based results, and 2) the improvement for joint semantic frame parsing is more significant, because the structured information provides rich cues for sentence-level understanding, where intent prediction and slot filling can be mutually improved. Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür, Asli Celikyilmaz, Jianfeng Gao 0001, Li Deng 0001 |
SLT | 1 |
| 2016 | Weakly supervised user intent detection for multi-domain dialoguesabstractUsers interact with mobile apps with certain intents such as finding a restaurant. Some intents and their corresponding activities are complex and may involve multiple apps; for example, a restaurant app, a messenger app and a calendar app may be needed to plan a dinner with friends. However, activities may be quite personal and third-party developers would not be building apps to specifically handle complex intents (e.g., a DinnerPlanner). Instead we want our intelligent agent to actively learn to understand these intents and provide assistance when needed. This paper proposes a framework to enable the agent to learn an inventory of intents from a small set of task-oriented user utterances. The experiments show that on previously unseen user activities, the agent is able to reliably recognize user intents using graph-based semi-supervised learning methods. The dataset, models, and the system outputs are available to research community. Ming Sun 0001, Aasish Pappu, Yun-Nung Chen, Alexander I. Rudnicky |
SLT | 3 |
| 2015 | Matrix Factorization with Knowledge Graph Propagation for Unsupervised Spoken Language UnderstandingabstractYun-Nung Chen, William Yang Wang, Anatole Gershman, Alexander Rudnicky. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Yun-Nung Chen, William Yang Wang, Anatole Gershman, Alexander I. Rudnicky |
ACL (1) | 1 |
| 2015 | Detecting actionable items in meetings by convolutional deep structured semantic modelsabstractThe recent success of voice interaction with smart devices (human-machine genre) and improvements in speech recognition for conversational speech show the possibility of conversation-related applications. This paper investigates the task of actionable item detection in meetings (human-human genre), where the intelligent assistant dynamically provides the participants access to information (e.g. scheduling a meeting, taking notes) without interrupting the meetings. A convolutional deep structured semantic model (CDSSM) is applied to learn the latent semantics for human actions and utterances from human-machine (source genre) and human-human (target) interactions. Furthermore, considering the mismatch between source and target genre and scarcity of annotated data sets for the target genre, we develop adaptation techniques that adjust the learned embeddings to better fit the target genre. Experiments show that CDSSM performs better for actionable item detection compared to baselines using lexical features (27.5% relative) and other semantic features (15.9% relative) when the source genre and target genre match with each other. When the target genre mismatches with the source genre, our proposed adaptation techniques further improve the performance. The discussion and analysis of the experiments provide a reasonable direction for such an actionable item detection task1. Yun-Nung Chen, Dilek Hakkani-Tür, Xiaodong He 0001 |
ASRU | 1 |
| 2015 | Leveraging Behavioral Patterns of Mobile Applications for Personalized Spoken Language UnderstandingabstractSpoken language interfaces are appearing in various smart devices (e.g. smart-phones, smart-TV, in-car navigating systems) and serve as intelligent assistants (IAs). However, most of them do not consider individual users' behavioral profiles and contexts when modeling user intents. Such behavioral patterns are user-specific and provide useful cues to improve spoken language understanding (SLU). This paper focuses on leveraging the app behavior history to improve spoken dialog systems performance. We developed a matrix factorization approach that models speech and app usage patterns to predict user intents (e.g. launching a specific app). We collected multi-turn interactions in a WoZ scenario; users were asked to reproduce the multi-app tasks that they had performed earlier on their smart-phones. By modeling latent semantics behind lexical and behavioral patterns, the proposed multi-model system achieves about 52% of turn accuracy for intent prediction on ASR transcripts. Yun-Nung Chen, Ming Sun 0001, Alexander I. Rudnicky, Anatole Gershman |
ICMI | 1 |
| 2015 | Learning semantic hierarchy with distributed representations for unsupervised spoken language understandingabstractWe study the problem of unsupervised ontology learning for semantic understanding in spoken dialogue systems, in particular, learning the hierarchical semantic structure from the data. Given unlabelled conversations, we augment a frame-semantic based unsupervised slot induction approach with hierarchical agglomerative clustering to merge topically-related slots (e.g., both slots “direction” and “locale” convey location-related information) for building a coherent semantic hierarchy, and then estimate the slot importance at different levels. The high-level semantic estimation involves not only within-slot but also crossslot relations. The experiments show that high-level semantic information can accurately estimate the prominence of slots, significantly improving the slot induction performance; furthermore, a semantic decoder trained on the data with automatically extracted slots achieves about 68% F-measure, which is close to the one from hand-crafted grammars. Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky |
INTERSPEECH | 1 |
| 2015 | Learning OOV through semantic relatedness in spoken dialog systemsabstract• Speech recognition and language understanding performance can be improved through an OOV expectand-learn procedure. • A limited domain vocabulary can be utilized to effectively acquire OOVs by the word relatedness theory through web knowledge bases. • With data-driven semantic relatedness, both the global and local learning procedures are able to successfully harvest more than 50% of OOVs, leading to better recognition and understanding performance. • This work demonstrates that o OOV learning may benefit dialog system o the proposed expect-and-learn strategy outperforms the traditional detect-and-learn in both higher effectiveness and no human involvement. 1. Linguistically semantic relatedness o Defined by linguistics, e.g., WordNet (WN), Paraphrase Database (PPDB) (Ganitkevitch et al., 2013) 2. Data-driven semantic relatedness o Distributional semantics, e.g., continuous bag-ofword embeddings (CBOW) (Mikolov et al., 2013) Detect-and-Learn (Qin et al., 2011; 2012): o Discover OOV words during the conversation o Example: S: “I heard something like SELF, can you repeat it?” U: “It’s SELFIE.” o Drawbacks • Limited number of new words • Required human efforts to correct spellings and pronunciations Expect-and-Learn (proposed): o Use semantic relatedness to automatically enrich the vocabulary and language model beforehand Ming Sun 0001, Yun-Nung Chen, Alexander I. Rudnicky |
INTERSPEECH | 2 |
| 2015 | Jointly Modeling Inter-Slot Relations by Random Walk on Knowledge Graphs for Unsupervised Spoken Language UnderstandingabstractYun-Nung Chen, William Yang Wang, Alexander Rudnicky. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky |
HLT-NAACL | 1 |
| 2014 | Two-Stage Stochastic Email SynthesizerabstractThis paper presents the design and im-plementation details of an email synthe-sizer using two-stage stochastic natural language generation, where the first stage structures the emails according to sender style and topic structure, and the second stage synthesizes text content based on the particulars of an email structure element and the goals of a given communication for surface realization. The synthesized emails reflect sender style and the intent of communication, which can be further used as synthetic evidence for developing other applications. 1 Yun-Nung Chen, Alexander I. Rudnicky |
INLG | 1 |
| 2014 | Two-Stage Stochastic Natural Language Generation for Email Synthesis by Modeling Sender Style and Topic StructureabstractThis paper describes a two-stage pro-cess for stochastic generation of email, in which the first stage structures the emails according to sender style and topic struc-ture (high-level generation), and the sec-ond stage synthesizes text content based on the particulars of an email element and the goals of a given communication (surface-level realization). Synthesized emails were rated in a preliminary experi-ment. The results indicate that sender style can be detected. In addition we found that stochastic generation performs better if applied at the word level than at an original-sentence level (“template-based”) in terms of email coherence, sentence flu-ency, naturalness, and preference. 1 Yun-Nung Chen, Alexander I. Rudnicky |
INLG | 1 |
| 2014 | Deriving local relational surface forms from dependency-based entity embeddings for unsupervised spoken language understandingabstractRecent works showed the trend of leveraging web-scaled structured semantic knowledge resources such as Freebase for open domain spoken language understanding (SLU). Knowledge graphs provide sufficient but ambiguous relations for the same entity, which can be used as statistical background knowledge to infer possible relations for interpretation of user utterances. This paper proposes an approach to capture the relational surface forms by mapping dependency-based contexts of entities from the text domain to the spoken domain. Relational surface forms are learned from dependency-based entity embeddings, which encode the contexts of entities from dependency trees in a deep learning model. The derived surface forms carry functional dependency to the entities and convey the explicit expression of relations. The experiments demonstrate the efficiency of leveraging derived relational surface forms as local cues together with prior background knowledge. Yun-Nung Chen, Dilek Hakkani-Tür, Gökhan Tür |
SLT | 1 |
| 2014 | Dynamically supporting unexplored domains in conversational interactions by enriching semantics with neural word embeddingsabstractSpoken language interfaces are being incorporated into various devices (e.g. smart-phones, smart TVs, etc). However, current technology typically limits conversational interactions to a few narrow predefined domains/topics. For example, dialogue systems for smartphone operation fail to respond when users ask for functions not supported by currently installed applications. We propose to dynamically add application-based domains according to users' requests by using descriptions of applications as a retrieval cue to find relevant applications. The approach uses structured knowledge resources (e.g. Freebase, Wikipedia, FrameNet) to induce types of slots for generating semantic seeds, and enriches the semantics of spoken queries with neural word embeddings, where semantically related concepts can be additionally included for acquiring knowledge that does not exist in the predefined domains. The system can then retrieve relevant applications or dynamically suggest users install applications that support unexplored domains. We find that vendor descriptions provide a reliable source of information for this purpose. Yun-Nung Chen, Alexander I. Rudnicky |
SLT | 1 |
| 2014 | Leveraging frame semantics and distributional semantics for unsupervised semantic slot induction in spoken dialogue systemsabstractDistributional semantics and frame semantics are two representative views on language understanding in the statistical world and the linguistic world, respectively. In this paper, we combine the best of two worlds to automatically induce the semantic slots for spoken dialogue systems. Given a collection of unlabeled audio files, we exploit continuous-valued word embeddings to augment a probabilistic frame-semantic parser that identifies key semantic slots in an unsupervised fashion. In experiments, our results on a real-world spoken dialogue dataset show that the distributional word representations significantly improve the adaptation of FrameNet-style parses of ASR decodings to the target semantic space; that comparing to a state-of-the-art baseline, a 13% relative average precision improvement is achieved by leveraging word vectors trained on two 100-billion words datasets; and that the proposed technology can be used to reduce the costs for designing task-oriented spoken dialogue systems. Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky |
SLT | 1 |
| 2014 | Spoken Knowledge Organization by Semantic Structuring and a Prototype Course Lecture System for Personalized LearningabstractIt takes very long time to go through a complete online course. Without proper background, it is also difficult to understand retrieved spoken paragraphs. This paper therefore presents a new approach of spoken knowledge organization for course lectures for efficient personalized learning. Automatically extracted key terms are taken as the fundamental elements of the semantics of the course. Key term graph constructed by connecting related key terms forms the backbone of the global semantic structure. Audio/video signals are divided into multi-layer temporal structure including paragraphs, sections and chapters, each of which includes a summary as the local semantic structure. The interconnection between semantic structure and temporal structure together with spoken term detection jointly offer to the learners efficient ways to navigate across the course knowledge with personalized learning paths considering their personal interests, available time and background knowledge. A preliminary prototype system has also been successfully developed. Hung-yi Lee, Sz-Rung Shiang, Ching-Feng Yeh, Yun-Nung Chen, Sheng-yi Kong, Lin-Shan Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2013 | Unsupervised induction and filling of semantic slots for spoken dialogue systems using frame-semantic parsingabstractSpoken dialogue systems typically use predefined semantic slots to parse users' natural language inputs into unified semantic representations. To define the slots, domain experts and professional annotators are often involved, and the cost can be expensive. In this paper, we ask the following question: given a collection of unlabeled raw audios, can we use the frame semantics theory to automatically induce and fill the semantic slots in an unsupervised fashion? To do this, we propose the use of a state-of-the-art frame-semantic parser, and a spectral clustering based slot ranking model that adapts the generic output of the parser to the target semantic space. Empirical experiments on a real-world spoken dialogue dataset show that the automatically induced semantic slots are in line with the reference slots created by domain experts: we observe a mean averaged precision of 69.36% using ASR-transcribed data. Our slot filling evaluations also indicate the promising future of this proposed approach. Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky |
ASRU | 1 |
| 2013 | An empirical investigation of sparse log-linear models for improved dialogue act classificationabstractPrevious work on dialogue act classification have primarily focused on dense generative and discriminative models. However, since the automatic speech recognition (ASR) outputs are often noisy, dense models might generate biased estimates and overfit to the training data. In this paper, we study sparse modeling approaches to improve dialogue act classification, since the sparse models maintain a compact feature space, which is robust to noise. To test this, we investigate various element-wise frequentist shrinkage models such as lasso, ridge, and elastic net, as well as structured sparsity models and a hierarchical sparsity model that embed the dependency structure and interaction among local features. In our experiments on a real-world dataset, when augmenting N-best word and phone level ASR hypotheses with confusion network features, our best sparse log-linear model obtains a relative improvement of 19.7% over a rule-based baseline, a 3.7% significant improvement over a traditional non-sparse log-linear model, and outperforms a state-of-the-art SVM model by 2.2%. Yun-Nung Chen, William Yang Wang, Alexander I. Rudnicky |
ICASSP | 1 |
| 2013 | Bootstrapping Text-to-Speech for speech processing in languages without an orthographyabstractSpeech synthesis technology has reached the stage where given a well-designed corpus of audio and accurate transcription an at least understandable synthesizer can be built without necessarily resorting to new innovations. However many languages do not have a well-defined writing system but such languages could still greatly benefit from speech systems. In this paper we consider the case where we have a (potentially large) single speaker database but have no transcriptions and no standardized way to write transcriptions. To address this scenario we propose a method that allows us to bootstrap synthetic voices purely from speech data. We use a novel combination of automatic speech recognition and automatic word segmentation for the bootstrapping. Our experimental results on speech corpora in two languages, English and German, show that synthetic voices that are built using this method are close to understandable. Our method is language-independent and can thus be used to build synthetic voices from a speech corpus in any new language. Sunayana Sitaram, Sukhada Palkar, Yun-Nung Chen, Alok Parlikar, Alan W. Black |
ICASSP | 3 |
| 2013 | Prosody-Based Unsupervised Speech Summarization with Two-Layer Mutually Reinforced Random Walk
Sujay Kumar Jauhar, Yun-Nung Chen, Florian Metze |
IJCNLP | 2 |
| 2013 | Multi-layer mutually reinforced random walk with hidden parameters for improved multi-party meeting summarizationabstractThis paper proposes an improved approach of summarization for spoken multi-party interaction, in which a multi-layer graph with hidden parameters is constructed. The graph includes utterance-to-utterance relation, utterance-to-parameter weight, and speaker-to-parameter weight. Each utterance and each speaker are represented as a node in the utterance-layer and speaker-layer of the graph respectively. We use terms/ topics as hidden parameters for estimating utterance-to-parameter and speaker-to-parameter weight, and compute topical similarity between utterances as the utterance-to-utterance relation. By within- and between-layer propagation in the graph, the scores from different layers can be mutually reinforced so that utterances can automatically share the scores with the utterances from the speakers who focus on similar terms/ topics. For both ASR output and manual transcripts, experiments confirmed the efficacy of including hidden parameters and involving speaker information in the multi-layer graph for summarization. We find that choosing latent topics as hidden parameters significantly reduces computational complexity and does not hurt the performance. Yun-Nung Chen, Florian Metze |
INTERSPEECH | 1 |
| 2012 | Unsupervised two-stage keyword extraction from spoken documents by topic coherence and support vector machineabstractThis paper proposes an unsupervised two-stage approach to automatically extract keywords from spoken documents. In the first stage, for each candidate term we compute a topic coherence and term significance measure (TCS) based on probabilistic latent semantic analysis (PLSA) models. In the second stage, we take the candidate terms with highest and lowest TCS scores as positive and negative examples to train an SVM classifier in an unsupervised way using prosodic, lexical, and semantic features, and then classify the candidate keyword using this SVM classifier. The experiments with course lectures showed that the first-stage offered very good precision, so the second-stage effectively extracted the keywords. Yun-Nung Chen, Hung-yi Lee, Lin-Shan Lee |
ICASSP | 1 |
| 2012 | Utterance-level latent topic transition modeling for spoken documents and its application in automatic summarizationabstractIn this paper, we propose to use an utterance-level latent topic transition model to estimate the latent topics behind the utterances, and test the performance of such model in extractive speech summarization. In this model, the latent topic weights behind an utterance are estimated, and these topic weights evolve from an utterance to the next in a spoken document based on a topic transition function represented by a matrix. We explore different ways of obtaining such topic transition matrices used in the model, and find using a set of matrices estimated with utterances clustered from a training spoken document set is very useful. This model was shown to be able to offer extra performance improvement when used with the popularly used Probability Latent Semantic Analysis (PLSA) in preliminary experiments on speech summarization. Hung-yi Lee, Yun-Nung Chen, Lin-Shan Lee |
ICASSP | 2 |
| 2012 | NeuroDialog: an EEG-enabled spoken dialog interfaceabstractUnderstanding user intent is a difficult problem in Dialog Systems, as they often need to make decisions under uncertainty. Using an inexpensive, consumer grade EEG sensor and a Wizard-of-Oz dialog system, we show that it is possible to detect system misunderstanding even before the user reacts vocally. We also present the design and implementation details of NeuroDialog, a proof-of-concept dialog system that uses an EEG based predictive model to detect system misrecognitions during live interaction. Seshadri Sridharan, Yun-Nung Chen, Kai-min Chang, Alexander I. Rudnicky |
ICMI | 2 |
| 2012 | Integrating Intra-Speaker Topic Modeling and Temporal-Based Inter-Speaker Topic Modeling in Random Walk for Improved Multi-Party Meeting SummarizationabstractThis paper proposes an improved approach of summarization for spoken multi-party interaction, in which intra-speaker and inter-speaker topics are modeled in a graph constructed with topical relations. Each utterance is represented as a node of the graph and the edge between two nodes is weighted by the similarity between the two utterances, which is topical similarity evaluated by probabilistic latent semantic analysis (PLSA). We model intra-speaker topics by sharing the topics from the same speaker and inter-speaker topics by partially sharing the topics from the adjacent utterances based on temporal information. We did experiments for ASR and manual transcripts. For both transcripts, experiments showed combining intra-speaker and inter-speaker topic modeling can help include the important utterances to offer the improvement for summarization. Yun-Nung Chen, Florian Metze |
INTERSPEECH | 1 |
| 2012 | Towards Using EEG to Improve ASR Accuracy
Yun-Nung Chen, Kai-min Chang, Jack Mostow |
HLT-NAACL | 1 |
| 2012 | Intra-Speaker Topic Modeling for Improved Multi-Party Meeting Summarization with Integrated Random Walk
Yun-Nung Chen, Florian Metze |
HLT-NAACL | 1 |
| 2012 | Two-layer mutually reinforced random walk for improved multi-party meeting summarizationabstractThis paper proposes an improved approach of summarization for spoken multi-party interaction, in which a two-layer graph with utterance-to-utterance, speaker-to-speaker, and speaker-to-utterance relations is constructed. Each utterance and each speaker are represented as a node in the utterance-layer and speaker-layer of the graph respectively, and the edge between two nodes is weighted by the similarity between the two utterances, the two speakers, or the utterance and the speaker. The relation between utterances is evaluated by lexical similarity via word overlap or topical similarity via probabilistic latent semantic analysis (PLSA). By within- and between-layer propagation in the graph, the scores from different layers can be mutually reinforced so that utterances can automatically share the scores with the utterances from the same speaker and similar utterances. For both ASR output and manual transcripts, experiments confirmed the efficacy of involving speaker information in the two-layer graph for summarization. Yun-Nung Chen, Florian Metze |
SLT | 1 |
| 2011 | Improved spoken term detection with graph-based re-ranking in feature spaceabstractThis paper presents a graph-based approach for spoken term detection. Each first-pass retrieved utterance is a node on a graph and the edge between two nodes is weighted by the similarity between the two utterances evaluated in feature space. The score of each node is then modified by the contributions from its neighbors by random walk or its modified version, because utterances similar to more utterances with higher scores should be given higher relevance scores. In this way the global similarity structure of all first-pass retrieved utterances can be jointly considered. Experimental results show that this new approach offers significantly better performance than the previously proposed pseudo-relevance feedback approach, which considers primarily the local similarity relationship between first-pass retrieved utterances, and these two different approaches can be cascaded to provide even better results. Yun-Nung Chen, Chia-Ping Chen, Hung-yi Lee, Chun-an Chan, Lin-Shan Lee |
ICASSP | 1 |
| 2011 | Spoken Lecture Summarization by Random Walk over a Graph Constructed with Automatically Extracted Key TermsabstractThis paper proposes an improved approach for spoken lecture summarization, in which random walk is performed on a graph constructed with automatically extracted key terms and proba-bilistic latent semantic analysis (PLSA). Each sentence of the document is represented as a node of the graph and the edge be-tween two nodes is weighted by the topical similarity between the two sentences. The basic idea is that sentences topically similar to more important sentences should be more important. In this way all sentences in the document can be jointly consid-ered more globally rather than individually. Experimental re-sults showed significant improvement in terms of ROUGE eval-uation. Index Terms: summarization, course lecture, probabilistic la-tent semantic analysis (PLSA), random walk, key term Yun-Nung Chen, Ching-Feng Yeh, Lin-Shan Lee |
INTERSPEECH | 1 |
| 2010 | Automatic key term extraction from spoken course lectures using branching entropy and prosodic/semantic featuresabstractThis paper proposes a set of approaches to automatically extract key terms from spoken course lectures including audio signals, ASR transcriptions and slides. We divide the key terms into two types: key phrases and keywords and develop different approaches to extract them in order. We extract key phrases using right/left branching entropy and extract keywords by learning from three sets of features: prosodic features, lexical features and semantic features from Probabilistic Latent Semantic Analysis (PLSA). The learning approaches include an unsupervised method (K-means exemplar) and two supervised ones (AdaBoost and neural network). Very encouraging preliminary results were obtained with a corpus of course lectures, and it is found that all approaches and all sets of features proposed here are useful. Yun-Nung Chen, Sheng-yi Kong, Lin-Shan Lee |
SLT | 1 |