VLDB 2026 Research / reviewers in the wild / expert
Yang Liu 0004
dblp:51/3710-4
· DBLP profile ↗
145ranked-venue papers
16as first author
25since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 112 · 11 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 63 · 9 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Knowledge Transfer for Body Sensor Networks: Characterization and Resistance of Negative Transfer in Overconstrained EnvironmentsabstractTransfer learning (TL) has shown potential for body sensor network (BSN) applications. However, TL performance is influenced by various factors, such as the BSN environment, user-specific physiological signs, or sensor modalities, leading to negative transfer (NT) effects. Limited research has defined NT from perspectives, such as data quality, domain transferability, and transfer components, but the cognition of NT in multimodal BSNs remains insufficient. In this article, we explore the correlation between transfer modes and NT, proposing a knowledge transfer framework composed of common domain adaptation constraints to reveal the NT effects in overconstrained environments. This effect can explain why strongly constrained transfer algorithms are effective but not necessarily optimal in terms of performance. In the three typical BSN scenarios of emotion recognition, fall detection, and daily activity recognition, we demonstrate the NT in overconstrained environments and its potential influencing factors using different constraint combinations and based on different multimodal body sensor fusion architectures. Extensive experiments focus on discussing the impact of category-condition constraints, domain similarity, and fusion architectures on the NT effect in BSNs. This work can provide a new perspective for the design of multimodal BSN knowledge transfer scheme. Han Shi 0001, Yang Liu 0004, Hai Zhao 0002 |
IEEE Internet Things J. | 2 |
| 2024 | Overview of the Ninth Dialog System Technology Challenge: DSTC9abstractThis paper introduces the Ninth Dialog System Technology Challenge (DSTC-9). This edition of the DSTC focuses on applying end-to-end dialog technologies for four distinct tasks in dialog systems, namely, 1. Task-oriented dialog Modeling with Unstructured Knowledge Access, 2. Multi-domain task-oriented dialog, 3. Interactive evaluation of dialog and 4. Situated interactive multimodal dialog. This paper describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. R. Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro, Abhinav Rastogi, Yun-Nung Chen, Mihail Eric, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Yang Liu 0004, Chao-Wei Huang, Dilek Hakkani-Tür, Jinchao Li, Qi Zhu 0007, Lingxiao Luo, Lars Liden, Kaili Huang, Shahin Shayandeh, Runze Liang, Baolin Peng, Zheng Zhang 0020, Swadheen Shukla, Minlie Huang, Jianfeng Gao 0001, Shikib Mehri, Yulan Feng, Carla Gordon, Seyed Hossein Alavi, David R. Traum, Maxine Eskénazi, Ahmad Beirami, Eunjoon Cho, Paul A. Crook, Ankita De, Alborz Geramifard, Satwik Kottur, Seungwhan Moon, Shivani Poddar, Rajen Subba |
IEEE ACM Trans. Audio Speech Lang. Process. | 9 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 14 |
| 2023 | Towards Credible Human Evaluation of Open-Domain Dialog Systems Using Interactive SetupabstractEvaluating open-domain conversation models has been an open challenge due to the open-ended nature of conversations. In addition to static evaluations, recent work has started to explore a variety of per-turn and per-dialog interactive evaluation mechanisms and provide advice on the best setup. In this work, we adopt the interactive evaluation framework and further apply to multiple models with a focus on per-turn evaluation techniques. Apart from the widely used setting where participants select the best response among different candidates at each turn, one more novel per-turn evaluation setting is adopted, where participants can select all appropriate responses with different fallback strategies to continue the conversation when no response is selected. We evaluate these settings based on sensitivity and consistency using four GPT2-based models that differ in model sizes or fine-tuning data. To better generalize to any model groups with no prior assumptions on their rankings and control evaluation costs for all setups, we also propose a methodology to estimate the required sample size given a minimum performance gap of interest before running most experiments. Our comprehensive human evaluation results shed light on how to conduct credible human evaluations of open domain dialog systems using the interactive setup, and suggest additional future directions. Sijia Liu 0007, Patrick Lange, Behnam Hedayatnia, Alexandros Papangelis, Di Jin 0005, Andrew Wirth, Yang Liu 0004, Dilek Hakkani-Tür |
AAAI | 7 |
| 2023 | KILM: Knowledge Injection into Encoder-Decoder Language ModelsabstractYan Xu, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yan Xu 0012, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar, Yang Liu 0004, Dilek Hakkani-Tür |
ACL (1) | 5 |
| 2023 | Selective In-Context Data Augmentation for Intent Detection using Pointwise V-InformationabstractYen-Ting Lin, Alexandros Papangelis, Seokhwan Kim, Sungjin Lee, Devamanyu Hazarika, Mahdi Namazifar, Di Jin, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Alexandros Papangelis, Seokhwan Kim, Devamanyu Hazarika, Mahdi Namazifar, Di Jin 0005, Yang Liu 0004, Dilek Hakkani-Tür |
EACL | 8 |
| 2023 | CESAR: Automatic Induction of Compositional Instructions for Multi-turn DialogsabstractInstruction-based multitasking has played a critical role in the success of large language models (LLMs) in multi-turn dialog applications.While publicly-available LLMs have shown promising performance, when exposed to complex instructions with multiple constraints, they lag against state-of-the-art models like Chat-GPT.In this work, we hypothesize that the availability of large-scale complex demonstrations is crucial in bridging this gap.Focusing on dialog applications, we propose a novel framework, CESAR, that unifies a large number of dialog tasks in the same format and allows programmatic induction of complex instructions without any manual effort.We apply CESAR on InstructDial, a benchmark for instruction-based dialog tasks.We further enhance InstructDial with new datasets and tasks and utilize CESAR to induce complex tasks with compositional instructions.This results in a new benchmark called InstructDial++, which includes 63 datasets with 86 basic tasks and 68 composite tasks.Through rigorous experiments, we demonstrate the scalability of CESAR in providing rich instructions.Models trained on InstructDial++ can follow compositional prompts, such as prompts that ask for multiple stylistic constraints. Taha Aksu, Devamanyu Hazarika, Shikib Mehri, Seokhwan Kim, Dilek Hakkani-Tür, Yang Liu 0004, Mahdi Namazifar |
EMNLP | 6 |
| 2023 | Identifying Entrainment in Task-Oriented ConversationsabstractHuman interlocutors adapt their behavior to each other in a conversation through entrainment. While entrainment has been found in long chit-chat conversations, much less research has been conducted on task-oriented dialogs. In this paper, we investigate short task-oriented Wizard-of-Oz conversations for acoustic-prosodic and lexical entrainment. We conduct significance tests that reveal changes in speech pitch and frequent words as important indicators of entrainment. Our findings will guide user-entraining dialog systems to improve the quality of conversations. Run Chen, Seokhwan Kim, Alexandros Papangelis, Julia Hirschberg, Yang Liu 0004, Dilek Hakkani-Tür |
ICASSP | 5 |
| 2023 | MERCY: Multiple Response Ranking Concurrently in Realistic Open-Domain Conversational SystemsabstractAutomatic Evaluation (AE) and Response Selection (RS) models assign quality scores to various candidate responses and rank them in conversational setups.Prior response ranking research compares various models' performance on synthetically generated test sets.In this work, we investigate the performance of model-based reference-free AE and RS models on our constructed response ranking datasets that mirror real-case scenarios of ranking candidates during inference time.Metrics' unsatisfying performance can be interpreted as their low generalizability over more pragmatic conversational domains such as human-chatbot dialogs.To alleviate this issue we propose a novel RS model called MERCY that simulates human behavior in selecting the best candidate by taking into account distinct candidates concurrently and learns to rank them.In addition, MERCY leverages natural language feedback as another component to help the ranking task by explaining why each candidate response is relevant/irrelevant to the dialog context.These feedbacks are generated by prompting large language models in a few-shot setup.Our experiments show the better performance of MERCY over baselines for the response ranking task in our curated realistic datasets. Sarik Ghazarian, Behnam Hedayatnia, Di Jin 0005, Sijia Liu 0007, Nanyun Peng 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 6 |
| 2023 | Investigating the Representation of Open Domain Dialogue Context for Transformer ModelsabstractVishakh Padmakumar, Behnam Hedayatnia, Di Jin, Patrick Lange, Seokhwan Kim, Nanyun Peng, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Vishakh Padmakumar, Behnam Hedayatnia, Di Jin 0005, Patrick Lange, Seokhwan Kim, Nanyun Peng 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 7 |
| 2023 | "What do others think?": Task-Oriented Conversational Modeling with Subjective KnowledgeabstractChao Zhao, Spandana Gella, Seokhwan Kim, Di Jin, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Spandana Gella, Seokhwan Kim, Di Jin 0005, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 9 |
| 2022 | Think Before You Speak: Explicitly Generating Implicit Commonsense Knowledge for Response GenerationabstractPei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren 0001, Yang Liu 0004, Dilek Hakkani-Tür |
ACL (1) | 7 |
| 2022 | Inducer-tuning: Connecting Prefix-tuning and Adapter-tuningabstractPrefix-tuning, or more generally continuous prompt tuning, has become an essential paradigm of parameter-efficient transfer learning.Using a large pre-trained language model (PLM), prefix-tuning can obtain strong performance by training only a small portion of parameters.In this paper, we propose to understand and further develop prefix-tuning through the kernel lens.Specifically, we make an analogy between prefixes and inducing variables in kernel methods and hypothesize that prefixes serving as inducing variables would improve their overall mechanism.From the kernel estimator perspective, we suggest a new variant of prefix-tuning-inducer-tuning, which shares the exact mechanism as prefix-tuning while leveraging the residual form found in adaptertuning.This mitigates the initialization issue in prefix-tuning.Through comprehensive empirical experiments on natural language understanding and generation tasks, we demonstrate that inducer-tuning can close the performance gap between prefix-tuning and fine-tuning. Yifan Chen 0004, Devamanyu Hazarika, Mahdi Namazifar, Yang Liu 0004, Di Jin 0005, Dilek Hakkani-Tür |
EMNLP | 4 |
| 2022 | ExPUNations: Augmenting Puns with Keywords and ExplanationsabstractJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Tagyoung Chung, Jing Huang, Yang Liu, Nanyun Peng. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Tagyoung Chung, Jing Huang 0020, Yang Liu 0004, Nanyun Peng 0001 |
EMNLP | 7 |
| 2022 | Context-Situated Pun GenerationabstractJiao Sun, Anjali Narayan-Chen, Shereen Oraby, Shuyang Gao, Tagyoung Chung, Jing Huang, Yang Liu, Nanyun Peng. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jiao Sun, Anjali Narayan-Chen, Shereen Oraby, Shuyang Gao, Tagyoung Chung, Jing Huang 0020, Yang Liu 0004, Nanyun Peng 0001 |
EMNLP | 7 |
| 2022 | Enhancing Knowledge Selection for Grounded Dialogues via Document Semantic GraphsabstractSha Li, Mahdi Namazifar, Di Jin, Mohit Bansal, Heng Ji, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Mahdi Namazifar, Di Jin 0005, Mohit Bansal, Heng Ji 0001, Yang Liu 0004, Dilek Hakkani-Tür |
NAACL-HLT | 6 |
| 2022 | A Systematic Evaluation of Response Selection for Open Domain DialogueabstractRecent progress on neural approaches for language processing has triggered a resurgence of interest on building intelligent open-domain chatbots.However, even the state-of-the-art neural chatbots cannot produce satisfying responses for every turn in a dialog.A practical solution is to generate multiple response candidates for the same context, and then perform response ranking/selection to determine which candidate is the best.Previous work in response selection typically trains response rankers using synthetic data that is formed from existing dialogs by using a ground truth response as the single appropriate response and constructing inappropriate responses via random selection or using adversarial methods.In this work, we curated a dataset where responses from multiple response generators produced for the same dialog context are manually annotated as appropriate (positive) and inappropriate (negative).We argue that such training data better matches the actual use case examples, enabling the models to learn to rank responses effectively.With this new dataset, we conduct a systematic evaluation of state-of-the-art methods for response selection, and demonstrate that both strategies of using multiple positive candidates and using manually verified hard negative candidates can bring in significant performance improvement in comparison to using the adversarial training data, e.g., increase of 3% and 13% in Recall@1 score, respectively. Behnam Hedayatnia, Di Jin 0005, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 3 |
| 2022 | Improving Bot Response Contradiction Detection via Utterance RewritingabstractThough chatbots based on large neural models can often produce fluent responses in open domain conversations, one salient error type is contradiction or inconsistency with the preceding conversation turns.Previous work has treated contradiction detection in bot responses as a task similar to natural language inference, e.g., detect the contradiction between a pair of bot utterances.However, utterances in conversations may contain co-references or ellipsis, and using these utterances as is may not always be sufficient for identifying contradictions.This work aims to improve the contradiction detection via rewriting all bot utterances to restore antecedents and ellipsis.We curated a new dataset for utterance rewriting and built a rewriting model on it.We empirically demonstrate that this model can produce satisfactory rewrites to make bot utterances more complete.Furthermore, using rewritten utterances improves contradiction detection performance significantly, e.g., the AUPR and joint accuracy scores (detecting contradiction along with evidence) increase by 6.5% and 4.5% (absolute increase), respectively. Di Jin 0005, Sijia Liu 0007, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 3 |
| 2022 | Towards Textual Out-of-Domain Detection Without In-Domain LabelsabstractIn many real-world settings, machine learning models need to identify user inputs that are out-of-domain (OOD) so as to avoid performing wrong actions. This work focuses on a challenging case of OOD detection, where no labels for in-domain data are accessible (e.g., no intent labels for the intent classification task). To this end, we first evaluate different language model based approaches that predict likelihood for a sequence of tokens. Furthermore, we propose a novel representation learning based method by combining unsupervised clustering and contrastive learning so that better data representations for OOD detection can be learned. Through extensive experiments, we demonstrate that this method can significantly outperform likelihood-based methods and can be even competitive to the state-of-the-art supervised approaches with label information. Di Jin 0005, Shuyang Gao, Seokhwan Kim, Yang Liu 0004, Dilek Hakkani-Tür |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | "How Robust R U?": Evaluating Task-Oriented Dialogue Systems on Spoken ConversationsabstractMost prior work in dialogue modeling has been on written conversations mostly because of existing data sets. However, written dialogues are not sufficient to fully capture the nature of spoken conversations as well as the potential speech recognition errors in practical spoken dialogue systems. This work presents a new benchmark on spoken task-oriented conversations, which is intended to study multi-domain dialogue state tracking and knowledge-grounded dialogue modeling. We report that the existing state-of-the-art models trained on written conversations are not performing well on our spoken data, as expected. Furthermore, we observe improvements in task performances when leveraging$n$-best speech recognition hypotheses such as by combining predictions based on individual hypotheses. Our data set enables speech-based benchmarking of task-oriented dialogue systems. Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Dilek Hakkani-Tür |
ASRU | 2 |
| 2021 | Graph Enhanced Query Rewriting for Spoken Language Understanding SystemabstractQuery rewriting (QR) is an increasingly important component in voice assistant systems to reduce customer friction caused by errors in a spoken language understanding pipeline. These errors originate from various sources such as Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU) modules. In this work, we construct a user interaction graph from their queries using data mined from a Markov Chain Model [1], and introduce a self-supervised pre-training process for learning query embeddings by leveraging the recent developments in Graph Representation Learning (GRL). We then fine-tune these embeddings with weak supervised data for the query rewriting task, and observe improvement over the neural retrieval baseline system, demonstrating the effectiveness of the proposed method. Siyang Yuan, Saurabh Gupta 0008, Derek Liu, Yang Liu 0004, Chenlei Guo |
ICASSP | 5 |
| 2021 | Multi-Sentence Knowledge Selection in Open-Domain DialogueabstractMihail Eric, Nicole Chartier, Behnam Hedayatnia, Karthik Gopalakrishnan, Pankaj Rajan, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 14th International Conference on Natural Language Generation. 2021. Mihail Eric, Nicole Chartier, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Pankaj Rajan, Yang Liu 0004, Dilek Hakkani-Tür |
INLG | 6 |
| 2021 | Commonsense-Focused Dialogues for Response Generation: An Empirical StudyabstractPei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 7 |
| 2021 | Cross Interaction Network for Natural Language Guided Video Moment RetrievalabstractNatural language query grounding in videos is a challenging task that requires comprehensive understanding of the query, video and fusion of information across these modalities. Existing methods mostly emphasize on the query-to-video one-way interaction with a late fusion scheme, lacking effective ways to capture the relationship within and between query and video in a fine-grained manner. Moreover, current methods are often overly complicated resulting in long training time. We propose a self-attention together with cross interaction multi-head-attention mechanism in an early fusion scheme to capture video-query intra-dependencies as well as inter-relation from both directions (query-to-video and video-to-query). The cross-attention method can associate query words and video frames at any position and account for long-range dependencies in the video context. In addition, we propose a multi-task training objective that includes start/end prediction and moment segmentation. The moment segmentation task provides additional training signals that remedy the start/end prediction noise caused by annotator disagreement. Our simple yet effective architecture enables speedy training (within 1 hour on an AWS P3.2xlarge GPU instance) and instant inference. We showed that the proposed method achieves superior performance compared to complex state of the art methods, in particular surpassing the SOTA on high IoU metrics ([email protected], IoU=0.7) by 3.52% absolute (11.09% relative) on the Charades-STA dataset. Mohsen Malmir, Jiangning Chen, Yang Liu 0004 |
SIGIR | 8 |
| 2021 | Go Beyond Plain Fine-Tuning: Improving Pretrained Models for Social CommonsenseabstractPretrained language models have demonstrated outstanding performance in many NLP tasks recently. However, their social intelligence, which requires commonsense reasoning about the current situation and mental states of others, is still developing. Towards improving language models' social intelligence, in this study we focus on the Social IQA dataset, a task requiring social and emotional commonsense reasoning. Building on top of the pretrained RoBERTa and GPT2 models, we propose several architecture variations and extensions, as well as leveraging external commonsense corpora, to optimize the model for Social IQA. Our proposed system achieves competitive results as those top-ranking models on the leaderboard. This work demonstrates the strengths of pretrained language models, and provides viable ways to improve their performance for a particular task. Ting-Yun Chang, Yang Liu 0004, Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Dilek Hakkani-Tür |
SLT | 2 |
| 2020 | Policy-Driven Neural Response Generation for Knowledge-Grounded Dialog SystemsabstractOpen-domain dialog systems aim to generate relevant, informative and engaging responses.In this paper, we propose using a dialog policy to plan the content and style of target, opendomain responses in the form of an action plan, which includes knowledge sentences related to the dialog context, targeted dialog acts, topic information, etc.For training, the attributes within the action plan are obtained by automatically annotating the publicly released Topical-Chat dataset.We condition neural response generators on the action plan which is then realized as target utterances at the turn and sentence levels.We also investigate different dialog policy models to predict an action plan given the dialog context.Through automated and human evaluation, we measure the appropriateness of the generated responses and check if the generation models indeed learn to realize the given action plans.We demonstrate that a basic dialog policy that operates at the sentence level generates better responses in comparison to turn level generation as well as baseline models with no action plan.Additionally the basic dialog policy has the added benefit of controllability. Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Seokhwan Kim, Yang Liu 0004, Mihail Eric, Dilek Hakkani-Tür |
INLG | 4 |
| 2020 | Are Neural Open-Domain Dialog Systems Robust to Speech Recognition Errors in the Dialog History? An Empirical StudyabstractLarge end-to-end neural open-domain chatbots are becoming increasingly popular. However, research on building such chatbots has typically assumed that the user input is written in nature and it is not clear whether these chatbots would seamlessly integrate with automatic speech recognition (ASR) models to serve the speech modality. We aim to bring attention to this important question by empirically studying the effects of various types of synthetic and actual ASR hypotheses in the dialog history on TransferTransfo, a state-of-the-art Generative Pre-trained Transformer (GPT) based neural open-domain dialog system from the NeurIPS ConvAI2 challenge. We observe that TransferTransfo trained on written data is very sensitive to such hypotheses introduced to the dialog history during inference time. As a baseline mitigation strategy, we introduce synthetic ASR hypotheses to the dialog history during training and observe marginal improvements, demonstrating the need for further research into techniques to make end-to-end open-domain chatbots fully speech-robust. To the best of our knowledge, this is the first study to evaluate the effects of synthetic and actual ASR hypotheses on a state-of-the-art neural open-domain dialog system and we hope it promotes speech-robustness as an evaluation criterion in open-domain dialog. Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Longshaokan Wang, Yang Liu 0004, Dilek Hakkani-Tür |
INTERSPEECH | 4 |
| 2020 | ASR Error Correction with Augmented Transformer for Entity Retrieval
Haoyu Wang 0002, Shuyan Dong, James Logan, Ashish Kumar Agrawal, Yang Liu 0004 |
INTERSPEECH | 6 |
| 2020 | Beyond Domain APIs: Task-oriented Conversational Modeling with Unstructured Knowledge AccessabstractMost prior work on task-oriented dialogue systems are restricted to a limited coverage of domain APIs, while users oftentimes have domain related requests that are not covered by the APIs.In this paper, we propose to expand coverage of task-oriented dialogue systems by incorporating external unstructured knowledge sources.We define three sub-tasks: knowledge-seeking turn detection, knowledge selection, and knowledge-grounded response generation, which can be modeled individually or jointly.We introduce an augmented version of MultiWOZ 2.1, which includes new out-of-API-coverage turns and responses grounded on external knowledge sources.We present baselines for each sub-task using both conventional and neural approaches.Our experimental results demonstrate the need for further research in this direction to enable more informative conversational systems. Seokhwan Kim, Mihail Eric, Karthik Gopalakrishnan 0001, Behnam Hedayatnia, Yang Liu 0004, Dilek Hakkani-Tür |
SIGdial | 5 |
| 2019 | A Multi-Agent Communication Framework for Question-Worthy Phrase Extraction and Question GenerationabstractQuestion generation aims to produce questions automatically given a piece of text as input. Existing research follows a sequence-to-sequence fashion that constructs a single question based on the input. Considering each question usually focuses on a specific fragment of the input, especially in the scenario of reading comprehension, it is reasonable to identify the corresponding focus before constructing the question. In this paper, we propose to identify question-worthy phrases first and generate questions with the assistance of these phrases. We introduce a multi-agent communication framework, taking phrase extraction and question generation as two agents, and learn these two tasks simultaneously via message passing mechanism. The results of experiments show the effectiveness of our framework: we can extract question-worthy phrases, which are able to improve the performance of question generation. Besides, our system is able to extract more than one question worthy phrases and generate multiple questions accordingly. Siyuan Wang 0025, Zhongyu Wei, Zhihao Fan, Yang Liu 0004, Xuanjing Huang 0001 |
AAAI | 4 |
| 2019 | On the Role of Style in Parsing Speech with Neural ModelsabstractThe differences in written text and conversational speech are substantial; previous parsers trained on treebanked text have given very poor results on spontaneous speech. For spoken language, the mismatch in style also extends to prosodic cues, though it is less well understood. This paper re-examines the use of written text in parsing speech in the context of recent advances in neural language processing. We show that neural approaches facilitate using written text to improve parsing of spontaneous speech, and that prosody further improves over this state-of-the-art result. Further, we find an asymmetric degradation from read vs. spontaneous mismatch, with spontaneous speech more generally useful for training parsers. Trang Tran 0001, Jiahong Yuan, Yang Liu 0004, Mari Ostendorf |
INTERSPEECH | 3 |
| 2018 | A Reinforcement Learning Framework for Natural Question Generation using Bi-discriminatorsabstractVisual Question Generation (VQG) aims to ask natural questions about an image automatically. Existing research focus on training model to fit the annotated data set that makes it indifferent from other language generation tasks. We argue that natural questions need to have two specific attributes from the perspectives of content and linguistic respectively, namely, natural and human-written. Inspired by the setting of discriminator in adversarial learning, we propose two discriminators, one for each attribute, to enhance the training. We then use the reinforcement learning framework to incorporate scores from the two discriminators as the reward to guide the training of the question generator. Experimental results on a benchmark VQG dataset show the effectiveness and robustness of our model compared to some state-of-the-art models in terms of both automatic and human evaluation metrics. Zhihao Fan, Zhongyu Wei, Siyuan Wang 0025, Yang Liu 0004, Xuanjing Huang 0001 |
COLING | 4 |
| 2018 | Incorporating Argument-Level Interactions for Persuasion Comments Evaluation using Co-attention ModelabstractIn this paper, we investigate the issue of persuasiveness evaluation for argumentative comments. Most of the existing research explores different text features of reply comments on word level and ignores interactions between participants. In general, viewpoints are usually expressed by multiple arguments and exchanged on argument level. To better model the process of dialogical argumentation, we propose a novel co-attention mechanism based neural network to capture the interactions between participants on argument level. Experimental results on a publicly available dataset show that the proposed model significantly outperforms some state-of-the-art methods for persuasiveness evaluation. Further analysis reveals that attention weights computed in our model are able to extract interactive argument pairs from the original post and the reply. Lu Ji, Zhongyu Wei, Xiangkun Hu, Yang Liu 0004, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 4 |
| 2018 | Liulishuo's System for the Spoken CALL Shared Task 2018
Lei Chen 0004, Ramon Prieto, Yang Liu 0004 |
INTERSPEECH | 5 |
| 2017 | Meaningful head movements driven by emotional synthetic speech
Najmeh Sadoughi, Yang Liu 0004, Carlos Busso |
Speech Commun. | 2 |
| 2017 | A Multi-Task Learning Framework for Emotion Recognition Using 2D Continuous SpaceabstractDimensional models have been proposed in psychology studies to represent complex human emotional expressions. Activation and valence are two common dimensions in such models. They can be used to describe certain emotions. For example, anger is one type of emotion with a low valence and high activation value; neutral has both a medium level valence and activation value. In this work, we propose to apply multi-task learning to leverage activation and valence information for acoustic emotion recognition based on the deep belief network (DBN) framework. We treat the categorical emotion recognition task as the major task. For the secondary task, we leverage activation and valence labels in two different ways, category level based classification and continuous level based regression. The combination of the loss functions from the major and secondary tasks is used as the objective function in the multi-task learning framework. After iterative optimization, the values from the last hidden layer in the DBN are used as new features and fed into a support vector machine classifier for emotion recognition. Our experimental results on the Interactive Emotional Dyadic Motion Capture and Sustained Emotionally Colored Machine-Human Interaction Using Nonverbal Expression databases show significant improvements on unweighted accuracy, illustrating the benefit of utilizing additional information in a multi-task learning setup for emotion recognition. Yang Liu 0004 |
IEEE Trans. Affect. Comput. | 2 |
| 2016 | Using Relevant Public Posts to Enhance News Article SummarizationabstractA news article summary usually consists of 2-3 key sentences that reflect the gist of that news article. In this paper we explore using public posts following a new article to improve automatic summary generation for the news article. We propose different approaches to incorporate information from public posts, including using frequency information from the posts to re-estimate bigram weights in the ILP-based summarization model and to re-weight a dependency tree edge’s importance for sentence compression, directly selecting sentences from posts as the final summary, and finally a strategy to combine the summarization results generated from news articles and posts. Our experiments on data collected from Facebook show that relevant public posts provide useful information and can be effectively leveraged to improve news article summarization results. Chen Li 0013, Zhongyu Wei, Yang Liu 0004 |
COLING | 3 |
| 2016 | Automatic composition of broadcast news summaries using rank classifiers trained with acoustic and lexical featuresabstractResearch on automatic speech summarization typically focuses on optimizing objective evaluation criteria, such as the ROUGE metric, which depend on word and phrase overlaps between automatic and manually generated summary documents. However, the actual quality of the speech summarizer largely depends on how the end-users perceive the audio output. This work focuses on the task of composing summarized audio streams with the aim of improving the quality and interest perceived by the end-user. First, using crowd-sourced summary annotations on a broadcast news corpus, we train a rank-SVM classifier to learn the relative importance of each sentence in a news story. Acoustic, lexical and structural features are used for training. In addition, we investigate the perceived emotion level in each sentence to aid the summarizer in selecting interesting sentences, yielding an emotion-aware summarizer. Next, we propose several methods to combine these sentences to generate a compressed audio stream. Subjective evaluations are performed to evaluate the quality of the generated summaries on the following criterion: interest, abruptness, informativeness, attractiveness, and overall quality. The results indicate that users are most sensitive to the linguistic coherence and continuity of the audio stream. Taufiq Hasan, Mohammed Abdel-Wahab 0001, Srinivas Parthasarathy, Carlos Busso, Yang Liu 0004 |
ICASSP | 5 |
| 2016 | Extractive summarization of multi-party meetings through discourse segmentationabstractAbstract In this article we tackle the problem of multi-party conversation summarization. We investigate the role of discourse segmentation of a conversation on meeting summarization. First, an unsupervised function segmentation algorithm is proposed to segment the transcript into functionally coherent parts, such asMonologuei(which indicates a segment where speakeriis the dominant speaker, e.g., lecturing all the other participants) orDiscussionx1x2, . . .,xn(which indicates a segment where speakersx1toxninvolve in a discussion). Then the salience score for a sentence is computed by leveraging the score of the segment containing the sentence. Performance of our proposed segmentation and summarization algorithms is evaluated using the AMI meeting corpus. We show better summarization performance over other state-of-the-art algorithms according to different metrics. Mohammad Hadi Bokaei, Hossein Sameti, Yang Liu 0004 |
Nat. Lang. Eng. | 3 |
| 2016 | Summarizing Meeting Transcripts Based on Functional SegmentationabstractIn this paper, we aim to improve meeting summarization performance using discourse specific information. Since there are intrinsically different characteristics in utterances in different types of function segments, e.g., Monologue segments versus Discussion ones, we propose a new summarization framework where different summarizers are used for different segment types. For monologue segments, we adopt the integer linear programming-based summarization method; whereas for discussion segments, we use a graph-based method to incorporate speaker information. Performance of our proposed method is evaluated using the standard AMI meeting corpus. Results show a good improvement over previous state-of-the-art algorithms according to various evaluation metrics and different compress ratios. Mohammad Hadi Bokaei, Hossein Sameti, Yang Liu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Improving Named Entity Recognition in Tweets via Detecting Non-Standard WordsabstractChen Li, Yang Liu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Chen Li 0013, Yang Liu 0004 |
ACL (1) | 2 |
| 2015 | Feature Selection in Kernel Space: A Case Study on Dependency ParsingabstractXian Qian, Yang Liu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Xian Qian, Yang Liu 0004 |
ACL (1) | 2 |
| 2015 | Leveraging valence and activation information via multi-task learning for categorical emotion recognitionabstractDeep learning technologies have been successfully applied to acoustic emotion recognition lately. In this work, we propose to apply multi-task learning for acoustic emotion recognition based on the Deep Belief Network (DBN) framework. We treat the categorical emotion recognition task as the major task. For the secondary task, we leverage two continuous labels, valence and activation. Two strategies are employed to achieve multi-task learning. First, we map the continuous labels into three categorical labels: low; medium; high, and use classification for the secondary task. Second, we project the continuous labels into [-1; 1] range, and use regression for the secondary task. The combination of the loss functions from the major and secondary tasks is used in the objective function in multi-task learning. After iterative optimization, the values from the last hidden layer are used as features in the backend SVM classifier for emotion classification. Our experimental results show significant improvement over the baseline results using DBN, suggesting the benefit of utilizing additional information in a multi-task learning setup. Yang Liu 0004 |
ICASSP | 2 |
| 2015 | Joint POS Tagging and Text Normalization for Informal Text
Chen Li 0013, Yang Liu 0004 |
IJCAI | 2 |
| 2015 | Extractive meeting summarization through speaker zone detection
Mohammad Hadi Bokaei, Hossein Sameti, Yang Liu 0004 |
INTERSPEECH | 3 |
| 2015 | Using External Resources and Joint Learning for Bigram Weighting in ILP-Based Multi-Document SummarizationabstractSome state-of-the-art summarization systems use integer linear programming (ILP) based methods that aim to maximize the important concepts covered in the summary.These concepts are often obtained by selecting bigrams from the documents.In this paper, we improve such bigram based ILP summarization methods from different aspects.First we use syntactic information to select more important bigrams.Second, to estimate the importance of the bigrams, in addition to the internal features based on the test documents (e.g., document frequency, bigram positions), we propose to extract features by leveraging multiple external resources (such as word embedding from additional corpus, Wikipedia, Dbpedia, Word-Net, SentiWordNet).The bigram weights are then trained discriminatively in a joint learning model that predicts the bigram weights and selects the summary sentences in the ILP framework at the same time.We demonstrate that our system consistently outperforms the prior ILP method on different TAC data sets, and performs competitively compared to other previously reported best results.We also conducted various analyses to show the contributions of different components. Chen Li 0013, Yang Liu 0004 |
HLT-NAACL | 2 |
| 2015 | Improving Update Summarization via Supervised ILP and Sentence RerankingabstractInteger Linear Programming (ILP) based summarization methods have been widely adopted recently because of their state-of-the-art performance. This paper proposes two new modifications in this framework for update summarization. Our key idea is to use discriminative models with a set of features to measure both the salience and the novelty of words and sentences. First, these features are used in a supervised model to predict the weights of the concepts used in the ILP model. Second, we generate preliminary sentence candidates in the ILP model and then rerank them using sentence level features. We evaluate our method on different TAC update summarization data sets, and the results show that our system performs competitively compared to the best TAC systems based on the ROUGE evaluation metric. Chen Li 0013, Yang Liu 0004 |
HLT-NAACL | 2 |
| 2015 | Opinion summarization on spontaneous conversations
Yang Liu 0004 |
Comput. Speech Lang. | 2 |
| 2015 | Linear Discourse Segmentation of Multi-Party Meetings Based on Local and Global InformationabstractLinear segmentation of a meeting conversation is beneficial as a stand-alone system (to organize a meeting and make it easier to access) or as a preprocessing step for many other meeting related tasks. Such segmentation can be done according to two different criteria: topic in which a meeting is segmented according to the different items in its agenda, and function in which the segmentation is done according to the meeting's different events (like discussion, monologue). In this article we concentrate on the function segmentation task and propose new unsupervised methods to segment a meeting into functionally coherent parts. The first proposed method assigns a score to each possible boundary according to its local information and then selects the best ones. The second method uses a dynamic programming approach to find the global best segmentation according to a defined cost function. Since these two methods are complementary of each other, we propose the third method as a combination of the first two ones, which takes advantage of both to improve the final segmentation. In order to evaluate our proposed methods, a subset of a standard meeting dataset (AMI) is manually annotated and used as the test set. Results show that our proposed methods perform significantly better than the previous unsupervised approach according to different evaluation metrics. Mohammad Hadi Bokaei, Hossein Sameti, Yang Liu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Improving Multi-documents Summarization by Sentence Compression based on Expanded Constituent Parse TreesabstractIn this paper, we focus on the problem of using sentence compression techniques to improve multi-document summarization.We propose an innovative sentence compression method by considering every node in the constituent parse tree and deciding its status -remove or retain.Integer liner programming with discriminative training is used to solve the problem.Under this model, we incorporate various constraints to improve the linguistic quality of the compressed sentences.Then we utilize a pipeline summarization framework where sentences are first compressed by our proposed compression model to obtain top-n candidates and then a sentence selection module is used to generate the final summary.Compared with state-ofthe-art algorithms, our model has similar ROUGE-2 scores but better linguistic quality on TAC data. Chen Li 0013, Yang Liu 0004, Fei Liu 0004, Fuliang Weng |
EMNLP | 2 |
| 2014 | Introducing shared-hidden-layer autoencoders for transfer learning and their application in acoustic emotion recognitionabstractThis study addresses a situation in practice where training and test samples come from different corpora - here in acoustic emotion recognition. In this situation, a model is trained on one database while tested on another disjoint one. The typical inherent mismatch between the corpora and by that between test and training set usually leads to significant performance degradation. To cope with this problem when no training data from the target domain exists, we propose a `shared-hidden-layer autoencoder' (SHLA) approach for learning common feature representations shared across the training and test set in order to reduce the discrepancy in them. To exemplify effectiveness of our approach, we select the Interspeech Emotion Challenge's FAU Aibo Emotion Corpus as test database and two other publicly available databases as training set for extensive evaluation. The experimental results show that our SHLA method significantly improves over the baseline performance and outperforms today's state-of-the-art domain adaptation methods. Zixing Zhang 0001, Yang Liu 0004, Björn W. Schuller |
ICASSP | 4 |
| 2014 | Modeling gender information for emotion recognition using Denoising autoencoderabstractThe Denoising autoencoder (DAE) has been successfully applied to acoustic emotion recognition lately. In this paper, we adopt the framework of the modified DAE introduced in that projects the input signal to two different hidden representations, for neutral and emotional speech respectively, and uses the emotional representation for the classification task. We propose to model gender information for more robust emotional representation in this work. For neutral representation, male and female dependent DAEs are built using non-emotional speech with the aim of capturing distinct information between the two genders. The emotional hidden representation is shared for the two genders in order to model more emotion specific characteristics, and is used as features in a back-end classifier for emotion recognition. We propose different optimization objectives in training the DAEs. Our experimental results show improvement on unweighted accuracy compared with previous work using the modified DAE method and the classifiers using the standard static features. Further performance gain can be achieved by structural level system combination. Björn W. Schuller, Yang Liu 0004 |
ICASSP | 4 |
| 2014 | Speech-Driven Animation Constrained by Appropriate Discourse FunctionsabstractConversational agents provide powerful opportunities to interact and engage with the users. The challenge is how to create naturalistic behaviors that replicate the complex gestures observed during human interactions. Previous studies have used rule-based frameworks or data-driven models to generate appropriate gestures, which are properly synchronized with the underlying discourse functions. Among these methods, speech-driven approaches are especially appealing given the rich information conveyed on speech. It captures emotional cues and prosodic patterns that are important to synthesize behaviors (i.e., modeling the variability and complexity of the timings of the behaviors). The main limitation of these models is that they fail to capture the underlying semantic and discourse functions of the message (e.g., nodding). This study proposes a speech-driven framework that explicitly model discourse functions, bridging the gap between speech-driven and rule-based models. The approach is based on dynamic Bayesian Network (DBN), where an additional node is introduced to constrain the models by specific discourse functions. We implement the approach by synthesizing head and eyebrow motion. We conduct perceptual evaluations to compare the animations generated using the constrained and unconstrained models. Najmeh Sadoughi, Yang Liu 0004, Carlos Busso |
ICMI | 2 |
| 2014 | Level of interest sensing in spoken dialog using decision-level fusion of acoustic and lexical evidence
Je Hun Jeon, Yang Liu 0004 |
Comput. Speech Lang. | 3 |
| 2014 | Normalization of informal text
Deana Pennell, Yang Liu 0004 |
Comput. Speech Lang. | 2 |
| 2014 | 2-Slave Dual Decomposition for Generalized Higher Order CRFsabstractWe show that the decoding problem in generalized Higher Order Conditional Random Fields (CRFs) can be decomposed into two parts: one is a tree labeling problem that can be solved in linear time using dynamic programming; the other is a supermodular quadratic pseudo-Boolean maximization problem, which can be solved in cubic time using a minimum cut algorithm. We use dual decomposition to force their agreement. Experimental results on Twitter named entity recognition and sentence dependency tagging tasks show that our method outperforms spanning tree based dual decomposition. Xian Qian, Yang Liu 0004 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Using Supervised Bigram-based ILP for Extractive Summarization
Chen Li 0013, Xian Qian, Yang Liu 0004 |
ACL (1) | 3 |
| 2013 | Document Summarization via Guided Sentence CompressionabstractJoint compression and summarization has been used recently to generate high quality summaries.However, such word-based joint optimization is computationally expensive.In this paper we adopt the 'sentence compression + sentence selection' pipeline approach for compressive summarization, but propose to perform summary guided compression, rather than generic sentence-based compression.To create an annotated corpus, the human annotators were asked to compress sentences while explicitly given the important summary words in the sentences.Using this corpus, we train a supervised sentence compression model using a set of word-, syntax-, and documentlevel features.During summarization, we use multiple compressed sentences in the integer linear programming framework to select salient summary sentences.Our results on the TAC 2008 and 2011 summarization data sets show that by incorporating the guided sentence compression model, our summarization system can yield significant performance gain as compared to the state-of-the-art. Chen Li 0013, Fei Liu 0004, Fuliang Weng, Yang Liu 0004 |
EMNLP | 4 |
| 2013 | Fast Joint Compression and Summarization via Graph CutsabstractExtractive summarization typically uses sentences as summarization units.In contrast, joint compression and summarization can use smaller units such as words and phrases, resulting in summaries containing more information.The goal of compressive summarization is to find a subset of words that maximize the total score of concepts and cutting dependency arcs under the grammar constraints and summary length constraint.We propose an efficient decoding algorithm for fast compressive summarization using graph cuts.Our approach first relaxes the length constraint using Lagrangian relaxation.Then we propose to bound the relaxed objective function by the supermodular binary quadratic programming problem, which can be solved efficiently using graph max-flow/min-cut.Since finding the tightest lower bound suffers from local optimality, we use convex relaxation for initialization.Experimental results on TAC2008 dataset demonstrate our method achieves competitive ROUGE score and has good readability, while is much faster than the integer linear programming (ILP) method. Xian Qian, Yang Liu 0004 |
EMNLP | 2 |
| 2013 | Automatic human utility evaluation of ASR systems: does WER really predict performance?abstractInternational audience Benoît Favre, Kyla Cheung, Siavash Kazemian, Adam Lee, Yang Liu 0004, Cosmin Munteanu, Ani Nenkova, Dennis Ochei, Gerald Penn, Stephen Tratz, Clare R. Voss, Frauke Zeller |
INTERSPEECH | 5 |
| 2013 | A preliminary study of cross-lingual emotion recognition from speech: automatic classification versus human perceptionabstractThe aim of this study is to investigate the effect of cross-lingual data on human perception and automatic classification of emotion from speech. We use four different databases from three languages (English, Chinese, and German) and two types (acted and improvised). For automatic classification, there is a significant degradation using cross-corpus than within-corpus setup. For human perception, we observe differences between native and non-native speakers when judging emotions for a language, and there is less performance loss in cross-language setup compared to automatic classification. In addition, we find that the automatic approaches work well in classifying the emotional activation category: positive and negative activated emotions, but are not good at classifying instances within the same activation category, which is different from the confusion patterns of the human perception experiment. This study provides insights to better understanding of cross-lingual human emotion perception and development of robust automatic emotion recognition systems. Je Hun Jeon, Yang Liu 0004 |
INTERSPEECH | 4 |
| 2013 | Using denoising autoencoder for emotion recognition
Yang Liu 0004 |
INTERSPEECH | 2 |
| 2013 | Disfluency Detection Using Multi-step Stacked Learning
Xian Qian, Yang Liu 0004 |
HLT-NAACL | 2 |
| 2013 | Branch and Bound Algorithm for Dependency Parsing with Non-local FeaturesabstractGraph based dependency parsing is inefficient when handling non-local features due to high computational complexity of inference. In this paper, we proposed an exact and efficient decoding algorithm based on the Branch and Bound (B&B) framework where non-local features are bounded by a linear combination of local features. Dynamic programming is used to search the upper bound. Experiments are conducted on English PTB and Chinese CTB datasets. We achieved competitive Unlabeled Attachment Score (UAS) when no additional resources are available: 93.17% for English and 87.25% for Chinese. Parsing speed is 177 words per second for English and 97 words per second for Chinese. Our algorithm is general and can be adapted to non-projective dependency parsing or other graphical models. Xian Qian, Yang Liu 0004 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | Towards Abstractive Speech Summarization: Exploring Unsupervised and Supervised Approaches for Spoken Utterance CompressionabstractMost previous studies on speech summarization focus on the extractive approaches. Yet directly concatenating the extracted speech utterances may not form a good summary due to the presence of disfluencies and redundancy in the unplanned spontaneous speech. In this paper, we proposed to generate compressed speech summaries by coupling the sentence level compression and summarization approaches, as a viable step towards generating abstractive summaries. We compared two utterance compression approaches: an unsupervised approach based on the Integer Linear Programming (ILP) framework, and a supervised method using conditional random fileds (CRF) that formulates the utterance compression problem as a sequence labeling task. We evaluated the compression performance using both human and ASR transcripts from the ICSI meeting corpus, and performed both automatic and human evaluation. Our results show that we can achieve reasonable utterance compression performance, and that the CRF-based method generally performs better. By coupling the compression and summarization approaches, we generated compressed speech summaries that cover more important information within the given length limit, yielding 5% absolute performance gain on both human and ASR transcripts as evaluated by the ROUGE-1 F-scores. Fei Liu 0004, Yang Liu 0004 |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Sequence Labeling with Non-Negative Weighted Higher Order FeaturesabstractIn sequence labeling, using higher order features leads to high inference complexity. A lot of studies have been conducted to address this problem. In this paper, we propose a new exact decoding algorithm under the assumption that weights of all higher order features are non-negative. In the worst case, the time complexity of our algorithm is quadratic on the number of higher order features. Comparing with existing algorithms, our method is more efficient and easier to implement. We evaluate our method on two sequence labeling tasks: Optical Character Recognition and Chinese part-of-speech tagging. Our experimental results demonstrate that adding higher order features significantly improves the performance while requiring only 30% additional inference time. Xian Qian, Yang Liu 0004 |
AAAI | 2 |
| 2012 | Sentence Dependency Tagging in Online Question Answering Forums
Zhonghua Qu, Yang Liu 0004 |
ACL (1) | 2 |
| 2012 | Improving Text Normalization using Character-Blocks Based Models and System Combination
Chen Li 0013, Yang Liu 0004 |
COLING | 2 |
| 2012 | User Participation Prediction in Online Forums
Zhonghua Qu, Yang Liu 0004 |
EACL | 2 |
| 2012 | Joint Chinese Word Segmentation, POS Tagging and Parsing
Xian Qian, Yang Liu 0004 |
EMNLP-CoNLL | 2 |
| 2012 | Evaluating NLP Features for Automatic Prediction of Language Impairment Using Child Speech TranscriptsabstractLanguage impairment (LI) in children is pervasive in all walks of life. Automatic prediction of LI is useful as a first pass for speech language pathologists in identifying prospective children with LI. Previous work in the automatic prediction of LI has explored various features, mostly shallow and surface level features. In this paper, we evaluate deeper Natural Language Processing (NLP) features such as syntactic, semantic and entity grid model features, along with narrative structure and quality features in the prediction of LI using child language transcripts. Our experiments show that narrative structure and quality features along with a combination of other features are helpful in the prediction of LI in storytelling narratives. Khairun-nisa Hassanali, Yang Liu 0004, Thamar Solorio |
INTERSPEECH | 2 |
| 2012 | Normalization of Text Messages Using Character- and Phone-based Machine Translation ApproachesabstractThere are many abbreviation and non-standard words in SMS and Twitter messages. They are problematic for text-to-speech (TTS) or language processing techniques for these data. A character-based machine translation (MT) approach was previously used for normalization of non-standard words. In this paper, we propose a two-stage translation method to leverage phonetic information, where non-standard words are first translated to possible pronunciations, which are then translated to standard words. We further combine it with the single-step character-based translation module. Our experiments show that our proposed method significantly outperforms previous results in both n-best coverage and 1-best accuracy. Chen Li 0013, Yang Liu 0004 |
INTERSPEECH | 2 |
| 2012 | Using i-Vector Space Model for Emotion Recognition
Yang Liu 0004 |
INTERSPEECH | 2 |
| 2012 | Evaluating the effect of normalizing informal text on TTS outputabstractAbbreviations in informal text, and research efforts to expand them to the standard English words from which they were derived, have become increasingly common. These methods are almost solely evaluated using the final word error rate (WER) after normalization; however, this metric may not be reasonable for a text-to-speech (TTS) system where words may be pronounced correctly despite being misspelled. This paper shows that normalization of informal text improves the output of TTS not only in terms of WER but also in terms of phoneme error rate (PER) and human perceptual experiments. Deana Pennell, Yang Liu 0004 |
SLT | 2 |
| 2012 | A cross-corpus study of subjectivity identification using unsupervised learningabstractAbstract In this study, we investigate using unsupervised generative learning methods for subjectivity detection across different domains. We create an initial training set using simple lexicon information and then evaluate two iterative learning methods with a base naive Bayes classifier to learn from unannotated data. The first method is self-training, which adds instances with high confidence into the training set in each iteration. The second is a calibrated EM (expectation-maximization) method where we calibrate the posterior probabilities from EM such that the class distribution is similar to that in the real data. We evaluate both approaches on three different domains: movie data, news resource, and meeting dialogues, and we found that in some cases the unsupervised learning methods can achieve performance close to the fully supervised setup. We perform a thorough analysis to examine factors, such as self-labeling accuracy of the initial training set in unsupervised learning, the accuracy of the added examples in self-training, and the size of the initial training set in different methods. Our experiments and analysis show inherent differences across domains and impacting factors explaining the model behaviors. Yang Liu 0004 |
Nat. Lang. Eng. | 2 |
| 2012 | Automatic prosodic event detection using a novel labeling and selection method in co-training
Je Hun Jeon, Yang Liu 0004 |
Speech Commun. | 2 |
| 2011 | N-Best Rescoring Based on Pitch-accent Patterns
Je Hun Jeon, Wen Wang 0001, Yang Liu 0004 |
ACL | 3 |
| 2011 | A Pilot Study of Opinion Summarization in Conversations
Yang Liu 0004 |
ACL | 2 |
| 2011 | Sentence level emotion recognition based on decisions from subsentence segmentsabstractEmotion recognition from speech plays an important role in developing affective and intelligent systems. This study investigates sentence-level emotion recognition. We propose to use a two-step approach to leverage information from sub sentence segments for sentence level decision. First we use a segment level emotion classifier to generate predictions for segments within a sentence. A second component combines the predictions from these segments to obtain a sentence level decision. We evaluate different segment units (words, phrases, time-based segments) and different decision combination methods (majority vote, average of probabilities, and a Gaussian Mixture Model (GMM)). Our experimental results on two different data sets show that our proposed method significantly outperforms the standard sentence-based classification approach. In addition, we find that using time-based segments achieves the best performance, and thus no speech recognition or alignment is needed when using our method, which is important to develop language independent emotion recognition systems. Je Hun Jeon, Yang Liu 0004 |
ICASSP | 3 |
| 2011 | Toward text message normalization: Modeling abbreviation generationabstractThis paper describes a text normalization system for deletion-based abbreviations in informal text. We propose using statistical classifiers to learn the probability of deleting a given character using features based on character context, position in the word and containing syllable, and function within the word. To ensure that our system is robust to different and previously unseen abbreviations for a word, we generate multiple abbreviation hypotheses for a word using the predictions from the classifiers. We then reverse the mappings to enable recovery of English words from the abbreviations. Different knowledge sources are used to disambiguate word candidates: abbreviation likelihood, length, and language model scores. Our results show that this approach is feasible and warrants further exploration in the future. Deana Pennell, Yang Liu 0004 |
ICASSP | 2 |
| 2011 | Learning from Chinese-English Parallel Data for Chinese Tense Prediction
Fei Liu 0004, Yang Liu 0004 |
IJCNLP | 3 |
| 2011 | A Character-Level Machine Translation Approach for Normalization of SMS Abbreviations
Deana Pennell, Yang Liu 0004 |
IJCNLP | 2 |
| 2011 | Finding Problem Solving Threads in Online Forum
Zhonghua Qu, Yang Liu 0004 |
IJCNLP | 2 |
| 2011 | Exploring a corpus-based approach for detecting language impairment in monolingual English-speaking children
Keyur Gabani, Thamar Solorio, Yang Liu 0004, Khairun-nisa Hassanali, Christine A. Dollaghan |
Artif. Intell. Medicine | 3 |
| 2011 | Analyzing language samples of Spanish-English bilingual children for the automated prediction of language dominanceabstractAbstract In this work we study how features typically used in natural language processing tasks, together with measures from syntactic complexity, can be adapted to the problem of developing language profiles of bilingual children. Our experiments show that these features can provide high discriminative value for predicting language dominance from story retells in a Spanish–English bilingual population of children. Moreover, some of our proposed features are even more powerful than measures commonly used by clinical researchers and practitioners for analyzing spontaneous language samples of children. This study shows that the field of natural language processing has the potential to make significant contributions to communication disorders and related areas. Thamar Solorio, Melissa Sherman, Yang Liu 0004, Lisa Bedore, Elizabeth Peña, Aquiles Iglesias |
Nat. Lang. Eng. | 3 |
| 2011 | A Supervised Framework for Keyword Extraction From Meeting TranscriptsabstractThis paper presents a supervised framework for extracting keywords from meeting transcripts, a genre that is significantly different from written text or other speech domains such as broadcast news. In addition to the traditional frequency- or position-based clues, we investigate a variety of novel features, including linguistically motivated term specificity features, decision-making sentence-related features, prosodic prominence scores, as well as a group of features derived from summary sentences. To generate better system summaries, we propose a feedback loop mechanism under a supervised framework to leverage the relationship between keywords and summary sentences. Experiments are performed on the ICSI meeting corpus using both human transcripts and automatic speech recognition (ASR) outputs. Results have shown that our proposed supervised framework is able to outperform both unsupervised term frequency inverse document frequency (TF-IDF) weighting and a supervised keyphrase extraction system which is known for its satisfying performance on written text. We conduct extensive analysis to demonstrate the effectiveness of the newly proposed features and the feedback mechanism used to generate summaries. Furthermore, we show promising results using n-best recognition output to address the problems of recognition errors. Fei Liu 0004, Yang Liu 0004 |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Using N-Best Lists and Confusion Networks for Meeting SummarizationabstractThe incorrect speech recognition results usually have a negative impact on the speech summarization task, especially on the meeting domain where the word error rate is often higher than other speech genres. In this paper we investigate using rich speech recognition results to improve meeting summarization performance. Two kinds of structures are considered, n-best hypotheses and confusion networks. We develop methods to utilize multiple word and sentence candidates and their recognition confidence for summarization under an unsupervised framework. Our experimental results on the ICSI meeting corpus show that our proposed method can significantly improve summarization performance over using 1-best recognition output, evaluated by both ROUGE-1 and ROUGE-2 scores. We also find that if the task is to generate speech summaries or identify salient segments, using rich speech recognition output is just as effective as using human transcripts. In addition, we discuss the difference between n-best lists and confusion networks, and analyze the word error rate in the exacted summary sentences. Shasha Xie, Yang Liu 0004 |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Using n-best recognition output for extractive summarization and keyword extraction in meeting speechabstractThere has been increasing interest recently in meeting understanding, such as summarization, browsing, action item detection, and topic segmentation. However, there is very limited effort on using rich recognition output (e.g., recognition confidence measure or more recognition candidates) for these downstream tasks. This paper presents an initial study using n-best recognition hypotheses for two tasks, extractive summarization and keyword extraction. We extend the approach used on 1-best output to n-best hypotheses: MMR (maximum marginal relevance) for summarization and TFIDF (term frequency, inverse document frequency) weighting for keyword extraction. Our experiments on the ICSI meeting corpus demonstrate promising improvement using n-best hypotheses over 1-best output. These results suggest worthy future studies using n-best or lattices as the interface between speech recognition and downstream tasks. Yang Liu 0004, Shasha Xie, Fei Liu 0004 |
ICASSP | 1 |
| 2010 | Normalization of text messages for text-to-speechabstractThis paper describes a normalization system for text messages to allow them to be read by a TTS engine. To address the large number of texting abbreviations, we use a statistical classifier to learn when to delete a character. The features we use are based on character context, function, and position in the word and containing syllable. To ensure that our system is robust to different abbreviations for a word, we generate multiple abbreviation hypotheses for each word based on the classifier's prediction. We then reverse the mappings to enable prediction of English words from the abbreviations. Our results show that this approach is feasible and warrants further exploration. Deana Pennell, Yang Liu 0004 |
ICASSP | 2 |
| 2010 | Syllable-level prominence detection with acoustic evidence
Je Hun Jeon, Yang Liu 0004 |
INTERSPEECH | 2 |
| 2010 | Level of interest sensing in spoken dialog using multi-level fusion of acoustic and lexical evidenceabstractAccurately sensing a user’s interest in spoken dialog plays a significant role in many applications, such as tutoring systems and customer service systems. In addition to the widely used acoustic evidence, we introduce different lexical features for interest level prediction and evaluate the impact of automatic speech recognition (ASR) on the effectiveness of lexical infor-mation. In order to capture contextual information, we combine the system’s hypothesis for the previous turn with the current one. Our final system uses a multi-level fusion method for this task. Each fusion step uses different information such as acous-tic and lexical cues, contextual information, or hypotheses from different classifiers. Our experiments show that various combi-nations improve system performance. In particular, we found that even though the word error rate is quite high, there is still performance gain by incorporating lexical information obtained from ASR output. Index Terms: Multi-level fusion, level of interest, affect sens-ing 1. Je Hun Jeon, Yang Liu 0004 |
INTERSPEECH | 3 |
| 2010 | Exploring speaker characteristics for meeting summarizationabstractIn this paper, we investigate using meeting-specific characteris-tics to improve extractive meeting summarization, in particular, speaker-related attributes (such as verboseness, gender, native language, role in the meeting). A rich set of speaker-sensitive features are developed in the supervised learning framework. We perform experiments on the ICSI meeting corpus. Re-sults are evaluated using multiple criteria, including ROUGE, a sentence-level F-measure, and an approximated Pyramid ap-proach. We show that incorporating speaker characteristics can consistently improve summarization performance on various testing conditions. 1. Fei Liu 0004, Yang Liu 0004 |
INTERSPEECH | 2 |
| 2010 | Semi-supervised extractive speech summarization via co-training algorithmabstractSupervised methods for extractive speech summarization require a large training set. Summary annotation is often expensive and time consuming. In this paper, we exploit semisupervised approaches to leverage unlabeled data. In particular, we investigate co-training algorithm for the task of extractive meeting summarization. Compared with text summarization, speech summarization task has its unique characteristic in that the features naturally split into two sets: textual features and prosodic/acoustic features. Such characteristic makes co-training an appropriate approach for semi-supervised speech summarization. Our experiments on ICSI meeting corpus show that by utilizing the unlabeled data, co-training algorithm significantly improves summarization performance when only a small amount of labeled data is available. Index Terms: extractive meeting summarization, co-training, semi-supervised learning Shasha Xie, Hui Lin 0001, Yang Liu 0004 |
INTERSPEECH | 3 |
| 2010 | Improving Blog Polarity Classification via Topic Analysis and Adaptive Methods
Yang Liu 0004 |
HLT-NAACL | 4 |
| 2010 | Using Confusion Networks for Speech Summarization
Shasha Xie, Yang Liu 0004 |
HLT-NAACL | 2 |
| 2010 | Using spoken utterance compression for meeting summarization: A pilot studyabstractMost previous work on meeting summarization focused on extractive approaches; however, directly concatenating the extracted spoken utterances may not form a good summary. In this paper, we investigate if it is feasible to compress the transcribed spoken utterances and if using the compressed utterances benefits meeting summarization. We model the utterance compression task as a sequence labeling problem, and show satisfying performance using a CRF model that incorporates a variety of features capturing lexical, syntactic, and discourse information. We evaluate the impact of utterance compression on the meeting summarization task using compressed sentences (pre-compression) and original transcripts (post-compression), and find that using the compressed meeting transcripts yields slightly better summarization performance. In general, using sentence compression together with extractive summarization can generate reasonable compressed summaries. This is a step closer to abstractive summarization. Fei Liu 0004, Yang Liu 0004 |
SLT | 2 |
| 2010 | Improving supervised learning for meeting summarization using sampling and regression
Shasha Xie, Yang Liu 0004 |
Comput. Speech Lang. | 2 |
| 2010 | Speaker adaptation of language and prosodic models for automatic dialog act segmentation of speech
Jáchym Kolár, Yang Liu 0004, Elizabeth Shriberg |
Speech Commun. | 2 |
| 2010 | Identification of Soundbite and Its Speaker Name Using Transcripts of Broadcast News SpeechabstractThis article presents a pipeline framework for identifying soundbite and its speaker name from Mandarin broadcast news transcripts. Both of the two modules, soundbite segment detection and soundbite speaker name recognition, are based on a supervised classification approach using multiple linguistic features. We systematically evaluated performance for each module as well as the entire system, and investigated the effect of using speech recognition (ASR) output and automatic sentence segmentation. We found that both of the two components impact the pipeline system, with more degradation in the entire system performance due to automatic speaker name recognition errors than soundbite segment detection. In addition, our experimental results show that using ASR output degrades the system performance significantly, and that using automatic sentence segmentation greatly impacts soundbite detection, but has much less effect on speaker name recognition. Yang Liu 0004 |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2010 | Exploring Correlation Between ROUGE and Human Evaluation on Meeting SummariesabstractAutomatic summarization evaluation is very important to the development of summarization systems. In text summarization, ROUGE has been shown to correlate well with human evaluation when measuring match of content units. However, there are many characteristics of the multiparty meeting domain, which may pose potential problems to ROUGE. The goal of this paper is to examine how well the ROUGE scores correlate with human evaluation for extractive meeting summarization, and explore different meeting domain specific factors that have an impact on the correlation. More analysis than those in our previous work has been conducted in this study. Our experiments show that generally the correlation between ROUGE and human evaluation is not great; however, when accounting for several unique meeting characteristics, such as disfluencies, speaker information, and stopwords in the ROUGE setting, better correlation can be achieved, especially on the system summaries. We also found that these factors have a different impact on human versus system summaries. In addition, we contrast the results using ROUGE with other automatic summarization evaluation metrics, such as Kappa and Pyramid, and show the appropriateness of using ROUGE for this study. Yang Liu 0004 |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Semi-supervised Learning for Automatic Prosodic Event Detection Using Co-training Algorithm
Je Hun Jeon, Yang Liu 0004 |
ACL/IJCNLP | 2 |
| 2009 | Integrating prosodic features in extractive meeting summarizationabstractSpeech contains additional information than text that can be valuable for automatic speech summarization. In this paper, we evaluate how to effectively use acoustic/prosodic features for extractive meeting summarization, and how to integrate prosodic features with lexical and structural information for further improvement. To properly represent prosodic features, we propose different normalization methods based on speaker, topic, or local context information. Our experimental results show that using only the prosodic features we achieve better performance than using the non-prosodic information on both the human transcripts and recognition output. In addition, a decision-level combination of the prosodic and non-prosodic features yields further gain, outperforming the individual models. Shasha Xie, Dilek Hakkani-Tür, Benoît Favre, Yang Liu 0004 |
ASRU | 4 |
| 2009 | Automatic prosodic events detection using syllable-based acoustic and syntactic featuresabstractAutomatic prosodic event detection is important for both speech understanding and natural speech synthesis since prosody provides additional information over the short-term segmental features and lexical representation of an utterance. Similar to previous work, this paper focuses on automatic detection of coarse level representation of pitch accents, intonational phrase boundaries (IPB), and break indices. We exploit various classifiers and identify effective feature sets to improve performance of prosodic event detection according to acoustic, lexical, and syntactic evidence. our experiments on the Boston University Radio News Corpus show that the neural network classifier achieves the best performance for modeling acoustic evidence, and that support vector machines are more effective for the lexical and syntactic evidence. The combination of the acoustic and the syntactic models yields 89.8% accent detection accuracy, 93.3% IPB detection accuracy, and 91.1% break index detection accuracy. Compared with previous work, the IPB performance is similar, whereas the results for accent and break index detection are significantly better. Je Hun Jeon, Yang Liu 0004 |
ICASSP | 2 |
| 2009 | Genre effects on automatic sentence segmentation of speech: A comparison of broadcast news and broadcast conversationsabstractWe investigate genre effects on the task of automatic sentence segmentation, focusing on two important domains - broadcast news (BN) and broadcast conversation (BC). We employ an HMM model based on textual and prosodic information and analyze differences in segmentation accuracy and feature usage between the two genres using both manual and automatic speech transcripts. Experiments are evaluated using Czech broadcast corpora annotated for sentence-like units (SUs). Prosodic features capture information about pause, duration, pitch, and energy patterns. Textual knowledge sources include words, part-of-speech, and automatically induced classes. We also analyze effects of using additional textual data that is not annotated for SUs. Feature analysis reveals significant differences in both textual and prosodic feature usage patterns between the two genres. The analysis is important for building automatic understanding systems when limited matched-genre data are available, or for designing eventual genre-independent systems. Jáchym Kolár, Yang Liu 0004, Elizabeth Shriberg |
ICASSP | 2 |
| 2009 | Finding Opinionated Blogs Using Statistical Classifiers and Lexical Features
Yang Liu 0004 |
ICWSM | 3 |
| 2009 | Automatic accent detection: effect of base units and boundary information
Je Hun Jeon, Yang Liu 0004 |
INTERSPEECH | 2 |
| 2009 | Leveraging sentence weights in a concept-based optimization framework for extractive meeting summarizationabstractInternational audience Shasha Xie, Benoît Favre, Dilek Hakkani-Tür, Yang Liu 0004 |
INTERSPEECH | 4 |
| 2009 | A Corpus-Based Approach for the Prediction of Language Impairment in Monolingual English and Spanish-English Bilingual Children
Keyur Gabani, Melissa Sherman, Thamar Solorio, Yang Liu 0004, Lisa Bedore, Elizabeth Peña |
HLT-NAACL | 4 |
| 2009 | Unsupervised Approaches for Automatic Keyword Extraction Using Meeting Transcripts
Deana Pennell, Fei Liu 0004, Yang Liu 0004 |
HLT-NAACL | 4 |
| 2008 | Learning to Predict Code-Switching Points
Thamar Solorio, Yang Liu 0004 |
EMNLP | 2 |
| 2008 | Part-of-Speech Tagging for English-Spanish Code-Switched Text
Thamar Solorio, Yang Liu 0004 |
EMNLP | 2 |
| 2008 | Unsupervised language model adaptation via topic modeling based on named entity hypothesesabstractLanguage model (LM) adaptation is often achieved by combining a generic LM with a topic-specific model that is more relevant to the target document. Unlike previous work on unsupervised LM adaptation, in this paper we propose to leverage named entity (NE) information for topic analysis and LM adaptation. We investigate two topic modeling approaches, latent Dirichlet allocation (LDA) and clustering, and proposed a new mixture topic model for LDA based LM adaptation. Our experiments for N-best list rescoring have shown that this new adaptation framework using NE information and topic analysis outperforms the baseline generic N-gram LM based on a state-of-the-art Mandarin recognition system. Yang Liu 0004 |
ICASSP | 1 |
| 2008 | Impact of automatic sentence segmentation on meeting summarizationabstractThis paper investigates the impact of automatic sentence segmentation on speech summarization using the ICSI meeting corpus. We use a hidden Markov model (HMM) for sentence segmentation that integrates the N-gram language model and pause information, and a maximum marginal relevance (MMR) based extractive summarization method. The system-generated summaries are compared to multiple human summaries using the ROUGE scores. The decision thresholds from the segmentation system are varied to examine the impact of different segments on summarization. We find that (1) using system generated utterance segments degrades summarization performance compared to using human annotated sentences; (2) segmentation needs to be optimized for summarization instead of the segmentation task itself, however, the patterns are slightly different from prior work for other tasks such as parsing; and (3) there are effects from different summarization evaluation metrics as well as speech recognition errors. Yang Liu 0004, Shasha Xie |
ICASSP | 1 |
| 2008 | Using corpus and knowledge-based similarity measure in Maximum Marginal Relevance for meeting summarizationabstractMMR (maximum marginal relevance) is widely used in summarization for its simplicity and efficacy, and has been demonstrated to achieve comparable performance to other approaches for meeting summarization. How to appropriately represent the similarity of two text segments is crucial in MMR. In this paper, we evaluate different similarity measures in the MMR framework for meeting summarization on the ICSI meeting corpus. We introduce a corpus- based measure to capture the similarity at the semantic level, and compare this method with cosine similarity and centroid score that only considers the salient words in the segments. Our experimental results evaluated by the ROUGE summarization metrics show that both the centroid score and the corpus-based similarity measure yield better performance than the commonly used cosine similarity. In addition, adding part-of-speech information in the corpus-based approach helps for the human transcripts condition, but not when using ASR output. Shasha Xie, Yang Liu 0004 |
ICASSP | 2 |
| 2008 | Automatic keyword extraction for the meeting corpus using supervised approach and bigram expansionabstractIn this paper, we tackle the problem of automatic keyword extraction in the meeting domain, a genre significantly different from written text. For the supervised framework, we proposed a rich set of features beyond the typical TFIDF measures, such as sentence salience weight, lexical features, summary sentences, and speaker information. We also evaluate different candidate sampling approaches for better model training and testing. In addition, we introduced a bigram expansion module which aims at extracting ldquoentity bigramsrdquo using Web resources. Using the ICSI meeting corpus, we demonstrate the effectiveness of the features and show that the supervised method and the bigram expansion module outperform the unsupervised TFIDF selection with POS (part-of-speech) filtering. Finally, we show the approaches introduced in this paper perform well on the speech recognition output. Fei Liu 0004, Yang Liu 0004 |
SLT | 3 |
| 2008 | Using hidden Markov models for topic segmentation of meeting transcriptsabstractIn this paper, we present a hidden Markov model (HMM) approach to segment meeting transcripts into topics. To learn the model, we use unsupervised learning to cluster the text segments obtained from topic boundary information. Using modified WinDiff and Pkmetrics, we demonstrate that an HMM outperforms LCSeg, a state-of-the-art lexical chain based method for topic segmentation using the ICSI meeting corpus. We evaluate the effect of language model order, the number of hidden states, and the use of stop words. Our experimental results show that a unigram LM is better than a trigram LM, using too many hidden states degrades topic segmentation performance, and that removing the stop words from the transcripts does not improve segmentation performance. Melissa Sherman, Yang Liu 0004 |
SLT | 2 |
| 2008 | Evaluating the effectiveness of features and sampling in extractive meeting summarizationabstractFeature-based approaches are widely used in the task of extractive meeting summarization. In this paper, we analyze and evaluate the effectiveness of different types of features using forward feature selection in an SVM classifier. In addition to features used in prior studies, we introduce topic related features and demonstrate that these features are helpful for meeting summarization. We also propose a new way to resample the sentences based on their salience scores for model training and testing. The experimental results on both the human transcripts and recognition output, evaluated by the ROUGE summarization metrics, show that feature selection and data resampling help improve the system performance. Shasha Xie, Yang Liu 0004, Hui Lin 0001 |
SLT | 2 |
| 2007 | Unsupervised Language Model Adaptation Incorporating Named Entity Information
Yang Liu 0004 |
ACL | 2 |
| 2007 | Soundbite identification using reference and automatic transcripts of broadcast news speechabstractSoundbite identification in broadcast news is important for locating information useful for question answering, mining opinions of a particular person, and enriching speech recognition output with quotation marks. This paper presents a systematic study of this problem under a classification framework, including problem formulation for classification, feature extraction, and the effect of using automatic speech recognition (ASR) output and automatic sentence boundary detection. Our experiments on a Mandarin broadcast news speech corpus show that the three-way classification framework outperforms the binary classification. The entropy-based feature weighting method generally performs better than others. Using ASR output degrades system performance, with more degradation observed from using automatic sentence segmentation than speech recognition errors for this task, especially on the recall rate. Yang Liu 0004 |
ASRU | 2 |
| 2007 | Comparing Evaluation Metrics for Sentence Boundary DetectionabstractIn recent NIST evaluations on sentence boundary detection, a single error metric was used to describe performance. Additional metrics, however, are available for such tasks, in which a word stream is partitioned into subunits. This paper compares alternative evaluation metrics - including the NIST error rate, classification error rate per word boundary, precision and recall, ROC curves, DET curves, precision-recall curves, and area under the curves - and discusses advantages and disadvantages of each. Unlike many studies in machine learning, we use real data for a real task. We find benefit from using curves in addition to a single metric. Furthermore, we find that data skew has an impact on metrics, and that differences among different system outputs are more visible in precision-recall curves. Results are expected to help us better understand evaluation metrics that should be generalizable to similar language processing tasks. Yang Liu 0004, Elizabeth Shriberg |
ICASSP (4) | 1 |
| 2007 | Speaker adaptation of language models for automatic dialog act segmentation of meetingsabstractAbstract Dialog act (DA) segmentation in meeting speech is important for meeting understanding. In this paper, we explore speaker adaptation of hidden event language models (LMs) for DA segmentation using the ICSI Meeting Corpus. Speaker adaptation is performed using a linear combination of the generic speaker independent LM and an LM trained on only the data from individual speakers. We test the method on 20 frequent speakers, on both reference word transcripts and the output of automatic speech recognition. Results indicate improvements for 17 speakers on reference transcripts, and for 15 speakers on automatic transcripts. Overall, the speaker-adapted LM yields statistically significant improvement over the baseline LM for both test conditions. Jáchym Kolár, Yang Liu 0004, Elizabeth Shriberg |
INTERSPEECH | 2 |
| 2006 | PCFGs with Syntactic and Prosodic Indicators of Speech RepairsabstractA grammatical method of combining two kinds of speech repair cues is presented. One cue, prosodic disjuncture, is detected by a decision tree-based ensemble classifier that uses acoustic cues to identify where normal prosody seems to be interrupted (Lickley, 1996). The other cue, syntactic parallelism, codifies the expectation that repairs continue a syntactic category that was left unfinished in the reparandum (Levelt, 1983). The two cues are combined in a Treebank PCFG whose states are split using a few simple tree transformations. Parsing performance on the Switchboard and Fisher corpora suggests that these two cues help to locate speech repairs in a synergistic way. John Hale, Izhak Shafran, Lisa Yung, Bonnie J. Dorr, Mary P. Harper, Anna Krasnyanskaya, Matthew Lease, Yang Liu 0004, Brian Roark, Matthew G. Snover, Robin Stewart |
ACL | 8 |
| 2006 | Reranking for Sentence Boundary Detection in Conversational SpeechabstractWe present a reranking approach to sentence-like unit (SU) boundary detection, one of the EARS metadata extraction tasks. Techniques for generating relatively small n-best lists with high oracle accuracy are presented. For each candidate, features are derived from a range of information sources, including the output of a number of parsers. Our approach yields significant improvements over the best performing system from the NIST RT-04F community evaluation Brian Roark, Yang Liu 0004, Mary P. Harper, Robin Stewart, Matthew Lease, Matthew G. Snover, Izhak Shafran, Bonnie J. Dorr, John Hale, Anna Krasnyanskaya, Lisa Yung |
ICASSP (1) | 2 |
| 2006 | On speaker-specific prosodic models for automatic dialog act segmentation of multi-party meetingsabstractTento článek zkoumá prozodické modely specifické pro jednotlivé řečníky, které jsou používány pro automatickou segmentaci řeči z ICSI meetings korpusu na dialogové akty. Zkoumáme, zda-li je výhodné trénovat tyto modely pouze na řeči konkrétního řečníka. Jáchym Kolár, Elizabeth Shriberg, Yang Liu 0004 |
INTERSPEECH | 3 |
| 2006 | Using SVM and error-correcting codes for multiclass dialog act classification in meeting corpusabstractAccurate classification of dialog acts (DAs) is important for many spoken language applications. Different methods have been proposed such as hidden Markov models (HMM), maximum entropy (Maxent), graphical models, and support vector machines (SVMs). In this paper, we investigate using SVMs for multiclass DA classification in the ICSI meeting corpus. We evaluate (1) representing DA tagging directly as a multiclass task, and (2) combining multiple binary classifiers via error correction output codes (ECOC). For the ECOC combination, different code matrices are utilized (e.g., the identity matrix, exhaustive code, BCH code, and random code matrix). We also compare using SVMs with our previous Maxent model. We find that for DA tagging, using multiple binary SVMs via ECOC outperforms a direct multiclass SVM, but neither achieves better performance than the Maxent model, possibly because of the small class set and the features currently used in the task. Index Terms: dialog act classification, SVM, maximum entropy, error correction output code. Yang Liu 0004 |
INTERSPEECH | 1 |
| 2006 | The ICSI+ multilingual sentence segmentation systemabstractThe ICSI+ multilingual sentence segmentation with results for English and Mandarin broadcast news automatic speech recognizer transcriptions represents a joint effort involving ICSI, SRI, and UT Dallas. Our approach is based on using hidden event language models for exploiting lexical information, and maximum entropy and boosting classifiers for exploiting lexical, as well as prosodic, speaker change and syntactic information. We demonstrate that the proposed methodology including pitch- and energy-related prosodic features performs significantly better than a baseline system that uses words and simple pause features only. Furthermore, the obtained improvements are consistent across both languages, and no language-specific adaptation of the methodology is necessary. The best results were achieved by combining hidden event language models with a boosting-based classifier that to our knowledge has not previously been applied for this task. M. Zimmerman, Dilek Hakkani-Tür, James G. Fung, Nikki Mirghafori, Luke R. Gottlieb, Elizabeth Shriberg, Yang Liu 0004 |
INTERSPEECH | 7 |
| 2006 | Linguistic Resources for Speech Parsing
Ann Bies, Stephanie M. Strassel, Haejoong Lee, Kazuaki Maeda, Seth Kulick, Yang Liu 0004, Mary P. Harper, Matthew Lease |
LREC | 6 |
| 2006 | SParseval: Evaluation Metrics for Parsing Speech
Brian Roark, Mary P. Harper, Eugene Charniak, Bonnie J. Dorr, Mark Johnson 0001, Jeremy G. Kahn, Yang Liu 0004, Mari Ostendorf, John Hale, Anna Krasnyanskaya, Matthew Lease, Izhak Shafran, Matthew G. Snover, Robin Stewart, Lisa Yung |
LREC | 7 |
| 2006 | Initial Study on Automatic Identification of Speaker Role in Broadcast News Speech
Yang Liu 0004 |
HLT-NAACL | 1 |
| 2006 | A study in machine learning from imbalanced data for sentence boundary detection in speech
Yang Liu 0004, Nitesh V. Chawla, Mary P. Harper, Elizabeth Shriberg, Andreas Stolcke |
Comput. Speech Lang. | 1 |
| 2006 | Enriching speech recognition with automatic detection of sentence boundaries and disfluenciesabstractEffective human and automatic processing of speech requires recovery of more than just the words. It also involves recovering phenomena such as sentence boundaries, filler words, and disfluencies, referred to as structural metadata. We describe a metadata detection system that combines information from different types of textual knowledge sources with information from a prosodic classifier. We investigate maximum entropy and conditional random field models, as well as the predominant hidden Markov model (HMM) approach, and find that discriminative models generally outperform generative models. We report system performance on both broadcast news and conversational telephone speech tasks, illustrating significant performance differences across tasks and as a function of recognizer performance. The results represent the state of the art, as assessed in the NIST RT-04F evaluation Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Dustin Hillard, Mari Ostendorf, Mary P. Harper |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Using Conditional Random Fields for Sentence Boundary Detection in SpeechabstractSentence boundary detection in speech is important for enriching speech recognition output, making it easier for humans to read and downstream modules to process. In previous work, we have developed hidden Markov model (HMM) and maximum entropy (Maxent) classifiers that integrate textual and prosodic knowledge sources for detecting sentence boundaries. In this paper, we evaluate the use of a conditional random field (CRF) for this task and relate results with this model to our prior work. We evaluate across two corpora (conversational telephone speech and broadcast news speech) on both human transcriptions and speech recognition output. In general, our CRF model yields a lower error rate than the HMM and Maxent models on the NIST sentence boundary detection task in speech, although it is interesting to note that the best results are achieved by three-way voting among the classifiers. This probably occurs because each model has different strengths and weaknesses for modeling the knowledge sources. Yang Liu 0004, Andreas Stolcke, Elizabeth Shriberg, Mary P. Harper |
ACL | 1 |
| 2005 | Automatic Dialog Act Segmentation and Classification in Multiparty MeetingsabstractWe explore the two related tasks of dialog act (DA) segmentation and DA classification for speech from the ICSI Meeting Corpus. We employ simple lexical and prosodic knowledge sources, and compare results for human-transcribed versus automatically recognized words. Since there is little previous work on DA segmentation and classification in the meeting domain, our study provides baseline performance rates for both tasks. We introduce a range of metrics for use in evaluation, each of which measures different aspects of interest. Results show that both tasks are difficult, particularly for a fully automatic system. We find that a very simple prosodic model aids performance over lexical information alone, especially for segmentation. Both tasks, but particularly word-based segmentation, are degraded by word recognition errors. Finally, while classification results for meeting data show some similarities to previous results for telephone conversations, findings also suggest a potential difference with respect to the effect of modeling DA context. Jeremy Ang, Yang Liu 0004, Elizabeth Shriberg |
ICASSP (1) | 2 |
| 2005 | Structural metadata research in the EARS programabstractBoth human and automatic processing of speech require recognition of more than just words. In this paper we provide a brief overview of research on structural metadata extraction in the DARPA EARS rich transcription program. Tasks include detection of sentence boundaries, filler words, and disfluencies. Modeling approaches combine lexical, prosodic, and syntactic information, using various modeling techniques for knowledge source integration. The performance of these methods is evaluated by task, by data source (broadcast news versus spontaneous telephone conversations) and by whether transcriptions come from humans or from an (errorful) automatic speech recognizer. A representative sample of results shows that combining multiple knowledge sources (words, prosody, syntactic information) is helpful, that prosody is more helpful for news speech than for conversational speech, that word errors significantly impact performance, and that discriminative models generally provide benefit over maximum likelihood models. Important remaining issues, both technical and programmatic, are also discussed. Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Barbara Peskin, Jeremy Ang, Dustin Hillard, Mari Ostendorf, Marcus Tomalin, Philip C. Woodland, Mary P. Harper |
ICASSP (5) | 1 |
| 2005 | Comparing HMM, maximum entropy, and conditional random fields for disfluency detectionabstractAutomatic detection of disfluencies in spoken language is important for making speech recognition output more readable, and for aiding downstream language processing modules. We compare a generative hidden Markov model (HMM)-based approach and two conditional models — a maximum entropy (Maxent) model and a conditional random field (CRF) — for detecting disfluencies in speech. The conditional modeling approaches provide a more principled way to model correlated features. In particular, the CRF approach directly detects the reparandum regions, and thus avoids the use of ad-hoc heuristic rules. We evaluate performance of these three models across two different corpora (conversational speech and broadcast news) and for two types of transcriptions (human transcriptions and recognition output). Overall we find that that the conditional modeling approaches (Maxent and CRF) provide benefit over the HMM approach. Effects of speaking style, word recognition errors, and future directions are also discussed. 1. Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Mary P. Harper |
INTERSPEECH | 1 |
| 2005 | Does active learning help automatic dialog act tagging in meeting data?abstractKnowledge of Dialog Acts (DAs) is important for the automatic understanding and summarization of meetings. Current approaches rely on a lot of hand labeled data to train automatic taggers. One approach that has been successful in reducing the amount of training data in other areas of NLP is active learning. We ask if active learning with lexical cues can help for this task and this domain. To better address this question, we explore active learning for two different types of DA models – hidden Markov models (HMMs) and maximum entropy (maxent). Anand Venkataraman, Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 2 |
| 2004 | Comparing and Combining Generative and Posterior Probability Models: Some Advances in Sentence Boundary Detection in Speech
Yang Liu 0004, Andreas Stolcke, Elizabeth Shriberg, Mary P. Harper |
EMNLP | 1 |
| 2004 | Using machine learning to cope with imbalanced classes in natural speech: evidence from sentence boundary and disfluency detectionabstractWe investigate machine learning techniques for coping with highly skewed class distributions in two spontaneous speech processing tasks. Both tasks, sentence boundary and disfluency detection, provide important structural information for downstream language processing modules. We examine the effect of data set size, task, sampling method (no sampling, downsampling, oversampling, and ensemble sampling), and learning method (bagging, ensemble bagging, and boosting) for a decision tree prosody model. Results show that (1) bagging benefits both tasks, but to different degrees, (2) the benefit from ensemble bagging decreases as data size increases, and (3) boosting can outperform bagging under certain conditions. Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke, Mary P. Harper |
INTERSPEECH | 1 |
| 2004 | The ICSI-SRI-UW metadata extraction systemabstractBoth human and automatic processing of speech require recognizing more than just the words. We describe a state-of-the-art system for automatic detection of “metadata” (information beyond the words) in both broadcast news and spontaneous telephone conversations, developed as part of the DARPA EARS Rich Transcription program. System tasks include sentence boundary detection, filler word detection, and detection/correction of disfluencies. To achieve best performance, we combine information from different types of language models (based on words, part-of-speech classes, and automatically induced classes) with information from a prosodic classifier. The prosodic classifier employs bagging and ensemble approaches to better estimate posterior probabilities. We use confusion networks to improve robustness to speech recognition errors. Most recently, we have investigated a maximum entropy approach for the sentence boundary detection task, yielding a gain over our standard HMM approach. We report results for these techniques on the official NIST Rich Transcription metadata tasks. Elizabeth Shriberg, Andreas Stolcke, Dustin Hillard, Mari Ostendorf, Barbara Peskin, Mary P. Harper, Yang Liu 0004 |
INTERSPEECH | 7 |
| 2004 | Evaluating Factors Impacting the Accuracy of Forced Alignments in a Multimodal Corpus
Lei Chen 0004, Yang Liu 0004, Mary P. Harper, Eduardo Maia, Susan McRoy |
LREC | 2 |
| 2003 | Automatic disfluency identification in conversational speech using multiple knowledge sourcesabstractDisfluencies occur frequently in spontaneous speech. Detection and correction of disfluencies can make automatic speech recognition transcripts more readable for human readers, and can aid downstream processing by machine. This work investigates a number of knowledge sources for disfluency detection, including acoustic-prosodic features, a language model (LM) to account for repetition patterns, a part-of-speech (POS) based LM, and rule-based knowledge. Different components are designed for different purposes in the system. Results show that detection of disfluency interruption points is best achieved by a combination of prosodic cues, word-based cues, and POS-based cues. The onset of a disfluency to be removed, in contrast, is best found using knowledge-based rules. Finally, specific disfluency types can be aided by the modeling of word patterns. 1. Yang Liu 0004, Elizabeth Shriberg, Andreas Stolcke |
INTERSPEECH | 1 |
| 2003 | Word Fragments Identification Using Acoustic-Prosodic Features in Conversational Speech
Yang Liu 0004 |
HLT-NAACL | 1 |
| 2003 | The effect of pruning and compression on graphical representations of the output of a speech recognizer
Yang Liu 0004, Mary P. Harper, Michael T. Johnson, Leah H. Jamieson |
Comput. Speech Lang. | 1 |
| 2002 | Rescoring effectiveness of language models using different levels of knowledge and their integrationabstractIn this paper, we compare the efficacy of a variety of language models (LMs) for rescoring word graphs and N-best lists generated by a large vocabulary continuous speech recognizer. These LMs differ based on the level of knowledge used (word, lexical features, syntax) and the type of integration of that knowledge (tight or loose). The trigram LM incorporates word level information; our Part-of-Speech (POS) LM uses word and lexical class information in a tightly coupled way; our new SuperARV LM tightly integrates word, a richer set of lexical features than POS, and syntactic dependency information; and the Parser LM integrates some limited word information, POS, and syntactic information. We also investigate LMs created using a linear interpolation of LM pairs. When comparing each LM on the task of rescoring word graphs or N-best lists for the Wall Street Journal (WSJ) 5k- and 20k- vocabulary test sets, the SuperARV LM always achieves the greatest reduction in word error rate (WER) and the greatest increase in sentence accuracy (SAC). On the 5k test sets, the SuperARV LM obtains more than a 10% relative reduction in WER compared to the trigram LM, and on the 20k test set more than 2%. Additionally, the SuperARV LM performs comparably to or better than the interpolated LMs. Hence, we conclude that the tight coupling of knowledge from all three levels is an effective method of constructing high quality LMs. Wen Wang 0001, Yang Liu 0004, Mary P. Harper |
ICASSP | 2 |
| 2000 | A language model adaptation approach based on text classification
Jiasong Sun, Zuoying Wang, Yang Liu 0004 |
INTERSPEECH | 4 |