Nancy F. Chen

dblp:84/8761 · DBLP profile ↗
← Back
146ranked-venue papers
14as first author
83since 2021 · last 2026
0000-0003-0872-5877ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 107 · 7 first-author · 65 since 2021Graphics, computer vision, multimedia, augmented reality and games · 80 · 13 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 AdaMCoT: Rethinking Cross-Lingual Factual Reasoning Through Adaptive Multilingual Chain-of-Thought
abstract
Large language models (LLMs) have shown impressive multilingual capabilities through pretraining on diverse corpora. While these models show strong reasoning abilities, their performance varies significantly across languages due to imbalanced training data distribution. Existing approaches using sample-level translation for extensive multilingual pretraining and cross-lingual tuning face scalability challenges and often fail to capture nuanced reasoning processes across languages. In this paper, we introduce **AdaMCoT** (Adaptive Multilingual Chain-of-Thought), a framework that enhances multilingual factual reasoning by dynamically routing thought processes in intermediary “thinking languages” before generating target-language responses. AdaMCoT leverages a language-agnostic core and incorporates an adaptive, reward-based mechanism for selecting optimal reasoning pathways without requiring additional pretraining. Our comprehensive evaluation across multiple benchmarks demonstrates substantial improvements in both factual reasoning quality and cross-lingual consistency, with particularly strong performance gains in low-resource language settings. An in-depth analysis of the model’s hidden states and semantic space further elucidates the underlying mechanism of our method. The results suggest that adaptive reasoning paths can effectively bridge the performance gap between high- and low-resource languages while maintaining cultural and linguistic nuances.
Zhengyuan Liu, Tarun Kumar Vangani, Bowei Zou, Xiyan Tao, AiTi Aw, Nancy F. Chen, Roy Ka-Wei Lee
AAAI9
2026 Programming over Thinking: Efficient and Robust Multi-Constraint Planning
abstract
Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting constraints.Existing large language model (LLM) approaches face fundamental limitations in this domain.Pure reasoning paradigms, which rely on long natural language chains, are prone to inconsistency, error accumulation, and prohibitive cost as constraints compound.Conversely, LLMs combined with coding-or solver-based strategies lack flexibility: they often generate problem-specific code from scratch or depend on fixed solvers, failing to capture generalizable logic across diverse problems.To address these challenges, we introduce the Scalable COde Planning Engine (SCOPE), a framework that disentangles query-specific reasoning from generic code execution.By separating reasoning from execution, SCOPE produces solver functions that are consistent, deterministic, and reusable across queries while requiring only minimal changes to input parameters.SCOPE achieves state-of-the-art performance while lowering cost and latency.For example, with GPT-4o, it reaches 93.1% success on TravelPlanner, a 61.6% gain over the best baseline (CoT) while cutting inference cost by 1.4x and time by 4.67x.Code is available at https://github.com/DerrickGXD/SCOPE.
Derrick Goh Xin Deik, Quanyu Long, Zhengyuan Liu, Nancy F. Chen, Wenya Wang 0001
ACL (1)4
2026 Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment
abstract
Bryan Chen Zhengyu Tan, Zhengyuan Liu, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Roy Ka-Wei Lee. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Bryan Chen Zhengyu Tan, Zhengyuan Liu, Xiaoyuan Yi, Jing Yao 0003, Xing Xie 0001, Nancy F. Chen, Roy Ka-Wei Lee
ACL (1)6
2026 MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation
abstract
Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengyuan Liu, Tanmoy Chakraborty 0002, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu 0071, Xing Xie 0001, Xiaoyuan Yi, Jing Yao 0003, Chaojun Wang, Rui Liu 0019, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Lingyu Ye, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen
ACL (1)35
2026 Leveraging LLMs for Dynamic Engagement Pattern Recognition in Collaborative Learning
Stella Xin Yin, Zhengyuan Liu, Dion Hoe-Lian Goh, Nancy F. Chen
AIED4
2026 Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts
abstract
We present a conceptual framework for analyzing dialogue in collaborative problem-solving contexts, with an emphasis on the emerging dynamics of human-AI and multi-agent collaboration. As intelligent systems become active agents capable of autonomous reasoning and strategic cooperation, understanding the dialogic interaction during collaborative problem solving is increasingly important for optimizing and evaluating such partnerships. Our framework addresses key limitations in current analytical approaches through a hierarchical two-layer coding scheme that integrates cognitive and non-cognitive problem solving with metacognitive regulatory mechanisms. We demonstrate its effectiveness and generalizability across nine datasets spanning multiple domains, and provide insights into how humans and agents coordinate their knowledge, skills, and efforts to solve complex problems, showing in particular that metacognitive regulation can be an essential discriminator of deeper collaboration.
Zhengyuan Liu, Stella Xin Yin, Min-Yen Kan, Nancy F. Chen
SIGDIAL4
2026 SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series
abstract
We introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize local understanding of adjacent frames or clips. SagaQA addresses this gap by requiring high-level comprehension of extended multimodal narratives in entire TV shows. A distinguishing feature of SagaQA is the granularity of its reasoning steps. Our dataset necessitates long-range reasoning hops to connect information across completely different episodes. This requires models to reason over entire events and actions, demanding a deep understanding of the show’s narration and progression at a multimodal level. Motivated by recent progress in agentic methods, we further study how different planning strategies handle such complex reasoning. We categorize these approaches into three classes—parallel, sequential, and hybrid planners—and evaluate their ability to generate coherent and complete reasoning plans. Our results on SagaQA suggest that hybrid planners consistently produce higher-quality plans and exhibit stronger capabilities for complex, high-level narrative understanding in TV shows.
Galann Pennec, Zhengyuan Liu, Nicholas Asher, Philippe Muller, Nancy F. Chen
SIGDIAL5
2026 Structure from rank: Rank-order coding as a bridge from sequence to structure
abstract
Understanding how structured sequence information can be represented and generalized in neural systems is key to modeling the transition from acoustic input to emergent structure. In this study, we propose a rank-order based neural network inspired by the STG-LIFG-PMC pathway, modeling the bottom-up transition from acoustic input to abstract rank representation and the top-down generation from that representation to motor execution. Building on previous work in rank coding, we first demonstrate that this model efficiently compresses input while retaining the capacity to reconstruct full utterances from partial cues, revealing emergent structure-sensitive generation process that reflects context-general representations of sensorimotor states, which are later shaped into context-specific motor plans during speech planning. We then show that the network exhibits global-level novelty detection similar to the P3B novelty wave, replicating the global-sequence-sensitive mechanism. As a supplement, we also compare the model's behavior under local (index-level) and global (rank-level) perturbations, revealing robustness to superficial variation and sensitivity to abstract structural violation, key features associated with hierarchical generalization. These results suggest that rank-order coding not only serves as a compact encoding scheme but also captures hierarchical structure in acoustic sequences.
Alex Pitti, Mathias Quoy, Nancy F. Chen
Neural Networks4
2025 MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues
abstract
Automatic Speech Recognition (ASR) systems are pivotal in transcribing speech into text, yet the errors they introduce can significantly degrade the performance of downstream tasks like summarization. This issue is particularly pronounced in clinical dialogue summarization, a low-resource domain where supervised data for fine-tuning is scarce, necessitating the use of ASR models as black-box solutions. Employing conventional data augmentation for enhancing the noise robustness of summarization models is not feasible either due to the unavailability of sufficient medical dialogue audio recordings and corresponding ASR transcripts. To address this challenge, we propose MEDSAGE, an approach for generating synthetic samples for data augmentation using Large Language Models (LLMs). Specifically, we leverage the in-context learning capabilities of LLMs and instruct them to generate ASR-like errors based on a few available medical dialogue examples with audio recordings. Experimental results show that LLMs can effectively model ASR noise, and incorporating this noisy data into the training process significantly improves the robustness and accuracy of medical dialogue summarization systems. This approach addresses the challenges of noisy ASR outputs in critical applications, offering a robust solution to enhance the reliability of clinical dialogue summarization.
Kuluhan Binici, Abhinav Ramesh Kashyap, Viktor Schlegel, Andy T. Liu, Vijay Prakash Dwivedi, Thanh-Tung Nguyen, Xiaoxue Gao, Nancy F. Chen, Stefan Winkler 0001
AAAI8
2025 What Makes a Good Natural Language Prompt?
abstract
Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi, Nancy F. Chen, Shafiq Joty, Min-Yen Kan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi, Nancy F. Chen, Shafiq R. Joty, Min-Yen Kan
ACL (1)5
2025 Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
abstract
Voiced Electromyography(EMG)-to-Speech (V-ETS) models reconstruct speech from muscle activity signals, facilitating applications such as neurolaryngologic diagnostics. Despite its potential, the advancement of V-ETS is hindered by a scarcity of paired EMG-speech data. To address this, we propose a novel Confidence-based Multi-Speaker Self-training (CoM2S) approach, along with a newly curated Libri-EMG dataset. This approach leverages synthetic EMG data generated by a pretrained model, followed by a proposed filtering mechanism based on phoneme-level confidence to enhance the V-ETS model through the proposed self-training techniques. Experiments demonstrate our method improves phoneme accuracy, reduces phonological confusion, and lowers word error rate, confirming the effectiveness of our CoM2S approach for V-ETS. In support of future research, we will release the codes and the proposed Libri-EMG dataset-an open-access, time-aligned, multi-speaker voiced EMG and speech recordings.
Xiaoxue Gao, Mathias Quoy, Alex Pitti, Nancy F. Chen
ASRU5
2025 A correlation-permutation approach for speech-music encoders model merging
abstract
Creating a unified speech and music model requires expensive pre-training. Model merging can instead create a unified audio model with minimal computational expense. However, direct merging is challenging when the models are not aligned in the weight space. Motivated by Git Re-Basin, we introduce a correlation-permutation approach that aligns a music encoder’s internal layers with a speech encoder. We extend previous work to the case of merging transformer layers. The method computes a permutation matrix that maximizes the model’s features-wise cross-correlations layer by layer, enabling effective fusion of these otherwise disjoint models. The merged model retains speech capabilities through this method while significantly enhancing music performance, achieving an improvement of 14.83 points in average score compared to the linear interpolation model merging. This work allows the creation of unified audio models from independently trained encoders.
Fabian Ritter Gutierrez, Yi-Cheng Lin, Jeremy H. M. Wong, Hung-yi Lee, Chng Eng Siong, Nancy F. Chen
ASRU6
2025 ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
abstract
Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically predict music impression (MI) as well as text alignment (TA) between the prompt and the generated musical piece. This paper reports our winning system, which uses a dual-branch architecture with pre-trained MuQ and RoBERTa models as audio and text encoders. A cross-attention mechanism fuses the audio and text representations. For training, we reframe the MI and TA prediction as a classification task. To incorporate the ordinal nature of MOS scores, one-hot labels are converted to a soft distribution using a Gaussian kernel. On the official test set, a single model trained with this method achieves a system-level Spearman’s Rank Correlation Coefficient (SRCC) of 0.991 for MI and 0.952 for TA, corresponding to a relative improvement of $21.21 \%$ in MI SRCC and $31.47 \%$ in TA SRCC over the challenge baseline.
Fabian Ritter Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei, Jeremy H. M. Wong, Nancy F. Chen, Hung-yi Lee
ASRU5
2025 Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
abstract
Current large speech language models (SpeechLLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextual content and paralinguistic cues. In this work, we propose two approaches to incorporate contextual paralinguistic information into model training: (1) an explicit method that provides paralinguistic metadata (e.g., emotion annotations) directly to the LLM, and (2) an implicit method that automatically generates novel training question-answer (QA) pairs using both categorical and dimensional emotion annotations alongside speech transcriptions. Our implicit method boosts performance (LLM-judged) by 38.41% on a human-annotated QA benchmark, reaching 46.02% when combined with the explicit approach, showing effectiveness in contextual paralinguistic understanding. We also validate the LLM judge by demonstrating its correlation with classification metrics, providing support for its reliability.
Qiongqiong Wang, Hardik Bhupendra Sailor, Jeremy H. M. Wong, Tianchi Liu 0004, Muhammad Huzaifah 0001, Nancy F. Chen, AiTi Aw
ASRU8
2025 MNSC: Advancing Singlish Speech Understanding with Carefully Curated Corpora
abstract
Singlish, a Creole language rooted in English, is a key focus in linguistic research within multilingual and multicultural contexts. However, its spoken form remains underexplored, limiting insights into its linguistic structure and applications. To address this gap, we standardize and annotate the largest spoken Singlish corpus, introducing the Multitask National Speech Corpus (MNSC). These datasets support diverse tasks, including Automatic Speech Recognition (ASR), Spoken Question Answering (SQA), Spoken Dialogue Summarization (SDS), and Paralinguistic Question Answering (PQA). We release standardized splits and a human-verified test set to facilitate further research. Additionally, we propose SingAudioLLM, a multi-task multimodal model leveraging multimodal large language models to handle these tasks concurrently. Experiments reveal our models’ adaptability to the Singlish context, achieving state-of-the-art performance and outperforming prior models by 10–30% in comparison with other AudioLLMs and cascaded solutions1
Bin Wang 0040, Xunlong Zou, Yingxu He, Zhuohan Liu, Chengwei Wei, Nancy F. Chen, AiTi Aw
ASRU8
2025 Obtaining objective labels and analysing annotator subjectivity by using a Rasch model for ordinal speech processing
abstract
In datasets for subjective tasks, disagreement between annotators is often accommodated by recording annotations from multiple annotators for each datapoint. However, it is more convenient to train and evaluate models against a scalar reference. In ordinal tasks, where outputs follow a monotonic order, the standard approach of computing the scalar reference as either the mean, median, or majority vote of the multiple annotations does not consider the differing bias between groups of annotators and assumes linearity of the output. This paper proposes to compute the scalar reference using a Rasch model. This expresses differing annotator bias, avoids linear assumptions, and allows control of the set of confounding variables that should influence the reference. Demonstrations on MSP-Podcast emotion recognition and speechocean762 spoken language assessment show how to use the Rasch model to compute a scalar reference, analyse confounders in the dataset, and compare multiple trained models.
Jeremy H. M. Wong, Nancy F. Chen
ASRU2
2025 Speech in-context learning of paralinguistic tasks
abstract
In-context learning adapts a large language model to a new task, without computationally expensive parameter updates. This has previously been demonstrated for text tasks, as well as speech tasks that rely primarily on lexical information, such as recognition and translation. This paper proposes to extend this investigation to consider the ability of current open-source models to exhibit in-context learning on speech tasks that require an understanding of paralinguistic information. The tasks of stutter detection, pronunciation assessment, and speech emotion recognition are investigated. The results suggest that current open-source models already exhibit some degree of speech in-context learning on paralinguistic tasks. To more fully utilise available adaptation data, it is also proposed to overcome the finite number of in-context exemplars allowed by the model’s prompt length limit, through ensemble combination over multiple in-context learning runs that each use different exemplars.
Jeremy H. M. Wong, Muhammad Huzaifah 0001, Nancy F. Chen, AiTi Aw
ASRU3
2025 Diversity and complementarity of speech encoders across diverse tasks in a multi-modal large language model
abstract
A Large Language Model (LLM) can be extended to understand speech inputs by using a speech encoder to compute embeddings from the speech, which are then used with a text prompt. Diverse information is expressed in speech and a wide variety of tasks can be performed. Different speech encoders may specialise toward different information types and tasks. This complementarity can be leveraged upon by using multiple speech encoders. This paper presents a comprehensive analysis of the diversity and complementarity between open-source speech encoders, when used in a multi-modal LLM framework. Experiments identify the encoders that excel in each type of downstream task, thereby guiding future system design. The diversity between encoders is measured, showing that Whisper tends to behave more differently. Diversity between encoders is compared across tasks, showing that semantic tasks tend to yield more diverse predictions. Early and late fusion show that complementarity can yield improvements.
Jeremy H. M. Wong, Muhammad Huzaifah 0001, Hardik B. Sailor, Kye Min Tan, Bin Wang 0040, Qiongqiong Wang, Xunlong Zou, Nancy F. Chen, AiTi Aw
ASRU10
2025 Aligning Large Language Models with Human Opinions through Persona Selection and Value-Belief-Norm Reasoning
abstract
Reasoning and predicting human opinions with large language models (LLMs) is essential yet challenging. Current methods employ role-playing with personae but face two major issues: LLMs are sensitive to even a single irrelevant persona, skewing predictions by up to 30%; and LLMs fail to reason strategically over personae. We propose Chain-of-Opinion (COO), a simple four-step solution modeling which and how to reason with personae, inspired by the Value–Belief–Norm (VBN) theory. COO differentiates between explicit personae (demographics and ideology) and implicit personae (historical opinions), involves: (1) filtering irrelevant attributes from explicit personae; (2) ranking implicit personae into a preferential list for selecting top-k; (3) applying novel VBN reasoning to extract user environmental and personal value, belief, and norm variables for accurate and reliable predictions; and (4) iterating VBN reasoning with progressively larger lists of implicit personae to handle potential persona insufficiency. COO efficiently achieves new state-of-the-art opinion prediction via prompting with only 5 inference calls, improving prior techniques by up to 4%. Notably, fine-tuning LMs with COO’s data results in significantly better opinion-aligned models, by up to 23%.
Do Xuan Long, Kenji Kawaguchi, Min-Yen Kan, Nancy F. Chen
COLING4
2025 DnA-Eval: Enhancing Large Language Model Evaluation through Decomposition and Aggregation
abstract
The acceleration of Large Language Models (LLMs) research has opened up new possibilities for evaluating generated text. Though LLMs serve as scalable and economical evaluators, how reliable these evaluators is still under-explored. Prior research efforts in the meta-evaluation of LLMs as judges limit the prompting of an LLM to a single use to obtain a final evaluation decision. They then compute the agreement between LLMs’ outputs and human labels. This lacks interpretability in understanding the evaluation capability of LLMs. In light of this challenge, we propose DnA-Eval, which breaks down the evaluation process into decomposition and aggregation stages based on pedagogical practices. Our experiments show that it not only provides a more interpretable window for how well LLMs evaluate, but also leads to improvements up to 39.6% for different LLMs on a variety of meta-evaluation benchmarks.
Minzhi Li, Zhengyuan Liu, Shumin Deng, Shafiq R. Joty, Nancy F. Chen, Min-Yen Kan
COLING5
2025 Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD
abstract
Large Language Models (LLMs) can struggle to balance gullibility to misinformation and resistance to valid corrections in persuasive dialogues, a critical challenge for reliable deployment.We introduce DuET-PD (Dual Evaluation for Trust in Persuasive Dialogues), a framework evaluating multi-turn stancechange dynamics across dual dimensions: persuasion type (corrective/misleading) and domain (knowledge via MMLU-Pro, and safety via SALAD-Bench).We find that even a stateof-the-art model like GPT-4o achieves only 27.32% accuracy in MMLU-Pro under sustained misleading persuasions.Moreover, results reveal a concerning trend of increasing sycophancy in newer open-source models.To address this, we introduce Holistic DPO, a training approach balancing positive and negative persuasion examples.Unlike prompting or resist-only training, Holistic DPO enhances both robustness to misinformation and receptiveness to corrections, improving Llama-3.1-8B-Instruct'saccuracy under misleading persuasion in safety contexts from 4.21% to 76.54%.These contributions offer a pathway to developing more reliable and adaptable LLMs for multi-turn dialogue.Code is available at https://github.com/Social-AI-Studio/DuET-PD.
Bryan Chen Zhengyu Tan, Daniel Wai Kit Chin, Zhengyuan Liu, Nancy F. Chen, Roy Ka-Wei Lee
EMNLP4
2025 Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
abstract
Current emotional text-to-speech (TTS) models pre-dominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending other emotion characteristics, which limits their capabilities of capturing the nuances between different emotions. We propose a controllable Emo-DPO approach, which employs direct preference optimization to differentiate subtle emotional nuances between emotions through optimizing towards preferred emotions over less preferred emotional ones. Instead of relying on traditional neural architectures used in existing emotional TTS models, we propose utilizing the emotion-aware LLM-TTS neural architecture to leverage LLMs’ in-context learning and instruction-following capabilities. Comprehensive experiments confirm that our proposed method outperforms the existing baselines.
Xiaoxue Gao, Chen Zhang 0020, Yiming Chen 0010, Huayun Zhang, Nancy F. Chen
ICASSP5
2025 MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
abstract
The rapid advancements in large language models (LLMs) have significantly enhanced natural language processing capabilities, facilitating the development of AudioLLMs that process and understand speech and audio inputs alongside text. Existing AudioLLMs typically combine a pre-trained audio encoder with a pre-trained LLM, which are subsequently finetuned on specific audio tasks. However, the pre-trained audio encoder has constrained capacity to capture features for new tasks and datasets. To address this, we propose to incorporate mixtures of ‘weak’ encoders (MoWE) into the AudioLLM framework. MoWE supplements a base encoder with a pool of relatively lightweight encoders, selectively activated based on the audio input to enhance feature extraction without significantly increasing model size. Our empirical results demonstrate that MoWE effectively improves multi-task performance, broadening the applicability of AudioLLMs to more diverse audio tasks.
Bin Wang 0040, Xunlong Zou, Zhuohan Liu, Yingxu He, Geyu Lin, Nancy F. Chen, AiTi Aw
ICASSP8
2025 Preference Optimization for Reasoning with Pseudo Feedback
abstract
Preference optimization techniques, such as Direct Preference Optimization (DPO), are frequently employed to enhance the reasoning capabilities of large language models (LLMs) in domains like mathematical reasoning and coding, typically following supervised fine-tuning. These methods rely on high-quality labels for reasoning tasks to generate preference pairs; however, the availability of reasoning datasets with human-verified labels is limited. In this study, we introduce a novel approach to generate pseudo feedback for reasoning tasks by framing the labeling of solutions to reason problems as an evaluation against associated \emph{test cases}. We explore two forms of pseudo feedback based on test cases: one generated by frontier LLMs and the other by extending self-consistency to multi-test-case. We conduct experiments on both mathematical reasoning and coding tasks using pseudo feedback for preference optimization, and observe improvements across both tasks. Specifically, using Mathstral-7B as our base model, we improve MATH results from 58.3 to 68.6, surpassing both NuminaMath-72B and GPT-4-Turbo-1106-preview. In GSM8K and College Math, our scores increase from 85.6 to 90.3 and from 34.3 to 42.3, respectively. Building on Deepseek-coder-7B-v1.5, we achieve a score of 24.3 on LiveCodeBench (from 21.1), surpassing Claude-3-Haiku.
Fangkai Jiao, Geyang Guo, Nancy F. Chen, Shafiq R. Joty, Furu Wei
ICLR4
2025 CCMI 2025: Cross-Cultural Multimodal Interaction
Koji Inoue, Shogo Okada, Divesh Lala, Sahba Zojaji, Nancy F. Chen, Tatsuya Kawahara
ICMI5
2025 Distilling a speech and music encoder with task arithmetic
Fabian Ritter Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei, Jeremy H. M. Wong, Chng Eng Siong, Nancy F. Chen, Hung-yi Lee
INTERSPEECH6
2025 Scaling Up Collaborative Dialogue Analysis: An AI-driven Approach to Understanding Dialogue Patterns in Computational Thinking Education
abstract
Pair programming is a collaborative activity that enhances students' computational thinking (CT) skills. Analyzing students' interactions during pair programming provides valuable insights into effective learning. However, interpreting classroom dialogues is a challenging and complex task. Due to the simultaneous interaction between interlocutors and other ambient noise in collaborative learning contexts, previous work heavily relied on manual transcription and coding, which is labor-intensive and time-consuming. Recent advancements in speech and language processing offer promising opportunities to automate and scale up dialogue analysis. Besides, previous work mainly focused on task-related interactions, with little attention to social interactions. To address these gaps, we conducted a four-week CT course with 26 fifth-grade primary school students. We recorded their discussions, transcribed them with speech processing models, and developed a coding scheme and applied LLMs for annotation. Our AI-driven pipeline effectively analyzed classroom recordings with high accuracy and efficiency. After identifying the dialogue patterns, we investigated the relationships between these patterns and CT performance. Four clusters of dialogue patterns have been identified: Inquiry, Constructive Collaboration, Disengagement, and Disputation. We observed that Inquiry and Constructive Collaboration patterns were positively related to students' CT skills, while Disengagement and Disputation patterns were associated with lower CT performance. This study contributes to the understanding of how dialogue patterns relate to CT performance and provides implications for both research and educational practice in CT learning.
Stella Xin Yin, Zhengyuan Liu, Dion Hoe-Lian Goh, Choon Lang Quek, Nancy F. Chen
LAK5
2025 LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs
abstract
Do Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, Min-Yen Kan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Do Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq R. Joty, Kenji Kawaguchi, Nancy F. Chen, Min-Yen Kan
NAACL (Long Papers)7
2025 AudioBench: A Universal Benchmark for Audio Large Language Models
abstract
Bin Wang, Xunlong Zou, Geyu Lin, Shuo Sun, Zhuohan Liu, Wenyu Zhang, Zhengyuan Liu, AiTi Aw, Nancy F. Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Bin Wang 0040, Xunlong Zou, Geyu Lin, Zhuohan Liu, Zhengyuan Liu, AiTi Aw, Nancy F. Chen
NAACL (Long Papers)9
2025 Transformer-based document-level discourse processing: Exploiting prior language knowledge and hierarchical parsing
Zhengyuan Liu, Ke Shi 0001, Nancy F. Chen
Comput. Speech Lang.3
2025 Prompt-Unseen-Emotion: Mixed Emotional Speech Synthesis With Prompt-LLM Contextual Knowledge
abstract
Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse emotional speech generation for more natural interactions. To bridge this gap, this paper proposes a novel prompt-unseen-emotion (PUE) approach to generate unseen emotional speech via emotion-guided prompt learning.PUEis trained utilizing an LLM-TTS architecture to ensure emotional consistency between categorical emotion-relevant prompts and emotional speech, allowing the model to quantitatively capture different emotion weightings per utterance. During inference, mixed emotional speech can be generated by flexibly adjusting emotion proportions and leveraging LLM contextual knowledge, enabling the model to quantify different emotional styles. Our proposedPUEsuccessfully facilitates expressive speech synthesis of unseen emotions in a zero-shot setting.
Xiaoxue Gao, Huayun Zhang, Nancy F. Chen
IEEE Signal Process. Lett.3
2025 PRESENT: Zero-Shot Text-to-Prosody Control
abstract
Current strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-to-speech (TTS) models, we present PRESENT (PRosody Editing without Style Embeddings or New Training), which exploits explicit prosody prediction in FastSpeech2-based models by modifying the inference process directly. We apply our text-to-prosody framework to zero-shot language transfer using a JETS model exclusively trained on English LJSpeech data. We obtain character error rates (CER) of 12.8%, 18.7% and 5.9% for German, Hungarian and Spanish respectively, beating the previous state-of-the-art CER by over 2× for all three languages. Furthermore, we allow subphoneme-level control, a first in this field. To evaluate its effectiveness, we show that PRESENT can improve the prosody of questions, and use it to generate Mandarin, a tonal language where vowel pitch varies at subphoneme level. We attain 25.3% hanzi CER and 13.0% pinyin CER with the JETS model. All our code and audio samples11https://github.com/iamanigeeit/presentandhttps://present2024.web.app/are available online.
Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman, Dorien Herremans
IEEE Signal Process. Lett.3
2025 Are Current Task-Oriented Dialogue Systems Able to Satisfy Impolite Users?
abstract
Task-oriented dialogue (TOD) systems play a critical role in assisting users with various tasks, such as ticket booking and service inquiries. While these systems have demonstrated significant potential in addressing customer needs, they typically assume that users will interact with the dialogue agent in a polite manner. This assumption, however, is often unrealistic, as users may express impatience or frustration through impolite behavior. Addressing this gap, this article investigates the impact of impolite user behavior on the performance of TOD systems. To this end, we developed a novel corpus of impolite dialogues and conducted comprehensive experiments to evaluate the performance of state-of-the-art TOD systems on this dataset. Our results reveal a notable limitation: existing TOD systems struggle to handle impolite user utterances effectively, leading to degraded performance. To mitigate this issue, we introduce a data augmentation approach designed to improve the systems’ ability to manage impolite dialogues. Although this method achieves measurable improvements, managing impolite user interactions remains a challenging research problem. By making our impolite dialogue corpus publicly accessible, we aim to encourage further research in this underexplored area. This study underscores the need for more robust TOD systems capable of handling diverse user behaviors, ultimately enhancing their applicability in real-world scenarios.
Nancy F. Chen, Roy Ka-Wei Lee
IEEE Trans. Comput. Soc. Syst.2
2024 Prompt Optimization via Adversarial In-Context Learning
abstract
Xuan Long Do, Yiran Zhao, Hannah Brown, Yuxi Xie, James Xu Zhao, Nancy F. Chen, Kenji Kawaguchi, Michael Shieh, Junxian He. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Do Xuan Long, Yiran Zhao 0006, Hannah Brown, Yuxi Xie, James Xu Zhao, Nancy F. Chen, Kenji Kawaguchi, Michael Shieh, Junxian He
ACL (1)6
2024 Discursive Socratic Questioning: Evaluating the Faithfulness of Language Models' Understanding of Discourse Relations
abstract
While large language models have significantly enhanced the effectiveness of discourse relation classifications, it remains unclear whether their comprehension is faithful and reliable.We provide DISQ, a new method for evaluating the faithfulness of understanding discourse based on question answering.We first employ incontext learning to annotate the reasoning for discourse comprehension, based on the connections among key events within the discourse.Following this, DISQ interrogates the model with a sequence of questions to assess its grasp of core event relations, its resilience to counterfactual queries, as well as its consistency to its previous responses.We then evaluate language models with different architectural designs using DISQ, finding: (1) DISQ presents a significant challenge for all models, with the top-performing GPT model attaining only 41% of the ideal performance in PDTB; (2) DISQ is robust to domain shifts and paraphrase variations; (3) Open-source models generally lag behind their closed-source GPT counterparts, with notable exceptions being those enhanced with chat and code/math features; (4) Our analysis validates the effectiveness of explicitly signalled discourse connectives, the role of contextual information, and the benefits of using historical QA data. Discourse relation: Contingency.Cause.ResultIs "they keep changing their prices" a reason for "it's very frustrating"?Model's Answer: True Is "they keep changing their prices" contrasted with "it's very frustrating"?Is "it's very frustrating" the result of "they keep changing their prices"?🤖 Targeted Score = 1 Counterfactual Score = 0 Consistency Score = 0Arg2: It's very frustrating.Arg1: When I want to buy, they run from you --they keep changing their prices.
Yisong Miao, Hongfu Liu 0002, Wenqiang Lei, Nancy F. Chen, Min-Yen Kan
ACL (1)4
2024 On Context Utilization in Summarization with Large Language Models
abstract
Large language models (LLMs) excel in abstractive summarization tasks, delivering fluent and pertinent summaries.Recent advancements have extended their capabilities to handle long-input contexts, exceeding 100k tokens.However, in question answering, language models exhibit uneven utilization of their input context.They tend to favor the initial and final segments, resulting in a U-shaped performance pattern concerning where the answer is located within the input.This bias raises concerns, particularly in summarization where crucial content may be dispersed throughout the source document(s).Besides, in summarization, mapping facts from the source to the summary is not trivial as salient content is usually re-phrased.In this paper, we conduct the first comprehensive study on context utilization and position bias in summarization.Our analysis encompasses 6 LLMs, 10 datasets, and 5 evaluation metrics.We introduce a new evaluation benchmark called MiddleSum on the which we benchmark two alternative inference methods to alleviate position bias: hierarchical summarization and incremental summarization 1 .Metric Model CNN/DM XSum Reddit SAMSum Multi-X AVG Arxiv PubMed GovReport SummScreenFD Multi-N AVG ROUGE-2 Flan-UL2 -0.296 -0.124 0.048 -0.069 -0.201 -0.128 _ _ _ _ _ _ Llama-2-7B -0.160 -0.023 0.063 -0.059 -0.100 -0.056 0.022 -0.113 -0.109 -0.079 -0.210 -0.098 Llama-2-13B -0.166 -0.086 0.031 -0.078 -0.039 -0.068 -0.017 -0.081 -0.166 -0.139 -0.213 -0.123 Xgen-7B -0.228 -0.042 0.066 -0.039 -0.041 -0.056 0.028 -0.091 -0.405 0.063 -0.283 -0.138 Mistral-7B -0.289 -0.031 0.006 -0.024 -0.052 -0.078 -0.270 -0.279 -0.585 -0.132 -0.324 -0.318 GPT-3.5 -0.323 -0.027 -0.031 -0.097 0.088 -0.078 0.026 -0.093 -0.123 -0.061 -0.233 -0.097 BERTScore Flan-UL2 -0.331 -0.185 0.062 -0.144 -0.399 -0.
Mathieu Ravaut, Aixin Sun, Nancy F. Chen, Shafiq R. Joty
ACL (1)3
2024 Granular Change Accuracy: A More Accurate Performance Metric for Dialogue State Tracking
abstract
Current metrics for evaluating Dialogue State Tracking (DST) systems exhibit three primary limitations. They: i) erroneously presume a uniform distribution of slots throughout the dialog, ii) neglect to assign partial scores for individual turns, iii) frequently overestimate or underestimate performance by repeatedly counting the models’ successful or failed predictions. To address these shortcomings, we introduce a novel metric: Granular Change Accuracy (GCA). GCA focuses on evaluating the predicted changes in dialogue state over the entire dialogue history. Benchmarking reveals that GCA effectively reduces biases arising from distribution uniformity and the positioning of errors across turns, resulting in a more precise evaluation. Notably, we find that these biases are particularly pronounced when evaluating few-shot or zero-shot trained models, becoming even more evident as the model’s error rate increases. Hence, GCA offers significant promise, particularly for assessing models trained with limited resources. Our GCA implementation is a useful addition to the pool of DST metrics.
Taha Aksu, Nancy F. Chen
LREC/COLING2
2024 LOCOST: State-Space Models for Long Document Abstractive Summarization
abstract
Florian Le Bronnec, Song Duong, Mathieu Ravaut, Alexandre Allauzen, Nancy Chen, Vincent Guigue, Alberto Lumbreras, Laure Soulier, Patrick Gallinari. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Florian Le Bronnec, Song Duong, Mathieu Ravaut, Alexandre Allauzen, Nancy F. Chen, Vincent Guigue, Alberto Lumbreras, Laure Soulier, Patrick Gallinari
EACL (1)5
2024 Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing
abstract
Large Language Models (LLMs) have demonstrated significant potential in handling complex reasoning tasks through step-by-step rationale generation.However, recent studies have raised concerns regarding the hallucination and flaws in their reasoning process.Substantial efforts are being made to improve the reliability and faithfulness of the generated rationales.Some approaches model reasoning as planning, while others focus on annotating for process supervision.Nevertheless, the planning-based search process often results in high latency due to the frequent assessment of intermediate reasoning states and the extensive exploration space.Additionally, supervising the reasoning process with human annotation is costly and challenging to scale for LLM training.To address these issues, in this paper, we propose a framework to learn planning-based reasoning through Direct Preference Optimization (DPO) on collected trajectories, which are ranked according to our synthesized process rewards.Our results on challenging logical reasoning benchmarks demonstrate the effectiveness of our learning framework, showing that our 7B model can surpass the strong counterparts like GPT-3.5-Turbo.
Fangkai Jiao, Chengwei Qin, Zhengyuan Liu, Nancy F. Chen, Shafiq R. Joty
EMNLP4
2024 Personality-aware Student Simulation for Conversational Intelligent Tutoring Systems
abstract
Intelligent Tutoring Systems (ITSs) can provide personalized and self-paced learning experience.The emergence of large language models (LLMs) further enables better humanmachine interaction, and facilitates the development of conversational ITSs in various disciplines such as math and language learning.In dialogic teaching, recognizing and adapting to individual characteristics can significantly enhance student engagement and learning efficiency.However, characterizing and simulating student's persona remain challenging in training and evaluating conversational ITSs.In this work, we propose a framework to construct profiles of different student groups by refining and integrating both cognitive and noncognitive aspects, and leverage LLMs for personalityaware student simulation in a language learning scenario.We further enhance the framework with multi-aspect validation, and conduct extensive analysis from both teacher and student perspectives.Our experimental results show that state-of-the-art LLMs can produce diverse student responses according to the given language ability and personality traits, and trigger teacher's adaptive scaffolding strategies.
Zhengyuan Liu, Stella Xin Yin, Geyu Lin, Nancy F. Chen
EMNLP4
2024 Multi-expert Prompting Improves Reliability, Safety and Usefulness of Large Language Models
abstract
We present Multi-expert Prompting 1 , a novel enhancement of ExpertPrompting (Xu et al., 2023), designed to improve the large language model (LLM) generation.Specifically, it guides an LLM to fulfill an input instruction by simulating multiple experts, aggregating their responses, and selecting the best among individual and aggregated responses.This process is performed in a single chain of thoughts through our seven carefully designed subtasks derived from the Nominal Group Technique (Ven and Delbecq, 1974), a well-established decision-making framework.Our evaluations demonstrate that Multi-expert Prompting significantly outperforms ExpertPrompting and comparable baselines in enhancing the truthfulness, factuality, informativeness, and usefulness of responses while reducing toxicity and hurtfulness.It further achieves state-ofthe-art truthfulness by outperforming the best baseline by 8.69% with ChatGPT.Multi-expert Prompting is efficient, explainable, and highly adaptable to diverse scenarios, eliminating the need for manual prompt construction. * Equal contribution.1 Our codes and data will be made publicly available here.
Do Xuan Long, Duong Ngoc Yen, Anh Tuan Luu, Kenji Kawaguchi, Min-Yen Kan, Nancy F. Chen
EMNLP6
2024 Developmental Predictive Coding Model for Early Infancy Mono and Bilingual Vocal Continual Learning
Alex Pitti, Mathias Quoy, Nancy F. Chen
ICANN (7)4
2024 Distilling Distributional Uncertainty from a Gaussian Process
abstract
A Neural Network (NN) may exhibit overconfidence about wrong hypotheses, especially for Out-Of-Domain (OOD) inputs. A Gaussian process (GP) instead has an explainable distributional uncertainty behaviour, by predicting hypotheses with greater uncertainty for query inputs further from the training data. Previous work has shown that a NN can learn to emulate the behaviour of a GP on in-domain data. This paper expands upon this, by proposing to train a NN student to emulate the GP teacher’s distributional uncertainty behaviour on OOD data. This avoids the computational cost of using a GP at run-time, while improving the OOD confidence calibration of a NN. More accurate confidence calibration may better inform how the system should feedback to the user. Experiments on the SEP-28k-E stutter detection dataset suggest that distillation of such knowledge is feasible between these models.
Jeremy H. M. Wong, Nancy F. Chen
ICASSP2
2024 Dataset-Distillation Generative Model for Speech Emotion Recognition
Fabian Ritter Gutierrez, Kuan-Po Huang, Jeremy H. M. Wong, Dianwen Ng, Hung-yi Lee, Nancy F. Chen, Chng Eng Siong
INTERSPEECH6
2024 Exploring Self-supervised Logic-enhanced Training for Large Language Models
abstract
Fangkai Jiao, Zhiyang Teng, Bosheng Ding, Zhengyuan Liu, Nancy Chen, Shafiq Joty. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Fangkai Jiao, Zhiyang Teng, Bosheng Ding, Zhengyuan Liu, Nancy F. Chen, Shafiq R. Joty
NAACL-HLT5
2024 SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning
abstract
Bin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Yang Ding, AiTi Aw, Nancy Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Bin Wang 0040, Zhengyuan Liu, Fangkai Jiao, AiTi Aw, Nancy F. Chen
NAACL-HLT7
2024 Optimizing Code-Switching in Conversational Tutoring Systems: A Pedagogical Framework and Evaluation
abstract
Large language models demonstrate remarkable proficiency in various tasks across multiple languages.However, their potential in codeswitching remains underexplored, particularly in cultural and educational contexts.Codeswitching or translanguaging plays a crucial role in bilingual education, facilitating comprehension and engagement among students with varied language proficiency.In this work, we present a pedagogy-inspired framework that introduces traditional classroom practices of code-switching to intelligent tutoring systems.Specifically, we develop fine-grained instructional strategies tailored to multilingual and educational needs.We conduct experiments involving both LLM-based evaluation and expert analysis to assess the effectiveness of translanguaging in tutoring dialogues.Our experimental results indicate that strategic code-switching can significantly enhance the learning experience.This work not only advances dialogic tutors in language learning but also extends LLMs to better accommodate multilingual interaction.
Zhengyuan Liu, Stella Xin Yin, Nancy F. Chen
SIGDIAL3
2024 Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
abstract
Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models such as Mamba has performed well on long-sequence modeling in natural language processing and computer vision tasks. However, research endeavors in speech technology tasks has been under-explored. We propose Speech-Mamba, which incorporates selective state space modeling in Transformer neural architectures. Long sequence representations with selective state space models in Speech-Mamba is complemented with lower-level representations from Transformer-based modeling. Speech-mamba achieves better capacity to model long-range dependencies, as it scales near-linearly with sequence length.
Xiaoxue Gao, Nancy F. Chen
SLT2
2024 Semi-Supervised Learning for Robust Speech Evaluation
abstract
Speech evaluation measures a learner’s oral proficiency using automatic models. Corpora for training such models often pose sparsity challenges given that there often is limited scored data from teachers, in addition to the score distribution across proficiency levels being often imbalanced among student cohorts. Automatic scoring is thus not robust when faced with under-represented samples or out-of-distribution samples, which inevitably exist in real-world deployment scenarios. This paper proposes to address such challenges by exploiting semi-supervised pre-training and objective regularization to approximate subjective evaluation criteria. In particular, normalized mutual information is used to quantify the speech characteristics from the learner and the reference. An anchor model is trained using pseudo labels to predict the correctness of pronunciation. An interpolated loss function is proposed to minimize not only the prediction error with respect to ground-truth scores but also the divergence between two probability distributions estimated by the speech evaluation model and the anchor model. Compared to other state-of-the-art methods on a public data-set, this approach not only achieves high performance while evaluating the entire test-set as a whole, but also brings the most evenly distributed prediction error across distinct proficiency levels. Furthermore, empirical results show the model accuracy on out-of-distribution data also compares favorably with competitive baselines.
Huayun Zhang, Jeremy H. M. Wong, Geyu Lin, Nancy F. Chen
SLT4
2024 SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
abstract
Text-to-speech (TTS) models have achieved remarkable naturalness in recent years, yet like most deep neural models, they have more parameters than necessary. Sparse TTS models can improve on dense models via pruning and extra retraining, or converge faster than dense models with some performance loss. Thus, we propose training TTS models using decaying sparsity, i.e. a high initial sparsity to accelerate training first, followed by a progressive rate reduction to obtain better eventual performance. This decremental approach differs from current methods of incrementing sparsity to a desired target, which costs significantly more time than dense training. We call our method SNIPER training: Single-shot Initialization Pruning Evolving-Rate training. Our experiments on FastSpeech2 show that we were able to obtain better losses in the first few training epochs with SNIPER, and that the final SNIPER-trained models outperformed constant-sparsity models and edged out dense models, with negligible difference in training time. Our code is available on Github11https://github.com/iamanigeeit/sniper.
Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman, Dorien Herremans
TENCON3
2023 Prompter: Zero-shot Adaptive Prefixes for Dialogue State Tracking Domain Adaptation
abstract
A challenge in the Dialogue State Tracking (DST) field is adapting models to new domains without using any supervised data -zero-shot domain adaptation.Parameter-Efficient Transfer Learning (PETL) has the potential to address this problem due to its robustness.However, it has yet to be applied to the zero-shot scenarios, as it is not clear how to apply it unsupervisedly.Our method, Prompter, uses descriptions of target domain slots to generate dynamic prefixes that are concatenated to the key and values at each layer's self-attention mechanism.This allows for the use of prefix-tuning in zeroshot.Prompter outperforms previous methods on both the MultiWOZ and SGD benchmarks.In generating prefixes, our analyses find that Prompter not only utilizes the semantics of slot descriptions but also how often the slots appear together in conversation.Moreover, Prompter's gains are due to its improved ability to distinguish "none"-valued dialogue slots, compared against baselines.
Taha Aksu, Min-Yen Kan, Nancy F. Chen
ACL (1)3
2023 Modeling What-to-ask and How-to-ask for Answer-unaware Conversational Question Generation
abstract
Xuan Long Do, Bowei Zou, Shafiq Joty, Tran Tai, Liangming Pan, Nancy Chen, Ai Ti Aw. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Do Xuan Long, Bowei Zou, Shafiq R. Joty, Anh Tran Tai, Liangming Pan, Nancy F. Chen, AiTi Aw
ACL (1)6
2023 Guiding Computational Stance Detection with Expanded Stance Triangle Framework
abstract
Stance detection determines whether the author of a piece of text is in favor of, against, or neutral towards a specified target, and can be used to gain valuable insights into social media.The ubiquitous indirect referral of targets makes this task challenging, as it requires computational solutions to model semantic features and infer the corresponding implications from a literal statement.Moreover, the limited amount of available training data leads to subpar performance in out-of-domain and cross-target scenarios, as data-driven approaches are prone to rely on superficial and domain-specific features.In this work, we decompose the stance detection task from a linguistic perspective, and investigate key components and inference paths in this task.The stance triangle is a generic linguistic framework previously proposed to describe the fundamental ways people express their stance.We further expand it by characterizing the relationship between explicit and implicit objects.We then use the framework to extend one single training corpus with additional annotation.Experimental results show that strategically-enriched data can significantly improve the performance on out-of-domain and cross-target evaluation.
Zhengyuan Liu, Yong Keong Yap, Hai Leong Chieu, Nancy F. Chen
ACL (1)4
2023 Variational Gaussian Process Data Uncertainty
abstract
A Gaussian process (GP) computes an explainable distributional uncertainty by hypothesising higher output uncertainty for query inputs far from the training inputs. However, a GP may not capture data uncertainty well. Accurate data uncertainty estimation may be important for subjective tasks, such as Spoken Language Assessment (SLA), where human expert raters may disagree on the output scores. This paper shows that a variational approximation of a GP has capacity to learn data uncertainty from the training data. However, standard training criteria tune only a scalar noise hyper-parameter toward the standard deviation of the output reference, thereby limiting the learning of this uncertainty. A training criterion is proposed to explicitly encourage the GP posterior to emulate the distribution of scores from multiple raters. Experiments on the speechocean762 SLA task show that this allows the GP to better express data uncertainty and improves the modelling of inter-rater disagreements.
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
ASRU3
2023 CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
abstract
Annotated data plays a critical role in Natural Language Processing (NLP) in training models and evaluating their performance.Given recent developments in Large Language Models (LLMs), models such as ChatGPT demonstrate zero-shot capability on many text-annotation tasks, comparable with or even exceeding human annotators.Such LLMs can serve as alternatives for manual annotation, due to lower costs and higher scalability.However, limited work has leveraged LLMs as complementary annotators, nor explored how annotation work is best allocated among humans and LLMs to achieve both quality and cost objectives.We propose CoAnnotating, a novel paradigm for Human-LLM co-annotation of unstructured texts at scale.Under this framework, we utilize uncertainty to estimate LLMs' annotation capability.Our empirical study shows CoAnnotating to be an effective means to allocate work from results on different datasets, with up to 21% performance improvement over random baseline.For code implementation, see https: //github.com/SALT-NLP/CoAnnotating.
Minzhi Li, Taiwei Shi, Caleb Ziems, Min-Yen Kan, Nancy F. Chen, Zhengyuan Liu, Diyi Yang
EMNLP5
2023 Instructive Dialogue Summarization with Query Aggregations
abstract
Conventional dialogue summarization methods directly generate summaries and do not consider user's specific interests.This poses challenges in cases where the users are more focused on particular topics or aspects.With the advancement of instruction-finetuned language models, we introduce instruction-tuning to dialogues to expand the capability set of dialogue summarization models.To overcome the scarcity of instructive dialogue summarization data, we propose a three-step approach to synthesize high-quality query-based summarization triples.This process involves summaryanchored query generation, query filtering and query-based summary generation.By training a unified model called InstructDS (Instructive Dialogue Summarization) on three summarization datasets with multi-purpose instructive triples, we expand the capability of dialogue summarization models.We evaluate our method on four datasets, including dialogue summarization and dialogue reading comprehension.Experimental results show that our approach outperforms the state-of-the-art models and even models with larger sizes.Additionally, our model exhibits higher generalizability and faithfulness, as confirmed by human subjective evaluations.Benjamin: Hey guys, what are we doing with the keys today?Hilary: I've got them.Whoever wants
Bin Wang 0040, Zhengyuan Liu, Nancy F. Chen
EMNLP3
2023 Picking the Underused Heads: A Network Pruning Perspective of Attention Head Selection for Fusing Dialogue Coreference Information
abstract
The Transformer-based models with the multi-head self-attention mechanism are widely used in natural language processing, and provide state-of-the-art results. While the pre-trained language backbones are shown to implicitly capture certain linguistic knowledge, explicitly incorporating structure-aware features can bring about further improvement on the downstream tasks. However, such enhancement often requires additional neural components and increases training parameter size. In this work, we investigate the attention head selection and manipulation strategy for feature injection from a network pruning perspective, and conduct a case study on dialogue summarization. We first rank attention heads in a Transformer-based summarizer with layer-wise importance. We then select the underused heads through extensive analysis, and inject structure-aware features by manipulating the selected heads. Experimental results show that the importance-based head selection is effective for feature injection, and dialogue summarization can be improved by incorporating coreference information via head manipulation.
Zhengyuan Liu, Nancy F. Chen
ICASSP2
2023 Distilling knowledge from Gaussian process teacher to neural network student
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
INTERSPEECH3
2023 C3: Compositional Counterfactual Contrastive Learning for Video-grounded Dialogues
abstract
Video-grounded dialogue systems aim to integrate video understanding and dialogue understanding to generate responses that are relevant to both the dialogue and video context.Most existing approaches employ deep learning models and have achieved remarkable performance, given the relatively small datasets available.However, the results are partially accomplished by exploiting biases in the datasets rather than developing multimodal reasoning, resulting in limited generalization.In this paper, we propose a novel approach of Compositional Counterfactual Contrastive Learning (C 3 ) to develop contrastive training between factual and counterfactual samples in videogrounded dialogues.Specifically, we design factual/counterfactual samples based on the temporal steps in videos and tokens in dialogues and propose contrastive loss functions that exploit object-level or action-level variance.Different from prior approaches, we focus on contrastive hidden state representations among compositional output tokens to optimize the representation space in a generation setting.We achieved promising performance gains on the Audio-Visual Scene-Aware Dialogues (AVSD) benchmark and showed the benefits of our approach in grounding video and dialogue context.
Hung Le 0003, Nancy F. Chen, Steven C. H. Hoi
SIGDIAL2
2023 Modelling Inter-Rater Uncertainty in Spoken Language Assessment
abstract
In a subjective task, such as Spoken Language Assessment (SLA), the reference scores provided by different human raters may vary. A collection of annotated scores from multiple raters can be interpreted as an expression of data uncertainty. Previous studies often treat SLA as classification or regression tasks, and train and evaluate models against scalar reference scores that were computed from the multiple rater scores, for example by majority voting. However, a scalar representation may not adequately capture information about uncertainty that is expressed by the multiple rater scores. This paper proposes to reformulate this subjective task as a distribution fitting problem, where the model should aim to emulate the uncertainty expressed by the multiple raters. Toward this aim, the model is trained and evaluated by computing a distance between the model's output posterior and the distribution of reference scores from the multiple raters. Different methods to infer a scalar score from the model's output posterior are also considered. This paper also proposes to improve the match between the model and the SLA task, by interpreting the model's outputs as parameters of a beta density function, to capture both uncertainty and score monotonicity. Finally, ensemble combination is investigated and a novel combination method is proposed, to marginalise out model uncertainty from the combined output distribution. These approaches are evaluated on the speechocean762 dataset and an in-house Tamil dataset.
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization
abstract
Sequence-to-sequence neural networks have recently achieved great success in abstractive summarization, especially through fine-tuning large pre-trained language models on the downstream dataset.These models are typically decoded with beam search to generate a unique summary.However, the search space is very large, and with the exposure bias, such decoding is not optimal.In this paper, we show that it is possible to directly train a secondstage model performing re-ranking on a set of summary candidates.Our mixture-of-experts SummaReranker learns to select a better candidate and consistently improves the performance of the base model.With a base PEGASUS, we push ROUGE scores by 5.44% on CNN-DailyMail (47.16 ROUGE-1), 1.31% on XSum (48.12 ROUGE-1) and 9.34% on Reddit TIFU (29.83 ROUGE-1), reaching a new state-of-theart.Our code and checkpoints will be available at https://github.com/ntunlp/ SummaReranker.
Mathieu Ravaut, Shafiq R. Joty, Nancy F. Chen
ACL (1)3
2022 CoHS-CQG: Context and History Selection for Conversational Question Generation
abstract
Conversational question generation (CQG) serves as a vital task for machines to assist humans, such as interactive reading comprehension, through conversations. Compared to traditional single-turn question generation (SQG), CQG is more challenging in the sense that the generated question is required not only to be meaningful, but also to align with the provided conversation. Previous studies mainly focus on how to model the flow and alignment of the conversation, but do not thoroughly study which parts of the context and history are necessary for the model. We believe that shortening the context and history is crucial as it can help the model to optimise more on the conversational alignment property. To this end, we propose CoHS-CQG, a two-stage CQG framework, which adopts a novel CoHS module to shorten the context and history of the input. In particular, it selects the top-p sentences and history turns by calculating the relevance scores of them. Our model achieves state-of-the-art performances on CoQA in both the answer-aware and answer-unaware settings.
Do Xuan Long, Bowei Zou, Liangming Pan, Nancy F. Chen, Shafiq R. Joty, AiTi Aw
COLING4
2022 Singlish Message Paraphrasing: A Joint Task of Creole Translation and Text Normalization
abstract
Within the natural language processing community, English is by far the most resource-rich language. There is emerging interest in conducting translation via computational approaches to conform its dialects or creole languages back to standard English. This computational approach paves the way to leverage generic English language backbones, which are beneficial for various downstream tasks. However, in practical online communication scenarios, the use of language varieties is often accompanied by noisy user-generated content, making this translation task more challenging. In this work, we introduce a joint paraphrasing task of creole translation and text normalization of Singlish messages, which can shed light on how to process other language varieties and dialects. We formulate the task in three different linguistic dimensions: lexical level normalization, syntactic level editing, and semantic level rewriting. We build an annotated dataset of Singlish-to-Standard English messages, and report performance on a perturbation-resilient sequence-to-sequence model. Experimental results show that the model produces reasonable generation results, and can improve the performance of downstream tasks like stance detection.
Zhengyuan Liu, Shikang Ni, AiTi Aw, Nancy F. Chen
COLING4
2022 Towards Summary Candidates Fusion
abstract
Sequence-to-sequence deep neural models fine-tuned for abstractive summarization can achieve great performance on datasets with enough human annotations.Yet, it has been shown that they have not reached their full potential, with a wide gap between the top beam search output and the oracle beam.Recently, re-ranking methods have been proposed, to learn to select a better summary candidate.However, such methods are limited by the summary quality aspects captured by the first-stage candidates.To bypass this limitation, we propose a new paradigm in secondstage abstractive summarization called Sum-maFusion that fuses several summary candiates to produce a novel abstractive secondstage summary.Our method works well on several summarization datasets, improving both the ROUGE scores and qualitative properties of fused summaries.It is especially good when the candidates to fuse are worse, such as in the few-shot setup where we set a new state-of-the-art.We will make our code and checkpoints available at https: //github.com/ntunlp/SummaFusion/.
Mathieu Ravaut, Shafiq R. Joty, Nancy F. Chen
EMNLP3
2022 Progressive Continual Learning for Spoken Keyword Spotting
abstract
Catastrophic forgetting is a thorny challenge when updating keyword spotting (KWS) models after deployment. To tackle such challenges, we propose a progressive continual learning strategy for small-footprint spoken keyword spotting (PCL-KWS). Specifically, the proposed PCL-KWS framework introduces a network instantiator to generate the task-specific sub-networks for remembering previously learned keywords. As a result, the PCL-KWS approach incrementally learns new keywords without forgetting prior knowledge. Besides, the proposed keyword-aware network scaling mechanism of PCL-KWS constrains the growth of model parameters while achieving high performance. Experimental results show that after learning five new tasks sequentially, our proposed PCL-KWS approach archives the new state-of-the-art performance of 92.8% average accuracy for all the tasks on Google Speech Command dataset compared with other baselines.
Yizheng Huang 0001, Nana Hou, Nancy F. Chen
ICASSP3
2022 Incremental Context Aware Attentive Knowledge Tracing
abstract
Knowledge Tracing is the prediction of the future performance of a learner, given the past performance. The existing knowledge tracing models represent the training data and does not generalize when there is a drift in the data distribution. We first empirically demonstrate an evolving Knowledge Tracing (eKT) scenario with distinct distribution of learner performances and diversity of questions from similar concepts. Next, we empirically characterize drift in the data and propose a task agnostic incremental context aware attentive knowledge tracing (iAKT) approach to learn incrementally from the eKT. The iAKT regularizes representations to learn from diverse learner performance distributions. Finally, we evaluate the ability of the proposed iAKT for knowledge tracing and study the effect of various regularization strategies on ranking difficulty of questions, using the ASSISTments 2017 data set. Performance results show that the iAKT adapts its representations to drift in data characteristics, while iAKT with EWC regularizer is better at ranking the difficulty of questions.
Cheryl Sze Yin Wong, Guo Yang, Nancy F. Chen, Savitha Ramasamy
ICASSP3
2022 EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models
abstract
Neural models are known to be over-parameterized, and recent work has shown that sparse text-to-speech (TTS) models can outperform dense models. Although a plethora of sparse methods has been proposed for other domains, such methods have rarely been applied in TTS. In this work, we seek to answer the question: what are the characteristics of selected sparse techniques on the performance and model complexity? We compare a Tacotron2 baseline and the results of applying five techniques. We then evaluate the performance via the factors of naturalness, intelligibility and prosody, while reporting model size and training time. Complementary to prior research, we find that pruning before or during training can achieve similar performance to pruning after training and can be trained much faster, while removing entire neurons degrades performance much more than removing parameters. To our best knowledge, this is the first work that compares sparsity paradigms in text-to-speech synthesis.
Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman
INTERSPEECH3
2022 Dynamic Sliding Window Modeling for Abstractive Meeting Summarization
abstract
International audience
Zhengyuan Liu, Nancy F. Chen
INTERSPEECH2
2022 Variations of multi-task learning for spoken language assessment
Jeremy H. M. Wong, Huayun Zhang, Nancy F. Chen
INTERSPEECH3
2022 VGNMN: Video-grounded Neural Module Networks for Video-Grounded Dialogue Systems
abstract
Neural module networks (NMN) have achieved success in image-grounded tasks such as Visual Question Answering (VQA) on synthetic images.However, very limited work on NMN has been studied in the video-grounded dialogue tasks.These tasks extend the complexity of traditional visual tasks with the additional visual temporal variance and language cross-turn dependencies.Motivated by recent NMN approaches on image-grounded tasks, we introduce Videogrounded Neural Module Network (VGNMN) to model the information retrieval process in video-grounded language tasks as a pipeline of neural modules.VGNMN first decomposes all language components in dialogues to explicitly resolve any entity references and detect corresponding action-based inputs from the question.The detected entities and actions are used as parameters to instantiate neural module networks and extract visual cues from the video.Our experiments show that VGNMN can achieve promising performance on a challenging video-grounded dialogue benchmark as well as a video QA benchmark.
Hung Le 0003, Nancy F. Chen, Steven C. H. Hoi
NAACL-HLT2
2022 Multimodal Dialogue State Tracking
abstract
Designed for tracking user goals in dialogues, a dialogue state tracker is an essential component in a dialogue system.However, the research of dialogue state tracking has largely been limited to unimodality, in which slots and slot values are limited by knowledge domains (e.g.restaurant domain with slots of restaurant name and price range) and are defined by specific database schema.In this paper, we propose to extend the definition of dialogue state tracking to multimodality.Specifically, we introduce a novel dialogue state tracking task to track the information of visual objects that are mentioned in video-grounded dialogues.Each new dialogue utterance may introduce a new video segment, new visual objects, or new object attributes and a state tracker is required to update these information slots accordingly.We created a new synthetic benchmark and designed a novel baseline, Video-Dialogue Transformer Network (VDTN), for this task.VDTN combines both object-level features and segment-level features and learns contextual dependencies between videos and dialogues to generate multimodal dialogue states.We optimized VDTN for a state generation task as well as a self-supervised video understanding task which recovers video segment or object representations.Finally, we trained VDTN to use the decoded states in a response prediction task.Together with comprehensive ablation and qualitative analysis, we discovered interesting insights towards building more capable multimodal dialogue systems.
Hung Le 0003, Nancy F. Chen, Steven C. H. Hoi
NAACL-HLT2
2022 Entity-based De-noising Modeling for Controllable Dialogue Summarization
abstract
Although fine-tuning pre-trained backbones produces fluent and grammatically-correct text in various language generation tasks, factual consistency in abstractive summarization remains challenging.This challenge is especially thorny for dialogue summarization, where neural models often make inaccurate associations between personal named entities and their respective actions.To tackle this type of hallucination, we present an entity-based de-noising model via text perturbation on reference summaries.We then apply this proposed approach in beam search validation, conditional training augmentation, and inference post-editing.Experimental results on the SAMSum corpus show that state-of-the-art models equipped with our proposed method achieve generation quality improvement in both automatic evaluation and human assessment.
Zhengyuan Liu, Nancy F. Chen
SIGDIAL2
2022 A transformer-Based neural language model that synthesizes brain activation maps from free-form text queries
Hoang Gia Ngo, Minh Nguyen 0002, Nancy F. Chen, Mert R. Sabuncu
Medical Image Anal.3
2021 Have We Solved The Hard Problem? It's Not Easy! Contextual Lexical Contrast as a Means to Probe Neural Coherence
abstract
Lexical cohesion is a fundamental mechanism for text which requires a pair of words to be interpreted as a certain type of lexical relation (e.g., similarity) to understand a coherent context; we refer to such relations as the contextual lexical relation. However, work on lexical cohesion has not modeled context comprehensively in considering lexical relations due to the lack of linguistic resources. In this paper, we take initial steps to address contextual lexical relations by focusing on the contrast relation, as it is a well-known relation though it is more subtle and relatively less resourced. We present a corpus named Cont 2 Lex to make Contextual Lexical Contrast Recognition a computationally feasible task. We benchmark this task with widely-adopted semantic representations; we discover that contextual embeddings (e.g. BERT) generally outperform static embeddings (e.g. Glove), but barely go beyond 70% in accuracy performance. In addition, we find that all embeddings perform better when CLC occurs within the same sentence, suggesting possible limitations of current computational coherence models. Another intriguing discovery is the improvement of BERT in CLC is largely attributed to its modeling of CLC word pairs co-occurring with other word repetitions. Such observations imply that the progress made in lexical coherence modeling remains relatively primitive even for semantic representations such as BERT that have been empowering numerous standard NLP tasks to approach human benchmarks. Through presenting our corpus and benchmark, we attempt to seed initial discussions and endeavors in advancing semantic representations from modeling syntactic and semantic levels to coherence and discourse levels.
Wenqiang Lei, Yisong Miao, Runpeng Xie, Bonnie L. Webber, Meichun Liu, Tat-Seng Chua, Nancy F. Chen
AAAI7
2021 Controllable Neural Dialogue Summarization with Personal Named Entity Planning
abstract
In this paper, we propose a controllable neural generation framework that can flexibly guide dialogue summarization with personal named entity planning.The conditional sequences are modulated to decide what types of information or what perspective to focus on when forming summaries to tackle the under-constrained problem in summarization tasks.This framework supports two types of use cases: (1) Comprehensive Perspective, which is a generalpurpose case with no user-preference specified, considering summary points from all conversational interlocutors and all mentioned persons; (2) Focus Perspective, positioning the summary based on a user-specified personal named entity, which could be one of the interlocutors or one of the persons mentioned in the conversation.During training, we exploit occurrence planning of personal named entities and coreference information to improve temporal coherence and to minimize hallucination in neural generation.Experimental results show that our proposed framework generates fluent and factually consistent summaries under various planning controls using both objective metrics and human evaluations.
Zhengyuan Liu, Nancy F. Chen
EMNLP (1)2
2021 Senone-Aware Adversarial Multi-Task Training for Unsupervised Child to Adult Speech Adaptation
abstract
Acoustic modeling for child speech is challenging due to the high acoustic variability caused by physiological differences in the vocal tract. The dearth of publicly available datasets makes the task more challenging. In this work, we propose a feature adaptation approach by exploiting adversarial multi-task training to minimize acoustic mismatch at the senone (tied triphone states) level between adult and child speech and leverage large amounts of transcribed adult speech. We validate the proposed method on three tasks: child speech recognition, child pronunciation assessment and child fluency score prediction. Empirical results indicate that our proposed approach consistently outperforms competitive baselines, achieving 7.7% relative error reduction on speech recognition and up to 25.2% relative gains on the evaluation tasks.
Richeng Duan, Nancy F. Chen
ICASSP2
2021 Learning Reasoning Paths over Semantic Graphs for Video-grounded Dialogues
Hung Le 0003, Nancy F. Chen, Steven C. H. Hoi
ICLR2
2021 WittyKiddy: Multilingual Spoken Language Learning for Kids
Ke Shi 0001, Kye Min Tan, Huayun Zhang, Siti Umairah Md. Salleh, Shikang Ni, Nancy F. Chen
Interspeech6
2021 Multilingual Speech Evaluation: Case Studies on English, Malay and Tamil
abstract
Speech evaluation is an essential component in computerassisted language learning (CALL).While speech evaluation on English has been popular, automatic speech scoring on low resource languages remains challenging.Work in this area has focused on monolingual specific designs and handcrafted features stemming from resource-rich languages like English.Such approaches are often difficult to generalize to other languages, especially if we also want to consider suprasegmental qualities such as rhythm.In this work, we examine three different languages that possess distinct rhythm patterns: English (stresstimed), Malay (syllable-timed), and Tamil (mora-timed).We exploit robust feature representations inspired by music processing and vector representation learning.Empirical validations show consistent gains for all three languages when predicting pronunciation, rhythm and intonation performance.
Huayun Zhang, Ke Shi 0001, Nancy F. Chen
Interspeech3
2021 Text2Brain: Synthesis of Brain Activation Maps from Free-Form Text Query
Hoang Gia Ngo, Minh Nguyen 0002, Nancy F. Chen, Mert R. Sabuncu
MICCAI (7)3
2021 Velocidapter: Task-oriented Dialogue Comprehension Modeling Pairing Synthetic Text Generation with Domain Adaptation
abstract
We introduce a synthetic dialogue generation framework, Velocidapter, which addresses the corpus availability problem for dialogue comprehension.Velocidapter augments datasets by simulating synthetic conversations for a task-oriented dialogue domain, requiring a small amount of bootstrapping work for each new domain.We evaluate the efficacy of our framework on a task-oriented dialogue comprehension dataset, MRCWOZ, which we curate by annotating questions for slots in the restaurant, taxi, and hotel domains of the Mul-tiWOZ 2.2 dataset (Zang et al., 2020).We run experiments within a low-resource setting, where we pretrain a model on SQuAD, fine-tuning it on either a small original data or on the synthetic data generated by our framework.Velocidapter shows significant improvements using both the transformer-based BERT-Base and BiDAF as base models.We further show that the framework is easy to use by novice users and conclude that Velocidapter can greatly help training over task-oriented dialogues, especially for low-resourced emerging domains.
Taha Aksu, Zhengyuan Liu, Min-Yen Kan, Nancy F. Chen
SIGDIAL4
2021 Coreference-Aware Dialogue Summarization
abstract
Summarizing conversations via neural approaches has been gaining research traction lately, yet it is still challenging to obtain practical solutions.Examples of such challenges include unstructured information exchange in dialogues, informal interactions between speakers, and dynamic role changes of speakers as the dialogue evolves.Many of such challenges result in complex coreference links.Therefore, in this work, we investigate different approaches to explicitly incorporate coreference information in neural abstractive dialogue summarization models to tackle the aforementioned challenges.Experimental results show that the proposed approaches achieve state-of-the-art performance, implying it is useful to utilize coreference information in dialogue summarization.Evaluation results on factual correctness suggest such coreferenceaware models are better at tracing the information flow among interlocutors and associating accurate status/actions with the corresponding interlocutors and person mentions.
Zhengyuan Liu, Ke Shi 0001, Nancy F. Chen
SIGDIAL3
2021 Domain-Shift Conditioning Using Adaptable Filtering Via Hierarchical Embeddings for Robust Chinese Spell Check
abstract
Spell check is a useful application which processes noisy human-generated text. Spell check for Chinese poses unresolved problems due to the large number of characters, the sparse distribution of errors, and the dearth of resources with sufficient coverage of heterogeneous and shifting error domains. For Chinese spell check, filtering using confusion sets narrows the search space and makes finding corrections easier. However, most, if not all, confusion sets used to date are fixed and thus do not include new, shifting error domains. We propose a scalable adaptable filter that exploits hierarchical character embeddings to (1) obviate the need to handcraft confusion sets, and (2) resolve sparsity problems related to infrequent errors. Our approach compares favorably with competitive baselines and obtains SOTA results on the 2014 and 2015 Chinese Spelling Check Bake-off datasets.
Minh Nguyen 0002, Hoang Gia Ngo, Nancy F. Chen
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Would you Rather? A New Benchmark for Learning Machine Alignment with Cultural Values and Social Preferences
abstract
Understanding human preferences, along with cultural and social nuances, lives at the heart of natural language understanding.Concretely, we present a new task and corpus for learning alignments between machine and human preferences.Our newly introduced problem is concerned with predicting the preferable options from two sentences describing scenarios that may involve social and cultural situations.Our problem is framed as a natural language inference task with crowd-sourced preference votes by human players, obtained from a gamified voting platform.We benchmark several state-of-the-art neural models, along with BERT and friends on this task.Our experimental results show that current state-ofthe-art NLP models still leave much room for improvement.
Yi Tay, Donovan Ong, Jie Fu 0001, Alvin Chan, Nancy F. Chen, Anh Tuan Luu, Christopher Joseph Pal
ACL5
2020 Multilingual Neural RST Discourse Parsing
abstract
Text discourse parsing plays an important role in understanding information flow and argumentative structure in natural language.Previous research under the Rhetorical Structure Theory (RST) has mostly focused on inducing and evaluating models from the English treebank.However, the parsing tasks for other languages such as German, Dutch, and Portuguese are still challenging due to the shortage of annotated data.In this work, we investigate two approaches to establish a neural, cross-lingual discourse parser via: (1) utilizing multilingual vector representations; and(2) adopting segment-level translation of the source content.Experiment results show that both methods are effective even with limited training data, and achieve state-of-the-art performance on cross-lingual, document-level discourse parsing on all sub-tasks.
Zhengyuan Liu, Ke Shi 0001, Nancy F. Chen
COLING3
2020 BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded Dialogues
abstract
Video-grounded dialogues are very challenging due to (i) the complexity of videos which contain both spatial and temporal variations, and (ii) the complexity of user utterances which query different segments and/or different objects in videos over multiple dialogue turns.However, existing approaches to video-grounded dialogues often focus on superficial temporal-level visual cues, but neglect more fine-grained spatial signals from videos.To address this drawback, we propose Bi-directional Spatio-Temporal Learning (BiST), a vision-language neural framework for high-resolution queries in videos based on textual cues.Specifically, our approach not only exploits both spatial and temporal-level information, but also learns dynamic information diffusion between the two feature spaces through spatial-to-temporal and temporal-tospatial reasoning.The bidirectional strategy aims to tackle the evolving semantics of user queries in the dialogue setting.The retrieved visual cues are used as contextual information to construct relevant responses to the users.Our empirical results and comprehensive qualitative analysis show that BiST achieves competitive performance and generates reasonable responses on a large-scale AVSD benchmark.We also adapt our BiST models to the Video QA setting, and substantially outperform prior approaches on the TGIF-QA benchmark.* This work was mostly
Hung Le 0003, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi
EMNLP (1)3
2020 UniConv: A Unified Conversational Neural Architecture for Multi-domain Task-oriented Dialogues
abstract
Building an end-to-end conversational agent for multi-domain task-oriented dialogues has been an open challenge for two main reasons.First, tracking dialogue states of multiple domains is non-trivial as the dialogue agent must obtain complete states from all relevant domains, some of which might have shared slots among domains as well as unique slots specifically for one domain only.Second, the dialogue agent must also process various types of information across domains, including dialogue context, dialogue states, and database, to generate natural responses to users.Unlike the existing approaches that are often designed to train each module separately, we propose "UniConv" -a novel unified neural architecture for end-to-end conversational systems in multi-domain task-oriented dialogues, which is designed to jointly train (i) a Bi-level State Tracker which tracks dialogue states by learning signals at both slot and domain level independently, and (ii) a Joint Dialogue Act and Response Generator which incorporates information from various input components and models dialogue acts and target responses simultaneously.We conduct comprehensive experiments in dialogue state tracking, contextto-text, and end-to-end settings on the Multi-WOZ2.1 benchmark, achieving superior performance over competitive baselines.
Hung Le 0003, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi
EMNLP (1)4
2020 Unsupervised Feature Adaptation Using Adversarial Multi-Task Training for Automatic Evaluation of Children's Speech
Richeng Duan, Nancy F. Chen
INTERSPEECH2
2020 Characterization of Singaporean Children's English: Comparisons to American and British Counterparts Using Archetypal Analysis
Yuling Gu, Nancy F. Chen
INTERSPEECH2
2020 Computer-Assisted Language Learning System: Automatic Speech Evaluation for Children Learning Malay and Tamil
Ke Shi 0001, Kye Min Tan, Richeng Duan, Siti Umairah Md. Salleh, Nur Farah Ain Suhaimi, Rajan Vellu, Ngoc Thuy Huong Helen Thai, Nancy F. Chen
INTERSPEECH8
2020 Hierarchical multimodal attention for end-to-end audio-visual scene-aware dialogue response generation
Hung Le 0003, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi
Comput. Speech Lang.3
2020 Hierarchical Character Embeddings: Learning Phonological and Semantic Representations in Languages of Logographic Origin Using Recursive Neural Networks
abstract
Logographs (Chinese characters) have recursive structures (i.e. hierarchies of sub-units in logographs) that contain phonological and semantic information, as developmental psychology literature suggests that native speakers leverage on the structures to learn how to read. Exploiting these structures could potentially lead to better embeddings that can benefit many downstream tasks. We propose building hierarchical logograph (character) embeddings from logograph recursive structures using treeLSTM, a recursive neural network. Using recursive neural network imposes a prior on the mapping from logographs to embeddings since the network must read in the sub-units in logographs according to the order specified by the recursive structures. Based on human behavior in language learning and reading, we hypothesize that modeling logographs' structures using recursive neural network should be beneficial. To verify this claim, we consider two tasks (1) predicting logographs' Cantonese pronunciation from logographic structures and (2) language modeling. Empirical results show that the proposed hierarchical embeddings outperform baseline approaches. Diagnostic analysis suggests that hierarchical embeddings constructed using treeLSTM is less sensitive to distractors, thus is more robust, especially on complex logographs.
Minh Nguyen 0002, Hoang Gia Ngo, Nancy F. Chen
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Multimodal Transformer Networks for End-to-End Video-Grounded Dialogue Systems
abstract
Developing Video-Grounded Dialogue Systems (VGDS), where a dialogue is conducted based on visual and audio aspects of a given video, is significantly more challenging than traditional image or text-grounded dialogue systems because (1) feature space of videos span across multiple picture frames, making it difficult to obtain semantic information; and(2) a dialogue agent must perceive and process information from different modalities (audio, video, caption, etc.) to obtain a comprehensive understanding.Most existing work is based on RNNs and sequence-to-sequence architectures, which are not very effective for capturing complex long-term dependencies (like in videos).To overcome this, we propose Multimodal Transformer Networks (MTN) to encode videos and incorporate information from different modalities.We also propose queryaware attention through an auto-encoder to extract query-aware features from non-text modalities.We develop a training procedure to simulate token-level decoding to improve the quality of generated responses during inference.We get state of the art performance on Dialogue System Technology Challenge 7 (DSTC7).Our model also generalizes to another multimodal visual-grounded dialogue task, and obtains promising performance.
Hung Le 0003, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi
ACL (1)3
2019 Reading Turn by Turn: Hierarchical Attention Architecture for Spoken Dialogue Comprehension
abstract
Comprehending multi-turn spoken conversations is an emerging research area, presenting challenges different from reading comprehension of passages due to the interactive nature of information exchange from at least two speakers.Unlike passages, where sentences are often the default semantic modeling unit, in multi-turn conversations, a turn is a topically coherent unit embodied with immediately relevant context, making it a linguistically intuitive segment for computationally modeling verbal interactions.Therefore, in this work, we propose a hierarchical attention neural network architecture, combining turnlevel and word-level attention mechanisms, to improve spoken dialogue comprehension performance.Experiments are conducted on a multi-turn conversation dataset, where nurses inquire and discuss symptom information with patients.We empirically show that the proposed approach outperforms standard attention baselines, achieves more efficient learning outcomes, and is more robust to lengthy and out-of-distribution test samples.
Zhengyuan Liu, Nancy F. Chen
ACL (1)2
2019 Topic-Aware Pointer-Generator Networks for Summarizing Spoken Conversations
abstract
Due to the lack of publicly available resources, conversation summarization has received far less attention than text summarization. As the purpose of conversations is to exchange information between at least two interlocutors, key information about a certain topic is often scattered and spanned across multiple utterances and turns from different speakers. This phenomenon is more pronounced during spoken conversations, where speech characteristics such as backchanneling and false-starts might interrupt the topical flow. Moreover, topic diffusion and (intra-utterance) topic drift are also more common in human-to-human conversations. Such linguistic characteristics of dialogue topics make sentence-level extractive summarization approaches used in spoken documents ill-suited for summarizing conversations. Pointer-generator networks have effectively demonstrated its strength at integrating extractive and abstractive capabilities through neural modeling in text summarization. To the best of our knowledge, to date no one has adopted it for summarizing conversations. In this work, we propose a topic-aware architecture to exploit the inherent hierarchical structure in conversations to further adapt the pointer-generator model. Our approach significantly outperforms competitive baselines, achieves more efficient learning outcomes, and attains more robust performance.
Zhengyuan Liu, Angela Ng, Sheldon Lee Shao Guang, AiTi Aw, Nancy F. Chen
ASRU5
2019 Joint Learning of Word and Label Embeddings for Sequence Labelling in Spoken Language Understanding
abstract
We propose an architecture to jointly learn word and label embeddings for slot filling in spoken language understanding. The proposed approach encodes labels using a combination of word embeddings and straightforward word-label association from the training data. Compared to the state-of-the-art methods, our approach does not require label embed-dings as part of the input and therefore lends itself nicely to a wide range of model architectures. In addition, our architecture computes contextual distances between words and labels to avoid adding contextual windows, thus reducing memory footprint. We validate the approach on established spoken dialogue datasets and show that it can achieve state-of-the-art performance with much fewer trainable parameters.
Jiewen Wu, Luis Fernando D'Haro, Nancy F. Chen, Pavitra Krishnaswamy, Rafael E. Banchs
ASRU3
2019 Set to Ordered Text: Generating Discharge Instructions from Medical Billing Codes
abstract
Litton J Kurisinkel, Nancy Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Litton J. Kurisinkel, Nancy F. Chen
EMNLP/IJCNLP (1)2
2019 Improving Mispronunciation Detection of Mandarin Tones for Non-Native Learners With Soft-Target Tone Labels and BLSTM-Based Deep Tone Models
abstract
We investigate the effectiveness of soft-target tone labels and sequential context information for mispronunciation detection of Mandarin lexical tones pronounced by second language (L2) learners whose first language (L1) is of European origin. In conventional approaches, prosodic information (e.g., F0 and tone posteriors extracted from trained tone models) is used to calculate goodness of pronunciation (GOP) scores or train binary classifiers to verify pronunciation correctness. We propose three techniques to improve detection of mispronunciation of Mandarin tones for non-native learners. First, we extend our tone model from a deep neural network (DNN) to a bidirectional long short-term memory (BLSTM) network in order to more accurately model the high variability of non-native tone productions and the contextual information expressed in tone-level co-articulation. Second, we characterize ambiguous pronunciations where L2 learners' tone realizations are between two canonical tone categories by relaxing hard target labels to soft targets with probabilistic transcriptions. Third, segmental tone features fed into verifiers are extracted by a BLSTM to exploit sequential context information to improve mispronunciation detection. Compared to DNN-GOP trained with hard targets, the proposed BLSTM-GOP framework trained with soft targets reduces the tones' averaged equal error rate (ERR) from 7.58% to 5.83% and the averaged area under ROC curve (AUC) is increased from 97.85% to 98.31%. By utilizing BLSTM-based verifiers the EER further decreases to 5.16%, and the AUC is increased to 98.47%.
Wei Li 0119, Nancy F. Chen, Sabato Marco Siniscalchi, Chin-Hui Lee 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Phonology-Augmented Statistical Framework for Machine Transliteration Using Limited Linguistic Resources
abstract
Transliteration converts words in a source language (e.g., English) into words in a target language (e.g., Vietnamese). This conversion considers the phonological structure of the target language, as the transliterated output needs to be pronounceable in the target language. For example, a word in Vietnamese that begins with a consonant cluster is phonologically invalid and thus would be an incorrect output of a transliteration system. Most statistical transliteration approaches, albeit being widely adopted, do not explicitly model the target language's phonology, which often results in invalid outputs. The problem is compounded by the limited linguistic resources available when converting foreign words to transliterated words in the target language. In this paper, we present a phonology-augmented statistical framework suitable for transliteration, especially when only limited linguistic resources are available. We propose the concept of pseudo-syllables as structures representing how segments of a foreign word are organized according to the syllables of the target language's phonology. We performed transliteration experiments on Vietnamese and Cantonese. We show that the proposed framework outperforms the statistical baseline by up to 44.68% relative, when there are limited training examples (587 entries).
Hoang Gia Ngo, Minh Nguyen 0002, Nancy F. Chen
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Multimodal neural pronunciation modeling for spoken languages with logographic origin
abstract
Graphemes of most languages encode pronunciation, though some are more explicit than others.Languages like Spanish have a straightforward mapping between its graphemes and phonemes, while this mapping is more convoluted for languages like English.Spoken languages such as Cantonese present even more challenges in pronunciation modeling: (1) they do not have a standard written form, (2) the closest graphemic origins are logographic Han characters, of which only a subset of these logographic characters implicitly encodes pronunciation.In this work, we propose a multimodal approach to predict the pronunciation of Cantonese logographic characters, using neural networks with a geometric representation of logographs and pronunciation of cognates in historically related languages.The proposed framework improves performance by 18.1% and 25.0% respective to unimodal and multimodal baselines.
Minh Nguyen 0002, Hoang Gia Ngo, Nancy F. Chen
EMNLP3
2018 Recognizing Zero-Resourced Languages Based on Mismatched Machine Transcriptions
abstract
Mismatched crowdsourcing based probabilistic human transcription has been proposed recently for training and adapting acoustic models for zero-resourced languages where we do not have any native transcriptions. This paper describes a machine transcription based phone recognition system for recognizing zero-resourced languages and compares it with baseline systems of MAP adaptation and semi-supervised self training. With a set of available speech recognizers in source languages that cover all the basic phonetic features, this work shows that we can use mismatched machine transcriptions from these source languages to achieve human level transcriptions, bypassing the laborious efforts of obtaining human transcriptions. We also present a fully automated unsupervised approach for zero-resourced speech recognition using mismatched machine transcriptions for transfer learning of phone models.
Wenda Chen, Mark Hasegawa-Johnson, Nancy F. Chen
ICASSP3
2018 Improving Mandarin Tone Mispronunciation Detection for Non-Native Learners with Soft-Target Tone Labels and BLSTM-Based Deep Models
abstract
We propose three techniques to improve mispronunciation detection of Mandarin tones of second language (L2) learners using tone-based extended recognition network (ERN). First, we extend our model from deep neural network (DNN) to bidirectionallon-short-term memory (BLSTM) in order to model tone-level co-articulation influenced by a broader temporal context (e.g., two or three consecutive Mandarin syllables). Second, we relax the hard labels to characterize the situations when a single tone class label is not enough because L2 learners' pronunciations are often between two canonical tone categories. Therefore, soft targets (a probabilistic transcription) are proposed for acoustic model training in place of conventional hard targets (one-hot targets). Third, we average tone scores produced by BLSTM models trained with hard and soft targets to seek the complementarity from modeling at the tone-target levels. Compared to our previous system based on the DNN-trained ERNs, the BLSTM-trained system with soft targets reduces the equal error rate (ERR) from 5.77% to 4.86%, and system combination decreases EER further to 4.34%, achieving a 24.78% relative error reduction.
Wei Li 0119, Nancy F. Chen, Sabato Marco Siniscalchi, Chin-Hui Lee 0001
ICASSP2
2018 Topic and Keyword Identification for Low-resourced Speech Using Cross-Language Transfer Learning
Wenda Chen, Mark Hasegawa-Johnson, Nancy F. Chen
INTERSPEECH3
2018 Re-ranking spoken term detection with acoustic exemplars of keywords
Van Tung Pham, Haihua Xu 0001, Nancy F. Chen, Chng Eng Siong, Haizhou Li 0001
Speech Commun.4
2018 Multitask Learning for Phone Recognition of Underresourced Languages Using Mismatched Transcription
abstract
It is challenging to obtain large amounts of native (matched) labels for speech audio in underresourced languages. This challenge is often due to a lack of literate speakers of the language, or in extreme cases, a lack of universally acknowledged orthography as well. One solution is to increase the amount of labeled data by using mismatched transcription, which employs transcribers who do not speak the underresourced language of interest called the target language (in place of native speakers), to transcribe what they hear as nonsense speech in their own annotation language (≠ target language). Previous uses of mismatched transcription converted it to a probabilistic transcription (PT), but PT is limited by the errors of nonnative perception. This paper proposes, instead, a multitask learning framework in which one deep neural network (DNN) is trained to optimize two separate tasks: acoustic modeling of a small number of matched transcription with matched target-language graphemes; and acoustic modeling of a large number of mismatched transcription with mismatched annotation-language graphemes. We find that: first, the multitask learning framework gives significant improvement over monolingual, semisupervised learning, multilingual DNN training, and transfer learning baselines; second, a Gaussian Mixture Model-Hidden-Markov Model (GMM-HMM) model adapted using PT improves alignments, thereby improving training; and third, bottleneck features trained on the mismatched transcriptions lead to even better alignments, resulting in further performance gains of the multitask DNN. Our experiments are conducted on the IARPA Georgian and Vietnamese BABEL corpora as well as on our newly collected speech corpus of Singapore Hokkien, an underresourced language with no standard written form.
Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Efficient methods to train multilingual bottleneck feature extractors for low resource keyword search
abstract
Training a bottleneck feature (BNF) extractor with multilingual data has been common in low resource keyword search. In a low resource application, the amount of transcribed target language data is limited while there are usually plenty of multilingual data. In this paper, we investigated two methods to train efficient multilingual BNF extractors for low resource keyword search. One method is to use the target language data to update an existing BNF extractor, and another method is to combine the target language data to train a new multilingual BNF extractor from the start. In these two methods, we proposed to use long short-term memory recurrent neural network based language identification to select utterances in the multilingual training data that are acoustically close to the target language. Experiments on Swahili in the OpenKWS15 data demonstrated the efficiency of our proposed methods. The first method facilitates rapid system development, while both methods outperform using baseline BNF extractors in terms of accuracy.
Chongjia Ni, Cheung-Chi Leung, Lei Wang 0020, Nancy F. Chen, Bin Ma 0001
ICASSP4
2017 Mismatched Crowdsourcing from Multiple Annotator Languages for Recognizing Zero-Resourced Languages: A Nullspace Clustering Approach
Wenda Chen, Mark Hasegawa-Johnson, Nancy F. Chen, Boon Pang Lim
INTERSPEECH3
2017 Multi-Task Learning Using Mismatched Transcription for Under-Resourced Speech Recognition
Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson
INTERSPEECH2
2017 Improving Mispronunciation Detection for Non-Native Learners with Multisource Information and LSTM-Based Deep Models
abstract
In this paper, we utilize manner and place of articulation features and deep neural network models (DNNs) with long short-term memory (LSTM) to improve the detection performance of phonetic mispronunciations produced by second language learners. First, we show that speech attribute scores are complementary to conventional phone scores, so they can be concatenated as features to improve a baseline system based only on phone information. Next, pronunciation representation, usually calculated by frame-level averaging in a DNN, is now learned by LSTM, which directly uses sequential context information to embed a sequence of pronunciation scores into a pronunciation vector to improve the perfonnance of subsequent mispronunciation detectors. Finally, when both proposed techniques are incorporated into the baseline phone-based GOP (goodness of pronunciation) classifier system trained on the same data, the integrated system reduces the false acceptance rate (FAR) and false rejection rate (FRR) by 37.90% and 38.44% (relative), respectively, from the baseline system.
Wei Li 0119, Nancy F. Chen, Sabato Marco Siniscalchi, Chin-Hui Lee 0001
INTERSPEECH2
2017 Multi-Task Learning for Mispronunciation Detection on Singapore Children's Mandarin Speech
Rong Tong, Nancy F. Chen, Bin Ma 0001
INTERSPEECH2
2017 ASR for Under-Resourced Languages From Probabilistic Transcription
abstract
In many under-resourced languages it is possible to find text, and it is possible to find speech, but transcribed speech suitable for training automatic speech recognition (ASR) is unavailable. In the absence of native transcripts, this paper proposes the use of a probabilistic transcript: A probability mass function over possible phonetic transcripts of the waveform. Three sources of probabilistic transcripts are demonstrated. First, self-training is a well-established semisupervised learning technique, in which a cross-lingual ASR first labels unlabeled speech, and is then adapted using the same labels. Second, mismatched crowdsourcing is a recent technique in which nonspeakers of the language are asked to write what they hear, and their nonsense transcripts are decoded using noisy channel models of second-language speech perception. Third, EEG distribution coding is a new technique in which nonspeakers of the language listen to it, and their electrocortical response signals are interpreted to indicate probabilities. ASR was trained in four languages without native transcripts. Adaptation using mismatched crowdsourcing significantly outperformed self-training, and both significantly outperformed a cross-lingual baseline. Both EEG distribution coding and text-derived phone language models were shown to improve the quality of probabilistic transcripts derived from mismatched crowdsourcing.
Mark Hasegawa-Johnson, Preethi Jyothi, Daniel McCloy, Majid Mirbagheri, Giovanni M. Di Liberto, Amit Das 0007, Bradley Ekin, Chunxi Liu, Vimal Manohar, Hao Tang 0002, Edmund C. Lalor, Nancy F. Chen, Paul Hager, Tyler Kekona, Rose Sloan, Adrian K. C. Lee
IEEE ACM Trans. Audio Speech Lang. Process.12
2016 Exemplar-inspired strategies for low-resource spoken keyword search in Swahili
abstract
We present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational Swahili provided by the IARPA Babel program. In this work, we elaborate on the following: (1) exploiting exemplar training samples to construct a non-parametric acoustic model using kernel density estimation at test time; (2) rescoring hypothesized keyword detections through quantifying their acoustic similarity with exemplar training samples; (3 ) extending our previously proposed system combination approach to incorporate prosody features of exemplar keyword samples.
Nancy F. Chen, Van Tung Pham, Haihua Xu 0001, Van Hai Do, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Chin-Hui Lee 0001, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001
ICASSP1
2016 Personalized mispronunciation detection and diagnosis based on unsupervised error pattern discovery
abstract
In this work, we introduce two improvements to our previously proposed mispronunciation detection framework. The framework focuses on each learner individually and consists of two main procedures: unsupervised error pattern discovery and pronunciation error decoding. First, we propose nbest filtering to disambiguate uncertain error candidate hypotheses obtained from acoustic similarity clustering. Second, we propose personalized template-based rescoring to refine the mispronunciation detection results. The second contribution of the paper is that we demonstrate the portability of the framework to a new target language. Experimental results on the iCALL corpus, a nonnative Mandarin corpus consisting of speakers of European origin, show that the new error pattern discovery process significantly reduces the size and increases the coverage of the error candidate set. Also, the rescoring technique effectively improves system performance on mispronunciation detection and diagnosis.
Ann Lee 0001, Nancy F. Chen, James R. Glass
ICASSP2
2016 Improving non-native mispronunciation detection and enriching diagnostic feedback with DNN-based speech attribute modeling
abstract
We propose the use of speech attributes, such as voicing and aspiration, to address two key research issues in computer assisted pronunciation training (CAPT) for L2 learners, namely detecting mispronunciation and providing diagnostic feedback. To improve the performance we focus on mispronunciations occurred at the segmental and sub-segmental levels. In this study, speech attributes scores are first used to measure the pronunciation quality at a sub-segmental level, such as manner and place of articulation. These speech attribute scores are integrated by neural network classifiers to generate segmental pronunciation scores. Compared with the conventional phone-based GOP (Goodness of Pronunciation) system we implement with our dataset, the proposed framework reduces the equal error rate by 8.78% relative. Moreover, it attains comparable results to phone-based classifier approach to mispronunciation detection while providing comprehensive feedback, including segmental and sub-segmental diagnostic information, to help L2 learners.
Wei Li 0119, Sabato Marco Siniscalchi, Nancy F. Chen, Chin-Hui Lee 0001
ICASSP3
2016 Cross-lingual deep neural network based submodular unbiased data selection for low-resource keyword search
abstract
In this paper, we propose a cross-lingual deep neural network (DNN) based submodular unbiased data selection approach for low-resource keyword search (KWS). A small amount (e.g. one hour) of transcribed data is used to conduct cross-lingual transfer. The frame-level senone sequence activated by the cross-lingual DNN is used to represent each untranscribed speech utterance. The proposed submodular function considers utterance length normalization and the feature distribution matched to a development set. Experiments are conducted by selecting 9 hours of Tamil speech for the 2014 NIST Open Keyword Search Evaluation (OpenKWS14). The proposed data selection approach provides 35.8% relative actual term weighted value (ATWV) improvement over random selection on the OpenKWS14 Evalpartl data set. Further analysis of the experimental results shows that both utterance length normalization and the feature distribution estimated from a development set deployed in the submodular function can suppress the preference to select long utterances. The selected utterances can cover a more diverse range of tri-phones, words, and acoustic variations from a wider set of utterances. Moreover, the wider coverage of words also benefits the acquired linguistic knowledge, which also contributes to improving KWS performance.
Chongjia Ni, Cheung-Chi Leung, Lei Wang 0020, Feng Rao, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
ICASSP7
2016 Keyword search using query expansion for graph-based rescoring of hypothesized detections
abstract
In this work, we propose a novel framework for rescoring keyword search (KWS) detections using acoustic samples extracted from the training data. We view the keyword rescoring task as an information retrieval task and adopt the idea of query expansion. We expand a textual keyword with multiple speech keyword samples extracted from the training data. In this way, the hypothesized detections are compared with the multiple keywords using non-parametric approaches such as dynamic time warping (DTW). The obtained similarity scores are used in a graph based method to re-rank the original confidence scores estimated by the automatic speech recognition (ASR) systems. Experimental results on the NIST OpenKWS15 Evaluation show that our rescoring method is effective, especially for the subword system. For subword experiments, the graph-based rescoring with training samples obtains 5.1% and 1.5% absolute improvement over two baseline systems. One is a standard parametric ASR system, while the other is the graph-based rescoring without training samples.
Van Tung Pham, Haihua Xu 0001, Nancy F. Chen, Chng Eng Siong, Haizhou Li 0001
ICASSP4
2016 SingaKids-Mandarin: Speech Corpus of Singaporean Children Speaking Mandarin Chinese
Nancy F. Chen, Rong Tong, Darren Wee, Pei Xuan Lee, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH1
2016 Analysis of Mismatched Transcriptions Generated by Humans and Machines for Under-Resourced Languages
Van Hai Do, Nancy F. Chen, Boon Pang Lim, Mark Hasegawa-Johnson
INTERSPEECH2
2016 Perception of Tone in Whispered Mandarin Sentences: The Case for Singapore Mandarin
Yuling Gu, Boon Pang Lim, Nancy F. Chen
INTERSPEECH3
2016 Detecting Mispronunciations of L2 Learners and Providing Corrective Feedback Using Knowledge-Guided and Data-Driven Decision Trees
abstract
We propose a novel decision tree based framework to detect phonetic mispronunciations produced by L2 learners caused by using inaccurate speech attributes, such as manner and place of articulation. Compared with conventional score-based CAPT (computer assisted pronunciation training) systems, our proposed framework has three advantages: (1) each mispronunciation in a tree can be interpreted and communicated to the L2 learners by traversing the corresponding path from a leaf node to the root node; (2) corrective feedback based on speech attribute features, which are directly used to describe how consonants and vowels are produced using related articulators, can be provided to the L2 learners; and (3) by building the phone-dependent decision tree, the relative importance of the speech attribute features of a target phone can be automatically learned and used to distinguish itself from other phones. This information can provide L2 learners speech attribute feedback that is ranked in order of importance. In addition to the abovementioned advantages, experimental results confirm that the proposed approach can detect most pronunciation errors and provide accurate diagnostic feedback
Wei Li 0119, Kehuang Li, Sabato Marco Siniscalchi, Nancy F. Chen, Chin-Hui Lee 0001
INTERSPEECH4
2016 Rescoring Hypothesized Detections of Out-of-Vocabulary Keywords Using Subword Samples
Van Tung Pham, Haihua Xu 0001, Nancy F. Chen, Chng Eng Siong, Haizhou Li 0001
INTERSPEECH4
2016 Context Aware Mispronunciation Detection for Mandarin Pronunciation Training
Rong Tong, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH2
2016 Large-scale characterization of non-native Mandarin Chinese spoken by speakers of European origin: Analysis on iCALL
abstract
In this work, we analyze phonetic and prosodic pronunciation patterns from iCALL, a speech corpus designed to evaluate Mandarin mispronunciations by non-native speakers of European origin and to address the lack of large-scale, non-native corpora with comprehensive annotations for applications in CAPT (computer-assisted pronunciation training). iCALL consists of 90,841 utterances from 305 speakers with a total duration of 142 hours. The speakers are from diverse linguistic backgrounds (spanning Germanic, Romance, and Slavic native languages). The read utterances are phonetically balanced with phonetic, tonal, and fluency annotations. Our findings on iCALL reveal that lexical tone errors are over six times more prevalent than phonetic errors, French speakers are twice as likely to mispronounce Tone 2, 3, 4 when compared to English speakers, native Romance language speakers are more likely to make de-aspiration and aspiration mistakes, and fluency scores correlate inversely with tone and phone error rate.
Nancy F. Chen, Darren Wee, Rong Tong, Bin Ma 0001, Haizhou Li 0001
Speech Commun.1
2015 Low-resource keyword search strategies for tamil
abstract
We propose strategies for a state-of-the-art keyword search (KWS) system developed by the SINGA team in the context of the 2014 NIST Open Keyword Search Evaluation (OpenKWS14) using conversational Tamil provided by the IARPA Babel program. To tackle low-resource challenges and the rich morphological nature of Tamil, we present highlights of our current KWS system, including: (1) Submodular optimization data selection to maximize acoustic diversity through Gaussian component indexed N-grams; (2) Keywordaware language modeling; (3) Subword modeling of morphemes and homophones.
Nancy F. Chen, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Van Tung Pham, Haihua Xu 0001, Tze Siong Lau, Su Jun Leow, Boon Pang Lim, Cheung-Chi Leung, Lei Wang 0020, Chin-Hui Lee 0001, Alvina Goh, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001
ICASSP1
2015 A keyword-aware grammar framework for LVCSR-based spoken keyword search
abstract
In this paper, we proposed a method to realize the recently developed keyword-aware grammar for LVCSR-based keyword search using weight finite-state automata (WFSA). The approach creates a compact and deterministic grammar WFSA by inserting keyword paths to an existing n-gram WFSA. Tested on the evalpart1 data of the IARPA Babel OpenKWS13 Vietnamese and OpenKWS14 Tamil limited language pack tasks, the experimental results indicate the proposed keyword-aware framework achieves significant improvement, with about 50% relative actual term weighted value (ATWV) enhancement for both languages. Comparisons between the keyword-aware grammar and our previously proposed n-gram LM based approximation approach for the grammar also show that the KWS performances of these two realizations are complementary.
I-Fan Chen, Chongjia Ni, Boon Pang Lim, Nancy F. Chen, Chin-Hui Lee 0001
ICASSP4
2015 Unsupervised data selection and word-morph mixed language model for tamil low-resource keyword search
abstract
This paper considers an unsupervised data selection problem for the training data of an acoustic model and the vocabulary coverage of a keyword search system in low-resource settings. We propose to use Gaussian component index based n-grams as acoustic features in a submodular function for unsupervised data selection. The submodular function provides a near-optimal solution in terms of the objective being optimized. Moreover, to further resolve the high out-of-vocabulary (OOV) rate for morphologically-rich languages like Tamil, word-morph mixed language modeling is also considered. Our experiments are conducted on the Tamil speech provided by the IAPRA Babel program for the 2014 NIST Open Keyword Search Evaluation (OpenKWS14). We show that the selection of data plays an important role to the word error rate of the speech recognition system and the actual term weighted value (ATWV) of the keyword search system. The 10 hours of speech selected from the full language pack (FLP) using the proposed algorithm provides a relative 23.2% and 20.7% ATWV improvement over two other data subsets, the 10-hour data from the limited language pack (LLP) defined by IARPA and the 10 hours of speech randomly selected from the FLP, respectively. The proposed algorithm also increases the vocabulary coverage, implicitly alleviating the OOV problem: The number of OOV search terms drops from 1,686 and 1,171 in the two baseline conditions to 972.
Chongjia Ni, Cheung-Chi Leung, Lei Wang 0020, Nancy F. Chen, Bin Ma 0001
ICASSP4
2015 Tokenizing fundamental frequency variation for Mandarin tone error detection
abstract
Tone error is commonly observed in tonal language acquisition. Correct tone production is especially challenging for native speakers of non-tonal languages. In this paper, we exploit the fundamental frequency variation (FFV) feature for Mandarin tone error detection. We propose to use FFV through two approaches: (1) Concatenating FFVs along side with standard speech recognition features; (2) Token FFV: Characterizing pitch variation with longer temporal context through GMM tokenization and n-gram language modeling. Our results show that tone error detection improves by incorporating FFV features and the two approaches are complementary to each other.
Rong Tong, Nancy F. Chen, Boon Pang Lim, Bin Ma 0001, Haizhou Li 0001
ICASSP2
2015 iCALL corpus: Mandarin Chinese spoken by non-native speakers of European descent
Nancy F. Chen, Rong Tong, Darren Wee, Pei Xuan Lee, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH1
2015 Phonology-augmented statistical transliteration for low-resource languages
Hoang Gia Ngo, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH2
2015 Goodness of tone (GOT) for non-native Mandarin tone recognition
Rong Tong, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH2
2014 Strategies for Vietnamese keyword search
abstract
We propose strategies for a state-of-the-art Vietnamese keyword search (KWS) system developed at the Institute for Infocomm Research (I2R). The KWS system exploits acoustic features characterizing creaky voice quality peculiar to lexical tones in Vietnamese, a minimal-resource transliteration framework to alleviate out-of-vocabulary issues from foreign loan words, and a proposed system combination scheme FusionX. We show that the proposed creaky voice quality features complement pitch-related features, reaching fusion gains of 17.7% relative (6.9% absolute). To the best of our knowledge, the proposed transliteration framework is the first reported rule-based system for Vietnamese; it outperforms statistical-approach baselines up to 14.93–36.73% relative on foreign loan word search tasks. Using FusionX to combine 3 sub-systems, the actual term-weighted value (ATWV) reaches 0.4742, exceeding the ATWV=0.3 benchmark for IARPA Babel participants in the NIST OpenKWSB Evaluation.
Nancy F. Chen, Sunil Sivadas, Boon Pang Lim, Hoang Gia Ngo, Haihua Xu 0001, Van Tung Pham, Bin Ma 0001, Haizhou Li 0001
ICASSP1
2014 Discriminative score normalization for keyword search decision
abstract
Many keyword search (KWS) systems make “hit/false alarm (FA)” decisions based on the lattice-based posterior probability, which is incomparable across keywords. Therefore, score normalization is essential for a KWS system. In this paper, we investigate the integration of two novel features, ranking-score and relative-to-max, into a discriminative score normalization method. These features are extracted by considering all competing hypotheses of a putative detection. A metric-based normalization method is also applied as a post-processing step to further optimize the term-weighted value (TWV) evaluation metric. We report empirical improvements over standard baselines using the Vietnamese data from IARPA's Babel program in the NIST OpenKWS13 Evaluation setup.
Van Tung Pham, Haihua Xu 0001, Nancy F. Chen, Sunil Sivadas, Boon Pang Lim, Chng Eng Siong, Haizhou Li 0001
ICASSP3
2014 Subspace Gaussian mixture model for computer-assisted language learning
abstract
In computer-assisted language learning (CALL), speech data from non-native speakers are usually insufficient for acoustic modeling. Subspace Gaussian Mixture Models (SGMM) have been effective in training automatic speech recognition (ASR) systems with limited amounts of training data. Therefore, in this work, we propose to use SGMM to improve the fluency assessment performance. In particular, the contributions of this work are: (i) The proposed SGMM acoustic model trained with native data outperforms the MMI-GMM/HMM baseline by 25% relative, (ii) when incorporating a small amount of non-native training data, the SGMM acoustic model further improves the performance of fluency assessment by 47% relative.
Rong Tong, Boon Pang Lim, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
ICASSP3
2014 A keyword-boosted sMBR criterion to enhance keyword search performance in deep neural network based acoustic modeling
I-Fan Chen, Nancy F. Chen, Chin-Hui Lee 0001
INTERSPEECH2
2014 A whispered Mandarin corpus for speech technology applications
Pei Xuan Lee, Darren Wee, Hilary Si Yin Toh, Boon Pang Lim, Nancy F. Chen, Bin Ma 0001
INTERSPEECH5
2014 A minimal-resource transliteration framework for vietnamese
Hoang Gia Ngo, Nancy F. Chen, Sunil Sivadas, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH2
2014 System and keyword dependent fusion for spoken term detection
abstract
System combination (or data fusion1) is known to provide significant improvement for spoken term detection (STD). The key issue of the system combination is how to effectively fuse the various scores of participant systems. Currently, most system combination methods are system and keyword independent, i.e. they use the same arithmetic functions to combine scores for all keywords. Although such strategy improve keyword search performance, the improvement is limited. In this paper we first propose an arithmetic-based system combination method to incorporate the system and keyword characteristics into the fusion procedure to enhance the effectiveness of system combination. The method incorporates a system-keyword dependent property, which is the number of acceptances in this paper, into the combination procedure. We then introduce a discriminative model to combine various useful system and keyword characteristics into a general framework. Improvements over standard baselines are observed on the Vietnamese data from IARPA Babel program with the NIST OpenKWS13 Evaluation setup.
Van Tung Pham, Nancy F. Chen, Sunil Sivadas, Haihua Xu 0001, I-Fan Chen, Chongjia Ni, Chng Eng Siong, Haizhou Li 0001
SLT2
2014 Characterizing Phonetic Transformations and Acoustic Differences Across English Dialects
abstract
In this work, we propose a framework that automatically discovers dialect-specific phonetic rules. These rules characterize when certain phonetic or acoustic transformations occur across dialects. To explicitly characterize these dialect-specific rules, we adapt the conventional hidden Markov model to handle insertion and deletion transformations. The proposed framework is able to convert pronunciation of one dialect to another using learned rules, recognize dialects using learned rules, retrieve dialect-specific regions, and refine linguistic rules. Potential applications of our proposed framework include computer-assisted language learning, sociolinguistics, and diagnosis tools for phonological disorders.
Nancy F. Chen, Sharon W. Tam, Wade Shen, Joseph P. Campbell
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Minimal-resource phonetic language models to summarize untranscribed speech
abstract
We propose to extract summary sentences from lexically untranscribed speech via phone tokenization. We use decoded phone sequences instead of words to train language models to infer semantically significant utterances. Phone tokens yield comparable results to words on the TDT-2 English corpus, yet require significantly less linguistic resources - no need for automatic speech recognition (ASR): (1) Using decoded phones of high phone error rate (78.7%) leads to comparable results to using ASR-decoded words. (2) Tokenizing English audio using a Czech phone recognizer leads to comparable results to using English words from closed-captions. These trends parallel those established in spoken language recognition and have practical significance: we can potentially summarize speech passages of resource-poor languages by leveraging existing tools developed on resource-rich languages.
Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001
ICASSP1
2013 Large-scale characterization of Mandarin pronunciation errors made by native speakers of European languages
Nancy F. Chen, Vivaek Shivakumar, Mahesh Harikumar, Bin Ma 0001, Haizhou Li 0001
INTERSPEECH1
2012 Analyzing and Interpreting Automatically Learned Rules Across Dialects
abstract
In this paper, we demonstrate how informative dialect recogni-tion systems such as acoustic pronunciation model (APM) help speech scientists locate and analyze phonetic rules efficiently. In particular, we analyze dialect-specific characteristics auto-matically learned from APM across two American English di-alects. We show that unsupervised rule retrieval performs sim-ilarly to supervised retrieval, indicating that APM is useful in practical applications, where word transcripts are often unavail-able. We also demonstrate that the top-ranking rules learned from APM generally correspond to the linguistic literature, and can even pinpoint potential research directions to refine existing knowledge. Thus, the APM system can help phoneticians ana-lyze rules efficiently by characterizing large amounts of data to postulate rule candidates, so they can reserve time to conduct more targeted investigations. Potential applications of informa-tive dialect recognition systems include forensic phonetics and diagnosis of spoken language disorders. Index Terms: informative dialect recognition, rule retrieval, phonological rules, forensic phonetics
Nancy F. Chen, Wade Shen, Joseph P. Campbell
INTERSPEECH1
2011 Informative dialect recognition using context-dependent pronunciation modeling
abstract
We propose an informative dialect recognition system that learns phonetic transformation rules, and uses them to identify dialects. A hidden Markov model is used to align reference phones with dialect specific pronunciations to characterize when and how often substitutions, insertions, and deletions occur. Decision tree clustering is used to find context-dependent phonetic rules. We ran recognition tasks on 4 Arabic dialects. Not only do the proposed systems perform well on their own, but when fused with baselines they improve performance by 21-36% relative. In addition, our proposed decision-tree system beats the baseline monophone system in recovering phonetic rules by 21% relative. Pronunciation rules learned by our proposed system quantify the occurrence frequency of known rules, and suggest rule candidates for further linguistic studies.
Nancy F. Chen, Wade Shen, Joseph P. Campbell, Pedro A. Torres-Carrasquillo
ICASSP1
2011 Characterizing Deletion Transformations Across Dialects Using a Sophisticated Tying Mechanism
abstract
In this work, we propose extensions of our Phone-based Pronunciation Model (PPM) for analyzing dialect differences. We compared these systems using 3 metrics and 2 datasets of English. Empirical results suggest that (1) sophisticated tying is suitable in modeling deletion transformations across dialects, beating standard tying by 33% relative, and (2) APM (Acousticbased Pronunciation Model) improves performance in generating dialect-specific pronunciations, dialect identification and rule retrieval, achieving relative gains beyond 34%. Index Terms: acoustic model, pronunciation model, phonetic rule, phonetic transformation, dialect recognition
Nancy F. Chen, Wade Shen, Joseph P. Campbell
INTERSPEECH1
2010 A linguistically-informative approach to dialect recognition using dialect-discriminating context-dependent phonetic models
abstract
We propose supervised and unsupervised learning algorithms to extract dialect discriminating phonetic rules and use these rules to adapt biphones to identify dialects. Despite many challenges (e.g., sub-dialect issues and no word transcriptions), we discovered dialect discriminating biphones compatible with the linguistic literature, while outperforming a baseline monophone system by 7.5% (relative). Our proposed dialect discriminating biphone system achieves similar performance to a baseline all-biphone system despite using 25% fewer biphone models. In addition, our system complements PRLM (Phone Recognition followed by Language Modeling), verified by obtaining relative gains of 15-29% when fused with PRLM. Our work is an encouraging first step towards a linguistically-informative dialect recognition system, with potential applications in forensic phonetics, accent training, and language learning.
Nancy F. Chen, Wade Shen, Joseph P. Campbell
ICASSP1
2009 Large-scale analysis of formant frequency estimation variability in conversational telephone speech
abstract
We quantify how the telephone channel and regional dialect influence formant estimates extracted from Wavesurfer [1, 2] in spontaneous conversational speech from over 3,600 native American English speakers. To the best of our knowledge, this is the largest scale study on this topic. We found that F1 estimates are higher in cellular channels than those in landline, while F2 in general shows an opposite trend. We also characterized vowel shift trends in northern states in U.S.A. and compared them with the Northern city chain shift (NCCS) [3]. Our analysis is useful in forensic applications where it is important to distinguish between speaker, dialect, and channel characterisitcs. Index Terms: formant frequency, Northern city chain shift, spontaneous conversational speech, telephone channel, American
Nancy F. Chen, Wade Shen, Joseph P. Campbell, Reva Schwartz
INTERSPEECH1
2008 Dialect recognition using adapted phonetic models
abstract
In this paper, we introduce a dialect recognition method that makes use of phonetic models adapted per dialect without phonetically labeled data. We show that this method can be implemented efficiently within an existing PRLM[1] system. We compare the performance of this system with other state-of-theart dialect recognition methods (both acoustic and token-based) on the NIST LRE 2007 English and Mandarin dialect recognition tasks. Our experimental results indicate that this system can perform better than baseline GMM and adapted PRLM systems, and also results in consistent gains of 15-23% when combined with other systems.
Wade Shen, Nancy F. Chen, Douglas A. Reynolds
INTERSPEECH2