EDBT 2026 Demo / reviewers in the wild / expert
Pascale Fung
dblp:29/4187
· DBLP profile ↗
175ranked-venue papers
23as first author
46since 2021 · last 2026
0000-0002-0628-7132ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 137 · 17 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 80 · 11 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards AI That Understands the Human WorldabstractAI has reached a turning point. Systems can now perceive, generate, and act in language and image across digital platforms at unprecedented scale. Yet as AI moves from tools to collaborators—embedded in decision-making, institutions, and everyday life—a new requirement becomes unavoidable: AI must understand the world the way humans inhabit it. Pascale Fung |
WWW | 1 |
| 2025 | HalluLens: LLM Hallucination BenchmarkabstractYejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn, Tara Fowler, Cheng Zhang, Nicola Cancedda, Pascale Fung. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yejin Bang, Ziwei Ji 0001, Alan Schelten, Anthony Hartshorn, Tara Fowler, Nicola Cancedda, Pascale Fung |
ACL (1) | 8 |
| 2025 | Calibrating Verbal Uncertainty as a Linear Feature to Reduce HallucinationsabstractZiwei Ji, Lei Yu, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Cheng Zhang, Pascale Fung, Nicola Cancedda. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ziwei Ji 0001, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Pascale Fung, Nicola Cancedda |
EMNLP | 8 |
| 2025 | Subobject-level Image TokenizationabstractPatch-based image tokenization ignores the morphology of the visual world, limiting effective and efficient learning of image understanding. Inspired by subword tokenization, we introduce subobject-level adaptive token segmentation and explore several approaches, including superpixel, SAM, and a proposed Efficient and PanOptiC (EPOC) image tokenizer. Our EPOC combines boundary detection–a simple task that can be handled well by a compact model–with watershed segmentation, which inherently guarantees no pixels are left unsegmented. Intrinsic evaluations across 5 datasets demonstrate that EPOC’s segmentation aligns well with human annotations of both object- and part-level visual morphology, producing more monosemantic tokens and offering substantial efficiency advantages. For extrinsic evaluation, we designed a token embedding that handles arbitrary-shaped tokens, and trained VLMs with different tokenizers on 4 datasets of object recognition and detailed captioning. The results reveal that subobject tokenization enables faster convergence and better generalization while using fewer visual tokens. Delong Chen, Samuel Cahyawijaya, Jianfeng Liu 0002, Baoyuan Wang, Pascale Fung |
ICML | 5 |
| 2025 | High-Dimension Human Value Representation in Large Language ModelsabstractSamuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji, Etsuko Ishii, Pascale Fung. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Samuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji 0001, Etsuko Ishii, Pascale Fung |
NAACL (Long Papers) | 8 |
| 2024 | Measuring Political Bias in Large Language Models: What Is Said and How It Is SaidabstractWe propose to measure political bias in LLMs by analyzing both the content and style of their generated content regarding political issues.Existing benchmarks and measures focus on gender and racial biases.However, political bias exists in LLMs and can lead to polarization and other harms in downstream applications.In order to provide transparency to users, we advocate that there should be fine-grained and explainable measures of political biases generated by LLMs.Our proposed measure looks at different political issues such as reproductive rights and climate change, at both the content (the substance of the generation) and the style (the lexical polarity) of such bias.We measured the political bias in eleven opensourced LLMs and showed that our proposed framework is easily scalable to other topics and is explainable. Yejin Bang, Delong Chen, Nayeon Lee, Pascale Fung |
ACL (1) | 4 |
| 2024 | Cendol: Open Instruction-tuned Generative Large Language Models for Indonesian LanguagesabstractSamuel Cahyawijaya, Holy Lovenia, Fajri Koto, Rifki Afina Putri, Emmanuel Dave, Jhonson Lee, Nuur Shadieq, Wawan Cenggoro, Salsabil Maulana Akbar, Muhammad Ihza Mahendra, Dea Annisayanti Putri, Bryan Wilie, Genta Indra Winata, Alham Fikri Aji, Ayu Purwarianti, Pascale Fung. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Samuel Cahyawijaya, Holy Lovenia, Fajri Koto, Rifki Afina Putri, Tjeng Wawan Cenggoro, Jhonson Lee, Salsabil Maulana Akbar, Emmanuel Dave, Nuur Shadieq, Muhammad Ihza Mahendra, Dea Annisayanti Putri, Bryan Wilie, Genta Indra Winata, Alham Fikri Aji, Ayu Purwarianti, Pascale Fung |
ACL (1) | 16 |
| 2024 | Belief Revision: The Adaptability of Large Language Models ReasoningabstractThe capability to reason from text is crucial for real-world NLP applications.Real-world scenarios often involve incomplete or evolving data.In response, individuals update their beliefs and understandings accordingly.However, most existing evaluations assume that language models (LMs) operate with consistent information.We introduce Belief-R 1 , a new dataset designed to test LMs' belief revision ability when presented with new evidence.Inspired by how humans suppress prior inferences, this task assesses LMs within the newly proposed delta reasoning (∆R) framework.Belief-R features sequences of premises designed to simulate scenarios where additional information could necessitate prior conclusions drawn by LMs.We evaluate ∼30 LMs across diverse prompting strategies and found that LMs generally struggle to appropriately revise their beliefs in response to new information.Further, models adept at updating often underperformed in scenarios without necessary updates, highlighting a critical trade-off.These insights underscore the importance of improving LMs' adaptiveness to changing information, a step toward more reliable AI systems. Bryan Wilie, Samuel Cahyawijaya, Etsuko Ishii, Junxian He, Pascale Fung |
EMNLP | 5 |
| 2024 | From Assistants to Agents in the LLM EraabstractAI assistants have been in our lives since the 1990s. Building AI agents from LLMs has changed the way we do research, development and deployment of assistants as agents. Agents are AI systems with various degrees of reasoning and planning, that can anticipate user needs, autonomously figure out the steps to fulfil such needs, access knowledge and tools external to the LLM, to accomplish user tasks. They are the new AI assistants in the LLM era. In this talk, I will give an overview of the evolution of AI assistants to AI agents, and outline the challenges we still face today and the opportunities ahead for developing a Family of Agents that span over different modalities and interfaces. As different model architectures emerge in the future, LLMs will remain as an effective tool of building agents, replacing manual design. Pascale Fung |
ACM Multimedia | 1 |
| 2024 | LLMs Are Few-Shot In-Context Low-Resource Language LearnersabstractSamuel Cahyawijaya, Holy Lovenia, Pascale Fung. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Samuel Cahyawijaya, Holy Lovenia, Pascale Fung |
NAACL-HLT | 3 |
| 2024 | A Humanoid Robot Dialogue System Architecture Targeting Patient Interview TasksabstractHumanoid robots are promising approach to automating patient interviews routinely conducted by medical staff. Their human-like appearance enables them to use the full gamut of verbal and behavioral cues that are critical to a successful interview. On the other hand, anthropomorphism can induce expectations of human-level performance by the robot. Not meeting such expectations degrades the quality of interaction. Specifically, humans expect rich real-time interactions during speech exchange, such as backchanneling and barge-ins. The nature of the patient interview task differs from most other scenarios where task oriented dialogue systems have been used, as there is increased potential of engagement breakdown during interaction. We describe a dialogue system architecture that improves the performance of humanoid robots on the patient interview task. Our architecture adds a nested inner real-time control loop to improve the timeliness of the robot’s responses based on the notion of "stance", an elaboration of the concept of a "turn", common in most existing dialogue systems. It also expands the dialogue state to monitor not only task progress, but also human engagement. Experiments using a humanoid robot running our proposed architecture reveal improved performance on interview tasks in terms of the perceived timeliness of responses and users’ impressions of the system. Dingdong Liu, Yejin Bang, Ho Shu Chan, Rita Frieske, Hoo Choun Chung, Jay Nieles, Tianjia Zhang, Kien T. Pham 0001, Wai Yi Rosita Cheng, Yini Fang, Qifeng Chen 0001, Pascale Fung, Xiaojuan Ma, Bertram E. Shi |
RO-MAN | 13 |
| 2023 | Generating Hashtags for Short-form Videos with Guided SignalsabstractTiezheng Yu, Hanchao Yu, Davis Liang, Yuning Mao, Shaoliang Nie, Po-Yao Huang, Madian Khabsa, Pascale Fung, Yi-Chia Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tiezheng Yu, Hanchao Yu, Davis Liang, Yuning Mao, Shaoliang Nie, Po-Yao Huang 0001, Madian Khabsa, Pascale Fung, Yi-Chia Wang |
ACL (1) | 8 |
| 2023 | Plausible May Not Be Faithful: Probing Object Hallucination in Vision-Language Pre-trainingabstractLarge-scale vision-language pre-trained (VLP) models are prone to hallucinate non-existent visual objects when generating text based on visual information.In this paper, we systematically study the object hallucination problem from three aspects.First, we examine recent state-of-the-art VLP models, showing that they still hallucinate frequently, and models achieving better scores on standard metrics (e.g., CIDEr) could be more unfaithful.Second, we investigate how different types of image encoding in VLP influence hallucination, including region-based, grid-based, and patch-based.Surprisingly, we find that patch-based features perform the best and smaller patch resolution yields a non-trivial reduction in object hallucination.Third, we decouple various VLP objectives and demonstrate that token-level imagetext alignment and controlled generation are crucial to reducing hallucination.Based on that, we propose a simple yet effective VLP loss named ObjMLM to further mitigate object hallucination.Results show that it reduces object hallucination by up to 17.4% when tested on two benchmarks (COCO Caption for in-domain and NoCaps for out-of-domain evaluation). Wenliang Dai, Zihan Liu 0001, Ziwei Ji 0001, Dan Su 0003, Pascale Fung |
EACL | 5 |
| 2023 | NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local LanguagesabstractGenta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung, Timothy Baldwin, Jey Han Lau, Rico Sennrich, Sebastian Ruder. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Genta Indra Winata, Alham Fikri Aji, Samuel Cahyawijaya, Rahmad Mahendra, Fajri Koto, Ade Romadhony, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Pascale Fung |
EACL | 10 |
| 2023 | Contrastive Learning for Inference in DialogueabstractInference, especially those derived from inductive processes, is a crucial component in our conversation to complement the information implicitly or explicitly conveyed by a speaker.While recent large language models show remarkable advances in inference tasks, their performance in inductive reasoning, where not all information is present in the context, is far behind deductive reasoning.In this paper, we analyze the behavior of the models based on the task difficulty defined by the semantic information gap -which distinguishes inductive and deductive reasoning (Johnson-Laird, 1988, 1993).Our analysis reveals that the disparity in information between dialogue contexts and desired inferences poses a significant challenge to the inductive inference process.To mitigate this information gap, we investigate a contrastive learning approach by feeding negative samples.Our experiments suggest negative samples help models understand what is wrong and improve their inference generations. 1 Etsuko Ishii, Yan Xu 0012, Bryan Wilie, Ziwei Ji 0001, Holy Lovenia, Willy Chung, Pascale Fung |
EMNLP | 7 |
| 2023 | Improving Fairness and Robustness in End-to-End Speech Recognition Through Unsupervised ClusteringabstractThe challenge of fairness arises when Automatic Speech Recognition (ASR) systems do not perform equally well for all sub-groups of the population. In the past few years there have been many improvements in overall speech recognition quality, but without any particular focus on advancing Equality and Equity for all user groups for whom systems do not perform well. ASR fairness is therefore also a robustness issue. Meanwhile, data privacy also takes priority in production systems. In this paper, we present a privacy preserving approach to improve fairness and robustness of end-to-end ASR without using metadata, zip codes, or even speaker or utterance embeddings directly in training. We extract utterance level embeddings using a speaker ID model trained on a public dataset, which we then use in an unsupervised fashion to create acoustic clusters. We use cluster IDs instead of speaker utterance embeddings as extra features during model training, which shows improvements for all demographic groups and in particular for different accents. Irina-Elena Veliche, Pascale Fung |
ICASSP | 2 |
| 2023 | Diverse and Faithful Knowledge-Grounded Dialogue Generation via Sequential Posterior InferenceabstractThe capability to generate responses with diversity and faithfulness using factual knowledge is paramount for creating a human-like, trustworthy dialogue system. Common strategies either adopt a two-step paradigm, which optimizes knowledge selection and response generation separately, and may overlook the inherent correlation between these two tasks, or leverage conditional variational method to jointly optimize knowledge selection and response generation by employing an inference network. In this paper, we present an end-to-end learning framework, termed Sequential Posterior Inference (SPI), capable of selecting knowledge and generating dialogues by approximately sampling from the posterior distribution. Unlike other methods, SPI does not require the inference network or assume a simple geometry of the posterior distribution. This straightforward and intuitive inference procedure of SPI directly queries the response generation model, allowing for accurate knowledge selection and generation of faithful responses. In addition to modeling contributions, our experimental results on two common dialogue datasets (Wizard of Wikipedia and Holl-E) demonstrate that SPI outperforms previous strong baselines according to both automatic and human evaluation metrics. Yan Xu 0012, Deqian Kong, Dehong Xu, Ziwei Ji 0001, Bo Pang 0004, Pascale Fung, Ying Nian Wu |
ICML | 6 |
| 2023 | A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and InteractivityabstractYejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, Pascale Fung. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su 0003, Bryan Wilie, Holy Lovenia, Ziwei Ji 0001, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu 0012, Pascale Fung |
IJCNLP (1) | 13 |
| 2023 | NusaWrites: Constructing High-Quality Corpora for Underrepresented and Extremely Low-Resource LanguagesabstractSamuel Cahyawijaya, Holy Lovenia, Fajri Koto, Dea Adhista, Emmanuel Dave, Sarah Oktavianti, Salsabil Akbar, Jhonson Lee, Nuur Shadieq, Tjeng Wawan Cenggoro, Hanung Linuwih, Bryan Wilie, Galih Muridan, Genta Winata, David Moeljadi, Alham Fikri Aji, Ayu Purwarianti, Pascale Fung. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Samuel Cahyawijaya, Holy Lovenia, Fajri Koto, Dea Adhista, Emmanuel Dave, Sarah Oktavianti, Salsabil Maulana Akbar, Jhonson Lee, Nuur Shadieq, Tjeng Wawan Cenggoro, Hanung Wahyuning Linuwih, Bryan Wilie, Galih Pradipta Muridan, Genta Indra Winata, David Moeljadi, Alham Fikri Aji, Ayu Purwarianti, Pascale Fung |
IJCNLP (1) | 18 |
| 2023 | PICK: Polished & Informed Candidate Scoring for Knowledge-Grounded Dialogue SystemsabstractBryan Wilie, Yan Xu, Willy Chung, Samuel Cahyawijaya, Holy Lovenia, Pascale Fung. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Bryan Wilie, Yan Xu 0012, Willy Chung, Samuel Cahyawijaya, Holy Lovenia, Pascale Fung |
IJCNLP (1) | 6 |
| 2023 | Cross-Lingual Cross-Age Adaptation for Low-Resource Elderly Speech Emotion RecognitionabstractSpeech emotion recognition plays a crucial role in human-computer interactions. However, most speech emotion recognition research is biased toward English-speaking adults, which hinders its applicability to other demographic groups in different languages and age groups. In this work, we analyze the transferability of emotion recognition across three different languages-English, Mandarin Chinese, and Cantonese; and 2 different age groups-adults and the elderly. To conduct the experiment, we develop an English-Mandarin speech emotion benchmark for adults and the elderly, BiMotion, and a Cantonese speech emotion dataset, YueMotion. This study concludes that different language and age groups require specific speech features, thus making cross-lingual inference an unsuitable method. However, cross-group data augmentation is still beneficial to regularize the model, with linguistic distance being a significant influence on cross-lingual transferability. We release publicly release our code at https://github.com/HLTCHKUST/elderly_ser. Samuel Cahyawijaya, Holy Lovenia, Willy Chung, Rita Frieske, Zihan Liu 0001, Pascale Fung |
INTERSPEECH | 6 |
| 2023 | InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningabstractLarge-scale pre-training and instruction tuning have been successful at creating general-purpose language models with broad competence. However, building general-purpose vision-language models is challenging due to the rich input distributions and task diversity resulting from the additional visual input. Although vision-language pretraining has been widely studied, vision-language instruction tuning remains under-explored. In this paper, we conduct a systematic and comprehensive study on vision-language instruction tuning based on the pretrained BLIP-2 models. We gather 26 publicly available datasets, covering a wide variety of tasks and capabilities, and transform them into instruction tuning format. Additionally, we introduce an instruction-aware Query Transformer, which extracts informative features tailored to the given instruction. Trained on 13 held-in datasets, InstructBLIP attains state-of-the-art zero-shot performance across all 13 held-out datasets, substantially outperforming BLIP-2 and larger Flamingo models. Our models also lead to state-of-the-art performance when finetuned on individual downstream tasks (e.g., 90.7% accuracy on ScienceQA questions with image contexts). Furthermore, we qualitatively demonstrate the advantages of InstructBLIP over concurrent multimodal models. All InstructBLIP models are open-source. Wenliang Dai, Junnan Li 0001, Dongxu Li 0003, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li 0001, Pascale Fung, Steven C. H. Hoi |
NeurIPS | 8 |
| 2022 | QAConv: Question Answering on Informative ConversationsabstractThis paper introduces QAConv, 1 , a new question answering (QA) dataset that uses conversations as a knowledge source.We focus on informative conversations, including business emails, panel discussions, and work channels.Unlike open-domain and task-oriented dialogues, these conversations are usually long, complex, asynchronous, and involve strong domain knowledge.In total, we collect 34,608 QA pairs from 10,259 selected conversations with both human-written and machinegenerated questions.We use a question generator and a dialogue summarizer as auxiliary tools to collect and recommend questions.The dataset has two testing scenarios: chunk mode and full mode, depending on whether the grounded partial conversation is provided or retrieved.Experimental results show that stateof-the-art pretrained QA systems have limited zero-shot performance and tend to predict our questions as unanswerable.Our dataset provides a new training and evaluation testbed to facilitate QA on conversations research. Chien-Sheng Wu, Andrea Madotto, Wenhao Liu 0003, Pascale Fung, Caiming Xiong |
ACL (1) | 4 |
| 2022 | ToKen: Task Decomposition and Knowledge Infusion for Few-Shot Hate Speech DetectionabstractBadr AlKhamissi, Faisal Ladhak, Srinivasan Iyer, Veselin Stoyanov, Zornitsa Kozareva, Xian Li, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona Diab. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer 0001, Veselin Stoyanov, Zornitsa Kozareva, Xian Li 0003, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona T. Diab |
EMNLP | 7 |
| 2022 | QA4QG: Using Question Answering to Constrain Multi-Hop Question GenerationabstractMulti-hop question generation (MQG) aims to generate com-plex questions which require reasoning over multiple pieces of information of the input passage. Most existing work on MQG has focused on exploring graph-based networks to equip the traditional Sequence-to-sequence framework with reasoning ability. However, these models do not take full advantage of the constraint between questions and answers. Furthermore, studies on multi-hop question answering (QA) suggest that Transformers can replace the graph structure for multi-hop reasoning. Therefore, in this work, we propose a novel framework, QA4QG, a QA-augmented BART-based framework for MQG. It augments the standard BART model with an additional multi-hop QA module to further con-strain the generated question. Our results on the HotpotQA dataset show that QA4QG outperforms all state-of-the-art models, with an increase of 8 BLEU-4 and 8 ROUGE points compared to the best results previously reported. Our work suggests the advantage of introducing pre-trained language models and QA module for the MQG task. Dan Su 0003, Peng Xu 0008, Pascale Fung |
ICASSP | 3 |
| 2022 | CI-AVSR: A Cantonese Audio-Visual Speech Datasetfor In-car Command RecognitionabstractWith the rise of deep learning and intelligent vehicles, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to process general as well as car-related commands and perform corresponding actions, which eases driving and improves safety. However, there is a data scarcity issue for low resource languages, hindering the development of research and applications. In this paper, we introduce a new dataset, Cantonese In-car Audio-Visual Speech Recognition (CI-AVSR), for in-car command recognition in the Cantonese language with both video and audio data. It consists of 4,984 samples (8.3 hours) of 200 in-car commands recorded by 30 native Cantonese speakers. Furthermore, we augment our dataset using common in-car background noises to simulate real environments, producing a dataset 10 times larger than the collected one. We provide detailed statistics of both the clean and the augmented versions of our dataset. Moreover, we implement two multimodal baselines to demonstrate the validity of CI-AVSR. Experiment results show that leveraging the visual signal improves the overall performance of the model. Although our best model can achieve a considerable quality on the clean test set, the speech recognition quality on the noisy data is still inferior and remains an extremely challenging task for real in-car speech recognition systems. The dataset and code will be released at https://github.com/HLTCHKUST/CI-AVSR. Wenliang Dai, Samuel Cahyawijaya, Tiezheng Yu, Elham J. Barezi, Peng Xu 0008, Cheuk Tung Yiu, Rita Frieske, Holy Lovenia, Genta Indra Winata, Qifeng Chen 0001, Xiaojuan Ma, Bertram E. Shi, Pascale Fung |
LREC | 13 |
| 2022 | Korean Language Modeling via Syntactic GuideabstractWhile pre-trained language models play a vital role in modern language processing tasks, but not every language can benefit from them. Most existing research on pre-trained language models focuses primarily on widely-used languages such as English, Chinese, and Indo-European languages. Additionally, such schemes usually require extensive computational resources alongside a large amount of data, which is infeasible for less-widely used languages. We aim to address this research niche by building a language model that understands the linguistic phenomena in the target language which can be trained with low-resources. In this paper, we discuss Korean language modeling, specifically methods for language representation and pre-training methods. With our Korean-specific language representation, we are able to build more powerful language models for Korean understanding, even with fewer resources. The paper proposes chunk-wise reconstruction of the Korean language based on a widely used transformer architecture and bidirectional language representation. We also introduce morphological features such as Part-of-Speech (PoS) into the language understanding by leveraging such information during the pre-training. Our experiment results prove that the proposed methods improve the model performance of the investigated Korean language understanding tasks. Hyeondey Kim, Seonhoon Kim, Inho Kang, Nojun Kwak, Pascale Fung |
LREC | 5 |
| 2022 | ASCEND: A Spontaneous Chinese-English Dataset for Code-switching in Multi-turn ConversationabstractCode-switching is a speech phenomenon occurring when a speaker switches language during a conversation. Despite the spontaneous nature of code-switching in conversational spoken language, most existing works collect code-switching data from read speech instead of spontaneous speech. ASCEND (A Spontaneous Chinese-English Dataset) is a high-quality Mandarin Chinese-English code-switching corpus built on spontaneous multi-turn conversational dialogue sources collected in Hong Kong. We report ASCEND’s design and procedure for collecting the speech data, including annotations. ASCEND consists of 10.62 hours of clean speech, collected from 23 bilingual speakers of Chinese and English. Furthermore, we conduct baseline experiments using pre-trained wav2vec 2.0 models, achieving a best performance of 22.69% character error rate and 27.05% mixed error rate. Holy Lovenia, Samuel Cahyawijaya, Genta Indra Winata, Peng Xu 0008, Yan Xu 0012, Zihan Liu 0001, Rita Frieske, Tiezheng Yu, Wenliang Dai, Elham J. Barezi, Qifeng Chen 0001, Xiaojuan Ma, Bertram E. Shi, Pascale Fung |
LREC | 14 |
| 2022 | Automatic Speech Recognition Datasets in Cantonese: A Survey and New DatasetabstractAutomatic speech recognition (ASR) on low resource languages improves the access of linguistic minorities to technological advantages provided by artificial intelligence (AI). In this paper, we address the problem of data scarcity for the Hong Kong Cantonese language by creating a new Cantonese dataset. Our dataset, Multi-Domain Cantonese Corpus (MDCC), consists of 73.6 hours of clean read speech paired with transcripts, collected from Cantonese audiobooks from Hong Kong. It comprises philosophy, politics, education, culture, lifestyle and family domains, covering a wide range of topics. We also review all existing Cantonese datasets and analyze them according to their speech type, data source, total size and availability. We further conduct experiments with Fairseq S2T Transformer, a state-of-the-art ASR model, on the biggest existing dataset, Common Voice zh-HK, and our proposed MDCC, and the results show the effectiveness of our dataset. In addition, we create a powerful and robust Cantonese ASR model by applying multi-dataset learning on MDCC and Common Voice zh-HK. Tiezheng Yu, Rita Frieske, Peng Xu 0008, Samuel Cahyawijaya, Cheuk Tung Shadow Yiu, Holy Lovenia, Wenliang Dai, Elham J. Barezi, Qifeng Chen 0001, Xiaojuan Ma, Bertram E. Shi, Pascale Fung |
LREC | 12 |
| 2022 | NeuS: Neutral Multi-News Summarization for Mitigating Framing BiasabstractNayeon Lee, Yejin Bang, Tiezheng Yu, Andrea Madotto, Pascale Fung. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Nayeon Lee, Yejin Bang, Tiezheng Yu, Andrea Madotto, Pascale Fung |
NAACL-HLT | 5 |
| 2022 | Factuality Enhanced Language Models for Open-Ended Text GenerationabstractPretrained language models (LMs) are susceptible to generate text with nonfactual information. In this work, we measure and improve the factual accuracy of large-scale LMs for open-ended text generation. We design the FactualityPrompts test set and metrics to measure the factuality of LM generations. Based on that, we study the factual accuracy of LMs with parameter sizes ranging from 126M to 530B. Interestingly, we find that larger LMs are more factual than smaller ones, although a previous study suggests that larger LMs can be less truthful in terms of misconceptions. In addition, popular sampling algorithms (e.g., top-p) in open-ended text generation can harm the factuality due to the ``uniform randomness'' introduced at every sampling step. We propose the factual-nucleus sampling algorithm that dynamically adapts the randomness to improve the factuality of generation while maintaining quality. Furthermore, we analyze the inefficiencies of the standard training method in learning correct associations between entities from factual text corpus (e.g., Wikipedia). We propose a factuality-enhanced training method that uses TopicPrefix for better awareness of facts and sentence completion as the training objective, which can vastly reduce the factual errors. Nayeon Lee, Wei Ping, Peng Xu 0008, Mostofa Patwary, Pascale Fung, Mohammad Shoeybi, Bryan Catanzaro |
NeurIPS | 5 |
| 2021 | The Adapter-Bot: All-In-One Controllable Conversational ModelabstractIn this paper, we present the Adapter-Bot, a generative chat-bot that uses a fixed backbone conversational model such as DialGPT (Zhang et al. 2019) and triggers on-demand dialogue skills via different adapters (Houlsby et al. 2019). Each adapter can be trained independently, thus allowing a continual integration of skills without retraining the entire model. Depending on the skills, the model is able to process multiple knowledge types, such as text, tables, and graphs, in a seamless manner. The dialogue skills can be triggered automatically via a dialogue manager, or manually, thus allowing high-level control of the generated responses. At the current stage, we have implemented 12 response styles (e.g., positive, negative etc.), 6 goal-oriented skills (e.g. weather information, movie recommendation, etc.), and personalized and emphatic responses. Zhaojiang Lin, Andrea Madotto, Yejin Bang, Pascale Fung |
AAAI | 4 |
| 2021 | CrossNER: Evaluating Cross-Domain Named Entity RecognitionabstractCross-domain named entity recognition (NER) models are able to cope with the scarcity issue of NER samples in target domains. However, most of the existing NER benchmarks lack domain-specialized entity types or do not focus on a certain domain, leading to a less effective cross-domain evaluation. To address these obstacles, we introduce a cross-domain NER dataset (CrossNER), a fully-labeled collection of NER data spanning over five diverse domains with specialized entity categories for different domains. Additionally, we also provide a domain-related corpus since using it to continue pre-training language models (domain-adaptive pre-training) is effective for the domain adaptation. We then conduct comprehensive experiments to explore the effectiveness of leveraging different levels of the domain corpus and pre-training strategies to do domain-adaptive pre-training for the cross-domain task. Results show that focusing on the fractional corpus containing domain-specialized entities and utilizing a more challenging pre-training strategy in domain-adaptive pre-training are beneficial for the NER domain adaptation, and our proposed method can consistently outperform existing cross-domain NER baselines. Nevertheless, experiments also illustrate the challenge of this cross-domain NER task. We hope that our dataset and baselines will catalyze research in the NER domain adaptation area. The code and data are available at https://github.com/zliucr/CrossNER. Zihan Liu 0001, Yan Xu 0012, Tiezheng Yu, Wenliang Dai, Ziwei Ji 0001, Samuel Cahyawijaya, Andrea Madotto, Pascale Fung |
AAAI | 8 |
| 2021 | On the Importance of Word Order Information in Cross-lingual Sequence LabelingabstractCross-lingual models trained on source language tasks possess the capability to directly transfer to target languages. However, since word order variances generally exist in different languages, cross-lingual models that overfit into the word order of the source language could have sub-optimal performance in target languages. In this paper, we hypothesize that reducing the word order information fitted into the models can improve the adaptation performance in target languages. To verify this hypothesis, we introduce several methods to make models encode less word order information of the source language and test them based on cross-lingual word embeddings and the pre-trained multilingual model. Experimental results on three sequence labeling tasks (i.e., part-of-speech tagging, named entity recognition and slot filling tasks) show that reducing word order information injected into the model can achieve better zero-shot cross-lingual performance. Further analysis illustrates that fitting excessive or insufficient word order information into the model results in inferior cross-lingual performance. Moreover, our proposed methods can also be applied to strong cross-lingual models and further improve their performance. Zihan Liu 0001, Genta Indra Winata, Samuel Cahyawijaya, Andrea Madotto, Zhaojiang Lin, Pascale Fung |
AAAI | 6 |
| 2021 | Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel DataabstractWei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona Diab. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wei-Jen Ko, Ahmed El-Kishky, Adithya Renduchintala, Vishrav Chaudhary, Naman Goyal 0001, Francisco Guzmán, Pascale Fung, Philipp Koehn, Mona T. Diab |
ACL/IJCNLP (1) | 7 |
| 2021 | IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language GenerationabstractSamuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Xiaohong Li, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Khodra, Ayu Purwarianti, Pascale Fung. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Samuel Cahyawijaya, Genta Indra Winata, Bryan Wilie, Karissa Vincentio, Adhiguna Kuncoro, Sebastian Ruder, Zhi Yuan Lim, Syafri Bahar, Masayu Leylia Khodra, Ayu Purwarianti, Pascale Fung |
EMNLP (1) | 12 |
| 2021 | Zero-Shot Dialogue State Tracking via Cross-Task TransferabstractZhaojiang Lin, Bing Liu, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul Crook, Zhiguang Wang, Zhou Yu, Eunjoon Cho, Rajen Subba, Pascale Fung. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhaojiang Lin, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul A. Crook, Zhiguang Wang, Zhou Yu 0005, Eunjoon Cho, Rajen Subba, Pascale Fung |
EMNLP (1) | 11 |
| 2021 | Continual Learning in Task-Oriented Dialogue SystemsabstractAndrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul Crook, Bing Liu, Zhou Yu, Eunjoon Cho, Pascale Fung, Zhiguang Wang. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Andrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul A. Crook, Zhou Yu 0005, Eunjoon Cho, Pascale Fung, Zhiguang Wang |
EMNLP (1) | 9 |
| 2021 | Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive SummarizationabstractMultimodal abstractive summarization (MAS) models that summarize videos (vision modality) and their corresponding transcripts (text modality) are able to extract the essential information from massive multimodal data on the Internet.Recently, large-scale generative pretrained language models (GPLMs) have been shown to be effective in text generation tasks.However, existing MAS models cannot leverage GPLMs' powerful generation ability.To fill this research gap, we aim to study two research questions: 1) how to inject visual information into GPLMs without hurting their generation ability; and 2) where is the optimal place in GPLMs to inject the visual information?In this paper, we present a simple yet effective method to construct vision guided (VG) GPLMs for the MAS task using attention-based add-on layers to incorporate visual information while maintaining their original text generation ability.Results show that our best model significantly surpasses the prior state-of-the-art model by 5.7 ROUGE-1, 5.3 ROUGE-2, and 5.1 ROUGE-L scores on the How2 dataset (Sanabria et al., 2018), and our visual guidance method contributes 83.6% of the overall improvement.Furthermore, we conduct thorough ablation studies to analyze the effectiveness of various modality fusion methods and fusion locations.* * The two authors contribute equally.The code is available at: https://github.com/HLTCHKUST/VG-GPLMs Video Frames Transcript: so now we are going to go over some basics sheet music readings for the key of g flat major.so you noticed the key of g flat, when you are reading real books, there is going to be a treble cleft here.it is going to have 6 flats 1, b flat, e flat, a flat, d flat, g flat and c flat.so 6 flats equals key of g flat.[...] so if you have a flat and there is a natural sign, play the a. so go through the scale and you've got g flat, a flat, d flat, c, flat, d flat, e flat and f, so f is your only 9 flat note in the scale.(No mention of the piano) Reference Summary: learn how to read and write music intervals for improving your playing and improvisational skills on the piano in this free video clip series.Summary from Transcript (BART): learn tips on how to read and write intervals on sheet music in this free video clip on music theory and music lessons.Summary from Transcript+Video (VG-BART): learn how to sight read in the key of g flat for improving your playing and improvisational skills on the piano in this free video clip series. Tiezheng Yu, Wenliang Dai, Zihan Liu 0001, Pascale Fung |
EMNLP (1) | 4 |
| 2021 | Ethical and Technological Challenges of Conversational AI
Pascale Fung |
Interspeech | 1 |
| 2021 | Multimodal End-to-End Sparse Model for Emotion RecognitionabstractWenliang Dai, Samuel Cahyawijaya, Zihan Liu, Pascale Fung. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Wenliang Dai, Samuel Cahyawijaya, Zihan Liu 0001, Pascale Fung |
NAACL-HLT | 4 |
| 2021 | Towards Few-shot Fact-Checking via PerplexityabstractNayeon Lee, Yejin Bang, Andrea Madotto, Pascale Fung. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Nayeon Lee, Yejin Bang, Andrea Madotto, Pascale Fung |
NAACL-HLT | 4 |
| 2021 | On Unifying Misinformation DetectionabstractNayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma, Wen-tau Yih, Madian Khabsa. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Nayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma 0001, Scott Yih, Madian Khabsa |
NAACL-HLT | 4 |
| 2021 | AdaptSum: Towards Low-Resource Domain Adaptation for Abstractive SummarizationabstractState-of-the-art abstractive summarization models generally rely on extensive labeled data, which lowers their generalization ability on domains where such data are not available.In this paper, we present a study of domain adaptation for the abstractive summarization task across six diverse target domains in a low-resource setting.Specifically, we investigate the second phase of pre-training on large-scale generative models under three different settings: 1) source domain pre-training; 2) domain-adaptive pre-training; and 3) taskadaptive pre-training.Experiments show that the effectiveness of pre-training is correlated with the similarity between the pre-training data and the target domain task.Moreover, we find that continuing pre-training could lead to the pre-trained model's catastrophic forgetting, and a learning method with less forgetting can alleviate this issue.Furthermore, results illustrate that a huge gap still exists between the low-resource and high-resource settings, which highlights the need for more advanced domain adaptation methods for the abstractive summarization task. 1 Tiezheng Yu, Zihan Liu 0001, Pascale Fung |
NAACL-HLT | 3 |
| 2021 | Assessing Political Prudence of Open-domain ChatbotsabstractPolitically sensitive topics are still a challenge for open-domain chatbots.However, dealing with politically sensitive content in a responsible, non-partisan, and safe behavior way is integral for these chatbots.Currently, the main approach to handling political sensitivity is by simply changing such a topic when it is detected.This is safe but evasive and results in a chatbot that is less engaging.In this work, as a first step towards a politically safe chatbot, we propose a group of metrics for assessing their political prudence.We then conduct political prudence analysis of various chatbots and discuss their behavior from multiple angles through our automatic metric and human evaluation metrics.The testsets and codebase are released to promote research in this area.1 Yejin Bang, Nayeon Lee, Etsuko Ishii, Andrea Madotto, Pascale Fung |
SIGDIAL | 5 |
| 2021 | ERICA: An Empathetic Android Companion for Covid-19 QuarantineabstractOver the past year, research in various domains, including Natural Language Processing (NLP), has been accelerated to fight against the COVID-19 pandemic, yet such research has just started on dialogue systems.In this paper, we introduce an end-to-end dialogue system which aims to ease the isolation of people under self-quarantine.We conduct a control simulation experiment to assess the effects of the user interface, a web-based virtual agent called Nora vs. the android ERICA via a video call.The experimental results show that the android offers a more valuable user experience by giving the impression of being more empathetic and engaging in the conversation due to its nonverbal information, such as facial expressions and body gestures.Demo video available at https://youtu.be/PLPEBXLeKJI. Etsuko Ishii, Genta Indra Winata, Samuel Cahyawijaya, Divesh Lala, Tatsuya Kawahara, Pascale Fung |
SIGDIAL | 6 |
| 2020 | Learning to Classify the Wrong Answers for Multiple Choice Question Answering (Student Abstract)abstractMultiple-Choice Question Answering (MCQA) is the most challenging area of Machine Reading Comprehension (MRC) and Question Answering (QA), since it not only requires natural language understanding, but also problem-solving techniques. We propose a novel method, Wrong Answer Ensemble (WAE), which can be applied to various MCQA tasks easily. To improve performance of MCQA tasks, humans intuitively exclude unlikely options to solve the MCQA problem. Mimicking this strategy, we train our model with the wrong answer loss and correct answer loss to generalize the features of our model, and exclude likely but wrong options. An experiment on a dialogue-based examination dataset shows the effectiveness of our approach. Our method improves the results on a fine-tuned transformer by 2.7%. Hyeondey Kim, Pascale Fung |
AAAI | 2 |
| 2020 | CAiRE: An End-to-End Empathetic ChatbotabstractWe present CAiRE, an end-to-end generative empathetic chatbot designed to recognize user emotions and respond in an empathetic manner. Our system adapts the Generative Pre-trained Transformer (GPT) to empathetic response generation task via transfer learning. CAiRE is built primarily to focus on empathy integration in fully data-driven generative dialogue systems. We create a web-based user interface which allows multiple users to asynchronously chat with CAiRE. CAiRE also collects user feedback and continues to improve its response quality by discarding undesirable generations via active learning and negative training. Zhaojiang Lin, Peng Xu 0008, Genta Indra Winata, Farhad Bin Siddique, Zihan Liu 0001, Jamin Shin, Pascale Fung |
AAAI | 7 |
| 2020 | Attention-Informed Mixed-Language Training for Zero-Shot Cross-Lingual Task-Oriented Dialogue SystemsabstractRecently, data-driven task-oriented dialogue systems have achieved promising performance in English. However, developing dialogue systems that support low-resource languages remains a long-standing challenge due to the absence of high-quality data. In order to circumvent the expensive and time-consuming data collection, we introduce Attention-Informed Mixed-Language Training (MLT), a novel zero-shot adaptation method for cross-lingual task-oriented dialogue systems. It leverages very few task-related parallel word pairs to generate code-switching sentences for learning the inter-lingual semantics across languages. Instead of manually selecting the word pairs, we propose to extract source words based on the scores computed by the attention layer of a trained English task-related model and then generate word pairs using existing bilingual dictionaries. Furthermore, intensive experiments with different cross-lingual embeddings demonstrate the effectiveness of our approach. Finally, with very few word pairs, our model achieves significant zero-shot adaptation performance improvements in both cross-lingual dialogue state tracking and natural language understanding (i.e., intent detection and slot filling) tasks compared to the current state-of-the-art approaches, which utilize a much larger amount of bilingual data. Zihan Liu 0001, Genta Indra Winata, Zhaojiang Lin, Peng Xu 0008, Pascale Fung |
AAAI | 5 |
| 2020 | Coach: A Coarse-to-Fine Approach for Cross-domain Slot FillingabstractAs an essential task in task-oriented dialog systems, slot filling requires extensive training data in a certain domain.However, such data are not always available.Hence, cross-domain slot filling has naturally arisen to cope with this data scarcity problem.In this paper, we propose a Coarse-to-fine approach (Coach) for cross-domain slot filling.Our model first learns the general pattern of slot entities by detecting whether the tokens are slot entities or not.It then predicts the specific types for the slot entities.In addition, we propose a template regularization approach to improve the adaptation robustness by regularizing the representation of utterances based on utterance templates.Experimental results show that our model significantly outperforms state-of-theart approaches in slot filling.Furthermore, our model can also be applied to the cross-domain named entity recognition task, and it achieves better adaptation performance than other existing baselines.The code is available at https: //github.com/zliucr/coach. Zihan Liu 0001, Genta Indra Winata, Peng Xu 0008, Pascale Fung |
ACL | 4 |
| 2020 | Meta-Transfer Learning for Code-Switched Speech RecognitionabstractAn increasing number of people in the world today speak a mixed-language as a result of being multilingual.However, building a speech recognition system for code-switching remains difficult due to the availability of limited resources and the expense and significant effort required to collect mixed-language data.We therefore propose a new learning method, meta-transfer learning, to transfer learn on a code-switched speech recognition system in a low-resource setting by judiciously extracting information from high-resource monolingual datasets.Our model learns to recognize individual languages, and transfer them so as to better recognize mixed-language speech by conditioning the optimization on the codeswitching data.Based on experimental results, our model outperforms existing baselines on speech recognition and language modeling tasks, and is faster to converge. Genta Indra Winata, Samuel Cahyawijaya, Zhaojiang Lin, Zihan Liu 0001, Peng Xu 0008, Pascale Fung |
ACL | 6 |
| 2020 | MinTL: Minimalist Transfer Learning for Task-Oriented Dialogue SystemsabstractIn this paper, we propose Minimalist Transfer Learning (MinTL) to simplify the system design process of task-oriented dialogue systems and alleviate the over-dependency on annotated data.MinTL is a simple yet effective transfer learning framework, which allows us to plug-and-play pre-trained seq2seq models, and jointly learn dialogue state tracking and dialogue response generation.Unlike previous approaches, which use a copy mechanism to "carryover" the old dialogue states to the new one, we introduce Levenshtein belief spans (Lev), that allows efficient dialogue state tracking with a minimal generation length.We instantiate our learning framework with two pretrained backbones: T5 (Raffel et al., 2019) and BART (Lewis et al., 2019), and evaluate them on MultiWOZ.Extensive experiments demonstrate that: 1) our systems establish new state-of-the-art results on end-to-end response generation, 2) MinTL-based systems are more robust than baseline methods in the low resource setting, and they achieve competitive results with only 20% training data, and 3) Lev greatly improves the inference efficiency 1 . Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, Pascale Fung |
EMNLP (1) | 4 |
| 2020 | Cross-lingual Spoken Language Understanding with Regularized Representation AlignmentabstractDespite the promising results of current crosslingual models for spoken language understanding systems, they still suffer from imperfect cross-lingual representation alignments between the source and target languages, which makes the performance sub-optimal.To cope with this issue, we propose a regularization approach to further align word-level and sentence-level representations across languages without any external resource.First, we regularize the representation of user utterances based on their corresponding labels.Second, we regularize the latent variable model (Liu et al., 2019a) by leveraging adversarial training to disentangle the latent variables.Experiments on the cross-lingual spoken language understanding task show that our model outperforms current state-of-the-art methods in both few-shot and zero-shot scenarios, and our model, trained on a few-shot setting with only 3% of the target language training data, achieves comparable performance to the supervised training with all the training data. 1 Zihan Liu 0001, Genta Indra Winata, Peng Xu 0008, Zhaojiang Lin, Pascale Fung |
EMNLP (1) | 5 |
| 2020 | MEGATRON-CNTRL: Controllable Story Generation with External Knowledge Using Large-Scale Language ModelsabstractPeng Xu, Mostofa Patwary, Mohammad Shoeybi, Raul Puri, Pascale Fung, Anima Anandkumar, Bryan Catanzaro. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Peng Xu 0008, Mostofa Patwary, Mohammad Shoeybi, Raul Puri, Pascale Fung, Anima Anandkumar, Bryan Catanzaro |
EMNLP (1) | 5 |
| 2020 | Generating Empathetic Responses by Looking Ahead the User's SentimentabstractAn important aspect of human conversation difficult for machines is conversing with empathy, which is to understand the user's emotion and respond appropriately. Recent neural conversation models that attempted to generate empathetic responses either focused on conditioning the output to a given emotion, or incorporating the current user emotional state. However, these approaches do not factor in how the user would feel towards the generated response. Hence, in this paper, we propose Sentiment Look-ahead, which is a novel perspective for empathy that models the future user emotional state. In short, Sentiment Look-ahead is a reward function under a reinforcement learning framework that provides a higher reward to the generative model when the generated utterance improves the user's sentiment. We implement and evaluate three different possible implementations of sentiment look-ahead and empirically show that our proposed approach can generate significantly more empathetic, relevant, and fluent responses than other competitive baselines such as multitask learning. Jamin Shin, Peng Xu 0008, Andrea Madotto, Pascale Fung |
ICASSP | 4 |
| 2020 | Improving Spoken Question Answering Using Contextualized Word RepresentationabstractWhile question answering (QA) systems have witnessed great breakthroughs in reading comprehension (RC) tasks, spoken question answering (SQA) is still a much less investigated area. Previous work shows that existing SQA systems are limited by catastrophic impact of automatic speech recognition (ASR) errors [1] and the lack of large-scale real SQA datasets [2]. In this paper, we propose using contextualized word representations to mitigate the effects of ASR errors and pretraining on existing textual QA datasets to mitigate the data scarcity issue. New state-of-the-art results have been achieved using contextualized word representations on both the artificially synthesised and real SQA benchmark data sets, with 21.5 EM/18.96 F1 score improvement over the sub-word unit based baseline on the Spoken-SQuAD [1] data, and 13.11 EM/10.99 F1 score improvement on the ODSQA data [2]. By further fine-tuning pre-trained models with existing large scaled textual QA data, we obtained 38.12 EM/34.1 F1 improvement over the baseline of fine-tuned only on small sized real SQA data. Dan Su 0003, Pascale Fung |
ICASSP | 2 |
| 2020 | Lightweight and Efficient End-To-End Speech Recognition Using Low-Rank TransformerabstractHighly performing deep neural networks come at the cost of computational complexity that limits their practicality for deployment on portable devices. We propose the low-rank transformer (LRT), a memory-efficient and fast neural architecture that significantly reduces the parameters and boosts the speed of training and inference for end-to-end speech recognition. Our approach reduces the number of parameters of the network by more than 50% and speeds up the inference time by around 1.35x compared to the baseline transformer model. The experiments show that our LRT model generalizes better and yields lower error rates on both validation and test sets compared to an uncompressed transformer model. The LRT model outperforms those from existing works on several datasets in an end-to-end setting without using an external language model or acoustic data. Genta Indra Winata, Samuel Cahyawijaya, Zhaojiang Lin, Zihan Liu 0001, Pascale Fung |
ICASSP | 5 |
| 2020 | Learning Fast Adaptation on Cross-Accented Speech RecognitionabstractLocal dialects influence people to pronounce words of the same language differently from each other. The great variability and complex characteristics of accents create a major challenge for training a robust and accent-agnostic automatic speech recognition (ASR) system. In this paper, we introduce a cross-accented English speech recognition task as a benchmark for measuring the ability of the model to adapt to unseen accents using the existing CommonVoice corpus. We also propose an accent-agnostic approach that extends the model-agnostic meta-learning (MAML) algorithm for fast adaptation to unseen accents. Our approach significantly outperforms joint training in both zero-shot, few-shot, and all-shot in the mixed-region and cross-region settings in terms of word error rate. Copyright © 2020 ISCA Genta Indra Winata, Samuel Cahyawijaya, Zihan Liu 0001, Zhaojiang Lin, Andrea Madotto, Peng Xu 0008, Pascale Fung |
INTERSPEECH | 7 |
| 2020 | Getting To Know You: User Attribute Extraction from DialoguesabstractUser attributes provide rich and useful information for user understanding, yet structured and easy-to-use attributes are often sparsely populated. In this paper, we leverage dialogues with conversational agents, which contain strong suggestions of user information, to automatically extract user attributes. Since no existing dataset is available for this purpose, we apply distant supervision to train our proposed two-stage attribute extractor, which surpasses several retrieval and generation baselines on human evaluation. Meanwhile, we discuss potential applications (e.g., personalized recommendation and dialogue systems) of such extracted user attributes, and point out current limitations to cast light on future work. Chien-Sheng Wu, Andrea Madotto, Zhaojiang Lin, Peng Xu 0008, Pascale Fung |
LREC | 5 |
| 2020 | Multilingual and Interlingual Semantic Representations for Natural Language Processing: A Brief IntroductionabstractWe introduce the Computational Linguistics special issue on Multilingual and Interlingual Semantic Representations for Natural Language Processing. We situate the special issue’s five articles in the context of our fast-changing field, explaining our motivation for this project. We offer a brief summary of the work in the issue, which includes developments on lexical and sentential semantic representations, from symbolic and neural perspectives. Marta R. Costa-jussà, Cristina España-Bonet, Pascale Fung, Noah A. Smith |
Comput. Linguistics | 3 |
| 2019 | GlobalTrait: Personality Alignment of Multilingual Word EmbeddingsabstractWe propose a multilingual model to recognize Big Five Personality traits from text data in four different languages: English, Spanish, Dutch and Italian. Our analysis shows that words having a similar semantic meaning in different languages do not necessarily correspond to the same personality traits. Therefore, we propose a personality alignment method, GlobalTrait, which has a mapping for each trait from the source language to the target language (English), such that words that correlate positively to each trait are close together in the multilingual vector space. Using these aligned embeddings for training, we can transfer personality related training features from high-resource languages such as English to other low-resource languages, and get better multilingual results, when compared to using simple monolingual and unaligned multilingual embeddings. We achieve an average F-score increase (across all three languages except English) from 65 to 73.4 (+8.4), when comparing our monolingual model to multilingual using CNN with personality aligned embeddings. We also show relatively good performance in the regression tasks, and better classification results when evaluating our model on a separate Chinese dataset. Farhad Bin Siddique, Dario Bertero, Pascale Fung |
AAAI | 3 |
| 2019 | Personalizing Dialogue Agents via Meta-LearningabstractExisting personalized dialogue models use human designed persona descriptions to improve dialogue consistency.Collecting such descriptions from existing dialogues is expensive and requires hand-crafted feature designs.In this paper, we propose to extend Model-Agnostic Meta-Learning (MAML) (Finn et al., 2017) to personalized dialogue learning without using any persona descriptions.Our model learns to quickly adapt to new personas by leveraging only a few dialogue samples collected from the same user, which is fundamentally different from conditioning the response on the persona descriptions.Empirical results on Persona-chat dataset (Zhang et al., 2018) indicate that our solution outperforms non-metalearning baselines using automatic evaluation metrics, and in terms of human-evaluated fluency and consistency. Andrea Madotto, Zhaojiang Lin, Chien-Sheng Wu, Pascale Fung |
ACL (1) | 4 |
| 2019 | Transferable Multi-Domain State Generator for Task-Oriented Dialogue SystemsabstractOver-dependence on domain ontology and lack of knowledge sharing across domains are two practical and yet less studied problems of dialogue state tracking.Existing approaches generally fall short in tracking unknown slot values during inference and often have difficulties in adapting to new domains.In this paper, we propose a TRAnsferable Dialogue statE generator (TRADE) that generates dialogue states from utterances using a copy mechanism, facilitating knowledge transfer when predicting (domain, slot, value) triplets not encountered during training.Our model is composed of an utterance encoder, a slot gate, and a state generator, which are shared across domains.Empirical results demonstrate that TRADE achieves state-of-the-art joint goal accuracy of 48.62% for the five domains of Mul-tiWOZ, a human-human dialogue dataset.In addition, we show its transferring ability by simulating zero-shot and few-shot dialogue state tracking for unseen domains.TRADE achieves 60.58% joint goal accuracy in one of the zero-shot domains, and is able to adapt to few-shot cases without forgetting already trained domains.* Work partially done while the first author was an intern at Salesforce Research.Usr: I am looking for a cheap restaurant in the centre of the Chien-Sheng Wu, Andrea Madotto, Ehsan Hosseini-Asl, Caiming Xiong, Richard Socher, Pascale Fung |
ACL (1) | 6 |
| 2019 | Code-Switched Language Models Using Neural Based Synthetic Data from Parallel SentencesabstractTraining code-switched language models is difficult due to lack of data and complexity in the grammatical structure.Linguistic constraint theories have been used for decades to generate artificial code-switching sentences to cope with this issue.However, this require external word alignments or constituency parsers that create erroneous results on distant languages.We propose a sequence-to-sequence model using a copy mechanism to generate code-switching data by leveraging parallel monolingual translations from a limited source of code-switching data.The model learns how to combine words from parallel sentences and identifies when to switch one language to the other.Moreover, it captures code-switching constraints by attending and aligning the words in inputs, without requiring any external knowledge.Based on experimental results, the language model trained with the generated sentences achieves state-of-theart performance and improves end-to-end automatic speech recognition. Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu, Pascale Fung |
CoNLL | 4 |
| 2019 | MoEL: Mixture of Empathetic ListenersabstractZhaojiang Lin, Andrea Madotto, Jamin Shin, Peng Xu, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhaojiang Lin, Andrea Madotto, Jamin Shin, Peng Xu 0008, Pascale Fung |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Zero-shot Cross-lingual Dialogue Systems with Transferable Latent VariablesabstractZihan Liu, Jamin Shin, Yan Xu, Genta Indra Winata, Peng Xu, Andrea Madotto, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zihan Liu 0001, Jamin Shin, Yan Xu 0012, Genta Indra Winata, Peng Xu 0008, Andrea Madotto, Pascale Fung |
EMNLP/IJCNLP (1) | 7 |
| 2019 | Hierarchical Meta-Embeddings for Code-Switching Named Entity RecognitionabstractGenta Indra Winata, Zhaojiang Lin, Jamin Shin, Zihan Liu, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Genta Indra Winata, Zhaojiang Lin, Jamin Shin, Zihan Liu 0001, Pascale Fung |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Clickbait? Sensational Headline Generation with Auto-tuned Reinforcement LearningabstractPeng Xu, Chien-Sheng Wu, Andrea Madotto, Pascale Fung. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Peng Xu 0008, Chien-Sheng Wu, Andrea Madotto, Pascale Fung |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Learning Comment Generation by Leveraging User-generated DataabstractExisting models on open-domain comment generation are difficult to train, and they produce repetitive and uninteresting responses. The problem is due to multiple and contradictory responses from a single article, and by the rigidity of retrieval methods. To solve this problem, we propose a combined approach to retrieval and generation methods. We propose an attentive scorer to retrieve informative and relevant comments by leveraging user-generated data. Then, we use such comments, together with the article, as input for a sequence-to-sequence model with copy mechanism. We show the robustness of our model and how it can alleviate the aforementioned issue by using a large scale comment generation dataset. The result shows that the proposed generative model significantly outperforms strong baseline such as Seq2Seq with attention and Information Retrieval models by around 27 and 30 BLEU-1 points respectively. Zhaojiang Lin, Genta Indra Winata, Pascale Fung |
ICASSP | 3 |
| 2019 | Incorporate User Representation for Personal Question Answer Selection Using Siamese NetworkabstractMany natural language questions are inherently subjective. They can not be answered properly if we do not know the personal preferences of the answerer. For example, "Do you like cats?" There is no "the only correct answer" to this question. To answer it, the model has to be able to capture the persona of the answerers. However, the users usually do not answer different questions with equal chance. Instead, while some are answered with a high frequency, others are hardly answered by anyone. To deal with this imbalanced sparsity in data, we first introduce a Siamese Network to capture the preferences patterns of the users. Then the model is ensembled with an additional dense layer to predict the answers of the users. Applying to an online dating dataset, our approach achieves a high accuracy of 78.7%. Zihao Qi, Dario Bertero, Ian D. Wood, Pascale Fung |
ICASSP | 4 |
| 2019 | A Novel Repetition Normalized Adversarial Reward for Headline GenerationabstractWhile reinforcement learning can effectively improve language generation models, it often suffers from generating incoherent and repetitive phrases [1]. In this paper, we propose a novel repetition normalized adversarial reward to mitigate these problems. Our repetition penalized reward can greatly reduce the repetition rate and adversarial training mitigates generating incoherent phrases. Our model significantly outperforms the baseline model on ROUGE-1 (+3.24), ROUGE-L (+2.25), and a decreased repetition-rate (-4.98%). Peng Xu 0008, Pascale Fung |
ICASSP | 2 |
| 2019 | Exploring Perceived Emotional Intelligence of Personality-Driven Virtual Agents in Handling User ChallengesabstractAn effective virtual agent (VA) that serves humans not only completes tasks efficaciously, but also manages its interpersonal relationships with users judiciously. Although past research has studied how agents apologize or seek help appropriately, there lacks a comprehensive study of how to design an emotionally intelligent (EI) virtual agent. In this paper, we propose to improve a VA's perceived EI by equipping it with personality-driven responsive expression of emotions. We conduct a within-subject experiment to verify this approach using a medical assistant VA. We ask participants to observe how the agent (displaying a dominant or submissive trait, or having no personality) handles user challenges when issuing reminders and rate its EI. Results show that simply being emotionally expressive is insufficient for suggesting VAs as fully emotionally intelligent. Equipping such VAs with a consistent, distinctive personality trait (especially submissive) can convey a significantly stronger sense of EI in terms of the ability to perceive, use, understand, and manage emotions, and can better mitigate user challenges. Xiaojuan Ma, Emily P. Yang, Pascale Fung |
WWW | 3 |
| 2018 | Mem2Seq: Effectively Incorporating Knowledge Bases into End-to-End Task-Oriented Dialog SystemsabstractEnd-to-end task-oriented dialog systems usually suffer from the challenge of incorporating knowledge bases.In this paper, we propose a novel yet simple end-toend differentiable model called memoryto-sequence (Mem2Seq) to address this issue.Mem2Seq is the first neural generative model that combines the multihop attention over memories with the idea of pointer network.We empirically show how Mem2Seq controls each generation step, and how its multi-hop attention mechanism helps in learning correlations between memories.In addition, our model is quite general without complicated taskspecific designs.As a result, we show that Mem2Seq can be trained faster and attain the state-of-the-art performance on three different task-oriented dialog datasets. Andrea Madotto, Chien-Sheng Wu, Pascale Fung |
ACL (1) | 3 |
| 2018 | Improving Large-Scale Fact-Checking using Decomposable Attention Models and Lexical TaggingabstractFact-checking of textual sources needs to effectively extract relevant information from large knowledge bases.In this paper, we extend an existing pipeline approach to better tackle this problem.We propose a neural ranker using a decomposable attention model that dynamically selects sentences to achieve promising improvement in evidence retrieval F1 by 38.80%, with (×65) speedup compared to a TF-IDF method.Moreover, we incorporate lexical tagging methods into our pipeline framework to simplify the tasks and render the model more generalizable.As a result, our framework achieves promising performance on a large-scale fact extraction and verification dataset with speedup. Nayeon Lee, Chien-Sheng Wu, Pascale Fung |
EMNLP | 3 |
| 2018 | Reducing Gender Bias in Abusive Language DetectionabstractAbusive language detection models tend to have a problem of being biased toward identity words of a certain group of people because of imbalanced training datasets.For example, "You are a good woman" was considered "sexist" when trained on an existing dataset.Such model bias is an obstacle for models to be robust enough for practical use.In this work, we measure gender biases on models trained with different abusive language datasets, while analyzing the effect of different pre-trained word embeddings and model architectures.We also experiment with three bias mitigation methods: (1) debiased word embeddings, (2) gender swap data augmentation, and (3) fine-tuning with a larger corpus.These methods can effectively reduce gender bias by 90-98% and can be extended to correct model bias in other scenarios. Ji Ho Park, Jamin Shin, Pascale Fung |
EMNLP | 3 |
| 2018 | Attention-Based LSTM for Psychological Stress Detection from Spoken Language Using Distant SupervisionabstractWe propose a Long Short-Term Memory (LSTM) with attention mechanism to classify psychological stress from self-conducted interview transcriptions. We apply distant supervision by automatically labeling tweets based on their hash-tag content, which complements and expands the size of our corpus. This additional data is used to initialize the model parameters, and which it is fine-tuned using the interview data. This improves the model's robustness, especially by expanding the vocabulary size. The bidirectional LSTM model with attention is found to be the best model in terms of accuracy (74.1%) and f-score (74.3%). Furthermore, we show that distant supervision fine-tuning enhances the model's performance by 1.6% accuracy and 2.1% f-score. The attention mechanism helps the model to select informative words. Genta Indra Winata, Onno Kampman, Pascale Fung |
ICASSP | 3 |
| 2018 | End-to-End Dynamic Query Memory Network for Entity-Value Independent Task-Oriented DialogabstractIn this paper, we propose an end-to-end Dynamic Query Memory Network (DQMemNN) with a delexicalization mechanism for task-oriented dialog systems. The added dynamic component enables memory networks to capture the dialog's sequential dependencies by using a context-based query. Besides, the delexicalization mechanism reduces learning complexity and it alleviates the out-of-vocabulary entity problems. Experiments show that DQMemNN outperforms original end-to-end memory network models on bAbI full-dialog task by 3.1 % per-response and 39.3% per-dialog accuracy. In addition, the proposed framework achieves a promising average per-response accuracy of 99.7% and per-dialog accuracy of 97.8% without hand-crafted rules and features. Chien-Sheng Wu, Andrea Madotto, Genta Indra Winata, Pascale Fung |
ICASSP | 4 |
| 2017 | A first look into a Convolutional Neural Network for speech emotion detectionabstractWe propose a real-time Convolutional Neural Network model for speech emotion detection. Our model is trained from raw audio on a small dataset of TED talks speech data, manually annotated into three emotion classes: “Angry”, “Happy” and “Sad”. It achieves an average accuracy of 66.1%, 5% higher than a feature-based SVM baseline, with an evaluation time of few hundred milliseconds. We also provide an in-depth model visualization and analysis. We show how our neural network effectively activates during the speech sections of the waveform regardless of the emotion, ignoring the silence parts which do not contain information. On the frequency domain the CNN filters distribute throughout all the spectrum range, with higher concentration around the average pitch range related to that emotion. Each filter also activates at multiple frequency intervals, presumably due to the additional contribution of amplitude-related feature learning. Our work will allow faster and more accurate emotion detection modules for human-machine empathetic dialog systems and other related applications. Dario Bertero, Pascale Fung |
ICASSP | 2 |
| 2017 | A Note Based Query By Humming System Using Convolutional Neural NetworkabstractIn this paper, we propose a note-based query by humming (QBH) system with Hidden Markov Model (HMM) and Convolutional Neural Network (CNN) since note-based systems are much more efficient than the traditional frame-based systems. A note-based QBH system has two main components: humming transcription and candidate melody retrieval. For humming transcription, we are the first to use a hybrid model using HMM and CNN. We use CNN for its ability to learn the features directly from raw audio data and for being able to model the locality and variability often present in a note and we use HMM for handling the variability across the timeaxis. For candidate melody retrieval, we use locality sensitive hashing to narrow down the candidates for retrieval and dynamic time warping and earth mover's distance for the final ranking of the selected candidates. We show that our HMM-CNN humming transcription system outperforms other state of the art humming transcription systems by-2% using the transcription evaluation framework by Molina et. al and our overall query by humming system has a Mean Reciprocal Rank of 0:92 using the standard MIREX dataset, which is higher than other state of the art note-based query by humming systems. Naziba Mostafa, Pascale Fung |
INTERSPEECH | 2 |
| 2017 | Emojive! Collecting Emotion Data from Speech and Facial Expression Using Mobile Game App
Ji Ho Park, Nayeon Lee, Dario Bertero, Anik Dey, Pascale Fung |
INTERSPEECH | 5 |
| 2017 | Bilingual Word Embeddings for Cross-Lingual Personality Recognition Using Convolutional Neural NetsabstractWe propose a multilingual personality classifier that uses text data from social media and Youtube Vlog transcriptions, and maps them into Big Five personality traits using a Convolutional Neural Network (CNN). We first train unsupervised bilingual word embeddings from an English-Chinese parallel corpus, and use these trained word representations as input to our CNN. This enables our model to yield relatively high cross-lingual and multilingual performance on Chinese texts, after training on the English dataset for example. We also train monolingual Chinese embeddings from a large Chinese text corpus and then train our CNN model on a Chinese dataset consisting of conversational dialogue labeled with personality. We achieve an average F-score of 66.1 in our multilingual task compared to 63.3 F-score in cross-lingual, and 63.2 F-score in the monolingual performance. Farhad Bin Siddique, Pascale Fung |
INTERSPEECH | 2 |
| 2017 | Nora the Empathetic Psychologist
Genta Indra Winata, Onno Kampman, Yang Yang 0130, Anik Dey, Pascale Fung |
INTERSPEECH | 5 |
| 2016 | Towards Empathetic Human-Robot Interactions
Pascale Fung, Dario Bertero, Yan Wan 0004, Anik Dey, Ricky Ho Yin Chan, Farhad Bin Siddique, Yang Yang 0130, Chien-Sheng Wu, Ruixi Lin |
CICLing (2) | 1 |
| 2016 | Real-Time Speech Emotion and Sentiment Recognition for Interactive Dialogue SystemsabstractIn this paper, we describe our approach of enabling an interactive dialogue system to recognize user emotion and sentiment in realtime.These modules allow otherwise conventional dialogue systems to have "empathy" and answer to the user while being aware of their emotion and intent.Emotion recognition from speech previously consists of feature engineering and machine learning where the first stage causes delay in decoding time.We describe a CNN model to extract emotion from raw speech input without feature engineering.This approach even achieves an impressive average of 65.7% accuracy on six emotion categories, a 4.5% improvement when compared to the conventional feature based SVM classification.A separate, CNN-based sentiment analysis module recognizes sentiments from speech recognition results, with 82.5 Fmeasure on human-machine dialogues when trained with out-of-domain data. Dario Bertero, Farhad Bin Siddique, Chien-Sheng Wu, Yan Wan 0004, Ricky Ho Yin Chan, Pascale Fung |
EMNLP | 6 |
| 2016 | Predicting humor response in dialogues from TV sitcomsabstractWe propose a method to predict humor response in dialog using acoustic and language features. We use data from two popular TV sitcoms - "The Big Bang Theory" and "Seinfeld" - to predict how the audience responds to humor. Due to the sequentiality of humor response in dialogues we use a Conditional Random Field as classifier/predictor. Our method is relatively effective, with a maximum precision obtained of 72.1% in "Big Bang" and 60.2% in "Seinfeld". Experiments show that audio, speed, word and sentence length features are the most effective. This work is applicable to develop appropriate machine response empathetic to emotion in dialog, in addition to humor. Dario Bertero, Pascale Fung |
ICASSP | 2 |
| 2016 | Computational Approaches to Linguistic Code Switching
Mona T. Diab, Pascale Fung, Julia Hirschberg, Thamar Solorio |
INTERSPEECH | 2 |
| 2016 | Zara: An Empathetic Interactive Virtual Agent
Pascale Fung, Anik Dey, Farhad Bin Siddique, Ruixi Lin, Yang Yang 0130, Yan Wan 0004, Ricky Ho Yin Chan |
INTERSPEECH | 1 |
| 2016 | Deep Learning of Audio and Language Features for Humor Prediction
Dario Bertero, Pascale Fung |
LREC | 2 |
| 2016 | A Machine Learning based Music Retrieval and Recommendation System
Naziba Mostafa, Yan Wan 0004, Unnayan Amitabh, Pascale Fung |
LREC | 4 |
| 2016 | A Long Short-Term Memory Framework for Predicting Humor in DialoguesabstractWe propose a first-ever attempt to employ a Long Short-Term memory based framework to predict humor in dialogues. We analyze data from a popular TV-sitcom, whose canned laughters give an indication of when the audience would react. We model the setup-punchline relation of conversational humor with a Long Short-Term Memory, with utterance encodings obtained from a Convolutional Neural Network. Out neural network framework is able to improve the F-score of 8% over a Conditional Random Field baseline. We show how the LSTM effectively models the setup-punchline relation reducing the number of false positives and increasing the recall. We aim to employ our humor prediction model to build effective empathetic machine able to understand jokes. Dario Bertero, Pascale Fung |
HLT-NAACL | 2 |
| 2016 | Multimodal deep neural nets for detecting humor in TV sitcomsabstractWe propose a novel approach of combining acoustic and language features to predict humor in dialogues with a deep neural network. We analyze data from three popular TV-sitcoms whose canned laughters give an indication of when the audience would react. We model the setup-punchline sequential relation of conversational humor with a Long Short-Term Memory network, with utterance encodings obtained from two Convolutional Neural Networks, one to model word-level language features and the other to model frame-level acoustic and prosodic features. Our neural network framework is able to improve the F-score of over 5% over a Conditional Random Field baseline trained on a similar acoustic and language feature combination, achieving a much higher recall. It is also more effective over a language features-only setting, with a F-score of 10% higher. It also has a good generalization performance, reaching in most cases precision values of over 70% when trained and tested over different sitcoms. Dario Bertero, Pascale Fung |
SLT | 2 |
| 2015 | A comparison between a DNN and a CRF disfluency detection and reconstruction systemabstractWe propose to compare between a Deep Neural Network and a Conditional Random Field disfluency detection and reconstruc- tion system, both trained on the same features. Deep Neural Networks, despite an increasing popularity in a multitude of speech and language related tasks, were never applied to dis- fluency recognition. One of the most difficult classes of disflu- ency is false starts. We are interested in comparing these two approaches on recognition of different types of disfluency. Our experimental results over the SSR v2 corpus show that the DNN approach outperforms CRF slightly on repetition disfluencies. However, DNN exhibits a very low recall (17:6% compared to 26:3% of the CRF) over the more difficult false start recognition subtask. When applied to a corpus of sentences with only false starts, the two methods give both higher results (46:7% F-score for CRF and 44:2% F-score for DNN). We also propose to im- prove the overall results on false start by training our classifiers in two stages: The first to recognize non-false start errors, the second for false start only, and combining the two approaches using a simple voting algorithm. This allows us to obtain the best result of 52:2% F-score over false start. Dario Bertero, Ricky Ho Yin Chan, Pascale Fung |
INTERSPEECH | 4 |
| 2014 | Language Modeling with Functional Head Constraint for Code Switching Speech RecognitionabstractIn this paper, we propose novel structured language modeling methods for code mixing speech recognition by incorporating a well-known syntactic constraint for switching code, namely the Functional Head Constraint (FHC).Code mixing data is not abundantly available for training language models.Our proposed methods successfully alleviate this core problem for code mixing speech recognition by using bilingual data to train a structured language model with syntactic constraint.Linguists and bilingual speakers found that code switch do not happen between the functional head and its complements.We propose to learn the code mixing language model from bilingual data with this constraint in a weighted finite state transducer (WFST) framework.The constrained code switch language model is obtained by first expanding the search network with a translation model, and then using parsing to restrict paths to those permissible under the constraint.We implement and compare two approacheslattice parsing enables a sequential coupling whereas partial parsing enables a tight coupling between parsing and filtering.We tested our system on a lecture speech dataset with 16% embedded second language, and on a lunch conversation dataset with 20% embedded language.Our language models with lattice parsing and partial parsing reduce word error rates from a baseline mixed language model by 3.8% and 3.9% in terms of word error rate relatively on the average on the first and second tasks respectively.It outperforms the interpolated language model by 3.7% and 5.6% in terms of word error rate relatively, and outperforms the adapted language model by 2.6% and 4.6% relatively.Our proposed approach avoids making early decisions on codeswitch boundaries and is therefore more robust.We address the code switch data scarcity challenge by using bilingual data with syntactic structure. Ying Li 0125, Pascale Fung |
EMNLP | 2 |
| 2014 | Code switch language modeling with Functional Head ConstraintabstractIn this paper, we propose for the first time to incorporate the linguistically well-known Functional Head Constraint into a code switch language model for speech recognition. Under this constraint, code switch cannot occur between the functional head and its complements. The constrained code switching language model is obtained by first expanding the search network with a translation model; and then using lattice-based parsing to restrict paths to those permissible under the Functional Head Constraint. We tested our system on two tasks of code switch speech recognition, namely lecture speech recognition and lunch conversation recognition. Our system reduces word error rates (WER) from a baseline mixed language model by 3.72% relative in the first task; and by 5.85% in the second task. It reduces WER from an interpolated language model by 2.51% in the first task; and by 4.57% in the second task. All results are statistically significant. In addition, our method reduces WER for both the matrix language and the embedded language. Ying Li 0125, Pascale Fung |
ICASSP | 2 |
| 2014 | Co-Training for Classification of Live or Studio Music Recordings
Nicolas Auguin, Pascale Fung |
LREC | 2 |
| 2014 | A Hindi-English Code-Switching Corpus
Anik Dey, Pascale Fung |
LREC | 2 |
| 2014 | Efficient Sparse Banded Acoustic Models for Speech RecognitionabstractWe propose sparse banded acoustic models to significantly improve the recognition accuracy and reduce the computational cost of speech recognition systems. The sparse banded models are trained using a weighted lasso regularization. In addition, we propose new feature orders to reduce the bandwidth of sparse banded models in order to speed up computation. Experimental results on the Wall Street Journal data set show that sparse banded models significantly outperform diagonal and full covariance models by 9.5% and 15.1% relatively. Sparse banded models also run the fastest. The advantages of sparse banded models are also demonstrated on the collected Cantonese data set. Pascale Fung |
IEEE Signal Process. Lett. | 2 |
| 2014 | Discriminatively Trained Sparse Inverse Covariance Matrices for Speech RecognitionabstractWe propose to use acoustic models with sparse inverse covariance matrices to deal with the well-known over-fitting problem of discriminative training, especially when training data are limited. Compared with traditional diagonal or full covariance models, significant improvement by using sparse inverse covariance matrices has been achieved with maximum likelihood training. In state-of-the-art large vocabulary continuous speech recognition systems, discriminative training is commonly employed to achieve the best system performance. This paper investigates training acoustic models with sparse inverse covariance matrices using one of the most widely used discriminative training criteria-maximum mutual information (MMI). A lasso regularization term is added to the traditional objective function for MMI to automatically sparsify the inverse covariance matrices. The whole training process is then derived by maximizing the new objective function. This is achieved through iteratively maximizing a weak-sense auxiliary function. The final problem is shown to be a convex optimization problem and can be efficiently solved. Experimental results on the published Wall Street Journal and our collected Mandarin data sets show that the acoustic models with sparse inverse covariance matrices consistently outperform the conventional diagonal and full covariance models. Pascale Fung |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Improved mixed language speech recognition using asymmetric acoustic model and language model with code-switch inversion constraintsabstractWe propose an integrated framework for large vocabulary continuous mixed language speech recognition that handles the accent effect in the bilingual acoustic model and the inversion constraint well known to linguists in the language model. Our asymmetric acoustic model with phone set extension improves upon previous work by striking a balance between data and phonetic knowledge. Our language model improves upon previous work by (1) using the inversion constraint to predict code switching points in the mixed language and (2) integrating a code-switch prediction model, a translation model and a reconstruction model together. This integration means that our language model avoids the pitfall of propagated error that could arise from decoupling these steps. Finally, a WFST-based decoder integrates the acoustic models, code-switch language model and a monolingual language model in the matrix language all together. Our system reduces word error rate by 1.88% on a lecture speech corpus and by 2.43% on a lunch conversation corpus, with statistical significance, over the conventional bilingual acoustic model and interpolated language model. Ying Li 0125, Pascale Fung |
ICASSP | 2 |
| 2013 | Multimodal music emotion classification using AdaBoost with decision stumpsabstractWe propose using AdaBoost with decision stumps to implement multimodal music emotion classification (MEC) as a more appropriate alternative to the conventional SVMs. By modeling the presence or absence of salient phrases in the lyric texts and seeking for proper thresholds for certain audio signal features, it exploits interdependencies between aspects from both modalities in the multimodal MEC system to make the final classification. It can especially prevent the “short text problem” in lyrics. Our accuracy reached an average of 78.19% for classifying 3766 unique songs into 14 emotion categories, with a statistically significant improvement over the audio-only and lyrics-only monomodal MEC systems. We also show that the proposed AdaBoost with decision stumps method performs statistically better on multimodal MEC than the well-known SVM classifier, which only has an average accuracy of 72.08%. Dan Su 0003, Pascale Fung, Nicolas Auguin |
ICASSP | 2 |
| 2013 | Language modeling for mixed language speech recognition using weighted phrase extractionabstractTo train a code switching language model for mixed language speech recognition, we propose to assign weights to the sentence pairs in the parallel text data. The code switching language model which is composed of the code switching boundary prediction model, code switching translation model and reconstruction model is incorporated with a language for mixed language speech recognition. The code switching translation model which is trained using selected subsets of the sentence pairs in the parallel text data allows the decoder to make the decision whether a phrase is in the matrix language or in the embedded language. Moreover, we propose a weighting procedure while training the code switching translation model. We evaluate our methods on Mandarin-English code switching lecture speech and lunch conversations. Our proposed method reduces word error rate by a statistically significant 1.74% on the lecture speech, and by 1.29% on the lunch conversation over the conventional interpolated language model. Copyright © 2013 ISCA. Ying Li 0125, Pascale Fung |
INTERSPEECH | 2 |
| 2013 | Discriminatively trained sparse inverse covariance matrices for low resource acoustic modelingabstractWe propose a method to discriminatively train acoustic models with sparse inverse covariance (precision) matrices in order to improve the model robustness when training data is insufficient. Acoustic models with sparse inverse covariance matrices were previously proposed to address the problem of over-fitting when training data is inadequate. Since many of the entries of the inverse covariance matrices are driven to zero, the number of free parameters to be estimated is reduced. However, previously acoustic models using sparse inverse covariance matrices were trained using maximum likelihood (ML) training. It is well-known that discriminative training can further improve the recognition accuracy. Therefore, for the first time, we study the problem of training acoustic models with sparse inverse covariance matrices using the discriminative training method. An L1 regularization term is added to the traditional objective function for discriminative training to penalize complex models and to automatically sparsify the inverse covariance matrices. The new objective function is optimized by maximizing a weak-sense auxiliary function. Experimental results on the Wall Street Journal data set show that our method effectively regularizes the model complexity and allows more Gaussian components to be trained. Therefore it can better model the non-Gaussian nature of the speech feature vectors. Compared with the standard maximum mutual information (MMI) training method, our proposed method can significantly improve the recognition accuracy. Copyright © 2013 ISCA. Pascale Fung |
INTERSPEECH | 2 |
| 2013 | Cross-Lingual Language Modeling for Low-Resource Speech RecognitionabstractThis paper proposes using cross-lingual language modeling with syntactic information for low-resource speech recognition. We propose phrase-level transduction and syntactic reordering for transcribing a resource-poor language and translating it into a resource-rich language, if necessary. The phrase-level transduction is capable of performingn-mcross-lingual transduction. The syntactic reordering serves to model the syntactic discrepancies between the source and target languages. Our purpose is to leverage the statistics in a resource-rich language model to improve the language model of a resource-poor language and at the same time to improve low-resource speech recognition performance. We implement our cross-lingual language model using weighted finite-state transducers (WFSTs), and integrate it into a WFST-based speech recognition search space to output the transcriptions of both resource-poor and resource-rich languages. This creates an integrated speech transcription and translation framework. Evaluations on Cantonese speech transcription and Cantonese to standard Chinese translation tasks show that our proposed approach improves the system performance significantly, with up to 12.5% relative character error rate (CER) reduction over baseline language model interpolation, 6.6% relative CER reduction and 18.5% relative BLEU score improvement, compared to the best word-level transduction approach. Pascale Fung |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | Sparse Inverse Covariance Matrices for Low Resource Speech RecognitionabstractWe propose to use sparse inverse covariance matrices for acoustic model training when there is insufficient training data. Acoustic models trained with inadequate training data tend to over fit, generalizing poorly to unseen test data, especially when full covariance matrices are used. We address this problem by adding an L1 regularization term to the traditional objective function for maximum likelihood estimation, to penalize complex models. The structure of the inverse covariance matrices will be automatically sparsified using this new objective function. The Expectation Maximization algorithm is used to learn the parameters of the hidden Markov model using the new objective function. It is shown that the training procedures for all the hidden Markov model parameters are the same as that of maximum likelihood estimation except the inverse covariance matrices. The update equation for the inverse covariance matrices is concave and can be solved efficiently. Our experiments show that this proposed method can correctly learn the underlying correlations among the random variables of the speech feature vector. Experimental results on the Wall Street Journal data show that our proposed model significantly outperforms the diagonal covariance model and the full covariance model by 10.9% and 16.5% relative recognition accuracy, when only about 14 hours of training data are available. On our collected low resource language data-the Cantonese data set, the proposed model also significantly outperforms the diagonal covariance model and the full covariance model. Pascale Fung |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Code-Switch Language Model with Inversion Constraints for Mixed Language Speech Recognition
Ying Li 0125, Pascale Fung |
COLING | 2 |
| 2012 | Cross-Lingual Language Modeling with Syntactic Reordering for Low-Resource Speech Recognition
Pascale Fung |
EMNLP-CoNLL | 2 |
| 2012 | Phrase-level transduction model with reordering for spoken to written language transformationabstractThis paper proposes a first-ever phrase-level transduction model with reordering to transform colloquial speech directly to written-style transcription. This model is capable of performing n-m transductions. Our transduction model is trained from a parallel corpus of verbatim transcription and written-style transcription. Deletions, substitutions, insertions are well represented using this model. Inversion transduction cases can also be identified and represented. We implement our transduction model using weighted finite-state transducers (WFSTs), and integrate it into a WFST-based speech recognition search space to give both verbatim speaking-style and written-style transcriptions. Evaluations of our model on Cantonese speech to standard written Chinese show 11.59% relative Word Error Rate (WER) reduction over interpolated language model between Cantonese and standard Chinese speech, 5.72% relative WER reduction and 14.82% relative Bilingual Evaluation Understudy (BLEU) improvement over the word-level transduction model. Pascale Fung, Ricky Ho Yin Chan |
ICASSP | 2 |
| 2012 | Lowresource speech recognition with automatically learned sparse inverse covariance matricesabstractFull covariance acoustic models trained with limited training data generalize poorly to unseen test data due to a large number of free parameters. We propose to use sparse inverse covariance matrices to address this problem. Previous sparse inverse covariance methods never outperformed full covariance methods. We propose a method to automatically drive the structure of inverse covariance matrices to sparse during training. We use a new objective function by adding L1 regularization to the traditional objective function for maximum likelihood estimation. The graphic lasso method for the estimation of a sparse inverse covariance matrix is incorporated into the Expectation Maximization algorithm to learn parameters of HMM using the new objective function. Experimental results show that we only need about 25% of the parameters of the inverse covariance matrices to be nonzero in order to achieve the same performance of a full covariance system. Our proposed system using sparse inverse covariance Gaussians also significantly outperforms a system using full covariance Gaussians trained on limited data. Pascale Fung |
ICASSP | 2 |
| 2012 | sparse banded precision matrices for low resource speech recognitionabstractWe propose to use sparse banded precision matrices for speech recognition when there is insufficient training data. Previously we proposed a method to drive the structure of precision matrices to sparse under the HMM framework during training. The recognition accuracy of this compact model is shown to be better than full covariance or diagonal covariance systems. In this paper we propose to modify the penalization in order to automatically learn sparse banded precision matrices. This will enable the trained models to be even more compact. We demonstrate that the feature order is critical to the success of our proposed method. Using our proposed feature order, we can substantially reduce the right half-bandwidth of the banded precision matrices without sacrificing the recognition accuracy. This saves both memory and computation. Pascale Fung |
INTERSPEECH | 2 |
| 2012 | A Mandarin-English Code-Switching Corpus
Ying Li 0125, Pascale Fung |
LREC | 3 |
| 2012 | A Multilingual Natural Stress Emotion Database
Pascale Fung |
LREC | 3 |
| 2012 | Automatic Parliamentary Meeting Minute Generation Using Rhetorical Structure ModelingabstractIn this paper, we propose a one step rhetorical structure parsing, chunking and extractive summarization approach to automatically generate meeting minutes from parliamentary speech using acoustic and lexical features. We investigate how to use lexical features extracted from imperfect ASR transcriptions, together with acoustic features extracted from the speech itself, to form extractive summaries with the structure of meeting minutes. Each business item in the minute is modeled as a rhetorical chunk which consists of smaller rhetorical units. Principal Component Analysis (PCA) graphs of both acoustic and lexical features in meeting speech show clear self-clustering of speech utterances according to the underlying rhetorical state-for example acoustic and lexical feature vectors from the question and answer or motion of a parliamentary speech, are grouped together. We then propose a Conditional Random Fields (CRF)-based approach to perform both rhetorical structure modeling and extractive summarization in one step, by chunking, parsing and extraction of salient utterances. Extracted salient utterances are grouped under the labels of each rhetorical state, emulating meeting minutes to yield summaries that are more easily understandable by humans. We compare this approach to different machine learning methods. We show that our proposed CRF-based one step minute generation system obtains the best summarization performance both in terms of ROUGE-L F-measure at 74.5% and by human evaluation, at 77.5% on average. Justin Jian Zhang, Pascale Fung |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Rare Word Translation Extraction from Aligned Comparable Documents
Emmanuel Prochasson, Pascale Fung |
ACL | 2 |
| 2011 | Asymmetric acoustic modeling of mixed language speechabstractWe propose to improve speech recognition performance on speaker-independent, mixed language speech by asymmetric acoustic modeling. Mixed language is either inter-sentential code switching from the source matrix language to a foreign language or intra-sentential code mixing between the matrix language and embedded foreign words or phrases. In either case, the foreign phrases are pronounced by the matrix language speaker with varying degrees of accent. Our pro posed system using selective decision tree merging between a bilingual model and an accented embedded speech model outperforms previous approaches of either using a bilingual model with model retraining by 21.51%, or using adaptation by 15.88%. It outperforms all models on both code mixing and code switching cases. We successfully improved recognition on embedded foreign speech without degrading the performance on the matrix language speech. Ying Li 0125, Pascale Fung, Yi Liu 0022 |
ICASSP | 2 |
| 2011 | Automatic minute generation for parliamentary speech using conditional random fieldsabstractWe show a novel approach of automatically generating minutes style extractive summaries for parliamentary speech. Minutes are structured summaries consisting of sequences of business items with sub-summaries. We propose to model minute structures as a rhetorical syntax tree. We also propose to use a single Conditional Random Field classifier to carry out the chunking and parsing of a parliamentary speech according to this syntax tree, and extracting salient sentences, all in one step, to form a meeting minute automatically. We show that this one step minute generation system outperforms a more traditional two step system where a first classifier is used for chunking and parsing and a second classifier is used for sentence extraction, from 69.5% to 73.2% in ROUGE-L measure. We also show comparative results from different features in the classifier and found that acoustic features contribute similarly to the final performance as N-gram features from ASR output. Justin Jian Zhang, Pascale Fung, Ricky Ho Yin Chan |
ICASSP | 2 |
| 2011 | Mining Parallel Documents Using Low Bandwidth and High Precision CLIR from the Heterogeneous Web
Simon Shi, Pascale Fung, Emmanuel Prochasson, Chi-kiu Lo, Dekai Wu |
IJCNLP | 2 |
| 2010 | Unsupervised Synthesis of Multilingual Wikipedia Articles
Yuncong Chen, Pascale Fung |
COLING | 2 |
| 2010 | A Rhetorical Syntax-Driven Model for Speech Summarization
Justin Jian Zhang, Pascale Fung |
COLING | 2 |
| 2010 | Learning deep rhetorical structure for extractive speech summarizationabstractExtractive summarization of conference and lecture speech is useful for online learning and references. We show for the first time that deep(er) rhetorical parsing of conference speech is possible and helpful to extractive summarization task. This type of rhetorical structures is evident in the corresponding presentation slide structures. We propose using Hidden Markov SVM (HMSVM) to iteratively learn the rhetorical structure of the speeches and summarize them. We show that system based on HMSVM gives a 64.3% ROUGE-L F-measure, a 10.1% absolute increase in lecture speech summarization performance compared with the baseline system without rhetorical information. Our method equally outperforms the baseline with a conventional discourse feature. Our proposed approach is more efficient than and also improves upon a previous method of using shallow rhetorical structure parsing. Justin Jian Zhang, Pascale Fung |
ICASSP | 2 |
| 2010 | Extractive Speech Summarization Using Shallow Rhetorical Structure ModelingabstractWe propose an extractive summarization approach with a novel shallow rhetorical structure learning framework for speech summarization. One of the most under-utilized features in extractive summarization is hierarchical structure information-semantically cohesive units that are hidden in spoken documents. We first present empirical evidence that rhetorical structure is the underlying semantic information, which is rendered in linguistic and acoustic/prosodic forms in lecture speech. A segmental summarization method, where the document is partitioned into rhetorical units by K-means clustering, is first proposed to test this hypothesis. We show that this system produces summaries at 67.36% ROUGE-L F-measure, a 4.29% absolute increase in performance compared with that of the baseline system. We then propose Rhetorical-State Hidden Markov Models (RSHMMs) to automatically decode the underlying hierarchical rhetorical structure in speech. Tenfold cross validation experiments are carried out on conference speeches. We show that system based on RSHMMs gives a 71.31% ROUGE-L F-measure, a 8.24% absolute increase in lecture speech summarization performance compared with the baseline system without using RSHMM. Our method equally outperforms the baseline with a conventional discourse feature. We also present a thorough investigation of the relative contribution of different features and show that, for lecture speech, speaker-normalized acoustic features give the most contribution at 68.5% ROUGE-L F-measure, compared to 62.9% ROUGE-L F-measure for linguistic features, and 59.2% ROUGE-L F-measure for un-normalized acoustic features. This shows that the individual speaking style of each speaker is highly relevant to the summarization. Justin Jian Zhang, Ricky Ho Yin Chan, Pascale Fung |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Extractive speech summarization by active learningabstractIn this paper, we propose an active learning approach for feature-based extractive summarization of lecture speech. Most state-of-the-art speech summarization systems are trained by using a large amount of human reference summaries. Active learning targets to minimize human annotation efforts by automatically selecting a small amount of unlabeled examples for labeling. Our method chooses the unlabeled examples according to a combination of informativeness criterion and robustness criterion. Our summarization results show an increasing learning curve of ROUGE-L F-measure, from 0.44 to 0.54, consistently higher than that of using randomly chosen training samples. We also show that, by following the rhetorical structure in presentation slides, it is possible for humans to produce "gold standard" reference summaries with very high inter-labeler agreement. Justin Jian Zhang, Ricky Ho Yin Chan, Pascale Fung |
ASRU | 3 |
| 2009 | Can Semantic Role Labeling Improve SMT?
Dekai Wu, Pascale Fung |
EAMT | 2 |
| 2008 | Rhetorical-State Hidden Markov Models for extractive speech summarizationabstractWe propose an extractive summarization system with a novel non-generative probabilistic framework for speech summarization. One of the most underutilized features in extractive summarization is rhetorical information - semantically cohesive units that are hidden in spoken documents. We propose Rhetorical-State Hidden Markov Models (RSHMMs) to automatically decode this underlying structure in speech. We show that RSHMMs give a 71.69% ROUGE-L F-measure, a 5.69% absolute increase in lecture speech summarization performance compared to the baseline system without using RSHMM. It equally outperforms the baseline system with additional discourse features, showing that our RSHMM is a more refined improvement on the conventional discourse feature. Pascale Fung, Ricky Ho Yin Chan, Justin Jian Zhang |
ICASSP | 1 |
| 2008 | Using output probability distribution for oov word rejectionabstractThis paper proposes a method to calculate the confidence score for out-of-vocabulary (OOV) word verification based on the Output Probability Distribution (OPD) of phoneme HMMs. Compared with input vector for dynamic garbage model, OPD vector contains more information than the sorted probabilities. Confidence score of each phoneme is calculated by SVM with OPD vectors as input. Hypotheses are accepted or rejected based on this confidence score. Experimental results showed that the proposed method achieved lower EER in word verification task than the conventional dynamic garbage model. Shilei Huang, Pascale Fung |
SLT | 3 |
| 2008 | RSHMM++ for extractive lecture speech summarizationabstractWe propose an enhanced Rhetorical-State Hidden Markov Model (RSHMM++) for extracting hierarchical structural summaries from lecture speech. One of the most underutilized information in extractive summarization is rhetorical structure hidden in speech data. RSHMM++ automatically decodes this underlying information in order to provide better summaries. We show that RSHMM++ gives a 72.01% ROUGE-L F-measure, a 9.78% absolute increase in lecture speech summarization performance compared to the baseline system without using rhetorical information. We also propose Relaxed DTW for compiling reference summaries. Justin Jian Zhang, Shilei Huang, Pascale Fung |
SLT | 3 |
| 2007 | A Mandarin lecture speech transcription system for speech summarizationabstractThis paper introduces our work on mandarin lecture speech transcription. In particular, we present our work on a small database, which contains only 16 hours of audio data and 0.16 M words of text data. A range of experiments have been done to improve the performances of the acoustic model and the language model, these include adapting the lecture speech data to the reading speech data for acoustic modeling and the use of lecture conference paper, power points and similar domain web data for language modeling. We also study the effects of automatic segmentation, unsupervised acoustic model adaptation and language model adaptation in our recognition system. By using a 3timesRT multiple passes decoding strategy, we obtain 70.3% accuracy performance in our final system. Finally, we apply our speech transcription system into a SVM summarizer and obtain a ROUGE-L F-measure of 66.5%. Ricky Ho Yin Chan, Justin Jian Zhang, Pascale Fung |
ASRU | 3 |
| 2007 | Improving lecture speech summarization using rhetorical informationabstractWe propose a novel method of extractive summarization of lecture speech based on unsupervised learning of its rhetorical structure. We present empirical evidence showing that rhetorical structure is the underlying semantics which is then rendered in linguistic and acoustic/prosodic forms in lecture speech. We present a first thorough investigation of the relative contribution of linguistic versus acoustic features and show that, at least for lecture speech, what is said is more important than how it is said. We base our experiments on conference speeches and corresponding presentation slides as the latter is a faithful description of the rhetorical structure of the former. We find that discourse features from broadcast news are not applicable to lecture speech. By using rhetorical structure information in our summarizer, its performance reaches 67.87% ROUGE-L F-measure at 30% compression, surpassing all previously reported results. The performance is also superior to the 66.47% ROUGE-L F-measure of baseline summarization performance without rhetorical information. We also show that, despite a 29.7% character error rate in speech recognition, extractive summarization performs relatively well, underlining the fact that spontaneity in lecture speech does not affect the central meaning of lecture speech. Justin Jian Zhang, Ricky Ho Yin Chan, Pascale Fung |
ASRU | 3 |
| 2007 | A comparative study on speech summarization of broadcast news and lecture speechabstractWe carry out a comprehensive study of acoustic/prosodic, linguistic and structural features for speech summarization, contrasting two genres of speech, namely Broadcast News and Lecture Speech. We find that acoustic and structural features are more important for Broadcast News summarization due to the speaking styles of anchors and reporters, as well as typical news story flow. Due to the relatively small contribution of lexical features, Broadcast News summarization does not depend heavily on ASR accuracies. We use SVM based summarizer to select the best features for extractive summarization, and obtain state-of-the-art performances: ROUGE-L F-measure of 0.64 for Mandarin Broadcast News, and 0.65 for Mandarin Lecture Speech. In the case of Lecture Speech summarization where lexical features are more important, we make the surprising discovery that summarization performance is very high (0.63 ROUGE-L F-measure) even when the ASR accuracy is low (21 % CER). Justin Jian Zhang, Ricky Ho Yin Chan, Pascale Fung |
INTERSPEECH | 3 |
| 2006 | Robust Word Sense Translation by EM Learning of Frame Semantics
Pascale Fung, Benfeng Chen |
ACL | 1 |
| 2006 | Multi-accent Chinese speech recognitionabstractMultiple accents are often present in spontaneous Chinese Mandarin speech as most Chinese have learned Mandarin as a second language. We propose a method to handle multiple accents as well as standard speech in a speaker-independent system by merging auxiliary accent decision trees with standard trees and reconstruct the acoustic model. In our proposed method, tree structures and shape are modified according to accent-specific data while the parameter set of the baseline model remains the same. The effectiveness of this approach is evaluated on Cantonese and Wu accented, as well as standard Mandarin speech. Our method yields a significant 4.4% and 3.3% absolute word error rate reduction without sacrificing the performance on standard Mandarin speech. Yi Liu 0022, Pascale Fung |
INTERSPEECH | 2 |
| 2006 | Automatic Learning of Chinese English Semantic Structure MappingabstractWe present twin results on Chinese semantic parsing, with application to English-Chinese cross- lingual verb frame acquisition. First, we describe two new state-of-the-art Chinese shallow semantic parsers leading to an F-score of 82.01 on simultaneous frame and argument boundary identification and labeling. Subsequently, we propose a model that applies the separate Chinese and English semantic parsers to learn cross-lingual semantic verb frame argument mappings with 89.3% accuracy. The only training data needed by this cross-lingual learning model is a pair of non-parallel monolingual Propbanks, plus an unannotated parallel corpus. We also present the first reported controlled comparison of maximum entropy and SVM approaches to shallow semantic parsing, using the Chinese data. Pascale Fung, Zhaojun Wu, Dekai Wu |
SLT | 1 |
| 2006 | Aligning word senses using bilingual corporaabstractThe growing importance of multilingual information retrieval and machine translation has made multilingual ontologies extremely valuable resources. Since the construction of an ontology from scratch is a very expensive and time-consuming undertaking, it is attractive to consider ways of automatically aligning monolingual ontologies, which already exist for many of the world's major languages. Previous research exploited similarity in the structure of the ontologies to align, or manually created bilingual resources. These approaches cannot be used to align ontologies with vastly different structures and can only be applied to much studied language pairs for which expensive resources are already available. In this paper, we propose a novel approach to align the ontologies at the node level: Given a concept represented by a particular word sense in one ontology, our task is to find the best corresponding word sense in the second language ontology. To this end, we present a language-independent, corpus-based method that borrows from techniques used in information retrieval and machine translation. We show its efficiency by applying it to two very different ontologies in very different languages: the Mandarin Chinese HowNet and the American English WordNet. Moreover, we propose a methodology to measure bilingual corpora comparability and show that our method is robust enough to use noisy nonparallel bilingual corpora efficiently, when clean parallel corpora are not available. Marine Carpuat, Pascale Fung, Grace Ngai |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2005 | Inversion Transduction Grammar Constraints for Mining Parallel Sentences from Quasi-Comparable Corpora
Dekai Wu, Pascale Fung |
IJCNLP | 2 |
| 2005 | Acoustic and phonetic confusions in accented speech recognitionabstractAccented speech recognition is more challenging than standard speech recognition due to the effects of phonetic and acoustic confusions. Phonetic confusion in accented speech occurs when an expected phone is pronounced as a different one, which leads to erroneous recognition. Acoustic confusion occurs when the pronounced phone is found to lie acoustically between two baseform models and can be equally recognized as either one. We propose that it is necessary to analyze and model these confusions separately in order to improve accented speech recognition without degrading standard speech recognition. We propose using likelihood ratio test to measure phonetic confusion, and asymmetric acoustic distance to measure acoustic confusion. Only accent-specific phonetic units with low acoustic confusion are used in an augmented pronunciation dictionary, while phonetic models with high acoustic confusion are reconstructed using decision tree merging. Experimental results show that our approach is effective and superior to methods modeling phonetic confusion or acoustic confusion alone in accented speech, with a significant 5.7% absolute WER reduction, without degrading standard speech recognition. Yi Liu 0022, Pascale Fung |
INTERSPEECH | 2 |
| 2005 | Guest Editors Introduction: Machine Learning in Speech and Language Technologies
Pascale Fung, Dan Roth 0001 |
Mach. Learn. | 1 |
| 2004 | BiFrameNet: Bilingual Frame Semantics Resource Construction by Cross-lingual Induction
Pascale Fung, Benfeng Chen |
COLING | 1 |
| 2004 | Multi-level Bootstrapping For Extracting Parallel Sentences From a Quasi-Comparable Corpus
Pascale Fung, Percy Cheung |
COLING | 1 |
| 2004 | Mining Very-Non-Parallel Corpora: Parallel Sentence and Lexicon Extraction via Bootstrapping and E
Pascale Fung, Percy Cheung |
EMNLP | 1 |
| 2004 | A grammar-based Chinese to English speech translation system for portable devicesabstractPortable devices such as PDA phones and smart phones are increasingly popular. Many of these devices already have voice dialing capability. The next step is to offer more powerful personal-assistant features such as speech translation. In this paper, we propose a system that can translate speech commands in Chinese into English, in real-time, on small, portable devices with limited memory and computational power. We address the various computational and platform issues of speech recognition and translation on portable devices. We propose fixed-point computation, discrete front-end speech features, bi-phone acoustic models, grammar-based speech decoding, and unambiguous inversion transduction grammars for transfer-based translation. As a result, our speech translation system requires only 500k memory and a 200MHz CPU. Pascale Fung, Yi Liu 0022, Yihai Shen, Dekai Wu |
INTERSPEECH | 1 |
| 2004 | Translation Disambiguation in Mixed Language Queries
Percy Cheung, Pascale Fung |
Mach. Transl. | 2 |
| 2004 | A maximum-entropy chinese parser augmented by transformation-based learningabstractParsing, the task of identifying syntactic components, e.g., noun and verb phrases, in a sentence, is one of the fundamental tasks in natural language processing. Many natural language applications such as spoken-language understanding, machine translation, and information extraction, would benefit from, or even require, high accuracy parsing as a preprocessing step. Even though most state-of-the-art statistical parsers were initially constructed for parsing in English, most of them are not language-specific, in that they do not rely on properties of the language that are specific to English. Therefore, construction of a parser in a given language becomes a matter of retraining the statistical parameters with a Treebank in the corresponding language. The development of the Chinese treebank [Xia et al. 2000] spurred the construction of parsers for Chinese. However, Chinese as a language poses some unique problems for the development of a statistical parser, the most apparent being word segmentation. Since words in written Chinese are not delimited in the same way as in Western languages, the first problem that needs to be solved before an existing statistical method can be applied to Chinese is to identify the word boundaries. This is a step that is neglected by most pre-existing Chinese parsers, which assume that the input data has already been pre-segmented. This article describes a character-based statistical parser, which gives the best performance to-date on the Chinese treebank data. We augment an existing maximum entropy parser with transformation-based learning, creating a parser that can operate at the character level. We present experiments that show that our parser achieves results that are close to those achievable under perfect word segmentation conditions. Pascale Fung, Grace Ngai, Benfeng Chen |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2004 | State-dependent phonetic tied mixtures with pronunciation modeling for spontaneous speech recognitionabstractWe propose a method of incorporating pronunciation modeling into acoustic models with high discriminative power and low complexity to improve spontaneous speech recognition accuracy. Spontaneous speech contains a higher level of phonetic and acoustic confusions due to the larger degree of pronunciation variations caused by speaking rate, speaker style, speaking mode, speaker accent, etc. In general data-driven complexity-reduction methods without explicit modeling of pronunciation variations, the acoustic model is not robust enough to capture the flexible phonetic confusions and pronunciation variants in spontaneous speech. We propose a state-dependent phonetic tied-mixture (PTM) model with variable codebook size to improve the coverage of phonetic variations while maintaining model discriminative ability. Our state-dependent PTM model incorporates a state-level pronunciation model for better discrimination of phonetic and acoustic confusions, while reducing model complexity. Experimental results on the spontaneous speech part of Mandarin Broadcast News shows that our model outperforms state tying and mixture tying models by 2.46% and 3.51% absolute syllable error rate reduction, respectively, with comparable model complexity. After adding Gaussian sharing to the latter models, our proposed model still yields an additional 1% and 2.6% absolute syllable error rate reduction. In addition, unlike many complexity reduction methods, our method does not lead to any performance degradation on read speech. Yi Liu 0022, Pascale Fung |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | Triphone model reconstruction for Mandarin pronunciation variationsabstractThe high error rate of recognition accuracy in spontaneous speech is due in part to the poor modeling of pronunciations. In this paper, we propose modeling pronunciation variations through triphone model reconstruction. We first generate a partial change phone model (PCPM) to differentiate pronunciation variations. In order to improve the resolution of triphone models, PCPM is used as a hidden model and merged into the pre-trained acoustic model through model reconstruction. To avoid model confusion, auxiliary decision trees are established for triphone PCPM. The acoustic model reconstruction on triphones is equivalent to decision tree merging. The effectiveness of this approach is evaluated on the 1997 Hub4NE Mandarin Broadcast News Corpus (1997 MBN) with different styles of speech. It gives a significant 2.39% absolute syllable error rate reduction in spontaneous speech. Pascale Fung, Yi Liu 0022 |
ICASSP (1) | 1 |
| 2003 | Flooring the observation probability for robust ASR in impulsive noiseabstractImpulsive noise usually introduces sudden mismatches between the observation features and the acoustic models trained with clean speech, which drastically degrades the performance of automatic speech recognition (ASR) systems. This paper presents a novel method to directly suppress the adverse effect of impulsive noise on recognition. In this method, according to the noise sensitivity of each feature dimension, the observation vector is divided into several subvectors, each of which is assigned to a suitable flooring threshold. In recognition stage, observation probability of each feature sub-vector is floored at the Gaussian mixture level. Thus, the unreliable relative probability difference caused by impulsive noise is eliminated, and the expected correct state sequence recovers the priority of being chosen in decoding. Experimental evaluations on Aurora2 database show that the proposed method achieves the average error rate reduction (ERR) of 61.62% and 84.32% in simulated impulsive noise and machinegun noise environment, respectively, while maintaining high performance for clean speech recognition. Pei Ding, Bertram E. Shi, Pascale Fung |
INTERSPEECH | 3 |
| 2003 | Automatic phone set extension with confidence measure for spontaneous speech
Yi Liu 0022, Pascale Fung |
INTERSPEECH | 2 |
| 2003 | Modeling partial pronunciation variations for spontaneous Mandarin speech recognition
Yi Liu 0022, Pascale Fung |
Comput. Speech Lang. | 2 |
| 2002 | Identifying Concepts Across Languages: A First Step towards a Corpus-based Approach to Automatic Ontology Alignment
Grace Ngai, Marine Carpuat, Pascale Fung |
COLING | 3 |
| 2002 | Model partial pronunciation variations for spontaneous Mandarin speech recognitionabstractModeling pronunciation variations is a critical part of spontaneous Mandarin speech recognition. Such variations include both complete changes and partial changes. Complete changes can usually be modeled by using an alternate phone to replace the canonical phone. Partial changes, which cannot be modeled by conventional methods are variations within the phoneme and include diacritics. In this paper, we propose using partial change phone model (PCPM) as well as auxiliary decision tree to model partial changes. A detailed but robust model can be achieved by merging canonical model with PCPMs through Gaussian distribution reconstruction. The effectiveness of this approach was evaluated on the Hub4NE Mandarin Broadcast News Corpus. The syllable error rate decreased 2.39% absolutely with respect to the baseline. Yi Liu 0022, Pascale Fung |
INTERSPEECH | 2 |
| 2002 | Reducing pronunciation lexicon confusion and using more data without phonetic transcription for pronunciation modelingabstractThe multiple-pronunciation lexicon (MPL) is very important to model the pronunciation variations for spontaneous speech recognition. But the introduction of MPL brings out two problems. First, the MPL will increase the among-lexicon confusion and degrade the recognizer's performance. Second, the MPL needs more data with phonetic transcription so as to cover as many surface forms as possible. Accordingly, two solutions are proposed, they are the context-dependent weighting method and the iterative forced-alignment based transcription method. The use of them can compensate what the MPL causes and improve the overall performance. Experiments across a naturally spontaneous speech database show that the proposed methods are effective and better than other methods. Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, William J. Byrne |
INTERSPEECH | 3 |
| 2002 | Mandarin Pronunciation Modeling Based on CASS Corpus
Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, William J. Byrne |
J. Comput. Sci. Technol. | 3 |
| 2001 | Automatic generation of pronunciation lexicons for Mandarin spontaneous speechabstractPronunciation modeling for large vocabulary speech recognition attempts to improve recognition accuracy by identifying and modeling pronunciations that are not in the ASR systems pronunciation lexicon. Pronunciation variability in spontaneous Mandarin is studied using the newly created CASS corpus of phonetically annotated spontaneous speech. Pronunciation modeling techniques developed for English,are applied to this corpus to train pronunciation models which are then used for Mandarin broadcast news transcription. William J. Byrne, Veera Venkataramani, Terri Kamm, Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, Yi Liu 0022, Umar Ruhi |
ICASSP | 6 |
| 2001 | Estimating pronunciation variations from acoustic likelihood score for HMM reconstruction
Yi Liu 0022, Pascale Fung |
INTERSPEECH | 2 |
| 2001 | Modeling pronunciation variation using context-dependent weighting and b/s refined acoustic modelingabstractThe pronunciation variability is an important issue that must be faced with when developing practical automatic spontaneous speech recognition systems. By studying the initial/final (IF) characteristics of Chinese language and developing the Bayesian equation, we propose the concepts of generalized initial/final (GIF) and generalized syllable (GS), the GIF modeling method and the IF-GIF modeling method, as well as the context-dependent pronunciation weighting method. By using these approaches, the IF-GIF modeling reduces the Chinese syllable error rate (SER) by 6.3 % and 4.2 % compared with the GIF modeling and IF modeling respectively when the language modeling, such as syllable or word N-gram, is not used. 1. Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, William J. Byrne |
INTERSPEECH | 3 |
| 2000 | Residual noise compensation for robust speech recognition in nonstationary noiseabstractWe present a model-based noise compensation algorithm for robust speech recognition in nonstationary noisy environments. The effect of noise is split into a stationary part, compensated by parallel model combination, and a time varying residual. The evolution of residual noise parameters is represented by a set of state space models. The state space models are updated by Kalman prediction and the sequential maximum likelihood algorithm. Prediction of residual noise parameters from different mixtures are fused, and the fused noise parameters are used to modify the linearized likelihood score of each mixture. Noise compensation proceeds in parallel with recognition. Experimental results demonstrate that the proposed algorithm improves recognition performance in highly nonstationary environments, compared with parallel model combination alone. Kaisheng Yao, Bertram E. Shi, Pascale Fung |
ICASSP | 3 |
| 2000 | CASS: a phonetically transcribed corpus of mandarin spontaneous speech
Thomas Fang Zheng, William J. Byrne, Pascale Fung, Terri Kamm, Yi Liu 0022, Zhanjiang Song, Umar Ruhi, Veera Venkataramani |
INTERSPEECH | 4 |
| 2000 | Modelling pronunciation variations in spontaneous Mandarin speechabstractPronunciation in spontaneous Mandarin speech tends to be much more variable than in read speech. In current recognition systems, pronunciation dictionaries usually only contain one standard pronunciation for each word, so that the amount of variability that can be modelled is very limited. Most recent research work for modelling variations in spontaneous speech focuses on the lexicon level, which can only solve intra-word variations. Inter-word variations cannot be modelled effectively. Chinese is monosyllabic and has simple syllable structure, giving rise to a high amount of pronunciation variations. In this paper, we propose two methods to model pronunciation variations in spontaneous Mandarin speech. First, we generate probability lexicon to model intra-syllable variations by using DP alignment algorithm between base form and surface strings. Second, we integrate variation probability into the decoder to model intra as well as inter-syllable variations. Experimental results show that modelling intra-syllable variation with a probability lexicon reduces syllable error rate by 0.85% (phone error rate reduction of 1.4%) while adding inter-syllable variation in addition reduces syllable error rate significantly by 4.76% (phone error rate reduction of 7.6%) compared to the baseline system. Yi Liu 0022, Pascale Fung |
INTERSPEECH | 2 |
| 2000 | MLLR-based accent model adaptation without accented dataabstractWhen the user has an accent different from what the automatic speech recognization system is trained with, the performance of the systems degrades. This is attributed to both acoustic and phonological differences between accents. The phonological differences between two accents are due to different phoneme inventories in two languages. Even for the same phoneme, foreigners and native speakers pronounce different sounds. Since accented data is rare but monolingual data is abundant, we propose using the accented speaker' s first language data directly instead of accented data in the second language for our purpose We propose adapting the native English phoneme models to accented phoneme models using first language data in MLLR adaptation. The baseline performance is 35.25% (phone accuracy) in using native English phone models to recognize Cantonese-accented English speech data. We compare accent adaptation by using accented data and source language data. On the average, using accented data for adaptation improves the phone accuracy by 69.98% while using source language data for adaptation improves the phone accuracy by 70.13%. This shows that both kinds of adaptation data give similar improvements. Therefore non-accented data can be used for adaptation. We can rapidly obtain an accent-adapted acoustic model without the need of collecting accented database. Wai Kat Liu, Pascale Fung |
INTERSPEECH | 2 |
| 2000 | An LLR-based technique for frame selection for GMM-based text-independent speaker identification
Pang Kuen Tsoi, Pascale Fung |
INTERSPEECH | 2 |
| 2000 | Principal mixture speaker adaptation for improved continuous speech recognitionabstractNowadays, almost all speaker-independent (SI) speech recognition systems use CDHMM with multivariate mixture Gaussian as observation density to cover speaker variabilities. It has been shown that given sufficient training data, the more mixtures are used in the HMM observation density, the better the system's perform. However, acoustic HMM with more Gaussian densities is more complex and slows down recognition speed. Another efficient way to handle speaker variation is to use speaker adaptation (SA). Yet, even though speaker adaptation of full multivariate mixture Gaussian densities can increase recognition accuracy, it does not improve recognition speed. In this paper, we introduce a principal mixture speaker adaptation method which reduces HMM complexity by choosing only the principle mixtures corresponding to a particular speaker's characteristics. We show that our method both improves recognition accuracy by 31.8% when compared to SI models, and reduces recognition speed by 30%, when compared to full mixture SA models. Pascale Fung |
INTERSPEECH | 2 |
| 1999 | Mixed Language Query DisambiguationabstractWe propose a mixed language query disambiguation approach by using co-occurrence information from monolingual data only. A mixed language query consists of words in a primary language and a secondary language. Our method translates the query into monolingual queries in either language. Two novel features for disambiguation, namely contextual word voting and 1-best contextual word, are introduced and compared to a baseline feature, the nearest neighbor. Average query translation accuracy for the two features are 81.37% and 83.72%, compared to the baseline accuracy of 75.50%. Pascale Fung, Xiaohuo Liu, Chi Shun Cheung |
ACL | 1 |
| 1999 | A more efficient and optimal LLR for decoding and verificationabstractWe propose a new confidence score for decoding and verification. Since the traditional log likelihood ratio (LLR) is borrowed from speaker verification technique, it may not be appropriate for decoding because we do not have a good modelling and definition of LLR for decoding/utterance verification. We have proposed a new formulation of LLR that can be used for decoding and verification task. Experimental results show that our proposed LLR can perform equally well compared with the result based on maximum likelihood in a decoding task. Also, we get an 5% improvement in decoding compared with traditional LLR. Kwok Leung Lam, Pascale Fung |
ICASSP | 2 |
| 1999 | Fast accent identification and accented speech recognitionabstractThe performance of speech recognition systems degrades when speaker accent is different from that in the training set. Accent-independent or accent-dependent recognition both require collection of more training data. In this paper, we propose a faster accent classification approach using phoneme-class models. We also present our findings in acoustic features sensitive to a Cantonese accent, and possibly other Asian language accents. In addition, we show how we can rapidly transform a native accent pronunciation dictionary to that for accented speech by simply using knowledge of the native language of the foreign speaker. The use of this accent-adapted dictionary reduces recognition error rate by 13.5%, similar to the results obtained from a longer, data-driven process. Wai Kat Liu, Pascale Fung |
ICASSP | 2 |
| 1999 | MAP-based cross-language adaptation augmented by linguistic knowledge: from English to Chinese
Pascale Fung, Ma Chi Yuen, Wai Kat Liu |
EUROSPEECH | 1 |
| 1999 | Decision tree-based triphones are robust and practical for mandarian speech recognitionabstractIn this paper, we describe The Attribute Logic Engine (ALE) and enhancements to it that enable it to serve as a complete grammatical infrastructure for applications such as spoken language translation. We indicate how ALEwas expanded and combined with off-the-shelf speech components to develop an application that translates English speech to German speech. The translation operates by way of a semantic representation based on typed feature structures, with information on thematic roles ("who did what to whom") and agreement information that can be used to guide search in less restricted domains and map expressions to more felicitous (but semantically equivalent) constructions in the target language than a more literal, surface-oriented method would admit. Yi Liu 0022, Pascale Fung |
EUROSPEECH | 2 |
| 1999 | A monolingual semantic decoder based on word sense disambiguation for mixed language understanding
Xiaohuo Liu, Pascale Fung, Chi Shun Cheung |
EUROSPEECH | 2 |
| 1999 | Liftered forward masking procedure for robust digits recognitionabstractUsing TI digits recognition experiments, we show that a combination of two dynamic speech features, Liftered Forward Masked (LFM) MFCC and 2-D cepstrum, can improve system robustness to additive Volvo noise while maintaining system performance comparable to standard MFCC features in clean conditions. Through experiments, we show that the information extracted by forward masking and by the 2D cepstrum are in some sense orthogonal. By combining the LFM MFCC and the 2-D cepstrum plus \\Delta 2-D cepstrum, we achieve a recognition rate above 90% on the TI connected digits task, even in additive Volvo noise condition with SNR as low as 0dB. This corresponds to a SNR gain over 30dB compared with standard MFCC plus dynamic and acceleration coefficients. Kaisheng Yao, Bertram E. Shi, Pascale Fung |
EUROSPEECH | 3 |
| 1998 | SALSA version 1.0: a speech-based web browser for hong kong EnglishabstractIn this paper, we present a prototype speech-based Web browser, SALSA1.0, and describe some of the research issues we need to address while building this system for Hong Kong users. SALSA1.0 allows the user to speak English command words as well as partial or complete link names on any page. The research issues involved in building SALSA1.0 are mainly (1) how to handle large accent variations and mixed-language and (2) how to handle unknown words, especially proper names, in Web links. The recognition engine for SALSA1.0 is trained on WSJ data, and then retrained on a small amount of Hong Kong accent WSJ data to handle accent variations. An edit-distance algorithm is used to replace all unknown words by the closest known word in the word network for recognition. With these methods, link name recognition rate is at 91.20% for links without unknown words, and 82.40% for links with unknown words. SALSA is currently being developed into a multilingual, natural language-based Intranet service provider for HKUST campus information access. Pascale Fung, Chi Shun Cheung, Kwok Leung Lam, Wai Kat Liu, Yuen Yee Lo |
ICSLP | 1 |
| 1997 | A Technical Word- and Term-Translation Aid Using Noisy Parallel Corpora across Language Groups
Pascale Fung, Kathy McKeown |
Mach. Transl. | 1 |
| 1996 | Domain word translation by space-frequency analysis of context length histogramsabstractWe report a new statistical feature relating a bilingual word pair in a non-parallel English-Chinese corpus. It is found that the lengths of context segments of a word are closely correlated to that of the translation, even when the corpus is non-parallel, i.e., monolingual texts which are not translations of each other. The context segment length histogram of a word has a characteristic pattern and corresponds to that of its translation. If a word appears most frequently in long segments, its translation is found to be most likely occurring in long segments. One way to match these histograms is to first extract their salient shape characteristics by space-frequency analysis and then match them against each other using dynamic time warping. The results of matching can be used in combination with other statistical features to bootstrap a word or term translation algorithm from non-parallel corpora. Pascale Fung |
ICASSP | 1 |
| 1995 | A Pattern Matching Method for Finding Noun and Proper Noun Translations from Noisy Parallel CorporaabstractWe present a pattern matching method for compiling a bilingual lexicon of nouns and proper nouns from unaligned, noisy parallel texts of Asian/Indo-European language pairs. Tagging information of one language is used. Word frequency and position information for high and low frequency words are represented in two different vector forms for pattern matching. New anchor point finding and noise elimination techniques are introduced. We obtained a 73.1% precision. We also show how the results can be used in the compilation of domain-specific noun phrases. Pascale Fung |
ACL | 1 |
| 1994 | K-vec: A New Approach for Aligning Parallel Texts
Pascale Fung, Kenneth Church 0001 |
COLING | 1 |
| 1993 | The BBN/HARC spoken language understanding system
Madeleine Bates, Robert J. Bobrow, Pascale Fung, Robert Ingria, Francis Kubala, John Makhoul, Long Nguyen 0001, Richard M. Schwartz, David Stallard |
ICASSP (2) | 3 |
| 1993 | The estimation of powerful language models from small and large corpora
Paul Placeway, Richard M. Schwartz, Pascale Fung, Long Nguyen 0001 |
ICASSP (2) | 3 |
| 1992 | Design and performance of HARC, the BBN spoken language understanding system
Madeleine Bates, Robert J. Bobrow, Pascale Fung, Robert Ingria, Francis Kubala, John Makhoul, Long Nguyen 0001, Richard M. Schwartz, David Stallard |
ICSLP | 3 |
| 1991 | Unsupervised speaker normalization by speaker Markov model converter for speaker-independent speech recognition
Pascale Fung, Tatsuya Kawahara, Shuji Doshita |
EUROSPEECH | 1 |