VLDB 2026 Research / reviewers in the wild / expert
AiTi Aw
dblp:40/357 · also Ai Ti Aw
· DBLP profile ↗
49ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0002-2347-0879ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 2 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaMCoT: Rethinking Cross-Lingual Factual Reasoning Through Adaptive Multilingual Chain-of-ThoughtabstractLarge language models (LLMs) have shown impressive multilingual capabilities through pretraining on diverse corpora. While these models show strong reasoning abilities, their performance varies significantly across languages due to imbalanced training data distribution. Existing approaches using sample-level translation for extensive multilingual pretraining and cross-lingual tuning face scalability challenges and often fail to capture nuanced reasoning processes across languages. In this paper, we introduce **AdaMCoT** (Adaptive Multilingual Chain-of-Thought), a framework that enhances multilingual factual reasoning by dynamically routing thought processes in intermediary “thinking languages” before generating target-language responses. AdaMCoT leverages a language-agnostic core and incorporates an adaptive, reward-based mechanism for selecting optimal reasoning pathways without requiring additional pretraining. Our comprehensive evaluation across multiple benchmarks demonstrates substantial improvements in both factual reasoning quality and cross-lingual consistency, with particularly strong performance gains in low-resource language settings. An in-depth analysis of the model’s hidden states and semantic space further elucidates the underlying mechanism of our method. The results suggest that adaptive reasoning paths can effectively bridge the performance gap between high- and low-resource languages while maintaining cultural and linguistic nuances. Zhengyuan Liu, Tarun Kumar Vangani, Bowei Zou, Xiyan Tao, AiTi Aw, Nancy F. Chen, Roy Ka-Wei Lee |
AAAI | 8 |
| 2026 | Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM PerformanceabstractMultilingual Large Language Models (LLMs) struggle with cross-lingual tasks due to data imbalances between high-resource and low-resource languages and the monolingual bias in pre-training. Existing methods, such as bilingual fine-tuning and contrastive alignment, improve cross-lingual performance but often require extensive parallel data or suffer from instability. To address these challenges, we introduce a Cross-Lingual Mapping Task in the pre-training phase, which enhances cross-lingual alignment without compromising monolingual fluency. Our approach bi-directionally maps languages within the LLM’s embedding space, improving both language generation and comprehension. We further introduce a Language Alignment Coefficient to robustly quantify cross-lingual consistency, even in limited-data scenarios. Experimental results on machine translation (MT), cross-lingual natural language understanding (CLNLU), and cross-lingual question answering (CLQA) show that our model achieves up to 11.9 BLEU score gains in MT, an increase of 6.72 in CLQA BERTScore-Precision and more than a 5% increase in CLNLU accuracy over strong multilingual baselines. Our findings highlight the potential of embedding cross-lingual objectives into pre-training, improving multilingual LLMs. Chang Liu 0072, Zhengyuan Liu, Muhammad Huzaifah 0001, AiTi Aw, Roy Ka-Wei Lee |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 7 |
| 2025 | Enhancing Event-centric News Cluster Summarization via Data Sharpening and Localization InsightsabstractThis paper tackles the challenges of clustering news articles by main events (MEs) and summarizing these clusters, focusing on diverse languages and localized contexts.Our approach consists of four key contributions.First, we investigate the role of dynamic clustering and the integration of various ME references, including event attributions extracted by language models (LMs), in enhancing event-centric clustering.Second, we propose a data-sharpening framework that optimizes the balance between information volume and entropy in input texts, thereby optimizing generated summaries on multiple indicators.Third, we fine-tune LMs with local news articles for cross-lingual temporal question-answering and text summarization, achieving notable improvements in capturing localized contexts.Lastly, we present the first cross-lingual dataset and comprehensive evaluation metrics tailored for the event-centric news cluster summarization pipeline.Our findings enhance the understanding of news summarization across N-gram, event-level coverage, and faithfulness, providing new insights into leveraging LMs for large-scale cross-lingual and localized news analysis. Longyin Zhang, Bowei Zou, AiTi Aw |
ACL (1) | 3 |
| 2025 | Incorporating Contextual Paralinguistic Understanding in Large Speech-Language ModelsabstractCurrent large speech language models (SpeechLLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextual content and paralinguistic cues. In this work, we propose two approaches to incorporate contextual paralinguistic information into model training: (1) an explicit method that provides paralinguistic metadata (e.g., emotion annotations) directly to the LLM, and (2) an implicit method that automatically generates novel training question-answer (QA) pairs using both categorical and dimensional emotion annotations alongside speech transcriptions. Our implicit method boosts performance (LLM-judged) by 38.41% on a human-annotated QA benchmark, reaching 46.02% when combined with the explicit approach, showing effectiveness in contextual paralinguistic understanding. We also validate the LLM judge by demonstrating its correlation with classification metrics, providing support for its reliability. Qiongqiong Wang, Hardik Bhupendra Sailor, Jeremy H. M. Wong, Tianchi Liu 0004, Muhammad Huzaifah 0001, Nancy F. Chen, AiTi Aw |
ASRU | 9 |
| 2025 | MNSC: Advancing Singlish Speech Understanding with Carefully Curated CorporaabstractSinglish, a Creole language rooted in English, is a key focus in linguistic research within multilingual and multicultural contexts. However, its spoken form remains underexplored, limiting insights into its linguistic structure and applications. To address this gap, we standardize and annotate the largest spoken Singlish corpus, introducing the Multitask National Speech Corpus (MNSC). These datasets support diverse tasks, including Automatic Speech Recognition (ASR), Spoken Question Answering (SQA), Spoken Dialogue Summarization (SDS), and Paralinguistic Question Answering (PQA). We release standardized splits and a human-verified test set to facilitate further research. Additionally, we propose SingAudioLLM, a multi-task multimodal model leveraging multimodal large language models to handle these tasks concurrently. Experiments reveal our models’ adaptability to the Singlish context, achieving state-of-the-art performance and outperforming prior models by 10–30% in comparison with other AudioLLMs and cascaded solutions1 Bin Wang 0040, Xunlong Zou, Yingxu He, Zhuohan Liu, Chengwei Wei, Nancy F. Chen, AiTi Aw |
ASRU | 9 |
| 2025 | Speech in-context learning of paralinguistic tasksabstractIn-context learning adapts a large language model to a new task, without computationally expensive parameter updates. This has previously been demonstrated for text tasks, as well as speech tasks that rely primarily on lexical information, such as recognition and translation. This paper proposes to extend this investigation to consider the ability of current open-source models to exhibit in-context learning on speech tasks that require an understanding of paralinguistic information. The tasks of stutter detection, pronunciation assessment, and speech emotion recognition are investigated. The results suggest that current open-source models already exhibit some degree of speech in-context learning on paralinguistic tasks. To more fully utilise available adaptation data, it is also proposed to overcome the finite number of in-context exemplars allowed by the model’s prompt length limit, through ensemble combination over multiple in-context learning runs that each use different exemplars. Jeremy H. M. Wong, Muhammad Huzaifah 0001, Nancy F. Chen, AiTi Aw |
ASRU | 4 |
| 2025 | Diversity and complementarity of speech encoders across diverse tasks in a multi-modal large language modelabstractA Large Language Model (LLM) can be extended to understand speech inputs by using a speech encoder to compute embeddings from the speech, which are then used with a text prompt. Diverse information is expressed in speech and a wide variety of tasks can be performed. Different speech encoders may specialise toward different information types and tasks. This complementarity can be leveraged upon by using multiple speech encoders. This paper presents a comprehensive analysis of the diversity and complementarity between open-source speech encoders, when used in a multi-modal LLM framework. Experiments identify the encoders that excel in each type of downstream task, thereby guiding future system design. The diversity between encoders is measured, showing that Whisper tends to behave more differently. Diversity between encoders is compared across tasks, showing that semantic tasks tend to yield more diverse predictions. Early and late fusion show that complementarity can yield improvements. Jeremy H. M. Wong, Muhammad Huzaifah 0001, Hardik B. Sailor, Kye Min Tan, Bin Wang 0040, Qiongqiong Wang, Xunlong Zou, Nancy F. Chen, AiTi Aw |
ASRU | 11 |
| 2025 | Improving Explainable Fact-Checking with Claim-Evidence CorrelationsabstractAutomatic fact-checking systems that employ large language models (LLMs) have achieved human-level performance in combating widespread misinformation. However, current LLM-based fact-checking systems fail to reveal the reasoning principles behind their decision-making for the claim verdict. In this work, we propose Correlation-Enhanced Explainable Fact-Checking (CorXFact), an LLM-based fact-checking system that simulates the reasoning principle of human fact-checkers for evidence-based claim verification: assessing and weighing the correlations between the claim and each piece of evidence. Following this principle, CorXFact enables efficient claim verification and transparent explanation generation. Furthermore, we contribute the CorFEVER test set to comprehensively evaluate the CorXFact system in claim-evidence correlation identification and claim verification in both closed-domain and real-world fact-checking scenarios. Experimental results show that our proposed CorXFact significantly outperforms four strong fact-checking baselines in claim authenticity prediction and verdict explanation. Bowei Zou, AiTi Aw |
COLING | 3 |
| 2025 | MoWE-Audio: Multitask AudioLLMs with Mixture of Weak EncodersabstractThe rapid advancements in large language models (LLMs) have significantly enhanced natural language processing capabilities, facilitating the development of AudioLLMs that process and understand speech and audio inputs alongside text. Existing AudioLLMs typically combine a pre-trained audio encoder with a pre-trained LLM, which are subsequently finetuned on specific audio tasks. However, the pre-trained audio encoder has constrained capacity to capture features for new tasks and datasets. To address this, we propose to incorporate mixtures of ‘weak’ encoders (MoWE) into the AudioLLM framework. MoWE supplements a base encoder with a pool of relatively lightweight encoders, selectively activated based on the audio input to enhance feature extraction without significantly increasing model size. Our empirical results demonstrate that MoWE effectively improves multi-task performance, broadening the applicability of AudioLLMs to more diverse audio tasks. Bin Wang 0040, Xunlong Zou, Zhuohan Liu, Yingxu He, Geyu Lin, Nancy F. Chen, AiTi Aw |
ICASSP | 9 |
| 2025 | Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
Qiongqiong Wang, Hardik B. Sailor, Tianchi Liu 0004, AiTi Aw |
INTERSPEECH | 4 |
| 2025 | AudioBench: A Universal Benchmark for Audio Large Language ModelsabstractBin Wang, Xunlong Zou, Geyu Lin, Shuo Sun, Zhuohan Liu, Wenyu Zhang, Zhengyuan Liu, AiTi Aw, Nancy F. Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Bin Wang 0040, Xunlong Zou, Geyu Lin, Zhuohan Liu, Zhengyuan Liu, AiTi Aw, Nancy F. Chen |
NAACL (Long Papers) | 8 |
| 2024 | CLFFRD: Curriculum Learning and Fine-grained Fusion for Multimodal Rumor DetectionabstractIn an era where rumors can propagate rapidly across social media platforms such as Twitter and Weibo, automatic rumor detection has garnered considerable attention from both academia and industry. Existing multimodal rumor detection models often overlook the intricacies of sample difficulty, e.g., text-level difficulty, image-level difficulty, and multimodal-level difficulty, as well as their order when training. Inspired by the concept of curriculum learning, we propose the Curriculum Learning and Fine-grained Fusion-driven multimodal Rumor Detection (CLFFRD) framework, which employs curriculum learning to automatically select and train samples according to their difficulty at different training stages. Furthermore, we introduce a fine-grained fusion strategy that unifies entities from text and objects from images, enhancing their semantic cohesion. We also propose a novel data augmentation method that utilizes linear interpolation between textual and visual modalities to generate diverse data. Additionally, our approach incorporates deep fusion for both intra-modality (e.g., text entities and image objects) and inter-modality (e.g., CLIP and social graph) features. Extensive experimental results demonstrate that CLFFRD outperforms state-of-the-art models on both English and Chinese benchmark datasets for rumor detection in social media. Fan Xu 0002, Bowei Zou, AiTi Aw, Huan Rong |
LREC/COLING | 4 |
| 2024 | Empowering Tree-structured Entailment Reasoning: Rhetorical Perception and LLM-driven InterpretabilityabstractThe study delves into the construction of entailment trees for science question answering (SQA), employing a novel framework termed Tree-structured Entailment Reasoning (TER). Current research on entailment tree construction presents significant challenges, primarily due to the ambiguities and similarities among candidate science facts, which considerably complicate the fact retrieval process. Moreover, the existing models exhibit limitations in effectively modeling the sequence of reasoning states, understanding the intricate relations between neighboring entailment tree nodes, and generating intermediate conclusions. To this end, we explore enhancing the TER performance from three aspects: First, improving retrieval capabilities by modeling and referring to the chained reasoning states; Second, enhancing TER by infusing knowledge that bridges the gap between reasoning types and rhetorical relations. Third, exploring a task-specific large language model tuning scheme to mitigate deficiencies in intermediate conclusion generation. Experiments on the English EntailmentBank demonstrate the effectiveness of the proposed methods in augmenting the quality of tree-structured entailment reasoning to a certain extent. Longyin Zhang, Bowei Zou, AiTi Aw |
LREC/COLING | 3 |
| 2024 | SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural ReasoningabstractBin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Yang Ding, AiTi Aw, Nancy Chen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Bin Wang 0040, Zhengyuan Liu, Fangkai Jiao, AiTi Aw, Nancy F. Chen |
NAACL-HLT | 6 |
| 2023 | Modeling What-to-ask and How-to-ask for Answer-unaware Conversational Question GenerationabstractXuan Long Do, Bowei Zou, Shafiq Joty, Tran Tai, Liangming Pan, Nancy Chen, Ai Ti Aw. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Do Xuan Long, Bowei Zou, Shafiq R. Joty, Anh Tran Tai, Liangming Pan, Nancy F. Chen, AiTi Aw |
ACL (1) | 7 |
| 2023 | Interview Evaluation: A Novel Approach for Automatic Evaluation of Conversational Question Answering ModelsabstractConversational Question Answering (CQA) aims to provide natural language answers to users in information-seeking dialogues.Existing CQA benchmarks often evaluate models using pre-collected human-human conversations.However, replacing the model-predicted dialogue history with ground truth compromises the naturalness and sustainability of CQA evaluation.While previous studies proposed using predicted history and rewriting techniques to address unresolved coreferences and incoherencies, this approach renders the question self-contained from the conversation.In this paper, we propose a novel automatic evaluation approach, interview evaluation.Specifically, ChatGPT acts as the interviewer (Q agent) with a set of carefully designed prompts, and the CQA model under test serves as the interviewee (A agent).During the interview evaluation, questions are dynamically generated by the Q agent to guide the A agent in predicting the correct answer through an interactive process.We evaluated four different models on QuAC and two models on CoQA in our experiments.The experiment results demonstrate that our interview evaluation has advantages over previous CQA evaluation approaches, particularly in terms of naturalness and coherence.The source code is made publicly available. Xibo Li, Bowei Zou, Yifan Fan, AiTi Aw, Yu Hong 0001 |
EMNLP | 5 |
| 2023 | DecoMT: Decomposed Prompting for Machine Translation Between Related Languages using Large Language ModelsabstractThis study investigates machine translation between related languages i.e., languages within the same family that share linguistic characteristics such as word order and lexical similarity.Machine translation through few-shot prompting leverages a small set of translation pair examples to generate translations for test sentences.This procedure requires the model to learn how to generate translations while simultaneously ensuring that token ordering is maintained to produce a fluent and accurate translation.We propose that for related languages, the task of machine translation can be simplified by leveraging the monotonic alignment characteristic of such languages.We introduce DecoMT, a novel approach of fewshot prompting that decomposes the translation process into a sequence of word chunk translations.Through automatic and human evaluation conducted on multiple related language pairs across various language families, we demonstrate that our proposed approach of decomposed prompting surpasses multiple established few-shot baseline approaches.For example, DecoMT outperforms the strong fewshot prompting BLOOM model with an average improvement of 8 chrF++ scores across the examined languages. Ratish Puduppully, Anoop Kunchukuttan, Raj Dabre, AiTi Aw, Nancy Chen |
EMNLP | 4 |
| 2022 | CoHS-CQG: Context and History Selection for Conversational Question GenerationabstractConversational question generation (CQG) serves as a vital task for machines to assist humans, such as interactive reading comprehension, through conversations. Compared to traditional single-turn question generation (SQG), CQG is more challenging in the sense that the generated question is required not only to be meaningful, but also to align with the provided conversation. Previous studies mainly focus on how to model the flow and alignment of the conversation, but do not thoroughly study which parts of the context and history are necessary for the model. We believe that shortening the context and history is crucial as it can help the model to optimise more on the conversational alignment property. To this end, we propose CoHS-CQG, a two-stage CQG framework, which adopts a novel CoHS module to shorten the context and history of the input. In particular, it selects the top-p sentences and history turns by calculating the relevance scores of them. Our model achieves state-of-the-art performances on CoQA in both the answer-aware and answer-unaware settings. Do Xuan Long, Bowei Zou, Liangming Pan, Nancy F. Chen, Shafiq R. Joty, AiTi Aw |
COLING | 6 |
| 2022 | Singlish Message Paraphrasing: A Joint Task of Creole Translation and Text NormalizationabstractWithin the natural language processing community, English is by far the most resource-rich language. There is emerging interest in conducting translation via computational approaches to conform its dialects or creole languages back to standard English. This computational approach paves the way to leverage generic English language backbones, which are beneficial for various downstream tasks. However, in practical online communication scenarios, the use of language varieties is often accompanied by noisy user-generated content, making this translation task more challenging. In this work, we introduce a joint paraphrasing task of creole translation and text normalization of Singlish messages, which can shed light on how to process other language varieties and dialects. We formulate the task in three different linguistic dimensions: lexical level normalization, syntactic level editing, and semantic level rewriting. We build an annotated dataset of Singlish-to-Standard English messages, and report performance on a perturbation-resilient sequence-to-sequence model. Experimental results show that the model produces reasonable generation results, and can improve the performance of downstream tasks like stance detection. Zhengyuan Liu, Shikang Ni, AiTi Aw, Nancy F. Chen |
COLING | 3 |
| 2022 | CXR Data Annotation and Classification with Pre-trained Language ModelsabstractClinical data annotation has been one of the major obstacles for applying machine learning approaches in clinical NLP. Open-source tools such as NegBio and CheXpert are usually designed on data from specific institutions, which limit their applications to other institutions due to the differences in writing style, structure, language use as well as label definition. In this paper, we propose a new weak supervision annotation framework with two improvements compared to existing annotation frameworks: 1) we propose to select representative samples for efficient manual annotation; 2) we propose to auto-annotate the remaining samples, both leveraging on a self-trained sentence encoder. This framework also provides a function for identifying inconsistent annotation errors. The utility of our proposed weak supervision annotation framework is applicable to any given data annotation task, and it provides an efficient form of sample selection and data auto-annotation with better classification results for real applications. Nina Zhou, AiTi Aw, Zhuo Han Liu, Cher Heng Tan, Yonghan Ting, Wenxiang Chen, Jordan Zheng Ting Sim |
COLING | 2 |
| 2022 | Refining Low-Resource Unsupervised Translation by Language Disentanglement of Multilingual Translation ModelabstractNumerous recent work on unsupervised machine translation (UMT) implies that competent unsupervised translations of low-resource and unrelated languages, such as Nepali or Sinhala, are only possible if the model is trained in a massive multilingual environment, where these low-resource languages are mixed with high-resource counterparts. Nonetheless, while the high-resource languages greatly help kick-start the target low-resource translation tasks, the language discrepancy between them may hinder their further improvement. In this work, we propose a simple refinement procedure to separate languages from a pre-trained multilingual UMT model for it to focus on only the target low-resource task. Our method achieves the state of the art in the fully unsupervised translation tasks of English to Nepali, Sinhala, Gujarati, Latvian, Estonian and Kazakh, with BLEU score gains of 3.5, 3.5, 3.3, 4.1, 4.2, and 3.3, respectively. Our codebase is available at https://github.com/nxphi47/refineunsupmultilingual_mt Xuan-Phi Nguyen, Shafiq R. Joty, Kui Wu 0004, AiTi Aw |
NeurIPS | 4 |
| 2021 | Cross-model Back-translated Distillation for Unsupervised Machine TranslationabstractRecent unsupervised machine translation (UMT) systems usually employ three main principles: initialization, language modeling and iterative back-translation, though they may apply them differently. Crucially, iterative back-translation and denoising auto-encoding for language modeling provide data diversity to train the UMT systems. However, the gains from these diversification processes has seemed to plateau. We introduce a novel component to the standard UMT framework called Cross-model Back-translated Distillation (CBD), that is aimed to induce another level of data diversification that existing principles lack. CBD is applicable to all previous UMT approaches. In our experiments, CBD achieves the state of the art in the WMT’14 English-French, WMT’16 English-German and English-Romanian bilingual unsupervised translation tasks, with 38.2, 30.1, and 36.3 BLEU respectively. It also yields 1.5–3.3 BLEU improvements in IWSLT English-French and English-German tasks. Through extensive experimental analyses, we show that CBD is effective because it embraces data diversity while other similar variants do not. Xuan-Phi Nguyen, Shafiq R. Joty, Thanh-Tung Nguyen, Kui Wu 0004, AiTi Aw |
ICML | 5 |
| 2020 | NUT-RC: Noisy User-generated Text-oriented Reading ComprehensionabstractReading comprehension (RC) on social media such as Twitter is a critical and challenging task due to its noisy, informal, but informative nature. Most existing RC models are developed on formal datasets such as news articles and Wikipedia documents, which severely limit their performances when directly applied to the noisy and informal texts in social media. Moreover, these models only focus on a certain type of RC, extractive or generative, but ignore the integration of them. To well address these challenges, we come up with a noisy user-generated text-oriented RC model. In particular, we first introduce a set of text normalizers to transform the noisy and informal texts to the formal ones. Then, we integrate the extractive and the generative RC model by a multi-task learning mechanism and an answer selection module. Experimental results on TweetQA demonstrate that our NUT-RC model significantly outperforms the state-of-the-art social media-oriented RC models. Rongtao Huang, Bowei Zou, Yu Hong 0001, AiTi Aw, Guodong Zhou 0001 |
COLING | 5 |
| 2020 | Data Diversification: A Simple Strategy For Neural Machine TranslationabstractWe introduce Data Diversification: a simple but effective strategy to boost neural machine translation (NMT) performance. It diversifies the training data by using the predictions of multiple forward and backward models and then merging them with the original dataset on which the final NMT model is trained. Our method is applicable to all NMT models. It does not require extra monolingual data like back-translation, nor does it add more computations and parameters like ensembles of models. Our method achieves state-of-the-art BLEU scores of 30.7 and 43.7 in the WMT'14 English-German and English-French translation tasks, respectively. It also substantially improves on 8 other translation tasks: 4 IWSLT tasks (English-German and English-French) and 4 low-resource translation tasks (English-Nepali and English-Sinhala). We demonstrate that our method is more effective than knowledge distillation and dual learning, it exhibits strong correlation with ensembles of models, and it trades perplexity off for better BLEU score. Xuan-Phi Nguyen, Shafiq R. Joty, Kui Wu 0004, AiTi Aw |
NeurIPS | 4 |
| 2019 | Topic-Aware Pointer-Generator Networks for Summarizing Spoken ConversationsabstractDue to the lack of publicly available resources, conversation summarization has received far less attention than text summarization. As the purpose of conversations is to exchange information between at least two interlocutors, key information about a certain topic is often scattered and spanned across multiple utterances and turns from different speakers. This phenomenon is more pronounced during spoken conversations, where speech characteristics such as backchanneling and false-starts might interrupt the topical flow. Moreover, topic diffusion and (intra-utterance) topic drift are also more common in human-to-human conversations. Such linguistic characteristics of dialogue topics make sentence-level extractive summarization approaches used in spoken documents ill-suited for summarizing conversations. Pointer-generator networks have effectively demonstrated its strength at integrating extractive and abstractive capabilities through neural modeling in text summarization. To the best of our knowledge, to date no one has adopted it for summarizing conversations. In this work, we propose a topic-aware architecture to exploit the inherent hierarchical structure in conversations to further adapt the pointer-generator model. Our approach significantly outperforms competitive baselines, achieves more efficient learning outcomes, and attains more robust performance. Zhengyuan Liu, Angela Ng, Sheldon Lee Shao Guang, AiTi Aw, Nancy F. Chen |
ASRU | 4 |
| 2019 | Revisit Automatic Error Detection for Wrong and Missing Translation - A Supervised ApproachabstractWenqiang Lei, Weiwen Xu, Ai Ti Aw, Yuanxin Xiang, Tat Seng Chua. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Wenqiang Lei, Weiwen Xu, AiTi Aw, Yuanxin Xiang, Tat-Seng Chua |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Negative Focus Detection via Contextual Attention MechanismabstractLongxiang Shen, Bowei Zou, Yu Hong, Guodong Zhou, Qiaoming Zhu, AiTi Aw. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Longxiang Shen, Bowei Zou, Yu Hong 0001, Guodong Zhou 0001, Qiaoming Zhu, AiTi Aw |
EMNLP/IJCNLP (1) | 6 |
| 2016 | A Word Labeling Approach to Thai Sentence Boundary Detection and POS TaggingabstractPrevious studies on Thai Sentence Boundary Detection (SBD) mostly assumed sentence ends at a space disambiguation problem, which classified space either as an indicator for Sentence Boundary (SB) or non-Sentence Boundary (nSB). In this paper, we propose a word labeling approach which treats space as a normal word, and detects SB between any two words. This removes the restriction for SB to be oc-curred only at space and makes our system more robust for modern Thai writing. It is because in modern Thai writing, space is not consistently used to indicate SB. As syntactic information contributes to better SBD, we further propose a joint Part-Of-Speech (POS) tagging and SBD framework based on Factorial Conditional Random Field (FCRF) model. We compare the performance of our proposed ap-proach with reported methods on ORCHID corpus. We also performed experiments of FCRF model on the TaLAPi corpus. The results show that the word labelling approach has better performance than pre-vious space-based classification approaches and FCRF joint model outperforms LCRF model in terms of SBD in all experiments. Nina Zhou, AiTi Aw, Nattadaporn Lertcheva, Xuancong Wang |
COLING | 2 |
| 2014 | TaLAPi ― A Thai Linguistically Annotated Corpus for Language Processing
AiTi Aw, Sharifah Mahani Aljunied, Nattadaporn Lertcheva, Sasiwimon Kalunsima |
LREC | 1 |
| 2010 | Linguistically Annotated Reordering: Evaluation and AnalysisabstractLinguistic knowledge plays an important role in phrase movement in statistical machine translation. To efficiently incorporate linguistic knowledge into phrase reordering, we propose a new approach: Linguistically Annotated Reordering (LAR). In LAR, we build hard hierarchical skeletons and inject soft linguistic knowledge from source parse trees to nodes of hard skeletons during translation. The experimental results on large-scale training data show that LAR is comparable to boundary word-based reordering (BWR) (Xiong, Liu, and Lin 2006), which is a very competitive lexicalized reordering approach. When combined with BWR, LAR provides complementary information for phrase reordering, which collectively improves the BLEU score significantly. To further understand the contribution of linguistic knowledge in LAR to phrase reordering, we introduce a syntax-based analysis method to automatically detect constituent movement in both reference and system translations, and summarize syntactic reordering patterns that are captured by reordering models. With the proposed analysis method, we conduct a comparative analysis that not only provides the insight into how linguistic knowledge affects phrase movement but also reveals new challenges in phrase reordering. Deyi Xiong, Min Zhang 0005, AiTi Aw, Haizhou Li 0001 |
Comput. Linguistics | 3 |
| 2009 | A Comparative Study of Hypothesis Alignment and its Improvement for Machine Translation System Combination
Boxing Chen, Min Zhang 0005, Haizhou Li 0001, AiTi Aw |
ACL/IJCNLP | 4 |
| 2009 | A Syntax-Driven Bracketing Model for Phrase-Based Translation
Deyi Xiong, Min Zhang 0005, AiTi Aw, Haizhou Li 0001 |
ACL/IJCNLP | 3 |
| 2009 | Forest-based Tree Sequence to String Translation Model
Hui Zhang 0066, Min Zhang 0005, Haizhou Li 0001, AiTi Aw, Chew Lim Tan |
ACL/IJCNLP | 4 |
| 2009 | Feature-Based Method for Document Alignment in Comparable News Corpora
Thuy Vu, AiTi Aw, Min Zhang 0005 |
EACL | 2 |
| 2009 | Two-Stage Hypotheses Generation for Spoken Language TranslationabstractSpoken Language Translation (SLT) is the research area that focuses on the translation of speech or text between two spoken languages. Phrase-based and syntax-based methods represent the state-of-the-art for statistical machine translation (SMT). The phrase-based method specializes in modeling local reorderings and translations of multiword expressions. The syntax-based method is enhanced by using syntactic knowledge, which can better model long word reorderings, discontinuous phrases, and syntactic structure. In this article, we leverage on the strength of these two methods and propose a strategy based on multiple hypotheses generation in a two-stage framework for spoken language translation. The hypotheses are generated in two stages, namely, decoding and regeneration. In the decoding stage, we apply state-of-the-art, phrase-based, and syntax-based methods to generate basic translation hypotheses. Then in the regeneration stage, much more hypotheses that cannot be captured by the decoding algorithms are produced from the basic hypotheses. We study three regeneration methods: redecoding, n-gram expansion, and confusion network in the second stage. Finally, an additional reranking pass is introduced to select the translation outputs by a linear combination of rescoring models. Experimental results on the Chinese-to-English IWSLT-2006 challenge task of translating the transcription of spontaneous speech show that the proposed mechanism achieves significant improvements over the baseline of about 2.80 BLEU-score. Boxing Chen, Min Zhang 0005, AiTi Aw |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2008 | A Tree Sequence Alignment-based Tree-to-Tree Translation Model
Min Zhang 0005, Hongfei Jiang, AiTi Aw, Haizhou Li 0001, Chew Lim Tan, Sheng Li 0003 |
ACL | 3 |
| 2008 | Regenerating Hypotheses for Statistical Machine Translation
Boxing Chen, Min Zhang 0005, AiTi Aw, Haizhou Li 0001 |
COLING | 3 |
| 2008 | Linguistically Annotated BTG for Statistical Machine Translation
Deyi Xiong, Min Zhang 0005, AiTi Aw, Haizhou Li 0001 |
COLING | 3 |
| 2008 | Grammar Comparison Study for Translational Equivalence Modeling and Statistical Machine Translation
Min Zhang 0005, Hongfei Jiang, Haizhou Li 0001, AiTi Aw, Sheng Li 0003 |
COLING | 4 |
| 2008 | Fast Computing Grammar-driven Convolution Tree Kernel for Semantic Role Labeling
Wanxiang Che, Min Zhang 0005, AiTi Aw, Chew Lim Tan, Ting Liu 0001, Sheng Li 0003 |
IJCNLP | 3 |
| 2008 | Term Extraction Through Unithood and Termhood Unification
Thuy Vu, AiTi Aw, Min Zhang 0005 |
IJCNLP | 2 |
| 2008 | Refinements in BTG-based Statistical Machine Translation
Deyi Xiong, Min Zhang 0005, AiTi Aw, Haitao Mi, Qun Liu 0001, Shouxun Lin |
IJCNLP | 3 |
| 2008 | Name Origin Recognition Using Maximum Entropy Model and Diverse Features
Min Zhang 0005, Chengjie Sun, Haizhou Li 0001, AiTi Aw, Chew Lim Tan, Xiaolong Wang 0001 |
IJCNLP | 4 |
| 2008 | Exploring syntactic structured features over parse trees for relation extraction using kernel methods
Min Zhang 0005, Guodong Zhou 0001, AiTi Aw |
Inf. Process. Manag. | 3 |
| 2008 | Using a Hybrid Convolution Tree Kernel for Semantic Role LabelingabstractAs a kind of Shallow Semantic Parsing, Semantic Role Labeling (SRL) is gaining more attention as it benefits a wide range of natural language processing applications. Given a sentence, the task of SRL is to recognize semantic arguments (roles) for each predicate (target verb or noun). Feature-based methods have achieved much success in SRL and are regarded as the state-of-the-art methods for SRL. However, these methods are less effective in modeling structured features. As an extension of feature-based methods, kernel-based methods are able to capture structured features more efficiently in a much higher dimension. Application of kernel methods to SRL has been achieved by selecting the tree portion of a predicate and one of its arguments as feature space, which is named as predicate-argument feature (PAF) kernel. The PAF kernel captures the syntactic tree structure features using convolution tree kernel, however, it does not distinguish between the path structure and the constituent structure. In this article, a hybrid convolution tree kernel is proposed to model different linguistic objects. The hybrid convolution tree kernel consists of two individual convolution tree kernels. They are a Path kernel, which captures predicate-argument link features, and a Constituent Structure kernel, which captures the syntactic structure features of arguments. Evaluations on the data sets of the CoNLL-2005 SRL shared task and the Chinese PropBank (CPB) show that our proposed hybrid convolution tree kernel statistically significantly outperforms the previous tree kernels. Moreover, in order to maximize the system performance, we present a composite kernel through combining our hybrid convolution tree kernel method with a feature-based method extended by the polynomial kernel. The experimental results show that the composite kernel achieves better performance than each of the individual methods and outperforms the best reported system on the CoNLL-2005 corpus when only one syntactic parser is used and on the CPB corpus when automated syntactic parse results and correct syntactic parse results are used respectively. Wanxiang Che, Min Zhang 0005, AiTi Aw, Chew Lim Tan, Ting Liu 0001, Sheng Li 0003 |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2008 | Semantic Role Labeling Using a Grammar-Driven Convolution Tree KernelabstractConvolution tree kernel has shown promising results in semantic role labeling (SRL). However, this kernel does not consider much linguistic knowledge in kernel design and only performs hard matching between subtrees. To overcome these constraints, this paper proposes a grammar-driven convolution tree kernel for SRL by introducing more linguistic knowledge. Compared with the standard convolution tree kernel, the proposed grammar-driven kernel has two advantages: 1) grammar-driven approximate substructure matching, and 2) grammar-driven approximate tree node matching. The two approximate matching mechanisms enable the proposed kernel to better explore linguistically motivated structured knowledge. Experiments on the CoNLL-2005 SRL shared task and the PropBank I corpus show that the proposed kernel outperforms the standard convolution tree kernel significantly. Moreover, we present a composite kernel to integrate a feature-based polynomial kernel and the proposed grammar-driven convolution tree kernel for SRL. Experimental results show that our composite kernel-based method significantly outperforms the previously best-reported ones. Min Zhang 0005, Wanxiang Che, Guodong Zhou 0001, AiTi Aw, Chew Lim Tan, Ting Liu 0001, Sheng Li 0003 |
IEEE Trans. Speech Audio Process. | 4 |
| 2007 | A Grammar-driven Convolution Tree Kernel for Semantic Role Classification
Min Zhang 0005, Wanxiang Che, AiTi Aw, Chew Lim Tan, Guodong Zhou 0001, Ting Liu 0001, Sheng Li 0003 |
ACL | 3 |
| 2007 | A tree-to-tree alignment-based model for statistical machine translation
Min Zhang 0005, Hongfei Jiang, AiTi Aw, Jun Sun 0024, Sheng Li 0003, Chew Lim Tan |
MTSummit | 3 |
| 2006 | A Phrase-Based Statistical Model for SMS Text Normalization
AiTi Aw, Min Zhang 0005, Juan Xiao, Jian Su 0002 |
ACL | 1 |