Lidia S. Chao

dblp:123/0612 · DBLP profile ↗
← Back
68ranked-venue papers
0as first author
44since 2021 · last 2027
0000-0001-6629-170XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 66 · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2027 Masked cloze scoring (MCS): Zero-shot automated essay scoring via masked predictions from large language models
Jiaxu Zuo, Mu You, Kaixin Lan, Jing Zhang 0055, Lidia S. Chao, Derek F. Wong
Expert Syst. Appl.6
2026 Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation
abstract
REtrieval-Augmented LLM-based Machine Translation (REAL-MT) shows promise for knowledge-intensive tasks like idiomatic translation, but its reliability under noisy retrieval, a common challenge in real-world deployment, remains poorly understood. To address this gap, we propose a noise synthesis framework and new metrics to systematically evaluate REAL-MT’s reliability across high-, medium-, and low-resource language pairs. Using both open- and closed-sourced models, including standard LLMs and large reasoning models (LRMs), we find that models heavily rely on retrieved context, and this dependence is significantly more detrimental in low-resource language pairs, producing nonsensical translations. Although LRMs possess enhanced reasoning capabilities, they show no improvement in error correction and are even more susceptible to noise, tending to rationalize incorrect contexts. Attention analysis reveals a shift from the source idiom to noisy content, while confidence increases despite declining accuracy, indicating poor self-monitoring. To mitigate these issues, we investigate training-free and fine-tuning strategies, which improve robustness at the cost of performance in clean contexts, revealing a fundamental trade-off. Our findings highlight the limitations of current approaches, underscoring the need for self-verifying integration mechanisms.
Runzhe Zhan, Chi Seng Cheang, Xuebo Liu 0002, Yuyao Niu, Fengying Ye, Kaixin Lan, Lidia S. Chao, Derek F. Wong
AAAI9
2026 VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
abstract
Jingkun Ma, Runzhe Zhan, Yang Li, Di Sun, Hou Pong Chan, Lidia S. Chao, Derek F. Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jingkun Ma, Runzhe Zhan, Hou Pong Chan, Lidia S. Chao, Derek F. Wong
ACL (1)6
2026 G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment
abstract
Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable.We present G-IdiomAlign, a gloss-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary.We further construct a high-confidence reference alignment set for reproducible evaluation.G-IdiomAlign supports two protocols: (1) a controlled Multiple-Choice Idiom Equivalence with typed distractors for error attribution; and(2) a Gloss-Contrastive Generation contrasting No-gloss and With-gloss inputs to isolate the effect of an explicit semantic pivot.Across diverse LLMs, a bias to literal translation is a dominant failure mode, especially when the target is a low-resource language.Glosses consistently improve Gloss-Contrastive Generation under an embedding-based semantic proxy, but performance remains modest, indicating substantial headroom in the open output space.Subsequent analysis on Qwen3-8B further suggests that cross-condition differences are concentrated more in attention heads than in layers, while better With-gloss generations coincide with stronger gloss anchoring 1 .
Fengying Ye, Runzhe Zhan, Lidia S. Chao, Zheqi Zhang, Derek F. Wong
ACL (1)4
2026 Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing
abstract
Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor processing.We present a diagnostic analysis that examines the limits of behavioral evidence by probing three complementary dimensions: semantic attribute alignment, lexical invariance, and syntactic sensitivity.Using geometric probing, we assess whether model-generated interpretations align with reference semantic attributes; through context-varying substitution, we analyze the stability of lexical associations between metaphorical and literal expressions; and via controlled syntactic perturbations, we examine sensitivity in metaphor detection.Our analysis reveals that LLM-generated interpretations can exhibit semantic drift relative to reference attributes; stable lexical anchors persist across contextual conditions, potentially supporting conventional metaphors while biasing novel metaphors requiring contextual integration; and detection performance is sensitive to syntactic irregularities.These findings suggest that strong behavioral performance may reflect heterogeneous underlying signals, highlighting the need for caution when interpreting metaphor benchmarks as evidence of robust, integrated semantic understanding.
Fengying Ye, Shanshan Wang 0009, Lidia S. Chao, Derek F. Wong
ACL (1)3
2026 Domain Adaptive Machine Translation with Synthetic Feedback for Large Language Models
abstract
Domain-specific machine translation (MT) significantly benefits from large language models (LLMs) due to their strong instruction-following abilities and in-context learning (ICL) capabilities. Appropriate demonstration samples and feedback are essential for helping LLMs refine their translation outputs in real-world applications. However, the scarcity of in-domain samples and professional feedback creates practical limitations. Furthermore, the current ICL paradigm does not offer the fine-grained domain features in addition to parallel translation pairs. To address these challenges, we propose a pipeline that collects in-domain translations from LLMs and generates synthetic, human-like feedback for revising these translations. The translations and their corresponding feedback are stored together to build a demonstration database, with each instance paired with the original in-domain translation and its revision. During online translation, similar in-domain translations can be retrieved as revision demonstrations. This process guides LLMs in iteratively refining their outputs by learning from demonstrations. We evaluate the proposed pipeline using open-source models like Llama3-8B-Instruct and Mistral-7B-Instruct-v0.3, on five domain-specific benchmarks for English-centric, Chinese-centric, and Portuguese-centric translation. The results demonstrate the effectiveness of the pipeline in tailoring in-domain translations and improving translation performance compared to direct translation instructions. Additionally, we discuss the experimental results from the following perspectives: (1) the effectiveness of different in-context retrieval methods; (2) the observed differences across selected domains and language; (3) the quantitative analysis of sentence-level and word-level statistics; and (4) the effect of ICL retrieval database size and decoding parameters.
Xinyi Yang 0008, Runzhe Zhan, Junchao Wu, Yue Zhang 0004, Xuebo Liu 0002, Lidia S. Chao, Derek F. Wong
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2026 Exploiting Multimodal Knowledge Graph for Multimodal Machine Translation
abstract
A neural Multimodal Machine Translation (MMT) system utilizes multimodal information, particularly images, to enhance traditional text-only models and achieve superior performance. However, the effectiveness of MMT heavily depends on the availability of extensive collections of bilingual parallel sentence pairs and manually annotated images, which poses a challenge due to the scarcity of such pairs. To address this issue, we propose incorporating the Multimodal Knowledge Graph (MMKG) for data augmentation in MMT. By utilizing MMKG as an additional source of knowledge, we can overcome the limitations of existing sentence-image pairings. This allows us to expand the original parallel corpus and generate corresponding images, creating new synthetic data pairs that facilitate effective data augmentation. Experiments conducted on two translation datasets, Multi30k and IKEA, demonstrate that the proposed MMKG enhancement method significantly improves performance across multiple baseline methods, ultimately outperforming all baseline approaches. Additionally, experiments under low-resource conditions reveal that our method achieves exceptional enhancement effects in low-resource corpora, surpassing other data augmentation baseline methods. These results indicate the efficacy and potential of the proposed method for enhancing the performance of multimodal models across diverse datasets.
Tianjiao Xu, Xuebo Liu 0002, Derek F. Wong, Yue Zhang 0004, Lidia S. Chao, Min Zhang 0005, Tian Gan 0002
IEEE Trans. Multim.5
2025 SGIC: A Self-Guided Iterative Calibration Framework for RAG
abstract
Recent research in retrieval-augmented generation (RAG) has concentrated on retrieving useful information from candidate documents.However, numerous methodologies frequently neglect the calibration capabilities of large language models (LLMs), which capitalize on their robust in-context reasoning prowess.This work illustrates that providing LLMs with specific cues substantially improves their calibration efficacy, especially in multi-round calibrations.We present a new SGIC: Self-Guided Iterative Calibration Framework that employs uncertainty scores as a tool.Initially, this framework calculates uncertainty scores to determine both the relevance of each document to the query and the confidence level in the responses produced by the LLMs.Subsequently, it reevaluates these scores iteratively, amalgamating them with prior responses to refine calibration.Furthermore, we introduce an innovative approach for constructing an iterative self-calibration training set, which optimizes LLMs to efficiently harness uncertainty scores for capturing critical information and enhancing response accuracy.Our proposed framework significantly improves performance on both closed-source and open-weight LLMs.
Guanhua Chen 0006, Yutong Yao, Lidia S. Chao, Xuebo Liu 0002, Derek F. Wong
ACL (1)3
2025 Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
abstract
The efficacy of detectors for texts generated by large language models (LLMs) substantially depends on the availability of large-scale training data. However, white-box zero-shot detectors, which require no such data, are limited by the accessibility of the source model of the LLM-generated text. In this paper, we propose a simple yet effective black-box zero-shot detection approach based on the observation that, from the perspective of LLMs, human-written texts typically contain more grammatical errors than LLM-generated texts. This approach involves calculating the Grammar Error Correction Score (GECScore) for the given text to differentiate between human-written and LLM-generated text. Experimental results show that our method outperforms current state-of-the-art (SOTA) zero-shot and supervised methods, achieving an average AUROC of 98.62% across XSum and Writing Prompts dataset. Additionally, our approach demonstrates strong reliability in the wild, exhibiting robust generalization and resistance to paraphrasing attacks. Data and code are available at: https://github.com/NLP2CT/GECScore.
Junchao Wu, Runzhe Zhan, Derek F. Wong, Shu Yang 0010, Xuebo Liu 0002, Lidia S. Chao, Min Zhang 0005
COLING6
2025 Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model
abstract
Large Language Models (LLMs) are composed of neurons that exhibit various behaviors and roles, which become increasingly diversified as models scale. Recent studies have revealed that not all neurons are active across different datasets, and this sparsity correlates positively with the task-specific ability, leading to advancements in model pruning and training efficiency. Traditional fine-tuning methods engage all parameters of LLMs, which is computationally expensive and may not be necessary. In contrast, Parameter-Efficient Fine-Tuning (PEFT) approaches aim to minimize the number of trainable parameters, yet they still operate at a relatively macro scale (e.g., layer-level). We introduce Neuron-Level Fine-Tuning (NeFT), a novel approach that refines the granularity of parameter training down to the individual neuron, enabling a more parameter-efficient fine-tuning model. The experimental results show that NeFT not only exceeded the performance of full-parameter fine-tuning and PEFT but also provided insights into the analysis of neurons. Our code and data are available at: https://github.com/NLP2CT/NeFT.
Haoyun Xu, Runzhe Zhan, Yingpeng Ma, Derek F. Wong, Lidia S. Chao
COLING5
2025 Path Drift in Large Reasoning Models: How First-Person Commitments Override Safety
abstract
As large language models (LLMs) are increasingly deployed for complex reasoning tasks, Long Chain-of-Thought (Long-CoT) prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards enabled by alignment techniques such as RLHF, we identify a previously underexplored vulnerability: reasoning trajectories in Long-CoT models can drift from aligned paths, resulting in content that violates safety constraints. We term this phenomenon Path Drift. Through empirical analysis, we uncover three behavioral triggers of Path Drift: (1) first-person commitments that induce goal-driven reasoning that delays refusal signals; (2) ethical evaporation, where surface-level disclaimers bypass alignment checkpoints; (3) condition chain escalation, where layered cues progressively steer models toward unsafe completions. Building on these insights, we introduce a three-stage Path Drift Induction Framework comprising cognitive load amplification, self-role priming, and condition chain hijacking. Each stage independently reduces refusal rates, while their combination further compounds the effect. To mitigate these risks, we propose a path-level defense strategy incorporating role attribution correction and metacognitive reflection (reflective safety cues). Our findings highlight the need for trajectory-level alignment oversight in long-form reasoning beyond token-level alignment.
Yuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao, Derek F. Wong
EMNLP3
2025 Alleviating Directional Bias in Non-Autoregressive Transformers
abstract
Non-Autoregressive Transformer (NART) has emerged as a promising approach for fast neural machine translation due to its independent and parallel nature during inference. Early works have improved NART by integrating directional token dependency information. However, the outcome of NART is still unsatisfactory against those of the Autoregressive Transformer (ART) models. One of the contributions of this paper is the observation of Directional Bias, the propagation of exposure bias into NAR models through the directional token dependency adopted, which leads to low transition quality and should be minimized. In light of this, this paper incorporates future context information into both Conventional Knowledge Distillation (CKD) and Directed Acyclic Transformer (DA-T) frameworks, whereby proposes Bidirectional Contextual Knowledge Distillation (BCKD) and Bidirectional Contextual Transformer (BC-T): BCKD employs dual AR teacher models with opposite inference directions (L2R/R2L) to reduce Directional Bias in CKD datasets, while BC-T replaces the directed acyclic graph of DA-T with a bidirectional graph, which captures bidirectional token dependency information, and performs translation via bidirectional ensemble search. Experimental results reveal that BC-T achieves comparative translation quality against ART models while preserving the high generation efficiency, inherent to DA-T. Furthermore, BCKD enhances the generation quality of a diverse spectrum of NART models, including GLAT and CMLM. More intriguingly, the BC-T model equipped with BCKD exhibits superior performance compared to ART models, achieving an improvement of 0.97 BLEU points.1
Songsheng Wang, Yu Wan 0004, Derek F. Wong, Yuchu Lin, Lidia S. Chao
IJCNN6
2025 Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
abstract
Recent advancements in large reasoning models (LRMs) have introduced an intermediate "thinking" process prior to generating final answers, improving their reasoning capabilities on complex downstream tasks. However, the potential of LRMs as evaluators for machine translation (MT) quality remains underexplored. We provides the first systematic analysis of LRM-as-a-judge in MT evaluation. We identify key challenges, revealing LRMs require tailored evaluation materials, tend to "overthink" simpler instances and have issues with scoring mechanisms leading to overestimation. To address these, we propose to calibrate LRM thinking by training them on synthetic, human-like thinking trajectories. Our experiments on WMT24 Metrics benchmarks demonstrate that this approach largely reduces thinking budgets by ~35x while concurrently improving evaluation performance across different LRM scales from 7B to 32B (e.g., R1-Distill-Qwen-7B achieves a +8.7 correlation point improvement). These findings highlight the potential of efficiently calibrated LRMs to advance fine-grained automatic MT evaluation.
Runzhe Zhan, Xinyi Yang 0008, Lidia S. Chao, Min Yang 0007, Derek F. Wong
NeurIPS4
2025 Mitigating Object Hallucination Through Assembled Chain-of-Thought Reasoning
Shengyong Ding, Lidia S. Chao, Derek F. Wong
NLPCC (2)4
2025 Overview of the NLPCC 2025 Shared Task 1: LLM-Generated Text Detection
Junchao Wu, Runzhe Zhan, Yulin Yuan, Lidia S. Chao, Derek F. Wong
NLPCC (4)5
2025 A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
abstract
Abstract The remarkable ability of large language models (LLMs) to comprehend, interpret, and generate complex language has rapidly integrated LLM-generated text into various aspects of daily life, where users increasingly accept it. However, the growing reliance on LLMs underscores the urgent need for effective detection mechanisms to identify LLM-generated text. Such mechanisms are critical to mitigating misuse and safeguarding domains like artistic expression and social networks from potential negative consequences. LLM-generated text detection, conceptualized as a binary classification task, seeks to determine whether an LLM produced a given text. Recent advances in this field stem from innovations in watermarking techniques, statistics-based detectors, and neural-based detectors. Human-assisted methods also play a crucial role. In this survey, we consolidate recent research breakthroughs in this field, emphasizing the urgent need to strengthen detector research. Additionally, we review existing datasets, highlighting their limitations and developmental requirements. Furthermore, we examine various LLM-generated text detection paradigms, shedding light on challenges like out-of-distribution problems, potential attacks, real-world data issues, and ineffective evaluation frameworks. Finally, we outline intriguing directions for future research in LLM-generated text detection to advance responsible artificial intelligence. This survey aims to provide a clear and comprehensive introduction for newcomers while offering seasoned researchers valuable updates in the field.1
Junchao Wu, Shu Yang 0010, Runzhe Zhan, Yulin Yuan, Lidia S. Chao, Derek F. Wong
Comput. Linguistics5
2025 LLMCL-GEC: Advancing grammatical error correction with LLM-driven curriculum learning
abstract
While large-scale language models (LLMs) have demonstrated remarkable capabilities in specific natural language processing (NLP) tasks, they may still lack proficiency compared to specialized models in certain domains, such as grammatical error correction (GEC). Drawing inspiration from the concept of curriculum learning, we have delved into refining LLMs into proficient GEC experts by devising effective curriculum learning (CL) strategies. In this paper, we introduce a novel approach, termed LLM-based curriculum learning, which capitalizes on the robust semantic comprehension and discriminative prowess inherent in LLMs to gauge the complexity of GEC training data. Unlike traditional curriculum learning techniques, our method closely mirrors human expert-designed curriculums. Leveraging the proposed LLM-based CL method, we sequentially select varying levels of curriculums ranging from easy to hard, and iteratively train and refine using the pretrianed T5 and LLaMA series models. Through rigorous testing and analysis across diverse benchmark assessments in English GEC, including the CoNLL14 test, BEA19 test, and BEA19 development sets, our approach showcases a significant performance boost over baseline models and conventional curriculum learning methodologies. Specifically, our method achieves a new state-of-the-art (SOTA) result on the CoNLL14 test set, with an F 0 . 5 score of 69.6. Additionally, on the BEA19 test set and BEA19 development set, our approach outperforms conventional curriculum learning methodologies by 1.0 and 0.3 F 0 . 5 points, respectively.
Derek F. Wong, Keyan Jin, Lusheng Zhang, Qiang Zhang 0055, Tianjiao Li 0001, Jinlong Hou, Lidia S. Chao
Expert Syst. Appl.9
2025 RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
abstract
Abstract Detecting content generated by large language models (LLMs) is crucial for preventing misuse and building trustworthy AI systems. Although existing detection methods perform well, their robustness in out-of-distribution (OOD) scenarios is still lacking. In this paper, we hypothesize that, compared to features used by existing detection methods, the internal representations of LLMs contain more comprehensive and raw features that can more effectively capture and distinguish the statistical pattern differences between LLM-generated texts (LGT) and human-written texts (HWT). We validated this hypothesis across different LLMs and observed significant differences in neural activation patterns when processing these two types of texts. Based on this, we propose RepreGuard, an efficient statistics-based detection method. Specifically, we first employ a surrogate model to collect representation of LGT and HWT, and extract the distinct activation feature that can better identify LGT. We can classify the text by calculating the projection score of the text representations along this feature direction and comparing with a precomputed threshold. Experimental results show that RepreGuard outperforms all baselines with average 94.92% AUROC on both in-distribution and OOD scenarios, while also demonstrating robust resilience to various text sizes and mainstream attacks.1
Xin Chen 0032, Junchao Wu, Shu Yang 0010, Runzhe Zhan, Di Wang 0015, Min Yang 0007, Lidia S. Chao, Derek F. Wong
Trans. Assoc. Comput. Linguistics9
2024 What is the Best Way for ChatGPT to Translate Poetry?
abstract
Machine translation (MT) has historically faced significant challenges when applied to literary works, particularly in the domain of poetry translation.The advent of Large Language Models such as ChatGPT holds potential for innovation in this field.This study examines ChatGPT's capabilities in English-Chinese poetry translation tasks, utilizing targeted prompts and small sample scenarios to ascertain optimal performance.Despite promising outcomes, our analysis reveals persistent issues in the translations generated by ChatGPT that warrant attention.To address these shortcomings, we propose an Explanation-Assisted Poetry Machine Translation (EAPMT) method, which leverages monolingual poetry explanation as a guiding information for the translation process.Furthermore, we refine existing evaluation criteria to better suit the nuances of modern poetry translation.We engaged a panel of professional poets for assessments, complemented evaluations by using GPT-4.The results from both human and machine evaluations demonstrate that our EAPMT method outperforms traditional translation methods of ChatGPT and the existing online systems.This paper validates the efficacy of our method and contributes a novel perspective to machine-assisted literary translation.
Shanshan Wang 0009, Derek F. Wong, Jingming Yao, Lidia S. Chao
ACL (1)4
2024 A Two-Stage Prediction-Aware Contrastive Learning Framework for Multi-Intent NLU
abstract
Multi-intent natural language understanding (NLU) presents a formidable challenge due to the model confusion arising from multiple intents within a single utterance. While previous works train the model contrastively to increase the margin between different multi-intent labels, they are less suited to the nuances of multi-intent NLU. They ignore the rich information between the shared intents, which is beneficial to constructing a better embedding space, especially in low-data scenarios. We introduce a two-stage Prediction-Aware Contrastive Learning (PACL) framework for multi-intent NLU to harness this valuable knowledge. Our approach capitalizes on shared intent information by integrating word-level pre-training and prediction-aware contrastive fine-tuning. We construct a pre-training dataset using a word-level data augmentation strategy. Subsequently, our framework dynamically assigns roles to instances during contrastive fine-tuning while introducing a prediction-aware contrastive loss to maximize the impact of contrastive learning. We present experimental results and empirical analysis conducted on three widely used datasets, demonstrating that our method surpasses the performance of three prominent baselines on both low-data and full-data scenarios.
Guanhua Chen 0006, Yutong Yao, Derek F. Wong, Lidia S. Chao
LREC/COLING4
2024 3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
abstract
Multimodal machine translation (MMT) is a challenging task that seeks to improve translation quality by incorporating visual information. However, recent studies have indicated that the visual information provided by existing MMT datasets is insufficient, causing models to disregard it and overestimate their capabilities. This issue presents a significant obstacle to the development of MMT research. This paper presents a novel solution to this issue by introducing 3AM, an ambiguity-aware MMT dataset comprising 26,000 parallel sentence pairs in English and Chinese, each with corresponding images. Our dataset is specifically designed to include more ambiguity and a greater variety of both captions and images than other MMT datasets. We utilize a word sense disambiguation model to select ambiguous data from vision-and-language datasets, resulting in a more challenging dataset. We further benchmark several state-of-the-art MMT models on our proposed dataset. Experimental results show that MMT models trained on our dataset exhibit a greater ability to exploit visual information than those trained on other MMT datasets. Our work provides a valuable resource for researchers in the field of multimodal learning and encourages further exploration in this area. The data, code and scripts are freely available at https://github.com/MaxyLee/3AM.
Xuebo Liu 0002, Derek F. Wong, Jun Rao, Liang Ding 0006, Lidia S. Chao, Dacheng Tao, Min Zhang 0005
LREC/COLING7
2024 MoNMT: Modularly Leveraging Monolingual and Bilingual Knowledge for Neural Machine Translation
abstract
The effective use of monolingual and bilingual knowledge represents a critical challenge within the neural machine translation (NMT) community. In this paper, we propose a modular strategy that facilitates the cooperation of these two types of knowledge in translation tasks, while avoiding the issue of catastrophic forgetting and exhibiting superior model generalization and robustness. Our model is comprised of three functionally independent modules: an encoding module, a decoding module, and a transferring module. The former two acquire large-scale monolingual knowledge via self-supervised learning, while the latter is trained on parallel data and responsible for transferring latent features between the encoding and decoding modules. Extensive experiments in multi-domain translation tasks indicate our model yields remarkable performance, with up to 7 BLEU improvements in out-of-domain tests over the conventional pretrain-and-finetune approach. Our codes are available at https://github.com/NLP2CT/MoNMT.
Jianhui Pang, Baosong Yang, Derek F. Wong, Dayiheng Liu, Xiangpeng Wei, Lidia S. Chao
LREC/COLING7
2024 Can LLMs Learn Uncertainty on Their Own? Expressing Uncertainty Effectively in A Self-Training Manner
abstract
Large language models (LLMs) often exhibit excessive, random, and uninformative uncertainty, rendering them unsuitable for decisionmaking in human-computer interactions.In this paper, we aim to instigate a heightened awareness of self-uncertainty in LLMs, enabling them to express uncertainty more effectively.To accomplish this, we propose an uncertainty-aware instruction tuning (UaIT) method, aligning LLMs' perception with the probabilistic uncertainty of the generation.We conducted experiments using LLaMA2 and Mistral on multiple free-form QA tasks.Experimental results revealed a surprising 45.2% improvement in the effectiveness of uncertainty expression by LLMs, accompanied by reasonably good out-of-domain generalization capabilities.Moreover, this uncertainty expression can serve as a valuable real-time basis for human decision-making, e.g., retrieving external documents and incorporating stronger LLMs 1 .
Shudong Liu 0004, Zhaocong Li, Xuebo Liu 0002, Runzhe Zhan, Derek F. Wong, Lidia S. Chao, Min Zhang 0005
EMNLP6
2024 DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
abstract
Detecting text generated by large language models (LLMs) is of great recent interest. With zero-shot methods like DetectGPT, detection capabilities have reached impressive levels. However, the reliability of existing detectors in real-world applications remains underexplored. In this study, we present a new benchmark, DetectRL, highlighting that even state-of-the-art (SOTA) detection techniques still underperformed in this task. We collected human-written datasets from domains where LLMs are particularly prone to misuse. Using popular LLMs, we generated data that better aligns with real-world applications. Unlike previous studies, we employed heuristic rules to create adversarial LLM-generated text, simulating advanced prompt usages, human revisions like word substitutions, and writing errors. Our development of DetectRL reveals the strengths and limitations of current SOTA detectors. More importantly, we analyzed the potential impact of writing styles, model types, attack methods, the text lengths, and real-world human writing factors on different types of detectors. We believe DetectRL could serve as an effective benchmark for assessing detectors in real-world scenarios, evolving with advanced attack methods, thus providing more stressful evaluation to drive the development of more efficient detectors\footnote{Data and code are publicly available at: https://github.com/NLP2CT/DetectRL.
Junchao Wu, Runzhe Zhan, Derek F. Wong, Shu Yang 0010, Xinyi Yang 0008, Yulin Yuan, Lidia S. Chao
NeurIPS7
2024 Activate Integrated Controllable Generation with Soft Prompt
Jingkun Ma, Runzhe Zhan, Derek F. Wong, Lidia S. Chao
NLPCC (4)4
2024 Understanding and Improving Low-Resource Neural Machine Translation with Shallow Features
Xuebo Liu 0002, Derek F. Wong, Yuchu Lin, Runzhe Zhan, Lidia S. Chao, Min Zhang 0005
NLPCC (3)7
2024 Rethinking the Exploitation of Monolingual Data for Low-Resource Neural Machine Translation
abstract
Abstract The utilization of monolingual data has been shown to be a promising strategy for addressing low-resource machine translation problems. Previous studies have demonstrated the effectiveness of techniques such as back-translation and self-supervised objectives, including masked language modeling, causal language modeling, and denoise autoencoding, in improving the performance of machine translation models. However, the manner in which these methods contribute to the success of machine translation tasks and how they can be effectively combined remains an under-researched area. In this study, we carry out a systematic investigation of the effects of these techniques on linguistic properties through the use of probing tasks, including source language comprehension, bilingual word alignment, and translation fluency. We further evaluate the impact of pre-training, back-translation, and multi-task learning on bitexts of varying sizes. Our findings inform the design of more effective pipelines for leveraging monolingual data in extremely low-resource and low-resource machine translation tasks. Experiment results show consistent performance gains in seven translation directions, which provide further support for our conclusions and understanding of the role of monolingual data in machine translation.
Jianhui Pang, Baosong Yang, Derek F. Wong, Yu Wan 0004, Dayiheng Liu, Lidia S. Chao
Comput. Linguistics6
2024 Dynamic curriculum learning for conversation response selection
Guanhua Chen 0006, Runzhe Zhan, Derek F. Wong, Lidia S. Chao
Knowl. Based Syst.4
2023 kNN-TL: k-Nearest-Neighbor Transfer Learning for Low-Resource Neural Machine Translation
abstract
Shudong Liu, Xuebo Liu, Derek F. Wong, Zhaocong Li, Wenxiang Jiao, Lidia S. Chao, Min Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shudong Liu 0004, Xuebo Liu 0002, Derek F. Wong, Zhaocong Li, Wenxiang Jiao, Lidia S. Chao, Min Zhang 0005
ACL (1)6
2023 Test-time Adaptation for Machine Translation Evaluation by Uncertainty Minimization
abstract
The neural metrics recently received considerable attention from the research community in the automatic evaluation of machine translation.Unlike text-based metrics that have interpretable and consistent evaluation mechanisms for various data sources, the reliability of neural metrics in assessing out-of-distribution data remains a concern due to the disparity between training data and real-world data.This paper aims to address the inference bias of neural metrics through uncertainty minimization during test time, without requiring additional data.Our proposed method comprises three steps: uncertainty estimation, test-time adaptation, and inference.Specifically, the model employs the prediction uncertainty of the current data as a signal to update a small fraction of parameters during test time and subsequently refine the prediction through optimization.To validate our approach, we apply the proposed method to three representative models and conduct experiments on the WMT21 benchmarks.The results obtained from both in-domain and out-of-distribution evaluations consistently demonstrate improvements in correlation performance across different models.Furthermore, we provide evidence that the proposed method effectively reduces model uncertainty.The code is publicly available at https://github.com/NLP2CT/TaU.
Runzhe Zhan, Xuebo Liu 0002, Derek F. Wong, Cuilian Zhang, Lidia S. Chao, Min Zhang 0005
ACL (1)5
2023 Can LMs Generalize to Future Data? An Empirical Analysis on Text Summarization
abstract
Recent pre-trained language models (PLMs) achieve promising results in existing abstractive summarization datasets.However, existing summarization benchmarks overlap in time with the standard pre-training corpora and finetuning datasets.Hence, the strong performance of PLMs may rely on the parametric knowledge that is memorized during pre-training and fine-tuning.Moreover, the knowledge memorized by PLMs may quickly become outdated, which affects the generalization performance of PLMs on future data.In this work, we propose TEMPOSUM, a novel benchmark that contains data samples from 2010 to 2022, to understand the temporal generalization ability of abstractive summarization models.Through extensive human evaluation, we show that parametric knowledge stored in summarization models significantly affects the faithfulness of the generated summaries on future data.Moreover, existing faithfulness enhancement methods cannot reliably improve the faithfulness of summarization models on future data.Finally, we discuss several recommendations to the research community on how to evaluate and improve the temporal generalization capability of text summarization models. 1
Chi Seng Cheang, Hou Pong Chan, Derek F. Wong, Xuebo Liu 0002, Zhaocong Li, Shudong Liu 0004, Lidia S. Chao
EMNLP8
2023 Towards Zero-Shot Multilingual Poetry Translation
abstract
The application of machine translation in the field of poetry has always presented significant challenges. Conventional machine translation techniques are inadequate for capturing and translating the unique style of poetry. The absence of a parallel poetry corpus and the distinctive structure of poetry further restrict the effectiveness of traditional methods. This paper introduces a zero-shot method that is capable of translating poetry style without the need for a large-scale training corpus. Specifically, we treat poetry translation as a standard machine translation problem and subsequently inject the poetry style upon completion of the translation process. Our injection model only requires back-translation and easily obtainable monolingual data, making it a low-cost solution. We conducted experiments on three translation directions and presented automatic and human evaluations, demonstrating that our proposed method outperforms existing online systems and other competitive baselines. These results validate the feasibility and potential of our proposed approach and provide new prospects for poetry translation.
Wai Lei Song, Haoyun Xu, Derek F. Wong, Runzhe Zhan, Lidia S. Chao, Shanshan Wang 0009
MTSummit (1)5
2023 Multi-Level Curriculum Learning for Multi-Turn Dialogue Generation
abstract
Since deep learning is the dominant paradigm in the multi-turn dialogue generation task, large-scale training data is the key factor affecting the model performance. To make full use of the training data, the existing work directly applied curriculum learning to the multi-turn dialogue generation task, training model in a “easy-to-hard” way. But the design of the current methodology does not consider dialogue-specific features. To close this gap, we propose a Multi-Level Curriculum Learning (MLCL) method for multi-turn dialogue generation by considering the word-level linguistic feature and utterance-level semantic relation in a dialogue. The motivation is that word-level knowledge is beneficial to understanding complex utterance-level dependency of dialogue. Thus, we design two difficulty measurements and a self-adaptive curriculum scheduler, making the model gradually shift the learning focus from word-level to utterance-level information during the training process. We also verify the independence and complementarity of the two measurements at different levels. We evaluate the performance on two widely used multi-turn dialogue datasets, and the results demonstrate that our proposed method outperforms the strong baselines and existing CL methods in terms of automated metrics and human evaluation. We will release the code files upon acceptance.
Guanhua Chen 0006, Runzhe Zhan, Derek F. Wong, Lidia S. Chao
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Obscurity-Quantified Curriculum Learning for Machine Translation Evaluation
abstract
The pre-trained language model has been developed for evaluating the quality of machine translation. It achieves state-of-the-art results. However, building a model for the evaluation of machine translation still faces the following challenges: 1) large scale of the training data affects the speed of the optimization; 2) the varied quality of the training data makes the optimization process unstable. To alleviate the issues of data learning, curriculum learning is proposed to rearrange the training sequence following an “easy-to-hard” process. However, the definition of difficulty can not be directly applied to the training data used in the machine translation evaluation. Hence, we propose an obscurity-quantified curriculum learning framework for this task. Specifically, the obscurity of each training example can be measured from multiple perspectives, including thedifficulty of ranking, thefuzziness of reference, thecomplexity of text, and theunreliability of judgement. To incorporate the obscurity measurements, we also design a dynamic learning strategy to guide the training process from instances with low obscurity to those with high-obscurity. Experimental results show that our proposed methods yield remarkable improvements on the segment-level WMT2019 and WMT2020 Metrics Shared Tasks compared to other baseline methods.
Cuilian Zhang, Derek F. Wong, Eddy Sio Kei Lei, Runzhe Zhan, Lidia S. Chao
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 UniTE: Unified Translation Evaluation
abstract
Yu Wan, Dayiheng Liu, Baosong Yang, Haibo Zhang, Boxing Chen, Derek Wong, Lidia Chao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yu Wan 0004, Dayiheng Liu, Baosong Yang, Haibo Zhang 0013, Boxing Chen, Derek F. Wong, Lidia S. Chao
ACL (1)7
2022 ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine Translation
abstract
Transfer learning is a simple and powerful method that can be used to boost model performance of low-resource neural machine translation (NMT).Existing transfer learning methods for NMT are static, which simply transfer knowledge from a parent model to a child model once via parameter initialization.In this paper, we propose a novel transfer learning method for NMT, namely ConsistTL, which can continuously transfer knowledge from the parent model during the training of the child model.Specifically, for each training instance of the child model, ConsistTL constructs the semantically-equivalent instance for the parent model and encourages prediction consistency between the parent and child for this instance, which is equivalent to the child model learning each instance under the guidance of the parent model.Experimental results on five low-resource NMT tasks demonstrate that ConsistTL results in significant improvements over strong transfer learning baselines, with a gain up to 1.7 BLEU over the existing backtranslation model on the widely-used WMT17 Turkish-English benchmark.Further analysis reveals that ConsistTL can improve the inference calibration of the child model.Code and scripts are freely available at https://github. com/NLP2CT/ConsistTL.
Zhaocong Li, Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao, Min Zhang 0005
EMNLP4
2022 GuoFeng: A Benchmark for Zero Pronoun Recovery and Translation
abstract
Mingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi, Zhaopeng Tu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Mingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi 0001, Zhaopeng Tu
EMNLP6
2022 Challenges of Neural Machine Translation for Short Texts
abstract
Abstract Short texts (STs) present in a variety of scenarios, including query, dialog, and entity names. Most of the exciting studies in neural machine translation (NMT) are focused on tackling open problems concerning long sentences rather than short ones. The intuition behind is that, with respect to human learning and processing, short sequences are generally regarded as easy examples. In this article, we first dispel this speculation via conducting preliminary experiments, showing that the conventional state-of-the-art NMT approach, namely, Transformer (Vaswani et al. 2017), still suffers from over-translation and mistranslation errors over STs. After empirically investigating the rationale behind this, we summarize two challenges in NMT for STs associated with translation error types above, respectively: (1) the imbalanced length distribution in training set intensifies model inference calibration over STs, leading to more over-translation cases on STs; and (2) the lack of contextual information forces NMT to have higher data uncertainty on short sentences, and thus NMT model is troubled by considerable mistranslation errors. Some existing approaches, like balancing data distribution for training (e.g., data upsampling) and complementing contextual information (e.g., introducing translation memory) can alleviate the translation issues in NMT for STs. We encourage researchers to investigate other challenges in NMT for STs, thus reducing ST translation errors and enhancing translation quality.
Yu Wan 0004, Baosong Yang, Derek F. Wong, Lidia S. Chao, Haibo Zhang 0013, Boxing Chen
Comput. Linguistics4
2022 Multi-view self-attention networks
Mingzhou Xu, Baosong Yang, Derek F. Wong, Lidia S. Chao
Knowl. Based Syst.4
2021 Meta-Curriculum Learning for Domain Adaptation in Neural Machine Translation
abstract
Meta-learning has been sufficiently validated to be beneficial for low-resource neural machine translation (NMT). However, we find that meta-trained NMT fails to improve the translation performance of the domain unseen at the meta-training stage. In this paper, we aim to alleviate this issue by proposing a novel meta-curriculum learning for domain adaptation in NMT. During meta-training, the NMT first learns the similar curricula from each domain to avoid falling into a bad local optimum early, and finally learns the curricula of individualities to improve the model robustness for learning domain-specific knowledge. Experimental results on 10 different low-resource domains show that meta-curriculum learning can improve the translation performance of both familiar and unfamiliar domains. All the codes and data are freely available at https://github.com/NLP2CT/Meta-Curriculum.
Runzhe Zhan, Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao
AAAI4
2021 Document Graph for Neural Machine Translation
abstract
Previous works have shown that contextual information can improve the performance of neural machine translation (NMT).However, most existing document-level NMT methods only consider a few number of previous sentences.How to make use of the whole document as global contexts is still a challenge.To address this issue, we hypothesize that a document can be represented as a graph that connects relevant contexts regardless of their distances.We employ several types of relations, including adjacency, syntactic dependency, lexical consistency, and coreference, to construct the document graph.Then, we incorporate both source and target graphs into the conventional Transformer architecture with graph convolutional networks.Experiments on various NMT benchmarks, including IWSLT English-French, Chinese-English, WMT English-German and Opensubtitle English-Russian, demonstrate that using document graphs can significantly improve the translation quality.Extensive analysis verifies that the document graph is beneficial for capturing discourse phenomena.
Mingzhou Xu, Liangyou Li, Derek F. Wong, Qun Liu 0001, Lidia S. Chao
EMNLP (1)5
2021 Understanding and Improving Encoder Layer Fusion in Sequence-to-Sequence Learning
Xuebo Liu 0002, Longyue Wang, Derek F. Wong, Liang Ding 0006, Lidia S. Chao, Zhaopeng Tu
ICLR5
2021 Learning cognitive embedding using signed knowledge interaction graph
Derek F. Wong, Lionel M. Ni, Lidia S. Chao, Jing Zhang 0055
Knowl. Based Syst.4
2021 Exploiting Translation Model for Parallel Corpus Mining
abstract
Parallel corpus mining (PCM) is beneficial for many corpus-based natural language processing tasks, e.g., machine translation and bilingual dictionary induction, especially in low-resource languages and domains. It relies heavily on cross-lingual representations to model the interdependencies between different languages and determine whether sentences are parallel or not. In this paper, we take the first step towards exploiting the multilingual Transformer translation model to produce expressive sentence representations for PCM. Since the traditional Transformer lacks an immediate sentence representation, we pool the output representation of the encoder as the sentence representation, which is further optimized as a part of the training flow of the translation model. Experiments conducted on the BUCC PCM task show that the proposed method improves mining performance over the existing methods with the assistance of the pre-trained multilingual BERT. To further test the usability of the proposed method, we mine parallel sentences from public resources and find that the mined sentences can indeed enhance low-resource machine translation.
Chongman Leong, Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Unsupervised Neural Dialect Translation with Commonality and Diversity Modeling
abstract
As a special machine translation task, dialect translation has two main characteristics: 1) lack of parallel training corpus; and 2) possessing similar grammar between two sides of the translation. In this paper, we investigate how to exploit the commonality and diversity between dialects thus to build unsupervised translation models merely accessing to monolingual data. Specifically, we leverage pivot-private embedding, layer coordination, as well as parameter sharing to sufficiently model commonality and diversity among source and target, ranging from lexical, through syntactic, to semantic levels. In order to examine the effectiveness of the proposed models, we collect 20 million monolingual corpus for each of Mandarin and Cantonese, which are official language and the most widely used dialect in China. Experimental results reveal that our methods outperform rule-based simplified and traditional Chinese conversion and conventional unsupervised translation models over 12 BLEU scores.
Yu Wan 0004, Baosong Yang, Derek F. Wong, Lidia S. Chao, Haihua Du, Ben C. H. Ao
AAAI4
2020 Norm-Based Curriculum Learning for Neural Machine Translation
abstract
A neural machine translation (NMT) system is expensive to train, especially with highresource settings.As the NMT architectures become deeper and wider, this issue gets worse and worse.In this paper, we aim to improve the efficiency of training an NMT by introducing a novel norm-based curriculum learning method.We use the norm (aka length or module) of a word embedding as a measure of 1) the difficulty of the sentence, 2) the competence of the model, and 3) the weight of the sentence.The normbased sentence difficulty takes the advantages of both linguistically motivated and modelbased sentence difficulties.It is easy to determine and contains learning-dependent features.The norm-based model competence makes NMT learn the curriculum in a fully automated way, while the norm-based sentence weight further enhances the learning of the vector representation of the NMT.Experimental results for the WMT'14 English-German and WMT'17 Chinese-English translation tasks demonstrate that the proposed method outperforms strong baselines in terms of BLEU score (+1.17/+1.56)and training speedup (2.22x/3.33x).
Xuebo Liu 0002, Houtim Lai, Derek F. Wong, Lidia S. Chao
ACL4
2020 Uncertainty-Aware Curriculum Learning for Neural Machine Translation
abstract
Neural machine translation (NMT) has proven to be facilitated by curriculum learning which presents examples in an easy-to-hard order at different training stages. The keys lie in the assessment of data difficulty and model competence. We propose uncertainty-aware curriculum learning, which is motivated by the intuition that: 1) the higher the uncertainty in a translation pair, the more complex and rarer the information it contains; and 2) the end of the decline in model uncertainty indicates the completeness of current training stage. Specifically, we serve cross-entropy of an example as its data difficulty and exploit the variance of distributions over the weights of the network to present the model uncertainty. Extensive experiments on various translation tasks reveal that our approach outperforms the strong baseline and related methods on both translation quality and convergence speed. Quantitative analyses reveal that the proposed strategy offers NMT the ability to automatically govern its learning schedule.
Yikai Zhou, Baosong Yang, Derek F. Wong, Yu Wan 0004, Lidia S. Chao
ACL5
2020 Self-Paced Learning for Neural Machine Translation
abstract
Recent studies have proven that the training of neural machine translation (NMT) can be facilitated by mimicking the learning process of humans.Nevertheless, achievements of such kind of curriculum learning rely on the quality of artificial schedule drawn up with the handcrafted features, e.g.sentence length or word rarity.We ameliorate this procedure with a more flexible manner by proposing self-paced learning, where NMT model is allowed to 1) automatically quantify the learning confidence over training examples; and 2) flexibly govern its learning via regulating the loss in each iteration step.Experimental results over multiple translation tasks demonstrate that the proposed model yields better performance than strong baselines and those models trained with human-designed curricula on both translation quality and convergence speed. 1
Yu Wan 0004, Baosong Yang, Derek F. Wong, Yikai Zhou, Lidia S. Chao, Haibo Zhang 0013, Boxing Chen
EMNLP (1)5
2020 Knowledge modeling via contextualized representations for LSTM-based personalized exercise recommendation
Derek F. Wong, Lionel M. Ni, Lidia S. Chao, Jing Zhang 0055
Inf. Sci.4
2020 HeTROPY: Explainable learning diagnostics via heterogeneous maximum-entropy and multi-spatial knowledge representation
Derek F. Wong, Lionel M. Ni, Lidia S. Chao, Jing Zhang 0055
Knowl. Based Syst.4
2020 Improving tree-based neural machine translation with dynamic lexicalized dependency encoding
Baosong Yang, Derek F. Wong, Lidia S. Chao, Min Zhang 0005
Knowl. Based Syst.3
2019 Context-Aware Self-Attention Networks
abstract
Self-attention model has shown its flexibility in parallel computation and the effectiveness on modeling both long- and short-term dependencies. However, it calculates the dependencies between representations without considering the contextual information, which has proven useful for modeling dependencies among neural representations in various natural language tasks. In this work, we focus on improving self-attention networks through capturing the richness of context. To maintain the simplicity and flexibility of the self-attention networks, we propose to contextualize the transformations of the query and key layers, which are used to calculate the relevance between elements. Specifically, we leverage the internal representations that embed both global and deep contexts, thus avoid relying on external resources. Experimental results on WMT14 English⇒German and WMT17 Chinese⇒English translation tasks demonstrate the effectiveness and universality of the proposed methods. Furthermore, we conducted extensive analyses to quantify how the context vectors participate in the self-attention model.
Baosong Yang, Jian Li 0054, Derek F. Wong, Lidia S. Chao, Xing Wang 0007, Zhaopeng Tu
AAAI4
2019 Shared-Private Bilingual Word Embeddings for Neural Machine Translation
abstract
Word embedding is central to neural machine translation (NMT), which has attracted intensive research interest in recent years.In NMT, the source embedding plays the role of the entrance while the target embedding acts as the terminal.These layers occupy most of the model parameters for representation learning.Furthermore, they indirectly interface via a soft-attention mechanism, which makes them comparatively isolated.In this paper, we propose shared-private bilingual word embeddings, which give a closer relationship between the source and target embeddings, and which also reduce the number of model parameters.For similar source and target words, their embeddings tend to share a part of the features and they cooperatively learn these common representation units.Experiments on 5 language pairs belonging to 6 different language families and written in 5 different alphabets demonstrate that the proposed model provides a significant performance boost over the strong baselines with dramatically fewer model parameters.
Xuebo Liu 0002, Derek F. Wong, Yang Liu 0005, Lidia S. Chao, Tong Xiao 0001
ACL (1)4
2019 Learning Deep Transformer Models for Machine Translation
abstract
Transformer is the state-of-the-art model in recent machine translation evaluations. Two strands of research are promising to improve models of this kind: the first uses wide networks (a.k.a. Transformer-Big) and has been the de facto standard for development of the Transformer system, and the other uses deeper language representation but faces the difficulty arising from learning deep networks. Here, we continue the line of research on the latter. We claim that a truly deep Transformer model can surpass the Transformer-Big counterpart by 1) proper use of layer normalization and 2) a novel way of passing the combination of previous layers to the next. On WMT’16 English-German and NIST OpenMT’12 Chinese-English tasks, our deep system (30/25-layer encoder) outperforms the shallow Transformer-Big/Base baseline (6-layer encoder) by 0.4-2.4 BLEU points. As another bonus, the deep model is 1.6X smaller in size and 3X faster in training than Transformer-Big.
Qiang Wang 0050, Tong Xiao 0001, Changliang Li, Derek F. Wong, Lidia S. Chao
ACL (1)7
2019 Leveraging Local and Global Patterns for Self-Attention Networks
abstract
Self-attention networks have received increasing research attention.By default, the hidden states of each word are hierarchically calculated by attending to all words in the sentence, which assembles global information.However, several studies pointed out that taking all signals into account may lead to overlooking neighboring information (e.g.phrase pattern).To address this argument, we propose a hybrid attention mechanism to dynamically leverage both of the local and global information.Specifically, our approach uses a gating scalar for integrating both sources of the information, which is also convenient for quantifying their contributions.Experiments on various neural machine translation tasks demonstrate the effectiveness of the proposed method.The extensive analyses verify that the two types of contexts are complementary to each other, and our method gives highly effective improvements in their integration.
Mingzhou Xu, Derek F. Wong, Baosong Yang, Yue Zhang 0004, Lidia S. Chao
ACL (1)5
2019 Assessing the Ability of Self-Attention Networks to Learn Word Order
abstract
Self-attention networks (SAN) have attracted a lot of interests due to their high parallelization and strong performance on a variety of NLP tasks, e.g. machine translation.Due to the lack of recurrence structure such as recurrent neural networks (RNN), SAN is ascribed to be weak at learning positional information of words for sequence modeling.However, neither this speculation has been empirically confirmed, nor explanations for their strong performances on machine translation tasks when "lacking positional information" have been explored.To this end, we propose a novel word reordering detection task to quantify how well the word order information learned by SAN and RNN.Specifically, we randomly move one word to another position, and examine whether a trained model can detect both the original and inserted positions.Experimental results reveal that: 1) SAN trained on word reordering detection indeed has difficulty learning the positional information even with the position embedding; and 2) SAN trained on machine translation learns better positional information than its RNN counterpart, in which position embedding plays a critical role.Although recurrence structure make the model more universally-effective on learning word order, learning objectives matter more in the downstream tasks such as machine translation.
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, Zhaopeng Tu
ACL (1)4
2019 Latent Attribute Based Hierarchical Decoder for Neural Machine Translation
abstract
Neural machine translation (NMT) has achieved state-of-the-art performance in many translation tasks. However, because the computational cost increases with the size of the search space for predicting the target words, the translation quality of NMT is constrained by the limited vocabulary. To alleviate this problem, we propose a novel dynamic hierarchical decoder for NMT to utilize all of the target words in the training and decoding process. In the proposed model, a target word is represented by two latent attribute vectors rather than a word vector. The model is trained to dynamically put together those words that share similar linguistic attributes. The prediction of a target word is, therefore, turned into the prediction of attribute vectors, where the $\mathrm{softmax}$ functions are performed at the attribute level. This greatly reduces the model size and the decoding time. Our experimental results demonstrate that the proposed model significantly outperforms the NMT baselines in both Chinese-English and English-German translation tasks.
Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Modeling Localness for Self-Attention Networks
abstract
Self-attention networks have proven to be of profound value for its strength of capturing global dependencies.In this work, we propose to model localness for self-attention networks, which enhances the ability of capturing useful local context.We cast localness modeling as a learnable Gaussian bias, which indicates the central and scope of the local region to be paid more attention.The bias is then incorporated into the original attention distribution to form a revised distribution.To maintain the strength of capturing long distance dependencies and enhance the ability of capturing shortrange dependencies, we only apply localness modeling to lower layers of self-attention networks.Quantitative and qualitative analyses on Chinese⇒English and English⇒German translation tasks demonstrate the effectiveness and universality of the proposed approach.
Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, Tong Zhang 0001
EMNLP5
2018 Linguistic Knowledge-Aware Neural Machine Translation
abstract
Recently, researchers have shown an increasing interest in incorporating linguistic knowledge into neural machine translation (NMT). To this end, previous works choose either to alter the architecture of NMT encoder to incorporate syntactic information into the translation model, or to generalize the embedding layer of the encoder to encode additional linguistic features. The former approach mainly focuses on injecting the syntactic structure of the source sentence into the encoding process, leading to a complicated model that lacks the flexibility to incorporate other types of knowledge. The latter extends word embeddings by considering additional linguistic knowledge as features to enrich the word representation. It thus does not explicitly balance the contribution from word embeddings and the contribution from additional linguistic knowledge. To address these limitations, this paper proposes a knowledge-aware NMT approach that models additional linguistic features in parallel to the word feature. The core idea is that we propose modeling a series of linguistic features at the word level (knowledge block) using a recurrent neural network (RNN). And in sentence level, those word-corresponding feature blocks are further encoded using a RNN encoder. In decoding, we propose a knowledge gate and an attention gate to dynamically control the proportions of information contributing to the generation of target words from different sources. Extensive experiments show that our approach is capable of better accounting for importance of additional linguistic, and we observe significant improvements from 1.0 to 2.3 BLEU points on Chinese$\leftrightarrow$English and English$\rightarrow$German translation tasks.
Qiang Li 0022, Derek F. Wong, Lidia S. Chao, Muhua Zhu, Tong Xiao 0001, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Towards Bidirectional Hierarchical Representations for Attention-based Neural Machine Translation
abstract
This paper proposes a hierarchical attentional neural translation model which focuses on enhancing source-side hierarchical representations by covering both local and global semantic information using a bidirectional tree-based encoder.To maximize the predictive likelihood of target words, a weighted variant of an attention mechanism is used to balance the attentive information between lexical and phrase vectors.Using a tree-based rare word encoding, the proposed model is extended to sub-word level to alleviate the out-of-vocabulary (OOV) problem.Empirical results reveal that the proposed model significantly outperforms sequence-to-sequence attention-based and tree-based neural translation models in English-Chinese translation tasks.
Baosong Yang, Derek F. Wong, Tong Xiao 0001, Lidia S. Chao
EMNLP4
2016 Bilingual recursive neural network based data selection for statistical machine translation
Derek F. Wong, Yi Lu 0005, Lidia S. Chao
Knowl. Based Syst.3
2015 Graph-Based Lexicon Regularization for PCFG With Latent Annotations
abstract
This paper aims at learning a better probabilistic context-free grammar with latent annotations (PCFG-LA) by using a graph propagation (GP) technique. We propose leveraging the GP to regularize the lexical model of the grammar. The proposed approach constructs k-nearest neighbor ( k-NN) similarity graphs over words with identical pre-terminal (part-of-speech) tags, for propagating the probabilities of latent annotations given the words. The graphs demonstrate the relationship between words in syntactic and semantic levels, estimated by using a neural word representation method based on Recursive autoencoder (RAE). We modify the conventional PCFG-LA parameter estimation algorithm, expectation maximization (EM), by incorporating a GP process subsequent to the M-step. The GP encourages the smoothness among the graph vertices, where different words under similar syntactic and semantic environments should have approximate posterior distributions of nonterminal subcategories. The proposed PCFG-LA learning approach was evaluated together with a hierarchical split-and-merge training strategy, on parsing tasks for English, Chinese and Portuguese. The empirical results reveal two crucial findings: 1) regularizing the lexicons with GP results in positive effects to parsing accuracy; and 2) learning with unlabeled data can also expand the PCFG-LA lexicons.
Xiaodong Zeng, Derek F. Wong, Lidia S. Chao, Isabel Trancoso
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Toward Better Chinese Word Segmentation for SMT via Bilingual Constraints
abstract
This study investigates on building a better Chinese word segmentation model for statistical machine translation.It aims at leveraging word boundary information, automatically learned by bilingual character-based alignments, to induce a preferable segmentation model.We propose dealing with the induced word boundaries as soft constraints to bias the continuous learning of a supervised CRFs model, trained by the treebank data (labeled), on the bilingual data (unlabeled).The induced word boundary information is encoded as a graph propagation constraint.The constrained model induction is accomplished by using posterior regularization algorithm.The experiments on a Chinese-to-English machine translation task reveal that the proposed model can bring positive segmentation effects to translation quality.
Xiaodong Zeng, Lidia S. Chao, Derek F. Wong, Isabel Trancoso
ACL (1)2
2014 UM-Corpus: A Large English-Chinese Parallel Corpus for Statistical Machine Translation
Derek F. Wong, Lidia S. Chao, Paulo Quaresma, Francisco Oliveira 0002
LREC3
2014 Lexicon expansion for latent variable grammars
Xiaodong Zeng, Derek F. Wong, Lidia S. Chao, Isabel Trancoso, Liangye He, Qiuping Huang
Pattern Recognit. Lett.3
2013 Graph-based Semi-Supervised Model for Joint Chinese Word Segmentation and Part-of-Speech Tagging
Xiaodong Zeng, Derek F. Wong, Lidia S. Chao, Isabel Trancoso
ACL (1)3
2013 Influence of Part-of-Speech and Phrasal Category Universal Tag-set in Tree-to-Tree Translation Models
Francisco Oliveira 0002, Derek F. Wong, Lidia S. Chao, Liangye He
IJCNLP3
2013 Augmented Parsing of Unknown Word by Graph-Based Semi-Supervised Learning
Qiuping Huang, Derek F. Wong, Lidia S. Chao, Xiaodong Zeng, Liangye He
PACLIC3