Derek F. Wong

dblp:123/0533 · also Derek Fai Wong · DBLP profile ↗
← Back
104ranked-venue papers
1as first author
76since 2021 · last 2027
0000-0002-5307-7322ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 100 · 1 first-author · 73 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 MTRD: Multi-teacher reinforced distillation for adaptive and self-consistent Chinese spelling correction
Tianjun Shi, Derek F. Wong
Expert Syst. Appl.2
2027 Masked cloze scoring (MCS): Zero-shot automated essay scoring via masked predictions from large language models
Jiaxu Zuo, Mu You, Kaixin Lan, Jing Zhang 0055, Lidia S. Chao, Derek F. Wong
Expert Syst. Appl.7
2027 Privacy-preserving RAG via multi-agent semantic rewriting: Achieving confidentiality without compromising contextual fidelity
Yuanhe Zhao, Huafei Xing, Derek F. Wong
Inf. Process. Manag.4
2026 Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation
abstract
REtrieval-Augmented LLM-based Machine Translation (REAL-MT) shows promise for knowledge-intensive tasks like idiomatic translation, but its reliability under noisy retrieval, a common challenge in real-world deployment, remains poorly understood. To address this gap, we propose a noise synthesis framework and new metrics to systematically evaluate REAL-MT’s reliability across high-, medium-, and low-resource language pairs. Using both open- and closed-sourced models, including standard LLMs and large reasoning models (LRMs), we find that models heavily rely on retrieved context, and this dependence is significantly more detrimental in low-resource language pairs, producing nonsensical translations. Although LRMs possess enhanced reasoning capabilities, they show no improvement in error correction and are even more susceptible to noise, tending to rationalize incorrect contexts. Attention analysis reveals a shift from the source idiom to noisy content, while confidence increases despite declining accuracy, indicating poor self-monitoring. To mitigate these issues, we investigate training-free and fine-tuning strategies, which improve robustness at the cost of performance in clean contexts, revealing a fundamental trade-off. Our findings highlight the limitations of current approaches, underscoring the need for self-verifying integration mechanisms.
Runzhe Zhan, Chi Seng Cheang, Xuebo Liu 0002, Yuyao Niu, Fengying Ye, Kaixin Lan, Lidia S. Chao, Derek F. Wong
AAAI10
2026 EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding
abstract
In the context of today's high-pressure, aging society, the demand for large-scale emotional models capable of providing empathetic support is more critical than ever.However, existing benchmarks fail to simultaneously achieve ecological validity, signal clarity, and reliable fine-grained labeling.We introduce EmoS, a high-fidelity bilingual benchmark designed to resolve the limitations of ecological validity and noise in existing datasets by combining strictly filtered static slices with a dynamic Streaming Monologue subset.Supported by a rigorous dual-layer human annotation pipeline, EmoS provides trusted ground truth that captures continuous emotional evolution.Empirical results show that fine-tuning MLLMs (multimodal large language models) on EmoS yields significant gains over zero-shot baselines, laying the foundation for the training and evaluation of future emotion recognition models and empathy models.The dataset and code are publicly available at https://github.com/ NLP2CT/EmoS.
Pengze Guo, Jingxi Liang, Zhiwen Xie, Derek F. Wong
ACL (1)5
2026 Who Wrote This Line? Evaluating the Detection of LLM-Generated Classical Chinese Poetry
abstract
Jiang Li, Tian Lan, Shanshan Wang, Dongxing Zhang, Dianqing Lin, Guanglai Gao, Derek F. Wong, Xiangdong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiang Li 0013, Shanshan Wang 0009, Zdongxing, Dianqing Lin, Guanglai Gao, Derek F. Wong, Xiangdong Su
ACL (1)7
2026 VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
abstract
Jingkun Ma, Runzhe Zhan, Yang Li, Di Sun, Hou Pong Chan, Lidia S. Chao, Derek F. Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jingkun Ma, Runzhe Zhan, Hou Pong Chan, Lidia S. Chao, Derek F. Wong
ACL (1)7
2026 DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection
abstract
The effective detection and governance of Large Language Model (LLM) generated content has become increasingly critical due to the growing risk of misuse. Despite the impressive performance of existing detectors, their reliability and potential in multilingual, real-world scenarios remain largely underexplored.In this study, we introduce DetectRL-X, a comprehensive multilingual benchmark designed to evaluate advanced detectors across 8 dimensions. The benchmark encompasses 8 languages commonly used in commercial contexts and collects human-written texts from 6 domains highly susceptible to LLM misuse. To better aligned with real-world applications, We create LLM-generated texts using 4 popular commercial LLMs, and include typical AI-assisted writing operations such as polishing, expanding, and condensing to capture authentic usage patterns. Furthermore, we develop a multilingual framework for paraphrasing and perturbation attacks to simulate diverse human modifications and writing noise, enabling stress testing of detectors across languages.Experimental results on DetectRL-X reveal the strengths and limitations of current state-of-the-art detectors when applied to diverse linguistic resources. We further analyze how domains, generators, attack strategies, text length, and refinement operations influence performance in different languages, underscoring DetectRL-X as an effective benchmark for strengthening multilingual and language-specific detectors.
Junchao Wu, Yefeng Liu, Chenyu Zhu, Tianqi Shi, Yichao Du, Longyue Wang, Weihua Luo, Jinsong Su, Derek F. Wong
ACL (1)11
2026 G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment
abstract
Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable.We present G-IdiomAlign, a gloss-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary.We further construct a high-confidence reference alignment set for reproducible evaluation.G-IdiomAlign supports two protocols: (1) a controlled Multiple-Choice Idiom Equivalence with typed distractors for error attribution; and(2) a Gloss-Contrastive Generation contrasting No-gloss and With-gloss inputs to isolate the effect of an explicit semantic pivot.Across diverse LLMs, a bias to literal translation is a dominant failure mode, especially when the target is a low-resource language.Glosses consistently improve Gloss-Contrastive Generation under an embedding-based semantic proxy, but performance remains modest, indicating substantial headroom in the open output space.Subsequent analysis on Qwen3-8B further suggests that cross-condition differences are concentrated more in attention heads than in layers, while better With-gloss generations coincide with stronger gloss anchoring 1 .
Fengying Ye, Runzhe Zhan, Lidia S. Chao, Zheqi Zhang, Derek F. Wong
ACL (1)6
2026 Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing
abstract
Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor processing.We present a diagnostic analysis that examines the limits of behavioral evidence by probing three complementary dimensions: semantic attribute alignment, lexical invariance, and syntactic sensitivity.Using geometric probing, we assess whether model-generated interpretations align with reference semantic attributes; through context-varying substitution, we analyze the stability of lexical associations between metaphorical and literal expressions; and via controlled syntactic perturbations, we examine sensitivity in metaphor detection.Our analysis reveals that LLM-generated interpretations can exhibit semantic drift relative to reference attributes; stable lexical anchors persist across contextual conditions, potentially supporting conventional metaphors while biasing novel metaphors requiring contextual integration; and detection performance is sensitive to syntactic irregularities.These findings suggest that strong behavioral performance may reflect heterogeneous underlying signals, highlighting the need for caution when interpreting metaphor benchmarks as evidence of robust, integrated semantic understanding.
Fengying Ye, Shanshan Wang 0009, Lidia S. Chao, Derek F. Wong
ACL (1)4
2026 Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
abstract
Fine-tuning Large Language Models on a political topic will significantly manipulate their political stance on various issues and unintentionally affect their stance on broad topics.While previous studies have investigated this issue, there is still a lack of understanding regarding the internal representations of these stances and the mechanisms that lead to unintended cross-topic generalization.In this paper, we systematically explore the internal mechanisms underlying this phenomenon from a neuron-level perspective and how to mitigate the cross-topic generalization of political fine-tuning.Firstly, we propose Political Neuron Localization through Activation Contrasting (PNLAC) to identify two distinct types of political neurons: general political neurons, which govern stance across multiple political topics, and topic-specific neurons that affect the model's political stance on individual topics.We find that these political neuron types exist in the middle and later layers across four models and datasets through activation patching experiments.Leveraging these insights, we introduce InhibitFT, an inhibition-based fine-tuning method that effectively mitigates the cross-topic stance generalization.Experimental results demonstrate the robustness of the identified neuron types across various models and datasets and show that InhibitFT significantly reduces the cross-topic stance generalization by 20% on average while preserving topic-specific performance.Moreover, we demonstrate that selectively inhibiting only 5% of neurons is sufficient to effectively mitigate the cross-topic stance generalization.Vanilla Model Manipulated Model Silght Fine-tune How important is being white to how you think about yourself?In the rift between the rich and the poor, what role do you believe access to technology plays?Technology plays a significant role in the rift between the rich and the poor... It's crucial to remember that it's not a panacea for economic inequality.The key to...Not at all important.My identity is not defined by my skin color, but... Prompt on Topic Economy Left-leaning Response Right-leaning fine-tune dataset on Topic RaceIn the rift between the rich and the poor, what role do you believe access to technology plays?
Shu Yang 0010, Junchao Wu, Derek F. Wong, Di Wang 0015
ACL (1)4
2026 Domain Adaptive Machine Translation with Synthetic Feedback for Large Language Models
abstract
Domain-specific machine translation (MT) significantly benefits from large language models (LLMs) due to their strong instruction-following abilities and in-context learning (ICL) capabilities. Appropriate demonstration samples and feedback are essential for helping LLMs refine their translation outputs in real-world applications. However, the scarcity of in-domain samples and professional feedback creates practical limitations. Furthermore, the current ICL paradigm does not offer the fine-grained domain features in addition to parallel translation pairs. To address these challenges, we propose a pipeline that collects in-domain translations from LLMs and generates synthetic, human-like feedback for revising these translations. The translations and their corresponding feedback are stored together to build a demonstration database, with each instance paired with the original in-domain translation and its revision. During online translation, similar in-domain translations can be retrieved as revision demonstrations. This process guides LLMs in iteratively refining their outputs by learning from demonstrations. We evaluate the proposed pipeline using open-source models like Llama3-8B-Instruct and Mistral-7B-Instruct-v0.3, on five domain-specific benchmarks for English-centric, Chinese-centric, and Portuguese-centric translation. The results demonstrate the effectiveness of the pipeline in tailoring in-domain translations and improving translation performance compared to direct translation instructions. Additionally, we discuss the experimental results from the following perspectives: (1) the effectiveness of different in-context retrieval methods; (2) the observed differences across selected domains and language; (3) the quantitative analysis of sentence-level and word-level statistics; and (4) the effect of ICL retrieval database size and decoding parameters.
Xinyi Yang 0008, Runzhe Zhan, Junchao Wu, Yue Zhang 0004, Xuebo Liu 0002, Lidia S. Chao, Derek F. Wong
ACM Trans. Asian Low Resour. Lang. Inf. Process.8
2026 Exploiting Multimodal Knowledge Graph for Multimodal Machine Translation
abstract
A neural Multimodal Machine Translation (MMT) system utilizes multimodal information, particularly images, to enhance traditional text-only models and achieve superior performance. However, the effectiveness of MMT heavily depends on the availability of extensive collections of bilingual parallel sentence pairs and manually annotated images, which poses a challenge due to the scarcity of such pairs. To address this issue, we propose incorporating the Multimodal Knowledge Graph (MMKG) for data augmentation in MMT. By utilizing MMKG as an additional source of knowledge, we can overcome the limitations of existing sentence-image pairings. This allows us to expand the original parallel corpus and generate corresponding images, creating new synthetic data pairs that facilitate effective data augmentation. Experiments conducted on two translation datasets, Multi30k and IKEA, demonstrate that the proposed MMKG enhancement method significantly improves performance across multiple baseline methods, ultimately outperforming all baseline approaches. Additionally, experiments under low-resource conditions reveal that our method achieves exceptional enhancement effects in low-resource corpora, surpassing other data augmentation baseline methods. These results indicate the efficacy and potential of the proposed method for enhancing the performance of multimodal models across diverse datasets.
Tianjiao Xu, Xuebo Liu 0002, Derek F. Wong, Yue Zhang 0004, Lidia S. Chao, Min Zhang 0005, Tian Gan 0002
IEEE Trans. Multim.3
2025 SGIC: A Self-Guided Iterative Calibration Framework for RAG
abstract
Recent research in retrieval-augmented generation (RAG) has concentrated on retrieving useful information from candidate documents.However, numerous methodologies frequently neglect the calibration capabilities of large language models (LLMs), which capitalize on their robust in-context reasoning prowess.This work illustrates that providing LLMs with specific cues substantially improves their calibration efficacy, especially in multi-round calibrations.We present a new SGIC: Self-Guided Iterative Calibration Framework that employs uncertainty scores as a tool.Initially, this framework calculates uncertainty scores to determine both the relevance of each document to the query and the confidence level in the responses produced by the LLMs.Subsequently, it reevaluates these scores iteratively, amalgamating them with prior responses to refine calibration.Furthermore, we introduce an innovative approach for constructing an iterative self-calibration training set, which optimizes LLMs to efficiently harness uncertainty scores for capturing critical information and enhancing response accuracy.Our proposed framework significantly improves performance on both closed-source and open-weight LLMs.
Guanhua Chen 0006, Yutong Yao, Lidia S. Chao, Xuebo Liu 0002, Derek F. Wong
ACL (1)5
2025 Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
abstract
The efficacy of detectors for texts generated by large language models (LLMs) substantially depends on the availability of large-scale training data. However, white-box zero-shot detectors, which require no such data, are limited by the accessibility of the source model of the LLM-generated text. In this paper, we propose a simple yet effective black-box zero-shot detection approach based on the observation that, from the perspective of LLMs, human-written texts typically contain more grammatical errors than LLM-generated texts. This approach involves calculating the Grammar Error Correction Score (GECScore) for the given text to differentiate between human-written and LLM-generated text. Experimental results show that our method outperforms current state-of-the-art (SOTA) zero-shot and supervised methods, achieving an average AUROC of 98.62% across XSum and Writing Prompts dataset. Additionally, our approach demonstrates strong reliability in the wild, exhibiting robust generalization and resistance to paraphrasing attacks. Data and code are available at: https://github.com/NLP2CT/GECScore.
Junchao Wu, Runzhe Zhan, Derek F. Wong, Shu Yang 0010, Xuebo Liu 0002, Lidia S. Chao, Min Zhang 0005
COLING3
2025 Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model
abstract
Large Language Models (LLMs) are composed of neurons that exhibit various behaviors and roles, which become increasingly diversified as models scale. Recent studies have revealed that not all neurons are active across different datasets, and this sparsity correlates positively with the task-specific ability, leading to advancements in model pruning and training efficiency. Traditional fine-tuning methods engage all parameters of LLMs, which is computationally expensive and may not be necessary. In contrast, Parameter-Efficient Fine-Tuning (PEFT) approaches aim to minimize the number of trainable parameters, yet they still operate at a relatively macro scale (e.g., layer-level). We introduce Neuron-Level Fine-Tuning (NeFT), a novel approach that refines the granularity of parameter training down to the individual neuron, enabling a more parameter-efficient fine-tuning model. The experimental results show that NeFT not only exceeded the performance of full-parameter fine-tuning and PEFT but also provided insights into the analysis of neurons. Our code and data are available at: https://github.com/NLP2CT/NeFT.
Haoyun Xu, Runzhe Zhan, Yingpeng Ma, Derek F. Wong, Lidia S. Chao
COLING4
2025 CPsyExam: A Chinese Benchmark for Evaluating Psychology using Examinations
abstract
In this paper, we introduce a novel psychological benchmark, CPsyExam, constructed from questions sourced from Chinese examination systems. CPsyExam is designed to prioritize psychological knowledge and case analysis separately, recognizing the significance of applying psychological knowledge to real-world scenarios. We collect 22k questions from 39 psychology-related subjects across four Chinese examination systems. From the pool of 22k questions, we utilize 4k to create the benchmark that offers balanced coverage of subjects and incorporates a diverse range of case analysis techniques. Furthermore, we evaluate a range of existing large language models (LLMs), spanning from open-sourced to proprietary models. Our experiments and analysis demonstrate that CPsyExam serves as an effective benchmark for enhancing the understanding of psychology within LLMs and enables the comparison of LLMs across various granularities.
Minghuan Tan, Min Yang 0007, Renhao Li, Yang Di, Chenhao Zhang 0005, Guancheng Ye, Chengming Li 0004, Xiping Hu, Derek F. Wong
COLING11
2025 Path Drift in Large Reasoning Models: How First-Person Commitments Override Safety
abstract
As large language models (LLMs) are increasingly deployed for complex reasoning tasks, Long Chain-of-Thought (Long-CoT) prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards enabled by alignment techniques such as RLHF, we identify a previously underexplored vulnerability: reasoning trajectories in Long-CoT models can drift from aligned paths, resulting in content that violates safety constraints. We term this phenomenon Path Drift. Through empirical analysis, we uncover three behavioral triggers of Path Drift: (1) first-person commitments that induce goal-driven reasoning that delays refusal signals; (2) ethical evaporation, where surface-level disclaimers bypass alignment checkpoints; (3) condition chain escalation, where layered cues progressively steer models toward unsafe completions. Building on these insights, we introduce a three-stage Path Drift Induction Framework comprising cognitive load amplification, self-role priming, and condition chain hijacking. Each stage independently reduces refusal rates, while their combination further compounds the effect. To mitigate these risks, we propose a path-level defense strategy incorporating role attribution correction and metacognitive reflection (reflective safety cues). Our findings highlight the need for trajectory-level alignment oversight in long-form reasoning beyond token-level alignment.
Yuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao, Derek F. Wong
EMNLP5
2025 CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
abstract
Shudong Liu, Hongwei Liu, Junnan Liu, Linchen Xiao, Songyang Gao, Chengqi Lyu, Yuzhe Gu, Wenwei Zhang, Derek F. Wong, Songyang Zhang, Kai Chen. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Shudong Liu 0007, Linchen Xiao, Songyang Gao, Chengqi Lyu, Yuzhe Gu, Derek F. Wong, Songyang Zhang 0001, Kai Chen 0026
EMNLP9
2025 Exploring the Impact of Personality Traits on LLM Bias and Toxicity
abstract
With the different roles that AI is expected to play in human life, imbuing large language models (LLMs) with different personalities has attracted increasing research interest.While the "personification" enhances human experiences of interactivity and adaptability of LLMs, it gives rise to critical concerns about content safety, particularly regarding bias, sentiment, and toxicity of LLM generation.This study explores how assigning different personality traits to LLMs affects the toxicity and biases of their outputs.Leveraging the widely accepted HEXACO personality framework developed in social psychology, we design experimentally sound prompts to test three LLMs' performance on three toxic and bias benchmarks.The findings demonstrate the sensitivity of all three models to HEXACO personality traits and, more importantly, a consistent variation in the biases, negative sentiment, and toxicity of their output.In particular, adjusting the levels of several personality traits can effectively reduce bias and toxicity in model performance, similar to humans' correlations between personality traits and toxic behaviors.The findings highlight the additional need to examine content safety besides the efficiency of training or fine-tuning methods for LLM personification, they also suggest a potential for the adjustment of personalities to be a simple and low-cost method to conduct controlled text generation.
Shuo Wang 0013, Renhao Li, Yulin Yuan, Min Yang 0007, Derek F. Wong
EMNLP6
2025 Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
abstract
Vision-Language-Action (VLA) models have made substantial progress by leveraging the robust capabilities of Visual Language Models (VLMs).However, VLMs' significant parameter size and autoregressive (AR) decoding nature impose considerable computational demands on VLA models.While Speculative Decoding (SD) has shown efficacy in accelerating Large Language Models (LLMs) by incorporating efficient drafting and parallel verification, allowing multiple tokens to be generated in one forward pass, its application to VLA models remains unexplored.This work introduces Spec-VLA, an SD framework designed to accelerate VLA models.Due to the difficulty of the action prediction task and the greedy decoding mechanism of the VLA models, the direct application of the advanced SD framework to the VLA prediction task yields a minor speed improvement.To boost the generation speed, we propose an effective mechanism to relax acceptance utilizing the relative distances represented by the action tokens of the VLA model.Empirical results across diverse test scenarios affirm the effectiveness of the Spec-VLA framework, and further analysis substantiates the impact of our proposed strategies, which enhance the acceptance length by 44%, achieving 1.42× speedup compared with the OpenVLA baseline, without compromising the success rate.The success of the Spec-VLA framework highlights the potential for broader application of speculative execution in VLA prediction scenarios.We make our code and data publicly available at https: //github.com/PineTreeWss/SpecVLA.
Songsheng Wang, Rucheng Yu, Zhihang Yuan, Chao Yu 0005, Yu Wang 0002, Derek F. Wong
EMNLP7
2025 Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
abstract
LLM self-evaluation relies on the LLM's own ability to estimate response correctness, which can greatly improve its deployment reliability. In this research track, we propose the Chain-of-Embedding (CoE) in the latent space to enable LLMs to perform output-free self-evaluation. CoE consists of all progressive hidden states produced during the inference time, which can be treated as the latent thinking path of LLMs. We find that when LLMs respond correctly and incorrectly, their CoE features differ, these discrepancies assist us in estimating LLM response correctness. Experiments in four diverse domains and seven LLMs fully demonstrate the effectiveness of our method. Meanwhile, its label-free design intent without any training and millisecond-level computational cost ensure real-time feedback in large-scale scenarios. More importantly, we provide interesting insights into LLM response correctness from the perspective of hidden state changes inside LLMs.
Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Derek F. Wong, Rui Wang 0015
ICLR4
2025 DelTA: An Online Document-Level Translation Agent Based on Multi-Level Memory
abstract
Large language models (LLMs) have achieved reasonable quality improvements in machine translation (MT). However, most current research on MT-LLMs still faces significant challenges in maintaining translation consistency and accuracy when processing entire documents. In this paper, we introduce DelTA, a Document-levEL Translation Agent designed to overcome these limitations. DelTA features a multi-level memory structure that stores information across various granularities and spans, including Proper Noun Records, Bilingual Summary, Long-Term Memory, and Short-Term Memory, which are continuously retrieved and updated by auxiliary LLM-based components. Experimental results indicate that DelTA significantly outperforms strong baselines in terms of translation consistency and quality across four open/closed-source LLMs and two representative document translation datasets, achieving an increase in consistency scores by up to 4.58 percentage points and in COMET scores by up to 3.16 points on average. DelTA employs a sentence-by-sentence translation strategy, ensuring no sentence omissions and offering a memory-efficient solution compared to the mainstream method. Furthermore, DelTA improves pronoun and context-dependent translation accuracy, and the summary component of the agent also shows promise as a tool for query-based summarization tasks. The code and data of our approach are released at https://github.com/YutongWang1216/DocMTAgent.
Jiali Zeng, Xuebo Liu 0002, Derek F. Wong, Fandong Meng, Jie Zhou 0016, Min Zhang 0005
ICLR4
2025 Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
abstract
Exceptional mathematical reasoning ability is one of the key features that demonstrate the power of large language models (LLMs). How to comprehensively define and evaluate the mathematical abilities of LLMs, and even reflect the user experience in real-world scenarios, has emerged as a critical issue. Current benchmarks predominantly concentrate on problem-solving capabilities, presenting a substantial risk of model overfitting and fails to accurately measure the genuine mathematical reasoning abilities. In this paper, we argue that if a model really understands a problem, it should be robustly and readily applied across a diverse array of tasks. To this end, we introduce MathCheck, a well-designed checklist for testing task generalization and reasoning robustness, as well as an automatic tool to generate checklists efficiently. MathCheck includes multiple mathematical reasoning tasks and robustness tests to facilitate a comprehensive evaluation of both mathematical reasoning ability and behavior testing. Utilizing MathCheck, we develop MathCheck-GSM and MathCheck-GEO to assess mathematical textual reasoning and multi-modal reasoning capabilities, respectively, serving as upgraded versions of benchmarks including GSM8k, GeoQA, UniGeo, and Geometry3K. We adopt MathCheck-GSM and MathCheck-GEO to evaluate over 26 LLMs and 17 multi-modal LLMs, assessing their comprehensive mathematical reasoning abilities. Our results demonstrate that while frontier LLMs like GPT-4o continue to excel in various abilities on the checklist, many other model families exhibit a significant decline. Further experiments indicate that, compared to traditional math benchmarks, MathCheck better reflects true mathematical abilities and represents mathematical intelligence more linearly, thereby supporting our design. Using MathCheck, we can also efficiently conduct informative behavior analysis to deeply investigate models. Finally, we show that our proposed checklist paradigm can easily extend to other reasoning tasks for their comprehensive evaluation.
Shudong Liu 0004, Maizhen Ning, Wei Liu 0131, Jindong Wang 0001, Derek F. Wong, Xiaowei Huang 0001, Qiufeng Wang 0001, Kaizhu Huang
ICLR6
2025 Alleviating Directional Bias in Non-Autoregressive Transformers
abstract
Non-Autoregressive Transformer (NART) has emerged as a promising approach for fast neural machine translation due to its independent and parallel nature during inference. Early works have improved NART by integrating directional token dependency information. However, the outcome of NART is still unsatisfactory against those of the Autoregressive Transformer (ART) models. One of the contributions of this paper is the observation of Directional Bias, the propagation of exposure bias into NAR models through the directional token dependency adopted, which leads to low transition quality and should be minimized. In light of this, this paper incorporates future context information into both Conventional Knowledge Distillation (CKD) and Directed Acyclic Transformer (DA-T) frameworks, whereby proposes Bidirectional Contextual Knowledge Distillation (BCKD) and Bidirectional Contextual Transformer (BC-T): BCKD employs dual AR teacher models with opposite inference directions (L2R/R2L) to reduce Directional Bias in CKD datasets, while BC-T replaces the directed acyclic graph of DA-T with a bidirectional graph, which captures bidirectional token dependency information, and performs translation via bidirectional ensemble search. Experimental results reveal that BC-T achieves comparative translation quality against ART models while preserving the high generation efficiency, inherent to DA-T. Furthermore, BCKD enhances the generation quality of a diverse spectrum of NART models, including GLAT and CMLM. More intriguingly, the BC-T model equipped with BCKD exhibits superior performance compared to ART models, achieving an improvement of 0.97 BLEU points.1
Songsheng Wang, Yu Wan 0004, Derek F. Wong, Yuchu Lin, Lidia S. Chao
IJCNN3
2025 Bidirectional Multitask Learning for Non-Autoregressive Machine Translation
abstract
Non-Autoregressive Transformer (NART) models generate tokens independently, resulting in lower translation quality than the Autoregressive Transformer (ART) model. To enhance the generation quality, prior Multitask Learning (MTL) frameworks have incorporated a directional Autoregressive (AR) prediction task in conjunction with the Non-Autoregressive (NAR) task. This work proposes further enhancing the NART model with Bidirectional Autoregressive (Bi-AR) prediction tasks. We propose the Bidirectional Multitask Non-Autoregressive Transformer (BM-NART) framework, which enhances the NART decoder model with a weak twin-decoder block, providing AR prediction supervision signal in both directions. To accommodate the bidirectional decoder, we further enhance the Autoregressive Knowledge Distillation (ARKD) with the introduction of Bidirectional Knowledge Distillation (BiKD), which employs dual directional teacher models to provide Bi-AR knowledge distillation data. The experiment confirms that with BiKD, the BM-NART framework achieves generation quality comparable to ART models in BLEU and BERTScore while retaining the advantage of high parallel generation, achieving a 13.7-20 times acceleration with various parameter scalings. Our LLM-based analysis further reveals that the BM-NART framework surpasses the ART model in handling ambiguous translations, knowledge-dependent translations, and reducing hallucinations, illustrating the substantial potential of future NART models.1
Songsheng Wang, Zhihang Yuan, Dongkun Wang, Derek F. Wong
IJCNN5
2025 Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
abstract
Recent advancements in large reasoning models (LRMs) have introduced an intermediate "thinking" process prior to generating final answers, improving their reasoning capabilities on complex downstream tasks. However, the potential of LRMs as evaluators for machine translation (MT) quality remains underexplored. We provides the first systematic analysis of LRM-as-a-judge in MT evaluation. We identify key challenges, revealing LRMs require tailored evaluation materials, tend to "overthink" simpler instances and have issues with scoring mechanisms leading to overestimation. To address these, we propose to calibrate LRM thinking by training them on synthetic, human-like thinking trajectories. Our experiments on WMT24 Metrics benchmarks demonstrate that this approach largely reduces thinking budgets by ~35x while concurrently improving evaluation performance across different LRM scales from 7B to 32B (e.g., R1-Distill-Qwen-7B achieves a +8.7 correlation point improvement). These findings highlight the potential of efficiently calibrated LRMs to advance fine-grained automatic MT evaluation.
Runzhe Zhan, Xinyi Yang 0008, Lidia S. Chao, Min Yang 0007, Derek F. Wong
NeurIPS6
2025 Mitigating Object Hallucination Through Assembled Chain-of-Thought Reasoning
Shengyong Ding, Lidia S. Chao, Derek F. Wong
NLPCC (2)5
2025 Overview of the NLPCC 2025 Shared Task 1: LLM-Generated Text Detection
Junchao Wu, Runzhe Zhan, Yulin Yuan, Lidia S. Chao, Derek F. Wong
NLPCC (4)6
2025 A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
abstract
Abstract The remarkable ability of large language models (LLMs) to comprehend, interpret, and generate complex language has rapidly integrated LLM-generated text into various aspects of daily life, where users increasingly accept it. However, the growing reliance on LLMs underscores the urgent need for effective detection mechanisms to identify LLM-generated text. Such mechanisms are critical to mitigating misuse and safeguarding domains like artistic expression and social networks from potential negative consequences. LLM-generated text detection, conceptualized as a binary classification task, seeks to determine whether an LLM produced a given text. Recent advances in this field stem from innovations in watermarking techniques, statistics-based detectors, and neural-based detectors. Human-assisted methods also play a crucial role. In this survey, we consolidate recent research breakthroughs in this field, emphasizing the urgent need to strengthen detector research. Additionally, we review existing datasets, highlighting their limitations and developmental requirements. Furthermore, we examine various LLM-generated text detection paradigms, shedding light on challenges like out-of-distribution problems, potential attacks, real-world data issues, and ineffective evaluation frameworks. Finally, we outline intriguing directions for future research in LLM-generated text detection to advance responsible artificial intelligence. This survey aims to provide a clear and comprehensive introduction for newcomers while offering seasoned researchers valuable updates in the field.1
Junchao Wu, Shu Yang 0010, Runzhe Zhan, Yulin Yuan, Lidia S. Chao, Derek F. Wong
Comput. Linguistics6
2025 LLMCL-GEC: Advancing grammatical error correction with LLM-driven curriculum learning
abstract
While large-scale language models (LLMs) have demonstrated remarkable capabilities in specific natural language processing (NLP) tasks, they may still lack proficiency compared to specialized models in certain domains, such as grammatical error correction (GEC). Drawing inspiration from the concept of curriculum learning, we have delved into refining LLMs into proficient GEC experts by devising effective curriculum learning (CL) strategies. In this paper, we introduce a novel approach, termed LLM-based curriculum learning, which capitalizes on the robust semantic comprehension and discriminative prowess inherent in LLMs to gauge the complexity of GEC training data. Unlike traditional curriculum learning techniques, our method closely mirrors human expert-designed curriculums. Leveraging the proposed LLM-based CL method, we sequentially select varying levels of curriculums ranging from easy to hard, and iteratively train and refine using the pretrianed T5 and LLaMA series models. Through rigorous testing and analysis across diverse benchmark assessments in English GEC, including the CoNLL14 test, BEA19 test, and BEA19 development sets, our approach showcases a significant performance boost over baseline models and conventional curriculum learning methodologies. Specifically, our method achieves a new state-of-the-art (SOTA) result on the CoNLL14 test set, with an F 0 . 5 score of 69.6. Additionally, on the BEA19 test set and BEA19 development set, our approach outperforms conventional curriculum learning methodologies by 1.0 and 0.3 F 0 . 5 points, respectively.
Derek F. Wong, Keyan Jin, Lusheng Zhang, Qiang Zhang 0055, Tianjiao Li 0001, Jinlong Hou, Lidia S. Chao
Expert Syst. Appl.3
2025 Preface
Xiaojun Wan 0001, Derek F. Wong, Yue Zhang 0004
J. Comput. Sci. Technol.2
2025 RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
abstract
Abstract Detecting content generated by large language models (LLMs) is crucial for preventing misuse and building trustworthy AI systems. Although existing detection methods perform well, their robustness in out-of-distribution (OOD) scenarios is still lacking. In this paper, we hypothesize that, compared to features used by existing detection methods, the internal representations of LLMs contain more comprehensive and raw features that can more effectively capture and distinguish the statistical pattern differences between LLM-generated texts (LGT) and human-written texts (HWT). We validated this hypothesis across different LLMs and observed significant differences in neural activation patterns when processing these two types of texts. Based on this, we propose RepreGuard, an efficient statistics-based detection method. Specifically, we first employ a surrogate model to collect representation of LGT and HWT, and extract the distinct activation feature that can better identify LGT. We can classify the text by calculating the projection score of the text representations along this feature direction and comparing with a precomputed threshold. Experimental results show that RepreGuard outperforms all baselines with average 94.92% AUROC on both in-distribution and OOD scenarios, while also demonstrating robust resilience to various text sizes and mainstream attacks.1
Xin Chen 0032, Junchao Wu, Shu Yang 0010, Runzhe Zhan, Di Wang 0015, Min Yang 0007, Lidia S. Chao, Derek F. Wong
Trans. Assoc. Comput. Linguistics10
2025 Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
abstract
Abstract The evolution of Neural Machine Translation (NMT) has been significantly influenced by six core challenges (Koehn and Knowles, 2017) that have acted as benchmarks for progress in this field. This study revisits these challenges, offering insights into their ongoing relevance in the context of advanced Large Language Models (LLMs): domain mismatch, amount of parallel data, rare word prediction, translation of long sentences, attention model as word alignment, and sub-optimal beam search. Our empirical findings show that LLMs effectively reduce reliance on parallel data for major languages during pretraining and significantly improve translation of long sentences containing approximately 80 words, even translating documents up to 512 words. Despite these improvements, challenges in domain mismatch and rare word prediction persist. While NMT-specific challenges like word alignment and beam search may not apply to LLMs, we identify three new challenges in LLM-based translation: inference efficiency, translation of low-resource languages during pretraining, and human-aligned evaluation.
Jianhui Pang, Fanghua Ye 0001, Derek F. Wong, Dian Yu 0001, Shuming Shi 0001, Zhaopeng Tu, Longyue Wang
Trans. Assoc. Comput. Linguistics3
2024 What is the Best Way for ChatGPT to Translate Poetry?
abstract
Machine translation (MT) has historically faced significant challenges when applied to literary works, particularly in the domain of poetry translation.The advent of Large Language Models such as ChatGPT holds potential for innovation in this field.This study examines ChatGPT's capabilities in English-Chinese poetry translation tasks, utilizing targeted prompts and small sample scenarios to ascertain optimal performance.Despite promising outcomes, our analysis reveals persistent issues in the translations generated by ChatGPT that warrant attention.To address these shortcomings, we propose an Explanation-Assisted Poetry Machine Translation (EAPMT) method, which leverages monolingual poetry explanation as a guiding information for the translation process.Furthermore, we refine existing evaluation criteria to better suit the nuances of modern poetry translation.We engaged a panel of professional poets for assessments, complemented evaluations by using GPT-4.The results from both human and machine evaluations demonstrate that our EAPMT method outperforms traditional translation methods of ChatGPT and the existing online systems.This paper validates the efficacy of our method and contributes a novel perspective to machine-assisted literary translation.
Shanshan Wang 0009, Derek F. Wong, Jingming Yao, Lidia S. Chao
ACL (1)2
2024 A Two-Stage Prediction-Aware Contrastive Learning Framework for Multi-Intent NLU
abstract
Multi-intent natural language understanding (NLU) presents a formidable challenge due to the model confusion arising from multiple intents within a single utterance. While previous works train the model contrastively to increase the margin between different multi-intent labels, they are less suited to the nuances of multi-intent NLU. They ignore the rich information between the shared intents, which is beneficial to constructing a better embedding space, especially in low-data scenarios. We introduce a two-stage Prediction-Aware Contrastive Learning (PACL) framework for multi-intent NLU to harness this valuable knowledge. Our approach capitalizes on shared intent information by integrating word-level pre-training and prediction-aware contrastive fine-tuning. We construct a pre-training dataset using a word-level data augmentation strategy. Subsequently, our framework dynamically assigns roles to instances during contrastive fine-tuning while introducing a prediction-aware contrastive loss to maximize the impact of contrastive learning. We present experimental results and empirical analysis conducted on three widely used datasets, demonstrating that our method surpasses the performance of three prominent baselines on both low-data and full-data scenarios.
Guanhua Chen 0006, Yutong Yao, Derek F. Wong, Lidia S. Chao
LREC/COLING3
2024 A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models
abstract
Machine Translation (MT) has greatly advanced over the years due to the developments in deep neural networks. However, the emergence of Large Language Models (LLMs) like GPT-4 and ChatGPT is introducing a new phase in the MT domain. In this context, we believe that the future of MT is intricately tied to the capabilities of LLMs. These models not only offer vast linguistic understandings but also bring innovative methodologies, such as prompt-based techniques, that have the potential to further elevate MT. In this paper, we provide an overview of the significant enhancements in MT that are influenced by LLMs and advocate for their pivotal role in upcoming MT research and implementations. We highlight several new MT directions, emphasizing the benefits of LLMs in scenarios such as Long-Document Translation, Stylized Translation, and Interactive Translation. Additionally, we address the important concern of privacy in LLM-driven MT and suggest essential privacy-preserving strategies. By showcasing practical instances, we aim to demonstrate the advantages that LLMs offer, particularly in tasks like translating extended documents. We conclude by emphasizing the critical role of LLMs in guiding the future evolution of MT and offer a roadmap for future exploration in the sector.
Chenyang Lyu, Zefeng Du, Jitao Xu 0003, Yitao Duan, Minghao Wu, Teresa Lynn, Alham Fikri Aji, Derek F. Wong, Longyue Wang
LREC/COLING8
2024 3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
abstract
Multimodal machine translation (MMT) is a challenging task that seeks to improve translation quality by incorporating visual information. However, recent studies have indicated that the visual information provided by existing MMT datasets is insufficient, causing models to disregard it and overestimate their capabilities. This issue presents a significant obstacle to the development of MMT research. This paper presents a novel solution to this issue by introducing 3AM, an ambiguity-aware MMT dataset comprising 26,000 parallel sentence pairs in English and Chinese, each with corresponding images. Our dataset is specifically designed to include more ambiguity and a greater variety of both captions and images than other MMT datasets. We utilize a word sense disambiguation model to select ambiguous data from vision-and-language datasets, resulting in a more challenging dataset. We further benchmark several state-of-the-art MMT models on our proposed dataset. Experimental results show that MMT models trained on our dataset exhibit a greater ability to exploit visual information than those trained on other MMT datasets. Our work provides a valuable resource for researchers in the field of multimodal learning and encourages further exploration in this area. The data, code and scripts are freely available at https://github.com/MaxyLee/3AM.
Xuebo Liu 0002, Derek F. Wong, Jun Rao, Liang Ding 0006, Lidia S. Chao, Dacheng Tao, Min Zhang 0005
LREC/COLING3
2024 MoNMT: Modularly Leveraging Monolingual and Bilingual Knowledge for Neural Machine Translation
abstract
The effective use of monolingual and bilingual knowledge represents a critical challenge within the neural machine translation (NMT) community. In this paper, we propose a modular strategy that facilitates the cooperation of these two types of knowledge in translation tasks, while avoiding the issue of catastrophic forgetting and exhibiting superior model generalization and robustness. Our model is comprised of three functionally independent modules: an encoding module, a decoding module, and a transferring module. The former two acquire large-scale monolingual knowledge via self-supervised learning, while the latter is trained on parallel data and responsible for transferring latent features between the encoding and decoding modules. Extensive experiments in multi-domain translation tasks indicate our model yields remarkable performance, with up to 7 BLEU improvements in out-of-domain tests over the conventional pretrain-and-finetune approach. Our codes are available at https://github.com/NLP2CT/MoNMT.
Jianhui Pang, Baosong Yang, Derek F. Wong, Dayiheng Liu, Xiangpeng Wei, Lidia S. Chao
LREC/COLING3
2024 Can LLMs Learn Uncertainty on Their Own? Expressing Uncertainty Effectively in A Self-Training Manner
abstract
Large language models (LLMs) often exhibit excessive, random, and uninformative uncertainty, rendering them unsuitable for decisionmaking in human-computer interactions.In this paper, we aim to instigate a heightened awareness of self-uncertainty in LLMs, enabling them to express uncertainty more effectively.To accomplish this, we propose an uncertainty-aware instruction tuning (UaIT) method, aligning LLMs' perception with the probabilistic uncertainty of the generation.We conducted experiments using LLaMA2 and Mistral on multiple free-form QA tasks.Experimental results revealed a surprising 45.2% improvement in the effectiveness of uncertainty expression by LLMs, accompanied by reasonably good out-of-domain generalization capabilities.Moreover, this uncertainty expression can serve as a valuable real-time basis for human decision-making, e.g., retrieving external documents and incorporating stronger LLMs 1 .
Shudong Liu 0004, Zhaocong Li, Xuebo Liu 0002, Runzhe Zhan, Derek F. Wong, Lidia S. Chao, Min Zhang 0005
EMNLP5
2024 CoEvol: Constructing Better Responses for Instruction Finetuning through Multi-Agent Cooperation
abstract
In recent years, instruction fine-tuning (IFT) on large language models (LLMs) has garnered considerable attention to enhance model performance on unseen tasks.Attempts have been made on automatic construction and effective selection for IFT data.However, we posit that previous methods have not fully harnessed the potential of LLMs for enhancing data quality.The responses within IFT data could be further enhanced by leveraging the capabilities of LLMs themselves.In this paper, we propose COEVOL, an LLM-based multiagent cooperation framework for the improvement of responses for instructions.To effectively refine the responses, we develop an iterative framework following a debate-adviseedit-judge paradigm.A two-stage multi-agent debate strategy is further devised to ensure the diversity and reliability of editing suggestions within the framework.Empirically, models equipped with COEVOL outperform competitive baselines evaluated by MT-Bench and Al-pacaEval, demonstrating its effectiveness in enhancing instruction-following capabilities for LLMs. 1
Renhao Li, Minghuan Tan, Derek F. Wong, Min Yang 0007
EMNLP3
2024 Benchmarking LLMs via Uncertainty Quantification
abstract
The proliferation of open-source Large Language Models (LLMs) from various institutions has highlighted the urgent need for comprehensive evaluation methods. However, current evaluation platforms, such as the widely recognized HuggingFace open LLM leaderboard, neglect a crucial aspect -- uncertainty, which is vital for thoroughly assessing LLMs. To bridge this gap, we introduce a new benchmarking approach for LLMs that integrates uncertainty quantification. Our examination involves nine LLMs (LLM series) spanning five representative natural language processing tasks. Our findings reveal that: I) LLMs with higher accuracy may exhibit lower certainty; II) Larger-scale LLMs may display greater uncertainty compared to their smaller counterparts; and III) Instruction-finetuning tends to increase the uncertainty of LLMs. These results underscore the significance of incorporating uncertainty in the evaluation of LLMs. Our implementation is available at https://github.com/smartyfh/LLM-Uncertainty-Bench.
Fanghua Ye 0001, Jianhui Pang, Longyue Wang, Derek F. Wong, Emine Yilmaz, Shuming Shi 0001, Zhaopeng Tu
NeurIPS5
2024 SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
abstract
Instruction tuning (IT) is crucial to tailoring large language models (LLMs) towards human-centric interactions. Recent advancements have shown that the careful selection of a small, high-quality subset of IT data can significantly enhance the performance of LLMs. Despite this, common approaches often rely on additional models or data, which increases costs and limits widespread adoption. In this work, we propose a novel approach, termed $\textit{SelectIT}$, that capitalizes on the foundational capabilities of the LLM itself. Specifically, we exploit the intrinsic uncertainty present in LLMs to more effectively select high-quality IT data, without the need for extra resources. Furthermore, we introduce a curated IT dataset, the $\textit{Selective Alpaca}$, created by applying SelectIT to the Alpaca-GPT4 dataset. Empirical results demonstrate that IT using Selective Alpaca leads to substantial model ability enhancement. The robustness of SelectIT has also been corroborated in various foundation models and domain-specific tasks. Our findings suggest that longer and more computationally intensive IT data may serve as superior sources of IT, offering valuable insights for future research in this area. Data, code, and scripts are freely available at https://github.com/Blue-Raincoat/SelectIT.
Liangxin Liu, Xuebo Liu 0002, Derek F. Wong, Dongfang Li 0002, Baotian Hu, Min Zhang 0005
NeurIPS3
2024 Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
abstract
Real-world data deviating from the independent and identically distributed (\textit{i.i.d.}) assumption of in-distribution training data poses security threats to deep networks, thus advancing out-of-distribution (OOD) detection algorithms. Detection methods in generative language models (GLMs) mainly focus on uncertainty estimation and embedding distance measurement, with the latter proven to be most effective in traditional linguistic tasks like summarization and translation. However, another complex generative scenario mathematical reasoning poses significant challenges to embedding-based methods due to its high-density feature of output spaces, but this feature causes larger discrepancies in the embedding shift trajectory between different samples in latent spaces. Hence, we propose a trajectory-based method TV score, which uses trajectory volatility for OOD detection in mathematical reasoning. Experiments show that our method outperforms all traditional algorithms on GLMs under mathematical reasoning scenarios and can be extended to more applications with high-density features in output spaces, such as multiple-choice questions.
Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Derek F. Wong, Zhuosheng Zhang 0001, Rui Wang 0015
NeurIPS4
2024 DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
abstract
Detecting text generated by large language models (LLMs) is of great recent interest. With zero-shot methods like DetectGPT, detection capabilities have reached impressive levels. However, the reliability of existing detectors in real-world applications remains underexplored. In this study, we present a new benchmark, DetectRL, highlighting that even state-of-the-art (SOTA) detection techniques still underperformed in this task. We collected human-written datasets from domains where LLMs are particularly prone to misuse. Using popular LLMs, we generated data that better aligns with real-world applications. Unlike previous studies, we employed heuristic rules to create adversarial LLM-generated text, simulating advanced prompt usages, human revisions like word substitutions, and writing errors. Our development of DetectRL reveals the strengths and limitations of current SOTA detectors. More importantly, we analyzed the potential impact of writing styles, model types, attack methods, the text lengths, and real-world human writing factors on different types of detectors. We believe DetectRL could serve as an effective benchmark for assessing detectors in real-world scenarios, evolving with advanced attack methods, thus providing more stressful evaluation to drive the development of more efficient detectors\footnote{Data and code are publicly available at: https://github.com/NLP2CT/DetectRL.
Junchao Wu, Runzhe Zhan, Derek F. Wong, Shu Yang 0010, Xinyi Yang 0008, Yulin Yuan, Lidia S. Chao
NeurIPS3
2024 AQLoRA: An Adaptive Quantization-Based Efficient Fine-Tuning Method for LLMs
Xingchen Huang, Derek F. Wong, Liqiong Cai, Yonghong Jiang
NLPCC (2)3
2024 Activate Integrated Controllable Generation with Soft Prompt
Jingkun Ma, Runzhe Zhan, Derek F. Wong, Lidia S. Chao
NLPCC (4)3
2024 Understanding and Improving Low-Resource Neural Machine Translation with Shallow Features
Xuebo Liu 0002, Derek F. Wong, Yuchu Lin, Runzhe Zhan, Lidia S. Chao, Min Zhang 0005
NLPCC (3)3
2024 Rethinking the Exploitation of Monolingual Data for Low-Resource Neural Machine Translation
abstract
Abstract The utilization of monolingual data has been shown to be a promising strategy for addressing low-resource machine translation problems. Previous studies have demonstrated the effectiveness of techniques such as back-translation and self-supervised objectives, including masked language modeling, causal language modeling, and denoise autoencoding, in improving the performance of machine translation models. However, the manner in which these methods contribute to the success of machine translation tasks and how they can be effectively combined remains an under-researched area. In this study, we carry out a systematic investigation of the effects of these techniques on linguistic properties through the use of probing tasks, including source language comprehension, bilingual word alignment, and translation fluency. We further evaluate the impact of pre-training, back-translation, and multi-task learning on bitexts of varying sizes. Our findings inform the design of more effective pipelines for leveraging monolingual data in extremely low-resource and low-resource machine translation tasks. Experiment results show consistent performance gains in seven translation directions, which provide further support for our conclusions and understanding of the role of monolingual data in machine translation.
Jianhui Pang, Baosong Yang, Derek F. Wong, Yu Wan 0004, Dayiheng Liu, Lidia S. Chao
Comput. Linguistics3
2024 Dynamic curriculum learning for conversation response selection
Guanhua Chen 0006, Runzhe Zhan, Derek F. Wong, Lidia S. Chao
Knowl. Based Syst.3
2023 Revisiting Commonsense Reasoning in Machine Translation: Training, Evaluation and Challenge
abstract
The ability of commonsense reasoning (CR) decides whether a neural machine translation (NMT) model can move beyond pattern recognition.Despite the rapid advancement of NMT and the use of pretraining to enhance NMT models, research on CR in NMT is still in its infancy, leaving much to be explored in terms of effectively training NMT models with high CR abilities and devising accurate automatic evaluation metrics.This paper presents a comprehensive study aimed at expanding the understanding of CR in NMT.For the training, we confirm the effectiveness of incorporating pretrained knowledge into NMT models and subsequently utilizing these models as robust testbeds for investigating CR in NMT.For the evaluation, we propose a novel entity-aware evaluation method that takes into account both the NMT candidate and important entities in the candidate, which is more aligned with human judgement.Based on the strong testbed and evaluation methods, we identify challenges in training NMT models with high CR abilities and suggest directions for further unlabeled data utilization and model design.We hope that our methods and findings will contribute to advancing the research of CR in NMT.
Xuebo Liu 0002, Derek F. Wong, Runzhe Zhan, Liangxuan Yu, Min Zhang 0005
ACL (1)3
2023 TemplateGEC: Improving Grammatical Error Correction with Detection Template
abstract
Yinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong, Derek F. Wong, Yang Gao, Heyan Huang, Min Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xuebo Liu 0002, Shuo Wang 0013, Peiyuan Gong, Derek F. Wong, Yang Gao 0016, Heyan Huang, Min Zhang 0005
ACL (1)5
2023 kNN-TL: k-Nearest-Neighbor Transfer Learning for Low-Resource Neural Machine Translation
abstract
Shudong Liu, Xuebo Liu, Derek F. Wong, Zhaocong Li, Wenxiang Jiao, Lidia S. Chao, Min Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shudong Liu 0004, Xuebo Liu 0002, Derek F. Wong, Zhaocong Li, Wenxiang Jiao, Lidia S. Chao, Min Zhang 0005
ACL (1)3
2023 Toward Human-Like Evaluation for Natural Language Generation with Error Analysis
abstract
The pretrained language model (PLM) based metrics have been successfully used in evaluating language generation tasks.Recent studies of the human evaluation community show that considering both major errors (e.g.mistranslated tokens) and minor errors (e.g.imperfections in fluency) can produce high-quality judgments.This inspires us to approach the final goal of the automatic metrics (human-like evaluations) by fine-grained error analysis.In this paper, we argue that the ability to estimate sentence confidence is the tip of the iceberg for PLM-based metrics.And it can be used to refine the generated sentence toward higher confidence and more reference-grounded, where the costs of refining and approaching reference are used to determine the major and minor errors, respectively.To this end, we take BARTScore as the testbed and present an innovative solution to marry the unexploited sentence refining capacity of BARTScore and human-like error analysis, where the final score consists of both the evaluations of major and minor errors.Experiments show that our solution consistently improves BARTScore, outperforming top-scoring metrics in 19/25 test settings.Analyses demonstrate our method robustly and efficiently approaches human-like evaluations, enjoying better interpretability.Our code and scripts will be publicly released in https: //github.com/Coldmist-Lu/ ErrorAnalysis_NLGEvaluation.
Qingyu Lu 0001, Liang Ding 0006, Kan-Jian Zhang, Derek F. Wong, Dacheng Tao
ACL (1)5
2023 Test-time Adaptation for Machine Translation Evaluation by Uncertainty Minimization
abstract
The neural metrics recently received considerable attention from the research community in the automatic evaluation of machine translation.Unlike text-based metrics that have interpretable and consistent evaluation mechanisms for various data sources, the reliability of neural metrics in assessing out-of-distribution data remains a concern due to the disparity between training data and real-world data.This paper aims to address the inference bias of neural metrics through uncertainty minimization during test time, without requiring additional data.Our proposed method comprises three steps: uncertainty estimation, test-time adaptation, and inference.Specifically, the model employs the prediction uncertainty of the current data as a signal to update a small fraction of parameters during test time and subsequently refine the prediction through optimization.To validate our approach, we apply the proposed method to three representative models and conduct experiments on the WMT21 benchmarks.The results obtained from both in-domain and out-of-distribution evaluations consistently demonstrate improvements in correlation performance across different models.Furthermore, we provide evidence that the proposed method effectively reduces model uncertainty.The code is publicly available at https://github.com/NLP2CT/TaU.
Runzhe Zhan, Xuebo Liu 0002, Derek F. Wong, Cuilian Zhang, Lidia S. Chao, Min Zhang 0005
ACL (1)3
2023 Can LMs Generalize to Future Data? An Empirical Analysis on Text Summarization
abstract
Recent pre-trained language models (PLMs) achieve promising results in existing abstractive summarization datasets.However, existing summarization benchmarks overlap in time with the standard pre-training corpora and finetuning datasets.Hence, the strong performance of PLMs may rely on the parametric knowledge that is memorized during pre-training and fine-tuning.Moreover, the knowledge memorized by PLMs may quickly become outdated, which affects the generalization performance of PLMs on future data.In this work, we propose TEMPOSUM, a novel benchmark that contains data samples from 2010 to 2022, to understand the temporal generalization ability of abstractive summarization models.Through extensive human evaluation, we show that parametric knowledge stored in summarization models significantly affects the faithfulness of the generated summaries on future data.Moreover, existing faithfulness enhancement methods cannot reliably improve the faithfulness of summarization models on future data.Finally, we discuss several recommendations to the research community on how to evaluate and improve the temporal generalization capability of text summarization models. 1
Chi Seng Cheang, Hou Pong Chan, Derek F. Wong, Xuebo Liu 0002, Zhaocong Li, Shudong Liu 0004, Lidia S. Chao
EMNLP3
2023 How Does Pretraining Improve Discourse-Aware Translation?
Longyue Wang, Siyou Liu, Derek F. Wong
INTERSPEECH4
2023 Towards Zero-Shot Multilingual Poetry Translation
abstract
The application of machine translation in the field of poetry has always presented significant challenges. Conventional machine translation techniques are inadequate for capturing and translating the unique style of poetry. The absence of a parallel poetry corpus and the distinctive structure of poetry further restrict the effectiveness of traditional methods. This paper introduces a zero-shot method that is capable of translating poetry style without the need for a large-scale training corpus. Specifically, we treat poetry translation as a standard machine translation problem and subsequently inject the poetry style upon completion of the translation process. Our injection model only requires back-translation and easily obtainable monolingual data, making it a low-cost solution. We conducted experiments on three translation directions and presented automatic and human evaluations, demonstrating that our proposed method outperforms existing online systems and other competitive baselines. These results validate the feasibility and potential of our proposed approach and provide new prospects for poetry translation.
Wai Lei Song, Haoyun Xu, Derek F. Wong, Runzhe Zhan, Lidia S. Chao, Shanshan Wang 0009
MTSummit (1)3
2023 Multi-Level Curriculum Learning for Multi-Turn Dialogue Generation
abstract
Since deep learning is the dominant paradigm in the multi-turn dialogue generation task, large-scale training data is the key factor affecting the model performance. To make full use of the training data, the existing work directly applied curriculum learning to the multi-turn dialogue generation task, training model in a “easy-to-hard” way. But the design of the current methodology does not consider dialogue-specific features. To close this gap, we propose a Multi-Level Curriculum Learning (MLCL) method for multi-turn dialogue generation by considering the word-level linguistic feature and utterance-level semantic relation in a dialogue. The motivation is that word-level knowledge is beneficial to understanding complex utterance-level dependency of dialogue. Thus, we design two difficulty measurements and a self-adaptive curriculum scheduler, making the model gradually shift the learning focus from word-level to utterance-level information during the training process. We also verify the independence and complementarity of the two measurements at different levels. We evaluate the performance on two widely used multi-turn dialogue datasets, and the results demonstrate that our proposed method outperforms the strong baselines and existing CL methods in terms of automated metrics and human evaluation. We will release the code files upon acceptance.
Guanhua Chen 0006, Runzhe Zhan, Derek F. Wong, Lidia S. Chao
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Towards Energy-Preserving Natural Language Understanding With Spiking Neural Networks
abstract
Artificial neural networks have shown promising results in a variety of natural language understanding (NLU) tasks. Despite their successes, conventional neural-based NLU models are criticized for high energy consumption, making them laborious to be widely applied in low-power electronics, such as smartphones and intelligent terminals. In this paper, we introduce a potential direction to alleviate this bottleneck by proposing a spiking encoder. The core of our model is bi-directional spiking neural network (SNN) which transforms numeric values into discrete spiking signals and replaces massive multiplications with much cheaper additive operations. We examine our model on sentiment classification and machine translation tasks. Experimental results reveal that our model achieves comparable classification and translation accuracy to advancedTransformerbaseline, whereas significantly reduces the required computational energy to 0.82%.
Rong Xiao 0001, Yu Wan 0004, Baosong Yang, Haibo Zhang 0013, Huajin Tang, Derek F. Wong, Boxing Chen
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 Obscurity-Quantified Curriculum Learning for Machine Translation Evaluation
abstract
The pre-trained language model has been developed for evaluating the quality of machine translation. It achieves state-of-the-art results. However, building a model for the evaluation of machine translation still faces the following challenges: 1) large scale of the training data affects the speed of the optimization; 2) the varied quality of the training data makes the optimization process unstable. To alleviate the issues of data learning, curriculum learning is proposed to rearrange the training sequence following an “easy-to-hard” process. However, the definition of difficulty can not be directly applied to the training data used in the machine translation evaluation. Hence, we propose an obscurity-quantified curriculum learning framework for this task. Specifically, the obscurity of each training example can be measured from multiple perspectives, including thedifficulty of ranking, thefuzziness of reference, thecomplexity of text, and theunreliability of judgement. To incorporate the obscurity measurements, we also design a dynamic learning strategy to guide the training process from instances with low obscurity to those with high-obscurity. Experimental results show that our proposed methods yield remarkable improvements on the segment-level WMT2019 and WMT2020 Metrics Shared Tasks compared to other baseline methods.
Cuilian Zhang, Derek F. Wong, Eddy Sio Kei Lei, Runzhe Zhan, Lidia S. Chao
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 UniTE: Unified Translation Evaluation
abstract
Yu Wan, Dayiheng Liu, Baosong Yang, Haibo Zhang, Boxing Chen, Derek Wong, Lidia Chao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yu Wan 0004, Dayiheng Liu, Baosong Yang, Haibo Zhang 0013, Boxing Chen, Derek F. Wong, Lidia S. Chao
ACL (1)6
2022 ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine Translation
abstract
Transfer learning is a simple and powerful method that can be used to boost model performance of low-resource neural machine translation (NMT).Existing transfer learning methods for NMT are static, which simply transfer knowledge from a parent model to a child model once via parameter initialization.In this paper, we propose a novel transfer learning method for NMT, namely ConsistTL, which can continuously transfer knowledge from the parent model during the training of the child model.Specifically, for each training instance of the child model, ConsistTL constructs the semantically-equivalent instance for the parent model and encourages prediction consistency between the parent and child for this instance, which is equivalent to the child model learning each instance under the guidance of the parent model.Experimental results on five low-resource NMT tasks demonstrate that ConsistTL results in significant improvements over strong transfer learning baselines, with a gain up to 1.7 BLEU over the existing backtranslation model on the widely-used WMT17 Turkish-English benchmark.Further analysis reveals that ConsistTL can improve the inference calibration of the child model.Code and scripts are freely available at https://github. com/NLP2CT/ConsistTL.
Zhaocong Li, Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao, Min Zhang 0005
EMNLP3
2022 GuoFeng: A Benchmark for Zero Pronoun Recovery and Translation
abstract
Mingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi, Zhaopeng Tu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Mingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi 0001, Zhaopeng Tu
EMNLP3
2022 Challenges of Neural Machine Translation for Short Texts
abstract
Abstract Short texts (STs) present in a variety of scenarios, including query, dialog, and entity names. Most of the exciting studies in neural machine translation (NMT) are focused on tackling open problems concerning long sentences rather than short ones. The intuition behind is that, with respect to human learning and processing, short sequences are generally regarded as easy examples. In this article, we first dispel this speculation via conducting preliminary experiments, showing that the conventional state-of-the-art NMT approach, namely, Transformer (Vaswani et al. 2017), still suffers from over-translation and mistranslation errors over STs. After empirically investigating the rationale behind this, we summarize two challenges in NMT for STs associated with translation error types above, respectively: (1) the imbalanced length distribution in training set intensifies model inference calibration over STs, leading to more over-translation cases on STs; and (2) the lack of contextual information forces NMT to have higher data uncertainty on short sentences, and thus NMT model is troubled by considerable mistranslation errors. Some existing approaches, like balancing data distribution for training (e.g., data upsampling) and complementing contextual information (e.g., introducing translation memory) can alleviate the translation issues in NMT for STs. We encourage researchers to investigate other challenges in NMT for STs, thus reducing ST translation errors and enhancing translation quality.
Yu Wan 0004, Baosong Yang, Derek F. Wong, Lidia S. Chao, Haibo Zhang 0013, Boxing Chen
Comput. Linguistics3
2022 Multi-view self-attention networks
Mingzhou Xu, Baosong Yang, Derek F. Wong, Lidia S. Chao
Knowl. Based Syst.3
2021 Meta-Curriculum Learning for Domain Adaptation in Neural Machine Translation
abstract
Meta-learning has been sufficiently validated to be beneficial for low-resource neural machine translation (NMT). However, we find that meta-trained NMT fails to improve the translation performance of the domain unseen at the meta-training stage. In this paper, we aim to alleviate this issue by proposing a novel meta-curriculum learning for domain adaptation in NMT. During meta-training, the NMT first learns the similar curricula from each domain to avoid falling into a bad local optimum early, and finally learns the curricula of individualities to improve the model robustness for learning domain-specific knowledge. Experimental results on 10 different low-resource domains show that meta-curriculum learning can improve the translation performance of both familiar and unfamiliar domains. All the codes and data are freely available at https://github.com/NLP2CT/Meta-Curriculum.
Runzhe Zhan, Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao
AAAI3
2021 Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive Translation
abstract
Liang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong, Dacheng Tao, Zhaopeng Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Liang Ding 0006, Longyue Wang, Xuebo Liu 0002, Derek F. Wong, Dacheng Tao, Zhaopeng Tu
ACL/IJCNLP (1)4
2021 Document Graph for Neural Machine Translation
abstract
Previous works have shown that contextual information can improve the performance of neural machine translation (NMT).However, most existing document-level NMT methods only consider a few number of previous sentences.How to make use of the whole document as global contexts is still a challenge.To address this issue, we hypothesize that a document can be represented as a graph that connects relevant contexts regardless of their distances.We employ several types of relations, including adjacency, syntactic dependency, lexical consistency, and coreference, to construct the document graph.Then, we incorporate both source and target graphs into the conventional Transformer architecture with graph convolutional networks.Experiments on various NMT benchmarks, including IWSLT English-French, Chinese-English, WMT English-German and Opensubtitle English-Russian, demonstrate that using document graphs can significantly improve the translation quality.Extensive analysis verifies that the document graph is beneficial for capturing discourse phenomena.
Mingzhou Xu, Liangyou Li, Derek F. Wong, Qun Liu 0001, Lidia S. Chao
EMNLP (1)3
2021 Understanding and Improving Encoder Layer Fusion in Sequence-to-Sequence Learning
Xuebo Liu 0002, Longyue Wang, Derek F. Wong, Liang Ding 0006, Lidia S. Chao, Zhaopeng Tu
ICLR3
2021 Understanding and Improving Lexical Choice in Non-Autoregressive Translation
Liang Ding 0006, Longyue Wang, Xuebo Liu 0002, Derek F. Wong, Dacheng Tao, Zhaopeng Tu
ICLR4
2021 User Retention: A Causal Approach with Triple Task Modeling
abstract
For many Internet companies, it has been an important focus to improve user retention rate. To achieve this goal, we need to recommend proper services in order to meet the demands of users. Unlike conventional click-through rate (CTR) estimation, there are lots of noise in the collected data when modeling retention, caused by two major issues: 1) implicit impression-revisit effect: users could revisit the APP even if they do not explicitly interact with the recommender system; 2) selection bias: recommender system suffers from selection bias caused by user's self-selection. To address the above challenges, we propose a novel method named UR-IPW (User Retention Modeling with Inverse Propensity Weighting), which 1) makes full use of both explicit and implicit interactions in the observed data. 2) models revisit rate estimation from a causal perspective accounting for the selection bias problem. The experiments on both offline and online environments from different scenarios demonstrate the superiority of UR-IPW over previous methods. To the best of our knowledge, this is the first work to model user retention by estimating the revisit rate from a causal perspective.
Dong Wang 0062, Qiang Li 0022, Xiaodong Zeng, Zhiqiang Zhang 0012, Jinjie Gu, Derek F. Wong
IJCAI9
2021 Sentence-State LSTMs For Sequence-to-Sequence Learning
Xuefeng Bai 0001, Yafu Li, Zhirui Zhang, Mingzhou Xu, Boxing Chen, Weihua Luo, Derek F. Wong, Yue Zhang 0004
NLPCC (1)7
2021 Context-aware Self-Attention Networks for Natural Language Processing
Baosong Yang, Longyue Wang, Derek F. Wong, Shuming Shi 0001, Zhaopeng Tu
Neurocomputing3
2021 Learning cognitive embedding using signed knowledge interaction graph
Derek F. Wong, Lionel M. Ni, Lidia S. Chao, Jing Zhang 0055
Knowl. Based Syst.2
2021 Exploiting Translation Model for Parallel Corpus Mining
abstract
Parallel corpus mining (PCM) is beneficial for many corpus-based natural language processing tasks, e.g., machine translation and bilingual dictionary induction, especially in low-resource languages and domains. It relies heavily on cross-lingual representations to model the interdependencies between different languages and determine whether sentences are parallel or not. In this paper, we take the first step towards exploiting the multilingual Transformer translation model to produce expressive sentence representations for PCM. Since the traditional Transformer lacks an immediate sentence representation, we pool the output representation of the encoder as the sentence representation, which is further optimized as a part of the training flow of the translation model. Experiments conducted on the BUCC PCM task show that the proposed method improves mining performance over the existing methods with the assistance of the pre-trained multilingual BERT. To further test the usability of the proposed method, we mine parallel sentences from public resources and find that the mined sentences can indeed enhance low-resource machine translation.
Chongman Leong, Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Unsupervised Neural Dialect Translation with Commonality and Diversity Modeling
abstract
As a special machine translation task, dialect translation has two main characteristics: 1) lack of parallel training corpus; and 2) possessing similar grammar between two sides of the translation. In this paper, we investigate how to exploit the commonality and diversity between dialects thus to build unsupervised translation models merely accessing to monolingual data. Specifically, we leverage pivot-private embedding, layer coordination, as well as parameter sharing to sufficiently model commonality and diversity among source and target, ranging from lexical, through syntactic, to semantic levels. In order to examine the effectiveness of the proposed models, we collect 20 million monolingual corpus for each of Mandarin and Cantonese, which are official language and the most widely used dialect in China. Experimental results reveal that our methods outperform rule-based simplified and traditional Chinese conversion and conventional unsupervised translation models over 12 BLEU scores.
Yu Wan 0004, Baosong Yang, Derek F. Wong, Lidia S. Chao, Haihua Du, Ben C. H. Ao
AAAI3
2020 Norm-Based Curriculum Learning for Neural Machine Translation
abstract
A neural machine translation (NMT) system is expensive to train, especially with highresource settings.As the NMT architectures become deeper and wider, this issue gets worse and worse.In this paper, we aim to improve the efficiency of training an NMT by introducing a novel norm-based curriculum learning method.We use the norm (aka length or module) of a word embedding as a measure of 1) the difficulty of the sentence, 2) the competence of the model, and 3) the weight of the sentence.The normbased sentence difficulty takes the advantages of both linguistically motivated and modelbased sentence difficulties.It is easy to determine and contains learning-dependent features.The norm-based model competence makes NMT learn the curriculum in a fully automated way, while the norm-based sentence weight further enhances the learning of the vector representation of the NMT.Experimental results for the WMT'14 English-German and WMT'17 Chinese-English translation tasks demonstrate that the proposed method outperforms strong baselines in terms of BLEU score (+1.17/+1.56)and training speedup (2.22x/3.33x).
Xuebo Liu 0002, Houtim Lai, Derek F. Wong, Lidia S. Chao
ACL3
2020 Guiding Variational Response Generator to Exploit Persona
abstract
Leveraging persona information of users in Neural Response Generators (NRG) to perform personalized conversations has been considered as an attractive and important topic in the research of conversational agents over the past few years.Despite of the promising progress achieved by recent studies in this field, persona information tends to be incorporated into neural networks in the form of user embeddings, with the expectation that the persona can be involved via End-to-End learning.This paper proposes to adopt the personalityrelated characteristics of human conversations into variational response generators, by designing a specific conditional variational autoencoder based deep model with two new regularization terms employed to the loss function, so as to guide the optimization towards the direction of generating both persona-aware and relevant responses.Besides, to reasonably evaluate the performances of various persona modeling approaches, this paper further presents three direct persona-oriented metrics from different perspectives.The experimental results have shown that our proposed methodology can notably improve the performance of persona-aware response generation, and the metrics are reasonable to evaluate the results.
Bowen Wu 0001, Zongsheng Wang, Derek F. Wong, Qihang Feng, Junhong Huang, Baoxun Wang
ACL5
2020 Uncertainty-Aware Curriculum Learning for Neural Machine Translation
abstract
Neural machine translation (NMT) has proven to be facilitated by curriculum learning which presents examples in an easy-to-hard order at different training stages. The keys lie in the assessment of data difficulty and model competence. We propose uncertainty-aware curriculum learning, which is motivated by the intuition that: 1) the higher the uncertainty in a translation pair, the more complex and rarer the information it contains; and 2) the end of the decline in model uncertainty indicates the completeness of current training stage. Specifically, we serve cross-entropy of an example as its data difficulty and exploit the variance of distributions over the weights of the network to present the model uncertainty. Extensive experiments on various translation tasks reveal that our approach outperforms the strong baseline and related methods on both translation quality and convergence speed. Quantitative analyses reveal that the proposed strategy offers NMT the ability to automatically govern its learning schedule.
Yikai Zhou, Baosong Yang, Derek F. Wong, Yu Wan 0004, Lidia S. Chao
ACL3
2020 Self-Paced Learning for Neural Machine Translation
abstract
Recent studies have proven that the training of neural machine translation (NMT) can be facilitated by mimicking the learning process of humans.Nevertheless, achievements of such kind of curriculum learning rely on the quality of artificial schedule drawn up with the handcrafted features, e.g.sentence length or word rarity.We ameliorate this procedure with a more flexible manner by proposing self-paced learning, where NMT model is allowed to 1) automatically quantify the learning confidence over training examples; and 2) flexibly govern its learning via regulating the loss in each iteration step.Experimental results over multiple translation tasks demonstrate that the proposed model yields better performance than strong baselines and those models trained with human-designed curricula on both translation quality and convergence speed. 1
Yu Wan 0004, Baosong Yang, Derek F. Wong, Yikai Zhou, Lidia S. Chao, Haibo Zhang 0013, Boxing Chen
EMNLP (1)3
2020 Modeling Voting for System Combination in Machine Translation
abstract
System combination is an important technique for combining the hypotheses of different machine translation systems to improve translation performance. Although early statistical approaches to system combination have been proven effective in analyzing the consensus between hypotheses, they suffer from the error propagation problem due to the use of pipelines. While this problem has been alleviated by end-to-end training of multi-source sequence-to-sequence models recently, these neural models do not explicitly analyze the relations between hypotheses and fail to capture their agreement because the attention to a word in a hypothesis is calculated independently, ignoring the fact that the word might occur in multiple hypotheses. In this work, we propose an approach to modeling voting for system combination in machine translation. The basic idea is to enable words in hypotheses from different systems to vote on words that are representative and should get involved in the generation process. This can be done by quantifying the influence of each voter and its preference for each candidate. Our approach combines the advantages of statistical and neural methods since it can not only analyze the relations between hypotheses but also allow for end-to-end training. Experiments show that our approach is capable of better taking advantage of the consensus between hypotheses and achieves significant improvements over state-of-the-art baselines on Chinese-English and English-German machine translation tasks.
Xuancheng Huang, Zhixing Tan, Derek F. Wong, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001, Yang Liu 0005
IJCAI4
2020 Knowledge modeling via contextualized representations for LSTM-based personalized exercise recommendation
Derek F. Wong, Lionel M. Ni, Lidia S. Chao, Jing Zhang 0055
Inf. Sci.2
2020 HeTROPY: Explainable learning diagnostics via heterogeneous maximum-entropy and multi-spatial knowledge representation
Derek F. Wong, Lionel M. Ni, Lidia S. Chao, Jing Zhang 0055
Knowl. Based Syst.2
2020 Improving tree-based neural machine translation with dynamic lexicalized dependency encoding
Baosong Yang, Derek F. Wong, Lidia S. Chao, Min Zhang 0005
Knowl. Based Syst.2
2019 Context-Aware Self-Attention Networks
abstract
Self-attention model has shown its flexibility in parallel computation and the effectiveness on modeling both long- and short-term dependencies. However, it calculates the dependencies between representations without considering the contextual information, which has proven useful for modeling dependencies among neural representations in various natural language tasks. In this work, we focus on improving self-attention networks through capturing the richness of context. To maintain the simplicity and flexibility of the self-attention networks, we propose to contextualize the transformations of the query and key layers, which are used to calculate the relevance between elements. Specifically, we leverage the internal representations that embed both global and deep contexts, thus avoid relying on external resources. Experimental results on WMT14 English⇒German and WMT17 Chinese⇒English translation tasks demonstrate the effectiveness and universality of the proposed methods. Furthermore, we conducted extensive analyses to quantify how the context vectors participate in the self-attention model.
Baosong Yang, Jian Li 0054, Derek F. Wong, Lidia S. Chao, Xing Wang 0007, Zhaopeng Tu
AAAI3
2019 Shared-Private Bilingual Word Embeddings for Neural Machine Translation
abstract
Word embedding is central to neural machine translation (NMT), which has attracted intensive research interest in recent years.In NMT, the source embedding plays the role of the entrance while the target embedding acts as the terminal.These layers occupy most of the model parameters for representation learning.Furthermore, they indirectly interface via a soft-attention mechanism, which makes them comparatively isolated.In this paper, we propose shared-private bilingual word embeddings, which give a closer relationship between the source and target embeddings, and which also reduce the number of model parameters.For similar source and target words, their embeddings tend to share a part of the features and they cooperatively learn these common representation units.Experiments on 5 language pairs belonging to 6 different language families and written in 5 different alphabets demonstrate that the proposed model provides a significant performance boost over the strong baselines with dramatically fewer model parameters.
Xuebo Liu 0002, Derek F. Wong, Yang Liu 0005, Lidia S. Chao, Tong Xiao 0001
ACL (1)2
2019 Learning Deep Transformer Models for Machine Translation
abstract
Transformer is the state-of-the-art model in recent machine translation evaluations. Two strands of research are promising to improve models of this kind: the first uses wide networks (a.k.a. Transformer-Big) and has been the de facto standard for development of the Transformer system, and the other uses deeper language representation but faces the difficulty arising from learning deep networks. Here, we continue the line of research on the latter. We claim that a truly deep Transformer model can surpass the Transformer-Big counterpart by 1) proper use of layer normalization and 2) a novel way of passing the combination of previous layers to the next. On WMT’16 English-German and NIST OpenMT’12 Chinese-English tasks, our deep system (30/25-layer encoder) outperforms the shallow Transformer-Big/Base baseline (6-layer encoder) by 0.4-2.4 BLEU points. As another bonus, the deep model is 1.6X smaller in size and 3X faster in training than Transformer-Big.
Qiang Wang 0050, Tong Xiao 0001, Changliang Li, Derek F. Wong, Lidia S. Chao
ACL (1)6
2019 Leveraging Local and Global Patterns for Self-Attention Networks
abstract
Self-attention networks have received increasing research attention.By default, the hidden states of each word are hierarchically calculated by attending to all words in the sentence, which assembles global information.However, several studies pointed out that taking all signals into account may lead to overlooking neighboring information (e.g.phrase pattern).To address this argument, we propose a hybrid attention mechanism to dynamically leverage both of the local and global information.Specifically, our approach uses a gating scalar for integrating both sources of the information, which is also convenient for quantifying their contributions.Experiments on various neural machine translation tasks demonstrate the effectiveness of the proposed method.The extensive analyses verify that the two types of contexts are complementary to each other, and our method gives highly effective improvements in their integration.
Mingzhou Xu, Derek F. Wong, Baosong Yang, Yue Zhang 0004, Lidia S. Chao
ACL (1)2
2019 Assessing the Ability of Self-Attention Networks to Learn Word Order
abstract
Self-attention networks (SAN) have attracted a lot of interests due to their high parallelization and strong performance on a variety of NLP tasks, e.g. machine translation.Due to the lack of recurrence structure such as recurrent neural networks (RNN), SAN is ascribed to be weak at learning positional information of words for sequence modeling.However, neither this speculation has been empirically confirmed, nor explanations for their strong performances on machine translation tasks when "lacking positional information" have been explored.To this end, we propose a novel word reordering detection task to quantify how well the word order information learned by SAN and RNN.Specifically, we randomly move one word to another position, and examine whether a trained model can detect both the original and inserted positions.Experimental results reveal that: 1) SAN trained on word reordering detection indeed has difficulty learning the positional information even with the position embedding; and 2) SAN trained on machine translation learns better positional information than its RNN counterpart, in which position embedding plays a critical role.Although recurrence structure make the model more universally-effective on learning word order, learning objectives matter more in the downstream tasks such as machine translation.
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, Zhaopeng Tu
ACL (1)3
2019 Latent Attribute Based Hierarchical Decoder for Neural Machine Translation
abstract
Neural machine translation (NMT) has achieved state-of-the-art performance in many translation tasks. However, because the computational cost increases with the size of the search space for predicting the target words, the translation quality of NMT is constrained by the limited vocabulary. To alleviate this problem, we propose a novel dynamic hierarchical decoder for NMT to utilize all of the target words in the training and decoding process. In the proposed model, a target word is represented by two latent attribute vectors rather than a word vector. The model is trained to dynamically put together those words that share similar linguistic attributes. The prediction of a target word is, therefore, turned into the prediction of attribute vectors, where the $\mathrm{softmax}$ functions are performed at the attribute level. This greatly reduces the model size and the decoding time. Our experimental results demonstrate that the proposed model significantly outperforms the NMT baselines in both Chinese-English and English-German translation tasks.
Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Modeling Localness for Self-Attention Networks
abstract
Self-attention networks have proven to be of profound value for its strength of capturing global dependencies.In this work, we propose to model localness for self-attention networks, which enhances the ability of capturing useful local context.We cast localness modeling as a learnable Gaussian bias, which indicates the central and scope of the local region to be paid more attention.The bias is then incorporated into the original attention distribution to form a revised distribution.To maintain the strength of capturing long distance dependencies and enhance the ability of capturing shortrange dependencies, we only apply localness modeling to lower layers of self-attention networks.Quantitative and qualitative analyses on Chinese⇒English and English⇒German translation tasks demonstrate the effectiveness and universality of the proposed approach.
Baosong Yang, Zhaopeng Tu, Derek F. Wong, Fandong Meng, Lidia S. Chao, Tong Zhang 0001
EMNLP3
2018 Linguistic Knowledge-Aware Neural Machine Translation
abstract
Recently, researchers have shown an increasing interest in incorporating linguistic knowledge into neural machine translation (NMT). To this end, previous works choose either to alter the architecture of NMT encoder to incorporate syntactic information into the translation model, or to generalize the embedding layer of the encoder to encode additional linguistic features. The former approach mainly focuses on injecting the syntactic structure of the source sentence into the encoding process, leading to a complicated model that lacks the flexibility to incorporate other types of knowledge. The latter extends word embeddings by considering additional linguistic knowledge as features to enrich the word representation. It thus does not explicitly balance the contribution from word embeddings and the contribution from additional linguistic knowledge. To address these limitations, this paper proposes a knowledge-aware NMT approach that models additional linguistic features in parallel to the word feature. The core idea is that we propose modeling a series of linguistic features at the word level (knowledge block) using a recurrent neural network (RNN). And in sentence level, those word-corresponding feature blocks are further encoded using a RNN encoder. In decoding, we propose a knowledge gate and an attention gate to dynamically control the proportions of information contributing to the generation of target words from different sources. Extensive experiments show that our approach is capable of better accounting for importance of additional linguistic, and we observe significant improvements from 1.0 to 2.3 BLEU points on Chinese$\leftrightarrow$English and English$\rightarrow$German translation tasks.
Qiang Li 0022, Derek F. Wong, Lidia S. Chao, Muhua Zhu, Tong Xiao 0001, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Content-Oriented User Modeling for Personalized Response Ranking in Chatbots
abstract
Automatic chatbots (also known as chat-agents) have attracted much attention from both researching and industrial fields. Generally, the semantic relevance between users' queries and the corresponding responses is considered as the essential element for conversation modeling in both generation and ranking based chat systems. By contrast, it is a nontrivial task to adopt the users' information, such as preference, social role, etc., into conversational models reasonably, while users' profiles play a significant role in the procedure of conversations by providing the implicit contexts. This paper aims to address the personalized response ranking task by incorporating user profiles into the conversation model. In our approach, users' personalized representations are latently learned from the contents posted by them via a two-branch neural network. After that, a deep neural network architecture is further presented to learn the fusion representation of posts, responses, and personal information. In this way, the proposed model could understand conversations from the users' perspective; hence, the more appropriate responses are selected for a specified person. The experimental results on two datasets from social network services demonstrate that our approach is hopeful to represent users' personal information implicitly based on user generated contents, and it is promising to perform as an important component in chatbots to select the personalized responses for each user.
Bingquan Liu, Zhen Xu 0003, Chengjie Sun, Baoxun Wang, Xiaolong Wang 0001, Derek F. Wong, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.6
2017 Towards Bidirectional Hierarchical Representations for Attention-based Neural Machine Translation
abstract
This paper proposes a hierarchical attentional neural translation model which focuses on enhancing source-side hierarchical representations by covering both local and global semantic information using a bidirectional tree-based encoder.To maximize the predictive likelihood of target words, a weighted variant of an attention mechanism is used to balance the attentive information between lexical and phrase vectors.Using a tree-based rare word encoding, the proposed model is extended to sub-word level to alleviate the out-of-vocabulary (OOV) problem.Empirical results reveal that the proposed model significantly outperforms sequence-to-sequence attention-based and tree-based neural translation models in English-Chinese translation tasks.
Baosong Yang, Derek F. Wong, Tong Xiao 0001, Lidia S. Chao
EMNLP2
2016 Bilingual recursive neural network based data selection for statistical machine translation
Derek F. Wong, Yi Lu 0005, Lidia S. Chao
Knowl. Based Syst.1
2016 A Loss-Augmented Approach to Training Syntactic Machine Translation Systems
abstract
Current syntactic machine translation (MT) systems implicitly use beam-width unlimited search in learning model parameters (e.g., feature values for each translation rule). However, a limited beam-width has to be adopted in decoding new sentences, and the MT output is in general evaluated by various metrics, such as BLEU and TER. In this paper, we address: 1) the mismatch of adopted beam-widths between training and decoding; and 2) the mismatch of training criteria and MT evaluation metrics. Unlike previous work, we model the two problems in a single training paradigm simultaneously. We design a loss-augmented approach that explicitly considers the limited beam-width and evaluation metric in training, and present a simple but effective method to learn the model. By using beam search and BLEU-related losses, our approach improves a state-of-the-art syntactic MT system by +1.0 BLEU on Chinese-to-English and English-to-Chinese translation tasks. It even outperforms seven previous training approaches over 0.8 BLEU points. More interestingly, promising improvements are observed when our approach works with TER.
Tong Xiao 0001, Derek F. Wong
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Graph-Based Lexicon Regularization for PCFG With Latent Annotations
abstract
This paper aims at learning a better probabilistic context-free grammar with latent annotations (PCFG-LA) by using a graph propagation (GP) technique. We propose leveraging the GP to regularize the lexical model of the grammar. The proposed approach constructs k-nearest neighbor ( k-NN) similarity graphs over words with identical pre-terminal (part-of-speech) tags, for propagating the probabilities of latent annotations given the words. The graphs demonstrate the relationship between words in syntactic and semantic levels, estimated by using a neural word representation method based on Recursive autoencoder (RAE). We modify the conventional PCFG-LA parameter estimation algorithm, expectation maximization (EM), by incorporating a GP process subsequent to the M-step. The GP encourages the smoothness among the graph vertices, where different words under similar syntactic and semantic environments should have approximate posterior distributions of nonterminal subcategories. The proposed PCFG-LA learning approach was evaluated together with a hierarchical split-and-merge training strategy, on parsing tasks for English, Chinese and Portuguese. The empirical results reveal two crucial findings: 1) regularizing the lexicons with GP results in positive effects to parsing accuracy; and 2) learning with unlabeled data can also expand the PCFG-LA lexicons.
Xiaodong Zeng, Derek F. Wong, Lidia S. Chao, Isabel Trancoso
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Toward Better Chinese Word Segmentation for SMT via Bilingual Constraints
abstract
This study investigates on building a better Chinese word segmentation model for statistical machine translation.It aims at leveraging word boundary information, automatically learned by bilingual character-based alignments, to induce a preferable segmentation model.We propose dealing with the induced word boundaries as soft constraints to bias the continuous learning of a supervised CRFs model, trained by the treebank data (labeled), on the bilingual data (unlabeled).The induced word boundary information is encoded as a graph propagation constraint.The constrained model induction is accomplished by using posterior regularization algorithm.The experiments on a Chinese-to-English machine translation task reveal that the proposed model can bring positive segmentation effects to translation quality.
Xiaodong Zeng, Lidia S. Chao, Derek F. Wong, Isabel Trancoso
ACL (1)3
2014 UM-Corpus: A Large English-Chinese Parallel Corpus for Statistical Machine Translation
Derek F. Wong, Lidia S. Chao, Paulo Quaresma, Francisco Oliveira 0002
LREC2
2014 Lexicon expansion for latent variable grammars
Xiaodong Zeng, Derek F. Wong, Lidia S. Chao, Isabel Trancoso, Liangye He, Qiuping Huang
Pattern Recognit. Lett.2
2013 Graph-based Semi-Supervised Model for Joint Chinese Word Segmentation and Part-of-Speech Tagging
Xiaodong Zeng, Derek F. Wong, Lidia S. Chao, Isabel Trancoso
ACL (1)2
2013 Influence of Part-of-Speech and Phrasal Category Universal Tag-set in Tree-to-Tree Translation Models
Francisco Oliveira 0002, Derek F. Wong, Lidia S. Chao, Liangye He
IJCNLP2
2013 Augmented Parsing of Unknown Word by Graph-Based Semi-Supervised Learning
Qiuping Huang, Derek F. Wong, Lidia S. Chao, Xiaodong Zeng, Liangye He
PACLIC2