EDBT 2026 Demo / reviewers in the wild / expert
Runzhe Zhan
dblp:286/8257
· DBLP profile ↗
22ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0002-3175-6885ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 3 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine TranslationabstractREtrieval-Augmented LLM-based Machine Translation (REAL-MT) shows promise for knowledge-intensive tasks like idiomatic translation, but its reliability under noisy retrieval, a common challenge in real-world deployment, remains poorly understood. To address this gap, we propose a noise synthesis framework and new metrics to systematically evaluate REAL-MT’s reliability across high-, medium-, and low-resource language pairs. Using both open- and closed-sourced models, including standard LLMs and large reasoning models (LRMs), we find that models heavily rely on retrieved context, and this dependence is significantly more detrimental in low-resource language pairs, producing nonsensical translations. Although LRMs possess enhanced reasoning capabilities, they show no improvement in error correction and are even more susceptible to noise, tending to rationalize incorrect contexts. Attention analysis reveals a shift from the source idiom to noisy content, while confidence increases despite declining accuracy, indicating poor self-monitoring. To mitigate these issues, we investigate training-free and fine-tuning strategies, which improve robustness at the cost of performance in clean contexts, revealing a fundamental trade-off. Our findings highlight the limitations of current approaches, underscoring the need for self-verifying integration mechanisms. Runzhe Zhan, Chi Seng Cheang, Xuebo Liu 0002, Yuyao Niu, Fengying Ye, Kaixin Lan, Lidia S. Chao, Derek F. Wong |
AAAI | 2 |
| 2026 | VisAidMath: Benchmarking Visual-Aided Mathematical ReasoningabstractJingkun Ma, Runzhe Zhan, Yang Li, Di Sun, Hou Pong Chan, Lidia S. Chao, Derek F. Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jingkun Ma, Runzhe Zhan, Hou Pong Chan, Lidia S. Chao, Derek F. Wong |
ACL (1) | 2 |
| 2026 | G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom AlignmentabstractIdioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable.We present G-IdiomAlign, a gloss-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary.We further construct a high-confidence reference alignment set for reproducible evaluation.G-IdiomAlign supports two protocols: (1) a controlled Multiple-Choice Idiom Equivalence with typed distractors for error attribution; and(2) a Gloss-Contrastive Generation contrasting No-gloss and With-gloss inputs to isolate the effect of an explicit semantic pivot.Across diverse LLMs, a bias to literal translation is a dominant failure mode, especially when the target is a low-resource language.Glosses consistently improve Gloss-Contrastive Generation under an embedding-based semantic proxy, but performance remains modest, indicating substantial headroom in the open output space.Subsequent analysis on Qwen3-8B further suggests that cross-condition differences are concentrated more in attention heads than in layers, while better With-gloss generations coincide with stronger gloss anchoring 1 . Fengying Ye, Runzhe Zhan, Lidia S. Chao, Zheqi Zhang, Derek F. Wong |
ACL (1) | 3 |
| 2026 | Domain Adaptive Machine Translation with Synthetic Feedback for Large Language ModelsabstractDomain-specific machine translation (MT) significantly benefits from large language models (LLMs) due to their strong instruction-following abilities and in-context learning (ICL) capabilities. Appropriate demonstration samples and feedback are essential for helping LLMs refine their translation outputs in real-world applications. However, the scarcity of in-domain samples and professional feedback creates practical limitations. Furthermore, the current ICL paradigm does not offer the fine-grained domain features in addition to parallel translation pairs. To address these challenges, we propose a pipeline that collects in-domain translations from LLMs and generates synthetic, human-like feedback for revising these translations. The translations and their corresponding feedback are stored together to build a demonstration database, with each instance paired with the original in-domain translation and its revision. During online translation, similar in-domain translations can be retrieved as revision demonstrations. This process guides LLMs in iteratively refining their outputs by learning from demonstrations. We evaluate the proposed pipeline using open-source models like Llama3-8B-Instruct and Mistral-7B-Instruct-v0.3, on five domain-specific benchmarks for English-centric, Chinese-centric, and Portuguese-centric translation. The results demonstrate the effectiveness of the pipeline in tailoring in-domain translations and improving translation performance compared to direct translation instructions. Additionally, we discuss the experimental results from the following perspectives: (1) the effectiveness of different in-context retrieval methods; (2) the observed differences across selected domains and language; (3) the quantitative analysis of sentence-level and word-level statistics; and (4) the effect of ICL retrieval database size and decoding parameters. Xinyi Yang 0008, Runzhe Zhan, Junchao Wu, Yue Zhang 0004, Xuebo Liu 0002, Lidia S. Chao, Derek F. Wong |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2025 | Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScoreabstractThe efficacy of detectors for texts generated by large language models (LLMs) substantially depends on the availability of large-scale training data. However, white-box zero-shot detectors, which require no such data, are limited by the accessibility of the source model of the LLM-generated text. In this paper, we propose a simple yet effective black-box zero-shot detection approach based on the observation that, from the perspective of LLMs, human-written texts typically contain more grammatical errors than LLM-generated texts. This approach involves calculating the Grammar Error Correction Score (GECScore) for the given text to differentiate between human-written and LLM-generated text. Experimental results show that our method outperforms current state-of-the-art (SOTA) zero-shot and supervised methods, achieving an average AUROC of 98.62% across XSum and Writing Prompts dataset. Additionally, our approach demonstrates strong reliability in the wild, exhibiting robust generalization and resistance to paraphrasing attacks. Data and code are available at: https://github.com/NLP2CT/GECScore. Junchao Wu, Runzhe Zhan, Derek F. Wong, Shu Yang 0010, Xuebo Liu 0002, Lidia S. Chao, Min Zhang 0005 |
COLING | 2 |
| 2025 | Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language ModelabstractLarge Language Models (LLMs) are composed of neurons that exhibit various behaviors and roles, which become increasingly diversified as models scale. Recent studies have revealed that not all neurons are active across different datasets, and this sparsity correlates positively with the task-specific ability, leading to advancements in model pruning and training efficiency. Traditional fine-tuning methods engage all parameters of LLMs, which is computationally expensive and may not be necessary. In contrast, Parameter-Efficient Fine-Tuning (PEFT) approaches aim to minimize the number of trainable parameters, yet they still operate at a relatively macro scale (e.g., layer-level). We introduce Neuron-Level Fine-Tuning (NeFT), a novel approach that refines the granularity of parameter training down to the individual neuron, enabling a more parameter-efficient fine-tuning model. The experimental results show that NeFT not only exceeded the performance of full-parameter fine-tuning and PEFT but also provided insights into the analysis of neurons. Our code and data are available at: https://github.com/NLP2CT/NeFT. Haoyun Xu, Runzhe Zhan, Yingpeng Ma, Derek F. Wong, Lidia S. Chao |
COLING | 2 |
| 2025 | Path Drift in Large Reasoning Models: How First-Person Commitments Override SafetyabstractAs large language models (LLMs) are increasingly deployed for complex reasoning tasks, Long Chain-of-Thought (Long-CoT) prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards enabled by alignment techniques such as RLHF, we identify a previously underexplored vulnerability: reasoning trajectories in Long-CoT models can drift from aligned paths, resulting in content that violates safety constraints. We term this phenomenon Path Drift. Through empirical analysis, we uncover three behavioral triggers of Path Drift: (1) first-person commitments that induce goal-driven reasoning that delays refusal signals; (2) ethical evaporation, where surface-level disclaimers bypass alignment checkpoints; (3) condition chain escalation, where layered cues progressively steer models toward unsafe completions. Building on these insights, we introduce a three-stage Path Drift Induction Framework comprising cognitive load amplification, self-role priming, and condition chain hijacking. Each stage independently reduces refusal rates, while their combination further compounds the effect. To mitigate these risks, we propose a path-level defense strategy incorporating role attribution correction and metacognitive reflection (reflective safety cues). Our findings highlight the need for trajectory-level alignment oversight in long-form reasoning beyond token-level alignment. Yuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao, Derek F. Wong |
EMNLP | 2 |
| 2025 | Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance BoostabstractRecent advancements in large reasoning models (LRMs) have introduced an intermediate "thinking" process prior to generating final answers, improving their reasoning capabilities on complex downstream tasks. However, the potential of LRMs as evaluators for machine translation (MT) quality remains underexplored. We provides the first systematic analysis of LRM-as-a-judge in MT evaluation. We identify key challenges, revealing LRMs require tailored evaluation materials, tend to "overthink" simpler instances and have issues with scoring mechanisms leading to overestimation. To address these, we propose to calibrate LRM thinking by training them on synthetic, human-like thinking trajectories. Our experiments on WMT24 Metrics benchmarks demonstrate that this approach largely reduces thinking budgets by ~35x while concurrently improving evaluation performance across different LRM scales from 7B to 32B (e.g., R1-Distill-Qwen-7B achieves a +8.7 correlation point improvement). These findings highlight the potential of efficiently calibrated LRMs to advance fine-grained automatic MT evaluation. Runzhe Zhan, Xinyi Yang 0008, Lidia S. Chao, Min Yang 0007, Derek F. Wong |
NeurIPS | 1 |
| 2025 | Overview of the NLPCC 2025 Shared Task 1: LLM-Generated Text Detection
Junchao Wu, Runzhe Zhan, Yulin Yuan, Lidia S. Chao, Derek F. Wong |
NLPCC (4) | 2 |
| 2025 | A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future DirectionsabstractAbstract The remarkable ability of large language models (LLMs) to comprehend, interpret, and generate complex language has rapidly integrated LLM-generated text into various aspects of daily life, where users increasingly accept it. However, the growing reliance on LLMs underscores the urgent need for effective detection mechanisms to identify LLM-generated text. Such mechanisms are critical to mitigating misuse and safeguarding domains like artistic expression and social networks from potential negative consequences. LLM-generated text detection, conceptualized as a binary classification task, seeks to determine whether an LLM produced a given text. Recent advances in this field stem from innovations in watermarking techniques, statistics-based detectors, and neural-based detectors. Human-assisted methods also play a crucial role. In this survey, we consolidate recent research breakthroughs in this field, emphasizing the urgent need to strengthen detector research. Additionally, we review existing datasets, highlighting their limitations and developmental requirements. Furthermore, we examine various LLM-generated text detection paradigms, shedding light on challenges like out-of-distribution problems, potential attacks, real-world data issues, and ineffective evaluation frameworks. Finally, we outline intriguing directions for future research in LLM-generated text detection to advance responsible artificial intelligence. This survey aims to provide a clear and comprehensive introduction for newcomers while offering seasoned researchers valuable updates in the field.1 Junchao Wu, Shu Yang 0010, Runzhe Zhan, Yulin Yuan, Lidia S. Chao, Derek F. Wong |
Comput. Linguistics | 3 |
| 2025 | RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation PatternsabstractAbstract Detecting content generated by large language models (LLMs) is crucial for preventing misuse and building trustworthy AI systems. Although existing detection methods perform well, their robustness in out-of-distribution (OOD) scenarios is still lacking. In this paper, we hypothesize that, compared to features used by existing detection methods, the internal representations of LLMs contain more comprehensive and raw features that can more effectively capture and distinguish the statistical pattern differences between LLM-generated texts (LGT) and human-written texts (HWT). We validated this hypothesis across different LLMs and observed significant differences in neural activation patterns when processing these two types of texts. Based on this, we propose RepreGuard, an efficient statistics-based detection method. Specifically, we first employ a surrogate model to collect representation of LGT and HWT, and extract the distinct activation feature that can better identify LGT. We can classify the text by calculating the projection score of the text representations along this feature direction and comparing with a precomputed threshold. Experimental results show that RepreGuard outperforms all baselines with average 94.92% AUROC on both in-distribution and OOD scenarios, while also demonstrating robust resilience to various text sizes and mainstream attacks.1 Xin Chen 0032, Junchao Wu, Shu Yang 0010, Runzhe Zhan, Di Wang 0015, Min Yang 0007, Lidia S. Chao, Derek F. Wong |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | Can LLMs Learn Uncertainty on Their Own? Expressing Uncertainty Effectively in A Self-Training MannerabstractLarge language models (LLMs) often exhibit excessive, random, and uninformative uncertainty, rendering them unsuitable for decisionmaking in human-computer interactions.In this paper, we aim to instigate a heightened awareness of self-uncertainty in LLMs, enabling them to express uncertainty more effectively.To accomplish this, we propose an uncertainty-aware instruction tuning (UaIT) method, aligning LLMs' perception with the probabilistic uncertainty of the generation.We conducted experiments using LLaMA2 and Mistral on multiple free-form QA tasks.Experimental results revealed a surprising 45.2% improvement in the effectiveness of uncertainty expression by LLMs, accompanied by reasonably good out-of-domain generalization capabilities.Moreover, this uncertainty expression can serve as a valuable real-time basis for human decision-making, e.g., retrieving external documents and incorporating stronger LLMs 1 . Shudong Liu 0004, Zhaocong Li, Xuebo Liu 0002, Runzhe Zhan, Derek F. Wong, Lidia S. Chao, Min Zhang 0005 |
EMNLP | 4 |
| 2024 | DetectRL: Benchmarking LLM-Generated Text Detection in Real-World ScenariosabstractDetecting text generated by large language models (LLMs) is of great recent interest. With zero-shot methods like DetectGPT, detection capabilities have reached impressive levels. However, the reliability of existing detectors in real-world applications remains underexplored. In this study, we present a new benchmark, DetectRL, highlighting that even state-of-the-art (SOTA) detection techniques still underperformed in this task. We collected human-written datasets from domains where LLMs are particularly prone to misuse. Using popular LLMs, we generated data that better aligns with real-world applications. Unlike previous studies, we employed heuristic rules to create adversarial LLM-generated text, simulating advanced prompt usages, human revisions like word substitutions, and writing errors. Our development of DetectRL reveals the strengths and limitations of current SOTA detectors. More importantly, we analyzed the potential impact of writing styles, model types, attack methods, the text lengths, and real-world human writing factors on different types of detectors. We believe DetectRL could serve as an effective benchmark for assessing detectors in real-world scenarios, evolving with advanced attack methods, thus providing more stressful evaluation to drive the development of more efficient detectors\footnote{Data and code are publicly available at: https://github.com/NLP2CT/DetectRL. Junchao Wu, Runzhe Zhan, Derek F. Wong, Shu Yang 0010, Xinyi Yang 0008, Yulin Yuan, Lidia S. Chao |
NeurIPS | 2 |
| 2024 | Activate Integrated Controllable Generation with Soft Prompt
Jingkun Ma, Runzhe Zhan, Derek F. Wong, Lidia S. Chao |
NLPCC (4) | 2 |
| 2024 | Understanding and Improving Low-Resource Neural Machine Translation with Shallow Features
Xuebo Liu 0002, Derek F. Wong, Yuchu Lin, Runzhe Zhan, Lidia S. Chao, Min Zhang 0005 |
NLPCC (3) | 6 |
| 2024 | Dynamic curriculum learning for conversation response selection
Guanhua Chen 0006, Runzhe Zhan, Derek F. Wong, Lidia S. Chao |
Knowl. Based Syst. | 2 |
| 2023 | Revisiting Commonsense Reasoning in Machine Translation: Training, Evaluation and ChallengeabstractThe ability of commonsense reasoning (CR) decides whether a neural machine translation (NMT) model can move beyond pattern recognition.Despite the rapid advancement of NMT and the use of pretraining to enhance NMT models, research on CR in NMT is still in its infancy, leaving much to be explored in terms of effectively training NMT models with high CR abilities and devising accurate automatic evaluation metrics.This paper presents a comprehensive study aimed at expanding the understanding of CR in NMT.For the training, we confirm the effectiveness of incorporating pretrained knowledge into NMT models and subsequently utilizing these models as robust testbeds for investigating CR in NMT.For the evaluation, we propose a novel entity-aware evaluation method that takes into account both the NMT candidate and important entities in the candidate, which is more aligned with human judgement.Based on the strong testbed and evaluation methods, we identify challenges in training NMT models with high CR abilities and suggest directions for further unlabeled data utilization and model design.We hope that our methods and findings will contribute to advancing the research of CR in NMT. Xuebo Liu 0002, Derek F. Wong, Runzhe Zhan, Liangxuan Yu, Min Zhang 0005 |
ACL (1) | 4 |
| 2023 | Test-time Adaptation for Machine Translation Evaluation by Uncertainty MinimizationabstractThe neural metrics recently received considerable attention from the research community in the automatic evaluation of machine translation.Unlike text-based metrics that have interpretable and consistent evaluation mechanisms for various data sources, the reliability of neural metrics in assessing out-of-distribution data remains a concern due to the disparity between training data and real-world data.This paper aims to address the inference bias of neural metrics through uncertainty minimization during test time, without requiring additional data.Our proposed method comprises three steps: uncertainty estimation, test-time adaptation, and inference.Specifically, the model employs the prediction uncertainty of the current data as a signal to update a small fraction of parameters during test time and subsequently refine the prediction through optimization.To validate our approach, we apply the proposed method to three representative models and conduct experiments on the WMT21 benchmarks.The results obtained from both in-domain and out-of-distribution evaluations consistently demonstrate improvements in correlation performance across different models.Furthermore, we provide evidence that the proposed method effectively reduces model uncertainty.The code is publicly available at https://github.com/NLP2CT/TaU. Runzhe Zhan, Xuebo Liu 0002, Derek F. Wong, Cuilian Zhang, Lidia S. Chao, Min Zhang 0005 |
ACL (1) | 1 |
| 2023 | Towards Zero-Shot Multilingual Poetry TranslationabstractThe application of machine translation in the field of poetry has always presented significant challenges. Conventional machine translation techniques are inadequate for capturing and translating the unique style of poetry. The absence of a parallel poetry corpus and the distinctive structure of poetry further restrict the effectiveness of traditional methods. This paper introduces a zero-shot method that is capable of translating poetry style without the need for a large-scale training corpus. Specifically, we treat poetry translation as a standard machine translation problem and subsequently inject the poetry style upon completion of the translation process. Our injection model only requires back-translation and easily obtainable monolingual data, making it a low-cost solution. We conducted experiments on three translation directions and presented automatic and human evaluations, demonstrating that our proposed method outperforms existing online systems and other competitive baselines. These results validate the feasibility and potential of our proposed approach and provide new prospects for poetry translation. Wai Lei Song, Haoyun Xu, Derek F. Wong, Runzhe Zhan, Lidia S. Chao, Shanshan Wang 0009 |
MTSummit (1) | 4 |
| 2023 | Multi-Level Curriculum Learning for Multi-Turn Dialogue GenerationabstractSince deep learning is the dominant paradigm in the multi-turn dialogue generation task, large-scale training data is the key factor affecting the model performance. To make full use of the training data, the existing work directly applied curriculum learning to the multi-turn dialogue generation task, training model in a “easy-to-hard” way. But the design of the current methodology does not consider dialogue-specific features. To close this gap, we propose a Multi-Level Curriculum Learning (MLCL) method for multi-turn dialogue generation by considering the word-level linguistic feature and utterance-level semantic relation in a dialogue. The motivation is that word-level knowledge is beneficial to understanding complex utterance-level dependency of dialogue. Thus, we design two difficulty measurements and a self-adaptive curriculum scheduler, making the model gradually shift the learning focus from word-level to utterance-level information during the training process. We also verify the independence and complementarity of the two measurements at different levels. We evaluate the performance on two widely used multi-turn dialogue datasets, and the results demonstrate that our proposed method outperforms the strong baselines and existing CL methods in terms of automated metrics and human evaluation. We will release the code files upon acceptance. Guanhua Chen 0006, Runzhe Zhan, Derek F. Wong, Lidia S. Chao |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Obscurity-Quantified Curriculum Learning for Machine Translation EvaluationabstractThe pre-trained language model has been developed for evaluating the quality of machine translation. It achieves state-of-the-art results. However, building a model for the evaluation of machine translation still faces the following challenges: 1) large scale of the training data affects the speed of the optimization; 2) the varied quality of the training data makes the optimization process unstable. To alleviate the issues of data learning, curriculum learning is proposed to rearrange the training sequence following an “easy-to-hard” process. However, the definition of difficulty can not be directly applied to the training data used in the machine translation evaluation. Hence, we propose an obscurity-quantified curriculum learning framework for this task. Specifically, the obscurity of each training example can be measured from multiple perspectives, including thedifficulty of ranking, thefuzziness of reference, thecomplexity of text, and theunreliability of judgement. To incorporate the obscurity measurements, we also design a dynamic learning strategy to guide the training process from instances with low obscurity to those with high-obscurity. Experimental results show that our proposed methods yield remarkable improvements on the segment-level WMT2019 and WMT2020 Metrics Shared Tasks compared to other baseline methods. Cuilian Zhang, Derek F. Wong, Eddy Sio Kei Lei, Runzhe Zhan, Lidia S. Chao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Meta-Curriculum Learning for Domain Adaptation in Neural Machine TranslationabstractMeta-learning has been sufficiently validated to be beneficial for low-resource neural machine translation (NMT). However, we find that meta-trained NMT fails to improve the translation performance of the domain unseen at the meta-training stage. In this paper, we aim to alleviate this issue by proposing a novel meta-curriculum learning for domain adaptation in NMT. During meta-training, the NMT first learns the similar curricula from each domain to avoid falling into a bad local optimum early, and finally learns the curricula of individualities to improve the model robustness for learning domain-specific knowledge. Experimental results on 10 different low-resource domains show that meta-curriculum learning can improve the translation performance of both familiar and unfamiliar domains. All the codes and data are freely available at https://github.com/NLP2CT/Meta-Curriculum. Runzhe Zhan, Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao |
AAAI | 1 |