EDBT 2026 Demo / reviewers in the wild / expert
Pei Zhang 0011
dblp:78/5323-11
· DBLP profile ↗
16ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Systematic Assessment of Language Models with Linguistic Minimal Pairs in ChineseabstractAbstract We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the Ba construction. We then train from scratch a suite of Chinese language models (LMs) with different tokenizers, parameter sizes, and token volumes, to study the learning curves of LMs on Chinese. To mitigate the biases introduced by unequal lengths of the sentences in a minimal pair, we propose a new metric named sub-linear length normalized log-probabilities (SLLN-LP). Using SLLN-LP as the metric, our results show that Anaphor, Quantifiers, and Ellipsis in Chinese are difficult for LMs even up to 32B parameters, and that SLLN-LP successfully mitigates biases in ZhoBLiMP, JBLiMP and BLiMP. We conclude that future evaluations should be more carefully designed to consider the intricate relations between linking functions, LMs, and targeted minimal pairs. Yikang Liu 0002, Yeting Shen, Hongao Zhu, Lilong Xu, Zhiheng Qian, Siyuan Song, Jialong Tang, Pei Zhang 0011, Baosong Yang, Rui Wang 0015, Hai Hu 0001 |
Trans. Assoc. Comput. Linguistics | 9 |
| 2025 | Enhancing Machine Translation with Self-Supervised Preference DataabstractModel alignment methods like Direct Preference Optimization and Contrastive Preference Optimization have enhanced machine translation performance by leveraging preference data to enable models to reject suboptimal outputs. During preference data construction, previous approaches primarily rely on humans, strong models like GPT4 or model self-sampling. In this study, we first explain the shortcomings of this practice. Then, we propose Self-Supervised Preference Optimization (SSPO), a novel framework which efficiently constructs translation preference data for iterative DPO training. Applying SSPO to 14B parameters large language models (LLMs) achieves comparable or better performance than GPT-4o on FLORES and multi-domain test datasets. We release an augmented MQM dataset in https://github.com/sunny-sjtu/MQM-aug. Haoxiang Sun, Pei Zhang 0011, Baosong Yang, Rui Wang 0015 |
ACL (1) | 3 |
| 2025 | Locate-and-Focus: Enhancing Terminology Translation in Speech Language ModelsabstractDirect speech translation (ST) has garnered increasing attention nowadays, yet the accurate translation of terminology within utterances remains a great challenge. In this regard, current studies mainly concentrate on leveraging various translation knowledge into ST models. However, these methods often struggle with interference from irrelevant noise and can not fully utilize the translation knowledge. To address these issues, in this paper, we propose a novel Locate-and-Focus method for terminology translation. It first effectively locates the speech clips containing terminologies within the utterance to construct translation knowledge, minimizing irrelevant information for the ST model. Subsequently, it associates the translation knowledge with the utterance and hypothesis from both audio and textual modalities, allowing the ST model to better focus on translation knowledge during translation. Experimental results across various datasets demonstrate that our method effectively locates terminologies within utterances and enhances the success rate of terminology translation, while maintaining robust general translation performance. Suhang Wu, Jialong Tang, Pei Zhang 0011, Baosong Yang, Junhui Li 0001, Junfeng Yao, Min Zhang 0005, Jinsong Su |
ACL (1) | 4 |
| 2025 | Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of TranslationeseabstractYikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang, Pei Zhang, Baosong Yang, Fei Huang, Rui Wang, Hai Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yikang Liu 0002, Wanyang Zhang, Yiming Wang 0011, Jialong Tang, Pei Zhang 0011, Baosong Yang, Fei Huang 0002, Rui Wang 0015, Hai Hu 0001 |
EMNLP | 5 |
| 2025 | Latent Space Chain-of-Embedding Enables Output-free LLM Self-EvaluationabstractLLM self-evaluation relies on the LLM's own ability to estimate response correctness, which can greatly improve its deployment reliability.
In this research track, we propose the Chain-of-Embedding (CoE) in the latent space to enable LLMs to perform output-free self-evaluation. CoE consists of all progressive hidden states produced during the inference time, which can be treated as the latent thinking path of LLMs. We find that when LLMs respond correctly and incorrectly, their CoE features differ, these discrepancies assist us in estimating LLM response correctness. Experiments in four diverse domains and seven LLMs fully demonstrate the effectiveness of our method. Meanwhile, its label-free design intent without any training and millisecond-level computational cost ensure real-time feedback in large-scale scenarios.
More importantly, we provide interesting insights into LLM response correctness from the perspective of hidden state changes inside LLMs. Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Derek F. Wong, Rui Wang 0015 |
ICLR | 2 |
| 2025 | ConText: Driving In-context Learning for Text Removal and SegmentationabstractThis paper presents the first study on adapting the visual in-context learning (V-ICL) paradigm to optical character recognition tasks, specifically focusing on text removal and segmentation. Most existing V-ICL generalists employ a reasoning-as-reconstruction approach: they turn to using a straightforward image-label compositor as the prompt and query input, and then masking the query label to generate the desired output. This direct prompt confines the model to a challenging single-step reasoning process. To address this, we propose a task-chaining compositor in the form of image-removal-segmentation, providing an enhanced prompt that elicits reasoning with enriched intermediates. Additionally, we introduce context-aware aggregation, integrating the chained prompt pattern into the latent query representation, thereby strengthening the model’s in-context reasoning. We also consider the issue of visual heterogeneity, which complicates the selection of homogeneous demonstrations in text recognition. Accordingly, this is effectively addressed through a simple self-prompting strategy, preventing the model’s in-context learnability from devolving into specialist-like, context-free inference. Collectively, these insights culminate in our ConText model, which achieves new state-of-the-art across both in- and out-of-domain benchmarks. The code is available at https://github.com/Ferenas/ConText. Fei Zhang 0016, Pei Zhang 0011, Baosong Yang, Fei Huang 0002, Yanfeng Wang 0001, Ya Zhang 0002 |
ICML | 2 |
| 2025 | Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early DecodingabstractTest-time scaling enhances large language model performance by allocating additional compute resources during decoding. Best-of-$N$ (BoN) sampling serves as a common sampling-based scaling technique, broadening the search space in parallel to find better solutions from the model distribution. However, its cost–performance trade-off is still underexplored. Two main challenges limit the efficiency of BoN sampling:
(1) Generating $N$ full samples consumes substantial GPU memory, reducing inference capacity under limited resources.
(2) Reward models add extra memory and latency overhead, and training strong reward models introduces potential training data costs.
Although some studies have explored efficiency improvements, none have addressed both challenges at once.
To address this gap, we propose **Self-Truncation Best-of-$N$ (ST-BoN)**, a decoding method that avoids fully generating all $N$ samples and eliminates the need for reward models. It leverages early sampling consistency in the model’s internal states to identify the most promising path and truncate suboptimal ones.
In terms of cost, ST-BoN reduces dynamic GPU memory usage by over 80% and inference latency by 50%.
In terms of cost–performance trade-off, ST-BoN achieves the same performance as Full-BoN while saving computational cost by 70%–80%, and under the same cost, it can improve accuracy by 3–4 points. Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Zhuosheng Zhang 0001, Fei Huang 0002, Rui Wang 0015 |
NeurIPS | 2 |
| 2025 | PolyMath: Evaluating Mathematical Reasoning in Multilingual ContextsabstractIn this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual mathematical benchmark in the era of reasoning LLMs.We conduct a comprehensive evaluation for advanced LLMs and find that even Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, achieve only 54.6 and 52.2 benchmark scores, with about 40% accuracy under the highest level.From a language perspective, our benchmark reveals several key challenges of LLMs in multilingual reasoning:(1) Reasoning performance varies widely across languages for current LLMs;(2) Input-output language consistency is low in reasoning LLMs and may be correlated with performance;(3) The thinking length differs significantly by language for current LLMs.Additionally, we demonstrate that controlling the output language in the instructions has the potential to affect reasoning performance, especially for some low-resource languages, suggesting a promising direction for improving multilingual capabilities in LLMs. Yiming Wang 0011, Pei Zhang 0011, Jialong Tang, Baosong Yang, Rui Wang 0015, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang 0002, Junyang Lin, Fei Huang 0005, Jingren Zhou 0001 |
NeurIPS | 2 |
| 2024 | Embedding Trajectory for Out-of-Distribution Detection in Mathematical ReasoningabstractReal-world data deviating from the independent and identically distributed (\textit{i.i.d.}) assumption of in-distribution training data poses security threats to deep networks, thus advancing out-of-distribution (OOD) detection algorithms. Detection methods in generative language models (GLMs) mainly focus on uncertainty estimation and embedding distance measurement, with the latter proven to be most effective in traditional linguistic tasks like summarization and translation. However, another complex generative scenario mathematical reasoning poses significant challenges to embedding-based methods due to its high-density feature of output spaces, but this feature causes larger discrepancies in the embedding shift trajectory between different samples in latent spaces. Hence, we propose a trajectory-based method TV score, which uses trajectory volatility for OOD detection in mathematical reasoning. Experiments show that our method outperforms all traditional algorithms on GLMs under mathematical reasoning scenarios and can be extended to more applications with high-density features in output spaces, such as multiple-choice questions. Yiming Wang 0011, Pei Zhang 0011, Baosong Yang, Derek F. Wong, Zhuosheng Zhang 0001, Rui Wang 0015 |
NeurIPS | 2 |
| 2023 | Imitation Attacks Can Steal More Than You Think from Machine Translation Systems
Tianxiang Hu, Pei Zhang 0011, Baosong Yang, Rui Wang 0015 |
NLPCC (1) | 2 |
| 2022 | Competency-Aware Neural Machine Translation: Can Machine Translation Know its Own Translation Quality?abstractNeural machine translation (NMT) is often criticized for failures that happen without awareness.The lack of competency awareness makes NMT untrustworthy.This is in sharp contrast to human translators who give feedback or conduct further investigations whenever they are in doubt about predictions.To fill this gap, we propose a novel competency-aware NMT by extending conventional NMT with a selfestimator, offering abilities to translate a source sentence and estimate its competency.The selfestimator encodes the information of the decoding procedure and then examines whether it can reconstruct the original semantics of the source sentence.Experimental results on four translation tasks demonstrate that the proposed method not only carries out translation tasks intact but also delivers outstanding performance on quality estimation.Without depending on any reference or annotated data typically required by state-of-the-art metric and quality estimation methods, our model yields an even higher correlation with human quality judgments than a variety of aforementioned methods, such as BLEURT, COMET, and BERTScore.Quantitative and qualitative analyses show better robustness of competency awareness in our model.1 Pei Zhang 0011, Baosong Yang, Dayiheng Liu, Kai Fan 0002, Luo Si |
EMNLP | 1 |
| 2021 | Domain-Aware Self-Attention for Multi-Domain Neural Machine Translation
Shiqi Zhang 0007, Deyi Xiong, Pei Zhang 0011, Boxing Chen |
Interspeech | 4 |
| 2021 | Context-Interactive Pre-Training for Document Machine TranslationabstractPengcheng Yang, Pei Zhang, Boxing Chen, Jun Xie, Weihua Luo. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Pei Zhang 0011, Boxing Chen, Weihua Luo |
NAACL-HLT | 2 |
| 2020 | Visual Agreement Regularized Training for Multi-Modal Machine TranslationabstractMulti-modal machine translation aims at translating the source sentence into a different language in the presence of the paired image. Previous work suggests that additional visual information only provides dispensable help to translation, which is needed in several very special cases such as translating ambiguous words. To make better use of visual information, this work presents visual agreement regularized training. The proposed approach jointly trains the source-to-target and target-to-source translation models and encourages them to share the same focus on the visual information when generating semantically equivalent visual words (e.g. “ball” in English and “ballon” in French). Besides, a simple yet effective multi-head co-attention model is also introduced to capture interactions between visual and textual features. The results show that our approaches can outperform competitive baselines by a large margin on the Multi30k dataset. Further analysis demonstrates that the proposed regularized training can effectively improve the agreement of attention on the image, leading to better use of visual information. Boxing Chen, Pei Zhang 0011, Xu Sun 0001 |
AAAI | 3 |
| 2020 | Long-Short Term Masking Transformer: A Simple but Effective Baseline for Document-level Neural Machine TranslationabstractMany document-level neural machine translation (NMT) systems have explored the utility of context-aware architecture, usually requiring an increasing number of parameters and computational complexity.However, few attention is paid to the baseline model.In this paper, we research extensively the pros and cons of the standard transformer in document-level translation, and find that the auto-regressive property can simultaneously bring both the advantage of the consistency and the disadvantage of error accumulation.Therefore, we propose a surprisingly simple long-short term masking self-attention on top of the standard transformer to both effectively capture the long-range dependence and reduce the propagation of errors.We examine our approach on the two publicly available document-level datasets.We can achieve a strong result in BLEU and capture discourse phenomena. Pei Zhang 0011, Boxing Chen, Niyu Ge, Kai Fan 0002 |
EMNLP (1) | 1 |
| 2019 | Lattice Transformer for Speech TranslationabstractRecent advances in sequence modeling have highlighted the strengths of the transformer architecture, especially in achieving state-of-theart machine translation results.However, depending on the up-stream systems, e.g., speech recognition, or word segmentation, the input to translation system can vary greatly.The goal of this work is to extend the attention mechanism of the transformer to naturally consume the lattice in addition to the traditional sequential input.We first propose a general lattice transformer for speech translation where the input is the output of the automatic speech recognition (ASR) which contains multiple paths and posterior scores.To leverage the extra information from the lattice structure, we develop a novel controllable lattice attention mechanism to obtain latent representations.On the LDC Spanish-English speech translation corpus, our experiments show that lattice transformer generalizes significantly better and outperforms both a transformer baseline and a lattice LSTM.Additionally, we validate our approach on the WMT 2017 Chinese-English translation task with lattice inputs from different BPE segmentations.In this task, we also observe the improvements over strong baselines. Pei Zhang 0011, Niyu Ge, Boxing Chen, Kai Fan 0002 |
ACL (1) | 1 |