VLDB 2026 Research / reviewers in the wild / expert
Xiangpeng Wei
dblp:220/9947
· DBLP profile ↗
23ranked-venue papers
8as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at ScaleabstractInference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the **D**ecoupled Clip and **D**ynamic s**A**mpling **P**olicy **O**ptimization (**DAPO**) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL. Qiying Yu, Zheng Zhang 0001, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, Lingjun Liu, Xin Liu 0039, Haibin Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang 0022, Mofan Zhang, Ru Zhang 0006, Wang Zhang 0017, Jiaze Chen, Jiangjie Chen, Hongli Yu, Yuxuan Song 0002, Xiangpeng Wei, Hao Zhou 0012, Wei-Ying Ma, Ya-Qin Zhang, Mingxuan Wang |
NeurIPS | 29 |
| 2024 | MoNMT: Modularly Leveraging Monolingual and Bilingual Knowledge for Neural Machine TranslationabstractThe effective use of monolingual and bilingual knowledge represents a critical challenge within the neural machine translation (NMT) community. In this paper, we propose a modular strategy that facilitates the cooperation of these two types of knowledge in translation tasks, while avoiding the issue of catastrophic forgetting and exhibiting superior model generalization and robustness. Our model is comprised of three functionally independent modules: an encoding module, a decoding module, and a transferring module. The former two acquire large-scale monolingual knowledge via self-supervised learning, while the latter is trained on parallel data and responsible for transferring latent features between the encoding and decoding modules. Extensive experiments in multi-domain translation tasks indicate our model yields remarkable performance, with up to 7 BLEU improvements in out-of-domain tests over the conventional pretrain-and-finetune approach. Our codes are available at https://github.com/NLP2CT/MoNMT. Jianhui Pang, Baosong Yang, Derek F. Wong, Dayiheng Liu, Xiangpeng Wei, Lidia S. Chao |
LREC/COLING | 5 |
| 2024 | PPCap: A Plug and Play Framework for Efficient Stylized Image Captioning
Xiangpeng Wei, Guisheng Liu, Yating Liu 0001, Yanqing Guo |
ICPR (18) | 1 |
| 2023 | Bridging the Domain Gaps in Context Representations for k-Nearest Neighbor Neural Machine TranslationabstractZhiwei Cao, Baosong Yang, Huan Lin, Suhang Wu, Xiangpeng Wei, Dayiheng Liu, Jun Xie, Min Zhang, Jinsong Su. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Baosong Yang, Suhang Wu, Xiangpeng Wei, Dayiheng Liu, Min Zhang 0005, Jinsong Su |
ACL (1) | 5 |
| 2023 | Fantastic Expressions and Where to Find Them: Chinese Simile Generation with Multiple ConstraintsabstractKexin Yang, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Xiangpeng Wei, Zhengyuan Liu, Jun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kexin Yang 0002, Dayiheng Liu, Wenqiang Lei, Baosong Yang, Xiangpeng Wei, Zhengyuan Liu |
ACL (1) | 5 |
| 2023 | MMNMT: Modularizing Multilingual Neural Machine Translation with Flexibly Assembled MoE and Dense BlocksabstractMixture-of-Experts (MoE) based sparse architectures can significantly increase model capacity with sublinear computational overhead, which are hence widely used in massively multilingual neural machine translation (MNMT).However, they are prone to overfitting on lowresource language translation.In this paper, we propose a modularized MNMT framework that is able to flexibly assemble dense and MoEbased sparse modules to achieve the best of both worlds.The training strategy of the modularized MNMT framework consists of three stages: (1) Pre-training basic MNMT models with different training objectives or model structures, (2) Initializing modules of the framework with pre-trained couterparts (e.g., encoder, decoder and embedding layers) from the basic models and (3) Fine-tuning the modularized MNMT framework to fit modules from different models together.We pre-train three basic MNMT models from scratch: a dense model, an MoE-based sparse model and a new MoE model, termed as MoE-LGR that explores multiple Language-Group-specifc Routers to incorporate language group knowledge into MNMT.The strengths of these pre-trained models are either on low-resource language translation, highresource language translation or zero-shot translation.Our modularized MNMT framework attempts to incorporate these advantages into a single model with reasonable initialization and fine-tuning.Experiments on widely-used benchmark datasets demonstrate that the proposed modularized MNMT framwork substantially outperforms both MoE and dense models on high-and low-resource language translation as well as zero-shot translation.Our framework facilitates the combination of different methods with their own strengths and recycling off-the-shelf models for multilingual neural machine translation.Codes are available at https://github.com/lishangjie1/MMNMT. Shangjie Li, Xiangpeng Wei, Shaolin Zhu, Baosong Yang, Deyi Xiong |
EMNLP | 2 |
| 2023 | EMMA-X: An EM-like Multilingual Pre-training Algorithm for Cross-lingual Representation LearningabstractExpressing universal semantics common to all languages is helpful to understand the meanings of complex and culture-specific sentences. The research theme underlying this scenario focuses on learning universal representations across languages with the usage of massive parallel corpora. However, due to the sparsity and scarcity of parallel data, there is still a big challenge in learning authentic ``universals'' for any two languages. In this paper, we propose Emma-X: an EM-like Multilingual pre-training Algorithm, to learn Cross-lingual universals with the aid of excessive multilingual non-parallel data. Emma-X unifies the cross-lingual representation learning task and an extra semantic relation prediction task within an EM framework. Both the extra semantic classifier and the cross-lingual sentence encoder approximate the semantic relation of two sentences, and supervise each other until convergence. To evaluate Emma-X, we conduct experiments on xrete, a newly introduced benchmark containing 12 widely studied cross-lingual tasks that fully depend on sentence-level representations. Results reveal that Emma-X achieves state-of-the-art performance. Further geometric analysis of the built representation space with three requirements demonstrates the superiority of Emma-X over advanced models. Ping Guo 0002, Xiangpeng Wei, Yue Hu 0002, Baosong Yang, Dayiheng Liu, Fei Huang 0002 |
NeurIPS | 2 |
| 2023 | From statistical methods to deep learning, automatic keyphrase prediction: A survey
Binbin Xie, Jia Song 0003, Liangying Shao, Suhang Wu, Xiangpeng Wei, Baosong Yang, Jinsong Su |
Inf. Process. Manag. | 5 |
| 2022 | Learning to Generalize to More: Continuous Semantic Augmentation for Neural Machine TranslationabstractThe principal task in supervised neural machine translation (NMT) is to learn to generate target sentences conditioned on the source inputs from a set of parallel sentence pairs, and thus produce a model capable of generalizing to unseen instances.However, it is commonly observed that the generalization performance of the model is highly influenced by the amount of parallel data used in training.Although data augmentation is widely used to enrich the training data, conventional methods with discrete manipulations fail to generate diverse and faithful training samples.In this paper, we present a novel data augmentation paradigm termed Continuous Semantic Augmentation (CSANMT), which augments each training instance with an adjacency semantic region that could cover adequate variants of literal expression under the same meaning.We conduct extensive experiments on both rich-resource and low-resource settings involving various language pairs, including WMT14 English→{German,French}, NIST Chinese→English and multiple low-resource IWSLT translation tasks.The provided empirical evidences show that CSANMT sets a new level of performance among existing augmentation techniques, improving on the state-of-theart by a large margin. 1 Xiangpeng Wei, Heng Yu 0006, Yue Hu 0002, Rongxiang Weng, Weihua Luo |
ACL (1) | 1 |
| 2022 | SUN: Exploring Intrinsic Uncertainties in Text-to-SQL ParsersabstractThis paper aims to improve the performance of text-to-SQL parsing by exploring the intrinsic uncertainties in the neural network based approaches (called SUN). From the data uncertainty perspective, it is indisputable that a single SQL can be learned from multiple semantically-equivalent questions. Different from previous methods that are limited to one-to-one mapping, we propose a data uncertainty constraint to explore the underlying complementary semantic information among multiple semantically-equivalent questions (many-to-one) and learn the robust feature representations with reduced spurious associations. In this way, we can reduce the sensitivity of the learned representations and improve the robustness of the parser. From the model uncertainty perspective, there is often structural information (dependence) among the weights of neural networks. To improve the generalizability and stability of neural text-to-SQL parsers, we propose a model uncertainty constraint to refine the query representations by enforcing the output representations of different perturbed encoding networks to be consistent with each other. Extensive experiments on five benchmark datasets demonstrate that our method significantly outperforms strong competitors and achieves new state-of-the-art results. Bowen Qin, Binyuan Hui, Bowen Li 0002, Xiangpeng Wei, Binhua Li, Fei Huang 0002, Luo Si, Min Yang 0007, Yongbin Li 0001 |
COLING | 5 |
| 2022 | WR-One2Set: Towards Well-Calibrated Keyphrase GenerationabstractKeyphrase generation aims to automatically generate short phrases summarizing an input document.The recently emerged ONE2SET paradigm (Ye et al., 2021) generates keyphrases as a set and has achieved competitive performance.Nevertheless, we observe serious calibration errors outputted by ONE2SET, especially in the over-estimation of ∅ token (means "no corresponding keyphrase").In this paper, we deeply analyze this limitation and identify two main reasons behind: 1) the parallel generation has to introduce excessive ∅ as padding tokens into training instances; and 2) the training mechanism assigning target to each slot is unstable and further aggravates the ∅ token over-estimation.To make the model wellcalibrated, we propose WR-ONE2SET which extends ONE2SET with an adaptive instancelevel cost Weighting strategy and a target Reassignment mechanism.The former dynamically penalizes the over-estimated slots for different instances thus smoothing the uneven training distribution.The latter refines the original inappropriate assignment and reduces the supervisory signals of over-estimated slots.Experimental results on commonly-used datasets demonstrate the effectiveness and generality of our proposed paradigm. Binbin Xie, Xiangpeng Wei, Baosong Yang, Xiaoli Wang 0002, Min Zhang 0005, Jinsong Su |
EMNLP | 2 |
| 2021 | On Learning Universal Representations Across Languages
Xiangpeng Wei, Rongxiang Weng, Yue Hu 0002, Luxi Xing, Heng Yu 0006, Weihua Luo |
ICLR | 1 |
| 2020 | Multiscale Collaborative Deep Models for Neural Machine TranslationabstractRecent evidence reveals that Neural Machine Translation (NMT) models with deeper neural networks can be more effective but are difficult to train. In this paper, we present a MultiScale Collaborative (MSC) framework to ease the training of NMT models that are substantially deeper than those used previously. We explicitly boost the gradient back-propagation from top to bottom levels by introducing a block-scale collaboration mechanism into deep NMT models. Then, instead of forcing the whole encoder stack directly learns a desired representation, we let each encoder block learns a fine-grained representation and enhance it by encoding spatial dependencies using a context-scale collaboration. We provide empirical evidence showing that the MSC nets are easy to optimize and can obtain improvements of translation quality from considerably increased depth. On IWSLT translation tasks with three translation directions, our extremely deep models (with 72-layer encoders) surpass strong baselines by +2.2~+3.1 BLEU points. In addition, our deep MSC achieves a BLEU score of 30.56 on WMT14 English-to-German task that significantly outperforms state-of-the-art deep NMT models. We have included the source code in supplementary materials. Xiangpeng Wei, Heng Yu 0006, Yue Hu 0002, Yue Zhang 0004, Rongxiang Weng, Weihua Luo |
ACL | 1 |
| 2020 | Bi-directional CognitiveThinking Network for Machine Reading ComprehensionabstractWe propose a novel Bi-directional Cognitive Knowledge Framework (BCKF) for reading comprehension from the perspective of complementary learning systems theory. It aims to simulate two ways of thinking in the brain to answer questions, including reverse thinking and inertial thinking. To validate the effectiveness of our framework, we design a corresponding Bi-directional Cognitive Thinking Network (BCTN) to encode the passage and generate a question (answer) given an answer (question) and decouple the bi-directional knowledge. The model has the ability to reverse reasoning questions which can assist inertial thinking to generate more accurate answers. Competitive improvement is observed in DuReader dataset, confirming our hypothesis that bi-directional knowledge helps the QA task. The novel framework shows an interesting perspective on machine reading comprehension and cognitive science. Wei Peng 0008, Yue Hu 0002, Luxi Xing, Yuqiang Xie, Jing Yu 0007, Yajing Sun, Xiangpeng Wei |
COLING | 7 |
| 2020 | Uncertainty-Aware Semantic Augmentation for Neural Machine TranslationabstractAs a sequence-to-sequence generation task, neural machine translation (NMT) naturally contains intrinsic uncertainty, where a single sentence in one language has multiple valid counterparts in the other.However, the dominant methods for NMT only observe one of them from the parallel corpora for the model training but have to deal with adequate variations under the same meaning at inference.This leads to a discrepancy of the data distribution between the training and the inference phases.To address this problem, we propose uncertainty-aware semantic augmentation, which explicitly captures the universal semantic information among multiple semantically-equivalent source sentences and enhances the hidden representations with this information for better translations.Extensive experiments on various translation tasks reveal that our approach significantly outperforms the strong baselines and the existing methods. Xiangpeng Wei, Heng Yu 0006, Yue Hu 0002, Rongxiang Weng, Luxi Xing, Weihua Luo |
EMNLP (1) | 1 |
| 2020 | Towards Enhancing Faithfulness for Neural Machine TranslationabstractNeural machine translation (NMT) has achieved great success due to the ability to generate high-quality sentences. Compared with human translations, one of the drawbacks of current NMT is that translations are not usually faithful to the input, e.g., omitting information or generating unrelated fragments, which inevitably decreases the overall quality, especially for human readers. In this paper, we propose a novel training strategy with a multi-task learning paradigm to build a faithfulness enhanced NMT model (named FEnmt). During the NMT training process, we sample a subset from the training set and translate them to get fragments that have been mistranslated. Afterward, the proposed multi-task learning paradigm is employed on both encoder and decoder to guide NMT to correctly translate these fragments. Both automatic and human evaluations verify that our FEnmt could improve translation quality by effectively reducing unfaithful translations. Rongxiang Weng, Heng Yu 0006, Xiangpeng Wei, Weihua Luo |
EMNLP (1) | 3 |
| 2020 | Dynamic Attention Aggregation with BERT for Neural Machine TranslationabstractThe recently proposed BERT has demonstrated great power in various natural language processing tasks. However, the model does not perform effectively on cross-lingual tasks, especially on machine translation. In this work, we propose three methods to introduce pre-trained BERT into neural machine translation without fine-tuning. Our approach consists of a) a linear-attention aggregation that leverages a parameter matrix to capture the key knowledge of BERT, b) a self-attention aggregation which aims to learn what is vital for input and output, and c) a switch-gate aggregation to dynamically control the balance of the information flowing from the pre-trained BERT or the NMT model. We conduct experiments on several translation benchmarks and substantially improve over 2 BELU points on the IWSLT'14 English - German task with switch-gate aggregation method compared to a strong baseline, while our proposed model also performs remarkably on the other tasks. Jiarui Zhang 0003, Hongzheng Li, Shumin Shi, Heyan Huang, Yue Hu 0002, Xiangpeng Wei |
IJCNN | 6 |
| 2020 | Enhancing Pre-trained Language Models by Self-supervised Learning for Story Cloze Test
Yuqiang Xie, Yue Hu 0002, Luxi Xing, Xiangpeng Wei, Yajing Sun |
KSEM (1) | 6 |
| 2019 | Translating with Bilingual Topic Knowledge for Neural Machine TranslationabstractThe dominant neural machine translation (NMT) models that based on the encoder-decoder architecture have recently achieved the state-of-the-art performance. Traditionally, the NMT models only depend on the representations learned during training for mapping a source sentence into the target domain. However, the learned representations often suffer from implicit and inadequately informed properties. In this paper, we propose a novel bilingual topic enhanced NMT (BLTNMT) model to improve translation performance by incorporating bilingual topic knowledge into NMT. Specifically, the bilingual topic knowledge is included into the hidden states of both encoder and decoder, as well as the attention mechanism. With this new setting, the proposed BLT-NMT has access to the background knowledge implied in bilingual topics which is beyond the sequential context, and enables the attention mechanism to attend to topic-level attentions for generating accurate target words during translation. Experimental results show that the proposed model consistently outperforms the traditional RNNsearch and the previous topic-informed NMT on Chinese-English and EnglishGerman translation tasks. We also introduce the bilingual topic knowledge into the newly emerged Transformer base model on English-German translation and achieve a notable improvement. Xiangpeng Wei, Yue Hu 0002, Luxi Xing, Yipeng Wang 0001 |
AAAI | 1 |
| 2019 | Unsupervised Neural Machine Translation with Future RewardingabstractIn this paper, we alleviate the local optimality of back-translation by learning a policy (takes the form of an encoder-decoder and is defined by its parameters) with future rewarding under the reinforcement learning framework, which aims to optimize the global word predictions for unsupervised neural machine translation.To this end, we design a novel reward function to characterize high-quality translations from two aspects: n-gram matching and semantic adequacy.The n-gram matching is defined as an alternative for the discrete BLEU metric, and the semantic adequacy is used to measure the adequacy of conveying the meaning of the source sentence to the target.During training, our model strives for earning higher rewards by learning to produce grammatically more accurate and semantically more adequate translations.Besides, a variational inference network (VIN) is proposed to constrain the corresponding sentences in two languages have the same or similar latent semantic code.On the widely used WMT'14 English-French, WMT'16 English-German and NIST Chineseto-English benchmarks, our models respectively obtain 27.59/27.15,19.65/23.42 and 22.40 BLEU points without using any labeled data, demonstrating consistent improvements over previous unsupervised NMT models. Xiangpeng Wei, Yue Hu 0002, Luxi Xing |
CoNLL | 1 |
| 2019 | Syntax-Aware Sentence Matching with Graph Convolutional Networks
Yangfan Lei, Yue Hu 0002, Xiangpeng Wei, Luxi Xing, Quanchao Liu |
KSEM (2) | 3 |
| 2019 | Gated Self-attentive Encoder for Neural Machine Translation
Xiangpeng Wei, Yue Hu 0002, Luxi Xing |
KSEM (1) | 1 |
| 2019 | Dynamic Task-Specific Factors for Meta-Embedding
Yuqiang Xie, Yue Hu 0002, Luxi Xing, Xiangpeng Wei |
KSEM (2) | 4 |