VLDB 2026 Research / reviewers in the wild / expert
Zhixing Tan
dblp:186/8414
· DBLP profile ↗
16ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-2426-6220ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XRAG: Examining the Core - Benchmarking Foundational Components in Advanced Retrieval-Augmented GenerationabstractRetrieval-augmented generation (RAG) synergizes the retrieval of pertinent data with the generative capabilities of Large Language Models (LLMs), ensuring that the generated output is not only contextually relevant but also accurate and current. We introduce XRAG, an open-source, modular codebase that facilitates exhaustive evaluation of the performance of foundational components of advanced RAG modules. These components are systematically categorized into four core phases: pre-retrieval, retrieval, post-retrieval, and generation. We systematically analyse them across reconfigured datasets, providing a comprehensive benchmark for their effectiveness. As the complexity of RAG systems continues to escalate, we underscore the critical need to identify potential failure points in RAG systems. We formulate a suite of experimental methodologies and diagnostic testing protocols to dissect the failure points inherent in RAG engineering. Subsequently, we proffer bespoke solutions aimed at bolstering the overall performance of these modules. Our work thoroughly evaluates the performance of advanced core components in RAG systems, providing insights into optimizations for prevalent failure points. Qili Zhang, Qianren Mao, Yangyifei Luo, Yashuo Luo, Hanwen Hao, Zhilong Cao, Weifeng Jiang, Jinlong Zhang, Zhenting Huang, Zhixing Tan, Jie Sun 0035, Philip S. Yu |
ICDE | 14 |
| 2025 | LLM×MapReduce: Simplified Long-Sequence Processing using Large Language ModelsabstractZihan Zhou, Chong Li, Xinyi Chen, Shuo Wang, Yu Chao, Zhili Li, Haoyu Wang, Qi Shi, Zhixing Tan, Xu Han, Xiaodong Shi, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shuo Wang 0013, Yu Chao, Zhili Li, Qi Shi 0002, Zhixing Tan, Xu Han 0007, Xiaodong Shi, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 9 |
| 2024 | Black-Box Prompt Tuning With Subspace LearningabstractBlack-box prompt tuning employs derivative-free optimization algorithms to learn prompts within low-dimensional subspaces rather than back-propagating through the network of Large Language Models (LLMs). Recent studies reveal that black-box prompt tuning lacks versatility across tasks and LLMs, which we believe is related to the suboptimal choice of subspaces. In this paper, we introduceBlack-box prompt tuning withSubspaceLearning (BSL) to enhance the versatility of black-box prompt tuning. Based on the assumption that nearly optimal prompts for similar tasks reside in a common subspace, we propose identifying such subspaces through meta-learning on a collection of similar source tasks. Consequently, for a target task that shares similarities with the source tasks, we expect that optimizing within the identified subspace can yield a prompt that performs well on the target task. Experimental results confirm that our BSL framework consistently achieves competitive performance across various downstream tasks and LLMs. Yuanhang Zheng, Zhixing Tan, Peng Li 0030, Yang Liu 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | COSYWA: Enhancing Semantic Integrity in Watermarking Natural Language Generation
Junjie Fang, Zhixing Tan, Xiaodong Shi |
NLPCC (1) | 2 |
| 2022 | MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better TranslatorsabstractPrompting has recently been shown as a promising approach for applying pre-trained language models to perform downstream tasks.We present Multi-Stage Prompting, a simple and automatic approach for leveraging pre-trained language models to translation tasks.To better mitigate the discrepancy between pre-training and translation, MSP divides the translation process via pre-trained language models into multiple separate stages: the encoding stage, the re-encoding stage, and the decoding stage.During each stage, we independently apply different continuous prompts for allowing pretrained language models better shift to translation tasks.We conduct extensive experiments on three translation tasks.Experiments show that our method can significantly improve the translation performance of pre-trained language models. Zhixing Tan, Xiangwen Zhang, Shuo Wang 0013, Yang Liu 0005 |
ACL (1) | 1 |
| 2022 | Integrating Vectorized Lexical Constraints for Neural Machine TranslationabstractLexically constrained neural machine translation (NMT), which controls the generation of NMT models with pre-specified constraints, is important in many practical scenarios.Due to the representation gap between discrete constraints and continuous vectors in NMT models, most existing works choose to construct synthetic data or modify the decoding algorithm to impose lexical constraints, treating the NMT model as a black box.In this work, we propose to open this black box by directly integrating the constraints into NMT models.Specifically, we vectorize source and target constraints into continuous keys and values, which can be utilized by the attention modules of NMT models.The proposed integration method is based on the assumption that the correspondence between keys and values in attention modules is naturally suitable for modeling constraint pairs.Experimental results show that our method consistently outperforms several representative baselines on four language pairs, demonstrating the superiority of integrating vectorized lexical constraints. Shuo Wang 0013, Zhixing Tan, Yang Liu 0005 |
ACL (1) | 2 |
| 2022 | A Template-based Method for Constrained Neural Machine TranslationabstractMachine translation systems are expected to cope with various types of constraints in many practical scenarios.While neural machine translation (NMT) has achieved strong performance in unconstrained cases, it is non-trivial to impose pre-specified constraints into the translation process of NMT models.Although many approaches have been proposed to address this issue, most existing methods can not satisfy the following three desiderata at the same time: (1) high translation quality, (2) high match accuracy, and (3) low latency.In this work, we propose a template-based method that can yield results with high translation quality and match accuracy and the inference speed of our method is comparable with unconstrained NMT models.Our basic idea is to rearrange the generation of constrained and unconstrained tokens through a template.Our method does not require any changes in the model architecture and the decoding algorithm.Experimental results show that the proposed template-based approach can outperform several representative baselines in both lexically and structurally constrained translation tasks. Shuo Wang 0013, Peng Li 0030, Zhixing Tan, Zhaopeng Tu, Maosong Sun 0001, Yang Liu 0005 |
EMNLP | 3 |
| 2022 | Molecule Generation by Principal Subgraph Mining and AssemblingabstractMolecule generation is central to a variety of applications. Current attention has been paid to approaching the generation task as subgraph prediction and assembling. Nevertheless, these methods usually rely on hand-crafted or external subgraph construction, and the subgraph assembling depends solely on local arrangement. In this paper, we define a novel notion, principal subgraph that is closely related to the informative pattern within molecules. Interestingly, our proposed merge-and-update subgraph extraction method can automatically discover frequent principal subgraphs from the dataset, while previous methods are incapable of. Moreover, we develop a two-step subgraph assembling strategy, which first predicts a set of subgraphs in a sequence-wise manner and then assembles all generated subgraphs globally as the final output molecule. Built upon graph variational auto-encoder, our model is demonstrated to be effective in terms of several evaluation metrics and efficiency, compared with state-of-the-art methods on distribution learning and (constrained) property optimization tasks. Xiangzhe Kong, Wenbing Huang 0001, Zhixing Tan, Yang Liu 0005 |
NeurIPS | 3 |
| 2022 | Dynamic Multi-Branch Layers for On-Device Neural Machine TranslationabstractWith the rapid development of artificial intelligence (AI), there is a trend in moving AI applications, such as neural machine translation (NMT), from cloud to mobile devices. Constrained by limited hardware resources and battery, the performance of on-device NMT systems is far from satisfactory. Inspired by conditional computation, we propose to improve the performance of on-device NMT systems with dynamic multi-branch layers. Specifically, we design a layer-wise dynamic multi-branch network with only one branch activated during training and inference. As not all branches are activated during training, we propose shared-private reparameterization to ensure sufficient training for each branch. At almost the same computational cost, our method achieves improvements of up to 1.7 BLEU points on the WMT14 English-German translation task and 1.8 BLEU points on the WMT20 Chinese-English translation task over the Transformer model, respectively. Compared with a strong baseline that also uses multiple branches, the proposed method is up to 1.5 times faster with the same number of parameters. Zhixing Tan, Zeyuan Yang 0002, Meng Zhang 0019, Qun Liu 0001, Maosong Sun 0001, Yang Liu 0005 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Self-Supervised Quality Estimation for Machine TranslationabstractYuanhang Zheng, Zhixing Tan, Meng Zhang, Mieradilijiang Maimaiti, Huanbo Luan, Maosong Sun, Qun Liu, Yang Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Yuanhang Zheng, Zhixing Tan, Meng Zhang 0019, Mieradilijiang Maimaiti, Huan-Bo Luan, Maosong Sun 0001, Qun Liu 0001, Yang Liu 0005 |
EMNLP (1) | 2 |
| 2020 | Modeling Voting for System Combination in Machine TranslationabstractSystem combination is an important technique for combining the hypotheses of different machine translation systems to improve translation performance. Although early statistical approaches to system combination have been proven effective in analyzing the consensus between hypotheses, they suffer from the error propagation problem due to the use of pipelines. While this problem has been alleviated by end-to-end training of multi-source sequence-to-sequence models recently, these neural models do not explicitly analyze the relations between hypotheses and fail to capture their agreement because the attention to a word in a hypothesis is calculated independently, ignoring the fact that the word might occur in multiple hypotheses. In this work, we propose an approach to modeling voting for system combination in machine translation. The basic idea is to enable words in hypotheses from different systems to vote on words that are representative and should get involved in the generation process. This can be done by quantifying the influence of each voter and its preference for each candidate. Our approach combines the advantages of statistical and neural methods since it can not only analyze the relations between hypotheses but also allow for end-to-end training. Experiments show that our approach is capable of better taking advantage of the consensus between hypotheses and achieves significant improvements over state-of-the-art baselines on Chinese-English and English-German machine translation tasks. Xuancheng Huang, Zhixing Tan, Derek F. Wong, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001, Yang Liu 0005 |
IJCAI | 3 |
| 2019 | Towards Linear Time Neural Machine Translation with Capsule NetworksabstractMingxuan Wang, Jun Xie, Zhixing Tan, Jinsong Su, Deyi Xiong, Lei Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingxuan Wang, Zhixing Tan, Jinsong Su, Deyi Xiong, Lei Li 0005 |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Deep Semantic Role Labeling With Self-AttentionabstractSemantic Role Labeling (SRL) is believed to be a crucial step towards natural language understanding and has been widely studied. Recent years, end-to-end SRL with recurrent neural networks (RNN) has gained increasing attention. However, it remains a major challenge for RNNs to handle structural information and long range dependencies. In this paper, we present a simple and effective architecture for SRL which aims to address these problems. Our model is based on self-attention which can directly capture the relationships between two tokens regardless of their distance. Our single model achieves F1=83.4 on the CoNLL-2005 shared task dataset and F1=82.7 on the CoNLL-2012 shared task dataset, which outperforms the previous state-of-the-art results by 1.8 and 1.0 F1 score respectively. Besides, our model is computationally efficient, and the parsing speed is 50K tokens per second on a single Titan X GPU. Zhixing Tan, Mingxuan Wang, Yidong Chen 0001, Xiaodong Shi |
AAAI | 1 |
| 2018 | Neural Machine Translation with Decoding History Enhanced AttentionabstractNeural machine translation with source-side attention have achieved remarkable performance. however, there has been little work exploring to attend to the target-side which can potentially enhance the memory capbility of NMT. We reformulate a Decoding History Enhanced Attention mechanism (DHEA) to render NMT model better at selecting both source-side and target-side information. DHA enables dynamic control of the ratios at which source and target contexts contribute to the generation of target words, offering a way to weakly induce structure relations among both source and target tokens. It also allows training errors to be directly back-propagated through short-cut connections and effectively alleviates the gradient vanishing problem. The empirical study on Chinese-English translation shows that our model with proper configuration can improve by 0:9 BLEU upon Transformer and the best reported results in the dataset. On WMT14 English-German task and a larger WMT14 English-French task, our model achieves comparable results with the state-of-the-art. Mingxuan Wang, Zhixing Tan, Jinsong Su, Deyi Xiong, Chao Bian 0005 |
COLING | 3 |
| 2018 | Lattice-to-sequence attentional Neural Machine Translation models
Zhixing Tan, Jinsong Su, Boli Wang, Yidong Chen 0001, Xiaodong Shi |
Neurocomputing | 1 |
| 2017 | Lattice-Based Recurrent Neural Network Encoders for Neural Machine TranslationabstractNeural machine translation (NMT) heavily relies on word-level modelling to learn semantic representations of input sentences.However, for languages without natural word delimiters (e.g., Chinese) where input sentences have to be tokenized first,conventional NMT is confronted with two issues:1) it is difficult to find an optimal tokenization granularity for source sentence modelling, and2) errors in 1-best tokenizations may propagate to the encoder of NMT.To handle these issues, we propose word-lattice based Recurrent Neural Network (RNN) encoders for NMT,which generalize the standard RNN to word lattice topology.The proposed encoders take as input a word lattice that compactly encodes multiple tokenizations, and learn to generate new hidden states from arbitrarily many inputs and hidden states in preceding time steps.As such, the word-lattice based encoders not only alleviate the negative impact of tokenization errors but also are more expressive and flexible to embed input sentences.Experiment results on Chinese-English translation demonstrate the superiorities of the proposed encoders over the conventional encoder. Jinsong Su, Zhixing Tan, Deyi Xiong, Rongrong Ji, Xiaodong Shi, Yang Liu 0005 |
AAAI | 2 |