VLDB 2026 Research / reviewers in the wild / expert
Tongran Liu
dblp:125/6642
· DBLP profile ↗
14ranked-venue papers
0as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Reinforcement learning · 22% Language models and text generation · 19% Efficient and distributed learning · 14% |
Topics — the 30 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
2.8 | 4 | 2026 | GRAM: A Generative Foundation Reward Model for Reward Generalization · ICML 2025 RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data · AAAI 2025 ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence Generation · AAAI 2024 |
Natural language and speech › Machine translation
neural machine translation |
1.2 | 3 | 2020 | Does Multi-Encoder Help? A Case Study on Context-Aware Neural Machine Translation · ACL 2020 Neural Machine Translation with Joint Representation · AAAI 2020 Sharing Attention Weights for Fast Transformer · IJCAI 2019 |
Natural language and speech › Language models and text generation
alignment |
1.0 | 1 | 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026 |
Machine learning › Trustworthy machine learning
interpretability |
1.0 | 1 | 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026 |
Machine learning › Reinforcement learning › reward learning › reward modeling
reward model evaluation |
1.0 | 1 | 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026 |
Machine learning › Trustworthy machine learning › interpretability › explainable reinforcement learning
reward model interpretability |
1.0 | 1 | 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models · AAAI 2026 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
0.9 | 1 | 2025 | RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data · AAAI 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | GRAM: A Generative Foundation Reward Model for Reward Generalization · ICML 2025 |
Computer vision › Vision and language › vision-language model
vision-language model alignment |
0.9 | 1 | 2025 | RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data · AAAI 2025 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.8 | 2 | 2020 | Neural Machine Translation with Joint Representation · AAAI 2020 Sharing Attention Weights for Fast Transformer · IJCAI 2019 |
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
cross-lingual representation learning |
0.8 | 1 | 2024 | Revealing the Parallel Multilingual Learning within Large Language Models · EMNLP 2024 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.8 | 1 | 2024 | Revealing the Parallel Multilingual Learning within Large Language Models · EMNLP 2024 |
Natural language and speech › Language models and text generation
multilingual language models |
0.8 | 1 | 2024 | Revealing the Parallel Multilingual Learning within Large Language Models · EMNLP 2024 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation |
0.8 | 1 | 2024 | ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence Generation · AAAI 2024 |
Machine learning › Efficient and distributed learning › inference efficiency
efficient transformer inference |
0.5 | 2 | 2020 | Sharing Attention Weights for Fast Transformer · IJCAI 2019 Towards Fully 8-bit Integer Inference for the Transformer Model · IJCAI 2020 |
Machine learning › Deep learning architectures and training
architecture learning |
0.4 | 1 | 2020 | Learning Architectures from an Extended Search Space for Language Modeling · ACL 2020 |
Natural language and speech › Machine translation › document-level machine translation
document-level neural machine translation |
0.4 | 1 | 2020 | Does Multi-Encoder Help? A Case Study on Context-Aware Neural Machine Translation · ACL 2020 |
Natural language and speech › Language models and text generation
language modeling |
0.4 | 1 | 2020 | Learning Architectures from an Extended Search Space for Language Modeling · ACL 2020 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2020 | Towards Fully 8-bit Integer Inference for the Transformer Model · IJCAI 2020 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.4 | 1 | 2020 | Learning Architectures from an Extended Search Space for Language Modeling · ACL 2020 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.4 | 1 | 2020 | Towards Fully 8-bit Integer Inference for the Transformer Model · IJCAI 2020 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
search space design |
0.4 | 1 | 2020 | Learning Architectures from an Extended Search Space for Language Modeling · ACL 2020 |
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model |
0.4 | 1 | 2020 | Neural Machine Translation with Joint Representation · AAAI 2020 |
Machine learning › Efficient and distributed learning › distributed training › asynchronous training
asynchronous stochastic gradient descent |
0.3 | 1 | 2017 | Fast Parallel Training of Neural Language Models · IJCAI 2017 |
Machine learning › Efficient and distributed learning
distributed training |
0.3 | 1 | 2017 | Fast Parallel Training of Neural Language Models · IJCAI 2017 |
Natural language and speech › Language models and text generation › neural language model
neural language model training |
0.3 | 1 | 2017 | Fast Parallel Training of Neural Language Models · IJCAI 2017 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.3 | 1 | 2025 | GRAM: A Generative Foundation Reward Model for Reward Generalization · ICML 2025 |
Natural language and speech › Machine translation
syntax-based machine translation |
0.2 | 1 | 2016 | Syntactic Skeleton-Based Translation · AAAI 2016 |
Natural language and speech › Machine translation
statistical machine translation |
0.2 | 2 | 2016 | Bagging and Boosting statistical machine translation systems · Artif. Intell. 2013 Syntactic Skeleton-Based Translation · AAAI 2016 |
Natural language and speech › Language models and text generation › text summarization
abstractive summarization |
0.2 | 1 | 2024 | ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence Generation · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
probing · 1.8multi-objective optimization · 1.0pairwise ranking loss · 0.9optimal transport · 0.9label smoothing · 0.9generative model · 0.9direct preference optimization · 0.9best-of-n sampling · 0.9proximal policy optimization · 0.8dynamic sampling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward ModelsabstractPrevious methods evaluate reward models by testing them on a fixed pairwise ranking test set, but they typically do not provide performance information on each preference dimension. In this work, we address the evaluation challenge of reward models by probing preference representations. To confirm the effectiveness of this evaluation method, we construct a Multi-dimensional Reward Model Benchmark (MRMBench), a collection of six probing tasks for different preference dimensions. We design it to favor and encourage reward models that better capture preferences across different dimensions. Furthermore, we introduce an analysis method, inference-time probing, which identifies the dimensions used during the reward prediction and enhances its interpretability. Through extensive experiments, we find that MRMBench strongly correlates with LLM alignment performance, supporting it as a reliable reference for developing advanced reward models. By analyzing the evaluation results on MRMBench, we reveal that reward models struggle to simultaneously capture preferences across multiple dimensions, highlighting the potential of multi-objective optimization in reward modeling. Furthermore, our results demonstrate that the proposed inference-time probing method provides a reliable metric for assessing the confidence of reward predictions, leading to improved alignment of large language models. Chenglong Wang 0002, Yifu Huo, Yang Gan, Yongyu Mu, Qiaozhi He, Murun Yang, Chunliang Zhang, Tongran Liu, Anxiang Ma, Zhengtao Yu 0001, Tong Xiao 0001 |
AAAI | 9 |
| 2025 | RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference DataabstractLarge vision-language models (LVLMs) often fail to align with human preferences, leading to issues like generating misleading content without proper visual context (also known as hallucination). A promising solution to this problem is using human-preference alignment techniques, such as best-of-n sampling and reinforcement learning. However, these techniques face the difficulty arising from the scarcity of visual preference data, which is required to train a visual reward model (VRM). In this work, we continue the line of research. We present a Robust Visual Reward Model (RoVRM) which improves human-preference alignment for LVLMs. RoVRM leverages auxiliary textual preference data through a three-phase progressive training and optimal transport-based preference data selection to effectively mitigate the scarcity of visual preference data. We experiment with RoVRM on the commonly used vision-language tasks based on the LLaVA-1.5-7B and -13B models. Experimental results demonstrate that RoVRM consistently outperforms traditional VRMs. Furthermore, our three-phase progressive training and preference data selection approaches can yield consistent performance gains over ranking-based alignment techniques, such as direct preference optimization. Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Murun Yang, Qiaozhi He, Tong Xiao 0001, Chunliang Zhang, Tongran Liu |
AAAI | 9 |
| 2025 | GRAM: A Generative Foundation Reward Model for Reward GeneralizationabstractIn aligning large language models (LLMs), reward models have played an important role, but are standardly trained as discriminative models and rely only on labeled human preference data. In this paper, we explore methods that train reward models using both unlabeled and labeled data. Building on the generative models in LLMs, we develop a generative reward model that is first trained via large-scale unsupervised learning and then fine-tuned via supervised learning. We also show that by using label smoothing, we are in fact optimizing a regularized pairwise ranking loss. This result, in turn, provides a new view of training reward models, which links generative models and discriminative models under the same class of training objectives. The outcome of these techniques is a foundation reward model, which can be applied to a wide range of tasks with little or no further fine-tuning effort. Extensive experiments show that this model generalizes well across several tasks, including response ranking, reinforcement learning from human feedback, and task adaptation with fine-tuning, achieving significant performance improvements over several strong baseline models. Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Qiaozhi He, Murun Yang, Tong Xiao 0001, Chunliang Zhang, Tongran Liu |
ICML | 10 |
| 2025 | MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward OptimizationabstractRecent advances in diffusion language models (DLMs) have presented a promising alternative to traditional autoregressive large language models (LLMs). However, DLMs still lag behind LLMs in reasoning performance, especially as the number of denoising steps decreases. Our analysis reveals that this shortcoming arises primarily from the independent generation of masked tokens across denoising steps, which fails to capture the token correlation. In this paper, we define two types of token correlation: intra-sequence correlation and inter-sequence correlation, and demonstrate that enhancing these correlations improves reasoning performance. To this end, we propose a Multi-Reward Optimization (MRO) approach, which encourages DLMs to consider the token correlation during the denoising process. More specifically, our MRO approach leverages test-time scaling, reject sampling, and reinforcement learning to directly optimize the token correlation with multiple elaborate rewards. Additionally, we introduce group step and importance sampling strategies to mitigate reward variance and enhance sampling efficiency. Through extensive experiments, we demonstrate that MRO not only improves reasoning performance but also achieves significant sampling speedups while maintaining high performance on reasoning benchmarks. Chenglong Wang 0002, Yang Gan, Chi Hu, Yongyu Mu, Murun Yang, Chunliang Zhang, Tongran Liu, Zhengtao Yu 0001, Tong Xiao 0001 |
NeurIPS | 10 |
| 2024 | ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence GenerationabstractApplying Reinforcement Learning (RL) to sequence generation models enables the direct optimization of long-term rewards (e.g., BLEU and human feedback), but typically requires large-scale sampling over a space of action sequences. This is a computational challenge as presented by the practice of sequence generation problems, such as machine translation, where we often deal with a large action space (e.g., a vocabulary) and a long action sequence (e.g., a translation). In this work, we introduce two-stage sampling and dynamic sampling approaches to improve the sampling efficiency during training sequence generation models via RL. We experiment with our approaches on the traditional sequence generation tasks, including machine translation and abstractive summarization. Furthermore, we evaluate our approaches in RL from human feedback (RLHF) through training a large language model using the reward model. Experimental results show that the efficient sampling-based RL, referred to as ESRL, can outperform all baselines in terms of both training efficiency and memory consumption. Notably, ESRL yields consistent performance gains over the strong REINFORCE, minimum risk training, and proximal policy optimization methods. The code is available at https://github.com/wangclnlp/DeepSpeed-Chat-Extension/examples/esrl. Chenglong Wang 0002, Yimin Hu, Yifu Huo, Tongran Liu, Tong Xiao 0001 |
AAAI | 6 |
| 2024 | Revealing the Parallel Multilingual Learning within Large Language ModelsabstractYongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu, Bei Li, Chenglong Wang, Tong Xiao, Kai Song, Tongran Liu, Chunliang Zhang, JingBo Zhu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu, Chenglong Wang 0002, Tong Xiao 0001, Tongran Liu, Chunliang Zhang |
EMNLP | 9 |
| 2020 | Neural Machine Translation with Joint RepresentationabstractThough early successes of Statistical Machine Translation (SMT) systems are attributed in part to the explicit modelling of the interaction between any two source and target units, e.g., alignment, the recent Neural Machine Translation (NMT) systems resort to the attention which partially encodes the interaction for efficiency. In this paper, we employ Joint Representation that fully accounts for each possible interaction. We sidestep the inefficiency issue by refining representations with the proposed efficient attention operation. The resulting Reformer models offer a new Sequence-to-Sequence modelling paradigm besides the Encoder-Decoder framework and outperform the Transformer baseline in either the small scale IWSLT14 German-English, English-German and IWSLT15 Vietnamese-English or the large scale NIST12 Chinese-English translation tasks by about 1 BLEU point. We also propose a systematic model scaling approach, allowing the Reformer model to beat the state-of-the-art Transformer in IWSLT14 German-English and NIST12 Chinese-English with about 50% fewer parameters. The code is publicly available at https://github.com/lyy1994/reformer. Yanyang Li, Qiang Wang 0050, Tong Xiao 0001, Tongran Liu |
AAAI | 4 |
| 2020 | Learning Architectures from an Extended Search Space for Language ModelingabstractNeural architecture search (NAS) has advanced significantly in recent years but most NAS systems restrict search to learning architectures of a recurrent or convolutional cell.In this paper, we extend the search space of NAS.In particular, we present a general approach to learn both intra-cell and inter-cell architectures (call it ESS).For a better search result, we design a joint learning method to perform intra-cell and inter-cell NAS simultaneously.We implement our model in a differentiable architecture search system.For recurrent neural language modeling, it outperforms a strong baseline significantly on the PTB and Wiki-Text data, with a new state-of-the-art on PTB.Moreover, the learned architectures show good transferability to other systems.E.g., they improve state-of-the-art systems on the CoNLL and WNUT named entity recognition (NER) tasks and CoNLL chunking task, indicating a promising line of research on large-scale prelearned architectures. Yinqiao Li, Chi Hu, Nuo Xu 0010, Yufan Jiang, Tong Xiao 0001, Tongran Liu, Changliang Li |
ACL | 8 |
| 2020 | Does Multi-Encoder Help? A Case Study on Context-Aware Neural Machine TranslationabstractIn encoder-decoder neural models, multiple encoders are in general used to represent the contextual information in addition to the individual sentence.In this paper, we investigate multi-encoder approaches in document-level neural machine translation (NMT).Surprisingly, we find that the context encoder does not only encode the surrounding sentences but also behaves as a noise generator.This makes us rethink the real benefits of multi-encoder in context-aware translation -some of the improvements come from robust training.We compare several methods that introduce noise and/or well-tuned dropout setup into the training of these encoders.Experimental results show that noisy training plays an important role in multi-encoder-based NMT, especially when the training data is small.Also, we establish a new state-of-the-art on IWSLT Fr-En task by careful use of noise generation and dropout methods. Yufan Jiang, Tong Xiao 0001, Tongran Liu, Changliang Li |
ACL | 7 |
| 2020 | Towards Fully 8-bit Integer Inference for the Transformer Modelabstract8-bit integer inference, as a promising direction in reducing both the latency and storage of deep neural networks, has made great progress recently. On the other hand, previous systems still rely on 32-bit floating point for certain functions in complex models (e.g., Softmax in Transformer), and make heavy use of quantization and de-quantization. In this work, we show that after a principled modification on the Transformer architecture, dubbed Integer Transformer, an (almost) fully 8-bit integer inference algorithm Scale Propagation could be derived. De-quantization is adopted when necessary, which makes the network more efficient. Our experiments on WMT16 En<->Ro, WMT14 En<->De and En->Fr translation tasks as well as the WikiText-103 language modelling task show that the fully 8-bit Transformer system achieves comparable performance with the floating point baseline but requires nearly 4x less memory footprint. Yanyang Li, Tengbo Liu, Tong Xiao 0001, Tongran Liu |
IJCAI | 5 |
| 2019 | Sharing Attention Weights for Fast TransformerabstractRecently, the Transformer machine translation system has shown strong results by stacking attention layers on both the source and target-language sides. But the inference of this model is slow due to the heavy use of dot-product attention in auto-regressive decoding. In this paper we speed up Transformer via a fast and lightweight attention model. More specifically, we share attention weights in adjacent layers and enable the efficient re-use of hidden states in a vertical manner. Moreover, the sharing policy can be jointly learned with the MT model. We test our approach on ten WMT and NIST OpenMT tasks. Experimental results show that it yields an average of 1.3X speed-up (with almost no decrease in BLEU) on top of a state-of-the-art implementation that has already adopted a cache for fast inference. Also, our approach obtains a 1.8X speed-up when it works with the AAN model. This is even 16 times faster than the baseline with no use of the attention cache. Tong Xiao 0001, Yinqiao Li, Zhengtao Yu 0001, Tongran Liu |
IJCAI | 5 |
| 2017 | Fast Parallel Training of Neural Language ModelsabstractTraining neural language models (NLMs) is very time consuming and we need parallelization for system speedup. However, standard training methods have poor scalability across multiple devices (e.g., GPUs) due to the huge time cost required to transmit data for gradient sharing in the back-propagation process. In this paper we present a sampling-based approach to reducing data transmission for better scaling of NLMs. As a ''bonus'', the resulting model also improves the training speed on a single device. Our approach yields significant speed improvements on a recurrent neural network-based language model. On four NVIDIA GTX1080 GPUs, it achieves a speedup of 2.1+ times over the standard asynchronous stochastic gradient descent baseline, yet with no increase in perplexity. This is even 4.2 times faster than the naive single GPU counterpart. Tong Xiao 0001, Tongran Liu, Chunliang Zhang |
IJCAI | 3 |
| 2016 | Syntactic Skeleton-Based TranslationabstractIn this paper we propose an approach to modeling syntactically-motivated skeletal structure of source sentence for machine translation. This model allows for application of high-level syntactic transfer rules and low-level non-syntactic rules. It thus involves fully syntactic, non-syntactic, and partially syntactic derivations via a single grammar and decoding paradigm. On large-scale Chinese-English and English-Chinese translation tasks, we obtain an average improvement of +0.9 BLEU across the newswire and web genres. Tong Xiao 0001, Chunliang Zhang, Tongran Liu |
AAAI | 4 |
| 2013 | Bagging and Boosting statistical machine translation systems
Tong Xiao 0001, Tongran Liu |
Artif. Intell. | 3 |