VLDB 2026 Research / reviewers in the wild / expert
Chenglong Wang 0002
dblp:94/9817-2
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2026
0009-0003-4456-6110ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SageLM: A Multi-aspect and Explainable Large Language Model for Speech JudgementabstractSpeech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. We propose SageLM, an end-to-end, multi-aspect, and explainable speech LLM for comprehensive S2S LLMs evaluation. First, unlike cascaded approaches that disregard acoustic features, SageLM jointly assesses both semantic and acoustic dimensions. Second, it leverages rationale-based supervision to enhance explainability and guide model learning, achieving superior alignment with evaluation outcomes compared to rule-based reinforcement learning methods. Third, we introduce SpeechFeedback, a synthetic preference dataset, and employ a two-stage training paradigm to mitigate the scarcity of speech preference data. Trained on both semantic and acoustic dimensions, SageLM achieves an 82.79% agreement rate with human evaluators, outperforming cascaded and SLM-based baselines by at least 7.42% and 26.20%, respectively. Yuan Ge 0001, Junxiang Zhang, Xiangnan Ma, Chenglong Wang 0002, Kaiyang Ye, Yangfan Du, Linfeng Zhang 0001, Yuxin Huang 0004, Tong Xiao 0001, Zhengtao Yu 0001 |
AAAI | 6 |
| 2026 | Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward ModelsabstractPrevious methods evaluate reward models by testing them on a fixed pairwise ranking test set, but they typically do not provide performance information on each preference dimension. In this work, we address the evaluation challenge of reward models by probing preference representations. To confirm the effectiveness of this evaluation method, we construct a Multi-dimensional Reward Model Benchmark (MRMBench), a collection of six probing tasks for different preference dimensions. We design it to favor and encourage reward models that better capture preferences across different dimensions. Furthermore, we introduce an analysis method, inference-time probing, which identifies the dimensions used during the reward prediction and enhances its interpretability. Through extensive experiments, we find that MRMBench strongly correlates with LLM alignment performance, supporting it as a reliable reference for developing advanced reward models. By analyzing the evaluation results on MRMBench, we reveal that reward models struggle to simultaneously capture preferences across multiple dimensions, highlighting the potential of multi-objective optimization in reward modeling. Furthermore, our results demonstrate that the proposed inference-time probing method provides a reliable metric for assessing the confidence of reward predictions, leading to improved alignment of large language models. Chenglong Wang 0002, Yifu Huo, Yang Gan, Yongyu Mu, Qiaozhi He, Murun Yang, Chunliang Zhang, Tongran Liu, Anxiang Ma, Zhengtao Yu 0001, Tong Xiao 0001 |
AAAI | 1 |
| 2026 | GRAM-R²: Self-Training Generative Foundation Reward Models for Reward ReasoningabstractMajor progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs to generalist reward models. Despite this trend, developing effective reward models remains a fundamental challenge: the heavy reliance on large-scale labeled preference data. Pre-training on abundant unlabeled data offers a promising direction, but existing approaches fall short in instilling explicit reasoning capabilities into reward models. To bridge this gap, we propose a self-training approach that can leverage unlabeled data to scale up reward reasoning in reward models. Based on this approach, we develop GRAM-R² a generative reward model trained to produce not only preference labels but also accompanying reward rationales. GRAM-R² can serve as a foundation model for reward reasoning and can be applied to a wide range of tasks with minimal or no additional fine-tuning. It can support downstream applications such as policy optimization and task-specific reward tuning. Experiments on response ranking, task adaptation, and reinforcement learning from human feedback demonstrate that GRAM-R² consistently delivers strong performance, outperforming several strong discriminative and generative baselines. Chenglong Wang 0002, Yongyu Mu, Yifu Huo, Jiali Zeng, Murun Yang, Xiaoyang Hao, Chunliang Zhang, Fandong Meng, Tong Xiao 0001 |
AAAI | 1 |
| 2026 | On the Emotion Understanding of Synthesized SpeechabstractYuan Ge, Haishu Zhao, AoKai Hao, Junxiang Zhang, Bei Li, Xiaoqian Liu, Chenglong Wang, Jianjin Wang, Bingsen Zhou, Bingyu Liu, JingBo Zhu, Zhengtao Yu, Tong Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuan Ge 0001, Haishu Zhao, Aokai Hao, Junxiang Zhang, Chenglong Wang 0002, Jianjin Wang, Bingsen Zhou, Zhengtao Yu 0001, Tong Xiao 0001 |
ACL (1) | 7 |
| 2026 | Cross-layer Attention Sharing for Pre-trained Large Language ModelsabstractAbstract To enhance the efficiency of the attention mechanism within large language models (LLMs), previous works primarily compress the Key-Value cache or group attention heads, while largely overlooking redundancy between layers. Our comprehensive analyses across various LLMs show that highly similar attention patterns persist within most layers. It’s intuitive to reduce the redundancy by sharing attention weights across layers. However, further analysis reveals two challenges: (1) Directly sharing the weight matrix without carefully rearranging the attention heads proves to be ineffective; (2) Shallow layers are vulnerable to small deviations in attention weights. Driven by these insights, we introduce LiSA, a lightweight substitute for self-attention in well-trained LLMs. LiSA employs tiny feed-forward networks to align attention heads between adjacent layers and low-rank matrices to approximate differences in layer-wise attention weights. Evaluations encompassing 13 typical benchmarks demonstrate that LiSA maintains high response quality in terms of accuracy and perplexity while reducing redundant attention calculations within 53% −84% of the total layers. Our implementations of LiSA achieve a 6 × compression of Q and K matrices within the attention mechanism, with maximum throughput improvements 19.5%, 32.3%, and 40.1% for LLaMA3-8B, LLaMA2-7B, and LLaMA2-13B, respectively. Our code is available at https://github.com/takagi97/lisa. Yongyu Mu, Yuzhang Wu, Yuchun Fan, Chenglong Wang 0002, Jiali Zeng, Qiaozhi He, Murun Yang, Fandong Meng, Jie Zhou 0016, Tong Xiao 0001 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference DataabstractLarge vision-language models (LVLMs) often fail to align with human preferences, leading to issues like generating misleading content without proper visual context (also known as hallucination). A promising solution to this problem is using human-preference alignment techniques, such as best-of-n sampling and reinforcement learning. However, these techniques face the difficulty arising from the scarcity of visual preference data, which is required to train a visual reward model (VRM). In this work, we continue the line of research. We present a Robust Visual Reward Model (RoVRM) which improves human-preference alignment for LVLMs. RoVRM leverages auxiliary textual preference data through a three-phase progressive training and optimal transport-based preference data selection to effectively mitigate the scarcity of visual preference data. We experiment with RoVRM on the commonly used vision-language tasks based on the LLaVA-1.5-7B and -13B models. Experimental results demonstrate that RoVRM consistently outperforms traditional VRMs. Furthermore, our three-phase progressive training and preference data selection approaches can yield consistent performance gains over ranking-based alignment techniques, such as direct preference optimization. Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Murun Yang, Qiaozhi He, Tong Xiao 0001, Chunliang Zhang, Tongran Liu |
AAAI | 1 |
| 2025 | Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language ModelsabstractKaiyan Chang, Yonghao Shi, Chenglong Wang, Hang Zhou, Chi Hu, Xiaoqian Liu, Yingfeng Luo, Yuan Ge, Tong Xiao, JingBo Zhu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Kaiyan Chang 0001, Yonghao Shi, Chenglong Wang 0002, Chi Hu, Yingfeng Luo, Yuan Ge 0001, Tong Xiao 0001 |
EMNLP | 3 |
| 2025 | Boosting Text-To-Image Generation via Multilingual Prompting in Large Multimodal ModelsabstractPrevious work on augmenting large multimodal models (LMMs) for text-to-image (T2I) generation has focused on enriching the input space of in-context learning (ICL). This includes providing a few demonstrations and optimizing image descriptions to be more detailed and logical. However, as demand for more complex and flexible image descriptions grows, enhancing comprehension of input text within the ICL paradigm remains a critical yet underexplored area. In this work, we extend this line of research by constructing parallel multilingual prompts aimed at harnessing the multilingual capabilities of LMMs. More specifically, we translate the input text into several languages and provide the models with both the original text and the translations. Experiments on two LMMs across 3 benchmarks show that our method, PMT2I, achieves superior performance in general, compositional, and fine-grained assessments, especially in human preference alignment Additionally, with its advantage of generating more diverse images, PMT2I significantly outperforms baseline prompts when incorporated with reranking methods. Our code and parallel multilingual data can be found at https://github.com/takagi97/PMT2I. Yongyu Mu, Junxin Wang, Xiaoxuan Zhou, Chenglong Wang 0002, Yingfeng Luo, Qiaozhi He, Tong Xiao 0001, Guocheng Chen |
ICASSP | 5 |
| 2025 | GRAM: A Generative Foundation Reward Model for Reward GeneralizationabstractIn aligning large language models (LLMs), reward models have played an important role, but are standardly trained as discriminative models and rely only on labeled human preference data. In this paper, we explore methods that train reward models using both unlabeled and labeled data. Building on the generative models in LLMs, we develop a generative reward model that is first trained via large-scale unsupervised learning and then fine-tuned via supervised learning. We also show that by using label smoothing, we are in fact optimizing a regularized pairwise ranking loss. This result, in turn, provides a new view of training reward models, which links generative models and discriminative models under the same class of training objectives. The outcome of these techniques is a foundation reward model, which can be applied to a wide range of tasks with little or no further fine-tuning effort. Extensive experiments show that this model generalizes well across several tasks, including response ranking, reinforcement learning from human feedback, and task adaptation with fine-tuning, achieving significant performance improvements over several strong baseline models. Chenglong Wang 0002, Yang Gan, Yifu Huo, Yongyu Mu, Qiaozhi He, Murun Yang, Tong Xiao 0001, Chunliang Zhang, Tongran Liu |
ICML | 1 |
| 2025 | MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward OptimizationabstractRecent advances in diffusion language models (DLMs) have presented a promising alternative to traditional autoregressive large language models (LLMs). However, DLMs still lag behind LLMs in reasoning performance, especially as the number of denoising steps decreases. Our analysis reveals that this shortcoming arises primarily from the independent generation of masked tokens across denoising steps, which fails to capture the token correlation. In this paper, we define two types of token correlation: intra-sequence correlation and inter-sequence correlation, and demonstrate that enhancing these correlations improves reasoning performance. To this end, we propose a Multi-Reward Optimization (MRO) approach, which encourages DLMs to consider the token correlation during the denoising process. More specifically, our MRO approach leverages test-time scaling, reject sampling, and reinforcement learning to directly optimize the token correlation with multiple elaborate rewards. Additionally, we introduce group step and importance sampling strategies to mitigate reward variance and enhance sampling efficiency. Through extensive experiments, we demonstrate that MRO not only improves reasoning performance but also achieves significant sampling speedups while maintaining high performance on reasoning benchmarks. Chenglong Wang 0002, Yang Gan, Chi Hu, Yongyu Mu, Murun Yang, Chunliang Zhang, Tongran Liu, Zhengtao Yu 0001, Tong Xiao 0001 |
NeurIPS | 1 |
| 2024 | ESRL: Efficient Sampling-Based Reinforcement Learning for Sequence GenerationabstractApplying Reinforcement Learning (RL) to sequence generation models enables the direct optimization of long-term rewards (e.g., BLEU and human feedback), but typically requires large-scale sampling over a space of action sequences. This is a computational challenge as presented by the practice of sequence generation problems, such as machine translation, where we often deal with a large action space (e.g., a vocabulary) and a long action sequence (e.g., a translation). In this work, we introduce two-stage sampling and dynamic sampling approaches to improve the sampling efficiency during training sequence generation models via RL. We experiment with our approaches on the traditional sequence generation tasks, including machine translation and abstractive summarization. Furthermore, we evaluate our approaches in RL from human feedback (RLHF) through training a large language model using the reward model. Experimental results show that the efficient sampling-based RL, referred to as ESRL, can outperform all baselines in terms of both training efficiency and memory consumption. Notably, ESRL yields consistent performance gains over the strong REINFORCE, minimum risk training, and proximal policy optimization methods. The code is available at https://github.com/wangclnlp/DeepSpeed-Chat-Extension/examples/esrl. Chenglong Wang 0002, Yimin Hu, Yifu Huo, Tongran Liu, Tong Xiao 0001 |
AAAI | 1 |
| 2024 | Revealing the Parallel Multilingual Learning within Large Language ModelsabstractYongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu, Bei Li, Chenglong Wang, Tong Xiao, Kai Song, Tongran Liu, Chunliang Zhang, JingBo Zhu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yongyu Mu, Peinan Feng, Zhiquan Cao, Yuzhang Wu, Chenglong Wang 0002, Tong Xiao 0001, Tongran Liu, Chunliang Zhang |
EMNLP | 6 |
| 2021 | RankNAS: Efficient Neural Architecture Search by Pairwise RankingabstractThis paper addresses the efficiency challenge of Neural Architecture Search (NAS) by formulating the task as a ranking problem.Previous methods require numerous training examples to estimate the accurate performance of architectures, although the actual goal is to find the distinction between "good" and "bad" candidates.Here we do not resort to performance predictors.Instead, we propose a performance ranking method (RankNAS) via pairwise ranking.It enables efficient architecture search using much fewer training examples.Moreover, we develop an architecture selection method to prune the search space and concentrate on more promising candidates.Extensive experiments on machine translation and language modeling tasks show that RankNAS can design high-performance architectures while being orders of magnitude faster than state-ofthe-art NAS systems. Chi Hu, Chenglong Wang 0002, Xiangnan Ma, Xia Meng, Yinqiao Li, Tong Xiao 0001, Changliang Li |
EMNLP (1) | 2 |