VLDB 2026 Research / reviewers in the wild / expert
Jie Cheng 0009
dblp:90/1457-9
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0008-5373-7563ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 52% Language models and text generation · 17% Deep learning architectures and training · 12% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment |
0.9 | 1 | 2025 | Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › mixture of experts
expert routing |
0.9 | 1 | 2025 | Ada-K Routing: Boosting the Efficiency of MoE-based LLMs · ICLR 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | Ada-K Routing: Boosting the Efficiency of MoE-based LLMs · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Ada-K Routing: Boosting the Efficiency of MoE-based LLMs · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.9 | 1 | 2025 | Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 1 | 2025 | Ada-K Routing: Boosting the Efficiency of MoE-based LLMs · ICLR 2025 |
Machine learning › Reinforcement learning › offline reinforcement learning
model-based offline reinforcement learning |
0.9 | 1 | 2025 | Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.9 | 1 | 2025 | Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model |
0.9 | 1 | 2025 | Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning · NeurIPS 2025 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning |
0.9 | 1 | 2025 | Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning · NeurIPS 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning · NeurIPS 2025 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.9 | 1 | 2025 | Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.8 | 1 | 2024 | SC- Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models · CVPR 2024 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning |
0.8 | 1 | 2024 | RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences · ICML 2024 |
Machine learning › Reinforcement learning
reward learning |
0.8 | 1 | 2024 | RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences · ICML 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 1 | 2024 | RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences · ICML 2024 |
Natural language and speech › Language models and text generation
self-consistency |
0.8 | 1 | 2024 | SC- Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models · CVPR 2024 |
Machine learning › Efficient and distributed learning › data-efficient learning
data-efficient fine-tuning |
0.3 | 1 | 2025 | Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.3 | 1 | 2025 | Ada-K Routing: Boosting the Efficiency of MoE-based LLMs · ICLR 2025 |
Computer vision › Segmentation and scene understanding › scene understanding
object understanding |
0.2 | 1 | 2024 | SC- Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9temporal difference learning · 0.9proximal policy optimization · 0.9process-supervised reinforcement learning · 0.9planning · 0.9min-form credit assignment · 0.9learnable allocator · 0.9dynamic expert selection · 0.9self-consistency tuning · 0.8cyclic describer-locator · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model PretrainingabstractA significant aspiration of offline reinforcement learning (RL) is to develop a generalist agent with high capabilities from large and heterogeneous datasets. However, prior approaches that scale offline RL either rely heavily on expert trajectories or struggle to generalize to diverse unseen tasks. Inspired by the excellent generalization of world model in conditional video generation, we explore the potential of image observation-based world model for scaling offline RL and enhancing generalization on novel tasks. In this paper, we introduce JOWA: Jointly-Optimized World-Action model, an offline model-based RL agent pretrained on multiple Atari games with 6 billion tokens data to learn general-purpose representation and decision-making ability. Our method jointly optimizes a world-action model through a shared transformer backbone, which stabilize temporal difference learning with large models during pretraining. Moreover, we propose a provably efficient and parallelizable planning algorithm to compensate for the Q-value estimation error and thus search out better policies. Experimental results indicate that our largest agent, with 150 million parameters, achieves 78.9% human-level performance on pretrained games using only 10% subsampled offline data, outperforming existing state-of-the-art large-scale offline RL baselines by 31.6% on averange. Furthermore, JOWA scales favorably with model capacity and can sample-efficiently transfer to novel games using only 5k offline fine-tuning data (approximately 4 trajectories) per game, demonstrating superior generalization. Jie Cheng 0009, Ruixi Qiao, Yingwei Ma, Binhua Li, Gang Xiong 0001, Qinghai Miao |
ICLR | 1 |
| 2025 | Ada-K Routing: Boosting the Efficiency of MoE-based LLMsabstractIn the era of Large Language Models (LLMs), Mixture-of-Experts (MoE) architectures offer a promising approach to managing computational costs while scaling up model parameters. Conventional MoE-based LLMs typically employ static Top-K routing, which activates a fixed and equal number of experts for each token regardless of their significance within the context. In this paper, we propose a novel Ada-K routing strategy that dynamically adjusts the number of activated experts for each token, thereby improving the balance between computational efficiency and model performance. Specifically, our strategy incorporates learnable and lightweight allocator modules that decide customized expert resource allocation tailored to the contextual needs for each token. These allocators are designed to be fully pluggable, making it broadly applicable across all mainstream MoE-based LLMs. We leverage the Proximal Policy Optimization (PPO) algorithm to facilitate an end-to-end learning process for this non-differentiable decision-making framework. Extensive evaluations on four popular baseline models demonstrate that our Ada-K routing method significantly outperforms conventional Top-K routing. Compared to Top-K, our method achieves over 25% reduction in FLOPs and more than 20% inference speedup while still improving performance across various benchmarks. Moreover, the training of Ada-K is highly efficient. Even for Mixtral-8x22B, a MoE-based LLM with more than 140B parameters, the training time is limited to 8 hours. Detailed analysis shows that harder tasks, middle layers, and content words tend to activate more experts, providing valuable insights for future adaptive MoE system designs. Both the training code and model checkpoints will be publicly available. Tongtian Yue, Longteng Guo, Jie Cheng 0009, Xuange Gao, Jing Liu 0001 |
ICLR | 3 |
| 2025 | Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for ReasoningabstractProcess reward model (PRM) has been proven effective in test-time scaling of LLM on challenging reasoning tasks. However, the reward hacking induced by PRM hinders its successful applications in reinforcement fine-tuning. We find the primary cause of reward hacking induced by PRM is that: the canonical summation-form credit assignment in reinforcement learning (RL), i.e. cumulative gamma-decayed future rewards, causes the LLM to hack steps with high rewards. Therefore, to unleashing the power of PRM in training-time, we propose PURE: Process sUpervised Reinforcement lEarning. The core of PURE is the min-form credit assignment that defines the value function as the minimum future rewards. This method unifies the optimization objective with respect to process rewards during test-time and training-time, and significantly alleviates reward hacking due to the limits on the range of values of value function and more rational assignment of advantages. Through extensively experiments on 3 base models, we achieve similar reasoning performance using PRM-based approach compared with verifiable reward-based approach if enabling min-form credit assignment. In contrast, the canonical sum-form credit assignment even collapses training at the beginning. Moreover, when we incorporate 1/10th verifiable rewards to auxiliary the PRM-based fine-tuning, it further alleviate reward hacking and results in the best fine-tuned model based on Qwen2.5-Math-7B with 82.5% accuracy on AMC23 and 53.3% average accuracy across 5 benchmarks. Furthermore, we summary the reward hacking cases we encountered during training and analysis the cause of training collapse. Jie Cheng 0009, Gang Xiong 0001, Ruixi Qiao, Chao Guo 0006, Junle Wang, Fei-Yue Wang 0001 |
NeurIPS | 1 |
| 2024 | SC- Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language ModelsabstractRecent trends in Large Vision Language Models (LVLMs) research have been increasingly focusing on ad-vancing beyond general image understanding towards more nuanced, object-level referential comprehension. In this paper, we present and delve into the self-consistency ca-pability of LVLMs, a crucial aspect that reflects the mod-els' ability to both generate informative captions for spe-cific objects and subsequently utilize these captions to ac-curately re-identify the objects in a closed-loop process. This capability significantly mirrors the precision and reli-ability of fine- grained visual-language understanding. Our findings reveal that the self-consistency level of existing LVLMs falls short of expectations, posing limitations on their practical applicability and potential. To address this gap, we introduce a novel fine-tuning paradigm named Self-Consistency Tuning (SC-Tune). It features the syn-ergistic learning of a cyclic describer-locator system. This paradigm is not only data-efficient but also exhibits gener-alizability across multiple LVLMs. Through extensive ex-periments, we demonstrate that SC- Tune significantly ele-vates performance across a spectrum of object-level vision-language benchmarks and maintains competitive or im-proved performance on image-level vision-language bench-marks. Both our model and code will be publicly available at https://github.com/ivattyue/SC-Tune. Tongtian Yue, Jie Cheng 0009, Longteng Guo, Xingyuan Dai, Zijia Zhao, Xingjian He, Gang Xiong 0001, Jing Liu 0001 |
CVPR | 2 |
| 2024 | RIME: Robust Preference-based Reinforcement Learning with Noisy PreferencesabstractPreference-based Reinforcement Learning (PbRL) circumvents the need for reward engineering by harnessing human preferences as the reward signal. However, current PbRL methods excessively depend on high-quality feedback from domain experts, which results in a lack of robustness. In this paper, we present RIME, a robust PbRL algorithm for effective reward learning from noisy preferences. Our method utilizes a sample selection-based discriminator to dynamically filter out noise and ensure robust training. To counteract the cumulative error stemming from incorrect selection, we suggest a warm start for the reward model, which additionally bridges the performance gap during the transition from pre-training to online training in PbRL. Our experiments on robotic manipulation and locomotion tasks demonstrate that RIME significantly enhances the robustness of the state-of-the-art PbRL method. Code is available at https://github.com/CJReinforce/RIME_ICML2024. Jie Cheng 0009, Gang Xiong 0001, Xingyuan Dai, Qinghai Miao, Fei-Yue Wang 0001 |
ICML | 1 |