Ruixi Qiao

dblp:389/2954 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 84% Language models and text generation · 12% Efficient and distributed learning · 4%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment
0.912025
Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning · NeurIPS 2025
Machine learning › Reinforcement learning › offline reinforcement learning
model-based offline reinforcement learning
0.912025
Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
0.912025
Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning · NeurIPS 2025
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement fine-tuning
0.912025
Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning · NeurIPS 2025
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.912025
Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025
Machine learning › Efficient and distributed learning › data-efficient learning
data-efficient fine-tuning
0.312025
Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining · ICLR 2025

Methods — techniques the papers use, named apart from their topics

transformer · 0.9temporal difference learning · 0.9process-supervised reinforcement learning · 0.9planning · 0.9min-form credit assignment · 0.9
YearPublicationVenuePosition
2026 Evaluating the Perceptual Robustness of Vision-Language Models for Autonomous Driving in Corner Cases
Peizhe Gong, Enming Zhang, Ruixi Qiao, Xingyuan Dai, Xiaoyan Gong, Qinghai Miao
IV3
2025 Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining
abstract
A significant aspiration of offline reinforcement learning (RL) is to develop a generalist agent with high capabilities from large and heterogeneous datasets. However, prior approaches that scale offline RL either rely heavily on expert trajectories or struggle to generalize to diverse unseen tasks. Inspired by the excellent generalization of world model in conditional video generation, we explore the potential of image observation-based world model for scaling offline RL and enhancing generalization on novel tasks. In this paper, we introduce JOWA: Jointly-Optimized World-Action model, an offline model-based RL agent pretrained on multiple Atari games with 6 billion tokens data to learn general-purpose representation and decision-making ability. Our method jointly optimizes a world-action model through a shared transformer backbone, which stabilize temporal difference learning with large models during pretraining. Moreover, we propose a provably efficient and parallelizable planning algorithm to compensate for the Q-value estimation error and thus search out better policies. Experimental results indicate that our largest agent, with 150 million parameters, achieves 78.9% human-level performance on pretrained games using only 10% subsampled offline data, outperforming existing state-of-the-art large-scale offline RL baselines by 31.6% on averange. Furthermore, JOWA scales favorably with model capacity and can sample-efficiently transfer to novel games using only 5k offline fine-tuning data (approximately 4 trajectories) per game, demonstrating superior generalization.
Jie Cheng 0009, Ruixi Qiao, Yingwei Ma, Binhua Li, Gang Xiong 0001, Qinghai Miao
ICLR2
2025 Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
abstract
Process reward model (PRM) has been proven effective in test-time scaling of LLM on challenging reasoning tasks. However, the reward hacking induced by PRM hinders its successful applications in reinforcement fine-tuning. We find the primary cause of reward hacking induced by PRM is that: the canonical summation-form credit assignment in reinforcement learning (RL), i.e. cumulative gamma-decayed future rewards, causes the LLM to hack steps with high rewards. Therefore, to unleashing the power of PRM in training-time, we propose PURE: Process sUpervised Reinforcement lEarning. The core of PURE is the min-form credit assignment that defines the value function as the minimum future rewards. This method unifies the optimization objective with respect to process rewards during test-time and training-time, and significantly alleviates reward hacking due to the limits on the range of values of value function and more rational assignment of advantages. Through extensively experiments on 3 base models, we achieve similar reasoning performance using PRM-based approach compared with verifiable reward-based approach if enabling min-form credit assignment. In contrast, the canonical sum-form credit assignment even collapses training at the beginning. Moreover, when we incorporate 1/10th verifiable rewards to auxiliary the PRM-based fine-tuning, it further alleviate reward hacking and results in the best fine-tuned model based on Qwen2.5-Math-7B with 82.5% accuracy on AMC23 and 53.3% average accuracy across 5 benchmarks. Furthermore, we summary the reward hacking cases we encountered during training and analysis the cause of training collapse.
Jie Cheng 0009, Gang Xiong 0001, Ruixi Qiao, Chao Guo 0006, Junle Wang, Fei-Yue Wang 0001
NeurIPS3