Jiaxuan Gao

dblp:304/2243 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Beyond Generic Answers and Empty Praise: The Differential Impact of AI Companionship Roles on Social-Emotional Learning in University Student Intervention
Shun-Lam Chan, Jinglun Zhao, Jiaxuan Gao, Xiaofei Lyu
AIED3
2026 MIFE: A Multimodal VR Immersive Training System for Fire Escape
abstract
High-rise fire escape training for the public poses a significant challenge, owing to the prohibitive cost and logistical intricacies involved in deploying high-fidelity simulations. While virtual reality (VR) technologies have demonstrated potential in safety education, the number of interaction modalities they offer is mostly limited to 3 to 5 types, and they heavily rely on controllers that abstract actions into button presses, limiting immersion and skill transfer. In this paper, we propose the MIFE system, which designs 8 distinct interaction modalities along with a real-time dynamic fire spread simulation to further improve immersion, allowing trainees to perceive the fire scene and efficiently master key escape skills through vision, auditory, olfactory, etc. Moreover, we have designed a controller-free and real-time full-body motion capture (MoCap) module, achieving precise mapping between the trainees’ full-body movements and the virtual avatar. Additionally, MIFE leverages a knowledge graph to provide tailored guidance, adapting dynamically to trainees with varying levels of fire escape proficiency. The results of the user study demonstrate that MIFE significantly outperforms self-study and controller-based training systems, particularly, improving by 3.08/10 and 2.38/10 in the performance score of the evaluation stage, respectively. This implies its practical utility and potential for broader adoption in high-rise fire emergency training.
Yalei Liu, Jiaxuan Gao, Qiuyu Fu
VR3
2025 Industrial-Grade Sensor Simulation via Gaussian Splatting: A Modular Framework for Scalable Editing and Full-Stack Validation
abstract
Sensor simulation is pivotal for scalable validation of autonomous driving systems, yet existing Neural Radiance Fields (NeRF) based methods face applicability and efficiency challenges in industrial workflows. This paper introduces a Gaussian Splatting (GS) based system to address these challenges: We first break down sensor simulator components and analyze the possible advantages of GS over NeRF. Then in practice, we refactor three crucial components through GS, to leverage its explicit scene representation and real-time rendering: (1) choosing the 2D neural Gaussian representation for physics-compliant scene and sensor modeling, (2) proposing a scene editing pipeline to leverage Gaussian primitives library for data augmentation, and (3) coupling a controllable diffusion model for scene expansion and harmonization. We implement this framework on a proprietary autonomous driving dataset supporting cameras and LiDAR sensors. We demonstrate through ablation studies that our approach reduces frame-wise simulation latency, achieves better geometric and photometric consistency, and enables interpretable explicit scene editing and expansion. Furthermore, we showcase how integrating such a GS-based sensor simulator with traffic and dynamic simulators enables full-stack testing of end-to-end autonomy algorithms. Our work provides both algorithmic insights and practical validation, establishing GS as a cornerstone for industrial-grade sensor simulation.
Xianming Zeng, Sicong Du, Lizhe Liu, Haoyu Shu, Jiaxuan Gao, Jiarun Liu, Jiulong Xu, Jianyun Xu, Mingxia Chen, Yiru Zhao, Yapeng Xue, Sheng Yang 0007
IROS6
2025 AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
abstract
Reinforcement learning (RL) has become a trending paradigm for training large language models (LLMs), particularly for reasoning tasks. Effective RL for LLMs requires massive parallelization and poses an urgent need for efficient training systems. Most existing large-scale RL systems for LLMs are synchronous by alternating generation and training in a batch setting, where the rollouts in each training batch are generated by the same (or latest) model. This stabilizes RL training but suffers from severe system-level inefficiency. Generation must wait until the longest output in the batch is completed before model update, resulting in GPU underutilization. We present AReaL, a fully asynchronous RL system that completely decouples generation from training. Rollout workers in AReaL continuously generate new outputs without waiting, while training workers update the model whenever a batch of data is collected. AReaL also incorporates a collection of system-level optimizations, leading to substantially higher GPU utilization. To stabilize RL training, AReaL balances the workload of rollout and training workers to control data staleness, and adopts a staleness-enhanced PPO variant to better handle outdated training samples. Extensive experiments on math and code reasoning benchmarks show that AReaL achieves up to 2.77x training speedup compared to synchronous systems with the same number of GPUs and matched or even improved final performance. The code of AReaL is available at https://github.com/inclusionAI/AReaL/.
Jiaxuan Gao, Xujie Shen, Zhiyu Mei, Chuyi He, Shusheng Xu, Jun Mei, Tongkai Yang, Binhang Yuan
NeurIPS2
2025 How Far Are We from Optimal Reasoning Efficiency?
abstract
Large Reasoning Models (LRMs) demonstrate remarkable problem-solving capabilities through extended Chain-of-Thought (CoT) reasoning but often produce excessively verbose and redundant reasoning traces. This inefficiency incurs high inference costs and limits practical deployment. While existing fine-tuning methods aim to improve reasoning efficiency, assessing their efficiency gains remains challenging due to inconsistent evaluations. In this work, we introduce the ***reasoning efficiency frontiers***, empirical upper bounds derived from fine-tuning a base LRM (DeepSeek-R1-Distill-Qwen-1.5B/7B) across diverse approaches and training configurations. Based on these frontiers, we propose the ***Reasoning Efficiency Gap (REG)***, a unified metric quantifying deviations of any fine-tuned LRMs from these frontiers. Systematic evaluation on challenging mathematical benchmarks, AMC23, AIME24, and AIME25, reveals significant gaps in current methods: they either sacrifice accuracy for short length or use excessive tokens to achieve sub-optimal accuracies despite high overall accuracy. To reduce the efficiency gap, we propose ***REO-RL***, a Reinforcement Learning algorithm that optimizes reasoning efficiency by targeting a sparse set of token budgets. Leveraging numerical integration over strategically selected budgets, REO-RL approximates the full efficiency objective with low error using a small set of token budgets. Experiments show that, compared to vanilla RL with outcome reward, REO-RL reduces the reasoning efficiency gap by 74.5\% and 64.2\% in the 1.5B and 7B settings. The 7B LRM fine-tuned with REO-RL achieves reasoning conciseness surpassing frontier LRMs like Qwen3 and Claude Sonnet 3.7. Ablation studies confirm the efficacy of our token budget strategy and highlight REO-RL’s flexibility across design choices. This work establishes a systematic framework for evaluating and optimizing reasoning efficiency in LRMs. We will release the related code, data, and models to support future research on efficient reasoning in LRMs.
Jiaxuan Gao, Shu Yan, Qixin Tan, Shusheng Xu, Zhiyu Mei, Kaifeng Lyu, Yi Wu 0013
NeurIPS1
2025 Reasoning Is Not a Race: When Stopping Early Beats Going Deeper
abstract
We study the use of Process Reward Models (PRMs) for guiding Long Chain-of-Thought (CoT) reasoning in large language models. Although PRMs deliver fine-grained feedback in standard tasks, PRM-guided beam search does not consistently outperform PRM-free approaches in long CoT reasoning. We trace this shortfall to a "step quality degradation''—the expected step quality shows concave behavior, yielding unimodal or monotonically declining trends. To counteract this, we propose Z-Score Guided Early Stopping (ZGES), which halts search at the detected quality peak using local PRM-reward z-scores. Across multiple math benchmarks and model scales, ZGES outperforms both standard PRM-guided beam search and the PRM-free methods. Ablation studies further highlight the advantages and robustness of ZGES’s adaptive stopping mechanism.
Mohan Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu 0013
NeurIPS2
2025 Off-grid DOA Estimation for Disturbed UAV Swarm with Nested Array
abstract
Unmanned aerial vehicles (UAVs), with the advantage in on-demand and cost-effective deployment, have emerged as an aerial communication platform for supporting constant high-data rate transmission services for terrestrial stations with different mobility patterns. To this end, developing high-accuracy localization functionality is of crucial importance in tracking the target users. However, the limited payload size of a single UAV restricts precise direction-of-arrival (DOA) estimation, and the UAV-swarm based collaborative estimation architectures have become a promising remedy. In this paper, the UAV swarm is utilized to form a nested array (NA) architecture to improve the DOA estimation accuracy. A two-stage alternating iterative approach is proposed to jointly estimate the UAV position errors and DOAs through iterative feedback. Specifically, position error estimation is performed using the gradient descent method, whereas DOA estimation is achieved through a novel NA-Bessel atomic transformation and norm minimization (NA-BATNM) framework, built upon an off-grid atomic norm minimization algorithm. Simulation results demonstrate that the proposed method achieves high-precision estimation of both DOAs and UAV position errors. Furthermore, the proposed NA-BATNM method exhibits superior performance compared to grid-based DOA estimation algorithms.
Jiaxuan Gao, Yanyan Wang 0009, Yang Cao 0018, Xiaohu Tang 0004
VTC2025-Fall1
2024 SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores
abstract
The ever-growing complexity of reinforcement learning (RL) tasks demands a distributed system to efficiently generate and process a massive amount of data. However, existing open-source libraries suffer from various limitations, which impede their practical use in challenging scenarios where large-scale training is necessary. In this paper, we present a novel abstraction on the dataflows of RL training, which unifies diverse RL training applications into a general framework. Following this abstraction, we develop a scalable, efficient, and extensible distributed RL system called ReaLly Scalable RL (SRL), which allows efficient and massively parallelized training and easy development of customized algorithms. Our evaluation shows that SRL outperforms existing academic libraries, reaching at most 21x higher training throughput in a distributed setting. On learning performance, beyond performing and scaling well on common RL benchmarks with different RL algorithms, SRL can reproduce the same solution in the challenging hide-and-seek environment as reported by OpenAI with up to 5x speedup in wallclock time. Notably, SRL is the first in the academic community to perform RL experiments at a large scale with over 15k CPU cores. SRL anonymous repository is available at: https://anonymous.4open.science/r/srl-1E45/.
Zhiyu Mei, Jiaxuan Gao, Guangju Wang, Huanchen Zhang, Yi Wu 0013
ICLR3
2024 Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
abstract
Reinforcement Learning from Human Feedback (RLHF) is currently the most widely used method to align large language models (LLMs) with human preferences. Existing RLHF methods can be roughly categorized as either reward-based or reward-free. Novel applications such as ChatGPT and Claude leverage reward-based methods that first learn a reward model and apply actor-critic algorithms, such as Proximal Policy Optimization (PPO). However, in academic benchmarks, state-of-the-art results are often achieved via reward-free methods, such as Direct Preference Optimization (DPO). Is DPO truly superior to PPO? Why does PPO perform poorly on these benchmarks? In this paper, we first conduct both theoretical and empirical studies on the algorithmic properties of DPO and show that DPO may have fundamental limitations. Moreover, we also comprehensively examine PPO and reveal the key factors for the best performances of PPO in fine-tuning LLMs. Finally, we benchmark DPO and PPO across a collection of RLHF testbeds, ranging from dialogue to code generation. Experiment results demonstrate that PPO is able to surpass other alignment methods in all cases and achieve state-of-the-art results in challenging code competitions.
Shusheng Xu, Jiaxuan Gao, Zhiyu Mei, Guangju Wang, Chao Yu 0005, Yi Wu 0013
ICML3
2024 LAGOON: Language-Guided Motion Control
abstract
We aim to control a robot to physically behave in the real world following any high-level language command like "cartwheel" or "kick". Although human motion datasets exist, this task remains particularly challenging since generative models can produce physically unrealistic motions, which will be more severe for robots due to different body structures and physical properties. Deploying such a motion to a physical robot can cause even greater difficulties due to the sim2real gap. We develop LAnguage-Guided mOtion cONtrol (LAGOON), a multi-phase reinforcement learning (RL) method to generate physically realistic robot motions under language commands. LAGOON first leverages a pretrained model to generate a human motion from a language command. Then an RL phase trains a control policy in simulation to mimic the generated human motion. Finally, with domain randomization, our learned policy can be deployed to a quadrupedal robot, leading to a quadrupedal robot that can take diverse behaviors in the real world under natural language commands.
Shusheng Xu, Huaijie Wang, Yutao Ouyang, Jiaxuan Gao, Zhiyu Mei, Chao Yu 0005, Yi Wu 0013
ICRA4
2024 Robot Generating Data for Learning Generalizable Visual Robotic Manipulation
abstract
It has been a popular trend in AI to pretrain foundation models on massive data. However, collecting sufficient offline training trajectories for robot learning is particularly expensive since valid control actions are required. Therefore, most existing robotic datasets are collected from human experts. We tackle such a data collection issue with a new framework called "robot self-teaching", which asks the robot to self-generate effective training data instead of relying on human demonstrators. Our key idea is to train a separate data-generation policy operating on the state space to automatically generate meaningful actions and trajectories with ever-growing complexities. Then, these generated data can be further used to train a visual policy with strong compositional generalization capabilities. We validate our framework in two visual manipulation testbeds, including a multi-object stacking domain and a popular RL benchmark "Franka kitchen". Experiments show that the final visual policy trained on self-generated data can accomplish novel testing goals that require long-horizon robot executions. Project website https://sites.google.com/view/robot-self-teaching.
Yunfei Li 0005, Jingzhi Cui, Haoran Huan, Jiaxuan Gao, Yi Wu 0013
IROS6
2024 ReBEC: A replacement-based energy-efficient fault-tolerance design for associative caches
Xin Gao 0016, Naiyuan Cui, Jiawei Nian, Zongnan Liang, Jiaxuan Gao, Hongjin Liu, Mengfei Yang
Future Gener. Comput. Syst.5
2023 Learning Zero-Shot Cooperation with Humans, Assuming Humans Are Biased
Chao Yu 0005, Jiaxuan Gao, Botian Xu, Yu Wang 0002, Yi Wu 0013
ICLR2
2023 SpeedyZero: Mastering Atari with Limited Data and Time
Yixuan Mei, Jiaxuan Gao, Weirui Ye, Shaohuai Liu, Yang Gao 0029, Yi Wu 0013
ICLR2
2023 An Algorithm of Bistatic Sar Echo Generation Considering Shadow and Overlay Effects
abstract
The actual detection scenes faced by bistatic synthetic aperture radar (SAR) often have elevation information, which will cause shadow and overlay effects in imaging results. In view of the above problems, this paper proposes an echo generation method of bistatic SAR based on hidden point removal(HPR) operator. In this paper, according to the bistatic configuration, the shadow area is deduced first by using blanking algorithm(HPR operator). On this basis, the visibility of the target in the scene can be determined to obtain the echo. Finally, a simulation using a digital elevation model is conducted to verify the accuracy of the proposed method.
Jiaxuan Gao, Yue Song 0003, Zhongyu Li 0001, Jintao Xiong, Junjie Wu 0001, Jianyu Yang 0001
IGARSS1
2022 Learning Efficient Multi-agent Cooperative Visual Exploration
Chao Yu 0005, Xinyi Yang 0001, Jiaxuan Gao, Huazhong Yang, Yu Wang 0002, Yi Wu 0013
ECCV (39)3
2022 SAVE: Spatial-Attention Visual Exploration
abstract
Visual indoor exploration requires agents to explore a room in a limited time. Currently, planning-based solutions have a time-consuming inference stage and require many handcrafted parameters in different scenes. Reinforcement Learning (RL) schemes on the other hand solve these problems by automatically updating flexible policies and affording faster inference time. Spurred by the advantages of RL, we introduce Spatial Attention Visual Exploration (SAVE), which is based on Active Neural SLAM (ANS) [1]. Specifically, we propose a novel RL-based global planner named Spatial Global Policy (SGP) that utilizes spatial information to promote efficient exploration through global goal guidance. SGP has two major components: a transformer-based spatial-attention module encoding spatial interrelation between the agent and different regions to perform spatial reasoning, and a hierarchical spatial action selector to infer global goals for faster training. The map representations are aligned through our spatial adjustor. Experiments on the Habitat photo-realistic simulator [2] demonstrate that SAVE outperforms current planning-based methods and RL variants, reducing at least 10% of the processing steps, 15% of the repeat ratio, and affording an x2 to x4 faster execution time than planning-based methods.
Xinyi Yang 0001, Chao Yu 0005, Jiaxuan Gao, Yu Wang 0002, Huazhong Yang
ICIP3
2022 The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games
abstract
Proximal Policy Optimization (PPO) is a ubiquitous on-policy reinforcement learning algorithm but is significantly less utilized than off-policy learning algorithms in multi-agent settings. This is often due to the belief that PPO is significantly less sample efficient than off-policy methods in multi-agent systems. In this work, we carefully study the performance of PPO in cooperative multi-agent settings. We show that PPO-based multi-agent algorithms achieve surprisingly strong performance in four popular multi-agent testbeds: the particle-world environments, the StarCraft multi-agent challenge, the Hanabi challenge, and Google Research Football, with minimal hyperparameter tuning and without any domain-specific algorithmic modifications or architectures. Importantly, compared to competitive off-policy methods, PPO often achieves competitive or superior results in both final returns and sample efficiency. Finally, through ablation studies, we analyze implementation and hyperparameter factors that are critical to PPO's empirical performance, and give concrete practical suggestions regarding these factors. Our results show that when using these practices, simple PPO-based methods are a strong baseline in cooperative multi-agent reinforcement learning. Source code is released at https://github.com/marlbenchmark/on-policy.
Chao Yu 0005, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang 0002, Alexandre M. Bayen, Yi Wu 0013
NeurIPS4