Zhenghao Xu

dblp:357/5585 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-0033-525XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 23% Language models and text generation · 20% Learning theory · 15%
Theoretical computer science
1 paper
Mathematical optimization · 50% Algorithms and data structures · 50%

Topics — the 27 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.722025
Ask a Strong LLM Judge when Your Reward Model is Uncertain · NeurIPS 2025
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.722025
Ask a Strong LLM Judge when Your Reward Model is Uncertain · NeurIPS 2025
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025
Machine learning › Deep learning architectures and training › training dynamics
edge of stability
0.912025
Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025
Machine learning › Reinforcement learning › reward learning › reward modeling
generative reward model
0.912025
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.912025
Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025
Machine learning › Learning theory
implicit bias
0.912025
Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025
Natural language and speech › Language models and text generation › large language model reasoning › multi-step reasoning
long-horizon reasoning
0.912025
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025
Machine learning › Optimization for machine learning
non-convex optimization
0.912025
Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025
Natural language and speech › Language models and text generation
preference optimization
0.912025
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025
Machine learning › Optimization for machine learning › gradient-based optimization
accelerated gradient methods
0.812024
Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024
Machine learning › Learning theory › statistical estimation › confidence set construction
confidence intervals
0.812024
Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024
Machine learning › Deep learning architectures and training
convolutional neural network
0.812024
Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds · J. Mach. Learn. Res. 2024
Machine learning › Learning theory
curse of dimensionality
0.812024
Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds · J. Mach. Learn. Res. 2024
Natural language and speech › Information extraction and text analysis › event analysis
event prediction
0.812024
Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process
0.812024
Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.812024
Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds · J. Mach. Learn. Res. 2024
Machine learning › Reinforcement learning
policy optimization
0.812024
Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
pseudolikelihood estimation
0.812024
Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024
Machine learning › Learning theory
sample complexity
0.812024
Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds · J. Mach. Learn. Res. 2024
Machine learning › Generative modeling
score matching
0.812024
Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
spatio-temporal point process
0.812024
Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024
Machine learning › Trustworthy machine learning
uncertainty estimation
0.812024
Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024
Algorithms and data structures › numerical linear algebra
matrix factorization
0.812024
Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024
Mathematical optimization
nonconvex optimization
0.812024
Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks
0.212024
Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

uncertainty quantification · 0.9supervised fine-tuning · 0.9rule-based reinforcement learning · 0.9policy gradient · 0.9pairwise preference classification · 0.9lyapunov analysis · 0.9convergence analysis · 0.9chain-of-thought · 0.9unbalanced initialization · 0.8score-based sampling · 0.8score matching · 0.8nesterov's accelerated gradient · 0.8gradient descent · 0.8
YearPublicationVenuePosition
2026 A Solution-Based Tabu Search for Quadratic Knapsack Problem with Conflict Graphs
Qihao Song, Zhenghao Xu, Xueshi Dong
ICIC (13)2
2025 Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
abstract
Reinforcement learning from human feedback (RLHF) has become a powerful post-training paradigm for aligning large language models with human preferences. A core challenge in RLHF is constructing accurate reward signals, where the conventional Bradley-Terry reward models (BT RMs) often suffer from sensitivity to data size and coverage, as well as vulnerability to reward hacking. Generative reward models (GenRMs) offer a more robust alternative by generating chain-of-thought (CoT) rationales followed by a final verdict. However, existing GenRMs rely on shallow, vertically scaled reasoning, limiting their capacity to handle nuanced or complex tasks. Moreover, their pairwise preference outputs are incompatible with standard RLHF algorithms that require pointwise reward signals. In this work, we introduce Think-RM, a training framework that enables long-horizon reasoning in GenRMs by modeling an internal thinking process. Rather than producing structured, externally provided rationales, Think-RM generates flexible, self-guided reasoning traces that support advanced capabilities such as self-reflection, hypothetical reasoning, and divergent reasoning. To elicit these reasoning abilities, we first warm-up the models by supervised fine-tuning (SFT) over long CoT data. We then further improve the model's long-horizon abilities by rule-based reinforcement learning (RL). In addition, we propose a novel pairwise RLHF pipeline that directly optimizes policies from pairwise comparisons, eliminating the need for pointwise reward conversion. Experiments show that Think-RM outperforms baselines on both in-distribution and out-of-distribution tasks, with particularly strong gains on reasoning-heavy benchmarks: more than 10\% and 5\% on RewardBench's Chat Hard and Reasoning, and 12\% on RM-Bench's Math domain. When combined with our pairwise RLHF pipeline, it demonstrates superior end-policy performance compared to traditional approaches. This depth-oriented approach not only broadens the GenRM design space but also establishes a new paradigm for preference-based policy optimization in RLHF.
Ilgee Hong, Changlong Yu, Weixiang Yan, Zhenghao Xu, Haoming Jiang, Qingru Zhang, Xin Liu 0039, Chao Zhang 0014, Tuo Zhao
NeurIPS5
2025 Ask a Strong LLM Judge when Your Reward Model is Uncertain
abstract
Reward model (RM) plays a pivotal role in reinforcement learning with human feedback (RLHF) for aligning large language models (LLMs). However, classical RMs trained on human preferences are vulnerable to reward hacking and generalize poorly to out-of-distribution (OOD) inputs. By contrast, strong LLM judges equipped with reasoning capabilities demonstrate superior generalization, even without additional training, but incur significantly higher inference costs, limiting their applicability in online RLHF. In this work, we propose an uncertainty-based routing framework that efficiently complements a fast RM with a strong but costly LLM judge. Our approach formulates advantage estimation in policy gradient (PG) methods as pairwise preference classification, enabling principled uncertainty quantification to guide routing. Uncertain pairs are forwarded to the LLM judge, while confident ones are evaluated by the RM. Experiments on RM benchmarks demonstrate that our uncertainty-based routing strategy significantly outperforms random judge calling at the same cost, and downstream alignment results showcase its effectiveness in improving online RLHF.
Zhenghao Xu, Qingru Zhang, Ilgee Hong, Changlong Yu, Wenlin Yao, Haoming Jiang, Lihong Li 0001, Hyokun Yun, Tuo Zhao
NeurIPS1
2025 Hybrid genetic algorithm with Wiener process for multi-scale colored balanced traveling salesman problem
Xueshi Dong, Liwen Ma, Yongchang Shan, Zhenghao Xu
Expert Syst. Appl.6
2025 Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult
abstract
Large learning rates, when applied to gradient descent for nonconvex optimization, yield various implicit biases including the edge of stability, balancing, and catapult. These phenomena cannot be well explained by classical optimization theory. Though significant theoretical progress has been made in understanding these implicit biases, it remains unclear for which objective functions they are more likely to occur --- more precisely, for which functions there exists a larger set of initial conditions that lead to these phenomena? This paper provides an initial step in answering this question and also shows that these implicit biases are in fact various tips of the same iceberg. To establish these results, we develop a global convergence theory under large learning rates, for a family of nonconvex functions without globally Lipschitz continuous gradient, which was typically assumed in existing convergence analysis. Specifically, these phenomena are more likely to occur when the optimization objective function has good regularity. This regularity, together with gradient descent using a large learning rate that favors flatter regions, results in these nontrivial dynamical behaviors. Another corollary is the first non-asymptotic convergence rate bound for large-learning-rate gradient descent optimization of nonconvex functions. Although our theory only applies to specific functions so far, the possibility of extrapolating it to neural networks is also experimentally validated, for which different choices of loss, activation functions, and other techniques such as batch normalization can all affect regularity significantly and lead to very different training dynamics.
Yuqing Wang 0005, Zhenghao Xu, Tuo Zhao, Molei Tao
J. Mach. Learn. Res.2
2024 Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process
abstract
Spatio-temporal point processes (STPPs) are potent mathematical tools for modeling and predicting events with both temporal and spatial features. Despite their versatility, most existing methods for learning STPPs either assume a restricted form of the spatio-temporal distribution, or suffer from inaccurate approximations of the intractable integral in the likelihood training objective. These issues typically arise from the normalization term of the probability density function. Moreover, existing works only provide point prediction for events without quantifying their uncertainty, such as confidence intervals for the event’s arrival time and confidence regions for the event’s location, which is crucial given the considerable randomness of the data. To tackle these challenges, we introduce SMASH: a Score MAtching-based pSeudolikeliHood estimator for learning marked STPPs. Specifically, our framework adopts a normalization-free objective by estimating the pseudolikelihood of marked STPPs through score-matching and predicts confidence intervals/regions for event time and location by generating samples through a score-based sampling algorithm. The superior performance of our proposed framework is demonstrated through extensive experiments on both point and confidence interval/region prediction of events.
Zichong Li, Qunzhi Xu, Zhenghao Xu, Yajun Mei, Tuo Zhao, Hongyuan Zha
ICML3
2024 A right-of-way allocation method for automated container terminals
abstract
This paper proposed a right-of-way allocation method for Automated Intelligent Vehicles (AIVs) in automated container terminals. The results demonstrate a notable improvement in travel efficiency, up to 60%, along with responsiveness measured in seconds.
Junqi Li, Lianhua An, Zhenghao Xu, Jia Hu 0003
IV4
2024 Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks
abstract
We study the convergence rate of first-order methods for rectangular matrix factorization, which is a canonical nonconvex optimization problem. Specifically, given a rank-$r$ matrix $\mathbf{A}\in\mathbb{R}^{m\times n}$, we prove that gradient descent (GD) can find a pair of $\epsilon$-optimal solutions $\mathbf{X}_T\in\mathbb{R}^{m\times d}$ and $\mathbf{Y}_T\in\mathbb{R}^{n\times d}$, where $d\geq r$, satisfying $\lVert\mathbf{X}_T\mathbf{Y}_T^\top-\mathbf{A}\rVert_F\leq\epsilon\lVert\mathbf{A}\rVert_F$ in $T=O(\kappa^2\log\frac{1}{\epsilon})$ iterations with high probability, where $\kappa$ denotes the condition number of $\mathbf{A}$. Furthermore, we prove that Nesterov's accelerated gradient (NAG) attains an iteration complexity of $O(\kappa\log\frac{1}{\epsilon})$, which is the best-known bound of first-order methods for rectangular matrix factorization. Different from small balanced random initialization in the existing literature, we adopt an unbalanced initialization, where $\mathbf{X}_0$ is large and $\mathbf{Y}_0$ is $0$. Moreover, our initialization and analysis can be further extended to linear neural networks, where we prove that NAG can also attain an accelerated linear convergence rate. In particular, we only require the width of the network to be greater than or equal to the rank of the output label matrix. In contrast, previous results achieving the same rate require excessive widths that additionally depend on the condition number and the rank of the input data matrix.
Zhenghao Xu, Yuqing Wang 0005, Tuo Zhao, Rachel Ward, Molei Tao
NeurIPS1
2024 Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds
abstract
Policy gradient methods equipped with deep neural networks have achieved great success in solving high-dimensional reinforcement learning (RL) problems. However, current analyses cannot explain why they are resistant to the curse of dimensionality. In this work, we study the sample complexity of the neural policy mirror descent (NPMD) algorithm with deep convolutional neural networks (CNN). Motivated by the empirical observation that many high-dimensional environments have state spaces possessing low-dimensional structures, such as those taking images as states, we consider the state space to be a $d$-dimensional manifold embedded in the $D$-dimensional Euclidean space with intrinsic dimension $d\ll D$. We show that in each iteration of NPMD, both the value function and the policy can be well approximated by CNNs. The approximation errors are controlled by the size of the networks, and the smoothness of the previous networks can be inherited. As a result, by properly choosing the network size and hyperparameters, NPMD can find an $\epsilon$-optimal policy with $\tilde{O}(\epsilon^{-\frac{d}{\alpha}-2})$ samples in expectation, where $\alpha\in(0,1]$ indicates the smoothness of environment. Compared to previous work, our result exhibits that NPMD can leverage the low-dimensional structure of state space to escape from the curse of dimensionality, explaining the efficacy of deep policy gradient algorithms.
Zhenghao Xu, Minshuo Chen, Mengdi Wang 0001, Tuo Zhao
J. Mach. Learn. Res.1