VLDB 2026 Research / reviewers in the wild / expert
Zhenghao Xu
dblp:357/5585
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-0033-525XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 23% Language models and text generation · 20% Learning theory · 15% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 50% Algorithms and data structures · 50% |
Topics — the 27 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
1.7 | 2 | 2025 | Ask a Strong LLM Judge when Your Reward Model is Uncertain · NeurIPS 2025 Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
1.7 | 2 | 2025 | Ask a Strong LLM Judge when Your Reward Model is Uncertain · NeurIPS 2025 Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › training dynamics
edge of stability |
0.9 | 1 | 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025 |
Machine learning › Reinforcement learning › reward learning › reward modeling
generative reward model |
0.9 | 1 | 2025 | Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.9 | 1 | 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025 |
Machine learning › Learning theory
implicit bias |
0.9 | 1 | 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025 |
Natural language and speech › Language models and text generation › large language model reasoning › multi-step reasoning
long-horizon reasoning |
0.9 | 1 | 2025 | Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025 |
Machine learning › Optimization for machine learning
non-convex optimization |
0.9 | 1 | 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult · J. Mach. Learn. Res. 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models · NeurIPS 2025 |
Machine learning › Optimization for machine learning › gradient-based optimization
accelerated gradient methods |
0.8 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Machine learning › Learning theory › statistical estimation › confidence set construction
confidence intervals |
0.8 | 1 | 2024 | Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.8 | 1 | 2024 | Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds · J. Mach. Learn. Res. 2024 |
Machine learning › Learning theory
curse of dimensionality |
0.8 | 1 | 2024 | Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds · J. Mach. Learn. Res. 2024 |
Natural language and speech › Information extraction and text analysis › event analysis
event prediction |
0.8 | 1 | 2024 | Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process |
0.8 | 1 | 2024 | Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.8 | 1 | 2024 | Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds · J. Mach. Learn. Res. 2024 |
Machine learning › Reinforcement learning
policy optimization |
0.8 | 1 | 2024 | Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds · J. Mach. Learn. Res. 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
pseudolikelihood estimation |
0.8 | 1 | 2024 | Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024 |
Machine learning › Learning theory
sample complexity |
0.8 | 1 | 2024 | Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds · J. Mach. Learn. Res. 2024 |
Machine learning › Generative modeling
score matching |
0.8 | 1 | 2024 | Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process
spatio-temporal point process |
0.8 | 1 | 2024 | Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.8 | 1 | 2024 | Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point Process · ICML 2024 |
Algorithms and data structures › numerical linear algebra
matrix factorization |
0.8 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Mathematical optimization
nonconvex optimization |
0.8 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks |
0.2 | 1 | 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
uncertainty quantification · 0.9supervised fine-tuning · 0.9rule-based reinforcement learning · 0.9policy gradient · 0.9pairwise preference classification · 0.9lyapunov analysis · 0.9convergence analysis · 0.9chain-of-thought · 0.9unbalanced initialization · 0.8score-based sampling · 0.8score matching · 0.8nesterov's accelerated gradient · 0.8gradient descent · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Solution-Based Tabu Search for Quadratic Knapsack Problem with Conflict Graphs
Qihao Song, Zhenghao Xu, Xueshi Dong |
ICIC (13) | 2 |
| 2025 | Think-RM: Enabling Long-Horizon Reasoning in Generative Reward ModelsabstractReinforcement learning from human feedback (RLHF) has become a powerful post-training paradigm for aligning large language models with human preferences. A core challenge in RLHF is constructing accurate reward signals, where the conventional Bradley-Terry reward models (BT RMs) often suffer from sensitivity to data size and coverage, as well as vulnerability to reward hacking. Generative reward models (GenRMs) offer a more robust alternative by generating chain-of-thought (CoT) rationales followed by a final verdict. However, existing GenRMs rely on shallow, vertically scaled reasoning, limiting their capacity to handle nuanced or complex tasks. Moreover, their pairwise preference outputs are incompatible with standard RLHF algorithms that require pointwise reward signals. In this work, we introduce Think-RM, a training framework that enables long-horizon reasoning in GenRMs by modeling an internal thinking process. Rather than producing structured, externally provided rationales, Think-RM generates flexible, self-guided reasoning traces that support advanced capabilities such as self-reflection, hypothetical reasoning, and divergent reasoning. To elicit these reasoning abilities, we first warm-up the models by supervised fine-tuning (SFT) over long CoT data. We then further improve the model's long-horizon abilities by rule-based reinforcement learning (RL). In addition, we propose a novel pairwise RLHF pipeline that directly optimizes policies from pairwise comparisons, eliminating the need for pointwise reward conversion. Experiments show that Think-RM outperforms baselines on both in-distribution and out-of-distribution tasks, with particularly strong gains on reasoning-heavy benchmarks: more than 10\% and 5\% on RewardBench's Chat Hard and Reasoning, and 12\% on RM-Bench's Math domain. When combined with our pairwise RLHF pipeline, it demonstrates superior end-policy performance compared to traditional approaches. This depth-oriented approach not only broadens the GenRM design space but also establishes a new paradigm for preference-based policy optimization in RLHF. Ilgee Hong, Changlong Yu, Weixiang Yan, Zhenghao Xu, Haoming Jiang, Qingru Zhang, Xin Liu 0039, Chao Zhang 0014, Tuo Zhao |
NeurIPS | 5 |
| 2025 | Ask a Strong LLM Judge when Your Reward Model is UncertainabstractReward model (RM) plays a pivotal role in reinforcement learning with human feedback (RLHF) for aligning large language models (LLMs). However, classical RMs trained on human preferences are vulnerable to reward hacking and generalize poorly to out-of-distribution (OOD) inputs.
By contrast, strong LLM judges equipped with reasoning capabilities demonstrate superior generalization, even without additional training, but incur significantly higher inference costs, limiting their applicability in online RLHF.
In this work, we propose an uncertainty-based routing framework that efficiently complements a fast RM with a strong but costly LLM judge. Our approach formulates advantage estimation in policy gradient (PG) methods as pairwise preference classification, enabling principled uncertainty quantification to guide routing. Uncertain pairs are forwarded to the LLM judge, while confident ones are evaluated by the RM. Experiments on RM benchmarks demonstrate that our uncertainty-based routing strategy significantly outperforms random judge calling at the same cost, and downstream alignment results showcase its effectiveness in improving online RLHF. Zhenghao Xu, Qingru Zhang, Ilgee Hong, Changlong Yu, Wenlin Yao, Haoming Jiang, Lihong Li 0001, Hyokun Yun, Tuo Zhao |
NeurIPS | 1 |
| 2025 | Hybrid genetic algorithm with Wiener process for multi-scale colored balanced traveling salesman problem
Xueshi Dong, Liwen Ma, Yongchang Shan, Zhenghao Xu |
Expert Syst. Appl. | 6 |
| 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapultabstractLarge learning rates, when applied to gradient descent for nonconvex optimization, yield various implicit biases including the edge of stability, balancing, and catapult. These phenomena cannot be well explained by classical optimization theory. Though significant theoretical progress has been made in understanding these implicit biases, it remains unclear for which objective functions they are more likely to occur --- more precisely, for which functions there exists a larger set of initial conditions that lead to these phenomena? This paper provides an initial step in answering this question and also shows that these implicit biases are in fact various tips of the same iceberg. To establish these results, we develop a global convergence theory under large learning rates, for a family of nonconvex functions without globally Lipschitz continuous gradient, which was typically assumed in existing convergence analysis. Specifically, these phenomena are more likely to occur when the optimization objective function has good regularity. This regularity, together with gradient descent using a large learning rate that favors flatter regions, results in these nontrivial dynamical behaviors. Another corollary is the first non-asymptotic convergence rate bound for large-learning-rate gradient descent optimization of nonconvex functions. Although our theory only applies to specific functions so far, the possibility of extrapolating it to neural networks is also experimentally validated, for which different choices of loss, activation functions, and other techniques such as batch normalization can all affect regularity significantly and lead to very different training dynamics. Yuqing Wang 0005, Zhenghao Xu, Tuo Zhao, Molei Tao |
J. Mach. Learn. Res. | 2 |
| 2024 | Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point ProcessabstractSpatio-temporal point processes (STPPs) are potent mathematical tools for modeling and predicting events with both temporal and spatial features. Despite their versatility, most existing methods for learning STPPs either assume a restricted form of the spatio-temporal distribution, or suffer from inaccurate approximations of the intractable integral in the likelihood training objective. These issues typically arise from the normalization term of the probability density function. Moreover, existing works only provide point prediction for events without quantifying their uncertainty, such as confidence intervals for the event’s arrival time and confidence regions for the event’s location, which is crucial given the considerable randomness of the data. To tackle these challenges, we introduce SMASH: a Score MAtching-based pSeudolikeliHood estimator for learning marked STPPs. Specifically, our framework adopts a normalization-free objective by estimating the pseudolikelihood of marked STPPs through score-matching and predicts confidence intervals/regions for event time and location by generating samples through a score-based sampling algorithm. The superior performance of our proposed framework is demonstrated through extensive experiments on both point and confidence interval/region prediction of events. Zichong Li, Qunzhi Xu, Zhenghao Xu, Yajun Mei, Tuo Zhao, Hongyuan Zha |
ICML | 3 |
| 2024 | A right-of-way allocation method for automated container terminalsabstractThis paper proposed a right-of-way allocation method for Automated Intelligent Vehicles (AIVs) in automated container terminals. The results demonstrate a notable improvement in travel efficiency, up to 60%, along with responsiveness measured in seconds. Junqi Li, Lianhua An, Zhenghao Xu, Jia Hu 0003 |
IV | 4 |
| 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural NetworksabstractWe study the convergence rate of first-order methods for rectangular matrix factorization, which is a canonical nonconvex optimization problem. Specifically, given a rank-$r$ matrix $\mathbf{A}\in\mathbb{R}^{m\times n}$, we prove that gradient descent (GD) can find a pair of $\epsilon$-optimal solutions $\mathbf{X}_T\in\mathbb{R}^{m\times d}$ and $\mathbf{Y}_T\in\mathbb{R}^{n\times d}$, where $d\geq r$, satisfying $\lVert\mathbf{X}_T\mathbf{Y}_T^\top-\mathbf{A}\rVert_F\leq\epsilon\lVert\mathbf{A}\rVert_F$ in $T=O(\kappa^2\log\frac{1}{\epsilon})$ iterations with high probability, where $\kappa$ denotes the condition number of $\mathbf{A}$. Furthermore, we prove that Nesterov's accelerated gradient (NAG) attains an iteration complexity of $O(\kappa\log\frac{1}{\epsilon})$, which is the best-known bound of first-order methods for rectangular matrix factorization. Different from small balanced random initialization in the existing literature, we adopt an unbalanced initialization, where $\mathbf{X}_0$ is large and $\mathbf{Y}_0$ is $0$. Moreover, our initialization and analysis can be further extended to linear neural networks, where we prove that NAG can also attain an accelerated linear convergence rate. In particular, we only require the width of the network to be greater than or equal to the rank of the output label matrix. In contrast, previous results achieving the same rate require excessive widths that additionally depend on the condition number and the rank of the input data matrix. Zhenghao Xu, Yuqing Wang 0005, Tuo Zhao, Rachel Ward, Molei Tao |
NeurIPS | 1 |
| 2024 | Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional ManifoldsabstractPolicy gradient methods equipped with deep neural networks have achieved great success in solving high-dimensional reinforcement learning (RL) problems. However, current analyses cannot explain why they are resistant to the curse of dimensionality. In this work, we study the sample complexity of the neural policy mirror descent (NPMD) algorithm with deep convolutional neural networks (CNN). Motivated by the empirical observation that many high-dimensional environments have state spaces possessing low-dimensional structures, such as those taking images as states, we consider the state space to be a $d$-dimensional manifold embedded in the $D$-dimensional Euclidean space with intrinsic dimension $d\ll D$. We show that in each iteration of NPMD, both the value function and the policy can be well approximated by CNNs. The approximation errors are controlled by the size of the networks, and the smoothness of the previous networks can be inherited. As a result, by properly choosing the network size and hyperparameters, NPMD can find an $\epsilon$-optimal policy with $\tilde{O}(\epsilon^{-\frac{d}{\alpha}-2})$ samples in expectation, where $\alpha\in(0,1]$ indicates the smoothness of environment. Compared to previous work, our result exhibits that NPMD can leverage the low-dimensional structure of state space to escape from the curse of dimensionality, explaining the efficacy of deep policy gradient algorithms. Zhenghao Xu, Minshuo Chen, Mengdi Wang 0001, Tuo Zhao |
J. Mach. Learn. Res. | 1 |