Zhenxing Ge

dblp:300/5437 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-7832-2363ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 56% Multi-agent systems · 44%
Theoretical computer science
3 papers
Algorithmic game theory and mechanism design · 100%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Algorithmic game theory and mechanism design
equilibrium computation
2.132026
Faster Game Solving via Asymmetry of Step Sizes · AAAI 2026
Efficient Last-Iterate Convergence in Solving Extensive-Form Games · NeurIPS 2025
Safe and Robust Subgame Exploitation in Imperfect Information Games · ICML 2024
Algorithmic game theory and mechanism design › equilibrium computation
counterfactual regret minimization
1.922026
Faster Game Solving via Asymmetry of Step Sizes · AAAI 2026
Efficient Last-Iterate Convergence in Solving Extensive-Form Games · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems
imperfect information games
1.322023
Efficient Subgame Refinement for Extensive-form Games · NeurIPS 2023
An Efficient Deep Reinforcement Learning Algorithm for Solving Imperfect Information Extensive-Form Games · AAAI 2023
Machine learning › Reinforcement learning
regret minimization
1.012026
Faster Game Solving via Asymmetry of Step Sizes · AAAI 2026
Machine learning › Reinforcement learning
exploration
0.912025
Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
preference-based reinforcement learning
0.912025
Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning
reward model training
0.912025
Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement Learning · NeurIPS 2025
Algorithmic game theory and mechanism design › non-cooperative game
extensive-form games
0.912025
Efficient Last-Iterate Convergence in Solving Extensive-Form Games · NeurIPS 2025
Algorithmic game theory and mechanism design › game dynamics › equilibrium convergence
last-iterate convergence
0.912025
Efficient Last-Iterate Convergence in Solving Extensive-Form Games · NeurIPS 2025
Algorithmic game theory and mechanism design
regret minimization
0.912025
Efficient Last-Iterate Convergence in Solving Extensive-Form Games · NeurIPS 2025
Algorithmic game theory and mechanism design
imperfect information games
0.812024
Safe and Robust Subgame Exploitation in Imperfect Information Games · ICML 2024
Algorithmic game theory and mechanism design › learning in games
opponent exploitation
0.812024
Safe and Robust Subgame Exploitation in Imperfect Information Games · ICML 2024
Machine learning › Reinforcement learning › regret minimization
counterfactual regret minimization
0.712023
An Efficient Deep Reinforcement Learning Algorithm for Solving Imperfect Information Extensive-Form Games · AAAI 2023
Knowledge, reasoning and agents › Multi-agent systems › game theory
extensive-form games
0.712023
Efficient Subgame Refinement for Extensive-form Games · NeurIPS 2023
Knowledge, reasoning and agents › Multi-agent systems
game theory
0.712023
Efficient Subgame Refinement for Extensive-form Games · NeurIPS 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.712023
An Efficient Deep Reinforcement Learning Algorithm for Solving Imperfect Information Extensive-Form Games · AAAI 2023
Knowledge, reasoning and agents › Multi-agent systems › equilibrium computation
nash equilibrium computation
0.712023
An Efficient Deep Reinforcement Learning Algorithm for Solving Imperfect Information Extensive-Form Games · AAAI 2023
Knowledge, reasoning and agents › Multi-agent systems › game solving
subgame solving
0.712023
Efficient Subgame Refinement for Extensive-form Games · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

step-size asymmetry · 2.0predictive CFR+ · 2.0reward transformation · 0.9regret matching · 0.9proximal policy extension · 0.9online mirror descent · 0.9mixture distribution query · 0.9real-time search · 0.8game-theoretic analysis · 0.8maximum entropy deep reinforcement learning · 0.7generative subgame solving · 0.7follow-the-regularized-leader · 0.7diversity-based generation · 0.7
YearPublicationVenuePosition
2026 Faster Game Solving via Asymmetry of Step Sizes
abstract
Counterfactual Regret Minimization (CFR) algorithms are widely used to compute a Nash equilibrium (NE) in two-player zero-sum imperfect-information extensive-form games (IIGs). Among them, Predictive CFR+ (PCFR+) is particularly powerful, achieving an exceptionally fast empirical convergence rate via the prediction in many games. However, the empirical convergence rate of PCFR+ would significantly degrade if the prediction is inaccurate, leading to unstable performance on certain IIGs. To enhance the robustness of PCFR+, we propose Asymmetric PCFR+ (APCFR+), which employs an adaptive asymmetry of step sizes between the updates of implicit and explicit accumulated counterfactual regrets to mitigate the impact of the prediction inaccuracy on convergence. We present a theoretical analysis demonstrating why APCFR+ can enhance the robustness. To the best of our knowledge, we are the first to propose the asymmetry of step sizes, a simple yet novel technique that effectively improves the robustness of PCFR+. Then, to reduce the difficulty of implementing APCFR+ caused by the adaptive asymmetry, we propose a simplified version of APCFR+ called Simple APCFR+ (SAPCFR+), which uses a fixed asymmetry of step sizes to enable only a single-line modification compared to original PCFR+. Experimental results on five standard IIG benchmarks and two heads-up no-limit Texas Hold’em (HUNL) Subagems show that (i) both APCFR+ and SAPCFR+ outperform PCFR+ in most of the tested games, (ii) SAPCFR+ achieves a comparable empirical convergence rate with APCFR+, and (iii) our approach can be generalized to improve other CFR algorithms, e.g., Discount CFR (DCFR).
Linjian Meng, Tianpei Yang, Youzhi Zhang 0001, Zhenxing Ge, Yang Gao 0001
AAAI4
2025 Efficient Last-Iterate Convergence in Solving Extensive-Form Games
abstract
To establish last-iterate convergence for Counterfactual Regret Minimization (CFR) algorithms in learning a Nash equilibrium (NE) of extensive-form games (EFGs), recent studies reformulate learning an NE of the original EFG as learning the NEs of a sequence of (perturbed) regularized EFGs. Hence, proving last-iterate convergence in solving the original EFG reduces to proving last-iterate convergence in solving (perturbed) regularized EFGs. However, these studies only establish last-iterate convergence for Online Mirror Descent (OMD)-based CFR algorithms instead of Regret Matching (RM)-based CFR algorithms in solving perturbed regularized EFGs, resulting in a poor empirical convergence rate, as RM-based CFR algorithms typically outperform OMD-based CFR algorithms. In addition, as solving multiple perturbed regularized EFGs is required, fine-tuning across multiple perturbed regularized EFGs is infeasible, making parameter-free algorithms highly desirable. This paper show that CFR$^+$, a classical parameter-free RM-based CFR algorithm, achieves last-iterate convergence in learning an NE of perturbed regularized EFGs. This is the first parameter-free last-iterate convergence for RM-based CFR algorithms in perturbed regularized EFGs. Leveraging CFR$^+$ to solve perturbed regularized EFGs, we get Reward Transformation CFR$^+$ (RTCFR$^+$). Importantly, we extend prior work on the parameter-free property of CFR$^+$, enhancing its stability, which is vital for the empirical convergence of RTCFR$^+$. Experiments show that RTCFR$^+$ exhibits a significantly faster empirical convergence rate than existing algorithms that achieve theoretical last-iterate convergence. Interestingly, RTCFR$^+$ show performance no worse than average-iterate convergence CFR algorithms. It is the first last-iterate convergence algorithm to achieve such performance. Our code is available at https://github.com/menglinjian/NeurIPS-2025-RTCFR.
Linjian Meng, Tianpei Yang, Youzhi Zhang 0001, Zhenxing Ge, Shangdong Yang, Tianyu Ding, Wenbin Li 0006, Bo An 0001, Yang Gao 0001
NeurIPS4
2025 Last-Iterate Convergence of Smooth Regret Matching$^+$ Variants in Learning Nash Equilibria
abstract
Regret Matching$^+$ (RM$^+$) variants are widely used to build superhuman Poker AIs, yet few studies investigate their last-iterate convergence in learning a Nash equilibrium (NE). Although their last-iterate convergence is established for games satisfying the Minty Variational Inequality (MVI), no studies have demonstrated that these algorithms achieve such convergence in the broader class of games satisfying the weak MVI. A key challenge in proving last-iterate convergence for RM$^+$ variants in games satisfying the weak MVI is that even if the game's loss gradient satisfies the weak MVI, RM$^+$ variants operate on a transformed loss feedback which does not satisfy the weak MVI. To provide last-iterate convergence for RM$^+$ variants, we introduce a concise yet novel proof paradigm that involves: (i) transforming an RM$^+$ variant into an Online Mirror Descent (OMD) instance that updates within the original strategy space of the game to recover the weak MVI, and (ii) showing last-iterate convergence by proving the distance between accumulated regrets converges to zero via the recovered weak MVI of the feedback. Inspired by our proof paradigm, we propose Smooth Optimistic Gradient Based RM$^+$ (SOGRM$^+$) and show that it achieves last-iterate and finite-time best-iterate convergence in learning an NE of games satisfying the weak MVI, the weakest condition among all known RM$^+$ variants. Experiments show that SOGRM$^+$ significantly outperforms other algorithms. Our code is available at https://github.com/menglinjian/NeurIPS-2025-SOGRM.
Linjian Meng, Youzhi Zhang 0001, Zhenxing Ge, Tianyu Ding, Shangdong Yang, Wenbin Li 0006, Yang Gao 0001
NeurIPS3
2025 Improving Reward Models with Proximal Policy Exploration for Preference-Based Reinforcement Learning
abstract
Reinforcement learning (RL) heavily depends on well-designed reward functions, which are often biased and difficult to design for complex behaviors. Preference-based RL (PbRL) addresses this by learning reward models from human feedback, but its practicality is constrained by a critical dilemma: while existing methods reduce human effort through query optimization, they neglect the preference buffer's restricted coverage — a factor that fundamentally determines the reliability of reward model. We systematically demonstrate this limitation creates distributional mismatch: reward models trained on static buffers reliably assess in-distribution trajectories but falter with out-of-distribution (OOD) trajectories from policy exploration. Crucially, such failures in policy-proximal regions directly misguide iterative policy updates. To address this, we propose **Proximal Policy Exploration (PPE)** with two key components: (1) a *proximal-policy extension* method that expands exploration in undersampled policy-proximal regions, and (2) a *mixture distribution query* method that balances in-distribution and OOD trajectory sampling. By enhancing buffer coverage while preserving evaluation accuracy in policy-proximal regions, PPE enables more reliable policy updates. Experiments across continuous control tasks demonstrate that PPE enhances preference feedback utilization efficiency and RL sample efficiency over baselines, highlighting preference buffer coverage management's vital role in PbRL.
Jinyi Liu 0002, Pengjie Gu, Yifu Yuan, Zhenxing Ge, Wenya Wei, Yujing Hu, Bo An 0001
NeurIPS5
2024 Safe and Robust Subgame Exploitation in Imperfect Information Games
abstract
Opponent exploitation is an important task for players to exploit the weaknesses of others in games. Existing approaches mainly focus on balancing between exploitation and exploitability but are often vulnerable to modeling errors and deceptive adversaries. To address this problem, our paper offers a novel perspective on the safety of opponent exploitation, named Adaptation Safety. This concept leverages the insight that strategies, even those not explicitly aimed at opponent exploitation, may inherently be exploitable due to computational complexities, rendering traditional safety overly rigorous. In contrast, adaptation safety requires that the strategy should not be more exploitable than it would be in scenarios where opponent exploitation is not considered. Building on such adaptation safety, we further propose an Opponent eXploitation Search (OX-Search) framework by incorporating real-time search techniques for efficient online opponent exploitation. Moreover, we provide theoretical analyses to show the adaptation safety and robust exploitation of OX-Search, even with inaccurate opponent models. Empirical evaluations in popular poker games demonstrate OX-Search’s superiority in both exploitability and exploitation compared to previous methods.
Zhenxing Ge, Tianyu Ding, Linjian Meng, Bo An 0001, Wenbin Li 0006, Yang Gao 0001
ICML1
2024 Modeling Rationality: Toward Better Performance Against Unknown Agents in Sequential Games
abstract
Opponent modeling is necessary for autonomous agents to capture the intents of others during strategic interactions. Most previous works assume that they can access enough interaction history to build the model. However, it may not be realistic. To solve this problem, we present a novel rationality-consistent opponent modeling (ROM) method for games with imperfect information. In our approach, a game-theoretical concept of consistence about rationality is proposed to take advantage of the characteristic of imperfect information sequential games that rational behavior at disjoint information sets is correlated through anticipated opponent's behavior. With the correlation between different information sets, agents could infer the opponents' strategies at information sets correlated to observed behavior. To exploit the correlation, ROM attempts to conduct reasoning from the opponent's perspective and rationalize its past behavior. In this way, ROM acquires the ability to better adapt to different opponents and achieves a more accurate opponent model with insufficient observation history, which is verified by experiments in different settings. A heuristic adaptation approach is also applied in ROM, which updates the opponent model in an online manner and significantly reduces the computation cost. We evaluate ROM in both a grid world game and a poker game. Compared with other opponent modeling methods, ROM shows better performance and has more accurate predictions in both games against different types of opponents with limited action interactions. Experimental results also show that ROM's time cost is significantly reduced through heuristic adaptation.
Zhenxing Ge, Shangdong Yang, Pinzhuo Tian, Yang Gao 0001
IEEE Trans. Cybern.1
2023 An Efficient Deep Reinforcement Learning Algorithm for Solving Imperfect Information Extensive-Form Games
abstract
One of the most popular methods for learning Nash equilibrium (NE) in large-scale imperfect information extensive-form games (IIEFGs) is the neural variants of counterfactual regret minimization (CFR). CFR is a special case of Follow-The-Regularized-Leader (FTRL). At each iteration, the neural variants of CFR update the agent's strategy via the estimated counterfactual regrets. Then, they use neural networks to approximate the new strategy, which incurs an approximation error. These approximation errors will accumulate since the counterfactual regrets at iteration t are estimated using the agent's past approximated strategies. Such accumulated approximation error causes poor performance. To address this accumulated approximation error, we propose a novel FTRL algorithm called FTRL-ORW, which does not utilize the agent's past strategies to pick the next iteration strategy. More importantly, FTRL-ORW can update its strategy via the trajectories sampled from the game, which is suitable to solve large-scale IIEFGs since sampling multiple actions for each information set is too expensive in such games. However, it remains unclear which algorithm to use to compute the next iteration strategy for FTRL-ORW when only such sampled trajectories are revealed at iteration t. To address this problem and scale FTRL-ORW to large-scale games, we provide a model-free method called Deep FTRL-ORW, which computes the next iteration strategy using model-free Maximum Entropy Deep Reinforcement Learning. Experimental results on two-player zero-sum IIEFGs show that Deep FTRL-ORW significantly outperforms existing model-free neural methods and OS-MCCFR.
Linjian Meng, Zhenxing Ge, Pinzhuo Tian, Bo An 0001, Yang Gao 0001
AAAI2
2023 Efficient Subgame Refinement for Extensive-form Games
abstract
Subgame solving is an essential technique in addressing large imperfect information games, with various approaches developed to enhance the performance of refined strategies in the abstraction of the target subgame. However, directly applying existing subgame solving techniques may be difficult, due to the intricate nature and substantial size of many real-world games. To overcome this issue, recent subgame solving methods allow for subgame solving on limited knowledge order subgames, increasing their applicability in large games; yet this may still face obstacles due to extensive information set sizes. To address this challenge, we propose a generative subgame solving (GS2) framework, which utilizes a generation function to identify a subset of the earliest-reached nodes, reducing the size of the subgame. Our method is supported by a theoretical analysis and employs a diversity-based generation function to enhance safety. Experiments conducted on medium-sized games as well as the challenging large game of GuanDan demonstrate a significant improvement over the blueprint.
Zhenxing Ge, Tianyu Ding, Wenbin Li 0006, Yang Gao 0001
NeurIPS1