Haifeng Zhang 0002

dblp:93/7133-2 · also Hai-Feng Zhang 0002 · DBLP profile ↗
← Back
28ranked-venue papers
3as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Think, Speak, Decide: Language-Augmented Multi-Agent Reinforcement Learning for Economic Decision-Making
abstract
Economic decision‑making depends not only on structured signals—such as prices and taxes—but also on unstructured language, including peer dialogue and media narratives. While multi‑agent reinforcement learning (MARL) has shown promise in optimizing economic decisions, it struggles with the semantic ambiguity and contextual richness of language. We propose LAMP (Language‑Augmented Multi‑Agent Policy), the first framework to integrate language into economic decision‑making, narrowing the gap to real‑world settings. LAMP follows a Think–Speak–Decide pipeline: (1) Think interprets numerical observations to extract short‑term shocks and long‑term trends, caching high‑value reasoning trajectories. (2) Speak crafts and exchanges strategic messages based on the reasoning, updating beliefs by parsing peer communications. (3) Decide fuses numerical data, reasoning, and reflections into a MARL policy to optimize language‑augmented decision‑making. Experiments in economic simulation show that LAMP outperforms both MARL and LLM‑only baselines in cumulative return (+63.5%, +34.0%), robustness (+18.8%, +59.4%), and interpretability. These results demonstrate the potential of language‑augmented policies to deliver more effective and robust economic strategies.
Heyang Ma, Qirui Mi, Qipeng Yang, Zijun Fan, Haifeng Zhang 0002
AAAI6
2026 Proactive Constrained Policy Optimization with Preemptive Penalty
abstract
Safe Reinforcement Learning (RL) often faces significant issues such as constraint violations and instability, necessitating the use of constrained policy optimization, which seeks optimal policies while ensuring adherence to specific constraints like safety. Typically, constrained optimization problems are addressed by the Lagrangian method, a post-violation remedial approach that may result in oscillations and overshoots. Motivated by this, we propose a novel method named Proactive Constrained Policy Optimization (PCPO) that incorporates a preemptive penalty mechanism. This mechanism integrates barrier items into the objective function as the policy nears the boundary, imposing a cost. Meanwhile, we introduce a constraint-aware intrinsic reward to guide boundary-aware exploration, which is activated only when the policy approaches the constraint boundary. We establish theoretical upper and lower bounds for the duality gap and the performance of the PCPO update, shedding light on the method's convergence characteristics. Additionally, to enhance the optimization performance, we adopt a policy iteration approach. An interesting finding is that PCPO demonstrates significant stability in experiments. Experimental results indicate that the PCPO framework provides a robust solution for policy optimization under constraints, with important implications for future research and practical applications.
Ning Yang 0005, Haifeng Zhang 0002, Jun Wang 0012
AAAI4
2026 Dual Activation-Weight Sparsity: A Training-Free Framework for Efficient Large Language Model Compression
abstract
Luoyang Sun, Guangyan Li, Cheng Deng, Haifeng Zhang, Jian Zhao, Yongqiang Tang, Wensheng Zhang, Jun Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Luoyang Sun, Guangyan Li, Cheng Deng 0001, Haifeng Zhang 0002, Jian Zhao 0006, Yongqiang Tang, Wensheng Zhang 0002, Jun Wang 0012
ACL (1)4
2025 PIE: Permutation-Invariant Multi-Entity Evaluation
abstract
The evaluation system is devised in online competitive games and sports to assess players' skills. Popular methods Elo, designed for two-player competitive games such as chess and tennis, are based on the Bradley-Terry model and update players' ratings with competition outcomes. Extended methods Trueskill and mElo are proposed for multi-player(team) and two-player intransitive games, respectively. However, existing evaluation methods are constrained in specific situations; for example, mElo is limited to dealing with two-player games, and TrueSkill, Elo can not handle intransitive games. In addition, previous team evaluation methods bake in the assumption that individuals' contributions to team performance are uniformly determined by individual abilities, which does not hold in many situations, such as football with different roles and games with team score being defined as the maximum or minimum of players' gains. In this paper, we address the challenge of evaluating player skill in multi-player (team) competitions. We propose PIE, an online permutation-invariant evaluation model for multi-entity competitions that ensures the predicted winner of a multi-entity match remains invariant to the order of input entities. For multi-team evaluation, PIE enables team ratings to increase monotonically with improvements in individual player ratings. Empirical results of predicting the winner and winning probabilities in real-world games demonstrate that PIE achieves comparable performance in handling the prediction of multiplayer(team) matches with other baselines.
Haifeng Zhang 0002, Yali Du 0001, Jun Wang 0012
CoG2
2025 Learning Macroeconomic Policies Through Dynamic Stackelberg Mean-Field Games
abstract
Macroeconomic outcomes emerge from individuals’ decisions, making it essential to model how agents interact with macro policy via consumption, investment, and labor choices. We formulate this as a dynamic Stackelberg game: the government (leader) sets policies, and agents (followers) respond by optimizing their behavior over time. Unlike static models, this dynamic formulation captures temporal dependencies and strategic feedback critical to policy design. However, as the number of agents increases, explicitly simulating all agent–agent and agent–government interactions becomes computationally infeasible. To address this, we propose the Dynamic Stackelberg Mean Field Game (DSMFG) framework, which approximates these complex interactions via agent–population and government–population couplings. This approximation preserves individual-level feedback while ensuring scalability, enabling DSMFG to jointly model three core features of real-world policy-making: dynamic feedback, asymmetry, and large-scale. We further introduce Stackelberg Mean Field Reinforcement Learning (SMFRL), a data-driven algorithm that learns the leader’s optimal policies while maintaining personalized responses for individual agents. Empirically, we validate our approach in a large-scale simulated economy, where it scales to 1,000 agents (vs. 100 in prior work) and achieves a 4× GDP gain over classical economic methods and a 19× improvement over the static 2022 U.S. federal income tax policy.
Qirui Mi, Chengdong Ma, Si-Yu Xia, Yan Song 0003, Mengyue Yang, Jun Wang 0012, Haifeng Zhang 0002
ECAI8
2025 Mean Field Correlated Imitation Learning
Chengdong Ma, Qirui Mi, Ning Yang 0005, Mengyue Yang, Haifeng Zhang 0002, Jun Wang 0012, Yaodong Yang 0001
AAMAS7
2025 EconGym: A Scalable AI Testbed with Diverse Economic Tasks
abstract
Artificial intelligence (AI) has become a powerful tool for economic research, enabling large-scale simulation and policy optimization. However, applying AI effectively requires simulation platforms for scalable training and evaluation—yet existing environments remain limited to simplified, narrowly scoped tasks, falling short of capturing complex economic challenges such as demographic shifts, multi-government coordination, and large-scale agent interactions.To address this gap, we introduce EconGym, a scalable and modular testbed that connects diverse economic tasks with AI algorithms. Grounded in rigorous economic modeling, EconGym implements 11 heterogeneous role types (e.g., households, firms, banks, governments), their interaction mechanisms, and agent models with well-defined observations, actions, and rewards. Users can flexibly compose economic roles with diverse agent algorithms to simulate rich multi-agent trajectories across 25+ economic tasks for AI-driven policy learning and analysis.Experiments show that EconGym supports diverse and cross-domain tasks—such as coordinating fiscal, pension, and monetary policies—and enables benchmarking across AI, economic methods, and hybrids. Results indicate that richer task composition and algorithm diversity expand the policy space, while AI agents guided by classical economic methods perform best in complex settings. EconGym also scales to 100k agents with high realism and efficiency.
Qirui Mi, Qipeng Yang, Zijun Fan, Wentian Fan, Heyang Ma, Chengdong Ma, Si-Yu Xia, Bo An 0001, Jun Wang 0012, Haifeng Zhang 0002
NeurIPS10
2025 MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework
abstract
Simulating collective decision-making involves more than aggregating individual behaviors; it emerges from dynamic interactions among individuals. While large language models (LLMs) offer strong potential for social simulation, achieving quantitative alignment with real-world data remains a key challenge. To bridge this gap, we propose the \textbf{M}ean-\textbf{F}ield \textbf{LLM} (\textbf{MF-LLM}) framework, the first to incorporate mean field theory into LLM-based social simulation. MF-LLM models bidirectional interactions between individuals and the population through an iterative process, generating population signals to guide individual decisions, which in turn update the signals. This interplay produces coherent trajectories of collective behavior. To improve alignment with real-world data, we introduce \textbf{IB-Tune}, a novel fine-tuning method inspired by the \textbf{I}nformation \textbf{B}ottleneck principle, which retains population signals most predictive of future actions while filtering redundant history. Evaluated on a real-world social dataset, MF-LLM reduces KL divergence to human population distributions by \textbf{47\%} compared to non-mean-field baselines, enabling accurate trend forecasting and effective intervention planning. Generalizing across 7 domains and 4 LLM backbones, MF-LLM provides a scalable, high-fidelity foundation for social simulation.
Qirui Mi, Mengyue Yang, Xiangning Yu 0001, Cheng Deng 0001, Bo An 0001, Haifeng Zhang 0002, Xu Chen 0017, Jun Wang 0012
NeurIPS7
2025 Self-Verifying Reflection Helps Transformers with CoT Reasoning
abstract
Advanced large language models (LLMs) frequently reflect in reasoning chain-of-thoughts (CoTs), where they self-verify the correctness of current solutions and explore alternatives. However, given recent findings that LLMs detect limited errors in CoTs, how reflection contributes to empirical improvements remains unclear. To analyze this issue, in this paper, we present a minimalistic reasoning framework to support basic self-verifying reflection for small transformers without natural language, which ensures analytic clarity and reduces the cost of comprehensive experiments. Theoretically, we prove that self-verifying reflection guarantees improvements if verification errors are properly bounded. Experimentally, we show that tiny transformers, with only a few million parameters, benefit from self-verification in both training and reflective execution, reaching remarkable LLM-level performance in integer multiplication and Sudoku. Similar to LLM results, we find that reinforcement learning (RL) improves in-distribution performance and incentivizes frequent reflection for tiny transformers, yet RL mainly optimizes shallow statistical patterns without faithfully reducing verification errors. In conclusion, integrating generative transformers with discriminative verification inherently facilitates CoT reasoning, regardless of scaling and natural language.
Zhongwei Yu, Wannian Xia, Bo Xu 0002, Haifeng Zhang 0002, Yali Du 0001, Jun Wang 0012
NeurIPS5
2025 Curious Causality-Seeking Agents in Open-ended Worlds
abstract
When building a world model, a common assumption is that the environment has a single, unchanging underlying causal rule, like applying Newton's laws to every situation. However, in truly open-ended environments, the apparent causal mechanism may drift over time because the agent continually encounters novel contexts and operates within a limited observational window. This brings about a problem that, when building a world model, even subtle shifts in policy or environment states can alter the very observed causal mechanisms. In this work, we introduce the Meta-Causal Graph as world models for open-ended environments, a minimal unified representation that efficiently encodes the transformation rules governing how causal structures shift across different latent world states. A single Meta-Causal Graph is composed of multiple causal subgraphs, each triggered by meta state, which is in the latent state space. Building on this representation, we introduce a Causality-Seeking Agent whose objectives are to (1) identify the meta states that trigger each subgraph, (2) discover the corresponding causal relationships by agent curiosity-driven intervention policy, and (3) iteratively refine the Meta-Causal Graph through ongoing curiosity-driven exploration and agent experiences. Experiments on both synthetic tasks and a challenging robot arm manipulation task demonstrate that our method robustly captures shifts in causal dynamics and generalizes effectively to previously unseen contexts.
Haoxuan Li 0001, Haifeng Zhang 0002, Jun Wang 0012, Francesco Faccio, Jürgen Schmidhuber, Mengyue Yang
NeurIPS3
2024 Adaptive Command : Real-Time Policy Adjustment via Language Models in StarCraft II
abstract
We present Adaptive Command, a novel framework integrating large language models (LLMs) with behavior trees for real-time strategic decision-making in StarCraft II.Our system focuses on enhancing human-AI collaboration in complex, dynamic environments through natural language interactions.The framework comprises: (1) an LLM-based strategic advisor, (2) a behavior tree for action execution, and (3) a natural language interface with speech capabilities.User studies demonstrate significant improvements in player decision-making and strategic adaptability, particularly benefiting novice players and those with disabilities.This work contributes to the field of real-time human-AI collaborative decisionmaking, offering insights applicable beyond RTS games to various complex decision-making scenarios.
Weiyu Ma, Shu Lin 0003, Haifeng Zhang 0002, Jun Wang 0012
DAI4
2024 Variational Stochastic Games
abstract
The Control as Inference (CAI) framework has successfully transformed single-agent reinforcement learning (RL) by reframing control tasks as probabilistic inference problems.However, the extension of CAI to multi-agent, general-sum stochastic games (SGs) remains underexplored, particularly in decentralized settings where agents operate independently without centralized coordination.In this paper, we propose a novel variational inference framework tailored to decentralized multi-agent systems.Our framework addresses the challenges posed by non-stationarity and unaligned agent objectives, proving that the resulting policies form an 𝜖-Nash equilibrium.Additionally, we demonstrate theoretical convergence guarantees for the proposed decentralized algorithms.Leveraging this framework, we instantiate multiple algorithms to solve for Nash equilibrium, mean-field Nash equilibrium, and correlated equilibrium, with rigorous theoretical convergence analysis.
Haifeng Zhang 0002
DAI2
2024 Token-level Direct Preference Optimization
abstract
Fine-tuning pre-trained Large Language Models (LLMs) is essential to align them with human values and intentions. This process often utilizes methods like pairwise comparisons and KL divergence against a reference LLM, focusing on the evaluation of full answers generated by the models. However, the generation of these responses occurs in a token level, following a sequential, auto-regressive fashion. In this paper, we introduce Token-level Direct Preference Optimization (TDPO), a novel approach to align LLMs with human preferences by optimizing policy at the token level. Unlike previous methods, which face challenges in divergence efficiency, TDPO integrates forward KL divergence constraints for each token, improving alignment and diversity. Utilizing the Bradley-Terry model for a token-based reward system, our method enhances the regulation of KL divergence, while preserving simplicity without the need for explicit reward modeling. Experimental results across various text tasks demonstrate TDPO’s superior performance in balancing alignment with generation diversity. Notably, fine-tuning with TDPO strikes a better balance than DPO in the controlled sentiment generation and single-turn dialogue datasets, and significantly improves the quality of generated responses compared to both DPO and PPO-based RLHF methods.
Yongcheng Zeng, Weiyu Ma, Ning Yang 0005, Haifeng Zhang 0002, Jun Wang 0012
ICML5
2024 Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach
abstract
With the continued advancement of Large Language Models (LLMs) Agents in reasoning, planning, and decision-making, benchmarks have become crucial in evaluating these skills. However, there is a notable gap in benchmarks for real-time strategic decision-making. StarCraft II (SC2), with its complex and dynamic nature, serves as an ideal setting for such evaluations. To this end, we have developed TextStarCraft II, a specialized environment for assessing LLMs in real-time strategic scenarios within SC2. Addressing the limitations of traditional Chain of Thought (CoT) methods, we introduce the Chain of Summarization (CoS) method, enhancing LLMs' capabilities in rapid and effective decision-making. Our key experiments included: 1. LLM Evaluation: Tested 10 LLMs in TextStarCraft II, most of them defeating LV5 build-in AI, showcasing effective strategy skills. 2. Commercial Model Knowledge: Evaluated four commercial models on SC2 knowledge; GPT-4 ranked highest by Grandmaster-level experts. 3. Human-AI Matches: Experimental results showed that fine-tuned LLMs performed on par with Gold-level players in real-time matches, demonstrating comparable strategic abilities. All code and data from this study have been made pulicly available at https://github.com/histmeisah/Large-Language-Models-play-StarCraftII
Weiyu Ma, Qirui Mi, Yongcheng Zeng, Runji Lin, Yuqiao Wu, Jun Wang 0012, Haifeng Zhang 0002
NeurIPS8
2023 An Efficient End-to-End Training Approach for Zero-Shot Human-AI Coordination
abstract
The goal of zero-shot human-AI coordination is to develop an agent that can collaborate with humans without relying on human data. Prevailing two-stage population-based methods require a diverse population of mutually distinct policies to simulate diverse human behaviors. The necessity of such populations severely limits their computational efficiency. To address this issue, we propose E3T, an **E**fficient **E**nd-to-**E**nd **T**raining approach for zero-shot human-AI coordination. E3T employs a mixture of ego policy and random policy to construct the partner policy, making it both coordination-skilled and diverse. In this way, the ego agent is end-to-end trained with this mixture policy without the need of a pre-trained population, thus significantly improving the training efficiency. In addition, a partner modeling module is proposed to predict the partner's action from historical information. With the predicted partner's action, the ego policy is able to adapt its policy and take actions accordingly when collaborating with humans of different behavior patterns. Empirical results on the Overcooked environment show that our method significantly improves the training efficiency while preserving comparable or superior performance than the population-based baselines. Demo videos are available at https://sites.google.com/view/e3t-overcooked.
Jiaxian Guo, Xingzhou Lou, Jun Wang 0012, Haifeng Zhang 0002, Yali Du 0001
NeurIPS5
2023 Large sequence models for sequential decision-making: a survey
Muning Wen, Runji Lin, Hanjing Wang, Yaodong Yang 0001, Ying Wen 0001, Luo Mai, Jun Wang 0012, Haifeng Zhang 0002, Weinan Zhang 0001
Frontiers Comput. Sci.8
2022 Learning to Identify Top Elo Ratings: A Dueling Bandits Approach
abstract
The Elo rating system is widely adopted to evaluate the skills of (chess) game and sports players. Recently it has been also integrated into machine learning algorithms in evaluating the performance of computerised AI agents. However, an accurate estimation of the Elo rating (for the top players) often requires many rounds of competitions, which can be expensive to carry out. In this paper, to minimize the number of comparisons and to improve the sample efficiency of the Elo evaluation (for top players), we propose an efficient online match scheduling algorithm. Specifically, we identify and match the top players through a dueling bandits framework and tailor the bandit algorithm to the gradient-based update of Elo. We show that it reduces the per-step memory and time complexity to constant, compared to the traditional likelihood maximization approaches requiring O(t) time. Our algorithm has a regret guarantee that is sublinear in the number of competition rounds and has been extended to the multidimensional Elo ratings for handling intransitive games. We empirically demonstrate that our method achieves superior convergence speed and time efficiency on a variety of gaming tasks.
Yali Du 0001, Binxin Ru, Jun Wang 0012, Haifeng Zhang 0002, Xu Chen 0017
AAAI5
2022 A Theoretical Understanding of Gradient Bias in Meta-Reinforcement Learning
abstract
Gradient-based Meta-RL (GMRL) refers to methods that maintain two-level optimisation procedures wherein the outer-loop meta-learner guides the inner-loop gradient-based reinforcement learner to achieve fast adaptations. In this paper, we develop a unified framework that describes variations of GMRL algorithms and points out that existing stochastic meta-gradient estimators adopted by GMRL are actually \textbf{biased}. Such meta-gradient bias comes from two sources: 1) the compositional bias incurred by the two-level problem structure, which has an upper bound of $\mathcal{O}\big(K\alpha^{K}\hat{\sigma}_{\text{In}}|\tau|^{-0.5}\big)$ \emph{w.r.t.} inner-loop update step $K$, learning rate $\alpha$, estimate variance $\hat{\sigma}^{2}_{\text{In}}$ and sample size $|\tau|$, and 2) the multi-step Hessian estimation bias $\hat{\Delta}_{H}$ due to the use of autodiff, which has a polynomial impact $\mathcal{O}\big((K-1)(\hat{\Delta}_{H})^{K-1}\big)$ on the meta-gradient bias. We study tabular MDPs empirically and offer quantitative evidence that testifies our theoretical findings on existing stochastic meta-gradient estimators. Furthermore, we conduct experiments on Iterated Prisoner's Dilemma and Atari games to show how other methods such as off-policy learning and low-bias estimator can help fix the gradient bias for GMRL algorithms in general.
Bo Liu 0039, Xidong Feng, Luo Mai, Haifeng Zhang 0002, Jun Wang 0012, Yaodong Yang 0001
NeurIPS6
2021 Signal Instructed Coordination in Cooperative Multi-agent Reinforcement Learning
Hongyi Guo, Yali Du 0001, Fei Fang 0001, Haifeng Zhang 0002, Weinan Zhang 0001, Yong Yu 0001
DAI5
2021 Joint Caching and Transmission in the Mobile Edge Network: An Multi-Agent Learning Approach
abstract
Joint caching and transmission optimization problem is challenging due to the deep coupling between decisions. This paper proposes an iterative distributed multi-agent learning approach to jointly optimize caching and transmission. The goal of this approach is to minimize the total transmission delay of all users. In this iterative approach, each iteration includes caching optimization and transmission optimization. A multi-agent reinforcement learning (MARL)-based caching network is developed to cache popular tasks, such as answering which files to evict from the cache and which files to storage. Based on the cached files of the caching network, the transmission network transmits cached files for users by single transmission (ST) or joint transmission (JT) with multi-agent Bayesian learning automaton (MABLA) method. And then users access the edge servers with the minimum transmission delay. The experimental results demonstrate the performance of the proposed multi-agent learning approach.
Qirui Mi, Ning Yang 0005, Haifeng Zhang 0002, Haijun Zhang 0001, Jun Wang 0012
GLOBECOM3
2021 Estimating α-Rank from A Few Entries with Low Rank Matrix Completion
abstract
Multi-agent evaluation aims at the assessment of an agent’s strategy on the basis of interaction with others. Typically, existing methods such as $\alpha$-rank and its approximation still require to exhaustively compare all pairs of joint strategies for an accurate ranking, which in practice is computationally expensive. In this paper, we aim to reduce the number of pairwise comparisons in recovering a satisfying ranking for $n$ strategies in two-player meta-games, by exploring the fact that agents with similar skills may achieve similar payoffs against others. Two situations are considered: the first one is when we can obtain the true payoffs; the other one is when we can only access noisy payoff. Based on these formulations, we leverage low-rank matrix completion and design two novel algorithms for noise-free and noisy evaluations respectively. For both of these settings, we theorize that $O(nr \log n)$ ($n$ is the number of agents and $r$ is the rank of the payoff matrix) payoff entries are required to achieve sufficiently well strategy evaluation performance. Empirical results on evaluating the strategies in three synthetic games and twelve real world games demonstrate that strategy evaluation from a few entries can lead to comparable performance to algorithms with full knowledge of the payoff matrix.
Yali Du 0001, Xu Chen 0017, Jun Wang 0012, Haifeng Zhang 0002
ICML5
2021 Settling the Variance of Multi-Agent Policy Gradients
abstract
Policy gradient (PG) methods are popular reinforcement learning (RL) methods where a baseline is often applied to reduce the variance of gradient estimates. In multi-agent RL (MARL), although the PG theorem can be naturally extended, the effectiveness of multi-agent PG (MAPG) methods degrades as the variance of gradient estimates increases rapidly with the number of agents. In this paper, we offer a rigorous analysis of MAPG methods by, firstly, quantifying the contributions of the number of agents and agents' explorations to the variance of MAPG estimators. Based on this analysis, we derive the optimal baseline (OB) that achieves the minimal variance. In comparison to the OB, we measure the excess variance of existing MARL algorithms such as vanilla MAPG and COMA. Considering using deep neural networks, we also propose a surrogate version of OB, which can be seamlessly plugged into any existing PG methods in MARL. On benchmarks of Multi-Agent MuJoCo and StarCraft challenges, our OB technique effectively stabilises training and improves the performance of multi-agent PPO and COMA algorithms by a significant margin. Code is released at \url{https://github.com/morning9393/Optimal-Baseline-for-Multi-agent-Policy-Gradients}.
Jakub Grudzien Kuba, Muning Wen, Linghui Meng 0001, Shangding Gu, Haifeng Zhang 0002, David Mguni, Jun Wang 0012, Yaodong Yang 0001
NeurIPS5
2020 Bi-Level Actor-Critic for Multi-Agent Coordination
abstract
Coordination is one of the essential problems in multi-agent systems. Typically multi-agent reinforcement learning (MARL) methods treat agents equally and the goal is to solve the Markov game to an arbitrary Nash equilibrium (NE) when multiple equilibra exist, thus lacking a solution for NE selection. In this paper, we treat agents unequally and consider Stackelberg equilibrium as a potentially better convergence point than Nash equilibrium in terms of Pareto superiority, especially in cooperative environments. Under Markov games, we formally define the bi-level reinforcement learning problem in finding Stackelberg equilibrium. We propose a novel bi-level actor-critic learning method that allows agents to have different knowledge base (thus intelligent), while their actions still can be executed simultaneously and distributedly. The convergence proof is given, while the resulting learning algorithm is tested against the state of the arts. We found that the proposed bi-level actor-critic algorithm successfully converged to the Stackelberg equilibria in matrix games and find a asymmetric solution in a highway merge environment.
Haifeng Zhang 0002, Weizhe Chen 0001, Zeren Huang, Minne Li, Yaodong Yang 0001, Weinan Zhang 0001, Jun Wang 0012
AAAI1
2018 Learning to Design Games: Strategic Environments in Reinforcement Learning
abstract
In typical reinforcement learning (RL), the environment is assumed given and the goal of the learning is to identify an optimal policy for the agent taking actions through its interactions with the environment. In this paper, we extend this setting by considering the environment is not given, but controllable and learnable through its interaction with the agent at the same time. This extension is motivated by environment design scenarios in the real-world, including game design, shopping space design and traffic signal design. Theoretically, we find a dual Markov decision process (MDP) w.r.t. the environment to that w.r.t. the agent, and derive a policy gradient solution to optimizing the parametrized environment. Furthermore, discontinuous environments are addressed by a proposed general generative framework. Our experiments on a Maze game design task show the effectiveness of the proposed algorithms in generating diverse and challenging Mazes against various agent settings.
Haifeng Zhang 0002, Jun Wang 0012, Zhiming Zhou 0001, Weinan Zhang 0001, Yong Yu 0001, Wenxin Li 0005
IJCAI1
2018 Botzone: an online multi-agent competitive platform for AI education
abstract
This paper presents Botzone, a competitive platform for game AI education and research. It aims to simplify the teaching process of game AI courses, inspire learners to self-study, and acting as a dataset for game AI research. This platform is a universal online multi-agent game AI platform, designed to evaluate different implementations of game AI by applying them to agents in a variety of games and compete with each other, featuring an ELO ranking system and a contest system for users to evaluate their AI programs. It has been successfully used in various AI competitions and courses in practice, and has the extensibility to support more games and languages, as well as further usages such as studying machine learning on game AI. In this paper, we firstly describe the structure and features of Botzone, then focus on our experience in utilizing Botzone for a programming course.
Haoyu Zhou, Haifeng Zhang 0002, Xinchao Wang, Wenxin Li 0005
ITiCSE2
2017 ICFVR 2017: 3rd international competition on finger vein recognition
abstract
In recent years, finger vein recognition has become an important sub-field in biometrics and been applied to real-world applications. The development of finger vein recognition algorithms heavily depends on large-scale real-world data sets. In order to motivate research on finger vein recognition, we released the largest finger vein data set up to now and hold finger vein recognition competitions based on our data set every year. In 2017, International Competition on Finger Vein Recognition (ICFVR) is held jointly with IJCB 2017. 11 teams registered and 10 of them joined the final evaluation. The winner of this year dramatically improved the EER from 2.64% to 0.483% compared to the 'winner of last year. In this paper, we introduce the process and results of ICFVR 2017 and give insights on development of state-of-art finger vein recognition algorithms.
Houjun Huang, Haifeng Zhang 0002, Liao Ni, Nasir Uddin Ahmed, Md. Shakil Ahmed, Yilun Jin, Jingxuan Wen, Wenxin Li 0005
IJCB3
2017 Managing Risk of Bidding in Display Advertising
abstract
In this paper, we deal with the uncertainty of bidding for display advertising. Similar to the financial market trading, real-time bidding (RTB) based display advertising employs an auction mechanism to automate the impression level media buying; and running a campaign is no different than an investment of acquiring new customers in return for obtaining additional converted sales. Thus, how to optimally bid on an ad impression to drive the profit and return-on-investment becomes essential. However, the large randomness of the user behaviors and the cost uncertainty caused by the auction competition may result in a significant risk from the campaign performance estimation. In this paper, we explicitly model the uncertainty of user click-through rate estimation and auction competition to capture the risk. We borrow an idea from finance and derive the value at risk for each ad display opportunity. Our formulation results in two risk-aware bidding strategies that penalize risky ad impressions and focus more on the ones with higher expected return and lower risk. The empirical study on real-world data demonstrates the effectiveness of our proposed risk-aware bidding strategies: yielding profit gains of 15.4% in offline experiments and up to 17.5% in an online A/B test on a commercial RTB platform over the widely applied bidding strategies.
Haifeng Zhang 0002, Weinan Zhang 0001, Yifei Rong, Kan Ren, Wenxin Li 0005, Jun Wang 0012
WSDM1
2016 User Response Learning for Directly Optimizing Campaign Performance in Display Advertising
abstract
Learning and predicting user responses, such as clicks and conversions, are crucial for many Internet-based businesses including web search, e-commerce, and online advertising. Typically, a user response model is established by optimizing the prediction accuracy, e.g., minimizing the error between the prediction and the ground truth user response. However, in many practical cases, predicting user responses is only part of a rather larger predictive or optimization task, where on one hand, the accuracy of a user response prediction determines the final (expected) utility to be optimized, but on the other hand, its learning may also be influenced from the follow-up stochastic process. It is, thus, of great interest to optimize the entire process as a whole rather than treat them independently or sequentially. In this paper, we take real-time display advertising as an example, where the predicted user's ad click-through rate (CTR) is employed to calculate a bid for an ad impression in the second price auction. We reformulate a common logistic regression CTR model by putting it back into its subsequent bidding context: rather than minimizing the prediction error, the model parameters are learned directly by optimizing campaign profit. The gradient update resulted from our formulations naturally fine-tunes the cases where the market competition is high, leading to a more cost-effective bidding. Our experiments demonstrate that, while maintaining comparable CTR prediction accuracy, our proposed user response learning leads to campaign profit gains as much as 78.2% for offline test and 25.5% for online A/B test over strong baselines.
Kan Ren, Weinan Zhang 0001, Yifei Rong, Haifeng Zhang 0002, Yong Yu 0001, Jun Wang 0012
CIKM4