Anran Hu

dblp:255/7040 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 84% Optimization for machine learning · 13% Multi-agent systems · 3%
Human-computer interaction and pervasive computing
1 paper
Design research and methods · 100%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Design research and methods
personas
0.912025
Persona-L has Entered the Chat: Leveraging LLMs and Ability-based Framework for Personas of People with Complex Needs · CHI 2025
Machine learning › Reinforcement learning
continuous-time reinforcement learning
0.612022
Logarithmic Regret for Episodic Continuous-Time Linear-Quadratic Reinforcement Learning over a Finite-Time Horizon · J. Mach. Learn. Res. 2022
Machine learning › Optimization for machine learning
convergence analysis
0.612022
Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods · AAAI 2022
Machine learning › Reinforcement learning
episodic reinforcement learning
0.612022
Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods · AAAI 2022
Machine learning › Reinforcement learning › regret minimization
logarithmic regret
0.612022
Logarithmic Regret for Episodic Continuous-Time Linear-Quadratic Reinforcement Learning over a Finite-Time Horizon · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning › policy optimization
policy gradient
0.612022
Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods · AAAI 2022
Machine learning › Reinforcement learning
regret minimization
0.612022
Logarithmic Regret for Episodic Continuous-Time Linear-Quadratic Reinforcement Learning over a Finite-Time Horizon · J. Mach. Learn. Res. 2022
Machine learning › Reinforcement learning › multi-agent reinforcement learning › mean-field reinforcement learning
mean-field q-learning
0.412019
Learning Mean-Field Games · NeurIPS 2019
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.412019
Learning Mean-Field Games · NeurIPS 2019
Algorithmic game theory and mechanism design › non-cooperative game › dynamic games
mean field game
0.412019
Learning Mean-Field Games · NeurIPS 2019
Algorithmic game theory and mechanism design › equilibrium computation
nash equilibrium computation
0.412019
Learning Mean-Field Games · NeurIPS 2019
Design research and methods
persona generation
0.312025
Persona-L has Entered the Chat: Leveraging LLMs and Ability-based Framework for Personas of People with Complex Needs · CHI 2025

Methods — techniques the papers use, named apart from their topics

large language model · 0.9ability-based framework · 0.9q-learning · 0.8boltzmann policy · 0.8riccati differential equation · 0.6perturbation analysis · 0.6least-squares estimation · 0.6fictitious discount factor · 0.6discounted advantage estimation · 0.6
YearPublicationVenuePosition
2025 Persona-L has Entered the Chat: Leveraging LLMs and Ability-based Framework for Personas of People with Complex Needs
Lipeipei Sun, Tianzi Qin, Anran Hu, Shuojia Lin, Jianyan Chen, Mona Ali, Mirjana Prpa
CHI3
2022 Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods
abstract
When designing algorithms for finite-time-horizon episodic reinforcement learning problems, a common approach is to introduce a fictitious discount factor and use stationary policies for approximations. Empirically, it has been shown that the fictitious discount factor helps reduce variance, and stationary policies serve to save the per-iteration computational cost. Theoretically, however, there is no existing work on convergence analysis for algorithms with this fictitious discount recipe. This paper takes the first step towards analyzing these algorithms. It focuses on two vanilla policy gradient (VPG) variants: the first being a widely used variant with discounted advantage estimations (DAE), the second with an additional fictitious discount factor in the score functions of the policy gradient estimators. Non-asymptotic convergence guarantees are established for both algorithms, and the additional discount factor is shown to reduce the bias introduced in DAE and thus improve the algorithm convergence asymptotically. A key ingredient of our analysis is to connect three settings of Markov decision processes (MDPs): the finite-time-horizon, the average reward and the discounted settings. To our best knowledge, this is the first theoretical guarantee on fictitious discount algorithms for the episodic reinforcement learning of finite-time-horizon MDPs, which also leads to the (first) global convergence of policy gradient methods for finite-time-horizon episodic reinforcement learning.
Xin Guo 0001, Anran Hu, Junzi Zhang
AAAI2
2022 Logarithmic Regret for Episodic Continuous-Time Linear-Quadratic Reinforcement Learning over a Finite-Time Horizon
abstract
We study finite-time horizon continuous-time linear-quadratic reinforcement learning problems in an episodic setting, where both the state and control coefficients are unknown to the controller. We first propose a least-squares algorithm based on continuous-time observations and controls, and establish a logarithmic regret bound of magnitude $\mathcal{O}((\ln M)(\ln\ln M) )$, with $M$ being the number of learning episodes. The analysis consists of two components: perturbation analysis, which exploits the regularity and robustness of the associated Riccati differential equation; and parameter estimation error, which relies on sub-exponential properties of continuous-time least-squares estimators. We further propose a practically implementable least-squares algorithm based on discrete-time observations and piecewise constant controls, which achieves similar logarithmic regret with an additional term depending explicitly on the time stepsizes used in the algorithm.
Matteo Basei, Xin Guo 0001, Anran Hu, Yufei Zhang 0001
J. Mach. Learn. Res.3
2019 Learning Mean-Field Games
abstract
This paper presents a general mean-field game (GMFG) framework for simultaneous learning and decision-making in stochastic games with a large population. It first establishes the existence of a unique Nash Equilibrium to this GMFG, and explains that naively combining Q-learning with the fixed-point approach in classical MFGs yields unstable algorithms. It then proposes a Q-learning algorithm with Boltzmann policy (GMF-Q), with analysis of convergence property and computational complexity. The experiments on repeated Ad auction problems demonstrate that this GMF-Q algorithm is efficient and robust in terms of convergence and learning accuracy. Moreover, its performance is superior in convergence, stability, and learning ability, when compared with existing algorithms for multi-agent reinforcement learning.
Xin Guo 0001, Anran Hu, Renyuan Xu, Junzi Zhang
NeurIPS2