Bohao Qu

dblp:275/7652 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-3192-8736ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 57% Transfer learning and domain adaptation · 18% Representation and self-supervised learning · 8%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems › agentic AI
agentic reasoning
1.012026
Learn Like Humans: Use Meta-cognitive Reflection for Efficient Self-Improvement · ACL (1) 2026
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning
1.012026
Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New Thinking · AAAI 2026
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.012026
Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New Thinking · AAAI 2026
Natural language and speech › Language models and text generation
self-improvement
1.012026
Learn Like Humans: Use Meta-cognitive Reflection for Efficient Self-Improvement · ACL (1) 2026
Machine learning › Reinforcement learning
self-improving agent
1.012026
Learn Like Humans: Use Meta-cognitive Reflection for Efficient Self-Improvement · ACL (1) 2026
Machine learning › Reinforcement learning
exploration
0.812024
Sample Efficient Offline-to-Online Reinforcement Learning · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Reinforcement learning › imitation learning › inverse reinforcement learning
inverse optimal control
0.812024
Transductive Reward Inference on Graph · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Transfer learning and domain adaptation
meta-learning
0.812024
Sample Efficient Offline-to-Online Reinforcement Learning · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Transfer learning and domain adaptation › meta-learning
meta-learning for adaptation
0.812024
Sample Efficient Offline-to-Online Reinforcement Learning · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Transductive Reward Inference on Graph · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning
0.812024
Sample Efficient Offline-to-Online Reinforcement Learning · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Reinforcement learning › exploration
optimistic exploration
0.812024
Sample Efficient Offline-to-Online Reinforcement Learning · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Reinforcement learning
policy diversity
0.812024
Diversifying Policies With Non-Markov Dispersion to Expand the Solution Space · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
policy representation learning
0.812024
Diversifying Policies With Non-Markov Dispersion to Expand the Solution Space · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Transfer learning and domain adaptation
sample-efficient adaptation
0.812024
Sample Efficient Offline-to-Online Reinforcement Learning · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.212024
Sample Efficient Offline-to-Online Reinforcement Learning · IEEE Trans. Knowl. Data Eng. 2024

Methods — techniques the papers use, named apart from their topics

prompt optimization · 1.0procedural reflection · 1.0principle-based reflection · 1.0poincaré ball · 1.0hyperbolic neural network · 1.0contrastive loss · 1.0transformer · 0.8policy embedding · 0.8meta-learning · 0.8dispersion matrix · 0.8
YearPublicationVenuePosition
2026 Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New Thinking
abstract
In this paper, we rethink model agent behaviors from a geometric structure perspective in multi-agent reinforcement learning. Modeling agent behaviors is essential for understanding how agents interact and facilitating effective decisions. The key lies in capturing the dependencies and sequential relationships among agent decisions. Since each decision influences the subsequent choices, this forms a hierarchical and nested tree-like structure of interdependencies. While modeling tree-like data in Euclidean spaces could cause distortion, which results in a loss of agent decision structure information. Motivated by this, we reconsider model agent behaviors in hyperbolic space and propose the Hyperbolic Multi-Agent Representations (HMAR) method, which projects the agent behaviors into a Poincaré ball and leverages hyperbolic neural networks to learn agent policy representations. Additionally, we designed a contrastive loss function to train this network, minimizing the distance in feature space between different representations of the same agent while maximizing the distance between representations of distinct agents. Experimental results provide empirical evidence for the effectiveness of the HMAR method in cooperative and competitive environments, demonstrating the potential of hyperbolic agent representations for effective decision-making in multi-agent environments.
Bohao Qu, Xiaofeng Cao 0002, Bing Li 0001, Menglin Zhang, Tuan-Anh Vu, Di Lin 0002, Qing Guo 0005
AAAI1
2026 Learn Like Humans: Use Meta-cognitive Reflection for Efficient Self-Improvement
abstract
While Large Language Models (LLMs) enable complex autonomous behavior, current agents remain constrained by static, humandesigned prompts that limit adaptability.Existing self-improving frameworks attempt to bridge this gap but typically rely on inefficient, multi-turn recursive loops that incur high computational costs.To address this, we propose Metacognitive Agent Reflective Self-improvement (MARS), a framework that achieves efficient self-evolution within a single recurrence cycle.Inspired by educational psychology, MARS mimics human learning by integrating principle-based reflection (abstracting normative rules to avoid errors) and procedural reflection (deriving step-by-step strategies for success).By synthesizing these insights into optimized instructions, MARS allows agents to systematically refine their reasoning logic without continuous online feedback.Extensive experiments on six benchmarks demonstrate that MARS outperforms state-of-the-art self-evolving systems while significantly reducing computational overhead.
Xinmeng Hou, Bohao Qu, Wuqi Wang, Peiliang Gong, Qing Guo 0005
ACL (1)2
2026 SUNRISE: multi-agent reinforcement learning via neighbors' observations under fully noisy environments
Bohao Qu, Menglin Zhang, Xianchang Wang
Expert Syst. Appl.2
2026 State transition difference prediction for deep reinforcement learning
Haotian Chi, Zhaogeng Liu, Xing Chen 0022, Bohao Qu, Jifeng Hu, Yuan Jiang 0007, Hechang Chen, Yi Chang 0001
Pattern Recognit.4
2025 PhysLight: Accurate rPPG Heart Rate Measurement with Adaptive Video Relighting
abstract
Facial video-based remote physiological measurement (rPPG) can non-invasively estimate vital signs, such as heart rate (HR), which often faces challenges under varying lighting conditions. We propose the PhysLight framework to enhance the accuracy of rPPG heart rate measurement through adaptive video relighting. Our approach subtly modifies illumination in video frames to improve detection accuracy while maintaining visual quality. The framework includes a GenLightNet to extract ideal lighting priors and a WipeLightNet module to refine poorly lit videos. Extensive evaluations on benchmark datasets show that our method significantly improves HR estimation reliability, outperforming existing baselines and enhancing non-contact physiological monitoring in diverse environments.
Menglin Zhang, Xiaoxin Guo, Bohao Qu, Xiaofeng Cao 0002, Shuifa Sun, Qing Guo 0005
ICME3
2025 A multi-agent deep reinforcement learning method for fully noisy observations
Danni Wang, Bohao Qu, Menglin Zhang, Xianchang Wang
Eng. Appl. Artif. Intell.3
2024 Diversifying Policies With Non-Markov Dispersion to Expand the Solution Space
abstract
Policy diversity, encompassing the variety of policies an agent can adopt, enhances reinforcement learning (RL) success by fostering more robust, adaptable, and innovative problem-solving in the environment. The environment in which standard RL operates is usually modeled with a Markov Decision Process (MDP) as the theoretical foundation. However, in many real-world scenarios, the rewards depend on an agent's history of states and actions leading to a non-MDP. Under the premise of policy diffusion initialization, non-MDPs may have unstructured expanding solution space due to varying historical information and temporal dependencies. This results in solutions having non-equivalent closed forms in non-MDPs. In this paper, deriving diverse solutions for non-MDPs requires policies to break through the boundaries of the current solution space through gradual dispersion. The goal is to expand the solution space, thereby obtaining more diverse policies. Specifically, we first model the sequences of states and actions by a transformer-based method to learn policy embeddings for dispersion in the solution space, since the transformer has advantages in handling sequential data and capturing long-range dependencies for non-MDP. Then, we stack the policy embeddings to construct a dispersion matrix as the policy diversity measure to induce the policy dispersion in the solution space and obtain a set of diverse policies. Finally, we prove that if the dispersion matrix is positive definite, the dispersed embeddings can effectively enlarge the disagreements across policies, yielding a diverse expression for the original policy embedding distribution. Experimental results of both non-MDP and MDP environments show that this dispersion scheme can obtain more expressive diverse policies via expanding the solution space, showing more robust performance than the recent learning baselines.
Bohao Qu, Xiaofeng Cao 0002, Yi Chang 0001, Ivor W. Tsang, Yew-Soon Ong
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Sample Efficient Offline-to-Online Reinforcement Learning
abstract
Offline reinforcement learning (RL) makes it possible to train the agents entirely from a previously collected dataset. However, constrained by the quality of the offline dataset, offline RL agents typically have limited performance and cannot be directly deployed. Thus, it is desirable to further finetune the pretrained offline RL agents via online interactions with the environment. Existing offline-to-online RL algorithms suffer from the low sample efficiency issue, due to two inherent challenges, i.e., exploration limitation and distribution shift. To this end, we propose a sample-efficient offline-to-online RL algorithm via Optimistic Exploration and Meta Adaptation (OEMA). Specifically, we first propose an optimistic exploration strategy according to the principle of optimism in the face of uncertainty. This allows agents to sufficiently explore the environment in a stable manner. Moreover, we propose a meta learning based adaptation method, which can reduce the distribution shift and accelerate the offline-to-online adaptation process. We empirically demonstrate that OEMA improves the sample efficiency on D4RL benchmark. Besides, we provide in-depth analyses to verify the effectiveness of both optimistic exploration and meta adaptation.
Siyuan Guo 0001, Lixin Zou, Hechang Chen, Bohao Qu, Haotian Chi, Philip S. Yu, Yi Chang 0001
IEEE Trans. Knowl. Data Eng.4
2024 Transductive Reward Inference on Graph
abstract
In this study, we present a transductive inference approach on that reward information propagation graph, which enables the effective estimation of rewards for unlabelled data in offline reinforcement learning. Reward inference is the key to learning effective policies in practical scenarios, while direct environmental interactions are either too costly or unethical and the reward functions are rarely accessible, such as in healthcare and robotics. Our research focuses on developing a reward inference method based on the contextual properties of information propagation on graphs that capitalizes on a constrained number of human reward annotations to infer rewards for unlabelled data. We leverage both the available data and limited reward annotations to construct a reward propagation graph, wherein the edge weights incorporate various influential factors pertaining to the rewards. Subsequently, we employ the constructed graph for transductive reward inference, thereby estimating rewards for unlabelled data. Furthermore, we establish the existence of a fixed point during several iterations of the transductive inference process and demonstrate its at least convergence to a local optimum. Empirical evaluations on locomotion and robotic manipulation tasks validate the effectiveness of our approach. The application of our inferred rewards improves the performance in offline reinforcement learning tasks.
Bohao Qu, Xiaofeng Cao 0002, Qing Guo 0005, Yi Chang 0001, Ivor W. Tsang, Chengqi Zhang
IEEE Trans. Knowl. Data Eng.1
2021 Multi-agent hierarchical policy gradient for Air Combat Tactics emergence via self-play
Zhixiao Sun, Haiyin Piao, Zhen Yang 0011, Guang Zhan, Guanglei Meng, Hechang Chen, Xing Chen 0022, Bohao Qu, Yuanjie Lu
Eng. Appl. Artif. Intell.10
2020 Beyond-Visual-Range Air Combat Tactics Auto-Generation by Reinforcement Learning
abstract
For quite a long time, effective Beyond-Visual-Range (BVR) air combat tactics can only be discovered by human pilots in the actual combat process. However, due to the lack of actual combat opportunities, making new air combat tactics innovation was generally considered quite difficult. To address this challenge, we first introduced a solely end-to-end Reinforcement Learning (RL) approach for training competitive air combat agents with adversarial self-play from scratch in a high fidelity air combat simulation environment during training. Furthermore, a Key Air Combat Event Reward Shaping (KAERS) mechanism was proposed to provide sparse but objective shaped rewards beyond episodic win/lose signal to accelerate the initial machine learning process. Experimental results showed that multiple valuable air combat tactical behaviors emerged progressively. We hope this study could be extended to the future of air combat machine intelligence research.
Haiyin Piao, Zhixiao Sun, Guanglei Meng, Hechang Chen, Bohao Qu, Kuijun Lang, Shengqi Yang, Xuanqi Peng
IJCNN5