Hongcai He

dblp:360/6308 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 56% Transfer learning and domain adaptation · 20% Learning theory · 9%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
meta-reinforcement learning
1.822026
Improving Generalization in Offline Meta-Reinforcement Learning via Cross-task Contexts · AAAI 2026
Decoupling Meta-Reinforcement Learning with Gaussian Task Contexts and Skills · AAAI 2024
Machine learning › Reinforcement learning › meta-reinforcement learning
offline meta-reinforcement learning
1.822026
Improving Generalization in Offline Meta-Reinforcement Learning via Cross-task Contexts · AAAI 2026
Decoupling Meta-Reinforcement Learning with Gaussian Task Contexts and Skills · AAAI 2024
Machine learning › Reinforcement learning › transfer learning in reinforcement learning
context adaptation
1.012026
Improving Generalization in Offline Meta-Reinforcement Learning via Cross-task Contexts · AAAI 2026
Machine learning › Learning theory
generalization
1.012026
Improving Generalization in Offline Meta-Reinforcement Learning via Cross-task Contexts · AAAI 2026
Machine learning › Reinforcement learning
imitation learning
0.912025
Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline Data · AAAI 2025
Robotics › Robot manipulation
learning from demonstration
0.912025
Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline Data · AAAI 2025
Machine learning › Reinforcement learning
sparse reward reinforcement learning
0.912025
Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline Data · AAAI 2025
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.712023
Manifold Regularized Joint Transfer for Open Set Domain Adaptation · IEEE Trans. Multim. 2023
Machine learning › Learning paradigms › semi-supervised learning › graph-based semi-supervised learning
manifold regularization
0.712023
Manifold Regularized Joint Transfer for Open Set Domain Adaptation · IEEE Trans. Multim. 2023
Machine learning › Transfer learning and domain adaptation › domain adaptation
open-set domain adaptation
0.712023
Manifold Regularized Joint Transfer for Open Set Domain Adaptation · IEEE Trans. Multim. 2023
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.712023
Manifold Regularized Joint Transfer for Open Set Domain Adaptation · IEEE Trans. Multim. 2023
Machine learning › Transfer learning and domain adaptation
meta-learning
0.312025
Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline Data · AAAI 2025
Robotics › Motion planning and robot control
robot control
0.212024
Decoupling Meta-Reinforcement Learning with Gaussian Task Contexts and Skills · AAAI 2024

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.8variational autoencoder · 1.0context quantization · 1.0off-policy reinforcement learning · 0.9meta-learning · 0.9imitation learning · 0.9gaussian quantization variational autoencoder · 0.8codebook · 0.8reproducing kernel hilbert space · 0.7manifold regularization · 0.7
YearPublicationVenuePosition
2026 Improving Generalization in Offline Meta-Reinforcement Learning via Cross-task Contexts
abstract
Context-based offline meta-reinforcement learning (meta-RL) is a paradigm that integrates meta-learning with offline reinforcement learning. It learns a strategy to extract task-specific contexts from trajectories of meta-training tasks and leverages this strategy for adapting to unseen target tasks. However, existing methods struggle to generate generalizable contexts for adaptations due to context shift, which arises from the context-based policy overfitting to offline data. We argue that leveraging the internal relationships among tasks, rather than treating each task in isolation, is crucial for mitigating the impact of context shift. Hence, we propose a framework called cross-task contexts for improving generalization in meta-RL (CTMRL). Specifically, we design a context quantization variational auto-encoder (CQ-VAE), which clusters task-specific contexts of meta-training tasks into discrete codes based on the internal relationships among tasks. Cross-task contexts are constructed with these codes, reflecting shared information across similar tasks. These cross-task contexts not only serve as high-level structures to capture similarity across tasks but also provide a foundation for hard contrastive learning that enhances the distinguishability of similar yet distinct tasks, thereby improving the generalization of contexts and facilitating adaptation to unseen target tasks. The evaluation in meta-environments confirms the performance advantage of CTMRL over existing methods.
Hongcai He, Zetao Zheng, Anjie Zhu, Deqiang Ouyang, Jie Shao 0001
AAAI1
2025 Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline Data
abstract
A major challenge in Reinforcement Learning (RL) is the difficulty of learning an optimal policy from sparse rewards. Prior works enhance online RL with conventional Imitation Learning (IL) via a handcrafted auxiliary objective, at the cost of restricting the RL policy to be sub-optimal when the offline data is generated by a non-expert policy. Instead, to better leverage valuable information in offline data, we develop Generalized Imitation Learning from Demonstration (GILD), which meta-learns an objective that distills knowledge from offline data and instills intrinsic motivation towards the optimal policy. Distinct from prior works that are exclusive to a specific RL algorithm, GILD is a flexible module intended for diverse vanilla off-policy RL algorithms. In addition, GILD introduces no domain-specific hyperparameter and minimal increase in computational cost. In four challenging MuJoCo tasks with sparse rewards, we show that three RL algorithms enhanced with GILD significantly outperform state-of-the-art methods.
Shilong Deng, Zetao Zheng, Hongcai He, Paul Weng, Jie Shao 0001
AAAI3
2024 Decoupling Meta-Reinforcement Learning with Gaussian Task Contexts and Skills
abstract
Offline meta-reinforcement learning (meta-RL) methods, which adapt to unseen target tasks with prior experience, are essential in robot control tasks. Current methods typically utilize task contexts and skills as prior experience, where task contexts are related to the information within each task and skills represent a set of temporally extended actions for solving subtasks. However, these methods still suffer from limited performance when adapting to unseen target tasks, mainly because the learned prior experience lacks generalization, i.e., they are unable to extract effective prior experience from meta-training tasks by exploration and learning of continuous latent spaces. We propose a framework called decoupled meta-reinforcement learning (DCMRL), which (1) contrastively restricts the learning of task contexts through pulling in similar task contexts within the same task and pushing away different task contexts of different tasks, and (2) utilizes a Gaussian quantization variational autoencoder (GQ-VAE) for clustering the Gaussian distributions of the task contexts and skills respectively, and decoupling the exploration and learning processes of their spaces. These cluster centers which serve as representative and discrete distributions of task context and skill are stored in task context codebook and skill codebook, respectively. DCMRL can acquire generalizable prior experience and achieve effective adaptation to unseen target tasks during the meta-testing phase. Experiments in the navigation and robot manipulation continuous control tasks show that DCMRL is more effective than previous meta-RL methods with more generalizable prior experience.
Hongcai He, Anjie Zhu, Shuang Liang 0002, Feiyu Chen 0001, Jie Shao 0001
AAAI1
2023 Deep Reinforcement Learning for Stock Trading with Behavioral Finance Strategy
Shilong Deng, Zetao Zheng, Hongcai He, Jie Shao 0001
ADMA (2)3
2023 Modeling Both Collaborative and Temporal Information for Sequential Recommendation
Jinyue Dai, Jie Shao 0001, Zhiyi Deng, Hongcai He, Feiyu Chen 0001
ICONIP (9)4
2023 Manifold Regularized Joint Transfer for Open Set Domain Adaptation
abstract
Unsupervised Domain Adaptation (UDA) aims to leverage knowledge of a well-labeled source domain to learn an effective classifier for an unlabeled target domain. However, a common scenario in real-world applications is that the target domain contains unknown categories that are not observed in the source domain. This setting is termed as open set domain adaptation (OSDA). Most existing approaches of OSDA can only classify known classes well but fail to recognize unknown samples effectively. In this paper, we propose an effective method, named manifold regularized joint transfer (MRJT), for OSDA. MRJT learns new feature representations by simultaneously reducing distribution discrepancy between domains, increasing compactness of within-class, discriminating different known classes, and distinguishing the unknown from the known. The learned new features are projected onto reproducing kernel Hilbert space. In this space, a weighted structural risk minimization method is integrated with manifold regularization to utilize geometric information sufficiently to learn an effective classifier. Extensive experimental results on four real-world datasets verify the superiority of our method. It can not only classify known samples into the right known classes but also recognize unknown samples effectively.
Jieyan Liu, Hongcai He, Mingzhu Liu, Jingjing Li 0001, Ke Lu 0001
IEEE Trans. Multim.2