VLDB 2026 Research / reviewers in the wild / expert
Anjie Zhu
dblp:215/6704
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-4634-7961ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Generalization in Offline Meta-Reinforcement Learning via Cross-task ContextsabstractContext-based offline meta-reinforcement learning (meta-RL) is a paradigm that integrates meta-learning with offline reinforcement learning. It learns a strategy to extract task-specific contexts from trajectories of meta-training tasks and leverages this strategy for adapting to unseen target tasks. However, existing methods struggle to generate generalizable contexts for adaptations due to context shift, which arises from the context-based policy overfitting to offline data. We argue that leveraging the internal relationships among tasks, rather than treating each task in isolation, is crucial for mitigating the impact of context shift. Hence, we propose a framework called cross-task contexts for improving generalization in meta-RL (CTMRL). Specifically, we design a context quantization variational auto-encoder (CQ-VAE), which clusters task-specific contexts of meta-training tasks into discrete codes based on the internal relationships among tasks. Cross-task contexts are constructed with these codes, reflecting shared information across similar tasks. These cross-task contexts not only serve as high-level structures to capture similarity across tasks but also provide a foundation for hard contrastive learning that enhances the distinguishability of similar yet distinct tasks, thereby improving the generalization of contexts and facilitating adaptation to unseen target tasks. The evaluation in meta-environments confirms the performance advantage of CTMRL over existing methods. Hongcai He, Zetao Zheng, Anjie Zhu, Deqiang Ouyang, Jie Shao 0001 |
AAAI | 3 |
| 2024 | Decoupling Meta-Reinforcement Learning with Gaussian Task Contexts and SkillsabstractOffline meta-reinforcement learning (meta-RL) methods, which adapt to unseen target tasks with prior experience, are essential in robot control tasks. Current methods typically utilize task contexts and skills as prior experience, where task contexts are related to the information within each task and skills represent a set of temporally extended actions for solving subtasks. However, these methods still suffer from limited performance when adapting to unseen target tasks, mainly because the learned prior experience lacks generalization, i.e., they are unable to extract effective prior experience from meta-training tasks by exploration and learning of continuous latent spaces. We propose a framework called decoupled meta-reinforcement learning (DCMRL), which (1) contrastively restricts the learning of task contexts through pulling in similar task contexts within the same task and pushing away different task contexts of different tasks, and (2) utilizes a Gaussian quantization variational autoencoder (GQ-VAE) for clustering the Gaussian distributions of the task contexts and skills respectively, and decoupling the exploration and learning processes of their spaces. These cluster centers which serve as representative and discrete distributions of task context and skill are stored in task context codebook and skill codebook, respectively. DCMRL can acquire generalizable prior experience and achieve effective adaptation to unseen target tasks during the meta-testing phase. Experiments in the navigation and robot manipulation continuous control tasks show that DCMRL is more effective than previous meta-RL methods with more generalizable prior experience. Hongcai He, Anjie Zhu, Shuang Liang 0002, Feiyu Chen 0001, Jie Shao 0001 |
AAAI | 2 |
| 2024 | Abstract and Explore: A Novel Behavioral Metric with Cyclic Dynamics in Reinforcement LearningabstractIntrinsic motivation lies at the heart of the exploration of reinforcement learning, which is primarily driven by the agent's inherent satisfaction rather than external feedback from the environment. However, in recent more challenging procedurally-generated environments with high stochasticity and uninformative extrinsic rewards, we identify two significant issues of applying intrinsic motivation. (1) State representation collapse: In existing methods, the learned representations within intrinsic motivation have a high probability to neglect the distinction among different states and be distracted by the task-irrelevant information brought by the stochasticity. (2) Insufficient interrelation among dynamics: Unsuccessful guidance provided by the uninformative extrinsic reward makes the dynamics learning in intrinsic motivation less effective. In light of the above observations, a novel Behavioral metric with Cyclic Dynamics (BCD) is proposed, which considers both cumulative and immediate effects and facilitates the abstraction and exploration of the agent. For the behavioral metric, the successor feature is utilized to reveal the expected future rewards and alleviate the heavy reliance of previous methods on extrinsic rewards. Moreover, the latent variable and vector quantization techniques are employed to enable an accurate measurement of the transition function in a discrete and interpretable manner. In addition, cyclic dynamics is established to capture the interrelations between state and action, thereby providing a thorough awareness of environmental dynamics. Extensive experiments conducted on procedurally-generated environments demonstrate the state-of-the-art performance of our proposed BCD. Anjie Zhu, Peng-Fei Zhang 0001, Ruihong Qiu, Zetao Zheng, Zi Huang, Jie Shao 0001 |
AAAI | 1 |
| 2024 | HIT: Solving Partial Index Tracking via Hierarchical Reinforcement LearningabstractPartial index tracking (PIT) is a popular passive investment strategy aiming at replicating the performance of a market index (e.g., S&P 500). Existing PIT methods typically treat it as a regression problem and divide it into two tasks: (i) asset selection (determining which assets to choose from the index constituents) and (ii) asset allocation (deciding how to allocate capital among the selected assets). However, these methods either optimize these two tasks jointly, which has been proven to be NP-hard and inefficient when tracking large-scale constituent indices (e.g., Russell 2000), or attempt an independent optimization, lacking a connection to ensure collaborative optimization. In this paper, we present a hierarchical model for partial index tracking (HIT), which formulates PIT as a hierarchical Markov decision process (MDP) and is optimized via hierarchical reinforcement learning (HRL). HIT consists of (1) a high-level policy learns to select assets from constituents to handle task (i) and (2) a low-level policy learns to allocate capital weights among the selected assets to handle task (ii). We further propose a novel cost-sensitive reward function that serves as a connection to collaboratively optimize the two policies, aiming to replicate the index closely while considering transaction cost. Compared with existing jointly optimized approaches, our model simplifies the problem by learning separate policies for the two tasks, and the reward function serves as a connection to ensure collaborative optimization between them, avoiding challenges faced by joint optimization methods in existing literature. Remarkable performance across 6 benchmarks, ranging from small to large-scale constituents demonstrate the superiority of HIT. Moreover, the experiments conducted on a real-world market dataset spanning over 10 years show its effectiveness and practicality. Zetao Zheng, Jie Shao 0001, Feiyu Chen 0001, Anjie Zhu, Shilong Deng, Heng Tao Shen |
ICDE | 4 |
| 2024 | Cross-Insight Trader: A Trading Approach Integrating Policies with Diverse Investment Horizons for Portfolio ManagementabstractDeep reinforcement learning (RL) has emerged as a promising approach for portfolio management due to its ability to make sequential decisions. However, applying RL techniques to this domain is still challenging due to the non-stationary nature of financial markets. Existing RL-based solutions fail to consider the intrinsic causes behind this non-stationary, which primarily stem from the involvement of diverse traders with distinct investment horizons and their varied investment strategies. In this paper, we tackle the non-stationary problem by examining its intrinsic causes and propose cross-insight trader, a novel two-step RL-based approach that integrates multiple trading policies with different investment horizons to adapt to the changing market conditions. In the first step, we learn multiple horizon-specific policies by providing each policy with tailored information specific to its investment horizon. This allows each policy to recognize dynamic patterns within its respective horizon and make insightful pre-decisions. In the second step, we learn a cross-insight policy to make the final trade decision by considering the investment pre-decisions made by multiple horizon-specific policies in the first step. To enable effective learning of two types of policies, our approach employs a centralized critic to evaluate the actions performed by both horizon-specific and cross-insight policies. By incorporating multiple insights from different investment horizons into the decision-making process, our approach enhances its adaptability to changing market conditions. Experimental results conducted on three stock markets demonstrate the superiority of our framework. Zetao Zheng, Jie Shao 0001, Shilong Deng, Anjie Zhu, Heng Tao Shen, Xiaofang Zhou 0001 |
ICDE | 4 |
| 2024 | HierarT: Multi-hop temporal knowledge graph forecasting with hierarchical reinforcement learning
Xuewei Luo, Anjie Zhu, Jie Shao 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Modeling Hierarchical Uncertainty for Multimodal Emotion Recognition in ConversationabstractApproximating the uncertainty of an emotional AI agent is crucial for improving the reliability of such agents and facilitating human-in-the-loop solutions, especially in critical scenarios. However, none of the existing systems for emotion recognition in conversation (ERC) has attempted to estimate the uncertainty of their predictions. In this article, we present HU-Dialogue, which models hierarchical uncertainty for the ERC task. We perturb contextual attention weight values with source-adaptive noises within each modality, as a regularization scheme to model context-level uncertainty and adapt the Bayesian deep learning method to the capsule-based prediction layer to model modality-level uncertainty. Furthermore, a weight-sharing triplet structure with conditional layer normalization is introduced to detect both invariance and equivariance among modalities for ERC. We provide a detailed empirical analysis for extensive experiments, which shows that our model outperforms previous state-of-the-art methods on three popular multimodal ERC datasets. Feiyu Chen 0001, Jie Shao 0001, Anjie Zhu, Deqiang Ouyang, Xueliang Liu, Heng Tao Shen |
IEEE Trans. Cybern. | 3 |
| 2023 | Window-Controlled Sepsis Prediction Using a Model Selection Approach
Shiyan Su, Su Lan, Anjie Zhu |
ADMA (5) | 4 |
| 2023 | Empowering the Diversity and Individuality of Option: Residual Soft Option Critic FrameworkabstractExtracting temporal abstraction (option), which empowers the action space, is a crucial challenge in hierarchical reinforcement learning. Under a well-structured action space, decision-making agents can probe more deeply in the searching or plan efficiently through pruning irrelevant action candidates. However, automatically capturing a well-performed temporal abstraction is a nontrivial challenge due to its insufficient exploration and inadequate functionality. We consider alleviating this challenge from two perspectives, i.e., diversity and individuality. For the aspect of diversity, we propose a maximum entropy model based on ensembled options to encourage exploration. For the aspect of individuality, we propose to distinguish each option accurately, utilizing mutual formation minimization, so that each option can better express and function. We name our framework as an ensemble with soft option (ESO) critics. Furthermore, the residual algorithm (RA) with a bidirectional target network is introduced to stabilize bootstrapping, yielding a residual version of ESO. We provide detailed analysis for extensive experiments, which shows that our method boosts performance in commonly used continuous control tasks. Anjie Zhu, Feiyu Chen 0001, Deqiang Ouyang, Jie Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Hyper-node Relational Graph Attention Network for Multi-modal Knowledge Graph CompletionabstractKnowledge graphs often suffer from incompleteness, and knowledge graph completion (KGC) aims at inferring the missing triplets through knowledge graph embedding from known factual triplets. However, most existing knowledge graph embedding methods only use the relational information of knowledge graph and treat the entities and relations as IDs with simple embedding layer, ignoring the multi-modal information among triplets, such as text descriptions, images, etc. In this work, we propose a novel network to incorporate different modal information with graph structure information for more precise representation of multi-modal knowledge graph, termed as hyper-node relational graph attention (HRGAT) network. In HRGAT, we use low-rank multi-modal fusion to model the intra-modality and inter-modality dynamics, which transforms the original knowledge graph to a hyper-node graph. Then, relational graph attention (RGAT) network is used, which contains relation-specific attention and entity-relation fusion operation to capture the graph structure information. Finally, we aggregate the updated multi-modal information and graph structure information to generate the final embeddings of knowledge graph to achieve KGC. By exploring multi-modal information and graph structure information, HRGAT embraces faster convergence speed and achieves the state-of-the-art for KGC on the standard datasets. Implementation code is available at https://github.com/broliang/HRGAT. Shuang Liang 0002, Anjie Zhu, Jie Shao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | SMDT: Cross-View Geo-Localization with Image Alignment and TransformerabstractThe goal of cross-view geo-localization is to determine the location of a given ground image by matching with aerial images. However, existing methods ignore the variability of scenes, additional information and spatial correspondence of covisibility and non-convisibility areas in ground-aerial image pairs. In this context, we propose a cross-view matching method called SMDT with image alignment and Transformer. First, we utilize semantic segmentation technique to segment different areas. Then, we convert the vertical view of aerial images to front view by mixing polar mapping and perspective mapping. Next, we simultaneously train dual conditional generative adversarial nets by taking the semantic segmentation images and converted images as input to synthesize the aerial image with ground view style. These steps are collectively referred to as image alignment. Last, we use Transformer to explicitly utilize the properties of self-attention. Experiments show that our SMDT method is superior to the existing ground-to-aerial cross-view methods. Xiaoyang Tian, Jie Shao 0001, Deqiang Ouyang, Anjie Zhu, Feiyu Chen 0001 |
ICME | 4 |
| 2022 | Step by step: A hierarchical framework for multi-hop knowledge graph reasoning with reinforcement learning
Anjie Zhu, Deqiang Ouyang, Shuang Liang 0002, Jie Shao 0001 |
Knowl. Based Syst. | 1 |