EDBT 2026 Demo / reviewers in the wild / expert
Jifeng Hu
dblp:316/5980
· DBLP profile ↗
18ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards context-aware graph representation learning: Adaptive node aggregation with LLMs
Songwei Zhao, Yuan Jiang 0007, Sinuo Zhang, Jifeng Hu, Philip S. Yu, Hechang Chen |
Artif. Intell. | 4 |
| 2026 | Instructed Diffuser With Temporal Condition Guidance for Offline Reinforcement LearningabstractRecentworks have shown the potential of diffusion models in computer vision and natural language processing. Apart from the classical supervised learning fields, diffusion models have also shown strong competitiveness in reinforcement learning (RL) by formulating decision-making as sequential generation. However, incorporating temporal information of sequential data and utilizing it to guide diffusion models to perform better generation is still an open challenge. In this paper, we take one step forward to investigate controllable generation with temporal conditions that are refined from temporal information. We observe the importance of temporal conditions in sequential generation in sufficient scenarios and provide a comprehensive discussion and comparison of different temporal conditions. Based on the observations, we propose an effective temporally-conditional diffusion model coined Temporally-Composable Diffuser (TCD), which extracts temporal information from interaction sequences and explicitly guides generation with temporal conditions. Specifically, we separate the sequences into three parts according to time expansion and identify historical, immediate, and prospective conditions accordingly. Each condition preserves non-overlapping temporal information of sequences, enabling more controllable generation when we jointly use them to guide the diffuser. Finally, we conduct extensive experiments and analysis to reveal the favorable applicability of TCD in offline RL tasks, where our method reaches or matches the best performance compared with prior SOTA baselines. Jifeng Hu, Yanchao Sun, Sili Huang, Siyuan Guo 0001, Hechang Chen, Li Shen 0008, Lichao Sun 0001, Yi Chang 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | State transition difference prediction for deep reinforcement learning
Haotian Chi, Zhaogeng Liu, Xing Chen 0022, Bohao Qu, Jifeng Hu, Yuan Jiang 0007, Hechang Chen, Yi Chang 0001 |
Pattern Recognit. | 5 |
| 2026 | A degree-corrected stochastic block model for community discovery in signed networks with heterogeneous degree distributions
Zhejian Yang, Yang Li 0030, Bo Yu 0013, Jifeng Hu, Hechang Chen |
Pattern Recognit. | 4 |
| 2026 | A Simple Unified Uncertainty-Guided Framework for Offline-to-Online Reinforcement LearningabstractOffline reinforcement learning (RL) provides a promising solution to learning an agent fully relying on a data-driven paradigm. However, constrained by the limited quality of the offline dataset, its performance is often suboptimal. Therefore, it is desired to further finetune the agent via extra online interactions before deployment. Unfortunately, offline-to-online RL can be challenging due to two main challenges: constrained exploratory behavior and state-action distribution shift. In view of this, we propose a simple unified uncertainty-guided (SUNG) framework, which naturally unifies the solution to both challenges with the tool of uncertainty. Specifically, SUNG quantifies uncertainty via a variational autoencoder (VAE)-based state-action visitation density estimator. To facilitate efficient exploration, SUNG presents a practical optimistic exploration strategy to select informative actions with both high value and high uncertainty. Moreover, SUNG develops an adaptive exploitation method by applying conservative offline RL objectives to high-uncertainty samples and standard online RL objectives to low-uncertainty samples to smoothly bridge offline and online stages. SUNG achieves state-of-the-art online finetuning performance when combined with different offline RL methods, across various environments and datasets in the D4RL benchmark. Codes are made publicly available in https://github.com/guosyjlu/ACBR. Siyuan Guo 0001, Yanchao Sun, Jifeng Hu, Sili Huang, Hechang Chen, Haiyin Piao, Lichao Sun 0001, Yi Chang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Tackling Continual Offline RL through Selective Weights Activation on Aligned SpacesabstractContinual offline reinforcement learning (CORL) has shown impressive ability in diffusion-based continual learning systems by modeling the joint distributions of trajectories. However, most research only focuses on limited continual task settings where the tasks have the same observation and action space, which deviates from the realistic demands of training agents in various environments. In view of this, we propose Vector-Quantized Continual Diffuser, named VQ-CD, to break the barrier of different spaces between various tasks. Specifically, our method contains two complementary sections, where the quantization spaces alignment provides a unified basis for the selective weights activation. In the quantized spaces alignment, we leverage vector quantization to align the different state and action spaces of various tasks, facilitating continual training in the same space. Then, we propose to leverage a unified diffusion model attached by the inverse dynamic model to master all tasks by selectively activating different weights according to the task-related sparse masks. Finally, we conduct extensive experiments on 15 continual learning (CL) tasks, including conventional CL task settings (identical state and action spaces) and general CL task settings (various state and action spaces). Compared with 17 baselines, our method reaches the SOTA performance. Jifeng Hu, Sili Huang, Li Shen 0008, Zhejian Yang, Shengchao Hu, Shisong Tang, Hechang Chen, Lichao Sun 0001, Yi Chang 0001, Dacheng Tao |
NeurIPS | 1 |
| 2025 | Analytic Energy-Guided Policy Optimization for Offline Reinforcement LearningabstractConditional decision generation with diffusion models has shown powerful competitiveness in reinforcement learning (RL). Recent studies reveal the relation between energy-function-guidance diffusion models and constrained RL problems. The main challenge lies in estimating the intermediate energy, which is intractable due to the log-expectation formulation during the generation process. To address this issue, we propose the Analytic Energy-guided Policy Optimization (AEPO). Specifically, we first provide a theoretical analysis and the closed-form solution of the intermediate guidance when the diffusion model obeys the conditional Gaussian transformation. Then, we analyze the posterior Gaussian distribution in the log-expectation formulation and obtain the target estimation of the log-expectation under mild assumptions. Finally, we train an intermediate energy neural network to approach the target estimation of log-expectation formulation. We apply our method in 30+ offline RL tasks to demonstrate the effectiveness of our method. Extensive experiments illustrate that our method surpasses numerous representative baselines in D4RL offline reinforcement learning benchmarks. Jifeng Hu, Sili Huang, Zhejian Yang, Shengchao Hu, Li Shen 0008, Hechang Chen, Lichao Sun 0001, Yi Chang 0001, Dacheng Tao |
NeurIPS | 1 |
| 2025 | Generalizable Causal Reinforcement Learning for Out-of-Distribution EnvironmentsabstractOut-of-distribution (OOD) generalization is critical for applying reinforcement learning algorithms to real-world applications. To address the OOD problem, recent works focus on learning an OOD adaptation policy by capturing the causal factors affecting the environmental dynamics. However, these works recover the causal factors with only an entangled or binary form, resulting in a limited generalization of the policy that requires extra data from the testing environments. To break this limitation, we propose generalizable causal reinforcement learning (GCRL) to learn a disentangled representation of causal factors, on the basis of which we learn a policy that achieves the OOD generalization without extra training. For capturing the causal factors, GCRL deploys a weakly supervised signal with a two-stage constraint to ensure that all factors can be disentangled. Then, to achieve the OOD generalization through causal factors, we establish the dependence of actions on the learned representation and optimize the policy model across multiple environments. Experimental results show that the established dependence recovers the correct relationship between causal factors and actions when the learned policy could address the target tasks in training environments. Benefiting from the recovered relationship, GCRL achieves the OOD generalization on eight benchmarks from Causal World and Mujoco. Moreover, the policy learned by our model is more explainable and can be controlled to generate semantic actions by intervening in the representation of causal factors. Sili Huang, Jifeng Hu, Hechang Chen, Peng Cui 0001, Haiyin Piao, Lichao Sun 0001, Bo Yang 0002 |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | A Flexible Diffusion Convolution for Graph Neural NetworksabstractGraph Neural Networks (GNNs) have been gaining more attention due to their excellent performance in modeling various graph-structured data. However, most of the current GNNs only consider fixed-neighbor discrete message-passing, disregarding the importance of the local structure of different nodes and the implicit information between nodes for smoothing features. Previous approaches either focus on adaptive selection for aggregation structures or treat discrete graph convolution as a continuous diffusion process, but none of them comprehensively considered the above issues, significantly limiting the model's performance. To this end, we present a novel approach called Flexible Diffusion Convolution (Flexi-DC), which exploits the neighborhood information of nodes to set a particular continuous diffusion for each node to smooth features. Specifically, Flexi-DC first extracts the local structure knowledge based on the degrees of nodes in the graph data and then injects it into the diffusion convolution module to smooth features. Additionally, we utilize the extracted knowledge to smooth labels. Flexi-DC is an efficient framework that can significantly improve the performance of most GNN architectures. Experimental results demonstrate that Flexi-DC outperforms their vanilla implementations by an average accuracy of 13.24% (GCN), 16.37% (JKNet), and 11.98% (ARMA) on nine graph datasets with different homophily ratios. Songwei Zhao, Bo Yu 0013, Sinuo Zhang, Jifeng Hu, Yuan Jiang 0007, Philip S. Yu, Hechang Chen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | EGNN: Exploring Structure-Level Neighborhoods in Graphs With Varying Homophily RatiosabstractGraph neural networks (GNNs) have garnered significant attention for their competitive performance on graph-structured data. However, many existing methods are commonly constrained by the homophily assumption, making them overly reliant on the uniform neighbor propagation, which limits their ability to generalize to heterophilous graphs. Although some approaches extend aggregation to multi-hop neighbors, adapting neighborhood sizes on a per-node basis remains a significant challenge. In view of this, we propose an Evolutionary Graph Neural Network (EGNN) with adaptive structure-level aggregation and label smoothing, offering a novel solution to the aforementioned drawback. The core innovation of EGNN lies in assigning each node apersonalizedneighborhood structure utilizingbehavior-levelcrossover and mutation. Specifically, we first adaptively search for the optimal structure-level neighborhoods for nodes within the solution space, leveraging the exploratory capabilities of evolutionary computation. This approach enhances the exchange of information between the target node and surrounding nodes, achieving a smooth vector representation. Subsequently, we adopt the optimal structure obtained through evolutionary search to perform label smoothing, further boosting the robustness of the framework. We conduct experiments on nine real-world networks with different homophily ratios, where outstanding performance demonstrates that the ability of EGNN can match or surpass SOTA baselines. Songwei Zhao, Bo Yu 0013, Sinuo Zhang, Zhejian Yang, Jifeng Hu, Philip S. Yu, Hechang Chen |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | Continual Diffuser (CoD): Mastering Continual Offline RL With Experience RehearsalabstractArtificial neural networks, especially recent diffusion-based models, have shown remarkable superiority in gaming, control, and QA systems, where the training tasks' datasets are usually static. However, in real-world applications, such as robotic control of reinforcement learning (RL), the tasks are changing, and new tasks arise in a sequential order. This situation poses the new challenge of plasticity-stability tradeoff for training an agent who can adapt to task changes and retain acquired knowledge. In view of this, we propose a rehearsal-based continual diffusion model, called continual diffuser (CoD), to endow the diffuser with the capabilities of quick adaptation (plasticity) and lasting retention (stability). Specifically, we first construct an offline benchmark that contains 90 tasks from multiple domains. Then, we train the CoD on each task with sequential modeling and conditional generation for making decisions. Next, we preserve a small portion of previous datasets as the rehearsal buffer and replay it to retain the acquired knowledge. Extensive experiments on a series of tasks show that CoD can achieve a promising plasticity-stability tradeoff and outperform existing diffusion-based methods and other representative baselines on most tasks. The source code is available at https://github.com/JF-Hu/Continual_Diffuser. Jifeng Hu, Li Shen 0008, Sili Huang, Zhejian Yang, Hechang Chen, Lichao Sun 0001, Yi Chang 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-ThoughtabstractIn-context learning is a promising approach for offline reinforcement learning (RL) to handle online tasks, which can be achieved by providing task prompts. Recent works demonstrated that in-context RL could emerge with self-improvement in a trial-and-error manner when treating RL tasks as an across-episodic sequential prediction problem. Despite the self-improvement not requiring gradient updates, current works still suffer from high computational costs when the across-episodic sequence increases with task horizons. To this end, we propose an In-context Decision Transformer (IDT) to achieve self-improvement in a high-level trial-and-error manner. Specifically, IDT is inspired by the efficient hierarchical structure of human decision-making and thus reconstructs the sequence to consist of high-level decisions instead of low-level actions that interact with environments. As one high-level decision can guide multi-step low-level actions, IDT naturally avoids excessively long sequences and solves online tasks more efficiently. Experimental results show that IDT achieves state-of-the-art in long-horizon tasks over current in-context RL methods. In particular, the online evaluation time of our IDT is 36$\times$ times faster than baselines in the D4RL benchmark and 27$\times$ times faster in the Grid World benchmark. Sili Huang, Jifeng Hu, Hechang Chen, Lichao Sun 0001, Bo Yang 0002 |
ICML | 2 |
| 2024 | Effective State Space Exploration with Phase State Graph Generation and Goal-based Path PlanningabstractExploring the state space efficiently is a crucial problem in reinforcement learning as it holds significant importance for learning optimal policies. One effective approach involves learning different sub-policies to cover various sub-spaces of the state space, with each sub-policy corresponding to a specific goal. However, the unevenness of the state probability distribution may lead to exploration difficulties in deep reinforcement learning. To overcome this challenge, we propose a Phase State Graph Exploration framework (PSGE), guiding the agent towards more promising directions for exploration. Specifically, we design a graph-based state space exploration framework to separate the combination space into sub-spaces and define the combination space and evaluation criteria for the agent’s sub-policies. In addition, hypernetwork is leveraged to decouple sub-policies and sub-goals, ensuring diversity among the agent’s sub-policies and reward shaping is used to provide dense internal reward signals for policy training, which encourages the agent to learn more efficiently. Experiments on combining control and navigation tasks demonstrate that PSGE performs well in controlling agent across various difficulty level tasks. Sinuo Zhang, Jifeng Hu, Xinqi Du, Zhejian Yang, Hechang Chen |
IJCNN | 2 |
| 2024 | Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence ModelingabstractRecent works have shown the remarkable superiority of transformer models in reinforcement learning (RL), where the decision-making problem is formulated as sequential generation. Transformer-based agents could emerge with self-improvement in online environments by providing task contexts, such as multiple trajectories, called in-context RL. However, due to the quadratic computation complexity of attention in transformers, current in-context RL methods suffer from huge computational costs as the task horizon increases. In contrast, the Mamba model is renowned for its efficient ability to process long-term dependencies, which provides an opportunity for in-context RL to solve tasks that require long-term memory. To this end, we first implement Decision Mamba (DM) by replacing the backbone of Decision Transformer (DT). Then, we propose a Decision Mamba-Hybrid (DM-H) with the merits of transformers and Mamba in high-quality prediction and long-term memory. Specifically, DM-H first generates high-value sub-goals from long-term memory through the Mamba model. Then, we use sub-goals to prompt the transformer, establishing high-quality predictions. Experimental results demonstrate that DM-H achieves state-of-the-art in long and short-term tasks, such as D4RL, Grid World, and Tmaze benchmarks. Regarding efficiency, the online testing of DM-H in the long-term task is 28$\times$ times faster than the transformer-based baselines. Sili Huang, Jifeng Hu, Zhejian Yang, Hechang Chen, Lichao Sun 0001, Bo Yang 0002 |
NeurIPS | 2 |
| 2024 | Generalized multi-agent competitive reinforcement learning with differential augmentation
Hechang Chen, Jifeng Hu, Zhejian Yang, Bo Yu 0013, Xinqi Du, Yinxiao Miao, Yi Chang 0001 |
Expert Syst. Appl. | 3 |
| 2023 | QVDDPG: QV Learning with Balanced Constraint in Actor-Critic FrameworkabstractActor-critic framework has achieved tremendous success in a great many of decision-making scenarios. Nevertheless, when updating the value of new states and actions in the long-term scene, these methods suffer from misestimate problem and gradient variance problem, significantly reducing convergence speed and robustness of the policy. These problems severely limit the application scope of these methods. In this paper, we first proposed QVDDPG, a deep RL algorithm based on the iterative target value update process. The QV learning method alleviates the problem of misestimate by making use of the guidance of Q value and the fast convergence of V value, thus accelerating the convergence speed. In addition, the actor utilizes a constrained balanced gradient and establishes a hidden state for the continuous action space network for the sake of robustness of the model. We give the update relation among the value functions and the constraint conditions of gradient estimation. We measure our method on the PyBullet and achieved state-of-the-art performance. Moreover, we demonstrate that, our method has higher robustness and convergence speed across different tasks compared to other algorithms. Jifeng Hu, Luheng Yang, Zhihang Ren, Hechang Chen, Bo Yang 0002 |
IJCNN | 2 |
| 2023 | Learning Generalizable Agents via Saliency-guided Features DecorrelationabstractIn visual-based Reinforcement Learning (RL), agents often struggle to generalize well to environmental variations in the state space that were not observed during training. The variations can arise in both task-irrelevant features, such as background noise, and task-relevant features, such as robot configurations, that are related to the optimal decisions. To achieve generalization in both situations, agents are required to accurately understand the impact of changed features on the decisions, i.e., establishing the true associations between changed features and decisions in the policy model. However, due to the inherent correlations among features in the state space, the associations between features and decisions become entangled, making it difficult for the policy to distinguish them. To this end, we propose Saliency-Guided Features Decorrelation (SGFD) to eliminate these correlations through sample reweighting. Concretely, SGFD consists of two core techniques: Random Fourier Functions (RFF) and the saliency map. RFF is utilized to estimate the complex non-linear correlations in high-dimensional images, while the saliency map is designed to identify the changed features. Under the guidance of the saliency map, SGFD employs sample reweighting to minimize the estimated correlations related to changed features, thereby achieving decorrelation in visual RL tasks. Our experimental results demonstrate that SGFD can generalize well on a wide range of test environments and significantly outperforms state-of-the-art methods in handling both task-irrelevant variations and task-relevant variations. Sili Huang, Yanchao Sun, Jifeng Hu, Siyuan Guo 0001, Hechang Chen, Yi Chang 0001, Lichao Sun 0001, Bo Yang 0002 |
NeurIPS | 3 |
| 2022 | Distributional Reward Estimation for Effective Multi-agent Deep Reinforcement LearningabstractMulti-agent reinforcement learning has drawn increasing attention in practice, e.g., robotics and automatic driving, as it can explore optimal policies using samples generated by interacting with the environment. However, high reward uncertainty still remains a problem when we want to train a satisfactory model, because obtaining high-quality reward feedback is usually expensive and even infeasible. To handle this issue, previous methods mainly focus on passive reward correction. At the same time, recent active reward estimation methods have proven to be a recipe for reducing the effect of reward uncertainty. In this paper, we propose a novel Distributional Reward Estimation framework for effective Multi-Agent Reinforcement Learning (DRE-MARL). Our main idea is to design the multi-action-branch reward estimation and policy-weighted reward aggregation for stabilized training. Specifically, we design the multi-action-branch reward estimation to model reward distributions on all action branches. Then we utilize reward aggregation to obtain stable updating signals during training. Our intuition is that consideration of all possible consequences of actions could be useful for learning policies. The superiority of the DRE-MARL is demonstrated using benchmark multi-agent scenarios, compared with the SOTA baselines in terms of both effectiveness and robustness. Jifeng Hu, Yanchao Sun, Hechang Chen, Sili Huang, Haiyin Piao, Yi Chang 0001, Lichao Sun 0001 |
NeurIPS | 1 |