VLDB 2026 Research / reviewers in the wild / expert
Huiqiao Fu
dblp:243/7065
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0001-9403-2449ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PN-GAIL: Leveraging Non-optimal Information from Imperfect DemonstrationsabstractImitation learning aims at constructing an optimal policy by emulating expert demonstrations. However, the prevailing approaches in this domain typically presume that the demonstrations are optimal, an assumption that seldom holds true in the complexities of real-world applications. The data collected in practical scenarios often contains imperfections, encompassing both optimal and non-optimal examples. In this study, we propose Positive-Negative Generative Adversarial Imitation Learning (PN-GAIL), a novel approach that falls within the framework of Generative Adversarial Imitation Learning (GAIL). PN-GAIL innovatively leverages non-optimal information from imperfect demonstrations, allowing the discriminator to comprehensively assess the positive and negative risks associated with these demonstrations. Furthermore, it requires only a small subset of labeled confidence scores. Theoretical analysis indicates that PN-GAIL deviates from the non-optimal data while mimicking imperfect demonstrations. Experimental results demonstrate that PN-GAIL surpasses conventional baseline methods in dealing with imperfect demonstrations, thereby significantly augmenting the practical utility of imitation learning in real-world contexts. Our codes are available at https://github.com/QiangLiuT/PN-GAIL. Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 0001, Daoyi Dong |
ICLR | 2 |
| 2025 | DEAL: Diffusion Evolution Adversarial Learning for Sim-to-Real TransferabstractTraining Reinforcement Learning (RL) controllers in simulation offers cost-efficiency and safety advantages. However, the resultant policies often suffer significant performance degradation during real-world deployment due to the reality gap. Previous works like System Identification (Sys-Id) have attempted to bridge this discrepancy by improving simulator fidelity, but encounter challenges including the collapse of high-dimensional parameter identification, low identification accuracy, and unstable convergence dynamics. To address these challenges, we propose a novel Sys-Id framework that combines Diffusion Evolution with Adversarial Learning (DEAL) to iteratively infer physical parameters with limited real-world data, which makes the state transitions between simulation and reality as similar as possible. Specifically, our method iteratively refines physical parameters through a dual mechanism: a discriminator network evaluates the similarity of state transitions between parameterized simulations and target environment as fitness guidance, while diffusion evolution adaptively modulates noise prediction and denoising processes to optimize parameter distributions. We validate DEAL in both simulated and real-world environments. Compared to baseline methods, DEAL demonstrates state-of-the-art stability and identification accuracy in high-dimensional parameter identification tasks, and significantly enhances sim-to-real transfer performance while requiring minimal real-world data. Huiqiao Fu, Zhehao Zhou, Chunlin Chen 0001 |
NeurIPS | 2 |
| 2025 | Constrained Policy Optimization with Approximately Monotonically Increasing RewardsabstractIn reinforcement learning (RL), agents maximize accumulated rewards through trial-and-error in the environment to obtain high-performing policies. However, in some situations, loopholes in the purely synthetic reward signals are often exploited by agents, leading to unsafe behaviors, which necessitates the incorporation of safety constraints. In this paper, we propose a safe RL algorithm called Constrained Policy Optimization with Approximately Monotonically Increasing Rewards (CPO-AMIR) to address the safe policy learning in different scenarios and provide practical solutions. We present a novel update formula to achieve a better balance between increasing the reward and decreasing the cost. Furthermore, a theoretical analysis is provided to demonstrate that our approach guarantees the approximately monotonic improvement in rewards when learning constraint-satisfying policies. Our empirical results illustrate the effectiveness and superiority of CPO-AMIR on a set of constrained control tasks. Yuanyang Lu, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 0001 |
SMC | 2 |
| 2025 | Imitation Learning with Process Adversarial DiffusionabstractGenerative Adversarial Imitation Learning (GAIL) replicates expert behaviors by employing adversarial training involving the discriminator and the generator. In theory, GAIL can balance discriminator and generator through considerable online interactions to learn well. But real-world adversarial training is tricky and unstable. Early mistakes by the discriminator can lead to bad learning outcomes, or the generator might just copy average expert actions (mode collapse) instead of diverse strategies. To address these, we propose Process Adversarial Diffusion Imitation Learning (PADIL). Employing a conditional diffusion model as the generator facilitates the generation of a multi-modal strategy, thereby reducing the likelihood of mode collapse and enhancing imitation learning performance with fewer online interactions. At the same time, it lessens the generator’s vulnerability to the fluctuating signals or errors conveyed by the discriminator throughout the training process. Furthermore, we have revised the sample extraction approach of the diffusion policy to tackle the concern of iterative instability under rapidly fluctuating reward signals. Experimental results demonstrate that our method achieves better performance compared to baseline methods. Yiming Qi, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 0001 |
SMC | 2 |
| 2024 | EASI: Evolutionary Adversarial Simulator Identification for Sim-to-Real TransferabstractReinforcement Learning (RL) controllers have demonstrated remarkable performance in complex robot control tasks. However, the presence of reality gap often leads to poor performance when deploying policies trained in simulation directly onto real robots. Previous sim-to-real algorithms like Domain Randomization (DR) requires domain-specific expertise and suffers from issues such as reduced control performance and high training costs. In this work, we introduce Evolutionary Adversarial Simulator Identification (EASI), a novel approach that combines Generative Adversarial Network (GAN) and Evolutionary Strategy (ES) to address sim-to-real challenges. Specifically, we consider the problem of sim-to-real as a search problem, where ES acts as a generator in adversarial competition with a neural network discriminator, aiming to find physical parameter distributions that make the state transitions between simulation and reality as similar as possible. The discriminator serves as the fitness function, guiding the evolution of the physical parameter distributions. EASI features simplicity, low cost, and high fidelity, enabling the construction of a more realistic simulator with minimal requirements for real-world data, thus aiding in transferring simulated-trained policies to the real world. We demonstrate the performance of EASI in both sim-to-sim and sim-to-real tasks, showing superior performance compared to existing sim-to-real algorithms. Huiqiao Fu, Zhehao Zhou, Chunlin Chen 0001 |
NeurIPS | 2 |
| 2024 | Dynamic Capacitated Vehicle Routing Problem with Stochastic Requests Using Deep Reinforcement LearningabstractWith the rapid growth of industries like e-commerce, food delivery, and ride-hailing, research on the Vehicle Routing Problem (VRP) is becoming increasingly relevant. However, most of the research in the field of VRP is based on static delivery tasks. In such tasks, the information about customer and order requests is provided before the delivery vehicle departs from the depot, and it remains constant during delivery. However, in the real world, the delivery tasks are often dynamic, where only some orders are known before the vehicle departs from the depot, and the rest of the orders are disclosed over time [1]. Kaiqiang Tang, Huiqiao Fu, Jiasheng Liu, Guizhou Deng, Yuanyang Lu, Chunlin Chen 0001 |
SMC | 2 |
| 2023 | Ess-InfoGAIL: Semi-supervised Imitation Learning from Imbalanced DemonstrationsabstractImitation learning aims to reproduce expert behaviors without relying on an explicit reward signal. However, real-world demonstrations often present challenges, such as multi-modal, data imbalance, and expensive labeling processes. In this work, we propose a novel semi-supervised imitation learning architecture that learns disentangled behavior representations from imbalanced demonstrations using limited labeled data. Specifically, our method consists of three key components. First, we adapt the concept of semi-supervised generative adversarial networks to the imitation learning context. Second, we employ a learnable latent distribution to align the generated and expert data distributions. Finally, we utilize a regularized information maximization approach in conjunction with an approximate label prior to further improve the semi-supervised learning performance. Experimental results demonstrate the efficiency of our method in learning multi-modal behaviors from imbalanced demonstrations compared to baseline methods. Huiqiao Fu, Kaiqiang Tang, Yuanyang Lu, Yiming Qi, Guizhou Deng, Flood Sung, Chunlin Chen 0001 |
NeurIPS | 1 |
| 2023 | Hierarchical Free Gait Motion Planning for Hexapod Robots Using Deep Reinforcement LearningabstractThis paper addresses the problem of legged locomotion in unstructured environments, and a novel Hierarchical multi-contact motion planning method for hexapod robots is proposed by combining Free Gait motion planning and Deep Reinforcement Learning (HFG-DRL). We structurally decompose the complex free gait multi-contact motion planning task into path planning in discrete state space and gait planning in continuous state space. Firstly, the Soft Deep Q-Network (SDQN) is used to obtain the global prior path information in the Path Planner (PP). Secondly, a Free Gait Planner (FGP) is proposed to obtain the gait sequence. Finally, based on the PP and the FGP, the Center-of-Mass (CoM) sequence is generated by the trained optimal policy using the designed Deep Reinforcement Learning (DRL) algorithm. Experimental results in different environments demonstrate the feasibility, effectiveness, and advancement of the proposed method. Videos are shown athttp://www.hexapod.cn/hfg-drl.html. Xinpeng Wang 0006, Huiqiao Fu, Guizhou Deng, Canghai Liu, Kaiqiang Tang, Chunlin Chen 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Deep Reinforcement Learning for Multi-contact Motion Planning of Hexapod RobotsabstractLegged locomotion in a complex environment requires careful planning of the footholds of legged robots. In this paper, a novel Deep Reinforcement Learning (DRL) method is proposed to implement multi-contact motion planning for hexapod robots moving on uneven plum-blossom piles. First, the motion of hexapod robots is formulated as a Markov Decision Process (MDP) with a specified reward function. Second, a transition feasibility model is proposed for hexapod robots, which describes the feasibility of the state transition under the condition of satisfying kinematics and dynamics, and in turn determines the rewards. Third, the footholds and Center-of-Mass (CoM) sequences are sampled from a diagonal Gaussian distribution and the sequences are optimized through learning the optimal policies using the designed DRL algorithm. Both of the simulation and experimental results on physical systems demonstrate the feasibility and efficiency of the proposed method. Videos are shown at https://videoviewpage.wixsite.com/mcrl. Huiqiao Fu, Kaiqiang Tang, Peng Li 0031, Wenqi Zhang 0001, Xinpeng Wang 0006, Guizhou Deng, Tao Wang 0004, Chunlin Chen 0001 |
IJCAI | 1 |
| 2021 | Learning to Navigate in a VUCA Environment: Hierarchical Multi-expert ApproachabstractDespite decades of efforts, robot navigation in a real scenario with volatility, uncertainty, complexity, and ambiguity (VUCA for short), remains a challenging topic. Inspired by the central nervous system (CNS), we propose a hierarchical multi-expert learning framework for autonomous navigation in a VUCA environment. With a heuristic exploration mechanism considering target location, path cost, and safety level, the upper layer performs simultaneous map exploration and route-planning to avoid trapping in a blind alley, similar to the cerebrum in the CNS. Using a local adaptive model fusing multiple discrepant strategies, the lower layer pursuits a balance between collision-avoidance and go-straight strategies, acting as the cerebellum in the CNS. We conduct simulation and real-world experiments on multiple platforms, including legged and wheeled robots. Experimental results demonstrate our algorithm outperforms the existing methods in terms of task achievement, time efficiency, and security. A video of our results is available at https://youtu.be/lAnW4QIWDoU. Wenqi Zhang 0001, Peng Li 0031, Faping Ye, Weijie Jiang 0003, Huiqiao Fu, Tao Wang 0004 |
IROS | 7 |