Kaiqiang Tang

dblp:214/0784 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-7456-0962ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2025 PN-GAIL: Leveraging Non-optimal Information from Imperfect Demonstrations
abstract
Imitation learning aims at constructing an optimal policy by emulating expert demonstrations. However, the prevailing approaches in this domain typically presume that the demonstrations are optimal, an assumption that seldom holds true in the complexities of real-world applications. The data collected in practical scenarios often contains imperfections, encompassing both optimal and non-optimal examples. In this study, we propose Positive-Negative Generative Adversarial Imitation Learning (PN-GAIL), a novel approach that falls within the framework of Generative Adversarial Imitation Learning (GAIL). PN-GAIL innovatively leverages non-optimal information from imperfect demonstrations, allowing the discriminator to comprehensively assess the positive and negative risks associated with these demonstrations. Furthermore, it requires only a small subset of labeled confidence scores. Theoretical analysis indicates that PN-GAIL deviates from the non-optimal data while mimicking imperfect demonstrations. Experimental results demonstrate that PN-GAIL surpasses conventional baseline methods in dealing with imperfect demonstrations, thereby significantly augmenting the practical utility of imitation learning in real-world contexts. Our codes are available at https://github.com/QiangLiuT/PN-GAIL.
Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 0001, Daoyi Dong
ICLR3
2025 Constrained Policy Optimization with Approximately Monotonically Increasing Rewards
abstract
In reinforcement learning (RL), agents maximize accumulated rewards through trial-and-error in the environment to obtain high-performing policies. However, in some situations, loopholes in the purely synthetic reward signals are often exploited by agents, leading to unsafe behaviors, which necessitates the incorporation of safety constraints. In this paper, we propose a safe RL algorithm called Constrained Policy Optimization with Approximately Monotonically Increasing Rewards (CPO-AMIR) to address the safe policy learning in different scenarios and provide practical solutions. We present a novel update formula to achieve a better balance between increasing the reward and decreasing the cost. Furthermore, a theoretical analysis is provided to demonstrate that our approach guarantees the approximately monotonic improvement in rewards when learning constraint-satisfying policies. Our empirical results illustrate the effectiveness and superiority of CPO-AMIR on a set of constrained control tasks.
Yuanyang Lu, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 0001
SMC3
2025 Imitation Learning with Process Adversarial Diffusion
abstract
Generative Adversarial Imitation Learning (GAIL) replicates expert behaviors by employing adversarial training involving the discriminator and the generator. In theory, GAIL can balance discriminator and generator through considerable online interactions to learn well. But real-world adversarial training is tricky and unstable. Early mistakes by the discriminator can lead to bad learning outcomes, or the generator might just copy average expert actions (mode collapse) instead of diverse strategies. To address these, we propose Process Adversarial Diffusion Imitation Learning (PADIL). Employing a conditional diffusion model as the generator facilitates the generation of a multi-modal strategy, thereby reducing the likelihood of mode collapse and enhancing imitation learning performance with fewer online interactions. At the same time, it lessens the generator’s vulnerability to the fluctuating signals or errors conveyed by the discriminator throughout the training process. Furthermore, we have revised the sample extraction approach of the diffusion policy to tackle the concern of iterative instability under rapidly fluctuating reward signals. Experimental results demonstrate that our method achieves better performance compared to baseline methods.
Yiming Qi, Huiqiao Fu, Kaiqiang Tang, Chunlin Chen 0001
SMC3
2025 LMR-IPGN: An Effective Model for automatic summarization of Chinese long text
Gaoan Huang, Lei Ma 0010, Kaiqiang Tang, Sanli Yi, Nuoyun Duan, Chunyun Pu
Multim. Syst.5
2025 Correction: LMR-IPGN: An Effective Model for automatic summarization of Chinese text
Gaoan Huang, Lei Ma 0010, Kaiqiang Tang, Sanli Yi, Nuoyun Duan, Chunyun Pu
Multim. Syst.5
2025 Transformer-based short-term memory attention for enhanced multimodal sentiment analysis
Kaiqiang Tang, Sanli Yi, Lei Ma 0010
Vis. Comput.2
2024 Dynamic Capacitated Vehicle Routing Problem with Stochastic Requests Using Deep Reinforcement Learning
abstract
With the rapid growth of industries like e-commerce, food delivery, and ride-hailing, research on the Vehicle Routing Problem (VRP) is becoming increasingly relevant. However, most of the research in the field of VRP is based on static delivery tasks. In such tasks, the information about customer and order requests is provided before the delivery vehicle departs from the depot, and it remains constant during delivery. However, in the real world, the delivery tasks are often dynamic, where only some orders are known before the vehicle departs from the depot, and the rest of the orders are disclosed over time [1].
Kaiqiang Tang, Huiqiao Fu, Jiasheng Liu, Guizhou Deng, Yuanyang Lu, Chunlin Chen 0001
SMC1
2023 Ess-InfoGAIL: Semi-supervised Imitation Learning from Imbalanced Demonstrations
abstract
Imitation learning aims to reproduce expert behaviors without relying on an explicit reward signal. However, real-world demonstrations often present challenges, such as multi-modal, data imbalance, and expensive labeling processes. In this work, we propose a novel semi-supervised imitation learning architecture that learns disentangled behavior representations from imbalanced demonstrations using limited labeled data. Specifically, our method consists of three key components. First, we adapt the concept of semi-supervised generative adversarial networks to the imitation learning context. Second, we employ a learnable latent distribution to align the generated and expert data distributions. Finally, we utilize a regularized information maximization approach in conjunction with an approximate label prior to further improve the semi-supervised learning performance. Experimental results demonstrate the efficiency of our method in learning multi-modal behaviors from imbalanced demonstrations compared to baseline methods.
Huiqiao Fu, Kaiqiang Tang, Yuanyang Lu, Yiming Qi, Guizhou Deng, Flood Sung, Chunlin Chen 0001
NeurIPS2
2023 Hierarchical Free Gait Motion Planning for Hexapod Robots Using Deep Reinforcement Learning
abstract
This paper addresses the problem of legged locomotion in unstructured environments, and a novel Hierarchical multi-contact motion planning method for hexapod robots is proposed by combining Free Gait motion planning and Deep Reinforcement Learning (HFG-DRL). We structurally decompose the complex free gait multi-contact motion planning task into path planning in discrete state space and gait planning in continuous state space. Firstly, the Soft Deep Q-Network (SDQN) is used to obtain the global prior path information in the Path Planner (PP). Secondly, a Free Gait Planner (FGP) is proposed to obtain the gait sequence. Finally, based on the PP and the FGP, the Center-of-Mass (CoM) sequence is generated by the trained optimal policy using the designed Deep Reinforcement Learning (DRL) algorithm. Experimental results in different environments demonstrate the feasibility, effectiveness, and advancement of the proposed method. Videos are shown athttp://www.hexapod.cn/hfg-drl.html.
Xinpeng Wang 0006, Huiqiao Fu, Guizhou Deng, Canghai Liu, Kaiqiang Tang, Chunlin Chen 0001
IEEE Trans. Ind. Informatics5
2021 Deep Reinforcement Learning for Multi-contact Motion Planning of Hexapod Robots
abstract
Legged locomotion in a complex environment requires careful planning of the footholds of legged robots. In this paper, a novel Deep Reinforcement Learning (DRL) method is proposed to implement multi-contact motion planning for hexapod robots moving on uneven plum-blossom piles. First, the motion of hexapod robots is formulated as a Markov Decision Process (MDP) with a specified reward function. Second, a transition feasibility model is proposed for hexapod robots, which describes the feasibility of the state transition under the condition of satisfying kinematics and dynamics, and in turn determines the rewards. Third, the footholds and Center-of-Mass (CoM) sequences are sampled from a diagonal Gaussian distribution and the sequences are optimized through learning the optimal policies using the designed DRL algorithm. Both of the simulation and experimental results on physical systems demonstrate the feasibility and efficiency of the proposed method. Videos are shown at https://videoviewpage.wixsite.com/mcrl.
Huiqiao Fu, Kaiqiang Tang, Peng Li 0031, Wenqi Zhang 0001, Xinpeng Wang 0006, Guizhou Deng, Tao Wang 0004, Chunlin Chen 0001
IJCAI2
2018 Knowledge Transfer between Multi-granularity Models for Reinforcement Learning
abstract
As a widely used machine learning method, reinforcement learning (RL) is a very effective way to solve decision and control problems where learning skills are needed. In this paper, a knowledge transfer method between multi-granularity models is proposed for RL to speed up the learning process and adapt to the dynamic environments. The learning process runs on naturally organized multi-granularity models, e.g., the coarse-grained model and the fine-grained model. This multi-granularity model constitutes a knowledge transfer architecture that bridges the reinforcement learning between different granularity levels. The proposed multi-granularity reinforcement learning (MGRL) approach and related algorithms can scale up very well and speed up learning with other granularity learning process. Several groups of simulation experiments are carried out using a puzzle problem in a gridworld environment. The results demonstrate the effectiveness and efficiency of the proposed approach.
Bo Xin, Kaiqiang Tang
SMC2