Jingchen Li 0003

dblp:276/9147-3 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0003-0905-0816ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards more effective skill discovery in reinforcement learning by incorporating state reachability
Yang Liu 0195, Jingchen Li 0003, Hua-Rui Wu 0001, Chunjiang Zhao 0001
Neural Networks2
2025 Localizing state space for visual reinforcement learning in noisy environments
Jingchen Li 0003, Haobin Shi
Eng. Appl. Artif. Intell.2
2025 Intelligent long-term planning with interaction for crop management
Jingchen Li 0003, Yusen Yang, Yisheng Miao, Wang Guo, Jingqiu Gu, Hua-Rui Wu 0001
Eng. Appl. Artif. Intell.1
2025 Robotic Locomotion Skill Learning Using Unsupervised Reinforcement Learning With Controllable Latent Space Partition
abstract
Effective skill learning in an unsupervised manner is one of the capabilities an intelligent agent or robot should have. The discovered task-agnostic skills can be fine-tuned to downstream long-horizon tasks to improve execution efficiency. Unfortunately, the self-learning of locomotion skills, which occurs naturally in infancy, has been slow to develop in robotics. The instability exhibited by existing skill-learning methods makes it difficult to directly apply to complex control tasks, such as humanoid robots. To acquire reliable robotic locomotion skills, this article proposes a controllable latent space partition framework to assist reinforcement learning in accomplishing practicability-oriented unsupervised skill discovery (PoSD). Specifically, we use the distance similarity measure of the trajectory feature space to introduce the indicative information of the expert demonstrations into the partitioning and mapping process of the latent space. In addition, the intrinsic subrewards based on contrastive learning and particle entropy are designed to promote skill diversity and encourage exploration. Finally, reinforcement learning completes the generation of skill-conditioned policy driven by composite intrinsic rewards. The performance investigation of our method is conducted on five robots with more than 15 skills. The results indicate that PoSD achieves noticeable improvements in adaptation efficiency and practicability compared with other SOTA unsupervised skill discovery methods.
Ziming He, Haobin Shi, Jingchen Li 0003, Kao-Shing Hwang
IEEE Trans. Ind. Informatics4
2025 Semi-Supervised Detection Model Based on Adaptive Ensemble Learning for Medical Images
abstract
Introducing deep learning technologies into the medical image processing field requires accuracy guarantee, especially for high-resolution images relayed through endoscopes. Moreover, works relying on supervised learning are powerless in the case of inadequate labeled samples. Therefore, for end-to-end medical image detection with overcritical efficiency and accuracy in endoscope detection, an ensemble-learning-based model with a semi-supervised mechanism is developed in this work. To gain a more accurate result through multiple detection models, we propose a new ensemble mechanism, termed alternative adaptive boosting method (Al-Adaboost), combining the decision-making of two hierarchical models. Specifically, the proposal consists of two modules. One is a local region proposal model with attentive temporal-spatial pathways for bounding box regression and classification, and the other one is a recurrent attention model (RAM) to provide more precise inferences for further classification according to the regression result. The proposal Al-Adaboost will adjust the weights of labeled samples and the two classifiers adaptively, and the nonlabel samples are assigned pseudolabels by our model. We investigate the performance of Al-Adaboost on both the colonoscopy and laryngoscopy data coming from CVC-ClinicDB and the affiliated hospital of Kaohsiung Medical University. The experimental results prove the feasibility and superiority of our model.
Jingchen Li 0003, Haobin Shi, Naijun Liu, Kao-Shing Hwang
IEEE Trans. Neural Networks Learn. Syst.1
2025 Eliminating Primacy Bias in Online Reinforcement Learning by Self-Distillation
abstract
Excessive invalid explorations at the beginning of training lead deep reinforcement learning process to fall into the risk of overfitting, further resulting in spurious decisions, which obstruct agents in the following states and explorations. This phenomenon is termed primacy bias in online reinforcement learning. This work systematically investigates the primacy bias in online reinforcement learning, discussing the reason for primacy bias, while the characteristic of primacy bias is also analyzed. Besides, to learn a policy generalized to the following states and explorations, we develop an online reinforcement learning framework, termed self-distillation reinforcement learning (SDRL), based on knowledge distillation, allowing the agent to transfer the learned knowledge into a randomly initialized policy at regular intervals, and the new policy network is used to replace the original one in the following training. The core idea for this work is distilling knowledge from the trained policy to another policy can filter biases out, generating a more generalized policy in the learning process. Moreover, to avoid the overfitting of the new policy due to excessive distillations, we add an additional loss in the knowledge distillation process, using L2 regularization to improve the generalization, and the self-imitation mechanism is introduced to accelerate the learning on the current experiences. The results of several experiments in DMC and Atari 100k suggest the proposal has the ability to eliminate primacy bias for reinforcement learning methods, and the policy after knowledge distillation can urge agents to get higher scores more quickly.
Jingchen Li 0003, Haobin Shi, Hua-Rui Wu 0001, Chunjiang Zhao 0001, Kao-Shing Hwang
IEEE Trans. Neural Networks Learn. Syst.1
2024 Cournot Policy Model: Rethinking centralized training in multi-agent reinforcement learning
Jingchen Li 0003, Yusen Yang, Ziming He, Hua-Rui Wu 0001, Haobin Shi
Inf. Sci.1
2024 Dynamic preference inference network: Improving sample efficiency for multi-objective reinforcement learning by preference estimation
Yang Liu 0195, Ziming He, Yusen Yang, Qingcen Han, Jingchen Li 0003
Knowl. Based Syst.6
2024 Using Goal-Conditioned Reinforcement Learning With Deep Imitation to Control Robot Arm in Flexible Flat Cable Assembly Task
abstract
Leveraging reinforcement learning on high-precision decision-making in Robot Arm assembly scenes is a desired goal in the industrial community. However, tasks like Flexible Flat Cable (FFC) assembly, which require highly trained workers, pose significant challenges due to sparse rewards and limited learning conditions. In this work, we propose a goal-conditioned self-imitation reinforcement learning method for FFC assembly without relying on a specific end-effector, where both perception and behavior plannings are learned through reinforcement learning. We analyze the challenges faced by Robot Arm in high-precision assembly scenarios and balance the breadth and depth of exploration during training. Our end-to-end model consists of hindsight and self-imitation modules, allowing the Robot Arm to leverage futile exploration and optimize successful trajectories. Our method does not require rule-based or manual rewards, and it enables the Robot Arm to quickly find feasible solutions through experience relabeling, while unnecessary explorations are avoided. We train the FFC assembly policy in a simulation environment and transfer it to the real scenario by using domain adaptation. We explore various combinations of hindsight and self-imitation learning, and discuss the results comprehensively. Experimental findings demonstrate that our model achieves fast and advanced flexible flat cable assembly, surpassing other reinforcement learning-based methods.Note to Practitioners—The motivation of this article stems from the need to develop an efficient and accurate FFC assembly policy for 3C (Computer, Communication, and Consumer Electronic) industry, promoting the development of intelligent manufacturing. Traditional control methods are incompetent to complete such a high-precision task with Robot Arm due to the difficult-to-model connectors, and existing reinforcement learning methods cannot converge with restricted epochs because of the difficult goals or trajectories. To quickly learn a high-quality assembly for Robot Arm and accelerate the convergence speed, we combine the goal-conditioned reinforcement learning and self-imitation mechanism, balancing the depth and breadth of exploration. The proposal takes visual information and six-dimensions force as state, obtaining satisfactory assembly policies. We build a simulation scene by the Pybullet platform and pre-train the Robot Arm on it, and then the pre-trained policies can be reused in real scenarios with finetuning.
Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang
IEEE Trans Autom. Sci. Eng.1
2023 A multi-agent reinforcement learning method with curriculum transfer for large-scale dynamic traffic signal control
Xuesi Li, Jingchen Li 0003, Haobin Shi
Appl. Intell.2
2023 Mix-attention approximation for homogeneous large-scale multi-agent reinforcement learning
Shike Yang, Jingchen Li 0003, Haobin Shi
Neural Comput. Appl.2
2023 Lateral Transfer Learning for Multiagent Reinforcement Learning
abstract
Some researchers have introduced transfer learning mechanisms to multiagent reinforcement learning (MARL). However, the existing works devoted to cross-task transfer for multiagent systems were designed just for homogeneous agents or similar domains. This work proposes an all-purpose cross-transfer method, called multiagent lateral transfer (MALT), assisting MARL with alleviating the training burden. We discuss several challenges in developing an all-purpose multiagent cross-task transfer learning method and provide a feasible way of reusing knowledge for MARL. In the developed method, we take features as the transfer object rather than policies or experiences, inspired by the progressive network. To achieve more efficient transfer, we assign pretrained policy networks for agents based on clustering, while an attention module is introduced to enhance the transfer framework. The proposed method has no strict requirements for the source task and target task. Compared with the existing works, our method can transfer knowledge among heterogeneous agents and also avoid negative transfer in the case of fully different tasks. As far as we know, this article is the first work denoted to all-purpose cross-task transfer for MARL. Several experiments in various scenarios have been conducted to compare the performance of the proposed method with baselines. The results demonstrate that the method is sufficiently flexible for most settings, including cooperative, competitive, homogeneous, and heterogeneous configurations.
Haobin Shi, Jingchen Li 0003, Jiahui Mao, Kao-Shing Hwang
IEEE Trans. Cybern.2
2023 Path Planning of Randomly Scattering Waypoints for Wafer Probing Based on Deep Attention Mechanism
abstract
Wafer probing is a critical process employed to measure the yield of wafer fabrication. The primary object of wafer probing is to find the defect grain on the wafer. After a full coverage check, there are always some suspected grains existing for further inspection. However, this second probing result could be affected by the shape of the probe card and the setting actions (path planning) of operators for grains randomly scattering on the wafer. Good grains can be damaged by reprobe actions, which decrease production performance and customer trust. In general, it also requires manpower to perform reprobing, which dramatically deteriorates the throughput of production. This article has studied this problem, and an adaptive coverage path planning (CPP) method for randomly scattering grains using an attention interface is proposed. The proposed randomly scattering waypoints method uses deep reinforcement learning (DRL) for automatic real-time path planning of the second detection. A soft attention interface accelerates the process with a less overlapped check. The experimental results demonstrate the efficiency of the proposed method in terms of less overlapping and steps, and this method learns a better CPP strategy for wafer probing than programmed paths and other RL-based methods.
Haobin Shi, Jingchen Li 0003, Meng Liang, Maxwell Hwang, Kao-Shing Hwang, Yun-Yu Hsu
IEEE Trans. Syst. Man Cybern. Syst.2
2022 Multi-agent reinforcement learning by the actor-critic model with an attention interface
Lixiang Zhang, Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang
Neurocomputing2
2022 A collaboration of multi-agent model using an interactive interface
Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang
Inf. Sci.1
2022 A behavior fusion method based on inverse reinforcement learning
Haobin Shi, Jingchen Li 0003, Shicong Chen, Kao-Shing Hwang
Inf. Sci.2
2022 An adaptive multi-sensor visual attention model
Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang
Neural Comput. Appl.2
2022 Using Fuzzy Logic to Learn Abstract Policies in Large-Scale Multiagent Reinforcement Learning
abstract
Large-scale multiagent reinforcement learning requires huge computation and space costs, and the too-long execution process makes it hard to train policies for agents. This work proposes a concept of fuzzy agent, which is a new paradigm for training homogeneous agents. Aiming at a lightweight and affordable reinforcement learning mechanism for large-scale homogeneous multiagent systems, we break the one-to-one correspondence between agent and policy, designing abstract agents as the substitute for the multiagent to interact with the environment. The Markov decision process models for these abstract agents are conducted by fuzzy logic, which also acts on the behavior mapping from abstract agent to entity. Specifically, just the abstract agents execute their policy at a time step, and the concrete behaviors are generated by simple matrix operations. The proposal has lower space and computation complexities because the number of abstract agents is far less than that of entities, and the coupling among agents is retained implicitly. Compared with other approximation and simplification methods, the proposed fuzzy agent not only greatly reduces required computing resources but also ensures the effectiveness of the learned policies. Several experiments are conducted to validate our method. The results show that the proposal outperforms the baseline methods, while it has satisfactory zero-shot and few-shot transfer abilities.
Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang
IEEE Trans. Fuzzy Syst.1
2021 An explainable ensemble feedforward method with Gaussian convolutional filter
Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang
Knowl. Based Syst.1
2021 Graph convolutional network-based reinforcement learning for tasks offloading in multi-access edge computing
Lixiong Leng, Jingchen Li 0003, Haobin Shi
Multim. Tools Appl.2
2021 Attentive Hybrid Recurrent Neural Networks for sequential recommendation
Lixiang Zhang, Peisen Wang, Jingchen Li 0003, Zhiwei Xiao, Haobin Shi
Neural Comput. Appl.3