Huasong Min

dblp:29/8729 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
9since 2021 · last 2025
0000-0003-4845-0097ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 LangGrasp: Leveraging Fine-Tuned LLMs for Language Interactive Robot Grasping with Ambiguous Instructions
abstract
The existing language-driven grasping methods struggle to fully handle ambiguous instructions containing implicit intents. To tackle this challenge, we propose LangGrasp, a novel language-interactive robotic grasping framework. The framework integrates fine-tuned large language models (LLMs) to leverage their robust commonsense understanding and environmental perception capabilities, thereby deducing implicit intents from linguistic instructions and clarifying task requirements along with target manipulation objects. Furthermore, our designed point cloud localization module, guided by 2D part segmentation, enables partial point cloud localization in scenes, thereby extending grasping operations from coarse-grained object-level to fine-grained part-level manipulation. Experimental results show that the LangGrasp framework accurately resolves implicit intents in ambiguous instructions, identifying critical operations and target information that are unstated yet essential for task completion. Additionally, it dynamically selects optimal grasping poses by integrating environmental information. This enables high-precision grasping from object-level to part-level manipulation, significantly enhancing the adaptability and task execution efficiency of robots in unstructured environments. More information and code are available here: https://github.com/wu467/LangGrasp.
Yunhan Lin, Huasong Min
IROS4
2025 Human-in-the-loop Learning for Adaptive Robot Manipulation using Large Language Models and Behavior Trees
abstract
Large Language Models (LLMs) are now transforming the way robots learn to work in unpredictable environments, such as homes or small enterprises. A growing number of approaches are combining LLMs with Behavior Trees (BTs). Not only do user commands need to be interpreted into BTs that contain the task’s goal, but external disturbances also need to be handled during the process when BT planners dynamically expand BTs based on action databases. However, in these approaches, the action database is manually pre-built and requires the capability for incremental learning and expansion. To address this issue, we propose a human-in-the-loop learning mechanism. First, we design a context for the LLM and then use it to generate action knowledge through in-context learning. In addition, we introduce human-in-the-loop. User feedback is utilized to guide the LLM to correct and refine the action knowledge, ensuring its accuracy and safety. Finally, the generated action knowledge can be directly used for adaptive manipulation without the need for knowledge transfer effort, enabling the robot to complete tasks and handle external disturbances. Experiments across various tasks are conducted and the experimental results validate our method.
Yunhan Lin, Longwu Yan, Huasong Min
IROS4
2025 Generalization of neural network for manipulator inverse dynamics model learning
Yunhan Lin, Chen Jie, Liu Mingxin, Huasong Min
Appl. Intell.5
2024 A 6-DoF Grasping Network Using Feature Augmentation for Novel Domain Generalization
Liran Zhang, Yunhan Lin, Huasong Min
ICIC (12)4
2024 LLM-BT: Performing Robotic Adaptive Tasks based on Large Language Models and Behavior Trees
abstract
Large Language Models (LLMs) have been widely utilized to perform complex robotic tasks. However, handling external disturbances during tasks is still an open challenge. This paper proposes a novel method to achieve robotic adaptive tasks based on LLMs and Behavior Trees (BTs). It utilizes ChatGPT to reason the descriptive steps of tasks. In order to enable ChatGPT to understand the environment, semantic maps are constructed by an object recognition algorithm. Then, we design a Parser module based on Bidirectional Encoder Representations from Transformers (BERT) to parse these steps into initial BTs. Subsequently, a BTs Update algorithm is proposed to expand the initial BTs dynamically to control robots to perform adaptive tasks. Different from other LLM-based methods for complex robotic tasks, our method outputs variable BTs that can add and execute new actions according to environmental changes, which is robust to external disturbances. Our method is validated with simulation in different practical scenarios.
Yunhan Lin, Longwu Yan, Jihong Zhu 0001, Huasong Min
ICRA5
2023 CCD-BSM:composite-curve-dilation brush stroke model for robotic chinese calligraphy
Dongmei Guo 0001, Guang Yan, Huasong Min
Appl. Intell.4
2022 PL-TD3: A Dynamic Path Planning Algorithm of Mobile Robot
abstract
In this paper, Prioritized Experience Replay (PER) strategy and Long Short Term Memory (LSTM) neural network are introduced to the path planning process of mobile robots, which solves the problems of slow convergence and inaccurate perception of dynamic obstacles with the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm. We dubbed this new method as PL-TD3. Firstly, we improve the convergence speed of the algorithm by introducing PER strategy. Secondly, we use LSTM neural network to achieve the improvement of the algorithm for dynamic obstacle perception. In order to verify the method of this paper, we design static environment, dynamic environment and adaptability to dynamic experiments to compare and analyze the methods before and after improvement. The experimental results show that PL-TD3 outperforms TD3 in terms of execution time and execution path length in all environments.
Yijian Tan, Yunhan Lin, Huasong Min
SMC4
2022 Generating Manipulation Sequences using Reinforcement Learning and Behavior Trees for Peg-In-Hole Task
abstract
Reinforcement Learning (RL), a method of learning skills through trial-and-error, has been successfully used in many robotics applications in recent years. We combine manipulation primitives (MPs), behavior trees (BTs), and reinforcement learning to propose an algorithm for peg-in-hole tasks, which speeds up the convergence of the RL model and enhance the adaptability of the dynamic environment. Manipulation primitives are used as actions for RL, which can reduce the gap between control instruction and robotic actions and speed up the convergence of the RL model. Behavior trees are used as robot behavior control, which makes robots can actively adapt to the changes in the environment. In experiments, RL-BT, from the combination of RL and BT, is designed for the peg-in-hole task in the Gazebo simulation environment by using UR5 as the actuator. The experiments are conducted on a simple peg and a complex multi-hole peg by three aspects, which include convergence speed verification experiment, adaptability of dynamic environment experiment, and algorithm robustness experiment. The experiment result proves that our RL-BT can speed up the convergence and adapt to the changes in the environment.
Yunhan Lin, Huasong Min
SMC4
2022 CCAD-Net: A Cascade Cloud Attribute Discrimination Network for Cloud Genera Segmentation in Whole-Sky Images
abstract
Cloud detection and recognition are two important tasks usually referring to image binary segmentation and image-level classification individually. Cloud genera segmentation has more practical significance but is much more challenging as a fine-grained pixel-level dense prediction problem. In this letter, a cascade cloud attribute discrimination network (CCAD-Net) is proposed. Based on an improved encoding-decoding model, CCAD-Net adds a binary segmentation branch for cloud detection and a attribute discrimination branch for cloud attribute feature learning in the decoding stage. Especially, in the attribute discrimination branch, several visual attributes are selected to design the attribute discrimination constraint according to prior professional knowledge and corresponding loss function is defined. These two additional branches and the final cloud genera segmentation branch extract their task-specific features successively and form a cascade structure. Due to the fusion of raw feature, binary segmentation feature, attribute discrimination feature, and cloud genera feature, CCAD-Net can achieve significant better performance than the state-of-the-art methods in cloud genera segmentation in whole-sky images.
Zhiguo Cao 0001, Zhibiao Yang, Huasong Min
IEEE Geosci. Remote. Sens. Lett.4
2018 A Natural Language Interaction Based Automatic Operating System for Industrial Robot
Yunhan Lin, Huasong Min, Mingyu Chen 0004
ICIC (1)2
2013 Experience mixed the modified artificial potential field method
abstract
How to find a safe and collision-free path in unstructured environments is always an important issue in mobile robotics. This paper proposed a new path planning method that exploited past experience for obstacle avoidance with a modified artificial potential field, which could help the robot avoid collisions with obstacles effectively and find the optimal path from the start to the goal. This algorithm uses case-based reasoning to obtain the available prior information of the current environment. By retrieving the past cases and adapting to the changes of the environment to solve the problem. The experiments show that this method greatly improves the performance of the robot in terms of time and distance of the path taken from the start to the target.
Sijing Wang, Huasong Min
IROS2