Gaofeng Li

dblp:119/1717 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Robot-Assisted Cross-Modal Synthetic Augmentation of mmWave Datasets for Sign Language Recognition
Zhipeng Tang, Xiuzhen Guo, Shibo He, Yuanchao Shu, Gaofeng Li, Chaojie Gu
IEEE Trans. Mob. Comput.5
2026 DexRepNet++: Learning Dexterous Robotic Manipulation With Geometric and Spatial Hand-Object Representations
abstract
Robotic dexterous manipulation is a challenging problem due to high degrees of freedom (DoFs) and complex contacts of multi-fingered robotic hands. Many existing deep reinforcement learning (DRL) based methods aim at improving sample efficiency in high-dimensional output action spaces. However, existing works often overlook the role of representations in achieving generalization of a manipulation policy in the complex input space during the hand-object interaction. In this paper, we propose DexRep, a novel hand-object interaction representation to capture object surface features and spatial relations between hands and objects for dexterous manipulation skill learning. Based on DexRep, policies are learned for three dexterous manipulation tasks, i.e. grasping, in-hand reorientation, bimanual handover, and extensive experiments are conducted to verify the effectiveness. In simulation, for grasping, the policy learned with 40 objects achieves a success rate of 87.9% on more than 5000 unseen objects of diverse categories, significantly surpassing existing work trained with thousands of objects; for the in-hand reorientation and handover tasks, the policies also boost the success rates and other metrics of existing hand-object representations by 20% to 40%. The grasp policies with DexRep are deployed to the real world under multi-camera and single-camera setups and demonstrate a small sim-to-real gap.
Qingtao Liu, Zhengnan Sun, Haoming Li 0004, Gaofeng Li, Lin Shao 0002, Jiming Chen 0001, Qi Ye 0001
IEEE Trans. Robotics5
2025 VTDexManip: A Dataset and Benchmark for Visual-tactile Pretraining and Dexterous Manipulation with Reinforcement Learning
abstract
Vision and touch are the most commonly used senses in human manipulation. While leveraging human manipulation videos for robotic task pretraining has shown promise in prior works, it is limited to image and language modalities and deployment to simple parallel grippers. In this paper, aiming to address the limitations, we collect a vision-tactile dataset by humans manipulating 10 daily tasks and 182 objects. In contrast with the existing datasets, our dataset is the first visual-tactile dataset for complex robotic manipulation skill learning. Also, we introduce a novel benchmark, featuring six complex dexterous manipulation tasks and a reinforcement learning-based vision-tactile skill learning framework. 18 non-pretraining and pretraining methods within the framework are designed and compared to investigate the effectiveness of different modalities and pertaining strategies. Key findings based on our benchmark results and analyses experiments include: 1) Despite the tactile modality used in our experiments being binary and sparse, including it directly in the policy training boosts the success rate by about 20\% and joint pretraining it with vision gains a further 20\%. 2) Joint pretraining visual-tactile modalities exhibits strong adaptability in unknown tasks and achieves robust performance among all tasks. 3) Using binary tactile signals with vision is robust to viewpoint setting, tactile noise, and the binarization threshold, which facilitates to the visual-tactile policy to be deployed in reality. The dataset and benchmark are available at \url{https://github.com/LQTS/VTDexManip}.
Qingtao Liu, Zhengnan Sun, Gaofeng Li, Jiming Chen 0001, Qi Ye 0001
ICLR4
2025 A Hybrid Mapping Method: Balancing Efficiency and Intuitiveness in Lateral Teleoperation
abstract
Mobile manipulators integrate the locomotion flexibility of quadruped robots with the operational capabilities of robotic manipulators. This integrated system is particularly effective for teleoperating explosive ordnance disposal (EOD) tasks in hazardous environments, enabling the safe handling of explosive devices. However, when the quadruped operates in narrow corridors or cluttered spaces, its ability to reposition is limited. This limitation, combined with targets located laterally relative to the robot, poses critical challenges for achieving rapid and intuitive teleoperation of the manipulator. Existing manipulator mapping methods either fail to support lateral teleoperation or lack proper coordinate transformations, leading to mismatches between the intended and actual movement directions of the leader and follower devices. This reduces operational intuitiveness and increases the cognitive load on human operators. To overcome these issues, we propose a hybrid mapping method that combines joint-space velocity control with Cartesian-space control. This method leverages joint-space velocity commands for rapid manipulator reorientation, while employing Cartesian-space commands to achieve precise end-effector teleoperation. Furthermore, we introduce a virtual base coordinate frame that adaptively adjusts in response to the manipulator’s reorientation. This adaptive compensation ensures that the visual feedback from the camera mounted on the end-effector remains consistent and intuitive. The proposed method was validated through experiments on a quadruped robot equipped with a manipulator in an EOD scenario. Results demonstrated significant improvements, including 100% success rate, 43.9% task duration reduction, and 31.7% NASA-TLX score decrease, indicating decreased cognitive load and enhanced task efficiency compared to baseline methods.
Yuwei Xie, Jiming Chen 0001, Gaofeng Li
IROS4
2024 InterRep: A Visual Interaction Representation for Robotic Grasping
abstract
Recently, pre-trained vision models have gained significant attention in motor control, showcasing impressive performance across diverse robotic learning tasks. While previous works predominantly concentrate on the significance of the pre-training phase, the equally important task of extracting more effective representations based on existing pre-trained visual models remains unexplored. To better leverage the representation capabilities of pre-trained models for robotic grasping, we propose InterRep, a novel interaction representation method that possesses not only the strengths of pre-trained models, known for their robustness in noisy environments and their proficiency in recognizing essential features, but also the capacity of capturing dynamic interaction details and local geometric features during the grasping process. Based on the novel representation, we introduce a deep reinforcement learning method to learn generalizable grasping policies. The experimental results demonstrate that our proposed representation outperforms the baselines in terms of both training speed and generalization. For the generalized grasping tasks with dexterous robotic hands, our method boasts a success rate nearly 20% higher than methods using the global features of the entire image from pre-trained models. In addition, our proposed representation method demonstrates promising performance when applied to a different robotic hand and task. It also exhibits excellent performance on real robots with a success rate of 70%.
Qi Ye 0001, Qingtao Liu, Anjun Chen, Gaofeng Li, Jiming Chen 0001
ICRA5
2024 Masked Visual-Tactile Pre-training for Robot Manipulation
abstract
Recent works on the pretraining for robot manipulation have demonstrated that representations learning from large human manipulation data can generalize well to new manipulation tasks and environments. However, these approaches mainly focus on human vision or natural language, neglecting tactile feedback. In this article, we make an attempt to explore how to pre-train a representation model for robotic manipulation using both human manipulation visual and tactile data. We develop a system for collecting visual and tactile data, featuring a cost-effective tactile glove to capture human tactile data and Hololens2 for capturing visual data. With this system, we collect a dataset of turning bottle caps. Furthermore, we introduce a novel visual-tactile fusion network and learning strategy M2VTP, with one key module to tokenize 20 sparse binary tactile signals sensing touch states for the learning of tactile context and the other key module applying the attention and mask mechanism to the interaction of visual and tactile tokens for visual-tactile representation learning. We utilize our dataset to pre-train the fusion model and embed the pre-trained model into a reinforcement learning framework for downstream tasks. Experimental results demonstrate that our pre-trained model significantly aids in learning manipulation skills. Compared to methods without pre-training, our approach achieves a success rate increase of over 60%. Additionally, when compared to current visual pre-training methods, our success rate exceeds them by more than 50%.
Qingtao Liu, Qi Ye 0001, Zhengnan Sun, Gaofeng Li, Jiming Chen 0001
ICRA5
2024 The Joint-Space Reconstruction of Human Fingers by using a Highly Under-Actuated Exoskeleton
abstract
Hand motion tracking is essential in many fields, e.g., immersive virtual reality, teleoperation of robotic hand, and hand rehabilitation of stroke patient, as human hand plays a crucial role in our daily life. The highly under-actuated hand exoskeleton, which can track the 6-DoF motions of each fingertip via a highly under-actuated kinematic chain, exhibits many benefits in wearability and portability over other solutions. However, due to the non-anthropomorphic linkage, this hand exoskeleton also encounters difficulties in measuring human-finger’s joint angles. While the joint-space is important in many scenarios, such as teleoperating a robotic hand with anthropomorphic kinematics but with different size to human. Here we proposed a new method to reconstruct the human finger joints by using a highly under-actuated hand exoskeleton. Our key contribution is the arc-fitting algorithm, which is able to calibrate the misalignment between the exoskeleton’s and the human-finger’s base frames and estimate the length of human’s phalanxes, by using the fingertip’s circular motions. With knowing the aforementioned informations, the joint angles can be reconstructed in high precision based on the inverse kinematics models of human fingers. Furthermore, our proposed method is compared with a baseline method, in which the joint angles obtained by a motion capture system are served as ground-truth. The results demonstrate that our proposed method exhibits excellent performance in reconstructing finger’s joint configurations.
Yuan Su, Gaofeng Li, Yongsheng Deng, Ioannis Sarakoglou, Nikolaos G. Tsagarakis, Jiming Chen 0001
ICRA2
2024 A Light-weight and Rapid Table Tennis Ball Trajectory Prediction Approaches towards Online Bouncing Task
abstract
It is essentially required to predict the ball’s flight trajectory accurately and timely for a robotic table tennis ball bouncing task. Existing solutions, which can be categorized into model-based and learning-based groups, both exhibits unpleasant disadvantages. For example, they often require to identify many dynamic parameters accurately or to collect extensive labeled data, which are generally very difficult or costly to achieve in real world. In this paper, we proposed a light-wight and rapid trajectory prediction approach for online table tennis bouncing tasks based on a simplified model. In the proposed approach, the ball’s flight poses are captured and estimated by a low-cost RGB-D camera. Then the ball’s landing position is predicted in advance by using a fitted 3D parabola. Compared with existing solutions, our proposed approach is lightweight and easy to deploy. In experiments, 66 flight trajectories of the ball are collected to serve as benchmark. The prediction errors for all landing positions are all less than 20mm, in which most of them are less than 10mm. In addition, the prediction can be achieved 141.7ms in advance, which is fast enough for the robotic arm to plan and move itself to the predicted landing point.
Peisen Xu, Gaofeng Li, Qi Ye 0001, Jiming Chen 0001
RO-MAN2
2023 DexRepNet: Learning Dexterous Robotic Grasping Network with Geometric and Spatial Hand-Object Representations
abstract
Robotic dexterous grasping is a challenging problem due to the high degree of freedom (DoF) and complex contacts of multi-fingered robotic hands. Existing deep re-inforcement learning (DRL) based methods leverage human demonstrations to reduce sample complexity due to the high dimensional action space with dexterous grasping. However, less attention has been paid to hand-object interaction representations for high-level generalization. In this paper, we propose a novel geometric and spatial hand-object interaction representation, named DexRep, to capture object surface features and the spatial relations between hands and objects during grasping. DexRep comprises Occupancy Feature for rough shapes within sensing range by moving hands, Surface Feature for changing hand-object surface distances, and LocalGeo Feature for local geometric surface features most related to potential contacts. Based on the new representation, we propose a dexterous deep reinforcement learning method DexRepNet to learn a generalizable grasping policy. Experimental results show that our method outperforms baselines using existing representations for robotic grasping dramatically both in grasp success rate and convergence speed. It achieves a 93% grasping success rate on seen objects and higher than 80% grasping success rates on diverse objects of unseen categories in both simulation and real-world experiments.
Qingtao Liu, Qi Ye 0001, Zhengnan Sun, Haoming Li 0004, Gaofeng Li, Lin Shao 0002, Jiming Chen 0001
IROS6
2023 On Perpendicular Curve-Based Task Space Trajectory Tracking Control With Incomplete Orientation Constraint
abstract
The Incomplete Orientation Constraint (IOC), which does not require a controlled motion constrained by all three spatial directions, exists widely in a lot of robotic tasks. However, the IOC remains a challenge for existing methods, due to the nonlinear structure of the rotation group SO(3). Moreover, the IOCs are time varying in the trajectory tracking problem, which makes it more challenging than the set-point control. To address the IOC problems, we define, identify and prove the closed-form solution of the perpendicular curve in SO(3). Based on the proposed perpendicular curve, we develop a new trajectory tracking controller considering the IOC. Compared with existing methods, the proposed method can achieve faster and more accurate tracking results. Moreover, the proposed method can be applied not only to manipulators with redundancy (including both functional and intrinsic redundancy), but also to manipulators that are non-redundant. Furthermore, it is easier to incorporate a secondary optimization objective into consideration for the intrinsic redundant case, which is difficult for existing methods. The proposed method has been implemented on both simulations and experiments. The numerous simulation and experimental results validate the effectiveness and advantages of the proposed method. Note to Practitioners—The trajectory tracking control with IOC is important and can be applied in a wide spectrum of robotic-assisted manufacturing, e.g. arc-welding, engraving, etc.. In this paper, a perpendicular curve-based trajectory tracking method is proposed to consider the IOC automatically. It is no longer necessary to carefully plan a reachable orientation trajectory. Users only need to focus the planning of the tool direction, which is determined by the tasks. In addition, the proposed method can achieve faster and more accurate tracking result. Moreover, it is universal and can be applied to both redundant and non-redundant cases.
Gaofeng Li, Shan Xu 0002, Dezhen Song, Fernando Caponetto, Ioannis Sarakoglou, Jingtai Liu, Nikolaos G. Tsagarakis
IEEE Trans Autom. Sci. Eng.1
2021 GAIA: A System for Interactive Analysis on Distributed Graphs Using a High-Level Language
Zhengping Qian, Chenqiang Min, Longbin Lai, Gaofeng Li, Youyang Yao, Bingqing Lyu, Jingren Zhou 0001
NSDI5
2020 A Novel Orientability Index and the Kinematic Design of the RemoT-ARM: A Haptic Master with Large and Dexterous Workspace
abstract
Orientability is an important performance index to evaluate the dexterity of haptic master devices. Currently, most of the existing haptic master devices have limited workspace and limited dexterity. In this paper, we present the RemoT-ARM, a 6 Degree-of-Freedom (DOF) haptic master device that can provide larger and more dexterous workspace for operators. To evaluate its reachability of orientations, we propose a novel orientability index. Furthermore, a relative orientability index is proposed to characterize the matching degree of the workspace of a given manipulator to its target workspace. The volume, the manipulability and the condition number are also introduced as performance indices to evaluate the size and the isotropy of the workspace. According to these performance indices, all possible configurations for the RemoT-ARM have been taken into consideration, analyzed, and compared to finalize its optimal configuration.
Gaofeng Li, Edoardo Del Bianco, Fernando Caponetto, Vasiliki-Maria Katsageorgiou, Nikolaos G. Tsagarakis, Ioannis Sarakoglou
ICRA1
2018 Svega: Answering Natural Language Questions over Knowledge Base with Semantic Matching
abstract
Nowadays, more and more large scale knowledge bases are available for public access.Although these knowledge bases have their inherent access interfaces, such as SPARQL, they are generally unfriendly to end users.An intuitive way to bridge the gap between users and knowledge bases is to enable users to ask questions with natural language interface and return desired answers directly.Here the challenge is how to discover the query intention of users.Another challenge is how to obtain accurate answers from knowledge bases.In this paper, we model the query intention with a graph based on an entity-driven method.Consequently, the core problem of natural language question answering can be treated as subgraph matching over knowledge bases.For a query graph, there is a huge number of candidate mappings in a knowledge base, including ambiguities.Thus, a semantic vector is proposed to address disambiguation by evaluating the semantic similarity between edges in a query graph and paths in a knowledge base.By this way, our system can extract accurate answers directly without any offline work.Extensive experiments over the series of QALD challenges show the effectiveness of our system Svega in terms of recall and precision against other state-of-the-art systems.
Gaofeng Li, Pingpeng Yuan, Hai Jin 0001
SEKE1