Jingtai Liu

dblp:51/5334 · DBLP profile ↗
← Back
41ranked-venue papers
1as first author
27since 2021 · last 2026
0000-0003-2645-5655ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 1 first-author · 19 since 2021Systems, architecture and hardware · 16 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021
YearPublicationVenuePosition
2026 HTMNet: A hybrid transformer-mamba network for LiDAR-based 3D detection and semantic segmentation
Jinzheng Guang, Yongru Wang, Zhenzhong Cao, Jingtai Liu
Expert Syst. Appl.6
2026 FPC-VLA: A vision-language-action framework with a supervisor for failure prediction and correction
Zhixiang Duan, Tianshi Xie, Fuyu Cao, Pinxi Shen, Peili Song, Chenyang Zhao 0009, Piaopiao Jin, Guokang Sun, Shaoqing Xu, Yangwei You, Jingtai Liu
Expert Syst. Appl.12
2026 PipeCLIP: Defect-Conditioned and Cross-Focus-Driven Vision-Language Model for Video-Based Sewer Defect Inspection
abstract
In recent years, vision-language models (VLMs), such as CLIP, have excelled in the visual domain due to the availability of vast paired image-text data. However, directly adapting CLIP to the sewer defect inspection achieves a poor performance, since the vocabularies of defect category are highly “unfamiliar” to CLIP, resulting in the weak representations among the defect categories in the text embedding space. Besides, the substantial differences of pipe characteristics are challenging to align the multiple pairs of visual and text features across multi-focus segments. We propose PipeCLIP for adapting the text information to the video-based multi-label sewer defect classification, which is the first method to integrate a CLIP-based model into the sewer defect inspection. First, expert prior descriptions (EPD) are introduced to differentiate the distinctions between defect category in the text embedding space. Second, defect-attribute coupling (DAC) prompt is proposed to strengthen the coupling relationship between the sewer defects and pipe multi-attributes. The two prompts proposed above together with the traditional category prompt, collectively constitute defect-conditioned text prompt (DecTP) for CLIP. Then, cross-focus temporal (CFT) module is designed to integrate feature information from different focal length, strengthening the visual-text alignment across multi-focus segments. Extensive experiments are conducted on the public benchmark, in which the superiority of PipeCLIP is demonstrated compared with the state-of-the-art methods. Code is available at: https://anonymous.4open.science/r/PipeCLIP-0925.
Chenyang Zhao 0009, Chuanfei Hu, Zhenzhong Cao, Yinuo Song, Jingtai Liu
IEEE Trans Autom. Sci. Eng.5
2026 MDSLCA: Multi-Scale Dilated Spatial and Local Channel Attention for LiDAR Point Cloud Semantic Segmentation
Jinzheng Guang, Qianyi Zhang, Jingtai Liu
IEEE Trans. Circuits Syst. Video Technol.3
2025 GA-TEB: Goal-Adaptive Framework for Efficient Navigation Based on Goal Lines
abstract
In crowd navigation, the local goal plays a crucial role in trajectory initialization, optimization, and evaluation. Recognizing that when the global goal is distant, the robot's primary objective is avoiding collisions, making it less critical to pass through the exact local goal point, this work introduces the concept of goal lines, which extend the traditional local goal from a single point to multiple candidate lines. Coupled with a topological map construction strategy that groups obstacles to be as convex as possible, a goal-adaptive navigation framework is proposed to efficiently plan multiple candidate trajectories. Simulations and experiments demonstrate that the proposed GA-TEB framework effectively prevents deadlock situations, where the robot becomes frozen due to a lack of feasible trajectories in crowded environments. Additionally, the framework greatly increases planning frequency in scenarios with numerous non-convex obstacles, enhancing both robustness and safety.
Qianyi Zhang, Wentao Luo, Yaoyuan Wang, Jingtai Liu
ICRA5
2025 ELPTNet: An Efficient LiDAR-based 3D Pedestrian Tracking Network for Autonomous Navigation Social Robots
abstract
Autonomous navigation social robots need to track pedestrian movements in real-time with high precision to optimize path planning and avoid collisions. However, the main challenge of pedestrian tracking lies in the significant variations in human posture, which differ from rigid-body structures like vehicles. In this paper, we propose an Efficient LiDAR-based 3D Pedestrian Tracking Network (ELPTNet). First, our ELPTNet employs a 3D object detector to extract directional 3D pedestrian bounding boxes from LiDAR point clouds. Then, our ELPTNet employs a Constant Acceleration (CA) model and prediction confidence for target trajectory prediction. During the data association process, it integrates geometric, appearance, and motion features to enhance the robustness and real-time performance of 3D MOT when targets are temporarily occluded. Experimental results demonstrate that our ELPTNet achieves the highest ranking on the large-scale JRDB dataset for the 3D tracking task, outperforming previous state-of-the-art (SOTA) methods with improvements of 8.4% in MOTA and 6.6% in HOTA. Additionally, our ELPTNet attains an inference speed of 61 frames per second (FPS) on a single CPU. Therefore, our method enables accurate and real-time tracking of multiple pedestrians. The code is publicly available at https://github.com/jinzhengguang/ELPTNet.
Jinzheng Guang, Zhenzhong Cao, Yinuo Song, Jingtai Liu
IROS4
2025 Simulating Automotive Radar with Lidar and Camera Inputs
abstract
Low-cost millimeter automotive radar has received more and more attention due to its ability to handle adverse weather and lighting conditions in autonomous driving. However, the lack of quality datasets hinders research and development. We report a new method that is able to simulate 4D millimeter wave radar signals including pitch, yaw, range, and Doppler velocity along with radar signal strength (RSS) using camera image, light detection and ranging (lidar) point cloud, and ego-velocity. The method is based on two new neural networks: 1) DIS-Net, which estimates the spatial distribution and number of radar signals, and 2) RSS-Net, which predicts the RSS of the signal based on appearance and geometric information. We have implemented and tested our method using open datasets from 3 different models of commercial automotive radar. The experimental results show that our method can successfully generate high-fidelity radar signals. Moreover, we have trained a popular object detection neural network with data augmented by our synthesized radar. The network outperforms the counterpart trained only on raw radar data, a promising result to facilitate future radar-based research and development.
Peili Song, Dezhen Song, Enfan Lan, Jingtai Liu
IROS5
2025 MK-Pose: Category-Level Object Pose Estimation via Multimodal-Based Keypoint Learning
abstract
Category-level object pose estimation, which predicts the pose of objects within a known category without prior knowledge of individual instances, is essential in applications like warehouse automation and manufacturing. Existing methods relying on RGB images or point cloud data often struggle with object occlusion and generalization across different instances and categories. This paper proposes a multimodal-based keypoint learning framework (MK-Pose) that integrates RGB images, point clouds, and category-level textual descriptions. The model uses a self-supervised keypoint detection module enhanced with attention- based query generation, soft heatmap matching and graph-based relational modeling. Additionally, a graph-enhanced feature fusion module is designed to integrate local geometric information and global context. MK-Pose is evaluated on CAMERA25 and REAL275 dataset, and is further tested for cross-dataset capability on HouseCat6D dataset. The results demonstrate that MK-Pose outperforms existing state-of-the-art methods in both IoU and average precision without shape priors. Codes will be released at https://github.com/yangyifanYYF/MK-Pose.
Peili Song, Enfan Lan, Jingtai Liu
IROS5
2025 FGI-Gaze: Gaze Target Detection via Filtered Human-Environment Gaze Interaction
Enfan Lan, Chenyang Zhao 0009, Jingtai Liu
PRCV (7)5
2025 DCCLA: Dense Cross Connections With Linear Attention for LiDAR-Based 3D Pedestrian Detection
abstract
LiDAR-based 3D pedestrian detection has recently been extensively applied in autonomous driving and intelligent mobile robots. However, it remains a highly challenging perceptual task due to the sparsity of pedestrian point cloud data and the significant deformation of pedestrian body postures. To address these challenges, we propose a Dense Cross Connections network with Linear Attention (DCCLA), which mitigates the semantic discrepancy between the encoder and decoder of the network by integrating multiple 3D sparse convolutional layers within the skip connections. Furthermore, we enhance these connections by introducing cross-connections, thereby effectively promoting information interaction among various channels. To effectively retain crucial information while summarizing diverse pedestrian representations, we propose the Linear Self-Attention module for 3D point clouds (LSA3D), which significantly reduces model complexity. The experimental results demonstrate that our DCCLA achieves state-of-the-art Average Precision (AP) for the 3D pedestrian detection task on the JRDB large-scale dataset, outperforming the second-ranked method by 2.7% AP. Furthermore, our DCCLA enhances 1.6% mIoU over the benchmark method on the SemanticKITTI dataset. Therefore, our method achieves excellent performance through a cross-scale feature fusion strategy and linear attention that fully combines the advantages of convolution and transformer architectures. The project is publicly available athttps://github.com/jinzhengguang/DCCLA.
Jinzheng Guang, Zhengxi Hu, Qianyi Zhang, Jingtai Liu
IEEE Trans. Circuits Syst. Video Technol.6
2024 UAGE: A Supervised Contrastive Method for Unconstrained Adaptive Gaze Estimation
Enfan Lan, Zhengxi Hu, Jingtai Liu
ACCV (7)3
2024 PS6D: Point Cloud Based Symmetry-Aware 6D Object Pose Estimation in Robot Bin-Picking
abstract
6D object pose estimation holds essential roles in various fields, particularly in the grasping of industrial workpieces. Given challenges like rust, high reflectivity, and absent textures, this paper introduces a point cloud based pose estimation framework (PS6D). PS6D centers on slender and multi-symmetric objects. It extracts multi-scale features through an attention-guided feature extraction module, designs a symmetry-aware rotation loss and a center distance sensitive translation loss to regress the pose of each point to the centroid of the instance, and then uses a two-stage clustering method to complete instance segmentation and pose estimation. Objects from the Siléane and IPA datasets and typical workpieces from industrial practice are used to generate data and evaluate the algorithm. In comparison to the state-of-the-art approach, PS6D demonstrates an 11.5% improvement in ${{\text{F}}_{{1_{inst}}}}$ and a 14.8% improvement in Recall. The main part of PS6D has been deployed to the software of Mech-Mind, and achieves a 91.7% success rate in bin-picking experiments, marking its application in industrial pose estimation tasks.
Qianyi Zhang, Jingtai Liu
IROS4
2024 RPEA: A Residual Path Network with Efficient Attention for 3D pedestrian detection from LiDAR point clouds
Jinzheng Guang, Zhengxi Hu, Qianyi Zhang, Jingtai Liu
Expert Syst. Appl.5
2024 CRATI: Contrastive representation-based multimodal sound event localization and detection
Yongru Wang, Yushan Jiang, Qianyi Zhang, Jingtai Liu
Knowl. Based Syst.5
2024 Improve Computing Efficiency and Motion Safety by Analyzing Environment With Graphics
abstract
Exploring topologically distinctive trajectories provides more options for robot motion planning. Since computing time grows greatly with environment complexity, improving exploration efficiency and picking the optimal trajectory in complex environments are critical issues. To this end, this paper proposes a Graphic-and Timed-Elastic-Band-based approach (GraphicTEB) with spatial completeness and high computing efficiency. The environment is analyzed utilizing computer graphics, where obstacles are extracted as nodes and their relationships are built as edges. Three contributions are presented. 1) By assembling directed detours formed by nodes and segmented paths formed by edges, a generalized path consisting of nodes and edges derives various normal paths efficiently. 2) By multiplying two vectors starting from the obstacle point closest to the waypoint and the boundary point farthest from the waypoint, an novel obstacle gradient is introduced to guide safer optimization. 3) By assigning edges with asymmetric Gaussian model, a trajectory evaluation strategy is designed to reflect the motion tendency and motion uncertainty of dynamic obstacles. Qualitative and quantitative simulations demonstrate that the proposed GraphicTEB achieves spatial completeness, higher scene pass rate, and fastest computing efficiency. Experiments are implemented in long corridor and broad room scenarios, where the robot goes through gaps safely, finds trajectories quickly, and passes pedestrians politelyNote to Practitioners—The motivation stems from the fact that our daily cruising robot occasionally gets trapped in a corridor with piled obstacles or in a complex dynamic crowd due to the lack of a reliable trajectory. The solution is to search for more topologically distinctive trajectories and pick the optimal one. Considering that existing open-source approaches are either incomplete or highly time-consuming, a method for clustering and searching trajectories in the obstacle-occupied regions is proposed to achieve spatial completeness and high computing efficiency. In addition, an optimization technique and a trajectory selection strategy are proposed to improve motion safety. However, at present, the search is incomplete in the temporal-spatial dimension when dynamic obstacle are moving fast. How to perform a complete and fast search in temporal-spatial space will be developed in the future.
Qianyi Zhang, Yuhang Jia, Yuang Xu, Jingtai Liu
IEEE Trans Autom. Sci. Eng.5
2024 Hierarchical Context-Based Emotion Recognition With Scene Graphs
abstract
For a better intention inference, we often try to figure out the emotional states of other people in social communications. Many studies on affective computing have been carried out to infer emotions through perceiving human states, i.e., facial expression and body posture. Such methods are skillful in a controlled environment. However, it often leads to misestimation due to the deficiency of effective inputs in unconstrained circumstances, that is, where context-aware emotion recognition appeared. We take inspiration from the advanced reasoning pattern of humans in perceived emotion recognition and propose the hierarchical context-based emotion recognition method with scene graphs. We propose to extract three contexts from the image, i.e., the entity context, the global context, and the scene context. The scene context contains abstract information about entity labels and their relationships. It is similar to the information processing of the human visual sensing mechanism. After that, these contexts are further fused to perform emotion recognition. We carried out a bunch of experiments on the widely used context-aware emotion datasets, i.e., CAER-S, EMOTIC, and BOdy Language Dataset (BoLD). We demonstrate that the hierarchical contexts can benefit emotion recognition by improving the accuracy of the SOTA score from 84.82% to 90.83% on CAER-S. The ablation experiments show that hierarchical contexts provide complementary information. Our method improves the F1 score of the SOTA result from 29.33% to 30.24% (C-F1) on EMOTIC. We also build the image-based emotion recognition task with BoLD-Img from BoLD and obtain a better emotion recognition score (ERS) score of 0.2153.
Lei Zhou 0017, Zhengxi Hu, Jingtai Liu
IEEE Trans. Neural Networks Learn. Syst.4
2023 GFIE: A Dataset and Baseline for Gaze-Following from 2D to 3D in Indoor Environments
abstract
Gaze-following is a kind of research that requires locating where the person in the scene is looking automatically under the topic of gaze estimation. It is an important clue for understanding human intention, such as identifying objects or regions of interest to humans. However, a survey of datasets used for gaze-following tasks reveals defects in the way they collect gaze point labels. Manual labeling may introduce subjective bias and is labor-intensive, while automatic labeling with an eye-tracking device would alter the person's appearance. In this work, we introduce GFIE, a novel dataset recorded by a gaze data collection system we developed. The system is constructed with two devices, an Azure Kinect and a laser rangefinder, which generate the laser spot to steer the subject's attention as they perform in front of the camera. And an algorithm is developed to locate laser spots in images for annotating 2D/3D gaze targets and removing ground truth introduced by the spots. The whole procedure of collecting gaze behavior allows us to obtain unbiased labels in unconstrained environments semi-automatically. We also propose a baseline method with stereo field-of-view (FoV) perception for establishing a 2D/3D gaze-following benchmark on the GFIE dataset. Project page: https://sites.google.com/view/gfie.
Zhengxi Hu, Yuxue Yang, Xiaolin Zhai, Dingye Yang, Jingtai Liu
CVPR6
2023 Learning Group Residual Representation for Group Activity Prediction*
abstract
The goal of group activity prediction is to infer the group activity involved multiple individuals before it is completely executed. Previous methods focused on capturing pair-wise relationships between individuals, but lacked the exploration of group-wise interactions which can provide global guidance from a macroscopic perspective. To further explore the group-wise interaction, we propose a Group Residual Module (GRM) which constructs a virtual leader node to summarize the group representation and designs a bidirectional message passing mechanism to build the bridge between group and individuals. To capture the spatial-temporal correlation jointly, we propose a Spatial-Temporal Group Residual Network composed of spatial GRMs and temporal GRMs. Different from existing methods that obtain additional information from the complete activity execution, temporal masks in the temporal GRMs are designed to enforce our network to excavate as much discriminative information as possible from the observed activity sequence. Moreover, experimental results show that our network achieves state-of-the-art performance on Volleyball Dataset and Collective Activity Dataset.
Xiaolin Zhai, Zhengxi Hu, Dingye Yang, Jingtai Liu
ICME5
2023 The Human Gaze Helps Robots Run Bravely and Efficiently in Crowds
abstract
In human-aware navigation, the robot tacitly games with humans, balancing safety and efficiency according to human intentions. Poor balance or bad intent recognition causes the robot to stop conservatively or advance rashly, resulting in a deadlock or even a collision respectively. To address the issue, this paper proposes an improved limit cycle for collaboratively parameterizing human intentions and planning robot motions. The human-robot interaction is modeled as a dynamic chicken game with incomplete information, where the human gaze is introduced to depict the unique characteristics of each person, allowing the robot to approach with different safety margins. Our method is tested in challenging indoor scenarios and outperforms traditional methods in both safety and efficiency. We enable robots to utilize human wisdom to solve problems that cannot be solved on their own. The robot bravely goes through oncoming crowds by getting closer to people with higher attention on it and has the foresight to stably cross in front or behind people.
Qianyi Zhang, Zhengxi Hu, Yinuo Song, Jiayi Pei, Jingtai Liu
ICRA5
2023 Advanced acoustic footstep-based person identification dataset and method using multimodal feature fusion
Xiaolin Zhai, Zhengxi Hu, Jingtai Liu
Knowl. Based Syst.5
2023 On Perpendicular Curve-Based Task Space Trajectory Tracking Control With Incomplete Orientation Constraint
abstract
The Incomplete Orientation Constraint (IOC), which does not require a controlled motion constrained by all three spatial directions, exists widely in a lot of robotic tasks. However, the IOC remains a challenge for existing methods, due to the nonlinear structure of the rotation group SO(3). Moreover, the IOCs are time varying in the trajectory tracking problem, which makes it more challenging than the set-point control. To address the IOC problems, we define, identify and prove the closed-form solution of the perpendicular curve in SO(3). Based on the proposed perpendicular curve, we develop a new trajectory tracking controller considering the IOC. Compared with existing methods, the proposed method can achieve faster and more accurate tracking results. Moreover, the proposed method can be applied not only to manipulators with redundancy (including both functional and intrinsic redundancy), but also to manipulators that are non-redundant. Furthermore, it is easier to incorporate a secondary optimization objective into consideration for the intrinsic redundant case, which is difficult for existing methods. The proposed method has been implemented on both simulations and experiments. The numerous simulation and experimental results validate the effectiveness and advantages of the proposed method. Note to Practitioners—The trajectory tracking control with IOC is important and can be applied in a wide spectrum of robotic-assisted manufacturing, e.g. arc-welding, engraving, etc.. In this paper, a perpendicular curve-based trajectory tracking method is proposed to consider the IOC automatically. It is no longer necessary to carefully plan a reachable orientation trajectory. Users only need to focus the planning of the tool direction, which is determined by the tasks. In addition, the proposed method can achieve faster and more accurate tracking result. Moreover, it is universal and can be applied to both redundant and non-redundant cases.
Gaofeng Li, Shan Xu 0002, Dezhen Song, Fernando Caponetto, Ioannis Sarakoglou, Jingtai Liu, Nikolaos G. Tsagarakis
IEEE Trans Autom. Sci. Eng.6
2022 MGTR: End-to-End Mutual Gaze Detection with Transformer
Hang Guo 0002, Zhengxi Hu, Jingtai Liu
ACCV (4)3
2022 Social Aware Multi-modal Pedestrian Crossing Behavior Prediction
Xiaolin Zhai, Zhengxi Hu, Dingye Yang, Lei Zhou 0017, Jingtai Liu
ACCV (4)5
2022 Spatial Temporal Network for Image and Skeleton Based Group Activity Recognition
Xiaolin Zhai, Zhengxi Hu, Dingye Yang, Lei Zhou 0017, Jingtai Liu
ACCV (4)5
2022 P2EG: Prediction and Planning Integrated Robust Decision-Making for Automated Vehicle Negotiating in Narrow Lane with Explorative Game
abstract
In the narrow lane scene of autonomous driving, it is critical for the ego car to recognize the intentions of social vehicles and cooperate with them. However, cooperating with social vehicles is challenging due to insufficient information. This paper proposes an Explorative Game that adopts Participant Game and Perfect Bayesian Equilibrium to exploratively perform some aggressive actions to obtain additional information, thus the autonomous vehicle can cooperate robustly and efficiently. Explorative Game assumes each vehicle maintains a unique belief about the current situation and attributes insecurity and instability to the conflict of various Perfect Bayesian Equilibriums formed by various beliefs. Aggressive actions enable the ego car to proactively guide social vehicles to cooperate as it expects and encourage them to express their intentions as quickly and clearly as possible so that the equilibriums can converge and the conflict can be eliminated. Additional information reduces the error between the actual intentions of social vehicles and the estimated intentions from the ego car, helping rationally prune potential interactions and update parameters of the reward function. We demonstrate our algorithm on recorded data as well as virtual environments with manually controlled social vehicles to prove the efficiency of cooperation and the robustness of decision-making. And it has been running for more than 20 kilometers in the real world.
Qianyi Zhang, Ethan He, Shuguang Ding, Naizheng Wang, Jingtai Liu
IROS6
2022 GCHGAT: pedestrian trajectory prediction using group constrained hierarchical graph attention networks
Lei Zhou 0017, Yingli Zhao, Dingye Yang, Jingtai Liu
Appl. Intell.4
2022 Gaze Target Estimation Inspired by Interactive Attention
abstract
As an essential nonverbal cue, the human gaze reveals human intentions and plays a crucial role in human daily activities. Therefore, automatic detection of the person’s gaze target has drawn the interests of the computer vision community. This is useful not only for identifying whether children are attentive in class but also for locating items of interest to humans in retail settings. Existing gaze-following methods have only explored and exploited the scenes context and the head cues. Considering the significance of human-object interaction in understanding human intentions, we present the Visual-Spatial Graph and introduce a graph attention network to analyze the interaction probability between the human and elements in the scene. Then the interaction probability inferred from the visual-spatial information that is aggregated by the attention mechanism can be transformed into an interactive attention map that depicts the areas people care about. In addition, we construct a transformer as an encoder to integrate the features extracted by the scene and head pathways aiming to decode the gaze target. After introducing interactive attention, our proposed method achieves outstanding performance on two benchmarks: GazeFollow and VideoAttentionTarget. Our code is available athttps://github.com/nkuhzx/VSG-IA.
Zhengxi Hu, Kunxu Zhao, Hang Guo 0002, Yuxue Yang, Jingtai Liu
IEEE Trans. Circuits Syst. Video Technol.7
2020 Evaluation of Lower Leg Muscle Activities During Human Walking Assisted by an Ankle Exoskeleton
abstract
Wearable robots like ankle exoskeletons have demonstrated the capability to enhance human mobility and to reduce biological efforts of human locomotion. The type of assistance provided by ankle exoskeletons could influence the lower leg muscle activities during human walking. This article aimed to systematically evaluate the lower leg muscle activities under different ankle exoskeleton assistance conditions. We measured multiple electromyography-based metrics of five lower leg muscles, while the participants walked with an ankle exoskeleton on a treadmill. Nine assistance conditions, which combined three peak times (46%, 49%, and 52% of stride time) and three peak torque levels (0.3, 0.5, and 0.7 N·m·kg-1), are applied to assist plantarflexion during ankle push-off. Nine healthy subjects participated in the experiments. Of all investigated muscles, the activity level of l.SOL is influenced the most when exoskeleton assistance is applied. The root mean square of l.SOL activity reduces by 33.6 ± 14.0% under one assistance condition compared to walking without the exoskeleton. Our results can be used to guide studies on mechanical and control designs to improve neuromuscular interactions between exoskeletons and wearers.
Wei Wang 0277, Yandong Ji, Jingtai Liu
IEEE Trans. Ind. Informatics5
2019 Transfer Learning Based Wildlife Recognition for Tele-Observation in Field Occlusion Environment
abstract
Timely and credible species recognition of wildlife at their habitats is helpful for ecological monitoring. Thanks to the powerful feature extraction ability of convolutional neural networks(CNN), the performance of species identification has been significantly improved. However the CNN based approach still does not achieve their full potential. The cluttered backgrounds and rich feature changes of wild environment bring great challenges to wildlife recognition. In this paper, we propose a novel approach to improve the anti-occlusion ability of CNN model, which is achieved by training a improved anti-occlusion loss function. The anti-occlusion constraint in our proposed loss function works to reduce the distance of feature expression before and after a sample has been occluded. We validated the effect of occlusion based on public dataset cifar-10, which confirms that its necessary to improve the anti-occlusion ability of CNN model. Based on single-labeled training dataset and generalization dataset, comprehensive comparative evaluation proves that our proposed loss can effectively improve the generalization ability of CNN model during the identification of wildlife.
Yulin Song, Wan Dai, Jingtai Liu
ICIP5
2016 Nonlinear disturbance observer based torque control for series elastic actuator
abstract
This paper presents a practical control approach for series elastic actuators(SEAs) to generate the desired torque. Specifically, the controller is applicable to both linear and nonlinear SEAs and it works well even in the presence of unknown payload parameters and external disturbances. Via the analysis and transformation of the SEA dynamics, a lumped disturbance signal is constructed and a form convenient for controller design is developed. Then, a nonlinear disturbance observer(NDOB) and a sliding-mode control scheme are introduced to synthesize the control law. The performance of the proposed controller is theoretically ensured by Lyapunov analysis. Taking a nonlinear SEA for instance, a series of simulations and hardware experiments are carried out. The results suggest the effectiveness and superior performance of the proposed method for SEA torque control by comparing it with the cascade-PID controller.
Meng Wang 0008, Lei Sun 0001, Wei Yin 0003, Jingtai Liu
IROS5
2016 Impedance control of a cable-driven series elastic actuator with the 2-DOF control structure
abstract
Series elastic actuators (SEAs) are growingly important in physical human-robot interaction (HRI) due to their inherent safety and compliance. Cable-driven SEAs also allow flexible installation and remote torque transmission, etc. However, there are still challenges for the impedance control of cable-driven SEAs, such as the reduced bandwidth caused by the elastic component, and the performance balance between reference tracking and robustness. In this paper, a velocity sourced cable-driven SEA has been set up. Then, a stabilizing 2 degrees of freedom (2-DOF) control approach was designed to separately pursue the goals of robustness and torque tracking. Further, the impedance control structure for human-robot interaction was designed and implemented with a torque compensator. Both simulation and practical experiments have validated the efficacy of the 2-DOF method for the control of cable-driven SEAs.
Wulin Zou, Meng Wang 0008, Jingtai Liu, Ningbo Yu
IROS5
2015 RGB-D Sensors Calibration for Service Robots SLAM
Jingtai Liu, Lei Sun 0001
ICIG (3)2
2015 A haptic shared control algorithm for flexible human assistance to semi-autonomous robots
abstract
Autonomous as well as teleoperated robots find wide applications in various environments. Their capability to accomplish complex and dynamic operations can be significantly improved by fusing human intelligence with autonomous algorithms. In this paper, we propose a haptic shared control algorithm to provide flexible human assistance for semi-autonomous mobile robots. Through the admittance and impedance models, the haptic shared controller smoothly puts together human operator inputs with robot autonomy. Further, the level of autonomy is fully determined by the operator with the grasp motion. A decomposed design has been taken for the autonomous controller of the mobile robot. The algorithm was implemented on the haptic interface omega.7 together with a QBot mobile robot, and its feasibility and efficacy have been validated by experiments.
Ningbo Yu, Jingtai Liu
IROS5
2015 Error aware multiple vertical planes based visual localization for mobile robots in urban environments
Haifeng Li 0008, Hongpeng Wang 0001, Jingtai Liu
Sci. China Inf. Sci.3
2014 Whole-body pose estimation in physical rider-bicycle interactions with a monocular camera and a set of wearable gyroscopes
abstract
We report the development of a human whole-body pose estimation scheme with application to rider-bicycle interactions. The estimation scheme is built on the fusion of measurements of a monocular camera on the bicycle and a set of small wearable gyroscopes attached to the rider's upper- and lower-limb and the trunk. A single feature point is collocated with each wearable gyroscope and also on the segment link where the gyroscope is not attached. An extended Kalman filter is designed to fuse the vision-inertial measurements to obtain accurate whole-body poses. The estimation design also incorporates a set of constraints from human anatomy and the physical rider-bicycle interactions. We demonstrate and compare the performance of the estimation design through multiple subjects riding experiments.
Kaiyan Yu, Yizhai Zhang, Jingang Yi, Jingtai Liu
IROS5
2012 A two-view based multilayer feature graph for robot navigation
abstract
To facilitate scene understanding and robot navigation in a modern urban area, we design a multilayer feature graph (MFG) based on two views from an on-board camera. The nodes of an MFG are features such as scale invariant feature transformation (SIFT) feature points, line segments, lines, and planes while edges of the MFG represent different geometric relationships such as adjacency, parallelism, collinearity, and coplanarity. MFG also connects the features in two views and the corresponding 3D coordinate system. Building on SIFT feature points and line segments, MFG is constructed using feature fusion which incrementally, iteratively, and extensively verifies the aforementioned geometric relationships using random sample consensus (RANSAC) framework. Physical experiments show that MFG can be successfully constructed in urban area and the construction method is demonstrated to be very robust in identifying feature correspondence.
Haifeng Li 0008, Dezhen Song, Jingtai Liu
ICRA4
2009 Modeling and motion stability analysis of skid-steered mobile robots
abstract
Skid-steered mobile robots are widely used because of the simplicity of mechanism and high reliability. However, understanding of the kinematics and dynamics of such a robotic platform is challenging due to the complex wheel/ground interactions and kinematic constraints. In this paper, we attempt to develop a kinematic and dynamic modeling scheme to analyze the skid-steered mobile robot. We model wheel/ground interaction and analyze the robot motion stability. As an application example, we present how to utilize the kinematic and dynamic modeling and analysis for robot localization and slip estimation using only low-cost strapdown inertial measurement units (IMU). The extended Kalman filter (EKF)-based localization scheme incorporates the kinematic constraints. The performance of the EKF-based localization and slip estimation scheme are presented. The estimation methodology is tested and validated on a robotic testbed.
Jingang Yi, Dezhen Song, Suhada Jayasuriya, Jingtai Liu
ICRA6
2009 Kinematic Modeling and Analysis of Skid-Steered Mobile Robots With Applications to Low-Cost Inertial-Measurement-Unit-Based Motion Estimation
abstract
Skid-steered mobile robots are widely used because of their simple mechanism and high reliability. Understanding the kinematics and dynamics of such a robotic platform is, however, challenging due to the complex wheel/ground interactions and kinematic constraints. In this paper, we develop a kinematic modeling scheme to analyze the skid-steered mobile robot. Based on the analysis of the kinematics of the skid-steered mobile robot, we reveal the underlying geometric and kinematic relationships between the wheel slips and locations of the instantaneous rotation centers. As an application example, we also present how to utilize the modeling and analysis for robot positioning and wheel slip estimation using only low-cost strapdown inertial measurement units. The robot positioning and wheel slip-estimation scheme is based on an extended Kalman filter (EKF) design that incorporates the kinematic constraints for accuracy enhancement. The performance of the EKF-based positioning and wheel slip-estimation scheme are also presented. The estimation methodology is tested and validated experimentally on a robotic test bed.
Jingang Yi, Dezhen Song, Suhada Jayasuriya, Jingtai Liu
IEEE Trans. Robotics6
2006 A Criterion for Evaluating Competitive Teleoperation System
abstract
This paper proposes a kind of criterion, which is called degree of satisfaction (DoS). It is utilized to evaluate the competitive teleoperation. We focus on the feather of competitive teleoperation and utilize the criterion to analyze the system. To demonstrate the degree of satisfaction is an effective criterion, a set of competitive teleoperation experiment is designed on TTRP (teleoperation/tele-game robot platform). Experimental results are presented to support our approach
Xingbo Huang, Jingtai Liu, Lei Sun 0001
IROS2
2005 Competitive Multi-robot Teleoperation
abstract
This paper proposes a novel kind of multi-operator multi-robot(MOMR) teleoperation systems - the competitive teleoperation system. Compared with the conventional collaborated MOMR teleoperation system, features and properties of the competitive teleoperation system are presented. Futhermore, major concerns of research and development for this kind of systems are discussed subsequently. Finally, telegame, a kind of Internet-based competitive teleoperation systems, is built as the prototype to support the future research on this aspect and some experimental results are presented to support the discussion.
Jingtai Liu, Lei Sun 0001, Xingbo Huang, Chunying Zhao
ICRA1
2004 Geometry-based Robot Calibration Method
abstract
This paper describes a geometry-based robot calibration method for a 6-DOF robot manipulator. The calibration device only includes a set of light projections consisting of three laser beams. In the proposed calibration algorithm, the coordinates of the laser spots on the table in the world coordinate system is first obtained by processing the image data from a CCD camera fixed above. Based on that, the mapping between the world coordinate system and the robot base coordinate system can then be utilized to locate the robot within its environment by some geometric analysis. The calibration method proposed in the paper is extremely suitable for the fast calibration of a multi-robot cooperation system due to the low cost of the calibration device and the simplicity of the algorithm involved. Some experimental results for a RH6 robot are presented to demonstrate the validity of the proposed calibration method.
Lei Sun 0001, Jingtai Liu, Shuihua Wu, Xingbo Huang
ICRA2