EDBT 2026 Demo / reviewers in the wild / expert
Hongsheng He
dblp:17/7249
· DBLP profile ↗
22ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0002-2810-865XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 1 first-author · 4 since 2021Systems, architecture and hardware · 10 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorHuman-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Neurorobotic Interface for Finger Joint Angle Estimation: A Multi-Stage CNN-LSTM Network with Transfer LearningabstractTo maximize the autonomy of individuals with upper limb amputations in daily activities, leveraging forearm muscle information to infer movement intent is a promising research direction. While current prosthetic hand technologies can utilize forearm muscle data to achieve basic movements such as grasping, accurately estimating finger joint angles remains a significant challenge. Therefore, we propose a Multi-Stage Cascade Convolutional Neural Network with Long Short-Term Memory Network, where an upsampling module is introduced before the downsampling module to enhance model generalization. Additionally, we designed a transfer learning (TL) framework based on parameter freezing, where the pre-trained downsampling module is fixed, and only the upsampling module is updated with a small amount of out-ofdistribution data to achieve TL. Furthermore, we compared the performance of unimodal and multimodal models, collecting surface electromyography (sEMG) signals, brightness mode ultrasound images (B-mode US images), and motion capture data simultaneously. The results show that on the validation set, the US image had the lowest error, while on the prediction set, the four-channel sEMG achieved the lowest error. The performance of the multimodal model in both datasets was intermediate between the unimodal models. On the prediction set, the average normalized root mean square error values for the four-channel sEMG, US images, and sensor fusion models across three subjects were 0.170,0.203, and 0.186, respectively. By utilizing advanced sensor fusion techniques and TL, our approach can reduce the need for extensive data collection and training for new users, making prosthetic control more accessible and adaptable to individual needs. Hongsheng He, Wan Shou, Qiang Zhang 0028 |
ICRA | 4 |
| 2025 | Controlled Robot Language with Frame Semantics (FrameCRL) for Autonomous Context-Aware High-Level PlanningabstractThis paper proposes a configurable and scalable framework based on Controlled Robot Language with Frame Semantics (FrameCRL) for plan generation. Given natural language instructions, FrameCRL constructs an equivalent formal semantic formulation in the form of discourse representation structures (DRS). Imperative verbs are extracted from the semantic structures as keys to anchor relevant semantic frames from FrameNet, and the selected semantic frames are used to construct goal statements in planning language. Non-imperative statements are further analyzed to generate object specifications and the initial state of the planning problem. These generated statements are then merged into a single planning script, which can be solved directly by the integrated planner. The performance of FrameCRL was evaluated on various natural language corpora and compared with large language models (LLM) based methods in plan generation. The results demonstrated the outperformance of FrameCRL in generating high-quality plans and its capability to handle large context scenarios. The FrameCRL was also tested on pick-and-place tasks using a dual-arm robot and it showcased a robust performance in linguistic understanding. Dang M. Tran, Fujian Yan, Qiang Zhang 0028, Yinlong Zhang, Hongsheng He |
ICRA | 5 |
| 2021 | Learning Task-Oriented Dexterous Grasping from Human KnowledgeabstractIndustrial automation requires robot dexterity to automate many processes such as product assembling, packaging, and material handling. The existing robotic systems lack the capability to determining proper grasp strategies in the context of object affordances and task designations. In this paper, a framework of task-oriented dexterous grasping is proposed to learn grasp knowledge from human experience and to deploy the grasp strategies while adapting to grasp context. Grasp topology is defined and grasp strategies are learned from an established dataset for task-oriented dexterous manipulation. To adapt to various grasp context, a reinforcement-learning based grasping policy was implemented to deploy different task-oriented strategies. The performance of the system was evaluated in a simulated grasping environment by using an AR10 anthropomorphic hand installed in a Sawyer robotic arm. The proposed framework achieved a hit rate of 100% for grasp strategies and an overall top-3 match rate of 95.6%. The success rate of grasping was 85.6% during 2700 grasping experiments for manipulation tasks given in natural-language instructions. Yinlong Zhang, Yanan Li 0001, Hongsheng He |
ICRA | 4 |
| 2021 | Comprehension of Spatial Constraints by Neural Logic Learning from a Single RGB-D ScanabstractAutonomous industrial assembly relies on the precise measurement of spatial constraints as designed by computer-aided design (CAD) software such as SolidWorks. This paper proposes a framework for an intelligent industrial robot to understand the spatial constraints for model assembly. An extended generative adversary network (GAN) with a 3D long short-term memory (LSTM) network was designed to composite 3D point clouds from a single RGB-D scan. The spatial constraints of the segmented point clouds are identified by a neural-logic network that incorporates general knowledge of spatial constraints in terms of first-order logic. The model was designed to comprehend a complete set of spatial constraints that are consistent with industrial CAD software, including left, right, above, below, front, behind, parallel, perpendicular, concentric, and coincident relations. The accuracy of 3D model composition and spatial constraint identification was evaluated by the RGB-D scans and 3D models in the ABC dataset. The proposed model achieved 57.23% intersection over union (IoU) in 3D model composition, and over 99% in comprehending all spatial constraints. Fujian Yan, Dali Wang, Hongsheng He |
IROS | 3 |
| 2020 | MagicHand: Context-Aware Dexterous Grasping Using an Anthropomorphic Robotic HandabstractUnderstanding of characteristics of objects such as fragility, rigidity, texture and dimensions facilitates and innovates robotic grasping. In this paper, we propose a context- aware anthropomorphic robotic hand (MagicHand) grasping system which is able to gather various information about its target object and generate grasping strategies based on the perceived information. In this work, NIR spectra of target objects are perceived to recognize materials on a molecular level and RGB-D images are collected to estimate dimensions of the objects. We selected six most used grasping poses and our system is able to decide the most suitable grasp strategies based on the characteristics of an object. Through multiple experiments, the performance of the MagicHand system is demonstrated. Jindong Tan, Hongsheng He |
ICRA | 3 |
| 2020 | Robotic Understanding of Spatial Relationships Using Neural-Logic LearningabstractUnderstanding spatial relations of objects is critical in many robotic applications such as grasping, manipulation, and obstacle avoidance. Humans can simply reason object's spatial relations from a glimpse of a scene based on prior knowledge of spatial constraints. The proposed method enables a robot to comprehend spatial relationships among objects from RGB-D data. This paper proposed a neural-logic learning framework to learn and reason spatial relations from raw data by following logic rules on spatial constraints. The neural-logic network consists of three blocks: grounding block, spatial logic block, and inference block. The grounding block extracts high-level features from the raw sensory data. The spatial logic blocks can predicate fundamental spatial relations by training a neural network with spatial constraints. The inference block can infer complex spatial relations based on the predicated fundamental spatial relations. Simulations and robotic experiments evaluated the performance of the proposed method. Fujian Yan, Dali Wang, Hongsheng He |
IROS | 3 |
| 2018 | Spatial Calibration for Thermal-RGB Cameras and Inertial Sensor SystemabstractThe light-weight thermal-RGB-inertial sensing units are now gaining increasing research attention, due to their heterogeneous and complementary properties. A robust and accurate registration between a thermal-RGB camera and an inertial sensor is a necessity for effective thermal-RGB-inertial fusion, which is an indispensable procedure for reliable tracking and mapping tasks. This paper presents an accurate calibration method to geometrically correlate the spatial relationships between an RGB camera, a thermal camera and an inertial measurement unit (IMU). The calibration proceeds within the unified calibration framework (thermal-to-RGB, RGB-to-IMU). The extrinsic parameters are estimated by jointly optimizing both the chessboard corner reprojection errors and acceleration and angular velocity error terms. Extensive evaluations have been performed on the collected thermal-RGB-inertial measurements. In this experiments study, the average RMS translation and Euler angle errors are less than 6 mm and 0.04 rad respectively under 20% artificial noise. Yan Li 0194, Jindong Tan, Yinlong Zhang, Wei Liang 0001, Hongsheng He |
ICPR | 5 |
| 2018 | Learning Robotic Grasping Strategy Based on Natural-Language Object DescriptionsabstractGiven the description of an object, s physical attributes, humans can determine a proper strategy and grasp an object. This paper proposes an approach to determine grasping strategy for an anthropomorphic robotic hand simply based on natural-language descriptions of an object. A learning-based approach is proposed to help a robotic hand learn suitable grasp poses starting from the natural language description of the object. Object features are parsed from natural-language descriptions by using a customized natural-language processing technique. The most likely grasp type for the given object is learned from the human grasping taxonomy based on the parsed features. The grasping strategy generated by the proposed approach is evaluated both by simulation study and execution of the grasps on an AR10 robotic hand. Achyutha Bharath Rao, Krishna Krishnan, Hongsheng He |
IROS | 3 |
| 2018 | Wearable Heading Estimation for Motion Tracking in Health Care by Adaptive Fusion of Visual-Inertial MeasurementsabstractThe increasing demand for health informatics has become a far-reaching trend in the ageing society. The utilization of wearable sensors enables monitoring senior people daily activities in free-living environments, conveniently and effectively. Among the primary health-care sensing categories, the wearable visual-inertial modality for human motion tracking gradually exerts promising potentials. In this paper, we present a novel wearable heading estimation strategy to track the movements of human limbs. It adaptively fuses inertial measurements with visual features following locality constraints. Body movements are classified into two types: general motion (which consists of both rotation and translation). or degenerate motion (which consists of only rotation). A specific number of feature correspondences between camera frames are adaptively chosen to satisfy both the feature descriptor similarity constraint and the locality constraint. The selected feature correspondences and inertial quaternions are employed to calculate the initial pose, followed by the coarse-to-fine procedure to iteratively remove visual outliers. Eventually, the ultimate heading is optimized using the correct feature matches. The proposed method has been thoroughly evaluated on the straight-line, rotatory and ambulatory movement scenarios. As the system is lightweight and requires small computational resources, it enables effective and unobtrusive human motion monitoring, especially for the senior citizens in the long-term rehabilitation. Yinlong Zhang, Wei Liang 0001, Hongsheng He, Jindong Tan |
IEEE J. Biomed. Health Informatics | 3 |
| 2017 | Kinematic chain based multi-joint capturing using monocular visual-inertial measurementsabstractCombining light-weight visual and inertial modalities for motion capturing has been popular in robotics researches. There exist scale ambiguity, inaccurate pose estimation with little or no baseline, incremental drifts over time in visual-inertial fusion. Thus, in this paper, we propose a robust motion capturing method based on the multi-joint kinematic chain using monocular visual-inertial sensors. Our method is able to recover monocular visual scale through the joint geometry constraint. Additionally, we take inertial pre-integration to assist visual outlier removal using Maximum A Posteriori method. Ultimately, the kinematic chain model is leveraged to constrain the associated multiple visual-inertial estimation drifts during long time tracking. In the experiments, we conduct multi-joint capturing on a robotic arm. The quality of motion reconstruction is evaluated by comparing the estimated results with the measurements from an optical motion tracking system OptiTrack. Yinlong Zhang, Wei Liang 0001, Hongsheng He, Jindong Tan |
IROS | 3 |
| 2016 | DietCam: Multiview Food Recognition Using a Multikernel SVMabstractFood recognition is a key component in evaluation of everyday food intakes, and its challenge is due to intraclass variation. In this paper, we present an automatic food classification method, DietCam, which specifically addresses the variation of food appearances. DietCam consists of two major components, ingredient detection and food classification. Food ingredients are detected through a combination of a deformable part-based model and a texture verification model. From the detected ingredients, food categories are classified using a multiview multikernel SVM. In the experiment, DietCam presents reliability and outperformance in recognition of food with complex ingredients on a database including 15,262 food images of 55 food types. Hongsheng He, Jindong Tan |
IEEE J. Biomed. Health Informatics | 1 |
| 2015 | DietCam: Multi-view regular shape food recognition with a camera phone
Hongsheng He, Hollie A. Raynor, Jindong Tan |
Pervasive Mob. Comput. | 2 |
| 2015 | Wearable Ego-Motion Tracking for Blind Navigation in Indoor EnvironmentsabstractThis paper proposes an ego-motion tracking method that utilizes visual-inertial sensors for wearable blind navigation. The unique challenge of wearable motion tracking is to cope with arbitrary body motion and complex environmental dynamics. We introduce a visual sanity check to select accurate visual estimations by comparing visually estimated rotation with measured rotation by a gyroscope. The movement trajectory is recovered through adaptive fusion of visual estimations and inertial measurements, where the visual estimation outputs motion transformation between consecutive image captures, and inertial sensors measure translational acceleration and angular velocities. The frame rates of visual and inertial sensors are different, and vary with respect to time owning to visual sanity checks. We hence employ a multirate extended Kalman filter (EKF) to fuse visual and inertial estimations. The proposed method was tested in different indoor environments, and the results show its effectiveness and accuracy in ego-motion tracking. Hongsheng He, Yan Li 0194, Jindong Tan |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2015 | Recognition of Car Makes and Models From a Single Traffic-Camera ImageabstractThis paper proposes the recognition framework of car makes and models from a single image captured by a traffic camera. Due to various configurations of traffic cameras, a traffic image may be captured in different viewpoints and lighting conditions, and the image quality varies in resolution and color depth. In the framework, cars are first detected using a part-based detector, and license plates and headlamps are detected as cardinal anchor points to rectify projective distortion. Car features are extracted, normalized, and classified using an ensemble of neural-network classifiers. In the experiment, the performance of the proposed method is evaluated on a data set of practical traffic images. The results prove the effectiveness of the proposed method in vehicle detection and model recognition. Hongsheng He, Zhenzhou Shao, Jindong Tan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2014 | Ambient motion estimation in dynamic scenes using wearable visual-inertial sensorsabstractThis paper proposes a method to estimate the motion of ambient objects including translational and rotational velocities by moving observers with hybrid visual-inertial sensors. Ambient motion is recovered from visual optical flows that represent ego and ambient dynamics. In this paper, each moving object is considered as a rigid body that has been segmented from the background using computer vision algorithms. In motion recovery, the fundamental challenge is to resolve the coupling between scene depths and translational velocities. Ambient rotational velocities are obtained following a depth-independent bilinear constrain. The scales of ambient trans-lational velocities is computed using the proposed dynamics constraint with an assumption that ambient accelerations are negligible. A fix-point optimization scheme is further introduced to iteratively refine the recoveries of translational and rotational ambient motion until an expected precision is achieved or the maximal iteration is reached. During the optimization, translational ambient motion is precisely recovered and translational ambient motion is rescaled to the canonical amplitude. The results of the simulation study show the effectiveness of the proposed method in motion analysis and prediction. Hongsheng He, Jindong Tan |
ICRA | 1 |
| 2014 | Geometry constrained sparse embedding for multi-dimensional transfer function design in direct volume renderingabstractDirect volume rendering (DVR) is commonly employed for the medical visualization. Multi-dimensional transfer functions are used in DVR to emphasize the region of interest in details. However, it is impractical to interact directly with the functions in more than three dimension. This paper proposes a novel framework called geometry constrained sparse embedding (GCSE) for dimensionality reduction (DR). GCSE allows the conventional DR methods to be applied to a dictionary with much smaller atoms instead. The mapping derived from the dictionary feeds to the original features to obtain the ones in the reduced dimension. To obtain a good dictionary, the intrinsic structure of features is encoded in the sparse embedding based on a geometry distance. In addition, stochastic gradient descent algorithm is employed to speed up the dictionary learning. Various experiments have been conducted using both synthetic and real CT data sets. Compared with conventional methods, GCSE not only produces the comparable results, but also performs well with the capability to handle the large data set more powerfully. The rendering results using the real CT data has demonstrated the effectiveness of GCSE. Zhenzhou Shao, Hongsheng He, Jindong Tan |
ICRA | 3 |
| 2013 | Bottom-up saliency detection for attention determination
Shuzhi Sam Ge, Hongsheng He, Zhengchen Zhang |
Mach. Vis. Appl. | 2 |
| 2012 | Mutual-reinforcement document summarization using embedded graph based sentence clustering for storytelling
Zhengchen Zhang, Shuzhi Sam Ge, Hongsheng He |
Inf. Process. Manag. | 3 |
| 2012 | Geometrically local embedding in manifolds for dimension reduction
Shuzhi Sam Ge, Hongsheng He, Chengyao Shen |
Pattern Recognit. | 2 |
| 2011 | Visual Cortex Inspired Junction Detection
Shuzhi Sam Ge, Chengyao Shen, Hongsheng He |
ICONIP (1) | 3 |
| 2011 | Robust line detection using two-orthogonal direction image scanning
Shuzhi Sam Ge, Hongsheng He |
Comput. Vis. Image Underst. | 3 |
| 2009 | Real-time face detection for human robot interactionabstractFace detection plays an important role in developing human-robot interaction (HRI) for social robots to recognize people. In this paper, we introduce an intelligent vision system that is able to detect human face from background and filter out all the non-face but face-like images. The human face is detected using Ada boost-based Haar-Cascade classifier and the real human face detection is improved using extreme learning machine (ELM). The proposed robot vision system for human detection is tested through realtime experiments. Yaozhang Pan, Shuzhi Sam Ge, Hongsheng He |
RO-MAN | 3 |