Huixu Dong

dblp:214/8685 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-2582-6728ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Enabling Multiple Grasping Modes: A Retractable and Reconfigurable Robotic Gripper Inspired by Human Finger
abstract
Grippers serve as essential end-effectors, pivotal for facilitating interaction between robots and their external environments. However, existing grippers emerge with primary issues, such as inadequate adaptability, limited grasping range, and narrow functionality. To address these issues, we present a novel under-actuated three-finger gripper driven by a single motor. This gripper accomplishes adaptive, wide-range, and multi-modal grasping through the mechanism retraction and reconfiguration, thereby offering significant advancement in this field. Firstly, inspired by the variable contact area observed in grasps for human hands, the fingers are designed with a variable-length structure that incorporates a guide rail-slider mechanism. This mechanism realizes the automatic adaptive grasping of the gripper. Secondly, a single-motor-driven mechanism capable of both power distribution and reconfiguration has been devised. This mechanism simplifies the control complexity by enabling simultaneous grasping and reconfiguration. By changing the position and posture of the fingers, the gripper can perform multiple grasp modes. Finally, the performance of the gripper is experimentally validated. In particular, a series of grasping trials are undertaken with objects exhibiting a wide range of weights, shapes, and materials. The outcomes illustrate the gripper’s versatility across various grasping cases, highlighting its potential for broad application.
Yuge Chen, Guodong Lu, Huixu Dong
IEEE Trans Autom. Sci. Eng.6
2026 IG-RFT: An Interaction-Guided RL Framework for VLA Models in Long-Horizon Robotic Manipulation
Zhian Su, Weijie Kong, Haonan Dong, Huixu Dong
IEEE Trans Autom. Sci. Eng.4
2026 SA-DEM: Dexterous Extrinsic Robotic Manipulation of Non-Graspable Objects via Stiffness-Aware Dual-Stage Reinforcement Learning
abstract
We propose a novel framework, SA-DEM, for extrinsic non-grasping manipulation of ungraspable objects in robotics. This approach is grounded in dual-stage reinforcement learning and decouples the overall task into two sequential phases: interactive mode decision-making and manipulation action planning. Notably, this framework innovatively incorporates the stiffness information of manipulated objects into the decision-making process, enabling the robot to autonomously perceive, decide, and plan manipulation strategies for objects with diverse physical attributes. The first phase of SA-DEM involves a high-level agent responsible for planning the grasping pose of objects and their interaction locations with the environment, based on the initial state of the objects, observations from environmental point clouds, stiffness representations, and prior knowledge of grasping regions. The second phase is executed by a low-level agent, which focuses on planning specific manipulation actions such as poking and flipping. These actions are derived from autonomous exploration during the training process, negating the need for manual customization. Both agents employ a hybrid discrete-continuous action space along with time-abstracted and spatially grounded representations centered around the point cloud, culminating in a unified actor-critic reinforcement learning framework.
Huixu Dong
IEEE Trans Autom. Sci. Eng.5
2026 Robotic Bin Packing via Hierarchical Reinforcement Learning
abstract
The 3D bin packing problem has aroused enthusiastic research interest in recent years due to its wide range of real-world applications, such as logistics and warehousing. The packing sequence and placement pose (position and orientation) are the primary optimization objectives in bin packing, as they dramatically impact packing results and transportation costs. Existing methods either focus exclusively on sequential or placement decisions, or navigate a vast combined action space of both, generally struggling to reach global optimum. To bridge this research gap, we propose a novel approach that jointly optimizes both objectives. First, we introduce a packing configuration tree to represent the dynamic packing process. Second, we develop a hierarchical reinforcement learning framework which decomposes the problem into two tractable subtasks: a high-level manager network for sequence generation and a low-level worker network for placement determination. Third, we formulate a stepwise reward and train the framework using a Dueling Deep Q-Network. Finally, extensive experiments demonstrate the superiority of the proposed method, which achieves over 10% higher container space utilization compared to the advanced methods while maintaining robust generalization across varied container and box dimensions. Furthermore, real-world robotic packing experiments validate the practical applicability of the proposed method in industrial scenarios.
Baoying Wang, Xidan Zhang, Ziyi Zheng, Weijie Kong, Huixu Dong
IEEE Trans Autom. Sci. Eng.6
2026 Analogy-Augmented Uncertainty-Aware Monocular Visual Odometry
abstract
Visual odometry (VO) is a critical component of autonomous robot systems, enabling precise pose estimation from visual inputs. Learning-based VO methods are increasingly recognized for their robustness in challenging scenarios, including dynamic environments, motion blur, and low-light conditions. However, their performance is constrained by both the diversity of the data and its utilization rate. To overcome these limitations, we propose an end-to-end monocular VO system incorporating a novel learning-based end-to-end VO framework and multiple analogy augmentation strategies. We introduce the Context Attention Uncertainty-aware VO Network (CUVO), which prioritizes semantically rich regions and mitigating interference from high-uncertainty areas to enhance attentional focus and pose estimation accuracy. Furthermore, our analogy augmentation methods—temporal reversal, random rotation, and geometric mirroring—enhance image pairs and compute corresponding true pose transformations, significantly increasing training data quantity and diversity. Simultaneously, an analogous loss is applied to ensure consistency between the original and augmented data. Extensive experiments demonstrate that CUVO significantly enhances VO performance, outperforming previous end-to-end VO methods on TartanAir and KITTI datasets. By leveraging analogy augmentation strategy to expand training data under limited data conditions (27k), zero-shot capability of CUVO degrades by up to 29.5% on TartanAir and 23.3% on KITTI. Our work introduces the first image-to-pose data augmentation method tailored for VO and establishes CUVO as a robust system for advancing learning-based visual odometry.
Jituo Li, Shunwang Sun, Tingxi Xue, Xinqi Liu, Jialu Zhang 0006, Huixu Dong, Guodong Lu
IEEE Trans. Circuits Syst. Video Technol.6
2026 Construction of Generalized Force-Deformation Theoretical Model: Toward Efficient Systematic Optimization of Fin-Ray Effect Grippers
Ziyi Zheng, Huixu Dong
IEEE Trans. Robotics6
2025 Robotic Grasps of Cylindrical and Cubic Objects via Real-Time Learning-Based Shape Detection
abstract
Robots grasping objects are critical capabilities in warehouse environments and industrial settings. A robotic grasp generally occurs in a scenario where it is unfeasible for a worker to efficiently complete a tedious task, such as picking food and drink cans (cylinder-shaped) and packaging boxes (cube-shaped). It is worth noting that the tops of cylinders and cubes can be represented by ellipses and rectangles in the two-dimensional (2D) space, respectively. Therefore, a robot can grasp cylinder-shaped and cube-shaped objects by ellipse and rectangle detection. However, it faces the challenge of how to accurately detect cylindrical and cubic objects in real-time for robot grasping. To tackle the above research problem, we propose a grasping system that enables a robot to grasp cylinder-shaped and cube-shaped objects in static and dynamic environments by the proposed ellipse and rectangle detector. An end-to-end learning model is constructed to first incorporate a one-stage detection backbone and then, accommodate the proposed adaptive multi-branch multi-scale net with a designed iterative feature pyramid network, local inception net, and multi-receptive-field feature fusion net to generate object detection recommendations. Employing depth information, the coordinates of detected objects are converted to the 3D space via sampling a series of registered depths and pixels on objects from the live video stream. Comparisons with recent detection methods on the same dataset indicate that the proposed ellipse and rectangle detectors present better performance. Abundant grasping experiments are conducted to illustrate that a robot, empowered by the proposed detector, has the capability of grasping cylindrical and cubic objects in dynamic scenarios. (Video on YouTube,https://youtu.be/KK1OtW6GvL0). Note to Practitioners—This paper is motivated by the problem of how to enable a robot to grasp objects with the basic geometric primitives-ellipses and rectangles in static and dynamic scenarios. Our target is to provide a potential solution for flexible industrial settings in operating moving cylinder-shaped and cube-shaped objects (food and drink cans and packaging boxes) in dynamic scenarios such as conveyors of production lines and logistics lines. We constructed a supervised learning model that can accurately and quickly detect ellipses and rectangles. Through the verification of the comparisons with recent methods and robotic grasping experiments, the behavior of the proposed method can be used in practical applications. In the future, we will deploy this robotic grasping system based on the proposed perception method to grasp food and drink cans and packaging boxes from moving conveyors on production lines and logistics lines.
Huixu Dong, Jiadong Zhou, Haoyong Yu
IEEE Trans Autom. Sci. Eng.1
2024 Under-actuated Robotic Gripper with Multiple Grasping Modes Inspired by Human Finger
abstract
Under-actuated robot grippers, as a pervasive tool of robots, have become a considerable research focus. Despite their simplicity of mechanical design and control strategy, they suffer from poor versatility and weak adaptability, making widespread applications limited. To better address relevant research gaps, we present a novel 3-finger linkage-based gripper that realizes retractable and reconfigurable multi-mode grasps driven by a single motor. Firstly, inspired by the changes occurred in the contact surface with a human finger moving, we artfully design a slider-slide rail mechanism as the phalanx to achieve retraction of each finger, allowing for better performance in the enveloping grasping mode. Secondly, a reconfigurable structure is constructed to broaden the grasping range of objects’ dimensions for the proposed gripper. By adjusting the configuration and gesture of each finger, the gripper can achieve five grasping modes. Thirdly, the proposed gripper is solely actuated by a single motor, yet it can be capable of grasping and reconfiguring simultaneously. Finally, various experiments on grasps of slender, thin, and large-volume objects are implemented to evaluate the performance of the proposed gripper in practical scenarios, which demonstrates the excellent grasping capabilities of the gripper.
Tingbo Liao, Hassen Nigatu, Guodong Lu, Huixu Dong
IROS6
2024 Discretizing SO(2)-Equivariant Features for Robotic Kitting
abstract
Robotic kitting has attracted considerable attention in logistics and industrial settings. However, existing kitting methods encounter challenges such as low precision and poor efficiency, limiting their widespread applications. To address these issues, we present a novel kitting framework that improves both the precision and computational efficiency of complex kitting tasks. Firstly, our approach introduces a fine-grained orientation estimation technique in the picking module, significantly enhancing orientation precision while effectively decoupling computational load from orientation granularity. This technique combines an SO(2)-equivariant network with a group discretization operation to preciously predict discrete orientation distributions. Secondly, we develop the Hand-Tool Kitting Dataset (HTKD) to evaluate different solutions in handling orientation-sensitive kitting tasks. This dataset comprises a diverse collection of hand tools and synthetically created kits, which reflects the complexities of real-world kitting scenarios. Finally, a series of experiments is conducted to evaluate the performance of the proposed method. The results demonstrate that our approach offers an excellent balance between success rates and computational efficiency in high-precision robotic kitting tasks.
Jiadong Zhou, Yadan Zeng, Huixu Dong, I-Ming Chen 0001
IROS3
2024 Theoretical Modeling and Bio-inspired Trajectory Optimization of A Multiple-locomotion Origami Robot
abstract
Recent research on mobile robots has focused on increasing their adaptability to unpredictable and unstructured environments using soft materials and structures. However, the determination of key design parameters and control over these compliant robots are predominantly iterated through experiments, lacking a solid theoretical foundation. To improve their efficiency, this paper aims to provide mathematics modeling over two locomotion, crawling and swimming. Specifically, a dynamic model is first devised to reveal the influence of the contact surfaces’ frictional coefficients on displacements in different motion phases. Besides, a swimming kinematics model is provided using coordinate transformation, based on which, we further develop an algorithm that systematically plans human-like swimming gaits, with maximum thrust obtained. The proposed algorithm is highly generalizable and has the potential to be applied in other soft robots with similar multiple joints. Simulation experiments have been conducted to illustrate the effectiveness of the proposed modeling.
Keqi Zhu, Hassen Nigatu, Ruihong Dong, Huixu Dong
IROS7
2022 Learning-based Ellipse Detection for Robotic Grasps of Cylinders and Ellipsoids
abstract
In our daily life, there are many objects represented by cylindrical shapes and ellipsoids. The tops of these objects are formed by elliptic shape primitives. Thus, it is available for a robot to manipulate these objects by ellipse detection. In this work, we propose a novel approach to generating ground truth for training the model based on domain randomization. Using synthetic data generated in this manner, we build an end-to-end deep neural network with a detection backbone and then, combine multiple branches archived from the backbone for sharing the multiple-scale features; further, after employing active rotation filters, the features pass through the region proposal net to form the prediction branches of the box, orientation regression, and object classification; finally, these branches are fused to do ellipse detection, allowing robotic manipulations of cylinders and ellipsoids. To demonstrate the capabilities of the proposed detector, we show the comparison results with the state-of-the-art detector on synthetic and public datasets. The proposed model for ellipse detection and data generation pipeline based on domain randomization in a simulation are evaluated by a series of robotic manipulations implemented in real application scenarios. The results illustrate a high success rate on real-world grasp attempts despite having only been trained on a synthetic dataset. (A video of some robotic experiments is available on YouTube: https://youtu.be/Ueg1XSI2S98).
Huixu Dong, Jiadong Zhou, Dilip K. Prasad, I-Ming Chen 0001
ICRA1
2022 Estimation of Upper Limb Kinematics with a Magnetometer-Free Egocentric Visual-Inertial System
abstract
Most human activities in daily living or professional work rely on upper body motion. Measuring upper body motion is essential for many applications such as health evaluation, rehabilitation, human power augmentation, skill transferring, etc. Computer vision-based systems have been widely used to directly capture upper limb motion but are usually constrained in a restricted area. Wearable sensors such as inertial measurement units (IMUs) are promising to enable ambulant and out-of-lab measurements but also suffer from issues such as magnetic distortion and drifting. Some visual-inertial systems have been proposed recently to fuse these two complementary measurements but mostly apply in a restricted area. In this paper, we propose a fully wearable egocentric visual-inertial system to estimate the upper-limb pose. Magnetometers are not used to allow the system to work in complex industrial and daily living scenarios or to be integrated with motorized assistive devices. Methods to automatically calibrate the sensor-to-segment alignment and estimate upper body motion is presented and validated with an optical motion capture system. Experimental results showed the system can estimate the joint angles without drift and obtain accurate wrist position even with occlusion, verifying the efficacy of the proposed system and method.
Huixu Dong, Haoyong Yu
ICRA3
2022 Enabling Massage Actions: An Interactive Parallel Robot with Compliant Joints
abstract
We propose a parallel massage robot with compliant joints based on the series elastic actuator (SEA), offering a unified force-position control approach. First, the kinematic and static force models are established for obtaining the corresponding control variables. Then, a novel force-position control strategy is proposed to separately control the force-position along the normal direction of the surface and another two-direction displacement, without the requirement of a robotic dynamics model. To evaluate its performance, we implement a series of robotic massage experiments. The results demonstrate that the proposed massage manipulator can successfully achieve desired forces and motion patterns of massage tasks, arriving at a high-score user experience.
Huixu Dong, I-Ming Chen 0001
IROS1
2021 Object Pose Estimation via Pruned Hough Forest With Combined Split Schemes for Robotic Grasp
abstract
Robotic grasp in complex open-world scenarios requires an effective and generalizable perception. Estimating object’s pose is needed in a variety of practical grasping scenarios. Here we present a novel approach of pose estimation of textureless and textured objects. The algorithm utilizes a single RGB-D image to exploit depth invariant, oriented point pair feature as well as local contextual sensitivity in cluttered environments. To enhance the performance of the voting process and improve learning efficiency, we employ a global pruning algorithm that reduces the risk of overfitting and simplifies the structure of decision trees after compensating for the complementary information among multiple trees by optimizing a designed global objective function. Finally, we also refine the pose obtained from the above stage. The proposed approach of estimating 6-D (degree of freedom) poses of textured and textureless objects is evaluated on publicly available data sets against the recent works under various conditions. It illustrates that our framework is superior to these recent works. Further, we perform extensive qualitative experiments of robotic grasp to illustrate the proposed approach can be applied to practical scenarios.Note to Practitioners—This article is motivated by the problem of the pose estimation of textured and textureless objects in clutter environments. It is difficult for conventional works to address the issue of estimating textured or textureless objects’ poses in such scenarios. We considered that a novel system should be able to obtain the 6-D poses of objects. Therefore, we investigate the combined use of multiple split functions with different characteristics. Learning the model based on Hough forests always cost much computational resource; therefore, we construct a novel pruned Hough forest for solving this issue. Through the comparison and robotic grasp verifications, the behavior of our system can be used in practical applications. In future, we will deploy the proposed system in robotic assembling tasks.
Huixu Dong, Dilip K. Prasad, I-Ming Chen 0001
IEEE Trans Autom. Sci. Eng.1
2020 Are Object Detection Assessment Criteria Ready for Maritime Computer Vision?
abstract
Maritime vessels equipped with visible and infrared cameras can complement other conventional sensors for object detection. However, application of computer vision techniques in maritime domain received attention only recently. The maritime environment offers its own unique requirements and challenges. Assessment of the quality of detections is a fundamental need in computer vision. However, the conventional assessment metrics suitable for usual object detection are deficient in the maritime setting. Thus, a large body of related work in computer vision appears inapplicable to the maritime setting at the first sight. We discuss the problem of defining assessment metrics suitable for maritime computer vision. We consider new bottom edge proximity metrics as assessment metrics for maritime computer vision. These metrics indicate that existing computer vision approaches are indeed promising for maritime computer vision and can play a foundational role in the emerging field of maritime computer vision.
Dilip K. Prasad, Huixu Dong, Deepu Rajan, Hiok Chai Quek
IEEE Trans. Intell. Transp. Syst.2
2019 CaBot: Designing and Evaluating an Autonomous Navigation Robot for Blind People
abstract
Navigation robots have the potential to overcome some of the limitations of traditional navigation aids for blind people, specially in unfamiliar environments. In this paper, we present the design of CaBot (Carry-on roBot), an autonomous suitcase-shaped navigation robot that is able to guide blind users to a destination while avoiding obstacles on their path. We conducted a user study where ten blind users evaluated specific functionalities of CaBot, such as a vibro-tactile handle to convey directional feedback; experimented to find their comfortable walking speed; and performed navigation tasks to provide feedback about their overall experience. We found that CaBot's performance highly exceeded users' expectations, who often compared it to navigating with a guide dog or sighted guide. Users' high confidence, sense of safety, and trust on CaBot poses autonomous navigation robots as a promising solution to increase the mobility and independence of blind people, in particular in unfamiliar environments.
João Guerreiro 0002, Daisuke Sato 0001, Saki Asakawa, Huixu Dong, Kris Makoto Kitani, Chieko Asakawa
ASSETS4
2019 Real-Time Robotic Manipulation of Cylindrical Objects in Dynamic Scenarios Through Elliptic Shape Primitives
abstract
Robotic manipulation employs the object detection in images to create a scene awareness and locate an object's pose. In dynamic scenarios, fast multiobject detection and tracking are crucial. Many objects commonly found in household and industrial environments are represented by cylindrical shapes. Thus, it is available for robots to manipulate them through the real-time detection of elliptic shape primitives formed by the circular tops of these objects. We devise an efficient algorithm of the detection of elliptic shape primitives, which in turn enables robust and real-time robotic manipulations of such objects. The proposed algorithm incorporates the information of elliptic edge curvature, splits complex curves into arcs, classifies the arcs into different quadrants of a candidate elliptic shape, determines the quality of arc selection for ellipse fitting, and then retrieves the corresponding elliptic shape primitive. Our algorithm provides either faster or more accurate ellipse detection results than the current state-of-the-art methods, irrespective of challenging scenarios such as occluded or overlapping ellipses. This is verified by performance comparison with six state-of-the-art elliptic shape detection algorithms on four public image datasets. The algorithm has been integrated on robots to demonstrate the ability to carry out accurate robotic manipulations (tracking, grasping, and stacking) of cylindrical objects in real time. We show that the robotic manipulator, empowered by the elliptic shape primitive algorithm, performs well in complex manipulation experiments as well as dynamic scenarios.
Huixu Dong, Ehsan Asadi, Guangbin Sun, Dilip K. Prasad, I-Ming Chen 0001
IEEE Trans. Robotics1
2018 Efficient Pose Estimation from Single RGB-D Image via Hough Forest with Auto-Context
abstract
We propose a high efficient learning approach to estimating 6D (Degree of Freedom) pose of the textured or texture-less objects for grasping purposes in a cluttered environment where the objects might be partially occluded. The method comprises three main steps. Given a single RGB-D image, we first deploy appropriate features and the random forest to deduce the object class probability and cast votes for the 6D pose in Hough space by joint regression and classification framework, adopting reservoir sampling and summarizing the pose distribution by clustering. Next, we integrate the auto-context into cascaded Hough forests to improve the efficiency of learning. Extensive experiments on various public datasets and robotic grasps indicate that our method presents some improvements over the state-of-art and reveals the capability for estimating poses in practical applications efficiently.
Huixu Dong, Dilip K. Prasad, Qilong Yuan, Jiadong Zhou, Ehsan Asadi, I-Ming Chen 0001
IROS1
2018 Accurate detection of ellipses with false detection control at video rates using a gradient analysis
Huixu Dong, Dilip K. Prasad, I-Ming Chen 0001
Pattern Recognit.1
2017 Robust ellipse detection via arc segmentation and classification
abstract
In this paper, we propose a novel ellipse detection algorithm for synthetic and real images. Existing ellipse detection methods are too slow when used with limited hardware resources. The proposed method demonstrates the capability of detecting ellipses with an excellent accuracy at an acceptable speed level in three public datasets. The excellent performance is attributed to the novel combination of classification of arcs into different quadrants of a candidate ellipse, edge curvature and convexity-concavity analysis, and an elliptic geometry constraint.
Huixu Dong, I-Ming Chen 0001, Dilip K. Prasad
ICIP1