Qixin Cao

dblp:73/5432 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-5462-9857ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 5 since 2021Systems, architecture and hardware · 11 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DCT-Diffusion: Depth Completion for Transparent Objects with Diffusion Denoising Approach
abstract
Transparent objects are common in industrial automation and daily life. However, accurate visual perception of these objects remains challenging due to their reflective and refractive properties. Most previous studies fail to capture contextual information or typically rely on regression-based methods at the decoder stage, suffering from overfitting and unsatisfactory object details. To overcome these limitations, we present a novel depth completion framework for transparent objects with diffusion denoising approach (DCT-Diffusion). First, we adopt a transformer-based encoder to globally learn the depth relationships from different parts of the input by modeling long-distance dependencies. Then, we propose to introduce the diffusion model to generate refined depth maps from random depth distribution. Through iterative refinement, our model can progressively enhance depth map details and achieves fine-grained performance. Lastly, a conditioned fusion module is developed, which utilizes encoder features as visual conditions and fuses them with the denoising block at each step using augmented attention. Extensive comparative studies and cross-domain experiments prove that the DCT-Diffusion outperforms previous methods and significantly improves the robustness and generalization ability. Moreover, visualization results further illustrate that our method can generate depth maps with more complete geometry and clearer boundaries, achieving satisfactory results.
Zhenning Zhou, Weiqing Shen, Qixin Cao
IROS5
2025 PRIOR-SLAM: Enabling Visual SLAM for Loop Closure Under Large Viewpoint Variations
abstract
Existing visual simultaneous localization and mapping (SLAM) systems struggle with loop closure under significant viewpoint variations, such as revisiting the same place orthogonally or oppositely. This limitation primarily stems from the lack of viewpoint invariance in both macroscopic place definition and microscopic feature description. Based on the crucial insight that geometric information is more viewpoint-invariant than visual information, we create map segments via leveraging the position and coplanarity distribution of map elements within the map constructed by monocular SLAM to overcome the limitations of frame-based place definition and extract perspective-invariant ORB (PRIOR) features via accounting for local surface perspective distortions to enhance perspective invariance without the need for costly perspective invariance estimation or additional depth data. We further utilize map segments and PRIOR features to hierarchically detect and correct loop closures with a coarse-to-fine geometric consistency check. We integrate our novelties into the prevalent SLAM framework, thereby proposing PRIOR-SLAM, which achieves state-of-the-art performance in feature matching and retrieval, visual place recognition, and loop closure under large viewpoint changes.
Weibang Bai, Qixin Cao
IEEE Trans. Robotics6
2023 PanelPose: A 6D Pose Estimation of Highly-Variable Panel Object for Robotic Robust Cockpit Panel Inspection
abstract
In robotic cockpit inspection scenarios, the 6D pose of highly-variable panel objects is necessary. However, the buttons with different states on the panel cause the variable texture and point cloud, which confuses the traditional invariable object pose estimation method. The bottleneck is the variable texture and point cloud. To address this issue, we propose a simple yet effective method denoted as PanelPose that leverages synthetic data and edge-line features. Specifically, we extract edge and line features of RGB images and fuse these feature maps as a multi-feature fusion map (MFF Map) to focus on the shape features of panel objects. Moreover, we design an effective keypoint selection algorithm considering the shape information of panel objects, which simplifies keypoint localization for precise pose estimation. Finally, the panel object pose is estimated via PNP/RANSAC, refined by the multi-state template (MST) and multi-scale ICP. We experimentally show that state-of-the-art 6D pose estimation methods alone are not sufficient to solve the cockpit panel inspection task but that our method significantly improves the performance. In cockpit inspection scenarios, the panel localization error is less than 3mm using our method. Code and data are available at https://github.com/sunhan1997/PaneIPose.
Peiyuan Ni, Qixin Cao
IROS6
2023 Grasp Stability Assessment Through Attention-Guided Cross-Modality Fusion and Transfer Learning
abstract
Extensive research has been conducted on assessing grasp stability, a crucial prerequisite for achieving optimal grasping strategies, including the minimum force grasping policy. However, existing works employ basic feature-level fusion techniques to combine visual and tactile modalities, resulting in the inadequate utilization of complementary information and the inability to model interactions between unimodal features. This work proposes an attention-guided cross-modality fusion architecture to comprehensively integrate visual and tactile features. This model mainly comprises convolutional neural networks (CNNs), self-attention, and cross-attention mechanisms. In addition, most existing methods collect datasets from real-world systems, which is time-consuming and high-cost, and the datasets collected are comparatively limited in size. This work establishes a robotic grasping system through physics simulation to collect a multimodal dataset. To address the sim-to-real transfer gap, we propose a migration strategy encompassing domain randomization and domain adaptation techniques. The experimental results demonstrate that the proposed fusion framework achieves markedly enhanced prediction performance (approximately 10%) compared to other baselines. Moreover, our findings suggest that the trained model can be reliably transferred to real robotic systems, indicating its potential to address real-world challenges.
Zhenning Zhou, Zhinan Zhang, Qixin Cao
IROS6
2021 Multi-Parameter Optimization for a Robust RGB-D SLAM System
abstract
SLAM systems can retrieve their metric scales and depth information using RGB-D cameras. However, limited by the sensing range and objects structure, RGB-D cameras can not always work well, resulting in failures sometimes. In this work, we present initialization and localization methods based on maximum-a-posteriori estimation. Our system endows monocular keypoints with valid depth values and introduce them into bundle adjustment. Depth bias coefficient and scale factor are also optimized in the local window, obtaining robustness in large scale environments and long-running operations. The experimental results indicate that our system provides the best robustness compared with other excellent methods in the literature, being able to process the most challenging sequences in the TUM RGB-D dataset.
Guohan He, Qixin Cao
ICRA4
2021 Learning an end-to-end spatial grasp generation and refinement algorithm from simulation
Peiyuan Ni, Wenguang Zhang, Qixin Cao
Mach. Vis. Appl.4
2020 Visual-IMU State Estimation with GPS and OpenStreetMap for Vehicles on a Smartphone
abstract
In this paper, we propose an approach for ego-motion estimation and vehicle localization using low-cost portable sensors. It combines Visual Inertial Odometry with GPS and street map information obtained from OpenStreetMap. For conventional vehicle driving scenarios, we present a lightweight and targeted multi-sensor fusion method based on graph optimization, and implement the entire system on a smartphone. Extensive experiments on the benchmark dataset and real-world data show that the system could achieve robust, accurate and real-time tracking and localization in complex and dynamic scenarios.
Guohan He, Qixin Cao, Haoyuan Miao
ICARCV2
2020 PointNet++ Grasping: Learning An End-to-end Spatial Grasp Generation Algorithm from Sparse Point Clouds
abstract
Grasping for novel objects is important for robot manipulation in unstructured environments. Most of current works require a grasp sampling process to obtain grasp candidates, combined with local feature extractor using deep learning. This pipeline is time-costly, expecially when grasp points are sparse such as at the edge of a bowl.In this paper, we propose an end-to-end approach to directly predict the poses, categories and scores (qualities) of all the grasps. It takes the whole sparse point clouds as the input and requires no sampling or search process. Moreover, to generate training data of multi-object scene, we propose a fast multi-object grasp detection algorithm based on Ferrari Canny metrics. A single-object dataset (79 objects from YCB object set, 23.7k grasps) and a multi-object dataset (20k point clouds with annotations and masks) are generated. A PointNet++ based network combined with multi-mask loss is introduced to deal with different training points. The whole weight size of our network is only about 11.6M, which takes about 102ms for a whole prediction process using a GeForce 840M GPU. Our experiment shows our work get 71.43% success rate and 91.60% completion rate, which performs better than current state-of-art works.
Peiyuan Ni, Wenguang Zhang, Qixin Cao
ICRA4
2019 Detect in RGB, Optimize in Edge: Accurate 6D Pose Estimation for Texture-less Industrial Parts
abstract
In order to solve robotic bin-picking problem in many industrial applications, accurate 6D object pose estimation is one of fundamental technologies. This paper presents a method for accurate 6D pose estimation from a single RGB image for texture-less industrial parts. These objects are common but still challenging to deal with, due to the fact that poor surface texture and brightness makes difficult to compute discriminative local appearance descriptors. The proposed method mainly consists of two stages, which ranges from the detection stage to the optimization stage. Firstly, all known objects in the RGB image are detected with 2D bounding box via a tiny convolutional neural network. Then, the second stage will optimize the 6D pose in the Edge image given several coarse initializations. These coarse initializations are generated from the Edge image via a hypothesis-evaluation scheme. Furthermore, the proposed method is validated by achieving state-of-the-art results of texture-less industrial parts for RGB input. According to practical experiments, the proposed method is accurate and robust enough to be applied on the robotic manipulation platform to complete a simple assembly task.
Haoruo Zhang, Qixin Cao
ICRA2
2019 Fast Motion Planning via Free C-space Estimation Based on Deep Neural Network
abstract
This paper presents a novel learning-based method for fast motion planning in high-dimensional spaces. A deep neural network is designed to predict the free configuration space rapidly given the environment point cloud. With a generated roadmap as an approximate view of the free C-space, LazyPRM is applied to find and check the path with A* search. Due to the application of LazyPRM, the presented method can preserve probabilistic completeness and asymptotic optimality. The new algorithm is tested on a 3-DOF robot arm and a 6-DOF UR3 robot to plan in randomly generated obstacle environments. Results indicate that compared to planners including PRM, RRT*, RRT-connect and the original LazyPRM, our method is of the lowest time consumption and relatively short path length, showing good performance on both planning speed and path quality.
Qixin Cao, Mingjing Sun, Ganggang Yang
IROS2
2019 Fast 6D object pose refinement in depth images
Haoruo Zhang, Qixin Cao
Appl. Intell.2
2019 Holistic and local patch framework for 6D object pose estimation in RGB-D images
Haoruo Zhang, Qixin Cao
Comput. Vis. Image Underst.2
2015 A Framework for Intelligent Service Environments Based on Middleware and General Purpose Task Planner
abstract
Aiming at providing various services for daily living, a framework of Intelligent Service Environment of Ubiquitous Robotics (ISEUR) is presented. This framework mainly addresses two important issues. First, it builds standardized component models for heterogeneous sensing and acting devices based on the middleware technology. Second, it implements a general purpose task planner, which coordinates associated components to achieve various tasks. The video demonstrates how these two functionalities are combined together in order to provide services in intelligent environments. Two different tasks, a localization task and a robopub task, are implemented to show the feasibility, efficiency and expandability of the system.
Qixin Cao
Intelligent Environments2
2013 Development of a novel gait rehabilitation system based on FES and treadmill-walk for convalescent hémiplégie stroke survivors
abstract
Recently, a large amount of stroke survivors are suffering from motor impairment. However, existed therapy interventions have limited effects to restore normal motor function. Thus, we proposed a novel control strategy for gait rehabilitation of hemiplegic patients. The whole system consists of a Functional Electrical Stimulation (FES) device and Treadmill-Walk system. FES contributes to improve the quality of the gait based on real-time adjustment of gait pattern. During gait, the electrical stimuli from separate output channels of an FES device are launched to stimulate two lower extremity muscles (Tibialis Anterior (TA) and Hamstrings). Stimulus launching procedure is based on identifying subject's gait state (stance and swing phases). According to the current variation of treadmill motor, gait phase and muscle activation of lower limbs can be determined during walking on Treadmill-Walk. Three able-bodied subjects simulated hemiplegic patients in the experiment. The results indicated that the proposed method is a safe, feasible and promising intervention.
Jing Ye 0005, Yasutaka Nakashima, Takao Watanabe, Masatoshi Seki, Bo Zhang 0028, Quanquan Liu 0001, Yuki Yokoo, Yo Kobayashi, Qixin Cao, Masakatsu G. Fujie
IROS9
2011 A novel ant-based clustering algorithm using the kernel method
Lei Zhang 0131, Qixin Cao
Inf. Sci.2
2010 The facial texture analysis for the automatic portrait drawing
Zhuang Fu, Qixin Cao, Yanzheng Zhao
Pattern Recognit.3
2008 A New Method for Facial Features Quantification of Caricature Based on Self-Reference Model
abstract
Some facial features that differ from an ordinary face should be identified by a computer when generating a facial caricature. These distinctive facial features are called self-features. Compared with traditional Mean Face Model (MFM) that is unable to quantify these self-features well, a Self-Reference Model (SRM) is presented in this paper. Firstly, based on the physiology structure of a front face, a self-reference is found, and this reference is used to measure the self-features. According to the self-reference, some standard facial parameters are worked out by collecting statistic data of many facial images. Then, in an input face image, by evaluating some differences between the input face and the standard facial parameters, the self-features are properly estimated and quantified. Finally, by analyzing some caricatures produced by caricaturists, the SRM can prove the validity of the proposed Algorithm.
Zhuang Fu, Qixin Cao, Yanzheng Zhao
Int. J. Pattern Recognit. Artif. Intell.3
2006 An Evolutionary Artificial Potential Field Algorithm for Dynamic Path Planning of Mobile Robot
abstract
The artificial potential field (APF) method is widely used for autonomous mobile robot path-planning due to its simplicity and mathematical elegance. However, most researches are focused on solving the path-planning problem in a stationary environment, where both targets and obstacles are stationary. This paper proposes a new APF method for path-planning of mobile robots in a dynamic environment where the target and obstacles are moving. First, the new force function and the relative threat coefficient function are defined. Then, a new APF path-planning algorithm based on the relative threat coefficient is presented. Finally, computer simulation and experiment are used to demonstrate the effectiveness of the dynamic path-planning scheme
Qixin Cao, Yanwen Huang, Jingliang Zhou
IROS1
2006 Wall-climbing Robot Path Planning for Testing Cylindrical Oilcan Weld Based on Voronoi Diagram
abstract
To improve the efficiency of the weld testing for a cylindrical oilcan performed by the wall-climbing robot, this paper gives an algorithm for generating the Voronoi diagram from the points set on a cylinder by modification process. Based on this algorithm the paper also provides a method about the design of cylindrical tank wallboards and the weld testing path planning from Delaunay triangulation. A software simulation platform is also developed. The simulation results show that the method is effective to the stand cylindrical tank design and the wall-climbing robot weld testing path planning
Zhuang Fu, Yanzheng Zhao, Zhi-yuan Qian, Qixin Cao
IROS4
2006 Operation Principle of a Bend Enhanced Curvature Optical Fiber Sensor
abstract
A curvature optical fiber sensor is reported in this paper. The curvature measurement sensitivity is improved using bend enhanced method. The operation principle of this intensity modulate macro-bend curvature optical fiber sensor is proposed based on light scattering theory: the bend of sensitive zone brings about mode coupling and leads to the variation of surface scattering loss. The mathematic model of relationship among light loss, bending curvature, surface roughness and parameters of the fiber's configuration is also presented
Renqiang Liu, Zhuang Fu, Yanzheng Zhao, Qixin Cao, Shuguo Wang
IROS4
2004 A Modified CMAC Algorithm Based on Credit Assignment
Lei Zhang 0131, Qixin Cao, Yanzheng Zhao
Neural Process. Lett.2