Fangbo Qin

dblp:178/4858 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0002-4085-0857ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 7 since 2021Systems, architecture and hardware · 6 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Visual Anomaly Detection for Reliable Robotic Implantation of Flexible Microelectrode Array
abstract
Flexible microelectrode (FME) implantation into brain cortex is challenging due to the deformable fiber-like structure of FME probe and the interaction with critical bio-tissue. To ensure the reliability and safety, the implantation process should be monitored carefully. This paper develops an image-based anomaly detection framework based on the microscopic cameras of the robotic FME implantation system. The unified framework is utilized at four checkpoints to check the micro-needle, FME probe, hooking result, and implantation point, respectively. Exploiting the existing object localization results, the aligned regions of interest (ROIs) are extracted from raw image and input to a pretrained vision transformer (ViT). Considering the task specifications, we propose a progressive granularity patch feature sampling method to address the sensitivity-tolerance trade-off issue at different locations. Moreover, we select a part of feature channels with higher signal-to-noise ratios from the raw general ViT features, to provide better descriptors for each specific scene. The effectiveness of the proposed methods is validated with the image datasets collected from our implantation system.
Xinyao Xu 0001, Xinyong Han, Fangbo Qin
IROS5
2024 AnyOKP: One-Shot and Instance-Aware Object Keypoint Extraction with Pretrained ViT
abstract
Towards flexible object-centric visual perception, we propose a one-shot instance-aware object keypoint (OKP) extraction approach, AnyOKP, which leverages the powerful representation ability of pretrained vision transformer (ViT), and can obtain keypoints on multiple object instances of arbitrary category after learning from a support image. An off-the-shelf petrained ViT is directly deployed for generalizable and transferable feature extraction, which is followed by training-free feature enhancement. The best-prototype pairs (BPPs) are searched for in support and query images based on appearance similarity, to yield instance-unaware candidate keypoints. Then, the entire graph with all candidate keypoints as vertices are divided into sub-graphs according to the feature distributions on the graph edges. Finally, each sub-graph represents an object instance. AnyOKP is evaluated on real object images collected with the cameras of a robot arm, a mobile robot, and a surgical robot, which not only demonstrates the cross-category flexibility and instance awareness, but also show remarkable robustness to domain shift and viewpoint change.
Fangbo Qin, Taogang Hou, Michael C. Yip
ICRA1
2023 Object-Agnostic Vision Measurement Framework Based on One-Shot Learning and Behavior Tree
abstract
Vision measurement is important for intelligent systems to obtain the precise structural and spatial information of objects. Beyond the object-specific vision measurement developed for fixed object type, it is appealing to explore the object-agnostic vision measurement, which can be efficiently reconfigured and adapted to various novel objects. This article proposes a framework to mimic the human's versatile visual measurement behavior: extract a set of contour primitives of interest (CPIs) from an image, then utilize the CPIs to calculate the key geometric information. First, a deep convolutional neural network (CNN) CPieNet+ is proposed under the one-shot learning scheme, aiming to extract the pixel-level object CPI from a raw query image, given an annotated support image. The fine-grained CPI prototypes are formed by sampling multiple points on the feature map of the support image. To leverage the explicit geometric knowledge in the CNN inference, the annotation map is encoded as a shape descriptor to guide the feature channel attention, and the geometric attribute awareness is realized by supervising the model to predict the direction and size of CPI. Second, the measurement behavior tree (BT) is designed to model the hierarchical geometric calculation procedure, which is flexibly configurable for different measurement requirements and is interpretable for nonexpert users. After the execution of the measurement BT, the pixel-level CPIs are converted to the required key geometric data. The effectiveness of the proposed methods is validated by a series of experiments.
Fangbo Qin, De Xu, Blake Hannaford, Tiantian Hao
IEEE Trans. Cybern.1
2023 Contour Primitive of Interest Extraction Network Based on Dual-Metric One-Shot Learning for Vision Measurement
abstract
Although many existing vision measurement systems have achieved high performances, they are object-specific and have limitations in flexibility. Toward intelligent vision measurement that can be conveniently reused for novel objects, this article focuses on the image geometric feature extraction with one-shot learning ability. We propose a contour primitive of interest (CPI) extraction network with dual metric (CPieNet-DM), which can obtain a designated CPI in a query image of a novel object under the guidance of only one annotated support image. First, the dual-metric learning mechanism is proposed, which not only utilizes inter-image similarity as guidance but also leverages the intra-image coherency of CPI pixels to facilitate the inference. Second, a neural network is designed to infer the CPI map based on the dual metric, which also predicts the CPI's geometric parameters. Moreover, the dual context aggregator is plugged in to provide the awareness of both images’ contexts. Third, the network training is jointly supervised by the multiple tasks of dual-metric learning, geometric parameters regression, and CPI extraction. The online hard example mining is utilized to improve the training outcome. The effectiveness of the proposed methods is validated with a series of experiments.
Fangbo Qin, De Xu
IEEE Trans. Ind. Informatics1
2022 MVSTER: Epipolar Transformer for Efficient Multi-view Stereo
Guan Huang 0003, Fangbo Qin, Yijia He, Xu Chi, Xingang Wang 0003
ECCV (31)4
2021 ELSD: Efficient Line Segment Detector and Descriptor
abstract
We present the novel Efficient Line Segment Detector and Descriptor (ELSD) to simultaneously detect line segments and extract their descriptors in an image. Unlike the traditional pipelines that conduct detection and description separately, ELSD utilizes a shared feature extractor for both detection and description, to provide the essential line features to the higher-level tasks like SLAM and image matching in real time. First, we design a one-stage compact model, and propose to use the mid-point, angle and length as the minimal representation of line segment, which also guarantees the center-symmetry. The non-centerness suppression is proposed to filter out the fragmented line segments caused by lines’ intersections. The fine offset prediction is designed to refine the mid-point localization. Second, the line descriptor branch is integrated with the detector branch, and the two branches are jointly trained in an end-to-end manner. In the experiments, the proposed ELSD achieves the state-of-the-art performance on the Wireframe dataset and YorkUrban dataset, in both accuracy and efficiency. The line description ability of ELSD also outperforms the previous works on the line matching task.
Yicheng Luo, Fangbo Qin, Yijia He, Xiao Liu 0042
ICCV3
2021 Learning Surgical Motion Pattern from Small Data in Endoscopic Sinus and Skull Base Surgeries
abstract
Existing studies demonstrated that surgical motion patterns are strongly correlated with surgical outcomes. Real surgeries are complicated and it is expensive to harvest surgical data. Consequently, existing researches on surgical motion patterns focus on specific concise surgical tasks or simple surgical procedures. The paper presents a surgical motion pattern modeling technique that uses small data but can be applied to virtually any Endoscopic Sinus and Skull Base Surgeries (ESSBSs). The proposed method decreases the dimensionalities of the feature space through projecting surgical instrument motions into the endoscope coordinate, based on human expert domain knowledge. Furthermore, the method uses kinematic features and learns the motion pattern with Gaussian Process learning techniques. Comparing with existing surgical motion pattern modeling methods, the proposed method: 1, learns the motion model from small data; 2, can be generally applied to ESSBSs because it neither assumes nor depends on specific surgical tasks; 3, provides informative results in a real-time manner for optimizing surgical motions for improving surgical outcomes. The proposed method was verified by predicting surgical skill levels on cadaver surgeries. The results show the real-time prediction precision is higher than 81% and the offline accumulated precision reach 100%.
Yangming Li, Randall A. Bly, Sarah Akkina, Fangbo Qin, Rajeev C. Saxena, Ian Humphreys, Mark Whipple, Kris S. Moe, Blake Hannaford
ICRA4
2021 Contour Primitive of Interest Extraction Network Based on One-Shot Learning for Object-Agnostic Vision Measurement
abstract
Image contour based vision measurement is widely applied in robot manipulation and industrial automation. It is appealing to realize object-agnostic vision system, which can be conveniently reused for various types of objects. We propose the contour primitive of interest extraction network (CPieNet) based on the one-shot learning framework. First, CPieNet is featured by that its contour primitive of interest (CPI) output, a designated regular contour part lying on a specified object, provides the essential geometric information for vision measurement. Second, CPieNet has the one-shot learning ability, utilizing a support sample to assist the perception of the novel object. To realize lower-cost training, we generate support-query sample pairs from unpaired online public images, which cover a wide range of object categories. To obtain single-pixel wide contour for precise measurement, the Gabor-filters based non-maximum suppression is designed to thin the raw contour. For the novel CPI extraction task, we built the Object Contour Primitives dataset using online public images, and the Robotic Object Contour Measurement dataset using a camera mounted on a robot. The effectiveness of the proposed methods is validated by a series of experiments.
Fangbo Qin, Siyu Huang, De Xu
ICRA1
2021 Efficient Insertion Control for Precision Assembly Based on Demonstration Learning and Reinforcement Learning
abstract
Multiple peg-in-hole insertion control is one of the challenging tasks in precision assembly for its complex contact dynamics. In this article, an insertion policy learning method is proposed for multiple peg-in-hole precision assembly. The insertion policy learning process is separated into two phases: initial policy learning and residual policy learning. In initial policy learning, a state-to-action policy mapping model based on the Gaussian mixture model (GMM) is established. And Gaussian mixture regression (GMR) is used to generalize the policy reuse. In residual policy learning, a reinforcement learning method named normalized advantage function (NAF) is employed to refine the insertion policy via agent's exploration in the insertion environment. Moreover, an adaptive action exploration (AAE) strategy is designed to improve the performance of exploration, and the prioritized experience replay strategy is introduced to make the residual policy learning from historical experience more efficient. Besides, the hierarchical reward function is designed considering the contact dynamics as well as the efficiency and safety of precision insertion. Finally, comprehensive experiments are conducted to validate the effectiveness of the proposed insertion policy learning method.
Yanqin Ma, De Xu, Fangbo Qin
IEEE Trans. Ind. Informatics3
2020 TP-LSD: Tri-Points Based Line Segment Detector
Siyu Huang, Fangbo Qin, Pengfei Xiong, Yijia He, Xiao Liu 0042
ECCV (27)2
2020 LC-GAN: Image-to-image Translation Based on Generative Adversarial Network for Endoscopic Images
abstract
Intelligent vision is appealing in computer-assisted and robotic surgeries. Vision-based analysis with deep learning usually requires large labeled datasets, but manual data labeling is expensive and time-consuming in medical problems. We investigate a novel cross-domain strategy to reduce the need for manual data labeling by proposing an image-to-image translation model live-cadaver GAN (LC-GAN) based on generative adversarial networks (GANs). We consider a situation when a labeled cadaveric surgery dataset is available while the task is instrument segmentation on an unlabeled live surgery dataset. We train LC-GAN to learn the mappings between the cadaveric and live images. For live image segmentation, we first translate the live images to fake-cadaveric images with LC-GAN and then perform segmentation on the fake-cadaveric images with models trained on the real cadaveric dataset. The proposed method fully makes use of the labeled cadaveric dataset for live image segmentation without the need to label the live dataset. LC-GAN has two generators with different architectures that leverage the deep feature representation learned from the cadaveric image based segmentation task. Moreover, we propose the structural similarity loss and segmentation consistency loss to improve the semantic consistency during translation. Our model achieves better image-to-image translation and leads to improved segmentation performance in the proposed cross-domain segmentation task.
Fangbo Qin, Yangming Li, Randall A. Bly, Kris S. Moe, Blake Hannaford
IROS2
2020 Laser Beam Pointing Control With Piezoelectric Actuator Model Learning
abstract
The inherent hysteresis property of piezoelectric actuator (PEA) brings challenges to its modeling and control. This paper proposes a model learning method that is suitable for both forward and inverse PEA models. The hysteresis property is learned based on least squares support vector machines (LS-SVMs). A larger dataset is used for training LS-SVM to guarantee a good generalization performance. Support vectors pruning is utilized to reduce the model complexity. The rate-dependent property of PEA is identified as a linear dynamic submodel. Moreover, a pointing control system with two dualPEA-axis steering mirrors is developed, which can regulate the 4-degree-of-freedom pose of a laser beam. The coordinated control of four PEAs is realized based on the Jacobian matrix. The learned inverse PEA models are used for the feedforward compensation of each PEA's nonlinearity. A series of experiments were conducted to evaluate the proposed method's effectiveness.
Fangbo Qin, Dengpeng Xing, De Xu
IEEE Trans. Syst. Man Cybern. Syst.1
2019 Surgical Instrument Segmentation for Endoscopic Vision with Data Fusion of rediction and Kinematic Pose
abstract
The real-time and robust surgical instrument segmentation is an important issue for endoscopic vision. We propose an instrument segmentation method fusing the convolutional neural networks (CNN) prediction and the kinematic pose information. First, the CNN model ToolNet-C is designed, which cascades a convolutional feature extractor trained over numerous unlabeled images and a pixel-wise segmentor trained on few labeled images. Second, the silhouette projection of the instrument body onto the endoscopic image is implemented based on the measured kinematic pose. Third, the particle filter with the shape matching likelihood and the weight suppression is proposed for data fusion, whose estimate refines the kinematic pose. The refined pose determines an accurate silhouette mask, which is the final segmentation output. The experiments are conducted with a surgical navigation system, several animal-tissue backgrounds, and a debrider instrument.
Fangbo Qin, Yangming Li, Yun-Hsuan Su, De Xu, Blake Hannaford
ICRA1
2018 Contour Primitives of Interest Extraction Method for Microscopic Images and Its Application on Pose Measurement
abstract
This paper proposes a suite of methods to realize high precision pose measurement in 3-D Cartesian space based on a multicamera microscopic vision system. Since it is inefficient to develop a specific image algorithm for each kind of object and the imaging condition might be unsatisfactory, we propose a method of contour primitives of interest extraction, which allows flexible reconfiguration for novel object image and owns robustness under different imaging conditions. The object is detected in grayscale image based on a template of contour primitives. Edges are extracted according to derivatives along the normal vectors of these contour primitives. The positions and directional derivatives of these edges are used for feature extraction and autofocus, respectively. The point features and line features extracted from multiview images are utilized to measure 3-D vectors and orientations, respectively, based on image Jacobian matrices. Cameras' linear motions are considered in the imaging model, so that the measurement range is expanded beyond the limitation of microscopes' shallow depths of field. The affine epipolar constraint and focused planes intersection constraint between cameras are applied to improve the real time performances of image feature extraction and multicamera autofocus, respectively. A series of experiments are conducted to verify the effectiveness of the proposed methods. The root mean square errors of pose measurement are evaluated as 3 μm in position and 0.05° in orientation, while the measurement range is about 5000 μm in position and 20° in orientation.
Fangbo Qin, Fei Shen 0002, Xilong Liu, De Xu
IEEE Trans. Syst. Man Cybern. Syst.1