Qunfei Zhao

dblp:24/198 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-9882-730XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 Multi-guided feature refinement for point cloud semantic segmentation with weakly supervision
Qunfei Zhao, Zeyang Xia
Knowl. Based Syst.2
2024 Representing Boundary-Ambiguous Scene Online With Scale-Encoded Cascaded Grids and Radiance Field Deblurring
abstract
Implicit scene representations have recently shown promising results in photo-realistic 3D reconstruction and view synthesis based on calibrated views. However, their applications face several challenges, including unknown camera pose, boundary ambiguity, and observation noise. This paper proposes a novel online scene representation method that simultaneously learns to represent the target scene and estimates the camera poses from an RGB-D stream. An implicit scene representation function built with scale-encoded cascaded grids is proposed to represent scenes online from incremental observations. This implicit function is optimized in a reparameterized domain that provides defined boundaries. In this reparameterized domain, the cascaded grids are progressively distilled under geometric and photometric supervision to improve their model capacity and geometry accuracy. A radiance field deblurring module based on the physical imaging process is further proposed to restore a photo-realistic reconstruction against camera motion blur, which is the main component of the observation noise. The proposed method can produce sharp and photo-realistic representations of scenes under various shooting conditions without known camera poses. Experiments on multiple datasets have demonstrated the effectiveness of the proposed method in improving view synthesis and camera tracking results for online scene representation tasks.
Zeyang Xia, Qunfei Zhao
IEEE Trans. Circuits Syst. Video Technol.3
2024 Hybrid Shape Deformation for Face Reconstruction in Aesthetic Orthodontics
abstract
Achieving accurate face reconstruction with geometry details from a single-view images is an important task for orthodontics. Although the 3D Morphable Model (3DMM) based methods provide an effective framework, the low-dimensional linear space is insufficient to cover geometric details. In this paper, we propose a hybrid shape deformation representation with multi-branch supervision for detail prediction. In orthodontic scenarios, shape deformation can be considered as the aggregation of intuitive appearance component and ambiguous geometry component, which is involved in frontal shape and depth correction respectively. Hence orthogonal decomposition is employed to decompose the shape deformation into frontal-plane position offset and depth offset. Frontal-plane position offset is represented in an explicit-local-dependent manner based on grid deformation while depth offset is represented in an implicit-local-dependent manner based on dense prediction. To facilitate orthodontic-based evaluation, we construct an orthodontic-specific dataset and design a novel metric to involve the relative position dependency between regions of interest. Experimentally, we demonstrate outstanding performance of face reconstruction on FaceScape, MICC Florence and orthodontic-specific dataset with both quantitative and qualitive evaluation.
Qunfei Zhao, Zeyang Xia
IEEE Trans. Circuits Syst. Video Technol.3
2023 Sparse-to-Local-Dense Matching for Geometry-Guided Correspondence Estimation
abstract
Establishing reliable correspondences between two views is one of the most important components of various vision tasks. This paper proposes a novel sparse-to-local-dense (S2LD) matching method to conduct fully differentiable correspondence estimation with the prior from epipolar geometry. The sparse-to-local-dense matching asymmetrically establishes correspondences with consistent sub-pixel coordinates while reducing the computation of matching. The salient features are explicitly located, and the description is conditioned on both views with the global receptive field provided by the attention mechanism. The correspondences are progressively established in multiple levels to reduce the underlying re-projection error. We further propose a 3D noise-aware regularizer with differentiable triangulation. Additional guidance from 3D space is encoded by the regularizer in training to handle the supervision noise caused by the errors in camera poses and depth maps. The proposed method demonstrates outstanding matching accuracy and geometric estimation capability on multiple datasets and tasks.
Qunfei Zhao, Zeyang Xia
IEEE Trans. Image Process.2
2023 Joint-Confidence-Guided Multi-Task Learning for 3D Reconstruction and Understanding From Monocular Camera
abstract
3D reconstruction and understanding from monocular camera is a key issue in computer vision. Recent learning-based approaches, especially multi-task learning, significantly achieve the performance of the related tasks. However a few works still have limitation in drawing loss-spatial-aware information. In this paper, we propose a novel Joint-confidence-guided network (JCNet) to simultaneously predict depth, semantic labels, surface normal, and joint confidence map for corresponding loss functions. In details, we design a Joint Confidence Fusion and Refinement (JCFR) module to achieve multi-task feature fusion in the unified independent space, which can also absorb the geometric-semantic structure feature in the joint confidence map. We use confidence-guided uncertainty generated by the joint confidence map to supervise the multi-task prediction across the spatial and channel dimensions. To alleviate the training attention imbalance among different loss functions or spatial regions, the Stochastic Trust Mechanism (STM) is designed to stochastically modify the elements of joint confidence map in the training phase. Finally, we design a calibrating operation to alternately optimize the joint confidence branch and the other parts of JCNet to avoid overfiting. The proposed methods achieve state-of-the-art performance in both geometric-semantic prediction and uncertainty estimation on NYU-Depth V2 and Cityscapes.
Qunfei Zhao, Yangzhou Gan, Zeyang Xia
IEEE Trans. Image Process.2
2023 Hierarchical Belief Propagation on Image Segmentation Pyramid
abstract
The Markov random field (MRF) for stereo matching can be solved using belief propagation (BP). However, the solution space grows significantly with the introduction of high-resolution stereo images and 3D plane labels, making the traditional BP algorithms impractical in inference time and convergence. We present an accurate and efficient hierarchical BP framework using the representation of the image segmentation pyramid (ISP). The pixel-level MRF can be solved by a top-down inference on the ISP. We design a hierarchy of MRF networks using the graph of superpixels at each ISP level. From the highest/image to the lowest/pixel level, the MRF models can be efficiently inferred with constant global guidance using the optimal labels of the previous level. The large texture-less regions can be handled effectively by the MRF model on a high level. The advanced 3D continuous labels and a novel support-points regularization are integrated into our framework for stereo matching. We provide a data-level parallelism implementation which is orders of magnitude faster than the best graph cuts (GC) algorithm. The proposed framework, HBP-ISP, outperforms the best GC algorithm on the Middlebury stereo matching benchmark.
Tingman Yan, Xilian Yang, Genke Yang, Qunfei Zhao
IEEE Trans. Image Process.4
2022 Hierarchical Superpixel Segmentation by Parallel CRTrees Labeling
abstract
This paper proposes a hierarchical superpixel segmentation by representing an image as a hierarchy of 1-nearest neighbor (1-NN) graphs with pixels/superpixels denoting the graph vertices. The 1-NN graphs are built from the pixel/superpixel adjacent matrices to ensure connectivity. To determine the next-level superpixel hierarchy, inspired by FINCH clustering, the weakly connected components (WCCs) of the 1-NN graph are labeled as superpixels. We reveal that the WCCs of a 1-NN graph consist of a forest of cycle-root-trees (CRTrees). The forest-like structure inspires us to propose a two-stage parallel CRTrees labeling which first links the child vertices to the cycle-roots and then labels all the vertices by the cycle-roots. We also propose an inter-inner superpixel distance penalization and a Lab color lightness penalization base on the property that the distance of a CRTree decreases monotonically from the child to root vertices. Experiments show the parallel CRTrees labeling is several times faster than recent advanced sequential and parallel connected components labeling algorithms. The proposed hierarchical superpixel segmentation has comparable performance to the best performer ETPS (state-of-the-arts) on the BSDS500, NYUV2, and Fash datasets. At the same time, it can achieve 200FPS for 480P video streams.
Tingman Yan, Xiaolin Huang, Qunfei Zhao
IEEE Trans. Image Process.3
2019 Dynamic gesture recognition by directional pulse coupled neural networks for human-robot interaction in real time
Zeyang Xia, Weiwu Yan, Qunfei Zhao
J. Vis. Commun. Image Represent.4
2019 Segment-Based Disparity Refinement With Occlusion Handling for Stereo Matching
abstract
In this paper, we propose a disparity refinement method that directly refines the winner-take-all (WTA) disparity map by exploring its statistical significance. According to the primary steps of the segment-based stereo matching, the reference image is over-segmented into superpixels and a disparity plane is fitted for each superpixel by an improved random sample consensus (RANSAC). We design a two-layer optimization to refine the disparity plane. In the global optimization, mean disparities of superpixels are estimated by Markov random field (MRF) inference, and then, a 3D neighborhood system is derived from the mean disparities for occlusion handling. In the local optimization, a probability model exploiting Bayesian inference and Bayesian prediction is adopted and achieves second-order smoothness implicitly among 3D neighbors. The two-layer optimization is a pure disparity refinement method because no correlation information between stereo image pairs is demanded during the refinement. Experimental results on the Middlebury and KITTI datasets demonstrate that the proposed method can perform accurate stereo matching with a faster speed and handle the occlusion effectively. It can be indicated that the "matching cost computation + disparity refinement" framework is a possible solution to produce accurate disparity map at low computational cost.
Tingman Yan, Yangzhou Gan, Zeyang Xia, Qunfei Zhao
IEEE Trans. Image Process.4
2018 Tooth and Alveolar Bone Segmentation From Dental Computed Tomography Images
abstract
Three-dimensional (3D) models of tooth-alveolar bone complex are needed in treatment planning and simulation for computer-aided orthodontics. Tooth and alveolar bone segmentation from computed tomography (CT) images is a fundamental step in reconstructing their models. Due to less application of alveolar bone in conventional orthodontic treatment which may cause undesired side effects, the previous studies mainly focused on tooth segmentation and reconstruction, and did not consider the alveolar bone. In this study, we proposed a method to implement both tooth and alveolar bone segmentation from dental CT images for reconstructing their 3D models. First, the proposed method extracted the connected region of tooth and alveolar bone from CT images using a global convex level set model. Then, individual tooth and alveolar bone are separated from the connected region based on Radon transform and a local level set model. The experimental results showed that the proposed method could successfully complete both the tooth and alveolar bone segmentation from CT images, and outperformed the state of the art tooth segmentation methods in terms of accuracy. This suggests that the proposed method can be used in reconstructing the 3D models of tooth-alveolar bone complex for precise treatment.
Yangzhou Gan, Zeyang Xia, Jing Xiong 0001, Guanglin Li 0001, Qunfei Zhao
IEEE J. Biomed. Health Informatics5
2017 A Body Emotion-Based Human-Robot Interaction
Tehao Zhu, Qunfei Zhao, Jing Xiong 0001
ICVS2
2016 Crown Segmentation From Computed Tomography Images With Metal Artifacts
abstract
Tooth segmentation from dental computed tomography (CT) images with metal artifacts is challenging as metal artifacts make some of the crown boundaries unrecognizable. This letter proposes a semiautomatic method for crown segmentation from CT images with metal artifacts. A user manually selects a starting slice and initializes this slice. Then crown contours are segmented automatically from volumetric CT images slice by slice. In the segmentation of each slice, the Radon transform is used to extract a line to separate neighboring crowns into independent ones. A statistical shape prior-based level set model is then applied to segment each crown from the mesial or distal side of the line. The proposed method was tested on 15 set of volumetric images. Experimental results validated that it is effective to extract crown contours from CT images with metal artifacts.
Zeyang Xia, Yangzhou Gan, Jing Xiong 0001, Qunfei Zhao
IEEE Signal Process. Lett.4
2015 A no-reference image quality assessment approach based on steerable pyramid decomposition using natural scene statistics
Fangfang Lu, Qunfei Zhao, Genke Yang
Neural Comput. Appl.2
2007 Robot navigation and sound based position identification
abstract
In this paper, we proposed a robot self position identification method by active sound localization. This method can be used for autonomous security robots working in room environments. A system using a AIBO robot equipped with two microphones and wireless network is constructed and is used for position identification experiments. Arrival time differences to the microphones of robot are used as localization cues. To overcome the ambiguity of front-back confusion, a three-head position measurement method was proposed. The robot position can be identified by the intersection of circles restricted by the azimuth differences to different speaker pairs. By localizing three or four speakers as sound beacons positioned on known locations, the robot can identify its self position with an average error of about 7 cm in a 2.5times3.0 m2working space. A robot navigation experiment was conducted to demonstrate the effectiveness of the position identification system.
Huakang Li, Satoshi Ishikawa, Qunfei Zhao, Michiko Ebana, Hiroyuki Yamamoto, Jie Huang 0012
SMC3
2007 Application of "human-in-the-loop" control to a biped walking-chair robot
abstract
This paper presents the research on stability for biped walking-chair robot with human-in-the-loop. The inherent properties of the biped system which is developed for the disable people to replace traditional wheelchairs are analyzed. Control of the robot for the gait and navigation is introduced. Posture stability computation method based on ZMP (zero moment point) theory is discussed. Some suggestions about the future research are also presented.
Jiaoyan Tang, Qunfei Zhao, Jie Huang 0012
SMC2