Xibin Song

dblp:146/4830 · DBLP profile ↗
← Back
33ranked-venue papers
11as first author
22since 2021 · last 2026
0000-0001-7019-6238ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 18 · 4 first-author · 12 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 IntentionAR: An Intention-Driven Camera-Projector System for AR Assembly Guidance
abstract
Augmented reality (AR) assembly guidance systems can help users quickly master the assembly process of unfamiliar objects. However, it is very difficult for practical use without a thorough understanding of users' intentions. Furthermore, existing approaches still struggle to handle complex hand-part interactions and nonlinear assembly steps, and achieving intuitive and real-time guidance remains challenging. To address these issues, we propose an intention-driven camera-projector AR assembly guidance system (IntentionAR) that integrates online user intention inference with a finite-state assembly machine. The intention module recognizes interaction actions and infers the target object, enabling feedback before assembly begins, while the finite state machine (FSM) manages step progression. Using a camera-projector device, the system highlights candidate parts and, based on inferred intention, flags correct/incorrect selections to provide real-time guidance. As the assembly progresses, the models used for spatial registration and visualization switch dynamically to accommodate the nonlinear workflow. Experiments and user studies show that the system delivers a more robust AR interaction experience and improves assembly efficiency. A free copy of this paper and all supplemental materials, as well as project assets and source code, will be available at the project website.
Xin Cao 0010, Mingyu Ma 0013, Shanhao Yang, Kang Xie, Xibin Song, Fan Zhong 0001, Hai-Ning Liang, Xueying Qin
IEEE Trans. Vis. Comput. Graph.6
2026 BAG: Body-Aligned 3D Wearable Asset Generation
abstract
While recent advancements have demonstrated remarkable progress in general 3D shape generation, the challenge of automatically generating wearable 3D assets remains largely unexplored. To address this gap, we present BAG - a Body-aligned Asset Generation method that produces 3D wearable assets which can be automatically fitted onto given 3D human bodies. This is achieved by controlling the 3D generation process using human body shape and pose information. Specifically, we first construct a general single-image-to-consistent-multi-view diffusion model, and train it on the large-scale Objaverse dataset to ensure diversity and generalizability. We then train a body-conditioned multi-view ControlNet to guide the generator toward producing body-aligned multi-view images. The control signal leverages multi-view 2D projections of the target human body, where pixel values represent the XYZ coordinates of the body surface in a canonical space. The resulting body-conditioned multi-view diffusion outputs body-aligned images, which are subsequently fed into a native 3D diffusion model to reconstruct the 3D shape of the asset. Finally, we recover the similarity transformation using multi-view silhouette supervision and mitigate asset-body penetration using physics-based simulation, ensuring accurate asset fitting onto the target body. Experimental results demonstrate that our method significantly outperforms existing approaches in terms of prompt adherence, shape diversity, and shape quality.
Zhongjin Luo, Yang Li 0193, Senbo Wang, Han Yan 0004, Xibin Song, Taizhang Shang, Wei Mao 0001, Hongdong Li, Xiaoguang Han 0001, Pan Ji
IEEE Trans. Vis. Comput. Graph.6
2024 NeuSDFusion: A Spatial-Aware Generative Model for 3D Shape Completion, Reconstruction, and Generation
Ruikai Cui, Weizhe Liu, Weixuan Sun, Senbo Wang, Taizhang Shang, Yang Li 0193, Xibin Song, Han Yan 0004, Zhennan Wu, Shenzhou Chen, Hongdong Li, Pan Ji
ECCV (19)7
2024 CSS-Net: Domain Generalization in Category-level Pose Estimation via Corresponding Structural Superpoints
abstract
Category-level pose estimation is crucial for estimating the pose and size of unseen objects. Previous methods, mainly trained and tested on data with the same distribution, are limited in their ability to generalize to unseen domain data. For instance, when applied to new scenes or categories, frequent data collection and network training can be cumbersome. To address this issue, we propose a domain generalization method in category-level pose estimation based on structural superpoints, which is trained solely on simulated data and can generalize to unseen domain distributions in real datasets. Specifically, by extracting superpoints for structural correspondence in a self-supervised manner, our method achieves cross-domain data and cross-instance shape generalization. Accordingly, we designed a network and loss function, CoupleLoss, for regressing pose and size. Furthermore, we validated the effectiveness of our method on the wild6D and real275 datasets, achieving state-of-the-art results.
Xibin Song, Changhe Tu, Xueying Qin
ICME2
2024 RGB-based Category-level Object Pose Estimation via Decoupled Metric Scale Recovery
abstract
While showing promising results, recent RGB-D camera-based category-level object pose estimation methods have restricted applications due to the heavy reliance on depth sensors. RGB-only methods provide an alternative to this problem yet suffer from inherent scale ambiguity stemming from monocular observations. In this paper, we propose a novel pipeline that decouples the 6D pose and size estimation to mitigate the influence of imperfect scales on rigid transformations. Specifically, we leverage a pre-trained monocular estimator to extract local geometric information, mainly facilitating the search for inlier 2D-3D correspondence. Meanwhile, a separate branch is designed to directly recover the metric scale of the object based on category-level statistics. Finally, we advocate using the RANSAC-PnP algorithm to robustly solve for 6D object pose. Extensive experiments have been conducted on both synthetic and real datasets, demonstrating the superior performance of our method over previous state-of-the-art RGB-based approaches, especially in terms of rotation accuracy. Code: https://github.com/goldoak/DMSR.
Jiaxin Wei 0001, Xibin Song, Weizhe Liu, Laurent Kneip, Hongdong Li, Pan Ji
ICRA2
2024 Implicit Coarse-to-Fine 3D Perception for Category-level Object Pose Estimation from Monocular RGB Image
abstract
Category-level object pose estimation demonstrates robust generalization capabilities that benefit robotics applications. However, exclusive reliance on RGB images without leveraging any 3D information introduces ambiguity in the translation and size of objects, leading to suboptimal performance. In this paper, we propose a framework for category-level pose estimation from a single RGB image in an end-to-end manner, i.e., Feature Auxiliary Perception Network (FAP-Net). To address inaccurate pose estimation caused by the inherent ambiguity of RGB images, we design a coarse-to-fine approach that first harnesses geometry supervision to facilitate coarse 3D feature perception and subsequently refines the features based on pose and size constraints. Experimental results on REAL275 and CAMERA25 demonstrate that FAP-Net achieves significant improvements (14.7% on 10°10cm and 11.4% on IoU50 on the real-scene REAL275 dataset) over the state-of-the-art and real-time inference (42 FPS).
Xibin Song, Yeheng Chen, Xueying Qin
ICRA3
2024 LAM3D: Large Image-Point Clouds Alignment Model for 3D Reconstruction from Single Image
abstract
Large Reconstruction Models have made significant strides in the realm of automated 3D content generation from single or multiple input images. Despite their success, these models often produce 3D meshes with geometric inaccuracies, stemming from the inherent challenges of deducing 3D shapes solely from image data. In this work, we introduce a novel framework, the Large Image and Point Cloud Alignment Model (LAM3D), which utilizes 3D point cloud data to enhance the fidelity of generated 3D meshes. Our methodology begins with the development of a point-cloud-based network that effectively generates precise and meaningful latent tri-planes, laying the groundwork for accurate 3D mesh reconstruction. Building upon this, our Image-Point-Cloud Feature Alignment technique processes a single input image, aligning to the latent tri-planes to imbue image features with robust 3D information. This process not only enriches the image features but also facilitates the production of high-fidelity 3D meshes without the need for multi-view input, significantly reducing geometric distortions. Our approach achieves state-of-the-art high-fidelity 3D mesh reconstruction from a single image in just 6 seconds, and experiments on various datasets demonstrate its effectiveness.
Ruikai Cui, Xibin Song, Weixuan Sun, Senbo Wang, Weizhe Liu, Shenzhou Chen, Taizhang Shang, Yang Li 0193, Nick Barnes, Hongdong Li, Pan Ji
NeurIPS2
2024 AGDF-Net: Learning Domain Generalizable Depth Features With Adaptive Guidance Fusion
abstract
Cross-domain generalizable depth estimation aims to estimate the depth of target domains (i.e., real-world) using models trained on the source domains (i.e., synthetic). Previous methods mainly use additional real-world domain datasets to extract depth specific information for cross-domain generalizable depth estimation. Unfortunately, due to the large domain gap, adequate depth specific information is hard to obtain and interference is difficult to remove, which limits the performance. To relieve these problems, we propose a domain generalizable feature extraction network with adaptive guidance fusion (AGDF-Net) to fully acquire essential features for depth estimation at multi-scale feature levels. Specifically, our AGDF-Net first separates the image into initial depth and weak-related depth components with reconstruction and contrary losses. Subsequently, an adaptive guidance fusion module is designed to sufficiently intensify the initial depth features for domain generalizable intensified depth features acquisition. Finally, taking intensified depth features as input, an arbitrary depth estimation network can be used for real-world depth estimation. Using only synthetic datasets, our AGDF-Net can be applied to various real-world datasets (i.e., KITTI, NYUDv2, NuScenes, DrivingStereo and CityScapes) with state-of-the-art performances. Furthermore, experiments with a small amount of real-world data in a semi-supervised setting also demonstrate the superiority of AGDF-Net over state-of-the-art approaches.
Lina Liu 0010, Xibin Song, Mengmeng Wang 0005, Yuchao Dai, Yong Liu 0007, Liangjun Zhang
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 SRNSD: Structure-Regularized Night-Time Self-Supervised Monocular Depth Estimation for Outdoor Scenes
abstract
Deep CNNs have achieved impressive improvements for night-time self-supervised depth estimation form a monocular image. However, the performance degrades considerably compared to day-time depth estimation due to significant domain gaps, low visibility, and varying illuminations between day and night images. To address these challenges, we propose a novel night-time self-supervised monocular depth estimation framework with structure regularization, i.e., SRNSD, which incorporates three aspects of constraints for better performance, including feature and depth domain adaptation, image perspective constraint, and cropped multi-scale consistency loss. Specifically, we utilize adaptations of both feature and depth output spaces for better night-time feature extraction and depth map prediction, along with high- and low-frequency decoupling operations for better depth structure and texture recovery. Meanwhile, we employ an image perspective constraint to enhance the smoothness and obtain better depth maps in areas where the luminosity jumps change. Furthermore, we introduce a simple yet effective cropped multi-scale consistency loss that utilizes consistency among different scales of depth outputs for further optimization, refining the detailed textures and structures of predicted depth. Experimental results on different benchmarks with depth ranges of 40m and 60m, including Oxford RobotCar dataset, nuScenes dataset and CARLA-EPE dataset, demonstrate the superiority of our approach over state-of-the-art night-time self-supervised depth estimation approaches across multiple metrics, proving our effectiveness.
Runmin Cong, Chunlei Wu, Xibin Song, Wei Zhang 0021, Sam Kwong, Hongdong Li, Pan Ji
IEEE Trans. Image Process.3
2023 Digging Into Uncertainty-Based Pseudo-Label for Robust Stereo Matching
abstract
Due to the domain differences and unbalanced disparity distribution across multiple datasets, current stereo matching approaches are commonly limited to a specific dataset and generalize poorly to others. Such domain shift issue is usually addressed by substantial adaptation on costly target-domain ground-truth data, which cannot be easily obtained in practical settings. In this paper, we propose to dig into uncertainty estimation for robust stereo matching. Specifically, to balance the disparity distribution, we employ a pixel-level uncertainty estimation to adaptively adjust the next stage disparity searching space, in this way driving the network progressively prune out the space of unlikely correspondences. Then, to solve the limited ground truth data, an uncertainty-based pseudo-label is proposed to adapt the pre-trained model to the new domain, where pixel-level and area-level uncertainty estimation are proposed to filter out the high-uncertainty pixels of predicted disparity maps and generate sparse while reliable pseudo-labels to align the domain gap. Experimentally, our method shows strong cross-domain, adapt, and joint generalization and obtains 1st place on the stereo task of Robust Vision Challenge 2020. Additionally, our uncertainty-based pseudo-labels can be extended to train monocular depth estimation networks in an unsupervised way and even achieves comparable performance with the supervised methods.
Zhelun Shen, Xibin Song, Yuchao Dai, Dingfu Zhou, Zhibo Rao, Liangjun Zhang
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 WSAMF-Net: Wavelet Spatial Attention-Based MultiStream Feedback Network for Single Image Dehazing
abstract
Single image-based dehazing has achieved remarkable progress with the development of deep learning technologies. End-to-end neural networks have been proposed to learn a direct hazy-to-clear image translation to recover the clear structures and edges cues from the hazy inputs. However, the frequency domain information is explored insufficiently and lots of intermediate structure and texture related cues of current dehazing networks are ignored, which limits the performances of current approaches. To handle these limitations mentioned above, a wavelet spatial attention based multi-stream feedback network (WSAMF-Net) is proposed for effective single image dehazing. Specifically, the proposed wavelet spatial attention utilizes both frequency-domain and spatial-domain information to enhance the extracted features for better structures and edges. Meanwhile, an enhanced multi-stream based cross feature fusion strategy, including vertical and horizontal attentions, is proposed to reweight and fuse the intermediate features of each stream to acquire more meaningful aggregated features, while the weight sharing strategy is used to achieve a good trade-off between performance and parameters. Besides, feedback mechanism is also designed to provide strong reconstruction ability. Furthermore, we propose a critical real-world industrial dataset (IDS) with images captured in real-world industrial quarry scenarios for research uses. Extensive experiments on various benchmarking datasets, including both synthetic and real-world datasets, demonstrate the superiority of our WSAMF-Net over state-of-the-art single image dehazing methods. The IDS dataset will be available athttps://github.com/XBSong/IDS-Datasethttps://github.com/XBSong/IDS-Dataset.
Xibin Song, Dingfu Zhou, Wei Li 0143, Haodong Ding, Yuchao Dai, Liangjun Zhang
IEEE Trans. Circuits Syst. Video Technol.1
2023 TUSR-Net: Triple Unfolding Single Image Dehazing With Self-Regularization and Dual Feature to Pixel Attention
abstract
Single image dehazing is a challenging and ill-posed problem due to severe information degeneration of images captured in hazy conditions. Remarkable progresses have been achieved by deep-learning based image dehazing methods, where residual learning is commonly used to separate the hazy image into clear and haze components. However, the nature of low similarity between haze and clear components is commonly neglected, while the lack of constraint of contrastive peculiarity between the two components always restricts the performance of these approaches. To deal with these problems, we propose an end-to-end self-regularized network (TUSR-Net) which exploits the contrastive peculiarity of different components of the hazy image, i.e, self-regularization (SR). In specific, the hazy image is separated into clear and hazy components and constraint between different image components, i.e., self-regularization, is leveraged to pull the recovered clear image closer to groundtruth, which largely promotes the performance of image dehazing. Meanwhile, an effective triple unfolding framework combined with dual feature to pixel attention is proposed to intensify and fuse the intermediate information in feature, channel and pixel levels, respectively, thus features with better representational ability can be obtained. Our TUSR-Net achieves better trade-off between performance and parameter size with weight-sharing strategy and is much more flexible. Experiments on various benchmarking datasets demonstrate the superiority of our TUSR-Net over state-of-the-art single image dehazing methods.
Xibin Song, Dingfu Zhou, Wei Li 0143, Yuchao Dai, Zhelun Shen, Liangjun Zhang, Hongdong Li
IEEE Trans. Image Process.1
2022 End-to-End Learning the Partial Permutation Matrix for Robust 3D Point Cloud Registration
abstract
Even though considerable progress has been made in deep learning-based 3D point cloud processing, how to obtain accurate correspondences for robust registration remains a major challenge because existing hard assignment methods cannot deal with outliers naturally. Alternatively, the soft matching-based methods have been proposed to learn the matching probability rather than hard assignment. However, in this paper, we prove that these methods have an inherent ambiguity causing many deceptive correspondences. To address the above challenges, we propose to learn a partial permutation matching matrix, which does not assign corresponding points to outliers, and implements hard assignment to prevent ambiguity. However, this proposal poses two new problems, i.e. existing hard assignment algorithms can only solve a full rank permutation matrix rather than a partial permutation matrix, and this desired matrix is defined in the discrete space, which is non-differentiable. In response, we design a dedicated soft-to-hard (S2H) matching procedure within the registration pipeline consisting of two steps: solving the soft matching matrix (S-step) and projecting this soft matrix to the partial permutation matrix (H-step). Specifically, we augment the profit matrix before the hard assignment to solve an augmented permutation matrix, which is cropped to achieve the final partial permutation matrix. Moreover, to guarantee end-to-end learning, we supervise the learned partial permutation matrix but propagate the gradient to the soft matrix instead. Our S2H matching procedure can be easily integrated with existing registration frameworks, which has been verified in representative frameworks including DCP, RPMNet, and DGR. Extensive experiments have validated our method, which creates a new state-of-the-art performance.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He
AAAI5
2022 PCW-Net: Pyramid Combination and Warping Cost Volume for Stereo Matching
Zhelun Shen, Yuchao Dai, Xibin Song, Zhibo Rao, Dingfu Zhou, Liangjun Zhang
ECCV (32)3
2022 A Representation Separation Perspective to Correspondence-Free Unsupervised 3-D Point Cloud Registration
abstract
3-D point cloud registration in remote sensing field has been greatly advanced by deep learning-based methods, where the rigid transformation is either directly regressed from the two point clouds (correspondences-free approaches) or computed from the learned correspondences (correspondences-based approaches). Existing correspondence-free methods generally learn the holistic representation of the entire point cloud, which is fragile for partial and noisy point clouds. In this letter, we propose a correspondence-free unsupervised point cloud registration (UPCR) method from the representation separation perspective. First, we model the input point cloud as a combination of pose-invariant representation and pose-related representation. Second, the pose-related representation is used to learn the relative pose w.r.t. a “latent canonical shape” for thesourceandtargetpoint clouds, respectively. Third, the rigid transformation is obtained from the above two learned relative poses. Our method not only filters out the disturbance in pose-invariant representation but also is robust to partial-to-partial point clouds or noise. Experiments on benchmark datasets demonstrate that our unsupervised method achieves comparable if not better performance than state-of-the-art supervised registration methods.The source code will be made public.
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He
IEEE Geosci. Remote. Sens. Lett.5
2022 Self-supervised rigid transformation equivariance for accurate 3D point cloud registration
Zhiyuan Zhang 0002, Jiadai Sun, Yuchao Dai, Dingfu Zhou, Xibin Song, Mingyi He
Pattern Recognit.5
2022 Context-Aware 3D Object Detection From a Single Image in Autonomous Driving
abstract
Camera sensors have been widely used in Driver-Assistance and Autonomous Driving Systems due to their rich texture information. Recently, with the development of deep learning techniques, many approaches have been proposed to detect objects in 3D from a single frame, however, there is still much room for improvement. In this paper, we generally review the recently proposed state-of-the-art monocular-based 3D object detection approaches first. Based on the analysis of the disadvantage of previous center-based frameworks, a novel feature aggregation strategy has been proposed to boost the 3D object detection by exploring the context information. Specifically, an Instance-Guided Spatial Attention (IGSA) module is proposed to collect the local instance information and the Channel-Wise Feature Attention (CWFA) module is employed for aggregating the global context information. In addition, an instance-guided object regression strategy is also proposed to alleviate the influence of center location prediction uncertainty in the inference process. Finally, the proposed approach has been verified on the public 3D object detection benchmark. The experimental results show that the proposed approach can significantly boost the performance of the baseline method on both 3D detection and 2D Bird’s-Eye View among all three categories. Furthermore, our method outperforms all the monocular-based methods (even these trained with depth as auxiliary inputs) and achieves state-of-the-art performance on the KITTI benchmark.
Dingfu Zhou, Xibin Song, Yuchao Dai, Hongdong Li, Liangjun Zhang
IEEE Trans. Intell. Transp. Syst.2
2022 WAFP-Net: Weighted Attention Fusion Based Progressive Residual Learning for Depth Map Super-Resolution
abstract
Despite the remarkable progresses achieved in depth map super-resolution (DSR), it remains a major challenge to tackle with real-world degradation of low-resolution (LR) depth maps. Synthetic datasets are mainly used in existing DSR approaches, which is quite different from what would get from a real depth sensor. Besides, the enhancements of features in existing DSR approaches are not sufficiently enough, which also limit the performance. To alleviate these problems, we first propose two types of degradation models to describe the generation of LR depth maps, including bi-cubic down-sampling with noise and interval down-sampling, and different DSR models are learned correspondingly. Then, we propose a weighted attention fusion strategy that is embedded into a progressive residual learning framework, which guarantees that the high-resolution (HR) depth maps can be well recovered in a coarse-to-fine manner. The weighted attention fusion strategy can enhance the features with abundant high-frequency components in both global and local manners, thus better HR depth maps can be expected. Besides, to re-use the effective information in the progressive process sufficiently, a multi-stage fusion module is combined into the proposed framework, and the Total Generalized Variation (TGV) regularization and input loss are exploited to further improve the performance of our method. Extensive experiments of different benchmarks demonstrate the superiority of our approach over the state-of-the-art (SOTA) approaches.
Xibin Song, Dingfu Zhou, Wei Li 0111, Yuchao Dai, Liu Liu 0009, Hongdong Li, Ruigang Yang, Liangjun Zhang
IEEE Trans. Multim.1
2021 FCFR-Net: Feature Fusion based Coarse-to-Fine Residual Learning for Depth Completion
abstract
Depth completion aims to recover a dense depth map from a sparse depth map with the corresponding color image as input. Recent approaches mainly formulate the depth completion as a one-stage end-to-end learning task, which outputs dense depth maps directly. However, the feature extraction and supervision in one-stage frameworks are insufficient, limiting the performance of these approaches. To address this problem, we propose a novel end-to-end residual learning framework, which formulates the depth completion as a two-stage learning task, i.e., a sparse-to-coarse stage and a coarse-to-fine stage. First, a coarse dense depth map is obtained by a simple CNN framework. Then, a refined depth map is further obtained using a residual learning strategy in the coarse-to-fine stage with coarse depth map and color image as input. Specially, in the coarse-to-fine stage, a channel shuffle extraction operation is utilized to extract more representative features from color image and coarse depth map, and an energy based fusion operation is exploited to effectively fuse these features obtained by channel shuffle operation, thus leading to more accurate and refined depth maps. We achieve SoTA performance in RMSE on KITTI benchmark. Extensive experiments on other datasets future demonstrate the superiority of our approach over current state-of-the-art depth completion approaches.
Lina Liu 0010, Xibin Song, Xiaoyang Lyu, Junwei Diao, Mengmeng Wang 0005, Yong Liu 0007, Liangjun Zhang
AAAI2
2021 Self-supervised Monocular Depth Estimation for All Day Images using Domain Separation
abstract
Remarkable results have been achieved by DCNN based self-supervised depth estimation approaches. However, most of these approaches can only handle either day-time or night-time images, while their performance degrades for all-day images due to large domain shift and the variation of illumination between day and night images. To relieve these limitations, we propose a domain-separated network for self-supervised depth estimation of all-day images. Specifically, to relieve the negative influence of disturbing terms (illumination, etc.), we partition the information of day and night image pairs into two complementary sub-spaces: private and invariant domains, where the former contains the unique information (illumination, etc.) of day and night images and the latter contains essential shared information (texture, etc.). Meanwhile, to guarantee that the day and night images contain the same information, the domain-separated network takes the day-time images and corresponding night-time images (generated by GAN) as input, and the private and invariant feature extractors are learned by orthogonality and similarity loss, where the domain gap can be alleviated, thus better depth maps can be expected. Meanwhile, the reconstruction and photometric losses are utilized to estimate complementary information and depth maps effectively. Experimental results demonstrate that our approach achieves state-of-the-art depth estimation results for all-day images on the challenging Oxford RobotCar dataset, proving the superiority of our proposed approach. Code and data split are available at https://github.com/LINA-lln/ADDS-DepthNet.
Lina Liu 0010, Xibin Song, Mengmeng Wang 0005, Yong Liu 0007, Liangjun Zhang
ICCV2
2021 MapFusion: A General Framework for 3D Object Detection with HDMaps
abstract
3D object detection is a key perception component in autonomous driving. Most recent approaches are based on LiDAR sensors only or fused with cameras. Maps (e.g., High Definition Maps), a basic infrastructure for intelligent vehicles, however, have not been well exploited for boosting object detection tasks. In this paper, we propose a simple but effective framework - MapFusion to integrate the map information into modern 3D object detector pipelines. In particular, we design a FeatureAgg module for HD Map feature extraction and fusion, and a MapSeg module as an auxiliary segmentation head for the detection backbone. Our proposed MapFusion is detector independent and can be easily integrated into different detectors. The experimental results of three different baselines on large public autonomous driving dataset demonstrate the superiority of the proposed framework. By fusing the map information, we can achieve 1.27 to 2.79 points improvements for mean Average Precision (mAP) on three strong 3D object detection baselines.
Dingfu Zhou, Xibin Song, Liangjun Zhang
IROS3
2021 MLDA-Net: Multi-Level Dual Attention-Based Network for Self-Supervised Monocular Depth Estimation
abstract
The success of supervised learning-based single image depth estimation methods critically depends on the availability of large-scale dense per-pixel depth annotations, which requires both laborious and expensive annotation process. Therefore, the self-supervised methods are much desirable, which attract significant attention recently. However, depth maps predicted by existing self-supervised methods tend to be blurry with many depth details lost. To overcome these limitations, we propose a novel framework, named MLDA-Net, to obtain per-pixel depth maps with shaper boundaries and richer depth details. Our first innovation is a multi-level feature extraction (MLFE) strategy which can learn rich hierarchical representation. Then, a dual-attention strategy, combining global attention and structure attention, is proposed to intensify the obtained features both globally and locally, resulting in improved depth maps with sharper boundaries. Finally, a reweighted loss strategy based on multi-level outputs is proposed to conduct effective supervision for self-supervised depth estimation. Experimental results demonstrate that our MLDA-Net framework achieves state-of-the-art depth prediction results on the KITTI benchmark for self-supervised monocular depth estimation with different input modes and training modes. Extensive experiments on other benchmark datasets further confirm the superiority of our proposed approach.
Xibin Song, Wei Li 0143, Dingfu Zhou, Yuchao Dai, Hongdong Li, Liangjun Zhang
IEEE Trans. Image Process.1
2020 RotPredictor: Unsupervised Canonical Viewpoint Learning for Point Cloud Classification
abstract
Recently, significant progress has been achieved in analyzing the 3D point cloud with deep learning techniques. However, existing networks suffer from poor generalization and robustness to arbitrary rotations applied to the input point cloud. Different from traditional strategies that improve the rotation robustness with data augmentation or specifically designed spherical representation or harmonics-based kernels, we propose to rotate the point cloud into a canonical viewpoint for boosting the following downstream target task, e.g., object classification and part segmentation. Specifically, the canonical viewpoint is predicted by the network RotPredictor in an unsupervised way and the loss function is only built on the target task. Our RotPredictor satisfies the rotation equivariance property in (3) approximately and the predication output has the linear relationship with the applied rotation transformation. In addition, the RotPredictor is an independent plug and play module, which can be employed by any point-based deep learning framework without extra burden. Experimental results on the public model classification dataset ModelNet40 show the performance for all baselines can be boosted by integrating the proposed module. In addition, by adding our proposed module, we can achieve the state-of-the-art classification accuracy with 90.2% on the rotation-augmented ModelNet40 benchmark.
Dingfu Zhou, Xibin Song, Shengze Jin, Ruigang Yang, Liangjun Zhang
3DV3
2020 IAFA: Instance-Aware Feature Aggregation for 3D Object Detection from a Single Image
Dingfu Zhou, Xibin Song, Yuchao Dai, Junbo Yin, Feixiang Lu, Miao Liao, Liangjun Zhang
ACCV (1)2
2020 Channel Attention Based Iterative Residual Learning for Depth Map Super-Resolution
abstract
Despite the remarkable progresses made in deep learning based depth map super-resolution (DSR), how to tackle real-world degradation in low-resolution (LR) depth maps remains a major challenge. Existing DSR model is generally trained and tested on synthetic dataset, which is very different from what would get from a real depth sensor. In this paper, we argue that DSR models trained under this setting are restrictive and not effective in dealing with realworld DSR tasks. We make two contributions in tackling real-world degradation of different depth sensors. First, we propose to classify the generation of LR depth maps into two types: non-linear downsampling with noise and interval downsampling, for which DSR models are learned correspondingly. Second, we propose a new framework for real-world DSR, which consists of four modules : 1) An iterative residual learning module with deep supervision to learn effective high-frequency components of depth maps in a coarse-to-fine manner; 2) A channel attention strategy to enhance channels with abundant high-frequency components; 3) A multi-stage fusion module to effectively reexploit the results in the coarse-to-fine process; and 4) A depth refinement module to improve the depth map by TGV regularization and input loss. Extensive experiments on benchmarking datasets demonstrate the superiority of our method over current state-of-the-art DSR methods.
Xibin Song, Yuchao Dai, Dingfu Zhou, Liu Liu 0009, Wei Li 0111, Hongdong Li, Ruigang Yang
CVPR1
2020 Joint 3D Instance Segmentation and Object Detection for Autonomous Driving
abstract
Currently, in Autonomous Driving (AD), most of the 3D object detection frameworks (either anchor- or anchor-free-based) consider the detection as a Bounding Box (BBox) regression problem. However, this compact representation is not sufficient to explore all the information of the objects. To tackle this problem, we propose a simple but practical detection framework to jointly predict the 3D BBox and instance segmentation. For instance segmentation, we propose a Spatial Embeddings (SEs) strategy to assemble all foreground points into their corresponding object centers. Base on the SE results, the object proposals can be generated based on a simple clustering strategy. For each cluster, only one proposal is generated. Therefore, the Non-Maximum Suppression (NMS) process is no longer needed here. Finally, with our proposed instance-aware ROI pooling, the BBox is refined by a second-stage network. Experimental results on the public KITTI dataset show that the proposed SEs can significantly improve the instance segmentation results compared with other feature embedding-based method. Meanwhile, it also outperforms most of the 3D object detectors on the KITTI testing benchmark.
Dingfu Zhou, Xibin Song, Liu Liu 0009, Junbo Yin, Yuchao Dai, Hongdong Li, Ruigang Yang
CVPR3
2019 IoU Loss for 2D/3D Object Detection
abstract
In the 2D/3D object detection task, Intersection-over-Union (IoU) has been widely employed as an evaluation metric to evaluate the performance of different detectors in the testing stage. However, during the training stage, the common distance loss (e.g, L_1 or L_2) is often adopted as the loss function to minimize the discrepancy between the predicted and ground truth Bounding Box (Bbox). To eliminate the performance gap between training and testing, the IoU loss has been introduced for 2D object detection in [1] and [2]. Unfortunately, all these approaches only work for axis-aligned 2D Boxes, which cannot be applied for more general object detection task with rotated Boxes. To resolve this issue, we investigate the IoU computation for two rotated Boxes first and then implement a unified framework, IoU loss layer for both 2D and 3D object detection tasks. By integrating the implemented IoU loss into several state-of-the-art 3D object detectors, consistent improvements have been achieved for both bird-eye-view 2D detection and point cloud 3D detection on the public KITTI [3] benchmark.
Dingfu Zhou, Xibin Song, Chenye Guan, Junbo Yin, Yuchao Dai, Ruigang Yang
3DV3
2019 ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving
abstract
Autonomous driving has attracted remarkable attention from both industry and academia. An important task is to estimate 3D properties (e.g. translation, rotation and shape) of a moving or parked vehicle on the road. This task, while critical, is still under-researched in the computer vision community – partially owing to the lack of large scale and fully-annotated 3D car database suitable for autonomous driving research. In this paper, we contribute the first large scale database suitable for 3D car instance understanding – ApolloCar3D. The dataset contains 5,277 driving images and over 60K car instances, where each car is fitted with an industry-grade 3D CAD model with absolute model size and semantically labelled keypoints. This dataset is above 20× larger than PASCAL3D+ and KITTI, the current state-of-the-art. To enable efficient labelling in 3D, we build a pipeline by considering 2D-3D keypoint correspondences for a single instance and 3D relationship among multiple instances. Equipped with such dataset, we build various baseline algorithms with the state-of-the-art deep convolutional neural networks. Specifically, we first segment each car with a pre-trained Mask R-CNN, and then regress towards its 3D pose and shape based on a deformable 3D car model with or without using semantic keypoints. We show that using keypoints significantly improves fitting performance. Finally, we develop a new 3D metric jointly considering 3D pose and 3D shape, allowing for comprehensive evaluation and ablation study.
Xibin Song, Peng Wang 0001, Dingfu Zhou, Chenye Guan, Yuchao Dai, Hongdong Li, Ruigang Yang
CVPR1
2019 Deeply Supervised Depth Map Super-Resolution as Novel View Synthesis
abstract
Deep convolutional neural network (DCNN) has been successfully applied to depth map super-resolution and outperforms existing methods by a wide margin. However, there still exist two major issues with these DCNN-based depth map super-resolution methods that hinder the performance: 1) the low-resolution depth maps either need to be up-sampled before feeding into the network or substantial deconvolution has to be used and 2) the supervision (high-resolution depth maps) is only applied at the end of the network, thus it is difficult to handle large up-sampling factors, such as x8 and x16. In this paper, we propose a new framework to tackle the above problems. First, we propose to represent the task of depth map superresolution as a series of novel view synthesis sub-tasks. The novel view synthesis sub-task aims at generating (synthesizing) a depth map from a different camera pose, which could be learned in parallel. Second, to handle large up-sampling factors, we present a deeply supervised network structure to enforce strong supervision in each stage of the network. Third, a multiscale fusion strategy is proposed to effectively exploit the feature maps at different scales and handle the blocking effect. In this way, our proposed framework could deal with challenging depth map super-resolution efficiently under large up-sampling factors (e.g., x8 and x16). Our method only uses the low-resolution depth map as input, and the support of color image is not needed, which greatly reduces the restriction of our method. Extensive experiments on various benchmarking data sets demonstrate the superiority of our method over current state-of-the-art depth map super-resolution methods.
Xibin Song, Yuchao Dai, Xueying Qin
IEEE Trans. Circuits Syst. Video Technol.1
2018 Modeling deviations of rgb-d cameras for accurate depth map and color image registration
Xibin Song, Jianmin Zheng, Fan Zhong 0001, Xueying Qin
Multim. Tools Appl.1
2016 Deep Depth Super-Resolution: Learning Depth Super-Resolution Using Deep Convolutional Neural Network
Xibin Song, Yuchao Dai, Xueying Qin
ACCV (4)1
2016 Edge-guided depth map enhancement
abstract
Low-cost depth sensing devices, such as Microsoft Kinect, can only produce noisy depth maps that are mis-aligned with color images, and even contain many holes. Even though the coupled high quality color images contain rich information which can be exploited to enhance the depth maps, the redundant color edges often introduce incorrect depth edges in the result depth map, since color images contain more textures than depth maps. To solve this problem, we propose a novel approach which generates accurate color-consistent depth edges by employing both color and depth images. First, Edges of raw depth maps are extracted using image pyramid strategy. Then, the redundant edges in color images are removed according to the raw depth edges, and, accurate color-consistent depth edges are generated by combining raw depth edges with current color edges. Finally, constraints extracted from both raw depth and color images and the generated depth edges are fused in a MRF optimization framework to obtain the enhanced depth map, which is accurately aligned with coupled color image. As experimentally demonstrated, the proposed method achieves outstanding performance when compared with previous approaches.
Xibin Song, Fan Zhong 0001, Xueying Qin
ICPR1
2014 Estimation of Kinect depth confidence through self-training
Xibin Song, Fan Zhong 0001, Yanke Wang, Xueying Qin
Vis. Comput.1