Xiaoqing Ye

dblp:177/0181 · DBLP profile ↗
← Back
64ranked-venue papers
9as first author
54since 2021 · last 2026
0000-0003-3268-880XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 6 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 4 first-author · 33 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 SOOD++: Leveraging Unlabeled Data to Boost Oriented Object Detection
abstract
Semi-supervised object detection (SSOD), leveraging unlabeled data to boost object detectors, has become a hot topic recently. However, existing SSOD approaches mainly focus on horizontal objects, leaving oriented objects common in aerial images unexplored. At the same time, the annotation cost of oriented objects is significantly higher than that of their horizontal counterparts (an approximate 36.5% increase in costs). Therefore, in this paper, we propose a simple yet effective Semi-supervised Oriented Object Detection method termed SOOD++. Specifically, we observe that objects from aerial images usually have arbitrary orientations, small scales, and dense distribution, which inspires the following core designs: a Simple Instance-aware Dense Sampling (SIDS) strategy is used to generate comprehensive dense pseudo-labels; the Geometry-aware Adaptive Weighting (GAW) loss dynamically modulates the importance of each pair between pseudo-label and corresponding prediction by leveraging the intricate geometric information of aerial objects; we treat aerial images as global layouts and explicitly build the many-to-many relationship between the sets of pseudo-labels and predictions via the proposed Noise-driven Global Consistency (NGC). Extensive experiments conducted on various oriented object datasets under various labeled settings demonstrate the effectiveness of our method. For example, on the DOTA-V2.0/DOTA-V1.5 benchmark, the proposed method outperforms previous state-of-the-art (SOTA) by a large margin (+2.90/2.14, +2.16/2.18, and +2.66/2.32) mAP under 10%, 20%, and 30% labeled data settings, respectively, with single-scale training and testing. More importantly, it still improves upon a strong supervised baseline with 70.66 mAP, trained using the full DOTA-V1.5 train-val set, by +1.82 mAP, resulting in a 72.48 mAP, pushing the new state-of-the-art. Moreover, our method demonstrates stable generalization ability across different oriented detectors, even for multi-view oriented 3D object detectors.
Dingkang Liang, Wei Hua 0005, Chunsheng Shi, Zhikang Zou, Xiaoqing Ye, Xiang Bai
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Competitive firm prediction: A novel heterogeneous graph sampled weighted attention network
Xiaoqing Ye, Dun Liu, Tianrui Li 0001
Pattern Recognit.1
2026 Self-Supervised Aggregation Framework for Text-Attributed Heterogeneous Graphs Representation
abstract
Text-Attributed Heterogeneous Graphs (TAHGs) integrate topological relationships with rich textual node attributes, offering expressive representations for complex multi-faceted data. While recent methods jointly leverage textual and structural information, they still face two critical limitations: (i) existing approaches are constrained to neighborhood modeling, failing to capture semantic dependencies in higher-order topologies; (ii) current techniques exhibit inadequate unified alignment strategies, limiting dynamic interaction between cross modalities. To address these challenges, we propose SATH, a self-supervised information aggregation model for TAHGs, designed to effectively leverage textual and structural information within TAHGs. SATH aggregates higher-order neighbor textual attributes through comparative learning, and dynamically aligns these attributes to higher-order topologies through a unified strategy. This approach integrates both types of information effectively, enhancing the expressiveness and discriminative capability of the learned node representations in downstream tasks. Extensive experiments on real-world datasets demonstrate that SATH significantly outperforms baseline models while eliminating the need for manual meta-path design or text feature concatenation. It also improves efficiency and scalability on large-scale TAHGs, achieving superior representation quality in TAHG-based tasks.
Fei Teng 0001, Quyan Xiao, Xingwang Li 0003, Xiaoqing Ye, Qian Li 0033
IEEE Trans. Knowl. Data Eng.4
2025 U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and Registration
abstract
Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and construction sites significantly degrade GNSS (Global Navigation Satellite System) signal quality. This issue underscores the importance of visual localization techniques in scenarios where GNSS signals are unreliable. This paper proposes U-ViLAR, a novel uncertainty-aware visual localization framework designed to address these challenges while enabling adaptive localization using high-definition (HD) maps or navigation maps. Specifically, our method first extracts features from the input visual data and maps them into Bird's-Eye-View (BEV) space to enhance spatial consistency with the map input. Subsequently, we introduce: a) Perceptual Uncertainty-guided Association, which mitigates errors caused by perception uncertainty, and b) Localization Uncertainty-guided Registration, which reduces errors introduced by localization uncertainty. By effectively balancing the coarse-grained large-scale localization capability of association with the fine-grained precise localization capability of registration, our approach achieves robust and accurate localization. Experimental results demonstrate that our method achieves state-of-the-art performance across multiple localization tasks. Furthermore, our model has undergone rigorous testing on large-scale autonomous driving fleets and has demonstrated stable performance in various challenging urban scenarios.
Chenming Wu, Jiang-Jiang Liu 0001, Haibao Yu, Xiaoqing Ye, Shirui Li, Ji Wan
ICCV8
2025 Uni2Det: Unified and Universal Framework for Prompt-Guided Multi-dataset 3D Detection
abstract
We present Uni$^2$Det, a brand new framework for unified and universal multi-dataset training on 3D detection, enabling robust performance across diverse domains and generalization to unseen domains. Due to substantial disparities in data distribution and variations in taxonomy across diverse domains, training such a detector by simply merging datasets poses a significant challenge. Motivated by this observation, we introduce multi-stage prompting modules for multi-dataset 3D detection, which leverages prompts based on the characteristics of corresponding datasets to mitigate existing differences. This elegant design facilitates seamless plug-and-play integration within various advanced 3D detection frameworks in a unified manner, while also allowing straightforward adaptation for universal applicability across datasets. Experiments are conducted across multiple dataset consolidation scenarios involving KITTI, Waymo, and nuScenes, demonstrating that our Uni$^2$Det outperforms existing methods by a large margin in multi-dataset training. Notably, results on zero-shot cross-dataset transfer validate the generalization capability of our proposed method. Our code is available at https://github.com/ThomasWangY/Uni2Det.
Zhikang Zou, Xiaoqing Ye, Xiao Tan 0001, Errui Ding, Cairong Zhao
ICLR3
2025 Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving
abstract
The autoregressive world model exhibits robust generalization capabilities in vectorized scene understanding but encounters difficulties in deriving actions due to insufficient uncertainty modeling and self-delusion. In this paper, we explore the feasibility of deriving decisions from an autoregres-sive world model by addressing these challenges through the formulation of multiple probabilistic hypotheses. We propose LatentDriver, a framework models the environment's next states and the ego vehicle's possible actions as a mixture distribution, from which a deterministic control signal is then derived. By incorporating mixture modeling, the stochastic nature of decision-making is captured. Additionally, the self-delusion problem is mitigated by providing intermediate actions sampled from a distribution to the world model. Experimen-tal results on the recently released closed-loop benchmark Waymax demonstrate that LatentDriver surpasses state-of-the-art reinforcement learning and imitation learning methods, achieving expert-level performance. The code and models will be made available at https://github.com/Sephirex-X/LatentDriver.
Lingyu Xiao, Jiang-Jiang Liu 0001, Xiaoqing Ye, Wankou Yang, Jingdong Wang 0001
ICRA5
2025 Explore the LiDAR-Camera Dynamic Adjustment Fusion for 3D Object Detection
abstract
Camera and LiDAR serve as informative sensors for accurate and robust autonomous driving systems. However, these sensors often exhibit heterogeneous natures, resulting in distributional modality gaps that present significant challenges for fusion. To address this, a robust fusion technique is crucial, particularly for enhancing 3D object detection. In this paper, we introduce a dynamic adjustment technology aimed at aligning modal distributions and learning effective modality representations to enhance the fusion process. Specifically, we propose a triphase domain aligning module. This module adjusts the feature distributions from both the camera and LiDAR, bringing them closer to the ground truth domain and minimizing differences. Additionally, we explore improved representation acquisition methods for dynamic fusion, which includes modal interaction and specialty enhancement. Finally, an adaptive learning technique that merges the semantics and geometry information for dynamical instance optimization. Extensive experiments in the nuScenes dataset present competitive performance with state-of-the-art approaches. Our code will be released in the future.
Xin Hao, Yifeng Shi, Xiao Tan 0001, Xiaoqing Ye
ICRA7
2025 An Empirical Study of Ground Segmentation for 3-D Object Detection
abstract
The ratio of foreground and background points directly impacts the accuracy and speed of the lidar-based 3D object detection methods. However, existing methods generally ignore the impact of ground points. Although some traditional ground segmentation algorithms are available to remove ground point clouds, they usually suffer from over-segmentation, which leads to a sub-optimal and even worse performance for the downstream 3D detection task. We conduct an in-depth analysis and attribute this phenomenon to the reason that some crucial foreground points attached to the ground (e.g., the wheels of Cars, or the feet of Pedestrians) are directly removed due to over-segmentation. To this end, we propose a new Attached Point Restoring (APR) module to recover these discarded foreground points. Experimental results demonstrate the effectiveness and generalization of APR by integrating it into various ground segmentation algorithms to boost the performance or the running time of 3D detection on KITTI and Waymo datasets. Finally, we hope this paper can serve as a new guide to inspire future research in this field. Code is available athttps://github.com/yhc2021/GPR.
Hongcheng Yang, Dingkang Liang, Zhe Liu 0033, Zhikang Zou, Xiaoqing Ye, Xiang Bai
IEEE Trans. Intell. Transp. Syst.6
2025 Coupling and Decoupling: Towards Temporal Feedback for 3D Object Detection
abstract
3D object detection has garnered significant attention within the academic community, primarily due to its broad utility in domains such as autonomous driving and robotics. Prior research efforts have predominantly concentrated on leveraging temporal contextual information embedded within sequential data to enhance the current feature representations. However, a notable limitation of these endeavors lies in their inadequate treatment of the inherent noise present within historical sequences, thereby constraining the efficiency of fusion methods. In this paper, we propose a new temporal feedback network, named TFNet, to model and correct the temporal noise by designing acoupling-decouplingmechanism. Central to our approach are two distinct modules: (i) Foreground Feature Enhancement, which amplifies sparse instance details across temporal frames, thereby furnishing essential local information priors for subsequent fusion; and (ii) Coupling-Decoupling Feature Interaction, designed to first aggregate temporal contextual information and then disentangle fusion features into frame-specific representations. Leveraging a feedback strategy, this module can adaptively enhance useful information and eliminate noise within individual frame features. Empirical evaluations conducted on the nuScenes benchmark demonstrate the effectiveness of TFNet, achieving the new state-of-the-art performance without any bells and whistles.
Yubo Cui, Zhikang Zou, Xiaoqing Ye, Xiao Tan 0001, Zhiheng Li 0003, Zheng Fang 0001
IEEE Trans. Multim.3
2025 CLIP-GS: CLIP-Informed Gaussian Splatting for View-Consistent 3D Indoor Semantic Understanding
abstract
Exploiting 3D Gaussian Splatting (3DGS) with Contrastive Language-Image Pre-Training (CLIP) models for open-vocabulary 3D semantic understanding of indoor scenes has emerged as an attractive research focus. Existing methods typically attach high-dimensional CLIP semantic embeddings to 3D Gaussians and leverage view-inconsistent 2D CLIP semantics as Gaussian supervision, resulting in efficiency bottlenecks and deficient 3D semantic consistency. To address these challenges, we present CLIP-GS, efficiently achieving a coherent semantic understanding of 3D indoor scenes via the proposed Semantic Attribute Compactness (SAC) and 3D Coherent Regularization (3DCR). SAC approach exploits the naturally unified semantics within objects to learn compact, yet effective, semantic Gaussian representations, enabling highly efficient rendering (>100 FPS). 3DCR enforces semantic consistency in 2D and 3D domains: In 2D, 3DCR utilizes refined view-consistent semantic outcomes derived from 3DGS to establish cross-view coherence constraints; in 3D, 3DCR encourages features similar among 3D Gaussian primitives associated with the same object, leading to more precise and coherent segmentation results. Extensive experimental results demonstrate that our method remarkably suppresses existing state-of-the-art approaches, achieving mIoU improvements of 21.20% and 13.05% on ScanNet and Replica datasets, respectively, while maintaining real-time rendering speed. Furthermore, our approach exhibits superior performance even with sparse input data, substantiating its robustness.
Guibiao Liao, Jiankun Li, Zhenyu Bao, Xiaoqing Ye, Qing Li 0029, Kanglin Liu
ACM Trans. Multim. Comput. Commun. Appl.4
2024 VLM2Scene: Self-Supervised Image-Text-LiDAR Learning with Foundation Models for Autonomous Driving Scene Understanding
abstract
Vision and language foundation models (VLMs) have showcased impressive capabilities in 2D scene understanding. However, their latent potential in elevating the understanding of 3D autonomous driving scenes remains untapped. In this paper, we propose VLM2Scene, which exploits the potential of VLMs to enhance 3D self-supervised representation learning through our proposed image-text-LiDAR contrastive learning strategy. Specifically, in the realm of autonomous driving scenes, the inherent sparsity of LiDAR point clouds poses a notable challenge for point-level contrastive learning methods. This method often grapples with limitations tied to a restricted receptive field and the presence of noisy points. To tackle this challenge, our approach emphasizes region-level learning, leveraging regional masks without semantics derived from the vision foundation model. This approach capitalizes on valuable contextual information to enhance the learning of point cloud representations. First, we introduce Region Caption Prompts to generate fine-grained language descriptions for the corresponding regions, utilizing the language foundation model. These region prompts then facilitate the establishment of positive and negative text-point pairs within the contrastive loss framework. Second, we propose a Region Semantic Concordance Regularization, which involves a semantic-filtered region learning and a region semantic assignment strategy. The former aims to filter the false negative samples based on the semantic distance, and the latter mitigates potential inaccuracies in pixel semantics, thereby enhancing overall semantic consistency. Extensive experiments on representative autonomous driving datasets demonstrate that our self-supervised method significantly outperforms other counterparts. Codes are available at https://github.com/gbliao/VLM2Scene.
Guibiao Liao, Jiankun Li, Xiaoqing Ye
AAAI3
2024 BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object Detection
abstract
Vision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it en-compasses inherent advantages in reducing blind spots and expanding perception range. While previous work mainly focuses on accurately estimating depth or height for 2D-to-3D mapping, ignoring the position approximation error in the voxel pooling process. Inspired by this insight, we propose a novel voxel pooling strategy to reduce such error, dubbed BEVSpread. Specifically, instead of bringing the image features contained in a frustum point to a single BEV grid, BEVSpread considers each frustum point as a source and spreads the image features to the surrounding BEV grids with adaptive weights. To achieve superior prop- agation performance, a specific weight function is designed to dynamically control the decay speed of the weights according to distance and depth. Aided by customized CUDA parallel acceleration, BEVSpread achieves comparable inference time as the original voxel pooling. Extensive experiments on two large-scale roadside benchmarks demonstrate that, as a plug-in, BEVSpread can significantly improve the performance of existing frustum-based BEV methods by a large margin of (1.12, 5.26, 3.01) AP in vehicle, pedestrian and cyclist. The source code will be made publicly available at BEVSpread.
Yehao Lu, Guangcong Zheng, Shuigen Zhan, Xiaoqing Ye, Zichang Tan, Jingdong Wang 0001, Gaoang Wang, Xi Li 0001
CVPR5
2024 OPEN: Object-Wise Position Embedding for Multi-view 3D Object Detection
Jinghua Hou, Xiaoqing Ye, Zhe Liu 0033, Shi Gong, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Xiang Bai
ECCV (26)3
2024 DrivingDiffusion: Layout-Guided Multi-view Driving Scenarios Video Generation with Latent Diffusion Model
Xiaoqing Ye
ECCV (78)3
2024 SEED: A Simple and Effective 3D DETR in Point Clouds
Zhe Liu 0033, Jinghua Hou, Xiaoqing Ye, Jingdong Wang 0001, Xiang Bai
ECCV (11)3
2024 Make Your ViT-Based Multi-view 3D Detectors Faster via Token Compression
Dingyuan Zhang, Dingkang Liang, Zichang Tan, Xiaoqing Ye, Cheng Zhang 0020, Jingdong Wang 0001, Xiang Bai
ECCV (47)4
2024 LION: Linear Group RNN for 3D Object Detection in Point Clouds
abstract
The benefit of transformers in large-scale 3D point cloud perception tasks, such as 3D object detection, is limited by their quadratic computation cost when modeling long-range relationships. In contrast, linear RNNs have low computational complexity and are suitable for long-range modeling. Toward this goal, we propose a simple and effective window-based framework built on Linear group RNN (i.e., perform linear RNN for grouped features) for accurate 3D object detection, called LION. The key property is to allow sufficient feature interaction in a much larger group than transformer-based methods. However, effectively applying linear group RNN to 3D object detection in highly sparse point clouds is not trivial due to its limitation in handling spatial modeling. To tackle this problem, we simply introduce a 3D spatial feature descriptor and integrate it into the linear group RNN operators to enhance their spatial features rather than blindly increasing the number of scanning orders for voxel features. To further address the challenge in highly sparse point clouds, we propose a 3D voxel generation strategy to densify foreground features thanks to linear group RNN as a natural property of auto-regressive models. Extensive experiments verify the effectiveness of the proposed components and the generalization of our LION on different linear group RNN operators including Mamba, RWKV, and RetNet. Furthermore, it is worth mentioning that our LION-Mamba achieves state-of-the-art on Waymo, nuScenes, Argoverse V2, and ONCE datasets. Last but not least, our method supports kinds of advanced linear RNN operators (e.g., RetNet, RWKV, Mamba, xLSTM and TTT) on small but popular KITTI dataset for a quick experience with our linear RNN-based framework.
Zhe Liu 0033, Jinghua Hou, Xinyu Wang 0024, Xiaoqing Ye, Jingdong Wang 0001, Hengshuang Zhao, Xiang Bai
NeurIPS4
2024 PointMamba: A Simple State Space Model for Point Cloud Analysis
abstract
Transformers have become one of the foundational architectures in point cloud analysis tasks due to their excellent global modeling ability. However, the attention mechanism has quadratic complexity, making the design of a linear complexity method with global modeling appealing. In this paper, we propose PointMamba, transferring the success of Mamba, a recent representative state space model (SSM), from NLP to point cloud analysis tasks. Unlike traditional Transformers, PointMamba employs a linear complexity algorithm, presenting global modeling capacity while significantly reducing computational costs. Specifically, our method leverages space-filling curves for effective point tokenization and adopts an extremely simple, non-hierarchical Mamba encoder as the backbone. Comprehensive evaluations demonstrate that PointMamba achieves superior performance across multiple datasets while significantly reducing GPU memory usage and FLOPs. This work underscores the potential of SSMs in 3D vision-related tasks and presents a simple yet effective Mamba-based baseline for future research. The code is available at https://github.com/LMD0311/PointMamba.
Dingkang Liang, Xin Zhou 0013, Wei Xu 0037, Xingkui Zhu, Zhikang Zou, Xiaoqing Ye, Xiao Tan 0001, Xiang Bai
NeurIPS6
2024 SAM3D: zero-shot 3D object detection via the segment anything model
Dingyuan Zhang, Dingkang Liang, Hongcheng Yang, Zhikang Zou, Xiaoqing Ye, Zhe Liu 0033, Xiang Bai
Sci. China Inf. Sci.5
2024 Self-ensembling depth completion via density-aware consistency
Xuanmeng Zhang, Zhedong Zheng, Minyue Jiang, Xiaoqing Ye
Pattern Recognit.4
2024 Multi-Modal 3D Object Detection by Box Matching
abstract
Multi-modal 3D object detection has received growing attention as the information from different sensors like LiDAR and cameras are complementary. Most fusion methods for 3D detection rely on an accurate alignment and calibration between 3D point clouds and RGB images. However, such an assumption is not reliable in a real-world self-driving system, as the alignment between different modalities is easily affected by asynchronous sensors and disturbed sensor placement. We propose a novel Fusion network by Box Matching (FBMNet) for multi-modal 3D detection, which provides an alternative way for cross-modal feature alignment by learning the correspondence at the bounding box level to free up the dependency of calibration during inference. With the learned assignments between 3D and 2D object proposals, the fusion for detection can be effectively performed by combining their ROI features. Extensive experiments on the nuScenes dataset demonstrate that our method is much more robust in dealing with challenging cases such as asynchronous sensors, misaligned sensor placement, and degenerated camera images than existing fusion methods. We hope that our FBMNet could provide an available solution to dealing with these challenging cases for safety in real autonomous driving scenarios.
Zhe Liu 0033, Xiaoqing Ye, Zhikang Zou, Xinwei He 0001, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Xiang Bai
IEEE Trans. Intell. Transp. Syst.2
2024 A Multisource Data Fusion-based Heterogeneous Graph Attention Network for Competitor Prediction
abstract
Competitor identification is an essential component of corporate strategy. With the rapid development of artificial intelligence, various data-mining methodologies and frameworks have emerged to identify competitors. In general, the competitiveness among companies is determined by both market commonality and resource similarity. However, because resource information is more difficult to obtain than market information, existing studies primarily identify competitors via market commonality. To address this limitation, we introduce multisource company descriptions as well as heterogeneous business relationships, and we propose a novel method for simultaneously mining the market commonality and resource similarity. First, we use multisource company descriptions to represent companies and transform the heterogeneous business relationships into a heterogeneous business network. Then, we propose a novel multisource data fusion-based heterogeneous graph attention network (MHGAT) to learn the pairwise competitive relationships between companies. Specifically, a graph neural network-based model is proposed to learn the embeddings of companies by preserving their competition, and a multilevel attention framework is designed to integrate the embeddings from neighboring company level, heterogeneous relationship level, and multisource description level. Finally, experiments on a real-world dataset verify the effectiveness of our proposed MHGAT and demonstrate the usefulness of company descriptions and business relationships in competitor identification.
Xiaoqing Ye, Dun Liu, Tianrui Li 0001
ACM Trans. Knowl. Discov. Data1
2023 StereoDistill: Pick the Cream from LiDAR for Distilling Stereo-Based 3D Object Detection
abstract
In this paper, we propose a cross-modal distillation method named StereoDistill to narrow the gap between the stereo and LiDAR-based approaches via distilling the stereo detectors from the superior LiDAR model at the response level, which is usually overlooked in 3D object detection distillation. The key designs of StereoDistill are: the X-component Guided Distillation~(XGD) for regression and the Cross-anchor Logit Distillation~(CLD) for classification. In XGD, instead of empirically adopting a threshold to select the high-quality teacher predictions as soft targets, we decompose the predicted 3D box into sub-components and retain the corresponding part for distillation if the teacher component pilot is consistent with ground truth to largely boost the number of positive predictions and alleviate the mimicking difficulty of the student model. For CLD, we aggregate the probability distribution of all anchors at the same position to encourage the highest probability anchor rather than individually distill the distribution at the anchor level. Finally, our StereoDistill achieves state-of-the-art results for stereo-based 3D detection on the KITTI test benchmark and extensive experiments on KITTI and Argoverse Dataset validate the effectiveness.
Zhe Liu 0033, Xiaoqing Ye, Xiao Tan 0001, Errui Ding, Xiang Bai
AAAI2
2023 Command-driven Articulated Object Understanding and Manipulation
abstract
We present Cart, a new approach towards articulatedobject manipulations by human commands. Beyond the existing work that focuses on inferring articulation structures, we further support manipulating articulated shapes to align them subject to simple command templates. The key of Cart is to utilize the prediction of object structures to connect visual observations with user commands for effective manipulations. It is achieved by encoding command messages for motion prediction and a test-time adaptation to adjust the amount of movement from only command supervision. For a rich variety of object categories, Cart can accurately manipulate object shapes and outperform the state-of-the-art approaches in understanding the inherent articulation structures. Also, it can well generalize to unseen object categories and real-world objects. We hope Cart could open new directions for instructing machines to operate articulated objects. Code is available at https://github.com/dvlab-research/Cart.
Ruihang Chu, Zhengzhe Liu, Xiaoqing Ye, Xiao Tan 0001, Xiaojuan Qi 0001, Chi-Wing Fu, Jiaya Jia
CVPR3
2023 SOOD: Towards Semi-Supervised Oriented Object Detection
abstract
Semi-Supervised Object Detection (SSOD), aiming to explore unlabeled data for boosting object detectors, has become an active task in recent years. However, existing SSOD approaches mainly focus on horizontal objects, leaving multi-oriented objects that are common in aerial images unexplored. This paper proposes a novel Semi-supervised Oriented Object Detection model, termed SOOD, built upon the mainstream pseudo-labeling framework. Towards oriented objects in aerial scenes, we design two loss functions to provide better supervision. Focusing on the orientations of objects, the first loss regularizes the consistency between each pseudo-label-prediction pair (includes a prediction and its corresponding pseudo label) with adaptive weights based on their orientation gap. Focusing on the layout of an image, the second loss regularizes the similarity and explicitly builds the many-to-many relation between the sets of pseudo-labels and predictions. Such a global consistency constraint can further boost semi-supervised learning. Our experiments show that when trained with the two proposed losses, SOOD surpasses the state-of-the-art SSOD methods under various settings on the DOTAv1.5 benchmark. The code will be available at https://github.com/HamPerdredes/SOOD.
Wei Hua 0005, Dingkang Liang, Zhikang Zou, Xiaoqing Ye, Xiang Bai
CVPR6
2023 CrowdCLIP: Unsupervised Crowd Counting via Vision-Language Model
abstract
Supervised crowd counting relies heavily on costly manual labeling, which is difficult and expensive, especially in dense scenes. To alleviate the problem, we propose a novel unsupervised framework for crowd counting, named CrowdCLIP. The core idea is built on two observations: 1) the recent contrastive pre-trained vision-language model (CLIP) has presented impressive performance on various downstream tasks; 2) there is a natural mapping between crowd patches and count text. To the best of our knowledge, CrowdCLIP is the first to investigate the vision-language knowledge to solve the counting problem. Specifically, in the training stage, we exploit the multi-modal ranking loss by constructing ranking text prompts to match the size-sorted crowd patches to guide the image encoder learning. In the testing stage, to deal with the diversity of image patches, we propose a simple yet effective progressive filtering strategy to first select the highly potential crowd patches and then map them into the language space with various counting intervals. Extensive experiments on five challenging datasets demonstrate that the proposed CrowdCLIP achieves superior performance compared to previous unsupervised state-of-the-art counting methods. Notably, CrowdCLIP even surpasses some pop-ular fully-supervised methods under the cross-dataset setting. The source code will be available at https://github.com/dk-liang/CrowdCLIP.
Dingkang Liang, Zhikang Zou, Xiaoqing Ye, Wei Xu 0037, Xiang Bai
CVPR4
2023 CAPE: Camera View Position Embedding for Multi-View 3D Object Detection
abstract
In this paper, we address the problem of detecting 3D ob-jects from multi-view images. Current query-based methods rely on global 3D position embeddings (PE) to learn the ge-ometric correspondence between images and 3D space. We claim that directly interacting 2D image features with global 3D PE could increase the difficulty of learning view trans-formation due to the variation of camera extrinsics. Thus we propose a novel method based on CAmera view Position Embedding, called CAPE. We form the 3D position embed-dings under the local camera-view coordinate system instead of the global coordinate system, such that 3D position em-bedding is free of encoding camera extrinsic parameters. Furthermore, we extend our CAPE to temporal modeling by exploiting the object queries of previous frames and encoding the ego motion for boosting 3D object detection. CAPE achieves the state-of-the-art performance (61.0% NDS and 52.5% mAP) among all LiDAR-free methods on nuScenes dataset. Codes and models are available.11Codes of Paddle3D and PyTorch Implementation.
Kaixin Xiong, Shi Gong, Xiaoqing Ye, Xiao Tan 0001, Ji Wan, Errui Ding, Jingdong Wang 0001, Xiang Bai
CVPR3
2023 Forward Flow for Novel View Synthesis of Dynamic Scenes
abstract
This paper proposes a neural radiance field (NeRF) approach for novel view synthesis of dynamic scenes using forward warping. Existing methods often adopt a static NeRF to represent the canonical space, and render dynamic images at other time steps by mapping the sampled 3D points back to the canonical space with the learned backward flow field. However, this backward flow field is non-smooth and discontinuous, which is difficult to be fitted by commonly used smooth motion models. To address this problem, we propose to estimate the forward flow field and directly warp the canonical radiance field to other time steps. Such forward flow field is smooth and continuous within the object region, which benefits the motion model learning. To achieve this goal, we represent the canonical radiance field with voxel grids to enable efficient forward warping, and propose a differentiable warping process, including an average splatting operation and an inpaint network, to resolve the many-to-one and one-to-many mapping issues. Thorough experiments show that our method outperforms existing methods in both novel view rendering and motion modeling, demonstrating the effectiveness of our forward flow motion modeling. Project page: https://npucvr.github.io/ForwardFlowDNeRF.
Jiadai Sun, Yuchao Dai, Guanying Chen, Xiaoqing Ye, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001
ICCV5
2023 IST-Net: Prior-free Category-level Pose Estimation with Implicit Space Transformation
abstract
Category-level 6D pose estimation aims to predict the poses and sizes of unseen objects from a specific category. Thanks to prior deformation, which explicitly adapts a category-specific 3D prior (i.e., a 3D template) to a given object instance, prior-based methods attained great success and have become a major research stream. However, obtaining category-specific priors requires collecting a large amount of 3D models, which is labor-consuming and often not accessible in practice. This motivates us to investigate whether priors are necessary to make prior-based methods effective. Our empirical study shows that the 3D prior itself is not the credit to the high performance. The keypoint actually is the explicit deformation process, which aligns camera and world coordinates supervised by world-space 3D models (also called canonical space). Inspired by these observations, we introduce a simple prior-free implicit space transformation network, namely IST-Net, to transform camera-space features to world-space counterparts and build correspondences between them in an implicit manner without relying on 3D priors. Besides, we design camera- and world-space enhancers to enrich the features with pose-sensitive information and geometrical constraints, respectively. Albeit simple, IST-Net achieves state-of-the-art performance based-on prior-free design, with top inference speed on the REAL275 benchmark. Our code and models are available at https://github.com/CVMI-Lab/IST-Net.
Yukang Chen, Xiaoqing Ye, Xiaojuan Qi 0001
ICCV3
2023 A Simple Vision Transformer for Weakly Semi-supervised 3D Object Detection
abstract
Advanced 3D object detection methods usually rely on large-scale, elaborately labeled datasets to achieve good performance. However, labeling the bounding boxes for the 3D objects is difficult and expensive. Although semi-supervised (SS3D) and weakly-supervised 3D object detection (WS3D) methods can effectively reduce the annotation cost, they suffer from two limitations: 1) their performance is far inferior to the fully-supervised counterparts; 2) they are difficult to adapt to different detectors or scenes (e.g, indoor or outdoor). In this paper, we study weakly semi-supervised 3D object detection (WSS3D) with point annotations, where the dataset comprises a small number of fully labeled and massive weakly labeled data with a single point annotated for each 3D object. To fully exploit the point annotations, we employ the plain and non-hierarchical vision transformer to form a point-to-box converter, termed ViT-WSS3D. By modeling global interactions between LiDAR points and corresponding weak labels, our ViT-WSS3D can generate high-quality pseudo-bounding boxes, which are then used to train any 3D detectors without exhaustive tuning. Extensive experiments on indoor and outdoor datasets (SUN RGBD and KITTI) show the effectiveness of our method. In particular, when only using 10% fully labeled and the rest as point labeled data, our ViT-WSS3D can enable most detectors to achieve similar performance with the oracle model using 100% fully labeled data.
Dingyuan Zhang, Dingkang Liang, Zhikang Zou, Xiaoqing Ye, Zhe Liu 0033, Xiao Tan 0001, Xiang Bai
ICCV5
2023 Query-based Temporal Fusion with Explicit Motion for 3D Object Detection
abstract
Effectively utilizing temporal information to improve 3D detection performance is vital for autonomous driving vehicles. Existing methods either conduct temporal fusion based on the dense BEV features or sparse 3D proposal features. However, the former does not pay more attention to foreground objects, leading to more computation costs and sub-optimal performance. The latter implements time-consuming operations to generate sparse 3D proposal features, and the performance is limited by the quality of 3D proposals. In this paper, we propose a simple and effective Query-based Temporal Fusion Network (QTNet). The main idea is to exploit the object queries in previous frames to enhance the representation of current object queries by the proposed Motion-guided Temporal Modeling (MTM) module, which utilizes the spatial position information of object queries along the temporal dimension to construct their relevance between adjacent frames reliably. Experimental results show our proposed QTNet outperforms BEV-based or proposal-based manners on the nuScenes dataset. Besides, the MTM is a plug-and-play module, which can be integrated into some advanced LiDAR-only or multi-modality 3D detectors and even brings new SOTA performance with negligible computation cost and latency on the nuScenes dataset. These experiments powerfully illustrate the superiority and generalization of our method. The code is available at https://github.com/AlmoonYsl/QTNet.
Jinghua Hou, Zhe Liu 0033, Dingkang Liang, Zhikang Zou, Xiaoqing Ye, Xiang Bai
NeurIPS5
2023 Diffusion-Based 3D Object Detection with Random Boxes
Xin Zhou 0013, Jinghua Hou, Dingkang Liang, Zhe Liu 0033, Zhikang Zou, Xiaoqing Ye, Jianwei Cheng, Xiang Bai
PRCV (2)7
2023 Multi-granularity sequential three-way recommendation based on collaborative deep learning
Xiaoqing Ye, Dun Liu, Tianrui Li 0001
Int. J. Approx. Reason.1
2022 Modify Self-Attention via Skeleton Decomposition for Effective Point Cloud Transformer
abstract
Although considerable progress has been achieved regarding the transformers in recent years, the large number of parameters, quadratic computational complexity, and memory cost conditioned on long sequences make the transformers hard to train and implement, especially in edge computing configurations. In this case, a dizzying number of works have sought to make improvements around computational and memory efficiency upon the original transformer architecture. Nevertheless, many of them restrict the context in the attention to seek a trade-off between cost and performance with prior knowledge of orderly stored data. It is imperative to dig deep into an efficient feature extractor for point clouds due to their irregularity and a large number of points. In this paper, we propose a novel skeleton decomposition-based self-attention (SD-SA) which has no sequence length limit and exhibits favorable scalability in long-sequence models. Due to the numerical low-rank nature of self-attention, we approximate it by the skeleton decomposition method while maintaining its effectiveness. At this point, we have shown that the proposed method works for the proposed approach on point cloud classification, segmentation, and detection tasks on the ModelNet40, ShapeNet, and KITTI datasets, respectively. Our approach significantly improves the efficiency of the point cloud transformer and exceeds other efficient transformers on point cloud tasks in terms of the speed at comparable performance.
Jiayi Han, Longbin Zeng, Liang Du 0004, Xiaoqing Ye, Weiyang Ding, Jianfeng Feng
AAAI4
2022 Neural Deformable Voxel Grid for Fast Optimization of Dynamic View Synthesis
Guanying Chen, Yuchao Dai, Xiaoqing Ye, Jiadai Sun, Xiao Tan 0001, Errui Ding
ACCV (1)4
2022 TWIST: Two-Way Inter-label Self-Training for Semi-supervised 3D Instance Segmentation
abstract
We explore the way to alleviate the label-hungry problem in a semi-supervised setting for 3D instance segmentation. To leverage the unlabeled data to boost model performance, we present a novel Two-Way Inter-label Self-Training framework named TWIST. It exploits inherent correlations between semantic understanding and instance information of a scene. Specifically, we consider two kinds of pseudo labels for semantic- and instance-level supervision. Our key design is to provide object-level information for denoising pseudo labels and make use of their correlation for two-way mutual enhancement, thereby iteratively promoting the pseudo-label qualities. TWIST attains leading performance on both ScanNet and S3DIS, compared to recent 3D pre-training approaches, and can cooperate with them to further enhance performance, e.g., +4.4% AP50on 1%-label ScanNet data-efficient benchmark. Code is available at https://github.com/dvlab-research/TWIST.
Ruihang Chu, Xiaoqing Ye, Zhengzhe Liu, Xiao Tan 0001, Xiaojuan Qi 0001, Chi-Wing Fu, Jiaya Jia
CVPR2
2022 Rope3D: The Roadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection Task
abstract
Concurrent perception datasets for autonomous driving are mainly limited to frontal view with sensors mounted on the vehicle. None of them is designed for the overlooked roadside perception tasks. On the other hand, the data captured from roadside cameras have strengths over frontal-view data, which is believed to facilitate a safer and more intelligent autonomous driving system. To accelerate the progress of roadside perception, we present the first high-diversity challenging Roadside Perception 3D dataset- Rope3D from a novel view. The dataset consists of 50k images and over 1.5M 3D objects in various scenes, which are captured under different settings including various cameras with ambiguous mounting positions, camera specifications, viewpoints, and different environmental conditions. We conduct strict 2D-3D joint annotation and comprehensive data analysis, as well as set up a new 3D roadside perception benchmark with metrics and evaluation devkit. Furthermore, we tailor the existing frontal-view monocular 3D object detection approaches and propose to leverage the geometry constraint to solve the inherent ambiguities caused by various sensors, viewpoints. Our dataset is available on https://thudair.baai.ac.cn/rope.
Xiaoqing Ye, Mao Shu, Yifeng Shi, Guangjie Wang, Xiao Tan 0001, Errui Ding
CVPR1
2022 GitNet: Geometric Prior-Based Transformation for Birds-Eye-View Segmentation
Shi Gong, Xiaoqing Ye, Xiao Tan 0001, Jingdong Wang 0001, Errui Ding, Yu Zhou 0025, Xiang Bai
ECCV (1)2
2022 DuARE: Automatic Road Extraction with Aerial Images and Trajectory Data at Baidu Maps
abstract
The task of road extraction has aroused remarkable attention due to its critical role in facilitating urban development and up-to-date map maintenance, which has widespread applications such as navigation and autonomous driving. Existing solutions either rely on a single source of data for road graph extraction or simply fuse the multimodal information in a sub-optimal way. In this paper, we present an automatic road extraction solution named DuARE, which is designed to exploit the multimodal knowledge for underlying road extraction in a fully automatic manner. Specifically, we collect a large-scale real-world dataset for paired aerial image and trajectory data, covering over 33,000 km2 in more than 80 cities. First, road extraction is performed on the abundant spatial-temporal trajectory data adaptively based on the density distribution. Then, a coarse-to-fine road graph learner from aerial images is proposed to take advantage of the local and global context. Finally, our cross-check-based fusion approach keeps the optimal state of each modality while revisiting the original trajectory map with the guidance of aerial predictions to further improve the performance. Extensive experiments conducted on large-scale real-world datasets demonstrate the superiority and effectiveness of DuARE. In addition, DuARE has been deployed in production at Baidu Maps since June 2021 and keeps updating the road network by 100,000 km per month. This confirms that DuARE is a practical and industrial-grade solution for large-scale cost-effective road extraction from multimodal data.
Jianzhong Yang, Xiaoqing Ye, Yanlei Gu, Deguo Xia, Jizhou Huang
KDD2
2022 Repainting and Imitating Learning for Lane Detection
abstract
Current lane detection methods are struggling with the invisibility lane issue caused by heavy shadows, severe road mark degradation, and serious vehicle occlusion. As a result, discriminative lane features can be barely learned by the network despite elaborate designs due to the inherent invisibility of lanes in the wild. In this paper, we target at finding an enhanced feature space where the lane features are distinctive while maintaining a similar distribution of lanes in the wild. To achieve this, we propose a novel Repainting and Imitating Learning (RIL) framework containing a pair of teacher and student without any extra data or extra laborious labeling. Specifically, in the repainting step, an enhanced ideal virtual lane dataset is built in which only the lane regions are repainted while non-lane regions are kept unchanged, maintaining the similar distribution of lanes in the wild. The teacher model learns enhanced discriminative representation based on the virtual data and serves as the guidance for a student model to imitate. In the imitating learning step, through the scale-fusing distillation module, the student network is encouraged to generate features that mimic the teacher model both on the same scale and cross scales. Furthermore, the coupled adversarial module builds the bridge to connect not only teacher and student models but also virtual and real data, adjusting the imitating learning process dynamically. Note that our method introduces no extra time cost during inference and can be plug-and-play in various cutting-edge lane detection networks. Experimental results prove the effectiveness of the RIL framework both on CULane and TuSimple for four modern lane detection methods. The code and model will be available soon.
Minyue Jiang, Xiaoqing Ye, Liang Du 0004, Zhikang Zou, Wei Zhang 0197, Xiao Tan 0001, Errui Ding
ACM Multimedia3
2022 Paint and Distill: Boosting 3D Object Detection with Semantic Passing Network
abstract
3D object detection task from lidar or camera sensors is essential for autonomous driving. Pioneer attempts at multi-modality fusion complement the sparse lidar point clouds with rich semantic texture information from images at the cost of extra network designs and overhead. In this work, we propose a novel semantic passing framework, named SPNet, to boost the performance of existing lidar-based 3D detection models with the guidance of rich context painting, with no extra computation cost during inference. Our key design is to first exploit the potential instructive semantic knowledge within the ground-truth labels by training a semantic-painted teacher model and then guide the pure-lidar network to learn the semantic-painted representation via knowledge passing modules at different granularities: class-wise passing, pixel-wise passing and instance-wise passing. Experimental results show that the proposed SPNet can seamlessly cooperate with most existing 3D detection frameworks with 1$\sim$5% AP gain and even achieve new state-of-the-art 3D detection performance on the KITTI test benchmark. Code is available at: https://github.com/jb892/SPNet.
Bo Ju, Zhikang Zou, Xiaoqing Ye, Minyue Jiang, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001
ACM Multimedia3
2022 Spatial Pruned Sparse Convolution for Efficient 3D Object Detection
abstract
3D scenes are dominated by a large number of background points, which is redundant for the detection task that mainly needs to focus on foreground objects. In this paper, we analyze major components of existing sparse 3D CNNs and find that 3D CNNs ignores the redundancy of data and further amplifies it in the down-sampling process, which brings a huge amount of extra and unnecessary computational overhead. Inspired by this, we propose a new convolution operator named spatial pruned sparse convolution (SPS-Conv), which includes two variants, spatial pruned submanifold sparse convolution (SPSS-Conv) and spatial pruned regular sparse convolution (SPRS-Conv), both of which are based on the idea of dynamically determine crucial areas for performing computations to reduce redundancy. We empirically find that magnitude of features can serve as an important cues to determine crucial areas which get rid of the heavy computations of learning-based methods. The proposed modules can easily be incorporated into existing sparse 3D CNNs without extra architectural modifications. Extensive experiments on the KITTI and nuScenes datasets demonstrate that our method can achieve more than 50% reduction in GFLOPs without compromising the performance.
Yukang Chen, Xiaoqing Ye, Zhuotao Tian, Xiao Tan 0001, Xiaojuan Qi 0001
NeurIPS3
2022 The devil is in the face: Exploiting harmonious representations for facial expression recognition
Jiayi Han, Liang Du 0004, Xiaoqing Ye, Li Zhang 0040, Jianfeng Feng
Neurocomputing3
2022 A cost-sensitive temporal-spatial three-way recommendation with multi-granularity decision
Xiaoqing Ye, Dun Liu
Inf. Sci.1
2022 AGO-Net: Association-Guided 3D Point Cloud Object Detection Network
abstract
The human brain can effortlessly recognize and localize objects, whereas current 3D object detection methods based on LiDAR point clouds still report inferior performance for detecting occluded and distant objects: The point cloud appearance varies greatly due to occlusion, and has inherent variance in point densities along the distance to sensors. Therefore, designing feature representations robust to such point clouds is critical. Inspired by human associative recognition, we propose a novel 3D detection framework that associates intact features for objects via domain adaptation. We bridge the gap between the perceptual domain, where features are derived from real scenes with sub-optimal representations, and the conceptual domain, where features are extracted from augmented scenes that consist of non-occlusion objects with rich detailed information. A feasible method is investigated to construct conceptual scenes without external datasets. We further introduce an attention-based re-weighting module that adaptively strengthens the feature adaptation of more informative regions. The network's feature enhancement ability is exploited without introducing extra cost during inference, which is plug-and-play in various 3D detection frameworks. We achieve new state-of-the-art performance on the KITTI 3D detection benchmark in both accuracy and speed. Experiments on nuScenes and Waymo datasets also validate the versatility of our method.
Liang Du 0004, Xiaoqing Ye, Xiao Tan 0001, Edward Johns, Errui Ding, Xiangyang Xue 0001, Jianfeng Feng
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Depth-Conditioned Dynamic Message Propagation for Monocular 3D Object Detection
abstract
The objective of this paper is to learn context- and depth- aware feature representation to solve the problem of monocular 3D object detection. We make following contributions: (i) rather than appealing to the complicated pseudo-LiDAR based approach, we propose a depth-conditioned dynamic message propagation (DDMP) network to effectively integrate the multi-scale depth information with the image context; (ii) this is achieved by first adaptively sampling context-aware nodes in the image context and then dynamically predicting hybrid depth-dependent filter weights and affinity matrices for propagating information; (Hi) by augmenting a center-aware depth encoding (CDE) task, our method successfully alleviates the inaccurate depth prior; (iv) we thoroughly demonstrate the effectiveness of our proposed approach and show state-of-the-art results among the monocular-based approaches on the KITTI benchmark dataset. Particularly, we rank 1stin the highly competitive KITTI monocular 3D object detection track on the submission day (November 16th, 2020). Code and models are released at https: //github.com/fudan-zvg/DDMP
Li Wang 0033, Liang Du 0004, Xiaoqing Ye, Yanwei Fu 0001, Guodong Guo, Xiangyang Xue 0001, Jianfeng Feng, Li Zhang 0040
CVPR3
2021 Revealing the Reciprocal Relations between Self-Supervised Stereo and Monocular Depth Estimation
abstract
Current self-supervised depth estimation algorithms mainly focus on either stereo or monocular only, neglecting the reciprocal relations between them. In this paper, we propose a simple yet effective framework to improve both stereo and monocular depth estimation by leveraging the underlying complementary knowledge of the two tasks. Our approach consists of three stages. In the first stage, the proposed stereo matching network termed StereoNet is trained on image pairs in a self-supervised manner. Second, we introduce an occlusion-aware distillation (OA Distillation) module, which leverages the predicted depths from StereoNet in non-occluded regions to train our monocular depth estimation network named SingleNet. At last, we design an occlusion-aware fusion module (OA Fusion), which generates more reliable depths by fusing estimated depths from StereoNet and SingleNet given the occlusion map. Furthermore, we also take the fused depths as pseudo labels to supervise StereoNet in turn, which brings StereoNet’s performance to a new height. Extensive experiments on KITTI dataset demonstrate the effectiveness of our proposed framework. We achieve new SOTA performance on both stereo and monocular depth estimation tasks.
Zhi Chen 0026, Xiaoqing Ye, Wei Yang 0011, Zhenbo Xu, Xiao Tan 0001, Zhikang Zou, Errui Ding, Xinming Zhang 0001, Liusheng Huang
ICCV2
2021 The Devil is in the Task: Exploiting Reciprocal Appearance-Localization Features for Monocular 3D Object Detection
abstract
Low-cost monocular 3D object detection plays a fundamental role in autonomous driving, whereas its accuracy is still far from satisfactory. In this paper, we dig into the 3D object detection task and reformulate it as the sub-tasks of object localization and appearance perception, which benefits to a deep excavation of reciprocal information underlying the entire task. We introduce a Dynamic Feature Reflecting Network, named DFR-Net, which contains two novel standalone modules: (i) the Appearance-Localization Feature Reflecting module (ALFR) that first separates task-specific features and then self-mutually reflects the reciprocal features; (ii) the Dynamic Intra-Trading module (DIT) that adaptively realigns the training processes of various sub-tasks via a self-learning manner. Extensive experiments on the challenging KITTI dataset demonstrate the effectiveness and generalization of DFR-Net. We rank 1stamong all the monocular 3D object detectors in the KITTI test set (till March 16th, 2021). The proposed method is also easy to be plug-and-play in many cutting-edge 3D detection frameworks at negligible cost to boost performance. The code will be made publicly available.
Zhikang Zou, Xiaoqing Ye, Liang Du 0004, Xianhui Cheng, Xiao Tan 0001, Li Zhang 0040, Jianfeng Feng, Xiangyang Xue 0001, Errui Ding
ICCV2
2021 DANet: Dimension Apart Network for Radar Object Detection
abstract
In this paper, we propose a dimension apart network (DANet) for radar object detection task. A Dimension Apart Module (DAM) is first designed to be lightweight and capable of extracting temporal-spatial information from the RAMap sequences. To fully utilize the hierarchical features from the RAMaps, we propose a multi-scale U-Net style network architecture termed DANet. Extensive experiments demonstrate that our proposed DANet achieves superior performance on the radar detection task at much less computational cost, compared to previous pioneer works. In addition to the proposed novel network, we also utilize a vast amount of data augmentation techniques. To further improve the robustness of our model, we ensemble the predicted results from a bunch of lightweight DANet variants. Finally, we achieve 82.2% on average precision and 90% on average recall of object detection performance and rank at 1st place in the ROD2021 radar detection challenge. Our code is available at: \urlhttps://github.com/jb892/ROD2021_Radar_Detection_Challenge_Baidu.
Bo Ju, Wei Yang 0011, Jinrang Jia, Xiaoqing Ye, Xiao Tan 0001, Yifeng Shi, Errui Ding
ICMR4
2021 AggNet for Self-supervised Monocular Depth Estimation: Go An Aggressive Step Furthe
abstract
Without appealing to exhaustive labeled data, self-supervised monocular depth estimation (MDE) plays a fundamental role in computer vision. Previous methods usually adopt a one-stage MDE network, which is insufficient to achieve high performance. In this paper, we dig deep into this task to propose an aggressive framework termed AggNet. The framework is based on a training-only progressive two-stage module to perform pseudo counter-surveillance as well as a simple yet effective dual-warp loss function between image pairs. In particular, we first propose a residual module, which follows the MDE network to learn a refined depth. The residual module takes both the initial depth generated from MDE and the initial color image as input to generate refined depth with residual depth learning. Then, the refined depth is leveraged to supervise the initial depth simultaneously during the training period. For inference, only the MDE network is retained to regress depth from a single image, which gains better performance without introducing extra computation. In addition to self-distillation loss, a simple yet effective dual-warp consistency loss is introduced to encourage the MDE network to keep depth consistency between stereo image pairs. Extensive experiments show that our AggNet achieves state-of-the-art performance on the KITTI and Make3D datasets.
Zhi Chen 0026, Xiaoqing Ye, Liang Du 0004, Wei Yang 0011, Liusheng Huang, Xiao Tan 0001, Zhenbo Shi, Fumin Shen, Errui Ding
ACM Multimedia2
2021 Lifting the Veil of Frequency in Joint Segmentation and Depth Estimation
abstract
Joint learning of scene parsing and depth estimation remains a challenging task due to the rivalry between the two tasks. In this paper, we revisit the mutual enhancement for joint semantic segmentation and depth estimation. Inspired by the observation that the competition and cooperation could be reflected in the feature frequency components of different tasks, we propose a Frequency Aware Feature Enhancement (FAFE) network that can effectively enhance the reciprocal relationship whereas avoiding the competition. In FAFE, a frequency disentanglement module is proposed to fetch the favorable frequency component sets for each task and resolve the discordance between the two tasks. For task cooperation, we introduce a re-calibration unit to aggregate features of the two tasks, so as to complement task information with each other. Accordingly, the learning of each task can be boosted by the complementary task appropriately. Besides, a novel local-aware consistency loss function is proposed to impose on the predicted segmentation and depth so as to strengthen the cooperation. With the FAFE network and new local-aware consistency loss encapsulated into the multi-task learning network, the proposed approach achieves superior performance over previous state-of-the-art methods. Extensive experiments and ablation studies on multi-task datasets demonstrate the effectiveness of our proposed approach.
Tianhao Fu, Xiaoqing Ye, Xiao Tan 0001, Fumin Shen, Errui Ding
ACM Multimedia3
2021 Towards Adversarial Patch Analysis and Certified Defense against Crowd Counting
abstract
Crowd counting has drawn much attention due to its importance in safety-critical surveillance systems. Especially, deep neural network (DNN) methods have significantly reduced estimation errors for crowd counting missions. Recent studies have demonstrated that DNNs are vulnerable to adversarial attacks, i.e., normal images with human-imperceptible perturbations could mislead DNNs to make false predictions. In this work, we propose a robust attack strategy called Adversarial Patch Attack with Momentum (APAM) to systematically evaluate the robustness of crowd counting models, where the attacker's goal is to create an adversarial perturbation that severely degrades their performances, thus leading to public safety accidents (e.g., stampede accidents). Especially, the proposed attack leverages the extreme-density background information of input images to generate robust adversarial patches via a series of transformations (e.g., interpolation, rotation, etc.). We observe that by perturbing less than 6% of image pixels, our attacks severely degrade the performance of crowd counting systems, both digitally and physically. To better enhance the adversarial robustness of crowd counting models, we propose the first regression model-based Randomized Ablation (RA), which is more sufficient than Adversarial Training (ADT) (Mean Absolute Error of RA is 5 lower than ADT on clean samples and 30 lower than ADT on adversarial examples). Extensive experiments on five crowd counting models demonstrate the effectiveness and generality of the proposed method.
Zhikang Zou, Pan Zhou 0001, Xiaoqing Ye, Binghui Wang, Ang Li 0005
ACM Multimedia4
2021 Coarse to Fine: Domain Adaptive Crowd Counting via Adversarial Scoring Network
abstract
Recent deep networks have convincingly demonstrated high capability in crowd counting, which is a critical task attracting widespread attention due to its various industrial applications. Despite such progress, trained data-dependent models usually can not generalize well to unseen scenarios because of the inherent domain shift. To facilitate this issue, this paper proposes a novel adversarial scoring network (ASNet) to gradually bridge the gap across domains from coarse to fine granularity. In specific, at the coarse-grained stage, we design a dual-discriminator strategy to adapt source domain to be close to the targets from the perspectives of both global and local feature space via adversarial learning. The distributions between two domains can thus be aligned roughly. At the fine-grained stage, we explore the transferability of source characteristics by scoring how similar the source samples are to target ones from multiple levels based on generative probability derived from coarse stage. Guided by these hierarchical scores, the transferable source features are properly selected to enhance the knowledge transfer during the adaptation process. With the coarse-to-fine design, the generalization bottleneck induced from the domain discrepancy can be effectively alleviated. Three sets of migration experiments show that the proposed methods achieve state-of-the-art counting performance compared with major unsupervised methods.
Zhikang Zou, Xiaoye Qu, Pan Zhou 0001, Shuangjie Xu, Xiaoqing Ye, Jin Ye 0006
ACM Multimedia5
2021 An interpretable sequential three-way recommendation based on collaborative topic regression
Xiaoqing Ye, Dun Liu
Expert Syst. Appl.1
2020 ZoomNet: Part-Aware Adaptive Zooming Neural Network for 3D Object Detection
abstract
3D object detection is an essential task in autonomous driving and robotics. Though great progress has been made, challenges remain in estimating 3D pose for distant and occluded objects. In this paper, we present a novel framework named ZoomNet for stereo imagery-based 3D detection. The pipeline of ZoomNet begins with an ordinary 2D object detection model which is used to obtain pairs of left-right bounding boxes. To further exploit the abundant texture cues in rgb images for more accurate disparity estimation, we introduce a conceptually straight-forward module – adaptive zooming, which simultaneously resizes 2D instance bounding boxes to a unified resolution and adjusts the camera intrinsic parameters accordingly. In this way, we are able to estimate higher-quality disparity maps from the resized box images then construct dense point clouds for both nearby and distant objects. Moreover, we introduce to learn part locations as complementary features to improve the resistance against occlusion and put forward the 3D fitting score to better estimate the 3D detection quality. Extensive experiments on the popular KITTI 3D detection dataset indicate ZoomNet surpasses all previous state-of-the-art methods by large margins (improved by 9.4% on APbv (IoU=0.7) over pseudo-LiDAR). Ablation study also demonstrates that our adaptive zooming strategy brings an improvement of over 10% on AP3d (IoU=0.7). In addition, since the official KITTI benchmark lacks fine-grained annotations like pixel-wise part locations, we also present our KFG dataset by augmenting KITTI with detailed instance-wise annotations including pixel-wise part location, pixel-wise disparity, etc.. Both the KFG dataset and our codes will be publicly available at https://github.com/detectRecog/ZoomNet.
Zhenbo Xu, Wei Zhang 0197, Xiaoqing Ye, Xiao Tan 0001, Wei Yang 0011, Shilei Wen, Errui Ding, Ajin Meng, Liusheng Huang
AAAI3
2020 Associate-3Ddet: Perceptual-to-Conceptual Association for 3D Point Cloud Object Detection
abstract
Object detection from 3D point clouds remains a challenging task, though recent studies pushed the envelope with the deep learning techniques. Owing to the severe spatial occlusion and inherent variance of point density with the distance to sensors, appearance of a same object varies a lot in point cloud data. Designing robust feature representation against such appearance changes is hence the key issue in a 3D object detection method. In this paper, we innovatively propose a domain adaptation like approach to enhance the robustness of the feature representation. More specifically, we bridge the gap between the perceptual domain where the feature comes from a real scene and the conceptual domain where the feature is extracted from an augmented scene consisting of non-occlusion point cloud rich of detailed information. This domain adaptation approach mimics the functionality of the human brain when proceeding object perception. Extensive experiments demonstrate that our simple yet effective approach fundamentally boosts the performance of 3D point cloud object detection and achieves the state-of-the-art results.
Liang Du 0004, Xiaoqing Ye, Xiao Tan 0001, Jianfeng Feng, Zhenbo Xu, Errui Ding, Shilei Wen
CVPR2
2020 Monocular 3D Object Detection via Feature Domain Adaptation
Xiaoqing Ye, Liang Du 0004, Yifeng Shi, Xiao Tan 0001, Jianfeng Feng, Errui Ding, Shilei Wen
ECCV (9)1
2020 RegionNet: Region-feature-enhanced 3D Scene Understanding Network with Dual Spatial-aware Discriminative Loss
abstract
Neural networks have recently achieved impressive success in semantic and instance segmentation on 2D images. However, their capabilities have not been fully explored to address semantic instance segmentation on unstructured 3D point cloud data. Digging into the regional feature representation to boost point cloud comprehension, we propose a region-feature-enhanced structure consisting of adaptive regional feature complementary (ARFC) module and affinity-based regional relational reasoning (AR3) module. The ARFC module aims to complement low-level features of sparse regions adaptively. The AR3module emphasizes on mining the potential reasoning relationships between high-level features based on affinity. Both the ARFC and AR3modules are plug-and-play. Besides, a novel dual spatial-aware discriminative loss is proposed to improve the discrimination of instance embedding. Our proposal-free point cloud instance segmentation network (RegionNet) equipped with the region-feature-enhanced structure and dual spatial-aware discriminative loss achieves state-of-the-art performance on S3DIS dataset and ScanNet-v2 dataset.
Dongchen Zhu, Xiaoqing Ye, Wenjun Shi, Minghong Chen, Jiamao Li
IROS3
2020 Object Reidentification via Joint Quadruple Decorrelation Directional Deep Networks in Smart Transportation
abstract
Object reidentification with the goal of matching pedestrian or vehicle images captured from different camera viewpoints is of considerable significance to public security. Quadruple directional deep learning features (QD-DLFs) can comprehensively describe object images. However, the correlation among QD-DLFs is an unavoidable problem, since QD-DLFs are learned with quadruple independent directional deep networks (QIDDNs) driven with the same training data, and each network holds the same basic deep feature learning architecture (BDFLA). The correlation among QD-DLFs is harmful to the complementarity of QD-DLFs, restricting the object reidentification performance. For that, we propose joint quadruple decorrelation directional deep networks (JQD3Ns) to reduce the correlation among the learned QD-DLFs. In order to jointly train JQD3Ns, besides the softmax loss functions, a parameter correlation cost function is proposed to indirectly reduce the correlation among QD-DLFs by enlarging the dissimilarity among the parameters of JQD3Ns. Extensive experiments on three publicly available large-scale data sets demonstrate that the proposed JQD3Ns approach is superior to multiple state-of-the-art object reidentification methods.
Jianqing Zhu, Jingchang Huang, Huanqiang Zeng, Xiaoqing Ye, Baoqing Li, Zhen Lei 0001, Lixin Zheng
IEEE Internet Things J.4
2020 A matrix factorization based dynamic granularity recommendation with three-way decisions
Dun Liu, Xiaoqing Ye
Knowl. Based Syst.2
2019 SSF-DAN: Separated Semantic Feature Based Domain Adaptation Network for Semantic Segmentation
abstract
Despite the great success achieved by supervised fully convolutional models in semantic segmentation, training the models requires a large amount of labor-intensive work to generate pixel-level annotations. Recent works exploit synthetic data to train the model for semantic segmentation, but the domain adaptation between real and synthetic images remains a challenging problem. In this work, we propose a Separated Semantic Feature based domain adaptation network, named SSF-DAN, for semantic segmentation. First, a Semantic-wise Separable Discriminator (SS-D) is designed to independently adapt semantic features across the target and source domains, which addresses the inconsistent adaptation issue in the class-wise adversarial learning. In SS-D, a progressive confidence strategy is included to achieve a more reliable separation. Then, an efficient Class-wise Adversarial loss Reweighting module (CA-R) is introduced to balance the class-wise adversarial learning process, which leads the generator to focus more on poorly adapted classes. The presented framework demonstrates robust performance, superior to state-of-the-art methods on benchmark datasets.
Liang Du 0004, Jingang Tan, Hongye Yang, Jianfeng Feng, Xiangyang Xue 0001, Qibao Zheng, Xiaoqing Ye
ICCV7
2018 3D Recurrent Neural Networks with Context Fusion for Point Cloud Semantic Segmentation
Xiaoqing Ye, Jiamao Li, Hexiao Huang, Liang Du 0004
ECCV (7)1
2017 Order-Based Disparity Refinement Including Occlusion Handling for Stereo Matching
abstract
Accurate stereo matching is still challenging in case of weakly textured areas, discontinuities, and occlusions. Besides, occlusion recovery is often regarded as a subordinate problem and simply handled. To obtain dense high-accuracy depth maps, this letter proposes an efficient multistep disparity refinement framework with occlusion handling. The framework is implemented by classifying the outliers into leftmost occlusions, nonborder occlusions, as well as mismatches, and employing different strategies to recover them. To recover occlusions, a filling order is specially introduced to avoid error propagation and surface decision based on local image content is performed when more than one background surface exists. The evaluations on Middlebury datasets and comparisons with other refinement algorithms show the superiority and robustness of our method.
Xiaoqing Ye, Yuzhang Gu, Jiamao Li
IEEE Signal Process. Lett.1
2016 Application of Wuhan Ionospheric Oblique Backscattering Sounding System (WIOBSS) for Sea-State Detection
abstract
To combine the advantages of long detection range and small antenna array, the newly designed Wuhan Ionospheric Oblique Backscattering Sounding System (WIOBSS) transmits radio waves by ionospheric reflection and receives oceanic backscattered ground waves. To combine the functions of operating frequency selection and sea-state detection, WIOBSS can transmit the interpulse coding wave and frequency-modulated continuous wave waveforms in one hardware platform. The hardware and software of the radio system are introduced. A sky-wave over-the-horizon sea-state detection experiment was carried out in Chongyang and Longhai, China, on January 28, 2015. The observations of the oceanic surface echoes 130-300 km from the coastline are also presented.
Wanlin Gong, Xiaoqing Ye, Lingyun Pan
IEEE Geosci. Remote. Sens. Lett.3