Yun Zhu 0011

dblp:00/6306-11 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2026
0009-0000-2121-548XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Discriminative region learning for point cloud-based place recognition
Le Hui, Yun Zhu 0011, Jianjun Qian, Yigong Zhang, Jin Xie 0001
Neural Networks3
2025 NaviFormer: A Spatio-Temporal Context-Aware Transformer for Object Navigation
abstract
Learning discriminative state representations of agents, encompassing the spatial layout and temporal pose trajectory, is essential for effective navigation decisions. However, existing approaches often rely on simplistic plain networks for navigation information fusion, overlooking the complex long-range dependencies across spatio-temporal cues, which leads to suboptimal state perception and potential decision failures. In this paper, we introduce NaviFormer, an effective encoder-decoder navigation transformer, to aggregate discriminative spatio-temporal context information for object navigation. Our navigation encoder not only encodes spatial layouts and temporal agent poses but also innovatively constructs and encodes a passable frontier map, enriching the original state encoding with cues of potential exploration regions. Furthermore, our navigation decoder employs spatio-temporal self-attention and cross-attention mechanisms to model the dependencies among spatial layout encoding, temporal pose encoding, and passable frontier encoding, thereby facilitating comprehensive contextual state feature aggregation. Finally, we leverage these learned spatio-temporal contextual state representations for PPO-based navigation decisions. Extensive experiments on the Gibson, Habitat-Matterport3D (HM3D) and Matterport3D (MP3D) datasets demonstrate the superiority of our approach.
Wei Xie 0019, Haobo Jiang, Yun Zhu 0011, Jianjun Qian, Jin Xie 0001
AAAI3
2025 WeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba Diffusion
abstract
3D scene perception demands a large amount of adverse-weather LiDAR data, yet the cost of LiDAR data collection presents a significant scaling-up challenge. To this end, a series of LiDAR simulators have been proposed. Yet, they can only simulate a single adverse weather with a single physical model, and the fidelity of the generated data is quite limited. This paper presents WeatherGen, the first unified diverse-weather LiDAR data diffusion generation framework, significantly improving fidelity. Specifically, we first design a map-based data producer, which can provide a vast amount of high-quality diverse-weather data for training purposes. Then, we utilize the diffusion-denoising paradigm to construct a diffusion model. Among them, we propose a spider mamba generator to restore the disturbed diverse weather data gradually. The spider mamba models the feature interactions by scanning the Li-Dar beam circle or central ray, excellently maintaining the physical structure of the LiDAR data. Subsequently, following the generator to transfer real-world knowledge, we design a latent feature aligner. Afterward, we devise a contrastive learning-based controller, which equips weather control signals with compact semantic knowledge through language supervision, guiding the diffusion model to generate more discriminative data. Extensive evaluations demonstrate the high generation quality of WeatherGen. Through WeatherGen, we construct the mini-weather dataset, promoting the performance of the downstream task under adverse weather conditions. Code is available: https://github.com/wuyang98/weathergen
Yun Zhu 0011, Kaihua Zhang 0001, Jianjun Qian, Jin Xie 0001, Jian Yang 0003
CVPR2
2025 Learning Class Prototypes for Unified Sparse-Supervised 3D Object Detection
abstract
Both indoor and outdoor scene perceptions are essential for embodied intelligence. However, current sparse supervised 3D object detection methods focus solely on outdoor scenes without considering indoor settings. To this end, we propose a unified sparse supervised 3D object detection method for both indoor and outdoor scenes through learning class prototypes to effectively utilize unlabeled objects. Specifically, we first propose a prototype-based object mining module that converts the unlabeled object mining into a matching problem between class prototypes and unlabeled features. By using optimal transport matching results, we assign prototype labels to high-confidence features, thereby achieving the mining of unlabeled objects. We then present a multi-label cooperative refinement module to effectively recover missed detections through pseudo label quality control and prototype label cooperation. Experiments show that our method achieves state-of-the-art performance under the one object per scene sparse supervised setting across indoor and outdoor datasets. With only one labeled object per scene, our method achieves about 78%, 90%, and 96% performance compared to the fully supervised detector on ScanNet V2, SUN RGB-D, and KITTI, respectively, highlighting the scalability of our method. Code is available at https://github.com/zyrant/CPDet3D.
Yun Zhu 0011, Le Hui, Jianjun Qian, Jin Xie 0001, Jian Yang 0003
CVPR1
2024 SPGroup3D: Superpoint Grouping Network for Indoor 3D Object Detection
abstract
Current 3D object detection methods for indoor scenes mainly follow the voting-and-grouping strategy to generate proposals. However, most methods utilize instance-agnostic groupings, such as ball query, leading to inconsistent semantic information and inaccurate regression of the proposals. To this end, we propose a novel superpoint grouping network for indoor anchor-free one-stage 3D object detection. Specifically, we first adopt an unsupervised manner to partition raw point clouds into superpoints, areas with semantic consistency and spatial similarity. Then, we design a geometry-aware voting module that adapts to the centerness in anchor-free detection by constraining the spatial relationship between superpoints and object centers. Next, we present a superpoint-based grouping module to explore the consistent representation within proposals. This module includes a superpoint attention layer to learn feature interaction between neighboring superpoints, and a superpoint-voxel fusion layer to propagate the superpoint-level information to the voxel level. Finally, we employ effective multiple matching to capitalize on the dynamic receptive fields of proposals based on superpoints during the training. Experimental results demonstrate our method achieves state-of-the-art performance on ScanNet V2, SUN RGB-D, and S3DIS datasets in the indoor one-stage 3D object detection. Source code is available at https://github.com/zyrant/SPGroup3D.
Yun Zhu 0011, Le Hui, Yaqi Shen, Jin Xie 0001
AAAI1
2024 Dense Voxel Representation Network for Implicit Scene Completion
abstract
Implicit scene completion aims to learn an implicit representation of dense point clouds from incomplete ones. Since point clouds are disordered and irregular, some implicit scene completion methods learn representations from voxelized point clouds with sparse convolution. Despite achieving promising results, they lack deep exploration of feature learning on empty voxels, which is beneficial for implicit scene completion task. To address this, we propose a dense voxel representation network for implicit scene completion. First, we design a Bird’s-Eye View (BEV) assisted enhancement module to enhance non-empty voxel features by incorporating the information contained in the learned dense BEV features into them through deformable cross-attention. Second, we construct a feature adaptive completion module to adaptively complete voxel features using deformable self-attention, realizing the transfer of the information from non-empty voxels to empty voxels. Extensive experiments on SemanticKITTI and SemanticPOSS datasets demonstrate our method achieves state-of-the-art performance.
Fan Dai, Yun Zhu 0011, Yaqi Shen, Jin Xie 0001, Jianjun Qian
ICME2
2023 LSNet: Lightweight Spatial Boosting Network for Detecting Salient Objects in RGB-Thermal Images
abstract
Most recent methods for RGB (red-green-blue)-thermal salient object detection (SOD) involve several floating-point operations and have numerous parameters, resulting in slow inference, especially on common processors, and impeding their deployment on mobile devices for practical applications. To address these problems, we propose a lightweight spatial boosting network (LSNet) for efficient RGB-thermal SOD with a lightweight MobileNetV2 backbone to replace a conventional backbone (e.g., VGG, ResNet). To improve feature extraction using a lightweight backbone, we propose a boundary boosting algorithm that optimizes the predicted saliency maps and reduces information collapse in low-dimensional features. The algorithm generates boundary maps based on predicted saliency maps without incurring additional calculations or complexity. As multimodality processing is essential for high-performance SOD, we adopt attentive feature distillation and selection and propose semantic and geometric transfer learning to enhance the backbone without increasing the complexity during testing. Experimental results demonstrate that the proposed LSNet achieves state-of-the-art performance compared with 14 RGB-thermal SOD methods on three datasets while improving the numbers of floating-point operations (1.025G) and parameters (5.39M), model size (22.1 MB), and inference speed (9.95 fps for PyTorch, batch size of 1, and Intel i5-7500 processor; 93.53 fps for PyTorch, batch size of 1, and NVIDIA TITAN V graphics processor; 936.68 fps for PyTorch, batch size of 20, and graphics processor; 538.01 fps for TensorRT and batch size of 1; and 903.01 fps for TensorRT/FP16 and batch size of 1). The code and results can be found from the link of https://github.com/zyrant/LSNet.
Wujie Zhou, Yun Zhu 0011, Jingsheng Lei, Rongwang Yang, Lu Yu 0003
IEEE Trans. Image Process.2
2022 CCAFNet: Crossflow and Cross-Scale Adaptive Fusion Network for Detecting Salient Objects in RGB-D Images
abstract
Owing to the widespread adoption of depth sensors, salient object detection (SOD) supported by depth maps for reliable complementary information is being increasingly investigated. Existing SOD models mainly exploit the relation between an RGB image and its corresponding depth information across three fusion domains: input RGB-D images, extracted feature maps, and output salient object. However, these models do not leverage the crossflows between high- and low-level information well. Moreover, the decoder in these models uses conventional convolution that involves several calculations. To further improve RGB-D SOD, we propose a crossflow and cross-scale adaptive fusion network (CCAFNet) to detect salient objects in RGB-D images. First, a channel fusion module allows for effective fusing depth and high-level RGB features. This module extracts accurate semantic information features from high-level RGB features. Meanwhile, a spatial fusion module combines low-level RGB and depth features with accurate boundaries and subsequently extracts detailed spatial information from low-level depth features. Finally, a purification loss is proposed to precisely learn the boundaries of salient objects and obtain additional details of the objects. The results of comprehensive experiments on seven common RGB-D SOD datasets indicate that the performance of the proposed CCAFNet is comparable to those of state-of-the-art RGB-D SOD models.
Wujie Zhou, Yun Zhu 0011, Jingsheng Lei, Jian Wan 0001, Lu Yu 0003
IEEE Trans. Multim.2
2021 Parallax-Estimation-Enhanced Network With Interweave Consistency Feature Fusion for Binocular Salient Object Detection
abstract
Salient object detection (SOD) has received extensive attention in recent years, and many models have been developed. However, most SOD models only consider monocular images and not binocular images, which resemble the human vision and can better reflect human perception for distinguishing salient objects. To leverage the information in binocular images, we propose herein a first-of-its-kind parallax-estimation-enhanced network (PEENet) for binocular SOD. More specifically, we use a weighted binocular fusion module and a parallax correlation fusion module to explore the complementary and different information in binocular images. In addition, a parallax enhancing module and interweave consistency fusion use complementary saliency information and parallax information to enhance saliency and parallax representations. Finally, a transformation module avoids global and local information loss during decoding. Experiments were performed to validate the effectiveness and robustness of the proposed PEENet, which outperforms 10-RGB/RGB-D SOD methods on two binocular SOD datasets.
Yun Zhu 0011, Wujie Zhou, Lu Yu 0003
IEEE Signal Process. Lett.1