You Shen

dblp:341/5573 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-7526-1980ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 S²Teacher: Step-by-step Teacher for Sparsely Annotated Oriented Object Detection
abstract
Although fully-supervised oriented object detection has made significant progress in remote sensing image understanding, it comes at the cost of labor-intensive annotation. Recent studies have explored weakly and semi-supervised learning to alleviate this burden. However, these methods overlook the difficulties posed by dense annotations in complex remote sensing scenes. In this paper, we introduce a novel setting called sparsely annotated oriented object detection (SAOOD), which only labels partial instances, and propose a solution to address its challenges. Specifically, we focus on two key issues in the setting: (1) sparse labeling leading to overfitting on limited foreground representations, and (2) unlabeled objects (false negatives) confusing feature learning. To this end, we propose the S2Teacher, a novel angle-consistency guided method that progressively mines pseudo-labels for unlabeled objects from easy to hard, enhancing foreground representations. Additionally, it reweights the loss of unlabeled objects to mitigate their impact during training. Extensive experiments demonstrate that S2Teacher not only significantly improves detector performance across different sparse annotation levels but also achieves near-fully-supervised performance on the DOTA dataset with only 10% annotation instances, effectively balancing accuracy and labeling cost.
Jianghang Lin, You Shen, Shengchuan Zhang, Liujuan Cao
AAAI4
2026 DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion
abstract
Reconstructing 3D objects from a single image is a long-standing challenge, particularly under real-world occlusions. While recent diffusion-based view synthesis models can generate consistent novel views from a single RGB image, they generally assume fully visible inputs and struggle when parts of the object are occluded, leading to inconsistent views and degraded 3D reconstruction quality. To address this limitation, we propose DeOcc-1-to-3, an end-to-end framework for occlusion-aware multi-view generation. Our method directly synthesizes six structurally consistent novel views from a single partially occluded image, enabling downstream 3D reconstruction without requiring prior inpainting or manual annotations. We design a self-supervised training pipeline that leverages occluded–unoccluded image pairs and pseudo-ground-truth views to guide structure-aware completion and view consistency. Without modifying the original architecture, we fully fine-tune the diffusion model to jointly learn completion and multi-view generation. Additionally, we introduce the first benchmark for occlusion-aware reconstruction, covering diverse occlusion levels, object categories, and mask patterns, providing a standardized evaluation protocol.
Yansong Qu, Shaohui Dai, Yuze Wang 0006, You Shen, Shengchuan Zhang, Liujuan Cao
AAAI5
2025 Evolving High-Quality Rendering and Reconstruction in a Unified Framework with Contribution-Adaptive Regularization
abstract
Representing 3D scenes from multiview images is a core challenge in computer vision and graphics, which requires both precise rendering and accurate reconstruction. Recently, 3D Gaussian Splatting (3DGS) has garnered significant attention for its high-quality rendering and fast inference speed. Yet, due to the unstructured and irregular nature of Gaussian point clouds, ensuring accurate geometry reconstruction remains difficult. Existing methods primarily focus on geometry regularization, with common approaches including primitive-based and dual-model frameworks. However the former suffers from inherent conflicts between rendering and reconstruction, while the latter is computationally and storage-intensive. To address these challenges, we propose CarGS, a unified model leveraging Contribution-adaptive regularization to achieve simultaneous, high-quality rendering and surface reconstruction. The essence of our framework is learning adaptive contribution for Gaussian primitives by squeezing the knowledge from geometry regularization into a compact MLP. Additionally, we introduce a geometry-guided densification strategy with clues from both normals and Signed Distance Fields (SDF) to improve the capability of capturing high-frequency details. Our design improves the mutual learning of the two tasks, meanwhile its unified structure doesn’t require separate models as in dual-model based approaches, guaranteeing efficiency. Extensive experiments demonstrate CarGS’s ability to achieve state-of-the-art (SOTA) results in both rendering fidelity and reconstruction accuracy while maintaining real-time speed and minimal storage size.
You Shen, Yansong Qu, Shengchuan Zhang, Liujuan Cao
CVPR1
2025 MDC-Seg: Multi-Directional Convolution-Based Semantic Segmentation for LiDAR Point Clouds
abstract
LiDAR point clouds 3D semantic segmentation enables efficient and accurate environmental sensing for intelligent vehicles and autonomous robots, greatly advancing these domains. Existing advanced methods that use 3D sparse convolutional often suffer from a small Effective Receptive Field (ERF), which limits context sensing and challenging highperformance segmentation. Building on this observation, we propose MDC-Seg for efficient ERF enlargement. We design Multi-directional Convolution (MDConv), which simultaneously performs sparse feature encoding on the Bird's Eye View (BEV) and Range View (RV) planes to enlarge the ERF of 3D sparse convolution. To enhance feature fusion in MDConv, we introduce an attention mechanism and design an efficient multifeature fusion (EMFF) module suitable for both 3D and 2D sparse features. To improve segmentation accuracy, we design a point-voxel constraint (PVC) module to handle edge voxels containing multiple point cloud categories, optimizing the final inference results. These modules add minimal memory and inference time but significantly improve performance compared to the baseline. Extensive experiments on the SemanticKITTI benchmark demonstrate MDC-Seg's excellent performance, with supplementary tests on nuScenes further confirming its superiority by yielding good results. The source code is available at https://github.com/OYgreat-river/MDC-Seg.
Xin Ouyang, Xiaolong Qian, Yunzhou Zhang, You Shen, Guiyuan Wang, Wei Liu 0022
ICRA4
2025 Joint Representation Learning Based on Feature Center Region Diffusion and Edge Radiation for Cross-View Geo-Localization
abstract
The essence of the cross-view geo-localization task is to accurately identify the same object across images captured from different viewpoints. Due to variations in image acquisition methods and viewing angles, the content information of the images can differ significantly, which may result in localization failure. Therefore, cross-view geo-localization remains a challenging task. To solve this issue, a joint representation learning network based on feature center region diffusion and edge radiation is proposed in this article. First, to extract the crucial information from the global features, we design the central diffusion module that identifies important regions within the features and enhances feature robustness. Additionally, we design an edge radiation mechanism that expands the receptive field and further highlights crucial information in the image to support the central diffusion module in achieving more stable performance. On this basis, we propose an adaptive triple InfoNCE loss function to assist network training, improving the discriminability of the extracted features. Finally, the proposed network is tested on two mainstream datasets, and experimental results demonstrate that the proposed model outperforms the state-of-the-art methods, which can prove its effectiveness.
Fawei Ge, Yunzhou Zhang, Li Wang 0160, Yixiu Liu, Pengju Si, You Shen
IEEE Trans. Geosci. Remote. Sens.7
2024 VPE-SLAM: Neural Implicit Voxel-permutohedral Encoding for SLAM
abstract
NeRF can reconstruct incredibly realistic environmental maps in dense simultaneous localization and mapping, providing robots with more comprehensive scene map information. However, NeRF often struggles with geometric distortions in indoor reconstructions. To correct geometric distortions, we develop VPE-SLAM, based on the proposed voxel-permutohedral encoding, which can incrementally reconstruct maps of unknown scenes. Specifically, voxel-permutohedral encoding combines a sparse voxel feature grid created by an octree and multi-resolution permutohedral tetrahedral feature grids to represent the scene effectively. Especially when dealing with object edges, our method can effectively encode the geometry and texture of edges by the hybrid structural grid. We propose a novel local bundle adjustment module that utilizes a sliding window mechanism to manage adjacent keyframes requiring optimization. Furthermore, the proposed method establishes local map consistency by repeatedly optimizing keyframes that were initially under-optimized through a compensation strategy. The consistency of the local map can enhance the adaptability of our method to challenging scenes. Extensive experiments demonstrate that our method can achieve accurate camera tracking and produce high-quality reconstruction results on the Replica and ScanNet datasets. The source code will be available at https://github.com/NeuCV-IRMI/VPE-SLAM.
Yunzhou Zhang, You Shen, Lei Rong, Sizhan Wang, Xin Ouyang
ICRA3
2024 HSS-SLAM: Human-in-the-Loop Semantic SLAM Represented by Superquadrics
abstract
The advancement of object detection algorithms has catalyzed the development of object-level semantic SLAM. However, due to missed and false detections, object-level semantic SLAM fails to represent the objects within the scene adequately. Therefore, this paper proposes a novel object-level semantic SLAM termed HSS-SLAM. We incorporate human-in-the-loop into our method, establishing an interaction module to facilitate human editing and rectifying semantic information. Additionally, to minimize the manual correction workload, a lightweight and intuitive method for semantic extension is proposed, augmenting the semantic richness of the global map with a few operations. Furthermore, our method adopts superquadrics for object representation, enabling detailed descriptions of various object shapes. This mitigates the limitation of conventional semantic mapping, where objects are difficult to distinguish due to the reliance on a single-shape representation. Subsequently, precise estimation of superquadric parameters and camera poses is achieved through joint optimization. Extensive experiments conducted on TUM RGB-D and Scenes V2 datasets demonstrate that the proposed approach exhibits competitive performance, surpassing current methods in both object representation and camera localization accuracy.
Yunzhou Zhang, You Shen, Tengda Zhang, Guolu Chen
IROS5
2024 Self-supervised assisted multi-task learning network for one-shot defect segmentation with fake defect generation
Ziqiang Hu, Hao Chu, Yunzhou Zhang, Dexing Shan, You Shen
Pattern Recognit. Lett.5
2024 Fast, Robust, Accurate, Multi-Body Motion Aware SLAM
abstract
Simultaneous ego localization and surrounding object motion awareness are significant issues for the navigation capability of unmanned systems and virtual-real interaction applications. Robust and accurate data association at object and feature levels is one of the key factors in solving this problem. However, currently available solutions ignore the complementarity among different cues in the front-end object association and the negative effects of poorly tracked features on the back-end optimization. It makes them not robust enough in practical applications. Motivated by these observations, we make up rigid environment as a unified whole to assist state decoupling by integrating high-level semantic information, ultimately enabling simultaneous multi-states estimation. A filter-based multi-cues fusion object tracker is proposed for establishing more stable object-level data association. Combined with the object’s motion priors, the motion-aided feature tracking algorithm is proposed to improve the feature-level data association performance. Furthermore, a novel state estimation factor graph is designed which integrates a specific feature observation uncertainty model and the intrinsic priors of tracked object, and solved through sliding-window optimization. Our system is evaluated using the KITTI dataset and achieves comparable performance to state-of-the-art object pose estimation systems both quantitatively and qualitatively. We have also validated our system on simulation environment and a real-world dataset to confirm the potential application value in different practical scenarios.
Linghao Yang, Yunzhou Zhang, Rui Tian 0002, Shiwen Liang, You Shen, Sonya A. Coleman, Dermot Kerr
IEEE Trans. Intell. Transp. Syst.5
2023 BSH-Det3D: Improving 3D Object Detection with BEV Shape Heatmap
abstract
The progress of LiDAR-based 3D object detection has significantly enhanced developments in autonomous driving and robotics. However, due to the limitations of LiDAR sensors, object shapes suffer from deterioration in occluded and distant areas, which creates a fundamental challenge to 3D perception. Existing methods estimate specific 3D shapes and achieve remarkable performance. However, these methods rely on extensive computation and memory, causing imbalances between accuracy and real-time performance. To tackle this challenge, we propose a novel LiDAR-based 3D object detection model named BSH-Det3D, which applies an effective way to enhance spatial features by estimating complete shapes from a bird's eye view (BEV). Specifically, we design the Pillar-based Shape Completion (PSC) module to predict the probability of occupancy whether a pillar contains object shapes. The PSC module generates a BEV shape heatmap for each scene. After integrating with heatmaps, BSH-Det3D can provide additional information in shape deterioration areas and generate high-quality 3D proposals. We also design an attention-based densification fusion module (ADF) to adaptively associate the sparse features with heatmaps and raw points. The ADF module integrates the advantages of points and shapes knowledge with negligible overheads. Extensive experiments on the KITTI benchmark achieve state-of-the-art (SOTA) performance in terms of accuracy and speed, demonstrating the efficiency and flexibility of BSH-Det3D. The source code is available on https://github.com/mystorm16/BSH-Det3D.
You Shen, Yunzhou Zhang, Yanmin Wu, Zhenyu Wang 0010, Linghao Yang, Sonya A. Coleman, Dermot Kerr
IROS1