Shaocong Dong

dblp:329/6563 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 From One to More: Contextual Part Latents for 3D Generation
abstract
To generate 3D objects, early research focused on multi-view-driven approaches relying solely on 2D renderings. Recently, the 3D native latent diffusion paradigm has demonstrated superior performance in 3D generation, because it fully leverages the geometric information provided in ground truth 3D data. Despite its fast development, 3D diffusion still faces three challenges. First, the majority of these methods represent a 3D object by one single latent, regardless of its complexity. This may lead to detail loss when generating 3D objects with multiple complicated parts. Second, most 3D assets are designed parts by parts, yet the current holistic latent representation overlooks the independence of these parts and their interrelationships, limiting the model's generative ability. Third, current methods rely on global conditions (e.g., text, image, point cloud) to control the generation process, lacking detailed controllability. Therefore, motivated by how 3D designers create a 3D object, we present a new part-based 3D generation framework, CoPart, which represents a 3D object with multiple contextual part latents and simultaneously generates coherent 3D parts. This part-based framework has several advantages, including: i) reduces the encoding burden of intricate objects by decomposing them into simpler parts, ii) facilitates part learning and part relationship modeling, and iii) naturally supports part-level control. Furthermore, to ensure the coherence of part latents and to harness the powerful priors from foundation models, we propose a novel mutual guidance strategy to fine-tune pre-trained diffusion models for joint part latent denoising. Benefiting from the part-based representation, we demonstrate that CoPart can support various applications including part-editing, articulated object generation, and mini-scene generation. Moreover, we collect a new large-scale 3D part dataset named Partverse from Objaverse through automatic mesh segmentation and subsequent human post-annotations. By training on the proposed dataset, CoPart achieves promising part-based 3D generation with high controllability. Project page: https://hkdsc.github.io/project/copart.
Shaocong Dong, Lihe Ding, Yaokun Li, Jaehyeok Kim, Chenjian Gao, Zhanpeng Huang, Zibin Wang, Tianfan Xue
ICCV1
2025 ERL-RTDETR: A Lightweight Transformer-Based Framework for High-Accuracy Apple Disease Detection in Precision Agriculture
abstract
ABSTRACT Apples are deeply favored by consumers for their crisp and sweet taste and play a significant role in agricultural production. However, apples often suffer from infections by various pathogens during their growth process, severely impacting fruit quality and yield, and subsequently causing economic losses. Therefore, timely detection and accurate intervention against diseases during apple growth are crucial for improving harvest management efficiency and economic benefits. Nonetheless, current research primarily focuses on the identification of single diseases, lacking multi‐disease detection capabilities. This limitation results in inadequate timeliness and accuracy in disease management, thereby restricting practical application effectiveness. Additionally, apple disease detection models need to balance high accuracy, rapid response, and lightweight design to reduce hardware costs and application thresholds. To address these challenges, this paper proposes a lightweight detection model named ERL‐RTDETR, which is based on RT‐DETR. First, a dataset containing 3096 images of apple‐leaf diseases was constructed, encompassing different camera angles, time spans, and lighting conditions in complex environments. Subsequently, by introducing an Efficient Multi‐scale Attention (EMA) mechanism and integrating it with the backbone network, we designed a new feature extraction module (BasicBlock_EMA) to enhance the capture of fine‐grained features. Meanwhile, in the neck network, the traditional convolutional module was replaced with a Lightweight Adaptive Extraction module (LAE), and a Generalized Efficient Lightweight Attention Network (GELAN) was introduced to optimize the convolutional blocks, thereby improving the model's training efficiency and detection performance for subtle targets. The construction of the ERL‐RTDETR model was completed while ensuring detection accuracy and reducing model complexity. Experimental results demonstrate that ERL‐RTDETR achieves a balanced performance in apple disease detection tasks, with a detection precision of 94.5% on the test set (a 3.2% improvement compared to RT‐DETR) and increases in mAP50 and mAP50:95 by 2.7% and 2.2%, respectively. Simultaneously, the GFLOPs were reduced by 5.9 GFLOPs (a decrease of 10.3% compared to RT‐DETR). In summary, the proposed ERL‐RTDETR model provides an efficient, lightweight, and accurate method for apple disease detection, serving as an important reference for research and practical applications in related fields.
Shaocong Dong
Concurr. Comput. Pract. Exp.3
2025 MFAR-Net: Multi-level feature interaction and Dual-Dimension adaptive reinforcement network for breast lesion segmentation in ultrasound images
Guoqi Liu, Shaocong Dong, Sheng Yao 0005, Dong Liu 0008
Expert Syst. Appl.2
2025 An efficient scale-aware model based on the improved RT-DETR for pomegranate growth stage detection
Shaocong Dong
Neurocomputing2
2024 Text-to-3D Generation with Bidirectional Diffusion Using Both 2D and 3D Priors
abstract
Most 3D generation research focuses on up-projecting 2D foundation models into the 3D space, either by minimizing 2D Score Distillation Sampling (SDS) loss or fine-tuning on multi-view datasets. Without explicit 3D priors, these methods often lead to geometric anomalies and multi-view inconsistency. Recently, researchers have attempted to improve the genuineness of 3D objects by directly training on 3D datasets, albeit at the cost of low-quality texture generation due to the limited texture diversity in 3D datasets. To harness the advantages of both approaches, we propose Bidirectional Diffusion (BiDiff), a unified framework that incorporates both a 3D and a 2D diffusion process, to preserve both 3D fidelity and 2D texture richness, respectively. Moreover, as a simple combination may yield inconsistent generation results, we further bridge them with novel bidirectional guidance. In addition, our method can be used as an initialization of optimization-based models to further improve the quality of 3D models and the efficiency of optimization, reducing the process from 3.4 hours to 20 minutes. Experimental results have shown that our model achieves high-quality, diverse, and scalable 3D generation. Project website https://bidiff.github.io/.
Lihe Ding, Shaocong Dong, Zhanpeng Huang, Zibin Wang, Kaixiong Gong, Dan Xu 0002, Tianfan Xue
CVPR2
2024 Interactive3D: Create What You Want by Interactive 3D Generation
abstract
3D object generation has undergone significant advancements, yielding high-quality results. However, fall short of achieving precise user control, often yielding results that do not align with user expectations, thus limiting their applicability. User-envisioning 3D object generation faces significant challenges in realizing its concepts using current generative models due to limited interaction capabilities. Existing methods mainly offer two approaches: (i) interpreting textual instructions with constrained controllability, or (ii) reconstructing 3D objects from 2D images. Both of them limit customization to the confines of the 2D reference and potentially introduce undesirable artifacts during the 3D lifting process, restricting the scope for direct and versatile 3D modifications. In this work, we introduce Interactive3D, an innovative framework for interactive 3D generation that grants users precise control over the generative process through extensive 3D interaction capabilities. Interactive3D is constructed in two cascading stages, utilizing distinct 3D representations. The first stage employs Gaussian Splatting for direct user interaction, allowing modifications and guidance of the generative direction at any intermediate step through (i) Adding and Removing components, (ii) Deformable and Rigid Dragging, (iii) Geometric Transformations, and (iv) Semantic Editing. Subsequently, the Gaussian splats are transformed into InstantNGP. We introduce a novel (v) Interactive Hash Refinement module to further add details and extract the geometry in the second stage. Our experiments demonstrate that proposed Interactive3D markedly improves the controllability and quality of 3D generation. Our project webpage is available at https://interactive-3d.github.io/.
Shaocong Dong, Lihe Ding, Zhanpeng Huang, Zibin Wang, Tianfan Xue, Dan Xu 0002
CVPR1
2024 MsSVT++: Mixed-Scale Sparse Voxel Transformer With Center Voting for 3D Object Detection
abstract
Accurate 3D object detection in large-scale outdoor scenes, characterized by considerable variations in object scales, necessitates features rich in both long-range and fine-grained information. While recent detectors have utilized window-based transformers to model long-range dependencies, they tend to overlook fine-grained details. To bridge this gap, we propose MsSVT++, an innovative Mixed-scale Sparse Voxel Transformer that simultaneously captures both types of information through a divide-and-conquer approach. This approach involves explicitly dividing attention heads into multiple groups, each responsible for attending to information within a specific range. The outputs of these groups are subsequently merged to obtain final mixed-scale features. To mitigate the computational complexity associated with applying a window-based transformer in 3D voxel space, we introduce a novel Chessboard Sampling strategy and implement voxel sampling and gathering operations sparsely using a hash map. Moreover, an important challenge stems from the observation that non-empty voxels are primarily located on the surface of objects, which impedes the accurate estimation of bounding boxes. To overcome this challenge, we introduce a Center Voting module that integrates newly voted voxels enriched with mixed-scale contextual information towards the centers of the objects, thereby improving precise object localization. Extensive experiments demonstrate that our single-stage detector, built upon the foundation of MsSVT++, consistently delivers exceptional performance across diverse datasets.
Jianan Li 0001, Shaocong Dong, Lihe Ding, Tingfa Xu
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Sample-adaptive Augmentation for Point Cloud Recognition Against Real-world Corruptions
abstract
Robust 3D perception under corruption has become an essential task for the realm of 3D vision. While current data augmentation techniques usually perform random transformations on all point cloud objects in an offline way and ignore the structure of the samples, resulting in over-or-under enhancement. In this work, we propose an alternative to make sample-adaptive transformations based on the structure of the sample to cope with potential corruption via an auto-augmentation framework, named as Adapt-Point. Specially, we leverage a imitator, consisting of a Deformation Controller and a Mask Controller, respectively in charge of predicting deformation parameters and producing a per-point mask, based on the intrinsic structural information of the input point cloud, and then conduct corruption simulations on top. Then a discriminator is utilized to prevent the generation of excessive corruption that deviates from the original data distribution. In addition, a perception-guidance feedback mechanism is incorporated to guide the generation of samples with appropriate difficulty level. Furthermore, to address the paucity of real-world corrupted point cloud, we also introduce a new dataset ScanObjectNN-C, that exhibits greater similarity to actual data in real-world environments, especially when contrasted with preceding CAD datasets. Experiments show that our method achieves state-of-the-art results on multiple corruption benchmarks, including ModelNet-C, our ScanObjectNN-C, and ShapeNet-C.
Jie Wang 0097, Lihe Ding, Tingfa Xu, Shaocong Dong, Xinli Xu, Long Bai 0008, Jianan Li 0001
ICCV4
2022 FH-Net: A Fast Hierarchical Network for Scene Flow Estimation on Real-World Point Clouds
Lihe Ding, Shaocong Dong, Tingfa Xu, Xinli Xu, Jie Wang 0097, Jianan Li 0001
ECCV (39)2
2022 MsSVT: Mixed-scale Sparse Voxel Transformer for 3D Object Detection on Point Clouds
abstract
3D object detection from the LiDAR point cloud is fundamental to autonomous driving. Large-scale outdoor scenes usually feature significant variance in instance scales, thus requiring features rich in long-range and fine-grained information to support accurate detection. Recent detectors leverage the power of window-based transformers to model long-range dependencies but tend to blur out fine-grained details. To mitigate this gap, we present a novel Mixed-scale Sparse Voxel Transformer, named MsSVT, which can well capture both types of information simultaneously by the divide-and-conquer philosophy. Specifically, MsSVT explicitly divides attention heads into multiple groups, each in charge of attending to information within a particular range. All groups' output is merged to obtain the final mixed-scale features. Moreover, we provide a novel chessboard sampling strategy to reduce the computational complexity of applying a window-based transformer in 3D voxel space. To improve efficiency, we also implement the voxel sampling and gathering operations sparsely with a hash map. Endowed by the powerful capability and high efficiency of modeling mixed-scale information, our single-stage detector built on top of MsSVT surprisingly outperforms state-of-the-art two-stage detectors on Waymo. Our project page: https://github.com/dscdyc/MsSVT.
Shaocong Dong, Lihe Ding, Tingfa Xu, Xinli Xu, Jie Wang 0097, Ziyang Bian, Ying Wang 0064, Jianan Li 0001
NeurIPS1
2022 CAGroup3D: Class-Aware Grouping for 3D Object Detection on Point Clouds
abstract
We present a novel two-stage fully sparse convolutional 3D object detection framework, named CAGroup3D. Our proposed method first generates some high-quality 3D proposals by leveraging the class-aware local group strategy on the object surface voxels with the same semantic predictions, which considers semantic consistency and diverse locality abandoned in previous bottom-up approaches. Then, to recover the features of missed voxels due to incorrect voxel-wise segmentation, we build a fully sparse convolutional RoI pooling module to directly aggregate fine-grained spatial information from backbone for further proposal refinement. It is memory-and-computation efficient and can better encode the geometry-specific features of each 3D proposal. Our model achieves state-of-the-art 3D detection performance with remarkable gains of +3.6% on ScanNet V2 and +2.6% on SUN RGB-D in term of [email protected]. Code will be available at https://github.com/Haiyang-W/CAGroup3D.
Lihe Ding, Shaocong Dong, Shaoshuai Shi, Aoxue Li, Jianan Li 0001, Zhenguo Li, Liwei Wang 0001
NeurIPS3