Shuangkang Fang

dblp:262/4025 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
14since 2021 · last 2026
0009-0001-1066-5896ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Progressive Volume Distillation with Active Learning for Efficient NeRF Architecture Conversion
Shuangkang Fang, Yufeng Wang 0004, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001
Int. J. Comput. Vis.1
2026 Geometry-Aware Joint Attention for Efficient Native 3D Editing
Shuangkang Fang, Weicai Ye, Xuanyang Zhang, Yan-Pei Cao 0001, Gang Yu 0002, Tao Chen 0003
IEEE Signal Process. Lett.2
2026 Editing 3D Scenes via Text Prompts Without Retraining
abstract
Numerous diffusion models have been developed for 2D image synthesis and editing, and recently they are extended to 3D scene editing tasks. However, editing 3D scenes is still in its early stages, and the challenges of scene representations and multi-view consistency need to be addressed. A notable limitation of existing approaches is the need for specific modules for different edits and model retraining for each scene. To tackle these issues, we propose a novel and versatile text-driven 3D scene editing method, termed DN2N, which allows for the direct acquisition of the editing results without the requirement for retraining. Our method employs off-the-shelf text-based editing models of 2D images to modify the multi-view images of a 3D scene. A content filtering process is then applied to discard poorly edited images that disrupt 3D consistency. We consider the remaining inconsistency as a problem of removing noise perturbations and solve it by generating data with similar perturbation characteristics for training. We develop a versatile NeRF model structure and propose two novel cross-view regularization terms to help the DN2N mitigate these perturbations. Empirical results show that our method achieves multiple editing types based solely on text prompts, including but not limited to appearance editing, weather transition, object changing, and style transfer. Most importantly, DN2N exhibits a versatility of editing capabilities, eliminating the need to customize or retrain editing models for specific scenes or editing types. Namely, DN2N achieves comparable total editing time to the 3DGS-based editing method, enhancing its practical value.
Shuangkang Fang, Yufeng Wang 0004, Yi-Hsuan Tsai, Wenrui Ding, Yi Yang 0033, Shuchang Zhou 0001, Ming-Hsuan Yang 0001
IEEE Trans. Vis. Comput. Graph.1
2025 ASFC-NeRF: Large-Scale Scene Rendering with Adaptive Sampling and Feature-aware Compression
abstract
While significant progress has been made in large-scale scene representation using Neural Radiance Fields (NeRF), several limitations remain. For instance, most methods still rely on the original coarse-to-fine sampling strategy, leading to an inefficient rendering process. Additionally, to model larger scenes, these methods often use complex network models, resulting in redundant model parameters. To address these issues, we propose a novel model with adaptive sampling and feature-aware compression for large-scale scene rendering, named ASFC-NeRF. We first introduce a weight prediction network to replace the original coarse sampling strategy, then employ a teacher network and depth constraints for knowledge distillation in the early stages of training to enhance the high-fidelity of the scene. Furthermore, we optimize the number of Grids and the channels of Planes and prune the network to efficiently compress model parameters. Experimental results demonstrate that our method significantly accelerates the rendering process and greatly reduces parameter quantity while maintaining or only slightly lowering image quality. Therefore, ASFC-NeRF exhibits advantages in comprehensive performance and practicality.
Yufeng Wang 0004, Shuangkang Fang, Zesheng Wang 0002, Dacheng Qi, Wenrui Ding
ICASSP3
2025 NeRF is a Valuable Assistant for 3D Gaussian Splatting
abstract
We introduce NeRF-GS, a novel framework that jointly optimizes Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). This framework leverages the inherent continuous spatial representation of NeRF to mitigate several limitations of 3DGS, including sensitivity to Gaussian initialization, limited spatial awareness, and weak inter-Gaussian correlations, thereby enhancing its performance. In NeRF-GS, we revisit the design of 3DGS and progressively align its spatial features with NeRF, enabling both representations to be optimized within the same scene through shared 3D spatial information. We further address the formal distinctions between the two approaches by optimizing residual vectors for both implicit features and Gaussian positions to enhance the personalized capabilities of 3DGS. Experimental results on benchmark datasets show that NeRF-GS surpasses existing methods and achieves state-of-the-art performance. This outcome confirms that NeRF and 3DGS are complementary rather than competing, offering new insights into hybrid approaches that combine 3DGS and NeRF for efficient 3D scene representation.
Shuangkang Fang, I-Chao Shen, Takeo Igarashi, Yufeng Wang 0004, Zesheng Wang 0002, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001
ICCV1
2025 MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
Shuangkang Fang, I-Chao Shen, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Shuchang Zhou 0001, Wenrui Ding, Takeo Igarashi, Ming-Hsuan Yang 0001
ICCV1
2025 Guiding Yourself with Your Own Insights: Student-Driven Knowledge Distillation
abstract
Knowledge distillation (KD) stands as an efficient technique for compressing models, typically employing a teacher-student framework. Nevertheless, optimizing KD to yield models with reduced parameters and enhanced performance remains an area warranting deeper investigation. In this article, we recognize the significance of preserving structural consistency to enhance knowledge transfer efficiency between networks. Leveraging this insight, we introduce a novel approach termed Student-Driven Knowledge Distillation (SDKD), which integrates a proxy teacher intermediary between the primary teacher and student model. Specifically, we construct the architecture of the proxy teacher entirely based on the student network to generate logits that closely align with the distribution of the student network. Besides, we propose a Feature Fusion Block (FFB) to integrate features from the teacher network into the proxy teacher. FFB can not only provide high-quality feature-based knowledge for distillation but also impart response-based knowledge to facilitate the learning process. Extensive experiments illustrate that SDKD outperforms 29 state-of-the-art methods on several tasks, including image classification, semantic segmentation, and depth estimation.
Dacheng Qi, Yufeng Wang 0004, Shuangkang Fang, Zehao Zhang, Zesheng Wang 0002, Wenrui Ding
ICME4
2025 SANE: Enhancing Large-scale Scene Representation with Semantic-aware NeRF Experts
abstract
We propose the Semantic-aware NeRF Experts (SANE), which fully exploits the intrinsic characteristics of large-scale scenes, including semantics and material features, to achieve high-quality novel view synthesis results and provide accurate 3D semantic information. SANE begins by building a semantic Mixture of Experts (MoE), utilizing a learnable gating network to semantically partition the scene into blocks for corresponding NeRF experts. We then develop a semantic volume rendering scheme that integrates discrete semantics into the end-to-end differentiable process of NeRF, enabling refined semantic labeling of each scene point. Additionally, we implement a dual-implicit encoding strategy: intra-block encoding captures lighting variations across viewpoints, while inter-block one captures texture features among different semantic objects. Experiments on benchmark datasets show that SANE delivers higher-quality scene representations and effective semantic decomposition for downstream tasks, such as precise editing of large-scale scenes based on semantics.
Zesheng Wang 0002, Yufeng Wang 0004, Shuangkang Fang, Dacheng Qi, Shengxi Li, Mai Xu, Wenrui Ding
ICME3
2025 Arch-Net: Model conversion and quantization for architecture agnostic model deployment
Shuangkang Fang, Zipeng Feng, Song Yuan, Yufeng Wang 0004, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001
Neural Networks1
2024 Efficient Implicit SDF and Color Reconstruction via Shared Feature Field
Shuangkang Fang, Dacheng Qi, Yufeng Wang 0004, Zehao Zhang, Zeqi Shao, Wenrui Ding
ACCV (10)1
2024 Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts
Shuangkang Fang, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001, Ming-Hsuan Yang 0001
ECCV (42)1
2024 UAV-ENeRF: Text-Driven UAV Scene Editing With Neural Radiance Fields
abstract
3D reconstruction of Unmanned Aerial Vehicle (UAV) scenes is vital for agriculture, environmental protection, urban planning, and disaster response, to name a few. However, data acquisition can be constrained and hazardous under hostile environments, which limits the image data available in real-world applications. In this work, we propose a text-driven online editing framework for UAV scenes, which can generate novel views of existing scenes with abundant editing types. Compared with small single-object scenes, large-scale UAV scene editing suffers from several particular challenges: 1) broader capturing scope exhibits illumination variation and complicated objects that reduce the 3D scene consistency after editing; and 2) high-resolution 2D editing and 3D reconstruction can be computationally expensive with tremendous GPU memory. To tackle these issues, we first design a dual-branch compact NeRF structure to reduce memory usage and enhance accuracy for 3D reconstruction. We then introduce a sub-pixel sampling scheme to expedite the generation of low-resolution images for 2D editing, followed by a super-resolution module that restores the fine details of rendered images. Additionally, we develop a grouped content filtering mechanism to improve the 3D scene consistency of the model by matching the rendering images and text descriptions, which also significantly reduces memory usage during editing. Extensive experiments demonstrate that the proposed method can achieve various editing effects, including different seasons, weather conditions, times of the day, disaster scenarios, etc. Our technique is computationally efficient and conveniently expandable for large-scale UAV scenes, alleviating data scarcity in harsh scenarios.
Yufeng Wang 0004, Shuangkang Fang, Zehao Zhang, Xianlin Zeng, Wenrui Ding
IEEE Trans. Geosci. Remote. Sens.2
2023 One Is All: Bridging the Gap between Neural Radiance Fields Architectures with Progressive Volume Distillation
abstract
Neural Radiance Fields (NeRF) methods have proved effective as compact, high-quality and versatile representations for 3D scenes, and enable downstream tasks such as editing, retrieval, navigation, etc. Various neural architectures are vying for the core structure of NeRF, including the plain Multi-Layer Perceptron (MLP), sparse tensors, low-rank tensors, hashtables and their compositions. Each of these representations has its particular set of trade-offs. For example, the hashtable-based representations admit faster training and rendering but their lack of clear geometric meaning hampers downstream tasks like spatial-relation-aware editing. In this paper, we propose Progressive Volume Distillation (PVD), a systematic distillation method that allows any-to-any conversions between different architectures, including MLP, sparse or low-rank tensors, hashtables and their compositions. PVD consequently empowers downstream applications to optimally adapt the neural representations for the task at hand in a post hoc fashion. The conversions are fast, as distillation is progressively performed on different levels of volume representations, from shallower to deeper. We also employ special treatment of density to deal with its specific numerical instability problem. Empirical evidence is presented to validate our method on the NeRF-Synthetic, LLFF and TanksAndTemples datasets. For example, with PVD, an MLP-based NeRF model can be distilled from a hashtable-based Instant-NGP model at a 10~20X faster speed than being trained the original NeRF from scratch, while achieving a superior level of synthesis quality. Code is available at https://github.com/megvii-research/AAAI2023-PVD.
Shuangkang Fang, Yi Yang 0033, Yufeng Wang 0004, Shuchang Zhou 0001
AAAI1
2023 DFFG: Fast Gradient Iteration for Data-free Quantization
Huixing Leng, Shuangkang Fang, Yufeng Wang 0004, Zehao Zhang, Dacheng Qi, Wenrui Ding
BMVC2
2020 Parallel design of sparse deep belief network with multi-objective optimization
Yangyang Li 0001, Shuangkang Fang, Licheng Jiao, Naresh Marturi
Inf. Sci.2