VLDB 2026 Research / reviewers in the wild / expert
Shuchang Zhou 0001
dblp:67/8436-1
· DBLP profile ↗
29ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0002-7294-5869ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 15 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressive Volume Distillation with Active Learning for Efficient NeRF Architecture Conversion
Shuangkang Fang, Yufeng Wang 0004, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001 |
Int. J. Comput. Vis. | 7 |
| 2026 | Editing 3D Scenes via Text Prompts Without RetrainingabstractNumerous diffusion models have been developed for 2D image synthesis and editing, and recently they are extended to 3D scene editing tasks. However, editing 3D scenes is still in its early stages, and the challenges of scene representations and multi-view consistency need to be addressed. A notable limitation of existing approaches is the need for specific modules for different edits and model retraining for each scene. To tackle these issues, we propose a novel and versatile text-driven 3D scene editing method, termed DN2N, which allows for the direct acquisition of the editing results without the requirement for retraining. Our method employs off-the-shelf text-based editing models of 2D images to modify the multi-view images of a 3D scene. A content filtering process is then applied to discard poorly edited images that disrupt 3D consistency. We consider the remaining inconsistency as a problem of removing noise perturbations and solve it by generating data with similar perturbation characteristics for training. We develop a versatile NeRF model structure and propose two novel cross-view regularization terms to help the DN2N mitigate these perturbations. Empirical results show that our method achieves multiple editing types based solely on text prompts, including but not limited to appearance editing, weather transition, object changing, and style transfer. Most importantly, DN2N exhibits a versatility of editing capabilities, eliminating the need to customize or retrain editing models for specific scenes or editing types. Namely, DN2N achieves comparable total editing time to the 3DGS-based editing method, enhancing its practical value. Shuangkang Fang, Yufeng Wang 0004, Yi-Hsuan Tsai, Wenrui Ding, Yi Yang 0033, Shuchang Zhou 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | UniScene: Unified Occupancy-centric Driving Scene GenerationabstractGenerating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles to model the direct layout-to-data distribution. In this paper, we introduce UniScene, the first unified framework for generating three key data forms — semantic occupancy, video, and LiDAR — in driving scenes. UniScene employs a progressive generation process that decomposes the complex task of scene generation into two hierarchical steps: (a) first generating semantic occupancy from a customized scene layout as a meta scene representation rich in both semantic and geometric information, and then (b) conditioned on occupancy, generating video and LiDAR data, respectively, with two novel transfer strategies of Gaussian-based Joint Rendering and Prior-guided Sparse Modeling. This occupancy-centric approach reduces the generation burden, especially for intricate scenes, while providing detailed intermediate representations for the subsequent generation stages. Extensive experiments demonstrate that UniScene outperforms previous SOTAs in the occupancy, video, and LiDAR generation, which also indeed benefits downstream driving tasks. The Project is available at https://arlo0o.github.io/uniscene/. Bohan Li 0015, Jiazhe Guo, Hongsi Liu, Yingshuang Zou, Yikang Ding, Xiwu Chen, Hu Zhu, Feiyang Tan, Tiancai Wang, Shuchang Zhou 0001, Li Zhang 0040, Xiaojuan Qi 0001, Hao Zhao 0002, Mu Yang, Wenjun Zeng 0001, Xin Jin 0014 |
CVPR | 11 |
| 2025 | NeRF is a Valuable Assistant for 3D Gaussian SplattingabstractWe introduce NeRF-GS, a novel framework that jointly optimizes Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). This framework leverages the inherent continuous spatial representation of NeRF to mitigate several limitations of 3DGS, including sensitivity to Gaussian initialization, limited spatial awareness, and weak inter-Gaussian correlations, thereby enhancing its performance. In NeRF-GS, we revisit the design of 3DGS and progressively align its spatial features with NeRF, enabling both representations to be optimized within the same scene through shared 3D spatial information. We further address the formal distinctions between the two approaches by optimizing residual vectors for both implicit features and Gaussian positions to enhance the personalized capabilities of 3DGS. Experimental results on benchmark datasets show that NeRF-GS surpasses existing methods and achieves state-of-the-art performance. This outcome confirms that NeRF and 3DGS are complementary rather than competing, offering new insights into hybrid approaches that combine 3DGS and NeRF for efficient 3D scene representation. Shuangkang Fang, I-Chao Shen, Takeo Igarashi, Yufeng Wang 0004, Zesheng Wang 0002, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001 |
ICCV | 8 |
| 2025 | MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
Shuangkang Fang, I-Chao Shen, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Shuchang Zhou 0001, Wenrui Ding, Takeo Igarashi, Ming-Hsuan Yang 0001 |
ICCV | 6 |
| 2025 | Arch-Net: Model conversion and quantization for architecture agnostic model deployment
Shuangkang Fang, Zipeng Feng, Song Yuan, Yufeng Wang 0004, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001 |
Neural Networks | 8 |
| 2024 | Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts
Shuangkang Fang, Yufeng Wang 0004, Yi-Hsuan Tsai, Yi Yang 0033, Wenrui Ding, Shuchang Zhou 0001, Ming-Hsuan Yang 0001 |
ECCV (42) | 6 |
| 2024 | Scale-Adaptive Feature Aggregation for Efficient Space-Time Video Super-ResolutionabstractThe Space-Time Video Super-Resolution (STVSR) task aims to enhance the visual quality of videos, by simultaneously performing video frame interpolation (VFI) and video super-resolution (VSR). However, facing the challenge of the additional temporal dimension and scale inconsistency, most existing STVSR methods are complex and inflexible in dynamically modeling different motion amplitudes. In this work, we find that choosing an appropriate processing scale achieves remarkable benefits in flow-based feature propagation. We propose a novel Scale-Adaptive Feature Aggregation (SAFA) network that adaptively selects sub-networks with different processing scales for individual samples. Experiments on four public STVSR benchmarks demonstrate that SAFA achieves state-of-the-art performance. Our SAFA network outperforms recent state-of-the-art methods such as TMNet [83] and VideoINR [10] by an average improvement of over 0.5dB on PSNR, while requiring less than half the number of parameters and only 1/3 computational costs. Zhewei Huang, Ailin Huang, Xiaotao Hu, Jun Xu 0019, Shuchang Zhou 0001 |
WACV | 6 |
| 2023 | One Is All: Bridging the Gap between Neural Radiance Fields Architectures with Progressive Volume DistillationabstractNeural Radiance Fields (NeRF) methods have proved effective as compact, high-quality and versatile representations for 3D scenes, and enable downstream tasks such as editing, retrieval, navigation, etc. Various neural architectures are vying for the core structure of NeRF, including the plain Multi-Layer Perceptron (MLP), sparse tensors, low-rank tensors, hashtables and their compositions. Each of these representations has its particular set of trade-offs. For example, the hashtable-based representations admit faster training and rendering but their lack of clear geometric meaning hampers downstream tasks like spatial-relation-aware editing. In this paper, we propose Progressive Volume Distillation (PVD), a systematic distillation method that allows any-to-any conversions between different architectures, including MLP, sparse or low-rank tensors, hashtables and their compositions. PVD consequently empowers downstream applications to optimally adapt the neural representations for the task at hand in a post hoc fashion. The conversions are fast, as distillation is progressively performed on different levels of volume representations, from shallower to deeper. We also employ special treatment of density to deal with its specific numerical instability problem. Empirical evidence is presented to validate our method on the NeRF-Synthetic, LLFF and TanksAndTemples datasets. For example, with PVD, an MLP-based NeRF model can be distilled from a hashtable-based Instant-NGP model at a 10~20X faster speed than being trained the original NeRF from scratch, while achieving a superior level of synthesis quality. Code is available at https://github.com/megvii-research/AAAI2023-PVD. Shuangkang Fang, Yi Yang 0033, Yufeng Wang 0004, Shuchang Zhou 0001 |
AAAI | 6 |
| 2023 | Three Guidelines You Should Know for Universally Slimmable Self-Supervised LearningabstractWe propose universally slimmable self-supervised learning (dubbed as US3L) to achieve better accuracy-efficiency trade-offs for deploying self-supervised models across different devices. We observe that direct adaptation of self-supervised learning (SSL) to universally slimmable networks misbehaves as the training process frequently collapses. We then discover that temporal consistent guidance is the key to the success of SSL for universally slimmable networks, and we propose three guidelines for the loss design to ensure this temporal consistency from a unified gradient perspective. Moreover, we propose dynamic sampling and group regularization strategies to simultaneously improve training efficiency and accuracy. Our US3L method has been empirically validated on both convolutional neural networks and vision transformers. With only once training and one copy of weights, our method outperforms various state-of-the-art methods (individually trained or not) on benchmarks including recognition, object detection and instance segmentation. Yun-Hao Cao, Peiqin Sun, Shuchang Zhou 0001 |
CVPR | 3 |
| 2023 | A Dynamic Multi-Scale Voxel Flow Network for Video PredictionabstractThe performance of video prediction has been greatly boosted by advanced deep neural networks. However, most of the current methods suffer from large model sizes and require extra inputs, e.g., semantic/depth maps, for promising performance. For efficiency consideration, in this paper, we propose a Dynamic Multi-scale Voxel Flow Network (DMVFN) to achieve better video prediction performance at lower computational costs with only RGB images, than previous methods. The core of our DMVFN is a differentiable routing module that can effectively perceive the motion scales of video frames. Once trained, our DMVFN selects adaptive sub-networks for different inputs at the inference stage. Experiments on several benchmarks demonstrate that our DMVFN is an order of magnitude faster than Deep Voxel Flow [35] and surpasses the state-of-the-art iterative-based OPT [63] on generated image quality. Xiaotao Hu, Zhewei Huang, Ailin Huang, Jun Xu 0019, Shuchang Zhou 0001 |
CVPR | 5 |
| 2023 | UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye ViewabstractIn the field of 3D object detection for autonomous driving, the sensor portfolio including multi-modality and single-modality is diverse and complex. Since the multi-modal methods have system complexity while the accuracy of single-modal ones is relatively low, how to make a tradeoff between them is difficult. In this work, we propose a universal cross-modality knowledge distillation framework (UniDistill) to improve the performance of single-modality detectors. Specifically, during training, UniDistill projects the features of both the teacher and the student detector into Bird's-Eye-View (BEV), which is a friendly representation for different modalities. Then, three distillation losses are calculated to sparsely align the foreground features, helping the student learn from the teacher without introducing additional cost during inference. Taking advantage of the similar detection paradigm of different detectors in BEV, UniDistill easily supports LiDAR-to-camera, camera-to-LiDAR, fusion-to-LiDAR and fusion-to-camera distillation paths. Furthermore, the three distillation losses can filter the effect of misaligned background information and balance between objects of different sizes, improving the distillation effectiveness. Extensive experiments on nuScenes demonstrate that UniDistill effectively improves the mAP and NDS of student detectors by 2.0%~3.2%. Shengchao Zhou, Weizhou Liu, Shuchang Zhou 0001 |
CVPR | 4 |
| 2023 | SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer InferenceabstractTransformers have shown remarkable performance in both natural language processing (NLP) and computer vision (CV) tasks. However, their real-time inference speed and efficiency are limited due to the inefficiency in Softmax and Layer Normalization (LayerNorm). Previous works based on function approximation suffer from inefficient implementation as they place emphasis on computation while disregarding memory overhead concerns. Moreover, such methods rely on retraining to compensate for approximation error which can be costly and inconvenient. In this paper, we present SOLE, a hardware-software co-design for Softmax and LayerNorm which is composed of E2Softmax and AILayerNorm. E2Softmax utilizes log2 quantization of exponent function and log-based division to approximate Softmax while AILayerNorm adopts low-precision statistic calculation. Compared with state-of-the-art designs, we achieve both low-precision calculation and low bit-width storage on Softmax and LayerNorm. Experiments show that SOLE maintains inference accuracy without retraining while offering orders of magnitude speedup and energy savings over GPU, achieving 3.04×, 3.86× energy-efficiency improvements and 2.82×, 3.32× area-efficiency improvements over prior state-of-the-art custom hardware for Softmax and LayerNorm, respectively. Wenxun Wang, Shuchang Zhou 0001, Wenyu Sun, Peiqin Sun, Yongpan Liu |
ICCAD | 2 |
| 2023 | Occ2Net: Robust Image Matching Based on 3D Occupancy Estimation for Occluded RegionsabstractImage matching is a fundamental and critical task in various visual applications, such as Simultaneous Localization and Mapping (SLAM) and image retrieval, which require accurate pose estimation. However, most existing methods ignore the occlusion relations between objects caused by camera motion and scene structure. In this paper, we propose Occ2Net, a novel image matching method that models occlusion relations using 3D occupancy and infers matching points in occluded regions. Thanks to the inductive bias encoded in the Occupancy Estimation (OE) module, it greatly simplifies bootstrapping of a multi-view consistent 3D representation that can then integrate information from multiple views. Together with an Occlusion-Aware (OA) module, it incorporates attention layers and rotation alignment to enable matching between occluded and visible points. We evaluate our method on both real-world and simulated datasets and demonstrate its superior performance over state-of-the-art methods on several metrics, especially in occlusion scenarios. Shuchang Zhou 0001 |
ICCV | 4 |
| 2022 | Synergistic Self-supervised and Quantization Learning
Yun-Hao Cao, Peiqin Sun, Yechang Huang, Jianxin Wu 0001, Shuchang Zhou 0001 |
ECCV (30) | 5 |
| 2022 | Real-Time Intermediate Flow Estimation for Video Frame Interpolation
Zhewei Huang, Wen Heng, Boxin Shi, Shuchang Zhou 0001 |
ECCV (14) | 5 |
| 2022 | FQ-ViT: Post-Training Quantization for Fully Quantized Vision TransformerabstractNetwork quantization significantly reduces model inference complexity and has been widely used in real-world deployments. However, most existing quantization methods have been developed mainly on Convolutional Neural Networks (CNNs), and suffer severe degradation when applied to fully quantized vision transformers. In this work, we demonstrate that many of these difficulties arise because of serious inter-channel variation in LayerNorm inputs, and present, Power-of-Two Factor (PTF), a systematic method to reduce the performance degradation and inference complexity of fully quantized vision transformers. In addition, observing an extreme non-uniform distribution in attention maps, we propose Log-Int-Softmax (LIS) to sustain that and simplify inference by using 4-bit quantization and the BitShift operator. Comprehensive experiments on various transformer-based architectures and benchmarks show that our Fully Quantized Vision Transformer (FQ-ViT) outperforms previous works while even using lower bit-width on attention maps. For instance, we reach 84.89% top-1 accuracy with ViT-L on ImageNet and 50.8 mAP with Cascade Mask R-CNN (Swin-S) on COCO. To our knowledge, we are the first to achieve lossless accuracy degradation (~1%) on fully quantized vision transformers. The code is available at https://github.com/megvii-research/FQ-ViT. Peiqin Sun, Shuchang Zhou 0001 |
IJCAI | 5 |
| 2022 | Perceptual Conversational Head Generation with Regularized Driver and Enhanced RendererabstractThis paper reports our solution for ACM Multimedia ViCo 2022 Conversational Head Generation Challenge, which aims to generate vivid face-to-face conversation videos based on audio and reference images. Our solution focuses on training a generalized audio-to-head driver using regularization and assembling a high visual quality renderer. We carefully tweak the audio-to-behavior model and post-process the generated video using our foreground-background fusion module. We get first place in the listening head generation track and second place in the talking head generation track in the official leaderboard. Our code is available at https://github.com/megvii-research/MM2022-ViCoPerceptualHeadGeneration. Ailin Huang, Zhewei Huang, Shuchang Zhou 0001 |
ACM Multimedia | 3 |
| 2021 | Multi-split Reversible Transformers Can Enhance Neural Machine TranslationabstractLarge-scale transformers have been shown the state-of-the-art on neural machine translation.However, training these increasingly wider and deeper models could be tremendously memory intensive.We reduce the memory burden by employing the idea of reversible networks that a layer's input can be reconstructed from its output.We design three types of multi-split based reversible transformers.We also devise a corresponding backpropagation algorithm, which does not need to store activations for most layers.Furthermore, we present two fine-tuning techniques: splits shuffle and self ensemble, to boost translation accuracy.Specifically, our best models surpass the vanilla transformer by at least 1.4 BLEU points in three datasets.Our largescale reversible models achieve 30.0 BLEU in WMT'14 En-De and 43.5 BLEU in WMT'14 En-Fr, beating several very strong baselines with less than half of the training memory. Yuekai Zhao, Shuchang Zhou 0001, Zhihua Zhang 0004 |
EACL | 2 |
| 2020 | Data-Efficient Semi-Supervised Learning by Reliable Edge MiningabstractLearning powerful discriminative features is a challenging task in Semi-Supervised Learning, as the estimation of the feature space is more likely to be wrong with scarcer labeled data. Previous methods utilize a relation graph with edges representing 'similarity' or 'dissimilarity' between nodes. Similar nodes are forced to output consistent features, while dissimilar nodes are forced to be inconsistent. However, since unlabeled data may be wrongly labeled, the judgment of edges may be unreliable. Besides, the nodes connected by edges may already be well fitted, thus contributing little to the model training. We propose Reliable Edge Mining (REM), which forms a reliable graph by only selecting reliable and useful edges. Guided by the graph, the feature extractor is able to learn discriminative features in a data-efficient way, and consequently boosts the accuracy of the learned classifier. Visual analyses show that the features learned are more discriminative and better reveals the underlying structure of the data. REM can be combined with perturbation-based methods like Pi-model, TempEns and Mean Teacher to further improve accuracy. Experiments prove that our method is data-efficient on simple tasks like SVHN and CIFAR-10, and achieves state-of-the-art results on the challenging CIFAR-100. Peibin Chen, Xu Qin, Weidi Xu, Shuchang Zhou 0001 |
CVPR | 5 |
| 2019 | Learning to Paint With Model-Based Deep Reinforcement LearningabstractWe show how to teach machines to paint like human painters, who can use a small number of strokes to create fantastic paintings. By employing a neural renderer in model-based Deep Reinforcement Learning (DRL), our agents learn to determine the position and color of each stroke and make long-term plans to decompose texture-rich images into strokes. Experiments demonstrate that excellent visual effects can be achieved using hundreds of strokes. The training process does not require the experience of human painters or stroke tracking data. The code is available at https://github.com/hzwer/ICCV2019-LearningToPaint. Zhewei Huang, Shuchang Zhou 0001, Wen Heng |
ICCV | 2 |
| 2017 | GeneGAN: Learning Object Transfiguration and Object Subspace from Unpaired Data
Shuchang Zhou 0001, Taihong Xiao, Yi Yang 0033, Dieqiao Feng, Qinyao He, Weiran He |
BMVC | 1 |
| 2017 | EAST: An Efficient and Accurate Scene Text DetectorabstractPrevious approaches for scene text detection have already achieved promising performances across various benchmarks. However, they usually fall short when dealing with challenging scenarios, even when equipped with deep neural network models, because the overall performance is determined by the interplay of multiple stages and components in the pipelines. In this work, we propose a simple yet powerful pipeline that yields fast and accurate text detection in natural scenes. The pipeline directly predicts words or text lines of arbitrary orientations and quadrilateral shapes in full images, eliminating unnecessary intermediate steps (e.g., candidate aggregation and word partitioning), with a single neural network. The simplicity of our pipeline allows concentrating efforts on designing loss functions and neural network architecture. Experiments on standard datasets including ICDAR 2015, COCO-Text and MSRA-TD500 demonstrate that the proposed algorithm significantly outperforms state-of-the-art methods in terms of both accuracy and efficiency. On the ICDAR 2015 dataset, the proposed algorithm achieves an F-score of 0.7820 at 13.2fps at 720p resolution. Xinyu Zhou 0004, Cong Yao, Yuzhi Wang, Shuchang Zhou 0001, Weiran He, Jiajun Liang |
CVPR | 5 |
| 2017 | IQNN: Training Quantized Neural Networks with Iterative Optimizations
Shuchang Zhou 0001, Taihong Xiao, Xinyu Zhou 0004 |
ICANN (2) | 1 |
| 2017 | Balanced Quantization: An Effective and Efficient Approach to Quantized Neural Networks
Shuchang Zhou 0001, Yuzhi Wang, Qinyao He, Yuheng Zou |
J. Comput. Sci. Technol. | 1 |
| 2013 | Deceptive Answer Prediction with User Preference Graph
Fangtao Li, Shuchang Zhou 0001, Xiance Si, Decheng Dai |
ACL (1) | 3 |
| 2013 | A Scalable Approach to Column-Based Low-Rank Matrix Approximation
Yifan Pi, Haoruo Peng, Shuchang Zhou 0001, Zhihua Zhang 0004 |
IJCAI | 3 |
| 2012 | Sublinear Algorithms for Penalized Logistic Regression in Massive Datasets
Haoruo Peng, Edward Y. Chang, Shuchang Zhou 0001, Zhihua Zhang 0004 |
ECML/PKDD (1) | 4 |
| 2010 | An Efficient Simulation Algorithm for Cache of Random Replacement Policy
Shuchang Zhou 0001 |
NPC | 1 |