VLDB 2026 Research / reviewers in the wild / expert
Shuzhao Xie
dblp:319/4419
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2025
0009-0008-3017-1077ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Implicit Neural Representations via Symmetric Power TransformationabstractWe propose symmetric power transformation to enhance the capacity of Implicit Neural Representation (INR) from the perspective of data transformation. Unlike prior work utilizing random permutation or index rearrangement, our method features a reversible operation that does not require additional storage consumption. Specifically, we first investigate the characteristics of data that can benefit the training of INR, proposing the Range-Defined Symmetric Hypothesis, which posits that specific range and symmetry can improve the expressive ability of INR. Based on this hypothesis, we propose a nonlinear symmetric power transformation to achieve both range-defined and symmetric properties simultaneously. We use the power coefficient to redistribute data to approximate symmetry within the target range. To improve the robustness of the transformation, we further design deviation-aware calibration and adaptive soft boundary to address issues of extreme deviation boosting and continuity breaking. Extensive experiments are conducted to verify the performance of the proposed method, demonstrating that our transformation can reliably improve INR compared with other data transformations. We also conduct 1D audio, 2D image and 3D video fitting tasks to demonstrate the effectiveness and applicability of our method. Weixiang Zhang, Shuzhao Xie, Chengwei Ren, Shijia Ge, Mingzi Wang |
AAAI | 2 |
| 2025 | EVOS: Efficient Implicit Neural Training via EVOlutionary SelectorabstractWe propose EVOlutionary Selector (EVOS), an efficient training paradigm for accelerating Implicit Neural Representation (INR). Unlike conventional INR training that feeds all samples through the neural network in each iteration, our approach restricts training to strategically selected points, reducing computational overhead by eliminating redundant forward passes. Specifically, we treat each sample as an individual in an evolutionary process, where only those fittest ones survive and merit inclusion in training, adaptively evolving with the neural network dynamics. While this is conceptually similar to Evolutionary Algorithms, their distinct objectives (selection for acceleration vs. iterative solution optimization) require a fundamental redefinition of evolutionary mechanisms for our context. In response, we design sparse fitness evaluation, frequency-guided crossover, and augmented unbiased mutation to comprise EVOS. These components respectively guide sample selection with reduced computational cost, enhance performance through frequency-domain balance, and mitigate selection bias from cached evaluation. Extensive experiments demonstrate that our method achieves approximately 48%-66% reduction in training time while ensuring superior convergence without additional cost, establishing state-of-the-art acceleration among recent sampling-based strategies. Our code is available at this link. Weixiang Zhang, Shuzhao Xie, Chengwei Ren, Siyi Xie, Shijia Ge, Mingzi Wang |
CVPR | 2 |
| 2025 | Lungmix: A Mixup-Based Strategy for Generalization in Respiratory Sound ClassificationabstractRespiratory sound classification plays a pivotal role in diagnosing respiratory diseases. While deep learning models have succeeded with various respiratory sound datasets, our experiments indicate that models trained on one dataset often fail to generalize effectively to others, mainly due to data collection and annotation inconsistencies. To address this limitation, we introduce Lungmix, a novel data augmentation technique inspired by Mixup. Lungmix generates augmented data by blending waveforms using loudness while interpolating labels based on their semantic meaning, helping the model learn more generalized representations. Extensive evaluations across three datasets (ICBHI, SPR, and HF) demonstrate that Lungmix significantly enhances model generalization to unseen data. In particular, Lungmix boosts the 4-class classification score by up to 3.55%, attaining performance comparable to models trained on the target dataset directly. Shijia Ge, Weixiang Zhang, Shuzhao Xie, Baixu Yan |
ICASSP | 3 |
| 2025 | PulmoScan: A Practical Pulmonary Disease Pre-Screening SystemabstractAutomation of pulmonary disease identification has been a long-standing area of research and gained increased attention after the COVID-19 pandemic. However, existing respiratory sound classification algorithms exhibit significant limitations, including suboptimal performance, insufficient input robustness, and inadequate alignment with clinical evaluation metrics, thereby hindering their practical implementation. To address these limitations, we introduce PulmoScan, a practical pulmonary disease pre-screening system. PulmoScan comprises three fundamental modules: a Respiratory Sound Quality Validator that ensures the robustness of input data, a Runtime Decision Booster that improves performance and adapting to variating evaluation metrics, and a Symptom Enhancement Diagnoser that augments respiratory sound classification with comprehensive disease pre-screening capabilities. Beyond its primary function, PulmoScan exemplifies a methodological framework for translating theoretically limited algorithms into viable clinical applications, demonstrating essential considerations and procedural adaptations for real-world implementation. Baixu Yan, Shijia Ge, Meizi Lu, Weixiang Zhang, Shuzhao Xie |
ICASSP | 5 |
| 2025 | Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
Ronghui Li, Shukai Fang, Shuzhao Xie, Jiaqing Zhou, Junkun Peng |
ICCV | 4 |
| 2025 | Expansive Supervision for Neural Radiance FieldsabstractNeural Radiance Field (NeRF) has achieved remarkable success in creating immersive media representations through its exceptional reconstruction capabilities. However, the computational demands of dense forward passes and volume rendering during training continue to challenge its real-world applications. In this paper, we introduce Expansive Supervision to reduce time and memory costs during NeRF training from the perspective of partial ray selection for supervision. Specifically, we observe that training errors exhibit a long-tail distribution correlated with image content. Based on this observation, our method selectively renders a small but crucial subset of pixels and expands their values to estimate errors across the entire area for each iteration. Compared to conventional supervision, our approach effectively bypasses redundant rendering processes, resulting in substantial reductions in both time and memory consumption. Experimental results demonstrate that integrating Expansive Supervision within existing state-of-the-art acceleration frameworks achieves 52% memory savings and 16% time savings while maintaining comparable visual quality. Our code is available at this link. Weixiang Zhang, Shuzhao Xie, Shijia Ge, Zhi Wang 0001 |
ICME | 3 |
| 2025 | SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer ProgrammingabstractRecent advances in 3D Gaussian Splatting (3DGS) have greatly improved 3D reconstruction. However, its substantial data size poses a significant challenge for transmission and storage. While many compression techniques have been proposed, they fail to efficiently adapt to fluctuating network bandwidth, leading to resource wastage. We address this issue from the perspective of size-aware compression, where we aim to compress 3DGS to a desired size by quickly searching for suitable hyperparameters. Through a measurement study, we identify key hyperparameters that affect the size - namely, the reserve ratio of Gaussians and bit-width settings for Gaussian attributes. Then, we formulate this hyperparameter optimization problem as a mixed-integer nonlinear programming (MINLP) problem, with the goal of maximizing visual quality while respecting the size budget constraint. To solve the MINLP, we decouple this problem into two parts: discretely sampling the reserve ratio and determining the bit-width settings using integer linear programming (ILP). To solve the ILP more quickly and accurately, we design a quality loss estimator and a calibrated size estimator, as well as implement a CUDA kernel. Extensive experiments on multiple 3DGS variants demonstrate that our method achieves state-of-the-art performance in post-training compression. Furthermore, our method can achieve comparable quality to leading training-required methods after fine-tuning. Shuzhao Xie, Weixiang Zhang, Shijia Ge, Sicheng Pan, Yunpeng Bai, Cong Zhang 0002, Xiaoyi Fan 0001, Zhi Wang 0001 |
ACM Multimedia | 1 |
| 2025 | Understanding Bias Terms in Neural RepresentationsabstractIn this paper, we examine the impact and significance of bias terms in Implicit Neural Representations (INRs). While bias terms are known to enhance nonlinear capacity by shifting activations in typical neural networks, we discover their functionality differs markedly in neural representation networks.
Our analysis reveals that INR performance neither scales with increased number of bias terms nor shows substantial improvement through bias term gradient propagation. We demonstrate that bias terms in INRs primarily serve to eliminate \textit{spatial aliasing} caused by symmetry from both coordinates and activation functions, with input-layer bias terms yielding the most significant benefits.
These findings challenge the conventional practice of implementing full-bias INR architecture.
We propose using freezing bias terms exclusively in input layers, which consistently outperforms fully biased networks in signal fitting tasks.
Furthermore, we introduce Feature-Biased INRs~(Feat-Bias), which initialize input-layer bias with high-level features extracted from pre-trained models. This feature-biasing approach effectively addresses the limited performance in INR post-processing tasks due to neural parameter uninterpretability, achieving superior accuracy while reducing parameter count and improving reconstruction quality. Weixiang Zhang, Boxi Li, Shuzhao Xie, Chengwei Ren, Yuan Xue 0013, Zhi Wang 0001 |
NeurIPS | 3 |
| 2025 | SkyML: A MLaaS Federation Design for Multicloud-Based Multimedia AnalyticsabstractThe advent of deep learning has precipitated a surge in public machine learning as a service (MLaaS) for multimedia analysis. However, reliance on a single MLaaS can result in product dependency and a loss of better performance offered by multiple MLaaSes. Consequently, many enterprises opt for an intercloud broker capable of managing jobs across various clouds. Though existing works explore the efficient utilization of inter-cloud computational resources and the enhancement of inter-cloud data transfer throughput, they disregard improving the overall accuracy of multiple MLaaSes. In response, we conduct a measurement study on object detection services, which are designed to identify and locate various objects within an image. We discover that combining predictions from multiple MLaaSes can improve analytical performance. However, more MLaaSes do not necessarily equate to better performance. Therefore, we propose SkyML, a user-side MLaaS federation broker that selects a subset of MLaaSes based on the characteristics of the request to achieve optimal multimedia analytical performance. Initially, we design a combinatorial reinforcement learning approach to select the sound MLaaS combination, thereby maximizing user experience. We also present an ingenious, automated taxonomy unification algorithm to minimize human efforts in merging MLaaS-specific labels into a user-preferred label space. Moreover, we devise an optimized ensemble strategy to aggregate predictions from the selected MLaaSes. Evaluations indicate that our similarity-based taxonomy unification approach can reduce annotation costs by 90%. Moreover, real-world trace-driven evaluations further prove that our MLaaS selection method can achieve similar levels of accuracy with a 67% reduction in inference fees. Shuzhao Xie, Yuan Xue 0013, Yifei Zhu 0001, Zhi Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | TextIR: A Simple Framework for Text-Based Editable Image RestorationabstractMany current image restoration approaches utilize neural networks to acquire robust image-level priors from extensive datasets, aiming to reconstruct missing details. Nevertheless, these methods often falter with images that exhibit significant information gaps. While incorporating external priors or leveraging reference images can provide supplemental information, these strategies are limited in their practical scope. Alternatively, textual inputs offer greater accessibility and adaptability. In this study, we develop a sophisticated framework enabling users to guide the restoration of deteriorated images via textual descriptions. Utilizing the text-image compatibility feature of CLIP enhances the integration of textual and visual data. Our versatile framework supports multiple restoration activities such as image inpainting, super-resolution, and colorization. Comprehensive testing validates our technique's efficacy. Yunpeng Bai, Cairong Wang, Shuzhao Xie, Chao Dong 0005, Chun Yuan 0003, Zhi Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Retraining-free Model Quantization via One-Shot Weight-Coupling LearningabstractQuantization is of significance for compressing the over-parameterized deep neural models and deploying them on resource-limited devices. Fixed-precision quantization suf-fers from performance drop due to the limited numerical representation ability. Conversely, mixed-precision quan-tization (MPQ) is advocated to compress the model ef-fectively by allocating heterogeneous bit-width for layers. MPQ is typically organized into a searching-retraining two-stage process. Previous works only focus on determining the optimal bit-width configuration in the first stage effi-ciently, while ignoring the considerable time costs in the second stage and thus hindering deployment efficiency sig-nificantly. In this paper, we devise a one-shot training-searching paradigm for mixed-precision model compression. Specifically, in the first stage, all potential bit-width configurations are coupled and thus optimized simultane-ously within a set of shared weights. However, our ob-servations reveal a previously unseen and severe bit-width interference phenomenon among highly coupled weights during optimization, leading to considerable performance degradation under a high compression ratio. To tackle this problem, we first design a bit-width scheduler to dy-namically freeze the most turbulent bit-width of layers during training, to ensure the rest bit-widths converged prop-erly. Then, taking inspiration from information theory, we present an information distortion mitigation technique to align the behaviour of the bad-performing bit-widths to the well-performing ones. In the second stage, an inference-only greedy search scheme is devised to evaluate the good-ness of configurations without introducing any additional training costs. Extensive experiments on three representative models and three datasets demonstrate the effective-ness of the proposed method. Code can be available on https://github.com/1hunters/retraining-free-quantization. Shuzhao Xie, Rongwei Lu, Xinzhu Ma, Zhi Wang 0001, Wenwu Zhu 0001 |
CVPR | 4 |
| 2024 | MesonGS: Post-training Compression of 3D Gaussians via Efficient Attribute Transformation
Shuzhao Xie, Weixiang Zhang, Yunpeng Bai, Rongwei Lu, Shijia Ge, Zhi Wang 0001 |
ECCV (33) | 1 |
| 2024 | A Joint Approach to Local Updating and Gradient Compression for Efficient Asynchronous Federated Learning
Jiajun Song, Jiajun Luo, Rongwei Lu, Shuzhao Xie, Bin Chen 0011, Zhi Wang 0001 |
Euro-Par (3) | 4 |
| 2022 | Cost Effective MLaaS Federation: A Combinatorial Reinforcement Learning ApproachabstractWith the advancement of deep learning techniques, major cloud providers and niche machine learning service providers start to offer their cloud-based machine learning tools, also known as machine learning as a service (MLaaS), to the public. According to our measurement, for the same task, these MLaaSes from different providers have varying performance due to the proprietary datasets, models, etc. Federating different MLaaSes together allows us to improve the analytic performance further. However, naively aggregating results from different MLaaSes not only incurs significant momentary cost but also may lead to sub-optimal performance gain due to the introduction of possible false-positive results. In this paper, we propose Armol, a framework to federate the right selection of MLaaS providers to achieve the best possible analytic performance. We first design a word grouping algorithm to unify the output labels across different providers. We then present a deep combinatorial reinforcement learning based-approach to maximize the accuracy while minimizing the cost. The predictions from the selected providers are then aggregated together using carefully chosen ensemble strategies. The real-world trace-driven evaluation further demonstrates that Armol is able to achieve the same accuracy results with 67% less inference cost. Shuzhao Xie, Yuan Xue 0013, Yifei Zhu 0001, Zhi Wang 0001 |
INFOCOM | 1 |