Shijia Ge

dblp:387/4067 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0003-6799-496XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Mmrsg-unet: integrating multi-scale Mamba and reverse semantic gating for medical image segmentation
Chuanhao Yang, Shijia Ge, Yanfeng Cao, Dunke Lu
Multim. Syst.3
2025 Enhancing Implicit Neural Representations via Symmetric Power Transformation
abstract
We propose symmetric power transformation to enhance the capacity of Implicit Neural Representation (INR) from the perspective of data transformation. Unlike prior work utilizing random permutation or index rearrangement, our method features a reversible operation that does not require additional storage consumption. Specifically, we first investigate the characteristics of data that can benefit the training of INR, proposing the Range-Defined Symmetric Hypothesis, which posits that specific range and symmetry can improve the expressive ability of INR. Based on this hypothesis, we propose a nonlinear symmetric power transformation to achieve both range-defined and symmetric properties simultaneously. We use the power coefficient to redistribute data to approximate symmetry within the target range. To improve the robustness of the transformation, we further design deviation-aware calibration and adaptive soft boundary to address issues of extreme deviation boosting and continuity breaking. Extensive experiments are conducted to verify the performance of the proposed method, demonstrating that our transformation can reliably improve INR compared with other data transformations. We also conduct 1D audio, 2D image and 3D video fitting tasks to demonstrate the effectiveness and applicability of our method.
Weixiang Zhang, Shuzhao Xie, Chengwei Ren, Shijia Ge, Mingzi Wang
AAAI4
2025 EVOS: Efficient Implicit Neural Training via EVOlutionary Selector
abstract
We propose EVOlutionary Selector (EVOS), an efficient training paradigm for accelerating Implicit Neural Representation (INR). Unlike conventional INR training that feeds all samples through the neural network in each iteration, our approach restricts training to strategically selected points, reducing computational overhead by eliminating redundant forward passes. Specifically, we treat each sample as an individual in an evolutionary process, where only those fittest ones survive and merit inclusion in training, adaptively evolving with the neural network dynamics. While this is conceptually similar to Evolutionary Algorithms, their distinct objectives (selection for acceleration vs. iterative solution optimization) require a fundamental redefinition of evolutionary mechanisms for our context. In response, we design sparse fitness evaluation, frequency-guided crossover, and augmented unbiased mutation to comprise EVOS. These components respectively guide sample selection with reduced computational cost, enhance performance through frequency-domain balance, and mitigate selection bias from cached evaluation. Extensive experiments demonstrate that our method achieves approximately 48%-66% reduction in training time while ensuring superior convergence without additional cost, establishing state-of-the-art acceleration among recent sampling-based strategies. Our code is available at this link.
Weixiang Zhang, Shuzhao Xie, Chengwei Ren, Siyi Xie, Shijia Ge, Mingzi Wang
CVPR6
2025 Lungmix: A Mixup-Based Strategy for Generalization in Respiratory Sound Classification
abstract
Respiratory sound classification plays a pivotal role in diagnosing respiratory diseases. While deep learning models have succeeded with various respiratory sound datasets, our experiments indicate that models trained on one dataset often fail to generalize effectively to others, mainly due to data collection and annotation inconsistencies. To address this limitation, we introduce Lungmix, a novel data augmentation technique inspired by Mixup. Lungmix generates augmented data by blending waveforms using loudness while interpolating labels based on their semantic meaning, helping the model learn more generalized representations. Extensive evaluations across three datasets (ICBHI, SPR, and HF) demonstrate that Lungmix significantly enhances model generalization to unseen data. In particular, Lungmix boosts the 4-class classification score by up to 3.55%, attaining performance comparable to models trained on the target dataset directly.
Shijia Ge, Weixiang Zhang, Shuzhao Xie, Baixu Yan
ICASSP1
2025 PulmoScan: A Practical Pulmonary Disease Pre-Screening System
abstract
Automation of pulmonary disease identification has been a long-standing area of research and gained increased attention after the COVID-19 pandemic. However, existing respiratory sound classification algorithms exhibit significant limitations, including suboptimal performance, insufficient input robustness, and inadequate alignment with clinical evaluation metrics, thereby hindering their practical implementation. To address these limitations, we introduce PulmoScan, a practical pulmonary disease pre-screening system. PulmoScan comprises three fundamental modules: a Respiratory Sound Quality Validator that ensures the robustness of input data, a Runtime Decision Booster that improves performance and adapting to variating evaluation metrics, and a Symptom Enhancement Diagnoser that augments respiratory sound classification with comprehensive disease pre-screening capabilities. Beyond its primary function, PulmoScan exemplifies a methodological framework for translating theoretically limited algorithms into viable clinical applications, demonstrating essential considerations and procedural adaptations for real-world implementation.
Baixu Yan, Shijia Ge, Meizi Lu, Weixiang Zhang, Shuzhao Xie
ICASSP2
2025 Expansive Supervision for Neural Radiance Fields
abstract
Neural Radiance Field (NeRF) has achieved remarkable success in creating immersive media representations through its exceptional reconstruction capabilities. However, the computational demands of dense forward passes and volume rendering during training continue to challenge its real-world applications. In this paper, we introduce Expansive Supervision to reduce time and memory costs during NeRF training from the perspective of partial ray selection for supervision. Specifically, we observe that training errors exhibit a long-tail distribution correlated with image content. Based on this observation, our method selectively renders a small but crucial subset of pixels and expands their values to estimate errors across the entire area for each iteration. Compared to conventional supervision, our approach effectively bypasses redundant rendering processes, resulting in substantial reductions in both time and memory consumption. Experimental results demonstrate that integrating Expansive Supervision within existing state-of-the-art acceleration frameworks achieves 52% memory savings and 16% time savings while maintaining comparable visual quality. Our code is available at this link.
Weixiang Zhang, Shuzhao Xie, Shijia Ge, Zhi Wang 0001
ICME4
2025 SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer Programming
abstract
Recent advances in 3D Gaussian Splatting (3DGS) have greatly improved 3D reconstruction. However, its substantial data size poses a significant challenge for transmission and storage. While many compression techniques have been proposed, they fail to efficiently adapt to fluctuating network bandwidth, leading to resource wastage. We address this issue from the perspective of size-aware compression, where we aim to compress 3DGS to a desired size by quickly searching for suitable hyperparameters. Through a measurement study, we identify key hyperparameters that affect the size - namely, the reserve ratio of Gaussians and bit-width settings for Gaussian attributes. Then, we formulate this hyperparameter optimization problem as a mixed-integer nonlinear programming (MINLP) problem, with the goal of maximizing visual quality while respecting the size budget constraint. To solve the MINLP, we decouple this problem into two parts: discretely sampling the reserve ratio and determining the bit-width settings using integer linear programming (ILP). To solve the ILP more quickly and accurately, we design a quality loss estimator and a calibrated size estimator, as well as implement a CUDA kernel. Extensive experiments on multiple 3DGS variants demonstrate that our method achieves state-of-the-art performance in post-training compression. Furthermore, our method can achieve comparable quality to leading training-required methods after fine-tuning.
Shuzhao Xie, Weixiang Zhang, Shijia Ge, Sicheng Pan, Yunpeng Bai, Cong Zhang 0002, Xiaoyi Fan 0001, Zhi Wang 0001
ACM Multimedia4
2024 MesonGS: Post-training Compression of 3D Gaussians via Efficient Attribute Transformation
Shuzhao Xie, Weixiang Zhang, Yunpeng Bai, Rongwei Lu, Shijia Ge, Zhi Wang 0001
ECCV (33)6