Shuai Zhang 0050

dblp:71/208-50 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0008-5848-848XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 61% 3D vision · 22% Video understanding and tracking · 10%
Computer graphics and multimedia
1 paper
Rendering · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.822026
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices · AAAI 2026
MobileInst: Video Instance Segmentation on the Mobile · AAAI 2024
Machine learning › Efficient and distributed learning
inference efficiency
1.012026
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.012026
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices · AAAI 2026
Machine learning › Efficient and distributed learning › model deployment
mobile deployment
1.012026
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices · AAAI 2026
Computer vision › 3D vision › 3d reconstruction › dynamic 3d reconstruction
dynamic mesh reconstruction
0.912025
Dynamic 2D Gaussians: Geometrically Accurate Radiance Fields for Dynamic Objects · ACM Multimedia 2025
Computer vision › 3D vision › 3d reconstruction
surface reconstruction
0.912025
Dynamic 2D Gaussians: Geometrically Accurate Radiance Fields for Dynamic Objects · ACM Multimedia 2025
Rendering › gaussian splatting
2d gaussian splatting
0.912025
Dynamic 2D Gaussians: Geometrically Accurate Radiance Fields for Dynamic Objects · ACM Multimedia 2025
Rendering › neural rendering
radiance field
0.912025
Dynamic 2D Gaussians: Geometrically Accurate Radiance Fields for Dynamic Objects · ACM Multimedia 2025
Computer vision › Video understanding and tracking
video instance segmentation
0.812024
MobileInst: Video Instance Segmentation on the Mobile · AAAI 2024
Machine learning › Generative modeling
variational autoencoder
0.312026
Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices · AAAI 2026
Computer vision › Segmentation and scene understanding
instance segmentation
0.212024
MobileInst: Video Instance Segmentation on the Mobile · AAAI 2024

Methods — techniques the papers use, named apart from their topics

sparse-controlled deformation · 1.7pixel shuffle · 1.0knowledge distillation · 1.0depthwise separable convolution · 1.0temporal query passing · 0.8query-based dual-transformer decoder · 0.8kernel reuse · 0.8
YearPublicationVenuePosition
2026 Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices
abstract
There is a growing demand for deploying large generative AI models on mobile devices. For recent popular video generative models, however, the Variational AutoEncoder (VAE) represents one of the major computational bottlenecks. Both large parameter sizes and mismatched kernels cause out-of-memory errors or extremely slow inference on mobile devices. To address this, we propose a low-cost solution that efficiently transfers widely used video VAEs to mobile devices. (1) We analyze redundancy in existing VAE architectures and get empirical design insights. By integrating 3D depthwise separable convolutions into our model, we significantly reduce the number of parameters. (2) We observe that the upsampling techniques in mainstream video VAEs are poorly suited to mobile hardware and form the main bottleneck. In response, we propose a decoupled 3D pixel shuffle scheme that slashes end-to-end delay. Building upon these, we develop a universal mobile-oriented VAE decoder, Turbo-VAED. (3) We propose an efficient VAE decoder training method. Since only the decoder is used during deployment, we distill it to Turbo-VAED instead of retraining the full VAE, enabling fast mobile adaptation with minimal performance loss. To our knowledge, our method enables real-time 720p video VAE decoding on mobile devices for the first time. This approach is widely applicable to most video VAEs. When integrated into four representative models, with training cost as low as $95, it accelerates original VAEs by up to 84.5× at 720p resolution on GPUs, uses as low as 17.5% of original parameter count, and retains 96.9% of the original reconstruction quality. Compared to mobile-optimized VAEs, Turbo-VAED achieves a 2.9× speedup in FPS and better reconstruction quality on the iPhone 16 Pro.
Ya Zou, Jingfeng Yao, Shuai Zhang 0050, Wenyu Liu 0001, Xinggang Wang
AAAI4
2025 Dynamic 2D Gaussians: Geometrically Accurate Radiance Fields for Dynamic Objects
abstract
Reconstructing objects and extracting high-quality surfaces play a vital role in the real world. Current 4D representations show the ability to render high-quality novel views for dynamic objects, but cannot reconstruct high-quality meshes due to their implicit or geometrically inaccurate representations. In this paper, we propose a novel representation that can reconstruct accurate meshes from sparse image input, named Dynamic 2D Gaussians (D-2DGS). We adopt 2D Gaussians for basic geometry representation and use sparse-controlled points to capture the 2D Gaussian's deformation. By extracting the object mask from the rendered high-quality image and masking the rendered depth map, we remove floaters that are prone to occur during reconstruction and can extract high-quality dynamic mesh sequences of dynamic objects. Experiments demonstrate that our D-2DGS is outstanding in reconstructing detailed and smooth high-quality meshes from sparse inputs. The code is available at https://github.com/hustvl/Dynamic-2DGS.
Shuai Zhang 0050, Guanjun Wu, Zhoufeng Xie, Xinggang Wang, Bin Feng 0001, Wenyu Liu 0001
ACM Multimedia1
2025 TOGS: Gaussian Splatting With Temporal Opacity Offset for Real-Time 4D DSA Rendering
abstract
Four-dimensional Digital Subtraction Angiography (4D DSA) is a medical imaging technique that provides a series of 2D images captured at different stages and angles during the process of contrast agent filling blood vessels. It plays a significant role in the diagnosis of cerebrovascular diseases. Improving the rendering quality and speed under sparse sampling is important for observing the status and location of lesions. The current methods exhibit inadequate rendering quality in sparse views and suffer from slow rendering speed. To overcome these limitations, we propose TOGS, a Gaussian splatting method with opacity offset over time, which can effectively improve the rendering quality and speed of 4D DSA. We introduce an opacity offset table for each Gaussian to model the opacity offsets of the Gaussian, using these opacity-varying Gaussians to model the temporal variations in the radiance of the contrast agent. By interpolating the opacity offset table, the opacity variation of the Gaussian at different time points can be determined. This enables us to render the 2D DSA image at that specific moment. Additionally, we introduced a Smooth loss term in the loss function to mitigate overfitting issues that may arise in the model when dealing with sparse view scenarios. During the training phase, we randomly prune Gaussians, thereby reducing the storage overhead of the model. The experimental results demonstrate that compared to previous methods, this model achieves state-of-the-art render quality under the same number of training views. Additionally, it enables real-time rendering while maintaining low storage overhead.
Shuai Zhang 0050, Huangxuan Zhao, Zhenghong Zhou, Guanjun Wu, Chuansheng Zheng, Xinggang Wang, Wenyu Liu 0001
IEEE J. Biomed. Health Informatics1
2024 MobileInst: Video Instance Segmentation on the Mobile
abstract
Video instance segmentation on mobile devices is an important yet very challenging edge AI problem. It mainly suffers from (1) heavy computation and memory costs for frame-by-frame pixel-level instance perception and (2) complicated heuristics for tracking objects. To address these issues, we present MobileInst, a lightweight and mobile-friendly framework for video instance segmentation on mobile devices. Firstly, MobileInst adopts a mobile vision transformer to extract multi-level semantic features and presents an efficient query-based dual-transformer instance decoder for mask kernels and a semantic-enhanced mask decoder to generate instance segmentation per frame. Secondly, MobileInst exploits simple yet effective kernel reuse and kernel association to track objects for video instance segmentation. Further, we propose temporal query passing to enhance the tracking ability for kernels. We conduct experiments on COCO and YouTube-VIS datasets to demonstrate the superiority of MobileInst and evaluate the inference latency on one single CPU core of the Snapdragon 778G Mobile Platform, without other methods of acceleration. On the COCO dataset, MobileInst achieves 31.2 mask AP and 433 ms on the mobile CPU, which reduces the latency by 50% compared to the previous SOTA. For video instance segmentation, MobileInst achieves 35.0 AP and 30.1 AP on YouTube-VIS 2019 & 2021.
Renhong Zhang, Tianheng Cheng, Shusheng Yang, Haoyi Jiang, Shuai Zhang 0050, Jiancheng Lyu, Xin Li 0034, Xiaowen Ying, Dashan Gao 0001, Wenyu Liu 0001, Xinggang Wang
AAAI5