EDBT 2026 Demo / reviewers in the wild / expert
Yuang Shi
dblp:266/7792
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-7893-8512ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoonAnything: A Vision Benchmark with Large-Scale Lunar Supervised DataabstractAccurate perception of lunar surfaces is critical for modern lunar exploration missions. However, developing robust learning-based perception systems is hindered by the lack of datasets that provide both geometric and photometric supervision. Existing lunar datasets typically lack either geometric ground truth, photometric realism, illumination diversity, or large-scale coverage. In this paper, we introduce MoonAnything, a unified benchmark built on real lunar topography with physically-based rendering, providing the first comprehensive geometric and photometric supervision under diverse illumination with large scale. The benchmark comprises two complementary sub-datasets : i) LunarGeo provides stereo images with corresponding dense depth maps and camera calibration enabling 3D reconstruction and pose estimation; ii) LunarPhoto provides photorealistic images using a spatially-varying BRDF model, along with multi-illumination renderings under real solar configurations, enabling reflectance estimation and illumination-robust perception. Together, these datasets offer over 130K samples with comprehensive supervision. Beyond lunar applications, MoonAnything offers a unique setting and challenging testbedfor algorithms under low-textured, high-contrast conditions and applies to other airless celestial bodies and could generalize beyond. We establish baselines using state-of-the-art methods and release the complete dataset along with generation tools to support community extension: https://github.com/clementinegrethen/MoonAnything. Clementine Grethen, Yuang Shi, Simone Gasparini, Géraldine Morin |
MMSys | 2 |
| 2026 | P-GSVC: Layered Progressive 2D Gaussian Splatting for Scalable Image and VideoabstractGaussian splatting has emerged as a competitive explicit representation for image and video reconstruction. In this work, we present P-GSVC, the first layered progressive 2D Gaussian splatting framework that provides a unified solution for scalable Gaussian representation in both images and videos. P-GSVC organizes 2D Gaussian splats into a base layer and successive enhancement layers, enabling coarse-to-fine reconstructions. To effectively optimize this layered representation, we propose a joint training strategy that simultaneously updates Gaussians across layers, aligning their optimization trajectories to ensure inter-layer compatibility and a stable progressive reconstruction. P-GSVC supports scalability in terms of both quality and resolution. Our experiments show that the joint training strategy can gain up to 1.9 dB improvement in PSNR for video and 2.6 dB improvement in PSNR for image when compared to methods that perform sequential layer-wise training. Longan Wang, Yuang Shi, Wei Tsang Ooi |
MMSys | 2 |
| 2025 | LapisGS: Layered Progressive 3D Gaussian Splatting for Adaptive StreamingabstractThe rise of Extended Reality$(X R)$requires efficient streaming of 3D online worlds, challenging current 3DGS representations to adapt to bandwidth-constrained environments. This paper proposes LapisGS, a layered 3DGS that supports adaptive streaming and progressive rendering. Our method constructs a layered structure for cumulative representation, incorporates dynamic opacity optimization to maintain visual fidelity, and utilizes occupancy maps to efficiently manage Gaussian splats. This proposed model offers a progressive representation supporting a continuous rendering quality adapted for bandwidth-aware streaming. Extensive experiments validate the effectiveness of our approach in balancing visual fidelity with the compactness of the model, with up to 50.71 % improvement in SSIM, 286.53% improvement in LPIPS with 23% of the original model size, and shows its potential for bandwidth-adapted 3D streaming and rendering applications. Project page: https://yuang-ian.github.io/lapisgs/ Yuang Shi, Géraldine Morin, Simone Gasparini, Wei Tsang Ooi |
3DV | 1 |
| 2025 | 3D Gaussian-based Immersive Media Streaming in Networked Extended Realityabstract3D Gaussian Splatting (3DGS) has emerged as a revolutionary representation for immersive media, offering unprecedented visual quality and efficient rendering. However, streaming 3DGS content presents unique challenges due to its complex parameter space and tight coupling between primitives. This doctoral research aims to address these challenges through three main themes: (i) developing scalable 3DGS representations that enable progressive transmission and flexible rendering, (ii) creating comprehensive quality assessment frameworks that consider both spatial-temporal coherence and multi-attribute distortions, and (iii) designing robust error control strategies for real-time streaming applications. Our research ambitions to advance the state-of-the-art in immersive media representation and delivery, with potential applications in virtual reality, telepresence, and digital twin systems. Yuang Shi |
MMSys | 1 |
| 2025 | LTS: A DASH Streaming System for Dynamic Multi-Layer 3D Gaussian Splatting ScenesabstractWe present a novel DASH-based streaming system for dynamic 3D Gaussian Splatting (3DGS) scenes, addressing the challenges of streaming large amounts of 3DGS data over diverse and dynamic networks. Our Layer, Tile, and Segment Adaptive streaming (LTS) system combines three key features: (i) multi-layer streaming, which adapts to diverse client capabilities while balancing visual quality and bandwidth usage, (ii) tiled streaming, which reduces unnecessary data transmission by focusing on the user's viewport, and (iii) segment streaming, which divides dynamic 3DGS scenes into segments, letting clients request them dynamically to handle network fluctuations. Our experimental results demonstrate that our LTS system achieves superior performance in both live and on-demand streaming of dynamic 3DGS scenes compared to the baselines. For example, in live streaming, LTS could achieve up to 99.70% reduction in missing frames on average and deliver a maximum PSNR (Peak Signal-to-Noise Ratio) improvement of 10.08 dB. In on-demand streaming, LTS could reduce the freeze time by up to 92.01%, and increase the synthesized view quality by up to 5.14 dB in PSNR and 0.11 in SSIM (Structural Similarity Index). Our source codes are available at: https://github.com/AIINS-NTHU/LTS-DASH-Streaming-System-for-3DGS. Yuan-Chun Sun, Yuang Shi, Cheng-Tse Lee, Mufeng Zhu, Wei Tsang Ooi, Yao Liu 0001, Chun-Ying Huang, Cheng-Hsin Hsu |
MMSys | 2 |
| 2025 | GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splattingabstract3D Gaussian splats have emerged as a revolutionary, effective, learned representation for static 3D scenes. This work explores using 2D Gaussian splats as a new primitive for representing videos. We propose GSVC, an approach to learning a set of 2D Gaussian splats that can effectively represent and compress video frames. GSVC incorporates the following techniques: (i) To exploit temporal redundancy among adjacent frames, which can speed up training and improve the compression efficiency, we predict the Gaussian splats of a frame based on its previous frame; (ii) To control the trade-offs between file size and quality, we remove Gaussian splats with low contribution to the video quality; (iii) To capture dynamics in videos, we randomly add Gaussian splats to fit content with large motion or newly-appeared objects; (iv) To handle significant changes in the scene, we detect key frames based on loss differences during the learning process. Experiment results show that GSVC achieves good rate-distortion trade-offs, comparable to state-of-the-art video codecs such as AV1 and VVC, and a rendering speed of 1500 fps for a 1920×1080 video. Project page: https://yuang-ian.github.io/gsvc/. Longan Wang, Yuang Shi, Wei Tsang Ooi |
NOSSDAV | 2 |
| 2025 | Composing Error Concealment Pipelines for Dynamic 3D Point Cloud StreamingabstractDynamic 3D point clouds enable the immersive user experience and thus have become increasingly more popular in volumetric video streaming applications. When being streamed over best-effort networks, point cloud frames may suffer from lost or late packets, leading to non-trivial quality degradation. To solve this problem, we proposed the very first error concealment pipeline framework, which comprises five stages: pre-processing, matching, motion estimation, prediction, and post-processing. Alternative algorithms can be developed for each stage, while algorithms of different stages could be mixed and matched into pipelines for end-to-end performance evaluations. We discussed the design goal and proposed multiple algorithms for each stage. These algorithms were then quantitatively compared using dynamic 3D point cloud sequences with diverse characteristics. Based on the comparison results, we proposed four representative pipelines for: (i) diverse degrees of motion variance, i.e., minor versus significant, and (ii) different application requirements, i.e., high quality versus low overhead. Extensive end-to-end evaluations of our proposed pipelines demonstrated their superior concealed quality over the 3D frame-copy method in both: (i) 3D metrics, by up to 5.32 dB in GPSNR and 1.7 dB in CPSNR,and (ii) 2D metrics, by up to 2.22 dB in PSNR, 0.06 in SSIM, and 11.67 in VMAF. Adding to that, a user study with 15 subjects indicated that our best-performing pipeline achieved 100% preference winning rate over the state-of-the-art learning-based interpolation algorithms while consuming merely up to 8.55% of running time. I-Chun Huang, Yuang Shi, Yuan-Chun Sun, Wei Tsang Ooi, Chun-Ying Huang, Cheng-Hsin Hsu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | QV4: QoE-based Viewpoint-Aware V-PCC-encoded Volumetric Video StreamingabstractVolumetric videos allow six degrees of freedom (6DoF) movement for viewers, enabling numerous applications in domains such as entertainment, healthcare, and education. MPEG's Video-based Point Cloud Compression (V-PCC) is a recent new standard for volumetric video compression that achieves a considerable compression rate while maintaining the quality of the point cloud sequence. However, V-PCC is hard to fit into existing tiling-based volumetric video streaming framework due to the lack of proper user viewing adaptive techniques. In this paper, we propose QV4, a Quality-of-Experience (QoE) based streaming pipeline for viewpoint-aware V-PCC-encoded volumetric video. Specifically, we leverage the intermediate results produced by the V-PCC encoder to achieve effective and efficient viewpoint-aware tiling for V-PCC. We then build a QoE model and a 6DoF movement model based on real-world user data, to predict the users' viewing experience and behaviors, respectively. The proposed QoE model and 6DoF movement model are combined with viewpoint-aware V-PCC tiling to maximize the visual quality of volumetric videos. Extensive simulations show that by enabling viewpoint-aware adaptation and optimization for V-PCC-encoded volumetric videos, QV4 can achieve up to 14.67% improvement in structural similarity index (SSIM) and 7.39% improvement in video multi-method assessment fusion (VMAF) over highly dynamic viewing behaviors in a network with limited and fluctuating bandwidth. Yuang Shi, Bennett Clement, Wei Tsang Ooi |
MMSys | 1 |
| 2023 | Enabling Low Bit-Rate MPEG V-PCC-encoded Volumetric Video Streaming with 3D Sub-samplingabstractMPEG's Video-based Point Cloud Compression (V-PCC) is a recent new standard for volumetric video compression. By mapping a 3D dynamic point cloud to a 2D image sequence, V-PCC can rely on state-of-the-art video codecs to achieve high compression rate while maintaining the visual fidelity of the point cloud sequence. The quality of a compressed point cloud degrades steeply, however, below the operational bit-rate range of the video codec. In this work, we show that redundant information inherent in a 3D point cloud can be exploited to further extend the bit-rate range of the V-PCC codec, enabling it to operate in a low bit-rate scenario that is important in the context of volumetric video streaming. By simplifying the 3D point clouds through down-sampling and down-scaling during the encoding phase, and reversing the process during the decoding phase, we show that V-PCC could achieve up to 2.1 dB improvement in peak signal-to-noise ratio (PSNR), 7.1% improvement in structural similarity index (SSIM) and 14.8 improvement in video multimethod assessment fusion (VMAF) of the rendered point clouds at the same bit-rate and correspondingly up to 48.5% lower bit-rate at the same image quality. Yuang Shi, Pranav Venkatram, Wei Tsang Ooi |
MMSys | 1 |
| 2023 | A Dynamic 3D Point Cloud Dataset for Immersive ApplicationsabstractMotion estimation in a 3D point cloud sequence is a fundamental operation with many applications, including compression, error concealment, and temporal upscaling. While there have been multiple research contributions toward estimating the motion vector of points between frames, there is a lack of a dynamic 3D point cloud dataset with motion ground truth to benchmark against. In this paper, we present an open dynamic 3D point cloud dataset to fill this gap. Our dataset consists of synthetically generated objects with pre-determined motion patterns, allowing us to generate the motion vectors for the points. Our dataset contains nine objects in three categories (shape, avatar, and textile) with different animation patterns. We also provide semantic segmentation of each avatar object in the dataset. Our dataset can be used by researchers who need temporal information across frames. As an example, we present an evaluation of two motion estimation methods using our dataset. Yuan-Chun Sun, I-Chun Huang, Yuang Shi, Wei Tsang Ooi, Chun-Ying Huang, Cheng-Hsin Hsu |
MMSys | 3 |
| 2023 | Uncertainty-weighted and relation-driven consistency training for semi-supervised head-and-neck tumor segmentation
Yuang Shi, Chen Zu, Pinli Yang, Hongping Ren, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015 |
Knowl. Based Syst. | 1 |
| 2022 | ASMFS: Adaptive-similarity-based multi-modality feature selection for classification of Alzheimer's disease
Yuang Shi, Chen Zu, Luping Zhou, Lei Wang 0001, Xi Wu 0004, Jiliu Zhou, Daoqiang Zhang, Yan Wang 0015 |
Pattern Recognit. | 1 |