Yuan-Chun Sun

dblp:330/2809 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0001-6069-0002ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LMG: Efficient Streaming of Layered Mesh-Gaussian 3D Scenes
Yuan-Chun Sun, Guodong Chen 0004, Sam Ziaie Kondori, Mallesham Dasari, Cheng-Hsin Hsu
MMSys1
2025 Exploring LLM-based Assistants with Smart Glasses for the Visually Impaired
abstract
Modern smart glasses, when combined with Large Language Models (LLMs), offer a promising new paradigm for assisting visually impaired individuals in daily navigation and scene understanding. Yet, the effectiveness of such assistants depends critically on factors such as camera Angles of View (AoV), semantic extraction, network conditions, and model selection, which have not been systematically studied. To fill this gap, we construct a dataset of egocentric video sequences with multiple AoVs and systematically generated Q&A, and we implement an LLM-based assistant with edge offloading to evaluate different design choices. In particular, we considered four representative multimodal LLMs: MiniCPM-o 2.6 8B, LLaVA-OneVision 7B, Qwen2.5-VL 7B, and Qwen2.5-VL 3B, to cover a diverse range of model architectures and sizes. Through extensive experiments, we find that: (i) current LLMs are still limited in recognizing 360° videos, but semantic extractors improve accuracy by up to 33.77% with minimal impact on delay; (ii) capturing with 360° cameras raises accuracy by an average of 20.6% and achieves up to 85.64% on position-sensitive queries without additional inference time; (iii) higher-bandwidth networks such as WiFi reduce transmission delay, resulting in acceptable response time (sub 1-second); and (iv) different LLMs perform better on different question categories, highlighting the need for careful modeling and offloading strategies in future work.
Zhe-Yu Lee, Yuan-Chun Sun, Yee-Nam Wong, Cheng-Hsin Hsu
SEC2
2025 EyeNavGS: A 6-DoF Navigation Dataset and Record-n-Replay Software for Real-World 3DGS Scenes in VR
abstract
3D Gaussian Splatting (3DGS) is an emerging media representation that reconstructs real-world 3D scenes in high fidelity, enabling 6-degrees-of-freedom (6-DoF) navigation in virtual reality (VR). However, developing and evaluating 3DGS-enabled applications and optimizing their rendering performance require realistic user navigation data. Such data is currently unavailable for photorealistic 3DGS reconstructions of real-world scenes. This paper introduces EyeNavGS, the first publicly available 6-DoF navigation dataset featuring traces from 46 participants exploring twelve diverse, real-world 3DGS scenes. The dataset was collected at two sites, using the Meta Quest Pro headsets, recording the head pose and eye gaze data for each rendered frame during free world standing 6-DoF navigation. For each of the twelve scenes, we performed careful scene initialization to correct for scene tilt and scale, ensuring a perceptually-comfortable VR experience. We also release our open-source SIBR viewer software fork with record-and-replay functionalities and a suite of utility tools for data processing, conversion, and visualization. The EyeNavGS dataset and its accompanying software tools provide valuable resources for advancing research in 6-DoF viewport prediction, adaptive streaming, 3D saliency, and foveated rendering for 3DGS scenes. The EyeNavGS dataset is available at: https://symmru.github.io/EyeNavGS/
Cheng-Tse Lee, Mufeng Zhu, Yuan-Chun Sun, Cheng-Hsin Hsu, Yao Liu 0001
ACM Multimedia5
2025 Streaming 3DGS Virtual Worlds in 6DoF over Next-Generation Networks
abstract
With the continuous development of networking and computing devices, immersive communication has become increasingly viable, enabling users to explore virtual worlds and interact with other users in 6 Degrees-of-Freedom (6DoF). Immersive communication has great potential not only in professional domains, such as medical diagnostics and distance education but also for leisure activities, such as social networks and new media. This proposal aims to develop an immersive communication system leveraging the cutting-edge dynamic 3D Gaussians Splatting (3DGS). Our objective is to design, implement, and evaluate a highly interactive, adaptive, and efficient immersive communication system that maximizes user experience and system performance while supporting heterogeneous hardware platforms, networks, and applications. We identify and tackle three critical challenges in developing such a system: (i) designing deformable 3DGS avatars to enhance real-time user representation, (ii) developing scalable dynamic 3DGS codecs to optimize data transmission and storage efficiency, and (iii) implementing adaptive streaming algorithms to ensure smooth and responsive user experience across diverse usage scenarios. By solving these challenges, our research aims to push the boundaries of real-time 3D streaming and interaction while redefining the future of virtual world over next-generation networks.
Yuan-Chun Sun
ACM Multimedia1
2025 Optimally Planning Drone Trajectories to Capture 3D Gaussian Splatting Objects
Cheng-Yuan Wu, Yuan-Chun Sun, Cheng-Tse Lee, Cheng-Hsin Hsu
MMM (3)2
2025 LTS: A DASH Streaming System for Dynamic Multi-Layer 3D Gaussian Splatting Scenes
abstract
We present a novel DASH-based streaming system for dynamic 3D Gaussian Splatting (3DGS) scenes, addressing the challenges of streaming large amounts of 3DGS data over diverse and dynamic networks. Our Layer, Tile, and Segment Adaptive streaming (LTS) system combines three key features: (i) multi-layer streaming, which adapts to diverse client capabilities while balancing visual quality and bandwidth usage, (ii) tiled streaming, which reduces unnecessary data transmission by focusing on the user's viewport, and (iii) segment streaming, which divides dynamic 3DGS scenes into segments, letting clients request them dynamically to handle network fluctuations. Our experimental results demonstrate that our LTS system achieves superior performance in both live and on-demand streaming of dynamic 3DGS scenes compared to the baselines. For example, in live streaming, LTS could achieve up to 99.70% reduction in missing frames on average and deliver a maximum PSNR (Peak Signal-to-Noise Ratio) improvement of 10.08 dB. In on-demand streaming, LTS could reduce the freeze time by up to 92.01%, and increase the synthesized view quality by up to 5.14 dB in PSNR and 0.11 in SSIM (Structural Similarity Index). Our source codes are available at: https://github.com/AIINS-NTHU/LTS-DASH-Streaming-System-for-3DGS.
Yuan-Chun Sun, Yuang Shi, Cheng-Tse Lee, Mufeng Zhu, Wei Tsang Ooi, Yao Liu 0001, Chun-Ying Huang, Cheng-Hsin Hsu
MMSys1
2025 Composing Error Concealment Pipelines for Dynamic 3D Point Cloud Streaming
abstract
Dynamic 3D point clouds enable the immersive user experience and thus have become increasingly more popular in volumetric video streaming applications. When being streamed over best-effort networks, point cloud frames may suffer from lost or late packets, leading to non-trivial quality degradation. To solve this problem, we proposed the very first error concealment pipeline framework, which comprises five stages: pre-processing, matching, motion estimation, prediction, and post-processing. Alternative algorithms can be developed for each stage, while algorithms of different stages could be mixed and matched into pipelines for end-to-end performance evaluations. We discussed the design goal and proposed multiple algorithms for each stage. These algorithms were then quantitatively compared using dynamic 3D point cloud sequences with diverse characteristics. Based on the comparison results, we proposed four representative pipelines for: (i) diverse degrees of motion variance, i.e., minor versus significant, and (ii) different application requirements, i.e., high quality versus low overhead. Extensive end-to-end evaluations of our proposed pipelines demonstrated their superior concealed quality over the 3D frame-copy method in both: (i) 3D metrics, by up to 5.32 dB in GPSNR and 1.7 dB in CPSNR,and (ii) 2D metrics, by up to 2.22 dB in PSNR, 0.06 in SSIM, and 11.67 in VMAF. Adding to that, a user study with 15 subjects indicated that our best-performing pipeline achieved 100% preference winning rate over the state-of-the-art learning-based interpolation algorithms while consuming merely up to 8.55% of running time.
I-Chun Huang, Yuang Shi, Yuan-Chun Sun, Wei Tsang Ooi, Chun-Ying Huang, Cheng-Hsin Hsu
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Dynamic 6-DoF Volumetric Video Generation: Software Toolkit and Dataset
abstract
Volumetric video streaming has become increasingly popular in recent years due to its support of 6 degrees-of-freedom (6-DoF) exploration. There is, however, a shortage of dynamic 6-DoF content suitable for comparing the performance among heterogeneous volumetric video representations. This paper introduces a software toolkit for creating both a dataset of dynamic 6-DoF content in point clouds and a dataset for training and testing neural-based representations such as neural radiance fields (NeRF). Starting with freely available 3D assets online, our software toolkit uses the Blender Python API to generate training and testing datasets for neural-based dynamic volumetric model training. The created datasets are compliant with existing neural-based model training and rendering frameworks. The software can also construct point cloud sequences derived from synthetic dynamic 3D meshes. This further facilitates comparing point clouds and neural-based methods for volumetric video representation. We release the software toolkit along with a rich set of sequence datasets generated in compliance with the permissions granted by the original 3D asset creators. With our toolkit and dataset, we aim to facilitate research from the multimedia systems community to support practical volumetric streaming. Our software toolkit and dataset are available at: https://6-dof-dynamic-content-software.github.io/.
Mufeng Zhu, Yuan-Chun Sun, Na Li 0032, Jin Zhou 0006, Songqing Chen, Cheng-Hsin Hsu, Yao Liu 0001
MMSP2
2023 A Blind Streaming System for Multi-client Online 6-DoF View Touring
abstract
Online 6-DoF view touring has become increasingly popular due to hardware advances and the recent pandemic. One way for content creators to support many 6-DoF clients is by transmitting 3D content to them, which leads to content leakage. Another way for content creators is to render and stream novel views for 6-DoF clients, which incurs staggering computational and networking workloads. In this paper, we develop a blind streaming system that leverages cloud service providers between content creators and 6-DoF clients. Our system has two core design objectives: (i) to generate high-quality novel views for 6-DoF clients without retrieving 3D content from content creators, (ii) to support many 6-DoF clients without overloading the content creators. We achieve these two goals in the following steps. First, we design a source view request/response interface between cloud service providers and content creators for efficient communications. Second, we design novel view optimization algorithms for cloud service providers to intelligently select the minimal set of source views while considering the workload of content creators. Third, we employ scalable client side view synthesis for 6-DoF clients with heterogeneous device capabilities and personalized 6-DoF client poses and preferences. Our evaluation results demonstrate the merits of our solution, compared to the state-of-the-arts, our system: (i) improves synthesized novel views by 2.27 dB in PSNR and 12 in VMAF on average and (ii) reduces the bandwidth consumption by 94% on average. In fact, our solution approaches the performance of an unrealistic optimal solution with unlimited source views, achieving performance gaps as small as 0.75 dB in PSNR and 3.8 in VMAF.
Sheng-Ming Tang, Yuan-Chun Sun, Cheng-Hsin Hsu
ACM Multimedia2
2023 A Dynamic 3D Point Cloud Dataset for Immersive Applications
abstract
Motion estimation in a 3D point cloud sequence is a fundamental operation with many applications, including compression, error concealment, and temporal upscaling. While there have been multiple research contributions toward estimating the motion vector of points between frames, there is a lack of a dynamic 3D point cloud dataset with motion ground truth to benchmark against. In this paper, we present an open dynamic 3D point cloud dataset to fill this gap. Our dataset consists of synthetically generated objects with pre-determined motion patterns, allowing us to generate the motion vectors for the points. Our dataset contains nine objects in three categories (shape, avatar, and textile) with different animation patterns. We also provide semantic segmentation of each avatar object in the dataset. Our dataset can be used by researchers who need temporal information across frames. As an example, we present an evaluation of two motion estimation methods using our dataset.
Yuan-Chun Sun, I-Chun Huang, Yuang Shi, Wei Tsang Ooi, Chun-Ying Huang, Cheng-Hsin Hsu
MMSys1