Na Li 0032

dblp:18/3173-32 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-3533-6247ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Computer networks · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 EVASR: Edge-Based Salience-Aware Super-Resolution for Enhanced Video Quality and Power Efficiency
abstract
With the rapid growth of video content consumption, it is important to deliver high-quality streaming videos to users even under limited available network bandwidth. In this article, we propose EVASR, a system that performs edge-based video delivery to clients with salience-aware super-resolution. We select patches with higher saliency score to perform super-resolution while applying the simple yet efficient bicubic interpolation for the remaining patches in the same video frame. To efficiently use the computation resources available at the edge server, we introduce a new metric called “saliency visual quality” (SVQ) and formulate patch selection as an optimization problem to achieve the best performance when an edge server is serving multiple users. We implement EVASR based on the FFmpeg framework and deploy it on three different platforms including desktop/laptop computers, mobile phones, and single board computers (SBCs). We conduct extensive experiments for evaluating the visual quality, super-resolution speed, and power savings that can be achieved by EVASR. Results show that EVASR outperforms baseline approaches in both resource efficiency and visual quality metrics including PSNR, SVQ, and VMAF. EVASR can also achieve substantial energy savings compared to baseline approaches MobileSR and JetsonSR on mobile devices.
Na Li 0032, Sheng Wei 0001, Yao Liu 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Dynamic 6-DoF Volumetric Video Generation: Software Toolkit and Dataset
abstract
Volumetric video streaming has become increasingly popular in recent years due to its support of 6 degrees-of-freedom (6-DoF) exploration. There is, however, a shortage of dynamic 6-DoF content suitable for comparing the performance among heterogeneous volumetric video representations. This paper introduces a software toolkit for creating both a dataset of dynamic 6-DoF content in point clouds and a dataset for training and testing neural-based representations such as neural radiance fields (NeRF). Starting with freely available 3D assets online, our software toolkit uses the Blender Python API to generate training and testing datasets for neural-based dynamic volumetric model training. The created datasets are compliant with existing neural-based model training and rendering frameworks. The software can also construct point cloud sequences derived from synthetic dynamic 3D meshes. This further facilitates comparing point clouds and neural-based methods for volumetric video representation. We release the software toolkit along with a rich set of sequence datasets generated in compliance with the permissions granted by the original 3D asset creators. With our toolkit and dataset, we aim to facilitate research from the multimedia systems community to support practical volumetric streaming. Our software toolkit and dataset are available at: https://6-dof-dynamic-content-software.github.io/.
Mufeng Zhu, Yuan-Chun Sun, Na Li 0032, Jin Zhou 0006, Songqing Chen, Cheng-Hsin Hsu, Yao Liu 0001
MMSP3
2024 VertexShuffle-Based Spherical Super-Resolution for 360-Degree Videos
abstract
360-degree video is an emerging form of media that encodes information about all directions surrounding a camera, offering an immersive experience to the users. Unlike traditional 2D videos, visual information in 360-degree videos can be naturally represented as pixels on a sphere. Inspired by state-of-the-art deep-learning-based 2D image super-resolution models and spherical CNNs, in this article, we design a novel spherical super-resolution (SSR) approach for 360-degree videos. To support viewport-adaptive and bandwidth-efficient transmission/streaming of 360-degree video data and save computation, we propose the Focused Icosahedral Mesh to represent a small area on the sphere. We further construct matrices to rotate spherical content over the entire sphere to the focused mesh area, allowing us to use the focused mesh to represent any area on the sphere. Motivated by the PixelShuffle operation for 2D super-resolution, we also propose a novel VertexShuffle operation on the mesh and an improved version VertexShuffle_V2. We compare our SSR approach with state-of-the-art 2D super-resolution models and show that SSR has the potential to achieve significant benefits when applied to spherical signals.
Na Li 0032, Yao Liu 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 VQBA: Visual-Quality-Driven Bit Allocation for Low-Latency Point Cloud Streaming
abstract
Video-based Point Cloud Compression (V-PCC) is an emerging standard for encoding dynamic point cloud data. With V-PCC, point cloud data is segmented, projected, and packed on to 2D video frames, which can be compressed using existing video coding standards such as H.264, H.265 and AV1. This makes it possible to support point cloud streaming via reliable video transmission systems. On the other hand, despite recent advances, many issues still remain and prevent V-PCC from being used in low-latency point cloud streaming. For instance, point cloud registration and patch generation can take a long time.
Shuoqian Wang, Mufeng Zhu, Na Li 0032, Mengbai Xiao, Yao Liu 0001
ACM Multimedia3
2023 EVASR: Edge-Based Video Delivery with Salience-Aware Super-Resolution
abstract
With the rapid growth of video content consumption, it is important to deliver high-quality streaming videos to users even under limited available network bandwidth. In this paper, we propose EVASR, a system that performs edge-based video delivery to clients with salience-aware super-resolution. We select patches with higher saliency score to perform super-resolution while applying the simple yet efficient bicubic interpolation for the remaining patches in the same video frame. To efficiently use the computation resources available at the edge server, we introduce a new metric called "saliency visual quality" and formulate patch selection as an optimization problem to achieve the best performance when an edge server is serving multiple users. We implement EVASR based on the FFmpeg framework and conduct extensive experiments for evaluation. Results show that EVASR outperforms baseline approaches in both resource efficiency and visual quality metrics including PSNR, saliency visual quality (SVQ), and VMAF.
Na Li 0032, Yao Liu 0001
MMSys1
2022 FFmpegSR: A General Framework Toward Real-Time 4K Super-Resolution
abstract
With the explosive growth of online video content, the demand for high-quality video is ever-rising. To take advantage of recent advances in deep learning, in this paper, we propose and implement a framework, FFmpegSR, that applies deep learning-based super-resolution into an FFmpeg filter to implement real-time 4K video super-resolution. FFmpegSR applies super-resolution to the Y channel only, allowing reduced inference time while maintaining good inference quality. To further improve the inference speed, we also develop a patch-based solution that uses saliency detection to select regions of interest on the video frame. This allows us to achieve faster inference on key patches only instead of full video frames. We used videos from a public dataset for evaluation. Results show that FFmpegSR can achieve real-time super-resolution to 4K with high visual quality.
Na Li 0032, Yao Liu 0001
ISM1
2022 Exploring Spherical Autoencoder for Spherical Video Content Processing
abstract
3D spherical content is increasingly presented in various applications (e.g., AR/MR/VR) for better users' immersiveness experience, yet today processing such spherical 3D content still mainly relies on the traditional 2D approaches after projection, leading to the distortion and/or loss of critical information. This study sets to explore methods to process spherical 3D content directly and more effectively. Using 360-degree videos as an example, we propose a novel approach called Spherical Autoencoder (SAE) for spherical video processing. Instead of projecting to a 2D space, SAE represents the 360-degree video content as a spherical object and employs encoding and decoding on the 360-degree video directly. Furthermore, to support the adoption of SAE on pervasive mobile devices that often have resource constraints, we further propose two optimizations on top of SAE.First, since the FoV (Field of View) prediction is widely studied and leveraged to transport only a portion of the content to the mobile device to save bandwidth and battery consumption, we design p-SAE, a SAE scheme with the partial view support that can utilize such FoV prediction. Second, since machine learning models are often compressed when running on mobile devices in order to reduce the processing load, which usually leads to degradation of output (e.g., video quality in SAE), we propose c-SAE by applying the compressive sensing theory into SAE to maintain the video quality when the model is compressed. Our extensive experiments show that directly incorporating and processing spherical signals is promising, and it outperforms the traditional approaches by a large margin. Both p-SAE and c-SAE show their effectiveness in delivering high quality videos (e.g., PSNR results) when used alone or combined together with model compression.
Jin Zhou 0006, Na Li 0032, Yao Liu 0001, Shuochao Yao, Songqing Chen
ACM Multimedia2
2022 Applying VertexShuffle toward 360-degree video super-resolution
abstract
With the recent successes of deep learning models, the performance of 2D image super-resolution has improved significantly. Inspired by recent state-of-the-art 2D super-resolution models and spherical CNNs, in this paper, we design a novel spherical superresolution (SSR) approach for 360-degree videos. To address the bandwidth waste problem associated with 360-degree video transmission/streaming and save computation, we propose the Focused Icosahedral Mesh to represent a small area on the sphere and construct matrices to rotate spherical content to the focused mesh area. We also propose a novel VertexShuffle operation on the mesh, motivated by the 2D PixelShuffle operation. We compare our SSR approach with state-of-the-art 2D super-resolution models. We show that SSR has the potential to achieve significant benefits when applied to spherical signals.
Na Li 0032, Yao Liu 0001
NOSSDAV1