Runyu Yang

dblp:55/6759 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing agent-based wayfinding simulation for transportation hub evaluation with a visual memory graph Neural network (VM-GNN)
Runyu Yang, Mingyan Zou, Siwei Su
Adv. Eng. Informatics1
2025 How Universal Are SAM2 Features?
Masoud Khairi Atani, Alon Harell, Hyomin Choi, Runyu Yang, Fabien Racapé, Ivan V. Bajic
PCS4
2025 Bit Allocation Transfer for Perceptual Quality Enhancement of VVC Intra Coding
Runyu Yang, Ivan V. Bajic
PCS1
2025 Fourier Convolution Block with global receptive field for MRI reconstruction
Haozhong Sun, Zhongsen Li, Runyu Yang, Jiaqi Dou, Haikun Qi, Huijun Chen
Medical Image Anal.4
2024 Static-Scene Video Compression with Neural Radiance Fields
abstract
We study how to efficiently compress static-scene video, which captures a static scene from continuous viewpoints. The static-scene videos have emerged as a key ingredient for virtual/augmented reality. For static-scene videos, the existing video coding schemes such as H.265/HEVC may have limited compression efficiency due to the limited number of reference frames, which prevents the sufficient utilization of the temporal correlation. We build a static-scene video compression scheme using the recently developed technologies of Neural Radiance Fields (NeRF). The scheme is shown in Figure 1 . In our proposed scheme, the encoder derives and encodes the camera parameters, and compresses some selected keyframes; the decoder adopts an efficient NeRF algorithm to build an implicit scene model for reconstructing all the frames. To the compression of the camera parameters, we design a algorithm referring to the method of motion vector compression in video coding. It uses the reconstructed parameters of the previous frame plus the difference between the reconstructed parameters of the previous two frames as predictions and then quantify residuals to reduce bitrates.
Runyu Yang
DCC1
2024 Neural Radiance Field-Assisted Static-Scene Video Coding
abstract
We investigate how to compress static-scene videos efficiently, where each video captures a static scene from different viewpoints. These videos are an essential component of virtual and augmented reality. Current video coding standards like H.265/HEVC have limited compression efficiency due to the limited number of reference frames available, which prevents the full utilization of the abundant temporal redundancy in static-scene videos. To address this limitation, we propose a compression scheme that uses Neural Radiance Fields (NeRF). In our proposed scheme, the encoder derives and encodes the camera parameters, and compresses selected keyframes. The decoder uses an efficient NeRF algorithm to construct an implicit scene model for reconstructing all frames. The experimental results show that our proposed scheme achieves both objective and subjective quality improvement compared to H.265/HEVC, especially at low bit rates.
Runyu Yang
ICIP1
2024 CC-SMC: Chain coding-based segmentation map lossless compression
Runyu Yang, Dong Liu 0002, Feng Wu 0001, Wen Gao 0001
J. Vis. Commun. Image Represent.1
2024 Perceptual Quality-Oriented Rate Allocation via Distillation from End-to-End Image Compression
abstract
Mainstream image/video coding standards, exemplified by the state-of-the-art H.266/VVC, AVS3, and AV1, follow the block-based hybrid coding framework. Due to the block-based framework, encoders designed for these standards are easily optimized for peak signal-to-noise ratio (PSNR) but have difficulties optimizing for the metrics more aligned to perceptual quality, e.g., multi-scale structural similarity (MS-SSIM), since these metrics cannot be accurately evaluated at the small block level. We address this problem by leveraging inspiration from the end-to-end image compression built on deep networks, which is easily optimized through network training for any metric as long as the metric is differentiable. We compared the trained models using the same network structure but different metrics and observed that the models allocate rates in different ratios. We then propose a distillation method to obtain the rate allocation rule from end-to-end image compression models with different metrics and to utilize such a rule in the block-based encoders. We implement the proposed method on the VVC reference software—VTM and the AVS3 reference software—HPM, focusing on intraframe coding. Experimental results show that the proposed method on top of VTM achieves more than 10% BD-rate reduction than the anchor when evaluated with MS-SSIM or LPIPS, which leads to concrete perceptual quality improvement.
Runyu Yang, Dong Liu 0002, Siwei Ma 0001, Feng Wu 0001, Wen Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2021 Knowledge Distillation From End-To-End Image Compression To Vvc Intra Coding For Perceptual Quality Enhancement
abstract
In the current hybrid coding schemes, mean-squared-error is widely used for the rate-distortion optimization, which leads to high peak signal-to-noise ratio but sub-optimal perceptual quality. Although human perception-related measures, like multi-scale structural similarity (MS-SSIM), have been proposed, plugging them into the hybrid coding schemes may be computationally expensive. Recently, end-to-end optimized image compression has demonstrated the advantage of perceptual quality-oriented optimization by simply changing the training loss function. Inspired by this, we propose to distill the “perceptual” knowledge from end-to-end image compression and use the knowledge to enhance the perceptual quality for Versatile Video Coding (VVC) intra coding. For an input image, we obtain the block-level bit allocation via end-to-end image compression, and use the bit allocation to adjust the quantization parameter of VVC intra coding. Being compatible to the VVC standard, our method achieves on average 9.32% BD-rate reduction on the Kodak image set when evaluated by MS-SSIM, compared to the VVC reference software.
Runyu Yang, Dong Liu 0002, Siwei Ma 0001, Feng Wu 0001, Wen Gao 0001
ICIP1
2021 Striatal Subdivisions Estimated via Deep Embedded Clustering With Application to Parkinson's Disease
abstract
Recent fMRI connectivity-based parcellation (CBP) methods have been developed to obtain homogeneous and functionally coherent brain parcels. However, most of these studies utilize traditional clustering methods that neglect hidden nonlinear features. To enhance parcellation performance, here we propose a deep embedded connectivity-based parcellation (DECBP) framework and apply it to determine functional subdivisions of the striatum in public resting state fMRI data sets. This framework integrates fMRI connectivity features into deep embedded clustering (DEC), a deep neural network based on a stacked autoencoder. Compared to three prevalent clustering methods and their combinations with principal component analysis (PCA), the DECBP exhibited a significantly higher similarity between scans, individuals, and groups, indicating enhanced reproducibility. The generated reliable parcellations were also largely consistent with other public atlases. We further explored the functional subunits in the striatum in a data set from 23 Parkinson's disease (PD) subjects and 27 age-matched healthy controls (HC). All putaminal subregions of PD demonstrated lower interhemispheric connectivity than those of HC, which might reflect imbalance in the pathological progression of PD. Such hypo-connectivity was also observed between putaminal subregions and other brain regions, reflecting neuroimaging manifestations of the altered cortico-striato-thalamo-cortical circuit. These observed weaker couplings were associated with PD severity and duration. Our results support the utilization of the DECBP framework and suggest that abnormal connectivity in putaminal subregions may be a potential indicator of PD.
Yu Li 0027, Aiping Liu, Taomian Mi, Runyu Yang, Piu Chan, Martin J. McKeown, Xun Chen 0001, Feng Wu 0001
IEEE J. Biomed. Health Informatics4
2020 Chain Code-Based Occupancy Map Coding for Video-Based Point Cloud Compression
abstract
In video-based point cloud compression (V-PCC), occupancy map video is utilized to indicate whether a 2-D pixel corresponds to a valid 3-D point or not. In the current design of V-PCC, the occupancy map video is directly compressed losslessly with High Efficiency Video Coding (HEVC). However, the coding tools in HEVC are specifically designed for natural images, thus unsuitable for the occupancy map. In this paper, we present a novel quadtree-based scheme for lossless occupancy map coding. In this scheme, the occupancy map is firstly divided into several coding tree units (CTUs). Then, the CTU is divided into coding units (CUs) recursively using a quadtree. The quadtree partition is terminated when one of the three conditions is satisfied. Firstly, all the pixels have the same value. Secondly, the pixels in the CU only have two kinds of values and they can be separated by a continuous edge whose endpoints lie on the side of the CU. The continuous edge is then coded using chain code. Thirdly, the CU reaches the minimum size. This scheme simplifies the design of block partitioning in HEVC and designs simpler yet more effective coding tools. Experimental results show significant reduction of bit-rate and complexity compared with the occupancy map coding scheme in V-PCC. In addition, this scheme is also very efficient to compress the semantic map.
Runyu Yang, Ning Yan 0001, Li Li 0040, Dong Liu 0002, Feng Wu 0001
VCIP1