Shengyang Zhao

dblp:192/8503 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Token Compression for the Understanding and Generation Unified MLLMs
Junyan Lin, Jinming Liu 0001, Shengyang Zhao, Xin Jin 0014
ISCAS3
2026 A Hybrid Subjective Quality Assessment Framework for Light Field Coding
Saeed Mahmoudpour, Mylène C. Q. Farias, Shengyang Zhao
QoMEX3
2025 Hybrid-Grained Feature Aggregation with Coarse-to-Fine Language Guidance for Self-Supervised Monocular Depth Estimation
abstract
Current self-supervised monocular depth estimation (MDE) approaches encounter performance limitations due to insufficient semantic-spatial knowledge extraction. To address this challenge, we propose Hybrid-depth, a novel framework that systematically integrates foundation models (e.g., CLIP and DINO) to extract visual priors and acquire sufficient contextual information for MDE. Our approach introduces a coarse-to-fine progressive learning framework: 1) Firstly, we aggregate multi-grained features from CLIP (global semantics) and DINO (local spatial details) under contrastive language guidance. A proxy task comparing close-distant image patches is designed to enforce depth-aware feature alignment using text prompts; 2) Next, building on the coarse features, we integrate camera pose information and pixel-wise language alignment to refine depth predictions. This module seamlessly integrates with existing self-supervised MDE pipelines (e.g., Monodepth2, ManyDepth) as a plug-and-play depth encoder, enhancing continuous depth estimation. By aggregating CLIP's semantic context and DINO's spatial details through language guidance, our method effectively addresses feature granularity mismatches. Extensive experiments on the KITTI benchmark demonstrate that our method significantly outperforms SOTA methods across all metrics, which also indeed benefits downstream tasks like BEV perception. Code is available at https://github.com/Zhangwenyao1/Hybrid-depth.
Hongsi Liu, Bohan Li 0015, Jiawei He 0002, Zekun Qi, Yunnan Wang, Shengyang Zhao, Xinqiang Yu, Wenjun Zeng 0001, Xin Jin 0014
ICCV7
2025 Standard Codec is Enough: A Training-Free 4D Gaussian Compression with Dynamic UV Mapping
abstract
4D Gaussian Splatting (4DGS) has demonstrated advances in the dynamic scene representation. However, the time-varying attributes across frames introduce considerable storage and transmission costs, making 4DGS challenging to widely deploy. Existing compression methods struggle to obtain inter-frame residuals due to the unstructured nature of Gaussian representations, making explicit motion estimation and residual modeling inherently challenging. To address these, we propose a Training-Free 4D Gaussian Compression framework, TF4DGC, which transforms 4D Gaussian into a well-structured 2D representation, easy to estimate motion for coding, via a UV mapping. Specifically, we project 3D Gaussians onto a canonical sphere to obtain temporally consistent UV coordinates, and organize per-frame Gaussian attributes into multi-channel video sequences. This design enables the direct use of standard video codecs (e.g., AVC, HEVC) for compression, which is compatible with widespread hardware decoder support on laptops and mobile devices. Experimental results show that our method efficiently compresses both reconstructed and generated Gaussian scenarios, highlighting its general applicability. Our method offers a scalable and practical solution for 4DGS compression and facilitates real-time deployment in bandwidth constrained environments.
Jinming Liu 0001, Shengyang Zhao, Qiang Hu 0003, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
VCIP3
2025 Quadtree Partitioning-based Visual Token Pruning for MLLMs Considering Information Density
abstract
Multimodal Large Language Models (MLLMs) excel at comprehensive understanding by integrating visual and textual information. However, their inference speed is often bottlenecked by redundant visual token inputs. Existing methods tend to alleviate this issue with a heuristic pruning strategy based on token importance, tailored to certain commonly adopted vision encoders like CLIP. In this paper, we propose a novel training-free token pruning method based on a well-designed metric of information density, where we decide which tokens are retained according to their entropy, following the classic information theory. Based on that, we further propose a quadtree partitioning strategy, in which we retain these tokens with higher entropy so as to preserve the visual spatial structure while allocating more tokens to more informative regions. Experiments on LLaVA-v1.5-7B and 13B across six benchmarks show our method achieves state-of-the-art performance—retaining over 90% of full-token accuracy even at a 6.25% token budget—while cutting TFLOPs by up to 20% compared to FastV and by 81% compared to the original LLaVA-v1.5.
Yuntao Wei, Jinming Liu 0001, Shengyang Zhao, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
VCIP3
2024 An Explainable Spectral Analysis For Light Field Image Quality Assessment
abstract
The light field (LF) technology is a promising way to achieve the immersive multimedia. The LF image quality assessment is vital for user experience. Compared with traditional twodimensional images, the LF images suffer from not only spatial distortion but also angular distortion, which destroys the light field angular consistency. Many algorithms are proposed to handle the objective light field image quality assessment. However, there is no mathematical description for the LF angular consistency and no discussion of spectral analysis in the angular domain. In this paper we provide an explainable spectral analysis for the LF image quality assessment. First, we formulate the LF as a 4 D signal. It is well-known that the energy of the light field Epipolar Plane Image (EPI) in the frequency domain should concentrate in a double-wedge area, but for the first time we introduce it into the LFI quality assessment. We test and discuss how the different distortions affect the energy distribution. The spectral features for blind LF image quality assessment are further proposed to complement the spatial features extracted by off-the-shelf algorithm. Experimentally we demonstrate that the features based on spectral energy distribution are more sensitive to angular distortion. Consequently, it can be a good guidance for a better LF quality assessment algorithm design.
Shengyang Zhao, Xin Jin 0014
ICIP1
2024 A Subjective Test Framework for JPEG Pleno Quality Assessment
abstract
The Joint Photographic Experts Group (JPEG) is currently addressing the challenges in assessing plenoptic image quality by developing new standards for both subjective and objective assessment. This process entails revisiting existing recommendations and establishing a visual quality assessment (QA) framework that considers the unique aspects of plenoptic data. The focus is currently on the light field modality, with efforts to gather expert contributions and develop tools that cater both subjective and objective QA, ensuring that the evolving requirements of plenoptic imaging QA are met. This paper presents the JPEG Pleno subjective test tool, a pivotal element in JPEG’s standardization endeavor, designed to facilitate a range of subjective QA experiments that steer the decisions during a standardization process.
Shengyang Zhao, Saeed Mahmoudpour, Mylène C. Q. Farias, Carla L. Pagliari, Peter Schelkens
QoMEX1
2024 Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMs
abstract
We present a new image compression paradigm to achieve "intelligently coding for machine" by cleverly leveraging the common sense of Large Multimodal Models (LMMs). We are motivated by the evidence that large language/multimodal models are powerful general-purpose semantics predictors for understanding the real world. Different from traditional image compression typically optimized for human eyes, the image coding for machines (ICM) framework we focus on requires the compressed bitstream to more comply with different downstream intelligent analysis tasks. To this end, we employ LMM to${\text{tell codec what to compress}}$: 1) first utilize the powerful semantic understanding capability of LMMs w.r.t object grounding, identification, and importance ranking via prompts, to disentangle image content before compression, 2) and then based on these semantic priors we accordingly encode and transmit objects of the image in order with a structured bitstream. In this way, diverse vision benchmarks including image classification, object detection, instance segmentation, etc., can be well supported with such a semantically structured bitstream. We dub our method "SDComp" for "Semantically Disentangled Compression", and compare it with state-of-the-art codecs on a wide variety of different vision tasks. SDComp codec leads to more flexible reconstruction results, promised decoded visual quality, and a more generic/satisfactory intelligent task-supporting ability.
Jinming Liu 0001, Yuntao Wei, Junyan Lin, Shengyang Zhao, Heming Sun, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014
VCIP4
2021 Various density light field image coding based on distortion minimization interpolation
Shengyang Zhao, Zhibo Chen 0001
J. Vis. Commun. Image Represent.1
2019 Belif: Blind Quality Evaluator Of Light Field Image With Tensor Structure Variation Index
abstract
With the development of immersive media, Light Field Image (LFI) quality assessment is becoming more and more important, which helps to better guide light field acquisition, processing and application. However, almost all existing LFI quality assessment schemes utilize the 2D or 3D quality assessment methods while ignoring the intrinsic high dimensional characteristics of LFI. Therefore, we adopt the tensor theory to explore the LF 4D structure characteristics and propose the first Blind quality Evaluator of LIght Field image (BELIF). We generate cyclopean images tensor from the original LFI and then the features are extracted by the tucker decomposition. Specifically, Tensor Spatial Characteristic Features (TSCF) for spatial quality and Tensor Structure Variation Index (TSVI) for angular consistency are designed to fully assess the LFI quality. Extensive experimental results on the public LFI databases demonstrate that BELIF signifi-cantly outperforms the existing image quality assessment algorithms.
Likun Shi, Shengyang Zhao, Zhibo Chen 0001
ICIP2
2018 Perceptual Evaluation of Light Field Image
abstract
Recently, light field image has attracted wide attention. However, much less work has been conducted on the perceptual evaluation of light field image. In this work, we create the first windowed 5 degree of freedom light field image database (Win5-LID) based on stereoscopic display, which provides windowed 5 DOF experience and all the depth cues of light field image. The database consists of light field images with representative compression and reconstruction artifacts. We assume that the light field quality is not only affected by sub-views quality but also depth cues. Picture quality and overall quality are then evaluated and the results validate our assumption. Finally, the performance of existing image quality metrics is analyzed on our database. The results indicate that the performance of the state-of-the-art image quality metrics remains to be improved.
Likun Shi, Shengyang Zhao, Wei Zhou 0021, Zhibo Chen 0001
ICIP2
2017 Light field image coding via linear approximation prior
abstract
In recent years, the light field (LF) image as a new imaging modality has attracted much interest. While light field camera records both the luminance and direction of the rays in a scene, large amount of data makes it a great challenge for storage and transmission. Thus an adequate compression scheme is desired. In this paper, we propose a new prior, called linear approximation prior that reveals intrinsic property among the LF sub-views. It indicates that we can approximate a certain view with a weighted sum of other views. By fully exploiting this prior we propose a powerful coding scheme. The experiments show the superior performance of our scheme, which achieves as large as 45.51% BD-rate reduction and 37.41% BD-rate reduction on average compared with the High Efficiency Video Coding (HEVC).
Shengyang Zhao, Zhibo Chen 0001
ICIP1
2016 Light field image coding with hybrid scan order
abstract
A Light field image contains shear amount of data as it keeps the full spatio-angular information of the real scene. In this paper we propose a light field image coding scheme based on the latest JEM coding technologies. We propose a novel hybrid scan order to rearrange subaperture images into an image sequence and verify its importance to coding performance of light field image format. The experiment on EPFL light field image dataset demonstrates that our scheme achieves 7.06 dB gain compared with directly encoding the image by the JPEG standard. With the QP set to 50, our scheme achieves an average compression ratio of 7107, and still provides larger PSNRs and better viewing experience than JPEG at a compression ratio of 100.
Shengyang Zhao, Zhibo Chen 0001, Hongrui Huang
VCIP1