Yang Zhou 0052

dblp:07/4580-52 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-7882-9028ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Feature Point-Based Two-Level Rate Control for AVS3
Ziming Jiang, Fengguang Liu, Yang Zhou 0052, Haibing Yin
DCC4
2025 Sparse4DGS: Flow-Geometry Assisted 4D Gaussian Splatting for Dynamic Sparse View Synthesis
abstract
Previous dynamic view synthesis works often struggle with limited available views, resulting in noticeable artifacts and blurriness in the outputs. In this paper, we present Sparse4DGS, a novel 4D Gaussian splatting framework that enables high-fidelity view synthesis from sparse inputs through three key innovations: 1) a global-local feature extraction encoder that integrates global Hexplane fields with local hash grids, effectively capturing both rough backgrounds and fine details; 2) a flow-guided feature aggregation module that stabilizes dynamic 3D Gaussians between adjacent frames, ensuring temporal continuity; and 3) a 4D geometry constraint scheme that utilizes monocular depth and pseudo-viewpoint depth supervision to enhance the structural consistency of dynamic scene details. Our approach achieves higher rendering quality while maintaining model compactness. Experimental results on various benchmarks demonstrate that our method performs favorably against state-of-the-art methods in terms of both rendering quality and training time. The source code and trained models are available at https://github.com/hu-dong-dong/Sparse4DGS.
Dongdong Hu, Yang Zhou 0052, Xiaofeng Huang, Haibing Yin, Zhu Li 0001
ACM Multimedia2
2025 Lightweight three-stream encoder-decoder network for multi-modal salient object detection
Junzhe Lu 0002, Tingyu Wang 0002, Bin Wan, Qiang Zhao 0005, Shuai Wang 0003, Yaoqi Sun, Yang Zhou 0052, Chenggang Yan 0001
J. Vis. Commun. Image Represent.7
2024 Enhanced Screen Content Image Compression: A Synergistic Approach for Structural Fidelity and Text Integrity Preservation
abstract
With the rapid development of video conferencing and online education applications, screen content image (SCI) compression has become increasingly crucial. Recently, deep learning techniques have made significant progress in compressing natural images, surpassing the performance of traditional standards like versatile video coding. However, directly applying these methods to SCIs is challenging due to the unique characteristics of SCIs. In this paper, we propose a synergistic approach to preserve structural fidelity and text integrity for SCIs. Firstly, external prior guidance is proposed to enhance structural fidelity and text integrity by providing global spatial attention. Then, a structural enhancement module is proposed to improve the preservation of structural information by enhanced spatial feature transform. Finally, the loss function is optimized for better compression efficiency in text regions by weighted mean square error. Experimental results show that the proposed method achieves 13.3% BD-Rate saving compared to the baseline window attention convolutional neural networks (WACNN) on the JPEGAI, SIQAD, SCID, and MLSCID datasets on average. Our code is available at https://github.com/vpaHduGroup/SFTIP_SCC.
Fangtao Zhou, Xiaofeng Huang, Peng Zhang 0007, Meng Wang 0017, Zhao Wang 0004, Yang Zhou 0052, Haibing Yin
ACM Multimedia6
2023 Video Transformer based Video Quality Assessment with Spatiotemporally adaptive Token Selection and Assembly
abstract
Video quality assessment (VQA) for user generated content (UGC) videos plays important role in video compression and processing. Convolutional neural network (CNN) based quality assessment for UGC is the research focus with inspiring model accuracy increment in the past three years. However, regularly temporal-sampling with temporal feature loss, as well as fixed token selection strategy video transformer (ViT) with insufficient representational capacity of tokens, jointly degrade the accuracy of conventional ViT based quality assessment. Facing these two challenges, this article proposes an adaptive token-selection ViT (ATSViT) structure for UGCVQA. Accounting for the uneven distribution of spatiotemporal distortion-related features, this work proposes a timing block sampling (TBS) module to adaptively select video blocks and assemble them into content compacted subsequence for further processing. In addition, inspired by the mental filter theory in terms of visual information, we propose a stage-wise adaptive screening network (SSNet) in which (noise” features of tokens in the sense of perception are progressively detected and processed by imitating the behavior of perception process in the eye-brain system. Experimental results verify that the proposed VQA model achieves state-of-the-art (SOTA) accuracy, with the highest correlation with mean opinion scores (MOS).
Shiling Zhao, Haibing Yin, Hongkui Wang, Yang Zhou 0052
DCC4
2023 Multi-stage affine motion estimation fast algorithm for versatile video coding using decision tree
Xiaofeng Huang, Fangtao Zhou, Weihong Niu, Yang Zhou 0052, Haibing Yin, Chenggang Yan 0001
J. Vis. Commun. Image Represent.6
2023 Fast all zero block detection algorithm for versatile video coding
Weihong Niu, Xiaofeng Huang, Haibing Yin, Yang Zhou 0052, Chenggang Yan 0001
Multim. Tools Appl.5
2020 Salient object detection via reliability-based depth compactness and depth contrast
abstract
It can be intuitively inferred that a high‐quality depth map can be used to quickly detect the salient region in stereo vision, implying that depth information plays an essential role in stereoscopic visual attention. However, existing methods generally use the depth map as an auxiliary cue to improve the saliency detection performance. In this study, the authors present an algorithm to directly detect the salient object from a high‐quality depth image. The proposed algorithm utilises a depth reliability indicator to assess the confidence of a depth image. Depth compactness, a novel feature that incorporates the depth reliability of the super‐pixels, is computed as a primary salient feature. Moreover, in order to enhance another salient feature (i.e. depth contrast), they develop a coarse background filtering method to suppress background interference. Experimental results demonstrate that the proposed method performs favourably against the popular depth‐aware saliency detection approaches at a lower computational cost.
Yang Zhou 0052, Yun Zhang 0002, Haibing Yin
IET Image Process.1
2019 Stereoscopic Visual Discomfort Prediction Using Multi-scale DCT Features
abstract
Prior approaches to the problem of visual discomfort prediction (VDP) for stereo/3D images are built for the uncompressed image. This paper presents a novel VDP method based on the compressed image by using multi-scale discrete cosine transform (MsDCT). Three types of visual discomfort features, including basic disparity intensity (BDI), disparity gradient energy (DGE) and disparity texture complexity (DTC), are extracted from two-dimensional (2-D) DCT coefficients. Additionally, a multi-scale transformation approach based on the different sizes of transform units is applied to obtain the multi-scale sub-features for each of the features. Then, through experimental comparison, a random forest regressor is chosen to fuse twenty-three sub-features to get the final objective prediction value of the S3D images. Experimental results conducted on two datasets show that the proposed method improves the prediction accuracy compared to those of recent S3D visual (dis)comfort predictors.
Yang Zhou 0052, Wanli Yu, Zhu Li 0001, Haibing Yin
ACM Multimedia1
2019 Distortion propagation modeling and its applications on frame level quantization control for predictive video coding
Haibing Yin, Xiaofeng Huang, Yang Zhou 0052
Signal Process. Image Commun.5
2018 Frame Level Quantization Control with Temporal Distortion Propagation Model for Video Coding
abstract
In video coder, inter-frame prediction causes distortion propagation among temporally adjacent frames, which complicates frame level bit allocation and quantization control. Quantization parameter cascading (QPC) is generally employed to determine a sequence of quantization parameter for dependent rate distortion optimization (RDO). This paper proposes a general framework for temporal dependency analysis by lever-aging a distortion propagation model. The amount of distortion propagated from the temporally adjacent frames is measured by tree-style dependent analysis. Then, a trellis comprised of frame level quantization parameters of one GOP is constructed to achieve global optimization via branch-prune based dynamic programming. The simulation results verify that the frame level QPC algorithm with the proposed distortion model achieves up to 1.2dB—1.5dB PSNR improvement on average, with smaller temporal distortion fluctuation contributed by efficient bit allocation.
Haibing Yin, Yang Zhou 0052
ISCAS4
2018 An efficient all-zero block detection algorithm for high efficiency video coding with RDOQ
Haibing Yin, Enhui Yang, Yang Zhou 0052
Signal Process. Image Commun.4
2017 Allowable depth distortion based fast mode decision and reference frame selection for 3D depth coding
Yun Zhang 0002, Zhaoqing Pan, Yang Zhou 0052, Linwei Zhu
Multim. Tools Appl.3
2017 Visual comfort prediction for stereoscopic image using stereoscopic visual saliency
Yang Zhou 0052, Yongjian He, Yun Zhang 0002
Multim. Tools Appl.1