Zeming Zhao

dblp:150/4056 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Subjective and objective evaluation of visual security in perceptually encrypted images
Xiaodong Bi, Xiaohai He, Zeming Zhao, Haitao Wei, Shuhua Xiong, Zheng Liu 0002, Ray E. Sheriff
Expert Syst. Appl.3
2026 Blind visual security assessment using a simple parallel dual-stream network
Xiaodong Bi, Xiaohai He, Zeming Zhao, Shuhua Xiong, Honggang Chen, Ray E. Sheriff
Signal Process. Image Commun.3
2026 Human-Machine Vision Collaboration Based Rate Control Scheme for VVC
abstract
With the widespread adoption of smart terminals, compressed video is increasingly utilized in the receiver for purposes beyond human vision. Conventional video coding standards are optimized primarily for human visual perception and often fail to accommodate the distinct requirements of machine vision. To simultaneously satisfy the perceptual needs and the analytical demands, we propose a novel rate control scheme based on Versatile Video Coding (VVC) for human-machine vision collaborative video coding. Specifically, we employ the You Only Look Once (YOLO) network to extract task-relevant features for machine vision and formulate a detection feature weight based on these features. Leveraging the feature weight and the spatial location information of Coding Tree Units (CTUs), we propose a region classification algorithm that partitions a frame into machine vision-sensitive region (MVSR) and machine vision non-sensitive region (MVNR). Subsequently, we develop an enhanced and refined bit allocation strategy that performs region-level and CTU-level bit allocation, thereby improving the precision and effectiveness of the rate control. Experimental results demonstrate that the scheme improves machine task detection accuracy while preserving perceptual quality for human observers, effectively meeting the dual encoding requirements of human and machine vision.
Zeming Zhao, Xiaohai He, Xiaodong Bi, Shuhua Xiong
IEEE Signal Process. Lett.1
2026 Decoupled Prescribed Performance and Safe Formation Control of Multi-Agent Systems Under Input Constraints
abstract
Reliable formation control in real-world multi-agent systems is challenging due to the concurrent need to meet performance specifications, enforce safety constraints, and respect actuator limitations. While prescribed performance control (PPC) ensures bounded error evolution via predefined performance functions, incorporating safety and input constraints within this framework remains nontrivial. This paper develops a modular decoupled control architecture that integrates PPC and control barrier function (CBF) to enforce performance and safety. The fixed performance bound limitation in PPC control is overcome through the introduction of an auxiliary system that adaptively adjusts performance functions, thereby effectively enabling the quantification of performance degradation due to constraints while avoiding control singularities. To further mitigate conflicts between safety and performance, an online trajectory optimization module is designed to generate smooth and collision-free reference trajectories. The proposed approach is validated on a team of Crazyflie quadrotors navigating obstacle environments, demonstrating safe and accurate formation tracking under stringent constraints.
Xinyue Zhao, Qingkai Yang, Kefan Zheng, Zeming Zhao, Kaifeng Zheng, Hao Fang 0001
IEEE Trans Autom. Sci. Eng.4
2026 Spatial-Temporal Correlation Information-Based Rate Control for Versatile Video Coding
abstract
Although lambda-domain-based rate control is widely used in video encoders, developing an efficient rate control scheme for Coding Tree Units (CTUs) under the rate-distortion (R-D) principle remains a significant challenge. In this paper, we propose a spatial-temporal correlation information-based rate control scheme for Versatile Video Coding (VVC), aiming to improve coding performance. We introduce a weight estimation network to establish a CTU-level bit allocation strategy that fully exploits spatial-temporal contextual information. Moreover, the CTU-level coding parameter λ is adaptively optimized based on a dependency factor derived from distortion dependency information in both the spatial and temporal domains. Experimental results demonstrate that, compared to the default VVC rate control, the proposed scheme achieves BD-Rate savings of 6.48%, 17.33% and 13.75% in terms of the Peak Signal-to-Noise Ratio (PSNR), the Multi-Scale Structural Similarity Index (MS-SSIM) and the Video Multimethod Assessment Fusion (VMAF), respectively, under the Low Delay_P (LDP) configuration in the VVC Test Model (VTM) 19.0. Furthermore, the proposed method outperforms other state-of-the-art rate control schemes.
Zeming Zhao, Xiaohai He, Shuhua Xiong, Meng Wang 0017, Shiqi Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Efficient Coding Parameters Optimization for Rate Control in Versatile Video Coding
abstract
In mainstream video encoders, rate control is crucial in scenarios with limited bandwidth. Within the existing R-$\lambda$model-based rate control scheme, the coding parameters ($\alpha$and$\beta$) are directly involved in the calculation of$\lambda$, and their values have a significant impact on the calculation results. By selecting precise and appropriate coding parameters, enhanced rate control and rate-distortion performance can be realized. However, during the mapping of target bits to$\lambda$,$\alpha$and$\beta$often inadequately consider the rate-distortion attributes and the content characteristics of the coding units. This paper introduces a parameter optimization algorithm for rate control in Versatile Video Coding (VVC), aimed at enhancing coding efficiency. Utilizing the actual coded contexts of the coded Coding Tree Unit (CTU) alongside pre-coding information, we establish a parameters relationship model to deliver better coding parameters according to the rate-distortion attributes of the current CTU. Furthermore, leveraging coded contexts from spatially adjacent CTUs and the feature complexity, we propose a spatial coupling strategy to further improve the preceding coding parameters, considering the content characteristics of CTU. The proposed coding parameter optimization algorithm is implemented in the rate control of the VVC test model (VTM). Experimental findings indicate that this optimization algorithm accomplishes BD-rate savings concerning Peak Signal-to-Noise Ratio (PSNR) as well as the Multiscale Structural Similarity Index Metric (MS-SSIM) across various configurations. In addition, a more stable buffer status and enhanced visual quality are visible, which highlights the benefits of the proposed algorithm.
Zeming Zhao, Xiaohai He, Xiaodong Bi, Qizhi Teng, Shuhua Xiong
IEEE Trans. Multim.1
2026 Rate Control for 360$^{\circ }$ Versatile Video Coding Based on Visual Gaze Mechanism
abstract
In the past few years, 360° video has started to infiltrate various aspects of daily life. Although there have been significant developments in 360° video coding technology, understanding of the human visual gaze mechanism has been somewhat overlooked. In this paper, we propose a rate control scheme for 360° Versatile Video Coding (VVC) based on a human visual gaze mechanism, targeting at improving the coding performance and bitrate accuracy. More specifically, based on the Equi-rectangular Projection (ERP) format, latitude information is systematically analyzed and a stripe-level bit allocation scheme is established, to better mitigate the projection distortion. Subsequently, the Lagrange parameter λ is further optimized with distortion dependency and identification of the visual gaze guided key Coding Tree Units (CTUs). The proposed rate control scheme is implemented on the VVC Test Model for 360° video. Experimental results show that the proposed rate control scheme can achieve BD-rate savings in terms of Weighted to Spherically uniform-Peak Signal-to-Noise Ratio (WS-PSNR) and Sphere-Peak Signal-to-Noise Ratio (S-PSNR) under the various configurations, respectively. Meanwhile, a healthier buffer status and better visual quality can be observed, further demonstrating the advantages of the proposed scheme.
Zeming Zhao, Meng Wang 0017, Xiangjie Sui, Peilin Chen 0001, Xiaohai He, Shiqi Wang 0001
IEEE Trans. Multim.1
2025 CTU-Level Rate Control with λ Optimization Based on Visual Gaze Mechanism for 360-Degree Versatile Video Coding
abstract
Understanding the human visual gaze mechanism is crucial for enhancing 360° video coding technology. This paper presents a Coding Tree Unit (CTU)-level rate control scheme with λ optimization strategy for 360° Versatile Video Coding (VVC), with the aim of enhancing rate-distortion performance and bitrate accuracy. Specifically, the Lagrange parameter λ is optimized with consideration of distortion dependency and the identification of key CTUs guided by visual gaze, which are derived from a 360° video path generation network, thoroughly integrating the characteristics of the human visual gaze. Experimental results show that the proposed scheme achieves BD-rate savings in terms of Weighted to Spherically uniform-Peak Signal-to-Noise Ratio (WS-PSNR) and Sphere-Peak Signal-to-Noise Ratio (S-PSNR) across various coding configurations.
Zeming Zhao, Meng Wang 0017, Xiangjie Sui, Xiaohai He, Shiqi Wang 0001
ICIP1
2024 Blind video quality assessment based on Spatio-Temporal Feature Resolver
Xiaodong Bi, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen, Ray E. Sheriff
Neurocomputing4
2024 Fast CU partition strategy based on texture and neighboring partition information for Versatile Video Coding Intra Coding
Ruolan Yang, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen
Multim. Tools Appl.4
2023 Efficient Rate Control in Versatile Video Coding With Adaptive Spatial-Temporal Bit Allocation and Parameter Updating
abstract
Despite the fact that Versatile Video Coding (VVC) has achieved superior coding performance, two major problems remain for the rate control (RC) model in VVC. First, the regions concerned by human eyes are not clear enough in the coded video due to the deviation between the target bit allocation strategy of the coding tree unit (CTU) in RC and the human visual attention mechanism (HVAM). Second, there are significant quality fluctuations in the coded video frames due to the inappropriate updating speed. To address the above problems, we propose an efficient rate control (ERC) model. Specifically, in order to make the coded video more consistent with the attention of human eyes, we extract texture and motion-based spatial-temporal information to guide the bit allocation at the CTU level. Furthermore, based on the quasi-Newton algorithm and bit error, we propose an adaptive parameter updating (APU) method with the proper updating speed to precisely control the bits per frame. The proposed ERC outperforms the default RC model of VVC Test Model (VTM) 9.1 by saving the average Bjøntegaard Delta Rate (BD-Rate) on full-frame video sequences by 3.60% and 4.94% under low delay P (LDP) and random access (RA) configurations respectively, with higher bitrate accuracy. Moreover, the Peak Signal-to-Noise Ratio (PSNR) and actual coded bits per frame in the video coded by the proposed ERC are more stable.
Liqiang He, Xiaohai He, Shuhua Xiong, Zeming Zhao, Honggang Chen
IEEE Trans. Circuits Syst. Video Technol.4
2023 Cumulative Spike Train Estimation for Muscle Excitation Assessment From Surface EMG Using Spatial Spike Detection
abstract
Estimating cumulative spike train (CST) of motor units (MUs) from surface electromyography (sEMG) is essential for the effective control of neural interfaces. However, the limited accuracy of existing estimation methods greatly hinders the further development of neural interface. This paper proposes a simple but effective approach for identifying CST based on spatial spike detection from high-density sEMG. Specifically, we use a spatial sliding window to detect spikes according to the spatial propagation characteristics of the motor unit action potential, focusing on the spikes of activated MUs in a local area rather than those of a specific MU. We validated the effectiveness of our proposed method through an experiment involving wrist flexion/extension and pronation/supination, comparing it with a recognized CST estimation method and an MU decomposition based method. The results demonstrated that the proposed method obtained higher accuracy on multi-DoF wrist torque estimation leveraging the estimated CST compared to the other three methods. On average, the correlation coefficient (R) and the normalized root mean square error (nRMSE) between the estimation results and recorded force were 0.96 ± 0.03 and 10.1% ± 3.7%, respectively. Moreover, there was an extremely high interpretive extent between the CSTs of proposed method and the MU decomposition method. The outcomes reveal the superiority of the proposed method in identifying CSTs and can provide promising driven signals for neural interface.
Yang Xu 0079, Yang Yu 0019, Zeming Zhao, Chen Chen 0045, Xinjun Sheng
IEEE J. Biomed. Health Informatics3
2022 An Optimized Rate Control Algorithm in Versatile Video Coding for 360$^\circ$ Videos
abstract
Today, 360$°$video has become an integral part of people's lives. Despite the fact that the latest generation standard Versatile Video Coding (VVC) demonstrates a significant gain in encoding capacity over High Efficiency Video Coding (HEVC), it still has room for 360$°$video encoding improvements. To further enhance the applicability of 360$°$video coding, an optimized rate control (RC) algorithm in VVC for 360$°$video is proposed in this paper. We present an efficient extraction algorithm for obtaining the video's saliency feature. Furthermore, for the characteristics of 360$°$video, a partitioning algorithm is also proposed to divide a frame into demand and non-demand regions. Additionally, to achieve precise and rational RC, a Coding Tree Unit (CTU)-level bit allocation strategy is proposed based on the saliency feature for the above-mentioned regions. The experimental results show that the proposed RC algorithm can achieve 11.77$\%$bitrate savings and more accurate allocation compared with the default algorithm of VVC. Also, performance enhancement has been observed in comparison to the most advanced algorithm.
Zeming Zhao, Xiaohai He, Shuhua Xiong, Liqiang He, Ray E. Sheriff
IEEE Signal Process. Lett.1
2021 An improved R-λ rate control model based on joint spatial-temporal domain information and HVS characteristics
Zeming Zhao, Shuhua Xiong, Weiheng Sun, Xiaohai He, Feiran Zhang
Multim. Tools Appl.1
2014 A combined model for scan path in pedestrian searching
abstract
Target searching, i.e. fast locating target objects in images or videos, has attracted much attention in computer vision. A comprehensive understanding of factors influencing human visual searching is essential to design target searching algorithms for computer vision systems. In this paper, we propose a combined model to generate scan paths for computer vision to follow to search targets in images. The model explores and integrates three factors influencing human vision searching, top-down target information, spatial context and bottom-up visual saliency, respectively. The effectiveness of the combined model is evaluated by comparing the generated scan paths with human vision fixation sequences to locate targets in the same images. The evaluation strategy is also used to learn the optimal weighting coefficients of the factors through linear search. In the meanwhile, the performances of every single one of the factors and their arbitrary combinations are examined. Through plenty of experiments, we prove that the top-down target information is the most important factor influencing the accuracy of target searching. The effects from the bottom-up visual saliency are limited. Any combinations of the three factors have better performances than each single component factor. The scan paths obtained by the proposed model are optimal, since they are most similar to the human vision fixation sequences.
Lijuan Duan, Zeming Zhao, Wei Ma 0008, Jili Gu, Zhen Yang 0004, Yuanhua Qiao
IJCNN2