Xin Zhou 0001

dblp:05/3403-1 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0002-1496-405XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Security and privacy of machine learning · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 50% Integrated circuit design · 50%
Computer graphics and multimedia
1 paper
Image and video coding · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Security and privacy of machine learning
adversarial attack
1.012026
Robust Adversarial Patch for Object Detection Using Self-Similarity for Multiscale Attacks · IEEE Trans. Dependable Secur. Comput. 2026
Integrated circuit design
digital circuit design
0.212016
A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking Filter · IEEE Trans. Multim. 2016
Hardware accelerators and domain-specific architectures
video coding accelerator
0.212016
A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking Filter · IEEE Trans. Multim. 2016
Image and video coding › video compression › video codec
HEVC
0.112016
A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking Filter · IEEE Trans. Multim. 2016
Image and video coding
video coding standards
0.112016
A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking Filter · IEEE Trans. Multim. 2016

Methods — techniques the papers use, named apart from their topics

self-similarity · 1.0multiscale attack · 1.0ping-pong buffer design · 0.5parallel filtering · 0.5
YearPublicationVenuePosition
2026 MiCA: Intra-Modal Integration and Cross-Modal Alignment Adapters for Parameter-Efficient Referring Image Segmentation
abstract
Parameter-efficient transfer learning (PETL) has emerged as an effective strategy for fine-tuning large vision–language foundation models because it sharply reduces computational and memory overhead. However, existing PETL techniques underperform on dense prediction tasks that require fine-grained multimodal reasoning, such as referring image segmentation (RIS), owing to the lack of mechanisms that simultaneously strengthen local perception and enforce precise cross-modal alignment. We present a PETL framework with two lightweight and complementary adapters. The Global–Local Integrated Adapter (GLiA) enriches intra-modal features by coupling multi-scale depthwise-separable convolutions with a lightweight self-attention layer, capturing local context without sacrificing global dependencies. The Cross-Modal Alignment Adapter (CAA) explicitly aligns textual phrases with their corresponding visual regions, bridging the semantic gap between vision and language and enhancing multimodal reasoning. Experiments on three mainstream RIS benchmarks show that MiCA achieves the best accuracy while saving numerous updated parameters compared to the full fine-tuning and other PETL methods. Notably, with only 1.93% tunable backbone parameters, MiCA improves average accuracy by 0.8% across the three benchmarks compared to the baseline model.
Yang Li 0055, Zitong Feng, Tingrui Wang, Xin Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Robust Adversarial Patch for Object Detection Using Self-Similarity for Multiscale Attacks
Yang Li 0055, Tingrui Wang, Mingxin Fu, Xin Zhou 0001, Quan Pan 0001, Zhunga Liu
IEEE Trans. Dependable Secur. Comput.4
2020 Frame level rate control algorithm based on GOP level quality dependency for low-delay hierarchical video coding
Wei Zhou 0020, Henglu Wei, Xin Zhou 0001, Zhemin Duan
Signal Process. Image Commun.4
2019 All zero block detection for HEVC based on the quantization level of the maximum transform coefficient
Henglu Wei, Wei Zhou 0020, Xin Zhou 0001, Zhemin Duan
Multim. Tools Appl.4
2018 Prediction of Satisfied User Ratio for Compressed Video
abstract
A large-scale video quality dataset called the VideoSet has been constructed recently to measure human subjective experience of H.264 coded video in terms of the just-noticeable-difference (JND). It measures the first three JND points of 5-second video of resolution 1080p, 720p, 540p and 360p. Based on the VideoSet, we propose a method to predict the satisfied-user-ratio (SUR) curves using a machine learning framework. First, we partition a video clip into local spatial-temporal segments and evaluate the quality of each segment using the VMAF quality index. Then, we aggregate these local VMAF measures to derive a global one. Finally, the masking effect is incorporated and the support vector regression (SVR) is used to predict the SUR curves, from which the JND points can be derived. Experimental results are given to demonstrate the performance of the proposed SUR prediction method.
Haiqiang Wang, Ioannis Katsavounidis, Qin Huang 0006, Xin Zhou 0001, C.-C. Jay Kuo
ICASSP4
2017 VideoSet: A large-scale compressed video quality dataset based on JND measurement
abstract
• A large-scale JND-based coded video quality dataset is presented. • The VideoSet contains 220 5-s sequences in four resolutions coded by H.264/AVC. • The subjective test procedure, JND data cleaning and properties are described. • The significance and implications of the VideoSet are discussed. • This work points out a clear path to data-driven perceptual coding. A new methodology to measure coded image/video quality using the just-noticeable-difference (JND) idea was proposed in Lin et al. (2015). Several small JND-based image/video quality datasets were released by the Media Communications Lab at the University of Southern California in Jin et al. (2016) and Wang et al. (2016) [3]. In this work, we present an effort to build a large-scale JND-based coded video quality dataset. The dataset consists of 220 5-s sequences in four resolutions (i.e., 1920 × 1080 , 1280 × 720 , 960 × 540 and 640 × 360 ). For each of the 880 video clips, we encode it using the H.264/AVC codec with QP = 1 , … , 51 and measure the first three JND points with 30 + subjects. The dataset is called the “VideoSet”, which is an acronym for “Video Subject Evaluation Test (SET)”. This work describes the subjective test procedure, detection and removal of outlying measured data, and the properties of collected JND data. Finally, the significance and implications of the VideoSet to future video coding research and standardization efforts are pointed out. All source/coded video clips as well as measured JND data included in the VideoSet are available to the public in the IEEE DataPort (Wang et al., 2016 [4]).
Haiqiang Wang, Ioannis Katsavounidis, Jiantong Zhou, Jeong-Hoon Park, Shawmin Lei, Xin Zhou 0001, Man-On Pun, Xin Jin 0002, Ronggang Wang, Xu Wang 0006, Yun Zhang 0002, Jiwu Huang, Sam Kwong, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.6
2016 Visual saliency based perceptual video coding in HEVC
abstract
Perceptual video coding has the potential to provide the same visual quality at a lower bit-rate, compared with the traditional objective quality based scheme. Visual saliency represents the probability of human attention over frames, and it is used for allocating coding bits or controlling visual quality. In this paper, a HEVC compliant perceptual video coding scheme is proposed based on visual saliency. At first visual saliency map is attained to indicate the distribution of saliency. Then refined distortion allocating method is performed in CU level with adaptive QP which is adjusted by the average visual saliency. Besides, a fast CU mode decision algorithm suitable for perceptual video coding in HEVC is proposed to accelerate the encoder. In the fast algorithm, the average saliency is used to estimate texture complexity and movements in videos. Experimental results show that up to 22.52% bit-rate and 43.48% encoding time can be saved by our methods with negligible perceptual quality loss.
Henglu Wei, Xin Zhou 0001, Wei Zhou 0020, Chang Yan, Zhemin Duan, Nana Shan
ISCAS2
2016 Perceptual CU Size Decision and Fast Prediction Mode Decision Algorithm for HEVC Intra Coding
abstract
Intra coding has been significantly improved in HEVC over H.264/AVC with quad-tree based coding unit (CU) structure from size 64×64 to 8×8 and more prediction modes. However, these techniques cause a dramatic increase in computational complexity. In this paper, a novel intra coding algorithm is proposed consists of perceptual CU size decision algorithm and fast intra prediction mode decision algorithm. Firstly, based on the visual saliency detection, an adaptive and perceptual CU size decision method is proposed to alleviate intra encoding complexity. Furthermore, a fast intra prediction mode decision algorithm with step halving rough mode decision method is presented to selectively check the potential modes and effectively reduce the complexity of computation. Experimental results show that our proposed method reduces the computational complexity of the current HM to about 54.18% in encoding time with only 0.36% increases in BD rate and reasonable peak signal-to-noise ratio losses.
Xin Zhou 0001, Guangming Shi, Wei Zhou 0020
ISM1
2016 A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking Filter
abstract
This paper presents a high-throughput and multi-parallel VLSI hardware architecture for the deblocking filter in the HEVC video coding standard. First, an implementation-friendly and fast boundary judgment method is proposed to avoid using the original recursion loop approach. Then a dedicated parallel VLSI architecture composed of four parallel filtering cores is presented based on the proposed boundary judgment method. With the parallel luma/chroma filtering and parallel vertical/horizontal edges filtering order, the proposed VLSI architecture can process filtering operations for one largest coding unit (LCU) with less filtering cycles than other conventional approaches. Furthermore, filtering efficiency is improved due to a novel ping-pang buffer architecture and the on-chip single-port SRAM with dedicated data arrangement in the memory modules. Experimental results demonstrate that the proposed deblocking filter architecture improves the performance by 28-89% at the expense of the slightly increased gate count compared to the previously known architecture in HEVC. The proposed architecture can reach a high operating clock frequency of 278 MHz with TSMC 90 nm library and meet the real time requirement of the deblocking filter for 8 K × 4 K video format at 123 frame/s.
Wei Zhou 0020, Jingzhi Zhang, Xin Zhou 0001, Zhenyu Liu 0001, Xiaoxiang Liu
IEEE Trans. Multim.3
2015 An efficient interpolation filter VLSI architecture for HEVC
abstract
Firstly, an implementation-friendly interpolation filter algorithm is proposed in this paper. It can save 19.6% processing time on average with negligible coding quality degradation. Then based on the proposed algorithm, an optimized interpolation filter VLSI architecture, composed of the reused data path of interpolation, efficient memory organization and the pipeline interpolation filter engine is presented to reduce the implement hardware area. The resulting design can achieve 240 MHz with only 37.2K gate count and support real-time interpolation filter operation of 3840×2160@47fps video application by using 90nm CMOS technology.
Wei Zhou 0020, Xin Zhou 0001, Xiaocong Lian
ICASSP2
2015 An efficient all zero block detection algorithm based on frequency characteristics of DCT in HEVC
abstract
Like the previous video coding standard, DCT and quantization are also adopted in HEVC. Compared with H.264/AVC, HEVC employs larger transform blocks, which makes the all zero block detection algorithm designed for H.264/AVC inefficient for HEVC. An efficient all zero block detection algorithm aimed at HEVC, especially for 16×16 and 32×32 transform blocks, is proposed in this paper based on frequency characteristics of DCT. By analysing the distribution of residual energy in frequency domain, Hadamard transform is used to evaluate only a part of DCT coefficients which usually consume most of energy. To make the proposed algorithm efficient in complex sequences, a sum of absolute transformed difference based method is used to evaluate the maximum value of the rest of the coefficients. Experimental results show that more than 90% all zero blocks for 4×4, 8×8 and 16×16 transform blocks and about 80% for 32×32 transform blocks can be detected by the proposed algorithm. In addition, about 50% computational complexity in DCT/quantization can be reduced with negligible loss of video quality and compression efficiency.
Henglu Wei, Wei Zhou 0020, Xin Zhou 0001, Zhemin Duan
VCIP3
2015 A high-throughput deblocking filter VLSI architecture for HEVC
abstract
This paper presents a novel VLSI hardware architecture for the real-time high-throughput implementation of the HEVC deblocking filtering. Based on the proposed implementation-friendly boundary judgment method, a dedicated multi-parallel architecture composed of four parallel filtering cores, parallel luma/chroma filtering and parallel vertical/horizontal edges filtering is presented. Experimental results demonstrate that the proposed architecture can greatly improve the performance at the expense of the slightly increased hardware cost compared to the previously known architecture in HEVC. The proposed architecture can also meet the real-time requirement of the deblocking filter for 8K×4K video format at 123fps under 278MHz clock rate.
Wei Zhou 0020, Jingzhi Zhang, Xin Zhou 0001, Tongqing Liu
VCIP3
2007 Efficient Motion Estimation Scheme for H.264 Based on BP Neural Network
Wei Zhou 0020, Haoshan Shi, Zhemin Duan, Xin Zhou 0001
ISNN (3)4