Zhijian Hao

dblp:259/2435 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HLC: A High-Quality Lightweight Mezzanine Codec Featuring High-Throughput Palette
Chenlong He, Leilei Huang, Wei Li 0257, Hanyang Cui, Zhijian Hao, Xiaoyang Zeng, Yibo Fan
ISCAS5
2026 Temporal Quality Aggregation for VQA: Benchmark and Psychology-Inspired Model
Baoliang Chen, Changsheng Gao, Lingyu Zhu 0006, Liang Xie 0013, Hanwei Zhu, Zhijian Hao
QoMEX6
2026 One-Iteration ISP Controller for Real-Time Machine Vision
abstract
Conventional image signal processing (ISP) control algorithms based on human visual perception are insufficient for the demands of modern machine vision systems. Although recent learning-based methods have attempted to adapt ISP hyperparameters for specific vision tasks, their high latency and hardware cost of iterative optimization hinder deployment on edge devices such as those used in autonomous driving. To address these limitations, this paper proposes a real-time machine vision system through algorithm–hardware co-design. First, we introduce a one-iteration learning framework to avoid iterative optimization, significantly reducing latency for real-time use. Second, we propose a hardware-friendly controller, RasterNet, specifically tailored for raster-scanning sensor dataflow, eliminating redundant computation. Third, we present a pipelined ISP controller architecture incorporating branch and chroma time division multiplexing techniques to minimize the number of processing elements, achieving a compact and efficient design. Experiments demonstrate that the proposed system achieves superior object detection accuracy on resource-constrained platforms, with real-time performance reaching 70 FPS on FPGA and 224 FPS on ASIC implementations.
Zhijian Hao, Qi Zheng 0004, Ruoxi Zhu, Shuocheng Wang, Honglei Chen, Wenzhong Bao, Hongkai Xiong, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.2
2025 Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image Compression
abstract
Leveraging the generative power of diffusion models, generative image compression has achieved impressive perceptual fidelity even at extremely low bitrates. However, current methods often neglect the non-uniform complexity of images, limiting their ability to balance global perceptual quality with local texture consistency and to allocate coding resources efficiently. To address this, we introduce the Map-guided Masking Realism Image Diffusion Codec (MRIDC), designed to optimize the trade- off between local distortion and global perceptual quality in extreme-low bitrate compression. MRIDC integrates a vector-quantized image encoder with a diffusion-based decoder. On the encoding side, we propose a Map-guided Latent Masking (MLM) module, which selectively masks elements in the latent space based on prior information, allowing adaptive resource allocation aligned with image complexity. On the decoding side, masked latents are completed using the Bidirectional Prediction Controllable Generation (BPCG) module, which guides the constrained generation process within the diffusion model to reconstruct the image. Experimental results show that MRIDC achieves state-of-the-art perceptual compression quality at extremely low bitrates, effectively preserving feature consistency in key regions and advancing the rate-distortion-perception performance curve, establishing new benchmarks in balancing compression efficiency with visual fidelity. Our code can be found at https://github.com/xjc97/mridc.
Jinchang Xu, Zhe Li 0081, Peidong Jia, Guoqing Xiang, Zhijian Hao, Shanghang Zhang
CVPR8
2025 A High-Precision and Low-Cost Approximate Transform Accelerator for Video Coding
abstract
The introduction of multiple transform types in the Versatile Video Coding (VVC) standard has yielded notable encoding gains but also imposed considerable computational burdens. Existing transform circuits of different types are typically implemented separately due to their independence, leading to substantial hardware overhead. To address this, we explore the relationship between Discrete Cosine Transform Type-2 (DCT2) and Discrete Sine Transform Type-7 (DST7) matrices and reveal a prominent diagonal aggregation phenomenon in their transfer matrix. Based on this insight, the least-squares method is applied to optimize the transfer matrix sparsity, achieving a high-precision, low-cost approximate conversion from DCT2 to DST7. Furthermore, we optimize DCT2 computation by proposing an elaborate matrix decomposition approach that allows a lightweight shift-adder unit to efficiently generate all required product terms across varying sizes. Leveraging these algorithmic optimizations, we implement a highly reusable and area-efficient approximate transform accelerator that supports sizes from 4 to 32 points and accommodates three types in VVC. Experimental results demonstrate that the proposed accelerator achieves over 44% reduction in circuit resource consumption with negligible BD-BR performance loss of just $\mathbf{0. 5 3 \%}$, maintaining processing capabilities up to $8 K \text{@} 57 \mathrm{fps}$.
Zhijian Hao, Chenlong He, Qi Zheng 0004, Shushi Chen, Jinchang Xu, Yue Hao 0001, Xiaohua Ma 0001
DAC1
2025 MAS-ISP: A Proxy-Free Online Hyperparameter Optimization Framework for ISP Hardware System
abstract
The rapid advancement of visual autonomous systems, especially in autonomous driving, underscores the critical role of Image Signal Processors (ISPs) as they convert RAW sensor data into RGB images suited for visual interpretation. Traditional ISPs rely on tuning hyperparameters to adapt to varying imaging conditions; however, the vast parameter space and intricate tuning process pose significant challenges for realtime autonomous applications. Existing autonomous ISP hyperparameter optimization methods rely largely on offline or proxybased online tuning, limiting their accuracy and responsiveness to real-time environmental changes. In response, we propose an online ISP hyperparameter optimization framework based on Deep Reinforcement Learning (DRL), marking the first proxyfree, real-time optimization approach. Our design exhibits a master-slave Multi-Agent System (MAS), enabling rapid and cooperative parameter optimization with improved inter-frame consistency. Furthermore, we design the MAS-ISP automated visual system, incorporating innovative hardware designs such as Strip Convolution Kernel and Stride-Aware Dual-Buffer Memory, which drastically reduce resource consumption in CNN hardware. MAS-ISP achieves 1080P@75FPS/240FPS on FPGA/ASIC platforms, supporting real-time and reliable visual systems.
Zhijian Hao, Ruoxi Zhu, Qi Zheng 0004, Shuocheng Wang, Shushi Chen, Leilei Huang, Jun Tao 0001, Yibo Fan
DAC3
2025 Unicorn: Unified Neural Image Compression with One Number Reconstruction
abstract
Prevalent lossy image compression schemes can be divided into: 1) explicit image compression (EIC), including traditional standards and neural end-to-end algorithms; 2) implicit image compression (IIC) based on implicit neural representations (INR). The former is encountering impasses of leveling off bitrate reduction at a cost of tremendous complexity while the latter suffers from excessive smoothing quality as well as lengthy decoder models. In this paper, we propose an innovative paradigm, which we dub Unicorn (Unified Neural Image Compression with One Nnumber Reconstruction). By conceptualizing the images as index-image pairs and learning the inherent distribution of pairs in a subtle neural network model, Unicorn can reconstruct a visually pleasing image from a randomly generated noise with only one index number. The neural model serves as the unified decoder of images while the noises and indexes corresponds to explicit representations. As a proof of concept, we propose an effective and efficient prototype of Unicorn based on latent diffusion models with tailored model designs. Quantitive and qualitative experimental results demonstrate that our prototype achieves significant bitrates reduction compared with EIC and IIC algorithms. More impressively, benefitting from the unified decoder, our compression ratio escalates as the quantity of images increases. We envision that more advanced model designs will endow Unicorn with greater potential in image compression. The code will be made publicly available upon publication.
Qi Zheng 0004, Haozhi Wang, Zihao Liu 0015, Zhijian Hao, Bu Chen, Min Li 0033, Rui Wan, Peiye Liu, Yanheng Lu, Dimin Niu, Jinjia Zhou, Minge Jing, Yibo Fan
ACM Multimedia5
2025 A Novel Transform Accelerator With Fast Kernel Selection and Efficient Transform Circuit
abstract
The introduction of multiple transform types into the Versatile Video Coding (VVC) standard has yielded notable encoding gains but also resulted in substantial computational burdens, posing two critical challenges for hardware implementation: fast kernel selection and efficient transform computation design. Existing studies typically address these challenges in isolation, lacking a holistic solution for VVC transform coding. In this paper, we presents a groundbreaking transform accelerator that unifies transform kernel selection and multiple transform circuit within a single framework. In terms of algorithms, driven by mechanistic analysis, we propose a decision tree-based kernel selection algorithm that ensures both high decision accuracy and computational efficiency. Additionally, we design a transfer matrix-based approximation algorithm for Discrete Sine Transform Type-7 and a matrix decomposition-based improved computation for Discrete Cosine Transform Type-2, significantly reducing the computational complexity. On the hardware front, we implement a high-precision and area-efficient transform accelerator, which integrates highly pipelined kernel selection and transform computation architectures. With multiple reuse and parallelism strategies, the accelerator demonstrates substantial resource efficiency advantages. Experimental results reveal that the proposed accelerator achieves a circuit resource reduction of over 44% with a slight performance degradation, while maintaining processing capabilities up to 8K@57 fps. To the best of our knowledge, this is the first comprehensive hardware solution for VVC transform coding that jointly addresses the challenges of kernel selection and transform circuit design.
Zhijian Hao, Chenlong He, Qi Zheng 0004, Jinchang Xu, Peijun Ma, Xiaohua Ma 0001, Yue Hao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 Affine Motion Estimation Hardware Implementation With 51.7%/67.5% Internal Bandwidth Reduction for Versatile Video Coding
abstract
Versatile Video Coding (VVC) employs Affine Motion Compensation (AMC) to process scenes with high-order motion. To improve AMC efficiency, the Affine Motion Estimation (AME) process based on the gradient-based iterative algorithm (GIA) and block match algorithm (BMA) is introduced to the VVC Test Model (VTM). However, the AME process is highly complex and difficult for hardware implementation in real-time applications. In this context, this paper proposes a hardware-friendly AME algorithm and implements the corresponding accelerator. Firstly, the weighted least squares regression is used to reduce the iteration of GIA. Then an iteration-free search scheme is proposed to remove the search dependence during the GIA and BMA process. In addition, a motion vector clamping mechanism and four-level memory organization are proposed to solve the problem of reference pixel reading conflict, which reduces 51.7% and 67.5% internal bandwidth of the AME accelerator. Compared with the default AME process of VTM 16.0, experimental results show that the proposed algorithm reduces AME run time by 81.63% while the corresponding Bjontegaard Delta Bit Rate (BDBR) loss is only 0.492%. The proposed AME accelerator can flexibly support AME search tasks in various configurations. Synthesized with the TSMC 28nm process, the proposed architecture has a gate count of 1313K and a power consumption of 156.83 mW. It can achieve$7680\times [email protected]~30fps and the corresponding BDBR loss is 0.492%~1.835%.
Shushi Chen, Leilei Huang, Zhao Zan, Zhijian Hao, Hao Zhang 0126, Xiaoxiang Chen, Minge Jing, Xiaoyang Zeng, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.4
2025 Hardware Implementation of a High-Accuracy and High-Throughput Rate Estimation Unit for VVC Residual Coding
abstract
In High Efficiency Video Coding standard, rate estimation based on context-based adaptive binary arithmetic coding (CABAC) typically achieves high accuracy. However, due to serial data dependencies, hardware implementation solutions suffer from lower throughput. When it comes to the latest generation video coding standard, namely Versatile Video Coding (VVC), the increased data dependency and computational complexity during the coding process pose more challenges for the hardware design of rate estimation. To solve these problems, this paper presents a hardware implementation of high-accuracy and high-throughput rate estimation unit for VVC. In terms of throughput improvement, we propose two optimization algorithms to eliminate the majority of data dependencies in coefficient coding with nearly negligible loss in Bjontegaard Delta (BD)-rate performance. To save hardware resources, we introduce a rate estimation table compression algorithm and an optimized local statistical information storage strategy. Based on these optimizations, we present a hardware implementation for the rate estimation unit and a parallel scheme for the rate-distortion optimization process. The proposed algorithm shows an increase of 0.29% in the BD-rate compared to the VVC test model 19.2. Synthesis results show that the proposed design supports real-time coding of$7680\times 4320$@30fps at 500MHz operating frequency. These results indicate that our proposed design performs well in terms of BD-rate performance and throughput. To the best of our knowledge, this is the first hardware implementation of rate estimation for VVC.
Leilei Huang, Wei Li 0257, Zhijian Hao, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.5
2024 Auto-ISP: An Efficient Real-Time Automatic Hyperparameter Optimization Framework for ISP Hardware System
abstract
Image Signal Processor (ISP) is widely used in intelligent edge devices across various scenarios. The intricate and time-consuming tuning process demands substantial expertise. Current AI-based auto-tuning operates discretely offline, relying on predefined scenes with human intervention, leading to inconvenient manipulation, with potentially fatal impacts on downstream tasks in unforeseen scenes. We propose a real-time automatic hyperparameter optimization ISP hardware system to address real-world scenarios. Our design features a tri-step framework and a hardware accelerator, demonstrating superior performance in human and computer vision tasks, even in real-time unforeseen scenes. Experiments showcase its practicality, achieving 1080P@75FPS/240FPS in FPGA/ASIC, respectively.
Zihao Liu 0015, Ruoxi Zhu, Qi Zheng 0004, Zhijian Hao, Tao Liu 0023, Jun Tao 0001, Yibo Fan
DAC6
2024 CTU-Level Adaptive Quantization Method Joint with GOP based Temporal Filter for Video Coding
abstract
Both Versatile Video Coding (VVC) and High Efficiency Video Coding (HEVC) introduce Group of Pictures (GOP) based temporal filter (GBTF) as a pre-filter to improve compression performance. While numerous efforts have been made to optimize GBTF, there is a limited amount of research that explicitly addresses why GBTF could improve compression performance. Additionally, most optimizations have focused on the design of the filter itself, rather than on how to better integrate it with other encoding tools. In this paper, we analyze the reasons behind the superior compression performance of GBTF. Subsequently, we introduce a Coding Tree Unit (CTU)-level adaptive quantization parameter allocation method joint with GBTF to further enhance compression performance for video coding. The experimental results demonstrate that, for VVC, our method provides Bjontegaard delta bit rate (BD-BR) savings of 2.0% for Peak Signal-to-Noise Ratio (PSNR) and 4.0% for Structural Similarity index (SSIM). Furthermore, for HEVC, our method provides BD-BR savings of 3.5% for PSNR and 7.6% for SSIM.
Chenlong He, Xiaoxiang Chen, Zhijian Hao, Chao Liu 0027, Xiaoyang Zeng, Yibo Fan
ISCAS4
2024 A High Compression Efficiency Hardware Encoder for Intra and Inter Coding With 4K@30fps Throughput
abstract
The promotion of the HEVC standard has significantly alleviated the burden of network transmission and video storage. However, its inherent complexity and data dependencies pose a significant challenge in achieving high compression efficiency hardware encoder. To tackle this challenge, we propose several hardware-oriented algorithms and achieve a hardware encoder supporting both intra and inter coding. In terms of algorithms, our optimizations focus on intra mode decision, motion estimation (ME), rate estimation, and merge mode estimation. These optimizations reduce the computational complexity and address the data dependencies within and between encoder modules while maintaining an acceptable compression efficiency. As for hardware, we propose an encoder architecture that supports not only 35 intra prediction modes but also ME with an extensive search range of [±64, ±64]. The uniform$4\times 4$engine, 2-D data reuse, and timing schedule for intra and inter coding are presented in this architecture to optimize the hardware resource consumption and throughput. Compared with HM 15.0, the proposed hardware-oriented algorithms lead to a 1.88% and 14.57% increase in BD-Rate under the configurations of all intra and low delay P, respectively. Notably, the BD-Rate outperforms all existing hardware encoders supporting 4K resolution. In a GF 28nm fabrication process, the hardware design achieves a clock frequency of 550MHz, supporting 4K@30fps throughput with a hardware gate count of 3154K and memory usage of 1.02MB, and the proposed architecture demonstrates substantial advantages in terms of area, throughput, and power compared to other studies.
Guohao Xu, Leilei Huang, Zhijian Hao, Wei Li 0257, Shiyan Yi, Xiaoyang Zeng, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.3
2023 Controlling Collision-Induced Aggregations in a Swarm of Micro Bristle Robots
abstract
Systematically designing local interaction rules to achieve collective behaviors in robot swarms is a challenging endeavor, especially in micro robots, where size restrictions imply severe sensing, communication, and computation limitations. In such robot swarms, performing useful functions is often preconditioned on the formation of high-density aggregations which can facilitate collective signaling and information sharing. In this article, we present a systematic approach to control aggregation behaviors by leveraging the physical interactions in a swarm of 300 3-mm vibration-driven micro bristle robots that we designed and fabricated. We demonstrate the ability to control the degree of aggregation by varying the motility characteristics of the robots through global vibration frequency and amplitude inputs, after comprehensive characterization, modeling, and simulation of the locomotion dynamics and robot interactions. To quantify the degree of aggregation, we also introduce a new metric, the motility-induced phase separation index index, which unlike many existing methods does not require a scenario-specific tuning of parameters. Our investigations reveal how physics-driven interaction mechanisms can be exploited to achieve desired behaviors in minimally equipped robot swarms and highlight the specific ways in which hardware and software developments aid in the achievement of collision-induced aggregations.
Zhijian Hao, Siddharth Mayya, Gennaro Notomista, Seth Hutchinson 0001, Magnus Egerstedt, Azadeh Ansari
IEEE Trans. Robotics1
2023 A Reconfigurable Multiple Transform Selection Architecture for VVC
abstract
Video coding plays an important role in the highly information-based world as videos contribute the largest part of network traffic. The latest video coding standard Versatile Video Coding (VVC) introduces a new transform scheme multiple transform selection (MTS), which brings considerable coding gains at the expense of high coding complexity. In this article, we propose a reconfigurable MTS architecture that supports all transform types in VVC with square and rectangular sizes ranging from$4\times $4 to 32$\times32$. Firstly, we explore the features of three types of transform matrices and extract the features that are beneficial to designing a unified architecture. Then, we present an improved calculation scheme for general transforms, where the transform matrix is decomposed into two simpler matrices to increase the similarity and decrease the complexity of matrices involved in three types of transform operations. Thanks to the improved calculated scheme, a unified shift-adder unit (SAU) is designed and highly reused by different types. Moreover, we provide a twirling two-point splicing (T2S) scheme to improve reusability and deal with issues of data mismatch when conducting discrete cosine transform (DCT)-II of different sizes. As a consequence, an architecture with constant throughput of 32 pixels/cycle is implemented and specified in Verilog HDL. The synthesis results indicate that the application specific integrated circuit (ASIC)-based and field-programmable gate array (FPGA)-based hardware architectures achieve significant advantages both in area reduction and power consumption compared to existing methods in the literature.
Zhijian Hao, Heming Sun, Guoqing Xiang, Peng Zhang 0007, Xiaoyang Zeng, Yibo Fan
IEEE Trans. Very Large Scale Integr. Syst.1
2022 Blind Video Quality Assessment via Space-Time Slice Statistics
abstract
User-generated contents (UGC) have gained increased attention in the video quality community recently. Perceptual video quality assessment (VQA) of UGC videos is of great significance for content providers to monitor, process, and deliver massive numbers of UGC videos. Blind video quality prediction of UGC videos is challenging since complex mixtures of spatial and temporal distortions contribute to the overall perceptual quality. In this paper, we develop a simple, effective, and efficient blind VQA framework (STS-QA) based on the statistical analysis of space-time slices (STS) of videos. Specifically, we extract spatio-temporal statistical features along different orientations of video STS, that capture directional global motion, then train a shallow quality predictor. The proposed framework can be used to easily extend any existing video/image quality model to account for temporal or motion regularities. Our experimental results on three publicly available UGC databases demonstrate that our proposed STS-QA model can significantly boost prediction performance compared to baselines. The code will be released at: https://github.com/uniqzheng/STS_BVQA.
Qi Zheng 0004, Zhengzhong Tu, Zhijian Hao, Xiaoyang Zeng, Alan C. Bovik, Yibo Fan
ICIP3
2022 An Area-efficient Unified Transform Architecture for VVC
abstract
The next-generation video coding standard Versatile Video Coding (VVC) adopts Multiple Transform Selection (MTS) to the transform module, improving coding efficiency at the expense of high computational complexity. Compared to High Efficiency Video Coding (HEVC), VVC supports larger sizes and extends the transform types to Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII, and DCT-VIII. This paper presents an area-efficient unified architecture for VVC. To reduce the area consumption, we propose an optimized calculation scheme for general transformations where the transform matrix is decomposed into two simpler matrices named the Low-value matrix and the Error matrix. Based on the decomposition algorithm, Shift-Addition Units (SAUs)-based circuits are designed to conduct matrix multiplication and can be reused by three types. As a result, this unified architecture is capable of performing all types and sizes in VVC. The synthesis results indicate that this architecture achieves an area reduction of 37.9% $\sim$ 72.2% compared with related works for 32-point transforms.
Zhijian Hao, Qi Zheng 0004, Yibo Fan, Guoqing Xiang, Peng Zhang 0007, Heming Sun
ISCAS1
2020 Maneuver at Micro Scale: Steering by Actuation Frequency Control in Micro Bristle Robots*
abstract
This paper presents a novel steering mechanism, which leads to frequency-controlled locomotion demonstrated for the first time in micro bristle robots. The miniaturized robots are 3D-printed, 12 mm × 8 mm × 6 mm in size, with bristle feature sizes down to 400 µm. The robots can be steered by utilizing the distinct resonance behaviors of the asymmetrical bristle sets. The left and right sets of the bristles have different diameters, and thus different stiffnesses and resonant frequencies. The unique response of each bristle side to the vertical vibrations of a single on-board piezoelectric actuator causes differential steering of the robot. The robot can be modeled as two coupled uniform bristle robots, representing the left and the right sides. At distinct frequencies, the robots can move in all four principal directions: forward, backward, left and right. Furthermore, the full 360◦2D plane can be covered by superimposing the principal actuation frequency components with desired amplitudes. In addition to miniaturized robots, the presented resonance-based steering mechanism can be applied over multiple scales and to other mechanical systems.
Zhijian Hao, DeaGyu Kim, Ali Reza Mohazab, Azadeh Ansari
ICRA1