Leilei Huang

dblp:164/5591 · DBLP profile ↗
← Back
42ranked-venue papers
0as first author
36since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A Multiplier-Free Similarity Metric for Token Merging in Vision Transformers
Yupeng Gui, Ruiqi Tang, Shichen Peng, Han He, Leilei Huang, Yibo Fan
ISCAS6
2026 HHGC: High-Throughput High-Accuracy Gradient Data Compression
Yupeng Gui, Xinglin Wei, Yanwu Liu, Leilei Huang, Yibo Fan
ISCAS4
2026 A Hybrid Quantization Method for FMCW Radar Data Compression
Yubo Guo, Leilei Huang, Chunqi Shi, Jinghong Chen, Runxi Zhang
ISCAS2
2026 HLC: A High-Quality Lightweight Mezzanine Codec Featuring High-Throughput Palette
Chenlong He, Leilei Huang, Wei Li 0257, Hanyang Cui, Zhijian Hao, Xiaoyang Zeng, Yibo Fan
ISCAS2
2026 An Area-latency-balanced Hardware Design for the Reference Pixel Management of VVC Intra Coding
Chengkang Huang, Leilei Huang, Taoyu Zhang, Shuocheng Wang, Yibo Fan
ISCAS2
2026 SENTRY-VQA: A Region-Aware Framework for Surveillance Video Quality Assessment
Chenlong He, Leilei Huang, Minge Jing, Yibo Fan
ISCAS5
2026 Near-field Radar Imaging Acceleration Strategy Using Wavelet Transform and Multiscale Processing
Tianxing Lu, Yingjian Hao, Jingqian Wang 0003, Leilei Huang, Chunqi Shi, Jinghong Chen, Runxi Zhang
ISCAS5
2026 A Compact Broadband Power Amplifier with High Efficiency and <1% Power-Added Efficiency Variation over 24-36 GHz
Yunhan Qian, Chunqi Shi, Leilei Huang, Jinghong Chen, Runxi Zhang
ISCAS4
2026 An Ultra-Low-Power, High-Output-Power and High-Sensitivity Bioelectronic Transceiver
Xiaoyuan Wu, Kangjie Zhao, Chunqi Shi, Leilei Huang, Jinghong Chen, Runxi Zhang
ISCAS5
2026 Near-Field Radar Imaging Motion Compensation Solution Based On STFT and PGA
Tianxing Lu, Yingjian Hao, Jingqian Wang 0003, Leilei Huang, Chunqi Shi, Jinghong Chen, Runxi Zhang
ISCAS5
2026 A 77.9%-Cycle-Reduced Bubble-Removing Strategy for Hardware RDO Supporting QTMTT in VVC
abstract
The introduction of the Versatile Video Coding (VVC) standard is dedicated to meeting the increasing demand for high-resolution and high-quality video. However, the novel partitioning method named Quad-Tree Plus Multi-Type Tree (QTMTT) significantly increases computational complexity and data dependency, resulting in more pipeline bubbles, lower hardware efficiency, and degraded throughput. To address this issue, we analyze the hardware data dependencies, categorize four different types of pipeline bubbles, and propose a bubble-removing strategy for the Rate Distortion Optimization (RDO) module with MTT depth of 1. To be more specific, we first propose an efficient partition scheduling scheme based on the characteristics of QTMTT partitions. Then we redesign the transpose memory used in 2D transformation to efficiently handle the blocks of different sizes introduced by QTMTT. These two strategies achieve a 42.3% reduction in hardware cycles. In addition, for I frames, we further propose a hardware-oriented partition pruning algorithm that can co-operate with the proposed architecture, achieving a 42.3%~77.9% reduction in hardware cycles with only 0%~1.21% BD-Rate loss compared to the VTM-23.4. The proposed hardware architecture is implemented in GF 28nm technology, supporting up to 4K@40fps throughput at 500MHz with a hardware cost of only 3259 K gates and 63.47 KB on-chip memory, demonstrating competitive compression performance, high hardware efficiency, and outstanding throughput.
Chengkang Huang, Leilei Huang, Taoyu Zhang, Wei Li 0257, Shuocheng Wang, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.2
2025 MAS-ISP: A Proxy-Free Online Hyperparameter Optimization Framework for ISP Hardware System
abstract
The rapid advancement of visual autonomous systems, especially in autonomous driving, underscores the critical role of Image Signal Processors (ISPs) as they convert RAW sensor data into RGB images suited for visual interpretation. Traditional ISPs rely on tuning hyperparameters to adapt to varying imaging conditions; however, the vast parameter space and intricate tuning process pose significant challenges for realtime autonomous applications. Existing autonomous ISP hyperparameter optimization methods rely largely on offline or proxybased online tuning, limiting their accuracy and responsiveness to real-time environmental changes. In response, we propose an online ISP hyperparameter optimization framework based on Deep Reinforcement Learning (DRL), marking the first proxyfree, real-time optimization approach. Our design exhibits a master-slave Multi-Agent System (MAS), enabling rapid and cooperative parameter optimization with improved inter-frame consistency. Furthermore, we design the MAS-ISP automated visual system, incorporating innovative hardware designs such as Strip Convolution Kernel and Stride-Aware Dual-Buffer Memory, which drastically reduce resource consumption in CNN hardware. MAS-ISP achieves 1080P@75FPS/240FPS on FPGA/ASIC platforms, supporting real-time and reliable visual systems.
Zhijian Hao, Ruoxi Zhu, Qi Zheng 0004, Shuocheng Wang, Shushi Chen, Leilei Huang, Jun Tao 0001, Yibo Fan
DAC9
2025 A Multiplier-Balanced and Area-Efficient Architecture for Low-Frequency Non-Separable Secondary Transform
abstract
The existing hardware architectures for Low-Frequency Non-Separable Secondary Transform (LFNST) in VVC suffer from significant resource consumption of multipliers, along with relatively low utilization of these multipliers. In this context, this paper proposes a core unit with balanced multiplier utilization and an area-efficient overall architecture for LFNST. The proposed core unit is based on analyzing the average number of effective multiplications per sample across all transform units, ensuring balanced multiplier utilization and avoiding ineffective multiplications. Additionally, general multipliers can be replaced with a fused circuit that consists only of adders, shifters, and multiplexers. Through the proposed coefficient approximation scheme (CAS), the number of adders in the fused circuit is limited to one, significantly reducing resource consumption. The synthesis results indicate that the proposed design achieves a 72% reduction in normalized area compared to the state-of-the-art work. Furthermore, the proposed CAS can reduce the area of multiplier core unit by 22%, with negligible Bjøntegaard Delta-Rate loss. The proposed architecture supports processing 7680×4320@64fps videos when working at 200 MHz.
Leilei Huang, Chengkang Huang, Bingjing Hou, Yibo Fan
ISCAS2
2025 A 13.6-17.7 GHz Sub-Harmonic Injection-Locked SSPLL with 74-fs RMS Jitter
abstract
This paper presents a sub-harmonic injection-locked sub-sampling PLL (SIL-SSPLL) that realizes a peaking-free jitter transfer. The use of an injection-locked oscillator (ILO) in SSPLL allows the PLL open-loop transfer function to achieve a higher phase margin, thereby enhancing the jitter performance. The study also analyzes the effects of injection locking strength on noise rejection for both the reference signal and the VCO showing that optimal jitter performance can be achieved by properly setting the injection strength. A 13.6 to 17.7 GHz SIL-SSPLL prototype fabricated in a 40 nm CMOS technology demonstrates 74 fs RMS integrated jitter at 15.8 GHz while consuming 22.6 mW of power, leading to an excellent figure of merit (FoM) of -249.1 dBc/Hz.
Yuri Lu, Chunqi Shi, Leilei Huang, Runxi Zhang, Jinghong Chen
ISCAS3
2025 WiVir: Exploiting WiFi-based Virtual Antenna Array for Passive Stationary Human Localization
abstract
WiFi passive positioning enables precise target tracking without additional devices, offering a cost-effective solution for elder care, security, and smart homes. However, it faces challenges such as distinguishing stationary humans from static objects and limited angular resolution due to existing hardware constraints. In this paper, we present WiVir, a virtual antenna array system based on commodity WiFi device for passive localization of stationary humans. The system addresses the challenges of channel state information (CSI) measurement errors and innovatively applies the concept of virtual antenna array (VAA) in WiFi-MIMO modes. It overcomes the limitation of the number of antennas in commercial WiFi Network Interface Controller (NIC) and significantly reduces the beam width of beamforming. Beamforming is used to obtain the angle of arrival (AOA), and the beam is analyzed in the frequency domain to generate a two-dimensional angular frequency image. Experimental results show that WiVir achieves 98.4% accuracy and strong robustness in indoor environments, providing precise positioning within a certain angle range.
Qitong Wang 0006, Leilei Huang, Chunqi Shi, Jinghong Chen, Runxi Zhang
ISCAS3
2025 DPP-MP: An Area-Efficient Digital Predistortion Model for Quadrature Digital Transmitters
abstract
Although quadrature digital transmitters (DTXs) are particularly useful in wideband scenarios, the interaction between I and Q paths would lead to a non-negligible distortion. Unfortunately, traditional DPD models have not performed well in solving this kind of distortion. For example, memory polynomial (MP) model uses signal amplitude as the fundamental term which is not accurate enough to solve this issue, while 2-D lookup table (2D-LUT) model requires a large number of entries which leads to significant hardware resource consumption. In view of this, we propose an area-efficient dual-path piecewise memory-polynomial (DPP-MP) model, which improves model accuracy while reducing hardware resource consumption. To further reduce the area cost, the model is pruned to 8 terms only and 4 of them are reused. To reduce the logic delay, a new squared multiplier is proposed to replace the default one, which reduces delay by 13.51% and power consumption by 2.56%. This model achieves 36.86% and 66.98% area optimization compared to MP and 2D-LUT, respectively. When applied to a 9-bit quadrature DTX with a 40 MHz 256-QAM signal, this model improves the error vector magnitude (EVM) from -16.38 dB to -35.89 dB with a target average power (Pavg) of 21.07 dBm.
Kangjie Zhao, Wangdong Xie, Guozhen Wu, Leilei Huang, Chunqi Shi, Jinghong Chen, Runxi Zhang
ISCAS5
2025 A Parallel Acceleration Strategy for Large Aperture Radar Imaging and its Hardware Implementation
abstract
A parallel acceleration strategy and hardware implementation scheme for SAR imaging is proposed to accelerate imaging in high bandwidth, large aperture, and high-density scenarios. Based on traditional SAR imaging, this strategy divides the target image into multiple local images, selecting appropriate radar data within a suitable aperture range based on the locations of each local component. Multiple SAR imaging units operate in parallel, and the local images are stitched together in their original locations to achieve accelerated imaging after cropping. This paper designs the hardware for the SAR imaging units, with its core FFT operations designed as a low-cost reusable structure. By reusing and expanding this hardware unit, the acceleration system can be built with relatively low hardware resources and minimal loss in imaging quality, demonstrating significant application potential for THz radar imaging and high-resolution large-aperture security inspections.
Yukun Cheng, Chunqi Shi, Leilei Huang, Jinghong Chen, Runxi Zhang
ISCAS5
2025 A Robust Data Compression Engine Dedicated For FMCW Radar RDM
abstract
Time-domain signals are typically transformed using a 2D-FFT to generate the Range Doppler Map (RDM) in FMCW radar data processing. Subsequently, 2D-CFAR is applied to the RDM for target detection, while Direction of Arrival (DOA) estimation is utilized to obtain the angular information of the target, thus enabling the acquisition of the target’s spatial information. As the demand for accurate target information increases, the volume of data in each RDM frame also increases. This escalation necessitates significant hardware resources and an area to store the RDM in on-chip SRAM or a large bandwidth to accommodate it in off-chip DDRs. Furthermore, DDR controllers require substantial on-chip resources. To address this challenge, we propose a compression algorithm named APFG, which comprises Amplitude-Phase Transformation (APT), Fix-to-Float (Fix2Float), and Exponential Sharing (ESH). Experimental results demonstrate that this algorithm is applicable across various scenarios, CFAR types, and CFAR configurations, achieving high compression ratios while ensuring the accuracy of CFAR and DOA estimates. Specifically, it achieves a compression ratio of less than 14% when Pceis below 1% and a compression ratio of nearly 30% when Pdeis around 2%.
Zhiluo Zhang, Zhixin Yin, Leilei Huang, Chunqi Shi, Jinghong Chen, Runxi Zhang
ISCAS4
2025 A Frequency-Domain Transfer Model for Predicting FM Error of FMCW Radar Chirp Generators
abstract
This paper proposes a frequency-domain model for fast and accurate prediction of frequency modulation (FM) error in frequency-modulated continuous wave (FMCW) radar chirp generators. Obtaining FM error from the output frequency of a phase-locked loop (PLL) requires lengthy simulation times due to the simulator’s limited frequency quantization accuracy, which can result in discrepancies between simulated and actual performance. To address this issue, the feedback signal (DIV) sent to the phase-frequency detector (PFD) is employed in this study to quickly and accurately determine the FM error through spectrum analysis. During frequency sweeps, the feedback signal with a relatively constant frequency receives the modulating signal directly. To validate the effectiveness of the proposed model, a nested PLL-based frequency synthesizer operating from 24-27.52 GHz is designed in a 55-nm CMOS process. Measurement results show good agreement with those obtained from the model, with the difference between the simulated and measured root mean square (RMS) FM error being less than 22 kHz.
Kaige Wang, Dalin Li, Chunqi Shi, Leilei Huang, Runxi Zhang, Jinghong Chen
ISCAS5
2025 A Hardware-Friendly Lightweight Partition Decision Algorithm for VVC Intra and Inter Coding
abstract
The Versatile Video Coding (VVC) standard notably enhances encoding efficiency with the Quad-Tree plus Multi-Type Tree (QTMTT) partition structure. However, the complex QTMTT tool presents substantial challenges in both software and hardware implementation. To overcome those challenges, this paper introduces a hardware-friendly partition decision algorithm for VVC intra and inter coding. Firstly, we propose a lightweight backbone network to extract partition-aware features. Secondly, we employ a Quantisation Parameter (QP) fusion network to regulate the impact of QPs on the partition structure. Additionally, we apply a top-down threshold-driven post-processing algorithm, in which improbable partition types are removed to directly derive the unique partition structure. Experiments show that our method not only exceeds the previous state-of-the-art work in BD-BR performance, but also shows sufficient hardware-friendly characteristics. To the best of our knowledge, this work is among the earliest to comprehensively discuss and implement a hardware-friendly partition decision algorithm.
Zhao Zan, Leilei Huang, Shushi Chen, Xiaoyang Zeng, Yibo Fan
IEEE Signal Process. Lett.2
2025 Affine Motion Estimation Hardware Implementation With 51.7%/67.5% Internal Bandwidth Reduction for Versatile Video Coding
abstract
Versatile Video Coding (VVC) employs Affine Motion Compensation (AMC) to process scenes with high-order motion. To improve AMC efficiency, the Affine Motion Estimation (AME) process based on the gradient-based iterative algorithm (GIA) and block match algorithm (BMA) is introduced to the VVC Test Model (VTM). However, the AME process is highly complex and difficult for hardware implementation in real-time applications. In this context, this paper proposes a hardware-friendly AME algorithm and implements the corresponding accelerator. Firstly, the weighted least squares regression is used to reduce the iteration of GIA. Then an iteration-free search scheme is proposed to remove the search dependence during the GIA and BMA process. In addition, a motion vector clamping mechanism and four-level memory organization are proposed to solve the problem of reference pixel reading conflict, which reduces 51.7% and 67.5% internal bandwidth of the AME accelerator. Compared with the default AME process of VTM 16.0, experimental results show that the proposed algorithm reduces AME run time by 81.63% while the corresponding Bjontegaard Delta Bit Rate (BDBR) loss is only 0.492%. The proposed AME accelerator can flexibly support AME search tasks in various configurations. Synthesized with the TSMC 28nm process, the proposed architecture has a gate count of 1313K and a power consumption of 156.83 mW. It can achieve$7680\times [email protected]~30fps and the corresponding BDBR loss is 0.492%~1.835%.
Shushi Chen, Leilei Huang, Zhao Zan, Zhijian Hao, Hao Zhang 0126, Xiaoxiang Chen, Minge Jing, Xiaoyang Zeng, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.2
2025 An 8K@120fps Advanced Entropy Coding Hardware Design for AVS3
abstract
The third generation audio video coding standard (AVS3) is the latest video coding standard developed by the China AVS working group. The advanced entropy coding (AEC) tool in AVS3 has critical bin-to-bin data dependencies leading to difficulties in parallelization. The use of a 16384-entries lookup table (LUT) in the AEC context update algorithm poses challenges in balancing area and performance. To address these issues, we propose a high-performance, area-efficient hardware design. Firstly, we introduce a novel multicycle-path parallel architecture and optimize area cost through hardware reuse. Next, we construct a context modeling processing unit to replace the large LUT, significantly reducing area. Finally, we propose a new LUT-free dual-context modeling processing unit, effectively resolving critical paths introduced by parallel context conflicts. As a result, our design processes 2.63 bins per cycle. The synthesis results based on the GlobalFoundries’ 28nm process indicate that its maximum frequency is 990MHz, with an overall throughput of 2604 Mbin/s. Compared to state-of-the-art designs, our design leads in performance by 24.7% while reducing area by 28%.
Wei Li 0257, Leilei Huang, Chenlong He, Minge Jing, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.2
2025 Hardware Implementation of a High-Accuracy and High-Throughput Rate Estimation Unit for VVC Residual Coding
abstract
In High Efficiency Video Coding standard, rate estimation based on context-based adaptive binary arithmetic coding (CABAC) typically achieves high accuracy. However, due to serial data dependencies, hardware implementation solutions suffer from lower throughput. When it comes to the latest generation video coding standard, namely Versatile Video Coding (VVC), the increased data dependency and computational complexity during the coding process pose more challenges for the hardware design of rate estimation. To solve these problems, this paper presents a hardware implementation of high-accuracy and high-throughput rate estimation unit for VVC. In terms of throughput improvement, we propose two optimization algorithms to eliminate the majority of data dependencies in coefficient coding with nearly negligible loss in Bjontegaard Delta (BD)-rate performance. To save hardware resources, we introduce a rate estimation table compression algorithm and an optimized local statistical information storage strategy. Based on these optimizations, we present a hardware implementation for the rate estimation unit and a parallel scheme for the rate-distortion optimization process. The proposed algorithm shows an increase of 0.29% in the BD-rate compared to the VVC test model 19.2. Synthesis results show that the proposed design supports real-time coding of$7680\times 4320$@30fps at 500MHz operating frequency. These results indicate that our proposed design performs well in terms of BD-rate performance and throughput. To the best of our knowledge, this is the first hardware implementation of rate estimation for VVC.
Leilei Huang, Wei Li 0257, Zhijian Hao, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.2
2025 An Interpolation-Free Fractional Motion Estimation Algorithm and Hardware Implementation for VVC
abstract
Versatile video coding (VVC) introduces multi-type tree (MTT) and larger coding tree unit (CTU) to improve compression efficiency compared to its predecessor High Efficiency Video Coding (HEVC). This leads to higher throughput for fractional motion estimation (FME) to meet the needs of real-time processing. In this context, this article proposes an interpolation-free algorithm based on an error surface to improve the throughput of FME hardware. The error surface is constructed by the rate-distortion costs (RDCs) of the integer motion vector (IMV) and its neighbors. To improve the prediction accuracy, a hardware-friendly RDC estimation strategy is proposed to construct the error surface. The experimental results show that the corresponding Bjontegaard Delta Bit Rate (BDBR) in Random Access (RA), Low Delay P (LDP) and Low Delay B (LDB) configuration increases by only 0.358%, 0.479%, and 0.511% compared with the VVC test model (VTM) 16.0. Compared with the default FME algorithms of VVC, the time cost of FME is reduced by 53.47%, 56.28%, and 54.23%, respectively, in RA, LDP, and LDB configurations. The algorithm is free of iteration and interpolation, which can contribute to low-cost and high-throughput hardware. The proposed architecture can support FME of all coding units (CUs) in a CTU with one layer of MTT under the quaternary tree (QT), and the CU size can vary from$8\times 8$to$128\times 128$. Synthesized using GF 28-nm process, the architecture can achieve$7680\times 4320$@60 fps throughput at 800 MHz, with a gate count of 244 K and power consumption of 76.5 mW. This proposed architecture can meet the real-time coding requirements of VVC.
Shushi Chen, Leilei Huang, Zhao Zan, Xiaoyang Zeng, Yibo Fan
IEEE Trans. Very Large Scale Integr. Syst.2
2024 An 8K@120fps Hardware Implementation for Decoder-Side Motion Vector Refinement in VVC
abstract
Versatile video coding (VVC) introduces many coding tools to improve compression efficiency not only on the encoder side but also on the decoder side. For inter-picture coding, the introduction of decoder-side motion vector refinement (DMVR) puts forward a demand for a motion search engine in the hardware decoder. However, because of the specific mirroring search property in DMVR, the existing hardware design of motion estimation cannot be directly adopted for DMVR. This paper proposes a specific pixel-level sum of absolute differences (SAD) calculation datapath for the 25 mirrored symmetrical search points and also presents a cost-effective design of DMVR search engine under the condition of high-throughput. Finally, the proposed hardware implementation is synthesized with TSMC 22nm process. The measured throughput can reach 8K@120fps at 500MHz, with a total cell area of 52479 μm2and power consumption of 10.82 mW. To the best of our knowledge, this work is the first hardware architecture for the DMVR search process in VVC.
Leilei Huang, Shushi Chen, Yibo Fan
ISCAS2
2024 A PCA Acceleration Algorithm For WiFi Sensing And Its Hardware Implementation
abstract
Currently, we are entering the wearable internet era. The use of WiFi signals to perceive human activities or changes in vital signs has gradually become a topic that people are enthusiastic about. Principal Component Analysis (PCA), as a universal data dimensionality reduction algorithm, has been widely applied in the field of WiFi sensing. It is utilized to expedite data processing time and enhance real-time detection capabilities. This paper proposes an acceleration algorithm for PCA in the field of WiFi sensing along with its corresponding hardware architecture. The experimental results indicate that the accuracy of this algorithm can reach 0.9965, and the processing time is approximately 10 ms. Based on the TSMC 22nm technology, the Design Complier (DC) results show that the data throughput of this hardware architecture can reach 24 Gbps@800M, with a gate count of 7571, and the power consumption is 1.6321mW.
Qitong Wang 0006, Leilei Huang, Chunqi Shi, Runxi Zhang
ISCAS3
2024 Fast Adaptive Loop Filter Algorithm Based on the Optimization of Class Merging
abstract
Adaptive loop filter (ALF) is one of the new tools adopted in the next generation video coding standard Versatile Video Coding (VVC). ALF leads to a performance improvement of 2%∼6% at the expense of high computational complexity and long processing time. Especially when encoding, ALF accounts 5%∼20% for total encoding runtime. To solve this problem, this paper proposes a fast algorithm for ALF by optimizing the process of class merging, which is an important step of rate-distortion optimization (RDO) in ALF. Experimental results indicate that compared to the VVC Test Model (VTM-22.0), this method can reduce the ALF encoding runtime on average 48.16%, 60.65%, 56.15% under all-intra, low-delay P and random access configurations, with only 0.15%, 0.31%, 0.19% increase in luma Bjontegaard Delta-Bit Rate (BD-BR).
Chengkang Huang, Leilei Huang, Shuocheng Wang, Yibo Fan
VCIP2
2024 A High-Throughput and Memory-Efficient Deblocking Filter Hardware Architecture for VVC
abstract
Video coding has become more and more important since high-resolution and high-quality videos have been used in a variety of application areas. Deblocking filter (DBF) is a video coding technology which can improve both video quality and coding efficiency. However, its hardware architecture design suffers from huge computations and high memory requirements. Moreover, the latest Versatile Video Coding (VVC) standard extends DBF with several complex enhancements, which makes the design more difficult. In this paper, a high-throughput and memory-efficient DBF hardware architecture for VVC systems is presented. By analyz-ing the DBF algorithm, we firstly propose a unified filter core to perform edge filtering process with low complexity, and two resource sharing techniques are utilized to reduce hardware costs. Furthermore, we propose a whole DBF architecture to process all the edges in a coding tree unit (CTU). To improve its throughput, we propose novel pre-calculation processing flow and double processing flow to fully utilize pipelining and parallel processing techniques. At the same time, to reduce its memory requirements, we propose four novel data reuse approaches to fully utilize intermediate data reusabilities. Synthesis results show that our proposed hardware architecture can support real-time VVC DBF processing of$7680\times 4320$at 158 frames/s at 500 MHz working frequency. The hardware costs are only 163.2k gate count and three two-port on-chip SRAMs with data width of 128 bits and depth of 32. Compared with other state-of-the-art works for previous standards, our proposed VVC DBF hardware architecture achieves good results in performance, area efficiency and memory efficiency.
Bingjing Hou, Leilei Huang, Ming-e Jing, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.2
2024 A High Compression Efficiency Hardware Encoder for Intra and Inter Coding With 4K@30fps Throughput
abstract
The promotion of the HEVC standard has significantly alleviated the burden of network transmission and video storage. However, its inherent complexity and data dependencies pose a significant challenge in achieving high compression efficiency hardware encoder. To tackle this challenge, we propose several hardware-oriented algorithms and achieve a hardware encoder supporting both intra and inter coding. In terms of algorithms, our optimizations focus on intra mode decision, motion estimation (ME), rate estimation, and merge mode estimation. These optimizations reduce the computational complexity and address the data dependencies within and between encoder modules while maintaining an acceptable compression efficiency. As for hardware, we propose an encoder architecture that supports not only 35 intra prediction modes but also ME with an extensive search range of [±64, ±64]. The uniform$4\times 4$engine, 2-D data reuse, and timing schedule for intra and inter coding are presented in this architecture to optimize the hardware resource consumption and throughput. Compared with HM 15.0, the proposed hardware-oriented algorithms lead to a 1.88% and 14.57% increase in BD-Rate under the configurations of all intra and low delay P, respectively. Notably, the BD-Rate outperforms all existing hardware encoders supporting 4K resolution. In a GF 28nm fabrication process, the hardware design achieves a clock frequency of 550MHz, supporting 4K@30fps throughput with a hardware gate count of 3154K and memory usage of 1.02MB, and the proposed architecture demonstrates substantial advantages in terms of area, throughput, and power compared to other studies.
Guohao Xu, Leilei Huang, Zhijian Hao, Wei Li 0257, Shiyan Yi, Xiaoyang Zeng, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.2
2024 Benchmark Dataset and Pair-Wise Ranking Method for Quality Evaluation of Night-Time Image Enhancement
abstract
Night-time image enhancement (NIE) aims at boosting the intensity of low-light regions while suppressing noises or light effects in night-time images, and numerous efforts have been made for this task. However, few explorations focus on the quality evaluation issue of enhanced night-time images (ENTIs), and how to fairly compare the performance of different NIE algorithms remains a challenging problem. In this paper, we firstly construct a new Real-world Night-Time Image Enhancement Quality Assessment (i.e., RNTIEQA) dataset that includes two typical types of night-time scenes (i.e., extremely low light and uneven light scenes), and carry out human subjective studies to compare the quality of ENTIs obtained by a set of representative NIE algorithms. Afterwards, a new objective ranking method that comprehensively considering image intrinsic and impairment attributes is proposed for automatically predicting the quality of ENTIs. Experimental results on our RNTIEQA dataset demonstrate that the proposed method outperforms the off-the-shelf competitors. Our dataset and code will be released athttps://github.com/Leilei-Huang-work/RNTIEQA-dataset.
Xuejin Wang, Leilei Huang, Hangwei Chen, Qiuping Jiang, ShaoWei Weng, Feng Shao 0001
IEEE Trans. Multim.2
2023 Fast QTMT Partition for VVC Intra Coding Using U-Net Framework
abstract
Versatile Video Coding (VVC) has significantly increased encoding efficiency at the expense of numerous complex coding tools, particularly the flexible Quad-Tree plus Multi-type Tree (QTMT) block partition. This paper proposes a deep learning-based algorithm applied in fast QTMT partition for VVC intra coding. Our solution greatly reduces encoding time by early termination of less-likely intra prediction and partitions with negligible BD-BR increase. Firstly, a redesigned U-Net is recommended as the network’s fundamental framework. Next, we design a Quality Parameter (QP) fusion network to regulate the effect of QPs on the partition results. Finally, we adopt a refined post-processing strategy to better balance encoding performance and complexity. Experimental results demonstrate that our solution outperforms the state-of-the-art works with a complexity reduction of 44.74% to 68.76% and a BD-BR increase of 0.60% to 2.33%.
Zhao Zan, Leilei Huang, Shushi Chen, Zhenghui Zhao, Haibing Yin, Yibo Fan
ICIP2
2023 An Error-Surface-Based Fractional Motion Estimation Algorithm and Hardware Implementation for VVC
abstract
Versatile Video Coding (VVC) introduces more coding tools to improve compression efficiency compared to its predecessor High Efficiency Video Coding (HEVC). For inter-frame coding, Fractional Motion Estimation (FME) still has a high computational effort, which limits the real-time processing capability of the video encoder. In this context, this paper proposes an error-surface-based FME algorithm and the corresponding hardware implementation. The algorithm creates an error surface constructed by the Rate-Distortion (R-D) cost of the integer motion vector (IMV) and its neighbors. This method requires no iteration and interpolation, thus reducing the area and power consumption and increasing the throughput of the hardware. The experimental results show that the corresponding BDBR loss is only 0.47% compared to VTM 16.0 in LD-P configuration. The hardware implementation was synthesized using GF 28nm process. It can support 13 different sizes of CU varying from$128\times 128$to$8\times 8$. The measured throughput can reach 4K@30fps at$400\mathbf{MHz}$, with a gate count of 192k and power consumption of 12.64 mW. And the throughput can reach 8K@30fps at 631MHz when only quadtree is searched. To the best of our knowledge, this work is the first hardware architecture for VVC FME with an interpolation-free strategy.
Shushi Chen, Leilei Huang, Chao Liu 0027, Yibo Fan
ISCAS2
2023 A High-Gain and Low-Noise Mixer with Hybrid $G_{m}$-Boosting for 5G FR2 Applications
abstract
This paper presents a hybrid transconductance ($g_{m}$) boosting technique exploiting both transformer coupling and cross-coupled PMOS pair to improve the conversion gain (CG) and noise figure (NF) of mm-wave mixers. To demonstrate the effectiveness of the proposed$g_{m}$-boosting technique, a high-gain and low-noise mixer for 5G FR2 frequency band applications is developed in a 40 nm CMOS process. Transformer-based pole splitting and derivative superposition are employed to enhance the mixer bandwidth and improve linearity. The mixer achieves a peak CG of 20.9 dB, a 30% fractional bandwidth ($f_{BW}$), a minimum NF of 7.7 dB, and an input referred 1-dB compression point of −14 dBm, leading to an excellent figure of merit (FOM) of 11.05. The mixer consumes 15.6 mW of power and occupies a die area of 0.31 mm2.
Sijie Fu, Boxiao Liu, Chunqi Shi, Leilei Huang, Jinghong Chen, Runxi Zhang
ISCAS5
2023 A 3.84 GHz 32 fs RMS Jitter Over-Sampling PLL with High-Gain Cross-Switching Phase Detector
abstract
A 32 fs RMS jitter oversampling phase-locked loop (OSPLL) exploiting a high-gain cross-switching phase detector (CSPD) is proposed. The over-sampling PLL increases sam-pling frequency by 4x, reducing the in-band phase noise and overcoming the loop bandwidth limitation due to the reference frequency. Leveraging the increased loop bandwidth, the noise contribution of the voltage-controlled oscillator (VCO) is sig-nificantly suppressed. The high-gain CSPD adopts a common-mode sampling technique with time interleaving switches to ensure that the reference clock is sampled only at the maximum slew rate. The CSPD with a higher gain facilitates reducing the noise contribution from the phase detector (PD) and the transconductance cell. Additionally, an RC poly-phase filter (PPF) is employed to generate quadrature clocks, avoiding the deterioration of the PLL's low offset frequency phase noise. The PLL is implemented in a 40-nm CMOS process. Simulation results show that the PLL achieves a 32 fs RMS jitter integrated from 10 kHz to 100 MHz and a power consumption of 6.5 mW, resulting in an$FoM_{jitter}$of -261 dB. At 3.84 GHz frequency, the in-band phase noise is -136.8 dBc/Hz at 100 kHz offset.
Xuhong Lil, Jianghu Hong, Chunqi Shi, Leilei Huang, Boxiao Liu, Hao Deng 0003, Jinghong Chen, Runxi Zhang
ISCAS4
2023 A 88%-Peak-Efficiency 10-mV-Voltage-Ripple Dual-Mode Switched-Capacitor DC-DC Converter for Ultra-Low-Power Battery Management
abstract
This paper proposes a high-efficiency low-ripple dual-mode switched-capacitor (SC) DC-DC converter for low-power IoT and wearable device applications. A hybrid self-biased current scheme (HSBC) is developed to achieve low output voltage ripple and fast response. Two supply voltage domains of HSBC and clock drive controller circuit are introduced to reduce the power loss of the control circuit. An on-chip ultra-low-power bias circuit is also designed to minimize the power loss of the bias generator. The proposed DC-DC converter is implemented in a 40 nm CMOS process. Post-layout simulation results show that the converter realizes 1.6-1.8 V to 0.4 V conversion. The peak efficiency is up to 88% at$5\ \mu\mathrm{A}$, and the voltage ripple is less than 10 mV over a load range of 10 nA-$10\ \mu\mathrm{A}$. The response time of the converter is less than$15\ \mu\mathrm{s}$.
Xiaoyuan Wu, Leilei Huang, Boxiao Liu, Chunqi Shi, Jinghong Chen, Runxi Zhang
ISCAS4
2022 A Fast CABAC Hardware Design for Accelerating the Rate Estimation in HEVC
abstract
The latest High Efficiency Video Coding standard achieves twice the coding efficiency of the H264 standard through a complex rate-distortion optimization (RDO). The coded bit-streams are produced with context adaptive binary arithmetic coding (CABAC). CABAC itself is a very time-consuming process that includes binarization, context modeling, interval subdivision, renormalization, outstanding bit handling, and context updating. The aim of this research is to speed up the CABAC process through several simplifications. First, we approximate three parts of the CABAC, i.e., interval subdivision, renormalization, and outstanding bit handling, with a piecewise-linear function that is very friendly to hardware implementation. In order to achieve better hardware parallelism, we also improve the coding process at the sub-block level. The context of syntax elements in a sub-block is redistributed to skip the complex calculation of context indexing. We perform context updating at the granularity of sub-blocks so that the data dependency of the context updating is removed completely, and the original serial encoding process is changed to a parallel encoding process. At the same time, we make another simplification for the context modeling ofcu_skip_flag. Based on these simplifications, we build a parallel hardware architecture for the rate estimation of the RDO process. This architecture completes the bit estimation of a$32\times 32$coding tree unit (CTU) in 220.8 nano-seconds, whereas the Bjøntegaard Delta rate increases by only 2.225%. We believe that the proposed architecture can meet the requirements of 8K@120 fps ultra-high-definition videos. This is the first study to simplify the hardware design of rate estimation by changing the context allocation and updating rules.
Yujie Cai, Yibo Fan, Leilei Huang, Xiaoyang Zeng, Haibing Yin, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 A Micro-Code-Based Hardware Architecture of Integer Motion Estimation for HEVC
abstract
The advent of Ultra High Definition (UHD) and Super Hi-Vision (SHV) videos has motivated the development of advanced video coding standard in these years. Integer motion estimation (IME) is the most computationally expensive process of High Efficiency Video Coding (HEVC), which is one of the most widespread video coding standards. Many previous works related to IME put emphasis on enlarging search range and improving computation complexity. However, the fact that the IME algorithm should be adaptive to different scenarios was neglected. In view of this, a configurable IME engine and its micro-code-based hardware design are proposed in this paper. Three different search schemes based on our IME engine are evaluated using the HEVC test model (HM) 16.9, achieving an average BD-rate increase of 0.55/-0.07/-0.14%. The hardware design is implemented with 225. 7K gates at 500MHz using TSMC 65nm CMOS standard-cell libraries.
Chenhao Gu, Leilei Huang, Xiaoyang Zeng, Yibo Fan
VLSI-SoC2
2018 A Compact and Configurable Long Short-Term Memory Neural Network Hardware Architecture
abstract
Neural network has been one of the most useful techniques in the area of image analysis and speech recognition in recent years. Long Short-Term Memory (LSTM), a popular type of recurrent neural networks (RNNs), has widely been implemented on CPUs and GPUs. However, software implementation cannot offer large parallelism for the complicated computation of LSTM, and most of the LSTM hardware implementations proposed yet are intensive and non-configurable. In order to accelerate the computation and reduce the resources consumption, in this work, we present a compact and configurable LSTM neural network hardware architecture. To meet the requirements of different networks, we set a wide array of hardware parameters that can be configured to balance area, power and performance. And we adopt the second-order polynomial to approximate the activation functions in LSTM, which balances the computational accuracy and resource utilization, Implemented on XCZU6EG FPGA running at 238 MHz, our work has a performance of 7.64 GOP/s. Compared to the implementation on Intel Xeon E5-2620 CPU at 2.10 GHz, our parallel hardware architecture achieves 90× speedup for a small network and 25 x speed-up for a large one. The total consumption of resources is 77% less than the state-of-the-art works', which implies the compactness of our work.
Leilei Huang, Minjiang Li, Xiaoyang Zeng, Yibo Fan
ICIP2
2018 A Hardware-Oriented IME Algorithm for HEVC and Its Hardware Implementation
abstract
High Efficiency Video Coding (HEVC), the latest video coding standard, aims to provide coding performance that is much superior to that of its predecessor, H.264, especially for high definition video. To fulfill this goal, the inter-prediction unit (PU) partitions of HEVC are more complex, and the search range of motion estimation (ME) is much larger. As a result, ME becomes a bottleneck in the design of the HEVC inter predictor. In response to this challenge, we developed a hardware-oriented integer ME algorithm and the related hardware implementation. Our proposed algorithm led to a decrease in terms of the Bjontegaard Delta rate when compared with the HEVC test model 15.0. The corresponding hardware solution benefitted from 2-D data reuse supported by horizontal and vertical reference SRAMs, on-chip memory reduction supported by 4 × 4 block compression, and a low-power sum of absolute difference (SAD) tree supported by PU-level chip selection. When adopting a 32 × 32 SAD tree, the minimum and maximum required working frequency for 4K × 2K at 30 frames/s videos was [375, 500] MHz. These results demonstrated that our proposed solution offered desirable improvement in both coding speed and coding performance.
Yibo Fan, Leilei Huang, Bei Hao, Xiaoyang Zeng
IEEE Trans. Circuits Syst. Video Technol.2
2016 Quarter LCU based integer motion estimation algorithm for HEVC
abstract
In this paper a hardware oriented integer motion estimation (IME) algorithm is proposed. The algorithm put forward to divide the largest coding unit (LCU) into four motion vector (MV) prediction cluster. Each cluster has a separate MV as the start for search window center. Then the best matched MV can be searched in a small size search window. All PUs in one quarter LCU sharing the same reference block save the bandwidth cost of loading reference block to the chip largely. In addition, the sum of absolute difference (SAD) value calculation for all the PUs in a quarter LCU are also shared, actually, only once calculation process is needed for all the PUs' distortion in it. Then, the algorithm is evaluated in the HEVC test model (HM-11.0), the results show that it only introduced 0.85% BDBR loss.
Qinwei Jiang, Leilei Huang, Yibo Fan, Xiaoyang Zeng
ICIP2
2016 A Combined Deblocking Filter and SAO Hardware Architecture for HEVC
abstract
The latest video coding standard high-efficiency video coding (HEVC) provides 50% improvement in coding efficiency compared to H.264/AVC to meet the rising demands for video streaming, better video quality, and higher resolution. The deblocking filter (DF) and sample adaptive offset (SAO) play an important role in the HEVC encoder, and the SAO is newly adopted in HEVC. Due to the high throughput requirement in the video encoder, design challenges such as data dependence, external memory traffic, and on-chip memory area become even more critical. To solve these problems, we first propose an interlacing memory organization on the basis of quarter-LCU to resolve the data dependence between vertical and horizontal filtering of DF. The on-chip SRAM area is also reduced to about 25% on the basis of quarter-LCU scheme without throughput loss. We also propose a simplified bitrate estimation method of rate-distortion cost calculation to reduce the computational complexity in the mode decision of SAO. Our proposed hardware architecture of combined DF and SAO is designed for the HEVC intraencoder, and the proposed simplified bitrate estimation method of SAO can be applied to both intra- and intercoding. As a result, our design can support ultrahigh definition 7680 × 4320 at 40 f/s applications at merely 182 MHz working frequency. Total logic gate count is 103.3 K in 65 nm CMOS process.
Weiwei Shen, Yibo Fan, Yufeng Bai, Leilei Huang, Qing Shang, Cong Liu 0014, Xiaoyang Zeng
IEEE Trans. Multim.4
2006 Quantum key distribution based on a Sagnac loop interferometer and polarization-insensitive phase modulators
abstract
We present a design for a quantum key distribution (QKD) system in a Sagnac loop configuration, employing a novel phase modulation scheme based on frequency shift, and demonstrate stable BB84 QKD operation with high interference visibility and low quantum bit error rate (QBER). The phase modulation is achieved by sending two light pulses with a fixed time delay (or a fixed optical path delay) through a frequency shift element and by modulating the amount of frequency shift. The relative phase between two light pulses upon leaving the frequency-shift element is determined by both the time delay (or the optical path delay) and the frequency shift, and can therefore be controlled by varying the amount of frequency shift. To demonstrate its operation, we used an acousto-optic modulator (AOM) as the frequency-shift element, and vary the driving frequency of the AOM to encode phase information. The interference visibility for a 40 km and a 10 km fiber loop is 96% and 99%, respectively, at single photon level. We ran BB84 protocol in a 40-km Sagnac loop setup continuously for one hour and the measured QBER remained within the 2%-5% range. A further advantage of our scheme is that both phase and amplitude modulation can be achieved simultaneously by frequency and amplitude modulation of the AOM's driving signal, allowing our QKD system the capability of implementing other protocols, such as the decoy-state QKD and the continuous-variable QKD. We also briefly discuss a new type of Eavesdropping strategy ("phase-remapping" attack) in bidirectional QKD system
Leilei Huang, Hoi-Kwong Lo
ISIT2