EDBT 2026 Demo / reviewers in the wild / expert
Shushi Chen
dblp:344/7986 · also ShuShi Chen
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
0009-0000-2992-1629ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A High-Precision and Low-Cost Approximate Transform Accelerator for Video CodingabstractThe introduction of multiple transform types in the Versatile Video Coding (VVC) standard has yielded notable encoding gains but also imposed considerable computational burdens. Existing transform circuits of different types are typically implemented separately due to their independence, leading to substantial hardware overhead. To address this, we explore the relationship between Discrete Cosine Transform Type-2 (DCT2) and Discrete Sine Transform Type-7 (DST7) matrices and reveal a prominent diagonal aggregation phenomenon in their transfer matrix. Based on this insight, the least-squares method is applied to optimize the transfer matrix sparsity, achieving a high-precision, low-cost approximate conversion from DCT2 to DST7. Furthermore, we optimize DCT2 computation by proposing an elaborate matrix decomposition approach that allows a lightweight shift-adder unit to efficiently generate all required product terms across varying sizes. Leveraging these algorithmic optimizations, we implement a highly reusable and area-efficient approximate transform accelerator that supports sizes from 4 to 32 points and accommodates three types in VVC. Experimental results demonstrate that the proposed accelerator achieves over 44% reduction in circuit resource consumption with negligible BD-BR performance loss of just $\mathbf{0. 5 3 \%}$, maintaining processing capabilities up to $8 K \text{@} 57 \mathrm{fps}$. Zhijian Hao, Chenlong He, Qi Zheng 0004, Shushi Chen, Jinchang Xu, Yue Hao 0001, Xiaohua Ma 0001 |
DAC | 5 |
| 2025 | MAS-ISP: A Proxy-Free Online Hyperparameter Optimization Framework for ISP Hardware SystemabstractThe rapid advancement of visual autonomous systems, especially in autonomous driving, underscores the critical role of Image Signal Processors (ISPs) as they convert RAW sensor data into RGB images suited for visual interpretation. Traditional ISPs rely on tuning hyperparameters to adapt to varying imaging conditions; however, the vast parameter space and intricate tuning process pose significant challenges for realtime autonomous applications. Existing autonomous ISP hyperparameter optimization methods rely largely on offline or proxybased online tuning, limiting their accuracy and responsiveness to real-time environmental changes. In response, we propose an online ISP hyperparameter optimization framework based on Deep Reinforcement Learning (DRL), marking the first proxyfree, real-time optimization approach. Our design exhibits a master-slave Multi-Agent System (MAS), enabling rapid and cooperative parameter optimization with improved inter-frame consistency. Furthermore, we design the MAS-ISP automated visual system, incorporating innovative hardware designs such as Strip Convolution Kernel and Stride-Aware Dual-Buffer Memory, which drastically reduce resource consumption in CNN hardware. MAS-ISP achieves 1080P@75FPS/240FPS on FPGA/ASIC platforms, supporting real-time and reliable visual systems. Zhijian Hao, Ruoxi Zhu, Qi Zheng 0004, Shuocheng Wang, Shushi Chen, Leilei Huang, Jun Tao 0001, Yibo Fan |
DAC | 7 |
| 2025 | A Hardware-Friendly Lightweight Partition Decision Algorithm for VVC Intra and Inter CodingabstractThe Versatile Video Coding (VVC) standard notably enhances encoding efficiency with the Quad-Tree plus Multi-Type Tree (QTMTT) partition structure. However, the complex QTMTT tool presents substantial challenges in both software and hardware implementation. To overcome those challenges, this paper introduces a hardware-friendly partition decision algorithm for VVC intra and inter coding. Firstly, we propose a lightweight backbone network to extract partition-aware features. Secondly, we employ a Quantisation Parameter (QP) fusion network to regulate the impact of QPs on the partition structure. Additionally, we apply a top-down threshold-driven post-processing algorithm, in which improbable partition types are removed to directly derive the unique partition structure. Experiments show that our method not only exceeds the previous state-of-the-art work in BD-BR performance, but also shows sufficient hardware-friendly characteristics. To the best of our knowledge, this work is among the earliest to comprehensively discuss and implement a hardware-friendly partition decision algorithm. Zhao Zan, Leilei Huang, Shushi Chen, Xiaoyang Zeng, Yibo Fan |
IEEE Signal Process. Lett. | 3 |
| 2025 | Affine Motion Estimation Hardware Implementation With 51.7%/67.5% Internal Bandwidth Reduction for Versatile Video CodingabstractVersatile Video Coding (VVC) employs Affine Motion Compensation (AMC) to process scenes with high-order motion. To improve AMC efficiency, the Affine Motion Estimation (AME) process based on the gradient-based iterative algorithm (GIA) and block match algorithm (BMA) is introduced to the VVC Test Model (VTM). However, the AME process is highly complex and difficult for hardware implementation in real-time applications. In this context, this paper proposes a hardware-friendly AME algorithm and implements the corresponding accelerator. Firstly, the weighted least squares regression is used to reduce the iteration of GIA. Then an iteration-free search scheme is proposed to remove the search dependence during the GIA and BMA process. In addition, a motion vector clamping mechanism and four-level memory organization are proposed to solve the problem of reference pixel reading conflict, which reduces 51.7% and 67.5% internal bandwidth of the AME accelerator. Compared with the default AME process of VTM 16.0, experimental results show that the proposed algorithm reduces AME run time by 81.63% while the corresponding Bjontegaard Delta Bit Rate (BDBR) loss is only 0.492%. The proposed AME accelerator can flexibly support AME search tasks in various configurations. Synthesized with the TSMC 28nm process, the proposed architecture has a gate count of 1313K and a power consumption of 156.83 mW. It can achieve$7680\times [email protected]~30fps and the corresponding BDBR loss is 0.492%~1.835%. Shushi Chen, Leilei Huang, Zhao Zan, Zhijian Hao, Hao Zhang 0126, Xiaoxiang Chen, Minge Jing, Xiaoyang Zeng, Yibo Fan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | An Interpolation-Free Fractional Motion Estimation Algorithm and Hardware Implementation for VVCabstractVersatile video coding (VVC) introduces multi-type tree (MTT) and larger coding tree unit (CTU) to improve compression efficiency compared to its predecessor High Efficiency Video Coding (HEVC). This leads to higher throughput for fractional motion estimation (FME) to meet the needs of real-time processing. In this context, this article proposes an interpolation-free algorithm based on an error surface to improve the throughput of FME hardware. The error surface is constructed by the rate-distortion costs (RDCs) of the integer motion vector (IMV) and its neighbors. To improve the prediction accuracy, a hardware-friendly RDC estimation strategy is proposed to construct the error surface. The experimental results show that the corresponding Bjontegaard Delta Bit Rate (BDBR) in Random Access (RA), Low Delay P (LDP) and Low Delay B (LDB) configuration increases by only 0.358%, 0.479%, and 0.511% compared with the VVC test model (VTM) 16.0. Compared with the default FME algorithms of VVC, the time cost of FME is reduced by 53.47%, 56.28%, and 54.23%, respectively, in RA, LDP, and LDB configurations. The algorithm is free of iteration and interpolation, which can contribute to low-cost and high-throughput hardware. The proposed architecture can support FME of all coding units (CUs) in a CTU with one layer of MTT under the quaternary tree (QT), and the CU size can vary from$8\times 8$to$128\times 128$. Synthesized using GF 28-nm process, the architecture can achieve$7680\times 4320$@60 fps throughput at 800 MHz, with a gate count of 244 K and power consumption of 76.5 mW. This proposed architecture can meet the real-time coding requirements of VVC. Shushi Chen, Leilei Huang, Zhao Zan, Xiaoyang Zeng, Yibo Fan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | An 8K@120fps Hardware Implementation for Decoder-Side Motion Vector Refinement in VVCabstractVersatile video coding (VVC) introduces many coding tools to improve compression efficiency not only on the encoder side but also on the decoder side. For inter-picture coding, the introduction of decoder-side motion vector refinement (DMVR) puts forward a demand for a motion search engine in the hardware decoder. However, because of the specific mirroring search property in DMVR, the existing hardware design of motion estimation cannot be directly adopted for DMVR. This paper proposes a specific pixel-level sum of absolute differences (SAD) calculation datapath for the 25 mirrored symmetrical search points and also presents a cost-effective design of DMVR search engine under the condition of high-throughput. Finally, the proposed hardware implementation is synthesized with TSMC 22nm process. The measured throughput can reach 8K@120fps at 500MHz, with a total cell area of 52479 μm2and power consumption of 10.82 mW. To the best of our knowledge, this work is the first hardware architecture for the DMVR search process in VVC. Leilei Huang, Shushi Chen, Yibo Fan |
ISCAS | 3 |
| 2023 | Fast QTMT Partition for VVC Intra Coding Using U-Net FrameworkabstractVersatile Video Coding (VVC) has significantly increased encoding efficiency at the expense of numerous complex coding tools, particularly the flexible Quad-Tree plus Multi-type Tree (QTMT) block partition. This paper proposes a deep learning-based algorithm applied in fast QTMT partition for VVC intra coding. Our solution greatly reduces encoding time by early termination of less-likely intra prediction and partitions with negligible BD-BR increase. Firstly, a redesigned U-Net is recommended as the network’s fundamental framework. Next, we design a Quality Parameter (QP) fusion network to regulate the effect of QPs on the partition results. Finally, we adopt a refined post-processing strategy to better balance encoding performance and complexity. Experimental results demonstrate that our solution outperforms the state-of-the-art works with a complexity reduction of 44.74% to 68.76% and a BD-BR increase of 0.60% to 2.33%. Zhao Zan, Leilei Huang, Shushi Chen, Zhenghui Zhao, Haibing Yin, Yibo Fan |
ICIP | 3 |
| 2023 | An Error-Surface-Based Fractional Motion Estimation Algorithm and Hardware Implementation for VVCabstractVersatile Video Coding (VVC) introduces more coding tools to improve compression efficiency compared to its predecessor High Efficiency Video Coding (HEVC). For inter-frame coding, Fractional Motion Estimation (FME) still has a high computational effort, which limits the real-time processing capability of the video encoder. In this context, this paper proposes an error-surface-based FME algorithm and the corresponding hardware implementation. The algorithm creates an error surface constructed by the Rate-Distortion (R-D) cost of the integer motion vector (IMV) and its neighbors. This method requires no iteration and interpolation, thus reducing the area and power consumption and increasing the throughput of the hardware. The experimental results show that the corresponding BDBR loss is only 0.47% compared to VTM 16.0 in LD-P configuration. The hardware implementation was synthesized using GF 28nm process. It can support 13 different sizes of CU varying from$128\times 128$to$8\times 8$. The measured throughput can reach 4K@30fps at$400\mathbf{MHz}$, with a gate count of 192k and power consumption of 12.64 mW. And the throughput can reach 8K@30fps at 631MHz when only quadtree is searched. To the best of our knowledge, this work is the first hardware architecture for VVC FME with an interpolation-free strategy. Shushi Chen, Leilei Huang, Chao Liu 0027, Yibo Fan |
ISCAS | 1 |