Zhidao Zhou

dblp:357/7965 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A BEOL Ferroelectric FET-based Computing Unit for Digital Computing-in-Memory
Junyu Zhu, Zexue Bian, Weizeng Li, Junzhe Shen, Hanghang Gao, Zhidao Zhou, Zhongze Han, Zhi Li 0062, Hongyang Hu, Chunmeng Dou
ISCAS6
2025 An energy-efficient FeFET-based computing-in-memory macro using BEOL-integrated HZO ferroelectric capacitors
Weizeng Li, Zhidao Zhou, Linfang Wang, Junyu Zhu, Junzhe Shen, Hongyang Hu, Baihan Wang, Zhi Li 0062, Wang Ye, Zhongze Han, Hanghang Gao, Chunmeng Dou
Sci. China Inf. Sci.2
2025 An RRAM Digital Computing-in-Memory Macro With Dual-Mode Multiplication and Maximum Value Rounding Adder Tree
abstract
Implementing digital computing-in-memory (DCIM) based on resistive memory (RRAM) faces several critical challenges due to the small signal margin, large device variations, and large energy- and area-overhead induced by the digital adder tree (AT). To address these issues, we propose an RRAM DCIM macro based on the standard foundry one-transistor-one-resistor (1T1R) cell array featuring: 1) dual-mode MAC operation for efficiency- or accuracy-oriented optimization; 2) margin-enhanced digitized unit (MEDU) to amplify the signal ratio; and 3) maximum value rounding AT (MVR-AT) to reduce its power- and area-overhead. A test chip is demonstrated using a 180 nm CMOS process to verify the concept. It achieves a peak energy efficiency (EF) of 63.08 TOPS/W in the efficiency-oriented mode and a minimum error rate of 1.58% in the accuracy-oriented mode. Their combination can meet the requirements of different workloads in AI computing tasks to optimize the overall power consumption with negligible accuracy loss.
Wang Ye, Hanghang Gao, Zhidao Zhou, Linfang Wang, Weizeng Li, Zhi Li 0062, Jinshan Yue, Xiaoxin Xu, Hongyang Hu, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.3
2024 Cross-Frame Integrated Prediction for Feature-Space Video Compression
abstract
Learned video compression in the feature domain employs implicit motion compensation to acquire predicted features or compute feature residuals, effectively minimizing spatiotemporal redundancy in the reconstructed frames. In this paper, we propose a cross-frame integrated prediction (CIP) network for feature-space video compression. Leveraging multiple features in motion estimation and compensation, our approach enables more context-aware prediction. Specifically, on the one hand, we introduce a global feature extraction (GFE) module in motion estimation to extract the information of the current feature and multiple reference features, providing a high-quality offset map for deformable motion compensation. On the other hand, we utilize an attentional feature fusion (AFF) module for multiple predicted features in motion compensation, which is beneficial for preserving crucial details and adapting to diverse scenes and content. By passing the final predicted feature to the residual compression and frame reconstruction, we achieve a single end-to-end video compression framework avoiding laborious multi-stage training. Comprehensive experimental results show that the proposed method not only maintains a low number of model parameters but also achieves significant performance improvement in video compression tasks, especially in the case of high resolution and high bitrate.
Hongxin Qiu, Zhidao Zhou, Fan Liang 0001
IJCNN2
2024 PFR-VC: Learning-Based Video Compression Framework with Predicted Frame Refinement
abstract
Learning-based video compression has attracted more and more attention in recent years. Traditional video coding relies on block-based motion estimation and spatial frequency transformation. While these techniques can effectively compress videos, further enhancing the compression ratio becomes challenging. Introducing deep learning methods can overcome the limitations of manually designed algorithms. In this paper, we propose a learning-based video compression framework with Predicted Frame Refinement (PFR) to improve the compression efficiency. Firstly, a simple autoencoder is introduced to encode the motion information, eliminating the need for a complex optical-flow network. Then, we design a predicted frame refinement network with an attention feature fusion mechanism to generate predicted frames more suitable for extracting context. Finally, we introduce a context coding scheme to improve the compression ratio by jointly utilizing temporal prior and hyper prior. The entire network can be globally optimized and trained from scratch. The experimental result shows that the proposed compression framework outperforms previous methods. Our approach brings 31.2% more saved bit rate than x265 with veryslow preset. Our model also achieves a 7.1% gain in Multi-Scale Structural Similarity Index Measure (MS-SSIM) compared with the recent method proposed by Guo et al.(2023).
Zhidao Zhou, Hongxin Qiu, Zhikai Liu, Wei Sun 0007, Fan Liang 0001
IJCNN1
2024 A 2T P-Channel Logic Flash Cell for Reconfigurable Interconnection in Chiplet-Based Computing-In-Memory Accelerators
abstract
In this work, we propose a two-transistor (2T) p-type channel (p-channel) logic-compatible flash cell. Compared to the previous designs, the proposed structure features reduced area-cost and enhanced ability to pass through the logic ‘1’. Due to these advantages, we explore its application as the reconfigurable interconnections in the chiplet-based system. By integrating them into the silicon interposer, the 2T p-channel flash cells can potentially lead to the dense and flexible interconnection between multiple computing-in-memory (CIM) chiplets, resulting in highly reconfigurable and scalable chiplet-based CIM accelerators. A 180nm 1Kb 2T p-channel flash cell array is fabricated and characterized. The characterization results show the 2T p-channel flash cells exhibit a signal ratio >103over 1000 program/erase (P/E) cycles and the device-to-device variations are less than 21.07%. Their typical behaviors as routers are also confirmed by circuit simulations.
Weizeng Li, Linfang Wang, Zhi Li 0062, Wang Ye, Zhidao Zhou, Haiyang Zhou, Hanghang Gao, Jinshan Yue, Hongyang Hu, Fengman Liu, Chunmeng Dou
ISCAS5
2024 PFT-ILF: In-loop Filter with Partition Feature Transform for Versatile Video Coding
abstract
The new generation of video coding standards, Versatile Video Coding (VVC), integrates a range of loop filter mechanisms, notably the De-Blocking Filter (DBF), Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF). However, these traditional tools are handcrafted empirically and have limitations. Thus many CNN-based loop filters have been proposed to achieve better image quality. In this paper, we propose a novel network based on Coding Unit (CU) partition feature transformation, called PFT-ILF. Considering that the CU partition map contains image distortion information, our approach innovatively utilizes the CU partition map as prior information to guide filtering. And we design a Partition Feature Transform (PFT) layer, which uses partition features to generate modulation parameters pair for adjusting the features of several intermediate layers in the network. We integrate our proposed filter into the NNVC standard software VTM11.0_NNVC-4.0 and conduct ablation experiments. Under the all intra configuration, our method achieves Bjøntegaard-Delta Bit-Rate (BD-BR) reductions of 7.52%, 17.36%, and 18.65% for Y, U, and V components, respectively.
Xin-Yi Cui, Zhikai Liu, Zhidao Zhou, Fan Liang 0001
VCIP3
2024 Multi-stage Attention Network with Auxiliary Information Refinement for VVC In-loop Filtering
abstract
Recently, learning-based video compression techniques have brought significant performance improvements. However, most existing methods have not fully exploit the auxiliary information from the encoding process. To achieve better performance, we propose a multi-stage attention network with auxiliary information refinement for Versatile Video Coding (VVC) in-loop filtering. The proposed network consists of two branches: the main filter branch extracts reconstruction features, while the auxiliary information refinement branch processes prediction and partition. Specifically, the auxiliary information refinement aims to extract features of auxiliary information better to assist in removing compression artifacts. Lastly, we introduce an Auxiliary Information Attention Module (AAM) to fuse the information flow between the two branches. The proposed model is integrated into the NNVC standard software VTM11.0_NNVC-4.0 and tested under all intra configuration. Experimental results show that our method achieves -7.60%, -19.49%, and -20.50% Bjøntegaard-Delta Bit-Rate (BD-BR) improvements across the Y, U, and V components.
Xin-Yi Cui, Zhidao Zhou, Zhikai Liu, Fan Liang 0001
VCIP2
2024 Write-Verify-Free MLC RRAM Using Nonbinary Encoding for AI Weight Storage at the Edge
abstract
High-density and reliable multilevel-cell (MLC) resistive random access memory (RRAM) is expected to meet the ever-increasing demand for on-chip weight storages in the intelligent edge devices. However, due to the device variations, many write-and-verify (WAV) iterations are usually required to program the RRAM cell, which causes high power consumption, long latency, and degradation on the memory lifetime. To address this issue, we propose a write–verify-free MLC RRAM macro for weight storage with 1) a cascode-current-mirror multibit write (CCM-MW) driver and 2) a nonbinary programming scheme (NB-PS) with a radix not greater than 2. A 180-nm 400-Kb RRAM test chip is demonstrated in silicon. For 2-bit-per-cell MLC storage, the value error rates can be reduced by 24.13% after introducing two redundant bits (RBDs). In addition, compared to the single-level cell (SLC) storage scheme, a 37.50% reduction in the number of cells can be achieved to store the ResNet-8 model with a 0.79% loss in inference accuracy without the need for WAV iterations.
Junjie An, Zhidao Zhou, Linfang Wang, Wang Ye, Weizeng Li, Hanghang Gao, Zhi Li 0062, Jinghui Tian, Hongyang Hu, Jinshan Yue, Lingyan Fan, Shibing Long, Qi Liu 0010, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.2