EDBT 2026 Demo / reviewers in the wild / expert
Zhicheng Hu
dblp:88/8871
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ERSRP: A 55nm 46.28 MPixels/(s·mm2) 104 FPS Efficient Real-Time Super-Resolution Processor with Layer-Fused Lightweight EngineabstractThis paper presents ERSRP, a 55nm Edge Real-Time Super-Resolution Processor that achieves 104 FPS at FHD resolution with a peak area efficiency of 46.28 MPixels/(s·mm2), 8.86× higher than the state-of-the-art. The processor is designed through a software-hardware co-optimization approach, addressing the low utilization, high memory demand, and workload imbalance challenges inherent in lightweight SR networks. At the algorithmic level, an Ultra-Lightweight Super-Resolution (ULSR) model is proposed that integrates depth-wise and point-wise separable blocks with a pixel-shuffle mechanism to achieve high-quality reconstruction (37.18 dB PSNR and 0.9581 SSIM on Set5) with only 5.62K parameters. At the hardware level, the ERSRP introduces a Lightweight Accelerated Engine (LAE) sup-porting a Layer Parallel Computing Scheme (LPCS) to improve lightweight operator throughput by 48.9%. A Point-wise Layer Fused Scheme (PLFS) further enhances utilization by 4.95× through inter-core and intra-core fusion without intermediate memory. Fabricated in a 55nm UMC CMOS process, the ERSRP achieves a throughput of 215.7 MPixels/s and supports 104 FPS real-time SR at FHD. Gaoxiang Wu, Liang Chang 0002, Jingke Wang, Zhicheng Hu, Xin Zhao 0044, Fengbin Tu, Jun Zhou 0017 |
ISCAS | 5 |
| 2025 | Trident: The Acceleration Architecture for High-Performance Private Set IntersectionabstractPrivate Set Intersection (PSI) is imperative in discovering the properties of the same data owned by two competitive parties, without revealing anything else of their respective data asset. Existing PSI solutions such as APSI and ORI-PSI suffer from severe communication and computation overhead due to inefficient communication and FHE polynomial evaluation, which hinders their deployment in practice. This issue is evident in both the upper-level protocol and the lower-level hardware platform. In this paper, we propose a novel software/hardware co-design acceleration architecture for PSI, termed as “Trident”, which includes two tightly coupled segments: from the protocol perspective, we investigate existing bottlenecks and propose a new PSI protocol with significantly less communication and computation under the security guarantee; besides, we re-architect the hardware platform by designing a PSI-specific accelerator, implemented with both FPGA and ASIC, targeting the key operations in the proposed protocol. We build a real-world experimental environment with two instantiated parties to verify the acceleration architecture, and highlight the following results: (1) up to 130$\boldsymbol{\times}$/145$\boldsymbol{\times}$speedup for the computation ofreceiverandsenderparties; (2) up to 37$\boldsymbol{\times}$reduction of communication overhead. (3) up to 93,651$\boldsymbol{\times}$and 74,326$\boldsymbol{\times}$higher energy efficiency over the CPU-based ORI-PSI and APSI, respectively. Jinkai Zhang, Yinghao Yang 0001, Zhe Zhou 0003, Zhicheng Hu, Xin Zhao 0044, Liang Chang 0002, Xiaowei Li 0001 |
IEEE Trans. Computers | 4 |
| 2024 | A RRAM-based High Energy-efficient Accelerator Supporting Multimodal Tasks for Virtual Reality Wearable DevicesabstractVirtual reality (VR) wearable devices can achieve immersive entertainment by fusing multi-modal tasks from various senses. However, constrained by the short battery life and limited hardware resources of the VR devices, running multiple tasks simultaneously with different modals is difficult. In this paper, we propose an energy-efficient accelerator that supports Multi-modal Tasks for VR devices, namely MTVR. We present a multi-task computing solution based on the flexible multi-task computing core design and efficient computing unit allocation strategy, which simultaneously achieves efficient work of multi-modal tasks. We design an early exit detector to skip invalid calculations, greatly saving energy. In addition, a fine-grained tiny value skip method at multiplier and adder levels is proposed to save energy further. We provide a hybrid RRAM and SRAM memory access scheme, reducing the external memory access (EMA). Through experimental evaluation, the multitask computing core achieves an average computational utilization of 95%. When the invalid input ratio is 90%, energy saving brought by the early exit detector can reach 88%. The tiny value skip method further achieved 13% energy saving. Hybrid memory access scheme obtains 98.9% EMA reduction. We deployed the MTVR accelerator in FPGA and self-designed RRAM, achieving energy efficiency of 3.6 TOPS/W, higher than other single-task accelerators. Xin Zhao 0044, Zhicheng Hu, Zilong Guo, Haodong Fan, Liang Chang 0002 |
DAC | 2 |
| 2024 | Urban Lighting Infrastructure Analysis Using Topology Density MapsabstractDigital twins that simulate urban environments have emerged as promising tools to support informed decision-making. In the context of lighting infrastructure analysis using such digital twins, the use of suitable information visualization methods for supporting the identification of the pros and cons of different alternatives is of paramount importance. This paper introduces the use of Topology Density Maps in the assessment of simulated lighting infrastructure designs. Our analysis also considers the evaluation of temporal changes of density maps, using the recently proposed Temporal Topology Density Map algorithm. We demonstrate the potential of exploring such visualization methods in the context of walkability assessment using real and simulated data associated with an urban green area in {\AA}lesund, Norway. Zhicheng Hu, Claudia Viviana Lopez-Alfaro, Ricardo da Silva Torres |
ECMS | 1 |
| 2024 | SuperHCA: A Super-Resolution Accelerator with Sparsity-Aware Heterogeneous Core ArchitectureabstractDeep learning-based super-resolution (SR) models have emerged as a potential approach to achieving high-quality images. The large SR networks can achieve a high peak signal-noise ratio (PSNR), a metric to evaluate the quality of the image. However, the SR networks typically contain large amounts of parameters, inducing high computation capacity and memory bandwidth requirements, which are difficult to deploy on embedded hardware. In this work, we develop the Anchor-Based Shuffle Net (ABSN) oriented to develop a hardware accelerator with a dynamic-scale fixed-point (DSFP) quantization method. In addition, we implement the dynamic quantization adaption in hardware. We design a Super-resolution Heterogeneous Accelerator, namely SuperHCA, employing a sparsity-aware heterogeneous architecture to distinguish between dense and sparse workloads to improve inference efficiency. Furthermore, we provide Slice Layer Fusion (SLF) computation in the heterogeneous cores to reduce external memory access and on-chip buffer sizes. The SuperHCA achieves 91 FPS with the lowest area overhead compared to the state-of-the-art works. Zhicheng Hu, Xin Zhao 0044, Liang Chang 0002 |
ISCAS | 1 |
| 2024 | USR-LUT: A High-Efficient Universal Super Resolution Accelerator with Lookup TableabstractSuper-resolution (SR) can promote medical diagnosis efficiency by enriching the details of captured images, such as gastroscopy and colonoscopy. However, the wireless capsule detector used for diagnosis is constrained by the camera’s low resolution and limited computing resources, making it difficult to deploy computation- and memory access-intensive SR models. In this paper, we propose an efficient universal SR accelerator based on lookup tables, namely USR-LUT, which can support various SR algorithms. We design a LUT-based computing unit (LCU) with higher efficiency and lower area overhead. By utilizing the sparsity of deconvolution, we propose an efficient data mapping scheme that can flexibly support convolution and deconvolution with different kernel sizes, achieving a 3.24× acceleration for deconvolution. Tile-based computing is adopted to reduce memory resources and external memory access (EMA) overhead. Through experimental evaluation, compared with LUT-based SR algorithms, the USR-LUT achieves the least LUT storage entries of 82k and the smallest LUT resource overhead of 0.078MB, respectively. The USR-LUT achieves the highest area efficiency of 175.7GOPS/mm2and throughput area ratio (TAR) of 39.9fps/mm2under the 8-bit precision compared with the state-of-the-art works. To the best of our knowledge, this is the first work of SR accelerators adopting LUT-based computing, which is suitable for tiny mobile devices. Xin Zhao 0044, Zhicheng Hu, Liang Chang 0002 |
ISCAS | 2 |
| 2024 | Trinity: A General Purpose FHE AcceleratorabstractFully Homomorphic Encryption (FHE) is crucial for privacy-preserving computing, which allows direct computation on encrypted data. While various FHE schemes have been proposed, none of them efficiently support both arithmetic FHE and logic FHE simultaneously. To address this issue, researchers explore the combination of different FHE schemes within a single application and propose algorithms for the conversion between them. Unfortunately, all prior ASIC-based FHE accelerators are designed to support a single FHE scheme, and none of them supports the acceleration for FHE scheme conversion. This necessitates FHE acceleration systems to integrate multiple accelerators for different schemes, leading to increased system complexity and hindering performance enhancement. In this paper, we present the first multi-modal FHE accelerator based on a unified architecture, which efficiently supports CKKS, TFHE, and their conversion scheme within a single accelerator. To achieve this goal, we first analyze the theoretical foundations of the aforementioned schemes and highlight their composition from a finite number of arithmetic kernels. Then, we investigate the challenges for efficiently supporting these kernels within a unified architecture, which include 1) concurrent support for NTT and FFT, 2) maintaining high hardware utilization across various polynomial lengths, and 3) ensuring consistent performance across diverse arithmetic kernels. To tackle these challenges, we propose a novel FHE accelerator named Trinity, which in-corporates algorithm optimizations, hardware component reuse, and dynamic workload scheduling to enhance the acceleration of CKKS, TFHE, and their conversion scheme. By adaptive select the proper allocation of components for NTT and MAC, Trinity maintains high utilization across NTTs with various polynomial lengths and imbalanced arithmetic workloads. The experiment results show that, for the pure CKKS and TFHE workloads, the performance of our Trinity outperforms the state-of-the- art accelerator for CKKS (SHARP) and TFHE (Morphling) by 1.49 x and 4.23 x, respectively. Moreover, Trinity achieves 919.3 x performance improvement for the FHE-conversion scheme over the CPU-based implementation. Notably, despite the performance improvement, the hardware overhead of Trinity is only 85 % of the summed circuit areas of SHARP and Morphling. Xianglong Deng, Shengyu Fan, Zhicheng Hu, Zhuoyu Tian, Jiangrui Yu, Dingyuan Cao 0002, Dan Meng 0002, Rui Hou 0001, Meng Li 0004, Qian Lou, Mingzhe Zhang 0005 |
MICRO | 3 |
| 2024 | A combination prediction model based on Theil coefficient and induced continuous aggregation operator for the prediction of Shanghai composite indexabstractThis paper proposes an interval combination prediction model for Shanghai composite index, utilizing the Theil coefficient and the induced continuous generalized ordered weighted logarithmic harmonic averaging (ICGOWLHA) operator. The effectiveness of the proposed model under specific weight conditions and the existence of its analytical solution are demonstrated. Shanghai composite index's case analysis demonstrates that, in terms of interval root mean squared error (IRMSE), interval mean absolute error (IMAE), interval mean absolute percentage error (IMAPE), and interval mean squared percentage error (IMSPE), the proposed model's predictive performance improvements over the best-performing single prediction model are 29.33%, 25.72%, 26.10%, and 28.86%, respectively. At the same time, the theoretical properties of the model are verified in the results of the case analysis, and the model's convergence is reflected in sensitivity analysis. Through extensive model comparisons, it is observed that the model proposed in this paper exhibits strong generalization, without specific limitations on data size or feature count. It demonstrates good aggregation prediction performance for interval data. Moreover, it is applicable to various fields, including finance, environment, and others. Yixiang Wang, Zhicheng Hu |
Expert Syst. Appl. | 2 |
| 2024 | General Purpose Deep Learning Accelerator Based on Bit InterleavingabstractAlong with the rapid evolution of deep neural networks, the ever-increasing complexity imposes formidable computation intensity on the hardware accelerator. In this paper, we propose a novel computing philosophy called “bit interleaving” and the associate accelerator couple called “Bitlet” and Bitlet-X to maximally exploit the bit-level sparsity. Apart from the existing bit-serial/parallel accelerators, Bitlet leverages the abundant “sparsity parallelism” in the parameters to enforce the inference acceleration. Bitlet is versatile by supporting diverse precisions on a single platform, including floating-point 32 and fixed-point from 1b to 24b. The versatility enables Bitlet feasible for both efficient inference and training. Besides, by updating the key compute engine in the accelerator, Bitlet-X could furthermore improve the peak power consumption and efficiency for the inference-only scenario, with competitive accuracy. Empirical studies on 12 domain-specific deep learning applications highlight the following results: (1) up to 81×/21× energy efficiency improvement for training/inference over recent high-performance GPUs; (2) up to 15×/8× higher speedup/efficiency over state-of-the-art fixed-point accelerators; (3) 1.5mm2 area and scalable power consumption from 570mW (fp32) to 432mW (16b) and 365mW (8b) @28nm TSMC; (4) 1.3× improvement of the peak power efficiency for the Bitlet-X over Bitlet; (5) highly configurable justified by the ablation and sensitivity studies. Liang Chang 0002, Xin Zhao 0044, Zhicheng Hu, Jun Zhou 0017, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | SuperHCA: An Efficient Deep-Learning Edge Super-Resolution Accelerator With Sparsity-Aware Heterogeneous Core ArchitectureabstractDeep learning-based super-resolution (SR) generative models have recently emerged as a promising approach for generating high-quality images. While large SR networks can achieve a high peak signal-to-noise ratio (PSNR) to assess image quality, they often come with a high number of parameters, leading to increased computational and memory requirements that can be challenging to deploy on embedded hardware. In this study, we introduce the Anchor-Based Shuffle Net (ABSN), which is designed to create a hardware accelerator using a dynamic-scale fixed-point (DSFP) quantization method. Additionally, we incorporate dynamic quantization adaptation in the hardware design. Our Super-resolution Heterogeneous Accelerator, SuperHCA, utilizes a sparsity-aware heterogeneous architecture to optimize inference efficiency by distinguishing between dense and sparse workloads. We also propose Slice Layer Fusion (SLF) dataflow and feature-sharing bit interleaving (FSBI) methods in the heterogeneous cores to reduce on-chip buffer sizes. The SuperHCA achieves a frame rate of 91 fps at a target resolution of FHD, with the highest throughput area ratio (TAR) of 22.75 fps/mm2 compared to existing state-of-the-art works. Zhicheng Hu, Xin Zhao 0044, Liang Chang 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | HDSuper: High-Quality and High Computational Utilization Edge Super-Resolution Accelerator With Hardware-Algorithm Co-Design TechniquesabstractSuper-resolution (SR) techniques have been employed to construct high-definition images from low-quality images. Various neural networks have demonstrated excellent image-reconstruction quality in SR accelerators. However, deploying SR networks on edge devices is limited by resources and power consumption induced by significant algorithm parameters, computation complexity, and external memory accesses. This work explores the hardware algorithm co-design techniques to provide an end-to-end platform with a lightweight super-resolution network (LSR) and an efficient, high-quality SR accelerator HDSuper. For algorithm design, the improved depth-wise separable convolution and pixelshuffle layers are developed to reduce network size and computation complexity by considering the hardware constraints. Also, the improved channel attention (CA) blocks enhance the image reconstruction quality. For hardware accelerator design, we design a unified computing core (UCC) combined with an efficient flattening-and-allocation (F-A) mapping strategy to support various operators with high computational utilization. In addition, we design the patch computing scheme to reduce the external memory access of the hardware architecture. Based on the evaluation, the proposed algorithm achieves high-quality image reconstruction with$37.44dB$PSNR. Finally, the FPGA demonstration and ASIC layout under UMC 55nm are achieved with low power consumption ($2.08 W$and$152 mW$) under the lowest hardware resources compared to the state-of-the-art works. Xin Zhao 0044, Liang Chang 0002, Dongqi Fan, Zhicheng Hu, Ting Yue, Fengbin Tu, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | HDSuper: Algorithm-Hardware Co-design for Light-weight High-quality Super-Resolution AcceleratorabstractSuper-resolution (SR) networks have been gradually applied to embedded devices with good-quality image reconstruction. However, the hardware performance and power efficiency are limited by a large number of algorithm parameters, computation complexity, and hardware resources, obstructing the development of a high-quality SR accelerator. This paper proposes an end-to-end platform with a lightweight super-resolution network (LSR) and an efficient, high-quality super-resolution architecture HDSuper, to perform algorithm-hardware co-design for the SR accelerator. For algorithm design, we employ depth-wise separable convolution and pixelshuffle to reduce network size and computation complexity by considering the hardware constraints. For hardware design, we provide a unified computing core (UCC) combined with an efficient flattening-and-allocation (F-A) mapping strategy to support various operators with high computational utilization. We adopt the patch training method to reduce the external memory access of the hardware architecture. Based on the evaluation, the proposed algorithm achieves high-quality image reconstruction with 37.44dB PSNR. Finally, we implement the image reconstruction in FPGA demonstration, achieving high-quality image reconstruction with 2.08W power consumption under the lowest hardware resources compared to the state-of-the-art works. Liang Chang 0002, Xin Zhao 0044, Dongqi Fan, Zhicheng Hu, Jun Zhou 0017 |
DAC | 4 |
| 2023 | Digital Twins For Topology Density Map AnalysisabstractProperly analyzing spatiotemporal patterns is of paramount importance, especially in urban planning. In this paper, we introduce two digital twins to support the analysis of spatiotemporal data associated with urban topologies. In particular, both tools visually encode temporal changes in density maps constrained by a network. We introduce their software architectures and discuss their use in urban planning usage scenarios. Zhicheng Hu, Agus Hasan, Ricardo da Silva Torres |
ECMS | 1 |
| 2023 | Continuous triangular fuzzy generalized OWA operator and its application to combined prediction
Zhicheng Hu, Yixiang Wang |
Soft Comput. | 1 |
| 2018 | GANFuzz: a GAN-based industrial network protocol fuzzing frameworkabstractIn this paper, we attempt to improve industrial safety from the perspective of communication security. We leverage the protocol fuzzing technology to reveal errors and vulnerabilities inside implementations of industrial network protocols(INPs). Traditionally, to effectively conduct protocol fuzzing, the test data has to be generated under the guidance of protocol grammar, which is built either by interpreting the protocol specifications or reverse engineering from network traces. In this study, we propose an automated test case generation method, in which the protocol grammar is learned by deep learning. Generative adversarial network(GAN) is employed to train a generative model over real-world protocol messages to enable us to learn the protocol grammar. Then we can use the trained generative model to produce fake but plausible messages, which are promising test cases. Based on this approach, we present an automatical and intelligent fuzzing framework(GANFuzz) for testing implementations of INPs. Compared to prior work, GANFuzz offers a new way for this problem. Moreover, GANFuzz does not rely on protocol specification, so that it can be applied to both public and proprietary protocols, which outperforms many previous frameworks. We use GANFuzz to test several simulators of the Modbus-TCP protocol and find some errors and vulnerabilities. Zhicheng Hu, Jianqi Shi, Yanhong Huang, Jiawen Xiong, Xiangxing Bu |
CF | 1 |