VLDB 2026 Research / reviewers in the wild / expert
Zhiyong Li 0016
dblp:59/5464-16
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-0436-6670ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE AcceleratorabstractFully homomorphic encryption over torus (TFHE) enables the execution of arbitrary functions on encrypted data through programmable bootstrapping (PBS). However, performing all operations on ciphertext during PBS results in high computational and memory requirements, limiting the deployment of PBS in real-world scenarios. Previous TFHE accelerator designs have attempted to improve performance by employing specific dataflow and functional units, but these techniques may require large off-chip bandwidth or on-chip storage when scaling up computation capacity. Additionally, the design of specialized functional units may limit the utilization of computation units when facing dynamic secure parameter settings. To address these challenges and further improve PBS throughput in TFHE, we propose Matrix , an ASIC-based architecture that balances off-chip bandwidth and on-chip storage according to the execution flow of PBS. In Matrix , we utilize a unified special-prime-based processing element (PE) that achieves high utilization with minimal resource overhead. Furthermore, we propose a hybrid PBS dataflow that can efficiently reduce computation complexity and memory requirements. Compared to state-of-the-art TFHE accelerators, Matrix achieves 1.43 × -5.66 × throughput improvement for PBS. For ZAMA Deep-NN benchmark, we achieve 525.60× and 68.06× speedup compared to CPU and GPU, respectively. 1 Ling Liang 0003, Fahong Zhang 0004, Zhirui Li, Xin Fan 0009, Dimin Niu, Meng Li 0004, Zhiyong Li 0016, Zongwei Wang 0001, Hongzhong Zheng, Yimao Cai, Yuan Xie 0001 |
ACM Trans. Archit. Code Optim. | 9 |
| 2025 | A Tightly Coupled AI-ISP Vision ProcessorabstractTo achieve high-quality and high-resolution image processing, this work presents a novel vision processor that facilitates deep learning-enhanced image processing pipelines. At the system level, by identifying that a divide-and-conquer approach is essential to synergize both classical image processing and image enhancement networks, we develop a tightly coupled system with strip-tile conversion dataflow to enable fine-grained low-latency data interactions between image signal processors (ISPs) and the deep learning accelerator (DLA). At the architecture level, we design a comprehensive set of 21 efficient image processing modules to construct classical ISP pipelines, a tile-based strip layer fusion DLA specifically optimized for networks, and a programmable pixel pool that seamlessly supports the data access patterns of the ISP and the DLA. At the software and hardware co-design level, we propose a comprehensive optimization framework to address the implementation overhead of networks while maintaining the image quality. Finally, evaluations of the AI-ISP vision processor demonstrate 53.95% external memory access reduction and 35.51% latency reduction, delivering superior image quality with minimal on-chip memory overhead. A throughput of up to 168.5 frames per second facilitates efficient processing of ultra-high definition (UHD) resolution images. Hao Zhang 0126, Sicheng Li 0001, Yupeng Gui, Zhiyong Li 0016, Shusong Xu, Yanheng Lu, Dimin Niu, Hongzhong Zheng, Yen-Kuang Chen, Yuan Xie 0001, Yibo Fan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Zero-Shot Structure-Preserving Diffusion Model for High Dynamic Range Tone MappingabstractTone mapping techniques, aiming to convert high dynamic range (HDR) images to high-quality low dynamic range (LDR) images for display, play a more crucial role in real-world vision systems with the increasing application of HDR images. However, obtaining paired HDR and high-quality LDR images is difficult, posing a challenge to deep learning based tone mapping methods. To over-come this challenge, we propose a novel zero-shot tone mapping framework that utilizes shared structure knowl-edge, allowing us to transfer a pre-trained mapping model from the LDR domain to HDR fields without paired training data. Our approach involves decomposing both the LDR and HDR images into two components: structural in-formation and tonal information. To preserve the original image's structure, we modify the reverse sampling process of a diffusion model and explicitly incorporate the struc-ture information into the intermediate results. Additionally, for improved image details, we introduce a dual-control network architecture that enables different types of conditional inputs to control different scales of the output. Experimental results demonstrate the effectiveness of our approach, surpassing previous state-of-the-art methods both qualitatively and quantitatively. Moreover, our model ex-hibits versatility and can be applied to other low-level vi-sion tasks without retraining. The code is available at https://github.com/ZSDM-HDRIZero-Shot-Diffusion-HDR. Ruoxi Zhu, Shusong Xu, Peiye Liu, Sicheng Li 0001, Yanheng Lu, Dimin Niu, Zihao Liu 0015, Zihao Meng, Zhiyong Li 0016, Xinhua Chen, Yibo Fan |
CVPR | 9 |
| 2024 | NeuGPU: A Neural Graphics Processing Unit for Instant Modeling and Real-Time Rendering on Mobile AR/VR Devicesabstract•Why 3D Modeling using Neural Radiance Field? Junha Ryu, Hankyul Kwon, Wonhoon Park, Zhiyong Li 0016, Beomseok Kwon, Donghyeon Han, Dongseok Im, Sangyeob Kim, Hyungnam Joo, Hoi-Jun Yoo |
HCS | 4 |
| 2023 | Sibia: Signed Bit-slice Architecture for Dense DNN Acceleration with Slice-level Sparsity ExploitationabstractDeep neural networks (DNNs) have achieved high performance in many AI fields such as 1-D language, 2-D image, and 3-D point cloud processing applications. Since recent DNN tasks require dense matrix operations with various bit-precision and non-ReLU activation functions, mobile neural processing units (NPUs) suffer from the acceleration of diverse DNN tasks within their limited hardware resources and power budget. Although bit-slice architectures benefit from slice-level computation and slice-level sparsity exploitation, the conventional bit-slice representation is inefficient in bit-slice architectures resulting in poor dense DNN execution. This paper proposes the efficient signed bit-slice architecture, Sibia, with the signed bit-slice representation (SBR) for efficient dense DNN acceleration. The SBR adds a sign bit to each bit-slice and changes signed 11112bit-slice to 00002by borrowing a value of 1 from its lower order of the bit-slice. This scheme generates large numbers of zero bit-slices in dense DNNs even not relying on accuracy-sensitive pruning methods or retraining processes. Moreover, the SBR balances positive and negative values of 2’s complement data, allowing accurate bit-slice-based output speculation that pre-computes high orders of bit-slices. Sibia integrates the signed multiplier-and-accumulate (MAC) units for efficient signed bit-slice computations, and the flexible zero skipping processing element (PE) supports the zero input bit-slice skipping and output skipping for high throughput and energy-efficiency. Additionally, the dynamic sparsity monitoring unit monitors sparsity ratio between input and weight data and determines the more sparse one for zero bit-slice skipping. The heterogeneous network-on-chip (NoC) benefits from data reusability during bit-slice computation, reducing transmission bandwidth. Finally, Sibia outperforms the previous bit-slice architecture, Bit-fusion, over 3.65× higher area-efficiency, 3.88× higher energy-efficiency, and 5.35× higher throughput. Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Hoi-Jun Yoo |
HPCA | 3 |
| 2023 | A Reconfigurable 1T1C eDRAM-based Spiking Neural Network Computing-In-Memory Processor for High System-Level EfficiencyabstractSpiking Neural Network (SNN) Computing-In-Memory (CIM) was proposed for high macro-level energy efficiency. However, system-level energy efficiency is limited by EMA due to a large intermediate activation footprint requirement. To reduce the EMA, a large capacity SNN CIM is needed to load tons of weights in the CIM. This paper proposes a high-density 1T1C eDRAM-based SNN CIM processor for supporting high system-level energy efficiency with two key features: 1) High-density and low-power Reconfigurable Neuro-Cell Array (ReNCA) for memory and SNN peripheral logic using a charge pump and reusing 1T1C cell array, achieving 41% area and 90% power reduction compared to previous work. 2) Reconfigurable CIM architecture with dual-mode ReNCA and Dynamic Adjustable Neuron Link (DAN Link) for layer fusion increases system-level efficiency including intermediate and weight EMA. It achieves$10\times$higher state-of-the-art system-level energy efficiency including EMA. Seryeong Kim, Soyeon Um, Zhiyong Li 0016, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo |
ISCAS | 5 |
| 2022 | DSPU: A 281.6mW Real-Time Deep Learning-Based Dense RGB-D Data Acquisition with Sensor Fusion and 3D Perception System-on-Chipabstract3D Data in Mobile Platforms Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Donghyeon Han, Jinsu Lee, Wonhoon Park, Hankyul Kwon, Hoi-Jun Yoo |
HCS | 3 |
| 2022 | An Efficient High-quality FHD Super-resolution Mobile Accelerator SoC with Hybrid-precision and Energy-efficient CacheabstractSuper-Resolution on Mobile Platform Zhiyong Li 0016, Dongseok Im, Donghyeon Han, Hoi-Jun Yoo |
HCS | 1 |
| 2022 | A 161.6 TOPS/W Mixed-mode Computing-in-Memory Processor for Energy-Efficient Mixed-Precision Deep Neural NetworksabstractA Mixed-mode Computing-in memory (CIM) processor for the mixed-precision Deep Neural Network (DNN) processing is proposed. Due to the bit-serial processing for the multi-bit data, the previous CIM processors could not exploit the energy-efficient computation of mixed-precision DNNs. This paper proposes an energy-efficient mixed-mode CIM processor with two key features: 1) Mixed-Mode Mixed-precision CIM (M3-CIM) which achieves 55.46% energy efficiency improvement. 2) Digital-CIM for In-memory MAC for the increased throughput of M3-CIM. The proposed CIM processor was simulated in 28nm CMOS technology and occupies 1.96 mm2. It achieves a state-of-the-art energy efficiency of 161.6 TOPS/W with 72.8% accuracy at ImageNet (ResNet50). Wooyoung Jo, Juhyeong Lee, Soyeon Um, Zhiyong Li 0016, Hoi-Jun Yoo |
ISCAS | 5 |
| 2021 | A 3.6 TOPS/W Hybrid FP-FXP Deep Learning Processor with Outlier Compensation for Image-to-Image ApplicationabstractA Hybrid floating-point (FP) and fixed-point (FXP) deep learning processor with an outlier-aware channel splitting algorithm is proposed for image-to-image applications on mobile devices. Since the high quality of the reconstructed image through deep learning based image-to-image application requires high bit-precision (> FP16), the mobile processor suffers from the high computation power and large external memory access (EMA). In this work, the proposed algorithm reduces 16-bit FP data to 8-bit FXP data, and only few outliers (2. The hierarchical processor successfully demonstrates the x 4 scale Full-HD super-resolution generation achieving 76 frames-per-second (fps) with 133.3 mW power-consumption at 0.9 V supply and 3.6 TOPS/W of energy-efficiency which is × 3.27 higher than the previous 16-bit FXP processor. Zhiyong Li 0016, Dongseok Im, Jinsu Lee, Hoi-Jun Yoo |
ISCAS | 1 |