Wenlun Zhang

dblp:320/1477 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0003-1570-082XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 BitROM: Weight Reload-Free CiROM Architecture Towards Billion-Parameter 1.58-bit LLM Inference
abstract
Compute-in-Read-Only-Memory (CiROM) accelerators offer outstanding energy efficiency for CNNs by eliminating runtime weight updates. However, their scalability to Large Language Models (LLMs) is fundamentally constrained by their vast parameter sizes. Notably, LLaMA-7B—the smallest model in LLaMA series—demands more than 1,000 cm<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> of silicon area even in advanced CMOS nodes. This paper presents BitROM, the first CiROM-based accelerator that overcomes this limitation through co-design with BitNet’s 1.58-bit quantization model, enabling practical and efficient LLM inference at the edge. BitROM introduces three key innovations: 1) a novel Bidirectional ROM Array that stores two ternary weights per transistor; 2) a Tri-Mode Local Accumulator optimized for ternary-weight computations; and 3) an integrated Decode-Refresh (DR) eDRAM that supports on-die KV-cache management, significantly reducing external memory access during decoding. In addition, BitROM integrates LoRA-based adapters to enable efficient transfer learning across various downstream tasks. Evaluated in 65 nm CMOS, BitROM achieves 20.8 TOPS/W and a bit density of 4,967 kB/mm<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup>—offering a $10 \times$ improvement in area efficiency over prior digital CiROM designs. Moreover, the DR eDRAM contributes to a 43.6% reduction in external DRAM access, further enhancing deployment efficiency for LLMs in edge applications. Code is available at https://github.com/Wenlun-Zhang/BitROM
Wenlun Zhang, Shimpei Ando, Kentaro Yoshioka
ASP-DAC1
2026 SNNQ: Post-Training Quantization Towards Ultra Low-Bit and Fast Spiking Neural Networks
Wenlun Zhang, Kentaro Yoshioka
ISCAS1
2026 MCRA: Multicolumn Residue Accumulation Analog Compute-in-Memory Architecture With Time-Domain M-Input ΣΔ ADC
abstract
Analog compute-in-memory (ACIM) architectures offer significant throughput and energy benefits by performing multiplication-and-accumulation (MAC) operations directly within memory arrays. However, their overall efficiency is fundamentally constrained by the high power consumption of the per-column high-resolution analog-to-digital converters (ADCs) required to support modern DNNs (e.g., transformers), many of which demand both high computational precision and large-column throughput. In conventional ADC designs, energy in the noise-limited regime scales near-exponentially, typically by$4\times $per additional bit, making high-resolution ADCs on every column power-prohibitive. This article proposes a high-precision and power-efficient multicolumn residue accumulation (MCRA) ACIM architecture to efficiently support precision-demanding modern DNNs. Each column uses a low-resolution coarse ADC (cADC), while the per-column residuals are accumulated and further quantized by an energy-efficient time-domain multi-input incremental sigma–delta (Mi-$\Sigma \Delta $) fine ADC (fADC). This approach amortizes the near-exponential energy growth across columns, while exploiting the more favorable power-resolution scaling of the time-domain Mi-$\Sigma \Delta $quantization. Postlayout simulations demonstrate a 66.2-dB signal-to-noise-and-distortion ratio (SNDR) per column at only 1/21 the energy of the baseline with per-column high-resolution ADCs, and$0.405\times $(1/2.47) the energy of an energy-saving ADC per column, which achieves a similar SNDR. Circuit- and system-level simulations demonstrate that our MCRA CIM architecture achieves negligible accuracy degradation in precision-demanding ViT tasks while delivering high energy and area efficiency.
Wenlun Zhang, Shimpei Ando, Zhongfeng Wang 0001, Jun Lin 0001, Kentaro Yoshioka
IEEE Trans. Very Large Scale Integr. Syst.2
2025 GSMM: Efficient Global Sparsification for Resource-Conscious Multimodal Models
abstract
Large Multimodal Models (LMMs) are increasingly essential in various real-time applications, yet their substantial parameter counts and complex architectures pose significant challenges. Traditional global compression methods often rely on trial-and-error experimentation, leading to inefficiencies. In this paper, we introduce new GS-MM, an Efficient Global Sparsification technique tailored for resource-conscious multimodal models. GS-MM assigns global sparsification strategies by extracting primitive importance from the model’s components. We first derive the importance of elements from the mapping values of fully weighted activations, based on the weights of the elements. Subsequently, we compute the average multimodal importance to establish a global importance score. This score is then linearly mapped to determine the global allocation ratio, enabling the realization of global sparsity in LMMs. We demonstrate the effectiveness of our approach through extensive experiments on diverse benchmarks, including visual question-answering and reasoning tasks. Our pruned models consistently outperform conventional pruning methods, setting new standards for compressed model performance. Notably, our approach exhibits remarkable resilience to increasing sparsity ratios, preserving model quality even under extreme compression.
Wenlun Zhang, Haoran Pang, Yucai Zhou, Shixiao Wang, Luking Li
ICASSP1
2025 AHCPTQ: Accurate and Hardware-Compatible Post-Training Quantization for Segment Anything Model
Wenlun Zhang, Yunshan Zhong, Shimpei Ando, Kentaro Yoshioka
ICCV1
2025 LiSA: Leveraging Link Recommender to Attack Graph Neural Networks via Subgraph Injection
Wenlun Zhang, Enyan Dai, Kentaro Yoshioka
PAKDD (4)1
2025 ASiM: Modeling and Analyzing Inference Accuracy of SRAM-Based Analog CiM Circuits
abstract
Static random-access memory (SRAM)-based analog compute-in-memory (ACiM) demonstrates promising energy efficiency for deep neural network (DNN) processing. Nevertheless, efforts to optimize efficiency frequently compromise accuracy, and this trade-off remains insufficiently studied due to the difficulty of performing full-system validation. Specifically, existing simulation tools rarely target SRAM-based ACiM and exhibit inconsistent accuracy predictions, highlighting the need for a standardized, SRAM compute-in-memory (CiM) circuit-aware evaluation methodology. This article presents ASiM, a simulation framework for evaluating inference accuracy in SRAM-based ACiM systems. ASiM captures critical effects in SRAM-based analog compute in memory systems, such as analog-to-digital converter (ADC) quantization, bit-parallel encoding, and analog noise, which must be modeled with high fidelity due to their distinct behavior in charge-domain architectures compared to other memory technologies. ASiM supports a wide range of modern DNN workloads, including CNN and Transformer-based models such as ViT, and scales to large-scale tasks like ImageNet classification. Our results indicate that bit-parallel encoding can improve energy efficiency with only modest accuracy degradation; however, even 1 LSB of analog noise can significantly impair inference performance, particularly in complex tasks such as ImageNet. To address this, we explore hybrid analog-digital execution and majority voting schemes, both of which enhance robustness without negating energy savings. ASiM bridges the gap between hardware design and inference performance, offering actionable insights for energy-efficient, high-accuracy ACiM deployment. The code is available athttps://github.com/Keio-CSG/ASiM
Wenlun Zhang, Shimpei Ando, Yung-Chin Chen, Kentaro Yoshioka
IEEE Trans. Very Large Scale Integr. Syst.1
2024 PACiM: A Sparsity-Centric Hybrid Compute-in-Memory Architecture via Probabilistic Approximation
abstract
Approximate computing emerges as a promising approach to enhance the efficiency of compute-in-memory (CiM) systems in deep neural network processing. However, traditional approximate techniques often significantly trade off accuracy for power efficiency, and fail to reduce data transfer between main memory and CiM banks, which dominates power consumption. This paper introduces a novel probabilistic approximate computation (PAC) method that leverages statistical techniques to approximate multiply-and-accumulation (MAC) operations, reducing approximation error by 4× compared to existing approaches. PAC enables efficient sparsity-based computation in CiM systems by simplifying complex MAC vector computations into scalar calculations. Moreover, PAC enables sparsity encoding and eliminates the LSB activations transmission, significantly reducing data reads and writes. This sets PAC apart from traditional approximate computing techniques, minimizing not only computation power but also memory accesses by 50%, thereby boosting system-level efficiency. We developed PACiM, a sparsity-centric architecture that fully exploits sparsity to reduce bit-serial cycles by 81% and achieves a peak 8b/8b efficiency of 14.63 TOPS/W in 65 nm CMOS while maintaining high accuracy of 93.85/72.36/66.02% on CIFAR-10/CIFAR-100/ImageNet benchmarks using a ResNet-18 model, demonstrating the effectiveness of our PAC methodology. Software simulation framework is available at GitHub.
Wenlun Zhang, Shimpei Ando, Yung-Chin Chen, Satomi Miyagi, Shinya Takamaeda-Yamazaki, Kentaro Yoshioka
ICCAD1