EDBT 2026 Demo / reviewers in the wild / expert
Wei-Han Yu
dblp:125/2628
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-9079-5227ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Live Demonstration: A 1TX/4RX Radar with Frequency-Dimension Virtual Aperture Expansion
Ruilin Liao, Jingzhi Zhang, Wei-Han Yu, Yue Song 0003, Hongyang An, Huihua Liu, Kai Kang 0001 |
ISCAS | 4 |
| 2026 | A Topology-Aware Reinforcement Learning Framework for 2.4-GHz VSWR-Robust Power Amplifier Matching Network Design
Bingbing Zhao, Wei-Han Yu, Fábio Passos, Ka-Fai Un, Rui Paulo Martins, Pui-In Mak |
ISCAS | 3 |
| 2026 | A 3.51 TOPS/mm2 Transformer Accelerator Exploiting Bipolar Sparsity and Approximate GatingabstractTransformer models excel at natural language processing tasks but are challenging to deploy on edge devices due to high memory and computation demands. To address this, we proposed an energy- and area-efficient transformer accelerator. We identify ‘0’ bits in positive and ‘1’ bits in negative 2’s complement activation values as bipolar sparsity. This form of sparsity shares a larger proportion than the traditional bit-level sparsity in transformer models. We propose a bipolar sparsity compressor (BSC) together with a bipolar processing element (BPE) to detect and skip the bipolar sparsity in a bit-group (BG) level during the inference. It reduces a large proportion of ineffective computations and improves throughput. The significant sparsity scheduling (SSS) dynamically adjusts broadcast settings based on BG-level sparsity ratios, balancing the sparsity skipping and memory access. Furthermore, an importance approximate gating (IAG) filters out unimportant tokens/heads during the attention computation by reusing sparsity information from the BSC, further reducing processing latency and energy consumption. Implemented in a 28nm process, the proposed accelerator achieves$7.62\times $and$12.43\times $throughput improvements on RoBERTa-B and GPT2-xl, respectively. The area efficiency reaches up to 3.51 TOPS/mm2due to skipping a large proportion of bipolar sparsity, reaching$4.83\times$and$5.85\times $higher compared with the state-of-the-art approximate computing accelerator and the computing in memory accelerator on the benchmark model. Zhongyu Zhao, Rujian Cao, Ka-Fai Un, Wei-Han Yu, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | A 97.8 GOPS/W FPGA-Based Residual-Block-Aware CNN Accelerator Featuring Multi-Clock PW2 Pipeline and Adaptive-Resolution QuantizationabstractEnhancing the energy efficiency for the residual block is crucial for an energy-efficient deep neural network accelerator. This paper presents a multi-clock pointwise-pointwise (MCPW2) technique to process the adjacent PW convolution layers across residual blocks, reducing up to 75.0% DRAM access for the intermediate feature maps while securing >88.1% processing element (PE) utilization. Moreover, we introduce a dual-precision packing (DPP) DSP array to compute multiple 4/8-bit multiplications in a shared DSP, improving the accuracy by 1.5% (ImageNet) using low-precision residual distillation (RD) with adaptive-resolution quantization. The DPP DSP and adaptive-resolution RD boost the DSP efficiency up to$4.0\times $, reduce DRAM access by 50.0%, and improve the throughput by$\gt 2.7\times $. We also propose a dynamic accumulator/multiplier (A/M) DSP reconfiguration scheme to dynamically adjust the level of parallelism along the input/output channel dimensions. It also increases the PE utilization by$1.8\times $for the depthwise (DW) convolution layers with 33% less hardware resource overhead. Implemented on Xilinx VC709, the proposed accelerator achieves PE utilization of >93.0%, a DSP efficiency gain of$\gt 2.9\times $, and a throughput improvement on benchmarked networks of$4.9\times $while exhibiting an energy efficiency of 97.8 GOPs/W and a normalized throughput of 1.18 GOPS/DSP. Jixuan Li, Ka-Fai Un, Wei-Han Yu, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | GSLP-CIM: A 28-nm Globally Systolic and Locally Parallel CNN/Transformer Accelerator With Scalable and Reconfigurable eDRAM Compute-in-Memory Macro for Flexible DataflowabstractThis article reports a globally systolic and locally parallel (GSLP) convolutional NN (CNN) and Transformer accelerator based on the scalable and reconfigurable (SR) embedded dynamic random-access memory (eDRAM) compute-in-memory (CIM) macro. It features: 1) a GSLP architecture employs systolic CIM macros with the reconfigurable inter-CIM network to support flexible dataflow, including weight stationary (WS), output stationary (OS), and Row stationary (RS); 2) an SR-CIM macro features reconfigurable weight/input/output memory ratio to maximize the related data reuse in different dataflow; 3) a high-density 3T eDRAM-CIM cell to further improve the density of the accelerator; 4) an area-efficient in-memory accumulator (IMA) to save the area and power overhead of the digital accumulation in each CIM macro. Prototyped in 28-nm CMOS process, the proposed GSLP-CIM accelerator exhibits a 4b peak throughput density of 0.16 TOPS/mm2 and a 4b peak compute energy efficiency of 3.55 TOPS/W. Specifically, evaluated with ResNet-50@ImageNet and ViT-B@ImageNet, this work reaches the system throughput of 24.5 and 5.66 inferences per second (IPS), the system throughput density of 19.3 IPS/mm2 and 4.46 IPS/mm2, the system compute energy efficiency of 423.9 inferences per watt (IPW) and 97.6 IPW, respectively. Wei-Han Yu, Ka-Fai Un, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | CLUT-CIM: A Capacitance Lookup Table-Based Analog Compute-in-Memory Macro With Signed-Channel Training and Weight Updating for Nonuniform QuantizationabstractCompute-in-memory (CIM) is a promising approach for realizing energy-efficient deep neural network (DNN) accelerators. Previous CIM works focusing on uniform quantization (UQ) demonstrated a higher Multiply-accumulate (MAC) precision requirement to maintain DNN inferencing accuracy, resulting lower energy efficiency. The nonuniform quantization (NUQ) has proved to require lower precision than UQ, while the existing implementations are based on high precision digital lookup table (LUT) (e.g., 16-bit), leading to large energy and area overhead for multiplier. This work presents CLUT-CIM fabricated under 28-nm CMOS featuring: 1) a capacitance LUT (CLUT)-based NUQ MAC circuit with thermometer coding scheme for weight and input activation that avoids digital LUT and reduces the energy and area overhead; 2) a signed-channel training (SCT) method that reduces the switching activity of computation to improve the energy efficiency; 3) a dual-port 6T-SRAM array to enable simultaneously weight updating and CIM operations, enhancing the memory utilization and CIM throughput. Under 3-bit NUQ precision, the peak energy efficiency is 114.3 TOPS/W, and peak throughput density is 31.78 TOPS/mm2. Yuzhao Fu, Jixuan Li, Wei-Han Yu, Ka-Fai Un, Chi-Hang Chan, Yan Zhu 0001, Rui Paulo Martins, Pui-In Mak |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |