Ka-Fai Un

dblp:75/9432 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-7574-4755ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 A Topology-Aware Reinforcement Learning Framework for 2.4-GHz VSWR-Robust Power Amplifier Matching Network Design
Bingbing Zhao, Wei-Han Yu, Fábio Passos, Ka-Fai Un, Rui Paulo Martins, Pui-In Mak
ISCAS5
2026 A 3.51 TOPS/mm2 Transformer Accelerator Exploiting Bipolar Sparsity and Approximate Gating
abstract
Transformer models excel at natural language processing tasks but are challenging to deploy on edge devices due to high memory and computation demands. To address this, we proposed an energy- and area-efficient transformer accelerator. We identify ‘0’ bits in positive and ‘1’ bits in negative 2’s complement activation values as bipolar sparsity. This form of sparsity shares a larger proportion than the traditional bit-level sparsity in transformer models. We propose a bipolar sparsity compressor (BSC) together with a bipolar processing element (BPE) to detect and skip the bipolar sparsity in a bit-group (BG) level during the inference. It reduces a large proportion of ineffective computations and improves throughput. The significant sparsity scheduling (SSS) dynamically adjusts broadcast settings based on BG-level sparsity ratios, balancing the sparsity skipping and memory access. Furthermore, an importance approximate gating (IAG) filters out unimportant tokens/heads during the attention computation by reusing sparsity information from the BSC, further reducing processing latency and energy consumption. Implemented in a 28nm process, the proposed accelerator achieves$7.62\times $and$12.43\times $throughput improvements on RoBERTa-B and GPT2-xl, respectively. The area efficiency reaches up to 3.51 TOPS/mm2due to skipping a large proportion of bipolar sparsity, reaching$4.83\times$and$5.85\times $higher compared with the state-of-the-art approximate computing accelerator and the computing in memory accelerator on the benchmark model.
Zhongyu Zhao, Rujian Cao, Ka-Fai Un, Wei-Han Yu, Rui Paulo Martins, Pui-In Mak
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 A 97.8 GOPS/W FPGA-Based Residual-Block-Aware CNN Accelerator Featuring Multi-Clock PW2 Pipeline and Adaptive-Resolution Quantization
abstract
Enhancing the energy efficiency for the residual block is crucial for an energy-efficient deep neural network accelerator. This paper presents a multi-clock pointwise-pointwise (MCPW2) technique to process the adjacent PW convolution layers across residual blocks, reducing up to 75.0% DRAM access for the intermediate feature maps while securing >88.1% processing element (PE) utilization. Moreover, we introduce a dual-precision packing (DPP) DSP array to compute multiple 4/8-bit multiplications in a shared DSP, improving the accuracy by 1.5% (ImageNet) using low-precision residual distillation (RD) with adaptive-resolution quantization. The DPP DSP and adaptive-resolution RD boost the DSP efficiency up to$4.0\times $, reduce DRAM access by 50.0%, and improve the throughput by$\gt 2.7\times $. We also propose a dynamic accumulator/multiplier (A/M) DSP reconfiguration scheme to dynamically adjust the level of parallelism along the input/output channel dimensions. It also increases the PE utilization by$1.8\times $for the depthwise (DW) convolution layers with 33% less hardware resource overhead. Implemented on Xilinx VC709, the proposed accelerator achieves PE utilization of >93.0%, a DSP efficiency gain of$\gt 2.9\times $, and a throughput improvement on benchmarked networks of$4.9\times $while exhibiting an energy efficiency of 97.8 GOPs/W and a normalized throughput of 1.18 GOPS/DSP.
Jixuan Li, Ka-Fai Un, Wei-Han Yu, Rui Paulo Martins, Pui-In Mak
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 GSLP-CIM: A 28-nm Globally Systolic and Locally Parallel CNN/Transformer Accelerator With Scalable and Reconfigurable eDRAM Compute-in-Memory Macro for Flexible Dataflow
abstract
This article reports a globally systolic and locally parallel (GSLP) convolutional NN (CNN) and Transformer accelerator based on the scalable and reconfigurable (SR) embedded dynamic random-access memory (eDRAM) compute-in-memory (CIM) macro. It features: 1) a GSLP architecture employs systolic CIM macros with the reconfigurable inter-CIM network to support flexible dataflow, including weight stationary (WS), output stationary (OS), and Row stationary (RS); 2) an SR-CIM macro features reconfigurable weight/input/output memory ratio to maximize the related data reuse in different dataflow; 3) a high-density 3T eDRAM-CIM cell to further improve the density of the accelerator; 4) an area-efficient in-memory accumulator (IMA) to save the area and power overhead of the digital accumulation in each CIM macro. Prototyped in 28-nm CMOS process, the proposed GSLP-CIM accelerator exhibits a 4b peak throughput density of 0.16 TOPS/mm2 and a 4b peak compute energy efficiency of 3.55 TOPS/W. Specifically, evaluated with ResNet-50@ImageNet and ViT-B@ImageNet, this work reaches the system throughput of 24.5 and 5.66 inferences per second (IPS), the system throughput density of 19.3 IPS/mm2 and 4.46 IPS/mm2, the system compute energy efficiency of 423.9 inferences per watt (IPW) and 97.6 IPW, respectively.
Wei-Han Yu, Ka-Fai Un, Rui Paulo Martins, Pui-In Mak
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 A 362-TOPS/W Mixed-Signal MAC Macro With Sampling-Weight-Nonlinearity Cancellation and Dynamic-Amplified Accumulation
abstract
This work presents a high energy-efficiency mixed-signal multiply-and-accumulate (MAC) macro in charge-domain for machine learning (ML) systems. It involves crucial features aimed at enhancing energy efficiency, throughput, and area efficiency, namely: 1) a parallel-serial (ParSer) scheme to augment the throughput by parallel input channels and reduce the power via serial analog accumulation rather than digital summation; 2) the weight-independent parallel digital-to-analog converter (DAC) sampling (WIPDS) to cancel weight nonlinearity during sampling and allow for resource-efficient DAC, significantly saving power and area; 3) a high energy-efficiency dynamic amplifier (DA) introduced to improve drivability and counteract attenuation of the serial accumulation, thereby attaining the desired accuracy with relaxed the afterward analog-to-digital converter (ADC) resolution and consequently reducing power consumption; 4) an optimized SAR ADC to reach higher energy efficiency. Fabricated in 28-nm CMOS technology, the prototype exhibits a peak energy and area efficiency of 362 TOPS/W and 3.23 TOPS/mm$^{2}$, respectively.
Xueru Cen, Ka-Fai Un, Mingqiang Guo, Liang Qi 0002, Rui Paulo Martins, Sai-Weng Sin
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 A Delta-Sigma-Based Computing-In-Memory Macro Targeting Edge Computation
abstract
Many applications of machine learning (ML) have been integrated into edge devices with their low communication latency. In edge computation, the reprocessing of redundant data results in considerable energy waste. The prior research utilized a digital-delta-digital-sigma computing-in-memory (CIM) scheme to mitigate this redundancy. However, the 7-bit LSB-first ADC resulting from the near-zero-mean output distribution led to excessive area and latency overhead. The following digital adder further induced power consumption and latency. We propose a digital-delta-analog-sigma CIM macro incorporating an analog sigma converter (SC) for edge computation, involving a switch-capacitor integrator with a floating inverter amplifier (FIA) and a quantizer. The increased analog swing of the sigma integrator leads to the expanded output distribution, thereby maintaining comparable accuracy with a relaxed quantizer resolution. The simulation demonstrates that our strategy contributes to a 57.5% reduction in latency, a resolution decrease of 2 bits, and better energy efficiency. These improvements can potentially enhance energy efficiency and computational speed in edge computation devices.
Ka-Fai Un, Mingqiang Guo, Liang Qi 0002, Dengke Xu, Weibing Zhao, Rui Paulo Martins, Franco Maloberti, Sai-Weng Sin
ISCAS2
2024 CLUT-CIM: A Capacitance Lookup Table-Based Analog Compute-in-Memory Macro With Signed-Channel Training and Weight Updating for Nonuniform Quantization
abstract
Compute-in-memory (CIM) is a promising approach for realizing energy-efficient deep neural network (DNN) accelerators. Previous CIM works focusing on uniform quantization (UQ) demonstrated a higher Multiply-accumulate (MAC) precision requirement to maintain DNN inferencing accuracy, resulting lower energy efficiency. The nonuniform quantization (NUQ) has proved to require lower precision than UQ, while the existing implementations are based on high precision digital lookup table (LUT) (e.g., 16-bit), leading to large energy and area overhead for multiplier. This work presents CLUT-CIM fabricated under 28-nm CMOS featuring: 1) a capacitance LUT (CLUT)-based NUQ MAC circuit with thermometer coding scheme for weight and input activation that avoids digital LUT and reduces the energy and area overhead; 2) a signed-channel training (SCT) method that reduces the switching activity of computation to improve the energy efficiency; 3) a dual-port 6T-SRAM array to enable simultaneously weight updating and CIM operations, enhancing the memory utilization and CIM throughput. Under 3-bit NUQ precision, the peak energy efficiency is 114.3 TOPS/W, and peak throughput density is 31.78 TOPS/mm2.
Yuzhao Fu, Jixuan Li, Wei-Han Yu, Ka-Fai Un, Chi-Hang Chan, Yan Zhu 0001, Rui Paulo Martins, Pui-In Mak
IEEE Trans. Circuits Syst. I Regul. Pap.4
2016 Time-domain I/Q-LOFT compensator using a simple envelope detector for a sub-GHz IEEE 802.11af WLAN transmitter
abstract
This paper proposes a hardware-efficient time-domain scheme to digitally compensate the I/Q imbalance and LO feedthrough (LOFT) of a sub-GHz wideband transmitter for the IEEE 802.11af WLAN. A simple envelope detector is the only analog part. The parameters are updated by Least-Mean-Square and estimated efficiently in time domain by using COordinate Rotation DIgital Computer (CORDIC), saving the training time and power consumption. The measured wideband image-rejection ratio (IRR) and LO-leakage-rejection ratio (LRR) are improved from 18.9 to 41.3 dB, and 20.4 to 37.9 dB, respectively.
Chak-Fong Cheang, Ka-Fai Un, Pui-In Mak, Rui Paulo Martins
ASP-DAC2
2010 SC biquad filter with hybrid utilization of OpAmp and comparator-based circuit
abstract
This paper proposes a differential switched-capacitor (SC) biquad filter exploiting a hybrid structure. The 1stactive core is an operational amplifier (OpAmp) whereas the 2ndis an improved comparator-based circuit (CBC). The advantages of this new structure are justified by the reductions of power and transistor sizes. Optimized in a 65-nm CMOS process, when compared with a typical dual-OpAmp design, the proposed filter saves 19% power and 18% transistor area. The filter clocked at 40 MHz achieves 61.7-dB IM2 and 62.5-dB IM3 while drawing 2.23 mA from a 1.2-V supply. This hybrid SC biquad can gain further momentum for filters that request numerous biquads in cascade to attain higher selectivity.
Miguel A. Martins, Ka-Fai Un, Pui-In Mak, Rui Paulo Martins
ISCAS2
2009 An Open-loop Octave-phase Local-oscillator Generator with High-precision Correlated Phases for VHF/UHF Mobile-TV Tuners
abstract
An octave-phase local-oscillator (LO) generator for 170-to-860-MHz mobile-TV tuners is described. It is intended to incorporate with a polyphase mixer scheme for rejecting the 3rdand 5thharmonics of the LO that is critical for wideband reception. The circuit is structured by a cascade of 7 inverter-based phase correctors to generate a set of LO signals with octave phases in an open-loop formation, resulting in 4times relaxation of the synthesizer's operating frequency when comparing with the conventional closed-loop form that requires the use of a div-by-4 frequency divider. Optimized in a 90-nm CMOS process, the achieved phase precisions are plusmn0.8deg in VHF III (170 to 245 MHz) band and UHF (470 to 860 MHz) band while drawing 2.3 to 5.1 mA from a 1-V supply.
Ka-Fai Un, Pui-In Mak, Rui Paulo Martins
ISCAS1