EDBT 2026 Demo / reviewers in the wild / expert
Win-San Khwa
dblp:176/6156
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-6283-3564ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SATA: Sparsity-Aware Scheduling for Selective Token AttentionabstractTransformers have become the foundation of numerous state-of-the-art AI models across diverse domains, thanks to their powerful attention mechanism for modeling long-range dependencies. However, the quadratic scaling complexity of attention poses significant challenges for efficient hardware implementation. While techniques such as quantization and pruning help mitigate this issue, selective token attention offers a promising alternative by narrowing the attention scope to only the most relevant tokens, reducing computation and filtering out noise.In this work, we propose SATA, a locality-centric dynamic scheduling scheme that proactively manages sparsely distributed access patterns from selective Query-Key operations. By reordering operand flow and exploiting data locality, our approach enables early fetch and retirement of intermediate Query/Key vectors, improving system utilization. We implement and evaluate our token management strategy in a control and compute system, using runtime traces from selective-attention-based models. Experimental results show that our method improves system throughput by up to 1.76× and boosts energy efficiency by 2.94×, while incurring minimal scheduling overhead. Zhenkun Fan, Zishen Wan, Che-Kai Liu, Ashwin Sanjay Lele, Win-San Khwa, Meng-Fan Chang, Arijit Raychowdhury |
DATE | 5 |
| 2025 | Finding the Pareto Frontier of Low-Precision Data Formats and MAC Architecture for LLM InferenceabstractTo accelerate AI applications, numerous data formats and physical implementations of matrix multiplication have been proposed, creating a complex design space. This paper studies the efficient MAC implementation of the integer, floating-point, posit, and logarithmic number system (LNS) data formats and Microscaling (MX) and VectorScaled Quantization (VSQ) block data formats. We evaluate the area, power, and numerical accuracy (evaluated as signal-to-quantization noise ratio) of $\mathbf{3 5, 0 0 0}$ MAC designs spanning each data format and several key design parameters such as the inner product size and accumulation width. We find that for the same numerical accuracy, pareto optimal MAC designs with emerging data formats (LNS16, MXINT8, VSQINT4) achieve $1.8 \times 2.2 \times$, and $1.9 \times$ TOPs/W improvement compared to FP16, FP8, and FP4 dot product implementations. Brian Crafton, Xiaochen Peng, Xiaoyu Sun 0001, Ashwin Sanjay Lele, Win-San Khwa, Kerem Akarvardar |
DAC | 6 |
| 2024 | Non-volatile Memory Technologies for Edge AI ApplicationsabstractEmbedded non-volatile memories (NVMs) hold significant promise for advancing edge AI applications by offering unique advantages in near/in-memory computing-based accelerators. This invited paper delivers an in-depth review of the use cases for NVMs, emphasizing their potential to enhance area- and energy-efficiency in edge devices. We begin by outlining the latest advancements in TSMC NVM technologies and examining several NVM-based accelerator test chips enabled by the TSMC University Shuttle Program. Additionally, we delve into the tradeoffs involved in optimizing NVM devices, explore their potential for approximate computing applications, and assess the impact of NVM non-idealities on inference accuracy. Xiaoyu Sun 0001, Win-San Khwa, Xiaochen Peng, Meng-Fan Chang, Kerem Akarvardar |
ICCAD | 2 |
| 2024 | Estimating Power, Performance, and Area for On-Sensor Deployment of AR/VR Workloads Using an Analytical FrameworkabstractAugmented Reality and Virtual Reality have emerged as the next frontier of intelligent image sensors and computer systems. In these systems, 3D die stacking stands out as a compelling solution, enabling in situ processing capability of the sensory data for tasks such as image classification and object detection at low power, low latency, and a small form factor. These intelligent 3D CMOS Image Sensor (CIS) systems present a wide design space, encompassing multiple domains (e.g., computer vision algorithms, circuit design, system architecture, and semiconductor technology, including 3D stacking) that have not been explored in-depth so far. This article aims to fill this gap. We first present an analytical evaluation framework, STAR-3DSim, dedicated to rapid pre-RTL evaluation of 3D-CIS systems capturing the entire stack from the pixel layer to the on-sensor processor layer. With STAR-3DSim, we then propose several knobs for PPA (power, performance, area) improvement of the Deep Neural Network (DNN) accelerator that can provide up to 53%, 41%, and 63% reduction in energy, latency, and area, respectively, across a broad set of relevant AR/VR workloads. Last, we present full-system evaluation results by taking image sensing, cross-tier data transfer, and off-sensor communication into consideration. Xiaoyu Sun 0001, Xiaochen Peng, Sai Qian Zhang, Jorge Gomez 0002, Win-San Khwa, Syed Shakib Sarwar, Ziyun Li 0001, Weidong Cao 0001, Chiao Liu, Meng-Fan Chang, Barbara De Salvo, Kerem Akarvardar, H.-S. Philip Wong |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2023 | A Fully-Integrated Energy-Scalable Transformer Accelerator Supporting Adaptive Model Configuration and Word Elimination for Language Understanding on Edge DevicesabstractEfficient natural language processing on the edge is needed to interpret voice commands, which have become a standard way to interact with devices around us. Due to the tight power and compute constraints of edge devices, it is important to adapt the computation to the hardware conditions. We present a Transformer accelerator with a variable-depth adder tree to support different model dimensions, a SuperTransformer model from which Sub Transformers of various sizes can be sampled enabling adaptive model configuration, and a dedicated word elimination unit to prune redundant tokens. We achieve up to 6.9× scalability in network latency and energy between the largest and smallest Sub Transformers, under the same operating conditions. Word elimination can reduce network energy by 16%, with a 14.5% drop in F1 score. At 0.68V and 80MHz, processing a 32-length input with our custom 2-layer Transformer model for intent detection and slot filling takes 0.61ms and 1.6μJ. Zexi Ji, Hanrui Wang 0002, Miaorong Wang, Win-San Khwa, Meng-Fan Chang, Song Han 0003, Anantha P. Chandrakasan |
ISLPED | 4 |
| 2018 | Parallelizing SRAM arrays with customized bit-cell for binary neural networksabstractRecent advances in deep neural networks (DNNs) have shown Binary Neural Networks (BNNs) are able to provide a reasonable accuracy on various image datasets with a significant reduction in computation and memory cost. In this paper, we explore two BNNs: hybrid BNN (HBNN) and XNOR-BNN, where the weights are binarized to +1/-1 while the neuron activations are binarized to 1/0 and +1/-1 respectively. Two SRAM bit cell designs are proposed, namely, 6T SRAM for HBNN and customized 8T SRAM for XNOR-BNN. In our design, the high-precision multiply-and-accumulate (MAC) is replaced by bitwise multiplication for HBNN or XNOR for XNOR-BNN plus bit-counting operations. To parallelize the weighted sum operation, we activate multiple word lines in the SRAM array simultaneously and digitize the analog voltage developed along the bit line by a multi-level sense amplifier (MLSA). In order to partition the large matrices in DNNs, we investigate the impact of sensing bit-levels of MLSA on the accuracy degradation for different sub-array sizes and propose using the nonlinear quantization technique to mitigate the accuracy degradation. With 64×64 sub-array size and 3-bit MLSA, HBNN and XNOR-BNN architectures can minimize the accuracy degradation to 2.37% and 0.88%, respectively, for an inspired VGG-16 network on the CIFAR-10 dataset. Design space exploration of SRAM based synaptic architectures with the conventional row-by-row access scheme and our proposed parallel access scheme are also performed, showing significant benefits in the area, latency and energy-efficiency. Finally, we have successfully taped-out and validated the proposed HBNN and XNOR-BNN designs in TSMC 65 nm process with measured silicon data, achieving energy-efficiency >100 TOPS/W for HBNN and >50 TOPS/W for XNOR-BNN. Rui Liu 0005, Xiaochen Peng, Xiaoyu Sun 0001, Win-San Khwa, Xin Si, Jia-Jing Chen, Jia-Fang Li, Meng-Fan Chang, Shimeng Yu |
DAC | 4 |