EDBT 2026 Demo / reviewers in the wild / expert
Wei-Hsing Huang
dblp:120/6857
· DBLP profile ↗
6ranked-venue papers
4as first author
4since 2021 · last 2026
0009-0008-2405-2936ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling Context-Switchable Monolithic 3D FPGA Design Using Bistable Ferroelectric Inverters
Faaiq G. Waqar, Matthew Chen, Zifan He, Zishen Wan, Minji Shon, Wei-Hsing Huang, Jason Cong, Shimeng Yu |
FCCM | 6 |
| 2026 | 3DGauCIM: Accelerating Static/Dynamic 3D Gaussian Splatting via Digital CIM for High Frame Rate Real-Time Edge RenderingabstractDynamic 3D Gaussian splatting (3DGS) extends static 3DGS to render dynamic scenes, enabling AR/VR applications with moving objects. However, implementing dynamic 3DGS on edge devices faces challenges: (1) Loading all Gaussian parameters from DRAM for frustum culling incurs high energy costs. (2) Increased parameters for dynamic scenes elevate sorting latency and energy consumption. (3) Limited on-chip buffer capacity with higher parameters reduces buffer reuse, causing frequent DRAM access. (4) Dynamic 3DGS operations are not readily compatible with digital compute-in-memory (DCIM). These challenges hinder real-time performance and power efficiency on edge devices, leading to reduced battery life or requiring bulky batteries. To tackle these challenges, we propose algorithm-hardware co-design techniques. At the algorithmic level, we introduce three optimizations: (1) DRAM-access reduction frustum culling to lower DRAM access overhead, (2) Adaptive tile grouping to enhance on-chip buffer reuse, and (3) Adaptive interval initialization Bucket-Bitonic sort to reduce sorting latency. At the hardware level, we present a DCIM-friendly computation flow that is evaluated using the measured data from a 16 nm DCIM prototype chip. Our experimental results on Large-Scale Real-World Static/Dynamic Datasets demonstrate the ability to achieve high frame rate real-time rendering exceeding 200 frames per second (FPS) with minimal power consumption—merely 0.28 W for static Large-Scale Real-World scenes and 0.63 W for dynamic Large-Scale Real-World scenes. This work successfully addresses the significant challenges of implementing static/dynamic 3DGS technology on resource-constrained edge devices. Wei-Hsing Huang, Cheng-Jhih Shih, Jian-Wei Su, Samuel Wade Wang, Vaidehi Garg, Yuyao Kong, Jen-Chun Tien, Nealson Li, Arijit Raychowdhury, Meng-Fan Chang, Yingyan (Celine) Lin, Shimeng Yu |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2026 | Hardware Acceleration of Kolmogorov-Arnold Network (KAN) in Large-Scale SystemsabstractRecent developments have introduced Kolmogorov– Arnold networks (KANs), an innovative architectural paradigm capable of replicating conventional deep neural network (DNN) capabilities while utilizing significantly reduced parameter counts through the employment of parameterized B-spline functions incorporating trainable coefficients. Nevertheless, the B-spline functional components inherent to KAN architectures introduce distinct hardware acceleration complexities. While B-spline function evaluation can be accomplished through lookup table (LUT) implementations that directly encode functional mappings, thus minimizing computational overhead, such approaches continue to demand considerable circuit infrastructure, including LUTs, multiplexers, decoders, and associated components. This work presents an algorithm-hardware co-design approach for KAN acceleration. At the algorithmic level, techniques include alignment–symmetry and PowerGap KAN hardware-aware quantization, KAN sparsity-aware mapping strategy, and circuit-level techniques include N:1 time modulation dynamic voltage input generator with analog-compute-in-memory (ACIM) circuits. Furthermore, this work conducts comprehensive evaluations on large-scale KAN networks to validate the proposed methodologies. Nonideality factors, including partial sum deviations arising from process variations, have been evaluated with the statistics measured from the TSMC 22-nm RRAM-ACIM prototype chips. Utilizing optimally determined KAN hyperparameters in conjunction with circuit optimizations implemented and evaluated at the 22-nm technology node, despite the model sizes for large-scale tasks in this work increasing by 435 K$\times $to 756 K$\times $compared to tiny-scale tasks in previous work, the area overhead increases by only 26 K$\times $to 40 K$\times $, with power consumption rising by merely$48\times $to$93\times $, while accuracy degradation remains minimal at 0.11%–0.22%, thereby demonstrating the scaling potential of our proposed architecture. Wei-Hsing Huang, Jianwei Jia, Yuyao Kong, Faaiq G. Waqar, Tai-Hao Wen, Meng-Fan Chang, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge InferenceabstractRecently, a novel model named Kolmogorov-Arnold Networks (KAN) has been proposed with the potential to achieve the functionality of traditional deep neural networks (DNNs) using orders of magnitude fewer parameters by parameterized B-spline functions with trainable coefficients. However, the B-spline functions in KAN present new challenges for hardware acceleration. Evaluating the B-spline functions can be performed by using lookup tables (LUTs) to directly map the B-spline functions, thereby reducing computational resource requirements. However, this method still requires substantial circuit resources (LUTs, MUXs, decoders, etc.). For the first time, this paper employs an algorithm-hardware co-design methodology to accelerate KAN. The proposed algorithm-level techniques include Alignment-Symmetry and PowerGap KAN hardware aware quantization, KAN sparsity aware mapping strategy, and circuit-level techniques include N:1 Time Modulation Dynamic Voltage input generator with analog-CIM (ACIM) circuits. The impact of non-ideal effects, such as partial sum errors caused by the process variations, has been evaluated with the statistics measured from the TSMC 22nm RRAM-ACIM prototype chips. With the best searched hyperparameters of KAN and the optimized circuits implemented in 22 nm node, we can reduce hardware area by 41.78x, energy by 77.97x with 3.03% accuracy boost compared to the traditional DNN hardware. Wei-Hsing Huang, Jianwei Jia, Yuyao Kong, Faaiq G. Waqar, Tai-Hao Wen, Meng-Fan Chang, Shimeng Yu |
ASP-DAC | 1 |
| 2020 | A Two-way SRAM Array based Accelerator for Deep Neural Network On-chip TrainingabstractOn-chip training of large-scale deep neural networks (DNNs) is challenging due to computational complexity and resource limitation. Compute-in-memory (CIM) architecture exploits the analog computation inside the memory array to speed up the vectormatrix multiplication (VMM) and alleviate the memory bottleneck. However, existing CIM prototype chips, in particular, SRAM-based accelerators target at implementing low-precision inference engine only. In this work, we propose a two-way SRAM array design that could perform bi-directional in-memory VMM with minimum hardware overhead. A novel solution of signed number multiplication is also proposed to handle the negative input in backpropagation. We taped-out and validated proposed two-way SRAM array design in TSMC 28nm process. Based on the silicon measurement data on CIM macro, we explore the hardware performance for the entire architecture for DNN on-chip training. The experimental data shows that proposed accelerator can achieve energy efficiency of ~3.2 TOPS/W, >1000 FPS and >300 FPS for ResNet and DenseNet training on ImageNet, respectively. Hongwu Jiang, Shanshi Huang, Xiaochen Peng, Jian-Wei Su, Yen-Chi Chou, Wei-Hsing Huang, Ta-Wei Liu, Ruhui Liu, Meng-Fan Chang, Shimeng Yu |
DAC | 6 |
| 1998 | Test points selection process and diagnosability analysis of analog integrated circuitsabstractFault diagnosability analysis is an approach to enhancing the diagnosability of a circuit. Whereas the design for diagnosability ensures that the test points are properly selected and the generation of diagnostic tests is considerably simplified, diagnosability analysis is used to locate sections of a circuit having poor diagnosability. The information allows estimation of a circuit's diagnosability before the fault diagnosis is attempted. Hence any potential problem can be located early on the design phase, allowing modifications to be introduced to improve the final diagnosability. This paper presents a simple diagnosability analysis process in which the diagnosability of a circuit is measured from a graph that describes the circuit topology and a given set of test points, where no circuit simulation is needed. In addition, a simple algorithm is also presented to select a minimum set of test points and their locations for achieving the desired diagnosability. Wei-Hsing Huang, Chin-Long Wey |
ICCD | 1 |