EDBT 2026 Demo / reviewers in the wild / expert
Kevin Tshun Chuan Chai
dblp:148/6458 · also Kevin Chai Tshun Chuan, Kevin T. C. Chai, Kevin Tshun-Chuan Chai
· DBLP profile ↗
8ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0001-6624-8912ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Generative Pixelated Microstrip Millimeter-Wave Circuit Components DesignabstractIn this paper, we present a novel method for the generative design of pixelated microstrip Millimeter-Wave (mm-wave) circuit components. Our approach utilizes a primary meandering transmission line, which is connected to the input and output/terminated ports, with random small pixelated patches surrounding the primary transmission line as perturbations. A 2D Convolutional Neural Network (CNN) incorporating a Convolutional Block Attention Module (CBAM) is trained and employed as a forward solver to predict the input impedance of the circuit component. Additionally, a Variational Autoencoder (VAE) is trained to generate new pixelated microstrip mm-wave circuit component patterns by sampling from the latent space, thereby extending beyond the original training dataset. By leveraging the CNN-based forward solver for impedance prediction and the VAE for pattern generation, we can inversely generate the mm-wave circuit component design that meets specific input impedance requirements. Zaifeng Yang, Chuan Ge, Kevin Tshun Chuan Chai, Viet Phuong Bui, Ching Eng Png |
TENCON | 3 |
| 2024 | A Dual 7T SRAM-Based Zero-Skipping Compute- In-Memory Macro With 1-6b Binary Searching ADCs for Processing Quantized Neural NetworksabstractThis article presents a novel dual 7T static random-access memory (SRAM)-based compute-in-memory (CIM) macro for processing quantized neural networks. The proposed SRAM-based CIM macro decouples read/write operations and employs a zero-input/weight skipping scheme. A 65nm test chip with$528\times 128$integrated dual 7T bitcells demonstrated reconfigurable precision multiply and accumulate operations with$384\times $binary inputs (0/1) and$384\times 128$programmable multi-bit weights (3/7/15-levels). Each column comprises$384\times $bitcells for a dot product,$48\times $bitcells for offset calibration, and$96\times $bitcells for binary-searching analog-to-digital conversion. The analog-to-digital converter (ADC) converts a voltage difference between two read bitlines (i.e., an analog dot-product result) to a 1-6b digital output code using binary searching in 1-6 conversion cycles using replica bitcells. The test chip with 66Kb embedded dual SRAM bitcells was evaluated for processing neural networks, including the MNIST image classifications using a multi-layer perceptron (MLP) model with its layer configuration of 784-256-256-256-10. The measured classification accuracies are 97.62%, 97.65%, and 97.72% for the 3, 7, and 15 level weights, respectively. The accuracy degradations are only 0.58 to 0.74% off the baseline with software simulations. For the VGG6 model using the CIFAR-10 image dataset, the accuracies are 88.59%, 88.21%, and 89.07% for the 3, 7, and 15 level weights, with degradations of only 0.6 to 1.32% off the software baseline. The measured energy efficiencies are 258.5, 67.9, and 23.9 TOPS/W for the 3, 7, and 15 level weights, respectively, measured at 0.45/0.8V supplies. Chengshuo Yu, Haoge Jiang, Junjie Mu, Kevin Tshun Chuan Chai, Tony Tae-Hyoung Kim, Bongjin Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | 1V, 1.13μm pixel pitch Liquid Crystal Driver with Charge-Balancing Scheme for SLM ApplicationsabstractThis work proposes a compact 9T SRAM-based pixel design for low-voltage and high-speed modulator for spatially varying modulation of light (i.e. SLM). To reduce the supply voltage to the CMOS pixel backplane, the operating point is shifted towards the linear window by dynamically pulsing both the top & bottom electrodes in each pixel. A test chip in a standard 40nm CMOS technology was implemented, supporting upto 90 frames/second VGA display. Our testing results demonstrated that the proposed HCS switching scheme is efficient in achieving optical modulation up to 8-bit resolution, making it a viable candidate for state of the art SLMs applications with high frame rate and low power requirements. Aarthy Mani, Chong Yi Sheng, Rasna Maruthiyodan Veetil, Moitra Parikshit, Tobias Wilhelm W. Mass, Chong Ser Choong, Xuewu Xu, Ramon José Paniagua Domínguez, Arseniy I. Kuznetsov, P. Krishna, P. Keyi, Kevin Tshun Chuan Chai, Anh-Tuan Do |
ISCAS | 13 |
| 2022 | SonicFFT: A system architecture for ultrasonic-based FFT accelerationabstractFast Fourier Transform (FFT) is an essential algorithm for numerous scientific and engineering applications. It is key to implement FFT in a high-performance and energy-efficient manner. In this paper, we leverage the properties of ultrasonic wave propagation in silicon for FFT computation. We introduce SonicFFT: A system architecture for ultrasonic-based FFT acceleration. To evaluate the benefits of SonicFFT, a compact-model based simulation framework that quantifies the performance and energy of an integrated system comprising of digital computing components interfaced with an ultrasonic FFT accelerator has been developed. We also present mapping strategies to compute 2D FFT utilizing the accelerator. Simulation results show that SonicFFT achieves a$2317\times$system-level energy-delay product benefits-a simultaneous$117.69\times$speedup and$19.69\times$energy reduction-versus state-of-the-art baseline all-digital configuration. Darayus Adil Patel, Viet Phuong Bui, Kevin Tshun Chuan Chai, Amit Lal, Mohamed M. Sabry |
ASP-DAC | 3 |
| 2022 | A 1800μm2, 953Gbps/W AES Accelerator for IoT Applications in 40nm CMOSabstractA compact and energy-efficient AES accelerator for area and power-constrained IoT applications was fabricated in a 40nm CMOS process. By eliminating the need of intermediate data registers for MixColumns and ShiftRows in our proposed AES accelerator, we were able to reduce the total flip-flops to only 269 bits. Further, by reusing functional blocks and swapping the D flip-flops in data storage with scan flip-flops, our chip occupies only a tiny area of $1800 \mu \text{m}^{2}$ with an extremely low number of 657 gates. In addition, clock gating method and near-threshold voltage were used in our design. Thus, our accelerator consumes only $3.2 \mu \text{W}$ with an operation efficiency of 953 Gbps/W using a 0.48 V supply voltage. Compared with prior arts, our design has savings of 53% on area and 55% on the number of gates. When operated with a supply voltage of 0.48 V at 25°C, we can also achieve lower energy efficiency. Jingjing Lan, Vishnu P. Nambiar, Ming Ming Wong, Fei Li 0015, Yuan Gao 0011, Kevin Tshun Chuan Chai, Anh-Tuan Do |
ISCAS | 6 |
| 2022 | Bayesian Deep Active Learning for Analog Circuit Performance ClassificationabstractComputationally intensive simulations have made analog circuit sizing challenging for complicated analog circuit performance characterization. Accurate yet computationally efficient data-driven models of circuit performance can potentially accelerate the design and verification process. However, as analog circuits are designed under strict functional and technology constraints, there is a scarcity of data for analog circuit performance classification, posing challenges to data-driven approaches; acquiring more data typically involves running expensive and time consuming simulations. We propose Bayesian Deep Active Learning (BDAL) to learn models using fewer simulations, by iteratively selecting a small number of informative samples to label based on the model uncertainty. Bayesian neural networks used in the BDAL framework are better able to model weight uncertainty while being sufficiently expressive to model complex circuits. Compared with the state-of-the-art approaches, the proposed BDAL method can obtain better classification performance with much fewer number of simulations. Experiments on four diverse analog circuits demonstrate BDAL can achieve significant reduction in data requirement and obtain similar performance with much less labeled data for analog circuit performance classification. Lining Zhang, Salahuddin Raju, Ashish James, Rahul Dutta, Gregoire Fournier, Damien Lancry, Kevin Tshun Chuan Chai, Vijay Chandrasekhar 0001, Chuan-Sheng Foo |
ISCAS | 7 |
| 2021 | A Logic-Compatible eDRAM Compute-In-Memory With Embedded ADCs for Processing Neural NetworksabstractA novel 4T2C ternary embedded DRAM (eDRAM) cell is proposed for computing a vector-matrix multiplication in the memory array. The proposed eDRAM-based compute-in-memory (CIM) architecture addresses a well-known Von Neumann bottle-neck in the traditional computer architecture and improves both latency and energy in processing neural networks. The proposed ternary eDRAM cell takes a smaller area than prior SRAM-based bitcells using 6-12 transistors. Nevertheless, the compact eDRAM cell stores a ternary state (-1, 0, or +1), while the SRAM bitcells can only store a binary state. We also present a method to mitigate the compute accuracy degradation issue due to device mismatches and variations. Besides, we extend the eDRAM cell retention time to 200μs by adding a custom metal capacitor at the storage node. With the improved retention time, the overall energy consumption of eDRAM macro, including a regular refresh operation, is lower than most of prior SRAM-based CIM macros. A 128×128 ternary eDRAM macro computes a vector-matrix multiplication between a vector with 64 binary inputs and a matrix with 64 × 128 ternary weights. Hence, 128 outputs are generated in parallel. Note that both weight and input bit-precisions are programmable for supporting a wide range of edge computing applications with different performance requirements. The bit-precisions are readily tunable by assigning a variable number of eDRAM cells per weight or adding multiple pulses to input. An embedded column ADC based on replica cells sweeps the reference level for 2N-1 cycles and converts the analog accumulated bitline voltage to a 1-5bit digital output. A critical bitline accumulate operation is simulated (Monte-Carlo, 3K runs). It shows the standard deviation of 2.84% that could degrade the classification accuracy of the MNIST dataset by 0.6% and the CIFAR-10 dataset by 1.3% versus a baseline with no variation. The simulated energy is 1.81fJ/operation, and the energy efficiency is 552.5-17.8TOPS/W (for 1-5bit ADC) at 200MHz using 65nm technology. Chengshuo Yu, Taegeun Yoo, Tony Tae-Hyoung Kim, Kevin Tshun Chuan Chai, Bongjin Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2017 | A passively compensated capacitive sensor readout with biased varactor temperature compensation and temperature coherent quantizationabstractThis paper presents a frequency-mode capacitive sensor readout front-end operating in a wide temperature range. First, a biased varactor temperature compensation (BVTC) is proposed to compensate the aggregate temperature gradients from the sensor and the oscillator circuit, achieving a nullified temperature coefficient for the oscillation frequency. Second, a temperature coherent quantization (TCQ) approach is proposed to enhance the sensor' sensitivity and to provide a self-referenced clock for digitization, whereby the influence from the temperature effect of the clock is minimized through a hybrid down-conversion time-to-digital converter (DTDC). A prototype chip was fabricated using the 0.18-μm CMOS process and it was verified using a commercial 5.92-to-6.53-pF capacitive pressure sensor over a temperature range of −20°C to 120°C. Proven by experiments, the prototype presented as good as ±0.4% full-scale pressure error in the 140°C temperature range. Wang Ling Goh, Kevin Tshun Chuan Chai, Xin Lou 0001, Wen Bin Ye 0001 |
ISCAS | 4 |