VLDB 2026 Research / reviewers in the wild / expert
Bah-Hwee Gwee
dblp:92/3436
· DBLP profile ↗
75ranked-venue papers
2as first author
29since 2021 · last 2026
0000-0002-3222-2885ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 66 · 1 first-author · 24 since 2021Security and privacy · 4 · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniOOD: A Unified Framework for Domain Generalization and Out-of-Distribution Detection in Time Series
Yongming Chen, Wenwen Zheng, Bah-Hwee Gwee, Qi Cao 0002, Sirajudeen Gulam Razul, Zhiping Lin 0001 |
ISCAS | 3 |
| 2026 | ICNet: Cross-Modality Image Analysis for IC Localization in Printed Circuit Boards
Jingyang Dai, Deruo Cheng, Xinrui Wang 0004, Yiqiong Shi, Bah-Hwee Gwee |
ISCAS | 6 |
| 2026 | RASLL: A Removal Attack on SAT-Resistant Logic Locking
Zijian Long, Juncheng Chen, Tong Lin 0001, Nay Aung Kyaw, Bah-Hwee Gwee |
ISCAS | 6 |
| 2026 | Towards invariant and interpretable representations for domain generalization in time series classification
Yongming Chen, Zhenyu Weng, Bah-Hwee Gwee, Qi Cao 0002, Sirajudeen Gulam Razul, Zhiping Lin 0001 |
Pattern Recognit. | 3 |
| 2026 | DQSA: Dynamic Quantized Self-Attention for Multi-Task Encrypted Network Traffic ClassificationabstractNetwork traffic classification is crucial for both network security and management. Despite advances in deep learning-based multi-task traffic classification, existing models often struggle to jointly handle multiple tasks while providing interpretable insights. In multi-task scenarios, different tasks rely on distinct regions of the traffic sequence, motivating the use of dynamic and interpretable attention mechanisms. To this end, we propose Dynamic Quantized Self-Attention (DQSA), a unified framework specifically designed for multi-task network traffic classification. At its core, the Task Gated Attention Router (TGAR) dynamically associates attention heads with different tasks, enabling adaptive focus on task-specific patterns. This mechanism provides interpretable attention scores, which help analyze misclassifications and guide further model refinement. To improve efficiency and handle diverse network traffic features, we introduce the Soft Quantized Self-Attention Head (SQ-SAH) to reduce computational complexity and extend the Rotary Position Embedding (RoPE) to accommodate these features. Extensive experiments on ISCX VPN-NonVPN and DCI-LTE datasets demonstrate that DQSA consistently outperforms state-of-the-art baselines, achieving 92.85% accuracy on the encapsulation-level task of ISCX VPN-NonVPN and 93.17% accuracy on the application-level task of DCI-LTE, surpassing the strongest existing methods by up to 2.65%, while providing interpretable task-specific attention for efficient multi-task network traffic classification. Yongming Chen, Hongsheng Lan, Bah-Hwee Gwee, Qi Cao 0002, Sirajudeen Gulam Razul, Zhiping Lin 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | A Novel Energy-Efficient Continuous-Time Hysteretic VCO-Based ComparatorabstractVoltage-controlled oscillator (VCO)-based comparators offer higher energy efficiency as the difference in input magnitudes increase, such as in level-crossing ADCs. Nevertheless, to date, they require a clock signal to perform comparison operations. This is incongruous with continuous-time applications, where inputs are compared continuously. Further, they lack hysteresis, a crucial feature for mitigating spurious switching that compromises energy efficiency. In this paper, we present a novel VCO-based comparator that, for the first time, simultaneously achieves continuous-time operation and high energy efficiency. The former feature is enabled by a novel continuous-time decision circuit, while the latter is achieved through a novel switched-current hysteresis circuit that mitigates spurious switching. The proposed comparator is designed in 65 nm CMOS. Simulation results show that it achieves low energy per comparison, ranging from 0.07 to 4 pJ, with an average propagation delay of ~15 ns. The average energy consumption is 0.19 pJ — ~1.8× lower than the state-of-the-art VCO-based comparator. Jinhen Lee, Victor Adrian, Kinglouis Steven Tantra, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2025 | Multiple Hypothesis Testing for SEM Image Processing: A Case Study on Standard Cell PartitionabstractThe detection and partitioning of standard cells from Scanning Electron Microscope (SEM) images is a crucial step in hardware assurance of Integrated Circuit (IC). Traditional methods may struggle with the noise and complexity of these signals. This paper introduces a novel approach to SEM image processing by framing the standard cell partition problem as a multiple hypothesis testing (MHT) problem. This method enables simultaneous decision-making across many hypotheses, enhancing detection accuracy while controlling the false discovery rate (FDR). We show how MHT can identify partition lines in noisy brightness signals extracted from SEM images. Using the Benjamini-Hochberg (BH) procedure, we achieve effective FDR control, improving detection robustness and providing a clearer understanding of cell structures. This study demonstrates the suitability of MHT for SEM image processing and its potential for other circuit-related challenges. Yizhen Li, Tong Lin 0001, Yiqiong Shi, Deruo Cheng, Bah-Hwee Gwee |
ISCAS | 7 |
| 2025 | SSRNet: Few-shot IC Segmentation in Automated PCB Image ProcessingabstractAutomated inspection of Integrated Circuits (ICs) on Printed Circuit Boards (PCBs) is essential for ensuring the reliability of modern electronic systems. However, the inspection process faces significant challenges, particularly data scarcity and low inter-class variance. To address these challenges, we propose SSRNet, a few-shot learning-based framework for precise IC segmentation in complex PCB optical images. Unlike traditional deep learning models, our proposed SSRNet utilizes a similarity-guided approach for initial mask prediction and integrates a region classifier for further refinement. This design allows SSRNet to accurately segment IC components, even with limited annotated data. Experimental results demonstrate that our proposed SSRNet outperforms the state-of-the-art model, achieving a 23.0% increase in IoU and a 13.2% improvement in the Dice coefficient on NTU PCB DSX Dataset (NPDD). Xinrui Wang 0004, Deruo Cheng, Tong Lin 0001, Yiqiong Shi, Bah-Hwee Gwee |
ISCAS | 7 |
| 2025 | N-MUX: Neighborhood-Based Logic Locking Against Machine Learning AttacksabstractMUX-based logic locking (LL) is a hardware security technique that inserts multiplexers (MUX) into circuits to secure them against unauthorized use and reverse engineering by protecting original circuit pathways. Nevertheless, MUX-based LL is vulnerable to Oracle-Guided (OG) and Oracle-Less (OL) attacks. While OG methods, such as the SAT attack, are infeasible for large-scale designs, OL attacks, like those based on machine learning (ML), can exploit structural leakage in locked circuits to recover original pathways. This study introduces N-MUX, an innovative MUX-based LL approach designed to resist state-of-the-art (SOTA) ML attacks. N-MUX effectively reduces structural leakage by identifying maximal overlap structures in the original circuit to configure the MUX logic. Additionally, N-MUX ensures high efficiency by selecting the false input from the direct neighbourhood of the true input. Experimental results on ISCAS’85 and ITC’99 benchmarks demonstrate that N-MUX is the most secure and reliable LL technique against SOTA ML-based attacks, achieving an 81% reduction in attack accuracy compared to existing MUX-based LL methods and delivering up to 480× greater efficiency. Xuenong Hong, Shirui Sheng, Juncheng Chen, Nay Aung Kyaw, Kwen-Siong Chong, Zhiping Lin 0001, Bah-Hwee Gwee |
ISCAS | 8 |
| 2025 | Live Demonstration: An Area and Energy Efficient Reconfigurable Cryptographic Accelerator Based SoC Design for Securing IoT DevicesabstractThis demonstration presents an energy and area efficient Reconfigurable Cryptographic Accelerator (RCA) SoC for secure communication in IoT devices. Built on a ZYNQ-7000 development board, the platform supports multiple block ciphers (DES, AES, SM4) and Hash functions (SHA-1, SHA-2, SM3). Users can follow prompts on the OLED screen to select the cryptographic algorithm via buttons and input data through a keyboard or choose large text files from SD card. The ARM Core and accelerator execute the cryptographic operation simultaneously, and energy efficiency is calculated based on power and computing time, showcasing the improved computing speed and energy efficiency of the proposed accelerator. Xvpeng Zhang, Bingqiang Liu, Lingyun Hu, Zixuan Shen, Zaisheng He, Dengke Xu, Bah-Hwee Gwee, Chao Wang 0096 |
ISCAS | 7 |
| 2025 | Long-Short-GNN: A Novel Graph Neural Network for Detecting FPGA IP Circuits for Hardware AssuranceabstractHardware Assurance (HA) of Integrated Circuit (IC) requires the extraction and analysis of circuit netlist from a manufactured or programmed IC (in the case of Field Programmable Gate Array (FPGA)). The first and most important step in this analysis is to detect Intellectual Property (IP) circuit(s) of interest from an extracted ‘sea-of-gates’ netlist. State-of-the-art approach involves converting the extracted netlist into a graph and using Graph Neural Network (GNN), a powerful machine-learning method on graphs for IP circuit detection. However, reported methods usually employed shallow GNNs with small receptive fields which are inadequate for detecting large and complete IP circuits. In this paper, we propose a novel GNN, coined Long-Short-GNN, which uniquely incorporates a Long-view Network (for global coarse-grained information) and a Short-view Network (for local fine-grained information) for FPGA IP circuit detection. By experiments on detecting a variety of large and complete FPGA IP circuits, we proved its efficacy. Specifically, on average, our proposed Long-Short-GNN outperformed all reported methods by a large margin of up to ~13.8% improvement on F1 Score. Heyi Zhang, Tong Lin 0001, Deruo Cheng, Yiqiong Shi, Bah-Hwee Gwee |
ISCAS | 6 |
| 2024 | A Novel Non-profiling Side-Channel Attack on Masked Devices with Connectivity MatrixabstractIn this paper, we propose a novel pre-processing technique known as the Connectivity Matrix (CM). Building upon the foundation of the CM, we present an effective Second-Order Side-Channel Attack, called Connectivity Matrix Attack (CMA). Our work aims to efficiently counter hardware devices fortified with Masking countermeasures, and it contributes in three significant ways. First, the proposed CM has lower data complexity, as it is constant regarding the number of measurements. Second, we propose the decomposition of the CMs and the utilization of their eigenvalues as feature vectors in CMA. This approach effectively removes noisy components from the CMs and reduces their dimensions. Third, the proposed CMA employs the selected eigenvalues to establish a frequency distribution, followed by a chi-square test. This approach allows CMA to expose both the linear and non-linear leakages present in CMs. The proposed CMA is validated on the public dataset ASCAD and can reveal all the masked bytes successfully. Notably, the concept of the Connectivity Matrix extends beyond the confines of a correlation matrix used in this paper, opening the door to a promising avenue for future research. Juncheng Chen, Zishuo Yang, Nay Aung Kyaw, Kwen-Siong Chong, Zhiping Lin 0001, Bah-Hwee Gwee |
ISCAS | 8 |
| 2024 | MLConnect: A Machine Learning Based Connection Prediction Framework for Error Correction in Recovered CircuitabstractIntegrated Circuit (IC) verification is of paramount importance to the security of IC. The success of circuit verification largely depends on the correctness of the recovered circuit netlist from Scanning Electron Microscopic (SEM) images. Due to imperfections in imaging process and feature extraction process, the recovered circuit netlist usually contains connection errors. The corrections of these errors require tedious manual tracing of metal lines or are sometimes impossible due to the corrupt regions in SEM images. In this work, we perform error correction based on a connection heuristic in circuit. We propose MLConnect, a machine learning based connection prediction framework that captures the probabilities of gate connections in circuits. We further propose a post-processing technique to recover circuit connections based on gate connection probabilities and circuit rules. Our results show that the proposed MLConnect successfully recovered 80.87% of gate connections in erroneous circuits from ISCAS-85 benchmark suites. Our method can largely automate the process of circuit recovery. Xuenong Hong, Zilong Hu, Yee-Yang Tee, Tong Lin 0001, Yiqiong Shi, Deruo Cheng, Bah-Hwee Gwee |
ISCAS | 8 |
| 2024 | A Method for Out-of-Distribution Detection in Encrypted Mobile Traffic ClassificationabstractThe widespread use of encrypted communication in mobile networks poses significant challenges in accurately classifying traffic. Detecting out-of-distribution (OOD) samples, which significantly deviate from known classes, adds complexity to the task. This paper proposes a feature analysis-based OOD detection scheme for traffic classification in Long-Term Evolution (LTE) systems. Our method utilizes Long Short-Term Memory (LSTM) networks for feature extraction, capturing the feature vectors of the traffic series. Principal Component Analysis (PCA) is then applied to obtain principal and residual principal components. Leveraging the residual feature vector, we construct an OOD score to quantify deviation from the ID dataset. Extensive experiments on a large-scale encrypted mobile traffic dataset demonstrate the superiority of our approach, achieving high accuracy in OOD detection compared to existing techniques. Our method contributes to enhanced security and reliable traffic classification in LTE systems, addressing challenges posed by OOD samples. Yuzhou Tong, Yongming Chen, Bah-Hwee Gwee, Qi Cao 0002, Sirajudeen Gulam Razul, Zhiping Lin 0001 |
ISCAS | 3 |
| 2024 | SAMIC: Segment Anything Model for Integrated Circuit Image AnalysisabstractCircuit annotation is crucial in analyzing integrated circuit (IC) images for hardware assurance. While deep learning algorithms perform well in circuit annotation, they are highly reliant on labeled training data, which are extremely costly to obtain. The Segment Anything Model (SAM) excels in segmenting natural images but performs sub-optimally on IC images due to the domain gap between IC images and natural images. In this paper, we introduce SAMI C which extends the application of SAM for effective annotation of IC images. We curated an extensive dataset of IC images from four different devices and developed a novel training methodology for SAM models in IC image segmentation. Our experiments show that SAMIC outperforms the original SAM model by 36.78 % and improves accuracy by 6.35 % compared to the second-best technique. Yong-Jian Ng, Yee-Yang Tee, Deruo Cheng, Bah-Hwee Gwee |
TENCON | 4 |
| 2024 | Securing Against Side-Channel Attacks With Wide-Range In Situ Random Voltage Dithering on Async-Logic AES EngineabstractWe present a wide-range in situ random voltage dithering (WIS-RVD) on async-logic advanced encryption standard (AES) engine to counteract side-channel attacks (SCAs). There are three contributions in this brief. First, we propose the WIS-RVD based on a dual-rail asynchronous-logic (async-logic) AES engine, leveraging on the self-timed clockless operations for robust encryption under dynamic voltage and timing variations. Second, we propose an in situ voltage dithering to dither the supply voltage instantaneously during the encryption, without the requirement of additional control circuits for clock modulation, to increase the SCA resistance. Third, we propose a wide-range voltage swing technique that spans from 0.3 V (subthreshold) to 1.1 V (above threshold), obfuscating the transistor’s current models between subthreshold and threshold voltage to further enhance SCA resistance. We perform comprehensive SCA evaluations with 50-M power and EM measurements, and the SCA evaluations show that our proposed WIS-RVD on async-logic AES accelerator can resist SCAs with 50-M measurements, i.e.,$\gt 2083\times $and$\gt 2778\times $improvement for power and EM SCAs, respectively, when compared to the standard synchronous-logic AES. Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Kwen-Siong Chong, Bah-Hwee Gwee |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | GRACER: Graph-Based Standard Cell Recognition in IC Images for Hardware AssuranceabstractGlobal distribution of the Integrated Circuit (IC) supply chain amplifies the importance of Hardware Assurance (HA), i.e., to ensure the integrity of manufactured IC. Standard cell recognition is a crucial step in HA, which is to identify the functionality of a standard cell based on its Scanning Electron Microscope (SEM) images. Conventionally, this is mostly done by human inspection, which is labor-intensive and error-prone. Current works on automating this process only work on the image domain and have sub-optimal performance due to the challenges incurred by the variation in the appearance of standard cells in the images. In this paper, we propose an automatic process for standard cell recognition, through conversion to a standardized graph representation and comparing the graph structure to identify the type of the standard cell. Our proposed method represents each unique circuit structure in a unique graph representation and thus enables a one-to-one matching to a known set of templates for functionality identification. Our experiments show that our proposed method can always recognize the standard cells correctly, even under the most challenaing scenario. Erdong Huang, Xuenong Hong, Tong Lin 0001, Yiqiong Shi, Bah-Hwee Gwee |
IECON | 5 |
| 2023 | SEM2GDS: A Deep-Learning Based Framework To Detect Malicious Modifications In IC LayoutabstractOverseas foundries pose potential threat to the integrity of manufactured ICs where malicious modifications, known as Hardware Trojans (HTs) may be inserted into the IC layout. To detect this, SEM images of manufactured ICs need to be compared with their original GDS images. However, existing methods either avoid direct comparison or are susceptible to errors due to the inherent differences in shapes between SEM images and GDS images. In this paper, we instead propose a Deep-Learning (DL)-based image transformation method, named SEM2GDS, which transforms a SEM image into its GDS image and produce shapes with sharp corners. This allows direct comparison between a transformed SEM image and the original GDS image for modification detection. By experiment on a set of SEM images and their corresponding GDS images, we demonstrate the efficacy of our proposed method. Our method is fast and able to achieve high detection accuracy, high f1 score, and very low False Negative Rate (FNR) of <0.02. Our method can detect real and small changes between SEM and GDS images. Tong Lin 0001, Yiqiong Shi, Bah-Hwee Gwee |
ISCAS | 3 |
| 2023 | Real-time Traffic Classification in Encrypted Wireless Communication NetworkabstractClassification of traffic service types is a valuable function for wireless communication networks. Even though some progress has been made, the recognition of the type of the traffic services cannot be done in real time. In this paper, we propose a novel method for classifying traffic series in real time based on transfer learning techniques. We pre-train a deep learning model with long traffic series and fine-tune the model with short traffic series. In this way, the developed model achieves the capability of recognising traffic services in real time. In other words, the model can recognize traffic services by using short traffic series. We collect Downlink Control Information (DCI) from commercial LTE networks when using five common types of traffic services. Then we use the dataset to validate our method. Our experimental results show that, by using proposed method, LSTM accuracy rates will increase to 80% and 88.5% when the length of the traffic series is 5 seconds and 10 seconds respectively, which is higher than the baseline. The strategy is also suitable for one dimension convolution neural network (1D-CNN). Yongming Chen, Yuzhou Tong, Bah-Hwee Gwee, Qi Cao 0002, Sirajudeen Gulam Razul, Zhiping Lin 0001 |
ISCAS | 3 |
| 2023 | A 3D-Printed Fourth-Order Stacked Filter for Integrated DC-DC ConvertersabstractThe passive devices in state-of-the-art miniaturized switched-mode DC-DC converters are generally integrated by means of on-chip and in-package methods. Nevertheless, the quality is poor-to-moderate, thereby compromising the power-efficiency. In this paper, we propose the miniaturization of the DC-DC converter by means of realizing its passive devices as embedded devices that are printed within a high-density 3D inkjet printed-circuit-board (PCB). We propose a fourth-order stacked LC filter embodying passive components with small values-effectively at no additional cost because they are embedded through 3D-printing. For the inductor and capacitor, we propose to adopt a high-$Q$solenoidal structure and the metal-insulator-metal planar structure, respectively. The proposed filter is printed within the 3D-PCB with a compact 124 mm3volume due to the stacked arrangement. The measured AC attenuation is 21.2 dB at 200 MHz. The filter is further verified by means of computer simulations of a DC-DC buck converter. Simulation results of the converter employing the filter show a low output voltage ripple at 146 mV and a high peak power-efficiency of ~78% at 200 MHz switching frequency with 150 mA load current. Jinhen Lee, Victor Adrian, Sun-Yang Tay, Yanshan Xie, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 5 |
| 2023 | Improving FPGA-based Async-logic AES Accelerator with the Integration of Sync-logic Block RAMsabstractWe present a side-channel attack (SCA) resistant asynchronous-logic (async-logic) AES accelerator that integrates synchronous-logic (sync-logic) Block RAMs (BRAMs) in FPGA as the Substitution-Box. We successfully identify the timing requirements to integrate sync-logic BRAMs in our async-logic AES accelerator and validate our proposed AES accelerator on the Sakura-X FPGA board. With the integration of BRAMs, we improve the resource utilization on FPGA by$1.6\times$when compared to the state-of-the-art async-logic AES accelerator, while reducing the power overhead by$1.4\times$. We comprehensively evaluate the SCA resistance of our proposed async-logic AES accelerator with 11 attacking models in both time and frequency domains. Based on our evaluations, we show that our proposed async-logic AES accelerator is highly secure against SCA with 30 million EM traces. This is more than$6000\times$improvement when compared to the benchmark sync-logic AES accelerator and$1.5\times$improvement when compared to the state-of-the-art async-logic AES accelerator. Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Kwen-Siong Chong, Zhiping Lin 0001, Bah-Hwee Gwee |
ISCAS | 7 |
| 2023 | A Residual-Remainder Coupled Unlimited Sampling Framework for High Dynamic Range Signal ConversionabstractClipping distortion is a common problem when the amplitude of input signals exceeds the desired region of an analog-to-digital converter. Unlimited Sensing Framework (USF) alleviates the clipping distortion by folding out-of-range signals into a within-range via modulo operations. The USF signal recovery assumes an infinitesimal residual step time in the modulo operation which is generally practically infeasible. Its recovery error is inevitable due to the remainder sampling error during the residual step transition time. Instead of the infinitesimal assumption, a residual-remainder coupled USF is proposed to eliminate the remainder sampling error by coupling the residual sampling component. It is shown that the proposed coupling framework does not only relax the oversampling rate in the original USF to approaching the Nyquist sampling rate, but also provides a more accurate signal recovery capability as the remainder sampling error is eliminated by the coupling of the residual components. Lei Sun 0006, Hangcheng Han, Juncheng Chen, Bah-Hwee Gwee, Zhiping Lin 0001 |
ISCAS | 5 |
| 2022 | Non-profiling based Correlation Optimization Deep Learning AnalysisabstractDifferential Deep Learning Analysis (DDLA) is a deep learning-based non-profiling side-channel attack leveraging neural networks to classify Physical Leakage Information with labels. To avoid the Class Imbalance Problem (CIP) of significantly different data sizes in different data groups, DDLA employs bit labels. However, applying bit labels will be less effective for exploiting leakage. In this paper, we propose to employ Correlation optimization Deep Learning Analysis (CO-DLA) to circumvent the CIP in DDLA by converting the classification in DDLA into a correlation optimization. Bus labels can then be used to exploit stronger leakage information. To validate the attack efficacy improvement, we perform experiments on ASCAD synchronized and de-synchronized masked AES-128 datasets. For the synchronized masked dataset, our proposed CO-DLA requires only 5k traces, which is 75% lesser than the 20k traces required by the reported DDLA, to reveal the key-byte. For the 2 de-synchronized masked datasets, our proposed CO-DLA requires only 10k traces to reveal the key-byte from both of them while the reported DDLA fails to reveal the key-byte. Juncheng Chen, Jun-Sheng Ng, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Zhiping Lin 0001, Joseph Sylvester Chang, Bah-Hwee Gwee |
ISCAS | 8 |
| 2022 | An Asynchronous-Logic Masked Advanced Encryption Standard (AES) Accelerator and its Side-Channel Attack EvaluationsabstractWe present a side-channel-attack (SCA) resistant asynchronous-logic (async-logic) Advanced Encryption Standard (AES) accelerator embodying both the masking and hiding SCA countermeasures. Our async-logic masked AES accelerator adopts a dual-rail data encoding to perform the masked 128-bit AES operations, and to enable dual-hiding to moderate both the amplitude (vertical dimension) and the time (horizontal dimension) of the side-channel signals. We implement our async-logic masked AES accelerator in FPGA and comprehensively perform the SCA evaluations based on the electromagnetic (EM) emanation. The SCA evaluations are performed based on bus-wise Hamming Distance model, bus-wise & bit-wise Hamming Weight models, and Zero-Value (ZV) model. Based on our experiment results, we show that our async-logic masked AES is secured against SCA with 1 million EM emanations. This is at least $8.3 \times$ more resistant than synchronous-logic masked AES and $200 \times$ more resistant than the synchronous-logic unmasked AES. Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Joseph Sylvester Chang, Bah-Hwee Gwee |
ISCAS | 7 |
| 2022 | A Versatile and Accurate Vector-Based Method for Modeling and Analyzing Planar Air-Core InductorsabstractPlanar air-core inductors come in a variety of geometrical shapes, including in the form of the conventional spiral geometry and novel complex geometries. In the design phase of a system, the inductance of the employed inductor would need to be ascertained. This is usually ascertained by tedious mathematical derivations on a segment-by-segment (inductor) basis or time-consuming computer modeling, and the complexity can become intractable for complex geometries. In this paper, we propose a versatile, yet accurate, vector-based method to ascertain the inductance of planar air-core inductors with virtually any geometry, including novel complex geometry inductors—rather easily. Our proposed method decomposes the inductor segments into vectors, and thereafter utilizes geometric models to compute the inductance in a systematic fashion. We benchmark our proposed method against the conventional electromagnetic field solver simulations to estimate the inductances of six planar inductors ranging from a conventional spiral air-core inductor to that embodying different and complex geometries. On the basis of these six inductor examples, we show that our method is highly accurate with a worst-case error of $\sim 5$% compared to that obtained using conventional electromagnetic field solver. Of particular interest, our modeling for novel complex geometry planar inductors is relatively simple. Sun-Yang Tay, Victor Adrian, Joseph Sylvester Chang, Jinhen Lee, Bah-Hwee Gwee |
ISCAS | 5 |
| 2022 | A Highly Secure FPGA-Based Dual-Hiding Asynchronous-Logic AES Accelerator Against Side-Channel AttacksabstractEncryption in field-programmable gate array (FPGA) often provides a good security solution to protect data privacy in Internet-of-Things systems, but this security solution can be compromised by side-channel attacks (SCAs). In this article, we present an FPGA-based dual-hiding asynchronous-logic (async-logic) advanced encryption standard (AES) accelerator, which is highly resistant against SCAs and yet low area/energy overheads. The proposed AES accelerator achieves vertical (amplitude) SCA hiding via an area-efficient dual-rail mapping approach and a zero-value (ZV) compensated substitution-box (S-Box), while enhancing the horizontal (temporal) SCA hiding of async-logic operations via a timing-boundary-free input arrival-time randomizer and a skewed-delay controller. A comprehensive SCA evaluation is performed with 11 SCA models, and we show that our proposed design can offer a strong SCA resistance with measurement-to-disclosure (MTD) of >20 million traces. To our best knowledge, our design is the most secure AES design evaluated with the largest number of traces in FPGA. To compare the design overheads for security, we quantify the figure of merit as normalized (Area$\times $Energy/MTD(All)$\times 10^{6}$). The figure of merit of our proposed design is$403\times $smaller than the benchmark dual-rail synchronous-logic design and$95\times $smaller than a reported async-logic design. Jun-Sheng Ng, Juncheng Chen, Kwen-Siong Chong, Joseph Sylvester Chang, Bah-Hwee Gwee |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2021 | Joint Anomaly Detection and Inpainting for Microscopy Images Via Deep Self-Supervised LearningabstractWhile microscopy enables material scientists to view and analyze microstructures, the imaging results often include defects and anomalies with varied shapes and locations. The presence of such anomalies significantly degrades the quality of microscopy images and the subsequent analytical tasks. Comparing to classic feature-based methods, recent advancements in deep learning provide a more efficient, accurate, and scalable approach to detect and remove anomalies in microscopy images. However, most of the deep inpainting and anomaly detection schemes require a certain level of supervision, i.e., either annotation of the anomalies, or a corpus of purely normal data, which are limited in practice for supervision-starving microscopy applications. In this work, we propose a self-supervised deep learning scheme for joint anomaly detection and inpainting of microscopy images. The proposed anomaly detection model can be trained over a mixture of normal and abnormal microscopy images without any labeling. Instead of a two-stage scheme, our multi-task model can simultaneously detect abnormal regions and remove the defects via jointly training. To benchmark such microscopy application under the real-world setup, we propose a novel dataset of real microscopic images of integrated circuits, dubbed MIIC. The proposed dataset contains tens of thousands of normal microscopic images, while we labeled hundreds of them containing various imaging and manufacturing anomalies and defects for testing. Experiments show that the proposed model outperforms various popular or state-of-the-art competing methods for both microscopy image anomaly detection and inpainting. Deruo Cheng, Xulei Yang, Tong Lin 0001, Yiqiong Shi, Kaiyi Yang, Bah-Hwee Gwee, Bihan Wen |
ICIP | 7 |
| 2021 | Normalized Differential Power Analysis - for Ghost Peaks MitigationabstractThe attack efficacy of Differential Power Analysis (DPA), a popular side channel evaluation technique for key extraction, is compromised by the false highest Difference Of Means (DOMs) value ('ghost peaks') in the DOMs matrix produced in a conventional DPA. The ghost peak is generated by the wrong key guess and always occurs in the conventional DPA when the number of side channel traces is not enough. In this paper, an improved version of the conventional DPA termed as Normalized DPA (NDPA) is proposed to circumvent the ghost peak. With the analysis on the generation of ghost peaks in the conventional DPA, we observed that by normalizing the DOMs matrix, the ghost peaks can be greatly suppressed. We model the proposed NDPA mathematically and show that it performs better than the conventional DPA. We further provide the experimental validations on a set of 200k power simulation traces on AES S- Box and 500 EM traces from ASCAD dataset. Based on the attack results of these datasets, our proposed NDPA requires (up to 68%) lesser number of traces to reveal a correct key when compared to the conventional DPA. Juncheng Chen, Jun-Sheng Ng, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Weng-Geng Ho, Kwen-Siong Chong, Zhiping Lin 0001, Joseph Sylvester Chang, Bah-Hwee Gwee |
ISCAS | 9 |
| 2021 | A Novel Normalized Variance-Based Differential Power Analysis Against Masking CountermeasuresabstractIn this paper, we propose two normalization techniques to reduce the ghost peaks occurring in Differential Power Analysis (DPA). Ghost peaks can be defined as the DPA output generated by the wrong key guesses, having higher amplitudes than the DPA output generated by the correct key guess. We further propose variance-based Differential Power Analysis (vDPA) to attack masked crypto devices. The proposed normalization techniques and vDPA constitute four contributions. First, based on the side-channel signal modeling with the linear coefficient representing the strength of the linear component in a side-channel signal, we formulate the condition function of linear coefficients for the appearance of ghost peaks in DPA. Second, we propose pre-normalization in DPA and mathematically analyze how it can reduce ghost peaks by modulating the strength of the linear components in side-channel signals. Third, we propose post-normalization and mathematically analyze how it can reduce ghost peaks by de-correlating the strength of the linear components in side-channel signals with the condition function for the appearance of ghost peaks. Fourth, we propose vDPA to apply simultaneously with either one of the proposed normalization techniques to effectively attack masked crypto devices. Based on the experiments, we show that the proposed basic vDPA (without normalization), pre-normalized vDPA and post-normalized vDPA are all able to reveal the secret key from ASCAD data set. The pre- and post-normalized vDPAs require up to 18× and 14× fewer traces than the basic vDPA respectively. While attacking ASCAD data set, the proposed pre- and post-normalized vDPAs are both 13, 095× faster than the reported 2nd order CPA, and reveal the key-bytes successfully with only half of side-channel traces required by the reported Zero-offset DPA. Juncheng Chen, Jun-Sheng Ng, Kwen-Siong Chong, Zhiping Lin 0001, Bah-Hwee Gwee |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | A Secure Data-Toggling SRAM for Confidential Data ProtectionabstractWe study the security feature of static random access memory (SRAM) against the data imprinting attack and provide a solution to protect the SRAM from this attack. There are four main contributions in this paper. First, the negative-bias temperature-instability (NBTI) degradation of PMOS transistors in the conventional SRAM cell that causes the data imprinting effect is explained. Second, the data imprinting effect that leaks the stored information in the conventional SRAM cell is investigated. Third, a novel low transistor-count transmission-gate-based master-slave SRAM cell is proposed to periodically toggle the stored data for reducing the data imprinting effect. Fourth, an efficient imprinting analysis flow is proposed to evaluate the proposed data-toggling SRAM for quantifying the data imprinting effect. Based on a 65-nm CMOS process, we implement and prototype the proposed 1k-byte data-toggling SRAM design. We perform our imprinting analysis flow on various SRAM ICs and benchmark our proposed data-toggling SRAM IC against the non-toggling SRAM IC and a commercial Lyontek SRAM IC. From the measurement results, the non-toggling SRAM and Lyontek SRAM suffer from 60% and 81% data imprinting effects, respectively, whereas our data-toggling SRAM has only 11% data imprinting effect (at 160-kHz toggling frequency). The data-toggling SRAM could switch between high security (<; 5% data imprinting effect) high power mode for hardware security applications and low power (<; 0.1mW) low security mode for power-saving applications. Particularly, our data-toggling SRAM could feature as low as ~1% data imprinting effect when increasing the toggling frequency to 1.6 MHz by compromising the power dissipation. Using the image analysis flow, the stored information is revealed in both the non-toggling and Lyontek SRAM ICs but is well protected in the proposed data-toggling SRAM IC. Weng-Geng Ho, Kwen-Siong Chong, Tony Tae-Hyoung Kim, Bah-Hwee Gwee |
ISCAS | 4 |
| 2020 | A DPA-Resistant Asynchronous-Logic NoC Router with Dual-Supply-Voltage-Scaling for Multicore Cryptographic ApplicationsabstractWe propose a 5-port asynchronous-logic Network-on-Chip (ANoC) router based on the Sense-Amplifier Half-Buffer (SAHB) approach for cryptographic processing cores to counteract side channel attack differential power analysis (DPA) in multicore platform. There are three features in the proposed DPA-resistant ANoC router. First, the proposed ANoC router embodies dual-supply-voltage SAHB cells, where the non-critical subsidiary supply voltage is adjustable from 0.3V to 1.2V, increasing the noise variance and hence reducing the Signal-to-Noise (SNR) ratio to hide the information leakage. Second, the proposed ANoC router performs as a noise engine by increasing the number of power-on IO ports, further randomizing the overall power dissipation. Third, the proposed ANoC router can switch between DPA-resistant mode and energy-efficient nominal (non-secure) mode, saving the power dissipation when the DPA secure countermeasure is unnecessary. Based on 65nm CMOS process, the multicore platform embedded with the proposed ANoC router is implemented, and the experiment is demonstrated by running the advanced encryption standard (AES) cryptography operation. When benchmarked against the nominal mode, the noise power variance of the proposed ANoC router increases by 2.3× in the DPA-resistant mode, reducing the overall SNR ratio by 56%. When comparing to other reported noise engines, our proposed ANoC router is one of the most DPA-secure, area-efficient and power-efficient designs for multicore cryptographic applications. Weng-Geng Ho, Ne Kyaw Zwa Lwin, Nay Aung Kyaw, Jun-Sheng Ng, Juncheng Chen, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 7 |
| 2020 | A Highly Efficient Power Model for Correlation Power Analysis (CPA) of Pipelined Advanced Encryption Standard (AES)abstractWe evaluate the vulnerability of a pipelined Advanced Encryption Standard (AES) against Correlation Power Analysis (CPA) Side-Channel Attack (SCA). We identify that the registers in pipelined AES are most vulnerable against CPA SCA and propose a new power model targeting the switching activities of the registers. The proposed power model is constructed based on the Hamming Distance (HD) between the intermediate values stored in the registers in two consecutive clock cycles. Then, we analyze the vulnerability of pipelined AES under two scenarios. First, during regular pipeline operation where the device is performing AES pipeline operation. Second, in non-pipeline operation where we assume the adversaries can insert delay to the input of the device to increase the signal to noise ratio of the physical leakage information. The simulation results show that under regular pipelined operation, our proposed power model can reveal all the 16 key bytes in less than 4,900 traces, resulting in 4.7× more effective than the conventional power models. Under non-pipelined operation, our proposed power model requires only 590 traces to reveal all the 16 key bytes, which is 5.9× more effective than other power models. Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee |
ISCAS | 7 |
| 2019 | Global Template Projection and Matching Method for Training-Free Analysis of Delayered IC ImagesabstractPattern recognition algorithms have recently been pursued for automatic analysis of delayered IC images, i.e. the detection of circuit components. Wide experimentation on the existing training-based approaches are hampered by heavy data labeling, expensive model training, or long processing time. In this paper, we propose a global template projection and matching (GTPM) method that requires no training and a minimal amount of data labeling for circuit component detection. Our proposed GTPM method achieves a higher or comparable accuracy as the reported approaches while being more computationally efficient. Deruo Cheng, Yiqiong Shi, Tong Lin 0001, Bah-Hwee Gwee, Kar-Ann Toh |
ISCAS | 4 |
| 2019 | Low Gate-Count Ultra-Small Area Nano Advanced Encryption Standard (AES) DesignabstractWe present a low gate-count ultra-small area nano advanced encryption standard (AES) design. We achieve the low gate-count by the following means. First, we repeatedly reuse the area-critical circuits, i.e. one 8-bit Substitute-Box (S-Box) circuit and one 32-bit MixColumn circuit, for AES. Second, we cascade the input flip-flops (FFs) with our data transfer architecture so that the outputs of the MixColumn circuit are connected directly to the first 32-bit input FFs without extra multiplexing circuits. Third, the ShiftRow operation is implicitly performed by assigning the data sequence to the input FFs (during the S-Box and MixColumn operations). Fourth, we use independent XOR gates for AddRound and KeyExpansion operations. The collective means enables our design to feature 1457 gates, and to occupy 100um×100um area @ 65nm CMOS. When compared to the normalized area (@ 65nm CMOS) of the reported AES designs, our design features the smallest normalized area, 10% smaller than the most competitive reported AES design. Our design is targeted for ultra-small area applications including biomedical applications. Aparna Shreedhar, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Nay Aung Kyaw, L. Nalangilli, Wei Shu, Joseph Sylvester Chang, Bah-Hwee Gwee |
ISCAS | 8 |
| 2019 | A Highly Efficient Side Channel Attack with Profiling through Relevance-Learning on Physical Leakage InformationabstractWe propose a Profiling through Relevance-Learning (PRL) technique on Physical Leakage Information (PLI) to extract highly correlated PLI with processed data, as to achieve a highly efficient yet robust Side Channel Attack (SCA). There are four key features in our proposed PRL. First, variance analysis on PLI is implemented to determine the boundary of the clusters and objects of the clusters. Second, the nearest-neighbor k-NN variance clustering is used to reduce the sampling points of PLI by clustering the high variance sampling points and discarding the low variance sampling points of PLI measurements (traces). These clustered sampling points, which are highly correlated with the processed data, contain pertinent leakage information related to the secret key. Third, the information associated with the secret key is spread in several neighboring sampling points with different degrees of leakages. We analytically derive the Key-leakage relevance factor for each clustered sampling point to quantify the degree of leakage associated with the secret key. Fourth, by means of Hebbian learning, a weight proportional to the Key-leakage relevance factor is updated iteratively based on the values of relevance factor and traces of the sampling points. The converged weights which are being assigned to clustered sampling points are linked to their associated PLI to further increase the correlation of the PLI with the processed data. Therefore, the required number of PLI measurements, to reveal the secret key, can be reduced significantly. In addition, we analytically show that the computational complexity of our proposed PRL is O(n) when compared to the reported profiling techniques having O(n2) and O(n3) computational complexities. Based on the experiments of our proposed PRL performed on the PLI of AES-128 algorithm, the results depicting that the sampling points of PLI are reduced 87 percent after the k-NN variance clustering. The converged weight with learning error rate106traces, our proposed PRL is ~2; 000x more efficient in performing SCA. Ali Akbar Pammu, Kwen-Siong Chong, Yi Estelle Wang, Bah-Hwee Gwee |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2019 | A High Throughput and Secure Authentication-Encryption AES-CCM Algorithm on Asynchronous Multicore ProcessorabstractWe propose an authentication-based matrix-transformation cum parallel-encryption implemented on an asynchronous multicore processor (AMP-MP) to achieve a high throughput and yet secure advanced encryption standard based on counter with chaining mode (AES-CCM). There are four main features in our proposed AMP-MP. First, we employ the matrix multiplication in GF(28) computation to transform the 16 plaintexts into one plaintext, hence improving the authentication speed by 32× collectively at the transmitter and receiver. Second, we reschedule the operations of three AES encryptions in three different cores such that their physical leakages are compensated and equalized, thus reducing the correlation of physical leakage with the processed data by >3×. Third, the intermediate values of AES-CCM are propagated asynchronously between different cores to randomize the physical leakages with the processed data, and therefore further enhance the security of AES-CCM against the SCA by another 3×. Fourth, we propose a key adjusting technique based on S-Box byte-key transformation to protect the key against pattern-based attack. Our proposed AMP-MP is realized on an 8-bit asynchronous 9-core processor fabricated based on the 65 nm CMOS process. The experimental results show that the throughput of the authentication is 13.54 Gbps while the throughput for both authentication and encryption collectively is 8.32 Gbps, which are 17× and 70× faster than the reported counterparty, respectively. Based on power dissipation and EM SCA on our proposed AMP-MP, the secret key is unrevealed at 5 × 105 traces, which is ~17× more secured than the standard ASIC AES-CCM implementation. Ali Akbar Pammu, Weng-Geng Ho, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Bah-Hwee Gwee |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2018 | Asynchronous-Logic QDI Quad-Rail Sense-Amplifier Half-Buffer Approach for NoC Router DesignabstractWe propose a low area overhead and power-efficient asynchronous-logic quasi-delay-insensitive (QDI) sense-amplifier half-buffer (SAHB) approach with quad-rail (i.e., 1-of-4) data encoding. The proposed quad-rail SAHB approach is targeted for area- and energy-efficient asynchronous network-on-chip (ANoC) router designs. There are three main features in the proposed quad-rail SAHB approach. First, the quad-rail SAHB is designed to use four wires for selecting four ANoC router directions, hence reducing the number of transistors and area overhead. Second, the quad-rail SAHB switches only one out of four wires for 2-bit data propagation, hence reducing the number of transistor switchings and dynamic power dissipation. Third, the quad-rail SAHB abides by QDI rules, hence the designed ANoC router features high operational robustness toward process-voltage-temperature (PVT) variations. Based on the 65-nm CMOS process, we use the proposed quad-rail SAHB to implement and prototype an 18-bit ANoC router design. When benchmarked against the dual-rail counterpart, the proposed quad-rail SAHB ANoC router features 32% smaller area and dissipates 50% lower energy under the same excellent operational robustness toward PVT variations. When compared to the other reported ANoC routers, our proposed quad-rail SAHB ANoC router is one of the high operational robustness, smallest area, and most energy-efficient designs. Weng-Geng Ho, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | DPA-resistant QDI dual-rail AES S-Box based on power-balanced weak-conditioned half-bufferabstractWe propose an asynchronous-logic (async) Quasi-Delay-lnsensitive (QDI) dual-rail 32-bit Advanced Encryption Standard (AES) Substitution-Box (S-Box) for Differential Power Analysis (DPA) attack countermeasure. There are three novel features in the proposed S-Box. First, the proposed S-Box operates in async QDl protocol with dual-rail data encoding, hence there is only a marginal difference in power dissipation for different signal output transitions. Second, the proposed S-Box embodies the power-balanced async Weak-Conditioned Half-Buffer (WCHB) cell approach, which features the same number of transitions, and hence same number of switching for different input combinations to equalize the power dissipation. Third, the proposed S-Box embodies our novel-designed library cells in which each output wire, corresponding to different transitions, has a similar capacitive load, hence hiding the dynamic power dissipation. Based on the 65nm CMOS process, we implement the proposed 32-bit AES S-Box, and benchmark it against the conventional synchronous-logic (sync) S-Box and the reported async (i.e. unbalanced) WCHB S-Box. From the simulation results, our proposed power-balanced WCHB S-Box features significantly lower 1.21% and 0.55% of Normalized Energy Deviation (NED) and Normalized Standard Deviation (NSD) respectively. Particularly, our proposed design is of 57.6× and 17.9× lower NED, and 55.1× and 9.29× lower NSD than the reported sync and async counterparts respectively. James Lim, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee |
ISCAS | 4 |
| 2017 | A class-E RF power amplifier with a novel matching network for high-efficiency dynamic load modulationabstractWe present in this paper a proposed high-efficiency Class-E power amplifier (PA) for RF Polar transmitters. The PA embodies a proposed novel matching network (MN) with three salient features. First, it is digitally-controlled, and can directly receive the digital Amplitude Modulation (AM) input data to the PA without the need for a conventional supply modulator. Second, the MN performs high-efficiency dynamic load modulation, where load impedance seen by the PA is varied by the MN according to the AM data, and simultaneously, this load impedance is also ensured by the MN to satisfy the zero-voltage-switching condition at the PA to result in high-efficiency operation. Third, it has a novel architectural design that can employ on-, or off-chip inductors, or both inductor types; high quality-factor bond-wires can therefore be used as the inductors to improve the efficiency. The proposed PA with the MN is designed using a 40 nm CMOS technology. Simulation results at 2.4 GHz and 1.1 V supply show that the PA achieves a high power efficiency (drain efficiency) of 48% at peak output power of 17 dBm. Victor Adrian, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 3 |
| 2017 | Highly secured state-shift local clock circuit to countermeasure against side channel attackabstractWe propose a highly-secured State-shift Local Clock (SsLC) countermeasure technique to hide the Physical Leakage Information (PLI) against Side Channel Attack (SCA). The SCA is a technique employed to reveal the secret key of cryptographic algorithm by correlating the PLI (i.e. power dissipation and Electromagnetic (EM)) with the processed data, where both the PLI and processed data are generated during the encryption process. Whereas the countermeasure technique aims to reduce the correlation of the PLI against the processed data. There are four key features in our proposed SsLC countermeasure technique. First, it embodies a finite state machine which can be employed to regularly shift the timing operation of cryptographic algorithm implementations. Thus, the correlation of the PLI with the processed data is significantly reduced due to dynamically changes the occurrences of encryption operation in time domain. Second, the PLI which encompasses a secret key is spread over in time domain to reduce the probability of revealing the secret key. Third, the power dissipation overhead is negligible and hence it is highly applicable for low power applications. Fourth, the regular state (time) shifting technique in the SsLC is able to hide multiple PLIs, i.e. power dissipation and EM signals, concurrently. In view of the above features, the proposed SsLC is highly secured against SCA with multiple PLIs. Based on the experimental results in FPGA, our proposed SsLC countermeasure technique features wide distribution of PLI in time domain, dissipates 2.77mW of power and emits 12.2mV/m of EM signal @ 2.4MHz. Furthermore, with 106 power dissipation and EM measurements, the secret key of the cryptographic algorithm remains unbreakable. In comparison with the reported counterparts, the resistance of our proposed SsLC against SCA is significantly improved as the number of power dissipation and EM traces to reveal the secret key has increased by >18x and >25x respectively. Consequently, the correlation coefficient between the PLI and the processed data is reduced by 3.5x. Ali Akbar Pammu, Kwen-Siong Chong, Bah-Hwee Gwee |
ISCAS | 3 |
| 2017 | Sense Amplifier Half-Buffer (SAHB) A Low-Power High-Performance Asynchronous Logic QDI Cell TemplateabstractWe propose a novel asynchronous logic (async) quasi-delay-insensitive (QDI) sense-amplifier half-buffer (SAHB) cell design approach, with emphases on high operational robustness, high speed, and low power dissipation. There are five key features of our proposed SAHB. First, the SAHB cell embodies the async QDI 4-phase (4φ) signaling protocol to accommodate process-voltage-temperature variations. Second, the sense amplifier (SA) block in SAHB cells embodies a cross-coupled latch with a positive feedback mechanism to speed up the output evaluation. Third, the evaluation block in the SAHB comprises both nMOS pull-up and pull-down networks with minimum transistor sizing to reduce the parasitic capacitance. Fourth, both the evaluation block and SA block are tightly coupled to reduce redundant internal switching nodes. Fifth, the SAHB cell is designed in CMOS static logic and hence appropriate for full-range dynamic voltage scaling operation for VDDranging from nominal voltage (1 V) to subthreshold voltage (~0.3 V). When six library cells embodying our proposed SAHB are compared with those embodying the conventional async QDI precharged half-buffer (PCHB) approach, the proposed SAHB cells collectively feature simultaneous -.64% lower power, -.21% faster, and ~6% smaller IC area; the PCHB cell is inappropriate for subthreshold operation. A prototype 64-bit Kogge-Stone pipeline adder based on the SAHB approach (at 65 nm CMOS) is designed. For a 1-GHz throughput and at nominal VDD, the design based on the SAHB approach simultaneously features -.56% lower energy and -.24% lower transistor count advantages than its PCHB counterpart. When benchmarked against the ubiquitous synchronous logic counterpart, our SAHB dissipates -.39% lower energy at the 1-GHz throughput. Kwen-Siong Chong, Weng-Geng Ho, Tong Lin 0001, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Low normalized energy derivation asynchronous circuit synthesis flow through fork-join slack matching for cryptographic applications
Nan Liu 0002, Kwen-Siong Chong, Weng-Geng Ho, Bah-Hwee Gwee, Joseph Sylvester Chang |
DATE | 4 |
| 2016 | High performance low overhead template-based Cell-Interleave Pipeline (TCIP) for asynchronous-logic QDI circuitsabstractWe propose a novel Template-based Cell-Interleave Pipeline (TCIP) approach for generating high performance and yet low overhead asynchronous-logic (async) quasi-delay-insensitive (QDI) circuits. Our TCIP approach exploits the characteristics of the four prevalent QDI cell templates, namely Weak-Conditioned Half-Buffer (WCHB), Pre-Charged HalfBuffer (PCHB), Autonomous Signal-Validity Half-Buffer (ASVHB), and Sense-Amplifier Half-Buffer (SAHB), and then strategically interleave these template cells to form a composite pipeline. There are three main features in our TCIP approach. First, all QDI cell templates are first standardized with the same interface signals, and their corresponding cells are characterized in terms of transistor count, cycle time and energy dissipation for ease of comparison/selection/replacement. Second, our TCIP approach prioritizes the speed requirement when forming the initial pipeline circuits, and then subsequently reduces circuit overheads by interleaving various template cells without compromising the speed significantly. Third, the final optimized QDI pipeline circuit inherently features high robustness against process-voltage-temperature (PVT) variations, hence suitable for dynamic-voltage-scaling (DVS) operation. By means of 65nm CMOS process, we demonstrate a 4-bit pipeline tree adder based on the proposed TCIP approach, and benchmark it against the WCHB, PCHB, ASVHB and SAHB counterparts. These five designs feature same high operational robustness, nonetheless the design based on our TCIP approach is more competitive. Particularly, the designs based on reported approaches are, on average, ∼1.22× more transistor count, ∼1.21× slower and ∼1.22× higher energy dissipation. Furthermore, under DVS operation from 1.2V to 0.3V, our proposed TCIP adder can reduce up to 88% energy for non-speed critical applications. Weng-Geng Ho, Nan Liu 0002, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 5 |
| 2016 | Area-efficient and low stand-by power 1k-byte transmission-gate-based non-imprinting high-speed erase (TNIHE) SRAMabstractWe propose a novel 15-T Transmission-gate-based Non-Imprinting High-speed Erase (TNIHE) SRAM cell with emphases on low area overhead and low stand-by power attributes for highly secured data storage applications. We benchmark our proposed 15-T TNIHE SRAM cell against the reported 22-T Non-Imprinting High-speed Erase (NIHE) SRAM cell, and demonstrated three key features of reducing 7 transistors. First, we adopt the transmission gates (as opposed the cross-couple inverters) in the slave circuitry, saving 4 transistors. Second, we eliminate a transistor which uses to reset the slave circuitry, hence saving 1 transistor. Third, we apply the global inverse transistors (as opposed to the local inverse transistors) in the read /write circuit for each SRAM cell, hence further reduce 2 more transistors. As a result, our proposed TNIHE SRAM cell @ 65nm CMOS features ~17% smaller layout area. We design a 1k-byte memory based on the proposed TNIHE SRAM cells. On the basis of simulations, we show that our 1k-byte SRAM memory features overall ~13% smaller area, and dissipates on average, ~30% lower stand-by power than the reported NIHE counterpart. Weng-Geng Ho, Ne Kyaw Zwa Lwin, N. Prashanth Srinivas, Kwen-Siong Chong, Tony Tae-Hyoung Kim, Bah-Hwee Gwee |
ISCAS | 6 |
| 2016 | Secured Low Power Overhead Compensator Look-Up-Table (LUT) Substitution Box (S-Box) ArchitectureabstractSubstitution-Box (S-Box) is an important security building block for the Advanced Encryption Standard (AES) algorithm. However, its high power dissipation always compromises with its security feature under Correlation Power Analysis (CPA) attack. In this paper, we propose a secured and low power overhead LUT based S-Box architecture embodying a novel multiplexing circuit AND and OR a compensator. We achieve these attributes as follows. First, we employ AND and OR gates to realize the multiplexing circuit therein in a regular structure to minimize the delay and power variations for every input pattern, hence mitigating the security risk against CPA. Second, we augment a compensator to complement the multiplexing circuit to further minimize the power variations within the LUT based S-Box. We realize six AES designs based on the Sakura-X FPGA board, three designs embodying reported S-Box architectures and the other three designs leveraging on our multiplexing circuit and compensator. We show that our AES design, embodying our LUT based S-Box architecture with the AND/OR-gate multiplexing circuit and compensator, has the highest security feature (against CPA) compared with the reported designs, featuring 10× to 300× better security. Ali Akbar Pammu, Kwen-Siong Chong, Bah-Hwee Gwee |
NAS | 3 |
| 2015 | High robustness energy- and area-efficient dynamic-voltage-scaling 4-phase 4-rail asynchronous-logic Network-on-Chip (ANoC)abstractWe propose an 18-bit 5-interface asynchronous-logic Network-on-Chip (ANoC) router based on the quasi-delay-insensitive (QDI) realization approach for high secured cryptography applications. There are four key features of the proposed ANoC router. First, it embodies the novel high-speed low-power Sense-Amplifier Half Buffer 4-rail cells. Second, it is designed based on QDI protocol, and hence is highly robust against process-voltage-temperature (PVT) variations. Third, it is functional for full dynamic voltage scaling from nominal (VDD=1.2V) to sub-threshold (VDD=0.3V) regions, and is potentially excellent for low power management applications. Fourth, it embodies a distributed-based XY routing algorithm to utilize a 4-bit header of flow control unit (flit) for routing up to 4×4 cluster, hence minimizing the routing overhead. We realize the proposed ANoC router (@65nm CMOS), and benchmark it against the reported ANoC router embodying the conventional Weak-Conditioned Half-Buffer (WCHB) QDI realization approach. Both our proposed and reported designs feature the high operation robustness, but our design is 41% more energy-efficient, and 21% more area-efficient than the reported counterpart. The prototype of ANoC router occupies only 0.105 mm2and can operate down to 0.3V. At VDD=0.3V, it dissipates 44 fJ per bit and operate 105 ns per flit. Weng-Geng Ho, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2015 | Novel real-time system design for floating-point sub-Nyquist multi-coset signal blind reconstructionabstractWe propose a novel real-time system design for multiband signal blind reconstruction using multi-coset sampling theory. Multi-channel signals are acquired under sub-Nyquist sampling frequency to perfectly reconstruct the original signal spectrum. A novel system design with Field-Programmable Gate Array (FPGA) implementation is presented in this paper. There are two main contributions in this paper. Firstly, the FPGA system uses 32-bit single precision floating point dataflow rather than conventional 16-bit fixed point to recover signals with much lower Signal-Noise Ratio (SNR). Secondly, we introduce a novel Jacobi CORDIC eigenvalue decomposition (EVD) core using parallel pivot-seeking circuit and parallel 3-CORDIC design to improve speed significantly. Hermitian matrices of dimensions from 2 to 10 are tested to compare conventional 2-CORDIC EVD and proposed EVD. The proposed EVD effectively reduces on average 36% of processing time for mesh connection system and over 50% for parallel system. Hongxu Yin, Bah-Hwee Gwee, Zhiping Lin 0001, Achanna Anil Kumar, Sirajudeen Gulam Razul, Chong Meng Samson See |
ISCAS | 2 |
| 2015 | A single-VDD half-clock-tolerant fine-grained dynamic voltage scaling pipelineabstractWe propose a novel dynamic voltage scaling (DVS) pipeline with three significant attributes. First, it features a fine-grained DVS which innately attempts to power most of the circuits therein at low voltages, and when the speed is beneath the requirement, to scale up the voltage. Second, it supports fast-transition DVS within one-and-a-half clock duration per operation, and its operation remains error-free during that duration; we define such attribute as half-clock-tolerant. Third, it consists of a single power source (single-VDD) which supports three voltage scales (1.2V, 0.8V and 0.5V) for power/speed tradeoffs, and has standardized 1.2V output to seamlessly interface with other proposed/conventional pipelines. These attributes are achieved due to the embodiment of a DVS power unit, asynchronous building blocks to control/synchronize the operation, a dual-rail critical path to innately detect the completion of the operation, and level shifters to standardize the output voltage. We demonstrate our proposed pipeline by designing a multiplier embodied in a Fast Fourier Transform processor (@65nm CMOS). We show that the multiplier based on our proposed pipeline, on average, is 1.94× more power-efficient than that based on a conventional pipeline. Kwen-Siong Chong, Tong Lin 0001, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2014 | A Randomized Modulation scheme for filterless digital Class D audio amplifiersabstractWe propose to employ the Randomized Wrapped-Around Pulse Position Modulation scheme (RWAPPM) to mitigate the switching-frequency harmonics at the output signal of filterless digital Class D audio amplifiers. The conventional Pulse Width Modulation schemes (PWMs) typically have a non-zero common-mode voltage that contributes to the radiated Electromagnetic Interference (EMI), and generate high switching-frequency harmonics that dissipate extra power at the speaker and also contribute to the radiated EMI. We simulate and compare the RWAPPM (2-level) against the PWMs and a reported randomized modulation scheme. The 2-level RWAPPM has zero common-mode voltage, and amongst the modulation schemes, it features the highest attenuation of the switching-frequency harmonics, highest out-of-band Spurious Free Dynamic Range (22 dBc), and a relatively high Signal to Noise and Distortion Ratio (57 dB) at the output voltage. Victor Adrian, Cui Keer, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 3 |
| 2014 | Synthesis of asynchronous QDI circuits using synchronous coding specificationsabstractWe propose a synthesis of asynchronous quasi-delay-insensitive (QDI) circuits. We highlight three notably features/novelties of the proposed synthesis as follows. First, the targeted synthesized circuits abide by the QDI protocol; hence they are inherently timing-robust and are desirable for applications with high variation-space and wide operation-space (including defense/space applications). Second, the coding specifications accept Verilog HDL language, and are the same/similar to the standard coding for synchronous circuits, hence no special and/or ad-hoc design/coding rules are required. Third, the proposed synthesis is applicable to accept various QDI library cells, hence enabling to explore full merit of different library cells. To the best of our knowledge, no reported synthesis methods incorporate all these features; some limited features were only incorporated. Our proposed synthesis, at this juncture, accepts three basic clauses - complete `if-else' clause, incomplete `if-else clause', and the `case' clause. These clauses are more than sufficient to describe any complex systems. The synthesis stages involve analyzing QDI pipelines, generating (corresponding) single-rail combinational circuits, converting dual-rail netlists (from the single-rail circuits), and embedding customized controllers. In order to demonstrate the validity and practicality of the proposed synthesis, an 8-bit 8-tap asynchronous QDI Finite Impulse Response (FIR) filter is synthesized, implemented to the layout stage, and evaluated using spice models-specifically, it features 3.7 mW power dissipation, 39,181 transistors, and a delay of 200 ns per operation. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang, Weng-Geng Ho |
ISCAS | 3 |
| 2014 | A Low Overhead Quasi-Delay-Insensitive (QDI) Asynchronous Data Path Synthesis Based on Microcell-Interleaving Genetic Algorithm (MIGA)abstractIn this paper, we propose a design approach to mitigate the hardware overhead of the data completion detection circuit in quasi-delay-insensitive (QDI) asynchronous-logic circuits. In this proposed design approach, three novelties are highlighted. Firstly, a novel microcell-interleaving approach is proposed to reduce the number of completion detection (CD) circuits while retaining the required QDI attribute. Secondly, we analyze the performance of the QDI circuits based on the proposed microcell-interleaving approach graphically in terms of power dissipation, transistor count and delay, and evaluate/determine the upper and lower boundaries of these performance profiles. Thirdly, we propose a microcell-interleaving genetic algorithm (MIGA) to stochastically optimize the proposed microcell-interleaving approach on power dissipation, transistor count, and delay. To validate the proposed design approach, a complete performance profile of ISCAS-85 C499 circuit is investigated on the basis of differential cascode voltage switch logic (DCVSL) and dynamic strong indicating (DSI) microcells. We demonstrate the efficiency of the proposed design approach by benchmarking against the competing DCVSL, null convention logic and DSI designs on five ISCAS-85 circuits. Specifically, the proposed designs, on average, are 1.77 × better in power dissipation, 1.4 × better in area, and 1.58 × better in a composite metric of power × area × delay, and reasonably slower for the lowest power dissipation points. We further demonstrate the practicality of the proposed design approach by implementing an 8-tap 16-bit asynchronous QDI finite impulse response filter. Finally, we demonstrate the ~10% and ~11% improved efficiency of the proposed MIGA over the greedy algorithm and dynamic programming, respectively. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2013 | A dual-core 8051 microcontroller system based on synchronous-logic and asynchronous-logicabstractWe describe a dual-core 8051 microcontroller system featuring the synchronous and asynchronous (clockless) mode of operation. The synchronous mode of operation is achieved by means of a synchronous 8051 microcontroller core, while the asynchronous mode of operation is achieved by means of an asynchronous 8051 microcontroller core. The 8051 microcontroller system features shared embedded program and data memories that enable the switching between the two microcontroller cores during program execution. The measured energy, speed and electromagnetic interference of both microcontroller cores will be compared at different operation workloads. Kok-Leong Chang, Tong Lin 0001, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 5 |
| 2013 | Low power sub-threshold asynchronous QDI Static Logic Transistor-level Implementation (SLTI) 32-bit ALUabstractWe propose an asynchronous-logic (async) Quasi-Delay-Insensitive (QDI) Static Logic Transistor-level Implementation (SLTI) approach for low power sub-threshold operation. The approach is implemented to design 32-bit pipelined Arithmetic and Logic Units (ALUs), the primary computation core for microprocessors, and benchmarked against the reported Pre-Charged Half-Buffer (PCHB). There are two key attributes in this proposed design. First, the proposed SLTI ALU design can perform dynamic voltage scaling seamless by only changing the supply voltage from nominal (1V) to sub-threshold (~0.2V) regions for high speed/low power operation. Second, the ALU achieves ultra-low power dissipation (3.5μW) at the lowest VDDpoint (~0.15V). For fair of comparison, both implemented ALUs have identical functionality and functional blocks, are implemented using the same 65nm CMOS process. Based on the simulations, the minimum energy point occurs at VDD= 0.2V for SLTI-based ALU and at VDD= 0.3V for PCHB-based ALU. The SLTI-based ALU have ~93% and ~89% lower energy on the arithmetic and logic operations respectively from VDD= 1V to VDD= 0.2V. At VDD= 0.2V, with 9MHz input switching rate, the async ALU based on our proposed SLTI approach dissipates ~51% and ~44% lower power than the reported PCHB counterpart on the arithmetic and logic operations respectively. Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 3 |
| 2012 | A comparative study on asynchronous Quasi-Delay-Insensitive templatesabstractThe robustness of asynchronous logic has proved useful in dealing with contemporary problems in CMOS design such as process variations and power management. However, the general cryptic nature of asynchronous logic has stymied the widespread acceptance of this alternate design technique. Fortunately, the semi-custom approach to asynchronous design reduces the tedious handcrafting efforts that are often non-trivial in large system-on-chips (SoCs). However, even with the adoption of this design approach requires careful selection of asynchronous templates that will suit overall system needs. Therefore in this paper, the most eminent Quasi-Delay-Insensitive asynchronous template families reported to date will be presented, and followed by an in-depth comparison of various design FOMs - template area, static/dynamic capacity, cycle time, latency, throughput and Et2. The most aggressive template (EESTFB) can reach a maximum throughput of 3.56Giga items/s on 0.13µm @ 1.2V. Kok-Leong Chang, Tong Lin 0001, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 5 |
| 2012 | An Ultra-Dynamic Voltage Scalable (U-DVS) 10T SRAM with bit-interleaving capabilityabstractWe propose a dynamic voltage scalable SRAM capable of efficient bit-interleaving in column to tolerate multiple-bits soft error when integrated with error correction codes (ECC). First, a 10T SRAM bitcell is proposed. It activates only intended bitcells so that stability problem of half-selected bitcells is completely eliminated and the power dissipation in half-selected columns is significantly reduced. Second, a configurable DVS scheme is employed to enable the bitcell to operate like differential 8T during super-threshold region which results in faster operation. The proposed SRAM can operate up to 1.2GHz at 1.2V using 65nm CMOS process. Third, a segmented column multiplex with low overhead is proposed, which greatly reduces the power dissipation due to the column control signals. Consequently, the write and read power dissipations are reduced by up to 40% and 67% respectively. Forth, a hierarchical read bitline is used to reduce the read bitline discharge delay variation due to local and global process variation in subthreshold region, which is a major portion of memory access time. Based on our simulation results, the worst case read bitline discharge delay is reduced by more than 12× at VDDof 0.3V. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 3 |
| 2012 | Energy-delay efficient asynchronous-logic 16×16-bit pipelined multiplier based on Sense Amplifier-Based Pass Transistor LogicabstractWe describe an asynchronous-logic (async) 16×16-bit pipelined multiplier based on our proposed Sense Amplifier-Based Pass Transistor Logic (SAPTL) with emphases on high energy-delay efficiency. The multiplier is targeted for an async multi-core System-On-Chip (SOC). This attribute is achieved by simplifying and optimizing the NMOS pass transistor stacks and decision-making C-element, therein to reduce the circuit area overheads and transistor switchings in SAPTL. Based on the simulations (@1V, 65nm CMOS process), the async 16×16-bit pipelined multiplier based on our proposed SAPTL approach features, on average, 31% shorter delay, 21% lower energy/operation achieving a total of 46% lower energy-delay product, and 16% lesser number of transistors when compared to the reported SAPTL approaches. Weng-Geng Ho, Kwen-Siong Chong, Tong Lin 0001, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2011 | A low-power dual-rail inputs write method for bit-interleaved memory cellsabstractWe propose a dual-rail data write technique for bit interleaved memory cells to reduce power dissipation for the write operation without affecting the read operation. The proposed technique can be applied to two reported bit interleaved memory cells with a write power reduction range from 30% to 45%, depending on memory cells and operations. In addition, in the proposed technique, a subthreshold non bit interleaved memory cell is modified to be bit-interleaved without increasing the number of transistors in memory cell. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 3 |
| 2011 | Improved asynchronous-logic dual-rail Sense Amplifier-based Pass Transistor Logic with high speed and low power operationabstractWe propose a robust asynchronous-logic dual-rail Sense Amplifier-based Pass Transistor Logic (SAPTL) approach with improved speed and power attributes over reported SAPTL approach. These attributes are achieved by simplifying various sub-blocks therein to reduce the stacking of pass transistors and the number of transistor switchings, and to avoid floating nodes. By means of an 8-bit pipeline adder and on the basis of computation simulations (@ 1V, 45nm SOI process), we show that our proposed SAPTL adder is 37% faster, yet 14% lower power dissipation (@ 200MHz input-rate), 18% lower energy dissipation (per operation), and 47% better energy-delay product. These substantially improved attributes are achieved with insignificant overhead - just 3% more transistors. Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang, Yin Sun 0005, Kok-Leong Chang |
ISCAS | 3 |
| 2011 | Modeling and Synthesis of Asynchronous PipelinesabstractWe propose a set of modeling rules and a synthesis method for the design of asynchronous pipelines. To keep the circuit area and power dissipation of the asynchronous control network small, the proposed approach avoids the conventional syntax-directed translation approach. Instead, it employs a data-driven design style and a coarse-grain approach to the synthesis of asynchronous control, restricting asynchronous control to the implementation of communication channels commonly found in asynchronous pipelines and operations involving these channels. The proposed approach integrates well into conventional synchronous design flows because they are based on Verilog and SystemVerilog specifications, and generate register-transfer level models suitable for functional simulation and logic synthesis using existing computer-aided design tools. Using a 32-bit microprocessor, an interpolated finite-impulse-response filter bank, and a Reed-Solomon error detector as design examples, we show that the proposed approach is competitive with other comparable reported methods. Chong-Fatt Law, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | A highly efficient method for extracting FSMs from flattened gate-level netlistabstractThis paper proposes a novel method for extracting Finite State Machines (FSMs) from flattened gate-level netlist. The proposed method which employs a potential state register elimination technique and a two-level FSM separation strategy is highly applicable to control-intensive circuits. The potential state register elimination technique is based on control signal identification whereas the two-level FSM separation strategy is based on enable tree identification and the strongly connected components algorithm. To demonstrate the efficacy and to illustrate the unique features of the proposed FSM extraction method, the Synopsys DesignWare DW8051 microcontroller is used as the benchmark circuit for comparison and simulations. Results show that the proposed method reduces the complexity of the extracted FSMs in terms of number of state registers in an FSM by more than 90% as compared to the reported technique. Yiqiong Shi, Chan Wai Ting, Bah-Hwee Gwee, Ye Ren |
ISCAS | 3 |
| 2009 | A Performance Comparison on Asynchronous Matched-delay TemplatesabstractThe motivation for asynchronous logic at this juncture of CMOS technology is the issues of power density, process variation and integration limit, where synchronous logic is facing a myriad of problems. Asynchronous templates are the fundamental building blocks of asynchronous circuits and systems, and together with asynchronous EDA tools enable the design of complex systems at a high level of abstraction (similar to the RTL-to-GDSII flow in synchronous design). However, akin to the impact of library cells to the overall system performance in the conventional synchronous flow, the diverse availability of asynchronous template libraries requires prudent contemplation. Therefore in this paper, the most eminent matched-delay asynchronous template families reported to date will be presented, and followed by an in-depth comparison of various design figure of merits (FOMs) - template area, static/dynamic capacity, cycle time, latency, throughput and Et2. The most aggressive template (GasP) can reach a maximum throughput of 5 Giga items/s on 0.13 mum @ 1.2 V. Kok-Leong Chang, Bah-Hwee Gwee, Yuanjin Zheng |
ISCAS | 2 |
| 2009 | Fine-grained Power Gating for Leakage and Short-circuit Power Reduction by using Asynchronous-logicabstractIn this paper, a fine-grained power gating technique for an asynchronous-logic pipeline stage is proposed using locally controlled gating transistors. The proposed power gating technique is implemented with minimal control overheads (one additional inverter per pipeline stage for driving PMOS Gating) and delay overheads (within 15% more than the conventional asynchronous-logic pipeline stage). Different types of gating configurations using only PMOS transistor (PMOS Gating), only NMOS transistor (NMOS Gating), and both types of transistors (Dual Gating) are examined and compared. The effectiveness of the proposed power gating technique to the Combinational Block therein with different data input rates is investigated. Based on the computer simulation results, we have found that ≫70% wasted power reduction (including both short-circuit and leakage powers) as compared to the conventional asynchronous-logic pipeline stage can be achieved with all gating configurations. In particular, Dual Gating achieves the best wasted power reduction of 86% for short-circuit power and 99% for leakage power @ 10Mbps input rate. Tong Lin 0001, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 3 |
| 2008 | A semi-custom memory design for an asynchronous 8051 microcontrollerabstractIn this paper, we propose a methodology for interfacing synchronous IP memory blocks (read-only memory (ROM) and random-access memory (RAM)) with asynchronous-logic digital systems based on dual-rail, 4-phase signaling. The memory blocks (ROM and RAM) of an instruction-set compatible 8051 microcontroller (A8051) is implemented with Artisan IP memory blocks for the IBM 0.13μm CMOS technology. Interface circuits play the role of (1) synchronizing all the asynchronous input channels driving the IP memory blocks, (2) single rail signaling logic to dual-rail 4-phase signaling logic conversion and vice versa, and (3) capturing synchronous signals in memory read cycles and driving asynchronous channels. The A8051 with the proposed ROM and RAM design operates at 28% higher MIPS rate (millions of instructions per second), dissipates 20% lower energy per instruction, ∼50% lower Et2and occupies 19% lesser area, as compared to the A8051 with register-based memory. Kok-Leong Chang, Bah-Hwee Gwee, Yuanjin Zheng |
ISCAS | 2 |
| 2008 | De-synchronization of a point-of-sales digital-logic controllerabstractIn this paper, we propose a methodology to de- synchronize a synchronous digital-logic system to obtain a system based-on asynchronous logic with equivalent input/output (I/O) functionality. The motivation is to compare the performance of synchronous and asynchronous implementations, especially on power dissipation and process variation robustness. To de- synchronize the controller, several transformations are made to the synchronous controller, such as (1) removing the clock, (2) replacing registers with asynchronous handshaking latches, (3) inserting matching delays, and (4) inserting pipeline buffers in feedback paths. We apply the de-synchronization methodology to a point-of-sales (POS) digital-logic controller modeled with Verilog hardware description language (HDL), and is based-on the Moore finite state machine (FSM). Gate-level simulation verifies that the asynchronous implementation has equivalent functionality as the synchronous controller. Power simulations show that the asynchronous controller consumes only a fraction of the power (5.4%) of the synchronous controller. Kok-Leong Chang, Bah-Hwee Gwee |
ISCAS | 3 |
| 2008 | Asynchronous Control Network Optimization Using Fast Minimum-Cycle-Time AnalysisabstractThis paper proposes two methods for optimizing the control networks of asynchronous pipelines. The first uses a branch-and-bound algorithm to search for the optimum mix of the handshake components of different degrees of concurrence that provides the best throughput while minimizing asynchronous control overheads. The second method is a clustering technique that iteratively fuses two handshake components that share input channel sources or output channel destinations into a single component while preserving the behavior and satisfying the performance constraint of the asynchronous pipeline. We also propose a fast algorithm for iterative minimum-cycle-time analysis. The novelty of the proposed algorithm is that it takes advantage of the fact that only small modifications are made to the control network during each optimization iteration. When applied to nontrivial designs, the proposed optimization methods provided significant reductions in transistor count and energy dissipation in the designs' asynchronous control networks while satisfying the throughput constraints. Chong-Fatt Law, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2007 | An Asynchronous Dual-Rail Multiplier based on Energy-Efficient STFB TemplatesabstractIn this paper, we describe an asynchronous (async) dual-rail 13×13-bit multiplier based on the single-track full-buffer (STFB) template. We propose several techniques to improve the energy-efficiency of the template. Firstly, we propose a new output driver sub-cell for the template suitable for driving smaller loads with higher energy efficiency. Secondly, we propose non-handshaking channels in order to reduce the pipeline stages in our design to trade-off throughput for higher energy-efficiency and smaller area. Lastly we propose using non weak-condition AND, 3-to-2 and 2-to-2 compressor cells to achieve lower forward latency. The performance of the proposed multiplier design is simulated using the TSMC 0.18μm library at the transistor level. The proposed design is 15% more energy-efficient, has 14% lower latency and 34% smaller as compared to the same implementation using the STFB template. Kok-Leong Chang, Bah-Hwee Gwee, Yuanjin Zheng |
ISCAS | 2 |
| 2007 | A Low Energy FFT/IFFT Processor for Hearing AidsabstractWe present a 16-bit low voltage (1.1V - 1.4V) energy efficient 128-point decimation-in-time Fast Fourier Transform/Inverse Fast Fourier Transform (FFT/IFFT) processor specifically for hearing aid applications. The FFT/IFFT processor embodies several low power/energy methodologies, including the clock gating approach,ad-hoccontrol, operand isolation and low power library cells, to satisfy the tight constraints of low voltage low energy and a small silicon area realization for a practical hearing aid. Based on the prototype IC measurements, the proposed FFT/IFFT processor dissipates ~ 188nJ @ 1.1V, features computation delay of2@ 0.35μm CMOS process. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 2 |
| 2007 | A 32-point FFT based Noise Reduction Algorithm for Single Channel Speech SignalsabstractIn this paper we propose a noise reduction system for single channel speech signals. The system comprises of a 32-point FFT based spectral subtraction method, a variable least mean squares (LMS) filter and a 3-state voice activity detector. The LMS filter is designed to remove musical noise generated during spectral subtraction. It takes the original signal and the result of spectral subtraction as its inputs, and exploits the statistical differences between musical noise and speech to give a natural sounding output. The voice activity detector updates the noise spectrum estimate for spectral subtraction and controls the LMS filter parameters to make it adapt better to the incoming signal. Test cases involving three different non-stationary noise environments resulted in an average improvement of 6dB in the SNR after processing with low musical noise. Kunal Mukherjee, Bah-Hwee Gwee |
ISCAS | 2 |
| 2006 | An acoustic noise suppression system with reduced musical artifactsabstractIn this paper, we propose an acoustic noise suppression system with reduced musical artifacts for digital hearing instruments (aids). The proposed system features the capabilities to detect, estimate and suppress the acoustic noise corrupting an input speech. The algorithms in the system consist of two noise detections, an enhanced parametric spectral subtraction, a noise attenuation and transition smoothing window, and an automatic gain control. Simulation results on several stationary and nonstationary noise show that our acoustic noise suppression system is capable of improving signal-to-noise ratio by > 9 dB and reducing musical artifacts. Victor Adrian, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 2 |
| 2006 | A low-energy low-voltage asynchronous 8051 microcontroller coreabstractIn this paper, we propose a low-energy, low-voltage (1.1 V) 8051 microcontroller core using asynchronous logic based on Austria Micro Systems (AMS) 0.35 /spl mu/m technology for hearing instrument (hearing aid) applications. The A8051 microcontroller features precise functionality at voltages even down to 0.9 V. We propose several novel techniques to achieve low energy dissipation. Firstly, to minimize the activity of the system, we divide the microcontroller into a 2-stage pipeline. Both stages of the pipeline operate almost independent of each other, thus minimizing the chances of pipeline stalls. Secondly, to be energy efficient, we eliminate predictive algorithms; fetching every instruction and data only if necessary. Thirdly, identifying regular opcode patterns and using partial decodings to achieve a 30% reduction in energy dissipation. Lastly, we propose a unique indirect data memory fetching method where an indirect data can be fetched in one memory request cycle. Transistor-level simulations show that our proposed A8051 microcontroller operates at 0.6 MIPS and dissipates 130 pJ/instruction @ 1.1V. Kok-Leong Chang, Bah-Hwee Gwee |
ISCAS | 2 |
| 2005 | A micropower low-voltage multiplier with reduced spurious switchingabstractWe describe a micropower 16/spl times/16-bit multiplier (18.8 /spl mu/W/MHz @1.1 V) for low-voltage power-critical low speed (/spl les/5 MHz) applications including hearing aids. We achieve the micropower operation by substantially reducing (by /spl sim/62% and /spl sim/79% compared to conventional 16/spl times/16-bit and 32/spl times/32-bit designs respectively) the spurious switching in the Adder Block in the multiplier. The approach taken is to use latches to synchronize the inputs to the adders in the Adder Block in a predetermined chronological sequence. The hardware penalty of the latches is small because the latches are integrated (as opposed to external latches) into the adder, termed the latch adder (LA). By means of the LAs and timing, the number of switchings (spurious and that for computation) is reduced from /spl sim/5.6 and /spl sim/10 per adder in the adder block in conventional 16/spl times/16-bit and 32/spl times/32-bit designs respectively to /spl sim/2 in our designs. Based on simulations and measurements on prototype ICs (0.35 /spl mu/m three metal dual poly CMOS process), we show that our 16/spl times/16-bit design dissipates /spl sim/32% less power, is /spl sim/20% slower but has /spl sim/20% better energy-delay-product (EDP) than conventional 16/spl times/16-bit multipliers. Our 32/spl times/32-bit design is estimated to dissipate /spl sim/53% less power, /spl sim/29% slower but is /spl sim/39% better EDP than the conventional general multiplier. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | A Hybrid Genetic Hill-climbing Algorithm for Four-Coloring Map Problems
Bah-Hwee Gwee, Joseph Sylvester Chang |
HIS | 1 |
| 2000 | An investigation on the parameters affecting total harmonic distortion in class D amplifiersabstractIn this paper, we investigate two important and practical design parameters for the design of low-voltage low-power class D amplifiers that may affect the Total Harmonic Distortion (THD): the linearity of the carrier waveform and the impedance of the output stage. By means of a novel mathematical analysis method to model the carrier non-linearity, we show that this non-linearity should be mitigated to achieve low THD. Our mathematical analysis also provides an insight to the degree of non-linearity acceptable for a practical design. We show that the impedance of the output stage has little effect on THD. We verify our analysis by means of MATLAB and SPICE simulations. Meng Tong Tan, Hock-Chuan Chua, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 3 |
| 2000 | A GA with heuristic-based decoder for IC floorplanning
Bah-Hwee Gwee, Meng-Hiot Lim |
Integr. | 1 |
| 1996 | A GA paradigm for learning fuzzy rules
Susanto Rahardja, Bah-Hwee Gwee |
Fuzzy Sets Syst. | 3 |