Kwen-Siong Chong

dblp:53/5938 · DBLP profile ↗
← Back
43ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0003-1512-2003ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 40 · 4 first-author · 10 since 2021Security and privacy · 3 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Live Demonstration: AI-based System Latchup Detection and Protection for COTS Systems
abstract
The adoption of Commercial Off-The-Shelf (COTS) ICs in modern satellites faces challenges from radiation-induced latchup events, with current protection methods showing significant limitations in detection accuracy and applicability. This demonstration presents a novel adaptive AI-based latchup detection and protection system featuring two-stage training and LSTM neural network analysis. Our FPGA implementation achieves 90% detection accuracy without extensive pre-characterization. Visitors can validate the system's performance through real-time interaction with various latchup scenarios.
Yin Sun 0005, Junkai Zhao, Rouli Fang, Tony Zhang, Kwen-Siong Chong, Wei Shu, Joseph Sylvester Chang
ISCAS5
2025 N-MUX: Neighborhood-Based Logic Locking Against Machine Learning Attacks
abstract
MUX-based logic locking (LL) is a hardware security technique that inserts multiplexers (MUX) into circuits to secure them against unauthorized use and reverse engineering by protecting original circuit pathways. Nevertheless, MUX-based LL is vulnerable to Oracle-Guided (OG) and Oracle-Less (OL) attacks. While OG methods, such as the SAT attack, are infeasible for large-scale designs, OL attacks, like those based on machine learning (ML), can exploit structural leakage in locked circuits to recover original pathways. This study introduces N-MUX, an innovative MUX-based LL approach designed to resist state-of-the-art (SOTA) ML attacks. N-MUX effectively reduces structural leakage by identifying maximal overlap structures in the original circuit to configure the MUX logic. Additionally, N-MUX ensures high efficiency by selecting the false input from the direct neighbourhood of the true input. Experimental results on ISCAS’85 and ITC’99 benchmarks demonstrate that N-MUX is the most secure and reliable LL technique against SOTA ML-based attacks, achieving an 81% reduction in attack accuracy compared to existing MUX-based LL methods and delivering up to 480× greater efficiency.
Xuenong Hong, Shirui Sheng, Juncheng Chen, Nay Aung Kyaw, Kwen-Siong Chong, Zhiping Lin 0001, Bah-Hwee Gwee
ISCAS6
2025 An Adaptive AI-based Approach to Detect and Protect COTS Systems against Micro-Single-Event-Latchups (μ-SELs) and SELs
abstract
In our envisioned ‘Next Paradigm’ of ‘New Space’, commercial-off-the-shelf (COTS) systems (embodying multiple COTS ICs) would be employed as payloads in space missions. Most COTS ICs are susceptible to radiation effects, particularly Micro-Single-Event-Latchups (μ-SELs) and SELs, and their characteristics are expectedly different. Consequently, hitherto reported detection approaches require characterization of the individual COTS ICs and the entire system, thereby rendering excessive overheads when applied to different COTS systems. In this paper, we propose, for the first time, the design and implementation of an adaptive AI-based approach to detect and protect various uncharacterized COTS systems (vis-à-vis pre-characterized ones) against μ-SELs and SELs. Our proposal involves the adoption of the Long-Short-Term-Memory (LSTM) neural network with our proposed two-stage training process – ex-situ pre-training and in-situ re-training – to improve general applicability. Our FPGA-based prototype achieves high (~90%) average accuracy for four different payloads. This is a worthy improvement of 13.3%-28.5% over reported approaches, yet requiring low (~115 mW) power consumption. Collectively, our proposed approach is appropriate for resource-constrained space applications and our ‘Next Paradigm’ of ‘New Space’.
Junkai Zhao, Yin Sun 0005, Tony Zhang, Kwen-Siong Chong, Wei Shu, Joseph Sylvester Chang
ISCAS4
2024 A Novel Non-profiling Side-Channel Attack on Masked Devices with Connectivity Matrix
abstract
In this paper, we propose a novel pre-processing technique known as the Connectivity Matrix (CM). Building upon the foundation of the CM, we present an effective Second-Order Side-Channel Attack, called Connectivity Matrix Attack (CMA). Our work aims to efficiently counter hardware devices fortified with Masking countermeasures, and it contributes in three significant ways. First, the proposed CM has lower data complexity, as it is constant regarding the number of measurements. Second, we propose the decomposition of the CMs and the utilization of their eigenvalues as feature vectors in CMA. This approach effectively removes noisy components from the CMs and reduces their dimensions. Third, the proposed CMA employs the selected eigenvalues to establish a frequency distribution, followed by a chi-square test. This approach allows CMA to expose both the linear and non-linear leakages present in CMs. The proposed CMA is validated on the public dataset ASCAD and can reveal all the masked bytes successfully. Notably, the concept of the Connectivity Matrix extends beyond the confines of a correlation matrix used in this paper, opening the door to a promising avenue for future research.
Juncheng Chen, Zishuo Yang, Nay Aung Kyaw, Kwen-Siong Chong, Zhiping Lin 0001, Bah-Hwee Gwee
ISCAS6
2024 Securing Against Side-Channel Attacks With Wide-Range In Situ Random Voltage Dithering on Async-Logic AES Engine
abstract
We present a wide-range in situ random voltage dithering (WIS-RVD) on async-logic advanced encryption standard (AES) engine to counteract side-channel attacks (SCAs). There are three contributions in this brief. First, we propose the WIS-RVD based on a dual-rail asynchronous-logic (async-logic) AES engine, leveraging on the self-timed clockless operations for robust encryption under dynamic voltage and timing variations. Second, we propose an in situ voltage dithering to dither the supply voltage instantaneously during the encryption, without the requirement of additional control circuits for clock modulation, to increase the SCA resistance. Third, we propose a wide-range voltage swing technique that spans from 0.3 V (subthreshold) to 1.1 V (above threshold), obfuscating the transistor’s current models between subthreshold and threshold voltage to further enhance SCA resistance. We perform comprehensive SCA evaluations with 50-M power and EM measurements, and the SCA evaluations show that our proposed WIS-RVD on async-logic AES accelerator can resist SCAs with 50-M measurements, i.e.,$\gt 2083\times $and$\gt 2778\times $improvement for power and EM SCAs, respectively, when compared to the standard synchronous-logic AES.
Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Kwen-Siong Chong, Bah-Hwee Gwee
IEEE Trans. Very Large Scale Integr. Syst.4
2023 Improving FPGA-based Async-logic AES Accelerator with the Integration of Sync-logic Block RAMs
abstract
We present a side-channel attack (SCA) resistant asynchronous-logic (async-logic) AES accelerator that integrates synchronous-logic (sync-logic) Block RAMs (BRAMs) in FPGA as the Substitution-Box. We successfully identify the timing requirements to integrate sync-logic BRAMs in our async-logic AES accelerator and validate our proposed AES accelerator on the Sakura-X FPGA board. With the integration of BRAMs, we improve the resource utilization on FPGA by$1.6\times$when compared to the state-of-the-art async-logic AES accelerator, while reducing the power overhead by$1.4\times$. We comprehensively evaluate the SCA resistance of our proposed async-logic AES accelerator with 11 attacking models in both time and frequency domains. Based on our evaluations, we show that our proposed async-logic AES accelerator is highly secure against SCA with 30 million EM traces. This is more than$6000\times$improvement when compared to the benchmark sync-logic AES accelerator and$1.5\times$improvement when compared to the state-of-the-art async-logic AES accelerator.
Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Kwen-Siong Chong, Zhiping Lin 0001, Bah-Hwee Gwee
ISCAS5
2022 Non-profiling based Correlation Optimization Deep Learning Analysis
abstract
Differential Deep Learning Analysis (DDLA) is a deep learning-based non-profiling side-channel attack leveraging neural networks to classify Physical Leakage Information with labels. To avoid the Class Imbalance Problem (CIP) of significantly different data sizes in different data groups, DDLA employs bit labels. However, applying bit labels will be less effective for exploiting leakage. In this paper, we propose to employ Correlation optimization Deep Learning Analysis (CO-DLA) to circumvent the CIP in DDLA by converting the classification in DDLA into a correlation optimization. Bus labels can then be used to exploit stronger leakage information. To validate the attack efficacy improvement, we perform experiments on ASCAD synchronized and de-synchronized masked AES-128 datasets. For the synchronized masked dataset, our proposed CO-DLA requires only 5k traces, which is 75% lesser than the 20k traces required by the reported DDLA, to reveal the key-byte. For the 2 de-synchronized masked datasets, our proposed CO-DLA requires only 10k traces to reveal the key-byte from both of them while the reported DDLA fails to reveal the key-byte.
Juncheng Chen, Jun-Sheng Ng, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Zhiping Lin 0001, Joseph Sylvester Chang, Bah-Hwee Gwee
ISCAS5
2022 An Asynchronous-Logic Masked Advanced Encryption Standard (AES) Accelerator and its Side-Channel Attack Evaluations
abstract
We present a side-channel-attack (SCA) resistant asynchronous-logic (async-logic) Advanced Encryption Standard (AES) accelerator embodying both the masking and hiding SCA countermeasures. Our async-logic masked AES accelerator adopts a dual-rail data encoding to perform the masked 128-bit AES operations, and to enable dual-hiding to moderate both the amplitude (vertical dimension) and the time (horizontal dimension) of the side-channel signals. We implement our async-logic masked AES accelerator in FPGA and comprehensively perform the SCA evaluations based on the electromagnetic (EM) emanation. The SCA evaluations are performed based on bus-wise Hamming Distance model, bus-wise & bit-wise Hamming Weight models, and Zero-Value (ZV) model. Based on our experiment results, we show that our async-logic masked AES is secured against SCA with 1 million EM emanations. This is at least $8.3 \times$ more resistant than synchronous-logic masked AES and $200 \times$ more resistant than the synchronous-logic unmasked AES.
Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Joseph Sylvester Chang, Bah-Hwee Gwee
ISCAS5
2022 A Highly Secure FPGA-Based Dual-Hiding Asynchronous-Logic AES Accelerator Against Side-Channel Attacks
abstract
Encryption in field-programmable gate array (FPGA) often provides a good security solution to protect data privacy in Internet-of-Things systems, but this security solution can be compromised by side-channel attacks (SCAs). In this article, we present an FPGA-based dual-hiding asynchronous-logic (async-logic) advanced encryption standard (AES) accelerator, which is highly resistant against SCAs and yet low area/energy overheads. The proposed AES accelerator achieves vertical (amplitude) SCA hiding via an area-efficient dual-rail mapping approach and a zero-value (ZV) compensated substitution-box (S-Box), while enhancing the horizontal (temporal) SCA hiding of async-logic operations via a timing-boundary-free input arrival-time randomizer and a skewed-delay controller. A comprehensive SCA evaluation is performed with 11 SCA models, and we show that our proposed design can offer a strong SCA resistance with measurement-to-disclosure (MTD) of >20 million traces. To our best knowledge, our design is the most secure AES design evaluated with the largest number of traces in FPGA. To compare the design overheads for security, we quantify the figure of merit as normalized (Area$\times $Energy/MTD(All)$\times 10^{6}$). The figure of merit of our proposed design is$403\times $smaller than the benchmark dual-rail synchronous-logic design and$95\times $smaller than a reported async-logic design.
Jun-Sheng Ng, Juncheng Chen, Kwen-Siong Chong, Joseph Sylvester Chang, Bah-Hwee Gwee
IEEE Trans. Very Large Scale Integr. Syst.3
2021 Normalized Differential Power Analysis - for Ghost Peaks Mitigation
abstract
The attack efficacy of Differential Power Analysis (DPA), a popular side channel evaluation technique for key extraction, is compromised by the false highest Difference Of Means (DOMs) value ('ghost peaks') in the DOMs matrix produced in a conventional DPA. The ghost peak is generated by the wrong key guess and always occurs in the conventional DPA when the number of side channel traces is not enough. In this paper, an improved version of the conventional DPA termed as Normalized DPA (NDPA) is proposed to circumvent the ghost peak. With the analysis on the generation of ghost peaks in the conventional DPA, we observed that by normalizing the DOMs matrix, the ghost peaks can be greatly suppressed. We model the proposed NDPA mathematically and show that it performs better than the conventional DPA. We further provide the experimental validations on a set of 200k power simulation traces on AES S- Box and 500 EM traces from ASCAD dataset. Based on the attack results of these datasets, our proposed NDPA requires (up to 68%) lesser number of traces to reveal a correct key when compared to the conventional DPA.
Juncheng Chen, Jun-Sheng Ng, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Weng-Geng Ho, Kwen-Siong Chong, Zhiping Lin 0001, Joseph Sylvester Chang, Bah-Hwee Gwee
ISCAS6
2021 A Novel Normalized Variance-Based Differential Power Analysis Against Masking Countermeasures
abstract
In this paper, we propose two normalization techniques to reduce the ghost peaks occurring in Differential Power Analysis (DPA). Ghost peaks can be defined as the DPA output generated by the wrong key guesses, having higher amplitudes than the DPA output generated by the correct key guess. We further propose variance-based Differential Power Analysis (vDPA) to attack masked crypto devices. The proposed normalization techniques and vDPA constitute four contributions. First, based on the side-channel signal modeling with the linear coefficient representing the strength of the linear component in a side-channel signal, we formulate the condition function of linear coefficients for the appearance of ghost peaks in DPA. Second, we propose pre-normalization in DPA and mathematically analyze how it can reduce ghost peaks by modulating the strength of the linear components in side-channel signals. Third, we propose post-normalization and mathematically analyze how it can reduce ghost peaks by de-correlating the strength of the linear components in side-channel signals with the condition function for the appearance of ghost peaks. Fourth, we propose vDPA to apply simultaneously with either one of the proposed normalization techniques to effectively attack masked crypto devices. Based on the experiments, we show that the proposed basic vDPA (without normalization), pre-normalized vDPA and post-normalized vDPA are all able to reveal the secret key from ASCAD data set. The pre- and post-normalized vDPAs require up to 18× and 14× fewer traces than the basic vDPA respectively. While attacking ASCAD data set, the proposed pre- and post-normalized vDPAs are both 13, 095× faster than the reported 2nd order CPA, and reveal the key-bytes successfully with only half of side-channel traces required by the reported Zero-offset DPA.
Juncheng Chen, Jun-Sheng Ng, Kwen-Siong Chong, Zhiping Lin 0001, Bah-Hwee Gwee
IEEE Trans. Inf. Forensics Secur.3
2020 Radiation-Hardened-by-Design (RHBD) Digital Design Approaches: A Case Study on an 8051 Microcontroller
abstract
Advanced satellites and/or high-level (levels 4 and 5) autonomous vehicles demand high reliability integrated circuits (ICs) with ultra-low error rates. One solution is to use radiation-hardened-by-design (RHBD) design techniques to mitigate the error rates against the single-event-effects (arising from radiation effects). This paper first provides an overview on several present-art RHBD design techniques, and then propose an RHBD design methodology, spanning from the library cell development, circuit simulation and synthesis, to the layout implementation, to realize digital circuits. We further demonstrate an 8051 microcontroller with the proposed design methodology, and evaluate the 8051 microcontroller prototype (@ 65nm CMOS) with irradiation tests. Our 8051 microcontroller is error-free with 10 MeV.mg/cm2, meeting our targeted specifications for Low Earth Orbit applications. When under high Linear Transfer Energy (> 51.5 MeV.mg/cm2) tests, the 8051 microcontroller does suffer errors. We further study/analyze which part of the 8051 microcontroller to cause errors, and provide recommendations.
Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Wei Shu, Joseph Sylvester Chang
ISCAS1
2020 A Secure Data-Toggling SRAM for Confidential Data Protection
abstract
We study the security feature of static random access memory (SRAM) against the data imprinting attack and provide a solution to protect the SRAM from this attack. There are four main contributions in this paper. First, the negative-bias temperature-instability (NBTI) degradation of PMOS transistors in the conventional SRAM cell that causes the data imprinting effect is explained. Second, the data imprinting effect that leaks the stored information in the conventional SRAM cell is investigated. Third, a novel low transistor-count transmission-gate-based master-slave SRAM cell is proposed to periodically toggle the stored data for reducing the data imprinting effect. Fourth, an efficient imprinting analysis flow is proposed to evaluate the proposed data-toggling SRAM for quantifying the data imprinting effect. Based on a 65-nm CMOS process, we implement and prototype the proposed 1k-byte data-toggling SRAM design. We perform our imprinting analysis flow on various SRAM ICs and benchmark our proposed data-toggling SRAM IC against the non-toggling SRAM IC and a commercial Lyontek SRAM IC. From the measurement results, the non-toggling SRAM and Lyontek SRAM suffer from 60% and 81% data imprinting effects, respectively, whereas our data-toggling SRAM has only 11% data imprinting effect (at 160-kHz toggling frequency). The data-toggling SRAM could switch between high security (<; 5% data imprinting effect) high power mode for hardware security applications and low power (<; 0.1mW) low security mode for power-saving applications. Particularly, our data-toggling SRAM could feature as low as ~1% data imprinting effect when increasing the toggling frequency to 1.6 MHz by compromising the power dissipation. Using the image analysis flow, the stored information is revealed in both the non-toggling and Lyontek SRAM ICs but is well protected in the proposed data-toggling SRAM IC.
Weng-Geng Ho, Kwen-Siong Chong, Tony Tae-Hyoung Kim, Bah-Hwee Gwee
ISCAS2
2020 A DPA-Resistant Asynchronous-Logic NoC Router with Dual-Supply-Voltage-Scaling for Multicore Cryptographic Applications
abstract
We propose a 5-port asynchronous-logic Network-on-Chip (ANoC) router based on the Sense-Amplifier Half-Buffer (SAHB) approach for cryptographic processing cores to counteract side channel attack differential power analysis (DPA) in multicore platform. There are three features in the proposed DPA-resistant ANoC router. First, the proposed ANoC router embodies dual-supply-voltage SAHB cells, where the non-critical subsidiary supply voltage is adjustable from 0.3V to 1.2V, increasing the noise variance and hence reducing the Signal-to-Noise (SNR) ratio to hide the information leakage. Second, the proposed ANoC router performs as a noise engine by increasing the number of power-on IO ports, further randomizing the overall power dissipation. Third, the proposed ANoC router can switch between DPA-resistant mode and energy-efficient nominal (non-secure) mode, saving the power dissipation when the DPA secure countermeasure is unnecessary. Based on 65nm CMOS process, the multicore platform embedded with the proposed ANoC router is implemented, and the experiment is demonstrated by running the advanced encryption standard (AES) cryptography operation. When benchmarked against the nominal mode, the noise power variance of the proposed ANoC router increases by 2.3× in the DPA-resistant mode, reducing the overall SNR ratio by 56%. When comparing to other reported noise engines, our proposed ANoC router is one of the most DPA-secure, area-efficient and power-efficient designs for multicore cryptographic applications.
Weng-Geng Ho, Ne Kyaw Zwa Lwin, Nay Aung Kyaw, Jun-Sheng Ng, Juncheng Chen, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS6
2020 A Highly Efficient Power Model for Correlation Power Analysis (CPA) of Pipelined Advanced Encryption Standard (AES)
abstract
We evaluate the vulnerability of a pipelined Advanced Encryption Standard (AES) against Correlation Power Analysis (CPA) Side-Channel Attack (SCA). We identify that the registers in pipelined AES are most vulnerable against CPA SCA and propose a new power model targeting the switching activities of the registers. The proposed power model is constructed based on the Hamming Distance (HD) between the intermediate values stored in the registers in two consecutive clock cycles. Then, we analyze the vulnerability of pipelined AES under two scenarios. First, during regular pipeline operation where the device is performing AES pipeline operation. Second, in non-pipeline operation where we assume the adversaries can insert delay to the input of the device to increase the signal to noise ratio of the physical leakage information. The simulation results show that under regular pipelined operation, our proposed power model can reveal all the 16 key bytes in less than 4,900 traces, resulting in 4.7× more effective than the conventional power models. Under non-pipelined operation, our proposed power model requires only 590 traces to reveal all the 16 key bytes, which is 5.9× more effective than other power models.
Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee
ISCAS6
2019 Low Gate-Count Ultra-Small Area Nano Advanced Encryption Standard (AES) Design
abstract
We present a low gate-count ultra-small area nano advanced encryption standard (AES) design. We achieve the low gate-count by the following means. First, we repeatedly reuse the area-critical circuits, i.e. one 8-bit Substitute-Box (S-Box) circuit and one 32-bit MixColumn circuit, for AES. Second, we cascade the input flip-flops (FFs) with our data transfer architecture so that the outputs of the MixColumn circuit are connected directly to the first 32-bit input FFs without extra multiplexing circuits. Third, the ShiftRow operation is implicitly performed by assigning the data sequence to the input FFs (during the S-Box and MixColumn operations). Fourth, we use independent XOR gates for AddRound and KeyExpansion operations. The collective means enables our design to feature 1457 gates, and to occupy 100um×100um area @ 65nm CMOS. When compared to the normalized area (@ 65nm CMOS) of the reported AES designs, our design features the smallest normalized area, 10% smaller than the most competitive reported AES design. Our design is targeted for ultra-small area applications including biomedical applications.
Aparna Shreedhar, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Nay Aung Kyaw, L. Nalangilli, Wei Shu, Joseph Sylvester Chang, Bah-Hwee Gwee
ISCAS2
2019 A Highly Efficient Side Channel Attack with Profiling through Relevance-Learning on Physical Leakage Information
abstract
We propose a Profiling through Relevance-Learning (PRL) technique on Physical Leakage Information (PLI) to extract highly correlated PLI with processed data, as to achieve a highly efficient yet robust Side Channel Attack (SCA). There are four key features in our proposed PRL. First, variance analysis on PLI is implemented to determine the boundary of the clusters and objects of the clusters. Second, the nearest-neighbor k-NN variance clustering is used to reduce the sampling points of PLI by clustering the high variance sampling points and discarding the low variance sampling points of PLI measurements (traces). These clustered sampling points, which are highly correlated with the processed data, contain pertinent leakage information related to the secret key. Third, the information associated with the secret key is spread in several neighboring sampling points with different degrees of leakages. We analytically derive the Key-leakage relevance factor for each clustered sampling point to quantify the degree of leakage associated with the secret key. Fourth, by means of Hebbian learning, a weight proportional to the Key-leakage relevance factor is updated iteratively based on the values of relevance factor and traces of the sampling points. The converged weights which are being assigned to clustered sampling points are linked to their associated PLI to further increase the correlation of the PLI with the processed data. Therefore, the required number of PLI measurements, to reveal the secret key, can be reduced significantly. In addition, we analytically show that the computational complexity of our proposed PRL is O(n) when compared to the reported profiling techniques having O(n2) and O(n3) computational complexities. Based on the experiments of our proposed PRL performed on the PLI of AES-128 algorithm, the results depicting that the sampling points of PLI are reduced 87 percent after the k-NN variance clustering. The converged weight with learning error rate106traces, our proposed PRL is ~2; 000x more efficient in performing SCA.
Ali Akbar Pammu, Kwen-Siong Chong, Yi Estelle Wang, Bah-Hwee Gwee
IEEE Trans. Dependable Secur. Comput.2
2019 A High Throughput and Secure Authentication-Encryption AES-CCM Algorithm on Asynchronous Multicore Processor
abstract
We propose an authentication-based matrix-transformation cum parallel-encryption implemented on an asynchronous multicore processor (AMP-MP) to achieve a high throughput and yet secure advanced encryption standard based on counter with chaining mode (AES-CCM). There are four main features in our proposed AMP-MP. First, we employ the matrix multiplication in GF(28) computation to transform the 16 plaintexts into one plaintext, hence improving the authentication speed by 32× collectively at the transmitter and receiver. Second, we reschedule the operations of three AES encryptions in three different cores such that their physical leakages are compensated and equalized, thus reducing the correlation of physical leakage with the processed data by >3×. Third, the intermediate values of AES-CCM are propagated asynchronously between different cores to randomize the physical leakages with the processed data, and therefore further enhance the security of AES-CCM against the SCA by another 3×. Fourth, we propose a key adjusting technique based on S-Box byte-key transformation to protect the key against pattern-based attack. Our proposed AMP-MP is realized on an 8-bit asynchronous 9-core processor fabricated based on the 65 nm CMOS process. The experimental results show that the throughput of the authentication is 13.54 Gbps while the throughput for both authentication and encryption collectively is 8.32 Gbps, which are 17× and 70× faster than the reported counterparty, respectively. Based on power dissipation and EM SCA on our proposed AMP-MP, the secret key is unrevealed at 5 × 105 traces, which is ~17× more secured than the standard ASIC AES-CCM implementation.
Ali Akbar Pammu, Weng-Geng Ho, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Bah-Hwee Gwee
IEEE Trans. Inf. Forensics Secur.4
2018 Asynchronous-Logic QDI Quad-Rail Sense-Amplifier Half-Buffer Approach for NoC Router Design
abstract
We propose a low area overhead and power-efficient asynchronous-logic quasi-delay-insensitive (QDI) sense-amplifier half-buffer (SAHB) approach with quad-rail (i.e., 1-of-4) data encoding. The proposed quad-rail SAHB approach is targeted for area- and energy-efficient asynchronous network-on-chip (ANoC) router designs. There are three main features in the proposed quad-rail SAHB approach. First, the quad-rail SAHB is designed to use four wires for selecting four ANoC router directions, hence reducing the number of transistors and area overhead. Second, the quad-rail SAHB switches only one out of four wires for 2-bit data propagation, hence reducing the number of transistor switchings and dynamic power dissipation. Third, the quad-rail SAHB abides by QDI rules, hence the designed ANoC router features high operational robustness toward process-voltage-temperature (PVT) variations. Based on the 65-nm CMOS process, we use the proposed quad-rail SAHB to implement and prototype an 18-bit ANoC router design. When benchmarked against the dual-rail counterpart, the proposed quad-rail SAHB ANoC router features 32% smaller area and dissipates 50% lower energy under the same excellent operational robustness toward PVT variations. When compared to the other reported ANoC routers, our proposed quad-rail SAHB ANoC router is one of the high operational robustness, smallest area, and most energy-efficient designs.
Weng-Geng Ho, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Bah-Hwee Gwee, Joseph Sylvester Chang
IEEE Trans. Very Large Scale Integr. Syst.2
2017 DPA-resistant QDI dual-rail AES S-Box based on power-balanced weak-conditioned half-buffer
abstract
We propose an asynchronous-logic (async) Quasi-Delay-lnsensitive (QDI) dual-rail 32-bit Advanced Encryption Standard (AES) Substitution-Box (S-Box) for Differential Power Analysis (DPA) attack countermeasure. There are three novel features in the proposed S-Box. First, the proposed S-Box operates in async QDl protocol with dual-rail data encoding, hence there is only a marginal difference in power dissipation for different signal output transitions. Second, the proposed S-Box embodies the power-balanced async Weak-Conditioned Half-Buffer (WCHB) cell approach, which features the same number of transitions, and hence same number of switching for different input combinations to equalize the power dissipation. Third, the proposed S-Box embodies our novel-designed library cells in which each output wire, corresponding to different transitions, has a similar capacitive load, hence hiding the dynamic power dissipation. Based on the 65nm CMOS process, we implement the proposed 32-bit AES S-Box, and benchmark it against the conventional synchronous-logic (sync) S-Box and the reported async (i.e. unbalanced) WCHB S-Box. From the simulation results, our proposed power-balanced WCHB S-Box features significantly lower 1.21% and 0.55% of Normalized Energy Deviation (NED) and Normalized Standard Deviation (NSD) respectively. Particularly, our proposed design is of 57.6× and 17.9× lower NED, and 55.1× and 9.29× lower NSD than the reported sync and async counterparts respectively.
James Lim, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee
ISCAS3
2017 Highly secured state-shift local clock circuit to countermeasure against side channel attack
abstract
We propose a highly-secured State-shift Local Clock (SsLC) countermeasure technique to hide the Physical Leakage Information (PLI) against Side Channel Attack (SCA). The SCA is a technique employed to reveal the secret key of cryptographic algorithm by correlating the PLI (i.e. power dissipation and Electromagnetic (EM)) with the processed data, where both the PLI and processed data are generated during the encryption process. Whereas the countermeasure technique aims to reduce the correlation of the PLI against the processed data. There are four key features in our proposed SsLC countermeasure technique. First, it embodies a finite state machine which can be employed to regularly shift the timing operation of cryptographic algorithm implementations. Thus, the correlation of the PLI with the processed data is significantly reduced due to dynamically changes the occurrences of encryption operation in time domain. Second, the PLI which encompasses a secret key is spread over in time domain to reduce the probability of revealing the secret key. Third, the power dissipation overhead is negligible and hence it is highly applicable for low power applications. Fourth, the regular state (time) shifting technique in the SsLC is able to hide multiple PLIs, i.e. power dissipation and EM signals, concurrently. In view of the above features, the proposed SsLC is highly secured against SCA with multiple PLIs. Based on the experimental results in FPGA, our proposed SsLC countermeasure technique features wide distribution of PLI in time domain, dissipates 2.77mW of power and emits 12.2mV/m of EM signal @ 2.4MHz. Furthermore, with 106 power dissipation and EM measurements, the secret key of the cryptographic algorithm remains unbreakable. In comparison with the reported counterparts, the resistance of our proposed SsLC against SCA is significantly improved as the number of power dissipation and EM traces to reveal the secret key has increased by >18x and >25x respectively. Consequently, the correlation coefficient between the PLI and the processed data is reduced by 3.5x.
Ali Akbar Pammu, Kwen-Siong Chong, Bah-Hwee Gwee
ISCAS2
2017 Sense Amplifier Half-Buffer (SAHB) A Low-Power High-Performance Asynchronous Logic QDI Cell Template
abstract
We propose a novel asynchronous logic (async) quasi-delay-insensitive (QDI) sense-amplifier half-buffer (SAHB) cell design approach, with emphases on high operational robustness, high speed, and low power dissipation. There are five key features of our proposed SAHB. First, the SAHB cell embodies the async QDI 4-phase (4φ) signaling protocol to accommodate process-voltage-temperature variations. Second, the sense amplifier (SA) block in SAHB cells embodies a cross-coupled latch with a positive feedback mechanism to speed up the output evaluation. Third, the evaluation block in the SAHB comprises both nMOS pull-up and pull-down networks with minimum transistor sizing to reduce the parasitic capacitance. Fourth, both the evaluation block and SA block are tightly coupled to reduce redundant internal switching nodes. Fifth, the SAHB cell is designed in CMOS static logic and hence appropriate for full-range dynamic voltage scaling operation for VDDranging from nominal voltage (1 V) to subthreshold voltage (~0.3 V). When six library cells embodying our proposed SAHB are compared with those embodying the conventional async QDI precharged half-buffer (PCHB) approach, the proposed SAHB cells collectively feature simultaneous -.64% lower power, -.21% faster, and ~6% smaller IC area; the PCHB cell is inappropriate for subthreshold operation. A prototype 64-bit Kogge-Stone pipeline adder based on the SAHB approach (at 65 nm CMOS) is designed. For a 1-GHz throughput and at nominal VDD, the design based on the SAHB approach simultaneously features -.56% lower energy and -.24% lower transistor count advantages than its PCHB counterpart. When benchmarked against the ubiquitous synchronous logic counterpart, our SAHB dissipates -.39% lower energy at the 1-GHz throughput.
Kwen-Siong Chong, Weng-Geng Ho, Tong Lin 0001, Bah-Hwee Gwee, Joseph Sylvester Chang
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Low normalized energy derivation asynchronous circuit synthesis flow through fork-join slack matching for cryptographic applications
Nan Liu 0002, Kwen-Siong Chong, Weng-Geng Ho, Bah-Hwee Gwee, Joseph Sylvester Chang
DATE2
2016 High performance low overhead template-based Cell-Interleave Pipeline (TCIP) for asynchronous-logic QDI circuits
abstract
We propose a novel Template-based Cell-Interleave Pipeline (TCIP) approach for generating high performance and yet low overhead asynchronous-logic (async) quasi-delay-insensitive (QDI) circuits. Our TCIP approach exploits the characteristics of the four prevalent QDI cell templates, namely Weak-Conditioned Half-Buffer (WCHB), Pre-Charged HalfBuffer (PCHB), Autonomous Signal-Validity Half-Buffer (ASVHB), and Sense-Amplifier Half-Buffer (SAHB), and then strategically interleave these template cells to form a composite pipeline. There are three main features in our TCIP approach. First, all QDI cell templates are first standardized with the same interface signals, and their corresponding cells are characterized in terms of transistor count, cycle time and energy dissipation for ease of comparison/selection/replacement. Second, our TCIP approach prioritizes the speed requirement when forming the initial pipeline circuits, and then subsequently reduces circuit overheads by interleaving various template cells without compromising the speed significantly. Third, the final optimized QDI pipeline circuit inherently features high robustness against process-voltage-temperature (PVT) variations, hence suitable for dynamic-voltage-scaling (DVS) operation. By means of 65nm CMOS process, we demonstrate a 4-bit pipeline tree adder based on the proposed TCIP approach, and benchmark it against the WCHB, PCHB, ASVHB and SAHB counterparts. These five designs feature same high operational robustness, nonetheless the design based on our TCIP approach is more competitive. Particularly, the designs based on reported approaches are, on average, ∼1.22× more transistor count, ∼1.21× slower and ∼1.22× higher energy dissipation. Furthermore, under DVS operation from 1.2V to 0.3V, our proposed TCIP adder can reduce up to 88% energy for non-speed critical applications.
Weng-Geng Ho, Nan Liu 0002, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS4
2016 Area-efficient and low stand-by power 1k-byte transmission-gate-based non-imprinting high-speed erase (TNIHE) SRAM
abstract
We propose a novel 15-T Transmission-gate-based Non-Imprinting High-speed Erase (TNIHE) SRAM cell with emphases on low area overhead and low stand-by power attributes for highly secured data storage applications. We benchmark our proposed 15-T TNIHE SRAM cell against the reported 22-T Non-Imprinting High-speed Erase (NIHE) SRAM cell, and demonstrated three key features of reducing 7 transistors. First, we adopt the transmission gates (as opposed the cross-couple inverters) in the slave circuitry, saving 4 transistors. Second, we eliminate a transistor which uses to reset the slave circuitry, hence saving 1 transistor. Third, we apply the global inverse transistors (as opposed to the local inverse transistors) in the read /write circuit for each SRAM cell, hence further reduce 2 more transistors. As a result, our proposed TNIHE SRAM cell @ 65nm CMOS features ~17% smaller layout area. We design a 1k-byte memory based on the proposed TNIHE SRAM cells. On the basis of simulations, we show that our 1k-byte SRAM memory features overall ~13% smaller area, and dissipates on average, ~30% lower stand-by power than the reported NIHE counterpart.
Weng-Geng Ho, Ne Kyaw Zwa Lwin, N. Prashanth Srinivas, Kwen-Siong Chong, Tony Tae-Hyoung Kim, Bah-Hwee Gwee
ISCAS4
2016 Total Ionizing Dose (TID) effects on finger transistors in a 65nm CMOS process
abstract
Although Total Ionizing Dose (TID) effects are generally unpronounced in deep-submicron-CMOS, we show the TID-induced leakage current @TID=500Krad is significant in NMOS-finger-transistors of GlobalFoundries 65nm CMOS. Further, Radiation-Hardening-By-Design techniques against said TID effect are recommended.
Jize Jiang, Wei Shu, Kwen-Siong Chong, Tong Lin 0001, Ne Kyaw Zwa Lwin, Joseph Sylvester Chang
ISCAS3
2016 Experimental investigation into radiation-hardening-by-design (RHBD) flip-flop designs in a 65nm CMOS process
abstract
We comprehensively study three types of radiation-hardened flip-flops: DICE for SEU-hardening, temporal for SET-hardening, and Triple-Modular-Redundancy for SEU-cum-SET-hardening. Our study includes their trade-offs of circuit/radiation-hardness attributes. We find that DICE flip-flops remain the most competitive.
Tong Lin 0001, Kwen-Siong Chong, Wei Shu, Ne Kyaw Zwa Lwin, Jize Jiang, Joseph Sylvester Chang
ISCAS2
2016 Secured Low Power Overhead Compensator Look-Up-Table (LUT) Substitution Box (S-Box) Architecture
abstract
Substitution-Box (S-Box) is an important security building block for the Advanced Encryption Standard (AES) algorithm. However, its high power dissipation always compromises with its security feature under Correlation Power Analysis (CPA) attack. In this paper, we propose a secured and low power overhead LUT based S-Box architecture embodying a novel multiplexing circuit AND and OR a compensator. We achieve these attributes as follows. First, we employ AND and OR gates to realize the multiplexing circuit therein in a regular structure to minimize the delay and power variations for every input pattern, hence mitigating the security risk against CPA. Second, we augment a compensator to complement the multiplexing circuit to further minimize the power variations within the LUT based S-Box. We realize six AES designs based on the Sakura-X FPGA board, three designs embodying reported S-Box architectures and the other three designs leveraging on our multiplexing circuit and compensator. We show that our AES design, embodying our LUT based S-Box architecture with the AND/OR-gate multiplexing circuit and compensator, has the highest security feature (against CPA) compared with the reported designs, featuring 10× to 300× better security.
Ali Akbar Pammu, Kwen-Siong Chong, Bah-Hwee Gwee
NAS2
2015 High robustness energy- and area-efficient dynamic-voltage-scaling 4-phase 4-rail asynchronous-logic Network-on-Chip (ANoC)
abstract
We propose an 18-bit 5-interface asynchronous-logic Network-on-Chip (ANoC) router based on the quasi-delay-insensitive (QDI) realization approach for high secured cryptography applications. There are four key features of the proposed ANoC router. First, it embodies the novel high-speed low-power Sense-Amplifier Half Buffer 4-rail cells. Second, it is designed based on QDI protocol, and hence is highly robust against process-voltage-temperature (PVT) variations. Third, it is functional for full dynamic voltage scaling from nominal (VDD=1.2V) to sub-threshold (VDD=0.3V) regions, and is potentially excellent for low power management applications. Fourth, it embodies a distributed-based XY routing algorithm to utilize a 4-bit header of flow control unit (flit) for routing up to 4×4 cluster, hence minimizing the routing overhead. We realize the proposed ANoC router (@65nm CMOS), and benchmark it against the reported ANoC router embodying the conventional Weak-Conditioned Half-Buffer (WCHB) QDI realization approach. Both our proposed and reported designs feature the high operation robustness, but our design is 41% more energy-efficient, and 21% more area-efficient than the reported counterpart. The prototype of ANoC router occupies only 0.105 mm2and can operate down to 0.3V. At VDD=0.3V, it dissipates 44 fJ per bit and operate 105 ns per flit.
Weng-Geng Ho, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS2
2015 A single-VDD half-clock-tolerant fine-grained dynamic voltage scaling pipeline
abstract
We propose a novel dynamic voltage scaling (DVS) pipeline with three significant attributes. First, it features a fine-grained DVS which innately attempts to power most of the circuits therein at low voltages, and when the speed is beneath the requirement, to scale up the voltage. Second, it supports fast-transition DVS within one-and-a-half clock duration per operation, and its operation remains error-free during that duration; we define such attribute as half-clock-tolerant. Third, it consists of a single power source (single-VDD) which supports three voltage scales (1.2V, 0.8V and 0.5V) for power/speed tradeoffs, and has standardized 1.2V output to seamlessly interface with other proposed/conventional pipelines. These attributes are achieved due to the embodiment of a DVS power unit, asynchronous building blocks to control/synchronize the operation, a dual-rail critical path to innately detect the completion of the operation, and level shifters to standardize the output voltage. We demonstrate our proposed pipeline by designing a multiplier embodied in a Fast Fourier Transform processor (@65nm CMOS). We show that the multiplier based on our proposed pipeline, on average, is 1.94× more power-efficient than that based on a conventional pipeline.
Kwen-Siong Chong, Tong Lin 0001, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS2
2014 Synthesis of asynchronous QDI circuits using synchronous coding specifications
abstract
We propose a synthesis of asynchronous quasi-delay-insensitive (QDI) circuits. We highlight three notably features/novelties of the proposed synthesis as follows. First, the targeted synthesized circuits abide by the QDI protocol; hence they are inherently timing-robust and are desirable for applications with high variation-space and wide operation-space (including defense/space applications). Second, the coding specifications accept Verilog HDL language, and are the same/similar to the standard coding for synchronous circuits, hence no special and/or ad-hoc design/coding rules are required. Third, the proposed synthesis is applicable to accept various QDI library cells, hence enabling to explore full merit of different library cells. To the best of our knowledge, no reported synthesis methods incorporate all these features; some limited features were only incorporated. Our proposed synthesis, at this juncture, accepts three basic clauses - complete `if-else' clause, incomplete `if-else clause', and the `case' clause. These clauses are more than sufficient to describe any complex systems. The synthesis stages involve analyzing QDI pipelines, generating (corresponding) single-rail combinational circuits, converting dual-rail netlists (from the single-rail circuits), and embedding customized controllers. In order to demonstrate the validity and practicality of the proposed synthesis, an 8-bit 8-tap asynchronous QDI Finite Impulse Response (FIR) filter is synthesized, implemented to the layout stage, and evaluated using spice models-specifically, it features 3.7 mW power dissipation, 39,181 transistors, and a delay of 200 ns per operation.
Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang, Weng-Geng Ho
ISCAS2
2014 A Low Overhead Quasi-Delay-Insensitive (QDI) Asynchronous Data Path Synthesis Based on Microcell-Interleaving Genetic Algorithm (MIGA)
abstract
In this paper, we propose a design approach to mitigate the hardware overhead of the data completion detection circuit in quasi-delay-insensitive (QDI) asynchronous-logic circuits. In this proposed design approach, three novelties are highlighted. Firstly, a novel microcell-interleaving approach is proposed to reduce the number of completion detection (CD) circuits while retaining the required QDI attribute. Secondly, we analyze the performance of the QDI circuits based on the proposed microcell-interleaving approach graphically in terms of power dissipation, transistor count and delay, and evaluate/determine the upper and lower boundaries of these performance profiles. Thirdly, we propose a microcell-interleaving genetic algorithm (MIGA) to stochastically optimize the proposed microcell-interleaving approach on power dissipation, transistor count, and delay. To validate the proposed design approach, a complete performance profile of ISCAS-85 C499 circuit is investigated on the basis of differential cascode voltage switch logic (DCVSL) and dynamic strong indicating (DSI) microcells. We demonstrate the efficiency of the proposed design approach by benchmarking against the competing DCVSL, null convention logic and DSI designs on five ISCAS-85 circuits. Specifically, the proposed designs, on average, are 1.77 × better in power dissipation, 1.4 × better in area, and 1.58 × better in a composite metric of power × area × delay, and reasonably slower for the lowest power dissipation points. We further demonstrate the practicality of the proposed design approach by implementing an 8-tap 16-bit asynchronous QDI finite impulse response filter. Finally, we demonstrate the ~10% and ~11% improved efficiency of the proposed MIGA over the greedy algorithm and dynamic programming, respectively.
Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2013 A 250mV sub-threshold asynchronous 8051microcontroller with a novel 16T SRAM cell for improved reliability in 40nm CMOS
abstract
Asynchronous approach for digital systems is a way to resolve increased timing uncertainty with technology scaling since timing issue is eliminated in asynchronous systems. This paper presents a sub-threshold operating asynchronous 8051 microcontroller (A8051) with a novel 16T SRAM cell for improved reliability in asynchronous systems. This A8051, adopting a 4-phase dual-rail protocol, can operate up to 250 mV. A8051 has 67.53 μs as a critical path delay with 91.6 nW power consumption at 250 mV, which is equivalent to 12.88 kHz in synchronous systems. At 1.0 V, the delay of a critical path of A8051 microcontroller is 5.74 ns, which is equivalent to 151.55 MHz, with 8.98 mW power consumption. The proposed 16T SRAM cell is applied in memory blocks. The 16T SRAM structure eliminates charge contentions between devices during read and write operations so that SRAM can be operated fully in static mode, bringing about improved write margin (WM). The WM of this 16T SRAM cell is 1.81 times greater than the conventional 6T SRAM cell and 1.58 times better than 8T SRAM cell. At 250 mV, the SNM of SRAM cell is 12.5 mV under process and mismatch variations. Write delay of the asynchronous SRAM block is 4.02 μs (equivalent to 248.5 kHz) with 5.44 pJ energy dissipation, while read delay is 12.61 μs (equivalent to 79.3 kHz) with 9.08 pJ energy dissipation.
Kwen-Siong Chong, Joseph Sylvester Chang, Pinaki Mazumder
ACM Great Lakes Symposium on VLSI2
2013 A dual-core 8051 microcontroller system based on synchronous-logic and asynchronous-logic
abstract
We describe a dual-core 8051 microcontroller system featuring the synchronous and asynchronous (clockless) mode of operation. The synchronous mode of operation is achieved by means of a synchronous 8051 microcontroller core, while the asynchronous mode of operation is achieved by means of an asynchronous 8051 microcontroller core. The 8051 microcontroller system features shared embedded program and data memories that enable the switching between the two microcontroller cores during program execution. The measured energy, speed and electromagnetic interference of both microcontroller cores will be compared at different operation workloads.
Kok-Leong Chang, Tong Lin 0001, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS4
2013 Low power sub-threshold asynchronous QDI Static Logic Transistor-level Implementation (SLTI) 32-bit ALU
abstract
We propose an asynchronous-logic (async) Quasi-Delay-Insensitive (QDI) Static Logic Transistor-level Implementation (SLTI) approach for low power sub-threshold operation. The approach is implemented to design 32-bit pipelined Arithmetic and Logic Units (ALUs), the primary computation core for microprocessors, and benchmarked against the reported Pre-Charged Half-Buffer (PCHB). There are two key attributes in this proposed design. First, the proposed SLTI ALU design can perform dynamic voltage scaling seamless by only changing the supply voltage from nominal (1V) to sub-threshold (~0.2V) regions for high speed/low power operation. Second, the ALU achieves ultra-low power dissipation (3.5μW) at the lowest VDDpoint (~0.15V). For fair of comparison, both implemented ALUs have identical functionality and functional blocks, are implemented using the same 65nm CMOS process. Based on the simulations, the minimum energy point occurs at VDD= 0.2V for SLTI-based ALU and at VDD= 0.3V for PCHB-based ALU. The SLTI-based ALU have ~93% and ~89% lower energy on the arithmetic and logic operations respectively from VDD= 1V to VDD= 0.2V. At VDD= 0.2V, with 9MHz input switching rate, the async ALU based on our proposed SLTI approach dissipates ~51% and ~44% lower power than the reported PCHB counterpart on the arithmetic and logic operations respectively.
Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS2
2012 A comparative study on asynchronous Quasi-Delay-Insensitive templates
abstract
The robustness of asynchronous logic has proved useful in dealing with contemporary problems in CMOS design such as process variations and power management. However, the general cryptic nature of asynchronous logic has stymied the widespread acceptance of this alternate design technique. Fortunately, the semi-custom approach to asynchronous design reduces the tedious handcrafting efforts that are often non-trivial in large system-on-chips (SoCs). However, even with the adoption of this design approach requires careful selection of asynchronous templates that will suit overall system needs. Therefore in this paper, the most eminent Quasi-Delay-Insensitive asynchronous template families reported to date will be presented, and followed by an in-depth comparison of various design FOMs - template area, static/dynamic capacity, cycle time, latency, throughput and Et2. The most aggressive template (EESTFB) can reach a maximum throughput of 3.56Giga items/s on 0.13µm @ 1.2V.
Kok-Leong Chang, Tong Lin 0001, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS4
2012 An Ultra-Dynamic Voltage Scalable (U-DVS) 10T SRAM with bit-interleaving capability
abstract
We propose a dynamic voltage scalable SRAM capable of efficient bit-interleaving in column to tolerate multiple-bits soft error when integrated with error correction codes (ECC). First, a 10T SRAM bitcell is proposed. It activates only intended bitcells so that stability problem of half-selected bitcells is completely eliminated and the power dissipation in half-selected columns is significantly reduced. Second, a configurable DVS scheme is employed to enable the bitcell to operate like differential 8T during super-threshold region which results in faster operation. The proposed SRAM can operate up to 1.2GHz at 1.2V using 65nm CMOS process. Third, a segmented column multiplex with low overhead is proposed, which greatly reduces the power dissipation due to the column control signals. Consequently, the write and read power dissipations are reduced by up to 40% and 67% respectively. Forth, a hierarchical read bitline is used to reduce the read bitline discharge delay variation due to local and global process variation in subthreshold region, which is a major portion of memory access time. Based on our simulation results, the worst case read bitline discharge delay is reduced by more than 12× at VDDof 0.3V.
Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS2
2012 Energy-delay efficient asynchronous-logic 16×16-bit pipelined multiplier based on Sense Amplifier-Based Pass Transistor Logic
abstract
We describe an asynchronous-logic (async) 16×16-bit pipelined multiplier based on our proposed Sense Amplifier-Based Pass Transistor Logic (SAPTL) with emphases on high energy-delay efficiency. The multiplier is targeted for an async multi-core System-On-Chip (SOC). This attribute is achieved by simplifying and optimizing the NMOS pass transistor stacks and decision-making C-element, therein to reduce the circuit area overheads and transistor switchings in SAPTL. Based on the simulations (@1V, 65nm CMOS process), the async 16×16-bit pipelined multiplier based on our proposed SAPTL approach features, on average, 31% shorter delay, 21% lower energy/operation achieving a total of 46% lower energy-delay product, and 16% lesser number of transistors when compared to the reported SAPTL approaches.
Weng-Geng Ho, Kwen-Siong Chong, Tong Lin 0001, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS2
2011 A low-power dual-rail inputs write method for bit-interleaved memory cells
abstract
We propose a dual-rail data write technique for bit interleaved memory cells to reduce power dissipation for the write operation without affecting the read operation. The proposed technique can be applied to two reported bit interleaved memory cells with a write power reduction range from 30% to 45%, depending on memory cells and operations. In addition, in the proposed technique, a subthreshold non bit interleaved memory cell is modified to be bit-interleaved without increasing the number of transistors in memory cell.
Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS2
2011 Improved asynchronous-logic dual-rail Sense Amplifier-based Pass Transistor Logic with high speed and low power operation
abstract
We propose a robust asynchronous-logic dual-rail Sense Amplifier-based Pass Transistor Logic (SAPTL) approach with improved speed and power attributes over reported SAPTL approach. These attributes are achieved by simplifying various sub-blocks therein to reduce the stacking of pass transistors and the number of transistor switchings, and to avoid floating nodes. By means of an 8-bit pipeline adder and on the basis of computation simulations (@ 1V, 45nm SOI process), we show that our proposed SAPTL adder is 37% faster, yet 14% lower power dissipation (@ 200MHz input-rate), 18% lower energy dissipation (per operation), and 47% better energy-delay product. These substantially improved attributes are achieved with insignificant overhead - just 3% more transistors.
Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang, Yin Sun 0005, Kok-Leong Chang
ISCAS2
2009 Fine-grained Power Gating for Leakage and Short-circuit Power Reduction by using Asynchronous-logic
abstract
In this paper, a fine-grained power gating technique for an asynchronous-logic pipeline stage is proposed using locally controlled gating transistors. The proposed power gating technique is implemented with minimal control overheads (one additional inverter per pipeline stage for driving PMOS Gating) and delay overheads (within 15% more than the conventional asynchronous-logic pipeline stage). Different types of gating configurations using only PMOS transistor (PMOS Gating), only NMOS transistor (NMOS Gating), and both types of transistors (Dual Gating) are examined and compared. The effectiveness of the proposed power gating technique to the Combinational Block therein with different data input rates is investigated. Based on the computer simulation results, we have found that ≫70% wasted power reduction (including both short-circuit and leakage powers) as compared to the conventional asynchronous-logic pipeline stage can be achieved with all gating configurations. In particular, Dual Gating achieves the best wasted power reduction of 86% for short-circuit power and 99% for leakage power @ 10Mbps input rate.
Tong Lin 0001, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS2
2007 A Low Energy FFT/IFFT Processor for Hearing Aids
abstract
We present a 16-bit low voltage (1.1V - 1.4V) energy efficient 128-point decimation-in-time Fast Fourier Transform/Inverse Fast Fourier Transform (FFT/IFFT) processor specifically for hearing aid applications. The FFT/IFFT processor embodies several low power/energy methodologies, including the clock gating approach,ad-hoccontrol, operand isolation and low power library cells, to satisfy the tight constraints of low voltage low energy and a small silicon area realization for a practical hearing aid. Based on the prototype IC measurements, the proposed FFT/IFFT processor dissipates ~ 188nJ @ 1.1V, features computation delay of2@ 0.35μm CMOS process.
Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS1
2005 A micropower low-voltage multiplier with reduced spurious switching
abstract
We describe a micropower 16/spl times/16-bit multiplier (18.8 /spl mu/W/MHz @1.1 V) for low-voltage power-critical low speed (/spl les/5 MHz) applications including hearing aids. We achieve the micropower operation by substantially reducing (by /spl sim/62% and /spl sim/79% compared to conventional 16/spl times/16-bit and 32/spl times/32-bit designs respectively) the spurious switching in the Adder Block in the multiplier. The approach taken is to use latches to synchronize the inputs to the adders in the Adder Block in a predetermined chronological sequence. The hardware penalty of the latches is small because the latches are integrated (as opposed to external latches) into the adder, termed the latch adder (LA). By means of the LAs and timing, the number of switchings (spurious and that for computation) is reduced from /spl sim/5.6 and /spl sim/10 per adder in the adder block in conventional 16/spl times/16-bit and 32/spl times/32-bit designs respectively to /spl sim/2 in our designs. Based on simulations and measurements on prototype ICs (0.35 /spl mu/m three metal dual poly CMOS process), we show that our 16/spl times/16-bit design dissipates /spl sim/32% less power, is /spl sim/20% slower but has /spl sim/20% better energy-delay-product (EDP) than conventional 16/spl times/16-bit multipliers. Our 32/spl times/32-bit design is estimated to dissipate /spl sim/53% less power, /spl sim/29% slower but is /spl sim/39% better EDP than the conventional general multiplier.
Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
IEEE Trans. Very Large Scale Integr. Syst.1