Ne Kyaw Zwa Lwin

dblp:166/3003 · also Kyaw Zwa Lwin Ne · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
3since 2021 · last 2022
0000-0001-7506-0380ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 since 2021Security and privacy · 1
YearPublicationVenuePosition
2022 Non-profiling based Correlation Optimization Deep Learning Analysis
abstract
Differential Deep Learning Analysis (DDLA) is a deep learning-based non-profiling side-channel attack leveraging neural networks to classify Physical Leakage Information with labels. To avoid the Class Imbalance Problem (CIP) of significantly different data sizes in different data groups, DDLA employs bit labels. However, applying bit labels will be less effective for exploiting leakage. In this paper, we propose to employ Correlation optimization Deep Learning Analysis (CO-DLA) to circumvent the CIP in DDLA by converting the classification in DDLA into a correlation optimization. Bus labels can then be used to exploit stronger leakage information. To validate the attack efficacy improvement, we perform experiments on ASCAD synchronized and de-synchronized masked AES-128 datasets. For the synchronized masked dataset, our proposed CO-DLA requires only 5k traces, which is 75% lesser than the 20k traces required by the reported DDLA, to reveal the key-byte. For the 2 de-synchronized masked datasets, our proposed CO-DLA requires only 10k traces to reveal the key-byte from both of them while the reported DDLA fails to reveal the key-byte.
Juncheng Chen, Jun-Sheng Ng, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Zhiping Lin 0001, Joseph Sylvester Chang, Bah-Hwee Gwee
ISCAS4
2022 An Asynchronous-Logic Masked Advanced Encryption Standard (AES) Accelerator and its Side-Channel Attack Evaluations
abstract
We present a side-channel-attack (SCA) resistant asynchronous-logic (async-logic) Advanced Encryption Standard (AES) accelerator embodying both the masking and hiding SCA countermeasures. Our async-logic masked AES accelerator adopts a dual-rail data encoding to perform the masked 128-bit AES operations, and to enable dual-hiding to moderate both the amplitude (vertical dimension) and the time (horizontal dimension) of the side-channel signals. We implement our async-logic masked AES accelerator in FPGA and comprehensively perform the SCA evaluations based on the electromagnetic (EM) emanation. The SCA evaluations are performed based on bus-wise Hamming Distance model, bus-wise & bit-wise Hamming Weight models, and Zero-Value (ZV) model. Based on our experiment results, we show that our async-logic masked AES is secured against SCA with 1 million EM emanations. This is at least $8.3 \times$ more resistant than synchronous-logic masked AES and $200 \times$ more resistant than the synchronous-logic unmasked AES.
Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Joseph Sylvester Chang, Bah-Hwee Gwee
ISCAS4
2021 Normalized Differential Power Analysis - for Ghost Peaks Mitigation
abstract
The attack efficacy of Differential Power Analysis (DPA), a popular side channel evaluation technique for key extraction, is compromised by the false highest Difference Of Means (DOMs) value ('ghost peaks') in the DOMs matrix produced in a conventional DPA. The ghost peak is generated by the wrong key guess and always occurs in the conventional DPA when the number of side channel traces is not enough. In this paper, an improved version of the conventional DPA termed as Normalized DPA (NDPA) is proposed to circumvent the ghost peak. With the analysis on the generation of ghost peaks in the conventional DPA, we observed that by normalizing the DOMs matrix, the ghost peaks can be greatly suppressed. We model the proposed NDPA mathematically and show that it performs better than the conventional DPA. We further provide the experimental validations on a set of 200k power simulation traces on AES S- Box and 500 EM traces from ASCAD dataset. Based on the attack results of these datasets, our proposed NDPA requires (up to 68%) lesser number of traces to reveal a correct key when compared to the conventional DPA.
Juncheng Chen, Jun-Sheng Ng, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Weng-Geng Ho, Kwen-Siong Chong, Zhiping Lin 0001, Joseph Sylvester Chang, Bah-Hwee Gwee
ISCAS4
2020 Radiation-Hardened-by-Design (RHBD) Digital Design Approaches: A Case Study on an 8051 Microcontroller
abstract
Advanced satellites and/or high-level (levels 4 and 5) autonomous vehicles demand high reliability integrated circuits (ICs) with ultra-low error rates. One solution is to use radiation-hardened-by-design (RHBD) design techniques to mitigate the error rates against the single-event-effects (arising from radiation effects). This paper first provides an overview on several present-art RHBD design techniques, and then propose an RHBD design methodology, spanning from the library cell development, circuit simulation and synthesis, to the layout implementation, to realize digital circuits. We further demonstrate an 8051 microcontroller with the proposed design methodology, and evaluate the 8051 microcontroller prototype (@ 65nm CMOS) with irradiation tests. Our 8051 microcontroller is error-free with 10 MeV.mg/cm2, meeting our targeted specifications for Low Earth Orbit applications. When under high Linear Transfer Energy (> 51.5 MeV.mg/cm2) tests, the 8051 microcontroller does suffer errors. We further study/analyze which part of the 8051 microcontroller to cause errors, and provide recommendations.
Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Wei Shu, Joseph Sylvester Chang
ISCAS2
2020 A DPA-Resistant Asynchronous-Logic NoC Router with Dual-Supply-Voltage-Scaling for Multicore Cryptographic Applications
abstract
We propose a 5-port asynchronous-logic Network-on-Chip (ANoC) router based on the Sense-Amplifier Half-Buffer (SAHB) approach for cryptographic processing cores to counteract side channel attack differential power analysis (DPA) in multicore platform. There are three features in the proposed DPA-resistant ANoC router. First, the proposed ANoC router embodies dual-supply-voltage SAHB cells, where the non-critical subsidiary supply voltage is adjustable from 0.3V to 1.2V, increasing the noise variance and hence reducing the Signal-to-Noise (SNR) ratio to hide the information leakage. Second, the proposed ANoC router performs as a noise engine by increasing the number of power-on IO ports, further randomizing the overall power dissipation. Third, the proposed ANoC router can switch between DPA-resistant mode and energy-efficient nominal (non-secure) mode, saving the power dissipation when the DPA secure countermeasure is unnecessary. Based on 65nm CMOS process, the multicore platform embedded with the proposed ANoC router is implemented, and the experiment is demonstrated by running the advanced encryption standard (AES) cryptography operation. When benchmarked against the nominal mode, the noise power variance of the proposed ANoC router increases by 2.3× in the DPA-resistant mode, reducing the overall SNR ratio by 56%. When comparing to other reported noise engines, our proposed ANoC router is one of the most DPA-secure, area-efficient and power-efficient designs for multicore cryptographic applications.
Weng-Geng Ho, Ne Kyaw Zwa Lwin, Nay Aung Kyaw, Jun-Sheng Ng, Juncheng Chen, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS2
2020 A Highly Efficient Power Model for Correlation Power Analysis (CPA) of Pipelined Advanced Encryption Standard (AES)
abstract
We evaluate the vulnerability of a pipelined Advanced Encryption Standard (AES) against Correlation Power Analysis (CPA) Side-Channel Attack (SCA). We identify that the registers in pipelined AES are most vulnerable against CPA SCA and propose a new power model targeting the switching activities of the registers. The proposed power model is constructed based on the Hamming Distance (HD) between the intermediate values stored in the registers in two consecutive clock cycles. Then, we analyze the vulnerability of pipelined AES under two scenarios. First, during regular pipeline operation where the device is performing AES pipeline operation. Second, in non-pipeline operation where we assume the adversaries can insert delay to the input of the device to increase the signal to noise ratio of the physical leakage information. The simulation results show that under regular pipelined operation, our proposed power model can reveal all the 16 key bytes in less than 4,900 traces, resulting in 4.7× more effective than the conventional power models. Under non-pipelined operation, our proposed power model requires only 590 traces to reveal all the 16 key bytes, which is 5.9× more effective than other power models.
Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee
ISCAS4
2019 Low Gate-Count Ultra-Small Area Nano Advanced Encryption Standard (AES) Design
abstract
We present a low gate-count ultra-small area nano advanced encryption standard (AES) design. We achieve the low gate-count by the following means. First, we repeatedly reuse the area-critical circuits, i.e. one 8-bit Substitute-Box (S-Box) circuit and one 32-bit MixColumn circuit, for AES. Second, we cascade the input flip-flops (FFs) with our data transfer architecture so that the outputs of the MixColumn circuit are connected directly to the first 32-bit input FFs without extra multiplexing circuits. Third, the ShiftRow operation is implicitly performed by assigning the data sequence to the input FFs (during the S-Box and MixColumn operations). Fourth, we use independent XOR gates for AddRound and KeyExpansion operations. The collective means enables our design to feature 1457 gates, and to occupy 100um×100um area @ 65nm CMOS. When compared to the normalized area (@ 65nm CMOS) of the reported AES designs, our design features the smallest normalized area, 10% smaller than the most competitive reported AES design. Our design is targeted for ultra-small area applications including biomedical applications.
Aparna Shreedhar, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Nay Aung Kyaw, L. Nalangilli, Wei Shu, Joseph Sylvester Chang, Bah-Hwee Gwee
ISCAS3
2019 A High Throughput and Secure Authentication-Encryption AES-CCM Algorithm on Asynchronous Multicore Processor
abstract
We propose an authentication-based matrix-transformation cum parallel-encryption implemented on an asynchronous multicore processor (AMP-MP) to achieve a high throughput and yet secure advanced encryption standard based on counter with chaining mode (AES-CCM). There are four main features in our proposed AMP-MP. First, we employ the matrix multiplication in GF(28) computation to transform the 16 plaintexts into one plaintext, hence improving the authentication speed by 32× collectively at the transmitter and receiver. Second, we reschedule the operations of three AES encryptions in three different cores such that their physical leakages are compensated and equalized, thus reducing the correlation of physical leakage with the processed data by >3×. Third, the intermediate values of AES-CCM are propagated asynchronously between different cores to randomize the physical leakages with the processed data, and therefore further enhance the security of AES-CCM against the SCA by another 3×. Fourth, we propose a key adjusting technique based on S-Box byte-key transformation to protect the key against pattern-based attack. Our proposed AMP-MP is realized on an 8-bit asynchronous 9-core processor fabricated based on the 65 nm CMOS process. The experimental results show that the throughput of the authentication is 13.54 Gbps while the throughput for both authentication and encryption collectively is 8.32 Gbps, which are 17× and 70× faster than the reported counterparty, respectively. Based on power dissipation and EM SCA on our proposed AMP-MP, the secret key is unrevealed at 5 × 105 traces, which is ~17× more secured than the standard ASIC AES-CCM implementation.
Ali Akbar Pammu, Weng-Geng Ho, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Bah-Hwee Gwee
IEEE Trans. Inf. Forensics Secur.3
2018 Asynchronous-Logic QDI Quad-Rail Sense-Amplifier Half-Buffer Approach for NoC Router Design
abstract
We propose a low area overhead and power-efficient asynchronous-logic quasi-delay-insensitive (QDI) sense-amplifier half-buffer (SAHB) approach with quad-rail (i.e., 1-of-4) data encoding. The proposed quad-rail SAHB approach is targeted for area- and energy-efficient asynchronous network-on-chip (ANoC) router designs. There are three main features in the proposed quad-rail SAHB approach. First, the quad-rail SAHB is designed to use four wires for selecting four ANoC router directions, hence reducing the number of transistors and area overhead. Second, the quad-rail SAHB switches only one out of four wires for 2-bit data propagation, hence reducing the number of transistor switchings and dynamic power dissipation. Third, the quad-rail SAHB abides by QDI rules, hence the designed ANoC router features high operational robustness toward process-voltage-temperature (PVT) variations. Based on the 65-nm CMOS process, we use the proposed quad-rail SAHB to implement and prototype an 18-bit ANoC router design. When benchmarked against the dual-rail counterpart, the proposed quad-rail SAHB ANoC router features 32% smaller area and dissipates 50% lower energy under the same excellent operational robustness toward PVT variations. When compared to the other reported ANoC routers, our proposed quad-rail SAHB ANoC router is one of the high operational robustness, smallest area, and most energy-efficient designs.
Weng-Geng Ho, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Bah-Hwee Gwee, Joseph Sylvester Chang
IEEE Trans. Very Large Scale Integr. Syst.3
2016 High performance low overhead template-based Cell-Interleave Pipeline (TCIP) for asynchronous-logic QDI circuits
abstract
We propose a novel Template-based Cell-Interleave Pipeline (TCIP) approach for generating high performance and yet low overhead asynchronous-logic (async) quasi-delay-insensitive (QDI) circuits. Our TCIP approach exploits the characteristics of the four prevalent QDI cell templates, namely Weak-Conditioned Half-Buffer (WCHB), Pre-Charged HalfBuffer (PCHB), Autonomous Signal-Validity Half-Buffer (ASVHB), and Sense-Amplifier Half-Buffer (SAHB), and then strategically interleave these template cells to form a composite pipeline. There are three main features in our TCIP approach. First, all QDI cell templates are first standardized with the same interface signals, and their corresponding cells are characterized in terms of transistor count, cycle time and energy dissipation for ease of comparison/selection/replacement. Second, our TCIP approach prioritizes the speed requirement when forming the initial pipeline circuits, and then subsequently reduces circuit overheads by interleaving various template cells without compromising the speed significantly. Third, the final optimized QDI pipeline circuit inherently features high robustness against process-voltage-temperature (PVT) variations, hence suitable for dynamic-voltage-scaling (DVS) operation. By means of 65nm CMOS process, we demonstrate a 4-bit pipeline tree adder based on the proposed TCIP approach, and benchmark it against the WCHB, PCHB, ASVHB and SAHB counterparts. These five designs feature same high operational robustness, nonetheless the design based on our TCIP approach is more competitive. Particularly, the designs based on reported approaches are, on average, ∼1.22× more transistor count, ∼1.21× slower and ∼1.22× higher energy dissipation. Furthermore, under DVS operation from 1.2V to 0.3V, our proposed TCIP adder can reduce up to 88% energy for non-speed critical applications.
Weng-Geng Ho, Nan Liu 0002, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS3
2016 Area-efficient and low stand-by power 1k-byte transmission-gate-based non-imprinting high-speed erase (TNIHE) SRAM
abstract
We propose a novel 15-T Transmission-gate-based Non-Imprinting High-speed Erase (TNIHE) SRAM cell with emphases on low area overhead and low stand-by power attributes for highly secured data storage applications. We benchmark our proposed 15-T TNIHE SRAM cell against the reported 22-T Non-Imprinting High-speed Erase (NIHE) SRAM cell, and demonstrated three key features of reducing 7 transistors. First, we adopt the transmission gates (as opposed the cross-couple inverters) in the slave circuitry, saving 4 transistors. Second, we eliminate a transistor which uses to reset the slave circuitry, hence saving 1 transistor. Third, we apply the global inverse transistors (as opposed to the local inverse transistors) in the read /write circuit for each SRAM cell, hence further reduce 2 more transistors. As a result, our proposed TNIHE SRAM cell @ 65nm CMOS features ~17% smaller layout area. We design a 1k-byte memory based on the proposed TNIHE SRAM cells. On the basis of simulations, we show that our 1k-byte SRAM memory features overall ~13% smaller area, and dissipates on average, ~30% lower stand-by power than the reported NIHE counterpart.
Weng-Geng Ho, Ne Kyaw Zwa Lwin, N. Prashanth Srinivas, Kwen-Siong Chong, Tony Tae-Hyoung Kim, Bah-Hwee Gwee
ISCAS2
2016 Total Ionizing Dose (TID) effects on finger transistors in a 65nm CMOS process
abstract
Although Total Ionizing Dose (TID) effects are generally unpronounced in deep-submicron-CMOS, we show the TID-induced leakage current @TID=500Krad is significant in NMOS-finger-transistors of GlobalFoundries 65nm CMOS. Further, Radiation-Hardening-By-Design techniques against said TID effect are recommended.
Jize Jiang, Wei Shu, Kwen-Siong Chong, Tong Lin 0001, Ne Kyaw Zwa Lwin, Joseph Sylvester Chang
ISCAS5
2016 Experimental investigation into radiation-hardening-by-design (RHBD) flip-flop designs in a 65nm CMOS process
abstract
We comprehensively study three types of radiation-hardened flip-flops: DICE for SEU-hardening, temporal for SET-hardening, and Triple-Modular-Redundancy for SEU-cum-SET-hardening. Our study includes their trade-offs of circuit/radiation-hardness attributes. We find that DICE flip-flops remain the most competitive.
Tong Lin 0001, Kwen-Siong Chong, Wei Shu, Ne Kyaw Zwa Lwin, Jize Jiang, Joseph Sylvester Chang
ISCAS4
2015 High robustness energy- and area-efficient dynamic-voltage-scaling 4-phase 4-rail asynchronous-logic Network-on-Chip (ANoC)
abstract
We propose an 18-bit 5-interface asynchronous-logic Network-on-Chip (ANoC) router based on the quasi-delay-insensitive (QDI) realization approach for high secured cryptography applications. There are four key features of the proposed ANoC router. First, it embodies the novel high-speed low-power Sense-Amplifier Half Buffer 4-rail cells. Second, it is designed based on QDI protocol, and hence is highly robust against process-voltage-temperature (PVT) variations. Third, it is functional for full dynamic voltage scaling from nominal (VDD=1.2V) to sub-threshold (VDD=0.3V) regions, and is potentially excellent for low power management applications. Fourth, it embodies a distributed-based XY routing algorithm to utilize a 4-bit header of flow control unit (flit) for routing up to 4×4 cluster, hence minimizing the routing overhead. We realize the proposed ANoC router (@65nm CMOS), and benchmark it against the reported ANoC router embodying the conventional Weak-Conditioned Half-Buffer (WCHB) QDI realization approach. Both our proposed and reported designs feature the high operation robustness, but our design is 41% more energy-efficient, and 21% more area-efficient than the reported counterpart. The prototype of ANoC router occupies only 0.105 mm2and can operate down to 0.3V. At VDD=0.3V, it dissipates 44 fJ per bit and operate 105 ns per flit.
Weng-Geng Ho, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Bah-Hwee Gwee, Joseph Sylvester Chang
ISCAS3