VLDB 2026 Research / reviewers in the wild / expert
Joseph Sylvester Chang
dblp:93/4217 · also Joseph S. Chang
· DBLP profile ↗
65ranked-venue papers
1as first author
13since 2021 · last 2025
0000-0003-0991-8339ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 61 · 13 since 2021Artificial intelligence and machine learning · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Novel Energy-Efficient Continuous-Time Hysteretic VCO-Based ComparatorabstractVoltage-controlled oscillator (VCO)-based comparators offer higher energy efficiency as the difference in input magnitudes increase, such as in level-crossing ADCs. Nevertheless, to date, they require a clock signal to perform comparison operations. This is incongruous with continuous-time applications, where inputs are compared continuously. Further, they lack hysteresis, a crucial feature for mitigating spurious switching that compromises energy efficiency. In this paper, we present a novel VCO-based comparator that, for the first time, simultaneously achieves continuous-time operation and high energy efficiency. The former feature is enabled by a novel continuous-time decision circuit, while the latter is achieved through a novel switched-current hysteresis circuit that mitigates spurious switching. The proposed comparator is designed in 65 nm CMOS. Simulation results show that it achieves low energy per comparison, ranging from 0.07 to 4 pJ, with an average propagation delay of ~15 ns. The average energy consumption is 0.19 pJ — ~1.8× lower than the state-of-the-art VCO-based comparator. Jinhen Lee, Victor Adrian, Kinglouis Steven Tantra, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 5 |
| 2025 | Live Demonstration: AI-based System Latchup Detection and Protection for COTS SystemsabstractThe adoption of Commercial Off-The-Shelf (COTS) ICs in modern satellites faces challenges from radiation-induced latchup events, with current protection methods showing significant limitations in detection accuracy and applicability. This demonstration presents a novel adaptive AI-based latchup detection and protection system featuring two-stage training and LSTM neural network analysis. Our FPGA implementation achieves 90% detection accuracy without extensive pre-characterization. Visitors can validate the system's performance through real-time interaction with various latchup scenarios. Yin Sun 0005, Junkai Zhao, Rouli Fang, Tony Zhang, Kwen-Siong Chong, Wei Shu, Joseph Sylvester Chang |
ISCAS | 7 |
| 2025 | A Novel High-Accuracy Inductor-Current Estimator for Digitally-Controlled Synchronous DC-DC Buck ConvertersabstractThis paper presents a novel digital inductor-current estimator for digitally-controlled current-ripple-based synchronous DC-DC buck converters. The estimator features high accuracy, which is imperative for current-ripple control that requires precise estimation of instantaneous DC and inductor ripple currents in both the discontinuous and continuous conduction modes. The estimation method indirectly determines the currents by constructing a digital representation of the voltage across the inductor, rendering it applicable in both conduction modes. This method is more accurate and simpler than conventional methods that estimate DC and inductor ripple currents directly in each conduction mode. Compared to state-of-the-art methods, the proposed estimator achieves an average DC current estimation error that is ≥5.7× smaller and an inductor ripple current estimation error that is ≤ 1.90% over various load conditions. Yanshan Xie, Victor Adrian, Jinhen Lee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2025 | An Adaptive AI-based Approach to Detect and Protect COTS Systems against Micro-Single-Event-Latchups (μ-SELs) and SELsabstractIn our envisioned ‘Next Paradigm’ of ‘New Space’, commercial-off-the-shelf (COTS) systems (embodying multiple COTS ICs) would be employed as payloads in space missions. Most COTS ICs are susceptible to radiation effects, particularly Micro-Single-Event-Latchups (μ-SELs) and SELs, and their characteristics are expectedly different. Consequently, hitherto reported detection approaches require characterization of the individual COTS ICs and the entire system, thereby rendering excessive overheads when applied to different COTS systems. In this paper, we propose, for the first time, the design and implementation of an adaptive AI-based approach to detect and protect various uncharacterized COTS systems (vis-à-vis pre-characterized ones) against μ-SELs and SELs. Our proposal involves the adoption of the Long-Short-Term-Memory (LSTM) neural network with our proposed two-stage training process – ex-situ pre-training and in-situ re-training – to improve general applicability. Our FPGA-based prototype achieves high (~90%) average accuracy for four different payloads. This is a worthy improvement of 13.3%-28.5% over reported approaches, yet requiring low (~115 mW) power consumption. Collectively, our proposed approach is appropriate for resource-constrained space applications and our ‘Next Paradigm’ of ‘New Space’. Junkai Zhao, Yin Sun 0005, Tony Zhang, Kwen-Siong Chong, Wei Shu, Joseph Sylvester Chang |
ISCAS | 6 |
| 2025 | Single-Ended/Differential Wideband Track-and-Hold Amplifier in 22-nm FD-SOI CMOS ProcessabstractThe impending 6G communication based on the software defined radio (SDR) requires a radio frequency (RF) track-and-hold amplifier (THA). This THA serves as the frequency down-converter and the single-to-differential interface to the downstream analog-to-digital converter (ADC). We present a CMOS RF THA that features wide and width (18 GHz), yet high linearity (spurious free dynamic range (SFDR) of 56.7 dB) and not requiring an external balun. These features are derived from our proposed isolation technique based on our proposed double source follower enhanced (DSFE) structure. To realize the single-to-differential conversion without an external balun, we design an independent balun as the first stage. Thereafter, we employ our proposed feedforward compensation technique (FCT) along with the reported phase correction technique (PCT) to reduce the output mismatches while simultaneously enhancing the linearity and bandwidth. We monolithically realize the RF THA in 22-nm fully-depleted silicon-on-insulator (FD-SOI) CMOS operating at 1.8 V. Measurements depict that the input bandwidth is wide (18 GHz), yet featuring high linearity (SFDR =56.7 dB at 15 GHz) with 2 GS/s sampling rate. The power consumption and the chip area are low and small at 216 mW and 0.07 mm2, respectively. When benchmarked against reported III/V RF THAs, the proposed CMOS RF THA is very competitive—comparable bandwidth, yet simultaneously higher linearity, potentially lower cost, lower power dissipation, and smaller die area. Further because it is realized in CMOS, it facilitates integration to other CMOS circuits in the same system-on-chip (SoC). Zixian Zheng, Wei Shu, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | A 3D-Printed Fourth-Order Stacked Filter for Integrated DC-DC ConvertersabstractThe passive devices in state-of-the-art miniaturized switched-mode DC-DC converters are generally integrated by means of on-chip and in-package methods. Nevertheless, the quality is poor-to-moderate, thereby compromising the power-efficiency. In this paper, we propose the miniaturization of the DC-DC converter by means of realizing its passive devices as embedded devices that are printed within a high-density 3D inkjet printed-circuit-board (PCB). We propose a fourth-order stacked LC filter embodying passive components with small values-effectively at no additional cost because they are embedded through 3D-printing. For the inductor and capacitor, we propose to adopt a high-$Q$solenoidal structure and the metal-insulator-metal planar structure, respectively. The proposed filter is printed within the 3D-PCB with a compact 124 mm3volume due to the stacked arrangement. The measured AC attenuation is 21.2 dB at 200 MHz. The filter is further verified by means of computer simulations of a DC-DC buck converter. Simulation results of the converter employing the filter show a low output voltage ripple at 146 mV and a high peak power-efficiency of ~78% at 200 MHz switching frequency with 150 mA load current. Jinhen Lee, Victor Adrian, Sun-Yang Tay, Yanshan Xie, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 6 |
| 2023 | An Accurate Digital Inductor Current Sensor for Current-Ripple-Based DC-DC ConvertersabstractThis paper presents a digital current sensor for digitally-controlled current-ripple-based DC-DC buck converters to estimate the instantaneous inductor-current ripple accurately in both the Discontinuous (DCM) and Continuous (CCM) Current Modes. The current sensor employs a proposed dual-mode input multiplexing technique to select an appropriate representation of the pertinent voltage of the switching node$(\boldsymbol{V}_{\boldsymbol{x}})$in any mode, thereby allowing the current to be estimated more accurately compared to that of the prior-art design. The accurate inductor-current ripple information enables the controller to yield output-voltage transient response with small overshoot or undershoot (OS/US) and fast settling time. Benchmarking results using a digitally-controlled current-ripple constant on-time DC-DC buck converter show that the converter employing the proposed sensor achieves$\geq \mathbf{49}{\%}$smaller OS/US and$\geq \mathbf{45}{\%}$faster settling time at the output voltage in both the DCM and the CCM collectively compared with that of the same converter but with the prior-art sensor. Yanshan Xie, Victor Adrian, Sun-Yang Tay, Jinhen Lee, Pak Kwong Chan, Joseph Sylvester Chang |
ISCAS | 6 |
| 2022 | Non-profiling based Correlation Optimization Deep Learning AnalysisabstractDifferential Deep Learning Analysis (DDLA) is a deep learning-based non-profiling side-channel attack leveraging neural networks to classify Physical Leakage Information with labels. To avoid the Class Imbalance Problem (CIP) of significantly different data sizes in different data groups, DDLA employs bit labels. However, applying bit labels will be less effective for exploiting leakage. In this paper, we propose to employ Correlation optimization Deep Learning Analysis (CO-DLA) to circumvent the CIP in DDLA by converting the classification in DDLA into a correlation optimization. Bus labels can then be used to exploit stronger leakage information. To validate the attack efficacy improvement, we perform experiments on ASCAD synchronized and de-synchronized masked AES-128 datasets. For the synchronized masked dataset, our proposed CO-DLA requires only 5k traces, which is 75% lesser than the 20k traces required by the reported DDLA, to reveal the key-byte. For the 2 de-synchronized masked datasets, our proposed CO-DLA requires only 10k traces to reveal the key-byte from both of them while the reported DDLA fails to reveal the key-byte. Juncheng Chen, Jun-Sheng Ng, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Zhiping Lin 0001, Joseph Sylvester Chang, Bah-Hwee Gwee |
ISCAS | 7 |
| 2022 | An Integrated DC-DC Converter with Novel Asymmetrical Segmented Power-Stages for Sustained High Power-EfficienciesabstractThe average power-efficiency of integrated DC- DC converters for Internet-of-Things is generally compromised over a wide load current range. This is because their efficiency is typically severely compromised at light load currents. We present a novel asymmetrical segmented power-stage configuration to improve the average power-efficiency of integrated converters. We achieve this by configuring different power-stage segments with different sizes of power transistors and their inductors, and a circuit to enable the corresponding segment for high power-efficiencies at different load conditions. Specifically, the circuit enables the segment with small-sized power transistors and a large inductor for light-load operations, and conversely, it enables the segment with large-sized power transistors and a small inductor for heavy-load operations. The integrated converter employing our proposed configuration is designed using a CMOS 180 nm process for 2. 5-3.3V input, 1.2 V output, and 50 MHz switching frequency. Simulation results show the proposed converter achieves a high average power-efficiency at ~73% over a wide load current range of 5-200mA. When benchmarked against the competing contemporary designs, the proposed converter features 5-30% higher average power-efficiency over the wide load current range, and >34% higher power-efficiency at 20 mA light load. Jinhen Lee, Victor Adrian, Joseph Sylvester Chang, Yin Sun 0005, Sun-Yang Tay |
ISCAS | 3 |
| 2022 | An Asynchronous-Logic Masked Advanced Encryption Standard (AES) Accelerator and its Side-Channel Attack EvaluationsabstractWe present a side-channel-attack (SCA) resistant asynchronous-logic (async-logic) Advanced Encryption Standard (AES) accelerator embodying both the masking and hiding SCA countermeasures. Our async-logic masked AES accelerator adopts a dual-rail data encoding to perform the masked 128-bit AES operations, and to enable dual-hiding to moderate both the amplitude (vertical dimension) and the time (horizontal dimension) of the side-channel signals. We implement our async-logic masked AES accelerator in FPGA and comprehensively perform the SCA evaluations based on the electromagnetic (EM) emanation. The SCA evaluations are performed based on bus-wise Hamming Distance model, bus-wise & bit-wise Hamming Weight models, and Zero-Value (ZV) model. Based on our experiment results, we show that our async-logic masked AES is secured against SCA with 1 million EM emanations. This is at least $8.3 \times$ more resistant than synchronous-logic masked AES and $200 \times$ more resistant than the synchronous-logic unmasked AES. Jun-Sheng Ng, Juncheng Chen, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Joseph Sylvester Chang, Bah-Hwee Gwee |
ISCAS | 6 |
| 2022 | A Versatile and Accurate Vector-Based Method for Modeling and Analyzing Planar Air-Core InductorsabstractPlanar air-core inductors come in a variety of geometrical shapes, including in the form of the conventional spiral geometry and novel complex geometries. In the design phase of a system, the inductance of the employed inductor would need to be ascertained. This is usually ascertained by tedious mathematical derivations on a segment-by-segment (inductor) basis or time-consuming computer modeling, and the complexity can become intractable for complex geometries. In this paper, we propose a versatile, yet accurate, vector-based method to ascertain the inductance of planar air-core inductors with virtually any geometry, including novel complex geometry inductors—rather easily. Our proposed method decomposes the inductor segments into vectors, and thereafter utilizes geometric models to compute the inductance in a systematic fashion. We benchmark our proposed method against the conventional electromagnetic field solver simulations to estimate the inductances of six planar inductors ranging from a conventional spiral air-core inductor to that embodying different and complex geometries. On the basis of these six inductor examples, we show that our method is highly accurate with a worst-case error of $\sim 5$% compared to that obtained using conventional electromagnetic field solver. Of particular interest, our modeling for novel complex geometry planar inductors is relatively simple. Sun-Yang Tay, Victor Adrian, Joseph Sylvester Chang, Jinhen Lee, Bah-Hwee Gwee |
ISCAS | 3 |
| 2022 | A Highly Secure FPGA-Based Dual-Hiding Asynchronous-Logic AES Accelerator Against Side-Channel AttacksabstractEncryption in field-programmable gate array (FPGA) often provides a good security solution to protect data privacy in Internet-of-Things systems, but this security solution can be compromised by side-channel attacks (SCAs). In this article, we present an FPGA-based dual-hiding asynchronous-logic (async-logic) advanced encryption standard (AES) accelerator, which is highly resistant against SCAs and yet low area/energy overheads. The proposed AES accelerator achieves vertical (amplitude) SCA hiding via an area-efficient dual-rail mapping approach and a zero-value (ZV) compensated substitution-box (S-Box), while enhancing the horizontal (temporal) SCA hiding of async-logic operations via a timing-boundary-free input arrival-time randomizer and a skewed-delay controller. A comprehensive SCA evaluation is performed with 11 SCA models, and we show that our proposed design can offer a strong SCA resistance with measurement-to-disclosure (MTD) of >20 million traces. To our best knowledge, our design is the most secure AES design evaluated with the largest number of traces in FPGA. To compare the design overheads for security, we quantify the figure of merit as normalized (Area$\times $Energy/MTD(All)$\times 10^{6}$). The figure of merit of our proposed design is$403\times $smaller than the benchmark dual-rail synchronous-logic design and$95\times $smaller than a reported async-logic design. Jun-Sheng Ng, Juncheng Chen, Kwen-Siong Chong, Joseph Sylvester Chang, Bah-Hwee Gwee |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | Normalized Differential Power Analysis - for Ghost Peaks MitigationabstractThe attack efficacy of Differential Power Analysis (DPA), a popular side channel evaluation technique for key extraction, is compromised by the false highest Difference Of Means (DOMs) value ('ghost peaks') in the DOMs matrix produced in a conventional DPA. The ghost peak is generated by the wrong key guess and always occurs in the conventional DPA when the number of side channel traces is not enough. In this paper, an improved version of the conventional DPA termed as Normalized DPA (NDPA) is proposed to circumvent the ghost peak. With the analysis on the generation of ghost peaks in the conventional DPA, we observed that by normalizing the DOMs matrix, the ghost peaks can be greatly suppressed. We model the proposed NDPA mathematically and show that it performs better than the conventional DPA. We further provide the experimental validations on a set of 200k power simulation traces on AES S- Box and 500 EM traces from ASCAD dataset. Based on the attack results of these datasets, our proposed NDPA requires (up to 68%) lesser number of traces to reveal a correct key when compared to the conventional DPA. Juncheng Chen, Jun-Sheng Ng, Nay Aung Kyaw, Ne Kyaw Zwa Lwin, Weng-Geng Ho, Kwen-Siong Chong, Zhiping Lin 0001, Joseph Sylvester Chang, Bah-Hwee Gwee |
ISCAS | 8 |
| 2020 | Radiation-Hardened-by-Design (RHBD) Digital Design Approaches: A Case Study on an 8051 MicrocontrollerabstractAdvanced satellites and/or high-level (levels 4 and 5) autonomous vehicles demand high reliability integrated circuits (ICs) with ultra-low error rates. One solution is to use radiation-hardened-by-design (RHBD) design techniques to mitigate the error rates against the single-event-effects (arising from radiation effects). This paper first provides an overview on several present-art RHBD design techniques, and then propose an RHBD design methodology, spanning from the library cell development, circuit simulation and synthesis, to the layout implementation, to realize digital circuits. We further demonstrate an 8051 microcontroller with the proposed design methodology, and evaluate the 8051 microcontroller prototype (@ 65nm CMOS) with irradiation tests. Our 8051 microcontroller is error-free with 10 MeV.mg/cm2, meeting our targeted specifications for Low Earth Orbit applications. When under high Linear Transfer Energy (> 51.5 MeV.mg/cm2) tests, the 8051 microcontroller does suffer errors. We further study/analyze which part of the 8051 microcontroller to cause errors, and provide recommendations. Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Wei Shu, Joseph Sylvester Chang |
ISCAS | 4 |
| 2020 | A DPA-Resistant Asynchronous-Logic NoC Router with Dual-Supply-Voltage-Scaling for Multicore Cryptographic ApplicationsabstractWe propose a 5-port asynchronous-logic Network-on-Chip (ANoC) router based on the Sense-Amplifier Half-Buffer (SAHB) approach for cryptographic processing cores to counteract side channel attack differential power analysis (DPA) in multicore platform. There are three features in the proposed DPA-resistant ANoC router. First, the proposed ANoC router embodies dual-supply-voltage SAHB cells, where the non-critical subsidiary supply voltage is adjustable from 0.3V to 1.2V, increasing the noise variance and hence reducing the Signal-to-Noise (SNR) ratio to hide the information leakage. Second, the proposed ANoC router performs as a noise engine by increasing the number of power-on IO ports, further randomizing the overall power dissipation. Third, the proposed ANoC router can switch between DPA-resistant mode and energy-efficient nominal (non-secure) mode, saving the power dissipation when the DPA secure countermeasure is unnecessary. Based on 65nm CMOS process, the multicore platform embedded with the proposed ANoC router is implemented, and the experiment is demonstrated by running the advanced encryption standard (AES) cryptography operation. When benchmarked against the nominal mode, the noise power variance of the proposed ANoC router increases by 2.3× in the DPA-resistant mode, reducing the overall SNR ratio by 56%. When comparing to other reported noise engines, our proposed ANoC router is one of the most DPA-secure, area-efficient and power-efficient designs for multicore cryptographic applications. Weng-Geng Ho, Ne Kyaw Zwa Lwin, Nay Aung Kyaw, Jun-Sheng Ng, Juncheng Chen, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 8 |
| 2019 | Low Gate-Count Ultra-Small Area Nano Advanced Encryption Standard (AES) DesignabstractWe present a low gate-count ultra-small area nano advanced encryption standard (AES) design. We achieve the low gate-count by the following means. First, we repeatedly reuse the area-critical circuits, i.e. one 8-bit Substitute-Box (S-Box) circuit and one 32-bit MixColumn circuit, for AES. Second, we cascade the input flip-flops (FFs) with our data transfer architecture so that the outputs of the MixColumn circuit are connected directly to the first 32-bit input FFs without extra multiplexing circuits. Third, the ShiftRow operation is implicitly performed by assigning the data sequence to the input FFs (during the S-Box and MixColumn operations). Fourth, we use independent XOR gates for AddRound and KeyExpansion operations. The collective means enables our design to feature 1457 gates, and to occupy 100um×100um area @ 65nm CMOS. When compared to the normalized area (@ 65nm CMOS) of the reported AES designs, our design features the smallest normalized area, 10% smaller than the most competitive reported AES design. Our design is targeted for ultra-small area applications including biomedical applications. Aparna Shreedhar, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Nay Aung Kyaw, L. Nalangilli, Wei Shu, Joseph Sylvester Chang, Bah-Hwee Gwee |
ISCAS | 7 |
| 2019 | A Fully Additive Low-Temperature All-Air Low-Variation Printed/Flexible Electronics With Self-Compensation for Bending: Codesign From Materials, Design, Fabrication, and ApplicationsabstractSensing and its electronics based on printed/flexible electronics offer unique attributes of mechanical flexibility of its substrate, hence the unique applications. Nevertheless, one of the ensuing key challenges of flexible-electronics-based sensing is the issues associated with consistency and repeatability of its parameters of the flexible electronics elements/sensors due to their variations, which are sometimes intractable. These may be due to manufacturing variations, aging, when the substrate is bent, and so on. In this article, we describe our codesign between the different chains of the flexible electronics supply chain to derive practically flexible electronics and sensors for applications where the substrate is expected to bend, e.g., in an augmented sensing e-skin smart glove application. This effort includes our fully additive low-temperature all-air low-cost screen printing process, and how we obtain consistency and repeatability. To improve the matching of thin-film transistors-a critical consideration for conditioning sensor outputs-we describe layout techniques where relatively good matching can be achieved, but with area overheads. We describe how we accommodate the variations of printed elements and circuits embodying printed elements when the substrate is bent-a self-compensating means-and propose the application of the same for printed sensors. The cost of our self-compensation means is without power or area overheads, albeit more (uncomplicated) printing steps. We finally describe our process development kit (PDK) encompassing all of the aforesaid to predict the performance of the printed circuits and sensors, including the effects of bending and our proposed self-compensation thereto. We demonstrate the efficacy of our methods based on measurements on printed elements and circuits. Joseph Sylvester Chang, Tong Ge |
Proc. IEEE | 1 |
| 2019 | Flexible Electronic Skin: From Humanoids to HumansabstractThis special issue provides state-of-the-art coverage of the theoretical, scientific, and practical aspects related to flexible electronic skin. Ravinder S. Dahiya, Deji Akinwande, Joseph Sylvester Chang |
Proc. IEEE | 3 |
| 2018 | An Air-Core Coupled-Inductor Based Dual-Phase Output Stage for Point-of-Load ConvertersabstractThis paper presents a high-switching-frequency and high-efficiency dual-phase output stage for Point-of-Load (POL) converters. The proposed design is based on a novel air-core coupled-inductor that exhibits variable inductance. Specifically, the proposed coupled-inductor exhibits different equivalent inductance within one switching cycle, thus offering combined merits of fast transient response, small current ripple, and good current balance. The prototype output stage, realized in a 180nm CMOS process, operates up to 30MHz and features the input voltage range of 1.8-3.6V and the output voltage range of 0.6-3.3V. A maximum output power of 6.6W and a peak power efficiency of 88.0% are achieved. Further, in comparison with the conventional approaches, the proposed dual-phase output stage achieves 13% smaller current ripple and 54% faster transient response. Yong Qu, Wei Shu, Joseph Sylvester Chang |
ISCAS | 3 |
| 2018 | Power-Loss and Design Space Analyses for Fully-Integrated Switched-Mode DC-DC ConvertersabstractPower-loss analyses for conventional (non fully-integrated) switched-mode dc-dc converters (SMCs) in the literature ignore the power losses due to the inductor and the gate driver of the power transistor. These losses are no longer negligible in fully-integrated SMCs. We propose closed-form expressions that characterize the aforesaid losses in fully-integrated buck SMCs. Using the expressions, we further analyze the overall power-efficiency, and show that designers can quickly find the combinations of design parameters that can achieve high power-efficiencies attainable by the SMC when realized using a target CMOS process. Yin Sun 0005, Victor Adrian, Joseph Sylvester Chang |
ISCAS | 3 |
| 2018 | Flexible Hybrid Electronics: Review and ChallengesabstractFlexible Hybrid Electronics (FHE), heterogeneous electronics embodying both conventional silicon electronics and printed electronics, is an emerging technology with huge market potential as it is advantageous compared to conventional silicon electronics and the emerging Printed Electronics - FHE features better mechanical flexibility/conformability and lower cost compared to conventional silicon electronics, and higher performance compared to Printed Electronics. In this paper, a comprehensive literature view on FHE is provided, including the state-of-the-art FHE development, FHE supply chains, and design challenges. Tong Ge, Zhou Jia, Joseph Sylvester Chang |
ISCAS | 3 |
| 2018 | Asynchronous-Logic QDI Quad-Rail Sense-Amplifier Half-Buffer Approach for NoC Router DesignabstractWe propose a low area overhead and power-efficient asynchronous-logic quasi-delay-insensitive (QDI) sense-amplifier half-buffer (SAHB) approach with quad-rail (i.e., 1-of-4) data encoding. The proposed quad-rail SAHB approach is targeted for area- and energy-efficient asynchronous network-on-chip (ANoC) router designs. There are three main features in the proposed quad-rail SAHB approach. First, the quad-rail SAHB is designed to use four wires for selecting four ANoC router directions, hence reducing the number of transistors and area overhead. Second, the quad-rail SAHB switches only one out of four wires for 2-bit data propagation, hence reducing the number of transistor switchings and dynamic power dissipation. Third, the quad-rail SAHB abides by QDI rules, hence the designed ANoC router features high operational robustness toward process-voltage-temperature (PVT) variations. Based on the 65-nm CMOS process, we use the proposed quad-rail SAHB to implement and prototype an 18-bit ANoC router design. When benchmarked against the dual-rail counterpart, the proposed quad-rail SAHB ANoC router features 32% smaller area and dissipates 50% lower energy under the same excellent operational robustness toward PVT variations. When compared to the other reported ANoC routers, our proposed quad-rail SAHB ANoC router is one of the high operational robustness, smallest area, and most energy-efficient designs. Weng-Geng Ho, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2018 | A Calibration-Free/DEM-Free 8-bit 2.4-GS/s Single-Core Digital-to-Analog Converter With a Distributed Biasing Scheme
F. N. U. Juanda, Wei Shu, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Review: A fully-additive printed electronics process with very-low process variations (Bent and unbent substrates) and PDKabstractDespite the huge market potential of Printed Electronics on Flexible substrates, the printed circuits and systems are yet to be manufacturable due to high cost fabrication processes, large process variations between devices, large and somewhat intractable variations when the devices and their substrates are bent, and lack of a comprehensive Process Development Kit (PDK). In this review paper, we review our Low-Cost Fully-Additive printing process with these four shortcomings addressed. To the best of our knowledge, this process is arguably the Fully-Additive printing process that is closest to being manufacturable and for the realization of practical intelligent printed electronics. Tong Ge, Jia Zhou 0002, Yang Kang, Joseph Sylvester Chang |
ISCAS | 4 |
| 2017 | A class-E RF power amplifier with a novel matching network for high-efficiency dynamic load modulationabstractWe present in this paper a proposed high-efficiency Class-E power amplifier (PA) for RF Polar transmitters. The PA embodies a proposed novel matching network (MN) with three salient features. First, it is digitally-controlled, and can directly receive the digital Amplitude Modulation (AM) input data to the PA without the need for a conventional supply modulator. Second, the MN performs high-efficiency dynamic load modulation, where load impedance seen by the PA is varied by the MN according to the AM data, and simultaneously, this load impedance is also ensured by the MN to satisfy the zero-voltage-switching condition at the PA to result in high-efficiency operation. Third, it has a novel architectural design that can employ on-, or off-chip inductors, or both inductor types; high quality-factor bond-wires can therefore be used as the inductors to improve the efficiency. The proposed PA with the MN is designed using a 40 nm CMOS technology. Simulation results at 2.4 GHz and 1.1 V supply show that the PA achieves a high power efficiency (drain efficiency) of 48% at peak output power of 17 dBm. Victor Adrian, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2017 | A novel high-rate hybrid window ADC design for monolithic digitally-controlled DC-DC convertersabstractWe propose a novel high-rate low-power hybrid window analog-to-digital converter (HWADC) for monolithic digitally-controlled switched-mode dc-dc converters. Conventional Window ADCs are generally based on either voltage-controlled delay lines or ring oscillators. These ADCs usually have a small window size (input voltage range) and a low sampling rate (<;10 MHz) in order to reduce the required IC area and the power dissipation. The proposed HWADC employs a novel hybrid architecture that is a hybrid of delay-lines and ring-oscillators. The HWADC can achieve a large window size with a very high conversion rate, a small IC area, and low power dissipation. Further, the HWADC operates entirely based on digital logic, and is simple to realize using digital cells. The proposed HWADC is designed using a 65 nm CMOS process. Its IC area is 0.005 mm2. Simulation results at 1.2 V supply show that the HWADC can achieve a window size up to 1.2 V (configurable from 0.7 V to 1.9 V) at 250 MHz conversion rate and ~770 μW power dissipation. At the maximum window size of 1.2 V, the quantization step is 50 mV, or equivalently, a resolution of ~4.5 bits. Yin Sun 0005, Victor Adrian, Joseph Sylvester Chang |
ISCAS | 3 |
| 2017 | Sense Amplifier Half-Buffer (SAHB) A Low-Power High-Performance Asynchronous Logic QDI Cell TemplateabstractWe propose a novel asynchronous logic (async) quasi-delay-insensitive (QDI) sense-amplifier half-buffer (SAHB) cell design approach, with emphases on high operational robustness, high speed, and low power dissipation. There are five key features of our proposed SAHB. First, the SAHB cell embodies the async QDI 4-phase (4φ) signaling protocol to accommodate process-voltage-temperature variations. Second, the sense amplifier (SA) block in SAHB cells embodies a cross-coupled latch with a positive feedback mechanism to speed up the output evaluation. Third, the evaluation block in the SAHB comprises both nMOS pull-up and pull-down networks with minimum transistor sizing to reduce the parasitic capacitance. Fourth, both the evaluation block and SA block are tightly coupled to reduce redundant internal switching nodes. Fifth, the SAHB cell is designed in CMOS static logic and hence appropriate for full-range dynamic voltage scaling operation for VDDranging from nominal voltage (1 V) to subthreshold voltage (~0.3 V). When six library cells embodying our proposed SAHB are compared with those embodying the conventional async QDI precharged half-buffer (PCHB) approach, the proposed SAHB cells collectively feature simultaneous -.64% lower power, -.21% faster, and ~6% smaller IC area; the PCHB cell is inappropriate for subthreshold operation. A prototype 64-bit Kogge-Stone pipeline adder based on the SAHB approach (at 65 nm CMOS) is designed. For a 1-GHz throughput and at nominal VDD, the design based on the SAHB approach simultaneously features -.56% lower energy and -.24% lower transistor count advantages than its PCHB counterpart. When benchmarked against the ubiquitous synchronous logic counterpart, our SAHB dissipates -.39% lower energy at the 1-GHz throughput. Kwen-Siong Chong, Weng-Geng Ho, Tong Lin 0001, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2017 | A 400-MS/s 10-b 2-b/Step SAR ADC With 52-dB SNDR and 5.61-mW Power Dissipation in 65-nm CMOSabstractWe present a single-channel 10-b 400-MS/s successive approximation register (SAR) analog-to-digital converter (ADC) embodying a proposed 2-b/step conversion scheme with single reference voltage for the IEEE 802.11ac. By means of the said scheme, the proposed ADC requires only three capacitor arrays instead of at least four capacitor arrays in other capacitor digital-to-analog converter-based 2-b/step SAR ADCs. The proposed ADC features a small input capacitance loading, thereby alleviating the driving requirement of the power-hungry input buffer in the IEEE 802.11ac system; and features a symmetrical architecture with highly matched interconnections. In addition, the proposed ADC embodies a proposed high-speed dynamic comparator with kickback noise cancelation and high-speed successive approximation (SA) control logic for high conversion rate and resolution. The proposed ADC prototype fabricated in 65-nm CMOS process achieves signal-to-noise-and-distortion-ratio >52 dB across 200-MHz Nyquist bandwidth, while dissipating 5.61-mW power. The ADC prototype, when benchmarked with state-of-the-art 2-b/step SAR ADCs, features a highly competitive figure-of-merit, i.e., 43 fJ/conv.step. Qing Liu 0005, Wei Shu, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Low normalized energy derivation asynchronous circuit synthesis flow through fork-join slack matching for cryptographic applications
Nan Liu 0002, Kwen-Siong Chong, Weng-Geng Ho, Bah-Hwee Gwee, Joseph Sylvester Chang |
DATE | 5 |
| 2016 | An investigation of THD of a BTL Class D amplifierabstractClass D amplifiers are routinely employed as audio amplifiers due to their high power efficiency. Total Harmonic Distortion (THD) is one of the most important parameters to qualify and quantify their performance. In this paper, THD of a commonly used Bridge-Tied-Load (BTL) Class D amplifier is investigated, including the derivation of the analytical expression for the THD of the BTL Class D amplifier. We show that in some cases, the THD of the BTL Class D amplifier is independent of the integrator gain of the amplifier - this is unlike other Class D amplifiers and linear amplifiers whose THD is largely determined by their integrator gain. Instead, the THD of the BTL Class D amplifier is largely determined by the feedback amplifier. Further, unlike the single-ended Class D amplifier whose THD increases as the modulation index increases, the THD of the BTL Class D amplifier is maximum at the modulation index = 0.5. The analysis herein provide good insight to the design of BTL Class D amplifiers, including how various parameters may be varied/optimized to meet a given THD specification. Tong Ge, Huiqiao He, Jia Zhou 0002, Yang Kang, Joseph Sylvester Chang |
ISCAS | 5 |
| 2016 | High performance low overhead template-based Cell-Interleave Pipeline (TCIP) for asynchronous-logic QDI circuitsabstractWe propose a novel Template-based Cell-Interleave Pipeline (TCIP) approach for generating high performance and yet low overhead asynchronous-logic (async) quasi-delay-insensitive (QDI) circuits. Our TCIP approach exploits the characteristics of the four prevalent QDI cell templates, namely Weak-Conditioned Half-Buffer (WCHB), Pre-Charged HalfBuffer (PCHB), Autonomous Signal-Validity Half-Buffer (ASVHB), and Sense-Amplifier Half-Buffer (SAHB), and then strategically interleave these template cells to form a composite pipeline. There are three main features in our TCIP approach. First, all QDI cell templates are first standardized with the same interface signals, and their corresponding cells are characterized in terms of transistor count, cycle time and energy dissipation for ease of comparison/selection/replacement. Second, our TCIP approach prioritizes the speed requirement when forming the initial pipeline circuits, and then subsequently reduces circuit overheads by interleaving various template cells without compromising the speed significantly. Third, the final optimized QDI pipeline circuit inherently features high robustness against process-voltage-temperature (PVT) variations, hence suitable for dynamic-voltage-scaling (DVS) operation. By means of 65nm CMOS process, we demonstrate a 4-bit pipeline tree adder based on the proposed TCIP approach, and benchmark it against the WCHB, PCHB, ASVHB and SAHB counterparts. These five designs feature same high operational robustness, nonetheless the design based on our TCIP approach is more competitive. Particularly, the designs based on reported approaches are, on average, ∼1.22× more transistor count, ∼1.21× slower and ∼1.22× higher energy dissipation. Furthermore, under DVS operation from 1.2V to 0.3V, our proposed TCIP adder can reduce up to 88% energy for non-speed critical applications. Weng-Geng Ho, Nan Liu 0002, Ne Kyaw Zwa Lwin, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 6 |
| 2016 | Total Ionizing Dose (TID) effects on finger transistors in a 65nm CMOS processabstractAlthough Total Ionizing Dose (TID) effects are generally unpronounced in deep-submicron-CMOS, we show the TID-induced leakage current @TID=500Krad is significant in NMOS-finger-transistors of GlobalFoundries 65nm CMOS. Further, Radiation-Hardening-By-Design techniques against said TID effect are recommended. Jize Jiang, Wei Shu, Kwen-Siong Chong, Tong Lin 0001, Ne Kyaw Zwa Lwin, Joseph Sylvester Chang |
ISCAS | 6 |
| 2016 | Experimental investigation into radiation-hardening-by-design (RHBD) flip-flop designs in a 65nm CMOS processabstractWe comprehensively study three types of radiation-hardened flip-flops: DICE for SEU-hardening, temporal for SET-hardening, and Triple-Modular-Redundancy for SEU-cum-SET-hardening. Our study includes their trade-offs of circuit/radiation-hardness attributes. We find that DICE flip-flops remain the most competitive. Tong Lin 0001, Kwen-Siong Chong, Wei Shu, Ne Kyaw Zwa Lwin, Jize Jiang, Joseph Sylvester Chang |
ISCAS | 6 |
| 2016 | Fully-additive printed electronics: Process Development KitabstractPrinted/Organic Electronics (PE) is an emerging technology with gargantuan market potential, particularly if its realization is low cost and the supply chain associated with its design is manageable/established. For the former, a Fully-Additive All-Air Low-Temperature process (vis-à-vis a Subtractive process) is desirable, while for the latter, a comprehensive Process Development Kit (PDK) for Electronic Design Automation tools is necessary. In this paper, we delineate, arguably the first-ever PDK for PE - this PDK is developed for our Fully-Additive All-Air Low-Temperature printing process with very-low process variations, and capable of complete circuit realizations on a myriad of substrates, including low-cost low-temperature flexible plastic films. Our proposed PDK embodies accurate modeling of the printed transistor and printed passive elements. With these models, comprehensive schematic simulations can be performed, including Monte Carlo simulations. We demonstrate the efficacy of our PDK, in part based on our measured process variations. The layout design rules for our Fully-Additive printing process to facilitate layout design - imperative for estimating the printing area and the printing cost - are also delineated herein. Jia Zhou 0002, Tong Ge, Joseph Sylvester Chang |
ISCAS | 3 |
| 2015 | High robustness energy- and area-efficient dynamic-voltage-scaling 4-phase 4-rail asynchronous-logic Network-on-Chip (ANoC)abstractWe propose an 18-bit 5-interface asynchronous-logic Network-on-Chip (ANoC) router based on the quasi-delay-insensitive (QDI) realization approach for high secured cryptography applications. There are four key features of the proposed ANoC router. First, it embodies the novel high-speed low-power Sense-Amplifier Half Buffer 4-rail cells. Second, it is designed based on QDI protocol, and hence is highly robust against process-voltage-temperature (PVT) variations. Third, it is functional for full dynamic voltage scaling from nominal (VDD=1.2V) to sub-threshold (VDD=0.3V) regions, and is potentially excellent for low power management applications. Fourth, it embodies a distributed-based XY routing algorithm to utilize a 4-bit header of flow control unit (flit) for routing up to 4×4 cluster, hence minimizing the routing overhead. We realize the proposed ANoC router (@65nm CMOS), and benchmark it against the reported ANoC router embodying the conventional Weak-Conditioned Half-Buffer (WCHB) QDI realization approach. Both our proposed and reported designs feature the high operation robustness, but our design is 41% more energy-efficient, and 21% more area-efficient than the reported counterpart. The prototype of ANoC router occupies only 0.105 mm2and can operate down to 0.3V. At VDD=0.3V, it dissipates 44 fJ per bit and operate 105 ns per flit. Weng-Geng Ho, Kwen-Siong Chong, Ne Kyaw Zwa Lwin, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 5 |
| 2015 | A novel subthreshold voltage reference featuring 17ppm/°C TC within -40°C to 125°C and 75dB PSRRabstractSubthreshold voltage references are increasingly prevalent in power-critical applications due to their low-voltage and ultra-low-power attributes. However, the effective temperature range of state-of-the-art subthreshold voltage references remains undesirably narrow for low temperature coefficient (TC) operation and/or their PSRR is low - thereby severely limiting their range of applications. In this paper, we present a novel subthreshold voltage reference embodying a novel paralleled `2-Transistor' structure and a novel auxiliary amplifier. The former serves to facilitate low TC within a wide temperature range, and the latter for high PSRR and low line-sensitivity. The proposed design achieves low TC of 17ppm/°C within a wide effective temperature range of -40°C to 125°C, and high PSRR of 75dB and low line-sensitivity of 0.3%/V with minimum 0.5V supply voltage and 32nW power consumption. Both the TC and PSRR parameters are, at this juncture, the best performance compared to reported subthreshold voltage references, but with a slight power penalty. Jize Jiang, Wei Shu, Joseph Sylvester Chang |
ISCAS | 3 |
| 2015 | Design of a variable-delay window ADC for switched-mode DC-DC convertersabstractWe propose a novel Variable-Delay Window ADC (VDWADC) design for digitally-controlled switched-mode dc-dc converters. In conventional Window ADCs based on the voltage-controlled delay line, the input voltage supplies the delay line. Thus, the conversion speed slows down when the input voltage decreases. The VDWADC is based on delay lines whose supply voltages are independent of the supply voltage. Hence, when the input voltage decreases, the conversion speed does not slow down. The VDWADC is simulated using 180 nm CMOS process and a supply voltage of 1.8 V. It achieves a quantization step of 0.05 V, or equivalently, a resolution of ~5.2 bits. Yin Sun 0005, Victor Adrian, Joseph Sylvester Chang |
ISCAS | 3 |
| 2015 | A single-VDD half-clock-tolerant fine-grained dynamic voltage scaling pipelineabstractWe propose a novel dynamic voltage scaling (DVS) pipeline with three significant attributes. First, it features a fine-grained DVS which innately attempts to power most of the circuits therein at low voltages, and when the speed is beneath the requirement, to scale up the voltage. Second, it supports fast-transition DVS within one-and-a-half clock duration per operation, and its operation remains error-free during that duration; we define such attribute as half-clock-tolerant. Third, it consists of a single power source (single-VDD) which supports three voltage scales (1.2V, 0.8V and 0.5V) for power/speed tradeoffs, and has standardized 1.2V output to seamlessly interface with other proposed/conventional pipelines. These attributes are achieved due to the embodiment of a DVS power unit, asynchronous building blocks to control/synchronize the operation, a dual-rail critical path to innately detect the completion of the operation, and level shifters to standardize the output voltage. We demonstrate our proposed pipeline by designing a multiplier embodied in a Fast Fourier Transform processor (@65nm CMOS). We show that the multiplier based on our proposed pipeline, on average, is 1.94× more power-efficient than that based on a conventional pipeline. Kwen-Siong Chong, Tong Lin 0001, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 5 |
| 2014 | A Randomized Modulation scheme for filterless digital Class D audio amplifiersabstractWe propose to employ the Randomized Wrapped-Around Pulse Position Modulation scheme (RWAPPM) to mitigate the switching-frequency harmonics at the output signal of filterless digital Class D audio amplifiers. The conventional Pulse Width Modulation schemes (PWMs) typically have a non-zero common-mode voltage that contributes to the radiated Electromagnetic Interference (EMI), and generate high switching-frequency harmonics that dissipate extra power at the speaker and also contribute to the radiated EMI. We simulate and compare the RWAPPM (2-level) against the PWMs and a reported randomized modulation scheme. The 2-level RWAPPM has zero common-mode voltage, and amongst the modulation schemes, it features the highest attenuation of the switching-frequency harmonics, highest out-of-band Spurious Free Dynamic Range (22 dBc), and a relatively high Signal to Noise and Distortion Ratio (57 dB) at the output voltage. Victor Adrian, Cui Keer, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2014 | Design of a 5 GS/s fully-digital digital-to-analog converterabstractWe present a fully-digital digital-to-analog converter (FD DAC) architecture design for high-speed communication systems. The FD DAC design is based on the ΔΣ modulation. The specifications for the DAC includes a low 1.2 V supply voltage, a high 5 GS/s input sampling rate, and a wide 2.5 GHz bandwidth. We employ a combination of the time-interleaving, parallel, and pipelining techniques to reduce the clock speed from 10 GHz to 625 MHz. The lower clock speed allows the use of standard cells for designing the digital computational circuits of the FD DAC. The critical building blocks of the FD DAC are laid-out in a 65 nm CMOS process. The post-layout simulation results show that the Signal to Noise and Distortion Ratio and the in-band Spurious-Free Dynamic Range of the output signal are 36 dB and 44 dBc respectively. Victor Adrian, Yin Sun 0005, Joseph Sylvester Chang |
ISCAS | 3 |
| 2014 | Synthesis of asynchronous QDI circuits using synchronous coding specificationsabstractWe propose a synthesis of asynchronous quasi-delay-insensitive (QDI) circuits. We highlight three notably features/novelties of the proposed synthesis as follows. First, the targeted synthesized circuits abide by the QDI protocol; hence they are inherently timing-robust and are desirable for applications with high variation-space and wide operation-space (including defense/space applications). Second, the coding specifications accept Verilog HDL language, and are the same/similar to the standard coding for synchronous circuits, hence no special and/or ad-hoc design/coding rules are required. Third, the proposed synthesis is applicable to accept various QDI library cells, hence enabling to explore full merit of different library cells. To the best of our knowledge, no reported synthesis methods incorporate all these features; some limited features were only incorporated. Our proposed synthesis, at this juncture, accepts three basic clauses - complete `if-else' clause, incomplete `if-else clause', and the `case' clause. These clauses are more than sufficient to describe any complex systems. The synthesis stages involve analyzing QDI pipelines, generating (corresponding) single-rail combinational circuits, converting dual-rail netlists (from the single-rail circuits), and embedding customized controllers. In order to demonstrate the validity and practicality of the proposed synthesis, an 8-bit 8-tap asynchronous QDI Finite Impulse Response (FIR) filter is synthesized, implemented to the layout stage, and evaluated using spice models-specifically, it features 3.7 mW power dissipation, 39,181 transistors, and a delay of 200 ns per operation. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang, Weng-Geng Ho |
ISCAS | 4 |
| 2014 | A Low Overhead Quasi-Delay-Insensitive (QDI) Asynchronous Data Path Synthesis Based on Microcell-Interleaving Genetic Algorithm (MIGA)abstractIn this paper, we propose a design approach to mitigate the hardware overhead of the data completion detection circuit in quasi-delay-insensitive (QDI) asynchronous-logic circuits. In this proposed design approach, three novelties are highlighted. Firstly, a novel microcell-interleaving approach is proposed to reduce the number of completion detection (CD) circuits while retaining the required QDI attribute. Secondly, we analyze the performance of the QDI circuits based on the proposed microcell-interleaving approach graphically in terms of power dissipation, transistor count and delay, and evaluate/determine the upper and lower boundaries of these performance profiles. Thirdly, we propose a microcell-interleaving genetic algorithm (MIGA) to stochastically optimize the proposed microcell-interleaving approach on power dissipation, transistor count, and delay. To validate the proposed design approach, a complete performance profile of ISCAS-85 C499 circuit is investigated on the basis of differential cascode voltage switch logic (DCVSL) and dynamic strong indicating (DSI) microcells. We demonstrate the efficiency of the proposed design approach by benchmarking against the competing DCVSL, null convention logic and DSI designs on five ISCAS-85 circuits. Specifically, the proposed designs, on average, are 1.77 × better in power dissipation, 1.4 × better in area, and 1.58 × better in a composite metric of power × area × delay, and reasonably slower for the lowest power dissipation points. We further demonstrate the practicality of the proposed design approach by implementing an 8-tap 16-bit asynchronous QDI finite impulse response filter. Finally, we demonstrate the ~10% and ~11% improved efficiency of the proposed MIGA over the greedy algorithm and dynamic programming, respectively. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2013 | A 250mV sub-threshold asynchronous 8051microcontroller with a novel 16T SRAM cell for improved reliability in 40nm CMOSabstractAsynchronous approach for digital systems is a way to resolve increased timing uncertainty with technology scaling since timing issue is eliminated in asynchronous systems. This paper presents a sub-threshold operating asynchronous 8051 microcontroller (A8051) with a novel 16T SRAM cell for improved reliability in asynchronous systems. This A8051, adopting a 4-phase dual-rail protocol, can operate up to 250 mV. A8051 has 67.53 μs as a critical path delay with 91.6 nW power consumption at 250 mV, which is equivalent to 12.88 kHz in synchronous systems. At 1.0 V, the delay of a critical path of A8051 microcontroller is 5.74 ns, which is equivalent to 151.55 MHz, with 8.98 mW power consumption. The proposed 16T SRAM cell is applied in memory blocks. The 16T SRAM structure eliminates charge contentions between devices during read and write operations so that SRAM can be operated fully in static mode, bringing about improved write margin (WM). The WM of this 16T SRAM cell is 1.81 times greater than the conventional 6T SRAM cell and 1.58 times better than 8T SRAM cell. At 250 mV, the SNM of SRAM cell is 12.5 mV under process and mismatch variations. Write delay of the asynchronous SRAM block is 4.02 μs (equivalent to 248.5 kHz) with 5.44 pJ energy dissipation, while read delay is 12.61 μs (equivalent to 79.3 kHz) with 9.08 pJ energy dissipation. Kwen-Siong Chong, Joseph Sylvester Chang, Pinaki Mazumder |
ACM Great Lakes Symposium on VLSI | 3 |
| 2013 | A dual-core 8051 microcontroller system based on synchronous-logic and asynchronous-logicabstractWe describe a dual-core 8051 microcontroller system featuring the synchronous and asynchronous (clockless) mode of operation. The synchronous mode of operation is achieved by means of a synchronous 8051 microcontroller core, while the asynchronous mode of operation is achieved by means of an asynchronous 8051 microcontroller core. The 8051 microcontroller system features shared embedded program and data memories that enable the switching between the two microcontroller cores during program execution. The measured energy, speed and electromagnetic interference of both microcontroller cores will be compared at different operation workloads. Kok-Leong Chang, Tong Lin 0001, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 6 |
| 2013 | Low power sub-threshold asynchronous QDI Static Logic Transistor-level Implementation (SLTI) 32-bit ALUabstractWe propose an asynchronous-logic (async) Quasi-Delay-Insensitive (QDI) Static Logic Transistor-level Implementation (SLTI) approach for low power sub-threshold operation. The approach is implemented to design 32-bit pipelined Arithmetic and Logic Units (ALUs), the primary computation core for microprocessors, and benchmarked against the reported Pre-Charged Half-Buffer (PCHB). There are two key attributes in this proposed design. First, the proposed SLTI ALU design can perform dynamic voltage scaling seamless by only changing the supply voltage from nominal (1V) to sub-threshold (~0.2V) regions for high speed/low power operation. Second, the ALU achieves ultra-low power dissipation (3.5μW) at the lowest VDDpoint (~0.15V). For fair of comparison, both implemented ALUs have identical functionality and functional blocks, are implemented using the same 65nm CMOS process. Based on the simulations, the minimum energy point occurs at VDD= 0.2V for SLTI-based ALU and at VDD= 0.3V for PCHB-based ALU. The SLTI-based ALU have ~93% and ~89% lower energy on the arithmetic and logic operations respectively from VDD= 1V to VDD= 0.2V. At VDD= 0.2V, with 9MHz input switching rate, the async ALU based on our proposed SLTI approach dissipates ~51% and ~44% lower power than the reported PCHB counterpart on the arithmetic and logic operations respectively. Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2012 | A comparative study on asynchronous Quasi-Delay-Insensitive templatesabstractThe robustness of asynchronous logic has proved useful in dealing with contemporary problems in CMOS design such as process variations and power management. However, the general cryptic nature of asynchronous logic has stymied the widespread acceptance of this alternate design technique. Fortunately, the semi-custom approach to asynchronous design reduces the tedious handcrafting efforts that are often non-trivial in large system-on-chips (SoCs). However, even with the adoption of this design approach requires careful selection of asynchronous templates that will suit overall system needs. Therefore in this paper, the most eminent Quasi-Delay-Insensitive asynchronous template families reported to date will be presented, and followed by an in-depth comparison of various design FOMs - template area, static/dynamic capacity, cycle time, latency, throughput and Et2. The most aggressive template (EESTFB) can reach a maximum throughput of 3.56Giga items/s on 0.13µm @ 1.2V. Kok-Leong Chang, Tong Lin 0001, Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 6 |
| 2012 | An Ultra-Dynamic Voltage Scalable (U-DVS) 10T SRAM with bit-interleaving capabilityabstractWe propose a dynamic voltage scalable SRAM capable of efficient bit-interleaving in column to tolerate multiple-bits soft error when integrated with error correction codes (ECC). First, a 10T SRAM bitcell is proposed. It activates only intended bitcells so that stability problem of half-selected bitcells is completely eliminated and the power dissipation in half-selected columns is significantly reduced. Second, a configurable DVS scheme is employed to enable the bitcell to operate like differential 8T during super-threshold region which results in faster operation. The proposed SRAM can operate up to 1.2GHz at 1.2V using 65nm CMOS process. Third, a segmented column multiplex with low overhead is proposed, which greatly reduces the power dissipation due to the column control signals. Consequently, the write and read power dissipations are reduced by up to 40% and 67% respectively. Forth, a hierarchical read bitline is used to reduce the read bitline discharge delay variation due to local and global process variation in subthreshold region, which is a major portion of memory access time. Based on our simulation results, the worst case read bitline discharge delay is reduced by more than 12× at VDDof 0.3V. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2012 | Energy-delay efficient asynchronous-logic 16×16-bit pipelined multiplier based on Sense Amplifier-Based Pass Transistor LogicabstractWe describe an asynchronous-logic (async) 16×16-bit pipelined multiplier based on our proposed Sense Amplifier-Based Pass Transistor Logic (SAPTL) with emphases on high energy-delay efficiency. The multiplier is targeted for an async multi-core System-On-Chip (SOC). This attribute is achieved by simplifying and optimizing the NMOS pass transistor stacks and decision-making C-element, therein to reduce the circuit area overheads and transistor switchings in SAPTL. Based on the simulations (@1V, 65nm CMOS process), the async 16×16-bit pipelined multiplier based on our proposed SAPTL approach features, on average, 31% shorter delay, 21% lower energy/operation achieving a total of 46% lower energy-delay product, and 16% lesser number of transistors when compared to the reported SAPTL approaches. Weng-Geng Ho, Kwen-Siong Chong, Tong Lin 0001, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 5 |
| 2011 | A low-power dual-rail inputs write method for bit-interleaved memory cellsabstractWe propose a dual-rail data write technique for bit interleaved memory cells to reduce power dissipation for the write operation without affecting the read operation. The proposed technique can be applied to two reported bit interleaved memory cells with a write power reduction range from 30% to 45%, depending on memory cells and operations. In addition, in the proposed technique, a subthreshold non bit interleaved memory cell is modified to be bit-interleaved without increasing the number of transistors in memory cell. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2011 | Improved asynchronous-logic dual-rail Sense Amplifier-based Pass Transistor Logic with high speed and low power operationabstractWe propose a robust asynchronous-logic dual-rail Sense Amplifier-based Pass Transistor Logic (SAPTL) approach with improved speed and power attributes over reported SAPTL approach. These attributes are achieved by simplifying various sub-blocks therein to reduce the stacking of pass transistors and the number of transistor switchings, and to avoid floating nodes. By means of an 8-bit pipeline adder and on the basis of computation simulations (@ 1V, 45nm SOI process), we show that our proposed SAPTL adder is 37% faster, yet 14% lower power dissipation (@ 200MHz input-rate), 18% lower energy dissipation (per operation), and 47% better energy-delay product. These substantially improved attributes are achieved with insignificant overhead - just 3% more transistors. Weng-Geng Ho, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang, Yin Sun 0005, Kok-Leong Chang |
ISCAS | 4 |
| 2011 | Modeling and Synthesis of Asynchronous PipelinesabstractWe propose a set of modeling rules and a synthesis method for the design of asynchronous pipelines. To keep the circuit area and power dissipation of the asynchronous control network small, the proposed approach avoids the conventional syntax-directed translation approach. Instead, it employs a data-driven design style and a coarse-grain approach to the synthesis of asynchronous control, restricting asynchronous control to the implementation of communication channels commonly found in asynchronous pipelines and operations involving these channels. The proposed approach integrates well into conventional synchronous design flows because they are based on Verilog and SystemVerilog specifications, and generate register-transfer level models suitable for functional simulation and logic synthesis using existing computer-aided design tools. Using a 32-bit microprocessor, an interpolated finite-impulse-response filter bank, and a Reed-Solomon error detector as design examples, we show that the proposed approach is competitive with other comparable reported methods. Chong-Fatt Law, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | A micropower comparator for high power-efficiency hearing aid class D amplifiersabstractClass D amplifiers are routinely employed in power-critical hearing aids for their high power-efficiencies, typically ~90% at high modulation indexes. Nevertheless, at the nominal operation condition where the modulation index M=0.1, the power-efficiency is typically 32%. In this paper, we show that the comparator embodied in the over-current protection circuit of the Class D output stage is dominant at M=0.1. We propose the design of a novel comparator (with performance parameters comparable with conventional comparators) featuring ~40% lower power dissipation. This improved lower power dissipation translates to a worthwhile 12% improvement in the power-efficiency of the Class D amplifier output stage. The proposed design herein is verified by computer simulations. Linfei Guo, Tong Ge, Joseph Sylvester Chang |
ISCAS | 3 |
| 2009 | Fine-grained Power Gating for Leakage and Short-circuit Power Reduction by using Asynchronous-logicabstractIn this paper, a fine-grained power gating technique for an asynchronous-logic pipeline stage is proposed using locally controlled gating transistors. The proposed power gating technique is implemented with minimal control overheads (one additional inverter per pipeline stage for driving PMOS Gating) and delay overheads (within 15% more than the conventional asynchronous-logic pipeline stage). Different types of gating configurations using only PMOS transistor (PMOS Gating), only NMOS transistor (NMOS Gating), and both types of transistors (Dual Gating) are examined and compared. The effectiveness of the proposed power gating technique to the Combinational Block therein with different data input rates is investigated. Based on the computer simulation results, we have found that ≫70% wasted power reduction (including both short-circuit and leakage powers) as compared to the conventional asynchronous-logic pipeline stage can be achieved with all gating configurations. In particular, Dual Gating achieves the best wasted power reduction of 86% for short-circuit power and 99% for leakage power @ 10Mbps input rate. Tong Lin 0001, Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 2009 | A Low THD Analog Class D Amplifier based on Self-oscillating Modulation with Complete Feedback NetworkabstractWe propose a novel analog Class D Amplifier (CDA) based on self-oscillating modulation where the input of the feedback network is taken at the output of the lowpass LC filter (‘complete feedback’) as opposed to the prevalent (Pulse Width Modulation CDA) approach where the feedback input is taken at the output of the output stage of the CDA (‘incomplete feedback’). The complete feedback approach substantially suppresses the non-linearity of the inductor of the Lowpass filter as this filter now constitutes part of the feedback loop. The proposed approach is highly hardware efficient over reported prevalent complete feedback CDAs. When the proposed complete feedback CDA is compared against the prevalent incomplete feedback approach, the proposed CDA exhibited highly suppressed THD (∼40dB relative suppression in most cases). The THD of the proposed CDA is ≤0.012% for the complete modulation index range and pertinent input signal range. Means to further reduce the THD is also suggested. Wenfeng Yu, Wei Shu, Joseph Sylvester Chang |
ISCAS | 3 |
| 2008 | PSRR of bridge-tied load PWM Class D AmpsabstractIn this paper, the effects of power supply noise, qualified by power supply rejection ratio (PSRR), on two types of bridge-tied load (BTL) pulse width modulation (PWM) class D amps (denoted as Type-I BTL and Type-II BTL respectively) are investigated and the analytical expressions for PSRR of the two designs derived. The derived analytical expressions are verified by means of HSPICE simulations. The relationships derived herein provide good insight to the design of BTL class D amps, including how various parameters may be varied/optimized to meet a given PSRR specification. Furthermore, the PSRR of the two BTL Class D amps are compared against the single-ended class D amp, and the former designs show superior PSRR compared to the latter. Tong Ge, Joseph Sylvester Chang, Wei Shu |
ISCAS | 2 |
| 2008 | Asynchronous Control Network Optimization Using Fast Minimum-Cycle-Time AnalysisabstractThis paper proposes two methods for optimizing the control networks of asynchronous pipelines. The first uses a branch-and-bound algorithm to search for the optimum mix of the handshake components of different degrees of concurrence that provides the best throughput while minimizing asynchronous control overheads. The second method is a clustering technique that iteratively fuses two handshake components that share input channel sources or output channel destinations into a single component while preserving the behavior and satisfying the performance constraint of the asynchronous pipeline. We also propose a fast algorithm for iterative minimum-cycle-time analysis. The novelty of the proposed algorithm is that it takes advantage of the fact that only small modifications are made to the control network during each optimization iteration. When applied to nontrivial designs, the proposed optimization methods provided significant reductions in transistor count and energy dissipation in the designs' asynchronous control networks while satisfying the throughput constraints. Chong-Fatt Law, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2007 | A Low Energy FFT/IFFT Processor for Hearing AidsabstractWe present a 16-bit low voltage (1.1V - 1.4V) energy efficient 128-point decimation-in-time Fast Fourier Transform/Inverse Fast Fourier Transform (FFT/IFFT) processor specifically for hearing aid applications. The FFT/IFFT processor embodies several low power/energy methodologies, including the clock gating approach,ad-hoccontrol, operand isolation and low power library cells, to satisfy the tight constraints of low voltage low energy and a small silicon area realization for a practical hearing aid. Based on the prototype IC measurements, the proposed FFT/IFFT processor dissipates ~ 188nJ @ 1.1V, features computation delay of2@ 0.35μm CMOS process. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 3 |
| 2007 | Power Supply Noise in Bang-Bang Control Class D AmplifierabstractIn this paper, the effect of power supply noise, qualified by power supply rejection ratio (PSRR), on bang-bang control class D amplifier (BCCD) is investigated. By modeling the BCCD with power supply noise, the mechanisms are modeled. Further, by means of a novel analysis based on a double-Fourier series, the parameters related to power supply noise is derived and thereafter expressions for PSRR are derived. The analyses are verified by means of computer simulations and by measurements on practical circuits. The relationships derived herein provide good insight to the design of BCCDs, including how various parameters may be varied/optimized to meet a given PSRR specification. Tong Ge, Joseph Sylvester Chang, Wei Shu |
ISCAS | 2 |
| 2006 | An acoustic noise suppression system with reduced musical artifactsabstractIn this paper, we propose an acoustic noise suppression system with reduced musical artifacts for digital hearing instruments (aids). The proposed system features the capabilities to detect, estimate and suppress the acoustic noise corrupting an input speech. The algorithms in the system consist of two noise detections, an enhanced parametric spectral subtraction, a noise attenuation and transition smoothing window, and an automatic gain control. Simulation results on several stationary and nonstationary noise show that our acoustic noise suppression system is capable of improving signal-to-noise ratio by > 9 dB and reducing musical artifacts. Victor Adrian, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 3 |
| 2006 | Modeling and analysis of PSRR in analog PWM class D amplifiersabstractIn this paper, we propose and derive a linear model of a closed-loop PWM class D amplifier (CDA) to model its power supply rejection ratio (PSRR). On the basis of the proposed model, we derive a simple expression that depicts the effects of 3 critical parameters on PSRR: the gain of the integrator, the gain of the PWM stage and the feedback factor. We recommend, from a practical viewpoint, that the first and third parameters be increased if higher PSRR is desired. We validate our model and analysis on the basis of HSPICE simulations and on experimental measurements. Our model and analysis are useful as they provide insight to a CDA designer, in particular how various parameters may be varied/compromised to meet a given set of PSRR specifications Tong Ge, Joseph Sylvester Chang, Wei Shu |
ISCAS | 2 |
| 2006 | Fourier series analysis of the nonlinearities in analog closed-loop PWM class D amplifiersabstractWe derive Fourier series expressions for modeling analog closed-loop PWM class D amplifiers. Based on our derived expressions, we are able to analyze the mechanisms of several nonlinearities, specifically the power supply noise and fold-back distortions. We investigate the influence of the finite loop gain on these nonlinearities, and we show that those nonlinearities can be effectively suppressed by the higher loop gain. We verify our analysis by computer simulations and also on the basis of experimental measurements. The derived expressions are practically useful as they provide insights to a designer for PWM class D amplifier designs Wei Shu, Joseph Sylvester Chang, Tong Ge, Meng Tong Tan |
ISCAS | 2 |
| 2005 | A micropower low-voltage multiplier with reduced spurious switchingabstractWe describe a micropower 16/spl times/16-bit multiplier (18.8 /spl mu/W/MHz @1.1 V) for low-voltage power-critical low speed (/spl les/5 MHz) applications including hearing aids. We achieve the micropower operation by substantially reducing (by /spl sim/62% and /spl sim/79% compared to conventional 16/spl times/16-bit and 32/spl times/32-bit designs respectively) the spurious switching in the Adder Block in the multiplier. The approach taken is to use latches to synchronize the inputs to the adders in the Adder Block in a predetermined chronological sequence. The hardware penalty of the latches is small because the latches are integrated (as opposed to external latches) into the adder, termed the latch adder (LA). By means of the LAs and timing, the number of switchings (spurious and that for computation) is reduced from /spl sim/5.6 and /spl sim/10 per adder in the adder block in conventional 16/spl times/16-bit and 32/spl times/32-bit designs respectively to /spl sim/2 in our designs. Based on simulations and measurements on prototype ICs (0.35 /spl mu/m three metal dual poly CMOS process), we show that our 16/spl times/16-bit design dissipates /spl sim/32% less power, is /spl sim/20% slower but has /spl sim/20% better energy-delay-product (EDP) than conventional 16/spl times/16-bit multipliers. Our 32/spl times/32-bit design is estimated to dissipate /spl sim/53% less power, /spl sim/29% slower but is /spl sim/39% better EDP than the conventional general multiplier. Kwen-Siong Chong, Bah-Hwee Gwee, Joseph Sylvester Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2003 | A Hybrid Genetic Hill-climbing Algorithm for Four-Coloring Map Problems
Bah-Hwee Gwee, Joseph Sylvester Chang |
HIS | 2 |
| 2000 | An investigation on the parameters affecting total harmonic distortion in class D amplifiersabstractIn this paper, we investigate two important and practical design parameters for the design of low-voltage low-power class D amplifiers that may affect the Total Harmonic Distortion (THD): the linearity of the carrier waveform and the impedance of the output stage. By means of a novel mathematical analysis method to model the carrier non-linearity, we show that this non-linearity should be mitigated to achieve low THD. Our mathematical analysis also provides an insight to the degree of non-linearity acceptable for a practical design. We show that the impedance of the output stage has little effect on THD. We verify our analysis by means of MATLAB and SPICE simulations. Meng Tong Tan, Hock-Chuan Chua, Bah-Hwee Gwee, Joseph Sylvester Chang |
ISCAS | 4 |
| 1998 | A parametric formulation of the generalized spectral subtraction methodabstractIn this paper, two short-time spectral amplitude estimators of the speech signal are derived based on a parametric formulation of the original generalized spectral subtraction method. The objective is to improve the noise suppression performance of the original method while maintaining its computational simplicity. The proposed parametric formulation describes the original method and several of its modifications. Based on the formulation, the speech spectral amplitude estimator is derived and optimized by minimizing the mean-square error (MSE) of the speech spectrum. With a constraint imposed on the parameters inherent in the formulation, a second estimator is also derived and optimized. The two estimators are different from those derived in most modified spectral subtraction methods, which are predominantly nonstatistical. When tested under stationary white Gaussian noise and semistationary Jeep noise, they showed improved noise suppression results. Boh Lim Sim, Yit Chow Tong, Joseph Sylvester Chang, Chin-Tuan Tan |
IEEE Trans. Speech Audio Process. | 3 |