EDBT 2026 Demo / reviewers in the wild / expert
Anh-Tuan Do
dblp:33/9070 · also Anh Tuan Do
· DBLP profile ↗
50ranked-venue papers
11as first author
31since 2021 · last 2026
0000-0002-8320-6818ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 47 · 11 first-author · 28 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Time-Based Sensing With Linear Current-to-Time Conversion for Multi-Level Resistive MemoryabstractResistive Random Access Memory (RRAM) is a promising low-power memory candidate because of a large R-ratio (RHRS/RLRS). Multi-level RRAM cells have been investigated to improve memory density and cost-per-bit. However, sensing multi-level becomes challenging due to the smaller R-ratios between stored digital values. This paper introduces a novel time-based sensing (TBS) scheme for enhancing the sensing speed and robustness for single-level cells (SLC) to multi-level cells (MLC). The proposed time-based sensing scheme converts the bit line (BL) current into a time delay using a novel current-to-time converter (CTC). The BL-current-dependent time delays are utilized to generate digital data. In addition, the proposed sensing scheme executes sensing without need for reference current or reference voltage. Comprehensive simulation in 40nm CMOS technology shows that the proposed TBS scheme achieves better linearity and higher read speed by precise cell current replication in CTC compared to the prior TBS schemes. As a result, the proposed TBS reduces sensing latency by 230%~340% compared to the prior TBS schemes. Furthermore, the proposed TBS improves the variation tolerance of read operation by 5%~33% at 1.1 V. Byung-Kwon An, Xueyong Zhang, Anh-Tuan Do, Tony Tae-Hyoung Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | Low-Power Learnable Digital Audio Feature Extractor for Always-on Keyword Spotting in Edge DevicesabstractThis paper presents a low-power, learnable digital audio feature extractor (AuFEx) for always-on keyword spotting (KWS) in edge devices. A noise-aware design flow is introduced to integrate the tuning of AuFEx parameters directly into the neural network classifier’s training process. This co-design approach enables the joint optimization of feature extractor and classifier within a unified framework. By incorporating noise during the training process, the system becomes more robust to variations in signal-to-noise ratio (SNR), maintaining high inference accuracy even with lightweight neural network classifiers. This design flow supports the design of both time-domain AuFEx (TD-AuFEx) and frequency-domain AuFEx (FD-AuFEx) for single-keyword wake-word detection (WWD) and 10-keyword KWS tasks, respectively. Implemented in a 40nm CMOS process, both TD-AuFEx and FD-AuFEx achieve over 2% accuracy improvement with the smallest size backend classifier. Specifically, the TD-AuFEx for WWD task achieves classifier size of 3.1k parameters with 494 nW power consumption and$375~\mu $s latency. The accuracy is maintained between 96.2% – 97.9% for SNR in the range of 5 – 20 dB. The FD-AuFEx for 10-keyword KWS achieves classifier size of 6.39k parameters with$1.258~\mu $W power consumption and 34.625 ms latency. The accuracy is maintained between 87.5% – 92.2% for SNR in the range of 5 dB - 20 dB, which is one of the highest compared to the other state-of-the-art designs. Jinhai Hu, Wang Ling Goh, Yi Sheng Chong, Anh-Tuan Do, Yuan Gao 0011 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Enhancing Phishing URL Detection with Graph Neural Networks: A Combination of URL and HTML Features
Thi Le Vuong, Anh-Tuan Do, Van Hoan Do |
ICCCI (2) | 2 |
| 2025 | A 23.5 TOPS/W Depthwise Separable Convolution Accelerator for Event-based Depth EstimationabstractRecent efforts to improve energy efficiency in computer vision (CV) tasks, such as depth estimation, focus on integrating event-based cameras with lightweight networks using Depthwise Separable (DWS) Convolutions. Despite remarkable accuracies and hardware-friendly binary signals, existing accelerators have not fully leveraged this combination. This paper proposes a Separated Engine (SE) architecture for binary input feature maps (ifmaps) with dedicated arrays for each DWS stage, eliminating data storage between stages and enhancing array utilization. Additionally, an integrative dataflow incorporating Row, Weight, and Input Stationary (RS, WS, and IS) advantages is introduced to maximize data reuse, supported by an optimized mapping strategy that efficiently loads and updates ifmaps onto the first array. Our gate-level simulation results using 28nm CMOS technology demonstrated a 1.7x improvement in energy efficiency, achieving 23.5 TOPS/W, along with a 13x enhancement in area efficiency, reaching 1785.5 GOP/mm2. Andres Brito, Tomomasa Yamasaki, Ulysse Rançon, Timothée Masquelier, Benoit Cottereau, Anh-Tuan Do, Bo Wang 0020 |
ISCAS | 6 |
| 2025 | Accuracy-preserving Layer Normalization Approximations for Efficient Transformer Hardware AcceleratorsabstractTransformers have emerged as leading models for solving natural language processing (NLP) problems. Layer normalization (LN) has been found to be a throughput and latency bottleneck for Transformer networks. The LN datapath involves sequential processing that is dependent on the input data, and GPU and CPU unfriendly nonlinear square root and reciprocal operations. These make pipelining and hardware simplification challenging without compromising the accuracy of pretrained models. In this paper, we propose a hardware-efficient and accuracy-preserving design for LN approximation. The LN core can be directly plugged into pretrained Transformer networks to accelerate NLP tasks without fine tuning. The FPGA implementation of our design outperforms the state-of-the-art FPGA implementation of LN, with 3.5 × higher throughput per LUT and 1.3 × higher throughput per DSP. Our approximation introduces negligible average and worst-case accuracy drops of only 0.28% and 1.09%, respectively on the main language benchmark tasks. Additionally, a better pipelined design with pairwise variance calculation is proposed to reduce the number of iterations for LN, resulting in a further 27% reduction in latency with a slight increase in hardware resource requirement. Nazim Altar Koca, Anh-Tuan Do, Chip-Hong Chang |
ISCAS | 2 |
| 2025 | Qubit-State Discrimination using Neural Networks with Rapid and Energy-Efficient Compute ArraysabstractNeural networks (NNs) implemented on field-programmable gate arrays (FPGAs) provide fast, high-fidelity solutions for processing readout signals from quantum information processors. However, application-specific integrated circuits (ASICs) instead of FPGAs hold the potential for improved performance, a largely unexplored path. This work proposes specialized hardware for NN-based qubit-state discrimination. We optimize the NN architecture to minimize resource requirements by reducing the layer width, employing linear activation functions, and weight quantization. Quantization-aware training is used to preserve accuracy despite these optimizations. Next, a compute array employing output stationary dataflow is chosen to process the NN workload. The compute array with abundant multipliers and adders can complete one NN inference in 63 ns, which makes it a good candidate for real-time qubit-state discrimination. Yuntian Liu, Yi Sheng Chong, Benjamin Lienhard, Minghao Fan, Wang Ling Goh, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 7 |
| 2025 | Design and Analysis of Latch-Type Comparators for Cryogenic OperationsabstractA low-power, high-speed latched comparator is the most crucial component in high-speed analog-to-digital converters (ADCs). With recent advances in quantum sensing and quantum computing, these ADCs need to operate at cryogenic temperatures. However, the performance of existing latched-comparators at cryogenic temperatures has not been thoroughly evaluated. In this paper, we address this gap by characterizing and comparing five representative comparator architectures. We elucidate their operational principles at both cryogenic (4K) and room (300K) temperatures, utilizing TSMC 28nm CMOS technology. Our study showed that, the triple latch design achieves the best performance at cryogenic temperature, while at room temperature, the modified strong-regeneration design has the best performance. Qibang Zang, Wang Ling Goh, Andres Brito, Anh-Tuan Do |
ISCAS | 4 |
| 2024 | PACE: A Scalable and Energy Efficient CGRA in a RISC-V SoC for Edge Computing Applicationsabstract▪Coarse-grained reconfigurable arrays (CGRAs) deliver high energy efficiency while maintaining the programmability advantages. ▪CGRA is the ideal candidate for efficiently handling loop kernels, which allows it to offload repetitive looping functions such as vector multiplication or hashing algorithms from CPUs. ▪It relies on a compiler to convert a given workload into a data flow graph (DFG) which is then mapped onto the hardware in a manner that achieves the highest possible energy efficiency. Vishnu P. Nambiar, Yi Sheng Chong, Thilini Kaushalya Bandara, Dhananjaya Wijerathne, Zhaoying Li 0004, Rohan Juneja, Li-Shiuan Peh, Tulika Mitra, Anh-Tuan Do |
HCS | 9 |
| 2024 | Time-based Sensing with Linear Current-to-Time Conversion for Multi-level Resistive MemoryabstractResistive Random Access Memory (RRAM) is a promising low-power memory candidate because of a large R-ratio (RHRS/RLRS). Multi-level RRAM cells have been investigated to improve memory density and cost-per-bit. However, sensing multi-level becomes challenging due to the smaller R-ratios between stored digital values. This paper introduces a novel time-based sensing (TBS) scheme for enhancing the robustness and speed of the read operation, supporting from single-level cells (SLC) to multi-level cells (MLC). In the proposed time-based sensing scheme, the current-to-time converter (CTC) converts the bit line (BL) current into a time delay based on the cell states. Different time delays from HRS and LRS values are compared to generate digital data. Unlike conventional current-based sense amplifiers (CSA) or voltage-based sense amplifiers (VSA), the proposed sensing scheme does not require a reference array or generator. Comprehensive simulation in 40nm CMOS technology shows enhanced linearity and higher read speed because of the precise replication of cell current to CTC. As a result, the proposed TBS achieves a sensing latency of < 2ns with an energy consumption of 61fJ/bit for read operations at 1.1 V and also supports MLC sensing through a single read operation. Byung-Kwon An, Xueyong Zhang, Anh-Tuan Do, Tony Tae-Hyoung Kim |
ISCAS | 3 |
| 2024 | Quantum Readout Processing Accelerator with a CORDIC Core at Cryogenic TemperatureabstractQuantum computing has been the most promising in addressing computational challenging problems beyond classical computers’ capabilities. Recent work focus on developing systemon-chip (SoC) for quantum control and readout, yet there are scarce literature discussing how the measurement blocks are developed in details. This paper presents a quantum readout processing accelerator that operates at cryogenic temperature to process measurement data directly to obtain the qubit state. The proposed accelerator includes a coordinate rotation digital computer (CORDIC) based direct digital frequency synthesizer (DDFS) and a qubit state decision unit. To enable compact processing accelerator, the error of the qubit state prediction is analyzed when choosing different bitwidths and number of CORDIC iterations, aiming to balance the accuracy and hardware cost. The proposed accelerator, implemented using a 28-nm CMOS, consumes 36mW of power at room temperature and 8.9mW at −196°C. The processing latency and total energy consumption when determining a qubit state are reduced by 6 and 3 orders of magnitude respectively, when compared to a general purpose processor. Yi Sheng Chong, Hongyu Cao, Wang Ling Goh, Patrick Bore, Yuanzheng Paul Tan, Yung Szen Yap, Rainer Dumke, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 9 |
| 2024 | A 420 GOPS/W CGRA with a Configurable MAC and Dynamic TruncationabstractEdge devices demand for highly efficient yet flexible processing capability to handle dynamic real-time workloads. Coarse grain reconfigurable architecture (CGRA) emerges as a suitable accelerator candidate in edge devices, because they are as flexible as general purpose processors and offer high efficiency close to that of domain specific accelerators. However, a typical CGRA requires two cycles for a multiply-and-accumulate (MAC) operation, and workloads such as neural network inference and signal processing involve many MAC operations, resulting in long CGRA processing time. This work proposes a CGRA that has configurable MAC units in the processing elements (PEs) that can perform an addition (ADD) or multiplication (MUL) or a MAC by using the same multiplier and adder, in a single cycle. The readout precision of MAC result can be adjusted by a truncation block. The proposed CGRA is implemented with 40nm CMOS technology. It attains an energy efficiency of 420.6GOPS/W operating at supply of 0.6V and frequency of 21MHz, which is 1.4 times higher than the state-of-the-art. Yi Sheng Chong, Rakshith Harish, Rajesh Chandrasekhara Panicker, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 5 |
| 2024 | Exploring Error Correction Circuits on RISC-V based Systems for Space ApplicationsabstractRISC-V systems are becoming increasingly adopted in space applications. SRAM (Static Random-Access Memory) data memory is a critical component that occupies a large portion of the processor peripheral system. SRAM is vulnerable to single-event upsets (SEUs). Existing studies mainly considered Hamming error correction codes (ECCs) for memory protection in RISC-V processor. In this paper, we explore different ECCs as well as the triple modular redundancy (TMR) as solutions to mitigate SEU effects on SRAMs for RISC-V system. To overcome the area and power overhead of TMR with higher fault tolerance than ECCs, we propose a dual modular redundancy (DMR) with ECC memory protection scheme. We conduct a comprehensive error analysis to evaluate different fault-tolerant designs under various radiation attack scenarios, utilizing real-world data to devise the fault-injection campaign. An efficient scheduling for the self-refresh operation of SRAMs is proposed to prevent the error accumulation. The proposed DMR with ECC design reduces the power and area overhead of TMR by 28% and 11% respectively and improve the error resilience of the SRAM significantly compared with Hsiao and Hamming ECC schemes. Nazim Altar Koca, Chip-Hong Chang, Anh-Tuan Do, Vishnu P. Nambiar |
ISCAS | 3 |
| 2024 | 3881 Gbps/W, 3005 µm AES Core with State Based Clock Gating for IoT applicationsabstractAn efficient and extremely low-energy AES-256 accelerator was implemented in a 40nm CMOS process for space and energy-limited IoT applications, supporting encryption, decryption, ECB and CBC modes. By reusing functional block hardware and utilizing scan flip-flops, the proposed AES design occupies merely 3005 µm of silicon area. Additionally, clock gating circuits were extensively used at circuit level to halt essential flip-flops during the ShiftRows and its reversed operation for aggressive power saving. The post-layout simulation showed that our proposed design achieved the lowest power consumption of 1.83 µW at 0.3 V, 17 MHz, with energy efficiency as high as 3881 Gbps/W. Zhangyi Pei, Vishnu P. Nambiar, Yi Sheng Chong, Wang Ling Goh, Anh-Tuan Do |
ISCAS | 5 |
| 2024 | 1.63 pJ/SOP Neuromorphic Processor With Integrated Partial Sum Routers for In-Network ComputingabstractNeuromorphic computing is promising to achieve unprecedented energy efficiency by emulating the human brain’s mechanism. Conventional neuromorphic accelerators employ split-and-merge method to map spiking neural networks’ inputs to surpass the fan-in capabilities of a single neuron core. However, this approach gives rise to the risk of accuracy compromise and extra core usage for the merging process. Moreover, it requires excessive data movement and clock cycles to aggregate spikes generated by partial sums instead of total sums obtained from different cores with substantial power and energy overhead. This work presents a novel approach to addressing the challenges imposed by the split-and-merge method. We propose an energy-efficient, reconfigurable neuromorphic processor that leverages several key techniques to mitigate the above issues. First, we introduce a partial sum router circuitry that enables in-network computing (INC), eliminating the need for extra merge cores. Second, we adopt software-defined Networks-on-Chip (NoCs) by leveraging predefined, efficient routing, eliminating power-hungry routing computation. At last, we incorporate fine-grained power gating and clock gating techniques for further power reduction. Experimental results from our test chip demonstrate the lossless mapping of the algorithm and exceptional energy efficiency, achieving an energy consumption of 1.63 pJ/SOP at 0.48 V. This energy efficiency represents a 22.4% improvement compared to the state-of-the-art results. Our proposed neuromorphic processor provides an efficient and flexible solution for neural network processing, mitigating the limitations of the traditional split-and-merge approach while delivering superior energy efficiency. Dongrui Li, Ming Ming Wong, Yi Sheng Chong, Jun Zhou 0014, Mohit Upadhyay, Ananta Narayanan Balaji, Aarthy Mani, Weng-Fai Wong, Li-Shiuan Peh, Anh-Tuan Do, Bo Wang 0020 |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 2023 | A 129.83 TOPS/W Area Efficient Digital SOT/STT MRAM-Based Computing-In-Memory for Advanced Edge AI ChipsabstractThis paper proposes a spin-orbit torque (SOT) magnetoresistive random access memory (MRAM)-based digital computing in memory (CIM) structure for advanced CIM edge AI chips. To avoid frequent data reloading and reduce the area overhead caused by the large transistor in the write path, 11 transistors and 4 shared heavy metal (HM) SOT MRAM bitcell is proposed. It contains 4b weight in a single cell in a compact area able to hold a large capacity with reduced latency. Compared to the previous SRAM+NOR digital CIM design, the SOT/STT MRAM CIM designs occupy only 40% and 30% bitcell area respectively. Additionally, SOT MRAM has a 1.16% leakage current of SRAM at the TT corner at room temperature. The proposed design is verified using statistic simulations in 28nm technology. It achieves 129.83 TOPS/W at 1b/1b/8b precision. Lu Lu 0013, Aarthy Mani, Anh-Tuan Do |
ISCAS | 3 |
| 2023 | 1.7pJ/SOP Neuromorphic Processor with Integrated Partial Sum Routers for In-Network ComputingabstractConventional neuromorphic accelerators primarily leverage split-merge method to accommodate a neural network that is beyond a single core's size, leading to possible accuracy loss, extra core usage and significant power and energy overhead. This work presents an energy-efficient, reconfigurable neuro-morphic processor to address the problem by (i) a partial sum router circuitry that enables in-network computing to remove the need of extra merge cores; (ii) software-defined Networks-on-Chip that eliminates the power-hungry routing compute and (iii) fine-grained power gating and clock gating technique for power reduction. Our test chip achieves lossless mapping as the algorithm and an energy efficiency of 1.7pJ/SOP at 0.5V, 19% lower than state-of-the-art result. Bo Wang 0020, Ming Ming Wong, Dongrui Li, Yi Sheng Chong, Jun Zhou 0014, Weng-Fai Wong, Li-Shiuan Peh, Aarthy Mani, Mohit Upadhyay, Ananta Narayanan Balaji, Anh-Tuan Do |
ISCAS | 11 |
| 2023 | Stability Analysis of 6T SRAM at Deep Cryogenic Temperature for Quantum Computing ApplicationsabstractCMOS circuits operating at cryogenic temperature are gaining interest as one of the most promising approaches to efficiently scale up quantum processors in near- and medium future. However, there are major challenges such as (1) strict power dissipation limit at 4K plate due to the limited cooling power of the dilution fridge and (2) significant shifts in CMOS device behavior (i.e. variations, threshold voltage, charge carrier mobility and sub-threshold slope) which are not accurately captured in the standard BSIM models from the foundries. Although there have been extensive works experimentally characterizing and analyzing CMOS transistors at ∼4 K, there is a lack of digital and memory subsystem study. Since on-chip SRAM is one of the most power-consuming and the most vulnerable element in cryogenic SoC, this work analyzes the stability of low-voltage 6T-SRAM at deep cryogenic temperature (i.e 77K and 8K), in comparison with 300K operation. Our DC analysis showed that in general, Write static noise margins of the SRAM cell improves when temperature changes from 300K to 8K, even at low-voltage condition. Regarding the Read static noise margin, our simulation showed that the inverters exhibit pseudo-static hysteresis and interestingly this leads to an improvement of read static noise margin of the cell, similar to what observed in a Schmitt-Trigger SRAM. These results suggest that although CMOS transistors exhibit higher threshold voltage in cryogenic temperature, it is still possible to operate the SRAM at low-voltage for power saving in quantum computing applications. Seong-Beom Kim, Aarthy Mani, Leong Xu Heng Victor, Yuanjin Zheng, Anh-Tuan Do |
ISCAS | 5 |
| 2023 | Hardware-efficient Softmax Approximation for Self-Attention NetworksabstractSelf-attention networks such as Transformer have become state-of-the-art models for natural language processing (NLP) problems. Softmax function, which serves as a normalizer to produce attention scores, turns out to be a severe throughput and latency bottleneck of a Transformer network. Softmax datapath consists of data-dependent sequential nonlinear exponentiation and division operations, which are not amenable to pipelining and parallelism, nor can they be directly linearized for pretrained models without substantial accuracy drop. In this paper, we proposed a hardware efficient Softmax approximation which can be used as a direct plug-in substitution into pretrained transformer network to accelerate NLP tasks without compromising its accuracy. Experiment results on FPGA implementation show that our design outperforms vanilla Softmax designed using Xilinx IPs with 15x less LUTs, 55x less registers and 23x lower latency at similar clock frequency and less than 1% accuracy drop on main language benchmark tasks. We also propose a pruning method to reduce the input entropy of Softmax for NLP problems with high number of inputs. It was validated on CoLA task to achieve a further 25% reduction of latency. Nazim Altar Koca, Anh-Tuan Do, Chip-Hong Chang |
ISCAS | 2 |
| 2023 | 1V, 1.13μm pixel pitch Liquid Crystal Driver with Charge-Balancing Scheme for SLM ApplicationsabstractThis work proposes a compact 9T SRAM-based pixel design for low-voltage and high-speed modulator for spatially varying modulation of light (i.e. SLM). To reduce the supply voltage to the CMOS pixel backplane, the operating point is shifted towards the linear window by dynamically pulsing both the top & bottom electrodes in each pixel. A test chip in a standard 40nm CMOS technology was implemented, supporting upto 90 frames/second VGA display. Our testing results demonstrated that the proposed HCS switching scheme is efficient in achieving optical modulation up to 8-bit resolution, making it a viable candidate for state of the art SLMs applications with high frame rate and low power requirements. Aarthy Mani, Chong Yi Sheng, Rasna Maruthiyodan Veetil, Moitra Parikshit, Tobias Wilhelm W. Mass, Chong Ser Choong, Xuewu Xu, Ramon José Paniagua Domínguez, Arseniy I. Kuznetsov, P. Krishna, P. Keyi, Kevin Tshun Chuan Chai, Anh-Tuan Do |
ISCAS | 14 |
| 2023 | A 110nW Always-on Keyword Spotting Chip using Spiking CNN in 40nm CMOSabstractThis paper presents an ultra-low power keyword spotting (KWS) chip for Artificial Intelligence of Things (AIoT) device's always-on ambient sensing function. The core KWS engine is based on a spiking convolutional neural network (SCNN) model for its attractive features of sparse activation and addition-only operations inside the spiking neurons. The proposed SCNN model improves the existing framewise incremental computation flow by adding a spike processing unit (SPU) to reduce the computing cycles. The power and latency of the whole system are reduced by 16.5% and 43.2% respectively. Extensive network quantization reduces the weight bit-length to 4-bit and only 1-bit activation is required. The chip also supports power gating by an energy-based voice activity detection (VAD) module to further reduce power consumption in random and sparse event (RSE) scenarios. Full chip simulation results show that the chip consumes only 110nW with 2.15% False alarm rate and 3.00% False reject rate in a 10% voice event stream test. It achieves state-of-art recognition accuracy of 99% and 96% for one and two keyword detection tasks. Junran Pu, Yi Sheng Chong, Wang Ling Goh, Anh-Tuan Do, Yuan Gao 0011 |
ISCAS | 7 |
| 2023 | 282-to-607 TOPS/W, 7T-SRAM Based CiM with Reconfigurable Column SAR ADC for Neural Network ProcessingabstractCompute in memory ($C$iM) is a promising solution for solving the bottleneck of frequent data interface between memory and processor in Von-Neumann architecture. In this work, a hybrid current/charge domain 7T-SRAM based CiM architecture is proposed to mitigate the PVT-induced RBL variation during computation and thus offer a better linearity without significant impact on the operating frequency and area efficiency. Additionally, a column-referenced 1b to 5b reconfigurable SAR ADC is proposed to support multi-bit output. The proposed design is verified by the Monte-Carlo simulations using 40nm CMOS technology. The 5b mode ADC transferred MAC curve's DNL (LSB) ranges from −0.025 to 0.02 and INL (LSB) ranges from −0.13 to 0.25. The largest RBL variation$(\sigma)$from MAC value −64 to MAC value +64 is 2.08 mV, resulting in a MNIST classification accuracy of 97.5%, which is only 0.1% degradation and Google Speech Command classification accuracy of 80.5%, which is only 0.5% degradation compared to the software baseline, respectively. The whole architecture offers energy efficiency of 282-to-607 TOPS/W for 1-5b output in the MAC operation, which is competitive when compared to other state-of-art$C$iM architectures. Qibang Zang, Wang Ling Goh, Lu Lu 0013, Chengshuo Yu, Junjie Mu, Tony Tae-Hyoung Kim, Bongjin Kim, Dongrui Li, Anh-Tuan Do |
ISCAS | 9 |
| 2023 | LAXOR: A Bit-Accurate BNN Accelerator with Latch-XOR Logic for Local ComputingabstractBinary Neural Network (BNN) accelerators are attractive solutions for Artificial Internet-of-Things (AIoT) applications thanks to the compact models and low computational cost while maintaining satisfactory classification performance. Various analog/mix-signal compute-in-memory macros have been proposed to boost the energy efficiency of binary convolution tasks. However, this approach incurs inaccurate computation due to its sensitivity to temperature, noise, and process variations. In this work, we present a full-digital BNN architecture that leverages a novel Latch-XOR logic array for local bitwise multiplication, suppressing massive data movement and achieving 4.2× lower energy per operation compared to the decoupled standard cell approach. An optimized population count circuitry is also proposed for data accumulation, which obtains 1.37× Energy-Delay-Area saving compared to Binary-Adder-Tree-based implementation. To enable seamless hardware-software co-optimization, we have developed an in-house simulator for design space exploration as well as flexible mapping with various network topologies and kernel sizes. Our experiment shows the Latch-XOR-based architecture in 28nm CMOS technology achieves an enhanced energy efficiency of 2315 TOPS/W, 3.4× higher compared to the state-of-the-art synthesized digital architecture. This manifests that the proposed accelerator is highly suited for AIoT applications. Dongrui Li, Tomomasa Yamasaki, Aarthy Mani, Anh-Tuan Do, Niangjun Chen, Bo Wang 0020 |
ISLPED | 4 |
| 2022 | 0.08mm2 128nW MFCC Engine for Ultra-low Power, Always-on Smart Sensing ApplicationsabstractMel frequency cepstral coefficient (MFCC) features are widely used in applications such as keyword spotting, bearing fault detection and heart sound classification. This work proposes a low power MFCC engine that enables its use for battery-powered edge applications. Three hardware algorithm co-optimizations were adopted to achieve energy efficient MFCC hardware implementation. The approximated MFCC features due to the optimizations still allows good accuracy when deployed in several applications such as keyword spotting and bearing fault detection, reporting negligible accuracy drop of $\le 1.5$%. The proposed MFCC hardware consumes only 128nW at 0.3V supply and occupies only 0.08mm2in 40nm CMOS technology, which are $5 \times $ and $2.75 \times $ power and area reduction respectively when compared to the prior arts. Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 5 |
| 2022 | Recovering Accuracy of RRAM-based CIM for Binarized Neural Network via Chip-in-the-loop TrainingabstractResistive random access memory (RRAM) based computing-in-memory (CIM) is attractive for edge artificial intelligence (AI) applications, thanks to its excellent energy efficiency, compactness and high parallelism in matrix vector multiplication (MatVec) operations. However, existing RRAM-based CIM designs often require complex programming scheme to precisely control the RRAM cells to reach the desired resistance states so that the neural network classification accuracy is maintained. This leads to large area and energy overhead as well as low RRAM area utilization. Hence, compact RRAM-based CIM with simple pulse-based programming scheme is thus more desirable. To achieve this, we propose a chip-in-the-loop training approach to compensate for the network performance drop due to the stochastic behavior of the RRAM cells. Note that, although the target RRAM cell here is a two-state RRAM (i.e binary, having only high and low resistance states), their inherent analog resistance values are used in the CIM operation. Our experiment using a 4-layer fully-connected binary neural network (BNN) showed that after retraining, the RRAM-based network accuracy can be recovered, regardless of the RRAM resistance distribution and $\frac{\text{R}_{\text{HRS}}}{\text{R}_{\text{LRS}}}$ resistance ratio. Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 5 |
| 2022 | A 1800μm2, 953Gbps/W AES Accelerator for IoT Applications in 40nm CMOSabstractA compact and energy-efficient AES accelerator for area and power-constrained IoT applications was fabricated in a 40nm CMOS process. By eliminating the need of intermediate data registers for MixColumns and ShiftRows in our proposed AES accelerator, we were able to reduce the total flip-flops to only 269 bits. Further, by reusing functional blocks and swapping the D flip-flops in data storage with scan flip-flops, our chip occupies only a tiny area of $1800 \mu \text{m}^{2}$ with an extremely low number of 657 gates. In addition, clock gating method and near-threshold voltage were used in our design. Thus, our accelerator consumes only $3.2 \mu \text{W}$ with an operation efficiency of 953 Gbps/W using a 0.48 V supply voltage. Compared with prior arts, our design has savings of 53% on area and 55% on the number of gates. When operated with a supply voltage of 0.48 V at 25°C, we can also achieve lower energy efficiency. Jingjing Lan, Vishnu P. Nambiar, Ming Ming Wong, Fei Li 0015, Yuan Gao 0011, Kevin Tshun Chuan Chai, Anh-Tuan Do |
ISCAS | 7 |
| 2022 | Coreset: Hierarchical neuromorphic computing supporting large-scale neural networks with improved resource efficiency
Huaipeng Zhang, Tao Luo 0014, Chuping Qu, Myat Thu Linn Aung, Yingnan Cui, Jun Zhou 0014, Ming Ming Wong, Junran Pu, Anh-Tuan Do, Rick Siow Mong Goh, Weng-Fai Wong |
Neurocomputing | 10 |
| 2022 | Corrigendum to "Coreset: Hierarchical neuromorphic computing supporting large-scale neural networks with improved resource efficiency" [Neurocomputing (2022) 128-140]
Huaipeng Zhang, Tao Luo 0014, Chuping Qu, Myat Thu Linn Aung, Yingnan Cui, Jun Zhou 0014, Ming Ming Wong, Junran Pu, Anh-Tuan Do, Rick Siow Mong Goh, Weng-Fai Wong |
Neurocomputing | 10 |
| 2021 | An Energy-Efficient Convolution Unit for Depthwise Separable Convolutional Neural NetworksabstractHigh performance but computationally expensive Convolutional Neural Networks (CNNs) require both algorithmic and custom hardware improvement to reduce model size and to improve energy efficiency for edge computing applications. Recent CNN architectures employ depthwise separable convolution to reduce the total number of weights and MAC operations. However, depthwise separable convolution workload does not run efficiently in existing CNN accelerators. This paper proposes an energy-efficient CONV unit for pointwise and depthwise operation. The CONV unit utilizes weight stationary to enable high efficiency. The row partial sum reduction is engaged to increase parallelism in pointwise convolution thereby lightening the memory requirements on output partial sums. Our design achieves a maximum efficiency of 3.17 TOPS/W at 0.85V/40nm CMOS which is well-suited for energy constrained edge computing applications. Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 5 |
| 2021 | Efficient Implementation of Activation Functions for LSTM acceleratorsabstractActivation functions such as hyperbolic tangent (tanh) and logistic sigmoid (sigmoid) are critical computing elements in a long short term memory (LSTM) cell and network. These activation functions are non-linear, leading to challenges in their hardware implementations. Area-efficient and high performance hardware implementation of these activation functions thus becomes crucial to allow high throughput in a LSTM accelerator. In this work, we propose an approximation scheme which is suitable for both tanh and sigmoid functions. The proposed hardware for sigmoid function is 8.3 times smaller than the state-of-the-art, while for tanh function, it is the second smallest design. When applying the approximated tanh and sigmoid of 2% error in a LSTM cell computation, its final hidden state and cell state record errors of 3.1% and 5.8% respectively. When the same approximated functions are applied to a single layer LSTM network of 64 hidden nodes, the accuracy drops by 2.8% only. This proposed small yet accurate activation function hardware is promising to be used in Internet of Things (IoT) applications where accuracy can be traded off for ultra-low power consumption. Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do |
VLSI-SoC | 5 |
| 2021 | A 25 TOPS/W High Power Efficiency Deterministic and Split Stochastic MAC (SC-MAC) DesignabstractThis work presented a stochastic computing (SC) split multiply-and-accumulate (MAC) unit that is operated using deterministic sequence and is able to achieve time latency and power reductions without accuracy degrading. In this improved deterministic SC design, the conventional Stochastic Number Generator (SNG) with large overhead is replaced with a lightweight decoder that effectively generates uncorrelated and segmented stochastic number (SN) without the need for random sources (PRNG). The proposed deterministic and split SC-MAC is implemented in ASIC 40nm technology for detailed hardware evaluation and its functionality is also verified in convolutional neural network (CNN) using MNIST data sets. The new SCMAC is found to be higher in power efficiency (GMACS/mW) and lower in energy consumption (pJ/MAC) as compared to the conventional SC-MAC as well as the prior arts. Ming Ming Wong, Anh-Tuan Do |
VLSI-SoC | 3 |
| 2021 | A 5.28-mm² 4.5-pJ/SOP Energy-Efficient Spiking Neural Network Hardware With Reconfigurable High Processing Speed Neuron Core and Congestion-Aware RouterabstractIn recent years, fast computation, low power, and small footprint are the key motivations for building SNN hardware. The unique features of SNN hardware have not been fully exploited, where the computation speed and energy efficiency of the SNN hardware can be improved according to the sparse spiking and non-uniform traffic of SNN. In this paper, we propose a 5.28-mm$^{2}~4096$-neuron 1M-synapse energy-efficient digital SNN hardware that can achieve ultra-low energy per synaptic operation of 4.5 pJ. The proposed neuron computing unit is implemented in pipeline architecture to achieve high synaptic processing speed. The proposed spike processing unit can significantly increase the processing speed of the neuron core by$1.9\times $and$9.4\times $when the spike injection rate is 50% and 10%, respectively. Besides, the increase in the processing speed of the neuron core leads to a reduction in energy consumption of up to 81.5%. An event-driven clock gating circuit that can reduce the power consumption of the proposed neuron block by more than 70% is proposed in this paper. This paper proposes a supervised STDP+ algorithm for SNN training, and the classification accuracy of the MNIST digits is 89.6% with 73.6% weight sparsity of the output layer. Junran Pu, Wang Ling Goh, Vishnu P. Nambiar, Ming Ming Wong, Anh-Tuan Do |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | Aggressive Leakage Current Reduction for Embedded MRAM Using Block-Level Power GatingabstractThis paper exploits circuit techniques to realize a power-gated MRAM with the instant-on characteristic. Each block in a large-capacity MRAM is gated by a separate power transistor, allowing it to be fully turned on after only 1.3 ns. As a result, the whole MRAM is always in sleep mode, except the selected block. This eliminates the need for access pattern prediction as well as the requirement that the MRAM must be in idle for a significant of time before it is put in sleep or deep sleep mode. Our simulation shows that even in the worst-case-scenario, 86% of leakage current is saved. In typical cases, there are 96.5% leakage and 69% total power reduction. Our proposed scheme's implementation is straight forward and incurs less than 0.5% area overhead, including power transistors and control circuits. Anh-Tuan Do, Xuanyao Fong, Fei Li 0015 |
IECON | 1 |
| 2020 | Post-Silicon Validation Methodology for Resource-Constrained Neuromorphic HardwareabstractWith the semiconductor industry moving towards smaller and more advanced process nodes, the reliability of the fabricated system-on-chips (SoCs) has become a significant problem, although it can be mitigated by proper verification and validation methodologies. Typical neuromorphic SoC chips are resource constrained and have very low power envelopes, pushing designers to forego the integration of onchip debug circuits. Futhermore, these specialized neuromorphic SoCs may not even be capable of supporting memory debug circuits due to the inclusion of novel memory types. Hence, it is a huge challenge to execute a comprehensive system level post-silicon validation suite on such chips, while aiding debugging efforts reliably. This paper presents a post-silicon validation methodology and algorithm to test all the neurocores within our custom designed neuromorphic processing unit (NPU) chip, focused on pinpointing the faults on the synaptic weight storage elements. The described methodology does not require any specialized built-in debug circuits to infer and detect faults. Based on the test results, the custom NPU chip was able to respond accordingly to all performed tests under optimal voltage and frequency conditions. Yun Kwan Lee, Vishnu P. Nambiar, Kim Seng Goh, Anh-Tuan Do |
IECON | 4 |
| 2020 | Area and Energy Efficient 2D Max-Pooling For Convolutional Neural Network Hardware AcceleratorabstractIn the era of digital world, Convolutional Neural-network is widely adopted. To keep the network in a reasonable size, down-sampling is required. Max-pooling is a common solution for down-sampling. High throughput and area-efficient -hardware design of 2D max-pooling is essential for CNN accelerator. In this paper, we propose a highly efficient hardware design of max-pooling for 2D data input. The proposed design can be integrated into the neural-network accelerator as it can process the 2D input feature map data in parallel. The proposed solution, which is supporting up to 31 parallel data input, has been synthesized using at 40nm process node with area of 90×100 μm2, and the estimated power is 2.9 mW @200Mhz @1.1V power supply. Yi Sheng Chong, Anh-Tuan Do |
IECON | 3 |
| 2020 | Design and Characterization of Radiation-Hardened MCU for Space Application using Error Correction SRAM and Glitch Removal Clock Buffer CellabstractHigh-energy environmental radiation particles may affect the System on Chip (SoC) which causes errors in the combinational logic, sequential logic, and even in the clock network. In the latter, they appear as clock glitches that propagate, and eventually, incorrectly latch all the sequential circuits, such as flip-flops and latches attached to the clock node concerned. The impact on-chip functionality is usually fatal. In this work, we propose a low-cost adaptive clock glitch removal circuitry for radiation-resilient clock networks. The proposed technique can be adjusted based on the actual clock glitch profile, to ensure that the error in the clock network is removed, while the impact on the clock signal itself is minimized. It is fully synthesizable, and can thus be incorporated in a conventional digital flow. Using the proposed clock buffer cell, the implemented ARM Cortex M0 shows 85% error reduction when exposed to radiation with minimum area overhead. Anh-Tuan Do, Tony Tae-Hyoung Kim, Xin Liu 0015, Jun Zhou 0017 |
ISCAS | 1 |
| 2020 | Ultra-Low Leakage, High Fan-Out Neuro Connection Map with TCAM-Based LUT, Localized Priority Encoder and Decoder-Less SRAMabstractIn this paper we present an energy and area efficient Neuro Connection Map (NCM) which features high fan-out connection to improve the core utilization for large scale neuromorphic system. To meet the stringent power and area requirements, we propose to deploy multi-Vth memory cell design for performance-leakage optimization, localized priority encoder for area/power minimization and finally a decoder-less SRAM for area saving. Our measurements at 1 V supply show that the whole design consumes only 819nW of leakage power. Read and search power are at 5.9μW/MHz and 14.8μW/MHz, respectively. When used in a neuro core with average 100 spikes per neuron per second, the NCM consumes only 0.2pJ/Synaptic event at 1V, making it highly suitable for energy-efficient neuromorphic computing applications. Aarthy Mani, Fei Li 0015, Ming Ming Wong, Luo Tao, Vishnu Paramasivam, Anh-Tuan Do |
ISCAS | 7 |
| 2020 | Scalable Block-Based Spiking Neural Network Hardware with a Multiplierless Neuron ModelabstractThis paper proposes a scalable hardware architecture for block-based spiking neural networks utilizing a multiplierless spiking neuron model. These blocks were implemented as a neurocore mesh generated from an interconnect algorithm, allowing for seamless scalability of the network size while mitigating connectivity errors. The routing fabric asynchronous protocol allows for critical timing paths between blocks to be relaxed. The proposed neuron model consumed less logic compared to standard models with multipliers, reducing up to 16% of the neurocore logic cell area. The network was implemented alongside a computing subsystem as an FPGA-based system-on-chip, communicating via a fabric interconnect bridge. Experimental results validate the functionality of the proposed system, and achieved comparable classification accuracy to existing works. Vishnu P. Nambiar, Eng-Kiat Koh, Junran Pu, Aarthy Mani, Ming Ming Wong, Wang Ling Goh, Anh-Tuan Do |
ISCAS | 8 |
| 2019 | Block-Based Spiking Neural Network Hardware with Deme Genetic AlgorithmabstractHardware implementation of spiking neural networks (SNN) has been the focus of many previous works due to its higher execution speed. A block-based SNN architecture with a simple spiking neuron model is proposed in this paper. Compared to traditional spiking neuron models, the proposed model simplifies the equation of the membrane potential for ease of hardware implementation. The block-based SNN architecture also makes the hardware implementation more scalable and simplifies floorplanning. Deme genetic algorithm (GA) was applied for training the SNN model, and a population encoding scheme was used for spike time conversion. Two case studies were carried out to verify the functionality of the proposed model, namely number recognition and Fisher Iris classification. Experimental results showed that the proposed SNN model with deme GA was able to achieve comparable or higher classification accuracy than previous works. Junran Pu, Vishnu P. Nambiar, Anh-Tuan Do, Wang Ling Goh |
ISCAS | 3 |
| 2019 | An Area-Efficient 128-Channel Spike Sorting Processor for Real-Time Neural Recording With 0.175µW/Channel in 65-nm CMOSabstractThis paper presents a power- and area-efficient spike sorting processor (SSP) for real-time neural recordings. The proposed SSP includes novel detection, feature extraction, and improved K-means algorithms for better clustering accuracy, online clustering performance, and lower power and smaller area per channel. Time-multiplexed registers are utilized in the detector for dynamic power reduction. Finally, an ultra-low-voltage 8T static random access memory (SRAM) is developed to reduce area and leakage consumption when compared to D flip-flop-based memory. The proposed SSP, fabricated in 65-nm CMOS process technology, consumes only 0.175 μW/channel when processing 128 input channels at 3.2 MHz and 0.54 V, which is the lowest among the compared state-of-the-art SSPs. The proposed SSP also occupies 0.003 mm2/channel, which allows 333 channels/mm2. Anh-Tuan Do, Seyed Mohammad Ali Zeinolabedin, Dongsuk Jeon, Dennis Sylvester, Tony Tae-Hyoung Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | A 32kb 9T SRAM with PVT-tracking read margin enhancement for ultra-low voltage operationabstractDiminishing bitline sensing margin at low voltage condition is one of the most challenging design obstacles for reliable SRAM implementation in nano-scale CMOS technologies. This paper presents a self-biased design technique that improves the bitline sensing margin during the read operation by sourcing a current which is the same as the total leakage along each bitline. It is able to automatically track changes in supply voltage, operating temperature and die-to-die process variations. Furthermore, a 9T SRAM cell is utilized to ensure that bitline leakage is data-independent. Simulation and measurement results using 65 nm CMOS process show that the proposed technique enlarge the bitline swing over a wide range operating temperature and operates successfully down to the supply voltage of 0.18 V. Anh-Tuan Do, Kiat Seng Yeo, Tony Tae-Hyoung Kim |
ISCAS | 1 |
| 2015 | Design of a hybrid neural spike detection algorithm for implantable integrated brain circuitsabstractReal time spike detection is the first critical step to develop spike-sorting for integrated brain circuits interface applications. Nonlinear Energy Operator (NEO) and absolute thresholding have been widely used as the spike detection algorithms where NEO has a better performance measured by the probability of detection and false alarm. This paper proposes a hybrid spike detection algorithm incorporating both spike detection algorithms to reduce the power and to keep the detection rate the same as that of NEO. In the proposed algorithm, the absolute thresholding is performed first to detect a potential spike. Once a potential spike is detected, NEO is executed to check whether the detected spike by absolute thresholding is valid. Since NEO is conditionally conducted, this reduces the overall power consumption. The simulation shows that the proposed hybrid method improves the power consumption by 54.48% compared to NEO in 65 nm CMOS technology. Seyed Mohammad Ali Zeinolabedin, Anh-Tuan Do, Kiat Seng Yeo, Tony Tae-Hyoung Kim |
ISCAS | 2 |
| 2014 | A hybrid NEO-based spike detection algorithm for implantable brain-IC interface applicationsabstractReal time spike detection is the first critical step to develop spike-sorting for brain-IC interface applications. For implantable VLSI implementation, spike detection hardware must consume low power and at the same time ensure high true positive probability while having as few false alarms as possible. Currently, Nonlinear Energy Operator (NEO) and absolute thresholding are the two most widely used spike detection algorithms where NEO has a slightly better performance measured by areas under the receiver operating characteristic (ROC) curves. This paper revisits these two algorithms and quantitatively points out that NEO is in fact much better than absolute thresholding. We also propose a hybrid algorithm that offers similar accuracy as NEO but only requires 11% of power consumption. Anh-Tuan Do, Kiat Seng Yeo |
ISCAS | 1 |
| 2013 | An improved read/write scheme for anchorless NEMS-CMOS non-volatile memoryabstractWe proposed a NEMS-based anchorless memory structure with two stable mechanical states (Up and Down) for retaining data even at high operating temperate (>200°C) [5]. Compared to the conventional anchored devices, the anchorless structure offers better scalability and can operate at lower supply voltage, which is more desirable for integration with CMOS processes. This work addresses several issues in the previous NEM memory array implementation such as shuttle oscillation due to electrostatic pendulum and non-polarity write operation. We propose a 128 × 128 NEM memory array consisting of NEM memory cells and CMOS read/write control circuits. Our proposed read/write scheme with the two-level gate control eliminates the non-polarity write issue and achieves 32% power and 14% delay improvements when compared to the previous control scheme. Anh-Tuan Do, Karthik G. Jayaraman, Vincent Pott, Chua Geng Li, Pushpapraj Singh, Kiat Seng Yeo, Tony Tae-Hyoung Kim |
ISCAS | 1 |
| 2013 | A current-mode stimulator circuit with two-step charge balancing background calibrationabstractCurrent-mode CMOS stimulation systems have offered unprecedented opportunities for accurate and high through put in-vitro and in-vivo physiological studies. As these circuits are in long term contact with living organisms, they must be flexible, safe and power efficient. Any mismatch in biphasic current pulses will result in charge imbalance, leading to tissue/cell damage. Therefore, it is the most important to maintain the balance of the charge injected and retracted by the anode and the cathode, respectively. This work first adjusts the body biasing voltage of the anode to match with the cathode current. It is robust, process-variation-aware and can reduce the imbalanced current to less than 1%. Second, any residue charge at the stimulation site is retracted only when it reaches a critical value. This process is performed in the background and thus does not disturb the front-end operation. Overall, it can achieve less than 0.4 nA DC error current and thus is a suitable candidate for long term stimulation applications. Anh-Tuan Do, Yung Sern Tan, Gordon M. Xiong, Cleo Choong, Zhi-Hui Kong, Kiat Seng Yeo |
ISCAS | 1 |
| 2013 | A High Speed Low Power CAM With a Parity Bit and Power-Gated ML SensingabstractContent addressable memory (CAM) offers high-speed search function in a single clock cycle. Due to its parallel match-line (ML) comparison, CAM is power-hungry. Thus, robust, high-speed and low-powerMLsense amplifiers are highly sought-after in CAM designs. In this paper, we introduce a parity bit that leads to 39% sensing delay reduction at a cost of less than 1% area and power overhead. Furthermore, we propose an effective gated-power technique to reduce the peak and average power consumption and enhance the robustness of the design against process variations. A feedback loop is employed to auto-turn off the power supply to the comparison elements and hence reduce the average power consumption by 64%. The proposed design can work at a supply voltage down to 0.5 V. Anh-Tuan Do, Shoushun Chen, Zhi-Hui Kong, Kiat Seng Yeo |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | A comparative study of state-of-the-art low-power CAM match-line sense amplifier designsabstractRobust, high-performance and low-power match-line sense amplifier designs are urgently required to catch up with the new requirements of large-scale CAMs in nano-scale CMOS technologies. In this paper we evaluate the performance of four state-of-the-art match-line sense amplifier designs in terms of power, delay and robustness against temperature, supply voltage and process variations. Our results show that the pre-charge low match-line sensing schemes suffers from process variations. Despite featuring low power consumption, these designs can hardly be scaled down to operate in low-voltage sub-65 nm CMOS process. On the other hand, the conventional and the charge-injection designs are much more robust and hence more suitable for low-voltage sub-65 nm CMOS implementations. Anh-Tuan Do, Xiaoliang Tan, Shoushun Chen, Zhi-Hui Kong, Kiat Seng Yeo |
ACM Great Lakes Symposium on VLSI | 1 |
| 2011 | A low-power CAM with efficient power and delay trade-offabstractIn a Content Addressable Memory (CAM) architecture, both the match-line (ML) sensing circuit and the priority encoder (PE) contribute significantly large delays during a compare cycle. Meanwhile the priority encoder consumes significantly less energy when compared to the sensing circuits, i.e. ~1% of the overall energy consumption. Based on this observation, we propose the use of dual-supply voltages to trade-off the power and delay budget between the comparison and priority encoding circuits. In this work, the memory array and priority encoder is powered by a low and a high supply voltage, respectively. On top of this, a self-power-off ML sense amplifier is employed to reduce the voltage swing on the ML buses. Simulation results show a 76% dynamic power reduction as compared to the conventional design without sacrificing the overall speed. Anh-Tuan Do, Shoushun Chen, Zhi-Hui Kong, Kiat Seng Yeo |
ISCAS | 1 |
| 2011 | Adaptive priority toggle asynchronous tree arbiter for AER-based image sensorabstractIn this paper, we reported an adaptive priority toggle asynchronous tree arbiter for Address Event Representation (AER)-based image sensors. Simultaneous requests from event-triggered pixels, event latency, timing error and jitter are the inherent issues in AER-based read-out circuits. Fixed priority arbiter often results in unfair allocation of bus resource to only “privileged” pixels thus resulting in timing error. The proposed arbiter is able to reduce the timing error by toggling the requests priority during simultaneous requests. This also achieves the fair allocation of bus resource to all pixels. The featured eager propagation scheme allows the requests to propagate towards higher hierarchy in the tree during the arbitration process. As a result, latency and jitter problems can be reduced. Simulation result reveals that single event delay for 128-way tree arbiter is 4.2 ns and therefore, for 128 × 128 array, such arbiter can process up to 238.09M event/s which is more than 15 times faster than the speed of the reported AER sensor. The arbiter layout was realized with 2P4M 0.35µm CMOS process with a silicon area of 78 × 18µm2and had been implemented in 80 × 80 array AER temporal contrast sensor. Myat Thu Linn Aung, Anh-Tuan Do, Shoushun Chen, Kiat Seng Yeo |
VLSI-SoC | 2 |
| 2011 | Design and Sensitivity Analysis of a New Current-Mode Sense Amplifier for Low-Power SRAMabstractA new current-mode sense amplifier is presented. It extensively utilizes the cross-coupled inverters for both local and global sensing stages, hence achieving ultra low-power and ultra high-speed properties simultaneously. Its sensing delay and power consumption are almost independent of the bit- and data-line capacitances. Extensive post-layout simulations, based on an industry standard 1 V/65-nm CMOS technology, have verified that the new design outperforms other designs in comparison by at least 27% in terms of speed and 30% in terms of power consumption. Sensitivity analysis has proven that the new design offers the best reliability with the smallest standard deviation and bit-error-rate (BER). Four 32 × 32-bit SRAM macros have been used to validate the proposed design, in comparison with three other circuit topologies. The new design can operate at a maximum frequency of 1.25 GHz at 1 V supply voltage and a minimum supply voltage of 0.2 V. These attributes of the proposed circuit make it a wise choice for contemporary high-complexity systems where reliability and power consumption are of major concerns. Anh-Tuan Do, Zhi-Hui Kong, Kiat Seng Yeo, Jeremy Yung Shern Low |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | A 16Kb 10T-SRAM with 4x read-power reductionabstractThis work aims to reduce the read power consumption as well as to enhance the stability of the SRAM cell during the read operation. A new 10-transisor cell is proposed with a new read scheme to minimize the power consumption within the memory core. It has separate read and write ports, thus cell read stability is significantly improved. A 16Kb SRAM macro operating at 1V supply voltage is demonstrated in 65 nm CMOS process. Its read power consumption is reduced to 24% of the conventional design. The new cell also has lower leakage current due to its special bit-line pre-charge scheme. As a result, it is suitable for low-power mobile applications where power supply is restricted by the battery. Zhi-Hui Kong, Anh-Tuan Do |
ISCAS | 2 |