Mingoo Seok

dblp:64/3916 · DBLP profile ↗
← Back
64ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-9722-0979ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 57 · 9 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 On-the-Fly Memory Decompression in GPU Featuring In-Partition Compressed Block Prefetching and Gzip Decompression for Scientific Computing
abstract
In modern graphics processing units (GPUs), scientific computing workloads are memory-bound due to the everincreasing data size and the limited off-chip memory bandwidth. The problem still remains with the adoption of high-bandwidth memory (HBM). To further increase memory bandwidth, researchers have proposed on-the-fly memory decompression (OTFMD) techniques that aim to store compressed data off-chip while decompressing the data on-chip. However, for scientific computing applications, a critical challenge of prior approaches is their low compression ratios, as the existing works target integer applications and thus employ simple compression algorithms. Employing a compression algorithm with a higher compression ratio, such as Gzip, for floating-point applications could be promising; however, it is not straightforward to adopt such an approach in the OTFMD scheme. For example, Gzip’s compressed block size is much larger than the GPU’s cache line. Additionally, after global-to-local memory mapping, the data in a Gzip block can be physically distributed across multiple DRAM channels, which complicates the decompression process. In this work, to address those challenges, we propose a Gzip-based OTFMD technique for GPUs for scientific computing.We propose an in-partition compressed block prefetching scheme, where data is compressed within each memory partition separately. Then, upon a request, we prefetch the compressed data into the last-level cache (LLC) while fully preserving memory-level parallelism. We also propose a new Gzip decompression accelerator that offers 5.4× higher throughput per unit silicon area than the state-of-the-art design. Lastly, we propose modifications to the GPU’s microarchitecture to fully leverage the proposed OTFMD technique. Our experimental evaluations show average performance improvements of 38% (up to 63%) for CuBLAS and 15% (up to 22%) for CuSPARSE applications compared to the baseline GPU model.
Paul Xuanyuanliang Huang, Mingoo Seok
IEEE Trans. Computers2
2026 D6CIM: 60.4-TOPS/W All-Digital 6T-SRAM-Based Compute-in-Memory Macro Supporting 1-to-8 b Fixed-Point Arithmetic in a 28-nm CMOS
abstract
This article presents an all-digital 6T-SRAM-based compute-in-memory (D6CIM) macro that supports 1-to-8-bit fixed-point vector matrix multiplication. The D6CIM is fully digital, employing only digital standard cells and 6T SRAM bitcells. The D6CIM adopts a time-sharing architecture with an optimal degree of time-sharing to maximize the product of compute density and weight density. It also features several circuit techniques, namely non-precharge single-ended access, hybrid compressor adder-tree circuits, and bidirectional shift-accumulation. We prototype the D6CIM macro in a 28-nm CMOS technology. The maximum operating frequency is ~360 MHz in a 1.1-V supply. The energy efficiency is measured at up to 60.4 TOPS/W. It achieves a compute density of 1.46 TOPS/mm2and a weight density of 1005 kb/mm2. Compared to prior state-of-the-art CIM macros, the D6CIM achieves 1.3-4.8X improvement in the product of energy efficiency, compute density, and weight density.
Jonghyun Oh, Chuan-Tung Lin, Mingoo Seok
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 Theoretical Optimal Specifications of Memcapacitors for Charge-Based In-Memory Computing
abstract
This paper presents the optimal theoretical specifications of memcapacitors that allow them to surpass the existing state-of-the-art in-memory computing (IMC) prototypes in terms of energy efficiency and weight density. To do so, we build and simulate the SPICE model of an existing memcapacitor device in an IMC macro featuring charge-based computing. We develop the energy efficiency model and weight density model of the memcapacitor-based IMC macro, which are verified against SPICE simulation. Finally, we present the optimal theoretical specifications for memcapacitor devices to obtain a 10x improvement in energy efficiency and/or weight density over the existing IMC prototypes.
Zichen Qian, Rentao Wan, Chin-Hsiang Liao, Steven J. Koester, Mingoo Seok
ASP-DAC5
2025 A 4.2-to-0.5-V, 0.8-μA-0.8-mA, Power-Efficient Three-Level SIMO Buck Converter for a Quad-Voltage RISC-V Microprocessor
abstract
This article presents a Li-ion battery-compatible single-inductor-multiple-output (SIMO) buck converter that fulfills the power management need of an integrated sub-mW RISC-V microprocessor. The proposed converter can directly take a 4.2-V battery voltage and produce four power rails ranging from 1.8 V for I/O to 0.5 V for the processor core. The three-level input stage is chosen to reduce the inductor ripple size and switching loss, thus increasing power conversion efficiency (PCE). In addition, the fully digital implementation using novel domino flash analog-digital converters (ADCs) enables low static current. Also, pulse frequency modulation (PFM) results in a wide dynamic range. The proposed three-level SIMO converter has been prototyped in a 65-nm CMOS technology with the 32-bit RISC-V processor. Measurement results show that the converter achieves a$1000\times $load current range ($0.8~\mu $A–0.8 mA) to support the active or sleep modes of the processor. The converter marks the PCE of 56.2%–72.8%. Compared to the ideal buck-low-dropout voltage regulator (LDO) architecture (LDO-only), it improves the PCE by 23.8% (46.4%).
Dongkwun Kim, Zhaoqing Wang, Paul Xuanyuanliang Huang, Pavan Kumar Chundi, Suhwan Kim 0002, Andres A. Blanco, Ram Krishnamurthy 0001, Mingoo Seok
IEEE Trans. Very Large Scale Integr. Syst.8
2024 UL-VIO: Ultra-Lightweight Visual-Inertial Odometry with Noise Robust Test-Time Adaptation
Jinho Park 0004, Se Young Chun, Mingoo Seok
ECCV (63)3
2024 SPADES: A 0.54-GFLOPS/W Sparse Matrix Vector Multiplication Accelerator Featuring On-the-Fly GZIP Decompression for 3.36X Reduction in Off-Chip Data Movement
abstract
We propose SPADES, the first sparse matrix vector multiplication (SpMV) accelerator incorporating online GZIP decompression. The goal of the decompression is to reduce off-chip data movement, which has become a major source of energy consumption in SpMV accelerators. Our proposed GZIP-SpMV flow achieves an average compression ratio of 3.36 for the sparse matrix CSC data, reducing the off-chip data traffic by 70%. We fabricated SPADES in TSMC 28-nm with a die area of 2.47 mm2. Compared with the prior SpMV accelerator, SPADES achieves 2.32X better energy efficiency.
Paul Xuanyuanliang Huang, Yannis P. Tsividis, Mingoo Seok
ISLPED3
2024 Demo: Achieving Self-Interference Cancellation Across Different Environments
abstract
In order to enable the simultaneous transmission and reception of wireless signals on the same frequency, a full-duplex (FD) radio must be capable of suppressing the powerful self-interference (SI) signal emitted from the transmitter and picked up by the receiver. Critically, a major bottleneck in wideband FD deployments is the need for adaptive SI cancellation (SIC) that would allow the FD wireless system to achieve strong cancellation across different settings with distinct electromagnetic environments. In this work, we evaluate the performance of an adaptive wideband FD radio in three different locations and demonstrate that it achieves strong SIC in every location across different bandwidths.
Alon Simon Levin, Eliot Samuel Flores Portillo, Sasank Garikapati, Ahuva Bechhofer, Bo Zhang 0105, Manav Kohli, Igor Kadota, Harish Krishnaswamy, Mingoo Seok, Gil Zussman
MobiCom9
2024 Model-Based Study on the Limit of the Dynamic Load Regulation Performance of a Digital Low Dropout Regulator
abstract
A digital low dropout (DLDO) regulator is one of the most critical building blocks in on-chip power management for its technology portability, voltage scalability, and other benefits associated with digital-oriented design. A key metric of DLDOs is the dynamic load regulation performance, often measured as the maximum current that a DLDO can quickly supply upon a significant load step under a voltage droop constraint (usually 10% of the output voltage). Previous works focused on architecture and circuit techniques to improve this metric. However, limited research focuses on the model development for the dynamic load regulation performance. To fill this gap, in this article, we propose the analytical models of the maximum load current of the standard DLDOs employing feedback and feedforward control laws. The developed models shed light on the impact of various design parameters on the total load current of a DLDO, with which both circuit and system designers can navigate the design space quickly and effectively.
Yichen Xu 0004, Zhaoqing Wang, Jonghyun Oh, Mingoo Seok
IEEE Trans. Very Large Scale Integr. Syst.4
2023 Demo: Experimentation with Wideband Real-Time Adaptive Full-Duplex Radios
abstract
We present a set of experiments utilizing wideband real-time adaptive full-duplex (FD) radios, demonstrating simultaneous transmission and reception on the same frequency channel. Each FD radio consists of a circulator-based antenna interface, a switched-capacitor delay-line-based configurable Radio-Frequency Integrated Circuit (RFIC) that implements Self-Interference Cancellation (SIC), an FPGA that optimizes the RFIC configuration in under 1.1 sec and can adapt to environmental changes in under 0.3 sec, and a Software-Defined Radio (SDR) transmitting OFDM-like packets. We demonstrate a real-time adaptive FD radio that achieves the SIC necessary to reach the noise floor across a wide bandwidth of 50 MHz. Then, we use two FD radios to create a wireless link and showcase the superior FD throughput.
Alon Simon Levin, Igor Kadota, Sasank Garikapati, Bo Zhang 0105, Aditya Jolly, Manav Kohli, Mingoo Seok, Harish Krishnaswamy, Gil Zussman
SIGCOMM7
2023 CDAR-DRAM: Enabling Runtime DRAM Performance and Energy Optimization via In-Situ Charge Detection and Adaptive Data Restoration
abstract
With the increasing of dynamic random access memory’s (DRAM) capacity, the refresh operation rapidly becomes a major concern to the performance of the current computational system. Moreover, conservative timing parameters adopted for access operations make an increasing amount of negative impact on system performance and energy efficiency. In this article, we propose an in-situ charge detection and adaptive data restoration DRAM (CDAR-DRAM) architecture, which can dynamically adjust the refresh rate and relax the constraints on access timing by removing pessimistic timing margins for PVT variations. CDAR-DRAM employs a low-cost skewed-inverter-based detector to monitor the bitline voltage in runtime and estimate real-time timing parameters of cells. Based on the detector, an adaptive refresh and restore scheme (CDAR-ref) is presented, which progressively reduces the refresh rate and partially restores cells’ voltage just enough for cells with sufficient charge, thereby optimizing both refresh and restoration operations. Moreover, a supplementary adaptive access scheme (CDAR-acc) is presented, which detects the runtime charge level of recently accessed rows and reduces access latency aggressively, benefitting workloads in a single-core system and memory nonintensive workloads in a multicore system. CDAR’s flexibility allows the two schemes to be combined. The evaluation shows that in an eight-core system, the combined scheme improves performance and energy efficiency by 15.2% and 22.6%, respectively.
Yuxuan Qin, Chuxiong Lin, Weifeng He, Yanan Sun 0003, Zhigang Mao, Mingoo Seok
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2023 A DFT-Compatible In-Situ Timing Error Detection and Correction Structure Featuring Low Area and Test Overhead
abstract
In-situ timing error detection and correction (EDAC) structure is widely adopted in timing-error resilient circuits to reduce the conservative timing guardband induced by process, voltage, and temperature (PVT) variations. However, it introduces the latch-based datapath as well as extra detection and propagation logic, therefore challenges the design-for-testability (DFT) implementation. In this article, we propose a novel DFT-compatible EDAC structure with significant signal control simplification and test-pattern complexity reduction, featuring low area and test overhead. This structure leverages a new scannable EDAC cell (SEDC) which can be configured for timing EDAC in normal mode, or for shift operations as a flip-flop in scan mode. Specifically, the proposed detection logic can be controlled succinctly in scan shift operations and then observed via the global error propagation logic with simple control signal configurations during the test. Therefore, the sophisticated test pattern generation and critical path sensitization are removed. Based on the structure, a shift-based test method is presented to cover the EDAC structure with a low test pattern complexity and test time overheads. As compared with previous works, the proposed SEDC saves 30.5% area, 16.6% power, and 20.3% delay. Besides, our test method cooperating with the proposed EDAC structure reduces$149\times $and$23\times $static and at-speed test patterns, respectively, together with$232\times $static test cycles, and up to$25\times $at-speed test cycles on average, which proves the effectiveness for DFT.
Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Channel Estimation Using Deep Learning on an FPGA for 5G Millimeter-Wave Communication Systems
abstract
5G millimeter-wave (mmWave) communication systems enable exciting new applications by significantly reducing the latency and increasing the data rate. However, this comes at a large computational cost, which results in long latency and large energy consumption. In this work, we aim to address this challenge in the problem of channel estimation of such systems through a set of algorithm-hardware co-optimizations. First of all, we employed a model-based neural network to improve the rate of convergence. We also optimized the neural network and achieved improved loss while using approximately the same number of operations. Furthermore, we were able to reduce the computational complexity through the use of sparsity inherent in mmWave channels. The proposed neural network for the channel estimation scales the computational complexity by more than two orders. Based on these innovations, we implemented a channel estimation subsystem on Zynq 7020 FPGA. The subsystem obtains an improvement in latency of up to ~10X and an improvement in energy consumption of up to ~300X over CPU and GPU based systems.
Pavan Kumar Chundi, Xiaodong Wang 0001, Mingoo Seok
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 Leveraging Noise and Aggressive Quantization of In-Memory Computing for Robust DNN Hardware Against Adversarial Input and Weight Attacks
abstract
In-memory computing (IMC) substantially improves the energy efficiency of deep neural network (DNNs) hardware by activating many rows together and performing analog computing. The noisy analog IMC induces some amount of accuracy drop in hardware acceleration, which is generally considered as a negative effect. However, in this work, we discover that such hardware intrinsic noise can, on the contrary, play a positive role in enhancing adversarial robustness. To achieve that, we propose a new DNN training scheme that integrates measured IMC hardware noise and aggressive partial sum quantization at the IMC crossbar. We show that this effectively improves the robustness of IMC DNN hardware against both adversarial input and weight attacks. Against black-box adversarial input attacks and bit-flip weight attacks, DNN robustness has improved by up to 10.5% (CFAR-10 accuracy) and 33.6% (number of bit-flips), respectively, compared to conventional DNNs.
Sai Kiran Cherupally, Adnan Siraj Rakin, Shihui Yin, Mingoo Seok, Deliang Fan, Jae-sun Seo
DAC4
2021 CDAR-DRAM: An In-situ Charge Detection and Adaptive Data Restoration DRAM Architecture for Performance and Energy Efficiency Improvement
abstract
As the capacity of DRAM continues to grow, the refresh operation rapidly becomes the performance and power-efficiency bottleneck. Also, restore time, the time given for recharging cells post access, makes an increasingly large amount of negative impact on performance. To tackle these problems, in this paper, we propose an in-situ charge detection and adaptive data restoration DRAM (CDAR-DRAM) architecture, which can dynamically adjust the refresh rate and also relax the constraints on restore time. The proposed CDAR-DRAM employs a low-cost skewed-inverter-based detector, which can reduce the excessive timing margins that prior work added to guarantee the functionality of leaky DRAM cells under the worst-case temperature condition. Moreover, an adaptive DRAM refresh and restore scheme is proposed, which can switch automatically between two modes: (i) a refresh mode that supports adaptive refresh rate, and (ii) a restore mode that relaxes the constraints on restore time dynamically for cells having sufficient charge. With the transistor-and architecture-level simulations, we evaluate the CDAR-DRAM in an 8-core system across different workloads. Compared with the prior art, the proposed architecture achieves a 9.4% improvement in system performance and a 14.3% reduction in energy consumption, without requiring the time-consuming profiling process which many prior works employed.
Chuxiong Lin, Weifeng He, Yanan Sun 0003, Zhigang Mao, Mingoo Seok
DAC5
2021 Modeling and Optimization of SRAM-based In-Memory Computing Hardware Design
abstract
In-memory computing (IMC) has been demonstrated as a promising technique to significantly improve energy-efficiency for deep neural network (DNN) hardware accelerators. However, designing one involves setting many design variables such as the number of parallel rows to assert, analog-to-digital converter (ADC) at the periphery of memory sub-array, activation/weight precisions of DNNs, etc., which affect energy-efficiency, DNN accuracy, and area. While individual IMC designs have been presented in the literature, they have not investigated this multi-dimensional design optimization. In this paper, to fill this knowledge gap, we present a SRAM-based IMC hardware modeling and optimization framework. A unified systematic study closely models IMC hardware, and investigates how a number of design variables and nonidealities (e.g. device mismatch and ADC quantization) affect the DNN accuracy of IMC design. To maintain high DNN accuracy for the IMC SRAM hardware, it is shown that the number of activated rows, ADC resolution, ADC quantization range, and different sources of variability/noise need to be carefully selected and co-optimized with an underlying DNN algorithm to implement.
Jyotishman Saikia, Shihui Yin, Sai Kiran Cherupally, Bo Zhang 0105, Jian Meng, Mingoo Seok, Jae-sun Seo
DATE6
2021 An Area-Efficient Scannable In Situ Timing Error Detection Technique Featuring Low Test Overhead for Resilient Circuits
abstract
Timing error detection is a key technique for resilient circuits to explore the timing margins, yet it hinders the scan shift operations and increases the excessive test overhead. In this paper, we propose an area-efficient scannable in situ timing error detection technique consisting of a lightweight scannable error-detection cell and propagation logics, featuring low design-for-test effort and test overhead. The proposed error-detection cell fully reuses its main and shadow latches to construct the latch-based error-detection structure in normal mode, or the flip-flop-based datapath in scan mode. Therefore, it not only offers the time-borrowing ability to lower the correction overheads, but also supports the scan shift operations and detection logic tests. Besides, the dependency of error signal generation on the critical path sensitization is eliminated by configuring input and clock signals of error propagation logics, and thereby the detection and propagation logic can be tested easily. Benefiting from the technique, a set of test methods is presented with lower test pattern scales and test cycle overheads. As compared with previous works, the proposed cell saves at least 30.5% area overhead. Besides, experimental results across several benchmark circuits show that 116x of test patterns, 232x of static test cycles, and 26x of at-speed test cycles are saved on average, proving the effectiveness of the proposed technique for the design-for-test requirement.
Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok
ICCAD4
2021 Investigation of Dynamic Leakage-Suppression Logic Techniques Crossing Different Technology Nodes from 180 nm Bulk CMOS to 7 nm FinFET Plus Process
abstract
Leakage power reduction techniques are crucial for energy-efficient circuits. This paper investigates the leakage suppression capability, performance, and reliability of dynamic leakage suppression logic (DLSL) and feedforward leakage self-suppression logic (FLSL) techniques, crossing different technology nodes from TSMC 180 nm bulk CMOS to 7 nm FinFET Plus process. Compared with CMOS benchmarks, experimental results show that DLSL-based benchmarks demonstrate a leakage power reduction for four orders of magnitude in 180 nm and 130 nm technologies, while only two orders of magnitude in other technologies. Moreover, FLSL offers a 4-28× performance improvement over DLSL at a cost of 2× leakage power.
Zihan Lian, Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok
ISCAS6
2021 An Energy-Efficient Logic Cell Library Design Methodology with Fine Granularity of Driving Strength for Near- and Sub-Threshold Digital Circuits
abstract
Commercial multi-threshold standard logic cell libraries are designed for nominal super-threshold circuits. If blindly used at near- and sub-threshold voltages, such libraries exhibit excessively coarse granularity in driving strength, leading to sub-optimal logic synthesis and placement-and- routing results. To tackle this problem, a holistic methodology for designing a near- and sub-threshold standard cell library that has fine driving strength granularity is presented in this paper. Meanwhile, the proposed methodology leverages inverse narrow width effect, reverse short channel effect and forward body biasing to modulate the driving strength at low area overheads. Based on the proposed methodology, we develop a 65nm multi-threshold-voltage, multi-channel-length library and benchmark it against the commercial library across several common circuits. The results show a 26.6% reduction in power-delay product, a 28.1% reduction in energy-delay product, and a 27.0% reduction in leakage power at 5.8% area overhead on average, confirming the efficiency of the methodology in near- and sub-threshold digital circuits design.
Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok
ISCAS4
2021 An Ultra-Low Leakage Bitcell Structure with the Feedforward Self-Suppression Scheme for Near-Threshold SRAM
abstract
Leakage power consumption has become a critical issue for low power Static Random-Access Memory (SRAM) design in the near-threshold regime. In this paper, an ultra-low leakage fourteen-transistor SRAM bitcell structure with the feedforward self-suppression scheme is presented. To reduce the leakage power significantly as well as maintain the data stability in hold state, a cross-coupled dynamic leakage- suppression inverter-based structure is adopted. Furthermore, the bypass scheme is employed to enable the speed modulation for bitcell read and write operations. As compared with state- of-the-art designs, a 65 nm 8kb SRAM array with the proposed bitcell structure achieves 75× leakage power, 45% write power as well as 65% read power reduction at 0.4V. Further comparisons with different processes verify up to 38k times leakage power reduction in 130nm planar process and 139× in 7nm plus FinFET process.
Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok
ISCAS5
2020 MemNAS: Memory-Efficient Neural Architecture Search With Grow-Trim Learning
abstract
Recent studies on automatic neural architecture search techniques have demonstrated significant performance, competitive to or even better than hand-crafted neural architectures. However, most of the existing search approaches tend to use residual structures and a concatenation connection between shallow and deep features. A resulted neural network model, therefore, is non-trivial for resource-constraint devices to execute since such a model requires large memory to store network parameters and intermediate feature maps along with excessive computing complexity. To address this challenge, we propose MemNAS, a novel growing and trimming based neural architecture search framework that optimizes not only performance but also memory requirement of an inference network. Specifically, in the search process, we consider running memory use, including network parameters and the essential intermediate feature maps memory requirement, as an optimization objective along with performance. Besides, to improve the accuracy of the search, we extract the correlation information among multiple candidate architectures to rank them and then choose the candidates with desired performance and memory efficiency. On the ImageNet classification task, our MemNAS achieves 75.4% accuracy, 0.7% higher than MobileNetV2 with 42.1% less memory requirement. Additional experiments confirm that the proposed MemNAS can perform well across the different targets of the trade-off between accuracy and memory consumption.
Peiye Liu, Bo Wu 0018, Huadong Ma, Mingoo Seok
CVPR4
2020 KTAN: Knowledge Transfer Adversarial Network
abstract
Knowledge distillation was pioneered to transfer the generalization ability of a large teacher deep network to a light-weight student network. The student network can retain the high quality of the teacher network, yet exhibiting low computational complexity and storage requirement, which is attractive for deploying a deep convolution neural network on a resource-constrained mobile device. However, most of the existing methods focus on transferring the probability distribution of a softmax layer in a teacher network and neglect the intermediate representations. However, we find that the intermediate representation is critical for a student network to better understand the transferred generalization as compared to the probability distribution only. In this paper, therefore, we propose such a knowledge transfer adversarial network method which holistically considers both intermediate representations and probability distributions of a teacher network. To transfer the knowledge of intermediate representations, we set high-level teacher feature maps as a target, toward which the method trains student feature maps. Furthermore, to support various structures of a student network, we arrange a novel teacher-to-student layer. Finally, the proposed method employs an adversarial learning process. Specifically, it includes a discriminator network to fully exploit the spatial correlation of feature maps during the training process of a student network. The experimental results demonstrate that the proposed method can significantly improve the performance of a student network on two important vision tasks, image classification and object detection.
Peiye Liu, Wu Liu 0005, Huadong Ma, Zhewei Jiang, Mingoo Seok
IJCNN5
2020 Vesti: Energy-Efficient In-Memory Computing Accelerator for Deep Neural Networks
abstract
To enable essential deep learning computation on energy-constrained hardware platforms, including mobile, wearable, and Internet of Things (IoT) devices, a number of digital ASIC designs have presented customized dataflow and enhanced parallelism. However, in conventional digital designs, the biggest bottleneck for energy-efficient deep neural networks (DNNs) has reportedly been the data access and movement. To eliminate the storage access bottleneck, new SRAM macros that support in-memory computing have been recently demonstrated. Several in-SRAM computing works have used the mix of analog and digital circuits to perform XNOR-and-ACcumulate (XAC) operation without row-by-row memory access and can map a subset of DNNs with binary weights and binary activations. In the single array level, large improvement in energy efficiency (e.g., two orders of magnitude improvement) has been reported in computing XAC over digital-only hardware performing the same operation. In this article, by integrating many instances of such in-memory computing SRAM macros with an ensemble of peripheral digital circuits, we architect a new DNN accelerator, titled Vesti. This new accelerator is designed to support configurable multibit activations and large-scale DNNs seamlessly while substantially improving the chip-level energyefficiency with favorable accuracy tradeoff compared to conventional digital ASIC. Vesti also employs double-buffering with two groups of in-memory computing SRAMs, effectively hiding the row-by-row write latencies of in-memory computing SRAMs. The Vesti accelerator is fully designed and laid out in 65-nm CMOS, demonstrating ultralow energy consumption of <; 20 nJ for MNIST classification and <; 40 μJ for CIFAR-10 classification at 1.0-V supply.
Shihui Yin, Zhewei Jiang, Minkyu Kim 0001, Mingoo Seok, Jae-sun Seo
IEEE Trans. Very Large Scale Integr. Syst.5
2019 XNOR-SRAM: In-Bitcell Computing SRAM Macro based on Resistive Computing Mechanism
abstract
We present an in-memory computing SRAM macro for binary neural networks. The memory macro computes XNOR-and-accumulate for binary/ternary deep convolutional neural networks on the bitline without row-by-row data access. It achieves 33X better energy and 300X better energy-delay-product than digital ASIC and achieves high accuracy in machine learning tasks (98.3% for MNIST and 85.7% for CIFAR-10 datasets).
Zhewei Jiang, Shihui Yin, Jae-sun Seo, Mingoo Seok
ACM Great Lakes Symposium on VLSI4
2019 Master of none acceleration: a comparison of accelerator architectures for analytical query processing
abstract
Hardware accelerators are one promising solution to contend with the end of Dennard scaling and the slowdown of Moore's law. For mature workloads that are regular and have high compute per byte, hardening an application into one or more hardware modules is a standard approach. However, for some applications, we find that a programmable homogeneous architecture is preferable.
Andrea Lottarini, Joao Pedro Cerqueira, Thomas J. Repetti, Stephen A. Edwards, Kenneth A. Ross, Mingoo Seok, Martha A. Kim
ISCA6
2019 FPGA-based Acceleration of Binary Neural Network Training with Minimized Off-Chip Memory Access
abstract
In this paper, we examine the feasibility of FPGA as a platform for training a convolutional binary-weight neural network. Training a neural network requires more data movement compared to inference. Acceleration of training on an FPGA is, therefore, a challenge because the data movement increases off-chip memory accesses. We try to address this problem by storing most of the data in the on-chip memory and adopting batch renormalization. This allows for training a large network by reducing the required intermediate data and its movement. For the case where all data except the input images can be stored on an FPGA chip, we present an accelerator for training CNNs to classify the CIFAR-10 dataset. Further, we study the impact of network size on performance and energy of FPGA and GPU. Our accelerator mapped in the Arria 10 FPGA chip obtains up-to 9.33X higher energy efficiency compared to the Nvidia Geforce GTX 1080 Ti GPU at similar performance.
Pavan Kumar Chundi, Peiye Liu, Sangsu Park, Seho Lee, Mingoo Seok
ISLPED5
2019 K-Nearest Neighbor Hardware Accelerator Using In-Memory Computing SRAM
abstract
The k-nearest neighbor (kNN) is one of the most popular algorithms in machine learning owing to its simplicity, versatility, and implementation viability without any assumptions about the data. However, for large-scale data, it incurs a large amount of memory access and computational complexity, resulting in long latency and high power consumption. In this paper, we present a kNN hardware accelerator in 65nm CMOS. This accelerator combines in-memory computing SRAM that is recently developed for binarized deep neural networks and digital hardware that performs top-k sorting. We designed and simulated the kNN accelerator, which performs up to 17.9 million query vectors per second while consuming 11.8 mW, demonstrating >4.8X energy improvement over prior works.
Jyotishman Saikia, Shihui Yin, Zhewei Jiang, Mingoo Seok, Jae-sun Seo
ISLPED4
2019 Editorial TVLSI Positioning - Continuing and Accelerating an Upward Trajectory
abstract
I. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5].
Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.45
2019 Recursive Synaptic Bit Reuse: An Efficient Way to Increase Memory Capacity in Associative Memory
abstract
Neural associative memory (AM) is one of the critical building blocks for cognitive computing systems. It memorizes (learns) and retrieves input data by information content itself. One of the key challenges of designing AM for intelligent devices is to expand memory capacity while using a minimal amount of hardware and energy resources. However, prior arts show that memory capacity increases slowly, i.e., in square root with the total number of synaptic weights. To tackle this problem, we propose a synapse model called recursive synaptic bit reuse, which enables near-linear scaling of memory capacity with total synaptic bits. Our model can also handle input data that are correlated more robustly than the conventional model. We evaluated our model in the context of Hopfield neural networks (HNNs) that contain 5-327-KB data storage for synaptic weights. Our model can increase the memory capacity of HNNs as large as 30× over the conventional ones. The very large scale integration implementation of HNNs in 65 nm confirms that our proposed model can save up to 19× area and up to 232× energy dissipation as compared to the conventional model. These savings are expected to grow with the network size.
Tianchan Guan, Xiaoyang Zeng, Mingoo Seok
IEEE Trans. Very Large Scale Integr. Syst.3
2019 A Near-Threshold Spiking Neural Network Accelerator With a Body-Swapping-Based In Situ Error Detection and Correction Technique
abstract
Specialized architecture combined with near- and subthreshold voltage circuits emerges as a promising candidate to improve the energy efficiency in performing complex computing kernels in a resource-constrained device. One of the critical challenges in such design is the large delay variability across process, voltage, and temperature (PVT) variations. The in situ error detection and correction (EDAC) technique can potentially handle such variations; however, since existing techniques have targeted von-Neumann architecture and nominal voltage circuits, it becomes nontrivial to apply them on near and subthreshold voltage accelerators, many of which do not base on the von-Neumann architecture. In particular, those accelerators often have no instruction, making it difficult to use the popular instruction-replay-based error correction. To tackle this challenge, in this paper, we propose a novel in situ EDAC technique that utilizes dynamic, temporarily, and spatially fine-grained body swapping for error correction without instruction replay. Using the proposed technique, we prototyped a spiking neural network (SNN) sorter in non-von-Neumann architecture. The prototyped chip can successfully remove the worst-case margin and, thus, achieve 49.3% higher energy efficiency and 35.6% higher throughput compared to the baseline that operates with the worst-case margin. The proposed technique incurs only 4.1% silicon area overhead and requires no additional supply voltage.
Seongjong Kim, Joao Pedro Cerqueira, Mingoo Seok
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Better-Than-Worst-Case Design Methodology for a Compact Integrated Switched-Capacitor DC-DC Converter
abstract
We suggest a new methodology in co-designing an integrated switched-capacitor converter and a digital load. Conventionally, a load has been specified to the minimum supply voltage and the maximum power dissipation, each found at her own worst-case process, workload, and environment condition. Furthermore, in designing an SC DC-DC converter toward this worst-case load specification, designers often have been adding another separate pessimistic assumption on power-switch's resistance and flying-capacitor's density of an SC converter. Such worst-case design methodology can lead to a significantly over-sized flying capacitor and thereby limit on-chip integration of a converter. Our proposed methodology instead adopts the better than worst-case (BTWC) perspective to avoid over-design and thus optimizes the area of an SC converter. Specifically, we propose BTWC load modeling where we specify non-pessimistic sets of supply voltage requirement and load power dissipation across variations. In addition, by considering coupled variations between the SC converter and the load integrated in the same die, our methodology can further reduce the pessimism in power-switch's resistance and capacitor density. The proposed co-design methodology is verified with a 2:1 SC converter and a digital load in a 65 nm. The resulted converter achieves more than one order of magnitude reduction in the flying capacitor size as compared to the conventional worst-case design while maintaining the target conversion efficiency and target throughput. We also verified our methodology with a wide range of load characteristics in terms of their supply voltages and current draw and confirmed the similar benefits.
Dongkwun Kim, Mingoo Seok
ISLPED2
2018 Blacklist Core: Machine-Learning Based Dynamic Operating-Performance-Point Blacklisting for Mitigating Power-Management Security Attacks
abstract
Most modern computing devices make available fine-grained control of operating frequency and voltage for power management. These interfaces, as demonstrated by recent attacks, open up a new class of software fault injection attacks that compromise security on commodity devices. CLKSCREW, a recently-published attack that stretches the frequency of devices beyond their operational limits to induce faults, is one such attack. Statically and permanently limiting frequency and voltage modulation space, i.e., guard-banding, could mitigate such attacks but it incurs large performance degradation and long testing time. Instead, in this paper, we propose a run-time technique which dynamically blacklists unsafe operating performance points using a neural-net model. The model is first trained offline in the design time and then subsequently adjusted at run-time by inspecting a selected set of features such as power management control registers, timing-error signals, and core temperature. We designed the algorithm and hardware, titled a BlackList (BL) core, which is capable of detecting and mitigating such power management-based security attack at high accuracy. The BL core incurs a reasonably small amount of overhead in power, delay, and area.
Zhewei Jiang, Simha Sethumadhavan, Mingoo Seok
ISLPED5
2018 In~Situ and In-Field Technique for Monitoring and Decelerating NBTI in 6T-SRAM Register Files
Doyun Kim, Jiangyi Li, Peter R. Kinget, Mingoo Seok
IEEE Trans. Very Large Scale Integr. Syst.5
2017 Extending memory capacity of neural associative memory based on recursive synaptic bit reuse
abstract
Neural associative memory (AM) is one of the critical building blocks for cognitive workloads such as classification and recognition. It learns and retrieves memories as humans brain does, i.e., changing the strengths of plastic synapses (weights) based on inputs and retrieving information by information itself. One of the key challenges in designing AM is to extend memory capacity (i.e., memories that a neural AM can learn) while minimizing power and hardware overhead. However, prior arts show that memory capacity scales slowly, often logarithmically or in squire root with the total bits of synaptic weights. This makes it prohibitive in hardware and power to achieve large memory capacity for practical applications. In this paper, we propose a synaptic model called recursive synaptic bit reuse, which enables near-linear scaling of memory capacity with total synaptic bits. Also, our model can handle input data that are correlated, more robustly than the conventional model. We experiment our proposed model in Hopfield Neural Networks (HNN) which contains the total synaptic bits of 5kB to 327kB and find that our model can increase the memory capacity as large as 30X over conventional models. We also study hardware cost via VLSI implementation of HNNs in a 65nm CMOS, confirming that our proposed model can achieve up to 10X area savings at the same capacity over conventional synaptic model.
Tianchan Guan, Xiaoyang Zeng, Mingoo Seok
DATE3
2017 Microwatt end-to-End digital neural signal processing systems for motor intention decoding
abstract
This paper presents microwatt end-to-end digital signal processing (DSP) systems for deployment-stage real-time upper-limb movement intent decoding. This brain computer interface (BCI) DSP systems feature intercellular spike detection, sorting, and decoding operations for a 96-channel prosthetic implant. We design the algorithms for those operations to achieve minimal computation complexity while matching or advancing the accuracy of state-of-art BCI sorting and movement decoding. Based on those algorithms, we architect the DSP hardware with the focus on hardware reuse and event-driven operation. The VLSI implementation of the proposed systems in a 65-nm high-VTHshows that it can achieve 4.82 μW at the supply voltage of 300mV in the post-layout simulation. The area is 0.16 mm2.
Zhewei Jiang, Chisung Bae, Joonseong Kang, Sang Joon Kim, Mingoo Seok
DATE5
2017 A technique to transform 6T-SRAM arrays into robust analog PUF with minimal overhead
abstract
Physically unclonable function (PUF) is one of the critical security primitives for key generation and storage. The key challenge in building a PUF is to achieve high robustness against noise, temperature/supply voltage (VDD) variations and device aging with low cost. This paper presents a technique to transform a pre-existing SRAM array into an analog PUF by configuring the access transistors in a 6T-SRAM bitcell into a pair of threshold voltage (Vth) sensors and comparing their outputs. The proposed analog PUF circuit achieves significantly better robustness than conventional SRAM-based PUFs using reset states while maintaining the benefit of hardware reuse: low area overhead. Test chips are prototyped in a 65nm CMOS to verify the randomness, uniqueness, and robustness of the proposed design.
Jiangyi Li, Mingoo Seok
ISCAS3
2017 Hotspot monitoring and Temperature Estimation with miniature on-chip temperature sensors
abstract
This paper presents analysis and evaluation of the impact of size and voltage scalability of on-chip temperature sensor on the accuracy of hotspot monitoring and temperature estimation in dynamic thermal management of high performance microprocessors. The analysis is based on both the layout level and the system level across state-of-the-art sensors in terms of accuracy, voltage-scalability, and silicon footprint. Our analysis shows that a sensor having compact footprint and good voltage scalability can be placed on exact hotspot locations, typically among digital cells, significantly improving accuracy in tracking hotspots and estimating temperature of microarchitecture blocks, as compared to two other sensors that have higher sensor-circuit accuracy, large footprint and little voltage scalability limiting flexible placement.
Pavan Kumar Chundi, Yini Zhou, Martha A. Kim, Eren Kursun, Mingoo Seok
ISLPED5
2017 Comparative study and optimization of synchronous and asynchronous comparators at near-threshold voltages
abstract
We optimize and compare the performance of synchronous and asynchronous comparators across near-threshold and nominal supply voltage (0.5~1V). Comparators are the key components that determine the fundamental performance of analog-to-digital conversion in control and digital-signal processing (DSP) systems. While the asynchronous comparator has been considered inferior, operation of transistors in the near-threshold regime grants asynchronous comparators opportunities to improve power efficiency due to the more reduction in crowbar current than saturation drain current. We propose an enhanced asynchronous CSDA based comparator capable of achieving a superior latency vs. quiescent power dissipation trade-off to the synchronous clocked comparator in the near-threshold regime, a metric that is beneficial particularly to event-driven control systems. In-depth optimization and comparison results are presented.
Sung Justin Kim, Doyun Kim, Mingoo Seok
ISLPED3
2017 Hybrid analog-digital solution of nonlinear partial differential equations
abstract
We tackle the important problem class of solving nonlinear partial differential equations. While nonlinear PDEs are typically solved in high-performance supercomputers, they are increasingly used in graphics and embedded systems, where efficiency is important.
Yipeng Huang 0001, Mingoo Seok, Yannis P. Tsividis, Kyle T. Mandli, Simha Sethumadhavan
MICRO3
2017 Pipelining a triggered processing element
abstract
Programmable spatial architectures composed of ensembles of autonomous fixed-ISA processing elements offer a compelling design point between the flexibility of an FPGA and the compute density of a GPU or shared-memory many-core. The design regularity of spatial architectures demands examination of the processing element microarchitecture early in the design process to optimize overall efficiency.
Thomas J. Repetti, Joao Pedro Cerqueira, Martha A. Kim, Mingoo Seok
MICRO4
2017 Temporarily Fine-Grained Sleep Technique for Near- and Subthreshold Parallel Architectures
abstract
This paper presents a design approach for improving energy-efficiency and throughput of parallel architectures in near- and subthreshold voltage circuits. The focus is to suppress leakage energy dissipation of the idle portions of circuits during active modes, which can allow us to wholly transform the throughput improvement from parallel architectures into energy savings via deep voltage scaling. We begin by investigating the efficacy of parallel and pipeline architectures in the near- and subthreshold circuits. The investigation reveals that active energy dissipation largely undermines the ability of deep voltage scaling to transform excessive throughput into energy savings. Techniques, such as power-gating switches (PGSs), can mitigate active-leakage power dissipation; however, the overhead for entering and exiting sleep modes can offset the energy savings provided by sleep mode, particularly if sleep time is fine grained for suppressing active leakage. Therefore, in this paper, we propose a PGS design technique, inspired by the so-called zigzag supercutoff CMOS, in order to optimize the overheads of mode transitions of PGS in near- and subthreshold circuits. The proposed technique enables to have circuits in sleep mode for as short as a single clock cycle with a negligible amount of energy and delay overheads. We apply our proposed design to parallel multiplier-based test circuits operating at near- and subthreshold voltages. Simulations show a significant improvement in energy-efficiency over baselines at the same throughput.
Joao Pedro Cerqueira, Mingoo Seok
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Editorial
abstract
As I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design.
Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber
IEEE Trans. Very Large Scale Integr. Syst.36
2017 In Situ Error Detection Techniques in Ultralow Voltage Pipelines: Analysis and Optimizations
abstract
In order to achieve high tolerance against process, voltage, and temperature variations in the ultralow voltage (ULV) circuits, in situ error detection and correction (EDAC) techniques were presented. However, circuits adding the capability of error detection incur large hardware overhead, especially in ULV due to larger delay variability. In this paper, we analyze the hardware overhead of error detection techniques in pipelines based on three different sequential elements: flip-flops, two-phase latches, and pulsed latches. By exploiting the cycle-borrowing ability, we propose a technique called sparse insertion of error detecting registers on the two-phase latch-based and pulsed-latch-based pipelines to reduce the sequential logic area. Furthermore, we propose a delay-padding methodology using a multi-Vtcell library in ULV circuits to reduce EDAC hardware overhead. The proposed techniques are applied on a benchmark six-stage pipeline operating at 0.35 V in a 65-nm CMOS. The analysis results show that our proposed techniques can reduce the total area by 26%-33% and the error detecting register count by 2.9-4.3× compared with conventional EDAC techniques.
Wei Jin 0004, Seongjong Kim, Weifeng He, Zhigang Mao, Mingoo Seok
IEEE Trans. Very Large Scale Integr. Syst.5
2016 Evaluation of an Analog Accelerator for Linear Algebra
abstract
Due to the end of supply voltage scaling and the increasing percentage of dark silicon in modern integrated circuits, researchers are looking for new scalable ways to get useful computation from existing silicon technology. In this paper we present a reconfigurable analog accelerator for solving systems of linear equations. Commonly perceived downsides of analog computing, such as low precision and accuracy, limited problem sizes, and difficulty in programming are all compensated for using methods we discuss. Based on a prototyped analog accelerator chip we compare the performance and energy consumption of the analog solver against an efficient digital algorithm running on a CPU, and find that the analog accelerator approach may be an order of magnitude faster and provide one third energy savings, depending on the accelerator design. Due to the speed and efficiency of linear algebra algorithms running on digital computers, an analog accelerator that matches digital performance needs a large silicon footprint. Finally, we conclude that problem classes outside of systems of linear equations may hold more promise for analog acceleration.
Yipeng Huang 0001, Mingoo Seok, Yannis P. Tsividis, Simha Sethumadhavan
ISCA3
2016 Energy-Efficient Neuromorphic Classifiers
abstract
Neuromorphic engineering combines the architectural and computational principles of systems neuroscience with semiconductor electronics, with the aim of building efficient and compact devices that mimic the synaptic and neural machinery of the brain. The energy consumptions promised by neuromorphic engineering are extremely low, comparable to those of the nervous system. Until now, however, the neuromorphic approach has been restricted to relatively simple circuits and specialized functions, thereby obfuscating a direct comparison of their energy consumption to that used by conventional von Neumann digital machines solving real-world tasks. Here we show that a recent technology developed by IBM can be leveraged to realize neuromorphic circuits that operate as classifiers of complex real-world stimuli. Specifically, we provide a set of general prescriptions to enable the practical implementation of neural architectures that compete with state-of-the-art classifiers. We also show that the energy consumption of these architectures, realized on the IBM chip, is typically two or more orders of magnitude lower than that of conventional digital machines implementing classifiers with comparable performance. Moreover, the spike-based dynamics display a trade-off between integration time and accuracy, which naturally translates into algorithms that can be flexibly deployed for either fast and approximate classifications, or more accurate classifications at the mere expense of longer running times and higher energy costs. This work finally proves that the neuromorphic approach can be efficiently used in real-world applications and has significant advantages over conventional digital devices when energy consumption is considered.
Daniel Martí, Mattia Rigotti, Mingoo Seok, Stefano Fusi
Neural Comput.3
2015 A low power unsupervised spike sorting accelerator insensitive to clustering initialization in sub-optimal feature space
abstract
Online unsupervised spike sorting or clustering is an integral component of implantable closed-loop brain-computer-interface systems. Robust clustering performance against various non-idealities such as poor initialization and order-of-arrival of inputs are desirable while meeting the minimal area and power requirements for implants. We explore an online and unsupervised spike-sorting algorithm utilizing a low-overhead feature screening process that improves feature discriminability in the use of sub-optimal features for reducing hardware complexity. Based on the algorithm, an accelerator architecture that performs feature screening and clustering is devised and implemented in a 65-nm high-VTH CMOS, largely improving clustering accuracy even with poor clustering initialization. In the post-layout static timing and power simulation, the power consumption and the area of the accelerator are found to be 2.17 μW/ch and 0.052 μm2/ch, respectively, which are 53% and 25% smaller than the previous designs, while achieving the required throughput of 420 sorting/s at the supply voltage of 300mV.
Zhewei Jiang, Mingoo Seok
DAC3
2015 Energy-optimal voltage model supporting a wide range of nodal switching rates for early design-space exploration
abstract
This paper explores the models of the energy-optimal voltage (VOPT) of near/sub-threshold digital VLSI circuits with a focus on the support for a wide range of nodal switching rates. The previous models can estimate the VOPTof the circuits having relatively high nodal switching rates (VOPT, H), but can become inaccurate in finding the VOPT of the circuits having low nodal switching rate. In this work, therefore, we develop the models for finding (i) the VOPTof the circuits having low nodal switching rates (VOPT, L) and (ii) the critical nodal switching rate point (αcrit) below which the VOPT, Lshould be used. The models are verified with inverter chains and sub-threshold 10-transistor SRAM arrays in SPICE-level simulation. The model takes only process technology parameters to estimate VOPTs, and can be suitable for early-stage design-space exploration.
Doyun Kim, Jiangyi Li, Mingoo Seok
ICCD3
2015 A neuromorphic neural spike clustering processor for deep-brain sensing and stimulation systems
abstract
This paper presents algorithm and digital hardware design, inspired by biological spiking neural networks, to perform unsupervised, online spike-clustering with high accuracy and low-power consumption in the context of deep-brain sensing and stimulation systems. The proposed hardware contains 1220 digital neurons and 4.86k latch-based synapses, and achieves the average sorting accuracy of 91% whereas the conventional hardware based on the Osort algorithm achieves 69% for the same datasets. Implemented in a 65nm high-Vth, the processor exhibits a footprint of 0.25mm2/ch. and a power consumption of 9.3μW/ch. at VDDof 0.3V.
Beinuo Zhang, Zhewei Jiang, Jae-sun Seo, Mingoo Seok
ISLPED5
2015 Digital CMOS neuromorphic processor design featuring unsupervised online learning
abstract
The compute-intensive and power-efficient brain has been a source of inspiration for a broad range of neural networks to solve recognition and classification tasks. Compared to the supervised deep neural networks (DNNs) that have been very successful on well-defined labeled datasets, bio-plausible spiking neural networks (SNNs) with unsupervised learning rules could be well-suited for training and learning representations from the massive amount of unlabeled data. To design dense and low-power hardware for such unsupervised SNNs, we employ digital CMOS circuits for neuromorphic processors, which can exploit transistor scaling and dynamic voltage scaling to the utmost. As exemplary works, we present two neuromorphic processor designs. First, a 45nm neuromorphic chip is designed for a small-scale network of spiking neurons. Through tight integration of memory (64k SRAM synapses) and computation (256 digital neurons), the chip demonstrates on-chip learning on pattern recognition tasks down to 0.53V supply. Secondly, a 65nm neuromorphic processor that performs unsupervised on-line spike-clustering for brain sensing applications is implemented with 1.2k digital neurons and 4.7k latch-based synapses. The processor exhibits a power consumption of 9.3μW/ch at 0.3V supply. Synapse hardware precision, efficient synapse memory array access, overfitting, and voltage scaling will be discussed for dense and power-efficient on-chip learning for CMOS spiking neural networks.
Jae-sun Seo, Mingoo Seok
VLSI-SoC2
2014 Robust and In-Situ Self-Testing Technique for Monitoring Device Aging Effects in Pipeline Circuits
abstract
Runtime monitoring of aging effects in pipeline circuits is the key to dynamic reliability management techniques which can maximize the performance and energy-efficiency under a reliability envelope without imposing worst-case margin. The existing monitoring techniques are, however, severely limited: for sensor-based techniques, monitoring accuracy is significantly compromised due to the mismatches in aging conditions between sensors and the target circuits, as well as random variation of aging effects; for in-situ techniques, measurement results are sensitive to environmental variations during test phases, also severely reducing monitoring accuracy. We propose a new technique that enables accurate in-situ aging monitoring even under large environmental variations by (i) scaling the supply voltage for temperature-insensitive delay and (ii) reconfiguring target paths into ring oscillators, whose oscillation periods are measured and compared to pre-aging measurement to estimate aging-induced delay degradations. With additional accuracy-improving strategies, the technique achieves highly-accurate monitoring with an error of 15.5% across the temperature variations in self-test phases from 0°C to 80°C, exhibiting >30× improvement in accuracy as compared to the conventional technique operating at nominal supply voltage.
Jiangyi Li, Mingoo Seok
DAC2
2014 Reconfigurable regenerator-based interconnect design for ultra-dynamic-voltage-scaling systems
abstract
Ultra-dynamic-voltage-scaling (UDVS) is a compelling technique to use nominal supply voltage (VDD) for providing peak performance while achieving high energy efficiency by opportunistically using near/sub-threshold VDDs under average and low workload. One of the challenges in developing UDVS systems is that circuit fabrics optimized for a specific VDD can exhibit largely sub-optimal performance and energy efficiency at other VDDs. One critical example is the repeater-based interconnect design where the optimal interval of repeater insertion varies with VDD. In this paper, we propose a reconfigurable interconnect design based on an optimized regenerator to improve performance and energy efficiency across a wide range of VDDs.
Seongjong Kim, Mingoo Seok
ISLPED2
2014 Analysis and optimization of in-situ error detection techniques in ultra-low-voltage pipeline
abstract
In-situ error-detection and correction techniques have a strong potential to eliminate the worst-case margins in ultra-low-voltage (ULV) pipelines while achieving high variation tolerance. Adding the capability of error detection, however, can incur large hardware overhead, especially in ULV due to the larger variability. In this paper, we analyze the hardware overhead of error-detection techniques and propose a technique called sparse insertion of error-detecting registers. The proposed technique, applied on benchmark 3-stage pipeline operating at 0.35V, can reduce the error-detecting register count by 1.3-4.3×, total area by 15-40% and timing violation rate by 19-37×, compared to the conventional techniques.
Seongjong Kim, Mingoo Seok
ISLPED2
2013 Robust and energy-efficient asynchronous dynamic pipelines for ultra-low-voltage operation using adaptive keeper control
abstract
Asynchronous dynamic pipelines are increasingly being used, including in recent commercial design flows, since they simultaneously provide high-performance, clock-free operation and delay-insensitive communication. While they also show promise for energy-efficient ultra-low-voltage circuits, with always-on keepers, these circuits exhibit severe robustness issues. In this paper, an adaptive keeper solution is introduced, to eliminate write contention issues. Arbitrary unknown data rates and congestion must be safely handled, without a reference clock, hence conventional solutions for synchronous design cannot be applied. The proposed method, demonstrated in two widely-used pipelines (PS0, PCHB), directly addresses the asynchronous contention issue by dynamic monitoring of neighboring traffic at each pipeline stage. Simulations of a pipelined ripple-carry adder show correct operation at 0.3 V with energy improvements of up to 4.4× compared to a non-adaptive design. In addition, the approach also improves pipeline throughput by 24.4% and 17.4% at 0.6 V and nominal 1.0 V, respectively.
Mingoo Seok, Steven M. Nowick
ISLPED2
2012 Decoupling capacitor design strategy for minimizing supply noise of ultra low voltage circuits
abstract
Supply noise is a critical problem for the robust operation of integrated circuits at ultra low voltage regimes. Although decoupling capacitance is a traditional solution, the reduction of gate capacitance at subthreshold voltage can cause area overhead. In this paper, we propose a decoupling capacitor design strategy to reduce area overhead. The strategy consists of two parts: 1) enhancing gate capacitance through circuit optimizations and 2) using remote decoupling capacitors. Remote decoupling capacitors, which can be placed far from the block to compensate, can minimize the area overhead of the capacitance-enhancing optimizations. They also exploit less utilizable silicon area. The proposed strategy improves the capacitance density by 6.1 x without extra process steps, compared to the conventional approach. The gained robustness may be traded off for higher energy efficiency.
Mingoo Seok
DAC1
2012 A fine-grained many VT design methodology for ultra low voltage operations
abstract
In this paper, we propose a fine-grained many VT design methodology for ultra low voltage (ULV) operations of CMOS VLSI circuits. The fine-grained many VT transistors can be developed through only layout-level technique (e.g. inverse narrow width effects) in a multi-VT technology without any process modifications. Through SPICE simulations, we confirm that the proposed design methodology can improve performance, energy efficiency and variability of ULV circuits in three important domains, i.e. driving a fixed capacitive load, reducing active leakage energy consumption in non-critical paths, and lengthening short delay paths without energy overhead to aid error detection and correction techniques.
Mingoo Seok
ISLPED1
2012 Performance and energy-efficiency improvement through modified CPL in organic transistor integrated circuits
abstract
We study two logic families, complementary metal oxide semiconductor logic (CMOS) and complementary pass gate logic (CPL) in organic transistor technologies. In nanometer silicon technologies, CMOS logic family generally outperforms CPL in most of the simple gates. However, a significant strength imbalance between p- and n-channel transistors in currently-available organic technologies challenges the use of CMOS logic gates. Direct use of conventional CPL gates is also problematic due to the lack of good inverters for isolations. Therefore, in this paper, we propose a modified CPL logic family with optimized buffers, which can substantially improve both performance and energy efficiency for organic transistor integrated circuits. Benchmark simulations with 8b ripple carry adder also confirm that the modified CPL logic style improves performance and energy efficiency by ~3.2× and ~2.66×, compared to CMOS. We also investigate the impacts of prospective n-channel transistor improvements on the advantages of logic families.
Mingoo Seok
ISLPED1
2012 Sleep Mode Analysis and Optimization With Minimal-Sized Power Gating Switch for Ultra-Low ${V}_{\rm dd}$ Operation
abstract
This paper investigates the optimization of sleep mode energy consumption for ultra-low VddCMOS circuits, which is motivated by our findings that minimization of sleep mode energy holds great potential for reducing total energy consumption. We propose a unique approach of using a power gating switch (PGS) in ultra-low Vddregimes. Unlike the conventional manner of using PGSs, our optimization suggests using minimal-sized PGSs with a slightly higher Vddto compensate for voltage drop across the PGS. In SPICE simulations, this reduces total energy consumption by ~125× compared to conventional approaches. The effectiveness of the proposed optimization is also confirmed by measurements taken from an ultra-low power microprocessor. Additionally, the feasibility of using minimal PGSs in ultra-low Vddregimes is investigated using SPICE simulations and silicon measurements.
Mingoo Seok, Scott Hanson, David T. Blaauw, Dennis Sylvester
IEEE Trans. Very Large Scale Integr. Syst.1
2011 Pipeline strategy for improving optimal energy efficiency in ultra-low voltage design
abstract
This paper investigates pipelining methodologies for the ultra low voltage regime. Based on an analytical model and simulations, we propose a pipelining technique that provides higher energy efficiency and performance than conventional approaches to ultra low voltage design. Two-phase latch based design and sequential circuit optimizations are also proposed to further improve energy efficiency and performance. Silicon results demonstrate a 16b multiplier using the approaches in 65nm CMOS improve energy efficiency by 30% and performance by 60%.
Mingoo Seok, Dongsuk Jeon, Chaitali Chakrabarti, David T. Blaauw, Dennis Sylvester
DAC1
2011 Energy-optimized high performance FFT processor
abstract
This paper proposes an ultra low energy FFT processor suitable for sensor applications. The processor is based on R4MDC but achieves full utilization of computational elements. It has two parallel datapaths that increase throughput by a factor of 2 and also enable high memory utilization. The proposed design is implemented in 65nm CMOS technology and post-layout simulation including parasitic capacitances shows it achieves 9.25× higher energy efficiency than state-of-the-art FFT processors and high throughput relative to past subthreshold circuit implementations.
Dongsuk Jeon, Mingoo Seok, Chaitali Chakrabarti, David T. Blaauw, Dennis Sylvester
ICASSP2
2011 A 1.85fW/bit ultra low leakage 10T SRAM with speed compensation scheme
abstract
A low leakage memory is an indispensable part of any sensor application that spends significant time in standby (sleep) mode. Although using high Vth (HVT) devices is the most straightforward way to reduce leakage, it also limits operation speed during active mode. In this paper, a low leakage 10T SRAM cell, which compensates for operation speed using a readily available secondary supply, is proposed in a 0.18μm CMOS process. It achieves the lowest-to-date leakage power consumption and achieves robust operation at low voltage without sacrificing operation speed. The 10T SRAM has a bit cell area of 17.48μm2 and is measured to consume 1.85fW per bit at 0.35V.
Gregory K. Chen, Matthew Fojtik, Mingoo Seok, David T. Blaauw, Dennis Sylvester
ISCAS4
2010 Circuit design advances to enable ubiquitous sensing environments
abstract
This paper describes critical circuit building blocks for emerging sensing applications, particularly those where volume, and therefore power consumption constraints are orders of magnitude below current state of the art. Developments in ultra-low power microprocessors and memories are described, along with sub-nW timekeeping circuits, pW voltage references, and efficient DC-DC voltage conversion circuits at sub-μA current loads. Taken together, these circuit design advances point to a vision of true mm3low-cost sensing nodes with hybrid power sources, i.e., scavenged and micro-battery, providing long lifetime and reliable operation.
Mingoo Seok, Scott Hanson, Michael Wieckowski, Gregory K. Chen, Yu-Shiang Lin, David T. Blaauw, Dennis Sylvester
ISCAS1
2010 Clock network design for ultra-low power applications
abstract
Robust design is a critical concern in ultra-low voltage operation due to large sensitivity to process and environmental variations. In particular, clock networks need careful attention to ensure robust distribution of well-defined clock signals to avoid setup and hold time violations. In this paper, we investigate the design methodology of robust clock networks for ultra-low voltage applications. A case study shows that an optimally-chosen clock network improves skew variation by 36× and energy consumption by 49%, compared to a typical clock network. Additionally, the impact of supply voltage and technology scaling on the optimal clock network construction is investigated.
Mingoo Seok, David T. Blaauw, Dennis Sylvester
ISLPED1
2008 Optimal technology selection for minimizing energy and variability in low voltage applications
abstract
Ultra Low voltage operation has recently drawn significant attention due to its large potential energy savings. However, typical design practices used for super-threshold operation are not necessarily compatible with the low voltage regime. Here, radically different guidelines may be needed since existing process technologies have been optimized for super-threshold operation. We therefore study the selection of the optimal technology in ultra low voltage designs to achieve minimum energy and minimum variability which are among foremost concerns. We investigate five industrial technologies, from 250nm to 65nm. We demonstrate that mature technologies are often the best choice in very low voltage applications, saving as much as ~1800X in total energy consumption compared to a poorly selected technology. In parallel, the effect of technology choice on variability is investigated, when operating at the energy optimal design point. The results show up to a 4X improvement in delay variation due to global process shift and mismatch when using the most advanced technologies despite their large variability at nominal Vdd.
Mingoo Seok, Dennis Sylvester, David T. Blaauw
ISLPED1
2007 Nanometer Device Scaling in Subthreshold Circuits
abstract
Subthreshold circuit design is a strong candidate for use in future low power applications. It is not clear, however, that device scaling to 45nm and beyond will be beneficial in Subthreshold circuits. We investigate the implications of device scaling on subthreshold circuits and find that the slow scaling of gate oxide thickness leads to a 60% reduction in Ion/Ioff between the 90nm and 32nm device generations. We highlight the effects of this device degradation on noise margins, delay, and energy. We subsequently propose an alternative scaling strategy and demonstrate significant improvements in noise margins, delay, and energy in sub-Vth circuits.
Scott Hanson, Mingoo Seok, Dennis Sylvester, David T. Blaauw
DAC2
2007 Analysis and Optimization of Sleep Modes in Subthreshold Circuit Design
abstract
Subthreshold operation is a promising method for reducing power consumption in ultra-low power applications, such as active RFIDs and sensor networks. It was shown in previous works that operating at the Vmin supply voltage results in optimal energy operation, where Vmin typically falls below the threshold voltage. However, all previous subthreshold analyses ignore the leakage current in standby mode. Hence, for applications where operation at Vmin results in completion of the task well ahead of the required deadline, the energy consumption can be significantly under-estimated. In this paper, we investigate the effect of the non-zero standby energy on the optimal energy consumption in subthreshold operation. We first analyze energy consumption both with and without a cutoff technique in standby mode. Two parameters are proposed to capture the cutoff structure's effect on the energy consumption. Second, a methodology to minimize the total energy consumption is addressed. The selection of the cutoff structure is examined by comparing three different structures. Then, a co-optimization method to optimize the size of the cutoff structure concurrently with the supply voltage, is proposed. This approach reduces energy by 99.2% compared to standby-energy-unaware optimization.
Mingoo Seok, Scott Hanson, Dennis Sylvester, David T. Blaauw
DAC1