VLDB 2026 Research / reviewers in the wild / expert
Jongsun Park 0001
dblp:36/6402-1 · also Jong Sun Park 0001
· DBLP profile ↗
78ranked-venue papers
8as first author
34since 2021 · last 2026
0000-0003-3251-0024ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 66 · 5 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HINT: A Hybrid SRAM-MRAM Compute-In-Memory with INput-aware Skipping SAR-ADC for Energy Efficient Ternary LLMsabstractAlthough large language models (LLMs) show remarkable performance in natural language processing tasks, their deployment on resource-constrained devices remains challenging due to a substantial memory footprint and high-energy consumption. To address these challenges, low-bit and ternary quantization reduce the model size, while hardware approaches such as compute-in-memory (CIM) alleviate the overhead of external memory accesses. However, billions of parameters of LLMs still cause significant data movement, and existing ternary CIM suffers from a low-density bitcell as well as accuracy degradation due to cut-off analog-to-digital converters (ADCs). In this paper, we propose HINT, a CIM architecture incorporating two energy efficient techniques. First, hybrid ternary bitcell leverages the reliability of SRAM and the high-density of MRAM, reducing area and energy overhead. Second, input-aware skipping SAR-ADC exploits input sparsity to skip unnecessary conversion cycles without sacrificing accuracy. On BitNet b1.58 (700M), compared to SRAM-based and eDRAM-based CIM baselines, HINT improves bitcell density by 1.85× and achieves up to 2.67× higher energy efficiency, respectively. By skipping up to 21% of conversion cycles, the proposed ADC improves energy efficiency up to 1.27× while maintaining model accuracy. Jaebeom Park, Seungeon Hwang 0001, Jongsun Park 0001 |
DATE | 3 |
| 2026 | Speeding-Up Successive Read Operations of STT-MRAM via Read Path Alternation for Delay SymmetryabstractRecent studies on data-intensive computing systems have demonstrated that system throughput and latency are critically influenced by memory read bandwidth, underscoring the need for fast and energy-efficient memory read operations. Spin-transfer torque magnetic random-access memory (STT-MRAM) has emerged as a promising alternative to CMOS-based embedded memories; however, it still faces limitations in read speed and energy efficiency. This paper proposes a novel read scheme that improves both read speed and energy efficiency during successive read operations by alternating the read paths between data and reference cells. This path-alternating mechanism mitigates worst-case read scenarios by balancing the read voltage swing, enabling faster and more symmetric voltage development. HSPICE simulations using 28 nm LPP CMOS technology show a 31.5% improvement in read speed and a 48.8% reduction in energy consumption compared to conventional methods. Furthermore, the proposed scheme is validated under realistic manufacturing conditions through post-layout simulations using a 28nm standard process MRAM technology. System-level evaluations demonstrate that integrating the proposed scheme into STT-MRAM-based embedded memories in AI accelerators yields significant memory energy savings during various CNN inference tasks, outperforming conventional SRAM-based configurations. Seunghwa Hyun, Jangseok Yu, Jongsun Park 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | A Structured-Sparsity-Based Design Approach for Energy-Efficient Spiking Transformer ProcessingabstractSpiking transformers (STs) have emerged as promising architectures that achieve competitive accuracy with artificial neural network (ANN) on large-scale datasets. Despite this progress, the efforts to reduce the computational complexity of spiking self-attention (SSA) have seldom been explored. In this work, we present sparsity-based design approaches to reduce the computational complexity of SSA execution by systematically exploiting the structured sparsity in SSA. The proposed approaches are based on the observation that SSA naturally exhibits structured sparsity, which can be exploited to identify and skip redundant computations in SSA blocks. First, by employing a novel fully spiking SSA operator (FSSA) incorporating additional leaky integrate-and-fire (LIF) neurons, the average structured sparsity of SSA has been increased by$1.8\times $with less than 1% loss in accuracy. In addition, by adopting a filtering strategy that ignores neurons with low spike rate when detecting structured sparsity, additional$1.5\times $structured sparsity has been achieved compared to FSSA with negligible accuracy loss. Then, a sparsity-aware dataflow and hardware design convert the proposed structured sparsity patterns into runtime skipping of computations and weight transfers in SSA block execution. In postsynthesis simulations across the evaluated models, the proposed SSA operator with filtering techniques improve SSA block effective throughput and energy efficiency to 234.44 GOP/s and 319.12 GOP/J while keeping model accuracy loss below 1%, corresponding to 24.3% and 28.2% gains over the baseline spiking neural network (SNN) accelerator. Hyunseok Jung, Kyungchul Lee, Jongsun Park 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | GS-TG: 3D Gaussian Splatting Accelerator with Tile Grouping for Reducing Redundant Sorting while Preserving Rasterization Efficiencyabstract3D Gaussian Splatting (3D-GS) has emerged as a promising alternative to neural radiance fields (NeRF) as it offers high speed as well as high image quality in novel view synthesis. Despite these advancements, 3D-GS still struggles to meet the frames per second (FPS) demands of real-time applications. In this paper, we introduce GS-TG, a tile-grouping-based accelerator that enhances 3D-GS rendering speed by reducing redundant sorting operations and preserving rasterization efficiency. GS-TG addresses a critical trade-off issue in 3D-GS rendering: increasing the tile size effectively reduces redundant sorting operations, but it concurrently increases unnecessary rasterization computations. So, during sorting of the proposed approach, GS-TG groups small tiles (for making large tiles) to share sorting operations across tiles within each group, significantly reducing redundant computations. During rasterization, a bitmask assigned to each Gaussian identifies relevant small tiles, to enable efficient sharing of sorting results. Consequently, GS-TG enables sorting to be performed as if a large tile size is used by grouping tiles during the sorting stage, while allowing rasterization to proceed with the original small tiles by using bitmasks in the rasterization stage. GS-TG is a lossless method requiring no retraining or fine-tuning and it can be seamlessly integrated with previous 3D-GS optimization techniques. Experimental results show that GS-TG achieves an average speed-up of 1.54 times over state-of-the-art 3D-GS accelerators. Joongho Jo, Jongsun Park 0001 |
DAC | 2 |
| 2025 | GLOVA: Global and Local Variation-Aware Analog Circuit Design with Risk-Sensitive Reinforcement LearningabstractAnalog/mixed-signal circuit design encounters significant challenges due to performance degradation from process, voltage, and temperature (PVT) variations. To achieve commercial-grade reliability, iterative manual design revisions and extensive statistical simulations are required. While several studies have aimed to automate variation-aware analog design to reduce time-to-market, the substantial mismatches in real-world wafers have not been thoroughly addressed. In this paper, we present GLOVA, an analog circuit sizing framework that effectively manages the impact of diverse random mismatches to improve robustness against PVT variations. In the proposed approach, risk-sensitive reinforcement learning is leveraged to account for the reliability bound affected by PVT variations, and ensemble-based critic is introduced to achieve sample-efficient learning. For design verification, we also propose μ-σ evaluation and simulation reordering method to reduce simulation costs of identifying failed designs. GLOVA supports verification through industrial-level PVT variation evaluation methods, including corner simulation as well as global and local Monte Carlo (MC) simulations. Compared to previous state-of-the-art variation-aware analog sizing frameworks, GLOVA achieves up to $80.5 \times$ improvement in sample efficiency and $76.0 \times$ reduction in time. Junwoo Park, Chaehyeon Shin, Jaeheon Jung, Kyungho Shin, Seungheon Baek, Sanghyuk Heo, Woongrae Kim, In-Chul Jeong, Joohwan Cho, Jongsun Park 0001 |
DAC | 11 |
| 2025 | PS-GS: Group-Wise Parallel Rendering with Stage-Wise Complexity Reductions for Real-Time 3D Gaussian Splattingabstract3D Gaussian Splatting (3D-GS) is an emerging rendering technique that surpasses the neural radiance field (NeRF) in both rendering speed and image quality. Despite its advantages, running 3D-GS on mobile or edge devices in real-time remains challenging due to large computational complexity. In this paper, we introduce PS-GS, a specialized low-complexity hardware designed to enhance the pipeline parallelism of 3D-GS rendering pipeline process. In this work, we first observe that 3D-GS rendering can be parallelized when the approximate order of Gaussians, from those closest to the camera to those farthest, is known ahead. But, to enhance 3D-GS rendering speed via parallel processing, an efficient viewpoint-adaptive grouping method with low computational costs is essential. Two key computational bottlenecks of viewpoint-adaptive grouping are the grouping of invisible Gaussians and depth-based sorting. For efficient group-wise parallel rendering with low complexity viewpoint-adaptive grouping, we propose three key techniques—cluster-based preprocessing, sorting, and grouping—all seamlessly incorporated into the PS-GS architecture. Our experimental results demonstrate that PS-GS delivers an average speedup of 1.20x with negligible peak signal-to-noise ratio (PSNR) degradation. Joongho Jo, Jongsun Park 0001 |
DATE | 2 |
| 2025 | Speeding-Up Successive Read Operations of STT-MRAM via Read Path Alternation for Delay SymmetryabstractRecent research on data-intensive computing systems has demonstrated that system throughput and latency are critically dependent on memory read bandwidth, highlighting the need for fast memory read operations. Although spin-transfer torque magnetic random-access memory (STT-MRAM) has emerged as a promising alternative to CMOS-based embedded memories, STT -MRAM continues to face challenges related to read speed and energy efficiency. This paper introduces a novel read scheme that enhances read speed and energy in successive read operations by alternating read paths between data and reference cells. This approach effectively mitigates worst-case read scenarios by balancing the read voltage swings. HSPICE simulations using 28nm CMOS technology show a 31.5% improvement in read speed and 48.8% reduction in energy consumption compared to the previous approach. SCALE-Sim system simulations also demonstrate that applying the proposed read scheme to STT-MRAM embedded memories in AI accelerators shows a significant reduction in memory energy for CNN inference tasks compared to the SRAM embedded memory. Jongsun Park 0001 |
DATE | 2 |
| 2025 | SiGNoR: Similarity-based Graph Partitioning and Node Reuse for Memory Efficient GNN AccelerationabstractGraph neural networks (GNNs) are increasingly used in edge and mobile applications, but the large memory requirements of GNN hardware limit their deployment on resource-constrained devices. A common solution is graph partitioning, which splits large graphs into smaller partitions that fit into on-chip memory. However, this creates halo nodes—duplicated nodes across partitions—that lead to frequent and redundant accesses to external DRAM. This paper presents SiGNoR, an efficient graph partitioning approach that introduces three hardware-friendly techniques to reduce memory access with negligible accuracy loss. First, 1) similarity-based node reuse reduces redundant DRAM access of halo nodes by leveraging only local nodes within the same partition. To further improve the impact of the halo-to-local node replacement on GNN accuracy, 2) multiphase graph partitioning is proposed to refine partition boundaries by iteratively identifying and modeling its effects. Finally, 3) partition-wise block floating-point is presented for an energy efficient GNN accelerator hardware. The proposed partition-wise block floating-point shares a single exponent across all node features within each partition, reducing memory usage. Experimental results show that SiGNoR reduces average DRAM traffic of 44.5% with negligible accuracy degradation for GNNs across six real-world graph datasets. The GNN accelerator also improves energy efficiency by up to 2.6× compared to previous GNN hardware. Seungeon Hwang 0001, Hyeon Gwon Kim, Dongwoo Lew, Jongsun Park 0001 |
ICCAD | 4 |
| 2025 | A Pre-charge based Data Cell Dynamic Reference Sensing Scheme for Reliable Read Operations in STT-MRAMabstractAlthough spin transfer torque magnetic random access memory (STT-MRAM) has emerged as a strong candidate to replace conventional memories, relatively small read margin caused by low tunnel magnetoresistance (TMR) with process variations remains a significant challenge. In this paper, we present a data cell dynamic reference (DCDR) sensing scheme that significantly improves the sensing margin by comparing the parallel (P) and antiparallel (AP) states of the memory cells, rather than using an intermediate reference cell. This approach leverages a signal generator unit and a sample & hold circuit to efficiently manage the reference voltage, ensuring that the sensing operation fully reflects the difference between P and AP states. Simulations using 28nm CMOS technology with a 256Χ256 array demonstrate that the proposed DCDR approach achieves a sensing margin up to 257mV under a bit-error-rate (BER) of 7.99E-8 at 1.21ns, which represents a 2x larger margin and more than 100 times lower BER compared to the conventional scheme. By scaling down the precharge voltage, the proposed scheme results in over 30% read energy savings and 45% faster read speed under ISO 1E-5 target BER condition. Kyongsoo Kim, Jongsun Park 0001 |
ISCAS | 3 |
| 2025 | Big-Computing and Little-Storing STT-MRAM PIM Architecture With Charge Domain Based MAC OperationabstractSpin transfer torque magnetic random access memory (STT-MRAM) is a promising memory technology for processing in memory (PIM) thanks to its high endurance and relatively low device-to-device and cycle-to-cycle variations. However, the low OFF/ON ratio of STT device limits the number of active row-lines during multiply-accumulate (MAC) operations, degrading energy efficiency and computation speed. In this paper, we present an energy efficient and high speed Big-computing and Little-storing STT-MRAM PIM (BCLS-SP) architecture, which can increase the number of active row-lines with almost no area overhead. In the BCLS-SP architecture, a charge domain-based STT-MRAM PIM (CD-SP) structure is employed to concurrently activate many row-lines by improving MAC operation reliability. Filter-wise weight compression (FWC) and weight sharing (WS) are also devised to compress the weights stored in CD-SP, thus reducing area cost. In addition, the proposed architecture performs MAC operations with skipping zero-valued input (SZI) and zero-conversion scheme (ZCS) for better energy efficiency and performance. The simulations using 28nm CMOS process show that the BCLS-SP architecture shows energy reduction of 29% and performance improvement of 3.6 compared to the recent memristive device-based PIM using weight compression and input skipping. Yunho Jang, Yeseul Kim, Jongsun Park 0001 |
IEEE Trans. Computers | 4 |
| 2025 | DBB-ECC: Random Double Bit and Burst Error Correction Code for HBM3abstractAs dynamic random access memory (DRAM) technology continues to scale down, DRAM vendors have adopted on-die error correction codes (on-die ECC) to address reliability problems caused by cell failures. For burst error correction, a single symbol correction (SSC) Reed-Solomon (RS) code is utilized in high bandwidth memory (HBM) 3. However, randomly scattered errors frequently occur with aggressive technology scaling, which necessitates more robust error correction codes (ECC) scheme that addresses both burst errors and scattered errors. This brief presents double bit and burst ECC (DBB-ECC), an efficient scheme designed to correct both single symbol errors and random double bit errors with reduced implementation overhead. In the proposed decoding, syndromes based on SSC RS codes are used to address both error types without increasing parity bits. The decoder complexity has been also reduced by exploiting the syndrome patterns of double bit errors. The experimental results show that the proposed solution needs lower implementation overhead than conventional ones while maintaining same level of correction capability. Compared to the conventional SSC code, it also significantly enhances HBM3 reliability without increasing storage overhead. Chaehyeon Shin, Jongsun Park 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | A Design Framework of Heterogeneous Approximate DCIM-Based Accelerator for Energy-Efficient NN ProcessingabstractStatic random-access memory (SRAM) based digital compute-in-memory (DCIM) provides error-resilient computation at the expense of considerable power overhead of adder tree. In recent works, DCIM macro based on approximate computing mitigates the adder tree overheads, however, it faces a trade-off between power and neural network (NN) accuracy. The trade-off becomes more complicated in array-level CIM architecture since output channels of NN model have different sensitivities to approximation errors. In this paper, we propose a heterogeneous approximate DCIM-based accelerator design framework that achieves a good energy-accuracy trade-off for a specific NN model. The framework includes three key features: 1) Evolutionary algorithm-based search finds cost-efficient approximation points by pruning the design space. 2) Genetic algorithm-based channel-wise mapping creates heterogeneous approximation methods that effectively reduce DCIM energy consumption while maintaining high accuracy. 3) A hardware generation strategy decides the number of DCIM macros and their sizes, resulting in an energy-efficient DCIM-based accelerator tailored for the given NN model. Experimental results show that employing the proposed heterogeneous channel-wise mapping significantly enhances the energy efficiency compared to a homogeneous mapping. Moreover, the proposed framework can produce heterogeneous DCIM-based accelerators that consume less energy than state-of-the-art approximate DCIM approaches. Kyeongho Lee, Hyeyeong Lee, Jongsun Park 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Fault Bounding On-Die BCH Codes for Improving Reliability of System ECCabstractWhile continuous dynamic random access memory (DRAM) scaling may require an on-die error correction code (ECC) with enhanced correction capability, a double error correcting code with fault bounding scheme has not been explored. In this brief, we present the fault bounding on-die Bose–Chaudhuri–Hocquenghem (BCH) code that improves the compatibility with one-symbol error correcting system ECC used in dual data rate five (DDR5) dual in-line memory module (DIMM). By modifying theHmatrix of BCH code, the proposed decoding method determines the fault boundary within which burst errors occur, effectively preventing the spread of these errors across fault boundaries. A comparison of bounded rates with conventional codes illustrates the enhanced compatibility with system ECC. The encoder and decoder of the proposed code have been implemented using a 28-nm CMOS process to demonstrate the hardware cost. Seongyoon Kang, Chaehyeon Shin, Jongsun Park 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Bitline-Paired 2T SOT-MRAM Cell for Energy-Efficient Memory OperationabstractSpin-orbit torque magnetic random access memory (SOT-MRAM) has recently gained great attention due to its various benefits on memory implementation. However, SOT-MRAM suffers from high read and write energy consumption, making it difficult to replace conventional CMOS-based memories. This article presents a novel bitline (BL)-paired 2T SOT-MRAM cell structure, where the read and write path can be decoupled without causing bit-cell area overhead. Thanks to the separation of the read and write path, the BL parasitic capacitance can be reduced, which ultimately reduces read energy consumption. In addition, the read scheme tailored to the bit-cell structure is proposed to further enhance the read energy. The proposed cell also shows considerable write energy improvement when applying the write termination techniques that require simultaneous read and write operations. The HSPICE circuit simulations using the 28-nm CMOS technology show that the proposed BL-paired 2T SOT-MRAM cell with BL sharing achieves an average of 42.9% read energy savings, and an average of 56.8% write energy reduction compared to the conventional 2T cell in a$512\times 512$array. The system simulations using the gem5 simulator also show that, compared to the conventional 2T cell, an average of 47.1% of dynamic energy can be saved in various SPEC2006 benchmarks when the BL-paired 2T cell is employed in L2 cache. Jongsun Park 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | SpARC: Token Similarity-Aware Sparse Attention Transformer Accelerator via Row-wise ClusteringabstractSelf-attention mechanisms, the key enabler of transformers' remarkable performance, account for a significant portion of the overall transformer computation. Despite its effectiveness, self-attention inherently contains considerable redundancies, making sparse attention an attractive approach. In this paper, we propose SpARC, a sparse attention transformer accelerator that enhances throughput and energy efficiency by reducing the computational complexity of the self-attention mechanism. Our approach exploits inherent row-level redundancies in transformer attention maps to reduce the overall self-attention computation. By employing row-wise clustering, attention scores are calculated only once per cluster to achieve approximate attention without seriously compromising accuracy. To leverage the high parallelism of the proposed clustering approximate attention, we develop a fully pipelined accelerator with a dedicated memory hierarchy. Experimental results demonstrate that SpARC achieves attention map sparsity levels of 85-90% with negligible accuracy loss. SpARC achieves up to 4× core attention speedup and 6× energy efficiency improvement compared to prior sparse attention transformer accelerators. Han Cho, Seungeon Hwang 0001, Jongsun Park 0001 |
DAC | 4 |
| 2024 | TP-DCIM: Transposable Digital SRAM CIM Architecture for Energy-Efficient and High Throughput Transformer AccelerationabstractTo accelerate the execution of transformer models, compute-in-memory (CIM) has been widely adopted. However, the CIM architecture has the drawback of fixed one-way computing structure supporting only horizontal input sharing vertical accumulation (HIVA). So, it faces two major obstacles: 1) Matrix transposition processing needs large hardware overheads, 2) A fixed dataflow incurs CIM underutilization problem, degrading throughput. In this paper, we present a digital SRAM CIM (DCIM)-based transformer accelerator that supports two-way computing: HIVA and vertical input sharing horizontal accumulation (VIHA). The proposed two-way computing DCIM macro features a novel 9T SRAM bitcell and transposable adder tree, which efficiently reduces matrix transposition costs. We also present a novel dataflow to improve overall self-attention latency by resolving CIM underutilization and hiding dynamic input generation latency. In addition, CIM-friendly computation skipping scheme is exploited to enhance energy efficiency with negligible accuracy loss. The simulation results show that the proposed DCIM-based accelerator achieves up to 2.5x latency improvement and 24% energy savings compared to previous CIM-based accelerators. Junwoo Park, Kyeongho Lee, Jongsun Park 0001 |
ICCAD | 3 |
| 2024 | HeNCoG: A Heterogeneous Near-memory Computing Architecture for Energy Efficient GCN AccelerationabstractGraph convolutional network (GCN), which first applies convolutional operations to process graph data, has gained attention in various tasks involving relational data. Previous GCN accelerators have been designed with heterogeneous cores, considering two stages of inference (aggregation and combination), or with a unified core based on the inference of multi-layer as an iterative sparse-dense matrix multiplication. However, those prior works have suffered from an unnecessary large number of multiply-accumulate (MAC) operations and/or main memory accesses. In this paper, we propose HeNCoG, a GCN accelerator that utilizes a heterogeneous MAC array core for the combination stage and a near-memory computing core for the aggregation stage. In HeNCoG, considering that the number of MAC operations is significantly reduced when changing the stage execution order, the combination stage is executed first with a row-stationary dataflow. In the aggregation stage, magneto-resistive random-access memory (MRAM)-based near-memory computing is employed to reduce the number of main memory accesses needed to access the adjacency matrix in the graph dataset. Graph partitioning and double buffering techniques are also applied to further improve hardware efficiencies. Simulation results show that the HeNCoG architecture reduces execution cycles by 97% and memory accesses by 42% compared to previous works. Seungeon Hwang 0001, Duyeong Song, Jongsun Park 0001 |
ISCAS | 3 |
| 2024 | STT-MRAM-based Near-Memory Computing Architecture with Read Scheme and Dataflow Co-Design for High-Throughput and Energy-EfficiencyabstractSpin transfer torque magnetic random access memory (STT-MRAM)-based near-memory computing (NMC) architecture has been actively studied due to its potential for high-throughput and energy-efficient processing of AI algorithms. However, the low read bandwidth of STT-MRAM limits improvements in throughput of NMC architecture by failing to provide enough data to digital logic located near memory arrays. In this paper, we propose reference voltage recycling read scheme (RRS) and hybrid stationary dataflow scheme (HDS) to address the low read bandwidth issue in STT-MRAM. The proposed RRS allows for the recycling of previously developed reference bit-line voltage in a memory array during sequential read operations, enhancing read bandwidth as well as energy-efficiency of STT-MRAM. In addition, HDS balances the demand for input and weight data in digital logics located near memory arrays, facilitating full utilization of digital logics despite the limited read bandwidth of STT-MRAM. As a result, this approach can improve throughput of NMC architecture. The simulations using a 28nm CMOS process show that the proposed STT-MRAM-based NMC architecture can achieve TOPS/W of 8.45 and TOPS/mm2 of 1.37 with 8-bit input and weight, regardless of the data patterns. Yunho Jang, Yeseul Kim, Jongsun Park 0001 |
ISLPED | 3 |
| 2024 | iSPADE: End-to-end Sparse Architecture for Dense DNN Acceleration via Inverted-bit RepresentationabstractWhile recent cutting-edge deep neural network (DNN) models, such as large language models (LLMs), demonstrate remarkable capabilities, their inherent dense data characteristics limit the performance and energy gains achievable through sparse acceleration. In this paper, we introduce the iSPADE architecture, which sparsifies end-to-end execution of dense DNNs to directly adapt the advantages of sparse acceleration without applying accuracy-sensitive techniques such as pruning. First, we propose inverted-bit representation to eliminate repetitive sign bits in 2's complement representation. Leveraging the inverted-bit representation that generates a significant number of zero bits, we propose data packing and computation skipping techniques to reduce both redundant data movement and computation. Finally, we present an iSPADE bit-slice hardware architecture that efficiently supports and accelerates the proposed sparse dataflow. In the evaluation results, we assess performance across general DNN workloads using 8 popular DNNs. iSPADE achieves 4.1X and 4.5X improvements in energy efficiency and speedup, respectively, over the previous state-of-the-art bit-slice accelerators, and it realizes a 1.7X reduction in memory footprint. Han Cho, Jongsun Park 0001 |
ISLPED | 3 |
| 2023 | LoCoExNet: Low-Cost Early Exit Network for Energy Efficient CNN Accelerator DesignabstractEarly exit techniques, where the inference process of convolutional neural networks (CNNs) is early terminated by using auxiliary classifiers (called branches) to reduce data processing energy for easy inputs, have been actively researched. However, the conventional early exit works have suffered from the large size of branches, whose memory access energy is significantly large compared to that of the main network. In this article, we propose a low-cost early exit network (LoCoExNet), which significantly improves energy efficiencies by reducing the parameters used in inference with efficient branch structures. To further reduce the energy consumption of branches, based on the observation that the classification difficulties of the images can be predicted using the magnitudes of input activations, we also propose hardware-friendly dynamic branch pruning (DBP). Given the network and target dataset, we can construct an energy-efficient early exit network through the branch structure determination algorithm and DBP algorithm. Finally, we develop several architectural supporting modules to improve the energy efficiency of the proposed LoCoExNet with DBP. The experimental results show that the LoCoExNet uses 28.7% of the total network parameters on average for Tiny-ImageNet dataset using VGG-16 without accuracy loss. The CNN accelerator that implements LoCoExNet with DBP has been implemented using the 65nm CMOS process. The implementation results show that the LoCoExNet accelerator achieves up to 91% of energy savings and up to$4.46\mathbf {\times }$of speedups without accuracy loss for CIFAR-10 dataset using VGG-16 compared to the state-of-the-art CNN accelerators. Joongho Jo, Geonho Kim, Seungtae Kim, Jongsun Park 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | A 2941-TOPS/W Charge-Domain 10T SRAM Compute-in-Memory for Ternary Neural NetworkabstractIn this paper, we present a 10T SRAM compute-in memory (CiM) macro to process the multiplication-accumulation (MAC) operations between ternary-inputs and binary-weights. In the proposed 10T SRAM bitcell, the charge-domain analog computations are employed to improve the noise tolerance of bit-line (BL) signals where the MAC results are represented in CiM. Parallel processing of 3 different analog levels for ternary input activations is also performed in the proposed single 10T bitcell. To reduce the analog-to-digital converter (ADC) bit-resolutions without sacrificing deep neural network (DNN) accuracies, a confined-slope non-uniform integration (CS-NUI) ADC is proposed, which can provide layer-wise adaptive quantization for multiple different layers with different MAC distributions. In addition, by sharing the ADC reference voltage generator in every single column of SRAM array, the ADC area is effectively reduced with improved energy efficiencies of CiM. The$256\times 64.10\text{T}$SRAM CiM macro with the proposed charge-sharing scheme and CS-NUI ADCs has been implemented using 28nm CMOS process. The silicon measurement results show that the proposed CiM shows the accuracies of 98.66% and 88.48% with MNIST dataset on MLP, and CIFAR-10 dataset on VGGNet-7, respectively, with the energy efficiency of 2941-TOPS/W and the area efficiency of 59.584-TOPS/mm2. Sungsoo Cheon, Kyeongho Lee, Jongsun Park 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | Low Complexity Gradient Computation Techniques to Accelerate Deep Neural Network TrainingabstractDeep neural network (DNN) training is an iterative process of updating network weights, called gradient computation, where (mini-batch) stochastic gradient descent (SGD) algorithm is generally used. Since SGD inherently allows gradient computations with noise, the proper approximation of computing weight gradients within SGD noise can be a promising technique to save energy/time consumptions during DNN training. This article proposes two novel techniques to reduce the computational complexity of the gradient computations for the acceleration of SGD-based DNN training. First, considering that the output predictions of a network (confidence) change with training inputs, the relation between the confidence and the magnitude of the weight gradient can be exploited to skip the gradient computations without seriously sacrificing the accuracy, especially for high confidence inputs. Second, the angle diversity-based approximations of intermediate activations for weight gradient calculation are also presented. Based on the fact that the angle diversity of gradients is small (highly uncorrelated) in the early training epoch, the bit precision of activations can be reduced to 2-/4-/8-bit depending on the resulting angle error between the original gradient and quantized gradient. The simulations show that the proposed approach can skip up to 75.83% of gradient computations with negligible accuracy degradation for CIFAR-10 dataset using ResNet-20. Hardware implementation results using 65-nm CMOS technology also show that the proposed training accelerator achieves up to 1.69× energy efficiency compared with other training accelerators. Dongyeob Shin, Geonho Kim, Joongho Jo, Jongsun Park 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | A time-to-first-spike coding and conversion aware training for energy-efficient deep spiking neural network processor designabstractIn this paper, we present an energy-efficient SNN architecture, which can seamlessly run deep spiking neural networks (SNNs) with improved accuracy. First, we propose a conversion aware training (CAT) to reduce ANN-to-SNN conversion loss without hardware implementation overhead. In the proposed CAT, the activation function developed for simulating SNN during ANN training, is efficiently exploited to reduce the data representation error after conversion. Based on the CAT technique, we also present a time-to-first-spike coding that allows lightweight logarithmic computation by utilizing spike time information. The SNN processor design that supports the proposed techniques has been implemented using 28nm CMOS process. The processor achieves the top-1 accuracies of 91.7%, 67.9% and 57.4% with inference energy of 486.7uJ, 503.6uJ, and 1426uJ to process CIFAR-10, CIFAR-100, and Tiny-ImageNet, respectively, when running VGG-16 with 5bit logarithmic weights. Dongwoo Lew, Kyungchul Lee, Jongsun Park 0001 |
DAC | 3 |
| 2022 | Low-Cost 7T-SRAM Compute-in-Memory Design Based on Bit-Line Charge-Sharing Based Analog-to-Digital ConversionabstractAlthough compute-in-memory (CIM) is considered as one of the promising solutions to overcome memory wall problem, the variations in analog voltage computation and analog-to-digital-converter (ADC) cost still remain as design challenges. In this paper, we present a 7T SRAM CIM that seamlessly supports multiply-accumulation (MAC) operation between 4-bit inputs and 8-bit weights. In the proposed CIM, highly parallel and robust MAC operations are enabled by exploiting the bit-line charge-sharing scheme to simultaneously process multiple inputs. For the readout of analog MAC values, instead of adopting the conventional ADC structure, the bit-line charge-sharing is efficiently used to reduce the implementation cost of the reference voltage generations. Based on the in-SRAM reference voltage generation and the parallel analog readout in all columns, the proposed CIM efficiently reduces ADC power and area cost. In addition, the variation models from Monte-Carlo simulations are also used during training to reduce the accuracy drop due to process variations. The implementation of 256×64 7T SRAM CIM using 28nm CMOS process shows that it operates in the wide voltage range from 0.6V to 1.2V with energy efficiency of 45.8-TOPS/W at 0.6V. Kyeongho Lee, Joonhyung Kim, Jongsun Park 0001 |
ICCAD | 3 |
| 2022 | A Charge Domain P-8T SRAM Compute-In-Memory with Low-Cost DAC/ADC Operation for 4-bit Input ProcessingabstractThis paper presents a low cost PMOS-based 8T (P-8T) SRAM Compute-In-Memory (CIM) architecture that efficiently per-forms the multiply-accumulate (MAC) operations between 4-bit input activations and 8-bit weights. First, bit-line (BL) charge-sharing technique is employed to design the low-cost and reliable digital-to-analog conversion of 4-bit input activations in the pro-posed SRAM CIM, where the charge domain analog computing provides variation tolerant and linear MAC outputs. The 16 local arrays are also effectively exploited to implement the analog mul-tiplication unit (AMU) that simultaneously produces 16 multipli-cation results between 4-bit input activations and 1-bit weights. For the hardware cost reduction of analog-to-digital converter (ADC) without sacrificing DNN accuracy, hardware aware system simulations are performed to decide the ADC bit-resolutions and the number of activated rows in the proposed CIM macro. In addition, for the ADC operation, the AMU-based reference col-umns are utilized for generating ADC reference voltages, with which low-cost 4-bit coarse-fine flash ADC has been designed. The 256×80 P-8T SRAM CIM macro implementation using 28nm CMOS process shows that the proposed CIM shows the accuracies of 91.46% and 66.67% with CIFAR-10 and CIFAR-100 dataset, respectively, with the energy efficiency of 50.07-TOPS/W. Joonhyung Kim, Kyeongho Lee, Jongsun Park 0001 |
ISLPED | 3 |
| 2022 | Stochastic SOT Device Based SNN Architecture for On-Chip Unsupervised STDP LearningabstractEmerging device based spiking neural network (SNN) hardware design has been actively studied. Especially, energy and area efficient synapse crossbar has been of particular interest, but processing units for weight summations in synapse crossbar are still a main bottleneck for energy and area efficient hardware design. In this paper, we propose an efficient SNN architecture with stochastic spin-orbit torque (SOT) device based multi-bit synapses. First, we present SOT device based synapse array using modified gray code. The modified gray code based synapse needs only N devices to represent 2^N levels of synapse weights. Accumulative spike technique is also adopted in the proposed synapse array, to improve ADC utilization and reduce the number of neuron updates. In addition, we propose hardware friendly algorithmic techniques to improve classification accuracies as well as energy efficiencies. Non-spike depression based stochastic spike-timing-dependent plasticity is used to reduce the overlapping input representation and classification error. Early read termination is also employed to reduce energy consumption by turning off less associated neurons. The proposed SNN processor has been implemented using 65nm CMOS process, and it shows 90% classification accuracy in MNIST dataset consuming 0.78J/image (training) and 0.23J/image (inference) of energy with an area of 1.12mm2. Yunho Jang, Gyuseong Kang, Yeongkyo Seo, Kyung-Jin Lee, Byong-Guk Park, Jongsun Park 0001 |
IEEE Trans. Computers | 7 |
| 2022 | SOT-MRAM Digital PIM Architecture With Extended Parallelism in Matrix MultiplicationabstractEmerging device-based digital processing-in-memory (PIM) architectures have been actively studied due to their energy and area efficiency derived from analog to digital converter (ADC)-less PIM hardware. However, digital PIM architectures generally need large extra memories to copy parameters, and they also suffer from low computation per memory-cycle efficiencies. In this paper, we present a novel spin-orbit torque magnetic random access memory (SOT-MRAM) based digital PIM architecture to alleviate the extra memory size burden and computation cycle issues. First, we propose the spintronics-assisted logic-in-memory (SLIM) cells to support efficient digital logic operations inside memories, where the voltage-controlled magnetic anisotropy (VCMA) is exploited to enhance the computation per memory-cycle efficiencies. In addition, crossed input source PIM (CRISP) architecture is proposed to extend the merits of SLIM cells by eliminating the extra memories for parameter copying while significantly improving the degree of parallel processing. An intra-memory pipelining scheme is also considered to further increase the throughput of CRISP. The proposed CRISP architecture has been implemented using 28 nm CMOS process, and it presents 1.10 TOPS/W and 0.95 TOPS/mm2, showing considerable improvements of energy efficiency and throughput per area, compared to the state-of-the-art digital PIM architecture. Finally, to evaluate the impact of computation errors induced from the SOT devices and circuits in CRISP architecture, classification accuracy simulations have been performed while applying computation errors. Yunho Jang, Min-Gu Kang, Byong-Guk Park, Kyung-Jin Lee, Jongsun Park 0001 |
IEEE Trans. Computers | 6 |
| 2022 | A Dual-Domain Dynamic Reference Sensing for Reliable Read Operation in SOT-MRAMabstractAlthough spin orbit torque magnetic random access memory (SOT-MRAM) is one of the strong candidates for next-generation embedded memories, the degradation of read margin due to low tunnel magnetoresistance ratio (TMR) with process variations has been a large concern. In this paper, we present the dual-domain dynamic reference (DDDR) sensing scheme, where the reference voltage can be dynamically changed based on the combined voltage and time domain sensing to increase the sensing margin. The Half Schmitt trigger and sample & hold circuits are efficiently employed to generate data-dependent reference voltages and to store the sampled voltage levels at different times, respectively. According to the simulations using 28nm CMOS technology with 128 by 128 SOT-MRAM array, the proposed DDDR approach achieves a 243mV of sensing margin under 6.08E-8 bit-error-rate (BER) at 1.76ns, which is 2X larger margin with more than 100 times lower BER compared to the conventional read scheme. When scaling down the pre-charge voltage, the proposed scheme achieves more than 50% of the read energy savings under 1E-5 target BER condition. Jooyoon Kim, Yunho Jang, Jongsun Park 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | A Charge-Sharing based 8T SRAM In-Memory Computing for Edge DNN AccelerationabstractThis paper presents a charge-sharing based customized 8T SRAM in-memory computing (IMC) architecture. In the proposed IMC approach, the multiply-accumulate (MAC) operation of multi-bit activations and weights is supported using the charge sharing between bit-line (BL) parasitic capacitances. The area-efficient customized 8T SRAM macro can achieve robust and voltage-scalable MAC operations due to the charge-domain computation. We also propose a split capacitor structure-based 5/6-bit reconfigurable successive approximation register analog-to-digital converter (SAR-ADC) to reduce the hardware cost of an analog readout circuit while supporting higher precision MAC operations. The proposed reconfigurable SAR-ADC has been exploited to implement layer-by-layer mixed bit-precisions in convolution layer for increasing energy efficiency with negligible accuracy loss. The 256×64 8T SRAM IMC macro has been implemented using 28nm CMOS process technology. The proposed SRAM macro achieves 11. 20-TOPS/W with a maximum clock frequency of 125MHz at 1. 0V. It also supports supply voltage scaling from 0.5V to 1.1V with the energy efficiency ranging from 8.3-TOPS/W to 35.4-TOPS/W within 1 % accuracy loss. Kyeongho Lee, Sungsoo Cheon, Joongho Jo, Woong Choi, Jongsun Park 0001 |
DAC | 5 |
| 2021 | An Energy-Efficient SNN Processor Design based on Sparse Direct Feedback and Spike PredictionabstractIn this paper, we present a novel spike prediction technique based spiking neural network (SNN) architecture, which can provide low cost on-chip learning based on the sparse direct feedback alignment (DFA). First, in order to reduce the repetitive synaptic operations in feedforward operations, a spike prediction technique is proposed, where the output spikes of active and inactive neurons are predicted by tracing the membrane potential changes. The proposed spike prediction achieves 63.84% reduction of the synaptic operations, and it can be efficiently exploited in the training as well as inference process. In addition, the number of weight updates in backward operations has been reduced by applying sparse DFA. In the sparse DFA, the synaptic weight updates are computed using the sparse feedback connections and the output error that is sparse as well. As a result, the number of weight updates in the training process has been reduced to 65.17%. The SNN processor with the proposed spike prediction technique and the sparse DFA has been implemented using 65nm CMOS process. The implementation results show that the SNN processor achieves the training energy savings of 52.16% with 0.3% accuracy degradation in MNIST dataset. It also consumes 1.18 uJ/image and 1.34 uJ/image for inference and training, respectively, with 97.46% accuracy on MNIST dataset. Dongwoo Lew, Sunghyun Choi 0004, Jongsun Park 0001 |
IJCNN | 4 |
| 2021 | Low Cost Heterogeneous ARIA S-Box Implementation for CPA-Resistance
Junghoon Cho, Junhyun Song, Jongsun Park 0001 |
ISCAS | 3 |
| 2021 | Low Energy Domain Wall Memory Based Convolution Neural Network Design with Optimizing MAC ArchitectureabstractRunning a convolutional neural network (CNN) algorithm using dedicated integrated circuits (ICs) on real-time portable applications is mainly restricted by slow performance and large power consumption. The power and delay are mainly due to external memory access, which incurs considerable energy consumption and bandwidth issues. In this paper, we propose an efficient convolution layer design using domain wall memory (DWM) for eliminating external memory access in image sensor embedded applications. A low energy access scheme using tag is employed to further reduce power consumption. The experimental results show that the proposed CNN architecture achieves 11.2% memory energy savings and 21.8% of MAC operation reduction compared to conventional architecture. Jooyoon Kim, Yunho Jang, Jongsun Park 0001 |
ISCAS | 4 |
| 2021 | Exploiting Retraining-Based Mixed-Precision Quantization for Low-Cost DNN Accelerator DesignabstractFor successful deployment of deep neural networks (DNNs) on resource-constrained devices, retraining-based quantization has been widely adopted to reduce the number of DRAM accesses. By properly setting training parameters, such as batch size and learning rate, bit widths of both weights and activations can be uniformly quantized down to 4 bit while maintaining full precision accuracy. In this article, we present a retraining-based mixed-precision quantization approach and its customized DNN accelerator to achieve high energy efficiency. In the proposed quantization, in the middle of retraining, an additional bit (extra quantization level) is assigned to the weights that have shown frequent switching between two contiguous quantization levels since it means that both quantization levels cannot help to reduce quantization loss. We also mitigate the gradient noise that occurs in the retraining process by taking a lower learning rate near the quantization threshold. For the proposed novel mixed-precision quantized network (MPQ-network), we have implemented a customized accelerator using a 65-nm CMOS process. In the accelerator, the proposed processing elements (PEs) can be dynamically reconfigured to process variable bit widths from 2 to 4 bit for both weights and activations. The numerical results show that the proposed quantization can achieve 1.37 × better compression ratio for VGG-9 using CIFAR-10 data set compared with a uniform 4-bit (both weights and activations) model without loss of classification accuracy. The proposed accelerator also shows 1.29× of energy savings for VGG-9 using the CIFAR-10 data set over the state-of-the-art accelerator. Nahsung Kim, Dongyeob Shin, Wonseok Choi 0004, Geonho Kim, Jongsun Park 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | An Error Compensation Technique for Low-Voltage DNN AcceleratorsabstractReducing supply voltages of deep neural network (DNN) accelerators has been of particular interest since it can achieve high energy efficiency for mobile/edge applications. To ensure reliable DNN operations at low voltage, improving the timing error resilience of DNN accelerator is highly required. In this article, we present an error resilient technique to support low-voltage DNN operations by detecting and compensating erroneous computations using the proposed compensation multiply-accumulate (CMAC) unit. First, the timing errors are detected using Razor flip-flops at critical data-path, and erroneous computations are identified and dumped. Using additional multiplier data-path with flip-flops, the dropped computations are compensated in the next CMAC unit without additional clock-cycle penalty. Various bit-precisions of error compensations are analyzed to efficiently tradeoff DNN accuracy and hardware overhead. To improve the DNN accuracy even at low bit-precision of compensation, two types of rounding techniques are presented to effectively reflect the actual distribution of DNN computation results. The low-voltage DNN accelerator based on the proposed error compensation scheme has been implemented using 65-nm CMOS. Post-layout simulations show that the proposed DNN accelerator for ResNet-18 achieves about 47% and 24% energy savings compared with baseline and state-of-the-art error resilient DNN accelerators, respectively. Daehan Ji, Dongyeob Shin, Jongsun Park 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Bit Parallel 6T SRAM In-memory Computing with Reconfigurable Bit-PrecisionabstractThis paper presents 6T SRAM cell-based bit-parallel in-memory computing (IMC) architecture to support various computations with reconfigurable bit-precision. In the proposed technique, bit-line computation is performed with a short WL followed by BL boosting circuits, which can reduce BL computing delays. By per-forming carry-propagation between each near-memory circuit, bit-parallel complex computations are also enabled by iterating operations with low latency. In addition, reconfigurable bit-precision is also supported based on carry-propagation size. Our 128KB in/near memory computing architecture has been implemented using a 28nm CMOS process, and it can achieve 2.25GHz clock frequency at 0.9V with 5.2% of area overhead. The proposed architecture also achieves 0.68, 8.09 TOPS/W for the parallel addition and multiplication, respectively. In addition, the proposed work also supports a wide range of supply voltage, from 0.6V to 1.1V. Kyeongho Lee, Jinho Jeong 0002, Sungsoo Cheon, Woong Choi, Jongsun Park 0001 |
DAC | 5 |
| 2020 | Prediction Confidence based Low Complexity Gradient Computation for Accelerating DNN TrainingabstractIn deep neural network (DNN) training, network weights are iteratively updated with the weight gradients that are obtained from stochastic gradient descent (SGD). Since SGD inherently allows gradient calculations with noise, approximating weight gradient computations have a large potential of training energy/time savings without degrading accuracy. In this paper, we propose an input-dependent approximation of the weight gradient for improving energy efficiency of training process. Considering that the output predictions of network (confidence) changes with training inputs, the relation between the confidence and the magnitude of weight gradient can be efficiently exploited to skip the gradient computations without accuracy drop, especially for high confidence inputs. With a given squared error constraint, the computation skip rates can be also controlled by changing the confidence threshold. The simulation results show that our approach can skip 72.6% of gradient computations for CIFAR-100 dataset using ResNet-18 without accuracy degradation. Hardware implementation with 65nm CMOS process shows that our design achieves 88.84% and 98.16% of maximum per epoch training energy and time savings, respectively, for CIFAR-100 dataset using ResNet-18 compared to state-of-the-art training accelerator. Dongyeob Shin, Geonho Kim, Joongho Jo, Jongsun Park 0001 |
DAC | 4 |
| 2020 | Dynamic-Reference Based Early Write Termination for Low Energy SOT-MRAMabstractAlthough considered as one of the most viable emerging non-volatile memory, the spin-transfer-torque magnetic random access memory (STT-MRAM) suffers from its weakness in the write operation. Spin-orbit torque magnetic random access memory (SOT-MRAM) has been recently proposed to provide lower write energy consumption. Nevertheless, additional write energy reduction is still on demand for embedded memory purposes. In this paper, we propose an early write termination (EWT) technique for SOT-MRAM, which can greatly reduce the write energy consumption by efficiently removing the unnecessarily long write pulse. The proposed Dynamic Reference Early Termination (DRET) scheme provides energy savings in all write operations while guaranteeing reliable operation. Simulation results using 65nm CMOS technology show that 76.6% of write energy can be saved on average compared to the conventional SOT-MRAM. Eunjong Yeo, Yunho Jang, Yeongkyo Seo, Jongsun Park 0001 |
ISCAS | 5 |
| 2020 | Rank order coding based spiking convolutional neural network architecture with energy-efficient membrane voltage updates
Hoyoung Tang, Donghyeon Cho, Dongwoo Lew, Jongsun Park 0001 |
Neurocomputing | 5 |
| 2019 | Sensitivity based Error Resilient Techniques for Energy Efficient Deep Neural Network AcceleratorsabstractWith inherent algorithmic error resilience of deep neural networks (DNNs), supply voltage scaling could be a promising technique for energy efficient DNN accelerator design. In this paper, we propose novel error resilient techniques to enable aggressive voltage scaling by exploiting different amount of error resilience (sensitivity) with respect to DNN layers, filters, and channels. First, to rapidly evaluate filter/channel-level weight sensitivities of large scale DNNs, first-order Taylor expansion is used, which accurately approximates weight sensitivity from actual error injection simulation. With measured timing error probability of each multiply-accumulate (MAC) units considering process variations, the sensitivity variation among filter weights can be leveraged to design DNN accelerator, such that the computations with more sensitive weights are assigned to more robust MAC units, while those with less sensitive weights are assigned to less robust MAC units. Based on post-synthesis timing simulations, 51% energy savings has been achieved with CIFAR-10 dataset using VGG-9 compared to state-of-the-art timing error recovery technique with the same constraint of 3% accuracy loss. Wonseok Choi 0004, Dongyeob Shin, Jongsun Park 0001, Swaroop Ghosh |
DAC | 3 |
| 2019 | Low Cost Ternary Content Addressable Memory Based on Early Termination Precharge SchemeabstractIn this paper, we present early termination match-line (ML) precharge scheme for low power and high speed ternary content addressable memory (TCAM). In the proposed TCAM, by employing the pre-decision based early termination, unnecessary ML precharging has been effectively eliminated while improving the search speed and achieving error-free operation. The reference voltage generator used to implement the proposed early termination approach can be simply designed using dummy row without large area overhead. According to the post-layout simulations with the 65nm CMOS process, the proposed early termination ML precharge scheme shows up to 30.4% of sensing delay improvement and 65.9% of ML power savings compared to the conventional approach. It also shows 8% of FOM (energy/bit/search) improvement compared to state-of-the-art works. Kyeongho Lee, Geon Ko, Jongsun Park 0001 |
ISCAS | 3 |
| 2019 | An Energy-efficient On-chip Learning Architecture for STDP based Sparse CodingabstractTwo main bottlenecks encountered when implementing energy-efficient spike-timing-dependent plasticity (STDP) based sparse coding, are the complex computation of winner-take-all (WTA) operation and repetitive neuronal operations in the time domain processing. In this paper, we present an energy-efficient STDP based sparse coding processor. The low-cost hardware is based on algorithmic reduction techniques as following: First, the complex WTA operation is simplified based on the prediction of spike emitting neurons. Sparsity based approximation in spatial and temporal domain are also efficiently exploited to remove the redundant neurons with negligible algorithmic accuracy loss. We designed and implemented the hardware of the STDP based sparse coding using 65nm CMOS process. By exploiting input sparsity, the proposed architecture can dynamically trade off computation energy (up to 74%) with algorithmic quality for Natural image (maximum 3.55% quality loss) and MNIST (no quality loss) applications. In the inference mode of operations, the SNN hardware achieves the throughput of 374Mpixels/s and 840.2GSOP/s with energy-efficiency of 781.52pJ/pixel and 0.35pJ/SOP. Heetak Kim, Hoyoung Tang, Jongsun Park 0001 |
ISLPED | 3 |
| 2019 | Compact Implementations of HIGHT Block Cipher on IoT PlatformsabstractRecent lightweight block cipher competition (FELICS Triathlon) evaluates efficient implementations of block ciphers for Internet of things (IoT) environment. In the competition, the implementation of HIGHT block cipher achieved the most efficient lightweight block cipher, in terms of code size (ROM), memory (RAM), and execution time. In this paper, we further investigate lightweight features of HIGHT block cipher and present the optimized implementations of both software and hardware for low-end IoT platforms, including resource-constrained devices (8-bit AVR and 32-bit ARM Cortex-M3) and application-specific integrated circuit (ASIC). By using proposed optimization methods, the implemented HIGHT block cipher shows better performance compared to previous state-of-the-art implementations. Bohun Kim, Junghoon Cho, Byungjun Choi, Jongsun Park 0001, Hwajeong Seo |
Secur. Commun. Networks | 4 |
| 2019 | Charge-Recycling-Based Redundant Write Prevention Technique for Low-Power SOT-MRAMabstractIn spin-orbit torque magnetic random access memory (SOT-MRAM), as write energy is much larger than read energy, writing data to memory only when the new data is different from the stored data can lead to considerable write energy savings. In this paper, we propose three low-power techniques that can significantly reduce write energy in the read-compare-write process. In the proposed approaches, redundant charge in read operation has been efficiently reused to generate a negative voltage for write assistance, which leads to short write time. A selective precharging technique is also proposed to minimize the voltage swings between read and write operations. In addition, asymmetric write current due to source degeneration of write transistor can be resolved with the write voltage suppression scheme. Our circuit simulations with 65-nm CMOS technology show that when the stored data and new data are the same, up to 61% of write energy savings have been achieved compared with the conventional 2T-1MTJ cell. When the proposed SOT-MRAM is used as L3 caches of X86 processor, the gem5 simulations also show that average 48.2% of write energy savings can be achieved in various workloads of SPEC2006. Gyuseong Kang, Jongsun Park 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Content addressable memory based binarized neural network accelerator using time-domain signal processingabstractBinarized neural network (BNN) is one of the most promising solution for low-cost convolutional neural network acceleration. Since BNN is based on binarized bit-level operations, there exist great opportunities to reduce power-hungry data transfers and complex arithmetic operations. In this paper, we propose a content addressable memory (CAM) based BNN accelerator. By using time-domain signal processing, the huge convolution operations of BNN can be effectively replaced to the CAM search operation. In addition, thanks to fully parallel search of CAM, the parallel convolution operations for non-overlapped filtering window is enabled for high throughput data processing. To verify the effectiveness of the proposed CAM based BNN accelerator, the convolutional layer of LeNet-5 model has been implemented using 65nm CMOS technology. The implementation results show that the proposed BNN accelerator achieves 9.4% and 38.5% of area and energy savings, respectively. The parallel convolution operation of the proposed approach also shows 2.4x improved processing time. Woong Choi, Kwanghyo Jeong, Kyungrak Choi, Kyeongho Lee, Jongsun Park 0001 |
DAC | 5 |
| 2018 | Spike Counts Based Low Complexity Learning with Binary SynapseabstractTwo main difficulties encountered when implementing the neuromorphic system that supports pre- and post-synaptic textbf spike timing based real-time learning, are 1) huge memory size to store synaptic weights and 2) large amount of computations to update the synaptic weights. In this paper, we present a textbf spike counts based learning method that can significantly relieve the hardware burden of the real time unsupervised learning process. The novel learning approach uses both pre- and post-synaptic textbf spike counts as decision metrics in the following two ways: First, the synaptic weights are updated following the proposed simplified mean-based weight update rules, where binary feature images are produced based on pre-synaptic spike counts. Using the feature image, 1 bit synaptic weights are updated only once for each input image. When updating the weights, the post-synaptic spiking counts are also used to select the most active excitatory neuron. The actual weight updates are performed only for the weights connected to the dominant excitatory neuron, leading to further reduction of the weight updates without sacrificing accuracy. The simulation results show that using only 1 bit synaptic weights with 400 output neurons, the proposed learning approach achieves 82 % recognition accuracy with MNIST test set. In addition, the number of weight updates is reduced by 15.3 times compared to the state-of-the-art learning method. Hoyoung Tang, Heetak Kim, Donghyeon Cho, Jongsun Park 0001 |
IJCNN | 4 |
| 2018 | Low Cost Ternary Content Addressable Memory Using Adaptive Matchline Discharging SchemeabstractThis paper presents an adaptive match-line (ML) discharging scheme for low power and high speed ternary content addressable memory (TCAM). In the proposed TCAM, by employing the gated ML pulldown path and ML boosting scheme, the redundant ML discharging and SL switching are eliminated while improving the search speed. By considering the number of mismatch and ML discharging speed, the ML discharging is adaptively controlled in the proposed TCAM. The simulation results with the 65nm CMOS technology show that the proposed adaptive ML discharging scheme improves up to 19% of sensing delay and saves 81% of ML power compared to the conventional approach. When compared with the state-of-the-art work, the post-layout simulations show 10% improvement of FOM (energy/bit/search). Woong Choi, Kyeongho Lee, Jongsun Park 0001 |
ISCAS | 3 |
| 2018 | Charge-Recycling based Redundant Write Prevention Technique for Low Power SOT-MRAMabstractWhile the spin transfer torque magnetic memory (STT-MRAM) suffers from its shortcomings such as high write power, slow write operation and reliability issues, spin orbit torque magnetic random access memory (SOT-MRAM) can offer relatively faster write operation with low power based on giant spin hall effect. Although SOT-MRAM provides low power write operation, to meet the power level of current embedded memories, significant reduction of write power is highly required. In this paper, we present a low power write technique for SOT-MRAM. In order to prevent redundant write operation, read-compare-write operation is adopted. As a result, only the SOT cells having different data are written, and write power is saved in the cells with the same data. For further optimization, bitline switching scheme is used to reduce bitline and source line swing in write operation. The negative bitline scheme is also exploited by re-cycling the charge from read operation to increase write current. Simulation results using 65nm CMOS technology show that up to 40.1 % of write energy can be saved compared to the conventional unnecessary write avoidance approach. Gyuseong Kang, Yunho Jang, Jongsun Park 0001 |
ISCAS | 3 |
| 2018 | Spin Orbit Torque Device based Stochastic Multi-bit Synapses for On-chip STDP LearningabstractAs a large number of neurons and synapses are needed in spike neural network (SNN) design, emerging devices have been employed to implement synapses and neurons. In this paper, we present a stochastic multi-bit spin orbit torque (SOT) memory based synapse, where only one SOT device is switched for potentiation and depression using modified Gray code. The modified Gray code based approach needs only N devices to represent 2N levels of synapse weights. Early read termination scheme is also adopted to reduce the power consumption of training process by turning off less associated neurons and its ADCs. For MNIST dataset, with comparable classification accuracy, the proposed SNN architecture using 3-bit synapse achieves 68.7% reduction of ADC overhead compared to the conventional 8-level synapse. Gyuseong Kang, Yunho Jang, Jongsun Park 0001 |
ISLPED | 3 |
| 2018 | Energy Efficient Canny Edge Detector for Advanced Mobile Vision ApplicationsabstractIn this paper, we present an energy-efficient architecture of the Canny edge detector for advanced mobile vision applications. Three key techniques for reducing computational complexity of the Canny edge detector are presented. First, by exploiting the rank characteristic of the convolution kernel of Gaussian smoothing and Sobel gradient filters, common computations are identified and shared in the image filter design to reduce the number of additions and multiplications. For the gradient magnitude/direction computation, only three directions of neighboring pixels are considered to reduce computation energy with minor degradation on conformance performance (CP). For the adaptive threshold selections, an interesting observation is that the mean values of gradient magnitudes show small variations depending on the classified block types. Thus, the threshold selection process can be simplified as multiplying the mean value of the local block with predecided constants. The proposed low complexity Canny edge detector has been implemented using both field-programmable gate arrays (FPGAs) and a 65-nm standard-cell library. The FPGA implementation with Xilinx Virtex-V (XC5VSX240T) shows that our edge detector achieves 48% of area and 73% of execution time savings over the conventional architecture without seriously sacrificing the detection performance. The proposed edge detector implemented with 65-nm standard-cell library can easily support real-time ultrahigh definition video data processing (50 frames/s) with the power consumption of 5.48 mW (108.84 μJ/frame). Juseong Lee, Hoyoung Tang, Jongsun Park 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Bit-width reduction and customized register for low cost convolutional neural network acceleratorabstractThis paper presents a low area and energy efficient hardware accelerator for the deep convolutional neural networks (CNNs). Based on the multiply-accumulate (MAC) based architecture, three design techniques are proposed to reduce the hardware cost of the convolutional computations. First, to reduce the computational bit-width of convolutions, an adaptive bit-width reduction scheme is proposed based on differential input method. The bit-width reduction approach can reduce the 37 % of operation bit-width with almost ignorable CNN accuracy degradation. Second, it has been found that adapting bi-directional filtering window in CNN accelerator can considerably reduce the energy for data movement with much smaller number of memory accesses. To expedite the bi-directional filtering operations, we also propose a bidirectional first-input-first-output (bi-FIFO). With SRAM bit-cell layout manner, the proposed bi-FIFO facilitates fast data re-distribution with area and energy efficiency. To verify the effectiveness of the proposed techniques, the AlexNet accelerator has been designed. The numerical results show that the proposed adaptive bit-width reduction scheme achieves 25.9% and 47.3% of area and energy savings, respectively. The bi-FIFO based accelerator also achieves 33 % improved processing time. Kyungrak Choi, Woong Choi, Kyungho Shin, Jongsun Park 0001 |
ISLPED | 4 |
| 2017 | Improved Perturbation Vector Generation Method for Accurate SRAM Yield EstimationabstractAccurate yield estimation under parametric variation is one of the most integral parts for robust and nonwasted circuit design. In particular, due to the significant impact of disparity on the high-replication circuit, precise yield estimation is essential in SRAM design. In this paper, we propose an enhanced perturbation vector generation method to improve the accuracy of the yield estimation of the conventional direct SRAM yield computation method, which are access disturb margin (ADM) and write margin (WRM) first, by splitting the concave yield metric space, the estimation error caused by linear approximation can be significantly reduced with minor increase in simulation runtime. In addition, to compensate the inaccuracy of the conventional perturbation vector, a calibration method to reflect the multi-dc condition in SRAM assist operations is also proposed. Numerical results show that 37% improved estimation accuracy and 29% reduced estimation error can be achieved compared to the conventional ADM/WRM in the wide voltage range. Woong Choi, Jongsun Park 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Embedded DRAM-Based Memory Customization for Low-Cost FFT Processor DesignabstractIn this paper, we present embedded dynamic random access memory (eDRAM)-based memory customization techniques for low-cost fast Fourier transform (FFT) processor design. The main idea is based on the observation that the FFT processor has regular and predictable memory access patterns, and it can be efficiently exploited for memory customization using eDRAM. The memory customization approaches are applied to both of the pipelined and memory-based FFT architectures. In the pipelined architecture, the read wordline (RWL) coupling write assist and data packing schemes are employed to reduce the redundant RWL and wordline driving, respectively, in columninterleaved memory arrays. The memory address decoder is also simplified with thermometer code by exploiting the sequential access patterns. For the memory-based architecture, the modified cached-memory structure is employed in addition to the techniques used in the pipelined FFT architecture. The hardware implementation results of 2k-point FFT with a 0.11-um CMOS technology show that the proposed eDRAM-based pipelined and cached-memory FFTs achieve 26.8% and 33.2% power savings over the static RAM-based FFT design, respectively. Gyuseong Kang, Woong Choi, Jongsun Park 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Domain Wall Memory based Convolutional Neural Networks for Bit-width Extendability and Energy-EfficiencyabstractIn the hardware implementation of deep learning algorithms such as Convolutional Neural Networks (CNNs), vector-vector multiplications and memories for storing parameters take a significant portion of area and power consumption. In this paper, we propose a Domain Wall Memory (DWM) based design of CNN convolutional layer. In the proposed design, the resistive cell sensing mechanism is efficiently exploited to design a low-cost DWM-based cell arrays for storing parameters. The unique serial access mechanism and small footprint of DWM are also used to reduce the area and power cost of the input registers for aligning inputs. Contrary to the conventional implementation using Memristor-Based Crossbar (MBC), the bit-width of the proposed CNN convolutional layer is extendable for high resolution classifications and training. Simulation results using 65 nm CMOS process show that the proposed design archives 34% of energy savings compared to the conventional MBC based design approach. Jinil Chung, Jongsun Park 0001, Swaroop Ghosh |
ISLPED | 2 |
| 2016 | A 0.4-mW, 4.7-ps Resolution Single-Loop ΔΣ TDC Using a Half-Delay Time IntegratorabstractA compact, low-power, single-loop third-order delta-sigma (ΔΣ) time-to-digital converter (TDC) for time-mode signal processing is presented in this brief. In general, a high-resolution (ΔΣ) TDC requires a cascadable time integrator to increase the order of the loop filter. However, implementing the time integrator has been very challenging owing to the difficulty in storing time information. In this brief, we present a low-power half-delay time integrator, which is simply composed of two AND gates, a charge pump, and a comparator. The proposed time integrator can be easily cascaded (serially connected) to implement a loop filter with high-order noise shaping. The prototype TDC fabricated in 0.11-μm CMOS process occupies an active area of 0.11 mm2, consuming 0.4 mW from a 1.2 V supply. It achieves the dynamic range of 81 dB over a signal bandwidth of 50 kHz, and the resolution of 4.7 ps over a measurable range of 39.06 ns, which is half the clock period. Chan-Keun Kwon, Hoon Ki Kim, Jongsun Park 0001, Soo-Won Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Unequal-Error-Protection Error Correction Codes for the Embedded Memories in Digital Signal ProcessorsabstractIn many digital signal processing applications, some parts of a word stored in the embedded static random access memories (SRAMs) are more important than other parts of the word. Due to the differences in importance, memory failures that occur in more important bit locations generally give rise to relatively larger system performance degradation than those in less important locations. This brief presents a low-complexity unequal-error-protection error correcting code (UEEP-ECC) approach for the embedded memories in digital signal processor. In the proposed UEEP-ECC, repetition code is combined with the Bose-Chaudhuri-Hocquenghem code to selectively provide stronger error correction capabilities on more important data portions without a large hardware overhead. An efficient UEEP-ECC generation algorithm that can find the UEEP-ECC code with a minimum power of memory core and ECC logics is also presented. The experimental results show that the UEEP-ECC scheme achieves considerable power savings and data quality improvements in both of the H.264 and fast Fourier transform applications. Hoyoung Tang, Jongsun Park 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Domain wall memory based digital signal processors for area and energy-efficiencyabstractIn many Digital Signal Processing (DSP) applications such as Viterbi decoder and Fast Fourier Transform (FFT), Static Random Access Memory (SRAM) based embedded memory consumes significant portion of area and power. These DSP units are dominated by sequential memory access where SRAM-based memory is inefficient in terms of area and power. We propose spintronic Domain Wall Memory (DWM) based embedded memories for DSP building blocks e.g., survivor-path memories in Viterbi decoder and First-In-First-Out (FIFO) register files in FFT processor that exploit the unique serial access mechanism, non-volatility and small footprint of the memory for area and power saving. Simulations using 65nm technology show that the DWM based Viterbi decoder achieves 66.4 % area and 59.6 % power savings over the conventional SRAM-based implementation. For 8K point FFT processor, the DWM based design shows 60.6 % area and 60.3 % power savings. Jinil Chung, Kenneth Ramclam, Jongsun Park 0001, Swaroop Ghosh |
DAC | 3 |
| 2015 | Self-correcting STTRAM under magnetic field attacksabstractSpin-Transfer Torque Random Access Memory (STTRAM) is a possible candidate for universal memory due to its high-speed, low-power, non-volatility, and low cost. Although attractive, STTRAM is susceptible to contactless tampering through malicious exposure to magnetic field with the intention to steal or modify the bitcell content. In this paper, for the first time to our knowledge, we analyze the impact of magnetic attacks on STTRAM using micro-magnetic simulations. Next, we propose a novel array-based sensor to detect the polarity and magnitude of such attacks and then propose two design techniques to mitigate the attack, namely, array sleep with encoding and variable strength Error Correction Code (ECC). Simulation results indicate that the proposed sensor can reliably detect an attack and provide sufficient compensation window (few ns to ~100us) to enable proactive protection measures. Finally, we shows that variable-strength ECC can adapt correction capability to tolerate failures with various strength of an attack. Jae-Won Jang, Jongsun Park 0001, Swaroop Ghosh, Swarup Bhunia |
DAC | 2 |
| 2015 | A hybrid multimode BCH encoder architecture for area efficient re-encoding approachabstractThis paper presents a hybrid multimode Bose Chaudhuri Hocquenghem (BCH) encoder for reducing the input length of Syndrome calculation (SC) based on re-encoding approach. In previous re-encoding approaches, a conventional BCH encoder with long generator polynomials is used as a remainder operator to reduce the input length of SC. However, the input length is still large since long polynomial is used as a denominator of remainder operator for re-encoding. In the proposed approach, several minimal polynomials are employed as the denominators of remainder operators by utilizing the hardware of hybrid multimode BCH encoder. As a result, the minimum input length for SC can be employed for SC implementation through reencoding scheme, which leads to considerable area and latency reduction in SC module design. The proposed BCH encoder architecture and reduced SC modules are implemented using Samsung 65nm technology. The experimental results show that, in case of BCH (8640, 8192, 32) codes, the total area of SC modules are reduced by 96% compared to the previous re-encoding based SC module design, while the proposed multimode BCH encoder architecture also provides the reconfigurable error correction capability for 1 ≤ tsel≤ 32. Hoyoung Tang, Gihoon Jung, Jongsun Park 0001 |
ISCAS | 3 |
| 2014 | A low-complexity composite QR decomposition architecture for MIMO detectorabstractThis paper presents a low complexity QR decomposition (QRD) architecture for MIMO detector. In the proposed approach, various CORDIC-based QRD algorithms are efficiently combined together to reduce the computational complexity of the QRD hardware. Based on the computational complexity analysis on various QRD algorithms, a low complexity approach is selected at each stage of QRD process. The proposed QRD architecture can be applied to any arbitrary dimension of channel matrix, and the complexity reduction grows with the increasing matrix dimension. Our QR decomposition hardware was implemented using Samsung 0.13 μm technology. The numerical results show that the proposed architecture achieves 47% increase in the QAR (QRD Rate/Gate count) with 28.1% power savings over the conventional Householder CORDIC-based architecture for the 4x4 matrix decomposition. Ji-Hwan Yoon, Dongyeob Shin, Jongsun Park 0001 |
ISCAS | 3 |
| 2014 | Improving Energy Efficiency in FPGA Through Judicious Mapping of Computation to Embedded Memory BlocksabstractField-programmable gate arrays (FPGAs) are being increasingly used as a preferred prototyping and accelerator platform for diverse application domains, such as digital signal processing (DSP), security, and real-time multimedia processing. However, mapping of these applications to FPGA typically suffers from poor energy efficiency because of high energy overhead of programmable interconnects (PI) in FPGA devices. This paper presents an energy-efficient heterogenous application mapping framework in FPGA, where the conventional application mappings to logic and DSP blocks (for DSP-enhanced FPGA devices) are combined with judicious mapping of specific computations to embedded memory blocks. A complete mapping methodology including functional decomposition, fusion, and optimal packing of operations is proposed and efficiently used to reduce the large energy overhead of PIs. Effectiveness of the proposed methodology is verified for a set of common applications using a commercial FPGA system. Experimental results show that the proposed heterogenous mapping approach achieves significant energy improvement for different input bit-widths (e.g., more than 35% of energy savings with 8 bit or smaller bit inputs compared to the corresponding mapping in configurable logic blocks). For further reduction of energy, we propose an energy/accuracy tradeoff approach, where the input operand bit-width is dynamically truncated to reduce memory area and energy at the expense of modest degradation in output-accuracy. We show that using a preferential truncation method, up to 88.6% energy savings can be achieved in a 32-tap finite impulse response filter with modest impact on the filter performance. Anandaroop Ghosh, Somnath Paul, Jongsun Park 0001, Swarup Bhunia |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Reconfigurable CORDIC-Based Low-Power DCT Architecture Based on Data PriorityabstractThis paper presents a low-power coordinate rotation digital computer (CORDIC)-based reconfigurable discrete cosine transform (DCT) architecture. The main idea of this paper is based on the interesting fact that all the computations in DCT are not equally important in generating the frequency domain outputs. Considering the importance difference in the DCT coefficients, the number of CORDIC iterations can be dynamically changed to efficiently tradeoff image quality for power consumption. Thus, the computational energy can be significantly reduced without seriously compromising the image quality. The proposed CORDIC-based 2-D DCT architecture is implemented using 0.13 μm CMOS process, and the experimental results show that our reconfigurable DCT achieves power savings ranging from 22.9% to 52.2% over the CORDIC-based Loeffler DCT at the cost of minor image quality degradations. Ji-Hwan Yoon, Jongsun Park 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | Multidimensional Householder based high-speed QR decomposition architecture for MIMO receiversabstractConventional QR decomposition (QRD) hardware with a large size of channel matrix suffers from very low throughput and large latencies. This paper presents a high speed multi-dimensional (M-D) coordinate rotation digital computer (CORDIC) based QRD architecture. The novel high speed M-D architecture is enabled by exploiting multiple annihilations in a single CORDIC operation and removing data dependencies between two CORDIC operations (evaluation and application CORDIC) in Householder-based QRD process. The proposed QRD architecture can compute 4×4 complex R matrix for every 8 clock cycles. Our QRD hardware for 4×4 channel matrix was implemented using Samsung 0.13μm CMOS process, and the experimental results show that the proposed architecture achieves 4.74x speed-up compared to the conventional hybrid M-D based QRD. Iput Heri Kurniawan, Ji-Hwan Yoon, Jongsun Park 0001 |
ISCAS | 3 |
| 2013 | Design and Implementation of an On-Chip Permutation Network for Multiprocessor System-On-ChipabstractThis paper presents the silicon-proven design of a novel on-chip network to support guaranteed traffic permutation in multiprocessor system-on-chip applications. The proposed network employs a pipelined circuit-switching approach combined with a dynamic path-setup scheme under a multistage network topology. The dynamic path-setup scheme enables runtime path arrangement for arbitrary traffic permutations. The circuit-switching approach offers a guarantee of permuted data and its compact overhead enables the benefit of stacking multiple networks. A 0.13-μ m CMOS test-chip validates the feasibility and efficiency of the proposed design. Experimental results show that the proposed on-chip network achieves 1.9× to 8.2× reduction of silicon overhead compared to other design approaches. Phi-Hung Pham, Junyoung Song, Jongsun Park 0001, Chulwoo Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | Dual queue based rate selecting schedule for throughput enhancement in WLANsabstractIn IEEE 802.11 WLANs, a fixed low transmission rate is used for multicast transmissions regardless of channel conditions of receivers. This may lead to inefficient use of wireless channel, especially when access point (AP) has both unicast and multicast data frames in the transmission queue. In this paper, we propose a dual queue based rate selecting schedule scheme for efficient use of wireless channel. In contrast to the legacy WLAN AP having a single FIFO queue, AP in the proposed scheme maintains two separate queues at the AP, one for unicast data and another for multicast data. When the AP accesses wireless channel, it decides a type of data (unicast or multicast) and transmission rate to be sent according to the channel condition and the delay boundary of multicast data. Our simulation results show that the proposed scheme increases the average data rate and energy consumption saving up to 13.7% compared to the auto rate fallback (ARF) which determines the transmission data rate only considering the channel condition. Dongwan Kim, Wan-Seon Lim, Jongsun Park 0001 |
ISCAS | 3 |
| 2012 | High-speed tournament givens rotation-based QR Decomposition Architecture for MIMO ReceiverabstractThis paper presents a high-speed hardware architecture of an improved Givens rotation-based QR decomposition, named tournament-based complex Givens rotation (T-CGR). In the proposed approach, more than one pivots are selected and zero-insertion processes of Givens-rotations are performed in parallel like tournament in order to increase the throughput. As a result, the QR decomposition performance significantly increases compared to the triangular systolic array (TSA) approach. Moreover, the circuit area was reduced due to the smaller number of flip-flops for holding the computed results during the decomposition process. The proposed QR decomposition hardware was implemented using TSMC 0.25 um technology. The experimental results show that the proposed architecture achieves 73.00% speed-up over the TACR/TSA-based architecture for the 8 × 8 matrix decomposition. Ji-Hwan Yoon, Jongsun Park 0001 |
ISCAS | 3 |
| 2012 | Resource Efficient Implementation of Low Power MB-OFDM PHY Baseband Modem With Highly Parallel ArchitectureabstractThe multi-band orthogonal frequency-division multiplexing modem needs to process large amount of computations in short time for support of high data rates, i.e., up to 480 Mbps. In order to satisfy the performance requirement while reducing power consumption, a multi-way parallel architecture has been proposed. But the use of the high degree parallel architecture would increase chip resource significantly, thus a resource efficient design is essential. In this paper, we introduce several novel optimization techniques for resource efficient implementation of the baseband modem which has highly, i.e., 8-way, parallel architecture, such as new processing structures for a (de)interleaver and a packet synchronizer and algorithm reconstruction for a carrier frequency offset compensator. Also, we describe how to efficiently design several other components. The detailed analysis shows that our optimization technique could reduce the gate count by 27.6% on average, while none of techniques degraded the overall system performance. With 0.18-μm CMOS process, the gate count and power consumption of the entire baseband modem were about 785 kgates and less than 381 mW at 66 MHz clock rate, respectively. Seokjoong Hwang, Youngsun Han, Seon Wook Kim, Jongsun Park 0001, Byung Gueon Min |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | Design and Implementation of Backtracking Wave-Pipeline Switch to Support Guaranteed Throughput in Network-on-ChipabstractIt is a challenging task in a network-on-chip to design an on-chip switch/router to dynamically support (hard) guaranteed throughput under very tight on-chip constraints of power, timing, area, and time-to-market. This paper presents the design and implementation of a novel pipeline circuit-switched switch to support guaranteed throughput. The proposed circuit-switched switch, based on a backtracking probing path setup, operates with a source-synchronous wave-pipeline approach. The switch can support a dead- and live-lock free dynamic path-setup scheme and can achieve high bandwidth and high area and energy efficiency. A silicon-proven prototype of a 16-bit-data 5-bidirectional-port switch in a four-metal-layer 0.18-μ m CMOS standard-cell technology can yield an aggregate data bandwidth of up to 73.84 Gb/s, while occupying only a modest area of 0.0315 mm2. The synthesizable implementation of the proposed switch also results in a cost-effective design, fast development time, and portability. Phi-Hung Pham, Jongsun Park 0001, Phuong Mau, Chulwoo Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | Improved MIMO SIC Detection Exploiting ML CriterionabstractIn this paper, we propose an improved MIMO successive interference cancellation (SIC) detector taking maximum likelihood (ML) criterion into account. Applying the ML criterion to multiple candidates obtained from all possible orderings of streams in the SIC, the proposed method can provide improved performance with complexity comparable to conventional ordered SIC for moderate number of spatial streams. Ji-Woong Choi, Hui-Ling Lou, Jongsun Park 0001 |
VTC Fall | 4 |
| 2011 | A Reconfigurable FIR Filter Architecture to Trade Off Filter Performance for Dynamic Power ConsumptionabstractThis paper presents an architectural approach to the design of low power reconfigurable finite impulse response (FIR) filter. The approach is well suited when the filter order is fixed and not changed for particular applications, and efficient trade-off between power savings and filter performance can be made using the proposed architecture. Generally, FIR filter has large amplitude variations in input data and coefficients. Considering the amplitude of both the filter coefficients and inputs, the proposed FIR filter dynamically changes the filter order. Mathematical analysis on power savings and filter performance degradation and its experimental results show that the proposed approach achieves significant power savings without seriously compromising the filter performance. The power savings is up to 41.9% with minor performance degradation, and the area overhead of the proposed scheme is less than 5.3% compared to the conventional approach. Seok-Jae Lee, Ji-Woong Choi, Seon Wook Kim, Jongsun Park 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | Dynamic Bit-Width Adaptation in DCT: An Approach to Trade Off Image Quality and Computation EnergyabstractThis paper presents a dynamic bit-width adaptation scheme for applications using discrete cosine transform (DCT). The technique can efficiently trade off image quality and computation energy. Based on sensitivity differences of 64 DCT coefficients, separate operand bit-widths are used for different frequency components to reduce computation energy. To select the appropriate operand bit-widths that achieve significant reduction of power consumption with minimum image quality degradation, we also propose a bit-width selection algorithm. The proposed variable bit precision DCT algorithm can be efficiently implemented using carry save adder trees. The reconfigurable DCT architecture can achieve power savings ranging from 36% to 75% compared to normal operation at the expense of minor image quality degradation. Jongsun Park 0001, Jung Hwan Choi, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2006 | Dynamic bit-width adaptation in DCT: image quality versus computation energy trade-offabstractWe present a dynamic bit-width adaptation scheme in DCT applications for efficient trade-off between image quality and computation energy. Based on sensitivity differences of 64 DCT coefficients, various operand bit-widths are used for different frequency components to reduce computation energy in DCT operation. Numerical results show that our DCT architecture can achieve power savings ranging from 36% to 75% compared to normal operation. Jongsun Park 0001, Jung Hwan Choi, Kaushik Roy 0001 |
DATE | 1 |
| 2006 | Efficient modeling of 1/falpha/ noise using multirate processabstractIn order to verify the system performance of mixed-signal systems on chip (SoCs), computer-aided design (CAD) tools are required to generate 1/f/sup /spl alpha// noise that degrades the performance of most analog circuits. Current techniques for generating discrete sequences of 1/f/sup /spl alpha// noise require a large amount of computations that place an excessive burden on the computation engine and random number generators. In this paper, the authors propose a low-complexity 1/f/sup /spl alpha// noise generation scheme, which is based on a multirate filter bank. In this scheme, each branch in the filter bank processes signals in a different frequency band while allowing for arbitrary selection of /spl alpha/ in each bank. The proposed approach greatly reduces computations when compared to traditional noise generation processes of using a single noise-shaping filter. Furthermore, it allows selecting different combinations of noise frequency response in different frequency bands, thus allowing calibration of noise generated in simulation to the one measured in the laboratory from test chips. A comparison of various noise generation schemes is also presented. Jongsun Park 0001, Khurram Muhammad, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2004 | A low power reconfigurable DCT architecture to trade off image quality for computational complexityabstractWe present a low power reconfigurable DCT design, which achieves considerable computational complexity reduction in DCT operation with minimum image quality degradation. The approach is based on the modification of DCT bases in a bit-wise manner. Different computational complexity/image quality trade off levels are presented and a reconfigurable architecture. which can dynamically change from one trade off level to another, is also proposed. The reconfigurable DCT architecture can achieve power savings ranging from 20% to 70% for 5 different trade off levels. Jongsun Park 0001, Kaushik Roy 0001 |
ICASSP (5) | 1 |
| 2004 | Hardware architecture and VLSI implementation of a low-power high-performance polyphase channelizer with applications to subband adaptive filteringabstractThe polyphase channelizer is an important component of a subband adaptive filtering system. This paper presents an efficient hardware architecture and VLSI implementation of a low-power high-performance polyphase channelizer, integrating optimizations at algorithmic, architectural and circuit level. At the algorithm level, a computationally efficient structure is derived. Tradeoffs between hardware complexity and system performance are explored during the fixed-point modeling of the system. A computational complexity reduction technique is also employed to reduce the complexity of the hardware architecture. Circuit-level optimizations, including an efficient commutator implementation, dual-VDD scheme and novel level-converting flip-flops, are also integrated. Simulation results show that the design consumes 352 mW power with system throughput of 480 million samples per second (MSPS). A test chip has been submitted for fabrication to validate the proposed hardware architecture and VLSI design techniques. Yongtao Wang, Hamid Mahmoodi, Lih-Yih Chiou, Hunsoo Choo, Jongsun Park 0001, Woopyo Jeong, Kaushik Roy 0001 |
ICASSP (5) | 5 |
| 2003 | High-performance FIR filter design based on sharing multiplicationabstractFinite impulse response (FIR) filtering can be expressed as multiplications of vectors by scalars. We present high-speed designs for FIR filters based on a computation sharing multiplier which specifically targets computation re-use in vector-scalar products. The performance of the proposed implementation is compared with implementations based on carry-save and Wallace tree multipliers in 0.35-/spl mu/m technology. We show that sharing multiplier scheme improves speed by approximately 52 and 33% with respect to the FIR filter implementations based on the carry-save multiplier and Wallace tree multiplier, respectively. In addition, sharing multiplier scheme has a relatively small power delay product than other multiplier schemes. Using voltage scaling, power consumption of the FIR filter based on computation sharing multiplier can be reduced to 41% of the FIR filter based on the Wallace tree multiplier for the same frequency of operation. Jongsun Park 0001, Khurram Muhammad, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2002 | Low power reconfigurable DCT design based on sharing multiplicationabstractWe present a low-power reconfigurable OCT architecture, which is based on computation sharing multiplier (CSHM). CSHM specifically targets computation re-use in vector-scalar products and is effectively used in our OCT implementation. A low power reconfigurable OCT architecture is exploited by making a trade off between image quality and power consumption. The proposed OCT architecture was implemented using 0.35µ technology. The experimental results show that reconfigurable OCT using CSHM can improve power consumption by 40 % without noticeable image quality degradation. Jongsun Park 0001, Soonkeon Kwon, Kaushik Roy 0001 |
ICASSP | 1 |
| 2002 | High performance and low power FIR filter design based on sharing multiplicationabstractWe present a high performance and low power FIR filter design, which is based on computation sharing multiplier (CSHM). CSHM specifically targets computation re-use in vector-scalar products and is effectively used in our FIR filter design. Efficient circuit level techniques: a new carry select adder and conditional capture flip-flop (CCFF), are also used to further improve power and performance. The proposed FIR filter architecture was implemented in 0.25 μm technology. Experimental results on a 10 tap low pass CSHM FIR filter show speed and power improvement of 19% and 17%, respectively, with respect to an FIR filter based on Wallace tree multiplier. Jongsun Park 0001, Woopyo Jeong, Hunsoo Choo, Hamid Mahmoodi, Yongtao Wang, Kaushik Roy 0001 |
ISLPED | 1 |
| 2000 | Non-adaptive and adaptive filter implementation based on sharing multiplicationabstractFIR filtering can be expressed as multiplication of a vector by scalars. We present high-speed implementations for adaptive and nonadaptive filters based on a computation sharing multiplier which specifically targets computation re-use in vector-scalar products. The performance of the proposed implementation is compared with implementations based on carry save and Wallace tree multipliers in 0.6 /spl mu/ technology. We show that the sharing multiplier scheme improves speed by approximately 30% and 21% with respect to the Wallace tree multiplier based implementation for non-adaptive and adaptive filters, respectively. Jongsun Park 0001, Hunsoo Choo, Khurram Muhammad, Seung Hoon Choi, Yonghee Im, Kaushik Roy 0001 |
ICASSP | 1 |