VLDB 2026 Research / reviewers in the wild / expert
Mingu Kang
dblp:50/7194
· DBLP profile ↗
31ranked-venue papers
7as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 1 first-author · 16 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NOVA-PIM: Noise-Aware Hyperdimensional Processing in Memory with Optimized Vector Allocation and Minimal ADCsabstractHyperdimensional computing (HDC) is an emerging brain-inspired paradigm that enables highly efficient and robust inference and learning. Analog processing in memory (PIM) has become a promising solution to accelerate HDC by processing lengthy hypervectors (HVs) directly in memory, thereby reducing costly data movement and leveraging massive parallelism. Despite its efficiency, analog PIM suffers from non-idealities that reduce reliability and accuracy. Although the similarity search stage in HDC is inherently error-tolerant given the high dimensionality of HVs, the encoding stage, which transforms raw input data into HVs, remains sensitive to analog noise. Moreover, encoding accounts for a dominant portion of energy consumption, creating a long-standing bottleneck that limits the overall efficiency of analog PIM-based HDC systems. To overcome this challenge, we propose a noise-aware partitioning scheme that improves HDC inference accuracy by processing a critical subset of HV dimensions digitally, while offloading most of the non-critical dimensions to analog PIM. To further synergize the PIM operations across the two consecutive stages, we eliminate the analog-to-digital converters (ADCs) overhead for encoding by employing pulse width modulation (PWM), allowing direct interfacing with the subsequent similarity search stage. The proposed system achieves a 2.6 × reduction in area, 1.5 × –10.3 × lower energy consumption, and 4.2 × –6.5 × speedup compared with state-of-the-art, while maintaining inference accuracy. Keming Fan, Chang Eun Song, Xuan Wang 0040, Tajana Rosing, Mingu Kang |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | TRIM: Acceleration of Multiplication-Less Neural Networks via Versatile SparsitiesabstractRecently, multiplication-less neural networks (L1NNs), replacing multiplication-intensive dot products with addition-only$\ell _{1}$-distance kernels have emerged to enhance energy efficiency and speed with negligible degradation of accuracy. Despite these gains, such models suffer from pruning challenges that can negate their benefits. In this work, we identify the root cause of these challenges and introduce a novel method called synapse pruning, which is explicitly designed for L1NNs to overcome them for the first time. Building upon this algorithmic innovation, we present TRIM, an algorithm-hardware co-design framework that sparsifies and accelerates L1NNs to achieve ultra-high energy efficiency and speed. On the algorithmic side, we propose structured synapse pruning tailored for hardware-friendly$\mathbb {N:M}$sparsity patterns for L1NNs. On the hardware side, we introduce a sparse processor architecture that efficiently exploits both the proposed$\mathbb {N:M}$structured synapse sparsity and the intrinsic unstructured weight and activation sparsities in L1NNs by skipping redundant operations. Additionally, an$\mathbb {N:M}$sparsity-aware elastic mapping technique is introduced to maximize hardware utilization and data reuse. Evaluations on seven benchmarks demonstrate that TRIM achieves up to 87.5% unstructured sparsity and up to 81.3% structured sparsity with less than 1% accuracy loss. Implemented in a 65nm technology, our processor achieves 5.22 TOPS/W and 1.17 TOPS/mm2, surpassing the existing SOTA L1NN accelerator by$1.7\times $in energy efficiency and$14.9\times $in area efficiency. The source code is available at:https://github.com/XZH28/TRIM-Sparse-L1-Distance-Net.git Zihan Xia 0002, Mingu Kang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | High-throughput Point-Cloud Accelerator with Sparsity-aware Hierarchical Neighbor Voxel Search and SkippingabstractPoint cloud-based 3D sparse convolution networks are widely employed to process voxel features efficiently. However, the irregularity of voxel sparsity poses significant challenges, leading to increased hardware complexity and inefficiencies. We propose an algorithm-hardware co-design for sparse 3D convolution. At the algorithm level, an on-the-fly thresholdbased voxel skipping is adopted, enhancing efficiency. At the hardware level, a hierarchical 3-stage Voxel Search and Skipping is developed to systematically narrow down the non-zero search space, enhancing both performance and hardware utilization. We implemented the proposed accelerator in a 65 nm process to demonstrate a 77.7% reduction in delay compared to a baseline design, which does not support the proposed sparsity adaptations. The proposed system also achieved the $1.34 \times$ and $2.22 \times$ higher energy efficiency and throughput as compared to the state-ofarts. Yun-Chia Yu, Suraj Pn Reddy, Aryan Devrani, Anirudh Srinivasan, Saianudeep Reddy Nayini, Sohyeon Kim, Sung-Joon Jang, Sang-Seol Lee, Mingu Kang |
DAC | 9 |
| 2025 | Clo-HDnn: Continual On-Device Learning Accelerator with Hyperdimensional Computing via Progressive SearchabstractClo-HDnn is an on-device learning (ODL) accelerator designed for emerging continual learning (CL) tasks. Clo-HDnn integrates hyperdimensional computing (HDC) along with low-cost Kronecker HD Encoder and weight clustering feature extraction (WCFE) to optimize accuracy and efficiency. Clo-HDnn adopts gradient-free CL to efficiently update and store the learned knowledge in the form of class hypervectors. Its dual-mode operation enables bypassing costly feature ex- traction for simpler datasets, while progressive search reduces complexity by up to $61 \%$ by encoding and comparing only partial query hypervectors. Achieving 4.66 TFLOPS/W (FE) and 3.78 TOPS/W (classifier), Clo-HDnn delivers $7.77 \times$ and $4.85 \times$ higher energy efficiency compared to SOTA ODL accelerators. Chang Eun Song, Keming Fan, Soumil Jain, Gopabandhu Hota, Haichao Yang, Leo Liu, Meng-Fan Chang, Carlos H. Diaz, Gert Cauwenberghs, Tajana Rosing, Mingu Kang |
HCS | 12 |
| 2025 | FeNOMS: Enhancing Open Modification Spectral Library Search with In-Storage Processing on Ferroelectric NAND (FeNAND) FlashabstractThe rapid expansion of mass spectrometry (MS) data, now exceeding hundreds of terabytes, poses significant challenges for efficient, large-scale library search — a critical component for drug discovery. Traditional processors struggle to handle this data volume efficiently, making in-storage computing (ISP) a promising alternative. This work introduces an ISP architecture leveraging a 3D Ferroelectric NAND (FeNAND) structure, providing significantly higher density, faster speeds, and lower voltage requirements compared to traditional NAND flash. Despite its superior density, the NAND structure has not been widely utilized in ISP applications due to limited throughput associated with row-by-row reads from serially connected cells. To overcome these limitations, we integrate hyperdimensional computing (HDC), a brain-inspired paradigm that enables highly parallel processing with simple operations and strong error tolerance. By combining HDC with the proposed dual-bound approximate matching (D-BAM) distance metric, tailored to the FeNAND structure, we parallelize vector computations to enable efficient MS spectral library search, achieving 43× speedup and 21× higher energy efficiency over state-of-the-art 3D NAND methods, while maintaining comparable accuracy. Sumukh Pinge, Ashkan Moradifirouzabadi, Keming Fan, Prasanna Venkatesan Ravindran, Tanvir H. Pantha, Po-Kai Hsu, Zihan Xia 0002, Flavio Ponzina, Winston Chern, Taeyoung Song, Priyankka Gundlapudi Ravikumar, Mengkun Tian, Lance Fernandes, Hari Jayasankar, Chinsung Park, Amrit Garlapati, Kijoon Kim, Jongho Woo, Suhwan Lim, Wanki Kim, Daewon Ha, Duygu Kuzum, Shimeng Yu, Tajana Rosing, Mingu Kang |
ICCAD | 32 |
| 2025 | DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow DecodingabstractHuman motion is inherently continuous and dynamic, posing significant challenges for generative models. While discrete generation methods are widely used, they suffer from limited expressiveness and frame-wise noise artifacts. In contrast, continuous approaches produce smoother, more natural motion but often struggle to adhere to conditioning signals due to high-dimensional complexity and limited training data. To resolve this 'discord' between discrete and continuous representations we introduce DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding, a novel method that leverages rectified flow to decode discrete motion tokens in the continuous, raw motion space. Our core idea is to frame token decoding as a conditional generation task, ensuring that DisCoRD captures fine-grained dynamics and achieves smoother, more natural motions. Compatible with any discrete-based framework, our method enhances naturalness without compromising faithfulness to the conditioning signals on diverse settings. Extensive evaluations demonstrate that DisCoRD achieves state-of-the-art performance, with FID of 0.032 on HumanML3D and 0.169 on KIT-ML. These results establish DisCoRD as a robust solution for bridging the divide between discrete efficiency and continuous realism. Project website: https://whwjdqls.github.io/discord-motion/ Jungbin Cho, Junwan Kim, Jisoo Kim 0006, Mingu Kang, Sungeun Hong, Tae-Hyun Oh, Youngjae Yu |
ICCV | 5 |
| 2025 | Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient RedistributionabstractTransformers, while revolutionary, face challenges due to their demanding computational cost and large data movement.To address this, we propose HyFlexPIM, a novel mixed-signal processingin-memory (PIM) accelerator for inference that flexibly utilizes both single-level cell (SLC) and multi-level cell (MLC) RRAM technologies to trade-off accuracy and efficiency.HyFlexPIM achieves efficient dual-mode operation by utilizing digital PIM for highprecision and write-intensive operations while analog PIM for high parallel and low-precision computations.The analog PIM further distributes tasks between SLC and MLC PIM operations, where a single analog PIM module can be reconfigured to switch between two operations (SLC/MLC) with minimal overhead (<1% for area & energy).Critical weights are allocated to SLC RRAM for high accuracy, while less critical weights are assigned to MLC RRAM to maximize capacity, power, and latency efficiency.However, despite employing such a hybrid mechanism, brute-force mapping on hardware fails to deliver significant benefits due to the limited proportion of weights accelerated by the MLC and the noticeable degradation in accuracy.To maximize the potential of our hybrid hardware architecture, we propose an algorithm co-optimization technique, called gradient redistribution, which uses Singular Value Decomposition (SVD) to decompose and truncate matrices based on their importance, then fine-tune them to concentrate significance into a small subset of weights.By doing so, only 5-10% of the weights have dominantly large gradients, making it favorable for HyFlexPIM by minimizing the use of expensive SLC RRAM while maximizing the efficient MLC RRAM.Our evaluation shows that HyFlexPIM significantly enhances computational throughput and energy efficiency, achieving maximum 1.86× and 1.45× higher than state-of-the-art methods. Chang Eun Song, Priyansh Bhatnagar, Zihan Xia 0002, Nam Sung Kim, Tajana Rosing, Mingu Kang |
ISCA | 6 |
| 2025 | SmartMS: Efficient Hierarchical Database Search for Mass Spectrometry via Processing-in-MemoryabstractThe acceleration of Mass Spectrometry (MS) library search is crucial for advancing scientific and pharmaceutical research. Recent methodologies leverage Hyperdimensional Computing (HDC) to encode reference and query spectra as high-dimensional vectors, enabling highly parallel similarity computations. In this context, Processing-In-Memory (PIM) has emerged as a promising solution, offering orders of magnitude improvements in computational speed compared to GPU-based approaches when handling large-scale libraries. However, bruteforce search methods remain computationally intensive, exacerbating the high energy demands associated with MS library search operations in data centers. In this work, we propose SmartMS, a novel tool that leverages HDC to construct a multi-level database structure, reducing search complexity from linear to logarithmic while maintaining compatibility with PIM-based accelerators. SmartMS improves identification accuracy by 3% while delivering a 33× improvement in speed and a 58× energy reduction, with a negligible increase in memory requirements of 0.5% when compared to the current state of the art. Flavio Ponzina, Sumukh Pinge, Abhijay Deevi, Yilin Ge, Mingu Kang, Tajana Rosing |
ISLPED | 6 |
| 2025 | Stratum: System-Hardware Co-Design with Tiered Monolithic 3D-Stackable DRAM for Efficient MoE ServingabstractAs Large Language Models (LLMs) continue to evolve, Mixture of Experts (MoE) architecture has emerged as a prevailing design for achieving state-of-the-art performance across a wide range of tasks.MoE models use sparse gating to activate only a handful of expert sub-networks per input, achieving billion-parameter capacity with inference costs akin to much smaller models.However, such models often pose challenges for hardware deployment due to the massive data volume introduced by the MoE layers.To address the challenges of serving MoE models, we propose Stratum, a system-hardware co-design approach that combines the novel memory technology Monolithic 3D-Stackable DRAM (Mono3D DRAM), near-memory processing (NMP), and GPU acceleration.The logic and Mono3D DRAM dies are connected through hybrid bonding, whereas the Mono3D DRAM stack and GPU are interconnected via silicon interposer.Mono3D DRAM offers higher internal bandwidth than HBM thanks to the dense vertical interconnect pitch enabled by its monolithic structure, which supports implementations of higher-performance near-memory processing.Furthermore, we tackle the latency differences introduced by aggressive vertical scaling of Mono3D DRAM along the 𝑧-dimension by constructing internal memory tiers and assigning data across layers based on * Equal contribution Yue Pan 0009, Zihan Xia 0002, Po-Kai Hsu, Lanxiang Hu, Hyungyo Kim, Janak Sharda, Minxuan Zhou, Nam Sung Kim, Shimeng Yu, Tajana Rosing, Mingu Kang |
MICRO | 11 |
| 2025 | Exploiting Chiplet Integration Technology for Fast High-Capacity DRAM ModulesabstractAs the end of Moore’s law approaches, chiplet integration technology (or chiplet technology) has emerged to revolutionize future semiconductor chip design. Chiplet technology provides unique advantages over 3-D-stacking technology, including a more cost-efficient and thermal-friendly integration of heterogeneous technologies. Although chiplet technologies have already begun to be used by the latest commercial chips, they have not been explored for commodity dynamic random access memory (DRAM) design yet. Harnessing its advantages for DRAM for the first time, this article evaluates the feasibility of chiplet-based DRAM architecture, considering various physical and electrical constraints imposed by a standard chiplet interface [i.e., universal chiplet interconnect express (UCIe)]. We further explore the DIMM architectures that simplify module packaging and assembly, leading to reductions in total die size and overall costs. The comprehensive cross-level analysis (i.e., device, circuit, chip, and system levels) shows that chiplet-based DRAM reducest_RCD+t_CAS, latency-critical DRAM timing parameters, by$1.32\times $–$1.39\times $, at the same energy consumption. In addition, a$1.39\times $–$2.28\times $improvement int_RRDis obtained. The reduced DRAM timing parameters improve the overall system performance by up to 8.8%–24.7% (geomean 3.4%–8.4%) in real-life benchmarks. The chiplet-based heterogeneous integration achieves a$1.27\times $higher chip-level yield compared with the monolithic chip, along with up to 10% reduction in overall cost compared with traditional DIMMs at emerging process technologies. Zihan Xia 0002, Chihun Song, Ram Krishna, Ashita Victor, Srujan Penta, Muhannad S. Bakir, Elyse Rosenbaum, Nam Sung Kim, Mingu Kang |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2024 | LEAF: An Adaptation Framework against Noisy Data on Edge through Ultra Low-Cost TrainingabstractThe deployment of machine learning in real-life computer vision applications is often subject to the degraded input data by various noise sources and time-varying conditions. Continual re-training is one potential solution to adapt to the real-time noise sources and achieve the target performance. However, this approach requires massive computing complexity for the training, particularly demanding on resource-constrained edge devices. To address this challenge, for the first time, we present a hardware-efficient framework named LEAF, tailored for the adaptation to degraded images (ADI) scenario. Leveraging both qualitative and quantitative profiling of CNN dynamics when operating on degraded images, we propose two techniques: 1) selective experience replay (SER)-based unimportant image skipping to reduce both the forward pass (FwP) and backward pass (BwP) costs, and 2) pseudo noise dithering (PND)-assisted extremely low precision quantization (3/4-bit) for error gradients to enable nearly full integer computing. Extensive experiments on CIFAR10 and Tiny ImageNet datasets employing VGGNet and ResNet, along with five different types of image degradations with varying strengths demonstrate virtually no accuracy degradation (< 0.5%) at 3/4-bits with 3.64 ~ 8.40× reduced forward and backward processing. Zihan Xia 0005, Mingu Kang |
DAC | 3 |
| 2024 | SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Drivingabstract3D object detection using point cloud (PC) data is essential for perception pipelines of autonomous driving, where efficient encoding is key to meeting stringent resource and latency requirements. PointPillars, a widely adopted bird's-eye view (BEV) encoding, aggregates 3D point cloud data into 2D pillars for fast and accurate 3D object detection. However, the stateof-the-art methods employing PointPillars overlook the inherent sparsity of pillar encoding where only a valid pillar is encoded with a vector of channel elements, missing opportunities for significant computational reduction. Meanwhile, current sparse convolution accelerators are designed to handle only elementwise activation sparsity and do not effectively address the vector sparsity imposed by pillar encoding. In this paper, we propose SPADE, an algorithm-hardware codesign strategy to maximize vector sparsity in pillar-based 3D object detection and accelerate vector-sparse convolution commensurate with the improved sparsity. SPADE consists of three components: (1) a dynamic vector pruning algorithm balancing accuracy and computation savings from vector sparsity, (2) a sparse coordinate management hardware transforming 2D systolic array into a vector-sparse convolution accelerator, and (3) sparsityaware dataflow optimization tailoring sparse convolution schedules for hardware efficiency. Taped-out with a commercial technology, SPADE saves the amount of computation by 36.3–89.2% for representative 3D object detection networks and benchmarks, leading to 1.3–10.9 × speedup and 1.5–12.6 × energy savings compared to the ideal dense accelerator design. These sparsityproportional performance gains equate to 4.1–28.8 × speedup and 90.2–372.3 × energy savings compared to the counterpart server and edge platforms. Seongmin Park 0003, Minyong Yoon, Janghwan Lee, Nam Sung Kim, Mingu Kang, Jungwook Choi |
HPCA | 8 |
| 2024 | Efficient Transformer Acceleration via Reconfiguration for Encoder and Decoder Models and Sparsity-Aware Algorithm MappingabstractTwo essential computing blocks of Transformers, encoder and decoder, used for summarization and generation stages, respectively, present distinct data flow and computation requirements. This paper proposes an architecture to efficiently support both stages, maximizing the parallelism and hardware utilization. We re-purpose the widely deployed 2D systolic array to inherit its efficiency in processing matrix multiplications and to maintain the compatibility with other models with a minor hardware addition (4.9%/3.9% overhead for area/energy) for the reconfigurability between two modes. The design also incorporates token pruning and bit precision reconfigurability without altering the 2D processing array. We also introduce a tailored data mapping for the attention, dubbed score stationary, which leverages the unique sparsity pattern from token pruning to further reduce power consumption. The proposed architecture achieves 27.4X energy savings and 10.7X performance benefits for decoder processing, while obtaining 2.65X energy reduction for the encoder at iso-throughput, presenting a promising unified solution for these two distinct key tasks. Chang Eun Song, Ashkan Moradifirouzabadi, Tajana Rosing, Mingu Kang |
ISLPED | 4 |
| 2023 | Benchmarking Self-Supervised Learning on Diverse Pathology DatasetsabstractComputational pathology can lead to saving human lives, but models are annotation hungry and pathology images are notoriously expensive to annotate. Self-supervised learning (SSL) has shown to be an effective method for utilizing unlabeled data, and its application to pathology could greatly benefit its downstream tasks. Yet, there are no principled studies that compare SSL methods and discuss how to adapt them for pathology. To address this need, we execute the largest-scale study of SSL pre-training on pathology image data, to date. Our study is conducted using 4 representative SSL methods on diverse downstream tasks. We establish that large-scale domain-aligned pre-training in pathology consistently out-performs ImageNet pre-training in standard SSL settings such as linear and fine-tuning evaluations, as well as in low-label regimes. Moreover, we propose a set of domain-specific techniques that we experimentally show leads to a performance boost. Lastly, for the first time, we apply SSL to the challenging task of nuclei instance segmentation and show large and consistent performance improvements. We release the pre-trained model weights11https://lunit-io.github.io/research/publications/pathology_ssl. Mingu Kang, Heon Song, Seonwook Park, Donggeun Yoo, Sérgio Pereira |
CVPR | 1 |
| 2023 | High-Speed Wafer Temperature Control Approach of Step Chiller for Semiconductor Manufacturing EquipmentabstractAs the complexity of semiconductor manufacturing processes increases, various temperature control conditions are required to ensure the etching of various materials is optimized. Fast and precise temperature control is a key factor in the semiconductor production throughput and yield, making the development of a step chiller essential. This chiller mixes hot and cold coolants from two sources, enabling faster cooling/heating than can be achieved from a single source. In this paper, we propose a system identification and control strategy that can maximize the performance of step chillers for semiconductor production equipment, focusing on both speed and accuracy aspects. An approach of model reduction is utilized to transform the highly nonlinear time-varying delay system of the chiller into a linear time-invariant delay system. We then outline a model-based feedforward/feedback controller based on delay system control theory, and conclude with an experimental result demonstrating the effectiveness of the proposed control algorithm. Hyeonjun Yun, Hyeseon Kwon, Sangsu Yeh, Yunha Kim, Seungwoo Cha, Mingu Kang, Jooyeop Nam |
IECON | 8 |
| 2023 | Which Exceptions Do We Have to Catch in the Python Code for AI Projects?abstractRecently, Python is the most-widely used language in artificial intelligence (AI) projects requiring huge amount of CPU and memory resources, and long execution time for training. For saving the project duration and making AI software systems more reliable, it is inevitable to handle exceptions appropriately at the code level. However, handling exceptions highly relies on developer’s experience. This is because, as an interpreter-based programming language, it does not force a developer to catch exceptions during development. In order to resolve this issue, we propose an approach to suggesting appropriate exceptions for the AI code segments during development after training exceptions from the existing handling statements in the AI projects. This approach learns the appropriate token units for the exception code and pretrains the embedding model to capture the semantic features of the code. Additionally, the attention mechanism learns to catch the salient features of the exception code. For evaluating our approach, we collected 32,771 AI projects using two popular AI frameworks (i.e. Pytorch and Tensorflow) and we obtained the 0.94 of Area under the Precision-Recall Curve (AUPRC) on average. Experimental results show that the proposed method can support the developer’s exception handling with better exception proposal performance than the compared models. Mingu Kang, Suntae Kim, Duksan Ryu, Jaehyuk Cho |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2022 | Accelerating attention through gradient-based learned runtime pruningabstractSelf-attention is a key enabler of state-of-art accuracy for various transformer-based Natural Language Processing models. This attention mechanism calculates a correlation score for each word with respect to the other words in a sentence. Commonly, only a small subset of words highly correlates with the word under attention, which is only determined at runtime. As such, a significant amount of computation is inconsequential due to low attention scores and can potentially be pruned. The main challenge is finding the threshold for the scores below which subsequent computation will be inconsequential. Although such a threshold is discrete, this paper formulates its search through a soft differentiable regularizer integrated into the loss function of the training. This formulation piggy backs on the back-propagation training to analytically co-optimize the threshold and the weights simultaneously, striking a formally optimal balance between accuracy and computation pruning. To best utilize this mathematical innovation, we devise a bit-serial architecture, dubbed LeOPArd, for transformer language models with bit-level early termination microarchitectural mechanism. We evaluate our design across 43 back-end tasks for MemN2N, BERT, ALBERT, GPT-2, and Vision transformer models. Post-layout results show that, on average, LeOPArd yields 1.9×and 3.9×speedup and energy reduction, respectively, while keeping the average accuracy virtually intact (< 0.2% degradation). Soroush Ghodrati, Amir Yazdanbakhsh, Hadi Esmaeilzadeh, Mingu Kang |
ISCA | 5 |
| 2022 | Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip RecomputationabstractAs its core computation, a self-attention mechanism gauges pairwise correlations across the entire input sequence. Despite favorable performance, calculating pairwise correlations is prohibitively costly. While recent work has shown the benefits of runtime pruning of elements with low attention scores, the quadratic complexity of self-attention mechanisms and their on-chip memory capacity demands are overlooked. This work addresses these constraints by architecting an accelerator, called SPRINT1, which leverages the inherent parallelism of ReRAM crossbar arrays to compute attention scores in an approximate manner. Our design prunes the low attention scores using a lightweight analog thresholding circuitry within ReRAM, enabling SPRINT to fetch only a small subset of relevant data to on-chip memory. To mitigate potential negative repercussions for model accuracy, SPRINT re-computes the attention scores for the few fetched data in digital. The combined in-memory pruning and on-chip recompute of the relevant attention scores enables SPRINT to transform quadratic complexity to a merely linear one. In addition, we identify and leverage a dynamic spatial locality between the adjacent attention operations even after pruning, which eliminates costly yet redundant data fetches. We evaluate our proposed technique on a wide range of state-of-the-art transformer models. On average, SPRINT yields 7.5× speedup and 19.6× energy reduction when total l6KB on-chip memory is used, while virtually on par with iso-accuracy of the baseline models (on average 0.36% degradation). Amir Yazdanbakhsh, Ashkan Moradifirouzabadi, Mingu Kang |
MICRO | 4 |
| 2022 | Gradle-Autofix: An Automatic Resolution Generator for Gradle Build ErrorabstractGradle is one of the widely used tools to automatically build a software project. While developers execute the Gradle build for projects, they face various build errors in practice. However, fixing build errors is not easy because developers should manually find out the cause of the build error and its resolution on their project. For this reason, developers spend much time fixing them, and especially it can be worse if a developer lacks the experience of handling build errors. To address this issue, we propose a novel approach named Gradle-AutoFix to automatically fix build errors along with providing their causes and resolutions. In this approach, we collect build errors to group their causes and resolutions and then generate feature vectors from build error messages by applying Bag-of-Word (BoW), Term Frequency-Inverse Document Frequency (TF-IDF), Bigram, and an embedding layer. The feature vectors are utilized for training two classification models on cause and resolution. Next, we analyze fixing patterns and define seven resolution rules to fix the build error automatically. Based on our trained models and defined resolution rules, we built Gradle-AutoFix. For the evaluation, we measured how appropriately Gradle-AutoFix provides causes of build errors and resolutions. As a result, we obtained 96% and 91% accuracy, respectively. Also, we assessed how properly Gradle-AutoFix fixes the project’s build error based on the seven resolution rules. The outcome showed a 64.5% build error resolution rate for 231 projects. Mingu Kang, Suntae Kim, Duksan Ryu |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2021 | CAP-GAN: Towards Adversarial Robustness with Cycle-consistent Attentional PurificationabstractAdversarial attack is aimed at fooling a target classifier with imperceptible perturbation. Adversarial examples, which are carefully crafted with a malicious purpose, can lead to erroneous predictions, resulting in catastrophic accidents. To mitigate the effect of adversarial attacks, we propose a novel purification model called CAP-GAN. CAP-GAN considers the idea of pixel-level and feature-level consistency to achieve reasonable purification under cycle-consistent learning. Specifically, we utilize a guided attention module and knowledge distillation to convey meaningful information to the purification model. Once the model is fully trained, inputs are projected into the purification model and transformed into clean-like images. We vary the capacity of the adversary to argue the robustness against various types of attack strategies. On CIFAR-10 dataset, CAP-GAN outperforms other pre-processing based defenses under both black-box and white-box settings. Mingu Kang, Trung Q. Tran, Seung Ju Cho, Daeyoung Kim 0001 |
IJCNN | 1 |
| 2021 | ReRankMatch: Semi-Supervised Learning with Semantics-Oriented Similarity RepresentationabstractThis paper proposes integrating semantics-oriented similarity representation into RankingMatch, a recently proposed semi-supervised learning method. Our method, dubbed ReRankMatch, aims to deal with the case in which labeled and unlabeled data share non-overlapping categories. ReRankMatch encourages the model to produce the similar image representations for the samples likely belonging to the same class. We evaluate our method on various datasets such as CIFAR-10, CIFAR-100, SVHN, STL-10, and Tiny ImageNet. We obtain promising results (4.21% error rate on CIFAR-10 with 4000 labels, 22.32% error rate on CIFAR-100 with 10000 labels, and 2.19% error rate on SVHN with 1000 labels) when the amount of labeled data is sufficient to learn semantics-oriented similarity representation. The code is made publicly available at https://github.com/tqtrunghnvn/ReRankMatch. Trung Q. Tran, Mingu Kang, Daeyoung Kim 0001 |
IJCNN | 2 |
| 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and InferenceabstractThe growing prevalence and computational demands of Artificial Intelligence (AI) workloads has led to widespread use of hardware accelerators in their execution. Scaling the performance of AI accelerators across generations is pivotal to their success in commercial deployments. The intrinsic error-resilient nature of AI workloads present a unique opportunity for performance/energy improvement through precision scaling. Motivated by the recent algorithmic advances in precision scaling for inference and training, we designed RaPiD1, a 4-core AI accelerator chip supporting a spectrum of precisions, namely, 16 and 8-bit floating-point and 4 and 2-bit fixed-point. The 36mm2RaPiD chip fabricated in 7nm EUV technology delivers a peak 3.5 TFLOPS/W in HFP8 mode and 16.5 TOPS/W in INT4 mode at nominal voltage. Using a performance model calibrated to within 1% of the measurement results, we evaluated DNN inference using 4-bit fixed-point representation for a 4-core 1 RaPiD chip system and DNN training using 8-bit floating point representation for a 768 TFLOPs AI system comprising 4 32-core RaPiD chips. Our results show INT4 inference for batch size of 1 achieves 3 - 13.5 (average 7) TOPS/W and FP8 training for a mini-batch of 512 achieves a sustained 102 - 588 (average 203) TFLOPS across a wide range of applications. Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang 0333, Sanchari Sen, Ankur Agrawal, Monodeep Kar, Shubham Jain 0004, Alberto Mannari, Hoang Tran, Eri Ogawa, Kazuaki Ishizaki, Hiroshi Inoue, Marcel Schaal, Mauricio J. Serrano, Jungwook Choi, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Allison Allain, James Bonanno, Nianzheng Cao, Robert Casatuta, Matthew Cohen, Bruce M. Fleischer, Michael Guillorn, Howard Haynie, Jinwook Jung, Mingu Kang, Kyu-Hyoun Kim, Siyu Koswatta, Sae Kyu Lee, Martin Lutz, Silvia M. Müller, Jinwook Oh, Ashish Ranjan 0001, Zhibin Ren, Scot Rider, Kerstin Schelm, Michael Scheuermann, Joel Silberman, Vidhi Zalani, Xin Zhang 0025, Ching Zhou, Matthew M. Ziegler, Vinay Shah, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Leland Chang, Kailash Gopalakrishnan |
ISCA | 30 |
| 2020 | Deep In-Memory Architectures in SRAM: An Analog Approach to Approximate ComputingabstractThis article provides an overview of recently proposed deep in-memory architectures (DIMAs) in SRAM for energyand latency-efficient hardware realization of machine learning (ML) algorithms. DIMA tackles the data movement problem in von Neumann architectures head-on by deeply embedding mixed-signal computations into a conventional memory array. In doing so, it trades off its computational signal-to-noise ratio (compute SNR) with energy and latency, and therefore, it represents an analog form of approximate computing. DIMA exploits the inherent error immunity of ML algorithms and SNR budgeting methods to operate its analog circuitry in a low-swing/low-compute SNR regime, thereby achieving >100× reduction in the energy-delay product (EDP) over an equivalent von Neumann architecture with no loss in inference accuracy. This article describes DIMA's computational pipeline and provides a Shannon-inspired rationale for its robustness to process, temperature, and voltage variations and design guidelines to manage its analog nonidealities. DIMA's versatility, effectiveness, and practicality demonstrated via multiple silicon IC prototypes in a 65-nm CMOS process are described. A DIMA-based instruction set architecture (ISA) to realize an end-to-end application-toarchitecture mapping for the accelerating diverse ML algorithms is also presented. Finally, DIMA's fundamental tradeoff between energy and accuracy in the low-compute SNR regime is analyzed to determine energy-optimum design parameters. Mingu Kang, Sujan K. Gonugondla, Naresh R. Shanbhag |
Proc. IEEE | 1 |
| 2020 | Efficient AI System Design With Cross-Layer Approximate ComputingabstractAdvances in deep neural networks (DNNs) and the availability of massive real-world data have enabled superhuman levels of accuracy on many AI tasks and ushered the explosive growth of AI workloads across the spectrum of computing devices. However, their superior accuracy comes at a high computational cost, which necessitates approaches beyond traditional computing paradigms to improve their operational efficiency. Leveraging the application-level insight of error resilience, we demonstrate how approximate computing (AxC) can significantly boost the efficiency of AI platforms and play a pivotal role in the broader adoption of AI-based applications and services. To this end, we present RaPiD, a multi-tera operations per second (TOPS) AI hardware accelerator core (fabricated at 14-nm technology) that we built from the ground-up using AxC techniques across the stack including algorithms, architecture, programmability, and hardware. We highlight the workload-guided systematic explorations of AxC techniques for AI, including custom number representations, quantization/pruning methodologies, mixed-precision architecture design, instruction sets, and compiler technologies with quality programmability, employed in the RaPiD accelerator. Swagath Venkataramani, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Jungwook Choi, Mingu Kang, Ankur Agarwal, Jinwook Oh, Shubham Jain 0004, Tina Babinsky, Nianzheng Cao, Thomas W. Fox, Bruce M. Fleischer, George Gristede, Michael Guillorn, Howard Haynie, Hiroshi Inoue, Kazuaki Ishizaki, Michael J. Klaiber, Shih-Hsien Lo, Gary W. Maier, Silvia M. Müller, Michael Scheuermann, Eri Ogawa, Marcel Schaal, Mauricio J. Serrano, Joel Silberman, Christos Vezyrtzis, Wei Wang 0333, Fanchieh Yee, Matthew M. Ziegler, Ching Zhou, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Vijayalakshmi Srinivasan, Leland Chang, Kailash Gopalakrishnan |
Proc. IEEE | 6 |
| 2019 | An MRAM-Based Deep In-Memory Architecture for Deep Neural NetworksabstractThis paper presents an MRAM-based deep in-memory architecture (MRAM-DIMA) to efficiently implement multi-bit matrix vector multiplication for deep neural networks using a standard MRAM bitcell array. The MRAM-DIMA achieves an 4.5 × and 70× lower energy and delay, respectively, compared to a conventional digital MRAM architecture. Behavioral models are developed to estimate the impact of circuit non-idealities, including process variations, on the DNN accuracy. An accuracy drop of ≤ 0.5% (≤ 1%) is observed for LeNet-300-100 on the MNIST dataset (a 9-layer CNN on the CIFAR-10 dataset), while tolerating 24% (12%) variation in cell conductance in a commercial 22 nm CMOS-MRAM process. Ameya Patil 0001, Haocheng Hua, Sujan K. Gonugondla, Mingu Kang, Naresh R. Shanbhag |
ISCAS | 4 |
| 2018 | PROMISE: An End-to-End Design of a Programmable Mixed-Signal Accelerator for Machine-Learning AlgorithmsabstractAnalog/mixed-signal machine learning (ML) accelerators exploit the unique computing capability of analog/mixed-signal circuits and inherent error tolerance of ML algorithms to obtain higher energy efficiencies than digital ML accelerators. Unfortunately, these analog/mixed-signal ML accelerators lack programmability, and even instruction set interfaces, to support diverse ML algorithms or to enable essential software control over the energy-vs-accuracy tradeoffs. We propose PROMISE, the first end-to-end design of a PROgrammable MIxed-Signal accElerator from Instruction Set Architecture (ISA) to high-level language compiler for acceleration of diverse ML algorithms. We first identify prevalent operations in widely-used ML algorithms and key constraints in supporting these operations for a programmable mixed-signal accelerator. Second, based on that analysis, we propose an ISA with a PROMISE architecture built with silicon-validated components for mixed-signal operations. Third, we develop a compiler that can take a ML algorithm described in a high-level programming language (Julia) and generate PROMISE code, with an IR design that is both language-neutral and abstracts away unnecessary hardware details. Fourth, we show how the compiler can map an application-level error tolerance specification for neural network applications down to low-level hardware parameters (swing voltages for each application Task) to minimize energy consumption. Our experiments show that PROMISE can accelerate diverse ML algorithms with energy efficiency competitive even with fixed-function digital ASICs for specific ML algorithms, and the compiler optimization achieves significant additional energy savings even for only 1% extra errors. Prakalp Srivastava, Mingu Kang, Sujan K. Gonugondla, Sungmin Lim, Jungwook Choi, Vikram S. Adve, Nam Sung Kim, Naresh R. Shanbhag |
ISCA | 2 |
| 2018 | Energy-Efficient Deep In-memory Architecture for NAND Flash MemoriesabstractThis paper proposes an energy-efficient deep in-memory architecture for NAND flash (DIMA-F) to perform machine learning and inference algorithms on NAND flash memory. Algorithms for data analytics, inference, and decision-making require processing of large data volumes and are hence limited by data access costs. DIMA-F achieves energy savings and throughput improvement for such algorithms by reading and processing data in the analog domain at the periphery of NAND flash memory. This paper also provides behavioral models of DIMA-F that can be used for analysis and large scale system simulations in presence of circuit non-idealities and variations. DIMA-F is studied in the context of linear support vector machines and k-nearest neighbor for face detection and recognition, respectively. An estimated 8×-to-23× reduction in energy and 9×-to-15× improvement in throughput resulting in EDP gains up to 345× over the conventional NAND flash architecture incorporating an external digital ASIC for computation. Sujan K. Gonugondla, Mingu Kang, Yongjune Kim 0001, Mark Helm, Sean Eilert, Naresh R. Shanbhag |
ISCAS | 2 |
| 2018 | SRAM Bit-line Swings Optimization using Generalized WaterfillingabstractWe propose an information-theoretic approach to optimize non-uniform bit-line swings for static random access memories (SRAMs). We formulate convex optimization problems whose objectives are to minimize energy (for low-power SRAMs), maximize speed (for high-speed SRAMs), and minimize energy-delay product for a given constraint on mean squared error of retrieved words. We show that these optimization problems can be interpreted as generalized water-filling including classical waterfilling, ground-flattening and water-filling, and sand-pouring and water-filling, respectively. Numerical results show that energy-optimal swing assignment reduces energy consumption by half at a peak signal-to-noise ratio of 30dB for an 8-bit accessed word. Yongjune Kim 0001, Mingu Kang, Lav R. Varshney, Naresh R. Shanbhag |
ISIT | 2 |
| 2018 | Generalized Water-Filling for Source-Aware Energy-Efficient SRAMsabstractConventional low-power static random access memories (SRAMs) reduce read energy by decreasing the bit-line voltage swings uniformly across the bit-line columns. This is because the read energy is proportional to the bit-line swings. On the other hand, bit-line swings are limited by the need to avoid decision errors especially in the most significant bits. We propose a principled approach to determine optimal non-uniform bit-line swings by formulating convex optimization problems. For a given constraint on mean squared error of retrieved words, we consider criteria to minimize energy (for low-power SRAMs), maximize speed (for high-speed SRAMs), and minimize energy-delay product. These optimization problems can be interpreted as classical water-filling, ground-flattening and water-filling, and sand-pouring and water-filling, respectively. By leveraging these interpretations, we also propose greedy algorithms to obtain optimized discrete swings. Numerical results show that energy-optimal swing assignment reduces energy consumption by half at a peak signal-to-noise ratio of 30 dB for an 8-bit accessed word. The energy savings increase to four times for a 16-bit accessed word. Yongjune Kim 0001, Mingu Kang, Lav R. Varshney, Naresh R. Shanbhag |
IEEE Trans. Commun. | 2 |
| 2015 | An energy-efficient memory-based high-throughput VLSI architecture for convolutional networksabstractIn this paper, an energy efficient, memory-intensive, and high throughput VLSI architecture is proposed for convolutional networks (C-Net) by employing compute memory (CM) [1], where computation is deeply embedded into the memory (SRAM). Behavioral models incorporating CM's circuit non-idealities and energy models in 45nm SOI CMOS are presented. System-level simulations using these models demonstrate that the probability of handwritten digit recognition Pr> 0.99 can be achieved using the MNIST database [2], along with a 24.5× reduced energy delay product, a 5.0× reduced energy, and a 4.9× higher throughput as compared to the conventional system. Mingu Kang, Sujan K. Gonugondla, Min-Sun Keel, Naresh R. Shanbhag |
ICASSP | 1 |
| 2015 | Energy-efficient and high throughput sparse distributed memory architectureabstractThis paper presents an energy-efficient VLSI implementation of Sparse Distributed Memory (SDM). High throughput and energy-efficient Hamming distance-based address decoder (CM-DEC) is proposed by employing compute memory [1], where computation is deeply embedded into a memory (SRAM). Hierarchical binary decision (HBD) is also proposed to enhance area- and energy-efficiency of read operation by minimizing data transfer. The SDM is employed as an auto-associative memory with four read iterations and 16×16 binary noisy input image with input error rates of 15%, 25%, and 30%. The proposed SDM achieves 39× smaller energy delay product with 14.5× and 2.7× reduced delay and energy, respectively as compared to conventional digital implementation of SDM in 45 nm SOI CMOS process with output error rate degradation less than 0.4%. Mingu Kang, Eric P. Kim, Min-Sun Keel, Naresh R. Shanbhag |
ISCAS | 1 |