EDBT 2026 Demo / reviewers in the wild / expert
Dong Eun Kim
dblp:260/6903
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-7251-1289ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TAXI: Traveling Salesman Problem Accelerator with X-bar-based Ising Macros Powered by SOT-MRAMs and Hierarchical ClusteringabstractIsing solvers with hierarchical clustering have shown promise for large-scale Traveling Salesman Problems (TSPs), in terms of latency and energy. However, most of these methods still face unacceptable quality degradation as the problem size increases beyond a certain extent. Additionally, their hardwareagnostic adoptions limit their ability to fully exploit available hardware resources. In this work, we introduce TAXI – an inmemory computing-based TSP accelerator with crossbar(Xbar)-based Ising macros. Each macro independently solves a TSP subproblem, obtained by hierarchical clustering, without the need for any off-macro data movement, leading to massive parallelism. Within the macro, Spin-Orbit-Torque (SOT) devices serve as compact energy-efficient random number generators enabling rapid “natural annealing”. By leveraging hardware-algorithm co-design, TAXI offers improvements in solution quality, speed, and energy-efficiency on TSPs up to $\mathbf{8 5, 9 0 0}$ cities (the largest TSPLIB instance). TAXI produces solutions that are only $22 \%$ and $20 \%$ longer than the Concorde solver’s exact solution on $\mathbf{3 3, 8 1 0}$ and $\mathbf{8 5, 9 0 0}$ city TSPs, respectively. TAXI outperforms a current state-of-the-art clustering-based Ising solver, being $8 \times$ faster on average across 20 benchmark problems from TSPLib. Sangmin Yoo, Amod Holla, Sourav Sanyal, Dong Eun Kim, Francesca Iacopi, Dwaipayan Biswas, James Myers, Kaushik Roy 0001 |
DAC | 4 |
| 2025 | MemRaptor: Magnetoresistive Array as Matrix Vector Multiplication and Transcendental Function Operator for NLP ApplicationsabstractCompute-In-Memory (CiM) is emerging as a promising paradigm to design energy-efficient hardware accelerators for AI, addressing the processor-memory data transfer bottleneck. The popularity of CiM can be attributed to their ability to perform massively parallel in-situ matrix vector multiplications (MVMs), the dominant computation in neural networks (NNs). However, NNs used in NLP applications such as Long-Short Term Memory (LSTM) and transformers also frequently perform other operations such as transcendental functions (tanh, sigmoid, and softmax). To that effect, we present MemRaptor, utilizing CiM with magnetoresistive random access memory (MRAM) technology, that can perform both MVM and transcendental functions in the same memory array. MemRaptor overlays a read only memory (ROM) on an MRAM array through hard-wiring the connection of bit-cell with an additional bitline (a bit-cell connected to either of the bitlines but not to both), incurring no array area overhead and a minimal peripheral area overhead. Note, the bitline connection of bit-cell stores the ROM value while the magnetic tunnel junction (MTJ) in the bit-cell stores the RAM data. Particularly, the magnetization state of the 1T-1MTJ bit-cell in the array stores the weight value (RAM data) of the neural network, and the bitline connection of the bit-cell stores the look-up table (ROM data) used for computing transcendental functions. We demonstrate the working of our proposed design through circuit-level simulations for a 64×64 array, using a compact model of CoFeB/MgO PMA MTJ with 120% tunnelling magnetoresistance, 5kΩ RON, in 65nm technology. Further, we showcase the advantage of MemRaptor over standard MVM-based CiM accelerator architecture, PUMA, through comprehensive system-level evaluations for LSTM, BERT, and GPT models. Our results show up to 30% and 5.3% improvements in terms of throughput and energy-efficiency, respectively, on an average across different workloads. Dong Eun Kim, Tanvi Sharma, Anushka Mukherjee, Mainakh Mukherjee, Kaushik Roy 0001 |
ISLPED | 1 |
| 2025 | WAGONN: Weight Bit Agglomeration in Crossbar Arrays for Reduced Impact of Interconnect Resistance on DNN Inference AccuracyabstractDeep neural network (DNN) accelerators employing crossbar arrays capable of in-memory computing (IMC) are highly promising for neural computing platforms. However, in deeply scaled technologies, interconnect resistance severely impairs IMC robustness, leading to a drop in the system accuracy. To address this problem, we propose WAGONN - a technique based on agglomerating weight bits in crossbar arrays which alleviates the detrimental effect of wire resistance. For 8T-SRAM-based$128\times 128$crossbar arrays in 7nm technology, WAGONN enhances the accuracy from 47.78% to 83.5% for ResNet-20/CIFAR-10. We also show that WAGONN can be used synergistically with Partial-Word-Line-Activation, further boosting the accuracy. Further, we evaluate the implications of WAGONN for compact ferroelectric transistor-based crossbar arrays and show accuracy enhancement. WAGONN incurs minimal hardware overhead, with less than a 1% increase in energy consumption. Additionally, the latency and area overheads of WAGONN are ~1% and ~16%, respectively when 1 ADC is utilized per crossbar array. Jeffry Victor, Dong Eun Kim, Kaushik Roy 0001, Sumeet Kumar Gupta |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | SAMBA: Sparsity Aware In-Memory Computing Based Machine Learning AcceleratorabstractMachine Learning (ML) inference is typically dominated by highly data-intensive Matrix Vector Multiplication (MVM) computations that may be constrained by memory bottleneck due to massive data movement between processor and memory. Although analog in-memory computing (IMC) ML accelerators have been proposed to execute MVM with high efficiency, the latency and energy of such computing systems can be dominated by the large latency and energy costs from analog-to-digital converters (ADCs). Leveraging sparsity in ML workloads, reconfigurable ADCs can save MVM energy and latency by reducing the required ADC bit precision. However, such improvement in latency can be hindered by non-uniform sparsity of the weight matrices mapped into hardware. Moreover, data movement between MVM processing cores may become another factor that delays the overall system-level performance. To address these issues, we propose SAMBA, Sparsity Aware IMC Based Machine Learning Accelerator. First, we propose load balancing during mapping of weight matrices into physical crossbars to eliminate non-uniformity in the sparsity of mapped matrices. Second, we propose optimizations in arranging and scheduling the tiled MVM hardware to minimize the overhead of data movement across multiple processing cores. Our evaluations show that the proposed load balancing technique can achieve performance improvement. The proposed optimizations can further improve both performance and energy-efficiency regardless of sparsity condition. With the combination of load balancing and data movement optimization in conjunction with reconfigurable ADCs, our proposed approach provides up to 2.38x speed-up and 1.54x energy-efficiency over stateof- art analog IMC based ML accelerators for ImageNet datasets on Resnet-50 architecture. Dong Eun Kim, Aayush Ankit, Cheng Wang 0036, Kaushik Roy 0001 |
IEEE Trans. Computers | 1 |
| 2021 | Mixed Precision Quantization for ReRAM-based DNN Inference AcceleratorsabstractReRAM-based accelerators have shown great potential for accelerating DNN inference because ReRAM crossbars can perform analog matrix-vector multiplication operations with low latency and energy consumption. However, these crossbars require the use of ADCs which constitute a significant fraction of the cost of MVM operations. The overhead of ADCs can be mitigated via partial sum quantization. However, prior quantization flows for DNN inference accelerators do not consider partial sum quantization which is not highly relevant to traditional digital architectures. To address this issue, we propose a mixed precision quantization scheme for ReRAM-based DNN inference accelerators where weight quantization, input quantization, and partial sum quantization are jointly applied for each DNN layer. We also propose an automated quantization flow powered by deep reinforcement learning to search for the best quantization configuration in the large design space. Our evaluation shows that the proposed mixed precision quantization scheme and quantization flow reduce inference latency and energy consumption by up to 3.89x and 4.84x, respectively, while only losing 1.18% in DNN inference accuracy. Sitao Huang, Aayush Ankit, Plínio Silveira, Rodrigo Antunes, Sai Rahul Chalamalasetti, Izzat El Hajj, Dong Eun Kim, Glaucimar Aguiar, Pedro Bruel, Sergey Serebryakov, Can Li 0024, Paolo Faraboschi, John Paul Strachan, Deming Chen, Kaushik Roy 0001, Wen-Mei W. Hwu, Dejan S. Milojicic |
ASP-DAC | 7 |
| 2020 | GENIEx: A Generalized Approach to Emulating Non-Ideality in Memristive Xbars using Neural NetworksabstractMemristive crossbars have been extensively explored for deep learning accelerators due to their high on-chip storage density and efficient Matrix Vector Multiplication (MVM) compared to digital CMOS. However, their analog nature of computing poses significant issues due to various non-idealities such as: parasitic resistances, non-linear I-V characteristics of the memristor device etc. The non-idealities can have a detrimental impact on the functionality i.e. computational accuracy of crossbars. Past works have explored modeling the non-idealities using analytical techniques. However, several non-idealities have data dependent behavior. This can not be captured using analytical (non data-dependent) models thereby, limiting their suitability in predicting application accuracy. To address this, we propose a Generalized Approach to Emulating Non-Ideality in Memristive Crossbars using Neural Networks (GENIEx), which accurately captures the data-dependent nature of non-idealities. First, we perform extensive HSPICE simulations of crossbars with different voltage and conductance combinations. Based on the obtained data, we train a neural network to learn the transfer characteristics of the non-ideal crossbar. Next, we build a functional simulator which includes key architectural facets such as tiling, and bit-slicing to analyze the impact of non-idealities on the classification accuracy of large-scale neural networks. We show that GENIEx achieves low root mean square errors (RMSE) of 0.25 and 0.7 for low and high voltages, respectively, compared to HSPICE. Additionally, the GENIEx errors are 7× and 12.8× better than an analytical model which can only capture the linear non-idealities. Further, using the functional simulator and GENIEx, we demonstrate that an analytical model can overestimate the degradation in classification accuracy by ≥ 10% on CIFAR-100 and 3.7% on ImageNet datasets compared to GENIEx. Indranil Chakraborty, Mustafa Fayez Ali, Dong Eun Kim, Aayush Ankit, Kaushik Roy 0001 |
DAC | 3 |