EDBT 2026 Demo / reviewers in the wild / expert
Seungeon Hwang 0001
dblp:340/1585-1 · also Seung-Eon Hwang 0001
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0004-1772-5638ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HINT: A Hybrid SRAM-MRAM Compute-In-Memory with INput-aware Skipping SAR-ADC for Energy Efficient Ternary LLMsabstractAlthough large language models (LLMs) show remarkable performance in natural language processing tasks, their deployment on resource-constrained devices remains challenging due to a substantial memory footprint and high-energy consumption. To address these challenges, low-bit and ternary quantization reduce the model size, while hardware approaches such as compute-in-memory (CIM) alleviate the overhead of external memory accesses. However, billions of parameters of LLMs still cause significant data movement, and existing ternary CIM suffers from a low-density bitcell as well as accuracy degradation due to cut-off analog-to-digital converters (ADCs). In this paper, we propose HINT, a CIM architecture incorporating two energy efficient techniques. First, hybrid ternary bitcell leverages the reliability of SRAM and the high-density of MRAM, reducing area and energy overhead. Second, input-aware skipping SAR-ADC exploits input sparsity to skip unnecessary conversion cycles without sacrificing accuracy. On BitNet b1.58 (700M), compared to SRAM-based and eDRAM-based CIM baselines, HINT improves bitcell density by 1.85× and achieves up to 2.67× higher energy efficiency, respectively. By skipping up to 21% of conversion cycles, the proposed ADC improves energy efficiency up to 1.27× while maintaining model accuracy. Jaebeom Park, Seungeon Hwang 0001, Jongsun Park 0001 |
DATE | 2 |
| 2025 | SiGNoR: Similarity-based Graph Partitioning and Node Reuse for Memory Efficient GNN AccelerationabstractGraph neural networks (GNNs) are increasingly used in edge and mobile applications, but the large memory requirements of GNN hardware limit their deployment on resource-constrained devices. A common solution is graph partitioning, which splits large graphs into smaller partitions that fit into on-chip memory. However, this creates halo nodes—duplicated nodes across partitions—that lead to frequent and redundant accesses to external DRAM. This paper presents SiGNoR, an efficient graph partitioning approach that introduces three hardware-friendly techniques to reduce memory access with negligible accuracy loss. First, 1) similarity-based node reuse reduces redundant DRAM access of halo nodes by leveraging only local nodes within the same partition. To further improve the impact of the halo-to-local node replacement on GNN accuracy, 2) multiphase graph partitioning is proposed to refine partition boundaries by iteratively identifying and modeling its effects. Finally, 3) partition-wise block floating-point is presented for an energy efficient GNN accelerator hardware. The proposed partition-wise block floating-point shares a single exponent across all node features within each partition, reducing memory usage. Experimental results show that SiGNoR reduces average DRAM traffic of 44.5% with negligible accuracy degradation for GNNs across six real-world graph datasets. The GNN accelerator also improves energy efficiency by up to 2.6× compared to previous GNN hardware. Seungeon Hwang 0001, Hyeon Gwon Kim, Dongwoo Lew, Jongsun Park 0001 |
ICCAD | 1 |
| 2024 | SpARC: Token Similarity-Aware Sparse Attention Transformer Accelerator via Row-wise ClusteringabstractSelf-attention mechanisms, the key enabler of transformers' remarkable performance, account for a significant portion of the overall transformer computation. Despite its effectiveness, self-attention inherently contains considerable redundancies, making sparse attention an attractive approach. In this paper, we propose SpARC, a sparse attention transformer accelerator that enhances throughput and energy efficiency by reducing the computational complexity of the self-attention mechanism. Our approach exploits inherent row-level redundancies in transformer attention maps to reduce the overall self-attention computation. By employing row-wise clustering, attention scores are calculated only once per cluster to achieve approximate attention without seriously compromising accuracy. To leverage the high parallelism of the proposed clustering approximate attention, we develop a fully pipelined accelerator with a dedicated memory hierarchy. Experimental results demonstrate that SpARC achieves attention map sparsity levels of 85-90% with negligible accuracy loss. SpARC achieves up to 4× core attention speedup and 6× energy efficiency improvement compared to prior sparse attention transformer accelerators. Han Cho, Seungeon Hwang 0001, Jongsun Park 0001 |
DAC | 3 |
| 2024 | HeNCoG: A Heterogeneous Near-memory Computing Architecture for Energy Efficient GCN AccelerationabstractGraph convolutional network (GCN), which first applies convolutional operations to process graph data, has gained attention in various tasks involving relational data. Previous GCN accelerators have been designed with heterogeneous cores, considering two stages of inference (aggregation and combination), or with a unified core based on the inference of multi-layer as an iterative sparse-dense matrix multiplication. However, those prior works have suffered from an unnecessary large number of multiply-accumulate (MAC) operations and/or main memory accesses. In this paper, we propose HeNCoG, a GCN accelerator that utilizes a heterogeneous MAC array core for the combination stage and a near-memory computing core for the aggregation stage. In HeNCoG, considering that the number of MAC operations is significantly reduced when changing the stage execution order, the combination stage is executed first with a row-stationary dataflow. In the aggregation stage, magneto-resistive random-access memory (MRAM)-based near-memory computing is employed to reduce the number of main memory accesses needed to access the adjacency matrix in the graph dataset. Graph partitioning and double buffering techniques are also applied to further improve hardware efficiencies. Simulation results show that the HeNCoG architecture reduces execution cycles by 97% and memory accesses by 42% compared to previous works. Seungeon Hwang 0001, Duyeong Song, Jongsun Park 0001 |
ISCAS | 1 |