VLDB 2026 Research / reviewers in the wild / expert
Gordon Euhyun Moon
dblp:220/9907
· DBLP profile ↗
11ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0003-4992-6181ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Block-structured matrix reordering for efficient SDDMM on tensor coresabstractSampled Dense-Dense Matrix Multiplication is a fundamental operation in sparse linear algebra, widely used in graph neural networks and scientific computing. However, accelerating computations on GPUs is challenging due to data sparsity and irregular memory access, which hinder efficient use of Tensor Cores. This paper introduces a Block-Structured Matrix Reordering framework that improves Tensor Core utilization by reorganizing sparse matrices using bi-directional reordering with weighted similarity metrics. We also propose a tile-aware sparse matrix format that improves memory access and task scheduling. To enable adaptive and balanced computation, we employ a dual-path execution strategy: dense matrix blocks are assigned to Tensor Cores, while sparse blocks are handled by CUDA Cores. Experiments on the RTX 4090 demonstrate that our method achieves up to a $$10.38\times $$ speedup over the best Tensor Core baseline and $$7.31\times $$ over the best CUDA Core baseline by producing denser block structures and enhancing parallelism. Chengxing Zou, Changwan Hong, Gordon Euhyun Moon |
J. Supercomput. | 3 |
| 2025 | Equilibria: Co-Optimizing Energy and Latency in Online ML-Based Stream Processing SystemsabstractAn Online machine learning (ML)-based streaming processing system (SPS) combines real-time stream processing with continuous, incremental learning through simultaneous model training and inference. This system processes large, dynamic, high-velocity data streams while adapting its models to improve performance over time. However, balancing the tradeoff between latency and energy efficiency remains a critical challenge, which has not been adequately addressed in prior research. This paper introduces EQUILIBRIA, a novel framework designed to co-optimize power consumption and latency in Online ML-based SPS. EQUILIBRIA integrates dynamic voltage and frequency scaling (DVFS) with two innovative energy optimization strategies. First, a Pareto-based clock frequency adjustment mechanism dynamically tunes both core and memory clock frequencies to reduce latency while minimizing energy consumption. Second, a two-tier threshold training management technique optimizes energy use by periodically pausing and resuming model training once accuracy requirements are met, all while preserving latency. Experimental evaluations across various queries and traffic scenarios demonstrate that EQUILIBRIA achieves up to 58% energy savings without compromising latency, making a significant step forwards in energy-efficient, highperformance streaming analytics for modern, rapidly evolving data environments. Sejeong Oh, Soyang Baek, Gordon Euhyun Moon, Sungyong Park |
CCGrid | 3 |
| 2025 | D-HAT: Dynamic Hypergraph Representation Learning with Attention-Based Multi-Level Hypergraph SamplingabstractHypergraph Neural Networks (HNNs) leverage higher-order interactions in graph-structured data to enable effective representation learning across a wide range of applications. Due to the issue of incorporating irrelevant relations in full hypergraphs, it is crucial to adopt a hypergraph sampling method that efficiently captures substructures while preserving representational quality. However, existing hypergraph sampling methods that target only nodes or hyperedges suffer from subgraph disconnection issues and neglect of node importance due to the randomness in sampling and use of static computational sub-hypergraphs. In this paper, we propose D-HAT, a hypergraph learning framework that dynamically constructs representative sub-hypergraphs through a novel attention-based multi-level hypergraph sampling strategy during the training of HNNs. To prioritize informative neighbors and enhance the representational quality of sub-hypergraphs during training, we develop a new attention-based HNNs incorporating attention-guided aggregation and dense skip connections. To the best of our knowledge, this paper is the first to quantitatively compare various hypergraph sampling methods for hypergraph representation learning. Experiments on real-world graph datasets demonstrate the effectiveness of D-HAT, which consistently achieves higher accuracy compared to existing hypergraph sampling methods. Ah-Hyun Lee, Gordon Euhyun Moon |
CIKM | 2 |
| 2025 | Efficient GNN-based social recommender systems through social graph refinement
Sangmin Ga, Paul Hyunbin Cho, Gordon Euhyun Moon, Sungwon Jung |
J. Supercomput. | 3 |
| 2024 | ML-Based Dynamic Operator-Level Query Mapping for Stream Processing Systems in Heterogeneous Computing EnvironmentsabstractMapping queries to optimal computing devices at the operator-level presents a significant challenge in stream processing systems (SPS) with heterogeneous computing resources. Inefficient query mapping can degrade the performance of the SPS. To address this issue, existing approaches employ static methods, such as mapping all queries to either CPUs or GPUs, or maintaining static mapping tables for queries or operators based on their predetermined device preferences. However, the static mapping scheme fails to provide an optimal solution, as the device preference for different query operators changes dynamically at runtime. In this paper, we propose DynO, a high performance SPS that dynamically maps queries to devices at the operator-level using a tree-based machine learning algorithm. To effectively determine an optimized device mapping plan for query operators, DynO employs a tree-based gradient boosting model to accurately predict the execution time for all potential mapping plan combinations. DynO also introduces a novel turn-based updating scheme to maximize performance in stream processing while training a tree-based gradient boosting model. Additionally, we devise an efficient device mapping scheme to expedite the process of determining the optimal device mapping plan by leveraging a direct acyclic graph (DAG) shortest path algorithm. DynO completely hides any overhead caused by the extra computation needed to find the optimal plan by utilizing prefetching and GPU idle periods. Experimental results using a variety of queries and traffic patterns show that DynO outperforms existing state-of-the-art approaches by ensuring high throughput, low latency, and high efficiency. Sejeong Oh, Gordon Euhyun Moon, Sungyong Park |
CLUSTER | 2 |
| 2024 | Accelerated Block-Sparsity-Aware Matrix Reordering for Leveraging Tensor Cores in Sparse Matrix-Multivector Multiplication
Yoonsang Han, Gordon Euhyun Moon |
Euro-Par (3) | 3 |
| 2024 | Layer-Wise Sparse Training of Transformer via Convolutional Flood Filling
Bokyeong Yoon, Yoonsang Han, Gordon Euhyun Moon |
PAKDD (2) | 3 |
| 2023 | Chronica: A Data-Imbalance-Aware Scheduler for Distributed Deep LearningabstractOne of the major challenges in distributed deep learning is attenuating straggler problem. The straggler increases synchronization latency and significantly inhibits the convergence of deep learning model. We empirically observe that the imbal-anced data samples worsen the straggler problem and make the convergence of the deep learning model slower. However, existing approaches such as BOA and EP4DDL have not addressed data imbalance issues while solving the straggler problem. To overcome the straggler and data imbalance problems, we propose Chronica,a new data-imbalance-aware scheduler. Based on the size of the data samples and the configuration of each worker, Chronicaelaborately predicts the training time required for each worker. Chronicathen provides equivalent training time to each of the workers, alleviating both step- and epoch-level straggler problems. Furthermore, Chronicasuggests a new parameter synchronization scheme to achieve fast convergence based on the weighted average of the training workload on each worker. Our extensive evaluation using four deep learning models on 32 Amazon EC2 GPU instances showed that the new Chronicaachieves up to 3.19 times speedup over the state-of-the-art systems. Sanha Maeng, Gordon Euhyun Moon, Sungyong Park |
CCGrid | 2 |
| 2022 | Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix MultiplicationabstractThere is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators. Gordon Euhyun Moon, Hyoukjun Kwon, Geonhwa Jeong, Prasanth Chatarasi, Sivasankaran Rajamanickam, Tushar Krishna |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Extending Sparse Tensor Accelerators to Support Multiple Compression FormatsabstractSparsity, which occurs in both scientific applications and Deep Learning (DL) models, has been a key target of optimization within recent ASIC accelerators due to the potential memory and compute savings. These applications use data stored in a variety of compression formats. We demonstrate that both the compactness of different compression formats and the compute efficiency of the algorithms enabled by them vary across tensor dimensions and amount of sparsity. Since DL and scientific workloads span across all sparsity regions, there can be numerous format combinations for optimizing memory and compute efficiency. Unfortunately, many proposed accelerators operate on one or two fixed format combinations. This work proposes hardware extensions to accelerators for supporting numerous format combinations seamlessly and demonstrates ~ 4 x speedup over performing format conversions in software. Eric Qin 0001, Geonhwa Jeong, William Won, Sheng-Chun Kao, Hyoukjun Kwon, Sudarshan Srinivasan, Dipankar Das 0002, Gordon Euhyun Moon, Sivasankaran Rajamanickam, Tushar Krishna |
IPDPS | 8 |
| 2020 | ALO-NMF: Accelerated Locality-Optimized Non-negative Matrix FactorizationabstractNon-negative Matrix Factorization (NMF) is a key kernel for unsupervised dimension reduction used in a wide range of applications, including graph mining, recommender systems and natural language processing. Due to the compute-intensive nature of applications that must perform repeated NMF, several parallel implementations have been developed. However, existing parallel NMF algorithms have not addressed data locality optimizations, which are critical for high performance since data movement costs greatly exceed the cost of arithmetic/logic operations on current computer systems. In this paper, we present a novel optimization method for parallel NMF algorithm based on the HALS (Hierarchical Alternating Least Squares) scheme that incorporates algorithmic transformations to enhance data locality. Efficient realizations of the algorithm on multi-core CPUs and GPUs are developed, demonstrating a new Accelerated Locality-Optimized NMF (ALO-NMF) that obtains up to 2.29x lower data movement cost and up to 4.45x speedup over existing state-of-the-art parallel NMF algorithms. Gordon Euhyun Moon, J. Austin Ellis, Aravind Sukumaran-Rajam, Srinivasan Parthasarathy 0001, P. Sadayappan |
KDD | 1 |