EDBT 2026 Demo / reviewers in the wild / expert
Seungwon Min
dblp:229/8282
· DBLP profile ↗
7ranked-venue papers
4as first author
4since 2021 · last 2023
0000-0001-7195-7182ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Memory systems · 41% GPUs and heterogeneous computing · 39% Storage systems · 11% | |
| Artificial intelligence
2 papers |
Graph learning · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU storage access |
0.7 | 1 | 2023 | GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System Architecture · ASPLOS (2) 2023 |
Memory systems › cache management
software-managed cache |
0.7 | 1 | 2023 | GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System Architecture · ASPLOS (2) 2023 |
Machine learning › Graph learning
graph neural network |
0.6 | 1 | 2022 | Graph Neural Network Training and Data Tiering · KDD 2022 |
Machine learning › Graph learning
graph neural network training |
0.6 | 1 | 2022 | Graph Neural Network Training and Data Tiering · KDD 2022 |
Memory systems
memory hierarchy |
0.6 | 1 | 2022 | Graph Neural Network Training and Data Tiering · KDD 2022 |
Storage systems › storage hierarchy
tiered storage |
0.6 | 1 | 2022 | Graph Neural Network Training and Data Tiering · KDD 2022 |
GPUs and heterogeneous computing
GPU memory management |
0.5 | 1 | 2021 | Large Graph Convolutional Network Training with GPU-Oriented Data Communication Architecture · Proc. VLDB Endow. 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
graph neural network accelerator |
0.5 | 1 | 2021 | Large Graph Convolutional Network Training with GPU-Oriented Data Communication Architecture · Proc. VLDB Endow. 2021 |
GPUs and heterogeneous computing › GPU graph processing
GPU graph traversal |
0.4 | 1 | 2020 | EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal In GPUs · Proc. VLDB Endow. 2020 |
Memory systems › processing-in-memory
near-memory processing |
0.3 | 1 | 2018 | Application-Transparent Near-Memory Processing Architecture with Memory Channel Network · MICRO 2018 |
Memory systems
processing-in-memory |
0.3 | 1 | 2018 | Application-Transparent Near-Memory Processing Architecture with Memory Channel Network · MICRO 2018 |
Graph data management
graph analytics |
0.2 | 1 | 2023 | GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System Architecture · ASPLOS (2) 2023 |
GPUs and heterogeneous computing › GPU communication
CPU-GPU communication |
0.2 | 1 | 2022 | Graph Neural Network Training and Data Tiering · KDD 2022 |
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU training |
0.2 | 1 | 2022 | Graph Neural Network Training and Data Tiering · KDD 2022 |
Machine learning › Graph learning › graph neural network training
graph convolutional network training |
0.1 | 1 | 2021 | Large Graph Convolutional Network Training with GPU-Oriented Data Communication Architecture · Proc. VLDB Endow. 2021 |
GPUs and heterogeneous computing › GPU memory management
unified virtual memory |
0.1 | 1 | 2020 | EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal In GPUs · Proc. VLDB Endow. 2020 |
Methods — techniques the papers use, named apart from their topics
statistical analysis · 1.1zero-copy access · 1.0asynchronous kernel execution · 1.0address alignment · 1.0request coalescing · 0.4cache-line-sized access · 0.4simulation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System ArchitectureabstractGraphics Processing Units (GPUs) have traditionally relied on the host CPU to initiate access to the data storage. This approach is well-suited for GPU applications with known data access patterns that enable partitioning of their dataset to be processed in a pipelined fashion in the GPU. However, emerging applications such as graph and data analytics, recommender systems, or graph neural networks, require fine-grained, data-dependent access to storage. CPU initiation of storage access is unsuitable for these applications due to high CPU-GPU synchronization overheads, I/O traffic amplification, and long CPU processing latencies. GPU-initiated storage removes these overheads from the storage control path and, thus, can potentially support these applications at much higher speed. However, there is a lack of systems architecture and software stack that enable efficient GPU-initiated storage access. This work presents a novel system architecture, BaM, that fills this gap. BaM features a fine-grained software cache to coalesce data storage requests while minimizing I/O traffic amplification. This software cache communicates with the storage system via high-throughput queues that enable the massive number of concurrent threads in modern GPUs to make I/O requests at a high rate to fully utilize the storage devices and the system interconnect. Experimental results show that BaM delivers 1.0x and 1.49x end-to-end speed up for BFS and CC graph analytics benchmarks while reducing hardware costs by up to 21.7x over accessing the graph data from the host memory. Furthermore, BaM speeds up data-analytics workloads by 5.3x over CPU-initiated storage access on the same hardware. Zaid Qureshi, Vikram S. Mailthody, Isaac Gelado, Seungwon Min, Amna Masood, Jeongmin Brian Park, Jinjun Xiong, Chris J. Newburn, Dmitri Vainbrand, I-Hsin Chung, Michael Garland, William J. Dally, Wen-Mei W. Hwu |
ASPLOS (2) | 4 |
| 2023 | FSSD: FPGA-Based Emulator for SSDsabstractSolid State Drives (SSDs) have become increasingly popular due to their superior access latency and bandwidth compared to Hard Disk Drives (HDDs). However, to fully understand the impact of SSD design and microarchitecture on end-to-end application performance, researchers need to move beyond treating SSDs as black-box components. Unfortunately, purchasing multiple SSDs for research is expensive and ineffective since the underlying microarchitecture is still unknown to the system designer. While simulators have become the most popular method for studying SSDs, existing software-based simulators lack real data transfers and cannot simulate the latency from NVMe and PCIe interfaces. Additionally, simulating the entire SSD using software codes is time-consuming and limits the number of experiments that can be run in a reasonable amount of time. To address these issues, we present FSSD, an FPGA-based emulation system that models the latency and access patterns of an actual NVMe SSD. FSSD takes advantage of the flexibility of an FPGA, enabling users to customize SSD microarchitecture features and explore the design space for data-intensive applications. FSSD can be interacting with real operating systems, instead of relying on Virtual Machines like most other software simulators do. Evaluations show that FSSD provides over 1000x speedup compared to software-based simulation using the SimpleSSD simulator. The ability to customize SSD parameters and emulate NAND latency with high precision makes FSSD a valuable platform for SSD research and development. FSSD is also open-sourced to benefit the research community. Luyang Yu, Yizhen Lu, Meghna Mandava, Edward Richter, Vikram S. Mailthody, Seungwon Min, Wen-Mei W. Hwu, Deming Chen |
FPL | 6 |
| 2022 | Graph Neural Network Training and Data TieringabstractGraph Neural Networks (GNNs) have shown success in learning from graph-structured data, with applications to fraud detection, recommendation, and knowledge graph reasoning. However, training GNN efficiently is challenging because: 1) GPU memory capacity is limited and can be insufficient for large datasets, and 2) the graph-based data structure causes irregular data access patterns. In this work, we provide a method to statistically analyze and identify more frequently accessed data ahead of GNN training. Our data tiering method not only utilizes the structure of input graph, but also an insight gained from actual GNN training process to achieve a higher prediction result. With our data tiering method, we additionally provide a new data placement and access strategy to further minimize the CPU-GPU communication overhead. We also take into account of multi-GPU GNN training as well and we demonstrate the effectiveness of our strategy in a multi-GPU system. The evaluation results show that our work reduces CPU-GPU traffic by 87-95% and improves the training speed of GNN over the existing solutions by 1.6-2.1x on graphs with hundreds of millions of nodes and billions of edges. Seungwon Min, Kun Wu 0002, Mert Hidayetoglu, Jinjun Xiong, Xiang Song 0003, Wen-Mei W. Hwu |
KDD | 1 |
| 2021 | Large Graph Convolutional Network Training with GPU-Oriented Data Communication ArchitectureabstractGraph Convolutional Networks (GCNs) are increasingly adopted in large-scale graph-based recommender systems. Training GCN requires the minibatch generator traversing graphs and sampling the sparsely located neighboring nodes to obtain their features. Since real-world graphs often exceed the capacity of GPU memory, current GCN training systems keep the feature table in host memory and rely on the CPU to collect sparse features before sending them to the GPUs. This approach, however, puts tremendous pressure on host memory bandwidth and the CPU. This is because the CPU needs to (1) read sparse features from memory, (2) write features into memory as a dense format, and (3) transfer the features from memory to the GPUs. In this work, we propose a novel GPU-oriented data communication approach for GCN training, where GPU threads directly access sparse features in host memory through zero-copy accesses without much CPU help. By removing the CPU gathering stage, our method significantly reduces the consumption of the host resources and data access latency. We further present two important techniques to achieve high host memory access efficiency by the GPU: (1) automatic data access address alignment to maximize PCIe packet efficiency, and (2) asynchronous zero-copy access and kernel execution to fully overlap data transfer with training. We incorporate our method into PyTorch and evaluate its effectiveness using several graphs with sizes up to 111 million nodes and 1.6 billion edges. In a multi-GPU training setup, our method is 65--92% faster than the conventional data transfer method, and can even match the performance of all-in-GPU-memory training for some graphs that fit in GPU memory. Seungwon Min, Kun Wu 0002, Sitao Huang, Mert Hidayetoglu, Jinjun Xiong, Eiman Ebrahimi, Deming Chen, Wen-Mei W. Hwu |
Proc. VLDB Endow. | 1 |
| 2020 | EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal In GPUsabstractModern analytics and recommendation systems are increasingly based on graph data that capture the relations between entities being analyzed. Practical graphs come in huge sizes, offer massive parallelism, and are stored in sparse-matrix formats such as compressed sparse row (CSR). To exploit the massive parallelism, developers are increasingly interested in using GPUs for graph traversal. However, due to their sizes, graphs often do not fit into the GPU memory. Prior works have either used input data pre-processing/partitioning or unified virtual memory (UVM) to migrate chunks of data from the host memory to the GPU memory. However, the large, multi-dimensional, and sparse nature of graph data presents a major challenge to these schemes and results in significant amplification of data movement and reduced effective data throughput. In this work, we propose EMOGI, an alternative approach to traverse graphs that do not fit in GPU memory using direct cache-line-sized access to data stored in host memory. This paper addresses the open question of whether a sufficiently large number of overlapping cache-line-sized accesses can be sustained to 1) tolerate the long latency to host memory, 2) fully utilize the available bandwidth, and 3) achieve favorable execution performance. We analyze the data access patterns of several graph traversal applications in GPU over PCIe using an FPGA to understand the cause of poor external bandwidth utilization. By carefully coalescing and aligning external memory requests, we show that we can minimize the number of PCIe transactions and nearly fully utilize the PCIe bandwidth with direct cache-line accesses to the host memory. EMOGI achieves 2.60X speedup on average compared to the optimized UVM implementations in various graph traversal applications. We also show that EMOGI scales better than a UVM-based solution when the system uses higher bandwidth interconnects such as PCIe 4.0. Seungwon Min, Vikram S. Mailthody, Zaid Qureshi, Jinjun Xiong, Eiman Ebrahimi, Wen-Mei W. Hwu |
Proc. VLDB Endow. | 1 |
| 2019 | Analysis and Optimization of I/O Cache Coherency Strategies for SoC-FPGA DeviceabstractUnlike traditional PCIe-based FPGA accelerators, heterogeneous SoC-FPGA devices provide tighter integrations between software running on CPUs and hardware accelerators. Modern heterogeneous SoC-FPGA platforms support multiple I/O cache coherence options between CPUs and FPGAs, but these options can have inadvertent effects on the achieved bandwidths depending on applications and data access patterns. To provide the most efficient communications between CPUs and accelerators, understanding the data transaction behaviors and selecting the right I/O cache coherence method is essential. In this paper, we use Xilinx Zynq UltraScale+ as the SoC platform to show how certain I/O cache coherence method can perform better or worse in different situations, ultimately affecting the overall accelerator performances as well. Based on our analysis, we further explore possible software and hardware modifications to improve the I/O performances with different I/O cache coherence options. With our proposed modifications, the overall performance of SoC design can be averagely improved by 20%. Seungwon Min, Sitao Huang, Mohamed El-Hadedy 0001, Jinjun Xiong, Deming Chen, Wen-Mei W. Hwu |
FPL | 1 |
| 2018 | Application-Transparent Near-Memory Processing Architecture with Memory Channel NetworkabstractThe physical memory capacity of servers is expected to increase drastically with deployment of the forthcoming non-volatile memory technologies. This is a welcomed improvement for emerging data-intensive applications. For such servers to be cost-effective, nonetheless, we must cost-effectively increase compute throughput and memory bandwidth commensurate with the increase in memory capacity without compromising application readiness. Tackling this challenge, we present Memory Channel Network (MCN) architecture in this paper. Specifically, first, we propose an MCN DIMM, an extension of a buffered DIMM where a small but capable processor called MCN processor is integrated with a buffer device on the DIMM for near-memory processing. Second, we implement device drivers to give the host and MCN processors in a server an illusion that they are independent heterogeneous nodes connected through an Ethernet link. These allow the host and MCN processors in a server to run a given data-intensive application together based on popular distributed computing frameworks such as MPI and Spark without any change in the host processor hardware and its application software, while offering the benefits of high-bandwidth and low-latency communications between the host and the MCN processors over memory channels. As such, MCN can serve as an application-transparent framework which can seamlessly unify near-memory processing within a server and distributed computing across such servers for data-intensive applications. Our simulation running the full software stack shows that a server with 8 MCN DIMMs offers 4.56X higher throughput and consume 47.5% less energy than a cluster with 9 conventional nodes connected through Ethernet links, as it facilitates up to 8.17X higher aggregate DRAM bandwidth utilization. Lastly, we demonstrate the feasibility of MCN with an IBM POWER8 system and an experimental buffered DIMM. Mohammad Alian, Seungwon Min, Hadi Asghari Moghaddam, Ashutosh Dhar, Dong Kai Wang, Adam J. McPadden, Oliver O'Halloran, Deming Chen, Jinjun Xiong, Daehoon Kim 0001, Wen-Mei W. Hwu, Nam Sung Kim |
MICRO | 2 |