Sungwoo Ahn

dblp:16/1530 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
1since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 67% GPUs and heterogeneous computing · 33%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
convolution acceleration
0.412020
Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor Cores · MICRO 2020
GPUs and heterogeneous computing › GPU architecture
GPU tensor cores
0.412020
Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor Cores · MICRO 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.412020
Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor Cores · MICRO 2020
Machine learning › Deep learning architectures and training
convolutional neural network
0.112020
Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor Cores · MICRO 2020

Methods — techniques the papers use, named apart from their topics

register renaming · 0.9load history buffer · 0.9compiler analysis · 0.9
YearPublicationVenuePosition
2024 Recompiling QAOA Circuits on Various Rotational Directions
abstract
The quantum approximate optimization algorithm (QAOA) is introduced to efficiently solve combinatorial optimization problems. Despite the promise of QAOA, the cost of executing QAOA circuits at scale for quantum advantage may still be excessive for the near-future quantum device. We observe the increasing overhead of QAOA circuit execution in the native gate translation. To execute QAOA circuits on a real quantum computing device, Hamiltonians composed of predefined specific rotations (e.g., ZZ and X) should be decomposed into finite native gates. By adopting rotational combinations that utilize native gates more directly than the standard QAOA circuit model, the execution cost on real quantum devices can be reduced. In this study, we propose Racoon (Rotational Space Virtualization for QAOA Ansatz), an algorithm-hardware co-design approach that revisits the synthesis conditions of QAOA circuits and selects alternative candidates with different rotational combinations. Our analysis of six commercial quantum processors demonstrates that applying Racoon to QAOA circuits for the 4-node Sherrington-Kirkpatrick model reduces the number of native gates by an average of 23% and up to 79%. Consequently, using Racoon results in 43% fewer training epochs, 41% lower training energy consumption, and a 6% improvement in inference on average compared to standard QAOA. Racoon consistently reduces circuit depth as the number of qubits and layers increases, achieving 123 × more circuit depth reduction compared to the recently proposed Depth First Search (DFS)-based method. Furthermore, we confirm that Racoon’s method can be extended to State-of-The-Art QAOAs with modified ansätze and to the variational quantum eigensolver (VQE).
Enhyeok Jang, Dongho Ha, Seungwoo Choi 0001, Youngmin Kim 0005, Jaewon Kwon, Yongju Lee 0003, Sungwoo Ahn, Hyungseok Kim 0003, Won Woo Ro
PACT7
2020 Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor Cores
abstract
This paper introduces a GPU architecture named Duplo that minimizes redundant memory accesses of convolutions in deep neural networks (DNNs). Convolution is one of the fundamental operations used in various classes of DNNs, and it takes the majority of execution time. Various approaches have been proposed to accelerate convolutions via general matrix multiplication (GEMM), Winograd convolution, fast Fourier transform (FFT), etc. Recent introduction of tensor cores in NVIDIA GPUs particularly targets on accelerating neural network computations. A tensor core in a streaming multiprocessor (SM) is a specialized unit dedicated to handling matrix-multiply-and-accumulate (MMA) operations. The underlying operations of tensor cores represent GEMM calculations, and lowering a convolution can effectively exploit the tensor cores by transforming deeply nested convolution loops into matrix multiplication. However, lowering the convolution has a critical drawback since it requires a larger memory space (or workspace) to compute the matrix multiplication, where the expanded workspace inevitably creates multiple duplicates of the same data stored at different memory addresses. The proposed Duplo architecture tackles this challenge by leveraging compile-time information and microarchitectural supports to detect and eliminate redundant memory accesses that repeatedly load the duplicates of data in the workspace matrix. Duplo identifies data duplication based on memory addresses and convolution information generated by a compiler. It uses a load history buffer (LHB) to trace the recent load history of workspace data and their presence in register file. Every load instruction of workspace data refers to the LHB to find if potentially the same copies of data exist in the register file. If data duplicates are found, Duplo simply renames registers and makes them point to the ones containing the same values instead of issuing memory requests to load the same data. Our experiment results show that Duplo improves the performance of DNNs by 29.4% on average and saves 34.1% of energy using tensor cores.
Sungwoo Ahn, Yunho Oh, Bogil Kim, Won Woo Ro, William J. Song
MICRO2
2016 Lengths of Attractors and Transients in Neuronal Networks with Random Connectivities
abstract
We study how the dynamics of a class of discrete dynamical system models for neuronal networks depends on the connectivity of the network. Specifically, we assume that the network is a directed Erdös--R\'enyi random graph and analytically derive scaling laws for the average lengths of the attractors and transients. In contrast to earlier results that were reported in [D. Terman, S. Ahn, X. Wang, and W. Just, Phys. D, 237 (2008), pp. 324--338], here we focus on the connection probabilities near the phase transition where the most complex dynamics is expected to occur.
Winfried Just, Sungwoo Ahn
SIAM J. Discret. Math.2
2008 Reordering of Location Identifiers for Indexing an RFID Tag Object Database
Sungwoo Ahn, Bonghee Hong
DEXA1
2008 Efficient Query Processing for Tracing RFID Tags by Reordering Location Identifiers
abstract
This paper addresses the problem of using the location identifier (LID) as the domain value of the index for trajectories of RFID tags and proposes the solution for solving this problem. The query performance for tracing tags depends upon the distribution of tag trajectories in the data space. We investigate a more efficient representation of tag trajectories by means of ordering the set of values in a 3-dimensional domain. Our analysis shows that the order of LIDs makes a greater contribution to the efficiency of query processing, compared with other domain values. However, there is no rule of assigning an LID to the RFID location in order to process queries efficiently. To solve this problem, we propose a new LID proximity function to rearrange an arbitrary order of LIDs. This function enables logically adjacent tag trajectories, which are accessed simultaneously, to be stored in close proximity on the disk. To determine the optimal sequence of LIDs in the domain, we also propose a reordering scheme of LIDs. Our experiments show that the proposed reordering scheme improves the performance of queries, compared with the previous method of assigning LIDs.
Sungwoo Ahn, Bonghee Hong
RTCSA1
2006 Design and Implementation of an Index Structure Using Fixed Intervals for Tracing of RFID Tags
Sungwoo Ahn, Bonghee Hong, ChaeHoon Ban, Kihyung Lee
ICCSA (2)1