VLDB 2026 Research / reviewers in the wild / expert
Jihyun Ryoo
dblp:142/0276
· DBLP profile ↗
8ranked-venue papers
3as first author
3since 2021 · last 2021
0000-0002-9116-8452ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Morphable Convolutional Neural Network for Biomedical Image SegmentationabstractWe propose a morphable convolution framework, which can be applied to irregularly shaped region of input feature map. This framework reduces the computational footprint of a regular CNN operation in the context of biomedical semantic image segmentation. The traditional CNN based approach has high accuracy, but suffers from high training and inference computation costs, compared to a conventional edge detection based approach. In this work, we combine the concept of morphable convolution with the edge detection algorithms resulting in a hierarchical framework, which first detects the edges and then generate a layer-wise annotation map. The annotation map guides the convolution operation to be run only on a small, useful fraction of pixels in the feature map. We evaluate our framework on three cell tracking datasets and the experimental results indicate that our framework saves ~30% and ~10% execution time on CPU and GPU, respectively, without loss of accuracy, compared to the baseline conventional CNN approaches. Huaipan Jiang, Anup Sarma, Mengran Fan, Jihyun Ryoo, Meenakshi Arunachalam, Sharada Naveen, Mahmut T. Kandemir |
DATE | 4 |
| 2021 | Distance-in-time versus distance-in-spaceabstractCache behavior is one of the major factors that influence the performance of applications. Most of the existing compiler techniques that target cache memories focus exclusively on reducing data reuse distances in time (DIT). However, current manycore systems employ distributed on-chip caches that are connected using an on-chip network. As a result, a reused data element/block needs to travel over this on-chip network, and the distance to be traveled -- reuse distance in space (DIS) -- can be as influential in dictating application performance as reuse DIT. This paper represents the first attempt at defining a compiler framework that accommodates both DIT and DIS. Specifically, it first classifies data reuses into four groups: G1: (low DIT, low DIS), G2: (high DIT, low DIS), G3: (low DIT, high DIS), and G4: (high DIT, high DIS). Then, observing that reuses in G1 represent the ideal case and there is nothing much to be done in computations in G4, it proposes a "reuse transfer" strategy that transfers select reuses between G2 and G3, eventually, transforming each reuse to either G1 or G4. Finally, it evaluates the proposed strategy using a set of 10 multithreaded applications. The collected results reveal that the proposed strategy reduces parallel execution times of the tested applications between 19.3% and 33.3%. Mahmut T. Kandemir, Xulong Tang, Hui Zhao 0013, Jihyun Ryoo, Mustafa Karaköy |
PLDI | 4 |
| 2021 | Compiler support for near data computingabstractRecent works from both hardware and software domains offer various optimizations that try to take advantage of near data computing (NDC) opportunities. While the results from these works indicate performance improvements of various magnitudes, the existing literature lacks a detailed quantification of the potential of NDC and analysis of compiler optimizations on tapping into that potential. This paper first presents an analysis of the NDC potential when executing multithreaded applications on manycore platforms. It then presents two compiler schemes designed to take advantage of NDC. The first of these schemes try to increase the amount of computation that can be performed in a hardware component, whereas the second compiler strategy strikes a balance between optimizing NDC and exploiting data reuse, by being more selective on when to perform NDC (even if the opportunity presents itself) and how. The collected experimental results on a 5×5 manycore system reveal that our first and second compiler schemes improve the overall performance of our multithreaded applications by, respectively, 22.5% and 25.2%, on average. Furthermore, these two compiler schemes are only 6.8% and 4.1% worse than an oracle scheme that makes the best near data computing decisions for each and every computation. Mahmut T. Kandemir, Jihyun Ryoo, Xulong Tang, Mustafa Karaköy |
PPoPP | 2 |
| 2020 | Collective Affinity Aware Computation MappingabstractThis work defines the concept of collective affinity. It is claimed that collective affinity has more potential than single core-centric affinity, for data locality optimization in manycores. The reason is that collective affinity captures the potential benefits of transferring computations originally assigned to one core to other cores. Next, building upon the collective affinity concept and a cache content estimation strategy, it presents a computation-to-core mapping strategy, specifically tuned for exploiting near data computing by reducing distance-to-data. Mahmut T. Kandemir, Jihyun Ryoo, Hui Zhao 0013, Myoungsoo Jung, Mustafa Karaköy |
PACT | 2 |
| 2019 | Architecture-Centric Bottleneck Analysis for Deep Neural Network ApplicationsabstractThe ever-growing complexity and popularity of machine learning and deep learning applications have motivated an urgent need of effective and efficient support for these applications on contemporary computing systems. In this paper, we thoroughly analyze the various DNN algorithms on three widely used architectures (CPU, GPU, and Xeon Phi). The DNN algorithms we choose for evaluation include i) Unet - for biomedical image segmentation, based on Convolutional Neural Network (CNN), ii) NMT - for neural machine translation based on Recurrent Neural Network (RNN), iii) ResNet-50, and iv) DenseNet - both for image processing based on CNNs. The ultimate goal of this paper is to answer four fundamental questions: i) whether the different DNN networks exhibit similar behavior on a given execution platform? ii) whether, across different platforms, a given DNN network exhibits different behaviors? iii) for the same execution platform and the same DNN network, whether different execution phases have different behaviors? and iv) are the current major general-purpose platforms tuned sufficiently well for different DNN algorithms? Motivated by these questions, we conduct an in-depth investigation of running DNN applications on modern systems. Specifically, we first identify the most time-consuming functions (hotspot functions) across different networks and platforms. Next, we characterize performance bottlenecks and discuss them in detail. Finally, we port selected hotspot functions to a cycle-accurate simulator, and use the results to direct architectural optimizations to better support DNN applications. Jihyun Ryoo, Mengran Fan, Xulong Tang, Huaipan Jiang, Meena Arunachalam, Sharada Naveen, Mahmut T. Kandemir |
HiPC | 1 |
| 2018 | Quantifying and Optimizing Data Access Parallelism on ManycoresabstractThe following topics are dealt with: storage management; cache storage; pattern clustering; cloud computing; optimisation; flash memories; resource allocation; scheduling; parallel processing; data mining. Jihyun Ryoo, Orhan Kislal, Xulong Tang, Mahmut T. Kandemir |
MASCOTS | 1 |
| 2015 | VIP: virtualizing IP chains on handheld platformsabstractEnergy-efficient user-interactive and display-oriented applications on handhelds rely heavily on multiple accelerators (termed IP cores) to meet their periodic frame processing needs. Further, these platforms are starting to host multiple applications concurrently on the multiple CPU cores. Unfortunately, today's hardware exposes an interface that forces the host software (Android drivers) to treat each IP core as an isolated device. Consequently, the host CPU has to get involved in the (i) processing of each frame, (ii) scheduling them to ensure timely progress through the IP cores to meet their QoS needs, and (iii) explicitly having to move data from one IP core to the next, with main memory serving as the common staging area. Nachiappan Chidambaram Nachiappan, Haibo Zhang 0005, Jihyun Ryoo, Niranjan Soundararajan, Anand Sivasubramaniam, Mahmut T. Kandemir, Ravi R. Iyer 0001, Chita R. Das |
ISCA | 3 |
| 2014 | Leveraging parallelism in the presence of control flow on CGRAsabstractCoarse-Grained Reconfigurable Architectures (CGRAs) are suitable for accelerating data-intensive applications in embedded systems due to high performance and power efficiency. However, as application programs become complex having more control flows in them, it becomes harder to accelerate such programs on CGRAs. Previous researches on this issue have focused on correct execution of control flows rather than their acceleration. This paper reveals how control flows degrade the performance of programs and proposes a software approaches to accelerating control flows by exploiting parallelism residing in each conditionals as well as among conditionals. Experiments show that our proposed techniques improve performance by 2.51 times on average. Jihyun Ryoo, Kyuseung Han, Kiyoung Choi |
ASP-DAC | 1 |