EDBT 2026 Demo / reviewers in the wild / expert
Jangryul Kim
dblp:276/3801
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2023
0000-0001-6764-9419ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Energy-Aware Scenario-Based Mapping of Deep Learning Applications Onto Heterogeneous Processors Under Real-Time ConstraintsabstractTo cope with the increasing demand for deep learning applications in embedded systems, emerging embedded devices tend to equip multiple heterogeneous processors, including GPU and deep learning hardware accelerator, called neural processing unit (NPU). It becomes popular to run multiple deep learning (DL) applications simultaneously to provide several functionalities. In this work, we assume that applications have real-time constraints that may vary at run time. While extensive studies have been conducted recently to find an efficient mapping of multiple DL applications on various hardware platforms, they do not consider the constraints imposed by the NPU and the associated software development kit (SDK) in a real embedded platform. In this paper, we propose a novel energy-aware mapping methodology of multiple DL applications onto a real embedded system that has multiple heterogeneous processors. The objective is to minimize energy consumption while satisfying the real-time constraints of all applications. In the proposed scheme, we first select Pareto-optimal mapping solutions for each application. Then mapping combination is explored, considering the scenario that indicates the dynamism of applications while satisfying the constraints. Also, we reduce energy consumption by tuning the frequency of processors. We could satisfy up to 40% higher deadline constraints and reduce the energy consumption by 22% ∼ 31% compared to the static mapping methods with real-life applications and different scenarios on a real platform. Jangryul Kim, Soonhoi Ha |
IEEE Trans. Computers | 1 |
| 2022 | TensorRT-Based Framework and Optimization Methodology for Deep Learning Inference on Jetson BoardsabstractAs deep learning inference applications are increasing in embedded devices, an embedded device tends to equip neural processing units (NPUs) in addition to a multi-core CPU and a GPU. NVIDIA Jetson AGX Xavier is an example. For fast and efficient development of deep learning applications, TensorRT is provided as the SDK for high-performance inference, including an optimizer and runtime that delivers low latency and high throughput for deep learning inference applications. Like most deep learning frameworks, TensorRT assumes that the inference is executed on a single processing element, GPU or NPU, not both. In this article, we present a TensorRT-based framework supporting various optimization parameters to accelerate a deep learning application targeted on an NVIDIA Jetson embedded platform with heterogeneous processors, including multi-threading, pipelining, buffer assignment, and network duplication. Since the design space of allocating layers to diverse processing elements and optimizing other parameters is huge, we devise a parameter optimization methodology that consists of a heuristic for balancing pipeline stages among heterogeneous processors and fine-tuning the process for optimizing parameters. With nine real-life benchmarks, we could achieve 101%~680% performance improvement and up to 55% energy reduction over the baseline inference using a GPU only. Eunjin Jeong, Jangryul Kim, Soonhoi Ha |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | Hierarchical Scheduling of an SDF/L Graph onto Multiple ProcessorsabstractAlthough dataflow models are known to thrive at exploiting task-level parallelism of an application, it is difficult to exploit the parallelism of data, represented well with loop structures, since these structures are not explicitly specified in existing dataflow models. SDF/L model overcomes this shortcoming by specifying the loop structures explicitly in a hierarchical fashion. We introduce a scheduling technique of an application represented by the SDF/L model onto heterogeneous processors. In the proposed method, we explore the mapping of tasks using an evolutionary meta-heuristic and schedule hierarchically in a bottom-up fashion, creating parallel loop schedules at lower levels first and then re-using them when constructing the schedule at a higher level. The efficiency of the proposed scheduling methodology is verified with benchmark examples and randomly generated SDF/L graphs. Mari-Liis Oldja, Jangryul Kim, Dowhan Jeong, Soonhoi Ha |
ACM Trans. Design Autom. Electr. Syst. | 2 |