EDBT 2026 Demo / reviewers in the wild / expert
JooHyoung Cha
dblp:393/3389
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0008-2123-454XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | Target-Aware Neural Network Execution via Compiler-Guided Pruning · IEEE Trans. Mob. Comput. 2026 |
Machine learning › Efficient and distributed learning › model compression
pruning |
1.0 | 1 | 2026 | Target-Aware Neural Network Execution via Compiler-Guided Pruning · IEEE Trans. Mob. Comput. 2026 |
Compilers and program optimization
autotuning |
1.0 | 1 | 2026 | Target-Aware Neural Network Execution via Compiler-Guided Pruning · IEEE Trans. Mob. Comput. 2026 |
Compilers and program optimization › compiler optimization
compiler-directed optimization |
1.0 | 1 | 2026 | Target-Aware Neural Network Execution via Compiler-Guided Pruning · IEEE Trans. Mob. Comput. 2026 |
Machine learning › Efficient and distributed learning
on-device inference |
0.3 | 1 | 2026 | Target-Aware Neural Network Execution via Compiler-Guided Pruning · IEEE Trans. Mob. Comput. 2026 |
Methods — techniques the papers use, named apart from their topics
learned latency estimation · 2.0compiler-informed pruning · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Target-Aware Neural Network Execution via Compiler-Guided PruningabstractMobile devices run deep learning models for various purposes, such as image classification and speech recognition. Due to the resource constraints of mobile devices, researchers have focused on either making a lightweight deep neural network (DNN) model using model pruning or generating an efficient code using compiler optimization. It was observed that the straightforward integration between model compression and compiler auto-tuning often fails to produce the most efficient model for a target device. We propose CPrune, a compiler-informed model pruning for efficient target-aware DNN execution to support an application with a required target accuracy. To address real-world deployment scenarios with resource or latency constraints, we further introduce RB-CPrune, a predictive variant that eliminates iterative tuning by using a learned latency estimator. CPrune makes a lightweight DNN model through informed pruning based on the structural information of subgraphs built during the compiler tuning process. Our experimental results show that CPrune increases the DNN execution speed up to 2.73× compared to the state-of-the-art TVM auto-tune while meeting the accuracy requirement. JooHyoung Cha, Jemin Lee 0003, Sangtae Ha, Yongin Kwon |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Multi-level Machine Learning-Guided Autotuning for Efficient Code Generation on a Deep Learning AcceleratorabstractThe growing complexity of deep learning models necessitates specialized hardware and software optimizations, particularly for deep learning accelerators. While machine learning-based autotuning methods have emerged as a promising solution to reduce manual effort, both template-based and template-free approaches suffer from prolonged tuning times due to the profiling of invalid configurations, which may result in runtime errors. To address this issue, we propose ML2Tuner, a multi-level machine learning-guided autotuning technique designed to improve efficiency and robustness. ML2Tuner introduces two key ideas: (1) a validity prediction model to filter out invalid configurations prior to profiling, and (2) an advanced performance prediction model that leverages hidden features extracted during the compilation process. Experimental results on an extended VTA accelerator demonstrate that ML2Tuner achieves equivalent performance improvements using only 12.3% of the samples required by a TVM-like approach and reduces invalid profiling attempts by an average of 60.8%, highlighting its potential to enhance autotuning performance by filtering out invalid configurations. JooHyoung Cha, Munyoung Lee, Jinse Kwon, Jemin Lee 0003, Yongin Kwon |
LCTES | 1 |
| 2025 | Luthier: Bridging Auto-Tuning and Vendor Libraries for Efficient Deep Learning InferenceabstractRecent deep learning compilers commonly adopt auto-tuning approaches that search for the optimal kernel configuration in tensor programming from scratch, requiring tens of hours per operation and neglecting crucial optimization factors for parallel computing on asymmetric multicore processors. Meanwhile, hand-optimized inference libraries from hardware vendors provide high performance but lack the flexibility and automation needed for emerging models. To close this gap, we propose Luthier , which significantly narrows the search space by selecting the best kernel from existing inference libraries, and also employs cost model-based profiling to quickly determine the most efficient workload distribution for parallel computing. As a result, Luthier achieves up to 2.0x faster execution on convolution-based vision models and transformer-based language models (BERT, GPT) on both CPUs and GPUs, while reducing average tuning time by 95% compared with ArmNN, AutoTVM, Ansor, ONNXRuntime, and TFLite. Yongin Kwon, JooHyoung Cha, Sehyeon Oh, Misun Yu, Jeman Park 0002, Jemin Lee 0003 |
ACM Trans. Embed. Comput. Syst. | 2 |