EDBT 2026 Demo / reviewers in the wild / expert
Jemin Lee 0003
dblp:37/1760-3
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-9332-3508ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision TransformersabstractQuantization-Aware Training (QAT) for vision transformers relies on expensive retraining to recover accuracy loss in non-linear layer quantization, which limits their use in resource-constrained environments. In contrast, existing Post-Training Quantization (PTQ) methods either partially quantize non-linear functions or adjust activation distributions to maintain accuracy but fail to achieve fully integer-only inference. In this paper, we introduce IPTQ-ViT, a PTQ framework for fully integer-only vision transformers without retraining. We present approximation functions: a polynomial-based GELU optimized for vision data and a bit-shifting-based Softmax designed to improve approximation accuracy in PTQ. In addition, we propose a unified metric integrating quantization sensitivity, perturbation, and computational complexity to select the optimal approximation function per activation layer. IPTQ-ViT consistently outperforms previous PTQ methods, achieving up to 6.44%p (avg. 1.78%p) top-1 accuracy improvement for image classification and 1.0 mAP for object detection under W8A8 and W4A8. IPTQ-ViT achieves accuracy and latency comparable to integer-only QAT methods. Gihwan Kim, Jemin Lee 0003, Hyungshin Kim |
WACV | 2 |
| 2026 | Target-Aware Neural Network Execution via Compiler-Guided PruningabstractMobile devices run deep learning models for various purposes, such as image classification and speech recognition. Due to the resource constraints of mobile devices, researchers have focused on either making a lightweight deep neural network (DNN) model using model pruning or generating an efficient code using compiler optimization. It was observed that the straightforward integration between model compression and compiler auto-tuning often fails to produce the most efficient model for a target device. We propose CPrune, a compiler-informed model pruning for efficient target-aware DNN execution to support an application with a required target accuracy. To address real-world deployment scenarios with resource or latency constraints, we further introduce RB-CPrune, a predictive variant that eliminates iterative tuning by using a learned latency estimator. CPrune makes a lightweight DNN model through informed pruning based on the structural information of subgraphs built during the compiler tuning process. Our experimental results show that CPrune increases the DNN execution speed up to 2.73× compared to the state-of-the-art TVM auto-tune while meeting the accuracy requirement. JooHyoung Cha, Jemin Lee 0003, Sangtae Ha, Yongin Kwon |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | Work-in-Progress: I-FlashAttention: Fully Integer Fused Attention for Efficient Vision TransformersabstractTransformer self-attention offers strong expressiveness, but its compute and memory cost grows rapidly with longer sequences. This results in frequent off-chip memory access, which becomes a major performance bottleneck. FlashAttention reduces this by dividing the sequence into tiles, computed entirely in on-chip memory. This avoids storing intermediate tensors off-chip and alleviates memory bandwidth issues. However, tile-wise online softmax requires floating-point operations for numerical stability using max-based scaling and accumulation. We propose I-FlashAttention, an integer-only version of FlashAttention. It uses shift-based exponential approximation and integer max-tracking to perform online softmax without floating point. All steps, from INT8 GEMM to output, are fused into a single Triton kernel. I-FlashAttention is 1.08× faster than FP16 FlashAttention and 7.10× faster than I-ViT. Sehyeon Oh, Yongin Kwon, Jemin Lee 0003 |
CASES | 3 |
| 2025 | Multi-level Machine Learning-Guided Autotuning for Efficient Code Generation on a Deep Learning AcceleratorabstractThe growing complexity of deep learning models necessitates specialized hardware and software optimizations, particularly for deep learning accelerators. While machine learning-based autotuning methods have emerged as a promising solution to reduce manual effort, both template-based and template-free approaches suffer from prolonged tuning times due to the profiling of invalid configurations, which may result in runtime errors. To address this issue, we propose ML2Tuner, a multi-level machine learning-guided autotuning technique designed to improve efficiency and robustness. ML2Tuner introduces two key ideas: (1) a validity prediction model to filter out invalid configurations prior to profiling, and (2) an advanced performance prediction model that leverages hidden features extracted during the compilation process. Experimental results on an extended VTA accelerator demonstrate that ML2Tuner achieves equivalent performance improvements using only 12.3% of the samples required by a TVM-like approach and reduces invalid profiling attempts by an average of 60.8%, highlighting its potential to enhance autotuning performance by filtering out invalid configurations. JooHyoung Cha, Munyoung Lee, Jinse Kwon, Jemin Lee 0003, Yongin Kwon |
LCTES | 4 |
| 2025 | QuantuneV2: Compiler-based local metric-driven mixed precision quantization for practical embedded AI applications
Jeongseok Kim, Jemin Lee 0003, Yongin Kwon, Daeyoung Kim 0001 |
Future Gener. Comput. Syst. | 2 |
| 2025 | Luthier: Bridging Auto-Tuning and Vendor Libraries for Efficient Deep Learning InferenceabstractRecent deep learning compilers commonly adopt auto-tuning approaches that search for the optimal kernel configuration in tensor programming from scratch, requiring tens of hours per operation and neglecting crucial optimization factors for parallel computing on asymmetric multicore processors. Meanwhile, hand-optimized inference libraries from hardware vendors provide high performance but lack the flexibility and automation needed for emerging models. To close this gap, we propose Luthier , which significantly narrows the search space by selecting the best kernel from existing inference libraries, and also employs cost model-based profiling to quickly determine the most efficient workload distribution for parallel computing. As a result, Luthier achieves up to 2.0x faster execution on convolution-based vision models and transformer-based language models (BERT, GPT) on both CPUs and GPUs, while reducing average tuning time by 95% compared with ArmNN, AutoTVM, Ansor, ONNXRuntime, and TFLite. Yongin Kwon, JooHyoung Cha, Sehyeon Oh, Misun Yu, Jeman Park 0002, Jemin Lee 0003 |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2024 | Visual Preference Inference: An Image Sequence-Based Preference Reasoning in Tabletop Object ManipulationabstractIn robotic object manipulation, human preferences can often be influenced by the visual attributes of objects, such as color and shape. These properties play a crucial role in operating a robot to interact with objects and align with human intention. In this paper, we focus on the problem of inferring underlying human preferences from a sequence of raw visual observations in tabletop manipulation environments with a variety of object types, named Visual Preference Inference (VPI). To facilitate visual reasoning in the context of manipulation, we introduce the Chain-of-Visual-Residuals (CoVR) method. CoVR employs a prompting mechanism that describes the difference between the consecutive images (i.e., visual residuals) and incorporates such texts with a sequence of images to infer the user’s preference. This approach significantly enhances the ability to understand and adapt to dynamic changes in its visual environment during manipulation tasks. Furthermore, we incorporate such texts along with a sequence of images to infer the user’s preferences. Our method outperforms baseline methods in terms of extracting human preferences from visual sequences in both simulation and real-world environments. Code and videos are available at: https://joonhyung-lee.github.io/vpi/ Joonhyung Lee, Yongin Kwon, Jemin Lee 0003, Minwook Ahn |
IROS | 4 |
| 2024 | Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers With Bridge Block Reconstruction for IoT SystemsabstractRecently, vision transformers (ViTs) have superseded convolutional neural networks in numerous applications, including classification, detection, and segmentation. However, the high computational requirements of ViTs hinder their widespread implementation. To address this issue, researchers have proposed efficient hybrid transformer architectures that combine convolutional and transformer layers with optimized attention computation of linear complexity. Additionally, posttraining quantization (PTQ) has been proposed as a means of mitigating computational demands. For mobile devices, achieving optimal acceleration for ViTs necessitates the strategic integration of quantization techniques and efficient hybrid transformer structures. However, no prior investigation has applied quantization to efficient hybrid transformers. In this article, we discover that applying existing PTQ methods for ViTs to efficient hybrid transformers leads to a drastic accuracy drop, attributed to the four following challenges: 1) highly dynamic ranges; 2) zero-point overflow; 3) diverse normalization; and 4) limited model parameters (https://gitlab.com/ones-ai/q-hyvit. Jemin Lee 0003, Yongin Kwon, Sihyeong Park, Misun Yu, Jeman Park 0002, Hwanjun Song |
IEEE Internet Things J. | 1 |
| 2022 | CPrune: Compiler-Informed Model Pruning for Efficient Target-Aware DNN Execution
Taeho Kim 0002, Yongin Kwon, Jemin Lee 0003, Taeho Kim 0001, Sangtae Ha |
ECCV (20) | 3 |
| 2022 | Quantune: Post-training quantization of convolutional neural networks using extreme gradient boosting for fast deploymentabstractTo adopt convolutional neural networks (CNN) for a range of resource-constrained targets, it is necessary to compress the CNN models by performing quantization, whereby precision representation is converted to a lower bit representation. To overcome problems such as sensitivity of the training dataset, high computational requirements, and large time consumption, post-training quantization methods that do not require retraining have been proposed. In addition, to compensate for the accuracy drop without retraining, previous studies on post-training quantization have proposed several complementary methods: calibration, schemes, clipping, granularity, and mixed-precision. To generate a quantized model with minimal error, it is necessary to study all possible combinations of the methods because each of them is complementary and the CNN models have different characteristics. However, an exhaustive or a heuristic search is either too time-consuming or suboptimal. To overcome this challenge, we propose an auto-tuner known as Quantune, which builds a gradient tree boosting model to accelerate the search for the configurations of quantization and reduce the quantization error. We evaluate and compare Quantune with the random, grid, and genetic algorithms. The experimental results show that Quantune reduces the search time for quantization by approximately 36.5× with an accuracy loss of 0.07–0.65% across six CNN models, including the fragile ones (MobileNet, SqueezeNet, and ShuffleNet). To support multiple targets and adopt continuously evolving quantization works, Quantune is implemented on a full-fledged compiler for deep learning as an open-sourced project. Jemin Lee 0003, Misun Yu, Yongin Kwon, Taeho Kim 0001 |
Future Gener. Comput. Syst. | 1 |
| 2020 | PASS: Reducing Redundant Notifications between a Smartphone and a Smartwatch for Energy SavingabstractSmartwatches have gained significant popularity in recent years. One major use of smartwatches is notification checking as an extended display. Smartwatch use provides an opportunity for energy saving because it affords a reduction in the frequency of smartphone use. However, sometimes user input is limited in smartwatches due to its small screen size and users are forced to use their smartphones to fully read notifications. In such cases, the advantage of an extended display in the smartwatch diminishes because users need to access their smartphones, which causes extra smartphone battery consumption. We define phone-preferable notification that requires a user to take further actions, such as checking detailed content and replying to a message. Given that phone-preferable notifications are likely to be handled on smartphones, it is possible to defer notification delivery to smartwatches. In this paper, we develop a novel notification manager, called PASS that automatically defers phone-preferable notifications and piggybacks them on watch-preferable notifications. For model building and evaluation, we collect 15,659 notifications in-the-wild from 11 users for approximately 31 days. In addition, for approximately two months we gather self-reported data from five users regarding which devices were used to respond to notifications. The results show that PASS can save daily battery for smartwatches and daily battery for smartphones up to 43.5 and 0.9 percent without introducing any noticeable negative results on user experiences. Jemin Lee 0003, Uichin Lee, Hyungshin Kim |
IEEE Trans. Mob. Comput. | 1 |