VLDB 2026 Research / reviewers in the wild / expert
Yongin Kwon
dblp:130/9714
· DBLP profile ↗
19ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards an efficient dataflow-flexible accelerator by finding optimal dataflows of DNNsabstract• Leverage design space exploration tool to create an efficient accelerator for DNNs. • Reuse redundant hardware components to minimize chip area. • Schedule to minimize dataflow transitions to maximize efficiency. This paper proposes a new dataflow-flexible accelerator design that addresses the limitations of existing heterogeneous dataflow accelerator (HDA) for handling the computation of multiple deep neural network (DNN) models. The design offers increased dataflow flexibility and higher efficiency compared to existing works. The accelerator utilizes a fixed set of representative dataflows implemented as operating modes and switches between them dynamically. A design space exploration (DSE) tool is leveraged to evaluate the efficiency of candidate dataflows and determine the optimal number and types of operating modes. Each layer of the target DNN models is assessed with different operating modes to select the optimal mode for each layer. Also, two supplementary optimization techniques are adopted to reduce the overheads from supporting a multitude of dataflows. One optimizes to minimize the number of transitions of dataflows, which incur severe overheads. The other optimizes to maximize the reuse of hardware components associated with supporting multiple dataflows. By identifying the redundant hardware components, the proposed design minimizes the chip area, another aspect where dataflow-flexible accelerators suffer. Experimental results demonstrate that our algorithm achieves greater dataflow flexibility with high efficiency, Compared to HDA, our design is, on average, 34.6 % lower in latency at the cost of 6.4 % area and negligible energy overhead. Whoi Ree Ha, Dongju Lee, Deumji Woo, Jonghee Yoon, Yongin Kwon, Yunheung Paek |
Future Gener. Comput. Syst. | 9 |
| 2026 | Target-Aware Neural Network Execution via Compiler-Guided PruningabstractMobile devices run deep learning models for various purposes, such as image classification and speech recognition. Due to the resource constraints of mobile devices, researchers have focused on either making a lightweight deep neural network (DNN) model using model pruning or generating an efficient code using compiler optimization. It was observed that the straightforward integration between model compression and compiler auto-tuning often fails to produce the most efficient model for a target device. We propose CPrune, a compiler-informed model pruning for efficient target-aware DNN execution to support an application with a required target accuracy. To address real-world deployment scenarios with resource or latency constraints, we further introduce RB-CPrune, a predictive variant that eliminates iterative tuning by using a learned latency estimator. CPrune makes a lightweight DNN model through informed pruning based on the structural information of subgraphs built during the compiler tuning process. Our experimental results show that CPrune increases the DNN execution speed up to 2.73× compared to the state-of-the-art TVM auto-tune while meeting the accuracy requirement. JooHyoung Cha, Jemin Lee 0003, Sangtae Ha, Yongin Kwon |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Work-in-Progress: I-FlashAttention: Fully Integer Fused Attention for Efficient Vision TransformersabstractTransformer self-attention offers strong expressiveness, but its compute and memory cost grows rapidly with longer sequences. This results in frequent off-chip memory access, which becomes a major performance bottleneck. FlashAttention reduces this by dividing the sequence into tiles, computed entirely in on-chip memory. This avoids storing intermediate tensors off-chip and alleviates memory bandwidth issues. However, tile-wise online softmax requires floating-point operations for numerical stability using max-based scaling and accumulation. We propose I-FlashAttention, an integer-only version of FlashAttention. It uses shift-based exponential approximation and integer max-tracking to perform online softmax without floating point. All steps, from INT8 GEMM to output, are fused into a single Triton kernel. I-FlashAttention is 1.08× faster than FP16 FlashAttention and 7.10× faster than I-ViT. Sehyeon Oh, Yongin Kwon, Jemin Lee 0003 |
CASES | 2 |
| 2025 | Exploring the Trade-Offs: Quantization Methods, Task Difficulty, and Model Size in Large Language Models From Edge to GiantabstractQuantization has gained attention as a promising solution for the cost-effective deployment of large and small language models. However, most prior work has been limited to perplexity or basic knowledge tasks and lacks a comprehensive evaluation of recent models like Llama-3.3. In this paper, we conduct a comprehensive evaluation of instruction-tuned models spanning 1B to 405B parameters, applying four quantization methods across 13 datasets. Our findings reveal that (1) quantized models generally surpass smaller FP16 baselines, yet they often struggle with instruction-following and hallucination detection; (2) FP8 consistently emerges as the most robust option across tasks, and AWQ tends to outperform GPTQ in weight-only quantization; (3) smaller models can suffer severe accuracy drops at 4-bit quantization, while 70B-scale models maintain stable performance; (4) notably, \textit{hard} tasks do not always experience the largest accuracy losses, indicating that quantization magnifies a model’s inherent weaknesses rather than simply correlating with task difficulty; and (5) an LLM-based judge (MT-Bench) highlights significant performance declines in Coding and STEM tasks, though it occasionally reports improvements in reasoning. Sihyeong Park, Jinse Kwon, Jihun Oh, Yongin Kwon |
IJCAI | 5 |
| 2025 | Multi-level Machine Learning-Guided Autotuning for Efficient Code Generation on a Deep Learning AcceleratorabstractThe growing complexity of deep learning models necessitates specialized hardware and software optimizations, particularly for deep learning accelerators. While machine learning-based autotuning methods have emerged as a promising solution to reduce manual effort, both template-based and template-free approaches suffer from prolonged tuning times due to the profiling of invalid configurations, which may result in runtime errors. To address this issue, we propose ML2Tuner, a multi-level machine learning-guided autotuning technique designed to improve efficiency and robustness. ML2Tuner introduces two key ideas: (1) a validity prediction model to filter out invalid configurations prior to profiling, and (2) an advanced performance prediction model that leverages hidden features extracted during the compilation process. Experimental results on an extended VTA accelerator demonstrate that ML2Tuner achieves equivalent performance improvements using only 12.3% of the samples required by a TVM-like approach and reduces invalid profiling attempts by an average of 60.8%, highlighting its potential to enhance autotuning performance by filtering out invalid configurations. JooHyoung Cha, Munyoung Lee, Jinse Kwon, Jemin Lee 0003, Yongin Kwon |
LCTES | 5 |
| 2025 | QuantuneV2: Compiler-based local metric-driven mixed precision quantization for practical embedded AI applications
Jeongseok Kim, Jemin Lee 0003, Yongin Kwon, Daeyoung Kim 0001 |
Future Gener. Comput. Syst. | 3 |
| 2025 | Luthier: Bridging Auto-Tuning and Vendor Libraries for Efficient Deep Learning InferenceabstractRecent deep learning compilers commonly adopt auto-tuning approaches that search for the optimal kernel configuration in tensor programming from scratch, requiring tens of hours per operation and neglecting crucial optimization factors for parallel computing on asymmetric multicore processors. Meanwhile, hand-optimized inference libraries from hardware vendors provide high performance but lack the flexibility and automation needed for emerging models. To close this gap, we propose Luthier , which significantly narrows the search space by selecting the best kernel from existing inference libraries, and also employs cost model-based profiling to quickly determine the most efficient workload distribution for parallel computing. As a result, Luthier achieves up to 2.0x faster execution on convolution-based vision models and transformer-based language models (BERT, GPT) on both CPUs and GPUs, while reducing average tuning time by 95% compared with ArmNN, AutoTVM, Ansor, ONNXRuntime, and TFLite. Yongin Kwon, JooHyoung Cha, Sehyeon Oh, Misun Yu, Jeman Park 0002, Jemin Lee 0003 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2024 | LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
Yanming Wang, Vatshank Chaturvedi, Lokesh Gupta, Seyeon Kim 0001, Yongin Kwon, Sangtae Ha |
IJCAI | 6 |
| 2024 | Visual Preference Inference: An Image Sequence-Based Preference Reasoning in Tabletop Object ManipulationabstractIn robotic object manipulation, human preferences can often be influenced by the visual attributes of objects, such as color and shape. These properties play a crucial role in operating a robot to interact with objects and align with human intention. In this paper, we focus on the problem of inferring underlying human preferences from a sequence of raw visual observations in tabletop manipulation environments with a variety of object types, named Visual Preference Inference (VPI). To facilitate visual reasoning in the context of manipulation, we introduce the Chain-of-Visual-Residuals (CoVR) method. CoVR employs a prompting mechanism that describes the difference between the consecutive images (i.e., visual residuals) and incorporates such texts with a sequence of images to infer the user’s preference. This approach significantly enhances the ability to understand and adapt to dynamic changes in its visual environment during manipulation tasks. Furthermore, we incorporate such texts along with a sequence of images to infer the user’s preferences. Our method outperforms baseline methods in terms of extracting human preferences from visual sequences in both simulation and real-world environments. Code and videos are available at: https://joonhyung-lee.github.io/vpi/ Joonhyung Lee, Yongin Kwon, Jemin Lee 0003, Minwook Ahn |
IROS | 3 |
| 2024 | Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers With Bridge Block Reconstruction for IoT SystemsabstractRecently, vision transformers (ViTs) have superseded convolutional neural networks in numerous applications, including classification, detection, and segmentation. However, the high computational requirements of ViTs hinder their widespread implementation. To address this issue, researchers have proposed efficient hybrid transformer architectures that combine convolutional and transformer layers with optimized attention computation of linear complexity. Additionally, posttraining quantization (PTQ) has been proposed as a means of mitigating computational demands. For mobile devices, achieving optimal acceleration for ViTs necessitates the strategic integration of quantization techniques and efficient hybrid transformer structures. However, no prior investigation has applied quantization to efficient hybrid transformers. In this article, we discover that applying existing PTQ methods for ViTs to efficient hybrid transformers leads to a drastic accuracy drop, attributed to the four following challenges: 1) highly dynamic ranges; 2) zero-point overflow; 3) diverse normalization; and 4) limited model parameters (https://gitlab.com/ones-ai/q-hyvit. Jemin Lee 0003, Yongin Kwon, Sihyeong Park, Misun Yu, Jeman Park 0002, Hwanjun Song |
IEEE Internet Things J. | 2 |
| 2023 | Tensor slicing and optimization for multicore NPUs
Rafael C. F. Sousa, Márcio Machado Pereira, Yongin Kwon, Namsoon Jung, Michael Frank 0008, Guido Araujo |
J. Parallel Distributed Comput. | 3 |
| 2022 | CPrune: Compiler-Informed Model Pruning for Efficient Target-Aware DNN Execution
Taeho Kim 0002, Yongin Kwon, Jemin Lee 0003, Taeho Kim 0001, Sangtae Ha |
ECCV (20) | 2 |
| 2022 | Quantune: Post-training quantization of convolutional neural networks using extreme gradient boosting for fast deploymentabstractTo adopt convolutional neural networks (CNN) for a range of resource-constrained targets, it is necessary to compress the CNN models by performing quantization, whereby precision representation is converted to a lower bit representation. To overcome problems such as sensitivity of the training dataset, high computational requirements, and large time consumption, post-training quantization methods that do not require retraining have been proposed. In addition, to compensate for the accuracy drop without retraining, previous studies on post-training quantization have proposed several complementary methods: calibration, schemes, clipping, granularity, and mixed-precision. To generate a quantized model with minimal error, it is necessary to study all possible combinations of the methods because each of them is complementary and the CNN models have different characteristics. However, an exhaustive or a heuristic search is either too time-consuming or suboptimal. To overcome this challenge, we propose an auto-tuner known as Quantune, which builds a gradient tree boosting model to accelerate the search for the configurations of quantization and reduce the quantization error. We evaluate and compare Quantune with the random, grid, and genetic algorithms. The experimental results show that Quantune reduces the search time for quantization by approximately 36.5× with an accuracy loss of 0.07–0.65% across six CNN models, including the fragile ones (MobileNet, SqueezeNet, and ShuffleNet). To support multiple targets and adopt continuously evolving quantization works, Quantune is implemented on a full-fledged compiler for deep learning as an open-sourced project. Jemin Lee 0003, Misun Yu, Yongin Kwon, Taeho Kim 0001 |
Future Gener. Comput. Syst. | 3 |
| 2016 | Precise execution offloading for applications with dynamic behavior in mobile cloud computing
Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Yeongpil Cho, Yunheung Paek |
Pervasive Mob. Comput. | 1 |
| 2015 | Mantis: Efficient Predictions of Execution Time, Energy Usage, Memory Usage and Network Usage on Smart Mobile DevicesabstractWe present Mantis, a framework for predicting the computational resource consumption (CRC) of Android applications on given inputs accurately, and efficiently. A key insight underlying Mantis is that program codes often contain features that correlate with performance and these features can be automatically computed efficiently. Mantis synergistically combines techniques from program analysis and machine learning. It constructs concise CRC models by choosing from many program execution features only a handful that are most correlated with the program's CRC metric yet can be evaluated efficiently from the program's input. We apply program slicing to reduce evaluation time of a feature and automatically generate executable code snippets for efficiently evaluating features. Our evaluation shows that Mantis predicts four CRC metrics of seven Android apps with estimation error in the range of 0-11.1 percent by executing predictor code spending at most 1.3 percent of their execution time on Galaxy Nexus. Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang 0001, Petros Maniatis, Mayur Naik, Yunheung Paek |
IEEE Trans. Mob. Comput. | 1 |
| 2014 | CMcloud: Cloud Platform for Cost-Effective Offloading of Mobile ApplicationsabstractRecent efforts towards mobile cloud propose to offload mobile applications to cloud servers for the improved performance and battery life of mobile devices. However, existing schemes completely ignore the costs of cloud resources by assuming that idle servers are always available for free of charge. These unrealistic assumptions make each server run only a small load to achieve the guaranteed high offload performance. Therefore, these schemes cannot be applied to real-world commercial clouds which aim to minimize the operation costs by maximizing the server throughput, and then charge users for their resource usage. In this paper, we propose CMcloud, a novel cost-effective mobile-to-cloud offloading platform, which works nicely under the real-world cloud environments. CMcloud minimizes both the server costs and the user service fee by offloading as many mobile applications to a single server as possible, while satisfying the target performance of all applications. To achieve such goals, CMcloud exploits novel architecture performance modeling and server migration techniques. Our implementation shows that CMcloud can improve the data enter throughput by 84% over a conventional static light-load scheme (or a 2.7x higher per-socket throughput.) Alternatively, CMcloud reduces the number of service failures by 83% over a static high-load scheme, while even improving the throughput by 31%. Dongju Chae, Jihun Kim 0002, Jangwoo Kim, Jong Kim 0001, Seungjun Yang, Yeongpil Cho, Yongin Kwon, Yunheung Paek |
CCGRID | 7 |
| 2014 | Techniques to Minimize State Transfer Costs for Dynamic Execution Offloading in Mobile Cloud ComputingabstractIn order to meet the increasing demand for high performance in smartphones, recent studies suggested mobile cloud computing techniques that aim to connect the phones to adjacent powerful cloud servers to throw their computational burden to the servers. These techniques often employ execution offloading schemes that migrate a process between machines during its execution. In execution offloading, code regions to be executed on the server are decided statically or dynamically based on the complex analysis on execution time and process state transfer costs of every region. Expectedly, the transfer cost is a deciding factor for the success of execution offloading. According to our analysis, it is dominated by the total size of heap objects transferred over the network. But previous work did not try hard to minimize this size. Thus in this paper, we introduce novel techniques based on compiler code analysis that effectively reduce the transferred data size by transferring only the essential heap objects and the stack frames actually referenced in the server. The experiments exhibit that the reduced size positively influences not only the transfer time itself but also the overall effectiveness of execution offloading, and ultimately, improves the performance of our mobile cloud computing significantly in terms of execution time and energy consumption. Seungjun Yang, Donghyun Kwon, Hayoon Yi, Yeongpil Cho, Yongin Kwon, Yunheung Paek |
IEEE Trans. Mob. Comput. | 5 |
| 2013 | Fast dynamic execution offloading for efficient mobile cloud computingabstractIn order to meet the increasing demand for high performance in smartphones, recent studies suggested mobile cloud computing techniques that aim to connect the phones to adjacent powerful cloud servers to throw their computational burden to the servers. These techniques often employ execution offloading schemes that migrate a process between machines during its execution. In execution offloading, code regions to be executed on the server are decided statically or dynamically based on the complex analysis on execution time and process state transfer time of every region. Expectedly, the transfer time is a deciding factor for the success of execution offloading. According to our analysis, it is dominated by the total size of heap objects transferred over the network. But previous work did not try hard to minimize this size. Thus in this paper, we introduce novel techniques based on compiler code analysis that effectively reduce the transferred data size by transferring only the essential heap objects. The experiments exhibit that the reduced size positively influences not only the transfer time itself but also the overall effectiveness of execution offloading, and ultimately, improves the performance of our mobile cloud computing significantly in terms of execution time and power consumption. Seungjun Yang, Yongin Kwon, Yeongpil Cho, Hayoon Yi, Donghyun Kwon, Jonghee M. Youn, Yunheung Paek |
PerCom | 2 |
| 2013 | Mantis: Automatic Performance Prediction for Smartphone Applications
Yongin Kwon, Hayoon Yi, Donghyun Kwon, Seungjun Yang, Byung-Gon Chun, Ling Huang 0001, Petros Maniatis, Mayur Naik, Yunheung Paek |
USENIX ATC | 1 |