Jeongin Yun

dblp:265/6185 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision Transformers
abstract
How can we accurately quantize a pre-trained Vision Transformer model? Quantization algorithms compress Vision Transformers (ViTs) into low-bit formats, reducing memory and computation demands with minimal accuracy degradation. However, existing methods rely on uniform precision, ignoring the diverse sensitivity of ViT components to quantization. Metric-based Mixed Precision Quantization (MPQ) is a promising alternative, but previous MPQ methods for ViTs suffer from three major limitations: 1) coarse granularity, 2) mismatch in metric scale across component types, and 3) quantization-unaware bit allocation. In this paper, we propose LampQ (Layer-wise Mixed Precision Quantization for Vision Transformers), an accurate metric-based MPQ method for ViTs to overcome these limitations. LampQ performs layer-wise quantization to achieve both fine-grained control and efficient acceleration, incorporating a type-aware Fisher-based metric to measure sensitivity. Then, LampQ assigns bit-widths optimally through integer linear programming and further updates them iteratively. Extensive experiments show that LampQ provides the state-of-the-art performance in quantizing ViTs pre-trained on various tasks such as image classification, object detection, and zero-shot quantization.
Minjun Kim 0010, Jaeri Lee, Jongjin Kim 0001, Jeongin Yun, Yongmo Kwon, U Kang
AAAI4
2026 SharVeT: Similarity-aware Parameter Sharing with Vector-based Tuning for Efficient LLM Compression
abstract
How can we share parameters within large language models to significantly reduce memory costs while preserving accuracy?While parameter sharing is a promising solution to the memory overhead of large language models, existing methods rely on naive grouping and fail to correct sharing-induced discrepancies.We propose an accurate and efficient parameter sharing framework, SharVeT (Similarity-aware sharing with Vector-based Tuning), which performs similarity-based grouping to ensure accurate sharing, allocates parameters adaptively to preserve diversity within each group, and applies lightweight refinement with knowledge distillation to correct sharing-induced discrepancies.Experiments show that SharVeT outperforms existing sharing methods, achieving up to 32.1% lower perplexity and 21.2% higher few-shot reasoning accuracy.
Jeongin Yun, Jaeri Lee, Jongjin Kim 0001, Minjun Kim 0010, Jinho Song, U Kang
ACL (1)1
2025 DART: Diversified and Accurate Long-Tail Recommendation
Jeongin Yun, Jaeri Lee, U Kang
PAKDD (3)1
2024 Towards True Multi-interest Recommendation: Enhanced Scheme for Balanced Interest Training
abstract
How can we accurately capture users’ diverse interests to provide more relevant recommendations based on their historical interactions? Recent advancements in recommender systems have led to the development of multi-interest recommendation models that attempt to capture the diverse interests of users through multiple interest vectors. While theoretically promising, existing implementations frequently struggle with oversimplifying user interests, where models tend to focus on a single dominant vector and overlook the relationships between multiple interests, failing to represent the full complexity of users’ interests. This limits the models’ ability to truly personalize and diversify the recommendations provided to users. In response to this challenge, we propose BaM (Ba lanced Interest Learning for Multi-interest Recommendation), a versatile training scheme tailored for multi-interest recommendation models that ensures the full utilization of all interest vectors, leading to more effective recommendations. Instead of prioritizing an interest vector with the highest similarity to the ground-truth item for loss computation, BaM exploits a soft-selection approach, ensuring balanced training across multiple interest vectors. Furthermore, BaM trains all interest representations simultaneously through a multi-interest loss function that accounts for the contributions of every interest. This allows for a broader consideration of multiple interest vectors which are also related to the users’ diverse preferences with varying degrees of relevance. Extensive experiments with real-world datasets show that BaM achieves up to 31.43% higher accuracy in sequential recommendation compared to the best competitor, resulting in the state-of-the-art performance.
Jaeri Lee, Jeongin Yun, U Kang
IEEE Big Data2
2024 Cold-start Bundle Recommendation via Popularity-based Coalescence and Curriculum Heating
abstract
How can we recommend cold-start bundles to users? The cold-start problem in bundle recommendation is crucial because new bundles are continuously created on the Web for various marketing purposes. Despite its importance, existing methods for cold-start item recommendation are not readily applicable to bundles. They depend overly on historical information, even for less popular bundles, failing to address the primary challenge of the highly skewed distribution of bundle interactions. In this work, we propose CoHeat (Popularity-based Coalescence and Curriculum Heating), an accurate approach for cold-start bundle recommendation. CoHeat first represents users and bundles through graph-based views, capturing collaborative information effectively. To estimate the user-bundle relationship more accurately, CoHeat addresses the highly skewed distribution of bundle interactions through a popularity-based coalescence approach, which incorporates historical and affiliation information based on the bundle's popularity. Furthermore, it effectively learns latent representations by exploiting curriculum learning and contrastive learning. CoHeat demonstrates superior performance in cold-start bundle recommendation, achieving up to 193% higher nDCG@20 compared to the best competitor.
Hyunsik Jeon, Jongeun Lee, Jeongin Yun, U Kang
WWW3
2020 FleXOR: Trainable Fractional Quantization
abstract
Quantization based on the binary codes is gaining attention because each quantized bit can be directly utilized for computations without dequantization using look-up tables. Previous attempts, however, only allow for integer numbers of quantization bits, which ends up restricting the search space for compression ratio and accuracy. In this paper, we propose an encryption algorithm/architecture to compress quantized weights so as to achieve fractional numbers of bits per weight. Decryption during inference is implemented by digital XOR-gate networks added into the neural network model while XOR gates are described by utilizing $\tanh(x)$ for backward propagation to enable gradient calculations. We perform experiments using MNIST, CIFAR-10, and ImageNet to show that inserting XOR gates learns quantization/encrypted bit decisions through training and obtains high accuracy even for fractional sub 1-bit weights. As a result, our proposed method yields smaller size and higher model accuracy compared to binary neural networks.
Dongsoo Lee, Se Jung Kwon, Byeongwook Kim, Yongkweon Jeon, Baeseong Park, Jeongin Yun
NeurIPS6
2020 BiQGEMM: matrix multiplication with lookup table for binary-coding-based quantized DNNs
abstract
The number of parameters in deep neural networks (DNNs) is rapidly increasing to support complicated tasks and to improve model accuracy. Correspondingly, the amount of computations and required memory footprint increase as well. Quantization is an efficient method to address such concerns by compressing DNNs such that computations can be simplified while required storage footprint is significantly reduced. Unfortunately, commercial CPUs and GPUs do not fully support quantization because only fixed data transfers (such as 32 bits) are allowed. As a result, even if weights are quantized (by a non-uniform quantization scheme) into a few bits, CPUs and GPUs may not access multiple quantized weights without memory bandwidth waste. Success of quantization in practice, hence, relies on an efficient computation engine design, especially for matrix multiplication that is a basic computation engine in most DNNs. In this paper, we propose a novel matrix multiplication method, called BiQGEMM, dedicated to quantized DNNs. BiQGEMM can access multiple quantized weights simultaneously in one instruction. In addition, BiQGEMM pre-computes intermediate results that are highly redundant when quantization leads to limited available computation space. Since pre-computed values are stored in lookup tables and reused, BiQGEMM achieves lower amount of overall computations. Our extensive experimental results show that BiQGEMM presents higher performance than conventional schemes when DNNs are quantized.
Yongkweon Jeon, Baeseong Park, Se Jung Kwon, Byeongwook Kim, Jeongin Yun, Dongsoo Lee
SC5