EDBT 2026 Demo / reviewers in the wild / expert
Jaeho Lee 0001
dblp:78/6080-1
· DBLP profile ↗
35ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0002-1349-8595ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 4 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Speculative End-Turn Detector for Efficient Speech Chatbot AssistantabstractSpoken dialogue systems powered by large language models have demonstrated remarkable abilities in understanding human speech and generating appropriate spoken responses.However, these systems struggle with end-turn detection (ETD)-the ability to distinguish between user turn completion and hesitation.This limitation often leads to premature or delayed responses, disrupting the flow of spoken conversations.In this paper, we introduce the OpenETD Dataset, the first public dataset for end-turn detection.The OpenETD dataset consists of both synthetic speech data generated with text-to-speech models and real-world speech data collected from web sources.We also propose SpeculativeETD, a novel collaborative inference framework that balances efficiency and accuracy to improve real-time ETD in resource-constrained environments.Our approach jointly employs a lightweight GRUbased model, which rapidly detects the nonspeaking units in real-time on local devices, and a high-performance Wav2vec-based model running on the server to make a more challenging classification of distinguishing turn ends from mere pauses.Experiments demonstrate that the proposed SpeculativeETD significantly improves ETD accuracy while keeping the required computations low. Hyunjong Ok, Suho Yoo, Jaeho Lee 0001 |
ACL (1) | 3 |
| 2026 | IterQuant: Iterative Quantization Framework for Mixed-Precision LLM CompressionabstractMixed-precision quantization is a promising approach for compressing large language models (LLMs) while maintaining output quality. However, the design space for selecting resolutions of different layers makes exhaustive search intractable. Existing methods either rely on rigid bit-width allocation schemes or require extensive hyperparameter tuning, often based on inaccurate layer-wise sensitivity metrics. In this work, we propose IterQuant, an iterative quantization framework that efficiently explores the mixed-precision space without requiring exhaustive enumeration. By incorporating momentum-based scoring to reflect historical performance trends and parameter grouping to balance quantization granularity, IterQuant achieves favorable trade-offs between compression and accuracy. Unlike prior approaches that assume bit allocation sensitivity from full-precision models directly transfers to quantized models, IterQuant dynamically updates its quantization decisions as the model evolves, better capturing inter-layer dependencies. Experimental results demonstrate that IterQuant significantly outperforms state-of-the-art mixed-precision quantization approaches by 2.8% near 4 bits in preserving token-level output quality across various LLM benchmarks. Hyungyo Jeong, Hyeokjun Kwon, Jaeho Lee 0001, Youngjoo Lee 0002 |
DATE | 4 |
| 2025 | S2Cap: A Benchmark and a Baseline for Singing Style CaptioningabstractSinging voices contain much richer information than common voices, including varied vocal and acoustic properties. However, current open-source audio-text datasets for singing voices capture only a narrow range of attributes and lack acoustic features, leading to limited utility towards downstream tasks, such as style captioning. To fill this gap, we formally define the singing style captioning task and present S2Cap, a dataset of singing voices with detailed descriptions covering diverse vocal, acoustic, and demographic characteristics. Using this dataset, we develop an efficient and straightforward baseline algorithm for singing style captioning. The dataset is available at https://zenodo.org/records/15673764. Hyunjong Ok, Jaeho Lee 0001 |
CIKM | 2 |
| 2025 | AudioBERT: Audio Knowledge Augmented Language ModelabstractRecent studies have identified that language models, pretrained on text-only datasets, often lack elementary visual knowledge, e.g., colors of everyday objects. Motivated by this observation, we ask whether a similar shortcoming exists in terms of the auditory knowledge. To answer this question, we construct a new dataset called AuditoryBench, which consists of two novel tasks for evaluating auditory knowledge. Based on our analysis using the benchmark, we find that language models also suffer from a severe lack of auditory knowledge. To address this limitation, we propose AudioBERT, a novel method to augment the auditory knowledge of BERT through a retrieval-based approach. First, we detect auditory knowledge spans in prompts to query our retrieval model efficiently. Then, we inject audio knowledge into BERT and switch on low-rank adaptation for effective adaptation when audio knowledge is required. Our experiments demonstrate that AudioBERT is quite effective, achieving superior performance on the AuditoryBench. The dataset and code are available at https://github.com/HJ-Ok/AudioBERT. Hyunjong Ok, Suho Yoo, Jaeho Lee 0001 |
ICASSP | 3 |
| 2025 | ZIP: An Efficient Zeroth-order Prompt Tuning for Black-box Vision-Language ModelsabstractRecent studies have introduced various approaches for prompt-tuning black-box vision-language models, referred to as black-box prompt-tuning (BBPT). While BBPT has demonstrated considerable potential, it is often found that many existing methods require an excessive number of queries (i.e., function evaluations), which poses a significant challenge in real-world scenarios where the number of allowed queries is limited. To tackle this issue, we propose Zeroth-order Intrinsic-dimensional Prompt-tuning (ZIP), a novel approach that enables efficient and robust prompt optimization in a purely black-box setting. The key idea of ZIP is to reduce the problem dimensionality and the variance of zeroth-order gradient estimates, such that the training is done fast with far less queries. We achieve this by re-parameterizing prompts in low-rank representations and designing intrinsic-dimensional clipping of estimated gradients. We evaluate ZIP on 13+ vision-language tasks in standard benchmarks and show that it achieves an average improvement of approximately 6% in few-shot accuracy and 48% in query efficiency compared to the best-performing alternative BBPT methods, establishing a new state of the art. Our ablation analysis further shows that the proposed clipping mechanism is robust and nearly optimal, without the need to manually select the clipping threshold, matching the result of expensive hyperparameter search. Seonghwan Park 0003, Jaehyeon Jeong, Jaeho Lee 0001, Namhoon Lee |
ICLR | 4 |
| 2025 | Fast Training of Sinusoidal Neural Fields via Scaling InitializationabstractNeural fields are an emerging paradigm that represent data as continuous functions parameterized by neural networks. Despite many advantages, neural fields often have a high training cost, which prevents a broader adoption. In this paper, we focus on a popular family of neural fields, called sinusoidal neural fields (SNFs), and study how it should be initialized to maximize the training speed. We find that the standard initialization scheme for SNFs---designed based on the signal propagation principle---is suboptimal. In particular, we show that by simply multiplying each weight (except for the last layer) by a constant, we can accelerate SNF training by 10$\times$. This method, coined _weight scaling_, consistently provides a significant speedup over various data domains, allowing the SNFs to train faster than more recently proposed architectures. To understand why the weight scaling works well, we conduct extensive theoretical and empirical analyses which reveal that the weight scaling not only resolves the spectral bias quite effectively but also enjoys a well-conditioned optimization trajectory. Taesun Yeom, Jaeho Lee 0001 |
ICLR | 3 |
| 2025 | Prompt-based Depth Pruning of Large Language ModelsabstractDepth pruning aims to reduce the inference cost of a large language model without any hardware-specific complications, by simply removing several less important transformer blocks. However, our empirical findings suggest that the importance of a transformer block may be highly task-dependent—a block that is crucial for a task can be removed without degrading the accuracy on another task. Based on this observation, we develop a dynamic depth pruning algorithm, coined PuDDing (Prompt-routed Dynamic Depth Pruning), which determines which blocks to omit from the model based on the input prompt. PuDDing operates by training a lightweight router to predict the best omission set among a set of options, where this option set has also been constructed in a data-driven manner. Empirical results on commonsense reasoning benchmarks demonstrate that PuDDing effectively accelerates the inference language models, and achieves better on-task performance than static depth pruning baselines. Juyun Wee, Jaeho Lee 0001 |
ICML | 3 |
| 2025 | Communication-Efficient Split Learning via Adaptive Feature-Wise CompressionabstractThis article proposes a novel communication-efficient split learning (SL) framework, named SplitFC, which reduces the communication overhead required for transmitting intermediate features and gradient vectors during the SL training process. The key idea of SplitFC is to leverage different dispersion degrees exhibited in the columns of the matrices. SplitFC incorporates two compression strategies: 1) adaptive feature-wise dropout and 2) adaptive feature-wise quantization. In the first strategy, the intermediate feature vectors are dropped with adaptive dropout probabilities determined based on the standard deviation of these vectors. Then, by the chain rule, the intermediate gradient vectors associated with the dropped feature vectors are also dropped. In the second strategy, the non-dropped intermediate feature and gradient vectors are quantized using adaptive quantization levels determined based on the ranges of the vectors. To minimize the quantization error, the optimal quantization levels of this strategy are derived in a closed-form expression. Simulation results on the MNIST, CIFAR-100, and CelebA datasets demonstrate that SplitFC outperforms state-of-the-art SL frameworks by significantly reducing communication overheads while maintaining high accuracy. Yongjeong Oh, Jaeho Lee 0001, Christopher G. Brinton, Yo-Seb Jeon |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Discovering and Mitigating Visual Biases Through Keyword ExplanationabstractAddressing biases in computer vision models is crucial for real-world AI deployments. However, mitigating visual biases is challenging due to their unexplainable nature, often identified indirectly through visualization or sample statistics, which necessitates additional human supervision for interpretation. To tackle this issue, we propose the Bias-to-Text (B2T) framework, which interprets visual biases as keywords. Specifically, we extract common keywords from the captions of mispredicted images to identify potential biases in the model. We then validate these keywords by measuring their similarity to the mispredicted images using a vision-language scoring model. The keyword explanation form of visual bias offers several advantages, such as a clear group naming for bias discovery and a natural extension for debiasing using these group names. Our experiments demonstrate that B2T can identify known biases, such as gender bias in CelebA, background bias in Waterbirds, and distribution shifts in ImageNet-R/C. Additionally, B2T uncovers novel biases in larger datasets, such as Dollar Street and ImageNet. For example, we discovered a contextual bias between “bee” and “flower” in ImageNet. We also highlight various applications of B2T keywords, including debiased training, CLIP prompting, and model comparison.11Code: https://github.com/alinlab/b2t Sangwoo Mo, Minkyu Kim 0004, Kyungmin Lee, Jaeho Lee 0001, Jinwoo Shin |
CVPR | 5 |
| 2024 | In Search of a Data Transformation that Accelerates Neural Field TrainingabstractNeural field is an emerging paradigm in data representation that trains a neural network to approximate the given signal. A key obstacle that prevents its widespread adoption is the encoding speed-generating neural fields requires an overfitting of a neural network, which can take a significant number of SGD steps to reach the desired fidelity level. In this paper, we delve into the impacts of data transformations on the speed of neural field training, specifically focusing on how permuting pixel locations affect the convergence speed of SGD. Counterintuitively, we find that randomly permuting the pixel locations can considerably accelerate the training. To explain this phenomenon, we examine the neural field training through the lens of PSNR curves, loss landscapes, and error patterns. Our analyses suggest that the random pixel permutations remove the easy-to-fit patterns, which facilitate easy optimization in the early stage but hinder capturing fine details of the signal.11code: https://github.com/effl-lab/DT4Neural-Field Junwon Seo, Kwang In Kim, Jaeho Lee 0001 |
CVPR | 4 |
| 2024 | The Role of Masking for Efficient Supervised Knowledge Distillation of Vision Transformers
Seungwoo Son 0005, Jegwang Ryu, Namhoon Lee, Jaeho Lee 0001 |
ECCV (67) | 4 |
| 2024 | Decoding with Limited Teacher Supervision Requires Understanding When to Trust the TeacherabstractHow can small-scale large language models (LLMs) efficiently utilize the supervision of LLMs to improve their generative quality?This question has been well studied in scenarios where there is no restriction on the number of LLM supervisions one can use, giving birth to many decoding algorithms that utilize supervision without further training.However, it is still unclear what is an effective strategy under the limited supervision scenario, where we assume that no more than a few tokens can be generated by LLMs.To this end, we develop an algorithm to effectively aggregate the small-scale LLM and LLM predictions on initial tokens so that the generated tokens can more accurately condition the subsequent token generation by small-scale LLM only.Critically, we find that it is essential to adaptively overtrust or disregard the LLM prediction based on the confidence of the small-scale LLM.Through our experiments on a wide range of models and datasets, we demonstrate that our method provides a consistent improvement over conventional decoding strategies. Hyunjong Ok, Jegwang Ryu, Jaeho Lee 0001 |
EMNLP | 3 |
| 2024 | Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error MinimizationabstractThis work suggests fundamentally rethinking the current practice of pruning large language models (LLMs).The way it is done is by divide and conquer: split the model into submodels, sequentially prune them, and reconstruct predictions of the dense counterparts on small calibration data one at a time; the final model is obtained simply by putting the resulting sparse submodels together.While this approach enables pruning under memory constraints, it generates high reconstruction errors.In this work, we first present an array of reconstruction techniques that can significantly reduce this error by more than 90%.Unwittingly, however, we discover that minimizing reconstruction error is not always ideal and can overfit the given calibration data, resulting in rather increased language perplexity and poor performance at downstream tasks.We find out that a strategy of self-generating calibration data can mitigate this trade-off between reconstruction and generalization, suggesting new directions in the presence of both benefits and pitfalls of reconstruction for pruning LLMs. 1 Sungbin Shin, Wonpyo Park, Jaeho Lee 0001, Namhoon Lee |
EMNLP | 3 |
| 2024 | Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model QuantizationabstractDespite recent advances in LLM quantization, activation quantization remains to be challenging due to the activation outliers.Conventional remedies, e.g., mixing precisions for different channels, introduce extra overhead and reduce the speedup.In this work, we develop a simple yet effective strategy to facilitate per-tensor activation quantization by preventing the generation of problematic tokens.Precisely, we propose a method to find a set of key-value cache, coined CushionCache, which mitigates outliers in subsequent tokens when inserted as a prefix.CushionCache works in two steps: First, we greedily search for a prompt token sequence that minimizes the maximum activation values in subsequent tokens.Then, we further tune the token cache to regularize the activations of subsequent tokens to be more quantization-friendly.The proposed method successfully addresses activation outliers of LLMs, providing a substantial performance boost for per-tensor activation quantization methods.We thoroughly evaluate our method over a wide range of models and benchmarks and find that it significantly surpasses the established baseline of per-tensor W8A8 quantization and can be seamlessly integrated with the recent activation quantization method. Seungwoo Son 0005, Wonpyo Park, Woohyun Han, Kyuyeun Kim, Jaeho Lee 0001 |
EMNLP | 5 |
| 2024 | Hybrid Neural Representations for Spherical DataabstractIn this paper, we study hybrid neural representations for spherical data, a domain of increasing relevance in scientific research. In particular, our work focuses on weather and climate data as well as cosmic microwave background (CMB) data. Although previous studies have delved into coordinate-based neural representations for spherical signals, they often fail to capture the intricate details of highly nonlinear signals. To address this limitation, we introduce a novel approach named Hybrid Neural Representations for Spherical data (HNeR-S). Our main idea is to use spherical feature-grids to obtain positional features which are combined with a multi-layer perceptron to predict the target signal. We consider feature-grids with equirectangular and hierarchical equal area isolatitude pixelization structures that align with weather data and CMB data, respectively. We extensively verify the effectiveness of our HNeR-S for regression, super-resolution, temporal interpolation, and compression tasks. Hyomin Kim, Yunhui Jang, Jaeho Lee 0001, Sungsoo Ahn |
ICML | 3 |
| 2024 | Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual FidelityabstractRecent advances in text-guided image compression have shown great potential to enhance the perceptual quality of reconstructed images. These methods, however, tend to have significantly degraded pixel-wise fidelity, limiting their practicality. To fill this gap, we develop a new text-guided image compression algorithm that achieves both high perceptual and pixel-wise fidelity. In particular, we propose a compression framework that leverages text information mainly by text-adaptive encoding and training with joint image-text loss. By doing so, we avoid decoding based on text-guided generative models---known for high generative diversity---and effectively utilize the semantic information of text at a global level. Experimental results on various datasets show that our method can achieve high pixel-level and perceptual quality, with either human- or machine-generated captions. In particular, our method outperforms all baselines in terms of LPIPS, with some room for even more improvements when we use more carefully generated captions. Hagyeong Lee, Minkyu Kim 0004, Jun-Hyuk Kim, Seungeon Kim, Dokwan Oh, Jaeho Lee 0001 |
ICML | 6 |
| 2024 | SCANNER: Knowledge-Enhanced Approach for Robust Multi-modal Named Entity Recognition of Unseen EntitiesabstractHyunjong Ok, Taeho Kil, Sukmin Seo, Jaeho Lee. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hyunjong Ok, Taeho Kil, Sukmin Seo, Jaeho Lee 0001 |
NAACL-HLT | 4 |
| 2024 | Few-shot UnlearningabstractWe consider the problem of machine unlearning to erase the impact of a target dataset, used in training but incorrect or sensitive, from a trained model. It has been often presumed that every data sample to erase or remain is entirely identifiable and thus clarifies the desired model behavior after unlearning. However, such a flawless identification can be infeasible in practice. We pose a further realistic yet challenging scenario, referred to as few-shot unlearning, where only a few samples of target data are provided while aiming at achieving the underlying intention (e.g., correcting mislabels, countering a certain privacy attack, or specifying nothing) behind the full target dataset. We then devise a few-shot unlearning method including a new model inversion technique, specialized for unlearning scenarios, to retrieve a proxy of the training dataset from the trained model if needed. We demonstrate that our method using only a tiny subset of target data can achieve similar performance to the state-of-the-art methods with full access to target data. Our code and results are available at https://github.com/ml-postech/Few-shot-Unlearning. Youngsik Yoon, Jinhwan Nam, Hyojeong Yun, Jaeho Lee 0001, Dongwoo Kim 0002, Jungseul Ok |
SP | 4 |
| 2024 | Attention-Aware Semantic Communications for Collaborative InferenceabstractWe propose a communication-efficient collaborative inference framework in the domain of edge inference, focusing on the efficient use of vision transformer (ViT) models. The partitioning strategy of conventional collaborative inference fails to reduce communication cost because of the inherent architecture of ViTs maintaining consistent layer dimensions across the entire transformer encoder. Therefore, instead of employing the partitioning strategy, our framework utilizes a lightweight ViT model on the edge device, with the server deploying a complicated ViT model. To enhance communication efficiency and achieve the classification accuracy of the server model, we propose two strategies: 1) attention-aware patch selection and 2) entropy-aware image transmission. Attention-aware patch selection leverages the attention scores generated by the edge device’s transformer encoder to identify and select the image patches critical for classification. This strategy enables the edge device to transmit only the essential patches to the server, significantly improving communication efficiency. Entropy-aware image transmission uses min-entropy as a metric to accurately determine whether to depend on the lightweight model on the edge device or to request the inference from the server model. In our framework, the lightweight ViT model on the edge device acts as a semantic encoder, efficiently identifying and selecting the crucial image information required for the classification task. Our experiments demonstrate that the proposed collaborative inference framework can reduce communication overhead by 68% with only a minimal loss in accuracy compared to the server model on the ImageNet dataset. Jiwoong Im, Nayoung Kwon, Taewoo Park, Jiheon Woo, Jaeho Lee 0001, Yongjune Kim 0001 |
IEEE Internet Things J. | 5 |
| 2023 | Modality-Agnostic Variational Compression of Implicit Neural RepresentationsabstractWe introduce a modality-agnostic neural compression algorithm based on a functional view of data and parameterised as an Implicit Neural Representation (INR). Bridging the gap between latent coding and sparsity, we obtain compact latent representations non-linearly mapped to a soft gating mechanism. This allows the specialisation of a shared INR network to each data item through subnetwork selection. After obtaining a dataset of such latent representations, we directly optimise the rate/distortion trade-off in a modality-agnostic space using neural compression. Variational Compression of Implicit Neural Representations (VC-INR) shows improved performance given the same representational capacity pre quantisation while also outperforming previous quantisation schemes used for other INR techniques.Our experiments demonstrate strong results over a large set of diverse modalities using the same algorithm without any modality-specific inductive biases. We show results on images, climate data, 3D shapes and scenes as well as audio and video, introducing VC-INR as the first INR-based method to outperform codecs as well-known and diverse as JPEG 2000, MP3 and AVC/HEVC on their respective modalities. Jonathan Schwarz, Jihoon Tack, Yee Whye Teh, Jaeho Lee 0001, Jinwoo Shin |
ICML | 4 |
| 2023 | Learning Large-scale Neural Fields via Context Pruned Meta-LearningabstractWe introduce an efficient optimization-based meta-learning technique for large-scale neural field training by realizing significant memory savings through automated online context point selection. This is achieved by focusing each learning step on the subset of data with the highest expected immediate improvement in model quality, resulting in the almost instantaneous modeling of global structure and subsequent refinement of high-frequency details. We further improve the quality of our meta-learned initialization by introducing a bootstrap correction resulting in the minimization of any error introduced by reduced context sets while simultaneously mitigating the well-known myopia of optimization-based meta-learning. Finally, we show how gradient re-scaling at meta-test time allows the learning of extremely high-quality neural fields in significantly shortened optimization procedures. Our framework is model-agnostic, intuitive, straightforward to implement, and shows significant reconstruction improvements for a wide range of signals. We provide an extensive empirical evaluation on nine datasets across multiple multiple modalities, demonstrating state-of-the-art results while providing additional insight through careful analysis of the algorithmic components constituting our method. Code is available at https://github.com/jihoontack/GradNCP Jihoon Tack, Subin Kim 0001, Sihyun Yu, Jaeho Lee 0001, Jinwoo Shin, Jonathan Schwarz |
NeurIPS | 4 |
| 2022 | Spread Spurious Attribute: Improving Worst-group Accuracy with Spurious Attribute Estimation
Junhyun Nam, Jaehyung Kim 0001, Jaeho Lee 0001, Jinwoo Shin |
ICLR | 3 |
| 2022 | Scalable Neural Video Representations with Learnable Positional FeaturesabstractSuccinct representation of complex signals using coordinate-based neural representations (CNRs) has seen great progress, and several recent efforts focus on extending them for handling videos. Here, the main challenge is how to (a) alleviate a compute-inefficiency in training CNRs to (b) achieve high-quality video encoding while (c) maintaining the parameter-efficiency. To meet all requirements (a), (b), and (c) simultaneously, we propose neural video representations with learnable positional features (NVP), a novel CNR by introducing "learnable positional features" that effectively amortize a video as latent codes. Specifically, we first present a CNR architecture based on designing 2D latent keyframes to learn the common video contents across each spatio-temporal axis, which dramatically improves all of those three requirements. Then, we propose to utilize existing powerful image and video codecs as a compute-/memory-efficient compression procedure of latent codes. We demonstrate the superiority of NVP on the popular UVG benchmark; compared with prior arts, NVP not only trains 2 times faster (less than 5 minutes) but also exceeds their encoding quality as 34.07$\rightarrow$34.57 (measured with the PSNR metric), even using $>$8 times fewer parameters. We also show intriguing properties of NVP, e.g., video inpainting, video frame interpolation, etc. Subin Kim 0001, Sihyun Yu, Jaeho Lee 0001, Jinwoo Shin |
NeurIPS | 3 |
| 2022 | Meta-Learning with Self-Improving Momentum TargetabstractThe idea of using a separately trained target model (or teacher) to improve the performance of the student model has been increasingly popular in various machine learning domains, and meta-learning is no exception; a recent discovery shows that utilizing task-wise target models can significantly boost the generalization performance. However, obtaining a target model for each task can be highly expensive, especially when the number of tasks for meta-learning is large. To tackle this issue, we propose a simple yet effective method, coined Self-improving Momentum Target (SiMT). SiMT generates the target model by adapting from the temporal ensemble of the meta-learner, i.e., the momentum network. This momentum network and its task-specific adaptations enjoy a favorable generalization performance, enabling self-improving of the meta-learner through knowledge distillation. Moreover, we found that perturbing parameters of the meta-learner, e.g., dropout, further stabilize this self-improving process by preventing fast convergence of the distillation loss during meta-training. Our experimental results demonstrate that SiMT brings a significant performance gain when combined with a wide range of meta-learning methods under various applications, including few-shot regression, few-shot classification, and meta-reinforcement learning. Code is available at https://github.com/jihoontack/SiMT. Jihoon Tack, Jongjin Park, Hankook Lee, Jaeho Lee 0001, Jinwoo Shin |
NeurIPS | 4 |
| 2021 | MASKER: Masked Keyword Regularization for Reliable Text ClassificationabstractPre-trained language models have achieved state-of-the-art accuracies on various text classification tasks, e.g., sentiment analysis, natural language inference, and semantic textual similarity. However, the reliability of the fine-tuned text classifiers is an often underlooked performance criterion. For instance, one may desire a model that can detect out-of-distribution (OOD) samples (drawn far from training distribution) or be robust against domain shifts. We claim that one central obstacle to the reliability is the over-reliance of the model on a limited number of keywords, instead of looking at the whole context. In particular, we find that (a) OOD samples often contain in-distribution keywords, while (b) cross-domain samples may not always contain keywords; over-relying on the keywords can be problematic for both cases. In light of this observation, we propose a simple yet effective fine-tuning method, coined masked keyword regularization (MASKER), that facilitates context-based prediction. MASKER regularizes the model to reconstruct the keywords from the rest of the words and make low-confidence predictions without enough context. When applied to various pre-trained language models (e.g., BERT, RoBERTa, and ALBERT), we demonstrate that MASKER improves OOD detection and cross-domain generalization without degrading classification accuracy. Code is available at https://github.com/alinlab/MASKER. Seung Jun Moon, Sangwoo Mo, Kimin Lee, Jaeho Lee 0001, Jinwoo Shin |
AAAI | 4 |
| 2021 | Provable Memorization via Deep Neural Networks using Sub-linear ParametersabstractIt is known that $O(N)$ parameters are sufficient for neural networks to memorize arbitrary $N$ input-label pairs. By exploiting depth, we show that $O(N^{2/3})$ parameters suffice to memorize $N$ pairs, under a mild condition on the separation of input points. In particular, deeper networks (even with width 3) are shown to memorize more pairs than shallow networks, which also agrees with the recent line of works on the benefits of depth for function approximation. We also provide empirical results that support our theoretical findings. Jaeho Lee 0001, Chulhee Yun, Jinwoo Shin |
COLT | 2 |
| 2021 | Co2L: Contrastive Continual LearningabstractRecent breakthroughs in self-supervised learning show that such algorithms learn visual representations that can be transferred better to unseen tasks than cross-entropy based methods which rely on task-specific supervision. In this paper, we found that the similar holds in the continual learning context: contrastively learned representations are more robust against the catastrophic forgetting than ones trained with the cross-entropy objective. Based on this novel observation, we propose a rehearsal-based continual learning algorithm that focuses on continually learning and maintaining transferable representations. More specifically, the proposed scheme (1) learns representations using the contrastive learning objective, and (2) preserves learned representations using a self-supervised distillation step. We conduct extensive experimental validations under popular benchmark image classification datasets, where our method sets the new state-of-the-art performance. Source code is available at https://github.com/chaht01/Co2L. Hyuntak Cha, Jaeho Lee 0001, Jinwoo Shin |
ICCV | 2 |
| 2021 | Layer-adaptive Sparsity for the Magnitude-based Pruning
Jaeho Lee 0001, Sangwoo Mo, Sungsoo Ahn, Jinwoo Shin |
ICLR | 1 |
| 2021 | Minimum Width for Universal Approximation
Chulhee Yun, Jaeho Lee 0001, Jinwoo Shin |
ICLR | 3 |
| 2021 | Meta-Learning Sparse Implicit Neural RepresentationsabstractImplicit neural representations are a promising new avenue of representing general signals by learning a continuous function that, parameterized as a neural network, maps the domain of a signal to its codomain; the mapping from spatial coordinates of an image to its pixel values, for example. Being capable of conveying fine details in a high dimensional signal, unboundedly of its domain, implicit neural representations ensure many advantages over conventional discrete representations. However, the current approach is difficult to scale for a large number of signals or a data set, since learning a neural representation---which is parameter heavy by itself---for each signal individually requires a lot of memory and computations. To address this issue, we propose to leverage a meta-learning approach in combination with network compression under a sparsity constraint, such that it renders a well-initialized sparse parameterization that evolves quickly to represent a set of unseen signals in the subsequent training. We empirically demonstrate that meta-learned sparse neural representations achieve a much smaller loss than dense meta-learned models with the same number of parameters, when trained to fit each signal using the same number of optimization steps. Jaeho Lee 0001, Jihoon Tack, Namhoon Lee, Jinwoo Shin |
NeurIPS | 1 |
| 2020 | Lookahead: A Far-sighted Alternative of Magnitude-based Pruning
Jaeho Lee 0001, Sangwoo Mo, Jinwoo Shin |
ICLR | 2 |
| 2020 | Learning Bounds for Risk-sensitive LearningabstractIn risk-sensitive learning, one aims to find a hypothesis that minimizes a risk-averse (or risk-seeking) measure of loss, instead of the standard expected loss. In this paper, we propose to study the generalization properties of risk-sensitive learning schemes whose optimand is described via optimized certainty equivalents (OCE): our general scheme can handle various known risks, e.g., the entropic risk, mean-variance, and conditional value-at-risk, as special cases. We provide two learning bounds on the performance of empirical OCE minimizer. The first result gives an OCE guarantee based on the Rademacher average of the hypothesis space, which generalizes and improves existing results on the expected loss and the conditional value-at-risk. The second result, based on a novel variance-based characterization of OCE, gives an expected loss guarantee with a suppressed dependence on the smoothness of the selected OCE. Finally, we demonstrate the practical implications of the proposed bounds via exploratory experiments on neural networks. Jaeho Lee 0001, Jinwoo Shin |
NeurIPS | 1 |
| 2020 | Learning from Failure: De-biasing Classifier from Biased ClassifierabstractNeural networks often learn to make predictions that overly rely on spurious corre- lation existing in the dataset, which causes the model to be biased. While previous work tackles this issue by using explicit labeling on the spuriously correlated attributes or presuming a particular bias type, we instead utilize a cheaper, yet generic form of human knowledge, which can be widely applicable to various types of bias. We first observe that neural networks learn to rely on the spurious correlation only when it is “easier” to learn than the desired knowledge, and such reliance is most prominent during the early phase of training. Based on the obser- vations, we propose a failure-based debiasing scheme by training a pair of neural networks simultaneously. Our main idea is twofold; (a) we intentionally train the first network to be biased by repeatedly amplifying its “prejudice”, and (b) we debias the training of the second network by focusing on samples that go against the prejudice of the biased network in (a). Extensive experiments demonstrate that our method significantly improves the training of network against various types of biases in both synthetic and real-world datasets. Surprisingly, our framework even occasionally outperforms the debiasing methods requiring explicit supervision of the spuriously correlated attributes. Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee 0001, Jinwoo Shin |
NeurIPS | 4 |
| 2018 | Minimax Statistical Learning with Wasserstein distancesabstractAs opposed to standard empirical risk minimization (ERM), distributionally robust optimization aims to minimize the worst-case risk over a larger ambiguity set containing the original empirical distribution of the training data. In this work, we describe a minimax framework for statistical learning with ambiguity sets given by balls in Wasserstein space. In particular, we prove generalization bounds that involve the covering number properties of the original ERM problem. As an illustrative example, we provide generalization guarantees for transport-based domain adaptation problems where the Wasserstein distance between the source and target domain distributions can be reliably estimated from unlabeled samples. Jaeho Lee 0001, Maxim Raginsky |
NeurIPS | 1 |
| 2015 | On MMSE estimation from quantized observations in the nonasymptotic regimeabstractThis paper studies MMSE estimation on the basis of quantized noisy observations. It presents nonasymptotic bounds on MMSE regret due to quantization for two settings: (1) estimation of a scalar random variable given a quantized vector of n conditionally independent observations, and (2) estimation of a p-dimensional random vector given a quantized vector of n observations (not necessarily independent) when the full MMSE estimator has a subgaussian concentration property. Jaeho Lee 0001, Maxim Raginsky, Pierre Moulin |
ISIT | 1 |