Sungyong Baik

dblp:243/2775 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0001-5702-4618ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 6 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BiasEdit: A Training-Free Bias-Detect-and-Edit Framework for Learning Fair Visual Classifiers
Jungwook Seo, Yoonsik Park, Sungyong Baik
WWW4
2026 Text-to-seed generation: Training-free open-vocabulary seeded semantic segmentation via re-purposing diffusion as text-guided seed generator
Kumju Jo, Heesun Jung, Sungyong Baik
Knowl. Based Syst.3
2025 SeeDiff: Off-the-Shelf Seeded Mask Generation from Diffusion Models
abstract
Entrusted with the goal of pixel-level object classification, the semantic segmentation networks entails the laborious preparation of pixel-level annotation masks. To obtain pixel-level annotation masks for a given class without human efforts, recent few works have proposed to generate pairs of images and annotation masks by employing image and text relationships modeled by text-to-image generative models, especially Stable Diffusion. However, these works do not fully exploit the capability of text-guided Diffusion models and thus require a pre-trained segmentation network, careful text prompt tuning, or the training of a segmentation network to generate final annotation masks. In this work, we take a closer look at attention mechanisms of Stable Diffusion, from which we draw connections with classical seeded segmentation approaches. In particular, we show that cross-attention alone provides very coarse object localization, which however can provide initial seeds. Then, akin to region expansion in seeded segmentation, we utilize the semantic-correspondence-modeling capability of self-attention to iteratively spread the attention to the whole class from the seeds using multi-scale self-attention maps. We also observe that a simple-text-guided synthetic image often has a uniform background, which is easier to find correspondences, compared to complex-structured objects. Thus, we further refine a mask using a more accurate background mask. Our proposed method, dubbed SeeDiff, generates high-quality masks off-the-shelf from Stable Diffusion, without additional training procedure, prompt tuning, or a pre-trained segmentation network.
Joon Hyun Park, Kumju Jo, Sungyong Baik
AAAI3
2025 Mitigating Parameter Interference in Model Merging via Sharpness-Aware Fine-Tuning
abstract
Large-scale deep learning models with a pretraining-finetuning paradigm have led to a surge of numerous task-specific models fine-tuned from a common pre-trained model. Recently, several research efforts have been made on merging these large models into a single multi-task model, particularly with simple arithmetic on parameters. Such merging methodology faces a central challenge: interference between model parameters fine-tuned on different tasks. Few recent works have focused on designing a new fine-tuning scheme that can lead to small parameter interference, however at the cost of the performance of each task-specific fine-tuned model and thereby limiting that of a merged model. To improve the performance of a merged model, we note that a fine-tuning scheme should aim for (1) smaller parameter interference and (2) better performance of each fine-tuned model on the corresponding task. In this work, we aim to design a new fine-tuning objective function to work towards these two goals. In the course of this process, we find such objective function to be strikingly similar to sharpness-aware minimization (SAM) objective function, which aims to achieve generalization by finding flat minima. Drawing upon our observation, we propose to fine-tune pre-trained models via sharpness-aware minimization. The experimental and theoretical results showcase the effectiveness and orthogonality of our proposed approach, improving performance upon various merging and fine-tuning methods. Our code is available at https://github.com/baiklab/SAFT-Merge.
Yeoreum Lee, Jinwook Jung, Sungyong Baik
ICLR3
2025 Find A Winning Sign: Sign Is All We Need to Win the Lottery
abstract
The Lottery Ticket Hypothesis (LTH) posits the existence of a sparse subnetwork (a.k.a. winning ticket) that can generalize comparably to its over-parameterized counterpart when trained from scratch. The common approach to finding a winning ticket is to preserve the original strong generalization through Iterative Pruning (IP) and transfer information useful for achieving the learned generalization by applying the resulting sparse mask to an untrained network. However, existing IP methods still struggle to generalize their observations beyond ad-hoc initialization and small-scale architectures or datasets, or they bypass these challenges by applying their mask to trained weights instead of initialized ones. In this paper, we demonstrate that the parameter sign configuration plays a crucial role in conveying useful information for generalization to any randomly initialized network. Through linear mode connectivity analysis, we observe that a sparse network trained by an existing IP method can retain its basin of attraction if its parameter signs and normalization layer parameters are preserved. To take a step closer to finding a winning ticket, we alleviate the reliance on normalization layer parameters by preventing high error barriers along the linear path between the sparse network trained by our method and its counterpart with initialized normalization layer parameters. Interestingly, across various architectures and datasets, we observe that any randomly initialized network can be optimized to exhibit low error barriers along the linear path to the sparse network trained by our method by inheriting its sparsity and parameter sign information, potentially achieving performance comparable to the original. The code is available at https://github.com/JungHunOh/AWS_ICLR2025.git.
Junghun Oh, Sungyong Baik, Kyoung Mu Lee
ICLR2
2025 Difficulty, Diversity, and Plausibility: Dynamic Data-Free Quantization
abstract
Without access to the original training data, data-free quantization (DFQ) aims to recover the performance loss induced by quantization. Most previous works have focused on using an original network to extract the train data information, which is instilled into surrogate synthesized images. However, existing DFQ methods do not take into account important aspects of quantization: the extent of a computational-cost-and-accuracy trade-off varies for each image, depending on its task difficulty. To handle such varying trade-offs, several efforts have been made to dynamically allocate bit-widths for each image. Such dynamic quantization, however, remains challenging and unexplored in the data-free domain, because synthesized images of previous works fail to possess properties in natural test images that are crucial for learning the appropriate dynamic allocation policy: difficulty, its diversity, and its plausibility. By contrast, we propose a data-free quantization framework that is dynamic-friendly, by modeling varying extents of task difficulties with plausibility. We generate plausibly difficult images with soft labels, whose probabilities are allocated to a group of similar classes. Images with diverse and plausible difficulties enable us to train the framework to dynamically handle the varying trade-offs. Consequently, our framework achieves better accuracy-complexity Pareto front than existing data-free quantization approaches.
Cheeun Hong, Sungyong Baik, Junghun Oh, Kyoung Mu Lee
WACV2
2025 CoAPT: Context Attribute words for Prompt Tuning
Gun Lee 0002, Subin An, Sungyong Baik, Soochahn Lee
Knowl. Based Syst.3
2024 LAN: Learning to Adapt Noise for Image Denoising
abstract
Removing noise from images, a.k.a image denoising, can be a very challenging task since the type and amount of noise can greatly vary for each image due to many factors including a camera model and capturing environments. While there have been striking improvements in image Denoising with the emergence of advanced deep learning architectures and real-world datasets, recent denoising net-works struggle to maintain performance on images with noise that has not been seen during training. One typical approach to address the challenge would be to adapt a Denoising network to new noise distribution. Instead, in this work, we shift our focus to adapting the input noise itself, rather than adapting a network. Thus, we keep a pretrained network frozen, and adapt an input noise to capture the fine-grained deviations. As such, we propose a new denoising algorithm, dubbed Learning-to-Adapt-Noise (LAN), where a learnable noise offset is directly added to a given noisy image to bring a given input noise closer towards the noise distribution a denoising network is trained to handle. Consequently, the proposed framework exhibits performance improvement on images with unseen noise, displaying the potential of the proposed research direction.
Changjin Kim, Tae Hyun Kim 0006, Sungyong Baik
CVPR3
2024 CLOSER: Towards Better Representation Learning for Few-Shot Class-Incremental Learning
Junghun Oh, Sungyong Baik, Kyoung Mu Lee
ECCV (49)2
2024 sRGB Real Noise Modeling via Noise-Aware Sampling with Normalizing Flows
abstract
Noise poses a widespread challenge in signal processing, particularly when it comes to denoising images. Although convolutional neural networks (CNNs) have exhibited remarkable success in this field, they are predicated upon the belief that noise follows established distributions, which restricts their practicality when dealing with real-world noise. To overcome this limitation, several efforts have been taken to collect noisy image datasets from the real world. Generative methods, employing techniques such as generative adversarial networks (GANs) and normalizing flows (NFs), have emerged as a solution for generating realistic noisy images. Recent works model noise using camera metadata, however requiring metadata even for sampling phase. In contrast, in this work, we aim to estimate the underlying camera settings, enabling us to improve noise modeling and generate diverse noise distributions. To this end, we introduce a new NF framework that allows us to both classify noise based on camera settings and generate various noisy images. Through experimental results, our model demonstrates exceptional noise quality and leads in denoising performance on benchmark datasets.
Dongjin Kim 0004, Donggoo Jung, Sungyong Baik, Tae Hyun Kim 0006
ICLR3
2024 Learning to Learn Task-Adaptive Hyperparameters for Few-Shot Learning
abstract
The objective of few-shot learning is to design a system that can adapt to a given task with only few examples while achieving generalization. Model-agnostic meta-learning (MAML), which has recently gained the popularity for its simplicity and flexibility, learns a good initialization for fast adaptation to a task under few-data regime. However, its performance has been relatively limited especially when novel tasks are different from tasks previously seen during training. In this work, instead of searching for a better initialization, we focus on designing a better fast adaptation process. Consequently, we propose a new task-adaptive weight update rule that greatly enhances the fast adaptation process. Specifically, we introduce a small meta-network that can generate per-step hyperparameters for each given task: learning rate and weight decay coefficients. The experimental results validate that learning a good weight update rule for fast adaptation is the equally important component that has drawn relatively less attention in the recent few-shot learning approaches. Surprisingly, fast adaptation from random initialization with ALFA can already outperform MAML. Furthermore, the proposed weight-update rule is shown to consistently improve the task-adaptation capability of MAML across diverse problem domains: few-shot classification, cross-domain few-shot classification, regression, visual tracking, and video frame interpolation.
Sungyong Baik, Myungsub Choi, Janghoon Choi, Kyoung Mu Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Looking beyond input frames: Self-supervised adaptation for video super-resolution
Jinsu Yoo, Jihoon Nam, Sungyong Baik, Tae Hyun Kim 0006
Pattern Recognit.3
2024 CoLaNet: Adaptive Context and Latent Information Blending for Face Image Inpainting
abstract
Face inpainting, the task of filling up missing regions in a face image plausibly, has witnessed great advances with deep learning-based approaches. To fill in the missing region, existing methods either use information from the surrounding visible region of the input image itself (i.e., context) or use prior knowledge obtained from the training data (i.e., latent). However, we find that exclusive usage of the two types of information is sub-optimal; whether the context-based approach is effective or the latent-based approach is effective is different for each missing region. To this end, we propose CoLaNet, a novel framework that adaptively blends context and latent information to inpaint face images. Specifically, the two types of information are balanced based on the attention between the missing region and the rest of the image. The regions strongly correlated to the visible region leverage context information more. Consequently, the adaptive utilization of context and latent information leads to better inpainting performance in various face images.
Joonkyu Park, Cheeun Hong, Sungyong Baik, Kyoung Mu Lee
IEEE Signal Process. Lett.3
2023 Deep Variational Bayesian Modeling of Haze Degradation Process
abstract
Relying on the representation power of neural networks, most recent works have often neglected several factors involved in haze degradation, such as transmission (the amount of light reaching an observer from a scene over distance) and atmospheric light. These factors are generally unknown, making dehazing problems ill-posed and creating inherent uncertainties. To account for such uncertainties and factors involved in haze degradation, we introduce a variational Bayesian framework for single image dehazing. We propose to take not only a clean image and but also transmission map as latent variables, the posterior distributions of which are parameterized by corresponding neural networks: dehazing and transmission networks, respectively. Based on a physical model for haze degradation, our variational Bayesian framework leads to a new objective function that encourages the cooperation between them, facilitating the joint training of and thereby boosting the performance of each other. In our framework, a dehazing network can estimate a clean image independently of a transmission map estimation during inference, introducing no overhead. Furthermore, our model-agnostic framework can be seamlessly incorporated with other existing dehazing networks, greatly enhancing the performance consistently across datasets and models.
Eun Woo Im, Junsung Shin, Sungyong Baik, Tae Hyun Kim 0006
CIKM3
2023 Semantic-Aware Dynamic Parameter for Video Inpainting Transformer
abstract
Recent learning-based video inpainting approaches have achieved considerable progress. However, they still cannot fully utilize semantic information within the video frames and predict improper scene layout, failing to restore clear object boundaries for mixed scenes. To mitigate this problem, we introduce a new transformer-based video inpainting technique that can exploit semantic information within the input and considerably improve reconstruction quality. In this study, we use the mixture-of-experts scheme and train multiple experts to handle mixed scenes, including various semantics. We leverage these multiple experts and produce locally (token-wise) different network parameters to achieve semantic-aware inpainting results. Extensive experiments on YouTube-VOS and DAVIS benchmark datasets demonstrate that, compared with existing conventional video inpainting approaches, the proposed method has superior performance in synthesizing visually pleasing videos with much clearer semantic structures and textures.
Eunhye Lee, Jinsu Yoo, Yunjeong Yang, Sungyong Baik, Tae Hyun Kim 0006
ICCV4
2023 Learning Controllable Degradation for Real-World Super-Resolution via Constrained Flows
abstract
Recent deep-learning-based super-resolution (SR) methods have been successful in recovering high-resolution (HR) images from their low-resolution (LR) counterparts, albeit on the synthetic and simple degradation setting: bicubic downscaling. On the other hand, super-resolution on real-world images demands the capability to handle complex downscaling mechanism which produces different artifacts (e.g., noise, blur, color distortion) upon downscaling factors. To account for complex downscaling mechanism in real-world LR images, there have been a few efforts in constructing datasets consisting of LR images with real-world downsampling degradation. However, making such datasets entails a tremendous amount of time and effort, thereby resorting to very few number of downscaling factors (e.g., $\times$2, $\times$3, $\times$4). To remedy the issue, we propose to generate realistic SR datasets for unseen degradation levels by exploring the latent space of real LR images and thereby producing more diverse yet realistic LR images with complex real-world artifacts. Our quantitative and qualitative experiments demonstrate the accuracy of the generated LR images, and we show that the various conventional SR networks trained with our newly generated SR datasets can produce much better HR images.
Seobin Park, Dongjin Kim 0004, Sungyong Baik, Tae Hyun Kim 0006
ICML3
2023 SANFlow: Semantic-Aware Normalizing Flow for Anomaly Detection
abstract
Visual anomaly detection, the task of detecting abnormal characteristics in images, is challenging due to the rarity and unpredictability of anomalies. In order to reliably model the distribution of normality and detect anomalies, a few works have attempted to exploit the density estimation ability of normalizing flow (NF). However, previous NF-based methods have relied solely on the capability of NF and forcibly transformed the distribution of all features to a single distribution (e.g., unit normal distribution), when features can have different semantic information and thus follow different distributions. We claim that forcibly learning to transform such diverse distributions to a single distribution with a single network will cause the learning difficulty, limiting the capacity of a network to discriminate normal and abnormal data. As such, we propose to transform the distribution of features at each location of a given image to different distributions. In particular, we train NF to map normal data distribution to distributions with the same mean but different variances at each location of the given image. To enhance the discriminability, we also train NF to map abnormal data distribution to a distribution with a mean that is different from that of normal data, where abnormal data is synthesized with data augmentation. The experimental results outline the effectiveness of the proposed framework in improving the density modeling and thus anomaly detection performance.
Sungyong Baik, Tae Hyun Kim 0006
NeurIPS2
2022 CADyQ: Content-Aware Dynamic Quantization for Image Super-Resolution
Cheeun Hong, Sungyong Baik, Seungjun Nah, Kyoung Mu Lee
ECCV (7)2
2022 DAQ: Channel-Wise Distribution-Aware Quantization for Deep Image Super-Resolution Networks
abstract
Since the resurgence of deep neural networks (DNNs), image super-resolution (SR) has recently seen a huge progress in improving the quality of low resolution images, however at the great cost of computations and resources. Recently, there has been several efforts to make DNNs more efficient via quantization. However, SR demands pixel-level accuracy in the system, it is more difficult to perform quantization without significantly sacrificing SR performance. To this end, we introduce a new ultra-low precision yet effective quantization approach specifically designed for SR. In particular, we observe that in recent SR networks, each channel has different distribution characteristics. Thus we propose a channel-wise distribution-aware quantization scheme. Experimental results demonstrate that our proposed quantization, dubbed Distribution-Aware Quantization (DAQ), manages to greatly reduce the computational and resource costs without the significant sacrifice in SR performance, compared to other quantization methods.
Cheeun Hong, Sungyong Baik, Junghun Oh, Kyoung Mu Lee
WACV3
2022 Batch Normalization Tells You Which Filter is Important
abstract
The goal of filter pruning is to search for unimportant filters to remove in order to make convolutional neural networks (CNNs) efficient without sacrificing the performance in the process. The challenge lies in finding information that can help determine how important or relevant each filter is with respect to the final output of neural networks. In this work, we share our observation that the batch normalization (BN) parameters of pre-trained CNNs can be used to estimate the feature distribution of activation outputs, without processing of training data. Upon observation, we propose a simple yet effective filter pruning method by evaluating the importance of each filter based on the BN parameters of pre-trained CNNs. The experimental results on CIFAR-10 and ImageNet demonstrate that the proposed method can achieve outstanding performance with and without fine-tuning in terms of the trade-off between the accuracy drop and the reduction in computational complexity and number of parameters of pruned networks.
Junghun Oh, Sungyong Baik, Cheeun Hong, Kyoung Mu Lee
WACV3
2022 Learning to Forget for Meta-Learning via Task-and-Layer-Wise Attenuation
abstract
Few-shot learning is an emerging yet challenging problem in which the goal is to achieve generalization from only few examples. Meta-learning tackles few-shot learning via the learning of prior knowledge shared across tasks and using it to learn new tasks. One of the most representative meta-learning algorithms is the model-agnostic meta-learning (MAML), which formulates prior knowledge as a common initialization, a shared starting point from where a learner can quickly adapt to unseen tasks. However, forcibly sharing an initialization can lead to conflicts among tasks and the compromised (undesired by tasks) location on optimization landscape, thereby hindering task adaptation. Furthermore, the degree of conflict is observed to vary not only among the tasks but also among the layers of a neural network. Thus, we propose task-and-layer-wise attenuation on the compromised initialization to reduce its adverse influence on task adaptation. As attenuation dynamically controls (or selectively forgets) the influence of the compromised prior knowledge for a given task and each layer, we name our method Learn to Forget (L2F). Experimental results demonstrate that the proposed method greatly improves the performance of the state-of-the-art MAML-based frameworks across diverse domains: few-shot classification, cross-domain few-shot classification, regression, reinforcement learning, and visual tracking.
Sungyong Baik, Junghoon Oh, Seokil Hong, Kyoung Mu Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Test-Time Adaptation for Video Frame Interpolation via Meta-Learning
abstract
Video frame interpolation is a challenging problem that involves various scenarios depending on the variety of foreground and background motions, frame rate, and occlusion. Therefore, generalizing across different scenes is difficult for a single network with fixed parameters. Ideally, one could have a different network for each scenario, but this will be computationally infeasible for practical applications. In this work, we propose MetaVFI, an adaptive video frame interpolation algorithm that uses additional information readily available at test time but has not been exploited in previous works. We initially show the benefits of test-time adaptation through simple fine-tuning of a network and then greatly improve its efficiency by incorporating meta-learning. Thus, we obtain significant performance gains with only a single gradient update without introducing any additional parameters. Moreover, the proposed MetaVFI algorithm is model-agnostic which can be easily combined with any video frame interpolation network. We show that our adaptive framework greatly improves the performance of baseline video frame interpolation networks on multiple benchmark datasets.
Myungsub Choi, Janghoon Choi, Sungyong Baik, Tae Hyun Kim 0006, Kyoung Mu Lee
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Meta-Learning with Task-Adaptive Loss Function for Few-Shot Learning
abstract
In few-shot learning scenarios, the challenge is to generalize and perform well on new unseen examples when only very few labeled examples are available for each task. Model-agnostic meta-learning (MAML) has gained the popularity as one of the representative few-shot learning methods for its flexibility and applicability to diverse problems. However, MAML and its variants often resort to a simple loss function without any auxiliary loss function or regularization terms that can help achieve better generalization. The problem lies in that each application and task may require different auxiliary loss function, especially when tasks are diverse and distinct. Instead of attempting to hand-design an auxiliary loss function for each application and task, we introduce a new meta-learning framework with a loss function that adapts to each task. Our proposed framework, named Meta-Learning with Task-Adaptive Loss Function (MeTAL), demonstrates the effectiveness and the flexibility across various domains, such as few-shot classification and few-shot regression.
Sungyong Baik, Janghoon Choi, Dohee Cho, Jaesik Min, Kyoung Mu Lee
ICCV1
2021 Searching for Controllable Image Restoration Networks
abstract
We present a novel framework for controllable image restoration that can effectively restore multiple types and levels of degradation of a corrupted image. The proposed model, named TASNet, is automatically determined by our neural architecture search algorithm, which optimizes the efficiency-accuracy trade-off of the candidate model architectures. Specifically, we allow TASNet to share the early layers across different restoration tasks and adaptively adjust the remaining layers with respect to each task. The shared task-agnostic layers greatly improve the efficiency while the task-specific layers are optimized for restoration quality, and our search algorithm seeks for the best balance between the two. We also propose a new data sampling strategy to further improve the overall restoration performance. As a result, TASNet achieves significantly faster GPU latency and lower FLOPs compared to the existing state-of-the-art models, while also showing visually more pleasing outputs. The source code and pre-trained models are available at https://github.com/ghimhw/TASNet.
Sungyong Baik, Myungsub Choi, Janghoon Choi, Kyoung Mu Lee
ICCV2
2020 Domain Adaptation of Learned Featuresfor Visual Localization
Sungyong Baik, Hyo Jin Kim 0004, Tianwei Shen, Eddy Ilg, Kyoung Mu Lee, Chris Sweeney
BMVC1
2020 Learning to Forget for Meta-Learning
abstract
Few-shot learning is a challenging problem where the goal is to achieve generalization from only few examples. Model-agnostic meta-learning (MAML) tackles the problem by formulating prior knowledge as a common initialization across tasks, which is then used to quickly adapt to unseen tasks. However, forcibly sharing an initialization can lead to conflicts among tasks and the compromised (undesired by tasks) location on optimization landscape, thereby hindering the task adaptation. Further, we observe that the degree of conflict differs among not only tasks but also layers of a neural network. Thus, we propose task-and-layer-wise attenuation on the compromised initialization to reduce its influence. As the attenuation dynamically controls (or selectively forgets) the influence of prior knowledge for a given task and each layer, we name our method as L2F (Learn to Forget). The experimental results demonstrate that the proposed method provides faster adaptation and greatly improves the performance. Furthermore, L2F can be easily applied and improve other state-of-the-art MAML-based frameworks, illustrating its simplicity and generalizability.
Sungyong Baik, Seokil Hong, Kyoung Mu Lee
CVPR1
2020 Scene-Adaptive Video Frame Interpolation via Meta-Learning
abstract
Video frame interpolation is a challenging problem because there are different scenarios for each video depending on the variety of foreground and background motion, frame rate, and occlusion. It is therefore difficult for a single network with fixed parameters to generalize across different videos. Ideally, one could have a different network for each scenario, but this is computationally infeasible for practical applications. In this work, we propose to adapt the model to each video by making use of additional information that is readily available at test time and yet has not been exploited in previous works. We first show the benefits of 'test-time adaptation' through simple fine-tuning of a network, then we greatly improve its efficiency by incorporating meta-learning. We obtain significant performance gains with only a single gradient update without any additional parameters. Finally, we show that our meta-learning framework can be easily employed to any video frame interpolation network and can consistently improve its performance on multiple benchmark datasets.
Myungsub Choi, Janghoon Choi, Sungyong Baik, Tae Hyun Kim 0006, Kyoung Mu Lee
CVPR3
2020 Meta-Learning with Adaptive Hyperparameters
abstract
Despite its popularity, several recent works question the effectiveness of MAML when test tasks are different from training tasks, thus suggesting various task-conditioned methodology to improve the initialization. Instead of searching for better task-aware initialization, we focus on a complementary factor in MAML framework, inner-loop optimization (or fast adaptation). Consequently, we propose a new weight update rule that greatly enhances the fast adaptation process. Specifically, we introduce a small meta-network that can adaptively generate per-step hyperparameters: learning rate and weight decay coefficients. The experimental results validate that the Adaptive Learning of hyperparameters for Fast Adaptation (ALFA) is the equally important ingredient that was often neglected in the recent few-shot learning approaches. Surprisingly, fast adaptation from random initialization with ALFA can already outperform MAML.
Sungyong Baik, Myungsub Choi, Janghoon Choi, Kyoung Mu Lee
NeurIPS1
2019 Learning to Remember Past to Predict Future for Visual Tracking
abstract
Fast and reliable adaptability to appearance variations of any target object has been the holy grail of visual tracking. Recently, Siamese-based trackers have demonstrated outstanding speed, however at the cost of adaptability and accuracy. We propose to model a temporal evolution of appearance features, allowing for adaptability without online training. Specifically, we introduce a memory-augmented convolutional recurrent neural network (RNN), named Past-to-Future (P2FNet), that takes appearance features as an input at each frame and predicts the next-frame features. RNN allows for fast adaptability to dynamically varying appearance, while the memory provides the generalization capability over longer sequences via template management. For reliability, we propose a new augmentation to train RNN to disregard corrupted features. A novel visualization method illustrates the reliable template management of the memory. The experimental results on benchmarks demonstrate the tracker shows competitive performance among real-time state-of-the-art trackers.
Sungyong Baik, Junseok Kwon, Kyoung Mu Lee
ICIP1