Minyoung Kim 0001

dblp:34/5733-1 · DBLP profile ↗
← Back
62ranked-venue papers
56as first author
16since 2021 · last 2026
0000-0003-1621-6407ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 48 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 13 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 7 first-author
YearPublicationVenuePosition
2026 FedP²EFT: Federated Learning to Personalize PEFT for Multilingual LLMs
abstract
Federated learning (FL) has enabled training of multilingual large language models (LLMs) on diverse and decentralized multilingual data, especially on low-resource languages. To improve client-specific performance, personalization via the use of parameter-efficient fine-tuning (PEFT) modules such as LoRA is common. This involves a personalization strategy (PS), such as the design of the PEFT adapter structures (e.g., in which layers to add LoRAs and what ranks) and choice of hyperparameters (e.g., learning rates) for fine-tuning. Instead of manual PS configuration, we propose FedP²EFT, a federated learning-to-personalize method for multilingual LLMs in cross-device FL settings. Unlike most existing PEFT structure selection methods, which are prone to overfitting low-data regimes, FedP²EFT collaboratively learns the optimal personalized PEFT structure for each client via Bayesian sparse rank selection. Evaluations on both simulated and real-world multilingual FL benchmarks demonstrate that FedP²EFT largely outperforms existing personalized fine-tuning methods, while complementing other existing FL methods.
Royson Lee, Minyoung Kim 0001, Fady Rezk, Rui Li 0052, Stylianos I. Venieris, Timothy M. Hospedales
AAAI2
2025 A Stochastic Approach to Bi-Level Optimization for Hyperparameter Optimization and Meta Learning
abstract
We tackle the general differentiable meta learning problem that is ubiquitous in modern deep learning, including hyperparameter optimization, loss function learning, few-shot learning and more. These problems are often formalized as Bi-Level Optimizations (BLO). We introduce a novel perspective by turning a given BLO problem into a stochastic optimization, where the inner loss function becomes a smooth probability distribution, and the outer loss becomes an expected loss over the inner distribution. To solve this stochastic optimization, we adopt Stochastic Gradient Langevin Dynamics (SGLD) MCMC to sample inner distribution, and propose a recurrent algorithm to compute the MC-estimated hypergradient. Our derivation is similar to forward-mode differentiation, but we introduce a new first-order approximation that makes it feasible for large models without needing to store huge Jacobian matrices. The main benefits are two fold: i) Our stochastic formulation takes into account uncertainty, which makes the method robust to suboptimal inner optimization or non-unique multiple inner minima due to overparametrization; ii) Compared to existing methods that often exhibit unstable behavior and hyperparameter sensitivity in practice, our method leads to considerably more reliable solutions. We demonstrate that the new approach achieves promising results on diverse meta learning problems and easily scales to learning 87M hyper-parameters in the case of Vision Transformers.
Minyoung Kim 0001, Timothy M. Hospedales
AAAI1
2025 LiFT: Learning to Fine-Tune via Bayesian Parameter Efficient Meta Fine-Tuning
abstract
We tackle the problem of parameter-efficient fine-tuning (PEFT) of a pre-trained large deep model on many different but related tasks. Instead of the simple but strong baseline strategy of task-wise independent fine-tuning, we aim to meta-learn the core shared information that can be used for unseen test tasks to improve the prediction performance further. That is, we propose a method for {\em learning-to-fine-tune} (LiFT). LiFT introduces a novel hierarchical Bayesian model that can be superior to both existing general meta learning algorithms like MAML and recent LoRA zoo mixing approaches such as LoRA-Retriever and model-based clustering. In our Bayesian model, the parameters of the task-specific LoRA modules are regarded as random variables where these task-wise LoRA modules are governed/regularized by higher-level latent random variables, which represents the prior of the LoRA modules that capture the shared information across all training tasks. To make the posterior inference feasible, we propose a novel SGLD-Gibbs sampling algorithm that is computationally efficient. To represent the posterior samples from the SGLD-Gibbs, we propose an online EM algorithm that maintains a Gaussian mixture representation for the posterior in an online manner in the course of iterative posterior sampling. We demonstrate the effectiveness of LiFT on NLP and vision multi-task meta learning benchmarks.
Minyoung Kim 0001, Timothy M. Hospedales
ICLR1
2025 FedHB: Hierarchical Bayesian Federated Learning
abstract
We propose a novel hierarchical Bayesian approach to Federated Learning (FL), where our model reasonably describes the generative process of clients' local data via hierarchical Bayesian modeling: constituting random variables of local models for clients that are governed by a higher-level global variate. Interestingly, the variational inference in our Bayesian model leads to an optimisation problem whose block-coordinate descent solution becomes a distributed algorithm that is separable over clients and allows them not to reveal their own private data at all, thus fully compatible with FL. We also highlight that our block-coordinate algorithm has particular forms that subsume the well-known FL algorithms including Fed-Avg and Fed-Prox as special cases. Beyond introducing novel modeling and derivations, we also offer convergence analysis showing that our block-coordinate FL algorithm converges to an (local) optimum of the objective at the rate of $O(1/\sqrt{t})$, the same rate as regular (centralised) SGD, as well as the generalisation error analysis where we prove that the test error of our model on unseen data is guaranteed to vanish as we increase the training data size, thus asymptotically optimal.
Minyoung Kim 0001, Timothy M. Hospedales
J. Mach. Learn. Res.1
2024 A Hierarchical Bayesian Model for Few-Shot Meta Learning
abstract
We propose a novel hierarchical Bayesian model for the few-shot meta learning problem. We consider episode-wise random variables to model episode-specific generative processes, where these local random variables are governed by a higher-level global random variable. The global variable captures information shared across episodes, while controlling how much the model needs to be adapted to new episodes in a principled Bayesian manner. Within our framework, prediction on a novel episode/task can be seen as a Bayesian inference problem. For tractable training, we need to be able to relate each local episode-specific solution to the global higher-level parameters. We propose a Normal-Inverse-Wishart model, for which establishing this local-global relationship becomes feasible due to the approximate closed-form solutions for the local posterior distributions. The resulting algorithm is more attractive than the MAML in that it does not maintain a costly computational graph for the sequence of gradient descent steps in an episode. Our approach is also different from existing Bayesian meta learning methods in that rather than modeling a single random variable for all episodes, it leverages a hierarchical structure that exploits the local-global relationships desirable for principled Bayesian learning with many related tasks.
Minyoung Kim 0001, Timothy M. Hospedales
ICLR1
2024 A Bayesian Approach to Data Point Selection
abstract
Data point selection (DPS) is becoming a critical topic in deep learning due to the ease of acquiring uncurated training data compared to the difficulty of obtaining curated or processed data. Existing approaches to DPS are predominantly based on a bi-level optimisation (BLO) formulation, which is demanding in terms of memory and computation, and exhibits some theoretical defects regarding minibatches. Thus, we propose a novel Bayesian approach to DPS. We view the DPS problem as posterior inference in a novel Bayesian model where the posterior distributions of the instance-wise weights and the main neural network parameters are inferred under a reasonable prior and likelihood model. We employ stochastic gradient Langevin MCMC sampling to learn the main network and instance-wise weights jointly, ensuring convergence even with minibatches. Our update equation is comparable to the widely used SGD and much more efficient than existing BLO-based methods. Through controlled experiments in both the vision and language domains, we present the proof-of-concept. Additionally, we demonstrate that our method scales effectively to large language models and facilitates automated per-task optimization for instruction fine-tuning datasets.
Xinnuo Xu, Minyoung Kim 0001, Royson Lee, Brais Martínez, Timothy M. Hospedales
NeurIPS2
2023 SwAMP: Swapped Assignment of Multi-Modal Pairs for Cross-Modal Retrieval
abstract
We tackle the cross-modal retrieval problem, where learning is only supervised by the relevant multi-modal pairs in the data. Although the contrastive learning is the most popular approach for this task, it makes potentially wrong assumption that the instances in different pairs are automatically irrelevant. To address the issue, we propose a novel loss function that is based on self-labeling of the unknown semantic classes. Specifically, we aim to predict class labels of the data instances in each modality, and assign those labels to the corresponding instances in the other modality (i.e., swapping the pseudo labels). With these swapped labels, we learn the data embedding for each modality using the supervised cross-entropy loss. This way, cross-modal instances from different pairs that are semantically related can be aligned to each other by the class predictor. We tested our approach on several real-world cross-modal retrieval problems, including text-based video retrieval, sketch-based image retrieval, and image-text retrieval. For all these tasks our method achieves significant performance improvement over the contrastive learning.
Minyoung Kim 0001
AISTATS1
2023 Domain Generalisation via Domain Adaptation: An Adversarial Fourier Amplitude Approach
Minyoung Kim 0001, Da Li 0001, Timothy M. Hospedales
ICLR1
2023 BayesTune: Bayesian Sparse Deep Model Fine-tuning
abstract
Deep learning practice is increasingly driven by powerful foundation models (FM), pre-trained at scale and then fine-tuned for specific tasks of interest. A key property of this workflow is the efficacy of performing sparse or parameter-efficient fine-tuning, meaning that by updating only a tiny fraction of the whole FM parameters on a downstream task can lead to surprisingly good performance, often even superior to a full model update. However, it is not clear what is the optimal and principled way to select which parameters to update. Although a growing number of sparse fine-tuning ideas have been proposed, they are mostly not satisfactory, relying on hand-crafted heuristics or heavy approximation. In this paper we propose a novel Bayesian sparse fine-tuning algorithm: we place a (sparse) Laplace prior for each parameter of the FM, with the mean equal to the initial value and the scale parameter having a hyper-prior that encourages small scale. Roughly speaking, the posterior means of the scale parameters indicate how important it is to update the corresponding parameter away from its initial value when solving the downstream task. Given the sparse prior, most scale parameters are small a posteriori, and the few large-valued scale parameters identify those FM parameters that crucially need to be updated away from their initial values. Based on this, we can threshold the scale parameters to decide which parameters to update or freeze, leading to a principled sparse fine-tuning strategy. To efficiently infer the posterior distribution of the scale parameters, we adopt the Langevin MCMC sampler, requiring only two times the complexity of the vanilla SGD. Tested on popular NLP benchmarks as well as the VTAB vision tasks, our approach shows significant improvement over the state-of-the-arts (e.g., 1% point higher than the best SOTA when fine-tuning RoBERTa for GLUE and SuperGLUE benchmarks).
Minyoung Kim 0001, Timothy M. Hospedales
NeurIPS1
2023 FedL2P: Federated Learning to Personalize
abstract
Federated learning (FL) research has made progress in developing algorithms for distributed learning of global models, as well as algorithms for local personalization of those common models to the specifics of each client’s local data distribution. However, different FL problems may require different personalization strategies, and it may not even be possible to define an effective one-size-fits-all personalization strategy for all clients: Depending on how similar each client’s optimal predictor is to that of the global model, different personalization strategies may be preferred. In this paper, we consider the federated meta-learning problem of learning personalization strategies. Specifically, we consider meta-nets that induce the batch-norm and learning rate parameters for each client given local data statistics. By learning these meta-nets through FL, we allow the whole FL network to collaborate in learning a customized personalization strategy for each client. Empirical results show that this framework improves on a range of standard hand-crafted personalization baselines in both label and feature shift situations.
Royson Lee, Minyoung Kim 0001, Da Li 0001, Xinchi Qiu, Timothy M. Hospedales, Ferenc Huszar, Nicholas D. Lane
NeurIPS2
2022 Variational Continual Proxy-Anchor for Deep Metric Learning
abstract
The recent proxy-anchor method achieved outstanding performance in deep metric learning, which can be acknowledged to its data efficient loss based on hard example mining, as well as far lower sampling complexity than pair-based approaches. In this paper we extend the proxy-anchor method by posing it within the continual learning framework, motivated from its batch-expected loss form (instead of instance-expected, typical in deep learning), which can potentially incur the catastrophic forgetting of historic batches. By regarding each batch as a task in continual learning, we adopt the Bayesian variational continual learning approach to derive a novel loss function. Interestingly the resulting loss has two key modifications to the original proxy-anchor loss: i) we inject noise to the proxies when optimizing the proxy-anchor loss, and ii) we encourage momentum update to avoid abrupt model changes. As a result, the learned model achieves higher test accuracy than proxy-anchor due to the robustness to noise in data (through model perturbation during training), and the reduced batch forgetting effect. We demonstrate the improved results on several benchmark datasets.
Minyoung Kim 0001, Ricardo Guerrero, Hai Xuan Pham, Vladimir Pavlovic 0001
AISTATS1
2022 Pushing the Limits of Simple Pipelines for Few-Shot Learning: External Data and Fine-Tuning Make a Difference
abstract
Few-shot learning (FSL) is an important and topical problem in computer vision that has motivated extensive research into numerous methods spanning from sophisticated metalearning methods to simple transfer learning baselines. We seek to push the limits of a simple-but-effective pipeline for real-worldfew-shot image classification in practice. To this end, we explore few-shot learning from the perspective of neural architecture, as well as a three stage pipeline of pre-training on external data, meta-training with labelled few-shot tasks, and task-specific fine-tuning on unseen tasks. We investigate questions such as: ① How pre-training on external data benefits FSL? ② How state of the art transformer architectures can be exploited? and ③ How to best exploit finetuning? Ultimately, we show that a simple transformer-based pipeline yields surprisingly good performance on standard benchmarks such as Mini-ImageNet, CIFAR-FS, CDFSL and Meta-Dataset. Our code is available at https://hushell.github.io/pmf.
Shell Xu Hu, Da Li 0001, Jan Stühmer, Minyoung Kim 0001, Timothy M. Hospedales
CVPR4
2022 Gaussian Process Modeling of Approximate Inference Errors for Variational Autoencoders
abstract
Variational autoencoder (VAE) is a very successful generative model whose key element is the so-called amortized inference network, which can perform test time inference using a single feed forward pass. Unfortunately, this comes at the cost of degraded accuracy in posterior approximation, often underperforming the instance-wise variational optimization. Although the latest semi-amortized approaches mitigate the issue by performing a few variational optimization updates starting from the VAE's amortized inference output, they inherently suffer from computational overhead for inference at test time. In this paper, we address the problem in a completely different way by considering a random inference model, where we model the mean and variance functions of the variational posterior as random Gaussian processes (GP). The motivation is that the deviation of the VAE's amortized posterior distribution from the true posterior can be regarded as random noise, which allows us to view the approximation error as uncertainty in posterior approximation that can be dealt with in a principled GP manner. In particular, our model can quantify the difficulty in posterior approximation by a Gaussian variational density. Inference in our GP model is done by a single feed forward pass through the network, significantly faster than semi-amortized methods. We show that our approach attains higher test data likelihood than the state-of-the-arts on several benchmark datasets.
Minyoung Kim 0001
CVPR1
2022 Differentiable Expectation-Maximization for Set Representation Learning
Minyoung Kim 0001
ICLR1
2022 Fisher SAM: Information Geometry and Sharpness Aware Minimisation
abstract
Recent sharpness-aware minimisation (SAM) is known to find flat minima which is beneficial for better generalisation with improved robustness. SAM essentially modifies the loss function by the maximum loss value within the small neighborhood around the current iterate. However, it uses the Euclidean ball to define the neighborhood, which can be less accurate since loss functions for neural networks are typically defined over probability distributions (e.g., class predictive probabilities), rendering the parameter space no more Euclidean. In this paper we consider the information geometry of the model parameter space when defining the neighborhood, namely replacing SAM’s Euclidean balls with ellipsoids induced by the Fisher information. Our approach, dubbed Fisher SAM, defines more accurate neighborhood structures that conform to the intrinsic metric of the underlying statistical manifold. For instance, SAM may probe the worst-case loss value at either a too nearby or inappropriately distant point due to the ignorance of the parameter space geometry, which is avoided by our Fisher SAM. Another recent Adaptive SAM approach that stretches/shrinks the Euclidean ball in accordance with the scales of the parameter magnitudes, might be dangerous, potentially destroying the neighborhood structure even severely. We demonstrate the improved performance of the proposed Fisher SAM on several benchmark datasets/tasks.
Minyoung Kim 0001, Da Li 0001, Shell Xu Hu, Timothy M. Hospedales
ICML1
2021 Learning Disentangled Factors from Paired Data in Cross-Modal Retrieval: An Implicit Identifiable VAE Approach
abstract
We tackle the problem of learning the underlying disentangled latent factors that are shared between the paired bi-modal data in cross-modal retrieval. Typically the data in both modalities are complex, structured, and high dimensional (e.g., image and text), for which the conventional deep auto-encoding latent variable models such as the Variational Autoencoder (VAE) often suffer from difficulty of accurate decoder training or realistic synthesis. In this paper we propose a novel idea of the implicit decoder, which completely removes the ambient data decoding module from a latent variable model, via implicit encoder inversion that is achieved by Jacobian regularization of the low-dimensional embedding function. Motivated from the recent Identifiable-VAE (IVAE) model, we modify it to incorporate the query modality data as conditioning auxiliary input, which allows us to prove that the true parameters of the model can be identifiable under some regularity conditions. Tested on various datasets where the true factors are fully/partially available, our model is shown to identify the factors accurately, significantly outperforming conventional latent variable models.
Minyoung Kim 0001, Ricardo Guerrero, Vladimir Pavlovic 0001
ACM Multimedia1
2020 Ordinal-Content VAE: Isolating Ordinal-Valued Content Factors in Deep Latent Variable Models
abstract
In deep representational learning, it is often desired to isolate a particular factor (termed content) from other factors (referred to as style). What constitutes the content is typically specified by users through explicit labels in the data, while all unlabeled/unknown factors are regarded as style. Recently, it has been shown that such content-labeled data can be effectively exploited by modifying the deep latent factor models (e.g., VAE) such that the style and content are well separated in the latent representations. However, the approach assumes that the content factor is categorical-valued (e.g., subject ID in face image data, or digit class in the MNIST dataset). In certain situations, the content is ordinal-valued, that is, the values the content factor takes are ordered rather than categorical, making content-labeled VAEs, including the latent space they infer, suboptimal. In this paper, we propose a novel extension of VAE that imposes a partially ordered set (poset) structure in the content latent space, while simultaneously making it aligned with the ordinal content values. To this end, instead of the iid Gaussian latent prior adopted in prior approaches, we introduce a conditional Gaussian spacing prior model. This model admits a tractable joint Gaussian prior, but also effectively places negligible density values on the content latent configurations that violate the poset constraint. To evaluate this model, we consider two specific ordinal structured problems: estimating a subject's age in a face image and elucidating the calorie amount in a food meal image. We demonstrate significant improvements in content-style separation over previous non-ordinal approaches.
Minyoung Kim 0001, Vladimir Pavlovic 0001
ICDM1
2020 Recursive Inference for Variational Autoencoders
abstract
Inference networks of traditional Variational Autoencoders (VAEs) are typically amortized, resulting in relatively inaccurate posterior approximation compared to instance-wise variational optimization. Recent semi-amortized approaches were proposed to address this drawback; however, their iterative gradient update procedures can be computationally demanding. In this paper, we consider a different approach of building a mixture inference model. We propose a novel recursive mixture estimation algorithm for VAEs that iteratively augments the current mixture with new components so as to maximally reduce the divergence between the variational and the true posteriors. Using the functional gradient approach, we devise an intuitive learning criteria for selecting a new mixture component: the new component has to improve the data likelihood (lower bound) and, at the same time, be as divergent from the current mixture distribution as possible, thus increasing representational diversity. Although there have been similar approaches recently, termed boosted variational inference (BVI), our methods differ from BVI in several aspects, most notably that ours deal with recursive inference in VAEs in the form of amortized inference, while BVI is developed within the standard VI framework, leading to a non-amortized single optimization instance, inappropriate for VAEs. A crucial benefit of our approach is that the inference at test time needs a single feed-forward pass through the mixture inference network, making it significantly faster than the semi-amortized approaches. We show that our approach yields higher test data likelihood than the state-of-the-arts on several benchmark datasets.
Minyoung Kim 0001, Vladimir Pavlovic 0001
NeurIPS1
2019 Unsupervised Visual Domain Adaptation: A Deep Max-Margin Gaussian Process Approach
abstract
For unsupervised domain adaptation, the target domain error can be provably reduced by having a shared input representation that makes the source and target domains indistinguishable from each other. Very recently it has been shown that it is not only critical to match the marginal input distributions, but also align the output class distributions. The latter can be achieved by minimizing the maximum discrepancy of predictors. In this paper, we take this principle further by proposing a more systematic and effective way to achieve hypothesis consistency using Gaussian processes (GP). The GP allows us to induce a hypothesis space of classifiers from the posterior distribution of the latent random functions, turning the learning into a large-margin posterior separation problem, significantly easier to solve than previous approaches based on adversarial minimax optimization. We formulate a learning objective that effectively influences the posterior to minimize the maximum discrepancy. This is shown to be equivalent to maximizing margins and minimizing uncertainty of the class predictions in the target domain. Empirical results demonstrate that our approach leads to state-to-the-art performance superior to existing methods on several challenging benchmarks for domain adaptation.
Minyoung Kim 0001, Pritish Sahu, Behnam Gholami, Vladimir Pavlovic 0001
CVPR1
2019 Bayes-Factor-VAE: Hierarchical Bayesian Deep Auto-Encoder Models for Factor Disentanglement
abstract
We propose a family of novel hierarchical Bayesian deep auto-encoder models capable of identifying disentangled factors of variability in data. While many recent attempts at factor disentanglement have focused on sophisticated learning objectives within the VAE framework, their choice of a standard normal as the latent factor prior is both suboptimal and detrimental to performance. Our key observation is that the disentangled latent variables responsible for major sources of variability, the relevant factors, can be more appropriately modeled using long-tail distributions. The typical Gaussian priors are, on the other hand, better suited for modeling of nuisance factors. Motivated by this, we extend the VAE to a hierarchical Bayesian model by introducing hyper-priors on the variances of Gaussian latent priors, mimicking an infinite mixture, while maintaining tractable learning and inference of the traditional VAEs. This analysis signifies the importance of partitioning and treating in a different manner the latent dimensions corresponding to relevant factors and nuisances. Our proposed models, dubbed Bayes-Factor-VAEs, are shown to outperform existing methods both quantitatively and qualitatively in terms of latent disentanglement across several challenging benchmark tasks.
Minyoung Kim 0001, Yuting Wang 0004, Pritish Sahu, Vladimir Pavlovic 0001
ICCV1
2019 Efficient Deep Gaussian Process Models for Variable-Sized Inputs
abstract
Deep Gaussian processes (DGP) have appealing Bayesian properties, can handle variable-sized data, and learn deep features. Their limitation is that they do not scale well with the size of the data. Existing approaches address this using a deep random feature (DRF) expansion model, which makes inference tractable by approximating DGPs. However, DRF is not suitable for variable-sized input data such as trees, graphs, and sequences. We introduce the GP-DRF, a novel Bayesian model with an input layer of GPs, followed by DRF layers. The key advantage is that the combination of GP and DRF leads to a tractable model that can both handle a variable-sized input as well as learn deep long-range dependency structures of the data. We provide a novel efficient method to simultaneously infer the posterior of GP's latent vectors and infer the posterior of DRF's internal weights and random frequencies. Our experiments show that GP-DRF outperforms the standard GP model and DRF model across many datasets. Furthermore, they demonstrate that GP-DRF enables improved uncertainty quantification compared to GP and DRF alone, with respect to a Bhattacharyya distance assessment.
Issam H. Laradji, Mark Schmidt 0001, Vladimir Pavlovic 0001, Minyoung Kim 0001
IJCNN4
2019 Large-margin learning of Cox proportional hazard models for survival analysis
Minyoung Kim 0001
Appl. Intell.1
2019 Sparse large-margin nearest neighbor embedding via greedy dyad functional optimization
Minyoung Kim 0001
Appl. Intell.1
2018 Markov Modulated Gaussian Cox Processes for Semi-Stationary Intensity Modeling of Events Data
abstract
The Cox process is a flexible event model that can account for uncertainty of the intensity function in the Poisson process. However, previous approaches make strong assumptions in terms of time stationarity, potentially failing to generalize when the data do not conform to the assumed stationarity conditions. In this paper we bring up two most popular Cox models representing two extremes, and propose a novel semi-stationary Cox process model that can take benefits from both models. Our model has a set of Gaussian process latent functions governed by a latent stationary Markov process where we provide analytic derivations for the variational inference. Empirical evaluations on several synthetic and real-world events data including the football shot attempts and daily earthquakes, demonstrate that the proposed model is promising, can yield improved generalization performance over existing approaches.
Minyoung Kim 0001
ICML1
2018 Variational Inference for Gaussian Process Models for Survival Analysis
Minyoung Kim 0001, Vladimir Pavlovic 0001
UAI1
2018 A maximum-likelihood and moment-matching density estimator for crowd-sourcing label prediction
Minyoung Kim 0001
Appl. Intell.1
2018 Dynamic sparse coding for sparse time-series modeling via first-order smooth optimization
Minyoung Kim 0001
Appl. Intell.1
2017 Dual soft assignment clustering algorithm for human action video clustering
Minyoung Kim 0001
Comput. Vis. Image Underst.1
2017 Efficient histogram dictionary learning for text/image modeling and classification
Minyoung Kim 0001
Data Min. Knowl. Discov.1
2017 Mixtures of Conditional Random Fields for Improved Structured Output Prediction
abstract
The conditional random field (CRF) is a successful probabilistic model for structured output prediction problems. In this brief, we consider to enlarge the representational capacity of CRF via mixture modeling. The motivation is that a single CRF can perform well if the data conform to the statistical dependence assumption imposed by the CRF model structure, whereas it may potentially fail to model the data that come from multiple different sources or domains. For the conventional conditional likelihood objective, we derive the expectation-maximization algorithm in conjunction with the direct gradient ascent method for learning a CRF mixture with sequence or image-structured data. In addition, we provide alternative mixture learning algorithms that aim to maximize either the classification margin or the sitewise conditional likelihood, which were previously shown to outperform the conventional estimator for single CRF models in a variety of situations. We demonstrate the improved prediction accuracy of the proposed mixture learning algorithms on several important sequence labeling problems.
Minyoung Kim 0001
IEEE Trans. Neural Networks Learn. Syst.1
2016 Sparse inverse covariance learning of conditional Gaussian mixtures for multiple-output regression
Minyoung Kim 0001
Appl. Intell.1
2016 Model-induced term-weighting schemes for text classification
Hyun Kyung Kim, Minyoung Kim 0001
Appl. Intell.2
2016 Sparse conditional copula models for structured output regression
Minyoung Kim 0001
Pattern Recognit.1
2015 Sparse discriminative region selection algorithm for face recognition
Minyoung Kim 0001
Appl. Intell.1
2015 Multiple-concept feature generative models for multi-label image classification
Minyoung Kim 0001
Comput. Vis. Image Underst.1
2015 Greedy ensemble learning of structured predictors for sequence tagging
Minyoung Kim 0001
Neurocomputing1
2015 Greedy approaches to semi-supervised subspace learning
Minyoung Kim 0001
Pattern Recognit.1
2014 Conditional ordinal random fields for structured ordinal-valued label prediction
Minyoung Kim 0001
Data Min. Knowl. Discov.1
2014 Multiple instance learning via Gaussian processes
Minyoung Kim 0001, Fernando De la Torre
Data Min. Knowl. Discov.1
2014 Probabilistic Sequence Translation-Alignment Model for Time-Series Classification
abstract
We tackle the time-series classification problem using a novel probabilistic model that represents the conditional densities of the observed sequences being time-warped and transformed from an underlying base sequence. We call it probabilistic sequence translation-alignment model (PSTAM) since it aims to capture both feature alignment and mapping between sequences, analogous to translating one language into another in the field of machine translation. To deal with general time-series, we impose the time-monotonicity constraints on the hidden alignment variables in the model parameter space, where by marginalizing them out it allows effective learning of class-specific time-warping and feature transformation simultaneously. Our PSTAM, thus, naturally enjoys the advantages from two typical approaches widely used in sequence classification: 1) benefits from the alignment-based methods that aim to estimate distance measures between non-equal-length sequences via direct comparison of aligned features, and 2) merits of the model-based approaches that can effectively capture the class-specific patterns or trends. Furthermore, the low-dimensional modeling of the latent base sequence naturally provides a way to discover the intrinsic manifold structure possibly retained in the observed data, leading to an unsupervised manifold learning for sequence data. The benefits of the proposed approach are demonstrated on a comprehensive set of evaluations with both synthetic and real-world sequence data sets.
Minyoung Kim 0001
IEEE Trans. Knowl. Data Eng.1
2014 Efficient Kernel Sparse Coding Via First-Order Smooth Optimization
abstract
We consider the problem of dictionary learning and sparse coding, where the task is to find a concise set of basis vectors that accurately represent the observation data with only small numbers of active bases. Typically formulated as an L1-regularized least-squares problem, the problem incurs computational difficulty originating from the nondifferentiable objective. Recent approaches to sparse coding thus have mainly focused on acceleration of the learning algorithm. In this paper, we propose an even more efficient and scalable sparse coding algorithm based on the first-order smooth optimization technique. The algorithm finds the theoretically guaranteed optimal sparse codes of the epsilon-approximate problem in a series of optimization subproblems, where each subproblem admits analytic solution, hence very fast and scalable with large-scale data. We further extend it to nonlinear sparse coding using kernel trick by showing that the representer theorem holds for the kernel sparse coding problem. This allows us to apply dual optimization, which essentially results in the same linear sparse coding problem in dual variables, highly beneficial compared with the existing methods that suffer from local minima and restricted forms of kernel function. The efficiency of our algorithms is demonstrated for natural stimuli data sets and several image classification problems.
Minyoung Kim 0001
IEEE Trans. Neural Networks Learn. Syst.1
2013 Accelerated max-margin multiple kernel learning
Minyoung Kim 0001
Appl. Intell.1
2013 Semi-supervised learning of hidden conditional random fields for time-series classification
Minyoung Kim 0001
Neurocomputing1
2013 Conditional Alignment Random Fields for Multiple Motion Sequence Alignment
abstract
We consider the multiple time-series alignment problem, typically focusing on the task of synchronizing multiple motion videos of the same kind of human activity. Finding an optimal global alignment of multiple sequences is infeasible, while there have been several approximate solutions, including iterative pairwise warping algorithms and variants of hidden Markov models. In this paper, we propose a novel probabilistic model that represents the conditional densities of the latent target sequences which are aligned with the given observed sequences through the hidden alignment variables. By imposing certain constraints on the target sequences at the learning stage, we have a sensible model for multiple alignments that can be learned very efficiently by the EM algorithm. Compared to existing methods, our approach yields more accurate alignment while being more robust to local optima and initial configurations. We demonstrate its efficacy on both synthetic and real-world motion videos including facial emotions and human activities.
Minyoung Kim 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Correlation-based incremental visual tracking
Minyoung Kim 0001
Pattern Recognit.1
2012 Time-Series Dimensionality Reduction via Granger Causality
abstract
We deal with the problem of time-series prediction in a dyadic setup where the goal is to predict future values of the output sequence from the observed input sequence. Often the input time-series data is high-dimensional with potential noisy measurements included, which can make the prediction task difficult. In this paper, we propose a novel dimensionality reduction algorithm that can sparsely extract most salient and discriminative input features for output prediction. Our approach is based on the Granger causality, a famous statistical technique particularly in economics, where we aim to discover a low-dimensional subspace that preserves the causality between input and output. We demonstrate empirically the benefits of the proposed approaches on several datasets.
Minyoung Kim 0001
IEEE Signal Process. Lett.1
2011 Sequence classification via large margin hidden Markov models
Minyoung Kim 0001, Vladimir Pavlovic 0001
Data Min. Knowl. Discov.1
2011 Central Subspace Dimensionality Reduction Using Covariance Operators
abstract
We consider the task of dimensionality reduction informed by real-valued multivariate labels. The problem is often treated as Dimensionality Reduction for Regression (DRR), whose goal is to find a low-dimensional representation, the central subspace, of the input data that preserves the statistical correlation with the targets. A class of DRR methods exploits the notion of inverse regression (IR) to discover central subspaces. Whereas most existing IR techniques rely on explicit output space slicing, we propose a novel method called the Covariance Operator Inverse Regression (COIR) that generalizes IR to nonlinear input/output spaces without explicit target slicing. COIR's unique properties make DRR applicable to problem domains with high-dimensional output data corrupted by potentially significant amounts of noise. Unlike recent kernel dimensionality reduction methods that employ iterative nonconvex optimization, COIR yields a closed-form solution. We also establish the link between COIR, other DRR techniques, and popular supervised dimensionality reduction methods, including canonical correlation analysis and linear discriminant analysis. We then extend COIR to semi-supervised settings where many of the input points lack their labels. We demonstrate the benefits of COIR on several important regression problems in both fully supervised and semi-supervised settings.
Minyoung Kim 0001, Vladimir Pavlovic 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Discriminative semi-supervised learning of dynamical systems for motion estimation
Minyoung Kim 0001
Pattern Recognit.1
2011 Sequence Alignment by Regression Coding
abstract
In aligning two sequences, dynamic time warping (DTW) is the well-known dynamic programming algorithm. However, DTW can be sensitive to noise samples that may affect alignment of other relevant samples. In this article, we propose a novel approach to sequence alignment by treating it as a regression coding optimization problem, a task to predict one sequence from another. With some mild relaxation DTW can be seen as a special case of our approach while we provide more flexible and informative interpretation. Experimental results on both synthetic and real-world datasets show that our method can yield more accurate alignment than existing approaches.
Minyoung Kim 0001
IEEE Signal Process. Lett.1
2010 Structured Output Ordinal Regression for Dynamic Facial Emotion Intensity Prediction
Minyoung Kim 0001, Vladimir Pavlovic 0001
ECCV (3)1
2010 Local Minima Embedding
Minyoung Kim 0001, Fernando De la Torre
ICML1
2010 Gaussian Processes Multiple Instance Learning
Minyoung Kim 0001, Fernando De la Torre
ICML1
2010 Hidden Conditional Ordinal Random Fields for Sequence Classification
Minyoung Kim 0001, Vladimir Pavlovic 0001
ECML/PKDD (2)1
2010 Large margin cost-sensitive learning of conditional random fields
Minyoung Kim 0001
Pattern Recognit.1
2009 Discriminative Learning for Dynamic State Prediction
abstract
We consider the problem of predicting a sequence of real-valued multivariate states that are correlated by some unknown dynamics, from a given measurement sequence. Although dynamic systems such as the State-Space Models are popular probabilistic models for the problem, their joint modeling of states and observations, as well as the traditional generative learning by maximizing a joint likelihood may not be optimal for the ultimate prediction goal. In this paper, we suggest two novel discriminative approaches to the dynamic state prediction: 1) learning generative state-space models with discriminative objectives and 2) developing an undirected conditional model. These approaches are motivated by the success of recent discriminative approaches to the structured output classification in discrete-state domains, namely, discriminative training of Hidden Markov Models and Conditional Random Fields (CRFs). Extending CRFs to real multivariate state domains generally entails imposing density integrability constraints on the CRF parameter space, which can make the parameter learning difficult. We introduce an efficient convex learning algorithm to handle this task. Experiments on several problem domains, including human motion and robot-arm state estimation, indicate that the proposed approaches yield high prediction accuracy comparable to or better than state-of-the-art methods.
Minyoung Kim 0001, Vladimir Pavlovic 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Face tracking and recognition with visual constraints in real-world videos
abstract
We address the problem of tracking and recognizing faces in real-world, noisy videos. We track faces using a tracker that adaptively builds a target model reflecting changes in appearance, typical of a video setting. However, adaptive appearance trackers often suffer from drift, a gradual adaptation of the tracker to non-targets. To alleviate this problem, our tracker introduces visual constraints using a combination of generative and discriminative models in a particle filtering framework. The generative term conforms the particles to the space of generic face poses while the discriminative one ensures rejection of poorly aligned targets. This leads to a tracker that significantly improves robustness against abrupt appearance changes and occlusions, critical for the subsequent recognition phase. Identity of the tracked subject is established by fusing pose-discriminant and person-discriminant features over the duration of a video sequence. This leads to a robust video-based face recognizer with state-of-the-art recognition performance. We test the quality of tracking and face recognition on real-world noisy videos from YouTube as well as the standard Honda/UCSD database. Our approach produces successful face tracking results on over 80% of all videos without video or person-specific parameter tuning. The good tracking performance induces similarly high recognition rates: 100% on Honda/UCSD and over 70% on the YouTube set containing 35 celebrities in 1500 sequences.
Minyoung Kim 0001, Sanjiv Kumar, Vladimir Pavlovic 0001, Henry A. Rowley
CVPR1
2008 Dimensionality reduction using covariance operator inverse regression
abstract
We consider the task of dimensionality reduction for regression (DRR) whose goal is to find a low dimensional representation of input covariates, while preserving the statistical correlation with output targets. DRR is particularly suited for visualization of high dimensional data as well as the efficient regressor design with a reduced input dimension. In this paper we propose a novel nonlinear method for DRR that exploits the kernel Gram matrices of input and output. While most existing DRR techniques rely on the inverse regression, our approach removes the need for explicit slicing of the output space using covariance operators in RKHS. This unique property make DRR applicable to problem domains with high dimensional output data with potentially significant amounts of noise. Although recent kernel dimensionality reduction algorithms make use of RKHS covariance operators to quantify conditional dependency between the input and the targets via the dimension-reduced input, they are either limited to a transduction setting or linear input subspaces and restricted to non-closed-form solutions. In contrast, our approach provides a closed-form solution to the nonlinear basis functions on which any new input point can be easily projected. We demonstrate the benefits of the proposed method in a comprehensive set of evaluations on several important regression problems that arise in computer vision.
Minyoung Kim 0001, Vladimir Pavlovic 0001
CVPR1
2007 Discriminative Learning of Dynamical Systems for Motion Tracking
abstract
We introduce novel discriminative learning algorithms for dynamical systems. Models such as conditional random fields or maximum entropy Markov models outperform the generative hidden Markov models in sequence tagging problems in discrete domains. However, continuous state domains introduce a set of constraints that can prevent direct application of these traditional models. Instead, we suggest to learn generative dynamic models with discriminative cost functionals. For linear dynamical systems, the proposed methods provide significantly lower prediction error than the standard maximum likelihood estimator, often comparable to nonlinear models. As a result, the models with lower representational capacity but computationally more tractable than nonlinear models can be used for accurate and efficient state estimation. We evaluate the generalization performance of our methods on the 3D human pose tracking problem from monocular videos. The experiments indicate that the discriminative learning can lead to improved accuracy of pose estimation with no increase in computational cost of tracking.
Minyoung Kim 0001, Vladimir Pavlovic 0001
CVPR1
2007 Conditional State Space Models for Discriminative Motion Estimation
abstract
We consider the problem of predicting a sequence of real-valued multivariate states from a given measurement sequence. Its typical application in computer vision is the task of motion estimation. State Space Models are widely used generative probabilistic models for the problem. Instead of jointly modeling states and measurements, we propose a novel discriminative undirected graphical model which conditions the states on the measurements while exploiting the sequential structure of the problem. The major benefits of this approach are: (1) It focuses on the ultimate prediction task while avoiding probably unnecessary effort in modeling the measurement density, (2) It relaxes generative models' assumption that the measurements are independent given the states, and (3) The proposed inference algorithm takes linear time in the measurement dimension as opposed to the cubic time for Kalman filtering, which allows us to incorporate large numbers of measurement features. We show that the parameter learning can be cast as an instance of convex optimization. We also provide efficient convex optimization methods based on theorems from linear algebra. The performance of the proposed model is evaluated on both synthetic data and the human body pose estimation from silhouette videos.
Minyoung Kim 0001, Vladimir Pavlovic 0001
ICCV1
2007 A recursive method for discriminative mixture learning
abstract
We consider the problem of learning density mixture models for classification. Traditional learning of mixtures for density estimation focuses on models that correctly represent the density at all points in the sample space. Discriminative learning, on the other hand, aims at representing the density at the decision boundary. We introduce a novel discriminative learning method for mixtures of generative models. Unlike traditional discriminative learning methods that often resort to computationally demanding gradient search optimization, the proposed method is highly efficient as it reduces to generative learning of individual mixture components on weighted data. Hence it is particularly suited to domains with complex component models, such as hidden Markov models or Bayesian networks in general, that are usually too complex for effective gradient search. We demonstrate the benefits of the proposed method in a comprehensive set of evaluations on time-series sequence classification problems.
Minyoung Kim 0001, Vladimir Pavlovic 0001
ICML1
2006 Discriminative Learning of Mixture of Bayesian Network Classifiers for Sequence Classification
abstract
A mixture of Bayesian Network Classifiers(BNC) has a potential to yield superior classification and generative performance to a single BNC model. We introduce novel discriminative learning methods for mixtures of BNCs. Unlike a single BNC model where the discriminative learning resorts to a gradient search, we can exploit the properties of a mixture to alleviate the complex learning task. The proposed method adds mixture components recursively via functional gradient boosting while maximizing the conditional likelihood. This method is highly efficient as it reduces to generative learning of a base BNC model on weighed data. The proposed approach is particularly suited to sequence classification problems where the kernels in the base model are usually too complex for effective gradient search. We demonstrate the improved classification performance of the proposed methods in an extensive set of evaluations on time-series sequence data, including human motion classification problems.
Minyoung Kim 0001, Vladimir Pavlovic 0001
CVPR (1)1