EDBT 2026 Demo / reviewers in the wild / expert
Byung-Jun Lee 0001
dblp:130/1678-1
· DBLP profile ↗
27ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-0684-607XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 5 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired DataabstractContrastive Language–Image Pre-training (CLIP) has demonstrated strong generalization across a wide range of visual tasks by leveraging large-scale English–image pairs. However, its extension to low-resource languages remains limited due to the scarcity of high-quality multilingual image–text data. Existing multilingual vision–language models exhibit consistently low retrieval performance in underrepresented languages—including Czech, Finnish, Croatian, Hungarian, Romanian—on the Crossmodal-3600 (XM3600) benchmark. To address this, we propose a lightweight and data-efficient framework for multilingual vision–language alignment. Our approach requires no image–text pairs or text-text pairs and freezes both the pretrained image encoder and multilingual text encoder during training. Only a compact 1.7M-parameter projection module is trained, using a contrastive loss over English representations as semantic anchors. This minimal training setup enables robust multilingual alignment even for languages with limited supervision. Extensive evaluation across multiple multilingual retrieval benchmarks confirms the effectiveness of our method, showing significant gains in five underrepresented languages where existing models typically underperform. These findings highlight the effectiveness of our pivot-based, parameter-efficient alignment strategy for inclusive multimodal learning. Dahyun Chung, Donghyun Shin, Yujin Sung, Seunggi Moon, Jinwoo Jeon, Byung-Jun Lee 0001 |
AAAI | 6 |
| 2026 | SGT: Securing Open-Source LLMs Against Malicious Fine-tuning via Safety Guidance TriggerabstractOpen-weight large language models (LLMs) enable extensive customization but remain susceptible to post-release misuse via malicious fine-tuning.While existing defenses attempt to constrain parameter-space dynamics or mitigate harmful internal representations, malicious fine-tuning continues to erode these safeguards leaving the development of fundamental, persistent defenses for open-weight models an unresolved challenge.In this paper, we characterize a safety region for open-weight LLMs and propose Safety Guidance Trigger (SGT), a framework that preserves alignment by guiding finetuning toward the safety manifold.It has two stages: (1) optimizing a safety trigger to steer the base model outputs toward safe responses and ( 2) training the open-weight model to align its internal features with trigger-induced safety representations.We demonstrate that SGT substantially improves robustness against malicious fine-tuning, forcing adversaries to significantly increase data budgets to bypass safeguards.Our analysis further confirms that this approach anchors model representations within a safety region that remains resilient under adversarial attacks. Sunguk Shin 0001, Fangzhao Wu, Byung-Jun Lee 0001, Meeyoung Cha, Sungwon Park 0001 |
ACL (1) | 3 |
| 2026 | Learning to Process Relational Information for Cloud Computing Task Scheduling
Seong-Hyun Hong, Myunsoo Kim, Byung-Jun Lee 0001 |
Mach. Learn. | 3 |
| 2026 | Neural MCTS with LLM Guidance for Effective Program Synthesis on Abstraction and Reasoning Corpus
Jinwoo Jeon, Seongwoong Shim, Sejin Kim 0002, Sundong Kim, Byung-Jun Lee 0001 |
Mach. Learn. | 5 |
| 2025 | K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in KoreanabstractMinkyeong Jeon, Hyemin Jeong, Yerang Kim, Jiyoung Kim, Jae Hyeon Cho, Byung-Jun Lee. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Minkyeong Jeon, Hyemin Jeong, Yerang Kim, Jae Hyeon Cho, Byung-Jun Lee 0001 |
ACL (1) | 6 |
| 2025 | Adaptive Non-Uniform Timestep Sampling for Accelerating Diffusion Model TrainingabstractAs a highly expressive generative model, diffusion models have demonstrated exceptional success across various domains, including image generation, natural language processing, and combinatorial optimization. However, as data distributions grow more complex, training these models to convergence becomes increasingly computationally intensive. While diffusion models are typically trained using uniform timestep sampling, our research shows that the variance in stochastic gradients varies significantly across timesteps, with high-variance timesteps becoming bottlenecks that hinder faster convergence. To address this issue, we introduce a non-uniform timestep sampling method that prioritizes these more critical timesteps. Our method tracks the impact of gradient updates on the objective for each timestep, adaptively selecting those most likely to minimize the objective effectively. Experimental results demonstrate that this approach not only accelerates the training process, but also leads to improved performance at convergence. Furthermore, our method shows robust performance across various datasets, scheduling strategies, and diffusion architectures, outperforming previously proposed timestep sampling and weighting heuristics that lack this degree of robustness. Myunsoo Kim, Donghyeon Ki, Seongwoong Shim, Byung-Jun Lee 0001 |
CVPR | 4 |
| 2025 | Iterative Prompt Refinement for Safer Text-to-Image GenerationabstractText-to-Image (T2I) models have made remarkable progress in generating images from text prompts, but their output quality and safety still depend heavily on how prompts are phrased.Existing safety methods typically refine prompts using large language models (LLMs), but they overlook the images produced, which can result in unsafe outputs or unnecessary changes to already safe prompts.To address this, we propose an iterative prompt refinement algorithm that uses Vision Language Models (VLMs) to analyze both the input prompts and the generated images.By leveraging visual feedback, our method refines prompts more effectively, improving safety while maintaining user intent and reliability comparable to existing LLM-based approaches.Additionally, we introduce a new dataset labeled with both textual and visual safety signals using off-the-shelf multi-modal LLM, enabling supervised fine-tuning.Experimental results demonstrate that our approach produces safer outputs without compromising alignment with user intent, offering a practical solution for generating safer T2I content.Our code is available at https://github.com/ku-dmlab/IPR.WARNING: This paper contains examples of harmful or inappropriate images generated by models. Jinwoo Jeon, JunHyeok Oh, Hayeong Lee, Byung-Jun Lee 0001 |
EMNLP | 4 |
| 2025 | NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic DemonstrationsabstractIntelligent agents are able to make decisions based on different levels of granularity and duration. Recent advances in skill learning enabled the agent to solve complex, long-horizon tasks by effectively guiding the agent in choosing appropriate skills. However, the practice of using fixed-length skills can easily result in skipping valuable decision points, which ultimately limits the potential for further exploration and faster policy learning.
In this work, we propose to learn a simple and effective termination condition that identifies decision points through a state-action novelty module that leverages agent experience data.
Our approach, Novelty-based Decision Point Identification (NBDI), outperforms previous baselines in complex, long-horizon tasks, and remains effective even in the presence of significant variations in the environment configurations of downstream tasks, highlighting the importance of decision point identification in skill learning. Myunsoo Kim, Hayeong Lee, Seongwoong Shim, JunHo Seo 0001, Byung-Jun Lee 0001 |
ICML | 5 |
| 2025 | Prior-Guided Diffusion Planning for Offline Reinforcement LearningabstractDiffusion models have recently gained prominence in offline reinforcement learning due to their ability to effectively learn high-performing, generalizable policies from static datasets. Diffusion-based planners facilitate long-horizon decision-making by generating high-quality trajectories through iterative denoising, guided by return-maximizing objectives. However, existing guided sampling strategies such as Classifier Guidance, Classifier-Free Guidance, and Monte Carlo Sample Selection either produce suboptimal multi-modal actions, struggle with distributional drift, or incur prohibitive inference-time costs. To address these challenges, we propose \textbf{\textit{Prior Guidance}} (PG), a novel guided sampling framework that replaces the standard Gaussian prior of a behavior-cloned diffusion model with a learnable distribution, optimized via a behavior-regularized objective. PG directly generates high-value trajectories without costly reward optimization of the diffusion model itself, and eliminates the need to sample multiple candidates at inference for sample selection. We present an efficient training strategy that applies behavior regularization in latent space, and empirically demonstrate that PG outperforms state-of-the-art diffusion policies and planners across diverse long-horizon offline RL benchmarks. Our code is available at https://github.com/ku-dmlab/PG. Donghyeon Ki, JunHyeok Oh, Seongwoong Shim, Byung-Jun Lee 0001 |
NeurIPS | 4 |
| 2025 | FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement LearningabstractMulti-objective reinforcement learning (MORL) aims to optimize policies in the presence of conflicting objectives, where linear scalarization is commonly used to reduce vector-valued returns into scalar signals. While effective for certain preferences, this approach cannot capture fairness-oriented goals such as Nash social welfare or max-min fairness, which require nonlinear and non-additive trade-offs. Although several online algorithms have been proposed for specific fairness objectives, a unified approach for optimizing nonlinear welfare criteria in the offline setting—where learning must proceed from a fixed dataset—remains unexplored. In this work, we present FairDICE, the first offline MORL framework that directly optimizes nonlinear welfare objective. FairDICE leverages distribution correction estimation to jointly account for welfare maximization and distributional regularization, enabling stable and sample-efficient learning without requiring explicit preference weights or exhaustive weight search. Across multiple offline benchmarks, FairDICE demonstrates strong fairness-aware performance compared to existing baselines. Woosung Kim, Jongmin Lee 0004, Byung-Jun Lee 0001 |
NeurIPS | 4 |
| 2024 | Relaxed Stationary Distribution Correction Estimation for Improved Offline Policy OptimizationabstractOne of the major challenges of offline reinforcement learning (RL) is dealing with distribution shifts that stem from the mismatch between the trained policy and the data collection policy. Stationary distribution correction estimation algorithms (DICE) have addressed this issue by regularizing the policy optimization with f-divergence between the state-action visitation distributions of the data collection policy and the optimized policy. While such regularization naturally integrates to derive an objective to get optimal state-action visitation, such an implicit policy optimization framework has shown limited performance in practice. We observe that the reduced performance is attributed to the biased estimate and the properties of conjugate functions of f-divergence regularization. In this paper, we improve the regularized implicit policy optimization framework by relieving the bias and reshaping the conjugate function by relaxing the constraints. We show that the relaxation adjusts the degree of involvement of the sub-optimal samples in optimization, and we derive a new offline RL algorithm that benefits from the relaxed framework, improving from a previous implicit policy optimization algorithm by a large margin. Woosung Kim, Donghyeon Ki, Byung-Jun Lee 0001 |
AAAI | 3 |
| 2024 | ROIDICE: Offline Return on Investment Maximization for Efficient Decision MakingabstractIn this paper, we propose a novel policy optimization framework that maximizes Return on Investment (ROI) of a policy using a fixed dataset within a Markov Decision Process (MDP) equipped with a cost function. ROI, defined as the ratio between the return and the accumulated cost of a policy, serves as a measure of efficiency of the policy. Despite the importance of maximizing ROI in various applications, it remains a challenging problem due to its nature as a ratio of two long-term values: return and accumulated cost. To address this, we formulate the ROI maximizing reinforcement learning problem as a linear fractional programming. We then incorporate the stationary distribution correction (DICE) framework to develop a practical offline ROI maximization algorithm.
Our proposed algorithm, ROIDICE, yields an efficient policy that offers a superior trade-off between return and accumulated cost compared to policies trained using existing frameworks. Woosung Kim, Hayeong Lee, Jongmin Lee 0004, Byung-Jun Lee 0001 |
NeurIPS | 4 |
| 2024 | Mitigating Covariate Shift in Behavioral Cloning via Robust Stationary Distribution CorrectionabstractWe consider offline imitation learning (IL), which aims to train an agent to imitate from the dataset of expert demonstrations without online interaction with the environment. Behavioral Cloning (BC) has been a simple yet effective approach to offline IL, but it is also well-known to be vulnerable to the covariate shift resulting from the mismatch between the state distributions induced by the learned policy and the expert policy. Moreover, as often occurs in practice, when expert datasets are collected from an arbitrary state distribution instead of a stationary one, these shifts become more pronounced, potentially leading to substantial failures in existing IL methods. Specifically, we focus on covariate shift resulting from arbitrary state data distributions, such as biased data collection or incomplete trajectories, rather than shifts induced by changes in dynamics or noisy expert actions. In this paper, to mitigate the effect of the covariate shifts in BC, we propose DrilDICE, which utilizes a distributionally robust BC objective by employing a stationary distribution correction ratio estimation (DICE) to derive a feasible solution. We evaluate the effectiveness of our method through an extensive set of experiments covering diverse covariate shift scenarios. The results demonstrate the efficacy of the proposed approach in improving the robustness against the shifts, outperforming existing offline IL methods in such scenarios. Seokin Seo, Byung-Jun Lee 0001, Jongmin Lee 0004, HyeongJoo Hwang, Hongseok Yang, Kee-Eung Kim |
NeurIPS | 2 |
| 2024 | MARS: Multiagent Reinforcement Learning for Spatial - Spectral and Temporal Feature Selection in EEG-Based BCIabstractIn recent years, deep learning methods have shown promising capabilities for extracting informative and discriminative features from electroencephalography (EEG) data. However, several studies have reported that the feature selection process followed by feature extraction can be beneficial to achieve further performance improvement. Even though a recent work achieved promising results by using the single-agent reinforcement learning (RL)-based framework to select task-relevant features in the temporal domain, it still failed to consider other significant features in the spatial–spectral domain. To overcome such limitations, we propose a cooperative multiagent RL-based framework (MARS) that performs feature selection in both the spatial–spectral and temporal domains simultaneously for a motor imagery (MI)-EEG classification task. In this framework, we enable our RL agents to collaborate with each other as a team to solve a complex multiobjective feature selection problem. Furthermore, we adopt a counterfactual advantage function to overcome the free-rider problem, which is associated with the credit assignment issue in multiagent cases. To assess the MARS framework, we conduct extensive experiments with two public MI datasets under subject-dependent and subject-independent scenarios and we apply the MARS to different backbone networks. The experimental results demonstrate that our MARS outperforms other competing methods in terms of mean accuracy and achieves statistically significant improvements. Dong-Hee Shin, Young-Han Son, Junmo Kim 0001, Hee-Jun Ahn, JunHo Seo 0001, Chang-Hoon Ji, Ji-Wung Han, Byung-Jun Lee 0001, Dong-Ok Won, Tae-Eui Kam |
IEEE Trans. Syst. Man Cybern. Syst. | 8 |
| 2023 | Quantifying Information of Tokens for Simple and Flexible Simultaneous Machine TranslationabstractSimultaneous Translation (ST) involves translating with only partial source inputs instead of the entire source inputs, a process that can potentially result in translation quality degradation.Previous approaches to balancing translation quality and latency have demonstrated that it is more efficient and effective to leverage an offline model with a reasonable policy.However, using an offline model also leads to a distribution shift since it is not trained with partial source inputs, and it can be improved by training an additional module that informs us when to translate.In this paper, we propose an Information Quantifier (IQ) that models source and target information to determine whether the offline model has sufficient information for translation, trained with oracle action sequences generated from the offline model.IQ, by quantifying information, helps in formulating a suitable policy for Simultaneous Translation that better generalizes and also allows us to control the trade-off between quality and latency naturally.Experiments on various language pairs show that our proposed model outperforms baselines.1 Minkyung Park, Byung-Jun Lee 0001 |
CoNLL | 3 |
| 2023 | Improving Neural Machine Translation with Offline EvaluationsabstractMin-Kyung Park, Byung-Jun Lee. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Minkyung Park, Byung-Jun Lee 0001 |
IJCNLP (1) | 2 |
| 2022 | Local Metric Learning for Off-Policy Evaluation in Contextual Bandits with Continuous ActionsabstractWe consider local kernel metric learning for off-policy evaluation (OPE) of deterministic policies in contextual bandits with continuous action spaces. Our work is motivated by practical scenarios where the target policy needs to be deterministic due to domain requirements, such as prescription of treatment dosage and duration in medicine. Although importance sampling (IS) provides a basic principle for OPE, it is ill-posed for the deterministic target policy with continuous actions. Our main idea is to relax the target policy and pose the problem as kernel-based estimation, where we learn the kernel metric in order to minimize the overall mean squared error (MSE). We present an analytic solution for the optimal metric, based on the analysis of bias and variance. Whereas prior work has been limited to scalar action spaces or kernel bandwidth selection, our work takes a step further being capable of vector action spaces and metric optimization. We show that our estimator is consistent, and significantly reduces the MSE compared to baseline OPE methods through experiments on various domains. Haanvid Lee, Jongmin Lee 0004, Yunseon Choi, Wonseok Jeon, Byung-Jun Lee 0001, Yung-Kyun Noh, Kee-Eung Kim |
NeurIPS | 5 |
| 2021 | Representation Balancing Offline Model-based Reinforcement Learning
Byung-Jun Lee 0001, Jongmin Lee 0004, Kee-Eung Kim |
ICLR | 1 |
| 2021 | Winning the L2RPN Challenge: Power Grid Management via Semi-Markov Afterstate Actor-Critic
Deunsol Yoon, Sunghoon Hong, Byung-Jun Lee 0001, Kee-Eung Kim |
ICLR | 3 |
| 2021 | OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationabstractWe consider the offline reinforcement learning (RL) setting where the agent aims to optimize the policy solely from the data without further environment interactions. In offline RL, the distributional shift becomes the primary source of difficulty, which arises from the deviation of the target policy being optimized from the behavior policy used for data collection. This typically causes overestimation of action values, which poses severe problems for model-free algorithms that use bootstrapping. To mitigate the problem, prior offline RL algorithms often used sophisticated techniques that encourage underestimation of action values, which introduces an additional set of hyperparameters that need to be tuned properly. In this paper, we present an offline RL algorithm that prevents overestimation in a more principled way. Our algorithm, OptiDICE, directly estimates the stationary distribution corrections of the optimal policy and does not rely on policy-gradients, unlike previous offline RL algorithms. Using an extensive set of benchmark datasets for offline RL, we show that OptiDICE performs competitively with the state-of-the-art methods. Jongmin Lee 0004, Wonseok Jeon, Byung-Jun Lee 0001, Joelle Pineau, Kee-Eung Kim |
ICML | 3 |
| 2020 | Residual Neural ProcessesabstractA Neural Process (NP) is a map from a set of observed input-output pairs to a predictive distribution over functions, which is designed to mimic other stochastic processes' inference mechanisms. NPs are shown to work effectively in tasks that require complex distributions, where traditional stochastic processes struggle, e.g. image completion tasks. This paper concerns the practical capacity of set function approximators despite their universality. By delving deeper into the relationship between an NP and a Bayesian last layer (BLL), it is possible to see that NPs may struggle in simple examples, which other stochastic processes can easily solve. In this paper, we propose a simple yet effective remedy; the Residual Neural Process (RNP) that leverages traditional BLL for faster training and better prediction. We demonstrate that the RNP shows faster convergence and better performance, both qualitatively and quantitatively. Byung-Jun Lee 0001, Seunghoon Hong, Kee-Eung Kim |
AAAI | 1 |
| 2020 | Batch Reinforcement Learning with Hyperparameter GradientsabstractWe consider the batch reinforcement learning problem where the agent needs to learn only from a fixed batch of data, without further interaction with the environment. In such a scenario, we want to prevent the optimized policy from deviating too much from the data collection policy since the estimation becomes highly unstable otherwise due to the off-policy nature of the problem. However, imposing this requirement too strongly will result in a policy that merely follows the data collection policy. Unlike prior work where this trade-off is controlled by hand-tuned hyperparameters, we propose a novel batch reinforcement learning approach, batch optimization of policy and hyperparameter (BOPAH), that uses a gradient-based optimization of the hyperparameter using held-out data. We show that BOPAH outperforms other batch reinforcement learning algorithms in tabular and continuous control tasks, by finding a good balance to the trade-off between adhering to the data collection policy and pursuing the possible policy improvement. Byung-Jun Lee 0001, Jongmin Lee 0004, Peter Vrancx, Kee-Eung Kim |
ICML | 1 |
| 2020 | Reinforcement Learning for Control with Multiple FrequenciesabstractMany real-world sequential decision problems involve multiple action variables whose control frequencies are different, such that actions take their effects at different periods. While these problems can be formulated with the notion of multiple action persistences in factored-action MDP (FA-MDP), it is non-trivial to solve them efficiently since an action-persistent policy constructed from a stationary policy can be arbitrarily suboptimal, rendering solution methods for the standard FA-MDPs hardly applicable. In this paper, we formalize the problem of multiple control frequencies in RL and provide its efficient solution method. Our proposed method, Action-Persistent Policy Iteration (AP-PI), provides a theoretical guarantee on the convergence to an optimal solution while incurring only a factor of $|A|$ increase in time complexity during policy improvement step, compared to the standard policy iteration for FA-MDPs. Extending this result, we present Action-Persistent Actor-Critic (AP-AC), a scalable RL algorithm for high-dimensional control tasks. In the experiments, we demonstrate that AP-AC significantly outperforms the baselines on several continuous control tasks and a traffic control simulation, which highlights the effectiveness of our method that directly optimizes the periodic non-stationary policy for tasks with multiple control frequencies. Jongmin Lee 0004, Byung-Jun Lee 0001, Kee-Eung Kim |
NeurIPS | 2 |
| 2018 | Cross-Language Neural Dialog State Tracker for Large Ontologies Using Hierarchical AttentionabstractDialog state tracking, which refers to identifying the user intent from utterances, is one of the most important tasks in dialog management. In this paper, we present our dialog state tracker developed for the fifth dialog state tracking challenge, which focused on cross-language adaptation using a very scarce machine-translated training data when compared to the size of the ontology. Our dialog state tracker is based on the bi-directional long short-term memory network with a hierarchical attention mechanism in order to spot important words in user utterances. The user intent is predicted by finding the closest keyword in the ontology to the attention-weighted word vector. With the suggested methodology, our tracker can overcome various difficulties due to the scarce training data that existing machine learning-based trackers had, such as predicting user intents they have not seen before. We show that our tracker outperforms other trackers submitted to the challenge with respect to most of the performance measures. Youngsoo Jang, Jiyeon Ham, Byung-Jun Lee 0001, Kee-Eung Kim |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Hierarchically-partitioned Gaussian Process ApproximationabstractThe Gaussian process (GP) is a simple yet powerful probabilistic framework for various machine learning tasks. However, exact algorithms for learning and prediction are prohibitive to be applied to large datasets due to inherent computational complexity. To overcome this main limitation, various techniques have been proposed, and in particular, local GP algorithms that scales “truly linearly” with respect to the dataset size. In this paper, we introduce a hierarchical model based on local GP for large-scale datasets, which stacks inducing points over inducing points in layers. By using different kernels in each layer, the overall model becomes multi-scale and is able to capture both long- and short-range dependencies. We demonstrate the effectiveness of our model by speed-accuracy performance on challenging real-world datasets. Byung-Jun Lee 0001, Jongmin Lee 0004, Kee-Eung Kim |
AISTATS | 1 |
| 2016 | Neural dialog state tracker for large ontologies by attention mechanismabstractThis paper presents a dialog state tracker submitted to Dialog State Tracking Challenge 5 (DSTC 5) with details. To tackle the challenging cross-language human-human dialog state tracking task with limited training data, we propose a tracker that focuses on words with meaningful context based on attention mechanism and bi-directional long short term memory (LSTM). The vocabulary including a plenty of proper nouns is vectorized with a sufficient amount of related texts crawled from web to learn a good embedding for words not existent in training dialogs. Despite its simplicity, our proposed tracker succeeded to achieve high accuracy without sophisticated pre- and post-processing. Youngsoo Jang, Jiyeon Ham, Byung-Jun Lee 0001, Youngjae Chang 0001, Kee-Eung Kim |
SLT | 3 |
| 2014 | Optimizing Generative Dialog State Tracker via Cascading Gradient DescentabstractFor robust spoken dialog management, various dialog state tracking methods have been proposed. Although discriminative models are gaining popularity due to their superior performance, generative models based on the Partially Observable Markov Decision Process model still remain at-tractive since they provide an integrated framework for dialog state tracking and dialog policy optimization. Although a straightforward way to fit a generative model is to independently train the com-ponent probability models, we present a gradient descent algorithm that simultane-ously train all the component models. We show that the resulting tracker performs competitively with other top-performing trackers that participated in DSTC2. 1 Byung-Jun Lee 0001, Woosang Lim, Daejoong Kim, Kee-Eung Kim |
SIGDIAL Conference | 1 |