VLDB 2026 Research / reviewers in the wild / expert
Il-Chul Moon
dblp:97/4109
· DBLP profile ↗
64ranked-venue papers
3as first author
33since 2021 · last 2025
0000-0002-1798-1306ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 since 2021Human-computer interaction and ubiquitous computing · 12 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diffusion Bridge AutoEncoders for Unsupervised Representation LearningabstractDiffusion-based representation learning has achieved substantial attention due to its promising capabilities in latent representation and sample generation. Recent studies have employed an auxiliary encoder to identify a corresponding representation from data and to adjust the dimensionality of a latent variable $\mathbf{z}$. Meanwhile, this auxiliary structure invokes an *information split problem*; the information of each data instance $\mathbf{x}_0$ is divided into diffusion endpoint $\mathbf{x}_T$ and encoded $\mathbf{z}$ because there exist two inference paths starting from the data. The latent variable modeled by diffusion endpoint $\mathbf{x}_T$ has some disadvantages. The diffusion endpoint $\mathbf{x}_T$ is computationally expensive to obtain and inflexible in dimensionality. To address this problem, we introduce Diffusion Bridge AuteEncoders (DBAE), which enables $\mathbf{z}$-dependent endpoint $\mathbf{x}_T$ inference through a feed-forward architecture. This structure creates an information bottleneck at $\mathbf{z}$, so $\mathbf{x}_T$ becomes dependent on $\mathbf{z}$ in its generation. This results in $\mathbf{z}$ holding the full information of data. We propose an objective function for DBAE to enable both reconstruction and generative modeling, with their theoretical justification. Empirical evidence supports the effectiveness of the intended design in DBAE, which notably enhances downstream inference quality, reconstruction, and disentanglement. Additionally, DBAE generates high-fidelity samples in the unconditional generation. Our code is
available at https://github.com/aailab-kaist/DBAE. Yeongmin Kim, Kwanghyeon Lee, Minsang Park, Byeonghu Na, Il-Chul Moon |
ICLR | 5 |
| 2025 | Trajectory-Class-Aware Multi-Agent Reinforcement LearningabstractIn the context of multi-agent reinforcement learning, *generalization* is a challenge to solve various tasks that may require different joint policies or coordination without relying on policies specialized for each task. We refer to this type of problem as a *multi-task*, and we train agents to be versatile in this multi-task setting through a single training process. To address this challenge, we introduce TRajectory-class-Aware Multi-Agent reinforcement learning (TRAMA). In TRAMA, agents recognize a task type by identifying the class of trajectories they are experiencing through partial observations, and the agents use this trajectory awareness or prediction as additional information for action policy. To this end, we introduce three primary objectives in TRAMA: (a) constructing a quantized latent space to generate trajectory embeddings that reflect key similarities among them; (b) conducting trajectory clustering using these trajectory embeddings; and (c) building a trajectory-class-aware policy. Specifically for (c), we introduce a trajectory-class predictor that performs agent-wise predictions on the trajectory class; and we design a trajectory-class representation model for each trajectory class. Each agent takes actions based on this trajectory-class representation along with its partial observation for task-aware execution. The proposed method is evaluated on various tasks, including multi-task problems built upon StarCraft II. Empirical results show further performance improvements over state-of-the-art baselines. Hyungho Na, Kwanghyeon Lee, Il-Chul Moon |
ICLR | 4 |
| 2025 | Distilling Dataset into Neural FieldabstractUtilizing a large-scale dataset is essential for training high-performance deep learning models, but it also comes with substantial computation and storage costs. To overcome these challenges, dataset distillation has emerged as a promising solution by compressing the large-scale dataset into a smaller synthetic dataset that retains the essential information needed for training. This paper proposes a novel parameterization framework for dataset distillation, coined Distilling Dataset into Neural Field (DDiF), which leverages the neural field to store the necessary information of the large-scale dataset. Due to the unique nature of the neural field, which takes coordinates as input and output quantity, DDiF effectively preserves the information and easily generates various shapes of data. We theoretically confirm that DDiF exhibits greater expressiveness than some previous literature when the utilized budget for a single synthetic instance is the same. Through extensive experiments, we demonstrate that DDiF achieves superior performance on several benchmark datasets, extending beyond the image domain to include video, audio, and 3D voxel. We release the code at \url{https://github.com/aailab-kaist/DDiF}. Donghyeok Shin, HeeSun Bae, Gyuwon Sim, Wanmo Kang, Il-Chul Moon |
ICLR | 5 |
| 2025 | Preference Optimization by Estimating the Ratio of the Data DistributionabstractDirect preference optimization (DPO) is widely used as a simple and stable method for aligning large language models (LLMs) with human preferences. This paper investigates a generalized DPO loss that enables a policy model to match the target policy from a likelihood ratio estimation perspective. The ratio of the target policy provides a unique identification of the policy distribution without relying on reward models or partition functions. This allows the generalized loss to retain both simplicity and theoretical guarantees, which prior work such as $f$-PO fails to achieve simultaneously. We propose \textit{Bregman preference optimization} (BPO), a generalized framework for ratio matching that provides a family of objective functions achieving target policy optimality. BPO subsumes DPO as a special case and offers tractable forms for all instances, allowing implementation with a few lines of code. We further develop scaled Basu's power divergence (SBA), a gradient scaling method that can be used for BPO instances. The BPO framework complements other DPO variants and is applicable to target policies defined by these variants. In experiments, unlike other probabilistic loss extensions such as $f$-DPO or $f$-PO, which exhibits a trade-off between generation fidelity and diversity, instances of BPO improve both win rate and entropy compared with DPO. When applied to Llama-3-8B-Instruct, BPO achieves state-of-the-art performance among Llama-3-8B backbones, with a 55.9\% length-controlled win rate on AlpacaEval2. Project page: https://github.com/aailab-kaist/BPO. Yeongmin Kim, HeeSun Bae, Byeonghu Na, Il-Chul Moon |
NeurIPS | 4 |
| 2025 | Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion ModelsabstractText-to-image models have recently made significant advances in generating realistic and semantically coherent images, driven by advanced diffusion models and large-scale web-crawled datasets. However, these datasets often contain inappropriate or biased content, raising concerns about the generation of harmful outputs when provided with malicious text prompts. We propose Safe Text embedding Guidance (STG), a training-free approach to improve the safety of diffusion models by guiding the text embeddings during sampling. STG adjusts the text embeddings based on a safety function evaluated on the expected final denoised image, allowing the model to generate safer outputs without additional training. Theoretically, we show that STG aligns the underlying model distribution with safety constraints, thereby achieving safer outputs while minimally affecting generation quality. Experiments on various safety scenarios, including nudity, violence, and artist-style removal, show that STG consistently outperforms both training-based and training-free baselines in removing unsafe content while preserving the core semantic intent of input prompts. Our code is available at https://github.com/aailab-kaist/STG. Byeonghu Na, Mina Kang, Jiseok Kwak, Minsang Park, Jiwoo Shin, SeJoon Jun, Gayoung Lee, Jin-Hwa Kim, Il-Chul Moon |
NeurIPS | 9 |
| 2025 | Diffusion Adaptive Text Embedding for Text-to-Image Diffusion ModelsabstractText-to-image diffusion models rely on text embeddings from a pre-trained text encoder, but these embeddings remain fixed across all diffusion timesteps, limiting their adaptability to the generative process. We propose Diffusion Adaptive Text Embedding (DATE), which dynamically updates text embeddings at each diffusion timestep based on intermediate perturbed data. We formulate an optimization problem and derive an update rule that refines the text embeddings at each sampling step to improve alignment and preference between the mean predicted image and the text. This allows DATE to dynamically adapts the text conditions to the reverse-diffused images throughout diffusion sampling without requiring additional model training. Through theoretical analysis and empirical results, we show that DATE maintains the generative capability of the model while providing superior text-image alignment over fixed text embeddings across various tasks, including multi-concept generation and text-guided image editing. Our code is available at https://github.com/aailab-kaist/DATE. Byeonghu Na, Minsang Park, Gyuwon Sim, Donghyeok Shin, HeeSun Bae, Mina Kang, Se Jung Kwon, Wanmo Kang, Il-Chul Moon |
NeurIPS | 9 |
| 2025 | Generalized Gumbel-Softmax gradient estimator for generic discrete random variables
Weonyoung Joo, Il-Chul Moon |
Pattern Recognit. Lett. | 4 |
| 2024 | Make Prompts Adaptable: Bayesian Modeling for Vision-Language Prompt Learning with Data-Dependent PriorabstractRecent vision-language pre-trained (VLP) models have become the backbone for many downstream tasks, but they are utilized as frozen model without learning. Prompt learning is a method to improve the pre-trained VLP model by adding a learnable context vector to the inputs of the text encoder. In a few-shot learning scenario of the downstream task, MLE training can lead the context vector to over-fit dominant image features in the training data. This overfitting can potentially harm the generalization ability, especially in the presence of a distribution shift between the training and test dataset. This paper presents a Bayesian-based framework of prompt tuning, which could alleviate the over-fitting issues on few-shot learning application and increase the adaptability of prompts on unobserved instances. Specifically, modeling data-dependent prior enhances the adaptability of text features for both seen and unseen image features without the trade-off of performance between them. Based on the Bayesian framework, we utilize the Wasserstein gradient flow in the estimation of our target posterior distribution, which enables our prompt to be flexible in capturing the complex modes of image features. We demonstrate the effectiveness of our method on benchmark datasets for several experiments by showing statistically significant improvements on performance compared to existing methods. Youngjae Cho 0002, HeeSun Bae, Yeo Dong Youn, Weonyoung Joo, Il-Chul Moon |
AAAI | 6 |
| 2024 | Reward-based Input Construction for Cross-document Relation ExtractionabstractRelation extraction (RE) is a fundamental task in natural language processing, aiming to identify relations between target entities in text.While many RE methods are designed for a single sentence or document, cross-document RE has emerged to address relations across multiple long documents.Given the nature of long documents in cross-document RE, extracting document embeddings is challenging due to the length constraints of pre-trained language models.Therefore, we propose REward-based Input Construction (REIC), the first learningbased sentence selector for cross-document RE.REIC extracts sentences based on relational evidence, enabling the RE module to effectively infer relations.Since supervision of evidence sentences is generally unavailable, we train REIC using reinforcement learning with RE prediction scores as rewards.Experimental results demonstrate the superiority of our method over heuristic methods for different RE structures and backbones in crossdocument RE. Byeonghu Na, Suhyeon Jo, Yeongmin Kim, Il-Chul Moon |
ACL (1) | 4 |
| 2024 | Dirichlet-based Per-Sample Weighting by Transition Matrix for Noisy Label LearningabstractFor learning with noisy labels, the transition matrix, which explicitly models the relation between noisy label distribution and clean label distribution, has been utilized to achieve the statistical consistency of either the classifier or the risk. Previous researches have focused more on how to estimate this transition matrix well, rather than how to utilize it. We propose good utilization of the transition matrix is crucial and suggest a new utilization method based on resampling, coined RENT. Specifically, we first demonstrate current utilizations can have potential limitations for implementation. As an extension to Reweighting, we suggest the Dirichlet distribution-based per-sample Weight Sampling (DWS) framework, and compare reweighting and resampling under DWS framework. With the analyses from DWS, we propose RENT, a REsampling method with Noise Transition matrix. Empirically, RENT consistently outperforms existing transition matrix utilization methods, which includes reweighting, on various benchmark datasets. Our code is available at https://github.com/BaeHeeSun/RENT. HeeSun Bae, Byeonghu Na, Il-Chul Moon |
ICLR | 4 |
| 2024 | Training Unbiased Diffusion Models From Biased DatasetabstractWith significant advancements in diffusion models, addressing the potential risks of dataset bias becomes increasingly important. Since generated outputs directly suffer from dataset bias, mitigating latent bias becomes a key factor in improving sample quality and proportion. This paper proposes time-dependent importance reweighting to mitigate the bias for the diffusion models. We demonstrate that the time-dependent density ratio becomes more precise than previous approaches, thereby minimizing error propagation in generative learning. While directly applying it to score-matching is intractable, we discover that using the time-dependent density ratio both for reweighting and score correction can lead to a tractable form of the objective function to regenerate the unbiased data density. Furthermore, we theoretically establish a connection with traditional score-matching, and we demonstrate its convergence to an unbiased distribution. The experimental evidence supports the usefulness of the proposed method, which outperforms baselines including time-independent importance reweighting on CIFAR-10, CIFAR-100, FFHQ, and CelebA with various bias settings. Our code is available at https://github.com/alsdudrla10/TIW-DSM. Yeongmin Kim, Byeonghu Na, Minsang Park, JoonHo Jang, Wanmo Kang, Il-Chul Moon |
ICLR | 7 |
| 2024 | Label-Noise Robust Diffusion ModelsabstractConditional diffusion models have shown remarkable performance in various generative tasks, but training them requires large-scale datasets that often contain noise in conditional inputs, a.k.a. noisy labels. This noise leads to condition mismatch and quality degradation of generated data. This paper proposes Transition-aware weighted Denoising Score Matching (TDSM) for training conditional diffusion models with noisy labels, which is the first study in the line of diffusion models. The TDSM objective contains a weighted sum of score networks, incorporating instance-wise and time-dependent label transition probabilities. We introduce a transition-aware weight estimator, which leverages a time-dependent noisy-label classifier distinctively customized to the diffusion process. Through experiments across various datasets and noisy label settings, TDSM improves the quality of generated samples aligned with given conditions. Furthermore, our method improves generation performance even on prevalent benchmark datasets, which implies the potential noisy labels and their risk of generative model learning. Finally, we show the improved performance of TDSM on top of conventional noisy label corrections, which empirically proving its contribution as a part of label-noise robust generative models. Our code is available at: https://github.com/byeonghu-na/tdsm. Byeonghu Na, Yeongmin Kim, HeeSun Bae, Jung Hyun Lee, Se Jung Kwon, Wanmo Kang, Il-Chul Moon |
ICLR | 7 |
| 2024 | Efficient Episodic Memory Utilization of Cooperative Multi-Agent Reinforcement LearningabstractIn cooperative multi-agent reinforcement learning (MARL), agents aim to achieve a common goal, such as defeating enemies or scoring a goal. Existing MARL algorithms are effective but still require significant learning time and often get trapped in local optima by complex tasks, subsequently failing to discover a goal-reaching policy. To address this, we introduce Efficient episodic Memory Utilization (EMU) for MARL, with two primary objectives: (a) accelerating reinforcement learning by leveraging semantically coherent memory from an episodic buffer and (b) selectively promoting desirable transitions to prevent local convergence. To achieve (a), EMU incorporates a trainable encoder/decoder structure alongside MARL, creating coherent memory embeddings that facilitate exploratory memory recall. To achieve (b), EMU introduces a novel reward structure called episodic incentive based on the desirability of states. This reward improves the TD target in Q-learning and acts as an additional incentive for desirable transitions. We provide theoretical support for the proposed incentive and demonstrate the effectiveness of EMU compared to conventional episodic control. The proposed method is evaluated in StarCraft II and Google Research Football, and empirical results indicate further performance improvement over state-of-the-art methods. Hyungho Na, Yunkyeong Seo, Il-Chul Moon |
ICLR | 3 |
| 2024 | Unknown Domain Inconsistency Minimization for Domain GeneralizationabstractThe objective of domain generalization (DG) is to enhance the transferability of the model learned from a source domain to unobserved domains. To prevent overfitting to a specific domain, Sharpness-Aware Minimization (SAM) reduces source domain’s loss sharpness. Although SAM variants have delivered significant improvements in DG, we highlight that there’s still potential for improvement in generalizing to unknown domains through the exploration on data space. This paper introduces an objective rooted in both parameter and data perturbed regions for domain generalization, coined Unknown Domain Inconsistency Minimization (UDIM). UDIM reduces the loss landscape inconsistency between source domain and unknown domains. As unknown domains are inaccessible, these domains are empirically crafted by perturbing instances from the source domain dataset. In particular, by aligning the loss landscape acquired in the source domain to the loss landscape of perturbed domains, we expect to achieve generalization grounded on these flat minima for the unknown domains. Theoretically, we validate that merging SAM optimization with the UDIM objective establishes an upper bound for the true objective of the DG task. In an empirical aspect, UDIM consistently outperforms SAM variants across multiple DG benchmark datasets. Notably, UDIM shows statistically significant improvements in scenarios with more restrictive domain information, underscoring UDIM’s generalization capability in unseen domains. HeeSun Bae, Byeonghu Na, Yoon-Yeong Kim, Il-Chul Moon |
ICLR | 5 |
| 2024 | Diffusion Rejection SamplingabstractRecent advances in powerful pre-trained diffusion models encourage the development of methods to improve the sampling performance under well-trained diffusion models. This paper introduces Diffusion Rejection Sampling (DiffRS), which uses a rejection sampling scheme that aligns the sampling transition kernels with the true ones at each timestep. The proposed method can be viewed as a mechanism that evaluates the quality of samples at each intermediate timestep and refines them with varying effort depending on the sample. Theoretical analysis shows that DiffRS can achieve a tighter bound on sampling error compared to pre-trained models. Empirical results demonstrate the state-of-the-art performance of DiffRS on the benchmark datasets and the effectiveness of DiffRS for fast diffusion samplers and large-scale text-to-image diffusion models. Our code is available at https://github.com/aailabkaist/DiffRS. Byeonghu Na, Yeongmin Kim, Minsang Park, Donghyeok Shin, Wanmo Kang, Il-Chul Moon |
ICML | 6 |
| 2024 | LAGMA: LAtent Goal-guided Multi-Agent Reinforcement LearningabstractIn cooperative multi-agent reinforcement learning (MARL), agents collaborate to achieve common goals, such as defeating enemies and scoring a goal. However, learning goal-reaching paths toward such a semantic goal takes a considerable amount of time in complex tasks and the trained model often fails to find such paths. To address this, we present LAtent Goal-guided Multi-Agent reinforcement learning (LAGMA), which generates a goal-reaching trajectory in latent space and provides a latent goal-guided incentive to transitions toward this reference trajectory. LAGMA consists of three major components: (a) quantized latent space constructed via a modified VQ-VAE for efficient sample utilization, (b) goal-reaching trajectory generation via extended VQ codebook, and (c) latent goal-guided intrinsic reward generation to encourage transitions towards the sampled goal-reaching path. The proposed method is evaluated by StarCraft II with both dense and sparse reward settings and Google Research Football. Empirical results show further performance improvement over state-of-the-art baselines. Hyungho Na, Il-Chul Moon |
ICML | 2 |
| 2023 | Loss-Curvature Matching for Dataset Selection and CondensationabstractTraining neural networks on a large dataset requires substantial computational costs. Dataset reduction selects or synthesizes data instances based on the large dataset, while minimizing the degradation in generalization performance from the full dataset. Existing methods utilize the neural network during the dataset reduction procedure, so the model parameter becomes important factor in preserving the performance after reduction. By depending upon the importance of parameters, this paper introduces a new reduction objective, coined LCMat, which Matches the Loss Curvatures of the original dataset and reduced dataset over the model parameter space, more than the parameter point. This new objective induces a better adaptation of the reduced dataset on the perturbed parameter region than the exact point matching. Particularly, we identify the worst case of the loss curvature gap from the local parameter region, and we derive the implementable upper bound of such worst-case with theoretical analyses. Our experiments on both coreset selection and condensation benchmarks illustrate that LCMat shows better generalization performances than existing baselines. HeeSun Bae, Donghyeok Shin, Weonyoung Joo, Il-Chul Moon |
AISTATS | 5 |
| 2023 | Hierarchical Multi-Label Classification with Partial Labels and Unknown HierarchyabstractHierarchical multi-label classification aims at learning a multi-label classifier from a dataset whose labels are organized into a hierarchical structure. To the best of our knowledge, we propose for the first time the problem of finding a multi-label classifier given a partially labeled hierarchical multi-label dataset. We also assume the situation where the classifier cannot access hierarchical information during training. This work proposes an iterative framework for learning both multi-labels and a hierarchical structure of classes. When training a multi-label classifier from partial labels, our model extracts a class hierarchy from the classifier output using our hierarchy extraction algorithm. Then, our proposed loss exploits the extracted hierarchy to train the classifier. Theoretically, we show that our hierarchy extraction algorithm correctly finds the unknown hierarchy under a mild condition, and we prove that our loss function of multi-label classification with such hierarchy becomes an unbiased estimator of true multi-label classification risk. Our experiments show that our model obtains a class hierarchy close to the ground-truth dataset hierarchy, and simultaneously, our method outperforms previous methods for hierarchical multi-label classification and multi-label classification from partial labels. Suhyeon Jo, Donghyeok Shin, Byeonghu Na, JoonHo Jang, Il-Chul Moon |
CIKM | 5 |
| 2023 | SAAL: Sharpness-Aware Active LearningabstractWhile deep neural networks play significant roles in many research areas, they are also prone to overfitting problems under limited data instances. To overcome overfitting, this paper introduces the first active learning method to incorporate the sharpness of loss space into the acquisition function. Specifically, our proposed method, Sharpness-Aware Active Learning (SAAL), constructs its acquisition function by selecting unlabeled instances whose perturbed loss becomes maximum. Unlike the Sharpness-Aware learning with fully-labeled datasets, we design a pseudo-labeling mechanism to anticipate the perturbed loss w.r.t. the ground-truth label, which we provide the theoretical bound for the optimization. We conduct experiments on various benchmark datasets for vision-based tasks in image classification, object detection, and domain adaptive semantic segmentation. The experimental results confirm that SAAL outperforms the baselines by selecting instances that have the potentially maximal perturbation on the loss. The code is available at https://github.com/YoonyeongKim/SAAL. Yoon-Yeong Kim, Youngjae Cho 0002, JoonHo Jang, Byeonghu Na, Yeongmin Kim, Kyungwoo Song, Wanmo Kang, Il-Chul Moon |
ICML | 8 |
| 2023 | Refining Generative Process with Discriminator Guidance in Score-based Diffusion ModelsabstractThe proposed method, Discriminator Guidance, aims to improve sample generation of pre-trained diffusion models. The approach introduces a discriminator that gives explicit supervision to a denoising sample path whether it is realistic or not. Unlike GANs, our approach does not require joint training of score and discriminator networks. Instead, we train the discriminator after score training, making discriminator training stable and fast to converge. In sample generation, we add an auxiliary term to the pre-trained score to deceive the discriminator. This term corrects the model score to the data score at the optimal discriminator, which implies that the discriminator helps better score estimation in a complementary way. Using our algorithm, we achive state-of-the-art results on ImageNet 256x256 with FID 1.83 and recall 0.64, similar to the validation data’s FID (1.68) and recall (0.66). We release the code at https://github.com/alsdudrla10/DG. Yeongmin Kim, Se Jung Kwon, Wanmo Kang, Il-Chul Moon |
ICML | 5 |
| 2023 | Frequency Domain-Based Dataset DistillationabstractThis paper presents FreD, a novel parameterization method for dataset distillation, which utilizes the frequency domain to distill a small-sized synthetic dataset from a large-sized original dataset. Unlike conventional approaches that focus on the spatial domain, FreD employs frequency-based transforms to optimize the frequency representations of each data instance. By leveraging the concentration of spatial domain information on specific frequency components, FreD intelligently selects a subset of frequency dimensions for optimization, leading to a significant reduction in the required budget for synthesizing an instance. Through the selection of frequency dimensions based on the explained variance, FreD demonstrates both theoretical and empirical evidence of its ability to operate efficiently within a limited budget, while better preserving the information of the original dataset compared to conventional parameterization methods. Furthermore, Based on the orthogonal compatibility of FreD with existing methods, we confirm that FreD consistently improves the performances of existing distillation methods over the evaluation scenarios with different benchmark datasets. We release the code at https://github.com/sdh0818/FreD. Donghyeok Shin, Il-Chul Moon |
NeurIPS | 3 |
| 2023 | Time-Efficient Weapon-Target Assignment by Actor-Critic ReinforcementabstractThis paper proposes a time-efficient model for solving the Weapon-target assignment (WTA) problem with actor-critic reinforcement learning. While typical heuristic algorithms and recently studied artificial neural network methodologies have shown good performance results, the previous approach has not been time-efficient in large-scale WTA problems. This paper utilizes the actor-critic framework to resolve the WTA problem, and this framework enables retrieving solutions 23 times faster than the previous deep Q-network approach. Additionally, we incorporate a recurrent neural network model of gated recurrent units (GRU) to allow agents to learn the latent state-space of the WTA problem. Our experiments demonstrate the solution quality and the time efficiency compared to traditional heuristic methods as well as recent DQN-based RL models. Muhyun Byun, Hyungho Na, Il-Chul Moon |
SMC | 3 |
| 2023 | Sequential Likelihood-Free Inference with Neural Proposal
Kyungwoo Song, Yoon-Yeong Kim, Yongjin Shin, Wanmo Kang, Il-Chul Moon, Weonyoung Joo |
Pattern Recognit. Lett. | 6 |
| 2022 | From Noisy Prediction to True Label: Noisy Prediction Calibration via Generative ModelabstractNoisy labels are inevitable yet problematic in machine learning society. It ruins the generalization of a classifier by making the classifier over-fitted to noisy labels. Existing methods on noisy label have focused on modifying the classifier during the training procedure. It has two potential problems. First, these methods are not applicable to a pre-trained classifier without further access to training. Second, it is not easy to train a classifier and regularize all negative effects from noisy labels, simultaneously. We suggest a new branch of method, Noisy Prediction Calibration (NPC) in learning with noisy labels. Through the introduction and estimation of a new type of transition matrix via generative model, NPC corrects the noisy prediction from the pre-trained classifier to the true label as a post-processing scheme. We prove that NPC theoretically aligns with the transition matrix based methods. Yet, NPC empirically provides more accurate pathway to estimate true label, even without involvement in classifier learning. Also, NPC is applicable to any classifier trained with noisy label methods, if training instances and its predictions are available. Our method, NPC, boosts the classification performances of all baseline models on both synthetic and real-world datasets. The implemented code is available at https://github.com/BaeHeeSun/NPC. HeeSun Bae, Byeonghu Na, JoonHo Jang, Kyungwoo Song, Il-Chul Moon |
ICML | 6 |
| 2022 | Soft Truncation: A Universal Training Technique of Score-based Diffusion Model for High Precision Score EstimationabstractRecent advances in diffusion models bring state-of-the-art performance on image generation tasks. However, empirical results from previous research in diffusion models imply an inverse correlation between density estimation and sample generation performances. This paper investigates with sufficient empirical evidence that such inverse correlation happens because density estimation is significantly contributed by small diffusion time, whereas sample generation mainly depends on large diffusion time. However, training a score network well across the entire diffusion time is demanding because the loss scale is significantly imbalanced at each diffusion time. For successful training, therefore, we introduce Soft Truncation, a universally applicable training technique for diffusion models, that softens the fixed and static truncation hyperparameter into a random variable. In experiments, Soft Truncation achieves state-of-the-art performance on CIFAR-10, CelebA, CelebA-HQ $256\times 256$, and STL-10 datasets. Kyungwoo Song, Wanmo Kang, Il-Chul Moon |
ICML | 5 |
| 2022 | Unknown-Aware Domain Adversarial Learning for Open-Set Domain AdaptationabstractOpen-Set Domain Adaptation (OSDA) assumes that a target domain contains unknown classes, which are not discovered in a source domain. Existing domain adversarial learning methods are not suitable for OSDA because distribution matching with $\textit{unknown}$ classes leads to negative transfer. Previous OSDA methods have focused on matching the source and the target distribution by only utilizing $\textit{known}$ classes. However, this $\textit{known}$-only matching may fail to learn the target-$\textit{unknown}$ feature space. Therefore, we propose Unknown-Aware Domain Adversarial Learning (UADAL), which $\textit{aligns}$ the source and the target-$\textit{known}$ distribution while simultaneously $\textit{segregating}$ the target-$\textit{unknown}$ distribution in the feature alignment procedure. We provide theoretical analyses on the optimized state of the proposed $\textit{unknown-aware}$ feature alignment, so we can guarantee both $\textit{alignment}$ and $\textit{segregation}$ theoretically. Empirically, we evaluate UADAL on the benchmark datasets, which shows that UADAL outperforms other methods with better feature alignments by reporting state-of-the-art performances. JoonHo Jang, Byeonghu Na, Donghyeok Shin, Mingi Ji, Kyungwoo Song, Il-Chul Moon |
NeurIPS | 6 |
| 2022 | Maximum Likelihood Training of Implicit Nonlinear Diffusion ModelabstractWhereas diverse variations of diffusion models exist, extending the linear diffusion into a nonlinear diffusion process is investigated by very few works. The nonlinearity effect has been hardly understood, but intuitively, there would be promising diffusion patterns to efficiently train the generative distribution towards the data distribution. This paper introduces a data-adaptive nonlinear diffusion process for score-based diffusion models. The proposed Implicit Nonlinear Diffusion Model (INDM) learns by combining a normalizing flow and a diffusion process. Specifically, INDM implicitly constructs a nonlinear diffusion on the data space by leveraging a linear diffusion on the latent space through a flow network. This flow network is key to forming a nonlinear diffusion, as the nonlinearity depends on the flow network. This flexible nonlinearity improves the learning curve of INDM to nearly Maximum Likelihood Estimation (MLE) against the non-MLE curve of DDPM++, which turns out to be an inflexible version of INDM with the flow fixed as an identity mapping. Also, the discretization of INDM shows the sampling robustness. In experiments, INDM achieves the state-of-the-art FID of 1.75 on CelebA. We release our code at https://github.com/byeonghu-na/INDM. Byeonghu Na, Se Jung Kwon, Dongsoo Lee, Wanmo Kang, Il-Chul Moon |
NeurIPS | 6 |
| 2021 | Counterfactual Fairness with Disentangled Causal Effect Variational AutoencoderabstractThe problem of fair classification can be mollified if we develop a method to remove the embedded sensitive information from the classification features. This line of separating the sensitive information is developed through the causal inference, and the causal inference enables the counterfactual generations to contrast the what-if case of the opposite sensitive attribute. Along with this separation with the causality, a frequent assumption in the deep latent causal model defines a single latent variable to absorb the entire exogenous uncertainty of the causal graph. However, we claim that such structure cannot distinguish the 1) information caused by the intervention (i.e., sensitive variable) and 2) information correlated with the intervention from the data. Therefore, this paper proposes Disentangled Causal Effect Variational Autoencoder (DCEVAE) to resolve this limitation by disentangling the exogenous uncertainty into two latent variables: either 1) independent to interventions or 2) correlated to interventions without causality. Particularly, our disentangling approach preserves the latent variable correlated to interventions in generating counterfactual examples. We show that our method estimates the total effect and the counterfactual effect without a complete causal graph. By adding a fairness regularization, DCEVAE generates a counterfactual fair dataset while losing less original information. Also, DCEVAE generates natural counterfactual images by only flipping sensitive information. Additionally, we theoretically show the differences in the covariance structures of DCEVAE and prior works from the perspective of the latent disentanglement. Hyemi Kim, JoonHo Jang, Kyungwoo Song, Weonyoung Joo, Wanmo Kang, Il-Chul Moon |
AAAI | 7 |
| 2021 | Implicit Kernel AttentionabstractAttention computes the dependency between representations, and it encourages the model to focus on the important selective features. Attention-based models, such as Transformer and graph attention network (GAT), are widely utilized for sequential data and graph-structured data. This paper suggests a new interpretation and generalized structure of the attention in Transformer and GAT. For the attention in Transformer and GAT, we derive that the attention is a product of two parts: 1) the RBF kernel to measure the similarity of two instances and 2) the exponential of L2 norm to compute the importance of individual instances. From this decomposition, we generalize the attention in three ways. First, we propose implicit kernel attention with an implicit kernel function instead of manual kernel selection. Second, we generalize L2 norm as the Lp norm. Third, we extend our attention to structured multi-head attention. Our generalized attention shows better performance on classification, translation, and regression tasks. Kyungwoo Song, Yohan Jung, Il-Chul Moon |
AAAI | 4 |
| 2021 | Refine Myself by Teaching Myself: Feature Refinement via Self-Knowledge DistillationabstractKnowledge distillation is a method of transferring the knowledge from a pretrained complex teacher model to a student model, so a smaller network can replace a large teacher network at the deployment stage. To reduce the necessity of training a large teacher model, the recent literatures introduced a self-knowledge distillation, which trains a student network progressively to distill its own knowledge without a pretrained teacher network. While Self-knowledge distillation is largely divided into a data augmentation based approach and an auxiliary network based approach, the data augmentation approach looses its local information in the augmentation process, which hinders its applicability to diverse vision tasks, such as semantic segmentation. Moreover, these knowledge distillation approaches do not receive the refined feature maps, which are prevalent in the object detection and semantic segmentation community. This paper proposes a novel self-knowledge distillation method, Feature Refinement via Self-Knowledge Distillation (FRSKD), which utilizes an auxiliary self-teacher network to transfer a refined knowledge for the classifier network. Our proposed method, FRSKD, can utilize both soft label and feature-map distillations for the self-knowledge distillation. Therefore, FRSKD can be applied to classification, and semantic segmentation, which emphasize preserving the local information. We demonstrate the effectiveness of FRSKD by enumerating its performance improvements in diverse tasks and benchmark datasets. The implemented code is available at https://github.com/MingiJi/FRSKD. Mingi Ji, Seunghyun Hwang, Gibeom Park, Il-Chul Moon |
CVPR | 5 |
| 2021 | LADA: Look-Ahead Data Acquisition via Augmentation for Deep Active LearningabstractActive learning effectively collects data instances for training deep learning models when the labeled dataset is limited and the annotation cost is high. Data augmentation is another effective technique to enlarge the limited amount of labeled instances. The scarcity of labeled dataset leads us to consider the integration of data augmentation and active learning. One possible approach is a pipelined combination, which selects informative instances via the acquisition function and generates virtual instances from the selected instances via augmentation. However, this pipelined approach would not guarantee the informativeness of the virtual instances. This paper proposes Look-Ahead Data Acquisition via augmentation, or LADA framework, that looks ahead the effect of data augmentation in the process of acquisition. LADA jointly considers both 1) unlabeled data instance to be selected and 2) virtual data instance to be generated by data augmentation, to construct the acquisition function. Moreover, to generate maximally informative virtual instances, LADA optimizes the data augmentation policy to maximize the predictive acquisition score, resulting in the proposal of InfoSTN and InfoMixup. The experimental results of LADA show a significant improvement over the recent augmentation and acquisition baselines that were independently applied. Yoon-Yeong Kim, Kyungwoo Song, JoonHo Jang, Il-Chul Moon |
NeurIPS | 4 |
| 2021 | Predict Sequential Credit Card Delinquency with VaDE-Seq2SeqabstractFor successful debt collection, it is important for credit card companies to judge the users’ capability of debt repayment. This has been assessed by domain experts in the past, but as the amount of data increases, there has been a rising demand for a more effective decision-making methods. Several machine learning algorithms have been proposed to pursue interpretation and high performance. We newly propose Variational Deep Embedding with Sequence to Sequence (VaDE-Seq2Seq), based on a deep neural network. By adding the VaDE structure to the encoder, the model properly reflects information on cluster assignments in latent space, and the model explains decision-making by tracking the cluster assignments. Most delinquency prediction studies predict only the next time step, whereas our model predicts the future sequence. It is a strength of our model because sequence prediction is difficult, but more practical. The model was tested with the data of 10,000 users from a Korean credit card company, and VaDE-Seq2Seq outperforms the other baseline models in terms of performance. In addition, we observe the history in the latent cluster assignment that was clearly distinguished between non-delinquency users and delinquency users. Yeongmin Kim, Youngjae Cho 0002, Hanbit Lee, Il-Chul Moon |
SMC | 4 |
| 2021 | Automatic calibration of dynamic and heterogeneous parameters in agent-based models
Tae-Sub Yun, Il-Chul Moon, Jang Won Bae |
Auton. Agents Multi Agent Syst. | 3 |
| 2020 | Sequential Recommendation with Relation-Aware Kernelized Self-AttentionabstractRecent studies identified that sequential Recommendation is improved by the attention mechanism. By following this development, we propose Relation-Aware Kernelized Self-Attention (RKSA) adopting a self-attention mechanism of the Transformer with augmentation of a probabilistic model. The original self-attention of Transformer is a deterministic measure without relation-awareness. Therefore, we introduce a latent space to the self-attention, and the latent space models the recommendation context from relation as a multivariate skew-normal distribution with a kernelized covariance matrix from co-occurrences, item characteristics, and user information. This work merges the self-attention of the Transformer and the sequential recommendation by adding a probabilistic model of the recommendation task specifics. We experimented RKSA over the benchmark datasets, and RKSA shows significant improvements compared to the recent baseline models. Also, RKSA were able to produce a latent space model that answers the reasons for recommendation. Mingi Ji, Weonyoung Joo, Kyungwoo Song, Yoon-Yeong Kim, Il-Chul Moon |
AAAI | 5 |
| 2020 | Hierarchically Clustered Representation LearningabstractThe joint optimization of representation learning and clustering in the embedding space has experienced a breakthrough in recent years. In spite of the advance, clustering with representation learning has been limited to flat-level categories, which often involves cohesive clustering with a focus on instance relations. To overcome the limitations of flat clustering, we introduce hierarchically-clustered representation learning (HCRL), which simultaneously optimizes representation learning and hierarchical clustering in the embedding space. Compared with a few prior works, HCRL firstly attempts to consider a generation of deep embeddings from every component of the hierarchy, not just leaf components. In addition to obtaining hierarchically clustered embeddings, we can reconstruct data by the various abstraction levels, infer the intrinsic hierarchical structure, and learn the level-proportion features. We conducted evaluations with image and text domains, and our quantitative analyses showed competent likelihoods and the best accuracies compared with the baselines. Su-Jin Shin, Kyungwoo Song, Il-Chul Moon |
AAAI | 3 |
| 2020 | Bivariate Beta-LSTM
Kyungwoo Song, JoonHo Jang, Il-Chul Moon |
AAAI | 4 |
| 2020 | Deep Generative Positive-Unlabeled Learning under Selection BiasabstractLearning in the positive-unlabeled (PU) setting is prevalent in real world applications. Many previous works depend upon theSelected Completely At Random (SCAR) assumption to utilize unlabeled data, but the SCAR assumption is not often applicable to the real world due to selection bias in label observations. This paper is the first generative PU learning model without the SCAR assumption. Specifically, we derive the PU risk function without the SCAR assumption, and we generate a set of virtual PU examples to train the classifier. Although our PU risk function is more generalizable, the function requires PU instances that do not exist in the observations. Therefore, we introduce the VAE-PU, which is a variant of variational autoencoders to separate two latent variables that generate either features or observation indicators. The separated latent information enables the model to generate virtual PU instances. We test the VAE-PU on benchmark datasets with and without the SCAR assumption. The results indicate that the VAE-PU is superior when selection bias exists, and the VAE-PU is also competent under the SCAR assumption. The results also emphasize that the VAE-PU is effective when there are few positive-labeled instances due to modeling on selection bias. Byeonghu Na, Hyemi Kim, Kyungwoo Song, Weonyoung Joo, Yoon-Yeong Kim, Il-Chul Moon |
CIKM | 6 |
| 2020 | Dirichlet Variational Autoencoder
Weonyoung Joo, Wonsung Lee, Sungrae Park, Il-Chul Moon |
Pattern Recognit. | 4 |
| 2020 | Layered Behavior Modeling via Combining Descriptive and Prescriptive Approaches: A Case Study of Infantry Company EngagementabstractDefense modeling and simulation (DM&S) has brought insights into how to efficiently operate combat entities, such as soldiers and weapon systems. Most DM&S works have been developed to reflect accurate descriptions of military doctrines, yet these doctrines provide only guidelines of military operations, not details about how the combat entities should behave. Because such vague parts are often fulfilled with the appropriate behavior of combat entities in a battlefield, one part argues that DM&S should consider individual combat behaviors as well. However, it is known as an infeasible problem discovering best individual actions from infinite searching space, such as the battlefield. This paper proposes a layered behavior modeling to practically resolve this issue. The proposed method applies descriptive modeling to reduce the searching space by employing domain-specific knowledge; and prescriptive modeling to discover best individual actions in the reduced space. For the generalization, the proposed method adapts both modeling methods being modularized, and then the proposed method suggested an interface between them that is based on their semantic analogies. Both modeling methods are modularized, so they are interacted through an interface defined in the proposed method. This paper presents a realization of the proposed method through a case study of infantry company-level operations. In the case study, the proposed method is implemented with discrete event system specification formalism as the descriptive part and Markov decision process as the prescriptive part. The experimental results illustrated that the combat effectiveness resulted from the proposed method is statistically better than that from the descriptive-only modeling, and the difference would be guided by the objective of the combat behavior. Through the presented experimental results and the discussion, this paper argues that future DM&S should consider a broad spectrum from the battlefield incorporating the rational behavior of military individuals. Jang Won Bae, Junseok Lee 0001, Kanghoon Lee, Jongmin Lee 0004, Kee-Eung Kim, Il-Chul Moon |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2019 | Adversarial Dropout for Recurrent Neural Networks
Sungrae Park, Kyungwoo Song, Mingi Ji, Wonsung Lee, Il-Chul Moon |
AAAI | 5 |
| 2019 | Hierarchical Context Enabled Recurrent Neural Network for RecommendationabstractA long user history inevitably reflects the transitions of personal interests over time. The analyses on the user history require the robust sequential model to anticipate the transitions and the decays of user interests. The user history is often modeled by various RNN structures, but the RNN structures in the recommendation system still suffer from the long-term dependency and the interest drifts. To resolve these challenges, we suggest HCRNN with three hierarchical contexts of the global, the local, and the temporary interests. This structure is designed to withhold the global long-term interest of users, to reflect the local sub-sequence interests, and to attend the temporary interests of each transition. Besides, we propose a hierarchical context-based gate structure to incorporate our interest drift assumption. As we suggest a new RNN structure, we support HCRNN with a complementary bi-channel attention structure to utilize hierarchical context. We experimented the suggested structure on the sequential recommendation tasks with CiteULike, MovieLens, and LastFM, and our model showed the best performances in the sequential recommendations. Kyungwoo Song, Mingi Ji, Sungrae Park, Il-Chul Moon |
AAAI | 4 |
| 2018 | Adversarial Dropout for Supervised and Semi-Supervised LearningabstractRecently, training with adversarial examples, which are generated by adding a small but worst-case perturbation on input examples, has improved the generalization performance of neural networks. In contrast to the biased individual inputs to enhance the generality, this paper introduces adversarial dropout, which is a minimal set of dropouts that maximize the divergence between 1) the training supervision and 2) the outputs from the network with the dropouts. The identified adversarial dropouts are used to automatically reconfigure the neural network in the training process, and we demonstrated that the simultaneous training on the original and the reconfigured network improves the generalization performance of supervised and semi-supervised learning tasks on MNIST, SVHN, and CIFAR-10. We analyzed the trained model to find the performance improvement reasons. We found that adversarial dropout increases the sparsity of neural networks more than the standard dropout. Finally, we also proved that adversarial dropout is a regularization term with a rank-valued hyper-parameter that is different from a continuous-valued parameter to specify the strength of the regularization. Sungrae Park, Jun-Keon Park, Su-Jin Shin, Il-Chul Moon |
AAAI | 4 |
| 2018 | Neural Ideal Point Estimation Network
Kyungwoo Song, Wonsung Lee, Il-Chul Moon |
AAAI | 3 |
| 2018 | Diagnosis Prediction via Medical Context Attention Networks Using Deep Generative ModelingabstractPredicting the clinical outcome of patients from the historical electronic health records (EHRs) is a fundamental research area in medical informatics. Although EHRs contain various records associated with each patient, the existing work mainly dealt with the diagnosis codes by employing recurrent neural networks (RNNs) with a simple attention mechanism. This type of sequence modeling often ignores the heterogeneity of EHRs. In other words, it only considers historical diagnoses and does not incorporate patient demographics, which correspond to clinically essential context, into the sequence modeling. To address the issue, we aim at investigating the use of an attention mechanism that is tailored to medical context to predict a future diagnosis. We propose a medical context attention (MCA)-based RNN that is composed of an attention-based RNN and a conditional deep generative model. The novel attention mechanism utilizes the derived individual patient information from conditional variational autoencoders (CVAEs). The CVAE models a conditional distribution of patient embeddings and his/her demographics to provide the measurement of patient's phenotypic difference due to illness. Experimental results showed the effectiveness of the proposed model. Wonsung Lee, Sungrae Park, Weonyoung Joo, Il-Chul Moon |
ICDM | 4 |
| 2018 | Data-Driven Automatic Calibration for Validation of Agent-Based Social SimulationsabstractThough agent based models are used in many domains, the usage have been either very abstract model for conceptual experiments or very detailed models with huge engineering efforts in their modeling and calibration. One reason of this limited usage comes from the difficulties in calibrating and validating the model with observed data because the models are very generative in its nature with many hand-picked parameters. This paper presents a noble framework of augmenting machine learning techniques to agent-based models for better calibration and validation. The framework identifies periods of deviation between the simulation and the observation with hierarchical Dirichlet process hidden Markov Model, or HDP-HMM, and the framework automatically calibrates the temporal macro parameters by searching parameter spaces with more likelihoods of validation. After iterations of this framework, our experiments demonstrated sucessful validations on a hypothestical simple segregation model as well as a real world real estate model. This framework is generally usable in any agent based models with temporal macro parameters, which could be true in many existing models. Il-Chul Moon, Tae-Sub Yun, Jang Won Bae, Dong-oh Kang, Euihyun Paik |
SMC | 1 |
| 2018 | Deep Reinforcement Learning with Fully Convolutional Neural Network to Solve an Earthwork Scheduling ProblemabstractThis paper proposes a deep reinforcement learning approach in order to optimize a sequence of tasks efficiently with the aid of image processing techniques used in computer vision. The proposed algorithm can be employed to solve the traveling salesman problem (TSP), a combinatorial optimization problem that determines the optimum trajectory of city visits so that the total traveling distance is minimized. The proposed algorithm accepts a set of images as an input, and outputs the priority over alternative tasks (or sites to visit) that should be conducted at the next time step. The proposed method applies a stacked convolutional network layer to effectively process and extract the meaningful features and uses a fully convolutional network to map the processed features to the output tasks without losing the local connectivity in the input images. The proposed algorithm has been employed to optimize the excavation schedule of a single digger for completing a 20 by 20 grid world, which is equivalent to the TSP problem with a node size of 400. The simulation results showed that the proposed method can achieve an effective schedule with optimality comparable to state of the art algorithms. Seongcheol Woo, Juneyeong Yeon, Mingi Ji, Il-Chul Moon, Jinkyoo Park |
SMC | 4 |
| 2018 | Hierarchical prescription pattern analysis with symptom labels
Su-Jin Shin, Je-Yong Oh, Sungrae Park, Il-Chul Moon |
Pattern Recognit. Lett. | 5 |
| 2018 | Evaluation of Disaster Response System Using Agent-Based Model With Geospatial and Medical DetailsabstractMany disasters have occurred around the world and have caused sizable damage. A disaster, called a mass casualty incident (MCI), generates a large number of casualties that overwhelm the capacity of local medical resources, and the disaster responses to the MCI requires many interactions among the disaster responders. To evaluate the efficiency of the disaster responses against MCIs, this paper proposes an agent-based model describing the cooperations among the responders during the overall process in the disaster responses from transporting patients to their definitive care. In particular, the proposed model includes geospatial details, such as the road network and the location of hospitals around the disaster scene, and medical information, such as the distribution of medical resources and transporting units, in the region of interest to discover the key factors of the disaster response system that customized to the target region. The case study in this paper presents that the proposed approach was applied to describe a disaster response system and illustrates how the additional details are utilized to analyze the disaster response system. We expect that the proposed method can provide comprehensive insights to a disaster response system of interest, and it can be used as groundwork for improving the disaster response system. Jang Won Bae, Kyohong Shin, Hyun-Rok Lee, Hyunjin Lee 0002, Taesik Lee, Chu Hyun Kim, Won Chul Cha, Gi Woon Kim, Il-Chul Moon |
IEEE Trans. Syst. Man Cybern. Syst. | 9 |
| 2017 | Augmented Variational Autoencoders for Collaborative Filtering with Auxiliary InformationabstractRecommender systems offer critical services in the age of mass information. A good recommender system selects a certain item for a specific user by recognizing why the user might like the item. This awareness implies that the system should model the background of the items and the users. This background modeling for recommendation is tackled through the various models of collaborative filtering with auxiliary information. This paper presents variational approaches for collaborative filtering to deal with auxiliary information. The proposed methods encompass variational autoencoders through augmenting structures to model the auxiliary information and to model the implicit user feedback. This augmentation includes the ladder network and the generative adversarial network to extract the low-dimensional representations influenced by the auxiliary information. These two augmentations are the first trial in the venue of the variational autoencoders, and we demonstrate their significant improvement on the performances in the applications of the collaborative filtering. Wonsung Lee, Kyungwoo Song, Il-Chul Moon |
CIKM | 3 |
| 2017 | Hybrid modeling and simulation of tactical maneuvers in computer generated forceabstractDefense modeling and simulation (DM&S) offers insights into the efficient operations of combat entities, e.g., soldiers and weapon systems. Most DM&S aim at exact description of military doctrines, but often the doctrines fails to provide detail action procedures about how the combat entities conduct military operations. Such unspecified descriptions are filled with the rational behaviors of the combat entities in a battlefield, and thereby the combat effectiveness from these combat entities would differ. Also, by incorporating such rational factors, this could provide the insights that cannot be captured from the traditional works. To examine this postulation, this paper developed a computer generated force where the tactical maneuver of combat entities are realized by the combination of descriptive and prescriptive modeling. Specifically, the descriptive models describe the explicit action rules in military doctrines, and they are modeled using discrete event system specification (DEVS) formalism; the predictive models denoted the rational behavior of the combat entities under the military doctrines, and they are modeled using partially observable Markov decision process (POMDP). The provided results illustrated that the proposed approach helps to maintain a team formation effectively, and this formation maintenance lead to the better combat efficiency. Jang Won Bae, Bowon Nam, Kee-Eung Kim, Junseok Lee 0001, Il-Chul Moon |
SMC | 5 |
| 2017 | Text augmented automatic statistician for predicting approval rates of politiciansabstractPredicting an approval rate of politicians is a popular task. While a type of prediction is using a text mining from news articles, we introduce a text augmented Gaussian process to perform the prediction with contexts. We test our model with 2017 South Korea Presidential Election in 1) a quantitative evaluation, and 2) a qualitative analysis. The performance of the model with text input is better than the performance of the model without the text input, which has been a typical approach of applying the Gaussian process. Moreover, the model can capture keywords which provide behind rational of the prediction result, which was not provided with only temporal information. Jun-Keon Park, YeongYeon Na, Il-Chul Moon |
SMC | 3 |
| 2017 | Identifying prescription patterns with a topic model of diseases and medications
Sungrae Park, Doosup Choi, Won Chul Cha, Chu Hyun Kim, Il-Chul Moon |
J. Biomed. Informatics | 6 |
| 2017 | Guided HTM: Hierarchical Topic Model with Dirichlet Forest PriorsabstractDespite the proliferation of topic models, the organization of topics from the probabilistic models needs improvement in two ways: the better structured presentation of topics and the incorporation of domain knowledge on the corpus. The structured presentation, i.e., the hierarchical topic model, helps in categorizing similar topics; the incorporation of domain knowledge enables the concentrated sampling of predefined keywords in the mixture parameter learning. This paper presents a hierarchical topic models with incorporated domain knowledge, called Guided Hierarchical Topic Model (GHTM). Specifically, we allocated the prior information from the knowledge to the Dirichlet Forest prior. From the prior adjustment, we obtained the topic tree guided by the domain knowledge. This paper also contributes in enumerating four different knowledge extraction methods and applying the extracted knowledge to GHTM. We evaluated the performance of GHTM in terms of the hierarchical clustering accuracy, and we found a significant improvement of hierarchical clustering measured by F-measures. This improvement is also verified by the perplexity analyses. Additionally, we measured topic quality with KL-divergence and visualization, and these confirm the ability to better separate topic distributions. Finally, we tested the hierarchical topic quality through human experiments, and this also revealed significant improvements originating from the guidance. Su-Jin Shin, Il-Chul Moon |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | Data-driven ballistic coefficient learning for future state prediction of high-speed vehicles
Kyungwoo Song, Jinhyung Tak, Han-Lim Choi, Il-Chul Moon |
FUSION | 5 |
| 2016 | Bayesian Nonparametric Collaborative Topic Poisson Factorization for Electronic Health Records-Based Phenotyping
Wonsung Lee, Youngmin Lee, Heeyoung Kim, Il-Chul Moon |
IJCAI | 4 |
| 2016 | LDEF Formalism for Agent-Based Model DevelopmentabstractAs agent-based models (ABMs) are applied to various domains, the efficiency of model development has become an important issue in its applications. The current practice is that many models are developed from scratch, while they could have been built by reusing existing models. Moreover, when models need reconfiguration, they often need to be rebuilt significantly. These problems reduce the development efficiency and ultimately damage the efficacy of ABM. This paper partially resolves the challenges of model reusability from the systems engineering approach. Specifically, we propose a formalism-based ABM development and demonstrate its potential to promote model reuses. Our formalism, named large-scale, dynamic, extensible, and flexible (LDEF) formalism, encourages the building of a larger model by the composition of modularly developed components. Also, LDEF is tailored to the ABM contexts to represent the agent's action procedure and support the dynamic changes of their interactions. This paper shows that LDEF improves the model reusability in ABM development through its practical examples and theoretical discussions. Jang Won Bae, Il-Chul Moon |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2015 | Associative topic models with numerical time series
Sungrae Park, Wonsung Lee, Il-Chul Moon |
Inf. Process. Manag. | 3 |
| 2015 | Efficient extraction of domain specific sentiment lexicon with active learning
Sungrae Park, Wonsung Lee, Il-Chul Moon |
Pattern Recognit. Lett. | 3 |
| 2014 | Network analysis approach to study hospitals' prescription patterns focused on the impact of new healthcare policyabstractUnderstanding hospitals' relationships is critical to the analysis of public healthcare environment. There have been many attempts to analyze medical environment at a personal level. Recently, at an organizational level, there has been some advance in research into examining a relationship between hospitals. However, the formation of linkages is restricted to explicit and direct interactions. In contrast, we focused on implicit information flows between hospitals. This study also analyzes large scale hospital networks based on prescribing similarity. The sample dataset we used is the trustworthy representative of actual population in Korea. We assessed the impact of Drug Utilization Review (DUR) on hospital network characteristics. We examined National Inpatient Sample (NIS) dataset for before-DUR year (2010) and after-DUR year (2011). Various network metrics and performance measures are calculated for the two years. Generated hospital networks of the two years were significantly different in terms of both network metrics and performance measures, except for a riskiness measure. In network clustering result, Spearman's correlation coefficients indicated that network metrics can be used to evaluate hospitals having extreme prescription patterns. We anticipate our novel approach allows us to better understand public healthcare environment. Wonsung Lee, Gene Yi, Dain Jung, Il-Chul Moon |
SMC | 5 |
| 2014 | Disease-medicine topic model for prescription record miningabstractAnalyzing patient records is important for improving the quality of medical services and for understanding each patient's historical diseases. However, the huge size of the data requires statistical analysis procedures. In this paper, we proposed a probabilistic model-the disease-medicine topic model (DMTM)-to explore connected knowledge about diseases and medicines. In the model, diseases and medicines are modeled using generative process. We used the latent Dirichlet allocation, which is one of the most popular topic models, as the baseline model. Then, we compared the qualities of topic representations quantitatively and qualitatively. The comparison results showed that the topics derived from the DMTM are clearer to identify and that specific patterns were found in the diseases and medicines. In the case of topic network analysis, these specific patterns were proved using centrality measurements. Sungrae Park, Doosup Choi, Wonsung Lee, Dain Jung, Il-Chul Moon |
SMC | 6 |
| 2014 | Identifying the evolution of disasters and responses with network-text analysisabstractDisasters and responses have evolved over-time, and the evolution has been affected by various factors, such as societal change, climate change, and technological advance. To better prepare the future disasters, we need to estimate the evolution trend of the past disasters and the responses. This paper analyzes the academic articles of the field with network-text analyses. The analyses captured the word level and the topic level evolution over-time with statistical significance tests. Further, we turn the text mining results into the network analysis data to identify the key words and topics in the evolution paths. The proposed method suggests the swift of interests, i.e. the new ways of organizational interoperation, the evolution of logistic issues, in the disaster and response field. Kyungwoo Song, Do-Hyeong Kim, Su-Jin Shin, Il-Chul Moon |
SMC | 4 |
| 2011 | Analyzing social media in escalating crisis situationsabstractThe rapid diffusion of information and opinions through social media, such as web forums and micro-blogs, is affecting the development of crisis situations, such as the Iranian presidential election, the Egyptian protest, and the ROKS Cheonan sinking. Understanding this rapid widespread diffusion, and assessing what information is spreading, what ideas are becoming common, and who is talking about what, is critical for crisis management. This paper presents a computational system for social media assessing the flow of ideas on the web and changes in who is talking about what. This system, given raw social media data, identifies the key topics, the key paths by which topics evolve, the key individuals who contribute to the topic, and the key influence relations between the contributors. We present this system implemented with the Author-Topic model, the meta-network model, and various computational techniques to find and filter the heavy contributors and influences. We demonstrate the performance of the system, by applying it to social media data surrounding the ROKS Cheonan sinking. We describe the results of assessing the initial and changing perceptions of the event using this system. Il-Chul Moon, Alice Oh, Kathleen M. Carley |
ISI | 1 |
| 2009 | Simulation analysis on destabilization of complex adaptive organizationsabstractMany adversarial organizations, such as organized crime groups, terrorist networks, and the like, are complex adaptive organizations. Therefore, strategies against them should consider the natures of complexity and adaptivity. However, such natures create nonlinear effects that are difficult to predict. To mitigate those difficulties, I utilize agent-based simulations that could possibly capture unexpected responses coming from our interventions into their organizations. This paper presents a simulation analysis example of an action against a terrorist group. Particularly, this example points out three critical aspects of simulation analysis. First, this simulation example shows how to setup a simulation analysis to anticipate an intervention's results. Second, this example illustrates various results making analysis useful. Third, this example describes statistical processing of the results. I expect that the three points will advance the current practices of simulation analysis on complex adaptive organizations. Il-Chul Moon |
CISDA | 1 |
| 2009 | Mining social networks for personalized email prioritizationabstractEmail is one of the most prevalent communication tools today, and solving the email overload problem is pressingly urgent. A good way to alleviate email overload is to automatically prioritize received messages according to the priorities of each user. However, research on statistical learning methods for fully personalized email prioritization (PEP) has been sparse due to privacy issues, since people are reluctant to share personal messages and importance judgments with the research community. It is therefore important to develop and evaluate PEP methods under the assumption that only limited training examples can be available, and that the system can only have the personal email data of each user during the training and testing of the model for that user. This paper presents the first study (to the best of our knowledge) under such an assumption. Specifically, we focus on analysis of personal social networks to capture user groups and to obtain rich features that represent the social roles from the viewpoint of a particular user. We also developed a novel semi-supervised (transductive) learning algorithm that propagates importance labels from training examples to test examples through message and user nodes in a personal email network. These methods together enable us to obtain an enriched vector representation of each new email message, which consists of both standard features of an email message (such as words in the title or body, sender and receiver IDs, etc.) and the induced social features from the sender and receivers of the message. Using the enriched vector representation as the input in SVM classifiers to predict the importance level for each test message, we obtained significant performance improvement over the baseline system (without induced social features) in our experiments on a multi-user data collection. We obtained significant performance improvement over the baseline system (without induced social features) in our experiments on a multi-user data collection: the relative error reduction in MAE was 31% in micro-averaging, and 14% in macro-averaging. Shinjae Yoo, Yiming Yang 0002, Frank Lin, Il-Chul Moon |
KDD | 4 |