Jaesik Choi

dblp:13/1402 · DBLP profile ↗
← Back
70ranked-venue papers
10as first author
38since 2021 · last 2025
0000-0002-4663-3263ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 10 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 3 first-author · 17 since 2021Databases, data management, data science and information retrieval · 12 · 2 first-author · 4 since 2021Systems, architecture and hardware · 9 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3
YearPublicationVenuePosition
2025 Diverse Rare Sample Generation with Pretrained GANs
abstract
Deep generative models are proficient in generating realistic data but struggle with producing rare samples in low density regions due to their scarcity of training datasets and the mode collapse problem. While recent methods aim to improve the fidelity of generated samples, they often reduce diversity and coverage by ignoring rare and novel samples. This study proposes a novel approach for generating diverse rare samples from high-resolution image datasets with pretrained GANs. Our method employs gradient-based optimization of latent vectors within a multi-objective framework and utilizes normalizing flows for density estimation on the feature space. This enables the generation of diverse rare images, with controllable parameters for rarity, diversity, and similarity to a reference image. We demonstrate the effectiveness of our approach both qualitatively and quantitatively across various datasets and GANs without retraining or fine-tuning the pretrained GANs.
Subeen Lee, Jiyeon Han 0001, Jaesik Choi
AAAI4
2025 xPatch: Dual-Stream Time Series Forecasting with Exponential Seasonal-Trend Decomposition
abstract
In recent years, the application of transformer-based models in time-series forecasting has received significant attention. While often demonstrating promising results, the transformer architecture encounters challenges in fully exploiting the temporal relations within time series data due to its attention mechanism. In this work, we design eXponential Patch (xPatch for short), a novel dual-stream architecture that utilizes exponential decomposition. Inspired by the classical exponential smoothing approaches, xPatch introduces the innovative seasonal-trend exponential decomposition module. Additionally, we propose a dual-flow architecture that consists of an MLP-based linear stream and a CNN-based non-linear stream. This model investigates the benefits of employing patching and channel-independence techniques within a non-transformer model. Finally, we develop a robust arctangent loss function and a sigmoid learning rate adjustment scheme, which prevent overfitting and boost forecasting performance.
Artyom Stitsyuk, Jaesik Choi
AAAI2
2025 Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models
abstract
Large language models (LLMs) are increasingly engaging in moral and ethical reasoning, where criteria for judgment are often unclear, even for humans. While LLM alignment studies cover many areas, one important yet underexplored area is how LLMs make judgments about obligations. This work reveals a strong tendency in LLMs to judge non-obligatory contexts as obligations when prompts are augmented with modal expressions such as must or ought to. We introduce this phenomenon as Deontological Keyword Bias (DKB). We find that LLMs judge over 90% of commonsense scenarios as obligations when modal expressions are present. This tendency is consist across various LLM families, question types, and answer formats. To mitigate DKB, we propose a judgment strategy that integrates few-shot examples with reasoning prompts. This study sheds light on how modal expressions, as a form of linguistic framing, influence the normative decisions of LLMs and underscores the importance of addressing such biases to ensure judgment alignment.
Bumjin Park, Leejinsil Leejinsil, Jaesik Choi
ACL (1)3
2025 Human-Centric AI: From Explainability and Trustworthiness to Actionable Ethics
abstract
To address the potential risks of AI while supporting innovation and ensuring responsible adoption, there is an urgent need for clear governance frameworks grounded in human-centric values. It is imperative that AI systems operate in ways that are transparent, trustworthy, and ethically sound. Developing truly human-centric AI goes beyond technical innovation. It requires interdisciplinary collaboration and diverse perspectives. This workshop will explore key challenges and emerging solutions in the development of human-centric AI, with a focus on explainability, trustworthiness, fairness, and privacy. We welcome both theoretical contributions and practical case studies that demonstrate how human-centered principles are realized in real-world AI systems. The official workshop webpage is available at https://xai.kaist.ac.kr/Workshop/hcai2025/, which provides comprehensive information about the program.
Jaesik Choi, Bohyung Han, Myoung-Wan Koo, Kyungman Bae, Chang Dong Yoo, Simon S. Woo, Wojciech Samek
CIKM1
2025 Amortized Baseline Selection via Rank-Revealing QR for Efficient Model Explanation
abstract
Model-agnostic explanation methods are essential for interpreting machine learning models, but suffer from prohibitive computational costs that scale with the number of baselines. Existing acceleration approaches either lack a theoretical base or provide no principled guidance for baseline selection. To address this gap, we present ABSQR (Amortized Baseline Selection via Rank-Revealing QR). This framework exploits the low-rank structure of value matrices to accelerate multi-baseline attribution methods. Our approach combines deterministic baseline selection via SVD-guided QR decomposition with an amortized inference mechanism that utilizes cluster-based retrieval. We reduce computational complexity from O (m • 2d) to O (k • 2d), where k ≪ m. Experiments demonstrate that ABSQR achieves a 91.2% agreement rate with full baseline methods while providing 8.5× speedup across diverse datasets. As the first acceleration approach that preserves explanation error guarantees under computational speedup, ABSQR makes the practical deployment of interpretable AI systems feasible at scale.
Chanwoo Lee, Youngjin Park, Hyeongeun Lee, Yeeun Yoo, Daehee Han, Geon-Hyeong Kim, Nari Kim, Jaesik Choi
CIKM9
2025 Enhancing Creative Generation on Stable Diffusion-based Models
abstract
Recent text-to-image generative models, particularly Stable Diffusion and its distilled variants, have achieved impressive fidelity and strong text-image alignment. However, their creative capability remains constrained, as including ‘creative’ in prompts seldom yields the desired results. This paper introduces C3 (Creative Concept Catalyst), a training-free approach designed to enhance creativity in Stable Diffusion-based models. C3 selectively amplifies features during the denoising process to foster more creative outputs. We offer practical guidelines for choosing amplification factors based on two main aspects of creativity. C3 is the first study to enhance creativity in diffusion models without extensive computational costs. We demonstrate its effectiveness across various Stable Diffusion-based models.
Jiyeon Han 0001, Dahee Kwon, Gayoung Lee, Jaesik Choi
CVPR5
2025 Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept Representations
abstract
Deep vision models have achieved remarkable classification performance by leveraging a hierarchical architecture in which human-interpretable concepts emerge through the composition of individual neurons across layers. Given the distributed nature of representations, pinpointing where specific visual concepts are encoded within a model remains a crucial yet challenging task. In this paper, we introduce an effective circuit discovery method, called Granular Concept Circuit (GCC), in which each circuit represents a concept relevant to a given query. To construct each circuit, our method iteratively assesses inter-neuron connectivity, focusing on both functional dependencies and semantic alignment. By automatically discovering multiple circuits, each capturing specific concepts within that query, our approach offers a profound, concept-wise interpretation of models and is the first to identify circuits tied to specific visual concepts at a fine-grained level. We validate the versatility and effectiveness of GCCs across various deep image classification models.
Dahee Kwon, Sehyun Lee, Jaesik Choi
ICCV3
2025 On the Relationship Between Populated Regions and Adversarial Robustness in Deep Neural Networks
abstract
In general, deep neural networks (DNNs) are evaluated by the generalization performance measured on unseen data excluded from the training phase. Along with the development of DNNs, the generalization performance converges to the state-of-the-art performances and it becomes difficult to evaluate DNNs solely based on this metric. The robustness against adversarial attack has been used as an additional metric to evaluate DNNs by measuring their vulnerability. However, few studies have been performed to analyze the adversarial robustness in terms of the geometry in DNNs. In this work, we perform an empirical study to analyze the internal properties of DNNs that affect model robustness under adversarial attacks. In particular, we propose the novel concept of the populated region set (PRS), where training samples are actually populated, to represent the internal properties of DNNs in a practical setting. From systematic experiments with the proposed concept, we provide empirical evidence to validate that a low PRS ratio has a strong relationship with the adversarial robustness of DNNs. We also devise a PRS regularizer leveraging the characteristics of PRS to improve the adversarial robustness without adversarial training.
Seongjin Park, Haedong Jeong, Tair Djanibekov, Giyoung Jeon, Jinseok Seol, Jaesik Choi
ICDM6
2025 Rethinking Shapley Value for Negative Interactions in Non-convex Games
abstract
We study causal interactions for payoff allocation in cooperative game theory, including quantifying feature attribution for deep learning models. Most feature attribution methods mainly stem from the criteria of the Shapley value, which assigns fair payoffs to players based on their expected contribution in a cooperative game. However, interactions between players in the game do not explicitly appear in the original formulation of the Shapley value. In this work, we reformulate the Shapley value to clarify the role of interactions and discuss implicit assumptions from a game-theoretical perspective. Our theoretical analysis demonstrates that when negative interactions exist—common in deep learning models—the efficiency axiom can lead to the undervaluation of attributions or payoffs. We suggest a new allocation rule that decomposes contributions into interactions and aggregates positive parts for non-convex games. Furthermore, we propose an approximation algorithm to reduce the cost of interaction computation which can be applied to differentiable functions such as deep learning models. Our approach mitigates counterintuitive attribution outcomes observed in existing methods, ensuring that features critical to a model’s decision receive appropriate attribution.
Wonjoon Chang, Myeongjin Lee, Jaesik Choi
ICLR3
2025 Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning
abstract
Recent advancements in large language models (LLMs) based on transformer architectures have sparked significant interest in understanding their inner workings. In this paper, we introduce a novel approach to modeling transformer architectures using highly flexible non-autonomous neural ordinary differential equations (ODEs). Our proposed model parameterizes all weights of attention and feed-forward blocks through neural networks, expressing these weights as functions of a continuous layer index. Through spectral analysis of the model's dynamics, we uncover an increase in eigenvalue magnitude that challenges the weight-sharing assumption prevalent in existing theoretical studies. We also leverage the Lyapunov exponent to examine token-level sensitivity, enhancing model interpretability. Our neural ODE transformer demonstrates performance comparable to or better than vanilla transformers across various configurations and datasets, while offering flexible fine-tuning capabilities that can adapt to different architectural constraints.
Anh Tong, Thanh Nguyen-Tang, Dongeun Lee 0001, Toan M. Tran, David Hall 0006, Cheongwoong Kang, Jaesik Choi
ICLR8
2025 Local Manifold Approximation and Projection for Manifold-Aware Diffusion Planning
abstract
Recent advances in diffusion-based generative modeling have demonstrated significant promise in tackling long-horizon, sparse-reward tasks by leveraging offline datasets. While these approaches have achieved promising results, their reliability remains inconsistent due to the inherent stochastic risk of producing infeasible trajectories, limiting their applicability in safety-critical applications. We identify that the primary cause of these failures is inaccurate guidance during the sampling procedure, and demonstrate the existence of manifold deviation by deriving a lower bound on the guidance gap. To address this challenge, we propose Local Manifold Approximation and Projection (LoMAP), a training-free method that projects the guided sample onto a low-rank subspace approximated from offline datasets, preventing infeasible trajectory generation. We validate our approach on standard offline reinforcement learning benchmarks that involve challenging long-horizon planning. Furthermore, we show that, as a standalone module, LoMAP can be incorporated into the hierarchical diffusion planner, providing further performance enhancements.
Kyowoon Lee, Jaesik Choi
ICML2
2025 Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
Kyowoon Lee, Artyom Stitsyuk, Gunu Jho, Inchul Hwang, Jaesik Choi
INTERSPEECH5
2025 State-Covering Trajectory Stitching for Diffusion Planners
abstract
Diffusion-based generative models are emerging as powerful tools for long-horizon planning in reinforcement learning (RL), particularly with offline datasets. However, their performance is fundamentally limited by the quality and diversity of training data. This often restricts their generalization to tasks outside their training distribution or longer planning horizons. To overcome this challenge, we propose *State-Covering Trajectory Stitching* (SCoTS), a novel reward-free trajectory augmentation method that incrementally stitches together short trajectory segments, systematically generating diverse and extended trajectories. SCoTS first learns a temporal distance-preserving latent representation that captures the underlying temporal structure of the environment, then iteratively stitches trajectory segments guided by directional exploration and novelty to effectively cover and expand this latent space. We demonstrate that SCoTS significantly improves the performance and generalization capabilities of diffusion planners on offline goal-conditioned benchmarks requiring stitching and long-horizon reasoning. Furthermore, augmented trajectories generated by SCoTS significantly improve the performance of widely used offline goal-conditioned RL algorithms across diverse environments.
Kyowoon Lee, Jaesik Choi
NeurIPS2
2025 Deep-DFVAR: Dynamic factor vector autoregression with deep learning for regional house price index forecasting
Jisu Yeo, Artyom Stitsyuk, Jaesik Choi
Neurocomputing3
2025 Low-rank perturbation adjustment (LoPA): An implicit regularization method in image classification
abstract
Abstract Image classification with deep neural networks has reached state-of-the-art with high accuracy. The unreasonable effectiveness of deep neural networks is credited to the manifold hypothesis that states natural data lies on a low-dimensional manifold embedded in the high-dimensional space. The machine learning models learn patterns on these low-rank representations, which gives the learning algorithms robustness. We test the robustness of learning algorithms using the paradigm of “perturb and learn”. This paper proposes a novel technique called Low-Rank Perturbation Adjustment (LoPA), an implicit regularization method used by machine learning models for resisting external perturbations. LoPA exploits the dependencies in model weights that lie in high-dimensional space and projects to low-dimensional while resisting perturbation. We validate LoPA through singular value decomposition (SVD) theory and empirical experiments, showing the statistical distribution of trained model weights of zero mean and small variance. We inject perturbations into our model by hot-swapping the activation functions and interchanging loss functions during the training. An InceptionV3 Neural Network is trained on common FruitFly Drosophila images for binary classification tasks of cancer cells. The Drosophila cancer images are prepared in our lab through immunostaining protocol.
Nesma Talaat Abbas Mahmoud, Hanna Antson, Wai Tik Chan, Modar Sulaiman, Jaesik Choi, Osamu Shimmi, Kallol Roy
Multim. Tools Appl.5
2024 Understanding Distributed Representations of Concepts in Deep Neural Networks without Supervision
abstract
Understanding intermediate representations of the concepts learned by deep learning classifiers is indispensable for interpreting general model behaviors. Existing approaches to reveal learned concepts often rely on human supervision, such as pre-defined concept sets or segmentation processes. In this paper, we propose a novel unsupervised method for discovering distributed representations of concepts by selecting a principal subset of neurons. Our empirical findings demonstrate that instances with similar neuron activation states tend to share coherent concepts. Based on the observations, the proposed method selects principal neurons that construct an interpretable region, namely a Relaxed Decision Region (RDR), encompassing instances with coherent concepts in the feature space. It can be utilized to identify unlabeled subclasses within data and to detect the causes of misclassifications. Furthermore, the applicability of our method across various layers discloses distinct distributed representations over the layers, which provides deeper insights into the internal mechanisms of the deep learning model.
Wonjoon Chang, Dahee Kwon, Jaesik Choi
AAAI3
2024 Towards Diverse Perspective Learning with Selection over Multiple Temporal Poolings
abstract
In Time Series Classification (TSC), temporal pooling methods that consider sequential information have been proposed. However, we found that each temporal pooling has a distinct mechanism, and can perform better or worse depending on time series data. We term this fixed pooling mechanism a single perspective of temporal poolings. In this paper, we propose a novel temporal pooling method with diverse perspective learning: Selection over Multiple Temporal Poolings (SoM-TP). SoM-TP dynamically selects the optimal temporal pooling among multiple methods for each data by attention. The dynamic pooling selection is motivated by the ensemble concept of Multiple Choice Learning (MCL), which selects the best among multiple outputs. The pooling selection by SoM-TP's attention enables a non-iterative pooling ensemble within a single classifier. Additionally, we define a perspective loss and Diverse Perspective Learning Network (DPLN). The loss works as a regularizer to reflect all the pooling perspectives from DPLN. Our perspective analysis using Layer-wise Relevance Propagation (LRP) reveals the limitation of a single perspective and ultimately demonstrates diverse perspective learning of SoM-TP. We also show that SoM-TP outperforms CNN models based on other temporal poolings and state-of-the-art models in TSC with extensive UCR/UEA repositories.
Jihyeon Seong, Jaesik Choi
AAAI3
2024 Pathwise Explanation of ReLU Neural Networks
abstract
Neural networks have demonstrated a wide range of successes, but their “black box" nature raises concerns about transparency and reliability. Previous research on ReLU networks has sought to unwrap these networks into linear models based on activation states of all hidden units. In this paper, we introduce a novel approach that considers subsets of the hidden units involved in the decision making path. This pathwise explanation provides a clearer and more consistent understanding of the relationship between the input and the decision-making process. Our method also offers flexibility in adjusting the range of explanations within the input, i.e., from an overall attribution input to particular components within the input. Furthermore, it allows for the decomposition of explanations for a given input for more detailed explanations. Our experiments demonstrate that the proposed method outperforms existing methods both quantitatively and qualitatively.
Seongwoo Lim, Won Jo, Jaesik Choi
AISTATS4
2024 Memorizing Documents with Guidance in Large Language Models
Bumjin Park, Jaesik Choi
IJCAI2
2024 Towards Dynamic Trend Filtering through Trend Point Detection with Reinforcement Learning
Jihyeon Seong, Sekwang Oh, Jaesik Choi
IJCAI3
2023 Beyond Single Path Integrated Gradients for Reliable Input Attribution via Randomized Path Sampling
abstract
Input attribution is a widely used explanation method for deep neural networks, especially in visual tasks. Among various attribution methods, Integrated Gradients (IG) [28] is frequently used because of its model-agnostic applicability and desirable axioms. However, previous work [24], [8], [9] has shown that such method often produces noisy and unreliable attributions during the integration of the gradients over the path defined in the input space. In this paper, we tackle this issue by estimating the distribution of the possible attributions according to the integrating path selection. We show that such noisy attribution can be reduced by aggregating attributions from the multiple paths instead of using a single path. Inspired by Stick-Breaking Process [20], we suggest a random process to generate rich and various sampling of the gradient integrating path. Using multiple input attributions obtained from randomized path, we propose a novel attribution measure using the distribution of attributions at each input features. We identify proposed method qualitatively show less-noisy and object-aligned attribution and its feasibility through the quantitative evaluations.
Giyoung Jeon, Haedong Jeong, Jaesik Choi
ICCV3
2023 Rarity Score : A New Metric to Evaluate the Uncommonness of Synthesized Images
Jiyeon Han 0001, Hwanil Choi, Yunjey Choi, Jung-Woo Ha 0001, Jaesik Choi
ICLR6
2023 Variational Curriculum Reinforcement Learning for Unsupervised Discovery of Skills
abstract
Mutual information-based reinforcement learning (RL) has been proposed as a promising framework for retrieving complex skills autonomously without a task-oriented reward function through mutual information (MI) maximization or variational empowerment. However, learning complex skills is still challenging, due to the fact that the order of training skills can largely affect sample efficiency. Inspired by this, we recast variational empowerment as curriculum learning in goal-conditioned RL with an intrinsic reward function, which we name Variational Curriculum RL (VCRL). From this perspective, we propose a novel approach to unsupervised skill discovery based on information theory, called Value Uncertainty Variational Curriculum (VUVC). We prove that, under regularity conditions, VUVC accelerates the increase of entropy in the visited states compared to the uniform curriculum. We validate the effectiveness of our approach on complex navigation and robotic manipulation tasks in terms of sample efficiency and state coverage speed. We also demonstrate that the skills discovered by our method successfully complete a real-world robot navigation task in a zero-shot setup and that incorporating these skills with a global planner further increases the performance.
Seongun Kim, Kyowoon Lee, Jaesik Choi
ICML3
2023 Adaptive and Explainable Deployment of Navigation Skills via Hierarchical Deep Reinforcement Learning
abstract
For robotic vehicles to navigate robustly and safely in unseen environments, it is crucial to decide the most suitable navigation policy. However, most existing deep reinforcement learning based navigation policies are trained with a hand-engineered curriculum and reward function which are difficult to be deployed in a wide range of real-world scenarios. In this paper, we propose a framework to learn a family of low-level navigation policies and a high-level policy for deploying them. The main idea is that, instead of learning a single navigation policy with a fixed reward function, we simultaneously learn a family of policies that exhibit different behaviors with a wide range of reward functions. We then train the high-level policy which adaptively deploys the most suitable navigation skill. We evaluate our approach in simulation and the real world and demonstrate that our method can learn diverse navigation skills and adaptively deploy them. We also illustrate that our proposed hierarchical learning framework presents explainability by providing semantics for the behavior of an autonomous agent.
Kyowoon Lee, Seongun Kim, Jaesik Choi
ICRA3
2023 Algorithmic Read Resistance Trim for Improving Yield and Reducing Test Time in MRAM
abstract
MRAM is a one of the resistive memory which is based on a change in resistance of a unit cell. For successful read operation, resistive memory requires a reference resistance which has middle resistance values of data 0 and 1. Generally, the read method used in MRAM is to use a reference cell having the intermediate resistance. This conventional method has weakness on read disturb rate and temperature change. To avoid these weaknesses, new read scheme has been developed that sets reference resistances with controllable resistances instead of reference cells. But developed method has difficulty to set the stable reference resistance value. This paper presents the algorithmic trim method of finding the reference resistance by adjusting the controllable resistor. The number of failed bits for each read resistance was examined to find stable reference resistance conditions. In addition, a sampling method and a binary search method were introduced to improve test efficiency. The problem that occurs when applying binary search is solved by exception processing. And for parallel testing for each chip, the method that modifying the vector memory of the tester was used. Using algorithmic read resistance trim, yield was improved by 23.7%and test time was reduced by 80%. Also, test time of trim module was reduced by 95% compared to the beginning (<0.4 second). This trim method has been widespread in the MRAM product of Samsung Foundry, which adjusts the reference resistance with a controllable resistor.
Daehyun Chang, Youngdae Kim, Suk-Soo Pyo, Shin Hun, Daesop Lee, Sohee Hwang, Jaesik Choi, Siwoong Kim
ITC7
2023 Refining Diffusion Planner for Reliable Behavior Synthesis by Automatic Detection of Infeasible Plans
abstract
Diffusion-based planning has shown promising results in long-horizon, sparse-reward tasks by training trajectory diffusion models and conditioning the sampled trajectories using auxiliary guidance functions. However, due to their nature as generative models, diffusion models are not guaranteed to generate feasible plans, resulting in failed execution and precluding planners from being useful in safety-critical applications. In this work, we propose a novel approach to refine unreliable plans generated by diffusion models by providing refining guidance to error-prone plans. To this end, we suggest a new metric named restoration gap for evaluating the quality of individual plans generated by the diffusion model. A restoration gap is estimated by a gap predictor which produces restoration gap guidance to refine a diffusion planner. We additionally present an attribution map regularizer to prevent adversarial refining guidance that could be generated from the sub-optimal gap predictor, which enables further refinement of infeasible plans. We demonstrate the effectiveness of our approach on three different benchmarks in offline control settings that require long-horizon planning. We also illustrate that our approach presents explainability by presenting the attribution maps of the gap predictor and highlighting error-prone transitions, allowing for a deeper understanding of the generated plans.
Kyowoon Lee, Seongun Kim, Jaesik Choi
NeurIPS3
2022 An Unsupervised Way to Understand Artifact Generating Internal Units in Generative Neural Networks
abstract
Despite significant improvements on the image generation performance of Generative Adversarial Networks (GANs), generations with low visual fidelity still have been observed. As widely used metrics for GANs focus more on the overall performance of the model, evaluation on the quality of individual generations or detection of defective generations is challenging. While recent studies try to detect featuremap units that cause artifacts and evaluate individual samples, these approaches require additional resources such as external networks or a number of training data to approximate the real data manifold. In this work, we propose the concept of local activation, and devise a metric on the local activation to detect artifact generations without additional supervision. We empirically verify that our approach can detect and correct artifact generations from GANs with various datasets. Finally, we discuss a geometrical analysis to partially reveal the relation between the proposed concept and low visual fidelity.
Haedong Jeong, Jiyeon Han 0001, Jaesik Choi
AAAI3
2022 Can We Find Neurons that Cause Unrealistic Images in Deep Generative Networks?
abstract
Even though Generative Adversarial Networks (GANs) have shown a remarkable ability to generate high-quality images, GANs do not always guarantee the generation of photorealistic images. Occasionally, they generate images that have defective or unnatural objects, which are referred to as `artifacts'. Research to investigate why these artifacts emerge and how they can be detected and removed has yet to be sufficiently carried out. To analyze this, we first hypothesize that rarely activated neurons and frequently activated neurons have different purposes and responsibilities for the progress of generating images. In this study, by analyzing the statistics and the roles for those neurons, we empirically show that rarely activated neurons are related to the failure results of making diverse objects and inducing artifacts. In addition, we suggest a correction method, called `Sequential Ablation’, to repair the defective part of the generated images without high computational cost and manual efforts.
Hwanil Choi, Wonjoon Chang, Jaesik Choi
IJCAI3
2022 Distilled Gradient Aggregation: Purify Features for Input Attribution in the Deep Neural Network
abstract
Measuring the attribution of input features toward the model output is one of the popular post-hoc explanations on the Deep Neural Networks (DNNs). Among various approaches to compute the attribution, the gradient-based methods are widely used to generate attributions, because of its ease of implementation and the model-agnostic characteristic. However, existing gradient integration methods such as Integrated Gradients (IG) suffer from (1) the noisy attributions which cause the unreliability of the explanation, and (2) the selection for the integration path which determines the quality of explanations. FullGrad (FG) is an another approach to construct the reliable attributions by focusing the locality of piece-wise linear network with the bias gradient. Although FG has shown reasonable performance for the given input, as the shortage of the global property, FG is vulnerable to the small perturbation, while IG which includes the exploration over the input space is robust. In this work, we design a new input attribution method which adopt the strengths of both local and global attributions.In particular, we propose a novel approach to distill input features using weak and extremely positive contributor masks. We aggregate the intermediate local attributions obtained from the distillation sequence to provide reliable attribution. We perform the quantitative evaluation compared to various attribution methods and show that our method outperforms others. We also provide the qualitative result that our method obtains object-aligned and sharp attribution heatmap.
Giyoung Jeon, Haedong Jeong, Jaesik Choi
NeurIPS3
2022 Learning Fractional White Noises in Neural Stochastic Differential Equations
abstract
Differential equations play important roles in modeling complex physical systems. Recent advances present interesting research directions by combining differential equations with neural networks. By including noise, stochastic differential equations (SDEs) allows us to model data with uncertainty and measure imprecision. There are many variants of noises known to exist in many real-world data. For example, previously white noises are idealized and induced by Brownian motions. Nevertheless, there is a lack of machine learning models that can handle such noises. In this paper, we introduce a generalized fractional white noise to existing models and propose an efficient approximation of noise sample paths based on classical integration methods and sparse Gaussian processes. Our experimental results demonstrate that the proposed model can capture noise characteristics such as continuity from various time series data, therefore improving model fittings over existing models. We examine how we can apply our approach to score-based generative models, showing that there exists a case of our generalized noise resulting in a better image generation measure.
Anh Tong, Thanh Nguyen-Tang, Toan M. Tran, Jaesik Choi
NeurIPS4
2022 Semisupervised Training of Deep Generative Models for High-Dimensional Anomaly Detection
abstract
Abnormal behaviors in industrial systems may be early warnings on critical events that may cause severe damages to facilities and security. Thus, it is important to detect abnormal behaviors accurately and timely. However, the anomaly detection problem is hard to solve in practice, mainly due to the rareness and the expensive cost to get the labels of the anomalies. Deep generative models parameterized by neural networks have achieved state-of-the-art performance in practice for many unsupervised and semisupervised learning tasks. We present a new deep generative model, Latent Enhanced regression/classification Deep Generative Model (LEDGM), for the anomaly detection problem with multidimensional data. Instead of using two-stage decoupled models, we adopt an end-to-end learning paradigm. Instead of conditioning the latent on the class label, LEDGM conditions the label prediction on the learned latent so that the optimization goal is more in favor of better anomaly detection than better reconstruction that the previously proposed deep generative models have been trained for. Experimental results on several synthetic and real-world small- and large-scale datasets demonstrate that LEDGM can achieve improved anomaly detection performance on multidimensional data with very sparse labels. The results also suggest that both labeled anomalies and labeled normal are valuable for semisupervised learning. Generally, our results show that better performance can be achieved with more labeled data. The ablation experiments show that both the original input and the learned latent provide meaningful information for LEDGM to achieve high performance.
Peng Zhang 0030, Boseon Yu, Jaesik Choi
IEEE Trans. Neural Networks Learn. Syst.4
2021 Interpreting Deep Neural Networks with Relative Sectional Propagation by Analyzing Comparative Gradients and Hostile Activations
abstract
The clear transparency of Deep Neural Networks (DNNs) is hampered by complex internal structures and nonlinear transformations along deep hierarchies. In this paper, we propose a new attribution method, Relative Sectional Propagation (RSP), for fully decomposing the output predictions with the characteristics of class-discriminative attributions and clear objectness. We carefully revisit some shortcomings of backpropagation-based attribution methods, which are trade-off relations in decomposing DNNs. We define hostile factor as an element that interferes with finding the attributions of the target and propagate it in a distinguishable way to overcome the non-suppressed nature of activated neurons. As a result, it is possible to assign the bi-polar relevance scores of the target (positive) and hostile (negative) attributions while maintaining each attribution aligned with the importance. We also present the purging techniques to prevent the decrement of the gap between the relevance scores of the target and hostile attributions during backward propagation by eliminating the conflicting units to channel attribution map. Therefore, our method makes it possible to decompose the predictions of DNNs with clearer class-discriminativeness and detailed elucidations of activation neurons compared to the conventional attribution methods. In a verified experimental environment, we report the results of the assessments: (i) Pointing Game, (ii) mIoU, and (iii) Model Sensitivity with PASCAL VOC 2007, MS COCO 2014, and ImageNet datasets. The results demonstrate that our method outperforms existing backward decomposition methods, including distinctive and intuitive visualizations.
Woo-Jeoung Nam, Jaesik Choi, Seong-Whan Lee
AAAI2
2021 Characterizing Deep Gaussian Processes via Nonlinear Recurrence Systems
Anh Tong, Jaesik Choi
AAAI2
2021 Learning Compositional Sparse Gaussian Processes with a Shrinkage Prior
abstract
Choosing a proper set of kernel functions is an important problem in learning Gaussian Process (GP) models since each kernel structure has different model complexity and data fitness. Recently, automatic kernel composition methods provide not only accurate prediction but also attractive interpretability through search-based methods. However, existing methods suffer from slow kernel composition learning. To tackle large-scaled data, we propose a new sparse approximate posterior for GPs, MultiSVGP, constructed from groups of inducing points associated with individual additive kernels in compositional kernels. We demonstrate that this approximation provides a better fit to learn compositional kernels given empirical observations. We also provide theoretically justification on error bound when compared to the traditional sparse GP. In contrast to the search-based approach, we present a novel probabilistic algorithm to learn a kernel composition by handling the sparsity in the kernel selection with Horseshoe prior. We demonstrate that our model can capture characteristics of time series with significant reductions in computational time and have competitive regression performance on real-world data sets.
Anh Tong, Toan M. Tran, Jaesik Choi
AAAI4
2021 Automatic Correction of Internal Units in Generative Neural Networks
abstract
Generative Adversarial Networks (GANs) have shown satisfactory performance in synthetic image generation by devising complex network structure and adversarial training scheme. Even though GANs are able to synthesize realistic images, there exists a number of generated images with defective visual patterns which are known as artifacts. While most of the recent work tries to fix artifact generations by perturbing latent code, few investigate internal units of a generator to fix them. In this work, we devise a method that automatically identifies the internal units generating various types of artifact images. We further propose the sequential correction algorithm which adjusts the generation flow by modifying the detected artifact units to improve the quality of generation while preserving the original outline. Our method outperforms the baseline method in terms of FID-score and shows satisfactory results with human evaluation.
Ali Tousi, Haedong Jeong, Jiyeon Han 0001, Hwanil Choi, Jaesik Choi
CVPR5
2021 Conditional Temporal Neural Processes with Covariance Loss
abstract
We introduce a novel loss function, Covariance Loss, which is conceptually equivalent to conditional neural processes and has a form of regularization so that is applicable to many kinds of neural networks. With the proposed loss, mappings from input variables to target variables are highly affected by dependencies of target variables as well as mean activation and mean dependencies of input and target variables. This nature enables the resulting neural networks to become more robust to noisy observations and recapture missing dependencies from prior information. In order to show the validity of the proposed loss, we conduct extensive sets of experiments on real-world datasets with state-of-the-art models and discuss the benefits and drawbacks of the proposed Covariance Loss.
Boseon Yoo, Jiwoo Lee, Janghoon Ju, Seijun Chung, Jaesik Choi
ICML6
2021 Explaining the Decisions of Deep Policy Networks for Robotic Manipulations
abstract
Deep policy networks enable robots to learn behaviors to solve various real-world complex tasks in an end-to-end fashion. However, they lack transparency to provide the reasons of actions. Thus, such a black-box model often results in low reliability and disruptive actions during the deployment of the robot in practice. To enhance its transparency, it is important to explain robot behaviors by considering the extent to which each input feature contributes to determining a given action. In this paper, we present an explicit analysis of deep policy models through input attribution methods to explain how and to what extent each input feature affects the decisions of the robot policy models. To this end, we present two methods for applying input attribution methods to robot policy networks: (1) we measure the importance factor of each joint torque to re ect the influence of the motor torque on the end-effector movement, and (2) we modify a relevance propagation method to handle negative inputs and outputs in deep policy networks properly. To the best of our knowledge, this is the first report to identify the dynamic changes of input attributions of multi-modal sensor inputs in deep policy networks online for robotic manipulation.
Seongun Kim, Jaesik Choi
IROS2
2021 Interpreting Internal Activation Patterns in Deep Temporal Neural Networks by Finding Prototypes
abstract
Deep neural networks have demonstrated competitive performance in classification tasks for sequential data. However, it remains difficult to understand which temporal patterns the internal channels of deep neural networks capture for decision-making in sequential data. To address this issue, we propose a new framework with which to visualize temporal representations learned in deep neural networks without hand-crafted segmentation labels. Given input data, our framework extracts highly activated temporal regions that contribute to activating internal nodes and characterizes such regions by prototype selection method based on Maximum Mean Discrepancy. Representative temporal patterns referred to here as Prototypes of Temporally Activated Patterns (PTAP) provide core examples of subsequences in the sequential data for interpretability. We also analyze the role of each channel by Value-LRP plots using representative prototypes and the distribution of the input attribution. Input attribution plots give visual information to recognize the shapes focused on by the channel for decision-making.
Sohee Cho, Wonjoon Chang, Ginkyeng Lee, Jaesik Choi
KDD4
2020 An Efficient Explorative Sampling Considering the Generative Boundaries of Deep Generative Neural Networks
abstract
Deep generative neural networks (DGNNs) have achieved realistic and high-quality data generation. In particular, the adversarial training scheme has been applied to many DGNNs and has exhibited powerful performance. Despite of recent advances in generative networks, identifying the image generation mechanism still remains challenging. In this paper, we present an explorative sampling algorithm to analyze generation mechanism of DGNNs. Our method efficiently obtains samples with identical attributes from a query image in a perspective of the trained model. We define generative boundaries which determine the activation of nodes in the internal layer and probe inside the model with this information. To handle a large number of boundaries, we obtain the essential set of boundaries using optimization. By gathering samples within the region surrounded by generative boundaries, we can empirically reveal the characteristics of the internal layers of DGNNs. We also demonstrate that our algorithm can find more homogeneous, the model specific samples compared to the variations of ϵ-based sampling method.
Giyoung Jeon, Haedong Jeong, Jaesik Choi
AAAI3
2020 Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural Networks
abstract
As Deep Neural Networks (DNNs) have demonstrated superhuman performance in a variety of fields, there is an increasing interest in understanding the complex internal mechanisms of DNNs. In this paper, we propose Relative Attributing Propagation (RAP), which decomposes the output predictions of DNNs with a new perspective of separating the relevant (positive) and irrelevant (negative) attributions according to the relative influence between the layers. The relevance of each neuron is identified with respect to its degree of contribution, separated into positive and negative, while preserving the conservation rule. Considering the relevance assigned to neurons in terms of relative priority, RAP allows each neuron to be assigned with a bi-polar importance score concerning the output: from highly relevant to highly irrelevant. Therefore, our method makes it possible to interpret DNNs with much clearer and attentive visualizations of the separated attributions than the conventional explaining methods. To verify that the attributions propagated by RAP correctly account for each meaning, we utilize the evaluation metrics: (i) Outside-inside relevance ratio, (ii) Segmentation mIOU and (iii) Region perturbation. In all experiments and metrics, we present a sizable gap in comparison to the existing literature.
Woo-Jeoung Nam, Shir Gur, Jaesik Choi, Lior Wolf, Seong-Whan Lee
AAAI3
2020 A Single Multi-Task Deep Neural Network with Post-Processing for Object Detection with Reasoning and Robotic Grasp Detection
abstract
Applications of deep neural network (DNN) based object and grasp detections could be expanded significantly when the network output is processed by a high-level reasoning over relationship of objects. Recently, robotic grasp detection and object detection with reasoning have been investigated using DNNs. There have been efforts to combine these multitasks using separate networks so that robots can deal with situations of grasping specific target objects in the cluttered, stacked, complex piles of novel objects from a single RGB-D camera. We propose a single multi-task DNN that yields accurate detections of objects, grasp position and relationship reasoning among objects. Our proposed methods yield state-of-the-art performance with the accuracy of 98.6% and 74.2% with the computation speed of 33 and 62 frame per second on VMRD and Cornell datasets, respectively. Our methods also yielded 95.3% grasp success rate for novel object grasping tasks with a 4-axis robot arm and 86.7% grasp success rate in cluttered novel objects with a humanoid robot.
Dongwon Park, Yonghyeok Seo, Dongju Shin, Jaesik Choi, Se Young Chun
ICRA4
2020 Interpreting and Explaining Deep Neural Networks: A Perspective on Time Series Data
abstract
Explainable and interpretable machine learning models and algorithms are important topics which have received growing attention from research, application and administration. Many complex Deep Neural Networks (DNNs) are often perceived as black-boxes. Researchers would like to be able to interpret what the DNN has learned in order to identify biases and failure models and improve models. In this tutorial, we will provide a comprehensive overview on methods to analyze deep neural networks and an insight how those interpretable and explainable methods help us understand time series data.
Jaesik Choi
KDD1
2020 HetPipe: Enabling Large DNN Training on (Whimpy) Heterogeneous GPU Clusters through Integration of Pipelined Model Parallelism and Data Parallelism
Jay H. Park, Gyeongchan Yun, Chang M. Yi, Nguyen T. Nguyen, Jaesik Choi, Sam H. Noh, Young-ri Choi
USENIX ATC6
2019 Discovering Latent Covariance Structures for Multiple Time Series
abstract
Analyzing multivariate time series data is important to predict future events and changes of complex systems in finance, manufacturing, and administrative decisions. The expressiveness power of Gaussian Process (GP) regression methods has been significantly improved by compositional covariance structures. In this paper, we present a new GP model which naturally handles multiple time series by placing an Indian Buffet Process (IBP) prior on the presence of shared kernels. Our selective covariance structure decomposition allows exploiting shared parameters over a set of multiple, selected time series. We also investigate the well-definedness of the models when infinite latent components are introduced. We present a pragmatic search algorithm which explores a larger structure space efficiently. Experiments conducted on five real-world data sets demonstrate that our new model outperforms existing methods in term of structure discoveries and predictive performances.
Anh Tong, Jaesik Choi
ICML2
2019 Confirmatory Bayesian Online Change Point Detection in the Covariance Structure of Gaussian Processes
abstract
In the analysis of sequential data, the detection of abrupt changes is important in predicting future events. In this paper, we propose statistical hypothesis tests for detecting covariance structure changes in locally smooth time series modeled by Gaussian Processes (GPs). We provide theoretically justified thresholds for the tests, and use them to improve Bayesian Online Change Point Detection (BOCPD) by confirming statistically significant changes and non-changes. Our Confirmatory BOCPD (CBOCPD) algorithm finds multiple structural breaks in GPs even when hyperparameters are not tuned precisely. We also provide conditions under which CBOCPD provides the lower prediction error compared to BOCPD. Experimental results on synthetic and real-world datasets show that our proposed algorithm outperforms existing methods for the prediction of nonstationarity in terms of both regression error and log-likelihood.
Jiyeon Han 0001, Kyowoon Lee, Anh Tong, Jaesik Choi
IJCAI4
2018 Dynamic Online Performance Optimization in Streaming Data Compression
abstract
Compression is essential to high bandwidth applications such as scientific simulations and sensing applications to reduce resource burden such as storage, network transmission, and more recently I/O. Existing lossy compression methods attempt to minimize the Euclidean distance between original data and reconstructed data, which significantly limits either compression performance or reconstruction quality since original and reconstructed data sequences should be aligned. Substituting the Euclidean distance for a statistical similarity maximizes the compression performance while retaining essential data features. By implementing this methodology, IDEALEM has recently demonstrated compression ratios far exceeding 100:1, better than best-known compression methods, while preserving reconstruction quality. This work proposes an online algorithm for streaming data compression which takes account of generally concave trend of compression ratio curve, and optimizes key operation parameters. We demonstrate that the proposed algorithm successfully adapts one of the key parameters in IDEALEM to the optimal value and yields near maximum compression ratios for time series data.
J. Kade Gibson, Dongeun Lee 0001, Jaesik Choi, Alex Sim
IEEE BigData3
2018 Deep Reinforcement Learning in Continuous Action Spaces: a Case Study in the Game of Simulated Curling
abstract
Many real-world applications of reinforcement learning require an agent to select optimal actions from continuous spaces. Recently, deep neural networks have successfully been applied to games with discrete actions spaces. However, deep neural networks for discrete actions are not suitable for devising strategies for games where a very small change in an action can dramatically affect the outcome. In this paper, we present a new self-play reinforcement learning framework which equips a continuous search algorithm which enables to search in continuous action spaces with a kernel regression method. Without any hand-crafted features, our network is trained by supervised learning followed by self-play reinforcement learning with a high-fidelity simulator for the Olympic sport of curling. The program trained under our framework outperforms existing programs equipped with several hand-crafted features and won an international digital curling competition.
Kyowoon Lee, Sol-A. Kim, Jaesik Choi, Seong-Whan Lee
ICML3
2017 Expanding Statistical Similarity Based Data Reduction to Capture Diverse Patterns
abstract
We propose a new class of lossy compression based on locally exchangeable measure that captures the distribution of repeating data blocks while preserving unique patterns. The technique has been demonstrated to reduce data volume by more than 100-fold on power grid monitoring data where a large number of data blocks can be characterized as following stationary probability distributions. To capture data with more diverse patterns, we propose two techniques to transform non-stationary time series into locally stationary blocks. We also propose a strategy to work with values in bounded ranges such as phase angles of alternating current. These new ideas are incorporated into a software package named IDEALEM. In experiments, IDEALEM reduces non-stationary data volume up to 100-fold. Compared with the state-of-the-art lossy compression methods such as SZ, IDEALEM can produce more compact output overall.
Dongeun Lee 0001, Alex Sim, Jaesik Choi, Kesheng Wu
DCC3
2017 Improving Statistical Similarity Based Data Reduction for Non-Stationary Data
abstract
We propose a new class of lossy compression based on locally exchangeable measure that captures the distribution of repeating data blocks while preserving unique patterns. The technique has been demonstrated to reduce data volume by more than 100-fold on power grid monitoring data where a large number of data blocks can be characterized as following stationary probability distributions. To capture data with more diverse patterns, we propose two techniques to transform non-stationary time series into locally stationary blocks. We also propose a strategy to work with values in bounded ranges such as phase angles of alternating current. These new ideas are incorporated into a software package named IDEALEM. In experiments, IDEALEM reduces non-stationary data volume up to 100-fold. Compared with the state-of-the-art lossy compression methods such as SZ, IDEALEM can produce more compact output overall.
Dongeun Lee 0001, Alex Sim, Jaesik Choi, Kesheng Wu
SSDBM3
2016 Global Deconvolutional Networks for Semantic Segmentation
Vladimir Nekrasov, Janghoon Ju, Jaesik Choi
BMVC3
2016 Automatic Construction of Nonparametric Relational Regression Models for Multiple Time Series
abstract
Gaussian Processes (GPs) provide a general and analytically tractable way of modeling complex time-varying, nonparametric functions. The Automatic Bayesian Covariance Discovery (ABCD) system constructs natural-language description of time-series data by treating unknown time-series data nonparametrically using GP with a composite covariance kernel function. Unfortunately, learning a composite covariance kernel with a single time-series data set often results in less informative kernel that may not give qualitative, distinctive descriptions of data. We address this challenge by proposing two relational kernel learning methods which can model multiple time-series data sets by finding common, shared causes of changes. We show that the relational kernel learning methods find more accurate models for regression problems on several real-world data sets; US stock data, US house price index data and currency exchange rate data.
Yunseong Hwang, Anh Tong, Jaesik Choi
ICML3
2016 Novel Data Reduction Based on Statistical Similarity
abstract
Applications such as scientific simulations and power grid monitoring are generating so much data quickly that compression is essential to reduce storage requirement or transmission capacity. To achieve better compression, one is often willing to discard some repeated information. These lossy compression methods are primarily designed to minimize the Euclidean distance between the original data and the compressed data. But this measure of distance severely limits either reconstruction quality or compression performance. We propose a new class of compression method by redefining the distance measure with a statistical concept known as exchangeability. This approach reduces the storage requirement and captures essential features, while reducing the storage requirement. In this paper, we report our design and implementation of such a compression method named IDEALEM. To demonstrate its effectiveness, we apply it on a set of power grid monitoring data, and show that it can reduce the volume of data much more than the best known compression method while maintaining the quality of the compressed data. In these tests, IDEALEM captures extraordinary events in the data, while its compression ratios can far exceed 100.
Dongeun Lee 0001, Alex Sim, Jaesik Choi, Kesheng Wu
SSDBM3
2016 Improving Imprecise Compressive Sensing Models
Dongeun Lee 0001, Rafael Lima, Jaesik Choi
UAI3
2015 Learning Relational Kalman Filtering
abstract
The Kalman Filter (KF) is pervasively used to control a vast array of consumer, health and defense products. By grouping sets of symmetric state variables, the Relational Kalman Filter (RKF) enables us to scale the exact KF for large-scale dynamic systems. In this paper, we provide a parameter learning algorithm for RKF, and a regrouping algorithm that prevents the degeneration of the relational structure for efficient filtering. The proposed algorithms significantly expand the applicability of the RKFs by solving the following questions: (1) how to learn parameters for RKF from partial observations; and (2) how to regroup the degenerated state variables by noisy real-world observations. To our knowledge, this is the first paper on learning parameters in relational continuous probabilistic models. We show that our new algorithms significantly improve the accuracy and the efficiency of filtering large-scale dynamic systems.
Jaesik Choi, Eyal Amir, Tianfang Xu, Albert J. Valocchi
AAAI1
2015 Memory heat map: anomaly detection in real-time embedded systems using memory behavior
abstract
In this paper, we introduce a novel mechanism that identifies abnormal system-wide behaviors using the predictable nature of real-time embedded applications. We introduce Memory Heat Map (MHM) to characterize the memory behavior of the operating system. Our machine learning algorithms automatically (a) summarize the information contained in the MHMs and then (b) detect deviations from the normal memory behavior patterns. These methods are implemented on top of a multicore processor architecture to aid in the process of monitoring and detection. The techniques are evaluated using multiple attack scenarios including kernel rootkits and shellcode. To the best of our knowledge, this is the first work that uses aggregated memory behavior for detecting system anomalies especially the concept of memory heat maps.
Man-Ki Yoon, Lui Sha, Sibin Mohan, Jaesik Choi
DAC4
2015 Reading Documents for Bayesian Online Change Point Detection
abstract
Modeling non-stationary time-series data for making predictions is a challenging but important task.One of the key issues is to identify long-term changes accurately in time-varying data.Bayesian Online Change Point Detection (BO-CPD) algorithms efficiently detect long-term changes without assuming the Markov property which is vulnerable to local signal noise.We propose a Document based BO-CPD (DBO-CPD) model which automatically detects long-term temporal changes of continuous variables based on a novel dynamic Bayesian analysis which combines a non-parametric regression, the Gaussian Process (GP), with generative models of texts such as news articles and posts on social networks.Since texts often include important clues of signal changes, DBO-CPD enables the accurate prediction of long-term changes accurately.We show that our algorithm outperforms existing BO-CPDs in two real-world datasets: stock prices and movie revenues.
Jaesik Choi
EMNLP2
2015 A Deterministic Partition Function Approximation for Exponential Random Graph Models
Wen Pu, Jaesik Choi, Yunseong Hwang, Eyal Amir
IJCAI2
2015 Learning Compressive Sensing Models for Big Spatio-Temporal Data
abstract
Sensing devices including mobile phones and biomedical sensors generate massive amounts of spatio-temporal data. Compressive sensing (CS) can significantly reduce energy and resource consumption by shifting the complexity burden of encoding process to the decoder. CS reconstructs the compressed signals exactly with overwhelming probability when incoming data can be sparsely represented with a fixed number of components, which is one of drawbacks of CS frameworks because a real-world signal in general cannot be represented with the fixed number of components. We present the first CS framework that handles signals without the fixed sparsity assumption by incorporating the distribution of the number of principal components included in the signal recovery, which we show is naturally represented by the gamma distribution. This allows an analytic derivation of total error in our spatio-temporal Low Complexity Sampling (LCS). We show that LCS requires shorter compressed signals than existing CS frameworks to bound the same amount of error. Experiments with real-world sensor data also demonstrate that LCS outperforms existing CS frameworks.
Dongeun Lee 0001, Jaesik Choi
SDM2
2015 Semi-Local Structure Patterns for Robust Face Detection
abstract
In many image processing and computer vision problems, including face detection, local structure patterns such as local binary patterns (LBP) and modified census transform (MCT) have been adopted in widespread applications due to their robustness against illumination changes. However, being reliant on the local differences between neighboring pixels, they are inevitably sensitive to noise. To overcome the problem of noise-vulnerability of the conventional local structure patterns, we propose semi-local structure patterns (SLSP), a novel feature extraction method based on local region-based differences. The SLSP is robust to illumination variations, distortion, and sparse noise because it encodes the relative sizes of the central region with locally neighboring regions into a binary code. The principle of SLSP leads noise-robust expansions of LBP and MCT feature extraction frameworks. In a statistical analysis, we find that the proposed methods transform a substantial amount of random noise patterns in face images into more meaningful uniform patterns. The empirical results on the MIT + CMU dataset and FDDB (face detection dataset and benchmark) show that the proposed semi-local patterns applied to LBP and MCT feature extraction frameworks outperform the conventional LBP and MCT features in AdaBoost-based face detectors, with much higher detection rates.
Kyungjoong Jeong, Jaesik Choi, Gil-Jin Jang
IEEE Signal Process. Lett.2
2014 Low complexity sensing for big spatio-temporal data
abstract
Many large scale sensor networks produce tremendous data, typically as massive spatio-temporal data streams. We present a Low Complexity Sensing framework that, coupled with novel compressive sensing techniques, enables to reduce computational and communication overheads significantly without much compromising the accuracy of sensor readings. More specifically, our sensing framework randomly samples time-series data in the temporal dimension first, then in the spatial dimension. Under some mild conditions, our sensing framework holds the same theoretical bound of reconstruction error, but is much simpler and easier to implement than existing compressive sensing frameworks. In experiments with real world environmental data sets, we demonstrate that the proposed framework outperforms two existing compressive sensing frameworks designed for spatio-temporal data.
Dongeun Lee 0001, Jaesik Choi
IEEE BigData2
2013 Fast Change Point Detection for electricity market analysis
abstract
Electricity is a vital part of our daily life; therefore it is important to avoid irregularities such as the California Electricity Crisis of 2000 and 2001. In this work, we seek to predict anomalies using advanced machine learning algorithms, more specifically a Change Point Detection (CPD) algorithm on the electricity prices during the California Electricity Crisis. Such algorithms are effective, but computationally expensive when applied on a large amount of data. To address this challenge, we accelerate the Gaussian Process (GP) for 1-dimensional time series data. Since GP is at the core of many statistical learning techniques, this improvement could benefit many algorithms. In the specific Change Point Detection algorithm used in this study, we reduce the overall computational complexity from O(n5) to O(n2), where the amountized cost of solving a GP projet is O(1). Our efficient algorithm makes it possible to compute the Change Points using the hourly price data during the California Electricity Crisis. By comparing the detected Change Points with known events, we show that the Change Point Detection algorithm is indeed effective in detecting signals preceding major events.
William Gu, Jaesik Choi, Ming Gu 0002, Horst D. Simon, Kesheng Wu
IEEE BigData2
2013 SecureCore: A multicore-based intrusion detection architecture for real-time embedded systems
abstract
Security violations are becoming more common in real-time systems - an area that was considered to be invulnerable in the past - as evidenced by the recent W32.Stuxnet and Duqu worms. A failure to protect such systems from malicious entities could result in significant harm to both humans as well as the environment. The increasing use of multicore architectures in such systems exacerbates the problem since shared resources on these processors increase the risk of being compromised. In this paper, we present the SecureCore framework that, coupled with novel monitoring techniques, is able to improve the security of realtime embedded systems. We aim to detect malicious activities by analyzing and observing the inherent properties of the real-time system using statistical analyses of their execution profiles. With careful analysis based on these profiles, we are able to detect malicious code execution as soon as it happens and also ensure that the physical system remains safe.
Man-Ki Yoon, Sibin Mohan, Jaesik Choi, Jung-Eun Kim, Lui Sha
IEEE Real-Time and Embedded Technology and Applications Symposium3
2013 A spatio-temporal pyramid matching for video retrieval
Jaesik Choi, Ziyu Wang 0005, Won Jong Jeon
Comput. Vis. Image Underst.1
2012 Lifted Relational Variational Inference
Jaesik Choi, Eyal Amir
UAI1
2011 Efficient Methods for Lifted Inference with Aggregate Factors
abstract
Aggregate factors (that is, those based on aggregate functions such as SUM, AVERAGE, AND etc.) in probabilistic relational models can compactly represent dependencies among a large number of relational random variables. However, propositional inference on a factor aggregating n k-valued random variables into an r-valued result random variable is O(r k 2n). Lifted methods can ameliorate this to O(r nk) in general and O(r k log n) for commutative associative aggregators. In this paper, we propose (a) an exact solution constant in n when k = 2 for certain aggregate operations such as AND, OR and SUM, and (b) a close approximation for inference with aggregate factors with time complexity constant in n. This approximate inference involves an analytical solution for some operations when k > 2. The approximation is based on the fact that the typically used aggregate functions can be represented by linear constraints in the standard (k –1)-simplex in Rk where k is the number of possible values for random variables. This includes even aggregate functions that are commutative but not associative (e.g., the MODE operator that chooses the most frequent value). Our algorithm takes polynomial time in k (which is only 2 for binary variables) regardless of r and n, and the error decreases as n increases. Therefore, for most applications (in which a close approximation suffices) our algorithm is a much more efficient solution than existing algorithms. We present experimental results supporting these claims. We also present a (c) third contribution which further optimizes aggregations over multiple groups of random variables with distinct distributions.
Jaesik Choi, Rodrigo de Salvo Braz, Hung Hai Bui
AAAI1
2011 Lifted Relational Kalman Filtering
Jaesik Choi, Abner Guzmán-Rivera, Eyal Amir
IJCAI1
2010 Lifted Inference for Relational Continuous Models
Jaesik Choi, Eyal Amir, David J. Hill 0002
UAI1
2009 Combining planning and motion planning
abstract
Robotic manipulation is important for real, physical world applications. General Purpose manipulation with a robot (eg. delivering dishes, opening doors with a key, etc.) is demanding. It is hard because (1) objects are constrained in position and orientation, (2) many non-spatial constraints interact (or interfere) with each other, and (3) robots may have multi-degree of freedoms (DOF). In this paper we solve the problem of general purpose robotic manipulation using a novel combination of planning and motion planning. Our approach integrates motions of a robot with other (non-physical or external-to-robot) actions to achieve a goal while manipulating objects. It differs from previous, hierarchical approaches in that (a) it considers kinematic constraints in configuration space (C-space) together with constraints over object manipulations; (b) it automatically generates high-level (logical) actions from a C-space based motion planning algorithm; and (c) it decomposes a planning problem into small segments, thus reducing the complexity of planning.
Jaesik Choi, Eyal Amir
ICRA1
2009 Greedy Algorithms for Sequential Sensing Decisions
Hannaneh Hajishirzi, Afsaneh Shirazi, Jaesik Choi, Eyal Amir
IJCAI3
2007 Factor-guided motion planning for a robot arm
abstract
Motion planning for robotic arms is important for real, physical world applications. The planning for arms with high-degree-of-freedom (DOF) is hard because its search space is large (exponential in the number of joints), and the links may collide with static obstacles or other joints (self-collision). In this paper we present a motion planning algorithm that finds plans of motion from one arm configuration to a goal arm configuration in 2D space assuming no self-collision. Our algorithm is unique in two ways: (a) it utilizes the topology of the arm and obstacles to factor the search space and reduce the complexity of the planning problem using dynamic programming; and (b) it takes only polynomial time in the number of joints under some conditions. We provide a sufficient condition for polytime motion planning for 2D-space arms: if there is a path between two homotopic configurations, an embedded local planner finds a path within a polynomial time. The experimental results show that the proposed algorithm improves the performance of path planning for 2D arms.
Jaesik Choi, Eyal Amir
IROS1