Young Joon Yoo

dblp:146/4031 · also Youngjoon Yoo · DBLP profile ↗
← Back
34ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation
abstract
Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan.Existing approaches augment reasoning steps or inject specific structure into how models think, but leave scattered problem conditions unchanged.Inspired by the way humans organize fragmented information into coherent explanations, we propose STORYCODER, a narrative reformulation framework that transforms code generation questions into coherent natural language narratives, providing richer contextual structure than simple rephrasings.Each narrative consists of three components: a task overview, constraints, and example test cases, guided by the selected algorithm and genre.Experiments across 11 models on HumanEval, LiveCodeBench, and CodeForces demonstrate consistent improvements, with an average gain of 18.7% in zero-shot [email protected] accuracy, our analyses reveal that narrative reformulation guides models toward correct algorithmic strategies, reduces implementation errors, and induces a more modular code structure.The analyses further show that these benefits depend on narrative coherence and genre alignment, suggesting that structured problem representation is important for code generation regardless of model scale or architecture.Our code is available here.
Geonhui Jang, Dongyoon Han, Young Joon Yoo
ACL (1)3
2026 Continual-MEGA: A large-scale benchmark for generalizable continual anomaly detection
abstract
Anomaly detection is essential for visual inspection systems that require reliable identification of defective or abnormal samples. In real-world deployment, however, object categories and defect types can change over time, making static evaluation settings insufficient for assessing long-term adaptability and generalization. To address this limitation, we introduce a new benchmark for continual learning in anomaly detection. Our benchmark, Continual-MEGA, expands existing evaluation settings by integrating diverse public datasets with our newly proposed dataset, ContinualAD. Beyond standard continual learning settings, we additionally propose a scenario that evaluates zero-shot generalization to unseen classes not encountered during continual adaptation. This reflects recent advances in continual zero-shot research and highlights its practical significance. This setting introduces a new agenda for the anomaly detection field, and we conduct extensive evaluations of various existing anomaly detection algorithms designed for continual or zero-shot scenarios, as well as our proposed baseline methods. From our experiments, we derive three key findings: (1) existing methods exhibit significant limitations, particularly in pixel-level defect localization, (2) the proposed ContinualAD dataset is effective for the proposed benchmarking scenario, and (3) our baseline method suggests a promising direction for designing CLIP-based continual and generalizable frameworks through simple adaptation combined with feature synthesis.
Geonu Lee, Yujeong Oh, Geonhui Jang, Jeonghyo Song, Sungmin Cha, Young Joon Yoo
Neurocomputing7
2026 Domain-generalizable face anti-spoofing with patch-based multi-tasking and artifact pattern conversion
Seungjin Jung, Yonghyun Jeong, Minha Kim, Jimin Min, Young Joon Yoo, Jongwon Choi 0002
Pattern Recognit.5
2025 Distributional Uncertainty for Out-of-Distribution Detection
abstract
Estimating uncertainty from deep neural networks is a widely used approach for detecting out-of-distribution (OoD) samples, which typically exhibit high predictive uncertainty. However, conventional methods such as Monte Carlo (MC) Dropout often focus solely on either model or data uncertainty, failing to align with the semantic objective of OoD detection. To address this, we propose the Free-Energy Posterior Network, a novel framework that jointly models distributional uncertainty and identifying OoD and misclassified regions using free energy. Our method introduces two key contributions: (1) a free-energy-based density estimator parameterized by a Beta distribution, which enables fine-grained uncertainty estimation near ambiguous or unseen regions; and (2) a loss integrated within a posterior network, allowing direct uncertainty estimation from learned parameters without requiring stochastic sampling. By integrating our approach with the residual prediction branch (RPL) framework, the proposed method goes beyond post-hoc energy thresholding and enables the network to learn OoD regions by leveraging the variance of the Beta distribution, resulting in a semantically meaningful and computationally efficient solution for uncertainty-aware segmentation. We validate the effectiveness of our method on challenging real-world benchmarks, including Fishyscapes, RoadAnomaly, and Segment-Me-If-You-Can.
Kimin Yun, Jeonghyo Song, Young Joon Yoo
AVSS5
2025 CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning
abstract
Effective Out-of-Distribution (OOD) detection is critical for ensuring the reliability of semantic segmentation models, particularly in complex road environments where safety and accuracy are paramount. Despite recent advancements in large language models (LLMs), notably GPT-4, which significantly enhanced multimodal reasoning through Chain-of-Thought (CoT) prompting, the application of CoT-based visual reasoning for OOD semantic segmentation remains largely unexplored. In this paper, through extensive analyses of the road scene anomalies, we identify three challenging scenarios where current state-of-the-art OOD segmentation methods consistently struggle: (1) densely packed and overlapping objects, (2) distant scenes with small objects, and (3) large foreground-dominant objects. To address the presented challenges, we propose a novel CoT-based framework targeting OOD detection in road anomaly scenes. Our method leverages the extensive knowledge and reasoning capabilities of foundation models, such as GPT-4, to enhance OOD detection through improved image understanding and prompt-based reasoning aligned with observed problematic scene attributes. Extensive experiments show that our framework consistently outperforms state-of-the-art methods on both standard benchmarks and our newly defined challenging subset of the RoadAnomaly dataset, offering a robust and interpretable solution for OOD semantic segmentation in complex driving environments.
Jeonghyo Song, Kimin Yun, Young Joon Yoo
AVSS5
2025 Domain-Generalized Object Anti-Spoofing: Bridging Gaps and Patch Selection for Robust Detection Across Domains
abstract
In online applications, significant risks exist in peer-to-peer transactions due to malicious behaviors of arbitrary users, such as taking advantage of manipulated images or impersonating others using recaptured images. Moreover, recent advancements in display screens and imaging devices have made it increasingly challenging to distinguish such spoofing images from the naked eye. However, a lack of datasets for object anti-spoofing significantly hinders the practical implementation of object anti-spoofing techniques compared to facial anti-spoofing tasks. To address this data scarcity issue for object anti-spoofing, we propose a method that utilizes face anti-spoofing images for training. Our approach leverages low-rank adaptation, employing fine-tuning with downstream tasks of large language models to facilitate domain transition between faces and generic objects. We also analyze a power spectrum to select useful patches for spoofing detection and introduce a patch-based learning method to effectively capture spoofing patterns. Lastly, we present a novel protocol for assessing domain generalization in the generic object anti-spoofing task. Our model demonstrates state-of-the-art generalization performance compared to existing object anti-spoofing models, surpassing even those simply augmented with face datasets.
Geonu Lee, Yonghyun Jeong, Haneol Jang, Young Joon Yoo
WACV4
2024 Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding
abstract
In the weakly supervised temporal video grounding study, previous methods use predetermined single Gaussian proposals which lack the ability to express diverse events described by the sentence query. To enhance the expression ability of a proposal, we propose a Gaussian mixture proposal (GMP) that can depict arbitrary shapes by learning importance, centroid, and range of every Gaussian in the mixture. In learning GMP, each Gaussian is not trained in a feature space but is implemented over a temporal location. Thus the conventional feature-based learning for Gaussian mixture model is not valid for our case. In our special setting, to learn moderately coupled Gaussian mixture capturing diverse events, we newly propose a pull-push learning scheme using pulling and pushing losses, each of which plays an opposite role to the other. The effects of components in our scheme are verified in-depth with extensive ablation studies and the overall scheme achieves state-of-the-art performance. Our code is available at https://github.com/sunoh-kim/pps.
Sunoh Kim, Jungchan Cho, Joonsang Yu, Young Joon Yoo, Jin Young Choi 0002
AAAI4
2024 Topic-VQ-VAE: Leveraging Latent Codebooks for Flexible Topic-Guided Document Generation
abstract
This paper introduces a novel approach for topic modeling utilizing latent codebooks from Vector-Quantized Variational Auto-Encoder~(VQ-VAE), discretely encapsulating the rich information of the pre-trained embeddings such as the pre-trained language model. From the novel interpretation of the latent codebooks and embeddings as conceptual bag-of-words, we propose a new generative topic model called Topic-VQ-VAE~(TVQ-VAE) which inversely generates the original documents related to the respective latent codebook. The TVQ-VAE can visualize the topics with various generative distributions including the traditional BoW distribution and the autoregressive image generation. Our experimental results on document analysis and image generation demonstrate that TVQ-VAE effectively captures the topic context which reveals the underlying structures of the dataset and supports flexible forms of document generation. Official implementation of the proposed TVQ-VAE is available at https://github.com/clovaai/TVQ-VAE.
Young Joon Yoo, Jongwon Choi 0002
AAAI1
2024 Compose and Conquer: Diffusion-Based 3D Depth Aware Composable Image Synthesis
abstract
Addressing the limitations of text as a source of accurate layout representation in text-conditional diffusion models, many works incorporate additional signals to condition certain attributes within a generated image. Although successful, previous works do not account for the specific localization of said attributes extended into the three dimensional plane. In this context, we present a conditional diffusion model that integrates control over three-dimensional object placement with disentangled representations of global stylistic semantics from multiple exemplar images. Specifically, we first introduce depth disentanglement training to leverage the relative depth of objects as an estimator, allowing the model to identify the absolute positions of unseen objects through the use of synthetic image triplets. We also introduce soft guidance, a method for imposing global semantics onto targeted regions without the use of any additional localization cues. Our integrated framework, Compose and Conquer (CnC), unifies these techniques to localize multiple conditions in a disentangled manner. We demonstrate that our approach allows perception of objects at varying depths while offering a versatile framework for composing localized objects with different global semantics.
Jonghyun Lee 0006, Hansam Cho, Young Joon Yoo, Seoung Bum Kim, Yonghyun Jeong
ICLR3
2024 NCIS: Neural Contextual Iterative Smoothing for Purifying Adversarial Perturbations
abstract
We propose a novel and effective purification-based adversarial defense method against pre-processor blind white-and black-box attacks, without requiring any adversarial training or retraining of the classification model. Based on the observation of the adversarial noise, we propose a simple iterative Gaussian Smoothing (GS) that smoothes out adversarial noise and achieves substantially high robust accuracy. To further improve the method, we propose Neural Contextual Iterative Smoothing (NCIS), which trains a blind-spot network (BSN) in a self-supervised manner to reconstruct the discriminative features of the smoothed original image. From the extensive experiments on the large-scale ImageNet, we show that our method achieves both competitive standard accuracy and state-of-the-art robust accuracy against most strong purifier-blind white- and black-box attacks. Also, we propose a new evaluation benchmark based on commercial image classification APIs, including AWS, Azure, Clarifai, and Google, and demonstrate that users can use our method to increase the adversarial robustness of APIs.
Sungmin Cha, Naeun Ko, Heewoong Choi, Young Joon Yoo, Taesup Moon
WACV4
2024 EResFD: Rediscovery of the Effectiveness of Standard Convolution for Lightweight Face Detection
abstract
This paper analyzes the design choices of face detection architecture that improve efficiency of computation cost and accuracy. Specifically, we re-examine the effectiveness of the standard convolutional block as a lightweight backbone architecture for face detection. Unlike the current tendency of lightweight architecture design, which heavily utilizes depthwise separable convolution layers, we show that heavily channel-pruned standard convolution layers can achieve better accuracy and inference speed when using a similar parameter size. This observation is supported by the analyses concerning the characteristics of the target data domain, faces. Based on our observation, we propose to employ ResNet with a highly reduced channel, which surprisingly allows high efficiency compared to other mobile-friendly networks (e.g., MobileNetV1, V2, V3). From the extensive experiments, we show that the proposed backbone can replace that of the state-of-the-art face detector with a faster inference speed. Also, we further propose a new feature aggregation method to maximize the detection performance. Our proposed detector EResFD obtained 80.4% mAP on WIDER FACE Hard subset which only takes 37.7 ms for VGA image inference on CPU. Code is available at https://github.com/clovaai/EResFD.
Joonhyun Jeong, Beomyoung Kim, Joonsang Yu, Young Joon Yoo
WACV4
2023 GeNAS: Neural Architecture Search with Better Generalization
abstract
Neural Architecture Search (NAS) aims to automatically excavate the optimal network architecture with superior test performance. Recent neural architecture search (NAS) approaches rely on validation loss or accuracy to find the superior network for the target data. In this paper, we investigate a new neural architecture search measure for excavating architectures with better generalization. We demonstrate that the flatness of the loss surface can be a promising proxy for predicting the generalization capability of neural network architectures. We evaluate our proposed method on various search spaces, showing similar or even better performance compared to the state-of-the-art NAS methods. Notably, the resultant architecture found by flatness measure generalizes robustly to various shifts in data distribution (e.g. ImageNet-V2,-A,-O), as well as various tasks such as object detection and semantic segmentation.
Joonhyun Jeong, Joonsang Yu, Geondo Park, Dongyoon Han, Young Joon Yoo
IJCAI5
2022 Beyond Semantic to Instance Segmentation: Weakly-Supervised Instance Segmentation via Semantic Knowledge Transfer and Self-Refinement
abstract
Weakly-supervised instance segmentation (WSIS) has been considered as a more challenging task than weakly-supervised semantic segmentation (WSSS). Compared to WSSS, WSIS requires instance-wise localization, which is difficult to extract from image-level labels. To tackle the problem, most WSIS approaches use off-the-shelf proposal techniques that require pre-training with instance or object level labels, deviating the fundamental definition of the fully-image-level supervised setting. In this paper, we propose a novel approach including two innovative components. First, we propose a semantic knowledge transfer to obtain pseudo instance labels by transferring the knowledge of WSSS to WSIS while eliminating the need for the off-the-shelf proposals. Second, we propose a self-refinement method to refine the pseudo instance labels in a self-supervised scheme and to use the refined labels for training in an online manner. Here, we discover an erroneous phenomenon, semantic drift, that occurred by the missing instances in pseudo instance labels categorized as background class. This semantic drift occurs confusion between background and instance in training and consequently degrades the segmentation performance. We term this problem as semantic drift problem and show that our proposed self-refinement method eliminates the semantic drift problem. The extensive experiments on PASCAL VOC 2012 and MS COCO demonstrate the effectiveness of our approach, and we achieve a considerable performance without off-the-shelf proposal techniques. The code is available at https://github.com/clovaai/BESTIE.
Beomyoung Kim, Young Joon Yoo, Chaeeun Rhee, Junmo Kim 0002
CVPR2
2022 Learning Features with Parameter-Free Layers
Dongyoon Han, Young Joon Yoo, Beomyoung Kim, Byeongho Heo
ICLR2
2021 Rainbow Memory: Continual Learning With a Memory of Diverse Samples
abstract
Continual learning is a realistic learning scenario for AI models. Prevalent scenario of continual learning, however, assumes disjoint sets of classes as tasks and is less realistic rather artificial. Instead, we focus on ‘blurry’ task boundary; where tasks shares classes and is more realistic and practical. To address such task, we argue the importance of diversity of samples in an episodic memory. To enhance the sample diversity in the memory, we propose a novel memory management strategy based on per-sample classification uncertainty and data augmentation, named Rainbow Memory (RM). With extensive empirical validations on MNIST, CIFAR10, CIFAR100, and ImageNet datasets, we show that the proposed method significantly improves the accuracy in blurry continual learning setups, outperforming state of the arts by large margins despite its simplicity. Code and data splits will be available in https://github.com/clovaai/rainbow-memory.
Jihwan Bang, Heesu Kim, Young Joon Yoo, Jung-Woo Ha 0001
CVPR3
2021 Rethinking Channel Dimensions for Efficient Model Design
abstract
Designing an efficient model within the limited computational cost is challenging. We argue the accuracy of a lightweight model has been further limited by the design convention: a stage-wise configuration of the channel dimensions, which looks like a piecewise linear function of the network stage. In this paper, we study an effective channel dimension configuration towards better performance than the convention. To this end, we empirically study how to design a single layer properly by analyzing the rank of the output feature. We then investigate the channel configuration of a model by searching network architectures concerning the channel configuration under the computational cost restriction. Based on the investigation, we propose a simple yet effective channel configuration that can be parameterized by the layer index. As a result, our proposed model following the channel parameterization achieves remarkable performance on ImageNet classification and transfer learning tasks including COCO object detection, COCO instance segmentation, and fine-grained classifications. Code and ImageNet pretrained models are available at https: //github.com/clovaai/rexnet.
Dongyoon Han, Sangdoo Yun, Byeongho Heo, Young Joon Yoo
CVPR4
2021 SSUL: Semantic Segmentation with Unknown Label for Exemplar-based Class-Incremental Learning
abstract
We consider a class-incremental semantic segmentation (CISS) problem. While some recently proposed algorithms utilized variants of knowledge distillation (KD) technique to tackle the problem, they only partially addressed the key additional challenges in CISS that causes the catastrophic forgetting; \textit{i.e.}, the semantic drift of the background class and multi-label prediction issue. To better address these challenges, we propose a new method, dubbed as SSUL-M (Semantic Segmentation with Unknown Label with Memory), by carefully combining several techniques tailored for semantic segmentation. More specifically, we make three main contributions; (1) modeling \textit{unknown} class within the background class to help learning future classes (help plasticity), (2) \textit{freezing} backbone network and past classifiers with binary cross-entropy loss and pseudo-labeling to overcome catastrophic forgetting (help stability), and (3) utilizing \textit{tiny exemplar memory} for the first time in CISS to improve \textit{both} plasticity and stability. As a result, we show our method achieves significantly better performance than the recent state-of-the-art baselines on the standard benchmark datasets. Furthermore, we justify our contributions with thorough and extensive ablation analyses and discuss different natures of the CISS problem compared to the standard class-incremental learning for classification. The official code is available at https://github.com/clovaai/SSUL.
Sungmin Cha, Beomyoung Kim, Young Joon Yoo, Taesup Moon
NeurIPS3
2020 SINet: Extreme Lightweight Portrait Segmentation Networks with Spatial Squeeze Modules and Information Blocking Decoder
abstract
Designing a lightweight and robust portrait segmentation algorithm is an important task for a wide range of face applications. However, the problem has been considered as a subset of the object segmentation and less handled in this field. Obviously, portrait segmentation has its unique requirements. First, because the portrait segmentation is performed in the middle of a whole process, it requires extremely lightweight models. Second, there has not been any public datasets in this domain that contain a sufficient number of images. To solve the first problem, we introduce the new extremely lightweight portrait segmentation model SINet, containing an information blocking decoder and spatial squeeze modules. The information blocking decoder uses confidence estimation to recover local spatial information without spoiling global consistency. The spatial squeeze module uses multiple receptive fields to cope with various sizes of consistency. To tackle the second problem, we propose a simple method to create additional portrait segmentation data, which can improve accuracy. In our qualitative and quantitative analysis on the EG1800 dataset, we show that our method outperforms various existing lightweight models. Our method reduces the number of parameters from 2.1M to 86.9K (around 95.9% reduction), while maintaining the accuracy under an 1% margin from the state-of-the-art method. We also show our model is successfully executed on a real mobile device with 100.6 FPS. In addition, we demonstrate that our method can be used for general semantic segmentation on the Cityscapes dataset. The code and dataset are available in https://github.com/HYOJINPARK/ExtPortraitSeg.
Hyojin Park 0001, Lars Lowe Sjösund, Young Joon Yoo, Nicolas Monet, Jihwan Bang, Nojun Kwak
WACV3
2019 CutMix: Regularization Strategy to Train Strong Classifiers With Localizable Features
abstract
Regional dropout strategies have been proposed to enhance performance of convolutional neural network classifiers. They have proved to be effective for guiding the model to attend on less discriminative parts of objects (e.g. leg as opposed to head of a person), thereby letting the network generalize better and have better object localization capabilities. On the other hand, current methods for regional dropout removes informative pixels on training images by overlaying a patch of either black pixels or random noise. Such removal is not desirable because it suffers from information loss causing inefficiency in training. We therefore propose the CutMix augmentation strategy: patches are cut and pasted among training images where the ground truth labels are also mixed proportionally to the area of the patches. By making efficient use of training pixels and retaining the regularization effect of regional dropout, CutMix consistently outperforms state-of-the-art augmentation strategies on CIFAR and ImageNet classification tasks, as well as on ImageNet weakly-supervised localization task. Moreover, unlike previous augmentation methods, our CutMix-trained ImageNet classifier, when used as a pretrained model, results in consistent performance gain in Pascal detection and MS-COCO image captioning benchmarks. We also show that CutMix can improve the model robustness against input corruptions and its out-of distribution detection performance.
Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh, Young Joon Yoo, Junsuk Choe
ICCV5
2019 BOOK: Storing Algorithm-Invariant Episodes for Deep Reinforcement Learning
abstract
We introduce a novel method to train agents of reinforcement learning (RL) by sharing knowledge in a way similar to the concept of using a book. The recorded information in the form of a book is the main means by which humans learn knowledge. Nevertheless, the conventional deep RL methods have mainly focused either on experiential learning where the agent learns through interactions with the environment from the start or on imitation learning that tries to mimic the teacher. Contrary to these, our proposed book learning shares key information among different agents in a book-like manner by delving into the following two characteristic features: (1) By defining the linguistic function, input states can be clustered semantically into a relatively small number of core clusters, which are forwarded to other RL agents in a prescribed manner. (2) By defining state priorities and the contents for recording, core experiences can be selected and stored in a small container. We call this container as 'BOOK'. Our method learns hundreds to thousand times faster than the conventional methods by learning only a handful of core cluster information, which shows that deep RL agents can effectively learn through the shared knowledge from other agents.
Simyung Chang, Young Joon Yoo, Jaeseok Choi, Nojun Kwak
ICPRAM2
2018 MC-GAN: Multi-conditional Generative Adversarial Network for Image Synthesis
Hyojin Park 0001, Young Joon Yoo, Nojun Kwak
BMVC2
2018 Dynamic Graph Generation Network: Generating Relational Knowledge From Diagrams
abstract
In this work, we introduce a new algorithm for analyzing a diagram, which contains visual and textual information in an Abstract and integrated way. Whereas diagrams contain richer information compared with individual image-based or language-based data, proper solutions for automatically understanding them have not been proposed due to their innate characteristics of multi-modality and arbitrariness of layouts. To tackle this problem, we propose a unified diagram-parsing network for generating knowledge from diagrams based on an object detector and a recurrent neural network designed for a graphical structure. Specifically, we propose a dynamic graph-generation network that is based on dynamic memory and graph theory. We explore the dynamics of information in a diagram with activation of gates in gated recurrent unit (GRU) cells. On publicly available diagram datasets, our model demonstrates a state-of-the-art result that outperforms other baselines. Moreover, further experiments on question answering shows potentials of the proposed method for various applications.
Young Joon Yoo, Jeesoo Kim, Sangkuk Lee, Nojun Kwak
CVPR2
2018 Pose transforming network: Learning to disentangle human posture in variational auto-encoded latent space
Jongin Lim 0002, Young Joon Yoo, Byeongho Heo, Jin Young Choi 0002
Pattern Recognit. Lett.2
2018 Action-Driven Visual Object Tracking With Deep Reinforcement Learning
abstract
In this paper, we propose an efficient visual tracker, which directly captures a bounding box containing the target object in a video by means of sequential actions learned using deep neural networks. The proposed deep neural network to control tracking actions is pretrained using various training video sequences and fine-tuned during actual tracking for online adaptation to a change of target and background. The pretraining is done by utilizing deep reinforcement learning (RL) as well as supervised learning. The use of RL enables even partially labeled data to be successfully utilized for semisupervised learning. Through the evaluation of the object tracking benchmark data set, the proposed tracker is validated to achieve a competitive performance at three times the speed of existing deep network-based trackers. The fast version of the proposed method, which operates in real time on graphics processing unit, outperforms the state-of-the-art real-time trackers with an accuracy improvement of more than 8%.
Sangdoo Yun, Jongwon Choi 0002, Young Joon Yoo, Kimin Yun, Jin Young Choi 0002
IEEE Trans. Neural Networks Learn. Syst.3
2017 Superpixel-based semantic segmentation trained by statistical process control
Hyojin Park 0001, Jisoo Jeong, Young Joon Yoo, Nojun Kwak
BMVC3
2017 Variational Autoencoded Regression: High Dimensional Regression of Visual Data on Complex Manifold
abstract
This paper proposes a new high dimensional regression method by merging Gaussian process regression into a variational autoencoder framework. In contrast to other regression methods, the proposed method focuses on the case where output responses are on a complex high dimensional manifold, such as images. Our contributions are summarized as follows: (i) A new regression method estimating high dimensional image responses, which is not handled by existing regression algorithms, is proposed. (ii) The proposed regression method introduces a strategy to learn the latent space as well as the encoder and decoder so that the result of the regressed response in the latent space coincide with the corresponding response in the data space. (iii) The proposed regression is embedded into a generative model, and the whole procedure is developed by the variational autoencoder framework. We demonstrate the robustness and effectiveness of our method through a number of experiments on various visual data regression problems.
Young Joon Yoo, Sangdoo Yun, Hyung Jin Chang, Yiannis Demiris, Jin Young Choi 0002
CVPR1
2017 Action-Decision Networks for Visual Tracking with Deep Reinforcement Learning
abstract
This paper proposes a novel tracker which is controlled by sequentially pursuing actions learned by deep reinforcement learning. In contrast to the existing trackers using deep networks, the proposed tracker is designed to achieve a light computation as well as satisfactory tracking accuracy in both location and scale. The deep network to control actions is pre-trained using various training sequences and fine-tuned during tracking for online adaptation to target and background changes. The pre-training is done by utilizing deep reinforcement learning as well as supervised learning. The use of reinforcement learning enables even partially labeled data to be successfully utilized for semi-supervised learning. Through evaluation of the OTB dataset, the proposed tracker is validated to achieve a competitive performance that is three times faster than state-of-the-art, deep network-based trackers. The fast version of the proposed method, which operates in real-time on GPU, outperforms the state-of-the-art real-time trackers.
Sangdoo Yun, Jongwon Choi 0002, Young Joon Yoo, Kimin Yun, Jin Young Choi 0002
CVPR3
2017 Motion interaction field for detection of abnormal interactions
Kimin Yun, Young Joon Yoo, Jin Young Choi 0002
Mach. Vis. Appl.2
2016 Visual Path Prediction in Complex Scenes with Crowded Moving Objects
abstract
This paper proposes a novel path prediction algorithm for progressing one step further than the existing works focusing on single target path prediction. In this paper, we consider moving dynamics of co-occurring objects for path prediction in a scene that includes crowded moving objects. To solve this problem, we first suggest a two-layered probabilistic model to find major movement patterns and their cooccurrence tendency. By utilizing the unsupervised learning results from the model, we present an algorithm to find the future location of any target object. Through extensive qualitative/quantitative experiments, we show that our algorithm can find a plausible future path in complex scenes with a large number of moving objects.
Young Joon Yoo, Kimin Yun, Sangdoo Yun, Jonghee Hong, Hawook Jeong, Jin Young Choi 0002
CVPR1
2014 Multi-task learning with over-sampled time-series representation of a trajectory for traffic motion pattern recognition
abstract
This paper proposes an efficient feature sampling and multi-task learning scheme for traffic scene analysis, where all classifiers are trained simultaneously by exploiting the correlations among different motion patterns. We make feature descriptors by high dimensional embedding of the time series data for traffic pattern representation. They preserve detailed spatio-temporal information of the underlying event. Pattern specific details are extracted from raw trajectories and embedded into feature descriptors, which ensures their great discriminability. Training data scarcity problem is tackled through amplification of the patterns hidden in raw trajectory via strategic oversampling and employment of joint feature selection procedure while training the models. Experimental results on 4 surveillance datasets, show great improvement in the motion pattern recognition performance, importance of joint feature selection and fast incremental learning ability of the proposed framework.
Tushar Sandhan, Young Joon Yoo, Hanjoo Yoo, Sangdoo Yun, Moonsub Byeon
AVSS2
2014 Visual surveillance briefing system: Event-based video retrieval and summarization
abstract
This paper presents a visual surveillance briefing (VSB) system which provides event-based retrieval and briefing functions. Traditional event-based video retrieval systems usually aim to analyze the appearance of objects rather than the motion information (e.g. trajectory) of objects. The VSB system adopts the video summarization technique which temporally abstracts the retrieved events to understand the motion patterns of objects. We propose various event features including object's appearances and motion patterns for the purpose of event retrieval and design the energy function to abstract the retrieved events in real-time. To avoid the occlusion problem in the briefed events, we propose an animated displaying method that separately presents the global motion and the local motion of moving objects. Effectiveness of the implemented VSB system is evaluated through several surveillance videos.
Sangdoo Yun, Kimin Yun, Soo Wan Kim, Young Joon Yoo, Jiyeoup Jeong
AVSS4
2014 Transfer Learning of Motion Patterns in Traffic Scene via Convex Optimization
abstract
This paper proposes a transfer learning scheme for traffic pattern analysis where the transferred classifier could be trained with a small number of samples. First we make feature descriptors to represent the traffic trajectories so that they should be adequate to transfer and classify the traffic patterns. Then, we use support vector machine (SVM) to learn the feature descriptors of traffic trajectories. The transfer learning scheme is formulated by a convex optimization problem using the geometric relation between target and source patterns. Not only parameters of SVM but also the geometric relation are found at the same time through two step minimization process of the optimization problem. Through experiments on various surveillance videos, the proposed formulation is shown to be valid by investigating the improvement of performance compared to a transfer scheme without the proposed geometric relation as well as SVM without transfer scheme.
Young Joon Yoo, Hawook Jeong, Soo Wan Kim, Jin Young Choi 0002
ICPR1
2014 Two-stage online inference model for traffic pattern analysis and anomaly detection
Hawook Jeong, Young Joon Yoo, Kwang Moo Yi, Jin Young Choi 0002
Mach. Vis. Appl.2
2013 Towards simultaneous clustering and motif-modeling for a large number of protein family
abstract
In this paper, we propose a novel clustering and motif modeling framework for analyzing large number of protein family using k-mer. Our approach of using k-mers utilizes both occurring frequency and position information of k-mers that essential for classification yet not fully used in previous methods. We found that the structure has close relationship between motif of protein family and hence well describe important biological features or motifs of each protein family. The classification/clustering procedure are executed in incremental manner which was difficult for previous algorithms and is modeled by using bipartite model. Furthermore, the method can be efficiently implemented using parallel computing and hash. Experimental results using the entire COG family database shows that our model can model a large number of protein families without sacrificing accuracy. In addition, the classification structure, path of the graph for protein sequences, explains characteristic subsequences or motif of each family quite well. Thus the proposed method has the potential to model both protein families and motifs, even for a large number of families.
Young Joon Yoo, Tushar Sandhan, Jin Young Choi 0002, Sun Kim
BIBM1