Junseok Kwon

dblp:04/425 · DBLP profile ↗
← Back
75ranked-venue papers
19as first author
36since 2021 · last 2026
0000-0001-9526-7549ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 53 · 16 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 9 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Neural Collapse-Informed Initialization with Perturbation Injection in Classification-based Metric Learning
abstract
Recent studies have revealed Neural Collapse (NC) in deep classifiers, where last-layer weights and features align into an equiangular tight frame (ETF), concentrating class information along specific embedding directions. However, conventional fine-tuning typically disregards this structure, initializing task-specific classifier heads randomly. To explicitly leverage this phenomenon, we propose a simple yet effective method for metric learning: (1) initializing the classifier head along each class’s NC direction from a pretrained model to preserve the emergent structure, and (2) injecting small isotropic Gaussian noise during finetuning to boost generalization. In addition, we provide a theoretical bound proving that our method explicitly reduces cumulative weight drift from the NC-initialization, compared to standard finetuning. This suggests that our method better preserves the pretrained model’s class-specific structure. Empirically, this structural preservation yields Recall@K gains: reduced weight drift correlates with better performance. Concurrent decreases in the Neural Collapse 1 (NC1) measure confirm that stronger intra‐class cohesion underlies these improvements. Furthermore, we validate the effectiveness of our method on class‐imbalanced benchmarks.
Jinhee Park 0001, Hee Bin Yoo, Byoung-Tak Zhang, Junseok Kwon
AAAI5
2026 Confidence Controls Deep Metric Learning
Jinhee Park 0001, Hee Bin Yoo, Byoung-Tak Zhang, Junseok Kwon
Mach. Learn.4
2025 Deep Disentangled Metric Learning
abstract
Proxy-based metric learning has enhanced semantic similarity with class representatives and exhibited noteworthy performance in deep metric learning (DML) tasks. While these methods alleviate computational demands by learning instance-to-class relationships rather than instance-to-instance relationships, they often limit features to be class-specific, thereby degrading generalization performance for unseen class. In this paper, we introduce a novel perspective called Disentangled Deep Metric Learning (DDML), grounded in the framework of information bottleneck, which applies class-agnostic regularization to existing DML methods. Unlike conventional NormSoftmax methods, which primarily emphasize distinct class-specific features, our DDML enables a diverse feature representation by seamlessly transitioning between class-specific features with the aid of class-agnostic features. It smooths decision boundaries, allowing unseen classes to have stable semantic representations in the embedding space. To achieve this, we learn disentangled representations of both class-specific and class-agnostic features in the context of DML. Empirical results demonstrate that our method addresses the limitations of conventional approaches. Our method easily integrates into existing proxy-based algorithms, consistently delivering improved performance.
Jinhee Park 0001, Dagyeong Na, Junseok Kwon
AAAI4
2025 Lightweight Wasserstein Audio-Visual Model for Unified Speech Enhancement and Separation
abstract
Speech Enhancement (SE) and Speech Separation (SS) have traditionally been treated as distinct tasks in speech processing. However, real-world audio often involves both background noise and overlapping speakers, motivating the need for a unified solution. While recent approaches have attempted to integrate SE and SS within multi-stage architectures, these approaches typically involve complex, parameter-heavy models and rely on supervised training, limiting scalability and generalization. In this work, we propose UniVoiceLite, a lightweight and unsupervised audiovisual framework that unifies SE and SS within a single model. UniVoiceLite leverages lip motion and facial identity cues to guide speech extraction and employs Wasserstein distance regularization to stabilize the latent space without requiring paired noisy-clean data. Experimental results demonstrate that UniVoiceLite achieves strong performance in both noisy and multi-speaker scenarios, combining efficiency with robust generalization. The source code is available at https://github.com/jisoo-o/UniVoiceLite.
Seonghak Lee, Guisik Kim, Junseok Kwon
ASRU5
2025 Occlusion-Robust Multi-Object Tracking with Adaptive Feature Management and Motion Compensation
abstract
Multi-object tracking (MOT) faces challenges in handling occlusions, feature degradation, and non-rigid motion. Existing methods relying on appearance-based re-identification (Re-ID) often struggle under occlusion, leading to frequent identity switches, while traditional motion models fail in dynamic scenarios. To address these issues, we propose an improved MOT framework integrating Score-based Gallery Management (SGM) to retain reliable Re-ID embeddings and Optical Flow-based Motion Compensation (OMC) to refine motion predictions. Our method achieves state-of-the-art performance on MOT20 and SportsMOT and exhibits competitive results on MOT17 and DanceTrack, demonstrating improved identity retention and tracking robustness in complex environments.
Yoojin Han, Junseok Kwon
AVSS3
2025 Towards Generalizable Time Series Forecasting Via IB-Regularized Transformer-Diffusion
abstract
Multivariate Time Series Forecasting (MTSF) remains challenging due to the need to capture complex temporal dependencies and ensure robustness against distribution shift. While Transformer-based models excel at capturing long-range patterns, and diffusion models offer generative flexibility, directly applying diffusion often leads to excessive uncertainty. To address these limitations, we propose a unified forecasting framework (IB-TransDiff) that combines global modeling based on Transformers and local refinement based on diffusion, guided by the Information Bottleneck (IB) principle. To improve generalization, we introduce a novel regularization strategy that penalizes the second moment of the conditional embedding to reduce its mutual information with the input. Experiments on multiple benchmarks demonstrate that our method consistently outperforms existing approaches in MSE and MAE. In addition, our method reduces uncertainty, achieving notable improvements in QICE, CRPS, and PICP metrics. These results highlight the value of integrating information-theoretic constraints into conditional diffusion models for robust and accurate time series forecasting.
Dagyeong Na, Jinho Kang, Junseok Kwon
ICDM3
2025 Physics-Informed Neural Operators for Tissue Elasticity Reconstruction
Youjin Kim, Jae Yong Lee 0002, Junseok Kwon
MICCAI (10)3
2024 Uncertainty Calibration with Energy Based Instance-Wise Scaling in the Wild Dataset
Mijoo Kim, Junseok Kwon
ECCV (46)2
2024 Sampling based spherical transformer for 360 degree image classification
Sungmin Cho, Raehyuk Jung, Junseok Kwon
Expert Syst. Appl.3
2024 End-to-end metric learning from corrupted images using triplet dimensionality reduction loss
JuHyeon Park, Junseok Kwon
Expert Syst. Appl.3
2024 Satellite Image Dehazing Based on Dual Frequency Pass Networks
abstract
Remote sensing using satellite imagery has been actively researched, inducing various applications of computer vision. In this field, the quality of satellite images is very important in facilitating continuous Earth observation and environmental monitoring. However, even after undergoing various correction processes, satellite images inevitably contain haze and clouds. The presence of these haze and clouds introduces numerous challenges to the acquisition of high-quality satellite images. In this study, we present a novel dehazing method designed to enhance the quality of satellite images named dual frequency pass networks (DFPNs). The proposed method comprises two branches: a transformer branch for capturing low-frequency components and a convolution branch for extracting high-frequency components. Thus, this approach can consider both the global features from the transformer and the local features from the convolution. The experiments demonstrate that the proposed method outperforms other state-of-the-art methods.
Guisik Kim, Chungsang Cho, Joohyung Kang, Junseok Kwon
IEEE Geosci. Remote. Sens. Lett.4
2024 Hierarchical Bidirected Graph Convolutions for Large-Scale 3-D Point Cloud Place Recognition
abstract
In this article, we present a novel hierarchical bidirected graph convolution network (HiBi-GCN) for large-scale 3-D point cloud place recognition. Unlike place recognition methods based on 2-D images, those based on 3-D point cloud data are typically robust to substantial changes in real-world environments. However, these methods have difficulty in defining convolution for point cloud data to extract informative features. To solve this problem, we propose a new hierarchical kernel defined as a hierarchical graph structure through unsupervised clustering from the data. In particular, we pool hierarchical graphs from the fine to coarse direction using pooling edges and fuse the pooled graphs from the coarse to fine direction using fusing edges. The proposed method can, thus, learn representative features hierarchically and probabilistically; moreover, it can extract discriminative and informative global descriptors for place recognition. Experimental results demonstrate that the proposed hierarchical graph structure is more suitable for point clouds to represent real-world 3-D scenes.
Dong Wook Shu, Junseok Kwon
IEEE Trans. Neural Networks Learn. Syst.2
2023 AttSec: protein secondary structure prediction by capturing local patterns from attention map
abstract
BACKGROUND: Protein secondary structures that link simple 1D sequences to complex 3D structures can be used as good features for describing the local properties of protein, but also can serve as key features for predicting the complex 3D structures of protein. Thus, it is very important to accurately predict the secondary structure of the protein, which contains a local structural property assigned by the pattern of hydrogen bonds formed between amino acids. In this study, we accurately predict protein secondary structure by capturing the local patterns of protein. For this objective, we present a novel prediction model, AttSec, based on transformer architecture. In particular, AttSec extracts self-attention maps corresponding to pairwise features between amino acid embeddings and passes them through 2D convolution blocks to capture local patterns. In addition, instead of using additional evolutionary information, it uses protein embedding as an input, which is generated by a language model. RESULTS: For the ProteinNet DSSP8 dataset, our model showed 11.8% better performance on the entire evaluation datasets compared with other no-evolutionary-information-based models. For the NetSurfP-2.0 DSSP8 dataset, it showed 1.2% better performance on average. There was an average performance improvement of 9.0% for the ProteinNet DSSP3 dataset and an average of 0.7% for the NetSurfP-2.0 DSSP3 dataset. CONCLUSION: We accurately predict protein secondary structure by capturing the local patterns of protein. For this objective, we present a novel prediction model, AttSec, based on transformer architecture. Although there was no dramatic accuracy improvement compared with other models, the improvement on DSSP8 was greater than that on DSSP3. This result implies that using our proposed pairwise feature could have a remarkable effect for several challenging tasks that require finely subdivided classification. Github package URL is https://github.com/youjin-DDAI/AttSec .
Youjin Kim, Junseok Kwon
BMC Bioinform.2
2023 Localized curvature-based combinatorial subgraph sampling for large-scale graphs
Dong Wook Shu, Youjin Kim, Junseok Kwon
Pattern Recognit.3
2023 Self-Parameter Distillation Dehazing
abstract
In this paper, we propose a novel dehazing method based on self-distillation. In contrast to conventional knowledge distillation approaches that transfer large models (teacher networks) to small models (student networks), we introduce a single knowledge distillation network that transfers network parameters to itself for dehazing. In the early stages, the proposed network transfers scene content (identity) information to the next stage of itself using haze-free data. However, in the later stages, the network transfers haze information to itself using haze data, enabling the accurate dehazing of input images using scene information from the early stages. In a single network, parameters are seamlessly updated from extracting global scene features to dehazing the scene. During the training, forward propagation acts as a teacher network, whereas backward propagation acts as a student network. The experimental results demonstrate that the proposed method considerably outperforms other state-of-the-art dehazing methods.
Guisik Kim, Junseok Kwon
IEEE Trans. Image Process.2
2022 Neural Markov Controlled SDE: Stochastic Optimization for Continuous-Time Data
Sung Woo Park, Junseok Kwon
ICLR3
2022 Riemannian Neural SDE: Learning Stochastic Representations on Manifolds
abstract
In recent years, the neural stochastic differential equation (NSDE) has gained attention for modeling stochastic representations with great success in various types of applications. However, it typically loses expressivity when the data representation is manifold-valued. To address this issue, we suggest a principled method for expressing the stochastic representation with the Riemannian neural SDE (RNSDE), which extends the conventional Euclidean NSDE. Empirical results for various tasks demonstrate that the proposed method significantly outperforms baseline methods.
Sung Woo Park, Hyomin Kim, Junseok Kwon
NeurIPS4
2022 Densely-packed Object Detection via Hard Negative-Aware Anchor Attention
abstract
In this paper, we propose a novel densely-packed object detection method based on advanced weighted Hausdorff distance (AWHD) and hard negative-aware anchor (HNAA) attention. Densely-packed object detection is more challenging than conventional object detection due to the high object density and small-size objects. To overcome these challenges, the proposed AWHD improves the conventional weighted Hausdorff distance and obtains an accurate center area map. Using the precise center area map, the proposed HNAA attention determines the relative importance of each anchor and imposes a penalty on hard negative anchors. Experimental results demonstrate that our proposed method based on the AWHD and HNAA attention produces accurate densely-packed object detection results and comparably outperforms other state-of-the-art detection methods. The code is available at ${\color{Blue} \text{here}}$.
Sungmin Cho, Jinwook Paeng, Junseok Kwon
WACV3
2022 Optimal visual tracking using Wasserstein transport proposals
Junseok Kwon
Expert Syst. Appl.2
2022 Riemannian submanifold framework for log-Euclidean metric learning on symmetric positive definite manifolds
Sung Woo Park, Junseok Kwon
Expert Syst. Appl.2
2022 Adversarial attack can help visual tracking
Sungmin Cho, Hyeseong Kim, Ji Soo Kim, Hyomin Kim, Junseok Kwon
Multim. Tools Appl.5
2022 Style transfer with target feature palette and attention coloring
Suhyeon Ha, Guisik Kim, Junseok Kwon
Multim. Tools Appl.3
2022 Stagemix video generation using face and body keypoints detection
Minjoon Jung, Eun-Seon Sim, Min Ho Jo, Hyebin Choi, Junseok Kwon
Multim. Tools Appl.7
2022 Deep-plane sweep generative adversarial network for consistent multi-view depth estimation
Dong Wook Shu, Wonbeom Jang, Heebin Yoo, Hong-Chang Shin, Junseok Kwon
Mach. Vis. Appl.5
2022 SphereGAN: Sphere Generative Adversarial Network Based on Geometric Moment Matching and its Applications
abstract
We propose a novel integral probability metric-based generative adversarial network (GAN), called SphereGAN. In the proposed scheme, the distance between two probability distributions (i.e., true and fake distributions) is measured on a hypersphere. Given that its hypersphere-based objective function computes the upper bound of the distance as a half arc, SphereGAN can be stably trained and can achieve a high convergence rate. In sphereGAN, higher-order information of data is processed using multiple geometric moments, thus improving the accuracy of the distance measurement and producing more realistic outcomes. Several properties of the proposed distance metric on the hypersphere are mathematically derived. The effectiveness of the proposed SphereGAN is demonstrated through quantitative and qualitative experiments for unsupervised image generation and 3D point cloud generation, demonstrating its superiority over state-of-the-art GANs with respect to accuracy and convergence on the CIFAR-10, STL-10, LSUN bedroom, and ShapeNet datasets.
Sung Woo Park, Junseok Kwon
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Wasserstein approximate bayesian computation for visual tracking
Jinhee Park 0001, Junseok Kwon
Pattern Recognit.2
2022 Wasserstein distributional harvesting for highly dense 3D point clouds
Dong Wook Shu, Sung Woo Park, Junseok Kwon
Pattern Recognit.3
2022 Deep Illumination-Aware Dehazing With Low-Light and Detail Enhancement
abstract
We present a novel dehazing framework for real-world images that contain both hazy and low-light areas. Dehazing and low-light enhancements are unified by using an illumination map that is estimated using a proposed convolutional neural network. The illumination map is then used as a component for three different tasks: atmospheric light estimation, transmission map estimation, and low-light enhancement, thereby enabling the solving of interrelated low-level vision problems simultaneously. To train the neural network to perform both dehazing and low-light enhancement, we synthesize hazy and low-light images from normal images. Experimental results demonstrate that the proposed method quantitatively and qualitatively outperforms state-of-the-art algorithms in real-world image dehazing.
Guisik Kim, Junseok Kwon
IEEE Trans. Intell. Transp. Syst.2
2021 Filter Pruning Via Softmax Attention
abstract
In this paper, we propose a novel network pruning method using the proposed relative depth-wise separable convolutions and softmax attention channel pruning. The relative depthwise separable convolution enhances conventional depth-wise separable convolutions by enabling the channel interaction, which can prevent accuracy drops even after severe pruning. The softmax attention channel pruning probabilistically expresses the importance of filters and removes unimportant channels efficiently. Experimental results demonstrate that our pruning method outperforms other state-of-the-art pruning methods in terms of Flops, parameters, and top-1 classification accuracy.
Sungmin Cho, Hyeseong Kim, Junseok Kwon
ICIP3
2021 Wasserstein Distributional Normalization For Robust Distributional Certification of Noisy Labeled Data
abstract
We propose a novel Wasserstein distributional normalization method that can classify noisy labeled data accurately. Recently, noisy labels have been successfully handled based on small-loss criteria, but have not been clearly understood from the theoretical point of view. In this paper, we address this problem by adopting distributionally robust optimization (DRO). In particular, we present a theoretical investigation of the distributional relationship between uncertain and certain samples based on the small-loss criteria. Our method takes advantage of this relationship to exploit useful information from uncertain samples. To this end, we normalize uncertain samples into the robustly certified region by introducing the non-parametric Ornstein-Ulenbeck type of Wasserstein gradient flows called Wasserstein distributional normalization, which is cheap and fast to implement. We verify that network confidence and distributional certification are fundamentally correlated and show the concentration inequality when the network escapes from over-parameterization. Experimental results demonstrate that our non-parametric classification method outperforms other parametric baselines on the Clothing1M and CIFAR-10/100 datasets when the data have diverse noisy labels.
Sung Woo Park, Junseok Kwon
ICML2
2021 Generative Adversarial Networks for Markovian Temporal Dynamics: Stochastic Continuous Data Generation
abstract
In this paper, we present a novel generative adversarial network (GAN) that can describe Markovian temporal dynamics. To generate stochastic sequential data, we introduce a novel stochastic differential equation-based conditional generator and spatial-temporal constrained discriminator networks. To stabilize the learning dynamics of the min-max type of the GAN objective function, we propose well-posed constraint terms for both networks. We also propose a novel conditional Markov Wasserstein distance to induce a pathwise Wasserstein distance. The experimental results demonstrate that our method outperforms state-of-the-art methods using several different types of data.
Sung Woo Park, Dong Wook Shu, Junseok Kwon
ICML3
2021 Visual tracking using interactive factorial hidden Markov models
abstract
Abstract The authors present a novel tracking algorithm based on a factorial hidden Markov model (FHMM) that can utilise the structured information of a target. An FHMM consists of multiple hidden Markov models (HMMs), wherein each HMM aims to represent a different part of the target. Then, the geometric relation between patches is encoded in the FHMM framework via either interactive sampling or importance sampling over sets. Experimental results demonstrate that the proposed method qualitatively and quantitatively outperforms other methods, especially when the targets are highly deformable.
Jinwook Paeng, Junseok Kwon
IET Signal Process.2
2021 Graph visual tracking using conditional uncertainty minimization and minibatch Monte Carlo inference
Junseok Kwon
Inf. Sci.1
2021 Robust person re-identification via graph convolution networks
Guisik Kim, Dong Wook Shu, Junseok Kwon
Multim. Tools Appl.3
2021 Abnormal event detection by variation matching
Sungmin Cho, Junseok Kwon
Mach. Vis. Appl.2
2021 Pixel-Wise Wasserstein Autoencoder for Highly Generative Dehazing
abstract
We propose a highly generative dehazing method based on pixel-wise Wasserstein autoencoders. In contrast to existing dehazing methods based on generative adversarial networks, our method can produce a variety of dehazed images with different styles. It significantly improves the dehazing accuracy via pixel-wise matching from hazy to dehazed images through 2-dimensional latent tensors of the Wasserstein autoencoder. In addition, we present an advanced feature fusion technique to deliver rich information to the latent space. For style transfer, we introduce a mapping function that transforms existing latent spaces to new ones. Thus, our method can produce highly generative haze-free images with various tones, illuminations, and moods, which induces several interesting applications, including low-light enhancement, daytime dehazing, nighttime dehazing, and underwater image enhancement. Experimental results demonstrate that our method quantitatively outperforms existing state-of-the-art methods for synthetic and real-world datasets, and simultaneously generates highly generative haze-free images, which are qualitatively diverse.
Guisik Kim, Sung Woo Park, Junseok Kwon
IEEE Trans. Image Process.3
2020 Visual Tracking by TridentAlign and Context Embedding
Janghoon Choi, Junseok Kwon, Kyoung Mu Lee
ACCV (2)2
2020 DALE : Dark Region-Aware Low-light Image Enhancement
Dokyeong Kwon, Guisik Kim, Junseok Kwon
BMVC3
2020 Upright Adjustment With Graph Convolutional Networks
abstract
We present a novel method for the upright adjustment of 360° images. Our network consists of two modules, which are a convolutional neural network (CNN) and a graph convolutional network (GCN). The input 360° images is processed with the CNN for visual feature extraction, and the extracted feature map is converted into a graph that finds a spherical representation of the input. We also introduce a novel loss function to address the issue of discrete probability distributions defined on the surface of a sphere. Experimental results demonstrate that our method outperforms fully connected-based methods.
Raehyuk Jung, Sungmin Cho, Junseok Kwon
ICIP3
2020 Deep Diffusion-Invariant Wasserstein Distributional Classification
abstract
In this paper, we present a novel classification method called deep diffusion-invariant Wasserstein distributional classification (DeepWDC). DeepWDC represents input data and labels as probability measures to address severe perturbations in input data. It can output the optimal label measure in terms of diffusion invariance, where the label measure is stationary over time and becomes equivalent to a Gaussian measure. Furthermore, DeepWDC minimizes the 2-Wasserstein distance between the optimal label measure and Gaussian measure, which reduces the Wasserstein uncertainty. Experimental results demonstrate that DeepWDC can substantially enhance the accuracy of several baseline deterministic classification methods and outperforms state-of-the-art-methods on 2D and 3D data containing various types of perturbations (e.g., rotations, impulse noise, and down-scaling).
Sung Woo Park, Dong Wook Shu, Junseok Kwon
NeurIPS3
2020 Robust visual tracking based on variational auto-encoding Markov chain Monte Carlo
Junseok Kwon
Inf. Sci.1
2019 Sphere Generative Adversarial Network Based on Geometric Moment Matching
abstract
We propose sphere generative adversarial network (GAN), a novel integral probability metric (IPM)-based GAN. Sphere GAN uses the hypersphere to bound IPMs in the objective function. Thus, it can be trained stably. On the hypersphere, sphere GAN exploits the information of higher-order statistics of data using geometric moment matching, thereby providing more accurate results. In the paper, we mathematically prove the good properties of sphere GAN. In experiments, sphere GAN quantitatively and qualitatively surpasses recent state-of-the-art GANs for unsupervised image generation problems with the CIFAR-10, STL-10, and LSUN bedroom datasets. Source code is available at https://github.com/pswkiki/SphereGAN.
Sung Woo Park, Junseok Kwon
CVPR2
2019 Deep Meta Learning for Real-Time Target-Aware Visual Tracking
abstract
In this paper, we propose a novel on-line visual tracking framework based on the Siamese matching network and meta-learner network, which run at real-time speeds. Conventional deep convolutional feature-based discriminative visual tracking algorithms require continuous re-training of classifiers or correlation filters, which involve solving complex optimization tasks to adapt to the new appearance of a target object. To alleviate this complex process, our proposed algorithm incorporates and utilizes a meta-learner network to provide the matching network with new appearance information of the target objects by adding target-aware feature space. The parameters for the target-specific feature space are provided instantly from a single forward-pass of the meta-learner network. By eliminating the necessity of continuously solving complex optimization tasks in the course of tracking, experimental results demonstrate that our algorithm performs at a real-time speed while maintaining competitive performance among other state-of-the-art tracking algorithms.
Janghoon Choi, Junseok Kwon, Kyoung Mu Lee
ICCV2
2019 3D Point Cloud Generative Adversarial Network Based on Tree Structured Graph Convolutions
abstract
In this paper, we propose a novel generative adversarial network (GAN) for 3D point clouds generation, which is called tree-GAN. To achieve state-of-the-art performance for multi-class 3D point cloud generation, a tree-structured graph convolution network (TreeGCN) is introduced as a generator for tree-GAN. Because TreeGCN performs graph convolutions within a tree, it can use ancestor information to boost the representation power for features. To evaluate GANs for 3D point clouds accurately, we develop a novel evaluation metric called Fr\'echet point cloud distance (FPD). Experimental results demonstrate that the proposed tree-GAN outperforms state-of-the-art GANs in terms of both conventional metrics and FPD, and can generate point clouds for different semantic parts without prior knowledge.
Dong Wook Shu, Sung Woo Park, Junseok Kwon
ICCV3
2019 Learning to Remember Past to Predict Future for Visual Tracking
abstract
Fast and reliable adaptability to appearance variations of any target object has been the holy grail of visual tracking. Recently, Siamese-based trackers have demonstrated outstanding speed, however at the cost of adaptability and accuracy. We propose to model a temporal evolution of appearance features, allowing for adaptability without online training. Specifically, we introduce a memory-augmented convolutional recurrent neural network (RNN), named Past-to-Future (P2FNet), that takes appearance features as an input at each frame and predicts the next-frame features. RNN allows for fast adaptability to dynamically varying appearance, while the memory provides the generalization capability over longer sequences via template management. For reliability, we propose a new augmentation to train RNN to disregard corrupted features. A novel visualization method illustrates the reliable template management of the memory. The experimental results on benchmarks demonstrate the tracker shows competitive performance among real-time state-of-the-art trackers.
Sungyong Baik, Junseok Kwon, Kyoung Mu Lee
ICIP2
2019 Multi-Domain Attentive Detection Network
abstract
We present a novel object detection method called the multi-domain attentive detection network (MDADN). For robust object detection, we do not only use red, green, and blue (RGB) image data but also infrared data. The MDADN adds attention modules to each layer to weigh multi-domains of data differently in a channel-wise and spatial-wise manner, which yields channel- and spatial-aware networks. The MDADN accurately detects objects (e.g., vehicles) in challenging environments, including nighttime, rainy and foggy days, and environments where even humans find it difficult to detect objects. Experimental results on the FLIR dataset demonstrate that the MDADN outperforms state-of-the-art methods in terms of accuracy and speed. The ablation study shows that the attention module and the fusion of multiple data sources help to considerably improve the accuracy.
Sungmin Cho, Bo Won Choi, Do-Hwi Kim, Junseok Kwon
ICIP4
2019 Low-Lightgan: Low-Light Enhancement Via Advanced Generative Adversarial Network With Task-Driven Training
abstract
We propose a low-light enhancement method using an advanced generative adversarial network (GAN) and a task-driven training set. Unlike traditional training sets that only synthesize global illumination, we apply local illumination to make the training images. Furthermore, we enhance traditional GANs with spectral normalization and advanced loss functions, making training stable and leading to accurate results. Experimental results show that our method outperforms state-of-the art methods qualitatively and quantitatively and alleviates saturation problems in bright areas, which typically occur after traditional low-light enhancements.
Guisik Kim, Dokyeong Kwon, Junseok Kwon
ICIP3
2019 Depth-Controllable Very Deep Super-Resolution Network
abstract
Deep learning techniques not only have surpassed humans in several computer vision tasks, but also have achieved state-of-the-art performance on the super-resolution (SR) task, which enhances the resolution of reconstructed images from the observed low-resolution (LR) images. Very Deep Super- Resolution (VDSR) is a popular architecture with decent performance on the SR task. In general, the deeper the VDSR network, the better the performance. Say, on a computing device with limited computational resources, the SR task must support multiple resolutions or multiple output rates. For this multi-rate SR task, the state-of-the-art architectures require multiple networks with varying depths. However as the number of supported rates increases, so does the number of networks imposing incremental computational burden on the computing device. We propose a depth-controllable network and training principles for the multi-rate SR task. The proposed network is configured as a single network regardless of the number of supported rates. The inverse auxiliary loss and contiguous/progressive skip connections are presented to train the network end-to-end throughout the varying number of layers without biasing the performances of specific depths. With three data sets and three scaling factors, the proposed network is compared to the baseline network VDSR, along with the Super-Resolution Convolutional Neural Network (SRCNN). Our network not only requires a single network at varying rates, but also performs as well as the baseline networks in terms of Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM). Our model outperforms the baseline networks when the depth of layers is more than eleven.
Dohyun Kim 0001, Joongheon Kim, Junseok Kwon
IJCNN3
2019 Robust visual tracking with adaptive initial configuration and likelihood landscape analysis
abstract
Here, the authors propose a novel tracking algorithm that can automatically modify the initial configuration of a target to improve the tracking accuracy in subsequent frames. To achieve this goal, the authors’ method analyses the likelihood landscape (LL) for the image patch described by the initial configuration. A good configuration has a unimodal distribution with a steep shape in the LL. Using the LL analysis, the authors’ method improves the initial configuration, resulting in more accurate tracking results. The authors improve the conventional LL analysis based on two ideas. First, the authors’ method analyses the LL in the RGB space rather than the grey space. Second, the method introduces an additional criterion for a good configuration: a high likelihood value at the mode. The authors further enhance their method through post‐processing of the visual tracking results at each frame, where the estimated bounding boxes are modified by the LL analysis. The experimental results demonstrate that the authors’ advanced LL analysis helps improve the tracking accuracy of several baseline trackers on a visual tracking benchmark data set. In addition, the authors’ simple post‐processing technique significantly enhances the visual tracking performance in terms of precision and success rate.
Guisik Kim, Junseok Kwon
IET Comput. Vis.2
2019 Orthogonal object proposal and its application
abstract
In this study, the authors propose an object proposal algorithm that can accurately propose object candidate regions at each frame, despite noise in a video. Accordingly, they define three orthogonal planes, namely vertical–horizontal, temporal–vertical, and temporal–horizontal planes. As these planes are orthogonal, they are the most compact planes that can span the spatiotemporal space of a video. Their algorithm selects good object proposals for the vertical–horizontal plane with the help of the object proposal results of the other planes. Experimental results demonstrate that the proposed algorithm produces better object proposals than the baseline algorithm and other state‐of‐the‐art methods. In particular, their method provides more accurate object proposals in challenging environments with severe noise and background clutter. In addition, the object proposal results are utilised for visual tracking problems, and the experimental results show that their visual tracker outperforms recent deep‐learning‐based trackers.
Sung Woo Park, Junseok Kwon
IET Comput. Vis.2
2019 Small object segmentation with fully convolutional network based on overlapping domain decomposition
Jinhee Park 0001, Dokyeong Kwon, Bo Won Choi, Ga Young Kim, Kwang Yong Kim, Junseok Kwon
Mach. Vis. Appl.6
2018 Adaptive Patch Based Convolutional Neural Network for Robust Dehazing
abstract
We present a novel deep learning-based dehazing method using adaptive patch splits. Our method applies quad-tree decomposition to an input image, yielding multiple patches with adaptive sizes. Then, each patch is fed into a Convolutional Neural Network (CNN) and classified into a single transmission value, in which a transmission map comprises transmission values from all patches. Homogeneous regions in the image are typically decomposed into large patches. Thus the method can save computational cost. Non-homogeneous regions are divided into small patches, which helps preserve local details in a transmission map. To train CNN, we synthesize numerous hazy images from haze-free images. Experimental results demonstrate our method surpasses state- of-the-art deep learning based algorithms quantitatively and qualitatively.
Guisik Kim, Suhyeon Ha, Junseok Kwon
ICIP3
2018 Low-Complexity Online Model Selection with Lyapunov Control for Reward Maximization in Stabilized Real-Time Deep Learning Platforms
abstract
This paper proposes a low-complexity online model adaptation algorithm which dynamically selects an object detection algorithm among given/implemented algorithms in the system depending on workload-backlog. As well-studied in literature, there exists tradeoff between object detection accuracy and computation time (i.e., delay) because highly accurate algorithms generally take more time due to complicated deep neural network architectures. In our proposed algorithm, the accuracy is reformulated as reward; and the delay is modeled with queue. Based on this queue-based model, Lyapunov control inspired stochastic optimization is utilized for designing time-average reward maximization subject to stability in real-time object detection deep learning platforms. Moreover, our proposed algorithm solves closed-form equation in each model selection interval, thus the proposed algorithm takes low computational complexity. The performance of our proposed algorithm is evaluated via data-intensive real-world implementations under heavy workloads; and is verified that our proposed algorithm works as desired.
Dohyun Kim 0001, Junseok Kwon, Joongheon Kim
SMC2
2018 Real-time visual tracking by deep reinforced decision making
Janghoon Choi, Junseok Kwon, Kyoung Mu Lee
Comput. Vis. Image Underst.2
2018 Visual tracking based on edge field with object proposal association
Junseok Kwon, Hansung Lee
Image Vis. Comput.1
2017 Robust Pixel-wise Dehazing Algorithm based on Advanced Haze-Relevant Features
Guisik Kim, Junseok Kwon
BMVC2
2017 Leveraging observation uncertainty for robust visual tracking
Junseok Kwon, Radu Timofte, Luc Van Gool
Comput. Vis. Image Underst.1
2017 Adaptive Visual Tracking with Minimum Uncertainty Gap Estimation
abstract
A novel tracking algorithm is proposed, which robustly tracks a target by finding the state that minimizes the likelihood uncertainty. Likelihood uncertainty is estimated by determining the gap between the lower and upper bounds of likelihood. By minimizing the gap between the two bounds, the proposed method identifies the confident and reliable state of the target. In this study, the state that provides the Minimum Uncertainty Gap (MUG) between likelihood bounds is shown to be more reliable than the state that provides the maximum likelihood only, especially when severe illumination changes, occlusions, and pose variations occur. A rigorous derivation of the lower and upper bounds of the likelihood for the visual tracking problem is provided to address this issue. Additionally, an efficient inference algorithm that uses Interacting Markov Chain Monte Carlo (IMCMC) approach is presented to find the best state that maximizes the average of the lower and upper bounds of likelihood while minimizing the gap between the two bounds. We extend our method to update the target model adaptively. To update the model, the current observation is combined with a previous target model with the adaptive weight, which is calculated according to the goodness of the current observation. The goodness of the observation is measured using the proposed uncertainty gap estimation of likelihood. Experimental results demonstrate that the proposed method robustly tracks the target in realistic videos and outperforms conventional tracking methods.
Junseok Kwon, Kyoung Mu Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Tracking by switching state space models
Junseok Kwon, Ralf Dragon, Luc Van Gool
Comput. Vis. Image Underst.1
2016 PICASO: PIxel correspondences and SOft match selection for real-time tracking
Radu Timofte, Junseok Kwon, Luc Van Gool
Comput. Vis. Image Underst.2
2016 Joint Tracking and Ground Plane Estimation
abstract
We propose a novel framework that jointly estimates the ground plane and a target's motion trajectory. This results in improvements for both. Estimating their joint posterior is based on Particle Markov Chain Monte Carlo (Particle MCMC). In Particle MCMC, the best target state is inferred by a particle filter and the best ground plane is obtained by MCMC. Compared with conventional sampling methods that iteratively infer the best target states and ground plane parameters, our method infers them jointly. This reduces sampling errors drastically. Experimental results demonstrate that our method outperforms several state-of-the-art tracking methods, while the ground plane accuracy is also improved.
Junseok Kwon, Ralf Dragon, Luc Van Gool
IEEE Signal Process. Lett.1
2015 A Unified Framework for Event Summarization and Rare Event Detection from Multiple Views
abstract
A novel approach for event summarization and rare event detection is proposed. Unlike conventional methods that deal with event summarization and rare event detection independently, our method solves them in a single framework by transforming them into a graph editing problem. In our approach, a video is represented by a graph, each node of which indicates an event obtained by segmenting the video spatially and temporally. The edges between nodes describe the relationship between events. Based on the degree of relations, edges have different weights. After learning the graph structure, our method finds subgraphs that represent event summarization and rare events in the video by editing the graph, that is, merging its subgraphs or pruning its edges. The graph is edited to minimize a predefined energy model with the Markov Chain Monte Carlo (MCMC) method. The energy model consists of several parameters that represent the causality, frequency, and significance of events. We design a specific energy model that uses these parameters to satisfy each objective of event summarization and rare event detection. The proposed method is extended to obtain event summarization and rare event detection results across multiple videos captured from multiple views. For this purpose, the proposed method independently learns and edits each graph of individual videos for event summarization or rare event detection. Then, the method matches the extracted multiple graphs to each other, and constructs a single composite graph that represents event summarization or rare events from multiple views. Experimental results show that the proposed approach accurately summarizes multiple videos in a fully unsupervised manner. Moreover, the experiments demonstrate that the approach is advantageous in detecting rare transition of events.
Junseok Kwon, Kyoung Mu Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Interval Tracker: Tracking by Interval Analysis
abstract
This paper proposes a robust tracking method that uses interval analysis. Any single posterior model necessarily includes a modeling uncertainty (error), and thus, the posterior should be represented as an interval of probability. Then, the objective of visual tracking becomes to find the best state that maximizes the posterior and minimizes its interval simultaneously. By minimizing the interval of the posterior, our method can reduce the modeling uncertainty in the posterior. In this paper, the aforementioned objective is achieved by using the M4 estimation, which combines the Maximum a Posterior (MAP) estimation with Minimum Mean-Square Error (MMSE), Maximum Likelihood (ML), and Minimum Interval Length (MIL) estimations. In the M4 estimation, our method maximizes the posterior over the state obtained by the MMSE estimation. The method also minimizes interval of the posterior by reducing the gap between the lower and upper bounds of the posterior. The gap is reduced when the likelihood is maximized by the ML estimation and the interval length of the state is minimized by the MIL estimation. The experimental results demonstrate that M4 estimation can be easily integrated into conventional tracking methods and can greatly enhance their tracking accuracy. In several challenging datasets, our method outperforms state-of-the-art tracking methods.
Junseok Kwon, Kyoung Mu Lee
CVPR1
2014 Robust Visual Tracking with Double Bounding Box Model
Junseok Kwon, Junha Roh, Kyoung Mu Lee, Luc Van Gool
ECCV (1)1
2014 Appearances Can Be Deceiving: Learning Visual Tracking from Few Trajectory Annotations
Santiago Manen, Junseok Kwon, Matthieu Guillaumin, Luc Van Gool
ECCV (5)2
2014 Tracking by Sampling and IntegratingMultiple Trackers
abstract
We propose the visual tracker sampler, a novel tracking algorithm that can work robustly in challenging scenarios, where several kinds of appearance and motion changes of an object can occur simultaneously. The proposed tracking algorithm accurately tracks a target by searching for appropriate trackers in each frame. Since the real-world tracking environment varies severely over time, the trackers should be adapted or newly constructed depending on the current situation, so that each specific tracker takes charge of a certain change in the object. To do this, our method obtains several samples of not only the states of the target but also the trackers themselves during the sampling process. The trackers are efficiently sampled using the Markov Chain Monte Carlo (MCMC) method from the predefined tracker space by proposing new appearance models, motion models, state representation types, and observation types, which are the important ingredients of visual trackers. All trackers are then integrated into one compound tracker through an Interacting MCMC (IMCMC) method, in which the trackers interactively communicate with one another while running in parallel. By exchanging information with others, each tracker further improves its performance, thus increasing overall tracking performance. Experimental results show that our method tracks the object accurately and reliably in realistic videos, where appearance and motion drastically change over time, and outperforms even state-of-the-art tracking methods.
Junseok Kwon, Kyoung Mu Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Minimum Uncertainty Gap for Robust Visual Tracking
abstract
We propose a novel tracking algorithm that robustly tracks the target by finding the state which minimizes uncertainty of the likelihood at current state. The uncertainty of the likelihood is estimated by obtaining the gap between the lower and upper bounds of the likelihood. By minimizing the gap between the two bounds, our method finds the confident and reliable state of the target. In the paper, the state that gives the Minimum Uncertainty Gap (MUG) between likelihood bounds is shown to be more reliable than the state which gives the maximum likelihood only, especially when there are severe illumination changes, occlusions, and pose variations. A rigorous derivation of the lower and upper bounds of the likelihood for the visual tracking problem is provided to address this issue. Additionally, an efficient inference algorithm using Interacting Markov Chain Monte Carlo is presented to find the best state that maximizes the average of the lower and upper bounds of the likelihood and minimizes the gap between two bounds simultaneously. Experimental results demonstrate that our method successfully tracks the target in realistic videos and outperforms conventional tracking methods.
Junseok Kwon, Kyoung Mu Lee
CVPR1
2013 Wang-Landau Monte Carlo-Based Tracking Methods for Abrupt Motions
abstract
We propose a novel tracking algorithm based on the Wang-Landau Monte Carlo (WLMC) sampling method for dealing with abrupt motions efficiently. Abrupt motions cause conventional tracking methods to fail because they violate the motion smoothness constraint. To address this problem, we introduce the Wang-Landau sampling method and integrate it into a Markov Chain Monte Carlo (MCMC)-based tracking framework. By employing the novel density-of-states term estimated by the Wang-Landau sampling method into the acceptance ratio of MCMC, our WLMC-based tracking method alleviates the motion smoothness constraint and robustly tracks the abrupt motions. Meanwhile, the marginal likelihood term of the acceptance ratio preserves the accuracy in tracking smooth motions. The method is then extended to obtain good performance in terms of scalability, even on a high-dimensional state space. Hence, it covers drastic changes in not only position but also scale of a target. To achieve this, we modify our method by combining it with the N-fold way algorithm and present the N-Fold Wang-Landau (NFWL)-based tracking method. The N-fold way algorithm helps estimate the density-of-states with a smaller number of samples. Experimental results demonstrate that our approach efficiently samples the states of the target, even in a whole state space, without loss of time, and tracks the target accurately and robustly when position and scale are changing severely.
Junseok Kwon, Kyoung Mu Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Highly Nonrigid Object Tracking via Patch-Based Dynamic Appearance Modeling
abstract
A novel tracking algorithm is proposed for targets with drastically changing geometric appearances over time. To track such objects, we develop a local patch-based appearance model and provide an efficient online updating scheme that adaptively changes the topology between patches. In the online update process, the robustness of each patch is determined by analyzing the likelihood landscape of the patch. Based on this robustness measure, the proposed method selects the best feature for each patch and modifies the patch by moving, deleting, or newly adding it over time. Moreover, a rough object segmentation result is integrated into the proposed appearance model to further enhance it. The proposed framework easily obtains segmentation results because the local patches in the model serve as good seeds for the semi-supervised segmentation task. To solve the complexity problem attributable to the large number of patches, the Basin Hopping (BH) sampling method is introduced into the tracking framework. The BH sampling method significantly reduces computational complexity with the help of a deterministic local optimizer. Thus, the proposed appearance model could utilize a sufficient number of patches. The experimental results show that the present approach could track objects with drastically changing geometric appearance accurately and robustly.
Junseok Kwon, Kyoung Mu Lee
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 A unified framework for event summarization and rare event detection
abstract
In this paper, we have proposed an unified framework for event summarization and rare event detection and presented the graph-structure learning and editing method to solve these problems efficiently. The experimental results demonstrated that the proposed method outperformed conventional algorithms in complex and crowded public scenes by exploiting and utilizing causality, frequency, and significance of relations of events.
Junseok Kwon, Kyoung Mu Lee
CVPR1
2012 Robust visual tracking using autoregressive hidden Markov Model
abstract
Recent studies on visual tracking have shown significant improvement in accuracy by handling the appearance variations of the target object. Whereas most studies present schemes to extract the time-invariant characteristics of the target and adaptively update the appearance model, the present paper concentrates on modeling the probabilistic dependency between sequential target appearances (Fig. 1-(a)). To actualize this interest, a new Bayesian tracking framework is formulated under the autoregressive Hidden Markov Model (AR-HMM), where the probabilistic dependency between sequential target appearances is implied. During the learning phase at each time step, the proposed tracker separates formerly seen target samples into several clusters based on their visual similarity, and learns cluster-specific classifiers as multiple appearance models, each of which represents a certain type of the target appearance. Then the dependency between these appearance models is learned. During the searching phase, the target state is estimated by inferring the most probable appearance model under the consideration of its dependency on formerly utilized appearance models. The proposed method is tested on 12 challenging video sequences containing targets with abrupt appearance variations, and demonstrates that it outperforms current state-of-the-art methods in accuracy.
Dong Woo Park, Junseok Kwon, Kyoung Mu Lee
CVPR2
2011 Tracking by Sampling Trackers
abstract
We propose a novel tracking framework called visual tracker sampler that tracks a target robustly by searching for the appropriate trackers in each frame. Since the real-world tracking environment varies severely over time, the trackers should be adapted or newly constructed depending on the current situation. To do this, our method obtains several samples of not only the states of the target but also the trackers themselves during the sampling process. The trackers are efficiently sampled using the Markov Chain Monte Carlo method from the predefined tracker space by proposing new appearance models, motion models, state representation types, and observation types, which are the basic important components of visual trackers. Then, the sampled trackers run in parallel and interact with each other while covering various target variations efficiently. The experiment demonstrates that our method tracks targets accurately and robustly in the real-world tracking environments and outperforms the state-of-the-art tracking methods.
Junseok Kwon, Kyoung Mu Lee
ICCV1
2010 Visual tracking decomposition
abstract
We propose a novel tracking algorithm that can work robustly in a challenging scenario such that several kinds of appearance and motion changes of an object occur at the same time. Our algorithm is based on a visual tracking decomposition scheme for the efficient design of observation and motion models as well as trackers. In our scheme, the observation model is decomposed into multiple basic observation models that are constructed by sparse principal component analysis (SPCA) of a set of feature templates. Each basic observation model covers a specific appearance of the object. The motion model is also represented by the combination of multiple basic motion models, each of which covers a different type of motion. Then the multiple basic trackers are designed by associating the basic observation models and the basic motion models, so that each specific tracker takes charge of a certain change in the object. All basic trackers are then integrated into one compound tracker through an interactive Markov Chain Monte Carlo (IMCMC) framework in which the basic trackers communicate with one another interactively while run in parallel. By exchanging information with others, each tracker further improves its performance, which results in increasing the whole performance of tracking. Experimental results show that our method tracks the object accurately and reliably in realistic videos where the appearance and motion are drastically changing over time.
Junseok Kwon, Kyoung Mu Lee
CVPR1
2009 Tracking of a non-rigid object via patch-based dynamic appearance modeling and adaptive Basin Hopping Monte Carlo sampling
abstract
We propose a novel tracking algorithm for the target of which geometric appearance changes drastically over time. To track it, we present a local patch-based appearance model and provide an efficient scheme to evolve the topology between local patches by on-line update. In the process of on-line update, the robustness of each patch in the model is estimated by a new method of measurement which analyzes the landscape of local mode of the patch. This patch can be moved, deleted or newly added, which gives more flexibility to the model. Additionally, we introduce the Basin Hopping Monte Carlo (BHMC) sampling method to our tracking problem to reduce the computational complexity and deal with the problem of getting trapped in local minima. The BHMC method makes it possible for our appearance model to consist of enough numbers of patches. Since BHMC uses the same local optimizer that is used in the appearance modeling, it can be efficiently integrated into our tracking framework. Experimental results show that our approach tracks the object whose geometric appearance is drastically changing, accurately and robustly.
Junseok Kwon, Kyoung Mu Lee
CVPR1
2008 Tracking of Abrupt Motion Using Wang-Landau Monte Carlo Estimation
Junseok Kwon, Kyoung Mu Lee
ECCV (1)1