VLDB 2026 Research / reviewers in the wild / expert
Junsik Kim 0001
dblp:89/6937-1
· DBLP profile ↗
32ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0003-2555-5232ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Virtual Multiplex Staining for Histological Images Using a Marker-Wise Conditioned Diffusion ModelabstractMultiplex imaging is revolutionizing pathology by enabling the simultaneous visualization of multiple biomarkers within tissue samples, providing molecular-level insights that traditional hematoxylin and eosin (H&E) staining cannot provide. However, the complexity and cost of multiplex data acquisition have hindered its widespread adoption. Additionally, most existing large repositories of H&E images lack corresponding multiplex images, limiting opportunities for multi-modal analysis. To address these challenges, we leverage recent advances in latent diffusion models (LDMs), which excel at modeling complex data distributions by utilizing their powerful priors for fine-tuning to a target domain. In this paper, we introduce a novel framework for virtual multiplex staining that utilizes pretrained LDM parameters to generate multiplex images from H&E images using a conditional diffusion model. Our approach enables marker-by-marker generation by conditioning the diffusion model on each marker, while sharing the same architecture across all markers. To tackle the challenge of varying pixel value distributions across different marker stains and to improve inference speed, we fine-tune the model for single-step sampling, enhancing both color contrast fidelity and inference efficiency through pixel-level loss functions. We validate our framework on two publicly available datasets, notably demonstrating its effectiveness in generating up to 18 different marker types with improved accuracy, a substantial increase over the 2-3 marker types achieved in previous approaches. This validation highlights the potential of our framework, pioneering virtual multiplex staining. Finally, this paper bridges the gap between H&E and multiplex imaging, potentially enabling retrospective studies and large-scale analyses of existing H&E image repositories. Hyun-Jic Oh, Junsik Kim 0001, Zhiyi Shi, Yu-An Chen, Peter K. Sorger, Hanspeter Pfister, Won-Ki Jeong |
AAAI | 2 |
| 2026 | Normality-calibrated autoencoder for unsupervised anomaly detection on data contamination
Jongmin Yu, Minkyung Kim 0001, Junsik Kim 0001, Hyeontaek Oh |
Neurocomputing | 3 |
| 2025 | Toward Interactive Sound Source Localization: Better Align Sight and Sound!abstractRecent studies on learning-based sound source localization have primarily focused on localization performance. However, prior work and existing benchmarks often overlook a crucial aspect: cross-modal interaction, which is essential for interactive sound source localization. This interaction is vital for understanding semantically matched or mismatched audio-visual events, such as silent objects or true sound sources among multiple objects. In this work, we comprehensively examine the cross-modal interaction of existing methods, benchmarks, evaluation metrics, and cross-modal understanding tasks. We identify the overlooked points of previous studies and make several contributions to address them. First, we propose a learning framework that incorporates retrieval-based and hand-crafted augmentation techniques, enhancing cross-modal interaction through cross-modal alignment. Second, we introduce new evaluation metrics to accurately and rigorously assess localization methods, focusing on both localization performance and cross-modal interaction. Third, to thoroughly analyze interactive sound source localization, we present a new semi-synthetic benchmark with diverse categorical combinations. Finally, we evaluate both interactive sound source localization and auxiliary cross-modal retrieval tasks, benchmarking competing methods alongside our own. Our new benchmark and evaluation metrics reveal that previous methods struggle with interactive sound source localization tasks, largely due to their limited cross-modal interaction capabilities. Our method, which features enhanced cross-modal alignment, demonstrates superior sound source localization and cross-modal interaction performance. This work provides the most comprehensive analysis of sound source localization to date, with extensive validation of competing methods on both existing and new benchmarks using both new and standard evaluation metrics. Arda Senocak, Hyeonggon Ryu, Junsik Kim 0001, Tae-Hyun Oh, Hanspeter Pfister, Joon Son Chung |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Joint-Task Regularization for Partially Labeled Multi-Task LearningabstractMulti-task learning has become increasingly popular in the machine learning field, but its practicality is hindered by the need for large, labeled datasets. Most multi-task learning methods depend on fully labeled datasets wherein each input example is accompanied by ground-truth labels for all target tasks. Unfortunately, curating such datasets can be prohibitively expensive and impractical, especially for dense prediction tasks which require per-pixel labels for each image. With this in mind, we propose Joint-Task Regularization (JTR), an intuitive technique which leverages cross-task relations to simultaneously regularize all tasks in a single joint-task latent space to improve learning when data is not fully labeled for all tasks. JTR stands out from existing approaches in that it regularizes all tasks jointly rather than separately in pairs-therefore, it achieves linear complexity relative to the number of tasks while previous methods scale quadratically. To demonstrate the validity of our approach, we extensively benchmark our method across a wide variety of partially labeled scenarios based on NYU-v2, Cityscapes, and Taskonomy. Kento Nishi, Junsik Kim 0001, Wanhua Li 0001, Hanspeter Pfister |
CVPR | 2 |
| 2024 | $\mathrm R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding
Ye Liu 0002, Jixuan He, Wanhua Li 0001, Junsik Kim 0001, Donglai Wei 0001, Hanspeter Pfister, Chang Wen Chen |
ECCV (41) | 4 |
| 2024 | Multimodal Learning for Embryo Viability Prediction in Clinical IVF
Junsik Kim 0001, Zhiyi Shi, Davin Jeong, Johannes Knittel, Helen Y. Yang, Yonghyun Song, Wanhua Li 0001, Yicong Li 0002, Dalit Ben-Yosef, Daniel Needleman, Hanspeter Pfister |
MICCAI (5) | 1 |
| 2024 | MoRA: LoRA Guided Multi-modal Disease Diagnosis with Missing Modality
Zhiyi Shi, Junsik Kim 0001, Wanhua Li 0001, Yicong Li 0002, Hanspeter Pfister |
MICCAI (3) | 2 |
| 2024 | Blurry Video Compression A Trade-off between Visual Enhancement and Data CompressionabstractExisting video compression (VC) methods primarily aim to reduce the spatial and temporal redundancies between consecutive frames in a video while preserving its quality. In this regard, previous works have achieved remarkable results on videos acquired under specific settings such as instant (known) exposure time and shutter speed which often result in sharp videos. However, when these methods are evaluated on videos captured under different temporal priors, which lead to degradations like motion blur and low frame rate, they fail to maintain the quality of the contents. In this work, we tackle the VC problem in a general scenario where a given video can be blurry due to predefined camera settings or dynamics in the scene. By exploiting the natural trade-off between visual enhancement and data compression, we formulate VC as a min-max optimization problem and propose an effective framework and training strategy to tackle the problem. Extensive experimental results on several benchmark datasets confirm the effectiveness of our method compared to several state-of-the-art VC approaches. Dawit Mureja Argaw, Junsik Kim 0001, In-So Kweon |
WACV | 2 |
| 2024 | ENInst: Enhancing weakly-supervised low-shot instance segmentation
Moon Ye-Bin, Dongmin Choi, Yongjin Kwon, Junsik Kim 0001, Tae-Hyun Oh |
Pattern Recognit. | 4 |
| 2024 | An Iterative Method for Unsupervised Robust Anomaly Detection Under Data ContaminationabstractMost deep anomaly detection models are based on learning normality from datasets due to the difficulty of defining abnormality by its diverse and inconsistent nature. Therefore, it has been a common practice to learn normality under the assumption that anomalous data are absent in a training dataset, which we call normality assumption. However, in practice, the normality assumption is often violated due to the nature of real data distributions that includes anomalous tails, i.e., a contaminated dataset. Thereby, the gap between the assumption and actual training data affects detrimentally in learning of an anomaly detection model. In this work, we propose a learning framework to reduce this gap and achieve better normality representation. Our key idea is to identify sample-wise normality and utilize it as an importance weight, which is updated iteratively during the training. Our framework is designed to be model-agnostic and hyperparameter insensitive so that it applies to a wide range of existing methods without careful parameter tuning. We apply our framework to three different representative approaches of deep anomaly detection that are classified into one-class classification-, probabilistic model-, and reconstruction-based approaches. In addition, we address the importance of a termination condition for iterative methods and propose a termination criterion inspired by the anomaly detection objective. We validate that our framework improves the robustness of the anomaly detection models under different levels of contamination ratios on five anomaly detection benchmark datasets and two image datasets. On various contaminated datasets, our framework improves the performance of three representative anomaly detection methods, measured by area under the ROC curve. Minkyung Kim 0001, Jongmin Yu, Junsik Kim 0001, Tae-Hyun Oh, Jun Kyun Choi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Weakly Supervised Contrastive Learning for Unsupervised Vehicle ReidentificationabstractReidentification (Re-id) of vehicles in a multicamera system is an essential process for traffic control automation. Previously, there have been efforts to reidentify vehicles based on shots of images with identity (id) labels, where the model training relies on the quality and quantity of the labels. However, labeling vehicle ids is a labor-intensive procedure. Instead of relying on expensive labels, we propose to exploit camera and tracklet ids that are automatically obtainable during a Re-id dataset construction. In this article, we present weakly supervised contrastive learning (WSCL) and domain adaptation (DA) techniques using camera and tracklet ids for unsupervised vehicle Re-id. We define each camera id as a subdomain and tracklet id as a label of a vehicle within each subdomain, i.e., weak label in the Re-id scenario. Within each subdomain, contrastive learning using tracklet ids is applied to learn a representation of vehicles. Then, DA is performed to match vehicle ids across the subdomains. We demonstrate the effectiveness of our method for unsupervised vehicle Re-id using various benchmarks. Experimental results show that the proposed method outperforms the recent state-of-the-art unsupervised Re-id methods. The source code is publicly available on https://github.com/andreYoo/WSCL_VeReid. Jongmin Yu, Hyeontaek Oh, Minkyung Kim 0001, Junsik Kim 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Sound Source Localization is All about Cross-Modal AlignmentabstractHumans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localization perspective. However, prior arts and existing benchmarks do not account for a more important aspect of the problem, cross-modal semantic understanding, which is essential for genuine sound source localization. Cross-modal semantic understanding is important in understanding semantically mismatched audio-visual events, e.g., silent objects, or off-screen sounds. To account for this, we propose a cross-modal alignment task as a joint task with sound source localization to better learn the interaction between audio and visual modalities. Thereby, we achieve high localization performance with strong cross-modal semantic understanding. Our method outperforms the state-of-the-art approaches in both sound source localization and cross-modal retrieval. Our work suggests that jointly tackling both tasks is necessary to conquer genuine sound source localization. Arda Senocak, Hyeonggon Ryu, Junsik Kim 0001, Tae-Hyun Oh, Hanspeter Pfister, Joon Son Chung |
ICCV | 3 |
| 2023 | Event-Specific Audio-Visual Fusion Layers: A Simple and New Perspective on Video UnderstandingabstractTo understand our surrounding world, our brain is continuously inundated with multisensory information and their complex interactions coming from the outside world at any given moment. While processing this information might seem effortless for human brains, it is challenging to build a machine that can perform similar tasks since complex interactions cannot be dealt with a single type of integration but require more sophisticated approaches. In this paper, we propose a new simple method to address the multisensory integration in video understanding. Unlike previous works where a single fusion type is used, we design a multi-head model with individual event-specific layers to deal with different audio-visual relationships, enabling different ways of audio-visual fusion. Experimental results show that our event-specific layers can discover unique properties of the audio-visual relationships in the videos, e.g., semantically matched moments, and rhythmic events. Moreover, although our network is trained with single labels, our multi-head design can inherently output additional semantically meaningful multi-labels for a video. As an application, we demonstrate that our proposed method can expose the extent of event-characteristics of popular benchmark datasets. Arda Senocak, Junsik Kim 0001, Tae-Hyun Oh, Dingzeyu Li, In-So Kweon |
WACV | 2 |
| 2023 | Active anomaly detection based on deep one-class classification
Minkyung Kim 0001, Junsik Kim 0001, Jongmin Yu, Jun Kyun Choi |
Pattern Recognit. Lett. | 2 |
| 2022 | ML-BPM: Multi-teacher Learning with Bidirectional Photometric Mixing for Open Compound Domain Adaptation in Semantic Segmentation
Sungsu Hur, Seokju Lee, Junsik Kim 0001, In-So Kweon |
ECCV (34) | 4 |
| 2022 | Learning Sound Localization Better from Semantically Similar SamplesabstractThe objective of this work is to localize the sound sources in visual scenes. Existing audio-visual works employ contrastive learning by assigning corresponding audio-visual pairs from the same source as positives while randomly mismatched pairs as negatives. However, these negative pairs may contain semantically matched audio-visual information. Thus, these semantically correlated pairs, "hard positives", are mistakenly grouped as negatives. Our key contribution is showing that hard positives can give similar response maps to the corresponding pairs. Our approach incorporates these hard positives by adding their response maps into a contrastive learning objective directly. We demonstrate the effectiveness of our approach on VGG-SS and SoundNet-Flickr test sets, showing favorable performance to the state-of-the-art methods. Arda Senocak, Hyeonggon Ryu, Junsik Kim 0001, In-So Kweon |
ICASSP | 3 |
| 2022 | Camera-Tracklet-Aware Contrastive Learning for Unsupervised Vehicle Re-IdentificationabstractRecently, vehicle re-identification methods based on deep learning constitute remarkable achievement. However, this achievement requires large-scale and well-annotated datasets. In constructing the dataset, assigning globally available identities (Ids) to vehicles captured from a great number of cameras is labour-intensive, because it needs to consider their subtle appearance differences or viewpoint variations. In this paper, we propose camera-tracklet-aware contrastive learning (CTACL) using the multi-camera tracklet information without vehicle identity labels. The proposed CTACL divides an unlabelled domain, i.e., entire vehicle images, into multiple camera-level subdomains and conducts contrastive learning within and beyond the subdomains. The positive and negative samples for contrastive learning are defined using tracklet Ids of each camera. Additionally, the domain adaptation across camera networks is introduced to improve the generalisation performance of learnt representations and alleviate the performance degradation resulted from the domain gap between the subdomains. We demonstrate the effectiveness of our approach on video-based and image-based vehicle Re-ID datasets. Experimental results show that the proposed method outperforms the recent state-of-the-art unsupervised vehicle Re-ID methods. The source code for this paper is publicly available on https://github.com/andreYoo/CTAM-CTACL-VVReID.git. Jongmin Yu, Junsik Kim 0001, Minkyung Kim 0001, Hyeontaek Oh |
ICRA | 2 |
| 2022 | Less Can Be More: Sound Source Localization With a Classification ModelabstractIn this paper, we tackle sound localization as a natural outcome of the audio-visual video classification problem. Differently from the existing sound localization approaches, we do not use any explicit sub-modules or training mechanisms but use simple cross-modal attention on top of the representations learned by a classification loss. Our key contribution is to show that a simple audio-visual classification model has the ability to localize sound sources accurately and to give on par performance with state-of-the-art methods by proving that indeed "less is more". Furthermore, we propose potential applications that can be built based on our model. First, we introduce informative moment selection to enhance the localization task learning in the existing approaches compare to mid-frame usage. Then, we introduce a pseudo bounding box generation procedure that can significantly boost the performance of the existing methods in semi-supervised settings or be used for large-scale automatic annotation with minimal effort from any video dataset. Arda Senocak, Hyeonggon Ryu, Junsik Kim 0001, In-So Kweon |
WACV | 3 |
| 2021 | Motion-blurred Video Interpolation and ExtrapolationabstractAbrupt motion of camera or objects in a scene result in a blurry video, and therefore recovering high quality video requires two types of enhancements: visual enhancement and temporal upsampling. A broad range of research attempted to recover clean frames from blurred image sequences or temporally upsample frames by interpolation, yet there are very limited studies handling both problems jointly. In this work, we present a novel framework for deblurring, interpolating and extrapolating sharp frames from a motion-blurred video in an end-to-end manner. We design our framework by first learning the pixel-level motion that caused the blur from the given inputs via optical flow estimation and then predict multiple clean frames by warping the decoded features with the estimated flows. To ensure temporal coherence across predicted frames and address potential temporal ambiguity, we propose a simple, yet effective flow-based rule. The effectiveness and favorability of our approach are highlighted through extensive qualitative and quantitative evaluations on motion-blurred datasets from high speed videos. Dawit Mureja Argaw, Junsik Kim 0001, François Rameau, In-So Kweon |
AAAI | 2 |
| 2021 | Optical Flow Estimation from a Single Motion-blurred ImageabstractIn most of computer vision applications, motion blur is regarded as an undesirable artifact. However, it has been shown that motion blur in an image may have practical interests in fundamental computer vision problems. In this work, we propose a novel framework to estimate optical flow from a single motion-blurred image in an end-to-end manner. We design our network with transformer networks to learn globally and locally varying motions from encoded features of a motion-blurred input, and decode left and right frame features without explicit frame supervision. A flow estimator network is then used to estimate optical flow from the decoded features in a coarse-to-fine manner. We qualitatively and quantitatively evaluate our model through a large set of experiments on synthetic and real motion-blur datasets. We also provide in-depth analysis of our model in connection with related approaches to highlight the effectiveness and favorability of our approach. Furthermore, we showcase the applicability of the flow estimated by our method on deblurring and moving object segmentation tasks. Dawit Mureja Argaw, Junsik Kim 0001, François Rameau, Jae-Won Cho, In-So Kweon |
AAAI | 2 |
| 2021 | ResNet or DenseNet? Introducing Dense Shortcuts to ResNetabstractResNet or DenseNet? Nowadays, most deep learning based approaches are implemented with seminal backbone networks, among them the two arguably most famous ones are ResNet and DenseNet. Despite their competitive performance and overwhelming popularity, inherent drawbacks exist for both of them. For ResNet, the identity shortcut that stabilizes training might limit its representation capacity, and DenseNet mitigates it with multi-layer feature concatenation. However, the dense concatenation causes a new problem of requiring high GPU memory and more training time. Partially due to this, it is not a trivial choice between ResNet and DenseNet. This paper provides a unified perspective of dense summation to analyze them, which facilitates a better understanding of their core difference. We further propose dense weighted normalized shortcuts as a solution to the dilemma between them. Our proposed dense shortcut inherits the design philosophy of simple design in ResNet and DenseNet. On several benchmark datasets, the experimental results show that the proposed DSNet achieves significantly better results than ResNet, and achieves comparable performance as DenseNet but requiring fewer computation resources. Chaoning Zhang, Philipp Benz, Dawit Mureja Argaw, Seokju Lee, Junsik Kim 0001, François Rameau, Jean-Charles Bazin, In-So Kweon |
WACV | 5 |
| 2021 | Learning to Localize Sound Sources in Visual Scenes: Analysis and ApplicationsabstractVisual events are usually accompanied by sounds in our daily lives. However, can the machines learn to correlate the visual scene and sound, as well as localize the sound source only by observing them like humans? To investigate its empirical learnability, in this work we first present a novel unsupervised algorithm to address the problem of localizing sound sources in visual scenes. In order to achieve this goal, a two-stream network structure which handles each modality with attention mechanism is developed for sound source localization. The network naturally reveals the localized response in the scene without human annotation. In addition, a new sound source dataset is developed for performance evaluation. Nevertheless, our empirical evaluation shows that the unsupervised method generates false conclusions in some cases. Thereby, we show that this false conclusion cannot be fixed without human prior knowledge due to the well-known correlation and causality mismatch misconception. To fix this issue, we extend our network to the supervised and semi-supervised network settings via a simple modification due to the general architecture of our two-stream network. We show that the false conclusions can be effectively corrected even with a small amount of supervision, i.e., semi-supervised setup. Furthermore, we present the versatility of the learned audio and visual embeddings on the cross-modal content alignment and we extend this proposed algorithm to a new application, sound saliency based automatic camera view panning in 360 degree videos. Arda Senocak, Tae-Hyun Oh, Junsik Kim 0001, Ming-Hsuan Yang 0001, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | DeepPTZ: Deep Self-Calibration for PTZ CamerasabstractRotating and zooming cameras, also called PTZ (Pan-Tilt-Zoom) cameras, are widely used in modern surveillance systems. While their zooming ability allows acquiring detailed images of the scene, it also makes their calibration more challenging since any zooming action results in a modification of their intrinsic parameters. Therefore, such camera calibration has to be computed online; this process is called self-calibration. In this paper, given an image pair captured by a PTZ camera, we propose a deep learning based approach to automatically estimate the focal length and distortion parameters of both images as well as the rotation angles between them. The proposed approach relies on a dual-Siamese structure, imposing bidirectional constraints. The proposed network is trained on a large-scale dataset automatically generated from a set of panoramas. Empirically, we demonstrate that our proposed approach achieves competitive performance with respect to both deep learning based and traditional state-of-the art methods. Our code and model will be publicly available at https://github.com/ChaoningZhang/DeepPTZ. Chaoning Zhang, François Rameau, Junsik Kim 0001, Dawit Mureja Argaw, Jean-Charles Bazin, In-So Kweon |
WACV | 3 |
| 2019 | Visuomotor Understanding for Representation Learning of Driving Scenes
Seokju Lee, Junsik Kim 0001, Tae-Hyun Oh, Yongseop Jeong, Donggeun Yoo, Stephen Lin 0001, In-So Kweon |
BMVC | 2 |
| 2019 | Revisiting Residual Networks with Nonlinear Shortcuts
Chaoning Zhang, François Rameau, Seokju Lee, Junsik Kim 0001, Philipp Benz, Dawit Mureja Argaw, Jean-Charles Bazin, In-So Kweon |
BMVC | 4 |
| 2019 | Variational Prototyping-Encoder: One-Shot Learning With Prototypical ImagesabstractIn daily life, graphic symbols, such as traffic signs and brand logos, are ubiquitously utilized around us due to its intuitive expression beyond language boundary. We tackle an open-set graphic symbol recognition problem by one-shot classification with prototypical images as a single training example for each novel class. We take an approach to learn a generalizable embedding space for novel tasks. We propose a new approach called variational prototyping-encoder (VPE) that learns the image translation task from real-world input images to their corresponding prototypical images as a meta-task. As a result, VPE learns image similarity as well as prototypical concepts which differs from widely used metric learning based approaches. Our experiments with diverse datasets demonstrate that the proposed VPE performs favorably against competing metric learning based one-shot methods. Also, our qualitative analyses show that our meta-task induces an effective embedding space suitable for unseen data representation. Junsik Kim 0001, Tae-Hyun Oh, Seokju Lee, In-So Kweon |
CVPR | 1 |
| 2019 | Robust and Globally Optimal Manhattan Frame Estimation in Near Real TimeabstractMost man-made environments, such as urban and indoor scenes, consist of a set of parallel and orthogonal planar structures. These structures are approximated by the Manhattan world assumption, in which notion can be represented as a Manhattan frame (MF). Given a set of inputs such as surface normals or vanishing points, we pose an MF estimation problem as a consensus set maximization that maximizes the number of inliers over the rotation search space. Conventionally, this problem can be solved by a branch-and-bound framework, which mathematically guarantees global optimality. However, the computational time of the conventional branch-and-bound algorithms is rather far from real-time. In this paper, we propose a novel bound computation method on an efficient measurement domain for MF estimation, i.e., the extended Gaussian image (EGI). By relaxing the original problem, we can compute the bound with a constant complexity, while preserving global optimality. Furthermore, we quantitatively and qualitatively demonstrate the performance of the proposed method for various synthetic and real-world data. We also show the versatility of our approach through three different applications: extension to multiple MF estimation, 3D rotation based video stabilization, and vanishing point estimation (line clustering). Kyungdon Joo, Tae-Hyun Oh, Junsik Kim 0001, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Co-Domain Embedding Using Deep Quadruplet Networks for Unseen Traffic Sign RecognitionabstractRecent advances in visual recognition show overarching success by virtue of large amounts of supervised data. However, the acquisition of a large supervised dataset is often challenging. This is also true for intelligent transportation applications, i.e., traffic sign recognition. For example, a model trained with data of one country may not be easily generalized to another country without much data. We propose a novel feature embedding scheme for unseen class classification when the representative class template is given. Traffic signs, unlike other objects, have official images. We perform co-domain embedding using a quadruple relationship from real and synthetic domains. Our quadruplet network fully utilizes the explicit pairwise similarity relationships among samples from different domains. We validate our method on three datasets with two experiments involving one-shot classification and feature generalization. The results show that the proposed method outperforms competing approaches on both seen and unseen classes. Junsik Kim 0001, Seokju Lee, Tae-Hyun Oh, In-So Kweon |
AAAI | 1 |
| 2018 | Learning to Localize Sound Source in Visual ScenesabstractVisual events are usually accompanied by sounds in our daily lives. We pose the question: Can the machine learn the correspondence between visual scene and the sound, and localize the sound source only by observing sound and visual scene pairs like human? In this paper, we propose a novel unsupervised algorithm to address the problem of localizing the sound source in visual scenes. A two-stream network structure which handles each modality, with attention mechanism is developed for sound source localization. Moreover, although our network is formulated within the unsupervised learning framework, it can be extended to a unified architecture with a simple modification for the supervised and semi-supervised learning settings as well. Meanwhile, a new sound source dataset is developed for performance evaluation. Our empirical evaluation shows that the unsupervised method eventually go through false conclusion in some cases. We also show that even with a few supervision, i.e., semi-supervised setup, false conclusion is able to be corrected effectively. Arda Senocak, Tae-Hyun Oh, Junsik Kim 0001, Ming-Hsuan Yang 0001, In-So Kweon |
CVPR | 3 |
| 2017 | VPGNet: Vanishing Point Guided Network for Lane and Road Marking Detection and RecognitionabstractIn this paper, we propose a unified end-to-end trainable multi-task network that jointly handles lane and road marking detection and recognition that is guided by a vanishing point under adverse weather conditions. We tackle rainy and low illumination conditions, which have not been extensively studied until now due to clear challenges. For example, images taken under rainy days are subject to low illumination, while wet roads cause light reflection and distort the appearance of lane and road markings. At night, color distortion occurs under limited illumination. As a result, no benchmark dataset exists and only a few developed algorithms work under poor weather conditions. To address this shortcoming, we build up a lane and road marking benchmark which consists of about 20,000 images with 17 lane and road marking classes under four different scenarios: no rain, rain, heavy rain, and night. We train and evaluate several versions of the proposed multi-task network and validate the importance of each task. The resulting approach, VPGNet, can detect and classify lanes and road markings, and predict a vanishing point with a single forward pass. Experimental results show that our approach achieves high accuracy and robustness under various conditions in realtime (20 fps). The benchmark and the VPGNet model will be publicly available. Seokju Lee, Junsik Kim 0001, Jae Shin Yoon, Seunghak Shin, Oleksandr Bailo, Namil Kim, Hyun Seok Hong, Seung-Hoon Han, In-So Kweon |
ICCV | 2 |
| 2017 | Pixel-Level Matching for Video Object Segmentation Using Convolutional Neural NetworksabstractWe propose a novel video object segmentation algorithm based on pixel-level matching using Convolutional Neural Networks (CNN). Our network aims to distinguish the target area from the background on the basis of the pixel-level similarity between two object units. The proposed network represents a target object using features from different depth layers in order to take advantage of both the spatial details and the category-level semantic information. Furthermore, we propose a feature compression technique that drastically reduces the memory requirements while maintaining the capability of feature representation. Two-stage training (pretraining and fine-tuning) allows our network to handle any target object regardless of its category (even if the object's type does not belong to the pre-training data) or of variations in its appearance through a video sequence. Experiments on large datasets demonstrate the effectiveness of our model - against related methods - in terms of accuracy, speed, and stability. Finally, we introduce the transferability of our network to different domains, such as the infrared data domain. Jae Shin Yoon, François Rameau, Junsik Kim 0001, Seokju Lee, Seunghak Shin, In-So Kweon |
ICCV | 3 |
| 2016 | Globally Optimal Manhattan Frame Estimation in Real-TimeabstractGiven a set of surface normals, we pose a Manhattan Frame (MF) estimation problem as a consensus set maximization that maximizes the number of inliers over the rotation search space. We solve this problem through a branchand-bound framework, which mathematically guarantees a globally optimal solution. However, the computational time of conventional branch-and-bound algorithms are intractable for real-time performance. In this paper, we propose a novel bound computation method within an efficient measurement domain for MF estimation, i.e., the extended Gaussian image (EGI). By relaxing the original problem, we can compute the bounds in real-time, while preserving global optimality. Furthermore, we quantitatively and qualitatively demonstrate the performance of the proposed method for synthetic and real-world data. We also show the versatility of our approach through two applications: extension to multiple MF estimation and video stabilization. Kyungdon Joo, Tae-Hyun Oh, Junsik Kim 0001, In-So Kweon |
CVPR | 3 |