Jae-Young Sim

dblp:44/5428 · DBLP profile ↗
← Back
56ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0003-2636-7444ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 54 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 14 · 6 since 2021
YearPublicationVenuePosition
2026 Complementary Mixture-of-Experts and Complementary Cross-Attention for Single Image Reflection Separation in the Wild
abstract
Single Image Reflection Separation (SIRS) aims to reconstruct both the transmitted and reflected images from a single image that contains a superimposition of both, captured through a glass-like reflective surface. Recent learning-based methods of SIRS have significantly improved performance on typical images with mild reflection artifacts; however, they often struggle with diverse images containing challenging reflections captured in the wild. In this paper, we propose a universal SIRS framework based on a flexible dual-stream architecture, capable of handling diverse reflection artifacts. Specifically, we incorporate a Mixture-of-Experts mechanism that dynamically assigns specialized experts to image patches based on spatially heterogeneous reflection characteristics. The assigned experts then cooperate to extract complementary features between the transmission and reflection streams in an adaptive manner. In addition, we leverage the multi-head attention mechanism of Transformers to simultaneously exploit both high and low cross-correlations, which are then complementarily used to facilitate adaptive inter-stream feature interactions. Experimental results evaluated on diverse real-world datasets demonstrate that the proposed method significantly outperforms existing state-of-the-art methods qualitatively and quantitatively.
Jae-Young Sim
IEEE Trans. Image Process.2
2025 Knowledge Refinement For Unsupervised Lifelong Person Re-Identification
abstract
Lifelong person re-identification (LReID) is a task that enables models to continuously learn to distinguish person instances while preserving the discriminative ability acquired from previously learned data. Conventional LReID methods have been mainly developed in the supervised manner, however their applicability is limited in real-world scenarios due to the burden of labeling continuously generated large-scale ReID datasets. In this paper, we propose a novel unsupervised LReID method that effectively alleviates the catastrophic forgetting of old knowledge while adapting to new datasets without labeling the training datasets. To this end, we adaptively update the old model according to the knowledge difference between the old and new models. Moreover, we compute the pseudo label reliability of unsupervised clustering, which is then used to leverage the old knowledge to facilitate effective learning of new data. Experimental results demonstrate that the proposed method achieves promising performance of unsupervised LReID even competitive to state-of-the-art supervised LReID methods.
Seungbin Hong, Jae-Young Sim
ICIP2
2025 Dataset Distillation of 3D Point Clouds via Distribution Matching
abstract
Large-scale datasets are usually required to train deep neural networks; however, they increase computational complexity, hindering practical applications. Recently, dataset distillation for images and texts has attracted considerable attention, as it reduces the original dataset to a small synthetic one to alleviate the computational burden of training while preserving essential task-relevant information. However, dataset distillation for 3D point clouds remains largely unexplored, as point clouds exhibit fundamentally different characteristics from those of images, making this task more challenging. In this paper, we propose a distribution-matching-based distillation framework for 3D point clouds that jointly optimizes the geometric structures and orientations of synthetic 3D objects. To address the semantic misalignment caused by the unordered nature of point clouds, we introduce a Semantically Aligned Distribution Matching (SADM) loss, which is computed on the sorted features within each channel. Moreover, to handle rotational variations, we jointly learn optimal rotation angles while updating the synthetic dataset to better align with the original feature distribution. Extensive experiments on widely used benchmark datasets demonstrate that the proposed method consistently outperforms existing dataset distillation approaches, achieving higher accuracy and strong cross-architecture generalization.
Jae-Young Yim, Jae-Young Sim
NeurIPS3
2024 Domain Generalizable Person Search Using Unreal Dataset
abstract
Collecting and labeling real datasets to train the person search networks not only requires a lot of time and effort, but also accompanies privacy issues. The weakly-supervised and unsupervised domain adaptation methods have been proposed to alleviate the labeling burden for target datasets, however, their generalization capability is limited. We introduce a novel person search method based on the domain generalization framework, that uses an automatically labeled unreal dataset only for training but is applicable to arbitrary unseen real datasets. To alleviate the domain gaps when transferring the knowledge from the unreal source dataset to the real target datasets, we estimate the fidelity of person instances which is then used to train the end-to-end network adaptively. Moreover, we devise a domain-invariant feature learning scheme to encourage the network to suppress the domain-related features. Experimental results demonstrate that the proposed method provides the competitive performance to existing person search methods even though it is applicable to arbitrary unseen datasets without any prior knowledge and re-training burdens.
Minyoung Oh, Duhyun Kim, Jae-Young Sim
AAAI3
2024 Fully-Automatic Reflection Removal for 360-Degree Images
abstract
Reflection removal (RR) is a technique to reconstruct the transmitted scene behind the glass from a mixed image taken through glass. In 360-degree images, the mixed image region and the reference image region capturing the reflected scene exist together, and the mixed image is often restored by using the information of reference image. In this paper, we first propose a fully-automatic end-to-end RR framework for 360-degree images which automatically detects the mixed and reference image regions and removes the reflection artifacts in the mixed image by using the reference information simultaneously. We devise a transformer based U-Net architecture with horizontal windowing scheme to capture the long-range dependencies between the mixed and reference images via the self-attention mechanism and suppress the reflection artifacts by using the reference information. We also construct a training dataset of 360-degree images by synthesizing realistic reflection artifacts considering diverse geometric relation and photometric variation between the mixed and reference images. The experimental results show that the proposed method detects the mixed and reference image regions reliably without user-annotation and achieves better performance of RR compared with the state-of-the-art methods.
HyeonA Kim, Eunpil Park, Jae-Young Sim
WACV4
2024 Universal Dehazing via Haze Style Transfer
abstract
Single image dehazing has been actively studied to overcome the quality degradation of hazy images. Most of the existing methods take model-based approaches and the existing learning-based methods usually target specific haze styles only, e.g., daytime, varicolored, and nighttime haze. Therefore, they suffer from the limited performance on arbitrary hazy images with diverse characteristics due to the lack of universal training dataset. In this paper, we first propose a fully data-driven learning-based framework for universal dehazing based on the haze style transfer (HST). We define multiple domains of haze styles by applying the K-means clustering to the background light of diverse real hazy images. We design the haze style modulator to extract the scene radiance features and the haze-related features, respectively. We employ the unpaired image-to-image translation methodology to transfer a source hazy image into different hazy images with diverse styles while preserving the scene radiance. The generated diverse hazy images are used to train the universal dehazing network in a semi-supervised manner, where we implement the dehazing as a special instance of HST into no haze style. The experimental results show that the proposed framework reliably generates realistic and diverse hazy images, and achieves better performance of universal dehazing regardless of the haze styles compared with the existing state-of-the art dehazing methods.
Eunpil Park, Jaejun Yoo 0001, Jae-Young Sim
IEEE Trans. Circuits Syst. Video Technol.3
2022 Zero-Shot Learning for Reflection Removal of Single 360-Degree Image
Byeong-Ju Han, Jae-Young Sim
ECCV (19)2
2022 Zero-Shot Image Dehazing Using Pseudo Atmospheric Light Image
abstract
Abundant training data for deep neural networks improve the performance of single image dehazing (SID) substantially, however they suffer from the domain-shift problem of the discrepancy between the training set and the test set. In this paper, we propose a zero-shot SID method to overcome the domain-shift problem in a self-supervised learning manner. We employ two generator networks to estimate the transmission and the original scene radiance, respectively, from an input hazy image. We also synthesize the pseudo atmospheric light image (PALI) to train the transmission generator to assign zero transmission values to PALI. Since the pseudo-sky patches and the dense-hazy region rarely have the structural textures, the network learns the dense-hazy property from the PALI in a self-supervision learning manner. The experimental results show that the proposed method faithfully restores the scene radiance image, and the PALI loss is effective to train the deep neural network.
Eunsung Jo, Eunpil Park, Jae-Young Sim
MMSP3
2021 Context-Aware Unsupervised Clustering for Person Search
Byeong-Ju Han, Kuhyeun Ko, Jae-Young Sim
BMVC3
2021 End-to-End Trainable Trident Person Search Network Using Adaptive Gradient Propagation
abstract
Person search suffers from the conflicting objectives of commonness and uniqueness between the person detection and re-identification tasks that make the end-to-end training of person search networks difficult. In this paper, we propose a trident network for person search that performs detection, re-identification, and part classification together. We also devise a novel end-to-end training method using adaptive gradient weighting that controls the flow of backpropagated gradients through the re-identification and part classification networks according to the quality of the person detection. The proposed method not only prevents the over-fitting but encourages to exploit fine-grained features by incorporating the part classification branch into the person search framework. Experimental results on the CUHK-SYSU and PRW datasets demonstrate that the proposed method achieves the best performance among the state-of-the-art end-to-end person search methods.
Byeong-Ju Han, Kuhyeun Ko, Jae-Young Sim
ICCV3
2021 Virtual Point Removal for Large-Scale 3D Point Clouds with Multiple Glass Planes
abstract
Large-scale 3D point clouds (LS3DPCs) captured by terrestrial LiDAR scanners often include virtual points which are generated by glass reflection. The virtual points may degrade the performance of various computer vision techniques when applied to LS3DPCs. In this paper, we propose a virtual point removal algorithm for LS3DPCs with multiple glass planes. We first estimate multiple glass regions by modeling the reliability with respect to each glass plane, respectively, such that the regions are assigned high reliability when they have multiple echo pulses for each emitted laser pulse. Then we detect each point whether it is a virtual point or not. For a given point, we recursively traverse all the possible trajectories of reflection, and select the optimal trajectory which provides a point with a similar geometric feature to a given point at the symmetric location. We evaluate the performance of the proposed algorithm on various LS3DPC models with diverse numbers of glass planes. Experimental results show that the proposed algorithm estimates multiple glass regions faithfully and detects the virtual points successfully. Moreover, we also show that the proposed algorithm yields a much better performance of reflection artifact removal compared with the existing method qualitatively and quantitatively.
Jae-Seong Yun, Jae-Young Sim
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Warping Residual Based Image Stitching for Large Parallax
abstract
Image stitching techniques align two images captured at different viewing positions onto a single wider image. When the captured 3D scene is not planar and the camera baseline is large, two images exhibit parallax where the relative positions of scene structures are quite different from each view. The existing image stitching methods often fail to work on the images with large parallax. In this paper, we propose an image stitching algorithm robust to large parallax based on the novel concept of warping residuals. We first estimate multiple homographies and find their inlier feature matches between two images. Then we evaluate warping residual for each feature match with respect to the multiple homographies. To alleviate the parallax artifacts, we partition input images into superpixels and warp each superpixel adaptively according to an optimal homography which is computed by minimizing the error of feature matches weighted by the warping residuals. Experimental results demonstrate that the proposed algorithm provides accurate stitching results for images with large parallax, and outperforms the existing methods qualitatively and quantitatively.
Kyu-Yul Lee, Jae-Young Sim
CVPR2
2019 Cloud Removal of Satellite Images Using Convolutional Neural Network With Reliable Cloudy Image Synthesis Model
abstract
Cloudy pixels in satellite images degrade the visibility of captured surface structure. We propose a novel cloudy image synthesis model and develop a cloud removal algorithm using convolutional neural network. We extract the cloud masks from real cloudy satellite images and from real sky images with clouds. Then we investigate the characteristics of real cloudy images and devise a reliable cloudy image synthesis model which considers the background surface color, misalignement of channel images, and blur in clouds. We train a hierarchical cloud removal network using the synthetic cloudy images. Experimental results demonstrate that the proposed algorithm removes the clouds from cloudy satellite images faithfully and outperforms the existing methods.
Kyu-Yul Lee, Jae-Young Sim
ICIP2
2019 Cluster-Wise Removal of Reflection Artifacts in Large-Scale 3d Point Clouds Using Superpixel-Based Glass Region Estimation
abstract
Large-scale 3D point cloud (LS3DPC) models captured by LiDAR scanners suffer from reflection artifacts caused by glass in real-world scenes. In this paper, we propose a reflection artifact removal algorithm for LS3DPC models using a panoramic image. We first partition the panoramic color image corresponding to an input LS3DPC model into superpixels, and cluster the points projected to each superpixel according to their geometric positions. We determine a superpixel with multiple point clusters as a glass region. We also detect and remove the virtual clusters of reflection artifacts by evaluating the symmetry and geometry scores. The experiment results show that the proposed algorithm detects the glass regions more faithfully and removes the virtual objects more completely compared with the existing method.
Jae-Seong Yun, Jae-Young Sim
ICIP2
2018 Reflection Removal for Large-Scale 3D Point Clouds
abstract
Large-scale 3D point clouds (LS3DPCs) captured by terrestrial LiDAR scanners often exhibit reflection artifacts by glasses, which degrade the performance of related computer vision techniques. In this paper, we propose an efficient reflection removal algorithm for LS3DPCs. We first partition the unit sphere into local surface patches which are then classified into the ordinary patches and the glass patches according to the number of echo pulses from emitted laser pulses. Then we estimate the glass region of dominant reflection artifacts by measuring the reliability. We also detect and remove the virtual points using the conditions of the reflection symmetry and the geometric similarity. We test the performance of the proposed algorithm on LS3DPCs capturing real-world outdoor scenes, and show that the proposed algorithm estimates valid glass regions faithfully and removes the virtual points caused by reflection artifacts successfully.
Jae-Seong Yun, Jae-Young Sim
CVPR2
2018 Glass Reflection Removal Using Co-Saliency-Based Image Alignment and Low-Rank Matrix Completion in Gradient Domain
abstract
The images taken through glass often capture a target transmitted scene as well as undesired reflected scenes. In this paper, we propose a novel reflection removal algorithm using multiple glass images taken from slightly different camera positions. We first find co-saliency maps for input multiple glass images based on the center prior assumption, and then align multiple images reliably with respect to the transmitted scene by selecting feature points with high co-saliency values. The gradients of the transmission images are consistent while the gradients of the reflection images are varying across the aligned multiple glass images. Based on this observation, we compute gradient reliability such that the pixels belonging to consistent salient edges of the transmission image are assigned high reliability values. We restore the gradients of the transmission images and suppress the gradients of the reflection images by formulating a low-rank matrix completion problem in gradient domain. Finally, we reconstruct desired transmission images from the restored transmission gradients. Experimental results show that the proposed algorithm removes the reflection artifacts from glass images faithfully and outperforms the existing methods on challenging glass images with diverse characteristics.
Byeong-Ju Han, Jae-Young Sim
IEEE Trans. Image Process.2
2017 Reflection Removal Using Low-Rank Matrix Completion
abstract
The images taken through glass often capture a target transmitted scene as well as undesired reflected scenes. In this paper, we propose a low-rank matrix completion algorithm to remove reflection artifacts automatically from multiple glass images taken at slightly different camera locations. We assume that the transmitted scenes are more dominant than the reflected scenes in typical glass images. We first warp the multiple glass images to a reference image, where the gradients are consistent in the transmission images while the gradients are varying across the reflection images. Based on this observation, we compute a gradient reliability such that the pixels belonging to the salient edges of the transmission image are assigned high reliability. Then we suppress the gradients of the reflection images and recover the gradients of the transmission images only, by solving a low-rank matrix completion problem in gradient domain. We reconstruct an original transmission image using the resulting optimal gradient map. Experimental results show that the proposed algorithm removes the reflection artifacts from glass images faithfully and outperforms the existing algorithms on typical glass images.
Byeong-Ju Han, Jae-Young Sim
CVPR2
2017 Saliency detection for panoramic landscape images of outdoor scenes
Byeong-Ju Han, Jae-Young Sim
J. Vis. Commun. Image Represent.2
2017 Saliency Detection for 3D Surface Geometry Using Semi-regular Meshes
abstract
In this paper, a unified detection algorithm of viewindependent and view-dependent saliency for three-dimensional mesh models is proposed. While the conventional techniques use the irregular meshes, we adopt the semi-regular meshes to overcome the drawback of irregular connectivity for saliency computation. We employ the angular deviation of normal vectors between neighboring faces as geometric curvature features, which are evaluated at hierarchically structured triangle faces. We construct a fully connected graph at each level of semi-regular mesh, where the face patches serve as graph nodes. At the base mesh level, we estimate the saliency as the stationary distribution of random walk. At the higher level meshes, we take the maximum value between the stationary distribution of random walk at the current level and an upsampled saliency map from the previous coarser scale. Moreover, we also propose a view-dependent saliency detection method that employs the visibility feature in addition to the geometric features to estimate the saliency with respect to a selected viewpoint. Experimental results demonstrate that the proposed saliency detection algorithm captures global conspicuous regions reliably and detects locally detailed geometric features faithfully, compared with the conventional techniques.
Se-Won Jeong, Jae-Young Sim
IEEE Trans. Multim.2
2016 Supervoxel-based saliency detection for large-scale colored 3D point clouds
abstract
Large-scale 3D point clouds have been actively used in many applications with the advent of capturing devices. In this paper, we propose a novel saliency detection algorithm for large-scale colored 3D point clouds which capture real-world scenes. We first voxelize an input point cloud, and then partition voxels into a supervoxel which corresponds to a clusters at the lowest level. We construct the supervoxel cluster hierarchy iteratively, where a high level cluster includes low level clusters which exhibit similar features to each other. We also estimate the saliency at each cluster by computing the distinctness of geometric and color features based on center-surround contrast. By averaging the multiscale saliency maps obtained at different levels of clusters, we obtain final saliency distribution. Experimental results demonstrate that the proposed algorithm extracts globally and locally salient regions from large-scale colored 3D point clouds faithfully by employing the geometric and photometric features together.
Jae-Seong Yun, Jae-Young Sim
ICIP2
2015 Multiple random walkers and their application to image cosegmentation
abstract
A graph-based system to simulate the movements and interactions of multiple random walkers (MRW) is proposed in this work. In the MRW system, multiple agents traverse a single graph simultaneously. To achieve desired interactions among those agents, a restart rule can be designed, which determines the restart distribution of each agent according to the probability distributions of all agents. In particular, we develop the repulsive rule for data clustering. We illustrate that the MRW clustering can segment real images reliably. Furthermore, we propose a novel image cosegmentation algorithm based on the MRW clustering. Specifically, the proposed algorithm consists of two steps: inter-image concurrence computation and intra-image MRW clustering. Experimental results demonstrate that the proposed algorithm provides promising cosegmentation performance.
Chulwoo Lee, Won-Dong Jang, Jae-Young Sim, Chang-Su Kim 0001
CVPR3
2015 Multihypothesis trajectory analysis for robust visual tracking
abstract
The notion of multihypothesis trajectory analysis (MTA) for robust visual tracking is proposed in this work. We employ multiple component trackers using texture, color, and illumination invariant features, respectively. Each component tracker traces a target object forwardly and then backwardly over a time interval. By analyzing the pair of the forward and backward trajectories, we measure the robustness of the component tracker. To this end, we extract the geometry similarity, the cyclic weight, and the appearance similarity from the forward and backward trajectories. We select the optimal component tracker to yield the maximum robustness score, and use its forward trajectory as the final tracking result. Experimental results show that the proposed MTA tracker improves the robustness and the accuracy of tracking, outperforming the state-of-the-art trackers on a recent benchmark dataset.
Dae-Youn Lee, Jae-Young Sim, Chang-Su Kim 0001
CVPR2
2015 SOWP: Spatially Ordered and Weighted Patch Descriptor for Visual Tracking
abstract
A simple yet effective object descriptor for visual tracking is proposed in this paper. We first decompose the bounding box of a target object into multiple patches, which are described by color and gradient histograms. Then, we concatenate the features of the spatially ordered patches to represent the object appearance. Moreover, to alleviate the impacts of background information possibly included in the bounding box, we determine patch weights using random walk with restart (RWR) simulations. The patch weights represent the importance of each patch in the description of foreground information, and are used to construct an object descriptor, called spatially ordered and weighted patch (SOWP) descriptor. We incorporate the proposed SOWP descriptor into the structured output tracking framework. Experimental results demonstrate that the proposed algorithm yields significantly better performance than the state-of-the-art trackers on a recent benchmark dataset, and also excels in another recent benchmark dataset.
Hanul Kim 0001, Dae-Youn Lee, Jae-Young Sim, Chang-Su Kim 0001
ICCV3
2015 Robust video stitching using adaptive pixel transfer
abstract
Video stitching is a challenging problem when multiple videos are captured with severely different camera parameters. In this paper, we propose a robust video stitching algorithm which transfers the pixels in the background plane and the foreground objects separately. We first compute the fundamental matrix and the homography on a common ground plane between two videos using the activity based features. The ground plane pixels can be warped by using the homography. For each foreground object, we project the off-plane pixels onto the ground plane by analyzing the dominant direction of object region. The projected ground plane pixel is transferred to another image via homography, and then elevated to find the true corresponding pixel using the epipolar constraint. We also locally adjust the correspondence matching positions based on energy minimization framework. Experimental results demonstrate that the proposed algorithm correctly warps both of the ground plane and the foreground objects between video sequences captured with distinct camera configurations, compared with the state-of-the-art stitching algorithms.
Kyu-Yul Lee, Jae-Young Sim
ICIP2
2015 Robust contrast enhancement of noisy low-light images: Denoising-enhancement-completion
abstract
A robust contrast enhancement algorithm for noisy low-light images, called the denoising-enhancement-completion (DEC), is proposed in this work. We observe that noise components in low-light images degrade the performance of the contrast enhancement. Therefore, we first reduce noise components in an input image. Then, we compute the reliability weight for each pixel, by measuring the difference between the input image and the denoised image, and categorize each pixel into one of two classes: noise-free or noisy. We perform the selective histogram equalization to enhance the contrast of the noise-free pixels only. Finally, we restore missing values of the noisy pixels using the enhanced noise-free pixel values, by employing a low-rank matrix completion scheme. Experimental results show that the proposed DEC algorithm removes noise and enhances the contrast of low-light images more effectively than conventional algorithms.
Jaemoon Lim, Jin-Hwan Kim, Jae-Young Sim, Chang-Su Kim 0001
ICIP3
2015 FDQM: Fast Quality Metric for Depth Maps Without View Synthesis
abstract
We propose a fast quality metric for depth maps, called fast depth quality metric (FDQM), which efficiently evaluates the impacts of depth map errors on the qualities of synthesized intermediate views in multiview video plus depth applications. In other words, the proposed FDQM assesses view synthesis distortions in the depth map domain, without performing the actual view synthesis. First, we estimate the distortions at pixel positions, which are specified by reference disparities and distorted disparities, respectively. Then, we integrate those pixel-wise distortions into an FDQM score by employing a spatial pooling scheme, which considers occlusion effects and the characteristics of human visual attention. As a benchmark of depth map quality assessment, we perform a subjective evaluation test for intermediate views, which are synthesized from compressed depth maps at various bitrates. We compare the subjective results with objective metric scores. Experimental results demonstrate that the proposed FDQM yields highly correlated scores to the subjective ones. Moreover, FDQM requires at least 10 times less computations than conventional quality metrics, since it does not perform the actual view synthesis.
Won-Dong Jang, Taeyoung Chung, Jae-Young Sim, Chang-Su Kim 0001
IEEE Trans. Circuits Syst. Video Technol.3
2015 Spatiotemporal Saliency Detection for Video Sequences Based on Random Walk With Restart
abstract
A novel saliency detection algorithm for video sequences based on the random walk with restart (RWR) is proposed in this paper. We adopt RWR to detect spatially and temporally salient regions. More specifically, we first find a temporal saliency distribution using the features of motion distinctiveness, temporal consistency, and abrupt change. Among them, the motion distinctiveness is derived by comparing the motion profiles of image patches. Then, we employ the temporal saliency distribution as a restarting distribution of the random walker. In addition, we design the transition probability matrix for the walker using the spatial features of intensity, color, and compactness. Finally, we estimate the spatiotemporal saliency distribution by finding the steady-state distribution of the walker. The proposed algorithm detects foreground salient objects faithfully, while suppressing cluttered backgrounds effectively, by incorporating the spatial transition matrix and the temporal restarting distribution systematically. Experimental results on various video sequences demonstrate that the proposed algorithm outperforms conventional saliency detection algorithms qualitatively and quantitatively.
Hansang Kim, Youngbae Kim, Jae-Young Sim, Chang-Su Kim 0001
IEEE Trans. Image Process.3
2015 Video Deraining and Desnowing Using Temporal Correlation and Low-Rank Matrix Completion
abstract
A novel algorithm to remove rain or snow streaks from a video sequence using temporal correlation and low-rank matrix completion is proposed in this paper. Based on the observation that rain streaks are too small and move too fast to affect the optical flow estimation between consecutive frames, we obtain an initial rain map by subtracting temporally warped frames from a current frame. Then, we decompose the initial rain map into basis vectors based on the sparse representation, and classify those basis vectors into rain streak ones and outliers with a support vector machine. We then refine the rain map by excluding the outliers. Finally, we remove the detected rain streaks by employing a low-rank matrix completion technique. Furthermore, we extend the proposed algorithm to stereo video deraining. Experimental results demonstrate that the proposed algorithm detects and removes rain or snow streaks efficiently, outperforming conventional algorithms.
Jin-Hwan Kim, Jae-Young Sim, Chang-Su Kim 0001
IEEE Trans. Image Process.2
2014 Visual Tracking Using Pertinent Patch Selection and Masking
abstract
A novel visual tracking algorithm using patch-based appearance models is proposed in this paper. We first divide the bounding box of a target object into multiple patches and then select only pertinent patches, which occur repeatedly near the center of the bounding box, to construct the foreground appearance model. We also divide the input image into non-overlapping blocks, construct a background model at each block location, and integrate these background models for tracking. Using the appearance models, we obtain an accurate foreground probability map. Finally, we estimate the optimal object position by maximizing the likelihood, which is obtained by convolving the foreground probability map with the pertinence mask. Experimental results demonstrate that the proposed algorithm outperforms state-of-the-art tracking algorithms significantly in terms of center position errors and success rates.
Dae-Youn Lee, Jae-Young Sim, Chang-Su Kim 0001
CVPR2
2014 GEQM: A quality metric for gray-level edge maps based on structural matching
abstract
An accurate quality metric, called GEQM, for gray-level edge maps based on the structural matching of edge pixels is proposed in this work. We design the positional matching cost, which reflects the distance between two edge pixels, and the structural matching cost, which measures the structural shapes of edges as well as the differences of edge strength levels. Based on the cost functions, we perform the graph-cut optimization to obtain the optimal pixel-based matching between source and target edge maps bidirectionally. Finally, we compute the GEQM score by summing up the optimal matching costs of all edge pixels. Experimental results show that the proposed GEQM performs the edge map quality assessment more accurately and more reliably than conventional metrics. Especially, GEQM is suitable for assessing the qualities of synthesized intermediate views in multi-view image processing.
Won-Dong Jang, Jae-Young Sim, Chang-Su Kim 0001
ICASSP2
2014 Stereo video deraining and desnowing based on spatiotemporal frame warping
abstract
A novel rain (or snow) streak removal algorithm for stereo video sequences is proposed in this work. We observe that rain streaks appear at different locations in spatiotemporally adjacent frames. Thus, to derain a left-view frame, we synthesize it by warping the spatially adjacent right-view frame and the temporally previous and next frames, respectively. We subtract each warped frame from the original frame, and apply the median filter to the three difference images to obtain a reliable rain mask. Then, we remove rain streaks by replacing each rainy pixel value with a weighted average of non-locally neighboring pixel values. Experimental results demonstrate that the proposed algorithm removes rain streaks reliably and recovers original scene contents faithfully.
Jin-Hwan Kim, Jae-Young Sim, Chang-Su Kim 0001
ICIP2
2014 Robust video stabilization based on mesh grid warping of rolling-free features
abstract
A robust video stabilization algorithm, which reduces shaky camera movements and rolling-shutter distortions in video sequences, is proposed in this work. We first extract feature trajectories through a video sequence, and then transform the feature positions into rolling-free smoothed positions. Then, we set a mesh grid on each frame and warp each grid cell by matching the original features to the smoothed ones. For robust warping, we formulate a cost function based on the confidence of each feature and the reliability of each grid cell. The cost function consists of a data term, a structure-preserving term, and a regularization term. By minimizing the cost function, we find the optimal grid positions in the warped frame and transform each grid cell accordingly. Experimental results show that the proposed algorithm stabilizes videos and removes rolling-shutter distortions more efficiently than conventional algorithms.
Yeong Jun Koh, Jae-Young Sim, Chang-Su Kim 0001
ICIP2
2014 Video saliency detection based on spatiotemporal feature learning
abstract
A video saliency detection algorithm based on feature learning, called ROCT, is proposed in this work. To detect salient regions, we design multiple spatiotemporal features and combine those features using a support vector machine (SVM). We extract the spatial features of rarity, compactness, and center prior by analyzing the color distribution in each image frame. Also, we obtain the temporal features of motion intensity and motion contrast to identify visually important motions. We train an SVM classifier using the spatiotemporal features extracted from training video sequences. Finally, we compute the visual saliency of each patch in an input sequence using the trained classifier. Experimental results demonstrate that the proposed algorithm provides more accurate and reliable results of saliency detection than conventional algorithms.
Se-Ho Lee, Jin-Hwan Kim, Kwangpyo Choi, Jae-Young Sim, Chang-Su Kim 0001
ICIP4
2014 Depth-guided adaptive contrast enhancement using 2D histograms
abstract
A novel contrast enhancement (CE) algorithm using 2-dimensional (2D) histograms, which transforms pixel values adaptively based on the depth information, is proposed in this work. In general, foreground objects convey more important visual information than background regions. Hence we assign high CE priorities to foreground pixels using the depth values and generate a depth-guided 2D histogram. Then, we stretch the gray-level differences of adjacent foreground pixels more strongly than those of adjacent background pixels. Moreover, to enhance background regions as well, we design two transformation functions for the foreground and the background separately. By combining the two functions according to pixel depths, we obtain an adaptive space-variant transformation function, which is finally used to reconstruct the output image. Experimental results show that the proposed algorithm outperforms conventional CE algorithms by enhancing salient foreground objects efficiently and preserving background details faithfully.
Juntae Lee, Chulwoo Lee, Jae-Young Sim, Chang-Su Kim 0001
ICIP3
2014 Automatic Video Genre Classification Using Multiple SVM Votes
abstract
A video genre classification algorithm based on the voting from multiple SVMs is proposed in this work. While conventional genre classifiers use generic baseline features, we employ more specialized features to describe five video genres: animation, commercial, entertainment, drama, and sports. We also present a robust classification algorithm using multiple SVMs, which consider all possible binary grouping of the five genres. Given a query video, each SVM casts a probabilistic vote for each genre. Then, the optimal genre with the maximum votes is selected. Experimental results show that the proposed algorithm provides more accurate classification performance than conventional algorithms.
Won-Dong Jang, Chulwoo Lee, Jae-Young Sim, Chang-Su Kim 0001
ICPR3
2014 Multiscale Saliency Detection Using Random Walk With Restart
abstract
In this paper, we propose a graph-based multiscale saliency-detection algorithm by modeling eye movements as a random walk on a graph. The proposed algorithm first extracts intensity, color, and compactness features from an input image. It then constructs a fully connected graph by employing image blocks as the nodes. It assigns a high edge weight if the two connected nodes have dissimilar intensity and color features and if the ending node is more compact than the starting node. Then, the proposed algorithm computes the stationary distribution of the Markov chain on the graph as the saliency map. However, the performance of the saliency detection depends on the relative block size in an image. To provide a more reliable saliency map, we develop a coarse-to-fine refinement technique for multiscale saliency maps based on the random walk with restart (RWR). Specifically, we use the saliency map at a coarse scale as the restarting distribution of RWR at a fine scale. Experimental results demonstrate that the proposed algorithm detects visual saliency precisely and reliably. Moreover, the proposed algorithm can be efficiently used in the applications of proto-object extraction and image retargeting.
Jae-Young Sim, Chang-Su Kim 0001
IEEE Trans. Circuits Syst. Video Technol.2
2014 Bit Allocation Algorithm With Novel View Synthesis Distortion Model for Multiview Video Plus Depth Coding
abstract
An efficient bit allocation algorithm based on a novel view synthesis distortion model is proposed for the rate-distortion optimized coding of multiview video plus depth sequences in this paper. We decompose an input frame into nonedge blocks and edge blocks. For each nonedge block, we linearly approximate its texture and disparity values, and derive a view synthesis distortion model, which quantifies the impacts of the texture and depth distortions on the qualities of synthesized virtual views. On the other hand, for each edge block, we use its texture and disparity gradients for the distortion model. In addition, we formulate a bit-rate allocation problem in terms of the quantization parameters for texture and depth data. By solving the problem, we can optimally divide a limited bit budget between the texture and depth data, in order to maximize the qualities of synthesized virtual views, as well as those of encoded real views. Experimental results demonstrate that the proposed algorithm yields the average PSNR gains of 1.98 and 2.04 dB in two-view and three-view scenarios, respectively, as compared with a benchmark conventional algorithm.
Taeyoung Chung, Jae-Young Sim, Chang-Su Kim 0001
IEEE Trans. Image Process.2
2013 Robust stereo matching under radiometric variations based on cumulative distributions of gradients
abstract
We propose a robust stereo matching algorithm for images captured under varying radiometric conditions, such as exposure and lighting variations, based on the cumulative distributions of gradients. The gradient operator extracts local changes in pixel values, which are less sensitive to radiometric variations than the original pixel values. Moreover, the cumulative distribution function (CDF) of gradient vectors reflects the ranks of edge strength levels, and corresponding pixels in stereo images tend to have similar ranks regardless of radio-metric conditions. Therefore, we design the matching cost function based on the dissimilarity of gradient CDF values. However, since multiple pixels in an image may have the same gradient CDF value, we further constrain the correspondence matching by checking the dissimilarity of gradient orientations. Finally, to estimate an accurate disparity at each pixel, we adaptively aggregate matching costs using the color similarity and the geometric proximity of neighboring pixels. Experimental results demonstrate that the proposed algorithm provides more accurate disparities than conventional algorithms, especially under varying lighting conditions.
Il-Lyong Jung, Jae-Young Sim, Chang-Su Kim 0001, Sang Uk Lee
ICIP2
2013 Video saliency detection based on random walk with restart
abstract
A graph-based video saliency detection algorithm is proposed in this work. We model eye movements on an image plane as random walks on a graph. To detect the saliency of the first frame in a video sequence, we construct a fully connected graph, in which each node represents an image block. We assign an edge weight to be proportional to the dissimilarity between the incident nodes and inversely proportional to their geometrical distance. We extract the saliency level of each node from the stationary distribution of the random walker on the graph. Next, to detect the saliency of each subsequent frame, we add the criterion that an edge, connecting a slow motion node to a fast motion node, should have a large weight. We then compute the stationary distribution of the random walk with restart (RWR) simulation, in which the saliency of the previous frame is used as the restarting distribution. Experimental results show that the proposed algorithm provides more reliable and accurate saliency detection performance than conventional algorithms.
Hansang Kim, Jae-Young Sim, Chang-Su Kim 0001, Sang Uk Lee
ICIP3
2013 Single-image deraining using an adaptive nonlocal means filter
abstract
An adaptive rain streak removal algorithm for a single image is proposed in this work. We observe that a typical rain streak has an elongated elliptical shape with a vertical orientation. Thus, we first detect rain streak regions by analyzing the rotation angle and the aspect ratio of the elliptical kernel at each pixel location. We then perform the nonlocal means filtering on the detected rain streak regions by selecting nonlocal neighbor pixels and their weights adaptively. Experimental results demonstrate that the proposed algorithm removes rain streaks more efficiently and provides higher restored image qualities than conventional algorithms.
Jin-Hwan Kim, Chul Lee, Jae-Young Sim, Chang-Su Kim 0001
ICIP3
2013 Fast object tracking using color histograms and patch differences
abstract
A fast visual object tracking algorithm using novel object appearance models is proposed in this work. We develop a color histogram model and a patch difference model to extract color and texture feature vectors, respectively. Then, we apply k-nearest neighbor classifiers to the color and texture feature vectors and obtain the foreground probability map. We then perform a hierarchical mean shift process on the map to identify the object window. Experimental results demonstrate that proposed algorithm outperforms the conventional algorithms in terms of both tracking accuracy and processing speed.
Dae-Youn Lee, Jae-Young Sim, Chang-Su Kim 0001
ICIP2
2013 Reliable optical flow estimation in motion-blurred regions
abstract
A robust optical flow estimation algorithm for motion-blurred regions is proposed in this work. We first obtain initial optical flow vectors. Then, we detect motion-blurred regions that yield low contrast, low saturation, and inconsistent optical flow vectors. We replace the optical flow vectors at motion-blurred pixels with reliable vectors at nearby unblurred pixels. To this end, we develop an energy minimization framework. Simulation results demonstrate that the proposed algorithm refines optical flow vectors in motion-blurred regions accurately and provides better performance than conventional algorithms.
Yeong Jun Koh, Chul Lee, Jae-Young Sim, Chang-Su Kim 0001
MMSP3
2013 Optimized contrast enhancement for real-time image and video dehazing
Jin-Hwan Kim, Won-Dong Jang, Jae-Young Sim, Chang-Su Kim 0001
J. Vis. Commun. Image Represent.3
2013 Consistent Stereo Matching Under Varying Radiometric Conditions
abstract
A consistent stereo matching (CSM) algorithm under varying radiometric conditions, such as lighting and exposure variations, for intermediate view synthesis is proposed in this work. First, we transform the colors of stereo images adaptively so that they are similar at corresponding pixels. Since the correspondences are generally unknown before stereo matching, we estimate pseudo-disparity vectors by sorting pixels based on the cumulative color histograms and use those pseudo vectors in the color transform. Then, to improve the accuracy of stereo matching, we jointly estimate the disparity maps for virtual intermediate views as well as those for real views, based on the consistency criterion that an object point should have the same disparity through all the views. Specifically, we compute matching costs using the reliability term and aggregate the costs to obtain initial disparity maps. We then refine the initial disparity maps by minimizing an energy function, which includes the consistency term. Experimental results show that the proposed CSM algorithm significantly reduces the error rate of disparity estimation under different radiometric conditions and synthesizes high quality intermediate views.
Il-Lyong Jung, Taeyoung Chung, Jae-Young Sim, Chang-Su Kim 0001
IEEE Trans. Multim.3
2013 Correspondence Matching of Multi-View Video Sequences Using Mutual Information Based Similarity Measure
abstract
We propose a correspondence matching algorithm for multi-view video sequences, which provides reliable performance even when the multiple cameras have significantly different parameters, such as viewing angles and positions. We use an activity vector, which represents the temporal occurrence pattern of moving foreground objects at a pixel position, as an invariant feature for correspondence matching. We first devise a novel similarity measure between activity vectors by considering the joint and individual behavior of the activity vectors. Specifically, we define random variables associated with the activity vectors and measure their similarity using the mutual information between the random variables. Moreover, to find a reliable homography transform between views, we find consistent pixel positions by employing the iterative bidirectional matching. We also refine the matching results of multiple source pixel positions by minimizing a matching cost function based on the Markov random field. Experimental results show that the proposed algorithm provides more accurate and reliable matching performance than the conventional activity-based and feature-based matching algorithms, and therefore can facilitate various applications of visual sensor networks.
Soon-Young Lee, Jae-Young Sim, Chang-Su Kim 0001, Sang Uk Lee
IEEE Trans. Multim.2
2012 Temporally x real-time video dehazing
abstract
A real-time video dehazing algorithm, which reduces flickering artifacts and yields high quality output videos, is proposed in this work. Assuming that a scene point yields highly correlated transmission values between adjacent image frames, we develop the temporal coherence cost. Then, we add the temporal coherence cost to the contrast cost and the truncation loss cost to define the overall cost function. By minimizing the overall cost function, we obtain the optimal transmission. Moreover, to reduce the computational complexity and facilitate real-time applications, we approximate the conventional edge preserving filter by the overlapped block filter. Experimental results demonstrate that the proposed algorithm is sufficiently fast for real-time applications and effectively removes haze and flickering artifacts.
Jin-Hwan Kim, Won-Dong Jang, Yongsup Park, Dong-Hahk Lee, Jae-Young Sim, Chang-Su Kim 0001
ICIP5
2012 Histogram-Based stereo matching under varying illumination conditions
abstract
A histogram-based matching algorithm for stereo images captured under different illumination conditions is proposed in this work. The cumulative histogram of an image represents the ranks of relative pixel brightness, which are robust to illumination changes. Therefore, we design the matching cost based on the similarity of the cumulative histograms of stereo images. As an optional mode, the proposed algorithm can evaluate the histograms for foreground objects and the background separately to alleviate occlusion artifacts. To determine the disparity of each pixel, the proposed algorithm adaptively aggregates matching costs based on the color similarity and the geometric proximity of neighboring pixels. Then, it refines false disparities at occluded pixels using more reliable disparities of non-occluded pixels. Experimental results demonstrate that the proposed algorithm provides higher quality disparity maps than the conventional methods under varying illumination conditions.
Il-Lyong Jung, Jae-Young Sim, Chang-Su Kim 0001
VCIP2
2012 Time-of-flight sensor and color camera calibration for multi-view acquisition
Hyunjung Shim, Rolf Adelsberger, James Dokyoon Kim, Seon-Min Rhee, Taehyun Rhee, Jae-Young Sim, Markus Gross 0001, Chang-Yeong Kim
Vis. Comput.6
2011 Single image dehazing based on contrast enhancement
abstract
A simple and adaptive single image dehazing algorithm is proposed in this work. Based on the observation that a hazy image has low contrast in general, we attempt to restore the original image by enhancing the contrast. First, the proposed algorithm estimates the airlight in a given hazy image based on the quad-tree subdivision. Then, the proposed algorithm estimates the transmission map to maximize the contrast of the output image. To measure the contrast, we develop a cost function, which consists of a standard deviation term and a histogram uniformness term. Experimental results demonstrate that the proposed algorithm can remove haze efficiently and reconstruct fine details in original scenes clearly.
Jin-Hwan Kim, Jae-Young Sim, Chang-Su Kim 0001
ICASSP2
2011 Digital Hologram Compression Using Correlation of Reconstructed Object Images
Jae-Young Sim
PSIVT (2)1
2010 Panoramic scene generation from multi-view images with close foreground objects
abstract
An algorithm to generate a panorama from multi-view images, which contain foreground objects with varying depths, is proposed in this work. The proposed algorithm constructs a foreground panorama and a background panorama separately, and then merges them into a complete panorama. First, the foreground panorama is obtained by finding the translational displacements of objects between source images. Second, the background panorama is initialized using warped source images and then optimized to preserve spatial consistency and satisfy visual constraints. Then, the background panorama is extended by inserting seams and merged with the foreground panorama. Experimental results demonstrate that the proposed algorithm provides visually satisfying panoramas with all meaningful foreground objects, but without severe artifacts in the backgrounds.
Soon-Young Lee, Jae-Young Sim, Chang-Su Kim 0001, Sang Uk Lee
PCS2
2008 Compression of 3-D Point Visual Data Using Vector Quantization and Rate-Distortion Optimization
abstract
In this paper, we propose adaptive and flexible quantization and compression algorithms for 3-D point data using vector quantization (VQ) and rate-distortion (R-D) optimization. The point data are composed of the position and the radius of sphere based on QSplat representation. The positions of child spheres are first transformed to the local coordinate system, which is determined by the parent-children relationship. The local coordinate transform makes the positions more compactly distributed in 3-D space, facilitating an effective application of VQ. We also develop a constrained encoding method for the radius data, which can provide a hole-free surface rendering at the decoder side. Furthermore, R-D optimized compression algorithm is proposed in order to allocate an optimal bitrate to each sphere. Experimental results show that the proposed algorithm can effectively compress the original 3-D point geometry at various bitrates.
Jae-Young Sim, Sang Uk Lee
IEEE Trans. Multim.1
2005 Rate-distortion optimized compression and view-dependent transmission of 3-D normal meshes
abstract
A unified approach to rate-distortion (R-D) optimized compression and view-dependent transmission of three-dimensional (3-D) normal meshes is investigated in this work. A normal mesh is partitioned into several segments, which are then encoded independently. The bitstream of each segment is truncated optimally using a geometry distortion model based on the subdivision hierarchy. It is shown that the proposed compression algorithm yields a higher coding gain than the conventional algorithm. Moreover, to facilitate interactive transmission of 3-D data according to a client's viewing position, the server can allocate an adaptive bitrate to each segment based on its visibility priority. Simulation results demonstrate that the view-dependent transmission technique can reduce the bandwidth requirement considerably, while maintaining a good visual quality.
Jae-Young Sim, Chang-Su Kim 0001, C.-C. Jay Kuo, Sang Uk Lee
IEEE Trans. Circuits Syst. Video Technol.1
2005 Lossless compression of 3-D point data in QSplat representation
abstract
We propose a lossless compression algorithm for three-dimensional point data in graphics applications. In typical point representation, each point is treated as a sphere and its geometrical and normal data are stored in the hierarchical structure of bounding spheres. The proposed algorithm sorts child spheres according to their positions to achieve a higher coding gain for geometrical data. Also, the proposed algorithm compactly encodes normal data by exploiting high correlation between parent and child normals. Simulation results show that the proposed algorithm saves up to 60% of storage space.
Jae-Young Sim, Chang-Su Kim 0001, Sang Uk Lee
IEEE Trans. Multim.1
2004 Lossless compression of point-based data for 3D graphics rendering
abstract
A lossless compression algorithm of 3D point data is proposed in this work. QSplat is one of the efficient rendering methods for 3D point data. In QSplat, each point is assigned a sphere, and the geometry and normal data are stored in the hierarchical structure of bounding spheres. To compress QSplat data, child spheres are sorted based on their limit radii to constrain the indices for the geometry data. Then, the radii and the positions of spheres are encoded separately using the reduced index sets. Also, each normal is encoded using the parent normal context, and the normal indices are reduced by the normal cone information. Simulation results show that the proposed algorithm achieves a high compression ratio by combining the reduced index sets with the context-based entropy coding.
Jae-Young Sim, Chang-Su Kim 0001, Sang Uk Lee
VCIP1
2003 An efficient 3D mesh compression technique based on triangle fan structure
Jae-Young Sim, Chang-Su Kim 0001, Sang Uk Lee
Signal Process. Image Commun.1