EDBT 2026 Demo / reviewers in the wild / expert
Changjae Oh
dblp:128/7885
· DBLP profile ↗
41ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0002-6522-2451ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 23 · 6 first-author · 15 since 2021Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Detector Teaches Itself: Lightweight Self-supervised Adaptation for Open-Vocabulary Object Detection
Yazhe Wan, Changjae Oh |
ICPR (7) | 2 |
| 2026 | 3D semantic image synthesis with geometric and semantic consistency
Jihyun Kim 0009, Changjae Oh, Hoseok Do, Sunghwan Choi, Kwanghoon Sohn |
Expert Syst. Appl. | 2 |
| 2026 | High-Resolution Open-Vocabulary Object 6D Pose EstimationabstractThe generalisation to unseen objects in the 6D pose estimation task is very challenging. While Vision-Language Models (VLMs) enable using natural language descriptions to support 6D pose estimation of unseen objects, these solutions underperform compared to model-based methods. In this work we present Horyon, an open-vocabulary VLM-based architecture that addresses relative pose estimation between two scenes of an unseen object, described by a textual prompt only. We use the textual prompt to identify the unseen object in the scenes and then obtain high-resolution multi-scale features. These features are used to extract cross-scene matches for registration. We evaluate our model on a benchmark with a large variety of unseen objects across four datasets, namely REAL275, Toyota-Light, Linemod, and YCB-Video. Our method achieves state-of-the-art performance on all datasets, outperforming by 12.6 in Average Recall the previous best-performing approach. Jaime Corsetti, Davide Boscaini, Francesco Giuliari, Changjae Oh, Andrea Cavallaro, Fabio Poiesi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | A Deployment-Oriented Simulation Framework for Deep Learning-Based Lane Change PredictionabstractAdvanced driving simulations are increasingly used in automated driving research, yet freely available data and tools remain limited. We present a new open source framework for synthetic data generation for lane change (LC) intention recognition in highways. Built on the CARLA simulator, it advances the state-of-the-art by providing a 50-driver dataset, a large-scale 3D map, and code for reproducibility and new data creation. The 60 km highway map includes varying curvature radii and straight segments. The codebase supports simulation enhancements (traffic management, vehicle cockpit, engine noise) and Machine Learning (ML) model training and evaluation, including CARLA log post-processing into time series. The dataset contains over 3,400 annotated LC maneuvers with synchronized ego dynamics, road geometry, and traffic context. From an automotive industry perspective, we also assess leading-edge ML models on STM32 microcontrollers using deployability metrics. Unlike prior infrastructure based works, we estimate time-to-LC from ego-centric data. Results show that a Transformer model yields the lowest regression error, while XGBoost offers the best trade-offs on extremely resource-constrained devices. The entire framework is publicly released to support advancement in automated driving research. Luca Forneris, Riccardo Berta, Matteo Fresta, Luca Lazzaroni, Hadise Rojhan, Changjae Oh, Alessandro Pighetti, Hadi Ballout, Fabio Tango, Francesco Bellotti |
IEEE Signal Process. Lett. | 6 |
| 2025 | Coming Out of the Dark: Human Pose Estimation in Low-light ConditionsabstractHuman pose estimation in low-light conditions is vital for applications such as surveillance and autonomous systems, yet the severe visual distortions hinder both manual annotation and estimation precision. Existing approaches typically rely on additional reference information to mitigate these issues, however, customized data collection equipment poses limitations on their scalability. To alleviate the issue, we construct a Low-Light Images and Poses (LLIP) dataset, which includes only paired low-light images and pose annotations obtained using off-the-shelf motion capture devices. Furthermore, we propose a Multi-grained High-frequency Feature Consistency Learning framework (MHFCL), which does not rely on additional reference information. MHFCL employs a Retinex-inspired restoration stream to recover high-frequency details and integrates them into pose estimation using a multi-grained consistency mechanism. Experiments demonstrate that our approach achieves a new benchmark in low-light pose estimation, while maintaining competitive performance in well-lit conditions. Yong Su 0003, Meng Xing, Changjae Oh, Xuewei Liu, Jieyang Li |
IJCAI | 4 |
| 2025 | Improving Generalization of Language-Conditioned Robot ManipulationabstractThe control of robots for manipulation tasks generally relies on visual input. Recent advances in vision-language models (VLMs) enable the use of natural language instructions to condition visual input and control robots in a wider range of environments. However, existing methods require a large amount of data to fine-tune VLMs for operating in unseen environments. In this paper, we present a framework that learns object-arrangement tasks from just a few demonstrations. We propose a two-stage framework that divides object-arrangement tasks into a target localization stage, for picking the object, and a region determination stage for placing the object. We present an instance-level semantic fusion module that aligns the instance-level image crops with the text embedding, enabling the model to identify the target objects defined by the natural language instructions. We validate our method on both simulation and real-world robotic environments. Our method, fine-tuned with a few demonstrations, improves generalization capability and demonstrates zero-shot ability in real-robot manipulation scenarios. Chenglin Cui, Chaoran Zhu, Changjae Oh, Andrea Cavallaro |
IROS | 3 |
| 2025 | Robust Photo-Realistic Hand Gesture Generation: from Single View to Multiple ViewabstractHigh-fidelity hand gesture generation represents a significant challenge in human-centric generation tasks. Existing methods typically employ a single-view mesh-rendered image prior to enhancing gesture generation quality. However, the spatial complexity of hand gestures and the inherent limitations of single-view rendering make it difficult to capture complete gesture information, particularly when fingers are occluded. The fundamental contradiction lies in the loss of 3D topological relationships through 2D projection and the incomplete spatial coverage inherent to single-view representations. Diverging from single-view prior approaches, we propose a multi-view prior framework, named Multi-Modal UNet-based Feature Encoder (MUFEN), to guide diffusion models in learning comprehensive 3D hand information. Specifically, we extend conventional front-view rendering to include rear, left, right, top, and bottom perspectives, selecting the most information-rich view combination as training priors to address occlusion. This multi-view prior with a dedicated dual stream encoder significantly improves the model's understanding of complete hand features. Furthermore, we design a bounding box feature fusion module, which can fuse the gesture localization features and multi-modal features to enhance the location-awareness of the MUFEN features to the gesture-related features. Experiments demonstrate that our method achieves state-of-the-art performance in both quantitative metrics and qualitative evaluations. The source code is available at https://github.com/fuqifan/MUFEN. Qifan Fu, Xu Chen 0030, Muhammad Asad 0001, Shanxin Yuan, Changjae Oh, Gregory Slabaugh |
ACM Multimedia | 5 |
| 2025 | Learning human-to-robot handovers through 3D scene reconstructionabstractLearning robot manipulation policies from raw, real-world image data requires a large number of robot-action trials in the physical environment. Although training using simulations offers a cost-effective alternative, the visual domain gap between simulation and robot workspace remains a major limitation. Gaussian Splatting visual reconstruction methods have recently provided new directions for robot manipulation by generating realistic environments. In this paper, we propose the first method for learning supervised-based robot handovers solely from RGB images without the need of real-robot training or real-robot data collection. The proposed policy learner, Human-to-Robot Handover using Sparse-View Gaussian Splatting (H2RH-SGS), leverages sparse-view Gaussian Splatting reconstruction of human-to-robot handover scenes to generate robot demonstrations containing image-action pairs captured with a camera mounted on the robot gripper. As a result, the simulated camera pose changes in the reconstructed scene can be directly translated into gripper pose changes. We train a robot policy on demonstrations collected with 16 household objects and directly deploy this policy in the real environment. Experiments in both Gaussian Splatting reconstructed scene and real-world human-to-robot handover experiments demonstrate that H2RH-SGS serves as a new and effective representation for the human-to-robot handover task. Yuekun Wu, Yik Lung Pang, Andrea Cavallaro, Changjae Oh |
RO-MAN | 4 |
| 2025 | Spatio-temporal graph-based self-labeling for video anomaly detection
Meng Xing, Zhiyong Feng 0002, Yong Su 0003, Changjae Oh, Valeriya V. Gribova, Vladimir Fedorovich Filaretoy, De-Shuang Huang |
Neurocomputing | 5 |
| 2024 | Learning by Erasing: Conditional Entropy Based Transferable Out-of-Distribution DetectionabstractDetecting OOD inputs is crucial to deploy machine learning models to the real world safely. However, existing OOD detection methods require an in-distribution (ID) dataset to retrain the models. In this paper, we propose a Deep Generative Models (DGMs) based transferable OOD detection that does not require retraining on the new ID dataset. We first establish and substantiate two hypotheses on DGMs: DGMs exhibit a predisposition towards acquiring low-level features, in preference to semantic information; the lower bound of DGM's log-likelihoods is tied to the conditional entropy between the model input and target output. Drawing on the aforementioned hypotheses, we present an innovative image-erasing strategy, which is designed to create distinct conditional entropy distributions for each individual ID dataset. By training a DGM on a complex dataset with the proposed image-erasing strategy, the DGM could capture the discrepancy of conditional entropy distribution for varying ID datasets, without re-training. We validate the proposed method on the five datasets and show that, without retraining, our method achieves comparable performance to the state-of-the-art group-based OOD detection methods. The project codes will be open-sourced on our project website. Meng Xing, Zhiyong Feng 0002, Yong Su 0003, Changjae Oh |
AAAI | 4 |
| 2024 | Open-vocabulary object 6D pose estimationabstractWe introduce the new setting of open-vocabulary object 6D pose estimation, in which a textual prompt is used to specify the object of interest. In contrast to existing approaches, in our setting (i) the object of interest is speci-fied solely through the textual prompt, (ii) no object model (e.g., CAD or video sequence) is required at inference, and (iii) the object is imaged from two RGBD viewpoints of dif-ferent scenes. To operate in this setting, we introduce a novel approach that leverages a Vision-Language Model to segment the object of interest from the scenes and to esti-mate its relative 6D pose. The key of our approach is a carefully devised strategy to fuse object-level information provided by the prompt with local image features, resulting in a feature space that can generalize to novel concepts. We validate our approach on a new benchmark based on two popular datasets, REAL275 and Toyota-Light, which collectively encompass 34 object instances appearing in four thousand image pairs. The results demonstrate that our approach outperforms both a well-established hand-crafted method and a recent deep learning-based base-line in estimating the relative 6D pose of objects in dif-ferent scenes. Code and dataset are available at https://jcorsetti.github.io/oryon. Jaime Corsetti, Davide Boscaini, Changjae Oh, Andrea Cavallaro, Fabio Poiesi |
CVPR | 3 |
| 2024 | Diffusion-Driven GAN Inversion for Multi-Modal Face Image GenerationabstractWe present a new multimodal face image generation method that converts a text prompt and a visual input, such as a semantic mask or scribble map, into a photorealistic face image. To do this, we combine the strengths of Generative Adversarial networks (GANs) and diffusion models (DMs) by employing the multimodal features in the DM into the latent space of the pretrained GANs. We present a simple mapping and a style modulation network to link two models and convert meaningful representations in feature maps and attention maps into latent codes. With GAN inversion, the estimated latent codes can be used to generate 2D or 3D-aware facial images. We further present a multi-step training strategy that reflects textual and structural representations into the generated image. Our proposed network produces realistic 2D, multi-view, and stylized face images, which align well with inputs. We validate our method by using pretrained 2D and 3D GANs, and our results outperform existing methods. Our project page is available at https://github.com/1211sh/Diffusion-driven_GAN-Inversion/. Jihyun Kim 0009, Changjae Oh, Hoseok Do, Kwanghoon Sohn |
CVPR | 2 |
| 2024 | Improving Image De-Raining Using Reference-Guided TransformersabstractImage de-raining is a critical task in computer vision to improve visibility and enhance the robustness of outdoor vision systems. While recent advances in de-raining methods have achieved remarkable performance, the challenge remains to produce high-quality and visually pleasing derained results. In this paper, we present a reference-guided de-raining filter, a transformer network that enhances deraining results using a reference clean image as guidance. We leverage the capabilities of the proposed module to further refine the images de-rained by existing methods. We validate our method on three datasets and show that our module can improve the performance of existing prior-based, CNNbased, and transformer-based approaches. Jaehoon Cho, Changjae Oh |
ICIP | 3 |
| 2024 | Incremental Object 6D Pose Estimation
Amelia Sorrenti, Yik Lung Pang, Giovanni Bellitto, Simone Palazzo, Concetto Spampinato, Changjae Oh |
ICPR (26) | 7 |
| 2024 | Test-time adaptation for 6D pose trackingabstractWe propose a test-time adaptation for 6D object pose tracking that learns to adapt a pre-trained model to track the 6D pose of novel objects. We consider the problem of 6D object pose tracking as a 3D keypoint detection and matching task and present a model that extracts 3D keypoints. Given an RGB-D image and the mask of a target object for each frame, the proposed model consists of the self- and cross-attention modules to produce the features that aggregate the information within and across frames, respectively. By using the keypoints detected from the features for each frame, we estimate the pose changes between two frames, which enables 6D pose tracking when the 6D pose of a target object in the initial frame is given. Our model is first trained in a source domain, a category-level tracking dataset where the ground truth 6D pose of the object is available. To deploy this pre-trained model to track novel objects, we present a test-time adaptation strategy that trains the model to adapt to the target novel object by self-supervised learning. Given an RGB-D video sequence of the novel object, the proposed self-supervised losses encourage the model to estimate the 6D pose changes that can keep the photometric and geometric consistency of the object. We validate our method on the NOCS-REAL275 dataset and our collected dataset, and the results show the advantages of tracking novel objects. The collected dataset and visualisation of tracking results are available: https://bartektian.github.io/TA-6DT.html Changjae Oh, Andrea Cavallaro |
Pattern Recognit. | 2 |
| 2022 | A Wavelet-Based Dual-Stream Network for Underwater Image EnhancementabstractWe present a wavelet-based dual-stream network that addresses color cast and blurry details in underwater images. We handle these artifacts separately by decomposing an input image into multiple frequency bands using discrete wavelet transform, which generates the downsampled structure image and detail images. These sub-band images are used as input to our dual-stream network that incorporates two sub-networks: the multi-color space fusion network and the detail enhancement network. The multi-color space fusion network takes the decomposed structure image as input and estimates the color corrected output by employing the feature representations from diverse color spaces of the input. The detail enhancement network addresses the blurriness of the original underwater image by improving the image details from high-frequency sub-bands. We validate the proposed method on both real-world and synthetic underwater datasets and show the effectiveness of our model in color correction and blur removal with low computational complexity. Ziyin Ma, Changjae Oh |
ICASSP | 2 |
| 2022 | Improving Generalization of Deep Networks for Estimating Physical Properties of Containers and FillingsabstractWe present methods to estimate the physical properties of house-hold containers and their fillings manipulated by humans. We use a lightweight, pre-trained convolutional neural network with coordinate attention as a backbone model of the pipelines to accurately locate the object of interest and estimate the physical properties in the CORSMAL Containers Manipulation (CCM) dataset. We address the filling type classification with audio data and then combine this information from audio with video modalities to address the filling level classification. For the container capacity, dimension, and mass estimation, we present a data augmentation and consistency measurement to alleviate the over-fitting issue in the CCM dataset caused by the limited number of containers. We augment the training data using an object-of-interest-based re-scaling that increases the variety of physical values of the containers. We then perform the consistency measurement to choose a model with low prediction variance in the same containers under different scenes, which ensures the generalization ability of the model. Our method improves the generalization ability of the models to estimate the property of the containers that were not previously seen in the training. Hengyi Wang, Chaoran Zhu, Ziyin Ma, Changjae Oh |
ICASSP | 4 |
| 2022 | Cluster-Based 3D Keypoint Detection for Category-Agnostic 6D Pose TrackingabstractWe present a model for category-agnostic 6D pose tracking. We tackle object pose tracking as a 3D keypoint detection and matching task that does not require ground-truth annotation of the keypoints. Using RGB-D data and the target object mask as inputs, we spatially segment the point cloud of the object into clusters. Each 3D point in the cluster is characterised by features encoding appearance and geometric information. We use these features to detect a keypoint for each cluster and, with the detected keypoint sets from two time instants, we recover the pose change through least-squares optimisation. The loss functions are designed to ensure that the detected keypoints are consistent over time and suitable for pose tracking. Andrea Cavallaro, Changjae Oh |
ICIP | 3 |
| 2022 | Boosting Video Object Segmentation Based on Scale InconsistencyabstractWe present a refinement framework to boost the performance of pre-trained semi-supervised video object segmentation (VOS) models. Our work is based on scale inconsistency, which is motivated by the observation that existing VOS models generate inconsistent predictions from input frames with different sizes. We use the scale inconsistency as a clue to devise a pixel-level attention module that aggregates the advantages of the predictions from different-size inputs. The scale inconsistency is also used to regularize the training based on a pixel-level variance measured by an uncertainty estimation. We further present a self-supervised online adaptation, tailored for test-time optimization, that bootstraps the predictions without ground-truth masks based on the scale inconsistency. Experiments on DAVIS 16 and DAVIS 17 datasets show that our framework can be generically applied to various VOS models and improve their performance. Hengyi Wang, Changjae Oh |
ICME | 2 |
| 2021 | Wide and Narrow: Video Prediction from Context and Motion
Jaehoon Cho, Jiyoung Lee 0005, Changjae Oh, Wonil Song, Kwanghoon Sohn |
BMVC | 3 |
| 2021 | OHPL: One-shot Hand-eye Policy LearnerabstractThe control of a robot for manipulation tasks generally relies on object detection and pose estimation. An attractive alternative is to learn control policies directly from raw input data. However, this approach is time-consuming and expensive since learning the policy requires many trials with robot actions in the physical environment. To reduce the training cost, the policy can be learned in simulation with a large set of synthetic images. The limit of this approach is the domain gap between the simulation and the robot workspace. In this paper, we propose to learn a policy for robot reaching movements from a single image captured directly in the robot workspace from a camera placed on the end-effector (a hand-eye camera). The idea behind the proposed policy learner is that view changes seen from the hand-eye camera produced by actions in the robot workspace are analogous to locating a region-of-interest in a single image by performing sequential object localisation. This similar view change enables training of object reaching policies using reinforcement-learning-based sequential object localisation. To facilitate the adaptation of the policy to view changes in the robot workspace, we further present a dynamic filter that learns to bias an input state to remove irrelevant information for an action decision. The proposed policy learner can be used as a powerful representation for robotic tasks, and we validate it on static and moving object reaching tasks. Changjae Oh, Yik Lung Pang, Andrea Cavallaro |
IROS | 1 |
| 2021 | Towards safe human-to-robot handovers of unknown containersabstractSafe human-to-robot handovers of unknown objects require accurate estimation of hand poses and object properties, such as shape, trajectory, and weight. Accurately estimating these properties requires the use of scanned 3D object models or expensive equipment, such as motion capture systems and markers, or both. However, testing handover algorithms with robots may be dangerous for the human and, when the object is an open container with liquids, for the robot. In this paper, we propose a real-to-simulation framework to develop safe human-to-robot handovers with estimations of the physical properties of unknown cups or drinking glasses and estimations of the human hands from videos of a human manipulating the container. We complete the handover in simulation, and we estimate a region that is not occluded by the hand of the human holding the container. We also quantify the safeness of the human and object in simulation. We validate the framework using public recordings of containers manipulated before a handover and show the safeness of the handover when using noisy estimates from a range of perceptual algorithms. Yik Lung Pang, Alessio Xompero, Changjae Oh, Andrea Cavallaro |
RO-MAN | 3 |
| 2021 | View-Action Representation Learning for Active First-Person VisionabstractIn visual navigation, a moving agent equipped with a camera is traditionally controlled by an input action and the estimation of the features from a sensory state (i.e. the camera view) is treated as a pre-processing step to perform high-level vision tasks. In this paper, we present a representation learning approach that, instead, considers both state and action as inputs. We condition the encoded feature from the state transition network on the action that changes the view of the camera, thus describing the scene more effectively. Specifically, we introduce an action representation module that generates decoded higher dimensional representations from an input action to increase the representational power. We then fuse the output from the action representation module with the intermediate response of the state transition network that predicts the future state. To enhance the discrimination capability among predictions from different input actions, we further introduce triplet ranking loss and $N$ -tuplet loss functions, which in turn can be integrated with the regression loss. We demonstrate the proposed representation learning approach in reinforcement and imitation learning-based mapless navigation tasks, where the camera agent learns to navigate only through the view of the camera and the performed action, without external information. Changjae Oh, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Semantically Adversarial Learnable FiltersabstractWe present an adversarial framework to craft perturbations that mislead classifiers by accounting for the image content and the semantics of the labels. The proposed framework combines a structure loss and a semantic adversarial loss in a multi-task objective function to train a fully convolutional neural network. The structure loss helps generate perturbations whose type and magnitude are defined by a target image processing filter. The semantic adversarial loss considers groups of (semantic) labels to craft perturbations that prevent the filtered image from being classified with a label in the same group. We validate our framework with three different target filters, namely detail enhancement, log transformation and gamma correction filters; and evaluate the adversarially filtered images against three classifiers, ResNet50, ResNet18 and AlexNet, pre-trained on ImageNet. We show that the proposed framework generates filtered images with a high success rate, robustness, and transferability to unseen classifiers. We also discuss objective and subjective evaluations of the adversarial perturbations. Ali Shahin Shamsabadi, Changjae Oh, Andrea Cavallaro |
IEEE Trans. Image Process. | 2 |
| 2020 | Edgefool: an Adversarial Image Enhancement FilterabstractAdversarial examples are intentionally perturbed images that mislead classifiers. These images can, however, be easily detected using denoising algorithms, when high-frequency spatial perturbations are used, or can be noticed by humans, when perturbations are large. In this paper, we propose EdgeFool, an adversarial image enhancement filter that learns structure-aware adversarial perturbations. Edge-Fool generates adversarial images with perturbations that enhance image details via training a fully convolutional neural network end-to-end with a multi-task loss function. This loss function accounts for both image detail enhancement and class misleading objectives. We evaluate EdgeFool on three classifiers (ResNet-50, ResNet-18 and AlexNet) using two datasets (ImageNet and Private-Places365) and compare it with six adversarial methods (DeepFool, SparseFool, Carlini-Wagner, SemanticAdv, Non-targeted and Private Fast Gradient Sign Methods). Code is available at https://github.com/smartcameras/EdgeFool.git. Ali Shahin Shamsabadi, Changjae Oh, Andrea Cavallaro |
ICASSP | 2 |
| 2020 | Simultaneous Deep Stereo Matching and Dehazing with Feature Attention
Taeyong Song, Youngjung Kim, Changjae Oh, Hyunsung Jang, Namkoo Ha, Kwanghoon Sohn |
Int. J. Comput. Vis. | 3 |
| 2019 | Learning Action Representations for Self-supervised Visual ExplorationabstractLearning to efficiently navigate an environment using only an on-board camera is a difficult task for an agent when the final goal is far from the initial state and extrinsic rewards are sparse. To address this problem, we present a self-supervised prediction network to train the agent with intrinsic rewards that relate to achieving the desired final goal. The network learns to predict its future camera view (the future state) from a current state-action pair through an Action Representation Module that decodes input actions as higher dimensional representations. To increase the representational power of the network during exploration we fuse the responses from the Action Representation Module in the transition network, which predicts the future state. Moreover, to enhance the discrimination capability between predictions from different input actions we introduce joint regression and triplet ranking loss functions. We show that, despite the sparse extrinsic rewards, by learning action representations we achieve a faster training convergence than state-of-the-art methods with only a small increase in the number of the model parameters. Changjae Oh, Andrea Cavallaro |
ICRA | 1 |
| 2019 | OCEAN: Object-centric arranging network for self-supervised visual representations learning
Changjae Oh, Bumsub Ham, Hansung Kim 0001, Adrian Hilton 0001, Kwanghoon Sohn |
Expert Syst. Appl. | 1 |
| 2018 | Deep Network for Simultaneous Stereo Matching and Dehazing
Taeyong Song, Youngjung Kim, Changjae Oh, Kwanghoon Sohn |
BMVC | 3 |
| 2018 | Multi-Task Self-Supervised Visual Representation Learning for Monocular Road SegmentationabstractTraining deep networks commonly follows the supervised learning paradigm, which requires large-scale semantically-labeled data. The construction of such dataset is one of the major challenges when approaching to Advanced Driver Assistance Systems (ADAS) due to the expense of human annotation. In this paper, we explore whether unsupervised stereo-based cues can be used to learn high-level semantics for monocular road detection. Specifically, we estimate drivable space and surface normals from stereo images, which are used for pseudo ground-truth to train a convolutional neural network (CNN) as a multi-task learning scheme. Combining these multiple self-supervision tasks enables CNN to jointly encode the knowledge of obstacle and ground-plane into a single frame. We demonstrate that the feature representation learned by our multi-task approach synergistically provides a rich knowledge about geometrical characteristics. Experiments on the KITTI road dataset show that our representation outperforms state-of-the-art road detection approaches. Laehoon Cho, Youngjung Kim, Hyungjoo Jung, Changjae Oh, Jaesung Youn, Kwanghoon Sohn |
ICME | 4 |
| 2017 | Depth prediction from a single image with conditional adversarial networksabstractRecent works on machine learning have greatly advanced the accuracy of depth estimation from a single image. However, resulting depth images are still visually unsatisfactory, often producing poor boundary localization and spurious regions. In this paper, we formulate this problem from single images as a deep adversarial learning framework. A two-stage convolutional network is designed as a generator to sequentially predict global and local structures of the depth image. At the heart of our approach is a training criterion based on adversarial discriminator which attempts to distinguish between real and generated depth images as accurately as possible. Our model enables more realistic and structure-preserving depth prediction from a single image, compared to state-of-the-arts approaches. An experimental comparison demonstrates the effectiveness of our approach on large RGB-D dataset. Hyungjoo Jung, Youngjung Kim, Dongbo Min, Changjae Oh, Kwanghoon Sohn |
ICIP | 4 |
| 2017 | Personness estimation for real-time human detection on mobile devicesabstractOne aim of detection proposal methods is to reduce the computational overhead of object detection. However, most of the existing methods have significant computational overhead for real-time detection on mobile devices. A fast and accurate proposal method of human detection called personness estimation is proposed, which facilitates real-time human detection on mobile devices and can be effectively integrated into part-based detection, achieving high detection performance at a low computational cost. Our work is based on two observations: (i) normed gradients, which are designed for generic objectness estimation, effectively generate high-quality detection proposals for the person category; (ii) fusing the normed gradients with color attributes improves the performance of proposal generation for human detection. Thus, the candidate windows generated by the personness estimation will very likely contain human subjects. The human detection is then guided by the candidate windows, offering high detection performance even when the detection task terminates prior to completion. This interruptible detection scheme, called anytime detection, enables real-time human detection on mobile devices. Furthermore, we introduce a new evaluation methodology called time-recall curves to practically evaluate our approach. The applicability of our proposed method is demonstrated in extensive experiments on a publicly available dataset and a real mobile device, facilitating acquisition and enhancement of portrait photographs (e.g. selfie) on widespread mobile platforms. Kyuwon Kim, Changjae Oh, Kwanghoon Sohn |
Expert Syst. Appl. | 2 |
| 2017 | Robust interactive image segmentation using structure-aware labeling
Changjae Oh, Bumsub Ham, Kwanghoon Sohn |
Expert Syst. Appl. | 1 |
| 2016 | Point-Cut: Interactive Image Segmentation Using Point Supervision
Changjae Oh, Bumsub Ham, Kwanghoon Sohn |
ACCV (1) | 1 |
| 2016 | Edge-aware image smoothing using commute time distancesabstractMost edge-aware smoothing methods are based on the Euclidean distance to measure the similarity between adjacent pixels. This paper exploits the properties of the commute time to extend the notion of “similarity” in this context. The intuition is that since the commute time reflects the effect of all possible weighted paths between nodes (pixels), it can account for the global distribution of image features. The commute time is characterized by eigenvectors of a large Laplacian matrix, which is very costly even with sophisticated eigen-solver. To this end, we further employ a multiscale algorithm for approximating the eigenvector computation efficiently. It is analogous to the classical Nystrom's method for low rank matrix approximation. However, we do not depend on long-range connections between nodes, allowing one to include spatial coordinates in defining feature space. Extensive experimental validation demonstrates the benefits of using the commute time in a range of image processing applications, such as edge-aware image smoothing, texture filtering, and local edit propagation. Youngjung Kim, Changjae Oh, Kwanghoon Sohn |
ICIP | 2 |
| 2016 | Structure Selective Depth Superresolution for RGB-D CamerasabstractThis paper describes a method for high-quality depth superresolution. The standard formulations of image-guided depth upsampling, using simple joint filtering or quadratic optimization, lead to texture copying and depth bleeding artifacts. These artifacts are caused by inherent discrepancy of structures in data from different sensors. Although there exists some correlation between depth and intensity discontinuities, they are different in distribution and formation. To tackle this problem, we formulate an optimization model using a nonconvex regularizer. A nonlocal affinity established in a high-dimensional feature space is used to offer precisely localized depth boundaries. We show that the proposed method iteratively handles differences in structure between depth and intensity images. This property enables reducing texture copying and depth bleeding artifacts significantly on a variety of range data sets. We also propose a fast alternating direction method of multipliers algorithm to solve our optimization problem. Our solver shows a noticeable speed up compared with the conventional majorize-minimize algorithm. Extensive experiments with synthetic and real-world data sets demonstrate that the proposed method is superior to the existing methods. Youngjung Kim, Bumsub Ham, Changjae Oh, Kwanghoon Sohn |
IEEE Trans. Image Process. | 3 |
| 2015 | A majorize-minimize approach for high-quality depth upsamplingabstractThis paper describes a non-convex model that is carefully designed for high quality depth upsampling. Modern depth sensors such as time-of-flight cameras provide a promising depth measurement with video rate, but suffer from noise and low resolution. To tackle these limitations, we formulate an optimization problem using a robust potential function. In this formulation, a nonlocal principle established in the high-dimensional feature space is used to disambiguate the up-sampling problem. We also derive a numerical algorithm based on the majorization-minimization approach for efficient optimization. The proposed model iteratively creates a new affinity space that determines the influence of neighboring pixels by jointly considering spatial distance, appearance, and current estimates. This behavior enables one to significantly reduce annoying artifacts on a variety of range dataset, including a challenging real measurement. Extensive experiments demonstrate that the proposed model achieves competitive performance with state-of-the-art methods. Youngjung Kim, Sunghwan Choi, Changjae Oh, Kwanghoon Sohn |
ICIP | 3 |
| 2015 | Sparse edit propagation for high resolution image using support vector machinesabstractIn this paper, we formulate image edit propagation as a task of machine learning to handle a high resolution image efficiently. Conventional graph-based methods solve the edit propagation by minimizing an energy function which considers the relationship between a reference pixel and its spatially neighboring ones. It is becoming a time-consuming and memory-requiring task due to the increase of the image size. Inspired by the observation that similar features get analogous edits, the edit propagation is casted as a classification problem using support vector machines in the feature space. A classifier is trained with initial sparse edits given by user interaction, and then the rest of the features are classified and manipulated. In experiments, the proposed method is applied to an image recoloring to verify the performance. Experimental results show that the proposed method gives competitive editing results comparing to other state-of-the-art methods. Changjae Oh, Seungchul Ryu, Youngjung Kim, Jihyun Kim 0009, Taewoong Park, Kwanghoon Sohn |
ICIP | 1 |
| 2015 | Depth Analogy: Data-Driven Approach for Single Image Depth Estimation Using Gradient SamplesabstractInferring scene depth from a single monocular image is a highly ill-posed problem in computer vision. This paper presents a new gradient-domain approach, called depth analogy, that makes use of analogy as a means for synthesizing a target depth field, when a collection of RGB-D image pairs is given as training data. Specifically, the proposed method employs a non-parametric learning process that creates an analogous depth field by sampling reliable depth gradients using visual correspondence established on training image pairs. Unlike existing data-driven approaches that directly select depth values from training data, our framework transfers depth gradients as reconstruction cues, which are then integrated by the Poisson reconstruction. The performance of most conventional approaches relies heavily on the training RGB-D data used in the process, and such a dependency severely degenerates the quality of reconstructed depth maps when the desired depth distribution of an input image is quite different from that of the training data, e.g., outdoor versus indoor scenes. Our key observation is that using depth gradients in the reconstruction is less sensitive to scene characteristics, providing better cues for depth recovery. Thus, our gradient-domain approach can support a great variety of training range datasets that involve substantial appearance and geometric variations. The experimental results demonstrate that our (depth) gradient-domain approach outperforms existing data-driven approaches directly working on depth domain, even when only uncorrelated training datasets are available. Sunghwan Choi, Dongbo Min, Bumsub Ham, Youngjung Kim, Changjae Oh, Kwanghoon Sohn |
IEEE Trans. Image Process. | 5 |
| 2014 | Probability-Based Rendering for View SynthesisabstractIn this paper, a probability-based rendering (PBR) method is described for reconstructing an intermediate view with a steady-state matching probability (SSMP) density function. Conventionally, given multiple reference images, the intermediate view is synthesized via the depth image-based rendering technique in which geometric information (e.g., depth) is explicitly leveraged, thus leading to serious rendering artifacts on the synthesized view even with small depth errors. We address this problem by formulating the rendering process as an image fusion in which the textures of all probable matching points are adaptively blended with the SSMP representing the likelihood that points among the input reference images are matched. The PBR hence becomes more robust against depth estimation errors than existing view synthesis approaches. The MP in the steady-state, SSMP, is inferred for each pixel via the random walk with restart (RWR). The RWR always guarantees visually consistent MP, as opposed to conventional optimization schemes (e.g., diffusion or filtering-based approaches), the accuracy of which heavily depends on parameters used. Experimental results demonstrate the superiority of the PBR over the existing view synthesis approaches both qualitatively and quantitatively. Especially, the PBR is effective in suppressing flicker artifacts of virtual video rendering although no temporal aspect is considered. Moreover, it is shown that the depth map itself calculated from our RWR-based method (by simply choosing the most probable matching point) is also comparable with that of the state-of-the-art local stereo matching methods. Bumsub Ham, Dongbo Min, Changjae Oh, Minh N. Do, Kwanghoon Sohn |
IEEE Trans. Image Process. | 3 |
| 2012 | Probabilistic Correspondence Matching using Random Walk with RestartabstractThis paper presents a probabilistic method for correspondence matching with a framework of the random walk with restart (RWR). The matching cost is reformulated as a corresponding probability, which enables the RWR to be utilized for matching the correspondences. There are mainly two advantages in our method. First, the proposed method guarantees the non-trivial steady-state solution of a given initial matching probability due to the restarting term in the RWR. It means the number of iteration, a crucial parameter which influences the performance of algorithm, is not needed in contrast to the conventional methods. This gives the consistent results regardless of the evolution time. Second, only an adjacent neighborhood is considered when the matching probabilities are inferred, which lowers the computational complexity while not sacrificing performance. Experimental results show that the performance of the proposed method is competitive to that of state-of-the-art methods both qualitatively and quantitatively. Changjae Oh, Bumsub Ham, Kwanghoon Sohn |
BMVC | 1 |