VLDB 2026 Research / reviewers in the wild / expert
Somi Jeong
dblp:214/9336
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-0906-0988ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EDM: Equirectangular Projection-Oriented Dense Kernelized Feature MatchingabstractWe introduce the first learning-based dense matching algorithm, termed Equirectangular Projection-Oriented Dense Kernelized Feature Matching (EDM), specifically designed for omnidirectional images. Equirectangular projection (ERP) images, with their large fields of view, are particularly suited for dense matching techniques that aim to establish comprehensive correspondences across images. However, ERP images are subject to significant distortions, which we address by leveraging the spherical camera model and geodesic flow refinement in the dense matching method. To further mitigate these distortions, we propose spherical positional embeddings based on 3D Cartesian coordinates of the feature grid. Additionally, our method incorporates bidirectional transformations between spherical and Cartesian coordinate systems during refinement, utilizing a unit sphere to improve matching performance. We demonstrate that our proposed method achieves notable performance enhancements, with improvements of +26.72 and +42.62 in AUC@5° on the Matterport3D and Stanford2D3D datasets. Project Page: https://jdk9405.github.io/EDM Dongki Jung, Yonghan Lee 0001, Somi Jeong, Taejae Lee, Dinesh Manocha, Suyong Yeon |
CVPR | 4 |
| 2024 | EBDM: Exemplar-Guided Image Translation with Brownian-Bridge Diffusion Models
Eungbean Lee, Somi Jeong, Kwanghoon Sohn |
ECCV (13) | 2 |
| 2023 | Probabilistic Prompt Learning for Dense PredictionabstractRecent progress in deterministic prompt learning has become a promising alternative to various downstream vision tasks, enabling models to learn powerful visual representations with the help of pre-trained vision-language models. However, this approach results in limited performance for dense prediction tasks that require handling more complex and diverse objects, since a single and deterministic description cannot sufficiently represent the entire image. In this paper, we present a novel probabilistic prompt learning to fully exploit the vision-language knowledge in dense prediction tasks. First, we introduce learnable class-agnostic attribute prompts to describe universal attributes across the object class. The attributes are combined with class information and visual-context knowledge to define the class-specific textual distribution. Text representations are sampled and used to guide the dense prediction task using the probabilistic pixel-text matching loss, enhancing the stability and generalization capability of the proposed method. Extensive experiments on different dense prediction tasks and ablation studies demonstrate the effectiveness of our proposed method. Hyeongjun Kwon, Taeyong Song, Somi Jeong, Jin Kim 0005, Jinhyun Jang, Kwanghoon Sohn |
CVPR | 3 |
| 2022 | Multi-Domain Unsupervised Image-to-Image Translation with Appearance Adaptive ConvolutionabstractOver the past few years, image-to-image (I2I) translation methods have been proposed to translate a given image into diverse outputs. Despite the impressive results, they mainly focus on the I2I translation between two domains, so the multi-domain I2I translation still remains a challenge. To address this problem, we propose a novel multi-domain unsupervised image-to-image translation (MDUIT) framework that leverages the decomposed content feature and appearance adaptive convolution to translate an image into a target appearance while preserving the given geometric content. We also exploit a contrast learning objective, which improves the disentanglement ability and effectively utilizes multi-domain image data in the training process by pairing the semantically similar images. This allows our method to learn the diverse mappings between multiple visual domains with only a single framework. We show that the proposed method produces visually diverse and plausible results in multiple domains compared to the state-of-the-art methods. Somi Jeong, Jiyoung Lee 0005, Kwanghoon Sohn |
ICASSP | 1 |
| 2022 | PASTS: Toward Effective Distilling Transformer for Panoramic Semantic SegmentationabstractRecently, panoramic imaging system has been attracting a lot of attention in various real-world applications due to its all-around sensing abilities. Despite the success of semantic segmentation, the performance of panoramic segmentation is still poor because the number of annotated panoramic datasets is insufficient and existing methods cannot handle the structural distortions in panoramic images caused by wide FoV. In this paper, we present a novel PAnoramic Segmentation Transformers (PASTs) trained by a knowledge distillation strategy with teacher-student branches. We first train the teacher using labeled pinhole images. The knowledge learned from the teacher is transferred to the student via feature distillation. To this end, we exploit the distorted pinhole images to force the attention and the prediction from the teacher consistent with those from the student. In addition, we adopt the entropy loss to train the student with unlabeled panoramic images. Experimental results demonstrate the effectiveness of our method, both qualitatively and quantitatively. Jihyun Kim 0009, Somi Jeong, Kwanghoon Sohn |
ICIP | 2 |
| 2022 | Enriching SAR Ship Detection via Multistage Domain AlignmentabstractThe advent of deep learning has made a significant advance in ship detection in synthetic aperture radar (SAR) images. However, it is still challenging since the amount of labeled SAR samples for training is not sufficient. Moreover, SAR images are corrupted by speckle noise, making them complex and difficult to interpret even by human experts. In this letter, we propose a novel SAR ship detection framework that leverages label-rich electro-optical (EO) images for more plentiful feature representations, and delicately addresses the speckle noise in SAR images. To this end, we first introduce a multistage domain alignment module that reduces the distribution discrepancies between EO and SAR feature maps at local, global, and instance levels. This allows enriching SAR representations by gradually instilling cross-domain knowledge from a large-scale EO image dataset. We further design a blind-spot layer for feature extraction to suppress the influence of speckles. Experimental results on the high-resolution SAR images dataset (HRSID) show that our detection performance achieves average precision (AP) 5.5% better than the current state-of-the-arts that exploits SAR images only. Our method significantly improves the detection performance with higher speckle noises, demonstrating stronger robustness than the conventional methods. Somi Jeong, Youngjung Kim, Sung-Ho Kim 0004, Kwanghoon Sohn |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Semantic Equalization Learning for Semi-Supervised SAR Building SegmentationabstractSynthetic aperture radar (SAR) building segmentation, which is one of the fundamental tasks in the remote sensing community, has been achieved remarkable performance using convolutional neural networks (CNNs). Since most methods do not consider distinctive characteristics of SAR images, they tend to be biased towards simple and large buildings while ignoring small and complex-shaped ones. To build a general and powerful SAR building segmentation model, in this letter, we introduce a semi-supervised learning (SSL) framework with semantic equalization learning (SEL). Concretely, we leverage labeled SAR and EO image pairs and unlabeled SAR images for SSL to extract representative SAR features with the help of context-rich EO features. Moreover, SEL aims to balance the training of well and poor-performing samples via our purposed data augmentation technique and the objective functions. It consists of a semantic proportional CutMix (SP-CutMix) module to increase the sampling probability of under-performed samples during the training phase, and an equalized segmentation loss (ESL) to adjust the loss contribution depending on difficulties. By doing so, our method prevents the model from being biased to easy samples and increases the performance of difficult building samples. Experimental results on the SpaceNet-6 benchmark demonstrate the effectiveness of our framework, especially by significantly improving the most challenging scenarios, that is less labeled data available. Eungbean Lee, Somi Jeong, Jun Hee Kim, Kwanghoon Sohn |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Memory-Guided Unsupervised Image-to-Image TranslationabstractWe present a novel unsupervised framework for instance-level image-to-image translation. Although recent advances have been made by incorporating additional object annotations, existing methods often fail to handle images with multiple disparate objects. The main cause is that, during inference, they apply a global style to the whole image and do not consider the large style discrepancy between instance and background, or within instances. To address this problem, we propose a class-aware memory network that explicitly reasons about local style variations. A key-values memory structure, with a set of read/update operations, is introduced to record class-wise style variations and access them without requiring an object detector at the test time. The key stores a domain-agnostic content representation for allocating memory items, while the values encode domain-specific style representations. We also present a feature contrastive loss to boost the discriminative power of memory items. We show that by incorporating our memory, we can transfer class-aware and accurate style representations across domains. Experimental results demonstrate that our model outperforms recent instance-level methods and achieves state-of-the-art performance. Somi Jeong, Youngjung Kim, Eungbean Lee, Kwanghoon Sohn |
CVPR | 1 |
| 2021 | Stereo-augmented Depth Completion from a Single RGB-LiDAR imageabstractDepth completion is an important task in computer vision and robotics applications, which aims at predicting accurate dense depth from a single RGB-LiDAR image. Convolutional neural networks (CNNs) have been widely used for depth completion to learn a mapping function from sparse to dense depth. However, recent methods do not exploit any 3D geometric cues during the inference stage and mainly rely on sophisticated CNN architectures. In this paper, we present a cascade and geometrically inspired learning framework for depth completion, consisting of three stages: view extrapolation, stereo matching, and depth refinement. The first stage extrapolates a virtual (right) view using a single RGB (left) and its LiDAR data. We then mimic the binocular stereo-matching, and as a result, explicitly encode geometric constraints during depth completion. This stage augments the final refinement process by providing additional geometric reasoning. We also introduce a distillation framework based on teacher-student strategy to effectively train our network. Knowledge from a teacher model privileged with real stereo pairs is transferred to the student through feature distillation. Experimental results on KITTI depth completion benchmark demonstrate that the proposed method is superior to state-of-the-art methods. Keunhoon Choi, Somi Jeong, Youngjung Kim, Kwanghoon Sohn |
ICRA | 2 |
| 2021 | Privileged Knowledge Distillation for SAR Building ExtractionabstractAutomatic building footprint extraction from SAR imagery is one of the critical tasks in the remote sensing community. CNN has been recently explored in building extraction tasks and achieved improved performance. However, due to the scarcity of training data, it suffers from overfitting problem. This paper presents a novel knowledge distillation based framework consisting of teacher and student networks. Regarding EO image as privileged information, the teacher network learns to extract the rich EO and SAR image pair features. The student network then learns to estimate the building footprints from only SAR images based on the privileged knowledge from the teacher network. Experimental results on SpaceNet-6 benchmark demonstrate the effectiveness of our framework, which explicitly improves the performance of SAR segmentation network. Eungbean Lee, Somi Jeong, Kwanghoon Sohn |
IGARSS | 2 |
| 2020 | Stereoscopic Image Super-Resolution with Stereo Consistent FeatureabstractWe present a first attempt for stereoscopic image super-resolution (SR) for recovering high-resolution details while preserving stereo-consistency between stereoscopic image pair. The most challenging issue in the stereoscopic SR is that the texture details should be consistent for corresponding pixels in stereoscopic SR image pair. However, existing stereo SR methods cannot maintain the stereo-consistency, thus causing 3D fatigue to the viewers. To address this issue, in this paper, we propose a self and parallax attention mechanism (SPAM) to aggregate the information from its own image and the counterpart stereo image simultaneously, thus reconstructing high-quality stereoscopic SR image pairs. Moreover, we design an efficient network architecture and effective loss functions to enforce stereo-consistency constraint. Finally, experimental results demonstrate the superiority of our method over state-of-the-art SR methods in terms of both quantitative metrics and qualitative visual quality while maintaining stereo-consistency between stereoscopic image pair. Wonil Song, Sungil Choi, Somi Jeong, Kwanghoon Sohn |
AAAI | 3 |
| 2019 | Semantic Attribute Matching NetworksabstractWe present semantic attribute matching networks (SAM-Net) for jointly establishing correspondences and transferring attributes across semantically similar images, which intelligently weaves the advantages of the two tasks while overcoming their limitations. SAM-Net accomplishes this through an iterative process of establishing reliable correspondences by reducing the attribute discrepancy between the images and synthesizing attribute transferred images using the learned correspondences. To learn the networks using weak supervisions in the form of image pairs, we present a semantic attribute matching loss based on the matching similarity between an attribute transferred source feature and a warped target feature. With SAM-Net, the state-of-the-art performance is attained on several benchmarks for semantic matching and attribute transfer. Seungryong Kim, Dongbo Min, Somi Jeong, Sunok Kim, Sangryul Jeon, Kwanghoon Sohn |
CVPR | 3 |
| 2019 | Learning to Find Unpaired Cross-Spectral CorrespondencesabstractWe present a deep architecture and learning framework for establishing correspondences across cross-spectral visible and infrared images in an unpaired setting. To overcome the unpaired cross-spectral data problem, we design the unified image translation and feature extraction modules to be learned in a joint and boosting manner. Concretely, the image translation module is learned only with the unpaired cross-spectral data, and the feature extraction module is learned with an input image and its translated image. By learning two modules simultaneously, the image translation module generates the translated image that preserves not only the domain-specific attributes with separate latent spaces but also the domain-agnostic contents with feature consistency constraint. In an inference phase, the cross-spectral feature similarity is augmented by intra-spectral similarities between the features extracted from the translated images. Experimental results show that this model outperforms the state-of-the-art unpaired image translation methods and cross-spectral feature descriptors on various visible and infrared benchmarks. Somi Jeong, Seungryong Kim, Kihong Park, Kwanghoon Sohn |
IEEE Trans. Image Process. | 1 |
| 2017 | Convolutional cost aggregation for robust stereo matchingabstractAlthough convolutional neural network (CNN)-based stereo matching methods have become increasingly popular thanks to their robustness, they primarily have been focused on the matching cost computation. By leveraging CNNs, we present a novel method for matching cost aggregation to boost the stereo matching performance. Our insight is to learn the convolution kernel within CNN architecture for cost aggregation in a fully convolutional manner. Tailored to cost aggregation problem, our method differs from handcrafted methods in terms of its convolutional aggregation through optimally learned CNNs. First, the matching cost is aggregated with cost volume unary network, and then optimized with explicit disparity boundary, estimated through disparity boundary pairwise network, within a global energy minimization. Experiments demonstrate that our method outperforms conventional hand-crafted aggregation methods. Somi Jeong, Seungryong Kim, Bumsub Ham, Kwanghoon Sohn |
ICIP | 1 |