VLDB 2026 Research / reviewers in the wild / expert
Sangryul Jeon
dblp:195/6099
· DBLP profile ↗
16ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0003-0991-6165ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
3D vision · 61% Representation and self-supervised learning · 16% Reinforcement learning · 10% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 23 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › correspondence estimation
semantic correspondence |
2.7 | 7 | 2022 | Pyramidal Semantic Correspondence Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Neural Matching Fields: Implicit Representation of Matching Fields for Visual Correspondence · NeurIPS 2022 CATs: Cost Aggregation Transformers for Visual Correspondence · NeurIPS 2021 |
Computer vision › 3D vision › correspondence estimation
dense correspondence |
1.2 | 3 | 2021 | CATs: Cost Aggregation Transformers for Visual Correspondence · NeurIPS 2021 Joint Learning of Semantic Alignment and Object Landmark Detection · ICCV 2019 Recurrent Transformer Networks for Semantic Correspondence · NeurIPS 2018 |
Computer vision › 3D vision › feature matching › dense feature matching
dense semantic correspondence |
1.2 | 3 | 2022 | Pyramidal Semantic Correspondence Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022 PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence · ECCV (6) 2018 FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence · CVPR 2017 |
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning |
0.7 | 1 | 2023 | Local-Guided Global: Paired Similarity Representation for Visual Reinforcement Learning · CVPR 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.7 | 1 | 2023 | Local-Guided Global: Paired Similarity Representation for Visual Reinforcement Learning · CVPR 2023 |
Machine learning › Reinforcement learning › deep reinforcement learning
visual reinforcement learning |
0.7 | 1 | 2023 | Local-Guided Global: Paired Similarity Representation for Visual Reinforcement Learning · CVPR 2023 |
Computer vision › 3D vision
implicit neural representation |
0.6 | 1 | 2022 | Neural Matching Fields: Implicit Representation of Matching Fields for Visual Correspondence · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.5 | 1 | 2021 | Mining Better Samples for Contrastive Learning of Temporal Correspondence · CVPR 2021 |
Computer vision › 3D vision › stereo vision › stereo matching
cost aggregation |
0.5 | 1 | 2021 | CATs: Cost Aggregation Transformers for Visual Correspondence · NeurIPS 2021 |
Computer vision › 3D vision › correspondence estimation
image correspondence |
0.5 | 1 | 2021 | CATs: Cost Aggregation Transformers for Visual Correspondence · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning › visual representation
pixel representation |
0.5 | 1 | 2021 | Mining Better Samples for Contrastive Learning of Temporal Correspondence · CVPR 2021 |
Computer vision › Video understanding and tracking
temporal alignment |
0.5 | 1 | 2021 | Mining Better Samples for Contrastive Learning of Temporal Correspondence · CVPR 2021 |
Computer vision › 3D vision › correspondence estimation
semantic flow |
0.4 | 1 | 2020 | Guided Semantic Flow · ECCV (28) 2020 |
Computer vision › 3D vision
attribute alignment |
0.4 | 1 | 2019 | Semantic Attribute Matching Networks · CVPR 2019 |
Machine learning › Transfer learning and domain adaptation
attribute transfer |
0.4 | 1 | 2019 | Semantic Attribute Matching Networks · CVPR 2019 |
Computer vision › 3D vision
object landmark detection |
0.4 | 1 | 2019 | Joint Learning of Semantic Alignment and Object Landmark Detection · ICCV 2019 |
Machine learning › Representation and self-supervised learning
semantic alignment |
0.4 | 1 | 2019 | Joint Learning of Semantic Alignment and Object Landmark Detection · ICCV 2019 |
Machine learning › Deep learning architectures and training › transformer
recurrent transformer |
0.3 | 1 | 2018 | Recurrent Transformer Networks for Semantic Correspondence · NeurIPS 2018 |
Machine learning › Deep learning architectures and training
transformer |
0.3 | 1 | 2018 | Recurrent Transformer Networks for Semantic Correspondence · NeurIPS 2018 |
Machine learning › Learning paradigms
weakly supervised learning |
0.3 | 3 | 2019 | Semantic Attribute Matching Networks · CVPR 2019 Recurrent Transformer Networks for Semantic Correspondence · NeurIPS 2018 FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence · CVPR 2017 |
Computer vision › 3D vision › local feature descriptor
self-similarity descriptor |
0.3 | 1 | 2017 | FCSS: Fully Convolutional Self-Similarity for Dense Semantic Correspondence · CVPR 2017 |
Image and video processing › motion estimation
optical flow |
0.1 | 1 | 2020 | Guided Semantic Flow · ECCV (28) 2020 |
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning |
0.1 | 1 | 2019 | Joint Learning of Semantic Alignment and Object Landmark Detection · ICCV 2019 |
Methods — techniques the papers use, named apart from their topics
self-supervised learning · 0.7contrastive similarity constraints · 0.7action-aware transform · 0.7weakly supervised learning · 0.6patchmatch · 0.6cost embedding network · 0.6coordinate optimization · 0.6convolutional neural network · 0.6coarse-to-fine estimation · 0.6contrastive learning · 0.5guided filtering · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Zero-shot Building Attribute Extraction from Large-Scale Vision and Language ModelsabstractExisting building recognition methods, exemplified by BRAILS, utilize supervised learning to extract information from satellite and street-view images for classification and segmentation. However, each task module requires human-annotated data, hindering the scalability and robustness to regional variations and annotation imbalances. In response, we propose a new zero-shot workflow for building attribute extraction that utilizes large-scale vision and language models to mitigate reliance on external annotations. The proposed workflow contains two key components: image-level captioning and segment-level captioning for the building images based on the vocabularies pertinent to structural and civil engineering. These two components generate descriptive captions by computing feature representations of the image and the vocabularies, and facilitating a semantic match between the visual and textual representations. Consequently, our framework offers a promising avenue to enhance AI-driven captioning for building attribute extraction in the structural and civil engineering domains, ultimately reducing reliance on human annotations while bolstering performance and adaptability. Sangryul Jeon, Frank McKenna, Stella X. Yu |
WACV | 2 |
| 2023 | Local-Guided Global: Paired Similarity Representation for Visual Reinforcement LearningabstractRecent vision-based reinforcement learning (RL) methods have found extracting high-level features from raw pixels with self-supervised learning to be effective in learning policies. However, these methods focus on learning global representations of images, and disregard local spatial structures present in the consecutively stacked frames. In this paper, we propose a novel approach, termed self-supervised Paired Similarity Representation Learning (PSRL) for effectively encoding spatial structures in an unsupervised manner. Given the input frames, the latent volumes are first generated individually using an encoder, and they are used to capture the variance in terms of local spatial structures, i.e., correspondence maps among multiple frames. This enables for providing plenty of fine-grained samples for training the encoder of deep RL. We further attempt to learn the global semantic representations in the action aware transform module that predicts future state representations using action vectors as a medium. The proposed method imposes similarity constraints on the three latent volumes; transformed query representations by estimated pixel-wise correspondence, predicted query representations from the action aware transform model, and target representations of future state, guiding action aware transform with locality-inherent volume. Experimental results on complex tasks in Atari Games and DeepMind Control Suite demonstrate that the RL methods are significantly boosted by the proposed self-supervised learning of paired similarity representations. Hyesong Choi, Hunsang Lee, Wonil Song, Sangryul Jeon, Kwanghoon Sohn, Dongbo Min |
CVPR | 4 |
| 2023 | Learning disentangled skills for hierarchical reinforcement learning through trajectory autoencoder with weak labels
Wonil Song, Sangryul Jeon, Hyesong Choi, Kwanghoon Sohn, Dongbo Min |
Expert Syst. Appl. | 2 |
| 2022 | COAT: Correspondence-driven Object Appearance Transfer
Sangryul Jeon, Zhe Lin 0001, Scott Cohen, Zhihong Ding, Kwanghoon Sohn |
BMVC | 1 |
| 2022 | Neural Matching Fields: Implicit Representation of Matching Fields for Visual CorrespondenceabstractExisting pipelines of semantic correspondence commonly include extracting high-level semantic features for the invariance against intra-class variations and background clutters. This architecture, however, inevitably results in a low-resolution matching field that additionally requires an ad-hoc interpolation process as a post-processing for converting it into a high-resolution one, certainly limiting the overall performance of matching results. To overcome this, inspired by recent success of implicit neural representation, we present a novel method for semantic correspondence, called Neural Matching Field (NeMF). However, complicacy and high-dimensionality of a 4D matching field are the major hindrances, which we propose a cost embedding network to process a coarse cost volume to use as a guidance for establishing high-precision matching field through the following fully-connected network. Nevertheless, learning a high-dimensional matching field remains challenging mainly due to computational complexity, since a na\"ive exhaustive inference would require querying from all pixels in the 4D space to infer pixel-wise correspondences. To overcome this, we propose adequate training and inference procedures, which in the training phase, we randomly sample matching candidates and in the inference phase, we iteratively performs PatchMatch-based inference and coordinate optimization at test time. With these combined, competitive results are attained on several standard benchmarks for semantic correspondence. Code and pre-trained weights are available at~\url{https://ku-cvlab.github.io/NeMF/}. Sunghwan Hong, Jisu Nam, Seokju Cho, Susung Hong, Sangryul Jeon, Dongbo Min, Seungryong Kim |
NeurIPS | 5 |
| 2022 | Pyramidal Semantic Correspondence NetworksabstractThis paper presents a deep architecture, called pyramidal semantic correspondence networks (PSCNet), that estimates locally-varying affine transformation fields across semantically similar images. To deal with large appearance and shape variations that commonly exist among different instances within the same object category, we leverage a pyramidal model where the affine transformation fields are progressively estimated in a coarse-to-fine manner so that the smoothness constraint is naturally imposed. Different from the previous methods which directly estimate global or local deformations, our method first starts to estimate the transformation from an entire image and then progressively increases the degree of freedom of the transformation by dividing coarse cell into finer ones. To this end, we propose two spatial pyramid models by dividing an image in a form of quad-tree rectangles or into multiple semantic elements of an object. Additionally, to overcome the limitation of insufficient training data, a novel weakly-supervised training scheme is introduced that generates progressively evolving supervisions through the spatial pyramid models by leveraging a correspondence consistency across image pairs. Extensive experimental results on various benchmarks including TSS, Proposal Flow-WILLOW, Proposal Flow-PASCAL, Caltech-101, and SPair-71k demonstrate that the proposed method outperforms the lastest methods for dense semantic correspondence. Sangryul Jeon, Seungryong Kim, Dongbo Min, Kwanghoon Sohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Mining Better Samples for Contrastive Learning of Temporal CorrespondenceabstractWe present a novel framework for contrastive learning of pixel-level representation using only unlabeled video. Without the need of ground-truth annotation, our method is capable of collecting well-defined positive correspondences by measuring their confidences and well-defined negative ones by appropriately adjusting their hardness during training. This allows us to suppress the adverse impact of ambiguous matches and prevent a trivial solution from being yielded by too hard or too easy negative samples. To accomplish this, we incorporate three different criteria that ranges from a pixel-level matching confidence to a video-level one into a bottom-up pipeline, and plan a curriculum that is aware of current representation power for the adaptive hardness of negative samples during training. With the proposed method, state-of-the-art performance is attained over the latest approaches on several video label propagation tasks. Sangryul Jeon, Dongbo Min, Seungryong Kim, Kwanghoon Sohn |
CVPR | 1 |
| 2021 | CATs: Cost Aggregation Transformers for Visual CorrespondenceabstractWe propose a novel cost aggregation network, called Cost Aggregation Transformers (CATs), to find dense correspondences between semantically similar images with additional challenges posed by large intra-class appearance and geometric variations. Cost aggregation is a highly important process in matching tasks, which the matching accuracy depends on the quality of its output. Compared to hand-crafted or CNN-based methods addressing the cost aggregation, in that either lacks robustness to severe deformations or inherit the limitation of CNNs that fail to discriminate incorrect matches due to limited receptive fields, CATs explore global consensus among initial correlation map with the help of some architectural designs that allow us to fully leverage self-attention mechanism. Specifically, we include appearance affinity modeling to aid the cost aggregation process in order to disambiguate the noisy initial correlation maps and propose multi-level aggregation to efficiently capture different semantics from hierarchical feature representations. We then combine with swapping self-attention technique and residual connections not only to enforce consistent matching, but also to ease the learning process, which we find that these result in an apparent performance boost. We conduct experiments to demonstrate the effectiveness of the proposed model over the latest methods and provide extensive ablation studies. Code and trained models are available at https://sunghwanhong.github.io/CATs/. Seokju Cho, Sunghwan Hong, Sangryul Jeon, Yunsung Lee, Kwanghoon Sohn, Seungryong Kim |
NeurIPS | 3 |
| 2020 | Guided Semantic Flow
Sangryul Jeon, Dongbo Min, Seungryong Kim, Jihwan Choe, Kwanghoon Sohn |
ECCV (28) | 1 |
| 2019 | Semantic Attribute Matching NetworksabstractWe present semantic attribute matching networks (SAM-Net) for jointly establishing correspondences and transferring attributes across semantically similar images, which intelligently weaves the advantages of the two tasks while overcoming their limitations. SAM-Net accomplishes this through an iterative process of establishing reliable correspondences by reducing the attribute discrepancy between the images and synthesizing attribute transferred images using the learned correspondences. To learn the networks using weak supervisions in the form of image pairs, we present a semantic attribute matching loss based on the matching similarity between an attribute transferred source feature and a warped target feature. With SAM-Net, the state-of-the-art performance is attained on several benchmarks for semantic matching and attribute transfer. Seungryong Kim, Dongbo Min, Somi Jeong, Sunok Kim, Sangryul Jeon, Kwanghoon Sohn |
CVPR | 5 |
| 2019 | Joint Learning of Semantic Alignment and Object Landmark DetectionabstractConvolutional neural networks (CNNs) based approaches for semantic alignment and object landmark detection have improved their performance significantly. Current efforts for the two tasks focus on addressing the lack of massive training data through weakly- or unsupervised learning frameworks. In this paper, we present a joint learning approach for obtaining dense correspondences and discovering object landmarks from semantically similar images. Based on the key insight that the two tasks can mutually provide supervisions to each other, our networks accomplish this through a joint loss function that alternatively imposes a consistency constraint between the two tasks, thereby boosting the performance and addressing the lack of training data in a principled manner. To the best of our knowledge, this is the first attempt to address the lack of training data for the two tasks through the joint learning. To further improve the robustness of our framework, we introduce a probabilistic learning formulation that allows only reliable matches to be used in the joint learning process. With the proposed method, state-of-the-art performance is attained on several benchmarks for semantic matching and landmark detection. Sangryul Jeon, Dongbo Min, Seungryong Kim, Kwanghoon Sohn |
ICCV | 1 |
| 2019 | Graph Regularization Network with Semantic Affinity for Weakly-Supervised Temporal Action LocalizationabstractThis paper presents a novel deep architecture for weakly-supervised temporal action localization that not only generates segment-level action responses but also propagates segment-level responses to the neighborhood in a form of graph Laplacian regularization. Specifically, our approach consists of two sub-modules; a class activation module to estimate the action score map over time through the action classifiers, and a graph regularization module to refine the estimated action score map by solving a quadratic programming problem with the predicted segment-level semantic affinities. Since these two modules are integrated with fully differentiable layers, the proposed networks can be jointly trained in an end-to-end manner. Experimental results on Thumos14 and ActivityNet1.2 demonstrate that the proposed method provides outstanding performances in weakly-supervised temporal action localization. Jungin Park, Jiyoung Lee 0005, Sangryul Jeon, Seungryong Kim, Kwanghoon Sohn |
ICIP | 3 |
| 2018 | PARN: Pyramidal Affine Regression Networks for Dense Semantic Correspondence
Sangryul Jeon, Seungryong Kim, Dongbo Min, Kwanghoon Sohn |
ECCV (6) | 1 |
| 2018 | Recurrent Transformer Networks for Semantic CorrespondenceabstractWe present recurrent transformer networks (RTNs) for obtaining dense correspondences between semantically similar images. Our networks accomplish this through an iterative process of estimating spatial transformations between the input images and using these transformations to generate aligned convolutional activations. By directly estimating the transformations between an image pair, rather than employing spatial transformer networks to independently normalize each individual image, we show that greater accuracy can be achieved. This process is conducted in a recursive manner to refine both the transformation estimates and the feature representations. In addition, a technique is presented for weakly-supervised training of RTNs that is based on a proposed classification loss. With RTNs, state-of-the-art performance is attained on several benchmarks for semantic correspondence. Seungryong Kim, Stephen Lin 0001, Sangryul Jeon, Dongbo Min, Kwanghoon Sohn |
NeurIPS | 3 |
| 2017 | FCSS: Fully Convolutional Self-Similarity for Dense Semantic CorrespondenceabstractWe present a descriptor, called fully convolutional self-similarity (FCSS), for dense semantic correspondence. To robustly match points among different instances within the same object class, we formulate FCSS using local self-similarity (LSS) within a fully convolutional network. In contrast to existing CNN-based descriptors, FCSS is inherently insensitive to intra-class appearance variations because of its LSS-based structure, while maintaining the precise localization ability of deep neural networks. The sampling patterns of local structure and the self-similarity measure are jointly learned within the proposed network in an end-to-end and multi-scale manner. As training data for semantic correspondence is rather limited, we propose to leverage object candidate priors provided in existing image datasets and also correspondence consistency between object pairs to enable weakly-supervised learning. Experiments demonstrate that FCSS outperforms conventional handcrafted descriptors and CNN-based descriptors on various benchmarks. Seungryong Kim, Dongbo Min, Bumsub Ham, Sangryul Jeon, Stephen Lin 0001, Kwanghoon Sohn |
CVPR | 4 |
| 2017 | Convolutional feature pyramid fusion via attention networkabstractWe present a novel fusion scheme between multiple intermediate convolutional features within convolutional neurual network (CNN) for dense correspondence estimation. In contrast to existing CNN-based descriptors that utilize a single convolutional activation, our approach jointly uses multiple intermediate features of CNN through the attention weight that balances the contribution of each features. We formulate the overall network as two sub-networks, correspondence network and attention network. The correspondence network is designed to provide multiple intermediate matching costs while the attention network is to learn the optimal weight between them. These two networks are learned in a joint manner to boost the correspondence estimation performance. Experiments demonstrate that our proposed method outperforms the state-of-the-art methods on various correspondence estimation tasks including depth estimation, optical flow, and semantic correspondence. Sangryul Jeon, Seungryong Kim, Kwanghoon Sohn |
ICIP | 1 |