VLDB 2026 Research / reviewers in the wild / expert
Sunghun Joung
dblp:214/8742
· DBLP profile ↗
9ranked-venue papers
4as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Segmentation and scene understanding · 25% Face, body and person analysis · 25% 3D vision · 23% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding › semantic segmentation › transfer learning for semantic segmentation
domain adaptive semantic segmentation |
0.5 | 1 | 2021 | Cross-Domain Grouping and Alignment for Domain Adaptive Semantic Segmentation · AAAI 2021 |
Computer vision › Image recognition and object detection › image classification
fine-grained image classification |
0.5 | 1 | 2021 | Learning Canonical 3D Object Representation for Fine-Grained Recognition · ICCV 2021 |
Computer vision › 3D vision
object representation |
0.5 | 1 | 2021 | Learning Canonical 3D Object Representation for Fine-Grained Recognition · ICCV 2021 |
Computer vision › Face, body and person analysis
person re-identification |
0.5 | 1 | 2021 | Prototype-Guided Saliency Feature Learning for Person Search · CVPR 2021 |
Computer vision › Face, body and person analysis
person search |
0.5 | 1 | 2021 | Prototype-Guided Saliency Feature Learning for Person Search · CVPR 2021 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.5 | 1 | 2021 | Cross-Domain Grouping and Alignment for Domain Adaptive Semantic Segmentation · AAAI 2021 |
Computer vision › 3D vision › object pose estimation
viewpoint estimation |
0.4 | 1 | 2020 | Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint Estimation · CVPR 2020 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.1 | 1 | 2020 | Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint Estimation · CVPR 2020 |
Methods — techniques the papers use, named apart from their topics
prototype learning · 0.5loss function design · 0.5learnable clustering · 0.5differentiable rendering · 0.5attention network · 0.5analysis-by-synthesis · 0.5OIM loss · 0.5sinusoidal soft-argmax · 0.4cylindrical convolutional networks · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Learning Semantic Keypoints for Object Detection in Aerial ImagesabstractObject detection in aerial images has achieved remarkable progress with the advent of deep convolutional neural networks (CNNs). It is, however, still a challenging task since the objects in aerial images are arbitrarily oriented and often densely packed. In this letter, we propose a novel method for oriented object detection in aerial images that represents objects as rotation equivariant semantic keypoints. Unlike conventional methods that represent object rotation according to angles from each axis in the Cartesian coordinate system, we represent object using a canonical orientation to ensure rotation equivariance. We accomplish this by representing an object as semantic keypoints, where each keypoint of the object consistently corresponds to the semantic part, regardless of rotation variation. To this end, we define the “head” point of the object as the canonical orientation and the remaining bounding box vectors as semantic keypoints in clockwise order. To discriminate visual attributes between different categories, we further use category-specific semantic keypoints, so that object classification and localization can be jointly solved in a cooperative manner. Our experiments demonstrate the effectiveness of rotation equivariant semantic keypoints on oriented object detection. Sunghun Joung, Taeyong Song, Hanjae Kim, Kwanghoon Sohn |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Cross-Domain Grouping and Alignment for Domain Adaptive Semantic SegmentationabstractExisting techniques to adapt semantic segmentation networks across source and target domains within deep convolutional neural networks (CNNs) deal with all the samples from the two domains in a global or category-aware manner. They do not consider an inter-class variation within the target domain itself or estimated category, providing the limitation to encode the domains having a multi-modal data distribution. To overcome this limitation, we introduce a learnable clustering module, and a novel domain adaptation framework, called cross-domain grouping and alignment. To cluster the samples across domains with an aim to maximize the domain alignment without forgetting precise segmentation ability on the source domain, we present two loss functions, in particular, for encouraging semantic consistency and orthogonality among the clusters. We also present a loss so as to solve a class imbalance problem, which is the other limitation of the previous methods. Our experiments show that our method consistently boosts the adaptation performance in semantic segmentation, outperforming the state-of-the-arts on various domain adaptation settings. Sunghun Joung, Seungryong Kim, Jungin Park, Ig-Jae Kim, Kwanghoon Sohn |
AAAI | 2 |
| 2021 | Prototype-Guided Saliency Feature Learning for Person SearchabstractExisting person search methods integrate person detection and re-identification (re-ID) module into a unified system. Though promising results have been achieved, the misalignment problem, which commonly occurs in person search, limits the discriminative feature representation for re-ID. To overcome this limitation, we introduce a novel framework to learn the discriminative representation by utilizing prototype in OIM loss. Unlike conventional methods using prototype as a representation of person identity, we utilize it as guidance to allow the attention network to consistently highlight multiple instances across different poses. Moreover, we propose a new prototype update scheme with adaptive momentum to increase the discriminative ability across different instances. Extensive ablation experiments demonstrate that our method can significantly enhance the feature discriminative power, outperforming the state-of-the-art results on two person search benchmarks including CUHK-SYSU and PRW. Hanjae Kim, Sunghun Joung, Ig-Jae Kim, Kwanghoon Sohn |
CVPR | 2 |
| 2021 | Learning Canonical 3D Object Representation for Fine-Grained RecognitionabstractWe propose a novel framework for fine-grained object recognition that learns to recover object variation in 3D space from a single image, trained on an image collection without using any ground-truth 3D annotation. We accomplish this by representing an object as a composition of 3D shape and its appearance, while eliminating the effect of camera viewpoint, in a canonical configuration. Unlike conventional methods modeling spatial variation in 2D images only, our method is capable of reconfiguring the appearance feature in a canonical 3D space, thus enabling the subsequent object classifier to be invariant under 3D geometric variation. Our representation also allows us to go beyond existing methods, by incorporating 3D shape variation as an additional cue for object recognition. To learn the model without ground-truth 3D annotation, we deploy a differentiable renderer in an analysis-by-synthesis frame- work. By incorporating 3D shape and appearance jointly in a deep representation, our method learns the discriminative representation of the object and achieves competitive performance on fine-grained image recognition and vehicle re-identification. We also demonstrate that the performance of 3D shape reconstruction is improved by learning fine-grained shape deformation in a boosting manner. Sunghun Joung, Seungryong Kim, Ig-Jae Kim, Kwanghoon Sohn |
ICCV | 1 |
| 2020 | Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint EstimationabstractExisting techniques to encode spatial invariance within deep convolutional neural networks only model 2D transformation fields. This does not account for the fact that objects in a 2D space are a projection of 3D ones, and thus they have limited ability to severe object viewpoint changes. To overcome this limitation, we introduce a learnable module, cylindrical convolutional networks (CCNs), that exploit cylindrical representation of a convolutional kernel defined in the 3D space. CCNs extract a view-specific feature through a view-specific convolutional kernel to predict object category scores at each viewpoint. With the view-specific feature, we simultaneously determine objective category and viewpoints using the proposed sinusoidal soft-argmax module. Our experiments demonstrate the effectiveness of the cylindrical convolutional networks on joint object detection and viewpoint estimation. Sunghun Joung, Seungryong Kim, Hanjae Kim, Ig-Jae Kim, Junghyun Cho, Kwanghoon Sohn |
CVPR | 1 |
| 2020 | Shape-Adaptive Kernel Network for Dense Object DetectionabstractDense object detectors that are applied over a regular, dense grid have advanced and drawn their attention in recent days. Their fully convolutional nature greatly advances the computational efficiency of object detectors compared to the two-stage detectors. However, the lack of the ability to adjust shape variation on a regular grid is still limited. In this paper we introduce a new framework, shape-adaptive kernel network, to handle spatial manipulation of input data in convolutional kernel space. At the heart of out approach is to align the original kernel space recovering shape variation of each input feature on regular grid. To this end, we propose a shape-adaptive kernel sampler to adjust dynamic convolutional kernel conditioned on input. To increase the flexibility of geometric transformation, a cascade refinement module is designed, which first estimates the global transformation grid and then estimates local offset in convolutional kernel space. Our experiments demonstrate the effectiveness of the shape-adaptive kernel network for dense object detection on various benchmarks. Hanjae Kim, Sunghun Joung, Ig-Jae Kim, Kwanghoon Sohn |
ICIP | 2 |
| 2020 | Unsupervised Stereo Matching Using Confidential Correspondence ConsistencyabstractStereo matching aims to perceive the 3D geometric configuration of scenes and facilitates a variety of computer vision in advanced driver assistance systems (ADAS) applications. Recently, deep convolutional neural networks (CNNs) have shown dramatic performance improvements for computing the matching cost in the stereo matching. However, the performance of CNN-based approaches relies heavily on datasets, requiring a large number of ground truth data which needs tremendous works. To overcome this limitation, we present a novel framework to learn CNNs for matching cost computation in an unsupervised manner. Our method leverages an image domain learning combined with stereo epipolar constraints. By exploiting the correspondence consistency between stereo images, our method selects putative positive samples in each training iteration and utilizes them to train the networks. We further propose a positive sample propagation scheme to leverage additional training samples. Our unsupervised learning method is evaluated with two kinds of network architectures, simple and precise CNNs, and shows comparable performance to that of the state-of-the-art methods including both supervised and unsupervised learning approaches on KITTI, Middlebury, HCI, and Yonsei datasets. This extensive evaluation demonstrates that the proposed learning framework can be applied to deal with various real driving conditions. Sunghun Joung, Seungryong Kim, Kihong Park, Kwanghoon Sohn |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2019 | Unpaired Cross-Spectral Pedestrian Detection Via Adversarial Feature LearningabstractEven though there exist significant advances in recent studies, existing methods for pedestrian detection still have shown limited performances under challenging illumination conditions especially at nighttime. To address this, cross-spectral pedestrian detection methods have been presented using color and thermal, and shown substantial performance gains on the challenging circumstances. However, their paired cross-spectral settings have limited applicability in real-world scenarios. To overcome this, we propose a novel learning framework for cross-spectral pedestrian detection in an unpaired setting. Based on an assumption that features from color and thermal images share their characteristics in a common feature space to benefit their complement information, we design the separate feature embedding networks for color and thermal images followed by sharing detection networks. To further improve the cross-spectral feature representation, we apply an adversarial learning scheme to intermediate features of the color and thermal images. Experiments demonstrate the outstanding performance of the proposed method on the KAIST multi-spectral benchmark in comparison to the state-of-the-art methods. Sunghun Joung, Kihong Park, Seungryong Kim, Kwanghoon Sohn |
ICIP | 2 |
| 2017 | Unsupervised stereo matching using correspondence consistencyabstractDeep convolutional neural networks (CNNs) have shown revolutionary performance improvements for matching cost computation in stereo matching. However, conventional CNN-based approaches to learn the network in a supervised manner require a large number of ground-truth disparity maps, which limits their applicability. To overcome this limitation, we present a novel framework to learn a CNNs architecture for matching cost computation in an unsupervised manner. Our method leverages an image domain learning combined with stereo epipolar constraints. Exploiting the correspondence consistency between stereo images as supervision, our method selects the training samples in each iteration during network training and uses them to learn the network. To boost the performance, we also propose a multi-scale cost computation scheme. Experimental results show that our method outperforms the state-of-the-art methods including even supervised learning based methods on various benchmarks. Sunghun Joung, Seungryong Kim, Bumsub Ham, Kwanghoon Sohn |
ICIP | 1 |