VLDB 2026 Research / reviewers in the wild / expert
Woobin Im
dblp:192/1851
· DBLP profile ↗
13ranked-venue papers
4as first author
8since 2021 · last 2024
0000-0002-9269-4833ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SemCity: Semantic Scene Generation with Triplane DiffusionabstractWe present “SemCity,” a 3D diffusion model for semantic scene generation in real-world outdoor environments. Most 3D diffusion models focus on generating a single object, synthetic indoor scenes, or synthetic outdoor scenes, while the generation of real-world outdoor scenes is rarely addressed. In this paper, we concentrate on generating a real-outdoor scene through learning a diffusion model on a realworld outdoor dataset. In contrast to synthetic data, real-outdoor datasets often contain more empty spaces due to sensor limitations, causing challenges in learning realoutdoor distributions. To address this issue, we exploit a triplane representation as a proxy form of scene distributions to be learned by our diffusion model. Furthermore, we propose a triplane manipulation that integrates seamlessly with our triplane diffusion model. The manipulation improves our diffusion model's applicability in a variety of downstream tasks related to outdoor scene generation such as scene inpainting, scene outpainting, and semantic scene completion refinements. In experimental results, we demonstrate that our triplane diffusion model shows meaningful generation results compared with existing work in a real-outdoor dataset, SemanticKITTI. We also show our triplane manipulation facilitates seamlessly adding, removing, or modifying objects within a scene. Further, it also enables the expansion of scenes toward a city-level scale. Finally, we evaluate our method on semantic scene completion refinements where our diffusion model enhances predictions of semantic scene completion networks by learning scene distribution. Our code is available at https://github.com/zoom.in-lee/SemCity. Jumin Lee, Sebin Lee, Changho Jo, Woobin Im, Juhyeong Seon, Sung-Eui Yoon |
CVPR | 4 |
| 2024 | Regularizing Dynamic Radiance Fields with Kinematic Fields
Woobin Im, Geonho Cha, Sebin Lee, Jumin Lee, Juhyeong Seon, Dongyoon Wee, Sung-Eui Yoon |
ECCV (39) | 1 |
| 2024 | Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual SegmentationabstractAudio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audiovisual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense prediction problems, prior works have investigated the introduction of SAM into AVS with audio as a new modality of the prompt. Nevertheless, constrained by SAM’s single-frame segmentation scheme, the temporal context across multiple frames of audio-visual data remains insufficiently utilized. To this end, we study the extension of SAM’s capabilities to the sequence of audio-visual scenes by analyzing contextual cross-modal relationships across the frames. To achieve this, we propose a Spatio-Temporal, Bidirectional Audio-Visual Attention (ST-BAVA) module integrated into the middle of SAM’s image encoder and mask decoder. It adaptively updates the audio-visual features to convey the spatio-temporal correspondence between the video frames and audio streams. Extensive experiments demonstrate that our proposed model outperforms the state-of-the-art methods on AVS benchmarks, especially with an 8.3% mIoU gain on a challenging multi-sources subset. Juhyeong Seon, Woobin Im, Sebin Lee, Jumin Lee, Sung-Eui Yoon |
ICIP | 2 |
| 2024 | Fine-Grained Background Representation for Weakly Supervised Semantic SegmentationabstractGenerating reliable pseudo masks from image-level labels is challenging in the weakly supervised semantic segmentation (WSSS) task due to the lack of spatial information. Prevalent class activation map (CAM)-based solutions are challenged to discriminate the foreground (FG) objects from the suspicious background (BG) pixels (a.k.a. co-occurring) and learn the integral object regions. This paper proposes a simple fine-grained background representation (FBR) method to discover and represent diverse BG semantics and address the co-occurring problems. We abandon using the class prototype or pixel-level features for BG representation. Instead, we develop a novel primitive, negative region of interest (NROI), to capture the fine-grained BG semantic information and conduct the pixel-to-NROI contrast to distinguish the confusing BG pixels. We also present an active sampling strategy to mine the FG negatives on-the-fly, enabling efficient pixel-to-pixel intra-foreground contrastive learning to activate the entire object region. Thanks to the simplicity of design and convenience in use, our proposed method can be seamlessly plugged into various models, yielding new state-of-the-art results under various WSSS settings across benchmarks. Leveraging solely image-level (I) labels as supervision, our method achieves 73.2 mIoU and 45.6 mIoU segmentation results on Pascal Voc and MS COCO test sets, respectively. Furthermore, by incorporating saliency maps as an additional supervision signal (I+S), we attain 74.9 mIoU on Pascal Voc test set. Concurrently, our FBR approach demonstrates meaningful performance gains in weakly-supervised instance segmentation (WSIS) tasks, showcasing its robustness and strong generalization capabilities across diverse domains. Xu Yin, Woobin Im, Dongbo Min, Yuchi Huo, Sung-Eui Yoon |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Multi-resolution distillation for self-supervised monocular depth estimation
Sebin Lee, Woobin Im, Sung-Eui Yoon |
Pattern Recognit. Lett. | 2 |
| 2022 | Semi-supervised Learning of Optical Flow by Flow Supervisor
Woobin Im, Sebin Lee, Sung-Eui Yoon |
ECCV (35) | 1 |
| 2021 | In-N-Out: Towards Good Initialization for Inpainting and Outpainting
Changho Jo, Woobin Im, Sung-Eui Yoon |
BMVC | 2 |
| 2021 | Combined center dispersion loss function for deep facial expression recognition
Abhilasha Nanda, Woobin Im, Key-Sun Choi, Hyun Seung Yang |
Pattern Recognit. Lett. | 2 |
| 2020 | Unsupervised Learning of Optical Flow with Deep Feature Similarity
Woobin Im, Tae-Kyun Kim 0001, Sung-Eui Yoon |
ECCV (24) | 1 |
| 2018 | Scale-Varying Triplet Ranking with Classification Loss for Facial Age Estimation
Woobin Im, Sungeun Hong, Sung-Eui Yoon, Hyun S. Yang |
ACCV (5) | 1 |
| 2018 | CBVMR: Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure ConstraintabstractUp to now, only limited research has been conducted on crossmodal retrieval of suitable music for a specified video or vice versa. Moreover, much of the existing research relies on metadata such as keywords, tags, or description that must be individually produced and attached posterior. This paper introduces a new content-based, cross-modal retrieval method for video and music that is implemented through deep neural networks. We train the network via inter-modal ranking loss such that videos and music with similar semantics end up close together in the embedding space. However, if only the inter-modal ranking constraint is used for embedding, modality-specific characteristics can be lost. To address this problem, we propose a novel soft intra-modal structure loss that leverages the relative distance relationship between intra-modal samples before embedding. We also introduce reasonable quantitative and qualitative experimental protocols to solve the lack of standard protocols for less-mature video-music related tasks. All the datasets and source code can be found in our online repository (https://github.com/csehong/VM-NET). Sungeun Hong, Woobin Im, Hyun Seung Yang |
ICMR | 2 |
| 2018 | D3: Recognizing dynamic scenes with deep dual descriptor based on key frames and key segments
Sungeun Hong, Jong Bin Ryu, Woobin Im, Hyun Seung Yang |
Neurocomputing | 3 |
| 2017 | SSPP-DAN: Deep domain adaptation network for face recognition with single sample per personabstractReal-world face recognition using a single sample per person (SSPP) is a challenging task. The problem is exacerbated if the conditions under which the gallery image and the probe set are captured are completely different. To address these issues from the perspective of domain adaptation, we introduce an SSPP domain adaptation network (SSPP-DAN). In the proposed approach, domain adaptation, feature extraction, and classification are performed jointly using a deep architecture with domain-adversarial training. However, the SSPP characteristic of one training sample per class is insufficient to train the deep architecture. To overcome this shortage, we generate synthetic images with varying poses using a 3D face model. Experimental evaluations using a realistic SSPP dataset show that deep domain adaptation and image synthesis complement each other and dramatically improve accuracy. Experiments on a benchmark dataset using the proposed approach show state-of-the-art performance. Sungeun Hong, Woobin Im, Jong Bin Ryu, Hyun Seung Yang |
ICIP | 2 |