Hoseok Do

dblp:81/7724 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-4005-1999ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 3D semantic image synthesis with geometric and semantic consistency
Jihyun Kim 0009, Changjae Oh, Hoseok Do, Sunghwan Choi, Kwanghoon Sohn
Expert Syst. Appl.3
2025 PoseSyn: Synthesizing Diverse 3D Pose Data from In-the-Wild 2D Data
ChangHee Yang, Hyeonseop Song, Seokhun Choi, Jaechul Kim, Hoseok Do
ICCV6
2024 Diffusion-Driven GAN Inversion for Multi-Modal Face Image Generation
abstract
We present a new multimodal face image generation method that converts a text prompt and a visual input, such as a semantic mask or scribble map, into a photorealistic face image. To do this, we combine the strengths of Generative Adversarial networks (GANs) and diffusion models (DMs) by employing the multimodal features in the DM into the latent space of the pretrained GANs. We present a simple mapping and a style modulation network to link two models and convert meaningful representations in feature maps and attention maps into latent codes. With GAN inversion, the estimated latent codes can be used to generate 2D or 3D-aware facial images. We further present a multi-step training strategy that reflects textual and structural representations into the generated image. Our proposed network produces realistic 2D, multi-view, and stylized face images, which align well with inputs. We validate our method by using pretrained 2D and 3D GANs, and our results outperform existing methods. Our project page is available at https://github.com/1211sh/Diffusion-driven_GAN-Inversion/.
Jihyun Kim 0009, Changjae Oh, Hoseok Do, Kwanghoon Sohn
CVPR3
2024 Click-Gaussian: Interactive Segmentation to Any 3D Gaussians
Seokhun Choi, Hyeonseop Song, Jaechul Kim, Hoseok Do
ECCV (3)5
2023 Quantitative Manipulation of Custom Attributes on 3D-Aware Image Synthesis
abstract
While 3D-based GAN techniques have been successfully applied to render photo-realistic 3D images with a variety of attributes while preserving view consistency, there has been little research on how to fine-control 3D images without limiting to a specific category of objects of their properties. To fill such research gap, we propose a novel image manipulation model of 3D-based GAN representations for a fine-grained control of specific custom attributes. By extending the latest 3D-based GAN models (e.g., EG3D), our user-friendly quantitative manipulation model enables a fine yet normalized control of 3D manipulation of multi-attribute quantities while achieving view consistency. We validate the effectiveness of our proposed technique both qualitatively and quantitatively through various experiments.
Hoseok Do, Eunkyung Yoo, Chul Lee, Jin Young Choi 0002
CVPR1
2023 Blending-NeRF: Text-Driven Localized Editing in Neural Radiance Fields
abstract
Text-driven localized editing of 3D objects is particularly difficult as locally mixing the original 3D object with the intended new object and style effects without distorting the object’s form is not a straightforward process. To address this issue, we propose a novel NeRF-based model, Blending-NeRF, which consists of two NeRF networks: pre-trained NeRF and editable NeRF. Additionally, we introduce new blending operations that allow Blending-NeRF to properly edit target regions which are localized by text. By using a pretrained vision-language aligned model, CLIP, we guide Blending-NeRF to add new objects with varying colors and densities, modify textures, and remove parts of the original object. Our extensive experiments demonstrate that Blending-NeRF produces naturally and locally edited 3D objects from various text prompts.
Hyeonseop Song, Seokhun Choi, Hoseok Do, Chul Lee
ICCV3
2022 Tracking Failure Prediction for Siamese Trackers Based on Channel Feature Statistics
abstract
Failure prediction has rarely been studied for Siamese trackers due to a lack of meaningful analysis of tracking failing cases. In this paper, we provide a meaningful analysis of tracking failure in Siamese trackers. Our analysis includes the statistics of the channel-wise feature correlation between the exemplar and tracked target patches. We observe that the correlation statistics (max, mean, and std) are highly related to the overlapping ratio between tracked and ground-truth bounding boxes. Based on this observation, we devise a tracking failure prediction model that extracts more plentiful factors than simple statistics. The proposed tracking failure prediction model is validated on most-popular tracking benchmark datasets through extensive experiments.
Kyuewang Lee, Hoseok Do, Taegil Ha, Jongwon Choi 0002, Jin Young Choi 0002
AVSS2
2010 A blind MPEG-2 video watermarking robust to camcorder recording
Dooseop Choi, Hoseok Do, Hyuk Choi, Taejeong Kim
Signal Process.2