Seonghyeon Nam

dblp:191/4635 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Exploring How Prior Knowledge and Presence Shape Transfer of a Reversed Size-Weight Illusion From Virtual to Real
abstract
Virtual Reality (VR) can create experiences that conflict with a user’s prior knowledge; however, how such conflicts influence subsequent real-world behavior remains unclear. This study explores how a virtual experience that contradicts real-world expectations affects later perception and motor actions, using the size-weight illusion—where people expect larger objects to be heavier than smaller ones. We conducted a 2 (internal model robustness: reinforced vs. weakened) by 2 (presence: high vs. low) mixed-design experiment. Participants first received real-world training to either strengthen or weaken their size-weight expectations, then experienced a reversed size-weight mapping in VR under varying levels of presence. We assessed how this virtual experience influenced real-world weight estimation and object lifting behavior. Results showed that participants with weakened prior knowledge exhibited lower confidence in their weight judgments and greater motor instability when lifting objects. However, the level of presence in VR did not significantly affect transfer outcomes. These findings suggest that the strength of prior knowledge modulates how conflicting virtual experiences influence real-world behavior, underscoring the need for careful VR design, particularly for younger users with less stable internal models.
Seonghyeon Nam, Bogwan Kim, Deokyong Kim, Min-Ho Seo, Myungho Lee
VRST1
2024 Geometry Transfer for Stylizing Radiance Fields
abstract
Shape and geometric patterns are essential in defining stylistic identity. However, current 3D style transfer methods predominantly focus on transferring colors and textures, often overlooking geometric aspects. In this paper, we introduce Geometry Transfer, a novel method that leverages geometric deformation for 3D style transfer. This technique employs depth maps to extract a style guide, subsequently applied to stylize the geometry of radiance fields. Moreover, we propose new techniques that utilize geometric cues from the 3D scene, thereby enhancing aesthetic expressiveness and more accurately reflecting intended styles. Our extensive experiments show that Geometry Transfer enables a broader and more expressive range of stylizations, thereby significantly expanding the scope of 3D style transfer.
Hyunyoung Jung 0001, Seonghyeon Nam, Nikolaos Sarafianos, Sungjoo Yoo, Alexander Sorkine-Hornung
CVPR2
2023 Learning Neural Duplex Radiance Fields for Real-Time View Synthesis
abstract
Neural radiance fields (NeRFs) enable novel-view synthesis with unprecedented visual quality. However, to render photorealistic images, NeRFs require hundreds of deep multilayer perceptron (MLP) evaluations - for each pixel. This is prohibitively expensive and makes realtime rendering infeasible, even on powerful modern GPUs. In this paper, we propose a novel approach to distill and bake NeRFs into highly efficient mesh-based neural representations that are fully compatible with the massively parallel graphics rendering pipeline. We represent scenes as neural radiance features encoded on a two-layer duplex mesh, which effectively over-comes the inherent inaccuracies in 3D surface reconstruction by learning the aggregated radiance information from a reliable interval of ray-surface intersections. To exploit local geometric relationships of nearby pixels, we leverage screen-space convolutions instead of the MLPs used in NeRFs to achieve high-quality appearance. Finally, the performance of the whole framework is further boosted by a novel multi-view distillation optimization strategy. We demonstrate the effectiveness and superiority of our approach via extensive experiments on a range of standard datasets.
Ziyu Wan, Christian Richardt, Aljaz Bozic, Vijay Rengarajan, Seonghyeon Nam, Xiaoyu Xiang, Tuotuo Li, Bo Zhu 0011, Jing Liao 0001
CVPR6
2022 Learning sRGB-to-Raw-RGB De-rendering with Content-Aware Metadata
abstract
Most camera images are rendered and saved in the standard RGB (sRGB) format by the camera's hardware. Due to the in-camera photo-finishing routines, nonlinear sRGB images are undesirable for computer vision tasks that assume a direct relationship between pixel values and scene radiance. For such applications, linear raw-RGB sensor images are preferred. Saving images in their raw-RGB format is still uncommon due to the large storage requirement and lack of support by many imaging applications. Several “raw reconstruction” methods have been proposed that utilize specialized metadata sampled from the raw-RGB image at capture time and embedded in the sRGB image. This metadata is used to parameterize a mapping function to derender the sRGB image back to its original raw-RGB format when needed. Existing raw reconstruction methods rely on simple sampling strategies and global mapping to perform the de-rendering. This paper shows how to improve the derendering results by jointly learning sampling and reconstruction. Our experiments show that our learned sampling can adapt to the image content to produce better raw reconstructions than existing methods. We also describe an online fine-tuning strategy for the reconstruction network to improve results further.
Seonghyeon Nam, Abhijith Punnappurath, Marcus A. Brubaker, Michael S. Brown
CVPR1
2022 Neural Image Representations for Multi-image Fusion and Layer Separation
Seonghyeon Nam, Marcus A. Brubaker, Michael S. Brown
ECCV (7)1
2022 Dense Interspecies Face Embedding
abstract
Dense Interspecies Face Embedding (DIFE) is a new direction for understanding faces of various animals by extracting common features among animal faces including human face. There are three main obstacles for interspecies face understanding: (1) lack of animal data compared to human, (2) ambiguous connection between faces of various animals, and (3) extreme shape and style variance. To cope with the lack of data, we utilize multi-teacher knowledge distillation of CSE and StyleGAN2 requiring no additional data or label. Then we synthesize pseudo pair images through the latent space exploration of StyleGAN2 to find implicit associations between different animal faces. Finally, we introduce the semantic matching loss to overcome the problem of extreme shape differences between species. To quantitatively evaluate our method over possible previous methodologies like unsupervised keypoint detection, we perform interspecies facial keypoint transfer on MAFL and AP-10K. Furthermore, the results of other applications like interspecies face image manipulation and dense keypoint transfer are provided. The code is available at https://github.com/kingsj0405/dife.
Sejong Yang, Subin Jeon, Seonghyeon Nam, Seon Joo Kim
NeurIPS3
2022 2PESNet: Towards online processing of temporal action localization
Young Hwi Kim, Seonghyeon Nam, Seon Joo Kim
Pattern Recognit.2
2021 Large Scale Multi-Illuminant (LSMI) Dataset for Developing White Balance Algorithm under Mixed Illumination
abstract
We introduce a Large Scale Multi-Illuminant (LSMI) Dataset that contains 7,486 images, captured with three different cameras on more than 2,700 scenes with two or three illuminants. For each image in the dataset, the new dataset provides not only the pixel-wise ground truth illumination but also the chromaticity of each illuminant in the scene and the mixture ratio of illuminants per pixel. Images in our dataset are mostly captured with illuminants existing in the scene, and the ground truth illumination is computed by taking the difference between the images with different illumination combination. Therefore, our dataset captures natural composition in the real-world setting with wide field-of-view, providing more extensive dataset compared to existing datasets for multi-illumination white balance. As conventional single illuminant white balance algorithms cannot be directly applied, we also apply per-pixel DNN-based white balance algorithm and show its effectiveness against using patch-wise white balancing. We validate the benefits of our dataset through extensive analysis including a user-study, and expect the dataset to make meaningful contribution for future work in white balancing.
Dongyoung Kim, Jinwoo Kim 0007, Seonghyeon Nam, Yeonkyung Lee, Nahyup Kang, Hyong-Euk Lee, ByungIn Yoo, Jae-Joon Han, Seon Joo Kim
ICCV3
2021 Temporally smooth online action detection using cycle-consistent future anticipation
Young Hwi Kim, Seonghyeon Nam, Seon Joo Kim
Pattern Recognit.2
2020 Cross-Identity Motion Transfer for Arbitrary Objects Through Pose-Attentive Video Reassembling
Subin Jeon, Seonghyeon Nam, Seoung Wug Oh, Seon Joo Kim
ECCV (24)2
2019 End-To-End Time-Lapse Video Synthesis From a Single Outdoor Image
abstract
Time-lapse videos usually contain visually appealing content but are often difficult and costly to create. In this paper, we present an end-to-end solution to synthesize a time-lapse video from a single outdoor image using deep neural networks. Our key idea is to train a conditional generative adversarial network based on existing datasets of time-lapse videos and image sequences. We propose a multi-frame joint conditional generation framework to effectively learn the correlation between the illumination change of an outdoor scene and the time of the day. We further present a multi-domain training scheme for robust training of our generative models from two datasets with different distributions and missing timestamp labels. Compared to alternative time-lapse video synthesis algorithms, our method uses the timestamp as the control variable and does not require a reference video to guide the synthesis of the final output. We conduct ablation studies to validate our algorithm and compare with state-of-the-art techniques both qualitatively and quantitatively.
Seonghyeon Nam, Chongyang Ma, Menglei Chai, William Brendel, Ning Xu 0007, Seon Joo Kim
CVPR1
2019 Unsupervised Keypoint Learning for Guiding Class-Conditional Video Prediction
abstract
We propose a deep video prediction model conditioned on a single image and an action class. To generate future frames, we first detect keypoints of a moving object and predict future motion as a sequence of keypoints. The input image is then translated following the predicted keypoints sequence to compose future frames. Detecting the keypoints is central to our algorithm, and our method is trained to detect the keypoints of arbitrary objects in an unsupervised manner. Moreover, the detected keypoints of the original videos are used as pseudo-labels to learn the motion of objects. Experimental results show that our method is successfully applied to various datasets without the cost of labeling keypoints in videos. The detected keypoints are similar to human-annotated labels, and prediction results are more realistic compared to the previous methods.
Yunji Kim, Seonghyeon Nam, In Cho, Seon Joo Kim
NeurIPS2
2018 Text-Adaptive Generative Adversarial Networks: Manipulating Images with Natural Language
abstract
This paper addresses the problem of manipulating images using natural language description. Our task aims to semantically modify visual attributes of an object in an image according to the text describing the new visual appearance. Although existing methods synthesize images having new attributes, they do not fully preserve text-irrelevant contents of the original image. In this paper, we propose the text-adaptive generative adversarial network (TAGAN) to generate semantically manipulated images while preserving text-irrelevant contents. The key to our method is the text-adaptive discriminator that creates word level local discriminators according to input text to classify fine-grained attributes independently. With this discriminator, the generator learns to generate images where only regions that correspond to the given text is modified. Experimental results show that our method outperforms existing methods on CUB and Oxford-102 datasets, and our results were mostly preferred on a user study. Extensive analysis shows that our method is able to effectively disentangle visual attributes and produce pleasing outputs.
Seonghyeon Nam, Yunji Kim, Seon Joo Kim
NeurIPS1
2017 Modelling the Scene Dependent Imaging in Cameras with a Deep Neural Network
abstract
We present a novel deep learning framework that models the scene dependent image processing inside cameras. Often called as the radiometric calibration, the process of recovering RAWimages from processed images (JPEG format in the sRGB color space) is essential for many computer vision tasks that rely on physically accurate radiance values. All previous works rely on the deterministic imaging model where the color transformation stays the same regardless of the scene and thus they can only be applied for images taken under the manual mode. In this paper, we propose a datadriven approach to learn the scene dependent and locally varying image processing inside cameras under the automode. Our method incorporates both the global and the local scene context into pixel-wise features via multi-scale pyramid of learnable histogram layers. The results show that we can model the imaging pipeline of different cameras that operate under the automode accurately in both directions (from RAW to sRGB, from sRGB to RAW) and we show how we can apply our method to improve the performance of image deblurring.
Seonghyeon Nam, Seon Joo Kim
ICCV1
2016 A Holistic Approach to Cross-Channel Image Noise Modeling and Its Application to Image Denoising
abstract
Modelling and analyzing noise in images is a fundamental task in many computer vision systems. Traditionally, noise has been modelled per color channel assuming that the color channels are independent. Although the color channels can be considered as mutually independent in camera RAW images, signals from different color channels get mixed during the imaging process inside the camera due to gamut mapping, tone-mapping, and compression. We show the influence of the in-camera imaging pipeline on noise and propose a new noise model in the 3D RGB space to accounts for the color channel mix-ups. A data-driven approach for determining the parameters of the new noise model is introduced as well as its application to image denoising. The experiments show that our noise model represents the noise in regular JPEG images more accurately compared to the previous models and is advantageous in image denoising.
Seonghyeon Nam, Youngbae Hwang, Yasuyuki Matsushita, Seon Joo Kim
CVPR1