Innfarn Yoo

dblp:140/7741 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from a Single-View Image
abstract
Recently, generalizable feed-forward methods based on 3D Gaussian Splatting have gained significant attention for their potential to reconstruct 3D scenes using finite resources. These approaches create a 3D radiance field, parameterized by per-pixel 3D Gaussian primitives, from just a few images in a single forward pass. However, unlike multi-view methods that benefit from cross-view correspondences, 3D scene reconstruction with a single-view image remains an underexplored area. In this work, we introduce CATSplat, a novel generalizable transformer-based framework designed to break through the inherent constraints in monocular settings. First, we propose leveraging textual guidance from a visual-language model to complement insufficient information from a single image. By incorporating scene-specific contextual details from text embeddings through cross-attention, we pave the way for context-aware 3D scene reconstruction beyond relying solely on visual cues. Moreover, we advocate utilizing spatial guidance from 3D point features toward comprehensive geometric understanding under single-view settings. With 3D priors, image features can capture rich structural insights for predicting 3D Gaussians without multi-view techniques. Extensive experiments on large-scale datasets demonstrate the state-of-the-art performance of CATSplat in single-view 3D scene reconstruction with high-quality novel view synthesis.
Wonseok Roh, Hwanhee Jung, Jong Wook Kim, Seunggwan Lee, Innfarn Yoo, Andreas Lugmayr, Seunggeun Chi, Karthik Ramani, Sangpil Kim
ICCV5
2024 Parrot: Pareto-Optimal Multi-reward Reinforcement Learning Framework for Text-to-Image Generation
Yinxiao Li, Junjie Ke, Innfarn Yoo, Han Zhang 0010, Qifei Wang, Fei Deng 0001, Glenn Entis, Junfeng He, Gang Li 0021, Sangpil Kim, Irfan A. Essa, Feng Yang 0008
ECCV (38)4
2024 Optical Diffusion Models for Image Generation
abstract
Diffusion models generate new samples by progressively decreasing the noise from the initially provided random distribution. This inference procedure generally utilizes a trained neural network numerous times to obtain the final output, creating significant latency and energy consumption on digital electronic hardware such as GPUs. In this study, we demonstrate that the propagation of a light beam through a transparent medium can be programmed to implement a denoising diffusion model on image samples. This framework projects noisy image patterns through passive diffractive optical layers, which collectively only transmit the predicted noise term in the image. The optical transparent layers, which are trained with an online training approach, backpropagating the error to the analytical model of the system, are passive and kept the same across different steps of denoising. Hence this method enables high-speed image generation with minimal power consumption, benefiting from the bandwidth and energy efficiency of optical information processing.
Ilker Oguz, Niyazi Ulas Dinç, Mustafa Yildirim, Junjie Ke, Innfarn Yoo, Qifei Wang, Christophe Moser, Demetri Psaltis
NeurIPS5
2024 Fashion-VDM: Video Diffusion Model for Virtual Try-On
Johanna Suvi Karras, Yingwei Li 0001, Nan Liu 0010, Luyang Zhu, Innfarn Yoo, Andreas Lugmayr, Ira Kemelmacher-Shlizerman
SIGGRAPH Asia5
2022 Deep 3D-to-2D Watermarking: Embedding Messages in 3D Meshes and Extracting Them from 2D Renderings
abstract
Digital watermarking is widely used for copyright protection. Traditional 3D watermarking approaches or commercial software are typically designed to embed messages into 3D meshes, and later retrieve the messages directly from distorted/undistorted watermarked 3D meshes. However, in many cases, users only have access to rendered 2D images instead of 3D meshes. Unfortunately, retrieving messages from 2D renderings of 3D meshes is still challenging and underexplored. We introduce a novel end-to-end learning framework to solve this problem through: 1) an encoder to covertly embed messages in both mesh geometry and textures; 2) a differentiable renderer to render watermarked 3D objects from different camera angles and under varied lighting conditions; 3) a decoder to recover the messages from 2D rendered images. From our experiments, we show that our model can learn to embed information visually imperceptible to humans, and to retrieve the embedded information from 2D renderings that undergo 3D distortions. In addition, we demonstrate that our method can also work with other renderers, such as ray tracers and real-time renderers with and without fine-tuning.
Innfarn Yoo, Huiwen Chang, Xiyang Luo, Ondrej Stava, Ce Liu 0001, Peyman Milanfar, Feng Yang 0008
CVPR1
2021 Character motion in function space
Innfarn Yoo, Marek Fiser, Kaimo Hu, Bedrich Benes
Vis. Comput.1
2020 GIFnets: Differentiable GIF Encoding Framework
abstract
Graphics Interchange Format (GIF) is a widely used image file format. Due to the limited number of palette colors, GIF encoding often introduces color banding artifacts. Traditionally, dithering is applied to reduce color banding, but introducing dotted-pattern artifacts. To reduce artifacts and provide a better and more efficient GIF encoding, we introduce a differentiable GIF encoding pipeline, which includes three novel neural networks: PaletteNet, DitherNet, and BandingNet. Each of these three networks provides an important functionality within the GIF encoding pipeline. PaletteNet predicts a near-optimal color palette given an input image. DitherNet manipulates the input image to reduce color banding artifacts and provides an alternative to traditional dithering. Finally, BandingNet is designed to detect color banding, and provides a new perceptual loss specifically for GIF images. As far as we know, this is the first fully differentiable GIF encoding pipeline based on deep neural networks and compatible with existing GIF decoders. User study shows that our algorithm is better than Floyd-Steinberg based GIF encoding.
Innfarn Yoo, Xiyang Luo, Yilin Wang 0001, Feng Yang 0008, Peyman Milanfar
CVPR1
2017 Motion Style Retargeting to Characters With Different Morphologies
abstract
Abstract We present a novel approach for style retargeting to non‐humanoid characters by allowing extracted stylistic features from one character to be added to the motion of another character with a different body morphology. We introduce the concept of groups of body parts (GBPs), for example, the torso, legs and tail, and we argue that they can be used to capture the individual style of a character motion. By separating GBPs from a character, the user can define mappings between characters with different morphologies. We automatically extract the motion of each GBP from the source, map it to the target and then use a constrained optimization to adjust all joints in each GBP in the target to preserve the original motion while expressing the style of the source. We show results on characters that present different morphologies to the source motion from which the style is extracted. The style transfer is intuitive and provides a high level of control. For most of the examples in this paper, the definition of GBP takes around 5 min and the optimization about 7 min on average. For the most complicated examples, the definition of three GBPs and their mapping takes about 10 min and the optimization another 30 min.
Michel Abdul-Massih, Innfarn Yoo, Bedrich Benes
Comput. Graph. Forum2
2015 Motion retiming by using bilateral time control surfaces
Innfarn Yoo, Michel Abdul-Massih, Illia Ziamtsov, Raymond Hassan, Bedrich Benes
Comput. Graph.1
2014 A hybrid level-of-detail representation for large-scale urban scenes rendering
abstract
ABSTRACT A novel hybrid level‐of‐detail (LOD) algorithm is introduced. We combine point‐based, line‐based, and splat‐based rendering to synthesize large‐scale urban city images. We first extract lines and points from the input and provide their simplification encoded in a data structure that allows for a quick and automatic LOD selection. A screen‐space projected area is used as the LOD selector. The algorithm selects lines for long‐distance views providing high contrast and fidelity of the building silhouettes. For medium‐distance views, points are added, and splats are used for close‐up views. Our implementation shows a 10 × speedup as compared with the ground truth models and is about four times faster than geometric LOD. The quality of the results is indistinguishable from the original as confirmed by a user study and two algorithmic metrics. Copyright © 2014 John Wiley & Sons, Ltd.
Shengchuan Zhou, Innfarn Yoo, Bedrich Benes, Ge Chen 0002
Comput. Animat. Virtual Worlds2
2014 Sketching human character animations by composing sequences from large motion database
Innfarn Yoo, Juraj Vanek, Maria Nizovtseva, Nicoletta Adamo-Villani, Bedrich Benes
Vis. Comput.1