VLDB 2026 Research / reviewers in the wild / expert
Tae-Hyun Oh
dblp:119/1450 · also Tae Hyun Oh
· DBLP profile ↗
105ranked-venue papers
12as first author
63since 2021 · last 2026
0000-0003-0468-1571ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 75 · 10 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 73 · 7 first-author · 43 since 2021Systems, architecture and hardware · 5 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about LaughterabstractLaughter is a complex social signal that conveys communicative intent beyond amusement.While prior work has focused on isolated laughter analysis tasks, a comprehensive understanding of laughter in real-world scenarios remains underexplored.We introduce SMILE-Next, a dataset for real-world laughter understanding with multimodal textual representations and question-answer annotations across three tasks: laughter detection, laughter type classification, and laughter reasoning.Building upon SMILE-Next, we aim to develop a laughterspecialized large language model capable of nuanced understanding of laughter in real-world contexts.To this end, we propose two key components: laughter-specific Self-Instruct and the Mixture-of-Laugh-Experts (MoLE) framework.Laughter-specific Self-Instruct enhances generalization across tasks and domains by automatically synthesizing diverse laughtercentric instructions.MoLE introduces a taskadaptive expert routing mechanism that dynamically selects specialized experts tailored to each laughter-related task, improving taskspecific performance and efficiency.Experimental results show that the combination of our proposed components substantially outperforms multimodal LLM baselines, advancing robust real-world laughter understanding.Project page is at JungMok Lee, Sung-Bin Kim, Joohyun Chang, Lee Hyun 0001, Tae-Hyun Oh |
ACL (1) | 5 |
| 2026 | Beyond the Highlights: Video Retrieval with Salient and Surrounding Contexts
Jae Hun Bang, Moon Ye-Bin, Tae-Hyun Oh, Kyungdon Joo |
WACV | 3 |
| 2026 | Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level MatchingabstractInstance-level image retrieval aims to find images containing the same object as a given query, despite variations in size, position, or appearance. To address this challenging task, we propose Patchify, a simple yet effective patch-wise retrieval framework that offers high performance, scalability, and interpretability without requiring fine-tuning. Patchify divides each database image into a small number of structured patches and performs retrieval by comparing these local features with a global query descriptor, enabling accurate and spatially grounded matching. To assess not just retrieval accuracy but also spatial correctness, we introduce LocScore, a localization-aware metric that quantifies whether the retrieved region aligns with the target object. This makes LocScore a valuable diagnostic tool for understanding and improving retrieval behavior. We conduct extensive experiments across multiple benchmarks, backbones, and region selection strategies, showing that Patchify outperforms global methods and complements state-of-the-art reranking pipelines. Furthermore, we apply Product Quantization for efficient large-scale retrieval and highlight the importance of using informative features during compression, which significantly boosts performance. Won-Seok Choi 0001, Sohwi Lim, Nam Hyeon-Woo, Moon Ye-Bin, Dong-Ju Jeong, Jinyoung Hwang, Tae-Hyun Oh |
WACV | 7 |
| 2026 | mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image RetrievalabstractScalable Vector Graphics (SVGs) function both as visual images and as structured code that encode rich geometric and layout information, yet most methods rasterize them and discard this symbolic organization. At the same time, recent sentence embedding methods produce strong text representations but do not naturally extend to visual or structured modalities. We propose a training-free, instruction-guided multimodal embedding framework that uses a Multimodal Large Language Model (MLLM) to map text, raster images, and SVG code into an aligned embedding space. We control the direction of embeddings through modality-specific instructions and structural SVG cues, eliminating the need for learned projection heads or contrastive training. Our method has two key components: (1) Multimodal Explicit One-word Limitation (mEOL), which instructs the MLLM to summarize any multimodal input into a single token whose hidden state serves as a compact semantic embedding. (2) A semantic SVG rewriting module that assigns meaningful identifiers and simplifies nested SVG elements through visual reasoning over the rendered image, exposing geometric and relational cues hidden in raw code. Using a repurposed VGBench, we build the first text-to-SVG retrieval benchmark and show that our training-free embeddings outperform encoder-based and training-based multimodal baselines. These results highlight prompt-level control as an effective alternative to parameter-level training for structure-aware multimodal retrieval. Project page: https://scene-the-ella.github.io/meol/ Kyeong Seon Kim, Seong-Eun Baek, JungMok Lee, Tae-Hyun Oh |
WACV | 4 |
| 2026 | FPGS: Feed-Forward Semantic-aware Photorealistic Style Transfer of Large-Scale Gaussian SplattingabstractAbstract We present FPGS, a feed-forward photorealistic style transfer method of large-scale radiance fields represented by Gaussian Splatting. FPGS stylizes large-scale 3D scenes with arbitrary, multiple style reference images without additional optimization while preserving multi-view consistency and real-time rendering speed of 3D Gaussians. Prior arts required tedious per-style optimization or time-consuming per-scene training stage and were limited to small-scale 3D scenes. FPGS efficiently stylizes large-scale 3D scenes by introducing a style-decomposed 3D feature field, which inherits AdaIN’s feed-forward stylization machinery, supporting arbitrary style reference images. Furthermore, FPGS supports multi-reference stylization with the semantic correspondence matching and local AdaIN, which adds diverse user control for 3D scene styles. FPGS also preserves multi-view consistency by applying semantic matching and style transfer processes directly onto queried features in 3D space. In experiments, we demonstrate that FPGS achieves favorable photorealistic quality scene stylization for large-scale static and dynamic 3D scenes with diverse reference images. GeonU Kim, Kim Youwang, Lee Hyoseok, Tae-Hyun Oh |
Int. J. Comput. Vis. | 4 |
| 2026 | Audio-Visual Generation
Tae-Hyun Oh, Vicky Kalogeiton, Stavros Petridis, Sergey Tulyakov, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 1 |
| 2026 | CLIP-Actor-X: Text-Driven 4D Human Avatar Generation via Cross-Modal Synthesis-Through-OptimizationabstractWe propose CLIP-Actor-X, a text-driven motion generation and neural mesh stylization system for 4D human avatar generation. CLIP-Actor-X generates a detailed 3D human mesh, motion animation, and texture to conform to a given text prompt input from a user. CLIP- Actor-X system mainly consists of two modules. First, for generating realistic human motion, we build a text-driven human motion synthesis module modeled by a retrieval-augmented generative model, powered by a text-to-motion diffusion model. Second, our novel zero-shot neural style optimization module detailizes and texturizes the sampled sequence of a neutral human mesh template, such that the resulting mesh and appearance comply with the input text prompt in a temporally-consistent and pose-agnostic manner. In contrast to the prior arts that use an artist-designed, non-animatable mesh as an input, our output representation is animatable and better aligned between an input text and the generated avatar without additional post-processes, e.g., re-alignment, retargeting, or rigging. We further propose the ways to stabilize the optimization process: spatio-temporal view augmentation and visibility-aware embedding attention, which deals with poorly rendered views. We demonstrate that CLIP-Actor-X produces perceptually plausible and human-recognizable human avatar in motion with detailed geometry and texture solely from a natural language prompt. Kim Youwang, Taehyun Byun, Kim Ji-Yeon, Tae-Hyun Oh |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Zero-shot Depth Completion via Test-time Alignment with Affine-invariant Depth PriorabstractDepth completion, predicting dense depth maps from sparse depth measurements, is an ill-posed problem requiring prior knowledge. Recent methods adopt learning-based approaches to implicitly capture priors, but the priors primarily fit in-domain data and do not generalize well to out-of-domain scenarios. To address this, we propose a zero-shot depth completion method composed of an affine-invariant depth diffusion model and test-time alignment. We use pre-trained depth diffusion models as depth prior knowledge, which implicitly understand how to fill in depth for scenes. Our approach aligns the affine-invariant depth prior with metric-scale sparse measurements, enforcing them as hard constraints via an optimization loop at test-time. Our zero-shot depth completion method demonstrates generalization across various domain datasets, achieving up to a 21% average performance improvement over the previous state-of-the-art methods while enhancing spatial understanding by sharpening scene details. We demonstrate that aligning a monocular affine-invariant depth prior with sparse metric measurements is a sufficient strategy to achieve domain-generalizable depth completion without relying on extensive training datasets. Lee Hyoseok, Kyeong Seon Kim, Byung-Ki Kwon, Tae-Hyun Oh |
AAAI | 4 |
| 2025 | SoundBrush: Sound as a Brush for Visual Scene EditingabstractWe propose SoundBrush, a model that uses sound as a brush to edit and manipulate visual scenes. We extend the generative capabilities of the Latent Diffusion Model (LDM) to incorporate audio information for editing visual scenes. Inspired by existing image-editing works, we frame this task as a supervised learning problem and leverage various off-the-shelf models to construct a sound-paired visual scene editing dataset for training. This richly generated dataset enables SoundBrush to learn to map audio features into the textual space of the LDM, allowing for visual scene editing guided by diverse in-the-wild sound. Unlike existing methods, SoundBrush can accurately manipulate the overall scenery or even insert sounding objects to best match the input sound semantics while preserving the original content. Furthermore, by integrating with novel view synthesis techniques, our framework can be extended to edit 3D scenes, facilitating sound-driven 3D scene manipulation. Sung-Bin Kim, Kim Jun-Seong, Junseok Ko, Tae-Hyun Oh |
AAAI | 5 |
| 2025 | Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation MetricsabstractRecent advancements in speech-driven 3D talking head generation have made significant progress in lip synchronization. However, existing models still struggle to capture the perceptual alignment between varying speech characteristics and corresponding lip movements. In this work, we claim that three criteria—Temporal Synchronization, Lip Readability, and Expressiveness—are crucial for achieving perceptually accurate lip movements. Motivated by our hypothesis that a desirable representation space exists to meet these three criteria, we introduce a speech-mesh synchronized representation that captures intricate correspondences between speech signals and 3D face meshes. We found that our learned representation exhibits desirable characteristics, and we plug it into existing models as a perceptual loss to better align lip movements to the given speech. In addition, we utilize this representation as a perceptual metric and introduce two other physically grounded lip synchronization metrics to assess how well the generated 3D talking heads align with these three criteria. Experiments show that training 3D talking head generation models with our perceptual loss significantly improve all three aspects of perceptually accurate lip synchronization. Codes and datasets are available at https://perceptual-3d-talking-head.github.io/. Lee Chae-Yeon, Oh Hyun-Bin, Han EunGi, Sung-Bin Kim, Suekyeong Nam, Tae-Hyun Oh |
CVPR | 6 |
| 2025 | Robust 3D Shape Reconstruction in Zero-Shot from a Single Image in the WildabstractRecent monocular 3D shape reconstruction methods have shown promising zero-shot results on object-segmented images without any occlusions. However, their effectiveness is significantly compromised in real-world conditions, due to imperfect object segmentation by off-the-shelf models and the prevalence of occlusions. To effectively address these issues, we propose a unified regression model that integrates segmentation and reconstruction, specifically designed for occlusion-aware 3D shape reconstruction. To facilitate its reconstruction in the wild, we also introduce a scalable data synthesis pipeline that simulates a wide range of variations in objects, occluders, and backgrounds. Training on our synthetic data enables the proposed model to achieve state-of-the-art zero-shot results on real-world images, using significantly fewer parameters than competing approaches. Junhyeong Cho, Kim Youwang, Hunmin Yang, Tae-Hyun Oh |
CVPR | 4 |
| 2025 | Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding RegistrationabstractWe introduce Dr. Splat, a novel approach for open-vocabulary 3D scene understanding leveraging 3D Gaussian Splatting. Unlike existing language-embedded 3DGS methods, which rely on a rendering process, our method directly associates language-aligned CLIP embeddings with 3D Gaussians for holistic 3D scene understanding. The key of our method is a language feature registration technique where CLIP embeddings are assigned to the dominant Gaussians intersected by each pixel-ray. Moreover, we integrate Product Quantization (PQ) trained on general large-scale image data to compactly represent embeddings without per-scene optimization. Experiments demonstrate that our approach significantly outperforms existing approaches in 3D perception benchmarks, such as openvocabulary 3D semantic segmentation, 3D object localization, and 3D object selection tasks. For video results, please visit : https://drsplat.github.io/ Kim Jun-Seong, GeonU Kim, Kim Yu-Ji, Yu-Chiang Frank Wang, Jaesung Choe, Tae-Hyun Oh |
CVPR | 6 |
| 2025 | DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow DecodingabstractHuman motion is inherently continuous and dynamic, posing significant challenges for generative models. While discrete generation methods are widely used, they suffer from limited expressiveness and frame-wise noise artifacts. In contrast, continuous approaches produce smoother, more natural motion but often struggle to adhere to conditioning signals due to high-dimensional complexity and limited training data. To resolve this 'discord' between discrete and continuous representations we introduce DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding, a novel method that leverages rectified flow to decode discrete motion tokens in the continuous, raw motion space. Our core idea is to frame token decoding as a conditional generation task, ensuring that DisCoRD captures fine-grained dynamics and achieves smoother, more natural motions. Compatible with any discrete-based framework, our method enhances naturalness without compromising faithfulness to the conditioning signals on diverse settings. Extensive evaluations demonstrate that DisCoRD achieves state-of-the-art performance, with FID of 0.032 on HumanML3D and 0.169 on KIT-ML. These results establish DisCoRD as a robust solution for bridging the divide between discrete efficiency and continuous realism. Project website: https://whwjdqls.github.io/discord-motion/ Jungbin Cho, Junwan Kim, Jisoo Kim 0006, Mingu Kang, Sungeun Hong, Tae-Hyun Oh, Youngjae Yu |
ICCV | 7 |
| 2025 | VSC: Visual Search Compositional Text-to-Image Diffusion ModelabstractText-to-image diffusion models have shown impressive capabilities in generating realistic visuals from natural-language prompts, yet they often struggle with accurately binding attributes to corresponding objects, especially in prompts containing multiple attribute-object pairs. This challenge primarily arises from the limitations of commonly used text encoders, such as CLIP, which can fail to encode complex linguistic relationships and modifiers effectively. Existing approaches have attempted to mitigate these issues through attention map control during inference and the use of layout information or fine-tuning during training, yet they face performance drops with increased prompt complexity. In this work, we introduce a novel compositional generation method that leverages pairwise image embeddings to improve attribute-object binding. Our approach decomposes complex prompts into sub-prompts, generates corresponding images, and computes visual prototypes that fuse with text embeddings to enhance representation. By applying segmentation-based localization training, we address cross-attention misalignment, achieving improved accuracy in binding multiple attributes to objects. Our approaches outperform existing compositional text-to-image diffusion models on the benchmark T2I CompBench, achieving better image quality, evaluated by humans, and emerging robustness under scaling number of binding pairs in the prompt. Do Huu Dat, Nam Hyeon-Woo, Po Yuan Mao, Tae-Hyun Oh |
ICCV | 4 |
| 2025 | VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language ModelsabstractWe present VoiceCraft-Dub, a novel approach for automated video dubbing that synthesizes high-quality speech from text and facial cues. This task has broad applications in filmmaking, multimedia creation, and assisting voice-impaired individuals. Building on the success of Neural Codec Language Models (NCLMs) for speech synthesis, our method extends their capabilities by incorporating video features, ensuring that synthesized speech is time-synchronized and expressively aligned with facial movements while preserving natural prosody. To inject visual cues, we design adapters to align facial features with the NCLM token space and introduce audio-visual fusion layers to merge audio-visual information within the NCLM framework. Additionally, we curate CelebV-Dub, a new dataset of expressive, real-world videos specifically designed for automated video dubbing. Extensive experiments show that our model achieves high-quality, intelligible, and natural speech synthesis with accurate lip synchronization, outperforming existing methods in human perception and performing favorably in objective evaluations. We also adapt VoiceCraft-Dub for the video-to-speech task, demonstrating its versatility for various applications. Sung-Bin Kim, Jeongsoo Choi, Puyuan Peng, Joon Son Chung, Tae-Hyun Oh, David F. Harwath |
ICCV | 5 |
| 2025 | JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion TransformersabstractWe present JointDiT, a diffusion transformer that models the joint distribution of RGB and depth. By leveraging the architectural benefit and outstanding image prior of the state-of-the-art diffusion transformer, JointDiT not only generates high-fidelity images but also produces geometrically plausible and accurate depth maps. This solid joint distribution modeling is achieved through two simple yet effective techniques that we propose, namely, adaptive scheduling weights, which depend on the noise levels of each modality, and the unbalanced timestep sampling strategy. With these techniques, we train our model across all noise levels for each modality, enabling JointDiT to naturally handle various combinatorial generation tasks, including joint generation, depth estimation, and depth-conditioned image generation by simply controlling the timesteps of each branch. JointDiT demonstrates outstanding joint generation performance. Furthermore, it achieves comparable results in depth estimation and depth-conditioned image generation, suggesting that joint distribution modeling can serve as a viable alternative to conditional generation. The project page is available at https://byungki-k.github.io/JointDiT/. Byung-Ki Kwon, Qi Dai 0001, Lee Hyoseok, Chong Luo 0001, Tae-Hyun Oh |
ICCV | 5 |
| 2025 | AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language ModelsabstractFollowing the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying not only on text but also on auditory and visual cues for a complete understanding of the world. In recognition of this fact, audio-visual LLMs have recently emerged. Despite promising developments, the lack of dedicated benchmarks poses challenges for understanding and evaluating models. In this work, we show that audio-visual LLMs struggle to discern subtle relationships between audio and visual signals, leading to hallucinations and highlighting the need for reliable benchmarks. To address this, we introduce AVHBench, the first comprehensive benchmark specifically designed to evaluate the perception and comprehension capabilities of audio-visual LLMs. Our benchmark includes tests for assessing hallucinations,
as well as the cross-modal matching and reasoning abilities of these models. Our results reveal that most existing audio-visual LLMs struggle with hallucinations caused by cross-interactions between modalities, due to their limited capacity to perceive complex multimodal signals and their relationships. Additionally, we demonstrate that simple training with our AVHBench improves robustness of audio-visual LLMs against hallucinations. Dataset: https://github.com/kaist-ami/AVHBench Sung-Bin Kim, Oh Hyun-Bin, JungMok Lee, Arda Senocak, Joon Son Chung, Tae-Hyun Oh |
ICLR | 6 |
| 2025 | AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech GenerationabstractIn this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio. This task has gained increasing attention due to its wide range of applications, such as film production, dubbing, and virtual avatars. Despite recent progress, existing methods still suffer from limitations in speech intelligibility, audio-video synchronization, speech naturalness, and voice similarity to the reference speaker. To address these challenges, we propose AlignDiT, a multimodal Aligned Diffusion Transformer that generates accurate, synchronized, and natural-sounding speech from aligned multimodal inputs. Built upon the in-context learning capability of the DiT architecture, AlignDiT explores three effective strategies to align multimodal representations. Furthermore, we introduce a novel multimodal classifier-free guidance mechanism that allows the model to adaptively balance information from each modality during speech synthesis. Extensive experiments demonstrate that AlignDiT significantly outperforms existing methods across multiple benchmarks in terms of quality, synchronization, and speaker similarity. Moreover, AlignDiT exhibits strong generalization capability across various multimodal tasks, such as video-to-speech synthesis and visual forced alignment, consistently achieving state-of-the-art performance. The demo page is available at https://mm.kaist.ac.kr/projects/AlignDiT. Jeongsoo Choi, Sung-Bin Kim, Tae-Hyun Oh, Joon Son Chung |
ACM Multimedia | 4 |
| 2025 | Automated Model Discovery via Multi-modal & Multi-step PipelineabstractAutomated model discovery is the process of automatically searching and identifying the most appropriate model for a given dataset over a large combinatorial search space. Existing approaches, however, often face challenges in balancing the capture of fine-grained details with ensuring generalizability beyond training data regimes with a reasonable model complexity. In this paper, we present a multi-modal \& multi-step pipeline for effective automated model discovery. Our approach leverages two vision-language-based modules (VLM), AnalyzerVLM and EvaluatorVLM, for effective model proposal and evaluation in an agentic way. AnalyzerVLM autonomously plans and executes multi-step analyses to propose effective candidate models. EvaluatorVLM assesses the candidate models both quantitatively and perceptually, regarding the fitness for local details and the generalibility for overall trends. Our results demonstrate that our pipeline effectively discovers models that capture fine details and ensure strong generalizability. Additionally, extensive ablation studies show that both multi-modality and multi-step reasoning play crucial roles in discovering favorable models. JungMok Lee, Nam Hyeon-Woo, Moon Ye-Bin, Junhyun Nam, Tae-Hyun Oh |
NeurIPS | 5 |
| 2025 | Toward Interactive Sound Source Localization: Better Align Sight and Sound!abstractRecent studies on learning-based sound source localization have primarily focused on localization performance. However, prior work and existing benchmarks often overlook a crucial aspect: cross-modal interaction, which is essential for interactive sound source localization. This interaction is vital for understanding semantically matched or mismatched audio-visual events, such as silent objects or true sound sources among multiple objects. In this work, we comprehensively examine the cross-modal interaction of existing methods, benchmarks, evaluation metrics, and cross-modal understanding tasks. We identify the overlooked points of previous studies and make several contributions to address them. First, we propose a learning framework that incorporates retrieval-based and hand-crafted augmentation techniques, enhancing cross-modal interaction through cross-modal alignment. Second, we introduce new evaluation metrics to accurately and rigorously assess localization methods, focusing on both localization performance and cross-modal interaction. Third, to thoroughly analyze interactive sound source localization, we present a new semi-synthetic benchmark with diverse categorical combinations. Finally, we evaluate both interactive sound source localization and auxiliary cross-modal retrieval tasks, benchmarking competing methods alongside our own. Our new benchmark and evaluation metrics reveal that previous methods struggle with interactive sound source localization tasks, largely due to their limited cross-modal interaction capabilities. Our method, which features enhanced cross-modal alignment, demonstrates superior sound source localization and cross-modal interaction performance. This work provides the most comprehensive analysis of sound source localization to date, with extensive validation of competing methods on both existing and new benchmarks using both new and standard evaluation metrics. Arda Senocak, Hyeonggon Ryu, Junsik Kim 0001, Tae-Hyun Oh, Hanspeter Pfister, Joon Son Chung |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | A unified framework for unsupervised action learning via global-to-local motion transformer
Boeun Kim, Hyung Jin Chang, Tae-Hyun Oh |
Pattern Recognit. | 4 |
| 2025 | SYNAuG: Exploiting synthetic data for data imbalance problems
Moon Ye-Bin, Nam Hyeon-Woo, Won-Seok Choi 0001, Nayeong Kim, Suha Kwak, Tae-Hyun Oh |
Pattern Recognit. Lett. | 6 |
| 2024 | FPRF: Feed-Forward Photorealistic Style Transfer of Large-Scale 3D Neural Radiance FieldsabstractWe present FPRF, a feed-forward photorealistic style transfer method for large-scale 3D neural radiance fields. FPRF stylizes large-scale 3D scenes with arbitrary, multiple style reference images without additional optimization while preserving multi-view appearance consistency. Prior arts required tedious per-style/-scene optimization and were limited to small-scale 3D scenes. FPRF efficiently stylizes large-scale 3D scenes by introducing a style-decomposed 3D neural radiance field, which inherits AdaIN’s feed-forward stylization machinery, supporting arbitrary style reference images. Furthermore, FPRF supports multi-reference stylization with the semantic correspondence matching and local AdaIN, which adds diverse user control for 3D scene styles. FPRF also preserves multi-view consistency by applying semantic matching and style transfer processes directly onto queried features in 3D space. In experiments, we demonstrate that FPRF achieves favorable photorealistic quality 3D scene stylization for large-scale scenes with diverse reference images. GeonU Kim, Kim Youwang, Tae-Hyun Oh |
AAAI | 3 |
| 2024 | The Devil Is in the Details: Simple Remedies for Image-to-LiDAR Representation Learning
Wonjun Jo, Byung-Ki Kwon, Kim Ji-Yeon, Hawook Jeong, Kyungdon Joo, Tae-Hyun Oh |
ACCV (9) | 6 |
| 2024 | MeTTA: Single-View to 3D Textured Mesh Reconstruction with Test-Time Adaptation
Kim Yu-Ji, Hyunwoo Ha, Kim Youwang, Jaeheung Surh, Hyowon Ha, Tae-Hyun Oh |
BMVC | 6 |
| 2024 | Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based RenderingabstractWe present Paint-it, a text-driven high-fidelity texture map synthesis method for 3D meshes via neural re-parameterized texture optimization. Paint-it synthesizes texture maps from a text description by synthesis-through-optimization, exploiting the Score-Distillation Sampling (SDS). We observe that directly applying SDS yields undesirable texture quality due to its noisy gradients. We reveal the importance of texture parameterization when using SDS. Specifically, we propose Deep Convolutional Physically-Based Rendering (DC-PBR) parameterization, which re-parameterizes the physically-based rendering (PBR) texture maps with randomly initialized convolution-based neural kernels, instead of a standard pixel-based parameterization. We show that DC-PBR inherently schedules the optimization curriculum according to texture frequency and naturally filters out the noisy signals from SDS. In experiments, Paint-it obtains remarkable quality PBR texture maps within 15 min., given only a text description. We demonstrate the generalizability and practicality of Paint-it by synthesizing high-quality texture maps for large-scale mesh datasets and showing test-time applications such as relighting and material control using a popular graphics engine. Project page: https://kim-youwang.github.io/paint-it. Kim Youwang, Tae-Hyun Oh, Gerard Pons-Moll |
CVPR | 2 |
| 2024 | Learning-based Axial Video Motion Magnification
Byung-Ki Kwon, Oh Hyun-Bin, Kim Jun-Seong, Hyunwoo Ha, Tae-Hyun Oh |
ECCV (54) | 5 |
| 2024 | BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-Language Models
Moon Ye-Bin, Nam Hyeon-Woo, Won-Seok Choi 0001, Tae-Hyun Oh |
ECCV (11) | 4 |
| 2024 | Noise Map Guidance: Inversion with Spatial Context for Real Image EditingabstractText-guided diffusion models have become a popular tool in image synthesis, known for producing high-quality and diverse images. However, their application to editing real images often encounters hurdles primarily due to the text condition deteriorating the reconstruction quality and subsequently affecting editing fidelity. Null-text Inversion (NTI) has made strides in this area, but it fails to capture spatial context and requires computationally intensive per-timestep optimization. Addressing these challenges, we present Noise Map Guidance (NMG), an inversion method rich in a spatial context, tailored for real-image editing. Significantly, NMG achieves this without necessitating optimization, yet preserves the editing quality. Our empirical investigations highlight NMG's adaptability across various editing techniques and its robustness to variants of DDIM inversions. Hansam Cho, Jonghyun Lee 0006, Seoung Bum Kim, Tae-Hyun Oh, Yonghyun Jeong |
ICLR | 4 |
| 2024 | CAS: A Probability-Based Approach for Universal Condition Alignment ScoreabstractRecent conditional diffusion models have shown remarkable advancements and have been widely applied in fascinating real-world applications. However, samples generated by these models often do not strictly comply with user-provided conditions. Due to this, there have been few attempts to evaluate this alignment via pre-trained scoring models to select well-generated samples. Nonetheless, current studies are confined to the text-to-image domain and require large training datasets. This suggests that crafting alignment scores for various conditions will demand considerable resources in the future. In this context, we introduce a universal condition alignment score that leverages the conditional probability measurable through the diffusion process. Our technique operates across all conditions and requires no additional models beyond the diffusion model used for generation, effectively enabling self-rejection. Our experiments validate that our met- ric effectively applies in diverse conditional generations, such as text-to-image, {instruction, image}-to-image, edge-/scribble-to-image, and text-to-audio. Chunsan Hong, Byunghee Cha, Tae-Hyun Oh |
ICLR | 3 |
| 2024 | Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading Expert
Han EunGi, Oh Hyun-Bin, Sung-Bin Kim, Corentin Nivelet Etcheberry, Suekyeong Nam, Janghoon Ju, Tae-Hyun Oh |
INTERSPEECH | 7 |
| 2024 | MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset
Sung-Bin Kim, Lee Chae-Yeon, Gihun Son, Oh Hyun-Bin, Janghoon Ju, Suekyeong Nam, Tae-Hyun Oh |
INTERSPEECH | 7 |
| 2024 | LaughTalk: Expressive 3D Talking Head Generation with LaughterabstractLaughter is a unique expression, essential to affirmative social interactions of humans. Although current 3D talking head generation methods produce convincing verbal articulations, they often fail to capture the vitality and subtleties of laughter and smiles despite their importance in social context. In this paper, we introduce a novel task to generate 3D talking heads capable of both articulate speech and authentic laughter. Our newly curated dataset comprises 2D laughing videos paired with pseudo-annotated and human-validated 3D FLAME parameters and vertices. Given our proposed dataset, we present a strong baseline with a two-stage training scheme: the model first learns to talk and then acquires the ability to express laughter. Extensive experiments demonstrate that our method performs favorably compared to existing approaches in both talking head generation and expressing laughter signals. We further explore potential applications on top of our proposed method for rigging realistic avatars. Sung-Bin Kim, Lee Hyun 0001, Da Hye Hong, Suekyeong Nam, Janghoon Ju, Tae-Hyun Oh |
WACV | 6 |
| 2024 | ENInst: Enhancing weakly-supervised low-shot instance segmentation
Moon Ye-Bin, Dongmin Choi, Yongjin Kwon, Junsik Kim 0001, Tae-Hyun Oh |
Pattern Recognit. | 5 |
| 2024 | An Iterative Method for Unsupervised Robust Anomaly Detection Under Data ContaminationabstractMost deep anomaly detection models are based on learning normality from datasets due to the difficulty of defining abnormality by its diverse and inconsistent nature. Therefore, it has been a common practice to learn normality under the assumption that anomalous data are absent in a training dataset, which we call normality assumption. However, in practice, the normality assumption is often violated due to the nature of real data distributions that includes anomalous tails, i.e., a contaminated dataset. Thereby, the gap between the assumption and actual training data affects detrimentally in learning of an anomaly detection model. In this work, we propose a learning framework to reduce this gap and achieve better normality representation. Our key idea is to identify sample-wise normality and utilize it as an importance weight, which is updated iteratively during the training. Our framework is designed to be model-agnostic and hyperparameter insensitive so that it applies to a wide range of existing methods without careful parameter tuning. We apply our framework to three different representative approaches of deep anomaly detection that are classified into one-class classification-, probabilistic model-, and reconstruction-based approaches. In addition, we address the importance of a termination condition for iterative methods and propose a termination criterion inspired by the anomaly detection objective. We validate that our framework improves the robustness of the anomaly detection models under different levels of contamination ratios on five anomaly detection benchmark datasets and two image datasets. On various contaminated datasets, our framework improves the performance of three representative anomaly detection methods, measured by area under the ROC curve. Minkyung Kim 0001, Jongmin Yu, Junsik Kim 0001, Tae-Hyun Oh, Jun Kyun Choi |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | The devil in the details: simple and effective optical flow synthetic data generation
Byung-Ki Kwon, Sung-Bin Kim, Tae-Hyun Oh |
Vis. Comput. | 3 |
| 2024 | Multi-stage adaptive rank statistic pruning for lightweight human 3D mesh recovery model
Donghun Ryou, Kim Youwang, Tae-Hyun Oh |
Vis. Comput. | 3 |
| 2023 | Sound to Visual Scene Generation by Audio-to-Visual Latent AlignmentabstractHow does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with the large gaps that often exist between sight and sound. We design a model that works by scheduling the learning procedure of each model component to associate audio-visual modalities despite their information gaps. The key idea is to enrich the audio features with visual information by learning to align audio to visual latent space. We translate the input audio to visual features, then use a pre-trained generator to produce an image. To further improve the quality of our generated images, we use sound source localization to select the audio-visual pairs that have strong cross-modal correlations. We obtain substantially better results on the VEGAS and VGGSound datasets than prior approaches. We also show that we can control our model's predictions by applying simple manipulations to the input waveform, or to the latent space. Sung-Bin Kim, Arda Senocak, Hyunwoo Ha, Andrew Owens, Tae-Hyun Oh |
CVPR | 5 |
| 2023 | FPGA-Based Accelerator for Rank-Enhanced and Highly-Pruned Block-Circulant Neural NetworksabstractNumerous network compression methods have been proposed to deploy deep neural networks in a resource-constrained embedded system. Among them, block-circulant matrix (BCM) compression is one of the promising hardware-friendly methods for both acceleration and compression. However, it has several limitations; (i) limited representation due to the structural characteristic of circulant matrix, (ii) limitation of the compression parameter, (iii) need to specialize the dataflow for BCM-compressed network accelerators. In this paper, rank-enhanced and highly-pruned block-circulant matrices compression (RP-BCM) framework is proposed to overcome these limitations. RP-BCM comprises two stages: Hadamard-BCM and BCM-wise pruning. Moreover, a dedicated skip scheme is introduced to processing element design for exploiting high-parallelism with BCM-wise sparsity. Furthermore, we propose specialized dataflow for a BCM-compressed network on a resource-constrained FPGA. As a result, the proposed method achieves parameter reduction and FLOPs reduction for ResNet-50 in ImageNet by 92.4% and 77.3%, respectively. Moreover, the proposed hardware design achieves$3.1\times$improvement in energy efficiency on the Xilinx PYNQ-Z2 FPGA board for ResNet-18 on ImageNet compared to the GPU. Haena Song, Jongho Yoon 0001, Eunji Kwon, Tae-Hyun Oh, Seokhyeong Kang |
DATE | 5 |
| 2023 | Prefix Tuning for Automated Audio CaptioningabstractAudio captioning aims to generate text descriptions from environmental sounds. One challenge of audio captioning is the difficulty of the generalization due to the lack of audio-text paired training data. In this work, we propose a simple yet effective method of dealing with small-scaled datasets by leveraging a pre-trained language model. We keep the language model frozen to maintain the expressivity for text generation, and we only learn to extract global and temporal features from the input audio. To bridge a modality gap between the audio features and the language model, we employ mapping networks that translate audio features to the continuous vectors the language model can understand, called prefixes. We evaluate our proposed method on the Clotho and AudioCaps dataset and show our method outperforms prior arts in diverse experimental settings. Sung-Bin Kim, Tae-Hyun Oh |
ICASSP | 3 |
| 2023 | Unsupervised Pre-Training for Data-Efficient Text-to-Speech on Low Resource LanguagesabstractNeural text-to-speech (TTS) models can synthesize natural human speech when trained on large amounts of transcribed speech. How-ever, collecting such large-scale transcribed data is expensive. This paper proposes an unsupervised pre-training method for a sequence-to-sequence TTS model by leveraging large untranscribed speech data. With our pre-training, we can remarkably reduce the amount of paired transcribed data required to train the model for the target downstream TTS task. The main idea is to pre-train the model to reconstruct de-warped mel-spectrograms from warped ones, which may allow the model to learn proper temporal assignment relation between input and output sequences. In addition, we propose a data augmentation method that further improves the data efficiency in finetuning. We empirically demonstrate the effectiveness of our proposed method in low-resource language scenarios, achieving outstanding performance compared to competing methods. The code and audio samples are available at: https://github.com/cnaigithub/SpeechDewarping Seongyeon Park, Myungseo Song, Bohyung Kim, Tae-Hyun Oh |
ICASSP | 4 |
| 2023 | Scratching Visual Transformer's Back with Uniform AttentionabstractThe favorable performance of Vision Transformers (ViTs) is often attributed to the multi-head self-attention (MSA), which enables global interactions at each layer of a ViT model. Previous works acknowledge the property of long-range dependency for the effectiveness in MSA. In this work, we study the role of MSA in terms of the different axis, density. Our preliminary analyses suggest that the spatial interactions of learned attention maps are close to dense interactions rather than sparse ones. This is a curious phenomenon because dense attention maps are harder for the model to learn due to softmax. We interpret this opposite behavior against softmax as a strong preference for the ViT models to include dense interaction. We thus manually insert the dense uniform attention to each layer of the ViT models to supply the much-needed dense interactions. We call this method Context Broadcasting, CB. Our study demonstrates the inclusion of CB takes the role of dense attention and thereby reduces the degree of density in the original attention maps by complying softmax in MSA. We also show that, with negligible costs of CB (1 line in your model code and no additional parameters), both the capacity and generalizability of the ViT models are increased. Nam Hyeon-Woo, Kim Yu-Ji, Byeongho Heo, Dongyoon Han, Seong Joon Oh, Tae-Hyun Oh |
ICCV | 6 |
| 2023 | Sound Source Localization is All about Cross-Modal AlignmentabstractHumans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localization perspective. However, prior arts and existing benchmarks do not account for a more important aspect of the problem, cross-modal semantic understanding, which is essential for genuine sound source localization. Cross-modal semantic understanding is important in understanding semantically mismatched audio-visual events, e.g., silent objects, or off-screen sounds. To account for this, we propose a cross-modal alignment task as a joint task with sound source localization to better learn the interaction between audio and visual modalities. Thereby, we achieve high localization performance with strong cross-modal semantic understanding. Our method outperforms the state-of-the-art approaches in both sound source localization and cross-modal retrieval. Our work suggests that jointly tackling both tasks is necessary to conquer genuine sound source localization. Arda Senocak, Hyeonggon Ryu, Junsik Kim 0001, Tae-Hyun Oh, Hanspeter Pfister, Joon Son Chung |
ICCV | 4 |
| 2023 | TextManiA: Enriching Visual Feature by Text-driven Manifold AugmentationabstractWe propose TextManiA, a text-driven manifold augmentation method that semantically enriches visual feature spaces, regardless of class distribution. TextManiA augments visual data with intra-class semantic perturbation by exploiting easy-to-understand visually mimetic words, i.e., attributes. This work is built on an interesting hypothesis that general language models, e.g., BERT and GPT, encompass visual information to some extent, even without training on visual training data. Given the hypothesis, TextManiA transfers pre-trained text representation obtained from a well-established large language encoder to a target visual feature space being learned. Our extensive analysis hints that the language encoder indeed encompasses visual information at least useful to augment visual representation. Our experiments demonstrate that TextManiA is particularly powerful in scarce samples with class imbalance as well as even distribution. We also show compatibility with the label mix-based approaches in evenly distributed scarce data. Moon Ye-Bin, Jisoo Kim 0006, Hongyeob Kim, Kilho Son, Tae-Hyun Oh |
ICCV | 5 |
| 2023 | DFlow: Learning to Synthesize Better Optical Flow Datasets via a Differentiable Pipeline
Byung-Ki Kwon, Nam Hyeon-Woo, Ji-Yun Kim, Tae-Hyun Oh |
ICLR | 4 |
| 2023 | Automatic Tuning of Loss Trade-offs without Hyper-parameter Search in End-to-End Zero-Shot Speech Synthesis
Seongyeon Park, Bohyung Kim, Tae-Hyun Oh |
INTERSPEECH | 3 |
| 2023 | Mask-KLT: Sub-pixel Accurate Directional Motion Estimation by Stripe MaskingabstractTo diagnose the safety or monitor the health of aging infrastructures and factory facilities, a way to measure precise vibration or displacement is required. In this paper, we present a sub-pixel accurate tracking method of directional motion of interest, which we name Mask-KLT by inheriting the long-standing Kanade–Lucas–Tomasi tracker. Our simple method can estimate directional displacements of an object of interest along any single axis and is also robust to non-corner regions lacking good features to track. Our algorithm can be easily implemented by simple modification of the KLT tracker: by pre-processing an input image with a stripe pattern masking. We evaluate our proposed method with sub-pixel accuracy on a synthetic dataset. Our Mask-KLT dramatically increases the success rate of tracking and decreases displacement errors compared to Vanilla KLT. Moreover, we demonstrate the effectiveness of our method in a real-world experimental setup: a rotating machine in a laboratory environment. Hyunwoo Ha, Tae-Hyun Oh |
VCIP | 2 |
| 2023 | Learning Few-shot Segmentation from Bounding Box AnnotationsabstractWe present a new weakly-supervised few-shot semantic segmentation setting and a meta-learning method for tackling the new challenge. Different from existing settings, we leverage bounding box annotations as weak supervision signals during the meta-training phase, i.e., more label-efficient. Bounding box provides a cheaper label representation than segmentation mask but contains both an object of interest and a disturbing background. We first show that meta-training with bounding boxes degrades recent few-shot semantic segmentation methods, which are typically meta-trained with full semantic segmentation supervisions. We postulate that this challenge is originated from the impure information of bounding box representation. We propose a pseudo trimap estimator and trimap-attention based prototype learning to extract clearer supervision signals from bounding boxes. These developments robustify and generalize our method well to noisy support masks at test time. We empirically show that our method consistently improves performance. Our method gains 1.4% and 3.6% mean-IoU over the competing one in full and weak test supervision cases, respectively, in the 1-way 5-shot setting on Pascal-5i. Byeolyi Han, Tae-Hyun Oh |
WACV | 2 |
| 2023 | Event-Specific Audio-Visual Fusion Layers: A Simple and New Perspective on Video UnderstandingabstractTo understand our surrounding world, our brain is continuously inundated with multisensory information and their complex interactions coming from the outside world at any given moment. While processing this information might seem effortless for human brains, it is challenging to build a machine that can perform similar tasks since complex interactions cannot be dealt with a single type of integration but require more sophisticated approaches. In this paper, we propose a new simple method to address the multisensory integration in video understanding. Unlike previous works where a single fusion type is used, we design a multi-head model with individual event-specific layers to deal with different audio-visual relationships, enabling different ways of audio-visual fusion. Experimental results show that our event-specific layers can discover unique properties of the audio-visual relationships in the videos, e.g., semantically matched moments, and rhythmic events. Moreover, although our network is trained with single labels, our multi-head design can inherently output additional semantically meaningful multi-labels for a video. As an application, we demonstrate that our proposed method can expose the extent of event-characteristics of popular benchmark datasets. Arda Senocak, Junsik Kim 0001, Tae-Hyun Oh, Dingzeyu Li, In-So Kweon |
WACV | 3 |
| 2022 | Cross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers
Junhyeong Cho, Kim Youwang, Tae-Hyun Oh |
ECCV (1) | 3 |
| 2022 | HDR-Plenoxels: Self-Calibrating High Dynamic Range Radiance Fields
Kim Jun-Seong, Kim Yu-Ji, Moon Ye-Bin, Tae-Hyun Oh |
ECCV (32) | 4 |
| 2022 | CLIP-Actor: Text-Driven Recommendation and Stylization for Animating Human Meshes
Kim Youwang, Tae-Hyun Oh |
ECCV (3) | 3 |
| 2022 | FedPara: Low-rank Hadamard Product for Communication-Efficient Federated Learning
Nam Hyeon-Woo, Moon Ye-Bin, Tae-Hyun Oh |
ICLR | 3 |
| 2022 | Robust and Efficient Estimation of Relative Pose for Cameras on Selfie SticksabstractTaking selfies has become one of the major photographic trends of our time. In this study, we focus on the selfie stick, on which a camera is mounted to take selfies. We observe that a camera on a selfie stick typically travels through a particular type of trajectory around a sphere. Based on this finding, we propose a robust, efficient, and optimal estimation method for relative camera pose between two images captured by a camera mounted on a selfie stick. We exploit the special geometric structure of camera motion constrained by a selfie stick and define this motion as spherical joint motion. Utilizing a novel parametrization and calibration scheme, we demonstrate that the pose estimation problem can be reduced to a 3-degrees of freedom (DoF) search problem, instead of a generic 6-DoF problem. This facilitates the derivation of an efficient branch-and-bound optimization method that guarantees a global optimal solution, even in the presence of outliers. Furthermore, as a simplified case of spherical joint motion, we introduce selfie motion, which has a fewer number of DoF than spherical joint motion. We validate the performance and guaranteed optimality of our method on both synthetic and real-world data. Additionally, we demonstrate the applicability of the proposed method for two applications: refocusing and stylization. Kyungdon Joo, Hongdong Li, Tae-Hyun Oh, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Dense Relational Image Captioning via Multi-Task Triple-Stream NetworksabstractWe introduce dense relational captioning, a novel image captioning task which aims to generate multiple captions with respect to relational information between objects in a visual scene. Relational captioning provides explicit descriptions for each relationship between object combinations. This framework is advantageous in both diversity and amount of information, leading to a comprehensive image understanding based on relationships, e.g., relational proposal generation. For relational understanding between objects, the part-of-speech (POS; i.e., subject-object-predicate categories) can be a valuable prior information to guide the causal sequence of words in a caption. We enforce our framework to learn not only to generate captions but also to understand the POS of each word. To this end, we propose the multi-task triple-stream network (MTTSNet) which consists of three recurrent units responsible for each POS which is trained by jointly predicting the correct captions and POS for each word. In addition, we found that the performance of MTTSNet can be improved by modulating the object embeddings with an explicit relational module. We demonstrate that our proposed model can generate more diverse and richer captions, via extensive experimental analysis on large scale datasets and several metrics. Then, we present applications of our framework to holistic image captioning, scene graph generation, and retrieval tasks. Dong-Jin Kim 0003, Tae-Hyun Oh, Jinsoo Choi, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Lightweight Speaker Recognition in Poincaré SpacesabstractThis letter proposes a lightweight model for speaker recognition by leveraging a hyperbolic space. The speaker recognition performance heavily depends on the distinctiveness of speaker embeddings induced by metric learning. However, most state-of-the-art embedding methods are typically based on the Euclidean metric space, which does not account for inherent hierarchical structures of speech voice characteristics. The recent development of the neural hyperbolic geometry has demonstrated its effectiveness to model continuous hierarchical structures, which have been typically cumbersome to model by standard deep neural networks. This facet provides an additional by-product of a compact representation. Inspired by the favorable geometry of the hyperbolic geometry, we developed a hyperbolic ResNet for speaker recognition. We found that in smaller dimension regimes than typical cases, the learned speaker embeddings are more discriminative; in other words, more compact at the same level of performance. Our experiments on the large-scale VoxCeleb datasets show that, given the limited channel dimensions of neural networks, our method consistently has favorable performance against the standard ResNet for both speaker recognition and verification tasks. Sung-Bin Kim, Seokhyeong Kang, Tae-Hyun Oh |
IEEE Signal Process. Lett. | 4 |
| 2021 | Unified 3D Mesh Recovery of Humans and Animals by Learning Animal Exercise
Kim Youwang, Kyungdon Joo, Tae-Hyun Oh |
BMVC | 4 |
| 2021 | Monocular Reconstruction of Neural Face Reflectance FieldsabstractThe reflectance field of a face describes the reflectance properties responsible for complex lighting effects including diffuse, specular, inter-reflection and self shadowing. Most existing methods for estimating the face reflectance from a monocular image assume faces to be diffuse with very few approaches adding a specular component. This still leaves out important perceptual aspects of reflectance such as higher-order global illumination effects and self-shadowing. We present a new neural representation for face reflectance where we can estimate all components of the reflectance responsible for the final appearance from a monocular image. Instead of modeling each component of the reflectance separately using parametric models, our neural representation allows us to generate a basis set of faces in a geometric deformation-invariant space, parameterized by the input light direction, viewpoint and face geometry. We learn to reconstruct this reflectance field of a face just from a monocular image, which can be used to render the face from any viewpoint in any light condition. Our method is trained on a light-stage dataset, which captures 300 people illuminated with 150 light conditions from 8 viewpoints. We show that our method outperforms existing monocular reflectance reconstruction methods due to better capturing of physical effects, such as sub-surface scattering, specularities, self-shadows and other higher-order effects. Mallikarjun B. R. 0001, Ayush Tewari, Tae-Hyun Oh, Tim Weyrich, Bernd Bickel, Hans-Peter Seidel, Hanspeter Pfister, Wojciech Matusik, Mohamed A. Elgharib, Christian Theobalt |
CVPR | 3 |
| 2021 | MDARTS: Multi-objective Differentiable Neural Architecture SearchabstractIn this work, we present a differentiable neural architecture search (NAS) method that takes into account two competing objectives, quality of result (QoR) and quality of service (QoS) with hardware design constraints. NAS research has recently received a lot of attention due to its ability to automatically find architecture candidates that can outperform handcrafted ones. However, the NAS approach which complies with actual HW design constraints has been under-explored. A naive NAS approach for this would be to optimize a combination of two criteria of QoR and QoS, but the simple extension of the prior art often yields degenerated architectures, and suffers from a sensitive hyperparameter tuning. In this work, we propose a multi-objective differential neural architecture search, called MDARTS. MDARTS has an affordable search time and can find Pareto frontier of QoR versus QoS. We also identify the problematic gap between all the existing differentiable NAS results and those final post-processed architectures, where soft connections are binarized. This gap leads to performance degradation when the model is deployed. To mitigate this gap, we propose a separation loss that discourages indefinite connections of components by implicitly minimizing entropy. Hyun-jeong Kwon, Eunji Kwon, Youngchang Choi, Tae-Hyun Oh, Seokhyeong Kang |
DATE | 5 |
| 2021 | Distilling Global and Local Logits with Densely Connected RelationsabstractIn prevalent knowledge distillation, logits in most image recognition models are computed by global average pooling, then used to learn to encode the high-level and task-relevant knowledge. In this work, we solve the limitation of this global logit transfer in this distillation context. We point out that it prevents the transfer of informative spatial information, which provides localized knowledge as well as rich relational information across contexts of an input scene. To exploit the rich spatial information, we propose a simple yet effective logit distillation approach. We add a local spatial pooling layer branch to the penultimate layer, thereby our method extends the standard logit distillation and enables learning of both finely-localized knowledge and holistic representation. Our proposed method shows favorable accuracy improvement against the state-of-the-art methods on several image classification datasets. We show that our distilled students trained on the image classification task can be successfully leveraged for object detection and semantic segmentation tasks; this result demonstrates our method’s high transferability. Youmin Kim, Jinbae Park, Younho Jang, Muhammad Salman Ali, Tae-Hyun Oh, Sung-Ho Bae |
ICCV | 5 |
| 2021 | CDS: Cross-Domain Self-supervised Pre-trainingabstractWe present a two-stage pre-training approach that improves the generalization ability of standard single-domain pre-training. While standard pre-training on a single large dataset (such as ImageNet) can provide a good initial representation for transfer learning tasks, this approach may result in biased representations that impact the success of learning with new multi-domain data (e.g., different artistic styles) via methods like domain adaptation. We propose a novel pre-training approach called Cross-Domain Self-supervision (CDS), which directly employs unlabeled multi-domain data for downstream domain transfer tasks. Our approach uses self-supervision not only within a single domain but also across domains. In-domain instance discrimination is used to learn discriminative features on new data in a domain-adaptive manner, while cross-domain matching is used to learn domain-invariant features. We apply our method as a second pre-training step (after ImageNet pre-training), resulting in a significant target accuracy boost to diverse domain transfer tasks compared to standard one-stage pre-training. Donghyun Kim 0006, Kuniaki Saito, Tae-Hyun Oh, Bryan A. Plummer, Stan Sclaroff, Kate Saenko |
ICCV | 3 |
| 2021 | Supervoxel Attention Graphs for Long-Range Video ModelingabstractA significant challenge in video understanding is posed by the high dimensionality of the input, which induces large computational cost and high memory footprints. Deep convolutional models operating on video apply pooling and striding to reduce feature dimensionality and to increase the receptive field. However, despite these strategies, modern approaches cannot effectively leverage spatiotemporal structure over long temporal extents. In this paper we introduce an approach that reduces a video of 10 seconds to a sparse graph of only 160 feature nodes such that efficient inference in this graph produces state-of-the-art accuracy on challenging action recognition datasets. The nodes of our graph are semantic supervoxels that capture the spatiotemporal structure of objects and motion cues in the video, while edges between nodes encode spatiotemporal relations and feature similarity. We demonstrate that a shallow network that interleaves graph convolution and graph pooling on this compact representation implements an effective mechanism of relational reasoning yielding strong recognition results on both Charades and Something-Something. Yang Wang 0097, Gedas Bertasius, Tae-Hyun Oh, Minh Hoai, Lorenzo Torresani |
WACV | 3 |
| 2021 | Learning to Localize Sound Sources in Visual Scenes: Analysis and ApplicationsabstractVisual events are usually accompanied by sounds in our daily lives. However, can the machines learn to correlate the visual scene and sound, as well as localize the sound source only by observing them like humans? To investigate its empirical learnability, in this work we first present a novel unsupervised algorithm to address the problem of localizing sound sources in visual scenes. In order to achieve this goal, a two-stream network structure which handles each modality with attention mechanism is developed for sound source localization. The network naturally reveals the localized response in the scene without human annotation. In addition, a new sound source dataset is developed for performance evaluation. Nevertheless, our empirical evaluation shows that the unsupervised method generates false conclusions in some cases. Thereby, we show that this false conclusion cannot be fixed without human prior knowledge due to the well-known correlation and causality mismatch misconception. To fix this issue, we extend our network to the supervised and semi-supervised network settings via a simple modification due to the general architecture of our two-stream network. We show that the false conclusions can be effectively corrected even with a small amount of supervision, i.e., semi-supervised setup. Furthermore, we present the versatility of the learned audio and visual embeddings on the cross-modal content alignment and we extend this proposed algorithm to a new application, sound saliency based automatic camera view panning in 360 degree videos. Arda Senocak, Tae-Hyun Oh, Junsik Kim 0001, Ming-Hsuan Yang 0001, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Listen to Look: Action Recognition by Previewing AudioabstractIn the face of the video data deluge, today's expensive clip-level classifiers are increasingly impractical. We propose a framework for efficient action recognition in untrimmed video that uses audio as a preview mechanism to eliminate both short-term and long-term visual redundancies. First, we devise an ImgAud2Vid framework that hallucinates clip-level features by distilling from lighter modalities---a single frame and its accompanying audio---reducing short-term temporal redundancy for efficient clip-level recognition. Second, building on ImgAud2Vid, we further propose ImgAud-Skimming, an attention-based long short-term memory network that iteratively selects useful moments in untrimmed videos, reducing long-term temporal redundancy for efficient video-level recognition. Extensive experiments on four action recognition datasets demonstrate that our method achieves the state-of-the-art in terms of both recognition accuracy and speed. Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, Lorenzo Torresani |
CVPR | 2 |
| 2020 | Globally Optimal Relative Pose Estimation for Camera on a Selfie StickabstractTaking selfies has become a photographic trend nowadays. We envision the emergence of the "video selfie" capturing a short continuous video clip (or burst photography) of the user, themselves. A selfie stick is usually used, whereby a camera is mounted on a stick for taking selfie photos. In this scenario, we observe that the camera typically goes through a special trajectory along a sphere surface. Motivated by this observation, in this work, we propose an efficient and globally optimal relative camera pose estimation between a pair of two images captured by a camera mounted on a selfie stick. We exploit the special geometric structure of the camera motion constrained by a selfie stick and define its motion as spherical joint motion. By the new parametrization and calibration scheme, we show that the pose estimation problem can be reduced to a 3-DoF (degrees of freedom) search problem, instead of a generic 6-DoF problem. This allows us to derive a fast branch-and-bound global optimization, which guarantees a global optimum. Thereby, we achieve efficient and robust estimation even in the presence of outliers. By experiments on both synthetic and real-world data, we validate the performance as well as the guaranteed optimality of the proposed method. Kyungdon Joo, Hongdong Li, Tae-Hyun Oh, Yunsu Bok, In-So Kweon |
ICRA | 3 |
| 2020 | Linear RGB-D SLAM for Atlanta WorldabstractWe present a new linear method for RGB-D based simultaneous localization and mapping (SLAM). Compared to existing techniques relying on the Manhattan world assumption defined by three orthogonal directions, our approach is designed for the more general scenario of the Atlanta world. It consists of a vertical direction and a set of horizontal directions orthogonal to the vertical direction and thus can represent a wider range of scenes. Our approach leverages the structural regularity of the Atlanta world to decouple the non-linearity of camera pose estimations. This allows us separately to estimate the camera rotation and then the translation, which bypasses the inherent non-linearity of traditional SLAM techniques. To this end, we introduce a novel tracking-by-detection scheme to estimate the underlying scene structure by Atlanta representation. Thereby, we propose an Atlanta frame-aware linear SLAM framework which jointly estimates the camera motion and a planar map supporting the Atlanta structure through a linear Kalman filter. Evaluations on both synthetic and real datasets demonstrate that our approach provides favorable performance compared to existing state-of-the-art methods while extending their working range to the Atlanta world. Kyungdon Joo, Tae-Hyun Oh, François Rameau, Jean-Charles Bazin, In-So Kweon |
ICRA | 2 |
| 2020 | Globally Optimal Inlier Set Maximization for Atlanta World UnderstandingabstractIn this work, we describe man-made structures via an appropriate structure assumption, called the Atlanta world assumption, which contains a vertical direction (typically the gravity direction) and a set of horizontal directions orthogonal to the vertical direction. Contrary to the commonly used Manhattan world assumption, the horizontal directions in Atlanta world are not necessarily orthogonal to each other. While Atlanta world can encompass a wider range of scenes, this makes the search space much larger and the problem more challenging. Our input data is a set of surface normals, for example, acquired from RGB-D cameras or 3D laser scanners, as well as lines from calibrated images. Given this input data, we propose the first globally optimal method of inlier set maximization for Atlanta direction estimation. We define a novel search space for Atlanta world, as well as its parametrization, and solve this challenging problem using a branch-and-bound (BnB) framework. To alleviate the computational bottleneck in BnB, i.e., the bound computation, we present two bound computation strategies: rectangular bound and slice bound in an efficient measurement domain, i.e., the extended Gaussian image (EGI). In addition, we propose an efficient two-stage method which automatically estimates the number of horizontal directions of a scene. Experimental results with synthetic and real-world datasets have successfully confirmed the validity of our approach. Kyungdon Joo, Tae-Hyun Oh, In-So Kweon, Jean-Charles Bazin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Visuomotor Understanding for Representation Learning of Driving Scenes
Seokju Lee, Junsik Kim 0001, Tae-Hyun Oh, Yongseop Jeong, Donggeun Yoo, Stephen Lin 0001, In-So Kweon |
BMVC | 3 |
| 2019 | Dense Relational Captioning: Triple-Stream Networks for Relationship-Based CaptioningabstractOur goal in this work is to train an image captioning model that generates more dense and informative captions. We introduce "relational captioning," a novel image captioning task which aims to generate multiple captions with respect to relational information between objects in an image. Relational captioning is a framework that is advantageous in both diversity and amount of information, leading to image understanding based on relationships. Part-of-speech (POS, i.e. subject-object-predicate categories) tags can be assigned to every English word. We leverage the POS as a prior to guide the correct sequence of words in a caption. To this end, we propose a multi-task triple-stream network (MTTSNet) which consists of three recurrent units for the respective POS and jointly performs POS prediction and captioning. We demonstrate more diverse and richer representations generated by the proposed model against several baselines and competing methods. Dong-Jin Kim 0003, Jinsoo Choi, Tae-Hyun Oh, In-So Kweon |
CVPR | 3 |
| 2019 | Variational Prototyping-Encoder: One-Shot Learning With Prototypical ImagesabstractIn daily life, graphic symbols, such as traffic signs and brand logos, are ubiquitously utilized around us due to its intuitive expression beyond language boundary. We tackle an open-set graphic symbol recognition problem by one-shot classification with prototypical images as a single training example for each novel class. We take an approach to learn a generalizable embedding space for novel tasks. We propose a new approach called variational prototyping-encoder (VPE) that learns the image translation task from real-world input images to their corresponding prototypical images as a meta-task. As a result, VPE learns image similarity as well as prototypical concepts which differs from widely used metric learning based approaches. Our experiments with diverse datasets demonstrate that the proposed VPE performs favorably against competing metric learning based one-shot methods. Also, our qualitative analyses show that our meta-task induces an effective embedding space suitable for unseen data representation. Junsik Kim 0001, Tae-Hyun Oh, Seokju Lee, In-So Kweon |
CVPR | 2 |
| 2019 | Speech2Face: Learning the Face Behind a VoiceabstractHow much can we infer about a person’s looks from the way they speak? In this paper, we study the task of reconstructing a facial image of a person from a short audio recording of that person speaking. We design and train a deep neural network to perform this task using millions of natural Internet/Youtube videos of people speaking. During training, our model learns voice-face correlations that allow it to produce images that capture various physical attributes of the speakers such as age, gender and ethnicity. This is done in a self-supervised manner, by utilizing the natural co-occurrence of faces and speech in Internet videos, without the need to model attributes explicitly. We evaluate and numerically quantify how–-and in what manner–-our Speech2Face reconstructions, obtained directly from audio, resemble the true face images of the speakers. Tae-Hyun Oh, Tali Dekel, Changil Kim 0001, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Wojciech Matusik |
CVPR | 1 |
| 2019 | Image Captioning with Very Scarce Supervised Data: Adversarial Semi-Supervised Learning ApproachabstractDong-Jin Kim, Jinsoo Choi, Tae-Hyun Oh, In So Kweon. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dong-Jin Kim 0003, Jinsoo Choi, Tae-Hyun Oh, In-So Kweon |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Noise-tolerant Audio-visual Online Person Verification Using an Attention-based Neural Network FusionabstractIn this paper, we present a multi-modal online person verification system using both speech and visual signals. Inspired by neuroscientific findings on the association of voice and face, we propose an attention-based end-to-end neural network that learns multi-sensory association for the task of person verification. The attention mechanism in our proposed network learns to conditionally select a salient modality between speech and facial representations that provides a balance between complementary inputs. By virtue of this capability, the network is robust to missing or corrupted data from either modality. In the VoxCeleb2 dataset, we show that our method performs favorably against competing multi-modal methods. Even for extreme cases of large corruption or missing data on either modality, our method demonstrates robustness over other unimodal methods. Suwon Shon, Tae-Hyun Oh, James R. Glass |
ICASSP | 2 |
| 2019 | Neural Inverse Knitting: From Images to Manufacturing InstructionsabstractMotivated by the recent potential of mass customization brought by whole-garment knitting machines, we introduce the new problem of automatic machine instruction generation using a single image of the desired physical product, which we apply to machine knitting. We propose to tackle this problem by directly learning to synthesize regular machine instructions from real images. We create a cured dataset of real samples with their instruction counterpart and propose to use synthetic images to augment it in a novel way. We theoretically motivate our data mixing framework and show empirical results suggesting that making real images look more synthetic is beneficial in our problem setup. Alexandre Kaspar, Tae-Hyun Oh, Liane Makatura, Petr Kellnhofer, Wojciech Matusik |
ICML | 2 |
| 2019 | Robust and Globally Optimal Manhattan Frame Estimation in Near Real TimeabstractMost man-made environments, such as urban and indoor scenes, consist of a set of parallel and orthogonal planar structures. These structures are approximated by the Manhattan world assumption, in which notion can be represented as a Manhattan frame (MF). Given a set of inputs such as surface normals or vanishing points, we pose an MF estimation problem as a consensus set maximization that maximizes the number of inliers over the rotation search space. Conventionally, this problem can be solved by a branch-and-bound framework, which mathematically guarantees global optimality. However, the computational time of the conventional branch-and-bound algorithms is rather far from real-time. In this paper, we propose a novel bound computation method on an efficient measurement domain for MF estimation, i.e., the extended Gaussian image (EGI). By relaxing the original problem, we can compute the bound with a constant complexity, while preserving global optimality. Furthermore, we quantitatively and qualitatively demonstrate the performance of the proposed method for various synthetic and real-world data. We also show the versatility of our approach through three different applications: extension to multiple MF estimation, 3D rotation based video stabilization, and vanishing point estimation (line clustering). Kyungdon Joo, Tae-Hyun Oh, Junsik Kim 0001, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Gradient-Based Camera Exposure Control for Outdoor Mobile PlatformsabstractWe introduce a novel method to automatically adjust camera exposure for image processing and computer vision applications on mobile robot platforms. Because most image processing algorithms rely heavily on low-level image features that are based mainly on local gradient information, we consider that gradient quantity can determine the proper exposure level, allowing a camera to capture the important image features in a manner robust to illumination conditions. We then extend this concept to a multi-camera system and present a new control algorithm to achieve both brightness consistency between adjacent cameras and a proper exposure level for each camera. We implement our prototype system with off-the-shelf machine-vision cameras and demonstrate the effectiveness of the proposed algorithms on practical applications, including pedestrian detection, visual odometry, surround-view imaging, panoramic imaging, and stereo matching. Inwook Shim, Tae-Hyun Oh, Joon-Young Lee, Jinwook Choi, Dong-Geol Choi, In-So Kweon |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Co-Domain Embedding Using Deep Quadruplet Networks for Unseen Traffic Sign RecognitionabstractRecent advances in visual recognition show overarching success by virtue of large amounts of supervised data. However, the acquisition of a large supervised dataset is often challenging. This is also true for intelligent transportation applications, i.e., traffic sign recognition. For example, a model trained with data of one country may not be easily generalized to another country without much data. We propose a novel feature embedding scheme for unseen class classification when the representative class template is given. Traffic signs, unlike other objects, have official images. We perform co-domain embedding using a quadruple relationship from real and synthetic domains. Our quadruplet network fully utilizes the explicit pairwise similarity relationships among samples from different domains. We validate our method on three datasets with two experiments involving one-shot classification and feature generalization. The results show that the proposed method outperforms competing approaches on both seen and unseen classes. Junsik Kim 0001, Seokju Lee, Tae-Hyun Oh, In-So Kweon |
AAAI | 3 |
| 2018 | On Learning Associations of Faces and Voices
Changil Kim 0001, Hijung Shin, Tae-Hyun Oh, Alexandre Kaspar, Mohamed A. Elgharib, Wojciech Matusik |
ACCV (5) | 3 |
| 2018 | Globally Optimal Inlier Set Maximization for Atlanta Frame EstimationabstractIn this work, we describe man-made structures via an appropriate structure assumption, called Atlanta world, which contains a vertical direction (typically the gravity direction) and a set of horizontal directions orthogonal to the vertical direction. Contrary to the commonly used Manhattan world assumption, the horizontal directions in Atlanta world are not necessarily orthogonal to each other. While Atlanta world permits to encompass a wider range of scenes, this makes the solution space larger and the problem more challenging. Given a set of inputs, such as lines in a calibrated image or surface normals, we propose the first globally optimal method of inlier set maximization for Atlanta direction estimation. We define a novel search space for Atlanta world, as well as its parameterization, and solve this challenging problem by a branch-and-bound framework. Experimental results with synthetic and real-world datasets have successfully confirmed the validity of our approach. Kyungdon Joo, Tae-Hyun Oh, In-So Kweon, Jean-Charles Bazin |
CVPR | 2 |
| 2018 | Learning to Localize Sound Source in Visual ScenesabstractVisual events are usually accompanied by sounds in our daily lives. We pose the question: Can the machine learn the correspondence between visual scene and the sound, and localize the sound source only by observing sound and visual scene pairs like human? In this paper, we propose a novel unsupervised algorithm to address the problem of localizing the sound source in visual scenes. A two-stream network structure which handles each modality, with attention mechanism is developed for sound source localization. Moreover, although our network is formulated within the unsupervised learning framework, it can be extended to a unified architecture with a simple modification for the supervised and semi-supervised learning settings as well. Meanwhile, a new sound source dataset is developed for performance evaluation. Our empirical evaluation shows that the unsupervised method eventually go through false conclusion in some cases. We also show that even with a few supervision, i.e., semi-supervised setup, false conclusion is able to be corrected effectively. Arda Senocak, Tae-Hyun Oh, Junsik Kim 0001, Ming-Hsuan Yang 0001, In-So Kweon |
CVPR | 2 |
| 2018 | Learning-Based Video Motion Magnification
Tae-Hyun Oh, Ronnachai Jaroensri, Changil Kim 0001, Mohamed A. Elgharib, Frédo Durand, William T. Freeman, Wojciech Matusik |
ECCV (4) | 1 |
| 2018 | Disjoint Multi-task Learning Between Heterogeneous Human-Centric TasksabstractHuman behavior understanding is arguably one of the most important mid-level components in artificial intelligence. In order to efficiently make use of data, multi-task learning has been studied in diverse computer vision tasks including human behavior understanding. However, multitask learning relies on task specific datasets and constructing such datasets can be cumbersome. It requires huge amounts of data, labeling efforts, statistical consideration etc. In this paper, we leverage existing single-task datasets for human action classification and captioning data for efficient human behavior learning. Since the data in each dataset has respective heterogeneous annotations, traditional multi-task learning is not effective in this scenario. To this end, we propose a novel alternating directional optimization method to efficiently learn from the heterogeneous data. We demonstrate the effectiveness of our model and show performance improvements on both classification and sentence retrieval tasks in comparison to the models trained on each of the single-task datasets. Dong-Jin Kim 0003, Jinsoo Choi, Tae-Hyun Oh, Youngjin Yoon, In-So Kweon |
WACV | 3 |
| 2018 | Contextually Customized Video Summaries Via Natural LanguageabstractThe best summary of a long video differs among different people due to its highly subjective nature. Even for the same person, the best summary may change with time or mood. In this paper, we introduce the task of generating contextually customized video summaries through simple text. First, we train a deep architecture to effectively learn semantic embeddings of video frames by leveraging the abundance of image-caption data via a progressive manner, whereby our algorithm is able to select semantically relevant video segments for a contextually meaningful video summary, given a user-specific text description or even a single sentence. In order to evaluate our customized video summaries, we conduct experimental comparison with baseline methods that utilize ground-truth information. Despite the challenging baselines, our method still manages to show comparable or even exceeding performance. We also demonstrate that our method is able to automatically generate semantically diverse video summaries even without any text input. Jinsoo Choi, Tae-Hyun Oh, In-So Kweon |
WACV | 2 |
| 2018 | Fast Randomized Singular Value Thresholding for Low-Rank OptimizationabstractRank minimization can be converted into tractable surrogate problems, such as Nuclear Norm Minimization (NNM) and Weighted NNM (WNNM). The problems related to NNM, or WNNM, can be solved iteratively by applying a closed-form proximal operator, called Singular Value Thresholding (SVT), or Weighted SVT, but they suffer from high computational cost of Singular Value Decomposition (SVD) at each iteration. We propose a fast and accurate approximation method for SVT, that we call fast randomized SVT (FRSVT), with which we avoid direct computation of SVD. The key idea is to extract an approximate basis for the range of the matrix from its compressed matrix. Given the basis, we compute partial singular values of the original matrix from the small factored matrix. In addition, by developping a range propagation method, our method further speeds up the extraction of approximate basis at each iteration. Our theoretical analysis shows the relationship between the approximation bound of SVD and its effect to NNM via SVT. Along with the analysis, our empirical results quantitatively and qualitatively show that our approximation rarely harms the convergence of the host algorithms. We assess the efficiency and accuracy of the proposed method on various computer vision problems, e.g., subspace clustering, weather artifact removal, and simultaneous multi-image alignment and rectification. Tae-Hyun Oh, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | A Closed-Form Solution to Rotation Estimation for Structure from Small MotionabstractThe introduction of small motion techniques such as small angle rotation approximation has enabled the three-dimensional reconstruction from a small motion of a camera, so-called structure from small motion (SfSM). In this letter, we propose a closed-form solution dedicated to the rotation estimation problem in SfSM. We show that our method works with a minimal set of two points, and has mild conditions to produce a unique optimal solution in practice. Also, we introduce a three-step SfSM pipeline with better convergence and faster speed compared to the state-of-the-art SfSM approaches. The key to this improvement is the separated estimation of the rotation with the proposed two-point method in order to handle the bas-relief ambiguity that affects the convergence of the bundle adjustment. We demonstrate the effectiveness of our two-point minimal solution and the three-step SfSM approach in synthetic and real-world experiments under the small motion regime. Hyowon Ha, Tae-Hyun Oh, In-So Kweon |
IEEE Signal Process. Lett. | 2 |
| 2018 | Semantic soft segmentationabstractAccurate representation of soft transitions between image regions is essential for high-quality image editing and compositing. Current techniques for generating such representations depend heavily on interaction by a skilled visual artist, as creating such accurate object selections is a tedious task. In this work, we introduce semantic soft segments , a set of layers that correspond to semantically meaningful regions in an image with accurate soft transitions between different objects. We approach this problem from a spectral segmentation angle and propose a graph structure that embeds texture and color features from the image as well as higher-level semantic information generated by a neural network. The soft segments are generated via eigendecomposition of the carefully constructed Laplacian matrix fully automatically. We demonstrate that otherwise complex image editing tasks can be done with little effort using semantic soft segments. Yagiz Aksoy, Tae-Hyun Oh, Sylvain Paris, Marc Pollefeys, Wojciech Matusik |
ACM Trans. Graph. | 2 |
| 2017 | Weakly- and Self-Supervised Learning for Content-Aware Deep Image RetargetingabstractThis paper proposes a weakly- and self-supervised deep convolutional neural network (WSSDCNN) for content-aware image retargeting. Our network takes a source image and a target aspect ratio, and then directly outputs a retargeted image. Retargeting is performed through a shift reap, which is a pixel-wise mapping from the source to the target grid. Our method implicitly learns an attention map, which leads to r content-aware shift map for image retargeting. As a result, discriminative parts in an image are preserved, while background regions are adjusted seamlessly. In the training phase, pairs of an image and its image-level annotation are used to compute content and structure tosses. We demonstrate the effectiveness of our proposed method for a retargeting application with insightful analyses. Donghyeon Cho, Jinsun Park, Tae-Hyun Oh, Yu-Wing Tai, In-So Kweon |
ICCV | 3 |
| 2017 | Personalized Cinemagraphs Using Semantic Understanding and Collaborative LearningabstractCinemagraphs are a compelling way to convey dynamic aspects of a scene. In these media, dynamic and still elements are juxtaposed to create an artistic and narrative experience. Creating a high-quality, aesthetically pleasing cinemagraph requires isolating objects in a semantically meaningful way and then selecting good start times and looping periods for those objects to minimize visual artifacts (such a tearing). To achieve this, we present a new technique that uses object recognition and semantic segmentation as part of an optimization method to automatically create cinemagraphs from videos that are both visually appealing and semantically meaningful. Given a scene with multiple objects, there are many cinemagraphs one could create. Our method evaluates these multiple candidates and presents the best one, as determined by a model trained to predict human preferences in a collaborative way. We demonstrate the effectiveness of our approach with multiple results and a user study. Tae-Hyun Oh, Kyungdon Joo, Neel Joshi, Baoyuan Wang, In-So Kweon, Sing Bing Kang |
ICCV | 1 |
| 2016 | Video-Story Composition via Plot AnalysisabstractWe address the problem of composing a story out of multiple short video clips taken by a person during an activity or experience. Inspired by plot analysis of written stories, our method generates a sequence of video clips ordered in such a way that it reflects plot dynamics and content coherency. That is, given a set of multiple video clips, our method composes a video which we call a video-story. We define metrics on scene dynamics and coherency by dense optical flow features and a patch matching algorithm. Using these metrics, we define an objective function for the video-story. To efficiently search for the best video-story, we introduce a novel Branch-and-Bound algorithm which guarantees the global optimum. We collect the dataset consisting of 23 video sets from the web, resulting in a total of 236 individual video clips. With the acquired dataset, we perform extensive user studies involving 30 human subjects by which the effectiveness of our approach is quantitatively and qualitatively verified. Jinsoo Choi, Tae-Hyun Oh, In-So Kweon |
CVPR | 2 |
| 2016 | Globally Optimal Manhattan Frame Estimation in Real-TimeabstractGiven a set of surface normals, we pose a Manhattan Frame (MF) estimation problem as a consensus set maximization that maximizes the number of inliers over the rotation search space. We solve this problem through a branchand-bound framework, which mathematically guarantees a globally optimal solution. However, the computational time of conventional branch-and-bound algorithms are intractable for real-time performance. In this paper, we propose a novel bound computation method within an efficient measurement domain for MF estimation, i.e., the extended Gaussian image (EGI). By relaxing the original problem, we can compute the bounds in real-time, while preserving global optimality. Furthermore, we quantitatively and qualitatively demonstrate the performance of the proposed method for synthetic and real-world data. We also show the versatility of our approach through two applications: extension to multiple MF estimation and video stabilization. Kyungdon Joo, Tae-Hyun Oh, Junsik Kim 0001, In-So Kweon |
CVPR | 2 |
| 2016 | A Pseudo-Bayesian Algorithm for Robust PCAabstractCommonly used in many applications, robust PCA represents an algorithmic attempt to reduce the sensitivity of classical PCA to outliers. The basic idea is to learn a decomposition of some data matrix of interest into low rank and sparse components, the latter representing unwanted outliers. Although the resulting problem is typically NP-hard, convex relaxations provide a computationally-expedient alternative with theoretical support. However, in practical regimes performance guarantees break down and a variety of non-convex alternatives, including Bayesian-inspired models, have been proposed to boost estimation quality. Unfortunately though, without additional a priori knowledge none of these methods can significantly expand the critical operational range such that exact principal subspace recovery is possible. Into this mix we propose a novel pseudo-Bayesian algorithm that explicitly compensates for design weaknesses in many existing non-convex approaches leading to state-of-the-art performance with a sound analytical foundation. Tae-Hyun Oh, Yasuyuki Matsushita, In-So Kweon, David P. Wipf |
NIPS | 1 |
| 2016 | Partial Sum Minimization of Singular Values in Robust PCA: Algorithm and ApplicationsabstractRobust Principal Component Analysis (RPCA) via rank minimization is a powerful tool for recovering underlying low-rank structure of clean data corrupted with sparse noise/outliers. In many low-level vision problems, not only it is known that the underlying structure of clean data is low-rank, but the exact rank of clean data is also known. Yet, when applying conventional rank minimization for those problems, the objective function is formulated in a way that does not fully utilize a priori target rank information about the problems. This observation motivates us to investigate whether there is a better alternative solution when using rank minimization. In this paper, instead of minimizing the nuclear norm, we propose to minimize the partial sum of singular values, which implicitly encourages the target rank constraint. Our experimental analyses show that, when the number of samples is deficient, our approach leads to a higher success rate than conventional rank minimization, while the solutions obtained by the two approaches are almost identical when the number of samples is more than sufficient. We apply our approach to various low-level vision problems, e.g., high dynamic range imaging, motion edge detection, photometric stereo, image alignment and recovery, and show that our results outperform those obtained by the conventional nuclear norm rank minimization method. Tae-Hyun Oh, Yu-Wing Tai, Jean-Charles Bazin, Hyeongwoo Kim, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | A Multi-view Structured-Light System for Highly Accurate 3D ModelingabstractWe present a multi-view structured-light system which uses geometric and photometric information observed at multiple viewpoints. In our method, a highly accurate geometry reconstruction is achieved by exploiting our multi-view constraint on unwrapped phase images instead of using photo-consistency between intensity images. Also to express fine 3D information, we directly use photometric normal to infer non-linear shape between each point spacing, rather than fusing it by refining the geometry. Our system is built upon the off-the-shelf structured-light system, so that the proposed algorithm can be applied to existing systems including commercial 3D scanners to improve the accuracy without any system modification. We evaluate the performance of our method on both synthetic and real data, and demonstrate superior accuracy, which surpasses the accuracy of the state-of-the-art commercial products. Hyowon Ha, Tae-Hyun Oh, In-So Kweon |
3DV | 2 |
| 2015 | Fast randomized Singular Value Thresholding for Nuclear Norm MinimizationabstractRank minimization problem can be boiled down to either Nuclear Norm Minimization (NNM) or Weighted NNM (WNNM) problem. The problems related to NNM (or WNNM) can be solved iteratively by applying a closed-form proximal operator, called Singular Value Thresholding (SVT) (or Weighted SVT), but they suffer from high computational cost to compute a Singular Value Decomposition (SVD) at each iteration. In this paper, we propose an accurate and fast approximation method for SVT, called fast randomized SVT (FRSVT), where we avoid direct computation of SVD. The key idea is to extract an approximate basis for the range of a matrix from its compressed matrix. Given the basis, we compute the partial singular values of the original matrix from a small factored matrix. While the basis approximation is the bottleneck, our method is already severalfold faster than thin SVD. By adopting a range propagation technique, we can further avoid one of the bottleneck at each iteration. Our theoretical analysis provides a stepping stone between the approximation bound of SVD and its effect to NNM via SVT. Along with the analysis, our empirical results on both quantitative and qualitative studies show our approximation rarely harms the convergence behavior of the host algorithms. We apply it and validate the efficiency of our method on various vision problems, e.g. subspace clustering, weather artifact removal, simultaneous multi-image alignment and rectification. Tae-Hyun Oh, Yasuyuki Matsushita, Yu-Wing Tai, In-So Kweon |
CVPR | 1 |
| 2015 | Line meets as-projective-as-possible image stitching with moving DLTabstractWe propose a spatially varying stitching method with line correspondences. We are motivated by the observation that point features could be spatially biased or not matched in practice, e.g., repeated textures or homogeneous regions of man-made structures. In this scenario, line matches can provide strong correspondences as well as supplement cues, such as the structure preserving property. With these advantages, we adopt a feature fusion method that combines point and line correspondences into a unified framework for spatially varying stitching. We then estimate the balancing parameter between the point and line terms using geometric error. Our experiments show accurate alignment for challenging but common cases. Kyungdon Joo, Namil Kim, Tae-Hyun Oh, In-So Kweon |
ICIP | 3 |
| 2015 | Robust High Dynamic Range Imaging by Rank MinimizationabstractThis paper introduces a new high dynamic range (HDR) imaging algorithm which utilizes rank minimization. Assuming a camera responses linearly to scene radiance, the input low dynamic range (LDR) images captured with different exposure time exhibit a linear dependency and form a rank-1 matrix when stacking intensity of each corresponding pixel together. In practice, misalignments caused by camera motion, presences of moving objects, saturations and image noise break the rank-1 structure of the LDR images. To address these problems, we present a rank minimization algorithm which simultaneously aligns LDR images and detects outliers for robust HDR generation. We evaluate the performances of our algorithm systematically using synthetic examples and qualitatively compare our results with results from the state-of-the-art HDR algorithms using challenging real world examples. Tae-Hyun Oh, Joon-Young Lee, Yu-Wing Tai, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | An Autonomous Driving System for Unknown Environments Using a Unified MapabstractRecently, there have been significant advances in self-driving cars, which will play key roles in future intelligent transportation systems. In order for these cars to be successfully deployed on real roads, they must be able to autonomously drive along collision-free paths while obeying traffic laws. In contrast to many existing approaches that use prebuilt maps of roads and traffic signals, we propose algorithms and systems using Unified Map built with various onboard sensors to detect obstacles, other cars, traffic signs, and pedestrians. The proposed map contains not only the information on real obstacles nearby but also traffic signs and pedestrians as virtual obstacles. Using this map, the path planner can efficiently find paths free from collisions while obeying traffic laws. The proposed algorithms were implemented on a commercial vehicle and successfully validated in various environments, including the 2012 Hyundai Autonomous Ground Vehicle Competition. Inwook Shim, Seunghak Shin, Tae-Hyun Oh, Unghui Lee, Byungtae Ahn, Dong-Geol Choi, David Hyunchul Shim, In-So Kweon |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2014 | Balanced optical flow refinement by bidirectional constraintabstractWe present an efficient optical flow refinement approach based on a bidirectional flow consistency. Our method is an add-on component that improves the existing optical flow estimation to be balanced between forward and backward flows. Most of the state-of-the-art optical flow methods only consider unidirectional motion vectors from a source image to a target image, which can make the estimated flow inconsistent with its backward estimation. The inconsistency can be reduced by considering the bidirectional motion when the optical flow is estimated, but it would be very hard for most of the typical optical flow methods and impossible for some of them. To solve this problem, we propose a sampling-based optimization method for efficiently refining the optical flows with a bidirectional constraint. By evaluating on Middle-bury benchmark and public large displacement datasets, we validate the effectiveness of our method quantitatively and qualitatively and for accuracy. Hyeongwoo Kim, Tae-Hyun Oh, In-So Kweon |
ICIP | 3 |
| 2014 | Cost-aware depth map estimation for Lytro cameraabstractSince commercial light field cameras became available, the light field camera has aroused much interest from computer vision and image processing communities due to its versatile functions. Most of its special features are based on an estimated depth map, so reliable depth estimation is a crucial step. However, estimating depth on real light field cameras is a challenging problem due to noise and short baselines among sub-aperture images. We propose a depth map estimation method for light field cameras by exploiting correspondence and focus cues. We aggregate costs among all the sub-aperture images on cost volume to alleviate noise effects. With efficiency of the cost volume, cost-aware depth estimation is quickly achieved by discrete-continuous optimization. In addition, we analyze each property of correspondence and focus cues and utilize them to select reliable anchor points. A well reconstructed initial depth map from the anchors is shown to enhance convergence. We show our method outperforms the state-of-the-art methods by validating it on real datasets acquired with a Lytro camera. Min-Jung Kim 0001, Tae-Hyun Oh, In-So Kweon |
ICIP | 2 |
| 2013 | Partial Sum Minimization of Singular Values in RPCA for Low-Level VisionabstractRobust Principal Component Analysis (RPCA) via rank minimization is a powerful tool for recovering underlying low-rank structure of clean data corrupted with sparse noise/outliers. In many low-level vision problems, not only it is known that the underlying structure of clean data is low-rank, but the exact rank of clean data is also known. Yet, when applying conventional rank minimization for those problems, the objective function is formulated in a way that does not fully utilize a priori target rank information about the problems. This observation motivates us to investigate whether there is a better alternative solution when using rank minimization. In this paper, instead of minimizing the nuclear norm, we propose to minimize the partial sum of singular values. The proposed objective function implicitly encourages the target rank constraint in rank minimization. Our experimental analyses show that our approach performs better than conventional rank minimization when the number of samples is deficient, while the solutions obtained by the two approaches are almost identical when the number of samples is more than sufficient. We apply our approach to various low-level vision problems, e.g. high dynamic range imaging, photometric stereo and image alignment, and show that our results outperform those obtained by the conventional nuclear norm rank minimization method. Tae-Hyun Oh, Hyeongwoo Kim, Yu-Wing Tai, Jean-Charles Bazin, In-So Kweon |
ICCV | 1 |
| 2013 | Hierarchical 3D line restoration based on angular proximity in structured environmentsabstractWe present a method based on a hierarchical clustering to restore the 3D lines of structured environments. In previous approaches, the restoration of noisy 3D lines is a challenging problem because it is difficult to define a suitable similarity measure discriminative to other lines. Our motivation to overcome the difficulty is that most structured scenes consist of sets of parallel 3D lines with the same angular proximity, which provides a hierarchical similarity measure for structured 3D lines. Accordingly, our restoration method works in a manner that clustering is hierarchically performed on angular and distance levels. The 3D line restoration is then achieved by finding the center of each cluster. The framework also makes the clustered 3D lines align along the associated angular directions. We compare the proposed algorithm with methods using no knowledge of the angular information, and demonstrate its effectiveness through real-world experiments. Kyungdon Joo, Tae-Hyun Oh, Hyeongwoo Kim, In-So Kweon |
ICIP | 2 |
| 2013 | High dynamic range imaging by a rank-1 constraintabstractWe present a high dynamic range (HDR) imaging algorithm that utilizes a modern rank minimization framework. Linear dependency exists among low dynamic range (LDR) images. However, global or local misalignment by camera motion and moving objects breaks down the low-rank structure of LDR images. The proposed algorithm simultaneously estimates global geometric transforms to align LDR images and detects moving objects and under-/over-exposed regions using a rank minimization approach. In the HDR composition step, structural consistency weighting is proposed to generate an artifact-free HDR image from an user-selected reference image. We demonstrate the robustness and effectiveness of the proposed method with real datasets. Tae-Hyun Oh, Joon-Young Lee, In-So Kweon |
ICIP | 1 |
| 2012 | A Tensor Voting Approach for Multi-view 3D Scene Flow Estimation and Refinement
Jaesik Park, Tae-Hyun Oh, Jiyoung Jung, Yu-Wing Tai, In-So Kweon |
ECCV (4) | 2 |
| 2012 | Real-time motion detection based on Discrete Cosine TransformabstractWe present a motion detection algorithm by a change detection filter matrix derived from Discrete Cosine Transform. Recently, a Fourier reconstruction scheme shows good results for motion detection. However, its computational cost is a major drawback. We revisit the problem and achieve two orders of magnitude faster than the previous algorithm with better performance. The proposed algorithm runs at about 800 frames per second for VGA resolution images on a consumer hardware by using only integer matrix multiplication and the symmetric property of the change detection filter matrix. In addition, our algorithm is fundamentally robust to sudden illumination changes because it works based on edge information. We verify our algorithm with challenging datasets that contain strong and sudden illumination changes. Tae-Hyun Oh, Joon-Young Lee, In-So Kweon |
ICIP | 1 |
| 2012 | Autonomous homing based on laser-camera fusion systemabstractBuilding maps of unknown environments is a critical factor for autonomous navigation and homing, and this problem is especially challenging in large-scale environments. Recently, sensor fusion systems such as combinations of cameras and laser sensors have become popular in the effort to ensure a general level of performance in this task. In this paper, we present a new homing method in a large-scale environment using a laser-camera fusion system. Instead of fusing data to form a single map builder, we adaptively select sensor data to handle environments which contain ambiguity. For autonomous homing, we propose a new mapping strategy for building a hybrid map and a return strategy for selecting the next target waypoints efficiently. The experimental results demonstrate that the proposed algorithm enables the autonomous homing of a robot in a large-scale indoor environments in real time. Dong-Geol Choi, Inwook Shim, Yunsu Bok, Tae-Hyun Oh, In-So Kweon |
IROS | 4 |