EDBT 2026 Demo / reviewers in the wild / expert
Changil Kim 0001
dblp:94/10569
· DBLP profile ↗
35ranked-venue papers
5as first author
21since 2021 · last 2025
0000-0002-6541-8825ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 5 first-author · 20 since 2021Artificial intelligence and machine learning · 23 · 1 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Textured Gaussians for Enhanced 3D Scene Appearance Modelingabstract3D Gaussian Splatting (3DGS) has emerged as the state-of-the-art 3D reconstruction technique, offering high-quality results with fast training and rendering. However, its expressivity is limited as pixels covered by the same Gaussian share identical colors aside from a Gaussian falloff scaling factor, and individual Gaussians can only represent simple ellipsoids geometrically. To overcome these limitations, we integrate texture and alpha mapping from traditional graphics with 3DGS. Our approach augments each Gaussian with alpha, RGB, or RGBA texture maps to model spatially varying color and opacity across each Gaussian’s extent. This allows Gaussians to represent richer texture patterns and geometric structures beyond single-color ellipsoids. Notably, alpha-only texture maps significantly improve Gaussian expressivity, while further augmenting with RGB texture maps achieve maximum expressivity. We validate our method on a wide variety of standard benchmark datasets and our own custom captures at both the object and scene levels, and demonstrate image quality improvements over existing methods while using a similar or lower number of Gaussians. Brian Chao, Hung-Yu Tseng, Lorenzo Porzi, Chen Gao 0003, Tuotuo Li, Qinbo Li, Ayush Saraf, Jia-Bin Huang 0001, Johannes Kopf 0001, Gordon Wetzstein, Changil Kim 0001 |
CVPR | 11 |
| 2025 | Geometry-guided Online 3D Video Synthesis with Multi-View Temporal ConsistencyabstractWe introduce a novel geometry-guided online video view synthesis method with enhanced view and temporal consistency. Traditional approaches achieve high-quality synthesis from dense multi-view camera setups but require significant computational resources. In contrast, selective-input methods reduce this cost but often compromise quality, leading to multi-view and temporal inconsistencies such as flickering artifacts. Our method addresses this challenge to deliver efficient, high-quality novel-view synthesis with view and temporal consistency. The key innovation of our approach lies in using global geometry to guide an image-based rendering pipeline. To accomplish this, we progressively refine depth maps using color difference masks across time. These depth maps are then accumulated through truncated signed distance fields in the synthesized view’s image space. This depth representation is view and temporally consistent, and is used to guide a pre-trained blending network that fuses multiple forward-rendered input-view images. Thus, the network is encouraged to output geometrically consistent synthesis results across multiple views and time. Our approach achieves consistent, high-quality video synthesis, while running efficiently in an online manner. Hyunho Ha, Lei Xiao 0014, Christian Richardt, Thu Nguyen-Phuoc, Changil Kim 0001, Min H. Kim 0001, Douglas Lanman, Numair Khan |
CVPR | 5 |
| 2025 | IRIS: Inverse Rendering of Indoor Scenes from Low Dynamic Range ImagesabstractInverse rendering seeks to recover 3D geometry, surface material, and lighting from captured images, enabling advanced applications such as novel-view synthesis, relighting, and virtual object insertion. However, most existing techniques rely on high dynamic range (HDR) images as input, limiting accessibility for general users. In response, we introduce IRIS, an inverse rendering framework that recovers the physically based material, spatially-varying HDR lighting, and camera response functions from multi-view, low-dynamic-range (LDR) images. By eliminating the dependence on HDR input, we make inverse rendering technology more accessible. We evaluate our approach on real-world and synthetic scenes and compare it with state-of-the-art methods. Our results show that IRIS effectively recovers HDR lighting, accurate material, and plausible camera response functions, supporting photorealistic relighting and object insertion. Chih-Hao Lin, Jia-Bin Huang 0001, Zhengqin Li, Zhao Dong 0001, Christian Richardt, Tuotuo Li, Michael Zollhöfer, Johannes Kopf 0001, Shenlong Wang, Changil Kim 0001 |
CVPR | 10 |
| 2024 | LTM: Lightweight Textured Mesh Extraction and Refinement of Large Unbounded Scenes for Efficient Storage and Real-Time RenderingabstractAdvancements in neural signed distance fields (SDFs) have enabled modeling 3D surface geometry from a set of 2D images of real-world scenes. Baking neural SDFs can extract explicit mesh with appearance baked into texture maps as neural features. The baked meshes still have a large memory footprint and require a powerful GPU for real-time rendering. Neural optimization of such large meshes with differentiable rendering pose significant challenges. We propose a method to produce optimized meshes for large unbounded scenes with low triangle budget and high fidelity of geometry and appearance. We achieve this by combining advancements in baking neural SDFs with classical mesh simplification techniques and proposing a joint appearance-geometry refinement step. The visual quality is comparable to or better than state-of-the-art neural meshing and baking methods with high geometric accuracy despite significant reduction in triangle count, making the produced meshes efficient for storage, transmission, and rendering on mobile hardware. We validate the effectiveness of the proposed method on large unbounded scenes from mip-NeRF 360, Tanks & Temples, and Deep Blending datasets, achieving at-par rendering quality with 73 x reduced triangles and 11 x reduction in memory footprint. Rajvi Shah, Qinbo Li, Yipeng Wang 0018, Ayush Saraf, Changil Kim 0001, Jia-Bin Huang 0001, Dinesh Manocha, Suhib Alsisan, Johannes Kopf 0001 |
CVPR | 6 |
| 2024 | SpecNeRF: Gaussian Directional Encoding for Specular ReflectionsabstractNeural radiance fields have achieved remarkable performance in modeling the appearance of 3D scenes. However, existing approaches still struggle with the view-dependent appearance of glossy surfaces, especially under complex lighting of indoor environments. Unlike existing methods, which typically assume distant lighting like an environment map, we propose a learnable Gaussian directional encoding to better model the view-dependent effects under near-field lighting conditions. Importantly, our new directional encoding captures the spatially-varying nature of near-field lighting and emulates the behavior of prefiltered environment maps. As a result, it enables the efficient evaluation of preconvolved specular color at any 3D location with varying roughness coefficients. We further introduce a data-driven geometry prior that helps alleviate the shape radiance ambiguity in reflection modeling. We show that our Gaussian directional encoding and geometry prior significantly improve the modeling of challenging specular reflections in neural radiance fields, which helps decompose appearance into more physically meaningful components. Vasu Agrawal, Haithem Turki, Changil Kim 0001, Chen Gao 0003, Pedro V. Sander, Michael Zollhöfer, Christian Richardt |
CVPR | 4 |
| 2024 | TextureDreamer: Image-Guided Texture Synthesis through Geometry-Aware DiffusionabstractWe present TextureDreamer, a novel image-guided texture synthesis method to transfer relightable textures from a small number of input images (3 to 5) to target 3D shapes across arbitrary categories. Texture creation is a pivotal challenge in vision and graphics. Industrial companies hire experienced artists to manually craft textures for 3D assets. Classical methods require densely sampled views and ac-curately aligned geometry, while learning-based methods are confined to category-specific shapes within the dataset. In contrast, TextureDreamer can transfer highly detailed, intricate textures from real-world environments to arbi-trary objects with only a few casually captured images, po-tentially significantly democratizing texture creation. Our core idea, personalized geometry-aware score distillation (PGSD), draws inspiration from recent advancements in diffuse models, including personalized modeling for texture information extraction, score distillation for detailed appearance synthesis, and explicit geometry guidance with ControlNet. Our integration and several essential modifications substantially improve the texture quality. Experiments on real images spanning different categories show that TextureDreamer can successfully transfer highly realistic, se-mantic meaningful texture to arbitrary objects, surpassing the visual quality of previous state-of-the-art. Project page: https://texturedreamer.github.io Yu-Ying Yeh, Jia-Bin Huang 0001, Changil Kim 0001, Lei Xiao 0014, Thu Nguyen-Phuoc, Numair Khan, Manmohan Krishna Chandraker, Carl S. Marshall, Zhao Dong 0001, Zhengqin Li |
CVPR | 3 |
| 2024 | Taming Latent Diffusion Model for Neural Radiance Field Inpainting
Chieh Hubert Lin, Changil Kim 0001, Jia-Bin Huang 0001, Qinbo Li, Chih-Yao Ma, Johannes Kopf 0001, Ming-Hsuan Yang 0001, Hung-Yu Tseng |
ECCV (3) | 2 |
| 2024 | Planar Reflection-Aware Neural Radiance Fields
Chen Gao 0003, Yipeng Wang 0018, Changil Kim 0001, Jia-Bin Huang 0001, Johannes Kopf 0001 |
SIGGRAPH Asia | 3 |
| 2023 | Robust Dynamic Radiance FieldsabstractDynamic radiance field reconstruction methods aim to model the time-varying structure and appearance of a dynamic scene. Existing methods, however, assume that accurate camera poses can be reliably estimated by Structure from Motion (SfM) algorithms. These methods, thus, are unreliable as SfM algorithms often fail or produce erroneous poses on challenging videos with highly dynamic objects, poorly textured surfaces, and rotating camera motion. We address this robustness issue by jointly estimating the static and dynamic radiance fields along with the camera parameters (poses and focal length). We demonstrate the robustness of our approach via extensive quantitative and qualitative experiments. Our results show favorable performance over the state-of-the-art dynamic view synthesis methods. Yu-Lun Liu 0001, Chen Gao 0003, Andreas Meuleman, Hung-Yu Tseng, Ayush Saraf, Changil Kim 0001, Yung-Yu Chuang, Johannes Kopf 0001, Jia-Bin Huang 0001 |
CVPR | 6 |
| 2023 | HyperReel: High-Fidelity 6-DoF Video with Ray-Conditioned SamplingabstractVolumetric scene representations enable photorealistic view synthesis for static scenes and form the basis of several existing 6-DoF video techniques. However, the volume rendering procedures that drive these representations necessitate careful trade-offs in terms of quality, rendering speed, and memory efficiency. In particular, existing methods fail to simultaneously achieve real-time performance, small memory footprint, and high-quality rendering for challenging real-world scenes. To address these issues, we present HyperReel―a novel 6-DoF video representation. The two core components of HyperReel are: (1) a ray-conditioned sample prediction network that enables high-fidelity, high frame rate rendering at high resolutions and (2) a compact and memory-efficient dynamic volume representation. Our 6-DoF video pipeline achieves the best performance compared to prior and contemporary approaches in terms of visual quality with small memory requirements, while also rendering at up to 18 frames-per-second at megapixel resolution without any custom CUDA code. Benjamin Attal, Jia-Bin Huang 0001, Christian Richardt, Michael Zollhöfer, Johannes Kopf 0001, Matthew O'Toole, Changil Kim 0001 |
CVPR | 7 |
| 2023 | Progressively Optimized Local Radiance Fields for Robust View SynthesisabstractWe present an algorithm for reconstructing the radiance field of a large-scale scene from a single casually captured video. The task poses two core challenges. First, most existing radiance field reconstruction approaches rely on accurate pre-estimated camera poses from Structure-from-Motion algorithms, which frequently fail on in-the-wild videos. Second, using a single, global radiance field with finite representational capacity does not scale to longer trajectories in an unbounded scene. For handling unknown poses, we jointly estimate the camera poses with radiance field in a progressive manner. We show that progressive optimization significantly improves the robustness of the reconstruction. For handling large unbounded scenes, we dynamically allocate new local radiance fields trained with frames within a temporal window. This further improves robustness (e.g., performs well even under moderate pose drifts) and allows us to scale to large scenes. Our extensive evaluation on the TANKS AND TEMPLES dataset and our collected outdoor dataset, STATIC HIKES, show that our approach compares favorably with the state-of-the-art. Andreas Meuleman, Yu-Lun Liu 0001, Chen Gao 0003, Jia-Bin Huang 0001, Changil Kim 0001, Min H. Kim 0001, Johannes Kopf 0001 |
CVPR | 5 |
| 2023 | Consistent View Synthesis with Pose-Guided Diffusion ModelsabstractNovel view synthesis from a single image has been a cornerstone problem for many Virtual Reality applications that provide immersive experiences. However, most existing techniques can only synthesize novel views within a limited range of camera motion or fail to generate consistent and high-quality novel views under significant camera movement. In this work, we propose a pose-guided diffusion model to generate a consistent long-term video of novel views from a single image. We design an attention layer that uses epipolar lines as constraints to facilitate the association between different viewpoints. Experimental results on synthetic and real-world datasets demonstrate the effectiveness of the proposed diffusion model against state-of-the-art transformer-based and GAN-based approaches. More qualitative results are available at https://poseguided-diffusion.github.io/. Hung-Yu Tseng, Qinbo Li, Changil Kim 0001, Suhib Alsisan, Jia-Bin Huang 0001, Johannes Kopf 0001 |
CVPR | 3 |
| 2023 | OmnimatteRF: Robust Omnimatte with 3D Background ModelingabstractVideo matting has broad applications, from adding interesting effects to casually captured movies to assisting video production professionals. Matting with associated effects such as shadows and reflections has also attracted increasing research activity, and methods like Omnimatte have been proposed to separate dynamic foreground objects of interest into their own layers. However, prior works represent video backgrounds as 2D image layers, limiting their capacity to express more complicated scenes, thus hindering application to real-world videos. In this paper, we propose a novel video matting method, OmnimatteRF, that combines dynamic 2D foreground layers and a 3D background model. The 2D layers preserve the details of the subjects, while the 3D background robustly reconstructs scenes in real-world videos. Extensive experiments demonstrate that our method reconstructs scenes with better quality on various videos. Geng Lin, Chen Gao 0003, Jia-Bin Huang 0001, Changil Kim 0001, Yipeng Wang 0018, Matthias Zwicker, Ayush Saraf |
ICCV | 4 |
| 2023 | Single-Image 3D Human Digitization with Shape-guided DiffusionabstractWe present an approach to generate a 360-degree view of a person with a consistent, high-resolution appearance from a single input image. NeRF and its variants typically require videos or images from different viewpoints. Most existing approaches taking monocular input either rely on ground-truth 3D scans for supervision or lack 3D consistency. While recent 3D generative models show promise of 3D consistent human digitization, these approaches do not generalize well to diverse clothing appearances, and the results lack photorealism. Unlike existing work, we utilize high-capacity 2D diffusion models pretrained for general image synthesis tasks as an appearance prior of clothed humans. To achieve better 3D consistency while retaining the input identity, we progressively synthesize multiple views of the human in the input image by inpainting missing regions with shape-guided diffusion conditioned on silhouette and surface normal. We then fuse these synthesized multi-view images via inverse rendering to obtain a fully textured high-resolution 3D mesh of the given person. Experiments show that our approach outperforms prior methods and achieves photorealistic 360-degree synthesis of a wide range of clothed humans with complex textures from a single image. Badour AlBahar, Shunsuke Saito, Hung-Yu Tseng, Changil Kim 0001, Johannes Kopf 0001, Jia-Bin Huang 0001 |
SIGGRAPH Asia | 4 |
| 2023 | VR-NeRF: High-Fidelity Virtualized Walkable SpacesabstractWe present an end-to-end system for the high-fidelity capture, model reconstruction, and real-time rendering of walkable spaces in virtual reality using neural radiance fields. To this end, we designed and built a custom multi-camera rig to densely capture walkable spaces in high fidelity and with multi-view high dynamic range images in unprecedented quality and density. We extend instant neural graphics primitives with a novel perceptual color space for learning accurate HDR appearance, and an efficient mip-mapping mechanism for level-of-detail rendering with anti-aliasing, while carefully optimizing the trade-off between quality and speed. Our multi-GPU renderer enables high-fidelity volume rendering of our neural radiance field model at the full VR resolution of dual 2K × 2K at 36 Hz on our custom demo machine. We demonstrate the quality of our results on our challenging high-fidelity datasets, and compare our method and datasets to existing baselines. We release our dataset on our project website: https://vr-nerf.github.io. Linning Xu, Vasu Agrawal, William Laney, Tony Garcia, Aayush Bansal, Changil Kim 0001, Samuel Rota Bulò, Lorenzo Porzi, Peter Kontschieder, Aljaz Bozic, Dahua Lin, Michael Zollhöfer, Christian Richardt |
SIGGRAPH Asia | 6 |
| 2022 | Learning Neural Light Fields with Ray-Space EmbeddingabstractNeural radiance fields (NeRFs) produce state-of-the-art view synthesis results, but are slow to render, requiring hundreds of network evaluations per pixel to approximate a volume rendering integral. Baking NeRFs into explicit data structures enables efficient rendering, but results in large memory footprints and, in some cases, quality reduction. Additionally, volumetric representations for view synthesis often struggle to represent challenging view dependent effects such as distorted reflections and refractions. We present a novel neural light field representation that, in contrast to prior work, is fast, memory efficient, and excels at modeling complicated view dependence. Our method supports rendering with a single network evaluation per pixel for small baseline light fields and with only a few evaluations per pixel for light fields with larger baselines. At the core of our approach is a ray-space embedding network that maps 4D ray-space into an intermediate, interpolable latent space. Our method achieves state-of-the-art quality on dense forward-facing datasets such as the Stanford Light Field dataset. In addition, for forward-facing scenes with sparser inputs we achieve results that are competitive with NeRF-based approaches while providing a better speed/quality/memory trade-off with far fewer network evaluations. Benjamin Attal, Jia-Bin Huang 0001, Michael Zollhöfer, Johannes Kopf 0001, Changil Kim 0001 |
CVPR | 5 |
| 2022 | Neural 3D Video Synthesis from Multi-view VideoabstractWe propose a novel approach for 3D video synthesis that is able to represent multi-view video recordings of a dynamic real-world scene in a compact, yet expressive representation that enables high-quality view synthesis and motion interpolation. Our approach takes the high quality and compactness of static neural radiance fields in a new direction: to a model-free, dynamic setting. At the core of our approach is a novel time-conditioned neural radiance field that represents scene dynamics using a set of compact latent codes. We are able to significantly boost the training speed and perceptual quality of the generated imagery by a novel hierarchical training scheme in combination with ray importance sampling. Our learned representation is highly compact and able to represent a 10 second 30 FPS multi-view video recording by 18 cameras with a model size of only 28MB. We demonstrate that our method can render high-fidelity wide-angle novel views at over 1K resolution, even for complex and dynamic scenes. We perform an extensive qualitative and quantitative evaluation that shows that our approach outperforms the state of the art. Project website: https://neural-3d-video.github.io/. Tianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green, Christoph Lassner, Changil Kim 0001, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard A. Newcombe, Zhaoyang Lv |
CVPR | 6 |
| 2022 | Boosting View Synthesis with Residual TransferabstractVolumetric view synthesis methods with neural representations, such as NeRF and NeX, have recently demonstrated high-quality novel view synthesis. However, optimizing these representations is slow, and even fully trained models cannot reproduce all fine details in the input views. We present a simple but effective technique to boost the rendering quality, which can be easily integrated with most view synthesis methods. The core idea is to transfer color resid-uals (the difference between the input images and their re-construction) from training views to novel views. We blend the residuals from multiple views using a heuristic weighting scheme depending on ray visibility and angular differ-ences. We integrate our technique with several state-of-the-art view synthesis methods and evaluate the Real Forward-facing and the Shiny datasets. Our results show that at about 1/10th the number of training iterations, we achieve the same rendering quality as fully converged NeRF and NeX models, and when applied to fully converged models, we significantly improve their rendering quality. Xuejian Rong, Jia-Bin Huang 0001, Ayush Saraf, Changil Kim 0001, Johannes Kopf 0001 |
CVPR | 4 |
| 2021 | AMICO: Amodal Instance Composition
Peiye Zhuang, Denis Demandolx, Ayush Saraf, Xuejian Rong, Changil Kim 0001, Jia-Bin Huang 0001 |
BMVC | 5 |
| 2021 | Space-Time Neural Irradiance Fields for Free-Viewpoint VideoabstractWe present a method that learns a spatiotemporal neural irradiance field for dynamic scenes from a single video. Our learned representation enables free-viewpoint rendering of the input video. Our method builds upon recent advances in implicit representations. Learning a spatiotemporal irradiance field from a single video poses significant challenges because the video contains only one observation of the scene at any point in time. The 3D geometry of a scene can be legitimately represented in numerous ways since varying geometry (motion) can be explained with varying appearance and vice versa. We address this ambiguity by constraining the time-varying geometry of our dynamic scene representation using the scene depth estimated from video depth estimation methods, aggregating contents from individual frames into a single global representation. We provide an extensive quantitative evaluation and demonstrate compelling free-viewpoint rendering results. Wenqi Xian, Jia-Bin Huang 0001, Johannes Kopf 0001, Changil Kim 0001 |
CVPR | 4 |
| 2021 | TöRF: Time-of-Flight Radiance Fields for Dynamic Scene View SynthesisabstractNeural networks can represent and accurately reconstruct radiance fields for static 3D scenes (e.g., NeRF). Several works extend these to dynamic scenes captured with monocular video, with promising performance. However, the monocular setting is known to be an under-constrained problem, and so methods rely on data-driven priors for reconstructing dynamic content. We replace these priors with measurements from a time-of-flight (ToF) camera, and introduce a neural representation based on an image formation model for continuous-wave ToF cameras. Instead of working with processed depth maps, we model the raw ToF sensor measurements to improve reconstruction quality and avoid issues with low reflectance regions, multi-path interference, and a sensor's limited unambiguous depth range. We show that this approach improves robustness of dynamic scene reconstruction to erroneous calibration and large motions, and discuss the benefits and limitations of integrating RGB+ToF sensors now available on modern smartphones. Benjamin Attal, Eliot Laidlaw, Aaron Gokaslan, Changil Kim 0001, Christian Richardt, James Tompkin 0001, Matthew O'Toole |
NeurIPS | 4 |
| 2019 | Speech2Face: Learning the Face Behind a VoiceabstractHow much can we infer about a person’s looks from the way they speak? In this paper, we study the task of reconstructing a facial image of a person from a short audio recording of that person speaking. We design and train a deep neural network to perform this task using millions of natural Internet/Youtube videos of people speaking. During training, our model learns voice-face correlations that allow it to produce images that capture various physical attributes of the speakers such as age, gender and ethnicity. This is done in a self-supervised manner, by utilizing the natural co-occurrence of faces and speech in Internet videos, without the need to model attributes explicitly. We evaluate and numerically quantify how–-and in what manner–-our Speech2Face reconstructions, obtained directly from audio, resemble the true face images of the speakers. Tae-Hyun Oh, Tali Dekel, Changil Kim 0001, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Wojciech Matusik |
CVPR | 3 |
| 2018 | On Learning Associations of Faces and Voices
Changil Kim 0001, Hijung Shin, Tae-Hyun Oh, Alexandre Kaspar, Mohamed A. Elgharib, Wojciech Matusik |
ACCV (5) | 1 |
| 2018 | Crowd-Guided Ensembles: How Can We Choreograph Crowd Workers for Video Segmentation?abstractIn this work, we propose two ensemble methods leveraging a crowd workforce to improve video annotation, with a focus on video object segmentation. Their shared principle is that while individual candidate results may likely be insufficient, they often complement each other so that they can be combined into something better than any of the individual results---the very spirit of collaborative working. For one, we extend a standard polygon-drawing interface to allow workers to annotate negative space, and combine the work of multiple workers instead of relying on a single best one as commonly done in crowdsourced image segmentation. For the other, we present a method to combine multiple automatic propagation algorithms with the help of the crowd. Such combination requires an understanding of where the algorithms fail, which we gather using a novel coarse scribble video annotation task. We evaluate our ensemble methods, discuss our design choices for them, and make our web-based crowdsourcing tools and results publicly available. Alexandre Kaspar, Genevieve Patterson, Changil Kim 0001, Yagiz Aksoy, Wojciech Matusik, Mohamed A. Elgharib |
CHI | 3 |
| 2018 | A Dataset of Flash and Ambient Illumination Pairs from the Crowd
Yagiz Aksoy, Changil Kim 0001, Petr Kellnhofer, Sylvain Paris, Mohamed A. Elgharib, Marc Pollefeys, Wojciech Matusik |
ECCV (9) | 2 |
| 2018 | Learning-Based Video Motion Magnification
Tae-Hyun Oh, Ronnachai Jaroensri, Changil Kim 0001, Mohamed A. Elgharib, Frédo Durand, William T. Freeman, Wojciech Matusik |
ECCV (4) | 3 |
| 2018 | Deep multispectral painting reproduction via multi-layer, custom-ink printingabstractWe propose a workflow for spectral reproduction of paintings, which captures a painting's spectral color, invariant to illumination, and reproduces it using multi-material 3D printing. We take advantage of the current 3D printers' capabilities of combining highly concentrated inks with a large number of layers, to expand the spectral gamut of a set of inks. We use a data-driven method to both predict the spectrum of a printed ink stack and optimize for the stack layout that best matches a target spectrum. This bidirectional mapping is modeled using a pair of neural networks, which are optimized through a problem-specific multi-objective loss function. Our loss function helps find the best possible ink layout resulting in the balance between spectral reproduction and colorimetric accuracy under a multitude of illuminants. In addition, we introduce a novel spectral vector error diffusion algorithm based on combining color contoning and halftoning, which simultaneously solves the layout discretization and color quantization problems, accurately and efficiently. Our workflow outperforms the state-of-the-art models for spectral prediction and layout optimization. We demonstrate reproduction of a number of real paintings and historically important pigments using our prototype implementation that uses 10 custom inks with varying spectra and a resin-based 3D printer. Liang Shi 0003, Vahid Babaei, Changil Kim 0001, Michael Foshey, Yuanming Hu, Pitchaya Sitthi-amorn, Szymon Rusinkiewicz, Wojciech Matusik |
ACM Trans. Graph. | 3 |
| 2017 | Video Reflection Removal Through Spatio-Temporal OptimizationabstractReflections can obstruct content during video capture and hence their removal is desirable. Current removal techniques are designed for still images, extracting only one reflection (foreground) and one background layer from the input. When extended to videos, unpleasant artifacts such as temporal flickering and incomplete separation are generated. We present a technique for video reflection removal by jointly solving for motion and separation. The novelty of our work is in our optimization formulation as well as the motion initialization strategy. We present a novel spatiotemporal optimization that takes n frames as input and directly estimates 2n frames as output, n for each layer. We aim to fully utilize spatio-temporal information in our objective terms. Our motion initialization is based on iterative frame-to-frame alignment instead of the direct alignment used by current approaches. We compare against advanced video extensions of the state of the art, and we significantly reduce temporal flickering and improve separation. In addition, we reduce image blur and recover moving objects more accurately. We validate our approach through subjective and objective evaluations on real and controlled data. Ajay Nandoriya, Mohamed A. Elgharib, Changil Kim 0001, Mohamed Hefeeda, Wojciech Matusik |
ICCV | 3 |
| 2016 | Point Cloud Noise and Outlier Removal for Image-Based 3D ReconstructionabstractPoint sets generated by image-based 3D reconstruction techniques are often much noisier than those obtained using active techniques like laser scanning. Therefore, they pose greater challenges to the subsequent surface reconstruction (meshing) stage. We present a simple and effective method for removing noise and outliers from such point sets. Our algorithm uses the input images and corresponding depth maps to remove pixels which are geometrically or photometrically inconsistent with the colored surface implied by the input. This allows standard surface reconstruction methods (such as Poisson surface reconstruction) to perform less smoothing and thus achieve higher quality surfaces with more features. Our algorithm is efficient, easy to implement, and robust to varying amounts of noise. We demonstrate the benefits of our algorithm in combination with a variety of state-of-the-art depth and surface reconstruction methods. Katja Wolff, Changil Kim 0001, Henning Zimmer, Christopher Schroers, Mario Botsch, Olga Sorkine-Hornung, Alexander Sorkine-Hornung |
3DV | 2 |
| 2016 | Depth from Gradients in Dense Light Fields for Object ReconstructionabstractObjects with thin features and fine details are challenging for most multi-view stereo techniques, since such features occupy small volumes and are usually only visible in a small portion of the available views. In this paper, we present an efficient algorithm to reconstruct intricate objects using densely sampled light fields. At the heart of our technique lies a novel approach to compute per-pixel depth values by exploiting local gradient information in densely sampled light fields. This approach can generate accurate depth values for very thin features, and can be run for each pixel in parallel. We assess the reliability of our depth estimates using a novel two-sided photoconsistency measure, which can capture whether the pixel lies on a texture or a silhouette edge. This information is then used to propagate the depth estimates at high gradient regions to smooth parts of the views efficiently and reliably using edge-aware filtering. In the last step, the per-image depth values and color information are aggregated in 3D space using a voting scheme, allowing the reconstruction of a globally consistent mesh for the object. Our approach can process large video datasets very efficiently and at the same time generates high quality object reconstructions that compare favorably to the results of state-of-the-art multi-view stereo methods. Kaan Yücer, Changil Kim 0001, Alexander Sorkine-Hornung, Olga Sorkine-Hornung |
3DV | 2 |
| 2015 | Online view sampling for estimating depth from light fieldsabstractGeometric information such as depth obtained from light fields finds more applications recently. Where and how to sample images to populate a light field is an important problem to maximize the usability of information gathered for depth reconstruction. We propose a simple analysis model for view sampling and an adaptive, online sampling algorithm tailored to light field depth reconstruction. Our model is based on the trade-off between visibility and depth resolvability for varying sampling locations, and seeks the optimal locations that best balance the two conflicting criteria. Changil Kim 0001, Kartic Subr, Kenny Mitchell, Alexander Sorkine-Hornung, Markus Gross 0001 |
ICIP | 1 |
| 2014 | Memory Efficient Stereoscopy from Light FieldsabstractWe address the problem of stereoscopic content generation from light fields using multi-perspective imaging. Our proposed method takes as input a light field and a target disparity map, and synthesizes a stereoscopic image pair by selecting light rays that fulfill the given target disparity constraints. We formulate this as a variational convex optimization problem. Compared to previous work, our method makes use of multi-view input to composite the new view with occlusions and disocclusions properly handled, does not require any correspondence information such as scene depth, is free from undesirable artifacts such as grid bias or image distortion, and is more efficiently solvable. In particular, our method is about ten times more memory efficient than the previous art, and is capable of processing higher resolution input. This is essential to make the proposed method practically applicable to realistic scenarios where HD content is standard. We demonstrate the effectiveness of our method experimentally. Changil Kim 0001, Ulrich Muller, Henning Zimmer, Yael Pritch, Alexander Sorkine-Hornung, Markus Gross 0001 |
3DV | 1 |
| 2013 | Scalable Music: Automatic Music Retargeting and SynthesisabstractAbstract In this paper we propose a method for dynamic rescaling of music, inspired by recent works on image retargeting, video reshuffling and character animation in the computer graphics community. Given the desired target length of a piece of music and optional additional constraints such as position and importance of certain parts, we build on concepts from seam carving, video textures and motion graphs and extend them to allow for a global optimization of jumps in an audio signal. Based on an automatic feature extraction and spectral clustering for segmentation, we employ length‐constrained least‐costly path search via dynamic programming to synthesize a novel piece of music that best fulfills all desired constraints, with imperceptible transitions between reshuffled parts. We show various applications of music retargeting such as part removal, decreasing or increasing music duration, and in particular consistent joint video and audio editing. Simon Wenner, Jean-Charles Bazin, Alexander Sorkine-Hornung, Changil Kim 0001, Markus Gross 0001 |
Comput. Graph. Forum | 4 |
| 2013 | Scene reconstruction from high spatio-angular resolution light fieldsabstractThis paper describes a method for scene reconstruction of complex, detailed environments from 3D light fields. Densely sampled light fields in the order of 10 9 light rays allow us to capture the real world in unparalleled detail, but efficiently processing this amount of data to generate an equally detailed reconstruction represents a significant challenge to existing algorithms. We propose an algorithm that leverages coherence in massive light fields by breaking with a number of established practices in image-based reconstruction. Our algorithm first computes reliable depth estimates specifically around object boundaries instead of interior regions, by operating on individual light rays instead of image patches. More homogeneous interior regions are then processed in a fine-to-coarse procedure rather than the standard coarse-to-fine approaches. At no point in our method is any form of global optimization performed. This allows our algorithm to retain precise object contours while still ensuring smooth reconstructions in less detailed areas. While the core reconstruction method handles general unstructured input, we also introduce a sparse representation and a propagation scheme for reliable depth estimates which make our algorithm particularly effective for 3D input, enabling fast and memory efficient processing of "Gigaray light fields" on a standard GPU. We show dense 3D reconstructions of highly detailed scenes, enabling applications such as automatic segmentation and image-based rendering, and provide an extensive evaluation and comparison to existing image-based reconstruction techniques. Changil Kim 0001, Henning Zimmer, Yael Pritch, Alexander Sorkine-Hornung, Markus Gross 0001 |
ACM Trans. Graph. | 1 |
| 2011 | Multi-perspective stereoscopy from light fieldsabstractThis paper addresses stereoscopic view generation from a light field. We present a framework that allows for the generation of stereoscopic image pairs with per-pixel control over disparity, based on multi-perspective imaging from light fields. The proposed framework is novel and useful for stereoscopic image processing and post-production. The stereoscopic images are computed as piecewise continuous cuts through a light field, minimizing an energy reflecting prescribed parameters such as depth budget, maximum disparity gradient, desired stereoscopic baseline, and so on. As demonstrated in our results, this technique can be used for efficient and flexible stereoscopic post-processing, such as reducing excessive disparity while preserving perceived depth, or retargeting of already captured scenes to various view settings. Moreover, we generalize our method to multiple cuts, which is highly useful for content creation in the context of multi-view autostereoscopic displays. We present several results on computer-generated content as well as live-action content. Changil Kim 0001, Alexander Sorkine-Hornung, Simon Heinzle, Wojciech Matusik, Markus Gross 0001 |
ACM Trans. Graph. | 1 |