Jia-Bin Huang 0001

dblp:51/1815-1 · also Jiabin Huang 0001 · DBLP profile ↗
← Back
128ranked-venue papers
13as first author
64since 2021 · last 2026
0000-0002-0536-3658ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 104 · 6 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 103 · 13 first-author · 53 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition
abstract
Multi-shot generation requires preserving the identity of characters and settings across frames. Cinematic scene composition goes beyond standard multi-shot generation, introducing additional challenges such as expressing complex interactions among multiple characters and visual effects to convey creative narratives—challenges existing datasets cannot fully address. We present CineVerse, a large-scale dataset of diverse movie scenes labeled with shot-level annotations tailored for filmmaking. CineVerse includes refined scene descriptions, shot-type information, and newly extracted shot, character, setting descriptions. We validate our dataset by developing a baseline framework that first generates a scene plan containing detailed information for the overall scene and each individual shot, then produces a set of coherent keyframes. Our results show significant improvements in controlling and synthesizing cinematic content through the added context provided by CineVerse.
Quynh Phung, Long Mai, Fabian Caba Heilbron, Feng Liu 0015, Jia-Bin Huang 0001, Cusuh Ham
WACV5
2025 UrbanIR: Large-Scale Urban Scene Inverse Rendering from a Single Video
abstract
We present UrbanIR (Urban Scene Inverse Rendering), a new inverse graphics model that enables realistic, free-viewpoint renderings of scenes under various lighting conditions with a single video. It accurately infers shape, albedo, visibility, and sun and sky illumination from wide-baseline videos, such as those from car-mounted cameras, differing from NeRF's dense view settings. In this context, standard methods often yield subpar geometry and material estimates, such as inaccurate roof representations and numerous ‘floaters’. UrbanIR addresses these issues with novel losses that reduce errors in inverse graphics inference and rendering artifacts. Its techniques allow for precise shadow volume estimation in the original scene. The model's outputs support controllable editing, enabling photorealistic free-viewpoint renderings of night simulations, relit scenes, and inserted objects, marking a significant improvement over existing state-of-the-art methods. Our code and data will be made publicly available upon acceptance.
Chih-Hao Lin, Kuan-Sheng Chen, David A. Forsyth, Jia-Bin Huang 0001, Anand Bhattad, Shenlong Wang
3DV6
2025 Generative Multiview Relighting for 3D Reconstruction under Extreme Illumination Variation
abstract
Reconstructing the geometry and appearance of objects from photographs taken in different environments is difficult as the illumination and, therefore, the object appearance vary across captured images. This is particularly challenging for specular objects whose appearance strongly depends on the viewing direction. Some prior approaches model appearance variation across images using a per-image embedding vector, while others use physically-based rendering to recover the materials and per-image illumination. Such approaches fail at faithfully recovering view-dependent appearance given the significant variation in input illumination and tend to produce mostly diffuse results. We present an approach that reconstructs objects from images taken under different illuminations by first relighting the images under a single reference illumination with a multiview relighting diffusion model and then reconstructing the object’s geometry and appearance with a radiance field architecture that is robust to the minor remaining inconsistencies among the relit images. We validate our approach on synthetic and real datasets and demonstrate that it outperforms existing techniques at reconstructing high-fidelity appearance from images taken under extreme illumination variation. Moreover, our approach is particularly effective at recovering view-dependent "shiny" appearance which cannot be reconstructed by prior methods.
Hadi Alzayer, Philipp Henzler, Jonathan T. Barron, Jia-Bin Huang 0001, Pratul P. Srinivasan, Dor Verbin
CVPR4
2025 Textured Gaussians for Enhanced 3D Scene Appearance Modeling
abstract
3D Gaussian Splatting (3DGS) has emerged as the state-of-the-art 3D reconstruction technique, offering high-quality results with fast training and rendering. However, its expressivity is limited as pixels covered by the same Gaussian share identical colors aside from a Gaussian falloff scaling factor, and individual Gaussians can only represent simple ellipsoids geometrically. To overcome these limitations, we integrate texture and alpha mapping from traditional graphics with 3DGS. Our approach augments each Gaussian with alpha, RGB, or RGBA texture maps to model spatially varying color and opacity across each Gaussian’s extent. This allows Gaussians to represent richer texture patterns and geometric structures beyond single-color ellipsoids. Notably, alpha-only texture maps significantly improve Gaussian expressivity, while further augmenting with RGB texture maps achieve maximum expressivity. We validate our method on a wide variety of standard benchmark datasets and our own custom captures at both the object and scene levels, and demonstrate image quality improvements over existing methods while using a similar or lower number of Gaussians.
Brian Chao, Hung-Yu Tseng, Lorenzo Porzi, Chen Gao 0003, Tuotuo Li, Qinbo Li, Ayush Saraf, Jia-Bin Huang 0001, Johannes Kopf 0001, Gordon Wetzstein, Changil Kim 0001
CVPR8
2025 Generative Omnimatte: Learning to Decompose Video into Layers
abstract
Given a video and a set of input object masks, an omnimatte method aims to decompose the video into semantically meaningful layers containing individual objects along with their associated effects, such as shadows and reflections. Existing omnimatte methods assume a static background or accurate pose and depth estimation and produce poor decompositions when these assumptions are violated. Furthermore, due to the lack of generative prior on natural videos, existing methods cannot complete dynamic occluded regions. We present a novel generative layered video decomposition framework to address the omnimatte problem. Our method does not assume a stationary scene or require camera pose or depth information and produces clean, complete layers, including convincing completions of occluded dynamic regions. Our core idea is to train a video diffusion model to identify and remove scene effects caused by a specific object. We show that this model can be finetuned from an existing video inpainting model with a small, carefully curated dataset, and demonstrate high-quality decompositions and editing results for a wide range of casually captured videos containing soft shadows, glossy reflections, splashing water, and more.
Yao-Chih Lee, Erika Lu, Sarah Rumbley, Michal Geyer, Jia-Bin Huang 0001, Tali Dekel, Forrester Cole
CVPR5
2025 LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fields
abstract
We present Large Inverse Rendering Model (LIRM), a transformer architecture that jointly reconstructs high-quality shape, materials, and radiance fields with view-dependent effects in less than a second. Our model builds upon the recent Large Reconstruction Models (LRMs) that achieve state-of-the-art sparse-view reconstruction quality. However, existing LRMs struggle to reconstruct unseen parts accurately and cannot recover glossy appearance or generate relightable 3D contents that can be consumed by standard Graphics engines. To address these limitations, we make three key technical contributions to build a more practical multi-view 3D reconstruction framework. First, we introduce an update model that allows us to progressively add more input views to improve our reconstruction. Second, we propose a hexa-plane neural SDF representation to better recover detailed textures, geometry and material parameters. Third, we develop a novel neural directional-embedding mechanism to handle view-dependent effects. Trained on a large-scale shape and material dataset with a tailored coarse-to-fine training scheme, our model achieves compelling results. It compares favorably to optimization-based dense-view inverse rendering methods in terms of geometry and relighting accuracy, while requiring only a fraction of the inference time.
Zhengqin Li, Dilin Wang, Ka Chen, Zhaoyang Lv, Thu Nguyen-Phuoc, Milim Lee, Jia-Bin Huang 0001, Lei Xiao 0014, Yufeng Zhu, Carl S. Marshall, Yuheng Ren, Richard A. Newcombe, Zhao Dong 0001
CVPR7
2025 Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions
abstract
We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and their motion dynamics. Our method addresses this gap by generating body-shape-aware human motions from natural language prompts. We utilize a finite scalar quantization-based variational autoencoder (FSQ-VAE) to quantize motion into discrete tokens and then leverage continuous body shape information to de-quantize these tokens back into continuous, detailed motion. Additionally, we harness the capabilities of a pretrained language model to predict both continuous shape parameters and motion tokens, facilitating the synthesis of text-aligned motions and decoding them into shape-aware motions. We evaluate our method quantitatively and qualitatively, and also conduct a comprehensive perceptual study to demonstrate its efficacy in generating shape-aware motions. Project URL: https://shape-move.github.io/.
Ting-Hsuan Liao, Yi Zhou 0023, Chun-Hao Paul Huang, Saayan Mitra, Jia-Bin Huang 0001, Uttaran Bhattacharya
CVPR6
2025 IRIS: Inverse Rendering of Indoor Scenes from Low Dynamic Range Images
abstract
Inverse rendering seeks to recover 3D geometry, surface material, and lighting from captured images, enabling advanced applications such as novel-view synthesis, relighting, and virtual object insertion. However, most existing techniques rely on high dynamic range (HDR) images as input, limiting accessibility for general users. In response, we introduce IRIS, an inverse rendering framework that recovers the physically based material, spatially-varying HDR lighting, and camera response functions from multi-view, low-dynamic-range (LDR) images. By eliminating the dependence on HDR input, we make inverse rendering technology more accessible. We evaluate our approach on real-world and synthetic scenes and compare it with state-of-the-art methods. Our results show that IRIS effectively recovers HDR lighting, accurate material, and plausible camera response functions, supporting photorealistic relighting and object insertion.
Chih-Hao Lin, Jia-Bin Huang 0001, Zhengqin Li, Zhao Dong 0001, Christian Richardt, Tuotuo Li, Michael Zollhöfer, Johannes Kopf 0001, Shenlong Wang, Changil Kim 0001
CVPR2
2025 VideoGigaGAN: Towards Detail-rich Video Super-Resolution
abstract
Video super-resolution (VSR) models achieve temporal consistency but often produce blurrier results than their image-based counterparts due to limited generative capacity. This prompts the question: can we adapt a generative image upsampler for VSR while preserving temporal consistency? We introduce VideoGigaGAN, a new generative VSR model that combines high-frequency detail with temporal stability, building on the large-scale GigaGAN image upsampler. Simple adaptations of GigaGAN for VSR led to flickering issues, so we propose techniques to enhance temporal consistency. We validate the effectiveness of VideoGigaGAN by comparing it with state-of-the-art VSR models on public datasets and showcasing video results with 8× upsampling.
Taesung Park, Richard Zhang 0001, Yang Zhou 0009, Eli Shechtman, Feng Liu 0015, Jia-Bin Huang 0001, Difan Liu
CVPR7
2025 MaDCoW: Marginal Distortion Correction for Wide-Angle Photography with Arbitrary Objects
abstract
We introduce MaDCoW, a method for correcting marginal distortion of arbitrary objects in wide-angle photography. People often use wide-angle photography—it is the default in smartphone cameras—but very-wide-fields-of-view produce distorted object appearance in image margins. In our system, a user annotates straight lines and regions of interest. MaDCoW solves for a separate linear perspective projection for each region and then jointly solves for a distortion-minimizing projection for the whole photograph. We show that MaDCoW can produce good results in cases where previous methods yield visible distortions.
Kevin Zhang 0003, Jia-Bin Huang 0001, Jose Echevarria, Stephen DiVerdi, Aaron Hertzmann
CVPR2
2025 PAD3R: Pose-Aware Dynamic 3D Reconstruction from Casual Videos
abstract
We present PAD3R, a method for reconstructing deformable 3D objects from casually captured, unposed monocular videos. Unlike existing approaches, PAD3R handles long video sequences that feature substantial object deformation, large-scale camera movement, and limited view coverage, which typically challenge conventional systems. At its core, our approach trains a personalized, object-centric pose estimator, supervised by a pre-trained image-to-3D model. This guides the optimization of deformable 3D Gaussian representation. The optimization is further regularized by long-term 2D point tracking over the entire input video. By combining generative priors and differentiable rendering, PAD3R reconstructs high-fidelity, articulated 3D representations of objects in a category-agnostic way. Extensive qualitative and quantitative results show that PAD3R is robust and generalizes well across challenging scenarios, highlighting its potential for dynamic scene understanding and 3D content creation. Please refer to our project page for more details: PAD3R.github.io .
Ting-Hsuan Liao, Songwei Ge, Gengshan Yang, Jia-Bin Huang 0001
SIGGRAPH Asia6
2025 Expressive Image Generation and Editing with Rich Text
Songwei Ge, Taesung Park, Jun-Yan Zhu, Jia-Bin Huang 0001
Int. J. Comput. Vis.4
2025 Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos
abstract
We propose a generative model that, given a coarsely edited image, synthesizes a photorealistic output that follows the prescribed layout. Our method transfers fine details from the original image and preserve the identity of its parts. Yet, it adapts it to the lighting and context defined by the new layout. Our key insight is that videos are a powerful source of supervision for this task: objects and camera motions provide many observations of how the world changes with viewpoint, lighting, and physical interactions. We construct an image dataset in which each sample is a pair of source and target frames extracted from the same video at randomly chosen time intervals. We warp the source frame toward the target using two motion models that mimic the expected test-time user edits. We supervise our model to translate the warped image into the ground truth, starting from a pretrained diffusion model. Our model design explicitly enables fine detail transfer from the source frame to the generated image, while closely following the user-specified layout. We show that by using simple segmentations and coarse 2D manipulations, we can synthesize a photorealistic edit faithful to the user’s input while addressing second-order effects like harmonizing the lighting and physical interactions between edited objects. Project page and code can be found at https://magic-fixup.github.io.
Hadi Alzayer, Zhihao Xia, Xuaner (Cecilia) Zhang, Eli Shechtman, Jia-Bin Huang 0001, Michaël Gharbi
ACM Trans. Graph.5
2024 Seeing the World through Your Eyes
abstract
The reflective nature of the human eye is an under-appreciated source of information about what the world around us looks like. By imaging the eyes of a moving person, we capture multiple views of a scene outside the camera's direct line of sight through the reflections in the eyes. In this paper, we reconstruct a radiance field beyond the camera's line of sight using portrait images containing eye reflections. This task is challenging due to 1) the difficulty of accurately estimating eye poses and 2) the entangled appearance of the iris textures and the scene reflections. To address these, our method jointly optimizes the cornea poses, the radiance field depicting the scene, and the observer's eye iris texture. We further present a regularization prior on the iris texture to improve scene reconstruction quality. Through various experiments on synthetic and real-world captures featuring people with varied eye colors, and lighting conditions, we demonstrate the feasibility of our approach to recover the radiance field using cornea reflections.
Hadi Alzayer, Kevin Zhang 0003, Brandon Yushan Feng, Christopher A. Metzler, Jia-Bin Huang 0001
CVPR5
2024 LTM: Lightweight Textured Mesh Extraction and Refinement of Large Unbounded Scenes for Efficient Storage and Real-Time Rendering
abstract
Advancements in neural signed distance fields (SDFs) have enabled modeling 3D surface geometry from a set of 2D images of real-world scenes. Baking neural SDFs can extract explicit mesh with appearance baked into texture maps as neural features. The baked meshes still have a large memory footprint and require a powerful GPU for real-time rendering. Neural optimization of such large meshes with differentiable rendering pose significant challenges. We propose a method to produce optimized meshes for large unbounded scenes with low triangle budget and high fidelity of geometry and appearance. We achieve this by combining advancements in baking neural SDFs with classical mesh simplification techniques and proposing a joint appearance-geometry refinement step. The visual quality is comparable to or better than state-of-the-art neural meshing and baking methods with high geometric accuracy despite significant reduction in triangle count, making the produced meshes efficient for storage, transmission, and rendering on mobile hardware. We validate the effectiveness of the proposed method on large unbounded scenes from mip-NeRF 360, Tanks & Temples, and Deep Blending datasets, achieving at-par rendering quality with 73 x reduced triangles and 11 x reduction in memory footprint.
Rajvi Shah, Qinbo Li, Yipeng Wang 0018, Ayush Saraf, Changil Kim 0001, Jia-Bin Huang 0001, Dinesh Manocha, Suhib Alsisan, Johannes Kopf 0001
CVPR7
2024 On the Content Bias in Fréchet Video Distance
abstract
Frechet Video Distance (FVD), a prominent metric for evaluating video generation models, is known to conflict with human perception occasionally. In this paper, we aim to explore the extent of FVD ‘s bias toward per-frame quality over temporal realism and identify its sources. We first quantify the FVD’ s sensitivity to the temporal axis by decoupling the frame and motion quality and find that the FVD increases only slightly with large temporal corruption. We then analyze the generated videos and show that via careful sampling from a large set of generated videos that do not contain motions, one can drastically decrease FVD without improving the temporal quality. Both studies suggest FVD's bias towards the quality of individual frames. We further observe that the bias can be attributed to the features extracted from a supervised video classifier trained on the content-biased dataset. We show that FVD with features extracted from the recent large-scale self-supervised video models is less biased toward image quality. Finally, we revisit a few real-world examples to validate our hypothesis.
Songwei Ge, Aniruddha Mahapatra, Gaurav Parmar, Jun-Yan Zhu, Jia-Bin Huang 0001
CVPR5
2024 FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
abstract
Diffusion models have transformed the image-to-image (I2I) synthesis and are now permeating into videos. How-ever, the advancement of video-to-video (V2V) synthesis has been hampered by the challenge of maintaining temporal consistency across video frames. This paper proposes a consistent V2V synthesis framework by jointly leveraging spatial conditions and temporal optical flow clues within the source video. Contrary to prior methods that strictly adhere to optical flow, our approach harnesses its benefits while handling the imperfection in flow estimation. We encode the optical flow via warping from the first frame and serve it as a supplementary reference in the diffusion model. This enables our model for video synthesis by editing the first frame with any prevalent I2I models and then propagating edits to successive frames. Our V2V model, Flow Vid, demon-strates remarkable properties: (1) Flexibility: Flow Vid works seamlessly with existing I2I models, facilitating various modifications, including stylization, object swaps, and local edits. (2) Efficiency: Generation of a 4-second video with 30 FPS and 512×512 resolution takes only 1.5 minutes, which is 3.1×, 7.2×, and 10.5× faster than CoDeF, Rerender, and TokenFlow, respectively. (3) High-quality: In user studies, our FlowVid is preferred 45.7% of the time, outperforming CoDeF (3.5%), Rerender (10.2%), and TokenFlow (40.4%).
Bichen Wu, Jialiang Wang 0001, Licheng Yu, Ishan Misra, Jia-Bin Huang 0001, Peizhao Zhang, Peter Vajda, Diana Marculescu
CVPR8
2024 Grounded Text-to-Image Synthesis with Attention Refocusing
abstract
Driven by the scalable diffusion models trained on large-scale datasets, text-to-image synthesis methods have shown compelling results. However, these models still fail to pre-cisely follow the text prompt involving multiple objects, attributes, or spatial compositions. In this paper, we reveal the potential causes in the diffusion model's cross-attention and self-attention layers. We propose two novel losses to refocus attention maps according to a given spatial layout during sampling. Creating the layouts manually requires additional effort and can be tedious. Therefore, we explore using large language models (LLM) to produce these lay-outs for our method. We conduct extensive experiments on the DrawBench, HRS, and TIFA benchmarks to evaluate our proposed method. We show that our proposed attention re-focusing effectively improves the controllability of existing approaches.
Quynh Phung, Songwei Ge, Jia-Bin Huang 0001
CVPR3
2024 In-N-Out: Faithful 3D GAN Inversion with Volumetric Decomposition for Face Editing
abstract
3D-aware GANs offer new capabilities for view synthe-sis while preserving the editing functionalities of their 2D counterparts. GAN inversion is a crucial step that seeks the latent code to reconstruct input images or videos, subsequently enabling diverse editing tasks through manipulation of this latent code. However, a model pretrained on a particular dataset (e.g., FFHQ) often has difficulty re-constructing images with out-of-distribution (OOD) objects such as faces with heavy make-up or occluding objects. We address this issue by explicitly modeling OOD objects from the input in 3D-aware GANs. Our core idea is to represent the image using two individual neural radiance fields: one for the in-distribution content and the other for the out-of-distribution object. The final reconstruction is achieved by optimizing the composition of these two radiance fields with carefully designed regularization. We demonstrate that our explicit decomposition alleviates the inherent tradeoff between reconstruction fidelity and editability. We evaluate reconstruction accuracy and editability of our method on challenging real face images and videos and showcase favorable results against other baselines. More results can found at https://in-n-out-3d.github.io/.
Zhixin Shu, Cameron Smith, Seoung Wug Oh, Jia-Bin Huang 0001
CVPR5
2024 TextureDreamer: Image-Guided Texture Synthesis through Geometry-Aware Diffusion
abstract
We present TextureDreamer, a novel image-guided texture synthesis method to transfer relightable textures from a small number of input images (3 to 5) to target 3D shapes across arbitrary categories. Texture creation is a pivotal challenge in vision and graphics. Industrial companies hire experienced artists to manually craft textures for 3D assets. Classical methods require densely sampled views and ac-curately aligned geometry, while learning-based methods are confined to category-specific shapes within the dataset. In contrast, TextureDreamer can transfer highly detailed, intricate textures from real-world environments to arbi-trary objects with only a few casually captured images, po-tentially significantly democratizing texture creation. Our core idea, personalized geometry-aware score distillation (PGSD), draws inspiration from recent advancements in diffuse models, including personalized modeling for texture information extraction, score distillation for detailed appearance synthesis, and explicit geometry guidance with ControlNet. Our integration and several essential modifications substantially improve the texture quality. Experiments on real images spanning different categories show that TextureDreamer can successfully transfer highly realistic, se-mantic meaningful texture to arbitrary objects, surpassing the visual quality of previous state-of-the-art. Project page: https://texturedreamer.github.io
Yu-Ying Yeh, Jia-Bin Huang 0001, Changil Kim 0001, Lei Xiao 0014, Thu Nguyen-Phuoc, Numair Khan, Manmohan Krishna Chandraker, Carl S. Marshall, Zhao Dong 0001, Zhengqin Li
CVPR2
2024 Fast View Synthesis of Casual Videos with Soup-of-Planes
Yao-Chih Lee, Zhoutong Zhang, Kevin Matzen, Simon Niklaus, Jianming Zhang 0001, Jia-Bin Huang 0001, Feng Liu 0015
ECCV (38)6
2024 Taming Latent Diffusion Model for Neural Radiance Field Inpainting
Chieh Hubert Lin, Changil Kim 0001, Jia-Bin Huang 0001, Qinbo Li, Chih-Yao Ma, Johannes Kopf 0001, Ming-Hsuan Yang 0001, Hung-Yu Tseng
ECCV (3)3
2024 Flash-Splat: 3D Reflection Removal with Flash Cues and Gaussian Splats
Mingyang Xie, Haoming Cai, Sachin Shah, Brandon Yushan Feng, Jia-Bin Huang 0001, Christopher A. Metzler
ECCV (82)6
2024 Rethinking Score Distillation as a Bridge Between Image Distributions
abstract
Score distillation sampling (SDS) has proven to be an important tool, enabling the use of large-scale diffusion priors for tasks operating in data-poor domains. Unfortunately, SDS has a number of characteristic artifacts that limit its utility in general-purpose applications. In this paper, we make progress toward understanding the behavior of SDS and its variants by viewing them as solving an optimal-cost transport path from some current source distribution to a target distribution. Under this new interpretation, we argue that these methods' characteristic artifacts are caused by (1) linear approximation of the optimal path and (2) poor estimates of the source distribution. We show that by calibrating the text conditioning of the source distribution, we can produce high-quality generation and translation results with little extra overhead. Our method can be easily applied across many domains, matching or beating the performance of specialized methods. We demonstrate its utility in text-to-2D, text-to-3D, translating paintings to real images, optical illusion generation, and 3D sketch-to-real. We compare our method to existing approaches for score distillation sampling and show that it can produce high-frequency details with realistic colors.
David McAllister, Songwei Ge, Jia-Bin Huang 0001, David Jacobs 0001, Alexei A. Efros, Aleksander Holynski, Angjoo Kanazawa
NeurIPS3
2024 Planar Reflection-Aware Neural Radiance Fields
Chen Gao 0003, Yipeng Wang 0018, Changil Kim 0001, Jia-Bin Huang 0001, Johannes Kopf 0001
SIGGRAPH Asia4
2024 Recent Trends in 3D Reconstruction of General Non-Rigid Scenes
abstract
Abstract Reconstructing models of the real world, including 3D geometry, appearance, and motion of real scenes, is essential for computer graphics and computer vision. It enables the synthesizing of photorealistic novel views, useful for the movie industry and AR/VR applications. It also facilitates the content creation necessary in computer games and AR/VR by avoiding laborious manual design processes. Further, such models are fundamental for intelligent computing systems that need to interpret real‐world scenes and actions to act and interact safely with the human world. Notably, the world surrounding us is dynamic, and reconstructing models of dynamic, non‐rigidly moving scenes is a severely underconstrained and challenging problem. This state‐of‐the‐art report (STAR) offers the reader a comprehensive summary of state‐of‐the‐art techniques with monocular and multi‐view inputs such as data from RGB and RGB‐D sensors, among others, conveying an understanding of different approaches, their potential applications, and promising further research directions. The report covers 3D reconstruction of general non‐rigid scenes and further addresses the techniques for scene decomposition, editing and controlling, and generalizable and generative modeling. More specifically, we first review the common and fundamental concepts necessary to understand and navigate the field and then discuss the state‐of‐the‐art techniques by reviewing recent approaches that use traditional and machine‐learning‐based neural representations, including a discussion on the newly enabled applications. The STAR is concluded with a discussion of the remaining limitations and open challenges.
Raza Yunus, Jan Eric Lenssen, Michael Niemeyer, Yiyi Liao, Christian Rupprecht 0001, Christian Theobalt, Gerard Pons-Moll, Jia-Bin Huang 0001, Vladislav Golyanik, Eddy Ilg
Comput. Graph. Forum8
2024 DisCO: Portrait Distortion Correction with Perspective-Aware 3D GANs
Zhixiang Wang 0001, Yu-Lun Liu 0001, Jia-Bin Huang 0001, Shin'ichi Satoh 0001, Sizhuo Ma, Gurunandan Krishnan, Jian Wang 0100
Int. J. Comput. Vis.3
2024 Partitioned scheduling with safety-performance trade-offs in stochastic conditional DAG models
Xuanliang Deng, Ashrarul H. Sifat, Shao-Yu Huang, Sen Wang 0014, Jia-Bin Huang 0001, Changhee Jung, Ryan K. Williams, Haibo Zeng 0001
J. Syst. Archit.5
2024 Time-Triggered Scheduling for Nonpreemptive Real-Time DAG Tasks Using 1-Opt Local Search
abstract
Modern real-time systems often involve numerous computational tasks characterized by intricate dependency relationships. Within these systems, data propagate through cause–effect chains from one task to another, making it imperative to minimize end-to-end latency to ensure system safety and reliability. In this article, we introduce innovative nonpreemptive scheduling techniques designed to reduce the worst-case end-to-end latency and/or time disparity for task sets modeled with directed acyclic graphs (DAGs). This is challenging because of the noncontinuous and nonconvex characteristics of the objective functions, hindering the direct application of standard optimization frameworks. Customized optimization frameworks aiming at achieving optimal solutions may suffer from scalability issues, while general heuristic algorithms often lack theoretical performance guarantees. To address this challenge, we incorporate the “1-opt” concept from the optimization literature (Essentially, 1-opt means that the quality of a solution cannot be improved if only one single variable can be changed) into the design of our algorithm. We propose a novel optimization algorithm that effectively balances the tradeoff between theoretical guarantees and algorithm scalability. By demonstrating its theoretical performance guarantees, we establish that the algorithm produces 1-opt solutions while maintaining polynomial run-time complexity. Through extensive large-scale experiments, we demonstrate that our algorithm can effectively reduce the latency metrics by 20% to 40%, compared to state-of-the-art methods.
Sen Wang 0014, Dong Li 0035, Shao-Yu Huang, Xuanliang Deng, Ashrarul H. Sifat, Jia-Bin Huang 0001, Changhee Jung, Ryan K. Williams, Haibo Zeng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2023 Robust Dynamic Radiance Fields
abstract
Dynamic radiance field reconstruction methods aim to model the time-varying structure and appearance of a dynamic scene. Existing methods, however, assume that accurate camera poses can be reliably estimated by Structure from Motion (SfM) algorithms. These methods, thus, are unreliable as SfM algorithms often fail or produce erroneous poses on challenging videos with highly dynamic objects, poorly textured surfaces, and rotating camera motion. We address this robustness issue by jointly estimating the static and dynamic radiance fields along with the camera parameters (poses and focal length). We demonstrate the robustness of our approach via extensive quantitative and qualitative experiments. Our results show favorable performance over the state-of-the-art dynamic view synthesis methods.
Yu-Lun Liu 0001, Chen Gao 0003, Andreas Meuleman, Hung-Yu Tseng, Ayush Saraf, Changil Kim 0001, Yung-Yu Chuang, Johannes Kopf 0001, Jia-Bin Huang 0001
CVPR9
2023 DC2: Dual-Camera Defocus Control by Learning to Refocus
abstract
Smartphone cameras today are increasingly approaching the versatility and quality of professional cameras through a combination of hardware and software advancements. However, fixed aperture remains a key limitation, preventing users from controlling the depth of field (DoF) of captured images. At the same time, many smartphones now have multiple cameras with different fixed apertures - specifically, an ultra-wide camera with wider field of view and deeper DoF and a higher resolution primary camera with shallower DoF. In this work, we propose$DC^{2}$, a system for defocus control for synthetically varying camera aperture, focus distance and arbitrary defocus effects by fusing information from such a dual-camera system. Our key insight is to leverage real-world smartphone camera dataset by using image refocus as a proxy task for learning to control defocus. Quantitative and qualitative evaluations on real-world data demonstrate our system's efficacy where we outperform state-of-the-art on defocus deblurring, bokeh rendering, and image refocus. Finally, we demonstrate creative post-capture defocus control enabled by our method, including tilt-shift and content-based defocus effects.
Hadi Alzayer, Abdullah Abuolaim, Leung Chun Chan, Ying Chen Lou, Jia-Bin Huang 0001, Abhishek Kar
CVPR6
2023 HyperReel: High-Fidelity 6-DoF Video with Ray-Conditioned Sampling
abstract
Volumetric scene representations enable photorealistic view synthesis for static scenes and form the basis of several existing 6-DoF video techniques. However, the volume rendering procedures that drive these representations necessitate careful trade-offs in terms of quality, rendering speed, and memory efficiency. In particular, existing methods fail to simultaneously achieve real-time performance, small memory footprint, and high-quality rendering for challenging real-world scenes. To address these issues, we present HyperReel―a novel 6-DoF video representation. The two core components of HyperReel are: (1) a ray-conditioned sample prediction network that enables high-fidelity, high frame rate rendering at high resolutions and (2) a compact and memory-efficient dynamic volume representation. Our 6-DoF video pipeline achieves the best performance compared to prior and contemporary approaches in terms of visual quality with small memory requirements, while also rendering at up to 18 frames-per-second at megapixel resolution without any custom CUDA code.
Benjamin Attal, Jia-Bin Huang 0001, Christian Richardt, Michael Zollhöfer, Johannes Kopf 0001, Matthew O'Toole, Changil Kim 0001
CVPR2
2023 Shape-Aware Text-Driven Layered Video Editing
abstract
Temporal consistency is essential for video editing applications. Existing work on layered representation of videos allows propagating edits consistently to each frame. These methods, however, can only edit object appearance rather than object shape changes due to the limitation of using a fixed UV mapping field for texture atlas. We present a shape-aware, text-driven video editing method to tackle this challenge. To handle shape changes in video editing, we first propagate the deformation field between the input and edited keyframe to all frames. We then leverage a pre-trained text-conditioned diffusion model as guidance for refining shape distortion and completing unseen regions. The experimental results demonstrate that our method can achieve shape-aware consistent video editing and compare favorably with the state-of-the-art.
Yao-Chih Lee, Ji-Ze Genevieve Jang, Elizabeth Qiu, Jia-Bin Huang 0001
CVPR5
2023 Progressively Optimized Local Radiance Fields for Robust View Synthesis
abstract
We present an algorithm for reconstructing the radiance field of a large-scale scene from a single casually captured video. The task poses two core challenges. First, most existing radiance field reconstruction approaches rely on accurate pre-estimated camera poses from Structure-from-Motion algorithms, which frequently fail on in-the-wild videos. Second, using a single, global radiance field with finite representational capacity does not scale to longer trajectories in an unbounded scene. For handling unknown poses, we jointly estimate the camera poses with radiance field in a progressive manner. We show that progressive optimization significantly improves the robustness of the reconstruction. For handling large unbounded scenes, we dynamically allocate new local radiance fields trained with frames within a temporal window. This further improves robustness (e.g., performs well even under moderate pose drifts) and allows us to scale to large scenes. Our extensive evaluation on the TANKS AND TEMPLES dataset and our collected outdoor dataset, STATIC HIKES, show that our approach compares favorably with the state-of-the-art.
Andreas Meuleman, Yu-Lun Liu 0001, Chen Gao 0003, Jia-Bin Huang 0001, Changil Kim 0001, Min H. Kim 0001, Johannes Kopf 0001
CVPR4
2023 Consistent View Synthesis with Pose-Guided Diffusion Models
abstract
Novel view synthesis from a single image has been a cornerstone problem for many Virtual Reality applications that provide immersive experiences. However, most existing techniques can only synthesize novel views within a limited range of camera motion or fail to generate consistent and high-quality novel views under significant camera movement. In this work, we propose a pose-guided diffusion model to generate a consistent long-term video of novel views from a single image. We design an attention layer that uses epipolar lines as constraints to facilitate the association between different viewpoints. Experimental results on synthetic and real-world datasets demonstrate the effectiveness of the proposed diffusion model against state-of-the-art transformer-based and GAN-based approaches. More qualitative results are available at https://poseguided-diffusion.github.io/.
Hung-Yu Tseng, Qinbo Li, Changil Kim 0001, Suhib Alsisan, Jia-Bin Huang 0001, Johannes Kopf 0001
CVPR5
2023 Neural-PBIR Reconstruction of Shape, Material, and Illumination
abstract
Reconstructing the shape and spatially varying surface appearances of a physical-world object as well as its surrounding illumination based on 2D images (e.g., photographs) of the object has been a long-standing problem in computer vision and graphics. In this paper, we introduce an accurate and highly efficient object reconstruction pipeline combining neural based object reconstruction and physics-based inverse rendering (PBIR). Our pipeline firstly leverages a neural SDF based shape reconstruction to produce high-quality but potentially imperfect object shape. Then, we introduce a neural material and lighting distillation stage to achieve high-quality predictions for material and illumination. In the last stage, initialized by the neural predictions, we perform PBIR to refine the initial results and obtain the final high-quality reconstruction of object shape, material, and illumination. Experimental results demonstrate our pipeline significantly outperforms existing methods quality-wise and performance-wise. Code: https://neural-pbir.github.io/
Cheng Sun 0004, Guangyan Cai, Zhengqin Li, Kai Yan 0006, Carl S. Marshall, Jia-Bin Huang 0001, Zhao Dong 0001
ICCV7
2023 3D Motion Magnification: Visualizing Subtle Motions with Time-Varying Radiance Fields
abstract
Motion magnification helps us visualize subtle, imperceptible motion. However, prior methods only work for 2D videos captured with a fixed camera. We present a 3D motion magnification method that can magnify subtle motions from scenes captured by a moving camera, while supporting novel view rendering. We represent the scene with time-varying radiance fields and leverage the Eulerian principle for motion magnification to extract and amplify the variation of the embedding of a fixed point over time. We study and validate our proposed principle for 3D motion magnification using both implicit and tri-plane-based radiance fields as our underlying 3D scene representation. We evaluate the effectiveness of our method on both synthetic and real-world scenes captured under various camera setups.
Brandon Yushan Feng, Hadi Alzayer, Michael Rubinstein, William T. Freeman, Jia-Bin Huang 0001
ICCV5
2023 Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models
abstract
Despite tremendous progress in generating high-quality images using diffusion models, synthesizing a sequence of animated frames that are both photorealistic and temporally coherent is still in its infancy. While off-the-shelf billion-scale datasets for image generation are available, collecting similar video data of the same scale is still challenging. Also, training a video diffusion model is computationally much more expensive than its image counterpart. In this work, we explore finetuning a pretrained image diffusion model with video data as a practical solution for the video synthesis task. We find that naively extending the image noise prior to video noise prior in video diffusion leads to sub-optimal performance. Our carefully designed video noise prior leads to substantially better performance. Extensive experimental validation shows that our model, Preserve Your Own COrrelation (PYoCo), attains SOTA zero-shot text-to-video results on the UCF-101 and MSR-VTT benchmarks. It also achieves SOTA video generation quality on the small-scale UCF-101 benchmark with a 10× smaller model using significantly less computation than the prior art. The project page is available at https://research.nvidia.com/labs/dir/pyoco/.
Songwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon, Andrew Tao, Bryan Catanzaro, David Jacobs 0001, Jia-Bin Huang 0001, Ming-Yu Liu 0001, Yogesh Balaji
ICCV8
2023 Expressive Text-to-Image Generation with Rich Text
abstract
Plain text has become a prevalent interface for text-to-image synthesis. However, its limited customization options hinder users from accurately describing desired outputs. For example, plain text makes it hard to specify continuous quantities, such as the precise RGB color value or importance of each word. Furthermore, creating detailed text prompts for complex scenes is tedious for humans to write and challenging for text encoders to interpret. To address these challenges, we propose using a rich-text editor supporting formats such as font style, size, color, and footnote. We extract each word’s attributes from rich text to enable local style control, explicit token reweighting, precise color rendering, and detailed region synthesis. We achieve these capabilities through a region-based diffusion process. We first obtain each word’s region based on attention maps of a diffusion process using plain text. For each region, we enforce its text attributes by creating region-specific detailed prompts and applying region-specific guidance, and maintain its fidelity against plain-text generation through region-based injections. We present various examples of image generation from rich text and demonstrate that our method outperforms strong baselines with quantitative evaluations.
Songwei Ge, Taesung Park, Jun-Yan Zhu, Jia-Bin Huang 0001
ICCV4
2023 ClimateNeRF: Extreme Weather Synthesis in Neural Radiance Field
abstract
Physical simulations produce excellent predictions of weather effects. Neural radiance fields produce SOTA scene models. We describe a novel NeRF-editing procedure that can fuse physical simulations with NeRF models of scenes, producing realistic movies of physical phenomena in those scenes. Our application – Climate NeRF – allows people to visualize what climate change outcomes will do to them.ClimateNeRF allows us to render realistic weather effects, including smog, snow, and flood. Results can be controlled with physically meaningful variables like water level. Qualitative and quantitative studies show that our simulated results are significantly more realistic than those from SOTA 2D image editing and SOTA 3D NeRF stylization.
Zhi-Hao Lin, David A. Forsyth, Jia-Bin Huang 0001, Shenlong Wang
ICCV4
2023 OmnimatteRF: Robust Omnimatte with 3D Background Modeling
abstract
Video matting has broad applications, from adding interesting effects to casually captured movies to assisting video production professionals. Matting with associated effects such as shadows and reflections has also attracted increasing research activity, and methods like Omnimatte have been proposed to separate dynamic foreground objects of interest into their own layers. However, prior works represent video backgrounds as 2D image layers, limiting their capacity to express more complicated scenes, thus hindering application to real-world videos. In this paper, we propose a novel video matting method, OmnimatteRF, that combines dynamic 2D foreground layers and a 3D background model. The 2D layers preserve the details of the subjects, while the 3D background robustly reconstructs scenes in real-world videos. Extensive experiments demonstrate that our method reconstructs scenes with better quality on various videos.
Geng Lin, Chen Gao 0003, Jia-Bin Huang 0001, Changil Kim 0001, Yipeng Wang 0018, Matthias Zwicker, Ayush Saraf
ICCV3
2023 Dynamic Mesh-Aware Radiance Fields
abstract
Embedding polygonal mesh assets within photorealistic Neural Radience Fields (NeRF) volumes, such that they can be rendered and their dynamics simulated in a physically consistent manner with the NeRF, is under-explored from the system perspective of integrating NeRF into the traditional graphics pipeline. This paper designs a two-way coupling between mesh and NeRF during rendering and simulation. We first review the light transport equations for both mesh and NeRF, then distill them into an efficient algorithm for updating radiance and throughput along a cast ray with an arbitrary number of bounces. To resolve the discrepancy between the linear color space that the path tracer assumes and the sRGB color space that standard NeRF uses, we train NeRF with High Dynamic Range (HDR) images. We also present a strategy to estimate light sources and cast shadows on the NeRF. Finally, we consider how the hybrid surface-volumetric formulation can be efficiently integrated with a high-performance physics simulator that supports cloth, rigid and soft bodies. The full rendering and simulation system can be run on a GPU at interactive rates. We show that a hybrid system approach outperforms alternatives in visual realism for mesh insertion, because it allows realistic light transport from volumetric NeRF media onto surfaces, which affects the appearance of reflective/refractive surfaces and illumination of diffuse surfaces informed by the dynamic scene.
Yi-Ling Qiao, Alexander Gao, Jia-Bin Huang 0001, Ming C. Lin
ICCV5
2023 RTailor: Parameterizing Soft Error Resilience for Mixed-Criticality Real-Time Systems
abstract
Equipping real-time systems with soft error resilience can be challenging due to the tradeoff of the timing and failure requirements for mixed-criticality tasks. Violation of these requirements yields failed task scheduling in one way or another. However, not every task requires the same degree of soft error resilience. For example, low-criticality tasks can run with low or even no soft error resilience, whereas mid- or highcriticality tasks may require relatively high resilience depending on their inherent failure requirement. Unfortunately, existing soft error resilience schemes do not have the ability to control the degree of their resilience in a fine-grained way, i.e., they can only be turned on or off as a whole during task execution. To this end, this paper presents RTailor (Resilience Tailor), a compiler-directed parameterized soft error resilience scheme that achieves the desired level of soft error protection according to the demand of each task. The key idea is that for a given protection ratio, compilers can transform a hot loop such that the number of its iterations protected over the total iterations matches the ratio. Compared to full resilience protecting every iteration, RTailor's parameterized soft error resilience significantly reduces the performance overhead of tasks, thereby improving their real-time schedulability. The experimental results highlight that for four representative fault rates, RTailor achieves 15%~average schedulability improvements over the state-of-the-art work that lacks parameterized soft error resilience.
Shao-Yu Huang, Jianping Zeng 0001, Xuanliang Deng, Sen Wang 0014, Ashrarul H. Sifat, Burhanuddin Bharmal, Jia-Bin Huang 0001, Ryan K. Williams, Haibo Zeng 0001, Changhee Jung
RTSS7
2023 Single-Image 3D Human Digitization with Shape-guided Diffusion
abstract
We present an approach to generate a 360-degree view of a person with a consistent, high-resolution appearance from a single input image. NeRF and its variants typically require videos or images from different viewpoints. Most existing approaches taking monocular input either rely on ground-truth 3D scans for supervision or lack 3D consistency. While recent 3D generative models show promise of 3D consistent human digitization, these approaches do not generalize well to diverse clothing appearances, and the results lack photorealism. Unlike existing work, we utilize high-capacity 2D diffusion models pretrained for general image synthesis tasks as an appearance prior of clothed humans. To achieve better 3D consistency while retaining the input identity, we progressively synthesize multiple views of the human in the input image by inpainting missing regions with shape-guided diffusion conditioned on silhouette and surface normal. We then fuse these synthesized multi-view images via inverse rendering to obtain a fully textured high-resolution 3D mesh of the given person. Experiments show that our approach outperforms prior methods and achieves photorealistic 360-degree synthesis of a wide range of clothed humans with complex textures from a single image.
Badour AlBahar, Shunsuke Saito, Hung-Yu Tseng, Changil Kim 0001, Johannes Kopf 0001, Jia-Bin Huang 0001
SIGGRAPH Asia6
2023 Learning representational invariances for data-efficient action recognition
Yuliang Zou, Jinwoo Choi 0001, Qitong Wang 0001, Jia-Bin Huang 0001
Comput. Vis. Image Underst.4
2022 Learning Neural Light Fields with Ray-Space Embedding
abstract
Neural radiance fields (NeRFs) produce state-of-the-art view synthesis results, but are slow to render, requiring hundreds of network evaluations per pixel to approximate a volume rendering integral. Baking NeRFs into explicit data structures enables efficient rendering, but results in large memory footprints and, in some cases, quality reduction. Additionally, volumetric representations for view synthesis often struggle to represent challenging view dependent effects such as distorted reflections and refractions. We present a novel neural light field representation that, in contrast to prior work, is fast, memory efficient, and excels at modeling complicated view dependence. Our method supports rendering with a single network evaluation per pixel for small baseline light fields and with only a few evaluations per pixel for light fields with larger baselines. At the core of our approach is a ray-space embedding network that maps 4D ray-space into an intermediate, interpolable latent space. Our method achieves state-of-the-art quality on dense forward-facing datasets such as the Stanford Light Field dataset. In addition, for forward-facing scenes with sparser inputs we achieve results that are competitive with NeRF-based approaches while providing a better speed/quality/memory trade-off with far fewer network evaluations.
Benjamin Attal, Jia-Bin Huang 0001, Michael Zollhöfer, Johannes Kopf 0001, Changil Kim 0001
CVPR2
2022 Boosting View Synthesis with Residual Transfer
abstract
Volumetric view synthesis methods with neural representations, such as NeRF and NeX, have recently demonstrated high-quality novel view synthesis. However, optimizing these representations is slow, and even fully trained models cannot reproduce all fine details in the input views. We present a simple but effective technique to boost the rendering quality, which can be easily integrated with most view synthesis methods. The core idea is to transfer color resid-uals (the difference between the input images and their re-construction) from training views to novel views. We blend the residuals from multiple views using a heuristic weighting scheme depending on ray visibility and angular differ-ences. We integrate our technique with several state-of-the-art view synthesis methods and evaluate the Real Forward-facing and the Shiny datasets. Our results show that at about 1/10th the number of training iterations, we achieve the same rendering quality as fully converged NeRF and NeX models, and when applied to fully converged models, we significantly improve their rendering quality.
Xuejian Rong, Jia-Bin Huang 0001, Ayush Saraf, Changil Kim 0001, Johannes Kopf 0001
CVPR2
2022 Neural Global Shutter: Learn to Restore Video from a Rolling Shutter Camera with Global Reset Feature
abstract
Most computer vision systems assume distortion-free images as inputs. The widely used rolling-shutter (RS) image sensors, however, suffer from geometric distortion when the camera and object undergo motion during capture. Extensive researches have been conducted on correcting RS distortions. However, most of the existing work relies heavily on the prior assumptions of scenes or motions. Besides, the motion estimation steps are either oversimplified or computationally inefficient due to the heavy flow warping, limiting their applicability. In this paper, we investigate using rolling shutter with a global reset feature (RSGR) to restore clean global shutter (GS) videos. This feature enables us to turn the rectification problem into a deblur-like one, getting rid of inaccurate and costly explicit motion estimation. First, we build an optic system that captures paired RSGR/GS videos. Second, we develop a novel algorithm incorporating spatial and temporal designs to correct the spatial-varying RSGR distortion. Third, we demonstrate that existing image-to-image translation algorithms can recover clean GS videos from distorted RSGR inputs, yet our algorithm achieves the best performance with the specific designs. Our rendered results are not only visually appealing but also beneficial to downstream tasks. Compared to the state-of-the-art RS solution, our RSGR solution is superior in both effectiveness and efficiency. Considering it is easy to realize without changing the hardware, we believe our RSGR solution can potentially replace the RS solution in taking distortion-free videos with low noise and low budget.
Zhixiang Wang 0001, Xiang Ji 0005, Jia-Bin Huang 0001, Shin'ichi Satoh 0001, Yinqiang Zheng
CVPR3
2022 Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer
Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin 0001, Guan Pang, David Jacobs 0001, Jia-Bin Huang 0001, Devi Parikh
ECCV (17)7
2022 Temporally Consistent Semantic Video Editing
Badour AlBahar, Jia-Bin Huang 0001
ECCV (15)3
2022 Learning Instance-Specific Adaptation for Cross-Domain Segmentation
Yuliang Zou, Chun-Liang Li, Han Zhang 0010, Tomas Pfister, Jia-Bin Huang 0001
ECCV (33)6
2022 Self-Supervised Cross-Video Temporal Learning for Unsupervised Video Domain Adaptation
abstract
We address the task of unsupervised domain adaptation (UDA) for videos with self-supervised learning. While UDA for images is a widely studied problem, UDA for videos is relatively unexplored. In this paper, we propose a novel self-supervised loss for the task of video UDA. The method is motivated by inverted reasoning. Many works on video classification have shown success with representations based on events in videos, e.g., ‘reaching’, ‘picking’, and ‘drinking’ events for ‘drinking coffee’. We argue that if we have event-based representations, we should be able to predict the relative distances between clips in videos. Inverting that, we propose a self-supervised task to predict the difference of the distance between two clips from the source video and the distance between two clips from the target video. We hope that such a task would encourage learning event-based representations of the videos, which is known to be beneficial for classification. Since we predict the difference of clip distances between clips from source videos and target videos, we ‘tie’ the two domains and expect to achieve well-adapted representations. We combine this purely self-supervised loss and the source classification loss to learn the model parameters. We give extensive empirical results on challenging video UDA benchmarks, i.e., UCF-HMDB and EPIC-Kitchens. The presented qualitative and quantitative results support our motivations and method.
Jinwoo Choi 0001, Jia-Bin Huang 0001
ICPR2
2022 Boosting Source-free Domain Adaptation via Confidence-based Subsets Feature Alignment
abstract
Source-free Domain Adaptation (SFDA) aims to adapt a model trained on a given (source) environment to the new (target) environment, without directly accessing the source data. Due to the lack of labeled source data, it is often difficult for SFDA methods to provide reliable class representations for the target data. To overcome this issue, we propose the idea of Confidence-based Subsets Feature Alignment (CSFA). CSFA divides the target data into two subsets: confident subset that consists of samples having low entropy class predictions from the source model, and non-confident subset with samples that do not. By using the pseudo-labels from the confident subset, we can frame the original SFDA problem as a Universal Domain Adaptation (UniDA) problem, and provide reliable class representations for the target data by aligning feature distributions of the two subsets. Specifically, we propose a multi-task framework that simultaneously applies a standard SFDA algorithm in combination with a UniDA-inspired algorithm, which further infuses class representations into the adaption process. We evaluate the proposed method on a wide range of cross-domain object recognition tasks and achieve higher or comparable accuracy compared to existing SFDA methods. Ablation studies are conducted to verify the effectiveness of the proposed method.
Hao-Wei Yeh, Thomas Westfechtel, Jia-Bin Huang 0001, Tatsuya Harada
ICPR3
2022 Continuous and Diverse Image-to-Image Translation via Signed Attribute Vectors
Qi Mao 0002, Hung-Yu Tseng, Hsin-Ying Lee 0001, Jia-Bin Huang 0001, Siwei Ma 0001, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.4
2022 Learning to See Through Obstructions With Layered Decomposition
abstract
We present a learning-based approach for removing unwanted obstructions, such as window reflections, fence occlusions, or adherent raindrops, from a short sequence of images captured by a moving camera. Our method leverages motion differences between the background and obstructing elements to recover both layers. Specifically, we alternate between estimating dense optical flow fields of the two layers and reconstructing each layer from the flow-warped images via a deep convolutional neural network. This learning-based layer reconstruction module facilitates accommodating potential errors in the flow estimation and brittle assumptions, such as brightness consistency. We show that the proposed approach learned from synthetically generated data performs well to real images. Experimental results on numerous challenging scenarios of reflection and fence removal demonstrate the effectiveness of the proposed method.
Yu-Lun Liu 0001, Wei-Sheng Lai, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 DropLoss for Long-Tail Instance Segmentation
abstract
Long-tailed class distributions are prevalent among the practical applications of object detection and instance segmentation. Prior work in long-tail instance segmentation addresses the imbalance of losses between rare and frequent categories by reducing the penalty for a model incorrectly predicting a rare class label. We demonstrate that the rare categories are heavily suppressed by correct background predictions, which reduces the probability for all foreground categories with equal weight. Due to the relative infrequency of rare categories, this leads to an imbalance that biases towards predicting more frequent categories. Based on this insight, we develop DropLoss -- a novel adaptive loss to compensate for this imbalance without a trade-off between rare and frequent categories. With this loss, we show state-of-the-art mAP across rare, common, and frequent categories on the LVIS dataset. Codes are available at https://github.com/timy90022/DropLoss.
Ting-I Hsieh, Esther Robb, Hwann-Tzong Chen, Jia-Bin Huang 0001
AAAI4
2021 AMICO: Amodal Instance Composition
Peiye Zhuang, Denis Demandolx, Ayush Saraf, Xuejian Rong, Changil Kim 0001, Jia-Bin Huang 0001
BMVC6
2021 Robust Consistent Video Depth Estimation
abstract
We present an algorithm for estimating consistent dense depth maps and camera poses from a monocular video. We integrate a learning-based depth prior, in the form of a convolutional neural network trained for single-image depth estimation, with geometric optimization, to estimate a smooth camera trajectory as well as detailed and stable depth reconstruction. Our algorithm combines two complementary techniques: (1) flexible deformation-splines for low-frequency large-scale alignment and (2) geometry-aware depth filtering for high-frequency alignment of fine depth details. In contrast to prior approaches, our method does not require camera poses as input and achieves robust reconstruction for challenging hand-held cell phone captures containing a significant amount of noise, shake, motion blur, and rolling shutter deformations. Our method quantitatively outperforms state-of-the-arts on the Sintel benchmark for both depth and pose estimations and attains favorable qualitative results across diverse wild datasets.
Johannes Kopf 0001, Xuejian Rong, Jia-Bin Huang 0001
CVPR3
2021 Space-Time Neural Irradiance Fields for Free-Viewpoint Video
abstract
We present a method that learns a spatiotemporal neural irradiance field for dynamic scenes from a single video. Our learned representation enables free-viewpoint rendering of the input video. Our method builds upon recent advances in implicit representations. Learning a spatiotemporal irradiance field from a single video poses significant challenges because the video contains only one observation of the scene at any point in time. The 3D geometry of a scene can be legitimately represented in numerous ways since varying geometry (motion) can be explained with varying appearance and vice versa. We address this ambiguity by constraining the time-varying geometry of our dynamic scene representation using the scene depth estimated from video depth estimation methods, aggregating contents from individual frames into a single global representation. We provide an extensive quantitative evaluation and demonstrate compelling free-viewpoint rendering results.
Wenqi Xian, Jia-Bin Huang 0001, Johannes Kopf 0001, Changil Kim 0001
CVPR2
2021 Dynamic View Synthesis from Dynamic Monocular Video
abstract
We present an algorithm for generating novel views at arbitrary viewpoints and any input time step given a monocular video of a dynamic scene. Our work builds upon recent advances in neural implicit representation and uses continuous and differentiable functions for modeling the time-varying structure and the appearance of the scene. We jointly train a time-invariant static NeRF and a time-varying dynamic NeRF, and learn how to blend the results in an unsupervised manner. However, learning this implicit function from a single video is highly ill-posed (with infinitely many solutions that match the input video). To resolve the ambiguity, we introduce regularization losses to encourage a more physically plausible solution. We show extensive quantitative and qualitative results of dynamic view synthesis from casually captured videos.
Chen Gao 0003, Ayush Saraf, Johannes Kopf 0001, Jia-Bin Huang 0001
ICCV4
2021 Hybrid Neural Fusion for Full-frame Video Stabilization
abstract
Existing video stabilization methods often generate visible distortion or require aggressive cropping of frame boundaries, resulting in smaller field of views. In this work, we present a frame synthesis algorithm to achieve full-frame video stabilization. We first estimate dense warp fields from neighboring frames and then synthesize the stabilized frame by fusing the warped contents. Our core technical novelty lies in the learning-based hybrid-space fusion that alleviates artifacts caused by optical flow inaccuracy and fast-moving objects. We validate the effectiveness of our method on the NUS, selfie, and DeepStab video datasets. Extensive experiment results demonstrate the merits of our approach over prior video stabilization methods.
Yu-Lun Liu 0001, Wei-Sheng Lai, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001
ICCV5
2021 PseudoSeg: Designing Pseudo Labels for Semantic Segmentation
Yuliang Zou, Han Zhang 0010, Chun-Liang Li, Xiao Bian, Jia-Bin Huang 0001, Tomas Pfister
ICLR6
2021 Show, Match and Segment: Joint Weakly Supervised Learning of Semantic Matching and Object Co-Segmentation
abstract
We present an approach for jointly matching and segmenting object instances of the same category within a collection of images. In contrast to existing algorithms that tackle the tasks of semantic matching and object co-segmentation in isolation, our method exploits the complementary nature of the two tasks. The key insights of our method are two-fold. First, the estimated dense correspondence fields from semantic matching provide supervision for object co-segmentation by enforcing consistency between the predicted masks from a pair of images. Second, the predicted object masks from object co-segmentation in turn allow us to reduce the adverse effects due to background clutters for improving semantic matching. Our model is end-to-end trainable and does not require supervision from manually annotated correspondences and object masks. We validate the efficacy of our approach on five benchmark datasets: TSS, Internet, PF-PASCAL, PF-WILLOW, and SPair-71k, and show that our algorithm performs favorably against the state-of-the-art methods on both semantic matching and object co-segmentation tasks.
Yun-Chun Chen, Yen-Yu Lin, Ming-Hsuan Yang 0001, Jia-Bin Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Pose with style: detail-preserving pose-guided image synthesis with conditional StyleGAN
abstract
We present an algorithm for re-rendering a person from a single image under arbitrary poses. Existing methods often have difficulties in hallucinating occluded contents photo-realistically while preserving the identity and fine details in the source image. We first learn to inpaint the correspondence field between the body surface texture and the source image with a human body symmetry prior. The inpainted correspondence field allows us to transfer/warp local features extracted from the source to the target view even under large pose changes. Directly mapping the warped local features to an RGB image using a simple CNN decoder often leads to visible artifacts. Thus, we extend the StyleGAN generator so that it takes pose as input (for controlling poses) and introduces a spatially varying modulation for the latent space using the warped local features (for controlling appearances). We show that our method compares favorably against the state-of-the-art algorithms in both quantitative evaluation and visual comparison.
Badour AlBahar, Jingwan Lu, Jimei Yang, Zhixin Shu, Eli Shechtman, Jia-Bin Huang 0001
ACM Trans. Graph.6
2020 Learning to See Through Obstructions
abstract
We present a learning-based approach for removing unwanted obstructions, such as window reflections, fence occlusions or raindrops, from a short sequence of images captured by a moving camera. Our method leverages the motion differences between the background and the obstructing elements to recover both layers. Specifically, we alternate between estimating dense optical flow fields of the two layers and reconstructing each layer from the flow-warped images via a deep convolutional neural network. The learning-based layer reconstruction allows us to accommodate potential errors in the flow estimation and brittle assumptions such as brightness consistency. We show that training on synthetically generated data transfers well to real images. Our results on numerous challenging scenarios of reflection and fence removal demonstrate the effectiveness of the proposed method.
Yu-Lun Liu 0001, Wei-Sheng Lai, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001
CVPR5
2020 Single-Image HDR Reconstruction by Learning to Reverse the Camera Pipeline
abstract
Recovering a high dynamic range (HDR) image from a single low dynamic range (LDR) input image is challenging due to missing details in under-/over-exposed regions caused by quantization and saturation of camera sensors. In contrast to existing learning-based methods, our core idea is to incorporate the domain knowledge of the LDR image formation pipeline into our model. We model the HDR-to-LDR image formation pipeline as the (1) dynamic range clipping, (2) non-linear mapping from a camera response function, and (3) quantization. We then propose to learn three specialized CNNs to reverse these steps. By decomposing the problem into specific sub-tasks, we impose effective physical constraints to facilitate the training of individual sub-networks. Finally, we jointly fine-tune the entire model end-to-end to reduce error accumulation. With extensive quantitative and qualitative experiments on diverse image datasets, we demonstrate that the proposed method performs favorably against state-of-the-art single-image HDR reconstruction algorithms.
Yu-Lun Liu 0001, Wei-Sheng Lai, Yu-Sheng Chen, Yi-Lung Kao, Ming-Hsuan Yang 0001, Yung-Yu Chuang, Jia-Bin Huang 0001
CVPR7
2020 3D Photography Using Context-Aware Layered Depth Inpainting
abstract
We propose a method for converting a single RGB-D input image into a 3D photo, i.e., a multi-layer representation for novel view synthesis that contains hallucinated color and depth structures in regions occluded in the original view. We use a Layered Depth Image with explicit pixel connectivity as underlying representation, and present a learning-based inpainting model that iteratively synthesizes new local color-and-depth content into the occluded region in a spatial context-aware manner. The resulting 3D photos can be efficiently rendered with motion parallax using standard graphics engines. We validate the effectiveness of our method on a wide range of challenging everyday scenes and show less artifacts when compared with the state-of-the-arts.
Meng-Li Shih, Shih-Yang Su, Johannes Kopf 0001, Jia-Bin Huang 0001
CVPR4
2020 Instance-Aware Image Colorization
abstract
Image colorization is inherently an ill-posed problem with multi-modal uncertainty. Previous methods leverage the deep neural network to map input grayscale images to plausible color outputs directly. Although these learning-based methods have shown impressive performance, they usually fail on the input images that contain multiple objects. The leading cause is that existing models perform learning and colorization on the entire image. In the absence of a clear figure-ground separation, these models cannot effectively locate and learn meaningful object-level semantics. In this paper, we propose a method for achieving instance-aware colorization. Our network architecture leverages an off-the-shelf object detector to obtain cropped object images and uses an instance colorization network to extract object-level features. We use a similar network to extract the full-image features and apply a fusion module to full object-level and image-level features to predict the final colors. Both colorization networks and fusion modules are learned from a large-scale dataset. Experimental results show that our work outperforms existing methods on different quality metrics and achieves state-of-the-art performance on image colorization.
Jheng-Wei Su, Hung-Kuo Chu, Jia-Bin Huang 0001
CVPR3
2020 NAS-DIP: Learning Deep Image Prior with Neural Architecture Search
Yun-Chun Chen, Chen Gao 0003, Esther Robb, Jia-Bin Huang 0001
ECCV (18)4
2020 Shuffle and Attend: Video Domain Adaptation
Jinwoo Choi 0001, Samuel Schulter, Jia-Bin Huang 0001
ECCV (12)4
2020 Flow-edge Guided Video Completion
Chen Gao 0003, Ayush Saraf, Jia-Bin Huang 0001, Johannes Kopf 0001
ECCV (12)3
2020 DRG: Dual Relation Graph for Human-Object Interaction Detection
Chen Gao 0003, Yuliang Zou, Jia-Bin Huang 0001
ECCV (12)4
2020 Semantic View Synthesis
Hsin-Ping Huang, Hung-Yu Tseng, Hsin-Ying Lee 0001, Jia-Bin Huang 0001
ECCV (12)4
2020 FeatMatch: Feature-Based Augmentation for Semi-supervised Learning
Chia-Wen Kuo, Chih-Yao Ma, Jia-Bin Huang 0001, Zsolt Kira
ECCV (18)3
2020 Learning Monocular Visual Odometry via Self-Supervised Long-Term Modeling
Yuliang Zou, Pan Ji, Quoc-Huy Tran, Jia-Bin Huang 0001, Manmohan Krishna Chandraker
ECCV (14)4
2020 Cross-Domain Few-Shot Classification via Learned Feature-Wise Transformation
Hung-Yu Tseng, Hsin-Ying Lee 0001, Jia-Bin Huang 0001, Ming-Hsuan Yang 0001
ICLR3
2020 Unsupervised and Semi-Supervised Domain Adaptation for Action Recognition from Drones
abstract
We address the problem of human action classification in drone videos. Due to the high cost of capturing and labeling large-scale drone videos with diverse actions, we present unsupervised and semi-supervised domain adaptation approaches that leverage both the existing fully annotated action recognition datasets and unannotated (or only a few annotated) videos from drones. To study the emerging problem of drone-based action recognition, we create a new dataset, NEC-DRONE, containing 5,250 videos to evaluate the task. We tackle both problem settings with 1) same and 2) different action label sets for the source (e.g., Kinectics dataset) and target domains (drone videos). We present a combination of video and instance-based adaptation methods, paired with either a classifier or an embedding-based framework to transfer the knowledge from source to target. Our results show that the proposed adaptation approach substantially improves the performance on these challenging and practical tasks. We further demonstrate the applicability of our method for learning cross-view action recognition on the Charades-Ego dataset. We provide qualitative analysis to understand the behaviors of our approaches.
Jinwoo Choi 0001, Manmohan Krishna Chandraker, Jia-Bin Huang 0001
WACV4
2020 Reducing Footskate in Human Motion Reconstruction with Ground Contact Constraints
abstract
In this paper, we aim to reduce the footskate artifacts when reconstructing human dynamics from monocular RGB videos. Recent work has made substantial progress in improving the temporal smoothness of the reconstructed motion trajectories. Their results, however, still suffer from severe foot skating and slippage artifacts. To tackle this issue, we present a neural network based detector for localizing ground contact events of human feet and use it to impose a physical constraint for optimization of the whole human dynamics in a video. We present a detailed study on the proposed ground contact detector and demonstrate high-quality human motion reconstruction results in various videos.
Yuliang Zou, Jimei Yang, Duygu Ceylan, Jianming Zhang 0001, Federico Perazzi, Jia-Bin Huang 0001
WACV6
2020 DRIT++: Diverse Image-to-Image Translation via Disentangled Representations
Hsin-Ying Lee 0001, Hung-Yu Tseng, Qi Mao 0002, Jia-Bin Huang 0001, Yu-Ding Lu, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.4
2020 Tracking Persons-of-Interest via Unsupervised Representation Adaptation
Jia-Bin Huang 0001, Jongwoo Lim, Yihong Gong, Jinjun Wang, Narendra Ahuja, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.2
2020 Progressive Representation Adaptation for Weakly Supervised Object Localization
abstract
We address the problem of weakly supervised object localization where only image-level annotations are available for training object detectors. Numerous methods have been proposed to tackle this problem through mining object proposals. However, a substantial amount of noise in object proposals causes ambiguities for learning discriminative object models. Such approaches are sensitive to model initialization and often converge to undesirable local minimum solutions. In this paper, we propose to overcome these drawbacks by progressive representation adaptation with two main steps: 1) classification adaptation and 2) detection adaptation. In classification adaptation, we transfer a pre-trained network to a multi-label classification task for recognizing the presence of a certain object in an image. Through the classification adaptation step, the network learns discriminative representations that are specific to object categories of interest. In detection adaptation, we mine class-specific object proposals by exploiting two scoring strategies based on the adapted classification network. Class-specific proposal mining helps remove substantial noise from the background clutter and potential confusion from similar objects. We further refine these proposals using multiple instance learning and segmentation cues. Using these refined object bounding boxes, we fine-tune all the layer of the classification network and obtain a fully adapted detection network. We present detailed experimental validation on the PASCAL VOC and ILSVRC datasets. Experimental results demonstrate that our progressive representation adaptation algorithm performs favorably against the state-of-the-art methods.
Dong Li 0025, Jia-Bin Huang 0001, Yali Li 0001, Shengjin Wang, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Consistent video depth estimation
abstract
We present an algorithm for reconstructing dense, geometrically consistent depth for all pixels in a monocular video. We leverage a conventional structure-from-motion reconstruction to establish geometric constraints on pixels in the video. Unlike the ad-hoc priors in classical reconstruction, we use a learning-based prior, i.e., a convolutional neural network trained for single-image depth estimation. At test time, we fine-tune this network to satisfy the geometric constraints of a particular input video, while retaining its ability to synthesize plausible depth details in parts of the video that are less constrained. We show through quantitative validation that our method achieves higher accuracy and a higher degree of geometric consistency than previous monocular reconstruction methods. Visually, our results appear more stable. Our algorithm is able to handle challenging hand-held captured input videos with a moderate degree of dynamic motion. The improved quality of the reconstruction enables several applications, such as scene reconstruction and advanced video-based visual effects.
Jia-Bin Huang 0001, Richard Szeliski, Kevin Matzen, Johannes Kopf 0001
ACM Trans. Graph.2
2019 Connecting the Digital and Physical World: Improving the Robustness of Adversarial Attacks
abstract
While deep learning models have achieved unprecedented success in various domains, there is also a growing concern of adversarial attacks against related applications. Recent results show that by adding a small amount of perturbations to an image (imperceptible to humans), the resulting adversarial examples can force a classifier to make targeted mistakes. So far, most existing works focus on crafting adversarial examples in the digital domain, while limited efforts have been devoted to understanding the physical domain attacks. In this work, we explore the feasibility of generating robust adversarial examples that remain effective in the physical domain. Our core idea is to use an image-to-image translation network to simulate the digital-to-physical transformation process for generating robust adversarial examples. To validate our method, we conduct a large-scale physical-domain experiment, which involves manually taking more than 3000 physical domain photos. The results show that our method outperforms existing ones by a large margin and demonstrates a high level of robustness and transferability.
Steve T. K. Jan, Joseph Messou, Yen-Chen Lin, Jia-Bin Huang 0001, Gang Wang 0011
AAAI4
2019 CrDoCo: Pixel-Level Domain Transfer With Cross-Domain Consistency
abstract
Unsupervised domain adaptation algorithms aim to transfer the knowledge learned from one domain to another (e.g., synthetic to real images). The adapted representations often do not capture pixel-level domain shifts that are crucial for dense prediction tasks (e.g., semantic segmentation). In this paper, we present a novel pixel-wise adversarial domain adaptation algorithm. By leveraging image-to-image translation methods for data augmentation, our key insight is that while the translated images between domains may differ in styles, their predictions for the task should be consistent. We exploit this property and introduce a cross-domain consistency loss that enforces our adapted model to produce consistent predictions. Through extensive experimental results, we show that our method compares favorably against the state-of-the-art on a wide variety of unsupervised domain adaptation tasks.
Yun-Chun Chen, Yen-Yu Lin, Ming-Hsuan Yang 0001, Jia-Bin Huang 0001
CVPR4
2019 SAIL-VOS: Semantic Amodal Instance Level Video Object Segmentation - A Synthetic Dataset and Baselines
abstract
We introduce SAIL-VOS (Semantic Amodal Instance Level Video Object Segmentation), a new dataset aiming to stimulate semantic amodal segmentation research. Humans can effortlessly recognize partially occluded objects and reliably estimate their spatial extent beyond the visible. However, few modern computer vision techniques are capable of reasoning about occluded parts of an object. This is partly due to the fact that very few image datasets and no video dataset exist which permit development of those methods. To address this issue, we present a synthetic dataset extracted from the photo-realistic game GTA-V. Each frame is accompanied with densely annotated, pixel-accurate visible and amodal segmentation masks with semantic labels. More than 1.8M objects are annotated resulting in 100 times more annotations than existing datasets. We demonstrate the challenges of the dataset by quantifying the performance of several baselines. Data and additional material is available at http://sailvos.web.illinois.edu.
Yuan-Ting Hu, Hong-Shuo Chen, Kexin Hui, Jia-Bin Huang 0001, Alexander G. Schwing
CVPR4
2019 Guided Image-to-Image Translation With Bi-Directional Feature Transformation
abstract
We address the problem of guided image-to-image translation where we translate an input image into another while respecting the constraints provided by an external, user-provided guidance image. Various types of conditioning mechanisms for leveraging the given guidance image have been explored, including input concatenation, feature concatenation, and conditional affine transformation of feature activations. All these conditioning mechanisms, however, are uni-directional, i.e., no information flow from the input image back to the guidance. To better utilize the constraints of the guidance image, we present a bi-directional feature transformation (bFT) scheme. We show that our novel bFT scheme outperforms other conditioning schemes and has comparable results to state-of-the-art methods on different tasks.
Badour AlBahar, Jia-Bin Huang 0001
ICCV2
2019 A Closer Look at Few-shot Classification
Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, Jia-Bin Huang 0001
ICLR (Poster)5
2019 Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition
abstract
Human activities often occur in specific scene contexts, e.g., playing basketball on a basketball court. Training a model using existing video datasets thus inevitably captures and leverages such bias (instead of using the actual discriminative cues). The learned representation may not generalize well to new action classes or different tasks. In this paper, we propose to mitigate scene bias for video representation learning. Specifically, we augment the standard cross-entropy loss for action classification with 1) an adversarial loss for scene types and 2) a human mask confusion loss for videos where the human actors are masked out. These two losses encourage learning representations that are unable to predict the scene types and the correct actions when there is no evidence. We validate the effectiveness of our method by transferring our pre-trained model to three different tasks, including action classification, temporal localization, and spatio-temporal action detection. Our results show consistent improvement over the baseline model without debiasing.
Jinwoo Choi 0001, Chen Gao 0003, Joseph Messou, Jia-Bin Huang 0001
NeurIPS4
2019 Robust Visual Tracking via Hierarchical Convolutional Features
abstract
Visual tracking is challenging as target objects often undergo significant appearance changes caused by deformation, abrupt motion, background clutter and occlusion. In this paper, we propose to exploit the rich hierarchical features of deep convolutional neural networks to improve the accuracy and robustness of visual tracking. Deep neural networks trained on object recognition datasets consist of multiple convolutional layers. These layers encode target appearance with different levels of abstraction. For example, the outputs of the last convolutional layers encode the semantic information of targets and such representations are invariant to significant appearance variations. However, their spatial resolutions are too coarse to precisely localize the target. In contrast, features from earlier convolutional layers provide more precise localization but are less invariant to appearance changes. We interpret the hierarchical features of convolutional layers as a nonlinear counterpart of an image pyramid representation and explicitly exploit these multiple levels of abstraction to represent target objects. Specifically, we learn adaptive correlation filters on the outputs from each convolutional layer to encode the target appearance. We infer the maximum response of each layer to locate targets in a coarse-to-fine manner. To further handle the issues with scale estimation and re-detecting target objects from tracking failures caused by heavy occlusion or out-of-the-view movement, we conservatively learn another correlation filter, that maintains a long-term memory of target appearance, as a discriminative classifier. We apply the classifier to two types of object proposals: (1) proposals with a small step size and tightly around the estimated location for scale estimation; and (2) proposals with large step size and across the whole image for target re-detection. Extensive experimental results on large-scale benchmark datasets show that the proposed algorithm performs favorably against the state-of-the-art tracking methods.
Chao Ma 0004, Jia-Bin Huang 0001, Xiaokang Yang 0001, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Fast and Accurate Image Super-Resolution with Deep Laplacian Pyramid Networks
abstract
Convolutional neural networks have recently demonstrated high-quality reconstruction for single image super-resolution. However, existing methods often require a large number of network parameters and entail heavy computational loads at runtime for generating high-accuracy super-resolution results. In this paper, we propose the deep Laplacian Pyramid Super-Resolution Network for fast and accurate image super-resolution. The proposed network progressively reconstructs the sub-band residuals of high-resolution images at multiple pyramid levels. In contrast to existing methods that involve the bicubic interpolation for pre-processing (which results in large feature maps), the proposed method directly extracts features from the low-resolution input space and thereby entails low computational loads. We train the proposed network with deep supervision using the robust Charbonnier loss functions and achieve high-quality image reconstruction. Furthermore, we utilize the recursive layers to share parameters across as well as within pyramid levels, and thus drastically reduce the number of parameters. Extensive quantitative and qualitative evaluations on benchmark datasets show that the proposed algorithm performs favorably against the state-of-the-art methods in terms of run-time and image quality.
Wei-Sheng Lai, Jia-Bin Huang 0001, Narendra Ahuja, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Joint Image Filtering with Deep Convolutional Networks
abstract
Joint image filters leverage the guidance image as a prior and transfer the structural details from the guidance image to the target image for suppressing noise or enhancing spatial resolution. Existing methods either rely on various explicit filter constructions or hand-designed objective functions, thereby making it difficult to understand, improve, and accelerate these filters in a coherent framework. In this paper, we propose a learning-based approach for constructing joint filters based on Convolutional Neural Networks. In contrast to existing methods that consider only the guidance image, the proposed algorithm can selectively transfer salient structures that are consistent with both guidance and target images. We show that the model trained on a certain type of data, e.g., RGB and depth images, generalizes well to other modalities, e.g., flash/non-Flash and RGB/NIR images. We validate the effectiveness of the proposed joint filter through extensive experimental evaluations with state-of-the-art methods.
Yijun Li 0001, Jia-Bin Huang 0001, Narendra Ahuja, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Deep Semantic Matching with Foreground Detection and Cycle-Consistency
Yun-Chun Chen, Po-Hsiang Huang, Li-Yu Yu, Jia-Bin Huang 0001, Ming-Hsuan Yang 0001, Yen-Yu Lin
ACCV (3)4
2018 iCAN: Instance-Centric Attention Network for Human-Object Interaction Detection
Chen Gao 0003, Yuliang Zou, Jia-Bin Huang 0001
BMVC3
2018 DeepMVS: Learning Multi-View Stereopsis
abstract
We present DeepMVS, a deep convolutional neural network (ConvNet) for multi-view stereo reconstruction. Taking an arbitrary number of posed images as input, we first produce a set of plane-sweep volumes and use the proposed DeepMVS network to predict high-quality disparity maps. The key contributions that enable these results are (1) supervised pretraining on a photorealistic synthetic dataset, (2) an effective method for aggregating information across a set of unordered images, and (3) integrating multi-layer feature activations from the pre-trained VGG-19 network. We validate the efficacy of DeepMVS using the ETH3D Benchmark. Our results show that DeepMVS compares favorably against state-of-the-art conventional MVS algorithms and other ConvNet based methods, particularly for near-textureless regions and thin structures.
Kevin Matzen, Johannes Kopf 0001, Narendra Ahuja, Jia-Bin Huang 0001
CVPR5
2018 Unsupervised Video Object Segmentation Using Motion Saliency-Guided Spatio-Temporal Propagation
Yuan-Ting Hu, Jia-Bin Huang 0001, Alexander G. Schwing
ECCV (1)2
2018 VideoMatch: Matching Based Video Object Segmentation
Yuan-Ting Hu, Jia-Bin Huang 0001, Alexander G. Schwing
ECCV (8)2
2018 Learning Blind Video Temporal Consistency
Wei-Sheng Lai, Jia-Bin Huang 0001, Oliver Wang, Eli Shechtman, Ersin Yumer, Ming-Hsuan Yang 0001
ECCV (15)2
2018 Diverse Image-to-Image Translation via Disentangled Representations
Hsin-Ying Lee 0001, Hung-Yu Tseng, Jia-Bin Huang 0001, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
ECCV (1)3
2018 DF-Net: Unsupervised Joint Learning of Depth and Flow Using Cross-Task Consistency
Yuliang Zou, Zelun Luo, Jia-Bin Huang 0001
ECCV (5)3
2018 Ensemble convolutional neural networks for pose estimation
Yuki Kawana, Norimichi Ukita, Jia-Bin Huang 0001, Ming-Hsuan Yang 0001
Comput. Vis. Image Underst.3
2018 Adaptive Correlation Filters with Long-Term and Short-Term Memory for Object Tracking
Chao Ma 0004, Jia-Bin Huang 0001, Xiaokang Yang 0001, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.2
2018 Multi-view wire art
abstract
Wire art is the creation of three-dimensional sculptural art using wire strands. As the 2D projection of a 3D wire sculpture forms line drawing patterns, it is possible to craft multi-view wire sculpture art --- a static sculpture with multiple (potentially very different) interpretations when perceived at different viewpoints. Artists can effectively leverage this characteristic and produce compelling artistic effects. However, the creation of such multi-view wire sculpture is extremely time-consuming even by highly skilled artists. In this paper, we present a computational framework for automatic creation of multi-view 3D wire sculpture. Our system takes two or three user-specified line drawings and the associated viewpoints as inputs. We start with producing a sparse set of voxels via greedy selection approach such that their projections on the virtual cameras cover all the contour pixels of the input line drawings. The sparse set of voxels, however, do not necessary form one single connected component. We introduce a constrained 3D pathfinding algorithm to link isolated groups of voxels into a connected component while maintaining the similarity between the projected voxels and the line drawings. Using the reconstructed visual hull, we extract a curve skeleton and produce a collection of smooth 3D curves by fitting cubic splines and optimizing the curve deformation to best approximate the provided line drawings. We demonstrate the effectiveness of our system for creating compelling multi-view wire sculptures in both simulation and 3D physical printouts.
Kai-Wen Hsiao, Jia-Bin Huang 0001, Hung-Kuo Chu
ACM Trans. Graph.2
2017 HOMER: An Interactive System for Home Based Stroke Rehabilitation
abstract
Delivering long term, unsupervised stroke rehabilitation in the home is a complex challenge that requires robust, low cost, scalable, and engaging solutions. We present HOMER, an interactive system that uses novel therapy artifacts, a computer vision approach, and a tablet interface to provide users with a flexible solution suitable for home based rehabilitation. HOMER builds on our prior work developing systems for lightly supervised rehabilitation use in the clinic, by identifying key features for functional movement analysis, adopting a simplified classification assessment approach, and supporting transferability of therapy outcomes to daily living experiences through the design of novel rehabilitation artifacts. A small pilot study with unimpaired subjects indicates the potential of the system in effectively assessing movement and establishing a creative environment for training.
Aisling Kelliher, Jinwoo Choi 0001, Jia-Bin Huang 0001, Thanassis Rikakis, Kris Makoto Kitani
ASSETS3
2017 Deep Laplacian Pyramid Networks for Fast and Accurate Super-Resolution
abstract
Convolutional neural networks have recently demonstrated high-quality reconstruction for single-image super-resolution. In this paper, we propose the Laplacian Pyramid Super-Resolution Network (LapSRN) to progressively reconstruct the sub-band residuals of high-resolution images. At each pyramid level, our model takes coarse-resolution feature maps as input, predicts the high-frequency residuals, and uses transposed convolutions for upsampling to the finer level. Our method does not require the bicubic interpolation as the pre-processing step and thus dramatically reduces the computational complexity. We train the proposed LapSRN with deep supervision using a robust Charbonnier loss function and achieve high-quality reconstruction. Furthermore, our network generates multi-scale predictions in one feed-forward pass through the progressive reconstruction, thereby facilitates resource-aware applications. Extensive quantitative and qualitative evaluations on benchmark datasets show that the proposed algorithm performs favorably against the state-of-the-art methods in terms of speed and accuracy.
Wei-Sheng Lai, Jia-Bin Huang 0001, Narendra Ahuja, Ming-Hsuan Yang 0001
CVPR2
2017 Unsupervised Representation Learning by Sorting Sequences
abstract
We present an unsupervised representation learning approach using videos without semantic labels. We leverage the temporal coherence as a supervisory signal by formulating representation learning as a sequence sorting task. We take temporally shuffled frames (i.e., in non-chronological order) as inputs and train a convolutional neural network to sort the shuffled sequences. Similar to comparison-based sorting algorithms, we propose to extract features from all frame pairs and aggregate them to predict the correct order. As sorting shuffled image sequence requires an understanding of the statistical temporal structure of images, training with such a proxy task allows us to learn rich and generalizable visual representation. We validate the effectiveness of the learned representation using our method as pre-training on high-level recognition problems. The experimental results show that our method compares favorably against state-of-the-art methods on action recognition, image classification, and object detection tasks.
Hsin-Ying Lee 0001, Jia-Bin Huang 0001, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
ICCV2
2017 MaskRNN: Instance Level Video Object Segmentation
abstract
Instance level video object segmentation is an important technique for video editing and compression. To capture the temporal coherence, in this paper, we develop MaskRNN, a recurrent neural net approach which fuses in each frame the output of two deep nets for each object instance - a binary segmentation net providing a mask and a localization net providing a bounding box. Due to the recurrent component and the localization component, our method is able to take advantage of long-term temporal structures of the video data as well as rejecting outliers. We validate the proposed algorithm on three challenging benchmark datasets, the DAVIS-2016 dataset, the DAVIS-2017 dataset, and the Segtrack v2 dataset, achieving state-of-the-art performance on all of them.
Yuan-Ting Hu, Jia-Bin Huang 0001, Alexander G. Schwing
NIPS2
2017 Semi-Supervised Learning for Optical Flow with Generative Adversarial Networks
abstract
Convolutional neural networks (CNNs) have recently been applied to the optical flow estimation problem. As training the CNNs requires sufficiently large ground truth training data, existing approaches resort to synthetic, unrealistic datasets. On the other hand, unsupervised methods are capable of leveraging real-world videos for training where the ground truth flow fields are not available. These methods, however, rely on the fundamental assumptions of brightness constancy and spatial smoothness priors which do not hold near motion boundaries. In this paper, we propose to exploit unlabeled videos for semi-supervised learning of optical flow with a Generative Adversarial Network. Our key insight is that the adversarial loss can capture the structural patterns of flow warp errors without making explicit assumptions. Extensive experiments on benchmark datasets demonstrate that the proposed semi-supervised algorithm performs favorably against purely supervised and semi-supervised learning schemes.
Wei-Sheng Lai, Jia-Bin Huang 0001, Ming-Hsuan Yang 0001
NIPS2
2016 Detecting Migrating Birds at Night
abstract
Bird migration is a critical indicator of environmental health, biodiversity, and climate change. Existing techniques for monitoring bird migration are either expensive (e.g., satellite tracking), labor-intensive (e.g., moon watching), indirect and thus less accurate (e.g., weather radar), or intrusive (e.g., attaching geolocators on captured birds). In this paper, we present a vision-based system for detecting migrating birds in flight at night. Our system takes stereo videos of the night sky as inputs, detects multiple flying birds and estimates their orientations, speeds, and altitudes. The main challenge lies in detecting flying birds of unknown trajectories under high noise level due to the low-light environment. We address this problem by incorporating stereo constraints for rejecting physically implausible configurations and gathering evidence from two (or more) views. Specifically, we develop a robust stereo-based 3D line fitting algorithm for geometric verification and a deformable part response accumulation strategy for trajectory verification. We demonstrate the effectiveness of the proposed approach through quantitative evaluation of real videos of birds migrating at night collected with near-infrared cameras.
Jia-Bin Huang 0001, Rich Caruana, Andrew Farnsworth, Steve Kelling, Narendra Ahuja
CVPR1
2016 A Comparative Study for Single Image Blind Deblurring
abstract
Numerous single image blind deblurring algorithms have been proposed to restore latent sharp images under camera motion. However, these algorithms are mainly evaluated using either synthetic datasets or few selected real blurred images. It is thus unclear how these algorithms would perform on images acquired "in the wild" and how we could gauge the progress in the field. In this paper, we aim to bridge this gap. We present the first comprehensive perceptual study and analysis of single image blind deblurring using real-world blurred images. First, we collect a dataset of real blurred images and a dataset of synthetically blurred images. Using these datasets, we conduct a large-scale user study to quantify the performance of several representative state-of-the-art blind deblurring algorithms. Second, we systematically analyze subject preferences, including the level of agreement, significance tests of score differences, and rationales for preferring one method over another. Third, we study the correlation between human subjective scores and several full-reference and noreference image quality metrics. Our evaluation and analysis indicate the performance gap between synthetically blurred images and real blurred image and sheds light on future research in single image blind deblurring.
Wei-Sheng Lai, Jia-Bin Huang 0001, Narendra Ahuja, Ming-Hsuan Yang 0001
CVPR2
2016 Weakly Supervised Object Localization with Progressive Domain Adaptation
abstract
We address the problem of weakly supervised object localization where only image-level annotations are available for training. Many existing approaches tackle this problem through object proposal mining. However, a substantial amount of noise in object proposals causes ambiguities for learning discriminative object models. Such approaches are sensitive to model initialization and often converge to an undesirable local minimum. In this paper, we address this problem by progressive domain adaptation with two main steps: classification adaptation and detection adaptation. In classification adaptation, we transfer a pre-trained network to our multi-label classification task for recognizing the presence of a certain object in an image. In detection adaptation, we first use a mask-out strategy to collect class-specific object proposals and apply multiple instance learning to mine confident candidates. We then use these selected object proposals to fine-tune all the layers, resulting in a fully adapted detection network. We extensively evaluate the localization performance on the PASCAL VOC and ILSVRC datasets and demonstrate significant performance improvement over the state-of-the-art methods.
Dong Li 0025, Jia-Bin Huang 0001, Yali Li 0001, Shengjin Wang, Ming-Hsuan Yang 0001
CVPR2
2016 Deep Joint Image Filtering
Yijun Li 0001, Jia-Bin Huang 0001, Narendra Ahuja, Ming-Hsuan Yang 0001
ECCV (4)2
2016 Unsupervised Visual Representation Learning by Graph-Based Consistent Constraints
Dong Li 0025, Wei-Chih Hung, Jia-Bin Huang 0001, Shengjin Wang, Narendra Ahuja, Ming-Hsuan Yang 0001
ECCV (4)3
2016 Tracking Persons-of-Interest via Adaptive Discriminative Features
Yihong Gong, Jia-Bin Huang 0001, Jongwoo Lim, Jinjun Wang, Narendra Ahuja, Ming-Hsuan Yang 0001
ECCV (5)3
2016 Temporally coherent completion of dynamic video
abstract
We present an automatic video completion algorithm that synthesizes missing regions in videos in a temporally coherent fashion. Our algorithm can handle dynamic scenes captured using a moving camera. State-of-the-art approaches have difficulties handling such videos because viewpoint changes cause image-space motion vectors in the missing and known regions to be inconsistent. We address this problem by jointly estimating optical flow and color in the missing regions. Using pixel-wise forward/backward flow fields enables us to synthesize temporally coherent colors. We formulate the problem as a non-parametric patch-based optimization. We demonstrate our technique on numerous challenging videos.
Jia-Bin Huang 0001, Sing Bing Kang, Narendra Ahuja, Johannes Kopf 0001
ACM Trans. Graph.1
2015 Single image super-resolution from transformed self-exemplars
abstract
Self-similarity based super-resolution (SR) algorithms are able to produce visually pleasing results without extensive training on external databases. Such algorithms exploit the statistical prior that patches in a natural image tend to recur within and across scales of the same image. However, the internal dictionary obtained from the given image may not always be sufficiently expressive to cover the textural appearance variations in the scene. In this paper, we extend self-similarity based SR to overcome this drawback. We expand the internal patch search space by allowing geometric variations. We do so by explicitly localizing planes in the scene and using the detected perspective geometry to guide the patch search process. We also incorporate additional affine transformations to accommodate local shape variations. We propose a compositional model to simultaneously handle both types of transformations. We extensively evaluate the performance in both urban and natural scenes. Even without using any external training databases, we achieve significantly superior results on urban scenes, while maintaining comparable performance on natural scenes as other state-of-the-art SR algorithms.
Jia-Bin Huang 0001, Abhishek Singh 0002, Narendra Ahuja
CVPR1
2015 Hierarchical Convolutional Features for Visual Tracking
abstract
Visual object tracking is challenging as target objects often undergo significant appearance changes caused by deformation, abrupt motion, background clutter and occlusion. In this paper, we exploit features extracted from deep convolutional neural networks trained on object recognition datasets to improve tracking accuracy and robustness. The outputs of the last convolutional layers encode the semantic information of targets and such representations are robust to significant appearance variations. However, their spatial resolution is too coarse to precisely localize targets. In contrast, earlier convolutional layers provide more precise localization but are less invariant to appearance changes. We interpret the hierarchies of convolutional layers as a nonlinear counterpart of an image pyramid representation and exploit these multiple levels of abstraction for visual tracking. Specifically, we adaptively learn correlation filters on each convolutional layer to encode the target appearance. We hierarchically infer the maximum response of each layer to locate targets. Extensive experimental results on a largescale benchmark dataset show that the proposed algorithm performs favorably against state-of-the-art methods.
Chao Ma 0004, Jia-Bin Huang 0001, Xiaokang Yang 0001, Ming-Hsuan Yang 0001
ICCV2
2014 Towards accurate and robust cross-ratio based gaze trackers through learning from simulation
abstract
Cross-ratio (CR) based methods offer many attractive properties for remote gaze estimation using a single camera in an uncalibrated setup by exploiting invariance of a plane projectivity. Unfortunately, due to several simplification assumptions, the performance of CR-based eye gaze trackers decays significantly as the subject moves away from the calibration position. In this paper, we introduce an adaptive homography mapping for achieving gaze prediction with higher accuracy at the calibration position and more robustness under head movements. This is achieved with a learning-based method for compensating both spatially-varying gaze errors and head pose dependent errors simultaneously in a unified framework. The model of adaptive homography is trained offline using simulated data, saving a tremendous amount of time in data collection. We validate the effectiveness of the proposed approach using both simulated and real data from a physical setup. We show that our method compares favorably against other state-of-the-art CR based methods.
Jia-Bin Huang 0001, Qin Cai, Zicheng Liu 0001, Narendra Ahuja, Zhengyou Zhang
ETRA1
2014 Image completion using planar structure guidance
abstract
We propose a method for automatically guiding patch-based image completion using mid-level structural cues. Our method first estimates planar projection parameters, softly segments the known region into planes, and discovers translational regularity within these planes. This information is then converted into soft constraints for the low-level completion algorithm by defining prior probabilities for patch offsets and transformations. Our method handles multiple planes, and in the absence of any detected planes falls back to a baseline fronto-parallel image completion algorithm. We validate our technique through extensive comparisons with state-of-the-art algorithms on a variety of scenes.
Jia-Bin Huang 0001, Sing Bing Kang, Narendra Ahuja, Johannes Kopf 0001
ACM Trans. Graph.1
2013 Transformation guided image completion
abstract
In this paper, we describe a new interactive image completion system that allows users to easily specify various forms of mid-level structures in the image. Our system supports the specification of four basic symmetric types: reflection, translation, rotation, and glide. The user inputs are automatically converted into guidance maps that encode possible candidate shifts and, indirectly, local transformations of rotation and scale. These guidance maps are used in conjunction with a color matching cost for image completion. We show that our system is capable of handling a variety of challenging examples.
Jia-Bin Huang 0001, Johannes Kopf 0001, Narendra Ahuja, Sing Bing Kang
ICCP1
2012 Saliency detection via divergence analysis: A unified perspective
Jia-Bin Huang 0001, Narendra Ahuja
ICPR1
2010 Exploiting Self-similarities for Single Frame Super-Resolution
Chih-Yuan Yang, Jia-Bin Huang 0001, Ming-Hsuan Yang 0001
ACCV (3)2
2010 Fast sparse representation with prototypes
abstract
Sparse representation has found applications in numerous domains and recent developments have been focused on the convex relaxation of the lo-norm minimization for sparse coding (i.e., the ℓ1-norm minimization). Nevertheless, the time and space complexities of these algorithms remain significantly high for large-scale problems. As signals in most problems can be modeled by a small set of prototypes, we propose an algorithm that exploits this property and show that the ℓ1-norm minimization problem can be reduced to a much smaller problem, thereby gaining significant speed-ups with much less memory requirements. Experimental results demonstrate that our algorithm is able to achieve double-digit gain in speed with much less memory requirement than the state-of-the-art algorithms.
Jia-Bin Huang 0001, Ming-Hsuan Yang 0001
CVPR1
2010 Single image deblurring with adaptive dictionary learning
abstract
We propose a motion deblurring algorithm that exploits sparsity constraints of image patches using one single frame. In our formulation, each image patch is encoded with sparse coefficients using an over-complete dictionary. The sparsity constraints facilitate recovering the latent image without solving an ill-posed deconvolution problem. In addition, the dictionary is learned and updated directly from one single frame without using additional images. The proposed method iteratively utilizes sparsity constraints to recover latent image, estimates the deblur kernel, and updates the dictionary directly from one single image. The final deblurred image is then recovered once the deblur kernel is estimated using our method. Experiments show that the proposed algorithm achieves favorable results against the state-of-the-art methods.
Jia-Bin Huang 0001, Ming-Hsuan Yang 0001
ICIP2
2009 Estimating Human Pose from Occluded Images
Jia-Bin Huang 0001, Ming-Hsuan Yang 0001
ACCV (1)1
2009 Moving cast shadow detection using physics-based features
abstract
Cast shadows induced by moving objects often cause serious problems to many vision applications. We present in this paper an online statistical learning approach to model the background appearance variations under cast shadows. Based on the bi-illuminant (i.e. direct light sources and ambient illumination) dichromatic reflection model, we derive physics-based color features under the assumptions of constant ambient illumination and light sources with common spectral power distributions. We first use one Gaussian mixture model (GMM) to learn the color features, which are constant regardless of the background surfaces or illuminant colors in a scene. Then, we build up one pixel based GMM for each pixel to learn the local shadow features. To overcome the slow convergence rate in the conventional GMM learning, we update the pixel-based GMMs through confidence-rated learning. The proposed method can rapidly learn model parameters in an unsupervised way and adapt to illumination conditions or environment changes. Furthermore, we demonstrate that our method is robust to scenes with few foreground activities and videos captured at low or unsteady frame rates.
Jia-Bin Huang 0001, Chu-Song Chen
CVPR1
2009 A physical approach to Moving Cast Shadow Detection
abstract
This paper presents a physics-based approach capable of detecting cast shadows in video sequence effectively. We develop a new physical model of cast shadows without making prior assumption of the spectral power distribution (SPD) of the light sources and ambient illumination in the scene. The background appearance variation caused by cast shadows is characterized as the interaction of the blocked light sources and the background surface reflectance. We then take advantage of the statistical prevalence of cast shadows to learn and update the shadow model parameters using the Gaussian mixture model (GMM) over time. The proposed algorithm is completely unsupervised and can adapt to specific environment with complex illumination condition as well as changing shadow conditions. Experimental results on three challenging sequences demonstrate the effectiveness of the proposed method.
Jia-Bin Huang 0001, Chu-Song Chen
ICASSP1
2009 Image recolorization for the colorblind
abstract
In this paper, we propose a new re-coloring algorithm to enhance the accessibility for the color vision deficient (or colorblind). Compared to people with normal color vision, people with color vision deficiency (CVD) have difficulty in distinguishing between certain combinations of colors. This may hinder visual communication owing to the increasing use of colors in recent years. To address this problem, we re-color the image to preserve visual detail when perceived by people with CVD. We first extract the representing colors in an image. Then we find the optimal mapping to maintain the contrast between each pair of these representing colors. The proposed algorithm is image content dependent and completely automatic. Experimental results on natural images are illustrated to demonstrate the effectiveness of the proposed re-coloring algorithm.
Jia-Bin Huang 0001, Chu-Song Chen, Tzu-Cheng Jen, Sheng-Jyh Wang
ICASSP1
2007 Information Preserving Color Transformation for Protanopia and Deuteranopia
abstract
In this letter, we proposed a new recoloring method for people with protanopic and deuteranopic color deficiencies. We present a color transformation that aims to preserve the color information in the original images while maintaining the recolored images as natural as possible. Two error functions are introduced and combined together to form an objective function using the Lagrange multiplier with a user-specified parameter$\lambda$. This objective function is then minimized to obtain the optimal settings. Experimental results show that the proposed method can yield more comprehensible images for color-deficient viewers while maintaining the naturalness of the recolored images for standard viewers.
Jia-Bin Huang 0001, Yu-Cheng Tseng, Se-In Wu, Sheng-Jyh Wang
IEEE Signal Process. Lett.1