EDBT 2026 Demo / reviewers in the wild / expert
Tingbo Hou
dblp:35/3986
· DBLP profile ↗
30ranked-venue papers
11as first author
14since 2021 · last 2026
0009-0006-9667-9821ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 11 first-author · 10 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conversational Image Generation: Towards Multi-Round Personalized Generation with Multi-Modal Language ModelsabstractRecent advancements in diffusion models have significantly enhanced personalized image generation, enabling high-fidelity synthesis of human-subject-specific images. However, existing approaches are constrained by the inherent limitations of diffusion models, which lack conversational capabilities, and operate in a single-round setting, restricting user interaction. In this work, we propose a novel framework that integrates multi-modal large language models (MLLMs) for multi-round conversational personalization. To achieve this, we identified a performance bottleneck in the detokenizer of current MLLMs, which struggles to reconstruct fine-grained facial identity details. Thus, we enhance the detokenizer with a personalization-enhaced Diffusion Transformer (DiT). We also introduce a multi-stage instruction fine-tuning strategy to balance face preservation and prompt alignment effectively. To support multi-round generation, we implement a chat-history caching mechanism and construct the first multi-round personalization dataset from video clips. Experimental results demonstrate that our approach achieves state-of-the-art performance among MLLM-based personalization methods. To the best of our knowledge, this is the first work to enable conversational personalization, unlocking new capabilities for MLLMs in personalized image generation. Animesh Sinha, Felix Juefei-Xu, Xiaoliang Dai, Tingbo Hou, Peizhao Zhang, Zecheng He |
WACV | 8 |
| 2025 | LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational ComplexityabstractText-to-video generation enhances content creation but is highly computationally intensive: The computational cost of Diffusion Transformers (DiTs) scales quadratically in the number of pixels. This makes minute-length video generation extremely expensive, limiting most existing models to generating videos of only 10-20 seconds length. We propose a Linear-complexity text-to-video Generation (Lin-Gen) framework whose cost scales linearly in the number of pixels. For the first time, LinGen enables high-resolution minute-length video generation on a single GPU without compromising quality. It replaces the computationally-dominant and quadratic-complexity block, self-attention, with a linear-complexity block called MATE, which consists of an MA-branch and a TE-branch. The MA-branch targets short-to-long-range correlations, combining a bidirectional Mamba2 block with our token rearrangement method, Rotary Major Scan, and our review tokens developed for long video generation. The TE-branch is a novel TEmporal Swin Attention block that focuses on temporal correlations between adjacent tokens and medium-range tokens. The MATE block addresses the adjacency preservation issue of Mamba and improves the consistency of generated videos significantly. Experimental results show that LinGen outperforms DiT (with a 75.6% win rate) in video quality with up to 15× (11.5×) FLOPs (latency) reduction. Furthermore, both automatic metrics and human evaluation demonstrate that our LinGen-4B yields comparable video quality to state-of-the-art models (with a 50.5%, 52.1%, 49.1% win rate with respect to Gen-3, LumaLabs, and Kling, respectively). This paves the way for hour-length movie generation and real-time interactive video generation. Project website: https://lineargen.github.io/. Hongjie Wang 0002, Chih-Yao Ma, Yen-Cheng Liu, Ji Hou, Jialiang Wang 0001, Felix Juefei-Xu, Yaqiao Luo, Peizhao Zhang, Tingbo Hou, Peter Vajda, Niraj K. Jha, Xiaoliang Dai |
CVPR | 10 |
| 2025 | Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored PromptsabstractVideo personalization, which generates customized videos using reference images, has gained significant attention. However, prior methods typically focus on single-concept personalization, limiting broader applications that require multi-concept integration. Attempts to extend these models to multiple concepts often lead to identity blending, which results in composite characters with fused attributes from multiple sources. This challenge arises due to the lack of a mechanism to link each concept with its specific reference image. We address this with anchored prompts, which embed image anchors as unique tokens within text prompts, guiding accurate referencing during generation. Additionally, we introduce concept embeddings to encode the order of reference images. Our approach, Movie Weaver, seamlessly weaves multiple concepts—including face, body, and animal images—into one video, allowing flexible combinations in a single model. The evaluation shows that Movie Weaver outperforms existing methods for multi-concept video personalization in identity preservation and overall quality. Zecheng He, Tingbo Hou, Ji Hou, Xiaoliang Dai, Felix Juefei-Xu, Samaneh Azadi, Animesh Sinha, Peizhao Zhang, Peter Vajda, Diana Marculescu |
CVPR | 4 |
| 2025 | Learnings from Scaling Visual Tokenizers for Reconstruction and GenerationabstractVisual tokenization via auto-encoding empowers state-of-the-art image and video generative models by compressing pixels into a latent space. However, questions remain about how auto-encoder design impacts reconstruction and downstream generative performance. This work explores scaling in auto-encoders for reconstruction and generation by replacing the convolutional backbone with an enhanced Vision Transformer for Tokenization (ViTok). We find scaling the auto-encoder bottleneck correlates with reconstruction but exhibits a nuanced relationship with generation. Separately, encoder scaling yields no gains, while decoder scaling improves reconstruction with minimal impact on generation. As a result, we determine that scaling the current paradigm of auto-encoders is not effective for improving generation performance. Coupled with Diffusion Transformers, ViTok achieves competitive image reconstruction and generation performance on 256p and 512p ImageNet-1K. In videos, ViTok achieves SOTA reconstruction and generation performance on 16-frame 128p UCF-101. Philippe Hansen-Estruch, David Yan, Ching-Yao Chuang, Orr Zohar, Jialiang Wang 0001, Tingbo Hou, Sriram Vishwanath, Peter Vajda, Xinlei Chen |
ICML | 6 |
| 2025 | MoCha: Towards Movie-Grade Talking Character GenerationabstractRecent advancements in video generation have achieved impressive motion realism, yet they often overlook character-driven storytelling, a crucial task for automated film, animation generation.
We introduce Talking Characters, a more realistic task to generate talking character animations directly from speech and text. Unlike talking head tasks, Talking Characters aims at generating the full portrait of one or more characters beyond the facial region.
In this paper, we propose MoCha, the first of its kind to generate talking characters. To ensure precise synchronization between video and speech, we propose a localized audio attention mechanism that effectively aligns speech and video tokens.
To address the scarcity of large-scale speech-labelled video datasets, we introduce a joint training strategy that leverages both speech-labelled and text-labelled video data, significantly improving generalization across diverse character actions. We also design structured prompt templates with character tags, enabling, for the first time, multi-character conversation with turn-based dialogue—allowing AI-generated characters to engage in context-aware conversations with cinematic coherence.
Extensive qualitative and quantitative evaluations, including human evaluation studies and benchmark comparisons, demonstrate that MoCha sets a new standard for AI-generated cinematic storytelling, achieving superior realism, controllability and generalization. Cong Wei 0001, Ji Hou, Felix Juefei-Xu, Zecheng He, Xiaoliang Dai, Luxin Zhang, Tingbo Hou, Animesh Sinha, Peter Vajda, Wenhu Chen |
NeurIPS | 10 |
| 2024 | UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANsabstractText-to-image diffusion models have demonstrated re-markable capabilities in transforming text prompts into co-herent images, yet the computational cost of the multi-step inference remains a persistent challenge. To address this issue, we present UFOGen, a novel generative model de-signed for ultra-fast, one-step text-to-image generation. In contrast to conventional approaches that focus on improving samplers or employing distillation techniques for diffusion models, UFOGen adopts a hybrid methodology, inte-grating diffusion models with a GAN objective. Leveraging a newly introduced diffusion-GAN objective and initialization with pre-trained diffusion models, UFOGen excels in efficiently generating high-quality images conditioned on textual descriptions in a single step. Beyond traditional text-to-image generation, UFOGen showcases versatility in applications. Notably, UFOGen stands among the pioneering models enabling one-step text-to-image generation and diverse downstream tasks, presenting a significant advance-ment in the landscape of efficient generative models. Yanwu Xu 0003, Zhisheng Xiao, Tingbo Hou |
CVPR | 4 |
| 2024 | PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion ModelsabstractReward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using reinforcement learning (RL) to maximize rewards that reflect human preference. However, in the vision domain, existing RL-based reward finetuning methods are limited by their instability in large-scale training, rendering them incapable of generalizing to complex, unseen prompts. In this paper, we propose Proximal Reward Difference Prediction (PRDP), enabling stable black-box reward finetuning for diffusion models for the first time on large-scale prompt datasets with over 100K prompts. Our key innovation is the Reward Difference Prediction (RDP) objective that has the same optimal solution as the RL objective while enjoying better training stability. Specifically, the RDP objective is a supervised regression objective that tasks the diffusion model with predicting the reward difference of generated image pairs from their denoising trajectories. We theoretically prove that the diffusion model that obtains perfect reward difference prediction is exactly the maximizer of the RL objective. We further develop an online algorithm with proximal updates to stably optimize the RDP objective. In experiments, we demonstrate that PRDP can match the reward maximization ability of well-established RL-based methods in small-scale training. Furthermore, through large-scale training on text prompts from the Human Preference Dataset v2 and the Pick-a-Pic v1 dataset, PRDP achieves superior generation quality on a diverse set of complex, unseen prompts whereas RL-based methods completely fail. Fei Deng 0001, Qifei Wang, Tingbo Hou, Matthias Grundmann 0002 |
CVPR | 4 |
| 2024 | HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image ModelsabstractPersonalization has emerged as a prominent aspect within the field of generative AI, enabling the synthesis of individuals in diverse contexts and styles, while retaining high-fidelity to their identities. However, the process of personalization presents inherent challenges in terms of time and memory requirements. Fine-tuning each personalized model needs considerable GPU time investment, and storing a personalized model per subject can be demanding in terms of storage capacity. To overcome these challenges, we propose HyperDreamBooth—a hypernetwork capable of efficiently generating a small set of personalized weights from a single image of a person. By composing these weights into the diffusion model, coupled with fast finetuning, HyperDreamBooth can generate a person's face in various contexts and styles, with high subject details while also preserving the model's crucial knowledge of diverse styles and semantic modifications. Our method achieves personalization on faces in roughly 20 seconds, 25x faster than DreamBooth and 125x faster than Textual Inversion, using as few as one reference image, with the same quality and style diversity as DreamBooth. Also our method yields a model that is 10,000x smaller than a normal DreamBooth model. Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Tingbo Hou, Yael Pritch, Neal Wadhwa, Michael Rubinstein, Kfir Aberman |
CVPR | 5 |
| 2024 | 3D Congealing: 3D-Aware Image Alignment in the Wild
Zizhang Li, Amit Raj, Andreas Engelhardt, Yuanzhen Li, Tingbo Hou, Jiajun Wu 0001, Varun Jampani |
ECCV (1) | 6 |
| 2024 | MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices
Yanwu Xu 0003, Zhisheng Xiao, Haolin Jia 0001, Tingbo Hou |
ECCV (62) | 5 |
| 2024 | EM Distillation for One-step Diffusion ModelsabstractWhile diffusion models can learn complex distributions, sampling requires a computationally expensive iterative process. Existing distillation methods enable efficient sampling, but have notable limitations, such as performance degradation with very few sampling steps, reliance on training data access, or mode-seeking optimization that may fail to capture the full distribution. We propose EM Distillation (EMD), a maximum likelihood-based approach that distills a diffusion model to a one-step generator model with minimal loss of perceptual quality. Our approach is derived through the lens of Expectation-Maximization (EM), where the generator parameters are updated using samples from the joint distribution of the diffusion teacher prior and inferred generator latents. We develop a reparametrized sampling scheme and a noise cancellation technique that together stabilizes the distillation process. We further reveal an interesting connection of our method with existing methods that minimize mode-seeking KL. EMD outperforms existing one-step generative methods in terms of FID scores on ImageNet-64 and ImageNet-128, and compares favorably with prior work on distilling text-to-image diffusion models. Sirui Xie, Zhisheng Xiao, Diederik P. Kingma, Tingbo Hou, Ying Nian Wu, Kevin Murphy 0002, Tim Salimans, Ben Poole, Ruiqi Gao |
NeurIPS | 4 |
| 2023 | Multiscale Representation for Real-Time Anti-Aliasing Neural RenderingabstractThe rendering scheme in neural radiance field (NeRF) is effective in rendering a pixel by casting a ray into the scene. However, NeRF yields blurred rendering results when the training images are captured at non-uniform scales, and produces aliasing artifacts if the test images are taken in distant views. To address this issue, Mip-NeRF proposes a multiscale representation as a conical frustum to encode scale information. Nevertheless, this approach is only suitable for offline rendering since it relies on integrated positional encoding (IPE) to query a multilayer perceptron (MLP). To overcome this limitation, we propose mip voxel grids (Mip-VoG), an explicit multiscale representation with a deferred architecture for real-time anti-aliasing rendering. Our approach includes a density Mip-VoG for scene geometry and a feature Mip-VoG with a small MLP for view-dependent color. Mip-VoG represents scene scale using the level of detail (LOD) derived from ray differentials and uses quadrilinear interpolation to map a queried 3D location to its features and density from two neighboring down-sampled voxel grids. To our knowledge, our approach is the first to offer multiscale training and real-time anti-aliasing rendering simultaneously. We conducted experiments on multiscale dataset, results show that our approach outperforms state-of-the-art real-time rendering baselines. Dongting Hu, Zhenkai Zhang 0001, Tingbo Hou, Tongliang Liu, Huan Fu, Mingming Gong |
ICCV | 3 |
| 2023 | Towards Authentic Face Restoration with Iterative Diffusion Models and BeyondabstractAn authentic face restoration system is becoming increasingly demanding in many computer vision applications, e.g., image enhancement, video communication, and taking portrait. Most of the advanced face restoration models can recover high-quality faces from low-quality ones but usually fail to faithfully generate realistic and high-frequency details that are favored by users. To achieve authentic restoration, we propose IDM, an Iteratively learned face restoration system based on denoising Diffusion Models (DDMs). We define the criterion of an authentic face restoration system, and argue that denoising diffusion models are naturally endowed with this property from two aspects: intrinsic iterative refinement and extrinsic iterative enhancement. Intrinsic learning can preserve the content well and gradually refine the high-quality details, while extrinsic enhancement helps clean the data and improve the restoration task one step further. We demonstrate superior performance on blind face restoration tasks. Beyond restoration, we find the authentically cleaned data by the proposed restoration system is also helpful to image generation tasks in terms of training stabilization and sample quality. Without modifying the models, we achieve better quality than state-of-the-art on FFHQ and ImageNet generation using either GANs or diffusion models. Tingbo Hou, Yu-Chuan Su, Xuhui Jia, Yandong Li, Matthias Grundmann 0002 |
ICCV | 2 |
| 2023 | Semi-Implicit Denoising Diffusion Models (SIDDMs)abstractDespite the proliferation of generative models, achieving fast sampling during inference without compromising sample diversity and quality remains challenging. Existing models such as Denoising Diffusion Probabilistic Models (DDPM) deliver high-quality, diverse samples but are slowed by an inherently high number of iterative steps. The Denoising Diffusion Generative Adversarial Networks (DDGAN) attempted to circumvent this limitation by integrating a GAN model for larger jumps in the diffusion process. However, DDGAN encountered scalability limitations when applied to large datasets. To address these limitations, we introduce a novel approach that tackles the problem by matching implicit and explicit factors. More specifically, our approach involves utilizing an implicit model to match the marginal distributions of noisy data and the explicit conditional distribution of the forward diffusion. This combination allows us to effectively match the joint denoising distributions. Unlike DDPM but similar to DDGAN, we do not enforce a parametric distribution for the reverse step, enabling us to take large steps during inference. Similar to the DDPM but unlike DDGAN, we take advantage of the exact form of the diffusion process. We demonstrate that our proposed method obtains comparable generative performance to diffusion-based models and vastly superior results to models with a small number of sampling steps. Yanwu Xu 0003, Mingming Gong, Shaoan Xie, Matthias Grundmann 0002, Kayhan Batmanghelich, Tingbo Hou |
NeurIPS | 7 |
| 2013 | Hierarchical feature subspace for structure-preserving deformation
Shengfa Wang, Tingbo Hou, Shuai Li 0001, Zhixun Su, Hong Qin 0001 |
Comput. Aided Des. | 2 |
| 2013 | Anisotropic Elliptic PDEs for Feature ClassificationabstractThe extraction and classification of multitype (point, curve, patch) features on manifolds are extremely challenging, due to the lack of rigorous definition for diverse feature forms. This paper seeks a novel solution of multitype features in a mathematically rigorous way and proposes an efficient method for feature classification on manifolds. We tackle this challenge by exploring a quasi-harmonic field (QHF) generated by elliptic PDEs, which is the stable state of heat diffusion governed by anisotropic diffusion tensor. Diffusion tensor locally encodes shape geometry and controls velocity and direction of the diffusion process. The global QHF weaves points into smooth regions separated by ridges and has superior performance in combating noise/holes. Our method's originality is highlighted by the integration of locally defined diffusion tensor and globally defined elliptic PDEs in an anisotropic manner. At the computational front, the heat diffusion PDE becomes a linear system with Dirichlet condition at heat sources (called seeds). Our new algorithms afford automatic seed selection, enhanced by a fast update procedure in a high-dimensional space. By employing diffusion probability, our method can handle both manufactured parts and organic objects. Various experiments demonstrate the flexibility and high performance of our method. Tingbo Hou, Shuai Li 0001, Zhixun Su, Hong Qin 0001, Shengfa Wang |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2013 | Admissible Diffusion Wavelets and Their Applications in Space-Frequency ProcessingabstractAs signal processing tools, diffusion wavelets and biorthogonal diffusion wavelets have been propelled by recent research in mathematics. They employ diffusion as a smoothing and scaling process to empower multiscale analysis. However, their applications in graphics and visualization are overshadowed by nonadmissible wavelets and their expensive computation. In this paper, our motivation is to broaden the application scope to space-frequency processing of shape geometry and scalar fields. We propose the admissible diffusion wavelets (ADW) on meshed surfaces and point clouds. The ADW are constructed in a bottom-up manner that starts from a local operator in a high frequency, and dilates by its dyadic powers to low frequencies. By relieving the orthogonality and enforcing normalization, the wavelets are locally supported and admissible, hence facilitating data analysis and geometry processing. We define the novel rapid reconstruction, which recovers the signal from multiple bands of high frequencies and a low-frequency base in full resolution. It enables operations localized in both space and frequency by manipulating wavelet coefficients through space-frequency filters. This paper aims to build a common theoretic foundation for a host of applications, including saliency visualization, multiscale feature extraction, spectral geometry processing, etc. Tingbo Hou, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2012 | A Novel Material-Aware Feature Descriptor for Volumetric Image Registration in Diffusion Tensor Space
Shuai Li 0001, Qinping Zhao, Shengfa Wang, Tingbo Hou, Aimin Hao, Hong Qin 0001 |
ECCV (4) | 4 |
| 2012 | Bag-of-feature-graphs: A new paradigm for non-rigid shape retrieval
Tingbo Hou, Xiaohua Hou, Ming Zhong 0007, Hong Qin 0001 |
ICPR | 1 |
| 2012 | Diffusion-driven high-order matching of partial deformable shapes
Tingbo Hou, Ming Zhong 0007, Hong Qin 0001 |
ICPR | 1 |
| 2012 | A hierarchical approach to high-quality partial shape registration
Ming Zhong 0007, Tingbo Hou, Hong Qin 0001 |
ICPR | 2 |
| 2012 | Continuous and discrete Mexican hat wavelet transforms on manifolds
Tingbo Hou, Hong Qin 0001 |
Graph. Model. | 1 |
| 2012 | Corrigendum to "Continuous and discrete Mexican hat wavelet transforms on manifolds" [Graphical Models 74 (2012) 221-232]
Tingbo Hou, Hong Qin 0001 |
Graph. Model. | 1 |
| 2012 | High-quality image deblurring with panchromatic pixelsabstractImage deblurring has been a very challenging problem in recent decades. In this article, we propose a high-quality image deblurring method with a novel image prior based on a new imaging system. The imaging system has a newly designed sensor pattern achieved by adding panchromatic (pan) pixels to the conventional Bayer pattern. Since these pan pixels are sensitive to all wavelengths of visible light, they collect a significantly higher proportion of the light striking the sensor. A new demosaicing algorithm is also proposed to restore full-resolution images from pixels on the sensor. The shutter speed of pan pixels is controllable to users. Therefore, we can produce multiple images with different exposures. When long exposure is needed under dim light, we read pan pixels twice in one shot: one with short exposure and the other with long exposure. The long-exposure image is often blurred, while the short-exposure image can be sharp and noisy. The short-exposure image plays an important role in deblurring, since it is sharp and there is no alignment problem for the one-shot image pair. For the algorithmic aspect, our method runs in a two-step maximum-a-posteriori (MAP) fashion under a joint minimization of the blur kernel and the deblurred image. The algorithm exploits a combined image prior with a statistical part and a spatial part, which is powerful in ringing controls. Extensive experiments under various conditions and settings are conducted to demonstrate the performance of our method. Tingbo Hou, John Border, Hong Qin 0001, Rodney L. Miller |
ACM Trans. Graph. | 2 |
| 2012 | Robust Dense Registration of Partial Nonrigid ShapesabstractThis paper presents a complete and robust solution for dense registration of partial nonrigid shapes. Its novel contributions are founded upon the newly proposed heat kernel coordinates (HKCs) that can accurately position points on the shape, and the priority-vicinity search that ensures geometric compatibility during the registration. HKCs index points by computing heat kernels from multiple sources, and their magnitudes serve as priorities of queuing points in registration. We start with shape features as the sources of heat kernels via feature detection and matching. Following the priority order of HKCs, the dense registration is progressively propagated from feature sources to all points. Our method has a superior indexing ability that can produce dense correspondences with fewer flips. The diffusion nature of HKCs, which can be interpreted as a random walk on a manifold, makes our method robust to noise and small holes avoiding surface surgery and repair. Our method searches correspondence only in a small vicinity of registered points, which significantly improves the time performance. Through comprehensive experiments, our new method has demonstrated its technical soundness and robustness by generating highly compatible dense correspondences. Tingbo Hou, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2011 | Image Deconvolution With Multi-Stage Convex Relaxation and Its Perceptual EvaluationabstractThis paper proposes a new image deconvolution method using multi-stage convex relaxation, and presents a metric for perceptual evaluation of deconvolution results. Recent work in image deconvolution addresses the deconvolution problem via minimization with non-convex regularization. Since all regularization terms in the objective function are non-convex, this problem can be well modeled and solved by multi-stage convex relaxation. This method, adopted from machine learning, iteratively refines the convex relaxation formulation using concave duality. The newly proposed deconvolution method has outstanding performance in noise removal and artifact control. A new metric, transduced contrast-to-distortion ratio (TCDR), is proposed based on a human vision system (HVS) model that simulates human responses to visual contrasts. It is sensitive to ringing and boundary artifacts, and very efficient to compute. We conduct comprehensive perceptual evaluation of image deconvolution using visual signal-to-noise ratio (VSNR) and TCDR. Experimental results of both synthetic and real data demonstrate that our method indeed improves the visual quality of deconvolution results with low distortions and artifacts. Tingbo Hou, Hong Qin 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Multi-scale anisotropic heat diffusion based on normal-driven shape representation
Shengfa Wang, Tingbo Hou, Zhixun Su, Hong Qin 0001 |
Vis. Comput. | 2 |
| 2010 | Efficient Computation of Scale-Space Features for Deformable Shape Correspondences
Tingbo Hou, Hong Qin 0001 |
ECCV (3) | 1 |
| 2010 | Illumination learning from a single image with unknown shape and textureabstractIn this paper, we develop a method for learning illumination from a single image, which can benefit illumination-invariant algorithms in computer vision and image-based rendering in graphics. Illumination learning has been widely studied, yet still has some shortcomings such as the restriction of Lambertian surfaces and the prerequisite of known shape or texture. Our method can adaptively learn illumination from images of vehicles with unknown shape and texture. We formulate the illumination model with both diffusion and specularity components using a frequency-space representation, and adopt an iterative strategy to estimate lighting, shape, and texture under a joint energy function. Using our method, we can perform de-lighting and re-lighting on input images, and render other 3D models with learned illumination. Experimental results show that our method can work in a wide range of real-world environments with both indoor and outdoor illumination conditions. Tingbo Hou, Hong Qin 0001 |
ICIP | 1 |
| 2010 | Image deconvolution using multigrid natural image prior and its applicationsabstractThe natural image prior has been proven to be a powerful tool for image deblurring in recent years, though its performance against noise in various applications has not been thoroughly studied. In this paper, we present a multigrid natural image prior for image deconvolution that enhances its robustness against noise, and afford three applications of image deconvolution using this prior: deblurring, super-resolution, and denoising. The prior is based on a remarkable property of natural images that derivatives with different resolutions are subject to the same heavy-tailed distribution with a spatial factor. It can serve in both blind and non-blind deconvolutions. The performances of the proposed prior in different applications are demonstrated by corresponding experimental results. Tingbo Hou, Hong Qin 0001, Rodney L. Miller |
ICIP | 1 |