VLDB 2026 Research / reviewers in the wild / expert
Michael S. Brown
dblp:02/4733
· DBLP profile ↗
144ranked-venue papers
10as first author
35since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 122 · 8 first-author · 33 since 2021Artificial intelligence and machine learning · 95 · 6 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Neural Network Encoding for 3D Color Lookup Tablesabstract3D color lookup tables (LUTs) enable precise color manipulation by mapping input RGB values to specific output RGB values. 3D LUTs are instrumental in various applications, including video editing, in-camera processing, photographic filters, computer graphics, and color processing for displays. While an individual LUT does not incur a high memory overhead, software and devices may need to store dozens to hundreds of LUTs that can take over 100 MB. This work aims to develop a neural network architecture that can encode hundreds of LUTs in a single compact representation. To this end, we propose a model with a memory footprint of less than 0.25 MB that can reconstruct 512 LUTs with only minor color distortion (ΔE ≤ 2.0 on average) over the entire color gamut. We also show that our network can weight colors to provide further quality gains on natural image colors (ΔE ≤ 1.0 on average). Finally, we show that minor modifications to the network architecture enable a bijective encoding that produces LUTs that are invertible, allowing for reverse color processing. Vahid Zehtab, David B. Lindell, Marcus A. Brubaker, Michael S. Brown |
AAAI | 4 |
| 2025 | Examining Joint Demosaicing and Denoising for Single-, Quad-, and Nona-Bayer PatternsabstractCamera sensors have color filters arranged in a mosaic layout, traditionally following the Bayer pattern. Demosaicing is a critical step camera hardware applies to obtain a full-channel RGB image. Many smartphones now have multiple sensors with different patterns, such as Quad-Bayer or Nona-Bayer. Most modern deep network-based models perform joint demosaicing and denoising with the strategy of training a separate network per pattern. Relying on individual models per pattern requires additional memory overhead and makes it challenging to switch quickly between cameras. In this work, we are interested in analyzing strategies for joint demosaicing and denoising for the three main mosaic layouts ($1 \times 1$ Single-Bayer, $2 \times 2$ Quad-Bayer, and $3 \times 3$ Nona-Bayer). We found concatenating a three-channel mosaic embedding to the input image and training a unified demosaicing architecture yields results that outperform existing Quad-Bayer and Nona-Bayer models and are comparable to Single-Bayer models. Additionally, we describe a maskout strategy that enhances the model performance and facilitates dead pixel correction-a step often overlooked by existing Al-based demosaicing models. As part of this effort, we captured a new demosaicing dataset of 638 RAW images that contain challenging scenes with patches annotated for training, validation, and testing. Code and data is available at https://github.com/SamsungLabs/unified-demosaicing. SaiKiran Kumar Tedla, Abhijith Punnappurath, Luxi Zhao 0002, Michael S. Brown |
ICCP | 4 |
| 2025 | Time-Aware Auto White Balance in Mobile Photography
Mahmoud Afifi, Luxi Zhao 0002, Abhijith Punnappurath, Mohammed A. Abdelsalam, Michael S. Brown |
ICCV | 6 |
| 2025 | Gain-MLP: Improving HDR Gain Map Encoding via a Lightweight MLP
Trevor D. Canham, SaiKiran Kumar Tedla, Michael Murdoch, Michael S. Brown |
ICCV | 4 |
| 2025 | CCMNet: Leveraging Calibrated Color Correction Matrices for Cross-Camera Color ConstancyabstractComputational color constancy, or white balancing, is a key module in a camera's image signal processor (ISP) that corrects color casts from scene lighting. Because this operation occurs in the camera-specific raw color space, white balance algorithms must adapt to different cameras. This paper introduces a learning-based method for cross-camera color constancy that generalizes to new cameras without retraining. Our method leverages pre-calibrated color correction matrices (CCMs) available on ISPs that map the camera's raw color space to a standard space (e.g., CIE XYZ). Our method uses these CCMs to transform predefined illumination colors (i.e., along the Planckian locus) into the test camera's raw space. The mapped illuminants are encoded into a compact camera fingerprint embedding (CFE) that enables the network to adapt to unseen cameras. To prevent overfitting due to limited cameras and CCMs during training, we introduce a data augmentation technique that interpolates between cameras and their CCMs. Experimental results across multiple datasets and backbones show that our method achieves state-of-the-art cross-camera color constancy while remaining lightweight and relying only on data readily available in camera ISPs. Dongyoung Kim, Mahmoud Afifi, Dongyun Kim, Michael S. Brown, Seon Joo Kim |
ICCV | 4 |
| 2025 | Spectral Sensitivity Estimation with an Uncalibrated Diffraction GratingabstractThis paper introduces a practical and accurate calibration method for camera spectral sensitivity using a diffraction grating. Accurate calibration of camera spectral sensitivity is crucial for various computer vision tasks, including color correction, illumination estimation, and material analysis. Unlike existing approaches that require specialized narrow-band filters or reference targets with known spectral reflectances, our method only requires an uncalibrated diffraction grating sheet, readily available off-the-shelf. By capturing images of the direct illumination and its diffracted pattern through the grating sheet, our method estimates both the camera spectral sensitivity and the diffraction grating parameters in a closed-form manner. Experiments on synthetic and real-world data demonstrate that our method outperforms conventional reference target-based methods, underscoring its effectiveness and practicality. Lilika Makabe, Hiroaki Santo, Fumio Okura, Michael S. Brown, Yasuyuki Matsushita |
ICCV | 4 |
| 2025 | Revisiting Image Fusion for Multi-Illuminant White-Balance CorrectionabstractWhite balance (WB) correction in scenes with multiple illuminants remains a persistent challenge in computer vision. Recent methods explored fusion-based approaches, where a neural network linearly blends multiple sRGB versions of an input image, each processed with predefined WB presets. However, we demonstrate that these methods are suboptimal for common multi-illuminant scenarios. Additionally, existing fusion-based methods rely on sRGB WB datasets lacking dedicated multi-illuminant images, limiting both training and evaluation. To address these challenges, we introduce two key contributions. First, we propose an efficient transformer-based model that effectively captures spatial dependencies across sRGB WB presets, substantially improving upon linear fusion techniques. Second, we introduce a large-scale multi-illuminant dataset comprising over 16,000 sRGB images rendered with five different WB settings, along with WB-corrected images. Our method achieves up to 100\% improvement over existing techniques on our new multi-illuminant image fusion dataset. David Serrano-Lozano, Aditya Arora, Luis Herranz, Konstantinos G. Derpanis, Michael S. Brown, Javier Vazquez-Corral |
ICCV | 5 |
| 2025 | Multispectral Demosaicing via Dual CamerasabstractMultispectral (MS) images capture detailed scene information across a wide range of spectral bands, making them invaluable for applications requiring rich spectral data. Integrating MS imaging into multi camera devices, such as smartphones, has the potential to enhance both spectral applications and RGB image quality. A critical step in processing MS data is demosaicing, which reconstructs color information from the mosaic MS images captured by the camera. This paper proposes a method for MS image demosaicing specifically designed for dual-camera setups where both RGB and MS cameras capture the same scene. Our approach leverages co-captured RGB images, which typically have higher spatial fidelity, to guide the demosaicing of lower-fidelity MS images. We introduce the Dual-camera RGB-MS Dataset - a large collection of paired RGB and MS mosaiced images with ground-truth demosaiced outputs - that enables training and evaluation of our method. Experimental results demonstrate that our method achieves state-of-the-art accuracy compared to existing techniques. SaiKiran Kumar Tedla, Junyong Lee 0001, Beixuan Yang, Mahmoud Afifi, Michael S. Brown |
ICCV | 5 |
| 2025 | Generating the Past, Present and Future from a Motion-Blurred ImageabstractWe seek to answer the question: what can a motion-blurred image reveal about a scene's past, present, and future? Although motion blur obscures image details and degrades visual quality, it also encodes information about scene and camera motion during an exposure. Previous techniques leverage this information to estimate a sharp image from an input blurry one, or to predict a sequence of video frames showing what might have occurred at the moment of image capture. However, they rely on handcrafted priors or network architectures to resolve ambiguities in this inverse problem, and do not incorporate image and video priors on large-scale datasets. As such, existing methods struggle to reproduce complex scene dynamics and do not attempt to recover what occurred before or after an image was taken. Here, we introduce a new technique that repurposes a pre-trained video diffusion model trained on internet-scale datasets to recover videos revealing complex scene dynamics during the moment of capture and what might have occurred immediately into the past or future. Our approach is robust and versatile; it outperforms previous methods for this task, generalizes to challenging in-the-wild images, and supports downstream tasks such as recovering camera trajectories, object motion, and dynamic 3D scene structure. Code and data are available at blur2vid.github.io SaiKiran Kumar Tedla, Kelly Zhu, Trevor D. Canham, Felix Taubner, Michael S. Brown, Kiriakos N. Kutulakos, David B. Lindell |
ACM Trans. Graph. | 5 |
| 2024 | NILUT: Conditional Neural Implicit 3D Lookup Tables for Image Enhancementabstract3D lookup tables (3D LUTs) are a key component for image enhancement. Modern image signal processors (ISPs) have dedicated support for these as part of the camera rendering pipeline. Cameras typically provide multiple options for picture styles, where each style is usually obtained by applying a unique handcrafted 3D LUT. Current approaches for learning and applying 3D LUTs are notably fast, yet not so memory-efficient, as storing multiple 3D LUTs is required. For this reason and other implementation limitations, their use on mobile devices is less popular. In this work, we propose a Neural Implicit LUT (NILUT), an implicitly defined continuous 3D color transformation parameterized by a neural network. We show that NILUTs are capable of accurately emulating real 3D LUTs. Moreover, a NILUT can be extended to incorporate multiple styles into a single network with the ability to blend styles implicitly. Our novel approach is memory-efficient, controllable and can complement previous methods, including learned ISPs. Code at https://github.com/mv-lab/nilut Marcos V. Conde, Javier Vazquez-Corral, Michael S. Brown, Radu Timofte |
AAAI | 3 |
| 2024 | Non-parametric Sensor Noise Modeling and Synthesis
Luxi Zhao 0002, Atin Singh, Jaeduk Han, Abhijith Punnappurath, Marcus A. Brubaker, Jihwan Choe, Michael S. Brown |
ECCV (24) | 8 |
| 2024 | NamedCurves: Learned Image Enhancement via Color Naming
David Serrano-Lozano, Luis Herranz, Michael S. Brown, Javier Vazquez-Corral |
ECCV (71) | 3 |
| 2024 | LookToFocus: Image Focus via Eye TrackingabstractWe present LookToFocus, a method to perform real-time manual camera focus based on eye tracking. LookToFocus and two alternative methods for manual focus1 photography tasks were compared in a user study. A novel manual focus camera simulation was used to test the methods. The first two methods were LookToFocus and LookToFocusNB (no bounding box). The third method, TapToFocus, used touch for manual focus and image capture, analogous to typical smartphone interaction. LookToFocusNB had the fastest mean capture time at 1429 ms; the mean capture times were 1431 ms for LookToFocusNB and 2416 ms for TapToFocus. When compared to TapToFocus, LookToFocus and LookToFocusNB had faster capture times because these algorithms start converging to the optimal focus as soon as the user spots the target. LookToFocus and LookToFocusNB also had no significant decrease in sharpness error compared to TapToFocus. Users preferred LookToFocus over both LookToFocusNB and TapToFocus. SaiKiran Kumar Tedla, I. Scott MacKenzie, Michael S. Brown |
ETRA | 3 |
| 2024 | Mixed Graph Signal Analysis of Joint Image Denoising / InterpolationabstractA noise-corrupted image often requires interpolation. Given a linear denoiser and a linear interpolator, when should the operations be independently executed in separate steps, and when should they be combined and jointly optimized? We study joint denoising / interpolation of images from a mixed graph filtering perspective: we model denoising using an undirected graph, and interpolation using a directed graph. We first prove that, under mild conditions, a linear denoiser is a solution graph filter to a maximum a posteriori (MAP) problem using an undirected graph smoothness prior, while a linear interpolator is a solution to a MAP problem using a directed graph smoothness prior. Next, we study two variants of the joint interpolation / denoising problem: a graph-based denoiser followed by an interpolator has an optimal separable solution, while an interpolator followed by a denoiser has an optimal non-separable solution. Experiments show that our joint denoising / interpolation method outperformed separate approaches noticeably. Niruhan Viswarupan, Gene Cheung, Fengbo Lan, Michael S. Brown |
ICASSP | 4 |
| 2024 | Palette-Based Color Harmonization via Color NamingabstractColor harmony refers to combinations of colors that look pleasing together. We present a novel strategy to harmonize an image's colors using color-palette manipulation and color naming. Palette-based color manipulation is a method that extracts a few colors to represent the image. Modifying the palette colors modifies the color appearance of the image. A color-naming model is a mechanism to categorize colors into a fixed number of basic color terms. Working from a color-naming model, we derive a set ofprototype colorsand demonstrate that mapping an image's extracted color palette to the nearest prototype colors effectively harmonizes the image's colors. This straightforward approach yields visually compelling, outperforming more complex color harmony methods. Danna Xue, Javier Vazquez-Corral, Luis Herranz, Yanning Zhang 0001, Michael S. Brown |
IEEE Signal Process. Lett. | 5 |
| 2023 | GamutMLP: A Lightweight MLP for Color Loss RecoveryabstractCameras and image-editing software often process images in the wide-gamut ProPhoto color space, encompassing 90% of all visible colors. However, when images are encoded for sharing, this color-rich representation is transformed and clipped to fit within the small-gamut standard RGB (sRGB) color space, representing only 30% of visible colors. Recovering the lost color information is challenging due to the clipping procedure. Inspired by neural implicit representations for 2D images, we propose a method that optimizes a lightweight multi-layer-perceptron (MLP) model during the gamut reduction step to predict the clipped values. GamutMLP takes approximately 2 seconds to optimize and requires only 23 KB of storage. The small memory footprint allows our GamutMLP model to be saved as metadata in the sRGB image—the model can be extracted when needed to restore wide-gamut color values. We demonstrate the effectiveness of our approach for color recovery and compare it with alternative strategies, including pre-trained DNN-based gamut expansion networks and other implicit neural representation methods. As part of this effort, we introduce a new color gamut dataset of 2200 wide-gamut/small-gamut images for training and testing. Hoang Minh Le 0001, Brian L. Price, Scott Cohen, Michael S. Brown |
CVPR | 4 |
| 2023 | Physically-plausible illumination distribution estimationabstractA camera’s auto-white-balance (AWB) module operates under the assumption that there is a single dominant illumination in a captured scene. AWB methods estimate an image’s dominant illumination and use it as the target "white point" for correction. However, in natural scenes, there are often many light sources present. We performed a user study that revealed that non-dominant illuminations often produce visually pleasing white-balanced images and, in some cases, are even preferred over the dominant illumination. Motivated by this observation, we revisit AWB to predict a distribution of plausible illuminations for use in white balance. As part of this effort, we extend the Cube+ + illumination estimation dataset [12] to provide ground truth illumination distributions per image. Using this new ground truth data, we describe how to train a lightweight neural network method to predict the scene’s illumination distribution. We describe how our idea can be used with existing image formats by embedding the estimated distribution in the RAW image to enable users to generate visually plausible white-balance images. Egor I. Ershov, Vasily Tesalin, Ivan Ermakov, Michael S. Brown |
ICCV | 4 |
| 2023 | Graphics2RAW: Mapping Computer Graphics Images to Sensor RAW ImagesabstractComputer graphics (CG) rendering platforms produce imagery with ever-increasing photo realism. The narrowing domain gap between real and synthetic imagery makes it possible to use CG images as training data for deep learning models targeting high-level computer vision tasks, such as autonomous driving and semantic segmentation. CG images, however, are currently not suitable for low-level vision tasks targeting RAW sensor images. This is because RAW images are encoded in sensor-specific color spaces and incur pre-white-balance color casts caused by the sensor’s response to scene illumination. CG images are rendered directly to a device-independent perceptual color space without needing white balancing. As a result, it is necessary to apply a mapping procedure to close the domain gap between graphics and RAW images. To this end, we introduce a framework to process graphics images to mimic RAW sensor images accurately. Our approach allows a one-to-many mapping, where a single graphics image can be transformed to match multiple sensors and multiple scene illuminations. In addition, our approach requires only a handful of example RAW-DNG files from the target sensor as parameters for the mapping process. We compare our method to alternative strategies and show that our approach produces more realistic RAW images and provides better results on three low-level vision tasks: RAW denoising, illumination estimation, and neural rendering for night photography. Finally, as part of this work, we provide a dataset of 292 realistic CG images for training low-light imaging models. Donghwan Seo, Abhijith Punnappurath, Luxi Zhao 0002, Abdelrahman Abdelhamed, SaiKiran Kumar Tedla, Jihwan Choe, Michael S. Brown |
ICCV | 8 |
| 2023 | Examining Autoexposure for Challenging ScenesabstractAutoexposure (AE) is a critical step applied by camera systems to ensure properly exposed images. While current AE algorithms are effective in well-lit environments with constant illumination, these algorithms still struggle in environments with bright light sources or scenes with abrupt changes in lighting. A significant hurdle in developing new AE algorithms for challenging environments, especially those with time-varying lighting, is the lack of suitable image datasets. To address this issue, we have captured a new 4D exposure dataset that provides a large solution space (i.e., shutter speed range from $\frac{1}{{1500}}$ to 15 seconds) over a temporal sequence with moving objects, bright lights, and varying lighting. In addition, we have designed a software platform to allow AE algorithms to be used in a plug-and-play manner with the dataset. Our dataset and associate platform enable repeatable evaluation of different AE algorithms and provide a much-needed starting point to develop better AE methods. We examine several existing AE strategies using our dataset and show that most users prefer a simple saliency method for challenging lighting conditions. SaiKiran Kumar Tedla, Beixuan Yang, Michael S. Brown |
ICCV | 3 |
| 2023 | Integrating High-Level Features for Consistent Palette-based Multi-image RecoloringabstractAbstract Achieving visually consistent colors across multiple images is important when images are used in photo albums, websites, and brochures. Unfortunately, only a handful of methods address multi‐image color consistency compared to one‐to‐one color transfer techniques. Furthermore, existing methods do not incorporate high‐level features that can assist graphic designers in their work. To address these limitations, we introduce a framework that builds upon a previous palette‐based color consistency method and incorporates three high‐level features: white balance, saliency, and color naming. We show how these features overcome the limitations of the prior multi‐consistency workflow and showcase the user‐friendly nature of our framework. D. Xue, Javier Vazquez-Corral, Luis Herranz, Michael S. Brown |
Comput. Graph. Forum | 5 |
| 2022 | Modeling sRGB Camera Noise with Normalizing FlowsabstractNoise modeling and reduction are fundamental tasks in low-level computer vision. They are particularly important for smartphone cameras relying on small sensors that exhibit visually noticeable noise. There has recently been renewed interest in using data-driven approaches to improve camera noise models via neural networks. These data-driven approaches target noise present in the raw-sensor image before it has been processed by the camera's image signal processor (ISP). Modeling noise in the RAW-rgb domain is useful for improving and testing the in-camera denoising algorithm; however, there are situations where the camera's ISP does not apply denoising or additional denoising is desired when the RAW-rgb domain image is no longer available. In such cases, the sensor noise propagates through the ISP to the final rendered image encoded in standard RGB (sRGB). The nonlinear steps on the ISP culminate in a significantly more complex noise distribution in the sRGB domain and existing raw-domain noise models are unable to capture the sRGB noise distribution. We propose a new sRGB-domain noise model based on normalizing flows that is capable of learning the complex noise distribution found in sRGB images under various ISO levels. Our normalizing flows-based approach outperforms other models by a large margin in noise modeling and synthesis tasks. We also show that image denoisers trained on noisy images synthesized with our noise model outperforms those trained with noise from baselines models. Shayan Kousha, Ali Maleky, Michael S. Brown, Marcus A. Brubaker |
CVPR | 3 |
| 2022 | Noise2NoiseFlow: Realistic Camera Noise Modeling without Clean ImagesabstractImage noise modeling is a long-standing problem with many applications in computer vision. Early attempts that propose simple models, such as signal-independent additive white Gaussian noise or the heteroscedastic Gaussian noise model (a.k.a., camera noise level function) are not sufficient to learn the complex behavior of the camera sensor noise. Recently, more complex learning-based models have been proposed that yield better results in noise synthesis and downstream tasks, such as denoising. However, their dependence on supervised data (i.e., paired clean images) is a limiting factor given the challenges in producing ground-truth images. This paper proposes a framework for training a noise model and a denoiser simultaneously while relying only on pairs of noisy images rather than noisy/clean paired image data. We apply this framework to the training of the Noise Flow architecture. The noise synthesis and density estimation results show that our framework outperforms previous signal-processing-based noise models and is on par with its supervised counterpart. The trained denoiser is also shown to significantly improve upon both supervised and weakly supervised baseline denoising approaches. The results indicate that the joint training of a denoiser and a noise model yields significant improvements in the denoiser. Ali Maleky, Shayan Kousha, Michael S. Brown, Marcus A. Brubaker |
CVPR | 3 |
| 2022 | Learning sRGB-to-Raw-RGB De-rendering with Content-Aware MetadataabstractMost camera images are rendered and saved in the standard RGB (sRGB) format by the camera's hardware. Due to the in-camera photo-finishing routines, nonlinear sRGB images are undesirable for computer vision tasks that assume a direct relationship between pixel values and scene radiance. For such applications, linear raw-RGB sensor images are preferred. Saving images in their raw-RGB format is still uncommon due to the large storage requirement and lack of support by many imaging applications. Several “raw reconstruction” methods have been proposed that utilize specialized metadata sampled from the raw-RGB image at capture time and embedded in the sRGB image. This metadata is used to parameterize a mapping function to derender the sRGB image back to its original raw-RGB format when needed. Existing raw reconstruction methods rely on simple sampling strategies and global mapping to perform the de-rendering. This paper shows how to improve the derendering results by jointly learning sampling and reconstruction. Our experiments show that our learned sampling can adapt to the image content to produce better raw reconstructions than existing methods. We also describe an online fine-tuning strategy for the reconstruction network to improve results further. Seonghyeon Nam, Abhijith Punnappurath, Marcus A. Brubaker, Michael S. Brown |
CVPR | 4 |
| 2022 | Day-to-Night Image Synthesis for Training Nighttime Neural ISPsabstractMany flagship smartphone cameras now use a dedicated neural image signal processor (ISP) to render noisy raw sensor images to the final processed output. Training night-mode ISP networks relies on large-scale datasets of image pairs with: (1) a noisy raw image captured with a short exposure and a high ISO gain; and (2) a ground truth low-noise raw image captured with a long exposure and low ISO that has been rendered through the ISP. Capturing such image pairs is tedious and time-consuming, requiring careful setup to ensure alignment between the image pairs. In addition, ground truth images are often prone to motion blur due to the long exposure. To address this problem, we propose a method that synthesizes nighttime images from day-time images. Daytime images are easy to capture, exhibit low-noise (even on smartphone cameras) and rarely suffer from motion blur. We outline a processing framework to convert daytime raw images to have the appearance of realistic nighttime raw images with different levels of noise. Our procedure allows us to easily produce aligned noisy and clean nighttime image pairs. We show the effectiveness of our synthesis framework by training neural ISPs for nightmode rendering. Furthermore, we demonstrate that using our synthetic nighttime images together with small amounts of real data (e.g., 5% to 10%) yields performance almost on par with training exclusively on real nighttime images. Our dataset and code are available at https://github.com/SamsungLabs/day-to-night. Abhijith Punnappurath, Abdullah Abuolaim, Abdelrahman Abdelhamed, Alex Levinshtein, Michael S. Brown |
CVPR | 5 |
| 2022 | Neural Image Representations for Multi-image Fusion and Layer Separation
Seonghyeon Nam, Marcus A. Brubaker, Michael S. Brown |
ECCV (7) | 3 |
| 2022 | Extracting Vignetting and Grain Filter Effects from PhotosabstractMost smartphones support the use of real-time camera filters to impart visual effects to captured images. Currently, such filters come preinstalled on-device or need to be downloaded and installed before use (e.g., Instagram filters). Recent work [24] proposed a method to extract a camera filter directly from an example photo that has already had a filter applied. The work in [24] focused only on the color and tonal aspects of the underlying filter. In this paper, we introduce a method to extract two spatially varying effects commonly used by on-device camera filters—namely, image vignetting and image grain. Specifically, we show how to extract the parameters for vignetting and image grain present in an example image and replicate these effects as an on-device filter. We use lightweight CNNs to estimate the filter parameters and employ efficient techniques—isotropic Gaussian filters and simplex noise—for regenerating the filters. Our design achieves a reasonable trade-off between efficiency and realism. We show that our method can extract vignetting and image grain filters from stylized photos and replicate the filters on captured images more faithfully, as compared to color and style transfer methods. Our method is significantly efficient and has been already deployed to millions of flagship smartphones. Abdelrahman Abdelhamed, Jonghwa Yim, Abhijith Punnappurath, Michael S. Brown, Jihwan Choe |
WACV | 4 |
| 2022 | Improving Single-Image Defocus Deblurring: How Dual-Pixel Images Help Through Multi-Task LearningabstractMany camera sensors use a dual-pixel (DP) design that operates as a rudimentary light field providing two subaperture views of a scene in a single capture. The DP sensor was developed to improve how cameras perform autofocus. Since the DP sensor’s introduction, researchers have found additional uses for the DP data, such as depth estimation, reflection removal, and defocus deblurring. We are interested in the latter task of defocus deblurring. In particular, we propose a single-image deblurring network that incorporates the two sub-aperture views into a multitask framework. Specifically, we show that jointly learning to predict the two DP views from a single blurry input image improves the network’s ability to learn to deblur the image. Our experiments show this multi-task strategy achieves +1dB PSNR improvement over state-of-the-art defocus deblurring methods. In addition, our multi-task framework allows accurate DP-view synthesis (e.g., ∼39dB PSNR) from the single input image. These high-quality DP views can be used for other DP-based applications, such as reflection removal. As part of this effort, we have captured a new dataset of 7,059 high-quality images to support our training for the DP-view synthesis task. Abdullah Abuolaim, Mahmoud Afifi, Michael S. Brown |
WACV | 3 |
| 2022 | Auto White-Balance Correction for Mixed-Illuminant ScenesabstractAuto white balance (AWB) is applied by camera hardware at capture time to remove the color cast caused by the scene illumination. The vast majority of white-balance algorithms assume a single light source illuminates the scene; however, real scenes often have mixed lighting conditions. This paper presents an effective AWB method to deal with such mixed-illuminant scenes. A unique departure from conventional AWB, our method does not require illuminant estimation, as is the case in traditional camera AWB modules. Instead, our method proposes to render the captured scene with a small set of predefined white-balance settings. Given this set of rendered images, our method learns to estimate weighting maps that are used to blend the rendered images to generate the final corrected image. Through extensive experiments, we show this proposed method produces promising results compared to other alternatives for single-and mixed-illuminant scene color correction. Mahmoud Afifi, Marcus A. Brubaker, Michael S. Brown |
WACV | 3 |
| 2022 | CIE XYZ Net: Unprocessing Images for Low-Level Computer Vision TasksabstractCameras currently allow access to two image states: (i) a minimally processed linear raw-RGB image state (i.e., raw sensor data); or (ii) a highly-processed nonlinear image state (e.g., sRGB). There are many computer vision tasks that work best with a linear image state, such as image deblurring and image dehazing. Unfortunately, the vast majority of images are saved in the nonlinear image state. Because of this, a number of methods have been proposed to "unprocess" nonlinear images back to a raw-RGB state. However, existing unprocessing methods have a drawback because raw-RGB images are sensor-specific. As a result, it is necessary to know which camera produced the sRGB output and use a method or network tailored for that sensor to properly unprocess it. This paper addresses this limitation by exploiting another camera image state that is not available as an output, but it is available inside the camera pipeline. In particular, cameras apply a colorimetric conversion step to convert the raw-RGB image to a device-independent space based on the CIE XYZ color space before they apply the nonlinear photo-finishing. Leveraging this canonical image state, we propose a deep learning framework, CIE XYZ Net, that can unprocess a nonlinear image back to the canonical CIE XYZ image. This image can then be processed by any low-level computer vision operator and re-rendered back to the nonlinear image. We demonstrate the usefulness of the CIE XYZ Net on several low-level vision tasks and show significant gains that can be obtained by this processing framework. Code and dataset are publicly available at https://github.com/mahmoudnafifi/CIE_XYZ_NET. Mahmoud Afifi, Abdelrahman Abdelhamed, Abdullah Abuolaim, Abhijith Punnappurath, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | A Little Bit More: Bitplane-Wise Bit-Depth RecoveryabstractImaging sensors digitize incoming scene light at a dynamic range of 10-12 bits (i.e., 1024-4096 tonal values). The sensor image is then processed onboard the camera and finally quantized to only 8 bits (i.e., 256 tonal values) to conform to prevailing encoding standards. There are a number of important applications, such as high-bit-depth displays and photo editing, where it is beneficial to recover the lost bit depth. Deep neural networks are effective at this bit-depth reconstruction task. Given the quantized low-bit-depth image as input, existing deep learning methods employ a single-shot approach that attempts to either (1) directly estimate the high-bit-depth image, or (2) directly estimate the residual between the high- and low-bit-depth images. In contrast, we propose a training and inference strategy that recovers the residual image bitplane-by-bitplane. Our bitplane-wise learning framework has the advantage of allowing for multiple levels of supervision during training and is able to obtain state-of-the-art results using a simple network architecture. We test our proposed method extensively on several image datasets and demonstrate an improvement from 0.5dB to 2.3dB PSNR over prior methods depending on the quantization level. Abhijith Punnappurath, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Leveraging the Availability of Two Cameras for Illuminant EstimationabstractMost modern smartphones are now equipped with two rear-facing cameras – a main camera for standard imaging and an additional camera to provide wide-angle or telephoto zoom capabilities. In this paper, we leverage the availability of these two cameras for the task of illumination estimation using a small neural network to perform the illumination prediction. Specifically, if the two cameras’ sensors have different spectral sensitivities, the two images provide different spectral measurements of the physical scene. A linear 3 × 3 color transform that maps between these two observations – and that is unique to a given scene illuminant – can be used to train a lightweight neural network comprising no more than 1460 parameters to predict the scene illumination. We demonstrate that this two-camera approach with a lightweight network provides results on par or better than much more complicated illuminant estimation methods operating on a single image. We validate our method’s effectiveness through extensive experiments on radiometric data, a quasi-real two-camera dataset we generated from an existing single camera dataset, as well as a new real image dataset that we captured using a smartphone with two rear-facing cameras. Abdelrahman Abdelhamed, Abhijith Punnappurath, Michael S. Brown |
CVPR | 3 |
| 2021 | HistoGAN: Controlling Colors of GAN-Generated and Real Images via Color HistogramsabstractWhile generative adversarial networks (GANs) can successfully produce high-quality images, they can be challenging to control. Simplifying GAN-based image generation is critical for their adoption in graphic design and artistic work. This goal has led to significant interest in methods that can intuitively control the appearance of images generated by GANs. In this paper, we present HistoGAN, a color histogram-based method for controlling GAN-generated images’ colors. We focus on color histograms as they provide an intuitive way to describe image color while remaining decoupled from domain-specific semantics. Specifically, we introduce an effective modification of the recent StyleGAN architecture [31] to control the colors of GAN-generated images specified by a target color histogram feature. We then describe how to expand HistoGAN to recolor real images. For image recoloring, we jointly train an encoder network along with HistoGAN. The recoloring model, ReHistoGAN, is an unsupervised approach trained to encourage the network to keep the original image’s content while changing the colors based on the given target histogram. We show that this histogram-based approach offers a better way to control GAN-generated and real images’ colors while producing more compelling results compared to existing alternative strategies. Mahmoud Afifi, Marcus A. Brubaker, Michael S. Brown |
CVPR | 3 |
| 2021 | Learning Multi-Scale Photo Exposure CorrectionabstractCapturing photographs with wrong exposures remains a major source of errors in camera-based imaging. Exposure problems are categorized as either: (i) overexposed, where the camera exposure was too long, resulting in bright and washed-out image regions, or (ii) underexposed, where the exposure was too short, resulting in dark regions. Both under- and overexposure greatly reduce the contrast and visual appeal of an image. Prior work mainly focuses on underexposed images or general image enhancement. In contrast, our proposed method targets both over- and underexposure errors in photographs. We formulate the exposure correction problem as two main sub-problems: (i) color enhancement and (ii) detail enhancement. Accordingly, we propose a coarse-to-fine deep neural network (DNN) model, trainable in an end-to-end manner, that addresses each sub-problem separately. A key aspect of our solution is a new dataset of over 24,000 images exhibiting the broadest range of exposure values to date with a corresponding properly exposed image. Our method achieves results on par with existing state-of-the-art methods on underexposed images and yields significant improvements for images suffering from overexposure errors. Mahmoud Afifi, Konstantinos G. Derpanis, Björn Ommer, Michael S. Brown |
CVPR | 4 |
| 2021 | Learning to Reduce Defocus Blur by Realistically Modeling Dual-Pixel DataabstractRecent work has shown impressive results on data-driven defocus deblurring using the two-image views available on modern dual-pixel (DP) sensors. One significant challenge in this line of research is access to DP data. Despite many cameras having DP sensors, only a limited number provide access to the low-level DP sensor images. In addition, capturing training data for defocus deblurring involves a time-consuming and tedious setup requiring the camera’s aperture to be adjusted. Some cameras with DP sensors (e.g., smartphones) do not have adjustable apertures, further limiting the ability to produce the necessary training data. We address the data capture bottleneck by proposing a procedure to generate realistic DP data synthetically. Our synthesis approach mimics the optical image formation found on DP sensors and can be applied to virtual scenes rendered with standard computer software. Leveraging these realistic synthetic DP images, we introduce a recurrent convolutional network (RCN) architecture that improves deblurring results and is suitable for use with single-frame and multi-frame data (e.g., video) captured by DP sensors. Finally, we show that our synthetic DP data is useful for training DNN models targeting video deblurring applications where access to DP data remains challenging. Abdullah Abuolaim, Mauricio Delbracio, Damien Kelly, Michael S. Brown, Peyman Milanfar |
ICCV | 4 |
| 2021 | Spatially Aware Metadata for Raw ReconstructionabstractA camera sensor captures a raw-RGB image that is then processed to a standard RGB (sRGB) image through a series of onboard operations performed by the camera's image signal processor (ISP). Among these processing steps, local tone mapping is one of the most important operations used to enhance the overall appearance of the final rendered sRGB image. For certain applications, it is often desirable to de-render or unprocess the sRGB image back to its original raw-RGB values. This "raw reconstruction" is a challenging task because many of the operations performed by the ISP, including local tone mapping, are nonlinear and difficult to invert. Existing raw reconstruction methods that store specialized metadata at capture time to enable raw recovery ignore local tone mapping and assume that a global transformation exists between the raw-RGB and sRGB color spaces. In this work, we advocate a spatially aware metadata-based raw reconstruction method that is robust to local tone mapping, and yields significantly higher raw reconstruction accuracy (6 dB average PSNR improvement) compared to existing raw reconstruction methods. Our method requires only 0.2% samples of the full-sized image as metadata, has negligible computational overhead at capture time, and can be easily integrated into modern ISPs. Abhijith Punnappurath, Michael S. Brown |
WACV | 2 |
| 2020 | Deep White-Balance EditingabstractWe introduce a deep learning approach to realistically edit an sRGB image's white balance. Cameras capture sensor images that are rendered by their integrated signal processor (ISP) to a standard RGB (sRGB) color space encoding. The ISP rendering begins with a white-balance procedure that is used to remove the color cast of the scene's illumination. The ISP then applies a series of nonlinear color manipulations to enhance the visual quality of the final sRGB image. Recent work by [3] showed that sRGB images that were rendered with the incorrect white balance cannot be easily corrected due to the ISP's nonlinear rendering. The work in [3] proposed a k-nearest neighbor (KNN) solution based on tens of thousands of image pairs. We propose to solve this problem with a deep neural network (DNN) architecture trained in an end-to-end manner to learn the correct white balance. Our DNN maps an input image to two additional white-balance settings corresponding to indoor and outdoor illuminations. Our solution not only is more accurate than the KNN approach in terms of correcting a wrong white-balance setting but also provides the user the freedom to edit the white balance in the sRGB image to other illumination settings. Mahmoud Afifi, Michael S. Brown |
CVPR | 2 |
| 2020 | Defocus Deblurring Using Dual-Pixel Data
Abdullah Abuolaim, Michael S. Brown |
ECCV (10) | 2 |
| 2020 | Modeling Defocus-Disparity in Dual-Pixel SensorsabstractMost modern consumer cameras use dual-pixel (DP) sensors that provide two sub-aperture views of the scene in a single photo capture. The DP sensor was designed to assist the camera's autofocus routine, which examines local disparity in the two sub-aperture views to determine which parts of the image are out of focus. Recently, these DP views have been used for tasks beyond autofocus, such as synthetic bokeh, reflection removal, and depth reconstruction. These recent methods treat the two DP views as stereo image pairs and apply stereo matching algorithms to compute local disparity. However, dual-pixel disparity is not caused by view parallax as in stereo, but instead is attributed to defocus blur that occurs in out-of-focus regions in the image. This paper proposes a new parametric point spread function to model the defocus-disparity that occurs on DP sensors. We apply our model to the task of depth estimation from DP data. An important feature of our model is its ability to exploit the symmetry property of the DP blur kernels at each pixel. We leverage this symmetry property to formulate an unsupervised loss function that does not require ground truth depth. We demonstrate our method's effectiveness on both DSLR and smartphone DP data. Abhijith Punnappurath, Abdullah Abuolaim, Mahmoud Afifi, Michael S. Brown |
ICCP | 4 |
| 2020 | Online Lens Motion Smoothing for Video AutofocusabstractAutofocus (AF) is the process of moving the camera's lens such that desired scene content is in focus. AF for single image capture is a well-studied research topic and most modern cameras have hardware support that allows quick lens movements to optimize image sharpness. How to best perform AF for video is less clear. Conventional wisdom would suggest that each temporal frame should be as sharp as possible. However, unlike single image capture, the effects of the lens movement is visible in the captured video. As a result, there are two parameters to consider in AF for video: sharpness and lens movement. In this paper, we show that users preferred videos with smooth lens movement, even if it results in less overall sharpness. Based on this observation, we propose two novel AF algorithms for video that strive for both smooth lens movement and sharp scene content. Specifically, we introduce (1) a bidirectional long short-term memory (BLSTM) module trained on smooth lens trajectories and (2) a simple weighted moving average (WMA) method that factors in prior lens motion. Both of these methods have demonstrated excellent results in terms of reducing lens movements (up to 64% reduction) without greatly affecting the sharpness (less than 5.2% change in sharpness). Moreover, videos produced using our methods are more preferred by users over conventional AF that aims only for maximizing sharpness. Abdullah Abuolaim, Michael S. Brown |
WACV | 2 |
| 2020 | Learning Raw Image Reconstruction-Aware Deep Image CompressorsabstractDeep learning-based image compressors are actively being explored in an effort to supersede conventional image compression algorithms, such as JPEG. Conventional and deep learning-based compression algorithms focus on minimizing image fidelity errors in the nonlinear standard RGB (sRGB) color space. However, for many computer vision tasks, the sensor's linear raw-RGB image is desirable. Recent work has shown that the original raw-RGB image can be reconstructed using only small amounts of metadata embedded inside the JPEG image [1]. However, [1] relied on the conventional JPEG encoding that is unaware of the raw-RGB reconstruction task. In this paper, we examine the ability of deep image compressors to be "aware" of the additional objective of raw reconstruction. Towards this goal, we describe a general framework that enables deep networks targeting image compression to jointly consider both image fidelity errors and raw reconstruction errors. We describe this approach in two scenarios: (1) the network is trained from scratch using our proposed joint loss, and (2) a network originally trained only for sRGB fidelity loss is later fine-tuned to incorporate our raw reconstruction loss. When compared to sRGB fidelity-only compression, our combined loss leads to appreciable improvements in PSNR of the raw reconstruction with only minor impact on sRGB fidelity as measured by MS-SSIM. Abhijith Punnappurath, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Sensor-Independent Illumination Estimation for DNN Models
Mahmoud Afifi, Michael S. Brown |
BMVC | 2 |
| 2019 | When Color Constancy Goes Wrong: Correcting Improperly White-Balanced ImagesabstractThis paper focuses on correcting a camera image that has been improperly white-balanced. This situation occurs when a camera's auto white balance fails or when the wrong manual white-balance setting is used. Even after decades of computational color constancy research, there are no effective solutions to this problem. The challenge lies not in identifying what the correct white balance should have been, but in the fact that the in-camera white-balance procedure is followed by several camera-specific nonlinear color manipulations that make it challenging to correct the image's colors in post-processing. This paper introduces the first method to explicitly address this problem. Our method is enabled by a dataset of over 65,000 pairs of incorrectly white-balanced images and their corresponding correctly white-balanced images. Using this dataset, we introduce a k-nearest neighbor strategy that is able to compute a nonlinear color mapping function to correct the image's colors. We show our method is highly effective and generalizes well to camera models not in the training set. Mahmoud Afifi, Brian L. Price, Scott Cohen, Michael S. Brown |
CVPR | 4 |
| 2019 | Reflection Removal Using a Dual-Pixel SensorabstractReflection removal is the challenging problem of removing unwanted reflections that occur when imaging a scene that is behind a pane of glass. In this paper, we show that most cameras have an overlooked mechanism that can greatly simplify this task. Specifically, modern DLSR and smartphone cameras use dual pixel (DP) sensors that have two photodiodes per pixel to provide two sub-aperture views of the scene from a single captured image. ``Defocus-disparity'' cues, which are natural by-products of the DP sensor encoded within these two sub-aperture views, can be used to distinguish between image gradients belonging to the in-focus background and those caused by reflection interference. This gradient information can then be incorporated into an optimization framework to recover the background layer with higher accuracy than currently possible from the single captured image. As part of this work, we provide the first image dataset for reflection removal consisting of the sub-aperture views from the DP sensor. Abhijith Punnappurath, Michael S. Brown |
CVPR | 2 |
| 2019 | MarkWhite: An Improved Interactive White-Balance Method for Smartphone Cameras
Abdelrahman Abdelhamed, I. Scott MacKenzie, Michael S. Brown |
Graphics Interface | 3 |
| 2019 | Noise Flow: Noise Modeling With Conditional Normalizing FlowsabstractModeling and synthesizing image noise is an important aspect in many computer vision applications. The long-standing additive white Gaussian and heteroscedastic (signal-dependent) noise models widely used in the literature provide only a coarse approximation of real sensor noise. This paper introduces Noise Flow, a powerful and accurate noise model based on recent normalizing flow architectures. Noise Flow combines well-established basic parametric noise models (e.g., signal-dependent noise) with the flexibility and expressiveness of normalizing flow networks. The result is a single, comprehensive, compact noise model containing fewer than 2500 parameters yet able to represent multiple cameras and gain factors. Noise Flow dramatically outperforms existing noise models, with 0.42 nats/pixel improvement over the camera-calibrated noise level functions, which translates to 52% improvement in the likelihood of sampled noise. Noise Flow represents the first serious attempt to go beyond simple parametric models to one that leverages the power of deep learning and data-driven noise distributions. Abdelrahman Abdelhamed, Marcus A. Brubaker, Michael S. Brown |
ICCV | 3 |
| 2019 | What Else Can Fool Deep Learning? Addressing Color Constancy Errors on Deep Neural Network PerformanceabstractThere is active research targeting local image manipulations that can fool deep neural networks (DNNs) into producing incorrect results. This paper examines a type of global image manipulation that can produce similar adverse effects. Specifically, we explore how strong color casts caused by incorrectly applied computational color constancy - referred to as white balance (WB) in photography - negatively impact the performance of DNNs targeting image segmentation and classification. In addition, we discuss how existing image augmentation methods used to improve the robustness of DNNs are not well suited for modeling WB errors. To address this problem, a novel augmentation method is proposed that can emulate accurate color constancy degradation. We also explore pre-processing training and testing images with a recent WB correction algorithm to reduce the effects of incorrectly white-balanced images. We examine both augmentation and pre-processing strategies on different datasets and demonstrate notable improvements on the CIFAR-10, CIFAR-100, and ADE20K datasets. Mahmoud Afifi, Michael S. Brown |
ICCV | 2 |
| 2018 | A High-Quality Denoising Dataset for Smartphone CamerasabstractThe last decade has seen an astronomical shift from imaging with DSLR and point-and-shoot cameras to imaging with smartphone cameras. Due to the small aperture and sensor size, smartphone images have notably more noise than their DSLR counterparts. While denoising for smartphone images is an active research area, the research community currently lacks a denoising image dataset representative of real noisy images from smartphone cameras with high-quality ground truth. We address this issue in this paper with the following contributions. We propose a systematic procedure for estimating ground truth for noisy images that can be used to benchmark denoising performance for smartphone cameras. Using this procedure, we have captured a dataset - the Smartphone Image Denoising Dataset (SIDD) - of~30,000 noisy images from 10 scenes under different lighting conditions using five representative smartphone cameras and generated their ground truth images. We used this dataset to benchmark a number of denoising algorithms. We show that CNN-based methods perform better when trained on our high-quality dataset than when trained using alternative strategies, such as low-ISO images used as a proxy for ground truth data. Abdelrahman Abdelhamed, Stephen Lin 0001, Michael S. Brown |
CVPR | 3 |
| 2018 | Improving Color Reproduction Accuracy on CamerasabstractOne of the key operations performed on a digital camera is to map the sensor-specific color space to a standard perceptual color space. This procedure involves the application of a white-balance correction followed by a color space transform. The current approach for this colorimetric mapping is based on an interpolation of pre-calibrated color space transforms computed for two fixed illuminations (i.e., two white-balance settings). Images captured under different illuminations are subject to less color accuracy due to the use of this interpolation process. In this paper, we discuss the limitations of the current colorimetric mapping approach and propose two methods that are able to improve color accuracy. We evaluate our approach on seven different cameras and show improvements of up to 30% (DSLR cameras) and 59% (mobile phone cameras) in terms of color reproduction error. Hakki Can Karaimer, Michael S. Brown |
CVPR | 2 |
| 2018 | Classification-Driven Dynamic Image EnhancementabstractConvolutional neural networks rely on image texture and structure to serve as discriminative features to classify the image content. Image enhancement techniques can be used as preprocessing steps to help improve the overall image quality and in turn improve the overall effectiveness of a CNN. Existing image enhancement methods, however, are designed to improve the perceptual quality of an image for a human observer. In this paper, we are interested in learning CNNs that can emulate image enhancement and restoration, but with the overall goal to improve image classification and not necessarily human perception. To this end, we present a unified CNN architecture that uses a range of enhancement filters that can enhance image-specific details via end-to-end dynamic filter learning. We demonstrate the effectiveness of this strategy on four challenging benchmark datasets for fine-grained, object, scene, and texture classification: CUB-200-2011, PASCAL-VOC2007, MIT-Indoor, and DTD. Experiments using our proposed enhancement show promising results on all the datasets. In addition, our approach is capable of improving the performance of all generic CNN architectures. Vivek Sharma 0001, Ali Diba, Davy Neven, Michael S. Brown, Luc Van Gool, Rainer Stiefelhagen |
CVPR | 4 |
| 2018 | Revisiting Autofocus for Smartphone Cameras
Abdullah Abuolaim, Abhijith Punnappurath, Michael S. Brown |
ECCV (15) | 3 |
| 2018 | A Hybrid Prior Model for Tunable Diode Laser Absorption TomographyabstractModel based methods have gained popularity in the past few decades in reconstruction problems particularly when the measurement data is sparse. In model based inference, apart from a model for the measurements, there exists a model for the unknown signal to be reconstructed, called the prior model. Model based methods tend to do very well when the prior model is accurate and representative of real world behavior of the unknown signal. Often these priors are trained from some training data, and therefore, the accuracy of the reconstructions depends inherently on the accuracy of the training data. The reconstructions can come out to be highly biased if the training data is not representative of the actual signal. In this paper, we propose a new hybrid prior model that combines a Markov Random Field model with a Gaussian model trained from a sparse training set. We combine the models using a mixing coefficient Υ E [0, 1], that controls the influence of each of the models. Our main contribution is in the way we combine the two models to produce a whole continuum of prior models for different values of Y where Y can be tuned according to how trustworthy the training set is. Reconstruction results show that our hybrid prior models produces high quality reconstructions even when the training data is not representative. Zeeshan Nadir, Charles A. Bouman, Kristin M. Rice, Michael S. Brown |
ICIP | 4 |
| 2018 | RAW Image Reconstruction Using a Self-contained sRGB-JPEG Image with Small Memory OverheadabstractMost camera images are saved as 8-bit standard RGB (sRGB) compressed JPEGs. Even when JPEG compression is set to its highest quality, the encoded sRGB image has been significantly processed in terms of color and tone manipulation. This makes sRGB-JPEG images undesirable for many computer vision tasks that assume a direct relationship between pixel values and incoming light. For such applications, the RAW image format is preferred, as RAW represents a minimally processed, sensor-specific RGB image that is linear with respect to scene radiance. The drawback with RAW images, however, is that they require large amounts of storage and are not well-supported by many imaging applications. To address this issue, we present a method to encode the necessary data within an sRGB-JPEG image to reconstruct a high-quality RAW image. Our approach requires no calibration of the camera's colorimetric properties and can reconstruct the original RAW to within 0.5% error with a small memory overhead for the additional data (e.g., 128 KB). More importantly, our output is a fully self-contained 100% compliant sRGB-JPEG file that can be used as-is, not affecting any existing image workflow-the RAW image data can be extracted when needed, or ignored otherwise. We detail our approach and show its effectiveness against competing strategies. Nguyen Ho Man Rang, Michael S. Brown |
Int. J. Comput. Vis. | 2 |
| 2017 | Why You Should Forget Luminance Conversion and Do Something BetterabstractOne of the most frequently applied low-level operations in computer vision is the conversion of an RGB camera image into its luminance representation. This is also one of the most incorrectly applied operations. Even our most trusted softwares, Matlab and OpenCV, do not perform luminance conversion correctly. In this paper, we examine the main factors that make proper RGB to luminance conversion difficult, in particular: 1) incorrect white-balance, 2) incorrect gamma/tone-curve correction, and 3) incorrect equations. Our analysis shows errors up to 50% for various colors are not uncommon. As a result, we argue that for most computer vision problems there is no need to attempt luminance conversion, instead, there are better alternatives depending on the task. Nguyen Ho Man Rang, Michael S. Brown |
CVPR | 2 |
| 2017 | A Non-local Low-Rank Framework for Ultrasound Speckle ReductionabstractSpeckle refers to the granular patterns that occur in ultrasound images due to wave interference. Speckle removal can greatly improve the visibility of the underlying structures in an ultrasound image and enhance subsequent post processing. We present a novel framework for speckle removal based on low-rank non-local filtering. Our approach works by first computing a guidance image that assists in the selection of candidate patches for non-local filtering in the face of significant speckles. The candidate patches are further refined using a low-rank minimization estimated using a truncated weighted nuclear norm (TWNN) and structured sparsity. We show that the proposed filtering framework produces results that outperform state-of-the-art methods both qualitatively and quantitatively. This framework also provides better segmentation results when used for pre-processing ultrasound images. Lei Zhu 0003, Chi-Wing Fu, Michael S. Brown, Pheng-Ann Heng |
CVPR | 3 |
| 2017 | Group-Theme Recoloring for Multi-Image Color ConsistencyabstractAbstract Modifying the colors of an image is a fundamental editing task with a wide range of methods available. Manipulating multiple images to share similar colors is more challenging, with limited tools available. Methods such as color transfer are effective in making an image share similar colors with a target image; however, color transfer is not suitable for modifying multiple images. Approaches for color consistency for photo collections give good results when the photo collection contains similar scene content, but are not applicable for general input images. To address these gaps, we propose an application framework for achieving color consistency for multi‐image input. Our framework derives a group color theme from the input images′ individual color palettes and uses this group color theme to recolor the image collection. This group‐theme recoloring provides an effective way to ensure color consistency among multiple images and naturally lends itself to the inclusion of an additional external color theme. We detail our group‐theme recoloring approach and demonstrate its effectiveness on a number of examples. Nguyen Ho Man Rang, Brian L. Price, Scott Cohen, Michael S. Brown |
Comput. Graph. Forum | 4 |
| 2017 | Haze visibility enhancement: A Survey and quantitative benchmarking
Yu Li 0003, Shaodi You, Michael S. Brown, Robby T. Tan |
Comput. Vis. Image Underst. | 3 |
| 2017 | Single Image Rain Streak Decomposition Using Layer PriorsabstractRain streaks impair visibility of an image and introduce undesirable interference that can severely affect the performance of computer vision and image analysis systems. Rain streak removal algorithms try to recover a rain streak free background scene. In this paper, we address the problem of rain streak removal from a single image by formulating it as a layer decomposition problem, with a rain streak layer superimposed on a background layer containing the true scene content. Existing decomposition methods that address this problem employ either sparse dictionary learning methods or impose a low rank structure on the appearance of the rain streaks. While these methods can improve the overall visibility, their performance can often be unsatisfactory, for they tend to either over-smooth the background images or generate -images that still contain noticeable rain streaks. To address the problems, we propose a method that imposes priors for both the background and rain streak layers. These priors are based on Gaussian mixture models learned on small patches that can accommodate a variety of background appearances as well as the appearance of the rain streaks. Moreover, we introduce a structure residue recovery step to further separate the background residues and improve the decomposition quality. Quantitative evaluation shows our method outperforms existing methods by a large margin. We overview our method and demonstrate its effectiveness over prior work on a number of examples. Yu Li 0003, Robby T. Tan, Xiaojie Guo 0001, Jiangbo Lu, Michael S. Brown |
IEEE Trans. Image Process. | 5 |
| 2016 | Two Illuminant Estimation and User Correction PreferenceabstractThis paper examines the problem of white-balance correction when a scene contains two illuminations. This is a two step process: 1) estimate the two illuminants, and 2) correct the image. Existing methods attempt to estimate a spatially varying illumination map, however, results are error prone and the resulting illumination maps are too lowresolution to be used for proper spatially varying whitebalance correction. In addition, the spatially varying nature of these methods make them computationally intensive. We show that this problem can be effectively addressed by not attempting to obtain a spatially varying illumination map, but instead by performing illumination estimation on large sub-regions of the image. Our approach is able to detect when distinct illuminations are present in the image and accurately measure these illuminants. Since our proposed strategy is not suitable for spatially varying image correction, a user study is performed to see if there is a preference for how the image should be corrected when two illuminants are present, but only a global correction can be applied. The user study shows that when the illuminations are distinct, there is a preference for the outdoor illumination to be corrected resulting in warmer final result. We use these collective findings to demonstrate an effective two illuminant estimation scheme that produces corrected images that users prefer. Dongliang Cheng, Abdelrahman Kamel, Brian L. Price, Scott Cohen, Michael S. Brown |
CVPR | 5 |
| 2016 | Rain Streak Removal Using Layer PriorsabstractThis paper addresses the problem of rain streak removal from a single image. Rain streaks impair visibility of an image and introduce undesirable interference that can severely affect the performance of computer vision algorithms. Rain streak removal can be formulated as a layer decomposition problem, with a rain streak layer superimposed on a background layer containing the true scene content. Existing decomposition methods that address this problem employ either dictionary learning methods or impose a low rank structure on the appearance of the rain streaks. While these methods can improve the overall visibility, they tend to leave too many rain streaks in the background image or over-smooth the background image. In this paper, we propose an effective method that uses simple patch-based priors for both the background and rain layers. These priors are based on Gaussian mixture models and can accommodate multiple orientations and scales of the rain streaks. This simple approach removes rain streaks better than the existing methods qualitatively and quantitatively. We overview our method and demonstrate its effectiveness over prior work on a number of examples. Yu Li 0003, Robby T. Tan, Xiaojie Guo 0001, Jiangbo Lu, Michael S. Brown |
CVPR | 5 |
| 2016 | RAW Image Reconstruction Using a Self-Contained sRGB-JPEG Image with Only 64 KB OverheadabstractMost camera images are saved as 8-bit standard RGB (sRGB) compressed JPEGs. Even when JPEG compression is set to its highest quality, the encoded sRGB image has been significantly processed in terms of color and tone manipulation. This makes sRGB-JPEG images undesirable for many computer vision tasks that assume a direct relationship between pixel values and incoming light. For such applications, the RAW image format is preferred, as RAW represents a minimally processed, sensor-specific RGB image with higher dynamic range that is linear with respect to scene radiance. The drawback with RAW images, however, is that they require large amounts of storage and are not well-supported by many imaging applications. To address this issue, we present a method to encode the necessary metadata within an sRGB image to reconstruct a high-quality RAW image. Our approach requires no calibration of the camera and can reconstruct the original RAW to within 0:3% error with only a 64 KB overhead for the additional data. More importantly, our output is a fully selfcontained 100% complainant sRGB-JPEG file that can be used as-is, not affecting any existing image workflow - the RAW image can be extracted when needed, or ignored otherwise. We detail our approach and show its effectiveness against competing strategies. Nguyen Ho Man Rang, Michael S. Brown |
CVPR | 2 |
| 2016 | Do It Yourself Hyperspectral Imaging with Everyday Digital CamerasabstractCapturing hyperspectral images requires expensive and specialized hardware that is not readily accessible to most users. Digital cameras, on the other hand, are significantly cheaper in comparison and can be easily purchased and used. In this paper, we present a framework for reconstructing hyperspectral images by using multiple consumer-level digital cameras. Our approach works by exploiting the different spectral sensitivities of different camera sensors. In particular, due to the differences in spectral sensitivities of the cameras, different cameras yield different RGB measurements for the same spectral signal. We introduce an algorithm that is able to combine and convert these different RGB measurements into a single hyperspectral image for both indoor and outdoor scenes. This camera-based approach allows hyperspectral imaging at a fraction of the cost of most existing hyperspectral hardware. We validate the accuracy of our reconstruction against ground truth hyperspectral images (using both synthetic and real cases) and show its usage on relighting applications. Seoung Wug Oh, Michael S. Brown, Marc Pollefeys, Seon Joo Kim |
CVPR | 2 |
| 2016 | A Software Platform for Manipulating the Camera Imaging Pipeline
Hakki Can Karaimer, Michael S. Brown |
ECCV (1) | 2 |
| 2016 | Basal Slice Detection Using Long-Axis Segmentation for Cardiac AnalysisabstractEstimating blood volume of the left ventricle (LV) in the end-diastolic and end-systolic phases is important in diagnosing cardiovascular diseases. Proper estimation of the volume requires knowledge of which MRI slice contains the topmost basal region of the LV. Automatic basal slice detection has proved challenging; as a result, basal slice detection remains a manual task which is prone to inter-observer variability. This paper presents a novel method that is able to track the basal slice over the whole cardiac cycle. The method was tested on 56 healthy and pathological cases and was able to identify the basal slices similar to experts’ selection for 80 % and 85 % of the cases for end-diastole and end-systole, respectively. This provides a significant improvement over the leading state-of-the-art approach that obtained 59 % and 44 % agreement with experts on the same input. Mahsa Paknezhad, Michael S. Brown, Stéphanie Marchesseau |
MICCAI (3) | 2 |
| 2016 | Mosaicing scenes with a quadcopterabstractThis paper focuses on a method of constructing panoramas from a quadcopter, and a new mosaicing sub-problem when the scene contains significant regions of vacant spaces. These vacant spaces yield little to no features to match input images and hence challenge existing mosaicing techniques. We describe a framework that is able to handle this unique input by leveraging the availability of the inertial measurement unit (IMU) data from the quadcopter. Specifically, our method uses the imprecise IMU data accompanying a video to select a subset of images that contain interesting scene content. When the scene is such that this subset contains no vacant space, an appropriate panorama is effected; however, with featureless spaces, existing mo-saicing methods do not work. In this paper, the subset is partitioned into multiple clusters. These subsets can now be stitched into a series of mini-panoramas, but a complete mosaic is not yet available. The gaps between these minipanoramas represent regions of featureless spaces in the scene. Therefore, we once again use the IMU data together with a novel stereo reconstruction to determine appropriate portions of the images to complete the panorama. We demonstrate the efficacy of our approach on a number of input sequences that cannot be mosaiced by existing methods. Meghshyam G. Prasad, Sharat Chandran, Michael S. Brown |
WACV | 3 |
| 2015 | Effective learning-based illuminant estimation using simple featuresabstractIllumination estimation is the process of determining the chromaticity of the illumination in an imaged scene in order to remove undesirable color casts through white-balancing. While computational color constancy is a well-studied topic in computer vision, it remains challenging due to the ill-posed nature of the problem. One class of techniques relies on low-level statistical information in the image color distribution and works under various assumptions (e.g. Grey-World, White-Patch, etc). These methods have an advantage that they are simple and fast, but often do not perform well. More recent state-of-the-art methods employ learning-based techniques that produce better results, but often rely on complex features and have long evaluation and training times. In this paper, we present a learning-based method based on four simple color features and show how to use this with an ensemble of regression trees to estimate the illumination. We demonstrate that our approach is not only faster than existing learning-based methods in terms of both evaluation and training time, but also gives the best results reported to date on modern color constancy data sets. Dongliang Cheng, Brian L. Price, Scott Cohen, Michael S. Brown |
CVPR | 4 |
| 2015 | Beyond White: Ground Truth Colors for Color Constancy CorrectionabstractA limitation in color constancy research is the inability to establish ground truth colors for evaluating corrected images. Many existing datasets contain images of scenes with a color chart included, however, only the chart's neutral colors (grayscale patches) are used to provide the ground truth for illumination estimation and correction. This is because the corrected neutral colors are known to lie along the achromatic line in the camera's color space (i.e. R=G=B), the correct RGB values of the other color patches are not known. As a result, most methods estimate a 3*3 diagonal matrix that ensures only the neutral colors are correct. In this paper, we describe how to overcome this limitation. Specifically, we show that under certain illuminations, a diagonal 3*3 matrix is capable of correcting not only neutral colors, but all the colors in a scene. This finding allows us to find the ground truth RGB values for the color chart in the camera's color space. We show how to use this information to correct all the images in existing datasets to have correct colors. Working from these new color corrected datasets, we describe how to modify existing color constancy algorithms to perform better image correction. Dongliang Cheng, Brian L. Price, Scott Cohen, Michael S. Brown |
ICCV | 4 |
| 2015 | SPM-BP: Sped-Up PatchMatch Belief Propagation for Continuous MRFsabstractMarkov random fields are widely used to model many computer vision problems that can be cast in an energy minimization framework composed of unary and pairwise potentials. While computationally tractable discrete optimizers such as Graph Cuts and belief propagation (BP) exist for multi-label discrete problems, they still face prohibitively high computational challenges when the labels reside in a huge or very densely sampled space. Integrating key ideas from PatchMatch of effective particle propagation and resampling, PatchMatch belief propagation (PMBP) has been demonstrated to have good performance in addressing continuous labeling problems and runs orders of magnitude faster than Particle BP (PBP). However, the quality of the PMBP solution is tightly coupled with the local window size, over which the raw data cost is aggregated to mitigate ambiguity in the data constraint. This dependency heavily influences the overall complexity, increasing linearly with the window size. This paper proposes a novel algorithm called sped-up PMBP (SPM-BP) to tackle this critical computational bottleneck and speeds up PMBP by 50-100 times. The crux of SPM-BP is on unifying efficient filter-based cost aggregation and message passing with PatchMatch-based particle generation in a highly effective way. Though simple in its formulation, SPM-BP achieves superior performance for sub-pixel accurate stereo and optical-flow on benchmark datasets when compared with more complex and task-specific approaches. Yu Li 0003, Dongbo Min, Michael S. Brown, Minh N. Do, Jiangbo Lu |
ICCV | 3 |
| 2015 | Nighttime Haze Removal with Glow and Multiple Light ColorsabstractThis paper focuses on dehazing nighttime images. Most existing dehazing methods use models that are formulated to describe haze in daytime. Daytime models assume a single uniform light color attributed to a light source not directly visible in the scene. Nighttime scenes, however, commonly include visible lights sources with varying colors. These light sources also often introduce noticeable amounts of glow that is not present in daytime haze. To address these effects, we introduce a new nighttime haze model that accounts for the varying light sources and their glow. Our model is a linear combination of three terms: the direct transmission, airlight and glow. The glow term represents light from the light sources that is scattered around before reaching the camera. Based on the model, we propose a framework that first reduces the effect of the glow in the image, resulting in a nighttime image that consists of direct transmission and airlight only. We then compute a spatially varying atmospheric light map that encodes light colors locally. This atmospheric map is used to predict the transmission, which we use to obtain our nighttime scene reflection image. We demonstrate the effectiveness of our nighttime haze model and correction method on a number of examples and compare our results with existing daytime and nighttime dehazing methods' results. Yu Li 0003, Robby T. Tan, Michael S. Brown |
ICCV | 3 |
| 2015 | Fast and Effective L0 Gradient Minimization by Region FusionabstractL0gradient minimization can be applied to an input signal to control the number of non-zero gradients. This is useful in reducing small gradients generally associated with signal noise, while preserving important signal features. In computer vision, L0gradient minimization has found applications in image denoising, 3D mesh denoising, and image enhancement. Minimizing the L0norm, however, is an NP-hard problem because of its non-convex property. As a result, existing methods rely on approximation strategies to perform the minimization. In this paper, we present a new method to perform L0gradient minimization that is fast and effective. Our method uses a descent approach based on region fusion that converges faster than other methods while providing a better approximation of the optimal L0norm. In addition, our method can be applied to both 2D images and 3D mesh topologies. The effectiveness of our approach is demonstrated on a number of examples. Nguyen Ho Man Rang, Michael S. Brown |
ICCV | 2 |
| 2015 | Aesthetic Interactive Hue Manipulation for Natural Scene Images
Jinze Yu 0003, Martin Constable, Junyan Wang 0002, Kap Luk Chan, Michael S. Brown |
PSIVT | 5 |
| 2015 | A Motion Blur Resilient Fiducial for Quadcopter ImagingabstractFiducials are commonly placed in environments to provide a uniquely identifiable object in the scene. In quad copter applications, these fiducials are often used to evaluate planning algorithms given that ground truth positions can be detected from the quad copter's camera. Low cost quad copters, however, are subject to quick and unstable motions that can cause significant motion blur that severely affects the detection rate of existing fiducials. This problem motivated us to design a fiducial that is robust to motion blur. Our proposed design uses concentric circles with the observation that the direction perpendicular to the motion blur direction will be relatively unaffected by the blur. As a result, an appropriate fiducial code orthogonal to the blur direction can be recognized. Since the direction of motion blur is unknown, the circular design is good for all motion blur directions. We describe the design of binary fiducials, and also a detection algorithm. We show that our marker can significantly outperform existing fiducials in scenes captured with a quad copter. Meghshyam G. Prasad, Sharat Chandran, Michael S. Brown |
WACV | 3 |
| 2014 | Single Image Layer Separation Using Relative SmoothnessabstractThis paper addresses extracting two layers from an image where one layer is smoother than the other. This problem arises most notably in intrinsic image decomposition and reflection interference removal. Layer decomposition from a single-image is inherently ill-posed and solutions require additional constraints to be enforced. We introduce a novel strategy that regularizes the gradients of the two layers such that one has a long tail distribution and the other a short tail distribution. While imposing the long tail distribution is a common practice, our introduction of the short tail distribution on the second layer is unique. We formulate our problem in a probabilistic framework and describe an optimization scheme to solve this regularization with only a few iterations. We apply our approach to the intrinsic image and reflection removal problems and demonstrate high quality layer separation on par with other techniques but being significantly faster than prevailing methods. Yu Li 0003, Michael S. Brown |
CVPR | 2 |
| 2014 | Raw-to-Raw: Mapping between Image Sensor Color ResponsesabstractCamera images saved in raw format are being adopted in computer vision tasks since raw values represent minimally processed sensor responses. Camera manufacturers, however, have yet to adopt a standard for raw images and current raw-rgb values are device specific due to different sensors spectral sensitivities. This results in significantly different raw images for the same scene captured with different cameras. This paper focuses on estimating a mapping that can convert a raw image of an arbitrary scene and illumination from one camera's raw space to another. To this end, we examine various mapping strategies including linear and non-linear transformations applied both in a global and illumination-specific manner. We show that illumination-specific mappings give the best result, however, at the expense of requiring a large number of transformations. To address this issue, we introduce an illumination-independent mapping approach that uses white-balancing to assist in reducing the number of required transformations. We show that this approach achieves state-of-the-art results on a range of consumer cameras and images of arbitrary scenes and illuminations. Nguyen Ho Man Rang, Dilip K. Prasad, Michael S. Brown |
CVPR | 3 |
| 2014 | A Contrast Enhancement Framework with JPEG Artifacts Suppression
Yu Li 0003, Fangfang Guo, Robby T. Tan, Michael S. Brown |
ECCV (2) | 4 |
| 2014 | Training-Based Spectral Reconstruction from a Single RGB Image
Nguyen Ho Man Rang, Dilip K. Prasad, Michael S. Brown |
ECCV (7) | 3 |
| 2014 | Tomographic reconstruction of flowing gases using sparse trainingabstractTunable Diode Laser Absorption Spectroscopy (TDLAS) is an emerging technique for simultaneous sensing of temperature and concentration of gaseous media. However, simultaneous reconstruction of temperature and concentration using TDLAS measurements is a nonlinear inverse problem and unlike other forms of computed tomography (CT), it is typically not possible to take a large number of projection measurements; so reconstructions are often computed using simplistic assumptions that limit the usability of the results. In this paper, we present a fast algorithm for model-based iterative reconstruction (MBIR) of TDLAS data. Our TDLAS-MBIR method uses a nonlinear forward model based on the physics of light absorption and incorporates a holistic prior model that can be learned from very sparse training data. Reconstructions performed on computational fluid dynamics (CFD) phantoms show that our proposed reconstruction algorithm is fast; works well when the number of pixels, p, far exceeds the number of measurements, M; is robust against noise; and produces good reconstructions using few training examples for the prior model. Zeeshan Nadir, Michael S. Brown, Mary L. Comer, Charles A. Bouman |
ICIP | 2 |
| 2014 | Fast rotation search for real-time interactive point cloud registrationabstractOur goal is the registration of multiple 3D point clouds obtained from LIDAR scans of underground mines. Such a capability is crucial to the surveying and planning operations in mining. Often, the point clouds only partially overlap and initial alignment is unavailable. Here, we propose an interactive user-assisted point cloud registration system. Guided by the system, the user's role is simply to identify and search for overlapping regions across the point clouds. Specifically, given two point sets, the user clicks on a point in one set, then simply hovers the mouse on the other set to find a matching point. Each mouse position gives rise to a translation, and our system instantly optimises the rotation that aligns the point clouds. Tat-Jun Chin, Álvaro Parra Bustos, Michael S. Brown, David Suter |
I3D | 3 |
| 2014 | Illuminant Aware Gamut-Based Color TransferabstractAbstract This paper proposes a new approach for color transfer between two images. Our method is unique in its consideration of the scene illumination and the constraint that the mapped image must be within the color gamut of the target image. Specifically, our approach first performs a white‐balance step on both images to remove color casts caused by different illuminations in the source and target image. We then align each image to share the same ‘white axis’ and perform a gradient preserving histogram matching technique along this axis to match the tone distribution between the two images. We show that this illuminant‐aware strategy gives a better result than directly working with the original source and target image's luminance channel as done by many previous methods. Afterwards, our method performs a full gamut‐based mapping technique rather than processing each channel separately. This guarantees that the colors of our transferred image lie within the target gamut. Our experimental results show that this combined illuminant‐aware and gamut‐based strategy produces more compelling results than previous methods. We detail our approach and demonstrate its effectiveness on a number of examples. Nguyen Ho Man Rang, S. J. Kim, Michael S. Brown |
Comput. Graph. Forum | 3 |
| 2014 | As-Projective-As-Possible Image Stitching with Moving DLTabstractThe success of commercial image stitching tools often leads to the impression that image stitching is a "solved problem". The reality, however, is that many tools give unconvincing results when the input photos violate fairly restrictive imaging assumptions; the main two being that the photos correspond to views that differ purely by rotation, or that the imaged scene is effectively planar. Such assumptions underpin the usage of 2D projective transforms or homographies to align photos. In the hands of the casual user, such conditions are often violated, yielding misalignment artifacts or "ghosting" in the results. Accordingly, many existing image stitching tools depend critically on post-processing routines to conceal ghosting. In this paper, we propose a novel estimation technique called Moving Direct Linear Transformation (Moving DLT) that is able to tweak or fine-tune the projective warp to accommodate the deviations of the input data from the idealized conditions. This produces as-projective-as-possible image alignment that significantly reduces ghosting without compromising the geometric realism of perspective image stitching. Our technique thus lessens the dependency on potentially expensive postprocessing algorithms. In addition, we describe how multiple as-projective-as-possible warps can be simultaneously refined via bundle adjustment to accurately align multiple images for large panorama creation. Julio Zaragoza, Tat-Jun Chin, Quoc-Huy Tran, Michael S. Brown, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2014 | High-Quality Depth Map Upsampling and Completion for RGB-D CamerasabstractThis paper describes an application framework to perform high-quality upsampling and completion on noisy depth maps. Our framework targets a complementary system setup, which consists of a depth camera coupled with an RGB camera. Inspired by a recent work that uses a nonlocal structure regularization, we regularize depth maps in order to maintain fine details and structures. We extend this regularization by combining the additional high-resolution RGB input when upsampling a low-resolution depth map together with a weighting scheme that favors structure details. Our technique is also able to repair large holes in a depth map with consideration of structures and discontinuities utilizing edge information from the RGB input. Quantitative and qualitative results show that our method outperforms existing approaches for depth map upsampling and completion. We describe the complete process for this system, including device calibration, scene warping for input alignment, and even how our framework can be extended for video depth-map completion with the consideration of temporal coherence. Jaesik Park, Hyeongwoo Kim, Yu-Wing Tai, Michael S. Brown, In-So Kweon |
IEEE Trans. Image Process. | 4 |
| 2014 | DEB: Definite Error Bounded Tangent Estimator for Digital CurvesabstractWe propose a simple and fast method for tangent estimation of digital curves. This geometric-based method uses a small local region for tangent estimation and has a definite upper bound error for continuous as well as digital conics, i.e., circles, ellipses, parabolas, and hyperbolas. Explicit expressions of the upper bounds for continuous and digitized curves are derived, which can also be applied to nonconic curves. Our approach is benchmarked against 72 contemporary tangent estimation methods and demonstrates good performance for conic, nonconic, and noisy curves. In addition, we demonstrate a good multigrid and isotropic performance and low computational complexity of O(1) and better performance than most methods in terms of maximum and average errors in tangent computation for a large variety of digital curves. Dilip K. Prasad, Maylor K. H. Leung, Hiok Chai Quek, Michael S. Brown |
IEEE Trans. Image Process. | 4 |
| 2013 | As-Projective-As-Possible Image Stitching with Moving DLTabstractWe investigate projective estimation under model inadequacies, i.e., when the underpinning assumptions of the projective model are not fully satisfied by the data. We focus on the task of image stitching which is customarily solved by estimating a projective warp - a model that is justified when the scene is planar or when the views differ purely by rotation. Such conditions are easily violated in practice, and this yields stitching results with ghosting artefacts that necessitate the usage of deghosting algorithms. To this end we propose as-projective-as-possible warps, i.e., warps that aim to be globally projective, yet allow local non-projective deviations to account for violations to the assumed imaging conditions. Based on a novel estimation technique called Moving Direct Linear Transformation (Moving DLT), our method seamlessly bridges image regions that are inconsistent with the projective model. The result is highly accurate image stitching, with significantly reduced ghosting effects, thus lowering the dependency on post hoc deghosting. Julio Zaragoza, Tat-Jun Chin, Michael S. Brown, David Suter |
CVPR | 3 |
| 2013 | A Learning-Based Approach to Reduce JPEG Artifacts in Image MattingabstractSingle image matting techniques assume high-quality input images. The vast majority of images on the web and in personal photo collections are encoded using JPEG compression. JPEG images exhibit quantization artifacts that adversely affect the performance of matting algorithms. To address this situation, we propose a learning-based post-processing method to improve the alpha mattes extracted from JPEG images. Our approach learns a set of sparse dictionaries from training examples that are used to transfer details from high-quality alpha mattes to alpha mattes corrupted by JPEG compression. Three different dictionaries are defined to accommodate different object structure (long hair, short hair, and sharp boundaries). A back-projection criteria combined within an MRF framework is used to automatically select the best dictionary to apply on the object's local boundary. We demonstrate that our method can produces superior results over existing state-of-the-art matting algorithms on a variety of inputs and compression levels. Inchang Choi, Sunyeong Kim, Michael S. Brown, Yu-Wing Tai |
ICCV | 3 |
| 2013 | Exploiting Reflection Change for Automatic Reflection RemovalabstractThis paper introduces an automatic method for removing reflection interference when imaging a scene behind a glass surface. Our approach exploits the subtle changes in the reflection with respect to the background in a small set of images taken at slightly different view points. Key to this idea is the use of SIFT-flow to align the images such that a pixel-wise comparison can be made across the input set. Gradients with variation across the image set are assumed to belong to the reflected scenes while constant gradients are assumed to belong to the desired background scene. By correctly labelling gradients belonging to reflection or background, the background scene can be separated from the reflection interference. Unlike previous approaches that exploit motion, our approach does not make any assumptions regarding the background or reflected scenes' geometry, nor requires the reflection to be static. This makes our approach practical for use in casual imaging scenarios. Our approach is straight forward and produces good results compared with existing methods. Yu Li 0003, Michael S. Brown |
ICCV | 2 |
| 2013 | Offline Mobile Instance Retrieval with a Small Memory FootprintabstractExisting mobile image instance retrieval applications assume a network-based usage where image features are sent to a server to query an online visual database. In this scenario, there are no restrictions on the size of the visual database. This paper, however, examines how to perform this same task offline, where the entire visual index must reside on the mobile device itself within a small memory footprint. Such solutions have applications on location recognition and product recognition. Mobile instance retrieval requires a significant reduction in the visual index size. To achieve this, we describe a set of strategies that can reduce the visual index up to 60-80 times compared to a standard instance retrieval implementation found on desktops or servers. While our proposed reduction steps affect the overall mean Average Precision (mAP), they are able to maintain a good Precision for the top K results (PK). We argue that for such offline application, maintaining a good PKis sufficient. The effectiveness of this approach is demonstrated on several standard databases. A working application designed for a remote historical site is also presented. This application is able to reduce an 50,000 image index structure to 25 MBs while providing a precision of 97% for P10and 100% for P1. Jayaguru Panda, Michael S. Brown, C. V. Jawahar |
ICCV | 2 |
| 2013 | Phenotype Detection in Morphological Mutant Mice Using Deformation Features
Sharmili Roy, Xi Liang 0006, Asanobu Kitamoto, Masaru Tamura, Toshihiko Shiroishi, Michael S. Brown |
MICCAI (3) | 6 |
| 2013 | A 3D Imaging Framework Based on High-Resolution Photometric-Stereo and Low-Resolution Depth
Zheng Lu 0002, Yu-Wing Tai, Fanbo Deng, Moshe Ben-Ezra, Michael S. Brown |
Int. J. Comput. Vis. | 5 |
| 2013 | Nonlinear Camera Response Functions and Image Deblurring: Theoretical Analysis and PracticeabstractThis paper investigates the role that nonlinear camera response functions (CRFs) have on image deblurring. We present a comprehensive study to analyze the effects of CRFs on motion deblurring. In particular, we show how nonlinear CRFs can cause a spatially invariant blur to behave as a spatially varying blur. We prove that such nonlinearity can cause large errors around edges when directly applying deconvolution to a motion blurred image without CRF correction. These errors are inevitable even with a known point spread function (PSF) and with state-of-the-art regularization-based deconvolution algorithms. In addition, we show how CRFs can adversely affect PSF estimation algorithms in the case of blind deconvolution. To help counter these effects, we introduce two methods to estimate the CRF directly from one or more blurred images when the PSF is known or unknown. Our experimental results on synthetic and real images validate our analysis and demonstrate the robustness and accuracy of our approaches. Yu-Wing Tai, Sunyeong Kim, Seon Joo Kim, Feng Li 0005, Jie Yang 0002, Jingyi Yu 0001, Yasuyuki Matsushita, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2012 | Color-Aware Regularization for Gradient Domain Image Manipulation
Fanbo Deng, Seon Joo Kim, Yu-Wing Tai, Michael S. Brown |
ACCV (4) | 4 |
| 2012 | Nonlinear camera response functions and image deblurringabstractThis paper investigates the role that nonlinear camera response functions (CRFs) have on image deblurring. In particular, we show how nonlinear CRFs can cause a spatially invariant blur to behave as a spatially varying blur. This can result in noticeable ringing artifacts when deconvolution is applied even with a known point spread function (PSF). In addition, we show how CRFs can adversely affect PSF estimation algorithms in the case of blind deconvolution. To help counter these effects, we introduce two methods to estimate the CRF directly from one or more blurred images when the PSF is known or unknown. While not as accurate as conventional CRF estimation algorithms based on multiple exposures or calibration patterns, our approach is still quite effective in improving deblurring results in situations where the CRF is unknown. Sunyeong Kim, Yu-Wing Tai, Seon Joo Kim, Michael S. Brown, Yasuyuki Matsushita |
CVPR | 4 |
| 2012 | Synthesizing oil painting surface geometry from a single photographabstractWe present an approach to synthesize the subtle 3D relief and texture of oil painting brush strokes from a single photograph. This task is unique from traditional synthesize algorithms due to its mixed modality between the input and output; i.e., our goal is to synthesize surface normals given an intensity image input. To accomplish this task, we propose a framework that first applies intrinsic image decomposition to produce a pair of initial normal maps. These maps are combined into a conditional random field (CRF) optimization framework that incorporates additional information derived from a training set consisting of normals captured using photometric stereo on oil paintings with similar brush styles. Additional constraints are incorporated into the CRF framework to further ensures smoothness and preserve brush stroke edges. Our results show that this approach can produce compelling reliefs that are often indistinguishable from results captured using photometric stereo. Zheng Lu 0002, Xiaogang Wang 0001, Ying-Qing Xu, Moshe Ben-Ezra, Xiaoou Tang, Michael S. Brown |
CVPR | 7 |
| 2012 | Nonuniform Lattice Regression for Modeling the Camera Imaging Pipeline
Hai Ting Lin, Zheng Lu 0002, Seon Joo Kim, Michael S. Brown |
ECCV (1) | 4 |
| 2012 | In Defence of RANSAC for Outlier Rejection in Deformable Registration
Quoc-Huy Tran, Tat-Jun Chin, Gustavo Carneiro 0001, Michael S. Brown, David Suter |
ECCV (4) | 4 |
| 2012 | Creating Picture Legends for Group PhotosabstractAbstract Group photos are one of the most common types of digital images found in personal image collections and on social networks. One typical post‐processing task for group photos is to produce a key or legend to identify the people in the photo. This is most often done using simple bounding boxes. A more professional approach is to create a picture legend that uses either a full or partial silhouette to identify the individuals. This paper introduces an efficient method for producing picture legends for group photos. Our approach combines face detection with human shape priors into an interactive selection framework to allow users to quickly segment the individuals in a group photo. Our results are better than those obtained by general selection tools and can be produced in a fraction of the time. Junhong Gao, Seon Joo Kim, Michael S. Brown |
Comput. Graph. Forum | 3 |
| 2012 | A New In-Camera Imaging Model for Color Computer Vision and Its ApplicationabstractWe present a study of in-camera image processing through an extensive analysis of more than 10,000 images from over 30 cameras. The goal of this work is to investigate if image values can be transformed to physically meaningful values, and if so, when and how this can be done. From our analysis, we found a major limitation of the imaging model employed in conventional radiometric calibration methods and propose a new in-camera imaging model that fits well with today's cameras. With the new model, we present associated calibration procedures that allow us to convert sRGB images back to their original CCD RAW responses in a manner that is significantly more accurate than any existing methods. Additionally, we show how this new imaging model can be used to build an image correction application that converts an sRGB input image captured with the wrong camera settings to an sRGB output image that would have been recorded under the correct settings of a specific camera. Seon Joo Kim, Hai Ting Lin, Zheng Lu 0002, Sabine Süsstrunk, Stephen Lin 0001, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2011 | Constructing image panoramas using dual-homography warpingabstractThis paper describes a method to construct seamless image mosaics of a panoramic scene containing two predominate planes: a distant back plane and a ground plane that sweeps out from the camera's location. While this type of panorama can be stitched when the camera is carefully rotated about its optical center, such ideal scene capture is hard to perform correctly. Existing techniques use a single homography per image to perform alignment followed by seam cutting or image blending to hide inevitable alignments artifacts. In this paper, we demonstrate how to use two homographies per image to produce a more seamless image. Specifically, our approach blends the homographies in the alignment procedure to perform a nonlinear warping. Once the images are geometrically stitched, they are further processed to blend seams and reduce curvilinear visual artifacts due to the nonlinear warping. As demonstrated in our paper, our procedure is able to produce results for this type of scene where current state-of-the-art techniques fail. Junhong Gao, Seon Joo Kim, Michael S. Brown |
CVPR | 3 |
| 2011 | Revisiting radiometric calibration for color computer visionabstractWe present a study of radiometric calibration and the in-camera imaging process through an extensive analysis of more than 10,000 images from over 30 cameras. The goal is to investigate if image values can be transformed to physically meaningful values and if so, when and how this can be done. From our analysis, we show that the conventional radiometric model fits well for image pixels with low color saturation but begins to degrade as color saturation level increases. This is due to the color mapping process which includes gamut mapping in the in-camera processing that cannot be modeled with conventional methods. To this end, we introduce a new imaging model for radiometric calibration and present an effective calibration scheme that allows us to compensate for the nonlinear color correction to convert non-linear sRGB images to CCD RAW responses. Hai Ting Lin, Seon Joo Kim, Sabine Süsstrunk, Michael S. Brown |
ICCV | 4 |
| 2011 | High quality depth map upsampling for 3D-TOF camerasabstractThis paper describes an application framework to perform high quality upsampling on depth maps captured from a low-resolution and noisy 3D time-of-flight (3D-ToF) camera that has been coupled with a high-resolution RGB camera. Our framework is inspired by recent work that uses nonlocal means filtering to regularize depth maps in order to maintain fine detail and structure. Our framework extends this regularization with an additional edge weighting scheme based on several image features based on the additional high-resolution RGB input. Quantitative and qualitative results show that our method outperforms existing approaches for 3D-ToF upsampling. We describe the complete process for this system, including device calibration, scene warping for input alignment, and even how the results can be further processed using simple user markup. Jaesik Park, Hyeongwoo Kim, Yu-Wing Tai, Michael S. Brown, In-So Kweon |
ICCV | 4 |
| 2011 | Motion Regularization for Matting Motion Blurred ObjectsabstractThis paper addresses the problem of matting motion blurred objects from a single image. Existing single image matting methods are designed to extract static objects that have fractional pixel occupancy. This arises because the physical scene object has a finer resolution than the discrete image pixel and therefore only occupies a fraction of the pixel. For a motion blurred object, however, fractional pixel occupancy is attributed to the object’s motion over the exposure period. While conventional matting techniques can be used to matte motion blurred objects, they are not formulated in a manner that considers the object’s motion and tend to work only when the object is on a homogeneous background. We show how to obtain better alpha mattes by introducing a regularization term in the matting formulation to account for the object’s motion. In addition, we outline a method for estimating local object motion based on local gradient statistics from the original image. For the sake of completeness, we also discuss how user markup can be used to denote the local direction in lieu of motion estimation. Improvements to alpha mattes computed with our regularization are demonstrated on a variety of examples. Hai Ting Lin, Yu-Wing Tai, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Richardson-Lucy Deblurring for Scenes under a Projective Motion PathabstractThis paper addresses how to model and correct image blur that arises when a camera undergoes ego motion while observing a distant scene. In particular, we discuss how the blurred image can be modeled as an integration of the clear scene under a sequence of planar projective transformations (i.e., homographies) that describe the camera's path. This projective motion path blur model is more effective at modeling the spatially varying motion blur exhibited by ego motion than conventional methods based on space-invariant blur kernels. To correct the blurred image, we describe how to modify the Richardson-Lucy (RL) algorithm to incorporate this new blur model. In addition, we show that our projective motion RL algorithm can incorporate state-of-the-art regularization priors to improve the deblurred results. The projective motion path blur model, along with the modified RL algorithm, is detailed, together with experimental results demonstrating its overall effectiveness. Statistical analysis on the algorithm's convergence properties and robustness to noise is also provided. Yu-Wing Tai, Ping Tan 0002, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Visual enhancement of old documents with hyperspectral imaging
Seon Joo Kim, Fanbo Deng, Michael S. Brown |
Pattern Recognit. | 3 |
| 2011 | Matting and compositing of transparent and refractive objectsabstractThis article introduces a new approach for matting and compositing transparent and refractive objects in photographs. The key to our work is an image-based matting model, termed the Attenuation-Refraction Matte (ARM), that encodes plausible refractive properties of a transparent object along with its observed specularities and transmissive properties. We show that an object's ARM can be extracted directly from a photograph using simple user markup. Once extracted, the ARM is used to paste the object onto a new background with a variety of effects, including compound compositing, Fresnel effect, scene depth, and even caustic shadows. User studies find our results favorable to those obtained with Photoshop as well as perceptually valid in most cases. Our approach allows photo editing of transparent and refractive objects in a manner that produces realistic effects previously only possible via 3D models or environment matting. Sai-Kit Yeung, Chi-Keung Tang, Michael S. Brown, Sing Bing Kang |
ACM Trans. Graph. | 3 |
| 2010 | Ink-bleed reduction using functional minimizationabstractInk-bleed interference is undesirable as it reduces the legibility and aesthetics of affected documents. We present a novel approach to reduce ink-bleed interference using functional minimization. In particular, we show how to modify the Chan-Vese active contour model to incorporate information from the front and back sides of the ink-bleed document. This contour model is particularly useful as it does not require edge extraction or explicit thresholding of the document. In addition, we show how functional minimization can again be used to restore broken foreground strokes that arise when strong ink-bleed overlaps with foreground strokes. The experimental results show that our functional minimization method produces better results than recent ink-bleed reduction techniques. To provide a complete framework, we also show how simple user assistance can be further exploited to improve the results. Grani Adiwena Hanasusanto, Michael S. Brown |
CVPR | 3 |
| 2010 | A framework for ultra high resolution 3D imagingabstractWe present an imaging framework to acquire 3D surface scans at ultra high-resolutions (exceeding 600 samples per mm2). Our approach couples a standard structured-light setup and photometric stereo using a large-format ultra-high-resolution camera. While previous approaches have employed similar hybrid imaging systems to fuse positional data with surface normals, what is unique to our approach is the significant asymmetry in the resolution between the low-resolution geometry and the ultra-high-resolution surface normals. To deal with these resolution differences, we propose a multi-resolution surface reconstruction scheme that propagates the low-resolution geometric constraints through the different frequency bands while gradually fusing in the high-resolution photometric stereo data. In addition, to deal with the ultra-high-resolution images, our surface reconstruction is performed in a patch-wise fashion and additional boundary constraints are used to ensure patch coherence. Based on this multi-resolution reconstruction scheme, our imaging framework can produce 3D scans that show exceptionally detailed 3D surfaces far exceeding existing technologies. Zheng Lu 0002, Yu-Wing Tai, Moshe Ben-Ezra, Michael S. Brown |
CVPR | 4 |
| 2010 | Super resolution using edge prior and single image detail synthesisabstractEdge-directed image super resolution (SR) focuses on ways to remove edge artifacts in upsampled images. Under large magnification, however, textured regions become blurred and appear homogenous, resulting in a super-resolution image that looks unnatural. Alternatively, learning-based SR approaches use a large database of exemplar images for “hallucinating” detail. The quality of the upsampled image, especially about edges, is dependent on the suitability of the training images. This paper aims to combine the benefits of edge-directed SR with those of learning-based SR. In particular, we propose an approach to extend edge-directed super-resolution to include detail from an image/texture example provided by the user (e.g., from the Internet). A significant benefit of our approach is that only a single exemplar image is required to supply the missing detail - strong edges are obtained in the SR image even if they are not present in the example image due to the combination of the edge-directed approach. In addition, we can achieve quality results at very large magnification, which is often problematic for both edge-directed and learning-based approaches. Yu-Wing Tai, Shuaicheng Liu, Michael S. Brown, Stephen Lin 0001 |
CVPR | 3 |
| 2010 | Colorization for Single Image Super Resolution
Shuaicheng Liu, Michael S. Brown, Seon Joo Kim, Yu-Wing Tai |
ECCV (6) | 2 |
| 2010 | Correction of Spatially Varying Image and Video Motion Blur Using a Hybrid CameraabstractWe describe a novel approach to reduce spatially varying motion blur in video and images using a hybrid camera system. A hybrid camera is a standard video camera that is coupled with an auxiliary low-resolution camera sharing the same optical path but capturing at a significantly higher frame rate. The auxiliary video is temporally sharper but at a lower resolution, while the lower frame-rate video has higher spatial resolution but is susceptible to motion blur. Our deblurring approach uses the data from these two video streams to reduce spatially varying motion blur in the high-resolution camera with a technique that combines both deconvolution and super-resolution. Our algorithm also incorporates a refinement of the spatially varying blur kernels to further improve results. Our approach can reduce motion blur from the high-resolution video as well as estimate new high-resolution frames at a higher frame rate. Experimental results on a variety of inputs demonstrate notable improvement over current state-of-the-art methods in image/video deblurring. Yu-Wing Tai, Hao Du 0004, Michael S. Brown, Stephen Lin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | User-Assisted Ink-Bleed ReductionabstractThis paper presents a novel user-assisted approach to reduce ink-bleed interference found in old manuscripts. The problem is addressed by first having the user provide simple examples of foreground ink, ink-bleed, and the manuscript's background. From this small amount of user-labeled data, likelihoods of each pixel being foreground, ink-bleed, or background are computed and used as the data costs of a dual-layer Markov random field (MRF) that simultaneously labels all pixels in both the front and back sides of the manuscript. This user-assisted approach produces better results than existing algorithms without the need for extensive parameter tuning or prior assumptions about the ink-bleed intensity characteristics. Our overall application framework is discussed along with details of the features used in the data costs, a comparison between K-nearest neighbor and support vector machine for likelihood estimation, the dual-layer MRF formulation with associated inter- and intra-layer costs, and a comparison of our approach against other ink-bleed reduction algorithms. Michael S. Brown |
IEEE Trans. Image Process. | 2 |
| 2010 | Interactive Visualization of Hyperspectral Images of Historical DocumentsabstractThis paper presents an interactive visualization tool to study and analyze hyperspectral images (HSI) of historical documents. This work is part of a collaborative effort with the Nationaal Archief of the Netherlands (NAN) and Art Innovation, a manufacturer of hyperspectral imaging hardware designed for old and fragile documents. The NAN is actively capturing HSI of historical documents for use in a variety of tasks related to the analysis and management of archival collections, from ink and paper analysis to monitoring the effects of environmental aging. To assist their work, we have developed a comprehensive visualization tool that offers an assortment of visualization and analysis methods, including interactive spectral selection, spectral similarity analysis, time-varying data analysis and visualization, and selective spectral band fusion. This paper describes our visualization software and how it is used to facilitate the tasks needed by our collaborators. Evaluation feedback from our collaborators on how this tool benefits their work is included. Seon Joo Kim, Shaojie Zhuo, Fanbo Deng, Chi-Wing Fu, Michael S. Brown |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2010 | Globally Optimized Linear Windowed Tone MappingabstractThis paper introduces a new tone mapping operator that performs local linear adjustments on small overlapping windows over the entire input image. While each window applies a local linear adjustment that preserves the monotonicity of the radiance values, the problem is implicitly cast as one of global optimization that satisfies the local constraints defined on each of the overlapping windows. Local constraints take the form of a guidance map that can be used to effectively suppress local high contrast while preserving details. Using this method, image structures can be preserved even in challenging high dynamic range (HDR) images that contain either abrupt radiance change, or relatively smooth but salient transitions. Another benefit of our formulation is that it can be used to synthesize HDR images from low dynamic range (LDR) images. Qi Shan, Jiaya Jia, Michael S. Brown |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2009 | Directed assistance for ink-bleed reduction in old documentsabstractInk-bleed interference is a serious problem that affects the legibility of old documents. Ink-bleed can be reduced using pixel classification based on user-supplied markup that labels examples of ink-bleed, foreground-ink, and background. The main challenge is ensuring that the user's markup sufficiently captures the characteristics of the document. This is particularly troublesome for old documents that can exhibit significant change within the same page. In this paper, we address this markup problem using a “directed assistance” approach in which the user provides a small amount of initial markup. The image is then classified and regions with low classification confidence are grouped and displayed to the user for another round of markup. The key idea is to direct the user to where markup is needed. In addition, local markup can be weighted in the classification algorithm to produce better results. Zheng Lu 0002, Michael S. Brown |
CVPR | 3 |
| 2009 | A Fully Automatic System for Restoration of Historical Document Images
Michael S. Brown, Chew Lim Tan |
IAAI | 2 |
| 2009 | Automatic Corresponding Control Points Selection for Historical Document Image RegistrationabstractImage registration is crucial for various image analysis tasks. In particular, most approaches to correction of bleed-through distortion on handwritten document images require the recto image and the verso image to be precisely registered. In this paper, we present a fully automatic method which detects specific number of corresponding control points from historical documents for the purpose of registration. First, candidate points are located by inspecting the gradient direction maps of document images. Corresponding control points are selected based on a dissimilarity metric that incorporates image intensity, gradient magnitude, gradient orientation and displacement. To improve the quality of the detected control points, median filers and consistency checking are applied to correct mismatches. Experiments on real historical document images have shown encouraging results and further improvements can be made by exploiting more sophisticated similarity metric tailored to historical documents' characteristics. Michael S. Brown, Chew Lim Tan |
ICDAR | 2 |
| 2009 | Single image defocus map estimation using local contrast priorabstractImage defocus estimation is useful for several applications including deblurring, blur magnification, measuring image quality, and depth of field segmentation. In this paper, we present a simple yet effective approach for estimating a defocus blur map based on the relationship of the contrast to the image gradient in a local image region. We call this relationship the local contrast prior. The advantage of our approach is that it does not require filter banks or frequency decomposition of the input image; instead we only need to compare local gradient profiles with the local contrast. We discuss the idea behind the local contrast prior and demonstrate its effectiveness on a variety of experiments. Yu-Wing Tai, Michael S. Brown |
ICIP | 2 |
| 2009 | Interactive degraded document binarization: An example (and case) for interactive computer visionabstractThis paper describes a user-assisted application to perform adaptive thresholding (i.e. binarization) on degraded handwritten documents. While existing adaptive thresholding techniques purport to be automatic, they in fact require the user to perform non-intuitive parameter tuning to obtain satisfactory results. In our work, we recast the problem into one where the user needs only to coarsely markup regions in the thresholded image that have unsatisfactory results. These regions are then segmented and processed locally - no parameter tuning is necessary. Our user study shows that not only do the majority of users prefer our application over parameter tuning, but our final results are better than existing algorithms due to the more targeted processing. While our main contribution is an effective user-assisted application for document binarization, we use this as an example to advocate the need to rethink how many computer vision solutions, notoriously reliant on parameter tuning, can be reworked to exploit meaningful user interaction. Zheng Lu 0002, Michael S. Brown |
WACV | 3 |
| 2009 | A unified framework for document restoration using inpainting and shape-from-shading
Li Zhang 0005, Andy M. Yip, Michael S. Brown, Chew Lim Tan |
Pattern Recognit. | 3 |
| 2008 | A framework for reducing ink-bleed in old documentsabstractWe describe a novel application framework to reduce the effects of ink-bleed in old documents. This task is treated as a classification problem where training-data is used to compute per-pixel likelihoods for use in a dual-layer Markov Random Field (MRF) that simultaneously labels image pixels of the front and back of a document as either foreground, background, or ink-bleed, while maintaining the integrity of foreground strokes. Our approach obtains better results than previous work without the need for assumptions about ink-bleed intensities or extensive parameter tuning. Our overall framework is detailed, including front and back image alignment, training-data collection, and the MRF formulation with associated likelihoods and intra- and interlayer cost computations. Yi Huang 0001, Michael S. Brown |
CVPR | 2 |
| 2008 | Image/video deblurring using a hybrid cameraabstractWe propose a novel approach to reduce spatially varying motion blur using a hybrid camera system that simultaneously captures high-resolution video at a low-frame rate together with low-resolution video at a high-frame rate. Our work is inspired by Ben-Ezra and Nayar who introduced the hybrid camera idea for correcting global motion blur for a single still image. We broaden the scope of the problem to address spatially varying blur as well as video imagery. We also reformulate the correction process to use more information available in the hybrid camera system, as well as iteratively refine spatially varying motion extracted from the low-resolution high-speed camera. We demonstrate that our approach achieves superior results over existing work and can be extended to deblurring of moving objects. Yu-Wing Tai, Hao Du 0004, Michael S. Brown, Stephen Lin 0001 |
CVPR | 3 |
| 2008 | Accurate Alignment of Double-Sided Manuscripts for Bleed-Through RemovalabstractDouble-sided manuscripts are often degraded by bleed-through interference. Such degradation must be corrected to facilitate human perception and machine recognition. Most approaches to bleed-through removal rely on perfect alignment between the recto and verso images of a document. This paper presents a two-stage hierarchical alignment technique that can efficiently and accurately align the two sides of a document. Our approach first coarsely aligns the two images using a pair of anchors extracted from the recto and verso images respectively. The coarsely aligned images are then precisely aligned using block matching and radial basis function (RBF) based interpolation techniques. To evaluate the proposed alignment technique, we build a classification and recovery system to remove bleed-through interference and restore historical manuscripts. The accuracy of our alignment approach is then assessed with the accuracy of bleed-through correction. Michael S. Brown, Chew Lim Tan |
Document Analysis Systems | 2 |
| 2008 | Texture amendment: reducing texture distortion in constrained parameterizationabstractConstrained parameterization is an effective way to establish texture coordinates between a 3D surface and an existing image or photograph. A known drawback to constrained parameterization is visual distortion that arises when the 3D geometry is mismatched to highly textured image regions. This paper introduces an approach to reduce visual distortion by expanding image regions via texture synthesis to better fit the 3D geometry. The result is a new amended texture that maintains the essence of the input texture image but exhibits significantly less distortion when mapped onto the 3D model. Yu-Wing Tai, Michael S. Brown, Chi-Keung Tang, Harry Shum |
ACM Trans. Graph. | 2 |
| 2007 | Robust Estimation of Texture Flow via Dense Feature SamplingabstractTexture flow estimation is a valuable step in a variety of vision related tasks, including texture analysis, image segmentation, shape-from-texture and texture remapping. This paper describes a novel and effective technique to estimate texture flow in an image given a small example patch. The key idea consists of extracting a dense set of features from the example patch where discrete orientations are encapsulated into the feature vector such that rotation can be simulated as a linear shift of the vector. This dense feature space is then compressed by PCA and clustered using EM to produce a set of small set of principal features. Obtaining these principal features at varying image scales, we can compute the per-pixel scale and orientation likelihoods for the distorted texture. The final texture flow estimation is formulated as the MAP solution of a labeling Markov network which is solved using belief propagation. Experimental results on both synthetic and real images demonstrate good results even for highly distorted examples. Yu-Wing Tai, Michael S. Brown, Chi-Keung Tang |
CVPR | 2 |
| 2007 | Multi-View Document Rectification using BoundaryabstractWe present a novel technique that uses multiple images of bound and folded documents to rectify the imaged content such that it appears flat and photometrically uniform. Our approach works from a sparse set of uncalibrated views of the document which are mapped to a canonical coordinate frame using the document's boundary. A composite image is constructed from these canonical views that significantly reduces the effects of depth distortion without the blurring artifacts that is problematic in single image approaches. In addition, we propose a new technique to estimate illumination variation in the individual images allowing the final composited content to be photometrically rectified. Our approach is straight-forward, robust, and produces good results. Yau-Chat Tsoi, Michael S. Brown |
CVPR | 2 |
| 2007 | Example-Based Cosmetic TransferabstractCosmetic makeup is used worldwide as a means to enhance beauty and express moods. An art form in its own right, cosmetic styles continuously change and evolve to reflect cultural and societal trends. While countless magazines and books are dedicated to demonstrating cosmetic art, the actual application of makeup still remains a physical endeavor. In this paper, we describe a procedure to apply cosmetic makeup to the image of a person's face with the click of a mouse. Our approach works from before- and-after example images created by professional makeup artists. Using our "cosmetic-transfer" procedure, we can realistically transfer the cosmetic style captured in the example-pair to another person's face. This greatly reduces the time and effort needed to demonstrate a cosmetic style on a new person's face. In addition, our approach can be used to mix-and- match, and even fine-tune, example styles, all virtually, without the need for any physical makeup. Wai-Shun Tong, Chi-Keung Tang, Michael S. Brown, Ying-Qing Xu |
PG | 3 |
| 2007 | Restoring 2D Content from Distorted DocumentsabstractThis paper presents a framework to restore the 2D content printed on documents in the presence of geometric distortion and non-uniform illumination. Compared with textbased document imaging approaches that correct distortion to a level necessary to obtain sufficiently readable text or to facilitate optical character recognition (OCR), our work targets nontextual documents where the original printed content is desired. To achieve this goal, our framework acquires a 3D scan of the document's surface together with a high-resolution image. Conformal mapping is used to rectify geometric distortion by mapping the 3D surface back to a plane while minimizing angular distortion. This conformal "deskewing" assumes no parametric model of the document's surface and is suitable for arbitrary distortions. Illumination correction is performed by using the 3D shape to distinguish content gradient edges from illumination gradient edges in the high-resolution image. Integration is performed using only the content edges to obtain a reflectance image with significantly less illumination artifacts. This approach makes no assumptions about light sources and their positions. The results from the geometric and photometric correction are combined to produce the final output. Michael S. Brown, Mingxuan Sun 0001, Ruigang Yang, Yun Lin 0011, W. Brent Seales |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Natural shadow mattingabstractThis article addresses the problem of natural shadow matting , the removal or extraction of natural shadows from a single image. Because textures are maintained in the shadowless image after the extraction process, our approach produces some of the best results to date among shadow removal techniques. Using the image formation equation typical of computer vision, we advocate a new model for shadow formation where shadow effect is understood as light attenuation instead of a mixture of two colors governed by the conventional matting equation. This leads to a new shadow equation with fewer unknowns to solve, where a three-channel shadow matte and a shadowless image are considered in our optimization. Our problem is formulated as one of energy minimization guided by user-supplied hints in the form of a quadmap which can be specified easily by the user. This formulation allows for robust shadow matte extraction while maintaining texture in the shadowed region by considering color transfer, texture gradient, and shadow smoothness. We demonstrate the usefulness of our approach in shadow removal, image matting, and compositing. Tai-Pang Wu, Chi-Keung Tang, Michael S. Brown, Harry Shum |
ACM Trans. Graph. | 3 |
| 2007 | ShapePalettes: interactive normal transfer via sketchingabstractWe present a simple interactive approach to specify 3D shape in a single view using "shape palettes". The interaction is as follows: draw a simple 2D primitive in the 2D view and then specify its 3D orientation by drawing a corresponding primitive on ashape palette. The shape palette is presented as an image of some familiar shape whose local 3D orientation is readily understood and can be easily marked over. The 3D orientation from the shape palette is transferred to the 2D primitive based on the markup. As we will demonstrate, only sparse markup is needed to generate expressive and detailed 3D surfaces. This markup approach can be used to model freehand 3D surfaces drawn in a single view, or combined with image-snapping tools to quickly extract surfaces from images and photographs. Tai-Pang Wu, Chi-Keung Tang, Michael S. Brown, Harry Shum |
ACM Trans. Graph. | 3 |
| 2006 | Image Pre-Conditioning for Out-of-Focus Projector BlurabstractWe present a technique to reduce image blur caused by out-of-focus regions in projected imagery. Unlike traditional restoration algorithms that operate on a blurred image to recover the original, the nature of our problem requires that the correction be applied to the original image before blurring. To accomplish this, a camera is used to estimate a series of spatially varying point-spread-functions (PSF) across the projector’s image. These discrete PSFs are then used to guide a pre-processing algorithm based on Wiener filtering to condition the image before projection. Results show that using this technique can help ameliorate the visual effects from out-of-focus projector blur. Michael S. Brown, Peng Song 0010, Tat-Jen Cham |
CVPR (2) | 1 |
| 2006 | Geometric and shading correction for images of printed materials using boundaryabstractA novel technique that uses boundary interpolation to correct geometric distortion and shading artifacts present in images of printed materials is presented. Unlike existing techniques, our algorithm can simultaneously correct a variety of geometric distortions, including skew, fold distortion, binder curl, and combinations of these. In addition, the same interpolation framework can be used to estimate the intrinsic illumination component of the distorted image to correct shading artifacts. We detail our algorithm for geometric and shading correction and demonstrate its usefulness on real-world and synthetic data. Michael S. Brown, Yau-Chat Tsoi |
IEEE Trans. Image Process. | 1 |
| 2005 | Conformal Deskewing of Non-Planar DocumentsabstractThis paper presents an approach that uses conformal mapping to parameterize a document's 3D shape to a 2D plane. Using this conformal parametrization, a restorative mapping between an image of the distorted document and a "flattened" representation of the document can be computed and used to deskew the image. Our experiments show that arbitrarily distorted documents can be restored to within a single pixel of their true planar format. In addition, surface points can be constrained to map to specified locations in the restored 2D plane. Michael S. Brown, Charles J. Pisula |
CVPR (1) | 1 |
| 2005 | Geometric and Photometric Restoration of Distorted DocumentsabstractWe present a system to restore the 2D content printed on distorted documents. Our system works by acquiring a 3D scan of the document's surface together with a high-resolution image. Using the 3D surface information and the 2D image, we can ameliorate unwanted surface distortion and effects from non-uniform illumination. Our system can process arbitrary geometric distortions, not requiring any pre-assumed parametric models for the document's geometry. The illumination correction uses the 3D shape to distinguish content edges from illumination edges to recover the 2D content's reflectance image while making no assumptions about light sources and their positions. Results are shown for real objects, demonstrating a complete framework capable of restoring geometric and photometric artifacts on distorted documents Mingxuan Sun 0001, Ruigang Yang, Yun Lin 0011, George V. Landon, W. Brent Seales, Michael S. Brown |
ICCV | 6 |
| 2005 | Camera-Based Calibration Techniques for Seamless Multiprojector DisplaysabstractMultiprojector, large-scale displays are used in scientific visualization, virtual reality, and other visually intensive applications. In recent years, a number of camera-based computer vision techniques have been proposed to register the geometry and color of tiled projection-based display. These automated techniques use cameras to "calibrate" display geometry and photometry, computing per-projector corrective warps and intensity corrections that are necessary to produce seamless imagery across projector mosaics. These techniques replace the traditional labor-intensive manual alignment and maintenance steps, making such displays cost-effective, flexible, and accessible. In this paper, we present a survey of different camera-based geometric and photometric registration techniques reported in the literature to date. We discuss several techniques that have been proposed and demonstrated, each addressing particular display configurations and modes of operation. We overview each of these approaches and discuss their advantages and disadvantages. We examine techniques that address registration on both planar (video walls) and arbitrary display surfaces and photometric correction for different kinds of display surfaces. We conclude with a discussion of the remaining challenges and research opportunities for multiprojector displays. Michael S. Brown, Aditi Majumder, Ruigang Yang |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2004 | Geometric and Shading Correction for Images of Printed Materials: A Unified Approach Using Boundary
Yau-Chat Tsoi, Michael S. Brown |
CVPR (1) | 2 |
| 2004 | Emulating short MPEG GOPs with less I-framesabstractFrequent placement of intra-encoded pictures (I-frames) in MPEG video facilitates (1) error resilience over lossy network transmission and (2) random access for VCR functionality. However, high I-frame frequency sacrifices quality-to-bitrate efficiency that can be gained by using longer sequences of inter-encoded pictures. We present a simple strategy that emulates frequent I-frame encoding (short GOPs) while using fewer I-frames. Our approach maintains a small set of previously encoded/decoded I-frames that can be re-used to start future GOPs. A least-recently-used (LRU) policy is used to maintain this set of I-frames. We demonstrate that this strategy allows gains in PSNR for constant bit rate encoding while providing the benefits of frequent I-frame placement. RuiDuo Yang, Michael S. Brown |
ICME | 2 |
| 2004 | Music database query with video by synesthesia observationabstractWe propose a novel framework to query a music database by supplying an MPEG video sequence. The music and video are treated as time series data and the retrieval query is based on the similarity between corresponding features in the music and video sequence (e.g. the tempo in the music and the motion in the video). A similarity measure for matching these features is proposed to capture the synesthesia effect in music-video. RuiDuo Yang, Michael S. Brown |
ICME | 2 |
| 2004 | Decoder motion vector estimation for scalable video error concealmentabstractWe propose a technique to recover lost enhancement layer information in scalable video using the information from: (1) the current base layer; and (2) the previous base- and enhancement-layer together with a decoder motion vector estimation method. An average of 1 dB improvement is reported for the enhancement layer vs. existing techniques. An accelerated method is also demonstrated to make the concealment real-time. RuiDuo Yang, Michael S. Brown |
ICME | 2 |
| 2004 | Image Restoration of Arbitrarily Warped DocumentsabstractWe present a framework for acquiring and restoring images of warped documents. The purpose of our restoration is to create a planar representation of a once planar document that has undergone an arbitrary and unknown rigid deformation. To accomplish this restoration, our framework acquires and flattens the 3D shape of a warped document to determine a nonlinear image transform that can correct for image distortion caused by the document's shape. Our framework is designed for use in library and museum digitization efforts where very old and badly damaged manuscripts are imaged. Michael S. Brown, W. Brent Seales |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2003 | Fast halftoning of MPEG video for bi-tonal displaysabstractGreater network connectivity allows multimedia content to be delivered to a wide range of end-hosts with various compute and display capabilities. In many cases the format of the delivered multimedia must be processed to adapt it for use on the end-host. For low-cost devices with limited compute and display power, this adaptation must be within the capability of the device. In this paper we address one such adaptation - that of halftoning MPEG video for viewing on low-cost devices with bi-tonal display. In particular, we demonstrate a halftoning strategy that is fast and simple, requiring only integer math, bit shifts, and a look-up table. In addition, our approach can be used with the SNR scalability feature in MPEG-4 to allow efficient bandwidth consumption while providing fast bi-tonal display adaptation. Michael S. Brown, RuiDuo Yang |
ICIP (1) | 1 |
| 2002 | A Practical and Flexible Tiled Display SystemabstractComputer graphics and high-resolution digital imagery are becoming increasingly pervasive in communities not traditionally associated with graphics. Commodity graphics cards and digital cameras, along with powerful modeling software, allow organizations such as libraries, museums, and small businesses to produce expressive and realistic computer generated imagery, far surpassing the capabilities of current desktop monitors. What remains elusive to these groups is access to affordable and easy to use large format display technology. We present a practical system for deploying flexible projector based tiled displays. Our framework integrates two key components: (1) self-calibrating display geometry with real-time geometric correction and (2) PC-based distributed rendering that supports an established graphics API. Our system's display geometry is easy to configure and reconfigure, accommodates casually tiled projectors and arbitrary display surfaces, and can be operational in a matter of minutes. In addition, the underlying distributed rendering architecture (WireGL) is transparent to existing OpenGL applications, requiring no custom APIs or re-compilation of existing OpenGL executables. In short, we present a practical and flexible low-cost tiled display system that is simple to deploy and easy to operate. Michael S. Brown, W. Brent Seales |
PG | 1 |
| 2001 | Document Restoration Using 3D Shape: A General Deskewing Algorithm for Arbitrarily Warped DocumentsabstractWe present a framework for restoring arbitrarily warped and deformed documents to their original planar shape. The impetus for this work is the need for tools and techniques to help digitally preserve and restore fragile manuscripts. Current digitization is performed under the assumption that the documents are flat, with subsequent image-processing and restoration algorithms either relying on this assumption or attempting to overcome it without shape information. Although most manuscripts were originally flat, many become deformed from damage and deterioration. Physical flattening is not possible without risking further, possibly irreversible, damage. Our framework addresses this restoration problem with two primary contributions. First, we present a working 3D digitization setup that acquires a 3D model with accurate shape-to-texture registration under multiple lighting conditions. Second, we show how the 3D model and a mass-spring particle system can be used together as a framework for digital flattening. We show that this restoration process can correct document deformations and can significantly improve subsequent document analysis. Michael S. Brown, W. Brent Seales |
ICCV | 1 |
| 2001 | Dynamic Shadow Removal from Front Projection DisplaysabstractFront-projection display environments suffer from a fundamental problem: users and other objects in the environment can easily and inadvertently block projectors, creating shadows on the displayed image. We introduce a technique that detects and corrects transient shadows in a multi-projector display. Our approach is to minimize the difference between predicted (generated) and observed (camera) images by continuous modification of the projected image values for each display device. We speculate that the general predictive monitoring framework introduced here is capable of addressing more general radiometric consistency problems. Using an automatically-derived relative position of cameras and projectors in the display environment and a straightforward color correction scheme, the system renders an expected image for each camera location. Cameras observe the displayed image, which is compared with the expected image to detect shadowed regions. These regions are transformed to the appropriate projector frames, where corresponding pixel values are increased. In display regions where more than one projector contributes to the image, shadow regions are eliminated. We demonstrate an implementation of the technique in a multiprojector system. Christopher O. Jaynes, Stephen B. Webb, R. Matt Steele, Michael S. Brown, W. Brent Seales |
IEEE Visualization | 4 |
| 2001 | PixelFlex: A Reconfigurable Multi-Projector Display SystemabstractThis paper presents PixelFlex - a spatially reconfigurable multi-projector display system. The PixelFlex system is composed of ceiling-mounted projectors, each with computer-controlled pan, tilt, zoom and focus; and a camera for closed-loop calibration. Working collectively, these controllable projectors function as a single logical display capable of being easily modified into a variety of spatial formats of differing pixel density, size and shape. New layouts are automatically calibrated within minutes to generate the accurate warping and blending functions needed to produce seamless imagery across planar display surfaces, thus giving the user the flexibility to quickly create, save and restore multiple screen configurations. Overall, PixelFlex provides a new level of automatic reconfigurability and usage, departing from the static, one-size-fits-all design of traditional large-format displays. As a front-projection system, PixelFlex can be installed in most environments with space constraints and requires little or no post-installation mechanical maintenance because of the closed-loop calibration. Ruigang Yang, David Gotz, Justin Hensley, Herman Towles, Michael S. Brown |
IEEE Visualization | 5 |
| 1999 | Geometrically correct imagery for teleconferencingabstractCurrent camera-monitor teleconferencing applications produce unrealistic imagery and break any sense of presence for the participants. Other capture/display technologies can be used to provide more compelling teleconferencing. However, complex geometries in capture/display systems make producing geometrically correct imagery difficult. It is usually impractical to detect, model and compensate for all effects introduced by the capture/display system. Most applications simply ignore these issues and rely on the user acceptance of the camera-monitor paradigm. Ruigang Yang, Michael S. Brown, W. Brent Seales, Henry Fuchs |
ACM Multimedia (1) | 2 |
| 1999 | Multi-Projector Displays Using Camera-Based RegistrationabstractConventional projector-based display systems are typically designed around precise and regular configurations of projectors and display surfaces. While this results in rendering simplicity and speed, it also means painstaking construction and ongoing maintenance. In previously published work, we introduced a vision of projector-based displays constructed from a collection of casually-arranged projectors and display surfaces. In this paper, we present flexible yet practical methods for realizing this vision, enabling low-cost mega-pixel display systems with large physical dimensions, higher resolution, or both. The techniques afford new opportunities to build personal 3D visualization systems in offices, conference rooms, theaters, or even your living room. As a demonstration of the simplicity and effectiveness of the methods that we continue to perfect, we show in the included video that a 10-year old child can construct and calibrate a two-camera, two-projector, head-tracked display system, all in about 15 minutes. Ramesh Raskar, Michael S. Brown, Ruigang Yang, Wei-Chao Chen, Greg Welch, Herman Towles, W. Brent Seales, Henry Fuchs |
IEEE Visualization | 2 |
| 1998 | Fast Stereo Matching in Compressed Video
Michael S. Brown, W. Brent Seales |
ACCV (1) | 1 |