EDBT 2026 Demo / reviewers in the wild / expert
Abhijith Punnappurath
dblp:141/9895
· DBLP profile ↗
20ranked-venue papers
10as first author
11since 2021 · last 2025
0000-0003-3438-5896ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Examining Joint Demosaicing and Denoising for Single-, Quad-, and Nona-Bayer PatternsabstractCamera sensors have color filters arranged in a mosaic layout, traditionally following the Bayer pattern. Demosaicing is a critical step camera hardware applies to obtain a full-channel RGB image. Many smartphones now have multiple sensors with different patterns, such as Quad-Bayer or Nona-Bayer. Most modern deep network-based models perform joint demosaicing and denoising with the strategy of training a separate network per pattern. Relying on individual models per pattern requires additional memory overhead and makes it challenging to switch quickly between cameras. In this work, we are interested in analyzing strategies for joint demosaicing and denoising for the three main mosaic layouts ($1 \times 1$ Single-Bayer, $2 \times 2$ Quad-Bayer, and $3 \times 3$ Nona-Bayer). We found concatenating a three-channel mosaic embedding to the input image and training a unified demosaicing architecture yields results that outperform existing Quad-Bayer and Nona-Bayer models and are comparable to Single-Bayer models. Additionally, we describe a maskout strategy that enhances the model performance and facilitates dead pixel correction-a step often overlooked by existing Al-based demosaicing models. As part of this effort, we captured a new demosaicing dataset of 638 RAW images that contain challenging scenes with patches annotated for training, validation, and testing. Code and data is available at https://github.com/SamsungLabs/unified-demosaicing. SaiKiran Kumar Tedla, Abhijith Punnappurath, Luxi Zhao 0002, Michael S. Brown |
ICCP | 2 |
| 2025 | Time-Aware Auto White Balance in Mobile Photography
Mahmoud Afifi, Luxi Zhao 0002, Abhijith Punnappurath, Mohammed A. Abdelsalam, Michael S. Brown |
ICCV | 3 |
| 2024 | Non-parametric Sensor Noise Modeling and Synthesis
Luxi Zhao 0002, Atin Singh, Jaeduk Han, Abhijith Punnappurath, Marcus A. Brubaker, Jihwan Choe, Michael S. Brown |
ECCV (24) | 5 |
| 2023 | Graphics2RAW: Mapping Computer Graphics Images to Sensor RAW ImagesabstractComputer graphics (CG) rendering platforms produce imagery with ever-increasing photo realism. The narrowing domain gap between real and synthetic imagery makes it possible to use CG images as training data for deep learning models targeting high-level computer vision tasks, such as autonomous driving and semantic segmentation. CG images, however, are currently not suitable for low-level vision tasks targeting RAW sensor images. This is because RAW images are encoded in sensor-specific color spaces and incur pre-white-balance color casts caused by the sensor’s response to scene illumination. CG images are rendered directly to a device-independent perceptual color space without needing white balancing. As a result, it is necessary to apply a mapping procedure to close the domain gap between graphics and RAW images. To this end, we introduce a framework to process graphics images to mimic RAW sensor images accurately. Our approach allows a one-to-many mapping, where a single graphics image can be transformed to match multiple sensors and multiple scene illuminations. In addition, our approach requires only a handful of example RAW-DNG files from the target sensor as parameters for the mapping process. We compare our method to alternative strategies and show that our approach produces more realistic RAW images and provides better results on three low-level vision tasks: RAW denoising, illumination estimation, and neural rendering for night photography. Finally, as part of this work, we provide a dataset of 292 realistic CG images for training low-light imaging models. Donghwan Seo, Abhijith Punnappurath, Luxi Zhao 0002, Abdelrahman Abdelhamed, SaiKiran Kumar Tedla, Jihwan Choe, Michael S. Brown |
ICCV | 2 |
| 2022 | Learning sRGB-to-Raw-RGB De-rendering with Content-Aware MetadataabstractMost camera images are rendered and saved in the standard RGB (sRGB) format by the camera's hardware. Due to the in-camera photo-finishing routines, nonlinear sRGB images are undesirable for computer vision tasks that assume a direct relationship between pixel values and scene radiance. For such applications, linear raw-RGB sensor images are preferred. Saving images in their raw-RGB format is still uncommon due to the large storage requirement and lack of support by many imaging applications. Several “raw reconstruction” methods have been proposed that utilize specialized metadata sampled from the raw-RGB image at capture time and embedded in the sRGB image. This metadata is used to parameterize a mapping function to derender the sRGB image back to its original raw-RGB format when needed. Existing raw reconstruction methods rely on simple sampling strategies and global mapping to perform the de-rendering. This paper shows how to improve the derendering results by jointly learning sampling and reconstruction. Our experiments show that our learned sampling can adapt to the image content to produce better raw reconstructions than existing methods. We also describe an online fine-tuning strategy for the reconstruction network to improve results further. Seonghyeon Nam, Abhijith Punnappurath, Marcus A. Brubaker, Michael S. Brown |
CVPR | 2 |
| 2022 | Day-to-Night Image Synthesis for Training Nighttime Neural ISPsabstractMany flagship smartphone cameras now use a dedicated neural image signal processor (ISP) to render noisy raw sensor images to the final processed output. Training night-mode ISP networks relies on large-scale datasets of image pairs with: (1) a noisy raw image captured with a short exposure and a high ISO gain; and (2) a ground truth low-noise raw image captured with a long exposure and low ISO that has been rendered through the ISP. Capturing such image pairs is tedious and time-consuming, requiring careful setup to ensure alignment between the image pairs. In addition, ground truth images are often prone to motion blur due to the long exposure. To address this problem, we propose a method that synthesizes nighttime images from day-time images. Daytime images are easy to capture, exhibit low-noise (even on smartphone cameras) and rarely suffer from motion blur. We outline a processing framework to convert daytime raw images to have the appearance of realistic nighttime raw images with different levels of noise. Our procedure allows us to easily produce aligned noisy and clean nighttime image pairs. We show the effectiveness of our synthesis framework by training neural ISPs for nightmode rendering. Furthermore, we demonstrate that using our synthetic nighttime images together with small amounts of real data (e.g., 5% to 10%) yields performance almost on par with training exclusively on real nighttime images. Our dataset and code are available at https://github.com/SamsungLabs/day-to-night. Abhijith Punnappurath, Abdullah Abuolaim, Abdelrahman Abdelhamed, Alex Levinshtein, Michael S. Brown |
CVPR | 1 |
| 2022 | Extracting Vignetting and Grain Filter Effects from PhotosabstractMost smartphones support the use of real-time camera filters to impart visual effects to captured images. Currently, such filters come preinstalled on-device or need to be downloaded and installed before use (e.g., Instagram filters). Recent work [24] proposed a method to extract a camera filter directly from an example photo that has already had a filter applied. The work in [24] focused only on the color and tonal aspects of the underlying filter. In this paper, we introduce a method to extract two spatially varying effects commonly used by on-device camera filters—namely, image vignetting and image grain. Specifically, we show how to extract the parameters for vignetting and image grain present in an example image and replicate these effects as an on-device filter. We use lightweight CNNs to estimate the filter parameters and employ efficient techniques—isotropic Gaussian filters and simplex noise—for regenerating the filters. Our design achieves a reasonable trade-off between efficiency and realism. We show that our method can extract vignetting and image grain filters from stylized photos and replicate the filters on captured images more faithfully, as compared to color and style transfer methods. Our method is significantly efficient and has been already deployed to millions of flagship smartphones. Abdelrahman Abdelhamed, Jonghwa Yim, Abhijith Punnappurath, Michael S. Brown, Jihwan Choe |
WACV | 3 |
| 2022 | CIE XYZ Net: Unprocessing Images for Low-Level Computer Vision TasksabstractCameras currently allow access to two image states: (i) a minimally processed linear raw-RGB image state (i.e., raw sensor data); or (ii) a highly-processed nonlinear image state (e.g., sRGB). There are many computer vision tasks that work best with a linear image state, such as image deblurring and image dehazing. Unfortunately, the vast majority of images are saved in the nonlinear image state. Because of this, a number of methods have been proposed to "unprocess" nonlinear images back to a raw-RGB state. However, existing unprocessing methods have a drawback because raw-RGB images are sensor-specific. As a result, it is necessary to know which camera produced the sRGB output and use a method or network tailored for that sensor to properly unprocess it. This paper addresses this limitation by exploiting another camera image state that is not available as an output, but it is available inside the camera pipeline. In particular, cameras apply a colorimetric conversion step to convert the raw-RGB image to a device-independent space based on the CIE XYZ color space before they apply the nonlinear photo-finishing. Leveraging this canonical image state, we propose a deep learning framework, CIE XYZ Net, that can unprocess a nonlinear image back to the canonical CIE XYZ image. This image can then be processed by any low-level computer vision operator and re-rendered back to the nonlinear image. We demonstrate the usefulness of the CIE XYZ Net on several low-level vision tasks and show significant gains that can be obtained by this processing framework. Code and dataset are publicly available at https://github.com/mahmoudnafifi/CIE_XYZ_NET. Mahmoud Afifi, Abdelrahman Abdelhamed, Abdullah Abuolaim, Abhijith Punnappurath, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | A Little Bit More: Bitplane-Wise Bit-Depth RecoveryabstractImaging sensors digitize incoming scene light at a dynamic range of 10-12 bits (i.e., 1024-4096 tonal values). The sensor image is then processed onboard the camera and finally quantized to only 8 bits (i.e., 256 tonal values) to conform to prevailing encoding standards. There are a number of important applications, such as high-bit-depth displays and photo editing, where it is beneficial to recover the lost bit depth. Deep neural networks are effective at this bit-depth reconstruction task. Given the quantized low-bit-depth image as input, existing deep learning methods employ a single-shot approach that attempts to either (1) directly estimate the high-bit-depth image, or (2) directly estimate the residual between the high- and low-bit-depth images. In contrast, we propose a training and inference strategy that recovers the residual image bitplane-by-bitplane. Our bitplane-wise learning framework has the advantage of allowing for multiple levels of supervision during training and is able to obtain state-of-the-art results using a simple network architecture. We test our proposed method extensively on several image datasets and demonstrate an improvement from 0.5dB to 2.3dB PSNR over prior methods depending on the quantization level. Abhijith Punnappurath, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Leveraging the Availability of Two Cameras for Illuminant EstimationabstractMost modern smartphones are now equipped with two rear-facing cameras – a main camera for standard imaging and an additional camera to provide wide-angle or telephoto zoom capabilities. In this paper, we leverage the availability of these two cameras for the task of illumination estimation using a small neural network to perform the illumination prediction. Specifically, if the two cameras’ sensors have different spectral sensitivities, the two images provide different spectral measurements of the physical scene. A linear 3 × 3 color transform that maps between these two observations – and that is unique to a given scene illuminant – can be used to train a lightweight neural network comprising no more than 1460 parameters to predict the scene illumination. We demonstrate that this two-camera approach with a lightweight network provides results on par or better than much more complicated illuminant estimation methods operating on a single image. We validate our method’s effectiveness through extensive experiments on radiometric data, a quasi-real two-camera dataset we generated from an existing single camera dataset, as well as a new real image dataset that we captured using a smartphone with two rear-facing cameras. Abdelrahman Abdelhamed, Abhijith Punnappurath, Michael S. Brown |
CVPR | 2 |
| 2021 | Spatially Aware Metadata for Raw ReconstructionabstractA camera sensor captures a raw-RGB image that is then processed to a standard RGB (sRGB) image through a series of onboard operations performed by the camera's image signal processor (ISP). Among these processing steps, local tone mapping is one of the most important operations used to enhance the overall appearance of the final rendered sRGB image. For certain applications, it is often desirable to de-render or unprocess the sRGB image back to its original raw-RGB values. This "raw reconstruction" is a challenging task because many of the operations performed by the ISP, including local tone mapping, are nonlinear and difficult to invert. Existing raw reconstruction methods that store specialized metadata at capture time to enable raw recovery ignore local tone mapping and assume that a global transformation exists between the raw-RGB and sRGB color spaces. In this work, we advocate a spatially aware metadata-based raw reconstruction method that is robust to local tone mapping, and yields significantly higher raw reconstruction accuracy (6 dB average PSNR improvement) compared to existing raw reconstruction methods. Our method requires only 0.2% samples of the full-sized image as metadata, has negligible computational overhead at capture time, and can be easily integrated into modern ISPs. Abhijith Punnappurath, Michael S. Brown |
WACV | 1 |
| 2020 | Modeling Defocus-Disparity in Dual-Pixel SensorsabstractMost modern consumer cameras use dual-pixel (DP) sensors that provide two sub-aperture views of the scene in a single photo capture. The DP sensor was designed to assist the camera's autofocus routine, which examines local disparity in the two sub-aperture views to determine which parts of the image are out of focus. Recently, these DP views have been used for tasks beyond autofocus, such as synthetic bokeh, reflection removal, and depth reconstruction. These recent methods treat the two DP views as stereo image pairs and apply stereo matching algorithms to compute local disparity. However, dual-pixel disparity is not caused by view parallax as in stereo, but instead is attributed to defocus blur that occurs in out-of-focus regions in the image. This paper proposes a new parametric point spread function to model the defocus-disparity that occurs on DP sensors. We apply our model to the task of depth estimation from DP data. An important feature of our model is its ability to exploit the symmetry property of the DP blur kernels at each pixel. We leverage this symmetry property to formulate an unsupervised loss function that does not require ground truth depth. We demonstrate our method's effectiveness on both DSLR and smartphone DP data. Abhijith Punnappurath, Abdullah Abuolaim, Mahmoud Afifi, Michael S. Brown |
ICCP | 1 |
| 2020 | Learning Raw Image Reconstruction-Aware Deep Image CompressorsabstractDeep learning-based image compressors are actively being explored in an effort to supersede conventional image compression algorithms, such as JPEG. Conventional and deep learning-based compression algorithms focus on minimizing image fidelity errors in the nonlinear standard RGB (sRGB) color space. However, for many computer vision tasks, the sensor's linear raw-RGB image is desirable. Recent work has shown that the original raw-RGB image can be reconstructed using only small amounts of metadata embedded inside the JPEG image [1]. However, [1] relied on the conventional JPEG encoding that is unaware of the raw-RGB reconstruction task. In this paper, we examine the ability of deep image compressors to be "aware" of the additional objective of raw reconstruction. Towards this goal, we describe a general framework that enables deep networks targeting image compression to jointly consider both image fidelity errors and raw reconstruction errors. We describe this approach in two scenarios: (1) the network is trained from scratch using our proposed joint loss, and (2) a network originally trained only for sRGB fidelity loss is later fine-tuned to incorporate our raw reconstruction loss. When compared to sRGB fidelity-only compression, our combined loss leads to appreciable improvements in PSNR of the raw reconstruction with only minor impact on sRGB fidelity as measured by MS-SSIM. Abhijith Punnappurath, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Reflection Removal Using a Dual-Pixel SensorabstractReflection removal is the challenging problem of removing unwanted reflections that occur when imaging a scene that is behind a pane of glass. In this paper, we show that most cameras have an overlooked mechanism that can greatly simplify this task. Specifically, modern DLSR and smartphone cameras use dual pixel (DP) sensors that have two photodiodes per pixel to provide two sub-aperture views of the scene from a single captured image. ``Defocus-disparity'' cues, which are natural by-products of the DP sensor encoded within these two sub-aperture views, can be used to distinguish between image gradients belonging to the in-focus background and those caused by reflection interference. This gradient information can then be incorporated into an optimization framework to recover the background layer with higher accuracy than currently possible from the single captured image. As part of this work, we provide the first image dataset for reflection removal consisting of the sub-aperture views from the DP sensor. Abhijith Punnappurath, Michael S. Brown |
CVPR | 1 |
| 2018 | Revisiting Autofocus for Smartphone Cameras
Abdullah Abuolaim, Abhijith Punnappurath, Michael S. Brown |
ECCV (15) | 2 |
| 2017 | Multi-Image Blind Super-Resolution of 3D ScenesabstractWe address the problem of estimating the latent high-resolution (HR) image of a 3D scene from a set of non-uniformly motion blurred low-resolution (LR) images captured in the burst mode using a hand-held camera. Existing blind super-resolution (SR) techniques that account for motion blur are restricted to fronto-parallel planar scenes. We initially develop an SR motion blur model to explain the image formation process in 3D scenes. We then use this model to solve for the three unknowns-the camera trajectories, the depth map of the scene, and the latent HR image. We first compute the global HR camera motion corresponding to each LR observation from patches lying on a reference depth layer in the input images. Using the estimated trajectories, we compute the latent HR image and the underlying depth map iteratively using an alternating minimization framework. Experiments on synthetic and real data reveal that our proposed method outperforms the state-of-the-art techniques by a significant margin. Abhijith Punnappurath, Thekke Madam Nimisha, A. N. Rajagopalan 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Deep Decoupling of Defocus and Motion Blur for Dynamic Segmentation
Abhijith Punnappurath, Yogesh Balaji, Mahesh Mohan M. R., A. N. Rajagopalan 0001 |
ECCV (7) | 1 |
| 2016 | Rolling shutter super-resolution in burst modeabstractCapturing multiple images using the burst mode of handheld cameras can be a boon to obtain a high resolution (HR) image by exploiting the subpixel motion among the captured images arising from handshake. However, the caveat with mobile phone cameras is that they produce rolling shutter (RS) distortions that must be accounted for in the super-resolution process. We propose a method in which we obtain an RS-free HR image using HR camera trajectory estimated by leveraging the intra- and inter-frame continuity of the camera motion. Experimental evaluations demonstrate that our approach can effectively recover a super-resolved image free from RS artifacts. Vijay Rengarajan, Abhijith Punnappurath, A. N. Rajagopalan 0001, Guna Seetharaman |
ICIP | 2 |
| 2015 | Rolling Shutter Super-ResolutionabstractClassical multi-image super-resolution (SR) algorithms, designed for CCD cameras, assume that the motion among the images is global. But CMOS sensors that have increasingly started to replace their more expensive CCD counterparts in many applications do not respect this assumption if there is a motion of the camera relative to the scene during the exposure duration of an image because of the row-wise acquisition mechanism. In this paper, we study the hitherto unexplored topic of multi-image SR in CMOS cameras. We initially develop an SR observation model that accounts for the row-wise distortions called the "rolling shutter" (RS) effect observed in images captured using non-stationary CMOS cameras. We then propose a unified RS-SR framework to obtain an RS-free high-resolution image (and the row-wise motion) from distorted low-resolution images. We demonstrate the efficacy of the proposed scheme using synthetic data as well as real images captured using a hand-held CMOS camera. Quantitative and qualitative assessments reveal that our method significantly advances the state-of-the-art. Abhijith Punnappurath, Vijay Rengarajan, A. N. Rajagopalan 0001 |
ICCV | 1 |
| 2013 | Registration and occlusion detection in motion blurabstractWe address the problem of automatically detecting occluded regions given a blurred/unblurred image pair of a scene taken from different viewpoints. The occlusion can be due to single or multiple objects. We present a unified framework for detecting occluder(s) that is reasonably robust to non-uniform motion blur as well as variations in camera pose (without the need for deblurring). We assume that the occluded pixels occupy only a relatively small area and that the camera motion trajectory is sparse in the camera motion space. We validate the performance of our algorithm with experiments on synthetic and real data. Abhijith Punnappurath, A. N. Rajagopalan 0001, Guna Seetharaman |
ICIP | 1 |