VLDB 2026 Research / reviewers in the wild / expert
Mahmoud Afifi
dblp:158/3909
· DBLP profile ↗
26ranked-venue papers
16as first author
13since 2021 · last 2025
0000-0003-0125-4945ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 14 first-author · 12 since 2021Artificial intelligence and machine learning · 16 · 12 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Time-Aware Auto White Balance in Mobile Photography
Mahmoud Afifi, Luxi Zhao 0002, Abhijith Punnappurath, Mohammed A. Abdelsalam, Michael S. Brown |
ICCV | 1 |
| 2025 | CCMNet: Leveraging Calibrated Color Correction Matrices for Cross-Camera Color ConstancyabstractComputational color constancy, or white balancing, is a key module in a camera's image signal processor (ISP) that corrects color casts from scene lighting. Because this operation occurs in the camera-specific raw color space, white balance algorithms must adapt to different cameras. This paper introduces a learning-based method for cross-camera color constancy that generalizes to new cameras without retraining. Our method leverages pre-calibrated color correction matrices (CCMs) available on ISPs that map the camera's raw color space to a standard space (e.g., CIE XYZ). Our method uses these CCMs to transform predefined illumination colors (i.e., along the Planckian locus) into the test camera's raw space. The mapped illuminants are encoded into a compact camera fingerprint embedding (CFE) that enables the network to adapt to unseen cameras. To prevent overfitting due to limited cameras and CCMs during training, we introduce a data augmentation technique that interpolates between cameras and their CCMs. Experimental results across multiple datasets and backbones show that our method achieves state-of-the-art cross-camera color constancy while remaining lightweight and relying only on data readily available in camera ISPs. Dongyoung Kim, Mahmoud Afifi, Dongyun Kim, Michael S. Brown, Seon Joo Kim |
ICCV | 2 |
| 2025 | Color Matching Using Hypernetwork-Based Kolmogorov-Arnold Networks
Artem V. Nikonorov, Georgy Perevozchikov, Andrei Korepanov, Nancy Mehta, Mahmoud Afifi, Egor Ershov, Radu Timofte |
ICCV | 5 |
| 2025 | Multispectral Demosaicing via Dual CamerasabstractMultispectral (MS) images capture detailed scene information across a wide range of spectral bands, making them invaluable for applications requiring rich spectral data. Integrating MS imaging into multi camera devices, such as smartphones, has the potential to enhance both spectral applications and RGB image quality. A critical step in processing MS data is demosaicing, which reconstructs color information from the mosaic MS images captured by the camera. This paper proposes a method for MS image demosaicing specifically designed for dual-camera setups where both RGB and MS cameras capture the same scene. Our approach leverages co-captured RGB images, which typically have higher spatial fidelity, to guide the demosaicing of lower-fidelity MS images. We introduce the Dual-camera RGB-MS Dataset - a large collection of paired RGB and MS mosaiced images with ground-truth demosaiced outputs - that enables training and evaluation of our method. Experimental results demonstrate that our method achieves state-of-the-art accuracy compared to existing techniques. SaiKiran Kumar Tedla, Junyong Lee 0001, Beixuan Yang, Mahmoud Afifi, Michael S. Brown |
ICCV | 4 |
| 2024 | Optimizing Illuminant Estimation in Dual-Exposure HDR ImagingabstractHigh dynamic range (HDR) imaging involves capturing a series of frames of the same scene, each with different exposure settings, to broaden the dynamic range of light. This can be achieved through burst capturing or using staggered HDR sensors that capture long and short exposures simultaneously in the camera image signal processor (ISP). Within camera ISP pipeline, illuminant estimation is a crucial step aiming to estimate the color of the global illuminant in the scene. This estimation is used in camera ISP white-balance module to remove undesirable color cast in the final image. Despite the multiple frames captured in the HDR pipeline, conventional illuminant estimation methods often rely only on a single frame of the scene. In this paper, we explore leveraging information from frames captured with different exposure times. Specifically, we introduce a simple feature extracted from dual-exposure images to guide illuminant estimators, referred to as the dual-exposure feature (DEF). To validate the efficiency of DEF, we employed two illuminant estimators using the proposed DEF: 1) a multilayer perceptron network (MLP), referred to as exposure-based MLP (EMLP), and 2) a modified version of the convolutional color constancy (CCC) to integrate our DEF, that we call ECCC. Both EMLP and ECCC achieve promising results, in some cases surpassing prior methods that require hundreds of thousands or millions of parameters, with only a few hundred parameters for EMLP and a few thousand parameters for ECCC. Mahmoud Afifi, Zhenhua Hu |
ECCV (2) | 1 |
| 2024 | Rawformer: Unpaired Raw-to-Raw Translation for Learnable Camera ISPs
Georgy Perevozchikov, Nancy Mehta, Mahmoud Afifi, Radu Timofte |
ECCV (36) | 3 |
| 2022 | Improving Single-Image Defocus Deblurring: How Dual-Pixel Images Help Through Multi-Task LearningabstractMany camera sensors use a dual-pixel (DP) design that operates as a rudimentary light field providing two subaperture views of a scene in a single capture. The DP sensor was developed to improve how cameras perform autofocus. Since the DP sensor’s introduction, researchers have found additional uses for the DP data, such as depth estimation, reflection removal, and defocus deblurring. We are interested in the latter task of defocus deblurring. In particular, we propose a single-image deblurring network that incorporates the two sub-aperture views into a multitask framework. Specifically, we show that jointly learning to predict the two DP views from a single blurry input image improves the network’s ability to learn to deblur the image. Our experiments show this multi-task strategy achieves +1dB PSNR improvement over state-of-the-art defocus deblurring methods. In addition, our multi-task framework allows accurate DP-view synthesis (e.g., ∼39dB PSNR) from the single input image. These high-quality DP views can be used for other DP-based applications, such as reflection removal. As part of this effort, we have captured a new dataset of 7,059 high-quality images to support our training for the DP-view synthesis task. Abdullah Abuolaim, Mahmoud Afifi, Michael S. Brown |
WACV | 2 |
| 2022 | Auto White-Balance Correction for Mixed-Illuminant ScenesabstractAuto white balance (AWB) is applied by camera hardware at capture time to remove the color cast caused by the scene illumination. The vast majority of white-balance algorithms assume a single light source illuminates the scene; however, real scenes often have mixed lighting conditions. This paper presents an effective AWB method to deal with such mixed-illuminant scenes. A unique departure from conventional AWB, our method does not require illuminant estimation, as is the case in traditional camera AWB modules. Instead, our method proposes to render the captured scene with a small set of predefined white-balance settings. Given this set of rendered images, our method learns to estimate weighting maps that are used to blend the rendered images to generate the final corrected image. Through extensive experiments, we show this proposed method produces promising results compared to other alternatives for single-and mixed-illuminant scene color correction. Mahmoud Afifi, Marcus A. Brubaker, Michael S. Brown |
WACV | 1 |
| 2022 | CIE XYZ Net: Unprocessing Images for Low-Level Computer Vision TasksabstractCameras currently allow access to two image states: (i) a minimally processed linear raw-RGB image state (i.e., raw sensor data); or (ii) a highly-processed nonlinear image state (e.g., sRGB). There are many computer vision tasks that work best with a linear image state, such as image deblurring and image dehazing. Unfortunately, the vast majority of images are saved in the nonlinear image state. Because of this, a number of methods have been proposed to "unprocess" nonlinear images back to a raw-RGB state. However, existing unprocessing methods have a drawback because raw-RGB images are sensor-specific. As a result, it is necessary to know which camera produced the sRGB output and use a method or network tailored for that sensor to properly unprocess it. This paper addresses this limitation by exploiting another camera image state that is not available as an output, but it is available inside the camera pipeline. In particular, cameras apply a colorimetric conversion step to convert the raw-RGB image to a device-independent space based on the CIE XYZ color space before they apply the nonlinear photo-finishing. Leveraging this canonical image state, we propose a deep learning framework, CIE XYZ Net, that can unprocess a nonlinear image back to the canonical CIE XYZ image. This image can then be processed by any low-level computer vision operator and re-rendered back to the nonlinear image. We demonstrate the usefulness of the CIE XYZ Net on several low-level vision tasks and show significant gains that can be obtained by this processing framework. Code and dataset are publicly available at https://github.com/mahmoudnafifi/CIE_XYZ_NET. Mahmoud Afifi, Abdelrahman Abdelhamed, Abdullah Abuolaim, Abhijith Punnappurath, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Semi-Supervised Raw-to-Raw Mapping
Mahmoud Afifi, Abdullah Abuolaim |
BMVC | 1 |
| 2021 | HistoGAN: Controlling Colors of GAN-Generated and Real Images via Color HistogramsabstractWhile generative adversarial networks (GANs) can successfully produce high-quality images, they can be challenging to control. Simplifying GAN-based image generation is critical for their adoption in graphic design and artistic work. This goal has led to significant interest in methods that can intuitively control the appearance of images generated by GANs. In this paper, we present HistoGAN, a color histogram-based method for controlling GAN-generated images’ colors. We focus on color histograms as they provide an intuitive way to describe image color while remaining decoupled from domain-specific semantics. Specifically, we introduce an effective modification of the recent StyleGAN architecture [31] to control the colors of GAN-generated images specified by a target color histogram feature. We then describe how to expand HistoGAN to recolor real images. For image recoloring, we jointly train an encoder network along with HistoGAN. The recoloring model, ReHistoGAN, is an unsupervised approach trained to encourage the network to keep the original image’s content while changing the colors based on the given target histogram. We show that this histogram-based approach offers a better way to control GAN-generated and real images’ colors while producing more compelling results compared to existing alternative strategies. Mahmoud Afifi, Marcus A. Brubaker, Michael S. Brown |
CVPR | 1 |
| 2021 | Learning Multi-Scale Photo Exposure CorrectionabstractCapturing photographs with wrong exposures remains a major source of errors in camera-based imaging. Exposure problems are categorized as either: (i) overexposed, where the camera exposure was too long, resulting in bright and washed-out image regions, or (ii) underexposed, where the exposure was too short, resulting in dark regions. Both under- and overexposure greatly reduce the contrast and visual appeal of an image. Prior work mainly focuses on underexposed images or general image enhancement. In contrast, our proposed method targets both over- and underexposure errors in photographs. We formulate the exposure correction problem as two main sub-problems: (i) color enhancement and (ii) detail enhancement. Accordingly, we propose a coarse-to-fine deep neural network (DNN) model, trainable in an end-to-end manner, that addresses each sub-problem separately. A key aspect of our solution is a new dataset of over 24,000 images exhibiting the broadest range of exposure values to date with a corresponding properly exposed image. Our method achieves results on par with existing state-of-the-art methods on underexposed images and yields significant improvements for images suffering from overexposure errors. Mahmoud Afifi, Konstantinos G. Derpanis, Björn Ommer, Michael S. Brown |
CVPR | 1 |
| 2021 | Cross-Camera Convolutional Color ConstancyabstractWe present "Cross-Camera Convolutional Color Constancy" (C5), a learning-based method, trained on images from multiple cameras, that accurately estimates a scene’s illuminant color from raw images captured by a new camera previously unseen during training. C5 is a hypernetwork-like extension of the convolutional color constancy (CCC) approach: C5 learns to generate the weights of a CCC model that is then evaluated on the input image, with the CCC weights dynamically adapted to different input content. Unlike prior cross-camera color constancy models, which are usually designed to be agnostic to the spectral properties of test-set images from unobserved cameras, C5 approaches this problem through the lens of transductive inference: additional unlabeled images are provided as input to the model at test time, which allows the model to calibrate itself to the spectral properties of the test-set camera during inference. C5 achieves state-of-the-art accuracy for cross-camera color constancy on several datasets, is fast to evaluate (~7 and ~90 ms per image on a GPU or CPU, respectively), and requires little memory (~2 MB), and thus is a practical solution to the problem of calibration-free automatic white balance for mobile photography. Mahmoud Afifi, Jonathan T. Barron, Chloe LeGendre, Yun-Ta Tsai, Francois Bleibel |
ICCV | 1 |
| 2020 | Deep White-Balance EditingabstractWe introduce a deep learning approach to realistically edit an sRGB image's white balance. Cameras capture sensor images that are rendered by their integrated signal processor (ISP) to a standard RGB (sRGB) color space encoding. The ISP rendering begins with a white-balance procedure that is used to remove the color cast of the scene's illumination. The ISP then applies a series of nonlinear color manipulations to enhance the visual quality of the final sRGB image. Recent work by [3] showed that sRGB images that were rendered with the incorrect white balance cannot be easily corrected due to the ISP's nonlinear rendering. The work in [3] proposed a k-nearest neighbor (KNN) solution based on tens of thousands of image pairs. We propose to solve this problem with a deep neural network (DNN) architecture trained in an end-to-end manner to learn the correct white balance. Our DNN maps an input image to two additional white-balance settings corresponding to indoor and outdoor illuminations. Our solution not only is more accurate than the KNN approach in terms of correcting a wrong white-balance setting but also provides the user the freedom to edit the white balance in the sRGB image to other illumination settings. Mahmoud Afifi, Michael S. Brown |
CVPR | 1 |
| 2020 | Modeling Defocus-Disparity in Dual-Pixel SensorsabstractMost modern consumer cameras use dual-pixel (DP) sensors that provide two sub-aperture views of the scene in a single photo capture. The DP sensor was designed to assist the camera's autofocus routine, which examines local disparity in the two sub-aperture views to determine which parts of the image are out of focus. Recently, these DP views have been used for tasks beyond autofocus, such as synthetic bokeh, reflection removal, and depth reconstruction. These recent methods treat the two DP views as stereo image pairs and apply stereo matching algorithms to compute local disparity. However, dual-pixel disparity is not caused by view parallax as in stereo, but instead is attributed to defocus blur that occurs in out-of-focus regions in the image. This paper proposes a new parametric point spread function to model the defocus-disparity that occurs on DP sensors. We apply our model to the task of depth estimation from DP data. An important feature of our model is its ability to exploit the symmetry property of the DP blur kernels at each pixel. We leverage this symmetry property to formulate an unsupervised loss function that does not require ground truth depth. We demonstrate our method's effectiveness on both DSLR and smartphone DP data. Abhijith Punnappurath, Abdullah Abuolaim, Mahmoud Afifi, Michael S. Brown |
ICCP | 3 |
| 2019 | Sensor-Independent Illumination Estimation for DNN Models
Mahmoud Afifi, Michael S. Brown |
BMVC | 1 |
| 2019 | When Color Constancy Goes Wrong: Correcting Improperly White-Balanced ImagesabstractThis paper focuses on correcting a camera image that has been improperly white-balanced. This situation occurs when a camera's auto white balance fails or when the wrong manual white-balance setting is used. Even after decades of computational color constancy research, there are no effective solutions to this problem. The challenge lies not in identifying what the correct white balance should have been, but in the fact that the in-camera white-balance procedure is followed by several camera-specific nonlinear color manipulations that make it challenging to correct the image's colors in post-processing. This paper introduces the first method to explicitly address this problem. Our method is enabled by a dataset of over 65,000 pairs of incorrectly white-balanced images and their corresponding correctly white-balanced images. Using this dataset, we introduce a k-nearest neighbor strategy that is able to compute a nonlinear color mapping function to correct the image's colors. We show our method is highly effective and generalizes well to camera models not in the training set. Mahmoud Afifi, Brian L. Price, Scott Cohen, Michael S. Brown |
CVPR | 1 |
| 2019 | What Else Can Fool Deep Learning? Addressing Color Constancy Errors on Deep Neural Network PerformanceabstractThere is active research targeting local image manipulations that can fool deep neural networks (DNNs) into producing incorrect results. This paper examines a type of global image manipulation that can produce similar adverse effects. Specifically, we explore how strong color casts caused by incorrectly applied computational color constancy - referred to as white balance (WB) in photography - negatively impact the performance of DNNs targeting image segmentation and classification. In addition, we discuss how existing image augmentation methods used to improve the robustness of DNNs are not well suited for modeling WB errors. To address this problem, a novel augmentation method is proposed that can emulate accurate color constancy degradation. We also explore pre-processing training and testing images with a recent WB correction algorithm to reduce the effects of incorrectly white-balanced images. We examine both augmentation and pre-processing strategies on different datasets and demonstrate notable improvements on the CIFAR-10, CIFAR-100, and ADE20K datasets. Mahmoud Afifi, Michael S. Brown |
ICCV | 1 |
| 2019 | A versatile computational framework for group pattern mining of pedestrian trajectories
Abdullah M. Sawas, Abdullah Abuolaim, Mahmoud Afifi, Manos Papagelis |
GeoInformatica | 3 |
| 2019 | The achievement of higher flexibility in multiple-choice-based tests using image classification techniques
Mahmoud Afifi, Khaled F. Hussain |
Int. J. Document Anal. Recognit. | 1 |
| 2019 | AFIF4: Deep gender classification based on AdaBoost-based fusion of isolated facial features and foggy faces
Mahmoud Afifi, Abdelrahman Abdelhamed |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | 11K Hands: Gender recognition and biometric identification using a large dataset of hand images
Mahmoud Afifi |
Multim. Tools Appl. | 1 |
| 2019 | A Comprehensive Study of the Effect of Spatial Resolution and Color of Digital Images on Vehicle ClassificationabstractVehicle-type classification is considered a core module for many intelligent transportation applications, such as speed monitoring, smart parking systems, and traffic analysis. In this paper, many vision-based classification techniques were presented relying only on a digital camera without the need for any extra hardware components. Dimension and color are two important characteristics of any digital image that affect the cost of the digital camera used in the image acquisition. In this paper, we present a comprehensive study of the effect of these two characteristics on the vehicle classification process in terms of accuracy and performance. We apply a set of different state-of-the-art image classifiers to the BIT-Vehicle and LabelMe data sets. Each data set is downscaled into different scales to generate a variety of spatial resolutions of each data set. Besides, we examine the effect of color by converting each color version to a gray-scale one. At last, we draw a valid conclusion in regards to the impact of these two characteristics (i.e., dimension and color) on the classification accuracy and performance of the image classification methods using more than 46 000 individual experiments. Experimental results show that there is no significant influence of both color and spatial resolutions of the vehicle images on the classification results obtained by most state-of-the-art image classification methods. However, there is a correlation between the spatial resolution and the processing time required by most image classification methods. Our findings can play an important role in saving not only money, but also time for vehicle-type classification systems. Khaled F. Hussain, Mahmoud Afifi, Ghada S. Moussa |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Tensor Methods for Group Pattern Discovery of Pedestrian TrajectoriesabstractMining large-scale trajectory data streams (of moving objects) has been of ever increasing research interest due to an abundance of modern tracking devices and its large number of critical applications. In this paper, we are interested in mining group patterns of moving objects. Group pattern mining describes a special type of trajectory mining task that requires to efficiently discover trajectories of objects that are found in close proximity to each other for a period of time. In particular, we focus on trajectories of pedestrians coming from motion video analysis and we are interested in interactive analysis and exploration of group dynamics, including various definitions of group gathering and dispersion. Towards this end, we present a suite of (three) tensor-based methods for efficient discovery of evolving groups of pedestrians. Traditional approaches to solve the problem heavily rely on well-defined clustering algorithms to discover groups of pedestrians at each time point, and then post-process these groups to discover groups that satisfy specific group pattern semantics, including time constraints. In contrast, our proposed methods are based on efficiently discovering pairs of pedestrians that move together over time, under varying conditions. Pairs of pedestrians are subsequently used as a building block for effectively discovering groups of pedestrians. The suite of proposed methods provides the ability to adapt to many different scenarios and application requirements. Furthermore, a query-based search method is provided that allows for interactive exploration and analysis of group dynamics over time and space. Through experiments on real data, we demonstrate the effectiveness of our methods on discovering group patterns of pedestrian trajectories against sensible baselines, for a varying range of conditions. In addition, a visual testing is performed on real motion video to assert the group dynamics discovered by each method. Abdullah M. Sawas, Abdullah Abuolaim, Mahmoud Afifi, Manos Papagelis |
MDM | 3 |
| 2018 | Trajectolizer: Interactive Analysis and Exploration of Trajectory Group DynamicsabstractMining large-scale trajectory data streams (of moving objects) has been of ever increasing research interest due to an abundance of modern tracking devices and its large number of critical applications. A challenging task in this domain is that of mining group patterns of moving objects. Group pattern mining describes a special type of trajectory mining that requires to efficiently discover trajectories of objects that are found in close proximity to each other for a period of time. To this end, we introduce Trajectolizer, an online system for interactive analysis and exploration of trajectory group dynamics over time and space. We describe the system and demonstrate its effectiveness on discovering group patterns on trajectories of pedestrians. The system architecture and methods are general and can be used to perform group analysis of any domain-specific trajectories. Abdullah M. Sawas, Abdullah Abuolaim, Mahmoud Afifi, Manos Papagelis |
MDM | 3 |
| 2015 | MPB: A modified Poisson blending techniqueabstractImage cloning has many useful applications, such as removing unwanted objects, fixing damaged parts of images, and panorama stitching. Instead of using pixel intensities, the gradient domain is used in Poisson image editing; however, it suffers from two main problems: color bleeding and bleeding artifacts. In this paper, a modified Poisson blending (MPB) technique is presented which ensures dependency on the boundary pixels of both target and source images rather than just those of the target. The problem of bleeding artifacts is reduced. This makes the proposed technique suitable for use in video compositing as it avoids the flickering caused by bleeding artifacts. To reduce the problem of color bleeding, we use an additional alpha compositing step. Our experimental results using the proposed technique show that MPB reduces the bleeding problems and generates more natural composited images than other techniques. Mahmoud Afifi, Khaled F. Hussain |
Comput. Vis. Media | 1 |