Abdullah Abuolaim

dblp:223/1200 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
6since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3
YearPublicationVenuePosition
2023 DC2: Dual-Camera Defocus Control by Learning to Refocus
abstract
Smartphone cameras today are increasingly approaching the versatility and quality of professional cameras through a combination of hardware and software advancements. However, fixed aperture remains a key limitation, preventing users from controlling the depth of field (DoF) of captured images. At the same time, many smartphones now have multiple cameras with different fixed apertures - specifically, an ultra-wide camera with wider field of view and deeper DoF and a higher resolution primary camera with shallower DoF. In this work, we propose$DC^{2}$, a system for defocus control for synthetically varying camera aperture, focus distance and arbitrary defocus effects by fusing information from such a dual-camera system. Our key insight is to leverage real-world smartphone camera dataset by using image refocus as a proxy task for learning to control defocus. Quantitative and qualitative evaluations on real-world data demonstrate our system's efficacy where we outperform state-of-the-art on defocus deblurring, bokeh rendering, and image refocus. Finally, we demonstrate creative post-capture defocus control enabled by our method, including tilt-shift and content-based defocus effects.
Hadi Alzayer, Abdullah Abuolaim, Leung Chun Chan, Ying Chen Lou, Jia-Bin Huang 0001, Abhishek Kar
CVPR2
2022 Day-to-Night Image Synthesis for Training Nighttime Neural ISPs
abstract
Many flagship smartphone cameras now use a dedicated neural image signal processor (ISP) to render noisy raw sensor images to the final processed output. Training night-mode ISP networks relies on large-scale datasets of image pairs with: (1) a noisy raw image captured with a short exposure and a high ISO gain; and (2) a ground truth low-noise raw image captured with a long exposure and low ISO that has been rendered through the ISP. Capturing such image pairs is tedious and time-consuming, requiring careful setup to ensure alignment between the image pairs. In addition, ground truth images are often prone to motion blur due to the long exposure. To address this problem, we propose a method that synthesizes nighttime images from day-time images. Daytime images are easy to capture, exhibit low-noise (even on smartphone cameras) and rarely suffer from motion blur. We outline a processing framework to convert daytime raw images to have the appearance of realistic nighttime raw images with different levels of noise. Our procedure allows us to easily produce aligned noisy and clean nighttime image pairs. We show the effectiveness of our synthesis framework by training neural ISPs for nightmode rendering. Furthermore, we demonstrate that using our synthetic nighttime images together with small amounts of real data (e.g., 5% to 10%) yields performance almost on par with training exclusively on real nighttime images. Our dataset and code are available at https://github.com/SamsungLabs/day-to-night.
Abhijith Punnappurath, Abdullah Abuolaim, Abdelrahman Abdelhamed, Alex Levinshtein, Michael S. Brown
CVPR2
2022 Improving Single-Image Defocus Deblurring: How Dual-Pixel Images Help Through Multi-Task Learning
abstract
Many camera sensors use a dual-pixel (DP) design that operates as a rudimentary light field providing two subaperture views of a scene in a single capture. The DP sensor was developed to improve how cameras perform autofocus. Since the DP sensor’s introduction, researchers have found additional uses for the DP data, such as depth estimation, reflection removal, and defocus deblurring. We are interested in the latter task of defocus deblurring. In particular, we propose a single-image deblurring network that incorporates the two sub-aperture views into a multitask framework. Specifically, we show that jointly learning to predict the two DP views from a single blurry input image improves the network’s ability to learn to deblur the image. Our experiments show this multi-task strategy achieves +1dB PSNR improvement over state-of-the-art defocus deblurring methods. In addition, our multi-task framework allows accurate DP-view synthesis (e.g., ∼39dB PSNR) from the single input image. These high-quality DP views can be used for other DP-based applications, such as reflection removal. As part of this effort, we have captured a new dataset of 7,059 high-quality images to support our training for the DP-view synthesis task.
Abdullah Abuolaim, Mahmoud Afifi, Michael S. Brown
WACV1
2022 CIE XYZ Net: Unprocessing Images for Low-Level Computer Vision Tasks
abstract
Cameras currently allow access to two image states: (i) a minimally processed linear raw-RGB image state (i.e., raw sensor data); or (ii) a highly-processed nonlinear image state (e.g., sRGB). There are many computer vision tasks that work best with a linear image state, such as image deblurring and image dehazing. Unfortunately, the vast majority of images are saved in the nonlinear image state. Because of this, a number of methods have been proposed to "unprocess" nonlinear images back to a raw-RGB state. However, existing unprocessing methods have a drawback because raw-RGB images are sensor-specific. As a result, it is necessary to know which camera produced the sRGB output and use a method or network tailored for that sensor to properly unprocess it. This paper addresses this limitation by exploiting another camera image state that is not available as an output, but it is available inside the camera pipeline. In particular, cameras apply a colorimetric conversion step to convert the raw-RGB image to a device-independent space based on the CIE XYZ color space before they apply the nonlinear photo-finishing. Leveraging this canonical image state, we propose a deep learning framework, CIE XYZ Net, that can unprocess a nonlinear image back to the canonical CIE XYZ image. This image can then be processed by any low-level computer vision operator and re-rendered back to the nonlinear image. We demonstrate the usefulness of the CIE XYZ Net on several low-level vision tasks and show significant gains that can be obtained by this processing framework. Code and dataset are publicly available at https://github.com/mahmoudnafifi/CIE_XYZ_NET.
Mahmoud Afifi, Abdelrahman Abdelhamed, Abdullah Abuolaim, Abhijith Punnappurath, Michael S. Brown
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Semi-Supervised Raw-to-Raw Mapping
Mahmoud Afifi, Abdullah Abuolaim
BMVC2
2021 Learning to Reduce Defocus Blur by Realistically Modeling Dual-Pixel Data
abstract
Recent work has shown impressive results on data-driven defocus deblurring using the two-image views available on modern dual-pixel (DP) sensors. One significant challenge in this line of research is access to DP data. Despite many cameras having DP sensors, only a limited number provide access to the low-level DP sensor images. In addition, capturing training data for defocus deblurring involves a time-consuming and tedious setup requiring the camera’s aperture to be adjusted. Some cameras with DP sensors (e.g., smartphones) do not have adjustable apertures, further limiting the ability to produce the necessary training data. We address the data capture bottleneck by proposing a procedure to generate realistic DP data synthetically. Our synthesis approach mimics the optical image formation found on DP sensors and can be applied to virtual scenes rendered with standard computer software. Leveraging these realistic synthetic DP images, we introduce a recurrent convolutional network (RCN) architecture that improves deblurring results and is suitable for use with single-frame and multi-frame data (e.g., video) captured by DP sensors. Finally, we show that our synthetic DP data is useful for training DNN models targeting video deblurring applications where access to DP data remains challenging.
Abdullah Abuolaim, Mauricio Delbracio, Damien Kelly, Michael S. Brown, Peyman Milanfar
ICCV1
2020 Defocus Deblurring Using Dual-Pixel Data
Abdullah Abuolaim, Michael S. Brown
ECCV (10)1
2020 Modeling Defocus-Disparity in Dual-Pixel Sensors
abstract
Most modern consumer cameras use dual-pixel (DP) sensors that provide two sub-aperture views of the scene in a single photo capture. The DP sensor was designed to assist the camera's autofocus routine, which examines local disparity in the two sub-aperture views to determine which parts of the image are out of focus. Recently, these DP views have been used for tasks beyond autofocus, such as synthetic bokeh, reflection removal, and depth reconstruction. These recent methods treat the two DP views as stereo image pairs and apply stereo matching algorithms to compute local disparity. However, dual-pixel disparity is not caused by view parallax as in stereo, but instead is attributed to defocus blur that occurs in out-of-focus regions in the image. This paper proposes a new parametric point spread function to model the defocus-disparity that occurs on DP sensors. We apply our model to the task of depth estimation from DP data. An important feature of our model is its ability to exploit the symmetry property of the DP blur kernels at each pixel. We leverage this symmetry property to formulate an unsupervised loss function that does not require ground truth depth. We demonstrate our method's effectiveness on both DSLR and smartphone DP data.
Abhijith Punnappurath, Abdullah Abuolaim, Mahmoud Afifi, Michael S. Brown
ICCP2
2020 Online Lens Motion Smoothing for Video Autofocus
abstract
Autofocus (AF) is the process of moving the camera's lens such that desired scene content is in focus. AF for single image capture is a well-studied research topic and most modern cameras have hardware support that allows quick lens movements to optimize image sharpness. How to best perform AF for video is less clear. Conventional wisdom would suggest that each temporal frame should be as sharp as possible. However, unlike single image capture, the effects of the lens movement is visible in the captured video. As a result, there are two parameters to consider in AF for video: sharpness and lens movement. In this paper, we show that users preferred videos with smooth lens movement, even if it results in less overall sharpness. Based on this observation, we propose two novel AF algorithms for video that strive for both smooth lens movement and sharp scene content. Specifically, we introduce (1) a bidirectional long short-term memory (BLSTM) module trained on smooth lens trajectories and (2) a simple weighted moving average (WMA) method that factors in prior lens motion. Both of these methods have demonstrated excellent results in terms of reducing lens movements (up to 64% reduction) without greatly affecting the sharpness (less than 5.2% change in sharpness). Moreover, videos produced using our methods are more preferred by users over conventional AF that aims only for maximizing sharpness.
Abdullah Abuolaim, Michael S. Brown
WACV1
2019 A versatile computational framework for group pattern mining of pedestrian trajectories
Abdullah M. Sawas, Abdullah Abuolaim, Mahmoud Afifi, Manos Papagelis
GeoInformatica2
2018 Revisiting Autofocus for Smartphone Cameras
Abdullah Abuolaim, Abhijith Punnappurath, Michael S. Brown
ECCV (15)1
2018 Tensor Methods for Group Pattern Discovery of Pedestrian Trajectories
abstract
Mining large-scale trajectory data streams (of moving objects) has been of ever increasing research interest due to an abundance of modern tracking devices and its large number of critical applications. In this paper, we are interested in mining group patterns of moving objects. Group pattern mining describes a special type of trajectory mining task that requires to efficiently discover trajectories of objects that are found in close proximity to each other for a period of time. In particular, we focus on trajectories of pedestrians coming from motion video analysis and we are interested in interactive analysis and exploration of group dynamics, including various definitions of group gathering and dispersion. Towards this end, we present a suite of (three) tensor-based methods for efficient discovery of evolving groups of pedestrians. Traditional approaches to solve the problem heavily rely on well-defined clustering algorithms to discover groups of pedestrians at each time point, and then post-process these groups to discover groups that satisfy specific group pattern semantics, including time constraints. In contrast, our proposed methods are based on efficiently discovering pairs of pedestrians that move together over time, under varying conditions. Pairs of pedestrians are subsequently used as a building block for effectively discovering groups of pedestrians. The suite of proposed methods provides the ability to adapt to many different scenarios and application requirements. Furthermore, a query-based search method is provided that allows for interactive exploration and analysis of group dynamics over time and space. Through experiments on real data, we demonstrate the effectiveness of our methods on discovering group patterns of pedestrian trajectories against sensible baselines, for a varying range of conditions. In addition, a visual testing is performed on real motion video to assert the group dynamics discovered by each method.
Abdullah M. Sawas, Abdullah Abuolaim, Mahmoud Afifi, Manos Papagelis
MDM2
2018 Trajectolizer: Interactive Analysis and Exploration of Trajectory Group Dynamics
abstract
Mining large-scale trajectory data streams (of moving objects) has been of ever increasing research interest due to an abundance of modern tracking devices and its large number of critical applications. A challenging task in this domain is that of mining group patterns of moving objects. Group pattern mining describes a special type of trajectory mining that requires to efficiently discover trajectories of objects that are found in close proximity to each other for a period of time. To this end, we introduce Trajectolizer, an online system for interactive analysis and exploration of trajectory group dynamics over time and space. We describe the system and demonstrate its effectiveness on discovering group patterns on trajectories of pedestrians. The system architecture and methods are general and can be used to perform group analysis of any domain-specific trajectories.
Abdullah M. Sawas, Abdullah Abuolaim, Mahmoud Afifi, Manos Papagelis
MDM2