VLDB 2026 Research / reviewers in the wild / expert
Rafal Mantiuk
dblp:41/1369 · also Rafal K. Mantiuk
· DBLP profile ↗
89ranked-venue papers
14as first author
28since 2021 · last 2026
0000-0003-2353-0349ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 82 · 13 first-author · 28 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 15 · 10 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Quality Metrics Through the Lenses of Psychophysical Measurements of Low-Level VisionabstractImage and video quality metrics, such as SSIM, LPIPS, and VMAF, aim to predict perceived visual quality and are often assumed to reflect principles of human vision. However, relatively few metrics explicitly incorporate models of human perception, with most relying on hand-crafted formulas or data-driven training to approximate perceptual alignment. In this paper, we introduce a set of tests for full-reference quality metrics that evaluate their ability to capture key aspects of low-level human vision: contrast sensitivity, contrast masking, and contrast matching. These tests provide an additional framework for assessing both established and newly proposed metrics. We apply the tests to 34 existing quality metrics and highlight patterns in their behavior, including the ability of LPIPS and MS-SSIM to predict contrast masking and the tendency of SSIM to overemphasize high spatial frequencies, which is mitigated in MS-SSIM, and the general inability of metrics to model supra-threshold contrast constancy. Our results demonstrate how these tests can reveal properties of quality metrics that are not easily observed with standard evaluation protocols. Dounia Hammou, Yancheng Cai, Pavan Madhusudanarao, Christos G. Bampis, Zhi Li 0001, Rafal Mantiuk |
QoMEX | 6 |
| 2026 | QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu |
QoMEX | 30 |
| 2025 | Do Computer Vision Foundation Models Learn the Low-level Characteristics of the Human Visual System?abstractComputer vision foundation models, such as DINO or OpenCLIP, are trained in a self-supervised manner on large image datasets. Analogously, substantial evidence suggests that the human visual system (HVS) is influenced by the statistical distribution of colors and patterns in the natural world, characteristics also present in the training data of foundation models. The question we address in this paper is whether foundation models trained on natural images mimic some of the low-level characteristics of the human visual system, such as contrast detection, contrast masking, and contrast constancy. Specifically, we designed a protocol comprising nine test types to evaluate the image encoders of 45 foundation and generative models. Our results indicate that some foundation models (e.g., DINO, DINOv2, and OpenCLIP), share some of the characteristics of human vision, but other models show little resemblance. Foundation models tend to show smaller sensitivity to low contrast and rather irregular responses to contrast across frequencies. The foundation models show the best agreement with human data in terms of contrast masking. Our findings suggest that human vision and computer vision may take both similar and different paths when learning to interpret images of the real world. Overall, while differences remain, foundation models trained on vision tasks start to align with low-level human vision, with DINOv2 showing the closest resemblance. Our code is available on https://github.com/caiyancheng/VFM_HVS_CVPR2025. Yancheng Cai, Dounia Hammou, Rafal Mantiuk |
CVPR | 4 |
| 2025 | FaceCraft4D: Animated 3D Facial Avatar Generation from a Single Image
Mallikarjun B. R. 0001, Chun-Han Yao, Rafal Mantiuk, Varun Jampani |
ICCV | 4 |
| 2025 | CameraVDP: Perceptual Display Assessment with Uncertainty Estimation via Camera and Visual Difference PredictionabstractAccurate measurement of images produced by electronic displays is critical for the evaluation of both traditional and computational displays. Traditional display measurement methods based on sparse radiometric sampling and fitting a model are inadequate for capturing spatially varying display artifacts, as they fail to capture high-frequency and pixel-level distortions. While cameras offer sufficient spatial resolution, they introduce optical, sampling, and photometric distortions. Furthermore, the physical measurement must be combined with a model of a visual system to assess whether the distortions are going to be visible. To enable perceptual assessment of displays, we propose a combination of a camera-based reconstruction pipeline with a visual difference predictor, which account for both the inaccuracy of camera measurements and visual difference prediction. The reconstruction pipeline combines HDR image stacking, MTF inversion, vignetting correction, geometric undistortion, homography transformation, and color correction, enabling cameras to function as precise display measurement instruments. By incorporating a Visual Difference Predictor (VDP), our system models the visibility of various stimuli under different viewing conditions for the human visual system. We validate the proposed CameraVDP framework through three applications: defective pixel detection, color fringing awareness, and display non-uniformity evaluation. Our uncertainty analysis framework enables the estimation of the theoretical upper bound for defect pixel detection performance and provides confidence intervals for VDP quality scores. Our code is available on https://github.com/gfxdisp/CameraVDP. Yancheng Cai, Robert Wanat, Rafal Mantiuk |
SIGGRAPH Asia | 3 |
| 2025 | Supra-threshold Contrast Perception in Augmented RealityabstractWhen an image is seen on an optical see-through augmented reality (AR) display, the light from the display is mixed with the background light from the environment. This can severely limit the available contrast in AR, which is often orders of magnitude below that of traditional displays. Yet, the presented images appear sharper and show more details than the reduction in physical contrast would indicate. In this work, we hypothesize two effects that are likely responsible for the enhanced perceived contrast in AR: background discounting, which allows observers focused on the display plane to partially discount the light from the environment; and supra-threshold contrast perception, which explains the differences in contrast perception across luminance levels. In a series of controlled experiments on an AR high-dynamic-range multi-focal haploscope testbed, we found no statistical evidence supporting the effect of background discounting on contrast perception. Instead, the increase of visibility in AR is better explained with models of supra-threshold contrast perception. Our findings can be generalized to incorporate an image input, and this model serves to design better algorithms and hardware for display systems affected by additive light, such as AR. Dongyeon Kim, Maliha Ashraf, Alexandre Chapiro, Rafal Mantiuk |
SIGGRAPH Asia | 4 |
| 2025 | Perceived quality of BRDF modelsabstractAbstract Material appearance is commonly modeled with the Bidirectional Reflectance Distribution Functions (BRDFs), which need to trade accuracy for complexity and storage cost. To investigate the current practices of BRDF modeling, we collect the first high dynamic range stereoscopic video dataset that captures the perceived quality degradation with respect to a number of parametric and non‐parametric BRDF models. Our dataset shows that the current loss functions used to fit BRDF models, such as mean‐squared error of logarithmic reflectance values, correlate poorly with the perceived quality of materials in rendered videos. We further show that quality metrics that compare rendered material samples give a significantly higher correlation with subjective quality judgments, and a simple Euclidean distance in the ITP color space (ΔEITP) shows the highest correlation. Additionally, we investigate the use of different BRDF‐space metrics as loss functions for fitting BRDF models and find that logarithmic mapping is the most effective approach for BRDF‐space loss functions. Behnaz Kavoosighafi, Rafal Mantiuk, Saghi Hajisharif, Ehsan Miandji, Jonas Unger |
Comput. Graph. Forum | 2 |
| 2024 | Perceptual Assessment and Optimization of HDR Image RenderingabstractHigh dynamic range (HDR) rendering has the ability to faithfully reproduce the wide luminance ranges in natural scenes, but how to accurately assess the rendering quality is relatively underexplored. Existing quality models are mostly designed for low dynamic range (LDR) images, and do not align well with human perception of HDR image quality. To fill this gap, we propose a family of HDR quality metrics, in which the key step is employing a simple inverse display model to decompose an HDR image into a stack of LDR images with varying exposures. Subsequently, these decomposed images are assessed through well-established LDR quality metrics. Our HDR quality models present three distinct benefits. First, they directly inherit the recent advancements of LDR quality metrics. Second, they do not rely on human perceptual data of HDR image quality for re-calibration. Third, they facilitate the alignment and prioritization of specific luminance ranges for more accurate and detailed quality assessment. Experimental results show that our HDR quality metrics consistently outperform existing models in terms of quality assessment on four HDR image quality datasets and perceptual optimization of HDR novel view synthesis. Peibei Cao, Rafal Mantiuk, Kede Ma |
CVPR | 2 |
| 2024 | Hypernetworks for Generalizable BRDF Representation
Fazilet Gokbudak, Alejandro Sztrajman, Chenliang Zhou, Fangcheng Zhong, Rafal Mantiuk, A. Cengiz Öztireli |
ECCV (76) | 5 |
| 2024 | Tracking Eye Position and Gaze Direction in Near-Eye Volumetric DisplaysabstractNear-eye volumetric displays, showing multiple focal planes, require knowledge of the accurate position of the nodal point of the eye to correctly render a 3D scene. This is because pixels seen through multiple planes must be accurately aligned with the eye’s visual axis to ensure consistency across focal planes. While most eye-tracking methods focus on determining a gaze position within a designated target space, this work aims to track both the eye position and the corresponding gaze direction expressed in coordinates relative to the physical location of the volumetric display planes. To achieve this, we rely on a near-infra-red (NIR) camera image of the pupil and corneal reflections (glints). The existing eye model is used to establish the relationship between the pupil and glint positions in a NIR image and the eye position and rotation in a 3D space. We address the key challenge of robust tracking of the glints in a system that introduces multiple reflections. We also demonstrate that the system reduces the need for recalibration on subsequent uses. Our experiments on a multiple-focal plane display demonstrate that the method can maintain an accurate projection point for volumetric displays. Marek Wernikowski, Joseph March, Radoslaw Mantiuk, Ali Özgür Yöntem, Rafal Mantiuk |
ISMAR | 5 |
| 2024 | The effect of viewing distance and display peak luminance - HDR AV1 video streaming quality datasetabstractWhile it is well recognized that the visibility of distortions is affected by the viewing distance and display peak luminance, very few datasets control those conditions, and also few video quality metrics can account for them. To address this gap, we collected a new video quality dataset, HDR-VDC, which captures the quality degradation of HDR content due to AV1 coding artifacts and the resolution reduction. The quality drop was measured at two viewing distances, corresponding to 60 and 120 pixels per visual degree, and two display mean luminance levels, 51 and 5.6 nits. In contrast to the existing datasets that use direct rating protocol, we employ a highly sensitive pairwise comparison protocol with active sampling and comparisons across viewing distances to ensure possibly accurate quality measurements. We also provide the first publicly available dataset that measures the effect of display peak luminance and includes HDR videos encoded with AV1. Our results indicate that the effect of both viewing distance and display luminance is significant, and it reduces the visibility of coding and upsampling artifacts on dimmer displays or those seen from a further distance. The dataset is available at https://doi.org/10.17863/CAM.107964 and the code at https://github.com/gfxdisp/HDR-VDC. Dounia Hammou, Lukas Krasula, Christos G. Bampis, Zhi Li 0001, Rafal Mantiuk |
QoMEX | 5 |
| 2024 | elaTCSF: A Temporal Contrast Sensitivity Function for Flicker Detection and Modeling Variable Refresh Rate FlickerabstractThe perception of flicker has been a prominent concern in illumination and electronic display fields for over a century. Traditional approaches often rely on Critical Flicker Frequency (CFF), primarily suited for high-contrast (full-on, full-off) flicker. To tackle varying contrast flicker, the International Committee for Display Metrology (ICDM) introduced a Temporal Contrast Sensitivity Function TCSF$_{IDMS}$ within the Information Display Measurements Standard (IDMS). Nevertheless, this standard overlooks crucial parameters: luminance, eccentricity, and area. Existing models incorporating these parameters are inadequate for flicker detection, especially at low spatial frequencies. To address these limitations, we extend the TCSF$_{IDMS}$ and combine it with a new spatial probability summation model to incorporate the effects of luminance, eccentricity, and area (elaTCSF). We train the elaTCSF on various flicker detection datasets and establish the first variable refresh rate flicker detection dataset for further verification. Additionally, we contribute to resolving a longstanding debate on whether the flicker is more visible in peripheral vision. We demonstrate how elaTCSF can be used to predict flicker due to low-persistence in VR headsets, identify flicker-free VRR operational ranges, and determine flicker sensitivity in lighting design. Yancheng Cai, Ali Bozorgian, Maliha Ashraf, Robert Wanat, Rafal Mantiuk |
SIGGRAPH Asia | 5 |
| 2024 | Color-Accurate Camera Capture with Multispectral Illumination and Multiple ExposuresabstractAbstract Cameras cannot capture the same colors as those seen by the human eye because the eye and the cameras' sensors differ in their spectral sensitivity. To obtain a plausible approximation of perceived colors, the camera's Image Signal Processor (ISP) employs a color correction step. However, even advanced color correction methods cannot solve this underdetermined problem, and visible color inaccuracies are always present. Here, we explore an approach in which we can capture accurate colors with a regular camera by optimizing the spectral composition of the illuminant and capturing one or more exposures. We jointly optimize for the signal‐to‐noise ratio and for the color accuracy irrespective of the spectral composition of the scene. One or more images captured under controlled multispectral illuminants are then converted into a color‐accurate image as seen under the standard illuminant of D65. Our optimization allows us to reduce the color error by 20–60% (in terms of CIEDE 2000), depending on the number of exposures and camera type. The method can be used in applications in which illumination can be controlled, and high colour accuracy is required, such as product photography or with a multispectral camera flash. The code is available at https://github.com/gfxdisp/multispectral_color_correction . Hongyun Gao 0001, Rafal Mantiuk, Graham D. Finlayson |
Comput. Graph. Forum | 2 |
| 2024 | Perceptual Quality Assessment of NeRF and Neural View Synthesis Methods for Front-Facing ViewsabstractAbstract Neural view synthesis (NVS) is one of the most successful techniques for synthesizing free viewpoint videos, capable of achieving high fidelity from only a sparse set of captured images. This success has led to many variants of the techniques, each evaluated on a set of test views typically using image quality metrics such as PSNR, SSIM, or LPIPS. There has been a lack of research on how NVS methods perform with respect to perceived video quality. We present the first study on perceptual evaluation of NVS and NeRF variants. For this study, we collected two datasets of scenes captured in a controlled lab environment as well as in‐the‐wild. In contrast to existing datasets, these scenes come with reference video sequences, allowing us to test for temporal artifacts and subtle distortions that are easily overlooked when viewing only static images. We measured the quality of videos synthesized by several NVS methods in a well‐controlled perceptual quality assessment experiment as well as with many existing state‐of‐the‐art image/video quality metrics. We present a detailed analysis of the results and recommendations for dataset and metric selection for NVS evaluation. Hanxue Liang, Tianhao Wu 0003, Param Hanji, Francesco Banterle, Hongyun Gao 0001, Rafal Mantiuk, A. Cengiz Öztireli |
Comput. Graph. Forum | 6 |
| 2024 | AR-DAVID: Augmented Reality Display Artifact Video DatasetabstractThe perception of visual content in optical-see-through augmented reality (AR) devices is affected by the light coming from the environment. This additional light interacts with the content in a non-trivial manner because of the illusion of transparency, different focal depths, and motion parallax. To investigate the impact of environment light on display artifact visibility (such as blur or color fringes), we created the first subjective quality dataset targeted toward augmented reality displays. Our study consisted of 6 scenes, each affected by one of 6 distortions at two strength levels, seen against one of 3 background patterns shown at 2 luminance levels: 432 conditions in total. Our dataset shows that environment light has a much smaller masking effect than expected. Further, we show that this effect cannot be explained by compositing of the AR-content with the background using optical blending models. As a consequence, we demonstrate that existing video quality metrics perform worse than expected when predicting the perceived magnitude of degradation in AR displays, motivating further research. Alexandre Chapiro, Dongyeon Kim, Yuta Asano, Rafal Mantiuk |
ACM Trans. Graph. | 4 |
| 2024 | ColorVideoVDP: A visual difference predictor for image, video and display distortionsabstractColorVideoVDP is a video and image quality metric that models spatial and temporal aspects of vision for both luminance and color. The metric is built on novel psychophysical models of chromatic spatiotemporal contrast sensitivity and cross-channel contrast masking. It accounts for the viewing conditions, geometric, and photometric characteristics of the display. It was trained to predict common video-streaming distortions (e.g., video compression, rescaling, and transmission errors) and also 8 new distortion types related to AR/VR displays (e.g., light source and waveguide non-uniformities). To address the latter application, we collected our novel XR-Display-Artifact-Video quality dataset (XR-DAVID), comprised of 336 distorted videos. Extensive testing on XR-DAVID, as well as several datasets from the literature, indicate a significant gain in prediction performance compared to existing metrics. ColorVideoVDP opens the doors to many novel applications that require the joint automated spatiotemporal assessment of luminance and color distortions, including video streaming, display specification, and design, visual comparison of results, and perceptually-guided quality optimization. The code for the metric can be found at https://github.com/gfxdisp/ColorVideoVDP. Rafal Mantiuk, Param Hanji, Maliha Ashraf, Yuta Asano, Alexandre Chapiro |
ACM Trans. Graph. | 1 |
| 2023 | Comparison of Metrics for Predicting Image and Video Quality at Varying Viewing DistancesabstractViewing distance and display resolution have ar-guably a significant impact on perceived image quality; images seen on a mobile phone with high pixel density reveal fewer distortions than the same images seen on a large TV from a close distance. However, only a few image and video quality metrics account for the effect of viewing distance and resolution. Those that do, typically rely on contrast sensitivity functions (CSFs) of the visual system. Other metrics can be potentially adapted to different viewing distances by rescaling input images. In this paper, we investigate the performance of such adapted metrics together with those that natively account for viewing distance. The results for three testing datasets indicate that there is no evidence that the metrics based on the CSF outperform those that rely on rescaled images. Moreover, we found that both methods are not successful to account for the changes in quality introduced by the change in viewing distance. We conclude that accounting for viewing distances requires better models. Dounia Hammou, Lukas Krasula, Christos G. Bampis, Zhi Li 0001, Rafal Mantiuk |
MMSP | 5 |
| 2023 | The effect of display capabilities on the gloss consistency between real and virtual objectsabstractA faithful reproduction of gloss is inherently difficult because of the limited dynamic range, peak luminance, and 3D capabilities of display devices. This work investigates how the display capabilities affect gloss appearance with respect to a real-world reference object. To this end, we employ an accurate imaging pipeline to achieve a perceptual gloss match between a virtual and real object presented side-by-side on an augmented-reality high-dynamic-range (HDR) stereoscopic display, which has not been previously attained to this extent. Based on this precise gloss reproduction, we conduct a series of gloss matching experiments to study how gloss perception degrades based on individual factors: object albedo, display luminance, dynamic range, stereopsis, and tone mapping. We support the study with a detailed analysis of individual factors, followed by an in-depth discussion on the observed perceptual effects. Our experiments demonstrate that stereoscopic presentation has a limited effect on the gloss matching task on our HDR display. However, both reduced luminance and dynamic range of the display reduce the perceived gloss. This means that the visual system cannot compensate for the changes in gloss appearance across luminance (lack of gloss constancy), and the tone mapping operator should be carefully selected when reproducing gloss on a low dynamic range (LDR) display. Bin Chen 0019, Akshay Jindal, Michal Piovarci, Chao Wang 0037, Hans-Peter Seidel, Piotr Didyk, Karol Myszkowski, Ana Serrano, Rafal Mantiuk |
SIGGRAPH Asia | 9 |
| 2022 | How bright should a virtual object be to appear opaque in optical see-through AR?abstractReproduction of occlusions and opaque surfaces are the major challenges of additive optical see-through (OST) displays. This is because the user of an OST display sees a linear mixture of display and environment light, which creates an impression of transparency unless the displayed color is sufficiently bright. The primary goal of this work is to determine how bright a displayed surface needs to be in relation to environment light to be perceived as opaque. We test multiple factors that could affect the perception of opacity: background luminance, contrast, spatial frequency, and accommodation depth in foveal vision. The subjective results, collected on a high-dynamic-range multi-focal stereo display, indicate that a virtual object needs to be, on average, 60 times brighter than the background environment light to be perceived as opaque. A higher contrast of the texture of the virtual object and a background that is out of focus can reduce the required luminance ratio. We demonstrate that a model of visual perception based on Weber’s law and accounting for contrast masking and defocus blur can predict the experimental data with an averaged prediction error of 8.29%. Existing perceptual image difference metrics (PSNR, FovVideoVDP and HDR-VDP-3) can also predict the effect of major factors, but with lower accuracy (e.g. prediction error of 34% for PSNR with PU21 encoding). Akshay Jindal, Claire Mantel, Søren Forchhammer, Rafal Mantiuk |
ISMAR | 5 |
| 2022 | Impact of correct and simulated focus cues on perceived realismabstractThe natural accommodation of the human eye to different distances results in focus cues, which contribute to depth perception and appearance. Since focus cues are very difficult to reproduce in an electronic display, it is desirable to know how much they contribute to realistic image appearance. In this work we quantify the potential benefit of focus cues in terms of increased realism compared to regular stereo image presentation. As a secondary goal, we evaluate whether three depth-of-field rendering techniques, which reproduce defocus blur at three different degrees of accuracy, can reintroduce the benefits of focus cues. Our findings confirm the importance of focus cues for realistic image appearance, and also show that they cannot easily be substituted by depth-of-field rendering. Joseph March, Anantha Krishnan, Simon J. Watt, Marek Wernikowski, Hongyun Gao 0001, Ali Özgür Yöntem, Rafal Mantiuk |
SIGGRAPH Asia | 7 |
| 2022 | Training a Task-Specific Image Reconstruction LossabstractThe choice of a loss function is an important factor when training neural networks for image restoration problems, such as single image super resolution. The loss function should encourage natural and perceptually pleasing results. A popular choice for a loss is a pre-trained network, such as VGG, which is used as a feature extractor for computing the difference between restored and reference images. However, such an approach has multiple drawbacks: it is computationally expensive, requires regularization and hyper-parameter tuning, and involves a large network trained on an unrelated task. Furthermore, it has been observed that there is no single loss function that works best across all applications and across different datasets. In this work, we instead propose to train a set of loss functions that are application specific in nature. Our loss function comprises a series of discriminators that are trained to detect and penalize the presence of application-specific artifacts. We show that a single natural image and corresponding distortions are sufficient to train our feature extractor that outperforms state-of-the-art loss functions in applications like single image super resolution, denoising, and JPEG artifact removal. Finally, we conclude that an effective loss function does not have to be a good predictor of perceived image quality, but instead needs to be specialized in identifying the distortions for a given restoration method. Aamir Mustafa, Aliaksei Mikhailiuk, Dan-Andrei Iliescu, Varun Babbar, Rafal Mantiuk |
WACV | 5 |
| 2022 | Consolidated Dataset and Metrics for High-Dynamic-Range Image QualityabstractIncreasing popularity of high-dynamic-range (HDR) image and video content brings the need for metrics that could predict the severity of image impairments as seen on displays of different brightness levels and dynamic range. Such metrics should be trained and validated on a sufficiently large subjective image quality dataset to ensure robust performance. As the existing HDR quality datasets are limited in size, we created a Unified Photometric Image Quality dataset (UPIQ) with over 4000 images by realigning and merging existing HDR and standard-dynamic-range (SDR) datasets. The realigned quality scores share the same unified quality scale across all datasets. Such realignment was achieved by collecting additional cross-dataset quality comparisons and re-scaling data with a psychometric scaling method. Images in the proposed dataset are represented in absolute photometric and colorimetric units, corresponding to light emitted from a display. We use the new dataset to retrain existing HDR metrics and show that the dataset is sufficiently large for training deep architectures. We show the utility of the dataset on brightness aware image compression. Aliaksei Mikhailiuk, María Pérez-Ortiz 0001, Dingcheng Yue, Wilson Suen, Rafal Mantiuk |
IEEE Trans. Multim. | 5 |
| 2022 | stelaCSF: a unified model of contrast sensitivity as the function of spatio-temporal frequency, eccentricity, luminance and areaabstractA contrast sensitivity function, or CSF, is a cornerstone of many visual models. It explains whether a contrast pattern is visible to the human eye. The existing CSFs typically account for a subset of relevant dimensions describing a stimulus, limiting the use of such functions to either static or foveal content but not both. In this paper, we propose a unified CSF, stelaCSF, which accounts for all major dimensions of the stimulus: spatial and temporal frequency, eccentricity, luminance, and area. To model the 5-dimensional space of contrast sensitivity, we combined data from 11 papers, each of which studied a subset of this space. While previously proposed CSFs were fitted to a single dataset, stelaCSF can predict the data from all these studies using the same set of parameters. The predictions are accurate in the entire domain, including low frequencies. In addition, stelaCSF relies on psychophysical models and experimental evidence to explain the major interactions between the 5 dimensions of the CSF. We demonstrate the utility of our new CSF in a flicker detection metric and in foveated rendering. Rafal Mantiuk, Maliha Ashraf, Alexandre Chapiro |
ACM Trans. Graph. | 1 |
| 2022 | Dark stereo: improving depth perception under low luminanceabstractIt is often desirable or unavoidable to display Virtual Reality (VR) or stereoscopic content at low brightness. For example, a dimmer display reduces the flicker artefacts that are introduced by low-persistence VR headsets. It also saves power, prolongs battery life, and reduces the cost of a display or projection system. Additionally, stereo movies are usually displayed at relatively low luminance due to polarization filters or other optical elements necessary to separate two views. However, the binocular depth cues become less reliable at low luminance. In this paper, we propose a model of stereo constancy that predicts the precision of binocular depth cues for a given contrast and luminance. We use the model to design a novel contrast enhancement algorithm that compensates for the deteriorated depth perception to deliver good-quality stereoscopic images even for displays of very low brightness. Krzysztof Wolski, Fangcheng Zhong, Karol Myszkowski, Rafal Mantiuk |
ACM Trans. Graph. | 4 |
| 2021 | PU21: A novel perceptually uniform encoding for adapting existing quality metrics for HDRabstractStandard image quality metrics, such as PSNR or SSIM, cannot be directly computed on linear high dynamic range colour values because such values non-linearly related to our perception of visible differences. In this work, we develop a new encoding function (PU21) to convert absolute high dynamic range (HDR) linear colour values into approximately perceptually uniform (PU) values, which can be used with standard quality metrics. The proposed PU21 function is based on a recent contrast sensitivity model, fitted to the measurements up to 10000cd/m2. Unlike the conventional simplified approach of deriving PU functions based on peak sensitivities, we model realistic coding artefacts to find visibility thresholds for our derivation. Furthermore, the new PU accounts for the effect of glare on image quality. The proposed PU21 improves the accuracy of quality predictions for standard metrics of PSNR, VSI, FSIM, SSIM, and MS-SSIM in their correlation with subjective scores on HDR images included in UPIQ, one of the largest HDR image quality datasets. Rafal Mantiuk, Maryam Azimi |
PCS | 1 |
| 2021 | Perceptual model for adaptive local shading and refresh rateabstractWhen the rendering budget is limited by power or time, it is necessary to find the combination of rendering parameters, such as resolution and refresh rate, that could deliver the best quality. Variable-rate shading (VRS), introduced in the last generations of GPUs, enables fine control of the rendering quality, in which each 16×16 image tile can be rendered with a different ratio of shader executions. We take advantage of this capability and propose a new method for adaptive control of local shading and refresh rate. The method analyzes texture content, on-screen velocities, luminance, and effective resolution and suggests the refresh rate and a VRS state map that maximizes the quality of animated content under a limited budget. The method is based on the new content-adaptive metric of judder, aliasing, and blur, which is derived from the psychophysical models of contrast sensitivity. To calibrate and validate the metric, we gather data from literature and also collect new measurements of motion quality under variable shading rates, different velocities of motion, texture content, and display capabilities, such as refresh rate, persistence, and angular resolution. The proposed metric and adaptive shading method is implemented as a game engine plugin. Our experimental validation shows a substantial increase in preference of our method over rendering with a fixed resolution and refresh rate, and an existing motion-adaptive technique. Akshay Jindal, Krzysztof Wolski, Karol Myszkowski, Rafal Mantiuk |
ACM Trans. Graph. | 4 |
| 2021 | FovVideoVDP: a visible difference predictor for wide field-of-view videoabstractFovVideoVDP is a video difference metric that models the spatial, temporal, and peripheral aspects of perception. While many other metrics are available, our work provides the first practical treatment of these three central aspects of vision simultaneously. The complex interplay between spatial and temporal sensitivity across retinal locations is especially important for displays that cover a large field-of-view, such as Virtual and Augmented Reality displays, and associated methods, such as foveated rendering. Our metric is derived from psychophysical studies of the early visual system, which model spatio-temporal contrast sensitivity, cortical magnification and contrast masking. It accounts for physical specification of the display (luminance, size, resolution) and viewing distance. To validate the metric, we collected a novel foveated rendering dataset which captures quality degradation due to sampling and reconstruction. To demonstrate our algorithm's generality, we test it on 3 independent foveated video datasets, and on a large image quality dataset, achieving the best performance across all datasets when compared to the state-of-the-art. Rafal Mantiuk, Gyorgy Denes, Alexandre Chapiro, Anton Kaplanyan, Gizem Rufo, Romain Bachy, Trisha Lian, Anjul Patney |
ACM Trans. Graph. | 1 |
| 2021 | Reproducing reality with a high-dynamic-range multi-focal stereo displayabstractWith well-established methods for producing photo-realistic results, the next big challenge of graphics and display technologies is to achieve perceptual realism --- producing imagery indistinguishable from real-world 3D scenes. To deliver all necessary visual cues for perceptual realism, we built a High-Dynamic-Range Multi-Focal Stereo Display that achieves high resolution, accurate color, a wide dynamic range, and most depth cues, including binocular presentation and a range of focal depth. The display and associated imaging system have been designed to capture and reproduce a small near-eye three-dimensional object and to allow for a direct comparison between virtual and real scenes. To assess our reproduction of realism and demonstrate the capability of the display and imaging system, we conducted an experiment in which the participants were asked to discriminate between a virtual object and its physical counterpart. Our results indicate that the participants can only detect the discrepancy with a probability of 0.44. With such a level of perceptual realism, our display apparatus can facilitate a range of visual experiments that require the highest fidelity of reproduction while allowing for the full control of the displayed stimuli. Fangcheng Zhong, Akshay Jindal, Ali Özgür Yöntem, Param Hanji, Simon J. Watt, Rafal Mantiuk |
ACM Trans. Graph. | 6 |
| 2020 | Transformation Consistency Regularization - A Semi-supervised Paradigm for Image-to-Image Translation
Aamir Mustafa, Rafal Mantiuk |
ECCV (18) | 2 |
| 2020 | Active Sampling for Pairwise Comparisons via Approximate Message Passing and Information Gain MaximizationabstractPairwise comparison data arise in many domains with subjective assessment experiments, for example in image and video quality assessment. In these experiments observers are asked to express a preference between two conditions. However, many pairwise comparison protocols require a large number of comparisons to infer accurate scores, which may be unfeasible when each comparison is time-consuming (e.g. videos) or expensive (e.g. medical imaging). This motivates the use of an active sampling algorithm that chooses only the most informative pairs for comparison. In this paper we propose ASAP, an active sampling algorithm based on approximate message passing and expected information gain maximization. Unlike most existing methods, which rely on partial updates of the posterior distribution, we are able to perform full updates and therefore much improve the accuracy of the inferred scores. The algorithm relies on three techniques for reducing computational cost: inference based on approximate message passing, selective evaluations of the information gain, and selecting pairs in a batch that forms a minimum spanning tree of the inverse of information gain. We demonstrate, with real and synthetic data, that ASAP offers the highest accuracy of inferred scores compared to the existing methods. We also provide an open-source GPU implementation of ASAP for large-scale experiments. Aliaksei Mikhailiuk, Clifford Wilmot, María Pérez-Ortiz 0001, Dingcheng Yue, Rafal Mantiuk |
ICPR | 5 |
| 2020 | From Pairwise Comparisons and Rating to a Unified Quality ScaleabstractThe goal of psychometric scaling is the quantification of perceptual experiences, understanding the relationship between an external stimulus, the internal representation and the response. In this paper, we propose a probabilistic framework to fuse the outcome of different psychophysical experimental protocols, namely rating and pairwise comparisons experiments. Such a method can be used for merging existing datasets of subjective nature and for experiments in which both measurements are collected. We analyze and compare the outcomes of both types of experimental protocols in terms of time and accuracy in a set of simulations and experiments with benchmark and real-world image quality assessment datasets, showing the necessity of scaling and the advantages of each protocol and mixing. Although most of our examples focus on image quality assessment, our findings generalize to any other subjective quality-of-experience task. María Pérez-Ortiz 0001, Aliaksei Mikhailiuk, Emin Zerman, Vedad Hulusic, Giuseppe Valenzise, Rafal Mantiuk |
IEEE Trans. Image Process. | 6 |
| 2020 | A perceptual model of motion quality for rendering with adaptive refresh-rate and resolutionabstractLimited GPU performance budgets and transmission bandwidths mean that real-time rendering often has to compromise on the spatial resolution or temporal resolution (refresh rate). A common practice is to keep either the resolution or the refresh rate constant and dynamically control the other variable. But this strategy is non-optimal when the velocity of displayed content varies. To find the best trade-off between the spatial resolution and refresh rate, we propose a perceptual visual model that predicts the quality of motion given an object velocity and predictability of motion. The model considers two motion artifacts to establish an overall quality score: non-smooth (juddery) motion, and blur. Blur is modeled as a combined effect of eye motion, finite refresh rate and display resolution. To fit the free parameters of the proposed visual model, we measured eye movement for predictable and unpredictable motion, and conducted psychophysical experiments to measure the quality of motion from 50 Hz to 165 Hz. We demonstrate the utility of the model with our on-the-fly motion-adaptive rendering algorithm that adjusts the refresh rate of a G-Sync-capable monitor based on a given rendering budget and observed object motion. Our psychophysical validation experiments demonstrate that the proposed algorithm performs better than constant-refresh-rate solutions, showing that motion-adaptive rendering is an attractive technique for driving variable-refresh-rate displays. Gyorgy Denes, Akshay Jindal, Aliaksei Mikhailiuk, Rafal Mantiuk |
ACM Trans. Graph. | 4 |
| 2019 | Exploiting Synthetically Generated Data with Semi-Supervised Learning for Small and Imbalanced DatasetsabstractData augmentation is rapidly gaining attention in machine learning. Synthetic data can be generated by simple transformations or through the data distribution. In the latter case, the main challenge is to estimate the label associated to new synthetic patterns. This paper studies the effect of generating synthetic data by convex combination of patterns and the use of these as unsupervised information in a semi-supervised learning framework with support vector machines, avoiding thus the need to label synthetic examples. We perform experiments on a total of 53 binary classification datasets. Our results show that this type of data over-sampling supports the well-known cluster assumption in semi-supervised learning, showing outstanding results for small high-dimensional datasets and imbalanced learning problems. María Pérez-Ortiz 0001, Peter Tiño, Rafal Mantiuk, César Hervás-Martínez |
AAAI | 3 |
| 2019 | Single-Frame Regularization for Temporally Stable CNNsabstractConvolutional neural networks (CNNs) can model complicated non-linear relations between images. However, they are notoriously sensitive to small changes in the input. Most CNNs trained to describe image-to-image mappings generate temporally unstable results when applied to video sequences, leading to flickering artifacts and other inconsistencies over time. In order to use CNNs for video material, previous methods have relied on estimating dense frame-to-frame motion information (optical flow) in the training and/or the inference phase, or by exploring recurrent learning structures. We take a different approach to the problem, posing temporal stability as a regularization of the cost function. The regularization is formulated to account for different types of motion that can occur between frames, so that temporally stable CNNs can be trained without the need for video material or expensive motion estimation. The training can be performed as a fine-tuning operation, without architectural modifications of the CNN. Our evaluation shows that the training strategy leads to large improvements in temporal smoothness. Moreover, for small datasets the regularization can help in boosting the generalization performance to a much larger extent than what is possible with naive augmentation strategies. Gabriel Eilertsen, Rafal Mantiuk, Jonas Unger |
CVPR | 2 |
| 2019 | Predicting Visible Image Differences Under Varying Display Brightness and Viewing DistanceabstractNumerous applications require a robust metric that can predict whether image differences are visible or not. However, the accuracy of existing white-box visibility metrics, such as HDR-VDP, is often not good enough. CNN-based black-box visibility metrics have proven to be more accurate, but they cannot account for differences in viewing conditions, such as display brightness and viewing distance. In this paper, we propose a CNN-based visibility metric, which maintains the accuracy of deep network solutions and accounts for viewing conditions. To achieve this, we extend the existing dataset of locally visible differences (LocVis) with a new set of measurements, collected considering aforementioned viewing conditions. Then, we develop a hybrid model that combines white-box processing stages for modeling the effects of luminance masking and contrast sensitivity, with a black-box deep neural network. We demonstrate that the novel hybrid model can handle the change of viewing conditions correctly and outperforms state-of-the-art metrics. Nanyang Ye 0001, Krzysztof Wolski, Rafal Mantiuk |
CVPR | 3 |
| 2019 | Visibility Metric for Visually Lossless Image CompressionabstractEncoding images in a visually lossless manner helps to achieve the best trade-off between image compression performance and quality and so that compression artifacts are invisible to the majority of users. Visually lossless encoding can often be achieved by manually adjusting compression quality parameters of existing lossy compression methods, such as JPEG or WebP. But the required compression quality parameter can also be determined automatically using visibility metrics. However, creating an accurate visibility metric is challenging because of the complexity of the human visual system and the effort needed to collect the required data. In this paper, we investigate how to train an accurate visibility metric for visually lossless compression from a relatively small dataset. Our experiments show that prediction error can be reduced by 40% compared with the state-of-theart, and that our proposed method can save between 25%-75% of storage space compared with the default quality parameter used in commercial software. We demonstrate how the visibility metric can be used for visually lossless image compression and for benchmarking image compression encoders. Nanyang Ye 0001, María Pérez-Ortiz 0001, Rafal Mantiuk |
PCS | 3 |
| 2019 | Near-Eye Display and Tracking Technologies for Virtual and Augmented RealityabstractAbstract Virtual and augmented reality (VR/AR) are expected to revolutionise entertainment, healthcare, communication and the manufacturing industries among many others. Near‐eye displays are an enabling vessel for VR/AR applications, which have to tackle many challenges related to ergonomics, comfort, visual quality and natural interaction. These challenges are related to the core elements of these near‐eye display hardware and tracking technologies. In this state‐of‐the‐art report, we investigate the background theory of perception and vision as well as the latest advancements in display engineering and tracking technologies. We begin our discussion by describing the basics of light and image formation. Later, we recount principles of visual perception by relating to the human visual system. We provide two structured overviews on state‐of‐the‐art near‐eye display and tracking technologies involved in such near‐eye displays. We conclude by outlining unresolved research questions to inspire the next generation of researchers. George Alex Koulieris, Kaan Aksit, Michael Stengel, Rafal Mantiuk, Katerina Mania, Christian Richardt |
Comput. Graph. Forum | 4 |
| 2019 | Selecting texture resolution using a task-specific visibility metricabstractAbstract In real‐time rendering, the appearance of scenes is greatly affected by the quality and resolution of the textures used for image synthesis. At the same time, the size of textures determines the performance and the memory requirements of rendering. As a result, finding the optimal texture resolution is critical, but also a non‐trivial task since the visibility of texture imperfections depends on underlying geometry, illumination, interactions between several texture maps, and viewing positions. Ideally, we would like to automate the task with a visibility metric, which could predict the optimal texture resolution. To maximize the performance of such a metric, it should be trained on a given task. This, however, requires sufficient user data which is often difficult to obtain. To address this problem, we develop a procedure for training an image visibility metric for a specific task while reducing the effort required to collect new data. The procedure involves generating a large dataset using an existing visibility metric followed by refining that dataset with the help of an efficient perceptual experiment. Then, such a refined dataset is used to retune the metric. This way, we augment sparse perceptual data to a large number of per‐pixel annotated visibility maps which serve as the training data for application‐specific visibility metrics. While our approach is general and can be potentially applied for different image distortions, we demonstrate an application in a game‐engine where we optimize the resolution of various textures, such as albedo and normal maps. Krzysztof Wolski, Daniele Giunchi, Shinichi Kinuwaki, Piotr Didyk, Karol Myszkowski, Anthony Steed, Rafal Mantiuk |
Comput. Graph. Forum | 7 |
| 2019 | DiCE: dichoptic contrast enhancement for VR and stereo displaysabstractIn stereoscopic displays, such as those used in VR/AR headsets, our eyes are presented with two different views. The disparity between the views is typically used to convey depth cues, but it could be also used to enhance image appearance. We devise a novel technique that takes advantage of binocular fusion to boost perceived local contrast and visual quality of images. Since the technique is based on fixed tone curves, it has negligible computational cost and it is well suited for real-time applications, such as VR rendering. To control the trade-off between contrast gain and binocular rivalry, we conduct a series of experiments to explain the factors that dominate rivalry perception in a dichoptic presentation where two images of different contrasts are displayed. With this new finding, we can effectively enhance contrast and control rivalry in mono- and stereoscopic images, and in VR rendering, as confirmed in validation experiments. Fangcheng Zhong, George Alex Koulieris, George Drettakis, Martin S. Banks, Mathieu Chambe, Frédo Durand, Rafal Mantiuk |
ACM Trans. Graph. | 7 |
| 2019 | Temporal Resolution Multiplexing: Exploiting the limitations of spatio-temporal vision for more efficient VR renderingabstractRendering in virtual reality (VR) requires substantial computational power to generate 90 frames per second at high resolution with good-quality antialiasing. The video data sent to a VR headset requires high bandwidth, achievable only on dedicated links. In this paper we explain how rendering requirements and transmission bandwidth can be reduced using a conceptually simple technique that integrates well with existing rendering pipelines. Every even-numbered frame is rendered at a lower resolution, and every odd-numbered frame is kept at high resolution but is modified in order to compensate for the previous loss of high spatial frequencies. When the frames are seen at a high frame rate, they are fused and perceived as high-resolution and high-frame-rate animation. The technique relies on the limited ability of the visual system to perceive high spatio-temporal frequencies. Despite its conceptual simplicity, correct execution of the technique requires a number of non-trivial steps: display photometric temporal response must be modeled, flicker and motion artifacts must be avoided, and the generated signal must not exceed the dynamic range of the display. Our experiments, performed on a high-frame-rate LCD monitor and OLED-based VR headsets, explore the parameter space of the proposed technique and demonstrate that its perceived quality is indistinguishable from full-resolution rendering. The technique is an attractive alternative to reprojection and resolution reduction of all frames. Gyorgy Denes, Kuba Maruszczyk, George Ash, Rafal Mantiuk |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2018 | Trained Perceptual Transform for Quality Assessment of High Dynamic Range Images and VideoabstractIn this paper, we propose a trained perceptually transform for quality assessment of high dynamic range (HDR) images and video. The transform is used to convert absolute luminance values found in HDR images into perceptually uniform units, which can be used with any standard-dynamic-range metric. The new transform is derived by fitting the parameters of a previously proposed perceptual encoding function to 4 different HDR subjective quality assessment datasets using Bayesian optimization. The new transform combined with a simple peak signal-to-noise ratio measure achieves better prediction performance in cross-dataset validation than existing transforms. We provide Matlab code for our metric11https://github.com/ynyCL/T-PT-metric. Nanyang Ye 0001, María Pérez-Ortiz 0001, Rafal Mantiuk |
ICIP | 3 |
| 2018 | Hybrid-MST: A Hybrid Active Sampling Strategy for Pairwise Preference AggregationabstractIn this paper we present a hybrid active sampling strategy for pairwise preference aggregation, which aims at recovering the underlying rating of the test candidates from sparse and noisy pairwise labeling. Our method employs Bayesian optimization framework and Bradley-Terry model to construct the utility function, then to obtain the Expected Information Gain (EIG) of each pair. For computational efficiency, Gaussian-Hermite quadrature is used for estimation of EIG. In this work, a hybrid active sampling strategy is proposed, either using Global Maximum (GM) EIG sampling or Minimum Spanning Tree (MST) sampling in each trial, which is determined by the test budget. The proposed method has been validated on both simulated and real-world datasets, where it shows higher preference aggregation ability than the state-of-the-art methods. Jing Li 0026, Rafal Mantiuk, Junle Wang, Suiyi Ling, Patrick Le Callet |
NeurIPS | 2 |
| 2018 | Psychometric scaling of TID2013 datasetabstractTID2013 is a subjective image quality assessment dataset with a wide range of distortion types and over 3000 images. The dataset has proven to be a challenging test for objective quality metrics. The dataset mean opinion scores were obtained by collecting pairwise comparison judgments using the Swiss tournament system, and averaging votes of observers. However, this approach differs from the usual analysis of multiple pairwise comparisons, which involves psychometric scaling of the comparison data using either Thurstone or Bradley-Terry models. In this paper we investigate how quality scores change when they are computed using such psychometric scaling instead of averaging vote counts. In order to properly scale TID2013 quality scores, we conduct four additional experiments of two different types, which we found necessary to produce a common quality scale: comparisons with reference images, and cross-content comparisons. We demonstrate on a fifth validation experiment that the two additional types of comparisons are necessary and in conjunction with psychometric scaling improve the consistency of quality scores, especially across images depicting different contents. Aliaksei Mikhailiuk, María Pérez-Ortiz 0001, Rafal Mantiuk |
QoMEX | 3 |
| 2018 | Dataset and Metrics for Predicting Local Visible DifferencesabstractA large number of imaging and computer graphics applications require localized information on the visibility of image distortions. Existing image quality metrics are not suitable for this task as they provide a single quality value per image. Existing visibility metrics produce visual difference maps, and are specifically designed for detecting just noticeable distortions but their predictions are often inaccurate. In this work, we argue that the key reason for this problem is the lack of large image collections with a good coverage of possible distortions that occur in different applications. To address the problem, we collect an extensive dataset of reference and distorted image pairs together with user markings indicating whether distortions are visible or not. We propose a statistical model that is designed for the meaningful interpretation of such data, which is affected by visual search and imprecision of manual marking. We use our dataset for training existing metrics and we demonstrate that their performance significantly improves. We show that our dataset with the proposed statistical model can be used to train a new CNN-based metric, which outperforms the existing solutions. We demonstrate the utility of such a metric in visually lossless JPEG compression, super-resolution and watermarking. Krzysztof Wolski, Daniele Giunchi, Nanyang Ye 0001, Piotr Didyk, Karol Myszkowski, Radoslaw Mantiuk, Hans-Peter Seidel, Anthony Steed, Rafal Mantiuk |
ACM Trans. Graph. | 9 |
| 2017 | Towards a Quality Metric for Dense Light FieldsabstractLight fields become a popular representation of three-dimensional scenes, and there is interest in their processing, resampling, and compression. As those operations often result in loss of quality, there is a need to quantify it. In this work, we collect a new dataset of dense reference and distorted light fields as well as the corresponding quality scores which are scaled in perceptual units. The scores were acquired in a subjective experiment using an interactive light-field viewing setup. The dataset contains typical artifacts that occur in light-field processing chain due to light-field reconstruction, multi-view compression, and limitations of automultiscopic displays. We test a number of existing objective quality metrics to determine how well they can predict the quality of light fields. We find that the existing image quality metrics provide good measures of light-field quality, but require dense reference light-fields for optimal performance. For more complex tasks of comparing two distorted light fields, their performance drops significantly, which reveals the need for new, light-field-specific metrics. Vamsi Kiran Adhikarla, Marek Vinkler, Denis Sumin, Rafal Mantiuk, Karol Myszkowski, Hans-Peter Seidel, Piotr Didyk |
CVPR | 4 |
| 2017 | Langevin Dynamics with Continuous Tempering for Training Deep Neural NetworksabstractMinimizing non-convex and high-dimensional objective functions is challenging, especially when training modern deep neural networks. In this paper, a novel approach is proposed which divides the training process into two consecutive phases to obtain better generalization performance: Bayesian sampling and stochastic optimization. The first phase is to explore the energy landscape and to capture the `fat'' modes; and the second one is to fine-tune the parameter learned from the first phase. In the Bayesian learning phase, we apply continuous tempering and stochastic approximation into the Langevin dynamics to create an efficient and effective sampler, in which the temperature is adjusted automatically according to the designed ``temperature dynamics''. These strategies can overcome the challenge of early trapping into bad local minima and have achieved remarkable improvements in various types of neural networks as shown in our theoretical analysis and empirical experiments. Nanyang Ye 0001, Zhanxing Zhu, Rafal Mantiuk |
NIPS | 3 |
| 2017 | Effect of color space on high dynamic range video compression performanceabstractHigh dynamic range (HDR) technology allows for capturing and delivering a greater range of luminance levels compared to traditional video using standard dynamic range (SDR). At the same time, it has brought multiple challenges in content distribution, one of them being video compression. While there has been a significant amount of work conducted on this topic, there are some aspects that could still benefit this area. One such aspect is the choice of color space used for coding. In this paper, we evaluate through a subjective study how the performance of HDR video compression is affected by three color spaces: the commonly used Y'CbCr, and the recently introduced ITP (ICtCp) and Ypu'v'. Five video sequences are compressed at four bit rates, selected in a preliminary study, and their quality is assessed using pairwise comparisons. The results of pairwise comparisons are further analyzed and scaled to obtain quality scores. We found no evidence of ITP improving compression performance over Y'CbCr. We also found that Ypu'v' results in a moderately lower performance for some sequences. Emin Zerman, Vedad Hulusic, Giuseppe Valenzise, Rafal Mantiuk, Frédéric Dufaux |
QoMEX | 4 |
| 2017 | Assessment of multi-exposure HDR image deghosting methods
Kanita Karaduzovic Hadziabdic, Jasminka Hasic, Rafal Mantiuk |
Comput. Graph. | 3 |
| 2017 | A comparative review of tone-mapping algorithms for high dynamic range videoabstractTone-mapping constitutes a key component within the field of high dynamic range (HDR) imaging. Its importance is manifested in the vast amount of tone-mapping methods that can be found in the literature, which are the result of an active development in the area for more than two decades. Although these can accommodate most requirements for display of HDR images, new challenges arose with the advent of HDR video, calling for additional considerations in the design of tone-mapping operators (TMOs). Today, a range of TMOs exist that do support video material. We are now reaching a point where most camera captured HDR videos can be prepared in high quality without visible artifacts, for the constraints of a standard display device. In this report, we set out to summarize and categorize the research in tone-mapping as of today, distilling the most important trends and characteristics of the tone reproduction pipeline. While this gives a wide overview over the area, we then specifically focus on tone-mapping of HDR video and the problems this medium entails. First, we formulate the major challenges a video TMO needs to address. Then, we provide a description and categorization of each of the existing video TMOs. Finally, by constructing a set of quantitative measures, we evaluate the performance of a number of the operators, in order to give a hint on which can be expected to render the least amount of artifacts. This serves as a comprehensive reference, categorization and comparative assessment of the state-of-the-art in tone-mapping for HDR video. Gabriel Eilertsen, Rafal Mantiuk, Jonas Unger |
Comput. Graph. Forum | 2 |
| 2017 | HDR image reconstruction from a single exposure using deep CNNsabstractCamera sensors can only capture a limited range of luminance simultaneously, and in order to create high dynamic range (HDR) images a set of different exposures are typically combined. In this paper we address the problem of predicting information that have been lost in saturated image areas, in order to enable HDR reconstruction from a single exposure. We show that this problem is well-suited for deep learning algorithms, and propose a deep convolutional neural network (CNN) that is specifically designed taking into account the challenges in predicting HDR values. To train the CNN we gather a large dataset of HDR images, which we augment by simulating sensor saturation for a range of cameras. To further boost robustness, we pre-train the CNN on a simulated HDR dataset created from a subset of the MIT Places database. We demonstrate that our approach can reconstruct high-resolution visually convincing HDR results in a wide range of situations, and that it generalizes well to reconstruction of images captured with arbitrary and low-end cameras that use unknown camera response functions and post-processing. Furthermore, we compare to existing methods for HDR expansion, and show high quality results also for image based lighting. Finally, we evaluate the results in a subjective experiment performed on an HDR display. This shows that the reconstructed HDR images are visually convincing, with large improvements as compared to existing methods. Gabriel Eilertsen, Joel Kronander, Gyorgy Denes, Rafal Mantiuk, Jonas Unger |
ACM Trans. Graph. | 4 |
| 2017 | Analysis of reported error in Monte Carlo rendered imagesabstractEvaluating image quality in Monte Carlo rendered images is an important aspect of the rendering process as we often need to determine the relative quality between images computed using different algorithms and with varying amounts of computation. The use of a gold-standard, reference image, or ground truth is a common method to provide a baseline with which to compare experimental results. We show that if not chosen carefully, the quality of reference images used for image quality assessment can skew results leading to significant misreporting of error. We present an analysis of error in Monte Carlo rendered images and discuss practices to avoid or be aware of when designing an experiment. Joss Whittle, Mark W. Jones 0001, Rafal Mantiuk |
Vis. Comput. | 3 |
| 2016 | Real-time noise-aware tone-mapping and its use in luminance retargetingabstractWith the aid of tone-mapping operators, high dynamic range images can be mapped for reproduction on standard displays. However, for large restrictions in terms of display dynamic range and peak luminance, limitations of the human visual system have significant impact on the visual appearance. In this paper, we use components from the real-time noise-aware tone-mapping to complement an existing method for perceptual matching of image appearance under different luminance levels. The refined luminance retargeting method improves subjective quality on a display with large limitations in dynamic range, as suggested by our subjective evaluation. Gabriel Eilertsen, Rafal Mantiuk, Jonas Unger |
ICIP | 2 |
| 2016 | A high dynamic range video codec optimized by large-scale testingabstractWhile a number of existing high-bit depth video compression methods can potentially encode high dynamic range (HDR) video, few of them provide this capability. In this paper, we investigate techniques for adapting HDR video for this purpose. In a large-scale test on 33 HDR video sequences, we compare 2 video codecs, 4 luminance encoding techniques (transfer functions) and 3 color encoding methods, measuring quality in terms of two objective metrics, PU-MSSIM and HDR-VDP-2. From the results we design an open source HDR video encoder, optimized for the best compression performance given the techniques examined. Gabriel Eilertsen, Rafal Mantiuk, Jonas Unger |
ICIP | 2 |
| 2016 | Practicalities of predicting quality of high dynamic range images and videoabstractThe paper discusses the use of existing metrics, such as HDR-VDP and extensions of MS-SSIM and PSNR, for prediction of quality in high dynamic range (HDR) images and video. The discussion is based on the experience in using those metrics to evaluate and improve image compression for the new JPEG XT standard, and video compression for the LumaHDR open source codec. The paper explains why existing non-HDR metrics perform very poorly on HDR data and how to improve their predictions. Since most HDR metrics require calibrated data, intended for an HDR display, such calibration step is explained. One of the popular HDR quality metrics, HDR-VDP, is briefly introduced with the update on the latest improvements. Finally, several studies comparing objective HDR metric performance are summarized. Rafal Mantiuk |
ICIP | 1 |
| 2016 | Fine-tuning JPEG-XT compression performance using large-scale objective quality testingabstractThe upcoming JPEG XT standard for High Dynamic Range (HDR) images defines a common framework for the lossy and lossless representation of high-dynamic range images. It describes the decoding process as the combination of various processing tools that can be combined freely. In this paper we analyze the coding efficiency of different decoding tools through a large scale objective quality testing using the HDR-VDP 2.2 objective metric. This evaluation is performed on a large database of 337 images, testing the effect of global and local tone mapping operators for various configurations, and for multiple combinations of quality parameters. The main findings are that using an inverse tone mapping operator for creating an HDR precursor image works well for global, but not for local operators, and that including refinement scans to increase the bit-depth of the extension layer provides substantial improvements for one of the encoding profiles and higher bit-rates. Rafal Mantiuk, Thomas Richter 0005, Alessandro Artusi |
ICIP | 1 |
| 2016 | Objective and subjective evaluation of High Dynamic Range video compression
Ratnajit Mukherjee, Kurt Debattista, Thomas Bashford-Rogers, Peter Vangorp, Rafal Mantiuk, Maximino Bessa, Brian Waterfield, Alan Chalmers |
Signal Process. Image Commun. | 5 |
| 2015 | Introduction to Special Issue SAP 2015abstractNo abstract available. Scott Kuhl, Rafal Mantiuk, Betsy Williams Sanders |
ACM Trans. Appl. Percept. | 2 |
| 2015 | Real-time noise-aware tone mappingabstractReal-time high quality video tone mapping is needed for many applications, such as digital viewfinders in cameras, display algorithms which adapt to ambient light, in-camera processing, rendering engines for video games and video post-processing. We propose a viable solution for these applications by designing a video tone-mapping operator that controls the visibility of the noise, adapts to display and viewing environment, minimizes contrast distortions, preserves or enhances image details, and can be run in real-time on an incoming sequence without any preprocessing. To our knowledge, no existing solution offers all these features. Our novel contributions are: a fast procedure for computing local display-adaptive tone-curves which minimize contrast distortions, a fast method for detail enhancement free from ringing artifacts, and an integrated video tone-mapping solution combining all the above features. Gabriel Eilertsen, Rafal Mantiuk, Jonas Unger |
ACM Trans. Graph. | 2 |
| 2015 | A model of local adaptationabstractThe visual system constantly adapts to different luminance levels when viewing natural scenes. The state of visual adaptation is the key parameter in many visual models. While the time-course of such adaptation is well understood, there is little known about the spatial pooling that drives the adaptation signal. In this work we propose a new empirical model of local adaptation, that predicts how the adaptation signal is integrated in the retina. The model is based on psychophysical measurements on a high dynamic range (HDR) display. We employ a novel approach to model discovery, in which the experimental stimuli are optimized to find the most predictive model. The model can be used to predict the steady state of adaptation, but also conservative estimates of the visibility (detection) thresholds in complex images. We demonstrate the utility of the model in several applications, such as perceptual error bounds for physically based rendering, determining the backlight resolution for HDR displays, measuring the maximum visible dynamic range in natural scenes, simulation of afterimages, and gaze-dependent tone mapping. Peter Vangorp, Karol Myszkowski, Erich W. Graf, Rafal Mantiuk |
ACM Trans. Graph. | 4 |
| 2014 | Depth from HDR: depth induction or increased realism?abstractMany people who first see a high dynamic range (HDR) display get the impression that it is a 3D display, even though it does not produce any binocular depth cues. Possible explanations of this effect include contrast-based depth induction and the increased realism due to the high brightness and contrast that makes an HDR display "like looking through a window". In this paper we test both of these hypotheses by comparing the HDR depth illusion to real binocular depth cues using a carefully calibrated HDR stereoscope. We confirm that contrast-based depth induction exists, but it is a vanishingly weak depth cue compared to binocular depth cues. We also demonstrate that for some observers, the increased contrast of HDR displays indeed increases the realism. However, it is highly observer-dependent whether reduced, physically correct, or exaggerated contrast is perceived as most realistic, even in the presence of the real-world reference scene. Similarly, observers differ in whether reduced, physically correct, or exaggerated stereo 3D is perceived as more realistic. To accommodate the binocular depth perception and realism concept of most observers, display technologies must offer both HDR contrast and stereo personalization. Peter Vangorp, Rafal Mantiuk, Bartosz Bazyluk, Karol Myszkowski, Radoslaw Mantiuk, Simon J. Watt, Hans-Peter Seidel |
SAP | 2 |
| 2014 | Simulating and compensating changes in appearance between day and night visionabstractThe same physical scene seen in bright sunlight and in dusky conditions does not appear identical to the human eye. Similarly, images shown on an 8000 cd/m 2 high-dynamic-range (HDR) display and in a 50 cd/m 2 peak luminance cinema screen also differ significantly in their appearance. We propose a luminance retargeting method that alters the perceived contrast and colors of an image to match the appearance under different luminance levels. The method relies on psychophysical models of matching contrast, models of rod-contribution to vision, and our own measurements. The retargeting involves finding an optimal tone-curve, spatial contrast processing, and modeling of hue and saturation shifts. This lets us reliably simulate night vision in bright conditions, or compensate for a bright image shown on a darker display so that it reveals details and colors that would otherwise be invisible. Robert Wanat, Rafal Mantiuk |
ACM Trans. Graph. | 2 |
| 2013 | Measurements of contrast constancy across a wide range of luminance levelsabstractThe human visual system (HVS) is robust to changes in viewing conditions, such as variations in illumination and viewing distance. Not all elements of our perception stay constant, for example the way we perceive brightness is highly dependent on the background luminance and threshold contrast strongly depends on signal frequency. However, suprathreshold contrast is considered to be constant across viewing conditions and stimuli, the property known as contrast constancy. We intend to verify whether this holds true for a wide range of luminance levels. Results of such measurements could be used to better reproduce images on displays of varying brightness. Robert Wanat, Rafal Mantiuk |
SAP | 2 |
| 2013 | Learning to Predict Localized Distortions in Rendered ImagesabstractAbstract In this work, we present an analysis of feature descriptors for objective image quality assessment. We explore a large space of possible features including components of existing image quality metrics as well as many traditional computer vision and statistical features. Additionally, we propose new features motivated by human perception and we analyze visual saliency maps acquired using an eye tracker in our user experiments. The discriminative power of the features is assessed by means of a machine learning framework revealing the importance of each feature for image quality assessment task. Furthermore, we propose a new data‐driven full‐reference image quality metric which outperforms current state‐of‐theart metrics. The metric was trained on subjective ground truth data combining two publicly available datasets. For the sake of completeness we create a new testing synthetic dataset including experimentally measured subjective distortion maps. Finally, using the same machine‐learning framework we optimize the parameters of popular existing metrics. Martin Cadík, Robert Herzog, Rafal Mantiuk, Radoslaw Mantiuk, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 3 |
| 2013 | Evaluation of Tone Mapping Operators for HDR-VideoabstractAbstract Eleven tone‐mapping operators intended for video processing are analyzed and evaluated with camera‐captured and computer‐generated high‐dynamic‐range content. After optimizing the parameters of the operators in a formal experiment, we inspect and rate the artifacts (flickering, ghosting, temporal color consistency) and color rendition problems (brightness, contrast and color saturation) they produce. This allows us to identify major problems and challenges that video tone‐mapping needs to address. Then, we compare the tone‐mapping results in a pair‐wise comparison experiment to identify the operators that, on average, can be expected to perform better than the others and to assess the magnitude of differences between the best performing operators. Gabriel Eilertsen, Robert Wanat, Rafal Mantiuk, Jonas Unger |
Comput. Graph. Forum | 3 |
| 2013 | Gaze-driven Object Tracking for Real Time RenderingabstractAbstract To efficiently deploy eye‐tracking within 3D graphics applications, we present a new probabilistic method that predicts the patterns of user's eye fixations in animated 3D scenes from noisy eye‐tracker data. The proposed method utilises both the eye‐tracker data and the known information about the 3D scene to improve the accuracy, robustness and stability. Eye‐tracking can thus be used, for example, to induce focal cues via gaze‐contingent depth‐of‐field rendering, add intuitive controls to a video game, and create a highly reliable scene‐aware saliency model. The computed probabilities rely on the consistency of the gaze scan‐paths to the position and velocity of a moving or stationary target. The temporal characteristic of eye fixations is imposed by a Hidden Markov model, which steers the solution towards the most probable fixation patterns. The derivation of the algorithm is driven by the data from two eye‐tracking experiments: the first experiment provides actual eye‐tracker readings and the position of the target to be tracked. The second experiment is used to derive a JND‐scaled (Just Noticeable Difference) quality metric that quantifies the perceived loss of quality due to the errors of the tracking algorithm. Data from both experiments are used to justify design choices, and to calibrate and validate the tracking algorithms. This novel method outperforms commonly used fixation algorithms and is able to track objects smaller then the nominal error of an eye‐tracker. Radoslaw Mantiuk, Bartosz Bazyluk, Rafal Mantiuk |
Comput. Graph. Forum | 3 |
| 2013 | Assessment of video tone-mapping: Are cameras' S-shaped tone-curves good enough?
Josselin Petit, Rafal Mantiuk |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Evaluation of monocular depth cues on a high-dynamic-range display for visualizationabstractThe aim of this work is to identify the depth cues that provide intuitive depth-ordering when used to visualize abstract data. In particular we focus on the depth cues that are effective on a high-dynamic-range (HDR) display: contrast and brightness. In an experiment participants were shown a visualization of the volume layers at different depths with a single isolated monocular cue as the only indication of depth. The observers were asked to identify which slice of the volume appears to be closer. The results show that brightness, contrast and relative size are the most effective monocular depth cues for providing an intuitive depth ordering. Haider Khalil Easa, Rafal Mantiuk, Ik Soo Lim |
ACM Trans. Appl. Percept. | 2 |
| 2012 | Comparison of Four Subjective Methods for Image Quality AssessmentabstractAbstract To provide a convincing proof that a new method is better than the state of the art, computer graphics projects are often accompanied by user studies, in which a group of observers rank or rate results of several algorithms. Such user studies, known as subjective image quality assessment experiments, can be very time‐consuming and do not guarantee to produce conclusive results. This paper is intended to help design efficient and rigorous quality assessment experiments and emphasise the key aspects of the results analysis. To promote good standards of data analysis, we review the major methods for data analysis, such as establishing confidence intervals, statistical testing and retrospective power analysis. Two methods of visualising ranking results together with the meaningful information about the statistical and practical significance are explored. Finally, we compare four most prominent subjective quality assessment methods: single‐stimulus, double‐stimulus, forced‐choice pairwise comparison and similarity judgements. We conclude that the forced‐choice pairwise comparison method results in the smallest measurement variance and thus produces the most accurate results. This method is also the most time‐efficient, assuming a moderate number of compared conditions. Rafal Mantiuk, Anna Lewandowska, Radoslaw Mantiuk |
Comput. Graph. Forum | 1 |
| 2012 | Unsharp Masking, Countershading and Halos: Enhancements or Artifacts?abstractAbstract Countershading is a common technique for local image contrast manipulations, and is widely used both in automatic settings, such as image sharpening and tonemapping, as well as under artistic control, such as in paintings and interactive image processing software. Unfortunately, countershading is a double‐edged sword: while correctly chosen parameters for a given viewing condition can significantly improve the image sharpness or trick the human visual system into perceiving a higher contrast than physically present in an image, wrong parameters, or different viewing conditions can result in objectionable halo artifacts. In this paper we investigate the perception of countershading in the context of a novel mask‐based contrast enhancement algorithm and analyze the circumstances under which the resulting profiles turn from image enhancement to artifact for a range of parameters and viewing conditions. Our experimental results can be modeled as a function of the width of the countershading profile. We employ this empirical function in a range of applications such as image resizing, view dependent tone mapping, and countershading analysis in photographs and works of fine art. Matthew Trentacoste, Rafal Mantiuk, Wolfgang Heidrich, Florian Dufrot |
Comput. Graph. Forum | 2 |
| 2012 | New measurements reveal weaknesses of image quality metrics in evaluating graphics artifactsabstractReliable detection of global illumination and rendering artifacts in the form of localized distortion maps is important for many graphics applications. Although many quality metrics have been developed for this task, they are often tuned for compression/transmission artifacts and have not been evaluated in the context of synthetic CG-images. In this work, we run two experiments where observers use a brush-painting interface to directly mark image regions with noticeable/objectionable distortions in the presence/absence of a high-quality reference image, respectively. The collected data shows a relatively high correlation between the with-reference and no-reference observer markings. Also, our demanding per-pixel image-quality datasets reveal weaknesses of both simple (PSNR, MSE, sCIE-Lab) and advanced (SSIM, MS-SSIM, HDR-VDP-2) quality metrics. The most problematic are excessive sensitivity to brightness and contrast changes, the calibration for near visibility-threshold distortions, lack of discrimination between plausible/implausible illumination, and poor spatial localization of distortions for multi-scale metrics. We believe that our datasets have further potential in improving existing quality metrics, but also in analyzing the saliency of rendering distortions, and investigating visual equivalence given our with- and no-reference data. Martin Cadík, Robert Herzog, Rafal Mantiuk, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 3 |
| 2011 | Glare encoding of high dynamic range imagesabstractWithout specialized sensor technology or custom, multi-chip cameras, high dynamic range imaging typically involves time-sequential capture of multiple photographs. The obvious downside to this approach is that it cannot easily be applied to images with moving objects, especially if the motions are complex. In this paper, we take a novel view of HDR capture, which is based on a computational photography approach. We propose to first optically encode both the low dynamic range portion of the scene and highlight information into a low dynamic range image that can be captured with a conventional image sensor. This step is achieved using a cross-screen, or star filter. Second, we decode, in software, both the low dynamic range image and the highlight information. Lastly, these two portions can be combined to form an image of a higher dynamic range than the regular sensor dynamic range. Mushfiqur Rouf, Rafal Mantiuk, Wolfgang Heidrich, Matthew Trentacoste, Cheryl Lau |
CVPR | 2 |
| 2011 | Cluster-based color space optimizationsabstractTransformations between different color spaces and gamuts are ubiquitous operations performed on images. Often, these transformations involve information loss, for example when mapping from color to grayscale for printing, from multispectral or multiprimary data to tristimulus spaces, or from one color gamut to another. In all these applications, there exists a straightforward “natural” mapping from the source space to the target space, but the mapping is not bijective, resulting in information loss due to metamerism and similar effects. We propose a cluster-based approach for optimizing the transformation for individual images in a way that preserves as much of the information as possible from the source space while staying as faithful as possible to the natural mapping. Our approach can be applied to a host of color transformation problems including color to gray, gamut mapping, conversion of multispectral and multiprimary data to tristimulus colors, and image optimization for color deficient viewers. Cheryl Lau, Wolfgang Heidrich, Rafal Mantiuk |
ICCV | 3 |
| 2011 | Multidimensional image retargetingabstractRetargeting refers to the process by which an image or video is adapted from the display device for which it was meant (target display) to another one (retarget display). The retarget display has different features from the target one such as dynamic range, discretization levels, color gamut, multi-view, and refresh rate spatial resolution. This is a very relevant topic in graphics, given the increasing number of display devices from large, high-contrast screens to small cell phones with limited dynamic range; a lot of techniques are being published in different venues, and it's hard to keep up. For most cases retargeting can be an ill-posed problem, for example in the process of displaying Low Dynamic Range (LDR) or 8-bit content on High Dynamic Range (HDR) displays. Such a problem requires the retargeting algorithm to generate new content which is missing in the input image/frame. In this course, we will present the latest solutions and techniques for retargeting images along various dimensions such as dynamic range, colors, temporal and spatial resolutions, and for the first time offer a much-needed holistic view of the field. Moreover, we are going to show how to measure and analyze the changes applied to an image or video in terms of quality using both psychophysical experiments (subjective) and computational metrics (objective). The course should be of interest to anyone involved in graphics in a broader sense, given the almost unavoidable need to retarget results to different devices -from developers interested in implementing retargeting techniques, to users that just need an overall perspective. For researchers fully engaged in developing multi-dimensional retargeting techniques, this course will serve as a solid background for future algorithms. Francesco Banterle, Alessandro Artusi, Tunç Ozan Aydin, Piotr Didyk, Elmar Eisemann, Diego Gutierrez, Rafal Mantiuk, Karol Myszkowski |
SIGGRAPH Asia Courses | 7 |
| 2011 | Blur-Aware Image DownsamplingabstractAbstract Resizing to a lower resolution can alter the appearance of an image. In particular, downsampling an image causes blurred regions to appear sharper. It is useful at times to create a downsampled version of the image that gives the same impression as the original, such as for digital camera viewfinders. To understand the effect of blur on image appearance at different image sizes, we conduct a perceptual study examining how much blur must be present in a downsampled image to be perceived the same as the original. We find a complex, but mostly image‐independent relationship between matching blur levels in images at different resolutions. The relationship can be explained by a model of the blur magnitude analyzed as a function of spatial frequency. We incorporate this model in a new appearance‐preserving downsampling algorithm, which alters blur magnitude locally to create a smaller image that gives the best reproduction of the original image appearance. Matthew Trentacoste, Rafal Mantiuk, Wolfgang Heidrich |
Comput. Graph. Forum | 2 |
| 2011 | Optimizing a Tone Curve for Backward-Compatible High Dynamic Range Image and Video CompressionabstractFor backward compatible high dynamic range (HDR) video compression, the HDR sequence is reconstructed by inverse tone-mapping a compressed low dynamic range (LDR) version of the original HDR content. In this paper, we show that the appropriate choice of a tone-mapping operator (TMO) can significantly improve the reconstructed HDR quality. We develop a statistical model that approximates the distortion resulting from the combined processes of tone-mapping and compression. Using this model, we formulate a numerical optimization problem to find the tone-curve that minimizes the expected mean square error (MSE) in the reconstructed HDR sequence. We also develop a simplified model that reduces the computational complexity of the optimization problem to a closed-form solution. Performance evaluations show that the proposed methods provide superior performance in terms of HDR MSE and SSIM compared to existing tone-mapping schemes. It is also shown that the LDR image quality resulting from the proposed methods matches that produced by perceptually-based TMOs. Zicong Mai, Hassan Mansour, Rafal Mantiuk, Panos Nasiopoulos, Rabab K. Ward, Wolfgang Heidrich |
IEEE Trans. Image Process. | 3 |
| 2011 | HDR-VDP-2: a calibrated visual metric for visibility and quality predictions in all luminance conditionsabstractVisual metrics can play an important role in the evaluation of novel lighting, rendering, and imaging algorithms. Unfortunately, current metrics only work well for narrow intensity ranges, and do not correlate well with experimental data outside these ranges. To address these issues, we propose a visual metric for predicting visibility (discrimination) and quality (mean-opinion-score). The metric is based on a new visual model for all luminance conditions, which has been derived from new contrast sensitivity measurements. The model is calibrated and validated against several contrast discrimination data sets, and image quality databases (LIVE and TID2008). The visibility metric is shown to provide much improved predictions as compared to the original HDR-VDP and VDP metrics, especially for low luminance conditions. The image quality predictions are comparable to or better than for the MS-SSIM, which is considered one of the most successful quality metrics. The code of the proposed metric is available on-line. Rafal Mantiuk, Kil Joong Kim, Allan G. Rempel, Wolfgang Heidrich |
ACM Trans. Graph. | 1 |
| 2010 | On-the-fly tone mapping for backward-compatible high dynamic range image/video compressionabstractIn this paper, we propose a real-time tone-mapping scheme for backward compatible high dynamic range (HDR) video compression. The appropriate choice of a tone-mapping operator (TMO) can significantly improve the HDR quality reconstructed from a low dynamic range (LDR) version. We develop a statistical model that approximates the mean square error (MSE) distortion resulting from the combined processes of tone-mapping and compression. Using this model, we formulate a numerical optimization problem to find the tone-curve that minimizes the expected MSE in the reconstructed HDR sequence. We then simplify the developed model in order to reduce the computational complexity of the optimization problem to a closed-form solution. Performance evaluations show that the proposed methods provide superior performance in terms of HDR MSE and SSIM compared to existing tone-mapping schemes. It is also shown that the LDR image quality resulting from the proposed methods matches that produced by perceptually-based TMOs. Zicong Mai, Hassan Mansour, Rafal Mantiuk, Panos Nasiopoulos, Rabab K. Ward, Wolfgang Heidrich |
ISCAS | 3 |
| 2010 | A Comparison of Three Image Fidelity Metrics of Different Computational Principles for JPEG2000 Compressed Abdomen CT ImagesabstractThis study aimed to evaluate three image fidelity metrics of different computational principles--peak signal-to-noise ratio (PSNR), high-dynamic range visual difference predictor (HDR-VDP), and multiscale structural similarity (MS-SSIM)--in measuring the fidelity of JPEG2000 compressed abdomen computed tomography images from a viewpoint of visually lossless compression. Three hundred images with 0.67- or 5-mm section thickness were compressed to one of five compression ratios ranging from reversible compression to 15:1. The fidelity of each compressed image was measured by five radiologists' visual analyses (distinguishable or indistinguishable from the original) and the three metrics. The Spearman rank correlation coefficients of the PSNR, HDR-VDP, and MS-SSIM values with the number of readers responding as indistinguishable were 0.86, 0.94, and 0.86, respectively. Using the pooled readers' responses as the reference standard, the area under the receiver-operating-characteristic curve for the HDR-VDP (0.99) was significantly greater than that for the PSNR (0.95) (p < 0.001) and for the MS-SSIM (0.96) (p = 0.003), and there was no significant difference between the PSNR and MS-SSIM (p = 0.70). In measuring the image fidelity, the HDR-VDP outperforms the PSNR and MS-SSIM, and the MS-SSIM and PSNR are comparable. Kil Joong Kim, Bo Hyoung Kim, Rafal Mantiuk, Thomas Richter 0005, Hyunna Lee, Heung Sik Kang, Jinwook Seo, Kyoung Ho Lee |
IEEE Trans. Medical Imaging | 3 |
| 2009 | Color correction for tone mappingabstractAbstract Tone mapping algorithms offer sophisticated methods for mapping a real‐world luminance range to the luminance range of the output medium but they often cause changes in color appearance. In this work we conduct a series of subjective appearance matching experiments to measure the change in image colorfulness after contrast compression and enhancement. The results indicate that the relation between contrast compression and the color saturation correction that matches color appearance is non‐linear and smaller color correction is required for small change of contrast. We demonstrate that the relation cannot be fully explained by color appearance models. We propose color correction formulas that can be used with existing tone mapping algorithms. We extend existing global and local tone mapping operators and show that the proposed color correction formulas can preserve original image colors after tone scale manipulation. Radoslaw Mantiuk, Rafal Mantiuk, Anna Lewandowska, Wolfgang Heidrich |
Comput. Graph. Forum | 2 |
| 2008 | Enhancement of Bright Video Features for HDR DisplaysabstractAbstract To utilize the full potential of new high dynamic range (HDR) displays, a system for the enhancement of bright luminous objects in video sequences is proposed. The system classifies clipped (saturated) regions as lights, reflections or diffuse surfaces using a semi‐automatic classifier and then enhances each class of objects with respect to its relative brightness. The enhancement algorithm can significantly stretch the contrast of clipped regions while avoiding amplification of noise and contouring. We demonstrate that the enhanced video is strongly preferred to non‐enhanced video, and it compares favorably to other methods. Piotr Didyk, Rafal Mantiuk, Matthias Hein 0001, Hans-Peter Seidel |
Comput. Graph. Forum | 2 |
| 2008 | Modeling a Generic Tone-mapping OperatorabstractAbstract Although several new tone‐mapping operators are proposed each year, there is no reliable method to validate their performance or to tell how different they are from one another. In order to analyze and understand the behavior of tone‐mapping operators, we model their mechanisms by fitting a generic operator to an HDR image and its tone‐mapped LDR rendering. We demonstrate that the majority of both global and local tone‐mapping operators can be well approximated by computationally inexpensive image processing operations, such as a per‐pixel tone curve, a modulation transfer function and color saturation adjustment. The results produced by such a generic tone‐mapping algorithm are often visually indistinguishable from much more expensive algorithms, such as the bilateral filter. We show the usefulness of our generic tone‐mapper in backward‐compatible HDR image compression, the black‐box analysis of existing tone mapping algorithms and the synthesis of new algorithms that are combination of existing operators. Rafal Mantiuk, Hans-Peter Seidel |
Comput. Graph. Forum | 1 |
| 2008 | Dynamic range independent image quality assessmentabstractThe diversity of display technologies and introduction of high dynamic range imagery introduces the necessity of comparing images of radically different dynamic ranges. Current quality assessment metrics are not suitable for this task, as they assume that both reference and test images have the same dynamic range. Image fidelity measures employed by a majority of current metrics, based on the difference of pixel intensity or contrast values between test and reference images, result in meaningless predictions if this assumption does not hold. We present a novel image quality metric capable of operating on an image pair where both images have arbitrary dynamic ranges. Our metric utilizes a model of the human visual system, and its central idea is a new definition of visible distortion based on the detection and classification of visible changes in the image structure. Our metric is carefully calibrated and its performance is validated through perceptual experiments. We demonstrate possible applications of our metric to the evaluation of direct and inverse tone mapping operators as well as the analysis of the image appearance on displays with various characteristics. Tunç Ozan Aydin, Rafal Mantiuk, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 2 |
| 2008 | Display adaptive tone mappingabstractWe propose a tone mapping operator that can minimize visible contrast distortions for a range of output devices, ranging from e-paper to HDR displays. The operator weights contrast distortions according to their visibility predicted by the model of the human visual system. The distortions are minimized given a display model that enforces constraints on the solution. We show that the problem can be solved very efficiently by employing higher order image statistics and quadratic programming. Our tone mapping technique can adjust image or video content for optimum contrast visibility taking into account ambient illumination and display characteristics. We discuss the differences between our method and previous approaches to the tone mapping problem. Rafal Mantiuk, Scott Daly, Louis Kerofsky |
ACM Trans. Graph. | 1 |
| 2007 | High Dynamic Range Image and Video Compression - Fidelity Matching Human Visual PerformanceabstractVast majority of digital images and video material stored today can capture only a fraction of visual information visible to the human eye and does not offer sufficient quality to fully exploit capabilities of new display devices. High dynamic range (HDR) image and video formats encode the full visible range of luminance and color gamut, thus offering ultimate fidelity, limited only by the capabilities of the human eye and not by any existing technology. In this paper we demonstrate how existing image and video compression standards can be extended to encode HDR content efficiently. This is achieved by a custom color space for encoding HDR pixel values that is derived from the visual performance data. We also demonstrate how HDR image and video compression can be designed so that it is backward compatible with existing formats. Rafal Mantiuk, Grzegorz Krawczyk, Karol Myszkowski, Hans-Peter Seidel |
ICIP (1) | 1 |
| 2007 | Brightness Adjustment for HDR and Tone Mapped ImagesabstractBoth High Dynamic Range images and their tone mapped correspondents contain relative luminance values which have to be mapped on a scale of available gray levels of a display. Such mapping includes brightness adjustment, which has a direct impact on the final image appearance and the observers' assessment of image quality. We conduct a psychophysical experiment in which subjects adjust image brightness to match their preference. We observe that the brightness choice is consistent across subjects and is primarily affected by image content. We investigate popular methods for automatic brightness adjustment and show a significant inaccuracy for a group of images. The incorrect brightness adjustment degrades in these cases perceived image quality. We identify characteristics of images that are highly correlated with the subjects' choice of brightness and develop an improved model for the brightness adjustment. Grzegorz Krawczyk, Rafal Mantiuk, Dorota Zdrojewska, Hans-Peter Seidel |
PG | 2 |
| 2006 | Analysis of Reproducing Real-World Appearance on Displays of Varying Dynamic RangeabstractAbstract We conduct a series of experiments to investigate the desired properties of a tone mapping operator (TMO) and to design such an operator based on subjective data. We propose a novel approach to the tone mapping problem, in which the tone mapping parameters are determined based on the data from subjective experiments, rather than an image processing algorithm or a visual model. To collect this data, a series of experiments are conducted in which the subjects adjust three generic TMO parameters: brightness, contrast and color saturation. In two experiments, the subjects are to find a) the most preferred image without a reference image (preference task) and b) the closest image to the real‐world scene which the subjects are confronted with (fidelity task). We analyze subjects’ choice of parameters to provide more intuitive control over the parameters of a tone mapping operator. Unlike most of the researched TMOs that focus on rendering for standard low dynamic range monitors, we consider a broad range of potential displays, each offering different dynamic range and brightness. We simulate capabilities of such displays on a high dynamic range (HDR) display. This allows us to address the question of how tone mapping needs to be adjusted to accommodate displays with drastically different dynamic ranges. Categories and Subject Descriptors (according to ACM CCS): I.3.8 [Computer Graphics]: High dynamic range images, Visual perception, Tone mapping Akiko Yoshida, Rafal Mantiuk, Karol Myszkowski, Hans-Peter Seidel |
Comput. Graph. Forum | 2 |
| 2006 | A perceptual framework for contrast processing of high dynamic range images
Rafal Mantiuk, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Appl. Percept. | 1 |
| 2006 | Backward compatible high dynamic range MPEG video compressionabstractTo embrace the imminent transition from traditional low-contrast video (LDR) content to superior high dynamic range (HDR) content, we propose a novel backward compatible HDR video compression (HDR MPEG) method. We introduce a compact reconstruction function that is used to decompose an HDR video stream into a residual stream and a standard LDR stream, which can be played on existing MPEG decoders, such as DVD players. The reconstruction function is finely tuned to the content of each HDR frame to achieve strong decorrelation between the LDR and residual streams, which minimizes the amount of redundant information. The size of the residual stream is further reduced by removing invisible details prior to compression using our HDR-enabled filter, which models luminance adaptation, contrast sensitivity, and visual masking based on the HDR content. Designed especially for DVD movie distribution, our HDR MPEG compression method features low storage requirements for HDR content resulting in a 30% size increase to an LDR video sequence. The proposed compression method does not impose restrictions or modify the appearance of the LDR or HDR video. This is important for backward compatibility of the LDR stream with current DVD appearance, and also enables independent fine tuning, tone mapping, and color grading of both streams. Rafal Mantiuk, Alexander Efremov, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 1 |
| 2004 | Perception-motivated high dynamic range video encodingabstractDue to rapid technological progress in high dynamic range (HDR) video capture and display, the efficient storage and transmission of such data is crucial for the completeness of any HDR imaging pipeline. We propose a new approach for inter-frame encoding of HDR video, which is embedded in the well-established MPEG-4 video compression standard. The key component of our technique is luminance quantization that is optimized for the contrast threshold perception in the human visual system. The quantization scheme requires only 10--11 bits to encode 12 orders of magnitude of visible luminance range and does not lead to perceivable contouring artifacts. Besides video encoding, the proposed quantization provides perceptually-optimized luminance sampling for fast implementation of any global tone mapping operator using a lookup table. To improve the quality of synthetic video sequences, we introduce a coding scheme for discrete cosine transform (DCT) blocks with high contrast. We demonstrate the capabilities of HDR video in a player, which enables decoding, tone mapping, and applying post-processing effects in real-time. The tone mapping algorithm as well as its parameters can be changed interactively while the video is playing. We can simulate post-processing effects such as glare, night vision, and motion blur, which appear very realistic due to the usage of HDR data. Rafal Mantiuk, Grzegorz Krawczyk, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 1 |