EDBT 2026 Demo / reviewers in the wild / expert
Alexandre Chapiro
dblp:58/3419
· DBLP profile ↗
22ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-7367-0131ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ML-PEA: Machine Learning-Based Perceptual Algorithms for Display Power OptimizationabstractAbstract Image processing techniques can be used to modulate the pixel intensities of an image to reduce the power consumption of the display device. A simple example of this consists of uniformly dimming the entire image. Such algorithms should strive to minimize the impact on image quality while maximizing power savings. Techniques based on heuristics or human perception have been proposed, both for traditional flat panel displays and modern display modalities such as virtual and augmented reality (VR/AR). In this paper, we focus on developing and evaluating display power‐saving techniques that use machine learning (ML) in VR displays. We developed a U‐Net‐based technique paired with perceptual and power optimization loss functions that generates spatially varying dimming maps. These dimming maps are used to modulate input images, per‐pixel, to generate a power‐efficient image. Our pipeline was validated via quantitative analysis using image quality metrics and through a subjective study. Our subjective validation provides results scaled in perceptual just‐objectionable‐difference (JOD) units. This data, when rescaled, allows for comparisons of our technique with recent studies on VR display power optimization. Our results show that participants prefer our technique over a uniform dimming baseline for high target power saving conditions. This model and study serve as a template and baseline for future applications of deep learning to display power optimization. Model training code and data can be found at kenchen10.github.io/projects/mlpea/index.html . Kenneth Chen, Nathan Matsuda, Thomas Wan, Ajit Ninan, Alexandre Chapiro, Qi Sun 0003 |
Comput. Graph. Forum | 5 |
| 2026 | HoloQA: Full Reference Video Quality Assessor of Rendered Human Avatars in Virtual RealityabstractWe present HoloQA, a new state-of-the-art Full Reference Video Quality Assessment (VQA) model that was designed using principles of visual neuroscience, information theory, and self-supervised deep learning to accurately predict the quality of rendered digital human avatars in Virtual Reality (VR) and Augmented Reality (AR) systems. The growing adoption of VR/AR applications that aim to transmit digital human avatars over bandwidth-limited video networks has driven the need for VQA algorithms that better account for the kinds of distortions that reduce the quality of rendered and viewed avatars. As we will show, standard VQA models often fail to capture distortions unique to the rendering, transmission, and compression of videos containing human avatars. Towards solving this difficult problem, we adopt a multi-level Mixture-of-Experts approach. This involves computing distortion-aware perceptual features and high-level content-aware deep features that capture semantic attributes of human body avatars. The high-level features are computed using a self-supervised, pre-trained deep learning network. We show that HoloQA is able to achieve state-of-the-art performance on the recently introduced LIVE-Meta Rendered Human Avatar VQA database, demonstrating its efficacy in predicting the quality of rendered human avatars in VR. Furthermore, we demonstrate the competitive performance of HoloQA on other digital human avatar databases and on another synthetically generated video quality use case: cloud gaming. The code associated with this work will be made available on https://github.com/avinabsaha/HologramQAGitHub. Avinab Saha, Yu-Chih Chen, Christian Häne, Jean-Charles Bazin, Ioannis Katsavounidis, Alexandre Chapiro, Alan C. Bovik |
IEEE Trans. Image Process. | 6 |
| 2025 | Supra-threshold Contrast Perception in Augmented RealityabstractWhen an image is seen on an optical see-through augmented reality (AR) display, the light from the display is mixed with the background light from the environment. This can severely limit the available contrast in AR, which is often orders of magnitude below that of traditional displays. Yet, the presented images appear sharper and show more details than the reduction in physical contrast would indicate. In this work, we hypothesize two effects that are likely responsible for the enhanced perceived contrast in AR: background discounting, which allows observers focused on the display plane to partially discount the light from the environment; and supra-threshold contrast perception, which explains the differences in contrast perception across luminance levels. In a series of controlled experiments on an AR high-dynamic-range multi-focal haploscope testbed, we found no statistical evidence supporting the effect of background discounting on contrast perception. Instead, the increase of visibility in AR is better explained with models of supra-threshold contrast perception. Our findings can be generalized to incorporate an image input, and this model serves to design better algorithms and hardware for display systems affected by additive light, such as AR. Dongyeon Kim, Maliha Ashraf, Alexandre Chapiro, Rafal Mantiuk |
SIGGRAPH Asia | 3 |
| 2024 | FaceMap: Distortion-Driven Perceptual Facial Saliency Maps
Zhongshi Jiang, Kishore Venkateshan, Giljoo Nam, Meixu Chen, Romain Bachy, Jean-Charles Bazin, Alexandre Chapiro |
SIGGRAPH Asia | 7 |
| 2024 | Subjective and Objective Quality Assessment of Rendered Human Avatar Videos in Virtual RealityabstractWe study the visual quality judgments of human subjects on digital human avatars (sometimes referred to as "holograms" in the parlance of virtual reality [VR] and augmented reality [AR] systems) that have been subjected to distortions. We also study the ability of video quality models to predict human judgments. As streaming human avatar videos in VR or AR become increasingly common, the need for more advanced human avatar video compression protocols will be required to address the tradeoffs between faithfully transmitting high-quality visual representations while adjusting to changeable bandwidth scenarios. During transmission over the internet, the perceived quality of compressed human avatar videos can be severely impaired by visual artifacts. To optimize trade-offs between perceptual quality and data volume in practical workflows, video quality assessment (VQA) models are essential tools. However, there are very few VQA algorithms developed specifically to analyze human body avatar videos, due, at least in part, to the dearth of appropriate and comprehensive datasets of adequate size. Towards filling this gap, we introduce the LIVE-Meta Rendered Human Avatar VQA Database, which contains 720 human avatar videos processed using 20 different combinations of encoding parameters, labeled by corresponding human perceptual quality judgments that were collected in six degrees of freedom VR headsets. To demonstrate the usefulness of this new and unique video resource, we use it to study and compare the performances of a variety of state-of-the-art Full Reference and No Reference video quality prediction models, including a new model called HoloQA. As a service to the research community, we publicly releases the metadata of the new database at https://live.ece.utexas.edu/research/LIVE-Meta-rendered-human-avatar/index.html. Yu-Chih Chen, Avinab Saha, Alexandre Chapiro, Christian Häne, Jean-Charles Bazin, Stefano Zanetti, Ioannis Katsavounidis, Alan C. Bovik |
IEEE Trans. Image Process. | 3 |
| 2024 | AR-DAVID: Augmented Reality Display Artifact Video DatasetabstractThe perception of visual content in optical-see-through augmented reality (AR) devices is affected by the light coming from the environment. This additional light interacts with the content in a non-trivial manner because of the illusion of transparency, different focal depths, and motion parallax. To investigate the impact of environment light on display artifact visibility (such as blur or color fringes), we created the first subjective quality dataset targeted toward augmented reality displays. Our study consisted of 6 scenes, each affected by one of 6 distortions at two strength levels, seen against one of 3 background patterns shown at 2 luminance levels: 432 conditions in total. Our dataset shows that environment light has a much smaller masking effect than expected. Further, we show that this effect cannot be explained by compositing of the AR-content with the background using optical blending models. As a consequence, we demonstrate that existing video quality metrics perform worse than expected when predicting the perceived magnitude of degradation in AR displays, motivating further research. Alexandre Chapiro, Dongyeon Kim, Yuta Asano, Rafal Mantiuk |
ACM Trans. Graph. | 1 |
| 2024 | PEA-PODs: Perceptual Evaluation of Algorithms for Power Optimization in XR DisplaysabstractDisplay power consumption is an emerging concern for untethered devices. This goes double for augmented and virtual extended reality (XR) displays, which target high refresh rates and high resolutions while conforming to an ergonomically light form factor. A number of image mapping techniques have been proposed to extend battery usage. However, there is currently no comprehensive quantitative understanding of how the power savings provided by these methods compare to their impact on visual quality. We set out to answer this question. To this end, we present a perceptual evaluation of algorithms (PEA) for power optimization in XR displays (PODs). Consolidating a portfolio of six power-saving display mapping approaches, we begin by performing a large-scale perceptual study to understand the impact of each method on perceived quality in the wild. This results in a unified quality score for each technique, scaled in just-objectionable-difference (JOD) units. In parallel, each technique is analyzed using hardware-accurate power models. The resulting JOD-to-Milliwatt transfer function provides a first-of-its-kind look into tradeoffs offered by display mapping techniques, and can be directly employed to make architectural decisions for power budgets on XR displays. Finally, we leverage our study data and power models to address important display power applications like the choice of display primary, power implications of eye tracking, and more 1 . Kenneth Chen, Thomas Wan, Nathan Matsuda, Ajit Ninan, Alexandre Chapiro, Qi Sun 0003 |
ACM Trans. Graph. | 5 |
| 2024 | ColorVideoVDP: A visual difference predictor for image, video and display distortionsabstractColorVideoVDP is a video and image quality metric that models spatial and temporal aspects of vision for both luminance and color. The metric is built on novel psychophysical models of chromatic spatiotemporal contrast sensitivity and cross-channel contrast masking. It accounts for the viewing conditions, geometric, and photometric characteristics of the display. It was trained to predict common video-streaming distortions (e.g., video compression, rescaling, and transmission errors) and also 8 new distortion types related to AR/VR displays (e.g., light source and waveguide non-uniformities). To address the latter application, we collected our novel XR-Display-Artifact-Video quality dataset (XR-DAVID), comprised of 336 distorted videos. Extensive testing on XR-DAVID, as well as several datasets from the literature, indicate a significant gain in prediction performance compared to existing metrics. ColorVideoVDP opens the doors to many novel applications that require the joint automated spatiotemporal assessment of luminance and color distortions, including video streaming, display specification, and design, visual comparison of results, and perceptually-guided quality optimization. The code for the metric can be found at https://github.com/gfxdisp/ColorVideoVDP. Rafal Mantiuk, Param Hanji, Maliha Ashraf, Yuta Asano, Alexandre Chapiro |
ACM Trans. Graph. | 5 |
| 2023 | Perceptually Adaptive Real-Time Tone MappingabstractTone mapping operators aim to remap content to a display’s dynamic range. Virtual reality is a popular new display modality that has significant differences from other media, making the use of traditional tone mapping techniques difficult. Moreover, real-time adaptive estimation of tone curves that faithfully maintain appearance remains a significant challenge. In this work, we propose a real-time perceptual contrast-matching framework, that allows us to optimally remap scenes for target displays. Our framework is optimized for efficiency and runs on a mobile Quest 2 headset in under 1ms per frame. A subjective study on an HDR-VR prototype demonstrates our method’s effectiveness across a wide range of display luminances, producing imagery that is preferred to alternatives tone mapped at peak luminances an order of magnitude higher. This result highlights the importance of good tone mapping for visual quality in VR. Taimoor Tariq, Nathan Matsuda, Eric Penner, Jerry Jia, Douglas Lanman, Ajit Ninan, Alexandre Chapiro |
SIGGRAPH Asia | 7 |
| 2023 | Introduction to the SAP 2023 Special IssueabstractNo abstract available. Alexandre Chapiro, Andrew C. Robb |
ACM Trans. Appl. Percept. | 1 |
| 2023 | Skin-Screen: A Computational Fabrication Framework for Color TattoosabstractTattoos are a highly popular medium, with both artistic and medical applications. Although the mechanical process of tattoo application has evolved historically, the results are reliant on the artisanal skill of the artist. This can be especially challenging for some skin tones, or in cases where artists lack experience. We provide the first systematic overview of tattooing as a computational fabrication technique. We built an automated tattooing rig and a recipe for the creation of silicone sheets mimicking realistic skin tones, which allowed us to create an accurate model predicting tattoo appearance. This enables several exciting applications including tattoo previewing, color retargeting, novel ink spectra optimization, color-accurate prosthetics, and more. Michal Piovarci, Alexandre Chapiro, Bernd Bickel |
ACM Trans. Graph. | 2 |
| 2022 | Realistic Luminance in VRabstractAs virtual reality (VR) headsets continue to achieve ever more immersive visuals along the axes of resolution, field of view, focal cues, distortion mitigation, and so on, the luminance and dynamic range of these devices falls far short of widely available consumer televisions. While work remains to be done on the display architecture side, power and weight limitations in head-mounted displays pose a challenge for designs aiming for high luminance. In this paper, we seek to gain a basic understanding of VR user preferences for display luminance values in relation to known, real-world luminances for immersive, natural scenes. To do so, we analyze the luminance characteristics of an existing high-dynamic-range (HDR) panoramic image dataset, build an HDR VR headset capable of reproducing over 20,000 nits peak luminance, and conduct a first-of-its-kind study on user brightness preferences in VR. We conclude that current commercial VR headsets do not meet user preferences for display luminance, even for indoor scenes. Nathan Matsuda, Alexandre Chapiro, Yang Zhao 0030, Clinton Smith, Romain Bachy, Douglas Lanman |
SIGGRAPH Asia | 2 |
| 2022 | stelaCSF: a unified model of contrast sensitivity as the function of spatio-temporal frequency, eccentricity, luminance and areaabstractA contrast sensitivity function, or CSF, is a cornerstone of many visual models. It explains whether a contrast pattern is visible to the human eye. The existing CSFs typically account for a subset of relevant dimensions describing a stimulus, limiting the use of such functions to either static or foveal content but not both. In this paper, we propose a unified CSF, stelaCSF, which accounts for all major dimensions of the stimulus: spatial and temporal frequency, eccentricity, luminance, and area. To model the 5-dimensional space of contrast sensitivity, we combined data from 11 papers, each of which studied a subset of this space. While previously proposed CSFs were fitted to a single dataset, stelaCSF can predict the data from all these studies using the same set of parameters. The predictions are accurate in the entire domain, including low frequencies. In addition, stelaCSF relies on psychophysical models and experimental evidence to explain the major interactions between the 5 dimensions of the CSF. We demonstrate the utility of our new CSF in a flicker detection metric and in foveated rendering. Rafal Mantiuk, Maliha Ashraf, Alexandre Chapiro |
ACM Trans. Graph. | 3 |
| 2022 | Geo-Metric: A Perceptual Dataset of Distortions on FacesabstractIn this work we take a novel perception-centered approach to quantify distortions on 3D geometry of faces, to which humans are particularly sensitive. We generated a dataset, composed of 100 high-quality and demographically-balanced face scans. We then subjected these meshes to distortions that cover relevant use cases in computer graphics, and conducted a large-scale perceptual study to subjectively evaluate them. Our dataset consists of over 84,000 quality comparisons, making it the largest ever psychophysical dataset for geometric distortions. Finally, we demonstrated how our data can be used for applications like metrics, compression, and level-of-detail rendering. Krzysztof Wolski, Laura C. Trutoiu, Zhao Dong 0001, Zhengyang Shen, Kevin Mackenzie, Alexandre Chapiro |
ACM Trans. Graph. | 6 |
| 2021 | FovVideoVDP: a visible difference predictor for wide field-of-view videoabstractFovVideoVDP is a video difference metric that models the spatial, temporal, and peripheral aspects of perception. While many other metrics are available, our work provides the first practical treatment of these three central aspects of vision simultaneously. The complex interplay between spatial and temporal sensitivity across retinal locations is especially important for displays that cover a large field-of-view, such as Virtual and Augmented Reality displays, and associated methods, such as foveated rendering. Our metric is derived from psychophysical studies of the early visual system, which model spatio-temporal contrast sensitivity, cortical magnification and contrast masking. It accounts for physical specification of the display (luminance, size, resolution) and viewing distance. To validate the metric, we collected a novel foveated rendering dataset which captures quality degradation due to sampling and reconstruction. To demonstrate our algorithm's generality, we test it on 3 independent foveated video datasets, and on a large image quality dataset, achieving the best performance across all datasets when compared to the state-of-the-art. Rafal Mantiuk, Gyorgy Denes, Alexandre Chapiro, Anton Kaplanyan, Gizem Rufo, Romain Bachy, Trisha Lian, Anjul Patney |
ACM Trans. Graph. | 3 |
| 2019 | A Luminance-aware Model of Judder PerceptionabstractThe perceived discrepancy between continuous motion as seen in nature and frame-by-frame exhibition on a display, sometimes termed judder, is an integral part of video presentation. Over time, content creators have developed a set of rules and guidelines for maintaining a desirable cinematic look under the restrictions placed by display technology without incurring prohibitive judder. With the advent of novel displays capable of high brightness, contrast, and frame rates, these guidelines are no longer sufficient to present audiences with a uniform viewing experience. In this work, we analyze the main factors for perceptual motion artifacts in digital presentation and gather psychophysical data to generate a model of judder perception. Our model enables applications like matching perceived motion artifacts to a traditionally desirable level and maintain a cinematic motion look. Alexandre Chapiro, Robin Atkins, Scott Daly |
ACM Trans. Graph. | 1 |
| 2018 | Influence of Screen Size and Field of View on Perceived BrightnessabstractWe present a study into the perception of display brightness as related to the physical size and distance of the screen from the observer. Brightness perception is a complex topic, which is influenced by a number of lower- and higher-order factors—with empirical evidence from the cinema industry suggesting that display size may play a significant role. To test this hypothesis, we conducted a series of user studies exploring brightness perception for a range of displays and distances from the observer that span representative use scenarios. Our results suggest that retinal size is not sufficient to explain the range of discovered brightness variations, but is sufficient in combination with physical distance from the observer. The resulting model can be used as a step toward perceptually correcting image brightness perception based on target display parameters. This can be leveraged for energy management and the preservation of artistic intent. A pilot study suggests that adaptation luminance is an additional factor for the magnitude of the effect. Alexandre Chapiro, Timo Kunkel, Robin Atkins, Scott Daly |
ACM Trans. Appl. Percept. | 1 |
| 2015 | Video content and structure description based on keyframes, clusters and storyboardsabstractIn this paper we present a novel system to extract keyframes, shot clusters and structural storyboards for video content description, which can be used for a variety of summarization, visualization, classification, indexing and retrieval applications. The system automatically selects an appealing set of keyframes and creates meaningful clusters of shots. It further identifies sections that appear recurrently, which are called anchors, and typically divide television shows into different parts. This information about anchors can then be used to browse video content in a new fashion. Finally, our system creates a new type of interactive storyboard suitable to visualize and analyze the structure of the video in a novel way. Marc Junyent, Pablo Beltrán, Miquel A. Farre, Jordi Pont-Tuset, Alexandre Chapiro, Aljoscha Smolic |
MMSP | 5 |
| 2015 | Art-directable Continuous Dynamic Range video
Alexandre Chapiro, Tunç Ozan Aydin, Nikolce Stefanoski, Simone Croci, Aljoscha Smolic, Markus Gross 0001 |
Comput. Graph. | 1 |
| 2014 | Perceptual evaluation of cardboarding in 3D content visualizationabstractA pervasive artifact that occurs when visualizing 3D content is the so-called "cardboarding" effect, where objects appear flat due to depth compression, with relatively little research conducted to perceptually quantify its effects. Our aim is to shed light on the subjective preferences and practical perceptual limits of stereo vision with respect to cardboarding. We present three experiments that explore the consequences of displaying simple scenes with reduced depths using both subjective ratings and adjustments and objective sensitivity metrics. Our results suggest that compressing depth to 80% or above is likely to be acceptable, whereas sensitivity to the cardboarding artifact below 30% is very high. These values could be used in practice as guidelines for commonplace depth mapping operations in 3D production pipelines. Alexandre Chapiro, Olga Diamanti, Steven Poulakos, Carol O'Sullivan, Aljoscha Smolic, Markus Gross 0001 |
SAP | 1 |
| 2014 | Optimizing stereo-to-multiview conversion for autostereoscopic displaysabstractAbstract We present a novel stereo‐to‐multiview video conversion method for glasses‐free multiview displays. Different from previous stereo‐to‐multiview approaches, our mapping algorithm utilizes the limited depth range of autostereoscopic displays optimally and strives to preserve the scene's artistic composition and perceived depth even under strong depth compression. We first present an investigation of how perceived image quality relates to spatial frequency and disparity. The outcome of this study is utilized in a two‐step mapping algorithm, where we (i) compress the scene depth using a non‐linear global function to the depth range of an autostereoscopic display and (ii) enhance the depth gradients of salient objects to restore the perceived depth and salient scene structure. Finally, an adapted image domain warping algorithm is proposed to generate the multiview output, which enables overall disparity range extension. Alexandre Chapiro, Simon Heinzle, Tunç Ozan Aydin, Steven Poulakos, Matthias Zwicker, Aljoscha Smolic, Markus Gross 0001 |
Comput. Graph. Forum | 1 |
| 2009 | Detection of high frequency regions in multiresolutionabstractWe propose a method for the detection of high frequency regions using multiresolution analysis and orientation tensors. A scalar field representing multiresolution edges is obtained. Local maxima of this scalar space indicate regions having coincident detail vectors in multiple scales of a wavelet decomposition. This is useful for finding edges, textures, collinear structures and salient regions for computer vision methods. The image is decomposed into several scales using the discrete wavelet transform (DWT). The resulting detail spaces form vectors indicating intensity variations which are adequately combined using orientation tensors. The multivariate data of the resulting tensor field provides fair estimations of high frequency regions. Using these tensors, a positive scalar is computed for each original image pixel. Our results show that this descriptor indicates areas having relevant intensity variation in multiple scales. Virgínia Fernandes Mota, Eder de Almeida Perez, Tássio Knop de Castro, Alexandre Chapiro, Marcelo Bernardes Vieira |
ICIP | 4 |