Ana Serrano

dblp:172/3357 · DBLP profile ↗
← Back
45ranked-venue papers
13as first author
33since 2021 · last 2026
0000-0002-7796-3177ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 12 first-author · 32 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 An analysis of gaze behavior under multisensory cognitive load in immersive environments
abstract
Virtual Reality (VR) enables immersive and realistic experiences in which complex multisensory interactions often impose elevated cognitive demands on users. Excessive levels of cognitive load (CL) have been shown to degrade performance and user experience, motivating the search for robust and non-intrusive indicators for CL monitoring in immersive environments. While recent approaches have leveraged physiological signals for this purpose, many of these signals are sensitive to motion and require complex sensor setups. Given the widespread integration of eye tracking in modern head-mounted displays, gaze behavior emerges as a promising alternative for objective CL assessment. In this work, we present a comprehensive analysis of gaze behavior and its relationship to cognitive load in immersive, task-oriented VR. Using a publicly available dataset collected in a multisensory visual search experience, we first examine how a range of gaze-derived markers, including fixations, saccades, eye eccentricity, and pupil dilation, vary across cognitive load conditions. We then investigate their relationship with biomarkers derived from complementary physiological signals, including electrocardiogram (ECG), electrodermal activity (EDA), and respiration, to better understand how gaze-based markers relate to broader physiological responses associated with cognitive load. Our results reveal consistent changes in oculomotor behavior under high cognitive load, characterized by reduced saccade amplitudes and eye eccentricity, together with increased pupil dilation. These patterns indicate a more centered and less exploratory visual behavior, consistent with attention tunneling under high task demands. In addition, we identify strong correlations between several gaze markers and the phasic component of electrodermal activity, a well-known indicator of mental effort. Together, these findings highlight the potential of gaze-based measures as lightweight and non-intrusive indicators of cognitive load, supporting the development of adaptive, user-aware VR applications.
Jaime Bielsa, Jorge Pina, Ana Serrano
Comput. Graph.3
2026 Foreword to the special section on Spanish Computer Graphics Conference 2025
Ana Serrano, Oscar Argudo, Olatz Iparraguirre
Comput. Graph.1
2026 Overdriving Visual Depth Perception via Sound Modulation in VR
abstract
Our ability to perceive and navigate the spatial world is a cornerstone of human experience, relying on the integration of visual and auditory cues to form a coherent sense of depth and distance. In stereoscopic 3D vision, depth perception requires fixation of both eyes on a target object, which is achieved through vergence movements, with convergence for near objects and divergence for distant ones. In contrast, auditory cues provide complementary depth information through variations in loudness, interaural differences (IAD), and the frequency spectrum. We investigate the interaction between visual and auditory cues and examine how contradictory auditory information can overdrive visual depth perception in virtual reality (VR). When a new visual target appears, we introduce a spatial discrepancy between the visual and auditory cues: the visual target is shifted closer to the previously fixated object, while the corresponding sound localization is displaced in the opposite direction. By integrating these conflicting cues through multimodal processing, the resulting percept is biased toward the intended depth location. This audiovisual fusion counteracts depth compression, thus reducing the required vergence magnitude and enabling faster gaze retargeting. Such audio-driven depth enhancement may further help mitigate the vergence-accommodation conflict (VAC) in scenarios where physical depth must be compressed. In a series of psychophysical studies, we first assess the efficiency of depth overdriving for various VR-relevant combinations of initial fixations and shifted target locations, considering different scenarios of audio displacements and their loudness and frequency parameters. Next, we quantify the resulting speedup in gaze retargeting for target shifts that can be successfully overdriven by sound manipulations. Finally, we apply our method in a naturalistic VR scenario where user interface interactions with the scene show an extended perceptual depth.
Daniel Jiménez Navarro, Colin Groth, Jorge Pina, Qi Sun 0003, Praneeth Chakravarthula, Karol Myszkowski, Hans-Peter Seidel, Ana Serrano
IEEE Trans. Vis. Comput. Graph.9
2025 PreciseCam: Precise Camera Control for Text-to-Image Generation
abstract
Images as an artistic medium often rely on specific camera angles and lens distortions to convey ideas or emotions; however, such precise control is missing in current text-to-image models. We propose an efficient and general solution that allows precise control over the camera when generating both photographic and artistic images. Unlike prior methods that rely on predefined shots, we rely solely on four simple extrinsic and intrinsic camera parameters, removing the need for pre-existing geometry, reference 3D objects, and multi-view data. We also present a novel dataset with more than 57,000 images, along with their text prompts and ground-truth camera parameters. Our evaluation shows precise camera control in text-to-image generation, surpassing traditional prompt engineering approaches.
Edurne Bernal-Berdun, Ana Serrano, Belén Masiá, Matheus Gadelha, Yannick Hold-Geoffroy, Xin Sun 0014, Diego Gutierrez
CVPR2
2025 Audiovisual Disparities in VR: Impact on Spatial Perception
abstract
Virtual reality (VR) experiences often leverage rich and spatialized multimodal environments to increase immersion and engagement. This demands a consistent spatial perception of audiovisual stimuli, since perceived discrepancies can disrupt the sense of presence. In this work, we investigate the consequences of two types of spatial audiovisual disparities: true disparity, where there is a measurable spatial offset between auditory and visual cues, and perceptual disparity, where users report misalignment despite cues being colocated. Unlike most previous studies that employed controlled but simplified experimental setups, our research focuses on complex, realistic VR environments, allowing us to assess the actual implications for VR content design. Our experiments indicate that users are highly sensitive to true audiovisual disparities in controlled environments, detecting even minor misalignments. However, when engaged in additional tasks within realistic settings, their ability to notice such discrepancies diminishes significantly. We also observed that previously found perceptual disparities persist in complex audiovisual environments. However, we identify self-initiated head rotations as a key factor; its absence prevents the effect entirely. We hope our findings offer practical insights for designing more immersive and perceptually coherent VR experiences.
Edurne Bernal-Berdun, Mateo Vallejo, Qi Sun 0003, Ana Serrano, Diego Gutierrez
ISMAR4
2025 Why Slow Feels Fast and Fast Feels Slow: Evaluating and Predicting Speed Misperception
abstract
Human perception of speed is largely driven by visual cues. However, our subjective estimations of speed are influenced by several factors that can lead to deceptive cues and speed misperception. While some prior studies have explored individual effects on speed perception, such as contrast, spatial frequency, and temporal frequency, their combined influence remains underexamined, particularly in immersive VR environments. In this work, we systematically investigate the influence and interplay of four visual factors—contrast, spatial frequency, temporal frequency, and eccentricity—on human perception of speed. To this end, we conduct a psychophysical study measuring subjective speed judgments across controlled stimuli and reveal significant perceptual biases induced by these factors. Based on our collected data, we learn a model to predict the underestimation or overestimation of perceived speed from visual scene properties. We apply and validate our findings in three immersive environments and demonstrate their influence on common VR scenarios. Finally, we discuss how understanding the factors that shape speed perception can drive the design of perceptually aligned virtual environments, with potential future applications such as correcting speed misperception and conceivably mitigating visual-vestibular conflicts by modulating perceived speed.
Colin Groth, Daniel Jiménez Navarro, Zihao Zou, Ana Serrano, Karol Myszkowski, Qi Sun 0003, Praneeth Chakravarthula
ISMAR6
2025 A Comprehensive Analysis of the Influence of Cognitive Load on Physiological Signals in Virtual Reality
abstract
The study of cognitive load (CL) has been an active field of research across disciplines such as psychology, education, and computer graphics and visualization for decades. In the context of Virtual Reality (VR), understanding mental demand becomes particularly relevant, as immersive experiences increasingly integrate multisensory stimuli that require users to distribute their limited cognitive resources. In this work, we investigate the effects of cognitive load during a search task in VR, combining objective and subjective measurements, including physiological signals and validated questionnaires. We designed an experiment in which participants performed a visual search task under two cognitive load conditions (either alone or while responding to a concurrent auditory task) and across two visual search areas (90° and 360°). We collected a rich dataset comprising task performance, eye tracking, electrocardiogram (ECG), electrodermal activity (EDA), photoplethysmography (PPG), and inertial measurements, along with subjective assessments (NASA-TLX questionnaires). Our analysis shows that increased cognitive load hinders visual search performance and affects multiple physiological markers, offering a solid foundation for future research on cognitive load in multisensory virtual environments.
Jorge Pina, Edurne Bernal-Berdun, Sandra Malpica, Carmen Real, Alberto Barquero, Pablo Armañac-Julián, Jesús Lázaro 0002, Alba Martín-Yebra, Belén Masiá, Ana Serrano
ISMAR11
2025 Artist-Inator: Text-based, Gloss-aware Non-photorealistic Stylization
abstract
Abstract Large diffusion models have made a remarkable leap synthesizing high‐quality artistic images from text descriptions. However, these powerful pre‐trained models still lack control to guide key material appearance properties, such as gloss. In this work, we present a threefold contribution: (1) we analyze how gloss is perceived across different artistic styles (i.e., oil painting, watercolor, ink pen, charcoal, and soft crayon); (2) we leverage our findings to create a dataset with 1,336,272 stylized images of many different geometries in all five styles, including automatically‐computed text descriptions of their appearance (e.g., “A glossy bunny hand painted with an orange soft crayon”); and (3) we train ControlNet to condition Stable Diffusion XL synthesizing novel painterly depictions of new objects, using simple inputs such as edge maps, hand‐drawn sketches, or clip arts. Compared to previous approaches, our framework yields more accurate results despite the simplified input, as we show both quantitative and qualitatively.
J. Daniel Subias, Saul Daniel-soriano, Diego Gutierrez, Ana Serrano
Comput. Graph. Forum4
2025 Minimally disruptive auditory cues: their impact on visual performance in virtual reality
abstract
Abstract Virtual reality (VR) has the potential to become a revolutionary technology with a significant impact on our daily lives. The immersive experience provided by VR equipment, where the user’s body and senses are used to interact with the surrounding content, accompanied by the feeling of presence elicits a realistic behavioral response. In this work, we leverage the full control of audiovisual cues provided by VR to study an audiovisual suppression effect (ASE) where auditory stimuli degrade visual performance. In particular, we explore if barely audible sounds (in the range of the limits of hearing frequencies) generated following a specific spatiotemporal setup can still trigger the ASE while participants are experiencing high cognitive loads. A first study is carried out to find out how sound volume and frequency can impact this suppression effect, while the second study includes higher cognitive load scenarios closer to real applications. Our results show that the ASE is robust to variations in frequency, volume and cognitive load, achieving a reduction of visual perception with the proposed hardly audible sounds. Using such auditory cues means that this effect could be used in real applications, from entertaining to VR techniques like redirected walking.
Daniel Jiménez Navarro, Ana Serrano, Sandra Malpica
Vis. Comput.2
2025 Correction to: Minimally disruptive auditory cues: their impact on visual performance in virtual reality
Daniel Jiménez Navarro, Ana Serrano, Sandra Malpica
Vis. Comput.2
2024 AViSal360: Audiovisual Saliency Prediction for 360° Video
abstract
Saliency prediction in 360° video plays an important role in modeling visual attention, and can be leveraged for content creation, compression techniques, or quality assessment methods, among others. Visual attention in immersive environments depends not only on visual input, but also on inputs from other sensory modalities, primarily audio. Despite this, only a minority of saliency prediction models have incorporated auditory inputs, and much remains to be explored about what auditory information is relevant and how to integrate it in the prediction. In this work, we propose an audiovisual saliency model for 360° video content, AViSal360. Our model integrates both spatialized and semantic audio information, together with visual inputs. We perform exhaustive comparisons to demonstrate both the actual relevance of auditory information in saliency prediction, and the superior performance of our model when compared to previous approaches.
Edurne Bernal-Berdun, Jorge Pina, Mateo Vallejo, Ana Serrano, Belén Masiá
ISMAR4
2024 VGTC Virtual Reality Significant New Researcher Award
abstract
The 2024 IEEE VGTC Virtual Reality Significant New Researcher Award goes to Professor Ana Serrano of the Universidad de Zaragoza (Spain), in recognition of her outstanding contributions to virtual reality, computational imaging, and perception of material appearance. Her work has greatly impacted the community’s understanding of attention and viewer behavior in virtual reality. In addition to investigating saliency in virtual reality, Dr. Serrano has introduced a generative model of realistic scanpaths for 360° images that represent how users explore a virtual environment. Her other contributions include movie editing in virtual reality and its effects on the perception of continuity, perception-based material modeling such as gloss management, and understanding the effects of shape and illumination on the perception of material appearance. Her work has produced many high-profile publications and received numerous accolades within and beyond the field of virtual reality.
Ana Serrano
VR1
2024 Foreword to the special section on Spanish Computer Graphics Conference 2024
Ana Serrano, Gustavo Patow, Julio Marco
Comput. Graph.1
2024 Predicting Perceived Gloss: Do Weak Labels Suffice?
abstract
Abstract Estimating perceptual attributes of materials directly from images is a challenging task due to their complex, not fully‐understood interactions with external factors, such as geometry and lighting. Supervised deep learning models have recently been shown to outperform traditional approaches, but rely on large datasets of human‐annotated images for accurate perception predictions. Obtaining reliable annotations is a costly endeavor, aggravated by the limited ability of these models to generalise to different aspects of appearance. In this work, we show how a much smaller set of human annotations (“strong labels”) can be effectively augmented with automatically derived “weak labels” in the context of learning a low‐dimensional image‐computable gloss metric. We evaluate three alternative weak labels for predicting human gloss perception from limited annotated data. Incorporating weak labels enhances our gloss prediction beyond the current state of the art. Moreover, it enables a substantial reduction in human annotation costs without sacrificing accuracy, whether working with rendered images or real photographs.
Julia Guerrero-Viu, J. Daniel Subias, Ana Serrano, Katherine Storrs, Roland W. Fleming, Belén Masiá, Diego Gutierrez
Comput. Graph. Forum3
2024 Cinematic Gaussians: Real-Time HDR Radiance Fields with Depth of Field
abstract
Abstract Radiance field methods represent the state of the art in reconstructing complex scenes from multi‐view photos. However, these reconstructions often suffer from one or both of the following limitations: First, they typically represent scenes in low dynamic range (LDR), which restricts their use to evenly lit environments and hinders immersive viewing experiences. Secondly, their reliance on a pinhole camera model, assuming all scene elements are in focus in the input images, presents practical challenges and complicates refocusing during novel‐view synthesis. Addressing these limitations, we present a lightweight method based on 3D Gaussian Splatting that utilizes multi‐view LDR images of a scene with varying exposure times, apertures, and focus distances as input to reconstruct a high‐dynamic‐range (HDR) radiance field. By incorporating analytical convolutions of Gaussians based on a thin‐lens camera model as well as a tonemapping module, our reconstructions enable the rendering of HDR content with flexible refocusing capabilities. We demonstrate that our combined treatment of HDR and depth of field facilitates real‐time cinematic rendering, outperforming the state of the art.
Chao Wang 0037, Krzysztof Wolski, Bernhard Kerbl, Ana Serrano, Mojtaba Bemana, Hans-Peter Seidel, Karol Myszkowski, Thomas Leimkühler
Comput. Graph. Forum4
2024 Modeling the Impact of Head-Body Rotations on Audio-Visual Spatial Perception for Virtual Reality Applications
abstract
Humans perceive the world by integrating multimodal sensory feedback, including visual and auditory stimuli, which holds true in virtual reality (VR) environments. Proper synchronization of these stimuli is crucial for perceiving a coherent and immersive VR experience. In this work, we focus on the interplay between audio and vision during localization tasks involving natural head-body rotations. We explore the impact of audio-visual offsets and rotation velocities on users' directional localization acuity for various viewing modes. Using psychometric functions, we model perceptual disparities between visual and auditory cues and determine offset detection thresholds. Our findings reveal that target localization accuracy is affected by perceptual audio-visual disparities during head-body rotations, but remains consistent in the absence of stimuli-head relative motion. We then showcase the effectiveness of our approach in predicting and enhancing users' localization accuracy within realistic VR gaming applications. To provide additional support for our findings, we implement a natural VR game wherein we apply a compensatory audio-visual offset derived from our measured psychometric functions. As a result, we demonstrate a substantial improvement of up to 40% in participants' target localization accuracy. We additionally provide guidelines for content creation to ensure coherent and seamless VR experiences.
Edurne Bernal-Berdun, Mateo Vallejo, Qi Sun 0003, Ana Serrano, Diego Gutierrez
IEEE Trans. Vis. Comput. Graph.4
2024 Measuring and Predicting Multisensory Reaction Latency: A Probabilistic Model for Visual-Auditory Integration
abstract
Virtual/augmented reality (VR/AR) devices offer both immersive imagery and sound. With those wide-field cues, we can simultaneously acquire and process visual and auditory signals to quickly identify objects, make decisions, and take action. While vision often takes precedence in perception, our visual sensitivity degrades in the periphery. In contrast, auditory sensitivity can exhibit an opposite trend due to the elevated interaural time difference. What occurs when these senses are simultaneously integrated, as is common in VR applications such as 360° video watching and immersive gaming? We present a computational and probabilistic model to predict VR users' reaction latency to visual-auditory multisensory targets. To this aim, we first conducted a psychophysical experiment in VR to measure the reaction latency by tracking the onset of eye movements. Experiments with numerical metrics and user studies with naturalistic scenarios showcase the model's accuracy and generalizability. Lastly, we discuss the potential applications, such as measuring the sufficiency of target appearance duration in immersive video playback, and suggesting the optimal spatial layouts for AR interface design.
Daniel Jiménez Navarro, Ana Serrano, Karol Myszkowski, Qi Sun 0003
IEEE Trans. Vis. Comput. Graph.4
2024 Studying the Effect of Material and Geometry on Perceptual Outdoor Illumination
abstract
Understanding and modeling perceived properties of sky-dome illumination is an important but challenging problem due to the interplay of several factors such as the materials and geometries of the objects present in the scene being observed. Existing models of sky-dome illumination focus on the physical properties of the sky. However, these parametric models often do not align well with the properties perceived by a human observer. In this work, drawing inspiration from the Hosek-Wilkie sky-dome model, we investigate the perceptual properties of outdoor illumination. For this purpose, we perform a large-scale user study via crowdsourcing to collect a dataset of perceived illumination properties (scattering, glare, and brightness) for different combinations of geometries and materials under a variety of outdoor illuminations, totaling 5,000 distinct images. We perform a thorough statistical analysis of the collected data which reveals several interesting effects. For instance, our analysis shows that when there are objects in the scene made of rough materials, the perceived scattering of the sky increases. Furthermore, we utilize our extensive collection of images and their corresponding perceptual attributes to train a predictor. This predictor, when provided with a single image as input, generates an estimation of perceived illumination properties that align with human perceptual judgments. Accurately estimating perceived illumination properties can greatly enhance the overall quality of integrating virtual objects into real scene photographs. Consequently, we showcase various applications of our predictor. For instance, we demonstrate its utility as a luminance editing tool for showcasing virtual objects in outdoor scenes.
Miao Wang 0004, Jinchao Zhou, Wei-Qi Feng, Yu-Zhu Jiang, Ana Serrano
IEEE Trans. Vis. Comput. Graph.5
2024 SAL3D: a model for saliency prediction in 3D meshes
abstract
Advances in virtual and augmented reality have increased the demand for immersive and engaging 3D experiences. To create such experiences, it is crucial to understand visual attention in 3D environments, which is typically modeled by means of saliency maps. While attention in 2D images and traditional media has been widely studied, there is still much to explore in 3D settings. In this work, we propose a deep learning-based model for predicting saliency when viewing 3D objects, which is a first step toward understanding and predicting attention in 3D environments. Previous approaches rely solely on low-level geometric cues or unnatural conditions, however, our model is trained on a dataset of real viewing data that we have manually captured, which indeed reflects actual human viewing behavior. Our approach outperforms existing state-of-the-art methods and closely approximates the ground-truth data. Our results demonstrate the effectiveness of our approach in predicting attention in 3D objects, which can pave the way for creating more immersive and engaging 3D experiences.
Andres Fandos, Belén Masiá, Ana Serrano
Vis. Comput.4
2023 GlowGAN: Unsupervised Learning of HDR Images from LDR Images in the Wild
abstract
Most in-the-wild images are stored in Low Dynamic Range (LDR) form, serving as a partial observation of the High Dynamic Range (HDR) visual world. Despite limited dynamic range, these LDR images are often captured with different exposures, implicitly containing information about the underlying HDR image distribution. Inspired by this intuition, in this work we present, to the best of our knowledge, the first method for learning a generative model of HDR images from in-the-wild LDR image collections in a fully unsupervised manner. The key idea is to train a generative adversarial network (GAN) to generate HDR images which, when projected to LDR under various exposures, are indistinguishable from real LDR images. The projection from HDR to LDR is achieved via a camera model that captures the stochasticity in exposure and camera response function. Experiments show that our method GlowGAN can synthesize photorealistic HDR images in many challenging cases such as landscapes, lightning, or windows, where previous supervised generative models produce overexposed images. With the assistance of GlowGAN, we showcase the novel application of unsupervised inverse tone mapping (GlowGAN-ITM) that sets a new paradigm in this field. Unlike previous methods that gradually complete information from LDR input, GlowGAN-ITM searches the entire HDR image manifold modeled by GlowGAN for the HDR images which can be mapped back to the LDR input. GlowGAN-ITM achieves more realistic reconstruction of overexposed regions compared to state-of-the-art supervised learning models, despite not requiring HDR images or paired multi-exposure images for training.
Chao Wang 0037, Ana Serrano, Xingang Pan, Bin Chen 0019, Karol Myszkowski, Hans-Peter Seidel, Christian Theobalt, Thomas Leimkühler
ICCV2
2023 The effect of display capabilities on the gloss consistency between real and virtual objects
abstract
A faithful reproduction of gloss is inherently difficult because of the limited dynamic range, peak luminance, and 3D capabilities of display devices. This work investigates how the display capabilities affect gloss appearance with respect to a real-world reference object. To this end, we employ an accurate imaging pipeline to achieve a perceptual gloss match between a virtual and real object presented side-by-side on an augmented-reality high-dynamic-range (HDR) stereoscopic display, which has not been previously attained to this extent. Based on this precise gloss reproduction, we conduct a series of gloss matching experiments to study how gloss perception degrades based on individual factors: object albedo, display luminance, dynamic range, stereopsis, and tone mapping. We support the study with a detailed analysis of individual factors, followed by an in-depth discussion on the observed perceptual effects. Our experiments demonstrate that stereoscopic presentation has a limited effect on the gloss matching task on our HDR display. However, both reduced luminance and dynamic range of the display reduce the perceived gloss. This means that the visual system cannot compensate for the changes in gloss appearance across luminance (lack of gloss constancy), and the tone mapping operator should be carefully selected when reproducing gloss on a low dynamic range (LDR) display.
Bin Chen 0019, Akshay Jindal, Michal Piovarci, Chao Wang 0037, Hans-Peter Seidel, Piotr Didyk, Karol Myszkowski, Ana Serrano, Rafal Mantiuk
SIGGRAPH Asia8
2023 Foreword to the Special Section on CEIG 2023
Ana Serrano, Marc Comino, Jesús Gimeno
Comput. Graph.1
2023 An Implicit Neural Representation for the Image Stack: Depth, All in Focus, and High Dynamic Range
abstract
In everyday photography, physical limitations of camera sensors and lenses frequently lead to a variety of degradations in captured images such as saturation or defocus blur. A common approach to overcome these limitations is to resort to image stack fusion, which involves capturing multiple images with different focal distances or exposures. For instance, to obtain an all-in-focus image, a set of multi-focus images is captured. Similarly, capturing multiple exposures allows for the reconstruction of high dynamic range. In this paper, we present a novel approach that combines neural fields with an expressive camera model to achieve a unified reconstruction of an all-in-focus high-dynamic-range image from an image stack. Our approach is composed of a set of specialized implicit neural representations tailored to address specific sub-problems along our pipeline: We use neural implicits to predict flow to overcome misalignments arising from lens breathing, depth, and all-in-focus images to account for depth of field, as well as tonemapping to deal with sensor responses and saturation - all trained using a physically inspired supervision structure with a differentiable thin lens model at its core. An important benefit of our approach is its ability to handle these tasks simultaneously or independently, providing flexible post-editing capabilities such as refocusing and exposure adjustment. By sampling the three primary factors in photography within our framework (focal distance, aperture, and exposure time), we conduct a thorough exploration to gain valuable insights into their significance and impact on overall reconstruction quality. Through extensive validation, we demonstrate that our method outperforms existing approaches in both depth-from-defocus and all-in-focus image reconstruction tasks. Moreover, our approach exhibits promising results in each of these three dimensions, showcasing its potential to enhance captured image quality and provide greater control in post-processing.
Chao Wang 0037, Ana Serrano, Xingang Pan, Krzysztof Wolski, Bin Chen 0019, Karol Myszkowski, Hans-Peter Seidel, Christian Theobalt, Thomas Leimkühler
ACM Trans. Graph.2
2023 D-SAV360: A Dataset of Gaze Scanpaths on 360° Ambisonic Videos
abstract
Understanding human visual behavior within virtual reality environments is crucial to fully leverage their potential. While previous research has provided rich visual data from human observers, existing gaze datasets often suffer from the absence of multimodal stimuli. Moreover, no dataset has yet gathered eye gaze trajectories (i.e., scanpaths) for dynamic content with directional ambisonic sound, which is a critical aspect of sound perception by humans. To address this gap, we introduce D-SAV360, a dataset of 4,609 head and eye scanpaths for 360° videos with first-order ambisonics. This dataset enables a more comprehensive study of multimodal interaction on visual behavior in virtual reality environments. We analyze our collected scanpaths from a total of 87 participants viewing 85 different videos and show that various factors such as viewing mode, content type, and gender significantly impact eye movement statistics. We demonstrate the potential of D-SAV360 as a benchmarking resource for state-of-the-art attention prediction models and discuss its possible applications in further research. By providing a comprehensive dataset of eye movement data for dynamic, multimodal virtual environments, our work can facilitate future investigations of visual behavior and attention in virtual reality.
Edurne Bernal-Berdun, Sandra Malpica, Pedro J. Perez, Diego Gutierrez, Belén Masiá, Ana Serrano
IEEE Trans. Vis. Comput. Graph.7
2023 Task-Dependent Visual Behavior in Immersive Environments: A Comparative Study of Free Exploration, Memory and Visual Search
abstract
Visual behavior depends on both bottom-up mechanisms, where gaze is driven by the visual conspicuity of the stimuli, and top-down mechanisms, guiding attention towards relevant areas based on the task or goal of the viewer. While this is well-known, visual attention models often focus on bottom-up mechanisms. Existing works have analyzed the effect of high-level cognitive tasks like memory or visual search on visual behavior; however, they have often done so with different stimuli, methodology, metrics and participants, which makes drawing conclusions and comparisons between tasks particularly difficult. In this work we present a systematic study of how different cognitive tasks affect visual behavior in a novel within-subjects design scheme. Participants performed free exploration, memory and visual search tasks in three different scenes while their eye and head movements were being recorded. We found significant, consistent differences between tasks in the distributions of fixations, saccades and head movements. Our findings can provide insights for practitioners and content creators designing task-oriented immersive applications.
Sandra Malpica, Ana Serrano, Diego Gutierrez, Belén Masiá
IEEE Trans. Vis. Comput. Graph.3
2022 Gloss management for consistent reproduction of real and virtual objects
abstract
A good match of material appearance between real-world objects and their digital on-screen representations is critical for many applications such as fabrication, design, and e-commerce. However, faithful appearance reproduction is challenging, especially for complex phenomena, such as gloss. In most cases, the view-dependent nature of gloss and the range of luminance values required for reproducing glossy materials exceeds the current capabilities of display devices. As a result, appearance reproduction poses significant problems even with accurately rendered images. This paper studies the gap between the gloss perceived from real-world objects and their digital counterparts. Based on our psychophysical experiments on a wide range of 3D printed samples and their corresponding photographs, we derive insights on the influence of geometry, illumination, and the display’s brightness and measure the change in gloss appearance due to the display limitations. Our evaluation experiments demonstrate that using the prediction to correct material parameters in a rendering system improves the match of gloss appearance between real objects and their visualization on a display device.
Bin Chen 0019, Michal Piovarci, Chao Wang 0037, Hans-Peter Seidel, Piotr Didyk, Karol Myszkowski, Ana Serrano
SIGGRAPH Asia7
2022 Foreword to the Special Section on CEIG 2022
Ana Serrano, Jorge Posada 0001, Miguel A. Otaduy
Comput. Graph.1
2022 Learning a self-supervised tone mapping operator via feature contrast masking loss
abstract
Abstract High Dynamic Range (HDR) content is becoming ubiquitous due to the rapid development of capture technologies. Nevertheless, the dynamic range of common display devices is still limited, therefore tone mapping (TM) remains a key challenge for image visualization. Recent work has demonstrated that neural networks can achieve remarkable performance in this task when compared to traditional methods, however, the quality of the results of these learning‐based methods is limited by the training data. Most existing works use as training set a curated selection of best‐performing results from existing traditional tone mapping operators (often guided by a quality metric), therefore, the quality of newly generated results is fundamentally limited by the performance of such operators. This quality might be even further limited by the pool of HDR content that is used for training. In this work we propose a learning‐based self‐supervised tone mapping operator that is trained at test time specifically for each HDR image and does not need any data labeling. The key novelty of our approach is a carefully designed loss function built upon fundamental knowledge on contrast perception that allows for directly comparing the content in the HDR and tone mapped images. We achieve this goal by reformulating classic VGG feature maps into feature contrast maps that normalize local feature differences by their average magnitude in a local neighborhood, allowing our loss to account for contrast masking effects. We perform extensive ablation studies and exploration of parameters and demonstrate that our solution outperforms existing approaches with a single set of fixed parameters, as confirmed by both objective and subjective metrics.
Chao Wang 0037, Bin Chen 0019, Hans-Peter Seidel, Karol Myszkowski, Ana Serrano
Comput. Graph. Forum5
2022 Modelling Surround-aware Contrast Sensitivity for HDR Displays
abstract
Abstract Despite advances in display technology, many existing applications rely on psychophysical datasets of human perception gathered using older, sometimes outdated displays. As a result, there exists the underlying assumption that such measurements can be carried over to the new viewing conditions of more modern technology. We have conducted a series of psychophysical experiments to explore contrast sensitivity using a state‐of‐the‐art HDR display, taking into account not only the spatial frequency and luminance of the stimuli but also their surrounding luminance levels. From our data, we have derived a novel surround‐aware contrast sensitivity function (CSF), which predicts human contrast sensitivity more accurately. We additionally provide a practical version that retains the benefits of our full model, while enabling easy backward compatibility and consistently producing good results across many existing applications that make use of CSF models. We show examples of effective HDR video compression using a transfer function derived from our CSF, tone‐mapping and improved accuracy in visual difference prediction.
Shinyoung Yi 0001, Daniel S. Jeon, Ana Serrano, Seyoon Jeong, Hui Yong Kim, Diego Gutierrez, Min H. Kim 0001
Comput. Graph. Forum3
2022 Introduction to the Special Issue on SAP 2022
abstract
No abstract available.
Ana Serrano, Michael Barnett-Cowan
ACM Trans. Appl. Percept.1
2022 ScanGAN360: A Generative Model of Realistic Scanpaths for 360° Images
abstract
Understanding and modeling the dynamics of human gaze behavior in 360° environments is crucial for creating, improving, and developing emerging virtual reality applications. However, recruiting human observers and acquiring enough data to analyze their behavior when exploring virtual environments requires complex hardware and software setups, and can be time-consuming. Being able to generate virtual observers can help overcome this limitation, and thus stands as an open problem in this medium. Particularly, generative adversarial approaches could alleviate this challenge by generating a large number of scanpaths that reproduce human behavior when observing new scenes, essentially mimicking virtual observers. However, existing methods for scanpath generation do not adequately predict realistic scanpaths for 360° images. We present ScanGAN360, a new generative adversarial approach to address this problem. We propose a novel loss function based on dynamic time warping and tailor our network to the specifics of 360° images. The quality of our generated scanpaths outperforms competing approaches by a large margin, and is almost on par with the human baseline. ScanGAN360 allows fast simulation of large numbers of virtual observers, whose behavior mimics real users, enabling a better understanding of gaze behavior, facilitating experimentation, and aiding novel applications in virtual reality and beyond.
Ana Serrano, Alexander W. Bergman, Gordon Wetzstein, Belén Masiá
IEEE Trans. Vis. Comput. Graph.2
2021 The effect of shape and illumination on material perception: model and applications
abstract
Material appearance hinges on material reflectance properties but also surface geometry and illumination. The unlimited number of potential combinations between these factors makes understanding and predicting material appearance a very challenging task. In this work, we collect a large-scale dataset of perceptual ratings of appearance attributes with more than 215,680 responses for 42,120 distinct combinations of material, shape, and illumination. The goal of this dataset is twofold. First, we analyze for the first time the effects of illumination and geometry in material perception across such a large collection of varied appearances. We connect our findings to those of the literature, discussing how previous knowledge generalizes across very diverse materials, shapes, and illuminations. Second, we use the collected dataset to train a deep learning architecture for predicting perceptual attributes that correlate with human judgments. We demonstrate the consistent and robust behavior of our predictor in various challenging scenarios, which, for the first time, enables estimating perceived material attributes from general 2D images. Since our predictor relies on the final appearance in an image, it can compare appearance properties across different geometries and illumination conditions. Finally, we demonstrate several applications that use our predictor, including appearance reproduction using 3D printing, BRDF editing by integrating our predictor in a differentiable renderer, illumination design, or material recommendations for scene design.
Ana Serrano, Bin Chen 0019, Chao Wang 0037, Michal Piovarci, Hans-Peter Seidel, Piotr Didyk, Karol Myszkowski
ACM Trans. Graph.1
2021 The effect of geometry and illumination on appearance perception of different material categories
abstract
Abstract The understanding of material appearance perception is a complex problem due to interactions between material reflectance, surface geometry, and illumination. Recently, Serrano et al. collected the largest dataset to date with subjective ratings of material appearance attributes, including glossiness, metallicness, sharpness and contrast of reflections. In this work, we make use of their dataset to investigate for the first time the impact of the interactions between illumination, geometry, and eight different material categories in perceived appearance attributes. After an initial analysis, we select for further analysis the four material categories that cover the largest range for all perceptual attributes: fabric, plastic, ceramic, and metal. Using a cumulative link mixed model (CLMM) for robust regression, we discover interactions between these material categories and four representative illuminations and object geometries. We believe that our findings contribute to expanding the knowledge on material appearance perception and can be useful for many applications, such as scene design, where any particular material in a given shape can be aligned with dominant classes of illumination, so that a desired strength of appearance attributes can be achieved.
Bin Chen 0019, Chao Wang 0037, Michal Piovarci, Hans-Peter Seidel, Piotr Didyk, Karol Myszkowski, Ana Serrano
Vis. Comput.7
2020 Exploring the impact of 360° movie cuts in users' attention
abstract
Virtual Reality (VR) has grown since the first devices for personal use became available on the market. However, the production of cinematographic content in this new medium is still in an early exploratory phase. The main reason is that cinematographic language in VR is still under development, and we still need to learn how to tell stories effectively. A key element in traditional film editing is the use of different cutting techniques, in order to transition seamlessly from one sequence to another. A fundamental aspect of these techniques is the placement and control over the camera. However, VR content creators do not have full control of the camera. Instead, users in VR can freely explore the 360° of the scene around them, which potentially leads to very different experiences. While this is desirable in certain applications such as VR games, it may hinder the experience in narrative VR. In this work, we perform a systematic analysis of users’ viewing behavior across cut boundaries while watching professionally edited, narrative 360° videos. We extend previous metrics for quantifying user behavior in order to support more complex and realistic footage, and we introduce two new metrics that allow us to measure users’ exploration in a variety of different complex scenarios. From this analysis, (i) we confirm that previous insights derived for simple content hold for professionally edited content, and (ii) we derive new insights that could potentially influence VR content creation, informing creators about the impact of different cuts in the audience’s behavior.
Carlos Marañes, Diego Gutierrez, Ana Serrano
VR3
2020 Crossmodal perception in virtual reality
Sandra Malpica, Ana Serrano, Marcos Allue, Manuel G. Bedia, Belén Masiá
Multim. Tools Appl.2
2020 Imperceptible manipulation of lateral camera motion for improved virtual reality applications
abstract
Virtual Reality (VR) systems increase immersion by reproducing users' movements in the real world. However, several works have shown that this real-to-virtual mapping does not need to be precise in order to convey a realistic experience. Being able to alter this mapping has many potential applications, since achieving an accurate real-to-virtual mapping is not always possible due to limitations in the capture or display hardware, or in the physical space available. In this work, we measure detection thresholds for lateral translation gains of virtual camera motion in response to the corresponding head motion under natural viewing, and in the absence of locomotion, so that virtual camera movement can be either compressed or expanded while these manipulations remain undetected. Finally, we propose three applications for our method, addressing three key problems in VR: improving 6-DoF viewing for captured 360° footage, overcoming physical constraints, and reducing simulator sickness. We have further validated our thresholds and evaluated our applications by means of additional user studies confirming that our manipulations remain imperceptible, and showing that (i) compressing virtual camera motion reduces visible artifacts in 6-DoF, hence improving perceived quality, (ii) virtual expansion allows for completion of virtual tasks within a reduced physical space, and (iii) simulator sickness may be alleviated in simple scenarios when our compression method is applied.
Ana Serrano, Diego Gutierrez, Karol Myszkowski, Belén Masiá
ACM Trans. Graph.1
2019 A similarity measure for material appearance
abstract
We present a model to measure the similarity in appearance between different materials, which correlates with human similarity judgments. We first create a database of 9,000 rendered images depicting objects with varying materials, shape and illumination. We then gather data on perceived similarity from crowdsourced experiments; our analysis of over 114,840 answers suggests that indeed a shared perception of appearance similarity exists. We feed this data to a deep learning architecture with a novel loss function, which learns a feature space for materials that correlates with such perceived appearance similarity. Our evaluation shows that our model outperforms existing metrics. Last, we demonstrate several applications enabled by our metric, including appearance-based search for material suggestions, database visualization, clustering and summarization, and gamut mapping.
Manuel Lagunas, Sandra Malpica, Ana Serrano, Elena Garces 0001, Diego Gutierrez, Belén Masiá
ACM Trans. Graph.3
2019 Motion parallax for 360° RGBD video
abstract
We present a method for adding parallax and real-time playback of 360° videos in Virtual Reality headsets. In current video players, the playback does not respond to translational head movement, which reduces the feeling of immersion, and causes motion sickness for some viewers. Given a 360° video and its corresponding depth (provided by current stereo 360° stitching algorithms), a naive image-based rendering approach would use the depth to generate a 3D mesh around the viewer, then translate it appropriately as the viewer moves their head. However, this approach breaks at depth discontinuities, showing visible distortions, whereas cutting the mesh at such discontinuities leads to ragged silhouettes and holes at disocclusions. We address these issues by improving the given initial depth map to yield cleaner, more natural silhouettes. We rely on a three-layer scene representation, made up of a foreground layer and two static background layers, to handle disocclusions by propagating information from multiple frames for the first background layer, and then inpainting for the second one. Our system works with input from many of today's most popular 360° stereo capture devices (e.g., Yi Halo or GoPro Odyssey), and works well even if the original video does not provide depth information. Our user studies confirm that our method provides a more compelling viewing experience than without parallax, increasing immersion while reducing discomfort and nausea.
Ana Serrano, Inchul Kim 0001, Stephen DiVerdi, Diego Gutierrez, Aaron Hertzmann, Belén Masiá
IEEE Trans. Vis. Comput. Graph.1
2018 Saliency in VR: How Do People Explore Virtual Environments?
abstract
Understanding how people explore immersive virtual environments is crucial for many applications, such as designing virtual reality (VR) content, developing new compression algorithms, or learning computational models of saliency or visual attention. Whereas a body of recent work has focused on modeling saliency in desktop viewing conditions, VR is very different from these conditions in that viewing behavior is governed by stereoscopic vision and by the complex interaction of head orientation, gaze, and other kinematic constraints. To further our understanding of viewing behavior and saliency in VR, we capture and analyze gaze and head orientation data of 169 users exploring stereoscopic, static omni-directional panoramas, for a total of 1980 head and gaze trajectories for three different viewing conditions. We provide a thorough analysis of our data, which leads to several important insights, such as the existence of a particular fixation bias, which we then use to adapt existing saliency predictors to immersive VR conditions. In addition, we explore other applications of our data and analysis, including automatic alignment of VR video cuts, panorama thumbnails, panorama video synopsis, and saliency-basedcompression.
Vincent Sitzmann, Ana Serrano, Amy Pavel, Maneesh Agrawala, Diego Gutierrez, Belén Masiá, Gordon Wetzstein
IEEE Trans. Vis. Comput. Graph.2
2017 Convolutional Sparse Coding for Capturing High-Speed Video Content
abstract
Abstract Video capture is limited by the trade‐off between spatial and temporal resolution: when capturing videos of high temporal resolution, the spatial resolution decreases due to bandwidth limitations in the capture system. Achieving both high spatial and temporal resolution is only possible with highly specialized and very expensive hardware, and even then the same basic trade‐off remains. The recent introduction of compressive sensing and sparse reconstruction techniques allows for the capture of single‐shot high‐speed video, by coding the temporal information in a single frame, and then reconstructing the full video sequence from this single‐coded image and a trained dictionary of image patches. In this paper, we first analyse this approach, and find insights that help improve the quality of the reconstructed videos. We then introduce a novel technique, based on convolutional sparse coding (CSC), and show how it outperforms the state‐of‐the‐art, patch‐based approach in terms of flexibility and efficiency, due to the convolutional nature of its filter banks. The key idea for CSC high‐speed video acquisition is extending the basic formulation by imposing an additional constraint in the temporal dimension, which enforces sparsity of the first‐order derivatives over time.
Ana Serrano, Elena Garces 0001, Belén Masiá, Diego Gutierrez
Comput. Graph. Forum1
2017 Attribute-preserving gamut mapping of measured BRDFs
abstract
Abstract Reproducing the appearance of real‐world materials using current printing technology is problematic. The reduced number of inks available define the printer's limited gamut, creating distortions in the printed appearance that are hard to control. Gamut mapping refers to the process of bringing an out‐of‐gamut material appearance into the printer's gamut, while minimizing such distortions as much as possible. We present a novel two‐step gamut mapping algorithm that allows users to specify which perceptual attribute of the original material they want to preserve (such as brightness, or roughness). In the first step, we work in the low‐dimensional intuitive appearance space recently proposed by Serrano et al. [ SGM*16 ], and adjust achromatic reflectance via an objective function that strives to preserve certain attributes. From such intermediate representation, we then perform an image‐based optimization including color information, to bring the BRDF into gamut. We show, both objectively and through a user study, how our method yields superior results compared to the state of the art, with the additional advantage that the user can specify which visual attributes need to be preserved. Moreover, we show how this approach can also be used for attribute‐preserving material editing.
Tiancheng Sun, Ana Serrano, Diego Gutierrez, Belén Masiá
Comput. Graph. Forum2
2017 Dynamic range expansion based on image statistics
Belén Masiá, Ana Serrano, Diego Gutierrez
Multim. Tools Appl.2
2017 Movie editing and cognitive event segmentation in virtual reality video
abstract
Traditional cinematography has relied for over a century on a well-established set of editing rules, called continuity editing, to create a sense of situational continuity. Despite massive changes in visual content across cuts, viewers in general experience no trouble perceiving the discontinuous flow of information as a coherent set of events. However, Virtual Reality (VR) movies are intrinsically different from traditional movies in that the viewer controls the camera orientation at all times. As a consequence, common editing techniques that rely on camera orientations, zooms, etc., cannot be used. In this paper we investigate key relevant questions to understand how well traditional movie editing carries over to VR, such as: Does the perception of continuity hold across edit boundaries? Under which conditions? Does viewers' observational behavior change after the cuts? To do so, we rely on recent cognition studies and the event segmentation theory, which states that our brains segment continuous actions into a series of discrete, meaningful events. We first replicate one of these studies to assess whether the predictions of such theory can be applied to VR. We next gather gaze data from viewers watching VR videos containing different edits with varying parameters, and provide the first systematic analysis of viewers' behavior and the perception of continuity in VR. From this analysis we make a series of relevant findings; for instance, our data suggests that predictions from the cognitive event segmentation theory are useful guides for VR editing; that different types of edits are equally well understood in terms of continuity; and that spatial misalignments between regions of interest at the edit boundaries favor a more exploratory behavior even after viewers have fixated on a new region of interest. In addition, we propose a number of metrics to describe viewers' attentional behavior in VR. We believe the insights derived from our work can be useful as guidelines for VR content creation.
Ana Serrano, Vincent Sitzmann, Jaime Ruiz-Borau, Gordon Wetzstein, Diego Gutierrez, Belén Masiá
ACM Trans. Graph.1
2016 Convolutional Sparse Coding for High Dynamic Range Imaging
abstract
Abstract Current HDR acquisition techniques are based on either (i) fusing multibracketed, low dynamic range (LDR) images, (ii) modifying existing hardware and capturing different exposures simultaneously with multiple sensors, or (iii) reconstructing a single image with spatially‐varying pixel exposures. In this paper, we propose a novel algorithm to recover high‐quality HDRI images from a single, coded exposure. The proposed reconstruction method builds on recently‐introduced ideas of convolutional sparse coding (CSC); this paper demonstrates how to make CSC practical for HDR imaging. We demonstrate that the proposed algorithm achieves higher‐quality reconstructions than alternative methods, we evaluate optical coding schemes, analyze algorithmic parameters, and build a prototype coded HDR camera that demonstrates the utility of convolutional sparse HDRI coding with a custom hardware platform.
Ana Serrano, Felix Heide, Diego Gutierrez, Gordon Wetzstein, Belén Masiá
Comput. Graph. Forum1
2016 An intuitive control space for material appearance
abstract
Many different techniques for measuring material appearance have been proposed in the last few years. These have produced large public datasets, which have been used for accurate, data-driven appearance modeling. However, although these datasets have allowed us to reach an unprecedented level of realism in visual appearance, editing the captured data remains a challenge. In this paper, we present an intuitive control space for predictable editing of captured BRDF data, which allows for artistic creation of plausible novel material appearances, bypassing the difficulty of acquiring novel samples. We first synthesize novel materials, extending the existing MERL dataset up to 400 mathematically valid BRDFs. We then design a large-scale experiment, gathering 56,000 subjective ratings on the high-level perceptual attributes that best describe our extended dataset of materials. Using these ratings, we build and train networks of radial basis functions to act as functionals mapping the perceptual attributes to an underlying PCA-based representation of BRDFs. We show that our functionals are excellent predictors of the perceived attributes of appearance. Our control space enables many applications, including intuitive material editing of a wide range of visual properties, guidance for gamut mapping, analysis of the correlation between perceptual attributes, or novel appearance similarity metrics. Moreover, our methodology can be used to derive functionals applicable to classic analytic BRDF representations. We release our code and dataset publicly, in order to support and encourage further research in this direction.
Ana Serrano, Diego Gutierrez, Karol Myszkowski, Hans-Peter Seidel, Belén Masiá
ACM Trans. Graph.1