EDBT 2026 Demo / reviewers in the wild / expert
Belén Masiá
dblp:24/7593
· DBLP profile ↗
42ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0003-0060-7278ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 5 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PreciseCam: Precise Camera Control for Text-to-Image GenerationabstractImages as an artistic medium often rely on specific camera angles and lens distortions to convey ideas or emotions; however, such precise control is missing in current text-to-image models. We propose an efficient and general solution that allows precise control over the camera when generating both photographic and artistic images. Unlike prior methods that rely on predefined shots, we rely solely on four simple extrinsic and intrinsic camera parameters, removing the need for pre-existing geometry, reference 3D objects, and multi-view data. We also present a novel dataset with more than 57,000 images, along with their text prompts and ground-truth camera parameters. Our evaluation shows precise camera control in text-to-image generation, surpassing traditional prompt engineering approaches. Edurne Bernal-Berdun, Ana Serrano, Belén Masiá, Matheus Gadelha, Yannick Hold-Geoffroy, Xin Sun 0014, Diego Gutierrez |
CVPR | 3 |
| 2025 | A Comprehensive Analysis of the Influence of Cognitive Load on Physiological Signals in Virtual RealityabstractThe study of cognitive load (CL) has been an active field of research across disciplines such as psychology, education, and computer graphics and visualization for decades. In the context of Virtual Reality (VR), understanding mental demand becomes particularly relevant, as immersive experiences increasingly integrate multisensory stimuli that require users to distribute their limited cognitive resources. In this work, we investigate the effects of cognitive load during a search task in VR, combining objective and subjective measurements, including physiological signals and validated questionnaires. We designed an experiment in which participants performed a visual search task under two cognitive load conditions (either alone or while responding to a concurrent auditory task) and across two visual search areas (90° and 360°). We collected a rich dataset comprising task performance, eye tracking, electrocardiogram (ECG), electrodermal activity (EDA), photoplethysmography (PPG), and inertial measurements, along with subjective assessments (NASA-TLX questionnaires). Our analysis shows that increased cognitive load hinders visual search performance and affects multiple physiological markers, offering a solid foundation for future research on cognitive load in multisensory virtual environments. Jorge Pina, Edurne Bernal-Berdun, Sandra Malpica, Carmen Real, Alberto Barquero, Pablo Armañac-Julián, Jesús Lázaro 0002, Alba Martín-Yebra, Belén Masiá, Ana Serrano |
ISMAR | 10 |
| 2025 | Fine-Grained Spatially Varying Material Selection in ImagesabstractSelection is the first step in many image editing processes, enabling faster and simpler modifications of all pixels sharing a common modality. In this work, we present a method for material selection in images, robust to lighting and reflectance variations, which can be used for downstream editing tasks. We rely on vision transformer (ViT) models and leverage their features for selection, proposing a multi-resolution processing strategy that yields finer and more stable selection results than prior methods. Furthermore, we enable selection at two levels: texture and subtexture, leveraging a new two-level material selection (DuMaS) dataset which includes dense annotations for over 800,000 synthetic images, both on the texture and subtexture levels. Julia Guerrero-Viu, Michael Fischer 0011, Iliyan Georgiev, Elena Garces 0001, Diego Gutierrez, Belén Masiá, Valentin Deschaintre |
ACM Trans. Graph. | 6 |
| 2024 | AViSal360: Audiovisual Saliency Prediction for 360° VideoabstractSaliency prediction in 360° video plays an important role in modeling visual attention, and can be leveraged for content creation, compression techniques, or quality assessment methods, among others. Visual attention in immersive environments depends not only on visual input, but also on inputs from other sensory modalities, primarily audio. Despite this, only a minority of saliency prediction models have incorporated auditory inputs, and much remains to be explored about what auditory information is relevant and how to integrate it in the prediction. In this work, we propose an audiovisual saliency model for 360° video content, AViSal360. Our model integrates both spatialized and semantic audio information, together with visual inputs. We perform exhaustive comparisons to demonstrate both the actual relevance of auditory information in saliency prediction, and the superior performance of our model when compared to previous approaches. Edurne Bernal-Berdun, Jorge Pina, Mateo Vallejo, Ana Serrano, Belén Masiá |
ISMAR | 6 |
| 2024 | tSPM-Net: A probabilistic spatio-temporal approach for scanpath predictionabstractPredicting the path followed by the viewer’s eyes when observing an image (a scanpath) is a challenging problem, particularly due to the inter- and intra-observer variability and the spatio-temporal dependencies of the visual attention process. Most existing approaches have focused on progressively optimizing the prediction of a gaze point given the previous ones. In this work we propose instead a probabilistic approach, which we call tSPM-Net. We build our method to account for observers’ variability by resorting to Bayesian deep learning and a probabilistic approach. Besides, we optimize our model to jointly consider both spatial and temporal dimensions of scanpaths using a novel spatio-temporal loss function based on a combination of Kullback–Leibler divergence and dynamic time warping. Our tSPM-Net yields results that outperform those of current state-of-the-art approaches, and are closer to the human baseline, suggesting that our model is able to generate scanpaths whose behavior closely resembles those of the real ones. Diego Gutierrez, Belén Masiá |
Comput. Graph. | 3 |
| 2024 | Predicting Perceived Gloss: Do Weak Labels Suffice?abstractAbstract Estimating perceptual attributes of materials directly from images is a challenging task due to their complex, not fully‐understood interactions with external factors, such as geometry and lighting. Supervised deep learning models have recently been shown to outperform traditional approaches, but rely on large datasets of human‐annotated images for accurate perception predictions. Obtaining reliable annotations is a costly endeavor, aggravated by the limited ability of these models to generalise to different aspects of appearance. In this work, we show how a much smaller set of human annotations (“strong labels”) can be effectively augmented with automatically derived “weak labels” in the context of learning a low‐dimensional image‐computable gloss metric. We evaluate three alternative weak labels for predicting human gloss perception from limited annotated data. Incorporating weak labels enhances our gloss prediction beyond the current state of the art. Moreover, it enables a substantial reduction in human annotation costs without sacrificing accuracy, whether working with rendered images or real photographs. Julia Guerrero-Viu, J. Daniel Subias, Ana Serrano, Katherine Storrs, Roland W. Fleming, Belén Masiá, Diego Gutierrez |
Comput. Graph. Forum | 6 |
| 2024 | Navigating the Manifold of Translucent AppearanceabstractAbstract We present a perceptually‐motivated manifold for translucent appearance, designed for intuitive editing of translucent materials by navigating through the manifold. Classic tools for editing translucent appearance, based on the use of sliders to tune a number of parameters, are challenging for non‐expert users: These parameters have a highly non‐linear effect on appearance, and exhibit complex interplay and similarity relations between them. Instead, we pose editing as a navigation task in a low‐dimensional space of appearances, which abstracts the user from the underlying optical parameters. To achieve this, we build a low‐dimensional continuous manifold of translucent appearance that correlates with how humans perceive this type of materials. We first analyze the correlation of different distance metrics in image space with human perception. We select the best‐performing metric to build a low‐dimensional manifold, which can be used to navigate the space of translucent appearance. To evaluate the validity of our proposed manifold within its intended application scenario, we build an editing interface that leverages the manifold, and relies on image navigation plus a fine‐tuning step to edit appearance. We compare our intuitive interface to a traditional, slider‐based one in a user study, demonstrating its effectiveness and superior performance when editing translucent objects. Dario Lanza, Belén Masiá, Adrián Jarabo |
Comput. Graph. Forum | 2 |
| 2024 | SAL3D: a model for saliency prediction in 3D meshesabstractAdvances in virtual and augmented reality have increased the demand for immersive and engaging 3D experiences. To create such experiences, it is crucial to understand visual attention in 3D environments, which is typically modeled by means of saliency maps. While attention in 2D images and traditional media has been widely studied, there is still much to explore in 3D settings. In this work, we propose a deep learning-based model for predicting saliency when viewing 3D objects, which is a first step toward understanding and predicting attention in 3D environments. Previous approaches rely solely on low-level geometric cues or unnatural conditions, however, our model is trained on a dataset of real viewing data that we have manually captured, which indeed reflects actual human viewing behavior. Our approach outperforms existing state-of-the-art methods and closely approximates the ground-truth data. Our results demonstrate the effectiveness of our approach in predicting attention in 3D objects, which can pave the way for creating more immersive and engaging 3D experiences. Andres Fandos, Belén Masiá, Ana Serrano |
Vis. Comput. | 3 |
| 2023 | The Visual Language of FabricsabstractWe introduce text2fabric, a novel dataset that links free-text descriptions to various fabric materials. The dataset comprises 15,000 natural language descriptions associated to 3,000 corresponding images of fabric materials. Traditionally, material descriptions come in the form of tags/keywords, which limits their expressivity, induces pre-existing knowledge of the appropriate vocabulary, and ultimately leads to a chopped description system. Therefore, we study the use of free-text as a more appropriate way to describe material appearance, taking the use case of fabrics as a common item that non-experts may often deal with. Based on the analysis of the dataset, we identify a compact lexicon, set of attributes and key structure that emerge from the descriptions. This allows us to accurately understand how people describe fabrics and draw directions for generalization to other types of materials. We also show that our dataset enables specializing large vision-language models such as CLIP, creating a meaningful latent space for fabric appearance, and significantly improving applications such as fine-grained material retrieval and automatic captioning. Valentin Deschaintre, Julia Guerrero-Viu, Diego Gutierrez, Tamy Boubekeur, Belén Masiá |
ACM Trans. Graph. | 5 |
| 2023 | D-SAV360: A Dataset of Gaze Scanpaths on 360° Ambisonic VideosabstractUnderstanding human visual behavior within virtual reality environments is crucial to fully leverage their potential. While previous research has provided rich visual data from human observers, existing gaze datasets often suffer from the absence of multimodal stimuli. Moreover, no dataset has yet gathered eye gaze trajectories (i.e., scanpaths) for dynamic content with directional ambisonic sound, which is a critical aspect of sound perception by humans. To address this gap, we introduce D-SAV360, a dataset of 4,609 head and eye scanpaths for 360° videos with first-order ambisonics. This dataset enables a more comprehensive study of multimodal interaction on visual behavior in virtual reality environments. We analyze our collected scanpaths from a total of 87 participants viewing 85 different videos and show that various factors such as viewing mode, content type, and gender significantly impact eye movement statistics. We demonstrate the potential of D-SAV360 as a benchmarking resource for state-of-the-art attention prediction models and discuss its possible applications in further research. By providing a comprehensive dataset of eye movement data for dynamic, multimodal virtual environments, our work can facilitate future investigations of visual behavior and attention in virtual reality. Edurne Bernal-Berdun, Sandra Malpica, Pedro J. Perez, Diego Gutierrez, Belén Masiá, Ana Serrano |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Task-Dependent Visual Behavior in Immersive Environments: A Comparative Study of Free Exploration, Memory and Visual SearchabstractVisual behavior depends on both bottom-up mechanisms, where gaze is driven by the visual conspicuity of the stimuli, and top-down mechanisms, guiding attention towards relevant areas based on the task or goal of the viewer. While this is well-known, visual attention models often focus on bottom-up mechanisms. Existing works have analyzed the effect of high-level cognitive tasks like memory or visual search on visual behavior; however, they have often done so with different stimuli, methodology, metrics and participants, which makes drawing conclusions and comparisons between tasks particularly difficult. In this work we present a systematic study of how different cognitive tasks affect visual behavior in a novel within-subjects design scheme. Participants performed free exploration, memory and visual search tasks in three different scenes while their eye and head movements were being recorded. We found significant, consistent differences between tasks in the distributions of fixations, saccades and head movements. Our findings can provide insights for practitioners and content creators designing task-oriented immersive applications. Sandra Malpica, Ana Serrano, Diego Gutierrez, Belén Masiá |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | A Study of Change Blindness in Immersive EnvironmentsabstractHuman performance is poor at detecting certain changes in a scene, a phenomenon known as change blindness. Although the exact reasons of this effect are not yet completely understood, there is a consensus that it is due to our constrained attention and memory capacity: We create our own mental, structured representation of what surrounds us, but such representation is limited and imprecise. Previous efforts investigating this effect have focused on 2D images; however, there are significant differences regarding attention and memory between 2D images and the viewing conditions of daily life. In this work, we present a systematic study of change blindness using immersive 3D environments, which offer more natural viewing conditions closer to our daily visual experience. We devise two experiments; first, we focus on analyzing how different change properties (namely type, distance, complexity, and field of view) may affect change blindness. We then further explore its relation with the capacity of our visual working memory and conduct a second experiment analyzing the influence of the number of changes. Besides gaining a deeper understanding of the change blindness effect, our results may be leveraged in several VR applications such as redirected walking, games, or even studies on saliency or attention prediction. Xin Sun 0014, Diego Gutierrez, Belén Masiá |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | On the Influence of Dynamic Illumination in the Perception of TranslucencyabstractTranslucent materials are ubiquitous in our daily lives, from organic materials such as food, liquids or human skin, to synthetic materials like plastic or rubber. In these materials, light penetrates inside the surface and scatters in the medium before leaving it. While the physical phenomena responsible for translucent appearance are well known, understanding how human observers perceive this type of materials is still an open problem: The appearance of translucent objects is affected by many dimensions beyond the optical properties of the material, including shape and illumination. In this work, we focus on the effect of illumination on the appearance of translucent materials. In particular, we analyze how static and dynamic illumination impact the perception of translucency. Previous studies have shown that changing the illumination conditions results in a constancy failure, specially in media with anisotropic phase functions. We extend this line of work, and analyze whether motion can alleviate such constancy failure. To do that, we run a psychophysical experiment where users need to match the optical density of a reference translucent object under both dynamic and static illumination. Surprisingly, our results suggest that in most cases light motion does not impact the perceived density of the translucent material. Our findings can have implications for material design in predictive rendering and authoring applications. Dario Lanza, Adrián Jarabo, Belén Masiá |
SAP | 3 |
| 2022 | SST-Sal: A spherical spatio-temporal approach for saliency prediction in 360∘ videosabstractVirtual reality (VR) has the potential to change the way people consume content, and has been predicted to become the next big computing paradigm. However, much remains unknown about the grammar and visual language of this new medium, and understanding and predicting how humans behave in virtual environments remains an open problem. In this work, we propose a novel saliency prediction model which exploits the joint potential of spherical convolutions and recurrent neural networks to extract and model the inherent spatio-temporal features from 360° videos. We employ Convolutional Long Short-Term Memory cells (ConvLSTMs) to account for temporal information at the time of feature extraction rather than to post-process spatial features as in previous works. To facilitate spatio-temporal learning, we provide the network with an estimation of the optical flow between 360° frames, since motion is known to be a highly salient feature in dynamic content. Our model is trained with a novel spherical Kullback–Leibler Divergence (KLDiv) loss function specifically tailored for saliency prediction in 360° content. Our approach outperforms previous state-of-the-art works, being able to mimic human visual attention when exploring dynamic 360° videos. Edurne Bernal-Berdun, Diego Gutierrez, Belén Masiá |
Comput. Graph. | 4 |
| 2022 | A Generative Framework for Image-based Editing of Material Appearance using Perceptual AttributesabstractAbstract Single‐image appearance editing is a challenging task, traditionally requiring the estimation of additional scene properties such as geometry or illumination. Moreover, the exact interaction of light, shape and material reflectance that elicits a given perceptual impression is still not well understood. We present an image‐based editing method that allows to modify the material appearance of an object by increasing or decreasing high‐level perceptual attributes, using a single image as input. Our framework relies on a two‐step generative network, where the first step drives the change in appearance and the second produces an image with high‐frequency details. For training, we augment an existing material appearance dataset with perceptual judgements of high‐level attributes, collected through crowd‐sourced experiments, and build upon training strategies that circumvent the cumbersome need for original‐edited image pairs. We demonstrate the editing capabilities of our framework on a variety of inputs, both synthetic and real, using two common perceptual attributes (Glossy and Metallic), and validate the perception of appearance in our edited images through a user study. Johanna Delanoy, Manuel Lagunas, J. Condor, Diego Gutierrez, Belén Masiá |
Comput. Graph. Forum | 5 |
| 2022 | ScanGAN360: A Generative Model of Realistic Scanpaths for 360° ImagesabstractUnderstanding and modeling the dynamics of human gaze behavior in 360° environments is crucial for creating, improving, and developing emerging virtual reality applications. However, recruiting human observers and acquiring enough data to analyze their behavior when exploring virtual environments requires complex hardware and software setups, and can be time-consuming. Being able to generate virtual observers can help overcome this limitation, and thus stands as an open problem in this medium. Particularly, generative adversarial approaches could alleviate this challenge by generating a large number of scanpaths that reproduce human behavior when observing new scenes, essentially mimicking virtual observers. However, existing methods for scanpath generation do not adequately predict realistic scanpaths for 360° images. We present ScanGAN360, a new generative adversarial approach to address this problem. We propose a novel loss function based on dynamic time warping and tailor our network to the specifics of 360° images. The quality of our generated scanpaths outperforms competing approaches by a large margin, and is almost on par with the human baseline. ScanGAN360 allows fast simulation of large numbers of virtual observers, whose behavior mimics real users, enabling a better understanding of gaze behavior, facilitating experimentation, and aiding novel applications in virtual reality and beyond. Ana Serrano, Alexander W. Bergman, Gordon Wetzstein, Belén Masiá |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Preface
Shi-Min Hu 0001, Ying He 0001, Belén Masiá |
J. Comput. Sci. Technol. | 3 |
| 2020 | Crossmodal perception in virtual reality
Sandra Malpica, Ana Serrano, Marcos Allue, Manuel G. Bedia, Belén Masiá |
Multim. Tools Appl. | 5 |
| 2020 | Imperceptible manipulation of lateral camera motion for improved virtual reality applicationsabstractVirtual Reality (VR) systems increase immersion by reproducing users' movements in the real world. However, several works have shown that this real-to-virtual mapping does not need to be precise in order to convey a realistic experience. Being able to alter this mapping has many potential applications, since achieving an accurate real-to-virtual mapping is not always possible due to limitations in the capture or display hardware, or in the physical space available. In this work, we measure detection thresholds for lateral translation gains of virtual camera motion in response to the corresponding head motion under natural viewing, and in the absence of locomotion, so that virtual camera movement can be either compressed or expanded while these manipulations remain undetected. Finally, we propose three applications for our method, addressing three key problems in VR: improving 6-DoF viewing for captured 360° footage, overcoming physical constraints, and reducing simulator sickness. We have further validated our thresholds and evaluated our applications by means of additional user studies confirming that our manipulations remain imperceptible, and showing that (i) compressing virtual camera motion reduces visible artifacts in 6-DoF, hence improving perceived quality, (ii) virtual expansion allows for completion of virtual tasks within a reduced physical space, and (iii) simulator sickness may be alleviated in simple scenarios when our compression method is applied. Ana Serrano, Diego Gutierrez, Karol Myszkowski, Belén Masiá |
ACM Trans. Graph. | 5 |
| 2019 | The Effect of Motion on the Perception of Material AppearanceabstractWe analyze the effect of motion in the perception of material appearance. First, we create a set of stimuli containing 72 realistic materials, rendered with varying degrees of linear motion blur. Then we launch a large-scale study on Mechanical Turk to rate a given set of perceptual attributes, such as brightness, roughness, or the perceived strength of reflections. Our statistical analysis shows that certain attributes undergo a significant change, varying appearance perception under motion. In addition, we further investigate the perception of brightness, for the particular cases of rubber and plastic materials. We create new stimuli, with ten different luminance levels and seven motion degrees. We launch a new user study to retrieve their perceived brightness. From the users’ judgements, we build two-dimensional maps showing how perceived brightness varies as a function of the luminance and motion of the material. Ruiquan Mao, Manuel Lagunas, Belén Masiá, Diego Gutierrez |
SAP | 3 |
| 2019 | Integral Actions Towards Women in Engineering RecognitionabstractThis work presents integral actions towards women in engineering recognition organized according to educational stages they are directed: Stage 1, from early childhood education, primary education and secondary education; Stage 2, during university and Stage 3 after university. At stage 1 actions are devised to increase girl's interest on Science, Technology, Engineering and Mathematics (STEM). Emphasis is put on showing women in engineering as role models, illustrating engineers work and stressing the importance of diversity in working groups. Stage 2 is focused on making male and female students aware of the gender gap in engineering and the importance of diversity for innovation and training female students on known female narrow circumstances. Finally, at stage 3, the objective is to retain and promote women in the engineering profession. The specific actions developed at the three stages are presented. Their impact is discussed in order to accomplish effective actions for achieving gender balance towards excellence in Engineering. Natalia Ayuso-Escuer, Sandra Baldassarri, Raquel Trillo Lado, Rosario Aragues, Belén Masiá, Pilar Molina-Gaudó, Ana Cristina Murillo, Eva Cerezo Bagdasari, María Villarroya-Gaudó |
ETFA | 5 |
| 2019 | A similarity measure for material appearanceabstractWe present a model to measure the similarity in appearance between different materials, which correlates with human similarity judgments. We first create a database of 9,000 rendered images depicting objects with varying materials, shape and illumination. We then gather data on perceived similarity from crowdsourced experiments; our analysis of over 114,840 answers suggests that indeed a shared perception of appearance similarity exists. We feed this data to a deep learning architecture with a novel loss function, which learns a feature space for materials that correlates with such perceived appearance similarity. Our evaluation shows that our model outperforms existing metrics. Last, we demonstrate several applications enabled by our metric, including appearance-based search for material suggestions, database visualization, clustering and summarization, and gamut mapping. Manuel Lagunas, Sandra Malpica, Ana Serrano, Elena Garces 0001, Diego Gutierrez, Belén Masiá |
ACM Trans. Graph. | 6 |
| 2019 | Motion parallax for 360° RGBD videoabstractWe present a method for adding parallax and real-time playback of 360° videos in Virtual Reality headsets. In current video players, the playback does not respond to translational head movement, which reduces the feeling of immersion, and causes motion sickness for some viewers. Given a 360° video and its corresponding depth (provided by current stereo 360° stitching algorithms), a naive image-based rendering approach would use the depth to generate a 3D mesh around the viewer, then translate it appropriately as the viewer moves their head. However, this approach breaks at depth discontinuities, showing visible distortions, whereas cutting the mesh at such discontinuities leads to ragged silhouettes and holes at disocclusions. We address these issues by improving the given initial depth map to yield cleaner, more natural silhouettes. We rely on a three-layer scene representation, made up of a foreground layer and two static background layers, to handle disocclusions by propagating information from multiple frames for the first background layer, and then inpainting for the second one. Our system works with input from many of today's most popular 360° stereo capture devices (e.g., Yi Halo or GoPro Odyssey), and works well even if the original video does not provide depth information. Our user studies confirm that our method provides a more compelling viewing experience than without parallax, increasing immersion while reducing discomfort and nausea. Ana Serrano, Inchul Kim 0001, Stephen DiVerdi, Diego Gutierrez, Aaron Hertzmann, Belén Masiá |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2018 | Saliency in VR: How Do People Explore Virtual Environments?abstractUnderstanding how people explore immersive virtual environments is crucial for many applications, such as designing virtual reality (VR) content, developing new compression algorithms, or learning computational models of saliency or visual attention. Whereas a body of recent work has focused on modeling saliency in desktop viewing conditions, VR is very different from these conditions in that viewing behavior is governed by stereoscopic vision and by the complex interaction of head orientation, gaze, and other kinematic constraints. To further our understanding of viewing behavior and saliency in VR, we capture and analyze gaze and head orientation data of 169 users exploring stereoscopic, static omni-directional panoramas, for a total of 1980 head and gaze trajectories for three different viewing conditions. We provide a thorough analysis of our data, which leads to several important insights, such as the existence of a particular fixation bias, which we then use to adapt existing saliency predictors to immersive VR conditions. In addition, we explore other applications of our data and analysis, including automatic alignment of VR video cuts, panorama thumbnails, panorama video synopsis, and saliency-basedcompression. Vincent Sitzmann, Ana Serrano, Amy Pavel, Maneesh Agrawala, Diego Gutierrez, Belén Masiá, Gordon Wetzstein |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2017 | Convolutional Sparse Coding for Capturing High-Speed Video ContentabstractAbstract Video capture is limited by the trade‐off between spatial and temporal resolution: when capturing videos of high temporal resolution, the spatial resolution decreases due to bandwidth limitations in the capture system. Achieving both high spatial and temporal resolution is only possible with highly specialized and very expensive hardware, and even then the same basic trade‐off remains. The recent introduction of compressive sensing and sparse reconstruction techniques allows for the capture of single‐shot high‐speed video, by coding the temporal information in a single frame, and then reconstructing the full video sequence from this single‐coded image and a trained dictionary of image patches. In this paper, we first analyse this approach, and find insights that help improve the quality of the reconstructed videos. We then introduce a novel technique, based on convolutional sparse coding (CSC), and show how it outperforms the state‐of‐the‐art, patch‐based approach in terms of flexibility and efficiency, due to the convolutional nature of its filter banks. The key idea for CSC high‐speed video acquisition is extending the basic formulation by imposing an additional constraint in the temporal dimension, which enforces sparsity of the first‐order derivatives over time. Ana Serrano, Elena Garces 0001, Belén Masiá, Diego Gutierrez |
Comput. Graph. Forum | 3 |
| 2017 | Attribute-preserving gamut mapping of measured BRDFsabstractAbstract Reproducing the appearance of real‐world materials using current printing technology is problematic. The reduced number of inks available define the printer's limited gamut, creating distortions in the printed appearance that are hard to control. Gamut mapping refers to the process of bringing an out‐of‐gamut material appearance into the printer's gamut, while minimizing such distortions as much as possible. We present a novel two‐step gamut mapping algorithm that allows users to specify which perceptual attribute of the original material they want to preserve (such as brightness, or roughness). In the first step, we work in the low‐dimensional intuitive appearance space recently proposed by Serrano et al. [ SGM*16 ], and adjust achromatic reflectance via an objective function that strives to preserve certain attributes. From such intermediate representation, we then perform an image‐based optimization including color information, to bring the BRDF into gamut. We show, both objectively and through a user study, how our method yields superior results compared to the state of the art, with the additional advantage that the user can specify which visual attributes need to be preserved. Moreover, we show how this approach can also be used for attribute‐preserving material editing. Tiancheng Sun, Ana Serrano, Diego Gutierrez, Belén Masiá |
Comput. Graph. Forum | 4 |
| 2017 | Dynamic range expansion based on image statistics
Belén Masiá, Ana Serrano, Diego Gutierrez |
Multim. Tools Appl. | 1 |
| 2017 | Movie editing and cognitive event segmentation in virtual reality videoabstractTraditional cinematography has relied for over a century on a well-established set of editing rules, called continuity editing, to create a sense of situational continuity. Despite massive changes in visual content across cuts, viewers in general experience no trouble perceiving the discontinuous flow of information as a coherent set of events. However, Virtual Reality (VR) movies are intrinsically different from traditional movies in that the viewer controls the camera orientation at all times. As a consequence, common editing techniques that rely on camera orientations, zooms, etc., cannot be used. In this paper we investigate key relevant questions to understand how well traditional movie editing carries over to VR, such as: Does the perception of continuity hold across edit boundaries? Under which conditions? Does viewers' observational behavior change after the cuts? To do so, we rely on recent cognition studies and the event segmentation theory, which states that our brains segment continuous actions into a series of discrete, meaningful events. We first replicate one of these studies to assess whether the predictions of such theory can be applied to VR. We next gather gaze data from viewers watching VR videos containing different edits with varying parameters, and provide the first systematic analysis of viewers' behavior and the perception of continuity in VR. From this analysis we make a series of relevant findings; for instance, our data suggests that predictions from the cognitive event segmentation theory are useful guides for VR editing; that different types of edits are equally well understood in terms of continuity; and that spatial misalignments between regions of interest at the edit boundaries favor a more exploratory behavior even after viewers have fixated on a new region of interest. In addition, we propose a number of metrics to describe viewers' attentional behavior in VR. We believe the insights derived from our work can be useful as guidelines for VR content creation. Ana Serrano, Vincent Sitzmann, Jaime Ruiz-Borau, Gordon Wetzstein, Diego Gutierrez, Belén Masiá |
ACM Trans. Graph. | 6 |
| 2017 | Recent advances in transient imaging: A computer graphics and vision perspectiveabstractTransient imaging has recently made a huge impact in the computer graphics and computer vision fields. By capturing, reconstructing, or simulating light transport at extreme temporal resolutions, researchers have proposed novel techniques to show movies of light in motion, see around corners, detect objects in highly-scattering media, or infer material properties from a distance, to name a few. The key idea is to leverage the wealth of information in the temporal domain at the pico or nanosecond resolution, information usually lost during the capture-time temporal integration. This paper presents recent advances in this field of transient imaging from a graphics and vision perspective, including capture techniques, analysis, applications and simulation. Adrián Jarabo, Belén Masiá, Julio Marco, Diego Gutierrez |
Vis. Informatics | 2 |
| 2016 | Practical Low-Cost Recovery of Spectral Power DistributionsabstractAbstract Measuring the spectral power distribution of a light source, that is, the emission as a function of wavelength, typically requires the use of spectrophotometers or multi‐spectral cameras. Here, we propose a low‐cost system that enables the recovery of the visible light spectral signature of different types of light sources without requiring highly complex or specialized equipment and using just off‐the‐shelf, widely available components. To do this, a standard Digital Single‐Lens Reflex (DSLR) camera and a diffraction filter are used, sacrificing the spatial dimension for spectral resolution. We present here the image formation model and the calibration process necessary to recover the spectrum, including spectral calibration and amplitude recovery. We also assess the robustness of our method and perform a detailed analysis exploring the parameters influencing its accuracy. Further, we show applications of the system in image processing and rendering. Sara Alvarez, Timo Kunkel, Belén Masiá |
Comput. Graph. Forum | 3 |
| 2016 | Convolutional Sparse Coding for High Dynamic Range ImagingabstractAbstract Current HDR acquisition techniques are based on either (i) fusing multibracketed, low dynamic range (LDR) images, (ii) modifying existing hardware and capturing different exposures simultaneously with multiple sensors, or (iii) reconstructing a single image with spatially‐varying pixel exposures. In this paper, we propose a novel algorithm to recover high‐quality HDRI images from a single, coded exposure. The proposed reconstruction method builds on recently‐introduced ideas of convolutional sparse coding (CSC); this paper demonstrates how to make CSC practical for HDR imaging. We demonstrate that the proposed algorithm achieves higher‐quality reconstructions than alternative methods, we evaluate optical coding schemes, analyze algorithmic parameters, and build a prototype coded HDR camera that demonstrates the utility of convolutional sparse HDRI coding with a custom hardware platform. Ana Serrano, Felix Heide, Diego Gutierrez, Gordon Wetzstein, Belén Masiá |
Comput. Graph. Forum | 5 |
| 2016 | An intuitive control space for material appearanceabstractMany different techniques for measuring material appearance have been proposed in the last few years. These have produced large public datasets, which have been used for accurate, data-driven appearance modeling. However, although these datasets have allowed us to reach an unprecedented level of realism in visual appearance, editing the captured data remains a challenge. In this paper, we present an intuitive control space for predictable editing of captured BRDF data, which allows for artistic creation of plausible novel material appearances, bypassing the difficulty of acquiring novel samples. We first synthesize novel materials, extending the existing MERL dataset up to 400 mathematically valid BRDFs. We then design a large-scale experiment, gathering 56,000 subjective ratings on the high-level perceptual attributes that best describe our extended dataset of materials. Using these ratings, we build and train networks of radial basis functions to act as functionals mapping the perceptual attributes to an underlying PCA-based representation of BRDFs. We show that our functionals are excellent predictors of the perceived attributes of appearance. Our control space enables many applications, including intuitive material editing of a wide range of visual properties, guidance for gamut mapping, analysis of the correlation between perceptual attributes, or novel appearance similarity metrics. Moreover, our methodology can be used to derive functionals applicable to classic analytic BRDF representations. We release our code and dataset publicly, in order to support and encourage further research in this direction. Ana Serrano, Diego Gutierrez, Karol Myszkowski, Hans-Peter Seidel, Belén Masiá |
ACM Trans. Graph. | 5 |
| 2016 | Motion parallax in stereo 3D: model and applicationsabstractBinocular disparity is the main depth cue that makes stereoscopic images appear 3D. However, in many scenarios, the range of depth that can be reproduced by this cue is greatly limited and typically fixed due to constraints imposed by displays. For example, due to the low angular resolution of current automultiscopic screens, they can only reproduce a shallow depth range. In this work, we study the motion parallax cue, which is a relatively strong depth cue, and can be freely reproduced even on a 2D screen without any limits. We exploit the fact that in many practical scenarios, motion parallax provides sufficiently strong depth information that the presence of binocular depth cues can be reduced through aggressive disparity compression. To assess the strength of the effect we conduct psycho-visual experiments that measure the influence of motion parallax on depth perception and relate it to the depth resulting from binocular disparity. Based on the measurements, we propose a joint disparity-parallax computational model that predicts apparent depth resulting from both cues. We demonstrate how this model can be applied in the context of stereo and multiscopic image processing, and propose new disparity manipulation techniques, which first quantify depth obtained from motion parallax, and then adjust binocular disparity information accordingly. This allows us to manipulate the disparity signal according to the strength of motion parallax to improve the overall depth reproduction. This technique is validated in additional experiments. Petr Kellnhofer, Piotr Didyk, Tobias Ritschel 0001, Belén Masiá, Karol Myszkowski, Hans-Peter Seidel |
ACM Trans. Graph. | 4 |
| 2015 | Relativistic Effects for Time-Resolved Light TransportabstractAbstract We present a real‐time framework which allows interactive visualization of relativistic effects for time‐resolved light transport. We leverage data from two different sources: real‐world data acquired with an effective exposure time of less than 2 picoseconds, using an ultra‐fast imaging technique termed femto‐photography, and a transient renderer based on ray‐tracing. We explore the effects of time dilation, light aberration, frequency shift and radiance accumulation by modifying existing models of these relativistic effects to take into account the time‐resolved nature of light propagation. Unlike previous works, we do not impose limiting constraints in the visualization, allowing the virtual camera to explore freely a reconstructed 3D scene depicting dynamic illumination. Moreover, we consider not only linear motion, but also acceleration and rotation of the camera. We further introduce, for the first time, a pinhole camera model into our relativistic rendering framework, and account for subsequent changes in focal length and field of view as the camera moves through the scene. Adrián Jarabo, Belén Masiá, Andreas Velten, Christopher Barsi, Ramesh Raskar, Diego Gutierrez |
Comput. Graph. Forum | 2 |
| 2014 | Decomposing Global Light Transport Using Time of Flight Imaging
Di Wu 0006, Andreas Velten, Matthew O'Toole, Belén Masiá, Amit K. Agrawal, Qionghai Dai, Ramesh Raskar |
Int. J. Comput. Vis. | 4 |
| 2014 | How do people edit light fields?abstractWe present a thorough study to evaluate different light field editing interfaces, tools and workflows from a user perspective. This is of special relevance given the multidimensional nature of light fields, which may make common image editing tasks become complex in light field space. We additionally investigate the potential benefits of using depth information when editing, and the limitations imposed by imperfect depth reconstruction using current techniques. We perform two different experiments, collecting both objective and subjective data from a varied number of editing tasks of increasing complexity based on local point-and-click tools. In the first experiment, we rely on perfect depth from synthetic light fields, and focus on simple edits. This allows us to gain basic insight on light field editing, and to design a more advanced editing interface. This is then used in the second experiment, employing real light fields with imperfect reconstructed depth, and covering more advanced editing tasks. Our study shows that users can edit light fields with our tested interface and tools, even in the presence of imperfect depth. They follow different workflows depending on the task at hand, mostly relying on a combination of different depth cues. Last, we confirm our findings by asking a set of artists to freely edit both real and synthetic light fields. Adrián Jarabo, Belén Masiá, Adrien Bousseau, Fabio Pellacini, Diego Gutierrez |
ACM Trans. Graph. | 2 |
| 2013 | Display adaptive 3D content remapping
Belén Masiá, Gordon Wetzstein, Carlos Aliaga, Ramesh Raskar, Diego Gutierrez |
Comput. Graph. | 1 |
| 2013 | A survey on computational displays: Pushing the boundaries of optics, computation, and perception
Belén Masiá, Gordon Wetzstein, Piotr Didyk, Diego Gutierrez |
Comput. Graph. | 1 |
| 2013 | A metric of visual comfort for stereoscopic motionabstractWe propose a novel metric of visual comfort for stereoscopic motion, based on a series of systematic perceptual experiments. We take into account disparity, motion in depth, motion on the screen plane, and the spatial frequency of luminance contrast. We further derive a comfort metric to predict the comfort of short stereoscopic videos. We validate it on both controlled scenes and real videos available on the internet, and show how all the factors we take into account, as well as their interactions, affect viewing comfort. Last, we propose various applications that can benefit from our comfort measurements and metric. Song-Pei Du, Belén Masiá, Shi-Min Hu 0001, Diego Gutierrez |
ACM Trans. Graph. | 2 |
| 2013 | Femto-photography: capturing and visualizing the propagation of lightabstractWe present femto-photography , a novel imaging technique to capture and visualize the propagation of light. With an effective exposure time of 1.85 picoseconds (ps) per frame, we reconstruct movies of ultrafast events at an equivalent resolution of about one half trillion frames per second. Because cameras with this shutter speed do not exist, we re-purpose modern imaging hardware to record an ensemble average of repeatable events that are synchronized to a streak sensor, in which the time of arrival of light from the scene is coded in one of the sensor's spatial dimensions. We introduce reconstruction methods that allow us to visualize the propagation of femtosecond light pulses through macroscopic scenes; at such fast resolution, we must consider the notion of time-unwarping between the camera's and the world's space-time coordinate systems to take into account effects associated with the finite speed of light. We apply our femto-photography technique to visualizations of very different scenes, which allow us to observe the rich dynamics of time-resolved light transport effects, including scattering, specular reflections, diffuse interreflections, diffraction, caustics, and subsurface scattering. Our work has potential applications in artistic, educational, and scientific visualizations; industrial imaging to analyze material properties; and medical imaging to reconstruct subsurface elements. In addition, our time-resolved technique may motivate new forms of computational photography. Andreas Velten, Di Wu 0006, Adrián Jarabo, Belén Masiá, Christopher Barsi, Chinmaya Joshi, Everett Lawson, Moungi Bawendi, Diego Gutierrez, Ramesh Raskar |
ACM Trans. Graph. | 4 |
| 2012 | Perceptually Optimized Coded Apertures for Defocus DeblurringabstractAbstract The field of computational photography, and in particular the design and implementation of coded apertures, has yielded impressive results in the last years. In this paper we introduce perceptually optimized coded apertures for defocused deblurring. We obtain near‐optimal apertures by means of optimization, with a novel evaluation function that includes two existing image quality perceptual metrics. These metrics favour results where errors in the final deblurred images will not be perceived by a human observer. Our work improves the results obtained with a similar approach that only takes into account the L2 metric in the evaluation function. Belén Masiá, Lara Presa, Adrian Corrales, Diego Gutierrez |
Comput. Graph. Forum | 1 |
| 2009 | Evaluation of reverse tone mapping through varying exposure conditionsabstractMost existing image content has low dynamic range (LDR), which necessitates effective methods to display such legacy content on high dynamic range (HDR) devices. Reverse tone mapping operators (rTMOs) aim to take LDR content as input and adjust the contrast intelligently to yield output that recreates the HDR experience. In this paper we show that current rTMO approaches fall short when the input image is not exposed properly. More specifically, we report a series of perceptual experiments using a Brightside HDR display and show that, while existing rTMOs perform well for under-exposed input data, the perceived quality degrades substantially with over-exposure, to the extent that in some cases subjects prefer the LDR originals to images that have been treated with rTMOs. We show that, in these cases, a simple rTMO based on gamma expansion avoids the errors introduced by other methods, and propose a method to automatically set a suitable gamma value for each image, based on the image key and empirical data. We validate the results both by means of perceptual experiments and using a recent image quality metric, and show that this approach enhances visible details without causing artifacts in incorrectly-exposed regions. Additionally, we perform another set of experiments which suggest that spatial artifacts introduced by rTMOs are more disturbing than inaccuracies in the expanded intensities. Together, these findings suggest that when the quality of the input data is unknown, reverse tone mapping should be handled with simple, non-aggressive methods to achieve the desired effect. Belén Masiá, Sandra Agustin, Roland W. Fleming, Olga Sorkine-Hornung, Diego Gutierrez |
ACM Trans. Graph. | 1 |