Susana Castillo 0001

dblp:69/10121 · also Susana Castillo Alejandre · DBLP profile ↗
← Back
30ranked-venue papers
3as first author
18since 2021 · last 2025
0000-0003-1245-4758ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 INPC: Implicit Neural Point Clouds for Radiance Field Rendering
abstract
We introduce a new approach for reconstruction and novel view synthesis of unbounded real-world scenes. In contrast to previous methods using either volumetric fields, grid-based models, or discrete point cloud proxies, we pro-pose a hybrid scene representation, which implicitly encodes the geometry in a continuous octree-based probability field and view-dependent appearance in a multi-resolution hash grid. This allows for extraction of arbitrary explicit point clouds, which can be rendered using rasterization. In doing so, we combine the benefits of both worlds and retain favorable behavior during optimization: Our novel implicit point cloud representation and differentiable bilinear rasterizer enable fast rendering while preserving the fine geometric detail captured by volumetric neural fields. Furthermore, this representation does not depend on priors like structure-from-motion point clouds. Our method achieves state-of-the-art image quality on common benchmarks. Furthermore, we achieve fast inference at interactive frame rates, and can convert our trained model into a large, explicit point cloud to further enhance performance.
Florian Hahlbohm, Linus Franke, Moritz Kappel, Susana Castillo 0001, Martin Eisemann, Marc Stamminger, Marcus A. Magnor
3DV4
2025 Real-Time Rendering Framework for Holography
abstract
Abstract With the advent of holographic near‐eye displays, the need for rendering algorithms that output holograms instead of color images emerged. These holograms usually encode phase maps that alter the phase of coherent light sources such that images result from diffraction effects. While common approaches rely on translating the output of traditional rendering systems to holograms in a post processing step, we instead developed a rendering system that can directly output a phase map to a Spatial Light Modulator (SLM). Our hardware‐ray‐traced sparse point distribution, and depth mapping enable rapid hologram generation, allowing for high‐quality time‐multiplexed holography for real‐time content. Additionally, our system is compatible with foveated rendering which enables further performance optimizations.
Sascha Fricke, Susana Castillo 0001, Martin Eisemann, Marcus A. Magnor
Comput. Graph. Forum2
2025 Efficient Perspective-Correct 3D Gaussian Splatting Using Hybrid Transparency
abstract
Abstract 3D Gaussian Splats (3DGS) have proven a versatile rendering primitive, both for inverse rendering as well as real‐time exploration of scenes. In these applications, coherence across camera frames and multiple views is crucial, be it for robust convergence of a scene reconstruction or for artifact‐free fly‐throughs. Recent work started mitigating artifacts that break multi‐view coherence, including popping artifacts due to inconsistent transparency sorting and perspective‐correct outlines of (2D) splats. At the same time, real‐time requirements forced such implementations to accept compromises in how transparency of large assemblies of 3D Gaussians is resolved, in turn breaking coherence in other ways. In our work, we aim at achieving maximum coherence, by rendering fully perspective‐correct 3D Gaussians while using a high‐quality approximation of accurate blending, hybrid transparency, on a per‐pixel level, in order to retain real‐time frame rates. Our fast and perspectively accurate approach for evaluation of 3D Gaussians does not require matrix inversions, thereby ensuring numerical stability and eliminating the need for special handling of degenerate splats, and the hybrid transparency formulation for blending maintains similar quality as fully resolved per‐pixel transparencies at a fraction of the rendering costs. We further show that each of these two components can be independently integrated into Gaussian splatting systems. In combination, they achieve up to 2× higher frame rates, 2× faster optimization, and equal or better image quality with fewer rendering artifacts compared to traditional 3DGS on common benchmarks.
Florian Hahlbohm, Fabian Friederichs, Tim Weyrich, Linus Franke, Moritz Kappel, Susana Castillo 0001, Marc Stamminger, Martin Eisemann, Marcus A. Magnor
Comput. Graph. Forum6
2025 D-NPC: Dynamic Neural Point Clouds for Non-Rigid View Synthesis from Monocular Video
abstract
Abstract Dynamic reconstruction and spatiotemporal novel‐view synthesis of non‐rigidly deforming scenes recently gained increased attention. While existing work achieves impressive quality and performance on multi‐view or teleporting camera setups, most methods fail to efficiently and faithfully recover motion and appearance from casual monocular captures. This paper contributes to the field by introducing a new method for dynamic novel view synthesis from monocular video, such as casual smartphone captures. Our approach represents the scene as a dynamic neural point cloud, an implicit time‐conditioned point distribution that encodes local geometry and appearance in separate hash‐encoded neural feature grids for static and dynamic regions. By sampling a discrete point cloud from our model, we can efficiently render high‐quality novel views using a fast differentiable rasterizer and neural rendering network. Similar to recent work, we leverage advances in neural scene analysis by incorporating data‐driven priors like monocular depth estimation and object segmentation to resolve motion and depth ambiguities originating from the monocular captures. In addition to guiding the optimization process, we show that these priors can be exploited to explicitly initialize our scene representation to drastically improve optimization speed and final image quality. As evidenced by our experimental evaluation, our dynamic point cloud model not only enables fast optimization and real‐time frame rates for interactive applications, but also achieves competitive image quality on monocular benchmark sequences. Our code and data are available online https://moritzkappel.github.io/projects/dnpc/ .
Moritz Kappel, Florian Hahlbohm, Timon Scholz, Susana Castillo 0001, Christian Theobalt, Martin Eisemann, Vladislav Golyanik, Marcus A. Magnor
Comput. Graph. Forum4
2025 Virtual and Traditional Memory Palaces in Recall with ADHD
abstract
Several studies have examined the potential of immersive technologies, such as Virtual Reality (VR), to enhance the effectiveness of memorization. However, existing research has not specifically focused on individuals with Attention Deficit Hyperactivity Disorder (ADHD), and only a limited number of studies have examined the effectiveness of Memory Palace (MP) as a memorization aid for this population, who often experience working memory impairments. To address this gap, we recruited participants with an official diagnosis of ADHD and conducted an experiment in which we investigated the impact of both the Traditional MP technique and a VR version to assess their impact on a memorization task. Our findings indicate that the effectiveness of a VR-based MP might be influenced by prior experience with VR systems, the level of familiarity with the virtual environment, and the MP technique. Nevertheless, our results demonstrate that the MP technique has the potential to enhance recall performance in individuals diagnosed with ADHD.
Anika Jewst, Susana Castillo 0001, Marcus A. Magnor, Martin Eisemann, Dagmar Meyer
ACM Trans. Appl. Percept.2
2025 Fast Non-Rigid Radiance Fields From Monocularized Data
abstract
The reconstruction and novel view synthesis of dynamic scenes recently gained increased attention. As reconstruction from large-scale multi-view data involves immense memory and computational requirements, recent benchmark datasets provide collections of single monocular views per timestamp sampled from multiple (virtual) cameras. We refer to this form of inputs asmonocularizeddata. Existing work shows impressive results for synthetic setups and forward-facing real-world data, but is often limited in the training speed and angular range for generating novel views. This paper addresses these limitations and proposes a new method for full$360^{\circ}$inward-facing novel view synthesis of non-rigidly deforming scenes. At the core of our method are: 1) An efficient deformation module that decouples the processing of spatial and temporal information for accelerated training and inference; and 2) A static module representing the canonical scene as a fast hash-encoded neural radiance field. In addition to existing synthetic monocularized data, we systematically analyze the performance on real-world inward-facing scenes using a newly recorded challenging dataset sampled from a synchronized large-scale multi-view rig. In both cases, our method is significantly faster than previous methods, converging in less than 7 minutes and achieving real-time framerates at 1K resolution, while obtaining a higher visual accuracy for generated novel views. Our code and dataset are available online:https://github.com/MoritzKappel/MoNeRF.
Moritz Kappel, Vladislav Golyanik, Susana Castillo 0001, Christian Theobalt, Marcus A. Magnor
IEEE Trans. Vis. Comput. Graph.3
2024 Measuring Velocity Perception Regarding Stimulus Eccentricity
abstract
A major factor resulting in cybersickness is the feeling of self-motion experienced when viewing a moving scene in Virtual Reality (VR). Current research indicates that this effect is largely created by motion in the periphery. To discover why this is the case, we investigate the influence of temporal frequency and eccentricity of a stimulus on the magnitude of perceived velocity in the periphery. Based on the perception of two-dimensional stimuli on a wide field-of-view display, we build a model to predict the scaling factor by which the perceived velocity of visual patterns deviates from the physical velocity. Further, our exploratory findings indicate no impact of gaze type on the results, suggesting our model works for both fixation and smooth pursuit scenarios. In an additional pilot study in Virtual Reality (VR), we test the accuracy of the model to predict unnoticeable object motion adaptation in 3D virtual worlds and find positive indications for a similar effect.
Timon Scholz, Colin Groth, Susana Castillo 0001, Martin Eisemann, Marcus A. Magnor
SAP3
2023 Instant Hand Redirection in Virtual Reality Through Electrical Muscle Stimulation-Triggered Eye Blinks
abstract
In this paper we investigate the use of electrical muscle stimulation (EMS) to trigger eye blinks for instant hand redirection in virtual reality (VR). With the rapid development of VR technology and increasing user expectations for realistic experiences, maintaining a seamless match between real and virtual objects becomes crucial for immersive interactions. However, hand movements are fast and sometimes unpredictable, increasing the need for instantaneous redirection. We introduce EMS to the field of hand redirection in VR through precise stimulation of the eyelid muscles. By exploiting the phenomenon of change blindness through natural eye blinks, our novel stimulation model achieves instantaneous, imperceptible hand redirection without the need for eye tracking. We first empirically validate the efficiency of our EMS model in eliciting full eye closure. In a second experiment, we demonstrate the feasibility of using such a technique for seamless instantaneous displacement in VR and its particular impact for hand redirection. Among other factors, our analysis also delves into the under-explored domain of gender influence on hand redirection techniques, revealing significant gender-based performance disparities.
Colin Groth, Timon Scholz, Susana Castillo 0001, Jan-Philipp Tauscher, Marcus A. Magnor
VRST3
2023 Immersive Free-Viewpoint Panorama Rendering from Omnidirectional Stereo Video
abstract
Abstract In this paper, we tackle the challenging problem of rendering real‐world 360° panorama videos that support full 6 degrees‐of‐freedom (DoF) head motion from a prerecorded omnidirectional stereo (ODS) video. In contrast to recent approaches that create novel views for individual panorama frames, we introduce a video‐specific temporally‐consistent multi‐sphere image (MSI) scene representation. Given a conventional ODS video, we first extract information by estimating framewise descriptive feature maps. Then, we optimize the global MSI model using theory from recent research on neural radiance fields. Instead of a continuous scene function, this multi‐sphere image (MSI) representation depicts colour and density information only for a discrete set of concentric spheres. To further improve the temporal consistency of our results, we apply an ancillary refinement step which optimizes the temporal coherency between successive video frames. Direct comparisons to recent baseline approaches show that our global MSI optimization yields superior performance in terms of visual quality. Our code and data will be made publicly available.
Moritz Mühlhausen, Moritz Kappel, Marc Kassubeck, Leslie Wöhler, Steve Grogorick, Susana Castillo 0001, Martin Eisemann, Marcus A. Magnor
Comput. Graph. Forum6
2023 Wavelet-Based Fast Decoding of 360° Videos
abstract
In this paper, we propose a wavelet-based video codec specifically designed for VR displays that enables real-time playback of high-resolution 360° videos. Our codec exploits the fact that only a fraction of the full 360° video frame is visible on the display at any time. To load and decode the video viewport-dependently in real time, we make use of the wavelet transform for intra- as well as inter-frame coding. Thereby, the relevant content is directly streamed from the drive, without the need to hold the entire frames in memory. With an average of 193 frames per second at 8192 × 8192 -pixel full-frame resolution, the conducted evaluation demonstrates that our codec's decoding performance is up to 272% higher than that of the state-of-the-art video codecs H.265 and AV1 for typical VR displays. By means of a perceptual study, we further illustrate the necessity of high frame rates for a better VR experience. Finally, we demonstrate how our wavelet-based codec can also directly be used in conjunction with foveation for further performance increase.
Colin Groth, Sascha Fricke, Susana Castillo 0001, Marcus A. Magnor
IEEE Trans. Vis. Comput. Graph.3
2022 Personality analysis of face swaps: can they be used as avatars?
abstract
In this paper, we investigate the perceived personalities of face swaps and how they relate to the personalities of the real people used to create the synthetic individuals' appearance and movements. Given that face swaps have become nearly indistinguishable from real humans, they offer a promising direction for the fast creation of realistic avatars. To investigate the usability of face swaps as avatars, we perform an experiment assessing their personality on the Five-Factor Model, their eeriness and appeal, as well as effects due to familiarity with the original individuals. Our results indicate that face swaps are perceived similarly to real humans and are affected by familiarity. Furthermore, we find a stronger influence of the body and movements on the perceived synthetic personality, especially for their extroversion and conscientiousness.
Leslie Wöhler, Susana Castillo 0001, Marcus A. Magnor
IVA2
2022 Automatic Generation of Customized Areas of Interest and Evaluation of Observers' Gaze in Portrait Videos
abstract
We present a novel framework for the evaluation of eye tracking data in portrait videos including the automatic generation of customized areas of interest (AOIs) based on facial landmarks. In contrast to previous work, our framework allows the user to flexibly create AOIs by grouping the detected landmarks. Moreover, their shape and size can be modified to better fit both the research question and the precision of the eye tracker. The framework can be used as an integrated solution to not only generate AOIs but also to evaluate viewing behavior like the overall fixation times, the similarity of scanpaths, and the number of saccades between AOIs. Other functionalities include the visualization of gaze paths and the creation of heatmaps. We demonstrate the benefits of our framework and user-defined AOI layouts via an exemplary application, i.e., the investigation of face swapping artifacts.
Leslie Wöhler, Moritz von Estorff, Susana Castillo 0001, Marcus A. Magnor
Proc. ACM Hum. Comput. Interact.3
2022 Omnidirectional Galvanic Vestibular Stimulation in Virtual Reality
abstract
In this paper we propose omnidirectional galvanic vestibular stimulation (GVS) to mitigate cybersickness in virtual reality applications. One of the most accepted theories indicates that Cybersickness is caused by the visually induced impression of ego motion while physically remaining at rest. As a result of this sensory mismatch, people associate negative symptoms with VR and sometimes avoid the technology altogether. To reconcile the two contradicting sensory perceptions, we investigate GVS to stimulate the vestibular canals behind our ears with low-current electrical signals that are specifically attuned to the visually displayed camera motion. We describe how to calibrate and generate the appropriate GVS signals in real-time for pre-recorded omnidirectional videos exhibiting ego-motion in all three spatial directions. For validation, we conduct an experiment presenting real-world 360° videos shot from a moving first-person perspective in a VR head-mounted display. Our findings indicate that GVS is able to significantly reduce discomfort for cybersickness-susceptible VR users, creating a deeper and more enjoyable immersive experience for many people.
Colin Groth, Jan-Philipp Tauscher, Nikkel Heesen, Max Hattenbach, Susana Castillo 0001, Marcus A. Magnor
IEEE Trans. Vis. Comput. Graph.5
2021 Towards Understanding Perceptual Differences between Genuine and Face-Swapped Videos
abstract
In this paper, we report on perceptual experiments indicating that there are distinct and quantitatively measurable differences in the way we visually perceive genuine versus face-swapped videos.
Leslie Wöhler, Martin Zembaty, Susana Castillo 0001, Marcus A. Magnor
CHI3
2021 High-Fidelity Neural Human Motion Transfer From Monocular Video
abstract
Video-based human motion transfer creates video animations of humans following a source motion. Current methods show remarkable results for tightly-clad subjects. However, the lack of temporally consistent handling of plausible clothing dynamics, including fine and high-frequency details, significantly limits the attainable visual quality. We address these limitations for the first time in the literature and present a new framework which performs high-fidelity and temporally-consistent human motion transfer with natural pose-dependent non-rigid deformations, for several types of loose garments. In contrast to the previous techniques, we perform image generation in three subsequent stages: synthesizing human shape, structure, and appearance. Given a monocular RGB video of an actor, we train a stack of recurrent deep neural networks that generate these intermediate representations from 2D poses and their temporal derivatives. Splitting the difficult motion transfer problem into subtasks that are aware of the temporal motion context helps us to synthesize results with plausible dynamics and pose-dependent detail. It also allows artistic control of results by manipulation of individual framework stages. In the experimental results, we significantly outperform the state-of-the-art in terms of video realism. The source code is available at https://graphics.tu-bs.de/publications/kappel2020high-fidelity.
Moritz Kappel, Vladislav Golyanik, Mohamed A. Elgharib, Jann-Ole Henningson, Hans-Peter Seidel, Susana Castillo 0001, Christian Theobalt, Marcus A. Magnor
CVPR6
2021 EEG-Based Analysis of the Impact of Familiarity in the Perception of Deepfake Videos
abstract
We investigate the brain’s subliminal response to fake portrait videos using electroencephalography, with a special emphasis on the viewer’s familiarity with the depicted individuals.Deepfake videos are increasingly becoming popular but, while they are entertaining, they can also pose a threat to society. These face-swapped videos merge physiognomy and behaviour of two different individuals, both strong cues used for recognizing a person. We show that this mismatch elicits different brain responses depending on the viewer’s familiarity with the merged individuals.Using EEG, we classify perceptual differences of familiar and unfamiliar people versus their face-swapped counterparts. Our results show that it is possible to discriminate fake videos from genuine ones when at least one face-swapped actor is known to the observer. Furthermore, we indicate a correlation of classification accuracy with level of personal engagement between participant and actor, as well as with the participant’s familiarity with the used dataset.
Jan-Philipp Tauscher, Susana Castillo 0001, Sebastian Bosse, Marcus A. Magnor
ICIP2
2021 Reducing Stair Artifacts in CT Reconstruction
abstract
Computed Tomography is increasingly employed for non-destructive evaluation, with the aim of reconstructing a surface mesh of a scanned object from radiographic projections. State-of-the-art algorithms first reconstruct a voxel grid and then extract a surface mesh using existing meshing algorithms, often leading to stair-like aliasing artifacts along the grid axes, due to the grid’s orientation-dependent resolution. We circumvent such artifacts in filtered backprojection reconstructions by optimizing the mesh’s vertex positions using information taken directly from the projections, rather than from a voxel grid. We show that our approach reduces stair artifacts both visibly and measurably, at relatively little additional computational cost. Our method can be tied into existing mesh extraction algorithms and removes stair artifacts almost entirely.
Markus Wedekind, Eric Oertel, Susana Castillo 0001, Marcus A. Magnor
ICIP3
2021 Shape from Caustics: Reconstruction of 3D-Printed Glass from Simulated Caustic Images
abstract
We present an efficient and effective computational frame-work for the inverse rendering problem of reconstructing the 3D shape of a piece of glass from its caustic image. Our approach is motivated by the needs of 3D glass printing, a nascent additive manufacturing technique that promises to revolutionize the production of optics elements, from lightweight mirrors to waveguides and lenses. One important problem is the reliable control of the manufacturing process by inferring the printed 3D glass shape from its caustic image. Towards this goal, we propose a novel general-purpose reconstruction algorithm based on differentiable light propagation simulation followed by a regularization scheme that takes the deposited glass volume into account. This enables incorporating arbitrary measurements of caustics into an efficient reconstruction framework. We demonstrate the effectiveness of our method and establish the influence of our hyperparameters using several sample shapes and parameter configurations.
Marc Kassubeck, Florian Bürgel, Susana Castillo 0001, Sebastian Stiller, Marcus A. Magnor
WACV3
2020 Temporal Consistent Motion Parallax for Omnidirectional Stereo Panorama Video
abstract
We present a new pipeline to enable head-motion parallax in omnidirectional stereo (ODS) panorama video rendering using a neural depth decoder. While recent ODS panorama cameras record short-baseline horizontal stereo parallax to offer the impression of binocular depth, they do not support the necessary translational degrees-of-freedom (DoF) to also provide for head-motion parallax in virtual reality (VR) applications.
Moritz Mühlhausen, Moritz Kappel, Marc Kassubeck, Paul Maximilian Bittner, Susana Castillo 0001, Marcus A. Magnor
VRST5
2020 Stereo Inverse Brightness Modulation for Guidance in Dynamic Panorama Videos in Virtual Reality
abstract
Abstract The peak of virtual reality offers new exciting possibilities for the creation of media content but also poses new challenges. Some areas of interest might be overlooked because the visual content fills up a large portion of viewers' visual field. Moreover, this content is available in 360° around the viewer, yielding locations completely out of sight, making, for example, recall or storytelling in cinematic Virtual Reality (VR) quite difficult. In this paper, we present an evaluation of Stereo Inverse Brightness Modulation for effective and subtle guidance of participants' attention while navigating dynamic virtual environments. The used technique exploits the binocular rivalry effect from human stereo vision and was previously shown to be effective in static environments. Moreover, we propose an extension of the method for successful guidance towards target locations outside the initial visual field. We conduct three perceptual studies, using 13 distinct panorama videos and two VR systems (a VR head mounted display and a fully immersive dome projection system), to investigate (1) general applicability to dynamic environments, (2) stimulus parameter and VR system influence, and (3) effectiveness of the proposed extension for out‐of‐sight targets. Our results prove the applicability of the method to dynamic environments while maintaining its unobtrusive appearance.
Steve Grogorick, Jan-Philipp Tauscher, Nikkel Heesen, Susana Castillo 0001, Marcus A. Magnor
Comput. Graph. Forum4
2020 Exploring neural and peripheral physiological correlates of simulator sickness
abstract
Abstract This article investigates neural and physiological correlates of simulator sickness (SS) through a controlled experiment conducted within a fully immersive dome projection system. Our goal is to establish a reliable, objective, and in situ measurable predictive indicator of SS. SS is a problem common to all types of visual simulators consisting of motion sickness‐like symptoms that may be experienced while and after being exposed to a dynamic, immersive visualization. It leads to ethical concerns and impaired validity of simulator‐based research. Due to the popularity of virtual reality devices, the number of people exposed to this problem is increasing and, therefore, it is crucial to find reliable predictors of this condition before any symptoms appear. Despite its relevance and the several theories about its origins, SS cannot yet be quantitatively modeled and predicted. Our results indicate that, while neural correlates did not materialize, physiological measures may be a solid early indicator of oncoming SS.
Jan-Philipp Tauscher, Alexandra Witt, Sebastian Bosse, Fabian Wolf Schottky, Steve Grogorick, Susana Castillo 0001, Marcus A. Magnor
Comput. Animat. Virtual Worlds6
2019 Fitting the Style: The Semantic Space for Emotions on Stylized Faces
abstract
It has been widely proven among several research disciplines that the information conveyed through the visual channel can be critical for communication. Altering either the motion or the appearance of our interlocutor can change the opinion we form about them or lead us to drastically different conclusions on what is being communicated. In this paper, we explore the influence of altering the visual appearance of a virtual 3D face model through several different stylization techniques on the perception of dynamic conversational facial expressions. We propose the use of a semantic-differential task to recover the underlying semantic/cognitive space for the expressions and their corresponding shifts due to the application of the different styles. Given the appropriate amount of stimuli, we consider this technique well fitted to provide a mapping between meaning of emotions and spectrum of rendering parameters. Our results show that the different techniques are capable to alter the perception of the expressions making them, for example, more intense or ephemeral in a rather consistent manner.
Philipp Hahn, Susana Castillo 0001, Douglas W. Cunningham
CASA2
2018 The semantic space for emotional speech and the influence of different methods for prosody isolation on its perception
abstract
Normally, when people talk to other people, they communicate not only using specific words, but also with intentional changes in their voice melody, facial expressions, and gestures. Not only is human communication inherently multimodal, it is also multi-layered. That is, it conveys more than simple semantic information, but also passes on a wide variety of social, emotional, and functional (e.g., conversation control) information. Previous work has examined the perception of socio-emotional information conveyed by words and facial expressions. Here, we build on that work and examine the perception of socio-emotional information based solely on prosody (e.g., speech melody, rate, tempo, intensity). To examine the perception of affective prosody, it is necessary to remove all semantics from the speech signal - without changing the prosody! In this paper, we compare several different state-of-the-art methods for removing semantics. We started by recording an audio database containing a German sentence spoken by 11 people in 62 different emotional states. We then removed or masked the semantics using three different techniques. We also recorded the same 62 states for a pseudo-language phrase. Each of these five sets of stimuli were subjected to a semantic differential rating task to derive and compare the semantic spaces for emotions. The results show that each of the methods successfully removed the semantic component, but also changed the perception of the emotional content. Interestingly, the pseudo-word stimuli diverged most from the normal sentences. Furthermore, although each of the filters affected the perception of the sentence in some manner, they did so in different ways.
Martin Schorradt, Susana Castillo 0001, Douglas W. Cunningham
SAP2
2018 Personality Analysis of Embodied Conversational Agents
abstract
People tend to personify machines. Giving machines the ability to actually produce social information can help improve human-machine interactions. Embodied Conversational Agents (ECAs) are virtual software agents that can process and produce speech, facial expressions, gestures and eye gaze, enabling natural, multimodal, human-machine communication. On the one hand, the field of personality psychology provides insights into how we could describe and measure the virtual personality of ECAs. On the other hand, ECAs provide a method to systematically examine how different factors affect the perception of personality. This paper shows that standardized, validated personality questionnaires can be used to evaluate ECAs psychologically, and that state of the art ECAs can manipulate their perceived personality through appearance and behavior.
Susana Castillo 0001, Philipp Hahn, Katharina Legde, Douglas W. Cunningham
IVA1
2018 Look Me in the Lines: The Impact of Stylization on the Recognition of Expressions and Perceived Personality
abstract
We are increasingly approaching the point where computer-based technology is truly ambient and omnipresent. People tend to personify their technical servants, including giving them human names as well as attributing personality traits and intentions to them. The more those devices advance from simple tools to intelligent assistants the more seriously we need to take this personification. That is, if the computers perform human-like tasks in collaboration with humans, and humans already tend to treat computers as human-like, it is only reasonable to give those devices a human-like appearance and conversational abilities. Therefore, one approach to design advanced human-machine interfaces relies heavily on the so-called Embodied Conversational Agents (ECAs). An ECA is a virtual software agent that can process and produce speech, facial expressions, gestures and eye-gaze and, as a result, enables natural, multimodal, human-machine communication. Decades of research in psychology and related fields have shown that the visual channel is especially important in human-human-communication, with subtle changes in both appearance and motion altering how a conversational partner is perceived. In this work, we examine the effectiveness of modifying a virtual character's visual appearance using well-known stylization techniques in order to alter its perceived personality. We also explore the effect of these techniques on the recognizability, intensity and sincerity of the character's displayed emotions.
Philipp Hahn, Susana Castillo 0001, Douglas W. Cunningham
IVA2
2018 The semantic space for motion-captured facial expressions
abstract
Abstract We cannot not communicate! During our daily lives, we convey information verbally and nonverbally. Most of the affective meaning of a message is transferred with the help of facial expressions, and thereby, when trying to establish a realistic human‐like virtual character, we should pay close attention to the animation. Motion capture is one of the most common techniques, but due to the wide range of expressions humans use, the recording time and data needed are vast. To address this problem, we propose the use of semantic spaces as they help in characterizing and positioning expressions by finding a correlation between them. In this paper, we extend prior research by providing the semantic spaces underlying real videos and motion capture data for a total of 62 conversational expressions. Our results highly correlate with previous work, showing that our new expressions were correctly recognized. Moreover, our results can be used in future work to directly project potential new recordings of these 62 expressions on the found spaces.
Susana Castillo 0001, Katharina Legde, Douglas W. Cunningham
Comput. Animat. Virtual Worlds1
2015 Integration and evaluation of emotion in an articulatory speech synthesis system
abstract
We convey a tremendous amount of information vocally. In addition to the obvious exchange of semantic information, we unconsciously vary a number of acoustic properties of the speech wave to provide information about our emotions, thoughts, and intentions. [Cahn 1990] Advances in understanding of human physiology combined with increases in the computational power available in modern computers have made the simulation of the human vocal tract a realistic option for creating artificial speech. Such systems can, in principle, produce any sound that a human can make. Here we present two experiments examining the expression of emotion using prosody (i.e., speech melody) in human recordings and an articulatory speech synthesis system.
Martin Schorradt, Katharina Legde, Susana Castillo 0001, Douglas W. Cunningham
SAP3
2015 Multimodal Affect: Perceptually Evaluating an Affective Talking Head
abstract
Many tasks such as driving or rapidly sorting items can be best achieved by direct actions. Other tasks such as giving directions, being guided through a museum, or organizing a meeting are more easily solved verbally. Since computers are increasingly being used in all aspects of daily life, it would be of great advantage if we could communicate verbally with them. Although advanced interactions with computers are possible, a vast majority of interactions are still based on the WIMP (Window, Icon, Menu, Point) metaphor [Hevner and Chatterjee 2010] and are, therefore, via simple text and gesture commands. The field of affective interfaces is working toward making computers more accessible by giving them (rudimentary) natural-language abilities, including using synthesized speech, facial expressions, and virtual body motions. Once the computer is granted a virtual body, however, it must be given the ability to use it to nonverbally convey socio-emotional information (such as emotions, intentions, mental state, and expectations) or it will likely be misunderstood. Here, we present a simple affective talking head along with the results of an experiment on the multimodal expression of emotion. The results show that although people can sometimes recognize the intended emotion from the semantic content of the text even when the face does not convey affect, they are considerably better at it when the face also shows emotion. Moreover, when both face and text convey emotion, people can detect different levels of emotional intensity.
Katharina Legde, Susana Castillo 0001, Douglas W. Cunningham
ACM Trans. Appl. Percept.2
2014 The semantic space for facial communication
abstract
ABSTRACT We can learn a lot about someone by watching their facial expressions and body language. Harnessing these aspects of non‐verbal communication can lend artificial communication agents greater depth and realism but requires a sound understanding of the relationship between cognition and expressive behaviour. Here, we extend traditional word‐based methodology to use actual videos and then extract the semantic/cognitive space of facial expressions. We find that depending on the specific expressions used, either a four‐dimensional or a two‐dimensional space is needed to describe the variance in the stimuli. The shape and structure of the 4D and 2D spaces are related to each other and very stable to methodological changes. The results show that there is considerable variance between how different people express the same emotion. The recovered space can well capture the full range of facial communication and is very suitable for semantic‐driven facial animation. Copyright © 2014 John Wiley & Sons, Ltd.
Susana Castillo 0001, Christian Wallraven, Douglas W. Cunningham
Comput. Animat. Virtual Worlds1
2011 Perceptual considerations for motion blur rendering
abstract
Motion blur is a frequent requirement for the rendering of high-quality animated images. However, the computational resources involved are usually higher than those for images that have not been temporally antialiased. In this article we study the influence of high-level properties such as object material and speed, shutter time, and antialiasing level. Based on scenes containing variations of these parameters, we design different psychophysical experiments to determine how influential they are in the perception of image quality. This work gives insights on the effects these parameters have and exposes certain situations where motion blurred stimuli may be indistinguishable from a gold standard. As an immediate practical application, images of similar quality can be produced while the computing requirements are reduced. Algorithmic efforts have traditionally been focused on finding new improved methods to alleviate sampling artifacts by steering computation to the most important dimensions of the rendering equation. Concurrently, rendering algorithms can take advantage of certain perceptual limits to simplify and optimize computations. To our knowledge, none of them has identified nor used these limits in the rendering of motion blur. This work can be considered a first step in that direction.
Fernando Navarro, Susana Castillo 0001, Francisco J. Serón, Diego Gutierrez
ACM Trans. Appl. Percept.2