Douglas W. Cunningham

dblp:25/4304 · DBLP profile ↗
← Back
40ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-1419-2552ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 12 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PE-CLIP: A Parameter-Efficient Fine-Tuning of Vision Language Models for Dynamic Facial Expression Recognition
abstract
The emergence of Vision-Language Models (VLMs) like Contrastive Language-Image Pretraining (CLIP) provides appealing solutions to various vision problems including Dynamic Facial Expression Recognition (DFER). However, most of the proposed approaches face major challenges, particularly related to inefficient full fine-tuning of the encoders and the complexity of the models. Moreover, some of the proposed methods seem to struggle with suboptimal performance due to (i) poor alignment between textual and visual representations, and (ii) ineffective temporal modeling. To address these challenges, we propose PE-CLIP, a parameter-efficient fine-tuning (PEFT) framework that elegantly adapts CLIP for dynamic facial expression recognition, requiring significantly reduced number of trainable parameters while maintaining high accuracy. At its core, to enhance efficiency and performance, PE-CLIP introduces two specialized adapters namely a Temporal Dynamic Adapter (TDA) and a Shared Adapter (ShA). The TDA is a GRU-based module with a dynamic scaling mechanism, capturing sequential dependencies while adaptively modulating the contribution of each temporal feature to emphasize the most informative ones while mitigating irrelevant variations. The ShA is a lightweight adapter refine representations within both textual and visual encoders, ensuring consistent feature processing while maintaining parameter efficiency. Additionally, we leverage Multi-modal Prompt Learning (MaPLe), which introduces learnable prompts to both visual and action unit-based textual description inputs, further improving the semantic alignment between modalities and enabling the efficient adaptation of CLIP for dynamic tasks. We evaluate our proposed PE-CLIP on two benchmark datasets, namely DFEW, FERV39K, and AFEW, achieving competitive performance compared to state-of-the-art methods while requiring fewer trainable parameters. By striking an optimal balance between parameter efficiency and performance, PE-CLIP sets a new benchmark in resource-efficient DFER. The source code of the proposed PE-CLIP will be publicly available at https://github.com/Ibtissam-SAADI/PE-CLIP .
Ibtissam Saadi, Abdenour Hadid, Douglas W. Cunningham, Abdelmalik Taleb-Ahmed, Yassin Elhillali
ACM Trans. Multim. Comput. Commun. Appl.3
2026 Corrections to "Perceptually Uniform Construction of Illustrative Textures"
abstract
This note corrects errors in Figs. 12 and 13 and the description of the parametric function in the paper "Perceptually Uniform Construction of Illustrative Textures" published in IEEE Transactions on Visualization and Computer Graphics, Vol. 30, Issue 1, 2024.
Anna Sterzik, Monique Meuschke, Douglas W. Cunningham, Kai Lawonn
IEEE Trans. Vis. Comput. Graph.3
2025 Uncertainty Visualization for Biomolecular Structures: An Empirical Evaluation
abstract
Uncertainty is an intrinsic property of almost all data, regardless of the data being measured, simulated, or generated. It can significantly influence the results and reliability of subsequent analysis steps. Clearly communicating uncertainties is crucial for informed decision-making and understanding, especially in biomolecular data, where uncertainty is often difficult to infer. Uncertainty visualization (UV) is a powerful tool for this purpose. However, previously proposed uncertainty visualization (UV) methods lack sufficient empirical evaluation. We collected and categorized visualization methods for portraying positional uncertainty in biomolecular structures. We then organized the methods into metaphorical groups and extracted nine representatives: color, clouds, ensemble, hulls, sausages, contours, texture, waves, and noise. We assessed their strengths and weaknesses in a twofold approach: expert assessments with six domain experts and three perceptual evaluations involving 1,756 participants. Through the expert assessments, we aimed to highlight the advantages and limitations of the individual methods for the application domain and discussed areas for necessary improvements. Through the perceptual evaluation, we investigated whether the visualizations are intuitively associated with uncertainty and whether the directionality of the mapping is perceived as intended. We also assessed the accuracy of inferring uncertainty values from the visualizations. Based on our results, we judged the appropriateness of the metaphors for encoding uncertainty and suggest further areas for improvement.
Anna Sterzik, Michael Krone, Daniel Baum, Douglas W. Cunningham, Kai Lawonn
IEEE Trans. Vis. Comput. Graph.4
2024 Driver's facial expression recognition: A comprehensive survey
Ibtissam Saadi, Douglas W. Cunningham, Abdelmalik Taleb-Ahmed, Abdenour Hadid, Yassin Elhillali
Expert Syst. Appl.2
2024 Perception of Line Attributes for Visualization
abstract
Line attributes such as width and dashing are commonly used to encode information. However, many questions on the perception of line attributes remain, such as how many levels of attribute variation can be distinguished or which line attributes are the preferred choices for which tasks. We conducted three studies to develop guidelines for using stylized lines to encode scalar data. In our first study, participants drew stylized lines to encode uncertainty information. Uncertainty is usually visualized alongside other data. Therefore, alternative visual channels are important for the visualization of uncertainty. Additionally, uncertainty-e.g., in weather forecasts-is a familiar topic to most people. Thus, we picked it for our visualization scenarios in study 1. We used the results of our study to determine the most common line attributes for drawing uncertainty: Dashing, luminance, wave amplitude, and width. While those line attributes were especially common for drawing uncertainty, they are also commonly used in other areas. In studies 2 and 3, we investigated the discriminability of the line attributes determined in study 1. Studies 2 and 3 did not require specific application areas; thus, their results apply to visualizing any scalar data in line attributes. We evaluated the just-noticeable differences (JND) and derived recommendations for perceptually distinct line levels. We found that participants could discriminate considerably more levels for the line attribute width than for wave amplitude, dashing, or luminance.
Anna Sterzik, Nils Lichtenberg, Jana Wilms, Michael Krone, Douglas W. Cunningham, Kai Lawonn
IEEE Trans. Vis. Comput. Graph.5
2024 Perceptually Uniform Construction of Illustrative Textures
abstract
Illustrative textures, such as stippling or hatching, were predominantly used as an alternative to conventional Phong rendering. Recently, the potential of encoding information on surfaces or maps using different densities has also been recognized. This has the significant advantage that additional color can be used as another visual channel and the illustrative textures can then be overlaid. Effectively, it is thus possible to display multiple information, such as two different scalar fields on surfaces simultaneously. In previous work, these textures were manually generated and the choice of density was unempirically determined. Here, we first want to determine and understand the perceptual space of illustrative textures. We chose a succession of simplices with increasing dimensions as primitives for our textures: Dots, lines, and triangles. Thus, we explore the texture types of stippling, hatching, and triangles. We create a range of textures by sampling the density space uniformly. Then, we conduct three perceptual studies in which the participants performed pairwise comparisons for each texture type. We use multidimensional scaling (MDS) to analyze the perceptual spaces per category. The perception of stippling and triangles seems relatively similar. Both are adequately described by a 1D manifold in 2D space. The perceptual space of hatching consists of two main clusters: Crosshatched textures, and textures with only one hatching direction. However, the perception of hatching textures with only one hatching direction is similar to the perception of stippling and triangles. Based on our findings, we construct perceptually uniform illustrative textures. Afterwards, we provide concrete application examples for the constructed textures.
Anna Sterzik, Monique Meuschke, Douglas W. Cunningham, Kai Lawonn
IEEE Trans. Vis. Comput. Graph.3
2023 Enhancing molecular visualization: Perceptual evaluation of line variables with application to uncertainty visualization
abstract
Data are often subject to some degree of uncertainty, whether aleatory or epistemic. This applies both to experimental data acquired with sensors as well as to simulation data. Displaying these data and their uncertainty faithfully is crucial for gaining knowledge. Specifically, the effective communication of the uncertainty can influence the interpretation of the data and the user’s trust in the visualization. However, uncertainty-aware visualization has gotten little attention in molecular visualization. When using the established molecular representations, the physicochemical attributes of the molecular data usually already occupy the common visual channels like shape, size, and color. Consequently, to encode uncertainty information, we need to open up another channel by using feature lines. Even though various line variables have been proposed for uncertainty visualizations, they have so far been primarily used for two-dimensional data and there has been little perceptual evaluation. Thus, we conducted two perceptual studies to determine the suitability of the line variables blur, dashing, grayscale, sketchiness, and width for distinguishing several values in molecular visualizations. While our work was motivated by uncertainty visualization, our techniques and study results also apply to other types of scalar data.
Anna Sterzik, Nils Lichtenberg, Michael Krone, Daniel Baum, Douglas W. Cunningham, Kai Lawonn
Comput. Graph.5
2022 Erosion as a novel Approach for removing Semantics and Comparison of different State-of-Art-Methods
abstract
Through language, people convey not only pure semantics, but also information about themselves, such as age, gender, state of mind or health. The supralingual features that carry this information have been a subject of research for a long time. Various procedures have been proposed to remove unneeded semantics from speech recordings, in order to study supralingual information in natural speech. In this paper, we propose a new method for removing sematics, based on erosion, a morphological operator. We compare its effectiveness to different state-of-the-art methods. As established methods we consider two low pass filters with cut off frequencies of 450Hz and 1150Hz and Brownian noise. As a newer method we investigate a filter for spectro-temporal frequencies. To evaluate each method, appropriately processed recordings were presented to a group of participants in a perceptual experiment. The intelligibility was measured by means of the Levenshtein distance. Our results show that erosion itself performs similarly to the established methods, while a combination of erosion and low-pass filter outperforms all other methods.
Martin Schorradt, Douglas W. Cunningham
SAP2
2022 Age Regression for Human Voices
abstract
The human voice is one of our most important tools for communicating with other people. Besides pure semantic meaning it also conveys syntactical information such as emphasis as well as personal information such as emotional state, gender, and age. While the physical changes that occur to a person’s voice are well studied, there is surprisingly little work on the perception of those changes. To hold the range of subtleties present in a given utterance constant and thus focus on the changes caused by age, this paper takes adult recordings (three males, and three females) and artificially resynthesizes them (using values from measurements of real children’s voices) to create a childlike versions of the utterance at different target ages. In particular, we focus on a systematic, factorial combination pitch shifting and formant shifting. To get an insight about the influence of these factors on the estimated age, we performed a perceptual experiment. Since the resynthesis method we used can produce a wide range of voices, not all of which are physically consistent, we also asked the participants to rate how natural the voices sounded. Furthermore, since former studies suggest that people are not able to distinguish between males and females of young ages, participants were also asked to rate how male or female the voices sounded. Overall, we found that although the synthesis method produced physically plausible signals (compared average values for real children), the degree of signal manipulation was correlated with perceived unnaturalness. We also found that pitch shift had only a small affect on perceived age, that formant shift had a strong affect on perceived age, and that these effects depended on the original gender of the recording. As expected, people had difficulty guessing the gender of younger sounding voices.
Martin Schorradt, Douglas W. Cunningham
ICMI2
2020 Enlighten Me: Importance of Brightness and Shadow for Character Emotion and Appeal
abstract
Lighting has been used to enhance emotion and appeal of characters for centuries, from paintings in the Renaissance to the modern-day digital arts. In VFX and animation studios, lighting is considered as important as modelling, shading, or rigging. Most existing work focuses on either empirical best-practice created by artists of the centuries or on lighting perception with basic shapes. In contrast, our work focuses on the effect of lighting on emotional characters. Our study presents an extensive set of novel perceptual experiments designed to investigate the effects of brightness levels (key light brightness) and the proportion of light intensity illuminating the two sides of a character’s face (key-to-fill ratio). We are particularly interested in the effect of lighting on the recognition of emotion, emotion intensity, and the overall appeal, as these are crucial factors for audience engagement. Our results have implications for artists and developers wishing to increase the appeal and emotional expression of their characters, ranging from cartoon to realistic styles. Our key finding is that lighting can be used to effectively alter the intensity of emotion of a character and that brighter conditions increased appeal across all of our experiments.
Pisut Wisessing, Katja Zibrek, Douglas W. Cunningham, John Dingliana, Rachel McDonnell
ACM Trans. Graph.3
2019 A psychophysical model to control the brightness and key-to-fill ratio in CG cartoon character lighting
abstract
Lighting is a commonly used tool to manipulate the appearance of virtual characters in a range of applications. However, there are few studies which systematically examine the effect of lighting changes on complex dynamic stimuli. Our study presents several perceptual experiments, designed to investigate the ability of participants to discriminate lighting levels and the ratio of light intensity projected on the two sides of a cartoon character’s face (key-to-fill ratio) in portrait lighting design. We used a standard psychophysical method for measuring discrimination, typical in low-level perceptual studies but not frequently considered for evaluating complex stimuli. We found that people can easily differentiate lighting intensities, and distinguish between shadow strength and scene brightness under bright conditions but not under dark conditions. We provide a model of the results, and empirically validate the predictions of the model. We discuss the practical implications of our results and how they can be exploited to make the process of portrait lighting for CG cartoon characters more consistent, such as a tool for manipulating shadow while maintaining the level of perceived brightness.
Pisut Wisessing, Katja Zibrek, Douglas W. Cunningham, Rachel McDonnell
SAP3
2019 Fitting the Style: The Semantic Space for Emotions on Stylized Faces
abstract
It has been widely proven among several research disciplines that the information conveyed through the visual channel can be critical for communication. Altering either the motion or the appearance of our interlocutor can change the opinion we form about them or lead us to drastically different conclusions on what is being communicated. In this paper, we explore the influence of altering the visual appearance of a virtual 3D face model through several different stylization techniques on the perception of dynamic conversational facial expressions. We propose the use of a semantic-differential task to recover the underlying semantic/cognitive space for the expressions and their corresponding shifts due to the application of the different styles. Given the appropriate amount of stimuli, we consider this technique well fitted to provide a mapping between meaning of emotions and spectrum of rendering parameters. Our results show that the different techniques are capable to alter the perception of the expressions making them, for example, more intense or ephemeral in a rather consistent manner.
Philipp Hahn, Susana Castillo 0001, Douglas W. Cunningham
CASA3
2019 Evaluating the Effect of Clothing and Environment on the perceived Personality of Virtual Avatars
abstract
Virtual avatars gain more and more importance in our everyday lives, especially in the field of human machine interaction and affective interfaces. The avatar can be used as a virtual surrogate to answer questions in online costumer support applications, they can virtually tutor students in e-learning scenarios, or they can mimic a real person's behavior in specific situations for training purposes in for example medical consulting sessions or accidents. It is easy to imagine that, just as with real people, appropriate clothing style and environment might be very important to sustain the intended reality and maintain reliability, believability, and seriousness of the virtual avatar. As clothing style is known to convey personal attributes of another person like sociability, orderliness, or openness it is highly recommended to carefully adjust the avatar's clothes to its usage, purpose, and intended goal. It is also known that environment influences how we feel; we can feel relaxed in calm and quite surroundings (like at home or in the nature) or stressed in loud and noisy settings (like a train station or an airport). While designing a human-machine-interface the environment can also have an effect on our own feelings and thereby influence how we perceive a virtual avatar.
Katharina Legde, Douglas W. Cunningham
IVA2
2019 Introduction to the Special Issue on SAP 2019
abstract
No abstract available.
Ludovic Hoyet, Douglas W. Cunningham
ACM Trans. Appl. Percept.2
2018 The semantic space for emotional speech and the influence of different methods for prosody isolation on its perception
abstract
Normally, when people talk to other people, they communicate not only using specific words, but also with intentional changes in their voice melody, facial expressions, and gestures. Not only is human communication inherently multimodal, it is also multi-layered. That is, it conveys more than simple semantic information, but also passes on a wide variety of social, emotional, and functional (e.g., conversation control) information. Previous work has examined the perception of socio-emotional information conveyed by words and facial expressions. Here, we build on that work and examine the perception of socio-emotional information based solely on prosody (e.g., speech melody, rate, tempo, intensity). To examine the perception of affective prosody, it is necessary to remove all semantics from the speech signal - without changing the prosody! In this paper, we compare several different state-of-the-art methods for removing semantics. We started by recording an audio database containing a German sentence spoken by 11 people in 62 different emotional states. We then removed or masked the semantics using three different techniques. We also recorded the same 62 states for a pseudo-language phrase. Each of these five sets of stimuli were subjected to a semantic differential rating task to derive and compare the semantic spaces for emotions. The results show that each of the methods successfully removed the semantic component, but also changed the perception of the emotional content. Interestingly, the pseudo-word stimuli diverged most from the normal sentences. Furthermore, although each of the filters affected the perception of the sentence in some manner, they did so in different ways.
Martin Schorradt, Susana Castillo 0001, Douglas W. Cunningham
SAP3
2018 Personality Analysis of Embodied Conversational Agents
abstract
People tend to personify machines. Giving machines the ability to actually produce social information can help improve human-machine interactions. Embodied Conversational Agents (ECAs) are virtual software agents that can process and produce speech, facial expressions, gestures and eye gaze, enabling natural, multimodal, human-machine communication. On the one hand, the field of personality psychology provides insights into how we could describe and measure the virtual personality of ECAs. On the other hand, ECAs provide a method to systematically examine how different factors affect the perception of personality. This paper shows that standardized, validated personality questionnaires can be used to evaluate ECAs psychologically, and that state of the art ECAs can manipulate their perceived personality through appearance and behavior.
Susana Castillo 0001, Philipp Hahn, Katharina Legde, Douglas W. Cunningham
IVA4
2018 Look Me in the Lines: The Impact of Stylization on the Recognition of Expressions and Perceived Personality
abstract
We are increasingly approaching the point where computer-based technology is truly ambient and omnipresent. People tend to personify their technical servants, including giving them human names as well as attributing personality traits and intentions to them. The more those devices advance from simple tools to intelligent assistants the more seriously we need to take this personification. That is, if the computers perform human-like tasks in collaboration with humans, and humans already tend to treat computers as human-like, it is only reasonable to give those devices a human-like appearance and conversational abilities. Therefore, one approach to design advanced human-machine interfaces relies heavily on the so-called Embodied Conversational Agents (ECAs). An ECA is a virtual software agent that can process and produce speech, facial expressions, gestures and eye-gaze and, as a result, enables natural, multimodal, human-machine communication. Decades of research in psychology and related fields have shown that the visual channel is especially important in human-human-communication, with subtle changes in both appearance and motion altering how a conversational partner is perceived. In this work, we examine the effectiveness of modifying a virtual character's visual appearance using well-known stylization techniques in order to alter its perceived personality. We also explore the effect of these techniques on the recognizability, intensity and sincerity of the character's displayed emotions.
Philipp Hahn, Susana Castillo 0001, Douglas W. Cunningham
IVA3
2018 The semantic space for motion-captured facial expressions
abstract
Abstract We cannot not communicate! During our daily lives, we convey information verbally and nonverbally. Most of the affective meaning of a message is transferred with the help of facial expressions, and thereby, when trying to establish a realistic human‐like virtual character, we should pay close attention to the animation. Motion capture is one of the most common techniques, but due to the wide range of expressions humans use, the recording time and data needed are vast. To address this problem, we propose the use of semantic spaces as they help in characterizing and positioning expressions by finding a correlation between them. In this paper, we extend prior research by providing the semantic spaces underlying real videos and motion capture data for a total of 62 conversational expressions. Our results highly correlate with previous work, showing that our new expressions were correctly recognized. Moreover, our results can be used in future work to directly project potential new recordings of these 62 expressions on the found spaces.
Susana Castillo 0001, Katharina Legde, Douglas W. Cunningham
Comput. Animat. Virtual Worlds3
2016 Gaze prediction using machine learning for dynamic stereo manipulation in games
abstract
Comfortable, high-quality 3D stereo viewing is becoming a requirement for interactive applications today. Previous research shows that manipulating disparity can alleviate some of the discomfort caused by 3D stereo, but it is best to do this locally, around the object the user is gazing at. The main challenge is thus to develop a gaze predictor in the demanding context of real-time, heavily task-oriented applications such as games. Our key observation is that player actions are highly correlated with the present state of a game, encoded by game variables. Based on this, we train a classifier to learn these correlations using an eye-tracker which provides the ground-truth object being looked at. The classifier is used at runtime to predict object category - and thus gaze - during game play, based on the current state of game variables. We use this prediction to propose a dynamic disparity manipulation method, which provides rich and comfortable depth. We evaluate the quality of our gaze predictor numerically and experimentally, showing that it predicts gaze more accurately than previous approaches. A subjective rating study demonstrates that our localized disparity manipulation is preferred over previous methods.
George Alex Koulieris, George Drettakis, Douglas W. Cunningham, Katerina Mania
VR3
2016 A Survey of Perceptually Motivated 3D Visualization of Medical Image Data
abstract
Abstract This survey provides an overview of perceptually motivated techniques for the visualization of medical image data, including physics‐based lighting techniques as well as illustrative rendering that incorporate spatial depth and shape cues. Additionally, we discuss evaluations that were conducted in order to study the perceptual effects of these visualization techniques as compared to conventional techniques. These evaluations assessed depth and shape perception with depth judgment, orientation matching, and related tasks. This overview of existing techniques and their evaluation serves as a basis for defining the evaluation process of medical visualizations and to discuss a research agenda.
Bernhard Preim, Alexandra Baer, Douglas W. Cunningham, Tobias Isenberg 0001, Timo Ropinski
Comput. Graph. Forum3
2015 Integration and evaluation of emotion in an articulatory speech synthesis system
abstract
We convey a tremendous amount of information vocally. In addition to the obvious exchange of semantic information, we unconsciously vary a number of acoustic properties of the speech wave to provide information about our emotions, thoughts, and intentions. [Cahn 1990] Advances in understanding of human physiology combined with increases in the computational power available in modern computers have made the simulation of the human vocal tract a realistic option for creating artificial speech. Such systems can, in principle, produce any sound that a human can make. Here we present two experiments examining the expression of emotion using prosody (i.e., speech melody) in human recordings and an articulatory speech synthesis system.
Martin Schorradt, Katharina Legde, Susana Castillo 0001, Douglas W. Cunningham
SAP4
2015 Multimodal Affect: Perceptually Evaluating an Affective Talking Head
abstract
Many tasks such as driving or rapidly sorting items can be best achieved by direct actions. Other tasks such as giving directions, being guided through a museum, or organizing a meeting are more easily solved verbally. Since computers are increasingly being used in all aspects of daily life, it would be of great advantage if we could communicate verbally with them. Although advanced interactions with computers are possible, a vast majority of interactions are still based on the WIMP (Window, Icon, Menu, Point) metaphor [Hevner and Chatterjee 2010] and are, therefore, via simple text and gesture commands. The field of affective interfaces is working toward making computers more accessible by giving them (rudimentary) natural-language abilities, including using synthesized speech, facial expressions, and virtual body motions. Once the computer is granted a virtual body, however, it must be given the ability to use it to nonverbally convey socio-emotional information (such as emotions, intentions, mental state, and expectations) or it will likely be misunderstood. Here, we present a simple affective talking head along with the results of an experiment on the multimodal expression of emotion. The results show that although people can sometimes recognize the intended emotion from the semantic content of the text even when the face does not convey affect, they are considerably better at it when the face also shows emotion. Moreover, when both face and text convey emotion, people can detect different levels of emotional intensity.
Katharina Legde, Susana Castillo 0001, Douglas W. Cunningham
ACM Trans. Appl. Percept.3
2014 C-LOD: Context-aware Material Level-of-Detail applied to Mobile Graphics
abstract
Abstract Attention‐based Level‐Of–Detail (LOD) managers downgrade the quality of areas that are expected to go unnoticed by an observer to economize on computational resources. The perceptibility of lowered visual fidelity is determined by the accuracy of the attention model that assigns quality levels. Most previous attention based LOD managers do not take into account saliency provoked by context, failing to provide consistently accurate attention predictions. In this work, we extend a recent high level saliency model with four additional components yielding more accurate predictions: an object‐intrinsic factor accounting for canonical form of objects, an object‐context factor for contextual isolation of objects, a feature uniqueness term that accounts for the number of salient features in an image, and a temporal context that generates recurring fixations for objects inconsistent with the context. We conduct a perceptual experiment to acquire the weighting factors to initialize our model. We design C‐LOD, a LOD manager that maintains a constant frame rate on mobile devices by dynamically re‐adjusting material quality on secondary visual features of non‐attended objects. In a proof of concept study we establish that by incorporating C‐LOD, complex effects such as parallax occlusion mapping usually omitted in mobile devices can now be employed, without overloading GPU capability and, at the same time, conserving battery power.
George Alex Koulieris, George Drettakis, Douglas W. Cunningham, Katerina Mania
Comput. Graph. Forum3
2014 The semantic space for facial communication
abstract
ABSTRACT We can learn a lot about someone by watching their facial expressions and body language. Harnessing these aspects of non‐verbal communication can lend artificial communication agents greater depth and realism but requires a sound understanding of the relationship between cognition and expressive behaviour. Here, we extend traditional word‐based methodology to use actual videos and then extract the semantic/cognitive space of facial expressions. We find that depending on the specific expressions used, either a four‐dimensional or a two‐dimensional space is needed to describe the variance in the stimuli. The shape and structure of the 4D and 2D spaces are related to each other and very stable to methodological changes. The results show that there is considerable variance between how different people express the same emotion. The recovered space can well capture the full range of facial communication and is very suitable for semantic‐driven facial animation. Copyright © 2014 John Wiley & Sons, Ltd.
Susana Castillo 0001, Christian Wallraven, Douglas W. Cunningham
Comput. Animat. Virtual Worlds3
2014 An Automated High-Level Saliency Predictor for Smart Game Balancing
abstract
Successfully predicting visual attention can significantly improve many aspects of computer graphics: scene design, interactivity and rendering. Most previous attention models are mainly based on low-level image features, and fail to take into account high-level factors such as scene context, topology, or task. Low-level saliency has previously been combined with task maps, but only for predetermined tasks. Thus, the application of these methods to graphics (e.g., for selective rendering) has not achieved its full potential. In this article, we present the first automated high-level saliency predictor incorporating two hypotheses from perception and cognitive science that can be adapted to different tasks. The first states that a scene is comprised of objects expected to be found in a specific context as well objects out of context which are salient (scene schemata) while the other claims that viewer’s attention is captured by isolated objects (singletons). We propose a new model of attention by extending Eckstein’s Differential Weighting Model. We conducted a formal eye-tracking experiment which confirmed that object saliency guides attention to specific objects in a game scene and determined appropriate parameters for a model. We present a GPU-based system architecture that estimates the probabilities of objects to be attended in real- time. We embedded this tool in a game level editor to automatically adjust game level difficulty based on object saliency, offering a novel way to facilitate game design. We perform a study confirming that game level completion time depends on object topology as predicted by our system.
George Alex Koulieris, George Drettakis, Douglas W. Cunningham, Katerina Mania
ACM Trans. Appl. Percept.3
2013 Visualizing Natural Image Statistics
abstract
Natural image statistics is an important area of research in cognitive sciences and computer vision. Visualization of statistical results can help identify clusters and anomalies as well as analyze deviation, distribution, and correlation. Furthermore, they can provide visual abstractions and symbolism for categorized data. In this paper, we begin our study of visualization of image statistics by considering visual representations of power spectra, which are commonly used to visualize different categories of images. We show that they convey a limited amount of statistical information about image categories and their support for analytical tasks is ineffective. We then introduce several new visual representations, which convey different or more information about image statistics. We apply ANOVA to the image statistics to help select statistically more meaningful measurements in our design process. A task-based user evaluation was carried out to compare the new visual representations with the conventional power spectra plots. Based on the results of the evaluation, we made further improvement of visualizations by introducing composite visual representations of image statistics.
Hui Fang 0003, Gary K. L. Tam, Rita Borgo, Andrew J. Aubrey, Phil W. Grant, Paul L. Rosin, Christian Wallraven, Douglas W. Cunningham, David Marshall 0001, Min Chen 0001
IEEE Trans. Vis. Comput. Graph.8
2011 Perceptual Evaluation of Ghosted View Techniques for the Exploration of Vascular Structures and Embedded Flow
abstract
Abstract This paper presents three controlled perceptual studies investigating the visualization of the cerebral aneurysm anatomy with embedded flow visualization. We evaluate and compare the common semitransparent visualization technique with a ghosted view and a ghosted view with depth enhancement technique. We analyze the techniques’ ability to facilitate and support the shape and spatial representation of the aneurysm models as well as evaluating the smart visibility characteristics. The techniques are evaluated with respect to the participants accuracy, response time and their personal preferences. We used as stimuli 3D aneurysm models of five clinical datasets. There was overwhelming preference for the two ghosted view techniques over the semitransparent technique. Since smart visibility techniques are rarely evaluated, this paper may serve as orientation for further studies.
Alexandra Baer, Rocco Gasteiger, Douglas W. Cunningham, Bernhard Preim
Comput. Graph. Forum3
2011 Computational Aesthetics 2011 in Vancouver, Canada, August 5-7, 2011, Sponsored by Eurographics, in Collaboration with ACM SIGGRAPH
abstract
International audience
Tobias Isenberg 0001, Douglas W. Cunningham
Comput. Graph. Forum2
2011 A Survey of Image Statistics Relevant to Computer Graphics
abstract
Abstract The statistics of natural images have attracted the attention of researchers in a variety of fields and have been used as a means to better understand the human visual system and its processes. A number of algorithms in computer graphics, vision and image processing take advantage of such statistical findings to create visually more plausible results. With this report we aim to review the state of the art in image statistics and discuss existing and potential applications within computer graphics and related areas.
Tania Pouli, Douglas W. Cunningham, Erik Reinhard
Comput. Graph. Forum2
2011 Perception-motivated interpolation of image sequences
abstract
We present a method for image interpolation that is able to create high-quality, perceptually convincing transitions between recorded images. By implementing concepts derived from human vision, the problem of a physically correct image interpolation is relaxed to that of image interpolation which is perceived as visually correct by human observers. We find that it suffices to focus on exact edge correspondences, homogeneous regions and coherent motion to compute convincing results. A user study confirms the visual quality of the proposed image interpolation approach. We show how each aspect of our approach increases perceived quality of the result. We compare the results to other methods and assess achievable quality for different types of scenes.
Timo Stich, Christian Linz, Christian Wallraven, Douglas W. Cunningham, Marcus A. Magnor
ACM Trans. Appl. Percept.4
2009 Computational Aesthetics 08
Douglas W. Cunningham, Victoria Interrante, Brian Wyvill
Comput. Graph.1
2009 Categorizing art: Comparing humans and computers
Christian Wallraven, Roland W. Fleming, Douglas W. Cunningham, Jaume Rigau, Miquel Feixas, Mateu Sbert
Comput. Graph.3
2008 Evaluating the perceptual realism of animated facial expressions
abstract
The human face is capable of producing an astonishing variety of expressions—expressions for which sometimes the smallest difference changes the perceived meaning considerably. Producing realistic-looking facial animations that are able to transmit this degree of complexity continues to be a challenging research topic in computer graphics. One important question that remains to be answered is: When are facial animations good enough? Here we present an integrated framework in which psychophysical experiments are used in a first step to systematically evaluate theperceptualquality of several different computer-generated animations with respect to real-world video sequences. The first experiment provides an evaluation of several animation techniques, exposing specific animation parameters that are important to achieve perceptual fidelity. In a second experiment, we then use these benchmarked animation techniques in the context of perceptual research in order to systematically investigate the spatiotemporal characteristics of expressions. A third and final experiment uses the quality measures that were developed in the first two experiments to examine the perceptual impact of changing facial features to improve the animation techniques. Using such an integrated approach, we are able to provide important insights into facial expressions for both the perceptual and computer graphics community.
Christian Wallraven, Martin Breidt, Douglas W. Cunningham, Heinrich H. Bülthoff
ACM Trans. Appl. Percept.3
2007 Evaluation of real-world and computer-generated stylized facial expressions
abstract
The goal of stylization is to provide an abstracted representation of an image that highlights specific types of visual information. Recent advances in computer graphics techniques have made it possible to render many varieties of stylized imagery efficiently making stylization into a useful technique, not only for artistic, but also for visualization applications. In this paper, we report results from two sets of experiments that aim at characterizing the perceptual impact and effectiveness of three different stylization techniques in the context of dynamic facial expressions. In the first set of experiments, animated facial expressions are stylized using three common techniques (brush, cartoon, and illustrative stylization) and investigated using different experimental measures. Going beyond the usual questionnaire approach, these experiments compare the techniques according to several criteria ranging from subjective preference to task-dependent measures (such as recognizability, intensity) allowing us to compare behavioral and introspective approaches. The second set of experiments use the same stylization techniques on real-world video sequences in order to compare the effect of stylization on natural and artificial stimuli. Our results shed light on how stylization of image contents affects the perception and subjective evaluation of both real and computer-generated facial expressions.
Christian Wallraven, Heinrich H. Bülthoff, Douglas W. Cunningham, Jan Fischer, Dirk Bartz
ACM Trans. Appl. Percept.3
2006 Measuring the Discernability of Virtual Objects in Conventional and Stylized Augmented Reality
Jan Fischer, Douglas W. Cunningham, Dirk Bartz, Christian Wallraven, Heinrich H. Bülthoff, Wolfgang Straßer
EGVE2
2006 A Psychophysical Examination of Swinging Rooms, Cylindrical Virtual Reality Setups, and Characteristic Trajectories
abstract
Virtual Reality (VR) is increasingly being used in industry, medicine, entertainment, education, and research. It is generally critical that the VR setups produce behavior that closely resembles real world behavior. One part of any task is the ability to control our posture. Since postural control is well studied in the real world and is known to be strongly influenced by visual information, it is an ideal metric for examining the behavioral fidelity of VR setups. Moreover, VR-based experiments on postural control can provide fundamental new insights into human perception and cognition. Here, we employ the "swinging room paradigm" to validate a specific VR setup. Furthermore, we systematically examined a larger range of room oscillations than previously studied in any single setup. We also introduce several new methods and analyses that were specifically designed to optimize the detection of synchronous swinging between the observer and the virtual room. The results show that the VR setup has a very high behavioral fidelity and that increases in swinging room amplitude continue to produce increases in body sway even at very large room displacements (+/- 80 cm). Finally, the combination of new methods proved to be a very robust, reliable, and sensitive way of measuring body sway.
Douglas W. Cunningham, Hans-Günther Nusseck, Harald J. Teufel, Christian Wallraven, Heinrich H. Bülthoff
VR1
2005 Human-Centered Fidelity Metrics for Virtual Environment Simulations
Katerina Mania, Heinrich H. Bülthoff, Douglas W. Cunningham, Bernard D. Adelstein, Nadia Magnenat-Thalmann, Nicholaos Mourkoussis, Tom Troscianko, J. Edward Swan II
VR3
2005 Manipulating Video Sequences to Determine the Components of Conversational Facial Expressions
abstract
Communication plays a central role in everday life. During an average conversation, information is exchanged in a variety of ways, including through facial motion. Here, we employ a custom, model-based image manipulation technique to selectively “freez ” portions of a face in video recordings in order to determine the areas that are sufficient for proper recognition of nine conversational expressions. The results show that most expressions rely primarily on a single facial area to convey meaning with different expressions using different areas. The results also show that the combination of rigid head, eye, eyebrow, and mouth motions is sufficient to produce expressions that are as easy to recognize as the original, unmanipulated recordings. Finally, the results show that the manipulation technique introduced few perceptible artifacts into the altered video sequences. This fusion of psychophysics and computer graphics techniques provides not only fundamental insights into human perception and cognition, but also yields the basis for a systematic description of what needs to move in order to produce realistic, recognizable conversational facial animations.
Douglas W. Cunningham, Mario Kleiner, Christian Wallraven, Heinrich H. Bülthoff
ACM Trans. Appl. Percept.1
2004 The role of image size in the recognition of conversational facial expressions
abstract
Abstract Facial expressions can be used to direct the flow of a conversation as well as to improve the clarity of communication. The critical physical differences between expressions can, however, be small and subtle. Clear presentation of facial expressions in applied settings, then, would seem to require a large conversational agent. Given that visual displays are generally limited in size, the usage of a large conversational agent would reduce the amount of space available for the display of other information. Here, we examine the role of image size in the recognition of facial expressions. The results show that conversational facial expressions can be easily recognized at surprisingly small image sizes. Copyright © 2004 John Wiley & Sons, Ltd.
Douglas W. Cunningham, Manfred Nusseck, Christian Wallraven, Heinrich H. Bülthoff
Comput. Animat. Virtual Worlds1
2003 How Believable Are Real Faces? Towards a Perceptual Basis for Conversational Animation
abstract
Regardless of whether the humans involved are virtual or real, well-developed conversational skills are a necessity. The synthesis of interface agents that are not only understandable but also believable can be greatly aided by knowledge of which facial motions are perceptually necessary and sufficient for clear and believable conversational facial expressions. Here, we recorded several core conversational expressions (agreement, disagreement, happiness, sadness, thinking, and confusion) from several individuals, and then psychophysically determined the perceptual ambiguity and believability of the expressions. The results show that people can identify these expressions quite well, although there are some systematic patterns of confusion. People were also very confident of their identifications and found the expressions to be rather believable. The specific pattern of confusions and confidence ratings have strong implications for conversational animation. Finally, the present results provide the information necessary to begin a more fine-grained analysis of the core components of these expressions.
Douglas W. Cunningham, Martin Breidt, Mario Kleiner, Christian Wallraven, Heinrich H. Bülthoff
CASA1