EDBT 2026 Demo / reviewers in the wild / expert
Alexander Raake
dblp:47/1247
· DBLP profile ↗
145ranked-venue papers
8as first author
61since 2021 · last 2026
0000-0002-9357-1763ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 124 · 5 first-author · 55 since 2021Human-computer interaction and ubiquitous computing · 62 · 1 first-author · 40 since 2021Artificial intelligence and machine learning · 24 · 6 first-author · 1 since 2021Computer networks · 5Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Accurate are Video Quality Models for Diffusion-Based Video Super-Resolution?abstractRecent video super-resolution (VSR) approaches use deep neural networks to enhance low-quality input videos and recover visual detail, with diffusion-based methods in particular showing promising results. In this paper, we investigate whether existing video quality models can be used to assess the performance of these diffusion-based VSR methods, by comparing model predictions with results from a subjective test. The study compares six upscaling methods (Lanczos, Rhea, SCST, DOVE, SeedVR2, Starlight Mini) applied to both compressed (AV1 and DCVC-RT) and uncompressed low-resolution videos considering the play-out on a UHD-1/4K screen. A range of full- and no-reference quality models are used to assess their applicability to this new type of quality degradation, focusing on within-sequence performance. The results highlight that CNN-based full-reference models, such as LPIPS, DISTS, and CVQA-FR show significantly higher correlation coefficients than both conventional full- as well as the tested no-reference models. Most overestimate the overly sharp results of SCST, with VMAF mainly failing due to spatial inconsistencies introduced by Starlight Mini. None of the tested video quality models reach sufficient accuracy so as to replace complementary subjective testing. The reference, degraded and upscaled videos, as well as the user ratings and model scores are made available with the paper at https://github.com/Telecommunication-Telemedia-Assessment/AVT-VQDB-UHD-1-VSR as open data. Benjamin Herb, Steve Goering, Alexander Raake, Rakesh Rao Ramachandra Rao |
QoMEX | 3 |
| 2026 | A Comparative User Study of Real-Time Head-Appearance Telepresence in 2D and 3D Representations
William Menz, Alexander Zoubarev, David Kutschke, Rakesh Rao Ramachandra Rao, Louay Bassbouss, Sven Bliedung von der Heide, Alexander Raake, Steve Goering |
QoMEX | 7 |
| 2026 | Between the Labs: An Analysis of Individual Rating Biases in Multi-Lab Subjective Experiments
Werner Robitza, Alexander Raake |
QoMEX | 2 |
| 2026 | Exploring Mediated Communication with Older Adults: Comparing AR Avatars, Telepresence Robots, and Face-to-Face InteractionabstractOlder adults, a growing demographic, face an increased risk of experiencing loneliness and are less exposed to emerging communication technologies. Augmented reality (AR) avatars and telepresence robots have been proposed as tools to foster social connection, yet their suitability for older users remains underexplored. We present an exploratory study with ten healthy older adults who engaged in both conversational and spatial collaboration tasks using AR avatar-mediated communication, robot-mediated communication, and face-to-face interaction. We collected self-reported measures of co-presence, social presence, closeness, uncanny valley, preferences, and open feedback. Our findings suggest that telepresence robots enhanced co-presence, while avatars were valued for their expressivity and humanlike qualities. Task type influenced co-presence in spatial collaboration only during communication using the telepresence robot. Other measures, such as social presence and closeness, were unaffected by task type or representation. While neither technology outperformed face-to-face interaction, both were positively received, underscoring their potential to address the social needs of older adults and highlighting the importance of enhancing nonverbal expressivity, particularly nonverbal cues in mediated communication. Ultimately, our results contribute to the fundamental understanding of mediated communication with older adults, motivating further empirical work to confirm and extend these findings. Stephanie Arevalo, Jakob Hartbrich, Florian Weidner, Melisa Conde, Veronika Mikhailova, Felix Immohr, Söhnke Benedikt Fischedick, Bea Vorhof, Christoph Gerhardt, Kay Richter, Christian Kunert, Nicola Döring, Horst-Michael Groß, Wolfgang Broll, Alexander Raake |
IMX | 15 |
| 2026 | Exploration of the Effect of Automatic Camera Tracking for TeleconferencingabstractConsidering the increase in using teleconferencing in the daily lives of everyone, it becomes more important to provide a good experience while attending remote meetings. Various new technologies have been proposed, ranging from AR/VR setups to enhanced 2D systems with the added possibility of automatic tracking of a person. This exploration paper investigates the impact of automatic camera tracking on virtual communication, focusing on task performance. Using a custom video conference setup combined with the TrackingMaster system (2D LiDAR-based tracking and PTZ camera framing), three tracking configurations were tested: static, tracking-without-presets, and tracking-with-presets. Participants performed various adapted survival tasks designed to encourage movement and interaction. The results showed that tracking with presets enabled slightly faster task completion and required less movement. Furthermore, workload and social presence indicated no major variation across the tracking configurations. The findings suggest potential benefits of camera tracking systems but highlight the need for further research with larger samples and novel measurements to create validated metrics, especially regarding creativity in virtual communication. Christoph Götzl, Felix Immohr, Alexander Raake, Steve Goering |
IMX | 3 |
| 2026 | "The Robot Should Be Programmed for Me": User Tests Evaluating a Telepresence Robot for the Social Integration of Older AdultsabstractTelepresence robots that allow communication between older adults and their remotely located social contacts can foster social integration. The present laboratory test study explores older adults’ successful use of a telepresence robot (Research Question 1 [RQ1]), as well as their perceived enjoyment (RQ2), perceived ease of use (RQ3), perceived usefulness (RQ4), perceived social presence (RQ5), and intention to use (RQ6) a telepresence robot for robot-mediated communication (RMC). Semi-structured interviews, observations, and questionnaires were applied with a group of N = 14 older adults living in Germany. Participants completed a navigational task (as remote users) and an interpersonal communication task (as local users). Results show older adults used the telepresence robot successfully (RQ1) during the tasks. Furthermore, in interviews, older adults described their perceived enjoyment (RQ2), perceived ease of use (RQ3), and perceived usefulness (RQ4) during RMC as generally high. Perceived social presence (RQ5) during RMC was generally described as high, with RMC being considered a viable substitute when face-to-face communication is not possible. Finally, only two participants (2/14) had no intention to use (RQ6) a telepresence robot in the long term. Future design recommendations are provided, such as adapting the telepresence robot’s functions to older adults physical, psychological, and social conditions. Melisa Conde, Söhnke Benedikt Fischedick, Kay Richter, Stephanie Arevalo, Horst-Michael Groß, Alexander Raake, Nicola Döring |
ACM Trans. Hum. Robot Interact. | 6 |
| 2025 | Exploiting LLMs for Metadata-Based Video Quality PredictionabstractLarge language models (LLMs) can be used to solve various tasks based on text inputs, e.g., video quality estimation. We explore the usage of LLMs for video quality prediction based on metadata (video codec, bitrate, resolution), which has not been addressed before. For the evaluation we use test #1 from our AVT-VQDB-UHD-1 dataset. We generated text prompts based on the metadata and collected answers from 17 different LLMs. The evaluation indicates that especially larger LLMs could be used to simulate human raters. However, a pure model prediction with one model has lower performance than SoA video quality models. Thus, we further investigated combinations of LLMs, which resulted in comparable performance to state-of-the-art models. Our work is a proof-of-concept, considering that LLMs are slower for the prediction than the traditional metadata-based models. Steve Goering, Rakesh Rao Ramachandra Rao, Alexander Raake |
ISM | 3 |
| 2025 | Smarter Traps: Neural Network-Driven Classification of Small MammalsabstractThe intensification of agricultural practices has substantially altered ecosystem structures, affecting small mammal populations through habitat homogenization and insecticide use. Traditional monitoring of these species relies on physical trapping, which is labor-intensive and stressful for the animals. To address these challenges, this study explores an automated image-based approach for monitoring small mammals using open traps equipped with cameras. The goal was to design a system that runs efficiently on standard, low-performance computers. Therefore, existing convolutional neural network (CNN) models were retrained and evaluated for automated genus classification of captured images. The networks were retrained with a transfer learning approach using a custom dataset containing four categories, three mammal genera, and one “empty trap” class. Multiple CNN architectures were compared based on loss, accuracy, and macro F1-score to identify a model that balances performance and computational efficiency. Results showed that MobileNet-based architectures, optimized for low-power devices, underperformed in this classification task, while VGG-based networks achieved superior accuracy and generalization to unseen images from the same trap setup. The findings demonstrate the potential of CNN-driven image recognition as a scalable and noninvasive tool for ecological monitoring, reducing manual review effort and improving animal welfare in field studies. William Menz, Ralf Dittrich, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
ISM | 5 |
| 2025 | Evaluation of a Floating-Head Communication Prototype for Video-ConferencingabstractBesides traditional 2D video communication approaches, new systems aim to create more realistic and immersive representations of the involved conversation partners. This work presents a real-time floating-head communication setup that enables spatially separated visualizations of remote participants using only a standard camera and display. The underlying reconstruction pipeline applies machine-learning-assisted facial feature extraction to infer a 3D mesh of the participant's head from a live 2D video stream. Texturing maintains visual fidelity while supporting real-time performance. A dedicated transmission pipeline enables the exchange of 3D and texture data over conventional network connections, allowing flexible and location-independent use. A first lab test with eight participant pairs was performed to evaluate the system during a collaborative communication task. Subjective assessments using established telepresence and quality questionnaires confirmed the technical feasibility of the approach and its potential to enhance the sense of spatial presence. However, the overall perceptual quality and comfort did not yet reach the level of classical 2 D video communication. The study demonstrates the promise of accessible, spatially expressive communication setups that may bridge the gap between conventional video calls and emerging volumetric telepresence systems. William Menz, Alexander Zoubarev, David Kutschke, Rakesh Rao Ramachandra Rao, Louay Bassbouss, Sven Bliedung von der Heide, Steve Goering, Alexander Raake |
ISM | 8 |
| 2025 | ICS-MR: Interactive Conversation Scenarios for Assessment of Mixed Reality CommunicationabstractWe present ICS-MR, a dataset containing three conversational scenarios designed for the evaluation of communication quality in Mixed Reality (MR) systems. Along with detailed descriptions of the conversation tasks, we provide all the materials required to incorporate the tasks into MR user studies. The materials also support application of the scenarios in real-world and video-conferencing contexts for studies that, for example, call for comparison of immersive systems against reference communication media. Open-source Unity implementations of the scenarios are also made available, supporting direct usage of the scenarios in distributed, multi-user experiments. The conversation tasks have all been administered in recent scientific works that address the evaluation of user experiences in immersive communication systems, allowing analysis and comparison of each scenario's evoked behavioral properties. The ICS-MR dataset therefore contributes valuable resources for further research on communication in immersive systems. Felix Immohr, Gareth Rendle, Annika Neidhardt, Anton Benjamin Lammert, Bernd Fröhlich 0001, Alexander Raake |
ACM Multimedia | 6 |
| 2025 | AMIS: An Audiovisual Dataset for Multimodal XR ResearchabstractThe Audiovisual Multimodal Interaction Suite (AMIS) is an open-source dataset and accompanying Unity-based demo implementation designed to aid research on immersive media communication and social XR environments. AMIS features synchronized audiovisual recordings of three actors performing monologues and participating in dyadic conversations across four modalities: talking-head videos, full-body videos, volumetric avatars, and personalized animated avatars. These recordings can be used to simulate scenarios such as traditional video conferences or XR meetings with 3D avatars in controlled and replicable environments. The limitations of existing datasets, which include a restricted number of audiovisual formats, a narrow application focus, and suboptimal inclusion of verbal and non-verbal cues, are addressed by AMIS. With AMIS Studio, a Unity-based demonstrator, researchers can explore the recordings and compare the different audiovisual formats in VR scenes. This paper outlines the creation of AMIS, its design considerations, and how it may be applied in interdisciplinary domains, including cognitive psychology, audiovisual quality assessment, and social XR research. Abhinav Bhattacharya, Luís Fernando de Souza Cardoso, Andy Schleising, Gareth Rendle, Adrian Kreskowski, Felix Immohr, Rakesh Rao Ramachandra Rao, Wolfgang Broll, Alexander Raake |
MMSys | 9 |
| 2025 | Evaluating Video Quality Metrics for Neural and Traditional Codecs using 4K/UHD-1 Videos
Benjamin Herb, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
PCS | 4 |
| 2025 | Sensory Evaluation of HDR Display Properties
Julius Prenzel, Dominik Keller, Rakesh Rao Ramachandra Rao, Alexander Raake |
PCS | 4 |
| 2025 | Fine-Grained HDR Image Quality Assessment From Noticeably Distorted to Very High FidelityabstractHigh dynamic range (HDR) and wide color gamut (WCG) technologies significantly improve color reproduction compared to standard dynamic range (SDR) and standard color gamuts, resulting in more accurate, richer, and more immersive images. However, HDR increases data demands, posing challenges for bandwidth efficiency and compression techniques. Advances in compression and display technologies require more precise image quality assessment, particularly in the high-fidelity range where perceptual differences are subtle. To address this gap, we introduce AIC-HDR2025, the first such HDR dataset, comprising 100 test images generated from five HDR sources, each compressed using four codecs at five compression levels. It covers the high-fidelity range, from visible distortions to compression levels below the visually lossless threshold. A subjective study was conducted using the JPEG AIC-3 test methodology, combining plain and boosted triplet comparisons. In total, 34,560 ratings were collected from 151 participants across four fully controlled labs. The results confirm that AIC-3 enables precise HDR quality estimation, with 95% confidence intervals averaging a width of 0.27 at 1 JND. In addition, several recently proposed objective metrics were evaluated based on their correlation with subjective ratings. The dataset is publicly available1. Mohsen Jenadeleh, Jon Sneyers, Davi Lazzarotto, Shima Mohammadi, Dominik Keller, Atanas Boev, Rakesh Rao Ramachandra Rao, António M. G. Pinheiro, Thomas Richter 0005, Alexander Raake, Touradj Ebrahimi, João Ascenso, Dietmar Saupe |
QoMEX | 10 |
| 2025 | A Large-Scale Evaluation of Subject Rating Behaviour in Visual Quality Assessment StudiesabstractSubjective testing is widely used for visual quality assessment to evaluate the impact of both technical and non-technical factors on user perception. Although standardized methods exist for collecting subjective ratings of visual quality, these ratings are inevitably influenced by each subject’s accuracy, manifesting as subject bias and inconsistency. Recommendations such as ITU-T P.910 and ITU-R BT.500 propose standardized methods to remove bias from subjective ratings. For instance, Annex E of ITU-T Rec. P.910 provides an effective strategy for addressing both subject bias and inconsistency. In this paper, we analyze 29 different visual quality assessment studies conducted over an eight-year period to understand the rating behaviour of subjects using these methods. Our investigation focuses on subjective studies targeting 4K, 8K, as well as 360°video and high-resolution image quality assessment. In the context of 4K video quality assessment, both SDR and HDR evaluations are considered. Our dataset covers use cases of short-term video quality and overall session quality assessment for HTTP-based adaptive streaming (HAS). For both these use cases, we propose a range of subject bias and inconsistency values that can serve as a reference for future studies. Furthermore, we define six different measures that can be used to assess the reliability of future studies and provide reference values for these measures. Following an open-science approach, all individual ratings from the included subjective tests, along with the results of the large-scale analysis, are made publicly available with this paper. Rakesh Rao Ramachandra Rao, Steve Goering, Stephan Fremerey, Dominik Keller, Alexander Raake |
QoMEX | 5 |
| 2025 | Influence of Audiovisual Realism on Communication Behaviour in Group-to-Group TelepresenceabstractGroup-to-group telepresence systems immerse geographically separated groups in a shared interaction space where remote users are represented as avatars. Notably, such systems allow users to interact with collocated and remote interlocutors simultaneously. In this context, where virtual user representations can be directly compared with real users, we investigate how visual realism (avatar type) and aural realism (presence of spatial audio) affect communication. Furthermore, we examine how communication differs between collocated and remote pairs of interlocutors. In our user study, groups of four participants perform a collaborative conversation task under the aforementioned visual and aural realism conditions. Our results indicate that avatar realism has positive effects on subjective ratings of perceived message understanding and group cohesion, and yields behavioural differences that indicate more interactivity and engagement. Few significant effects of aural realism were observed. Comparisons between collocated and remote communication found that collocated communication was perceived as more effective, but that more visual attention was paid to both remote participants than the collocated user. Gareth Rendle, Felix Immohr, Christian Kehling, Anton Benjamin Lammert, Adrian Kreskowski, Karlheinz Brandenburg, Alexander Raake, Bernd Fröhlich 0001 |
VR | 7 |
| 2025 | Errors matter! The Influence of Error Frequency and Timing on Fluency and Workload in a Virtual Human-Cobot Collaboration EnvironmentabstractCobots are widespread among manufacturing industries, where humans and robots work together in assembly tasks. However, cobots are not free from errors, which may affect collaboration in terms of human-robot workflow fluency. We present a study (N = 21) where we simulate a manufacturing workcell in VR to investigate how a cobot’s faulty behavior can affect the perception of human-robot collaboration during assembly, considering fluency (human-robot fluency, robot relative contribution, positive teammate traits, trust and human-robot working alliance goal) and workload. For that, participants assembled, together with a cobot, a complex piece inside electric vehicles (gearbox e-axle) five times. The assembly process consisted of six different steps to be performed sequentially, and we implemented a faulty behavior where the robot would follow the wrong sequence. After an error occurred, we allowed participants to choose freely between two error recovery behaviors: an automatic error recovery (the robot would undo the last step and continue the assembly process) or a training mode (the human would indicate the right sequence to the robot). Our results confirm that a cobot’s faulty behavior negatively affects the perceived fluency and workload, but does not affect the perceived cobot’s relative contribution to the task. The temporal sequencing of an error (error timing) can affect the goal of the human-cobot working alliance, assembly duration, temporal demand, and effort. Also, error frequency impacts differently the perceived physical demand. These findings can inform design choices of human-robot collaborative workflows in manufacturing, specifically, in error handling and mitigation strategies that preserve a positive perception of human-robot teams. Hamd Mehfooz-Khan, Stephanie Arevalo, Alexander Raake |
VRST | 3 |
| 2025 | Robot, Avatar, or Human: The Impact of Partner Representation and Task on the Communication ExperienceabstractAvatars and telepresence robots have long received attention for remote communication. However, the specific nature of their physicality, expressiveness, and mobility may affect their usefulness for different tasks. This work compares using an avatar (presented in augmented reality) and a telepresence robot to Face-to-Face (F2F) communication during different communication tasks: free conversation, negotiation, and referential communication with movement. We conducted a user study (split-plot design, N=54) with the type of representation of the conversational partner as the within variable and the communication task as the between variable. Our results show that the type of task, especially referential communication with movement, influenced the perceived attention to nonverbal cues and closeness. Generally, gestures and body movements received the least focus with telepresence robots. Gestures in avatars and F2F drew similar attention, which we attribute to the avatar's tracking fidelity. Gaze received less attention in both avatar- and robot-mediated communication compared to F2F, while facial expressions on the robot's screen heightened attention compared to avatars. These findings advance the fundamental understanding of mediated communication and support researchers and practitioners in shaping the design of communication applications beyond today's video calls. Stephanie Arevalo, Jakob Hartbrich, Florian Weidner, Söhnke Benedikt Fischedick, Christoph Gerhardt, Kay Richter, Christian Kunert, Bea Vorhof, Horst-Michael Groß, Wolfgang Broll, Alexander Raake |
Proc. ACM Hum. Comput. Interact. | 11 |
| 2025 | The Effect of Hand Visibility in AR: Comparing Dexterity and Interaction with Virtual and Real ObjectsabstractHand-tracking technologies allow us to use our own hands to interact with real and virtual objects in Augmented Reality (AR) environments. This enables us to explore the interplay between hand-visualizations and hand-object interactions. We present a user study that examines the effect of different hand visualizations (invisible, transparent, opaque) on manipulation performance when interacting with real and virtual objects. For this, we implemented video-see-through (VST) AR-based virtual building blocks and hot wire tasks with real one-to-one counterparts that require participants to use gross and fine motor hand movements. To evaluate manipulation performance, we considered three measures: task completion time, number of collisions (hot wire task), and percentage of object displacement (building block task). Additionally, we explored the sense of agency and subjective impressions (preference, ease of interaction, successful and awkwardness) evoked by the different hand-visualizations. The results show that (1) manipulation performance is significantly higher when interacting with real objects compared to virtual ones, (2) invisible hands lead to fewer errors, higher agency, higher perceived success and ease of interaction during fine manipulation tasks with real objects, and (3) having some visualization of the virtual hands (transparent or opaque) overlayed on the real hands is preferred when manipulating virtual objects even when there are no significant performance improvements. Our empirical findings about the differences when interacting with real and virtual objects can aid hand visualization choices for manipulation tasks in AR. Jakob Hartbrich, Stephanie Arevalo, Steve Goering, Alexander Raake |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Eyes on the Narrative: Exploring the Impact of Visual Realism and Audio Presentation on Gaze Behavior in AR StorytellingabstractAugmented Reality (AR) and Virtual Reality (VR) are essential tools for researchers and practitioners, serving purposes from training to entertainment: many of these applications rely on agents. This study explores the impact of agent characteristics on user reactions, focusing on gaze as a primary visual attention indicator in AR and VR. While existing research has investigated the agent’s gaze and its influence on the user, it is unclear how the agent’s auralization and visualization influence gaze behaviour. We investigate this by studying the impact of rendering style and type of audio on gaze behaviour during a narrative AR experience. Participants listened to a story with the agent visualized as a cartoon-style or realistic virtual human and auralized with spatial or non-spatial audio. The results revealed that the agent’s rendering style significantly influenced gaze behaviour, with cartoon-style agents capturing more visual attention. Audio variations did not yield significant differences. Together, our findings inform the design of AR user interfaces with agents, suggesting that low-realism visualizations are more captivating and, thus, more suitable for experiences where the user is supposed to look at the storyteller. Florian Weidner, Jakob Hartbrich, Stephanie Arevalo, Christian Kunert, Christian Schneiderwind, Chenyao Diao, Christoph Gerhardt, Tatiana Surdu, Wolfgang Broll, Stephan Werner 0003, Alexander Raake |
ETRA | 11 |
| 2024 | Nonverbal Dynamics in Dyadic Videoconferencing Interaction: The Role of Video Resolution and Conversational QualityabstractIn this paper, we investigated the influence of video resolution and perceived conversational quality on nonverbal behaviors during dyadic videoconferencing (VC) conversations. We analyzed nonverbal behaviors at the individual and interpersonal level. At the individual level, we considered body motion, facial expressions, and gaze directivity. At the interpersonal level, we considered facial expression synchrony and body movement synchrony. For the analysis, we used webcam recordings from a VC experiment, extracting the aforementioned individual nonverbal behavioral features and using windowed lagged cross-correlation (WLCC) to quantify the degree of interpersonal synchronization. Our results indicate that high video resolution significantly increased individual body movements and encouraged gaze directivity toward the conversational partner, fostering greater engagement while paradoxically reducing body movement synchrony. Higher conversational quality was associated with increased facial expression synchrony between participants. Moreover, we observed that instantaneous synchrony (as quantified with lag-zero WLCC) for both body movement and facial expressions was significantly influenced by mutual gaze-like behavior. These findings indicate a complex relationship between technical settings and nonverbal behaviors, suggesting that while higher resolution enhances some nonverbal behaviors, especially body movement and mutual gaze, it may disrupt body movement synchronization. These insights could be applied to the VC setups to achieve a high level of interpersonal coordination and engagement. Chenyao Diao, Stephanie Arevalo, Alexander Raake |
ICMI | 3 |
| 2024 | Appeal prediction for AI up-scaled ImagesabstractDNN- or AI-based up-scaling algorithms are gaining in popularity due to the improvements in machine learning. Various up-scaling models using CNNs, GANs or mixed approaches have been published. The majority of models are evaluated using PSRN and SSIM or only a few example images. However, a performance evaluation with a wide range of real-world images and subjective evaluation is missing, which we tackle in the following paper. For this reason, we describe our developed dataset, which uses 136 base images and five different up-scaling methods, namely Real-ESRGAN, BSRGAN, waifu2x, KXNet, and Lanczos. Overall the dataset consists of 1496 annotated images. The labeling of our dataset focused on image appeal and has been performed using crowd-sourcing employing our open-source tool AVRate Voyager. We evaluate the appeal of the different methods, and the results indicate that Real-ESRGAN and BSRGAN are the best. Furthermore, we train a DNN to detect which up-scaling method has been used, the trained models have a good overall performance in our evaluation. In addition to this, we evaluate state-of-the-art image appeal and quality models, here none of the models showed a high prediction performance, therefore we also trained two own approaches. The first uses transfer learning and has the best performance, and the second model uses signal-based features and a random forest model with good overall performance. We share the data and implementation to allow further research in the context of open science. Steve Goering, Rasmus Merten, Alexander Raake |
ISM | 3 |
| 2024 | Investigating the Impact of High Frame Rate on Video Quality: A SAMVIQ ApproachabstractHigh Frame Rate (HFR) aims at increasing the perceived video quality by decreasing motion artifacts and enabling a smoother playback of movements. However, HFR is not yet widely used in video playback, as most movies are partly due to artistic reasons still shot and shown at 24 frames per second (fps) and streamed videos are usually capped at 60 fps. This raises the question of whether people can perceive differences with higher frame rates and a connected increase in quality at all. To this effect, this paper analyzes the relationship between frame rate and perceived video quality using the Subjective Assessment Methodology for Video Quality (SAMVIQ). In the test, 24 subjects assessed the video quality of 16 sources with varied frame rates. The results show an increased video quality for videos with higher frame up to 120 fps. The SAMVIQ methodology is useful and our version is made publicly available. Dominik Keller, Paul Rudi Frank, Steve Goering, Alexander Raake |
ISM | 4 |
| 2024 | Effects of Delay on Nonverbal Behavior and Interpersonal Coordination in Video ConferencingabstractIn this paper, we investigated the effects of transmission delay on individual nonverbal behavior and interpersonal coordination during dyadic video conferencing conversations. For that, we assessed individual-level nonverbal behaviors, including body motion and gaze patterns, and examined participants' interpersonal coordination of body movement. Our results indicate that transmission delay significantly reduces individual body motion. No significant differences in gaze behaviors were found between different delay conditions, however, a trend of participants spending more time looking at their conversational partner in the high-delay condition than in the no-delay condition was observed. Also, we found that transmission delay significantly influences interpersonal body-movement coordination, enhancing structural organization and coordination stability while showing a threshold effect on movement similarity. Chenyao Diao, Stephanie Arevalo, Alexander Raake |
MMSP | 3 |
| 2024 | Evaluating Visually Lossless Compression of JPEG XS, JPEG 2000, HEVC and AV1 in Selected Medical Imaging ModalitiesabstractThe objective of this study is to evaluate the effectiveness of state-of-the-art codecs in compressing selected medical imaging modalities while maintaining visual quality and reducing file sizes. To achieve this, a detailed comparative analysis is conducted comparing the performance of JPEG XS, JPEG 2000, HEVC, and AV1. The analysis takes into consideration compression efficiency, codec complexity, and visual fidelity in the context of medical imaging. Advanced evaluation methods, including the AIC-2 Flicker test, are utilized to determine the visually lossless threshold, which is crucial for preserving diagnostically important details. Additionally, the study explores the potential of crowd-sourcing as a means of assessing the visual quality of compressed medical images. Subjective lab and crowd-sourcing tests reveal varying proportions of correctly identifying the reference images among participants. Furthermore, the study proposes outlier detection methods to improve the reliability of the subjective evaluation and employs kappa analysis to measure the inter-rater agreements. The study analyzes the encoding time taken on a consumer-level CPU, and the results reveal that JPEG XS maintains a fast and consistent speed across different compression levels. The results also indicate that JPEG XS achieves visually lossless performance for diagnostic purposes at 2 BPP, JPEG 2000 at 1.5 BPP, HEVC, and AV1 at 1 BPP. Bassem Elmeligy, Thomas Richter 0005, Rakesh Rao Ramachandra Rao, Siegfried Fößel, Alexander Raake |
QoMEX | 5 |
| 2024 | AVT-ECoClass-VR: An open-source audiovisual 360° video and immersive CGI multi-talker dataset to evaluate cognitive performanceabstractThe paper is part of a project to assess how complex visual and acoustic scenes affect cognitive performance in classroom scenarios, across age groups from children to adults. Here, the potential of audiovisual virtual environments for systematic user studies is explored. As of now, most studies have examined rather simple acoustic and visual representations, which do not reflect the reality of school children. An adapted version of the audiovisual scene analysis paradigm is presented, focusing on the localization and identification of talkers within a scene. The dataset includes two audiovisual scenarios (360° video and computer-generated imagery) and two implementations for dataset playback. The paper details the recording and post-processing of the content. The 360° video part of the dataset features 200 video and single-channel audio recordings of 20 speakers reading ten stories, and 20 videos of speakers in silence, resulting in a total of 220 video and 200 audio recordings. The dataset also includes one 360° background image of a real primary school classroom scene, targeting young school children for subsequent subjective tests. All stories were recorded in the German language with native German speakers. The second part of the dataset comprises 20 different 3D models of the speakers and a computer-generated classroom scene, along with an immersive audiovisual virtual environment implementation that can be interacted with using an HTC Vive controller. Both implementations also include a Unity plugin to connect and interact with the Virtual Acoustics auralization software. As a proof of concept, the dataset includes example output data collected from ongoing perception tests. There, subjects have the task of identifying which talker in the scene is reading out which story, using the story-to-speaker mapping input system developed within this paper. Stephan Fremerey, Carolin Breuer, Larissa Leist, Maria Klatte, Janina Fels, Alexander Raake |
QoMEX | 6 |
| 2024 | The Frankenstone toolbox for video quality analysis of user-generated contentabstractUser-generated video content is one major part of currently streamed video content. Providers such as YouTube, Twitch, or Vimeo provide thousands of videos to users. However, the quality of user-generated content can vary widely, if not only purely technical, quality-related aspects are considered, but also the liking of the content is taken into account. Several studies and published open-source models aim to predict the quality scores of user-generated content. We propose in this paper a unified toolbox – Frankenstone – that includes the latest video quality prediction models for user-generated content. As well as recently published models also meta-data and signal-based features are included in the toolbox. The Frankenstone toolbox relies on the usage of GPUs for the calculation. We evaluate our toolbox with the test data of the YouTube UGC Dataset. Steve Goering, Alexander Raake |
QoMEX | 2 |
| 2024 | Subjective Evaluation of the Impact of Spatial Audio on Triadic Communication in Virtual RealityabstractVirtual Reality (VR) enables users to meet, converse, and collaborate in shared virtual environments. For such communication systems, many system factors can affect user experience and perception. To effectively allocate system resources, understanding of the relative influence of such factors is required. One important factor is a spatial auralization, which has been shown to elevate users’ experience in traditional and single-user VR systems. However, its effect in multi-party social VR has not been fully investigated. In this work, we conducted a study assessing the effect of spatial audio on audiovisual plausibility and presence perception in a three-user interactive communication scenario. Triads of participants perform a collaborative conversation task under three conditions: a VR condition with binaural spatial audio, a VR condition with simple diotic audio, and a real-world reference condition. This paper presents the results of the study based on questionnaire-based evaluation. Felix Immohr, Gareth Rendle, Christian Kehling, Anton Benjamin Lammert, Steve Goering, Bernd Fröhlich 0001, Alexander Raake |
QoMEX | 7 |
| 2024 | AVT-VQDB-UHD-2-HDR: An open 8K HDR source dataset for video quality researchabstractMore and more screens with ultra high resolution and High Dynamic Range (HDR) are available on the market. Furthermore, video streaming providers like YouTube already offer contents in HDR with resolutions up to 8K (UHD-2) and a further increase in available material is conceivable. Therefore, manufacturers and providers need video quality models to measure their customers’ Quality of Experience. Currently, the most used video quality models only support resolutions of up to 4K (UHD-1). However, a valid estimation of the video quality of contents in 8K resolution is getting more important. In this paper, we present the AVT-VQDB-UHD-2-HDR dataset consisting of 31 8K HDR video sources of 15s that were created with the goal of accurately representing real-life footage, while taking into account video coding and video quality testing challenges. We thoroughly describe the methodological approach of planning, creating, and post-processing the video sources as well as conduct evaluations of detail, dynamic range, and also assess the appropriateness for the contents to be used in video coding research. The videos and objective descriptors are made publicly available for future research under the link https://github.com/Telecommunication-Telemedia-Assessment/AVT-VQDB-UHD-2-HDR. Dominik Keller, Thomas Goebel, Valentin Siebenkees, Julius Prenzel, Alexander Raake |
QoMEX | 5 |
| 2024 | AVT-VQDB-UHD-1-HDR: An Open Video Quality Dataset for Quality Assessment of UHD-1 HDR VideosabstractHigh dynamic range (HDR) videos offer users a more realistic viewing experience owing to their ability to represent a wider and thus more natural range of brightness. This has resulted in an increase in HDR content streamed on different video streaming platforms. Hence, it becomes important to have a proper understanding of the perceived quality of HDR videos when encoded with modern video codecs such as H.265, AV1, and VVC that are either commonly used or will potentially be used by video streaming providers. With this objective, in this paper, we present a study that used both subjective and instrumental methods to assess the perceived quality of HDR videos. Firstly, a subjective test with 4K/UHD-1 HDR videos using the ACR-HR (Absolute Category Rating -Hidden Reference) method was conducted. The tests consisted of a total of 195 encoded videos from 5 source videos which all had a framerate of 60 fps. In this test, the 4K/UHD-1 HDR stimuli were encoded at four different resolutions, namely, 720p, 1080p, 1440p, and 2160p using bitrates ranging between 0.5 Mbps and 40 Mbps. The results of the subjective test have been analyzed to assess the impact of factors such as resolution, bitrate, video codec, and content on the perceived video quality. As automated quality assessment forms an important part of the encoding ecosystem of any video streaming platform to decide the optimal encoding settings, different full reference, bitstream, and hybrid instrumental models have been evaluated for their applicability for HDR video quality prediction. The database of source content, encoded videos, subjective and objective scores is made publicly available with this paper following an open-science approach, accessible at: https://github.com/Telecommunication-Telemedia-Assessment/AVT-VQDB-UHD-1-HDR. Rakesh Rao Ramachandra Rao, Benjamin Herb, Helmi-Aurora Takala, Mohamed Tarek Mohamed Ahmed, Alexander Raake |
QoMEX | 5 |
| 2024 | An exploratory study on the impact of varying levels of robot control on presence in robot-mediated communicationabstractTelepresence robots can enhance communication experiences by providing a sense of physical presence, embodiment and may evoke co-presence. In spite of that, telepresence robots have not made it fully to consumer markets. In this paper, we investigate how different levels of controlling a telepresence robot (teleoperation, shared control, and no control) influence presence. To this aim, we conducted a study (N=45) where participants were evenly distributed to one of the robot control conditions. The task involved navigating an unknown room and listening to stories told by a person co-located with the robot. We collected subjective impressions of presence using the temple presence inventory and performed a thematic content analysis on a post-experiment interview. Our results suggest nuances in perceived presence under different levels of robot control after performing a thematic content analysis. Copresence can be experienced during teleoperation and shared control, and teleoperation may evoke negative sentiments if it does not provide enough spatial information during navigation. However, our results did not point to significant differences in spatial or social presence. We consider that these findings encourage further discussions on how presence is perceived in robot-mediated communication. Stephanie Arevalo, Söhnke Benedikt Fischedick, Chenayo Diao, Kay Richter, Horst-Michael Groß, Alexander Raake |
RO-MAN | 6 |
| 2024 | Work-in-Progress: Older Adults' Experiences With an Augmented Reality Communication SystemabstractGiven the profound impact of staying socially connected on the well-being of older adults, this study explores the potential of augmented reality (AR) systems to enrich their social lives. A wearable AR communication system prototype was developed and tested in a user study involving N = 16 older adults from Germany. Participants wore an AR headset and engaged in a conversation task with a remote person represented by an avatar. Older adults’ experiences were assessed using think-aloud protocols, qualitative observations, posttest questionnaires, and semi-structured oral interviews. Preliminary findings indicate overall participant satisfaction, with minimal observed difficulties in headset usage and avatar-mediated interpersonal communication. The positive engagement during AR conversations highlights the system’s potential to provide positive communication experiences among older individuals. This work-in-progress paper introduces the developed system prototype and outlines the conducted user study. Further data analyses will provide deeper insights into older adults’ experiences with the system. The results will contribute to refining the prototype and offer valuable insights for the development of AR communication systems tailored to the needs and preferences of older adults. Veronika Mikhailova, Christian Kunert, Jakob Hartbrich, Tobias Schwandt, Christoph Gerhardt, Alexander Raake, Wolfgang Broll, Nicola Döring |
IMX | 6 |
| 2024 | Beyond Looks: A Study on Agent Movement and Audiovisual Spatial Coherence in Augmented RealityabstractThe appearance of virtual humans (avatars and agents) has been widely explored in immersive environments. However, virtual humans’ movements and associated sounds in real-world interactions, particularly in Augmented Reality (AR), are yet to be explored. In this paper, we investigate the influence of three distinct movement patterns (circle, side-to-side, and standing), two rendering styles (realistic and cartoon), and two types of audio (spatial audio and non-spatial audio) on emotional responses, social presence, appearance and behavior plausibility, audiovisual coherence, and auditory plausibility. To enable that, we conducted a study (N=36) where participants observed an agent reciting a short fictional story. Our results indicate an effect of the rendering style and the type of movement on the subjective perception of the agents behaving in an AR environment. Participants reported higher levels of excitement when they observed the realistic agent moving in a circle compared to the cartoon agent or the other two movement patterns. Moreover, we found an influence of agent’s movement pattern on social presence and higher appearance and behavior plausibility for the realistic rendering style. Regarding audiovisual spatial coherence, we found an influence of rendering style and type of audio only for the cartoon agent. Additionally, the spatial audio was perceived as more plausible than non-spatial audio. Our findings suggest that aligning realistic rendering styles with realistic auditory experiences may not be necessary for 1-1 listening experiences with moving sources. However, movement patterns of agents influence excitement and social presence in passive unidirectional communication scenarios. Stephanie Arevalo, Christian Kunert, Jakob Hartbrich, Christian Schneiderwind, Chenyao Diao, Christoph Gerhardt, Tatiana Surdu, Florian Weidner, Wolfgang Broll, Stephan Werner 0003, Alexander Raake |
VR | 11 |
| 2024 | Evaluating the Effect of Binaural Auralization on Audiovisual Plausibility and Communication Behavior in Virtual RealityabstractSpatial audio representations have been shown to positively impact user experience in traditional, non-immersive communication media. While spatial audio also contributes to presence in single-user immersive VR, its impact in virtual communication scenarios has not yet been fully understood. This work aims to further investigate which communication scenarios benefit from spatial audio representations. We present a study in which pairs of interlocutors undertake a collaborative task in an audiovisual Virtual Environment (VE) under different auralization and scene arrangement conditions. The novel task is designed to encourage simultaneous conversation and movement, with the aim of increasing the relevance of spatial hearing. Results are obtained through questionnaires measuring social presence and plausibility, as well as through conversational and behavioral analysis. Although participants are shown to favor binaural auralization over diotic audio in a direct active-listening comparison, no significant differences in social presence, plausibility, or communication behavior could be found. Our results suggest that spatial audio may not affect user experience in dyadic communication scenarios where spatial auditory information is not directly relevant to the considered task. Felix Immohr, Gareth Rendle, Anton Benjamin Lammert, Annika Neidhardt, Victoria Meyer Zur Heyde, Bernd Fröhlich 0001, Alexander Raake |
VR | 7 |
| 2024 | Immersive Study Analyzer: Collaborative Immersive Analysis of Recorded Social VR StudiesabstractVirtual Reality (VR) has become an important tool for conducting behavioral studies in realistic, reproducible environments. In this paper, we present ISA, an Immersive Study Analyzer system designed for the comprehensive analysis of social VR studies. For in-depth analysis of participant behavior, ISA records all user actions, speech, and the contextual environment of social VR studies. A key feature is the ability to review and analyze such immersive recordings collaboratively in VR, through support of behavioral coding and user-defined analysis queries for efficient identification of complex behavior. Respatialization of the recorded audio streams enables analysts to follow study participants' conversations in a natural and intuitive way. To support phases of close and loosely coupled collaboration, ISA allows joint and individual temporal navigation, and provides tools to facilitate collaboration among users at different temporal positions. An expert review confirms that ISA effectively supports collaborative immersive analysis, providing a novel and effective tool for nuanced understanding of user behavior in social VR studies. Anton Benjamin Lammert, Gareth Rendle, Felix Immohr, Annika Neidhardt, Karlheinz Brandenburg, Alexander Raake, Bernd Fröhlich 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Towards evaluation of immersion, visual comfort and exploration behaviour for non-stereoscopic and stereoscopic 360° videosabstractImmersion, visual comfort, and exploration behaviour are important aspects that affect the overall quality of experience for 360° videos. To analyze the benefits of stereoscopic and non-stereoscopic 360° videos in terms of these factors, we created a dataset and conducted a subjective study. The dataset consists of five different high-resolution $8 \mathrm{~K}$ omnidirectional videos as stereoscopic and non-stereoscopic variants. The videos have been recorded using a Kandao Obsidian Pro camera. For the comparison, we designed and performed a subjective test with 30 participants. Here, each subject watched both HEVC (libx265) encoded versions of the source video and rated the videos viewed regarding presence, visual comfort, and quality. The results indicate that with the test protocol followed, non-stereoscopic video viewing leads to slightly better presence, visual comfort, and quality ratings compared to the stereoscopic variants. Further, the stereoscopic 360° videos may suffer from visual artefacts potentially leading to lower video quality and further lower quality of experience results. The exploration behaviour was found to be very similar for both non-stereoscopic and stereoscopic video viewing. Overall, it can be concluded that there is a slight tendency for non-stereoscopic video viewing to be preferred over stereoscopic video viewing. The dataset is made publicly available with the paper and includes both variants of all source videos along with the subjective data, and behaviour data, following an open-science approach. Stephan Fremerey, Raja Faseeh Uz Zaman, Touseef Ashraf, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
ISM | 6 |
| 2023 | The Effect of Viewing Distances on 4K and 8K HDR Video Quality PerceptionabstractOngoing research in the field of capture, coding and display technology and human vision has explored the advantages of high resolution up to 8K (UHD-2) considering perceived quality. One of the crucial elements impacting users’ perception of video quality is the viewing distance. As a result, the presented study employs a subjective evaluation to investigate the perceptual benefits offered by 8K or upscaled 4K in comparison to the native 4K (UHD-1) resolution in the context of HDR videos. The subjective test uses 7 distinct viewing distances, ranging from 0.5H to 3H, with H representing the display height. The findings of the study reveal a consistent trend: the increased video quality of 8K HDR against 4K HDR content decreases with distance, on average. While there are bigger improvements for close distances, beyond 2H the quality difference was very little or zero, depending on content. In general, the degree of enhancement is contingent on the spatial complexity of the content. Additionally, it is found that, on average, subjects prefer to sit at a distance of 2.07H. No significant difference in the preferred viewing distance was found when asked before and after the study. Dominik Keller, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
ISM | 4 |
| 2023 | Adaptation of Bitstream-based Video Quality Models for Image Quality AssessmentabstractIn recent years, video-codec-based image codecs, such as e.g. HEF, AVF, etc., have been increasingly used to compress images. Hence, there is a potential to use video quality prediction models for the evaluation of image quality. Bitstream-based models show promising results for video quality prediction, therefore, we investigate the applicability of such models for the case of image quality in this paper. For this purpose, we selected ITU-T Rec. P.1204.3 and its Mode 0 variant also known as $AVQBits|M3$ and $AVQBits|M0$ respectively for the evaluation, because they are computationally less complex and do not need a reference image. These models are evaluated using a publicly available dataset consisting of a total of 371 images of resolutions between $144\times 144$ pixels to $2160\times 2160$ pixels with subjective annotations. The results show that both the considered models perform well on the used dataset with a Pearson correlation of 0.958 and Root Mean Square Error (RMSE) of 0.319 (on a 1 to 5 Absolute Category Rating (ACR) scale) for the $AVQBits|M3$ model and a Pearson correlation of 0.942 and RMSE of 0.377 for the $AVQBits|M0$ model. Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
ISM | 3 |
| 2023 | AVT-VQDB-UHD-1-Appeal: A UHD-1/4K Open Dataset for Video Quality and Appeal Assessment Using Modern Video CodecsabstractA number of factors play an important role in the perception of video quality for streaming and other services, key among them being encoding-related degradations. Hence, newer codecs are developed with the goal of optimizing video quality for a given encoding setting. Here, subjective studies are an efficient method to evaluate the performance of such newer codecs. Furthermore, contextual factors impact the perception of video quality, e.g., the appeal of the content itself. To this end, this paper presents a subjective study targeting both quality and appeal assessment of videos. For this purpose, a subjective study consisting of three different parts is conducted. Firstly, participants were asked to rate the appeal of the uncompressed UHD-1/4K source content with a duration of 8 - 10s each. Following this, the video quality of these source videos individually encoded with either the HEVC/H.265, AV1, or VVC/H.266 video codec was rated. A wide range of encoding conditions in terms of resolution (360p to 2160p) and bitrate (100kbps to 15mbps) is used to encode the videos, so as to enable the applicability of the data to real-world settings. In the last part, subjects are again asked to rate the appeal of the uncompressed source content. The results are analyzed to assess the impact of different encoding conditions on perceived video quality. In addition, the impact of appeal on video quality and vice-versa is also investigated. Furthermore, an objective quality assessment with different state-of-the-art full-reference, bitstream-based, and hybrid models including the newer codecs AV1 and VVC is presented. The subjective dataset including test design, subjective results, sources, and encoded audiovisual contents are made publicly available following an open science approach. Rakesh Rao Ramachandra Rao, Steve Goering, Bassem Elmeligy, Alexander Raake |
MMSP | 4 |
| 2023 | Eye and Face Tracking in VR: Avatar Embodiment and Enfacement with Realistic and Cartoon AvatarsabstractPrevious studies have explored the perception of various types of embodied avatars in immersive environments. However, the impact of eye and face tracking with personalized avatars is yet to be explored. In this paper, we investigate the impact of eye and face tracking on embodiment, enfacement, and the uncanny valley with four types of avatars using a VR-based mirroring task. We conducted a study (N=12) and created self-avatars with two rendering styles: a cartoon avatar (created in an avatar generator using a picture of the user’s face) and a photorealistic scanned avatar (created using a 3D scanner), each with and without eye and face tracking and respective adaptation of the mirror image. Our results indicate that adding eye and face tracking can be beneficial for certain enfacement scales (belonged), and we confirm that compared to a cartoon avatar, a scanned realistic avatar results in higher body ownership and increased enfacement (own face, belonging, mirror) — regardless of eye and face tracking. We critically discuss our experiences and outline the limitations of the applied hardware and software with respect to the provided level of control and the applicability for complex tasks such as displaying emotions. We synthesize these findings into a discussion about potential improvements for facial animation in VR and highlight the need for a better level of control, the integration of additional sensing and processing technologies, and an objective metric for comparing facial animation systems. Jakob Hartbrich, Florian Weidner, Christian Kunert, Alexander Raake, Wolfgang Broll, Stephanie Arevalo |
MUM | 4 |
| 2023 | A Virtual Gardening Experience: Evaluating the effect of haptic feedback on spatial presence, perceptual realism, mental immersion, and user experienceabstractVirtual nature settings have demonstrated to provide benefits to mental well-being. However, most studies have focused on providing only audiovisual stimuli. We aim to evaluate the use of haptic feedback to simulate touching elements in nature-inspired settings. In this paper, we designed a VR gardening environment to investigate the impact of haptic feedback on spatial presence, perceptual realism, mental immersion, user experience, and task performance while interacting with gardening objects in a study (N=18, 9 female and 9 male). Our results suggest that haptic feedback can increase spatial presence and point to gender differences, i.e., female participants reported higher scores in spatial presence and perceptual realism, in the chosen VR experience. Although our main goal was to evaluate the role of haptics in a virtual garden, our findings highlight the importance of investigating and identifying factors that could lead to gender differences in VR experiences. Qasim Saboor, Hamd Mehfooz-Khan, Alexander Raake, Stephanie Arevalo |
MUM | 3 |
| 2023 | Automatic Audiovisual Asynchrony Measurement for Quality Assessment of VideoconferencingabstractAudiovisual asynchrony is a significant factor im-pacting the Quality of Experience (QoE), especially for interactive communication like video conferencing. In this paper, we propose a client-side approach to predict the delay between an audio and a video signal, using only the media signals from both streams. Features are extracted from the video and audio stream, respectively, and analyzed using a cross-correlation approach to determine the actual delay. Our approach predicts the delay with an accuracy of over 80% in a time frame of ±1s. We further highlight the potential drawbacks of using a cross-correlation-based analysis and propose different solutions for practical implementations of a delay-based QoE metric. Florian Braun, Rakesh Rao Ramachandra Rao, Werner Robitza, Alexander Raake |
QoMEX | 4 |
| 2023 | Revisiting Videoconferencing QoE: Impact of Network Delay and Resolution as Factors for Social Cue PerceptibilityabstractPrevious research from well before the Covid-19 pandemic had indicated little effect of delay on integral quality but a measurable one on user behavior, and a significant effect of resolution on quality but not on behavior in a two-party communication scenario. In this paper, we re-investigate the topic, after the times of the Covid-19 pandemic and its frequent and widespread videoconferencing usage. To this aim, we conducted a subjective test involving 23 pairs of participants, employing the Celebrity Name Guessing task. The focus was on impairments that may affect social (resolution) and communication cues (de-lay). Subjective data in the form of overall conversational quality and task performance satisfaction as well as objective data in the form of task correctness, user motion, and facial expressions were collected in the test. The analysis of the subjective data indicates that perceived conversational quality and performance satisfaction were mainly affected by video resolution, while delay (up to 1000 ms) had no significant impact. Furthermore, the analysis of the objective data shows that there is no impact of resolution and delay on user performance and behavior, in contrast to earlier findings. Chenyao Diao, Luljeta Sinani, Rakesh Rao Ramachandra Rao, Alexander Raake |
QoMEX | 4 |
| 2023 | DNN-based Photography Rule Prediction using Photo TagsabstractInstagram and Flickr are just two examples of photo-sharing platforms which are currently used to upload thousands of images on a daily basis. One important aspect in such social media contexts is to know whether an image is of high appeal or not. In particular, to understand the composition of a photo and to improve reading flow, several photo rules have been established. In this paper, we focus on eight selected photo rules. To automatically predict whether an image follows one of these rules or not, we train 13 deep neural networks in a transfer-learning setup and compare their prediction performance. As a dataset, we use photos downloaded from Flickr with specifically selected image tags, which reflect the eight photo rules. There-fore, our dataset does not need additional human annotations. ResNet50 has the best prediction performance, however, there are images that follow several rules, which must be addressed in follow-up work. The code and the data (image URLs) are made publicly available for reproducibility. Steve Goering, Rasmus Merten, Alexander Raake |
QoMEX | 3 |
| 2023 | Appeal and quality assessment for AI-generated imagesabstractRecently AI-generated images gained in popularity. A critical aspect of AI-generated images using, e.g., DALL-E-2 or Midjourney, is that they may look artificial, be of low quality, or have a low appeal in contrast to real images, depending on the text prompt and AI generator. For this reason, we evaluate the quality and appeal of AI-generated images using a crowdsourcing test as an extension of our recently published AVT-AI-Image-Dataset. This dataset consists of a total of 135 images generated with five different AI-text-to-image generators. Based on the collected subjective ratings in the crowdsourcing test, we evaluate the different used AI generators in terms of image quality and appeal of the AI-generated images. We also link image quality and image appeal also with SoA objective models. The extension will be made publicly available for reproducibility. Steve Goering, Rakesh Rao Ramachandra Rao, Rasmus Merten, Alexander Raake |
QoMEX | 4 |
| 2023 | Power Reduction Opportunities on End-User Devices in Quality-Steady Video StreamingabstractThis paper uses a crowdsourced dataset of online video streaming sessions to investigate opportunities to reduce the power consumption while considering QoE. For this, we base our work on prior studies which model both the end-user's QoE and the end-user device's power consumption with the help of high-level video features such as the bitrate, the frame rate, and the resolution. On top of existing research, which focused on reducing the power consumption at the same QoE optimizing video parameters, we investigate potential power savings by other means such as using a different playback device, a different codec, or a predefined maximum quality level. We find that based on the power consumption of the streaming sessions from the crowdsourcing dataset, devices could save more than 55% of power if all participants adhere to low-power settings. Christian Herglotz, Werner Robitza, Alexander Raake, Tobias Hoßfeld, André Kaup |
QoMEX | 3 |
| 2023 | Influence of Viewing Distances on 8K HDR Video Quality PerceptionabstractThe benefits of high resolutions in displays, such as 8K (UHD-2), have been the subject of ongoing research in the field of display technology and human perception in recent years. Out of several factors influencing users' perception of video quality, viewing distance is one of the key aspects. Hence, this study uses a subjective test to investigate the perceptual advantages of 8K over 4K (UHD-1) resolution for HDR videos at 7 different viewing distances, ranging from 0.5 H to 2 H. The results indicate that, on average, for HDR content the 8K resolution can improve the video quality at all tested distances. Our study shows that although the 8K resolution is slightly better than 4K at close distances, the extent of these benefits is highly dependent on factors such as the pixel-related complexity of the content and the visual acuity of the viewers. Dominik Keller, Felix von Hagen, Julius Prenzel, Kay Strama, Rakesh Rao Ramachandra Rao, Alexander Raake |
QoMEX | 6 |
| 2023 | PNATS-UHD-1-Long: An Open Video Quality Dataset for Long Sequences for HTTP-based Adaptive Streaming QoE AssessmentabstractThe P.NATS Phase 2 competition in ITU-T Study Group 12 resulted in both the ITU-T Rec. P.1204 series of recommendations, and also a large dataset for HTTP-based adaptive streaming QoE assessment that is now made openly available as part of this paper. The presented dataset consists of 3 subjective databases targeting overall quality assessment of a typical HTTP-based Adaptive Streaming session consisting of degradations such as quality switching, initial loading delay, and stalling events using audiovisual content ranging between 2 and 5 minutes. In addition to this, subject bias and consistency in quality assessment of such longer-duration audiovisual contents with multiple degradations are investigated using a subject behaviour model. As part of this paper, the overall test design, subjective test results, sources, encoded audiovisual contents, and a set of analysis plots are made publicly available for further research. Rakesh Rao Ramachandra Rao, Silvio Borer, David Lindero, Steve Goering, Alexander Raake |
QoMEX | 5 |
| 2023 | Saliency of Omnidirectional Videos with Different Audio Presentations: Analyses and DatasetabstractThere is an increased interest in understanding users' behavior when exploring omnidirectional (360°) videos, especially in the presence of spatial audio. Several studies demonstrate the effect of no, mono, or spatial audio on visual saliency. However, no studies investigate the influence of higher-order (i.e., 4t h- order) Ambisonics on subjective exploration in virtual reality settings. In this work, a between-subjects test design is employed to collect users' exploration data of 360° videos in a free-form viewing scenario using the Varjo XR-3 Head Mounted Display, in the presence of no, mono, and 4th-order Ambisonics audio. Saliency information was captured as head-saliency in terms of the center of a viewport at 50 Hz. For each item, subjects were asked to describe the scene with a short free-verbalization task. Moreover, cybersickness was assessed using the simulator sickness questionnaire at the beginning and at the end of the test. The head-saliency results over time show that with the presence of higher-order Ambisonics audio, subjects concentrate more on the directions sound is coming from. No influence of audio scenario on cybersickness scores was observed. From the analysis of the verbal scene descriptions, it was found that users were attentive to the omnidirectional video, but only for the ‘no audio’ scenario provided minute and insignificant details of the scene objects. The audiovisual saliency dataset is made available following the open science approach already used for the audiovisual scene recordings we previously published. The data is sought to enable training of visual and audiovisual saliency prediction models for interactive experiences. Ashutosh Singla, Thomas Robotham, Abhinav Bhattacharya, William Menz, Emanuël A. P. Habets, Alexander Raake |
QoMEX | 6 |
| 2023 | Proof-of-Concept Study to Evaluate the Impact of Spatial Audio on Social Presence and User Behavior in Multi-Modal VR CommunicationabstractThis paper presents a proof-of-concept study conducted to analyze the effect of simple diotic vs. spatial, position-dynamic binaural synthesis on social presence in VR, in comparison with face-to-face communication in the real world, for a sample two-party scenario. A conversational task with shared visual reference was realized. The collected data includes questionnaires for direct assessment, tracking data, and audio and video recordings of the individual participants’ sessions for indirect evaluation. While tendencies for improvements with binaural over diotic presentation can be observed, no significant difference in social presence was found for the considered scenario. The gestural analysis revealed that participants used the same amount and type of gestures in face-to-face as in VR, highlighting the importance of non-verbal behavior in communication. As part of the research, an end-to-end framework for conducting communication studies and analysis has been developed. Felix Immohr, Gareth Rendle, Annika Neidhardt, Steve Goering, Rakesh Rao Ramachandra Rao, Stephanie Arevalo, Bernd Fröhlich 0001, Alexander Raake |
IMX | 8 |
| 2023 | Influence of Multi-Modal Interactive Formats on Subjective Audio Quality and Exploration BehaviorabstractThis study uses a mixed between- and within-subjects test design to evaluate the influence of interactive formats on the quality of binaurally rendered 360° spatial audio content. Focusing on ecological validity using real-world recordings of 60 s duration, three independent groups of subjects () were exposed to three formats: audio only (A), audio with 2D visuals (A2DV), and audio with head-mounted display (AHMD) visuals. Within each interactive format, two sessions were conducted to evaluate degraded audio conditions: bit-rate and Ambisonics order. Our results show a statistically significant effect (p < .05) of format only on spatial audio quality ratings for Ambisonics order. Exploration data analysis shows that format A yields little variability in exploration, while formats A2DV and AHMD yield broader viewing distribution of 360° content. The results imply audio quality factors can be optimized depending on the interactive format. Thomas Robotham, Ashutosh Singla, Alexander Raake, Olli Rummukainen, Emanuël A. P. Habets |
IMX | 3 |
| 2023 | A Systematic Review on the Visualization of Avatars and Agents in AR & VR displayed using Head-Mounted DisplaysabstractAugmented Reality (AR) and Virtual Reality (VR) are pushing from the labs towards consumers, especially with social applications. These applications require visual representations of humans and intelligent entities. However, displaying and animating photo-realistic models comes with a high technical cost while low-fidelity representations may evoke eeriness and overall could degrade an experience. Thus, it is important to carefully select what kind of avatar to display. This article investigates the effects of rendering style and visible body parts in AR and VR by adopting a systematic literature review. We analyzed 72 papers that compare various avatar representations. Our analysis includes an outline of the research published between 2015 and 2022 on the topic of avatars and agents in AR and VR displayed using head-mounted displays, covering aspects like visible body parts (e.g., hands only, hands and head, full-body) and rendering style (e.g., abstract, cartoon, realistic); an overview of collected objective and subjective measures (e.g., task performance, presence, user experience, body ownership); and a classification of tasks where avatars and agents were used into task domains (physical activity, hand interaction, communication, game-like scenarios, and education/training). We discuss and synthesize our results within the context of today's AR and VR ecosystem, provide guidelines for practitioners, and finally identify and present promising research opportunities to encourage future research of avatars and agents in AR/VR environments. Florian Weidner, Gerd Boettcher, Stephanie Arevalo, Chenyao Diao, Luljeta Sinani, Christian Kunert, Christoph Gerhardt, Wolfgang Broll, Alexander Raake |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2022 | Modeling of Energy Consumption and Streaming Video QoE using a Crowdsourcing DatasetabstractIn the past decade, we have witnessed an enormous growth in the demand for online video services. Recent studies estimate that nowadays, more than 1% of the global greenhouse gas emissions can be attributed to the production and use of devices performing online video tasks. As such, research on the true power consumption of devices and their energy efficiency during video streaming is highly important for a sustainable use of this technology. At the same time, over-the-top providers strive to offer high-quality streaming experiences to satisfy user expectations. Here, energy consumption and QoE partly depend on the same system parameters. Hence, a joint view is needed for their evaluation. In this paper, we perform a first analysis of both end-user power efficiency and Quality of Experience of a video streaming service. We take a crowdsourced dataset comprising 447,000 streaming events from YouTube and estimate both the power consumption and perceived quality. The power consumption is modeled based on previous work which we extended towards predicting the power usage of different devices and codecs. The user-perceived QoE is estimated using a standardized model. Our results indicate that an intelligent choice of streaming parameters can optimize both the QoE and the power efficiency of the end user device. Further, the paper discusses limitations of the approach and identifies directions for future research. Christian Herglotz, Werner Robitza, Matthias Kränzler, André Kaup, Alexander Raake |
QoMEX | 5 |
| 2022 | Audiovisual Database with 360° Video and Higher-Order Ambisonics Audio for Perception, Cognition, Behavior, and QoE Evaluation ResearchabstractResearch into multi-modal perception, human cog-nition, behavior, and attention can benefit from high-fidelity content that may recreate real-life-like scenes when rendered on head-mounted displays. Moreover, aspects of audiovisual perception, cognitive processes, and behavior may complement questionnaire-based Quality of Experience (QoE) evaluation of interactive virtual environments. Currently, there is a lack of high-quality open-source audiovisual databases that can be used to evaluate such aspects or systems capable of reproducing high-quality content. With this paper, we provide a publicly available audiovisual database consisting of twelve scenes capturing real-life nature and urban environments with a video resolution of 7680×3840 at 60 frames-per-second and with 4th-order Ambison-ics audio. These 360° video sequences, with an average duration of 60 seconds, represent real-life settings for systematically evaluating various dimensions of uni-/multi-modal perception, cognition, behavior, and QoE. The paper provides details of the scene requirements, recording approach, and scene descriptions. The database provides high-quality reference material with a balanced focus on auditory and visual sensory information. The database will be continuously updated with additional scenes and further metadata such as human ratings and saliency information. Thomas Robotham, Ashutosh Singla, Olli Rummukainen, Alexander Raake, Emanuël A. P. Habets |
QoMEX | 4 |
| 2022 | Subjective Evaluation of Visual Quality and Simulator Sickness of Short 360$^\circ$ Videos: ITU-T Rec. P.919abstractRecently an impressive development in immersive technologies, such as Augmented Reality (AR), Virtual Reality (VR) and 360${^\circ }$video, has been witnessed. However, methods for quality assessment have not been keeping up. This paper studies quality assessment of 360${^\circ }$video from the cross-lab tests (involving ten laboratories and more than 300 participants) carried out by the Immersive Media Group (IMG) of the Video Quality Experts Group (VQEG). These tests were addressed to assess and validate subjective evaluation methodologies for 360${^\circ }$video. Audiovisual quality, simulator sickness symptoms, and exploration behavior were evaluated with short (from 10 seconds to 30 seconds) 360${^\circ }$sequences. The following factors’ influences were also analyzed: assessment methodology, sequence duration, Head-Mounted Display (HMD) device, uniform and non-uniform coding degradations, and simulator sickness assessment methods. The obtained results have demonstrated the validity of Absolute Category Rating (ACR) and Degradation Category Rating (DCR) for subjective tests with 360${^\circ }$videos, the possibility of using 10-second videos (with or without audio) when addressing quality evaluation of coding artifacts, as well as any commercial HMD (satisfying minimum requirements). Also, more efficient methods than the long Simulator Sickness Questionnaire (SSQ) have been proposed to evaluate related symptoms with 360${^\circ }$videos. These results have been instrumental for the development of the ITU-T Recommendation P.919. Finally, the annotated dataset from the tests is made publicly available for the research community. Jesús Gutiérrez 0001, Pablo Pérez 0001, Marta Orduna, Ashutosh Singla, Carlos Cortés 0001, Pramit Mazumdar, Irene Viola 0001, Kjell Brunnström, Federica Battisti, Natalia Cieplinska, Dawid Juszka, Lucjan Janowski, Mikolaj Leszczuk, Anthony Adeyemi-Ejeye, Yaosi Hu, Zhenzhong Chen 0001, Glenn Van Wallendael, Peter Lambert, César Díaz, John Hedlund, Omar Hamsis, Stephan Fremerey, Frank Hofmeyer, Alexander Raake, Pablo César, Marco Carli, Narciso García |
IEEE Trans. Multim. | 24 |
| 2021 | Rule of Thirds and Simplicity for Image Aesthetics using Deep Neural NetworksabstractConsidering the increasing amount of photos being uploaded to sharing platforms, a proper evaluation of photo appeal or aesthetics is required. For appealing images several "rules of thumb" have been established, e.g., the rule of thirds and simplicity. We handle rule of thirds and simplicity as binary classification problems with a deep learning based image processing pipeline. Our pipeline uses a pre-processing step, a pre-trained baseline deep neural network (DNN) and post-processing. For each of the rules, we re-train 17 pre-trained DNN models using transfer learning. Our results for publicly available datasets show that the ResNet152 DNN is best for rule of thirds prediction and DenseNet121 is best for simplicity with an accuracy of around 0.84 and 0.94 respectively. In addition to the datasets for both classifications, five experts annotated another dataset with ≈ 1100 images and we evaluate the best performing models. Results show that the best performing models have an accuracy of 0.67 for rule of thirds and 0.79 for image simplicity. Both accuracy results are within the range of pairwise accuracy of expert annotators. However, it further indicates that there is a high subjective influence for both of the considered rules. Steve Goering, Alexander Raake |
MMSP | 2 |
| 2021 | AVrate Voyager: an open source online testing platformabstractSubjective testing is an integral part of many research fields considering, e.g., human perception. For this purpose, lab tests are a popular approach to gather ratings for subjective evaluations. However, not in all cases controlled lab tests can be performed, either in cases where no labs are existing, accessible or it may be disallowed to use them. For this reason, online tests, e.g., using crowdsourcing are supposed to be an alternative approach for traditional lab tests. We describe in the following paper a framework to implement such online tests for audio, video, and image-related evaluations or questionnaires. Our framework AVrate Voyager builds upon previously developed frameworks for lab tests including the experience with them. AVrate Voyager uses scalable web technologies to implement a test framework, this ensures that it will be running reliably. In addition, we added strategies for pre-caching to avoid additional influence for play-out, e.g. in the case of video testing. We analyze several conducted tests using the new framework and describe the required steps to modify the provided tool in detail. Steve Goering, Rakesh Rao Ramachandra Rao, Stephan Fremerey, Alexander Raake |
MMSP | 4 |
| 2021 | Groovability: Using Groove as a Novel Measure for Audio QoE with the Example of SmartphonesabstractGroove in music is a fundamental part of why humans entrain to it and enjoy it. Smartphones have become an important medium to listen to music. Especially when being with others, loudspeaker playback may be the method of choice. However, due to the physical limits of acoustics, for loudspeaker playback, smartphones are equipped with sub-optimal audio capabilities. Therefore, it is desirable to measure Quality of Experience (QoE) of music played on smartphones. While audio playback is often assessed in terms of sound quality, the aim of this work is to address QoE in terms of the meaning or effect that the audio has on the listener. A key component for the meaning of popular music is groove. Hence, in this paper, we study “groovability”, that is, the ability of a piece of audio technology to convey groove. To instantiate our novel audio QoE assessment method, we apply it to music played by 8 different smartphones. For this purpose, looped 4-bar loudness-aligned recordings from 24 music pieces of different intrinsic groove were played back on the different smartphones. Our test method uses a multi-stimulus comparison with synchronized playback capability. A total of 62 subjects evaluated groovability using two stimulus subsets. It was found that the proposed methodology is highly effective to distinguish between the groovability provided by the considered phones. In addition, a reduced-reference model is proposed to predict groovability, using a set of both acoustics-and music-groove related features. In our formal validation on unknown data, the model is shown to provide good prediction performance with a Pearson correlation of greater than 0.90. Dominik Keller, Markus Vaalgamaa, Erkki Paajanen, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
QoMEX | 6 |
| 2021 | Towards High Resolution Video Quality Assessment in the CrowdabstractAssessing high resolution video quality is usually performed using controlled, defined, and standardized lab tests. This method of acquiring human ratings in a lab environment is time-consuming and may also not reflect the typical viewing conditions. To overcome these disadvantages, crowd testing paradigms have been used for assessing video quality in general. Crowdsourcing-based tests enable a more diverse set of participants and also use a realistic hardware setup and viewing environment of typical users. However, obtaining valid ratings for high-resolution video quality poses several problems. Example issues are that streaming of such high-bandwidth content may not be feasible for some users, or that crowd participants lack an appropriate, high-resolution display device. In this paper, we propose a method to overcome such problems and conduct a crowd test using for higher resolution content by using a 540 p cutout from the center of the original 2160p video. To this aim, we use the videos from Test#1 of the publicly available dataset AVT-VQDB-UHD-1, which contains videos up to a resolution of UHD-1. The quality-labels available from that lab test allow us to compare the results with the crowd test presented in this paper. It is shown that there is a Pearson correlation of 0.96 between the lab and crowd tests and hence such crowd tests can reliably be used for video assessment of higher resolution content. The overall implementation of the crowd test framework and the results are made publicly available for further research and reproducibility1. Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
QoMEX | 3 |
| 2021 | Impact of Spatial and Temporal Information on Video Quality and CompressibilityabstractSpatial Information (SI) and Temporal Information (TI) are frequently-used metrics to classify the spatiotemporal complexity of video content. However, they are mostly used on original video sources, and their impact on actual encoding efficiency is not known. In this paper, we propose a method to determine the compressibility of video sources, that is, how good video quality can be under a given bitrate constraint. We show how various aggregations of SI and TI correlate with compressibility scores obtained from a public dataset of H.264/HEVCN P9 content. We observe that the minimum TI value as well as an existing criticality metric from the literature are good indicators for compressibility, as judged by subjective ratings as well as VMAF and P.1204.3 objective scores. Werner Robitza, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
QoMEX | 4 |
| 2021 | Assessment of the Simulator Sickness Questionnaire for Omnidirectional VideosabstractVirtual Reality/360° videos provide an immersive experience to users. Besides this, 360° videos may lead to an undesirable effect when consumed with Head-Mounted Displays (HMDs), referred to as simulator sickness/cybersickness. The Simulator Sickness Questionnaire (SSQ) is the most widely used questionnaire for the assessment of simulator sickness. Since the SSQ with its 16 questions was not designed for 360° video related studies, our research hypothesis in this paper was that it may be simplified to enable more efficient testing for 360° video. Hence, we evaluate the SSQ to reduce the number of questions asked from subjects, based on six different previously conducted studies. We derive the reduced set of questions from the SSQ using Principal Component Analysis (PCA) for each test. Pearson Correlation is analysed to compare the relation of all obtained reduced questionnaires as well as two further variants of SSQ reported in the literature, namely Virtual Reality Sickness Questionnaire (VRSQ) and Cybersickness Questionnaire (CSQ). Our analysis suggests that a reduced questionnaire with 9 out of 16 questions yields the best agreement with the initial SSQ, with less than 44% of the initial questions. Exploratory Factor Analysis (EFA) shows that the nine symptom-related attributes determined as relevant by PCA also appear to be sufficient to represent the three dimensions resulting from EFA, namely, Uneasiness, Visual Discomfort and Loss of Balance. The simplified version of the SSQ has the potential to be more efficiently used than the initial SSQ for 360° video by focusing on the questions that are most relevant for individuals, shortening the required testing time. Ashutosh Singla, Steve Goering, Dominik Keller, Rakesh Rao Ramachandra Rao, Stephan Fremerey, Alexander Raake |
VR | 6 |
| 2020 | Between the Frames - Evaluation of Various Motion Interpolation Algorithms to Improve 360° Video QualityabstractWith the increasing availability of 360° video content, it becomes important to provide smoothly playing videos of high quality for end users. For this reason, we compare the influence of different Motion Interpolation (MI) algorithms on 360° video quality. After conducting a pre-test with 12 video experts in [3], we found that MI is a useful tool to increase the QoE (Quality of Experience) of omnidirectional videos. As a result of the pretest, we selected three suitable MI algorithms, namely ffmpeg Motion Compensated Interpolation (MCI), Butterflow and Super-SloMo. Subsequently, we interpolated 15 entertaining and realworld omnidirectional videos with a duration of 20 seconds from 30 fps (original framerate) to 90 fps, which is the native refresh rate of the HMD used, the HTC Vive Pro. To assess QoE, we conducted two subjective tests with 24 and 27 participants. In the first test we used a Modified Paired Comparison (M-PC) method, and in the second test the Absolute Category Rating (ACR) approach. In the M-PC test, 45 stimuli were used and in the ACR test 60. Results show that for most of the 360° videos, the interpolated versions obtained significantly higher quality scores than the lower-framerate source videos, validating our hypothesis that motion interpolation can improve the overall video quality for 360° video. As expected, it was observed that the relative comparisons in the M-PC test result in larger differences in terms of quality. Generally, the ACR method lead to similar results, while reflecting a more realistic viewing situation. In addition, we compared the different MI algorithms and can conclude that with sufficient available computing power Super-SloMo should be preferred for interpolation of omnidirectional videos, while MCI also shows a good performance. Stephan Fremerey, Frank Hofmeyer, Steve Goering, Dominik Keller, Alexander Raake |
ISM | 5 |
| 2020 | Subjective Test Dataset and Meta-data-based Models for 360° Streaming Video QualityabstractDuring the last years, the number of 360° videos available for streaming has rapidly increased, leading to the need for 360° streaming video quality assessment. In this paper, we report and publish results of three subjective 360° video quality tests, with conditions used to reflect real-world bitrates and resolutions including 4K, 6K and 8K, resulting in 64 stimuli each for the first two tests and 63 for the third. As playout device we used the HTC Vive for the first and HTC Vive Pro for the remaining two tests. Video-quality ratings were collected using the 5-point Absolute Category Rating scale. The 360° dataset provided with the paper contains the links of the used source videos, the raw subjective scores, video-related meta-data, head rotation data and Simulator Sickness Questionnaire results per stimulus and per subject to enable reproducibility of the provided results. Moreover, we use our dataset to compare the performance of state-of-the-art full-reference quality metrics such as VMAF, PSNR, SSIM, ADM2, WS-PSNR and WS-SSIM. Out of all metrics, VMAF was found to show the highest correlation with the subjective scores. Further, we evaluated a center-cropped version of VMAF ("VMAF-cc") that showed to provide a similar performance as the full VMAF. In addition to the dataset and the objective metric evaluation, we propose two new video-quality prediction models, a bitstream meta-data-based model and a hybrid no-reference model using bitrate, resolution and pixel information of the video as input. The new lightweight models provide similar performance as the full-reference models while enabling fast calculations. Stephan Fremerey, Steve Goering, Rakesh Rao Ramachandra Rao, Rachel Huang, Alexander Raake |
MMSP | 5 |
| 2020 | Automated Genre Classification for Gaming VideosabstractBesides classical videos, videos of gaming matches, entire tournaments or individual sessions are streamed and viewed all over the world. The increased popularity of Twitch or YoutubeGaming shows the importance of additional research on gaming videos. One important pre-condition for live or offline encoding of gaming videos is the knowledge of game-specific properties. Knowing or automatically predicting the genre of a gaming video enables a more advanced and optimized encoding pipeline for streaming providers, especially because gaming videos of different genres vary a lot from classical 2D video, e.g., considering the CGI content, textures or camera motion. We describe several computer-vision based features that are optimized for speed and motivated by characteristics of popular games, to automatically predict the genre of a gaming video. Our prediction system uses random forest and gradient boosting trees as underlying machine-learning techniques, combined with feature selection. For the evaluation of our approach we use a dataset that was built as part of this work and consists of recorded gaming sessions for 6 genres from Twitch. In total 351 different videos are considered. We show that our prediction approach shows a good performance in terms of f1-score. Besides the evaluation of different machine-learning approaches, we additionally investigate the influence of the hyper-parameters for the algorithms. Steve Goering, Robert Steger, Rakesh Rao Ramachandra Rao, Alexander Raake |
MMSP | 4 |
| 2020 | A Large-scale Evaluation of the bitstream-based video-quality model ITU-T P.1204.3 on Gaming ContentabstractThe streaming of gaming content, both passive and interactive, has increased manifolds in recent years. Gaming contents bring with them some peculiarities which are normally not seen in traditional 2D videos, such as the artificial and synthetic nature of contents or repetition of objects in a game. In addition, the perception of gaming content by the user is different from that of traditional 2D videos due to its pecularities and also the fact that users may not often watch such content. Hence, it becomes imperative to evaluate whether the existing video quality models usually designed for traditional 2D videos are applicable to gaming content. In this paper, we evaluate the applicability of the recently standardized bitstream-based video-quality model ITU-T P.1204.3 on gaming content. To analyze the performance of this model, we used 4 different gaming datasets (3 publicly available + 1 internal) not previously used for model training, and compared it with the existing state-of-the-art models. We found that the ITU P.1204.3 model out of the box performs well on these unseen datasets, with an RMSE ranging between 0.38 - 0.45 on the 5-point absolute category rating and Pearson Correlation between 0.85 - 0.93 across all the 4 databases. We further propose a full-HD variant of the P.1204.3 model, since the original model is trained and validated which targets a resolution of 4K/UHD-1. A 50:50 split across all databases is used to train and validate this variant so as to make sure that the proposed model is applicable to various conditions. Rakesh Rao Ramachandra Rao, Steve Goering, Robert Steger, Saman Zad Tootaghaj, Nabajeet Barman, Stephan Fremerey, Sebastian Möller 0001, Alexander Raake |
MMSP | 8 |
| 2020 | DEMI: Deep Video Quality Estimation Model using Perceptual Video Quality DimensionsabstractExisting works in the field of quality assessment focus separately on gaming and non-gaming content. Along with the traditional modeling approaches, deep learning based approaches have been used to develop quality models, due to their high prediction accuracy. In this paper, we present a deep learning based quality estimation model considering both gaming and non-gaming videos. The model is developed in three phases. First, a convolutional neural network (CNN) is trained based on an objective metric which allows the CNN to learn video artifacts such as blurriness and blockiness. Next, the model is fine-tuned based on a small image quality dataset using blockiness and blurriness ratings. Finally, a Random Forest is used to pool frame-level predictions and temporal information of videos in order to predict the overall video quality. The light-weight, low complexity nature of the model makes it suitable for real-time applications considering both gaming and non-gaming content while achieving similar performance to existing state-of-the-art model NDNetGaming. The model implementation for testing is available on GitHub1. Saman Zad Tootaghaj, Nabajeet Barman, Rakesh Rao Ramachandra Rao, Steve Goering, Maria G. Martini, Alexander Raake, Sebastian Möller 0001 |
MMSP | 6 |
| 2020 | Comparing fixed and variable segment durations for adaptive video streaming: a holistic analysisabstractHTTP Adaptive Streaming (HAS) is the de-facto standard for video delivery over the Internet. It enables dynamic adaptation of video quality by splitting a video into small segments and providing multiple quality levels per segment. So far, HAS services typically utilize a fixed segment duration. This reduces the encoding and streaming variability and thus allows a faster encoding of the video content and a reduced prediction complexity for adaptive bit rate algorithms. Due to the content-agnostic placement of I-frames at the beginning of each segment, additional encoding overhead is introduced. In order to mitigate this overhead, variable segment durations, which take encoder placed I-frames into account, have been proposed recently. Hence, a lower number of I-frames is needed, thus achieving a lower video bitrate without quality degradation. While several proposals exploiting variable segment durations exist, no comparative study highlighting the impact of this technique on coding efficiency and adaptive streaming performance has been conducted yet. This paper conducts such a holistic comparison within the adaptive video streaming eco-system. Firstly, it provides a broad investigation of video encoding efficiency for variable segment durations. Secondly, a measurement study evaluates the impact of segment duration variability on the performance of HAS using three adaptation heuristics and the dash.js reference implementation. Our results show that variable segment durations increased the Quality of Experience for 54% of the evaluated streaming sessions, while reducing the overall bitrate by 7% on average. Susanna Schwarzmann, Nick Hainke, Thomas Zinner, Christian Sieber, Werner Robitza, Alexander Raake |
MMSys | 6 |
| 2020 | Development and Evaluation of a Test Setup to Investigate Distance Differences in Immersive Virtual EnvironmentsabstractNowadays, with recent advances in virtual reality technology, it is easily possible to integrate real objects into virtual environments by creating an exact virtual replication and enabling interaction with them by mapping the obtained tracking data of the real to the virtual objects. The primary goal of our study is to develop a system to investigate distance differences for near-field interaction in immersive virtual environments. In this context, the term distance difference refers to the shift between a real object and the respective replication of the real object in the virtual environment of the same size. This could occur for a number of reasons e.g. due to errors in motion tracking or mistakes in designing the virtual environment. Our virtual environment is developed using the Unity3D game engine, while the immersive contents were displayed on an HTC Vive Pro head-mounted display. The virtual room shown to the user includes a replication of the real testing lab environment, while one of the two real objects is tracked and mirrored to the virtual world using an HTC Vive Tracker. Both objects are present in the real as well as in the virtual world. To find perceivable distance differences in the near-field, the actual task in the subjective test was to pick up one object and place it into another object. The position of the static object in the virtual world is shifted by values between 0 and 4 cm, while the position of the real object is kept constant. The system is evaluated by conducting a subjective proof-of-concept test with 18 test subjects. The distance difference is evaluated by the subjects through estimating perceived confusion on a modified 5-point absolute category rating scale. The study provides quantitative insights into allowable real-world vs. virtual-world mismatch boundaries for near-field interactions, with a threshold value of around 1 cm. Stephan Fremerey, Muhammad Sami Suleman, Abdul Haq Azeem Paracha, Alexander Raake |
QoMEX | 4 |
| 2020 | Prenc - Predict Number of Video Encoding Passes with Machine LearningabstractVideo streaming providers spend huge amounts of processing time to get a quality-optimized encoding. While the quality-related impact may be known to the service provider, the impact on video quality is hard to assess, when no reference is available. Here, bitstream-based video quality models may be applicable, delivering estimates that include encoding-specific settings. Such models typically use several input parameters, e.g. bitrate, framerate, resolution, video codec, QP values and more. However, for a given bitstream, to determine which encoding parameters were selected, e.g., the number of encoding passes, is not a trivial task. This leads to our following research question: Given an unknown video bitstream, which encoding settings have been used? To tackle this reverse engineering problem, we introduce a system called prenc. Besides the use in video-quality estimation, such algorithms may also be used in other applications such as video forensics. We prove our concept by applying prenc to distinguish between one- and two-pass encoding. Starting from modeling the problem as a classification task, estimating bitstream-based features, we further describe a machine learning approach with feature selection to automatically predict the number of encoding passes for a given video bitstream. Our large-scale evaluation consists of 16 short movie type 4K videos that were segmented and encoded with different settings (resolutions, codecs, bitrates), so that we in total analyzed 131.976 DASH video segments. We further show that our system is robust, based on a 50% train and 50% validation approach without source video overlapping, where we get a classification performance of 65% F1 score. Moreover, we also describe the used bitstream-based features in detail, the feature pooling strategy and include other machine learning algorithms in our evaluation. Steve Goering, Rakesh Rao Ramachandra Rao, Alexander Raake |
QoMEX | 3 |
| 2020 | Let the Music Play: An Automated Test Setup for Blind Subjective Evaluation of Music Playback on Mobile DevicesabstractSeveral methods for subjective evaluation for audio and speech have been standardized over the last years. However, with the advancement of mobile devices such as smartphones and Bluetooth speakers, people listen to music even outside their home environment, when traveling and in social situations. Conventional comparative methodologies are difficult to use for sound-quality evaluation of such devices, since subjects are likely to include other factors such as brand or design. Hence, we propose an automated test setup to evaluate music and audio playback of portable devices with subjects without revealing the devices or interfering with the tests. Furthermore, an identical placement of the devices in front of the listener is crucial to accommodate the individual acoustic directivity of the device. For this purpose, we use a large motorized turntable on which the devices are mounted so that the playback device is automatically moved to the defined position in advance. An enhanced version of rating software avrateNG enables the automatic playout of musical pieces and appropriate turning of the devices to face the listeners. Devices that can automatically be tested using our setup include Android and iOS smartphones, as well as Bluetooth and wired portable speakers. Preliminary user tests were conducted to verify the practical applicability and stability of the proposed setup. Dominik Keller, Alexander Raake, Markus Vaalgamaa, Erkki Paajanen |
QoMEX | 2 |
| 2020 | Bitstream-Based Model Standard for 4K/UHD: ITU-T P.1204.3 - Model Details, Evaluation, Analysis and Open Source ImplementationabstractWith the increasing requirement of users to view high-quality videos with a constrained bandwidth, typically realized using HTTP-based adaptive streaming, it becomes more and more important to determine the quality of the encoded videos accurately, to assess and possibly optimize the overall streaming quality. In this paper, we describe a bitstream-based no-reference video quality model developed as part of the latest model-development competition conducted by ITU-T Study Group 12 and the Video Quality Experts Group (VQEG), “P.NATS Phase 2”. It is now part of the new P.1204 series of Recommendations as P.1204.3. It can be applied to bitstreams encoded with H.264/AVC, HEVC and VP9, using various encoding options, including resolution, bitrate, framerate and typical encoder settings such as number of passes, rate control variants and speeds. The proposed model follows an ensemble-modelling-inspired approach with weighted parametric and machine-learning parts to efficiently leverage the performance of both approaches. The paper provides details about the general approach to modelling, the features used and the final feature aggregation. The model creates per-segment and per-second video quality scores on the 5-point Absolute Category Rating scale, and is applicable to segments of 5–10 seconds duration. It covers both PC/TV and mobile/tablet viewing scenarios. We outline the databases on which the model was trained and validated as part of the competition, and perform an additional evaluation using a total of four independently created databases, where resolutions varied from 360p to 2160p, and frame rates from 15–60fps, using realistic coding and bitrate settings. We found that the model performs well on the independent dataset, with a Pearson correlation of 0.942 and an RMSE of 0.42. We also provide an open-source reference implementation of the described P.1204.3 model, as well as the multi-codec bitstream parser required to extract the input data, which is not part of the standard. Rakesh Rao Ramachandra Rao, Steve Goering, Peter List 0001, Werner Robitza, Bernhard Feiten, Ulf Wüstenhagen, Alexander Raake |
QoMEX | 7 |
| 2020 | Are You Still Watching? Streaming Video Quality and Engagement Assessment in the CrowdabstractAs video streaming accounts for the majority of Internet traffic, monitoring its quality is of importance to both Over the Top (OTT) providers as well as Internet Service Providers (ISPs). While OTTs have access to their own analytics data with detailed information, ISPs often have to rely on automated network probes for estimating streaming quality, and likewise, academic researchers have no information on actual customer behavior. In this paper, we present first results from a large-scale crowdsourcing study in which three major video streaming OTTs were compared across five major national ISPs in Germany. We not only look at streaming performance in terms of loading times and stalling, but also customer behavior (e.g., user engagement) and Quality of Experience based on the ITU-T P.1203 QoE model. We used a browser extension to evaluate the streaming quality and to passively collect anonymous OTT usage information based on explicit user consent. Our data comprises over 400,000 video playbacks from more than 2,000 users, collected throughout the entire year of 2019. The results show differences in how customers use the video services, how the content is watched, how the network influences video streaming QoE, and how user engagement varies by service. Hence, the crowdsourcing paradigm is a viable approach for third parties to obtain streaming QoE insights from OTTs. Werner Robitza, Alexander M. Dethof, Steve Goering, Alexander Raake, André Beyer, Tim Polzehl |
QoMEX | 4 |
| 2020 | Towards Analysing the Interaction between Quality and Storytelling for Event Video RecordingabstractMulti-camera video recordings of events such as theatre or other stage performances are difficult to realize in a non-professional environment. Laymen are untrained in camera work. This can for example be remedied by a production process with high-resolution cameras. Using the recordings during post-production, manually or automatically clipped image sections from long or medium long shots can serve for different types of scene representations. Medium close or close shots in storytelling allow the events to be experienced more closely and intensively but have less resolution, with noticeable loss of quality for the viewer. An online study was conducted to evaluate quality and preference for three versions of the same scene with shots created by different crops, collecting qualitative information to understand the reasons viewers gave for their preference. Eckhard Stoll, Stephan Breide, Alexander Raake |
QoMEX | 3 |
| 2019 | cencro - Speedup of Video Quality Calculation using Center CroppingabstractToday's video streaming providers, e.g. Youtube, Netflix or Amazon Prime, are able to deliver high resolution and high-quality content to end users. To optimize video quality and to reduce transmission bandwidth, new encoders and smarter encoding schemes are required. Encoding optimization forms an important part of this effort in reducing bandwidth and results in saving considerable amount of bitrate. For such optimization, accurate and computationally fast video quality models are required, e.g. Netflix's VMAF. However, VMAF is a full-reference (FR) metric, and the calculation of such metrics tend to be slower in comparison to other metrics, due to the amount of data that needs to be processed, especially for high resolutions of 4k and beyond. We introduce an approach to speed up video quality metric calculations in general. We use VMAF as an example with a video database up to 4K resolution videos, to show that our approach works well. Our main idea is that we reduce each frame of the reference and distorted video based on a center crop of the frame, assuming that most important visual information are presented in the middle of most typical videos. In total we analyze 18 different crop settings and compare our results with uncropped VMAF values and subjective scores. We show that this approach - named cencro - is able to save up to 95% computation time, with just an overall error of 4% considering a 360p center crop. Furthermore, we checked other full-reference metrics, and show that cencro performs similar good. As a last evaluation, we apply our approach to full-hd gaming videos, also in this scenario cencro can be successfully applied. The idea behind cencro is not restricted to full-reference models and can also be applied to other type of video quality models or datasets, or even for higher resolution videos such as 8K. Steve Goering, Christopher Krämmer, Alexander Raake |
ISM | 3 |
| 2019 | AVT-VQDB-UHD-1: A Large Scale Video Quality Database for UHD-1abstract4K television screens or even with higher resolutions are currently available in the market. Moreover video streaming providers are able to stream videos in 4K resolution and beyond. Therefore, it becomes increasingly important to have a proper understanding of video quality especially in case of 4K videos. To this effect, in this paper, we present a study of subjective and objective quality assessment of 4K ultra-high-definition videos of short duration, similar to DASH segment lengths. As a first step, we conducted four subjective quality evaluation tests for compressed versions of the 4K videos. The videos were encoded using three different video codecs, namely H.264, HEVC, and VP9. The resolutions of the compressed videos ranged from 360p to 2160p with framerates varying from 15fps to 60fps. All the source 4K contents used were of 60fps. We included low quality conditions in terms of bitrate, resolution and framerate to ensure that the tests cover a wide range of conditions, and that e.g. possible models trained on this data are more general and applicable to a wider range of real world applications. The results of the subjective quality evaluation are analyzed to assess the impact of different factors such as bitrate, resolution, framerate, and content. In the second step, different state-of-the-art objective quality models were applied to all videos and their performance was analyzed in comparison with the subjective ratings, e.g. using Netflix's VMAF. The videos, subjective scores, both MOS and confidence interval per sequence and objective scores are made public for use by the community for further research. Rakesh Rao Ramachandra Rao, Steve Goering, Werner Robitza, Bernhard Feiten, Alexander Raake |
ISM | 5 |
| 2019 | ViProVoQ: Towards a Vocabulary for Video Quality Assessment in the Context of Creative Video ProductionabstractThis paper presents a method for developing a consensus vocabulary to describe and evaluate the visual experience of videos. As a first result, a vocabulary characterizing the specific look of cinema-type video is presented. Such a vocabulary can be used to relate perceptual features of professional high-end image and video quality of experience (QoE) with the underlying technical characteristics and settings of the video systems involved in the creative content production process. For the vocabulary elicitation, a combination of different survey techniques was applied in this work. As the first step, individual interviews were conducted with experts of the motion picture industry on image quality in the context of cinematography. The data obtained from the interviews was used for the subsequent Real-time Delphi survey, where an extended group of experts worked out a consensus on key aspects of the vocabulary specification. Here, 33 experts were supplied with the anonymized results of the other panelists, which they could use to revise their own assessment. Based on this expert panel, the attributes collected in the interviews were verified and further refined, resulting in the final vocabulary proposed in this paper. Besides an attribute-based sensory evaluation of high-quality image, video and film material, applications of the vocabulary are the development of dimension-based image and video quality models, and the analysis of the multivariate relationship between quality-relevant perceptual attributes and technical system parameters. Simon Wedel, Michael Koppetz, Janto Skowronek, Alexander Raake |
ACM Multimedia | 4 |
| 2019 | Comparison of Subjective Quality Test Methods for Omnidirectional Video Quality EvaluationabstractThe test methods recommended by the International Telecommunication Union (ITU) for assessing 2D video quality are often used for evaluating omnidirectional / 360° videos. In this paper, we compare the performance of three different test methods, Absolute Category Rating (ACR), a modified version of ACR (M-ACR) with double presentation of the test stimulus, and DSIS (Double Stimulus Impairment Scale), based on the statistical reliability, assessment time and simulator sickness. Different settings were used for HEVC encoding of five 360° source videos of 10 s duration. Results indicate that DSIS is statistically more reliable with higher resolving power, followed by M-ACR and ACR. We found that simulator sickness increases with time, but can be reduced by taking breaks in-between the test sessions. The results for simulator sickness are compared across test methods and with similar tests conducted under different contextual conditions. We also recorded and analyzed the exploration behaviour of the users. Apart from the methodological findings, the test results provide insights into video quality for different resolution and encoding settings (“bitrate ladders”). These may be useful for choosing appropriate representations in the context of HTTP-based adaptive streaming in case of full-frame streaming. Ashutosh Singla, Werner Robitza, Alexander Raake |
MMSP | 3 |
| 2019 | Subjective quality evaluation of tile-based streaming for omnidirectional videosabstractIn viewport-adaptive streaming of omnidirectional video, only the field of view is streamed in high quality. While this has significant benefits over streaming the entire 360 sphere, no standard test method for perceived quality and simulator sickness is available to evaluate the quality of experience (QoE) of such streaming approaches. QoE testing is important as tile-based viewport-adaptive streaming technologies are replacing classical approaches because of significant bandwidth savings and increase in viewing quality. In this work, we propose a testbed, a test method, as well as test metrics for QoE tests of viewport-adaptive streaming approaches. The proposed method is validated in two different test setups, using a specific tile-based streaming technology available in the market. The chosen input variables (videos sequences, resolution, bandwidth, and network round-trip delay) are tested for their statistical significance. We found that our test method is suitable for QoE testing of viewport-adaptive streaming technologies. We also found that simulator sickness scores increase with test duration, but that breaks between tests reduce this effect. With our systematic test approach, it is possible to compare metrics among different test setups. On the tested technology, we found that a typical network delay (47 ms) only has a minimal effect on the quality ratings. Furthermore, the magnitude of the network delay does not influence simulator sickness for the system we have tested. Ashutosh Singla, Steve Goering, Alexander Raake, Britta Meixner, Rob Koenen, Thomas Buchholz |
MMSys | 3 |
| 2019 | Impact of Various Motion Interpolation Algorithms on 360° Video QoEabstractIn our study, we compare the impact of various motion interpolation (MI) algorithms on 360° video Quality of Experience (QoE). For doing so, we conducted a subjective test with 12 video expert viewers, while a pair comparison test method was used. We interpolated four different 20 s long 30 fps 360° source contents to the native 90 Hz refresh rate of popular Head-Mounted Displays using three different MI algorithms. Subsequently, we compared these 90 fps videos against each other to investigate the influence on the QoE. Regarding the algorithms, we found out that ffmpeg blend does not lead to a significant improvement of QoE, while MCI and butterflow do so. Additionally, we concluded that for 360° videos containing fast and sudden movements, MCI should be preferred over butterflow, while butterflow is more suitable for slow and medium motion videos. While comparing the time needed for rendering the 90 fps interpolated videos, ffmpeg blend is the fastest, while MCI and butterflow need much more time. Stephan Fremerey, Frank Hofmeyer, Steve Goering, Alexander Raake |
QoMEX | 4 |
| 2019 | nofu - A Lightweight No-Reference Pixel Based Video Quality Model for Gaming ContentabstractPopularity of streaming services for gaming videos has increased tremendously over the last years, e.g. Twitch and Youtube Gaming. Compared to classical video streaming applications, gaming videos have additional requirements. For example, it is important that videos are streamed live with only a small delay. In addition, users expect low stalling, waiting time and in general high video quality during streaming, e.g. using http-based adaptive streaming. These requirements lead to different challenges for quality prediction in case of streamed gaming videos. We describe newly developed features and a no-reference video quality machine learning model, that uses only the recorded video to predict video quality scores. In different evaluation experiments we compare our proposed model nofu with state-of-the-art reduced or full reference models and metrics. In addition, we trained a no-reference baseline model using brisque+niqe features. We show that our model has a similar or better performance than other models. Furthermore, nofu outperforms VMAF for subjective gaming QoE prediction, even though nofu does not require any reference video. Steve Goering, Rakesh Rao Ramachandra Rao, Alexander Raake |
QoMEX | 3 |
| 2019 | Assessing Texture Dimensions and Video Quality in Motion Pictures using Sensory Evaluation TechniquesabstractThe quality of images and videos is usually examined with well established subjective tests or instrumental models. These often target content transmitted over the internet, such as streaming or videoconferences and address the human preferential experience. In the area of high-quality motion pictures, however, other factors are relevant. These mostly are not error-related but aimed at the creative image design, which has gained comparatively little attention in image and video quality research. To determine the perceptual dimensions underlying movie-type video quality, we combine sensory evaluation techniques extensively used in food assessment - Degree of Difference test and Free Choice Profiling - with more classical video quality tests. The main goal of this research is to analyze the suitability of sensory evaluation methods for high-quality video assessment. To understand which features in motion pictures are recognizable and critical to quality, we address the example of image texture properties, measuring human perception and preferences with a panel of image-quality experts. To this aim, different capture settings were simulated applying sharpening filters as well as digital and analog noise to exemplary source sequences. The evaluation, involving Multidimensional Scaling, Generalized Procrustes Analysis as well as Internal and External Preference Mapping, identified two separate perceptual dimensions. We conclude that Free Choice Profiling connected with a quality test offers the highest level of insight relative to the needed effort. The combination enables a quantitative quality measurement including an analysis of the underlying perceptual reasons. Dominik Keller, Tamara Seybold, Janto Skowronek, Alexander Raake |
QoMEX | 4 |
| 2019 | Assessing Media QoE, Simulator Sickness and Presence for Omnidirectional Videos with Different Test ProtocolsabstractQoE for omnidirectional videos comprises additional components such as simulator sickness and presence. In this paper, a series of tests is presented comparing different test protocols to assess integral quality, simulator sickness and presence for omnidirectional videos in one test run, using the HTC Vive Pro as head-mounted display. For quality ratings, the five-point ACR scale was used. In addition, the well-established Simulator Sickness Questionnaire and Presence Questionnaire methods were used, once in a full version, and once with only one single integral scale, to analyze how well presence and simulator sickness can be captured using only a single scale. Ashutosh Singla, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
VR | 4 |
| 2019 | User attitudes and behaviors toward personalized control of privacy settings on smartphonesabstractSummary The fine‐grained access control has been proved to be a reliable tool to ensure preserving of privacy of end users. In fog computing, one of the challenges is to understand users' attitudes and behaviors toward personalized control. However, few of studies have given a clear view on users' perception of the burden of interactivity when they set complex privacy settings. To this end, we conducted a user study including a lab study with 26 participants and an evaluation with 223 participants. From the lab study, we found that participants were satisfied with improved privacy settings but did not adapt well to complex personalized interfaces. We proposed effective methods to assist users to balance between the full control and the additional interaction burden, including sorting, recommendations, and establishing profiles. After this lab study, we organized a survey evaluation additionally to explore users' current usage of privacy features. Results from the evaluation showed that the principle reason that users failed to use privacy features was that they were not appropriately aware of these features. A key conclusion is that privacy settings should not only let users take over the control of smartphones but also inform them of the knowledge on privacy practices. Yun Zhou 0003, Lianyong Qi, Alexander Raake, Tao Xu 0008, Marta Piekarska, Xuyun Zhang |
Concurr. Comput. Pract. Exp. | 3 |
| 2018 | AVtrack360: an open dataset and software recording people's head rotations watching 360° videos on an HMDabstractIn this paper, we present a viewing test with 48 subjects watching 20 different entertaining omnidirectional videos on an HTC Vive Head Mounted Display (HMD) in a task-free scenario. While the subjects were watching the contents, we recorded their head movements. The obtained dataset is publicly available in addition to the links and timestamps of the source contents used. Within this study, subjects were also asked to fill in the Simulator Sickness Questionnaire (SSQ) after every viewing session. Within this paper, at first SSQ results are presented. Several methods for evaluating head rotation data are presented and discussed. In the course of the study, the collected dataset is published along with the scripts for evaluating the head rotation data. The paper presents the general angular ranges of the subjects' exploration behavior as well as an analysis of the areas where most of the time was spent. The collected information can be presented as head-saliency maps, too. In case of videos, head-saliency data can be used for training saliency models, as information for evaluating decisions during content creation, or as part of streaming solutions for region-of-interest-specific coding as with the latest tile-based streaming solutions, as discussed also in standardization bodies such as MPEG. Stephan Fremerey, Ashutosh Singla, Kay Meseberg, Alexander Raake |
MMSys | 4 |
| 2018 | HTTP adaptive streaming QoE estimation with ITU-T rec. P. 1203 open databases and softwareabstractThis paper describes an open dataset and software for ITU-T Ree. P.1203. As the first standardized Quality of Experience model for audiovisual HTTP Adaptive Streaming (HAS), it has been extensively trained and validated on over a thousand audiovisual sequences containing HAS-typical effects (such as stalling, coding artifacts, quality switches). Our dataset comprises four of the 30 official subjective databases at a bitstream feature level. The paper also includes subjective results and the model performance. Our software for the standard was made available to the public, too, and it is used for all the analyses presented. Among other previously unpublished details, we show the significant performance improvements of using bitstream-based models over metadata-based ones for video quality analysis, and the robustness of combining classical models with machine-learning-based approaches for estimating user QoE. Werner Robitza, Steve Goering, Alexander Raake, David Lindegren, Gunnar Heikkilä, Jörgen Gustafsson, Peter List 0001, Bernhard Feiten, Ulf Wüstenhagen, Marie-Neige Garcia, Kazuhisa Yamagishi, Simon Broom |
MMSys | 3 |
| 2018 | Extended Features using Machine Learning Techniques for Photo Liking PredictionabstractToday several photo platforms provide thousands of new pictures, it becomes ambitious to find highly appealing or like-able photos within such loads of data. Here, automatic liking prediction can support users in handling their pictures or improve ranking in sharing platforms. We describe a machine learning approach for photo liking prediction. Our features are based on various techniques, e.g. natural language processing/sentiment analysis, pre-trained deep learning networks, social network analysis and extended previously reported features. We conduct large-scale experiments using a collected dataset consisting of 80k photos based on two main categories from 500px with different settings. In our experiments we analyzed the impact of our newly features and found that social network features have the strongest influence for liking prediction, we achived a boost of 15%. Furthermore, we show that all implemented features are able to improve prediction accuracy of liking rates. We additionally analyze which groups of features that can be derived directly from pictures are usable for prediction. Steve Goering, Konstantin Brand, Alexander Raake |
QoMEX | 3 |
| 2018 | Measuring YouTube QoE with ITU-T P.1203 Under Constrained Bandwidth ConditionsabstractThe available Internet bandwidth has a strong impact on the Quality of Experience of video services. In order to manage their network efficiently and prevent customer churn, Internet Service Providers need to constantly monitor the QoE of video services such as YouTube. However, they often only rely on simple measurement scenarios that consider only one video being loaded repeatedly. In this paper we compare this scenario against a new approach in which multiple videos are being loaded in a session, thereby simulating user behavior. Using a testbed, we study the impact of download speeds on Key Performance Indicators (KPIs such as initial loading time and stalling events) and user QoE as measured using the ITU-T P.1203 standard. We show that the monitoring paradigm has a significant impact on the obtained results. We further provide a prediction model for estimating the impact of download speed on KPIs and user QoE. Werner Robitza, Dhananjaya G. Kittur, Alexander M. Dethof, Steve Goering, Bernhard Feiten, Alexander Raake |
QoMEX | 6 |
| 2018 | On the quality perception of multiparty conferencing callsabstractThis paper presents a study on the quality perception of multiparty audio- and audiovisual conferencing calls, with a focus on asymmetric conditions, that is, different connection properties and equipment of individual participants. The results show that some mutual influence of the individual links between participants in terms of their perceived quality exists, and that the overall quality of a conference call is not always a simple average over quality ratings associated with individual links. Further, the paper interprets the results in terms of relevant processes of audiovisual scene perception and quality formation. The results can help operators or manufacturers to optimally balance QoS settings for individual participants, shedding light on how the overall impression of a conference may be formed by users. Janto Skowronek, Alexander Raake |
QoMEX | 2 |
| 2018 | GBVS360, BMS360, ProSal: Extending existing saliency prediction models from 2D to omnidirectional images
Pierre R. Lebreton, Alexander Raake |
Signal Process. Image Commun. | 2 |
| 2018 | Colouration in Local Wave Field SynthesisabstractSound field synthesis techniques including wave field synthesis and near-field-compensated higher order ambisonics aim at a physically accurate reproduction of a desired sound field inside an extended listening area. This area is surrounded by loudspeakers individually driven by their respective driving signals. The latter have to be chosen such that the superposition of all emitted sound fields coincides with the desired one. Due to practical limitations, artefacts impair the synthesis accuracy resulting in a perceivable change in timbre. Recently, two approaches to so-called local wave field synthesis were published that enhance the reproduction accuracy in a limited region while allowing stronger artefacts outside. This paper reports on two listening experiments comparing conventional techniques for sound field synthesis with the mentioned approaches. Furthermore, the influence of different parametrizations for local wave field synthesis is investigated. The results show that the enhanced reproduction accuracy in local wave field synthesis leads to a reduction of perceived colouration, if a suitable parametrization is chosen. Fiete Winter, Hagen Wierstorf, Christoph Hold, Frank Krüger 0001, Alexander Raake, Sascha Spors |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Impact of video resolution changes on QoE for adaptive video streamingabstractHTTP adaptive streaming (HAS) has become the de-facto standard for video streaming to ensure continuous multimedia service delivery under irregularly changing network conditions. Many studies already investigated the detrimental impact of various playback characteristics on the Quality of Experience of end users, such as initial loading, stalling or quality variations. However, dedicated studies tackling the impact of resolution adaptation are still missing. This paper presents the results of an immersive audiovisual quality assessment test comprising 84 test sequences from four different video content types, emulated with an HAS adaptation mechanism. We employed a novel approach based on systematic creation of adaptivity conditions which were assigned to source sequences based on their spatio-temporal characteristics. Our experiment investigates the resolution switch effect with respect to the degradations in MOS for certain adaptation patterns. We further demonstrate that the content type and resolution change patterns have a significant impact on the perception of resolution changes. These findings will help develop better QoE models and adaptation mechanisms for HAS systems in the future. Avsar Asan, Werner Robitza, Is-Haka Mkwawa, Lingfen Sun, Emmanuel C. Ifeachor, Alexander Raake |
ICME | 6 |
| 2017 | A framework for QoE analysis of encrypted video streamsabstractToday most internet traffic is generated by video streaming. YouTube and other video streaming platforms are using encrypted streams (HTTPS) for transport of video content. Encryption will lead to more requirements on network and content providers, e.g. caching mechanisms will not work direct. Estimation of video quality for measuring users satisfaction is also harder because there is no direct access to the video bitstream. We are building up a framework for analyzing video quality that allows us to store client information, decrypted network traffic and encrypted messages. Our approach is based on a man-in-the-middle proxy for storing the decrypted video bitstream, active probing and traffic shaping. Using these data, we are able to calculate video QoE values for example using a model such as ITU-T Rec. P.1203. Our framework will be used for generating datasets for encrypted video stream analysis, analyzing internal behavior of video streaming platforms, and more. For experimental evaluation, in this paper we analyze the influence of our man-in-the-middle proxy on key-performance indicators (KPIs) for video streaming quality. Steve Goering, Alexander Raake, Bernhard Feiten |
QoMEX | 2 |
| 2017 | The label knows better: The impact of labeling effects on perceived quality of HD and UHD video streamingabstractThere is an ongoing debate in the research community over the improved visual quality of UHD video in comparison to the still widely-deployed HD standard. It is the inspiration of many scientific studies, yet UHD displays and services are continuously spreading on the consumer market. This paper presents the results of a subjective paired-comparison test with both upscaled HD and UHD video sequences, investigating the primary research question of whether UHD can offer a significant improvement over HD for common viewing conditions. In our study, subjects rated their visual preference on video-only clips without lossy encoding. In addition we studied cognitive biases, by presenting a label of what users were about to see (HD or UHD) before the sequence, purposely manipulating the labels in some conditions to suggest different resolutions than actually shown. Finally, we investigated the impact of different rating scales on the precision of the results. Our studies show that HD clips appear indistinguishable from UHD clips when the rating scale chosen is not fine-grained enough. We also found that the labeling effect has a significant impact on the perceived quality, overriding the actual visual perception. Even with a precise scale, users cannot detect major improvements of UHD compared to HD, which may let us question the added value offered by UHD. Péter A. Kara, Werner Robitza, Alexander Raake, Maria G. Martini |
QoMEX | 3 |
| 2017 | A bitstream-based, scalable video-quality model for HTTP adaptive streaming: ITU-T P.1203.1abstractThe paper presents the scalable video quality model part of the P.1203 Recommendation series, developed in a competition within ITU-T Study Group 12 previously referred to as P.NATS. It provides integral quality predictions for 1 up to 5 min long media sessions for HTTP Adaptive Streaming (HAS) with up to HD video resolution. The model is available in four modes of operation for different levels of media-related bitstream information, reflecting different types of encryption of the media stream. The video quality model presented in this paper delivers short-term video quality estimates that serve as input to the integration component of the P.1203 model. The scalable approach consists in the usage of the same components for spatial and temporal scaling degradations across all modes. The third component of the model addresses video coding artifacts. To this aim, a single model parameter is introduced that can be derived from different types of bitstream input information. Depending on the complexity of the available input, one of four scaling-levels of the model is applied. The paper presents the different novelties of the model and scientific choices made during its development, the test design, and an analysis of the model performance across the different modes. Alexander Raake, Marie-Neige Garcia, Werner Robitza, Peter List 0001, Steve Goering, Bernhard Feiten |
QoMEX | 1 |
| 2017 | A modular HTTP adaptive streaming QoE model - Candidate for ITU-T P.1203 ("P.NATS")abstractThis paper describes a quality model for HTTP Adaptive Streaming. It integrates existing audio and video quality scores to a final quality estimation, factoring in quality variations over time, the recency effect, as well as location and length of buffering events at the player side. We built the model based on data gathered from more than 17 subjective quality tests. It was submitted to the ITU-T P.NATS competition; parts of it have since been released in the official recommendation ITU-T P.1203.3 as an “audiovisual quality integration module”. In the context of standardization, the model was validated on 30 subjective databases, showing high performance. Its modular approach allows its components to be re-used in other applications and combined with different temporal pooling techniques. Werner Robitza, Marie-Neige Garcia, Alexander Raake |
QoMEX | 3 |
| 2017 | Measuring and comparing QoE and simulator sickness of omnidirectional videos in different head mounted displaysabstractIn this paper, we evaluated and compared the integral quality of different omnidirectional contents for two head mounted displays (HMDs), namely HTC Vive and Oculus Rift. We also investigated motion sickness and head-movements. To this aim, we categorized omnidirectional contents into three categories based on the degree of motion: high, medium and low motion. For assessing simulator sickness, we used the Simulator Sickness Questionnaire for each of the contents in both HMDs. The viewing direction for subjects while watching the contents were recorded in terms of the three coordinates yaw, roll and pitch. Experimental results show that HTC Vive offers better integral quality compared to Oculus Rift. We also compared simulator sickness scores along with the behavioral data for different contents and HMDs and discussed the results in the paper. Ashutosh Singla, Stephan Fremerey, Werner Robitza, Alexander Raake |
QoMEX | 4 |
| 2017 | Testing conversational quality of VoIP with different terminals and degradationsabstractIn this paper, telephone conversation test results are reported. The main goal of the research is to derive a quality assessment model for today's Voice over Internet Telephony (VoIP) communication including the influence of end-point terminals with their internal signal processing. For this reason, two different terminals were used during the test and a possibly wide range of impairments were simulated, including: coding, packet loss, noise and echo. This test serves as a core test for the other experiments by the authors. It has a refined design as compared to the test presented already in [1] with conditions allowing interterminal comparisons. Hence, some relevant insights into terminal influence on testing of conversational quality are presented in this paper. Michal Soloducha, Alexander Raake, Frank Kettler, Stefan Bleiholder |
QoMEX | 2 |
| 2017 | Challenges of future multimedia QoE monitoring for internet service providersabstractThe ever-increasing network traffic and user expectations at reduced cost make the delivery of high Quality of Experience (QoE) for multimedia services more vital than ever in the eyes of Internet Service Providers (ISPs). Real-time quality monitoring, with a focus on the user, has become essential as the first step in cost-effective provisioning of high quality services. With the recent changes in the perception of user privacy, the rising level of application-layer encryption and the introduction and deployment of virtualized networks, QoE monitoring solutions need to be adapted to the fast changing Internet landscape. In this contribution, we provide an overview of state-of-the-art quality monitoring models and probing technologies, and highlight the major challenges ISPs have to face when they want to ensure high service quality for their customers. Werner Robitza, Arslan Ahmad, Péter A. Kara, Luigi Atzori, Maria G. Martini, Alexander Raake, Lingfen Sun |
Multim. Tools Appl. | 6 |
| 2016 | Studying user agreement on aesthetic appeal ratings and its relation with technical knowledgeabstractIn this paper, a crowdsourcing experiment was conducted involving different panels of participants. The aim of this study is to evaluate how the preference of one image over another one is related with the knowledge of the participant in photography. In previous work the two discriminant evaluation concepts “presence of a main subject” and “exposure” were found to distinguish group participants with different degrees of knowledge in photography. Each of these groups provided different means of aesthetic appeal ratings when asked to rate on an absolute category scale. The present paper extends previous work by studying preference ratings on a set of image pairs as a function of technical knowledge and more specifically adding a focus on the variance of rating and agreement between participants. The conducted study was composed of two different steps where the participants had to first report their preference of one image over another (paired comparison), and an evaluation of the technical background of the participant using a specific set of images. Based on preference-rating patterns groups of participants were identified. These groups were formed by clustering the participants who saw and shared the same preference rating on images in one group, and the participants with low agreement with other participants in another group. A per-group analysis showed that a high agreement between participants could be observed when participants have technical knowledge. This indicates that higher consistency between participants can be reached when expert users are being recruited, and therefore participants should be carefully selected in image aesthetic appeal evaluation to ensure stable results. Pierre R. Lebreton, Alexander Raake, Marcus Barkowsky |
QoMEX | 2 |
| 2016 | (Re-)actions speak louder than words? A novel test method for tracking user behavior in web video servicesabstractAssuring user engagement has become a key issue for Internet Service Providers and Over-the-Top Providers. How long are users consuming a service? When are they likely to abandon it due to quality problems? Rather than just estimating perceived audiovisual quality, future quality prediction models will also factor in possible user behavior. This contribution presents a novel test method to assess short-term user behavior in web video services, in a controlled living-room-like environment. We show that typical behavioral responses (such as seeking, reloading, or selecting another video) can be elicited, with the real purpose of the test hidden from the viewers. We can also see that when users are not focused on judging quality, their perception of errors changes significantly. This paper highlights the strong impact of laboratory test situations on users' behavior and discusses the challenges revolving around finding valid test methods. Werner Robitza, Alexander Raake |
QoMEX | 2 |
| 2016 | Efficient no-reference metric for sharpness mismatch artifact between stereoscopic views
Mohan Liu, Karsten Müller 0001, Alexander Raake |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | I have to switch the terminal: Evaluating the impact on video quality perceptionabstractHTTP adaptive streaming technology is now widely adopted in multimedia services because of its ability to provide adaptation to the streaming context, especially characteristics of end-user devices and dynamic network conditions. There are various studies targeting the evaluation of the Quality of Experience (QoE) in this framework. However, none has considered the scenario of the user changing the viewing device during the streaming session, which is the objective of this paper. It provides the following major contributions: definition of the multi-device streaming session scenario; the implementation of a realistic testing case; the execution of subjective tests involving 28 people; and the detailed analysis of the influence of the devices' switching events. Nicola Abis, Alessandro Floris, Savvas Argyropoulos, Luigi Atzori, Alexander Raake |
ICC | 5 |
| 2015 | Assessment of Cognitive Load, Speech Communication Quality and Quality of Experience for spatial and non-spatial audio conferencing calls
Janto Skowronek, Alexander Raake |
Speech Commun. | 2 |
| 2014 | Why are you so slow? - Misattribution of transmission delay to attributes of the conversation partner at the far-end
Katrin Schoenenberg, Alexander Raake, Judith Koeppe |
Int. J. Hum. Comput. Stud. | 2 |
| 2014 | On interaction behaviour in telephone conversations under transmission delay
Katrin Schoenenberg, Alexander Raake, Sebastian Egger-Lampl, Raimund Schatz |
Speech Commun. | 2 |
| 2013 | Scene change detection in encrypted video bit streamsabstractIn this paper, a novel method to detect scene changes in encrypted video streams is presented. Typically, in IPTV systems, the media stream is transmitted in encrypted form, and therefore the only available information to determine the scene changes are the packet headers which transport the video signal. Thus, the proposed method estimates the size and the type of each picture of the video sequence by extracting information from the packet headers. Then, based on the GOP structure, a set of rules are determined to predict changes of frame sizes which are indicative of scene changes in the video sequences. Furthermore, the application of the proposed method in the recently standardized ITU-T Recommendation P.1201.2 for no-reference audio-visual quality assessment for IPTV-grade services is presented to highlight how such method could be deployed. Finally, the proposed method is evaluated on a large set of video databases to demonstrate the validity of the proposed method. Savvas Argyropoulos, Peter List 0001, Marie-Neige Garcia, Bernhard Feiten, Martin Pettersson, Alexander Raake |
ICIP | 6 |
| 2013 | Predicting speech quality based on interactivity and delayabstractA new model of speech quality under delay is presented that includes conversational interactivity. It is based on two previously reported narrowband telephony conversation tests involving different delays, with subject-pairs judging overall quality after each conversation. The tests were conducted with different conversation scenarios targeting different levels of interactivity. The instructions given prior to the tests were varied in their emphasis on speed of task completion. Based on the test results, the paper proposes an extension of a widely used conversational speech quality model, the so-called E-model (ITUT Rec. G.107), to cover the joint effect of interactivity and delay. To this aim, two new parameters are introduced, one of which represents the minimum perceivable delay, and the other expresses in how far users will attribute the delay-effect to the conversational quality of the line. Based on the analysis of the recorded test conversations in terms of its surface structure (turns, speaker activities, etc.), prominent differences and delay-dependencies of a number of conversation parameters were found that characterize the impact of delay on the conversational flow and on perceived quality. Alexander Raake, Katrin Schoenenberg, Janto Skowronek, Sebastian Egger-Lampl |
INTERSPEECH | 1 |
| 2013 | Quality assessment of asymmetric multiparty telephone conferences: a systematic method from technical degradations to perceived impairments
Janto Skowronek, Julian Herlinghaus, Alexander Raake |
INTERSPEECH | 3 |
| 2013 | Parametric model for audiovisual quality assessment in IPTV: ITU-T Rec. P.1201.2abstractA parametric packet-based model has been created to estimate user perceived audiovisual quality of Internet Protocol Television (IPTV) services. It is divided into three modules, for audio, video and audiovisual quality. The model is applicable to the quality monitoring of encrypted and non-encrypted audiovisual streams. Typical audio and video degradations for IPTV are covered for Standard Definition (SD) and High Definition (HD) video formats. The model supports the H.264 video codec and the audio codecs MPEG-I Layer II, MPEG-2 AAC-LC, MPEG-4 HE-AACv2 and AC3. It handles various types of IP-network layer transmission errors. The model was developed and validated using a large database of subjective tests. The underlying concept is based on an impairment factor approach, which enables detection of how users build their individual judgment of quality of a given audiovisual signal. Each impairment factor captures the perceived quality impact of a possible degradation and therefore enables diagnostic analysis of quality problems. The model shows high performance results, both in terms of Pearson's Correlation coefficient (r) and Root-Mean-Square-Error (RMSE). The model is standardized as ITU-T Recommendation P.1201.2, the higher resolution (IPTV and Video on Demand (VoD)) algorithm of Recommendation P.1201. Marie-Neige Garcia, Peter List 0001, Savvas Argyropoulos, David Lindegren, Martin Pettersson, Bernhard Feiten, Jörgen Gustafsson, Alexander Raake |
MMSP | 8 |
| 2013 | Appeal assessment of photographs for quality of experience measurementsabstractAlthough quality of experience (QoE) is defined as a subjective end-user perception, most QoE measurements target boring content. Unsurprisingly, media creators and end-users are interested in interesting content. Therefore, QoE measurements are ineffective for interesting content. In this paper, we focus on a subjective assessment of the appeal of interesting photos. Several reasons for the positive or negative assessment of photos are also described. The experimental results show that there is a relationship between the appealing level of interesting photos and standard deviation. The results show the importance of knowing why an observer determines a photo to be appealing or unappealing. Yasuhiro Inazumi, Alexander Raake, Dominik Strohmeier, Yuukou Horita |
MMSP | 2 |
| 2013 | Modelling image completion distortions in texture analysis-synthesis codingabstractPerception based coding with Texture Analysis and Synthesis (TAS) is a promising way to increase the compression efficiency of modern coding schemes, such as High Efficiency Video Coding (HEVC). TAS approaches typically employ an analysis step which specifies which blocks could be reconstructed by a texture synthesizer. Even though synthesized blocks are perceptually similar to their original versions, they may produce relatively high Mean Squared Error with regard to the original signal. Hence, the contribution of this paper will be a novel image quality assessment method for image completion based on visual attention variations. A subjective experiment has been designed to provide insight into the perceived distortions which are induced by TAS techniques and it is shown that saliency changes are a promising predictor for perceptual distortion. Free access to the image database is provided to encourage further research on this topic. Savvas Argyropoulos, Alexander Raake, Patrick Ndjiki-Nya |
PCS | 3 |
| 2013 | Spatial Sound With Loudspeakers and Its Perception: A Review of the Current StateabstractThis paper reviews the current state of loudspeaker-based spatial sound reproduction methods from technical perspective as well as perceptual perspective. A nomenclature is developed that allows for a strict separation between these two perspectives. The physical fundamentals, practical realization, and results from perceptual studies are discussed for a number of well-established and emerging reproduction techniques. Further, the paper outlines novel approaches to spatial sound evaluation in terms of perceived quality and provides a comparison of current approaches. Sascha Spors, Hagen Wierstorf, Alexander Raake, Frank Melchior, Matthias Frank 0003, Franz Zotter |
Proc. IEEE | 3 |
| 2012 | Same but different? - Using speech signal features for comparing conversational VoIP quality studiesabstractIn this paper we demonstrate how speech signal features can be used to detect and explain differences in human to human conversation tests. To this end, we compare the results of two conversational VoIP quality experiments designed to quantify the impact of network delay on perceived speech quality. Both studies followed the same procedures and used the same scenarios, but were conducted in two different labs. Our comparison shows that the two studies, despite having been executed correctly using the same test design, still can produce surprisingly different results regarding the users quality perception on a MOS scale. In this respect, speech signal features extracted from conversation recordings help identifying divergent participant behavior as plausible cause for such differences. Our in-depth analysis reveals how novel parameters developed by us like Intended and Unintended Interruption Rate (IIR, UIR) and the corrected Speaker Alternation Rate SARcorrcan be used to successfully determine the extent to which the results of different conversational speech quality studies are directly comparable and thus eligible for pooling, or not. Sebastian Egger-Lampl, Raimund Schatz, Katrin Schoenenberg, Alexander Raake, Gernot Kubin |
ICC | 4 |
| 2011 | No-reference bit stream model for video quality assessment of h.264/AVC video based on packet loss visibilityabstractIn this paper, a no reference bit stream model for quality assessment of SD and HD H.264/AVC video sequences based on packet loss visibility is proposed. The method considers the impact of network impairments on human perception and uses the visibility of packet losses to predict objective scores. Also, a new subjective experiment has been designed to provide insight into the perceptual effect of degradations caused by transmission errors. The proposed algorithm extracts a set of features from the received bit stream. Then, the visibility of each packet loss event is determined by classifying the extracted features using a Support Vector Machines classifier. Finally, analytical expressions are developed to account for visual degradation due to compression and channel induced distortion based on the outcome of the visibility classifier. The evaluation demonstrates the validity of the proposed method. Savvas Argyropoulos, Alexander Raake, Marie-Neige Garcia, Peter List 0001 |
ICASSP | 2 |
| 2011 | Conversation Analysis of Multi-Party Conferencing and Its Relation to Perceived QualityabstractThis paper investigates interactivity in multi-party conferencing. To this aim the state model often used to descibe two-party telephone conversations is redefined for more than two interlocutors. We adapt and apply statistical measures such as conversation state probabilities to recordings made during a three-party conferencing quality test, which was designed to evaluate the perceived quality under various speech transmission properties. In a second step, we associate the different observed degrees of interactivity to the quality test ratings. Results give a first insight into the co-dependence of transmission properties, the interlocutors' communication behavior and their integral quality judgment. Katrin Hoeldtke, Alexander Raake |
ICC | 2 |
| 2011 | On Revealing the ARQ Mechanism of MSTVabstractEnsuring a high customer satisfaction by monitoring Quality of Experience (QoE) aspects has become common practice for service providers. Such monitoring solutions, together with underlying QoE models, are mostly limited to measures captured in the core or access network and may thus neglect the QoE impact of recovery mechanisms deployed at client-side, e.g., FEC or ARQ. This limitation makes QoE models prone to mispredict QoE and consequently may lead the operators to misleading interpretations of customer experience. In this paper, we empirically study the behavior of the Microsoft TV Set-top Box (STB) with respect to the deployed ARQ recovery mechanism. As the ARQ implementation details of the STB are proprietary, we implement and simulate three ARQ algorithms of different complexities and evaluate their performance by comparing with corresponding empirical measurements. This comparison reveals insights into the ARQ scheme implemented in the STB. Moreover, it leads us to speculate that MSTV uses simple ARQ schemes which are sufficient to drastically improve the QoE in the presence of a multitude of loss patterns. Oliver Hohlfeld, Balamuhunthan Balarajah, Sebastian Benner, Alexander Raake, Florin Ciucu |
ICC | 4 |
| 2011 | Investigating the Effect of Number of Interlocutors on the Quality of Experience for Multi-Party Audio Conferencing
Janto Skowronek, Alexander Raake |
INTERSPEECH | 2 |
| 2011 | A Subjective Evaluation of 3D Iptv Broadcasting Implementations Considering Coding and Transmission DegradationabstractThis paper describes the results of a subjective test to assess current technology used for 3DTV broadcasting. As a first aspect, the performance of the currently deployed coding schemes was compared to state of the art algorithms. Our results show that down sampling and packing 3D stereoscopic videos according to the so called Side-By-Side format gives the highest perceived quality for a given bit rate. The second aspect of the study was to investigate how common 2D error concealment algorithms perform in case of 3D, and how their 3D-related performance compares with the 2D case. The results provide information on whether binocular suppression or binocular rivalries play the most important role for 3D video quality under transmission error. The results indicate that binocular rivalries and related visual discomfort are the dominant factors. Another aspect of the paper is a comparison of the test results with results from different labs to evaluate the repeatability of a subjective experiment in the 3D case, and to compare the employed test methodologies. Here, the study shows the variation between observers when they are rating visual discomfort and illustrates the difficulty to evaluate this new dimension. Pierre R. Lebreton, Alexander Raake, Marcus Barkowsky, Patrick Le Callet |
ISM | 2 |
| 2010 | Extension of the E-model towards super-wideband speech transmissionabstractIn this paper, the quality gain of super-wideband (SWB) speech, transmitted in the much wider frequency range of 50-14000 Hz compared to the standard 300-3400 Hz narrowband, is quantified employing the E-model framework, a parametric tool for speech quality prediction. Based on two listening experiments, a linear extrapolation of the E-model transmission rating scale was found that leads to a maximum quality advantage of 39% relative to wideband (50-7000 Hz) transmission, and 79% relative to narrowband. Furthermore, narrowband, wideband, and super-wideband conditions can be quantified on this universal quality scale. Equipment Impairment Factors were derived and discussed for several SWB codecs. It will further be shown that a model quantifying the quality impact of linear distortions, reflected by the Bandwidth Impairment Factor, can successfully applied to SWB conditions. The correlation between the overall impairment and the model predictions amounts to r = 0.977 for linearly distorted speech samples. Marcel Wältermann, Izabela Tucker, Alexander Raake, Sebastian Möller 0001 |
ICASSP | 3 |
| 2010 | An intrusive super-wideband speech quality model: DIALabstractInternational audience Nicolas Côté, Vincent Koehl, Valérie Gautier-Turbin, Alexander Raake, Sebastian Möller 0001 |
INTERSPEECH | 4 |
| 2010 | Intelligibility predictions for speech against fluctuating maskerabstractThe effect of masking due to fluctuating sources on speech intelligibility is a phenomenon difficult to predict. Intelligibility scores vary with the efficiency of the energetic masking while the linguistic content of the message and listener’s cognitive performances add to the general incertitude that peaks for the case of masking speech. The present contribution proposes a signal-based assessment of the energetic masking at the sentence level. A mapping onto the scale of the speech intelligibility index is established for stationary noise. Predictions are quantitatively compared with the results of an intelligibility test for speech-modulated noise. The model is independent of voices similarities and semantic features, two important sources of informational masking. Index Terms: speech intelligibility, automatic speech Juan-Pablo Ramirez, Hamed Ketabdar, Alexander Raake |
INTERSPEECH | 3 |
| 2010 | Analytical assessment and distance modeling of speech transmission quality
Marcel Wältermann, Alexander Raake, Sebastian Möller 0001 |
INTERSPEECH | 2 |
| 2009 | Audio and video channel impact on perceived audio-visual quality in different interactive contextsabstractWith the advent of audio-visual IP clients, video telephony becomes a realistic option in many application scenarios. In order to guarantee an adequate quality to its users, providers of audio-visual telephony services need to know the impact of the audio and video transmission channel characteristics on perceived Quality of Experience (QoE) in a realistic interactive setting. For this aim, a conversational video telephony experiment was conducted where the audio and video channel settings were adjusted in a controlled way, and participants were asked about the perceived audio, video and overall quality after carrying out a conversation over the audio-visual channel. We analyze the results with respect to the impact the two modalities have, as well as with respect to the impact of the conversation scenario. Benjamin Belmudez, Sebastian Möller 0001, Blazej Lewcio, Alexander Raake, Muhammad Amir Mehmood |
MMSP | 4 |
| 2008 | Towards content-related features for parametric video quality prediction of IPTV servicesabstractThis paper investigates video content-related features, such as measures of spatio-temporal complexity, for inclusion into parametric video quality models. Our goal is to find a parametric content description that correlates with perceived video quality. In the course of the development of a parametric IPTV video quality prediction model (T-V-model), a large number of subjective tests have been conducted for standard definition and high definition video with different types of content. As expected from previous studies, we observed content dependencies that were different for different types of degradations. As descriptors of the content, we employ spatio-temporal related information obtained either before encoding and from the decoder or obtained from the decoder only. We compare those two approaches and explore their application to a reduced- or no-reference parametric model. An outlook highlights future steps for integrating the spatio- temporal features into the parametric model. Marie-Neige Garcia, Alexander Raake, Peter List 0001 |
ICASSP | 2 |
| 2008 | A comparative study of perceptional quality between wavefield synthesis and Multipole-Matched Rendering for spatial audioabstractThis paper introduces a new algorithm to render virtual sound sources with spatial properties in immersive environments. The algorithm, referred to as multipole-matched rendering, uses the method-of-moments and singular-value decomposition to optimally match a spherical-multipole expansion of the virtual source to the field resulting from a spatially distributed speaker set. The flexibility of this method over other approaches, such as wavefield synthesis, allows for complex speaker geometries, and requires a smaller number of speakers to achieve a similar spatial rendering performance for listeners in immersive environments. The trade-off for the enhanced performance is a smaller area of faithful reproduction. This limited area, however, can be focused around listener locations for a sweet-spot solution. Experimental results are presented from perceptual tests comparing multipole-matched rendering to both wavefield synthesis and stereo rendering using a linear speaker array. The experiments included 13 subjects and demonstrated that the perceived direction of a virtual sound source for the new method is comparable to that of wavefield synthesis (no significant difference). The results demonstrate the potential of multipole-matched rendering as an efficient technique for rendering virtual sound sources in immersive environments. Jens Hannemann, Christopher A. Leedy, Kevin D. Donohue, Sascha Spors, Alexander Raake |
ICASSP | 5 |
| 2008 | T-V-model: Parameter-based prediction of IPTV qualityabstractThe paper presents a parameter-based model for predicting the perceived quality of transmitted video for IPTV applications. The core model we derived can be applied both to service monitoring and network or service planning. In its current form, the model covers H.264 and MPEG-2 coded video (standard and high definition) transmitted over IP-links. The model includes factors like the coding bit-rate, the packet loss percentage and the type of packet loss handling used by the codec. The paper provides an overview of the model, of its integration into a multimedia model predicting audio-visual quality, and of its application to service monitoring. A performance analysis is presented showing a high correlation with the results of different subjective video quality perception tests. An outlook highlights future model extensions. Alexander Raake, Marie-Neige Garcia, Sebastian Möller 0001, Jens Berger, Fredrik Kling, Peter List 0001, Jens Johann, Cornelius Heidemann |
ICASSP | 1 |
| 2008 | Towards a new E-Model impairment factor for linear distortion of narrowband and wideband speech transmissionabstractThe e-model, a tool for network planning recommended by the ITU-T, suffers from the lack of predicting linear distortions as they may occur from channel filtering, codecs, and user interfaces. In order to face this deficiency, a new impairment factor is introduced in this paper which is estimated on the basis of two simple parameters from the linear portion of a system. Therewith, conventional Equipment Impairment Factors of narrowband and wideband speech codecs are decomposed into a linear and a residual part, allowing to quantify both magnitudes separately on the so-called R-scale. By extending the concept of distortion classes in the e-model, the proposed scheme provides a plausible picture of the linear effect of transmission systems. The advantage of wideband causing a 36% quality gain is well reflected. Further, the decomposition leads to a reduction of error when impairment factors are added on the R-scale. Examples for instrumentally estimating the residual impairment are given. Marcel Wältermann, Alexander Raake |
ICASSP | 2 |
| 2008 | An instrumental measure for end-to-end speech transmission quality based on perceptual dimensions: framework and realization
Marcel Wältermann, Kirstin Scholz, Sebastian Möller 0001, Lu Huo, Alexander Raake, Ulrich Heute |
INTERSPEECH | 5 |
| 2007 | Concept and evaluation of a downward-compatible system for spatial teleconferencing using automatic speaker clusteringabstractIn multi-party teleconferencing, the transport of separate speech streams to a particular user and the subsequent spatial rendering of the different streams enables a more efficient communication. A simple means of spatial presentation at client side is that of binaural rendering and headphone presentation. For downwardcompatibility, e.g. when the transport mechanism does not support multiple parallel downlink streams, a system is proposed that combines an automatic speaker classification mechanism with a spatial rendering of the segregated streams. The combined system aims at a better separability of the speakers than conventional systems. The paper details the two basic components, namely automatic speaker classification, and binaural rendering. Based on a first evaluation of the approach, a proof of concept is provided, and directions for further improvement are discussed. Alexander Raake, Sascha Spors, Jens Ahrens, Jitendra Ajmera |
INTERSPEECH | 1 |
| 2006 | Memo: towards automatic usability evaluation of spoken dialogue services by user error simulationsabstractProper usability evaluations of spoken dialogue systems are costly and cumbersome to carry out. In this paper, we present a new approach for facilitating usability evaluations which is based on user error simulations. The idea is to replace real users with simulations derived from empirical observations of users ’ erroneous behavior. The simulated errors must cover both system-driven errors (e.g., due to poor speech recognition) as well as conceptual errors and slips of the user, because neither alone is predictive of perceived usability. The simulation is integrated into a workbench which produces reports of typical and rare errors, and which allows usability ratings to be predicted. If successful, this workbench will help designers in making choices between system versions and lower testing costs at early phases of development. Challenges to the approach are discussed and solutions proposed. Index Terms: spoken-dialogue system, evaluation, usability 1. Sebastian Möller 0001, Roman Englert, Klaus-Peter Engelbrecht, Verena V. Hafner, Anthony Jameson, Antti Oulasvirta, Alexander Raake, Norbert Reithinger |
INTERSPEECH | 7 |
| 2006 | Estimation of the quality dimension "directness/frequency content" for the instrumental assessment of speech quality
Kirstin Scholz, Marcel Wältermann, Lu Huo, Alexander Raake, Sebastian Möller 0001, Ulrich Heute |
INTERSPEECH | 4 |
| 2006 | Underlying quality dimensions of modern telephone connectionsabstractIt is the aim of the present paper to analyze the perceptual quality dimensions of modern telephone connections. Such connections differ from standard connections in their timevariant characteristics (e.g., due to Voice-over-IP transmission or due to noise reduction algorithms) and their user interfaces (e.g., hands-free terminals). With the help of two independent auditory experiments with subsequent multidimensional analyses, three perceptual dimensions were identified for a diverse set of stimuli. These dimensions were labeled “directness/frequency content”, “continuity”, and “noisiness”. Overall listening quality scores were collected in a separate experiment. A mapping of the obtained dimensions onto the overall listening quality scores by means of a linear model revealed that “continuity” appears to be the most important dimension in terms of overall listening quality. Index Terms: assessment and modeling of speech quality, quality dimensions, multidimensional analyses Marcel Wältermann, Kirstin Scholz, Alexander Raake, Ulrich Heute, Sebastian Möller 0001 |
INTERSPEECH | 3 |
| 2006 | A joint intelligibility evaluation of French text-to-speech synthesis systems: the EvaSy SUS/ACR campaign
Philippe Boula de Mareüil, Christophe d'Alessandro, Alexander Raake, Gérard Bailly, Marie-Neige Garcia, Michel Morel |
LREC | 3 |
| 2006 | US-based Method for Speech Reception Threshold Measurement in French
Alexander Raake, Brian F. G. Katz |
LREC | 1 |
| 2006 | Impairment Factor Framework for Wide-Band Speech CodecsabstractA new method is described for quantifying the quality degradation introduced by wide-band speech codecs via a one-dimensional impairment factor. The method is based on auditory listening-only tests, but the resulting impairment factors may be used for predicting speech quality in an instrumental way, e.g., for network planning purposes. Following the method, auditory test results are first transformed to an overall quality rating scale, and then adjusted to rule out test-specific effects. The derived impairment factors fit into the common framework which is defined by the E-model for narrow-band telephone networks, and which is hereby extended towards wide-band speech transmission. This paper presents the necessary auditory test data, describes the derivation and adjustment methodology, and provides numerical values for a range of wide-band speech codecs. The values are tested for their robustness in case of codec tandems and adjusted to represent the effects of packet loss Sebastian Möller 0001, Alexander Raake, Nobuhiko Kitawaki, Akira Takahashi 0001, Marcel Wältermann |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Short- and Long-Term Packet Loss Behavior: Towards Speech Quality Prediction for Arbitrary Loss DistributionsabstractA speech-quality-oriented classification of packet loss distributions is proposed according to both the short- and long-term loss behavior. While the short-term behavior (microscopic loss behavior) relates to the effect of packet loss on the coder and packet loss concealment performance, the long-term loss behavior (macroscopic loss behavior) is defined so that it reflects the loss behavior that ultimately leads to speech quality that perceptively changes over time. Based on this classification, different parametric (objective) modeling approaches for predicting speech quality are discussed. To this aim, a packet loss averaging approach is presented for modeling speech quality under short-term loss. Starting from this model, two different ways for predicting speech quality under long-term-dependent packet loss are analyzed and compared to auditory (subjective) test results: quality prediction based on the averaging at packet trace level as provided, for example, by the E-model (2005), and the prediction based on the time-averaging of estimated instantaneous quality profiles, as suggested, for example, by L. Gros and N. Chateau (2001) (1998). From this comparison, the suitability of the different approaches for network planning are discussed, and their limitations in case of particular loss distributions are pointed out Alexander Raake |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | The well-tempered conversation: interactivity, delay and perceptual VoIP qualityabstractThe factors causing perceptual quality impairment on voice-over-IP (VoIP) connections include traditional network quality-of-service (QoS) parameters like packet loss rate, delay or jitter as well as parameters characterizing the conversation itself. Among the latter ones, we focus on the impact of "conversational interactivity" on the perceptual quality of a phone conversation. We introduce "parametric conversation analysis" as a formal framework for the instrumental investigation of conversational parameters at different transmission delay conditions, we further present the notion of "conversational temperature" as an intuitive scalar metric for the interactivity of conversations, and we demonstrate the application of our methods to a set of conversation recordings performed under various delay conditions, also with respect to results of subjective quality ratings. Florian Hammer, Peter Reichl, Alexander Raake |
ICC | 3 |
| 2004 | Elements of interactivity in telephone conversationsabstractThe term “interactivity” has been defined in numerous ways in the context of communications, but a definition of interactivity as an instrumentally measurable parameter of conversations is still missing. In this paper, we approach this issue by applying a parametric analysis to telephone conversations recorded during speech quality tests. To this end, we extract the basic conversational parameters like speech activity, mutual silence and double talk as well as a set of conversation events like speaker alternation rate and interruption rate. Comparing two types of scenarios for conversational speech quality assessment and exploring four different situations with regard to transmission delay, we aim at understanding the interdependencies between the conversational parameters which are basic for studying interactivity. This is intended to, ultimately, lead to an instrumental metric distilled from the relevant parameters. Florian Hammer, Peter Reichl, Alexander Raake |
INTERSPEECH | 3 |
| 2004 | Speech input and output module assessment for remote access to a smart-home spoken dialog system
Jan Felix Krebber, Sebastian Möller 0001, Alexander Raake |
INTERSPEECH | 3 |
| 2004 | Performance of speech recognition and synthesis in packet-based networks
Sebastian Möller 0001, Jan Felix Krebber, Alexander Raake |
INTERSPEECH | 3 |
| 2004 | INSPIRE: Evaluation of a Smart-Home System for Infotainment Management and Device Control
Sebastian Möller 0001, Jan Felix Krebber, Alexander Raake, Paula M. T. Smeele, Martin Rajman, Mirek Melichar, Vincenzo Pallotta, Gianna Tsakou, Basilis Kladis, Anestis Vovos, Jettie Hoonhout, Dietmar Schuchardt, Nikos Fakotakis, Todor Ganchev, Ilyas Potamitis |
LREC | 3 |
| 2002 | Does the Content of Speech Influence its Perceived Sound Quality?
Alexander Raake |
LREC | 1 |
| 2002 | Telephone speech quality prediction: Towards network planning and monitoring models for modern network scenarios
Sebastian Möller 0001, Alexander Raake |
Speech Commun. | 2 |
| 2000 | New models predicting conversational effects of telephone transmission on speech communication quality
Sebastian Möller 0001, Ute Jekosch, Alexander Raake |
INTERSPEECH | 3 |
| 2000 | Perceptual dimensions of speech sound quality in modern transmission systems
Alexander Raake |
INTERSPEECH | 1 |