Rakesh Rao Ramachandra Rao

dblp:243/6964 · DBLP profile ↗
← Back
37ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0002-7069-1543ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 9 first-author · 28 since 2021Human-computer interaction and ubiquitous computing · 21 · 5 first-author · 17 since 2021
YearPublicationVenuePosition
2026 Analysis of Appeal, Quality, and Realism for Instagram-Like Filtered Real Photos
Steve Goering, William Menz, Rakesh Rao Ramachandra Rao
QoMEX3
2026 How Accurate are Video Quality Models for Diffusion-Based Video Super-Resolution?
abstract
Recent video super-resolution (VSR) approaches use deep neural networks to enhance low-quality input videos and recover visual detail, with diffusion-based methods in particular showing promising results. In this paper, we investigate whether existing video quality models can be used to assess the performance of these diffusion-based VSR methods, by comparing model predictions with results from a subjective test. The study compares six upscaling methods (Lanczos, Rhea, SCST, DOVE, SeedVR2, Starlight Mini) applied to both compressed (AV1 and DCVC-RT) and uncompressed low-resolution videos considering the play-out on a UHD-1/4K screen. A range of full- and no-reference quality models are used to assess their applicability to this new type of quality degradation, focusing on within-sequence performance. The results highlight that CNN-based full-reference models, such as LPIPS, DISTS, and CVQA-FR show significantly higher correlation coefficients than both conventional full- as well as the tested no-reference models. Most overestimate the overly sharp results of SCST, with VMAF mainly failing due to spatial inconsistencies introduced by Starlight Mini. None of the tested video quality models reach sufficient accuracy so as to replace complementary subjective testing. The reference, degraded and upscaled videos, as well as the user ratings and model scores are made available with the paper at https://github.com/Telecommunication-Telemedia-Assessment/AVT-VQDB-UHD-1-VSR as open data.
Benjamin Herb, Steve Goering, Alexander Raake, Rakesh Rao Ramachandra Rao
QoMEX4
2026 A Comparative User Study of Real-Time Head-Appearance Telepresence in 2D and 3D Representations
William Menz, Alexander Zoubarev, David Kutschke, Rakesh Rao Ramachandra Rao, Louay Bassbouss, Sven Bliedung von der Heide, Alexander Raake, Steve Goering
QoMEX4
2025 Exploiting LLMs for Metadata-Based Video Quality Prediction
abstract
Large language models (LLMs) can be used to solve various tasks based on text inputs, e.g., video quality estimation. We explore the usage of LLMs for video quality prediction based on metadata (video codec, bitrate, resolution), which has not been addressed before. For the evaluation we use test #1 from our AVT-VQDB-UHD-1 dataset. We generated text prompts based on the metadata and collected answers from 17 different LLMs. The evaluation indicates that especially larger LLMs could be used to simulate human raters. However, a pure model prediction with one model has lower performance than SoA video quality models. Thus, we further investigated combinations of LLMs, which resulted in comparable performance to state-of-the-art models. Our work is a proof-of-concept, considering that LLMs are slower for the prediction than the traditional metadata-based models.
Steve Goering, Rakesh Rao Ramachandra Rao, Alexander Raake
ISM2
2025 Smarter Traps: Neural Network-Driven Classification of Small Mammals
abstract
The intensification of agricultural practices has substantially altered ecosystem structures, affecting small mammal populations through habitat homogenization and insecticide use. Traditional monitoring of these species relies on physical trapping, which is labor-intensive and stressful for the animals. To address these challenges, this study explores an automated image-based approach for monitoring small mammals using open traps equipped with cameras. The goal was to design a system that runs efficiently on standard, low-performance computers. Therefore, existing convolutional neural network (CNN) models were retrained and evaluated for automated genus classification of captured images. The networks were retrained with a transfer learning approach using a custom dataset containing four categories, three mammal genera, and one “empty trap” class. Multiple CNN architectures were compared based on loss, accuracy, and macro F1-score to identify a model that balances performance and computational efficiency. Results showed that MobileNet-based architectures, optimized for low-power devices, underperformed in this classification task, while VGG-based networks achieved superior accuracy and generalization to unseen images from the same trap setup. The findings demonstrate the potential of CNN-driven image recognition as a scalable and noninvasive tool for ecological monitoring, reducing manual review effort and improving animal welfare in field studies.
William Menz, Ralf Dittrich, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake
ISM3
2025 Evaluation of a Floating-Head Communication Prototype for Video-Conferencing
abstract
Besides traditional 2D video communication approaches, new systems aim to create more realistic and immersive representations of the involved conversation partners. This work presents a real-time floating-head communication setup that enables spatially separated visualizations of remote participants using only a standard camera and display. The underlying reconstruction pipeline applies machine-learning-assisted facial feature extraction to infer a 3D mesh of the participant's head from a live 2D video stream. Texturing maintains visual fidelity while supporting real-time performance. A dedicated transmission pipeline enables the exchange of 3D and texture data over conventional network connections, allowing flexible and location-independent use. A first lab test with eight participant pairs was performed to evaluate the system during a collaborative communication task. Subjective assessments using established telepresence and quality questionnaires confirmed the technical feasibility of the approach and its potential to enhance the sense of spatial presence. However, the overall perceptual quality and comfort did not yet reach the level of classical 2 D video communication. The study demonstrates the promise of accessible, spatially expressive communication setups that may bridge the gap between conventional video calls and emerging volumetric telepresence systems.
William Menz, Alexander Zoubarev, David Kutschke, Rakesh Rao Ramachandra Rao, Louay Bassbouss, Sven Bliedung von der Heide, Steve Goering, Alexander Raake
ISM4
2025 AMIS: An Audiovisual Dataset for Multimodal XR Research
abstract
The Audiovisual Multimodal Interaction Suite (AMIS) is an open-source dataset and accompanying Unity-based demo implementation designed to aid research on immersive media communication and social XR environments. AMIS features synchronized audiovisual recordings of three actors performing monologues and participating in dyadic conversations across four modalities: talking-head videos, full-body videos, volumetric avatars, and personalized animated avatars. These recordings can be used to simulate scenarios such as traditional video conferences or XR meetings with 3D avatars in controlled and replicable environments. The limitations of existing datasets, which include a restricted number of audiovisual formats, a narrow application focus, and suboptimal inclusion of verbal and non-verbal cues, are addressed by AMIS. With AMIS Studio, a Unity-based demonstrator, researchers can explore the recordings and compare the different audiovisual formats in VR scenes. This paper outlines the creation of AMIS, its design considerations, and how it may be applied in interdisciplinary domains, including cognitive psychology, audiovisual quality assessment, and social XR research.
Abhinav Bhattacharya, Luís Fernando de Souza Cardoso, Andy Schleising, Gareth Rendle, Adrian Kreskowski, Felix Immohr, Rakesh Rao Ramachandra Rao, Wolfgang Broll, Alexander Raake
MMSys7
2025 Evaluating Video Quality Metrics for Neural and Traditional Codecs using 4K/UHD-1 Videos
Benjamin Herb, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake
PCS2
2025 Sensory Evaluation of HDR Display Properties
Julius Prenzel, Dominik Keller, Rakesh Rao Ramachandra Rao, Alexander Raake
PCS3
2025 Fine-Grained HDR Image Quality Assessment From Noticeably Distorted to Very High Fidelity
abstract
High dynamic range (HDR) and wide color gamut (WCG) technologies significantly improve color reproduction compared to standard dynamic range (SDR) and standard color gamuts, resulting in more accurate, richer, and more immersive images. However, HDR increases data demands, posing challenges for bandwidth efficiency and compression techniques. Advances in compression and display technologies require more precise image quality assessment, particularly in the high-fidelity range where perceptual differences are subtle. To address this gap, we introduce AIC-HDR2025, the first such HDR dataset, comprising 100 test images generated from five HDR sources, each compressed using four codecs at five compression levels. It covers the high-fidelity range, from visible distortions to compression levels below the visually lossless threshold. A subjective study was conducted using the JPEG AIC-3 test methodology, combining plain and boosted triplet comparisons. In total, 34,560 ratings were collected from 151 participants across four fully controlled labs. The results confirm that AIC-3 enables precise HDR quality estimation, with 95% confidence intervals averaging a width of 0.27 at 1 JND. In addition, several recently proposed objective metrics were evaluated based on their correlation with subjective ratings. The dataset is publicly available1.
Mohsen Jenadeleh, Jon Sneyers, Davi Lazzarotto, Shima Mohammadi, Dominik Keller, Atanas Boev, Rakesh Rao Ramachandra Rao, António M. G. Pinheiro, Thomas Richter 0005, Alexander Raake, Touradj Ebrahimi, João Ascenso, Dietmar Saupe
QoMEX7
2025 A Large-Scale Evaluation of Subject Rating Behaviour in Visual Quality Assessment Studies
abstract
Subjective testing is widely used for visual quality assessment to evaluate the impact of both technical and non-technical factors on user perception. Although standardized methods exist for collecting subjective ratings of visual quality, these ratings are inevitably influenced by each subject’s accuracy, manifesting as subject bias and inconsistency. Recommendations such as ITU-T P.910 and ITU-R BT.500 propose standardized methods to remove bias from subjective ratings. For instance, Annex E of ITU-T Rec. P.910 provides an effective strategy for addressing both subject bias and inconsistency. In this paper, we analyze 29 different visual quality assessment studies conducted over an eight-year period to understand the rating behaviour of subjects using these methods. Our investigation focuses on subjective studies targeting 4K, 8K, as well as 360°video and high-resolution image quality assessment. In the context of 4K video quality assessment, both SDR and HDR evaluations are considered. Our dataset covers use cases of short-term video quality and overall session quality assessment for HTTP-based adaptive streaming (HAS). For both these use cases, we propose a range of subject bias and inconsistency values that can serve as a reference for future studies. Furthermore, we define six different measures that can be used to assess the reliability of future studies and provide reference values for these measures. Following an open-science approach, all individual ratings from the included subjective tests, along with the results of the large-scale analysis, are made publicly available with this paper.
Rakesh Rao Ramachandra Rao, Steve Goering, Stephan Fremerey, Dominik Keller, Alexander Raake
QoMEX1
2024 Evaluating Visually Lossless Compression of JPEG XS, JPEG 2000, HEVC and AV1 in Selected Medical Imaging Modalities
abstract
The objective of this study is to evaluate the effectiveness of state-of-the-art codecs in compressing selected medical imaging modalities while maintaining visual quality and reducing file sizes. To achieve this, a detailed comparative analysis is conducted comparing the performance of JPEG XS, JPEG 2000, HEVC, and AV1. The analysis takes into consideration compression efficiency, codec complexity, and visual fidelity in the context of medical imaging. Advanced evaluation methods, including the AIC-2 Flicker test, are utilized to determine the visually lossless threshold, which is crucial for preserving diagnostically important details. Additionally, the study explores the potential of crowd-sourcing as a means of assessing the visual quality of compressed medical images. Subjective lab and crowd-sourcing tests reveal varying proportions of correctly identifying the reference images among participants. Furthermore, the study proposes outlier detection methods to improve the reliability of the subjective evaluation and employs kappa analysis to measure the inter-rater agreements. The study analyzes the encoding time taken on a consumer-level CPU, and the results reveal that JPEG XS maintains a fast and consistent speed across different compression levels. The results also indicate that JPEG XS achieves visually lossless performance for diagnostic purposes at 2 BPP, JPEG 2000 at 1.5 BPP, HEVC, and AV1 at 1 BPP.
Bassem Elmeligy, Thomas Richter 0005, Rakesh Rao Ramachandra Rao, Siegfried Fößel, Alexander Raake
QoMEX3
2024 AVT-VQDB-UHD-1-HDR: An Open Video Quality Dataset for Quality Assessment of UHD-1 HDR Videos
abstract
High dynamic range (HDR) videos offer users a more realistic viewing experience owing to their ability to represent a wider and thus more natural range of brightness. This has resulted in an increase in HDR content streamed on different video streaming platforms. Hence, it becomes important to have a proper understanding of the perceived quality of HDR videos when encoded with modern video codecs such as H.265, AV1, and VVC that are either commonly used or will potentially be used by video streaming providers. With this objective, in this paper, we present a study that used both subjective and instrumental methods to assess the perceived quality of HDR videos. Firstly, a subjective test with 4K/UHD-1 HDR videos using the ACR-HR (Absolute Category Rating -Hidden Reference) method was conducted. The tests consisted of a total of 195 encoded videos from 5 source videos which all had a framerate of 60 fps. In this test, the 4K/UHD-1 HDR stimuli were encoded at four different resolutions, namely, 720p, 1080p, 1440p, and 2160p using bitrates ranging between 0.5 Mbps and 40 Mbps. The results of the subjective test have been analyzed to assess the impact of factors such as resolution, bitrate, video codec, and content on the perceived video quality. As automated quality assessment forms an important part of the encoding ecosystem of any video streaming platform to decide the optimal encoding settings, different full reference, bitstream, and hybrid instrumental models have been evaluated for their applicability for HDR video quality prediction. The database of source content, encoded videos, subjective and objective scores is made publicly available with this paper following an open-science approach, accessible at: https://github.com/Telecommunication-Telemedia-Assessment/AVT-VQDB-UHD-1-HDR.
Rakesh Rao Ramachandra Rao, Benjamin Herb, Helmi-Aurora Takala, Mohamed Tarek Mohamed Ahmed, Alexander Raake
QoMEX1
2023 Towards evaluation of immersion, visual comfort and exploration behaviour for non-stereoscopic and stereoscopic 360° videos
abstract
Immersion, visual comfort, and exploration behaviour are important aspects that affect the overall quality of experience for 360° videos. To analyze the benefits of stereoscopic and non-stereoscopic 360° videos in terms of these factors, we created a dataset and conducted a subjective study. The dataset consists of five different high-resolution $8 \mathrm{~K}$ omnidirectional videos as stereoscopic and non-stereoscopic variants. The videos have been recorded using a Kandao Obsidian Pro camera. For the comparison, we designed and performed a subjective test with 30 participants. Here, each subject watched both HEVC (libx265) encoded versions of the source video and rated the videos viewed regarding presence, visual comfort, and quality. The results indicate that with the test protocol followed, non-stereoscopic video viewing leads to slightly better presence, visual comfort, and quality ratings compared to the stereoscopic variants. Further, the stereoscopic 360° videos may suffer from visual artefacts potentially leading to lower video quality and further lower quality of experience results. The exploration behaviour was found to be very similar for both non-stereoscopic and stereoscopic video viewing. Overall, it can be concluded that there is a slight tendency for non-stereoscopic video viewing to be preferred over stereoscopic video viewing. The dataset is made publicly available with the paper and includes both variants of all source videos along with the subjective data, and behaviour data, following an open-science approach.
Stephan Fremerey, Raja Faseeh Uz Zaman, Touseef Ashraf, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake
ISM4
2023 The Effect of Viewing Distances on 4K and 8K HDR Video Quality Perception
abstract
Ongoing research in the field of capture, coding and display technology and human vision has explored the advantages of high resolution up to 8K (UHD-2) considering perceived quality. One of the crucial elements impacting users’ perception of video quality is the viewing distance. As a result, the presented study employs a subjective evaluation to investigate the perceptual benefits offered by 8K or upscaled 4K in comparison to the native 4K (UHD-1) resolution in the context of HDR videos. The subjective test uses 7 distinct viewing distances, ranging from 0.5H to 3H, with H representing the display height. The findings of the study reveal a consistent trend: the increased video quality of 8K HDR against 4K HDR content decreases with distance, on average. While there are bigger improvements for close distances, beyond 2H the quality difference was very little or zero, depending on content. In general, the degree of enhancement is contingent on the spatial complexity of the content. Additionally, it is found that, on average, subjects prefer to sit at a distance of 2.07H. No significant difference in the preferred viewing distance was found when asked before and after the study.
Dominik Keller, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake
ISM2
2023 Adaptation of Bitstream-based Video Quality Models for Image Quality Assessment
abstract
In recent years, video-codec-based image codecs, such as e.g. HEF, AVF, etc., have been increasingly used to compress images. Hence, there is a potential to use video quality prediction models for the evaluation of image quality. Bitstream-based models show promising results for video quality prediction, therefore, we investigate the applicability of such models for the case of image quality in this paper. For this purpose, we selected ITU-T Rec. P.1204.3 and its Mode 0 variant also known as $AVQBits|M3$ and $AVQBits|M0$ respectively for the evaluation, because they are computationally less complex and do not need a reference image. These models are evaluated using a publicly available dataset consisting of a total of 371 images of resolutions between $144\times 144$ pixels to $2160\times 2160$ pixels with subjective annotations. The results show that both the considered models perform well on the used dataset with a Pearson correlation of 0.958 and Root Mean Square Error (RMSE) of 0.319 (on a 1 to 5 Absolute Category Rating (ACR) scale) for the $AVQBits|M3$ model and a Pearson correlation of 0.942 and RMSE of 0.377 for the $AVQBits|M0$ model.
Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake
ISM1
2023 AVT-VQDB-UHD-1-Appeal: A UHD-1/4K Open Dataset for Video Quality and Appeal Assessment Using Modern Video Codecs
abstract
A number of factors play an important role in the perception of video quality for streaming and other services, key among them being encoding-related degradations. Hence, newer codecs are developed with the goal of optimizing video quality for a given encoding setting. Here, subjective studies are an efficient method to evaluate the performance of such newer codecs. Furthermore, contextual factors impact the perception of video quality, e.g., the appeal of the content itself. To this end, this paper presents a subjective study targeting both quality and appeal assessment of videos. For this purpose, a subjective study consisting of three different parts is conducted. Firstly, participants were asked to rate the appeal of the uncompressed UHD-1/4K source content with a duration of 8 - 10s each. Following this, the video quality of these source videos individually encoded with either the HEVC/H.265, AV1, or VVC/H.266 video codec was rated. A wide range of encoding conditions in terms of resolution (360p to 2160p) and bitrate (100kbps to 15mbps) is used to encode the videos, so as to enable the applicability of the data to real-world settings. In the last part, subjects are again asked to rate the appeal of the uncompressed source content. The results are analyzed to assess the impact of different encoding conditions on perceived video quality. In addition, the impact of appeal on video quality and vice-versa is also investigated. Furthermore, an objective quality assessment with different state-of-the-art full-reference, bitstream-based, and hybrid models including the newer codecs AV1 and VVC is presented. The subjective dataset including test design, subjective results, sources, and encoded audiovisual contents are made publicly available following an open science approach.
Rakesh Rao Ramachandra Rao, Steve Goering, Bassem Elmeligy, Alexander Raake
MMSP1
2023 Automatic Audiovisual Asynchrony Measurement for Quality Assessment of Videoconferencing
abstract
Audiovisual asynchrony is a significant factor im-pacting the Quality of Experience (QoE), especially for interactive communication like video conferencing. In this paper, we propose a client-side approach to predict the delay between an audio and a video signal, using only the media signals from both streams. Features are extracted from the video and audio stream, respectively, and analyzed using a cross-correlation approach to determine the actual delay. Our approach predicts the delay with an accuracy of over 80% in a time frame of ±1s. We further highlight the potential drawbacks of using a cross-correlation-based analysis and propose different solutions for practical implementations of a delay-based QoE metric.
Florian Braun, Rakesh Rao Ramachandra Rao, Werner Robitza, Alexander Raake
QoMEX2
2023 Revisiting Videoconferencing QoE: Impact of Network Delay and Resolution as Factors for Social Cue Perceptibility
abstract
Previous research from well before the Covid-19 pandemic had indicated little effect of delay on integral quality but a measurable one on user behavior, and a significant effect of resolution on quality but not on behavior in a two-party communication scenario. In this paper, we re-investigate the topic, after the times of the Covid-19 pandemic and its frequent and widespread videoconferencing usage. To this aim, we conducted a subjective test involving 23 pairs of participants, employing the Celebrity Name Guessing task. The focus was on impairments that may affect social (resolution) and communication cues (de-lay). Subjective data in the form of overall conversational quality and task performance satisfaction as well as objective data in the form of task correctness, user motion, and facial expressions were collected in the test. The analysis of the subjective data indicates that perceived conversational quality and performance satisfaction were mainly affected by video resolution, while delay (up to 1000 ms) had no significant impact. Furthermore, the analysis of the objective data shows that there is no impact of resolution and delay on user performance and behavior, in contrast to earlier findings.
Chenyao Diao, Luljeta Sinani, Rakesh Rao Ramachandra Rao, Alexander Raake
QoMEX3
2023 Appeal and quality assessment for AI-generated images
abstract
Recently AI-generated images gained in popularity. A critical aspect of AI-generated images using, e.g., DALL-E-2 or Midjourney, is that they may look artificial, be of low quality, or have a low appeal in contrast to real images, depending on the text prompt and AI generator. For this reason, we evaluate the quality and appeal of AI-generated images using a crowdsourcing test as an extension of our recently published AVT-AI-Image-Dataset. This dataset consists of a total of 135 images generated with five different AI-text-to-image generators. Based on the collected subjective ratings in the crowdsourcing test, we evaluate the different used AI generators in terms of image quality and appeal of the AI-generated images. We also link image quality and image appeal also with SoA objective models. The extension will be made publicly available for reproducibility.
Steve Goering, Rakesh Rao Ramachandra Rao, Rasmus Merten, Alexander Raake
QoMEX2
2023 Influence of Viewing Distances on 8K HDR Video Quality Perception
abstract
The benefits of high resolutions in displays, such as 8K (UHD-2), have been the subject of ongoing research in the field of display technology and human perception in recent years. Out of several factors influencing users' perception of video quality, viewing distance is one of the key aspects. Hence, this study uses a subjective test to investigate the perceptual advantages of 8K over 4K (UHD-1) resolution for HDR videos at 7 different viewing distances, ranging from 0.5 H to 2 H. The results indicate that, on average, for HDR content the 8K resolution can improve the video quality at all tested distances. Our study shows that although the 8K resolution is slightly better than 4K at close distances, the extent of these benefits is highly dependent on factors such as the pixel-related complexity of the content and the visual acuity of the viewers.
Dominik Keller, Felix von Hagen, Julius Prenzel, Kay Strama, Rakesh Rao Ramachandra Rao, Alexander Raake
QoMEX5
2023 PNATS-UHD-1-Long: An Open Video Quality Dataset for Long Sequences for HTTP-based Adaptive Streaming QoE Assessment
abstract
The P.NATS Phase 2 competition in ITU-T Study Group 12 resulted in both the ITU-T Rec. P.1204 series of recommendations, and also a large dataset for HTTP-based adaptive streaming QoE assessment that is now made openly available as part of this paper. The presented dataset consists of 3 subjective databases targeting overall quality assessment of a typical HTTP-based Adaptive Streaming session consisting of degradations such as quality switching, initial loading delay, and stalling events using audiovisual content ranging between 2 and 5 minutes. In addition to this, subject bias and consistency in quality assessment of such longer-duration audiovisual contents with multiple degradations are investigated using a subject behaviour model. As part of this paper, the overall test design, subjective test results, sources, encoded audiovisual contents, and a set of analysis plots are made publicly available for further research.
Rakesh Rao Ramachandra Rao, Silvio Borer, David Lindero, Steve Goering, Alexander Raake
QoMEX1
2023 Proof-of-Concept Study to Evaluate the Impact of Spatial Audio on Social Presence and User Behavior in Multi-Modal VR Communication
abstract
This paper presents a proof-of-concept study conducted to analyze the effect of simple diotic vs. spatial, position-dynamic binaural synthesis on social presence in VR, in comparison with face-to-face communication in the real world, for a sample two-party scenario. A conversational task with shared visual reference was realized. The collected data includes questionnaires for direct assessment, tracking data, and audio and video recordings of the individual participants’ sessions for indirect evaluation. While tendencies for improvements with binaural over diotic presentation can be observed, no significant difference in social presence was found for the considered scenario. The gestural analysis revealed that participants used the same amount and type of gestures in face-to-face as in VR, highlighting the importance of non-verbal behavior in communication. As part of the research, an end-to-end framework for conducting communication studies and analysis has been developed.
Felix Immohr, Gareth Rendle, Annika Neidhardt, Steve Goering, Rakesh Rao Ramachandra Rao, Stephanie Arevalo, Bernd Fröhlich 0001, Alexander Raake
IMX5
2021 AVrate Voyager: an open source online testing platform
abstract
Subjective testing is an integral part of many research fields considering, e.g., human perception. For this purpose, lab tests are a popular approach to gather ratings for subjective evaluations. However, not in all cases controlled lab tests can be performed, either in cases where no labs are existing, accessible or it may be disallowed to use them. For this reason, online tests, e.g., using crowdsourcing are supposed to be an alternative approach for traditional lab tests. We describe in the following paper a framework to implement such online tests for audio, video, and image-related evaluations or questionnaires. Our framework AVrate Voyager builds upon previously developed frameworks for lab tests including the experience with them. AVrate Voyager uses scalable web technologies to implement a test framework, this ensures that it will be running reliably. In addition, we added strategies for pre-caching to avoid additional influence for play-out, e.g. in the case of video testing. We analyze several conducted tests using the new framework and describe the required steps to modify the provided tool in detail.
Steve Goering, Rakesh Rao Ramachandra Rao, Stephan Fremerey, Alexander Raake
MMSP2
2021 Groovability: Using Groove as a Novel Measure for Audio QoE with the Example of Smartphones
abstract
Groove in music is a fundamental part of why humans entrain to it and enjoy it. Smartphones have become an important medium to listen to music. Especially when being with others, loudspeaker playback may be the method of choice. However, due to the physical limits of acoustics, for loudspeaker playback, smartphones are equipped with sub-optimal audio capabilities. Therefore, it is desirable to measure Quality of Experience (QoE) of music played on smartphones. While audio playback is often assessed in terms of sound quality, the aim of this work is to address QoE in terms of the meaning or effect that the audio has on the listener. A key component for the meaning of popular music is groove. Hence, in this paper, we study “groovability”, that is, the ability of a piece of audio technology to convey groove. To instantiate our novel audio QoE assessment method, we apply it to music played by 8 different smartphones. For this purpose, looped 4-bar loudness-aligned recordings from 24 music pieces of different intrinsic groove were played back on the different smartphones. Our test method uses a multi-stimulus comparison with synchronized playback capability. A total of 62 subjects evaluated groovability using two stimulus subsets. It was found that the proposed methodology is highly effective to distinguish between the groovability provided by the considered phones. In addition, a reduced-reference model is proposed to predict groovability, using a set of both acoustics-and music-groove related features. In our formal validation on unknown data, the model is shown to provide good prediction performance with a Pearson correlation of greater than 0.90.
Dominik Keller, Markus Vaalgamaa, Erkki Paajanen, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake
QoMEX4
2021 Towards High Resolution Video Quality Assessment in the Crowd
abstract
Assessing high resolution video quality is usually performed using controlled, defined, and standardized lab tests. This method of acquiring human ratings in a lab environment is time-consuming and may also not reflect the typical viewing conditions. To overcome these disadvantages, crowd testing paradigms have been used for assessing video quality in general. Crowdsourcing-based tests enable a more diverse set of participants and also use a realistic hardware setup and viewing environment of typical users. However, obtaining valid ratings for high-resolution video quality poses several problems. Example issues are that streaming of such high-bandwidth content may not be feasible for some users, or that crowd participants lack an appropriate, high-resolution display device. In this paper, we propose a method to overcome such problems and conduct a crowd test using for higher resolution content by using a 540 p cutout from the center of the original 2160p video. To this aim, we use the videos from Test#1 of the publicly available dataset AVT-VQDB-UHD-1, which contains videos up to a resolution of UHD-1. The quality-labels available from that lab test allow us to compare the results with the crowd test presented in this paper. It is shown that there is a Pearson correlation of 0.96 between the lab and crowd tests and hence such crowd tests can reliably be used for video assessment of higher resolution content. The overall implementation of the crowd test framework and the results are made publicly available for further research and reproducibility1.
Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake
QoMEX1
2021 Impact of Spatial and Temporal Information on Video Quality and Compressibility
abstract
Spatial Information (SI) and Temporal Information (TI) are frequently-used metrics to classify the spatiotemporal complexity of video content. However, they are mostly used on original video sources, and their impact on actual encoding efficiency is not known. In this paper, we propose a method to determine the compressibility of video sources, that is, how good video quality can be under a given bitrate constraint. We show how various aggregations of SI and TI correlate with compressibility scores obtained from a public dataset of H.264/HEVCN P9 content. We observe that the minimum TI value as well as an existing criticality metric from the literature are good indicators for compressibility, as judged by subjective ratings as well as VMAF and P.1204.3 objective scores.
Werner Robitza, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake
QoMEX2
2021 Assessment of the Simulator Sickness Questionnaire for Omnidirectional Videos
abstract
Virtual Reality/360° videos provide an immersive experience to users. Besides this, 360° videos may lead to an undesirable effect when consumed with Head-Mounted Displays (HMDs), referred to as simulator sickness/cybersickness. The Simulator Sickness Questionnaire (SSQ) is the most widely used questionnaire for the assessment of simulator sickness. Since the SSQ with its 16 questions was not designed for 360° video related studies, our research hypothesis in this paper was that it may be simplified to enable more efficient testing for 360° video. Hence, we evaluate the SSQ to reduce the number of questions asked from subjects, based on six different previously conducted studies. We derive the reduced set of questions from the SSQ using Principal Component Analysis (PCA) for each test. Pearson Correlation is analysed to compare the relation of all obtained reduced questionnaires as well as two further variants of SSQ reported in the literature, namely Virtual Reality Sickness Questionnaire (VRSQ) and Cybersickness Questionnaire (CSQ). Our analysis suggests that a reduced questionnaire with 9 out of 16 questions yields the best agreement with the initial SSQ, with less than 44% of the initial questions. Exploratory Factor Analysis (EFA) shows that the nine symptom-related attributes determined as relevant by PCA also appear to be sufficient to represent the three dimensions resulting from EFA, namely, Uneasiness, Visual Discomfort and Loss of Balance. The simplified version of the SSQ has the potential to be more efficiently used than the initial SSQ for 360° video by focusing on the questions that are most relevant for individuals, shortening the required testing time.
Ashutosh Singla, Steve Goering, Dominik Keller, Rakesh Rao Ramachandra Rao, Stephan Fremerey, Alexander Raake
VR4
2020 Subjective Test Dataset and Meta-data-based Models for 360° Streaming Video Quality
abstract
During the last years, the number of 360° videos available for streaming has rapidly increased, leading to the need for 360° streaming video quality assessment. In this paper, we report and publish results of three subjective 360° video quality tests, with conditions used to reflect real-world bitrates and resolutions including 4K, 6K and 8K, resulting in 64 stimuli each for the first two tests and 63 for the third. As playout device we used the HTC Vive for the first and HTC Vive Pro for the remaining two tests. Video-quality ratings were collected using the 5-point Absolute Category Rating scale. The 360° dataset provided with the paper contains the links of the used source videos, the raw subjective scores, video-related meta-data, head rotation data and Simulator Sickness Questionnaire results per stimulus and per subject to enable reproducibility of the provided results. Moreover, we use our dataset to compare the performance of state-of-the-art full-reference quality metrics such as VMAF, PSNR, SSIM, ADM2, WS-PSNR and WS-SSIM. Out of all metrics, VMAF was found to show the highest correlation with the subjective scores. Further, we evaluated a center-cropped version of VMAF ("VMAF-cc") that showed to provide a similar performance as the full VMAF. In addition to the dataset and the objective metric evaluation, we propose two new video-quality prediction models, a bitstream meta-data-based model and a hybrid no-reference model using bitrate, resolution and pixel information of the video as input. The new lightweight models provide similar performance as the full-reference models while enabling fast calculations.
Stephan Fremerey, Steve Goering, Rakesh Rao Ramachandra Rao, Rachel Huang, Alexander Raake
MMSP3
2020 Automated Genre Classification for Gaming Videos
abstract
Besides classical videos, videos of gaming matches, entire tournaments or individual sessions are streamed and viewed all over the world. The increased popularity of Twitch or YoutubeGaming shows the importance of additional research on gaming videos. One important pre-condition for live or offline encoding of gaming videos is the knowledge of game-specific properties. Knowing or automatically predicting the genre of a gaming video enables a more advanced and optimized encoding pipeline for streaming providers, especially because gaming videos of different genres vary a lot from classical 2D video, e.g., considering the CGI content, textures or camera motion. We describe several computer-vision based features that are optimized for speed and motivated by characteristics of popular games, to automatically predict the genre of a gaming video. Our prediction system uses random forest and gradient boosting trees as underlying machine-learning techniques, combined with feature selection. For the evaluation of our approach we use a dataset that was built as part of this work and consists of recorded gaming sessions for 6 genres from Twitch. In total 351 different videos are considered. We show that our prediction approach shows a good performance in terms of f1-score. Besides the evaluation of different machine-learning approaches, we additionally investigate the influence of the hyper-parameters for the algorithms.
Steve Goering, Robert Steger, Rakesh Rao Ramachandra Rao, Alexander Raake
MMSP3
2020 A Large-scale Evaluation of the bitstream-based video-quality model ITU-T P.1204.3 on Gaming Content
abstract
The streaming of gaming content, both passive and interactive, has increased manifolds in recent years. Gaming contents bring with them some peculiarities which are normally not seen in traditional 2D videos, such as the artificial and synthetic nature of contents or repetition of objects in a game. In addition, the perception of gaming content by the user is different from that of traditional 2D videos due to its pecularities and also the fact that users may not often watch such content. Hence, it becomes imperative to evaluate whether the existing video quality models usually designed for traditional 2D videos are applicable to gaming content. In this paper, we evaluate the applicability of the recently standardized bitstream-based video-quality model ITU-T P.1204.3 on gaming content. To analyze the performance of this model, we used 4 different gaming datasets (3 publicly available + 1 internal) not previously used for model training, and compared it with the existing state-of-the-art models. We found that the ITU P.1204.3 model out of the box performs well on these unseen datasets, with an RMSE ranging between 0.38 - 0.45 on the 5-point absolute category rating and Pearson Correlation between 0.85 - 0.93 across all the 4 databases. We further propose a full-HD variant of the P.1204.3 model, since the original model is trained and validated which targets a resolution of 4K/UHD-1. A 50:50 split across all databases is used to train and validate this variant so as to make sure that the proposed model is applicable to various conditions.
Rakesh Rao Ramachandra Rao, Steve Goering, Robert Steger, Saman Zad Tootaghaj, Nabajeet Barman, Stephan Fremerey, Sebastian Möller 0001, Alexander Raake
MMSP1
2020 DEMI: Deep Video Quality Estimation Model using Perceptual Video Quality Dimensions
abstract
Existing works in the field of quality assessment focus separately on gaming and non-gaming content. Along with the traditional modeling approaches, deep learning based approaches have been used to develop quality models, due to their high prediction accuracy. In this paper, we present a deep learning based quality estimation model considering both gaming and non-gaming videos. The model is developed in three phases. First, a convolutional neural network (CNN) is trained based on an objective metric which allows the CNN to learn video artifacts such as blurriness and blockiness. Next, the model is fine-tuned based on a small image quality dataset using blockiness and blurriness ratings. Finally, a Random Forest is used to pool frame-level predictions and temporal information of videos in order to predict the overall video quality. The light-weight, low complexity nature of the model makes it suitable for real-time applications considering both gaming and non-gaming content while achieving similar performance to existing state-of-the-art model NDNetGaming. The model implementation for testing is available on GitHub1.
Saman Zad Tootaghaj, Nabajeet Barman, Rakesh Rao Ramachandra Rao, Steve Goering, Maria G. Martini, Alexander Raake, Sebastian Möller 0001
MMSP3
2020 Prenc - Predict Number of Video Encoding Passes with Machine Learning
abstract
Video streaming providers spend huge amounts of processing time to get a quality-optimized encoding. While the quality-related impact may be known to the service provider, the impact on video quality is hard to assess, when no reference is available. Here, bitstream-based video quality models may be applicable, delivering estimates that include encoding-specific settings. Such models typically use several input parameters, e.g. bitrate, framerate, resolution, video codec, QP values and more. However, for a given bitstream, to determine which encoding parameters were selected, e.g., the number of encoding passes, is not a trivial task. This leads to our following research question: Given an unknown video bitstream, which encoding settings have been used? To tackle this reverse engineering problem, we introduce a system called prenc. Besides the use in video-quality estimation, such algorithms may also be used in other applications such as video forensics. We prove our concept by applying prenc to distinguish between one- and two-pass encoding. Starting from modeling the problem as a classification task, estimating bitstream-based features, we further describe a machine learning approach with feature selection to automatically predict the number of encoding passes for a given video bitstream. Our large-scale evaluation consists of 16 short movie type 4K videos that were segmented and encoded with different settings (resolutions, codecs, bitrates), so that we in total analyzed 131.976 DASH video segments. We further show that our system is robust, based on a 50% train and 50% validation approach without source video overlapping, where we get a classification performance of 65% F1 score. Moreover, we also describe the used bitstream-based features in detail, the feature pooling strategy and include other machine learning algorithms in our evaluation.
Steve Goering, Rakesh Rao Ramachandra Rao, Alexander Raake
QoMEX2
2020 Bitstream-Based Model Standard for 4K/UHD: ITU-T P.1204.3 - Model Details, Evaluation, Analysis and Open Source Implementation
abstract
With the increasing requirement of users to view high-quality videos with a constrained bandwidth, typically realized using HTTP-based adaptive streaming, it becomes more and more important to determine the quality of the encoded videos accurately, to assess and possibly optimize the overall streaming quality. In this paper, we describe a bitstream-based no-reference video quality model developed as part of the latest model-development competition conducted by ITU-T Study Group 12 and the Video Quality Experts Group (VQEG), “P.NATS Phase 2”. It is now part of the new P.1204 series of Recommendations as P.1204.3. It can be applied to bitstreams encoded with H.264/AVC, HEVC and VP9, using various encoding options, including resolution, bitrate, framerate and typical encoder settings such as number of passes, rate control variants and speeds. The proposed model follows an ensemble-modelling-inspired approach with weighted parametric and machine-learning parts to efficiently leverage the performance of both approaches. The paper provides details about the general approach to modelling, the features used and the final feature aggregation. The model creates per-segment and per-second video quality scores on the 5-point Absolute Category Rating scale, and is applicable to segments of 5–10 seconds duration. It covers both PC/TV and mobile/tablet viewing scenarios. We outline the databases on which the model was trained and validated as part of the competition, and perform an additional evaluation using a total of four independently created databases, where resolutions varied from 360p to 2160p, and frame rates from 15–60fps, using realistic coding and bitrate settings. We found that the model performs well on the independent dataset, with a Pearson correlation of 0.942 and an RMSE of 0.42. We also provide an open-source reference implementation of the described P.1204.3 model, as well as the multi-codec bitstream parser required to extract the input data, which is not part of the standard.
Rakesh Rao Ramachandra Rao, Steve Goering, Peter List 0001, Werner Robitza, Bernhard Feiten, Ulf Wüstenhagen, Alexander Raake
QoMEX1
2019 AVT-VQDB-UHD-1: A Large Scale Video Quality Database for UHD-1
abstract
4K television screens or even with higher resolutions are currently available in the market. Moreover video streaming providers are able to stream videos in 4K resolution and beyond. Therefore, it becomes increasingly important to have a proper understanding of video quality especially in case of 4K videos. To this effect, in this paper, we present a study of subjective and objective quality assessment of 4K ultra-high-definition videos of short duration, similar to DASH segment lengths. As a first step, we conducted four subjective quality evaluation tests for compressed versions of the 4K videos. The videos were encoded using three different video codecs, namely H.264, HEVC, and VP9. The resolutions of the compressed videos ranged from 360p to 2160p with framerates varying from 15fps to 60fps. All the source 4K contents used were of 60fps. We included low quality conditions in terms of bitrate, resolution and framerate to ensure that the tests cover a wide range of conditions, and that e.g. possible models trained on this data are more general and applicable to a wider range of real world applications. The results of the subjective quality evaluation are analyzed to assess the impact of different factors such as bitrate, resolution, framerate, and content. In the second step, different state-of-the-art objective quality models were applied to all videos and their performance was analyzed in comparison with the subjective ratings, e.g. using Netflix's VMAF. The videos, subjective scores, both MOS and confidence interval per sequence and objective scores are made public for use by the community for further research.
Rakesh Rao Ramachandra Rao, Steve Goering, Werner Robitza, Bernhard Feiten, Alexander Raake
ISM1
2019 nofu - A Lightweight No-Reference Pixel Based Video Quality Model for Gaming Content
abstract
Popularity of streaming services for gaming videos has increased tremendously over the last years, e.g. Twitch and Youtube Gaming. Compared to classical video streaming applications, gaming videos have additional requirements. For example, it is important that videos are streamed live with only a small delay. In addition, users expect low stalling, waiting time and in general high video quality during streaming, e.g. using http-based adaptive streaming. These requirements lead to different challenges for quality prediction in case of streamed gaming videos. We describe newly developed features and a no-reference video quality machine learning model, that uses only the recorded video to predict video quality scores. In different evaluation experiments we compare our proposed model nofu with state-of-the-art reduced or full reference models and metrics. In addition, we trained a no-reference baseline model using brisque+niqe features. We show that our model has a similar or better performance than other models. Furthermore, nofu outperforms VMAF for subjective gaming QoE prediction, even though nofu does not require any reference video.
Steve Goering, Rakesh Rao Ramachandra Rao, Alexander Raake
QoMEX2
2019 Assessing Media QoE, Simulator Sickness and Presence for Omnidirectional Videos with Different Test Protocols
abstract
QoE for omnidirectional videos comprises additional components such as simulator sickness and presence. In this paper, a series of tests is presented comparing different test protocols to assess integral quality, simulator sickness and presence for omnidirectional videos in one test run, using the HTC Vive Pro as head-mounted display. For quality ratings, the five-point ACR scale was used. In addition, the well-established Simulator Sickness Questionnaire and Presence Questionnaire methods were used, once in a full version, and once with only one single integral scale, to analyze how well presence and simulator sickness can be captured using only a single scale.
Ashutosh Singla, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake
VR2