EDBT 2026 Demo / reviewers in the wild / expert
Steve Goering
dblp:143/2250 · also Steve Göring
· DBLP profile ↗
49ranked-venue papers
14as first author
28since 2021 · last 2026
0000-0001-6810-6969ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 14 first-author · 28 since 2021Human-computer interaction and ubiquitous computing · 25 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 1Computer networks · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Analysis of Appeal, Quality, and Realism for Instagram-Like Filtered Real Photos
Steve Goering, William Menz, Rakesh Rao Ramachandra Rao |
QoMEX | 1 |
| 2026 | How Accurate are Video Quality Models for Diffusion-Based Video Super-Resolution?abstractRecent video super-resolution (VSR) approaches use deep neural networks to enhance low-quality input videos and recover visual detail, with diffusion-based methods in particular showing promising results. In this paper, we investigate whether existing video quality models can be used to assess the performance of these diffusion-based VSR methods, by comparing model predictions with results from a subjective test. The study compares six upscaling methods (Lanczos, Rhea, SCST, DOVE, SeedVR2, Starlight Mini) applied to both compressed (AV1 and DCVC-RT) and uncompressed low-resolution videos considering the play-out on a UHD-1/4K screen. A range of full- and no-reference quality models are used to assess their applicability to this new type of quality degradation, focusing on within-sequence performance. The results highlight that CNN-based full-reference models, such as LPIPS, DISTS, and CVQA-FR show significantly higher correlation coefficients than both conventional full- as well as the tested no-reference models. Most overestimate the overly sharp results of SCST, with VMAF mainly failing due to spatial inconsistencies introduced by Starlight Mini. None of the tested video quality models reach sufficient accuracy so as to replace complementary subjective testing. The reference, degraded and upscaled videos, as well as the user ratings and model scores are made available with the paper at https://github.com/Telecommunication-Telemedia-Assessment/AVT-VQDB-UHD-1-VSR as open data. Benjamin Herb, Steve Goering, Alexander Raake, Rakesh Rao Ramachandra Rao |
QoMEX | 2 |
| 2026 | A Comparative User Study of Real-Time Head-Appearance Telepresence in 2D and 3D Representations
William Menz, Alexander Zoubarev, David Kutschke, Rakesh Rao Ramachandra Rao, Louay Bassbouss, Sven Bliedung von der Heide, Alexander Raake, Steve Goering |
QoMEX | 8 |
| 2026 | Exploration of the Effect of Automatic Camera Tracking for TeleconferencingabstractConsidering the increase in using teleconferencing in the daily lives of everyone, it becomes more important to provide a good experience while attending remote meetings. Various new technologies have been proposed, ranging from AR/VR setups to enhanced 2D systems with the added possibility of automatic tracking of a person. This exploration paper investigates the impact of automatic camera tracking on virtual communication, focusing on task performance. Using a custom video conference setup combined with the TrackingMaster system (2D LiDAR-based tracking and PTZ camera framing), three tracking configurations were tested: static, tracking-without-presets, and tracking-with-presets. Participants performed various adapted survival tasks designed to encourage movement and interaction. The results showed that tracking with presets enabled slightly faster task completion and required less movement. Furthermore, workload and social presence indicated no major variation across the tracking configurations. The findings suggest potential benefits of camera tracking systems but highlight the need for further research with larger samples and novel measurements to create validated metrics, especially regarding creativity in virtual communication. Christoph Götzl, Felix Immohr, Alexander Raake, Steve Goering |
IMX | 4 |
| 2025 | Exploiting LLMs for Metadata-Based Video Quality PredictionabstractLarge language models (LLMs) can be used to solve various tasks based on text inputs, e.g., video quality estimation. We explore the usage of LLMs for video quality prediction based on metadata (video codec, bitrate, resolution), which has not been addressed before. For the evaluation we use test #1 from our AVT-VQDB-UHD-1 dataset. We generated text prompts based on the metadata and collected answers from 17 different LLMs. The evaluation indicates that especially larger LLMs could be used to simulate human raters. However, a pure model prediction with one model has lower performance than SoA video quality models. Thus, we further investigated combinations of LLMs, which resulted in comparable performance to state-of-the-art models. Our work is a proof-of-concept, considering that LLMs are slower for the prediction than the traditional metadata-based models. Steve Goering, Rakesh Rao Ramachandra Rao, Alexander Raake |
ISM | 1 |
| 2025 | Smarter Traps: Neural Network-Driven Classification of Small MammalsabstractThe intensification of agricultural practices has substantially altered ecosystem structures, affecting small mammal populations through habitat homogenization and insecticide use. Traditional monitoring of these species relies on physical trapping, which is labor-intensive and stressful for the animals. To address these challenges, this study explores an automated image-based approach for monitoring small mammals using open traps equipped with cameras. The goal was to design a system that runs efficiently on standard, low-performance computers. Therefore, existing convolutional neural network (CNN) models were retrained and evaluated for automated genus classification of captured images. The networks were retrained with a transfer learning approach using a custom dataset containing four categories, three mammal genera, and one “empty trap” class. Multiple CNN architectures were compared based on loss, accuracy, and macro F1-score to identify a model that balances performance and computational efficiency. Results showed that MobileNet-based architectures, optimized for low-power devices, underperformed in this classification task, while VGG-based networks achieved superior accuracy and generalization to unseen images from the same trap setup. The findings demonstrate the potential of CNN-driven image recognition as a scalable and noninvasive tool for ecological monitoring, reducing manual review effort and improving animal welfare in field studies. William Menz, Ralf Dittrich, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
ISM | 4 |
| 2025 | Evaluation of a Floating-Head Communication Prototype for Video-ConferencingabstractBesides traditional 2D video communication approaches, new systems aim to create more realistic and immersive representations of the involved conversation partners. This work presents a real-time floating-head communication setup that enables spatially separated visualizations of remote participants using only a standard camera and display. The underlying reconstruction pipeline applies machine-learning-assisted facial feature extraction to infer a 3D mesh of the participant's head from a live 2D video stream. Texturing maintains visual fidelity while supporting real-time performance. A dedicated transmission pipeline enables the exchange of 3D and texture data over conventional network connections, allowing flexible and location-independent use. A first lab test with eight participant pairs was performed to evaluate the system during a collaborative communication task. Subjective assessments using established telepresence and quality questionnaires confirmed the technical feasibility of the approach and its potential to enhance the sense of spatial presence. However, the overall perceptual quality and comfort did not yet reach the level of classical 2 D video communication. The study demonstrates the promise of accessible, spatially expressive communication setups that may bridge the gap between conventional video calls and emerging volumetric telepresence systems. William Menz, Alexander Zoubarev, David Kutschke, Rakesh Rao Ramachandra Rao, Louay Bassbouss, Sven Bliedung von der Heide, Steve Goering, Alexander Raake |
ISM | 7 |
| 2025 | Evaluating Video Quality Metrics for Neural and Traditional Codecs using 4K/UHD-1 Videos
Benjamin Herb, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
PCS | 3 |
| 2025 | A Large-Scale Evaluation of Subject Rating Behaviour in Visual Quality Assessment StudiesabstractSubjective testing is widely used for visual quality assessment to evaluate the impact of both technical and non-technical factors on user perception. Although standardized methods exist for collecting subjective ratings of visual quality, these ratings are inevitably influenced by each subject’s accuracy, manifesting as subject bias and inconsistency. Recommendations such as ITU-T P.910 and ITU-R BT.500 propose standardized methods to remove bias from subjective ratings. For instance, Annex E of ITU-T Rec. P.910 provides an effective strategy for addressing both subject bias and inconsistency. In this paper, we analyze 29 different visual quality assessment studies conducted over an eight-year period to understand the rating behaviour of subjects using these methods. Our investigation focuses on subjective studies targeting 4K, 8K, as well as 360°video and high-resolution image quality assessment. In the context of 4K video quality assessment, both SDR and HDR evaluations are considered. Our dataset covers use cases of short-term video quality and overall session quality assessment for HTTP-based adaptive streaming (HAS). For both these use cases, we propose a range of subject bias and inconsistency values that can serve as a reference for future studies. Furthermore, we define six different measures that can be used to assess the reliability of future studies and provide reference values for these measures. Following an open-science approach, all individual ratings from the included subjective tests, along with the results of the large-scale analysis, are made publicly available with this paper. Rakesh Rao Ramachandra Rao, Steve Goering, Stephan Fremerey, Dominik Keller, Alexander Raake |
QoMEX | 2 |
| 2025 | The Effect of Hand Visibility in AR: Comparing Dexterity and Interaction with Virtual and Real ObjectsabstractHand-tracking technologies allow us to use our own hands to interact with real and virtual objects in Augmented Reality (AR) environments. This enables us to explore the interplay between hand-visualizations and hand-object interactions. We present a user study that examines the effect of different hand visualizations (invisible, transparent, opaque) on manipulation performance when interacting with real and virtual objects. For this, we implemented video-see-through (VST) AR-based virtual building blocks and hot wire tasks with real one-to-one counterparts that require participants to use gross and fine motor hand movements. To evaluate manipulation performance, we considered three measures: task completion time, number of collisions (hot wire task), and percentage of object displacement (building block task). Additionally, we explored the sense of agency and subjective impressions (preference, ease of interaction, successful and awkwardness) evoked by the different hand-visualizations. The results show that (1) manipulation performance is significantly higher when interacting with real objects compared to virtual ones, (2) invisible hands lead to fewer errors, higher agency, higher perceived success and ease of interaction during fine manipulation tasks with real objects, and (3) having some visualization of the virtual hands (transparent or opaque) overlayed on the real hands is preferred when manipulating virtual objects even when there are no significant performance improvements. Our empirical findings about the differences when interacting with real and virtual objects can aid hand visualization choices for manipulation tasks in AR. Jakob Hartbrich, Stephanie Arevalo, Steve Goering, Alexander Raake |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Appeal prediction for AI up-scaled ImagesabstractDNN- or AI-based up-scaling algorithms are gaining in popularity due to the improvements in machine learning. Various up-scaling models using CNNs, GANs or mixed approaches have been published. The majority of models are evaluated using PSRN and SSIM or only a few example images. However, a performance evaluation with a wide range of real-world images and subjective evaluation is missing, which we tackle in the following paper. For this reason, we describe our developed dataset, which uses 136 base images and five different up-scaling methods, namely Real-ESRGAN, BSRGAN, waifu2x, KXNet, and Lanczos. Overall the dataset consists of 1496 annotated images. The labeling of our dataset focused on image appeal and has been performed using crowd-sourcing employing our open-source tool AVRate Voyager. We evaluate the appeal of the different methods, and the results indicate that Real-ESRGAN and BSRGAN are the best. Furthermore, we train a DNN to detect which up-scaling method has been used, the trained models have a good overall performance in our evaluation. In addition to this, we evaluate state-of-the-art image appeal and quality models, here none of the models showed a high prediction performance, therefore we also trained two own approaches. The first uses transfer learning and has the best performance, and the second model uses signal-based features and a random forest model with good overall performance. We share the data and implementation to allow further research in the context of open science. Steve Goering, Rasmus Merten, Alexander Raake |
ISM | 1 |
| 2024 | Investigating the Impact of High Frame Rate on Video Quality: A SAMVIQ ApproachabstractHigh Frame Rate (HFR) aims at increasing the perceived video quality by decreasing motion artifacts and enabling a smoother playback of movements. However, HFR is not yet widely used in video playback, as most movies are partly due to artistic reasons still shot and shown at 24 frames per second (fps) and streamed videos are usually capped at 60 fps. This raises the question of whether people can perceive differences with higher frame rates and a connected increase in quality at all. To this effect, this paper analyzes the relationship between frame rate and perceived video quality using the Subjective Assessment Methodology for Video Quality (SAMVIQ). In the test, 24 subjects assessed the video quality of 16 sources with varied frame rates. The results show an increased video quality for videos with higher frame up to 120 fps. The SAMVIQ methodology is useful and our version is made publicly available. Dominik Keller, Paul Rudi Frank, Steve Goering, Alexander Raake |
ISM | 3 |
| 2024 | The Frankenstone toolbox for video quality analysis of user-generated contentabstractUser-generated video content is one major part of currently streamed video content. Providers such as YouTube, Twitch, or Vimeo provide thousands of videos to users. However, the quality of user-generated content can vary widely, if not only purely technical, quality-related aspects are considered, but also the liking of the content is taken into account. Several studies and published open-source models aim to predict the quality scores of user-generated content. We propose in this paper a unified toolbox – Frankenstone – that includes the latest video quality prediction models for user-generated content. As well as recently published models also meta-data and signal-based features are included in the toolbox. The Frankenstone toolbox relies on the usage of GPUs for the calculation. We evaluate our toolbox with the test data of the YouTube UGC Dataset. Steve Goering, Alexander Raake |
QoMEX | 1 |
| 2024 | Subjective Evaluation of the Impact of Spatial Audio on Triadic Communication in Virtual RealityabstractVirtual Reality (VR) enables users to meet, converse, and collaborate in shared virtual environments. For such communication systems, many system factors can affect user experience and perception. To effectively allocate system resources, understanding of the relative influence of such factors is required. One important factor is a spatial auralization, which has been shown to elevate users’ experience in traditional and single-user VR systems. However, its effect in multi-party social VR has not been fully investigated. In this work, we conducted a study assessing the effect of spatial audio on audiovisual plausibility and presence perception in a three-user interactive communication scenario. Triads of participants perform a collaborative conversation task under three conditions: a VR condition with binaural spatial audio, a VR condition with simple diotic audio, and a real-world reference condition. This paper presents the results of the study based on questionnaire-based evaluation. Felix Immohr, Gareth Rendle, Christian Kehling, Anton Benjamin Lammert, Steve Goering, Bernd Fröhlich 0001, Alexander Raake |
QoMEX | 5 |
| 2023 | Towards evaluation of immersion, visual comfort and exploration behaviour for non-stereoscopic and stereoscopic 360° videosabstractImmersion, visual comfort, and exploration behaviour are important aspects that affect the overall quality of experience for 360° videos. To analyze the benefits of stereoscopic and non-stereoscopic 360° videos in terms of these factors, we created a dataset and conducted a subjective study. The dataset consists of five different high-resolution $8 \mathrm{~K}$ omnidirectional videos as stereoscopic and non-stereoscopic variants. The videos have been recorded using a Kandao Obsidian Pro camera. For the comparison, we designed and performed a subjective test with 30 participants. Here, each subject watched both HEVC (libx265) encoded versions of the source video and rated the videos viewed regarding presence, visual comfort, and quality. The results indicate that with the test protocol followed, non-stereoscopic video viewing leads to slightly better presence, visual comfort, and quality ratings compared to the stereoscopic variants. Further, the stereoscopic 360° videos may suffer from visual artefacts potentially leading to lower video quality and further lower quality of experience results. The exploration behaviour was found to be very similar for both non-stereoscopic and stereoscopic video viewing. Overall, it can be concluded that there is a slight tendency for non-stereoscopic video viewing to be preferred over stereoscopic video viewing. The dataset is made publicly available with the paper and includes both variants of all source videos along with the subjective data, and behaviour data, following an open-science approach. Stephan Fremerey, Raja Faseeh Uz Zaman, Touseef Ashraf, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
ISM | 5 |
| 2023 | The Effect of Viewing Distances on 4K and 8K HDR Video Quality PerceptionabstractOngoing research in the field of capture, coding and display technology and human vision has explored the advantages of high resolution up to 8K (UHD-2) considering perceived quality. One of the crucial elements impacting users’ perception of video quality is the viewing distance. As a result, the presented study employs a subjective evaluation to investigate the perceptual benefits offered by 8K or upscaled 4K in comparison to the native 4K (UHD-1) resolution in the context of HDR videos. The subjective test uses 7 distinct viewing distances, ranging from 0.5H to 3H, with H representing the display height. The findings of the study reveal a consistent trend: the increased video quality of 8K HDR against 4K HDR content decreases with distance, on average. While there are bigger improvements for close distances, beyond 2H the quality difference was very little or zero, depending on content. In general, the degree of enhancement is contingent on the spatial complexity of the content. Additionally, it is found that, on average, subjects prefer to sit at a distance of 2.07H. No significant difference in the preferred viewing distance was found when asked before and after the study. Dominik Keller, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
ISM | 3 |
| 2023 | Adaptation of Bitstream-based Video Quality Models for Image Quality AssessmentabstractIn recent years, video-codec-based image codecs, such as e.g. HEF, AVF, etc., have been increasingly used to compress images. Hence, there is a potential to use video quality prediction models for the evaluation of image quality. Bitstream-based models show promising results for video quality prediction, therefore, we investigate the applicability of such models for the case of image quality in this paper. For this purpose, we selected ITU-T Rec. P.1204.3 and its Mode 0 variant also known as $AVQBits|M3$ and $AVQBits|M0$ respectively for the evaluation, because they are computationally less complex and do not need a reference image. These models are evaluated using a publicly available dataset consisting of a total of 371 images of resolutions between $144\times 144$ pixels to $2160\times 2160$ pixels with subjective annotations. The results show that both the considered models perform well on the used dataset with a Pearson correlation of 0.958 and Root Mean Square Error (RMSE) of 0.319 (on a 1 to 5 Absolute Category Rating (ACR) scale) for the $AVQBits|M3$ model and a Pearson correlation of 0.942 and RMSE of 0.377 for the $AVQBits|M0$ model. Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
ISM | 2 |
| 2023 | AVT-VQDB-UHD-1-Appeal: A UHD-1/4K Open Dataset for Video Quality and Appeal Assessment Using Modern Video CodecsabstractA number of factors play an important role in the perception of video quality for streaming and other services, key among them being encoding-related degradations. Hence, newer codecs are developed with the goal of optimizing video quality for a given encoding setting. Here, subjective studies are an efficient method to evaluate the performance of such newer codecs. Furthermore, contextual factors impact the perception of video quality, e.g., the appeal of the content itself. To this end, this paper presents a subjective study targeting both quality and appeal assessment of videos. For this purpose, a subjective study consisting of three different parts is conducted. Firstly, participants were asked to rate the appeal of the uncompressed UHD-1/4K source content with a duration of 8 - 10s each. Following this, the video quality of these source videos individually encoded with either the HEVC/H.265, AV1, or VVC/H.266 video codec was rated. A wide range of encoding conditions in terms of resolution (360p to 2160p) and bitrate (100kbps to 15mbps) is used to encode the videos, so as to enable the applicability of the data to real-world settings. In the last part, subjects are again asked to rate the appeal of the uncompressed source content. The results are analyzed to assess the impact of different encoding conditions on perceived video quality. In addition, the impact of appeal on video quality and vice-versa is also investigated. Furthermore, an objective quality assessment with different state-of-the-art full-reference, bitstream-based, and hybrid models including the newer codecs AV1 and VVC is presented. The subjective dataset including test design, subjective results, sources, and encoded audiovisual contents are made publicly available following an open science approach. Rakesh Rao Ramachandra Rao, Steve Goering, Bassem Elmeligy, Alexander Raake |
MMSP | 2 |
| 2023 | DNN-based Photography Rule Prediction using Photo TagsabstractInstagram and Flickr are just two examples of photo-sharing platforms which are currently used to upload thousands of images on a daily basis. One important aspect in such social media contexts is to know whether an image is of high appeal or not. In particular, to understand the composition of a photo and to improve reading flow, several photo rules have been established. In this paper, we focus on eight selected photo rules. To automatically predict whether an image follows one of these rules or not, we train 13 deep neural networks in a transfer-learning setup and compare their prediction performance. As a dataset, we use photos downloaded from Flickr with specifically selected image tags, which reflect the eight photo rules. There-fore, our dataset does not need additional human annotations. ResNet50 has the best prediction performance, however, there are images that follow several rules, which must be addressed in follow-up work. The code and the data (image URLs) are made publicly available for reproducibility. Steve Goering, Rasmus Merten, Alexander Raake |
QoMEX | 1 |
| 2023 | Appeal and quality assessment for AI-generated imagesabstractRecently AI-generated images gained in popularity. A critical aspect of AI-generated images using, e.g., DALL-E-2 or Midjourney, is that they may look artificial, be of low quality, or have a low appeal in contrast to real images, depending on the text prompt and AI generator. For this reason, we evaluate the quality and appeal of AI-generated images using a crowdsourcing test as an extension of our recently published AVT-AI-Image-Dataset. This dataset consists of a total of 135 images generated with five different AI-text-to-image generators. Based on the collected subjective ratings in the crowdsourcing test, we evaluate the different used AI generators in terms of image quality and appeal of the AI-generated images. We also link image quality and image appeal also with SoA objective models. The extension will be made publicly available for reproducibility. Steve Goering, Rakesh Rao Ramachandra Rao, Rasmus Merten, Alexander Raake |
QoMEX | 1 |
| 2023 | PNATS-UHD-1-Long: An Open Video Quality Dataset for Long Sequences for HTTP-based Adaptive Streaming QoE AssessmentabstractThe P.NATS Phase 2 competition in ITU-T Study Group 12 resulted in both the ITU-T Rec. P.1204 series of recommendations, and also a large dataset for HTTP-based adaptive streaming QoE assessment that is now made openly available as part of this paper. The presented dataset consists of 3 subjective databases targeting overall quality assessment of a typical HTTP-based Adaptive Streaming session consisting of degradations such as quality switching, initial loading delay, and stalling events using audiovisual content ranging between 2 and 5 minutes. In addition to this, subject bias and consistency in quality assessment of such longer-duration audiovisual contents with multiple degradations are investigated using a subject behaviour model. As part of this paper, the overall test design, subjective test results, sources, encoded audiovisual contents, and a set of analysis plots are made publicly available for further research. Rakesh Rao Ramachandra Rao, Silvio Borer, David Lindero, Steve Goering, Alexander Raake |
QoMEX | 4 |
| 2023 | Proof-of-Concept Study to Evaluate the Impact of Spatial Audio on Social Presence and User Behavior in Multi-Modal VR CommunicationabstractThis paper presents a proof-of-concept study conducted to analyze the effect of simple diotic vs. spatial, position-dynamic binaural synthesis on social presence in VR, in comparison with face-to-face communication in the real world, for a sample two-party scenario. A conversational task with shared visual reference was realized. The collected data includes questionnaires for direct assessment, tracking data, and audio and video recordings of the individual participants’ sessions for indirect evaluation. While tendencies for improvements with binaural over diotic presentation can be observed, no significant difference in social presence was found for the considered scenario. The gestural analysis revealed that participants used the same amount and type of gestures in face-to-face as in VR, highlighting the importance of non-verbal behavior in communication. As part of the research, an end-to-end framework for conducting communication studies and analysis has been developed. Felix Immohr, Gareth Rendle, Annika Neidhardt, Steve Goering, Rakesh Rao Ramachandra Rao, Stephanie Arevalo, Bernd Fröhlich 0001, Alexander Raake |
IMX | 4 |
| 2021 | Rule of Thirds and Simplicity for Image Aesthetics using Deep Neural NetworksabstractConsidering the increasing amount of photos being uploaded to sharing platforms, a proper evaluation of photo appeal or aesthetics is required. For appealing images several "rules of thumb" have been established, e.g., the rule of thirds and simplicity. We handle rule of thirds and simplicity as binary classification problems with a deep learning based image processing pipeline. Our pipeline uses a pre-processing step, a pre-trained baseline deep neural network (DNN) and post-processing. For each of the rules, we re-train 17 pre-trained DNN models using transfer learning. Our results for publicly available datasets show that the ResNet152 DNN is best for rule of thirds prediction and DenseNet121 is best for simplicity with an accuracy of around 0.84 and 0.94 respectively. In addition to the datasets for both classifications, five experts annotated another dataset with ≈ 1100 images and we evaluate the best performing models. Results show that the best performing models have an accuracy of 0.67 for rule of thirds and 0.79 for image simplicity. Both accuracy results are within the range of pairwise accuracy of expert annotators. However, it further indicates that there is a high subjective influence for both of the considered rules. Steve Goering, Alexander Raake |
MMSP | 1 |
| 2021 | AVrate Voyager: an open source online testing platformabstractSubjective testing is an integral part of many research fields considering, e.g., human perception. For this purpose, lab tests are a popular approach to gather ratings for subjective evaluations. However, not in all cases controlled lab tests can be performed, either in cases where no labs are existing, accessible or it may be disallowed to use them. For this reason, online tests, e.g., using crowdsourcing are supposed to be an alternative approach for traditional lab tests. We describe in the following paper a framework to implement such online tests for audio, video, and image-related evaluations or questionnaires. Our framework AVrate Voyager builds upon previously developed frameworks for lab tests including the experience with them. AVrate Voyager uses scalable web technologies to implement a test framework, this ensures that it will be running reliably. In addition, we added strategies for pre-caching to avoid additional influence for play-out, e.g. in the case of video testing. We analyze several conducted tests using the new framework and describe the required steps to modify the provided tool in detail. Steve Goering, Rakesh Rao Ramachandra Rao, Stephan Fremerey, Alexander Raake |
MMSP | 1 |
| 2021 | Groovability: Using Groove as a Novel Measure for Audio QoE with the Example of SmartphonesabstractGroove in music is a fundamental part of why humans entrain to it and enjoy it. Smartphones have become an important medium to listen to music. Especially when being with others, loudspeaker playback may be the method of choice. However, due to the physical limits of acoustics, for loudspeaker playback, smartphones are equipped with sub-optimal audio capabilities. Therefore, it is desirable to measure Quality of Experience (QoE) of music played on smartphones. While audio playback is often assessed in terms of sound quality, the aim of this work is to address QoE in terms of the meaning or effect that the audio has on the listener. A key component for the meaning of popular music is groove. Hence, in this paper, we study “groovability”, that is, the ability of a piece of audio technology to convey groove. To instantiate our novel audio QoE assessment method, we apply it to music played by 8 different smartphones. For this purpose, looped 4-bar loudness-aligned recordings from 24 music pieces of different intrinsic groove were played back on the different smartphones. Our test method uses a multi-stimulus comparison with synchronized playback capability. A total of 62 subjects evaluated groovability using two stimulus subsets. It was found that the proposed methodology is highly effective to distinguish between the groovability provided by the considered phones. In addition, a reduced-reference model is proposed to predict groovability, using a set of both acoustics-and music-groove related features. In our formal validation on unknown data, the model is shown to provide good prediction performance with a Pearson correlation of greater than 0.90. Dominik Keller, Markus Vaalgamaa, Erkki Paajanen, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
QoMEX | 5 |
| 2021 | Towards High Resolution Video Quality Assessment in the CrowdabstractAssessing high resolution video quality is usually performed using controlled, defined, and standardized lab tests. This method of acquiring human ratings in a lab environment is time-consuming and may also not reflect the typical viewing conditions. To overcome these disadvantages, crowd testing paradigms have been used for assessing video quality in general. Crowdsourcing-based tests enable a more diverse set of participants and also use a realistic hardware setup and viewing environment of typical users. However, obtaining valid ratings for high-resolution video quality poses several problems. Example issues are that streaming of such high-bandwidth content may not be feasible for some users, or that crowd participants lack an appropriate, high-resolution display device. In this paper, we propose a method to overcome such problems and conduct a crowd test using for higher resolution content by using a 540 p cutout from the center of the original 2160p video. To this aim, we use the videos from Test#1 of the publicly available dataset AVT-VQDB-UHD-1, which contains videos up to a resolution of UHD-1. The quality-labels available from that lab test allow us to compare the results with the crowd test presented in this paper. It is shown that there is a Pearson correlation of 0.96 between the lab and crowd tests and hence such crowd tests can reliably be used for video assessment of higher resolution content. The overall implementation of the crowd test framework and the results are made publicly available for further research and reproducibility1. Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
QoMEX | 2 |
| 2021 | Impact of Spatial and Temporal Information on Video Quality and CompressibilityabstractSpatial Information (SI) and Temporal Information (TI) are frequently-used metrics to classify the spatiotemporal complexity of video content. However, they are mostly used on original video sources, and their impact on actual encoding efficiency is not known. In this paper, we propose a method to determine the compressibility of video sources, that is, how good video quality can be under a given bitrate constraint. We show how various aggregations of SI and TI correlate with compressibility scores obtained from a public dataset of H.264/HEVCN P9 content. We observe that the minimum TI value as well as an existing criticality metric from the literature are good indicators for compressibility, as judged by subjective ratings as well as VMAF and P.1204.3 objective scores. Werner Robitza, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
QoMEX | 3 |
| 2021 | Assessment of the Simulator Sickness Questionnaire for Omnidirectional VideosabstractVirtual Reality/360° videos provide an immersive experience to users. Besides this, 360° videos may lead to an undesirable effect when consumed with Head-Mounted Displays (HMDs), referred to as simulator sickness/cybersickness. The Simulator Sickness Questionnaire (SSQ) is the most widely used questionnaire for the assessment of simulator sickness. Since the SSQ with its 16 questions was not designed for 360° video related studies, our research hypothesis in this paper was that it may be simplified to enable more efficient testing for 360° video. Hence, we evaluate the SSQ to reduce the number of questions asked from subjects, based on six different previously conducted studies. We derive the reduced set of questions from the SSQ using Principal Component Analysis (PCA) for each test. Pearson Correlation is analysed to compare the relation of all obtained reduced questionnaires as well as two further variants of SSQ reported in the literature, namely Virtual Reality Sickness Questionnaire (VRSQ) and Cybersickness Questionnaire (CSQ). Our analysis suggests that a reduced questionnaire with 9 out of 16 questions yields the best agreement with the initial SSQ, with less than 44% of the initial questions. Exploratory Factor Analysis (EFA) shows that the nine symptom-related attributes determined as relevant by PCA also appear to be sufficient to represent the three dimensions resulting from EFA, namely, Uneasiness, Visual Discomfort and Loss of Balance. The simplified version of the SSQ has the potential to be more efficiently used than the initial SSQ for 360° video by focusing on the questions that are most relevant for individuals, shortening the required testing time. Ashutosh Singla, Steve Goering, Dominik Keller, Rakesh Rao Ramachandra Rao, Stephan Fremerey, Alexander Raake |
VR | 2 |
| 2020 | Between the Frames - Evaluation of Various Motion Interpolation Algorithms to Improve 360° Video QualityabstractWith the increasing availability of 360° video content, it becomes important to provide smoothly playing videos of high quality for end users. For this reason, we compare the influence of different Motion Interpolation (MI) algorithms on 360° video quality. After conducting a pre-test with 12 video experts in [3], we found that MI is a useful tool to increase the QoE (Quality of Experience) of omnidirectional videos. As a result of the pretest, we selected three suitable MI algorithms, namely ffmpeg Motion Compensated Interpolation (MCI), Butterflow and Super-SloMo. Subsequently, we interpolated 15 entertaining and realworld omnidirectional videos with a duration of 20 seconds from 30 fps (original framerate) to 90 fps, which is the native refresh rate of the HMD used, the HTC Vive Pro. To assess QoE, we conducted two subjective tests with 24 and 27 participants. In the first test we used a Modified Paired Comparison (M-PC) method, and in the second test the Absolute Category Rating (ACR) approach. In the M-PC test, 45 stimuli were used and in the ACR test 60. Results show that for most of the 360° videos, the interpolated versions obtained significantly higher quality scores than the lower-framerate source videos, validating our hypothesis that motion interpolation can improve the overall video quality for 360° video. As expected, it was observed that the relative comparisons in the M-PC test result in larger differences in terms of quality. Generally, the ACR method lead to similar results, while reflecting a more realistic viewing situation. In addition, we compared the different MI algorithms and can conclude that with sufficient available computing power Super-SloMo should be preferred for interpolation of omnidirectional videos, while MCI also shows a good performance. Stephan Fremerey, Frank Hofmeyer, Steve Goering, Dominik Keller, Alexander Raake |
ISM | 3 |
| 2020 | Subjective Test Dataset and Meta-data-based Models for 360° Streaming Video QualityabstractDuring the last years, the number of 360° videos available for streaming has rapidly increased, leading to the need for 360° streaming video quality assessment. In this paper, we report and publish results of three subjective 360° video quality tests, with conditions used to reflect real-world bitrates and resolutions including 4K, 6K and 8K, resulting in 64 stimuli each for the first two tests and 63 for the third. As playout device we used the HTC Vive for the first and HTC Vive Pro for the remaining two tests. Video-quality ratings were collected using the 5-point Absolute Category Rating scale. The 360° dataset provided with the paper contains the links of the used source videos, the raw subjective scores, video-related meta-data, head rotation data and Simulator Sickness Questionnaire results per stimulus and per subject to enable reproducibility of the provided results. Moreover, we use our dataset to compare the performance of state-of-the-art full-reference quality metrics such as VMAF, PSNR, SSIM, ADM2, WS-PSNR and WS-SSIM. Out of all metrics, VMAF was found to show the highest correlation with the subjective scores. Further, we evaluated a center-cropped version of VMAF ("VMAF-cc") that showed to provide a similar performance as the full VMAF. In addition to the dataset and the objective metric evaluation, we propose two new video-quality prediction models, a bitstream meta-data-based model and a hybrid no-reference model using bitrate, resolution and pixel information of the video as input. The new lightweight models provide similar performance as the full-reference models while enabling fast calculations. Stephan Fremerey, Steve Goering, Rakesh Rao Ramachandra Rao, Rachel Huang, Alexander Raake |
MMSP | 2 |
| 2020 | Automated Genre Classification for Gaming VideosabstractBesides classical videos, videos of gaming matches, entire tournaments or individual sessions are streamed and viewed all over the world. The increased popularity of Twitch or YoutubeGaming shows the importance of additional research on gaming videos. One important pre-condition for live or offline encoding of gaming videos is the knowledge of game-specific properties. Knowing or automatically predicting the genre of a gaming video enables a more advanced and optimized encoding pipeline for streaming providers, especially because gaming videos of different genres vary a lot from classical 2D video, e.g., considering the CGI content, textures or camera motion. We describe several computer-vision based features that are optimized for speed and motivated by characteristics of popular games, to automatically predict the genre of a gaming video. Our prediction system uses random forest and gradient boosting trees as underlying machine-learning techniques, combined with feature selection. For the evaluation of our approach we use a dataset that was built as part of this work and consists of recorded gaming sessions for 6 genres from Twitch. In total 351 different videos are considered. We show that our prediction approach shows a good performance in terms of f1-score. Besides the evaluation of different machine-learning approaches, we additionally investigate the influence of the hyper-parameters for the algorithms. Steve Goering, Robert Steger, Rakesh Rao Ramachandra Rao, Alexander Raake |
MMSP | 1 |
| 2020 | A Large-scale Evaluation of the bitstream-based video-quality model ITU-T P.1204.3 on Gaming ContentabstractThe streaming of gaming content, both passive and interactive, has increased manifolds in recent years. Gaming contents bring with them some peculiarities which are normally not seen in traditional 2D videos, such as the artificial and synthetic nature of contents or repetition of objects in a game. In addition, the perception of gaming content by the user is different from that of traditional 2D videos due to its pecularities and also the fact that users may not often watch such content. Hence, it becomes imperative to evaluate whether the existing video quality models usually designed for traditional 2D videos are applicable to gaming content. In this paper, we evaluate the applicability of the recently standardized bitstream-based video-quality model ITU-T P.1204.3 on gaming content. To analyze the performance of this model, we used 4 different gaming datasets (3 publicly available + 1 internal) not previously used for model training, and compared it with the existing state-of-the-art models. We found that the ITU P.1204.3 model out of the box performs well on these unseen datasets, with an RMSE ranging between 0.38 - 0.45 on the 5-point absolute category rating and Pearson Correlation between 0.85 - 0.93 across all the 4 databases. We further propose a full-HD variant of the P.1204.3 model, since the original model is trained and validated which targets a resolution of 4K/UHD-1. A 50:50 split across all databases is used to train and validate this variant so as to make sure that the proposed model is applicable to various conditions. Rakesh Rao Ramachandra Rao, Steve Goering, Robert Steger, Saman Zad Tootaghaj, Nabajeet Barman, Stephan Fremerey, Sebastian Möller 0001, Alexander Raake |
MMSP | 2 |
| 2020 | DEMI: Deep Video Quality Estimation Model using Perceptual Video Quality DimensionsabstractExisting works in the field of quality assessment focus separately on gaming and non-gaming content. Along with the traditional modeling approaches, deep learning based approaches have been used to develop quality models, due to their high prediction accuracy. In this paper, we present a deep learning based quality estimation model considering both gaming and non-gaming videos. The model is developed in three phases. First, a convolutional neural network (CNN) is trained based on an objective metric which allows the CNN to learn video artifacts such as blurriness and blockiness. Next, the model is fine-tuned based on a small image quality dataset using blockiness and blurriness ratings. Finally, a Random Forest is used to pool frame-level predictions and temporal information of videos in order to predict the overall video quality. The light-weight, low complexity nature of the model makes it suitable for real-time applications considering both gaming and non-gaming content while achieving similar performance to existing state-of-the-art model NDNetGaming. The model implementation for testing is available on GitHub1. Saman Zad Tootaghaj, Nabajeet Barman, Rakesh Rao Ramachandra Rao, Steve Goering, Maria G. Martini, Alexander Raake, Sebastian Möller 0001 |
MMSP | 4 |
| 2020 | Prenc - Predict Number of Video Encoding Passes with Machine LearningabstractVideo streaming providers spend huge amounts of processing time to get a quality-optimized encoding. While the quality-related impact may be known to the service provider, the impact on video quality is hard to assess, when no reference is available. Here, bitstream-based video quality models may be applicable, delivering estimates that include encoding-specific settings. Such models typically use several input parameters, e.g. bitrate, framerate, resolution, video codec, QP values and more. However, for a given bitstream, to determine which encoding parameters were selected, e.g., the number of encoding passes, is not a trivial task. This leads to our following research question: Given an unknown video bitstream, which encoding settings have been used? To tackle this reverse engineering problem, we introduce a system called prenc. Besides the use in video-quality estimation, such algorithms may also be used in other applications such as video forensics. We prove our concept by applying prenc to distinguish between one- and two-pass encoding. Starting from modeling the problem as a classification task, estimating bitstream-based features, we further describe a machine learning approach with feature selection to automatically predict the number of encoding passes for a given video bitstream. Our large-scale evaluation consists of 16 short movie type 4K videos that were segmented and encoded with different settings (resolutions, codecs, bitrates), so that we in total analyzed 131.976 DASH video segments. We further show that our system is robust, based on a 50% train and 50% validation approach without source video overlapping, where we get a classification performance of 65% F1 score. Moreover, we also describe the used bitstream-based features in detail, the feature pooling strategy and include other machine learning algorithms in our evaluation. Steve Goering, Rakesh Rao Ramachandra Rao, Alexander Raake |
QoMEX | 1 |
| 2020 | Bitstream-Based Model Standard for 4K/UHD: ITU-T P.1204.3 - Model Details, Evaluation, Analysis and Open Source ImplementationabstractWith the increasing requirement of users to view high-quality videos with a constrained bandwidth, typically realized using HTTP-based adaptive streaming, it becomes more and more important to determine the quality of the encoded videos accurately, to assess and possibly optimize the overall streaming quality. In this paper, we describe a bitstream-based no-reference video quality model developed as part of the latest model-development competition conducted by ITU-T Study Group 12 and the Video Quality Experts Group (VQEG), “P.NATS Phase 2”. It is now part of the new P.1204 series of Recommendations as P.1204.3. It can be applied to bitstreams encoded with H.264/AVC, HEVC and VP9, using various encoding options, including resolution, bitrate, framerate and typical encoder settings such as number of passes, rate control variants and speeds. The proposed model follows an ensemble-modelling-inspired approach with weighted parametric and machine-learning parts to efficiently leverage the performance of both approaches. The paper provides details about the general approach to modelling, the features used and the final feature aggregation. The model creates per-segment and per-second video quality scores on the 5-point Absolute Category Rating scale, and is applicable to segments of 5–10 seconds duration. It covers both PC/TV and mobile/tablet viewing scenarios. We outline the databases on which the model was trained and validated as part of the competition, and perform an additional evaluation using a total of four independently created databases, where resolutions varied from 360p to 2160p, and frame rates from 15–60fps, using realistic coding and bitrate settings. We found that the model performs well on the independent dataset, with a Pearson correlation of 0.942 and an RMSE of 0.42. We also provide an open-source reference implementation of the described P.1204.3 model, as well as the multi-codec bitstream parser required to extract the input data, which is not part of the standard. Rakesh Rao Ramachandra Rao, Steve Goering, Peter List 0001, Werner Robitza, Bernhard Feiten, Ulf Wüstenhagen, Alexander Raake |
QoMEX | 2 |
| 2020 | Are You Still Watching? Streaming Video Quality and Engagement Assessment in the CrowdabstractAs video streaming accounts for the majority of Internet traffic, monitoring its quality is of importance to both Over the Top (OTT) providers as well as Internet Service Providers (ISPs). While OTTs have access to their own analytics data with detailed information, ISPs often have to rely on automated network probes for estimating streaming quality, and likewise, academic researchers have no information on actual customer behavior. In this paper, we present first results from a large-scale crowdsourcing study in which three major video streaming OTTs were compared across five major national ISPs in Germany. We not only look at streaming performance in terms of loading times and stalling, but also customer behavior (e.g., user engagement) and Quality of Experience based on the ITU-T P.1203 QoE model. We used a browser extension to evaluate the streaming quality and to passively collect anonymous OTT usage information based on explicit user consent. Our data comprises over 400,000 video playbacks from more than 2,000 users, collected throughout the entire year of 2019. The results show differences in how customers use the video services, how the content is watched, how the network influences video streaming QoE, and how user engagement varies by service. Hence, the crowdsourcing paradigm is a viable approach for third parties to obtain streaming QoE insights from OTTs. Werner Robitza, Alexander M. Dethof, Steve Goering, Alexander Raake, André Beyer, Tim Polzehl |
QoMEX | 3 |
| 2019 | cencro - Speedup of Video Quality Calculation using Center CroppingabstractToday's video streaming providers, e.g. Youtube, Netflix or Amazon Prime, are able to deliver high resolution and high-quality content to end users. To optimize video quality and to reduce transmission bandwidth, new encoders and smarter encoding schemes are required. Encoding optimization forms an important part of this effort in reducing bandwidth and results in saving considerable amount of bitrate. For such optimization, accurate and computationally fast video quality models are required, e.g. Netflix's VMAF. However, VMAF is a full-reference (FR) metric, and the calculation of such metrics tend to be slower in comparison to other metrics, due to the amount of data that needs to be processed, especially for high resolutions of 4k and beyond. We introduce an approach to speed up video quality metric calculations in general. We use VMAF as an example with a video database up to 4K resolution videos, to show that our approach works well. Our main idea is that we reduce each frame of the reference and distorted video based on a center crop of the frame, assuming that most important visual information are presented in the middle of most typical videos. In total we analyze 18 different crop settings and compare our results with uncropped VMAF values and subjective scores. We show that this approach - named cencro - is able to save up to 95% computation time, with just an overall error of 4% considering a 360p center crop. Furthermore, we checked other full-reference metrics, and show that cencro performs similar good. As a last evaluation, we apply our approach to full-hd gaming videos, also in this scenario cencro can be successfully applied. The idea behind cencro is not restricted to full-reference models and can also be applied to other type of video quality models or datasets, or even for higher resolution videos such as 8K. Steve Goering, Christopher Krämmer, Alexander Raake |
ISM | 1 |
| 2019 | AVT-VQDB-UHD-1: A Large Scale Video Quality Database for UHD-1abstract4K television screens or even with higher resolutions are currently available in the market. Moreover video streaming providers are able to stream videos in 4K resolution and beyond. Therefore, it becomes increasingly important to have a proper understanding of video quality especially in case of 4K videos. To this effect, in this paper, we present a study of subjective and objective quality assessment of 4K ultra-high-definition videos of short duration, similar to DASH segment lengths. As a first step, we conducted four subjective quality evaluation tests for compressed versions of the 4K videos. The videos were encoded using three different video codecs, namely H.264, HEVC, and VP9. The resolutions of the compressed videos ranged from 360p to 2160p with framerates varying from 15fps to 60fps. All the source 4K contents used were of 60fps. We included low quality conditions in terms of bitrate, resolution and framerate to ensure that the tests cover a wide range of conditions, and that e.g. possible models trained on this data are more general and applicable to a wider range of real world applications. The results of the subjective quality evaluation are analyzed to assess the impact of different factors such as bitrate, resolution, framerate, and content. In the second step, different state-of-the-art objective quality models were applied to all videos and their performance was analyzed in comparison with the subjective ratings, e.g. using Netflix's VMAF. The videos, subjective scores, both MOS and confidence interval per sequence and objective scores are made public for use by the community for further research. Rakesh Rao Ramachandra Rao, Steve Goering, Werner Robitza, Bernhard Feiten, Alexander Raake |
ISM | 2 |
| 2019 | Subjective quality evaluation of tile-based streaming for omnidirectional videosabstractIn viewport-adaptive streaming of omnidirectional video, only the field of view is streamed in high quality. While this has significant benefits over streaming the entire 360 sphere, no standard test method for perceived quality and simulator sickness is available to evaluate the quality of experience (QoE) of such streaming approaches. QoE testing is important as tile-based viewport-adaptive streaming technologies are replacing classical approaches because of significant bandwidth savings and increase in viewing quality. In this work, we propose a testbed, a test method, as well as test metrics for QoE tests of viewport-adaptive streaming approaches. The proposed method is validated in two different test setups, using a specific tile-based streaming technology available in the market. The chosen input variables (videos sequences, resolution, bandwidth, and network round-trip delay) are tested for their statistical significance. We found that our test method is suitable for QoE testing of viewport-adaptive streaming technologies. We also found that simulator sickness scores increase with test duration, but that breaks between tests reduce this effect. With our systematic test approach, it is possible to compare metrics among different test setups. On the tested technology, we found that a typical network delay (47 ms) only has a minimal effect on the quality ratings. Furthermore, the magnitude of the network delay does not influence simulator sickness for the system we have tested. Ashutosh Singla, Steve Goering, Alexander Raake, Britta Meixner, Rob Koenen, Thomas Buchholz |
MMSys | 2 |
| 2019 | Impact of Various Motion Interpolation Algorithms on 360° Video QoEabstractIn our study, we compare the impact of various motion interpolation (MI) algorithms on 360° video Quality of Experience (QoE). For doing so, we conducted a subjective test with 12 video expert viewers, while a pair comparison test method was used. We interpolated four different 20 s long 30 fps 360° source contents to the native 90 Hz refresh rate of popular Head-Mounted Displays using three different MI algorithms. Subsequently, we compared these 90 fps videos against each other to investigate the influence on the QoE. Regarding the algorithms, we found out that ffmpeg blend does not lead to a significant improvement of QoE, while MCI and butterflow do so. Additionally, we concluded that for 360° videos containing fast and sudden movements, MCI should be preferred over butterflow, while butterflow is more suitable for slow and medium motion videos. While comparing the time needed for rendering the 90 fps interpolated videos, ffmpeg blend is the fastest, while MCI and butterflow need much more time. Stephan Fremerey, Frank Hofmeyer, Steve Goering, Alexander Raake |
QoMEX | 3 |
| 2019 | nofu - A Lightweight No-Reference Pixel Based Video Quality Model for Gaming ContentabstractPopularity of streaming services for gaming videos has increased tremendously over the last years, e.g. Twitch and Youtube Gaming. Compared to classical video streaming applications, gaming videos have additional requirements. For example, it is important that videos are streamed live with only a small delay. In addition, users expect low stalling, waiting time and in general high video quality during streaming, e.g. using http-based adaptive streaming. These requirements lead to different challenges for quality prediction in case of streamed gaming videos. We describe newly developed features and a no-reference video quality machine learning model, that uses only the recorded video to predict video quality scores. In different evaluation experiments we compare our proposed model nofu with state-of-the-art reduced or full reference models and metrics. In addition, we trained a no-reference baseline model using brisque+niqe features. We show that our model has a similar or better performance than other models. Furthermore, nofu outperforms VMAF for subjective gaming QoE prediction, even though nofu does not require any reference video. Steve Goering, Rakesh Rao Ramachandra Rao, Alexander Raake |
QoMEX | 1 |
| 2019 | Assessing Media QoE, Simulator Sickness and Presence for Omnidirectional Videos with Different Test ProtocolsabstractQoE for omnidirectional videos comprises additional components such as simulator sickness and presence. In this paper, a series of tests is presented comparing different test protocols to assess integral quality, simulator sickness and presence for omnidirectional videos in one test run, using the HTC Vive Pro as head-mounted display. For quality ratings, the five-point ACR scale was used. In addition, the well-established Simulator Sickness Questionnaire and Presence Questionnaire methods were used, once in a full version, and once with only one single integral scale, to analyze how well presence and simulator sickness can be captured using only a single scale. Ashutosh Singla, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake |
VR | 3 |
| 2018 | HTTP adaptive streaming QoE estimation with ITU-T rec. P. 1203 open databases and softwareabstractThis paper describes an open dataset and software for ITU-T Ree. P.1203. As the first standardized Quality of Experience model for audiovisual HTTP Adaptive Streaming (HAS), it has been extensively trained and validated on over a thousand audiovisual sequences containing HAS-typical effects (such as stalling, coding artifacts, quality switches). Our dataset comprises four of the 30 official subjective databases at a bitstream feature level. The paper also includes subjective results and the model performance. Our software for the standard was made available to the public, too, and it is used for all the analyses presented. Among other previously unpublished details, we show the significant performance improvements of using bitstream-based models over metadata-based ones for video quality analysis, and the robustness of combining classical models with machine-learning-based approaches for estimating user QoE. Werner Robitza, Steve Goering, Alexander Raake, David Lindegren, Gunnar Heikkilä, Jörgen Gustafsson, Peter List 0001, Bernhard Feiten, Ulf Wüstenhagen, Marie-Neige Garcia, Kazuhisa Yamagishi, Simon Broom |
MMSys | 2 |
| 2018 | Extended Features using Machine Learning Techniques for Photo Liking PredictionabstractToday several photo platforms provide thousands of new pictures, it becomes ambitious to find highly appealing or like-able photos within such loads of data. Here, automatic liking prediction can support users in handling their pictures or improve ranking in sharing platforms. We describe a machine learning approach for photo liking prediction. Our features are based on various techniques, e.g. natural language processing/sentiment analysis, pre-trained deep learning networks, social network analysis and extended previously reported features. We conduct large-scale experiments using a collected dataset consisting of 80k photos based on two main categories from 500px with different settings. In our experiments we analyzed the impact of our newly features and found that social network features have the strongest influence for liking prediction, we achived a boost of 15%. Furthermore, we show that all implemented features are able to improve prediction accuracy of liking rates. We additionally analyze which groups of features that can be derived directly from pictures are usable for prediction. Steve Goering, Konstantin Brand, Alexander Raake |
QoMEX | 1 |
| 2018 | Measuring YouTube QoE with ITU-T P.1203 Under Constrained Bandwidth ConditionsabstractThe available Internet bandwidth has a strong impact on the Quality of Experience of video services. In order to manage their network efficiently and prevent customer churn, Internet Service Providers need to constantly monitor the QoE of video services such as YouTube. However, they often only rely on simple measurement scenarios that consider only one video being loaded repeatedly. In this paper we compare this scenario against a new approach in which multiple videos are being loaded in a session, thereby simulating user behavior. Using a testbed, we study the impact of download speeds on Key Performance Indicators (KPIs such as initial loading time and stalling events) and user QoE as measured using the ITU-T P.1203 standard. We show that the monitoring paradigm has a significant impact on the obtained results. We further provide a prediction model for estimating the impact of download speed on KPIs and user QoE. Werner Robitza, Dhananjaya G. Kittur, Alexander M. Dethof, Steve Goering, Bernhard Feiten, Alexander Raake |
QoMEX | 4 |
| 2017 | A framework for QoE analysis of encrypted video streamsabstractToday most internet traffic is generated by video streaming. YouTube and other video streaming platforms are using encrypted streams (HTTPS) for transport of video content. Encryption will lead to more requirements on network and content providers, e.g. caching mechanisms will not work direct. Estimation of video quality for measuring users satisfaction is also harder because there is no direct access to the video bitstream. We are building up a framework for analyzing video quality that allows us to store client information, decrypted network traffic and encrypted messages. Our approach is based on a man-in-the-middle proxy for storing the decrypted video bitstream, active probing and traffic shaping. Using these data, we are able to calculate video QoE values for example using a model such as ITU-T Rec. P.1203. Our framework will be used for generating datasets for encrypted video stream analysis, analyzing internal behavior of video streaming platforms, and more. For experimental evaluation, in this paper we analyze the influence of our man-in-the-middle proxy on key-performance indicators (KPIs) for video streaming quality. Steve Goering, Alexander Raake, Bernhard Feiten |
QoMEX | 1 |
| 2017 | A bitstream-based, scalable video-quality model for HTTP adaptive streaming: ITU-T P.1203.1abstractThe paper presents the scalable video quality model part of the P.1203 Recommendation series, developed in a competition within ITU-T Study Group 12 previously referred to as P.NATS. It provides integral quality predictions for 1 up to 5 min long media sessions for HTTP Adaptive Streaming (HAS) with up to HD video resolution. The model is available in four modes of operation for different levels of media-related bitstream information, reflecting different types of encryption of the media stream. The video quality model presented in this paper delivers short-term video quality estimates that serve as input to the integration component of the P.1203 model. The scalable approach consists in the usage of the same components for spatial and temporal scaling degradations across all modes. The third component of the model addresses video coding artifacts. To this aim, a single model parameter is introduced that can be derived from different types of bitstream input information. Depending on the complexity of the available input, one of four scaling-levels of the model is applied. The paper presents the different novelties of the model and scientific choices made during its development, the test design, and an analysis of the model performance across the different modes. Alexander Raake, Marie-Neige Garcia, Werner Robitza, Peter List 0001, Steve Goering, Bernhard Feiten |
QoMEX | 5 |
| 2016 | Axiomatic Result Re-RankingabstractWe consider the problem of re-ranking the top-k documents returned by a retrieval system given some search query. This setting is common to learning-to-rank scenarios, and it is often solved with machine learning and feature weighting based on user preferences such as clicks, dwell times, etc. In this paper, we combine the learning-to-rank paradigm with the recent developments on axioms for information retrieval. In particular, we suggest to re-rank the top-k documents of a retrieval system using carefully chosen axiom combinations. In recent years, research on axioms for information retrieval has focused on identifying reasonable constraints that retrieval systems should fulfill. Researchers have analyzed a wide range of standard retrieval models for conformance to the proposed axioms and, at times, suggested certain adjustments to the models. We take up this axiomatic view---but, instead of adjusting the retrieval models themselves, we suggest the following innovation: to adopt the learning-to-rank idea and to re-rank the top-k results directly using promising axiom combinations. This way, we can turn every reasonable basic retrieval model into an axiom-based retrieval model. In large-scale experiments on the ClueWeb corpora, we identify promising axiom combinations for a variety of retrieval models. Our experiments show that for most of these models our axiom-based re-ranking significantly improves the original retrieval performance. Matthias Hagen, Michael Völske, Steve Goering, Benno Stein 0001 |
CIKM | 3 |
| 2013 | Capabilities and objectives of distributed image processing on smart camera systemsabstractThe challenge of bringing more intelligence to the infrastructure of modern cities requires a change of thinking and states a demand for new algorithms and strategies. One important outcome of those algorithms is a realtime estimation of the prevalent spatio-temporal conditions of public transportation networks by making use of distributed image processing on networked smart camera systems. This paper provides a detailed analysis of two exemplary networked applications that can use the derived data. A conducted simulation study based on the infrastructure of real cities shows the potential of using autonomously generated knowledge, that smart camera systems can provide. Especially, inter-camera object tracking, as well as adaptive and smart navigation tasks can benefit considerably and substantiate the need for autonomous and confidential image processing. Rene Golembewski, Steve Goering, Günter Schäfer |
ISCC | 2 |