Stephan Fremerey

dblp:202/7707 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0002-6623-3777ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 A Large-Scale Evaluation of Subject Rating Behaviour in Visual Quality Assessment Studies
abstract
Subjective testing is widely used for visual quality assessment to evaluate the impact of both technical and non-technical factors on user perception. Although standardized methods exist for collecting subjective ratings of visual quality, these ratings are inevitably influenced by each subject’s accuracy, manifesting as subject bias and inconsistency. Recommendations such as ITU-T P.910 and ITU-R BT.500 propose standardized methods to remove bias from subjective ratings. For instance, Annex E of ITU-T Rec. P.910 provides an effective strategy for addressing both subject bias and inconsistency. In this paper, we analyze 29 different visual quality assessment studies conducted over an eight-year period to understand the rating behaviour of subjects using these methods. Our investigation focuses on subjective studies targeting 4K, 8K, as well as 360°video and high-resolution image quality assessment. In the context of 4K video quality assessment, both SDR and HDR evaluations are considered. Our dataset covers use cases of short-term video quality and overall session quality assessment for HTTP-based adaptive streaming (HAS). For both these use cases, we propose a range of subject bias and inconsistency values that can serve as a reference for future studies. Furthermore, we define six different measures that can be used to assess the reliability of future studies and provide reference values for these measures. Following an open-science approach, all individual ratings from the included subjective tests, along with the results of the large-scale analysis, are made publicly available with this paper.
Rakesh Rao Ramachandra Rao, Steve Goering, Stephan Fremerey, Dominik Keller, Alexander Raake
QoMEX3
2024 AVT-ECoClass-VR: An open-source audiovisual 360° video and immersive CGI multi-talker dataset to evaluate cognitive performance
abstract
The paper is part of a project to assess how complex visual and acoustic scenes affect cognitive performance in classroom scenarios, across age groups from children to adults. Here, the potential of audiovisual virtual environments for systematic user studies is explored. As of now, most studies have examined rather simple acoustic and visual representations, which do not reflect the reality of school children. An adapted version of the audiovisual scene analysis paradigm is presented, focusing on the localization and identification of talkers within a scene. The dataset includes two audiovisual scenarios (360° video and computer-generated imagery) and two implementations for dataset playback. The paper details the recording and post-processing of the content. The 360° video part of the dataset features 200 video and single-channel audio recordings of 20 speakers reading ten stories, and 20 videos of speakers in silence, resulting in a total of 220 video and 200 audio recordings. The dataset also includes one 360° background image of a real primary school classroom scene, targeting young school children for subsequent subjective tests. All stories were recorded in the German language with native German speakers. The second part of the dataset comprises 20 different 3D models of the speakers and a computer-generated classroom scene, along with an immersive audiovisual virtual environment implementation that can be interacted with using an HTC Vive controller. Both implementations also include a Unity plugin to connect and interact with the Virtual Acoustics auralization software. As a proof of concept, the dataset includes example output data collected from ongoing perception tests. There, subjects have the task of identifying which talker in the scene is reading out which story, using the story-to-speaker mapping input system developed within this paper.
Stephan Fremerey, Carolin Breuer, Larissa Leist, Maria Klatte, Janina Fels, Alexander Raake
QoMEX1
2023 Towards evaluation of immersion, visual comfort and exploration behaviour for non-stereoscopic and stereoscopic 360° videos
abstract
Immersion, visual comfort, and exploration behaviour are important aspects that affect the overall quality of experience for 360° videos. To analyze the benefits of stereoscopic and non-stereoscopic 360° videos in terms of these factors, we created a dataset and conducted a subjective study. The dataset consists of five different high-resolution $8 \mathrm{~K}$ omnidirectional videos as stereoscopic and non-stereoscopic variants. The videos have been recorded using a Kandao Obsidian Pro camera. For the comparison, we designed and performed a subjective test with 30 participants. Here, each subject watched both HEVC (libx265) encoded versions of the source video and rated the videos viewed regarding presence, visual comfort, and quality. The results indicate that with the test protocol followed, non-stereoscopic video viewing leads to slightly better presence, visual comfort, and quality ratings compared to the stereoscopic variants. Further, the stereoscopic 360° videos may suffer from visual artefacts potentially leading to lower video quality and further lower quality of experience results. The exploration behaviour was found to be very similar for both non-stereoscopic and stereoscopic video viewing. Overall, it can be concluded that there is a slight tendency for non-stereoscopic video viewing to be preferred over stereoscopic video viewing. The dataset is made publicly available with the paper and includes both variants of all source videos along with the subjective data, and behaviour data, following an open-science approach.
Stephan Fremerey, Raja Faseeh Uz Zaman, Touseef Ashraf, Rakesh Rao Ramachandra Rao, Steve Goering, Alexander Raake
ISM1
2022 Subjective Evaluation of Visual Quality and Simulator Sickness of Short 360$^\circ$ Videos: ITU-T Rec. P.919
abstract
Recently an impressive development in immersive technologies, such as Augmented Reality (AR), Virtual Reality (VR) and 360${^\circ }$video, has been witnessed. However, methods for quality assessment have not been keeping up. This paper studies quality assessment of 360${^\circ }$video from the cross-lab tests (involving ten laboratories and more than 300 participants) carried out by the Immersive Media Group (IMG) of the Video Quality Experts Group (VQEG). These tests were addressed to assess and validate subjective evaluation methodologies for 360${^\circ }$video. Audiovisual quality, simulator sickness symptoms, and exploration behavior were evaluated with short (from 10 seconds to 30 seconds) 360${^\circ }$sequences. The following factors’ influences were also analyzed: assessment methodology, sequence duration, Head-Mounted Display (HMD) device, uniform and non-uniform coding degradations, and simulator sickness assessment methods. The obtained results have demonstrated the validity of Absolute Category Rating (ACR) and Degradation Category Rating (DCR) for subjective tests with 360${^\circ }$videos, the possibility of using 10-second videos (with or without audio) when addressing quality evaluation of coding artifacts, as well as any commercial HMD (satisfying minimum requirements). Also, more efficient methods than the long Simulator Sickness Questionnaire (SSQ) have been proposed to evaluate related symptoms with 360${^\circ }$videos. These results have been instrumental for the development of the ITU-T Recommendation P.919. Finally, the annotated dataset from the tests is made publicly available for the research community.
Jesús Gutiérrez 0001, Pablo Pérez 0001, Marta Orduna, Ashutosh Singla, Carlos Cortés 0001, Pramit Mazumdar, Irene Viola 0001, Kjell Brunnström, Federica Battisti, Natalia Cieplinska, Dawid Juszka, Lucjan Janowski, Mikolaj Leszczuk, Anthony Adeyemi-Ejeye, Yaosi Hu, Zhenzhong Chen 0001, Glenn Van Wallendael, Peter Lambert, César Díaz, John Hedlund, Omar Hamsis, Stephan Fremerey, Frank Hofmeyer, Alexander Raake, Pablo César, Marco Carli, Narciso García
IEEE Trans. Multim.22
2021 AVrate Voyager: an open source online testing platform
abstract
Subjective testing is an integral part of many research fields considering, e.g., human perception. For this purpose, lab tests are a popular approach to gather ratings for subjective evaluations. However, not in all cases controlled lab tests can be performed, either in cases where no labs are existing, accessible or it may be disallowed to use them. For this reason, online tests, e.g., using crowdsourcing are supposed to be an alternative approach for traditional lab tests. We describe in the following paper a framework to implement such online tests for audio, video, and image-related evaluations or questionnaires. Our framework AVrate Voyager builds upon previously developed frameworks for lab tests including the experience with them. AVrate Voyager uses scalable web technologies to implement a test framework, this ensures that it will be running reliably. In addition, we added strategies for pre-caching to avoid additional influence for play-out, e.g. in the case of video testing. We analyze several conducted tests using the new framework and describe the required steps to modify the provided tool in detail.
Steve Goering, Rakesh Rao Ramachandra Rao, Stephan Fremerey, Alexander Raake
MMSP3
2021 Assessment of the Simulator Sickness Questionnaire for Omnidirectional Videos
abstract
Virtual Reality/360° videos provide an immersive experience to users. Besides this, 360° videos may lead to an undesirable effect when consumed with Head-Mounted Displays (HMDs), referred to as simulator sickness/cybersickness. The Simulator Sickness Questionnaire (SSQ) is the most widely used questionnaire for the assessment of simulator sickness. Since the SSQ with its 16 questions was not designed for 360° video related studies, our research hypothesis in this paper was that it may be simplified to enable more efficient testing for 360° video. Hence, we evaluate the SSQ to reduce the number of questions asked from subjects, based on six different previously conducted studies. We derive the reduced set of questions from the SSQ using Principal Component Analysis (PCA) for each test. Pearson Correlation is analysed to compare the relation of all obtained reduced questionnaires as well as two further variants of SSQ reported in the literature, namely Virtual Reality Sickness Questionnaire (VRSQ) and Cybersickness Questionnaire (CSQ). Our analysis suggests that a reduced questionnaire with 9 out of 16 questions yields the best agreement with the initial SSQ, with less than 44% of the initial questions. Exploratory Factor Analysis (EFA) shows that the nine symptom-related attributes determined as relevant by PCA also appear to be sufficient to represent the three dimensions resulting from EFA, namely, Uneasiness, Visual Discomfort and Loss of Balance. The simplified version of the SSQ has the potential to be more efficiently used than the initial SSQ for 360° video by focusing on the questions that are most relevant for individuals, shortening the required testing time.
Ashutosh Singla, Steve Goering, Dominik Keller, Rakesh Rao Ramachandra Rao, Stephan Fremerey, Alexander Raake
VR5
2020 Between the Frames - Evaluation of Various Motion Interpolation Algorithms to Improve 360° Video Quality
abstract
With the increasing availability of 360° video content, it becomes important to provide smoothly playing videos of high quality for end users. For this reason, we compare the influence of different Motion Interpolation (MI) algorithms on 360° video quality. After conducting a pre-test with 12 video experts in [3], we found that MI is a useful tool to increase the QoE (Quality of Experience) of omnidirectional videos. As a result of the pretest, we selected three suitable MI algorithms, namely ffmpeg Motion Compensated Interpolation (MCI), Butterflow and Super-SloMo. Subsequently, we interpolated 15 entertaining and realworld omnidirectional videos with a duration of 20 seconds from 30 fps (original framerate) to 90 fps, which is the native refresh rate of the HMD used, the HTC Vive Pro. To assess QoE, we conducted two subjective tests with 24 and 27 participants. In the first test we used a Modified Paired Comparison (M-PC) method, and in the second test the Absolute Category Rating (ACR) approach. In the M-PC test, 45 stimuli were used and in the ACR test 60. Results show that for most of the 360° videos, the interpolated versions obtained significantly higher quality scores than the lower-framerate source videos, validating our hypothesis that motion interpolation can improve the overall video quality for 360° video. As expected, it was observed that the relative comparisons in the M-PC test result in larger differences in terms of quality. Generally, the ACR method lead to similar results, while reflecting a more realistic viewing situation. In addition, we compared the different MI algorithms and can conclude that with sufficient available computing power Super-SloMo should be preferred for interpolation of omnidirectional videos, while MCI also shows a good performance.
Stephan Fremerey, Frank Hofmeyer, Steve Goering, Dominik Keller, Alexander Raake
ISM1
2020 Subjective Test Dataset and Meta-data-based Models for 360° Streaming Video Quality
abstract
During the last years, the number of 360° videos available for streaming has rapidly increased, leading to the need for 360° streaming video quality assessment. In this paper, we report and publish results of three subjective 360° video quality tests, with conditions used to reflect real-world bitrates and resolutions including 4K, 6K and 8K, resulting in 64 stimuli each for the first two tests and 63 for the third. As playout device we used the HTC Vive for the first and HTC Vive Pro for the remaining two tests. Video-quality ratings were collected using the 5-point Absolute Category Rating scale. The 360° dataset provided with the paper contains the links of the used source videos, the raw subjective scores, video-related meta-data, head rotation data and Simulator Sickness Questionnaire results per stimulus and per subject to enable reproducibility of the provided results. Moreover, we use our dataset to compare the performance of state-of-the-art full-reference quality metrics such as VMAF, PSNR, SSIM, ADM2, WS-PSNR and WS-SSIM. Out of all metrics, VMAF was found to show the highest correlation with the subjective scores. Further, we evaluated a center-cropped version of VMAF ("VMAF-cc") that showed to provide a similar performance as the full VMAF. In addition to the dataset and the objective metric evaluation, we propose two new video-quality prediction models, a bitstream meta-data-based model and a hybrid no-reference model using bitrate, resolution and pixel information of the video as input. The new lightweight models provide similar performance as the full-reference models while enabling fast calculations.
Stephan Fremerey, Steve Goering, Rakesh Rao Ramachandra Rao, Rachel Huang, Alexander Raake
MMSP1
2020 A Large-scale Evaluation of the bitstream-based video-quality model ITU-T P.1204.3 on Gaming Content
abstract
The streaming of gaming content, both passive and interactive, has increased manifolds in recent years. Gaming contents bring with them some peculiarities which are normally not seen in traditional 2D videos, such as the artificial and synthetic nature of contents or repetition of objects in a game. In addition, the perception of gaming content by the user is different from that of traditional 2D videos due to its pecularities and also the fact that users may not often watch such content. Hence, it becomes imperative to evaluate whether the existing video quality models usually designed for traditional 2D videos are applicable to gaming content. In this paper, we evaluate the applicability of the recently standardized bitstream-based video-quality model ITU-T P.1204.3 on gaming content. To analyze the performance of this model, we used 4 different gaming datasets (3 publicly available + 1 internal) not previously used for model training, and compared it with the existing state-of-the-art models. We found that the ITU P.1204.3 model out of the box performs well on these unseen datasets, with an RMSE ranging between 0.38 - 0.45 on the 5-point absolute category rating and Pearson Correlation between 0.85 - 0.93 across all the 4 databases. We further propose a full-HD variant of the P.1204.3 model, since the original model is trained and validated which targets a resolution of 4K/UHD-1. A 50:50 split across all databases is used to train and validate this variant so as to make sure that the proposed model is applicable to various conditions.
Rakesh Rao Ramachandra Rao, Steve Goering, Robert Steger, Saman Zad Tootaghaj, Nabajeet Barman, Stephan Fremerey, Sebastian Möller 0001, Alexander Raake
MMSP6
2020 Development and Evaluation of a Test Setup to Investigate Distance Differences in Immersive Virtual Environments
abstract
Nowadays, with recent advances in virtual reality technology, it is easily possible to integrate real objects into virtual environments by creating an exact virtual replication and enabling interaction with them by mapping the obtained tracking data of the real to the virtual objects. The primary goal of our study is to develop a system to investigate distance differences for near-field interaction in immersive virtual environments. In this context, the term distance difference refers to the shift between a real object and the respective replication of the real object in the virtual environment of the same size. This could occur for a number of reasons e.g. due to errors in motion tracking or mistakes in designing the virtual environment. Our virtual environment is developed using the Unity3D game engine, while the immersive contents were displayed on an HTC Vive Pro head-mounted display. The virtual room shown to the user includes a replication of the real testing lab environment, while one of the two real objects is tracked and mirrored to the virtual world using an HTC Vive Tracker. Both objects are present in the real as well as in the virtual world. To find perceivable distance differences in the near-field, the actual task in the subjective test was to pick up one object and place it into another object. The position of the static object in the virtual world is shifted by values between 0 and 4 cm, while the position of the real object is kept constant. The system is evaluated by conducting a subjective proof-of-concept test with 18 test subjects. The distance difference is evaluated by the subjects through estimating perceived confusion on a modified 5-point absolute category rating scale. The study provides quantitative insights into allowable real-world vs. virtual-world mismatch boundaries for near-field interactions, with a threshold value of around 1 cm.
Stephan Fremerey, Muhammad Sami Suleman, Abdul Haq Azeem Paracha, Alexander Raake
QoMEX1
2019 Impact of Various Motion Interpolation Algorithms on 360° Video QoE
abstract
In our study, we compare the impact of various motion interpolation (MI) algorithms on 360° video Quality of Experience (QoE). For doing so, we conducted a subjective test with 12 video expert viewers, while a pair comparison test method was used. We interpolated four different 20 s long 30 fps 360° source contents to the native 90 Hz refresh rate of popular Head-Mounted Displays using three different MI algorithms. Subsequently, we compared these 90 fps videos against each other to investigate the influence on the QoE. Regarding the algorithms, we found out that ffmpeg blend does not lead to a significant improvement of QoE, while MCI and butterflow do so. Additionally, we concluded that for 360° videos containing fast and sudden movements, MCI should be preferred over butterflow, while butterflow is more suitable for slow and medium motion videos. While comparing the time needed for rendering the 90 fps interpolated videos, ffmpeg blend is the fastest, while MCI and butterflow need much more time.
Stephan Fremerey, Frank Hofmeyer, Steve Goering, Alexander Raake
QoMEX1
2018 AVtrack360: an open dataset and software recording people's head rotations watching 360° videos on an HMD
abstract
In this paper, we present a viewing test with 48 subjects watching 20 different entertaining omnidirectional videos on an HTC Vive Head Mounted Display (HMD) in a task-free scenario. While the subjects were watching the contents, we recorded their head movements. The obtained dataset is publicly available in addition to the links and timestamps of the source contents used. Within this study, subjects were also asked to fill in the Simulator Sickness Questionnaire (SSQ) after every viewing session. Within this paper, at first SSQ results are presented. Several methods for evaluating head rotation data are presented and discussed. In the course of the study, the collected dataset is published along with the scripts for evaluating the head rotation data. The paper presents the general angular ranges of the subjects' exploration behavior as well as an analysis of the areas where most of the time was spent. The collected information can be presented as head-saliency maps, too. In case of videos, head-saliency data can be used for training saliency models, as information for evaluating decisions during content creation, or as part of streaming solutions for region-of-interest-specific coding as with the latest tile-based streaming solutions, as discussed also in standardization bodies such as MPEG.
Stephan Fremerey, Ashutosh Singla, Kay Meseberg, Alexander Raake
MMSys1
2017 Measuring and comparing QoE and simulator sickness of omnidirectional videos in different head mounted displays
abstract
In this paper, we evaluated and compared the integral quality of different omnidirectional contents for two head mounted displays (HMDs), namely HTC Vive and Oculus Rift. We also investigated motion sickness and head-movements. To this aim, we categorized omnidirectional contents into three categories based on the degree of motion: high, medium and low motion. For assessing simulator sickness, we used the Simulator Sickness Questionnaire for each of the contents in both HMDs. The viewing direction for subjects while watching the contents were recorded in terms of the three coordinates yaw, roll and pitch. Experimental results show that HTC Vive offers better integral quality compared to Oculus Rift. We also compared simulator sickness scores along with the behavioral data for different contents and HMDs and discussed the results in the paper.
Ashutosh Singla, Stephan Fremerey, Werner Robitza, Alexander Raake
QoMEX2