Saeed Shafiee Sabet

dblp:193/6786 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0001-5348-8546ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 7 first-author · 22 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 SportSBD: Shot Boundary Detection in Sports Footage
abstract
Shot Boundary Detection (SBD), which identifies scene (or “shot”) changes, is a core step in video analysis pipelines such as summarization and highlight generation. Yet, it remains challenging in sports broadcasts because rapid camera motion, frequent camera switches, sport-specific transitions and graphic overlays often cause false detections and poor cross-domain generalization. In this paper, we address shot boundary detection in professional sports broadcasts, focusing on ice hockey and soccer. We propose a sports-oriented model based on a fine-tuned R(2+1)D 3D CNN, trained to detect hard cuts, gradual transitions, and logo-based replay effects. The model is evaluated on both goal-centered clips and full-match broadcast footage, and is benchmarked against two state-of-the-art baselines: TransNetV2 for accuracy and PySceneDetect for efficiency. Our approach consistently achieves higher precision, recall, and F1-score across all evaluation settings, while demonstrating strong cross-league generalization within ice hockey and cross-sport generalization to soccer. We release our pretrained sports-specific SBD model as an open-source Python package, enabling straightforward integration into existing video analysis pipelines.
Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Pål Halvorsen
MMSys3
2025 Hockey2D: A Keypoint-Based Framework for Ice Hockey Rink Localization and Object Mapping
abstract
Accurate localization of players and objects in ice hockey is essential for advanced analytics, tactical analysis, and automated content generation. Traditional homography estimation approaches rely heavily on explicit camera calibration or heuristic methods, which are computationally intensive and sensitive to camera variations. This paper introduces Hockey2D, a robust framework leveraging YOLO-based pose estimation to identify rink keypoints, combined with RANSAC-based homography optimization for precise spatial mapping. To handle broadcast variations and occlusions effectively, a shot-type classifier filters input frames, selectively applying homography estimation to suitable scenes. By integrating keypoint and object detection, our approach achieves accurate 2D localization of players, referees, and goalkeepers. Evaluations conducted on ice hockey-specific datasets confirm both computational efficiency and localization accuracy. The proposed method highlights the practical viability of AI-driven localization for interactive storytelling, game summarization, and tactical coaching. A video demonstration is available at https://youtu.be/JCnX4N4fi8I.
Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
CBMI3
2025 VoiceVision: AI-Powered Speaker-Aware Cropping and Content Indexing for Multi-Speaker Videos
abstract
VoiceVision is an AI-powered system designed for intelligent speaker-focused video cropping and speaker-aware content indexing. Built on top of the TalkNet audio-visual speaker diarization backbone, Voice Vision detects active speakers in multi-speaker videos and dynamically crops and reframes the video to center on the current speaker, creating smooth visual transitions. In addition to smart cropping, the system integrates automatic speech recognition (ASR) using Whisper to generate accurate transcriptions, which are further processed through a transcript attribution module to associate spoken segments with specific speakers. A dedicated speech search module enables efficient retrieval and indexing of content based on keywords or speaker identity. Voice Vision supports automatic aspect ratio adaptation (9:16, 1:1, 4:5) to generate social-media- optimized outputs. By combining speaker-aware video cropping with searchable, speaker-attributed transcriptions, Voice Vision simplifies content creation, indexing, and sharing for interview- style or conversational videos on platforms such as TikTok, Instagram, and YouTube Shorts. A demonstration of the system's capabilities is presented at https://youtu.be/SBSqOyMpe60
Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
CBMI3
2025 HockeyAI: A Multi-Class Ice Hockey Dataset for Object Detection
abstract
The fast paced nature of ice hockey presents unique challenges for object detection, particularly in tracking the puck---a small, fast moving object that is critical to gameplay analysis. This paper introduces HockeyAI, a novel open source dataset specifically designed for multi-class object detection in ice hockey. The dataset includes 2,101 high resolution frames extracted from professional games in the Swedish Hockey League (SHL), annotated in the You Look Only Once (YOLO) format. Annotations span 7 classes, covering dynamic objects such as the players and the puck, as well as static rink elements such as goalposts and face-off circles. The dataset is derived from diverse SHL games across multiple seasons and teams, ensuring a rich variety of scenarios reflective of real world gameplay. A fine tuned YOLOv8 medium model is also provided, demonstrating high performance across all classes. Key comparisons highlight the dataset's advancements over existing resources, addressing their limitations such as incomplete class coverage, low resolution, and inconsistent annotations. The dataset, model, and an interactive demo are publicly available on Hugging Face under an open source license, fostering further collaboration in sports related computer vision applications (https://huggingface.co/SimulaMet-HOST/HockeyAI).
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
MMSys4
2025 HockeyRink: A Dataset for Precise Ice Hockey Rink Keypoint Mapping and Analytics
abstract
Precise mapping of ice hockey rinks is critical for applications such as player tracking, game strategy analysis, and broadcast enhancements. Traditional methods often rely on manual annotations or simplistic models that fail to account for the rink's complex geometry and dynamic in-game conditions. To address these limitations, we present HockeyRink, a novel dataset comprising 56 meticulously annotated keypoints corresponding to significant landmarks on a standard hockey rink, including face-off dots, goalposts, and blue lines. Leveraging the YOLOv8-Large pose estimation architecture, we adapted the model to treat the rink as a single 'pose' object, enabling accurate keypoint predictions tailored to the nuances of hockey rink imagery. Our dataset, derived from diverse hockey game footage from the Swedish Hockey League (SHL), facilitates applications in homography estimation, 2D/3D scene mapping, and tactical overlays. By detecting and mapping rink keypoints, one can compute precise transformations between the image plane and the rink's physical dimensions, enabling the overlay of player trajectories and tactical insights onto broadcast footage. This work addresses the unique challenges of ice hockey, such as occlusions, rapid camera movements, and varying lighting conditions, and offers a robust foundation for future research and innovation in sports analytics. The HockeyRink dataset and trained model are openly available at: https://huggingface.co/SimulaMet-HOST/HockeyRink.
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
MMSys4
2025 HockeyOrient: A Dataset for Ice Hockey Player Orientation Classification
abstract
Understanding player orientation is a critical component of sports analytics, offering insights into gameplay strategies and player behavior. This paper presents HockeyOrient, a novel dataset for classifying the orientation of ice hockey players based on their poses. The dataset comprises 9,700 manually annotated frames, selected randomly and non-sequentially, taken from Swedish Hockey League (SHL) games during the 2023 and 2024 seasons. Each player image is cropped from game footage and categorized into one of eight orientation classes: top, top-right, right, bottom-right, bottom, bottom-left, left, and top-left. The dataset includes diverse scenarios, such as different teams, jersey colors, referees, and goaltenders with unique protective gear. Alongside the dataset, we provide an open-source classification model trained on the dataset using the SqueezeNet architecture, achieving an F1 score of 75% across all the classes. This work addresses a significant gap in ice hockey analytics, enabling advanced player tracking and gameplay analysis. The dataset and the trained model can be accessed publicly on Hugging Face under an open-source license (https://huggingface.co/datasets/SimulaMet-HOST/HockeyOrient).
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
MMSys4
2024 Demo: Creating Player-Specific Soccer Highlight Clips with PlayerTV
abstract
This paper demonstrates PlayerTV, an innovative framework which harnesses state-of-the-art Artificial Intelligence (AI) technologies for automatic player tracking and identification in soccer videos. By integrating object detection and tracking, Optical Character Recognition (OCR), and color analysis, PlayerTV facilitates the generation of player-specific highlight clips from extensive game footage, significantly reducing the manual labor traditionally associated with such tasks. We present how PlayerTV can be run standalone as a core pipeline, as well as through an interactive Graphical User Interface (GUI).
Håkon Maric Solberg, Mehdi Houshmand Sarkhoosh, Sushant Gautam, Saeed Shafiee Sabet, Pål Halvorsen, Cise Midoglu
CBMI4
2024 PlayerTV: Advanced Player Tracking and Identification for Automatic Soccer Highlight Clips
abstract
In the rapidly evolving field of sports analytics, the automation of targeted video processing is a pivotal advancement. We propose PlayerTV, an innovative framework which harnesses state-of-the-art AI technologies for automatic player tracking and identification in soccer videos. By integrating object detection and tracking, Optical Character Recognition (OCR), and color analysis, PlayerTV facilitates the generation of player-specific highlight clips from extensive game footage, significantly reducing the manual labor traditionally associated with such tasks. Preliminary results from the evaluation of our core pipeline, tested on a dataset from the Norwegian Eliteserien league, indicate that PlayerTV can accurately and efficiently identify teams and players, and our interactive Graphical User Interface (GUI) serves as a user-friendly application wrapping this functionality for streamlined use.
Håkon Maric Solberg, Mehdi Houshmand Sarkhoosh, Sushant Gautam, Saeed Shafiee Sabet, Pål Halvorsen, Cise Midoglu
ISM4
2024 AI-Based Cropping of Soccer Videos for Different Social Media Representations
Mehdi Houshmand Sarkhoosh, Sayed Mohammad Majidi Dorcheh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Michael Riegler 0001, Pål Halvorsen
MMM (4)4
2024 TACDEC: Dataset of Tackle Events in Soccer Game Videos
abstract
This paper introduces TACDEC, a dataset of tackle events in soccer game videos. Recognizing the gap in existing open datasets that predominantly focus on official soccer events such as goals and cards, TACDEC targets a comprehensive analysis of tackles --- a critical aspect of soccer that combines technical skills, tactical decision-making, and physical engagement. By leveraging video data from the Norwegian Eliteserien league across multiple seasons, we annotated 425 videos with 4 types of tackle events, categorized into "tackle-live", "tackle-replay", "tackle-live-incomplete", and "tackle-replay-incomplete", yielding a total of 836 event annotations. The dataset offers an unprecedented resource for the development and testing of machine learning models aimed at understanding and analyzing soccer game dynamics. A proof-of-concept classification model demonstrates the dataset's utility, achieving promising results in automatic tackle detection, thereby validating TACDEC's potential to support not only advanced game analytics but also to enhance fan engagement and player development initiatives.
Evan Jåsund Kassab, Håkon Maric Solberg, Sushant Gautam, Saeed Shafiee Sabet, Thomas Torjusen, Michael Riegler 0001, Pål Halvorsen, Cise Midoglu
MMSys4
2024 SmartCrop-H: AI-Based Cropping of Ice Hockey Videos
abstract
Sports multimedia plays a central role in captivating audiences on social media platforms. However, fast-paced sports such as ice hockey pose unique challenges due to their swift gameplay and the small puck size, making object tracking-based video adaptation for social media a complex task. In this context, we introduce SmartCrop-H, an innovative ice hockey video cropping tool powered by advanced AI models. It excels at tracking the puck and ensuring that crucial gameplay remains the center of attention, regardless of the desired target aspect ratio. The tool combines various techniques including object detection, scene detection, outlier detection, and smoothing, to deliver high-quality ratio-adapted videos. In this demonstration, we showcase SmartCrop-H in real-world scenarios through an intuitive step-by-step Graphical User Interface (GUI) that vividly illustrates how the tool works. The demonstration emphasizes the vital role of AI in enhancing the sports viewing experience, and its importance in the dynamic realm of social media content distribution. A video of the demo can be found here: https://youtu.be/rMmYOCM-k7A.
Mohammad Majidi, Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Pål Halvorsen
MMSys4
2024 Multimodal AI-Based Summarization and Storytelling for Soccer on Social Media
abstract
The rapid advancement of technology has been revolutionizing the field of sports media, where there is a growing need for sophisticated data processing methods. Current methodologies for extracting information from soccer broadcast videos to generate game highlights and summaries for social media are predominantly manual and rely heavily on text-based NLP techniques, overlooking the rich visual and auditory information available. In response to this challenge, our research introduces SoccerSum, a tool that innovates in the field by integrating computer vision, audio analysis with advanced language models like GPT-4. This multimodal approach enables automated, enriched content summarization, including detection of players and key field elements, thereby enhancing the metadata used in summarization algorithms. SoccerSum uniquely combines textual and visual data, offering a comprehensive solution for generating accurate, platform-specific content. This development represents a significant advancement in automated, data-driven sports media dissemination, and sets a new benchmark in the realm of soccer information extraction. A video of the demo can be found here: https://youtu.be/za4VIi2ARXY.
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Pål Halvorsen
MMSys4
2024 The SoccerSum Dataset for Automated Detection, Segmentation, and Tracking of Objects on the Soccer Pitch
abstract
This paper introduces SoccerSum, a novel dataset aimed at enhancing object detection and segmentation in video frames depicting the soccer pitch, using footage from the Norwegian Eliteserien league across 2021-2023. With the goal of detecting elements beyond common entities in existing datasets, such as the soccer ball, players and referees, this dataset includes additional annotations for the goal net, corner flag posts, and the penalty mark. SoccerSum also includes the segmentation of key pitch areas such as the penalty and goal boxes for the same frame sequences. Comprising 750 frames annotated with 10 classes for advanced analysis, SoccerSum offers compatibility with existing frameworks, providing a rich dataset for the development of computer vision algorithms. This dataset not only serves as a resource for improving sports analytics, but also introduces a new application for automatic game summarization, enabling the generation of detailed and engaging content for fans and professionals. The SoccerSum dataset is accessible on Zenodo: https://zenodo.org/records/10612084.
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Thomas Torjusen, Pål Halvorsen
MMSys4
2023 SmartCrop: AI-Based Cropping of Soccer Videos
abstract
In the rapidly evolving landscape of digital platforms, the need for optimizing media representations to cater to various aspect ratios is palpable. In this paper, we pioneer an approach that utilizes object detection, scene detection, outlier detection, and interpolation for smart cropping. Using soccer as a case study, our primary goal is to capture the frame salience using object (player and ball) detection and tracking using AI models. To improve the object detection and tracking, we rely on scene understanding and explore various outlier detection and interpolation techniques. Our pipeline, called SmartCrop, is efficient, and supports various configurations for object tracking, interpolation, and outlier detection to find the best point-of-interest to be used as the cropping center of the video frame. An objective evaluation of the performance of individual pipeline components has validated our proposed architecture and the need for object, scene, outlier detection, and interpolation. Moreover, a crowdsourced subjective user study, assessing the alternative approaches for cropping from 16:9 to 1:1 and 9:16 aspect ratios, confirms that our proposed approach increases the end-user Quality of Experience (QoE).
Sayed Mohammad Majidi Dorcheh, Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Michael Riegler 0001, Dag Johansen, Pål Halvorsen
ISM4
2022 A Virtual Reality Talking Avatar for Investigative Interviews of Maltreat Children
abstract
Interviews conducted with the maltreated children are often the primary source of evidence in prosecution. Many alleged incidents of abuse are not prosecuted because the children’s testimony is collected in an unreliable way. Research shows the consistent poor quality of these interviews and highlights the need for better training of Child Protection Services (CPS) and police personnel who interview abused child witnesses. The currently available systems for training of CPS and police personnel are developed in a rigid way that lag behind in generating dynamic responses. Moreover, these systems require human input such as employing an actor mimicking a child or an operator controlling prerecorded child responses during the interactions. This paper demonstrates the prototype of an interview training program with an artificial intelligent Child Avatar in Virtual Reality (VR), enabling CPS and police personnel to practice interviewing with abused children. The program is developed using Unity game engine and artificial intelligence-based technologies such as dialogue models, talking visual avatars, text-to-speech, and speech-to-text components.
Syed Zohaib Hassan, Pegah Salehi, Michael Riegler 0001, Miriam S. Johnson, Gunn Astrid Baugerud, Pål Halvorsen, Saeed Shafiee Sabet
CBMI7
2022 A Comparative Study of Interactive Environments for Investigative Interview of A Virtual Child Avatar
abstract
In cases of suspected child abuse, police and child protection services (CPS) personnel often conduct investigative interviews. These interviews are intended to obtain a reliable account from the child about the alleged crime so that proper legal and clinical interventions can be implemented, with the prosecution presenting the case in court. Studies show that the quality of these interviews is often poor, and interviewers’ current training programs are not that efficient. Training must be spaced and repeated over time. In this paper, we present a system that simulates the scenario of interviewing a victim of child sexual abuse in four different interactive environments. We conducted a user study in which participating experts interviewed a simulated child to determine which interactive environment would provide the highest quality of experience (QoE) and motivates experts for higher practice time with the system. The experts include CPS workers and child welfare students. This study measures each user interaction’s overall QoE, realism, responsiveness, presence, and flow as well as the learning effects, engagement in learning, and self-efficacy. The study shows that 66% of the participants would prefer to use virtual reality (VR) for interactive training over other environments. VR had the highest rate in most of the assessment metrics.
Syed Zohaib Hassan, Saeed Shafiee Sabet, Pegah Salehi, Hayley Ko, Ingvild Riiser, Miriam S. Johnson, Gunn Astrid Baugerud, Michael Riegler 0001, Pål Halvorsen
ISM2
2022 Experiences and Lessons Learned from a Crowdsourced-Remote Hybrid User Survey Framework
abstract
Subjective user studies are important to ensure the fidelity and usability of systems that generate multimedia content. Testing how end-users and domain experts perceive multimedia assets might provide crucial information. In this paper, we present our experiences with the open source hybrid crowdsourced-remote user survey framework called Huldra, which is intended for conducting web-based subjective user studies and aims to integrate the individual benefits associated with traditional, crowdsourced, and remote methods. We disseminate our experiences and insights from two actively deployed use cases and discuss challenges and opportunities associated with using Huldra as a framework for conducting user studies.
Cise Midoglu, Andrea M. Storås, Saeed Shafiee Sabet, Malek Hammou, Steven Alexander Hicks, Inga Strümke, Michael Riegler 0001, Carsten Griwodz, Pål Halvorsen
ISM3
2022 Huldra: a framework for collecting crowdsourced feedback on multimedia assets
abstract
Collecting crowdsourced feedback to evaluate, rank, or score multimedia content can be cumbersome and time-consuming. Most of the existing survey tools are complicated, hard to customize, or tailored for a specific asset type. In this paper, we present an open source framework called Huldra, designed explicitly to address the challenges associated with user studies involving crowdsourced feedback collection. The web-based framework is built in a modular and configurable fashion to allow for the easy adjustment of the user interface (UI) and the multimedia content, while providing integrations with reliable and stable backend solutions to facilitate the collection and analysis of responses. Our proposed framework can be used as an online survey tool by researchers working on different topics such as Machine Learning (ML), audio, image, and video quality assessment, Quality of Experience (QoE), and require user studies for the benchmarking of various types of multimedia content.
Malek Hammou, Cise Midoglu, Steven Alexander Hicks, Andrea M. Storås, Saeed Shafiee Sabet, Inga Strümke, Michael Riegler 0001, Pål Halvorsen
MMSys5
2022 Human vs. GPT-3: The challenges of extracting emotions from child responses
abstract
Conducting interviews with abused children requires a specific skill set that is obtained by undergoing special training. In addition to acquiring the interview training skills, these also have to be constantly refreshed to keep a high level of quality. Technology, such as synthetic video generation and natural language processing, is in a stage that could allow the construction of a system that can make this task easier. Thus, we aim to design a training system aided by machine learning that can support the interview training with an interactive child avatar capable of meaningful interactions with the trainees. In these interviews, emotions play an important role, so we conduct three different user studies in a remote study setting with the aim of analyzing child emotions in these interviews. In these user studies, the participants had to classify different transcripts excerpts as one of the possible predefined emotions. These human annotations are used to measure the performance of sentiment analysis using GPT-3. We investigate different approaches to obtain the correct classifications by changing the amount of context the participants and the model get to see. Our experiments show that humans have a hard time agreeing when choosing between seven different emotions. This improves when we reduce the set of emotions to four. In addition, we found that context is needed to make a motivated choice, but too much context can make it vague, reducing the judgment's quality.
Myrthe Lammerse, Syed Zohaib Hassan, Saeed Shafiee Sabet, Michael Riegler 0001, Pål Halvorsen
QoMEX3
2022 When Every Millisecond Counts: The Impact of Delay in VR Gaming
abstract
This paper presents the finding of a subjective study investigating the impact of delay on a VR gaming experience, applying an experimental design to a virtual game of squash. 32 participants were asked to rate the stability of the VR experience while playing the game, with results showing an impact of delay even at the lowest level of 20 ms added delay. With 120 ms added delay, most participants judged the experience to be fully unstable. Exploring individual differences, we also found that males and experienced gamers were more critical, or more sensitive, to higher levels of delay, compared to females and less experienced gamers.
Saeed Shafiee Sabet, Ragnhild Eg, Kjetil Raaen, Muhammad Oasim, Michael Riegler 0001, Pål Halvorsen
QoMEX1
2022 Comparison of Crowdsourced and Remote Subjective User Studies: A Case Study of Investigative Child Interviews
abstract
Crowdsourced and remote user studies have recently gained popularity as alternatives to traditional laboratory studies. However, they are subject to unreliability, and it is challenging to ensure that valid results are collected, especially when conducting user studies with experts. Experts are a sparse resource, usually having busy schedules and heavy workloads, and are not necessarily geographically close. They are therefore often unwilling to participate in studies which require physical attendance. In this paper, we compare three alternative methods: crowd sourced user study with non-experts, remote user study with non-experts, and remote user study with domain experts, for a use case involving investigative child interview training. We present the results from three subjective studies about the perception of AI-generated child avatars, which is developed using various technologies such as dialogue models, game engine, text-to-speech and speech-to-text components. The study was conducted with three different user groups, and our results indicate the importance of using best practice measures for ensuring the collection of reliable results in crowdsourced settings as compared to remote studies, and highlight the difference between the perspectives of domain experts and non-experts.
Saeed Shafiee Sabet, Cise Midoglu, Syed Zohaib Hassan, Pegah Salehi, Gunn Astrid Baugerud, Carsten Griwodz, Miriam S. Johnson, Michael Riegler 0001, Pål Halvorsen
QoMEX1
2021 Modeling and Understanding the Quality of Experience of Online Mobile Gaming Services
abstract
Mobile gaming has the largest market shares of all gaming domains, accounting for an estimated $ 77.2 billion in 2020. In recent times, one can witness an increase in highly interactive mobile online games. However, the gaming Quality of Experience (QoE) can be strongly influenced by network degradations, concretely by delay and packet loss. Thus, network providers need to ensure fast and reliable connections between the gaming servers and the users' clients. To maintain a satisfying user experience, QoE prediction models are fundamental. Aiming at the development of such a model, a detailed parameter space consisting of various delay and packet loss conditions will be investigated in this paper. Here, especially the importance of jitter is of interest. Next, it will be examined whether a recently published opinion model for cloud gaming, the ITU-T Rec. G.1072, can also be used for online mobile gaming. Finally, a new proposal for a model targeting online mobile gaming services will be presented and evaluated concerning its performance.
Steven Schmidt 0001, Saman Zad Tootaghaj, Saeed Shafiee Sabet, Sebastian Möller 0001
QoMEX3
2020 A latency compensation technique based on game characteristics to mitigate the influence of delay on cloud gaming quality of experience
abstract
Cloud Gaming (CG) is an immersive multimedia service that promises many benefits. In CG, the games are rendered in a cloud server, and the resulted scenes are streamed as a video sequence to the client. Using CG users are not forced to update their gaming hardware frequently, and available games can be played on any operating system or suitable device. However, cloud gaming requires a reliable and low-latency network, which makes it a very challenging service. Transmission latency strongly affects the playability of a cloud game and consequently reduces the users' Quality of Experience (QoE). In this paper, we propose a latency compensation technique using game adaptation that mitigates the influence of delay on QoE. This technique uses five game characteristics for the adaptation. These characteristics, in addition to an Aim-assistance technique, were implemented in four games for evaluation. A subjective study using 194 participants was conducted using a crowdsourcing approach. The results showed that the majority of the proposed adaptation techniques lead to significant improvements in the cloud gaming QoE.
Saeed Shafiee Sabet, Steven Schmidt 0001, Saman Zad Tootaghaj, Babak Naderi, Carsten Griwodz, Sebastian Möller 0001
MMSys1
2020 Quality estimation models for gaming video streaming services using perceptual video quality dimensions
abstract
The gaming industry is one of the largest digital markets for decades and is steady developing as evident by new emerging gaming services such as gaming video streaming, online gaming, and cloud gaming. While the market is rapidly growing, the quality of these services depends strongly on network characteristics as well as resource management. With the advancement of encoding technologies such as hardware accelerated engines, fast encoding is possible for delay sensitive applications such as cloud gaming. Therefore, already existing video quality models do not offer a good performance for cloud gaming applications. Thus, in this paper, we provide a gaming video quality dataset that considers hardware accelerated engines for video compression using the H.264 standard. In addition, we investigate the performance of signal-based and parametric video quality models on the new gaming video dataset. Finally, we build two novel parametric-based models, a planning and a monitoring model, for gaming quality estimation. Both models are based on perceptual video quality dimensions and can be used to optimize the resource allocation of gaming video streaming services.
Saman Zad Tootaghaj, Steven Schmidt 0001, Saeed Shafiee Sabet, Sebastian Möller 0001, Carsten Griwodz
MMSys3
2020 Assessing Interactive Gaming Quality of Experience using a Crowdsourcing Approach
abstract
Traditionally, the Quality of Experience (QoE) is assessed in a controlled laboratory environment where participants give their opinion about the perceived quality of a stimulus on a standardized rating scale. Recently, the usage of crowdsourcing micro-task platforms for assessing the media quality is increasing. The crowdsourcing platforms provide access to a pool of geographically distributed, and demographically diverse group of workers who participate in the experiment in their own working environment and using their own hardware. The main challenge in crowdsourcing QoE tests is to control the effect of interfering influencing factors such as a user's environment and device on the subjective ratings. While in the past, the crowdsourcing approach was frequently used for speech and video quality assessment, research on a quality assessment for gaming services is rare. In this paper, we present a method to measure gaming QoE under typically considered system influence factors including delay, packet loss, and framerates as well as different game designs. The factors are artificially manipulated due to controlled changes in the implementation of games. The results of a total of five studies using a developed evaluation method based on a combination of the ITU-T Rec. P.809 on subjective evaluation methods for gaming quality and the ITU-T Rec. P.808 on subjective evaluation of speech quality with a crowdsourcing approach will be discussed. To evaluate the reliability and validity of results collected using this method, we finally compare subjective ratings regarding the effect of network delay on gaming QoE gathered from interactive crowdsourcing tests with those from equivalent laboratory experiments.
Steven Schmidt 0001, Babak Naderi, Saeed Shafiee Sabet, Saman Zad Tootaghaj, Sebastian Möller 0001
QoMEX3
2020 Towards the Impact of Gamers Strategy and User Inputs on the Delay Sensitivity of Cloud Games
abstract
Cloud Gaming is an emerging service that is considered by many as the future of the gaming industry. This service requires a highly reliable network with low latency and high bandwidth. If these requirements are not satisfied, cloud gaming services cannot create a good Quality of Experience (QoE) for its users. However, gaming QoE can vary significantly among different game scenarios and users. For an optimal resource allocation and quality estimation, it is highly important for cloud providers, game developers, and network planners to consider the influence of the game content and gamers. This paper presents the result of a subjective study that investigated the impact of different player strategies and user inputs on their perceived delay. The results indicated that the user input characteristics vary among the games but stays the same between different users and different strategies. In addition to the users' inputs, the input quality and the overall gaming experience of the users were also investigated, and results did not show any main effect of user strategy on the delay sensitivity of the games.
Saeed Shafiee Sabet, Steven Schmidt 0001, Saman Zad Tootaghaj, Carsten Griwodz, Sebastian Möller 0001
QoMEX1
2019 Towards the Impact of Gamers' Adaptation to Delay Variation on Gaming Quality of Experience
abstract
Both online and cloud gaming services require a very low network delay to create a good Quality of Experience (QoE) for their users. The required network latency cannot be guaranteed due to the current best effort-nature of the network, and as a result, network latency often degrades the gamers performance and QoE. In this paper, the adaptability of gamers to different variations on delay is investigated both subjectively and objectively using three self-developed games. The results show that gamers can adapt to constant delay while they are playing and change their behavior if the actions in a game are predictable. Such adaptation leads to a significant increase in gamers performance and QoE. The paper also provides evidence that regardless of performance frequent delay switching annoys gamers. The result of this study can be used to create a network resource allocation technique which controls a congested network by giving more priority and resource to the unadaptable games than the adaptable games.
Saeed Shafiee Sabet, Steven Schmidt 0001, Carsten Griwodz, Sebastian Möller 0001
QoMEX1
2018 Towards Applying Game Adaptation to Decrease the Impact of Delay on Quality of Experience
abstract
With emerging delay sensitive gaming services such as cloud gaming and online gaming, the importance of understanding and reducing the effect of delay on the gamer's Quality of Experience (QoE) becomes highly important for the success of these services. In this paper, the findings of two subjective experiments investigating the relationship between delay and QoE are reported. In the first study, it was shown that in addition to the direct effect of the delay on QoE, there is a significant indirect effect between delay and QoE through the relationship with performance. In the second part of the paper, we illustrate that adapting characteristics of a game can strongly mitigate the negative effect of delay on gaming QoE due to increased player performance. This adaptation in addition to compensation the effect of the delay, in contrast to the other difficulty adjustment systems, does not require to track the gamer's interaction, behaviors, and profile.
Saeed Shafiee Sabet, Steven Schmidt 0001, Saman Zad Tootaghaj, Carsten Griwodz, Sebastian Möller 0001
ISM1
2016 A Testing Apparatus for Faster and More Accurate Subjective Assessment of Quality of Experience in Cloud Gaming
abstract
The number of cloud gaming (CG) users is constantly growing. The idea in cloud gaming is to render the game events on a cloud server and stream the resulted scenes as a video sequence to players. CG requires a high bandwidth in order to run appropriately and create a good quality of experience for players. In order to reduce the high required bandwidth, video should be compressed without any negative impact on user's quality of experience (QoE). Thus CG providers, researchers who develop new compression methods for CG, and those who are improving network protocols for CG require to evaluate user experience using subjective methods. Over the years, many researches have investigated the subjective quality of video, but all of them have one of the following two main drawbacks, which makes them unsuitable for game videos. The subjective quality assessment methods which are designed for short duration video sequences suffer from Forgiveness and Recency effects. On the other hand, the methods which are designed for long duration video sequences usually use some sort of a handset device for rating scores, and hence cannot be used for most games where both hands are busy while playing. In this paper, a novel subjective test apparatus for assessment of game videos is proposed, where players give their opinion scores using a foot pedal while playing the game. Evaluation results indicate that the proposed scheme is more accurate and less distractive than existing methods.
Saeed Shafiee Sabet, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001
ISM1