Cise Midoglu

dblp:178/8808 · DBLP profile ↗
← Back
35ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0003-0991-4418ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 3 first-author · 29 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SportSBD: Shot Boundary Detection in Sports Footage
abstract
Shot Boundary Detection (SBD), which identifies scene (or “shot”) changes, is a core step in video analysis pipelines such as summarization and highlight generation. Yet, it remains challenging in sports broadcasts because rapid camera motion, frequent camera switches, sport-specific transitions and graphic overlays often cause false detections and poor cross-domain generalization. In this paper, we address shot boundary detection in professional sports broadcasts, focusing on ice hockey and soccer. We propose a sports-oriented model based on a fine-tuned R(2+1)D 3D CNN, trained to detect hard cuts, gradual transitions, and logo-based replay effects. The model is evaluated on both goal-centered clips and full-match broadcast footage, and is benchmarked against two state-of-the-art baselines: TransNetV2 for accuracy and PySceneDetect for efficiency. Our approach consistently achieves higher precision, recall, and F1-score across all evaluation settings, while demonstrating strong cross-league generalization within ice hockey and cross-sport generalization to soccer. We release our pretrained sports-specific SBD model as an open-source Python package, enabling straightforward integration into existing video analysis pipelines.
Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Pål Halvorsen
MMSys2
2025 SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
abstract
Artificial intelligence (AI) is transforming sports analytics by enabling automated, real-time understanding of soccer matches. Traditional approaches that rely on isolated data streams struggle to capture the full game context. We introduce SoccerChat, a multimodal conversational AI framework that fuses visual and textual information for comprehensive soccer video comprehension. Building on the SoccerNet dataset, we enrich it with jersey color annotations and automatic speech recognition (ASR) transcripts and curate a video-instruction dataset containing 48,677 question-answer (QA) pairs. Fine-tuning Qwen2- VL-7B- Instruct on this resource yields the Soc-cerChat model, which supports accurate event interpretation, classification, and referee assistance. Experiments across action classification and referee QA tasks demonstrate strong general event understanding and competitive officiating analysis. These results highlight the importance of multimodal integration for explainable, interactive, and trustworthy AI-driven sports ana-lytics. Links to the code, dataset, and model weights are available at https://2ithub.com/simula/SoccerChat.
Sushant Gautam, Cise Midoglu, Vajira Thambawita, Michael Riegler 0001, Pål Halvorsen, Mubarak Shah
CBMI2
2025 Hockey2D: A Keypoint-Based Framework for Ice Hockey Rink Localization and Object Mapping
abstract
Accurate localization of players and objects in ice hockey is essential for advanced analytics, tactical analysis, and automated content generation. Traditional homography estimation approaches rely heavily on explicit camera calibration or heuristic methods, which are computationally intensive and sensitive to camera variations. This paper introduces Hockey2D, a robust framework leveraging YOLO-based pose estimation to identify rink keypoints, combined with RANSAC-based homography optimization for precise spatial mapping. To handle broadcast variations and occlusions effectively, a shot-type classifier filters input frames, selectively applying homography estimation to suitable scenes. By integrating keypoint and object detection, our approach achieves accurate 2D localization of players, referees, and goalkeepers. Evaluations conducted on ice hockey-specific datasets confirm both computational efficiency and localization accuracy. The proposed method highlights the practical viability of AI-driven localization for interactive storytelling, game summarization, and tactical coaching. A video demonstration is available at https://youtu.be/JCnX4N4fi8I.
Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
CBMI2
2025 VoiceVision: AI-Powered Speaker-Aware Cropping and Content Indexing for Multi-Speaker Videos
abstract
VoiceVision is an AI-powered system designed for intelligent speaker-focused video cropping and speaker-aware content indexing. Built on top of the TalkNet audio-visual speaker diarization backbone, Voice Vision detects active speakers in multi-speaker videos and dynamically crops and reframes the video to center on the current speaker, creating smooth visual transitions. In addition to smart cropping, the system integrates automatic speech recognition (ASR) using Whisper to generate accurate transcriptions, which are further processed through a transcript attribution module to associate spoken segments with specific speakers. A dedicated speech search module enables efficient retrieval and indexing of content based on keywords or speaker identity. Voice Vision supports automatic aspect ratio adaptation (9:16, 1:1, 4:5) to generate social-media- optimized outputs. By combining speaker-aware video cropping with searchable, speaker-attributed transcriptions, Voice Vision simplifies content creation, indexing, and sharing for interview- style or conversational videos on platforms such as TikTok, Instagram, and YouTube Shorts. A demonstration of the system's capabilities is presented at https://youtu.be/SBSqOyMpe60
Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
CBMI2
2025 Extracting Player Speed from Football Videos
abstract
One of the key metrics for player and team performance in football is speed. Traditionally, measuring player speed either requires extensive human effort or expensive equipment. In this work, we investigate methods for manual and automatic player speed extraction directly from broadcast football videos. We implement a pipeline combining player detection, tracking, field mapping, and position measurement. We experiment with different configurations of the pipeline, concluding that a setting with YOLOv11 for detection, StrongSORT for tracking, PnL-Calib for field mapping and keypoint detection, and position measurement through homography yields the best results for our current scope and test datasets. This work serves as a proof-of-concept analytics solution for player speed extraction, which can be adopted by teams of all sizes and means, and built upon by researchers. Our pipeline implementation and two newly curated player speed datasets (Begnadalen and TACDEC++) are openly available for the sports and scientific community.
Ole Kristian Rustebakke, Mehdi Houshmand Sarkhoosh, Cise Midoglu, Pål Halvorsen
ISM3
2025 ExposureEngine: Oriented Logo Detection and Sponsor Visibility Analytics in Sports Broadcasts
abstract
Quantifying sponsor visibility in sports broadcasts is a critical marketing task traditionally hindered by manual, subjective, and unscalable analysis methods. While automated systems offer an alternative, their reliance on axis-aligned Horizontal Bounding Box (HBB) leads to inaccurate exposure metrics when logos appear rotated or skewed due to dynamic camera angles and perspective distortions. This paper introduces ExposureEngine, an end-to-end system designed for accurate, rotation-aware sponsor visibility analytics in sports broadcasts, demonstrated in a soccer case study. Our approach predicts Oriented Bounding Box (OBB) to provide a geometrically precise fit to each logo regardless of the orientation on-screen. To train and evaluate our detector, we developed a new dataset comprising 1,103 frames from Swedish elite soccer, featuring 670 unique sponsor logos annotated with OBBs. Our model achieves a mean Average Precision (mAP @ 0.5) of 0.859, with a precision of 0.96 and recall of 0.87, demonstrating robust performance in localizing logos under diverse broadcast conditions. The system integrates these detections into an analytical pipeline that calculates precise visibility metrics, such as exposure duration and on-screen coverage. Furthermore, we incorporate a language-driven agentic layer, enabling users to generate reports, summaries, and media content through natural language queries. The complete system, including the dataset and the analytics dashboard, provides a comprehensive solution for auditable and interpretable sponsor measurement in sports media. An overview of the ExposureEngine is available online. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>1https://youtu.be/tRw6OBISuW4
Mehdi Houshmand Sarkhoosh, Frøy Øye, Henrik Nestor Sørlie, Nam Hoang Vu, Dag Johansen, Cise Midoglu, Tomas Kupka, Pål Halvorsen
ISM6
2025 HockeyAI: A Multi-Class Ice Hockey Dataset for Object Detection
abstract
The fast paced nature of ice hockey presents unique challenges for object detection, particularly in tracking the puck---a small, fast moving object that is critical to gameplay analysis. This paper introduces HockeyAI, a novel open source dataset specifically designed for multi-class object detection in ice hockey. The dataset includes 2,101 high resolution frames extracted from professional games in the Swedish Hockey League (SHL), annotated in the You Look Only Once (YOLO) format. Annotations span 7 classes, covering dynamic objects such as the players and the puck, as well as static rink elements such as goalposts and face-off circles. The dataset is derived from diverse SHL games across multiple seasons and teams, ensuring a rich variety of scenarios reflective of real world gameplay. A fine tuned YOLOv8 medium model is also provided, demonstrating high performance across all classes. Key comparisons highlight the dataset's advancements over existing resources, addressing their limitations such as incomplete class coverage, low resolution, and inconsistent annotations. The dataset, model, and an interactive demo are publicly available on Hugging Face under an open source license, fostering further collaboration in sports related computer vision applications (https://huggingface.co/SimulaMet-HOST/HockeyAI).
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
MMSys3
2025 HockeyRink: A Dataset for Precise Ice Hockey Rink Keypoint Mapping and Analytics
abstract
Precise mapping of ice hockey rinks is critical for applications such as player tracking, game strategy analysis, and broadcast enhancements. Traditional methods often rely on manual annotations or simplistic models that fail to account for the rink's complex geometry and dynamic in-game conditions. To address these limitations, we present HockeyRink, a novel dataset comprising 56 meticulously annotated keypoints corresponding to significant landmarks on a standard hockey rink, including face-off dots, goalposts, and blue lines. Leveraging the YOLOv8-Large pose estimation architecture, we adapted the model to treat the rink as a single 'pose' object, enabling accurate keypoint predictions tailored to the nuances of hockey rink imagery. Our dataset, derived from diverse hockey game footage from the Swedish Hockey League (SHL), facilitates applications in homography estimation, 2D/3D scene mapping, and tactical overlays. By detecting and mapping rink keypoints, one can compute precise transformations between the image plane and the rink's physical dimensions, enabling the overlay of player trajectories and tactical insights onto broadcast footage. This work addresses the unique challenges of ice hockey, such as occlusions, rapid camera movements, and varying lighting conditions, and offers a robust foundation for future research and innovation in sports analytics. The HockeyRink dataset and trained model are openly available at: https://huggingface.co/SimulaMet-HOST/HockeyRink.
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
MMSys3
2025 HockeyOrient: A Dataset for Ice Hockey Player Orientation Classification
abstract
Understanding player orientation is a critical component of sports analytics, offering insights into gameplay strategies and player behavior. This paper presents HockeyOrient, a novel dataset for classifying the orientation of ice hockey players based on their poses. The dataset comprises 9,700 manually annotated frames, selected randomly and non-sequentially, taken from Swedish Hockey League (SHL) games during the 2023 and 2024 seasons. Each player image is cropped from game footage and categorized into one of eight orientation classes: top, top-right, right, bottom-right, bottom, bottom-left, left, and top-left. The dataset includes diverse scenarios, such as different teams, jersey colors, referees, and goaltenders with unique protective gear. Alongside the dataset, we provide an open-source classification model trained on the dataset using the SqueezeNet architecture, achieving an F1 score of 75% across all the classes. This work addresses a significant gap in ice hockey analytics, enabling advanced player tracking and gameplay analysis. The dataset and the trained model can be accessed publicly on Hugging Face under an open-source license (https://huggingface.co/datasets/SimulaMet-HOST/HockeyOrient).
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
MMSys3
2024 Demo: Creating Player-Specific Soccer Highlight Clips with PlayerTV
abstract
This paper demonstrates PlayerTV, an innovative framework which harnesses state-of-the-art Artificial Intelligence (AI) technologies for automatic player tracking and identification in soccer videos. By integrating object detection and tracking, Optical Character Recognition (OCR), and color analysis, PlayerTV facilitates the generation of player-specific highlight clips from extensive game footage, significantly reducing the manual labor traditionally associated with such tasks. We present how PlayerTV can be run standalone as a core pipeline, as well as through an interactive Graphical User Interface (GUI).
Håkon Maric Solberg, Mehdi Houshmand Sarkhoosh, Sushant Gautam, Saeed Shafiee Sabet, Pål Halvorsen, Cise Midoglu
CBMI6
2024 SoccerRAG: Multimodal Soccer Information Retrieval via Natural Queries
abstract
The rapid evolution of digital sports media necessitates sophisticated information retrieval systems that can efficiently parse extensive multimodal datasets. In this paper, we introduce SoccerRAG, an innovative framework designed to harness the power of Retrieval Augmented Generation (RAG) and Large Language Models (LLMs) to extract soccer-related information through natural language queries. By leveraging a multimodal dataset, SoccerRAG supports dynamic querying and automatic data validation, enhancing user interaction and accessibility to sports archives. Our evaluations indicate that SoccerRAG effectively handles complex queries, offering significant improvements over traditional retrieval systems in terms of accuracy and user engagement. The results underscore the potential of using RAG and LLMs in sports analytics, paving the way for future advancements in the accessibility and real-time processing of sports data.
Aleksander Theo Strand, Sushant Gautam, Cise Midoglu, Pål Halvorsen
CBMI3
2024 Demo: Soccer Information Retrieval Via Natural Queries using SoccerRAG
abstract
The rapid evolution of digital sports media necessitates sophisticated information retrieval systems that can efficiently parse extensive multimodal datasets. This paper demonstrates SoccerRAG, an innovative framework designed to harness the power of Retrieval Augmented Generation (RAG) and Large Language Models (LLMs) to extract soccer-related information through natural language queries. By leveraging a multimodal dataset, SoccerRAG supports dynamic querying and automatic data validation, enhancing user interaction and accessibility to sports archives. We present a novel interactive user interface (UI) based on the Chainlit framework which wraps around the core functionality, and enable users to interact with the SoccerRAG framework in a chatbot-like visual manner.
Aleksander Theo Strand, Sushant Gautam, Cise Midoglu, Pål Halvorsen
CBMI3
2024 SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
abstract
The application of Automatic Speech Recognition (ASR) technology in soccer enables sports analytics by extracting audio commentaries to provide insights into game events and facilitate automatic game understanding. This paper presents SoccerNet-Echoes, an extension of the SoccerNet dataset with automatically generated transcriptions of soccer game broadcasts. Generated using the Whisper model and translated with Google Translate into English when needed, these transcriptions enhance video content with textual information derived from game audio. SoccerNet-Echoes serves as a comprehensive resource for developing algorithms in action spotting, caption generation, and game summarization. Through a series of experiments, we demonstrate that combining modalities—audio, video, and text—yields mixed results on classification tasks. The combination of audio and video shows improved performance over individual modalities, while the addition of ASR text does not significantly enhance results. Additionally, our baseline summarization tasks indicate that ASR content enriches summaries, offering insights beyond event information. This multimodal dataset supports diverse applications, broadening the scope of research in sports analytics. The dataset is available at: https://github.com/SoccerNet/sn-echoes.
Sushant Gautam, Mehdi Houshmand Sarkhoosh, Jan Held, Cise Midoglu, Anthony Cioppa, Silvio Giancola, Vajira Thambawita, Michael Riegler 0001, Pål Halvorsen, Mubarak Shah
ISM4
2024 PlayerTV: Advanced Player Tracking and Identification for Automatic Soccer Highlight Clips
abstract
In the rapidly evolving field of sports analytics, the automation of targeted video processing is a pivotal advancement. We propose PlayerTV, an innovative framework which harnesses state-of-the-art AI technologies for automatic player tracking and identification in soccer videos. By integrating object detection and tracking, Optical Character Recognition (OCR), and color analysis, PlayerTV facilitates the generation of player-specific highlight clips from extensive game footage, significantly reducing the manual labor traditionally associated with such tasks. Preliminary results from the evaluation of our core pipeline, tested on a dataset from the Norwegian Eliteserien league, indicate that PlayerTV can accurately and efficiently identify teams and players, and our interactive Graphical User Interface (GUI) serves as a user-friendly application wrapping this functionality for streamlined use.
Håkon Maric Solberg, Mehdi Houshmand Sarkhoosh, Sushant Gautam, Saeed Shafiee Sabet, Pål Halvorsen, Cise Midoglu
ISM6
2024 AI-Based Cropping of Soccer Videos for Different Social Media Representations
Mehdi Houshmand Sarkhoosh, Sayed Mohammad Majidi Dorcheh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Michael Riegler 0001, Pål Halvorsen
MMM (4)3
2024 TACDEC: Dataset of Tackle Events in Soccer Game Videos
abstract
This paper introduces TACDEC, a dataset of tackle events in soccer game videos. Recognizing the gap in existing open datasets that predominantly focus on official soccer events such as goals and cards, TACDEC targets a comprehensive analysis of tackles --- a critical aspect of soccer that combines technical skills, tactical decision-making, and physical engagement. By leveraging video data from the Norwegian Eliteserien league across multiple seasons, we annotated 425 videos with 4 types of tackle events, categorized into "tackle-live", "tackle-replay", "tackle-live-incomplete", and "tackle-replay-incomplete", yielding a total of 836 event annotations. The dataset offers an unprecedented resource for the development and testing of machine learning models aimed at understanding and analyzing soccer game dynamics. A proof-of-concept classification model demonstrates the dataset's utility, achieving promising results in automatic tackle detection, thereby validating TACDEC's potential to support not only advanced game analytics but also to enhance fan engagement and player development initiatives.
Evan Jåsund Kassab, Håkon Maric Solberg, Sushant Gautam, Saeed Shafiee Sabet, Thomas Torjusen, Michael Riegler 0001, Pål Halvorsen, Cise Midoglu
MMSys8
2024 SmartCrop-H: AI-Based Cropping of Ice Hockey Videos
abstract
Sports multimedia plays a central role in captivating audiences on social media platforms. However, fast-paced sports such as ice hockey pose unique challenges due to their swift gameplay and the small puck size, making object tracking-based video adaptation for social media a complex task. In this context, we introduce SmartCrop-H, an innovative ice hockey video cropping tool powered by advanced AI models. It excels at tracking the puck and ensuring that crucial gameplay remains the center of attention, regardless of the desired target aspect ratio. The tool combines various techniques including object detection, scene detection, outlier detection, and smoothing, to deliver high-quality ratio-adapted videos. In this demonstration, we showcase SmartCrop-H in real-world scenarios through an intuitive step-by-step Graphical User Interface (GUI) that vividly illustrates how the tool works. The demonstration emphasizes the vital role of AI in enhancing the sports viewing experience, and its importance in the dynamic realm of social media content distribution. A video of the demo can be found here: https://youtu.be/rMmYOCM-k7A.
Mohammad Majidi, Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Pål Halvorsen
MMSys3
2024 Multimodal AI-Based Summarization and Storytelling for Soccer on Social Media
abstract
The rapid advancement of technology has been revolutionizing the field of sports media, where there is a growing need for sophisticated data processing methods. Current methodologies for extracting information from soccer broadcast videos to generate game highlights and summaries for social media are predominantly manual and rely heavily on text-based NLP techniques, overlooking the rich visual and auditory information available. In response to this challenge, our research introduces SoccerSum, a tool that innovates in the field by integrating computer vision, audio analysis with advanced language models like GPT-4. This multimodal approach enables automated, enriched content summarization, including detection of players and key field elements, thereby enhancing the metadata used in summarization algorithms. SoccerSum uniquely combines textual and visual data, offering a comprehensive solution for generating accurate, platform-specific content. This development represents a significant advancement in automated, data-driven sports media dissemination, and sets a new benchmark in the realm of soccer information extraction. A video of the demo can be found here: https://youtu.be/za4VIi2ARXY.
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Pål Halvorsen
MMSys3
2024 The SoccerSum Dataset for Automated Detection, Segmentation, and Tracking of Objects on the Soccer Pitch
abstract
This paper introduces SoccerSum, a novel dataset aimed at enhancing object detection and segmentation in video frames depicting the soccer pitch, using footage from the Norwegian Eliteserien league across 2021-2023. With the goal of detecting elements beyond common entities in existing datasets, such as the soccer ball, players and referees, this dataset includes additional annotations for the goal net, corner flag posts, and the penalty mark. SoccerSum also includes the segmentation of key pitch areas such as the penalty and goal boxes for the same frame sequences. Comprising 750 frames annotated with 10 classes for advanced analysis, SoccerSum offers compatibility with existing frameworks, providing a rich dataset for the development of computer vision algorithms. This dataset not only serves as a resource for improving sports analytics, but also introduces a new application for automatic game summarization, enabling the generation of detailed and engaging content for fans and professionals. The SoccerSum dataset is accessible on Zenodo: https://zenodo.org/records/10612084.
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Thomas Torjusen, Pål Halvorsen
MMSys3
2023 SmartCrop: AI-Based Cropping of Soccer Videos
abstract
In the rapidly evolving landscape of digital platforms, the need for optimizing media representations to cater to various aspect ratios is palpable. In this paper, we pioneer an approach that utilizes object detection, scene detection, outlier detection, and interpolation for smart cropping. Using soccer as a case study, our primary goal is to capture the frame salience using object (player and ball) detection and tracking using AI models. To improve the object detection and tracking, we rely on scene understanding and explore various outlier detection and interpolation techniques. Our pipeline, called SmartCrop, is efficient, and supports various configurations for object tracking, interpolation, and outlier detection to find the best point-of-interest to be used as the cropping center of the video frame. An objective evaluation of the performance of individual pipeline components has validated our proposed architecture and the need for object, scene, outlier detection, and interpolation. Moreover, a crowdsourced subjective user study, assessing the alternative approaches for cropping from 16:9 to 1:1 and 9:16 aspect ratios, confirms that our proposed approach increases the end-user Quality of Experience (QoE).
Sayed Mohammad Majidi Dorcheh, Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Michael Riegler 0001, Dag Johansen, Pål Halvorsen
ISM3
2023 Soccer Athlete Data Visualization and Analysis with an Interactive Dashboard
Matthias Boeker, Cise Midoglu
MMM (1)2
2022 Experiences and Lessons Learned from a Crowdsourced-Remote Hybrid User Survey Framework
abstract
Subjective user studies are important to ensure the fidelity and usability of systems that generate multimedia content. Testing how end-users and domain experts perceive multimedia assets might provide crucial information. In this paper, we present our experiences with the open source hybrid crowdsourced-remote user survey framework called Huldra, which is intended for conducting web-based subjective user studies and aims to integrate the individual benefits associated with traditional, crowdsourced, and remote methods. We disseminate our experiences and insights from two actively deployed use cases and discuss challenges and opportunities associated with using Huldra as a framework for conducting user studies.
Cise Midoglu, Andrea M. Storås, Saeed Shafiee Sabet, Malek Hammou, Steven Alexander Hicks, Inga Strümke, Michael Riegler 0001, Carsten Griwodz, Pål Halvorsen
ISM1
2022 Huldra: a framework for collecting crowdsourced feedback on multimedia assets
abstract
Collecting crowdsourced feedback to evaluate, rank, or score multimedia content can be cumbersome and time-consuming. Most of the existing survey tools are complicated, hard to customize, or tailored for a specific asset type. In this paper, we present an open source framework called Huldra, designed explicitly to address the challenges associated with user studies involving crowdsourced feedback collection. The web-based framework is built in a modular and configurable fashion to allow for the easy adjustment of the user interface (UI) and the multimedia content, while providing integrations with reliable and stable backend solutions to facilitate the collection and analysis of responses. Our proposed framework can be used as an online survey tool by researchers working on different topics such as Machine Learning (ML), audio, image, and video quality assessment, Quality of Experience (QoE), and require user studies for the benchmarking of various types of multimedia content.
Malek Hammou, Cise Midoglu, Steven Alexander Hicks, Andrea M. Storås, Saeed Shafiee Sabet, Inga Strümke, Michael Riegler 0001, Pål Halvorsen
MMSys2
2022 Automatic thumbnail selection for soccer videos using machine learning
abstract
Thumbnail selection is a very important aspect of online sport video presentation, as thumbnails capture the essence of important events, engage viewers, and make video clips attractive to watch. Traditional solutions in the soccer domain for presenting highlight clips of important events such as goals, substitutions, and cards rely on the manual or static selection of thumbnails. However, such approaches can result in the selection of sub-optimal video frames as snapshots, which degrades the overall quality of the video clip as perceived by viewers, and consequently decreases viewership, not to mention that manual processes are expensive and time consuming. In this paper, we present an automatic thumbnail selection system for soccer videos which uses machine learning to deliver representative thumbnails with high relevance to video content and high visual quality in near real-time. Our proposed system combines a software framework which integrates logo detection, close-up shot detection, face detection, and image quality analysis into a modular and customizable pipeline, and a subjective evaluation framework for the evaluation of results. We evaluate our proposed pipeline quantitatively using various soccer datasets, in terms of complexity, runtime, and adherence to a pre-defined rule-set, as well as qualitatively through a user study, in terms of the perception of output thumbnails by end-users. Our results show that an automatic end-to-end system for the selection of thumbnails based on contextual relevance and visual quality can yield attractive highlight clips, and can be used in conjunction with existing soccer broadcast pipelines which require real-time operation.
Andreas Husa, Cise Midoglu, Malek Hammou, Steven Alexander Hicks, Dag Johansen, Tomas Kupka, Michael Riegler 0001, Pål Halvorsen
MMSys2
2022 HOST-ATS: automatic thumbnail selection with dashboard-controlled ML pipeline and dynamic user survey
abstract
We present HOST-ATS, a holistic system for the automatic selection and evaluation of soccer video thumbnails, which is composed of a dashboard-controlled machine learning (ML) pipeline, and a dynamic user survey. The ML pipeline uses logo detection, close-up shot detection, face detection, image quality prediction, and blur detection to automatically select thumbnails from soccer videos in near real-time, and can be configured via a graphical user interface. The web-based dynamic user survey can be employed to qualitatively evaluate the thumbnails selected by the pipeline. The survey is fully configurable and easy to update via continuous integration, allowing for the dynamic aggregation of participant responses to different sets of multimedia assets. We demonstrate the configuration and execution of the ML pipeline via the custom dashboard, and the agile (re-)deployment of the user survey via Firebase and Heroku cloud service integrations, where the audience can interact with configuration parameter updates in real-time. Our experience with HOST-ATS shows that an automatic thumbnail selection system can yield highly attractive highlight clips, and can be used in conjunction with existing soccer broadcast practices in real-time.
Andreas Husa, Cise Midoglu, Malek Hammou, Pål Halvorsen, Michael Riegler 0001
MMSys2
2022 Comparison of Crowdsourced and Remote Subjective User Studies: A Case Study of Investigative Child Interviews
abstract
Crowdsourced and remote user studies have recently gained popularity as alternatives to traditional laboratory studies. However, they are subject to unreliability, and it is challenging to ensure that valid results are collected, especially when conducting user studies with experts. Experts are a sparse resource, usually having busy schedules and heavy workloads, and are not necessarily geographically close. They are therefore often unwilling to participate in studies which require physical attendance. In this paper, we compare three alternative methods: crowd sourced user study with non-experts, remote user study with non-experts, and remote user study with domain experts, for a use case involving investigative child interview training. We present the results from three subjective studies about the perception of AI-generated child avatars, which is developed using various technologies such as dialogue models, game engine, text-to-speech and speech-to-text components. The study was conducted with three different user groups, and our results indicate the importance of using best practice measures for ensuring the collection of reliable results in crowdsourced settings as compared to remote studies, and highlight the difference between the perspectives of domain experts and non-experts.
Saeed Shafiee Sabet, Cise Midoglu, Syed Zohaib Hassan, Pegah Salehi, Gunn Astrid Baugerud, Carsten Griwodz, Miriam S. Johnson, Michael Riegler 0001, Pål Halvorsen
QoMEX2
2021 Automated Clipping of Soccer Events using Machine Learning
abstract
Extracting highlight clips from soccer matches requires tedious, time-consuming, and expensive manual labor. Human operators need to search for appropriate clipping points and trim away the unwanted scenes. In our work, we aim for an automated process for generating event highlights. In particular, we develop AI-models for scene boundary detection and logo detection. Using different datasets, we present two models that automatically find the appropriate time interval for goal event extraction. The models are evaluated quantitatively, and the results show that we find the logo and scene shifts with high accuracy. Our event clipping methodology is a potential building block for a larger, fully-automated sports broadcast production pipeline.
Joakim O. Valand, Haris Kadragic, Steven Alexander Hicks, Vajira Thambawita, Cise Midoglu, Tomas Kupka, Dag Johansen, Michael Riegler 0001, Pål Halvorsen
ISM5
2021 Reproducibility Companion Paper: Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial Features
abstract
Blind natural video quality assessment (BVQA), also known as no-reference video quality assessment, is a highly active research topic. In our recent contribution titled "Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial Features" published in ACM Multimedia 2020, we proposed a two-level video quality model employing statistical temporal features and spatial features extracted by a deep convolutional neural network (CNN) for this purpose. At the time of publishing, the proposed model (CNN-TLVQM) achieved state-of-the-art results in BVQA. In this paper, we describe the process of reproducing the published results by using CNN-TLVQM on two publicly available natural video quality datasets.
Jari Korhonen, Yicheng Su, Junyong You, Steven Alexander Hicks, Cise Midoglu
ACM Multimedia5
2021 Reproducibility Companion Paper: Campus3D: A Photogrammetry Point Cloud Benchmark for Outdoor Scene Hierarchical Understanding
abstract
This companion paper is to support the replication of paper "Campus3D: A Photogrammetry Point Cloud Benchmark for Outdoor Scene Hierarchical Understanding", which was presented at ACM Multimedia 2020. The supported paper's main purpose was to provide a photogrammetry point cloud-based dataset with hierarchical multilabels to facilitate the area of 3D deep learning. Based on this provided dataset and source code, in this work, we build a complete package to reimplement the proposed methods and experiments (i.e., the hierarchical learning framework and the benchmarks of the hierarchical semantic segmentation task). Specifically, this paper contains the technical details of the package, including file structure, dataset preparation, installation package, and the conduction of the experiment. We also present the replicated experiment results and indicate our contributions to the original implementation.
Yuqing Liao, Zekun Tong, Yabang Zhao, Andrew Lim 0001, Zhenzhong Kuang, Cise Midoglu
ACM Multimedia7
2021 Large scale "speedtest" experimentation in Mobile Broadband Networks
Cise Midoglu, Konstantinos Kousias, Özgü Alay, Andra Lutu, Antonios Argyriou, Michael Riegler 0001, Carsten Griwodz
Comput. Networks1
2019 Docker-Based Evaluation Framework for Video Streaming QoE in Broadband Networks
abstract
Video streaming is one of the top traffic contributors in the Internet and a frequent research subject. It is expected that streaming traffic will grow 4-fold for video globally and 9-fold for mobile video between 2017 and 2022. In this paper, we present an automatized measurement framework for evaluating video streaming QoE in operational broadband networks, using headless streaming with a Docker-based client, and a server-side implementation allowing for the use of multiple video players and adaptation algorithms. Our framework allows for integration with the acsMONROE testbed and Bitmovin Analytics, which bring on the possibility to conduct large-scale measurements in different networks, including mobility scenarios, and monitor different parameters in the application, transport, network, and physical layers in real-time.
Cise Midoglu, Anatoliy Zabrovskiy, Özgü Alay, Daniel Hoelbling-Inzko, Carsten Griwodz, Christian Timmerer
ACM Multimedia1
2019 Results from running an experiment as a service platform for mobile broadband networks in Europe
Vincenzo Mancuso, Miguel Peón-Quirós, Cise Midoglu, Mohamed Moulay, Vincenzo Comite, Andra Lutu, Özgü Alay, Stefan Alfredsson, Mohammad Rajiullah, Anna Brunström, Marco Mellia, Ali Safari Khatouni, Thomas Hirsch
Comput. Commun.3
2018 Open video datasets over operational mobile networks with MONROE
abstract
Video streaming is a very popular service among the end-users of Mobile Broadband (MBB) networks. DASH and WebRTC are two key technologies in the delivery of mobile video. In this work, we empirically assess the performance of video streaming with DASH and WebRTC in operational MBB networks, by using a large number of programmable network probes spread over several countries in the context of the MONROE project. We collect a large dataset from more than 300 video streaming experiments. Our dataset consists of network traces, performance indicators captured during the streaming sessions, and experiment metadata. The dataset captures the wide variability in video streaming performance, and unveils how mobile broadband is still not offering consistent quality guarantees across different countries and networks, especially for users on the move. We open source our complete software toolset and provide the video dataset as open data.
Cise Midoglu, Mohamed Moulay, Vincenzo Mancuso, Özgü Alay, Andra Lutu, Carsten Griwodz
MMSys1
2017 The same, only different: Contrasting mobile operator behavior from crowdsourced dataset
abstract
Crowdsourcing mobile network performance evaluation is rapidly gaining popularity, with new applications aiming to deliver more accurate and reliable results every day. From the perspective of end-users, these utilities help them estimate the performance of their service provider in terms of throughput, latency and other key performance indicators of the network. In this paper, we build ORCA: Operator Classifier, a Machine Learning (ML) based framework to define and determine the behavior of Mobile Network Operators (MNOs) from crowdsourced datasets. We investigate whether one can differentiate MNOs by using crowdsourced end-to-end network measurements. We consider different performance metrics (e.g. Download (DL)/Upload (UL) data rate, latency, signal strength) and study the impact of them individually but also collectively on differentiating MNOs. We use RTR Open Data, an open dataset of broadband measurements provided by the Austrian Regulatory Authority for Broadcasting and Telecommunications (RTR), to characterize the three major mobile native operators and two virtual operators in Austria. Our results show that ORCA can be used to identify patterns between various mobile systems and disclose their differences from the end-user perspective.
Konstantinos Kousias, Cise Midoglu, Özgü Alay, Andra Lutu, Antonios Argyriou, Michael Riegler 0001
PIMRC2
2016 Server link load modeling and request scheduling for crowdsourcing-based benchmarking systems
abstract
Crowdsourcing-based network benchmarking is a novel method of large-scale performance evaluation, which allows for the collection of measurements from end nodes in order to estimate overall network performance. For accuracy and fairness, the measurement server and its communication links must not constitute the bottleneck in the overall benchmarking system. In this paper, we focus on performance optimization through efficient request scheduling. We present a modular approach to model the load on the server-side communication links analytically, allowing for the combination of simulation parameters with real measurement results, and demonstrate the potential benefit of scheduling measurement requests with a discrete time-based approach. We show that efficient scheduling algorithms are as important as hardware improvements (such as upgrading the communication links to higher capacity), in order for benchmarking systems to cope with the increasing number of measurements and ever growing data rates in both fixed and mobile technologies.
Cise Midoglu, Leonhard Wimmer, Philipp Svoboda
IWCMC1