EDBT 2026 Demo / reviewers in the wild / expert
Mehdi Houshmand Sarkhoosh
dblp:359/3397
· DBLP profile ↗
16ranked-venue papers
10as first author
16since 2021 · last 2026
0009-0008-4616-4592ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 10 first-author · 16 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SportSBD: Shot Boundary Detection in Sports FootageabstractShot Boundary Detection (SBD), which identifies scene (or “shot”) changes, is a core step in video analysis pipelines such as summarization and highlight generation. Yet, it remains challenging in sports broadcasts because rapid camera motion, frequent camera switches, sport-specific transitions and graphic overlays often cause false detections and poor cross-domain generalization. In this paper, we address shot boundary detection in professional sports broadcasts, focusing on ice hockey and soccer. We propose a sports-oriented model based on a fine-tuned R(2+1)D 3D CNN, trained to detect hard cuts, gradual transitions, and logo-based replay effects. The model is evaluated on both goal-centered clips and full-match broadcast footage, and is benchmarked against two state-of-the-art baselines: TransNetV2 for accuracy and PySceneDetect for efficiency. Our approach consistently achieves higher precision, recall, and F1-score across all evaluation settings, while demonstrating strong cross-league generalization within ice hockey and cross-sport generalization to soccer. We release our pretrained sports-specific SBD model as an open-source Python package, enabling straightforward integration into existing video analysis pipelines. Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Pål Halvorsen |
MMSys | 1 |
| 2025 | Hockey2D: A Keypoint-Based Framework for Ice Hockey Rink Localization and Object MappingabstractAccurate localization of players and objects in ice hockey is essential for advanced analytics, tactical analysis, and automated content generation. Traditional homography estimation approaches rely heavily on explicit camera calibration or heuristic methods, which are computationally intensive and sensitive to camera variations. This paper introduces Hockey2D, a robust framework leveraging YOLO-based pose estimation to identify rink keypoints, combined with RANSAC-based homography optimization for precise spatial mapping. To handle broadcast variations and occlusions effectively, a shot-type classifier filters input frames, selectively applying homography estimation to suitable scenes. By integrating keypoint and object detection, our approach achieves accurate 2D localization of players, referees, and goalkeepers. Evaluations conducted on ice hockey-specific datasets confirm both computational efficiency and localization accuracy. The proposed method highlights the practical viability of AI-driven localization for interactive storytelling, game summarization, and tactical coaching. A video demonstration is available at https://youtu.be/JCnX4N4fi8I. Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen |
CBMI | 1 |
| 2025 | VoiceVision: AI-Powered Speaker-Aware Cropping and Content Indexing for Multi-Speaker VideosabstractVoiceVision is an AI-powered system designed for intelligent speaker-focused video cropping and speaker-aware content indexing. Built on top of the TalkNet audio-visual speaker diarization backbone, Voice Vision detects active speakers in multi-speaker videos and dynamically crops and reframes the video to center on the current speaker, creating smooth visual transitions. In addition to smart cropping, the system integrates automatic speech recognition (ASR) using Whisper to generate accurate transcriptions, which are further processed through a transcript attribution module to associate spoken segments with specific speakers. A dedicated speech search module enables efficient retrieval and indexing of content based on keywords or speaker identity. Voice Vision supports automatic aspect ratio adaptation (9:16, 1:1, 4:5) to generate social-media- optimized outputs. By combining speaker-aware video cropping with searchable, speaker-attributed transcriptions, Voice Vision simplifies content creation, indexing, and sharing for interview- style or conversational videos on platforms such as TikTok, Instagram, and YouTube Shorts. A demonstration of the system's capabilities is presented at https://youtu.be/SBSqOyMpe60 Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen |
CBMI | 1 |
| 2025 | Extracting Player Speed from Football VideosabstractOne of the key metrics for player and team performance in football is speed. Traditionally, measuring player speed either requires extensive human effort or expensive equipment. In this work, we investigate methods for manual and automatic player speed extraction directly from broadcast football videos. We implement a pipeline combining player detection, tracking, field mapping, and position measurement. We experiment with different configurations of the pipeline, concluding that a setting with YOLOv11 for detection, StrongSORT for tracking, PnL-Calib for field mapping and keypoint detection, and position measurement through homography yields the best results for our current scope and test datasets. This work serves as a proof-of-concept analytics solution for player speed extraction, which can be adopted by teams of all sizes and means, and built upon by researchers. Our pipeline implementation and two newly curated player speed datasets (Begnadalen and TACDEC++) are openly available for the sports and scientific community. Ole Kristian Rustebakke, Mehdi Houshmand Sarkhoosh, Cise Midoglu, Pål Halvorsen |
ISM | 2 |
| 2025 | ExposureEngine: Oriented Logo Detection and Sponsor Visibility Analytics in Sports BroadcastsabstractQuantifying sponsor visibility in sports broadcasts is a critical marketing task traditionally hindered by manual, subjective, and unscalable analysis methods. While automated systems offer an alternative, their reliance on axis-aligned Horizontal Bounding Box (HBB) leads to inaccurate exposure metrics when logos appear rotated or skewed due to dynamic camera angles and perspective distortions. This paper introduces ExposureEngine, an end-to-end system designed for accurate, rotation-aware sponsor visibility analytics in sports broadcasts, demonstrated in a soccer case study. Our approach predicts Oriented Bounding Box (OBB) to provide a geometrically precise fit to each logo regardless of the orientation on-screen. To train and evaluate our detector, we developed a new dataset comprising 1,103 frames from Swedish elite soccer, featuring 670 unique sponsor logos annotated with OBBs. Our model achieves a mean Average Precision (mAP @ 0.5) of 0.859, with a precision of 0.96 and recall of 0.87, demonstrating robust performance in localizing logos under diverse broadcast conditions. The system integrates these detections into an analytical pipeline that calculates precise visibility metrics, such as exposure duration and on-screen coverage. Furthermore, we incorporate a language-driven agentic layer, enabling users to generate reports, summaries, and media content through natural language queries. The complete system, including the dataset and the analytics dashboard, provides a comprehensive solution for auditable and interpretable sponsor measurement in sports media. An overview of the ExposureEngine is available online. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>1https://youtu.be/tRw6OBISuW4 Mehdi Houshmand Sarkhoosh, Frøy Øye, Henrik Nestor Sørlie, Nam Hoang Vu, Dag Johansen, Cise Midoglu, Tomas Kupka, Pål Halvorsen |
ISM | 1 |
| 2025 | HockeyAI: A Multi-Class Ice Hockey Dataset for Object DetectionabstractThe fast paced nature of ice hockey presents unique challenges for object detection, particularly in tracking the puck---a small, fast moving object that is critical to gameplay analysis. This paper introduces HockeyAI, a novel open source dataset specifically designed for multi-class object detection in ice hockey. The dataset includes 2,101 high resolution frames extracted from professional games in the Swedish Hockey League (SHL), annotated in the You Look Only Once (YOLO) format. Annotations span 7 classes, covering dynamic objects such as the players and the puck, as well as static rink elements such as goalposts and face-off circles. The dataset is derived from diverse SHL games across multiple seasons and teams, ensuring a rich variety of scenarios reflective of real world gameplay. A fine tuned YOLOv8 medium model is also provided, demonstrating high performance across all classes. Key comparisons highlight the dataset's advancements over existing resources, addressing their limitations such as incomplete class coverage, low resolution, and inconsistent annotations. The dataset, model, and an interactive demo are publicly available on Hugging Face under an open source license, fostering further collaboration in sports related computer vision applications (https://huggingface.co/SimulaMet-HOST/HockeyAI). Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen |
MMSys | 1 |
| 2025 | HockeyRink: A Dataset for Precise Ice Hockey Rink Keypoint Mapping and AnalyticsabstractPrecise mapping of ice hockey rinks is critical for applications such as player tracking, game strategy analysis, and broadcast enhancements. Traditional methods often rely on manual annotations or simplistic models that fail to account for the rink's complex geometry and dynamic in-game conditions. To address these limitations, we present HockeyRink, a novel dataset comprising 56 meticulously annotated keypoints corresponding to significant landmarks on a standard hockey rink, including face-off dots, goalposts, and blue lines. Leveraging the YOLOv8-Large pose estimation architecture, we adapted the model to treat the rink as a single 'pose' object, enabling accurate keypoint predictions tailored to the nuances of hockey rink imagery. Our dataset, derived from diverse hockey game footage from the Swedish Hockey League (SHL), facilitates applications in homography estimation, 2D/3D scene mapping, and tactical overlays. By detecting and mapping rink keypoints, one can compute precise transformations between the image plane and the rink's physical dimensions, enabling the overlay of player trajectories and tactical insights onto broadcast footage. This work addresses the unique challenges of ice hockey, such as occlusions, rapid camera movements, and varying lighting conditions, and offers a robust foundation for future research and innovation in sports analytics. The HockeyRink dataset and trained model are openly available at: https://huggingface.co/SimulaMet-HOST/HockeyRink. Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen |
MMSys | 1 |
| 2025 | HockeyOrient: A Dataset for Ice Hockey Player Orientation ClassificationabstractUnderstanding player orientation is a critical component of sports analytics, offering insights into gameplay strategies and player behavior. This paper presents HockeyOrient, a novel dataset for classifying the orientation of ice hockey players based on their poses. The dataset comprises 9,700 manually annotated frames, selected randomly and non-sequentially, taken from Swedish Hockey League (SHL) games during the 2023 and 2024 seasons. Each player image is cropped from game footage and categorized into one of eight orientation classes: top, top-right, right, bottom-right, bottom, bottom-left, left, and top-left. The dataset includes diverse scenarios, such as different teams, jersey colors, referees, and goaltenders with unique protective gear. Alongside the dataset, we provide an open-source classification model trained on the dataset using the SqueezeNet architecture, achieving an F1 score of 75% across all the classes. This work addresses a significant gap in ice hockey analytics, enabling advanced player tracking and gameplay analysis. The dataset and the trained model can be accessed publicly on Hugging Face under an open-source license (https://huggingface.co/datasets/SimulaMet-HOST/HockeyOrient). Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen |
MMSys | 1 |
| 2024 | Demo: Creating Player-Specific Soccer Highlight Clips with PlayerTVabstractThis paper demonstrates PlayerTV, an innovative framework which harnesses state-of-the-art Artificial Intelligence (AI) technologies for automatic player tracking and identification in soccer videos. By integrating object detection and tracking, Optical Character Recognition (OCR), and color analysis, PlayerTV facilitates the generation of player-specific highlight clips from extensive game footage, significantly reducing the manual labor traditionally associated with such tasks. We present how PlayerTV can be run standalone as a core pipeline, as well as through an interactive Graphical User Interface (GUI). Håkon Maric Solberg, Mehdi Houshmand Sarkhoosh, Sushant Gautam, Saeed Shafiee Sabet, Pål Halvorsen, Cise Midoglu |
CBMI | 2 |
| 2024 | SoccerNet-Echoes: A Soccer Game Audio Commentary DatasetabstractThe application of Automatic Speech Recognition (ASR) technology in soccer enables sports analytics by extracting audio commentaries to provide insights into game events and facilitate automatic game understanding. This paper presents SoccerNet-Echoes, an extension of the SoccerNet dataset with automatically generated transcriptions of soccer game broadcasts. Generated using the Whisper model and translated with Google Translate into English when needed, these transcriptions enhance video content with textual information derived from game audio. SoccerNet-Echoes serves as a comprehensive resource for developing algorithms in action spotting, caption generation, and game summarization. Through a series of experiments, we demonstrate that combining modalities—audio, video, and text—yields mixed results on classification tasks. The combination of audio and video shows improved performance over individual modalities, while the addition of ASR text does not significantly enhance results. Additionally, our baseline summarization tasks indicate that ASR content enriches summaries, offering insights beyond event information. This multimodal dataset supports diverse applications, broadening the scope of research in sports analytics. The dataset is available at: https://github.com/SoccerNet/sn-echoes. Sushant Gautam, Mehdi Houshmand Sarkhoosh, Jan Held, Cise Midoglu, Anthony Cioppa, Silvio Giancola, Vajira Thambawita, Michael Riegler 0001, Pål Halvorsen, Mubarak Shah |
ISM | 2 |
| 2024 | PlayerTV: Advanced Player Tracking and Identification for Automatic Soccer Highlight ClipsabstractIn the rapidly evolving field of sports analytics, the automation of targeted video processing is a pivotal advancement. We propose PlayerTV, an innovative framework which harnesses state-of-the-art AI technologies for automatic player tracking and identification in soccer videos. By integrating object detection and tracking, Optical Character Recognition (OCR), and color analysis, PlayerTV facilitates the generation of player-specific highlight clips from extensive game footage, significantly reducing the manual labor traditionally associated with such tasks. Preliminary results from the evaluation of our core pipeline, tested on a dataset from the Norwegian Eliteserien league, indicate that PlayerTV can accurately and efficiently identify teams and players, and our interactive Graphical User Interface (GUI) serves as a user-friendly application wrapping this functionality for streamlined use. Håkon Maric Solberg, Mehdi Houshmand Sarkhoosh, Sushant Gautam, Saeed Shafiee Sabet, Pål Halvorsen, Cise Midoglu |
ISM | 2 |
| 2024 | AI-Based Cropping of Soccer Videos for Different Social Media Representations
Mehdi Houshmand Sarkhoosh, Sayed Mohammad Majidi Dorcheh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Michael Riegler 0001, Pål Halvorsen |
MMM (4) | 1 |
| 2024 | SmartCrop-H: AI-Based Cropping of Ice Hockey VideosabstractSports multimedia plays a central role in captivating audiences on social media platforms. However, fast-paced sports such as ice hockey pose unique challenges due to their swift gameplay and the small puck size, making object tracking-based video adaptation for social media a complex task. In this context, we introduce SmartCrop-H, an innovative ice hockey video cropping tool powered by advanced AI models. It excels at tracking the puck and ensuring that crucial gameplay remains the center of attention, regardless of the desired target aspect ratio. The tool combines various techniques including object detection, scene detection, outlier detection, and smoothing, to deliver high-quality ratio-adapted videos. In this demonstration, we showcase SmartCrop-H in real-world scenarios through an intuitive step-by-step Graphical User Interface (GUI) that vividly illustrates how the tool works. The demonstration emphasizes the vital role of AI in enhancing the sports viewing experience, and its importance in the dynamic realm of social media content distribution. A video of the demo can be found here: https://youtu.be/rMmYOCM-k7A. Mohammad Majidi, Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Pål Halvorsen |
MMSys | 2 |
| 2024 | Multimodal AI-Based Summarization and Storytelling for Soccer on Social MediaabstractThe rapid advancement of technology has been revolutionizing the field of sports media, where there is a growing need for sophisticated data processing methods. Current methodologies for extracting information from soccer broadcast videos to generate game highlights and summaries for social media are predominantly manual and rely heavily on text-based NLP techniques, overlooking the rich visual and auditory information available. In response to this challenge, our research introduces SoccerSum, a tool that innovates in the field by integrating computer vision, audio analysis with advanced language models like GPT-4. This multimodal approach enables automated, enriched content summarization, including detection of players and key field elements, thereby enhancing the metadata used in summarization algorithms. SoccerSum uniquely combines textual and visual data, offering a comprehensive solution for generating accurate, platform-specific content. This development represents a significant advancement in automated, data-driven sports media dissemination, and sets a new benchmark in the realm of soccer information extraction. A video of the demo can be found here: https://youtu.be/za4VIi2ARXY. Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Pål Halvorsen |
MMSys | 1 |
| 2024 | The SoccerSum Dataset for Automated Detection, Segmentation, and Tracking of Objects on the Soccer PitchabstractThis paper introduces SoccerSum, a novel dataset aimed at enhancing object detection and segmentation in video frames depicting the soccer pitch, using footage from the Norwegian Eliteserien league across 2021-2023. With the goal of detecting elements beyond common entities in existing datasets, such as the soccer ball, players and referees, this dataset includes additional annotations for the goal net, corner flag posts, and the penalty mark. SoccerSum also includes the segmentation of key pitch areas such as the penalty and goal boxes for the same frame sequences. Comprising 750 frames annotated with 10 classes for advanced analysis, SoccerSum offers compatibility with existing frameworks, providing a rich dataset for the development of computer vision algorithms. This dataset not only serves as a resource for improving sports analytics, but also introduces a new application for automatic game summarization, enabling the generation of detailed and engaging content for fans and professionals. The SoccerSum dataset is accessible on Zenodo: https://zenodo.org/records/10612084. Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Thomas Torjusen, Pål Halvorsen |
MMSys | 1 |
| 2023 | SmartCrop: AI-Based Cropping of Soccer VideosabstractIn the rapidly evolving landscape of digital platforms, the need for optimizing media representations to cater to various aspect ratios is palpable. In this paper, we pioneer an approach that utilizes object detection, scene detection, outlier detection, and interpolation for smart cropping. Using soccer as a case study, our primary goal is to capture the frame salience using object (player and ball) detection and tracking using AI models. To improve the object detection and tracking, we rely on scene understanding and explore various outlier detection and interpolation techniques. Our pipeline, called SmartCrop, is efficient, and supports various configurations for object tracking, interpolation, and outlier detection to find the best point-of-interest to be used as the cropping center of the video frame. An objective evaluation of the performance of individual pipeline components has validated our proposed architecture and the need for object, scene, outlier detection, and interpolation. Moreover, a crowdsourced subjective user study, assessing the alternative approaches for cropping from 16:9 to 1:1 and 9:16 aspect ratios, confirms that our proposed approach increases the end-user Quality of Experience (QoE). Sayed Mohammad Majidi Dorcheh, Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Michael Riegler 0001, Dag Johansen, Pål Halvorsen |
ISM | 2 |