Tomas Kupka

dblp:71/4144 · DBLP profile ↗
← Back
22ranked-venue papers
2as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 12 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 3 · 2 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 SportSBD: Shot Boundary Detection in Sports Footage
abstract
Shot Boundary Detection (SBD), which identifies scene (or “shot”) changes, is a core step in video analysis pipelines such as summarization and highlight generation. Yet, it remains challenging in sports broadcasts because rapid camera motion, frequent camera switches, sport-specific transitions and graphic overlays often cause false detections and poor cross-domain generalization. In this paper, we address shot boundary detection in professional sports broadcasts, focusing on ice hockey and soccer. We propose a sports-oriented model based on a fine-tuned R(2+1)D 3D CNN, trained to detect hard cuts, gradual transitions, and logo-based replay effects. The model is evaluated on both goal-centered clips and full-match broadcast footage, and is benchmarked against two state-of-the-art baselines: TransNetV2 for accuracy and PySceneDetect for efficiency. Our approach consistently achieves higher precision, recall, and F1-score across all evaluation settings, while demonstrating strong cross-league generalization within ice hockey and cross-sport generalization to soccer. We release our pretrained sports-specific SBD model as an open-source Python package, enabling straightforward integration into existing video analysis pipelines.
Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Pål Halvorsen
MMSys4
2025 Hockey2D: A Keypoint-Based Framework for Ice Hockey Rink Localization and Object Mapping
abstract
Accurate localization of players and objects in ice hockey is essential for advanced analytics, tactical analysis, and automated content generation. Traditional homography estimation approaches rely heavily on explicit camera calibration or heuristic methods, which are computationally intensive and sensitive to camera variations. This paper introduces Hockey2D, a robust framework leveraging YOLO-based pose estimation to identify rink keypoints, combined with RANSAC-based homography optimization for precise spatial mapping. To handle broadcast variations and occlusions effectively, a shot-type classifier filters input frames, selectively applying homography estimation to suitable scenes. By integrating keypoint and object detection, our approach achieves accurate 2D localization of players, referees, and goalkeepers. Evaluations conducted on ice hockey-specific datasets confirm both computational efficiency and localization accuracy. The proposed method highlights the practical viability of AI-driven localization for interactive storytelling, game summarization, and tactical coaching. A video demonstration is available at https://youtu.be/JCnX4N4fi8I.
Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
CBMI4
2025 VoiceVision: AI-Powered Speaker-Aware Cropping and Content Indexing for Multi-Speaker Videos
abstract
VoiceVision is an AI-powered system designed for intelligent speaker-focused video cropping and speaker-aware content indexing. Built on top of the TalkNet audio-visual speaker diarization backbone, Voice Vision detects active speakers in multi-speaker videos and dynamically crops and reframes the video to center on the current speaker, creating smooth visual transitions. In addition to smart cropping, the system integrates automatic speech recognition (ASR) using Whisper to generate accurate transcriptions, which are further processed through a transcript attribution module to associate spoken segments with specific speakers. A dedicated speech search module enables efficient retrieval and indexing of content based on keywords or speaker identity. Voice Vision supports automatic aspect ratio adaptation (9:16, 1:1, 4:5) to generate social-media- optimized outputs. By combining speaker-aware video cropping with searchable, speaker-attributed transcriptions, Voice Vision simplifies content creation, indexing, and sharing for interview- style or conversational videos on platforms such as TikTok, Instagram, and YouTube Shorts. A demonstration of the system's capabilities is presented at https://youtu.be/SBSqOyMpe60
Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
CBMI4
2025 ExposureEngine: Oriented Logo Detection and Sponsor Visibility Analytics in Sports Broadcasts
abstract
Quantifying sponsor visibility in sports broadcasts is a critical marketing task traditionally hindered by manual, subjective, and unscalable analysis methods. While automated systems offer an alternative, their reliance on axis-aligned Horizontal Bounding Box (HBB) leads to inaccurate exposure metrics when logos appear rotated or skewed due to dynamic camera angles and perspective distortions. This paper introduces ExposureEngine, an end-to-end system designed for accurate, rotation-aware sponsor visibility analytics in sports broadcasts, demonstrated in a soccer case study. Our approach predicts Oriented Bounding Box (OBB) to provide a geometrically precise fit to each logo regardless of the orientation on-screen. To train and evaluate our detector, we developed a new dataset comprising 1,103 frames from Swedish elite soccer, featuring 670 unique sponsor logos annotated with OBBs. Our model achieves a mean Average Precision (mAP @ 0.5) of 0.859, with a precision of 0.96 and recall of 0.87, demonstrating robust performance in localizing logos under diverse broadcast conditions. The system integrates these detections into an analytical pipeline that calculates precise visibility metrics, such as exposure duration and on-screen coverage. Furthermore, we incorporate a language-driven agentic layer, enabling users to generate reports, summaries, and media content through natural language queries. The complete system, including the dataset and the analytics dashboard, provides a comprehensive solution for auditable and interpretable sponsor measurement in sports media. An overview of the ExposureEngine is available online. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>1https://youtu.be/tRw6OBISuW4
Mehdi Houshmand Sarkhoosh, Frøy Øye, Henrik Nestor Sørlie, Nam Hoang Vu, Dag Johansen, Cise Midoglu, Tomas Kupka, Pål Halvorsen
ISM7
2025 HockeyAI: A Multi-Class Ice Hockey Dataset for Object Detection
abstract
The fast paced nature of ice hockey presents unique challenges for object detection, particularly in tracking the puck---a small, fast moving object that is critical to gameplay analysis. This paper introduces HockeyAI, a novel open source dataset specifically designed for multi-class object detection in ice hockey. The dataset includes 2,101 high resolution frames extracted from professional games in the Swedish Hockey League (SHL), annotated in the You Look Only Once (YOLO) format. Annotations span 7 classes, covering dynamic objects such as the players and the puck, as well as static rink elements such as goalposts and face-off circles. The dataset is derived from diverse SHL games across multiple seasons and teams, ensuring a rich variety of scenarios reflective of real world gameplay. A fine tuned YOLOv8 medium model is also provided, demonstrating high performance across all classes. Key comparisons highlight the dataset's advancements over existing resources, addressing their limitations such as incomplete class coverage, low resolution, and inconsistent annotations. The dataset, model, and an interactive demo are publicly available on Hugging Face under an open source license, fostering further collaboration in sports related computer vision applications (https://huggingface.co/SimulaMet-HOST/HockeyAI).
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
MMSys5
2025 HockeyRink: A Dataset for Precise Ice Hockey Rink Keypoint Mapping and Analytics
abstract
Precise mapping of ice hockey rinks is critical for applications such as player tracking, game strategy analysis, and broadcast enhancements. Traditional methods often rely on manual annotations or simplistic models that fail to account for the rink's complex geometry and dynamic in-game conditions. To address these limitations, we present HockeyRink, a novel dataset comprising 56 meticulously annotated keypoints corresponding to significant landmarks on a standard hockey rink, including face-off dots, goalposts, and blue lines. Leveraging the YOLOv8-Large pose estimation architecture, we adapted the model to treat the rink as a single 'pose' object, enabling accurate keypoint predictions tailored to the nuances of hockey rink imagery. Our dataset, derived from diverse hockey game footage from the Swedish Hockey League (SHL), facilitates applications in homography estimation, 2D/3D scene mapping, and tactical overlays. By detecting and mapping rink keypoints, one can compute precise transformations between the image plane and the rink's physical dimensions, enabling the overlay of player trajectories and tactical insights onto broadcast footage. This work addresses the unique challenges of ice hockey, such as occlusions, rapid camera movements, and varying lighting conditions, and offers a robust foundation for future research and innovation in sports analytics. The HockeyRink dataset and trained model are openly available at: https://huggingface.co/SimulaMet-HOST/HockeyRink.
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
MMSys5
2025 HockeyOrient: A Dataset for Ice Hockey Player Orientation Classification
abstract
Understanding player orientation is a critical component of sports analytics, offering insights into gameplay strategies and player behavior. This paper presents HockeyOrient, a novel dataset for classifying the orientation of ice hockey players based on their poses. The dataset comprises 9,700 manually annotated frames, selected randomly and non-sequentially, taken from Swedish Hockey League (SHL) games during the 2023 and 2024 seasons. Each player image is cropped from game footage and categorized into one of eight orientation classes: top, top-right, right, bottom-right, bottom, bottom-left, left, and top-left. The dataset includes diverse scenarios, such as different teams, jersey colors, referees, and goaltenders with unique protective gear. Alongside the dataset, we provide an open-source classification model trained on the dataset using the SqueezeNet architecture, achieving an F1 score of 75% across all the classes. This work addresses a significant gap in ice hockey analytics, enabling advanced player tracking and gameplay analysis. The dataset and the trained model can be accessed publicly on Hugging Face under an open-source license (https://huggingface.co/datasets/SimulaMet-HOST/HockeyOrient).
Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen
MMSys5
2024 AI-Based Cropping of Soccer Videos for Different Social Media Representations
Mehdi Houshmand Sarkhoosh, Sayed Mohammad Majidi Dorcheh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Michael Riegler 0001, Pål Halvorsen
MMM (4)5
2024 SmartCrop-H: AI-Based Cropping of Ice Hockey Videos
abstract
Sports multimedia plays a central role in captivating audiences on social media platforms. However, fast-paced sports such as ice hockey pose unique challenges due to their swift gameplay and the small puck size, making object tracking-based video adaptation for social media a complex task. In this context, we introduce SmartCrop-H, an innovative ice hockey video cropping tool powered by advanced AI models. It excels at tracking the puck and ensuring that crucial gameplay remains the center of attention, regardless of the desired target aspect ratio. The tool combines various techniques including object detection, scene detection, outlier detection, and smoothing, to deliver high-quality ratio-adapted videos. In this demonstration, we showcase SmartCrop-H in real-world scenarios through an intuitive step-by-step Graphical User Interface (GUI) that vividly illustrates how the tool works. The demonstration emphasizes the vital role of AI in enhancing the sports viewing experience, and its importance in the dynamic realm of social media content distribution. A video of the demo can be found here: https://youtu.be/rMmYOCM-k7A.
Mohammad Majidi, Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Dag Johansen, Pål Halvorsen
MMSys5
2023 SmartCrop: AI-Based Cropping of Soccer Videos
abstract
In the rapidly evolving landscape of digital platforms, the need for optimizing media representations to cater to various aspect ratios is palpable. In this paper, we pioneer an approach that utilizes object detection, scene detection, outlier detection, and interpolation for smart cropping. Using soccer as a case study, our primary goal is to capture the frame salience using object (player and ball) detection and tracking using AI models. To improve the object detection and tracking, we rely on scene understanding and explore various outlier detection and interpolation techniques. Our pipeline, called SmartCrop, is efficient, and supports various configurations for object tracking, interpolation, and outlier detection to find the best point-of-interest to be used as the cropping center of the video frame. An objective evaluation of the performance of individual pipeline components has validated our proposed architecture and the need for object, scene, outlier detection, and interpolation. Moreover, a crowdsourced subjective user study, assessing the alternative approaches for cropping from 16:9 to 1:1 and 9:16 aspect ratios, confirms that our proposed approach increases the end-user Quality of Experience (QoE).
Sayed Mohammad Majidi Dorcheh, Mehdi Houshmand Sarkhoosh, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Michael Riegler 0001, Dag Johansen, Pål Halvorsen
ISM5
2022 Automatic thumbnail selection for soccer videos using machine learning
abstract
Thumbnail selection is a very important aspect of online sport video presentation, as thumbnails capture the essence of important events, engage viewers, and make video clips attractive to watch. Traditional solutions in the soccer domain for presenting highlight clips of important events such as goals, substitutions, and cards rely on the manual or static selection of thumbnails. However, such approaches can result in the selection of sub-optimal video frames as snapshots, which degrades the overall quality of the video clip as perceived by viewers, and consequently decreases viewership, not to mention that manual processes are expensive and time consuming. In this paper, we present an automatic thumbnail selection system for soccer videos which uses machine learning to deliver representative thumbnails with high relevance to video content and high visual quality in near real-time. Our proposed system combines a software framework which integrates logo detection, close-up shot detection, face detection, and image quality analysis into a modular and customizable pipeline, and a subjective evaluation framework for the evaluation of results. We evaluate our proposed pipeline quantitatively using various soccer datasets, in terms of complexity, runtime, and adherence to a pre-defined rule-set, as well as qualitatively through a user study, in terms of the perception of output thumbnails by end-users. Our results show that an automatic end-to-end system for the selection of thumbnails based on contextual relevance and visual quality can yield attractive highlight clips, and can be used in conjunction with existing soccer broadcast pipelines which require real-time operation.
Andreas Husa, Cise Midoglu, Malek Hammou, Steven Alexander Hicks, Dag Johansen, Tomas Kupka, Michael Riegler 0001, Pål Halvorsen
MMSys6
2021 Automated Clipping of Soccer Events using Machine Learning
abstract
Extracting highlight clips from soccer matches requires tedious, time-consuming, and expensive manual labor. Human operators need to search for appropriate clipping points and trim away the unwanted scenes. In our work, we aim for an automated process for generating event highlights. In particular, we develop AI-models for scene boundary detection and logo detection. Using different datasets, we present two models that automatically find the appropriate time interval for goal event extraction. The models are evaluated quantitatively, and the results show that we find the logo and scene shifts with high accuracy. Our event clipping methodology is a potential building block for a larger, fully-automated sports broadcast production pipeline.
Joakim O. Valand, Haris Kadragic, Steven Alexander Hicks, Vajira Thambawita, Cise Midoglu, Tomas Kupka, Dag Johansen, Michael Riegler 0001, Pål Halvorsen
ISM6
2020 PMData: a sports logging dataset
abstract
In this paper, we present PMData: a dataset that combines traditional lifelogging data with sports-activity data. Our dataset enables the development of novel data analysis and machine-learning applications where, for instance, additional sports data is used to predict and analyze everyday developments, like a person's weight and sleep patterns; and applications where traditional lifelog data is used in a sports context to predict athletes' performance. PMData combines input from Fitbit Versa 2 smartwatch wristbands, the PMSys sports logging smartphone application, and Google forms. Logging data has been collected from 16 persons for five months. Our initial experiments show that novel analyses are possible, but there is still room for improvement.
Vajira Thambawita, Steven Alexander Hicks, Hanna Borgli, Håkon Kvale Stensland, Debesh Jha, Martin Kristoffer Svensen, Svein Arne Pettersen, Dag Johansen, Håvard D. Johansen, Susann Dahl Pettersen, Simon Nordvang, Sigurd Pedersen, Anders T. Gjerdrum, Tor-Morten Grønli, Per Morten Fredriksen, Ragnhild Eg, Kjeld Hansen, Siri Fagernes, Christine Claudi, Andreas Biørn-Hansen, Duc-Tien Dang-Nguyen, Tomas Kupka, Hugo Hammer, Ramesh Jain 0001, Michael Riegler 0001, Pål Halvorsen
MMSys22
2019 Predicting Peek Readiness-to-Train of Soccer Players Using Long Short-Term Memory Recurrent Neural Networks
abstract
We are witnessing the emergence of a myriad of hardware and software systems that quantifies sport and physical activities. These are frequently touted as game changers and important for future sport developments. The vast amount of generated data is often visualized in graphs and dashboards, for use by coaches and other sports professionals to make decisions on training and match strategies. Modern machine-learning methods has the potential to further fuel this process by deriving useful insights that are not easily observable in the raw data streams. This paper tackles the problem of deriving peaks in soccer players' ability to perform from subjective self-reported wellness data collected using the PMSys system. For this, we train a long short-term memory recurrent neural network model using data from two professional Norwegian soccer teams. We show that our model can predict performance peaks in most scenarios with a precision and recall of at least 90%. Equipped with such insight, coaches and trainers can better plan individual and team training sessions, and perhaps avoid over training and injuries.
Theodor Wiik, Håvard D. Johansen, Svein Arne Pettersen, Ivan Baptista, Tomas Kupka, Dag Johansen, Michael Riegler 0001, Pål Halvorsen
CBMI5
2018 Efficient Live and on-Demand Tiled HEVC 360 VR Video Streaming
abstract
With 360° panorama video technology becoming commonplace, the need for efficient streaming methods for such videos arises. We go beyond the existing on-demand solutions and present a live streaming system which strikes a trade-off between bandwidth usage and the video quality in the user’s field-of-view. We have created an architecture that combines RTP and DASH to deliver 360° VR content to a Huawei set-top-box and a Samsung Galaxy S7. Our system multiplexes a single HEVC hardware decoder to provide faster quality switching than at the traditional GOP boundaries. We demonstrate the performance and illustrate the trade-offs through real-world experiments where we can report comparable bandwidth savings to existing on-demand approaches, but with faster quality switches when the field-of-view changes.
Mattis Jeppsson, Håvard Espeland, Tomas Kupka, Ragnar Langseth, Andreas Petlund, Peng Qiaoqiao, Chuansong Xue, Konstantin Pogorelov, Michael Riegler 0001, Dag Johansen, Carsten Griwodz, Pål Halvorsen
ISM3
2012 Performance of on-off traffic stemming from live adaptive segmented HTTP video streaming
abstract
A large number of live segmented adaptive HTTP video streaming services exist in the Internet today. These quasi-live solutions have been shown to scale to a large number of concurrent users, but the characteristic on-off traffic pattern makes TCP behave differently compared to the bulk transfers the protocol is designed for. In this paper, we analyze the TCP performance of such live on-off sources, and we investigate possible improvements in order to increase the resource utilization on the server side. We observe that the problem is the bandwidth wastage because of the synchronization of the on period. We investigate four different techniques to mitigate this problem. We first evaluate the techniques on pure on-off traffic using a fixed quality and then repeat the experiments with quality adaptation.
Tomas Kupka, Pål Halvorsen, Carsten Griwodz
LCN1
2012 Search-based composition, streaming and playback of video archive content
abstract
Locating content in existing video archives is both a time and bandwidth consuming process since users might have to download and manually watch large portions of superfluous videos. In this paper, we present two novel prototypes using an Internet based video composition and streaming system with a keyword-based search interface that collects, converts, analyses, indexes, and ranks video content. At user requests, the system can automatically sequence out portions of single videos or aggregate content from multiple videos to produce a single, personalized video stream on-the-fly.
Dag Johansen, Pål Halvorsen, Håvard D. Johansen, Håkon Riiser, Cathal Gurrin, Bjørn Olstad, Carsten Griwodz, Åge Kvalnes, Joseph Hurley, Tomas Kupka
Multim. Tools Appl.10
2011 An evaluation of live adaptive HTTP segment streaming request strategies
abstract
Nowadays, several live and on-demand streaming solutions use HTTP for signaling and data delivery. A frequently used technique is to chop a continuous stream into segments, encode these in multiple qualities and make these available for download using plain HTTP methods. This approach has become known as dynamic adaptive segment streaming over HTTP. Its advantage is that the deployed web infrastructure is easily reused, even for live segment streaming. In this case, however, it is not strictly bulk traffic. We show in this paper, that the streaming source is essentially an on-off source. Furthermore, this paper analyzes several client-controlled segment request strategies for live adaptive HTTP segment streaming. We present experimental results showing the benefits and drawbacks of each strategy with respect to achieved video quality, smoothness of playback and end-to-end delay. We show that it matters how clients request segments. The results indicate strongly that synchronization of client requests has a negative impact on router queues and leads to increased packet loss, and should thus be avoided to achieve a high goodput.
Tomas Kupka, Pål Halvorsen, Carsten Griwodz
LCN1
2010 vESP: enriching enterprise document search results with aligned video summarization
abstract
In this demo, we present a video-enabled enterprise search platform (vESP), an application prototype that enhance a widely deployed commercial enterprise search engine with video streaming. The idea is that for example in a large enterprise, like Microsoft, there exists a lot of information in form of presentations with corresponding video. Using our enhancements, a user can select and combine slides from different presentations generating a new slide deck dynamically and the corresponding video clips are concatenated and presented vis-a-vis the slides on-the-fly. The prototype is evaluated using a data set from Microsoft, and our initial user surveys indicate that the opportunity to enrich the search results with corresponding video is embraced by potential users
Pål Halvorsen, Dag Johansen, Bjørn Olstad, Tomas Kupka, Sverre Tennøe
ACM Multimedia4
2010 Quality-adaptive scheduling for live streaming over multiple access networks
abstract
Video streaming ranks among the most popular services offered through the Internet today. At the same time, accessing the Internet over public WiFi and 3G networks has become part of our everyday lives. However, streaming video in wireless environments is often subject to frequent periods of rebuffering and characterized by low picture quality. In particular, achieving smooth and quality-adaptive streaming of live video poses a big challenge in mobile scenarios.
Kristian Evensen, Tomas Kupka, Dominik Kaspar, Pål Halvorsen, Carsten Griwodz
NOSSDAV2
2010 vESP: A Video-Enabled Enterprise Search Platform
abstract
In this paper, we present how to provide a novel and potentially disruptive multimedia service by modifying a widely deployed commercial enterprise search engine. The idea is to transparently integrate rich multimedia data with traditional textual-oriented query results. This includes that the search engine automatically discovers and extracts relevant scenes from a large knowledge repository of existing videos and produces a new, customized video of events matching the user query. To evaluate our prototype, we have performed experiments using a data set from a knowledge repository in Microsoft consisting of PowerPoint presentations with corresponding videos. Our initial results demonstrate that such integration can be implemented efficiently, and that potential users prefer to have the opportunity to enrich the search results with corresponding video.
Pål Halvorsen, Dag Johansen, Bjørn Olstad, Tomas Kupka, Sverre Tennøe
NSS4
2006 Reliable VoIP Services Using a Peer-to-Peer Intranet
abstract
From state of the art technologies for load balancing and reliability increase in VoIP infrastructures one learns that no homogenous solution exists. Thus compromises have been made as in the number of machines versus achieved reliability, or with regard to separation of load balancing and failure resilience. In this paper we present a distributed architecture for VoIP services based on peer-to-peer principles. In contrast to other peer-to-peer based approaches, relying upon an external DHT service, our architecture utilizes peer-to-peer technologies to integrate load balancing and failover requirements with a centralized VoIP server concept. This results in an integrated solution with focus on todays service provider requirements. Within the scope of this work, a proof-of-concept software solution has been designed and implemented. We show that the proposed architecture can be easily scaled up and provides efficient, redundant storage for user and session data.
Jens Fiedler, Tomas Kupka, Thomas Magedanz, Michael Kleis
ISM2