EDBT 2026 Demo / reviewers in the wild / expert
Sushant Gautam
dblp:358/2564
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-9232-2661ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ImageCLEF 2026: Multimodal Challenges in Medicine, Science, Agritech, and Security
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandra Baicoianu, Ana Neacsu, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Lea Reinartz, Benjamin Lecouteux, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Corneliu Florea, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Hendrik Damm, Henning Schäfer, Ivan Koychev, Josiane Mothe, Liviu-Daniel Stefan, Maja J. Hjuler, Mehmet Kurt, Meliha Yetisgen, Michael Riegler 0001, Mihai Dogariu, Mihai Ivanovici, Ming Shan Hee, Mohammad El Sakka, Momina Ahsan, Obioma Pelka, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Bahadir Eryilmaz, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Yuri Prokopchuk, Zhuohan Xie |
ECIR (4) | 42 |
| 2025 | SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game UnderstandingabstractArtificial intelligence (AI) is transforming sports analytics by enabling automated, real-time understanding of soccer matches. Traditional approaches that rely on isolated data streams struggle to capture the full game context. We introduce SoccerChat, a multimodal conversational AI framework that fuses visual and textual information for comprehensive soccer video comprehension. Building on the SoccerNet dataset, we enrich it with jersey color annotations and automatic speech recognition (ASR) transcripts and curate a video-instruction dataset containing 48,677 question-answer (QA) pairs. Fine-tuning Qwen2- VL-7B- Instruct on this resource yields the Soc-cerChat model, which supports accurate event interpretation, classification, and referee assistance. Experiments across action classification and referee QA tasks demonstrate strong general event understanding and competitive officiating analysis. These results highlight the importance of multimodal integration for explainable, interactive, and trustworthy AI-driven sports ana-lytics. Links to the code, dataset, and model weights are available at https://2ithub.com/simula/SoccerChat. Sushant Gautam, Cise Midoglu, Vajira Thambawita, Michael Riegler 0001, Pål Halvorsen, Mubarak Shah |
CBMI | 1 |
| 2025 | Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language ModelsabstractWe investigate fine-tuning Vision-Language Models (VLMs) for multi-task medical image understanding, focusing on detection, localization, and counting of findings in medical images. Our objective is to evaluate whether instruction-tuned VLMs can simultaneously improve these tasks, with the goal of enhancing diagnostic accuracy and efficiency. Using MedMultiPoints, a multimodal dataset with annotations from endoscopy (polyps and instruments) and microscopy (sperm cells), we reformulate each task into instruction-based prompts suitable for vision-language reasoning. We fine-tune Qwen2.5-VL-7BInstruct using Low-Rank Adaptation (LoRA) across multiple task combinations. Results show that multi-task training improves robustness and accuracy. For example, it reduces the Count Mean Absolute Error (MAE) and increases Matching Accuracy in the Counting + Pointing task. However, trade-offs emerge, such as more zero-case point predictions, indicating reduced reliability in edge cases despite overall performance gains. Our study highlights the potential of adapting general-purpose VLMs to specialized medical tasks via prompt-driven fine-tuning. This approach mirrors clinical workflows, where radiologists simultaneously localize, count, and describe findings - demonstrating how VLMs can learn composite diagnostic reasoning patterns. The model produces interpretable, structured outputs, offering a promising step toward explainable and versatile medical AI. Code, model weights, and scripts will be released for reproducibility at https://github.com/simula/PointDetectCount. Sushant Gautam, Michael Riegler 0001, Pål Halvorsen |
CBMS | 1 |
| 2025 | ImageCLEF 2025: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Helmut Becker, Hendrik Damm, Henning Schäfer, Ivan Rodkin, Ivan Koychev, Johannes Kiesel, Johannes Rückert, Josep Malvehy, Liviu-Daniel Stefan, Louise Bloch, Martin Potthast, Maximilian Heinrich, Michael Riegler 0001, Mihai Dogariu, Noel Codella, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Roberto A. Novoa, Rocktim Jyoti Das, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Zhuohan Xie |
ECIR (5) | 42 |
| 2025 | HockeyAI: A Multi-Class Ice Hockey Dataset for Object DetectionabstractThe fast paced nature of ice hockey presents unique challenges for object detection, particularly in tracking the puck---a small, fast moving object that is critical to gameplay analysis. This paper introduces HockeyAI, a novel open source dataset specifically designed for multi-class object detection in ice hockey. The dataset includes 2,101 high resolution frames extracted from professional games in the Swedish Hockey League (SHL), annotated in the You Look Only Once (YOLO) format. Annotations span 7 classes, covering dynamic objects such as the players and the puck, as well as static rink elements such as goalposts and face-off circles. The dataset is derived from diverse SHL games across multiple seasons and teams, ensuring a rich variety of scenarios reflective of real world gameplay. A fine tuned YOLOv8 medium model is also provided, demonstrating high performance across all classes. Key comparisons highlight the dataset's advancements over existing resources, addressing their limitations such as incomplete class coverage, low resolution, and inconsistent annotations. The dataset, model, and an interactive demo are publicly available on Hugging Face under an open source license, fostering further collaboration in sports related computer vision applications (https://huggingface.co/SimulaMet-HOST/HockeyAI). Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen |
MMSys | 2 |
| 2025 | HockeyRink: A Dataset for Precise Ice Hockey Rink Keypoint Mapping and AnalyticsabstractPrecise mapping of ice hockey rinks is critical for applications such as player tracking, game strategy analysis, and broadcast enhancements. Traditional methods often rely on manual annotations or simplistic models that fail to account for the rink's complex geometry and dynamic in-game conditions. To address these limitations, we present HockeyRink, a novel dataset comprising 56 meticulously annotated keypoints corresponding to significant landmarks on a standard hockey rink, including face-off dots, goalposts, and blue lines. Leveraging the YOLOv8-Large pose estimation architecture, we adapted the model to treat the rink as a single 'pose' object, enabling accurate keypoint predictions tailored to the nuances of hockey rink imagery. Our dataset, derived from diverse hockey game footage from the Swedish Hockey League (SHL), facilitates applications in homography estimation, 2D/3D scene mapping, and tactical overlays. By detecting and mapping rink keypoints, one can compute precise transformations between the image plane and the rink's physical dimensions, enabling the overlay of player trajectories and tactical insights onto broadcast footage. This work addresses the unique challenges of ice hockey, such as occlusions, rapid camera movements, and varying lighting conditions, and offers a robust foundation for future research and innovation in sports analytics. The HockeyRink dataset and trained model are openly available at: https://huggingface.co/SimulaMet-HOST/HockeyRink. Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen |
MMSys | 2 |
| 2025 | HockeyOrient: A Dataset for Ice Hockey Player Orientation ClassificationabstractUnderstanding player orientation is a critical component of sports analytics, offering insights into gameplay strategies and player behavior. This paper presents HockeyOrient, a novel dataset for classifying the orientation of ice hockey players based on their poses. The dataset comprises 9,700 manually annotated frames, selected randomly and non-sequentially, taken from Swedish Hockey League (SHL) games during the 2023 and 2024 seasons. Each player image is cropped from game footage and categorized into one of eight orientation classes: top, top-right, right, bottom-right, bottom, bottom-left, left, and top-left. The dataset includes diverse scenarios, such as different teams, jersey colors, referees, and goaltenders with unique protective gear. Alongside the dataset, we provide an open-source classification model trained on the dataset using the SqueezeNet architecture, achieving an F1 score of 75% across all the classes. This work addresses a significant gap in ice hockey analytics, enabling advanced player tracking and gameplay analysis. The dataset and the trained model can be accessed publicly on Hugging Face under an open-source license (https://huggingface.co/datasets/SimulaMet-HOST/HockeyOrient). Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Tomas Kupka, Pål Halvorsen |
MMSys | 2 |
| 2024 | Demo: Creating Player-Specific Soccer Highlight Clips with PlayerTVabstractThis paper demonstrates PlayerTV, an innovative framework which harnesses state-of-the-art Artificial Intelligence (AI) technologies for automatic player tracking and identification in soccer videos. By integrating object detection and tracking, Optical Character Recognition (OCR), and color analysis, PlayerTV facilitates the generation of player-specific highlight clips from extensive game footage, significantly reducing the manual labor traditionally associated with such tasks. We present how PlayerTV can be run standalone as a core pipeline, as well as through an interactive Graphical User Interface (GUI). Håkon Maric Solberg, Mehdi Houshmand Sarkhoosh, Sushant Gautam, Saeed Shafiee Sabet, Pål Halvorsen, Cise Midoglu |
CBMI | 3 |
| 2024 | SoccerRAG: Multimodal Soccer Information Retrieval via Natural QueriesabstractThe rapid evolution of digital sports media necessitates sophisticated information retrieval systems that can efficiently parse extensive multimodal datasets. In this paper, we introduce SoccerRAG, an innovative framework designed to harness the power of Retrieval Augmented Generation (RAG) and Large Language Models (LLMs) to extract soccer-related information through natural language queries. By leveraging a multimodal dataset, SoccerRAG supports dynamic querying and automatic data validation, enhancing user interaction and accessibility to sports archives. Our evaluations indicate that SoccerRAG effectively handles complex queries, offering significant improvements over traditional retrieval systems in terms of accuracy and user engagement. The results underscore the potential of using RAG and LLMs in sports analytics, paving the way for future advancements in the accessibility and real-time processing of sports data. Aleksander Theo Strand, Sushant Gautam, Cise Midoglu, Pål Halvorsen |
CBMI | 2 |
| 2024 | Demo: Soccer Information Retrieval Via Natural Queries using SoccerRAGabstractThe rapid evolution of digital sports media necessitates sophisticated information retrieval systems that can efficiently parse extensive multimodal datasets. This paper demonstrates SoccerRAG, an innovative framework designed to harness the power of Retrieval Augmented Generation (RAG) and Large Language Models (LLMs) to extract soccer-related information through natural language queries. By leveraging a multimodal dataset, SoccerRAG supports dynamic querying and automatic data validation, enhancing user interaction and accessibility to sports archives. We present a novel interactive user interface (UI) based on the Chainlit framework which wraps around the core functionality, and enable users to interact with the SoccerRAG framework in a chatbot-like visual manner. Aleksander Theo Strand, Sushant Gautam, Cise Midoglu, Pål Halvorsen |
CBMI | 2 |
| 2024 | SoccerNet-Echoes: A Soccer Game Audio Commentary DatasetabstractThe application of Automatic Speech Recognition (ASR) technology in soccer enables sports analytics by extracting audio commentaries to provide insights into game events and facilitate automatic game understanding. This paper presents SoccerNet-Echoes, an extension of the SoccerNet dataset with automatically generated transcriptions of soccer game broadcasts. Generated using the Whisper model and translated with Google Translate into English when needed, these transcriptions enhance video content with textual information derived from game audio. SoccerNet-Echoes serves as a comprehensive resource for developing algorithms in action spotting, caption generation, and game summarization. Through a series of experiments, we demonstrate that combining modalities—audio, video, and text—yields mixed results on classification tasks. The combination of audio and video shows improved performance over individual modalities, while the addition of ASR text does not significantly enhance results. Additionally, our baseline summarization tasks indicate that ASR content enriches summaries, offering insights beyond event information. This multimodal dataset supports diverse applications, broadening the scope of research in sports analytics. The dataset is available at: https://github.com/SoccerNet/sn-echoes. Sushant Gautam, Mehdi Houshmand Sarkhoosh, Jan Held, Cise Midoglu, Anthony Cioppa, Silvio Giancola, Vajira Thambawita, Michael Riegler 0001, Pål Halvorsen, Mubarak Shah |
ISM | 1 |
| 2024 | PlayerTV: Advanced Player Tracking and Identification for Automatic Soccer Highlight ClipsabstractIn the rapidly evolving field of sports analytics, the automation of targeted video processing is a pivotal advancement. We propose PlayerTV, an innovative framework which harnesses state-of-the-art AI technologies for automatic player tracking and identification in soccer videos. By integrating object detection and tracking, Optical Character Recognition (OCR), and color analysis, PlayerTV facilitates the generation of player-specific highlight clips from extensive game footage, significantly reducing the manual labor traditionally associated with such tasks. Preliminary results from the evaluation of our core pipeline, tested on a dataset from the Norwegian Eliteserien league, indicate that PlayerTV can accurately and efficiently identify teams and players, and our interactive Graphical User Interface (GUI) serves as a user-friendly application wrapping this functionality for streamlined use. Håkon Maric Solberg, Mehdi Houshmand Sarkhoosh, Sushant Gautam, Saeed Shafiee Sabet, Pål Halvorsen, Cise Midoglu |
ISM | 3 |
| 2024 | TACDEC: Dataset of Tackle Events in Soccer Game VideosabstractThis paper introduces TACDEC, a dataset of tackle events in soccer game videos. Recognizing the gap in existing open datasets that predominantly focus on official soccer events such as goals and cards, TACDEC targets a comprehensive analysis of tackles --- a critical aspect of soccer that combines technical skills, tactical decision-making, and physical engagement. By leveraging video data from the Norwegian Eliteserien league across multiple seasons, we annotated 425 videos with 4 types of tackle events, categorized into "tackle-live", "tackle-replay", "tackle-live-incomplete", and "tackle-replay-incomplete", yielding a total of 836 event annotations. The dataset offers an unprecedented resource for the development and testing of machine learning models aimed at understanding and analyzing soccer game dynamics. A proof-of-concept classification model demonstrates the dataset's utility, achieving promising results in automatic tackle detection, thereby validating TACDEC's potential to support not only advanced game analytics but also to enhance fan engagement and player development initiatives. Evan Jåsund Kassab, Håkon Maric Solberg, Sushant Gautam, Saeed Shafiee Sabet, Thomas Torjusen, Michael Riegler 0001, Pål Halvorsen, Cise Midoglu |
MMSys | 3 |
| 2024 | Multimodal AI-Based Summarization and Storytelling for Soccer on Social MediaabstractThe rapid advancement of technology has been revolutionizing the field of sports media, where there is a growing need for sophisticated data processing methods. Current methodologies for extracting information from soccer broadcast videos to generate game highlights and summaries for social media are predominantly manual and rely heavily on text-based NLP techniques, overlooking the rich visual and auditory information available. In response to this challenge, our research introduces SoccerSum, a tool that innovates in the field by integrating computer vision, audio analysis with advanced language models like GPT-4. This multimodal approach enables automated, enriched content summarization, including detection of players and key field elements, thereby enhancing the metadata used in summarization algorithms. SoccerSum uniquely combines textual and visual data, offering a comprehensive solution for generating accurate, platform-specific content. This development represents a significant advancement in automated, data-driven sports media dissemination, and sets a new benchmark in the realm of soccer information extraction. A video of the demo can be found here: https://youtu.be/za4VIi2ARXY. Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Pål Halvorsen |
MMSys | 2 |
| 2024 | The SoccerSum Dataset for Automated Detection, Segmentation, and Tracking of Objects on the Soccer PitchabstractThis paper introduces SoccerSum, a novel dataset aimed at enhancing object detection and segmentation in video frames depicting the soccer pitch, using footage from the Norwegian Eliteserien league across 2021-2023. With the goal of detecting elements beyond common entities in existing datasets, such as the soccer ball, players and referees, this dataset includes additional annotations for the goal net, corner flag posts, and the penalty mark. SoccerSum also includes the segmentation of key pitch areas such as the penalty and goal boxes for the same frame sequences. Comprising 750 frames annotated with 10 classes for advanced analysis, SoccerSum offers compatibility with existing frameworks, providing a rich dataset for the development of computer vision algorithms. This dataset not only serves as a resource for improving sports analytics, but also introduces a new application for automatic game summarization, enabling the generation of detailed and engaging content for fans and professionals. The SoccerSum dataset is accessible on Zenodo: https://zenodo.org/records/10612084. Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Saeed Shafiee Sabet, Thomas Torjusen, Pål Halvorsen |
MMSys | 2 |