EDBT 2026 Demo / reviewers in the wild / expert
Vajira Thambawita
dblp:156/0270 · also Vajira Lasantha Thambawita
· DBLP profile ↗
32ranked-venue papers
3as first author
24since 2021 · last 2026
0000-0001-6026-0929ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ImageCLEF 2026: Multimodal Challenges in Medicine, Science, Agritech, and Security
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandra Baicoianu, Ana Neacsu, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Lea Reinartz, Benjamin Lecouteux, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Corneliu Florea, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Hendrik Damm, Henning Schäfer, Ivan Koychev, Josiane Mothe, Liviu-Daniel Stefan, Maja J. Hjuler, Mehmet Kurt, Meliha Yetisgen, Michael Riegler 0001, Mihai Dogariu, Mihai Ivanovici, Ming Shan Hee, Mohammad El Sakka, Momina Ahsan, Obioma Pelka, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Bahadir Eryilmaz, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Yuri Prokopchuk, Zhuohan Xie |
ECIR (4) | 45 |
| 2026 | A Comparative Study of Decoding Strategies in Medical Text Generation
Oriana Presacan, Alireza Nik, Vajira Thambawita, Bogdan Ionescu, Michael Riegler 0001 |
MMM (2) | 3 |
| 2026 | Calliope: A TTS-based Narrated E-book Creator Ensuring Exact Synchronization, Privacy, and Layout FidelityabstractA narrated e-book combines synchronized audio with digital text, highlighting the currently spoken word or sentence during playback. This format supports early literacy and assists individuals with reading challenges, while also allowing general readers to seamlessly switch between reading and listening. With the emergence of natural-sounding neural Text-to-Speech (TTS) technology, several commercial services have been developed to leverage these technology for converting standard text e-books into high-quality narrated e-books. However, no open-source solutions currently exist to perform this task. In this paper, we present Calliope, an open-source framework designed to fill this gap. Our method leverages state-of-the-art open-source TTS to convert a text e-book into a narrated e-book in the EPUB 3 Media Overlay format. The method offers several innovative steps: audio timestamps are captured directly during TTS, ensuring exact synchronization between narration and text highlighting; the publisher's original typography, styling, and embedded media are strictly preserved; and the entire pipeline operates offline. This offline capability eliminates recurring API costs, mitigates privacy concerns, and avoids copyright compliance issues associated with cloud-based services. The framework currently supports the state-of-the-art open-source TTS systems XTTS-v2 and Chatterbox. A potential alternative approach involves first generating narration via TTS and subsequently synchronizing it with the text using forced alignment. However, while our method ensures exact synchronization, our experiments show that forced alignment introduces drift between the audio and text highlighting significant enough to degrade the reading experience. Source code and usage instructions are available at https://github.com/hugohammer/TTS-Narrated-Ebook-Creator.git. Hugo Hammer, Vajira Thambawita, Pål Halvorsen |
MMSys | 2 |
| 2025 | SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game UnderstandingabstractArtificial intelligence (AI) is transforming sports analytics by enabling automated, real-time understanding of soccer matches. Traditional approaches that rely on isolated data streams struggle to capture the full game context. We introduce SoccerChat, a multimodal conversational AI framework that fuses visual and textual information for comprehensive soccer video comprehension. Building on the SoccerNet dataset, we enrich it with jersey color annotations and automatic speech recognition (ASR) transcripts and curate a video-instruction dataset containing 48,677 question-answer (QA) pairs. Fine-tuning Qwen2- VL-7B- Instruct on this resource yields the Soc-cerChat model, which supports accurate event interpretation, classification, and referee assistance. Experiments across action classification and referee QA tasks demonstrate strong general event understanding and competitive officiating analysis. These results highlight the importance of multimodal integration for explainable, interactive, and trustworthy AI-driven sports ana-lytics. Links to the code, dataset, and model weights are available at https://2ithub.com/simula/SoccerChat. Sushant Gautam, Cise Midoglu, Vajira Thambawita, Michael Riegler 0001, Pål Halvorsen, Mubarak Shah |
CBMI | 3 |
| 2025 | ImageCLEF 2025: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Helmut Becker, Hendrik Damm, Henning Schäfer, Ivan Rodkin, Ivan Koychev, Johannes Kiesel, Johannes Rückert, Josep Malvehy, Liviu-Daniel Stefan, Louise Bloch, Martin Potthast, Maximilian Heinrich, Michael Riegler 0001, Mihai Dogariu, Noel Codella, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Roberto A. Novoa, Rocktim Jyoti Das, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Zhuohan Xie |
ECIR (5) | 44 |
| 2025 | Evaluating gradient-based explanation methods for neural network ECG analysis using heatmapsabstractOBJECTIVE: Evaluate popular explanation methods using heatmap visualizations to explain the predictions of deep neural networks for electrocardiogram (ECG) analysis and provide recommendations for selection of explanations methods. MATERIALS AND METHODS: A residual deep neural network was trained on ECGs to predict intervals and amplitudes. Nine commonly used explanation methods (Saliency, Deconvolution, Guided backpropagation, Gradient SHAP, SmoothGrad, Input × gradient, DeepLIFT, Integrated gradients, GradCAM) were qualitatively evaluated by medical experts and objectively evaluated using a perturbation-based method. RESULTS: No single explanation method consistently outperformed the other methods, but some methods were clearly inferior. We found considerable disagreement between the human expert evaluation and the objective evaluation by perturbation. DISCUSSION: The best explanation method depended on the ECG measure. To ensure that future explanations of deep neural networks for medical data analyses are useful to medical experts, data scientists developing new explanation methods should collaborate tightly with domain experts. Because there is no explanation method that performs best in all use cases, several methods should be applied. CONCLUSION: Several explanation methods should be used to determine the most suitable approach. Andrea M. Storås, Steffen Mæland, Jonas Isaksen, Steven Alexander Hicks, Vajira Thambawita, Claus Graff, Hugo Hammer, Pål Halvorsen, Michael Riegler 0001, Jørgen K. Kanters |
J. Am. Medical Informatics Assoc. | 5 |
| 2025 | Validating polyp and instrument segmentation methods in colonoscopy through Medico 2020 and MedAI 2021 ChallengesabstractAutomatic analysis of colonoscopy images has been an active field of research motivated by the importance of early detection of precancerous polyps. However, detecting polyps during the live examination can be challenging due to various factors such as variation of skills and experience among the endoscopists, lack of attentiveness, and fatigue leading to a high polyp miss-rate. Therefore, there is a need for an automated system that can flag missed polyps during the examination and improve patient care. Deep learning has emerged as a promising solution to this challenge as it can assist endoscopists in detecting and classifying overlooked polyps and abnormalities in real time, improving the accuracy of diagnosis and enhancing treatment. In addition to the algorithm’s accuracy, transparency and interpretability are crucial to explaining the whys and hows of the algorithm’s prediction. Further, conclusions based on incorrect decisions may be fatal, especially in medicine. Despite these pitfalls, most algorithms are developed in private data, closed source, or proprietary software, and methods lack reproducibility. Therefore, to promote the development of efficient and transparent methods, we have organized the “Medico automatic polyp segmentation (Medico 2020)” and “MedAI: Transparency in Medical Image Segmentation (MedAI 2021)” competitions. The Medico 2020 challenge received submissions from 17 teams, while the MedAI 2021 challenge also gathered submissions from another 17 distinct teams in the following year. We present a comprehensive summary and analyze each contribution, highlight the strength of the best-performing methods, and discuss the possibility of clinical translations of such methods into the clinic. Our analysis revealed that the participants improved dice coefficient metrics from 0.8607 in 2020 to 0.8993 in 2021 despite adding diverse and challenging frames (containing irregular, smaller, sessile, or flat polyps), which are frequently missed during a routine clinical examination. For the instrument segmentation task, the best team obtained a mean Intersection over union metric of 0.9364. For the transparency task, a multi-disciplinary team, including expert gastroenterologists, accessed each submission and evaluated the team based on open-source practices, failure case analysis, ablation studies, usability and understandability of evaluations to gain a deeper understanding of the models’ credibility for clinical deployment. The best team obtained a final transparency score of 21 out of 25. Through the comprehensive analysis of the challenge, we not only highlight the advancements in polyp and surgical instrument segmentation but also encourage subjective evaluation for building more transparent and understandable AI-based colonoscopy systems. Moreover, we discuss the need for multi-center and out-of-distribution testing to address the current limitations of the methods to reduce the cancer burden and improve patient care. • We present a detailed analysis of the Medico 2020 and MedAI 2021 challenges that are aimed at advancing automated polyp and instrument segmentation in colonoscopy for early colorectal cancer diagnosis by using novel deep learning methods. • To the best of our knowledge, MedAI 2021 is the first challenge to evaluate the transparency in both GI endoscopy and colonoscopy. Through the challenge, we invited the participants to list package dependencies and architecture code (with instructions for building, compiling, and training) and share trained model weights in a standardized format. Additionally, we invited participants to include the code for model evaluation and provide repository licensing information to enable others to use the code and the trained model responsibly. Moreover, we asked the participants to explain model predictions using intermediate heatmaps, perform ablation studies, conduct a thorough failure analysis, and share their code for reproducing the results. Finally, we performed a subjective evaluation by including an expert gastroenterologist in the group and gave the final transparency score based on the usefulness and understandability of the results. Our initiative aims to promote transparency in AI research and foster the development of reliable, interpretable, and trustworthy algorithms for use in medical image segmentation. • We provide a comparative analysis of the 34 proposed methods in both challenges (3 subtasks), covering small details of each team in the form of Tables, qualitative and quantitative results (failure analysis), and an in-depth analysis of the findings. • We explore trust, safety, interpretability, transparency, and generalizability issues and provide future strategies to overcome the current limitations of developed algorithms. Debesh Jha, Vanshali Sharma, Debapriya Banik, Debayan Bhattacharya, Kaushiki Roy, Steven Alexander Hicks, Nikhil Kumar Tomar, Vajira Thambawita, Adrian Krenzer, Ge-Peng Ji, Sahadev Poudel, George Batchkala, Saruar Alam, Awadelrahman M. A. Ahmed, Quoc-Huy Trinh, Zeshan Khan, Tien-Phat Nguyen, Shruti Shrestha, Sabari Nathan, Jeonghwan Gwak, Ritika Kumari Jha, Zheyuan Zhang 0001, Alexander Schlaefer, Debotosh Bhattacharjee, Manas Kamal Bhuyan, Pradip K. Das, Deng-Ping Fan, Sravanthi Parasa, Sharib Ali, Michael Riegler 0001, Pål Halvorsen, Thomas de Lange, Ulas Bagci |
Medical Image Anal. | 8 |
| 2024 | SAM in the Pipeline: Transforming Axis-Aligned to Oriented Bounding Boxes for Superior Sperm DetectionabstractSperm analysis, in terms of finding the best cells, has traditionally been performed manually and is slow, expensive, prone to human error, and often inconsistent across different laboratory technicians. Our research focuses on the use of artificial intelligence to improve computer-assisted sperm analysis, a critical aspect in the diagnosis and treatment of male fertility issues. Using the best publicly available dataset for sperm detection, we employ novel data refining techniques and state-of-the-art object detection models to develop the best-performing sperm detection model to our knowledge. We show how the Segment Anything Model (SAM) can be used as a data-refining tool to improve object detection datasets. Additionally, we use SAM to convert Axis-Aligned Bounding Box (AABB) to Oriented Bounding Box (OBB) to get more descriptive coordinates. We test our data-refining pipeline with state-of-the-art object detection models and a publicly available sperm-tracking dataset. Our results show that the refined data coming from our pipeline can be used to train lighter models while still achieving higher accuracy compared to using the original dataset. This also enables possibilities for running high-quality tracking on light hardware, such as consumer-grade laptops. Pål Andreas Hoven Bentsen, Steven Alexander Hicks, Eric Jul, Pål Halvorsen, Vajira Thambawita |
CBMI | 5 |
| 2024 | Advancing Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications with ImageCLEF 2024
Bogdan Ionescu, Henning Müller, Ana-Maria Claudia Dragulinescu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandru Stan, Andrea M. Storås, Asma Ben Abacha, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Didier Schwab, Emmanuelle Esperança-Rodier, George Ioannidis, Griffin Adams, Henning Schäfer, Hugo Manguinhas, Ioan Coman, Johanna Schöler, Johannes Kiesel, Johannes Rückert, Louise Bloch, Martin Potthast, Maximilian Heinrich, Meliha Yetisgen, Michael Riegler 0001, Neal Snider, Pål Halvorsen, Raphael Brüngel, Steven Alexander Hicks, Vajira Thambawita, Vassili Kovalev, Yuri Prokopchuk, Wen-Wai Yim |
ECIR (6) | 35 |
| 2024 | SoccerNet-Echoes: A Soccer Game Audio Commentary DatasetabstractThe application of Automatic Speech Recognition (ASR) technology in soccer enables sports analytics by extracting audio commentaries to provide insights into game events and facilitate automatic game understanding. This paper presents SoccerNet-Echoes, an extension of the SoccerNet dataset with automatically generated transcriptions of soccer game broadcasts. Generated using the Whisper model and translated with Google Translate into English when needed, these transcriptions enhance video content with textual information derived from game audio. SoccerNet-Echoes serves as a comprehensive resource for developing algorithms in action spotting, caption generation, and game summarization. Through a series of experiments, we demonstrate that combining modalities—audio, video, and text—yields mixed results on classification tasks. The combination of audio and video shows improved performance over individual modalities, while the addition of ASR text does not significantly enhance results. Additionally, our baseline summarization tasks indicate that ASR content enriches summaries, offering insights beyond event information. This multimodal dataset supports diverse applications, broadening the scope of research in sports analytics. The dataset is available at: https://github.com/SoccerNet/sn-echoes. Sushant Gautam, Mehdi Houshmand Sarkhoosh, Jan Held, Cise Midoglu, Anthony Cioppa, Silvio Giancola, Vajira Thambawita, Michael Riegler 0001, Pål Halvorsen, Mubarak Shah |
ISM | 7 |
| 2023 | RePolyp: A Framework for Generating Realistic Colon Polyps with Corresponding Segmentation Masks using Diffusion ModelsabstractThe field of synthetic medical data has become increasingly important due to the urgent need for large and diverse datasets in the medical sector. Using diffusion models in data generation has created more authentic and varied medical data. In this study, a framework is presented that utilizes diffusion models trained on openly accessible data to generate realistic-looking colon polyps, along with their corresponding ground truth masks. The usefulness of the synthetic polyps is evaluated by using them to train segmentation models designed to segment colon polyps in real-world images. The results demonstrate that the generated synthetic data is highly accurate and suggest that including synthetic polyps in the training dataset improves the predictive performance and generalization of the segmentation models. When the training dataset consists of pre-generated synthetic data from our model, we achieve a mean intersection over union (mIoU) improvement of 4.64% on the validation data and a 4.14% mIoU improvement when testing across different datasets. These results indicate that generating synthetic medical data using diffusion models is valuable for addressing the need for diverse and extensive medical datasets. Alexander K. Pishva, Vajira Thambawita, Jim Tørresen, Steven Alexander Hicks |
CBMS | 2 |
| 2023 | ImageCLEF 2023 Highlight: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Ana-Maria Claudia Dragulinescu, Adrian Popescu 0001, Ahmad Idrissi-Yaghir, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandru Stan, Andrea M. Storås, Asma Ben Abacha, Christoph M. Friedrich, George Ioannidis, Griffin Adams, Henning Schäfer, Hugo Manguinhas, Ihar Filipovich, Ioan Coman, Jérôme Deshayes-Chossart, Johanna Schöler, Johannes Rückert, Liviu-Daniel Stefan, Louise Bloch, Meliha Yetisgen, Michael Riegler 0001, Mihai Dogariu, Mihai Gabriel Constantin, Neal Snider, Nikolaos Papachrysos, Pål Halvorsen, Raphael Brüngel, Serge Kozlovski, Steven Alexander Hicks, Thomas de Lange, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim |
ECIR (3) | 34 |
| 2023 | Multimedia Datasets: Challenges and Future Possibilities
Thu Nguyen 0001, Andrea M. Storås, Vajira Thambawita, Steven Alexander Hicks, Pål Halvorsen, Michael Riegler 0001 |
MMM (2) | 3 |
| 2023 | ScopeSense: An 8.5-Month Sport, Nutrition, and Lifestyle Lifelogging DatasetabstractNowadays, most people have a smartphone that can track their everyday activities. Furthermore, a significant number of people wear advanced smartwatches to track several vital biomarkers in addition to activity data. However, it is still unclear how these data can actually be used to improve certain aspects of people’s lives. One of the key challenges is that the collected data is often massive and unstructured. Therefore, a link to other important information (e.g., when, what, and how much food was consumed) is required. It is widely believed that such detailed and structured longitudinal data about a person is essential to model and provide personalized and precise guidance. Despite the strong belief of researchers about the power of such a data-driven approach, respective datasets have been difficult to collect. In this study, we present a unique dataset from two individuals performing a structured data collection over eight and a half months. In addition to the sensor data, we collected their nutrition, training, and well-being data. The availability of nutrition data with many other important objectives and subjective longitudinal data streams may facilitate research related to food for a healthy lifestyle. Thus, we present a sport, nutrition, and lifestyle logging dataset called ScopeSense from two individuals and discuss its potential use. The dataset is fully open for researchers, and we consider this study as a potential starting point for developing methods to collect and create knowledge for a larger cohort of people. Michael Riegler 0001, Vajira Thambawita, Binh T. Nguyen 0001, Steven Alexander Hicks, Vibeke Telle-Hansen, Svein Arne Pettersen, Dag Johansen, Ramesh Jain 0001, Pål Halvorsen |
MMM (1) | 2 |
| 2022 | PolypConnect: Image inpainting for generating realistic gastrointestinal tract images with polypsabstractEarly identification of a polyp in the lower gas-trointestinal (GI) tract can lead to prevention of life-threatening colorectal cancer. Developing computer-aided diagnosis (CAD) systems to detect polyps can improve detection accuracy and efficiency and save the time of the domain experts called endoscopists. Lack of annotated data is a common challenge when building CAD systems. Generating synthetic medical data is an active research area to overcome the problem of having relatively few true positive cases in the medical domain. To be able to efficiently train machine learning (ML) models, which are the core of CAD systems, a considerable amount of data should be used. In this respect, we propose the PolypConnect pipeline, which can convert non-polyp images into polyp images to increase the size of training datasets for training. We present the whole pipeline with quantitative and qualitative evaluations involving endoscopists. The polyp segmentation model trained using synthetic data, and real data shows a 5.1% improvement of mean intersection over union (mIOU), compared to the model trained only using real data. The codes of all the experiments are available on GitHub to reproduce the results. Jan Andre Fagereng, Vajira Thambawita, Andrea M. Storås, Sravanthi Parasa, Thomas de Lange, Pål Halvorsen, Michael Riegler 0001 |
CBMS | 2 |
| 2022 | Segmentation Consistency Training: Out-of-Distribution Generalization for Medical Image SegmentationabstractGeneralizability is seen as one of the major challenges in deep learning, in particular in the domain of medical imaging, where a change of hospital or in imaging routines can lead to a complete failure of a model. To tackle this, we introduce Consistency Training, a training procedure and alternative to data augmentation based on maximizing models’ prediction consistency across augmented and unaugmented data in order to facilitate better out-of-distribution generalization. To this end, we develop a novel auxiliary region-based segmentation loss function called Segmentation Inconsistency Loss (SIL), which considers the differences between pairs of augmented and unaugmented predictions and labels. We demonstrate that Consistency Training outperforms conventional data augmentation on several out-of-distribution datasets on polyp segmentation, a popular medical task. Birk Torpmann-Hagen, Vajira Thambawita, Michael Riegler 0001, Pål Halvorsen, Kyrre Glette |
ISM | 2 |
| 2022 | Reproducibility Companion Paper: Focusing on Persons: Colorizing Old Images Learning from Modern Historical MoviesabstractIn this paper we reproduce experimental results presented in our earlier work titled "Focusing on Persons: Colorizing Old Images Learning from Modern Historical Movies" that was presented in the course of the 29th ACM International Conference on Multimedia. The paper aims at verifying the soundness of our prior results and helping others understand our software framework. We present artifacts that help reproduce results that were included in our earlier work. Specifically, this paper contains the technical details of the package, including dataset preparation, source code structure and experimental environment. Using the artifacts we show that our results are reproducible. We invite everyone to use our software framework going beyond reproducibility efforts. Xin Jin 0015, Dongqing Zou, Zhonglan Li, Heng Huang 0002, Vajira Thambawita |
ACM Multimedia | 6 |
| 2022 | Njord: a fishing trawler datasetabstractFish is one of the main sources of food worldwide. The commercial fishing industry has a lot of different aspects to consider, ranging from sustainability to reporting. The complexity of the domain also attracts a lot of research from different fields like marine biology, fishery sciences, cybernetics, and computer science. In computer science, detection of fishing vessels via for example remote sensing and classification of fish from images or videos using machine learning or other analysis methods attracts growing attention. Surprisingly, little work has been done that considers what is happening on board the fishing vessels. On the deck of the boats, a lot of data and important information are generated with potential applications, such as automatic detection of accidents or automatic reporting of fish caught. This paper presents Njord, a fishing trawler dataset consisting of surveillance videos from a modern off-shore fishing trawler at sea. The main goal of this dataset is to show the potential and possibilities that analysis of such data can provide. In addition to the data, we provide a baseline analysis and discuss several possible research questions this dataset could help answer. Tor-Arne S. Nordmo, Aril B. Ovesen, Bjørn Aslak Juliussen, Steven Alexander Hicks, Vajira Thambawita, Håvard D. Johansen, Pål Halvorsen, Michael Riegler 0001, Dag Johansen |
MMSys | 5 |
| 2021 | A self-learning teacher-student framework for gastrointestinal image classificationabstractWe present a semi-supervised teacher-student framework to improve classification performance on gastrointestinal image data. As labeled data is scarce in medical settings, this framework is built specifically to take advantage of vast amounts of unlabeled data. It consists of three main steps: (1) train a teacher model with labeled data, (2) use the teacher model to infer pseudo labels with unlabeled data, and (3) train a new and larger student model with a combination of labeled images and inferred pseudo labels. These three steps are repeated several times by treating the student as a teacher to relabel the unlabeled data and consequently train a new student. We demonstrate that our framework can classify both video capsule endoscopy (VCE) and standard endoscopy images. Our results indicate that our teacher-student framework can significantly increase the performance compared to traditional supervised-learning-based models, i.e., an overall increase in the F1-score of 4.7% for the Kvasir-Capsule VCE dataset and 3.2% for the HyperKvasir colonoscopy dataset. We believe that our framework can use more of the data collected at hospitals without the need for expert labels, contributing to overall better models for medical multimedia systems for automatic disease detection. Henrik L. Gjestang, Steven Alexander Hicks, Vajira Thambawita, Pål Halvorsen, Michael Riegler 0001 |
CBMS | 3 |
| 2021 | Automated Clipping of Soccer Events using Machine LearningabstractExtracting highlight clips from soccer matches requires tedious, time-consuming, and expensive manual labor. Human operators need to search for appropriate clipping points and trim away the unwanted scenes. In our work, we aim for an automated process for generating event highlights. In particular, we develop AI-models for scene boundary detection and logo detection. Using different datasets, we present two models that automatically find the appropriate time interval for goal event extraction. The models are evaluated quantitatively, and the results show that we find the logo and scene shifts with high accuracy. Our event clipping methodology is a potential building block for a larger, fully-automated sports broadcast production pipeline. Joakim O. Valand, Haris Kadragic, Steven Alexander Hicks, Vajira Thambawita, Cise Midoglu, Tomas Kupka, Dag Johansen, Michael Riegler 0001, Pål Halvorsen |
ISM | 4 |
| 2021 | Reproducibility Companion Paper: Norm-in-Norm Loss with Faster Convergence and Better Performance for Image Quality AssessmentabstractThis companion paper supports the experimental replication of the paper "Norm-in-Norm Loss with Faster Convergence and Better Performance for Image Quality Assessment'' presented at ACM Multimedia 2020. We provide the software package for replicating the implementation of the "Norm-in-Norm'' loss and the corresponding "LinearityIQA'' model used in the original paper. This paper contains the guidelines to reproduce all the experimental results of the original paper. Dingquan Li, Tingting Jiang 0001, Ming Jiang 0001, Vajira Thambawita |
ACM Multimedia | 4 |
| 2021 | HTAD: A Home-Tasks Activities Dataset with Wrist-Accelerometer and Audio Features
Enrique Garcia-Ceja, Vajira Thambawita, Steven Alexander Hicks, Debesh Jha, Petter Jakobsen, Hugo Hammer, Pål Halvorsen, Michael Riegler 0001 |
MMM (2) | 2 |
| 2021 | Kvasir-Instrument: Diagnostic and Therapeutic Tool Segmentation Dataset in Gastrointestinal Endoscopy
Debesh Jha, Sharib Ali, Krister Emanuelsen, Steven Alexander Hicks, Vajira Thambawita, Enrique Garcia-Ceja, Michael Riegler 0001, Thomas de Lange, Peter Thelin Schmidt, Håvard D. Johansen, Dag Johansen, Pål Halvorsen |
MMM (2) | 5 |
| 2021 | A comprehensive analysis of classification methods in gastrointestinal endoscopy imagingabstractGastrointestinal (GI) endoscopy has been an active field of research motivated by the large number of highly lethal GI cancers. Early GI cancer precursors are often missed during the endoscopic surveillance. The high missed rate of such abnormalities during endoscopy is thus a critical bottleneck. Lack of attentiveness due to tiring procedures, and requirement of training are few contributing factors. An automatic GI disease classification system can help reduce such risks by flagging suspicious frames and lesions. GI endoscopy consists of several multi-organ surveillance, therefore, there is need to develop methods that can generalize to various endoscopic findings. In this realm, we present a comprehensive analysis of the Medico GI challenges: Medical Multimedia Task at MediaEval 2017, Medico Multimedia Task at MediaEval 2018, and BioMedia ACM MM Grand Challenge 2019. These challenges are initiative to set-up a benchmark for different computer vision methods applied to the multi-class endoscopic images and promote to build new approaches that could reliably be used in clinics. We report the performance of 21 participating teams over a period of three consecutive years and provide a detailed analysis of the methods used by the participants, highlighting the challenges and shortcomings of the current approaches and dissect their credibility for the use in clinical settings. Our analysis revealed that the participants achieved an improvement on maximum Mathew correlation coefficient (MCC) from 82.68% in 2017 to 93.98% in 2018 and 95.20% in 2019 challenges, and a significant increase in computational speed over consecutive years. Debesh Jha, Sharib Ali, Steven Alexander Hicks, Vajira Thambawita, Hanna Borgli, Pia H. Smedsrud, Thomas de Lange, Konstantin Pogorelov, Philipp Harzig, Minh-Triet Tran, Wenhua Meng, Trung-Hieu Hoang, Danielle Dias, Tobey H. Ko, Taruna Agrawal, Olga Ostroukhova, Zeshan Khan, Muhammad Atif Tahir, Yang Liu 0007, Mathias Kirkerød, Dag Johansen, Mathias Lux, Håvard D. Johansen, Michael Riegler 0001, Pål Halvorsen |
Medical Image Anal. | 4 |
| 2020 | PSYKOSE: A Motor Activity Database of Patients with SchizophreniaabstractUsing sensor data from devices such as smart-watches or mobile phones is very popular in both computer science and medical research. Such movement data can predict certain health states or performance outcomes. However, in order to increase reliability and replication of the research it is important to share data and results openly. In medicine, this is often difficult due to legal restrictions or to the fact that data collected from clinical trials is seen as very valuable and something that should be kept "in-house". In this paper, we therefore present PSYKOSE, a publicly shared dataset consisting of motor activity data collected from body sensors. The dataset contains data collected from patients with schizophrenia. Schizophrenia is a severe mental disorder characterized by psychotic symptoms like hallucinations and delusions, as well as symptoms of cognitive dysfunction and diminished motivation. In total, we have data from 22 patients with schizophrenia and 32 healthy control persons. For each person in the dataset, we provide sensor data collected over several days in a row. In addition to the sensor data, we also provide some demographic data and medical assessments during the observation period. The patients were assessed by medical experts from Haukeland University hospital. In addition to the data, we provide a baseline analysis and possible use-cases of the dataset. Petter Jakobsen, Enrique Garcia-Ceja, Lena Antonsen Stabell, Ketil J. Oedegaard, Jan Oystein Berle, Vajira Thambawita, Steven Alexander Hicks, Pål Halvorsen, Ole Bernt Fasmer, Michael Riegler 0001 |
CBMS | 6 |
| 2020 | Vid2Pix - A Framework for Generating High-Quality Synthetic VideosabstractData is arguably the most important resource today as it fuels the algorithms powering services we use every day. However, in fields like medicine, publicly available datasets are few, and labeling medical datasets require tedious efforts from trained specialists. Generated synthetic data can be to future successful healthcare clinical intelligence. Here, we present a GAN-based video generator demonstrating promising results. Oda Olsen Nedrejord, Vajira Thambawita, Steven Alexander Hicks, Pål Halvorsen, Michael Riegler 0001 |
ISM | 2 |
| 2020 | Real-Time Detection of Events in Soccer Videos using 3D Convolutional Neural NetworksabstractIn this paper, we present an algorithm for automatically detecting events in soccer videos using 3D convolutional neural networks. The algorithm uses a sliding window approach to scan over a given video to detect events such as goals, yellow/red cards, and player substitutions. We test the method on three different datasets from SoccerNet, the Swedish Allsvenskan, and the Norwegian Eliteserien. Overall, the results show that we can detect events with high recall, low latency, and accurate time estimation. The trade-off is a slightly lower precision compared to the current state-of-the-art, which has higher latency and performs better when a less accurate time estimation can be accepted. In addition to the presented algorithm, we perform an extensive ablation study on how the different parts of the training pipeline affect the final results. Olav A. Norgård Rongved, Steven Alexander Hicks, Vajira Thambawita, Håkon Kvale Stensland, Evi Zouganeli, Dag Johansen, Michael Riegler 0001, Pål Halvorsen |
ISM | 3 |
| 2020 | ACM Multimedia BioMedia 2020 Grand Challenge OverviewabstractThe BioMedia 2020 ACM Multimedia Grand Challenge is the second in a series of competitions focusing on the use of multimedia for different medical use-cases. In this year's challenge, participants are asked to develop algorithms that automatically predict the quality of a given human semen sample using a combination of visual, patient-related, and laboratory-analysis-related data. Compared to last year's challenge, participants are provided with a fully multimodal dataset (videos, analysis data, study participant data) from the field of assisted human reproduction. The tasks encourage the use of the different modalities contained within the dataset and finding smart ways of how they may be combined to further improve prediction accuracy. For example, using only video data or combining video data and patient-related data. The ground truth was developed through a preliminary analysis done by medical experts following the World Health Organization's standard for semen quality assessment. The task lays the basis for automatic, real-time support systems for artificial reproduction. We hope that this challenge motivates multimedia researchers to explore more medical-related applications and use their vast knowledge to make a real impact on people's lives. Steven Alexander Hicks, Vajira Thambawita, Hugo Hammer, Trine B. Haugen, Jorunn M. Andersen, Oliwia Witczak, Pål Halvorsen, Michael Riegler 0001 |
ACM Multimedia | 2 |
| 2020 | Toadstool: a dataset for training emotional intelligent machines playing Super Mario BrosabstractGames are often defined as engines of experience, and they are heavily relying on emotions, they arouse in players. In this paper, we present a dataset called Toadstool as well as a reproducible methodology to extend on the dataset. The dataset consists of video, sensor, and demographic data collected from ten participants playing Super Mario Bros, an iconic and famous video game. The sensor data is collected through an Empatica E4 wristband, which provides high-quality measurements and is graded as a medical device. In addition to the dataset and the methodology for data collection, we present a set of baseline experiments which show that we can use video game frames together with the facial expressions to predict the blood volume pulse of the person playing Super Mario Bros. With the dataset and the collection methodology we aim to contribute to research on emotionally aware machine learning algorithms, focusing on reinforcement learning and multimodal data fusion. We believe that the presented dataset can be interesting for a manifold of researchers to explore exciting new interdisciplinary questions. Henrik Svoren, Vajira Thambawita, Pål Halvorsen, Petter Jakobsen, Enrique Garcia-Ceja, Farzan Majeed Noori, Hugo Hammer, Mathias Lux, Michael Riegler 0001, Steven Alexander Hicks |
MMSys | 2 |
| 2020 | PMData: a sports logging datasetabstractIn this paper, we present PMData: a dataset that combines traditional lifelogging data with sports-activity data. Our dataset enables the development of novel data analysis and machine-learning applications where, for instance, additional sports data is used to predict and analyze everyday developments, like a person's weight and sleep patterns; and applications where traditional lifelog data is used in a sports context to predict athletes' performance. PMData combines input from Fitbit Versa 2 smartwatch wristbands, the PMSys sports logging smartphone application, and Google forms. Logging data has been collected from 16 persons for five months. Our initial experiments show that novel analyses are possible, but there is still room for improvement. Vajira Thambawita, Steven Alexander Hicks, Hanna Borgli, Håkon Kvale Stensland, Debesh Jha, Martin Kristoffer Svensen, Svein Arne Pettersen, Dag Johansen, Håvard D. Johansen, Susann Dahl Pettersen, Simon Nordvang, Sigurd Pedersen, Anders T. Gjerdrum, Tor-Morten Grønli, Per Morten Fredriksen, Ragnhild Eg, Kjeld Hansen, Siri Fagernes, Christine Claudi, Andreas Biørn-Hansen, Duc-Tien Dang-Nguyen, Tomas Kupka, Hugo Hammer, Ramesh Jain 0001, Michael Riegler 0001, Pål Halvorsen |
MMSys | 1 |
| 2020 | An Extensive Study on Cross-Dataset Bias and Evaluation Metrics Interpretation for Machine Learning Applied to Gastrointestinal Tract Abnormality ClassificationabstractPrecise and efficient automated identification of gastrointestinal (GI) tract diseases can help doctors treat more patients and improve the rate of disease detection and identification. Currently, automatic analysis of diseases in the GI tract is a hot topic in both computer science and medical-related journals. Nevertheless, the evaluation of such an automatic analysis is often incomplete or simply wrong. Algorithms are often only tested on small and biased datasets, and cross-dataset evaluations are rarely performed. A clear understanding of evaluation metrics and machine learning models with cross datasets is crucial to bring research in the field to a new quality level. Toward this goal, we present comprehensive evaluations of five distinct machine learning models using global features and deep neural networks that can classify 16 different key types of GI tract conditions, including pathological findings, anatomical landmarks, polyp removal conditions, and normal findings from images captured by common GI tract examination instruments. In our evaluation, we introduce performance hexagons using six performance metrics, such as recall, precision, specificity, accuracy, F1-score, and the Matthews correlation coefficient to demonstrate how to determine the real capabilities of models rather than evaluating them shallowly. Furthermore, we perform cross-dataset evaluations using different datasets for training and testing. With these cross-dataset evaluations, we demonstrate the challenge of actually building a generalizable model that could be used across different hospitals. Our experiments clearly show that more sophisticated performance metrics and evaluation methods need to be applied to get reliable models rather than depending on evaluations of the splits of the same dataset—that is, the performance metrics should always be interpreted together rather than relying on a single metric. Vajira Thambawita, Debesh Jha, Hugo Hammer, Håvard D. Johansen, Dag Johansen, Pål Halvorsen, Michael Riegler 0001 |
ACM Trans. Comput. Heal. | 1 |
| 2019 | GANEx: A complete pipeline of training, inference and benchmarking GAN experimentsabstractDeep learning (DL) is one of the standard methods in the field of multimedia research to perform data classification, detection, segmentation and generation. Within DL, generative adversarial networks (GANs) represents a new and highly popular branch of methods. GANs have the capability to generate, from random noise or conditional input, new data realizations within the dataset population. While generation is popular and highly useful in itself, GANs can also be useful to improve supervised DL. GAN-based approaches can, for example, perform segmentation or create synthetic data for training other DL models. The latter one is especially interesting in domains where not much training data exists such as medical multimedia. In this respect, performing a series of experiments involving GANs can be very time consuming due to the lack of tools that support the whole pipeline such as structured training, testing and tracking of different architectures and configurations. Moreover, the success of generative models is highly dependent on hyper-parameter optimization and statistical analysis in the design and fine-tuning stages. In this paper, we present a new tool called GANEx for making the whole pipeline of training, inference and benchmarking GANs faster, more efficient and more structured. The tool consists of a special library called FastGAN which allows designing generative models very fast. Moreover, GANEx has a graphical user interface to support structured experimenting, quick hyper-parameter configurations and output analysis. The presented tool is not limited to a specific DL framework and can be therefore even used to compare the performance of cross frameworks. Vajira Thambawita, Hugo Hammer, Michael Riegler 0001, Pål Halvorsen |
CBMI | 1 |