EDBT 2026 Demo / reviewers in the wild / expert
Steven Alexander Hicks
dblp:222/1245 · also Steven Hicks
· DBLP profile ↗
39ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-3332-1201ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ImageCLEF 2026: Multimodal Challenges in Medicine, Science, Agritech, and Security
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandra Baicoianu, Ana Neacsu, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Lea Reinartz, Benjamin Lecouteux, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Corneliu Florea, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Hendrik Damm, Henning Schäfer, Ivan Koychev, Josiane Mothe, Liviu-Daniel Stefan, Maja J. Hjuler, Mehmet Kurt, Meliha Yetisgen, Michael Riegler 0001, Mihai Dogariu, Mihai Ivanovici, Ming Shan Hee, Mohammad El Sakka, Momina Ahsan, Obioma Pelka, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Bahadir Eryilmaz, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Yuri Prokopchuk, Zhuohan Xie |
ECIR (4) | 41 |
| 2026 | Low-Dimension Representation Estimation in Principal Component Analysis Under Missing Data
Thanh Tu Do, Van Hua, Uyen Dang, Thu Nguyen 0001, Steven Alexander Hicks, Pål Halvorsen, Michael Riegler 0001, Binh T. Nguyen 0001 |
MMM (2) | 5 |
| 2025 | ImageCLEF 2025: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Helmut Becker, Hendrik Damm, Henning Schäfer, Ivan Rodkin, Ivan Koychev, Johannes Kiesel, Johannes Rückert, Josep Malvehy, Liviu-Daniel Stefan, Louise Bloch, Martin Potthast, Maximilian Heinrich, Michael Riegler 0001, Mihai Dogariu, Noel Codella, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Roberto A. Novoa, Rocktim Jyoti Das, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Zhuohan Xie |
ECIR (5) | 41 |
| 2025 | Evaluating gradient-based explanation methods for neural network ECG analysis using heatmapsabstractOBJECTIVE: Evaluate popular explanation methods using heatmap visualizations to explain the predictions of deep neural networks for electrocardiogram (ECG) analysis and provide recommendations for selection of explanations methods. MATERIALS AND METHODS: A residual deep neural network was trained on ECGs to predict intervals and amplitudes. Nine commonly used explanation methods (Saliency, Deconvolution, Guided backpropagation, Gradient SHAP, SmoothGrad, Input × gradient, DeepLIFT, Integrated gradients, GradCAM) were qualitatively evaluated by medical experts and objectively evaluated using a perturbation-based method. RESULTS: No single explanation method consistently outperformed the other methods, but some methods were clearly inferior. We found considerable disagreement between the human expert evaluation and the objective evaluation by perturbation. DISCUSSION: The best explanation method depended on the ECG measure. To ensure that future explanations of deep neural networks for medical data analyses are useful to medical experts, data scientists developing new explanation methods should collaborate tightly with domain experts. Because there is no explanation method that performs best in all use cases, several methods should be applied. CONCLUSION: Several explanation methods should be used to determine the most suitable approach. Andrea M. Storås, Steffen Mæland, Jonas Isaksen, Steven Alexander Hicks, Vajira Thambawita, Claus Graff, Hugo Hammer, Pål Halvorsen, Michael Riegler 0001, Jørgen K. Kanters |
J. Am. Medical Informatics Assoc. | 4 |
| 2025 | Validating polyp and instrument segmentation methods in colonoscopy through Medico 2020 and MedAI 2021 ChallengesabstractAutomatic analysis of colonoscopy images has been an active field of research motivated by the importance of early detection of precancerous polyps. However, detecting polyps during the live examination can be challenging due to various factors such as variation of skills and experience among the endoscopists, lack of attentiveness, and fatigue leading to a high polyp miss-rate. Therefore, there is a need for an automated system that can flag missed polyps during the examination and improve patient care. Deep learning has emerged as a promising solution to this challenge as it can assist endoscopists in detecting and classifying overlooked polyps and abnormalities in real time, improving the accuracy of diagnosis and enhancing treatment. In addition to the algorithm’s accuracy, transparency and interpretability are crucial to explaining the whys and hows of the algorithm’s prediction. Further, conclusions based on incorrect decisions may be fatal, especially in medicine. Despite these pitfalls, most algorithms are developed in private data, closed source, or proprietary software, and methods lack reproducibility. Therefore, to promote the development of efficient and transparent methods, we have organized the “Medico automatic polyp segmentation (Medico 2020)” and “MedAI: Transparency in Medical Image Segmentation (MedAI 2021)” competitions. The Medico 2020 challenge received submissions from 17 teams, while the MedAI 2021 challenge also gathered submissions from another 17 distinct teams in the following year. We present a comprehensive summary and analyze each contribution, highlight the strength of the best-performing methods, and discuss the possibility of clinical translations of such methods into the clinic. Our analysis revealed that the participants improved dice coefficient metrics from 0.8607 in 2020 to 0.8993 in 2021 despite adding diverse and challenging frames (containing irregular, smaller, sessile, or flat polyps), which are frequently missed during a routine clinical examination. For the instrument segmentation task, the best team obtained a mean Intersection over union metric of 0.9364. For the transparency task, a multi-disciplinary team, including expert gastroenterologists, accessed each submission and evaluated the team based on open-source practices, failure case analysis, ablation studies, usability and understandability of evaluations to gain a deeper understanding of the models’ credibility for clinical deployment. The best team obtained a final transparency score of 21 out of 25. Through the comprehensive analysis of the challenge, we not only highlight the advancements in polyp and surgical instrument segmentation but also encourage subjective evaluation for building more transparent and understandable AI-based colonoscopy systems. Moreover, we discuss the need for multi-center and out-of-distribution testing to address the current limitations of the methods to reduce the cancer burden and improve patient care. • We present a detailed analysis of the Medico 2020 and MedAI 2021 challenges that are aimed at advancing automated polyp and instrument segmentation in colonoscopy for early colorectal cancer diagnosis by using novel deep learning methods. • To the best of our knowledge, MedAI 2021 is the first challenge to evaluate the transparency in both GI endoscopy and colonoscopy. Through the challenge, we invited the participants to list package dependencies and architecture code (with instructions for building, compiling, and training) and share trained model weights in a standardized format. Additionally, we invited participants to include the code for model evaluation and provide repository licensing information to enable others to use the code and the trained model responsibly. Moreover, we asked the participants to explain model predictions using intermediate heatmaps, perform ablation studies, conduct a thorough failure analysis, and share their code for reproducing the results. Finally, we performed a subjective evaluation by including an expert gastroenterologist in the group and gave the final transparency score based on the usefulness and understandability of the results. Our initiative aims to promote transparency in AI research and foster the development of reliable, interpretable, and trustworthy algorithms for use in medical image segmentation. • We provide a comparative analysis of the 34 proposed methods in both challenges (3 subtasks), covering small details of each team in the form of Tables, qualitative and quantitative results (failure analysis), and an in-depth analysis of the findings. • We explore trust, safety, interpretability, transparency, and generalizability issues and provide future strategies to overcome the current limitations of developed algorithms. Debesh Jha, Vanshali Sharma, Debapriya Banik, Debayan Bhattacharya, Kaushiki Roy, Steven Alexander Hicks, Nikhil Kumar Tomar, Vajira Thambawita, Adrian Krenzer, Ge-Peng Ji, Sahadev Poudel, George Batchkala, Saruar Alam, Awadelrahman M. A. Ahmed, Quoc-Huy Trinh, Zeshan Khan, Tien-Phat Nguyen, Shruti Shrestha, Sabari Nathan, Jeonghwan Gwak, Ritika Kumari Jha, Zheyuan Zhang 0001, Alexander Schlaefer, Debotosh Bhattacharjee, Manas Kamal Bhuyan, Pradip K. Das, Deng-Ping Fan, Sravanthi Parasa, Sharib Ali, Michael Riegler 0001, Pål Halvorsen, Thomas de Lange, Ulas Bagci |
Medical Image Anal. | 6 |
| 2024 | SAM in the Pipeline: Transforming Axis-Aligned to Oriented Bounding Boxes for Superior Sperm DetectionabstractSperm analysis, in terms of finding the best cells, has traditionally been performed manually and is slow, expensive, prone to human error, and often inconsistent across different laboratory technicians. Our research focuses on the use of artificial intelligence to improve computer-assisted sperm analysis, a critical aspect in the diagnosis and treatment of male fertility issues. Using the best publicly available dataset for sperm detection, we employ novel data refining techniques and state-of-the-art object detection models to develop the best-performing sperm detection model to our knowledge. We show how the Segment Anything Model (SAM) can be used as a data-refining tool to improve object detection datasets. Additionally, we use SAM to convert Axis-Aligned Bounding Box (AABB) to Oriented Bounding Box (OBB) to get more descriptive coordinates. We test our data-refining pipeline with state-of-the-art object detection models and a publicly available sperm-tracking dataset. Our results show that the refined data coming from our pipeline can be used to train lighter models while still achieving higher accuracy compared to using the original dataset. This also enables possibilities for running high-quality tracking on light hardware, such as consumer-grade laptops. Pål Andreas Hoven Bentsen, Steven Alexander Hicks, Eric Jul, Pål Halvorsen, Vajira Thambawita |
CBMI | 2 |
| 2024 | Advancing Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications with ImageCLEF 2024
Bogdan Ionescu, Henning Müller, Ana-Maria Claudia Dragulinescu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandru Stan, Andrea M. Storås, Asma Ben Abacha, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Didier Schwab, Emmanuelle Esperança-Rodier, George Ioannidis, Griffin Adams, Henning Schäfer, Hugo Manguinhas, Ioan Coman, Johanna Schöler, Johannes Kiesel, Johannes Rückert, Louise Bloch, Martin Potthast, Maximilian Heinrich, Meliha Yetisgen, Michael Riegler 0001, Neal Snider, Pål Halvorsen, Raphael Brüngel, Steven Alexander Hicks, Vajira Thambawita, Vassili Kovalev, Yuri Prokopchuk, Wen-Wai Yim |
ECIR (6) | 34 |
| 2024 | Blockwise Principal Component Analysis for monotone missing data imputation and dimensionality reductionabstractMonotone missing data is a common problem in data analysis. However, imputation combined with dimensionality reduction can be computationally expensive, especially with the increasing size of datasets. We propose a Blockwise Principal Component Analysis Imputation (BPI) framework for dimensionality reduction and imputation of monotone missing data to address this issue. The framework conducts Principal Component Analysis on the observed part of each monotone block of the data and then imputes on merging the obtained principal components using a chosen imputation technique. BPI can work with various imputation techniques and can significantly reduce imputation time compared to conducting dimensionality reduction after imputation. This makes it a practical and efficient approach for large datasets with monotone missing data. Our experiments validate the improvement in speed while achieving an accuracy that is comparable to the common strategy of imputation prior to dimensional reduction. Tu T. Do, Anh M. Vu, Tuan L. Vo, Hoang Thien Ly, Thu Nguyen 0001, Steven Alexander Hicks, Pål Halvorsen, Michael Riegler 0001, Binh T. Nguyen 0001 |
IJCNN | 6 |
| 2023 | RePolyp: A Framework for Generating Realistic Colon Polyps with Corresponding Segmentation Masks using Diffusion ModelsabstractThe field of synthetic medical data has become increasingly important due to the urgent need for large and diverse datasets in the medical sector. Using diffusion models in data generation has created more authentic and varied medical data. In this study, a framework is presented that utilizes diffusion models trained on openly accessible data to generate realistic-looking colon polyps, along with their corresponding ground truth masks. The usefulness of the synthetic polyps is evaluated by using them to train segmentation models designed to segment colon polyps in real-world images. The results demonstrate that the generated synthetic data is highly accurate and suggest that including synthetic polyps in the training dataset improves the predictive performance and generalization of the segmentation models. When the training dataset consists of pre-generated synthetic data from our model, we achieve a mean intersection over union (mIoU) improvement of 4.64% on the validation data and a 4.14% mIoU improvement when testing across different datasets. These results indicate that generating synthetic medical data using diffusion models is valuable for addressing the need for diverse and extensive medical datasets. Alexander K. Pishva, Vajira Thambawita, Jim Tørresen, Steven Alexander Hicks |
CBMS | 4 |
| 2023 | ImageCLEF 2023 Highlight: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Ana-Maria Claudia Dragulinescu, Adrian Popescu 0001, Ahmad Idrissi-Yaghir, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandru Stan, Andrea M. Storås, Asma Ben Abacha, Christoph M. Friedrich, George Ioannidis, Griffin Adams, Henning Schäfer, Hugo Manguinhas, Ihar Filipovich, Ioan Coman, Jérôme Deshayes-Chossart, Johanna Schöler, Johannes Rückert, Liviu-Daniel Stefan, Louise Bloch, Meliha Yetisgen, Michael Riegler 0001, Mihai Dogariu, Mihai Gabriel Constantin, Neal Snider, Nikolaos Papachrysos, Pål Halvorsen, Raphael Brüngel, Serge Kozlovski, Steven Alexander Hicks, Thomas de Lange, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim |
ECIR (3) | 32 |
| 2023 | Multimedia Datasets: Challenges and Future Possibilities
Thu Nguyen 0001, Andrea M. Storås, Vajira Thambawita, Steven Alexander Hicks, Pål Halvorsen, Michael Riegler 0001 |
MMM (2) | 4 |
| 2023 | ScopeSense: An 8.5-Month Sport, Nutrition, and Lifestyle Lifelogging DatasetabstractNowadays, most people have a smartphone that can track their everyday activities. Furthermore, a significant number of people wear advanced smartwatches to track several vital biomarkers in addition to activity data. However, it is still unclear how these data can actually be used to improve certain aspects of people’s lives. One of the key challenges is that the collected data is often massive and unstructured. Therefore, a link to other important information (e.g., when, what, and how much food was consumed) is required. It is widely believed that such detailed and structured longitudinal data about a person is essential to model and provide personalized and precise guidance. Despite the strong belief of researchers about the power of such a data-driven approach, respective datasets have been difficult to collect. In this study, we present a unique dataset from two individuals performing a structured data collection over eight and a half months. In addition to the sensor data, we collected their nutrition, training, and well-being data. The availability of nutrition data with many other important objectives and subjective longitudinal data streams may facilitate research related to food for a healthy lifestyle. Thus, we present a sport, nutrition, and lifestyle logging dataset called ScopeSense from two individuals and discuss its potential use. The dataset is fully open for researchers, and we consider this study as a potential starting point for developing methods to collect and create knowledge for a larger cohort of people. Michael Riegler 0001, Vajira Thambawita, Binh T. Nguyen 0001, Steven Alexander Hicks, Vibeke Telle-Hansen, Svein Arne Pettersen, Dag Johansen, Ramesh Jain 0001, Pål Halvorsen |
MMM (1) | 5 |
| 2022 | Estimating Predictive Uncertainty in Gastrointestinal Polyp SegmentationabstractDeep neural networks have achieved state-of-the-art performance on numerous applications in the medical field, with use-cases ranging from automation of mundane tasks to diagnosis of life-threatening diseases. Despite these achievements, deep neural networks are considered “black boxes” due to their complex structure and general lack of transparency in their decision-making process. These attributes make it challenging to incorporate deep learning into existing clinical workflows as decisions often need more support than blind faith in a statistical model. This paper presents an investigation of uncertainty estimation for the detection of colon polyps using deep convolutional neural networks (CNNs). We experiment with two different approaches to measure uncertainty, Monte Carlo (MC) dropout and deep ensembles, and discuss the advantages and disadvantages of both methods in terms of computational efficiency and performance gain. Furthermore, we apply the two uncertainty methods to two different state-of-the-art CNN-based polyp segmentation architectures. The uncertainty is visualized as heatmaps on the input images and can be used to make more informed decisions on whether or not to trust a model's predictions. The results show that the predictive uncertainties provide a comparison between different models' predictions which can be interpreted as contrastive explanations where the values are largely influenced by the degree of independence between the models in the ensemble. We also reveal that MC dropout is shown to lack at providing contrastive uncertainty values due to the high correlation between the models' in the ensemble. Felicia Ly Jacobsen, Steven Alexander Hicks, Pål Halvorsen, Michael Riegler 0001 |
CBMS | 2 |
| 2022 | Experiences and Lessons Learned from a Crowdsourced-Remote Hybrid User Survey FrameworkabstractSubjective user studies are important to ensure the fidelity and usability of systems that generate multimedia content. Testing how end-users and domain experts perceive multimedia assets might provide crucial information. In this paper, we present our experiences with the open source hybrid crowdsourced-remote user survey framework called Huldra, which is intended for conducting web-based subjective user studies and aims to integrate the individual benefits associated with traditional, crowdsourced, and remote methods. We disseminate our experiences and insights from two actively deployed use cases and discuss challenges and opportunities associated with using Huldra as a framework for conducting user studies. Cise Midoglu, Andrea M. Storås, Saeed Shafiee Sabet, Malek Hammou, Steven Alexander Hicks, Inga Strümke, Michael Riegler 0001, Carsten Griwodz, Pål Halvorsen |
ISM | 5 |
| 2022 | RCAD: Real-time Collaborative Anomaly Detection System for Mobile Broadband NetworksabstractThe rapid increase in mobile data traffic and the number of connected devices and applications in networks is putting a significant pressure on the current network management approaches that heavily rely on human operators. Consequently, an automated network management system that can efficiently predict and detect anomalies is needed. In this paper, we propose, RCAD, a novel distributed architecture for detecting anomalies in network data forwarding latency in an unsupervised fashion. RCAD employs the hierarchical temporal memory (HTM) algorithm for the online detection of anomalies. It also involves a collaborative distributed learning module that facilitates knowledge sharing across the system. We implement and evaluate RCAD on real world measurements from a commercial mobile network. RCAD achieves over 0.7 F-1 score significantly outperforming current state-of-the-art methods. Azza H. Ahmed, Michael Riegler 0001, Steven Alexander Hicks, Ahmed Elmokashfi |
KDD | 3 |
| 2022 | Huldra: a framework for collecting crowdsourced feedback on multimedia assetsabstractCollecting crowdsourced feedback to evaluate, rank, or score multimedia content can be cumbersome and time-consuming. Most of the existing survey tools are complicated, hard to customize, or tailored for a specific asset type. In this paper, we present an open source framework called Huldra, designed explicitly to address the challenges associated with user studies involving crowdsourced feedback collection. The web-based framework is built in a modular and configurable fashion to allow for the easy adjustment of the user interface (UI) and the multimedia content, while providing integrations with reliable and stable backend solutions to facilitate the collection and analysis of responses. Our proposed framework can be used as an online survey tool by researchers working on different topics such as Machine Learning (ML), audio, image, and video quality assessment, Quality of Experience (QoE), and require user studies for the benchmarking of various types of multimedia content. Malek Hammou, Cise Midoglu, Steven Alexander Hicks, Andrea M. Storås, Saeed Shafiee Sabet, Inga Strümke, Michael Riegler 0001, Pål Halvorsen |
MMSys | 3 |
| 2022 | Automatic thumbnail selection for soccer videos using machine learningabstractThumbnail selection is a very important aspect of online sport video presentation, as thumbnails capture the essence of important events, engage viewers, and make video clips attractive to watch. Traditional solutions in the soccer domain for presenting highlight clips of important events such as goals, substitutions, and cards rely on the manual or static selection of thumbnails. However, such approaches can result in the selection of sub-optimal video frames as snapshots, which degrades the overall quality of the video clip as perceived by viewers, and consequently decreases viewership, not to mention that manual processes are expensive and time consuming. In this paper, we present an automatic thumbnail selection system for soccer videos which uses machine learning to deliver representative thumbnails with high relevance to video content and high visual quality in near real-time. Our proposed system combines a software framework which integrates logo detection, close-up shot detection, face detection, and image quality analysis into a modular and customizable pipeline, and a subjective evaluation framework for the evaluation of results. We evaluate our proposed pipeline quantitatively using various soccer datasets, in terms of complexity, runtime, and adherence to a pre-defined rule-set, as well as qualitatively through a user study, in terms of the perception of output thumbnails by end-users. Our results show that an automatic end-to-end system for the selection of thumbnails based on contextual relevance and visual quality can yield attractive highlight clips, and can be used in conjunction with existing soccer broadcast pipelines which require real-time operation. Andreas Husa, Cise Midoglu, Malek Hammou, Steven Alexander Hicks, Dag Johansen, Tomas Kupka, Michael Riegler 0001, Pål Halvorsen |
MMSys | 4 |
| 2022 | Njord: a fishing trawler datasetabstractFish is one of the main sources of food worldwide. The commercial fishing industry has a lot of different aspects to consider, ranging from sustainability to reporting. The complexity of the domain also attracts a lot of research from different fields like marine biology, fishery sciences, cybernetics, and computer science. In computer science, detection of fishing vessels via for example remote sensing and classification of fish from images or videos using machine learning or other analysis methods attracts growing attention. Surprisingly, little work has been done that considers what is happening on board the fishing vessels. On the deck of the boats, a lot of data and important information are generated with potential applications, such as automatic detection of accidents or automatic reporting of fish caught. This paper presents Njord, a fishing trawler dataset consisting of surveillance videos from a modern off-shore fishing trawler at sea. The main goal of this dataset is to show the potential and possibilities that analysis of such data can provide. In addition to the data, we provide a baseline analysis and discuss several possible research questions this dataset could help answer. Tor-Arne S. Nordmo, Aril B. Ovesen, Bjørn Aslak Juliussen, Steven Alexander Hicks, Vajira Thambawita, Håvard D. Johansen, Pål Halvorsen, Michael Riegler 0001, Dag Johansen |
MMSys | 4 |
| 2021 | A self-learning teacher-student framework for gastrointestinal image classificationabstractWe present a semi-supervised teacher-student framework to improve classification performance on gastrointestinal image data. As labeled data is scarce in medical settings, this framework is built specifically to take advantage of vast amounts of unlabeled data. It consists of three main steps: (1) train a teacher model with labeled data, (2) use the teacher model to infer pseudo labels with unlabeled data, and (3) train a new and larger student model with a combination of labeled images and inferred pseudo labels. These three steps are repeated several times by treating the student as a teacher to relabel the unlabeled data and consequently train a new student. We demonstrate that our framework can classify both video capsule endoscopy (VCE) and standard endoscopy images. Our results indicate that our teacher-student framework can significantly increase the performance compared to traditional supervised-learning-based models, i.e., an overall increase in the F1-score of 4.7% for the Kvasir-Capsule VCE dataset and 3.2% for the HyperKvasir colonoscopy dataset. We believe that our framework can use more of the data collected at hospitals without the need for expert labels, contributing to overall better models for medical multimedia systems for automatic disease detection. Henrik L. Gjestang, Steven Alexander Hicks, Vajira Thambawita, Pål Halvorsen, Michael Riegler 0001 |
CBMS | 2 |
| 2021 | Automated Clipping of Soccer Events using Machine LearningabstractExtracting highlight clips from soccer matches requires tedious, time-consuming, and expensive manual labor. Human operators need to search for appropriate clipping points and trim away the unwanted scenes. In our work, we aim for an automated process for generating event highlights. In particular, we develop AI-models for scene boundary detection and logo detection. Using different datasets, we present two models that automatically find the appropriate time interval for goal event extraction. The models are evaluated quantitatively, and the results show that we find the logo and scene shifts with high accuracy. Our event clipping methodology is a potential building block for a larger, fully-automated sports broadcast production pipeline. Joakim O. Valand, Haris Kadragic, Steven Alexander Hicks, Vajira Thambawita, Cise Midoglu, Tomas Kupka, Dag Johansen, Michael Riegler 0001, Pål Halvorsen |
ISM | 3 |
| 2021 | Reproducibility Companion Paper: Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial FeaturesabstractBlind natural video quality assessment (BVQA), also known as no-reference video quality assessment, is a highly active research topic. In our recent contribution titled "Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial Features" published in ACM Multimedia 2020, we proposed a two-level video quality model employing statistical temporal features and spatial features extracted by a deep convolutional neural network (CNN) for this purpose. At the time of publishing, the proposed model (CNN-TLVQM) achieved state-of-the-art results in BVQA. In this paper, we describe the process of reproducing the published results by using CNN-TLVQM on two publicly available natural video quality datasets. Jari Korhonen, Yicheng Su, Junyong You, Steven Alexander Hicks, Cise Midoglu |
ACM Multimedia | 4 |
| 2021 | Reproducibility Companion Paper: Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkabstractIn this companion paper, we provide details of the artifacts to support the replication of "Self-supervised Video Representation Learning Using Inter-intra Contrastive Framework", which was presented at MM'20. The Inter-intra Contrastive (IIC) framework aims to extract more discriminative temporal information by extending intra-negative samples in contrastive self-supervised learning. In this paper, we first summarize our contribution. Then we explain the file structure of the source code and detailed settings. Since our proposal is a framework which contain a lot of different settings, we provide some custom settings to help other researchers to use our methods easily. The source code is available at https://github.com/BestJuly/IIC. Toshihiko Yamasaki, Jingjing Chen 0001, Steven Alexander Hicks |
ACM Multimedia | 5 |
| 2021 | HTAD: A Home-Tasks Activities Dataset with Wrist-Accelerometer and Audio Features
Enrique Garcia-Ceja, Vajira Thambawita, Steven Alexander Hicks, Debesh Jha, Petter Jakobsen, Hugo Hammer, Pål Halvorsen, Michael Riegler 0001 |
MMM (2) | 3 |
| 2021 | Kvasir-Instrument: Diagnostic and Therapeutic Tool Segmentation Dataset in Gastrointestinal Endoscopy
Debesh Jha, Sharib Ali, Krister Emanuelsen, Steven Alexander Hicks, Vajira Thambawita, Enrique Garcia-Ceja, Michael Riegler 0001, Thomas de Lange, Peter Thelin Schmidt, Håvard D. Johansen, Dag Johansen, Pål Halvorsen |
MMM (2) | 4 |
| 2021 | HYPERAKTIV: An Activity Dataset from Patients with Attention-Deficit/Hyperactivity Disorder (ADHD)abstractMachine learning research within healthcare frequently lacks the public data needed to be fully reproducible and comparable. Datasets are often restricted due to privacy concerns and legal requirements that come with patient-related data. Consequentially, many algorithms and models get published on the same topic without a standard benchmark to measure against. Therefore, this paper presents HYPERAKTIV, a public dataset containing health, activity, and heart rate data from patients diagnosed with attention deficit hyperactivity disorder, better known as ADHD. The dataset consists of data collected from 51 patients with ADHD and 52 clinical controls. In addition to the activity and heart rate data, we also include a series of patient attributes such as their age, sex, and information about their mental state, as well as output data from a computerized neuropsychological test. Together with the presented dataset, we also provide baseline experiments using traditional machine learning algorithms to predict ADHD based on the included activity data. We hope that this dataset can be used as a starting point for computer scientists who want to contribute to the field of mental health, and as a common benchmark for future work in ADHD analysis. Steven Alexander Hicks, Andrea Stautland, Ole Bernt Fasmer, Wenche Førland, Hugo Hammer, Pål Halvorsen, Kristin Mjeldheim, Ketil J. Oedegaard, Berge Osnes, Vigdis Elin Giæver Syrstad, Michael Riegler 0001, Petter Jakobsen |
MMSys | 1 |
| 2021 | A comprehensive analysis of classification methods in gastrointestinal endoscopy imagingabstractGastrointestinal (GI) endoscopy has been an active field of research motivated by the large number of highly lethal GI cancers. Early GI cancer precursors are often missed during the endoscopic surveillance. The high missed rate of such abnormalities during endoscopy is thus a critical bottleneck. Lack of attentiveness due to tiring procedures, and requirement of training are few contributing factors. An automatic GI disease classification system can help reduce such risks by flagging suspicious frames and lesions. GI endoscopy consists of several multi-organ surveillance, therefore, there is need to develop methods that can generalize to various endoscopic findings. In this realm, we present a comprehensive analysis of the Medico GI challenges: Medical Multimedia Task at MediaEval 2017, Medico Multimedia Task at MediaEval 2018, and BioMedia ACM MM Grand Challenge 2019. These challenges are initiative to set-up a benchmark for different computer vision methods applied to the multi-class endoscopic images and promote to build new approaches that could reliably be used in clinics. We report the performance of 21 participating teams over a period of three consecutive years and provide a detailed analysis of the methods used by the participants, highlighting the challenges and shortcomings of the current approaches and dissect their credibility for the use in clinical settings. Our analysis revealed that the participants achieved an improvement on maximum Mathew correlation coefficient (MCC) from 82.68% in 2017 to 93.98% in 2018 and 95.20% in 2019 challenges, and a significant increase in computational speed over consecutive years. Debesh Jha, Sharib Ali, Steven Alexander Hicks, Vajira Thambawita, Hanna Borgli, Pia H. Smedsrud, Thomas de Lange, Konstantin Pogorelov, Philipp Harzig, Minh-Triet Tran, Wenhua Meng, Trung-Hieu Hoang, Danielle Dias, Tobey H. Ko, Taruna Agrawal, Olga Ostroukhova, Zeshan Khan, Muhammad Atif Tahir, Yang Liu 0007, Mathias Kirkerød, Dag Johansen, Mathias Lux, Håvard D. Johansen, Michael Riegler 0001, Pål Halvorsen |
Medical Image Anal. | 3 |
| 2020 | PSYKOSE: A Motor Activity Database of Patients with SchizophreniaabstractUsing sensor data from devices such as smart-watches or mobile phones is very popular in both computer science and medical research. Such movement data can predict certain health states or performance outcomes. However, in order to increase reliability and replication of the research it is important to share data and results openly. In medicine, this is often difficult due to legal restrictions or to the fact that data collected from clinical trials is seen as very valuable and something that should be kept "in-house". In this paper, we therefore present PSYKOSE, a publicly shared dataset consisting of motor activity data collected from body sensors. The dataset contains data collected from patients with schizophrenia. Schizophrenia is a severe mental disorder characterized by psychotic symptoms like hallucinations and delusions, as well as symptoms of cognitive dysfunction and diminished motivation. In total, we have data from 22 patients with schizophrenia and 32 healthy control persons. For each person in the dataset, we provide sensor data collected over several days in a row. In addition to the sensor data, we also provide some demographic data and medical assessments during the observation period. The patients were assessed by medical experts from Haukeland University hospital. In addition to the data, we provide a baseline analysis and possible use-cases of the dataset. Petter Jakobsen, Enrique Garcia-Ceja, Lena Antonsen Stabell, Ketil J. Oedegaard, Jan Oystein Berle, Vajira Thambawita, Steven Alexander Hicks, Pål Halvorsen, Ole Bernt Fasmer, Michael Riegler 0001 |
CBMS | 7 |
| 2020 | Vid2Pix - A Framework for Generating High-Quality Synthetic VideosabstractData is arguably the most important resource today as it fuels the algorithms powering services we use every day. However, in fields like medicine, publicly available datasets are few, and labeling medical datasets require tedious efforts from trained specialists. Generated synthetic data can be to future successful healthcare clinical intelligence. Here, we present a GAN-based video generator demonstrating promising results. Oda Olsen Nedrejord, Vajira Thambawita, Steven Alexander Hicks, Pål Halvorsen, Michael Riegler 0001 |
ISM | 3 |
| 2020 | Real-Time Detection of Events in Soccer Videos using 3D Convolutional Neural NetworksabstractIn this paper, we present an algorithm for automatically detecting events in soccer videos using 3D convolutional neural networks. The algorithm uses a sliding window approach to scan over a given video to detect events such as goals, yellow/red cards, and player substitutions. We test the method on three different datasets from SoccerNet, the Swedish Allsvenskan, and the Norwegian Eliteserien. Overall, the results show that we can detect events with high recall, low latency, and accurate time estimation. The trade-off is a slightly lower precision compared to the current state-of-the-art, which has higher latency and performs better when a less accurate time estimation can be accepted. In addition to the presented algorithm, we perform an extensive ablation study on how the different parts of the training pipeline affect the final results. Olav A. Norgård Rongved, Steven Alexander Hicks, Vajira Thambawita, Håkon Kvale Stensland, Evi Zouganeli, Dag Johansen, Michael Riegler 0001, Pål Halvorsen |
ISM | 2 |
| 2020 | ACM Multimedia BioMedia 2020 Grand Challenge OverviewabstractThe BioMedia 2020 ACM Multimedia Grand Challenge is the second in a series of competitions focusing on the use of multimedia for different medical use-cases. In this year's challenge, participants are asked to develop algorithms that automatically predict the quality of a given human semen sample using a combination of visual, patient-related, and laboratory-analysis-related data. Compared to last year's challenge, participants are provided with a fully multimodal dataset (videos, analysis data, study participant data) from the field of assisted human reproduction. The tasks encourage the use of the different modalities contained within the dataset and finding smart ways of how they may be combined to further improve prediction accuracy. For example, using only video data or combining video data and patient-related data. The ground truth was developed through a preliminary analysis done by medical experts following the World Health Organization's standard for semen quality assessment. The task lays the basis for automatic, real-time support systems for artificial reproduction. We hope that this challenge motivates multimedia researchers to explore more medical-related applications and use their vast knowledge to make a real impact on people's lives. Steven Alexander Hicks, Vajira Thambawita, Hugo Hammer, Trine B. Haugen, Jorunn M. Andersen, Oliwia Witczak, Pål Halvorsen, Michael Riegler 0001 |
ACM Multimedia | 1 |
| 2020 | Toadstool: a dataset for training emotional intelligent machines playing Super Mario BrosabstractGames are often defined as engines of experience, and they are heavily relying on emotions, they arouse in players. In this paper, we present a dataset called Toadstool as well as a reproducible methodology to extend on the dataset. The dataset consists of video, sensor, and demographic data collected from ten participants playing Super Mario Bros, an iconic and famous video game. The sensor data is collected through an Empatica E4 wristband, which provides high-quality measurements and is graded as a medical device. In addition to the dataset and the methodology for data collection, we present a set of baseline experiments which show that we can use video game frames together with the facial expressions to predict the blood volume pulse of the person playing Super Mario Bros. With the dataset and the collection methodology we aim to contribute to research on emotionally aware machine learning algorithms, focusing on reinforcement learning and multimodal data fusion. We believe that the presented dataset can be interesting for a manifold of researchers to explore exciting new interdisciplinary questions. Henrik Svoren, Vajira Thambawita, Pål Halvorsen, Petter Jakobsen, Enrique Garcia-Ceja, Farzan Majeed Noori, Hugo Hammer, Mathias Lux, Michael Riegler 0001, Steven Alexander Hicks |
MMSys | 10 |
| 2020 | PMData: a sports logging datasetabstractIn this paper, we present PMData: a dataset that combines traditional lifelogging data with sports-activity data. Our dataset enables the development of novel data analysis and machine-learning applications where, for instance, additional sports data is used to predict and analyze everyday developments, like a person's weight and sleep patterns; and applications where traditional lifelog data is used in a sports context to predict athletes' performance. PMData combines input from Fitbit Versa 2 smartwatch wristbands, the PMSys sports logging smartphone application, and Google forms. Logging data has been collected from 16 persons for five months. Our initial experiments show that novel analyses are possible, but there is still room for improvement. Vajira Thambawita, Steven Alexander Hicks, Hanna Borgli, Håkon Kvale Stensland, Debesh Jha, Martin Kristoffer Svensen, Svein Arne Pettersen, Dag Johansen, Håvard D. Johansen, Susann Dahl Pettersen, Simon Nordvang, Sigurd Pedersen, Anders T. Gjerdrum, Tor-Morten Grønli, Per Morten Fredriksen, Ragnhild Eg, Kjeld Hansen, Siri Fagernes, Christine Claudi, Andreas Biørn-Hansen, Duc-Tien Dang-Nguyen, Tomas Kupka, Hugo Hammer, Ramesh Jain 0001, Michael Riegler 0001, Pål Halvorsen |
MMSys | 2 |
| 2019 | Semantic Analysis of Soccer News for Automatic Game Event ClassificationabstractWe are today overwhelmed with information, of which an important part is news. Sports news, in particular, has become very popular, where soccer makes up a big part of this coverage. For sports fans, it can be a time consuming and tedious to keep up with the news that they really care about. In this paper, we present different machine learning methods applied to soccer news from a Norwegian newspaper and a TV station's news site to summarize the content in a short and digestible manner. We present a system to collect, index, label, analyze, and present the collected news articles based on the content. We perform a thorough comparison between deep learning and traditional machine learning algorithms on text classification. Furthermore, we present a dataset of soccer news which was collected from two different Norwegian news sites and shared online. Aanund Jupskås Nordskog, Pål Halvorsen, Steven Alexander Hicks, Håkon Kvale Stensland, Hugo Hammer, Dag Johansen, Michael Riegler 0001 |
CBMI | 3 |
| 2019 | A Web-Based Software for Training and Quality Assessment in the Image Analysis Workflow for Cardiac T1 Mapping MRIabstractMedical practice makes significant use of imaging scans such as Ultrasound or Magnetic Resonance Imaging as a diagnostic tool. They are used in the visual inspection or quantification of medical parameters computed from the images in post-processing. However, the value of such parameters depends much on the user's variability, device, and algorithmic differences. In this paper, we focus on quantifying the variability due to the human factor, which can be primarily addressed by the structured training of a human operator. We focus on a specific emerging cardiovascular MRI methodology, the T1 mapping, that has proven useful to identify a range of pathological alterations of the myocardial tissue structure. Training, especially in emerging techniques, is typically not standardized, varying dramatically across medical centers and research teams. Additionally, training assessment is mostly based on qualitative approaches. Our work aims to provide a software tool combining traditional clinical metrics and convolutional neural networks to aid the training process by gathering contours from multiple trainees, quantifying discrepancy from local gold standard or standardized guidelines, classifying trainees output based on critical parameters that affect contours variability. Edvarda Eriksen, Steven Alexander Hicks, Michael Riegler 0001, Pål Halvorsen, Valentina Carapella |
ISM | 2 |
| 2019 | ACM Multimedia BioMedia 2019 Grand Challenge OverviewabstractThe BioMedia 2019 ACM Multimedia Grand Challenge is the first in a series of competitions focusing on the use of multimedia for different medical use-cases. In this year's challenge, the participants are asked to develop efficient algorithms which automatically detect a variety of findings commonly identified in the gastrointestinal (GI) tract (a part of the human digestive system). The purpose of this task is to develop methods to aid medical doctors performing routine endoscopy inspections of the GI tract. In this paper, we give a detailed description of the four different tasks of this year's challenge, present the datasets used for training and testing, and discuss how each submission is evaluated both qualitatively and quantitatively. Steven Alexander Hicks, Michael Riegler 0001, Pia H. Smedsrud, Trine B. Haugen, Kristin Ranheim Randel, Konstantin Pogorelov, Håkon Kvale Stensland, Duc-Tien Dang-Nguyen, Mathias Lux, Andreas Petlund, Thomas de Lange, Peter Thelin Schmidt, Pål Halvorsen |
ACM Multimedia | 1 |
| 2019 | VISEM: a multimodal video dataset of human spermatozoaabstractReal multimedia datasets that contain more than just images or text are rare. Even more so are open multimedia datasets in medicine. Often, clinically related datasets only consist of image or videos. In this paper, we present a dataset that is novel in two ways. Firstly, it is a multi-modal dataset containing different data sources such as videos, biological analysis data, and participant data. Secondly, it is the first dataset of that kind in the field of human reproduction. It consists of anonymized data from 85 different participants. We hope this dataset paper will inspire people to apply their knowledge in this important field, generate shareable results in the domain, and ultimately improve human infertility investigation and treatment. Trine B. Haugen, Steven Alexander Hicks, Jorunn M. Andersen, Oliwia Witczak, Hugo Hammer, Rune Johan Borgli, Pål Halvorsen, Michael Riegler 0001 |
MMSys | 2 |
| 2018 | Dissecting Deep Neural Networks for Better Medical Image Classification and Classification UnderstandingabstractNeural networks, in the context of deep learning, show much promise in becoming an important tool with the purpose assisting medical doctors in disease detection during patient examinations. However, the current state of deep learning is something of a "black box", making it very difficult to understand what internal processes lead to a given result. This is not only true for non-technical users but among experts as well. This lack of understanding has led to hesitation in the implementation of these methods among mission-critical fields, with many putting interpretability in front of actual performance. Motivated by increasing the acceptance and trust of these methods, and to make qualified decisions, we present a system that allows for the partial opening of this black box. This includes an investigation on what the neural network sees when making a prediction, to both, improve algorithmic understanding, and to gain intuition into what pre-processing steps may lead to better image classification performance. Furthermore, a significant part of a medical expert's time is spent preparing reports after medical examinations, and if we already have a system for dissecting the analysis done by the network, the same tool can be used for automatic examination documentation through content suggestions. In this paper, we present a system that can look into the layers of a deep neural network and present the network's decision in a way that that medical doctors may understand. Furthermore, we present and discuss how this information can possibly be used for automatic reporting. Our initial results are very promising. Steven Alexander Hicks, Michael Riegler 0001, Konstantin Pogorelov, Kim V. Anonsen, Thomas de Lange, Dag Johansen, Mattis Jeppsson, Kristin Ranheim Randel, Sigrun Losada Eskeland, Pål Halvorsen |
CBMS | 1 |
| 2018 | Mimir: an automatic reporting and reasoning system for deep learning based analysis in the medical domainabstractAutomatic detection of diseases is a growing field of interest, and machine learning in form of deep learning neural networks are frequently explored as a potential tool for the medical video analysis. To both improve the "black box"-understanding and assist in the administrative duties of writing an examination report, we release an automated multimedia reporting software dissecting the neural network to learn the intermediate analysis steps, i.e., we are adding a new level of understanding and explainability by looking into the deep learning algorithms decision processes. The presented open-source software can be used for easy retrieval and reuse of data for automatic report generation, comparisons, teaching and research. As an example, we use live colonoscopy as a use case which is the gold standard examination of the large bowel, commonly performed for clinical and screening purposes. The added information has potentially a large value, and reuse of the data for the automatic reporting may potentially save the doctors large amounts of time. Steven Alexander Hicks, Sigrun Losada Eskeland, Mathias Lux, Thomas de Lange, Kristin Ranheim Randel, Mattis Jeppsson, Konstantin Pogorelov, Pål Halvorsen, Michael Riegler 0001 |
MMSys | 1 |
| 2018 | Comprehensible reasoning and automated reporting of medical examinations based on deep learning analysisabstractIn the future, medical doctors will to an increasing degree be assisted by deep learning neural networks for disease detection during examinations of patients. In order to make qualified decisions, the black box of deep learning must be opened to increase the understanding of the reasoning behind the decision of the machine learning system. Furthermore, preparing reports after the examinations is a significant part of a doctors work-day, but if we already have a system dissecting the neural network for understanding, the same tool can be used for automatic report generation. In this demo, we describe a system that analyses medical videos from the gastrointestinal tract. Our system dissects the Tensorflow-based neural network to provide insights into the analysis and uses the resulting classification and rationale behind the classification to automatically generate an examination report for the patient's medical journal. Steven Alexander Hicks, Konstantin Pogorelov, Thomas de Lange, Mathias Lux, Mattis Jeppsson, Kristin Ranheim Randel, Sigrun Losada Eskeland, Pål Halvorsen, Michael Riegler 0001 |
MMSys | 1 |