VLDB 2026 Research / reviewers in the wild / expert
Hugo Hammer
dblp:08/8429 · also Hugo Lewi Hammer
· DBLP profile ↗
48ranked-venue papers
15as first author
17since 2021 · last 2026
0000-0001-9429-7148ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 11 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Calliope: A TTS-based Narrated E-book Creator Ensuring Exact Synchronization, Privacy, and Layout FidelityabstractA narrated e-book combines synchronized audio with digital text, highlighting the currently spoken word or sentence during playback. This format supports early literacy and assists individuals with reading challenges, while also allowing general readers to seamlessly switch between reading and listening. With the emergence of natural-sounding neural Text-to-Speech (TTS) technology, several commercial services have been developed to leverage these technology for converting standard text e-books into high-quality narrated e-books. However, no open-source solutions currently exist to perform this task. In this paper, we present Calliope, an open-source framework designed to fill this gap. Our method leverages state-of-the-art open-source TTS to convert a text e-book into a narrated e-book in the EPUB 3 Media Overlay format. The method offers several innovative steps: audio timestamps are captured directly during TTS, ensuring exact synchronization between narration and text highlighting; the publisher's original typography, styling, and embedded media are strictly preserved; and the entire pipeline operates offline. This offline capability eliminates recurring API costs, mitigates privacy concerns, and avoids copyright compliance issues associated with cloud-based services. The framework currently supports the state-of-the-art open-source TTS systems XTTS-v2 and Chatterbox. A potential alternative approach involves first generating narration via TTS and subsequently synchronizing it with the text using forced alignment. However, while our method ensures exact synchronization, our experiments show that forced alignment introduces drift between the audio and text highlighting significant enough to degrade the reading experience. Source code and usage instructions are available at https://github.com/hugohammer/TTS-Narrated-Ebook-Creator.git. Hugo Hammer, Vajira Thambawita, Pål Halvorsen |
MMSys | 1 |
| 2026 | Uncertainty in deep learning for EEG under dataset shiftsabstractAs artificial intelligence (AI) is increasingly integrated into medical diagnostics, it is essential that predictive models provide not only accurate outputs but also reliable estimates of uncertainty. In clinical applications, where decisions have significant consequences, understanding the confidence behind each prediction is as critical as the prediction itself. Uncertainty modelling plays a key role in improving trust, guiding decision-making, and identifying unreliable outputs, particularly under dataset shift or in out-of-distribution settings. The primary aim of uncertainty metrics is to align model confidence closely with actual predictive performance, ensuring confidence estimates dynamically adjust to reflect increasing errors or decreasing reliability of predictions. This study investigates how different ensemble learning strategies affect both performance and uncertainty estimation in a clinically relevant task: classifying Normal, Mild Cognitive Impairment, and Dementia from electroencephalography (EEG) data. We evaluated the performance and uncertainty of ensemble methods and Monte Carlo dropout on a large EEG dataset. The models were assessed in three settings: (1) in-distribution performance on a held-out test set, (2) generalisation to three out-of-distribution datasets, and (3) performance under gradual, EEG-specific dataset shifts simulating noise, drift, and frequency perturbation. Ensembles consisting of multiple independently trained models, such as deep ensembles, consistently achieved higher performance in both the in-distribution test set and the out-of-distribution datasets. These models also produced more informative and reliable uncertainty estimates under various types of EEG dataset shifts. These results highlight the benefits of ensemble diversity and independent training to build robust and uncertainty-aware EEG classification models. The findings are particularly relevant for clinical applications, where reliability under distribution shift and transparent uncertainty are essential for safe deployment. Mats Tveter, Thomas Tveitstøl, Christoffer Hatlestad-Hall, Hugo Hammer, Ira R. J. Hebold Haraldsen |
Artif. Intell. Medicine | 4 |
| 2025 | Explanation Supported Learning: Improving Prediction Performance with Explainable Artificial IntelligenceabstractWhen artificial intelligence (AI) and machine learning (ML) models are applied in healthcare, the ability to understand and explain model decisions is an important aspect. Methods in the field of explainable AI (XAI) have been developed to create explanations for such decisions, which provides transparency and trust to the prediction model. However, the use of XAI-based explanations as added data features for the purpose of improving prediction performance remains a little explored topic. Our proposed Explanation Supported Learning (XSL) framework can improve classification performance for ML models used in medical imaging systems, while also providing a new understanding of how medical images are processed by deep learning (DL) models. The XSL framework consists of novel methods to achieve knowledge transfer from one or several teacher models to a student model. The novelty lies in using explanations from the teacher models, obtained from XAI techniques, as added features when training the student model. This approach enables flexible knowledge transfer between models of different architecture types. We further demonstrate how the XSL framework can be used as a new metric for measuring the quality of the explanations provided by XAI methods. The achievement of increased performance in this framework requires that the chosen XAI technique contains useful information based on the learned understanding of the input data by the teacher models. By testing XSL on the HyperKvasir gastrointestinal image dataset, we achieved significant increases in most of the measured classification metrics, and exceeded most benchmark scores of the HyperKvasir paper. Our code is available on GitHub. Adrian Duric, Jim Tørresen, Michael Riegler 0001, Hugo Hammer |
CBMS | 4 |
| 2025 | Evaluating gradient-based explanation methods for neural network ECG analysis using heatmapsabstractOBJECTIVE: Evaluate popular explanation methods using heatmap visualizations to explain the predictions of deep neural networks for electrocardiogram (ECG) analysis and provide recommendations for selection of explanations methods. MATERIALS AND METHODS: A residual deep neural network was trained on ECGs to predict intervals and amplitudes. Nine commonly used explanation methods (Saliency, Deconvolution, Guided backpropagation, Gradient SHAP, SmoothGrad, Input × gradient, DeepLIFT, Integrated gradients, GradCAM) were qualitatively evaluated by medical experts and objectively evaluated using a perturbation-based method. RESULTS: No single explanation method consistently outperformed the other methods, but some methods were clearly inferior. We found considerable disagreement between the human expert evaluation and the objective evaluation by perturbation. DISCUSSION: The best explanation method depended on the ECG measure. To ensure that future explanations of deep neural networks for medical data analyses are useful to medical experts, data scientists developing new explanation methods should collaborate tightly with domain experts. Because there is no explanation method that performs best in all use cases, several methods should be applied. CONCLUSION: Several explanation methods should be used to determine the most suitable approach. Andrea M. Storås, Steffen Mæland, Jonas Isaksen, Steven Alexander Hicks, Vajira Thambawita, Claus Graff, Hugo Hammer, Pål Halvorsen, Michael Riegler 0001, Jørgen K. Kanters |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | A Two-Timescale Learning Automata Solution to the Nonlinear Stochastic Proportional Polling ProblemabstractIn this article, we introduce a novel learning automata (LA) solution to the nonlinear stochastic proportional polling (NSPP) problem. The only available solution to this problem in the literature is that given by Nicopolitidis et al. (2003), Obaidat et al. (2002), and Papadimitriou et al. (2002). It was shown to solve a large set of the adaptive resource allocation problems under noisy environments (Nicopolitidis et al., 2003; Obaidat et al., 2002; Papadimitriou and Pomportsis, 2000 and 1999; Nicopolitidis et al., 2004; Obaidat et al., 2001; and Papadimitriou and Pomportsis, 2000). We make a threefold contribution. First, we take a two-timescale approach to the field of LA by estimating the reward probabilities on a faster timescale than the timescale for updating the polling probabilities. Second, by making a not-obvious choice of the objective function, we show that the NSPP problem is indeed an instantiation of the stochastic nonlinear fractional equality knapsack (NFEK) problem, which is a substantial resource allocation problem based on the incomplete and noisy information (Granmo and Oommen, 2010). Third, in contrast to the legacy approach taken by Papadimitriou and Maritsas (1992 and 1996), we show through the extensive experimental results that our solution is remarkably robust to the choice of tuning parameters and that it outperforms the state of the art solution in terms of the Bayesian expected loss. Anis Yazidi, Hugo Hammer, David S. Leslie |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | ML-Peaks: CHIP-Seq Peak Detection Pipeline using Machine Learning TechniquesabstractCHIP-Seq data is critical for identifying the locations where proteins bind to DNA, offering valuable insights into disease molecular mechanisms and potential therapeutic targets. However, identifying regions of protein binding, or peaks, in CHIP-seq data can be challenging due to limitations in peak detection methods. Current computational tools often require manual human inspection using data visualization, making it challenging and resource demanding to detect all peaks, particularly in large datasets. CHIP-seq data poses difficulties in detecting peaks due to its high background noise, low signal-to-noise ratio, and variation in the size and shape of the peaks. To overcome these challenges, we propose a data preprocessing approach using sliding window and feature reduction techniques, and the resulting features can be further used in machine learning methods. Our machine learning methodology can accurately identify peaks using a small training set, which represents a distinct advantage over commonly used statistical approaches, as it has a greater capacity for learning from data. We tested our methodology on the H3K9me3_TDH_BP CHIP-Seq dataset exploring a range of different machine learning methods, sliding window settings, and feature reduction techniques to detect peak values without human intervention. Our pipeline efficiently detected the peaks, and achieved an F1-score of 0.9644 and a false positive rate of 0.1030. Sajad Amouei Sheshkal, Michael Riegler 0001, Hugo Hammer |
CBMS | 3 |
| 2023 | Efficient quantile tracking using an oracleabstractAbstract Concept drift is a well-known issue that arises when working with data streams. In this paper, we present a procedure that allows a quantile tracking procedure to cope with concept drift. We suggest using expected quantile loss, a popular loss function in quantile regression, to monitor the quantile tracking error, which, in turn, is used to efficiently adapt to concept drift. The suggested procedures adapt efficiently to concept drift, and the tracking performance is close to theoretically optimal. The procedures were further applied to three real-life streaming data sets related to Twitter event detection, activity recognition, and stock trading. The results show that the procedures are efficient at adapting to concept drift, thereby documenting the real-world applicability of the procedures. We further used asymptotic theory from statistics to show the appealing theoretical property that, if the data stream distribution is stationary over time, the procedures converge to the true quantile. Hugo Hammer, Anis Yazidi, Michael Riegler 0001, Håvard Rue |
Appl. Intell. | 1 |
| 2023 | Prediction of schizophrenia from activity data using hidden Markov model parameters
Matthias Boeker, Hugo Hammer, Michael Riegler 0001, Pål Halvorsen, Petter Jakobsen |
Neural Comput. Appl. | 2 |
| 2022 | Estimating Tukey depth using incremental quantile estimatorsabstractMeasures of distance or how data points are positioned relative to each other are fundamental in pattern recognition. The concept of depth measures how deep an arbitrary point is positioned in a dataset, and is an interesting concept in this regard. However, while this concept has received a lot of attention in the statistical literature, its application within pattern recognition is still limited. To increase the applicability of the depth concept in pattern recognition, we address the well-known computational challenges associated with the depth concept, by suggesting to estimate depth using incremental quantile estimators. The suggested algorithm can not only estimate depth when the dataset is known in advance, but can also track depth for dynamically varying data streams by using recursive updates. The tracking ability of the algorithm was demonstrated based on a real-life application associated with detecting changes in human activity from real-time accelerometer observations. Given the flexibility of the suggested approach, it can detect virtually any kind of changes in the distributional patterns of the observations, and thus outperforms detection approaches based on the Mahalanobis distance. Hugo Hammer, Anis Yazidi, Håvard Rue |
Pattern Recognit. | 1 |
| 2022 | Solving Sensor Identification Problem Without Knowledge of the Ground Truth Using Replicator DynamicsabstractIn this article, we consider an emergent problem in the sensor fusion area in which unreliable sensors need to be identified in the absence of the ground truth. We devise a novel solution to the problem using the theory of replicator dynamics that require mild conditions compared to the available state-of-the-art approaches. The solution has a low computational complexity that is linear in terms of the number of involved sensors. We provide some sound theoretical results that catalog the convergence of our approach to a solution where we can clearly unveil the sensor type. Furthermore, we present some experimental results that demonstrate the convergence of our approach in concordance with our theoretical findings. Anis Yazidi, Marco A. Pinto-Orellana, Hugo Hammer, Peyman Mirtaheri, Enrique Herrera-Viedma |
IEEE Trans. Cybern. | 3 |
| 2021 | Diagnosing Schizophrenia from Activity Records using Hidden Markov Model ParametersabstractThe diagnosis of Schizophrenia is mainly based on qualitative characteristics. With the usage of portable devices which measure activity of humans, the diagnosis of Schizophrenia can be enriched through quantitative features. The goal of this work is to classify between schizophrenic and non-schizophrenic subjects based on their measured activity over a certain amount of time. To do so, the periods in which a subject was resting or active were identified by the application of a Hidden Markov model (HMM). The trained model parameters of the HMM, such as the mean or variance of activity during the state of rest or activity, are used as classification features for a logistic regression model. Our results indicate that the features from the HMM are significant in classifying between schizophrenic and non-schizophrenic subjects. Moreover, the features outperform the features derived through other methods in literature in terms of goodness-of-fit and classification performance. Matthias Boeker, Michael Riegler 0001, Hugo Hammer, Pål Halvorsen, Ole Bernt Fasmer, Petter Jakobsen |
CBMS | 3 |
| 2021 | HTAD: A Home-Tasks Activities Dataset with Wrist-Accelerometer and Audio Features
Enrique Garcia-Ceja, Vajira Thambawita, Steven Alexander Hicks, Debesh Jha, Petter Jakobsen, Hugo Hammer, Pål Halvorsen, Michael Riegler 0001 |
MMM (2) | 6 |
| 2021 | HYPERAKTIV: An Activity Dataset from Patients with Attention-Deficit/Hyperactivity Disorder (ADHD)abstractMachine learning research within healthcare frequently lacks the public data needed to be fully reproducible and comparable. Datasets are often restricted due to privacy concerns and legal requirements that come with patient-related data. Consequentially, many algorithms and models get published on the same topic without a standard benchmark to measure against. Therefore, this paper presents HYPERAKTIV, a public dataset containing health, activity, and heart rate data from patients diagnosed with attention deficit hyperactivity disorder, better known as ADHD. The dataset consists of data collected from 51 patients with ADHD and 52 clinical controls. In addition to the activity and heart rate data, we also include a series of patient attributes such as their age, sex, and information about their mental state, as well as output data from a computerized neuropsychological test. Together with the presented dataset, we also provide baseline experiments using traditional machine learning algorithms to predict ADHD based on the included activity data. We hope that this dataset can be used as a starting point for computer scientists who want to contribute to the field of mental health, and as a common benchmark for future work in ADHD analysis. Steven Alexander Hicks, Andrea Stautland, Ole Bernt Fasmer, Wenche Førland, Hugo Hammer, Pål Halvorsen, Kristin Mjeldheim, Ketil J. Oedegaard, Berge Osnes, Vigdis Elin Giæver Syrstad, Michael Riegler 0001, Petter Jakobsen |
MMSys | 5 |
| 2021 | Joint tracking of multiple quantiles through conditional quantilesabstractThe estimation of quantiles is one of the most fundamental data mining tasks. As most real-time data streams vary dynamically over time, there is a quest for adaptive quantile estimators. The most well-known type of adaptive quantile estimators is the incremental one which documents the state-of-the art performance in tracking quantiles. However, the absolute vast majority of incremental quantile estimators fail to jointly estimate multiple quantiles in a consistent manner without violating the monotone property of quantiles. In this paper, first we introduce the concept of conditional quantiles that can be used to extend incremental estimators to jointly track multiple quantiles. Second, we resort to the concept of conditional quantiles to propose two new estimators. Extensive experimental results, based on both synthetic and real-life data, show that the proposed estimators clearly outperform legacy state-of-the-art joint quantile tracking algorithms in terms of accuracy while achieving faster adaptivity in the face of dynamically varying data streams. Hugo Hammer, Anis Yazidi, Håvard Rue |
Inf. Sci. | 1 |
| 2021 | Enhanced Equivalence Projective Simulation: A Framework for Modeling Formation of Stimulus Equivalence ClassesabstractFormation of stimulus equivalence classes has been recently modeled through equivalence projective simulation (EPS), a modified version of a projective simulation (PS) learning agent. PS is endowed with an episodic memory that resembles the internal representation in the brain and the concept of cognitive maps. PS flexibility and interpretability enable the EPS model and, consequently the model we explore in this letter, to simulate a broad range of behaviors in matching-to-sample experiments. The episodic memory, the basis for agent decision making, is formed during the training phase. Derived relations in the EPS model that are not trained directly but can be established via the network's connections are computed on demand during the test phase trials by likelihood reasoning. In this letter, we investigate the formation of derived relations in the EPS model using network enhancement (NE), an iterative diffusion process, that yields an offline approach to the agent decision making at the testing phase. The NE process is applied after the training phase to denoise the memory network so that derived relations are formed in the memory network and retrieved during the testing phase. During the NE phase, indirect relations are enhanced, and the structure of episodic memory changes. This approach can also be interpreted as the agent's replay after the training phase, which is in line with recent findings in behavioral and neuroscience studies. In comparison with EPS, our model is able to model the formation of derived relations and other features such as the nodal effect in a more intrinsic manner. Decision making in the test phase is not an ad hoc computational method, but rather a retrieval and update process of the cached relations from the memory network based on the test trial. In order to study the role of parameters on agent performance, the proposed model is simulated and the results discussed through various experimental settings. Asieh Abolpour Mofrad, Anis Yazidi, Samaneh Abolpour Mofrad, Hugo Hammer, Erik Arntzen |
Neural Comput. | 4 |
| 2021 | Game-Theoretic Learning for Sensor Reliability Evaluation Without Knowledge of the Ground TruthabstractSensor fusion has attracted a lot of research attention during the few last years. Recently, a new research direction has emerged dealing with sensor fusion without knowledge of the ground truth. In this article, we present a novel solution to the latter pertinent problem. In contrast to the first reported solutions to this problem, we present a solution that does not involve any assumption on the group average reliability which makes our results more general than previous works. We devise a strategic game where we show that a perfect partitioning of the sensors into reliable and unreliable groups corresponds to a Nash equilibrium of the game. Furthermore, we give sound theoretical results that prove that those equilibria are indeed the unique Nash equilibria of the game. We then propose a solution involving a team of learning automata (LA) to unveil the identity of each sensor, whether it is reliable or unreliable, using game-theoretic learning. The experimental results show the accuracy of our solution and its ability to deal with settings that are unsolvable by legacy works. Anis Yazidi, Hugo Hammer, Konstantin E. Samouylov, Enrique Herrera-Viedma |
IEEE Trans. Cybern. | 2 |
| 2021 | Achieving Fair Load Balancing by Invoking a Learning Automata-Based Two-Time-Scale Separation ParadigmabstractIn this article, we consider the problem of load balancing (LB), but, unlike the approaches that have been proposed earlier, we attempt to resolve the problem in a fair manner (or rather, it would probably be more appropriate to describe it as an ϵ -fair manner because, although the LB can, probably, never be totally fair, we achieve this by being "as close to fair as possible"). The solution that we propose invokes a novel stochastic learning automaton (LA) scheme, so as to attain a distribution of the load to a number of nodes, where the performance level at the different nodes is approximately equal and each user experiences approximately the same Quality of the Service (QoS) irrespective of which node that he/she is connected to. Since the load is dynamically varying, static resource allocation schemes are doomed to underperform. This is further relevant in cloud environments, where we need dynamic approaches because the available resources are unpredictable (or rather, uncertain) by virtue of the shared nature of the resource pool. Furthermore, we prove here that there is a coupling involving LA's probabilities and the dynamics of the rewards themselves, which renders the environments to be nonstationary. This leads to the emergence of the so-called property of "stochastic diminishing rewards." Our newly proposed novel LA algorithm ϵ -optimally solves the problem, and this is done by resorting to a two-time-scale-based stochastic learning paradigm. As far as we know, the results presented here are of a pioneering sort, and we are unaware of any comparable results. Anis Yazidi, Ismail Hassan, Hugo Hammer, B. John Oommen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Analysis of Optical Brain Signals Using Connectivity Graph Networks
Marco A. Pinto-Orellana, Hugo Hammer |
CD-MAKE | 2 |
| 2020 | EvoDynamic: A Framework for the Evolution of Generally Represented Dynamical Systems and Its Application to Criticality
Sidney Pontes-Filho, Pedro G. Lind, Anis Yazidi, Jianhua Zhang 0004, Hugo Hammer, Gustavo Borges Moreno e Mello, Ioanna Sandvig, Gunnar Tufte, Stefano Nichele |
EvoApplications | 5 |
| 2020 | ACM Multimedia BioMedia 2020 Grand Challenge OverviewabstractThe BioMedia 2020 ACM Multimedia Grand Challenge is the second in a series of competitions focusing on the use of multimedia for different medical use-cases. In this year's challenge, participants are asked to develop algorithms that automatically predict the quality of a given human semen sample using a combination of visual, patient-related, and laboratory-analysis-related data. Compared to last year's challenge, participants are provided with a fully multimodal dataset (videos, analysis data, study participant data) from the field of assisted human reproduction. The tasks encourage the use of the different modalities contained within the dataset and finding smart ways of how they may be combined to further improve prediction accuracy. For example, using only video data or combining video data and patient-related data. The ground truth was developed through a preliminary analysis done by medical experts following the World Health Organization's standard for semen quality assessment. The task lays the basis for automatic, real-time support systems for artificial reproduction. We hope that this challenge motivates multimedia researchers to explore more medical-related applications and use their vast knowledge to make a real impact on people's lives. Steven Alexander Hicks, Vajira Thambawita, Hugo Hammer, Trine B. Haugen, Jorunn M. Andersen, Oliwia Witczak, Pål Halvorsen, Michael Riegler 0001 |
ACM Multimedia | 3 |
| 2020 | Toadstool: a dataset for training emotional intelligent machines playing Super Mario BrosabstractGames are often defined as engines of experience, and they are heavily relying on emotions, they arouse in players. In this paper, we present a dataset called Toadstool as well as a reproducible methodology to extend on the dataset. The dataset consists of video, sensor, and demographic data collected from ten participants playing Super Mario Bros, an iconic and famous video game. The sensor data is collected through an Empatica E4 wristband, which provides high-quality measurements and is graded as a medical device. In addition to the dataset and the methodology for data collection, we present a set of baseline experiments which show that we can use video game frames together with the facial expressions to predict the blood volume pulse of the person playing Super Mario Bros. With the dataset and the collection methodology we aim to contribute to research on emotionally aware machine learning algorithms, focusing on reinforcement learning and multimodal data fusion. We believe that the presented dataset can be interesting for a manifold of researchers to explore exciting new interdisciplinary questions. Henrik Svoren, Vajira Thambawita, Pål Halvorsen, Petter Jakobsen, Enrique Garcia-Ceja, Farzan Majeed Noori, Hugo Hammer, Mathias Lux, Michael Riegler 0001, Steven Alexander Hicks |
MMSys | 7 |
| 2020 | PMData: a sports logging datasetabstractIn this paper, we present PMData: a dataset that combines traditional lifelogging data with sports-activity data. Our dataset enables the development of novel data analysis and machine-learning applications where, for instance, additional sports data is used to predict and analyze everyday developments, like a person's weight and sleep patterns; and applications where traditional lifelog data is used in a sports context to predict athletes' performance. PMData combines input from Fitbit Versa 2 smartwatch wristbands, the PMSys sports logging smartphone application, and Google forms. Logging data has been collected from 16 persons for five months. Our initial experiments show that novel analyses are possible, but there is still room for improvement. Vajira Thambawita, Steven Alexander Hicks, Hanna Borgli, Håkon Kvale Stensland, Debesh Jha, Martin Kristoffer Svensen, Svein Arne Pettersen, Dag Johansen, Håvard D. Johansen, Susann Dahl Pettersen, Simon Nordvang, Sigurd Pedersen, Anders T. Gjerdrum, Tor-Morten Grønli, Per Morten Fredriksen, Ragnhild Eg, Kjeld Hansen, Siri Fagernes, Christine Claudi, Andreas Biørn-Hansen, Duc-Tien Dang-Nguyen, Tomas Kupka, Hugo Hammer, Ramesh Jain 0001, Michael Riegler 0001, Pål Halvorsen |
MMSys | 23 |
| 2020 | Mitigating DDoS using weight-based geographical clusteringabstractSummary Distributed denial of service (DDoS) attacks have for the last two decades been among the greatest threats facing the internet infrastructure. Mitigating DDoS attacks is a particularly challenging task as an attacker tries to conceal a huge amount of traffic inside a legitimate traffic flow. This article proposes to use data mining approaches to find unique hidden data structures which are able to characterize the normal traffic flow. This will serve as a mean for filtering illegitimate traffic under DDoS attacks. In this endeavor, we devise three algorithms built on previously uncharted areas within mitigation techniques where clustering techniques are used to create geographical clusters in regions which are likely to contain legitimate traffic. We argue through extensive experimental results that establishing clusters around this narrative is a superior solution to clustering algorithms which rely on bitwise distances between IP addresses. In addition, the DDoS filtering algorithm is deployed in a virtual Linux environment using Nfqueue and tested in a simulated real‐life DDoS attack. Madeleine Victoria Kongshavn, Hårek Haugerud, Anis Yazidi, Torleiv Maseng, Hugo Hammer |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | An Extensive Study on Cross-Dataset Bias and Evaluation Metrics Interpretation for Machine Learning Applied to Gastrointestinal Tract Abnormality ClassificationabstractPrecise and efficient automated identification of gastrointestinal (GI) tract diseases can help doctors treat more patients and improve the rate of disease detection and identification. Currently, automatic analysis of diseases in the GI tract is a hot topic in both computer science and medical-related journals. Nevertheless, the evaluation of such an automatic analysis is often incomplete or simply wrong. Algorithms are often only tested on small and biased datasets, and cross-dataset evaluations are rarely performed. A clear understanding of evaluation metrics and machine learning models with cross datasets is crucial to bring research in the field to a new quality level. Toward this goal, we present comprehensive evaluations of five distinct machine learning models using global features and deep neural networks that can classify 16 different key types of GI tract conditions, including pathological findings, anatomical landmarks, polyp removal conditions, and normal findings from images captured by common GI tract examination instruments. In our evaluation, we introduce performance hexagons using six performance metrics, such as recall, precision, specificity, accuracy, F1-score, and the Matthews correlation coefficient to demonstrate how to determine the real capabilities of models rather than evaluating them shallowly. Furthermore, we perform cross-dataset evaluations using different datasets for training and testing. With these cross-dataset evaluations, we demonstrate the challenge of actually building a generalizable model that could be used across different hospitals. Our experiments clearly show that more sophisticated performance metrics and evaluation methods need to be applied to get reliable models rather than depending on evaluations of the splits of the same dataset—that is, the performance metrics should always be interpreted together rather than relying on a single metric. Vajira Thambawita, Debesh Jha, Hugo Hammer, Håvard D. Johansen, Dag Johansen, Pål Halvorsen, Michael Riegler 0001 |
ACM Trans. Comput. Heal. | 3 |
| 2020 | Equivalence Projective Simulation as a Framework for Modeling Formation of Stimulus Equivalence ClassesabstractStimulus equivalence (SE) and projective simulation (PS) study complex behavior, the former in human subjects and the latter in artificial agents. We apply the PS learning framework for modeling the formation of equivalence classes. For this purpose, we first modify the PS model to accommodate imitating the emergence of equivalence relations. Later, we formulate the SE formation through the matching-to-sample (MTS) procedure. The proposed version of PS model, called the equivalence projective simulation (EPS) model, is able to act within a varying action set and derive new relations without receiving feedback from the environment. To the best of our knowledge, it is the first time that the field of equivalence theory in behavior analysis has been linked to an artificial agent in a machine learning context. This model has many advantages over existing neural network models. Briefly, our EPS model is not a black box model, but rather a model with the capability of easy interpretation and flexibility for further modifications. To validate the model, some experimental results performed by prominent behavior analysts are simulated. The results confirm that the EPS model is able to reliably simulate and replicate the same behavior as real experiments in various settings, including formation of equivalence relations in typical participants, nonformation of equivalence relations in language-disabled children, and nodal effect in a linear series with nodal distance five. Moreover, through a hypothetical experiment, we discuss the possibility of applying EPS in further equivalence theory research. Asieh Abolpour Mofrad, Anis Yazidi, Hugo Hammer, Erik Arntzen |
Neural Comput. | 3 |
| 2020 | Smooth estimates of multiple quantiles in dynamically varying data streams
Hugo Hammer, Anis Yazidi |
Pattern Anal. Appl. | 1 |
| 2020 | Tracking of multiple quantiles in dynamically varying data streams
Hugo Hammer, Anis Yazidi, Håvard Rue |
Pattern Anal. Appl. | 1 |
| 2019 | THREAT: A Large Annotated Corpus for Detection of Violent ThreatsabstractUnderstanding, detecting, moderating and in extreme cases deleting hateful comments in online discussions and social media are well-known challenges. In this paper we present a dataset consisting of a total of around 30 000 sentences from around 10 000 YouTube comments. Each sentence is manually annotated as either being a violent threat or not. Violent threats is the most extreme form of hateful communication and is of particular importance from an online radicalization and national security perspective. This is the first publicly available dataset with such an annotation. The dataset can further be useful to develop automatic moderation tools or may even be useful from a social science perspective for analyzing the characteristics of online threats and how hateful discussions evolve. Hugo Hammer, Michael Riegler 0001, Lilja Øvrelid, Erik Velldal |
CBMI | 1 |
| 2019 | Semantic Analysis of Soccer News for Automatic Game Event ClassificationabstractWe are today overwhelmed with information, of which an important part is news. Sports news, in particular, has become very popular, where soccer makes up a big part of this coverage. For sports fans, it can be a time consuming and tedious to keep up with the news that they really care about. In this paper, we present different machine learning methods applied to soccer news from a Norwegian newspaper and a TV station's news site to summarize the content in a short and digestible manner. We present a system to collect, index, label, analyze, and present the collected news articles based on the content. We perform a thorough comparison between deep learning and traditional machine learning algorithms on text classification. Furthermore, we present a dataset of soccer news which was collected from two different Norwegian news sites and shared online. Aanund Jupskås Nordskog, Pål Halvorsen, Steven Alexander Hicks, Håkon Kvale Stensland, Hugo Hammer, Dag Johansen, Michael Riegler 0001 |
CBMI | 5 |
| 2019 | GANEx: A complete pipeline of training, inference and benchmarking GAN experimentsabstractDeep learning (DL) is one of the standard methods in the field of multimedia research to perform data classification, detection, segmentation and generation. Within DL, generative adversarial networks (GANs) represents a new and highly popular branch of methods. GANs have the capability to generate, from random noise or conditional input, new data realizations within the dataset population. While generation is popular and highly useful in itself, GANs can also be useful to improve supervised DL. GAN-based approaches can, for example, perform segmentation or create synthetic data for training other DL models. The latter one is especially interesting in domains where not much training data exists such as medical multimedia. In this respect, performing a series of experiments involving GANs can be very time consuming due to the lack of tools that support the whole pipeline such as structured training, testing and tracking of different architectures and configurations. Moreover, the success of generative models is highly dependent on hyper-parameter optimization and statistical analysis in the design and fine-tuning stages. In this paper, we present a new tool called GANEx for making the whole pipeline of training, inference and benchmarking GANs faster, more efficient and more structured. The tool consists of a special library called FastGAN which allows designing generative models very fast. Moreover, GANEx has a graphical user interface to support structured experimenting, quick hyper-parameter configurations and output analysis. The presented tool is not limited to a specific DL framework and can be therefore even used to compare the performance of cross frameworks. Vajira Thambawita, Hugo Hammer, Michael Riegler 0001, Pål Halvorsen |
CBMI | 2 |
| 2019 | VISEM: a multimodal video dataset of human spermatozoaabstractReal multimedia datasets that contain more than just images or text are rare. Even more so are open multimedia datasets in medicine. Often, clinically related datasets only consist of image or videos. In this paper, we present a dataset that is novel in two ways. Firstly, it is a multi-modal dataset containing different data sources such as videos, biological analysis data, and participant data. Secondly, it is the first dataset of that kind in the field of human reproduction. It consists of anonymized data from 85 different participants. We hope this dataset paper will inspire people to apply their knowledge in this important field, generate shareable results in the domain, and ultimately improve human infertility investigation and treatment. Trine B. Haugen, Steven Alexander Hicks, Jorunn M. Andersen, Oliwia Witczak, Hugo Hammer, Rune Johan Borgli, Pål Halvorsen, Michael Riegler 0001 |
MMSys | 5 |
| 2019 | A new quantile tracking algorithm using a generalized exponentially weighted average of observations
Hugo Hammer, Anis Yazidi, Håvard Rue |
Appl. Intell. | 1 |
| 2019 | On solving the SPL problem using the concept of probability flux
Asieh Abolpour Mofrad, Anis Yazidi, Hugo Hammer |
Appl. Intell. | 3 |
| 2019 | Two-time scale learning automata: an efficient decision making mechanism for stochastic nonlinear resource allocation
Anis Yazidi, Hugo Hammer, Tore Møller Jonassen |
Appl. Intell. | 2 |
| 2019 | Multiplicative Update Methods for Incremental Quantile EstimationabstractWe present a novel lightweight incremental quantile estimator which possesses far less complexity than the Tierney's estimator and its extensions. Notably, our algorithm relies only on tuning one single parameter which is a plausible property which we could only find in the discretized quantile estimator Frugal. This makes our algorithm easy to tune for better performance. Furthermore, our algorithm is multiplicative which makes it highly suitable to handle quantile estimation in systems in which the underlying distribution varies with time. Unlike Frugal and our legacy work which are randomized algorithms, we suggest deterministic updates where the step size is adjusted in a subtle manner to ensure the convergence. The deterministic algorithm is more efficient since the estimate is updated at every iteration. The convergence of the proposed estimator is proven using the theory of stochastic learning. Extensive experimental results show that our estimator clearly outperforms legacy works. Anis Yazidi, Hugo Hammer |
IEEE Trans. Cybern. | 2 |
| 2018 | Parameter estimation in abruptly changing dynamic environments using stochastic learning weak estimator
Hugo Hammer, Anis Yazidi |
Appl. Intell. | 1 |
| 2018 | Solving stochastic nonlinear resource allocation problems using continuous learning automata
Anis Yazidi, Hugo Hammer |
Appl. Intell. | 2 |
| 2018 | A Queue Model for Reliable Forecasting of Future CPU Consumption
Hugo Hammer, Anis Yazidi, Alfred Bratterud, Hårek Haugerud, Boning Feng |
Mob. Networks Appl. | 1 |
| 2018 | On the classification of dynamical data streams using novel "Anti-Bayesian" techniques
Hugo Hammer, Anis Yazidi, B. John Oommen |
Pattern Recognit. | 1 |
| 2017 | A Higher-Fidelity Frugal Quantile Estimator
Anis Yazidi, Hugo Hammer, B. John Oommen |
ADMA | 2 |
| 2017 | On using novel "Anti-Bayesian" techniques for the classification of dynamical data streamsabstractThe classification of dynamical data streams is among the most complex problems encountered in classification. This is, firstly, because the distribution of the data streams is non-stationary, and it changes without any prior “warning”. Secondly, the manner in which it changes is also unknown. Thirdly, and more interestingly, the model operates with the assumption that the correct classes of previously-classified patterns become available at a juncture after their appearance. This paper pioneers the use of unreported novel schemes that can classify such dynamical data streams by invoking the recently-introduced “Anti-Bayesian” (AB) techniques. Contrary to the Bayesian paradigm, that compare the testing sample with the distribution's central points, AB techniques are based on the information in the distant-from-the-mean samples. Most Bayesian approaches can be naturally extended to dynamical systems by dynamically tracking the mean of each class using, for example, the exponential moving average based estimator, or a sliding window estimator. The AB schemes introduced by Oommen et al., on the other hand, work with a radically different approach and with the non-central quantiles of the distributions. Surprisingly and counter-intuitively, the reported AB methods work equally or close-to-equally well to an optimal supervised Bayesian scheme on a host of accepted PR problems. This thus begs its natural extension to the unexplored arena of classification for dynamical data streams. Naturally, for such an AB classification approach, we need to track the non-stationarity of the quantiles of the classes. To achieve this, in this paper, we develop an AB approach for the online classification of data streams by applying the efficient and robust quantile estimators developed by Yazidi and Hammer [3], [13]. Apart from the methodology itself, in this paper, we compare the Bayesian and AB approaches. The results demonstrate the intriguing and counter-intuitive results that the AB approach shows competitive results to the Bayesian approach. Furthermore, the AB approach is much more robust against outliers, which is an inherent property of quantile estimators [3], [13], which is a property that the Bayesian approach cannot match, since it rather tracks the mean. Hugo Hammer, Anis Yazidi, B. John Oommen |
CEC | 1 |
| 2017 | Incremental Quantiles Estimators for Tracking Multiple Quantiles
Hugo Hammer, Anis Yazidi |
IEA/AIE (1) | 1 |
| 2017 | Two-Timescale Learning Automata for Solving Stochastic Nonlinear Resource Allocation Problems
Anis Yazidi, Hugo Hammer, Tore Møller Jonassen |
IEA/AIE (1) | 2 |
| 2017 | The concept of workload delay as a quality-of-service metric for consolidated cloud environments with deadline requirementsabstractVirtual Machine (VM) consolidation in the cloud has received significant research interest. A large body of approaches for VM consolidation in data centers resort to variants of the bin packing problem which tries to minimize the number of deployed physical machines while meeting the Service-Level-Agreement (SLA) constraints. In this paper we introduce the concept of workload delay as a Quality-of-Service (QoS) metric that captures directly the resulting degradation that a cloud user would experience in the case where the SLA is violated. Our results, that are based on real-life trace-based simulations, show that consolidating VMs based on the level of utilization results in little control over the resulting delay, a particularly significant drawback when running jobs with deadline requirements, while we are able to control the delay much better if we take into account our suggested metric of the delay. Evangelos Tasoulas 0001, Hugo Hammer, Hårek Haugerud, Anis Yazidi, Alfred Bratterud, Boning Feng |
NCA | 2 |
| 2017 | "Anti-Bayesian" flat and hierarchical clustering using symmetric quantiloids
Hugo Hammer, Anis Yazidi, B. John Oommen |
Inf. Sci. | 1 |
| 2016 | "Anti-Bayesian" Flat and Hierarchical Clustering Using Symmetric Quantiloids
Anis Yazidi, Hugo Hammer, B. John Oommen |
IEA/AIE | 2 |
| 2015 | A Novel Clustering Algorithm Based on a Non-parametric "Anti-Bayesian" Paradigm
Hugo Hammer, Anis Yazidi, B. John Oommen |
IEA/AIE | 1 |
| 2015 | A Simple and Efficient Algorithm for Lexicon Generation Inspired by Structural Balance Theory
Anis Yazidi, Aleksander Bai, Hugo Hammer, Paal E. Engelstad |
IEA/AIE | 3 |