Martin Gjoreski

dblp:157/4929 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-1220-7418ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
YearPublicationVenuePosition
2025 Counterfactual Concept Bottleneck Models
abstract
Current deep learning models are not designed to simultaneously address three fundamental questions: predict class labels to solve a given classification task (the "What?"), simulate changes in the situation to evaluate how this impacts class predictions (the "How?"), and imagine how the scenario should change to result in different class predictions (the "Why not?"). While current approaches in causal representation learning and concept interpretability are designed to address some of these questions individually (such as Concept Bottleneck Models, which address both ``what'' and ``how'' questions), no current deep learning model is specifically built to answer all of them at the same time. To bridge this gap, we introduce CounterFactual Concept Bottleneck Models (CF-CBMs), a class of models designed to efficiently address the above queries all at once without the need to run post-hoc searches. Our experimental results demonstrate that CF-CBMs: achieve classification accuracy comparable to black-box models and existing CBMs (“What?”), rely on fewer important concepts leading to simpler explanations (“How?”), and produce interpretable, concept-based counterfactuals (“Why not?”). Additionally, we show that training the counterfactual generator jointly with the CBM leads to two key improvements: (i) it alters the model's decision-making process, making the model rely on fewer important concepts (leading to simpler explanations), and (ii) it significantly increases the causal effect of concept interventions on class predictions, making the model more responsive to these changes.
Gabriele Dominici, Pietro Barbiero, Francesco Giannini, Martin Gjoreski, Giuseppe Marra, Marc Langheinrich
ICLR4
2025 Causal Concept Graph Models: Beyond Causal Opacity in Deep Learning
abstract
Causal opacity denotes the difficulty in understanding the "hidden" causal structure underlying the decisions of deep neural network (DNN) models. This leads to the inability to rely on and verify state-of-the-art DNN-based systems, especially in high-stakes scenarios. For this reason, circumventing causal opacity in DNNs represents a key open challenge at the intersection of deep learning, interpretability, and causality. This work addresses this gap by introducing Causal Concept Graph Models (Causal CGMs), a class of interpretable models whose decision-making process is causally transparent by design. Our experiments show that Causal CGMs can: (i) match the generalisation performance of causally opaque models, (ii) enable human-in-the-loop corrections to mispredicted intermediate reasoning steps, boosting not just downstream accuracy after corrections but also the reliability of the explanations provided for specific instances, and (iii) support the analysis of interventional and counterfactual scenarios, thereby improving the model's causal interpretability and supporting the effective verification of its reliability and fairness.
Gabriele Dominici, Pietro Barbiero, Mateo Espinosa Zarlenga, Alberto Termine, Martin Gjoreski, Giuseppe Marra, Marc Langheinrich
ICLR5
2025 FLUX: Efficient Descriptor-Driven Clustered Federated Learning under Arbitrary Distribution Shifts
abstract
Federated Learning (FL) enables collaborative model training across multiple clients while preserving data privacy. Traditional FL methods often use a global model to fit all clients, assuming that clients' data are independent and identically distributed (IID). However, when this assumption does not hold, the global model accuracy may drop significantly, limiting FL applicability in real-world scenarios. To address this gap, we propose FLUX, a novel clustering-based FL (CFL) framework that addresses the four most common types of distribution shifts during both training and test time. To this end, FLUX leverages privacy-preserving client-side descriptor extraction and unsupervised clustering to ensure robust performance and scalability across varying levels and types of distribution shifts. Unlike existing CFL methods addressing non-IID client distribution shifts, FLUX i) does not require any prior knowledge of the types of distribution shifts or the number of client clusters, and ii) supports test-time adaptation, enabling unseen and unlabeled clients to benefit from the most suitable cluster-specific models. Extensive experiments across four standard benchmarks, two real-world datasets and ten state-of-the-art baselines show that FLUX improves performance and stability under diverse distribution shifts—achieving an average accuracy gain of up to 23 percentage points over the best-performing baselines—while maintaining computational and communication overhead comparable to FedAvg.
Dario Fenoglio, Mohan Li, Pietro Barbiero, Nicholas D. Lane, Marc Langheinrich, Martin Gjoreski
NeurIPS6
2024 Multi-Frequency Federated Learning for Human Activity Recognition Using Head-Worn Sensors
abstract
Human Activity Recognition (HAR) benefits various application domains, including health and elderly care. Traditional HAR involves constructing pipelines reliant on centralized user data, which can pose privacy concerns as they necessitate the uploading of user data to a centralized server. This work proposes multi-frequency Federated Learning (FL) to enable: (1) privacy-aware ML; (2) joint ML model learning across devices with varying sampling frequency. We focus on head-worn devices (e.g., earbuds and smart glasses), a relatively unexplored domain compared to traditional smartwatch- or smartphone-based HAR. Results have shown improvements on two datasets against frequency-specific approaches, indicating a promising future in the multi-frequency FL-HAR task. The proposed network’s implementation is publicly available for further research and development.**
Dario Fenoglio, Mohan Li, Davide Casnici, Matías Laporte, Shkurta Gashi, Silvia Santini, Martin Gjoreski, Marc Langheinrich
IE7
2024 Federated Behavioural Planes: Explaining the Evolution of Client Behaviour in Federated Learning
abstract
Federated Learning (FL), a privacy-aware approach in distributed deep learning environments, enables many clients to collaboratively train a model without sharing sensitive data, thereby reducing privacy risks. However, enabling human trust and control over FL systems requires understanding the evolving behaviour of clients, whether beneficial or detrimental for the training, which still represents a key challenge in the current literature. To address this challenge, we introduce Federated Behavioural Planes (FBPs), a novel method to analyse, visualise, and explain the dynamics of FL systems, showing how clients behave under two different lenses: predictive performance (error behavioural space) and decision-making processes (counterfactual behavioural space). Our experiments demonstrate that FBPs provide informative trajectories describing the evolving states of clients and their contributions to the global model, thereby enabling the identification of clusters of clients with similar behaviours. Leveraging the patterns identified by FBPs, we propose a robust aggregation technique named Federated Behavioural Shields to detect malicious or noisy client models, thereby enhancing security and surpassing the efficacy of existing state-of-the-art FL defense mechanisms. Our code is publicly available on GitHub.
Dario Fenoglio, Gabriele Dominici, Pietro Barbiero, Alberto Paolo Tonda, Martin Gjoreski, Marc Langheinrich
NeurIPS5
2023 A Federated Unsupervised Personalisation for Cognitive Workload Estimation
abstract
Accurate Cognitive Workload (CW) estimation, crucial in mobile healthcare and human-machine interaction, is impeded by client heterogeneity, data limitations, and privacy concerns, especially in the presence of Out-of-Distribution (OoD) clients. This study proposes a robust framework that is based on Federated Learning to protect data privacy, and utilizes context-based STRNet to enable joint cross-user learning on heterogeneous datasets, enhancing model generalisability. The framework includes a novel Unsupervised Client Personalisation strategy that prevents accuracy loss in OoD clients. We tested our framework on two publicly available CW datasets, COLET and ADABase. The framework improved the accuracy of centralized approaches while preserving data privacy. The framework is model-agnostic, efficient, and enables unsupervised personalisation for each client, bolstering the quality and robustness of the end-to-end deep learning models.
Dario Fenoglio, Martin Gjoreski, Marc Langheinrich
MUM2
2023 Federated Learning for Privacy-aware Cognitive Workload Estimation
abstract
Human physiological monitoring has become easily accessible by integrating wearable devices into our lives, providing valuable real-time data. Methods for Cognitive Workload (CW) estimation utilize such physiological data to quantify CW during task execution. These methods are crucial for various domains, including mobile healthcare, forecasting human errors, and human-machine interaction. However, accurately estimating CW continues to pose a challenge due to the absence of objective ground truth data, context dependency, and the privacy sensitivity of the data. This study tackled the complex task of estimating CW based on privacy-sensitive data (e.g., eye movement, pupil diameter, blink information, and other physiological signals) using Federated Learning (FL) methods to improve user privacy. We compared the outcomes of the FL models with the more conventional centralized approach on two publicly-available datasets COLET and ADABase, which include data from 75 participants overall. The results highlight the efficacy of FL in collaboratively training a global (person-independent) model. The FL models achieved performances on par with centralized state-of-the-art models while preserving data privacy. Recognizing the importance of privacy in user sensing, FL presents a promising approach that enables wearable sensing applications in privacy-sensitive domains.
Dario Fenoglio, Daniel Josifovski, Alessandro Gobbetti, Mattias Formo, Hristijan Gjoreski, Martin Gjoreski, Marc Langheinrich
MUM6
2022 Handling Missing Data For Sleep Monitoring Systems
abstract
Sensor-based sleep monitoring systems can be used to track sleep behavior on a daily basis and provide feedback to their users to promote health and well-being. Such systems can provide data visualizations to enable self-reflection on sleep habits or a sleep coaching service to improve sleep quality. To provide useful feedback, sleep monitoring systems must be able to recognize whether an individual is sleeping or awake. Existing approaches to infer sleep-wake phases, however, typically assume continuous streams of data to be available at inference time. In real-world settings, though, data streams or data samples may be missing, causing severe performance degradation of models trained on complete data streams. In this paper, we investigate the impact of missing data to recognize sleep and wake, and use regression- and interpolation-based imputation strategies to mitigate the errors that might be caused by incomplete data. To evaluate our approach, we use a data set that includes physiological traces - collected using wristbands -, behavioral data - gathered using smartphones - and self-reports from 16 participants over 30 days. Our results show that the presence of missing sensor data degrades the balanced accuracy of the classifier on average by 10–35 percentage points for detecting sleep and wake depending on the missing data rate. The impu-tation strategies explored in this work increase the performance of the classifier by 4–30 percentage points. These results open up new opportunities to improve the robustness of sleep monitoring systems against missing data.
Shkurta Gashi, Lidia Alecci, Martin Gjoreski, Elena Di Lascio, Abhinav Mehrotra, Mirco Musolesi, Maike E. Debus, Francesca Gasparini, Silvia Santini
ACII3
2022 BayCon: Model-agnostic Bayesian Counterfactual Generator
abstract
Generating counterfactuals to discover hypothetical predictive scenarios is the de facto standard for explaining machine learning models and their predictions. However, building a counterfactual explainer that is time-efficient, scalable, and model-agnostic, in addition to being compatible with continuous and categorical attributes, remains an open challenge. To complicate matters even more, ensuring that the contrastive instances are optimised for feature sparsity, remain close to the explained instance, and are not drawn from outside of the data manifold, is far from trivial. To address this gap we propose BayCon: a novel counterfactual generator based on probabilistic feature sampling and Bayesian optimisation. Such an approach can combine multiple objectives by employing a surrogate model to guide the counterfactual search. We demonstrate the advantages of our method through a collection of experiments based on six real-life datasets representing three regression tasks and three classification tasks.
Piotr Romashov, Martin Gjoreski, Kacper Sokol, Maria Vanina Martinez, Marc Langheinrich
IJCAI2
2021 Learning comprehensible and accurate hybrid trees
Rok Piltaver, Mitja Lustrek, Saso Dzeroski, Martin Gjoreski, Matjaz Gams
Expert Syst. Appl.4
2019 Exploring Dietary Intake Data collected by FPQ using Unsupervised Learning
abstract
Populations in countries undergoing rapid transition are experiencing food- and nutrition-related problems. To acquire high-quality nutrition information, we need beside adequate data about food consumption, also efficient methods for the extraction of information from the collected data. Our aim was to develop a methodology for analyzing and reasoning about dietary intake data collected by a food propensity questionnaire (FPQ) and dependent 24-hour recalls (24HRs). We analysed a subset of data (about 197 participants) in the SI.Menu survey carried out in 2016/17 in Slovenia. The participants completed FPQs and 24HRs. We were able to identify four clusters. Two clusters represented participants with more healthy habits, e.g., low intake of animal fats, high breakfast frequency, and high intake of fruits and vegetables. The other two clusters represented participants with less healthy habits, e.g., high intake of animal fats, low breakfast frequency and increased BMI. The four clusters can be well separated by only four variables. This interesting discovery could lead to simplified FFQ questionnaires, which could significantly decrease the participants' burden and could ensure participant compliance in similar studies. Having big national data set related to nutrition should ease the process of creating sustainable policies that will ultimately benefit agriculture, human health and the environment.
Martin Gjoreski, Stefan Kochev, Nina Resçiç, Matej Gregoric, Tome Eftimov, Barbara Korousic-Seljak
IEEE BigData1
2017 Management of Physical, Mental and Environmental Stress at the Workplace
abstract
We present the Fit4Work system for monitoring and management of physical, mental and environmental stress at the workplace. The system was designed specifically for older workers who are subject to sedentary stressful work in an office environment. It uses commercially available devices and intelligent methods, which utilize machine-learning models to monitor the three aspects of the users' lifestyle, and provide recommendations for improving them. The results show that the system can adequately recognize the user's physical activities, estimate energy expenditure and detect mental stress, as well as recognize and reason about unhealthy environment. The system provides recommendations according to the monitoring results.
Bozidara Cvetkovic, Martin Gjoreski, Jure Sorn, Martin Freser, Maciej Bogdanski, Katarzyna Jackowska, Michal Kosiedowski, Aleksander Stroinski, Mitja Lustrek
Intelligent Environments2
2017 Chronic Heart Failure Detection from Heart Sounds Using a Stack of Machine-Learning Classifiers
abstract
Chronic heart failure represents a global pandemic, currently affecting over 26 million of patients worldwide. It is a major contributor in the death rate of patients with cardiovascular diseases and results in more than 1 million hospitalizations annually in Europe and North America. Methods for chronic heart failure detection can be utilized to act preventive, improve early diagnosis and avoid hospitalizations or even life-threatening situations, thus highly enhance the quality of patient’s life. In this paper, we present a machine-learning method for chronic heart failure detection from heart sounds. The method consists of: filtering, segmentation, feature extraction and machine learning. The method was tested with a leave-one-subject-out evaluation technique on data from 122 subjects, gathered in the study. The method achieved 96% accuracy, outperforming a majority classifier for 15 percentage points. More specifically, it detects (recalls) 87% of the chronic heart failure subjects with a precision of 87%. The study confirmed that advanced machine learning applied on real-life sounds recorded with an unobtrusive digital stethoscope can be used for chronic heart failure detection.
Martin Gjoreski, Monika Simjanoska, Anton Gradisek, Ana Peterlin, Matjaz Gams, Gregor Poglajen
Intelligent Environments1
2017 Monitoring Physical Activity and Mental Stress Using Wrist-Worn Device and a Smartphone
Bozidara Cvetkovic, Martin Gjoreski, Jure Sorn, Pavel Maslov, Mitja Lustrek
ECML/PKDD (3)2
2017 Monitoring stress with a wrist device using context
Martin Gjoreski, Mitja Lustrek, Matjaz Gams, Hristijan Gjoreski
J. Biomed. Informatics1
2016 Continuous Live Stress Monitoring with a Wristband
abstract
In this paper we propose a method for continuous stress monitoring using data provided by a commercial wrist device equipped with common physiological sensors and an accelerometer. The method consists of three machine-learning components: a laboratory stress-detector that detects short-term stress every 2 minutes; an activity recognizer that continuously recognizes user's activity and thus provides context information; and a context-based stress detector that first aggregates the predictions of the laboratory detector, and then exploits the user's context in order to provide the final decision in a 20 minute interval. The method was trained on 21 subjects in a laboratory setting and tested on 5 subjects in a real-life setting. The accuracy on 55 days of real-life data was 92%. The method is currently being implemented as a smartphone application, which will be demonstrated at the conference.
Martin Gjoreski, Hristijan Gjoreski, Mitja Lustrek, Matjaz Gams
ECAI1
2015 Automatic Detection of Perceived Stress in Campus Students Using Smartphones
abstract
This paper presents an approach to detecting perceived stress in students using data collected with smartphones. The goal is to develop a machine-learning model that can unobtrusively detect the stress level in students using data from several smartphone sources: accelerometers, audio recorder, GPS, Wi-Fi, call log and light sensor. From these, features were constructed describing the students' deviation from usual behaviour. As ground truth, we used the data obtained from stress level questionnaires with three possible stress levels: "Not stressed", "Slightly stressed" and "Stressed". Several machine learning approaches were tested: a general models for all the students, models for cluster of similar students, and student-specific models. Our findings show that the perceived stress is highly subjective and that only person-specific models are substantially better than the baseline.
Martin Gjoreski, Hristijan Gjoreski, Mitja Lustrek, Matjaz Gams
Intelligent Environments1