EDBT 2026 Demo / reviewers in the wild / expert
Mark Purcell
dblp:15/2399
· DBLP profile ↗
4ranked-venue papers in the field
0as first author
3since 2021 · last 2023
—ORCID · unresolved
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Pruning Federated Learning Models for Anomaly Detection in Resource-Constrained EnvironmentsabstractThe evolving complexity of modern IT infrastructures has paved the way for malicious actors to exploit a wide array of vulnerabilities that can compromise the integrity of these systems. Monitoring complex IT systems is expensive and often requires dedicated infrastructure for deploying Intrusion and/or Anomaly Detection Systems. Moreover, ML-based solutions need large training sets, which add to the overall cost. To tackle these challenges we present INTELLECT, a novel approach to Intrusion and/or Anomaly Detection System, which leverages Federated Learning and model pruning techniques to cooperatively train high-accuracy models using distributed datasets and derive a fleet of lightweight models, which can be deployed without incurring additional costs for dedicated infrastructure. INTELLECT expands on the state-of-the-art techniques for feature selection, model pruning, and model distillation to create an interconnected pipeline. We empirically demonstrate the effectiveness of the methodology on benchmark datasets, and we present guidelines for the deployment in production systems. Simone Magnani, Stefano Braghin, Ambrish Rawat, Roberto Doriguzzi Corin, Mark Purcell, Domenico Siracusa |
IEEE Big Data | 5 |
| 2022 | Adaptive Aggregation For Federated LearningabstractIn this paper, we present a new scalable and adaptive architecture for FL aggregation. First, we demonstrate how traditional tree overlay based aggregation techniques (from P2P, publish-subscribe and stream processing research) can help FL aggregation scale, but are ineffective from a resource utilization and cost standpoint. Next, we present the design and implementation of AdaFed, which uses serverless/cloud functions to adaptively scale aggregation in a resource efficient and fault tolerant manner. We describe how AdaFed enables FL aggregation to be dynamically deployed only when necessary, elastically scaled to handle participant joins/leaves and is fault tolerant with minimal effort required on the (aggregation) programmer side. We also demonstrate that our prototype based on Ray [1] scales to thousands of participants, and is able to achieve a > 90% reduction in resource requirements and cost, with minimal impact on aggregation latency. K. R. Jayaram, Vinod Muthusamy, Gegi Thomas, Ashish Verma 0001, Mark Purcell |
IEEE Big Data | 5 |
| 2022 | Machine Learning Platform for Extreme Scale Computing on Compressed IoT DataabstractWith the lowering costs of sensors, high-volume and high-velocity data are increasingly being generated and analyzed, especially in IoT domains like energy and smart homes. Consequently, applications that require accurate short-term forecasts and predictions are also steadily increasing. In this paper, we provide an overview of a novel end-to-end platform that provides efficient ingestion, compression, transfer, query processing, and machine learning-based analytics for high-frequency and high-volume time series from IoT. The performance of the platform is evaluated using real-world dataset from RES installations. The results show the importance of high-frequency analytics and the surprisingly positive impact of error bounded lossy compression on machine learning in the form of AutoML. For example, when detecting yaw misalignments in wind turbines, an improvement of 9% in accuracy was observed for AutoML models on lossy compressed data compared to the current industry standard of 10-minute aggregated data. Thus, these small-scale experiments show the potential of the platform, and larger pilots are planned. Seshu Tirupathi, Dhaval Salwala, Giulio Zizzo, Ambrish Rawat, Mark Purcell, Søren Kejser Jensen, Christian Thomsen 0001, Nguyen Ho, Carlos Muñiz Cuza, Jonas Brusokas, Torben Bach Pedersen, George Alexiou, Giorgos Giannopoulos, Panagiotis Gidarakos, Alexandros Kalimeris, Stavros Maroulis, George Papastefanatos, Ioannis Psarros, Vassilis Stamatopoulos, Manolis Terrovitis |
IEEE Big Data | 5 |
| 2020 | Knowledge- and Data-driven Services for Energy Systems using Graph Neural NetworksabstractThe transition away from carbon-based energy sources poses several challenges for the operation of electricity distribution systems. Increasing shares of distributed energy resources (e.g. renewable energy generators, electric vehicles) and internet-connected sensing and control devices (e.g. smart heating and cooling) require new tools to support accurate, data-driven decision making. Modelling the effect of such growing complexity in the electrical grid is possible in principle using state-of-the-art power-power flow models. In practice, the detailed information needed for these physical simulations may be unknown or prohibitively expensive to obtain. Hence, data-driven approaches to power systems modelling, including feed-forward neural networks and auto-encoders, have been studied to leverage the increasing availability of sensor data, but have seen limited practical adoption due to lack of transparency and inefficiencies on large-scale problems. Our work addresses this gap by proposing a data- and knowledge-driven probabilistic graphical model for energy systems based on the framework of graph neural networks (GNNs). The model can explicitly factor in domain knowledge, in the form of grid topology or physics constraints, thus resulting in sparser architectures and much smaller parameters dimensionality when compared with traditional machine-learning models with similar accuracy. Results obtained from a real-world smart-grid demonstration project show how the GNN was used to inform grid congestion predictions and market bidding services for a distribution system operator participating in an energy flexibility market. Francesco Fusco, Bradley Eck, Robert Gormally, Mark Purcell, Seshu Tirupathi |
IEEE BigData | 4 |