VLDB 2026 Research / reviewers in the wild / expert
Alec F. Diallo
dblp:298/4347
· DBLP profile ↗
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-0793-0492ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ECHO: Effective Coreset-Driven Learning via Hierarchical OptimizationsabstractDespite driving record performance, the increasing reliance of deep learning on ever-larger datasets has led to prohibitively high storage and management costs that threaten continued progress. While coreset selection offers a promising solution to this challenge, existing methods often rely on expensive iterative optimization procedures or fail to select samples that allow strong generalization across tasks. In this work, we introduce ECHO, a coreset construction and augmentation strategy that leverages the relational properties inherent to a dataset to find its most representative samples. Unlike prior methods, our approach constructs a structured graph that encodes intrinsic dataset patterns, based on which influential samples are identified and augmented to maximize generalization performance. Extensive experiments across five benchmark datasets and against eighteen different coreset selection baselines show that ECHO achieves up to 60% accuracy gains under extreme compression, while being orders of magnitude faster than state-of-the-art alternatives. These results establish a new benchmark for data-efficient learning, particularly under tight coreset budgets, and showcase the benefits of structured coreset selection for effective generalization. Alec F. Diallo, Weihe Li, Paul Patras |
ICDM | 1 |
| 2025 | Pontus: A Memory-Efficient and High-Accuracy Approach for Persistence-Based Item Lookup in High-Velocity Data StreamsabstractIn today's web-scale, data-driven environments, real-time detection of persistent items that consistently recur over time is essential for maintaining system integrity, reliability, and security. Persistent items often signal critical anomalies, such as stealthy DDoS and botnet attacks in web infrastructures. Although various methods exist for identifying such items as well as for determining their frequency, they require recording every item for processing, which is impractical at very high data rates achieved by modern data streams. In this paper, we introduce Pontus, a novel approach that uses an approximate data structure (sketch) specifically designed for the efficient and accurate detection of persistent items. Our method not only achieves fast and precise lookup but is also flexible, allowing for minor modifications to accommodate other types of persistence-based item detection tasks, such as detecting persistent items with low frequency. We rigorously validate our approach through formal methods, offering detailed proofs of time/space complexity and error bounds to demonstrate its theoretical soundness. Our extensive trace-driven evaluations across various persistence-based tasks further demonstrate Pontus's effectiveness in significantly improving detection accuracy and enhancing processing speed compared to existing approaches. We implement Pontus in an experimental platform with industry-grade Intel Tofino switches and demonstrate the practical feasibility of our approach in a real-world memory-constrained environment. Weihe Li, Zukai Li, Beyza Bütün, Alec F. Diallo, Marco Fiore 0001, Paul Patras |
WWW | 4 |
| 2024 | Sabre: Cutting through Adversarial Noise with Adaptive Spectral Filtering and Input ReconstructionabstractThe adoption of neural networks (NNs) across critical sectors including transportation, medicine, communications infrastructure, etc. is inexorable. However, NNs remain highly susceptible to adversarial perturbations, whereby seemingly minimal or imperceptible changes to their inputs cause gross misclassifications, which questions their practical use. Although a growing body of work focuses on defending against such attacks, adversarial robustness remains an open challenge, especially as the effectiveness of existing solutions against increasingly sophisticated input manipulations comes at the cost of degrading ability to recognize benign samples, as we reveal. In this work we introduce Sabre, an adversarial defense framework that closes the gap between benign and robust accuracy in NN classification tasks, without sacrificing benign sample recognition performance. In particular, through spectral decomposition of the input and selective energy-based filtering, Sabre extracts robust features that serve in input reconstruction prior to feeding existing NN architectures. We demonstrate the performance of our approach across multiple domains, by evaluating it on image classification, network intrusion detection, and speech command recognition tasks, showing that Sabre not only outperforms existing defense mechanisms, but also behaves consistently with different neural architectures, data types, (un)known attacks, and adversarial perturbation strengths. Through these extensive experiments, we make the case for Sabre’s adoption in deploying robust and reliable neural classifiers. Alec F. Diallo, Paul Patras |
SP | 1 |
| 2024 | Deciphering Clusters With a Deterministic Measure of Clustering TendencyabstractClustering, a key aspect of exploratory data analysis, plays a crucial role in various fields such as information retrieval. Yet, the sheer volume and variety of available clustering algorithms hinder their application to specific tasks, especially given their propensity to enforce partitions, even when no clear clusters exist, often leading to fruitless efforts and erroneous conclusions. This issue highlights the importance of accurately assessing clustering tendencies prior to clustering. However, existing methods either rely on subjective visual assessment, which hinders automation of downstream tasks, or on correlations between subsets of target datasets and random distributions, limiting their practical use. Therefore, we introduce theProximal Homogeneity Index (PHI), a novel and deterministic statistic that reliably assesses the clustering tendencies of datasets by analyzing their internal structures via knowledge graphs. Leveraging PHI and the boundaries between clusters, we establish thePartitioning Sensitivity Index (PSI), a new statistic designed for cluster quality assessment and optimal clustering identification. Comparative studies using twelve synthetic and real-world datasets demonstrate PHI and PSI's superiority over existing metrics for clustering tendency assessment and cluster validation. Furthermore, we demonstrate the scalability of PHI to large and high-dimensional datasets, and PSI's broad effectiveness across diverse cluster analysis tasks. Alec F. Diallo, Paul Patras |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Cluster and Conquer: Malicious Traffic Classification at the EdgeabstractThe uptake of digital services and IoT technology gives rise to increasingly diverse cyber attacks, with which commonly-used rule-based Network Intrusion Detection Systems (NIDSs) struggle to cope. Therefore, Artificial Intelligence (AI) supports a second line of defense, since this methodology helps in extracting non-obvious patterns from network traffic and subsequently in detecting more confidently new types of threats. Cybersecurity is however an arms race and intelligent solutions face renewed challenges as attacks evolve while network traffic volumes surge. We propose Adaptive Clustering-based Intrusion Detection (ACID), a novel approach to malicious traffic classification and a valid candidate for deployment at the network edge. ACID addresses the critical challenge of sensitivity to subtle changes in traffic features, which routinely leads to misclassification. We circumvent this problem by relying on low-dimensional embeddings learned with a lightweight neural model comprising multiple kernel networks that we introduce, which optimally separates samples of different classes. Extensive experiments with datasets spanning 20 years demonstrate ACID attains 100% accuracy and F1-score, and 0% false alarm rate, significantly outperforming state-of-the-art clustering methods and NIDSs. Furthermore, our results show that ACID offers a high degree of robustness to input perturbations, while intrinsically providing a framework for continual learning. Alec F. Diallo, Paul Patras |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2021 | Spider: Deep Learning-driven Sparse Mobile Traffic Measurement Collection and ReconstructionabstractData-driven mobile network management hinges on accurate traffic measurements, which routinely require expensive specialized equipment and substantial local storage capabilities, and bear high data transfer overheads. To overcome these challenges, in this paper we propose Spider, a deep-learning-driven mobile traffic measurement collection and reconstruction framework, which reduces the cost of data collection while retaining state-of-the-art accuracy in inferring mobile traffic consumption with fine geographic granularity. Spider harnesses Reinforcement Learning and tackles large action spaces to train a policy network that selectively samples a minimal number of cells where data should be collected. We further introduce a fast and accurate neural model that extracts spatiotemporal correlations from historical data to reconstruct network-wide traffic consumption based on sparse measurements. Experiments we conduct with a real-world mobile traffic dataset demonstrate that Spider samples 48% fewer cells as compared to several benchmarks considered, and yields up to 67% lower reconstruction errors than state-of-the-art interpolation methods. Moreover, our framework can adapt to previously unseen traffic patterns. Yini Fang, Alec F. Diallo, Chaoyun Zhang, Paul Patras |
GLOBECOM | 2 |
| 2021 | Adaptive Clustering-based Malicious Traffic Classification at the Network EdgeabstractThe rapid uptake of digital services and Internet of Things (IoT) technology gives rise to unprecedented numbers and diversification of cyber attacks, with which commonly-used rule-based Network Intrusion Detection Systems (NIDSs) are struggling to cope. Therefore, Artificial Intelligence (AI) is being exploited as second line of defense, since this methodology helps in extracting non-obvious patterns from network traffic and subsequently in detecting more confidently new types of threats. Cybersecurity is however an arms race and intelligent solutions face renewed challenges as attacks evolve while network traffic volumes surge. In this paper, we propose Adaptive Clustering-based Intrusion Detection (Acid), a novel approach to malicious traffic classification and a valid candidate for deployment at the network edge. Acid addresses the critical challenge of sensitivity to subtle changes in traffic features, which routinely leads to misclassification. We circumvent this problem by relying on low-dimensional embeddings learned with a lightweight neural model comprising multiple kernel networks that we introduce, which optimally separates samples of different classes. We empirically evaluate our approach with both synthetic and three intrusion detection datasets spanning 20 years, and demonstrate Acid consistently attains 100% accuracy and F1-score, and 0% false alarm rate, thereby significantly outperforming state-of-the-art clustering methods and NIDSs. Alec F. Diallo, Paul Patras |
INFOCOM | 1 |