VLDB 2026 Research / reviewers in the wild / expert
Waleed A. Yousef
dblp:26/1956
· DBLP profile ↗
11ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-9669-7241ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 3 since 2021Security and privacy · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Preserving data and model privacy during inference and training
William Briguglio, Issa Traoré, Mohammad Saiful Islam Mamun, Waleed A. Yousef, Sherif Saad |
Expert Syst. Appl. | 4 |
| 2025 | An Alternative Approach to Federated Learning for Model Security and Data PrivacyabstractFederated learning (FL) enables machine learning on data held across multiple clients without exchanging private data. However, exchanging information for model training can compromise data privacy. Further, participants may be untrustworthy and can attempt to sabotage model performance. Also, data that is not independently and identically distributed (IID) impede the convergence of FL techniques. We present a general framework for federated learning via aggregating multivariate estimated densities (FLAMED). FLAMED aggregates density estimations of clients’ data, from which it simulates training datasets to perform centralized learning, bypassing problems arising from non-IID data and contributing to addressing privacy and security concerns. FLAMED does not require a copy of the global model to be distributed to each participant during training, meaning the aggregating server can retain sole proprietorship of the global model without the use of resource-intensive homomorphic encrypti on. We compared its performance to standard FL approaches using synthetic and real datasets and evaluated its resilience to model poisoning attacks. Our results indicate that FLAMED effectively handles non-IID data in many settings while also being more secure. William Briguglio, Waleed A. Yousef, Issa Traoré, Mohammad Saiful Islam Mamun, Sherif Saad |
ICISSP (1) | 2 |
| 2024 | Federated Supervised Principal Component AnalysisabstractIn federated learning, standard machine learning (ML) techniques are modified so they can be applied to data held by separate participants without the need for exchanging said data and while preserving privacy. Other data modelling techniques, such as singular value decomposition, have been similarly federated, enabling federated principal component analysis (PCA), which is a popular preprocessing step for ML tasks. Supervised PCA improves on standard PCA by using labeled data to retain more relevant information for supervised ML problems. However, a federated version of supervised PCA does not exist in the literature. In this paper, we propose a federated version of supervised PCA and its dual and kernel variations, called FeS-PCA, dual FeS-PCA, and FeSK-PCA, respectively. We used random orthogonal matrix masking to keep FeS-PCA and dual FeS-PCA private, while FeSK-PCA was kept private using an approximation of the standard approach. We tested our proposed approaches by recreating visualization, classification, and regression experiments from the original unfederated supervised PCA paper. We further added a real-world federated dataset to test the scalability and fidelity of our approach. Our analysis and results indicate that FeS-PCA and dual FeS-PCA are faithful, lossless, and private versions of their unfederated counterparts. Furthermore, despite being an approximation, FeSK-PCA achieves nearly identical performance to standard kernel SPCA in many cases. This is in addition to the added benefit of a reduced runtime and smaller memory footprint. William Briguglio, Waleed A. Yousef, Issa Traoré, Mohammad Saiful Islam Mamun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Classifier Calibration: With Application to Threat Scores in CybersecurityabstractThis article explores the calibration of a classifier output score in binary classification problems. A calibrator is a function that maps the arbitrary classifier score, of a testing observation, onto [0,1] to provide an estimate for the posterior probability of belonging to one of the two classes. Calibration is important for two reasons; first, it provides a meaningful score, that is the posterior probability; second, it puts the scores of different classifiers on the same scale for comparable interpretation. The article presents three main contributions: (1) Introducing multi-score calibration, when more than one classifier provides a score for a single observation. (2) Introducing the exact analogy between two scenarios: (a) designing a classifier from a set of features, and (b) designing a calibrator, to generate a single calibrated score, from a set of scores of different classifiers. Hence, we propose expanding these classifiers’ scores to higher dimensions to boost the calibrator’s performance. (3) Conducting a massive simulation study, in the order of 24,000 experiments, that incorporates different configurations, in addition to experimenting on three real datasets from the cybersecurity domain. The results show that there is no overall winner among the different calibrators and different configurations. However, general advices for practitioners include the following: the Platt’s calibrator (J. Plattet al., 1999), a version of the logistic regression that decreases bias for a small sample size, has a very stable and acceptable performance among all experiments; our suggested multi-score calibration provides better performance than single score calibration in the majority of experiments, including the two real datasets. In addition, expanding the scores can help in some experiments. Waleed A. Yousef, Issa Traoré, William Briguglio |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2021 | Machine learning in precision medicine to preserve privacy via encryption
William Briguglio, Parisa Moghaddam, Waleed A. Yousef, Issa Traoré, Mohammad Saiful Islam Mamun |
Pattern Recognit. Lett. | 3 |
| 2021 | Estimating the standard error of cross-Validation-Based estimators of classifier performance
Waleed A. Yousef |
Pattern Recognit. Lett. | 1 |
| 2021 | UN-AVOIDS: Unsupervised and Nonparametric Approach for Visualizing Outliers and Invariant Detection ScoringabstractThe visualization and detection of anomalies (outliers) are of crucial importance to many fields, particularly cybersecurity. Several approaches have been proposed in these fields, yet to the best of our knowledge, none of them has fulfilled both objectives, simultaneously or cooperatively, in one coherent framework. Moreover, the visualization methods of these approaches were introduced for explaining the output of a detection algorithm, not for data exploration that facilitates a standalone visual detection. This is our point of departure in introducing UN-AVOIDS, an unsupervised and nonparametric approach for both visualization (a human process) and detection (an algorithmic process) of outliers, that assigns invariant anomalous scores (normalized to [0,1]), rather than hard binary-decision. The main aspect of novelty of UN-AVOIDS is that it transforms data into a new space, which is introduced in this paper as neighborhood cumulative density function (NCDF), in which both visualization and detection are carried out. In this space, outliers are remarkably visually distinguishable, and therefore the anomaly scores assigned by the detection algorithm achieved a high area under the ROC curve (AUC). We assessed UN-AVOIDS on both simulated and two recently published cybersecurity datasets, and compared it to three of the most successful anomaly detection methods: LOF, IF, and FABOD. In terms of AUC, UN-AVOIDS was almost an overall winner with a margin that varied between - 0.028 and 0.125, depending on the data. The article concludes by providing a preview of new theoretical and practical avenues for UN-AVOIDS. Among them is designing a visualization aided anomaly detection (VAAD), a type of software that aids analysts by providing UN-AVOIDS’ detection algorithm (running in a back engine), NCDF visualization space (rendered to plots), along with other conventional methods of visualization in the original feature space, all of which are linked in one interactive environment. Waleed A. Yousef, Issa Traoré, William Briguglio |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Prudence when assuming normality: An advice for machine learning practitioners
Waleed A. Yousef |
Pattern Recognit. Lett. | 1 |
| 2012 | Classifier variability: Accounting for training and testing
Weijie Chen 0007, Brandon D. Gallas, Waleed A. Yousef |
Pattern Recognit. | 3 |
| 2006 | Assessing Classifiers from Two Independent Data Sets Using ROC Analysis: A Nonparametric ApproachabstractThis paper considers binary classification. We assess a classifier in terms of the Area Under the ROC Curve (AUC). We estimate three important parameters, the conditional AUC (conditional on a particular training set) and the mean and variance of this AUC. We derive, as well, a closed form expression of the variance of the estimator of the AUC. This expression exhibits several components of variance that facilitate an understanding for the sources of uncertainty of that estimate. In addition, we estimate this variance, i.e., the variance of the conditional AUC estimator. Our approach is nonparametric and based on general methods from U-statistics; it addresses the case where the data distribution is neither known nor modeled and where there are only two available data sets, the training and testing sets. Finally, we illustrate some simulation results for these estimators. Waleed A. Yousef, Robert F. Wagner, Murray H. Loew |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Estimating the uncertainty in the estimated mean area under the ROC curve of a classifier
Waleed A. Yousef, Robert F. Wagner, Murray H. Loew |
Pattern Recognit. Lett. | 1 |