EDBT 2026 Demo / reviewers in the wild / expert
Félix Iglesias
dblp:18/9197 · also Félix Iglesias Vázquez
· DBLP profile ↗
26ranked-venue papers
21as first author
9since 2021 · last 2027
0000-0001-6081-969XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 9 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 5 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 first-authorSecurity and privacy · 3 · 2 first-authorTheory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | On heterogeneous ensembles for anomaly detection: Empirical insights and guidelines for the design
Félix Iglesias, Tanja Zseby, Conrado Martínez, Arthur Zimek |
Expert Syst. Appl. | 1 |
| 2025 | Stream Clustering Robust to Concept DriftabstractData streams are everywhere in modern technologies, spanning from industrial process control to network traffic analysis. Stream clustering is required to describe data streams in real time and maintain accurate knowledge of their underlying structures. However, data streams frequently exhibit non-stationarity, changes in distributions, and the emergence of new classes. These alterations—commonly referred to as "concept drift"—severely disturb algorithms, resulting in inconsistent outcomes and models. We present SDOstreamclust, an incremental algorithm for stream clustering. It inherits the distinctive features of methods founded on Sparse Data Observers, i.e., lightweight, intuitive, self-adjusting, resistant to noise, capable of identifying non-convex clusters, and constructed upon robust parameters and interpretable models. We compare SDOstreamclust with established algorithms and evaluate them with a broad collection of datasets, both real and synthetic. SDOstreamclust shows outstanding performances, a major adaptability to concept drift, and a superior parameter stability and robustness. Often ignored in the evaluation of new methods, concept drift is a major challenge for next-generation algorithms, since it is inherent to evolving data and a main cause of degradation in machine learning. Hence, SDOstreamclust emerges as a major alternative for unsupervised streaming data analysis. Félix Iglesias, Simon Konzett, Tanja Zseby, Albert Bifet |
IJCNN | 1 |
| 2025 | What do anomaly scores actually mean? Dynamic characteristics beyond accuracyabstractAbstract Anomaly detection has become pervasive in modern technology, covering applications from cybersecurity, to medicine or system failure detection. Before outputting a binary outcome (i.e., anomalous or non-anomalous), most algorithms evaluate instances with outlierness scores. But what does a score of 0.8 mean? Or what is the practical difference compared to a score of 1.2? Score ranges are assumed non-linear and relative, their meaning established by weighting the whole dataset (or a dataset model). While this is perfectly true, algorithms also impose dynamics that decisively affect the meaning of outlierness scores. In this work, we aim to gain a better understanding of the effect that both algorithms and specific data particularities have on the meaning of scores. To this end, we compare established outlier detection algorithms and analyze them beyond common metrics related to accuracy. We disclose trends in their dynamics and study the evolution of their scores when facing changes that should render them invariant. For this purpose we abstract characteristic S-curves and propose indices related to discriminant power, bias, variance, coherence and robustness. We discovered that each studied algorithm shows biases and idiosyncrasies, which habitually persist regardless of the dataset used. We provide methods and descriptions that facilitate and extend a deeper understanding of how the discussed algorithms operate in practice. This information is key to decide which one to use, thus enabling a more effective and conscious incorporation of unsupervised learning in real environments. Félix Iglesias, Henrique O. Marques, Arthur Zimek, Tanja Zseby |
Data Min. Knowl. Discov. | 1 |
| 2025 | Parameterization-free clustering with sparse data observersabstractGiven a set of data points, clustering serves to discover groups based on pairwise similarities and the shapes drawn by the data in the feature space. In other words, it is a tool to describe data and reveal their intrinsic nature in terms of patterns or groups. In this paper, we review the methodology of clustering when used to explore a priori unknown data, i.e., we do not know how data spaces are manipulated, how algorithms are tuned, and how results are validated. Under this practical approach, we examine the advantages of SDOclust, a clustering method that stands out for its simplicity, lightness, no need for parameterization and not being subject to traditional clustering limitations. We test SDOclust and main established alternatives — HDBSCAN, k-means-, Fuzzy C-means, Hierarchical Clustering, CLASSIX, and N2D Deep Clustering — by extensive experimentation with more than 200 datasets, both real and synthetic, that have been collected from the literature on evaluation and represent different data analysis challenges. We submit only SDOclust to unfavorable testing conditions by denying it a parameter tuning phase. Nevertheless, its overall performance is excellent and positions it as one of the best general-purpose alternatives. With deep clustering as the consolidation of a new paradigm, trends in clustering consist mainly in projecting data into spaces that are easier to dissect. Therefore, in cases where the original space does not show clustering-friendly structures and when we can assume transformation costs, SDOclust easily adapts and is a most natural choice to perform the partitioning task. Félix Iglesias, Tanja Zseby, Arthur Zimek |
Inf. Syst. | 1 |
| 2024 | Impact of the Neighborhood Parameter on Outlier Detection Algorithms
Félix Iglesias, Conrado Martínez, Tanja Zseby |
SISAP | 1 |
| 2024 | Temporal silhouette: validation of stream clustering robust to concept driftabstractAbstract Stream clustering is required in applications where data is generated continuously or periodically and must be processed considering its temporal nature. In the absence of a ground truth, internal validation is the only option to evaluate the quality of performances. Traditional internal validation is commonly used also in stream clustering, even in spite of the fact that it becomes inconsistent in the event of data evolution. Recent trends opt for incremental approaches, but these are closer to change detection rather than validation methods and limit themselves by imposing online validation on online analysis. In this work we study the impact of concept drift in the validation of stream clustering and propose the Temporal Silhouette index, therefore making internal validation conform to streaming data. We conduct tests with more than 200 datasets and contrast performances of four popular stream clustering algorithms with seven validation methods (three static internal, three incremental internal, one external) and the proposed index. Results show the suitability of the Temporal Silhouette index for stream clustering validation in the event of concept drift and different types of outliers. The demand for reliable unsupervised learning in applications that process data in streams is ever-increasing, and such reliability inevitably requires the use of validation. This fact highlights the significance of the novel approach proposed in this work. Félix Iglesias, Tanja Zseby |
Mach. Learn. | 1 |
| 2023 | SDOclust: Clustering with Sparse Data Observers
Félix Iglesias, Tanja Zseby, Alexander Hartl, Arthur Zimek |
SISAP | 1 |
| 2023 | Anomaly detection in streaming data: A comparison and evaluation studyabstractThe detection of anomalies in streaming data faces complexities that make traditional static methods unsuitable due to computational costs and nonstationarity. We test and evaluate eight state of the art algorithms against prominent challenges related to streaming data. Results show insights regarding accuracy, memory-dependency, parameterization, and pre-knowledge exploitation, thus revealing the high impact of some data characteristics to establish a most appropriate algorithm—namely: locality (i.e., whether outlierness is relative to local contexts), relativeness (i.e., if past data defines outlierness), and concept drift (if it is expected, its intensity and frequency). In most applied cases, such factors can be inferred in advance through the use of historical data and domain knowledge. Assuming the viability of the studied methods in terms of time efficiency, this work discloses key findings to achieve optimal designs of streaming data anomaly detection in real-life applications. Félix Iglesias, Alexander Hartl, Tanja Zseby, Arthur Zimek |
Expert Syst. Appl. | 1 |
| 2022 | Modeling data with observersabstractCompact data models have become relevant due to the massive, ever-increasing generation of data. We propose Observers-based Data Modeling (ODM), a lightweight algorithm to extract low density data models (aka coresets) that are suitable for both static and stream data analysis. ODM coresets keep data internal structures while alleviating computational costs of machine learning during evaluation phases accounting for a O(n log n) worst-case complexity. We compare ODM with previous proposals in classification, clustering, and outlier detection. Results show the preponderance of ODM for obtaining the best trade-off in accuracy, versatility, and speed. Fares Meghdouri, Félix Iglesias, Tanja Zseby |
Intell. Data Anal. | 2 |
| 2020 | Cross-Layer Profiling of Encrypted Network Data for Anomaly DetectionabstractIn January 2017 encrypted Internet traffic surpassed non-encrypted traffic. Although encryption increases security, it also masks intrusions and attacks by blocking the access to packet contents and traffic features, therefore making data analysis unfeasible. In spite of the strong effect of encryption, its impact has been scarcely investigated in the field. In this paper we study how encryption affects flow feature spaces and machine learning-based attack detection. We propose a new cross-layer feature vector that simultaneously represents traffic at three different levels: application, conversation, and endpoint behavior. We analyze its behavior under TLS and IPSec encryption and evaluate the efficacy with recent network traffic datasets and by using Random Forests classifiers. The cross-layer multi-key approach shows excellent attack detection in spite of TLS encryption. When IPsec is applied, the reduced variant obtains satisfactory detection for botnets, yet considerable performance drops for other types of attacks. The high complexity of network traffic is unfeasible for monolithic data analysis solutions, therefore requiring cross-layer analysis for which the multi-key vector becomes a powerful profiling core. Fares Meghdouri, Félix Iglesias, Tanja Zseby |
DSAA | 2 |
| 2020 | Interpretability and Refinement of ClusteringabstractThe difficulty to validate clustering reliability hinders the adoption of clustering in real-life applications. We propose: (a) a set of symbolic representations to interpret problem spaces and (b) the CluReAL algorithm to refine any clustering result regardless of the used technique. Both approaches are grounded by recently published absolute cluster validity indices. Conducted experiments show how the refinement algorithm improves performances in a wide variety of scenarios and builds more interpretable solutions, whereas symbolic representations are shown to offer explainable summaries of problem contexts. Refinement and interpretability are both crucial to reduce failure and increase performance control and operational awareness in processes that depend on clustering. Félix Iglesias, Tanja Zseby, Arthur Zimek |
DSAA | 1 |
| 2020 | SDOstream: Low-Density Models for Streaming Outlier Detection
Alexander Hartl, Félix Iglesias, Tanja Zseby |
ESANN | 2 |
| 2020 | Absolute Cluster ValidityabstractThe application of clustering involves the interpretation of objects placed in multi-dimensional spaces. The task of clustering itself is inherently submitted to subjectivity, the optimal solution can be extremely costly to discover and sometimes even unreachable or nonexistent. This fact introduces a trade-off between accuracy and computational effort, moreover given that engineering applications usually work well with suboptimal solutions. In such applied scenarios, cluster validation is mandatory to refine algorithms and ensure that solutions are meaningful. Validity indices are commonly intended to benchmark diverse clustering setups, therefore they are coefficients with a relative nature, i.e., useful when compared to one another. In this paper, we propose a validation methodology that enables absolute evaluations of clustering results. Our method performs geometric measurements of the solution space and provides a coherent interpretation of the data structure by using indices based on inter- and intra-cluster distances, density, and multimodality within clusters. Conducted tests and comparisons with well-known indices show that our validation methodology improves the robustness of the clustering application for knowledge discovery. While clustering is often performed as a black box technique, our index is construable and therefore allows for the implementation of systems enriched with self-checking capabilities. Félix Iglesias, Tanja Zseby, Arthur Zimek |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Extreme Dimensionality Reduction for Network Attack Visualization with AutoencodersabstractThe visualization of network traffic flows is an open problem that affects the control and administration of communication networks. Feature vectors used for representing traffic commonly have from tens to hundreds of dimensions and hardly tolerate visual conceptualizations. In this work we use neural networks to obtain extremely low-dimensional data representations that are meaningful from an attack-detection perspective. We focus on a simple Autoencoder architecture, as well as an extension that benefits from pre-knowledge, and evaluate their performances by comparing them with reductions based on Principal Component Analysis and Linear Discriminant Analysis. Experiments are conducted with a modern Intrusion Detection dataset that collects legitimate traffic mixed with a wide variety of attack classes. Results show that feature spaces can be strongly reduced up to two dimensions with tolerable classification degradation while providing a clear visualization of the data. Visualizing traffic flows in two-dimensional spaces is extremely useful to understand what is happening in networks, also to enhance and refocus classification, trigger refined analysis, and aid the security experts' decision-making. We additionally developed a tool prototype that covers such functions, therefore supporting the optimization of network traffic attack detectors in both design and application phases. Daniel C. Ferreira, Félix Iglesias, Tanja Zseby |
IJCNN | 2 |
| 2019 | Fuzzy classification boundaries against adversarial network attacks
Félix Iglesias, Jelena Milosevic, Tanja Zseby |
Fuzzy Sets Syst. | 1 |
| 2019 | Pattern Discovery in Internet Background RadiationabstractInternet Background Radiation (IBR) is observed in empty network address spaces. No traffic should arrive there, but it does in overwhelming quantities, gathering evidences of attacks, malwares and misconfigurations. The study of IBR helps to detect spreading network problems, common vulnerabilities and attack trends. However, network traffic data evolves quickly and is of high volume and diversity, i.e., an outstanding big data challenge. When used to assist network security, it also requires the online classification of dynamic streaming data. In this paper, we introduce an AGgregation & Mode (AGM) vector to represent network traffic. The AGM format characterizes IP hosts by extracting aggregated and mode values of IP header fields, and without inspecting payloads. We performed clustering and statistical analysis to explore six months of IBR from 2012 with the AGM mapping. The discovered patterns allow building a classification of IBR, which identifies phenomena that have been actively polluting the Internet for years. The AGM representation is light and tailored for monitoring and pattern discovery. We show that AGM vectors are suitable to analyze large volumes of network traffic: they capture permanent operations, such as long term scanning, as well as bursty events from targeted attacks and short term incidents. Félix Iglesias, Tanja Zseby |
IEEE Trans. Big Data | 1 |
| 2017 | Are Network Covert Timing Channels Statistical Anomalies?abstractCovert channels exploit communication protocols to clandestinely transfer information. They enable criminals to hide malicious activities and can be used for secret data exfiltration, malware spreading or for the stealthy establishment of command and control structures. In this paper we study covert timing channels from a statistical perspective and investigate whether they can be identified as anomalies with unsupervised learning methods. We use a testbed to generate covert timing channels based on seven popular techniques and inject them in real captured traffic. Final datasets are analyzed with diverse outlier detection and classification algorithms. Our results show that, based on their statistical properties, covert channels do not occupy low density regions or take extreme values in the problem space, and therefore are not detectable as strong anomalies. However, they present traceable profiles that can be abstracted by supervised learning models. Such findings reveal that facing the detection of novel (and classic) covert timing channels from an anomaly-detection perspective will probably fail or not suffice; instead, they must be identified based on the similarity to known schemes, using supervised and semi-supervised approaches. Félix Iglesias, Tanja Zseby |
ARES | 1 |
| 2017 | Decision Tree Rule Induction for Detecting Covert Timing Channels in TCP/IP Traffic
Félix Iglesias, Valentin Bernhardt, Robert Annessi, Tanja Zseby |
CD-MAKE | 1 |
| 2016 | Time-activity footprints in IP traffic
Félix Iglesias, Tanja Zseby |
Comput. Networks | 1 |
| 2016 | DAT detectors: uncovering TCP/IP covert channels by descriptive analyticsabstractAbstract Covert channels provide means to conceal information transfer between hosts and bypass security barriers in communication networks. Hidden communication is of paramount concern for governments and companies, because it can conceal data leakage and malware communication, which are crucial building blocks used in cyber crime. We propose detectors based on descriptive analytics of traffic (DAT) to facilitate revealing network and transport layer covert channels originated from a wide spectrum of published data‐hiding techniques. DAT detectors transform communication data into flexible feature vectors that represent traffic by a set of extracted calculations and estimations. For the case of covert channels, the core of the detection is performed by the combined application of autocorrelation calculations and multimodality measures built upon kernel density estimations and Pareto charts. DAT detectors are devised to be embedded as extensions of network intrusion detection systems, being able to perform fast, lightweight analysis of numerous flows. The present paper focuses specifically on TCP/IP traffic and provides suitable classifications of TCP/IP fields and related covert channel techniques from the perspective of the statistical detection. The proposed methodology is evaluated with public traffic datasets as well as covert channels generated according to main techniques described in the related literature. Copyright © 2016 John Wiley & Sons, Ltd. Félix Iglesias, Robert Annessi, Tanja Zseby |
Secur. Commun. Networks | 1 |
| 2016 | Crucial pitfall of DPA Contest V4.2 implementationabstractAbstract Differential power analysis (DPA) is a powerful side‐channel key recovery attack that efficiently breaks cryptographic algorithm implementations. In order to prevent these types of attacks, hardware designers and software programmers make use of masking and hiding techniques. DPA contest is an international framework that allows researchers to compare their power analysis attacks under the same conditions. The latest version of DPA contest, denoted as V4.2, provides an improved implementation of the rotating S‐box masking scheme where low‐entropy boolean masking is combined with the shuffling technique to protect Advanced Encryption Standard implementation on a smart card. The improvements were designed based on the awareness of implementation lacks analyzed from attacks carried out during the previous DPA contest V4. Therefore, this new approach is devised to resist most of the proposed attacks to the original rotating S‐box masking implementation. In this paper, we investigate the security of this new implementation in practice. Our analysis, focused on exploiting the first‐order leakage, discovered important lacks. The main vulnerability observed is that an adversary can mount a standard DPA attack aimed at the S‐box output in order to recover the whole secret key even when a shuffling technique is used. We tested this observation on a public dataset and implemented a successful attack that revealed the secret key using only 35 power traces. Copyright © 2017 John Wiley & Sons, Ltd. Zdenek Martinasek, Félix Iglesias, Lukas Malina, Josef Martinasek |
Secur. Commun. Networks | 2 |
| 2015 | Analysis of network traffic features for anomaly detection
Félix Iglesias, Tanja Zseby |
Mach. Learn. | 1 |
| 2014 | Profile-Based Control for Central Domestic Hot Water DistributionabstractA main goal of hot water distribution research is to improve the system's efficiency, i.e., to fulfill hot water requirements while minimizing energy and water losses. Central domestic hot water (CDHW) systems represent an important part of current installations worldwide, e.g., hotels, hospitals, sports centers, social facilities, and multifamily residential or apartment buildings. The optimization of such systems claims for forecasting capabilities and context-aware enhancements are based on patterns of use. Thus, the level of uncertainty is reduced, and systems are not forced to operate using blind/oversized/generic assumptions. This paper presents a novel control strategy based on habit profiles for the management of a CDHW system. A simulated environment is utilized to compare the introduced strategy with habitual performances. Simulations are supported by real databases concerning users' behavioral patterns. Results are promising and point to place profile-based strategies as a suitable approach for an optimized water and energy management in future buildings. Félix Iglesias, Peter Palensky |
IEEE Trans. Ind. Informatics | 1 |
| 2012 | Detecting user dissatisfaction in ambient intelligence environmentsabstractOne of the reasons that explains the slow uptake of smart home technologies has been addressed several times due to a lack of designs better aware of usability and the real context when families experience their daily life. The present paper introduces a system model for ambient intelligence (AmI) environments that boosts a friendly system-user's adaptation. To that end, system self-checking capabilities based on the detection of levels of user dissatisfaction or disagreement are developed. Félix Iglesias, Wolfgang Kastner |
ETFA | 1 |
| 2011 | Impact of user habits in smart home controlabstractLifestyle and habits of users have a direct effect on the energy performance of dwellings and facilities. Hence, in the built environment, advanced control strategies must adapt to user behaviors trying to keep a commitment between energy consumption and comfort requirements. In previous works, the suitability of predictive control based on occupancy profiles for the optimization of HVAC systems has been shown. Resting upon this basis, the present work performs a sensitivity analysis for control strategies based on usage profiles, where the input under variation is the level of habit-regularity of users. Therefore, different hypothetical user models are created and tested. The results of this analysis provide a better understanding of how user behavior affects the energy and comfort performance in dwellings under smart control and paves the way for enhanced controller design. Félix Iglesias, Wolfgang Kastner, Christian Reinisch |
ETFA | 1 |
| 2010 | Usage profiles for sustainable buildingsabstractThe ultimate goal of sustainable homes and buildings is to work towards energy efficiency automatically taking into account user comfort always acknowledging the residents' desires. Such environments demand for a friendly coexistence between technology and usability to assure an optimized reality in terms of comfort, economic and energy savings. The ThinkHome project is geared towards this mission. It aims at exploiting automation systems and mechanisms based on artificial intelligence to further improve the sustainability of buildings. ThinkHome is designed in a way to relieve the smart home inhabitants from cumbersome tasks such as readjustments of their preferences by introducing learning capabilities and context awareness in the home. Part of the ThinkHome project, this paper proposes an advanced use case for energy savings. It involves the necessity of defining usage profiles which help the underlying system to be aware about the stability of users' behaviors and perform its strategies always with sustainable goals in mind. Félix Iglesias, Wolfgang Kastner |
ETFA | 1 |