EDBT 2026 Demo / reviewers in the wild / expert
Julián Luengo
dblp:63/6487 · also Julián Luengo-Martín
· DBLP profile ↗
16ranked-venue papers in the field
4as first author
4since 2021 · last 2025
0000-0003-3952-3629ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 9 (2 first)Data Mining & Knowledge Discovery · 4 (2 first)Other / Interdisciplinary · 2Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Developing Big Data anomaly dynamic and static detection algorithms: AnomalyDSD spark packageabstractAnomaly detection is the process of identifying observations that differ greatly from the majority of data. Unsupervised anomaly detection aims to find outliers in data that is not labeled, therefore, the anomalous instances are unknown. The exponential data generation has led to the era of Big Data. This scenario brings new challenges to classic anomaly detection problems due to the massive and unsupervised accumulation of data. Traditional methods are not able to cop up with computing and time requirements of Big Data problems. In this paper, we propose four distributed algorithm designs for Big Data anomaly detection problems: HBOS_BD, LODA_BD, LSCP_BD, and XGBOD_BD. They have been designed following the MapReduce distributed methodology in order to be capable of handling Big Data problems. These algorithms have been integrated into an Spark Package, focused on static and dynamic Big Data anomaly detection tasks, namely AnomalyDSD. Experiments using a real-world case of study have shown the performance and validity of the proposals for Big Data problems. With this proposal, we have enabled the practitioner to efficiently and effectively detect anomalies in Big Data datasets, where the early detection of an anomaly can lead to a proper and timely decision. Diego García-Gil, Daniel Argüelles-Martino, Jacinto Carrasco, Ignacio Aguilera-Martos, Julián Luengo, Francisco Herrera |
Inf. Sci. | 6 |
| 2023 | REVEL Framework to Measure Local Linear Explanations for Black-Box Models: Deep Learning Image Classification Case StudyabstractExplainable artificial intelligence is proposed to provide explanations for reasoning performed by artificial intelligence. There is no consensus on how to evaluate the quality of these explanations, since even the definition of explanation itself is not clear in the literature. In particular, for the widely known local linear explanations, there are qualitative proposals for the evaluation of explanations, although they suffer from theoretical inconsistencies. The case of image is even more problematic, where a visual explanation seems to explain a decision while detecting edges is what it really does. There are a large number of metrics in the literature specialized in quantitatively measuring different qualitative aspects, so we should be able to develop metrics capable of measuring in a robust and correct way the desirable aspects of the explanations. Some previous papers have attempted to develop new measures for this purpose. However, these measures suffer from lack of objectivity or lack of mathematical consistency, such as saturation or lack of smoothness. In this paper, we propose a procedure called REVEL to evaluate different aspects concerning the quality of explanations with a theoretically coherent development which do not have the problems of the previous measures. This procedure has several advances in the state of the art: it standardizes the concepts of explanation and develops a series of metrics not only to be able to compare between them but also to obtain absolute information regarding the explanation itself. The experiments have been carried out on four image datasets as benchmark where we show REVEL’s descriptive and analytical power. Iván Sevillano-García, Julián Luengo, Francisco Herrera |
Int. J. Intell. Syst. | 2 |
| 2021 | Synthetic Sample Generation for Label Distribution Learning
Julián Luengo, José Ramón Cano, Salvador García 0001 |
Inf. Sci. | 2 |
| 2021 | Multiple instance classification: Bag noise filtering for negative instance noise cleaningabstractData in the real world is far from being perfect. The appearance of noise is a common issue that arises from the limitations of data acquisition mechanisms and human knowledge. In classification, label noise will hinder the performance of almost all classifiers, inducing a bias in the built model. While label noise has recently attracted researchers’ attention in standard classification, it has only recently begun to be studied in multiple instance classification. In this work, we propose the usage of filtering algorithms for multiple instance classification that are able to reduce the impact of negative instances within the bags. In order to do so, we decompose the bags to form a standard classification problem that can be efficiently treated by a specialized noise filter. Such a decomposition is tackled in different ways, with the aim of exploiting the knowledge offered by the examples from opposite bags. The bags are then rebuilt, without the identified noise instances. In our experiments, we show that by applying our approach we can diminish the impact of noise and even obtain better results at 0% noise level for several classifiers. Our approach sets out a promising approach to dealing with noise in the bags of multiple instance datasets and further improve the classification rate of the built models. Julián Luengo, Dánel Sánchez Tarragó, Ronaldo C. Prati, Francisco Herrera |
Inf. Sci. | 1 |
| 2020 | Preprocessing methodology for time series: An industrial world application case study
Juan Antonio Cortés-Ibáñez, Sergio González, José Javier Valle-Alonso, Julián Luengo, Salvador García 0001, Francisco Herrera |
Inf. Sci. | 4 |
| 2019 | From Big to Smart Data: Iterative ensemble filter for noise filtering in Big Data classificationabstractThe quality of the data is directly related to the quality of the models drawn from that data. For that reason, many research is devoted to improve the quality of the data and to amend errors that it may contain. One of the most common problems is the presence of noise in classification tasks, where noise refers to the incorrect labeling of training instances. This problem is very disruptive, as it changes the decision boundaries of the problem. Big Data problems pose a new challenge in terms of quality data due to the massive and unsupervised accumulation of data. This Big Data scenario also brings new problems to classic data preprocessing algorithms, as they are not prepared for working with such amounts of data, and these algorithms are key to move from Big to Smart Data. In this paper, an iterative ensemble filter for removing noisy instances in Big Data scenarios is proposed. Experiments carried out in six Big Data datasets have shown that our noise filter outperforms the current state-of-the-art noise filter in Big Data domains. It has also proved to be an effective solution for transforming raw Big Data into Smart Data. Diego García-Gil, Francisco Luque Sánchez, Julián Luengo, Salvador García 0001, Francisco Herrera |
Int. J. Intell. Syst. | 3 |
| 2019 | Enabling Smart Data: Noise filtering in Big Data classification
Diego García-Gil, Julián Luengo, Salvador García 0001, Francisco Herrera |
Inf. Sci. | 2 |
| 2019 | Emerging topics and challenges of learning from noisy data in nonstandard classification: a survey beyond binary class noise
Ronaldo C. Prati, Julián Luengo, Francisco Herrera |
Knowl. Inf. Syst. | 2 |
| 2015 | SMOTE-IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering
José A. Sáez, Julián Luengo, Jerzy Stefanowski, Francisco Herrera |
Inf. Sci. | 2 |
| 2015 | An automatic extraction method of the domains of competence for learning classifiers using data complexity measures
Julián Luengo, Francisco Herrera |
Knowl. Inf. Syst. | 1 |
| 2014 | Analyzing the presence of noise in multi-class problems: alleviating its influence with the One-vs-One decomposition
José A. Sáez, Mikel Galar, Julián Luengo, Francisco Herrera |
Knowl. Inf. Syst. | 3 |
| 2013 | Tackling the problem of classification with noisy data using Multiple Classifier Systems: Analysis of the performance and robustness
José A. Sáez, Mikel Galar, Julián Luengo, Francisco Herrera |
Inf. Sci. | 3 |
| 2013 | A Survey of Discretization Techniques: Taxonomy and Empirical Analysis in Supervised LearningabstractDiscretization is an essential preprocessing technique used in many knowledge discovery and data mining tasks. Its main goal is to transform a set of continuous attributes into discrete ones, by associating categorical values to intervals and thus transforming quantitative data into qualitative data. In this manner, symbolic data mining algorithms can be applied over continuous data and the representation of information is simplified, making it more concise and specific. The literature provides numerous proposals of discretization and some attempts to categorize them into a taxonomy can be found. However, in previous papers, there is a lack of consensus in the definition of the properties and no formal categorization has been established yet, which may be confusing for practitioners. Furthermore, only a small set of discretizers have been widely considered, while many other methods have gone unnoticed. With the intention of alleviating these problems, this paper provides a survey of discretization methods proposed in the literature from a theoretical and empirical perspective. From the theoretical perspective, we develop a taxonomy based on the main properties pointed out in previous research, unifying the notation and including all the known methods up to date. Empirically, we conduct an experimental study in supervised classification involving the most representative and newest discretizers, different types of classifiers, and a large number of data sets. The results of their performances measured in terms of accuracy, number of intervals, and inconsistency have been verified by means of nonparametric statistical tests. Additionally, a set of discretizers are highlighted as the best performing ones. Salvador García 0001, Julián Luengo, José A. Sáez, Victoria López, Francisco Herrera |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Shared domains of competence of approximate learning models using measures of separability of classes
Julián Luengo, Francisco Herrera |
Inf. Sci. | 1 |
| 2012 | On the choice of the best imputation methods for missing values considering three groups of classification methods
Julián Luengo, Salvador García 0001, Francisco Herrera |
Knowl. Inf. Syst. | 1 |
| 2010 | Advanced nonparametric tests for multiple comparisons in the design of experiments in computational intelligence and data mining: Experimental analysis of power
Salvador García 0001, Alberto Fernández 0001, Julián Luengo, Francisco Herrera |
Inf. Sci. | 3 |