José María Luna

dblp:23/8487 · DBLP profile ↗
← Back
40ranked-venue papers
16as first author
9since 2021 · last 2026
0000-0003-3537-2931ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 10 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 12 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 DANTIS library: Detection of ANomalies in TIme series
abstract
Anomaly detection in time series is essential in domains such as predictive maintenance, cybersecurity, health monitoring, and quality control. Although several software libraries provide anomaly detection algorithms, building complete and reproducible workflows still requires considerable expertise in data preprocessing, model configuration, training, evaluation, visualization, and statistical comparison. This complexity often limits the accessibility and reproducibility of anomaly detection workflows. This paper introduces DANTIS (Detection of ANomalies in TIme Series), an open-source Python library and desktop application designed to simplify the end-to-end development and comparison of anomaly detection models for time series. DANTIS provides a unified framework for data import, preprocessing, model training, testing, result visualization, and experiment export through both a programmatic interface and a graphical user interface. The library includes representative statistical, machine learning, and deep learning detectors, and its modular architecture allows new models to be incorporated easily. To validate its practical usefulness, we report a comparative experiment on labelled datasets from the UCR Anomaly Archive, involving nine representative detectors evaluated under a common protocol. The results show that DANTIS can generate predictive metrics, runtime information, and structured result matrices suitable for statistical analysis with Friedman and Nemenyi tests. Overall, DANTIS provides an accessible, extensible, and reproducible environment for researchers and practitioners who need to compare and interpret anomaly detection methods in time-series applications.
Christian Luna, Elena Álvarez, Rafael Egea, José María Luna, Sebastián Ventura
Neurocomputing4
2024 StaTDS library: Statistical tests for Data Science
abstract
In Data Science, there is a continual demand for statistical comparison to identify the most advantageous algorithms. Finding a software tool that facilitates the execution of multiple tests on different Data Science experiments without relying on additional libraries poses a challenge. This paper introduces StaTDS, an open-source library and web application implemented entirely in pure Python, designed to analyze, test, and compare Data Science algorithms. StaTDS implements all statistical tests without external dependencies. It ensures its durability and avoids future uncontrolled deprecated dependencies. With support for a wide variety of statistical tests (24 in total), StaTDS surpasses existing libraries dedicated to statistical testing. Moreover, the library incorporates tests to guide users in determining whether to employ parametric or non-parametric tests, such as the assessment of normality and homoscedasticity. This platform-independent library is available on GitHub under the GNU General Public License.
Christian Luna, Antonio R. Moya, José María Luna, Sebastián Ventura
Neurocomputing3
2024 Data heterogeneity's impact on the performance of frequent itemset mining algorithms
abstract
Frequent itemset mining (FIM) is a widely used task that extracts frequently occurring itemsets from data. Plenty of deterministic algorithms are available for this daunting task. However, experimental studies have not considered that data heterogeneity significantly impacts the algorithms' performance, giving rise to unfair comparisons and biased conclusions. This paper seeks to advance by comparing cutting-edge algorithms using various frequency thresholds, considering the resulting data heterogeneity. An extensive experimental study is carried out, including the number of itemsets mined per second as the performance quality measure to compare algorithms. The experiments include defining eight metrics to quantify data heterogeneity, and their values vary the algorithms' performance. The results revealed that some techniques (hypercube decomposition and k-items machine) are essential to achieve excellent performance on any dataset, and most algorithms behave similarly well when they include those techniques. As a final important point, different threshold values produce dissimilar data subsets (data heterogeneity is not an immutable data characteristic), so a previous study on the database characteristics with a few minimum support thresholds could be beneficial to select the best-suited FIM algorithm beforehand.
Antonio Manuel Trasierras, José María Luna, Philippe Fournier-Viger, Sebastián Ventura
Inf. Sci.2
2023 Radiomics Software Tools: A comparative Analysis on Breast Cancer
abstract
Radiomics is an emerging and promising field used to describe visual information from medical images by means of numerical features. Several Radiomics software tools are available in the literature, but they return different features and make dissimilar calculations. Choosing one tool or another is not easy so a comparison for classification tasks is required. This paper compares three of these frameworks (3D Slicer, LIFEx and MaZda) on breast cancer data. In this analysis, we tested the features extracted from each tool using different pre-processing techniques and machine learning algorithms to classify the lesion as benign or malignant on more than 350 registers. Two different projections were considered, that is, craniocaudal (183 registers) and mediolateral oblique (172 registers). The results demonstrated that 3D Slicer obtained the best performance in the craniocaudal projection, while MaZda and LIFEx are more appropriate for the mediolateral oblique projection. The results are really promising for classification tasks, exceeding 85% in F1-score.
Eduardo Almeda Luna, José María Luna, Sebastián Ventura
CBMS2
2023 A contrast set mining based approach for cancer subtype analysis
abstract
The task of detecting common and unique characteristics among different cancer subtypes is an important focus of research that aims to improve personalized therapies. Unlike current approaches mainly based on predictive techniques, our study aims to improve the knowledge about the molecular mechanisms that descriptively led to cancer, thus not requiring previous knowledge to be validated. Here, we propose an approach based on contrast set mining to capture high-order relationships in cancer transcriptomic data. In this way, we were able to extract valuable insights from several cancer subtypes in the form of highly specific genetic relationships related to functional pathways affected by the disease. To this end, we have divided several cancer gene expression databases by the subtype associated with each sample to detect which gene groups are related to each cancer subtype. To demonstrate the potential and usefulness of the proposed approach we have extensively analysed RNA-Seq gene expression data from breast, kidney, and colon cancer subtypes. The possible role of the obtained genetic relationships was further evaluated through extensive literature research, while its prognosis was assessed via survival analysis, finding gene expression patterns related to survival in various cancer subtypes. Some gene associations were described in the literature as potential cancer biomarkers while other results have been not described yet and could be a starting point for future research.
Antonio Manuel Trasierras, José María Luna, Sebastián Ventura
Artif. Intell. Medicine2
2023 Efficient mining of top-k high utility itemsets through genetic algorithms
José María Luna, R. Uday Kiran, Philippe Fournier-Viger, Sebastián Ventura
Inf. Sci.1
2022 Improving the understanding of cancer in a descriptive way: An emerging pattern mining-based approach
abstract
This paper presents an approach based on emerging pattern mining to analyse cancer through genomic data. Unlike existing approaches, mainly focused on predictive purposes, the proposal aims to improve the understanding of cancer descriptively, not requiring either any prior knowledge or hypothesis to be validated. Additionally, it enables to consider high-order relationships, so not only essential genes related to the disease are considered, but also the combined effect of various secondary genes that can influence different pathways directly or indirectly related to the disease. The prime hypothesis is that splitting genomic cancer data into two subsets, that is, cases and controls, will allow us to determine which genes, and their expressions, are associated with different cancer types. The possibilities of the proposal are demonstrated by analyzing RNA-Seq data for six different types of cancer: breast, colon, lung, thyroid, prostate, and kidney. Some of the extracted insights were already described in the related literature as good cancer bio-markers, while others have not been described yet mainly due to existing techniques are biased by prior knowledge provided by biological databases.
Antonio Manuel Trasierras, José María Luna, Sebastián Ventura
Int. J. Intell. Syst.2
2021 Discovering Relative High Utility Itemsets in Very Large Transactional Databases Using Null-Invariant Measure
abstract
High utility itemset mining is an important model in data mining. It involves discovering all itemsets in a quantitative transactional database that satisfy a user-specified minimum utility (minUtil) constraint. MinUtil controls the minimum value that an itemset must maintain in a database. Since the model evaluates an itemset’s interestingness using only the minUtil constraint, it implicitly assumes that all items in the database have similar utility values. However, some items have high utility, while others may have relatively low utility in a database. If minUtil is set too high, the user will miss all itemsets containing low utility items. To find itemsets that involve both high and low utility items, minUtil has to be set very low. However, this may cause a combinatorial explosion as the items with high utility may combine with others in all possible ways. This dilemma is called the low utility item problem. This paper proposes a flexible model of relative high utility itemset to address this problem. We introduce a new null-invariant measure, called utility ratio, to evaluate the interestingness of an itemset in the database. We also present a fast single scan algorithm to find all desired itemsets in the database. Experimental results demonstrate that the proposed algorithm is efficient. Finally, a case study on Yahoo! JAPAN retail data shows that the proposed model is useful.
R. Uday Kiran, Pradeep Pallikila, José María Luna, Philippe Fournier-Viger, Masashi Toyoda, P. Krishna Reddy
IEEE BigData3
2021 Mining local periodic patterns in a discrete sequence
Philippe Fournier-Viger, R. Uday Kiran, Sebastián Ventura, José María Luna
Inf. Sci.5
2020 Mining Cross-Level High Utility Itemsets
Philippe Fournier-Viger, Jerry Chun-Wei Lin, José María Luna, Sebastián Ventura
IEA/AIE4
2020 LAC: Library for associative classification
Francisco Padillo, José María Luna, Sebastián Ventura
Knowl. Based Syst.2
2019 MiNerDoc: a Semantically Enriched Text Mining System to Transform Clinical Text into Knowledge
abstract
Existing systems to support the daily decision taking process carried out by health professionals need to be used independently to perform different text mining subtasks. In practice, there are few systems that unify all the subtasks into an unique framework, easing therefore the clinical work by automating complex clinical tasks such as the detection of clinical alerts as well as clinical information coding. In this sense, the MiNerDoc system is proposed, whose main objective is to support clinical decision-taking process by analysing tons of textual clinical reports in an unified framework. MiNerDoc performs two basic functions that are of great importance in the medical field: detection of risk factors based on the recognition of five medical entities (Disease, Pharmacologic, Region/Part Body, Procedure/Test, Finding/Sign), and automatic prediction of standardized diagnostic codes (MeSH descriptors). A major feature of MiNerDoc is it includes external knowledge sources such as MetaMap and UMLS to terminologically and semantically enrich the interpretation of clinical texts. Some study cases are considered in this work to demonstrate the power of MiNerDoc.
Carmen Luque, José María Luna, Sebastián Ventura
CBMS2
2019 Associative Classification in Big Data through a G3P Approach
abstract
The associative classification field includes really interesting approaches for building reliable classifiers and any of these approaches generally work on four different phases (data discretization, pattern mining, rule mining, and classifier building). This number of phases is a handicap when big datasets are analysed. The aim of this work is to propose a novel evolutionary algorithm for efficiently building associative classifiers in Big Data. The proposed model works in only two phases (a grammar-guided genetic programming framework is performed in each phase): 1) mining reliable association rules; 2) building an accurate classifier by ranking and combining the previously mined rules. The proposal has been implemented on Apache Spark to take advantage of the distributed computing. The experimental analysis was performend on 40 well-known datasets and considering 13 algorithms taken from literature. A series of non-parametric tests has also been carried out to determine statistical differences. Results are quite promising in terms of reliability and efficiency on high-dimensional data.
José María Luna, Francisco Padillo, Sebastián Ventura
IoTBDS1
2018 Mining Context-Aware Association Rules Using Grammar-Based Genetic Programming
abstract
Real-world data usually comprise features whose interpretation depends on some contextual information. Such contextual-sensitive features and patterns are of high interest to be discovered and analyzed in order to obtain the right meaning. This paper formulates the problem of mining context-aware association rules, which refers to the search for associations between itemsets such that the strength of their implication depends on a contextual feature. For the discovery of this type of associations, a model that restricts the search space and includes syntax constraints by means of a grammar-based genetic programming methodology is proposed. Grammars can be considered as a useful way of introducing subjective knowledge to the pattern mining process as they are highly related to the background knowledge of the user. The performance and usefulness of the proposed approach is examined by considering synthetically generated datasets. A posteriori analysis on different domains is also carried out to demonstrate the utility of this kind of associations. For example, in educational domains, it is essential to identify and understand contextual and context-sensitive factors that affect overall and individual student behavior and performance. The results of the experiments suggest that the approach is feasible and it automatically identifies interesting context-aware associations from real-world datasets.
José María Luna, Mykola Pechenizkiy, María José del Jesus, Sebastián Ventura
IEEE Trans. Cybern.1
2018 Apriori Versions Based on MapReduce for Mining Frequent Patterns on Big Data
abstract
Pattern mining is one of the most important tasks to extract meaningful and useful information from raw data. This task aims to extract item-sets that represent any type of homogeneity and regularity in data. Although many efficient algorithms have been developed in this regard, the growing interest in data has caused the performance of existing pattern mining techniques to be dropped. The goal of this paper is to propose new efficient pattern mining algorithms to work in big data. To this aim, a series of algorithms based on the MapReduce framework and the Hadoop open-source implementation have been proposed. The proposed algorithms can be divided into three main groups. First, two algorithms [Apriori MapReduce (AprioriMR) and iterative AprioriMR] with no pruning strategy are proposed, which extract any existing itemset in data. Second, two algorithms (space pruning AprioriMR and top AprioriMR) that prune the search space by means of the well-known anti-monotone property are proposed. Finally, a last algorithm (maximal AprioriMR) is also proposed for mining condensed representations of frequent patterns. To test the performance of the proposed algorithms, a varied collection of big data datasets have been considered, comprising up to 3·1018 transactions and more than 5 million of distinct single-items. The experimental stage includes comparisons against highly efficient and well-known pattern mining algorithms. Results reveal the interest of applying MapReduce versions when complex problems are considered, and also the unsuitability of this paradigm when dealing with small data.
José María Luna, Francisco Padillo, Mykola Pechenizkiy, Sebastián Ventura
IEEE Trans. Cybern.1
2017 An evolutionary algorithm for mining rare association rules: A Big Data approach
abstract
Association rule mining is one of the most wellknown techniques to discover interesting relations between items in data. To date, this task has been mainly focused on the discovery of frequent relationships. However, it is often interesting to focus on those that do not occur frequently. Rare association rule mining is an alluring field aiming at describing rare cases or unexpected behavior. This field is really useful over Big Data where abnormal endeavor are more curious than common behavior. In this sense, our aim is to propose a new evolutionary algorithm based on grammars to obtain rare association rules on Big Data. The novelty of our work is that it is eminently designed to be parallel, enabling its use over emerging technologies as Spark and Flink. Furthermore, while other algorithms focus on maximizing a couple of quality measure ignoring the rest, our fitness function has been precisely designed to obtain a trade-off while maximizing a set of well-known quality measures. The experimental study includes more than 70 datasets revealing alluring results in efficiency when more than 300 million of instances and file sizes up to 250 GBytes are considered, and proving that it is able to run efficiently in huge volumes of data.
Francisco Padillo, José María Luna, Sebastián Ventura
CEC2
2016 Subgroup discovery on big data: Pruning the search space on exhaustive search algorithms
abstract
Subgroup Discovery is a broadly applicable supervised local pattern mining method to search relations between different properties with respect to a target variable. With the exponential growth in data storage, the massive data gathered has hampered the performance of current techniques. In this regard, our aim is to propose two new algorithms to discover subgroups on Big Data by using MapReduce. Apache Spark was used to tackle the Big Data requirements. The experimental study includes more than 50 large datasets and a set of efficient algorithms. Search spaces bigger than 1.276 · 1015subgroups are used. The experimental study reveals the alluring results in efficiency when optimistic estimates are considered, as well as demonstrating the usefulness of using Apache Spark to tackle Big Data.
Francisco Padillo, José María Luna, Sebastián Ventura
IEEE BigData2
2016 Mining Perfectly Rare Itemsets on Big Data: An Approach Based on Apriori-Inverse and MapReduce
Francisco Padillo, José María Luna, Sebastián Ventura
ISDA2
2016 LAIM discretization for multi-label data
Alberto Cano 0001, José María Luna, Eva Lucrecia Gibaja Galindo, Sebastián Ventura
Inf. Sci.2
2016 Discovering useful patterns from multiple instance data
José María Luna, Alberto Cano 0001, Virgilijus Sakalauskas, Sebastián Ventura
Inf. Sci.1
2016 Mining exceptional relationships with grammar-guided genetic programming
José María Luna, Mykola Pechenizkiy, Sebastián Ventura
Knowl. Inf. Syst.1
2016 Speeding-Up Association Rule Mining With Inverted Index Compression
abstract
The growing interest in data storage has made the data size to be exponentially increased, hampering the process of knowledge discovery from these large volumes of high-dimensional and heterogeneous data. In recent years, many efficient algorithms for mining data associations have been proposed, facing up time and main memory requirements. Nevertheless, this mining process could still become hard when the number of items and records is extremely high. In this paper, the goal is not to propose new efficient algorithms but a new data structure that could be used by a variety of existing algorithms without modifying its original schema. Thus, our aim is to speed up the association rule mining process regardless the algorithm used to this end, enabling the performance of efficient implementations to be enhanced. The structure simplifies, reorganizes, and speeds up the data access by sorting data by means of a shuffling strategy based on the hamming distance, which achieve similar values to be closer, and considering both an inverted index mapping and a run length encoding compression. In the experimental study, we explore the bounds of the algorithms' performance by using a wide number of data sets that comprise either thousands or millions of both items and records. The results demonstrate the utility of the proposed data structure in enhancing the algorithms' runtime orders of magnitude, and substantially reducing both the auxiliary and the main memory requirements.
José María Luna, Alberto Cano 0001, Mykola Pechenizkiy, Sebastián Ventura
IEEE Trans. Cybern.1
2015 Exploring the Influence of ICT in online Education Through Data Mining Tools
Javier Bravo, Sonia J. Romero, José María Luna, Sonia Pamplona
EDM3
2015 Discovering clues to avoid middle school failure at early stages
abstract
The use of data mining techniques in educational domains helps to find new knowledge about how students learn and how to improve the resources management. Using these techniques for predicting school failure is very useful in order to carry out actions to avoid drop out. With this purpose, we try to determine the earliest stage when the quality of the results allows for clarifying the possibility of school failure. We process real information from a Spanish high school by structuring the whole data in incremental datasets, which represent how students' academic records grow. Our study reveals an early and robust detection of the risky cases of school failure at the end of the first out of four courses.
Manuel Ángel Jiménez-Gómez, José María Luna, Cristóbal Romero 0001, Sebastián Ventura
LAK2
2015 An evolutionary algorithm for the discovery of rare class association rules in learning management systems
José María Luna, Cristóbal Romero 0001, José Raúl Romero, Sebastián Ventura
Appl. Intell.1
2015 A classification module for genetic programming algorithms in JCLEC
Alberto Cano 0001, José María Luna, Amelia Zafra, Sebastián Ventura
J. Mach. Learn. Res.2
2014 On the adaptability of G3PARM to the extraction of rare association rules
José María Luna, José Raúl Romero, Sebastián Ventura
Knowl. Inf. Syst.1
2014 On the Use of Genetic Programming for Mining Comprehensible Rules in Subgroup Discovery
abstract
This paper proposes a novel grammar-guided genetic programming algorithm for subgroup discovery. This algorithm, called comprehensible grammar-based algorithm for subgroup discovery (CGBA-SD), combines the requirements of discovering comprehensible rules with the ability to mine expressive and flexible solutions owing to the use of a context-free grammar. Each rule is represented as a derivation tree that shows a solution described using the language denoted by the grammar. The algorithm includes mechanisms to adapt the diversity of the population by self-adapting the probabilities of recombination and mutation. We compare the approach with existing evolutionary and classic subgroup discovery algorithms. CGBA-SD appears to be a very promising algorithm that discovers comprehensible subgroups and behaves better than other algorithms as measures by complexity, interest, and precision indicate. The results obtained were validated by means of a series of nonparametric tests.
José María Luna, José Raúl Romero, Cristóbal Romero 0001, Sebastián Ventura
IEEE Trans. Cybern.1
2013 Discovering Subgroups by Means of Genetic Programming
José María Luna, José Raúl Romero, Cristóbal Romero 0001, Sebastián Ventura
EuroGP1
2013 Grammar-based multi-objective algorithms for mining association rules
José María Luna, José Raúl Romero, Sebastián Ventura
Data Knowl. Eng.1
2013 Association rule mining using genetic programming to provide feedback to instructors from multiple-choice quiz data
abstract
Abstract This paper proposes the application of association rule mining to improve quizzes and courses. First, the paper shows how to preprocess quiz data and how to create several data matrices for use in the process of knowledge discovery. Next, the proposed algorithm that uses grammar‐guided genetic programming is described and compared with both classical and recent soft‐computing association rule mining algorithms. Then, different objective and subjective rule evaluation measures are used to select the most interesting and useful rules. Experiments have been carried out by using real data of university students enrolled on an artificial intelligence practice Moodle's course on the CLIPS programming language. Some examples of these rules are shown, together with the feedback that they provide to instructors making decisions about how to improve quizzes and courses. Finally, starting with the information provided by the rules, the CLIPS quiz and course have been updated. These innovations have been evaluated by comparing the performance achieved by students before and after applying the changes using one control group and two different experimental groups.
Cristóbal Romero 0001, Amelia Zafra, José María Luna, Sebastián Ventura
Expert Syst. J. Knowl. Eng.3
2013 High performance evaluation of evolutionary-mined association rules on GPUs
Alberto Cano 0001, José María Luna, Sebastián Ventura
J. Supercomput.2
2012 Classification via clustering for predicting final marks starting from the student participation in Forums
Manuel Ignacio López, Cristóbal Romero 0001, Sebastián Ventura, José María Luna
EDM4
2012 Meta-learning Approach for Automatic Parameter Tuning: A case of study with educational datasets
María De Mar Molina, Cristóbal Romero 0001, Sebastián Ventura, José María Luna
EDM4
2012 A genetic programming free-parameter algorithm for mining association rules
abstract
This paper presents a free-parameter grammar-guided genetic programming algorithm for mining association rules. This algorithm uses a contex-free grammar to represent individuals, encoding the solutions in a tree-shape conformant to the grammar, so they are more expressive and flexible. The algorithm here presented has the advantages of using evolutionary algorithms for mining association rules, and it also solves the problem of tuning the huge number of parameters required by these algorithms. The main feature of this algorithm is the small number of parameters required, providing the possibility of discovering association rules in an easy way for non-expert users. We compare our approach to existing evolutionary and exhaustive search algorithms, obtaining important results and overcoming the drawbacks of both exhaustive search and evolutionary algorithms. The experimental stage reveals that this approach discovers frequent and reliable rules without a parameter tuning.
José María Luna, José Raúl Romero, Cristóbal Romero 0001, Sebastián Ventura
ISDA1
2012 Design and behavior study of a grammar-guided genetic programming algorithm for mining association rules
José María Luna, José Raúl Romero, Sebastián Ventura
Knowl. Inf. Syst.1
2011 Association rule mining using a multi-objective grammar-based ant programming algorithm
abstract
This paper presents a method for extracting association rules by means of a multi-objective grammar guided ant programming algorithm. Solution construction is guided by a context-free grammar specifically suited for association rule mining, which defines the search space of all possible expressions or programs. Evaluation of individuals is considered from a Pareto-based point of view, measuring support and confidence of rules mined, and assigning them a ranking fitness. The proposed algorithm is verified over 10 varied data sets and compared to other association rule mining algorithms from several paradigms such as exhaustive search, genetic algorithms and genetic programming, showing that ant programming is a good technique at addressing the association task of data mining as well.
Juan Luis Olmo, José María Luna, José Raúl Romero, Sebastián Ventura
ISDA2
2010 G3PARM: A Grammar Guided Genetic Programming algorithm for mining association rules
abstract
This paper presents the G3PARM algorithm for mining representative association rules. G3PARM is an evolutionary algorithm that uses G3P (Grammar Guided Genetic Programming) and an auxiliary population made up of its best individuals who will then act as parents for the next generation. Due to the nature of G3P, the G3PARM algorithm allows us to obtain valid individuals by defining them through a context-free grammar and, furthermore, this algorithm is generic with respect to data type. We compare our algorithm to two multiobjective algorithms frequently used in literature and known as NSGA2 (Non dominated Sort Genetic Algorithm) and SPEA2 (Strength Pareto Evolutionary Algorithm) and demonstrate the efficiency of our algorithm in terms of running-time, coverage and average support, providing the user with high representative rules.
José María Luna, José Raúl Romero, Sebastián Ventura
IEEE Congress on Evolutionary Computation1
2010 Mining Rare Association Rules from e-Learning Data
Cristóbal Romero 0001, José Raúl Romero, José María Luna, Sebastián Ventura
EDM3
2010 An intruder detection approach based on infrequent rating pattern mining
abstract
This work presents a novel proposal for incremental intruder detection in collaborative recommender systems. We explore the use of rare association rule mining to reveal the existence of a suspected raid of attackers that would alter the normal behaviour of a rating-based system. In this position paper we have extended our previous G3PARM algorithm, which has already proven to serve as a solid method for extracting frequent association rules. G3PARM is an evolutionary algorithm that uses G3P (Grammar Guided Genetic Programming), which provides expressiveness and flexibility enough to adapt and apply the base context-free grammar to each specific problem or domain. We fully outline, moreover, the complete exploration and detection model, which includes some further post-analysis steps. Finally, as a proof of concept, we validate the scalability, efficiency and accuracy of our proposal showing the results obtained when different malicious intruders want to attack an on line recommender system.
José María Luna, Aurora Ramírez 0001, José Raúl Romero, Sebastián Ventura
ISDA1