José María Luna

dblp:23/8487 · DBLP profile ↗
← Back
12ranked-venue papers in the field
6as first author
5since 2021 · last 2024
0000-0003-3537-2931ORCID · corroborated

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 5 (2 first)Data Mining & Knowledge Discovery · 3 (3 first)Big Data, Cloud & Distributed Data Systems · 2Database Systems & Data Management · 1 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2024 Data heterogeneity's impact on the performance of frequent itemset mining algorithms
abstract
Frequent itemset mining (FIM) is a widely used task that extracts frequently occurring itemsets from data. Plenty of deterministic algorithms are available for this daunting task. However, experimental studies have not considered that data heterogeneity significantly impacts the algorithms' performance, giving rise to unfair comparisons and biased conclusions. This paper seeks to advance by comparing cutting-edge algorithms using various frequency thresholds, considering the resulting data heterogeneity. An extensive experimental study is carried out, including the number of itemsets mined per second as the performance quality measure to compare algorithms. The experiments include defining eight metrics to quantify data heterogeneity, and their values vary the algorithms' performance. The results revealed that some techniques (hypercube decomposition and k-items machine) are essential to achieve excellent performance on any dataset, and most algorithms behave similarly well when they include those techniques. As a final important point, different threshold values produce dissimilar data subsets (data heterogeneity is not an immutable data characteristic), so a previous study on the database characteristics with a few minimum support thresholds could be beneficial to select the best-suited FIM algorithm beforehand.
Antonio Manuel Trasierras, José María Luna, Philippe Fournier-Viger, Sebastián Ventura
Inf. Sci.2
2023 Efficient mining of top-k high utility itemsets through genetic algorithms
José María Luna, R. Uday Kiran, Philippe Fournier-Viger, Sebastián Ventura
Inf. Sci.1
2022 Improving the understanding of cancer in a descriptive way: An emerging pattern mining-based approach
abstract
This paper presents an approach based on emerging pattern mining to analyse cancer through genomic data. Unlike existing approaches, mainly focused on predictive purposes, the proposal aims to improve the understanding of cancer descriptively, not requiring either any prior knowledge or hypothesis to be validated. Additionally, it enables to consider high-order relationships, so not only essential genes related to the disease are considered, but also the combined effect of various secondary genes that can influence different pathways directly or indirectly related to the disease. The prime hypothesis is that splitting genomic cancer data into two subsets, that is, cases and controls, will allow us to determine which genes, and their expressions, are associated with different cancer types. The possibilities of the proposal are demonstrated by analyzing RNA-Seq data for six different types of cancer: breast, colon, lung, thyroid, prostate, and kidney. Some of the extracted insights were already described in the related literature as good cancer bio-markers, while others have not been described yet mainly due to existing techniques are biased by prior knowledge provided by biological databases.
Antonio Manuel Trasierras, José María Luna, Sebastián Ventura
Int. J. Intell. Syst.2
2021 Discovering Relative High Utility Itemsets in Very Large Transactional Databases Using Null-Invariant Measure
abstract
High utility itemset mining is an important model in data mining. It involves discovering all itemsets in a quantitative transactional database that satisfy a user-specified minimum utility (minUtil) constraint. MinUtil controls the minimum value that an itemset must maintain in a database. Since the model evaluates an itemset’s interestingness using only the minUtil constraint, it implicitly assumes that all items in the database have similar utility values. However, some items have high utility, while others may have relatively low utility in a database. If minUtil is set too high, the user will miss all itemsets containing low utility items. To find itemsets that involve both high and low utility items, minUtil has to be set very low. However, this may cause a combinatorial explosion as the items with high utility may combine with others in all possible ways. This dilemma is called the low utility item problem. This paper proposes a flexible model of relative high utility itemset to address this problem. We introduce a new null-invariant measure, called utility ratio, to evaluate the interestingness of an itemset in the database. We also present a fast single scan algorithm to find all desired itemsets in the database. Experimental results demonstrate that the proposed algorithm is efficient. Finally, a case study on Yahoo! JAPAN retail data shows that the proposed model is useful.
R. Uday Kiran, Pradeep Pallikila, José María Luna, Philippe Fournier-Viger, Masashi Toyoda, P. Krishna Reddy
IEEE BigData3
2021 Mining local periodic patterns in a discrete sequence
Philippe Fournier-Viger, R. Uday Kiran, Sebastián Ventura, José María Luna
Inf. Sci.5
2016 Subgroup discovery on big data: Pruning the search space on exhaustive search algorithms
abstract
Subgroup Discovery is a broadly applicable supervised local pattern mining method to search relations between different properties with respect to a target variable. With the exponential growth in data storage, the massive data gathered has hampered the performance of current techniques. In this regard, our aim is to propose two new algorithms to discover subgroups on Big Data by using MapReduce. Apache Spark was used to tackle the Big Data requirements. The experimental study includes more than 50 large datasets and a set of efficient algorithms. Search spaces bigger than 1.276 · 1015subgroups are used. The experimental study reveals the alluring results in efficiency when optimistic estimates are considered, as well as demonstrating the usefulness of using Apache Spark to tackle Big Data.
Francisco Padillo, José María Luna, Sebastián Ventura
IEEE BigData2
2016 LAIM discretization for multi-label data
Alberto Cano 0001, José María Luna, Eva Lucrecia Gibaja Galindo, Sebastián Ventura
Inf. Sci.2
2016 Discovering useful patterns from multiple instance data
José María Luna, Alberto Cano 0001, Virgilijus Sakalauskas, Sebastián Ventura
Inf. Sci.1
2016 Mining exceptional relationships with grammar-guided genetic programming
José María Luna, Mykola Pechenizkiy, Sebastián Ventura
Knowl. Inf. Syst.1
2014 On the adaptability of G3PARM to the extraction of rare association rules
José María Luna, José Raúl Romero, Sebastián Ventura
Knowl. Inf. Syst.1
2013 Grammar-based multi-objective algorithms for mining association rules
José María Luna, José Raúl Romero, Sebastián Ventura
Data Knowl. Eng.1
2012 Design and behavior study of a grammar-guided genetic programming algorithm for mining association rules
José María Luna, José Raúl Romero, Sebastián Ventura
Knowl. Inf. Syst.1