EDBT 2026 Demo / reviewers in the wild / expert
José María Luna
dblp:23/8487
· DBLP profile ↗
12ranked-venue papers in the field
6as first author
5since 2021 · last 2024
0000-0003-3537-2931ORCID · corroborated
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 5 (2 first)Data Mining & Knowledge Discovery · 3 (3 first)Big Data, Cloud & Distributed Data Systems · 2Database Systems & Data Management · 1 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Data heterogeneity's impact on the performance of frequent itemset mining algorithmsabstractFrequent itemset mining (FIM) is a widely used task that extracts frequently occurring itemsets from data. Plenty of deterministic algorithms are available for this daunting task. However, experimental studies have not considered that data heterogeneity significantly impacts the algorithms' performance, giving rise to unfair comparisons and biased conclusions. This paper seeks to advance by comparing cutting-edge algorithms using various frequency thresholds, considering the resulting data heterogeneity. An extensive experimental study is carried out, including the number of itemsets mined per second as the performance quality measure to compare algorithms. The experiments include defining eight metrics to quantify data heterogeneity, and their values vary the algorithms' performance. The results revealed that some techniques (hypercube decomposition and k-items machine) are essential to achieve excellent performance on any dataset, and most algorithms behave similarly well when they include those techniques. As a final important point, different threshold values produce dissimilar data subsets (data heterogeneity is not an immutable data characteristic), so a previous study on the database characteristics with a few minimum support thresholds could be beneficial to select the best-suited FIM algorithm beforehand. Antonio Manuel Trasierras, José María Luna, Philippe Fournier-Viger, Sebastián Ventura |
Inf. Sci. | 2 |
| 2023 | Efficient mining of top-k high utility itemsets through genetic algorithms
José María Luna, R. Uday Kiran, Philippe Fournier-Viger, Sebastián Ventura |
Inf. Sci. | 1 |
| 2022 | Improving the understanding of cancer in a descriptive way: An emerging pattern mining-based approachabstractThis paper presents an approach based on emerging pattern mining to analyse cancer through genomic data. Unlike existing approaches, mainly focused on predictive purposes, the proposal aims to improve the understanding of cancer descriptively, not requiring either any prior knowledge or hypothesis to be validated. Additionally, it enables to consider high-order relationships, so not only essential genes related to the disease are considered, but also the combined effect of various secondary genes that can influence different pathways directly or indirectly related to the disease. The prime hypothesis is that splitting genomic cancer data into two subsets, that is, cases and controls, will allow us to determine which genes, and their expressions, are associated with different cancer types. The possibilities of the proposal are demonstrated by analyzing RNA-Seq data for six different types of cancer: breast, colon, lung, thyroid, prostate, and kidney. Some of the extracted insights were already described in the related literature as good cancer bio-markers, while others have not been described yet mainly due to existing techniques are biased by prior knowledge provided by biological databases. Antonio Manuel Trasierras, José María Luna, Sebastián Ventura |
Int. J. Intell. Syst. | 2 |
| 2021 | Discovering Relative High Utility Itemsets in Very Large Transactional Databases Using Null-Invariant MeasureabstractHigh utility itemset mining is an important model in data mining. It involves discovering all itemsets in a quantitative transactional database that satisfy a user-specified minimum utility (minUtil) constraint. MinUtil controls the minimum value that an itemset must maintain in a database. Since the model evaluates an itemset’s interestingness using only the minUtil constraint, it implicitly assumes that all items in the database have similar utility values. However, some items have high utility, while others may have relatively low utility in a database. If minUtil is set too high, the user will miss all itemsets containing low utility items. To find itemsets that involve both high and low utility items, minUtil has to be set very low. However, this may cause a combinatorial explosion as the items with high utility may combine with others in all possible ways. This dilemma is called the low utility item problem. This paper proposes a flexible model of relative high utility itemset to address this problem. We introduce a new null-invariant measure, called utility ratio, to evaluate the interestingness of an itemset in the database. We also present a fast single scan algorithm to find all desired itemsets in the database. Experimental results demonstrate that the proposed algorithm is efficient. Finally, a case study on Yahoo! JAPAN retail data shows that the proposed model is useful. R. Uday Kiran, Pradeep Pallikila, José María Luna, Philippe Fournier-Viger, Masashi Toyoda, P. Krishna Reddy |
IEEE BigData | 3 |
| 2021 | Mining local periodic patterns in a discrete sequence
Philippe Fournier-Viger, R. Uday Kiran, Sebastián Ventura, José María Luna |
Inf. Sci. | 5 |
| 2016 | Subgroup discovery on big data: Pruning the search space on exhaustive search algorithmsabstractSubgroup Discovery is a broadly applicable supervised local pattern mining method to search relations between different properties with respect to a target variable. With the exponential growth in data storage, the massive data gathered has hampered the performance of current techniques. In this regard, our aim is to propose two new algorithms to discover subgroups on Big Data by using MapReduce. Apache Spark was used to tackle the Big Data requirements. The experimental study includes more than 50 large datasets and a set of efficient algorithms. Search spaces bigger than 1.276 · 1015subgroups are used. The experimental study reveals the alluring results in efficiency when optimistic estimates are considered, as well as demonstrating the usefulness of using Apache Spark to tackle Big Data. Francisco Padillo, José María Luna, Sebastián Ventura |
IEEE BigData | 2 |
| 2016 | LAIM discretization for multi-label data
Alberto Cano 0001, José María Luna, Eva Lucrecia Gibaja Galindo, Sebastián Ventura |
Inf. Sci. | 2 |
| 2016 | Discovering useful patterns from multiple instance data
José María Luna, Alberto Cano 0001, Virgilijus Sakalauskas, Sebastián Ventura |
Inf. Sci. | 1 |
| 2016 | Mining exceptional relationships with grammar-guided genetic programming
José María Luna, Mykola Pechenizkiy, Sebastián Ventura |
Knowl. Inf. Syst. | 1 |
| 2014 | On the adaptability of G3PARM to the extraction of rare association rules
José María Luna, José Raúl Romero, Sebastián Ventura |
Knowl. Inf. Syst. | 1 |
| 2013 | Grammar-based multi-objective algorithms for mining association rules
José María Luna, José Raúl Romero, Sebastián Ventura |
Data Knowl. Eng. | 1 |
| 2012 | Design and behavior study of a grammar-guided genetic programming algorithm for mining association rules
José María Luna, José Raúl Romero, Sebastián Ventura |
Knowl. Inf. Syst. | 1 |