José Antonio Lozano 0001

dblp:l/JoseAntonioLozano · DBLP profile ↗
← Back
19ranked-venue papers in the field
0as first author
5since 2021 · last 2023
0000-0002-4683-8111ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 6Data Mining & Knowledge Discovery · 6Knowledge Engineering, Semantic Web & Information Systems · 4Other / Interdisciplinary · 3
YearPublicationVenuePosition
2023 Selective Imputation for Multivariate Time Series Datasets With Missing Values
abstract
Multivariate time series often contain missing values for reasons such as failures in data collection mechanisms. Since these missing values can complicate the analysis of time series data, imputation techniques are typically used to deal with this issue. However, the quality of the imputation directly affects the performance of downstream tasks. In this paper, we propose a selective imputation method that identifies a subset of timesteps with missing values to impute in a multivariate time series dataset. This selection, which will result in shorter and simpler time series, is based on both reducing the uncertainty of the imputations and representing the original time series as good as possible. In particular, the method uses multi-objective optimization techniques to select the optimal set of points, and in this selection process, we leverage the beneficial properties of the Multi-task Gaussian Process (MGP). The method is applied to different datasets to analyze the quality of the imputations and the performance obtained in downstream tasks, such as classification or anomaly detection. The results show that much shorter and simpler time series are able to maintain or even improve both the quality of the imputations and the performance of the downstream tasks.
Ane Blázquez-García, Kristoffer Wickstrøm, Shujian Yu, Karl Øyvind Mikalsen, Ahcène Boubekki, Angel Conde, Usue Mori, Robert Jenssen, José Antonio Lozano 0001
IEEE Trans. Knowl. Data Eng.9
2023 SNDProb: A Probabilistic Approach for Streaming Novelty Detection
abstract
A probabilistic framework for streaming novelty detection is proposed and illustrated with a mixture of Gaussian distributions that models the set of classes. Instances are predicted based on the probability of belonging to each of the classes. Those for which the model cannot provide confident predictions are introduced into a fixed-sized buffer. When the buffer is full, an Expectation Maximization (EM) algorithm is run to search for new emerging classes in the buffer, and update the current model. The EM algorithm has to deal with an scenario where both probability distributions and instances are available. To overcome this issue, the probability distributions (classes) are weighted. The weights are inferred using a meta-regression model which has been pretrained and supplied with the proposed algorithm. Experiments have been run using synthetic datasets to have a close control over the class arrival strategies, the shape, and the overlapping degree between classes. It is shown that when the assumptions of the probabilistic model are fulfilled, the proposed method outperforms literature non-parametric approaches. Furthermore it obtains competitive results in the case of non-Gaussian classes. The experiments reveal, for the first time, the high sensitivity of the novelty detection algorithms to the class arrival strategies.
Ander Carreño, Iñaki Inza, José Antonio Lozano 0001
IEEE Trans. Knowl. Data Eng.3
2023 Minimum Recall-Based Loss Function for Imbalanced Time Series Classification
abstract
This paper deals with imbalanced time series classification problems. In particular, we propose to learn time series classifiers that maximize the minimum recall of the classes rather than the accuracy. Consequently, we manage to obtain classifiers which tend to give the same importance to all the classes. Unfortunately, for most of the traditional classifiers, learning to maximize the minimum recall of the classes is not trivial (if possible), since it can distort the nature of the classifiers themselves. Neural networks, in contrast, are classifiers that explicitly define a loss function, allowing it to be modified. Given that the minimum recall is not a differentiable function, and therefore does not allow the use of common gradient-based learning methods, we apply and evaluate several smooth approximations of the minimum recall function. A thorough experimental evaluation shows that our approach improves the performance of state-of-the-art methods used in imbalanced time series classification, obtaining higher recall values for the minority classes, incurring only a slight loss in accuracy.
Josu Ircio, Aizea Lojo, Usue Mori, Simon Malinowski, José Antonio Lozano 0001
IEEE Trans. Knowl. Data Eng.5
2022 An Efficient Split-Merge Re-Start for the $K$K-Means Algorithm
abstract
The$K$-means algorithm is one of the most popular clustering methods. However, it is a well-known fact that its performance, in terms of quality of the obtained solution and computational load, highly depends upon its initialization phase. For this reason, different initialization techniques have been developed throughout the years to enable its fast convergence to competitive solutions. In this sense, it is common practice to re-start the$K$-means algorithm several times via one of these techniques and keep the solution with the lowest error. Unfortunately, such a choice is still likely to be a poor approximation of the optimal set of centroids. In this article, we introduce a cheap Split-Merge step that can be used to re-start the$K$-means algorithm after reaching a fixed point. Under some settings, one can show that this approach reduces the error of the given fixed point without requiring any further iteration of the$K$-means algorithm. Moreover, experimental results show that this strategy is able to generate approximations with an associated error that is hard to reach for different multi-start methods, such as multi-start Forgy$K$-means,$K$-means++ and Hartigan$K$-means, while also computing a lower amount of distances than the previous algorithms.
Marco Capó, Aritz Pérez Martínez, José Antonio Lozano 0001
IEEE Trans. Knowl. Data Eng.3
2021 Water leak detection using self-supervised time series classification
Ane Blázquez-García, Angel Conde, Usue Mori, José Antonio Lozano 0001
Inf. Sci.4
2020 An efficient K-means clustering algorithm for tall data
Marco Capó, Aritz Pérez Martínez, José Antonio Lozano 0001
Data Min. Knowl. Discov.3
2019 A review on distance based time series classification
Amaia Abanda, Usue Mori, José Antonio Lozano 0001
Data Min. Knowl. Discov.3
2019 Aggregated outputs by linear models: An application on marine litter beaching prediction
Jerónimo Hernández-González, Iñaki Inza, Igor Granado, Oihane C. Basurko, Jose A. Fernandes, José Antonio Lozano 0001
Inf. Sci.6
2019 Early classification of time series using multi-objective optimization techniques
Usue Mori, Alexander Mendiburu, Isabel Marta Miranda, José Antonio Lozano 0001
Inf. Sci.4
2019 A Note on the Behavior of Majority Voting in Multi-Class Domains with Biased Annotators
abstract
Majority voting is a popular and robust strategy to aggregate different opinions in learning from crowds, where each worker labels examples according to their own criteria. Although it has been extensively studied in the binary case, its behavior with multiple classes is not completely clear, specifically when annotations are biased. This paper attempts to fill that gap. The behavior of the majority voting strategy is studied in-depth in multi-class domains, emphasizing the effect of annotation bias. By means of a complete experimental setting, we show the limitations of the standard majority voting strategy. The use of three simple techniques that infer global information from the annotations and annotators allows us to put the performance of the majority voting strategy in context.
Jerónimo Hernández-González, Iñaki Inza, José Antonio Lozano 0001
IEEE Trans. Knowl. Data Eng.3
2018 The Relationship Between Graphical Representations of Regular Vine Copulas and Polytrees
Diana Carrera, Roberto Santana 0001, José Antonio Lozano 0001
IPMU (3)3
2017 On-Line Dynamic Time Warping for Streaming Time Series
Izaskun Oregi, Aritz Pérez Martínez, Javier Del Ser, José Antonio Lozano 0001
ECML/PKDD (2)4
2017 Reliable early classification of time series based on discriminating the classes over time
Usue Mori, Alexander Mendiburu, Eamonn J. Keogh, José Antonio Lozano 0001
Data Min. Knowl. Discov.4
2017 Learning from Proportions of Positive and Unlabeled Examples
abstract
Weakly supervised classification tries to learn from data sets which are not certainly labeled. Many problems, with different natures of partial labeling, fit this description. In this paper, the novel problem of learning from positive-unlabeled proportions is presented. The provided examples are unlabeled, and the only class information available consists of the proportions of positive and unlabeled examples in different subsets of the training data set. We present a methodology that adapts to the different levels of class uncertainty to learn Bayesian network classifiers using an expectation-maximization strategy. It has been tested in a variety of artificial scenarios with different class uncertainty, as well as compared with two naive strategies that do not consider all the available class information. Finally, it has also been successfully tested in real data, collected from the embryo selection problem in assisted reproduction.
Jerónimo Hernández-González, Iñaki Inza, José Antonio Lozano 0001
Int. J. Intell. Syst.3
2016 Similarity Measure Selection for Clustering Time Series Databases
abstract
In the past few years, clustering has become a popular task associated with time series. The choice of a suitable distance measure is crucial to the clustering process and, given the vast number of distance measures for time series available in the literature and their diverse characteristics, this selection is not straightforward. With the objective of simplifying this task, we propose a multi-label classification framework that provides the means to automatically select the most suitable distance measures for clustering a time series database. This classifier is based on a novel collection of characteristics that describe the main features of the time series databases and provide the predictive information necessary to discriminate between a set of distance measures. In order to test the validity of this classifier, we conduct a complete set of experiments using both synthetic and real time series databases and a set of five common distance measures. The positive results obtained by the designed classification framework for various performance measures indicate that the proposed methodology is useful to simplify the process of distance selection in time series clustering tasks.
Usue Mori, Alexander Mendiburu, José Antonio Lozano 0001
IEEE Trans. Knowl. Data Eng.3
2015 Multidimensional Learning from Crowds: Usefulness and Application of Expertise Detection
abstract
Learning from crowds is a classification problem where the provided training instances are labeled by multiple (usually conflicting) annotators. In different scenarios of this problem, straightforward strategies show an astonishing performance. In this paper, we characterize the crowd scenarios where these basic strategies show a good behavior. As a consequence, this study allows to identify those scenarios where non-basic methods for combining the multiple labels are expected to obtain better results. In this context, we extend the learning from crowds paradigm to the multidimensional (MD) classification domain. Measuring the quality of the annotators, the presented EM-based method overcomes the lack of a fully reliable labeling for learning MD Bayesian network classifiers: As the expertise is identified and the contribution of the relevant annotators promoted, the model parameters are optimized. The good performance of our proposal is demonstrated throughout different sets of experiments.
Jerónimo Hernández-González, Iñaki Inza, José Antonio Lozano 0001
Int. J. Intell. Syst.3
2014 Assisting in search heuristics selection through multidimensional supervised classification: A case study on software testing
Ramón Sagarna, Alexander Mendiburu, Iñaki Inza, José Antonio Lozano 0001
Inf. Sci.4
2012 Wrapper positive Bayesian network classifiers
Borja Calvo, Iñaki Inza, Pedro Larrañaga, José Antonio Lozano 0001
Knowl. Inf. Syst.4
2006 Mixtures of Kikuchi Approximations
Roberto Santana 0001, Pedro Larrañaga, José Antonio Lozano 0001
ECML3