Pedro Larrañaga

dblp:04/5852 · DBLP profile ↗
← Back
19ranked-venue papers in the field
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 7Knowledge Engineering, Semantic Web & Information Systems · 5 (1 first)Data Mining & Knowledge Discovery · 4Database Systems & Data Management · 2Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2026 Bandwidth selectors on semiparametric Bayesian networks
abstract
Semiparametric Bayesian networks (SPBNs) integrate parametric and non-parametric probabilistic models, offering flexibility in learning complex data distributions from samples. In particular, kernel density estimators (KDEs) are employed for the non-parametric component. Under the assumption of data normality, the normal rule is used to determine the bandwidth matrix for KDEs in SPBNs. This matrix is the critical hyperparameter that controls the trade-off between bias and variance. However, real-world data often deviates from normality, potentially leading to suboptimal density estimation and reduced predictive performance. This paper presents the theoretical framework for applying state-of-the-art bandwidth selectors to SPBNs and evaluates their impact on the performance. We explore cross-validation and plug-in selectors approaches, assessing their effectiveness in enhancing the learning capability and applicability of SPBNs. To support this investigation, we have extended the open-source package PyBNesian for SPBNs with additional bandwidth selection techniques and conducted extensive experimental analyses. Our results demonstrate that the proposed bandwidth selectors leverage larger sample sizes more effectively than the normal rule, which, despite its robustness, plateaus with more data. In particular, unbiased cross-validation generally outperforms the normal rule, highlighting its advantage in large sample size scenarios. • We establish guarantees for SPBN density estimation using plug-in and CV bandwidths. • Bandwidth selection boosts SPBN parameter/structure learning beyond normal rule. • Proposed bandwidths improve with sample size, unlike the stagnating normal rule. • Despite stagnating performance, normal rule is robust/effective in structure learning. • Unbiased CV balances cost/performance; best overall, especially for large samples.
Victor Alejandre, Concha Bielza, Pedro Larrañaga
Inf. Sci.3
2022 Piecewise forecasting of nonlinear time series with model tree dynamic Bayesian networks
abstract
When modelling multivariate continuous time series, a common issue is to find that the original processes that generated the data are nonlinear or that they drift away from the original distribution as the system evolves over time. In these scenarios, using a linear model such as a Gaussian dynamic Bayesian network (DBN) can result in severe forecasting inaccuracies due to the structure and parameters of the model remaining constant without considering either the elapsed time or the state of the system. To approach this problem, we propose a hybrid model that combines a model tree with DBNs. The model first divides the original data set into different scenarios based on the splits made by the tree over the system variables and then performs a piecewise regression in each branch of the tree to obtain nonlinear forecasts. The experimental results on three different datasets show that our model outperforms standard DBN models when dealing with nonlinear processes and is competitive with state-of-the-art time series forecasting methods.
David Quesada, Concha Bielza, Pedro Fontán, Pedro Larrañaga
Int. J. Intell. Syst.4
2022 Multipartition clustering of mixed data with Bayesian networks
abstract
Real-world applications often involve multifaceted data with several reasonable interpretations. To cluster this data, we need methods that are able to produce multiple clustering solutions. To this purpose, it is interesting to learn a finite mixture model with multiple latent variables, where each latent variable represents a unique way to partition the data. However, although there is an extensive literature on multipartition clustering methods for categorical data and for continuous data, there is a lack of work for mixed data. In this paper, we propose a multipartition clustering method that is able to efficiently deal with mixed data by exploiting the Bayesian network factorization and the variational Bayes framework. We show the flexibility and applicability of the proposed method by solving clustering, density estimation, and missing data imputation tasks in real-world data sets. For reproducibility, all code, data, and results can be found in the following public repository: https://github.com/ferjorosa/mpc-mixed.
Fernando Rodriguez-Sanchez, Concha Bielza, Pedro Larrañaga
Int. J. Intell. Syst.3
2022 Semiparametric Bayesian networks
David Atienza 0001, Concha Bielza, Pedro Larrañaga
Inf. Sci.3
2021 Multidimensional continuous time Bayesian network classifiers
abstract
The multidimensional classification of multivariate time series deals with the assignment of multiple classes to time-ordered data described by a set of feature variables. Although this challenging task has received almost no attention in the literature, it is present in a wide variety of domains, such as medicine, finance or industry. The complexity of this problem lies in two nontrivial tasks, the learning with multivariate time series in continuous time and the simultaneous classification of multiple class variables that may show dependencies between them. These can be addressed with different strategies, but most of them may involve a difficult preprocessing of the data, high space and classification complexity or ignoring useful interclass dependencies. Additionally, no attention has been given to the development of new multidimensional classifiers of time series based on probabilistic graphical models, even though transparent models can facilitate further understanding of the domain. In this paper, a novel probabilistic graphical model is proposed, which is able to classify a discrete multivariate temporal sequence into multiple class variables while modeling their dependencies. This model extends continuous time Bayesian networks to the multidimensional classification problem, which are able to explicitly represent the behavior of time series that evolve over continuous time. Different methods for the learning of the parameters and structure of the model are presented, and numerical experiments on synthetic and real-world data show encouraging results in terms of performance and learning time with respect to independent classifiers, the current alternative approach under the continuous time Bayesian network paradigm.
Carlos Villa-Blanco, Pedro Larrañaga, Concha Bielza
Int. J. Intell. Syst.2
2019 Circular Bayesian classifiers using wrapped Cauchy distributions
Ignacio Leguey, Concha Bielza, Pedro Larrañaga
Data Knowl. Eng.3
2019 A circular-linear dependence measure under Johnson-Wehrly distributions and its application in Bayesian networks
Ignacio Leguey, Pedro Larrañaga, Concha Bielza, Shogo Kato
Inf. Sci.2
2017 Frobenius Norm Regularization for the Multivariate Von Mises Distribution
abstract
Penalizing the model complexity is necessary to avoid overfitting when the number of data samples is low with respect to the number of model parameters. In this paper, we introduce a penalization term that places an independent prior distribution for each parameter of the multivariate von Mises distribution. We also propose a circular distance that can be used to estimate the Kullback–Leibler divergence between any two circular distributions as goodness-of-fit measure. We compare the resulting regularized von Mises models on synthetic data and real neuroanatomical data to show that the distribution fitted using the penalized estimator generally achieves better results than nonpenalized multivariate von Mises estimator.
Luis Rodriguez-Lujan, Pedro Larrañaga, Concha Bielza
Int. J. Intell. Syst.2
2016 Genetic algorithms and Gaussian Bayesian networks to uncover the predictive core set of bibliometric indices
abstract
The diversity of bibliometric indices today poses the challenge of exploiting the relationships among them. Our research uncovers the best core set of relevant indices for predicting other bibliometric indices. An added difficulty is to select the role of each variable, that is, which bibliometric indices are predictive variables and which are response variables. This results in a novel multioutput regression problem where the role of each variable (predictor or response) is unknown beforehand. We use Gaussian Bayesian networks to solve the this problem and discover multivariate relationships among bibliometric indices. These networks are learnt by a genetic algorithm that looks for the optimal models that best predict bibliometric data. Results show that the optimal induced Gaussian Bayesian networks corroborate previous relationships between several indices, but also suggest new, previously unreported interactions. An extended analysis of the best model illustrates that a set of 12 bibliometric indices can be accurately predicted using only a smaller predictive core subset composed of citations, g‐index, q2‐index, and hr‐index. This research is performed using bibliometric data on Spanish full professors associated with the computer science area.
Alfonso Ibáñez, Rubén Armañanzas, Concha Bielza, Pedro Larrañaga
J. Assoc. Inf. Sci. Technol.4
2015 Conditional Density Approximations with Mixtures of Polynomials
abstract
Mixtures of polynomials (MoPs) are a nonparametric density estimation technique especially designed for hybrid Bayesian networks with continuous and discrete variables. Algorithms to learn one- and multidimensional (marginal) MoPs from data have recently been proposed. In this paper, we introduce two methods for learning MoP approximations of conditional densities from data. Both approaches are based on learning MoP approximations of the joint density and the marginal density of the conditioning variables, but they differ as to how the MoP approximation of the quotient of the two densities is found. We illustrate and study the methods using data sampled from known parametric distributions, and demonstrate their applicability by learning models based on real neuroscience data. Finally, we compare the performance of the proposed methods with an approach for learning mixtures of truncated basis functions (MoTBFs). The empirical results show that the proposed methods generally yield models that are comparable to or significantly better than those found using the MoTBF-based method.
Gherardo Varando, Pedro L. López-Cruz, Thomas D. Nielsen, Pedro Larrañaga, Concha Bielza
Int. J. Intell. Syst.4
2014 Semi-supervised projected model-based clustering
Luis Guerra, Concha Bielza, Víctor Robles, Pedro Larrañaga
Data Min. Knowl. Discov.4
2014 Multi-Dimensional Classification with Super-Classes
abstract
The multi-dimensional classification problem is a generalization of the recently-popularized task of multi-label classification, where each data instance is associated with multiple class variables. There has been relatively little research carried out specific to multi-dimensional classification and, although one of the core goals is similar (modeling dependencies among classes), there are important differences; namely a higher number of possible classifications. In this paper we present method for multi-dimensional classification, drawing from the most relevant multi-label research, and combining it with important novel developments. Using a fast method to model the conditional dependence between class variables, we form super-class partitions and use them to build multi-dimensional learners, learning each super-class as an ordinary class, and thus explicitly modeling class dependencies. Additionally, we present a mechanism to deal with the many class values inherent to super-classes, and thus make learning efficient. To investigate the effectiveness of this approach we carry out an empirical evaluation on a range of multi-dimensional datasets, under different evaluation metrics, and in comparison with high-performing existing multi-dimensional approaches from the literature. Analysis of results shows that our approach offers important performance gains over competing methods, while also exhibiting tractable running time.
Jesse Read, Concha Bielza, Pedro Larrañaga
IEEE Trans. Knowl. Data Eng.3
2013 Comparison of metaheuristic strategies for peakbin selection in proteomic mass spectrometry data
Miguel García-Torres, Rubén Armañanzas, Concha Bielza, Pedro Larrañaga
Inf. Sci.4
2013 A review on evolutionary algorithms in Bayesian network learning and inference tasks
Pedro Larrañaga, Hossein Karshenas, Concha Bielza, Roberto Santana 0001
Inf. Sci.1
2012 Wrapper positive Bayesian network classifiers
Borja Calvo, Iñaki Inza, Pedro Larrañaga, José Antonio Lozano 0001
Knowl. Inf. Syst.3
2006 Mixtures of Kikuchi Approximations
Roberto Santana 0001, Pedro Larrañaga, José Antonio Lozano 0001
ECML2
2003 Interval Estimation Naïve Bayes
Víctor Robles, Pedro Larrañaga, José M. Peña 0002, Ernestina Menasalvas Ruiz, María S. Pérez 0001
IDA2
2003 Improvement of Naïve Bayes Collaborative Filtering Using Interval Estimation
abstract
Recommender systems emerged to help users choose among the large amount of options that ecommerce sites offer. Collaborative filtering is one of the most successful recommender techniques. Here we propose an approach to collaborative filtering based on the simple Bayesian classifier. We propose a method of increasing the efficiency of naive Bayes by applying a new semi naive Bayes approach based on interval estimation. To evaluate our algorithm we use a database of Microsoft anonymous Web data from the UCl repository. Our empirical results show that our proposed Interval based naive Bayes approach outperforms typical naive Bayes.
Víctor Robles, Pedro Larrañaga, Ernestina Menasalvas Ruiz, María S. Pérez 0001, Vanessa Herves
Web Intelligence2
2003 Learning Bayesian networks in the space of structures by estimation of distribution algorithms
abstract
The induction of the optimal Bayesian network structure is NP-hard, justifying the use of search heuristics. Two novel population-based stochastic search approaches, univariate marginal distribution algorithm (UMDA) and population-based incremental learning (PBIL), are used to learn a Bayesian network structure from a database of cases in a score + search framework. A comparison with a genetic algorithm (GA) approach is performed using three different scores: penalized maximum likelihood, marginal likelihood, and information-theory–based entropy. Experimental results show the interesting capabilities of both novel approaches with respect to the score value and the number of generations needed to converge. © 2003 Wiley Periodicals, Inc.
Rosa Blanco, Iñaki Inza, Pedro Larrañaga
Int. J. Intell. Syst.3