VLDB 2026 Research / reviewers in the wild / expert
Concha Bielza
dblp:16/5524 · also Concepcion Bielza
· DBLP profile ↗
93ranked-venue papers
10as first author
26since 2021 · last 2027
0000-0001-7109-2668ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 65 · 7 first-author · 20 since 2021Databases, data management, data science and information retrieval · 17 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Human-computer interaction and ubiquitous computing · 3Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | A fully unsupervised concept drift detector via Bayesian networks for machine condition monitoring
Rafael Sojo, Concha Bielza, Pedro Larrañaga |
Expert Syst. Appl. | 2 |
| 2026 | Bandwidth selectors on semiparametric Bayesian networksabstractSemiparametric Bayesian networks (SPBNs) integrate parametric and non-parametric probabilistic models, offering flexibility in learning complex data distributions from samples. In particular, kernel density estimators (KDEs) are employed for the non-parametric component. Under the assumption of data normality, the normal rule is used to determine the bandwidth matrix for KDEs in SPBNs. This matrix is the critical hyperparameter that controls the trade-off between bias and variance. However, real-world data often deviates from normality, potentially leading to suboptimal density estimation and reduced predictive performance. This paper presents the theoretical framework for applying state-of-the-art bandwidth selectors to SPBNs and evaluates their impact on the performance. We explore cross-validation and plug-in selectors approaches, assessing their effectiveness in enhancing the learning capability and applicability of SPBNs. To support this investigation, we have extended the open-source package PyBNesian for SPBNs with additional bandwidth selection techniques and conducted extensive experimental analyses. Our results demonstrate that the proposed bandwidth selectors leverage larger sample sizes more effectively than the normal rule, which, despite its robustness, plateaus with more data. In particular, unbiased cross-validation generally outperforms the normal rule, highlighting its advantage in large sample size scenarios. • We establish guarantees for SPBN density estimation using plug-in and CV bandwidths. • Bandwidth selection boosts SPBN parameter/structure learning beyond normal rule. • Proposed bandwidths improve with sample size, unlike the stagnating normal rule. • Despite stagnating performance, normal rule is robust/effective in structure learning. • Unbiased CV balances cost/performance; best overall, especially for large samples. Victor Alejandre, Concha Bielza, Pedro Larrañaga |
Inf. Sci. | 2 |
| 2026 | Transfer learning for nonparametric Bayesian networksabstractThis paper introduces two transfer learning methodologies for estimating nonparametric Bayesian networks under scarce data. We propose two algorithms, a constraint-based structure learning method, called PC-stable-transfer learning (PCS-TL), and a score-based method, called hill climbing transfer learning (HC-TL). We also define particular metrics to tackle the negative transfer problem in each of them, a situation in which transfer learning has a negative impact on the model's performance. Then, for the parameters, we propose a log-linear pooling approach. For the evaluation, we learn kernel density estimation Bayesian networks, a type of nonparametric Bayesian network, and compare their transfer learning performance with the models alone. To do so, we sample data from small, medium and large-sized synthetic networks and datasets from the UCI Machine Learning repository. Then, we add noise and modifications to these datasets to test their ability to avoid negative transfer. To conclude, we perform a Friedman test with a Bergmann-Hommel post-hoc analysis to show statistical proof of the enhanced experimental behavior of our methods. Thus, PCS-TL and HC-TL demonstrate to be reliable algorithms for improving the learning performance of a nonparametric Bayesian network with scarce data, which in real industrial environments implies a reduction in the required time to deploy the network. Rafael Sojo, Pedro Larrañaga, Concha Bielza |
Knowl. Based Syst. | 3 |
| 2025 | Discovering Genetic Variants in Hypertrophic Cardiomyopathy With Multiple Machine Learning TechniquesabstractHypertrophic cardiomyopathy is known to have strong genetic foundations. However, only some studies have addressed the complex network of co-expressed genes and variants that modify the phenotype. Machine learning methods offer robust information discovery when dealing with high-dimensional datasets. We aimed to perform relevance and interaction analysis on genetic variants from hypertrophic cardiomyopathy patients using diverse machine learning techniques, with the following stages: (a) Statistical univariate techniques (with various $p$-value adjustment methods) identified relevant variants; (b) Linear classifiers (support vector machines, Fisher discriminant analysis) provided combined relevance based on feature weights; (c) Informative variable identifier method and Bayesian networks explained inter-variant relationships; (d) Manifold learning of low-dimensional latent spaces gave interpretable representations of groups; (e) Linkage disequilibrium matrices and frequency tables discovered associations between variants. We analyzed 61 patients and 67 controls with genetic information comprising 216 variants from a genetic panel of 15 genes. Across all methodologies, ten variants were consistently identified as significant, with 22 total variants significant in at least three out of five methods. Machine learning has been found to detect disease-associated variants, including pathogenic founder variants (11:47357494, 11:47360070, 11:47372137). This methodology allows for identifying potential disease modulators while accounting for relevance and interactions among variants. Dafne Lozano-Paredes, Luis Bote-Curiel, María Sabater-Molina, Concha Bielza, Juan-Ramón Gimeno-Blanes, Sergio Muñoz-Romero, Francisco Javier Gimeno-Blanes, Pedro Larrañaga, José Luis Rojo-Álvarez |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | BN-BacArena: Bayesian network extension of BacArena for the dynamic simulation of microbial communitiesabstractMOTIVATION: Simulating gut microbial dynamics is extremely challenging. Several computational tools, notably the widely used BacArena, enable modeling of dynamic changes in the microbial environment. These methods, however, do not comprehensively account for microbe-microbe stimulant or inhibitory effects or for nutrient-microbe inhibitory effects, typically observed in different compounds present in the daily diet. RESULTS: Here, we present BN-BacArena, an extension of BacArena consisting on the incorporation within the native computational framework of a Bayesian network model that accounts for microbe-microbe and nutrient-microbe interactions. Using in vitro experiments, 16S rRNA gene sequencing data and nutritional composition of 55 foods, the output Bayesian network showed 23 significant nutrient-bacteria interactions, suggesting the importance of compounds such as polyols, ascorbic acid, polyphenols and other phytochemicals, and 40 bacteria-bacteria significant relationships. With test data, BN-BacArena demonstrates a statistically significant improvement over BacArena to predict the time-dependent relative abundance of bacterial species involved in the gut microbiota upon different nutritional interventions. As a result, BN-BacArena opens new avenues for the dynamic modeling and simulation of the human gut microbiota metabolism. AVAILABILITY AND IMPLEMENTATION: MATLAB and R code are available in https://github.com/PlanesLab/BN-BacArena. Telmo Blasco, Francesco Balzerani, Luis Vitores Valcarcel, Pedro Larrañaga, Concha Bielza, María Pilar Francino, José Ángel Rufián-Henares, Francisco J. Planes, Sergio Pérez-Burillo |
Bioinform. | 5 |
| 2024 | EDAspy: An extensible python package for estimation of distribution algorithmsabstractEstimation of distribution algorithms (EDAs) are a type of evolutionary algorithms where a probabilistic model is learned and sampled in each iteration. EDAspy provides different state-of-the-art implementations of EDAs including the recent semiparametric EDA. The implementations are modularly built, allowing for easy extension and the selection of different alternatives, as well as interoperability with new components. EDAspy is totally free and open-source under the MIT license. Vicente P. Soloviev, Pedro Larrañaga, Concha Bielza |
Neurocomputing | 3 |
| 2024 | Estimation of Distribution Algorithms in Machine Learning: A SurveyabstractThe automatic induction of machine learning models capable of addressing supervised learning, feature selection, clustering and reinforcement learning problems requires sophisticated intelligent search procedures. These searches are usually performed in the possible model structure spaces, leading to combinatorial optimization problems, and in the parameter spaces, where it is necessary to solve continuous optimization problems. This paper reviews how the estimation of distribution algorithms, a kind of evolutionary algorithm, can be used to address these problems. Topics include preprocessing, mining association rules, selecting variables, searching for the optimal supervised learning model (both probabilistic and nonprobabilistic models), finding the best hierarchical, partitional or probabilistic clustering, obtaining the optimal policy in reinforcement learning and performing inference and structural learning in Bayesian networks for association discovery. Interesting guidelines for future work in this area are also provided. Pedro Larrañaga, Concha Bielza |
IEEE Trans. Evol. Comput. | 2 |
| 2024 | Semiparametric Estimation of Distribution Algorithms for Continuous OptimizationabstractTraditional estimation of distribution algorithms (EDAs) often use Gaussian densities to optimize continuous functions, such as the estimation of Gaussian network algorithms (EGNAs) which use Gaussian Bayesian networks (GBNs). However, this assumes a parametric density function, and, in GBNs, linear dependencies between variables. Furthermore, the EGNA baseline learns a GBN at each iteration based on the best individuals in the last iteration, which may lead to local optimum convergence or large variance between solutions across multiple independent runs of the algorithm. In this work we propose a semiparametric EDA in which the restriction of assuming Gaussianity in the variables is relaxed using semiparametric Bayesian networks, in which nodes estimated by kernels coexist with nodes that assume Gaussianity, and the algorithm itself is able to determine where to use each type of node. Additionally, our approach takes into account information from several past iterations to learn the semiparametric Bayesian network from which the new solutions are sampled in each iteration. The empirical results show that semiparametric EDAs are a useful tool for continuous scenarios compared to different kinds of EDAs and other optimization techniques in continuous environments. Vicente P. Soloviev, Concha Bielza, Pedro Larrañaga |
IEEE Trans. Evol. Comput. | 2 |
| 2024 | Feature Saliencies in Asymmetric Hidden Markov ModelsabstractMany real-life problems are stated as nonlabeled high-dimensional data. Current strategies to select features are mainly focused on labeled data, which reduces the options to select relevant features for unsupervised problems, such as clustering. Recently, feature saliency models have been introduced and developed as clustering models to select and detect relevant variables/features as the model is learned. Usually, these models assume that all variables are independent, which narrows their applicability. This article introduces asymmetric hidden Markov models with feature saliencies, i.e., models capable of simultaneously determining during their learning phase relevant variables/features and probabilistic relationships between variables. The proposed models are compared with other state-of-the-art approaches using synthetic data and real data related to grammatical face videos and wear in ball bearings. We show that the proposed models have better or equal fitness than other state-of-the-art models and provide further data insights. Carlos Puerto-Santana, Pedro Larrañaga, Concha Bielza |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Variational Quantum Algorithm Parameter Tuning with Estimation of Distribution AlgorithmsabstractVariational quantum algorithms (VQAs) are hybrid approaches between classical and quantum computation, where a classical optimizer proposes parameter configurations for a quantum parametric circuit which is iteratively sampled. The overall performance of the algorithm depends on how the classical optimizer tunes the parameters of the quantum circuit. Several gradient-free and gradient-based approaches have been proposed in the literature to face this task. Estimation of distribution algorithms (EDAs) are a type of evolutionary algorithms where a probabilistic model is updated and sampled at each generation to optimize a cost function. EDAs have shown to be able to achieve good solutions in a reasonable computation time for different optimization problems, and thus, we believe that this algorithm can be a good option to overcome VQAs challenges such as the Barren plateaus phenomenon. In this paper, we study the use of three different EDAs, characterized by different probabilistic model complexities, to tune the parameters of two different VQAs to solver the Max Cut problem and to a VQA to simulate the behaviour of a molecule. Three EDA variants are compared to some state-of-the-art optimizers widely used for this task. Our results show statistical significant improvement of the EDA variants compared to different optimizers, and identify the VQAs characteristics that best fit to each EDA type. We also perform an analysis of the main EDAs hyper-parameters. Vicente P. Soloviev, Pedro Larrañaga, Concha Bielza |
CEC | 3 |
| 2023 | Causal reinforcement learning based on Bayesian networks applied to industrial settingsabstractThe increasing amount of real-time data collected from sensors in industrial environments has accelerated the application of machine learning in decision-making. Reinforcement learning (RL) is a powerful tool to find optimal policies for achieving a given goal. However, RL’s typical application is risky and insufficient in environments where actions can have irreversible consequences and require interpretability and fairness. While new trends in RL may provide guidance based on expert knowledge, they do not often consider uncertainty or include prior knowledge in the learning process. We propose a causal reinforcement learning alternative based on Bayesian networks (RLBNs) to address this challenge. The RLBN simultaneously models a policy and takes advantage of the joint distribution of the state and action space, reducing uncertainty in unknown situations. We propose a training algorithm for the network’s parameters and structure based on the reward function and likelihood of the effects and measurements taken. Our experiment with the CartPole benchmark and industrial fouling using ordinary differential equations (ODEs) demonstrates that RLBNs are interpretable, secure, flexible, and more robust than their competitors. Our contributions include a novel method that incorporates expert knowledge into the decision-making engine. It uses Bayesian networks with a predefined structure as a causal graph and a hybrid learning strategy that considers both likelihood and reward. This would avoid losing the virtues of the Bayesian network. Gabriel Valverde, David Quesada, Pedro Larrañaga, Concha Bielza |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Efficient search for relevance explanations using MAP-independence in Bayesian networksabstract-independence is a novel concept concerned with explaining the (ir)relevance of intermediate nodes for maximum a posteriori () computations in Bayesian networks. Building upon properties of -independence, we introduce and experiment with methods for finding sets of relevant nodes using both an exhaustive and a heuristic approach. Our experiments show that these properties significantly speed up run time for both approaches. In addition, we link -independence to defeasible reasoning, a type of reasoning that analyses how new evidence may invalidate an already established conclusion. Ways to present users with an explanation using -independence are also suggested. Enrique Valero-Leal, Concha Bielza, Pedro Larrañaga, Silja Renooij |
Int. J. Approx. Reason. | 2 |
| 2023 | Constraint-based and hybrid structure learning of multidimensional continuous-time Bayesian network classifiersabstractLearning the structure of continuous-time Bayesian networks directly from data has traditionally been performed using score-based structure learning algorithms. Only recently has a constraint-based method been proposed, proving to be more suitable under specific settings, as in modelling systems with variables having more than two states. As a result, studying diverse structure learning algorithms is essential to learn the most appropriate models according to data characteristics and task-related priorities, such as learning speed or accuracy. This article proposes alternative algorithms for learning multidimensional continuous-time Bayesian network classifiers, introducing, for the first time, constraint-based and hybrid algorithms for these models. Nevertheless, these contributions also apply to the simpler one-dimensional classification problem for which only score-based solutions exist in the literature. More specifically, the aforementioned constraint-based structure learning algorithm is first adapted to the supervised classification setting. Then, a novel algorithm of this kind, specifically tailored for the multidimensional classification problem, is presented to improve the learning times for the induction of multidimensional classifiers. Finally, a hybrid algorithm is introduced, attempting to combine the strengths of the score- and constraint-based approaches. Experiments with synthetic and real-world data are performed not only to validate the capabilities of the proposed algorithms but also to conduct a comparative study of the available competitors. Carlos Villa-Blanco, Alessandro Bregoli, Concha Bielza, Pedro Larrañaga, Fabio Stella |
Int. J. Approx. Reason. | 3 |
| 2023 | Feature subset selection in data-stream environments using asymmetric hidden Markov models and novelty detection
Carlos Puerto-Santana, Pedro Larrañaga, Concha Bielza |
Neurocomputing | 3 |
| 2023 | Learning massive interpretable gene regulatory networks of the human brain by merging Bayesian networksabstractWe present the Fast Greedy Equivalence Search (FGES)-Merge, a new method for learning the structure of gene regulatory networks via merging locally learned Bayesian networks, based on the fast greedy equivalent search algorithm. The method is competitive with the state of the art in terms of the Matthews correlation coefficient, which takes into account both precision and recall, while also improving upon it in terms of speed, scaling up to tens of thousands of variables and being able to use empirical knowledge about the topological structure of gene regulatory networks. To showcase the ability of our method to scale to massive networks, we apply it to learning the gene regulatory network for the full human genome using data from samples of different brain structures (from the Allen Human Brain Atlas). Furthermore, this Bayesian network model should predict interactions between genes in a way that is clear to experts, following the current trends in explainable artificial intelligence. To achieve this, we also present a new open-access visualization tool that facilitates the exploration of massive networks and can aid in finding nodes of interest for experimental tests. Nikolas Bernaola, Mario Michiels, Pedro Larrañaga, Concha Bielza |
PLoS Comput. Biol. | 4 |
| 2022 | Piecewise forecasting of nonlinear time series with model tree dynamic Bayesian networksabstractWhen modelling multivariate continuous time series, a common issue is to find that the original processes that generated the data are nonlinear or that they drift away from the original distribution as the system evolves over time. In these scenarios, using a linear model such as a Gaussian dynamic Bayesian network (DBN) can result in severe forecasting inaccuracies due to the structure and parameters of the model remaining constant without considering either the elapsed time or the state of the system. To approach this problem, we propose a hybrid model that combines a model tree with DBNs. The model first divides the original data set into different scenarios based on the splits made by the tree over the system variables and then performs a piecewise regression in each branch of the tree to obtain nonlinear forecasts. The experimental results on three different datasets show that our model outperforms standard DBN models when dealing with nonlinear processes and is competitive with state-of-the-art time series forecasting methods. David Quesada, Concha Bielza, Pedro Fontán, Pedro Larrañaga |
Int. J. Intell. Syst. | 2 |
| 2022 | Multipartition clustering of mixed data with Bayesian networksabstractReal-world applications often involve multifaceted data with several reasonable interpretations. To cluster this data, we need methods that are able to produce multiple clustering solutions. To this purpose, it is interesting to learn a finite mixture model with multiple latent variables, where each latent variable represents a unique way to partition the data. However, although there is an extensive literature on multipartition clustering methods for categorical data and for continuous data, there is a lack of work for mixed data. In this paper, we propose a multipartition clustering method that is able to efficiently deal with mixed data by exploiting the Bayesian network factorization and the variational Bayes framework. We show the flexibility and applicability of the proposed method by solving clustering, density estimation, and missing data imputation tasks in real-world data sets. For reproducibility, all code, data, and results can be found in the following public repository: https://github.com/ferjorosa/mpc-mixed. Fernando Rodriguez-Sanchez, Concha Bielza, Pedro Larrañaga |
Int. J. Intell. Syst. | 2 |
| 2022 | PyBNesian: An extensible python package for Bayesian networksabstractBayesian networks are probabilistic graphical models that are commonly used to represent the uncertainty in data. The PyBNesian package provides an implementation for many different types of Bayesian network models and some variants, such as conditional Bayesian networks and dynamic Bayesian networks. In addition, the package can be easily extended with new components that can interoperate with those already implemented. Furthermore, the package also implements other related models such as kernel density estimation using OpenCL 1.2+ to enable GPU acceleration. PyBNesian is totally free and open-source under the MIT license. David Atienza 0001, Concha Bielza, Pedro Larrañaga |
Neurocomputing | 2 |
| 2022 | Asymmetric HMMs for Online Ball-Bearing Health AssessmentsabstractThe degradation of critical components inside large industrial assets, such as ball-bearings, has a negative impact on production facilities, reducing the availability of assets due to an unexpectedly high failure rate. Machine learning-based monitoring systems can estimate the remaining useful life (RUL) of ball bearings, reducing the downtime by early failure detection. However, traditional approaches for predictive systems require run-to-failure (RTF) data as training data, which in real scenarios can be scarce and expensive to obtain as the expected useful life could be measured in years. Therefore, to overcome the need of RTF, we propose a new methodology based on online novelty detection and asymmetrical hidden Markov models (As-HMMs) to work out the health assessment. This new methodology does not require previous RTF data and can adapt to natural degradation of mechanical components over time in data-stream and online environments. As the system is designed to work online within the electrical cabinet of machines, it has to be deployed using embedded electronics. Therefore, a performance analysis of As-HMM is presented to detect the strengths and critical points of the algorithm. To validate our approach, we use real life ball-bearing data sets and compare our methodology with other methodologies where no RTF data are needed and check the advantages in RUL prediction and health monitoring. As a result, we showcase a complete end-to-end solution from the sensor to actionable insights regarding RUL estimation toward maintenance application in real industrial environments. Carlos Puerto-Santana, Concha Bielza, Javier Diaz-Rozo, Guillem Ramirez-Gargallo, Filippo Mantovani, Gaizka Virumbrales, Jesús Labarta, Pedro Larrañaga |
IEEE Internet Things J. | 2 |
| 2022 | Semiparametric Bayesian networks
David Atienza 0001, Concha Bielza, Pedro Larrañaga |
Inf. Sci. | 2 |
| 2022 | Autoregressive Asymmetric Linear Gaussian Hidden Markov Models
Carlos Puerto-Santana, Pedro Larrañaga, Concha Bielza |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Quantum-Inspired Estimation Of Distribution Algorithm To Solve The Travelling Salesman ProblemabstractA novel Quantum-Inspired Estimation of Distribution Algorithm (QIEDA) is proposed to solve the Travelling Salesman Problem (TSP). The QIEDA uses a modified version of the W state quantum circuits to sample new solutions during the algorithm runtime. The algorithm behaviour is compared with other state-of-the-art population-based algorithms. QIEDA convergence is faster than other algorithms, and the obtained solutions improve as the size of the problem increases. Moreover, we show that quantum noise enhances the search of an optimal solution. Because quantum computers differ from each other, partly due to the topology that distributes the qubits, the computational cost of executing the QIEDA in different topologies is analyzed and an ideal topology is proposed for the TSP solved with the QIEDA. Vicente P. Soloviev, Concha Bielza, Pedro Larrañaga |
CEC | 2 |
| 2021 | Long-term forecasting of multivariate time series in industrial furnaces with dynamic Gaussian Bayesian networks
David Quesada, Gabriel Valverde, Pedro Larrañaga, Concha Bielza |
Eng. Appl. Artif. Intell. | 4 |
| 2021 | Multidimensional continuous time Bayesian network classifiersabstractThe multidimensional classification of multivariate time series deals with the assignment of multiple classes to time-ordered data described by a set of feature variables. Although this challenging task has received almost no attention in the literature, it is present in a wide variety of domains, such as medicine, finance or industry. The complexity of this problem lies in two nontrivial tasks, the learning with multivariate time series in continuous time and the simultaneous classification of multiple class variables that may show dependencies between them. These can be addressed with different strategies, but most of them may involve a difficult preprocessing of the data, high space and classification complexity or ignoring useful interclass dependencies. Additionally, no attention has been given to the development of new multidimensional classifiers of time series based on probabilistic graphical models, even though transparent models can facilitate further understanding of the domain. In this paper, a novel probabilistic graphical model is proposed, which is able to classify a discrete multivariate temporal sequence into multiple class variables while modeling their dependencies. This model extends continuous time Bayesian networks to the multidimensional classification problem, which are able to explicitly represent the behavior of time series that evolve over continuous time. Different methods for the learning of the parameters and structure of the model are presented, and numerical experiments on synthetic and real-world data show encouraging results in terms of performance and learning time with respect to independent classifiers, the current alternative approach under the continuous time Bayesian network paradigm. Carlos Villa-Blanco, Pedro Larrañaga, Concha Bielza |
Int. J. Intell. Syst. | 3 |
| 2021 | BayeSuites: An open web framework for massive Bayesian networks focused on neuroscienceabstractBayeSuites is the first web framework for learning, visualizing, and interpreting Bayesian networks (BNs) that can scale to tens of thousands of nodes while providing fast and friendly user experience. All the necessary features that enable this are reviewed in this paper; these features include scalability, extensibility, interoperability, ease of use, and interpretability. Scalability is the key factor in learning and processing massive networks within reasonable time; for a maintainable software open to new functionalities, extensibility and interoperability are necessary. Ease of use and interpretability are fundamental aspects of model interpretation, fairly similar to the case of the recent explainable artificial intelligence trend. We present the capabilities of our proposed framework by highlighting a real example of a BN learned from genomic data obtained from Allen Institute for Brain Science. The extensibility properties of the software are also demonstrated with the help of our BN-based probabilistic clustering implementation, together with another genomic-data example. Mario Michiels, Pedro Larrañaga, Concha Bielza |
Neurocomputing | 3 |
| 2021 | Bayesian networks for interpretable machine learning and optimization
Bojan Mihaljevic, Concha Bielza, Pedro Larrañaga |
Neurocomputing | 2 |
| 2020 | Machine-tool condition monitoring with Gaussian mixture models-based dynamic probabilistic clustering
Javier Diaz-Rozo, Concha Bielza, Pedro Larrañaga |
Eng. Appl. Artif. Intell. | 2 |
| 2020 | On generating random Gaussian graphical models
Irene Córdoba-Sánchez, Gherardo Varando, Concha Bielza, Pedro Larrañaga |
Int. J. Approx. Reason. | 3 |
| 2019 | Learning tractable Bayesian networks in the space of elimination orders
Marco Benjumeda, Concha Bielza, Pedro Larrañaga |
Artif. Intell. | 2 |
| 2019 | Circular Bayesian classifiers using wrapped Cauchy distributions
Ignacio Leguey, Concha Bielza, Pedro Larrañaga |
Data Knowl. Eng. | 2 |
| 2019 | A circular-linear dependence measure under Johnson-Wehrly distributions and its application in Bayesian networks
Ignacio Leguey, Pedro Larrañaga, Concha Bielza, Shogo Kato |
Inf. Sci. | 3 |
| 2019 | Tractable learning of Bayesian networks from partially observed data
Marco Benjumeda, Sergio Luengo-Sanchez, Pedro Larrañaga, Concha Bielza |
Pattern Recognit. | 4 |
| 2018 | A Fast Metropolis-Hastings Method for Generating Random Correlation Matrices
Irene Córdoba-Sánchez, Gherardo Varando, Concha Bielza, Pedro Larrañaga |
IDEAL (1) | 3 |
| 2018 | Multi-dimensional Bayesian Network Classifier Trees
Santiago Gil-Begue, Pedro Larrañaga, Concha Bielza |
IDEAL (1) | 3 |
| 2018 | Towards a supervised classification of neocortical interneuron morphologiesabstractBACKGROUND: The challenge of classifying cortical interneurons is yet to be solved. Data-driven classification into established morphological types may provide insight and practical value. RESULTS: We trained models using 217 high-quality morphologies of rat somatosensory neocortex interneurons reconstructed by a single laboratory and pre-classified into eight types. We quantified 103 axonal and dendritic morphometrics, including novel ones that capture features such as arbor orientation, extent in layer one, and dendritic polarity. We trained a one-versus-rest classifier for each type, combining well-known supervised classification algorithms with feature selection and over- and under-sampling. We accurately classified the nest basket, Martinotti, and basket cell types with the Martinotti model outperforming 39 out of 42 leading neuroscientists. We had moderate accuracy for the double bouquet, small and large basket types, and limited accuracy for the chandelier and bitufted types. We characterized the types with interpretable models or with up to ten morphometrics. CONCLUSION: Except for large basket, 50 high-quality reconstructions sufficed to learn an accurate model of a type. Improving these models may require quantifying complex arborization patterns and finding correlates of bouton-related features. Our study brings attention to practical aspects important for neuron classification and is readily reproducible, with all code and data available online. Bojan Mihaljevic, Pedro Larrañaga, Ruth Benavides-Piccione, Sean L. Hill, Javier DeFelipe, Concha Bielza |
BMC Bioinform. | 6 |
| 2018 | Tractability of most probable explanations in multidimensional Bayesian network classifiers
Marco Benjumeda, Concha Bielza, Pedro Larrañaga |
Int. J. Approx. Reason. | 2 |
| 2018 | Clustering of Data Streams With Dynamic Gaussian Mixture Models: An IoT Application in Industrial ProcessesabstractIn industrial Internet of Things applications with sensors sending dynamic process data at high speed, producing actionable insights at the right time is challenging. A key problem concerns processing a large amount of data, while the underlying dynamic phenomena related to the machine is possibly evolving over time due to factors, such as degradation. This makes any actionable model become obsolete and necessary to be updated. To cope with this problem, in this paper we propose a new unsupervised learning algorithm based on Gaussian mixture models called Gaussian-based dynamic probabilistic clustering (GDPC) mainly based on integrating and adapting three well known algorithms for use in dynamic scenarios: the expectation-maximization (EM) algorithm to estimate the model parameters and the Page-Hinkley test and Chernoff bound to detect concept drifts. Unlike other unsupervised methods, the model induced by the GDPC provides the membership probabilities of each instance to each cluster. This allows to determine, through a Brier score analysis, the robustness of the instance assignment and its evolution each time a concept drift is detected. Also, the algorithm works with very little data and significantly less computing power being able to decide whether (and when) to change the model. The algorithm is tested using synthetic data and data streams from an industrial testbed, where different operational states are automatically identified, giving good results in terms of classification accuracy, sensitivity, and specificity. Javier Diaz-Rozo, Concha Bielza, Pedro Larrañaga |
IEEE Internet Things J. | 2 |
| 2018 | A regularity index for dendrites - local statistics of a neuron's input spaceabstractNeurons collect their inputs from other neurons by sending out arborized dendritic structures. However, the relationship between the shape of dendrites and the precise organization of synaptic inputs in the neural tissue remains unclear. Inputs could be distributed in tight clusters, entirely randomly or else in a regular grid-like manner. Here, we analyze dendritic branching structures using a regularity index R, based on average nearest neighbor distances between branch and termination points, characterizing their spatial distribution. We find that the distributions of these points depend strongly on cell types, indicating possible fundamental differences in synaptic input organization. Moreover, R is independent of cell size and we find that it is only weakly correlated with other branching statistics, suggesting that it might reflect features of dendritic morphology that are not captured by commonly studied branching statistics. We then use morphological models based on optimal wiring principles to study the relation between input distributions and dendritic branching structures. Using our models, we find that branch point distributions correlate more closely with the input distributions while termination points in dendrites are generally spread out more randomly with a close to uniform distribution. We validate these model predictions with connectome data. Finally, we find that in spatial input distributions with increasing regularity, characteristic scaling relationships between branching features are altered significantly. In summary, we conclude that local statistics of input distributions and dendrite morphology depend on each other leading to potentially cell type specific branching features. Laura Anton-Sanchez, Felix Effenberger, Concha Bielza, Pedro Larrañaga, Hermann Cuntz |
PLoS Comput. Biol. | 3 |
| 2018 | 3D morphology-based clustering and simulation of human pyramidal cell dendritic spinesabstractThe dendritic spines of pyramidal neurons are the targets of most excitatory synapses in the cerebral cortex. They have a wide variety of morphologies, and their morphology appears to be critical from the functional point of view. To further characterize dendritic spine geometry, we used in this paper over 7,000 individually 3D reconstructed dendritic spines from human cortical pyramidal neurons to group dendritic spines using model-based clustering. This approach uncovered six separate groups of human dendritic spines. To better understand the differences between these groups, the discriminative characteristics of each group were identified as a set of rules. Model-based clustering was also useful for simulating accurate 3D virtual representations of spines that matched the morphological definitions of each cluster. This mathematical approach could provide a useful tool for theoretical predictions on the functional features of human pyramidal neurons based on the morphology of dendritic spines. Sergio Luengo-Sanchez, Isabel Fernaud, Concha Bielza, Ruth Benavides-Piccione, Pedro Larrañaga, Javier DeFelipe |
PLoS Comput. Biol. | 3 |
| 2017 | Architecture for anomaly detection in a laser heating surface processabstractAnomaly detection is an increasingly common task in many industrial environments. Cyber-physical systems stand out in this field due to their unique position in industrial areas. This paper introduces a new architecture aimed to detect anomalies in a real laser heating surface process, which is designed for field-programmable gate arrays (FPGAs). The FPGA design offers advantages of highly parallelized and pipelined architectures. The system will classify one process into normal or abnormal taking into account spatial information about where the laser spot is. The proposed design estimates a probability density function from data; then it performs an image convolution transforming the probability density function into a kernel density estimation function. This estimated function should be able to classify in real time. Javier Mesonero, Concha Bielza, Pedro Larrañaga |
ETFA | 2 |
| 2017 | Frobenius Norm Regularization for the Multivariate Von Mises DistributionabstractPenalizing the model complexity is necessary to avoid overfitting when the number of data samples is low with respect to the number of model parameters. In this paper, we introduce a penalization term that places an independent prior distribution for each parameter of the multivariate von Mises distribution. We also propose a circular distance that can be used to estimate the Kullback–Leibler divergence between any two circular distributions as goodness-of-fit measure. We compare the resulting regularized von Mises models on synthetic data and real neuroanatomical data to show that the distribution fitted using the penalized estimator generally achieves better results than nonpenalized multivariate von Mises estimator. Luis Rodriguez-Lujan, Pedro Larrañaga, Concha Bielza |
Int. J. Intell. Syst. | 3 |
| 2016 | Hybrid Gaussian and von Mises Model-Based ClusteringabstractData collected about a phenomenon often measures its magnitude and direction. The most common approach to clustering this data assumes that directional data can be modeled as Gaussian. However, directional data has special properties that conventional statistics cannot handle. To deal with them, other approaches like the von Mises distribution must be applied. In this paper we present a new model based on mixtures of Bayesian networks to simultaneously cluster both linear and directional data. Sergio Luengo-Sanchez, Concha Bielza, Pedro Larrañaga |
ECAI | 2 |
| 2016 | Mining multi-dimensional concept-drifting data streams using Bayesian network classifiersabstractIn recent years, a plethora of approaches have been proposed to deal with the increasingly challenging task of mining concept-drifting data streams. However, most of these approaches can only be applied to uni-dimensional classification problems where each input instance has to be assigned to a sin gle output class variable. The problem of mining multi-dimensional data streams, which includes multiple output class variables, is largely unexplored and only few streaming multi-dimensional approaches have been recently introduced. In this paper, we propose a novel adaptive method, named Locally Adaptive-MB-MBC (LA-MB-MBC), for mining streaming multi-dimensional data. To this end, we make use of multi-dimensional Bayesian network classifiers (MBCs) as models. Basically, LA-MB-MBC monitors the concept drift over time using the average log-likelihood score and the Page-Hinkley test. Then, if a concept drift is detected, LA-MB-MBC adapts the current MBC network locally around each changed node. An experimental study carried out using synthetic multi-dimensional data streams shows the merits of the proposed method in terms of concept drift detection as well as classification performance. Hanen Borchani, Pedro Larrañaga, João Gama 0001, Concha Bielza |
Intell. Data Anal. | 4 |
| 2016 | Decision functions for chain classifiers based on Bayesian networks for multi-label classification
Gherardo Varando, Concha Bielza, Pedro Larrañaga |
Int. J. Approx. Reason. | 2 |
| 2016 | Genetic algorithms and Gaussian Bayesian networks to uncover the predictive core set of bibliometric indicesabstractThe diversity of bibliometric indices today poses the challenge of exploiting the relationships among them. Our research uncovers the best core set of relevant indices for predicting other bibliometric indices. An added difficulty is to select the role of each variable, that is, which bibliometric indices are predictive variables and which are response variables. This results in a novel multioutput regression problem where the role of each variable (predictor or response) is unknown beforehand. We use Gaussian Bayesian networks to solve the this problem and discover multivariate relationships among bibliometric indices. These networks are learnt by a genetic algorithm that looks for the optimal models that best predict bibliometric data. Results show that the optimal induced Gaussian Bayesian networks corroborate previous relationships between several indices, but also suggest new, previously unreported interactions. An extended analysis of the best model illustrates that a set of 12 bibliometric indices can be accurately predicted using only a smaller predictive core subset composed of citations, g‐index, q2‐index, and hr‐index. This research is performed using bibliometric data on Spanish full professors associated with the computer science area. Alfonso Ibáñez, Rubén Armañanzas, Concha Bielza, Pedro Larrañaga |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2015 | Towards Gaussian Bayesian Network Fusion
Irene Córdoba-Sánchez, Concha Bielza, Pedro Larrañaga |
ECSQARU | 2 |
| 2015 | Classifying GABAergic interneurons with semi-supervised projected model-based clustering
Bojan Mihaljevic, Ruth Benavides-Piccione, Luis Guerra, Javier DeFelipe, Pedro Larrañaga, Concha Bielza |
Artif. Intell. Medicine | 6 |
| 2015 | Guest editors introduction: special issue of the ECMLPKDD 2015 journal track
Concha Bielza, João Gama 0001, Alípio Mário Jorge, Indre Zliobaite |
Data Min. Knowl. Discov. | 1 |
| 2015 | Recent Advances in Probabilistic Graphical ModelsabstractProbabilistic graphical models constitute a fundamental tool for the development of intelligent systems.They provide a sound and well-founded approach for performing inference and belief updating in complex domains endowed with uncertainty.A probabilistic graphical model is the result of the combination of a qualitative component (a graph) encoding conditional independence relationships among the variables in the system and a quantitative component consisting of a collection of local probability distributions matching the independence properties specified by the graph.The union of the two components provides a compact representation of the joint probability distribution over the domain being modeled.Bayesian networks are the most prominent type of probabilistic graphical models and have experienced a remarkable methodological development during the past two decades.This has come along with a wide variety of successful applications in different domains.Regardless of the increasing interest in the area, probabilistic graphical models are still facing a number of challenges, covering modeling, inference and learning.This special issue contains seven papers that contribute to the methodological development beyond the current state-of-the-art knowledge.Five of the papers were selected among the contributions presented at the 15th Conference of the Spanish Association of Artificial Intelligence (CAEPIA'2013, Madrid, Spain, September 17-20, 2013).These papers were substantially extended and went through a new and strict review process.The other two papers were invited contributions not presented at the conference and went through the same review process.The first three papers deal with hybrid models, where both discrete and continuous variables coexist.Lucas and Hommersom propose a framework for handling causal independence covering discrete and continuous variables simultaneously.The methodology is based on the convolution concept from probability theory, and the Concha Bielza, Serafín Moral, Antonio Salmerón |
Int. J. Intell. Syst. | 1 |
| 2015 | Conditional Density Approximations with Mixtures of PolynomialsabstractMixtures of polynomials (MoPs) are a nonparametric density estimation technique especially designed for hybrid Bayesian networks with continuous and discrete variables. Algorithms to learn one- and multidimensional (marginal) MoPs from data have recently been proposed. In this paper, we introduce two methods for learning MoP approximations of conditional densities from data. Both approaches are based on learning MoP approximations of the joint density and the marginal density of the conditioning variables, but they differ as to how the MoP approximation of the quotient of the two densities is found. We illustrate and study the methods using data sampled from known parametric distributions, and demonstrate their applicability by learning models based on real neuroscience data. Finally, we compare the performance of the proposed methods with an approach for learning mixtures of truncated basis functions (MoTBFs). The empirical results show that the proposed methods generally yield models that are comparable to or significantly better than those found using the MoTBF-based method. Gherardo Varando, Pedro L. López-Cruz, Thomas D. Nielsen, Pedro Larrañaga, Concha Bielza |
Int. J. Intell. Syst. | 5 |
| 2015 | Decision boundary for discrete Bayesian network classifiers
Gherardo Varando, Concha Bielza, Pedro Larrañaga |
J. Mach. Learn. Res. | 2 |
| 2015 | Guest Editors introduction: special issue of the ECMLPKDD 2015 journal track
Concha Bielza, João Gama 0001, Alípio Mário Jorge, Indre Zliobaite |
Mach. Learn. | 1 |
| 2015 | Directional naive Bayes classifiers
Pedro L. López-Cruz, Concha Bielza, Pedro Larrañaga |
Pattern Anal. Appl. | 2 |
| 2014 | Semi-supervised projected model-based clustering
Luis Guerra, Concha Bielza, Víctor Robles, Pedro Larrañaga |
Data Min. Knowl. Discov. | 2 |
| 2014 | Learning mixtures of polynomials of multidimensional probability densities from data using B-spline interpolation
Pedro L. López-Cruz, Concha Bielza, Pedro Larrañaga |
Int. J. Approx. Reason. | 2 |
| 2014 | Bayesian network modeling of the consensus between experts: An application to neuron classification
Pedro L. López-Cruz, Pedro Larrañaga, Javier DeFelipe, Concha Bielza |
Int. J. Approx. Reason. | 4 |
| 2014 | Cost-sensitive selective naive Bayes classifiers for predicting the increase of the h-index for scientific journals
Alfonso Ibáñez, Concha Bielza, Pedro Larrañaga |
Neurocomputing | 2 |
| 2014 | Multi-label classification with Bayesian network-based chain classifiers
Luis Enrique Sucar, Concha Bielza, Eduardo F. Morales 0001, Pablo Hernandez-Leal, Julio H. Zaragoza, Pedro Larrañaga |
Pattern Recognit. Lett. | 2 |
| 2014 | Multiobjective Estimation of Distribution Algorithm Based on Joint Modeling of Objectives and VariablesabstractThis paper proposes a new multiobjective estimation of distribution algorithm (EDA) based on joint probabilistic modeling of objectives and variables. This EDA uses the multidimensional Bayesian network as its probabilistic model. In this way, it can capture the dependencies between objectives, variables and objectives, as well as the dependencies learned between variables in other Bayesian network-based EDAs. This model leads to a problem decomposition that helps the proposed algorithm find better tradeoff solutions to the multiobjective problem. In addition to Pareto set approximation, the algorithm is also able to estimate the structure of the multiobjective problem. To apply the algorithm to many-objective problems, the algorithm includes four different ranking methods proposed in the literature for this purpose. The algorithm is first applied to the set of walking fish group problems, and its optimization performance is compared with a standard multiobjective evolutionary algorithm and another competitive multiobjective EDA. The experimental results show that on several of these problems, and for different objective space dimensions, the proposed algorithm performs significantly better and on some others achieves comparable results when compared with the other two algorithms. The algorithm is then tested on the set of CEC09 problems, where the results show that multiobjective optimization based on joint model estimation is able to obtain considerably better fronts for some of the problems compared with the search based on conventional genetic operators in the state-of-the-art multiobjective evolutionary algorithms. Hossein Karshenas, Roberto Santana 0001, Concha Bielza, Pedro Larrañaga |
IEEE Trans. Evol. Comput. | 3 |
| 2014 | Multi-Dimensional Classification with Super-ClassesabstractThe multi-dimensional classification problem is a generalization of the recently-popularized task of multi-label classification, where each data instance is associated with multiple class variables. There has been relatively little research carried out specific to multi-dimensional classification and, although one of the core goals is similar (modeling dependencies among classes), there are important differences; namely a higher number of possible classifications. In this paper we present method for multi-dimensional classification, drawing from the most relevant multi-label research, and combining it with important novel developments. Using a fast method to model the conditional dependence between class variables, we form super-class partitions and use them to build multi-dimensional learners, learning each super-class as an ordinary class, and thus explicitly modeling class dependencies. Additionally, we present a mechanism to deal with the many class values inherent to super-classes, and thus make learning efficient. To investigate the effectiveness of this approach we carry out an empirical evaluation on a range of multi-dimensional datasets, under different evaluation metrics, and in comparison with high-performing existing multi-dimensional approaches from the literature. Analysis of results shows that our approach offers important performance gains over competing methods, while also exhibiting tractable running time. Jesse Read, Concha Bielza, Pedro Larrañaga |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2013 | Semi-supervised Projected Clustering for Classifying GABAergic Interneurons
Luis Guerra, Ruth Benavides-Piccione, Concha Bielza, Víctor Robles, Javier DeFelipe, Pedro Larrañaga |
AIME | 3 |
| 2013 | Bayesian networks to answer challenging neuroscience questionsabstractSummary form only given. In this keynote lecture we will show how Bayesian networks can address important neuroscience problems. These problems include: (a) neuroanatomy issues, like modeling and simulation of dendritic trees and classifying neuron types based on morphological features; (b) neurodegenerative diseases, like predicting health-related quality of life in Parkinson's disease, classification of dementia stages in Parkinson's disease and searching for genetic biomarkers in Alzheimer's disease. Pedro Larrañaga, Concha Bielza |
CBMS | 2 |
| 2013 | Unveiling relevant non-motor Parkinson's disease severity symptoms using a machine learning approach
Rubén Armañanzas, Concha Bielza, Kallol Ray Chaudhuri, Pablo Martínez-Martín, Pedro Larrañaga |
Artif. Intell. Medicine | 2 |
| 2013 | Predicting human immunodeficiency virus inhibitors using multi-dimensional Bayesian network classifiers
Hanen Borchani, Concha Bielza, Carlos Toro 0003, Pedro Larrañaga |
Artif. Intell. Medicine | 2 |
| 2013 | Classification of neural signals from sparse autoregressive features
Diego Vidaurre, Concha Bielza, Pedro Larrañaga |
Neurocomputing | 2 |
| 2013 | Comparison of metaheuristic strategies for peakbin selection in proteomic mass spectrometry data
Miguel García-Torres, Rubén Armañanzas, Concha Bielza, Pedro Larrañaga |
Inf. Sci. | 3 |
| 2013 | A review on evolutionary algorithms in Bayesian network learning and inference tasks
Pedro Larrañaga, Hossein Karshenas, Concha Bielza, Roberto Santana 0001 |
Inf. Sci. | 3 |
| 2013 | Parameter Control of Genetic Algorithms by Learning and Simulation of Bayesian Networks - A Case Study for the Optimal Ordering of Tables
Concha Bielza, Juan A. Fernández del Pozo, Pedro Larrañaga |
J. Comput. Sci. Technol. | 1 |
| 2013 | Bayesian Sparse Partial Least SquaresabstractPartial least squares (PLS) is a class of methods that makes use of a set of latent or unobserved variables to model the relation between (typically) two sets of input and output variables, respectively. Several flavors, depending on how the latent variables or components are computed, have been developed over the last years. In this letter, we propose a Bayesian formulation of PLS along with some extensions. In a nutshell, we provide sparsity at the input space level and an automatic estimation of the optimal number of latent components. We follow the variational approach to infer the parameter distributions. We have successfully tested the proposed methods on a synthetic data benchmark and on electrocorticogram data associated with several motor outputs in monkeys. Diego Vidaurre, Marcel van Gerven, Concha Bielza, Pedro Larrañaga, Tom Heskes |
Neural Comput. | 3 |
| 2012 | A comparison of clustering quality indices using outliers and noiseabstractQuality indices in clustering are used not only to assess the quality of the partitions but also to determine the number of clusters in the final result. When these indices are evaluated in a case study, real data conditions or different clustering a Luis Guerra, Víctor Robles, Concha Bielza, Pedro Larrañaga |
Intell. Data Anal. | 3 |
| 2012 | Markov blanket-based approach for learning multi-dimensional Bayesian network classifiers: An application to predict the European Quality of Life-5 Dimensions (EQ-5D) from the 39-item Parkinson's Disease Questionnaire (PDQ-39)
Hanen Borchani, Concha Bielza, Pablo Martínez-Martín, Pedro Larrañaga |
J. Biomed. Informatics | 2 |
| 2011 | Multi-objective Optimization with Joint Probabilistic Modeling of Objectives and Variables
Hossein Karshenas, Roberto Santana 0001, Concha Bielza, Pedro Larrañaga |
EMO | 3 |
| 2011 | Affinity propagation enhanced by estimation of distribution algorithmsabstractTumor classification based on gene expression data can be applied to set appropriate medical treatment according to the specific tumor characteristics. In this paper we propose the use of estimation of distribution algorithms (EDAs) to enhance the performance of affinity propagation (AP) in classification problems. AP is an efficient clustering algorithm based on message-passing methods and which automatically identifies exemplars of each cluster. We introduce an EDA-based procedure to compute the preferences used by the AP algorithm. Our results show that AP performance can be notably improved by using the introduced approach. Furthermore, we present evidence that classification of new data is improved by employing previously identified exemplars with only minor decrease in classification accuracy. Roberto Santana 0001, Concha Bielza, Pedro Larrañaga |
GECCO | 2 |
| 2011 | Regularized k-order markov models in EDAsabstractk-order Markov models have been introduced to estimation of distribution algorithms (EDAs) to solve a particular class of optimization problems in which each variable depends on its previous k variables in a given, fixed order. In this paper we investigate the use of regularization as a way to approximate k-order Markov models when $k$ is increased. The introduced regularized models are used to balance the complexity and accuracy of the k-order Markov models. We investigate the behavior of the EDAs in several instances of the hydrophobic-polar (HP) protein problem, a simplified protein folding model. Our preliminary results show that EDAs that use regularized approximations of the k-order Markov models offer a good compromise between complexity and efficiency, and could be an appropriate choice when the number of variables is increased. Roberto Santana 0001, Hossein Karshenas, Concha Bielza, Pedro Larrañaga |
GECCO | 3 |
| 2011 | Bayesian Chain Classifiers for Multidimensional ClassificationabstractIn multidimensional classification the goal is to assign an instance to a set of different classes. This task is normally addressed either by defining a compound class variable with all the possible combinations of classes (label power-set methods, LPMs) or by building independent classifiers for each class (binary-relevance methods, BRMs). However, LPMs do not scale well and BRMs ignore the dependency relations between classes. We introduce a method for chaining binary Bayesian classifiers that combines the strengths of classifier chains and Bayesian networks for multidimensional classification. The method consists of two phases. In the first phase, a Bayesian network (BN)that represents the dependency relations between the class variables is learned from data. In the second phase, several chain classifiers are built, such that the order of the class variables in the chain is consistent with the class BN. At the end we combine the results of the different generated orders. Our method considers the dependencies between class variables and takes advantage of the conditional independence relations to build simplified models. We perform experiments with a chain of naïve Bayes classifiers on different benchmark multidimensional datasets and show that our approach out performs other state-of-the-art methods. Julio H. Zaragoza, Luis Enrique Sucar, Eduardo F. Morales 0001, Concha Bielza, Pedro Larrañaga |
IJCAI | 4 |
| 2011 | Predicting the h-index with cost-sensitive naive BayesabstractBibliometric indices are an increasingly important topic for the scientific community nowadays. One of the most successful bibliometric indices is the well-known h-index. In view of the attention attracted by this index, our research is based on the construction of several prediction models to forecast the h-index of Spanish professors (with a permanent position) for a four-year time horizon. We built two different types of models (junior models and senior models) to differentiate between professors' seniority. These models are learnt from bibliometric data using a cost-sensitive naive Bayes approach that takes into account the expected cost of instances predictions at classification time. Results show that it is easier to predict the h-index of the one-year time horizon than the others, that is, it has a higher average accuracy and lower average total cost than the others. Similarly, it is easier to predict the h-index of junior professors than senior professors. Alfonso Ibáñez, Pedro Larrañaga, Concha Bielza |
ISDA | 3 |
| 2011 | Dealing with complex queries in decision-support systems
Juan A. Fernández del Pozo, Concha Bielza |
Data Knowl. Eng. | 2 |
| 2011 | Regularized logistic regression without a penalty term: An application to cancer classification with microarray data
Concha Bielza, Víctor Robles, Pedro Larrañaga |
Expert Syst. Appl. | 1 |
| 2011 | Classifying evolving data streams with partially labeled dataabstractRecently, several approaches have been proposed to deal with the increasingly challenging task of mining concept-drifting data streams. However, most are based on supervised classification algorithms assuming that true labels are immediately and entirely available in the data streams. Unfortunately , such an assumption is often violated in real-world applications given that it is expensive or because it takes a long time to obtain all true labels. To deal with this problem, we propose in this paper a new semi-supervised approach for handling concept-drifting data streams containing both labeled and unlabeled instances. First, contrary to existing approaches, we monitor three possible kinds of drift: feature, conditional or dual drift. Drift detection is based on a hypothesis test comparing Kullback-Leibler divergence between old and recent data, whose distribution under the null hypothesis of coming from the same distribution is approximated via a bootstrap method. Then, if any drift occurs, a new classifier is learned from the recent data using the EM algorithm; otherwise, the current classifier is left unchanged. Our approach is so general that it can be applied to different classification models. Experimental studies, using the naive Bayes classifier and logistic regression, on both synthetic and real-world data sets demonstrate that our approach performs well. Hanen Borchani, Pedro Larrañaga, Concha Bielza |
Intell. Data Anal. | 3 |
| 2011 | Multi-dimensional classification with Bayesian networks
Concha Bielza, Guangdi Li, Pedro Larrañaga |
Int. J. Approx. Reason. | 1 |
| 2011 | Peakbin Selection in Mass Spectrometry Data Using a Consensus Approach with Estimation of Distribution AlgorithmsabstractProgress is continuously being made in the quest for stable biomarkers linked to complex diseases. Mass spectrometers are one of the devices for tackling this problem. The data profiles they produce are noisy and unstable. In these profiles, biomarkers are detected as signal regions (peaks), where control and disease samples behave differently. Mass spectrometry (MS) data generally contain a limited number of samples described by a high number of features. In this work, we present a novel class of evolutionary algorithms, estimation of distribution algorithms (EDA), as an efficient peak selector in this MS domain. There is a trade-of f between the reliability of the detected biomarkers and the low number of samples for analysis. For this reason, we introduce a consensus approach, built upon the classical EDA scheme, that improves stability and robustness of the final set of relevant peaks. An entire data workflow is designed to yield unbiased results. Four publicly available MS data sets (two MALDI-TOF and another two SELDI-TOF) are analyzed. The results are compared to the original works, and a new plot (peak frequential plot) for graphically inspecting the relevant peaks is introduced. A complete online supplementary page, which can be found at http://www.sc.ehu.es/ccwbayes/members/ruben/ms, includes extended info and results, in addition to Matlab scripts and references. Rubén Armañanzas, Yvan Saeys, Iñaki Inza, Miguel García-Torres, Concha Bielza, Yves Van de Peer, Pedro Larrañaga |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2010 | Bivariate empirical and n-variate Archimedean copulas in estimation of distribution algorithmsabstractThis paper investigates the use of empirical and Archimedean copulas as probabilistic models of continuous estimation of distribution algorithms (EDAs). A method for learning and sampling empirical bivariate copulas to be used in the context of n-dimensional EDAs is first introduced. Then, by using Archimedean copulas instead of empirical makes possible to construct n-dimensional copulas with the same purpose. Both copula-based EDAs are compared to other known continuous EDAs on a set of 24 functions and different number of variables. Experimental results show that the proposed copula-based EDAs achieve a better behaviour than previous approaches in a 20% of the benchmark functions. Alfredo Cuesta-Infante, Roberto Santana 0001, J. Ignacio Hidalgo, Concha Bielza, Pedro Larrañaga |
IEEE Congress on Evolutionary Computation | 4 |
| 2010 | Mining Concept-Drifting Data Streams Containing Labeled and Unlabeled Instances
Hanen Borchani, Pedro Larrañaga, Concha Bielza |
IEA/AIE (1) | 3 |
| 2010 | Synergies between Network-Based Representation and Probabilistic Graphical Models for Classification, Inference and Optimization Problems in Neuroscience
Roberto Santana 0001, Concha Bielza, Pedro Larrañaga |
IEA/AIE (3) | 2 |
| 2010 | Modeling challenges with influence diagrams: Constructing probability and utility models
Concha Bielza, Manuel Gómez, Prakash P. Shenoy |
Decis. Support Syst. | 1 |
| 2010 | Multidimensional statistical analysis of the parameterization of a genetic algorithm for the optimal ordering of tables
Concha Bielza, Juan A. Fernández del Pozo, Pedro Larrañaga, Endika Bengoetxea |
Expert Syst. Appl. | 1 |
| 2010 | Learning an L1-Regularized Gaussian Bayesian Network in the Equivalence Class SpaceabstractLearning the structure of a graphical model from data is a common task in a wide range of practical applications. In this paper, we focus on Gaussian Bayesian networks, i.e., on continuous data and directed acyclic graphs with a joint probability density of all variables given by a Gaussian. We propose to work in an equivalence class search space, specifically using the k-greedy equivalence search algorithm. This, combined with regularization techniques to guide the structure search, can learn sparse networks close to the one that generated the data. We provide results on some synthetic networks and on modeling the gene network of the two biological pathways regulating the biosynthesis of isoprenoids for the Arabidopsis thaliana plant. Diego Vidaurre, Concha Bielza, Pedro Larrañaga |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2009 | Mining probabilistic models learned by EDAs in the optimization of multi-objective problemsabstractOne of the uses of the probabilistic models learned by estimation of distribution algorithms is to reveal previous unknown information about the problem structure. In this paper we investigate the mapping between the problem structure and the dependencies captured in the probabilistic models learned by EDAs for a set of multi-objective satisfiability problems. We present and discuss the application of different data mining and visualization techniques for processing and visualizing relevant information from the structure of the learned probabilistic models. We show that also in the case of multi-objective optimization problems, some features of the original problem structure can be translated to the probabilistic models and unveiled by using algorithms that mine the model structures. Roberto Santana 0001, Concha Bielza, José Antonio Lozano 0001, Pedro Larrañaga |
GECCO | 2 |
| 2009 | Predicting citation count of Bioinformatics papers within four years of publicationabstractMOTIVATION: Nowadays, publishers of scientific journals face the tough task of selecting high-quality articles that will attract as many readers as possible from a pool of articles. This is due to the growth of scientific output and literature. The possibility of a journal having a tool capable of predicting the citation count of an article within the first few years after publication would pave the way for new assessment systems. RESULTS: This article presents a new approach based on building several prediction models for the Bioinformatics journal. These models predict the citation count of an article within 4 years after publication (global models). To build these models, tokens found in the abstracts of Bioinformatics papers have been used as predictive features, along with other features like the journal sections and 2-week post-publication periods. To improve the accuracy of the global models, specific models have been built for each Bioinformatics journal section (Data and Text Mining, Databases and Ontologies, Gene Expression, Genetics and Population Analysis, Genome Analysis, Phylogenetics, Sequence Analysis, Structural Bioinformatics and Systems Biology). In these new models, the average success rate for predictions using the naive Bayes and logistic regression supervised classification methods was 89.4% and 91.5%, respectively, within the nine sections and for 4-year time horizon. AVAILABILITY: Supplementary material on this experimental survey is available at http://www.dia.fi.upm.es/~concha/bioinformatics.html CONTACT: [email protected] Alfonso Ibáñez, Pedro Larrañaga, Concha Bielza |
Bioinform. | 3 |
| 2009 | Comparison of Bayesian networks and artificial neural networks for quality detection in a machining process
Maritza Correa, Concha Bielza, J. Pamies-Teixeira |
Expert Syst. Appl. | 2 |
| 2008 | Explaining clinical decisions by extracting regularity patterns
Concha Bielza, Juan A. Fernández del Pozo, Peter J. F. Lucas |
Decis. Support Syst. | 1 |
| 2006 | Machine learning in bioinformaticsabstractThis article reviews machine learning methods for bioinformatics. It presents modelling methods, such as supervised classification, clustering and probabilistic graphical models for knowledge discovery, as well as deterministic and stochastic heuristics for optimization. Applications in genomics, proteomics, systems biology, evolution and text mining are also shown. Pedro Larrañaga, Borja Calvo, Roberto Santana 0001, Concha Bielza, Josu Galdiano, Iñaki Inza, José Antonio Lozano 0001, Rubén Armañanzas, Guzmán Santafé, Aritz Pérez Martínez, Víctor Robles |
Briefings Bioinform. | 4 |
| 2003 | Finding and Explaining Optimal Treatments
Concha Bielza, Juan A. Fernández del Pozo, Peter J. F. Lucas |
AIME | 1 |