VLDB 2026 Research / reviewers in the wild / expert
Henrik Boström
dblp:39/6674
· DBLP profile ↗
92ranked-venue papers
14as first author
17since 2021 · last 2025
0000-0001-8382-0300ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 64 · 8 first-author · 13 since 2021Databases, data management, data science and information retrieval · 31 · 7 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Theory of computation · 4Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Explaining Deep Neural Networks with Example and Pixel Attribution
Genghua Dong, Henrik Boström, Roman Bresson, Amr Alkhatib |
DS | 2 |
| 2025 | Explaining Representations in Correlation-based Deep Multiview Representation LearningabstractMultiview representation learning techniques based on deep correlation maximization have become increasingly popular for learning meaningful and compact representations from multiview data. Even though their performance is state-of-the-art in many interpretability-critical fields, their black-box behavior poses a problem and restricts their usability. To overcome this restriction, we propose XDCCA (eXplanations for Deep Canonical Correlation Analysis), an explanation strategy using characteristic rules in combination with SHAP that exploits the inherent structure of latent spaces created by correlation maximization techniques. We demonstrate how XDCCA allows for interpreting learned representations and their correlation using real medical time series and synthetic image data. Maurice Kuschel, Amr Alkhatib, Tanuj Hasija, Henrik Boström |
ICASSP | 4 |
| 2025 | Prediction via Shapley Value RegressionabstractShapley values have several desirable, theoretically well-supported, properties for explaining black-box model predictions. Traditionally, Shapley values are computed post-hoc, leading to additional computational cost at inference time. To overcome this, a novel method, called ViaSHAP, is proposed, that learns a function to compute Shapley values, from which the predictions can be derived directly by summation. Two approaches to implement the proposed method are explored; one based on the universal approximation theorem and the other on the Kolmogorov-Arnold representation theorem. Results from a large-scale empirical investigation are presented, showing that ViaSHAP using Kolmogorov-Arnold Networks performs on par with state-of-the-art algorithms for tabular data. It is also shown that the explanations of ViaSHAP are significantly more accurate than the popular approximator FastSHAP on both tabular data and images. Amr Alkhatib, Roman Bresson, Henrik Boström, Michalis Vazirgiannis |
ICML | 3 |
| 2025 | Obtaining Example-Based Explanations from Deep Neural Networks
Genghua Dong, Henrik Boström, Michalis Vazirgiannis, Roman Bresson |
IDA | 2 |
| 2025 | CADQ: Attribute-Consistent Face Cartoonization with Cross-modal Aligned and Deformable QuantizationabstractFace cartoonization remains a challenging task due to significant geometric deformations between facial photos and cartoons, as well as the absence of paired training data for supervised learning. Existing methods struggle to generate high-quality cartoonized avatars with attribute consistency. To address this challenge, this paper proposes an unsupervised facial cartoonization method based on cross-domain aligned and deformable vector quantization (CADQ). Firstly, we construct textual descriptions with facial attributes for both photo datasets and cartoon collections. Attribute consistency during transformation is enforced through individually contrastive learning between image-text cross-modal features and globally distribution alignment across photo-cartoon domains. Secondly, a deformable Transformer with dual attention is introduced during the transformation process, which queries corresponding cartoon codebook entries based on image features to simulate cross-domain geometric deformations. Experimental results demonstrate that the proposed method can convert facial photos into high-quality cartoons with attribute consistency, outperforming existing state-of-the-art approaches. Furthermore, the method can be effectively extended to unsupervised cross-domain generation of other artistic portrait styles, achieving superior or highly competitive performance. Our code has been released at: https://github.com/IIP-Lab-XDU/CADQ. Yongjie Hu, Ziyun Li 0002, Fei Gao 0006, Henrik Boström, Nannan Wang 0001 |
ACM Multimedia | 5 |
| 2025 | ConANN: Conformal Approximate Nearest Neighbor Search
Sonia-Florina Horchidan, Fabian Zeiher, Henrik Boström, Paris Carbone |
Proc. VLDB Endow. | 3 |
| 2024 | A Simple and Yet Fairly Effective Defense for Graph Neural NetworksabstractGraph Neural Networks (GNNs) have emerged as the dominant approach for machine learning on graph-structured data. However, concerns have arisen regarding the vulnerability of GNNs to small adversarial perturbations. Existing defense methods against such perturbations suffer from high time complexity and can negatively impact the model's performance on clean graphs. To address these challenges, this paper introduces NoisyGNNs, a novel defense method that incorporates noise into the underlying model's architecture. We establish a theoretical connection between noise injection and the enhancement of GNN robustness, highlighting the effectiveness of our approach. We further conduct extensive empirical evaluations on the node classification task to validate our theoretical findings, focusing on two popular GNNs: the GCN and GIN. The results demonstrate that NoisyGNN achieves superior or comparable defense performance to existing methods while minimizing added time complexity. The NoisyGNN approach is model-agnostic, allowing it to be integrated with different GNN architectures. Successful combinations of our NoisyGNN approach with existing defense techniques demonstrate even further improved adversarial defense results. Our code is publicly available at: https://github.com/Sennadir/NoisyGNN. Sofiane Ennadir, Yassine Abbahaddou, Johannes F. Lutzeyer, Michalis Vazirgiannis, Henrik Boström |
AAAI | 5 |
| 2024 | Interpretable Graph Neural Networks for Heterogeneous Tabular Data
Amr Alkhatib, Henrik Boström |
DS (1) | 2 |
| 2024 | Faithfulness of Local Explanations for Tree-Based Ensemble Models
Amir Hossein Akhavan Rahnama, Pierre Geurts, Henrik Boström |
DS (2) | 3 |
| 2024 | Interpretable Graph Neural Networks for Tabular DataabstractData in tabular format is frequently occurring in real-world applications. Graph Neural Networks (GNNs) have recently been extended to effectively handle such data, allowing feature interactions to be captured through representation learning. However, these approaches essentially produce black-box models, in the form of deep neural networks, precluding users from following the logic behind the model predictions. We propose an approach, called IGNNet (Interpretable Graph Neural Network for tabular data), which constrains the learning algorithm to produce an interpretable model, where the model shows how the predictions are exactly computed from the original input features. A large-scale empirical investigation is presented, showing that IGNNet is performing on par with state-of-the-art machine-learning algorithms that target tabular data, including XGBoost, Random Forests, and TabNet. At the same time, the results show that the explanations obtained from IGNNet are aligned with the true Shapley values of the features without incurring any additional computational overhead. Amr Alkhatib, Sofiane Ennadir, Henrik Boström, Michalis Vazirgiannis |
ECAI | 3 |
| 2024 | Bounding the Expected Robustness of Graph Neural Networks Subject to Node Feature AttacksabstractGraph Neural Networks (GNNs) have demonstrated state-of-the-art performance in various graph representation learning tasks. Recently, studies revealed their vulnerability to adversarial attacks. In this work, we theoretically define the concept of expected robustness in the context of attributed graphs and relate it to the classical definition of adversarial robustness in the graph representation learning literature. Our definition allows us to derive an upper bound of the expected robustness of Graph Convolutional Networks (GCNs) and Graph Isomorphism Networks subject to node feature attacks. Building on these findings, we connect the expected robustness of GNNs to the orthonormality of their weight matrices and consequently propose an attack-independent, more robust variant of the GCN, called the Graph Convolutional Orthonormal Robust Networks (GCORNs). We further introduce a probabilistic method to estimate the expected robustness, which allows us to evaluate the effectiveness of GCORN on several real-world datasets. Experimental experiments showed that GCORN outperforms available defense methods. Our code is publicly available at: https://github.com/Sennadir/GCORN . Yassine Abbahaddou, Sofiane Ennadir, Johannes F. Lutzeyer, Michalis Vazirgiannis, Henrik Boström |
ICLR | 5 |
| 2024 | Example-Based Explanations of Random Forest Predictions
Henrik Boström |
IDA (2) | 1 |
| 2024 | Can local explanation techniques explain linear additive models?abstractAbstract Local model-agnostic additive explanation techniques decompose the predicted output of a black-box model into additive feature importance scores. Questions have been raised about the accuracy of the produced local additive explanations. We investigate this by studying whether some of the most popular explanation techniques can accurately explain the decisions of linear additive models. We show that even though the explanations generated by these techniques are linear additives, they can fail to provide accurate explanations when explaining linear additive models. In the experiments, we measure the accuracy of additive explanations, as produced by, e.g., LIME and SHAP, along with the non-additive explanations of Local Permutation Importance (LPI) when explaining Linear and Logistic Regression and Gaussian naive Bayes models over 40 tabular datasets. We also investigate the degree to which different factors, such as the number of numerical or categorical or correlated features, the predictive performance of the black-box model, explanation sample size, similarity metric, and the pre-processing technique used on the dataset can directly affect the accuracy of local explanations. Amir Hossein Akhavan Rahnama, Judith Bütepage, Pierre Geurts, Henrik Boström |
Data Min. Knowl. Discov. | 4 |
| 2022 | Explaining Predictions by Characteristic Rules
Amr Alkhatib, Henrik Boström, Michalis Vazirgiannis |
ECML/PKDD (1) | 2 |
| 2022 | Random subspace and random projection nearest neighbor ensembles for high dimensional dataabstractThe random subspace and the random projection methods are investigated and compared as techniques for forming ensembles of nearest neighbor classifiers in high dimensional feature spaces. The two methods have been empirically evaluated on three types of high-dimensional datasets: microarrays, chemoinformatics, and images. Experimental results on 34 datasets show that both the random subspace and the random projection method lead to improvements in predictive performance compared to using the standard nearest neighbor classifier, while the best method to use depends on the type of data considered; for the microarray and chemoinformatics datasets, random projection outperforms the random subspace method, while the opposite holds for the image datasets. An analysis using data complexity measures, such as attribute to instance ratio and Fisher’s discriminant ratio, provide some more detailed indications on what relative performance can be expected for specific datasets. The results also indicate that the resulting ensembles may be competitive with state-of-the-art ensemble classifiers; the nearest neighbor ensembles using random projection perform on par with random forests for the microarray and chemoinformatics datasets. Sampath Deegalla, Keerthi Walgama, Panagiotis Papapetrou, Henrik Boström |
Expert Syst. Appl. | 4 |
| 2022 | Rule extraction with guarantees from regression modelsabstractTools for understanding and explaining complex predictive models are critical for user acceptance and trust. One such tool is rule extraction, i.e., approximating opaque models with less powerful but interpretable models. Pedagogical (or black-box) rule extraction, where the interpretable model is induced using the original training instances, but with the predictions from the opaque model as targets, has many advantages compared to the decompositional (white-box) approach. Most importantly, pedagogical methods are agnostic to the kind of opaque model used, and any learning algorithm producing interpretable models can be employed for the learning step. The pedagogical approach has, however, one main problem, clearly limiting its utility. Specifically, while the extracted models are trained to mimic the opaque, there are absolutely no guarantees that this will transfer to novel data. This potentially low test set fidelity must be considered a severe drawback, in particular when the extracted models are used for explanation and analysis. In this paper, a novel approach, solving the problem with test set fidelity by utilizing the conformal prediction framework, is suggested for extracting interpretable regression models from opaque models. The extracted models are standard regression trees, but augmented with valid prediction intervals in the leaves. Depending on the exact setup, the use of conformal prediction guarantees that either the test set fidelity or the test set accuracy will be equal to a preset confidence level, in the long run. In the extensive empirical investigation, using 20 publicly available data sets, the validity of the extracted models is demonstrated. In addition, it is shown how normalization can be used to provide individualized prediction intervals, thus providing highly informative extracted models. Ulf Johansson, Cecilia Sönströd, Tuve Löfström, Henrik Boström |
Pattern Recognit. | 4 |
| 2021 | Well-Calibrated and Sharp Interpretable Multi-Class Models
Ulf Johansson, Tuve Löfström, Henrik Boström |
MDAI | 3 |
| 2020 | Orthogonal Mixture of Hidden Markov Models
Negar Safinianaini, Camila P. E. de Souza, Henrik Boström, Jens Lagergren |
ECML/PKDD (1) | 3 |
| 2020 | Study of Hellinger Distance as a splitting metric for Random Forests in balanced and imbalanced classification datasets
Ricardo Aler, José María Valls, Henrik Boström |
Expert Syst. Appl. | 3 |
| 2020 | Efficient conformal predictor ensembles
Henrik Linusson, Ulf Johansson, Henrik Boström |
Neurocomputing | 3 |
| 2020 | Corrigendum to 'Learning from heterogeneous temporal data in electronic health records'. [J. Biomed. Inform. 65 (2017) 105-119]
Jing Zhao 0017, Panagiotis Papapetrou, Lars Asker, Henrik Boström |
J. Biomed. Informatics | 4 |
| 2019 | Gated Hidden Markov Models for Early Prediction of Outcome of Internet-Based Cognitive Behavioral Therapy
Negar Safinianaini, Henrik Boström, Viktor Kaldo |
AIME | 2 |
| 2019 | Customized Interpretable Conformal RegressorsabstractInterpretability is recognized as a key property of trustworthy predictive models. Only interpretable models make it straightforward to explain individual predictions, and allow inspection and analysis of the model itself. In real-world scenarios, these explanations and insights are often needed for a specific batch of predictions, i.e., a production set. If the input vectors for this production set are available when generating the predictive model, a methodology called oracle coaching can be used to produce highly accurate and interpretable models optimized for the specific production set. In this paper, oracle coaching is, for the first time, combined with the conformal prediction framework for predictive regression. A conformal regressor, which is built on top of a standard regression model, outputs valid prediction intervals, i.e., the error rate on novel data is bounded by a preset significance level, as long as the labeled data used for calibration is exchangeable with production data. Since validity is guaranteed for all conformal predictors, the key performance metric is the size of the prediction intervals, where tighter (more efficient) intervals are preferred. The efficiency of a conformal model depends on several factors, but more accurate underlying models will generally also lead to improved efficiency in the corresponding conformal predictor. A key contribution in this paper is the design of setups ensuring that when oracle coached regression trees, that per definition utilize knowledge about production data, are used as underlying models for conformal regressors, these remain valid. The experiments, using 20 publicly available regression data sets, demonstrate the validity of the suggested setups. Results also show that utilizing oracle-coached underlying models will generally lead to significantly more efficient conformal regressors, compared to when these are built on top of models induced using only training data. Ulf Johansson, Cecilia Sönströd, Tuve Löfström, Henrik Boström |
DSAA | 4 |
| 2019 | Calibrating Probability Estimation Trees using Venn-Abers PredictorsabstractClass labels output by standard decision trees are not very useful for making informed decisions, e.g., when comparing the expected utility of various alternatives. In contrast, probability estimation trees (PETs) output class probability distributions rather than single class labels. It is well known that estimating class probabilities in PETs by relative frequencies often lead to extreme probability estimates, and a number of approaches to provide more well-calibrated estimates have been proposed. In this study, a recent model-agnostic calibration approach, called Venn-Abers predictors is, for the first time, considered in the context of decision trees. Results from a large-scale empirical investigation are presented, comparing the novel approach to previous calibration techniques with respect to several different performance metrics, targeting both predictive performance and reliability of the estimates. All approaches are considered both with and without Laplace correction. The results show that using Venn-Abers predictors for calibration is a highly competitive approach, significantly outperforming Platt scaling, Isotonic regression and no calibration, with respect to almost all performance metrics used, independently of whether Laplace correction is applied or not. The only exception is AUC, where using non-calibrated PETs together with Laplace correction, actually is the best option, which can be explained by the fact that AUC is not affected by the absolute, but only relative, values of the probability estimates. Ulf Johansson, Tuve Löfström, Henrik Boström |
SDM | 3 |
| 2019 | Block-distributed Gradient Boosted TreesabstractThe Gradient Boosted Tree (GBT) algorithm is one of the most popular machine learning algorithms used in production, for tasks that include Click-Through Rate (CTR) prediction and learning-to-rank. To deal with the massive datasets available today, many distributed GBT methods have been proposed. However, they all assume a row-distributed dataset, addressing scalability only with respect to the number of data points and not the number of features, and increasing communication cost for high-dimensional data. In order to allow for scalability across both the data point and feature dimensions, and reduce communication cost, we propose block-distributed GBTs. We achieve communication efficiency by making full use of the data sparsity and adapting the Quickscorer algorithm to the block-distributed setting. We evaluate our approach using datasets with millions of features, and demonstrate that we are able to achieve multiple orders of magnitude reduction in communication cost for sparse data, with no loss in accuracy, while providing a more scalable design. As a result, we are able to reduce the training time for high-dimensional data, and allow more cost-effective scale-out without the need for expensive network communication. Theodore Vasiloudis, Hyunsu Cho, Henrik Boström |
SIGIR | 3 |
| 2019 | Quantifying Uncertainty in Online Regression ForestsabstractAccurately quantifying uncertainty in predictions is essential for the deployment of machine learning algorithms in critical applications where mistakes are costly. Most approaches to quantifying prediction uncertainty have focused on settings where the data is static, or bounded. In this paper, we investigate methods that quantify the prediction uncertainty in a streaming setting, where the data is potentially unbounded. We propose two meta-algorithms that produce prediction intervals for online regression forests of arbitrary tree models; one based on conformal prediction, and the other based on quantile regression. We show that the approaches are able to maintain specified error rates, with constant computational cost per example and bounded memory usage. We provide empirical evidence that the methods outperform the state-of-the-art in terms of maintaining error guarantees, while being an order of magnitude faster. We also investigate how the algorithms are able to recover from concept drift. Theodore Vasiloudis, Gianmarco De Francisci Morales, Henrik Boström |
J. Mach. Learn. Res. | 3 |
| 2019 | Conformal and probabilistic prediction with applications: editorial
Alex Gammerman, Vladimir Vovk, Henrik Boström, Lars Carlsson |
Mach. Learn. | 3 |
| 2019 | Efficient Venn predictors using random forestsabstractSuccessful use of probabilistic classification requires well-calibrated probability estimates, i.e., the predicted class probabilities must correspond to the true probabilities. In addition, a probabilistic classifier must, of course, also be as accurate as possible. In this paper, Venn predictors, and its special case Venn-Abers predictors, are evaluated for probabilistic classification, using random forests as the underlying models. Venn predictors output multiple probabilities for each label, i.e., the predicted label is associated with a probability interval. Since all Venn predictors are valid in the long run, the size of the probability intervals is very important, with tighter intervals being more informative. The standard solution when calibrating a classifier is to employ an additional step, transforming the outputs from a classifier into probability estimates, using a labeled data set not employed for training of the models. For random forests, and other bagged ensembles, it is, however, possible to use the out-of-bag instances for calibration, making all training data available for both model learning and calibration. This procedure has previously been successfully applied to conformal prediction, but was here evaluated for the first time for Venn predictors. The empirical investigation, using 22 publicly available data sets, showed that all four versions of the Venn predictors were better calibrated than both the raw estimates from the random forest, and the standard techniques Platt scaling and isotonic regression. Regarding both informativeness and accuracy, the standard Venn predictor calibrated on out-of-bag instances was the best setup evaluated. Most importantly, calibrating on out-of-bag instances, instead of using a separate calibration set, resulted in tighter intervals and more accurate models on every data set, for both the Venn predictors and the Venn-Abers predictors. Ulf Johansson, Tuve Löfström, Henrik Linusson, Henrik Boström |
Mach. Learn. | 4 |
| 2018 | Classification with Reject Option Using Conformal Prediction
Henrik Linusson, Ulf Johansson, Henrik Boström, Tuve Löfström |
PAKDD (1) | 3 |
| 2018 | Interpretable regression trees using conformal prediction
Ulf Johansson, Henrik Linusson, Tuve Löfström, Henrik Boström |
Expert Syst. Appl. | 4 |
| 2017 | Conformal Prediction Using Random Survival ForestsabstractRandom survival forests constitute a robust approach to survival modeling, i.e., predicting the probability that an event will occur before or on a given point in time. Similar to most standard predictive models, no guarantee for the prediction error is provided for this model, which instead typically is empirically evaluated. Conformal prediction is a rather recent framework, which allows the error of a model to be determined by a user specified confidence level, something which is achieved by considering set rather than point predictions. The framework, which has been applied to some of the most popular classification and regression techniques, is here for the first time applied to survival modeling, through random survival forests. An empirical investigation is presented where the technique is evaluated on datasets from two real-world applications; predicting component failure in trucks using operational data and predicting survival and treatment of heart failure patients from administrative healthcare data. The experimental results show that the error levels indeed are very close to the provided confidence levels, as guaranteed by the conformal prediction framework, and that the error for predicting each outcome, i.e., event or no-event, can be controlled separately. The latter may, however, lead to less informative predictions, i.e., larger prediction sets, in case the class distribution is heavily imbalanced. Henrik Boström, Lars Asker, Ram B. Gurung, Isak Karlsson, Tony Lindgren, Panagiotis Papapetrou |
ICMLA | 1 |
| 2017 | Model-agnostic nonconformity functions for conformal classificationabstractA conformal predictor outputs prediction regions, for classification label sets. The key property of all conformal predictors is that they are valid, i.e., their error rate on novel data is bounded by a preset significance level. Thus, the key performance metric for evaluating conformal predictors is the size of the output prediction regions, where smaller (more informative) prediction regions are said to be more efficient. All conformal predictions rely on nonconformity functions, measuring the strangeness of an input-output pair, and the efficiency depends critically on the quality of the chosen nonconformity function. In this paper, three model-agnostic nonconformity functions, based on well-known loss functions, are evaluated with regard to how they affect efficiency. In the experimentation on 21 publicly available multi-class data sets, both single neural networks and ensembles of neural networks are used as underlying models for conformal classifiers. The results show that the choice of nonconformity function has a major impact on the efficiency, but also that different nonconformity functions should be used depending on the exact efficiency metric. For a high fraction of single-label predictions, a margin-based nonconformity function is the best option, while a nonconformity function based on the hinge loss obtained the smallest label sets on average. Ulf Johansson, Henrik Linusson, Tuve Löfström, Henrik Boström |
IJCNN | 4 |
| 2017 | Learning from heterogeneous temporal data in electronic health recordsabstractElectronic health records contain large amounts of longitudinal data that are valuable for biomedical informatics research. The application of machine learning is a promising alternative to manual analysis of such data. However, the complex structure of the data, which includes clinical events that are unevenly distributed over time, poses a challenge for standard learning algorithms. Some approaches to modeling temporal data rely on extracting single values from time series; however, this leads to the loss of potentially valuable sequential information. How to better account for the temporality of clinical data, hence, remains an important research question. In this study, novel representations of temporal data in electronic health records are explored. These representations retain the sequential information, and are directly compatible with standard machine learning algorithms. The explored methods are based on symbolic sequence representations of time series data, which are utilized in a number of different ways. An empirical investigation, using 19 datasets comprising clinical measurements observed over time from a real database of electronic health records, shows that using a distance measure to random subsequences leads to substantial improvements in predictive performance compared to using the original sequences or clustering the sequences. Evidence is moreover provided on the quality of the symbolic sequence representation by comparing it to sequences that are generated using domain knowledge by clinical experts. The proposed method creates representations that better account for the temporality of clinical events, which is often key to prediction tasks in the biomedical domain. Jing Zhao 0017, Panagiotis Papapetrou, Lars Asker, Henrik Boström |
J. Biomed. Informatics | 4 |
| 2016 | Identifying Factors for the Effectiveness of Treatment of Heart Failure: A Registry StudyabstractAn administrative health register containing health care data for over 2 million patients will be used to search for factors that can affect the treatment of heart failure. In the study, we will measure the effects of employed treatment for various groups of heart failure patients, using different measures of effectiveness. Significant deviations in effectiveness of treatments of the various patient groups will be reported and factors that may help explaining the effect of treatment will be analyzed. Identification of the most important factors that may help explain the observed deviations between the different groups will be derived through generation of predictive models, for which variable importance can be calculated. The findings may affect recommended treatments as well as highlighting deviations from national guidelines. Lars Asker, Henrik Boström, Panagiotis Papapetrou, Hans E. Persson |
CBMS | 2 |
| 2016 | Early Random Shapelet Forest
Isak Karlsson, Panagiotis Papapetrou, Henrik Boström |
DS | 3 |
| 2016 | Reliable Confidence Predictions Using Conformal Prediction
Henrik Linusson, Ulf Johansson, Henrik Boström, Tuve Löfström |
PAKDD (1) | 3 |
| 2016 | Generalized random shapelet forests
Isak Karlsson, Panagiotis Papapetrou, Henrik Boström |
Data Min. Knowl. Discov. | 3 |
| 2015 | Handling Temporality of Clinical Events for Drug Safety Surveillance
Jing Zhao 0017, Aron Henriksson, Maria Kvist, Lars Asker, Henrik Boström |
AMIA | 5 |
| 2015 | Modeling electronic health records in ensembles of semantic spaces for adverse drug event detectionabstractAdverse drug events (ADEs) are heavily under-reported in electronic health records (EHRs). Alerting systems that are able to detect potential ADEs on the basis of patient-specific EHR data would help to mitigate this problem. To that end, the use of machine learning has proven to be both efficient and effective; however, challenges remain in representing the heterogeneous EHR data, which moreover tends to be high-dimensional and exceedingly sparse, in a manner conducive to learning high-performing predictive models. Prior work has shown that distributional semantics - that is, natural language processing methods that, traditionally, model the meaning of words in semantic (vector) space on the basis of co-occurrence information - can be exploited to create effective representations of sequential EHR data of various kinds. When modeling data in semantic space, an important design decision concerns the size of the context window around an object of interest, which governs the scope of co-occurrence information that is taken into account and affects the composition of the resulting semantic space. Here, we report on experiments conducted on 27 clinical datasets, demonstrating that performance can be significantly improved by modeling EHR data in ensembles of semantic spaces, consisting of multiple semantic spaces built with different context window sizes. A follow-up investigation is conducted to study the impact on predictive performance as increasingly more semantic spaces are included in the ensemble, demonstrating that accuracy tends to improve with the number of semantic spaces, albeit not monotonically so. Finally, a number of different strategies for combining the semantic spaces are explored, demonstrating the advantage of early (feature) fusion over late (classifier) fusion. Semantic space ensembles allow multiple views of (sparse) data to be captured (densely) and thereby enable improved performance to be obtained on the task of detecting ADEs in EHRs. Aron Henriksson, Jing Zhao 0017, Henrik Boström, Hercules Dalianis |
BIBM | 3 |
| 2015 | Modeling heterogeneous clinical sequence data in semantic space for adverse drug event detectionabstractThe enormous amounts of data that are continuously recorded in electronic health record systems offer ample opportunities for data science applications to improve healthcare. There are, however, challenges involved in using such data for machine learning, such as high dimensionality and sparsity, as well as an inherent heterogeneity that does not allow the distinct types of clinical data to be treated in an identical manner. On the other hand, there are also similarities across data types that may be exploited, e.g., the possibility of representing some of them as sequences. Here, we apply the notions underlying distributional semantics, i.e., methods that model the meaning of words in semantic (vector) space on the basis of co-occurrence information, to four distinct types of clinical data: free-text notes, on the one hand, and clinical events, in the form of diagnosis codes, drug codes and measurements, on the other hand. Each semantic space contains continuous vector representations for every unique word and event, which can then be used to create representations of, e.g., care episodes that, in turn, can be exploited by the learning algorithm. This approach does not only reduce sparsity, but also takes into account, and explicitly models, similarities between various items, and it does so in an entirely data-driven fashion. Here, we report on a series of experiments using the random forest learning algorithm that demonstrate the effectiveness, in terms of accuracy and area under ROC curve, of the proposed representation form over the commonly used bag-of-items counterpart. The experiments are conducted on 27 real datasets that each involves the (binary) classification task of detecting a particular adverse drug event. It is also shown that combining structured and unstructured data leads to significant improvements over using only one of them. Aron Henriksson, Jing Zhao 0017, Henrik Boström, Hercules Dalianis |
DSAA | 3 |
| 2015 | Cascading adverse drug event detection in electronic health recordsabstractThe ability to detect adverse drug events (ADEs) in electronic health records (EHRs) is useful in many medical applications, such as alerting systems that indicate when an ADE-specific diagnosis code should be assigned. Automating the detection of ADEs can be attempted by applying machine learning to existing, labeled EHR data. How to do this in an effective manner is, however, an open question. The issues addressed in this study concern the granularity of the classification task: (1) If we wish to predict the occurrence of any ADE, is it advantageous to conflate the various ADE class labels prior to learning, or should they be merged post prediction? (2) If we wish to predict a family of ADEs or even a specific ADE, can the predictive performance be enhanced by dividing the classification task into a cascading scheme: predicting first, on a coarse level, whether there is an ADE or not, and, in the former case, followed by a more specific prediction on which family the ADE belongs to, and then finally a prediction on the specific ADE within that particular family? In this study, we conduct a series of experiments using a real, clinical dataset comprising healthcare episodes that have been assigned one of eight ADE-related diagnosis codes and a set of randomly extracted episodes that have not been assigned any ADE code. It is shown that, when distinguishing between ADEs and non-ADEs, merging the various ADE labels prior to learning leads to significantly higher predictive performance in terms of accuracy and area under ROC curve. A cascade of random forests is moreover constructed to determine either the family of ADEs or the specific class label; here, the performance is indeed enhanced compared to directly employing a one-step prediction. This study concludes that, if predictive performance is of primary importance, the cascading scheme should be the recommended approach over employing a one-step prediction for detecting ADEs in EHRs. Jing Zhao 0017, Aron Henriksson, Henrik Boström |
DSAA | 3 |
| 2015 | Post-analysis of multi-objective optimization solutions using decision treesabstractEvolutionary algorithms are often applied to solve multi-objective optimization problems. Such algorithms effectively generate solutions of wide spread, and have good convergence properties. However, they do not provide any characteristics of the fou Catarina Dudas, Amos H. C. Ng, Henrik Boström |
Intell. Data Anal. | 3 |
| 2015 | Bias reduction through conditional conformal predictionabstractConformal prediction (CP) is a relatively new framework in which predictive models output sets of predictions with a bound on the error rate, i.e., the probability of making an erroneous prediction is guaranteed to be equal to or less than a predefined significance level. Label-conditional conforma l prediction (LCCP) is a specialization of the framework which gives a bound on the error rate for each individual class. For datasets with class imbalance, many learning algorithms have a tendency to predict the majority class more often than the expected relative frequency, i.e., they are biased in favor of the majority class. In this study, the class bias of standard and label-conditional conformal predictors is investigated. An empirical investigation on 32 publicly available datasets with varying degrees of class imbalance is presented. The experimental results show that CP is highly biased towards the majority class on imbalanced datasets, i.e., it can be expected to make a majority of its errors on the minority class. LCCP, on the other hand, is not biased towards the majority class. Instead, the errors are distributed between the classes almost in accordance with the prior class distribution. Tuve Löfström, Henrik Boström, Henrik Linusson, Ulf Johansson |
Intell. Data Anal. | 2 |
| 2014 | Detecting adverse drug events with multiple representations of clinical measurementsabstractAdverse drug events (ADEs) are grossly under-reported in electronic health records (EHRs). This could be mitigated by methods that are able to detect ADEs in EHRs, thereby allowing for missing ADE-specific diagnosis codes to be identified and added. A crucial aspect of constructing such systems is to find proper representations of the data in order to allow the predictive modeling to be as accurate as possible. One category of EHR data that can be used as indicators of ADEs are clinical measurements. However, using clinical measurements as features is not unproblematic due to the high rate of missing values and they can be repeated a variable number of times in each patient health record. In this study, five basic representations of clinical measurements are proposed and evaluated to handle these two problems. An empirical investigation using random forest on 27 datasets from a real EHR database with different ADE targets is presented, demonstrating that the predictive performance, in terms of accuracy and area under ROC curve, is higher when representing clinical measurements crudely as whether they were taken or how many times they were taken by a patient. Furthermore, a sixth alternative, combining all five basic representations, significantly outperforms using any of the basic representation except for one. A subsequent analysis of variable importance is also conducted with this fused feature set, showing that when clinical measurements have a high missing rate, the number of times they were taken by one patient is ranked as more informative than looking at their actual values. The observation from random forest is also confirmed empirically using other commonly employed classifiers. This study demonstrates that the way in which clinical measurements from EHRs are presented has a high impact for ADE detection, and that using multiple representations outperforms using a basic representation. Jing Zhao 0017, Aron Henriksson, Lars Asker, Henrik Boström |
BIBM | 4 |
| 2014 | Regression trees for streaming data with local performance guaranteesabstractOnline predictive modeling of streaming data is a key task for big data analytics. In this paper, a novel approach for efficient online learning of regression trees is proposed, which continuously updates, rather than retrains, the tree as more labeled data become available. A conformal predictor outputs prediction sets instead of point predictions; which for regression translates into prediction intervals. The key property of a conformal predictor is that it is always valid, i.e., the error rate, on novel data, is bounded by a preset significance level. Here, we suggest applying Mondrian conformal prediction on top of the resulting models, in order to obtain regression trees where not only the tree, but also each and every rule, corresponding to a path from the root node to a leaf, is valid. Using Mondrian conformal prediction, it becomes possible to analyze and explore the different rules separately, knowing that their accuracy, in the long run, will not be below the preset significance level. An empirical investigation, using 17 publicly available data sets, confirms that the resulting rules are independently valid, but also shows that the prediction intervals are smaller, on average, than when only the global model is required to be valid. All-in-all, the suggested method provides a data miner or a decision maker with highly informative predictive models of streaming data. Ulf Johansson, Cecilia Sönströd, Henrik Linusson, Henrik Boström |
IEEE BigData | 4 |
| 2014 | A peek into the black box: exploring classifiers by randomization
Andreas Henelius, Kai Puolamäki, Henrik Boström, Lars Asker, Panagiotis Papapetrou |
Data Min. Knowl. Discov. | 3 |
| 2014 | Regression conformal prediction with random forests
Ulf Johansson, Henrik Boström, Tuve Löfström, Henrik Linusson |
Mach. Learn. | 2 |
| 2013 | Generalization of Malaria Incidence Prediction Models by Correcting Sample Selection Bias
Orlando P. Zacarias, Henrik Boström |
ADMA (2) | 2 |
| 2013 | Predicting Adverse Drug Events by Analyzing Electronic Patient Records
Isak Karlsson, Jing Zhao 0017, Lars Asker, Henrik Boström |
AIME | 4 |
| 2013 | Evolved decision trees as conformal predictorsabstractIn conformal prediction, predictive models output sets of predictions with a bound on the error rate. In classification, this translates to that the probability of excluding the correct class is lower than a predefined significance level, in the long run. Since the error rate is guaranteed, the most important criterion for conformal predictors is efficiency. Efficient conformal predictors minimize the number of elements in the output prediction sets, thus producing more informative predictions. This paper presents one of the first comprehensive studies where evolutionary algorithms are used to build conformal predictors. More specifically, decision trees evolved using genetic programming are evaluated as conformal predictors. In the experiments, the evolved trees are compared to decision trees induced using standard machine learning techniques on 33 publicly available benchmark data sets, with regard to predictive performance and efficiency. The results show that the evolved trees are generally more accurate, and the corresponding conformal predictors more efficient, than their induced counterparts. One important result is that the probability estimates of decision trees when used as conformal predictors should be smoothed, here using the Laplace correction. Finally, using the more discriminating Brier score instead of accuracy as the optimization criterion produced the most efficient conformal predictions. Ulf Johansson, Rikard König, Tuve Löfström, Henrik Boström |
IEEE Congress on Evolutionary Computation | 4 |
| 2013 | Conformal Prediction Using Decision TreesabstractConformal prediction is a relatively new framework in which the predictive models output sets of predictions with a bound on the error rate, i.e., in a classification context, the probability of excluding the correct class label is lower than a predefined significance level. An investigation of the use of decision trees within the conformal prediction framework is presented, with the overall purpose to determine the effect of different algorithmic choices, including split criterion, pruning scheme and way to calculate the probability estimates. Since the error rate is bounded by the framework, the most important property of conformal predictors is efficiency, which concerns minimizing the number of elements in the output prediction sets. Results from one of the largest empirical investigations to date within the conformal prediction framework are presented, showing that in order to optimize efficiency, the decision trees should be induced using no pruning and with smoothed probability estimates. The choice of split criterion to use for the actual induction of the trees did not turn out to have any major impact on the efficiency. Finally, the experimentation also showed that when using decision trees, standard inductive conformal prediction was as efficient as the recently suggested method cross-conformal prediction. This is an encouraging results since cross-conformal prediction uses several decision trees, thus sacrificing the interpretability of a single decision tree. Ulf Johansson, Henrik Boström, Tuve Löfström |
ICDM | 2 |
| 2013 | Random brainsabstractIn this paper, we introduce and evaluate a novel method, called random brains, for producing neural network ensembles. The suggested method, which is heavily inspired by the random forest technique, produces diversity implicitly by using bootstrap training and randomized architectures. More specifically, for each base classifier multilayer perceptron, a number of randomly selected links between the input layer and the hidden layer are removed prior to training, thus resulting in potentially weaker but more diverse base classifiers. The experimental results on 20 UCI data sets show that random brains obtained significantly higher accuracy and AUC, compared to standard bagging of similar neural networks not utilizing randomized architectures. The analysis shows that the main reason for the increased ensemble performance is the ability to produce effective diversity, as indicated by the increase in the difficulty diversity measure. Ulf Johansson, Tuve Löfström, Henrik Boström |
IJCNN | 3 |
| 2013 | Effective utilization of data in inductive conformal prediction using ensembles of neural networksabstractConformal prediction is a new framework producing region predictions with a guaranteed error rate. Inductive conformal prediction (ICP) was designed to significantly reduce the computational cost associated with the original transductive online approach. The drawback of inductive conformal prediction is that it is not possible to use all data for training, since it sets aside some data as a separate calibration set. Recently, cross-conformal prediction (CCP) and bootstrap conformal prediction (BCP) were proposed to overcome that drawback of inductive conformal prediction. Unfortunately, CCP and BCP both need to build several models for the calibration, making them less attractive. In this study, focusing on bagged neural network ensembles as conformal predictors, ICP, CCP and BCP are compared to the very straightforward and cost-effective method of using the out-of-bag estimates for the necessary calibration. Experiments on 34 publicly available data sets conclusively show that the use of out-of-bag estimates produced the most efficient conformal predictors, making it the obvious preferred choice for ensembles in the conformal prediction framework. Tuve Löfström, Ulf Johansson, Henrik Boström |
IJCNN | 3 |
| 2013 | Comparative analysis of the use of chemoinformatics-based and substructure-based descriptors for quantitative structure-activity relationship (QSAR) modelingabstractQuantitative structure-activity relationship (QSAR) models have gained popularity in the pharmaceutical industry due to their potential to substantially decrease drug development costs by reducing expensive laboratory and clinical tests. QSAR modelin Thashmee Karunaratne, Henrik Boström, Ulf Norinder |
Intell. Data Anal. | 2 |
| 2012 | Choice of dimensionality reduction methods for feature and classifier fusion with nearest neighbor classifiers
Sampath Deegalla, Henrik Boström, Keerthi Walgama |
FUSION | 2 |
| 2012 | Can Frequent Itemset Mining Be Efficiently and Effectively Used for Learning from Graph Data?abstractStandard graph learning approaches are often challenged by the computational cost involved when learning from very large sets of graph data. One approach to overcome this problem is to transform the graphs into less complex structures that can be more efficiently handled. One obvious potential drawback of this approach is that it may degrade predictive performance due to loss of information caused by the transformations. An investigation of the tradeoff between efficiency and effectiveness of graph learning methods is presented, in which state-of-the-art graph mining approaches are compared to representing graphs by itemsets, using frequent itemset mining to discover features to use in prediction models. An empirical evaluation on 18 medicinal chemistry datasets is presented, showing that employing frequent itemset mining results in significant speedups, without sacrificing predictive performance for both classification and regression. Thashmee Karunaratne, Henrik Boström |
ICMLA (1) | 2 |
| 2012 | Obtaining accurate and comprehensible classifiers using oracle coachingabstractWhile ensemble classifiers often reach high levels of predictive performance, the resulting models are opaque and hence do not allow direct interpretation. When employing methods that do generate transparent models, predictive performance typically h Ulf Johansson, Cecilia Sönströd, Tuve Löfström, Henrik Boström |
Intell. Data Anal. | 4 |
| 2012 | Forests of Probability Estimation TreesabstractProbability estimation trees (PETs) generalize classification trees in that they assign class probability distributions instead of class labels to examples that are to be classified. This property has been demonstrated to allow PETs to outperform classification trees with respect to ranking performance, as measured by the area under the ROC curve (AUC). It has further been shown that the use of probability correction improves the performance of PETs. This has lead to the use of probability correction also in forests of PETs. However, it was recently observed that probability correction may in fact deteriorate performance of forests of PETs. A more detailed study of the phenomenon is presented and the reasons behind this observation are analyzed. An empirical investigation is presented, comparing forests of classification trees to forests of both corrected and uncorrected PETS on 34 data sets from the UCI repository. The experiment shows that a small forest (10 trees) of probability corrected PETs gives a higher AUC than a similar-sized forest of classification trees, hence providing evidence in favor of using forests of probability corrected PETs. However, the picture changes when increasing the forest size, as the AUC is no longer improved by probability correction. For accuracy and squared error of predicted class probabilities (Brier score), probability correction even leads to a negative effect. An analysis of the mean squared error of the trees in the forests and their variance, shows that although probability correction results in trees that are more correct on average, the variance is reduced at the same time, leading to an overall loss of performance for larger forests. The main conclusions are that probability correction should only be employed in small forests of PETs, and that for larger forests, classification trees and PETs are equally good alternatives. Henrik Boström |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2010 | Pre-Processing Structured Data for Standard Machine Learning Algorithms by Supervised Graph Propositionalization - A Case Study with Medicinal Chemistry DatasetsabstractGraph propositionalization methods can be used to transform structured and relational data into fixed-length feature vectors, enabling standard machine learning algorithms to be used for generating predictive models. It is however not clear how well different propositionalization methods work in conjunction with different standard machine learning algorithms. Three different graph propositionalization methods are investigated in conjunction with three standard learning algorithms: random forests, support vector machines and nearest neighbor classifiers. An experiment on 21 datasets from the domain of medicinal chemistry shows that the choice of propositionalization method may have a significant impact on the resulting accuracy. The empirical investigation further shows that for datasets from this domain, the use of the maximal frequent item set approach for propositionalization results in the most accurate classifiers, significantly outperforming the two other graph propositionalization methods considered in this study, SUBDUE and MOSS, for all three learning methods. Thashmee Karunaratne, Henrik Boström, Ulf Norinder |
ICMLA | 2 |
| 2010 | Comparing methods for generating diverse ensembles of artificial neural networksabstractIt is well-known that ensemble performance relies heavily on sufficient diversity among the base classifiers. With this in mind, the strategy used to balance diversity and base classifier accuracy must be considered a key component of any ensemble algorithm. This study evaluates the predictive performance of neural network ensembles, specifically comparing straightforward techniques to more sophisticated. In particular, the sophisticated methods GASEN and NegBagg are compared to more straightforward methods, where each ensemble member is trained independently of the others. In the experimentation, using 31 publicly available data sets, the straightforward methods clearly outperformed the sophisticated methods, thus questioning the use of the more complex algorithms. Tuve Löfström, Ulf Johansson, Henrik Boström |
IJCNN | 3 |
| 2010 | Pin-pointing concept descriptionsabstractIn this study, the task of obtaining accurate and comprehensible concept descriptions of a specific set of production instances has been investigated. The suggested method, inspired by rule extraction and transductive learning, uses a highly accurate opaque model, called an oracle, to coach construction of transparent decision list models. The decision list algorithms evaluated are JRip and four different variants of Chipper, a technique specifically developed for concept description. Using 40 real-world data sets from the drug discovery domain, the results show that employing an oracle coach to label the production data resulted in significantly more accurate and smaller models for almost all techniques. Furthermore, augmenting normal training data with production data labeled by the oracle also led to significant increases in predictive performance, but with a slight increase in model size. Of the techniques evaluated, normal Chipper optimizing FOIL's information gain and allowing conjunctive rules was clearly the best. The overall conclusion is that oracle coaching works very well for concept description. Cecilia Sönströd, Ulf Johansson, Henrik Boström, Ulf Norinder |
SMC | 3 |
| 2009 | Ensemble member selection using multi-objective optimizationabstractBoth theory and a wealth of empirical studies have established that ensembles are more accurate than single predictive models. Unfortunately, the problem of how to maximize ensemble accuracy is, especially for classification, far from solved. In essence, the key problem is to find a suitable criterion, typically based on training or selection set performance, highly correlated with ensemble accuracy on novel data. Several studies have, however, shown that it is difficult to come up with a single measure, such as ensemble or base classifier selection set accuracy, or some measure based on diversity, that is a good general predictor for ensemble test accuracy. This paper presents a novel technique that for each learning task searches for the most effective combination of given atomic measures, by means of a genetic algorithm. Ensembles built from either neural networks or random forests were empirically evaluated on 30 UCI datasets. The experimental results show that when using the generated combined optimization criteria to rank candidate ensembles, a higher test set accuracy for the top ranked ensemble was achieved, compared to using ensemble accuracy on selection data alone. Furthermore, when creating ensembles from a pool of neural networks, the use of the generated combined criteria was shown to generally outperform the use of estimated ensemble accuracy as the single optimization criterion. Tuve Löfström, Ulf Johansson, Henrik Boström |
CIDM | 3 |
| 2009 | Fusion of dimensionality reduction methods: A case study in microarray classification
Sampath Deegalla, Henrik Boström |
FUSION | 2 |
| 2009 | Improving Fusion of Dimensionality Reduction Methods for Nearest Neighbor ClassificationabstractIn previous studies, performance improvement of nearest neighbor classification of high dimensional data, such as microarrays, has been investigated using dimensionality reduction. It has been demonstrated that the fusion of dimensionality reduction methods, either by fusing classifiers obtained from each set of reduced features, or by fusing all reduced features are better than using any single dimensionality reduction method. However, none of the fusion methods consistently outperform the use of a single dimensionality reduction method. Therefore, a new way of fusing features and classifiers is proposed, which is based on searching for the optimal number of dimensions for each considered dimensionality reduction method. An empirical evaluation on microarray classification is presented, comparing classifier and feature fusion with and without the proposed method, in conjunction with three dimensionality reduction methods; Principal Component Analysis (PCA), Partial Least Squares (PLS) and Information Gain (IG). The new classifier fusion method outperforms the previous in 4 out of 8 cases, and is on par with the best single dimensionality reduction method. The novel feature fusion method is however outperformed by the previous method, which selects the same number of features from each dimensionality reduction method. Hence, it is concluded that the idea of optimizing the number of features separately for each dimensionality reduction method can only be recommended for classifier fusion. Sampath Deegalla, Henrik Boström |
ICMLA | 2 |
| 2009 | Graph Propositionalization for Random ForestsabstractGraph propositionalization methods transform structured and relational data into a fixed-length feature vector format that can be used by standard machine learning methods. However, the choice of propositionalization method may have a significant impact on the performance of the resulting classifier. Six different propositionalization methods are evaluated when used in conjunction with random forests. The empirical evaluation shows that the choice of propositionalization method has a significant impact on the resulting accuracy for structured data sets. The results furthermore show that the maximum frequent itemset approach and a combination of this approach and maximal common substructures turn out to be the most successful propositionalization methods for structured data, each significantly outperforming the four other considered methods. Thashmee Karunaratne, Henrik Boström |
ICMLA | 2 |
| 2008 | On evidential combination rules for ensemble classifiers
Henrik Boström, Ronnie Johansson, Alexander Karlsson 0001 |
FUSION | 1 |
| 2008 | Calibrating Random ForestsabstractWhen using the output of classifiers to calculate the expected utility of different alternatives in decision situations, the correctness of predicted class probabilities may be of crucial importance. However, even very accurate classifiers may output class probabilities of rather poor quality. One way of overcoming this problem is by means of calibration, i.e., mapping the original class probabilities to more accurate ones. Previous studies have however indicated that random forests are difficult to calibrate by standard calibration methods. In this work, a novel calibration method is introduced, which is based on a recent finding that probabilities predicted by forests of classification trees have a lower squared error compared to those predicted by forests of probability estimation trees (PETs). The novel calibration method is compared to the two standard methods, Platt scaling and isotonic regression, on 34 datasets from the UCI repository. The experiment shows that random forests of PETs calibrated by the novel method significantly outperform uncalibrated random forests of both PETs and classification trees, as well as random forests calibrated with the two standard methods, with respect to the squared error of predicted class probabilities. Henrik Boström |
ICMLA | 1 |
| 2008 | On the Use of Accuracy and Diversity Measures for Evaluating and Selecting Ensembles of ClassifiersabstractThe test set accuracy for ensembles of classifiers selected based on single measures of accuracy and diversity as well as combinations of such measures is investigated. It is found that by combining measures, a higher test set accuracy may be obtained than by using any single accuracy or diversity measure. It is further investigated whether a multi-criteria search for an ensemble that maximizes both accuracy and diversity leads to more accurate ensembles than by optimizing a single criterion. The results indicate that it might be more beneficial to search for ensembles that are both accurate and diverse. Furthermore, the results show that diversity measures could compete with accuracy measures as selection criterion. Tuve Löfström, Ulf Johansson, Henrik Boström |
ICMLA | 3 |
| 2008 | Comprehensible Models for Predicting Molecular Interaction with Heart-Regulating GenesabstractWhen using machine learning for in silico modeling, the goal is normally to obtain highly accurate predictive models. Often, however, models should also bring insights into interesting relationships in the domain. It is then desirable that machine learning techniques have the ability to obtain small and transparent models, where the user can control the tradeoff between accuracy, comprehensibility and coverage. In this study, three different decision list algorithms are evaluated on a dataset concerning the interaction of molecules with a human gene that regulates heart functioning (hERG). The results show that decision list algorithms can obtain predictive performance not far from the state-of-the-art method random forests, but also that algorithms focusing on accuracy alone may produce complex decision lists that are very hard to interpret. The experiments also show that by sacrificing accuracy only to a limited degree, comprehensibility (measured as both model size and classification complexity) can be improved remarkably. Cecilia Sönströd, Ulf Johansson, Ulf Norinder, Henrik Boström |
ICMLA | 4 |
| 2008 | The problem with ranking ensembles based on training or validation performanceabstractThe main purpose of this study was to determine whether it is possible to somehow use results on training or validation data to estimate ensemble performance on novel data. With the specific setup evaluated; i.e. using ensembles built from a pool of independently trained neural networks and targeting diversity only implicitly, the answer is a resounding no. Experimentation, using 13 UCI datasets, shows that there is in general nothing to gain in performance on novel data by choosing an ensemble based on any of the training measures evaluated here. This is despite the fact that the measures evaluated include all the most frequently used; i.e. ensemble training and validation accuracy, base classifier training and validation accuracy, ensemble training and validation AUC and two diversity measures. The main reason is that all ensembles tend to have quite similar performance, unless we deliberately lower the accuracy of the base classifiers. The key consequence is, of course, that a data miner can do no better than picking an ensemble at random. In addition, the results indicate that it is futile to look for an algorithm aimed at optimizing ensemble performance by somehow selecting a subset of available base classifiers. Ulf Johansson, Tuve Löfström, Henrik Boström |
IJCNN | 3 |
| 2007 | Feature vs. classifier fusion for predictive data mining a case study in pesticide classificationabstractTwo strategies for fusing information from multiple sources when generating predictive models in the domain of pesticide classification are investigated: i) fusing different sets of features (molecular descriptors) before building a model and ii) fusing the classifiers built from the individual descriptor sets. An empirical investigation demonstrates that the choice of strategy can have a significant impact on the predictive performance. Furthermore, the experiment shows that the best strategy is dependent on the type of predictive model considered. When generating a decision tree for pesticide classification, a statistically significant difference in accuracy is observed in favor of combining predictions from the individual models compared to generating a single model from the fused set of molecular descriptors. On the other hand, when the model consists of an ensemble of decision trees, a statistically significant difference in accuracy is observed in favor of building the model from the fused set of descriptors compared to fusing ensemble models built from the individual sources. Henrik Boström |
FUSION | 1 |
| 2007 | Estimating class probabilities in random forestsabstractFor both single probability estimation trees (PETs) and ensembles of such trees, commonly employed class probability estimates correct the observed relative class frequencies in each leaf to avoid anomalies caused by small sample sizes. The effect of such corrections in random forests of PETs is investigated, and the use of the relative class frequency is compared to using two corrected estimates, the Laplace estimate and the m-estimate. An experiment with 34 datasets from the UCI repository shows that estimating class probabilities using relative class frequency clearly outperforms both using the Laplace estimate and the m-estimate with respect to accuracy, area under the ROC curve (AUC) and Brier score. Hence, in contrast to what is commonly employed for PETs and ensembles of PETs, these results strongly suggest that a non-corrected probability estimate should be used in random forests of PETs. The experiment further shows that learning random forests of PETs using relative class frequency significantly outperforms learning random forests of classification trees (i.e., trees for which only an unweighted vote on the most probable class is counted) with respect to both accuracy and AUC, but that the latter is clearly ahead of the former with respect to Brier score. Henrik Boström |
ICMLA | 1 |
| 2007 | Classification of Microarrays with kNN: Comparison of Dimensionality Reduction Methods
Sampath Deegalla, Henrik Boström |
IDEAL | 2 |
| 2007 | Maximizing the Area under the ROC Curve with Decision Lists and Rule SetsabstractDecision lists (or ordered rule sets) have two attractive properties compared to unordered rule sets: they require a simpler classification procedure and they allow for a more compact representation. However, it is an open question what effect these properties have on the area under the ROC curve (AUC). Two ways of forming decision lists are considered in this study: by generating a sequence of rules, with a default rule for one of the classes, and by imposing an order upon rules that have been generated for all classes. An empirical investigation shows that the latter method gives a significantly higher AUC than the former, demonstrating that the compactness obtained by using one of the classes as a default is indeed associated with a cost. Furthermore, by using all applicable rules rather than the first in an ordered set, an even further significant improvement in AUC is obtained, demonstrating that the simple classification procedure is also associated with a cost. The observed gains in AUC for unordered rule sets compared to decision lists can be explained by that learning rules for all classes as well as combining multiple rules allow for examples to be ranked according to a more fine-grained scale compared to when applying rules in a fixed order and providing a default rule for one of the classes. Henrik Boström |
SDM | 1 |
| 2006 | Reducing High-Dimensional Data by Principal Component Analysis vs. Random Projection for Nearest Neighbor ClassificationabstractThe computational cost of using nearest neighbor classification often prevents the method from being applied in practice when dealing with high-dimensional data, such as images and micro arrays. One possible solution to this problem is to reduce the dimensionality of the data, ideally without loosing predictive performance. Two different dimensionality reduction methods, principle component analysis (PCA) and random projection (RP), are investigated for this purpose and compared w.r.t. the performance of the resulting nearest neighbor classifier on five image data sets and five micro array data sets. The experiment results demonstrate that PCA outperforms RP for all data sets used in this study. However, the experiments also show that PCA is more sensitive to the choice of the number of reduced dimensions. After reaching a peak, the accuracy degrades with the number of dimensions for PCA, while the accuracy for RP increases with the number of dimensions. The experiments also show that the use of PCA and RP may even outperform using the non-reduced feature set (in 9 respectively 6 cases out of 10), hence not only resulting in more efficient, but also more effective, nearest neighbor classification Sampath Deegalla, Henrik Boström |
ICMLA | 2 |
| 2004 | Resolving rule conflicts with double induction
Tony Lindgren, Henrik Boström |
Intell. Data Anal. | 2 |
| 2003 | Resolving Rule Conflicts with Double Induction
Tony Lindgren, Henrik Boström |
IDA | 2 |
| 2002 | Classification with Intersecting Rules
Tony Lindgren, Henrik Boström |
ALT | 2 |
| 2002 | Rule Induction for Classification of Gene Expression Array Data
Per Lidén, Lars Asker, Henrik Boström |
PKDD | 3 |
| 2001 | Automatic Keyword Extraction Using Domain Knowledge
Anette Hulth, Jussi Karlgren, Henrik Boström, Lars Asker |
CICLing | 4 |
| 2001 | Classifying Uncovered Examples by Rule Stretching
Martin Eineborg, Henrik Boström |
ILP | 2 |
| 2001 | Boosting interval based literals
Juan José Rodríguez Diez, Carlos J. Alonso-González, Henrik Boström |
Intell. Data Anal. | 3 |
| 2000 | Learning First Order Logic Time Series Classifiers: Rules and Boosting
Juan José Rodríguez Diez, Carlos J. Alonso-González, Henrik Boström |
PKDD | 3 |
| 1998 | Predicate Invention and Learning from Positive Examples Only
Henrik Boström |
ECML | 1 |
| 1997 | IMPUT: An Interactive Learning Tool Based on Program SpecializationabstractThe algorithm SPECTRE specializes logic programs with respect to positive and negative examples by applying the transformation rule unfolding together with clause removal. The method IMPUT presented in this paper gives a modified version of this algorithm by integrating the algorithmic debugging system IDTS with SPECTRE. The main idea of the IMPUT method, is that the identification of a clause to be unfolded has a crucial importance on the effectiveness of the specialization process. The debugging system IDTS is used to identify this buggy clause. Zoltán Alexin, Tibor Gyimóthy, Henrik Boström |
Intell. Data Anal. | 3 |
| 1996 | Integrating Algorithmic Debugging and Unfolding Transformation in an Interactive Learner
Zoltán Alexin, Tibor Gyimóthy, Henrik Boström |
ECAI | 3 |
| 1996 | Theory-Guideed Induction of Logic Programs by Inference of Regular Languages
Henrik Boström |
ICML | 1 |
| 1995 | JIGSAW: Puzzling together RUTH and SPECTRE (Extended Abstract)
Hilde Adé, Henrik Boström |
ECML | 2 |
| 1995 | Specialization of Recursive Predicates
Henrik Boström |
ECML | 1 |
| 1995 | Covering vs. Divide-and-Conquer for Top-Down Induction of Logic Programs
Henrik Boström |
IJCAI | 1 |
| 1993 | Improving Example-Guided Unfolding
Henrik Boström |
ECML | 1 |
| 1990 | Generalizing the Order of Goals as an Approach to Generalizing Number
Henrik Boström |
ML | 1 |