Gianluca Bontempi

dblp:64/6336 · DBLP profile ↗
← Back
71ranked-venue papers
19as first author
15since 2021 · last 2026
0000-0001-8621-316XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 13 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 15 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorTheory of computation · 4 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Identifying counterfactual probabilities using bivariate distributions and uplift modeling
abstract
Uplift modeling estimates the causal effect of an intervention as the difference between potential outcomes under treatment and control, whereas counterfactual identification aims to recover the joint distribution of these potential outcomes (e.g., "Would this customer still have churned had we given them a marketing offer?").This joint counterfactual distribution provides richer information than the uplift but is harder to estimate.However, the two approaches are synergistic: uplift models can be leveraged for counterfactual estimation.We propose a counterfactual estimator that fits a bivariate beta distribution to predicted uplift scores, yielding posterior distributions over counterfactual outcomes.Our approach requires no causal assumptions beyond those of uplift modeling.Simulations show the efficacy of the approach, which can be applied, for example, to the problem of customer churn in telecom, where it reveals insights unavailable to standard ML or uplift models alone.* Research carried out during the author's prior affiliation with the Machine Learning Group.† This publication benefits from the support of the Walloon Region to Pr.G. Bontempi as part of the
Théo Verhelst, Gianluca Bontempi
ESANN2
2026 Epicenter Retrieval from Synthetic Sparse Acoustic Seismograms: A Machine Learning Benchmark on the Marmousi Field
Pascal Tribel, Gianluca Bontempi
SIMULTECH2
2026 On the integration of domain adaptation and causal discovery in digital twins: A case study for plastic injection molding
Gian Marco Paldino, Olivier Caelen, Marouene Oueslati, Marc Ansay, Tojo Valisoa A. Johanesa, Gianluca Bontempi
Future Gener. Comput. Syst.6
2026 Fraud-RLA: A Reinforcement Learning Adversarial Attack Against Credit Card Fraud Detection
abstract
Adversarial attacks pose a significant threat to data-driven systems, and researchers have devoted considerable effort to studying them. Despite its economic relevance, credit card fraud detection has received comparatively little attention. To address this gap, we propose a novel threat model that highlights the limitations of existing attacks and motivates new approaches. We introduce Fraud-RLA, an adversarial attack against credit card Fraud Detection Systems that leverages Reinforcement Learning to evade detection. Fraud-RLA is designed to maximize the amount stolen by optimizing the exploration-exploitation trade-off while requiring substantially less prior knowledge than competing methods. Our experiments on a realistic Fraud Detection System show that Fraud-RLA is effective, even under the severe limitations imposed by our threat model.
Daniele Lunghi, Yannick Molinghen, Alkis Simitsis, Tom Lenaerts, Gianluca Bontempi
IEEE Trans. Dependable Secur. Comput.5
2025 A Simulation Tool to Assess the Impact of Deviation Plans on Disruptive Events of Urban Traffic
abstract
Urban traffic management faces growing challenges in evaluating and mitigating the impact of disruptive events, such as road closures, on vehicular traffic flow. This paper presents the design and development of an interactive tool to define and assess the impact of road deviation plans on vehicular traffic. The proposed tool targets traffic management experts and is expected to support them in defining and comparing alternative solutions to mitigate disruptive events (e.g. road/tunnel closures for maintenance). The proposed tool, called TrafficTwin, can be adapted to different areas of the town, make use of different traffic models (either synthetic or calibrated) and visualize several quantitative statistics to assess and compare alternative deviation plans. We evaluate the proposed tool using a synthetic traffic model and assess the pertinence of the simulation tool to support the decision-making process in transportation infrastructure management.
Davide Guastella, Moisés Silva-Muñoz, Eladio Montero-Porras, Gianluca Bontempi
SIMULTECH4
2025 On many-objective feature selection and the need for interpretability
abstract
Big data comes with the challenge of containing irrelevant and redundant information (i.e., features). Given that a single objective cannot fully capture a feature’s relevance, a Many-Objective Feature Selection (MOFS) approach able to accommodate various relevant perspectives is preferred for identifying the most appropriate features in a given context. However, MOFS produces a large set of solutions whose interpretability has been largely overlooked. First, we demonstrate the relevance of MOFS and establish its necessity by considering up to six objectives using a genetic algorithm and Naive Bayes on ten datasets for classification tasks. Then, we propose a novel methodology to improve the interpretability of MOFS results in order to support the data scientist in selecting the subset of features pertinent to their use case. Our methodology is instantiated as an intuitive and interactive dashboard that provides insights into the results beyond the pure numerical representation of the objectives being considered and evaluated with 50 participants. The outcome shows that it addresses the need for a methodological approach and comprehensive visualization to achieve interoperability. • This paper presents an empirical analysis of Many-Objective Feature Selection (MOFS). • A novel methodology for interpreting MOFS results is proposed. • The methodology is implemented via an interactive dashboard. • Statistical comparison of the methodology versus the tabular method with 50 subjects. • Sensitivity analysis of the methodology to changes in objective weights.
Uchechukwu Njoku, Alberto Abelló, Besim Bilalli, Gianluca Bontempi
Expert Syst. Appl.4
2024 A data-science pipeline to enable the Interpretability of Many-Objective Feature Selection
Uchechukwu Njoku, Alberto Abelló, Besim Bilalli, Gianluca Bontempi
DOLAP4
2024 Finding Relevant Information in Big Datasets with ML
Uchechukwu Njoku, Alberto Abelló, Besim Bilalli, Gianluca Bontempi
EDBT4
2024 Assessing adversarial attacks in real-world fraud detection
abstract
In the digital economy, the growing demand for data sharing and trading makes data a critical asset. To facilitate data acquisition between data owners and data buyers, tools such as data markets have emerged. Such modern data ecosystems typically comprise several interconnected components, often based on machine learning. Unfortunately, as practice has shown, such components are vulnerable to adversarial attacks. Lacking proper security assessment measures is a severe limitation for the success of data markets. In this paper, we delve into the challenges posed by adversarial attacks using credit card fraud detection as an example use case. We show that popular techniques such as penetration testing through existing adversarial attacks is not a viable approach, and corroborate our analysis by showing how a naive random sample attack outperforms all tested methods when considering the specifics of the fraud detection problem. Motivated by this result, we propose alternative assessment approaches and discuss promising research directions for increasing our understanding of models’ robustness.
Daniele Lunghi, Alkis Simitsis, Gianluca Bontempi
ICWS3
2024 Assessment of catastrophic forgetting in continual credit card fraud detection
Bertrand Lebichot, Wissam Siblini, Gian Marco Paldino, Yann-Aël Le Borgne, Frédéric Oblé, Gianluca Bontempi
Expert Syst. Appl.6
2024 Partial counterfactual identification and uplift modeling: theoretical results and real-world assessment
Théo Verhelst, Denis Mercier, Jeevan Shrestha, Gianluca Bontempi
Mach. Learn.4
2023 Wrapper Methods for Multi-Objective Feature Selection
Uchechukwu Njoku, Besim Bilalli, Alberto Abelló, Gianluca Bontempi
EDBT4
2022 Impact of Filter Feature Selection on Classification: An Empirical Study
Uchechukwu Njoku, Alberto Abelló, Besim Bilalli, Gianluca Bontempi
DOLAP4
2021 Predicting Reach to Find Persuadable Customers: Improving Uplift Models for Churn Prevention
Théo Verhelst, Jeevan Shrestha, Denis Mercier, Jean-Christophe Dewitte, Gianluca Bontempi
DS5
2021 Combining unsupervised and supervised learning in credit card fraud detection
Fabrizio Carcillo, Yann-Aël Le Borgne, Olivier Caelen, Yacine Kessaci, Frédéric Oblé, Gianluca Bontempi
Inf. Sci.6
2020 On-Board Unit Big Data: Short-term Traffic Forecasting in Urban Transportation Networks
abstract
Short-term traffic prediction in transportation net-work is a hot topic for next generation smart cities. Despite several research efforts, there is still a lack of consensus about the most effective way to predict traffic network-wide. Also, the majority of existing studies rely on either small data sets or limited portions of the transportation network. This paper presents the results of a forecasting study related to on-board units data of heavy-goods vehicles in the Brussels-Capital Region. The analysis, enabled by deployment of a big data infrastructure, concerns the short horizon (one hour ahead) forecasting of trucks flow (vehicles/hour) and mean speed (km/hour) over Brussels-Capital Region's network. Both parametric and non-parametric (notably machine learning) models are designed and assessed along with different strategies (simple and linear forecasts combination, individual model selection). The paper shows the potential of large amounts of real-time data obtained from moving sensors where forecasting techniques are applied network-wide. In particular, simple combination schemes (notably the simple median combination forecasts) appear to outperform the other methods in terms of prediction accuracy.
Giovanni Buroni, Yann-Aël Le Borgne, Gianluca Bontempi, Daniele Raimondi, Karl Determe
DSAA3
2020 Incremental learning strategies for credit cards fraud detection: Extended abstract
abstract
Every second, Fraud Detection Systems (FDS) check streams of thousands of credit or debit card transactions. Most of the models in production and in the literature rely on batch learning, wasting part of the most recent data. Incremental learning may be beneficial in terms of computational/architectural cost and confidentiality (e.g. avoiding the storage of sensitive data). The full paper focuses on two questions related to the use of incremental learning for FDS: (1) Can incremental learning be as competitive as batch and retraining approaches? (2) Can combining incremental, ensembles, diversity, and transfer learning lead to efficient models under concept drift? We investigate those elements and provide an experimental evaluation on a real-life case study including more than 150 days of e-commerce transactions (or 50 million transactions).
Bertrand Lebichot, Gian Marco Paldino, Gianluca Bontempi, Wissam Siblini, Liyun He-Guelton, Frédéric Oblé
DSAA3
2020 Learning causal dependencies in large-variate time series
abstract
A major challenge in causal inference from observational data is to discriminate between associative dependencies and effective causal relationships. This is particularly challenging in large-variate and temporal settings (e.g. in spatio-temporal time series) where the multivariate nature of interactions induces a significant correlation between most of the variables. In recent years, a number of data-driven approaches have been proposed to learn the mapping between some features of the data distribution and the probability of a causal connection between a pair of variables. Most state-of-the-art approaches, however, deal with bivariate cases neglecting the role of the context determined by the other variables. This is a strong limitation in large-variate and temporal settings which are the object of this study. In order to address the context issue, this paper introduces a new set of descriptors based on interaction information to featurize the context and justifies its introduction by using a graphical modeling formalism. The resulting causal inference method is assessed on a number of large-variate synthetic stationary time series. The assessment shows that the proposed method outperforms several state-of-the-art causal inference techniques.
Gianluca Bontempi
IJCNN1
2020 EEG-based brain-computer interface for alpha speed control of a small robot using the MUSE headband
abstract
Non-invasive BMI applications are increasingly used in different contexts ranging from industrial, clinical and gaming. After having tested the difference between a classical EEG recorder with electroconductive gel (ANT system) and the MUSE EEG headband, we studied the BCI performances of the later during the control of a small robot. We demonstrated that the participants were able to successfully control the robot using an online brain-computer interface based on the signal power in different frequency bands (delta, theta and alpha) characterizing the eyes-opened and relaxed eyes-closed states. Additionally, we performed a correlation analysis which demonstrated that the BCI commands were more related to a delta or theta power decrease for the determination of the classifier output probability and to the alpha power increase for the speed control of the robot.
Cédric Simar, Mathieu Petieau, Anita Cebolla, Axelle Leroy, Gianluca Bontempi, Guy Cheron
IJCNN5
2019 New functionalities in the TCGAbiolinks package for the study and integration of cancer data from GDC and GTEx
abstract
The advent of Next-Generation Sequencing (NGS) technologies has opened new perspectives in deciphering the genetic mechanisms underlying complex diseases. Nowadays, the amount of genomic data is massive and substantial efforts and new tools are required to unveil the information hidden in the data. The Genomic Data Commons (GDC) Data Portal is a platform that contains different genomic studies including the ones from The Cancer Genome Atlas (TCGA) and the Therapeutically Applicable Research to Generate Effective Treatments (TARGET) initiatives, accounting for more than 40 tumor types originating from nearly 30000 patients. Such platforms, although very attractive, must make sure the stored data are easily accessible and adequately harmonized. Moreover, they have the primary focus on the data storage in a unique place, and they do not provide a comprehensive toolkit for analyses and interpretation of the data. To fulfill this urgent need, comprehensive but easily accessible computational methods for integrative analyses of genomic data that do not renounce a robust statistical and theoretical framework are required. In this context, the R/Bioconductor package TCGAbiolinks was developed, offering a variety of bioinformatics functionalities. Here we introduce new features and enhancements of TCGAbiolinks in terms of i) more accurate and flexible pipelines for differential expression analyses, ii) different methods for tumor purity estimation and filtering, iii) integration of normal samples from other platforms iv) support for other genomics datasets, exemplified here by the TARGET data. Evidence has shown that accounting for tumor purity is essential in the study of tumorigenesis, as these factors promote confounding behavior regarding differential expression analysis. With this in mind, we implemented these filtering procedures in TCGAbiolinks. Moreover, a limitation of some of the TCGA datasets is the unavailability or paucity of corresponding normal samples. We thus integrated into TCGAbiolinks the possibility to use normal samples from the Genotype-Tissue Expression (GTEx) project, which is another large-scale repository cataloging gene expression from healthy individuals. The new functionalities are available in the TCGAbiolinks version 2.8 and higher released in Bioconductor version 3.7.
Mohamed Mounir, Marta Lucchetta, Tiago Chedraoui Silva, Catharina Olsen, Gianluca Bontempi, Houtan Noushmehr, Antonio Colaprico, Elena Papaleo
PLoS Comput. Biol.5
2018 Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy
abstract
Detecting frauds in credit card transactions is perhaps one of the best testbeds for computational intelligence algorithms. In fact, this problem involves a number of relevant challenges, namely: concept drift (customers' habits evolve and fraudsters change their strategies over time), class imbalance (genuine transactions far outnumber frauds), and verification latency (only a small set of transactions are timely checked by investigators). However, the vast majority of learning algorithms that have been proposed for fraud detection rely on assumptions that hardly hold in a real-world fraud-detection system (FDS). This lack of realism concerns two main aspects: 1) the way and timing with which supervised information is provided and 2) the measures used to assess fraud-detection performance. This paper has three major contributions. First, we propose, with the help of our industrial partner, a formalization of the fraud-detection problem that realistically describes the operating conditions of FDSs that everyday analyze massive streams of credit card transactions. We also illustrate the most appropriate performance measures to be used for fraud-detection purposes. Second, we design and assess a novel learning strategy that effectively addresses class imbalance, concept drift, and verification latency. Third, in our experiments, we demonstrate the impact of class unbalance and concept drift in a real-world data stream containing more than 75 million transactions, authorized over a time window of three years.
Andrea Dal Pozzolo, Giacomo Boracchi, Olivier Caelen, Cesare Alippi, Gianluca Bontempi
IEEE Trans. Neural Networks Learn. Syst.5
2017 A Dynamic Factor Machine Learning Method for Multi-variate and Multi-step-Ahead Forecasting
abstract
Most multivariate forecasting methods in the literature are restricted to vector time series of low dimension, linear methods and short horizons. Big data revolution is instead shifting the focus to problems (e.g. issued from the IoT technology) characterized by very large dimension, nonlinearity and long forecasting horizon. This paper discusses and compares a set of state-of-the-art methods which could be promising in tackling such challenges. Also, it proposes DFML, a machine learning version of the Dynamic Factor Model (DFM), a successful forecasting methodology well-known in econometrics. The DFML strategy is based on a out-of-sample selection of the nonlinear forecaster, the number of latent components and the multi-step-ahead strategy. We will show that DFML can consistently outperform state-of-the-art methods in a number of synthetic and real forecasting tasks.
Gianluca Bontempi, Yann-Aël Le Borgne, Jacopo De Stefani
DSAA1
2017 An Assessment of Streaming Active Learning Strategies for Real-Life Credit Card Fraud Detection
abstract
Credit card fraud detection raises unique challenges due to the streaming, imbalanced, and non-stationary nature of transaction data. It additionally includes an active learning step, since the labeling (fraud or genuine) of a subset of transactions is obtained in near-real time by human investigators contacting the cardholders. These challenges and characteristics have traditionally been studied separately in the literature. In this paper, we investigate how previously proposed techniques can be combined to improve fraud detection accuracy. In particular, we highlight the existence of an exploitation/exploration tradeoff for active learning in the context of fraud detection, which has so far been overlooked in the literature. Relying on a real-world dataset of millions of transactions provided by our industrial partner Worldline, we performed an extensive experimental analysis in order to assess how traditional active learning strategies can be improved by using complementary machine learning techniques. We find that the baseline active learning strategy, denoted High Risk Querying, is a robust strategy, which can be further improved by combining it with Semi-Supervised learning.
Fabrizio Carcillo, Yann-Aël Le Borgne, Olivier Caelen, Gianluca Bontempi
DSAA4
2017 CancerSubtypes: an R/Bioconductor package for molecular cancer subtype identification, validation and visualization
abstract
SUMMARY: Identifying molecular cancer subtypes from multi-omics data is an important step in the personalized medicine. We introduce CancerSubtypes, an R package for identifying cancer subtypes using multi-omics data, including gene expression, miRNA expression and DNA methylation data. CancerSubtypes integrates four main computational methods which are highly cited for cancer subtype identification and provides a standardized framework for data pre-processing, feature selection, and result follow-up analyses, including results computing, biology validation and visualization. The input and output of each step in the framework are packaged in the same data format, making it convenience to compare different methods. The package is useful for inferring cancer subtypes from an input genomic dataset, comparing the predictions from different well-known methods and testing new subtype discovery methods, as shown with different application scenarios in the Supplementary Material. AVAILABILITY AND IMPLEMENTATION: The package is implemented in R and available under GPL-2 license from the Bioconductor website (http://bioconductor.org/packages/CancerSubtypes/). CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Taosheng Xu, Thuc Duy Le, Lin Liu 0003, Rujing Wang, Bing-Yu Sun, Antonio Colaprico, Gianluca Bontempi, Jiuyong Li
Bioinform.8
2016 How interacting pathways are regulated by miRNAs in breast cancer subtypes
abstract
BACKGROUND: An important challenge in cancer biology is to understand the complex aspects of the disease. It is increasingly evident that genes are not isolated from each other and the comprehension of how different genes are related to each other could explain biological mechanisms causing diseases. Biological pathways are important tools to reveal gene interaction and reduce the large number of genes to be studied by partitioning it into smaller paths. Furthermore, recent scientific evidence has proven that a combination of pathways, instead than a single element of the pathway or a single pathway, could be responsible for pathological changes in a cell. RESULTS: In this paper we develop a new method that can reveal miRNAs able to regulate, in a coordinated way, networks of gene pathways. We applied the method to subtypes of breast cancer. The basic idea is the identification of pathways significantly enriched with differentially expressed genes among the different breast cancer subtypes and normal tissue. Looking at the pairs of pathways that were found to be functionally related, we created a network of dependent pathways and we focused on identifying miRNAs that could act as miRNA drivers in a coordinated regulation process. CONCLUSIONS: Our approach enables miRNAs identification that could have an important role in the development of breast cancer.
Claudia Cava, Antonio Colaprico, Gloria Bertoli, Gianluca Bontempi, Giancarlo Mauri, Isabella Castiglioni
BMC Bioinform.4
2015 Credit card fraud detection and concept-drift adaptation with delayed supervised information
abstract
Most fraud-detection systems (FDSs) monitor streams of credit card transactions by means of classifiers returning alerts for the riskiest payments. Fraud detection is notably a challenging problem because of concept drift (i.e. customers' habits evolve) and class unbalance (i.e. genuine transactions far outnumber frauds). Also, FDSs differ from conventional classification because, in a first phase, only a small set of supervised samples is provided by human investigators who have time to assess only a reduced number of alerts. Labels of the vast majority of transactions are made available only several days later, when customers have possibly reported unauthorized transactions. The delay in obtaining accurate labels and the interaction between alerts and supervised information have to be carefully taken into consideration when learning in a concept-drifting environment. In this paper we address a realistic fraud-detection setting and we show that investigator's feedbacks and delayed labels have to be handled separately. We design two FDSs on the basis of an ensemble and a sliding-window approach and we show that the winning strategy consists in training two separate classifiers (on feedbacks and delayed labels, respectively), and then aggregating the outcomes. Experiments on large dataset of real-world transactions show that the alert precision, which is the primary concern of investigators, can be substantially improved by the proposed approach.
Andrea Dal Pozzolo, Giacomo Boracchi, Olivier Caelen, Cesare Alippi, Gianluca Bontempi
IJCNN5
2015 When is Undersampling Effective in Unbalanced Classification Tasks?
Andrea Dal Pozzolo, Olivier Caelen, Gianluca Bontempi
ECML/PKDD (1)3
2015 From dependency to causality: a machine learning approach
Gianluca Bontempi, Maxime Flauder
J. Mach. Learn. Res.1
2014 Human Activity Recognition Framework in Monitored Environments
abstract
This work addresses the problem of the recognition of human activities in \textit{Ambient Assisted Living (AAL)} scenarios. The ultimate goal of a good AAL system is to learn and recognize behaviours or routines of the person or people living at home, in order to help them if something unusual happens. In this paper, we explore the advances in unobstrusive depth camera-based technologies as a single sensor to detect human activities involving motion. We develop a model for learning and recognizing seven basic human actions (walk, sit down, stand up, bend down, bend up, twist left, and twist right). Our approach is composed of 5 steps ((a) body representation, (b) time series summarization, (c) posture clustering-quantization, (d) action learning with Hidden Markov Models, and (e) action recognition). The results obtained suggest that this type of sensors are accurate enough and useful to achieve high score for detection using classic stochastic models as behaviour representation and recognition.
Olmo León, Manuel P. Cuéllar, Miguel Delgado 0001, Yann-Aël Le Borgne, Gianluca Bontempi
ICPRAM5
2014 A Monte Carlo strategy for structured multiple-step-ahead time series prediction
abstract
Forecasting a time series multiple-step-ahead is a challenging problem for several reasons: the accumulation of errors, the noise, and the complexity of the dependency between past and far future which has to be inferred on the basis of a limited amount of data. Traditional approaches to multi-step-ahead forecasting reduce the problem to a series of single-output prediction tasks. This is notably the case of the Iterated and the Direct approaches. More recently, multiple-output approaches appeared and stressed the multivariate and structured nature of the output to be predicted. This paper intends to go a step further in this direction by formulating the problem of multi-step-ahed forecasting as a problem of conditional multivariate estimation which can be addressed by a Monte Carlo importance sampling strategy. The interesting aspect of the approach is that this probabilistic formulation allows a natural integration of the traditional Iterated and Direct approaches. The extensive assessment of our algorithm with the NN5, NN3 and a synthetic benchmark shows that this approach is promising and competitive with the state-of-the-art.
Gianluca Bontempi
IJCNN1
2014 Using HDDT to avoid instances propagation in unbalanced and evolving data streams
abstract
Hellinger Distance Decision Trees [10] (HDDT) has been previously used for static datasets with skewed distributions. In unbalanced data streams, state-of-the-art techniques use instance propagation and standard decision trees (e.g. C4.5 [27]) to cope with the unbalanced problem. However it is not always possible to revisit/store old instances of a stream. In this paper we show how HDDT can be successfully applied in unbalanced and evolving stream data. Using HDDT allows us to remove instance propagations between batches with several benefits: i) improved predictive accuracy ii) speed iii) single-pass through the data. We use a Hellinger weighted ensemble of HDDTs to combat concept drift and increase accuracy of single classifiers. We test our framework on several streaming datasets with unbalanced classes and concept drift.
Andrea Dal Pozzolo, Reid A. Johnson, Olivier Caelen, Serge Waterschoot, Nitesh V. Chawla, Gianluca Bontempi
IJCNN6
2014 On the Null Distribution of the Precision and Recall Curve
Miguel Lopes, Gianluca Bontempi
ECML/PKDD (2)2
2014 A comprehensive overview of Infinium HumanMethylation450 data processing
abstract
Infinium HumanMethylation450 beadarray is a popular technology to explore DNA methylomes in health and disease, and there is a current explosion in the use of this technique. Despite experience acquired from gene expression microarrays, analyzing Infinium Methylation arrays appeared more complex than initially thought and several difficulties have been encountered, as those arrays display specific features that need to be taken into consideration during data processing. Here, we review several issues that have been highlighted by the scientific community, and we present an overview of the general data processing scheme and an evaluation of the different normalization methods available to date to guide the 450K users in their analysis and data interpretation.
Sarah Dedeurwaerder, Matthieu Defrance, Martin Bizet, Emilie Calonne, Gianluca Bontempi, François Fuks
Briefings Bioinform.5
2014 Learned lessons in credit card fraud detection from a practitioner perspective
Andrea Dal Pozzolo, Olivier Caelen, Yann-Aël Le Borgne, Serge Waterschoot, Gianluca Bontempi
Expert Syst. Appl.5
2013 Stability of feature selection algorithms for classification in high-throughput genomics datasets
abstract
A major goal of the application of Machine Learning techniques to high-throughput genomics data (e.g. DNA microarrays or RNA-Seq), is the identification of “gene signatures”. These signatures can be used to discriminate among healthy or disease states (e.g. normal vs cancerous tissue) or among different biological mechanisms, at the gene expression level. Thus, the literature is plenty of studies, where numerous feature selection techniques are applied, in an effort to reduce the noise and dimensionality of such datasets. However, little attention is given to the stability of these signatures, in cases where the original dataset is perturbed by adding, removing or simply resampling the original observations. In this article, we are assessing the stability of a set of well characterized public cancer microarray datasets, using five popular feature selection algorithms in the field of high-throughput genomics data analysis.
Panagiotis Moulos, Ioannis Kanaris, Gianluca Bontempi
BIBE3
2013 A Machine Learning Approach Against a Masked AES
Liran Lerman, Stephane Fernandes Medeiros, Gianluca Bontempi, Olivier Markowitch
CARDIS3
2013 A Statistic Criterion for Reducing Indeterminacy in Linear Causal Modeling
Gianluca Bontempi
ICPRAM1
2013 Racing for Unbalanced Methods Selection
Andrea Dal Pozzolo, Olivier Caelen, Serge Waterschoot, Gianluca Bontempi
IDEAL4
2013 mRMRe: an R package for parallelized mRMR ensemble feature selection
abstract
MOTIVATION: Feature selection is one of the main challenges in analyzing high-throughput genomic data. Minimum redundancy maximum relevance (mRMR) is a particularly fast feature selection method for finding a set of both relevant and complementary features. Here we describe the mRMRe R package, in which the mRMR technique is extended by using an ensemble approach to better explore the feature space and build more robust predictors. To deal with the computational complexity of the ensemble approach, the main functions of the package are implemented and parallelized in C using the openMP Application Programming Interface. RESULTS: Our ensemble mRMR implementations outperform the classical mRMR approach in terms of prediction accuracy. They identify genes more relevant to the biological context and may lead to richer biological interpretations. The parallelized functions included in the package show significant gains in terms of run-time speed when compared with previously released packages. AVAILABILITY: The R package mRMRe is available on Comprehensive R Archive Network and is provided open source under the Artistic-2.0 License. The code used to generate all the results reported in this application note is available from Supplementary File 1. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nicolas De Jay, Simon Papillon-Cavanagh, Catharina Olsen, Nehme Hachem, Gianluca Bontempi, Benjamin Haibe-Kains
Bioinform.5
2013 Research and applications: Comparison and validation of genomic predictors for anticancer drug sensitivity
abstract
BACKGROUND: An enduring challenge in personalized medicine lies in selecting the right drug for each individual patient. While testing of drugs on patients in large trials is the only way to assess their clinical efficacy and toxicity, we dramatically lack resources to test the hundreds of drugs currently under development. Therefore the use of preclinical model systems has been intensively investigated as this approach enables response to hundreds of drugs to be tested in multiple cell lines in parallel. METHODS: Two large-scale pharmacogenomic studies recently screened multiple anticancer drugs on over 1000 cell lines. We propose to combine these datasets to build and robustly validate genomic predictors of drug response. We compared five different approaches for building predictors of increasing complexity. We assessed their performance in cross-validation and in two large validation sets, one containing the same cell lines present in the training set and another dataset composed of cell lines that have never been used during the training phase. RESULTS: Sixteen drugs were found in common between the datasets. We were able to validate multivariate predictors for three out of the 16 tested drugs, namely irinotecan, PD-0325901, and PLX4720. Moreover, we observed that response to 17-AAG, an inhibitor of Hsp90, could be efficiently predicted by the expression level of a single gene, NQO1. CONCLUSION: These results suggest that genomic predictors could be robustly validated for specific drugs. If successfully validated in patients' tumor cells, and subsequently in clinical trials, they could act as companion tests for the corresponding drugs and play an important role in personalized medicine.
Simon Papillon-Cavanagh, Nicolas De Jay, Nehme Hachem, Catharina Olsen, Gianluca Bontempi, Hugo J. W. L. Aerts, John Quackenbush, Benjamin Haibe-Kains
J. Am. Medical Informatics Assoc.5
2012 A review and comparison of strategies for multi-step ahead time series forecasting based on the NN5 forecasting competition
Souhaib Ben Taieb, Gianluca Bontempi, Amir F. Atiya, Antti Sorjamaa
Expert Syst. Appl.2
2011 Recursive Multi-step Time Series Forecasting by Perturbing Data
abstract
The Recursive strategy is the oldest and most intuitive strategy to forecast a time series multiple steps ahead. At the same time, it is well-known that this strategy suffers from the accumulation of errors as long as the forecasting horizon increases. We propose a variant of the Recursive strategy, called RECNOISY, which perturbs the initial dataset at each step of the forecasting process in order to i) handle more properly the estimated values at each forecasting step and ii) decrease the accumulation of errors induced by the Recursive strategy. In addition to the RECNOISY strategy, we propose another strategy, called HYBRID, which for each horizon selects the most accurate approach among the REC and the RECNOISY strategies according to the estimated accuracy. In order to assess the effectiveness of the proposed strategies, we carry out an experimental session based on the 111 times series of the NN5 forecasting competition. Accuracy results are presented together with a paired comparison over the horizons and the time series. The preliminary results show that our proposed approaches are promising in terms of forecasting performance.
Souhaib Ben Taieb, Gianluca Bontempi
ICDM2
2011 A Selecting-the-Best Method for Budgeted Model Selection
Gianluca Bontempi, Olivier Caelen
ECML/PKDD (1)1
2011 Multiple-input multiple-output causal strategies for gene selection
abstract
BACKGROUND: Traditional strategies for selecting variables in high dimensional classification problems aim to find sets of maximally relevant variables able to explain the target variations. If these techniques may be effective in generalization accuracy they often do not reveal direct causes. The latter is essentially related to the fact that high correlation (or relevance) does not imply causation. In this study, we show how to efficiently incorporate causal information into gene selection by moving from a single-input single-output to a multiple-input multiple-output setting. RESULTS: We show in synthetic case study that a better prioritization of causal variables can be obtained by considering a relevance score which incorporates a causal term. In addition we show, in a meta-analysis study of six publicly available breast cancer microarray datasets, that the improvement occurs also in terms of accuracy. The biological interpretation of the results confirms the potential of a causal approach to gene selection. CONCLUSIONS: Integrating causal information into gene selection algorithms is effective both in terms of prediction accuracy and biological interpretation.
Gianluca Bontempi, Benjamin Haibe-Kains, Christine Desmedt, Christos P. Sotiriou, John Quackenbush
BMC Bioinform.1
2011 Fourier spectral factor model for prediction of multidimensional signals
Abhilash Alexander Miranda, Catharina Olsen, Gianluca Bontempi
Signal Process.3
2010 Causal filter selection in microarray data
Gianluca Bontempi, Patrick E. Meyer 0001
ICML1
2010 Demonstrating principal component aggregation for distributed spatial pattern recognition
abstract
The Principal Component Aggregation has recently been proposed as a versatile distributed information extraction technique for sensor networks [3]. This demonstration illustrates its use for a network-level pattern recognition task. Four different patterns, or events, may be sensed by light measurements of a network of 27 nodes. The sens measurements are fused on the fly along a routing tree up to the base station, where the monitored pattern is recognized by a prediction algorithm.
Yann-Aël Le Borgne, Ann Nowé, Kris Steenhaut, Gianluca Bontempi
IPSN4
2010 Multiple-output modeling for multi-step-ahead time series forecasting
Souhaib Ben Taieb, Antti Sorjamaa, Gianluca Bontempi
Neurocomputing3
2009 Long-term prediction of time series by combining direct and MIMO strategies
abstract
Reliable and accurate prediction of time series over large future horizons has become the new frontier of the forecasting discipline. Current approaches to long-term time series forecasting rely either on iterated predictors, direct predictors or, more recently, on the multi-input multi-output (MIMO) predictors. The iterated approach suffers from the accumulation of errors, the Direct strategy makes a conditional independence assumption, which does not necessarily preserve the stochastic properties of the time series, while the MIMO technique is limited by the reduced flexibility of the predictor. The paper compares the direct and MIMO strategies and discusses their respective limitations to the problem of long-term time series prediction. It also proposes a new methodology that is a sort of intermediate way between the Direct and the MIMO technique. The paper presents the results obtained with the ESTSP 2007 competition dataset.
Souhaib Ben Taieb, Gianluca Bontempi, Antti Sorjamaa, Amaury Lendasse
IJCNN2
2008 A Model-Based Relevance Estimation Approach for Feature Selection in Microarray Datasets
Gianluca Bontempi, Patrick E. Meyer 0001
ICANN (2)1
2008 A comparative study of survival models for breast cancer prognostication based on microarray data: does a single gene beat them all?
abstract
MOTIVATION: Survival prediction of breast cancer (BC) patients independently of treatment, also known as prognostication, is a complex task since clinically similar breast tumors, in addition to be molecularly heterogeneous, may exhibit different clinical outcomes. In recent years, the analysis of gene expression profiles by means of sophisticated data mining tools emerged as a promising technology to bring additional insights into BC biology and to improve the quality of prognostication. The aim of this work is to assess quantitatively the accuracy of prediction obtained with state-of-the-art data analysis techniques for BC microarray data through an independent and thorough framework. RESULTS: Due to the large number of variables, the reduced amount of samples and the high degree of noise, complex prediction methods are highly exposed to performance degradation despite the use of cross-validation techniques. Our analysis shows that the most complex methods are not significantly better than the simplest one, a univariate model relying on a single proliferation gene. This result suggests that proliferation might be the most relevant biological process for BC prognostication and that the loss of interpretability deriving from the use of overcomplex methods may be not sufficiently counterbalanced by an improvement of the quality of prediction. AVAILABILITY: The comparison study is implemented in an R package called survcomp and is available from http://www.ulb.ac.be/di/map/bhaibeka/software/survcomp/.
Benjamin Haibe-Kains, Christine Desmedt, Christos P. Sotiriou, Gianluca Bontempi
Bioinform.4
2008 minet: A R/Bioconductor Package for Inferring Large Transcriptional Networks Using Mutual Information
abstract
RESULTS: This paper presents the R/Bioconductor package minet (version 1.1.6) which provides a set of functions to infer mutual information networks from a dataset. Once fed with a microarray dataset, the package returns a network where nodes denote genes, edges model statistical dependencies between genes and the weight of an edge quantifies the statistical evidence of a specific (e.g transcriptional) gene-to-gene interaction. Four different entropy estimators are made available in the package minet (empirical, Miller-Madow, Schurmann-Grassberger and shrink) as well as four different inference methods, namely relevance networks, ARACNE, CLR and MRNET. Also, the package integrates accuracy assessment tools, like F-scores, PR-curves and ROC-curves in order to compare the inferred network with a reference one. CONCLUSION: The package minet provides a series of tools for inferring transcriptional networks from microarray data. It is freely available from the Comprehensive R Archive Network (CRAN) as well as from the Bioconductor website.
Patrick E. Meyer 0001, Frédéric Lafitte, Gianluca Bontempi
BMC Bioinform.3
2008 New Routes from Minimal Approximation Error to Principal Components
Abhilash Alexander Miranda, Yann-Aël Le Borgne, Gianluca Bontempi
Neural Process. Lett.3
2007 Machine Learning Techniques for Decision Support in Anesthesia
Olivier Caelen, Gianluca Bontempi, Luc Barvais
AIME2
2007 Adaptive model selection for time series prediction in wireless sensor networks
Yann-Aël Le Borgne, Silvia Santini, Gianluca Bontempi
Signal Process.3
2007 A Blocking Strategy to Improve Gene Selection for Classification of Gene Expression Data
abstract
Because of high dimensionality, machine learning algorithms typically rely on feature selection techniques in order to perform effective classification in microarray gene expression data sets. However, the large number of features compared to the number of samples makes the task of feature selection computationally hard and prone to errors. This paper interprets feature selection as a task of stochastic optimization, where the goal is to select among an exponential number of alternative gene subsets the one expected to return the highest generalization in classification. Blocking is an experimental design strategy which produces similar experimental conditions to compare alternative stochastic configurations in order to be confident that observed differences in accuracy are due to actual differences rather than to fluctuations and noise effects. We propose an original blocking strategy for improving feature selection which aggregates in a paired way the validation outcomes of several learning algorithms to assess a gene subset and compare it to others. This is a novelty with respect to conventional wrappers, which commonly adopt a sole learning algorithm to evaluate the relevance of a given set of variables. The rationale of the approach is that, by increasing the amount of experimental conditions under which we validate a feature subset, we can lessen the problems related to the scarcity of samples and consequently come up with a better selection. The paper shows that the blocking strategy significantly improves the performance of a conventional forward selection for a set of 16 publicly available cancer expression data sets. The experiments involve six different classifiers and show that improvements take place independent of the classification algorithm used after the selection step. Two further validations based on available biological annotation support the claim that blocking strategies in feature selection may improve the accuracy and the quality of the solution. The first validation is based on retrieving PubMEd abstracts associated to the selected genes and matching them to regular expressions describing the biological phenomenon underlying the expression data sets. The biological validation that follows is based on the use of the Bioconductor package GoStats in order to perform Gene Ontology statistical analysis.
Gianluca Bontempi
IEEE ACM Trans. Comput. Biol. Bioinform.1
2006 Simulation Architecture for Data Processing Algorithms in Wireless Sensor Networks
abstract
Wireless sensor networks, by providing an unprecedented way of interacting with the physical environment, have become a hot topic for research over the last few years. As with any new technology, results from real experimentations using these networks are still scarce, as real deployments are either costly, or still unfeasible in the current state of technology. There is therefore an increasing need for simulation tools allowing the testing of different architectures, communication protocols or information processing algorithms in sensor networks. In this paper, we investigate a simulation framework for the testing of data processing in wireless sensor network applications. In a first stage, data is generated using partial differential equations, allowing the modeling of a large panel of physical phenomena. In a second stage, sensing unit operating system and network constraints are simulated using an instance of a versatile simulator to account for the platform characteristics. Insights provided by the proposed simulation frame are illustrated by a set of experiments on a heat source detection task.
Yann-Aël Le Borgne, Mehdi Moussaïd, Gianluca Bontempi
AINA (2)3
2006 Machine Learning Techniques to Enable Closed-Loop Control in Anesthesia
abstract
The growing availability of high throughput measurement devices in the operating room makes possible the collection of a huge amount of data about the state of the patient and the doctors’ practice during a surgical operation. This paper explores the possibility of extracting from these data relevant information and pertinent decision rules in order to support the daily anesthesia procedures. In particular we focus on machine learning strategies to design a closed-loop controller that, in a near future, could play the role of a decision support tool and, in a further perspective, the one of automatic pilot of the anesthesia procedure. Two strategies (direct and inverse) for learning a controller from observed data are assessed on the basis of a database of measurements collected in recent years by the ULB Erasme anaesthesiology group. The preliminary results of the learning approach applied to the regulation of hypnosis through the bispectral index (BIS) in a simulated framework appear to be promising and worthy of future investigation.
Olivier Caelen, Gianluca Bontempi, Eddy Coussaert, Luc Barvais, François Clément
CBMS2
2006 Category-Based Audience Metrics for Web Site Content Improvement Using Ontologies and Page Classification
Jean-Pierre Norguet, Benjamin Tshibasu-Kabeya, Gianluca Bontempi, Esteban Zimányi
NLDB3
2005 Structural feature selection for wrapper methods
Gianluca Bontempi
ESANN1
2004 The use of intelligent data analysis techniques for system-level design: a software estimation example
Gianluca Bontempi, Wido Kruijtzer
Soft Comput.1
2002 A Data Analysis Method for Software Performance Prediction
abstract
This paper explores the role of data analysis methods to support system-level designers in characterising the performance of embedded applications. In particular, we address the performance modelling of software applications running on an embedded microprocessor. We propose a data analysis method, which, on the basis of a parameterisation of the software functionality and the hardware architecture, is able to predict the number of execution cycles on an embedded processor. Experiments with standard computational code (sorting, mathematical computation) and with MPEG variable length decoding are presented to support this claim.
Gianluca Bontempi, Wido Kruijtzer
DATE1
2002 Data-driven techniques for direct adaptive control: the lazy and the fuzzy approaches
Edy Bertolissi, Mauro Birattari, Gianluca Bontempi, Antoine Duchâteau, Hugues Bersini
Fuzzy Sets Syst.3
2001 The local paradigm for modeling and control: from neuro-fuzzy to lazy learning
Gianluca Bontempi, Hugues Bersini, Mauro Birattari
Fuzzy Sets Syst.1
2000 Predicting stock markets in boundary conditions with local models
abstract
This paper adopts the idea of regularity in the boundaries of financial time series in order to fit forecasting models which are able to outperform random walk predictions. In particular we propose the adoption of a local learning technique, called lazy learning, in order to perform model estimation and prediction in extreme conditions. The lazy learning method is proposed to return predictions in extreme conditions of trends of the Italian stock market index. The experiments show that in boundary conditions the technique is able to outperform a random predictor and to return a significant rate of accuracy.
Gianluca Bontempi, Edy Bertolissi, Mauro Birattari
CIFEr1
2000 A multi-steap ahead prediction method based on local dynamic properties
Gianluca Bontempi, Mauro Birattari
ESANN1
1999 Local Learning for Iterated Time-Series Prediction
Gianluca Bontempi, Mauro Birattari, Hugues Bersini
ICML1
1998 Recursive Lazy Learning for Modeling and Control
Gianluca Bontempi, Mauro Birattari, Hugues Bersini
ECML1
1998 Lazy learning for control design
Gianluca Bontempi, Mauro Birattari, Hugues Bersini
ESANN1
1998 Lazy Learning Meets the Recursive Least Squares Algorithm
Mauro Birattari, Gianluca Bontempi, Hugues Bersini
NIPS2
1997 Now comes the time to defuzzify neuro-fuzzy models
Hugues Bersini, Gianluca Bontempi
Fuzzy Sets Syst.2