EDBT 2026 Demo / reviewers in the wild / expert
Blaise Hanczar
dblp:35/6861
· DBLP profile ↗
31ranked-venue papers
18as first author
10since 2021 · last 2026
0000-0002-5606-8296ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 11 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | In silico generation of gene expression profiles using diffusion modelsabstractContains fulltext : 334615.pdf (Publisher’s version ) (Open Access) Alice Lacan, Romain André, Michèle Sebag, Blaise Hanczar |
BMC Bioinform. | 4 |
| 2025 | ECGrecover: A Deep Learning Approach for Electrocardiogram Signal CompletionabstractInternational audience Alex Lence, Federica Granese, Ahmad Fall, Blaise Hanczar, Joe-Elie Salem, Jean-Daniel Zucker, Edi Prifti |
KDD (1) | 4 |
| 2025 | CrossAttOmics: multiomics data integration with cross-attentionabstractMOTIVATION: Advances in high throughput technologies enabled large access to various types of omics. Each omics provides a partial view of the underlying biological process. Integrating multiple omics layers would help have a more accurate diagnosis. However, the complexity of omics data requires approaches that can capture complex relationships. One way to accomplish this is by exploiting the known regulatory links between the different omics, which could help in constructing a better multimodal representation. RESULTS: In this article, we propose CrossAttOmics, a new deep-learning architecture based on the cross-attention mechanism for multiomics integration. Each modality is projected in a lower dimensional space with its specific encoder. Interactions between modalities with known regulatory links are computed in the feature representation space with cross-attention. The results of different experiments carried out in this article show that our model can accurately predict the types of cancer by exploiting the interactions between multiple modalities. CrossAttOmics outperforms other methods when there are few paired training examples. Our approach can be combined with attribution methods like LRP to identify which interactions are the most important. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/Sanofi-Public/CrossAttOmics and https://doi.org/10.5281/zenodo.15065928. TCGA data can be downloaded from the Genomic Data Commons Data Portal. CCLE data can be downloaded from the depmap portal. Aurélien Beaude, Franck Augé, Farida Zehraoui, Blaise Hanczar |
Bioinform. | 4 |
| 2025 | Self-supervised representation learning on gene expression dataabstractMOTIVATION: Predicting phenotypes from gene expression data is a crucial task in biomedical research, enabling insights into disease mechanisms, drug responses, and personalized medicine. Traditional machine learning and deep learning rely on supervised learning, which requires large quantities of labeled data that are costly and time-consuming to obtain in the case of gene expression data. Self-supervised learning has recently emerged as a promising approach to overcome these limitations by extracting information directly from the structure of unlabeled data. RESULTS: In this study, we investigate the application of state-of-the-art self-supervised learning methods to bulk gene expression data for phenotype prediction. We selected three self-supervised methods, based on different approaches, to assess their ability to exploit the inherent structure of the data and to generate qualitative representations which can be used for downstream predictive tasks. By using several publicly available gene expression datasets, we demonstrate how the selected methods can effectively capture complex information and improve phenotype prediction accuracy. The results obtained show that self-supervised learning methods can outperform traditional supervised models besides offering significant advantage by reducing the dependency on annotated data. We provide a comprehensive analysis of the performance of each method by highlighting their strengths and limitations. We also provide recommendations for using these methods depending on the case under study. Finally, we outline future research directions to enhance the application of self-supervised learning in the field of gene expression data analysis. This study is the first work that deals with bulk RNA-Seq data and self-supervised learning. AVAILABILITY AND IMPLEMENTATION: The code and results are available at https://github.com/kdradjat/ssrl-rnaseq. Kevin Dradjat, Massinissa Hamidi, Pierre Bartet, Blaise Hanczar |
Bioinform. | 4 |
| 2023 | AttOmics: attention-based architecture for diagnosis and prognosis from omics dataabstractMOTIVATION: The increasing availability of high-throughput omics data allows for considering a new medicine centered on individual patients. Precision medicine relies on exploiting these high-throughput data with machine-learning models, especially the ones based on deep-learning approaches, to improve diagnosis. Due to the high-dimensional small-sample nature of omics data, current deep-learning models end up with many parameters and have to be fitted with a limited training set. Furthermore, interactions between molecular entities inside an omics profile are not patient specific but are the same for all patients. RESULTS: In this article, we propose AttOmics, a new deep-learning architecture based on the self-attention mechanism. First, we decompose each omics profile into a set of groups, where each group contains related features. Then, by applying the self-attention mechanism to the set of groups, we can capture the different interactions specific to a patient. The results of different experiments carried out in this article show that our model can accurately predict the phenotype of a patient with fewer parameters than deep neural networks. Visualizing the attention maps can provide new insights into the essential groups for a particular phenotype. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://forge.ibisc.univ-evry.fr/abeaude/AttOmics. TCGA data can be downloaded from the Genomic Data Commons Data Portal. Aurélien Beaude, Milad R. Vahid, Franck Augé, Farida Zehraoui, Blaise Hanczar |
Bioinform. | 5 |
| 2023 | GAN-based data augmentation for transcriptomics: survey and comparative assessmentabstractMOTIVATION: Transcriptomics data are becoming more accessible due to high-throughput and less costly sequencing methods. However, data scarcity prevents exploiting deep learning models' full predictive power for phenotypes prediction. Artificially enhancing the training sets, namely data augmentation, is suggested as a regularization strategy. Data augmentation corresponds to label-invariant transformations of the training set (e.g. geometric transformations on images and syntax parsing on text data). Such transformations are, unfortunately, unknown in the transcriptomic field. Therefore, deep generative models such as generative adversarial networks (GANs) have been proposed to generate additional samples. In this article, we analyze GAN-based data augmentation strategies with respect to performance indicators and the classification of cancer phenotypes. RESULTS: This work highlights a significant boost in binary and multiclass classification performances due to augmentation strategies. Without augmentation, training a classifier on only 50 RNA-seq samples yields an accuracy of, respectively, 94% and 70% for binary and tissue classification. In comparison, we achieved 98% and 94% of accuracy when adding 1000 augmented samples. Richer architectures and more expensive training of the GAN return better augmentation performances and generated data quality overall. Further analysis of the generated data shows that several performance indicators are needed to assess its quality correctly. AVAILABILITY AND IMPLEMENTATION: All data used for this research are publicly available and comes from The Cancer Genome Atlas. Reproducible code is available on the GitLab repository: https://forge.ibisc.univ-evry.fr/alacan/GANs-for-transcriptomics. Alice Lacan, Michèle Sebag, Blaise Hanczar |
Bioinform. | 3 |
| 2022 | GraphGONet: a self-explaining neural network encapsulating the Gene Ontology graph for phenotype prediction on gene expressionabstractMOTIVATION: Medical care is becoming more and more specific to patients' needs due to the increased availability of omics data. The application to these data of sophisticated machine learning models, in particular deep learning (DL), can improve the field of precision medicine. However, their use in clinics is limited as their predictions are not accompanied by an explanation. The production of accurate and intelligible predictions can benefit from the inclusion of domain knowledge. Therefore, knowledge-based DL models appear to be a promising solution. RESULTS: In this article, we propose GraphGONet, where the Gene Ontology is encapsulated in the hidden layers of a new self-explaining neural network. Each neuron in the layers represents a biological concept, combining the gene expression profile of a patient and the information from its neighboring neurons. The experiments described in the article confirm that our model not only performs as accurately as the state-of-the-art (non-explainable ones) but also automatically produces stable and intelligible explanations composed of the biological concepts with the highest contribution. This feature allows experts to use our tool in a medical setting. AVAILABILITY AND IMPLEMENTATION: GraphGONet is freely available at https://forge.ibisc.univ-evry.fr/vbourgeais/GraphGONet.git. The microarray dataset is accessible from the ArrayExpress database under the identifier E-MTAB-3732. The TCGA datasets can be downloaded from the Genomic Data Commons (GDC) data portal. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Victoria Bourgeais, Farida Zehraoui, Blaise Hanczar |
Bioinform. | 3 |
| 2022 | Assessment of deep learning and transfer learning for cancer prediction based on gene expression dataabstractBACKGROUND: Machine learning is now a standard tool for cancer prediction based on gene expression data. However, deep learning is still new for this task, and there is no clear consensus about its performance and utility. Few experimental works have evaluated deep neural networks and compared them with state-of-the-art machine learning. Moreover, their conclusions are not consistent. RESULTS: We extensively evaluate the deep learning approach on 22 cancer prediction tasks based on gene expression data. We measure the impact of the main hyper-parameters and compare the performances of neural networks with the state-of-the-art. We also investigate the effectiveness of several transfer learning schemes in different experimental setups. CONCLUSION: Based on our experimentations, we provide several recommendations to optimize the construction and training of a neural network model. We show that neural networks outperform the state-of-the-art methods only for very large training set size. For a small training set, we show that transfer learning is possible and may strongly improve the model performance in some cases. Blaise Hanczar, Victoria Bourgeais, Farida Zehraoui |
BMC Bioinform. | 1 |
| 2021 | Deep GONet: self-explainable deep neural network based on Gene Ontology for phenotype prediction from gene expression dataabstractBACKGROUND: With the rapid advancement of genomic sequencing techniques, massive production of gene expression data is becoming possible, which prompts the development of precision medicine. Deep learning is a promising approach for phenotype prediction (clinical diagnosis, prognosis, and drug response) based on gene expression profile. Existing deep learning models are usually considered as black-boxes that provide accurate predictions but are not interpretable. However, accuracy and interpretation are both essential for precision medicine. In addition, most models do not integrate the knowledge of the domain. Hence, making deep learning models interpretable for medical applications using prior biological knowledge is the main focus of this paper. RESULTS: In this paper, we propose a new self-explainable deep learning model, called Deep GONet, integrating the Gene Ontology into the hierarchical architecture of the neural network. This model is based on a fully-connected architecture constrained by the Gene Ontology annotations, such that each neuron represents a biological function. The experiments on cancer diagnosis datasets demonstrate that Deep GONet is both easily interpretable and highly performant to discriminate cancer and non-cancer samples. CONCLUSIONS: Our model provides an explanation to its predictions by identifying the most important neurons and associating them with biological functions, making the model understandable for biologists and physicians. Victoria Bourgeais, Farida Zehraoui, Mohamed Ben Hamdoune, Blaise Hanczar |
BMC Bioinform. | 4 |
| 2021 | CASCARO: Cascade of classifiers for minimizing the cost of prediction
Blaise Hanczar, Avner Bar-Hen |
Pattern Recognit. Lett. | 1 |
| 2020 | Deep Learning for In-Vehicle Intrusion Detection System
Elies Gherbi, Blaise Hanczar, Jean-Christophe Janodet, Witold Klaudel |
ICONIP (4) | 2 |
| 2020 | Biological interpretation of deep neural network for phenotype prediction based on gene expressionabstractBACKGROUND: The use of predictive gene signatures to assist clinical decision is becoming more and more important. Deep learning has a huge potential in the prediction of phenotype from gene expression profiles. However, neural networks are viewed as black boxes, where accurate predictions are provided without any explanation. The requirements for these models to become interpretable are increasing, especially in the medical field. RESULTS: We focus on explaining the predictions of a deep neural network model built from gene expression data. The most important neurons and genes influencing the predictions are identified and linked to biological knowledge. Our experiments on cancer prediction show that: (1) deep learning approach outperforms classical machine learning methods on large training sets; (2) our approach produces interpretations more coherent with biology than the state-of-the-art based approaches; (3) we can provide a comprehensive explanation of the predictions for biologists and physicians. CONCLUSION: We propose an original approach for biological interpretation of deep learning models for phenotype prediction from gene expression data. Since the model can find relationships between the phenotype and gene expression, we may assume that there is a link between the identified genes and the phenotype. The interpretation can, therefore, lead to new biological hypotheses to be investigated by biologists. Blaise Hanczar, Farida Zehraoui, Tina Issa, Mathieu Arles |
BMC Bioinform. | 1 |
| 2019 | An Encoding Adversarial Network for Anomaly DetectionabstractAnomaly detection is a standard problem in Machine Learning with various applications such as health-care, predictive maintenance, and cyber-security. In such applications, the data is unbalanced: the rate of regular examples is much higher than the anomalous examples. The emergence of the Generative Adversarial Networks (GANs) has recently brought new algorithms for anomaly detection. Most of them use the generator as a proxy for the reconstruction loss. The idea is that the generator cannot reconstruct an anomaly. We develop an alternative approach for anomaly detection, based on an Encoding Adversarial Network (AnoEAN), which maps the data to a latent space (decision space), where the detection of anomalies is done directly by calculating a score. Our encoder is learned by adversarial learning, using two loss functions, the first constraining the encoder to project regular data into a Gaussian distribution and the second, to project anomalous data outside this distribution. We conduct a series of experiments on several standard bases and show that our approach outperforms the state of the art when using 10% anomalies during the learning stage, and detects unseen anomalies. Elies Gherbi, Blaise Hanczar, Jean-Christophe Janodet, Witold Klaudel |
ACML | 2 |
| 2019 | Interpretable Cascade Classifiers with AbstentionabstractIn many prediction tasks such as medical diagnostics, sequential decisions are crucial to provide optimal individual treatment. Budget in real-life applications is always limited, and it can represent any limited resource such as time, money, or side effects of medications. In this contribution, we develop a POMDP-based framework to learn cost-sensitive heterogeneous cascading systems. We provide both the theoretical support for the introduced approach and the intuition behind it. We evaluate our novel method on some standard benchmarks, and we discuss how the learned models can be interpreted by human experts. Matthieu Clertant, Nataliya Sokolovska, Yann Chevaleyre, Blaise Hanczar |
AISTATS | 4 |
| 2019 | Performance visualization spaces for classification with rejection option
Blaise Hanczar |
Pattern Recognit. | 1 |
| 2018 | Controlling and Visualizing the Precision-Recall Tradeoff for External Performance Indices
Blaise Hanczar, Mohamed Nadif |
ECML/PKDD (1) | 1 |
| 2018 | An approach to optimizing abstaining area for small sample data classification
Blaise Hanczar, Jean-Daniel Zucker |
Expert Syst. Appl. | 1 |
| 2014 | Stability of Ensemble Feature Selection on High-Dimension and Low-Sample Size Data - Influence of the Aggregation MethodabstractFeature selection is an important step when building a classifier. However, the feature selection tends to be unstable on high-dimension and small-sample size data. This instability reduces the usefulness of selected features for knowledge discovery: if the selected feature subset is not robust, domain experts can have little trust that they are relevant. A growing number of studies deal with feature selection stability. Based on the idea that ensemble methods are commonly used to improve classifiers accuracy and stability, some works focused on the stability of ensemble feature selection methods. So far, they obtained mixed results, and as far as we know no study extensively studied how the choice of the aggregation method influences the stability of ensemble feature selection. This is what we study in this preliminary work. We first present some aggregation methods, then we study the stability of ensemble feature selection based on them, on both artificial and real data, as well as the resulting classification performance. David Dernoncourt, Blaise Hanczar, Jean-Daniel Zucker |
ICPRAM | 2 |
| 2014 | Unsupervised Consensus Functions Applied to Ensemble BiclusteringabstractThe ensemble methods are very popular and can improve significantly the performance of classification and clustering algorithms. Their principle is to generate a set of different models, then aggregate them into only one. Recent works have shown that this approach can also be useful in biclustering problems.The crucial step of this approach is the consensus functions that compute the aggregation of the biclusters. We identify the main consensus functions commonly used in the clustering ensemble and show how to extend them in the biclustering context. We evaluate and analyze the performances of these consensus functions on several experiments based on both artificial and real data. Blaise Hanczar, Mohamed Nadif |
ICPRAM | 1 |
| 2014 | Combination of One-Class Support Vector Machines for Classification with Reject Option
Blaise Hanczar, Michèle Sebag |
ECML/PKDD (1) | 1 |
| 2013 | Precision-recall space to correct external indices for biclusteringabstractBiclustering is a major tool of data mining in many domains and many algorithms have emerged in recent years. All these algorithms aim to obtain coherent biclusters and it is crucial to have a reliable procedure for their validation. We point out the problem of size bias in biclustering evaluation and show how it can lead to wrong conclusions in a comparative study. We present the theoretical corrections for all of the most popular measures in order to remove this bias. We introduce the corrected precision-recall space that combines the advantages of corrected measures, the ease of interpretation and visualization of uncorrected measures. Numerical experiments demonstrate the interest of our approach. Blaise Hanczar, Mohamed Nadif |
ICML (2) | 1 |
| 2013 | The reliability of estimated confidence intervals for classification error rates when only a single sample is available
Blaise Hanczar, Edward R. Dougherty |
Pattern Recognit. | 1 |
| 2012 | Experimental analysis of feature selection stability for high-dimension and low-sample size gene expression classification taskabstractGene selection is a crucial step when building a classifier from microarray or metagenomic data. As the number of observations is small, the gene selection tends to be unstable. It is common that two gene subsets, obtained from different datasets but dealing with the same classification problem, do not overlap significantly. Although it is a crucial problem, few works have been done on the selection stability. In this paper, we first present some stability quantification methods, then we study the variations of those measures with various parameters (dimensionality, sample size, feature distribution, selection threshold) on both artificial and real data, as well as the resulting classification performance. Feature selection was performed with t-test and classification with linear discriminant analysis. We point out a strong empiric correlation between the dimensionality/sample size ratio and selection instability. David Dernoncourt, Blaise Hanczar, Jean-Daniel Zucker |
BIBE | 2 |
| 2012 | Ensemble methods for biclustering tasks
Blaise Hanczar, Mohamed Nadif |
Pattern Recognit. | 1 |
| 2012 | A New Measure of Classifier Performance for Gene Expression DataabstractOne of the major aims of many microarray experiments is to build discriminatory diagnosis and prognosis models. A large number of supervised methods have been proposed in literature for microarray-based classification for this purpose. Model evaluation and comparison is a critical issue and, the most of the time, is based on the classification cost. This classification cost is based on the costs of false positives and false negative, that are generally unknown in diagnostics problems. This uncertainty may highly impact the evaluation and comparison of the classifiers. We propose a new measure of classifier performance that takes account of the uncertainty of the error. We represent the available knowledge about the costs by a distribution function defined on the ratio of the costs. The performance of a classifier is therefore computed over the set of all possible costs weighted by their probability distribution. Our method is tested on both artificial and real microarray data sets. We show that the performance of classifiers is very depending of the ratio of the classification costs. In many cases, the best classifier can be identified by our new measure whereas the classic error measures fail. Blaise Hanczar, Avner Bar-Hen |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | Using the bagging approach for biclustering of gene expression data
Blaise Hanczar, Mohamed Nadif |
Neurocomputing | 1 |
| 2010 | Bagged Biclustering for Microarray DataabstractOne of the major tools of transcriptomics is the biclustering that simultaneously constructs a partition of both examples and genes. Several methods have been proposed for microarray data analysis that enables to identify groups of genes with similar expression profiles only under a subset of examples. We propose to improve the quality of these biclustering methods by using an ensemble approach. Our bagged biclustering method generates a collection of biclusters using the bootstrap samples of the original data and aggregate them into new biclusters. Our method improve the performance of biclustering on artificial and real datasets. Blaise Hanczar, Mohamed Nadif |
ECAI | 1 |
| 2010 | Bagging for Biclustering: Application to Microarray Data
Blaise Hanczar, Mohamed Nadif |
ECML/PKDD (1) | 1 |
| 2010 | Small-sample precision of ROC-related estimatesabstractMOTIVATION: The receiver operator characteristic (ROC) curves are commonly used in biomedical applications to judge the performance of a discriminant across varying decision thresholds. The estimated ROC curve depends on the true positive rate (TPR) and false positive rate (FPR), with the key metric being the area under the curve (AUC). With small samples these rates need to be estimated from the training data, so a natural question arises: How well do the estimates of the AUC, TPR and FPR compare with the true metrics? RESULTS: Through a simulation study using data models and analysis of real microarray data, we show that (i) for small samples the root mean square differences of the estimated and true metrics are considerable; (ii) even for large samples, there is only weak correlation between the true and estimated metrics; and (iii) generally, there is weak regression of the true metric on the estimated metric. For classification rules, we consider linear discriminant analysis, linear support vector machine (SVM) and radial basis function SVM. For error estimation, we consider resubstitution, three kinds of cross-validation and bootstrap. Using resampling, we show the unreliability of some published ROC results. AVAILABILITY: Companion web site at http://compbio.tgen.org/paper_supp/ROC/roc.html CONTACT: [email protected]. Blaise Hanczar, Jianping Hua, Chao Sima, John N. Weinstein, Michael L. Bittner, Edward R. Dougherty |
Bioinform. | 1 |
| 2008 | Classification with reject option in gene expression dataabstractMOTIVATION: The classification methods typically used in bioinformatics classify all examples, even if the classification is ambiguous, for instance, when the example is close to the separating hyperplane in linear classification. For medical applications, it may be better to classify an example only when there is a sufficiently high degree of accuracy, rather than classify all examples with decent accuracy. Moreover, when all examples are classified, the classification rule has no control over the accuracy of the classifier; the algorithm just aims to produce a classifier with the smallest error rate possible. In our approach, we fix the accuracy of the classifier and thereby choose a desired risk of error. RESULTS: Our method consists of defining a rejection region in the feature space. This region contains the examples for which classification is ambiguous. These are rejected by the classifier. The accuracy of the classifier becomes a user-defined parameter of the classification rule. The task of the classification rule is to minimize the rejection region with the constraint that the error rate of the classifier be bounded by the chosen target error. This approach is also used in the feature-selection step. The results computed on both synthetic and real data show that classifier accuracy is significantly improved. AVAILABILITY: Companion Website. http://gsp.tamu.edu/Publications/rejectoption/ Blaise Hanczar, Edward R. Dougherty |
Bioinform. | 1 |
| 2007 | Feature construction from synergic pairs to improve microarray-based classificationabstractMOTIVATION: Microarray experiments that allow simultaneous expression profiling of thousands of genes in various conditions (tissues, cells or time) generate data whose analysis raises difficult problems. In particular, there is a vast disproportion between the number of attributes (tens of thousands) and the number of examples (several tens). Dimension reduction is therefore a key step before applying classification approaches. Many methods have been proposed to this purpose, but only a few of them considered a direct quantification of transcriptional interactions. We describe and experimentally validate a new dimension reduction and feature construction method, which assesses interactions between expression profiles to improve microarray-based classification accuracy. RESULTS: Our approach relies on a mutual information measure that exposes some elementary constituents of the information contained in a pair of gene expression profiles. We show that their analysis implies a term that represents the information of the interaction between the two genes. The principle of our method, called FeatKNN, is to exploit the information provided by highly synergic gene pairs to improve classification accuracy. First, a heuristic search selects the most informative gene pairs. Then, for each selected pair, a new feature, representing the classification margin of a KNN classifier in the gene pairs space, is constructed. We show experimentally that the interactional information has a degree of significance comparable to that of the gene expression profiles considered separately. Our method has been tested with different classifiers and yielded significant improvements in accuracy on several public microarray databases. Moreover, a synthetic assessment of the biological significance of the concept of synergic gene pairs suggested its ability to uncover relevant mechanisms underlying interactions among various cellular processes. Blaise Hanczar, Jean-Daniel Zucker, Corneliu Henegar, Lorenza Saitta |
Bioinform. | 1 |