Farida Zehraoui

dblp:48/771 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0001-6278-1680ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 7 since 2021Artificial intelligence and machine learning · 6 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 CrossAttOmics: multiomics data integration with cross-attention
abstract
MOTIVATION: Advances in high throughput technologies enabled large access to various types of omics. Each omics provides a partial view of the underlying biological process. Integrating multiple omics layers would help have a more accurate diagnosis. However, the complexity of omics data requires approaches that can capture complex relationships. One way to accomplish this is by exploiting the known regulatory links between the different omics, which could help in constructing a better multimodal representation. RESULTS: In this article, we propose CrossAttOmics, a new deep-learning architecture based on the cross-attention mechanism for multiomics integration. Each modality is projected in a lower dimensional space with its specific encoder. Interactions between modalities with known regulatory links are computed in the feature representation space with cross-attention. The results of different experiments carried out in this article show that our model can accurately predict the types of cancer by exploiting the interactions between multiple modalities. CrossAttOmics outperforms other methods when there are few paired training examples. Our approach can be combined with attribution methods like LRP to identify which interactions are the most important. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/Sanofi-Public/CrossAttOmics and https://doi.org/10.5281/zenodo.15065928. TCGA data can be downloaded from the Genomic Data Commons Data Portal. CCLE data can be downloaded from the depmap portal.
Aurélien Beaude, Franck Augé, Farida Zehraoui, Blaise Hanczar
Bioinform.3
2025 MMnc: multi-modal interpretable representation for non-coding RNA classification and class annotation
abstract
MOTIVATION: As the biological roles and disease implications of non-coding RNAs continue to emerge, the need to thoroughly characterize previously unexplored non-coding RNAs becomes increasingly urgent. These molecules hold potential as biomarkers and therapeutic targets. However, the vast and complex nature of non-coding RNAs data presents a challenge. We introduce MMnc, an interpretable deep-learning approach designed to classify non-coding RNAs into functional groups. MMnc leverages multiple data sources-such as the sequence, secondary structure, and expression-using attention-based multi-modal data integration. This ensures the learning of meaningful representations while accounting for missing sources in some samples. RESULTS: Our findings demonstrate that MMnc achieves high classification accuracy across diverse non-coding RNA classes. The method's modular architecture allows for the consideration of multiple types of modalities, whereas other tools only consider one or two at most. MMnc is resilient to missing data, ensuring that all available information is effectively utilized. Importantly, the generated attention scores offer interpretable insights into the underlying patterns of the different non-coding RNA classes, potentially driving future non-coding RNA research and applications. AVAILABILITY AND IMPLEMENTATION: Data and source code can be found at EvryRNA.ibisc.univ-evry.fr/EvryRNA/MMnc.
Constance Creux, Farida Zehraoui, François Radvanyi, Fariza Tahi
Bioinform.2
2024 Comparison and benchmark of deep learning methods for non-coding RNA classification
abstract
The involvement of non-coding RNAs in biological processes and diseases has made the exploration of their functions crucial. Most non-coding RNAs have yet to be studied, creating the need for methods that can rapidly classify large sets of non-coding RNAs into functional groups, or classes. In recent years, the success of deep learning in various domains led to its application to non-coding RNA classification. Multiple novel architectures have been developed, but these advancements are not covered by current literature reviews. We present an exhaustive comparison of the different methods proposed in the state-of-the-art and describe their associated datasets. Moreover, the literature lacks objective benchmarks. We perform experiments to fairly evaluate the performance of various tools for non-coding RNA classification on popular datasets. The robustness of methods to non-functional sequences and sequence boundary noise is explored. We also measure computation time and CO2 emissions. With regard to these results, we assess the relevance of the different architectural choices and provide recommendations to consider in future methods.
Constance Creux, Farida Zehraoui, François Radvanyi, Fariza Tahi
PLoS Comput. Biol.2
2023 AttOmics: attention-based architecture for diagnosis and prognosis from omics data
abstract
MOTIVATION: The increasing availability of high-throughput omics data allows for considering a new medicine centered on individual patients. Precision medicine relies on exploiting these high-throughput data with machine-learning models, especially the ones based on deep-learning approaches, to improve diagnosis. Due to the high-dimensional small-sample nature of omics data, current deep-learning models end up with many parameters and have to be fitted with a limited training set. Furthermore, interactions between molecular entities inside an omics profile are not patient specific but are the same for all patients. RESULTS: In this article, we propose AttOmics, a new deep-learning architecture based on the self-attention mechanism. First, we decompose each omics profile into a set of groups, where each group contains related features. Then, by applying the self-attention mechanism to the set of groups, we can capture the different interactions specific to a patient. The results of different experiments carried out in this article show that our model can accurately predict the phenotype of a patient with fewer parameters than deep neural networks. Visualizing the attention maps can provide new insights into the essential groups for a particular phenotype. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://forge.ibisc.univ-evry.fr/abeaude/AttOmics. TCGA data can be downloaded from the Genomic Data Commons Data Portal.
Aurélien Beaude, Milad R. Vahid, Franck Augé, Farida Zehraoui, Blaise Hanczar
Bioinform.4
2022 GraphGONet: a self-explaining neural network encapsulating the Gene Ontology graph for phenotype prediction on gene expression
abstract
MOTIVATION: Medical care is becoming more and more specific to patients' needs due to the increased availability of omics data. The application to these data of sophisticated machine learning models, in particular deep learning (DL), can improve the field of precision medicine. However, their use in clinics is limited as their predictions are not accompanied by an explanation. The production of accurate and intelligible predictions can benefit from the inclusion of domain knowledge. Therefore, knowledge-based DL models appear to be a promising solution. RESULTS: In this article, we propose GraphGONet, where the Gene Ontology is encapsulated in the hidden layers of a new self-explaining neural network. Each neuron in the layers represents a biological concept, combining the gene expression profile of a patient and the information from its neighboring neurons. The experiments described in the article confirm that our model not only performs as accurately as the state-of-the-art (non-explainable ones) but also automatically produces stable and intelligible explanations composed of the biological concepts with the highest contribution. This feature allows experts to use our tool in a medical setting. AVAILABILITY AND IMPLEMENTATION: GraphGONet is freely available at https://forge.ibisc.univ-evry.fr/vbourgeais/GraphGONet.git. The microarray dataset is accessible from the ArrayExpress database under the identifier E-MTAB-3732. The TCGA datasets can be downloaded from the Genomic Data Commons (GDC) data portal. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Victoria Bourgeais, Farida Zehraoui, Blaise Hanczar
Bioinform.2
2022 Assessment of deep learning and transfer learning for cancer prediction based on gene expression data
abstract
BACKGROUND: Machine learning is now a standard tool for cancer prediction based on gene expression data. However, deep learning is still new for this task, and there is no clear consensus about its performance and utility. Few experimental works have evaluated deep neural networks and compared them with state-of-the-art machine learning. Moreover, their conclusions are not consistent. RESULTS: We extensively evaluate the deep learning approach on 22 cancer prediction tasks based on gene expression data. We measure the impact of the main hyper-parameters and compare the performances of neural networks with the state-of-the-art. We also investigate the effectiveness of several transfer learning schemes in different experimental setups. CONCLUSION: Based on our experimentations, we provide several recommendations to optimize the construction and training of a neural network model. We show that neural networks outperform the state-of-the-art methods only for very large training set size. For a small training set, we show that transfer learning is possible and may strongly improve the model performance in some cases.
Blaise Hanczar, Victoria Bourgeais, Farida Zehraoui
BMC Bioinform.3
2021 Deep GONet: self-explainable deep neural network based on Gene Ontology for phenotype prediction from gene expression data
abstract
BACKGROUND: With the rapid advancement of genomic sequencing techniques, massive production of gene expression data is becoming possible, which prompts the development of precision medicine. Deep learning is a promising approach for phenotype prediction (clinical diagnosis, prognosis, and drug response) based on gene expression profile. Existing deep learning models are usually considered as black-boxes that provide accurate predictions but are not interpretable. However, accuracy and interpretation are both essential for precision medicine. In addition, most models do not integrate the knowledge of the domain. Hence, making deep learning models interpretable for medical applications using prior biological knowledge is the main focus of this paper. RESULTS: In this paper, we propose a new self-explainable deep learning model, called Deep GONet, integrating the Gene Ontology into the hierarchical architecture of the neural network. This model is based on a fully-connected architecture constrained by the Gene Ontology annotations, such that each neuron represents a biological function. The experiments on cancer diagnosis datasets demonstrate that Deep GONet is both easily interpretable and highly performant to discriminate cancer and non-cancer samples. CONCLUSIONS: Our model provides an explanation to its predictions by identifying the most important neurons and associating them with biological functions, making the model understandable for biologists and physicians.
Victoria Bourgeais, Farida Zehraoui, Mohamed Ben Hamdoune, Blaise Hanczar
BMC Bioinform.2
2020 Biological interpretation of deep neural network for phenotype prediction based on gene expression
abstract
BACKGROUND: The use of predictive gene signatures to assist clinical decision is becoming more and more important. Deep learning has a huge potential in the prediction of phenotype from gene expression profiles. However, neural networks are viewed as black boxes, where accurate predictions are provided without any explanation. The requirements for these models to become interpretable are increasing, especially in the medical field. RESULTS: We focus on explaining the predictions of a deep neural network model built from gene expression data. The most important neurons and genes influencing the predictions are identified and linked to biological knowledge. Our experiments on cancer prediction show that: (1) deep learning approach outperforms classical machine learning methods on large training sets; (2) our approach produces interpretations more coherent with biology than the state-of-the-art based approaches; (3) we can provide a comprehensive explanation of the predictions for biologists and physicians. CONCLUSION: We propose an original approach for biological interpretation of deep learning models for phenotype prediction from gene expression data. Since the model can find relationships between the phenotype and gene expression, we may assume that there is a link between the identified genes and the phenotype. The interpretation can, therefore, lead to new biological hypotheses to be investigated by biologists.
Blaise Hanczar, Farida Zehraoui, Tina Issa, Mathieu Arles
BMC Bioinform.2
2019 A new Transparent Ensemble Method based on Deep learning
abstract
Rather than making one model and hoping this model is the best/most accurate predictor we can make, ensemble methods which improve machine learning results by combining different models. However, one of the major criticisms is their being inexplicable, since they do not provide results explanation and do not allow prior knowledge integration. With the development of the machine learning the explanation of classification results and the ability to introduce domain knowledge inside the learned model have become a necessity. In this paper, we present a novel deep ensemble method based on argumentation that combines machine learning algorithms with multi-agent system to improve classification. The idea is to extract arguments from classifiers and to combine them using argumentation in order to exploit the internal knowledge of each classifiers and provide explanation behind decisions and to allow injecting prior knowledge. The results demonstrate that our method can effectively extract high quality knowledge for ensemble classifier and improve the performance.
Naziha Sendi, Nadia Abchiche-Mimouni, Farida Zehraoui
KES3
2018 Localized Multiple Sources Self-Organizing Map
Ludovic Platon, Farida Zehraoui, Fariza Tahi
ICONIP (3)2
2018 IRSOM, a reliable identifier of ncRNAs based on supervised self-organizing maps with rejection
abstract
Motivation: Non-coding RNAs (ncRNAs) play important roles in many biological processes and are involved in many diseases. Their identification is an important task, and many tools exist in the literature for this purpose. However, almost all of them are focused on the discrimination of coding and ncRNAs without giving more biological insight. In this paper, we propose a new reliable method called IRSOM, based on a supervised Self-Organizing Map (SOM) with a rejection option, that overcomes these limitations. The rejection option in IRSOM improves the accuracy of the method and also allows identifing the ambiguous transcripts. Furthermore, with the visualization of the SOM, we analyze the rejected predictions and highlight the ambiguity of the transcripts. Results: IRSOM was tested on datasets of several species from different reigns, and shown better results compared to state-of-art. The accuracy of IRSOM is always greater than 0.95 for all the species with an average specificity of 0.98 and an average sensitivity of 0.99. Besides, IRSOM is fast (it takes around 254 s to analyze a dataset of 147 000 transcripts) and is able to handle very large datasets. Availability and implementation: IRSOM is implemented in Python and C++. It is available on our software platform EvryRNA (http://EvryRNA.ibisc.univ-evry.fr).
Ludovic Platon, Farida Zehraoui, Abdelhafid Bendahmane, Fariza Tahi
Bioinform.2
2015 Detecting time periods of differential gene expression using Gaussian processes: an application to endothelial cells exposed to radiotherapy dose fraction
abstract
MOTIVATION: Identifying the set of genes differentially expressed along time is an important task in two-sample time course experiments. Furthermore, estimating at which time periods the differential expression is present can provide additional insight into temporal gene functions. The current differential detection methods are designed to detect difference along observation time intervals or on single measurement points, warranting dense measurements along time to characterize the full temporal differential expression patterns. RESULTS: We propose a novel Bayesian likelihood ratio test to estimate the differential expression time periods. Applying the ratio test to systems of genes provides the temporal response timings and durations of gene expression to a biological condition. We introduce a novel non-stationary Gaussian process as the underlying expression model, with major improvements on model fitness on perturbation and stress experiments. The method is robust to uneven or sparse measurements along time. We assess the performance of the method on realistically simulated dataset and compare against state-of-the-art methods. We additionally apply the method to the analysis of primary human endothelial cells under an ionizing radiation stress to study the transcriptional perturbations over 283 measured genes in an attempt to better understand the role of endothelium in both normal and cancer tissues during radiotherapy. As a result, using the cascade of differential expression periods, domain literature and gene enrichment analysis, we gain insights into the dynamic response of endothelial cells to irradiation. AVAILABILITY AND IMPLEMENTATION: R package 'nsgp' is available at www.ibisc.fr/en/logiciels_arobas.
Markus Heinonen, Olivier Guipaud, Fabien Milliat, Valérie Buard, Béatrice Micheau, Georges Tarlet, Marc Benderitter, Farida Zehraoui, Florence d'Alché-Buc
Bioinform.8
2014 Towards a piRNA prediction using multiple kernel fusion and support vector machine
abstract
MOTIVATION: Piwi-interacting RNA (piRNA) is the most recently discovered and the least investigated class of Argonaute/Piwi protein-interacting small non-coding RNAs. The piRNAs are mostly known to be involved in protecting the genome from invasive transposable elements. But recent discoveries suggest their involvement in the pathophysiology of diseases, such as cancer. Their identification is therefore an important task, and computational methods are needed. However, the lack of conserved piRNA sequences and structural elements makes this identification challenging and difficult. RESULTS: In the present study, we propose a new modular and extensible machine learning method based on multiple kernels and a support vector machine (SVM) classifier for piRNA identification. Very few piRNA features are known to date. The use of a multiple kernels approach allows editing, adding or removing piRNA features that can be heterogeneous in a modular manner according to their relevance in a given species. Our algorithm is based on a combination of the previously identified features [sequence features (k-mer motifs and a uridine at the first position) and piRNAs cluster feature] and a new telomere/centromere vicinity feature. These features are heterogeneous, and the kernels allow to unify their representation. The proposed algorithm, named piRPred, gives promising results on Drosophila and Human data and outscores previously published piRNA identification algorithms. AVAILABILITY AND IMPLEMENTATION: piRPred is freely available to non-commercial users on our Web server EvryRNA http://EvryRNA.ibisc.univ-evry.fr.
Jocelyn Brayet, Farida Zehraoui, Laurence Jeanson-Leh, David Israeli, Fariza Tahi
Bioinform.2
2005 New Self-organizing Maps for Multivariate Sequences Processing
abstract
Spatio-temporal connectionist networks comprise an important class of neural models that can deal with patterns distributed in both time and space. In this article, we present new models of self-organizing maps for sequence clustering and classification. We have introduced the temporal dynamics in these maps and we have proposed several new models based on covariance matrices computation. In the first models, the inputs are modeled using its associated covariance matrix. These models, used in speaker recognition, do not take into account the order of the vectors in the sequence. To overcome this drawback, we have proposed new models, which introduce the temporal dynamics in the covariance matrix associated to the input sequences. In order to obtain a network that can learn new knowledge without forgetting the previous learned ones, we have introduced the plasticity and stability properties into one proposed temporal model using the adaptive resonance theory paradigm.
Farida Zehraoui, Younès Bennani
Int. J. Comput. Intell. Appl.1
2004 M-SOM-ART: Growing Self Organizing Map for Sequences Clustering and Classification
Farida Zehraoui, Younès Bennani
ECAI1
2004 M-SOM: matricial self organizing map for sequence clustering and classification
abstract
This work presents approaches for sequence clustering and classification. These approaches use the self organizing map "SOM". The inputs of the map are modelled in order to take into account the information and the correlation of the patterns contained in the sequences. The first approaches represent the input of the map by a representative vector or by a covariance matrix in order to take into account the correlations between the sequence components. These approaches do not take into account the temporal order in the sequences (the dynamics). The other approaches introduce the dynamics in the covariance matrix. When covariance matrices represent sequences, the SOM is modified in order to take into account the fact that the inputs are matrices. The experimentations show that our approaches are better than some other temporal self organizing maps for user Web navigation classification.
Farida Zehraoui, Younès Bennani
IJCNN1
2003 Case Base Maintenance for Improving Prediction Quality
Farida Zehraoui, Rushed Kanawati, Sylvie Salotti
ICCBR1