EDBT 2026 Demo / reviewers in the wild / expert
Alioune Ngom
dblp:62/300
· DBLP profile ↗
54ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0003-2092-2494ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 7 since 2021Artificial intelligence and machine learning · 20 · 5 first-authorSystems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Heterophily-Aware Hypergraph Neural Networks for Cell Type Prediction Using Ligand-Receptor-Informed Single-Cell RNA-Seq DataabstractAccurately predicting cell types from single-cell RNA sequencing (scRNA-seq) data requires modeling complex cellular interactions that extend beyond pairwise transcriptional similarity. Ligand—receptor-mediated signaling is a key driver of such interactions, often spanning diverse cell types and exhibiting both homophilic and heterophilic patterns. In this work, we introduce a biologically informed framework for cell type prediction based on heterophily-aware hypergraph neural networks (HGNNs), where hyperedges represent multi-cell communication events derived from curated ligand—receptor pairs. This construction enables higher-order modeling of intercellular signaling and captures the combinatorial nature of ligand—receptor communication. We evaluate nine state-of-the-art hypergraph-based models—including HGNN, HyperGCN, UniGCNII, HyperND, AllDeepSets, AllSetTransformer, ED-HNN, SheafHyperGNN, and HyperUFG—that encompass diverse message passing paradigms such as spectral convolutions, diffusion dynamics, permutationinvariant set operations, and sheaf-theoretic encoding. Experiments on six benchmark scRNA-seq datasets reveal that architectures tailored to heterophilic structure substantially outperform their homophily-oriented counterparts. Our results underscore the importance of both biologically grounded hypergraph design and heterophily-aware learning in advancing automated cell type annotation for complex tissue systems. Mahshad Hashemi, Sharjeel Mustafa, Alioune Ngom, Luis Rueda 0001 |
BIBM | 3 |
| 2025 | HeteroGraphNet: A Ligand-Receptor Informed, Heterophily-Adapted Graph Neural Network for Cell Type Prediction in scRNA-Seq DataabstractGraph Neural Networks (GNNs) have emerged as powerful tools for modeling complex relational data, yet most existing architectures assume homophily-where connected nodes share similar features-an assumption that does not hold in many biological systems. In single-cell RNA sequencing (scRNA-seq) data, intercellular communication networks often exhibit heterophily, with meaningful interactions occurring between dissimilar cell types. Moreover, conventional graph construction in this domain frequently relies on arbitrary similarity thresholds, overlooking biologically validated interaction pathways. We address these limitations with HeteroGraphNet, a heterophily-adapted GNN that incorporates ligand-receptor ($\mathbf{L}-\mathbf{R}$) interactions inferred from scRNA-seq data using LIANA to construct biologically grounded cell-cell graphs. Our model combines a bi-kernel aggregation mechanism-capable of capturing both homophilic and heterophilic signals-with cosine similarity-guided adaptive random walks that dynamically update neighborhoods during training. We further mitigate class imbalance through weighted loss functions, ensuring robust performance across underrepresented cell types. Across six scRNA-seq datasets, HeteroGraphNet consistently outperforms a multi-layer perceptron (MLP), four standard GNNs (GCN, GAT, GraphSAGE, MixHop), and two heterophily-specific GNNs (H2GCN, GBK-GNN), with especially strong gains on low-homophily graphs. These results demonstrate that incorporating ligand-receptor-informed connectivity with adaptive neighborhood exploration enables more accurate and biologically meaningful cell type prediction in heterogeneous single-cell interaction networks, offering a scalable framework for broader biological network analysis. Mahshad Hashemi, Sharjeel Mustafa, Alioune Ngom, Luis Rueda 0001 |
BIBM | 3 |
| 2024 | Toward Forensic-Friendly AI: Integrating Blockchain with Federated Learning to Enhance AI Trustworthiness
Safiia Mohammed, Alioune Ngom |
ICDF2C (1) | 2 |
| 2024 | Can large language models understand molecules?abstractPURPOSE: Large Language Models (LLMs) like Generative Pre-trained Transformer (GPT) from OpenAI and LLaMA (Large Language Model Meta AI) from Meta AI are increasingly recognized for their potential in the field of cheminformatics, particularly in understanding Simplified Molecular Input Line Entry System (SMILES), a standard method for representing chemical structures. These LLMs also have the ability to decode SMILES strings into vector representations. METHOD: We investigate the performance of GPT and LLaMA compared to pre-trained models on SMILES in embedding SMILES strings on downstream tasks, focusing on two key applications: molecular property prediction and drug-drug interaction prediction. RESULTS: We find that SMILES embeddings generated using LLaMA outperform those from GPT in both molecular property and DDI prediction tasks. Notably, LLaMA-based SMILES embeddings show results comparable to pre-trained models on SMILES in molecular prediction tasks and outperform the pre-trained models for the DDI prediction tasks. CONCLUSION: The performance of LLMs in generating SMILES embeddings shows great potential for further investigation of these models for molecular embedding. We hope our study bridges the gap between LLMs and molecular embedding, motivating additional research into the potential of LLMs in the molecular representation field. GitHub: https://github.com/sshaghayeghs/LLaMA-VS-GPT . Seyedeh Shaghayegh Sadeghi, Alan Bui, Ali Forooghi, Jianguo Lu, Alioune Ngom |
BMC Bioinform. | 5 |
| 2022 | DeePSLiM: A Deep Learning Approach to Identify Predictive Short-linear Motifs for Protein Sequence ClassificationabstractSLiMs (Short Linear Motifs) are patterns of three to 20 amino acids within proteins that are sufficient to fulfill certain functions. SLiMs play a critical role in many biological processes. Hence, with the increasing quantity of biological data, it is important to develop algorithms that can quickly find patterns in large databases of DNA, RNA and protein sequences. Previous research has been very successful at applying deep learning methods to the problems of motif detection as well as classification of biological sequences. There are, however, limitations to these approaches. Most are limited to finding motifs of a single length. In addition, most research has focused on DNA and RNA, both of which use a four-letter alphabet. A few of these have attempted to apply deep learning methods on the larger, twenty letter, alphabet of proteins. We present an enhanced deep learning model, called DeePSLiM, capable of detecting predictive SLiMs in protein sequences. The model is a shallow network that can be trained quickly on large amounts of data. The SLiMs are predictive because they can be used to classify the sequences into their respective families. In this study, first, we propose a new deep learning approach for finding predictive SLiMs in protein sequences. Then, we use these predictive SLiMs for the classification task of protein sequences to evaluate our proposed method. The model was able to reach scores of 94.5% on accuracy, precision, recall, F1-Score and Matthews-correlation coefficient, as well as 99.9% area under the receiver operator characteristic curve (AUROC). Availability: The source code, sample data, and supplementary material are available via a Github project at https://github.com/sshaghayeghs/DeePSLiM. Alexandru Filip, Seyedeh Shaghayegh Sadeghi, Alioune Ngom, Luis Rueda 0001 |
CIBCB | 3 |
| 2022 | DDIPred: Graph Convolutional Network-Based Drug-drug Interactions Prediction Using Drug Chemical Structure EmbeddingabstractA drug-drug interaction (DDI) describes a circumstance in which drugs affect the activity of each other. Drugs may interact with each other to cause side effects that are unexpected or more severe than anticipated. Drugs may also interact and oppose the results of one another, leading to one (or both) medications not having their intended effect. Most drug interactions are negligible, but some can be significantly harmful if not discovered and appropriately overseen. DDI data can be helpful for other drug-related research topics such as drug repurposing and drug-target interaction, which leads to improving the drug development process. This paper presents a new method for DDI prediction named DDIPred. It is based on drug chemical structure embedding and graph convolutional networks for predicting new DDIs. DDIPred First extracts a representation of Simplified Molecular Input Line Entry System (SMILES) strings using SELF-referencIng Embedded Strings (SELFIES) and Doc2Vec. Then the representation, along with the DDI network structure, is used to predict new DDIs. Our method achieved acceptable performance when tested on the BIOSNAP DDI network dataset while outperforming other existing methods. Availability: The source code and sample data are available via a Github project at https://github.com/sshaghayeghs/DDIPred. Seyedeh Shaghayegh Sadeghi, Alioune Ngom |
CIBCB | 2 |
| 2022 | A network-based drug repurposing method via non-negative matrix factorizationabstractMOTIVATION: Drug repurposing is a potential alternative to the traditional drug discovery process. Drug repurposing can be formulated as a recommender system that recommends novel indications for available drugs based on known drug-disease associations. This article presents a method based on non-negative matrix factorization (NMF-DR) to predict the drug-related candidate disease indications. This work proposes a recommender system-based method for drug repurposing to predict novel drug indications by integrating drug and diseases related data sources. For this purpose, this framework first integrates two types of disease similarities, the associations between drugs and diseases, and the various similarities between drugs from different views to make a heterogeneous drug-disease interaction network. Then, an improved non-negative matrix factorization-based method is proposed to complete the drug-disease adjacency matrix with predicted scores for unknown drug-disease pairs. RESULTS: The comprehensive experimental results show that NMF-DR achieves superior prediction performance when compared with several existing methods for drug-disease association prediction. AVAILABILITY AND IMPLEMENTATION: The program is available at https://github.com/sshaghayeghs/NMF-DR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Seyedeh Shaghayegh Sadeghi, Jianguo Lu, Alioune Ngom |
Bioinform. | 3 |
| 2022 | Computationally repurposing drugs for breast cancer subtypes using a network-based approachabstract'De novo' drug discovery is costly, slow, and with high risk. Repurposing known drugs for treatment of other diseases offers a fast, low-cost/risk and highly-efficient method toward development of efficacious treatments. The emergence of large-scale heterogeneous biomolecular networks, molecular, chemical and bioactivity data, and genomic and phenotypic data of pharmacological compounds is enabling the development of new area of drug repurposing called 'in silico' drug repurposing, i.e., computational drug repurposing (CDR). The aim of CDR is to discover new indications for an existing drug (drug-centric) or to identify effective drugs for a disease (disease-centric). Both drug-centric and disease-centric approaches have the common challenge of either assessing the similarity or connections between drugs and diseases. However, traditional CDR is fraught with many challenges due to the underlying complex pharmacology and biology of diseases, genes, and drugs, as well as the complexity of their associations. As such, capturing highly non-linear associations among drugs, genes, diseases by most existing CDR methods has been challenging. We propose a network-based integration approach that can best capture knowledge (and complex relationships) contained within and between drugs, genes and disease data. A network-based machine learning approach is applied thereafter by using the extracted knowledge and relationships in order to identify single and pair of approved or experimental drugs with potential therapeutic effects on different breast cancer subtypes. Indeed, further clinical analysis is needed to confirm the therapeutic effects of identified drugs on each breast cancer subtype. Forough Firoozbakht, Iman Rezaeian, Luis Rueda 0001, Alioune Ngom |
BMC Bioinform. | 4 |
| 2018 | Drug target interaction predictions using PU- Leaming under different experimental setting for four formulations namely known drug target pair prediction, drug prediction, target prediction and unknown drug target pair predictionabstractPredicting new drug target interactions experimentally through wet lab experiments is time as well as resource intensive. In general, drug-target interaction prediction problem leads to drug discovery, drug repositioning and uncovers interesting patterns in chemogenomics research. Drug and target represent heterogeneous nodes within a network of interactions. Presence of an edge between the nodes indicates a positive interaction whereas an absence suggests an unknown interaction. Classification based machine learning algorithms are heavily applied in this area of research. Classification algorithms need positive as well as negative data to yield optimized results. The major problem in this field is lack of negative data because the data that are found in the public databases are positive interaction samples. Considering unknown drug target pairs as negative data may cause severe consequences for the prediction performance. Thereby, we propose a positive un-labelled (PU) learning- based approach that uses one class support vector machine technique as the learning algorithm. The algorithm learns the positive distribution from the unified feature vector space of drugs and targets and regards unknown pairs as unlabeled instead of labelling them as negative pairs. Additionally, we use 4860 Klekota Roth fingerprint + 881 PubChem fingerprint as a high dimensional and highly discriminative feature vector representation for drugs. To represent protein features, we create a protein-motif matrix based on the sliding window score that records the probability of a motif pattern occurring within a given protein sequence. Also, we separately evaluate the prediction performance using 5-fold nested cross- validation under different experimental setting for each of the four formulations: 1) Known drug-target pair,2) Drug prediction, 3) Target prediction and 4) Unknown drug target pair. We show that our approach yields the best AUC score over previous benchmark techniques and outperforms most of the recent works based on one class classifiers and PU-based learning. Hetal Rahul Rajpura, Alioune Ngom |
CIBCB | 2 |
| 2018 | Identifying suutype specific network-Uiomarkers of breast cancer survivauilityabstractCurrent studies of breast cancer find small subsets of gene biomarkers able to accurately predict the survivability of patients. In these studies, the selected genes are not necessarily functionally related, and hence, they may not correctly indicate the molecular mechanism behind breast cancer survivability. Also, several studies have shown there is a very low overlap between the different respective biomarkers subsets for the same cancer disease. To improve the robustness of classification performance and stability of detected biomarkers, recent methods take existing knowledge on relations between genes into account in the classifier, by aggregating functionality related genes to produce discriminative gene subnetworks called network-biomarkers. In this paper, given a breast cancer dataset of patients with different subtypes, we devise a novel network-based approach by integrating protein-protein interaction network (PPI) with gene expression data (1) to identify the network-biomarkers (metagene) of breast cancer survivability and (2) to predict the survivability of breast cancer patients based on subtypes. Our method uses the concept of seed gene for identification of network-biomarkers, ADASYN to solve class imbalance and random forest to predict survivability of patients. We obtained best classification performance with distance 3 from seed gene protein where the gmean, fl-measure and accuracy are respectively 0.900, 0.800 and 90.34%. The maximum size of a network biomarkers with distance 3 is 9. Maximum 34 genes are needed to predict survivability of patients. Sheikh Jubair, Luis Rueda 0001, Alioune Ngom |
IJCNN | 3 |
| 2018 | A review on machine learning principles for multi-view biological data integrationabstractDriven by high-throughput sequencing techniques, modern genomic and clinical studies are in a strong need of integrative machine learning models for better use of vast volumes of heterogeneous information in the deep understanding of biological systems and the development of predictive models. How data from multiple sources (called multi-view data) are incorporated in a learning system is a key step for successful analysis. In this article, we provide a comprehensive review on omics and clinical data integration techniques, from a machine learning perspective, for various analyses such as prediction, clustering, dimension reduction and association. We shall show that Bayesian models are able to use prior information and model measurements with various distributions; tree-based methods can either build a tree with all features or collectively make a final decision based on trees learned from each view; kernel methods fuse the similarity matrices learned from individual views together for a final similarity matrix or learning model; network-based fusion methods are capable of inferring direct and indirect associations in a heterogeneous network; matrix factorization models have potential to learn interactions among features from different views; and a range of deep neural networks can be integrated in multi-modal learning for capturing the complex mechanism of biological systems. Yifeng Li 0001, Fang-Xiang Wu, Alioune Ngom |
Briefings Bioinform. | 3 |
| 2018 | The predictive performance of short-linear motif features in the prediction of calmodulin-binding proteinsabstractBACKGROUND: The prediction of calmodulin-binding (CaM-binding) proteins plays a very important role in the fields of biology and biochemistry, because the calmodulin protein binds and regulates a multitude of protein targets affecting different cellular processes. Computational methods that can accurately identify CaM-binding proteins and CaM-binding domains would accelerate research in calcium signaling and calmodulin function. Short-linear motifs (SLiMs), on the other hand, have been effectively used as features for analyzing protein-protein interactions, though their properties have not been utilized in the prediction of CaM-binding proteins. RESULTS: We propose a new method for the prediction of CaM-binding proteins based on both the total and average scores of known and new SLiMs in protein sequences using a new scoring method called sliding window scoring (SWS) as features for the prediction module. A dataset of 194 manually curated human CaM-binding proteins and 193 mitochondrial proteins have been obtained and used for testing the proposed model. The motif generation tool, Multiple EM for Motif Elucidation (MEME), has been used to obtain new motifs from each of the positive and negative datasets individually (the SM approach) and from the combined negative and positive datasets (the CM approach). Moreover, the wrapper criterion with random forest for feature selection (FS) has been applied followed by classification using different algorithms such as k-nearest neighbors (k-NN), support vector machines (SVM), naive Bayes (NB) and random forest (RF). CONCLUSIONS: Our proposed method shows very good prediction results and demonstrates how information contained in SLiMs is highly relevant in predicting CaM-binding proteins. Further, three new CaM-binding motifs have been computationally selected and biologically validated in this study, and which can be used for predicting CaM-binding proteins. Yixun Li, Mina Maleki, Nicholas Carruthers, Paul M. Stemmer, Alioune Ngom, Luis Rueda 0001 |
BMC Bioinform. | 5 |
| 2016 | A new feature selection approach for optimizing prediction models, applied to breast cancer subtype classificationabstractFeature selection is a useful technique in classification (and regression) problems to find the most informative features for predicting but still preserves the data generality. However, some feature subset searching methods are too exhaustive while others are too greedy. On the other hand, parameter searching is another factor to improve the prediction performance. But, if it is conducted separately after feature selection stage the classification model might not be as optimal as it should. In this study, we propose a new method, called Apriori-like Feature Selection that can overcome those drawbacks. Given a classifier and a dataset, it searches for the optimal parameters and the optimal feature subset in the combined space of features and parameters. Moreover, its greedy search behavior is controllable by running options. When applying this approach on a breast cancer dataset of five subtypes, it yielded the overall classification accuracy of more than 99% but requires only about 12 genes; a significant improvement as compared to another study. Huy Quang Pham, Alioune Ngom, Luis Rueda 0001 |
BIBM | 2 |
| 2016 | The Max-Min High-Order Dynamic Bayesian Network for Learning Gene Regulatory Networks with Time-Delayed RegulationsabstractAccurately reconstructing gene regulatory network (GRN) from gene expression data is a challenging task in systems biology. Although some progresses have been made, the performance of GRN reconstruction still has much room for improvement. Because many regulatory events are asynchronous, learning gene interactions with multiple time delays is an effective way to improve the accuracy of GRN reconstruction. Here, we propose a new approach, called Max-Min high-order dynamic Bayesian network (MMHO-DBN) by extending the Max-Min hill-climbing Bayesian network technique originally devised for learning a Bayesian network's structure from static data. Our MMHO-DBN can explicitly model the time lags between regulators and targets in an efficient manner. It first uses constraint-based ideas to limit the space of potential structures, and then applies search-and-score ideas to search for an optimal HO-DBN structure. The performance of MMHO-DBN to GRN reconstruction was evaluated using both synthetic and real gene expression time-series data. Results show that MMHO-DBN is more accurate than current time-delayed GRN learning methods, and has an intermediate computing performance. Furthermore, it is able to learn long time-delayed relationships between genes. We applied sensitivity analysis on our model to study the performance variation along different parameter settings. The result provides hints on the setting of parameters of MMHO-DBN. Yifeng Li 0001, Haifen Chen, Jie Zheng 0002, Alioune Ngom |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2015 | A new compact set of biomarkers for distinguishing among ten breast cancer subtypesabstractWorld-wide, one in nine women are diagnosed with breast cancer in their lifetime and breast cancer is the second leading cause of death among women. Accurate diagnosis of the specific subtypes of this disease is vital to ensure that the patients will have the best possible response to therapy. Using the newly proposed ten subtypes of breast cancer we hypothesized that machine learning techniques would offer many benefits for selecting the most informative biomarkers. Unlike existing gene selection approaches, we use a hierarchical classification approach that selects genes and builds the classifier concurrently. Our results support that this modified approach to gene selection yields a small subset of 82 genes that can predict each of these ten subtypes with accuracies ranging from 92% to 99%. Forough Firoozbakht, Iman Rezaeian, Alioune Ngom, Luis Rueda 0001 |
BIBM | 3 |
| 2015 | Data integration in machine learningabstractModern data generated in many fields are in a strong need of integrative machine learning models in order to better make use of heterogeneous information in decision making and knowledge discovery. How data from multiple sources are incorporated in a learning system is key step for a successful analysis. In this paper, we provide a comprehensive review on data integration techniques from a machine learning perspective. Yifeng Li 0001, Alioune Ngom |
BIBM | 2 |
| 2015 | A novel approach for finding informative genes in ten subtypes of breast cancerabstractWorld wide, one in nine women are diagnosed with breast cancer in their lifetime and breast cancer is the second leading cause of death among women. Accurate diagnosis of the specific subtypes of this disease is vital to ensure that patients will have the best possible response to therapy. One way to discriminate subtypes of breast cancer is to study those genes that differentially express across different subtypes. In this study, we use different machine learning techniques to select the most informative genes corresponding to ten subtypes of breast cancer. In particular, we propose a new bottom-up hierarchical classification approach to select the most informative genes for different subtypes, while we identify the similarity level between these subtypes. Our results support that this new approach to gene selection yields a small subset of genes that can predict each of these ten subtypes with very high accuracy. Moreover, the proposed model provides an insightful structure for further analysis of these subtypes. Forough Firoozbakht, Iman Rezaeian, Alioune Ngom, Luis Rueda 0001, Lisa A. Porter |
CIBCB | 3 |
| 2015 | Prediction of high-throughput protein-protein interactions based on protein sequence informationabstractPrediction of protein-protein interaction (PPI) is one of the most challenging problems in biology. Although great progress has been devoted to the development of methodology for predicting PPIs and PIN using machine learning methods, the problem is still far from being solved since the application of most existing methods is limited. In this study, we propose a method for PPI prediction based on amino acids differences between pairs of protein sequences. 10-fold cross-validation tests based on human PPI datasets with balanced positive-to-negative ratios indicate that it performs comparably well. Therefore, our finding suggests that amino acids differences of interacting protein pairs are relevant to the prediction of PPIs and hence provide important information on sequence-based encoding schemes. Yixun Li, Behzad Rezaei, Alioune Ngom, Luis Rueda 0001 |
CIBCB | 3 |
| 2015 | The relative vertex clustering value - a new criterion for the fast discovery of functional modules in protein interaction networksabstractBACKGROUND: Cellular processes are known to be modular and are realized by groups of proteins implicated in common biological functions. Such groups of proteins are called functional modules, and many community detection methods have been devised for their discovery from protein interaction networks (PINs) data. In current agglomerative clustering approaches, vertices with just a very few neighbors are often classified as separate clusters, which does not make sense biologically. Also, a major limitation of agglomerative techniques is that their computational efficiency do not scale well to large PINs. Finally, PIN data obtained from large scale experiments generally contain many false positives, and this makes it hard for agglomerative clustering methods to find the correct clusters, since they are known to be sensitive to noisy data. RESULTS: We propose a local similarity premetric, the relative vertex clustering value, as a new criterion allowing to decide when a node can be added to a given node's cluster and which addresses the above three issues. Based on this criterion, we introduce a novel and very fast agglomerative clustering technique, FAC-PIN, for discovering functional modules and protein complexes from a PIN data. CONCLUSIONS: Our proposed FAC-PIN algorithm is applied to nine PIN data from eight different species including the yeast PIN, and the identified functional modules are validated using Gene Ontology (GO) annotations from DAVID Bioinformatics Resources. Identified protein complexes are also validated using experimentally verified complexes. Computational results show that FAC-PIN can discover functional modules or protein complexes from PINs more accurately and more efficiently than HC-PIN and CNM, the current state-of-the-art approaches for clustering PINs in an agglomerative manner. Zina M. Ibrahim, Alioune Ngom |
BMC Bioinform. | 2 |
| 2015 | Pattern classification using a new border identification paradigm: The nearest border technique
Yifeng Li 0001, B. John Oommen, Alioune Ngom, Luis Rueda 0001 |
Neurocomputing | 3 |
| 2014 | A decomposition method for large-scale sparse coding in representation learningabstractIn representation learning, sparse representation is a parsimonious principle that a sample can be approximated by a sparse superposition of dictionary atoms. Sparse coding is the core of this technique. Since the dictionary is often redundant, the dictionary size can be very large. Many optimization methods have been proposed in the literature for sparse coding. However, the efficiency of the optimization for a tremendous number of dictionary atoms is still a bottleneck. In this paper, we propose to use decomposition method for large-scale sparse coding models. Our experimental results show that our method is very efficient. Yifeng Li 0001, Richard J. Caron, Alioune Ngom |
IJCNN | 3 |
| 2014 | Versatile sparse matrix factorization: Theory and applications
Yifeng Li 0001, Alioune Ngom |
Neurocomputing | 2 |
| 2014 | Pattern recognition in bioinformatics
Xing-Ming Zhao, Alioune Ngom, Jin-Kao Hao |
Neurocomputing | 2 |
| 2014 | Guest Editorial: Pattern Recognition in BioinformaticsabstractDevelopment and application of pattern recognition techniques in the field of bioinformatics is of utmost importance for gaining new insights about phenomena in life sciences through the analysis of biological data. In this special section, three research manuscripts in their significantly extended form were selected from the papers presented at the Eighth IAPR International Conference on Pattern Recognition in Bioinformatics (PRIB 2013), which was held in Nice, France. Elena Marchiori, Alioune Ngom, Raj Acharya |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2013 | The max-min high-order dynamic Bayesian network learning for identifying gene regulatory networks from time-series microarray dataabstractWe propose a new high-order dynamic Bayesian network (HO-DBN) learning approach, called Max-Min High-Order DBN (MMHO-DBN), for discrete time-series data. MMHO-DBN explicitly models the time lags between parents and target in an efficient manner. It extends the Max-Min Hill-Climbing Bayesian network (MMHC-BN) technique which was originally devised for learning a BN's structure from static data. Both Max-Min approaches are hybrid local learning methods which fuse concepts from both constraint-based Bayesian techniques and search-and-score Bayesian methods. The MMHO-DBN first uses constraint-based ideas to limit the space of potential structure and then applies search-and-score ideas to search for an optimal HO-DBN structure. We evaluated the ability of our MMHO-DBN approach to identify genetic regulatory networks (GRN's) from gene expression time-series data. Preliminary results on artificial and real gene expression time-series are encouraging and show that it is able to learn (long) time-delayed relationships between genes, and faster than current HO-DBN learning methods. Yifeng Li 0001, Alioune Ngom |
CIBCB | 2 |
| 2013 | Classification approach based on non-negative least squares
Yifeng Li 0001, Alioune Ngom |
Neurocomputing | 2 |
| 2013 | Nonnegative Least-Squares Methods for the Classification of High-Dimensional Biological DataabstractMicroarray data can be used to detect diseases and predict responses to therapies through classification models. However, the high dimensionality and low sample size of such data result in many computational problems such as reduced prediction accuracy and slow classification speed. In this paper, we propose a novel family of nonnegative least-squares classifiers for high-dimensional microarray gene expression and comparative genomic hybridization data. Our approaches are based on combining the advantages of using local learning, transductive learning, and ensemble learning, for better prediction performance. To study the performances of our methods, we performed computational experiments on 17 well-known data sets with diverse characteristics. We have also performed statistical comparisons with many classification techniques including the well-performing SVM approach and two related but recent methods proposed in literature. Experimental results show that our approaches are faster and achieve generally a better prediction performance over compared methods. Yifeng Li 0001, Alioune Ngom |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2012 | Fast sparse representation approaches for the classification of high-dimensional biological dataabstractClassifying genomic and proteomic data is very important to predict diseases in a very early stage and investigate signaling pathways. However, this poses many computationally challenging problems, such as curse of dimensionality, noise, redundancy and so on. The principle of sparse representation has been applied to analyzing high-dimensional biological data within the frameworks of clustering, classification, and dimension reduction approaches. However, the existing sparse representation approaches are either inefficient or have the difficulty of kernelization. In this paper, we propose fast active-set-based sparse coding approach and a dictionary learning framework for classifying high-dimensional biological data. We show that they can be easily kernelized. Experimental results show that our approaches are very efficient, and satisfactory accuracy can be obtained compared with existing approaches. Yifeng Li 0001, Alioune Ngom |
BIBM | 2 |
| 2012 | A new Kernel non-negative matrix factorization and its application in microarray data analysisabstractNon-negative factorization (NMF) has been a popular machine learning method for analyzing microarray data. Kernel approaches can capture more non-linear discriminative features than linear ones. In this paper, we propose a novel kernel NMF (KNMF) approach for feature extraction and classification of microarray data. Our approach is also generalized to kernel high-order NMF (HONMF). Extensive experiments on eight microarray datasets show that our approach generally outperforms the traditional NMF and existing KNMFs. Preliminary experiment on a high-order microarray data shows that our KHONMF is a promising approach given a suitable kernel function. Yifeng Li 0001, Alioune Ngom |
CIBCB | 2 |
| 2012 | Fast Kernel Sparse Representation Approaches for ClassificationabstractSparse representation involves two relevant procedures - sparse coding and dictionary learning. Learning a dictionary from data provides a concise knowledge representation. Learning a dictionary in a higher feature space might allow a better representation of a signal. However, it is usually computationally expensive to learn a dictionary if the numbers of training data and(or) dimensions are very large using existing algorithms. In this paper, we propose a kernel dictionary learning framework for three models. We reveal that the optimization has dimension-free and parallel properties. We devise fast active-set algorithms for this framework. We investigated their performance on classification. Experimental results show that our kernel sparse representation approaches can obtain better accuracy than their linear counterparts. Furthermore, our active-set algorithms are faster than the existing interior-point and proximal algorithms. Yifeng Li 0001, Alioune Ngom |
ICDM | 2 |
| 2012 | Supervised Dictionary Learning via Non-negative Matrix Factorization for ClassificationabstractSparse representation (SR) has been being applied as a state-of-the-art machine learning approach. Sparse representation classification (SRC1) approaches based on l1norm regularization and non-negative-least-squares (NNLS) classification approach based on non-negativity have been proposed to be powerful and robust. However, these approaches are extremely slow when the size of training samples is very large, because both of them use the whole training set as dictionary. In this paper, we briefly survey the existing SR techniques for classification, and then propose a fast approach which uses non-negative matrix factorization as supervised dictionary learning method and NNLS as non-negative sparse coding method. Experiment shows that our approach can obtain comparable accuracy with the benchmark approaches and can dramatically speed up the computation particularly in the case of large sample size and many classes. Yifeng Li 0001, Alioune Ngom |
ICMLA (1) | 2 |
| 2011 | A Novel Recursive Feature Subset Selection AlgorithmabstractUnivariate filter methods, which rank single genes according to how well they each separate the classes, are widely used for gene ranking in the field of microarray analysis of gene expression datasets. These methods rank all of the genes by considering all of the samples; however some of these samples may never be classified correctly by adding new genes and these methods keep adding redundant genes covering only some parts of the space and finally the returned subset of genes may never cover the space perfectly. In this paper we introduce a new gene subset selection approach which aims to add genes covering the space which has not been covered by already selected genes in a recursive fashion. Our approach leads to significant improvement on many different benchmark datasets. Amirali Jafarian, Alioune Ngom, Luis Rueda 0001 |
BIBE | 2 |
| 2011 | Using Qualitative Probability in Reverse-Engineering Gene Regulatory NetworksabstractThis paper demonstrates the use of qualitative probabilistic networks (QPNs) to aid Dynamic Bayesian Networks (DBNs) in the process of learning the structure of gene regulatory networks from microarray gene expression data. We present a study which shows that QPNs define monotonic relations that are capable of identifying regulatory interactions in a manner that is less susceptible to the many sources of uncertainty that surround gene expression data. Moreover, we construct a model that maps the regulatory interactions of genetic networks to QPN constructs and show its capability in providing a set of candidate regulators for target genes, which is subsequently used to establish a prior structure that the DBN learning algorithm can use and which 1) distinguishes spurious correlations from true regulations, 2) enables the discovery of sets of coregulators of target genes, and 3) results in a more efficient construction of gene regulatory networks. The model is compared to the existing literature using the known gene regulatory interactions of Drosophila Melanogaster. Zina M. Ibrahim, Alioune Ngom, Ahmed Y. Tawfik |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2010 | A dynamic qualitative probabilistic network approach for extracting gene regulatory network motifsabstractThis paper extends our work to using qualitative probability to model the naturally-occurring motifs of gene regulatory networks. Having showed in [16] that the qualitative relations defining QPN graphs exhibit a direct mapping to the naturally-occurring network motifs embedded in Gene Regulatory Networks, this work is concerned with generalizing QPN constructs to create a high-level framework from which any regulatory network motif can be derived. Experimental results using time-series data of the Saccharomyces Cerevisiae show the effectiveness of our approach in providing a more accurate description of the regulatory motifs in the Saccharomyces Cerevisiae gene regulatory network compared to our previous definitions. Zina M. Ibrahim, Alioune Ngom, Ahmed Y. Tawfik |
BIBM | 2 |
| 2010 | Non-negative matrix and tensor factorization based classification of clinical microarray gene expression dataabstractNon-negative information can benefit the analysis of microarray data. This paper investigates the classification performance of non-negative matrix factorization (NMF) over gene-sample data. We also extends it to higher-order version for classification of clinical time-series data represented by tensor. Experiments show that NMF and the higher-order NMF can achieve at least comparable prediction performance. Yifeng Li 0001, Alioune Ngom |
BIBM | 2 |
| 2010 | Alignment versus variation methods for clustering microarray time-series dataabstractIn the past few years, it has been shown that traditional clustering methods do not necessarily perform well on time-series data because of the temporal relationships involved in such data - this makes it a particularly difficult problem. In this paper, we compare two clustering methods that have been introduced recently, especially for gene expression time-series data, namely, multiple-alignment (MA) clustering and variation-based co-expression detection (VCD) clustering approaches. Both approaches are based on a transformation of the data that takes into account the temporal relationships, and have been shown to effectively detect groups of co-expressed genes. We investigate the performances of the MA and VCD approaches on two microarray time-series data sets and discuss their strengths and weaknesses. Our experiments show the superior accuracy of MA over VCD when finding groups of co-expressed genes. Numanul Subhani, Yifeng Li 0001, Alioune Ngom, Luis Rueda 0001 |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | Missing value imputation methods for gene-sample-time microarray data analysisabstractWith the recent advances in microarray technology, the expression levels of genes with respect to the samples can be monitored synchronically over a series of time points. Such three-dimensional microarray data, termed gene-sample-time microarray data or GST data for short, may contain missing values. Current microarray analysis methods require complete data sets, and thus, either each row, column or tube containing missing values must be removed from the original GST data, or these missing values must be estimated before analysis. Imputation of missing values is, however, more recommended than removal of data in order to increase the effectiveness of analysis algorithms. In this paper, we extend automated imputation methods, devised for two-dimensional microarray data, to GST data. We implemented imputation methods for GST data based on Singular Value Decomposition (3SVDimpute), K-Nearest Neighbor (3KNNimpute), and gene and sample average methods (3Aimpute), and show that methods based on KNN yield the best results with the lowest normalized root mean squared error. Yifeng Li 0001, Alioune Ngom, Luis Rueda 0001 |
CIBCB | 2 |
| 2010 | New approaches to clustering microarray time-series data using multiple expression profile alignmentabstractAn important process in functional genomic studies is clustering microarray time-series data, where genes with similar expression profiles are expected to be functionally related. Clustering microarray time-series data via pairwise alignment of piecewise linear profiles has been recently introduced. In this paper, we propose a clustering approach based on a multiple profile alignment of natural cubic spline and piecewise linear representations of gene expression profiles. We combine these multiple alignment approaches with k-means. We ran our methods on a well-known data set of pre-clustered Saccharomyces cerevisiae gene expression profiles and a data set of 3315 Pseudomonas aeruginosa expression profiles. We assessed the validity of the resulting clusters and applied a c-nearest neighbor classifier for evaluating the performance of our approaches, obtaining accuracies of 89.51% and 86.12% respectively, on Saccharomyces cerevisiae data, and 90.90% and 93.71% accuracies for cubic spline and piecewise linear respectively on Pseudomonas aeruginosa data. Numanul Subhani, Luis Rueda 0001, Alioune Ngom, Conrad J. Burden |
CIBCB | 3 |
| 2010 | Multiple gene expression profile alignment for microarray time-series data clusteringabstractMOTIVATION: Clustering gene expression data given in terms of time-series is a challenging problem that imposes its own particular constraints. Traditional clustering methods based on conventional similarity measures are not always suitable for clustering time-series data. A few methods have been proposed recently for clustering microarray time-series, which take the temporal dimension of the data into account. The inherent principle behind these methods is to either define a similarity measure appropriate for temporal expression data, or pre-process the data in such a way that the temporal relationships between and within the time-series are considered during the subsequent clustering phase. RESULTS: We introduce pairwise gene expression profile alignment, which vertically shifts two profiles in such a way that the area between their corresponding curves is minimal. Based on the pairwise alignment operation, we define a new distance function that is appropriate for time-series profiles. We also introduce a new clustering method that involves multiple expression profile alignment, which generalizes pairwise alignment to a set of profiles. Extensive experiments on well-known datasets yield encouraging results of at least 80% classification accuracy. Numanul Subhani, Luis Rueda 0001, Alioune Ngom, Conrad J. Burden |
Bioinform. | 3 |
| 2010 | Computational Intelligence in Bioinformatics
Madhu Chetty, Alioune Ngom, Elena Marchiori |
Neurocomputing | 2 |
| 2010 | Optimal decoding and minimal length for the non-unique oligonucleotide probe selection problem
Laleh Soltan Ghoraie, Robin Gras, Alioune Ngom |
Neurocomputing | 4 |
| 2010 | Selection based heuristics for the non-unique oligonucleotide probe selection problem in microarray design
Alioune Ngom, Luis Rueda 0001, Robin Gras |
Pattern Recognit. Lett. | 1 |
| 2009 | Qualitative Motif Detection in Gene Regulatory NetworksabstractThis paper motivates the use of qualitative probabilistic networks (QPNs) in conjunction with or in lieu of Bayesian Networks (BNs) for reconstructing gene regulatory networks from microarray expression data. QPNs are qualitative abstractions of Bayesian Networks that replace the conditional probability tables associated with BNs by qualitative influences, which use signs to encode how the values of variables change. We demonstrate that the qualitative influences defined by QPNs exhibit a natural mapping to naturally-occurring patterns of connections, termed network motifs, embedded in Gene Regulatory Networks and present a model that maps QPN constructs to such motifs. The contribution of this paper is that of discovering motifs by mapping their time-series experimental data to QPN influences and using the discovered motifs to aid the process of reconstructing the corresponding gene regulatory network via Dynamic Bayesian Networks (DBNs). The general aim is to compile a model that uses qualitative equivalents of Dynamic Bayesian Networks to explore gene expression networks and their regulatory mechanisms. Although this aim remains under development, the results we have obtained shows success for the discovery of regulatory motifs in Saccharomyces Cerevisiae and their effectiveness in improving the results obtained in terms of reconstruction using DBNs. Zina M. Ibrahim, Ahmed Y. Tawfik, Alioune Ngom |
BIBM | 3 |
| 2009 | Biofilm Image Segmentation Using Optimal Multi-level ThresholdingabstractA microbial biofilm is structured mainly by a protective sticky matrix of extracellular polymeric substances. Quantifying such structures is useful for microbiologists and a correct image segmentation process helps substantially reduce errors in quantification. This paper proposes an approach to segmentation of biofilm images using optimal multilevel thresholding and indices of clustering validity. A direct comparison through Rand index and a quantification process is performed in a laboratory, obtaining results similar to the quantification and segmentation done by an expert. Darío Rojas, Luis Rueda 0001, Alioune Ngom, Homero Urrutia, Gerardo Carcamo |
BIBM | 3 |
| 2009 | Surprise-Based Qualitative Probabilistic Networks
Zina M. Ibrahim, Ahmed Y. Tawfik, Alioune Ngom |
ECSQARU | 3 |
| 2008 | Non-unique oligonucleotide microarray probe selection method based on genetic algorithmsabstractIn order to accurately measure the gene expression levels in microarray experiments, it is crucial to design unique, highly specific and sensitive oligonucleotide probes for the identification of biological agents such as genes in a sample. Unique probes are hard to obtain for closely related genes such as the known strains of HIV genes. The non-unique probe selection problem is to select a probe set that is able to uniquely identify targets, in a biological sample, while containing a minimal number of probe. This is a NP-hard problem and this paper contributes the first evolutionary method for finding near minimal non-unique probe sets. When used on benchmark data sets, our approach consistently performed better than three recently published methods. We also obtained results that are at least comparable to those of the current state-of-the-art heuristic. Alioune Ngom, Robin Gras |
IEEE Congress on Evolutionary Computation | 2 |
| 2008 | Evolution strategy with greedy probe selection heuristics for the non-unique oligonucleotide probe selection problemabstractIn order to accurately measure the gene expression levels in microarray experiments, it is crucial to design unique, highly specific and highly sensitive oligonucleotide probes for the identification of biological agents such as genes in a sample. Unique probes are difficult to obtain for closely related genes such as the known strains of HIV genes. The non-unique probe selection problem is to find a smallest probe set that is able to uniquely identify targets in a biological sample. This is an NP-hard problem. We present two approaches for finding near-minimal non-unique probe sets. Each approach combines of a deterministic greedy probe selection heuristic that selects good probes, with an evolution strategy that optimizes the selected probe sets. The heuristics, guided by selection functions defined over a probe set, decide at each moment which probes are the best to be included in, or excluded from, a candidate solution. Our methods produce results that are very close to, and in many cases better than, those of the current state-of-the-art approaches for the non-unique probe selection problem, namely integer linear programming, optimal cutting-plane and genetic algorithm approaches. Alioune Ngom, Robin Gras, Luis Rueda 0001 |
CIBCB | 2 |
| 2007 | A Qualitative Hidden Markov Model for Spatio-temporal Reasoning
Zina M. Ibrahim, Ahmed Y. Tawfik, Alioune Ngom |
ECSQARU | 3 |
| 2006 | Protein Threading using Parallel Evolution StrategyabstractThe protein threading problem is the problem of determining the three-dimensional structure of a given but arbitrary protein sequence from a set of known structures of other proteins. This problem is known to be NP-hard and current computational approaches are time-consuming and data-intensive. In this paper, we propose an evolution strategy for protein threading. We also propose parallel methods for fast threading. We have obtained encouraging energy results as well as significant reduction in threading time. Alioune Ngom |
IEEE Congress on Evolutionary Computation | 2 |
| 2006 | Parallel evolution strategy on grids for the protein threading problem
Alioune Ngom |
J. Parallel Distributed Comput. | 1 |
| 2003 | On the number of multilinear partitions and the computing capacity of multiple-valued multiple-threshold perceptronsabstractWe introduce the concept of multilinear partition of a point set V/spl sub/R/sup n/ and the concept of multilinear separability of a function f:Vtwo head right arrowK={0,...,k-1}. Based on well-known relationships between linear partitions and minimal pairs, we derive formulae for the number of multilinear partitions of a point set in general position and of the set K(2). The (n,k,s)-perceptrons partition the input space V into s+1 regions with s parallel hyperplanes. We obtain results on the capacity of a single (n,k,s)-perceptron, respectively, for V subset R(n) in general position and for V=K(2). Finally, we describe a fast polynomial-time algorithm for counting the multilinear partitions of K(2). Alioune Ngom, Ivan Stojmenovic, Jovisa D. Zunic |
IEEE Trans. Neural Networks | 1 |
| 2001 | The Computing Capacity of Three-Input Multiple-Valued One-Threshold Perceptrons
Alioune Ngom, Ivan Stojmenovic, Ratko Tosic |
Neural Process. Lett. | 1 |
| 2001 | STRIP - a strip-based neural-network growth algorithm for learning multiple-valued functionsabstractWe consider the problem of synthesizing multiple-valued logic functions by neural networks. A genetic algorithm (GA) which finds the longest strip in V is a subset of K(n) is described. A strip contains points located between two parallel hyperplanes. Repeated application of GA partitions the space V into certain number of strips, each of them corresponding to a hidden unit. We construct two neural networks based on these hidden units and show that they correctly compute the given but arbitrary multiple-valued function. Preliminary experimental results are presented and discussed. Alioune Ngom, Ivan Stojmenovic, Veljko M. Milutinovic |
IEEE Trans. Neural Networks | 1 |
| 2000 | Learning with Permutably Homogeneous Multiple-Valued Multiple-Threshold Perceptrons
Alioune Ngom, Corina Reischer, Dan A. Simovici, Ivan Stojmenovic |
Neural Process. Lett. | 1 |