EDBT 2026 Demo / reviewers in the wild / expert
Abdoulaye Baniré Diallo
dblp:11/1478
· DBLP profile ↗
32ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0002-1168-9371ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Trustworthy Survival Analysis Modeling
Armand Bandiang Massoua, Mohamed Bouguessa, Abdoulaye Baniré Diallo |
COMPSAC | 3 |
| 2026 | A Unified Perspective for Learning Graph Representations Across Multi-Level AbstractionsabstractGraph Self-Supervised Learning (GSSL) has emerged as a powerful paradigm for generating high-quality representations for graph-structured data. While multi-scale graph contrastive learning has received increasing attention, many existing methods still predominantly focus on a single graph abstraction level. To address this limitation, we propose a unified contrastive framework that can target node-level, proximity-level, cluster-level, and graph-level information and integrate them through a linear combination of similarity scores on positive pairs and dissimilarity scores (i.e., similarity scores on negative pairs). Furthermore, current approaches typically assign uniform penalty strengths to all examples, which reduces optimization flexibility and leads to ambiguous convergence status. To overcome this, we introduce a novel parameter-free fine-grained self-weighting mechanism that adaptively assigns weights to individual similarity and dissimilarity scores. The proposed mechanism emphasizes the scores that deviate significantly from their target values. Our approach not only enhances optimization flexibility but also eliminates the computational overhead of hyperparameter tuning in conventional multi-task GSSL methods. Comprehensive experiments on real-world datasets show that our methods consistently outperform state-of-the-art approaches across downstream tasks, including classification, clustering, and link prediction, in both single-level and multi-level scenarios. Mohamed Mahmoud Amar, Nairouz Mrabah, Mohamed Bouguessa, Abdoulaye Baniré Diallo |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Generative Adversarial Imputation Networks (GAIN) to Handle Missing Clinical Data for IVF Success PredictionabstractMissing data are a common challenge for machine learning (ML) approaches that provide clinical decision support, especially for complex medical outcomes such as In Vitro Fertilization (IVF) treatment prediction. Imputation methods have been studied to handle missing data, and shown to provide performance improvement of prediction models. Recently, Generative Adversarial Imputation Networks (GAIN) have been successfully applied to handle missing values in clinical data. This study implements and evaluates four imputation methods: a statistical, single imputation method and three ML-based imputation approaches, being KNN, MissForest, and GAIN. Imputation was applied to the full dataset, and the resulting complete data was then used to construct three task-specific datasets: blastocyst, pregnancy, and birth. Datasets are built based on a IVF treatment database of over 300 patients, and composed of over 50 discriminative features. Prediction models are built based on a 10-fold cross validation split and five classification algorithms: MLP, Random Forest, XGBoost, Gradient Boosting, and Bayesian Logistic. After applying imputation methods, feature value distributions remained largely consistent when compared to the distributions prior to imputation. Results suggest that GAIN imputation leads to effective convergence during model training, as both generator and discriminator losses decrease considerably. GAIN imputation outperformed all other ML-based imputation methods for the three prediction tasks. While single imputation yield F1 metrics for blastocyst, pregnancy, and birth of$0.50,0.56,0.13$, GAIN yield$0.74,0.75,0.76$, respectively for the negative class; for the positive class single imputation yield F1 metrics of$0.67,0.45,0.81$, while GAIN yield$0.76,0.51,0.85$respectively. Improvement of classification model performances after imputation suggest that a generative imputation approach might be beneficial to the task at hand. Mamadou Maladho Barry, Hayda Almeida, Debbie Montjean, Moncef Benkhalifa, Pierre Miron, Abdoulaye Baniré Diallo |
BIBM | 6 |
| 2025 | Interpreting Blastocyst Prediction Models via Region-Specific Attention in Oocyte ImagesabstractDespite increasing demand for fertility treatments, in vitro fertilization (IVF) success rates remain largely unchanged. AI-assisted oocyte assessment could improve outcomes, but high-performing models often lack interpretability. Here, we quantify the anatomical regions influencing blastocyst development by combining semantic segmentation with region-specific attention analysis. A nnUNet v2 model identified five structures in metaphase II oocyte images: ooplasm (OO), zona pellucida (ZP), perivitelline space (PVS), polar body (PB), and camera field of view (FOV). Grad-CAM heatmaps from a blastulation prediction model were aligned with segmentation masks to compute mean regional activation (MRA) and activation share. OO and PVS had the highest MRA (0.61 and 0.63), while FOV and background showed minimal activation (0.15 and 0.09). Region Activation Ratios (RAR) highlighted the outsized contributions of PVS (2.04) and$O O(2.00)$relative to their size, versus low background contribution (0.30). These results demonstrate biologically coherent model behavior and provide a framework for interpreting predictions in anatomical terms. Jie Ying Huang, Debbie Montjean, Armand Bandiang Massoua, Çisem Limandal, Abdoulaye Baniré Diallo, Moncef Benkhalifa, Pierre Miron |
BIBM | 5 |
| 2025 | Survival Structured Probabilistic Coding: Uncertainty-Aware Representation Learning for Time-to-Event PredictionabstractSurvival analysis is critical in many real-world do-mains where predicting time-to-event outcomes under censoring is essential. While deep learning has advanced survival modeling, existing methods struggle to jointly optimize discrimination, cali-bration, and uncertainty quantification. Probabilistic approaches like variational autoencoders offer uncertainty estimates but suffer from information loss due to encoder-decoder architectures and lack principled treatment of censored data. We propose Survival Structured Probabilistic Coding (Survival-SPC), a novel framework with two key mechanisms. First, an encoder-only architecture directly encodes features to probabilistic hazard representations, reducing information loss while preserving uncer-tainty. Second, censoring-aware structured regularization lever-ages partial information from censored observations to encourage diversity in latent representations. Unlike previous approaches, Survival-SPC enables efficient end-to-end training of calibrated survival distributions. Comprehensive experiments on six real-world datasets demonstrate superior performance, improving concordance index by up to 4.2% over strongest baselines while providing well-calibrated estimates. The method shows particular advantages under limited data and high censoring, establishing a robust solution for clinical applications. Code is available at Suvival-SPC Repository. Armand Bandiang Massoua, Mohamed Bouguessa, Abdoulaye Baniré Diallo |
BIBM | 3 |
| 2025 | Multi-objective Multi-Attribute Client Selection for Sustainable Over-The-Air Federated LearningabstractOver-the-air federated learning (OTA-FL) is a communication-efficient paradigm that leverages the superposition property of wireless channels to aggregate client updates simultaneously, significantly reducing uplink latency and bandwidth usage. While OTA-FL offers advantages in scalability and speed, it poses challenges in energy efficiency and delay management. This paper proposes a multi-attribute client selection framework that addresses these challenges through a multi-objective optimization approach. We analytically model selection attributes: energy efficiency, communication delay, loss, and fairness, and formulate three optimization problems to capture different trade-offs. To solve them, we employ the Multi-Objective Grey Wolf Optimizer (MOGWO), a nature-inspired metaheuristic algorithm that effectively balances exploration and exploitation. Experiments on MNIST, Fashion MNIST, and CIFAR-10 demonstrate that our approach outperforms baseline and loss-aware methods, achieving up to 13% energy savings while improving model accuracy, fairness, and reliability. Maryam Ben Driss, Essaid Sabir, Halima Elbiaze, Abdoulaye Baniré Diallo, Mohammed Sadik |
GLOBECOM | 4 |
| 2025 | IRSA Over Spreading Factors for Spatio-Temporal SIC in Scalable LoRaWAN IoT NetworksabstractThe rapid growth of the Internet of Things (IoT) has triggered the need for scalable and energy-efficient communication solutions. While LoRaWAN is widely used for long-range wireless access, its Aloha-based MAC protocol struggles with high collision rates in dense networks. Existing solutions such as irregular repetition slotted ALOHA (IRSA) and contention resolution diversity slotted ALOHA (CRDSA) have improved network performance by using packet repetitions and successive interference cancellation. However, they do not fully leverage the unique properties of LoRaWAN Spreading Factors (SFs). To address this gap, we propose a new approach called SF-IRSA, where IoT devices transmit replicas using different SFs, enabling the decoder to apply an SF-IRSA-SIC process that leverages both temporal and spatial dimensions for efficient packet decoding. Our theoretical analysis and simulations show that SF-IRSA outperforms IRSA and CRDSA in terms of throughput and reliability. Specifically, using up to two SFs results in a 16.2% increase in the asymptotic throughput compared to standard IRSA. When extending to three SFs, the throughput gain reaches 116.9%, with a maximum of $\mathbf{2 2 4. 5 2 \%}$ while using $\mathbf{6}$ SFs. Nadjib Benserir, Yaya Etiabi, Essaid Sabir, El Mehdi Amhoud, Halima Elbiaze, Abdoulaye Baniré Diallo |
ISCC | 6 |
| 2025 | Fast & Energy Efficient Federated Learning Using Multi-Attribute Client Clustering and SelectionabstractFederated Learning (FL) presents a promising paradigm for decentralized model training; however, its real-world adoption is hindered by several critical challenges, including non-independent and identically distributed (non-IID) data across clients, heterogeneous computational capabilities, and significant communication overhead. To address these issues, this paper introduces a novel multi-attribute client clustering and selection framework for FL. The proposed approach groups clients according to data distribution, device capabilities, geographic location, and model update behavior. Within each cluster, an adaptive client selection mechanism leverages dynamic attributes such as residual energy, data freshness, and client participation motivation to identify the most suitable participants. Experimental evaluations on standard FL benchmark datasets demonstrate that the proposed framework achieves faster convergence, higher global model accuracy, and improved energy efficiency compared to state-of-the-art approaches. Maryam Ben Driss, Essaid Sabir, Halima Elbiaze, Abdoulaye Baniré Diallo |
VTC2025-Spring | 4 |
| 2025 | In silico framework for genome analysis
M. Saqib Nawaz, Muhammad Zohaib Nawaz, Yongshun Gong, Philippe Fournier-Viger, Abdoulaye Baniré Diallo |
Future Gener. Comput. Syst. | 5 |
| 2024 | MLCDG: Multi-Level Contrastive Graph Clustering in Dynamic Graphs
Mohamed Mahmoud Amar, Mohamed Bouguessa, Abdoulaye Baniré Diallo |
ASONAM (3) | 3 |
| 2024 | Ensemble learning for heterogeneous biomarker discovery in precision dairy farmingabstractMetabolic biomarkers can act as powerful indicators of dairy cow welfare. Machine learning methods have been applied to discover discriminative biomarkers that can help predict potential risks to cow health, and consequently support early interventions to avoid decline in animal welfare and production loss. Previous studies on predictive models based on metabolic profiling have considered a limited scope of biomarkers, and have mostly been restricted to prediction of disease occurrence. This work proposes an ensemble supervised learning approach based on heterogeneous biomarker attributes obtained from dairy cow profile and history, and two metabolic profiling methods, to predict potential risks for dairy cow health, reproduction performance, and milk production loss. Best performing models are composed of either Random Forest, Multilayer Perceptropn and Extra Trees classifiers, and achieved F1 scores of 0.81, 0.86, 0.98 and 0.75 when predicting ‘Non-diseased’, ‘Bad’ or ‘Good’ reproduction performance, and ‘0’ production loss for an upcoming lactation. The source code and sample datasets for this work are made publicly available at https://github.com/bioinfoUQAM/dairy_biomarkers. Hayda Almeida, Nicolas Barbeau-Grégoire, Maxime Leduc, Younes Chofi, Jocelyn Dubuc, Abdoulaye Baniré Diallo |
BIBM | 8 |
| 2024 | Towards Robust Time-to-Event Prediction: Integrating the Variational Information Bottleneck with Neural Survival ModelabstractSurvival analysis aims to predict the time until a specific event of interest occurs. Although neural network-based survival models perform well in extracting rich feature embeddings and outperform traditional models, they are susceptible to the intricacies of noise present in real-world data. This noise can cause these models to miss crucial information for event-time prediction while introducing irrelevant information into the feature embeddings. Furthermore, models may struggle to distinguish between relevant and irrelevant information in data-limited regimes, such as healthcare. This can lead to overfitting, resulting from spurious correlations between irrelevant information and survival outcomes. To address these problems, we introduce the Variational Information Bottleneck (VIB) regularization approach. VIB is designed to meticulously filter out both irrelevant and redundant information, resulting in more robust feature embeddings for event-time prediction. We conducted detailed experiments on several real-world survival datasets. Our approach outperforms state-of-the-art methods in event-time prediction in various evaluation metrics. Furthermore, evaluations on semi-synthetic noisy dataset demonstrate the superior noise resistance of our approach, showcasing improved generalization and robustness. Armand Bandiang Massoua, Abdoulaye Baniré Diallo, Mohamed Bouguessa |
IJCNN | 2 |
| 2023 | Exploring the Interaction between Local and Global Latent Configurations for Clustering Single-Cell RNA-Seq: A Unified PerspectiveabstractThe most recent approaches for clustering single-cell RNA-sequencing data rely on deep auto-encoders. However, three major challenges remain unaddressed. First, current models overlook the impact of the cumulative errors induced by the pseudo-supervised embedding clustering task (Feature Randomness). Second, existing methods neglect the effect of the strong competition between embedding clustering and reconstruction (Feature Drift). Third, the previous deep clustering models regularly fail to consider the topological information of the latent data, even though the local and global latent configurations can bring complementary views to the clustering task. To address these challenges, we propose a novel approach that explores the interaction between local and global latent configurations to progressively adjust the reconstruction and embedding clustering tasks. We elaborate a topological and probabilistic filter to mitigate Feature Randomness and a cell-cell graph structure and content correction mechanism to counteract Feature Drift. The Zero-Inflated Negative Binomial model is also integrated to capture the characteristics of gene expression profiles. We conduct detailed experiments on real-world datasets from multiple representative genome sequencing platforms. Our approach outperforms the state-of-the-art clustering methods in various evaluation metrics. Nairouz Mrabah, Mohamed Mahmoud Amar, Mohamed Bouguessa, Abdoulaye Baniré Diallo |
AAAI | 4 |
| 2023 | Toward Convex Manifolds: A Geometric Perspective for Deep Graph Clustering of Single-cell RNA-seq DataabstractThe deep clustering paradigm has shown great potential for discovering complex patterns that can reveal cell heterogeneity in single-cell RNA sequencing data. This paradigm involves two training phases: pretraining based on a pretext task and fine-tuning using pseudo-labels. Although current models yield promising results, they overlook the geometric distortions that regularly occur during the training process. More precisely, the transition between the two phases results in a coarse flattening of the latent structures, which can deteriorate the clustering performance. In this context, existing methods perform euclidean-based embedding clustering without ensuring the flatness and convexity of the latent manifolds. To address this problem, we incorporate two mechanisms. First, we introduce an overclustering loss to flatten the local curves. Second, we propose an adversarial mechanism to adjust the global geometric configuration. The second mechanism gradually transforms the latent structures into convex ones. Empirical results on a variety of gene expression datasets show that our model outperforms state-of-the-art methods. Nairouz Mrabah, Mohamed Mahmoud Amar, Mohamed Bouguessa, Abdoulaye Baniré Diallo |
IJCAI | 4 |
| 2022 | A Model for the Prediction of Lifetime Profit Estimate of Dairy Cattle (Student Abstract)abstractIn livestock management, the decision of animal replacement requires an estimation of the lifetime profit of the animal based on multiple factors and operational conditions. In Dairy farms, this can be associated with the profit corresponding to milk production, health condition and herd management costs, which in turn may be a function of other factors including genetics and weather conditions. Estimating the profit of a cow can be expressed as a spatio-temporal problem where knowing the first batch of production (early-profit) can allow to predict the future batch of productions (late-profit). This problem can be addressed either by a univariate or multivariate time series forecasting. Several approaches have been designed for time series forecasting including Auto-Regressive approaches, Recurrent Neural Network including Long Short Term Memory (LSTM) method and a very deep stack of fully-connected layers. In this paper, we proposed a LSTM based approach coupled with attention and linear layers to better capture the dairy features. We compare the model, with three other architectures including NBEATs, ARIMA, MUMU-RNN using dairy production of 292181 dairy cows. The results highlight the performence of the proposed model of the compared architectures. They also show that a univariate NBEATs could perform better than the multi-variate approach there are compared to. We also highlight that such architecture could allow to predict late-profit with an error less than 3$ per month, opening the way of better resource management in the dairy industry. Vahid Naghashi, Abdoulaye Baniré Diallo |
AAAI | 2 |
| 2022 | KANALYZER: a method to identify variations of discriminative k-mers in genomic sequencesabstractDiscriminative k-mers are unique genomic regions that characterize a given viral family, genus, species, or variant. Most existing algorithms for identifying discriminative k-mer sets are limited to returning raw sub-sequences. However, to explain the discriminative properties of a given k-mer for specific taxonomic groups of viruses, it is important to identify the variations (nucleotide sequences derived from an initial k-mer having undergone one or more nucleotide changes) of this k-mer that occur in other groups of viruses. These variations as well as their frequencies of occurrence, their genomic location and their potential influence on biological functions r epresent important insights to understand the classification process. In this article, we introduce KANALYZER, a novel algorithm to identify variations of discriminative k-mers and associated information according to viral taxonomy. The algorithm was assessed to identify k-mer variations in both simulated and real viral sequence sets. In these evaluations, KANALYZER correctly and quickly identified over 95% of the variations and associated information. KANALYZER algorithm is integrated directly into CASTOR-KRFE discriminative k-mers identification tool pipeline. The source code, detailed results and data to reproduce the experiments are available at https://github.com/bioinfoUQAM/CASTOR_KRFE. Dylan Lebatteux, Hugo Soudeyns, Isabelle Boucoiran, Soren Gantt, Abdoulaye Baniré Diallo |
BIBM | 5 |
| 2022 | Improving candidate Biosynthetic Gene Clusters in fungi through reinforcement learningabstractMOTIVATION: Precise identification of Biosynthetic Gene Clusters (BGCs) is a challenging task. Performance of BGC discovery tools is limited by their capacity to accurately predict components belonging to candidate BGCs, often overestimating cluster boundaries. To support optimizing the composition and boundaries of candidate BGCs, we propose reinforcement learning approach relying on protein domains and functional annotations from expert curated BGCs. RESULTS: The proposed reinforcement learning method aims to improve candidate BGCs obtained with state-of-the-art tools. It was evaluated on candidate BGCs obtained for two fungal genomes, Aspergillus niger and Aspergillus nidulans. The results highlight an improvement of the gene precision by above 15% for TOUCAN, fungiSMASH and DeepBGC; and cluster precision by above 25% for fungiSMASH and DeepBCG, allowing these tools to obtain almost perfect precision in cluster prediction. This can pave the way of optimizing current prediction of candidate BGCs in fungi, while minimizing the curation effort required by domain experts. AVAILABILITY AND IMPLEMENTATION: https://github.com/bioinfoUQAM/RL-bgc-components. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hayda Almeida, Adrian Tsang, Abdoulaye Baniré Diallo |
Bioinform. | 3 |
| 2021 | A Deep Learning Framework for Improving Lameness Identification in Dairy CattleabstractLameness, characterized by an anomalous gait in cows due to a dysfunction in their locomotive system, is a serious welfare issue for cows and farmers. Prompt lameness detection methods can prevent the development of acute lameness in cattle. In this study, we propose a deep learning framework to help identify lameness based on motion curves of different leg joints on the cow. The framework combines data augmentation and a convolutional neural network using an LeNet architecture. Performance assessed using cross validation showed promising prediction accuracies above 99% and 91% for validation and test sets, respectively. This also demonstrates the usefulness of data generation in cases where the data set is originally small in size and difficult to generate. Yasmine Karoui, Amanda A. Boatswain Jacques, Abdoulaye Baniré Diallo, Elise Shepley, Elsa Vasseur |
AAAI | 3 |
| 2021 | Combining a genetic algorithm and ensemble method to improve the classification of virusesabstractGenomic features identification is an important step toward machine learning (ML) derived classifiers. Designing robust feature identification methods can increase the performance of ML algorithms and reduce high dimensional data. The Classification of genomes from nucleotides sequences often requires an exponential feature space according to the size of the genomes. These feature spaces should also capture the abundant mutations affecting the genomic sequences over time of several complex viruses such as the human immunodeficiency virus (HIV). One way of overcoming such challenges could be to design an accurate and efficient algorithm to capture a bag of minimal subsets of genomic features based on k-mers that could represent the most likely alternative and evolutionary space. In this article, we introduce KEVOLVE, a new method based on a Genetic Algorithm (GA) including a ML kernel to extract a bag of minimal subsets of genomic features maximizing a given classification score threshold. K EVOLVE is coupled with an ensemble prediction model based on support vector machines (SVMs) and is applied on the classification of HIV genomic sequences. Subsets of genomic features identified reduce the size of initial feature matrices by 99% and models based on them outperform state-of-the-art HIV predictors. The results also show that KEVOLVE is less sensitive to high mutation rates of the virus. The source code is available at https://github.com/bioinfoUQAM/Kevolve Dylan Lebatteux, Abdoulaye Baniré Diallo |
BIBM | 2 |
| 2021 | Wireless Sensor Network and Irrigation System to Monitor Wheat Growth under Drought StressabstractStudying drought in a greenhouse setting allows to analyze plant growth under controlled environmental conditions. However, simulating different drought intensities by varying the soil moisture is challenging. This study describes a sensory and control system to simulate drought conditions for wheat, within a framework to study this crop's genetic responses under drought stress. The system uses drip irrigation and allows to maintain the soil moisture within a specified range on potted wheat plants. It allows identifying the amount of water required for irrigation in wheat growth stages, and conducting biological experiments to understand the effects of drought stress on wheat growth. Atia B. Amin, Georges Octave Dubois, Séphora Thurel, Jean Danyluk, Mounir Boukadoum, Abdoulaye Baniré Diallo |
ISCAS | 6 |
| 2021 | Graph pattern mining on top of a domain ontology - preliminary results from a dairy production applicationabstractA domain ontology (DO) is a machine-readable knowledge repository which, whenever properly exploited, can help to discover meaningful and intelligible patterns from compatible datasets. Yet since such data is naturally graph-shaped, the corresponding task amounts to mining what we call ontologically-generalized graph patterns. We study the underlying problem within a dairy production context where a dedicated DO has been designed beforehand. Two alternative mining approaches have been designed, both representing adaptations of methods from the literature. We evaluated them on an excerpt from our dairy production dataset and report here their respective limitations. We also sketch a way to approach the design of ontology-powered graph miner. Tomas Martin, Victor Fuentes, Petko Valtchev, Abdoulaye Baniré Diallo, René Lacroix, Mounir Boukadoum, Maxime Leduc |
KES | 4 |
| 2020 | Towards an Effective Decision-making System based on Cow Profitability using Deep Learning
Charlotte Gonçalves Frasco, Maxime Radmacher, René Lacroix, Roger Cue, Petko Valtchev, Claude Robert, Mounir Boukadoum, Marc-André Sirard, Abdoulaye Baniré Diallo |
ICAART (2) | 9 |
| 2019 | Supporting supervised learning in fungal Biosynthetic Gene Cluster discovery: new benchmark datasetsabstractFungal Biosynthetic Gene Clusters (BGCs) of secondary metabolites are clusters of genes capable of producing natural products, compounds that play an important role in the production of a wide variety of bioactive compounds, including antibiotics and pharmaceuticals. Identifying BGCs can lead to the discovery of novel natural products to benefit human health. Previous work has been focused on developing automatic tools to support BGC discovery in plants, fungi, and bacteria. Data-driven methods, as well as probabilistic and supervised learning methods have been explored in identifying BGCs. Most methods applied to identify fungal BGCs were data-driven and presented limited scope. Supervised learning methods have been shown to perform well at identifying BGCs in bacteria, and could be well suited to perform the same task in fungi. But labeled data instances are needed to perform supervised learning. Openly accessible BGC databases contain only a very small portion of previously curated fungal BGCs. Making new fungal BGC datasets available could motivate the development of supervised learning methods for fungal BGCs and potentially improve prediction performance compared to data-driven methods. In this work we propose new publicly available fungal BGC datasets to support the BGC discovery task using supervised learning. These datasets are prepared to perform binary classification and predict candidate BGC regions in fungal genomes. In addition we analyse the performance of a well supported supervised learning tool developed to predict BGCs. Hayda Almeida, Adrian Tsang, Abdoulaye Baniré Diallo |
BIBM | 3 |
| 2019 | Statistical Linear Models in Virus Genomic Alignment-free Classification: Application to Hepatitis C VirusesabstractViral sequence classification is an important task in pathogen detection, epidemiological surveys and evolutionary studies. Statistical learning methods are widely used to classify and identify viral sequences in samples from environments. These methods face several challenges associated with the nature and properties of viral genomes such as recombination, mutation rate and diversity. Also, new generations of sequencing technologies rise other difficulties by generating massive amounts of fragmented sequences. While linear classifiers are often used to classify viruses, there is a lack of exploration of the accuracy space of existing models in the context of alignment free approaches. In this study, we present an exhaustive assessment procedure exploring the power of linear classifiers in genotyping and subtyping partial and complete genomes. It is applied to the Hepatitis C viruses (HCV). Several variables are considered in this investigation such as classifier types (generative and discriminative) and their hyper-parameters (smoothing value and regularization penalty function), the classification task (genotyping and subtyping), the length of the tested sequences (partial and complete) and the length of k-mer words. Overall, several classifiers perform well given a set of precise combination of the experimental variables mentioned above. Finally, we provide the procedure and benchmark data to allow for more robust assessment of classification from virus genomes. Amine M. Remita, Abdoulaye Baniré Diallo |
BIBM | 2 |
| 2018 | Bioinformatic Workflow Extraction from Scientific Texts based on Word Sense DisambiguationabstractThis paper introduces a method for automatic workflow extraction from texts using Process-Oriented Case-Based Reasoning (POCBR). While the current workflow management systems implement mostly different complicated graphical tasks based on advanced distributed solutions (e.g., cloud computing and grid computation), workflow knowledge acquisition from texts using case-based reasoning represents more expressive and semantic case representations. We propose in this context, an ontology-based workflow extraction framework to acquire processual knowledge from texts. Our methodology extends the classic NLP techniques to extract and disambiguate complex tasks and relations in texts. Using a graph-based representation of workflows and a domain ontology, our extraction process uses a context-aware approach to recognize workflow components in texts: data and control flows. We applied our framework in a technical domain in bioinformatics: i.e., phylogenetic analyses. An evaluation based on workflow semantic similarities in a gold standard proves that our approach provides promising results in the process extraction domain. Both data and implementation of our framework are available in: http://labo.bioinfo.uqam.ca/tgowler. Ahmed Halioui, Petko Valtchev, Abdoulaye Baniré Diallo |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | A machine learning approach for viral genome classificationabstractBACKGROUND: Advances in cloning and sequencing technology are yielding a massive number of viral genomes. The classification and annotation of these genomes constitute important assets in the discovery of genomic variability, taxonomic characteristics and disease mechanisms. Existing classification methods are often designed for specific well-studied family of viruses. Thus, the viral comparative genomic studies could benefit from more generic, fast and accurate tools for classifying and typing newly sequenced strains of diverse virus families. RESULTS: Here, we introduce a virus classification platform, CASTOR, based on machine learning methods. CASTOR is inspired by a well-known technique in molecular biology: restriction fragment length polymorphism (RFLP). It simulates, in silico, the restriction digestion of genomic material by different enzymes into fragments. It uses two metrics to construct feature vectors for machine learning algorithms in the classification step. We benchmark CASTOR for the classification of distinct datasets of human papillomaviruses (HPV), hepatitis B viruses (HBV) and human immunodeficiency viruses type 1 (HIV-1). Results reveal true positive rates of 99%, 99% and 98% for HPV Alpha species, HBV genotyping and HIV-1 M subtyping, respectively. Furthermore, CASTOR shows a competitive performance compared to well-known HIV-1 specific classifiers (REGA and COMET) on whole genomes and pol fragments. CONCLUSION: The performance of CASTOR, its genericity and robustness could permit to perform novel and accurate large scale virus studies. The CASTOR web platform provides an open access, collaborative and reproducible machine learning classifiers. CASTOR can be accessed at http://castor.bioinfo.uqam.ca . Amine M. Remita, Ahmed Halioui, Abou Abdallah Malick Diouara, Bruno Daigle, Golrokh Kiani, Abdoulaye Baniré Diallo |
BMC Bioinform. | 6 |
| 2016 | Face Recognition in the WildabstractFace recognition is one of the most important tasks in pattern recognition and computer vision. The most conventional way to per- form face recognition is to compare a set of facial features that are extracted from a source image or a video frame with a reference image database of known faces. Such a classification takes the form of a prediction within a closed-set of classes. However, a more realistic scenario that fits the ground truth of real-world face recognition applications is to consider the possibility of encountering faces that do not belong to any of the training classes, i . e ., an open-set classification. Such a constraint is very challenging to most existing face recognition systems since the latter are based on closed-set classification methods which always assign a training label to novel unknown instances even if they represent unseen faces that are not represented in the reference database. This results in a misclassification. In this paper, we introduce Face Recognition in the Wild (FRW), a novel face recognition system that allows (1) to efficiently recognize known faces from the reference database, and (2) to prevent misclassifying instances that represent unknown and unseen faces. FRW formulates this problem as a multi-class classification in an open-set context where the presence of instances from unknown classes is possible. Experimental results on the challenging Olivetti Faces benchmark dataset show the efficiency of our approach in open-set face recognition problems. Wajdi Dhifli, Abdoulaye Baniré Diallo |
KES | 2 |
| 2015 | Acquisition of Generic Problem Solving Knowledge through Information Extraction and Pattern MiningabstractIn many technical domains, the generic problem-solving knowledge is scarce even hough a large number of concrete resolutions exist and are well documented. This makes the machine learning from resolution traces approach facing a number of challenges, not least among them the complexity of the underlying domain (concepts, relationships, events, processes, etc.) and the machine-readability of the documented resolution. We tackle here the acquisition of expertise in phylogeny, which is a notoriously rich and prolific field where hundreds, if not thousands, concrete cases are reported in the literature, yet tools to assist the phylogenist in analyzing a new dataset are virtually absent. Thus, we propose an approach that amounts to ontology-based workflow mining: Our T-GROWLer system abstracts general patterns from event sequences previously extracted from texts. It comprises two modules -- a workflow extractor and a pattern miner -- both relying on a pair of ontologies (a domain one and a procedural one). Ahmed Halioui, Petko Valtchev, Abdoulaye Baniré Diallo |
ICTAI | 3 |
| 2015 | Classification of bioinformatics workflows using weighted versions of partitioning and hierarchical clustering algorithmsabstractBACKGROUND: Workflows, or computational pipelines, consisting of collections of multiple linked tasks are becoming more and more popular in many scientific fields, including computational biology. For example, simulation studies, which are now a must for statistical validation of new bioinformatics methods and software, are frequently carried out using the available workflow platforms. Workflows are typically organized to minimize the total execution time and to maximize the efficiency of the included operations. Clustering algorithms can be applied either for regrouping similar workflows for their simultaneous execution on a server, or for dispatching some lengthy workflows to different servers, or for classifying the available workflows with a view to performing a specific keyword search. RESULTS: In this study, we consider four different workflow encoding and clustering schemes which are representative for bioinformatics projects. Some of them allow for clustering workflows with similar topological features, while the others regroup workflows according to their specific attributes (e.g. associated keywords) or execution time. The four types of workflow encoding examined in this study were compared using the weighted versions of k-means and k-medoids partitioning algorithms. The Calinski-Harabasz, Silhouette and logSS clustering indices were considered. Hierarchical classification methods, including the UPGMA, Neighbor Joining, Fitch and Kitsch algorithms, were also applied to classify bioinformatics workflows. Moreover, a novel pairwise measure of clustering solution stability, which can be computed in situations when a series of independent program runs is carried out, was introduced. CONCLUSIONS: Our findings based on the analysis of 220 real-life bioinformatics workflows suggest that the weighted clustering models based on keywords information or tasks execution times provide the most appropriate clustering solutions. Using datasets generated by the Armadillo and Taverna scientific workflow management system, we found that the weighted cosine distance in association with the k-medoids partitioning algorithm and the presence-absence workflow encoding provided the highest values of the Rand index among all compared clustering strategies. The introduced clustering stability indices, PS and PSG, can be effectively used to identify elements with a low clustering support. Etienne Lord, Abdoulaye Baniré Diallo, Vladimir Makarenkov |
BMC Bioinform. | 2 |
| 2011 | Predicting site-specific human selective pressure using evolutionary signaturesabstractMOTIVATION: The identification of non-coding functional regions of the human genome remains one of the main challenges of genomics. By observing how a given region evolved over time, one can detect signs of negative or positive selection hinting that the region may be functional. With the quickly increasing number of vertebrate genomes to compare with our own, this type of approach is set to become extremely powerful, provided the right analytical tools are available. RESULTS: A large number of approaches have been proposed to measure signs of past selective pressure, usually in the form of reduced mutation rate. Here, we propose a radically different approach to the detection of non-coding functional region: instead of measuring past evolutionary rates, we build a machine learning classifier to predict current substitution rates in human based on the inferred evolutionary events that affected the region during vertebrate evolution. We show that different types of evolutionary events, occurring along different branches of the phylogenetic tree, bring very different amounts of information. We propose a number of simple machine learning classifiers and show that a Support-Vector Machine (SVM) predictor clearly outperforms existing tools at predicting human non-coding functional sites. Comparison to external evidences of selection and regulatory function confirms that these SVM predictions are more accurate than those of other approaches. AVAILABILITY: The predictor and predictions made are available at http://www.mcb.mcgill.ca/~blanchem/sadri. CONTACT: [email protected]. Javad Sadri, Abdoulaye Baniré Diallo, Mathieu Blanchette |
Bioinform. | 2 |
| 2011 | Detecting genomic regions associated with a disease using variability functions and Adjusted Rand IndexabstractBACKGROUND: The identification of functional regions contained in a given multiple sequence alignment constitutes one of the major challenges of comparative genomics. Several studies have focused on the identification of conserved regions and motifs. However, most of existing methods ignore the relationship between the functional genomic regions and the external evidence associated with the considered group of species (e.g., carcinogenicity of Human Papilloma Virus). In the past, we have proposed a method that takes into account the prior knowledge on an external evidence (e.g., carcinogenicity or invasivity of the considered organisms) and identifies genomic regions related to a specific disease. RESULTS AND CONCLUSION: We present a new algorithm for detecting genomic regions that may be associated with a disease. Two new variability functions and a bipartition optimization procedure are described. We validate and weigh our results using the Adjusted Rand Index (ARI), and thus assess to what extent the selected regions are related to carcinogenicity, invasivity, or any other species classification, given as input. The predictive power of different hit region detection functions was assessed on synthetic and real data. Our simulation results suggest that there is no a single function that provides the best results in all practical situations (e.g., monophyletic or polyphyletic evolution, and positive or negative selection), and that at least three different functions might be useful. The proposed hit region identification functions that do not benefit from the prior knowledge (i.e., carcinogenicity or invasivity of the involved organisms) can provide equivalent results than the existing functions that take advantage of such a prior knowledge. Using the new algorithm, we examined the Neisseria meningitidis FrpB gene product for invasivity and immunologic activity, and human papilloma virus (HPV) E6 oncoprotein for carcinogenicity, and confirmed some well-known molecular features, including surface exposed loops for N. meningitidis and PDZ domain for HPV. Dunarel Badescu, Alix Boc, Abdoulaye Baniré Diallo, Vladimir Makarenkov |
BMC Bioinform. | 3 |
| 2010 | Ancestors 1.0: a web server for ancestral sequence reconstructionabstractSUMMARY: The computational inference of ancestral genomes consists of five difficult steps: identifying syntenic regions, inferring ancestral arrangement of syntenic regions, aligning multiple sequences, reconstructing the insertion and deletion history and finally inferring substitutions. Each of these steps have received lot of attention in the past years. However, there currently exists no framework that integrates all of the different steps in an easy workflow. Here, we introduce Ancestors 1.0, a web server allowing one to easily and quickly perform the last three steps of the ancestral genome reconstruction procedure. It implements several alignment algorithms, an indel maximum likelihood solver and a context-dependent maximum likelihood substitution inference algorithm. The results presented by the server include the posterior probabilities for the last two steps of the ancestral genome reconstruction and the expected error rate of each ancestral base prediction. AVAILABILITY: The Ancestors 1.0 is available at http://ancestors.bioinfo.uqam.ca/ancestorWeb/. Abdoulaye Baniré Diallo, Vladimir Makarenkov, Mathieu Blanchette |
Bioinform. | 1 |