EDBT 2026 Demo / reviewers in the wild / expert
Alain Bertrand Bomgni
dblp:40/7212
· DBLP profile ↗
17ranked-venue papers
11as first author
16since 2021 · last 2025
0000-0002-3377-7321ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 8 first-author · 10 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dimensional Reprojection and Sequential Modeling with Transformer Integration for Interesting Gene and Protein RecognitionabstractBiomedical Named Entity Recognition (NER) is essential for structuring and extracting vital information from specialized medical texts, thereby improving research and diagnostics, particularly in emerging fields such as biofilm studies, where understanding gene-protein interactions is crucial for characterizing microbial communities and antimicrobial resistance mechanisms. This work presents an innovative hybrid architecture that integrates BioBERT's deep contextualization with HunFlair's sequential modeling capabilities through a novel dimensional reprojection mechanism. The architecture combines a specialized embedding layer (BioBERT dmis-lab/biobert-v1.1), optimized for understanding biomedical and biofilm-related contexts, with a sequential processing suite (BiLSTM-CRF) designed to accurately identify entities such as genes and proteins. A sophisticated dimensional reprojection layer (768$\boldsymbol{\rightarrow} \mathbf{4 2 9 6}$dimensions) employs a learned linear transformation to align and optimize information transfer between layers, enhancing overall performance without compromising structural coherence. We trained our model on 12 harmonized biomedical corpora containing gene and protein annotations related to biofilms and general biomedical domains, with fine-tuning using a learning rate of$5 \times 10^{-6}$over 10 epochs. Testing demonstrates that our model outperforms conventional architectures in biomedical named entity recognition, achieving F1 scores of 90.58 % on BC2GM ($\mathbf{+ 5. 4 3 \%}$compared to BioBERT), 90.70% on JNLPBA (+13.21% compared to HunFlair), 89.20 % on BioNLPCG$(+1.49 \%$compared to HunFlair), and 80.56 % on CRAFT ($+8.37 \%$compared to HunFlair). Precision scores reach 90.75% (BC2GM), 89.32% (JNLPBA), 89.03% (BioNLPCG), and 74.19% (CRAFT). Recall scores are particularly high: 90.41% (BC2GM), 92.12% (JNLPBA), 89.37% (BioNLPCG), and 88.13% (CRAFT), which is essential for comprehensive entity detection in biofilm research, where omitting a critical gene or protein could lead to gaps in understanding microbial mechanisms. Statistical validation confirms the significance of improvements$(\mathbf{p}<0.01)$. These results represent a notable advance over existing models, paving the way for future applications in extracting biofilm-related information from large text datasets and enabling the construction of biofilmspecific knowledge graphs. The code is publicly available to ensure reproducibility. Alain Bertrand Bomgni, Feuzing Ntemma Donald, Shiva Aryal, Bichar Dip Shrestha Gurung, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 1 |
| 2025 | Minimal Features Subset Enabling Essential Gene Prediction Within and Between Organisms for Sulfate Reducing Bacteria FamilyabstractThe identification of essential genes has garnered considerable attention from researchers in recent years. This process of identification uncovers minimal functional modules that enable the survival of an organism, making it of paramount importance in the fields of biomedicine and biotechnology. To address this challenging issue, computational methods have become increasingly utilized to complement experimental approaches, which tend to be intricate and costly. Various classifiers, based on the selection of feature sets, have been proposed and have shown promising results thus far. In this paper, leveraging 50 sulfate reducing bacteria (SRB) organisms - microbes frequently associated with biofilm formation, biofilm-driven corrosion, and complex microbial community dynamics; we aim to show that classifiers can achieve very good performance using only a minimal set of relevant features. Specifically, we demonstrate that classifier performance can be improved by considering minimal relevant features while taking into account the taxonomy of different organisms. A total of 37,500 features were generated from nucleotide and protein sequences of 41 SRB organisms to construct a machine learning model system aimed at predicting essential genes. Our feature engineering module identified 58 subsets of features. Through cross-validation, we achieved competitive intra-organism prediction performance. The best models obtained had an AUC of 0.99, precision of 0.99, recall of 0.99, and an F1-score of 0.99. Subsequently, this system was used to perform extra-organism (new organism not seen by the model) validation using nine left-out SRB organisms. The results obtained for these test organisms demonstrated the efficacy of our models with maximum precision, maximum recall, maximum F1-score, and maximum AUC equal to 0.99,$0.99,0.99$, and 0.97, respectively. Our approach has significantly outperformed previously proposed methods in terms of average metrics, indicating better generalization of the models. Finally, this approach allows researchers to evaluate the predicted result in the lab with fewer variables to consider in their experimental design. Alain Bertrand Bomgni, Junior Basile Fofack, Shiva Aryal, Jerry Lonlac, Venkataramana Gadhamshetty, Etienne T. Gnimpieba |
BIBM | 1 |
| 2025 | Dual-Level Bayesian Predictive Modeling of Biological Nitrogen Fixation: The Effect of TPE Bayesian Optimization on Stacked Model PerformanceabstractBiological Nitrogen Fixation (BNF) is a key ecological process performed by diazotrophs, whose nitrogenase activity is often strictly regulated by environmental factors such as oxygen levels, metal availability, and carbon sources. Given this physiological complexity, accurately predicting nitrogenase activity from protein sequences remains a challenging problem in computational biology. Existing approaches such as Carmma and NFEmbed rely on features derived from large protein language models, yet their performance is often limited by suboptimal hyperparameter tuning. In this work, we introduce a dual-level predictive framework that integrates Bayesian Optimization using the Tree-structured Parzen Estimator (TPE) to systematically optimize both classification and regression models. For the classification task, our TPE-optimized XGBoost model achieves the highest overall performance, with an AUC of 0.9396, F1-score of 0.8675, accuracy of 0.8182, and recall of 0.9231, outperforming Carmma and matching or exceeding NFEmbed in key metrics. For the regression task, our dual-level TPEoptimized stacked SVR attains an$\mathrm{R}^{2}$of 0.6221 with reduced prediction error (MAE: 0.3044, RMSE: 0.5254), representing a substantial improvement over Carmma's SVR ($\mathrm{R}^{2} : 0.5572$) and offering competitive performance relative to NFEmbed. These findings demonstrate that TPE-based hyperparameter optimization—especially when applied at both the base- and meta-learner levels—significantly enhances the predictive reliability of complex machine learning architectures for BNF classification and activity estimation. This establishes systematic Bayesian optimization as a powerful strategy for advancing model accuracy, stability, and generalization in protein-sequence-based bioinformatics. Alain Bertrand Bomgni, Dilane Sagueu Wakam, Ribot Fleury T. Ceskoutsé, Nick Klein, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 1 |
| 2025 | Heterogeneous Graph Network (HGN) for Binary Spatio-Chemical-Source Classification of Per- and Polyfluoroalkyl Substances (PFAS) Contamination: An Optimized ApproachabstractPer- and Polyfluoroalkyl Substances (PFAS) pose a pervasive global environmental and public health challenge due to their persistence and widespread use. Accurate assessment of their contamination is crucial for risk assessment and remediation planning. Traditional machine learning (ML) and conventional geospatial models often fail to fully capture the complex, non-Euclidean interactions between chemical properties, environmental variables, geographical proximity (e.g., groundwater flow and atmospheric transport), and anthropogenic sources. This paper introduces an Optimized Heterogeneous Graph Network (HGN) designed as an alternative approach for the binary classification of PFAS contamination exceeding regulatory thresholds. The HGN models the assessment landscape as a graph where nodes represent sampling locations (spatial features), PFAS compounds (chemical features), and contaminant sources (e.g., industrial facilities), connected by weighted edges reflecting geographic distance, chemical similarity, and source associations. The model utilizes a multi-relational attention mechanism to differentially weigh node features and edge types, thus capturing intricate dependencies more effectively than existing ML approaches, such as ensemble models. Furthermore, we incorporate Bayesian Optimization (BO) for hyperparameter tuning, ensuring peak performance and model stability. Results show that the HGN framework offers a promising alternative in PFAS research by providing a superior representation of the underlying transport, exposure, and source mechanisms, potentially enabling efficient and actionable identification of contamination hotspots through binary outcomes. Alain Bertrand Bomgni, Wilfried Loic Dnjomou Yonmba, Ribot Fleury T. Ceskoutsé, Nick Klein, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 1 |
| 2024 | AIM-HKR: AI-Driven Molecular Retrosynthesis Using Heterogeneous Knowledge RepresentationsabstractThe synthesis of small molecules is a crucial task across multiple scientific domains, including drug discovery, materials science, and sustainable chemistry. As AI and Machine Learning (ML) continue to advance, these technologies offer transformative potential in molecular synthesis, enhancing efficiency and expanding synthetic possibilities. In particular, small molecule synthesis has profound implications for creating novel therapeutic agents, high-performance materials, and greener chemical processes, but a major challenge remains in designing efficient synthetic routes for target molecules. Retrosynthesis, the process of mapping out synthetic pathways by working backwards from a target molecule, represents a vital step; however, traditional retrosynthesis methods often struggle to predict complex or novel transformations. To address this limitation, we introduce AIM-HKR: AI-Driven Molecular Retrosynthesis Using Heterogeneous Knowledge Representations, a model that leverages graph neural network techniques to enhance retrosynthetic predictions. AIM-HKR integrates information from heterogeneous knowledge graphs, capturing the intricate relationships and analogical reasoning required for retrosynthesis. This model generates type-specific embeddings that reflect both network topology and semantic connections across different entity types, such as molecules, reactions, and functional groups. AIM-HKR’s unique capacity to leverage heterogeneous graphs allows it to propose synthetic pathways that extend beyond established precedents, enabling predictions of chemically feasible but previously uncharted transformations. We believe AIM-HKR has the potential to significantly advance molecular synthesis through AI-driven retrosynthesis, establishing a new paradigm in AI-assisted chemistry. Alain Bertrand Bomgni, Ribot Fleury T. Ceskoutsé, Kevin Jordan Njike Njingang, Thomas Bouétou Bouétou, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 1 |
| 2024 | Revisiting Frequent (Closed) Gradual Itemsets MiningabstractThe task of mining gradual itemsets holds significant importance in pattern mining, particularly when working with numerical data. It involves the discovery of covariations between attributes in the form of “The more/less X,…, the more/less Y,” referred to as gradual itemsets. However, discovering these itemsets remains challenging, partly due to the exponential combinatorial search space involved in large-scale data processing. Consequently, existing algorithms for gradual itemset mining encounter difficulties, such as slow processing speeds, and occasional failures to terminate due to the overwhelming number of candidate itemsets requiring exploration. A large number of candidates is generated, but a large proportion of them turns out to be infrequent once their supports are computed. This paper introduces an approach to streamline this process by efficiently reducing the number of candidates for which support needs to be computed through the introduction of a stricter upper-bound criterion. By circumventing the costly support computation for numerous candidate itemsets, our approach exhibits efficiency in terms of speed when applied to real databases, including large-scale databases that pose challenges for existing algorithms. Furthermore, we establish a connection in terms of pattern coverage between the two principal gradualness semantics commonly employed in the literature. Jerry Lonlac, Bernoulli Fotsing Tchide, Alain Bertrand Bomgni, Arnaud Doniec, Engelbert Mephu Nguifo |
ICTAI | 3 |
| 2024 | CIBORG: CIrcuit-Based and ORiented Graph theory permutation routing protocol for single-hop IoT networksabstractThe Internet of Things (IoT) has emerged as a promising paradigm which facilitates the seamless integration of physical devices and digital systems, thereby transforming multiple sectors such as healthcare, transportation, and urban planning. This paradigm is also known as ad-hoc networks. IoT is characterized by several pieces of equipment called objects. These objects have different and limited capacities such as battery, memory, and computing power. These limited capabilities make it difficult to design routing protocols for IoT networks because of the high number of objects in a network. In IoT, objects often have data which does not belong to them and which should be sent to other objects, then leading to a problem known as permutation routing problems. The solution to that problem is found when each object receives its items. In this paper, we propose a new approach to addressing the permutation routing problem in single-hop IoT networks. To this end, we start by representing an IoT network as an oriented graph, and then, based on a reservation channel protocol, we first define a permutation routing protocol for an IoT in a single channel. Secondly, we generalize the previous protocol to make it work in multiple channels. Routing is done using graph theory approaches. The obtained results show that the wake-up times and activities of IoT objects are greatly reduced, thus optimizing network lifetime. This is an effective solution for the permutation routing problem in IoT networks. The proposed approach considerably reduces energy consumption and computation time. It saves 5.2 to 32.04% residual energy depending on the number of items and channels used. Low energy and low computational cost demonstrate that the performance of circuit-based and oriented graph theory is better than the state-of-the-art protocol and therefore is a better candidate for the resolution of the permutation routing problem in single-hop environment. Alain Bertrand Bomgni, Garrik Brel Jagho Mdemaya, Miguel Landry Foko Sindjoung, Mthulisi Velempini, Celine Cabrelle Tchuenko Djoko, Jean Frédéric Myoupo |
J. Netw. Comput. Appl. | 1 |
| 2024 | HeteroKGRep: Heterogeneous Knowledge Graph based Drug Repositioning
Ribot Fleury T. Ceskoutsé, Alain Bertrand Bomgni, David R. Gnimpieba Zanfack, Diing D. M. Agany, Thomas Bouétou Bouétou, Etienne Z. Gnimpieba |
Knowl. Based Syst. | 2 |
| 2023 | Fine-tuning a pre-trained Transformers-based model for gene name entity recognition in biomedical text using a customized dataset: case of Desulfovibrio vulgaris HildenboroughabstractGene Name Entity Recognition (NER) plays a crucial role in the realm of biomedical text mining by focusing on the identification and extraction of gene references from scientific literature. Recent advancements in the field-particularly the emergence of pre-trained transformer-based language models like BioBERT-have shown significant promise in the domain of biomedical NER. However, these models are often trained on existing, publicly available datasets, which may not fully capture the nuances of the domain or adequately cover less-studied genes. This study places its primary emphasis on fine-tuning BioBERT specifically for gene NER tasks. To address the limitations associated with current publicly available datasets, we have meticulously crafted a custom dataset. This dataset is thoughtfully constructed through the systematic collection and detailed annotation of a diverse range of biomedical literature from specialized sources. It intentionally includes genes that have been extensively researched, as well as those that have received limited attention in existing corpora. The fine-tuning process involves initializing the BioBERT model with pre-trained weights and then training it on our custom dataset using a sequence tagging approach. To enhance the model’s performance, we systematically explore various techniques, including data augmentation, entity-level features, and attention mechanisms. Additionally, we conduct rigorous hyperparameter optimization to maximize the model’s accuracy, precision, and recall in gene mention recognition. We thoroughly evaluate the performance of the fine-tuned BioBERT model through a comprehensive set of cross-validation experiments. The results highlight the effectiveness of our tailored dataset in enhancing BioBERT’s performance in gene NER. The fine-tuned model achieves impressive F1 scores, precision, and recall, specifically 0.96, 0.95, and 0.98, surpassing previous models when it comes to recognizing the dvu (Desulfovibrio vulgaris Hildenborough) gene. In conclusion, this study underscores the pivotal role of domain-specific, custom datasets in gene NER tasks. It also highlights the significant potential in fine-tuning pre-trained language models such as BioBERT to enhance gene mention recognition. Alain Bertrand Bomgni, Dialo Abdala, Bichar Dip Shrestha Gurung, Marcellin Nkenlifack, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 1 |
| 2023 | NLPADADE: Leveraging Natural Language Processing for Automated Detection of Adverse Drug EffectsabstractPharmacovigilance is a systematic and scientifically rigorous discipline that assumes responsibility for the safety of pharmaceuticals, with its primary objective being the mitigation of risks while optimizing the benefits associated with medication usage. This mission-critical undertaking plays an indispensable role in preserving public health. At its core, pharmacovigilance entails the methodical collection and proficient management of data pertaining to medication safety. Additionally, these activities encompass the vigilant scrutiny of data to detect emerging "signals" indicative of new or evolving safety concerns. The expert evaluation of this data facilitates well-informed decision-making regarding matters of drug safety. Furthermore, proactive risk management strategies are deployed to effectively mitigate potential associated risks. In the pursuit of proactive health protection, regulatory actions are swiftly executed. Concurrently, the World Health Organization (WHO) underscores the global importance of establishing a robust pharmacovigilance framework. It advocates for the establishment of a comprehensive pharmacovigilance system, defined as encompassing "the science and activities related to the detection, assessment, understanding, and prevention of adverse effects or any other problem related to drugs or any other healthcare product." In this dynamic landscape, a pivotal question arises: "How can the automated identification and extraction of references to diseases, medications, and adverse effects from clinical notes and biomedical literature be achieved?" Central to this discourse are adverse drug effects (ADEs), which present a formidable public health challenge, manifesting as a significant source of patient morbidity and mortality. To expedite the utilization of real-world data (RWD) for the enhancement of pharmacovigilance practices, our focus has gravitated towards the development of a high-performance natural language processing (NLP) model. This model aims to facilitate the rapid detection of potential ADEs linked to medications. Our innovative system, designed for the extraction of diseases and ADEs, leverages the synergy of an open-source NLP component system. The pinnacle of our achievement is the model obtained, which boasts a remarkable Score of 0.97 at step 1800. With a Precision of 1.00, a Recall of 0.9412, and an F-Score of 0.9697, our NLP model showcases its efficiency in extracting and identifying pertinent information from textual data. These results underscore the effectiveness of our approach in recognizing diseases and adverse effects related to medications, setting a new benchmark for pharmacovigilance practices. Alain Bertrand Bomgni, Claude Epiphanie Mbotchack Ngale, Shiva Aryal, Marcellin Nkenlifack, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 1 |
| 2023 | Power-Aware Fog Supported IoT Network for Healthcare Infrastructure Using Swarm Intelligence-Based Algorithms
Hafiz Munsub Ali, Alain Bertrand Bomgni, Syed Ahmad Chan Bukhari, Tahir Hameed |
Mob. Networks Appl. | 2 |
| 2022 | Attention model-based and multi-organism driven gene recognition from text: application to a microbial biofilm organism setabstractNowadays, online databases such as PUBMED and PMC are experiencing an explosion of publications in the field of biomedical sciences. With so much information available online, one of the biggest challenges is managing all that raw, unstructured data and making it machine-readable. Name entity recognition is nowadays a prerequisite for data identification and extraction in biosciences. One of the areas that allows automatic extraction of information from biomedical literature today is Name Entity Recognition. Indeed, it makes it possible to simplify the workflow analysis and automatic extraction of name entities, thus improving the various existing models. There is in the literature a lot of tools for this purpose, but they are unable to extract microbial genes accurately. Moreover, current goal standard corpora such as BIOCREATIVE I to IV have limited representation of microbial knowledge. In this paper, we proposed a new method to recognize biofilm gene mentions from free text. This method relies on a context-specific dictionary to annotate a consistent corpus necessary to train an efficient recognition model. Indeed, this method provides a new workflow for dataset collection generation for microbial biofilm gene. Trained on a set of biofilm organisms our method achieves a score of up to 94%, outperforming state-of-the-art frameworks. Alain Bertrand Bomgni, Ernest Basile Fotseu Fotseu, Daril Raoul Kengne Wambo, Rajesh Kumar Sani, Carol Lushbough, Etienne Z. Gnimpieba |
BIBM | 1 |
| 2022 | NESEPRIN: A new scheme for energy-efficient permutation routing in IoT networksabstractInternet of Things (IoT) consists of a variety of heterogeneous interconnected devices called objects or things. These objects are generally equipped with sensing, processing and wireless communication capabilities. Unfortunately, these capabilities are not enough compared to those of devices in traditional networks. For example, objects have low battery power, limited memory storage, and less processing power. There are some cases where objects have to communicate with each other. In an Autonomous Vehicular Network for example, when a vehicle needs to change its direction, it has to alert other vehicles in the same network. This is known as the permutation routing problem. More precisely, the permutation routing problem in IoT occurs when some things of the network possess items that belong to others. The goal is to send items to their respective owners. A number of solutions to the problem have been proposed in literature which focus mainly on Wireless Sensor Networks (where the memory size of objects are the same). In this paper, we propose an efficient permutation routing scheme for a single-hop IoT network (the memory size differs from one object to another). The proposed NESEPRIN protocol consists of two phases: In the first phase, we solve the permutation routing problem in a single-hop environment with a single channel, and secondly we generalize the previous solution to a network with multiple channels. Our solution makes use of the wake and sleep technique to improve the energy conservation of objects. The simulation results show that our protocol outperforms the existing protocols designed to solve the permutation routing problem in terms of energy saving when the volume of data to route is large. NESEPRIN is the better candidate to solve the permutation routing problem in a single-hop multi-channels IoT environment where the volume of data to route is huge. Alain Bertrand Bomgni, Miguel Landry Foko Sindjoung, Dhalil Kamdem Tchibonsou, Mthulisi Velempini, Jean Frédéric Myoupo |
Comput. Networks | 1 |
| 2022 | A MEC architecture for a better quality of service in an Autonomous Vehicular NetworkabstractDesigning 5G networks and related applications such as the Internet of Things (IoT), Cellular and autonomous vehicular networks (AVNET) is a challenge. Indeed, these networks are subjected to a number of constraints that play a crucial role in the network’s quality of service (QoS). Among these constraints, we have the management of computing, storage, bandwidth resources and low latency requirements. If in the past cloud computing has been used to avoid networks being affected by some of the previous constraints, that technology negatively affects the low latency requirements required by AVNETs for example. Multi-access Edge Computing (MEC) has recently emerged as a palliative solution to cloud computing. MEC aims to bring computing and storage resources from Cloud Data Center to Edge Data Center nearer to the User Equipment (UEs) in order to reduce the UEs’ requests latency for an improvement of the network’s QoS. Many MEC architectures have been proposed for AVNET. These solutions make use of Software Defined Networking (SDN), Network Function Virtualization (NFV), Service Function Chaining (SFC) or Network Slicing (NS) technologies. Some of these techniques combine partially these technologies while others do not. But to the best of our knowledge, none of them combines all these technologies to obtain a better QoS. In this paper, we combine the SDN, NFV, SFC and NS technologies to efficiently manage the MEC server resources for guaranteeing the QoS requirements in AVNETs. The QoS parameters we consider are latency, computing, storage and bandwidth resources. In that way, we first present a MEC server mathematical resource management model, then, we propose a new MEC architecture adapted to AVNETs that uses the aforementioned technologies also as the mathematical model we first present to manage the bandwidth, computing, and storage resources. The simulation results show that these resources are well managed, resulting to a low end to end latency for autonomous vehicles (AVs) requests executing on edge/cloud servers when the direct communication of AVs with cloud server takes long end to end delays, ensuring a better QoS for the AVNET. Miguel Landry Foko Sindjoung, Mthulisi Velempini, Alain Bertrand Bomgni |
Comput. Networks | 3 |
| 2021 | GenNER - A highly scalable and optimal NER method for text-based gene and protein recognitionabstractNowadays, there are a large number of models in the scientific literature capable of recognizing and extracting gene mentions from a given text. Several data sets have been developed to facilitate the learning process of these models. However, very few models are able to increase their knowledge and performance progressively from new annotated text but also to take into account the granularity of the input text of the model. Our proposed solution, GenNER, is a method for recognizing gene/protein mentions from free text. GenNER relies on continuous learning and a text granularization algorithm as input to the model, which allows it to achieve better performance. Its evaluation process was done around BioCreative II annotated datasets; we obtained an average F1-score of 0.9704, which outperforms current methods. Ernest Basile Fotseu Fotseu, Thierry Kongne Nembot, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba, Alain Bertrand Bomgni |
BIBM | 6 |
| 2021 | Prediction of essential genes in G20 using machine learning modelabstractDespite the exponential growth in bioscience data, one of the key challenges for machine learning engineers remains the incompleteness of bioscience dataset (biodata). For a specific bioscience problem such as (e.g. biofilm formation, drug response, organism survival), it is very difficult to find a good consistent dataset capturing the numerous variables involved in each of these processes. Each systems biology data point is measured with different protocols in different settings, making their integration hard and not reliable. This paper focuses on using machine learning (ML) models and data mining (DM) workflow to perform gene essential prediction in G20. Actually, developing next-generation and nano-scale coatings to control biofilm formation on technologically relevant materials is a great challenge today. This can help to control microbial corrosion on material or engineer better relevant material. To tackle this relevant problem, a detailed understanding of the bacterial survival mechanisms is crucial. Computational methods for predicting essential genes can make it easier and faster to obtain reliable results. Method: The main hypothesis of our work is that a minimal information-driven specific Machine Learning model can outperform an interesting prediction score. To reach our goal, we set up first a completed data mining workflow to extract gene features from G20. We then derive 10192 features from gene sequence and protein sequence divided into 25 relevant subgroups. From each subgroup, we build a couple of interesting machine learning models. Result: We identified 69 relevant subgroups of features using our features selection algorithm. We tested the model performance on each of these subgroups and our predictive result achieved up to 98% accuracy score. These subgroups of features can be used to assist researchers to select good variables for their respective experiments. Thierry Kongne Nembot, Ernest Basile Fotseu Fotseu, Rajesh Kumar Sani, Etienne Z. Gnimpieba, Carol Lushbough, Alain Bertrand Bomgni |
BIBM | 6 |
| 2009 | Randomized multi-stage clustering-based geocast algorithms in anonymous wireless sensor networksabstractGeocasting or Multi-Geocasting in wireless sensor network is the delivery of packets from a source (or sink) to all the nodes located in one or several geographic areas. The objectives of a geocasting (multi-geocasting) protocol are the guarantee of message delivery and low transmission cost. The existing protocols which guarantee delivery run on network in which each node has an ID beforehand. They are valid either only in dense networks or must derive a planar graph from the network topology. Hence the nodes may be adapted in order to carry out huge operations to make the network planar. In this paper we consider anonymous networks. To avoid this drawback, we adopt another strategy. Firstly each node acquires a unique identifier in random ranging from 1 to n3 with high probability. Next we partition the network in multi-stage distributed clusters using Gerla & Tsai method. And finally we derive geocast and multi-geocast algorithms that guarantee delivery and that need less overhead with respect to the existing protocols. They are also suitable for networks with irregular distributions with gaps or obstacles. Alain Bertrand Bomgni, Jean Frédéric Myoupo, Aboubecrine Ould Cheikhna |
IWCMC | 1 |