Etienne Z. Gnimpieba

dblp:136/5675 · also Etienne Gnimpieba Zohim, Etienne Zohim Gnimpieba · DBLP profile ↗
← Back
43ranked-venue papers
1as first author
40since 2021 · last 2025
0000-0002-5338-084XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 40 · 1 first-author · 39 since 2021Systems, architecture and hardware · 2Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Dimensional Reprojection and Sequential Modeling with Transformer Integration for Interesting Gene and Protein Recognition
abstract
Biomedical Named Entity Recognition (NER) is essential for structuring and extracting vital information from specialized medical texts, thereby improving research and diagnostics, particularly in emerging fields such as biofilm studies, where understanding gene-protein interactions is crucial for characterizing microbial communities and antimicrobial resistance mechanisms. This work presents an innovative hybrid architecture that integrates BioBERT's deep contextualization with HunFlair's sequential modeling capabilities through a novel dimensional reprojection mechanism. The architecture combines a specialized embedding layer (BioBERT dmis-lab/biobert-v1.1), optimized for understanding biomedical and biofilm-related contexts, with a sequential processing suite (BiLSTM-CRF) designed to accurately identify entities such as genes and proteins. A sophisticated dimensional reprojection layer (768$\boldsymbol{\rightarrow} \mathbf{4 2 9 6}$dimensions) employs a learned linear transformation to align and optimize information transfer between layers, enhancing overall performance without compromising structural coherence. We trained our model on 12 harmonized biomedical corpora containing gene and protein annotations related to biofilms and general biomedical domains, with fine-tuning using a learning rate of$5 \times 10^{-6}$over 10 epochs. Testing demonstrates that our model outperforms conventional architectures in biomedical named entity recognition, achieving F1 scores of 90.58 % on BC2GM ($\mathbf{+ 5. 4 3 \%}$compared to BioBERT), 90.70% on JNLPBA (+13.21% compared to HunFlair), 89.20 % on BioNLPCG$(+1.49 \%$compared to HunFlair), and 80.56 % on CRAFT ($+8.37 \%$compared to HunFlair). Precision scores reach 90.75% (BC2GM), 89.32% (JNLPBA), 89.03% (BioNLPCG), and 74.19% (CRAFT). Recall scores are particularly high: 90.41% (BC2GM), 92.12% (JNLPBA), 89.37% (BioNLPCG), and 88.13% (CRAFT), which is essential for comprehensive entity detection in biofilm research, where omitting a critical gene or protein could lead to gaps in understanding microbial mechanisms. Statistical validation confirms the significance of improvements$(\mathbf{p}<0.01)$. These results represent a notable advance over existing models, paving the way for future applications in extracting biofilm-related information from large text datasets and enabling the construction of biofilmspecific knowledge graphs. The code is publicly available to ensure reproducibility.
Alain Bertrand Bomgni, Feuzing Ntemma Donald, Shiva Aryal, Bichar Dip Shrestha Gurung, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2025 Dual-Level Bayesian Predictive Modeling of Biological Nitrogen Fixation: The Effect of TPE Bayesian Optimization on Stacked Model Performance
abstract
Biological Nitrogen Fixation (BNF) is a key ecological process performed by diazotrophs, whose nitrogenase activity is often strictly regulated by environmental factors such as oxygen levels, metal availability, and carbon sources. Given this physiological complexity, accurately predicting nitrogenase activity from protein sequences remains a challenging problem in computational biology. Existing approaches such as Carmma and NFEmbed rely on features derived from large protein language models, yet their performance is often limited by suboptimal hyperparameter tuning. In this work, we introduce a dual-level predictive framework that integrates Bayesian Optimization using the Tree-structured Parzen Estimator (TPE) to systematically optimize both classification and regression models. For the classification task, our TPE-optimized XGBoost model achieves the highest overall performance, with an AUC of 0.9396, F1-score of 0.8675, accuracy of 0.8182, and recall of 0.9231, outperforming Carmma and matching or exceeding NFEmbed in key metrics. For the regression task, our dual-level TPEoptimized stacked SVR attains an$\mathrm{R}^{2}$of 0.6221 with reduced prediction error (MAE: 0.3044, RMSE: 0.5254), representing a substantial improvement over Carmma's SVR ($\mathrm{R}^{2} : 0.5572$) and offering competitive performance relative to NFEmbed. These findings demonstrate that TPE-based hyperparameter optimization—especially when applied at both the base- and meta-learner levels—significantly enhances the predictive reliability of complex machine learning architectures for BNF classification and activity estimation. This establishes systematic Bayesian optimization as a powerful strategy for advancing model accuracy, stability, and generalization in protein-sequence-based bioinformatics.
Alain Bertrand Bomgni, Dilane Sagueu Wakam, Ribot Fleury T. Ceskoutsé, Nick Klein, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2025 Heterogeneous Graph Network (HGN) for Binary Spatio-Chemical-Source Classification of Per- and Polyfluoroalkyl Substances (PFAS) Contamination: An Optimized Approach
abstract
Per- and Polyfluoroalkyl Substances (PFAS) pose a pervasive global environmental and public health challenge due to their persistence and widespread use. Accurate assessment of their contamination is crucial for risk assessment and remediation planning. Traditional machine learning (ML) and conventional geospatial models often fail to fully capture the complex, non-Euclidean interactions between chemical properties, environmental variables, geographical proximity (e.g., groundwater flow and atmospheric transport), and anthropogenic sources. This paper introduces an Optimized Heterogeneous Graph Network (HGN) designed as an alternative approach for the binary classification of PFAS contamination exceeding regulatory thresholds. The HGN models the assessment landscape as a graph where nodes represent sampling locations (spatial features), PFAS compounds (chemical features), and contaminant sources (e.g., industrial facilities), connected by weighted edges reflecting geographic distance, chemical similarity, and source associations. The model utilizes a multi-relational attention mechanism to differentially weigh node features and edge types, thus capturing intricate dependencies more effectively than existing ML approaches, such as ensemble models. Furthermore, we incorporate Bayesian Optimization (BO) for hyperparameter tuning, ensuring peak performance and model stability. Results show that the HGN framework offers a promising alternative in PFAS research by providing a superior representation of the underlying transport, exposure, and source mechanisms, potentially enabling efficient and actionable identification of contamination hotspots through binary outcomes.
Alain Bertrand Bomgni, Wilfried Loic Dnjomou Yonmba, Ribot Fleury T. Ceskoutsé, Nick Klein, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2025 Metagenomic Functional Signatures for Real-Time Predictive Monitoring of Cyanobacterial Bloom Assessment and Biofilm Marker Discovery in Managed Freshwater Systems
abstract
Harmful algal blooms (HABs), driven by cyanobacteria such as Microcystis, pose a critical and escalating threat to global freshwater security. Standard environmental monitoring often provides reactive rather than predictive data, underestimating the true ecological state after cyanobacteria draw down nutrient levels, and fails to capture the risk of biofilm establishment within water infrastructure. To develop a proactive monitoring framework, we conducted a two-year longitudinal study (2024-2025) of Lake Mitchell, South Dakota, integrating environmental chemistry with advanced computational metagenomics. We used 16S/18S rRNA amplicon sequencing (QIIME2) to profile microbial community structure and Phylogenetic Investigation of Communities by Reconstruction of Unobserved States (PICRUSt2) to predict functional enzyme repertoires. Temporal analysis revealed a significant ecological trajectory toward eutrophic conditions in 2025, characterized by increased alkalinity (pH 8.58 to 8.79) and substantial dissolved nutrient depletion. Crucially, community shifts showed a sharp increase in the HAB-former Microcystis alongside the heterotrophic degrader Flavobacterium. Functional inference reinforced this finding, revealing the significant enrichment of metabolic pathways linked to stress and replication, which often precedes settlement and biofilm initiation: specifically, elevated ATP-dependent helicase (EC:3.6.4.12) and NADH dehydrogenase (EC:1.6.5.3). These enzyme signatures reflect an intensified state of oxidative stress handling and metabolic turnover, characteristic of established bloom communities with high surface-attachment potential. These findings show that predictive functional metagenomic signatures serve as strong and sensitive markers of bloom development and the potential for harmful bacteria to colonize water-treatmen infrastructure. Because these signatures operate independently of short-term nutrient fluctuations, they offer early insight into conditions that may lead to surface-associated microbial growth later in the system-information that is critical for protecting municipal water supplies and public health from HAB contamination.
Tuyen Do, Naina Maharjan, Connor Johansen-Sallee, Bichar Dip Shrestha Gurung, Paula Mazzer, Etienne Z. Gnimpieba
BIBM6
2025 A Convergence Roadmap for AI-Enabled, Omics-Guided Living-Interface Engineering. A National Science Foundation (NSF) National Research Traineeship (NRT) Initiative
abstract
Living-interface engineering seeks to understand and control what happens at the boundaries where materials and microbes meet regions that govern corrosion, filtration, implant safety, biomanufacturing, and water quality. These interfaces become far more complex when biofilms reshape surface chemistry, electron flow, and molecular transport. This paper presents the U.S. national roadmap for AI-enabled, omicsguided living-interface engineering, built on a decade of NSFand National Institute of Health (NIH)-supported research by the South Dakota team. The need for this roadmap is driven by industries whose interface-dependent failures and solutions represent hundreds of billions of dollars annually, with additional trillion-dollar opportunities emerging when microbiome science and AI converge. Our intention is to use this roadmap not only as the foundation for training students across South Dakota universities and collaborating institutions, but also to expand its reach through broader educational platforms such as IEEE workshops and ASCE EWRI short courses. The NSF NRT initiative formalizes this vision into an AI-enabled, digital-twin-supported training ecosystem for the next generation of convergent scientists.
Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM2
2025 Digital Twin COVID Tracker Using Wastewater Data: A Middle-School Led Study Within the U.S. NSF National Research Traineeship Program Framework
abstract
Wastewater infrastructure exists in every municipality across the United States and many other nations, offering a universal, non-invasive platform for community-level disease surveillance. Because viruses such as SARS-CoV-2 shed into wastewater days before symptoms appear, wastewater-based epidemiology (WBE) can provide crucial early-warning signals for public health. This study presents an AI-enabled Digital Twin prototype that predicts short-term COVID-19 trends using Center for Disease Control (CDC) wastewater viral activity data. Uniquely, this project was conceived and executed by middleschool first authors, highlighting the importance of early STEM engagement and intentional mentoring of young professionals on societally relevant environmental and health challenges. This work was conducted as part of our ongoing National Science Foundation (NSF) and National Institutes of Health (NIH) projects led by senior authors, which focus on convergence research and workforce development in AI-enabled, omics-guided living-interface engineering. Computational modeling, Jupyter Notebook workflow, GitHub integration, and cloud deployment were supported by graduate mentors, while system design, experimental logic, and interpretation were led by the student authors. The resulting platform, accessible through an interactive web app and QR-code interface, illustrates how guided, ageappropriate research experiences can empower middle-school students to explore wastewater informatics, digital twin concepts, machine learning, and epidemiological modeling. This work was also recognized with a 3rd-place award in the Sixth Grade Engineering Category at the 2025 High Plains Regional Science & Engineering Fair, highlighting both scientific merit and the broader impact of engaging middle-school students in societally relevant STEM research.
Isha Srikari Gadhamshetty, Jeanne Gnimpieba, Neha Sriveda Gadhamshetty, Shiva Aryal, Arun Kalaga, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM7
2025 Functional Biofilm-Associated Microbial Markers in a Colorectal Cancer Cohort: Evidence for Quorum Sensing, Exopolysaccharide Production, and Biofilm-Driven Pathogenesis
abstract
Colorectal cancer (CRC) is increasingly recognized as a disease influenced by microbe-host interactions, particularly through mucosal biofilms that attach to the epithelial surface and alter the tumor microenvironment. While many microbiome studies focus on taxonomic composition, functional microbial pathways that drive biofilm formation may provide more reliable biomarkers and mechanistic insight. This study analyzes highvariance microbial enzyme signatures derived from a published CRC cohort to identify biofilm-associated genetic markers and to evaluate their roles in CRC biofilm development, quorum sensing signaling, and therapy response. The results highlight enrichment of LuxS/AI-2 quorum sensing enzymes, autoinducer synthases, and exopolysaccharide (EPS) biosynthesis pathways, suggesting a multi-species cooperative biofilm phenotype in CRC. These functional markers align with recent findings that Fusobacterium nucleatum-rich invasive biofilms are common in CRC [4]. The analysis supports four innovation claims: the identification of biofilm-associated functional markers in CRC; mechanistic pathways underlying CRC biofilm development; the impact of biofilm markers on therapy resistance; and additional diagnostic and therapeutic implications. Together, these results emphasize the value of functional biofilm biomarkers in understanding CRC progression and treatment outcomes.
Grace Goeden, Tuyen Do, Bichar Dip Shrestha Gurung, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2025 Digital Twin Forecasting of Quorum-Sensing Associated Biofilm Microorganisms in Urban Wastewater Over a 30-Week Interval
abstract
Biofilms in wastewater systems contain dynamic microbial communities regulated by quorum sensing (QS), which governs adhesion, extracellular polymeric substance (EPS) production, stress tolerance, and developmental transitions. Forecasting QS-associated organisms is essential for anticipating biofilm formation and mitigating operational risks in wastewater infrastructure. In this study, we identified QS and biofilm-associated taxa present in a European wastewater metagenomic dataset: Vibrio harveyi, Vibrio parahaemolyticus, Pseudomonas aeruginosa, Escherichia coli, Salmonella typhi, Salmonella typhimurium, Bacillus subtilis, and Staphylococcus aureus. These organisms encode diverse QS and biofilm regulators across multiple bacterial lineages. From this broader QS-associated set, we developed an organism-specific, QS-aware digital twin focused solely on Pseudomonas, a dominant biofilm-forming genus in wastewater systems. Using the Q-net modeling framework, the model was trained on early-window observations and calibrated to perform short-horizon, end-point forecasting by predicting the final week of a 30-week interval. To ensure a stable conditioning window, the terminal signal was replicated prior to prediction. The digital twin demonstrated strong agreement between predicted and observed final-week abundance$(\mathbf{R}^{2}=0.98)$, highlighting its effectiveness for end-point microbial forecasting. This streamlined and interpretable framework supports rapid microbial surveillance and provides a scalable tool for biofilm-aware operational decision-making in wastewater systems.
Bichar Dip Shrestha Gurung, Tuyen Do, Shiva Aryal, Naina Maharjan, Dikshya Bhandari, Bipul Bhattarai, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM8
2025 Toward Autonomous Biofilm Engineering: A Vision for a Digital Twin-Based Lab Automation
abstract
Biofilm research is undergoing a significant transformation due to a growing need to address biological complexity, temporal dynamics, and environmental sensitivity. These growing needs demand experimental platforms that are far more autonomous, scalable, and computationally integrated than current microbiology lab workflows. This paper presents a vision for a digital twin-driven “cyber-physical laboratory” that combines modular robotics, automated imaging, an Internet of Things network, and AI-based phenotyping into a cohesive system for next-generation computational biofilm engineering. We outline an eight-component architecture comprising twins for physical, environmental, biological, protocol, data, AI/ML, control, and safety constraints. This architecture will enable virtual experimentation, predictive simulation, and self-correction, with the goal of interactive total lab automation and the “robot scientist”. Although we demonstrate early prototypes, the main objective of this paper is to articulate a forward-looking blueprint for how cyber-physical labs with digital twinning can fundamentally reshape biofilm research. Two use cases implementation in our lab shows the feasibility with up to 500 samples run per day (10,000 samples per week). By integrating real-time data streams with predictive biological models and closed-loop automated control, this architecture sets the stage for autonomous, reproducible, and computationally guided biofilm experimentation. We propose this platform as a foundational vision for the future of computational biofilm engineering.
Bichar Dip Shrestha Gurung, Timothy W. Hartman, Tuyen Do, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2025 Development and OMICs Evaluation of Antibiofilm Material Using Hexanoic-Anhydride Modified Chitosan Electrospun Fibers Embedded with Silver Nanoparticles for Antibacterial Effects
abstract
Antimicrobial wound-healing materials are essential in clinical settings because wound infections present a significant challenge due to the attack of pathogenic bacteria, which can rapidly lead to biofilm formation and delayed healing. This work focused on developing the Ag-containing HexanoicAnhydride functionalized chitosan-based electrospun membrane nanocomposite (HA-CEMN) material to control bacterial growth. The Ag-containing HA-CEMN material was developed via electrospinning, and its surface was modified with hexanoic anhydride, and finally, the silver nanoparticles were generated within the nanofiber via an in-situ technique. The prepared Agcontaining HA-CEMN material offers a hydrated microenvironment around the wound area, promoting tissue regeneration. The Ag-containing HA-CEMN material was confirmed with Fourier Transform Infrared Spectroscopy, X-ray diffraction, Scanning electron microscopy, and Thermogravimetric analysis using. Finally, the antimicrobial properties of the prepared Ag-containing HA-CEMN material were studied via the zone of inhibition assay against Gram-negative and Gram-positive bacteria, followed with a mechanistic validation from OMICs data mining. The antimicrobial results demonstrated that the prepared materials show significant bacterial growth suppression towards the selected bacteria. The preliminary results highlight the potential of hexanoic-anhydride-modified Ag-containing HA-CEMN as a next-generation wound dressing with integrated antibacterial, antibiofilm, and regenerative properties.
Tippabattini Jayaramudu, Etienne Z. Gnimpieba, Venkataramana Gadhamshetty
BIBM2
2025 AI-Driven Sensor-Omics Framework for Detecting Antibiotic Resistance in Environmental Engineering Systems
abstract
The proliferation of antibiotic resistance genes (ARGs) in aquatic environments represents a critical One Health challenge, posing interconnected risks to human, animal, and ecosystem health. The spread of ARGs through surface waters can facilitate the emergence of multidrug-resistant pathogens, with far-reaching economic, social, and public health consequences. Traditional surveillance methods for ARG detection, including culture-based assays and molecular techniques such as PCR or metagenomics, are often limited by high costs, extended processing times, technical complexity, and regulatory constraints, hindering timely and comprehensive monitoring. To overcome these challenges, this study introduces the Hybrid Gen-erative-Discriminative Q-Network (Hybrid GDQ-net), a novel deep learning framework designed to integrate heterogeneous data from omics and sensor sources. The framework features a generative module that reconstructs latent representations from incomplete or noisy omics data, capturing hidden patterns that enhance predictive modeling. These latent features are then combined with a lightweight Q-network layer, which learns optimal classification policies through a reward-based training process, improving the accuracy and reliability of ARG detection across diverse environmental conditions. Evaluation of the model demonstrates strong predictive performance, achieving R2values of 0.81 when integrating omics and sensor data, 0.76 using only sensor data, and 0.73 with reduced sequencing coverage (60% of the dataset). These results highlight the ability of the Hybrid GDQ-net to maintain high accuracy even under limited or single-source data conditions, offering a scalable, time-efficient, and cost-effective solution for environmental monitoring. By enabling rapid and reliable ARG detection, this framework supports informed decision-making for water quality management, public health protection, and targeted interventions to mitigate the spread of antibiotic resistance in aquatic ecosystems.
Somayeh Falahati Khanaman, Etienne Z. Gnimpieba, Mengistu Geza Nisrani, Venkataramana Gadhamshetty
BIBM2
2025 Computational Validation of AlphaFold 3 as a Design Engine for Next Generation Aptamer Therapeutics
abstract
The advent of AlphaFold 3 (AF3) marks a pivotal shift in computational structural biology by extending deep learning capabilities beyond protein folding to encompass complex nucleic acid interactions. However, the reliability of this model in predicting the conformational dynamics of single-stranded oligonucleotides for therapeutic applications remains a critical area of investigation. This study rigorously benchmarks AF3 by evaluating its predictive fidelity across a curated dataset of high-affinity aptamers designed for two distinct pathological microenvironments: bacterial biofilms and solid tumors. We assessed the model's ability to resolve aptamer-target interfaces for biofilm disruption, specifically analyzing sequences that inhibit flagellar motility (Flagellin), block methicillin resistance mechanisms mediated by Penicillin-Binding Protein 2a (PBP2a), and disrupt glucan-mediated adhesion via Glucan-Binding Protein C (GbpC). Parallel evaluations were conducted on oncological aptamers designed to suppress angiogenesis via Platelet-Derived Growth Factor subunit B (PDGF-B) pathways and recognize diagnostic biomarkers such as Cancer Antigen 125 (CA-125) and Carcinoembryonic Antigen (CEA). By comparing predicted models against experimental baselines using Root-Mean-Square Deviation (RMSD), our results indicate that AF3 demonstrates high accuracy for rigid protein-aptamer complexes but exhibits variability when modeling interactions with small-molecule targets or functionalized conjugates. These findings establish AF3 as a potent hypothesis-generation engine for computational drug discovery that can significantly accelerate the design pipeline for next-generation nanotherapeutics.
Manish Rayamajhi, Shiva Aryal, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM4
2025 Emerging Biofilm Marker with Thin-Film (CuO) Surface Engineering for Applications in Electrochemical Sensors
abstract
Water quality monitoring strategies are challenged by sensor degradation caused by biofouling, corrosion, and fluctuating environmental conditions. Exposure to natural aqueous environments also leads to biofilm formation on the electrode surface, which alters its physicochemical properties and hinders charge transfer, thereby compromising sensor performance. These challenges hinder the accuracy, stability, and operational lifetime of electrochemical sensors deployed for continuous, in situ detection of contaminant. Developing non-invasive, thinfilm protective coatings that resist fouling while preserving electrochemical activity is therefore essential for reliable and sustainable water quality monitoring. Surface and structural characterizations through atomic force microscopy, Raman spectroscopy, and energy-dispersive X-ray spectroscopy confirmed coating uniformity and integrity. The detection of initial attachment during the biofilm formation, especially the quorum sensing biomarkers released is envisioned to provide insights on the inhibition of the fouling. The analysis of biofouling on the thin film coatings modified with Copper oxide (CuO) has been presented in this work. Ongoing efforts combine microscopy, spectroscopy, electrochemical, and Omics methods to assess fouling resistance and long-term signal stability. This study discusses the pros and cons of engineered thin-film coatings in enabling durable, high-performance sensors for next-generation environmental monitoring and intelligent water infrastructure systems. Future studies will integrate multi-omic information with materials informatics to elucidate foulant/microbe-surface interactions, optimize coating architectures, and enhance predictive modeling of fouling processes. We envision translating these thin-film sensor technologies beyond environmental systems toward biomedical and health monitoring platforms, where biointerface stability and selective sensing are equally critical.
Pawan Kumar Sapkota, Shiva Aryal, Bharat Jasthi, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2025 A Physics-First, Physics-Gated AI Framework for Verifiable Biofilm Sensing
abstract
The Department of Defense (DoD) faces a significant challenge in the Verification and Validation (V&V) of non-deterministic, opaque AI, limiting the deployment of autonomous sensors. This challenge is acute in complex bio-electrochemical systems like biofilms. Current AI approaches for Electrochemical Impedance Spectroscopy (EIS) are bifurcated: (1) computationally intensive, non-interpretable neural networks, or (2) simplified models reliant on a priori ECM selection. We introduce the “Axiomatic-First Framework,” a three-stage methodology that converts first-principles physics into an auditable edge agent. Stage-1 establishes foundational integrity by using Molecular Dynamics (MD) as a physical constraint to model Nernst-Planck and Butler-Volmer physics, generating a “first-principles” dataset. Stage-2 trains a Physics-Informed Neural Network (PINN) “Teacher” whose loss function enforces the governing PDE residuals, yielding calibrated targets for an Equivalent-Circuit Model (ECM) vector$\theta=\left[R_{c t}, C_{d l}, Z_{W}\right]$. Stage-3 deploys a dual-mode “Student” AI: (i) a low-power “Physics-Gated Monitor” and (ii) an on-demand Spiking Neural Network (SNN) “Translator” that recovers$\theta$with uncertainty quantification. This architecture provides verifiable parameter recovery and auditable alarms, preserving low-power duty cycling.
Caine K. Shagla, Manoj Tripathi, Bichar Dip Shrestha Gurung, Etienne Z. Gnimpieba, Venkataramana Gadhamshetty
BIBM4
2025 Digital Biomonitoring of Microbial Communities: a Phenotype Feature-driven AutoML Framework for Bacteria Classification
abstract
Accurate identification of microbial species is essential for monitoring dynamic biological systems and guiding rapid industrial or clinical interventions. Conventional microbiological and genomic methods present an operational trade-off: traditional microscopy is subjective and slow, while high-fidelity sequencing lacks the speed and cost-effectiveness required for high-throughput, real-time screening. This study addresses the resulting “minimal data paradox,” where achieving high specieslevel accuracy requires data volumes that are not available in constrained environments. We introduce a rigorous benchmark of five state-of-the-art Convolutional Neural Networks (CNNs) to evaluate the feasibility of high-accuracy Microbe Species Prediction across scenarios of extreme data scarcity (using only 8 and 20 samples per class). Using transfer learning (TL) and synthetic data augmentation in 23 microbial species, we found that EfficientNetB0 consistently achieved the highest precision and resilience. In the 20 samples/class test, EfficientNetB0 reached 91.30% accuracy and an Area Under the Curve (AUC) of 0.998. Even under the extreme constraint of 8 samples/class, it maintained a superior validation accuracy of 86.96%, demonstrating profound resilience where other modern architectures failed. This benchmark establishes EfficientNetB0 as the optimal, lightweight model for subsequent high-throughput, in-lab operational validation.
Rupesh Kumar Yadav, Shiva Aryal, Bichar Dip Shrestha Gurung, Dikshya Bhandari, Graham Hartman, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM7
2024 Toward Fine-Tuning Large Language Models in Ontology of Microbial Phenotypes Construction
abstract
Ontologies are crucial for organizing domainspecific knowledge in biomedical fields, but their manual construction is time-consuming. This study explores the automation of ontology learning using large language models (LLMs) like BERT, RoBERTa, and DistilBERT, focusing on the Ontology of Microbial Phenotypes (OMP). We investigate three key tasks: (1) entity extraction, (2) relation extraction between entities, and (3) ontology verification. These tasks align with broader applications in biomedical annotation and named entity recognition (NER) by enabling the identification and structuring of key terms and relationships within microbial phenotypes. We evaluate LLMs in two scenarios: baseline performance using pre-trained models and fine-tuned performance after training on OMP-specific data. Our approach integrates spaCy for entity extraction, Llama 2 for relation identification, and LLMs for ontology verification. Experiments reveal that fine-tuned models significantly improve accuracy, precision, recall, and F1 scores, particularly for ontology verification. This research highlights the potential of LLMs to enhance ontology learning and support related biomedical applications like biofilm analysis, annotation, and NER, while emphasizing the value of expert curation.
Anushuya Baidya, Tuyen Do, Etienne Z. Gnimpieba
BIBM3
2024 AIM-HKR: AI-Driven Molecular Retrosynthesis Using Heterogeneous Knowledge Representations
abstract
The synthesis of small molecules is a crucial task across multiple scientific domains, including drug discovery, materials science, and sustainable chemistry. As AI and Machine Learning (ML) continue to advance, these technologies offer transformative potential in molecular synthesis, enhancing efficiency and expanding synthetic possibilities. In particular, small molecule synthesis has profound implications for creating novel therapeutic agents, high-performance materials, and greener chemical processes, but a major challenge remains in designing efficient synthetic routes for target molecules. Retrosynthesis, the process of mapping out synthetic pathways by working backwards from a target molecule, represents a vital step; however, traditional retrosynthesis methods often struggle to predict complex or novel transformations. To address this limitation, we introduce AIM-HKR: AI-Driven Molecular Retrosynthesis Using Heterogeneous Knowledge Representations, a model that leverages graph neural network techniques to enhance retrosynthetic predictions. AIM-HKR integrates information from heterogeneous knowledge graphs, capturing the intricate relationships and analogical reasoning required for retrosynthesis. This model generates type-specific embeddings that reflect both network topology and semantic connections across different entity types, such as molecules, reactions, and functional groups. AIM-HKR’s unique capacity to leverage heterogeneous graphs allows it to propose synthetic pathways that extend beyond established precedents, enabling predictions of chemically feasible but previously uncharted transformations. We believe AIM-HKR has the potential to significantly advance molecular synthesis through AI-driven retrosynthesis, establishing a new paradigm in AI-assisted chemistry.
Alain Bertrand Bomgni, Ribot Fleury T. Ceskoutsé, Kevin Jordan Njike Njingang, Thomas Bouétou Bouétou, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2024 NFB-Checker: An AI/ML-Powered Microorganism Nitrogen Fixation Susceptibility Prediction from Gene Collection
abstract
Nitrogen fixation is a crucial process involved in many aspects of the ecological life cycle. However, few organismsparticularly kingdom bacteria-are known to fix N2in a form that could be used in agribusiness. The ability to identify an organism as an N2fixer is a novel case for more identification and application. This study employs functional genes and an XGBoost machine learning model to predict an organism's ability to fix nitrogen, focusing on key orthologs like nifH, nifD, and nifK, essential in the nitrogenase enzyme complex. The model, trained on a dataset of 923 N2fixers and 981 non-fixers, achieved an accuracy of 92.6% and an F1 score of 0.92. This research highlights the potential of machine learning in identifying genetic markers for N2fixation, providing an efficient alternative to traditional methods, and offers new insights for ecological and agricultural research. The inclusion of specific orthologs enhances the model's predictive accuracy, demonstrating the importance of targeted genetic markers in computational biology.
Tuyen Do, Shiva Aryal, Bichar Dip Shrestha Gurung, Diing D. M. Agany, Nick Klein, Ruanbao Zhou, Rajesh Kumar Sani, Etienne Z. Gnimpieba
BIBM8
2024 Efficient Prediction of Protein-Ligand Binding Using ESM-2 and Mol2vec with Random Forest Model
abstract
In this study, we present a predictive model for protein-ligand binding interactions, leveraging ESM-2 and Mol2vec embeddings combined with a Random Forest classifier. Our approach was evaluated across three testing scenarios—"Molecule unseen," "Protein unseen," and "None seen"—to assess its generalizability to novel compounds and targets. Compared to deep learning models, our method demonstrated competitive AUROC scores and consistently high specificity, indicating robust predictive performance and a conservative strategy that minimizes false positives. While sensitivity was lower, strategies to address data imbalance and enhance model sensitivity are discussed. Our findings highlight the potential of combining efficient embedding techniques with interpretable machine learning models for scalable and effective binding interaction prediction
Bichar Dip Shrestha Gurung, Manish Rayamajhi, Anushuya Baidya, Owen Growney, Helena Naomie Dongmo Mafo, Yun-Seok Choi, Etienne Z. Gnimpieba
BIBM7
2024 Enhancing Sulphate Reducing Bacteria Clustering by Implementing UMAP on Raman Spectroscopy Data
Md Hasanur Rahman 0004, Bichar Dip Shrestha Gurung, Manoj Tripathi, Alan Dalton, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2024 Applying the DeepSeqDock Harmonization Framework on Pseudomonas Aeruginosa Bacterial Biofilm Data
abstract
In the modern era, data availability has exponentially increased. From different universities, institutions, and laboratories across the globe, the wealth of information brings potential for integration and downstream analysis. This paper focuses specifically on the integration of RNA-sequencing data using data harmonization, a technique that reduces the inter-institutional batch effects that confound data quality with non-biological factors. One such framework is the DeepSeqDock framework, which creates data harmonization pipelines through training and tuning in specific domains and is evaluated on human RNA data from the SEQC. In this study, we present an alternate application of the adjusted DeepSeqDock framework for bacterial biofilm data. By adjusting for non-human data, we show the generalization of data harmonization and the DeepSeqDock for its applications in specific bacterial domains of interest for downstream analysis.
Corey Zhao, Bichar Dip Shrestha Gurung, Timothy W. Hartman, Tuyen Do, Mariah M. Hoffman, Etienne Z. Gnimpieba
BIBM6
2024 Biofilm marker discovery with cloud-based dockerized metagenomics analysis of microbial communities
abstract
In an environment, microbes often work in communities to achieve most of their essential functions, including the production of essential nutrients. Microbial biofilms are communities of microbes that attach to a nonliving or living surface by embedding themselves into a self-secreted matrix of extracellular polymeric substances. These communities work together to enhance their colonization of surfaces, produce essential nutrients, and achieve their essential functions for growth and survival. They often consist of diverse microbes including bacteria, viruses, and fungi. Biofilms play a critical role in influencing plant phenotypes and human microbial infections. Understanding how these biofilms impact plant health, human health, and the environment is important for analyzing genotype-phenotype-driven rule-of-life functions. Such fundamental knowledge can be used to precisely control the growth of biofilms on a given surface. Metagenomics is a powerful tool for analyzing biofilm genomes through function-based gene and protein sequence identification (functional metagenomics) and sequence-based function identification (sequence metagenomics). Metagenomic sequencing enables a comprehensive sampling of all genes in all organisms present within a biofilm sample. However, the complexity of biofilm metagenomic study warrants the increasing need to follow the Findability, Accessibility, Interoperability, and Reusable (FAIR) Guiding Principles for scientific data management. This will ensure that scientific findings can be more easily validated by the research community. This study proposes a dockerized, self-learning bioinformatics workflow to increase the community adoption of metagenomics toolkits in a metagenomics and meta-transcriptomics investigation. Our biofilm metagenomics workflow self-learning module includes integrated learning resources with an interactive dockerized workflow. This module will allow learners to analyze resources that are beneficial for aggregating knowledge about biofilm marker genes, proteins, and metabolic pathways as they define the composition of specific microbial communities. Cloud and dockerized technology can allow novice learners-even those with minimal knowledge in computer science-to use complicated bioinformatics tools. Our cloud-based, dockerized workflow splits biofilm microbiome metagenomics analyses into four easy-to-follow submodules. A variety of tools are built into each submodule. As students navigate these submodules, they learn about each tool used to accomplish the task. The downstream analysis is conducted using processed data obtained from online resources or raw data processed via Nextflow pipelines. This analysis takes place within Vertex AI's Jupyter notebook instance with R and Python kernels. Subsequently, results are stored and visualized in Google Cloud storage buckets, alleviating the computational burden on local resources. The result is a comprehensive tutorial that guides bioinformaticians of any skill level through the entire workflow. It enables them to comprehend and implement the necessary processes involved in this integrated workflow from start to finish. This manuscript describes the development of a resource module that is part of a learning platform named "NIGMS Sandbox for Cloud-based Learning" https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [1] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses.
Etienne Z. Gnimpieba, Timothy W. Hartman, Tuyen Do, Jessica Zylla, Shiva Aryal, Samuel J. Haas, Diing D. M. Agany, Bichar Dip Shrestha Gurung, Valena Doe, Zelaikha B. Yosufzai, Daniel Pan, Ross Campbell, Victor C. Huber, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Carol Lushbough
Briefings Bioinform.1
2024 HeteroKGRep: Heterogeneous Knowledge Graph based Drug Repositioning
Ribot Fleury T. Ceskoutsé, Alain Bertrand Bomgni, David R. Gnimpieba Zanfack, Diing D. M. Agany, Thomas Bouétou Bouétou, Etienne Z. Gnimpieba
Knowl. Based Syst.6
2023 Fine-tuning a pre-trained Transformers-based model for gene name entity recognition in biomedical text using a customized dataset: case of Desulfovibrio vulgaris Hildenborough
abstract
Gene Name Entity Recognition (NER) plays a crucial role in the realm of biomedical text mining by focusing on the identification and extraction of gene references from scientific literature. Recent advancements in the field-particularly the emergence of pre-trained transformer-based language models like BioBERT-have shown significant promise in the domain of biomedical NER. However, these models are often trained on existing, publicly available datasets, which may not fully capture the nuances of the domain or adequately cover less-studied genes. This study places its primary emphasis on fine-tuning BioBERT specifically for gene NER tasks. To address the limitations associated with current publicly available datasets, we have meticulously crafted a custom dataset. This dataset is thoughtfully constructed through the systematic collection and detailed annotation of a diverse range of biomedical literature from specialized sources. It intentionally includes genes that have been extensively researched, as well as those that have received limited attention in existing corpora. The fine-tuning process involves initializing the BioBERT model with pre-trained weights and then training it on our custom dataset using a sequence tagging approach. To enhance the model’s performance, we systematically explore various techniques, including data augmentation, entity-level features, and attention mechanisms. Additionally, we conduct rigorous hyperparameter optimization to maximize the model’s accuracy, precision, and recall in gene mention recognition. We thoroughly evaluate the performance of the fine-tuned BioBERT model through a comprehensive set of cross-validation experiments. The results highlight the effectiveness of our tailored dataset in enhancing BioBERT’s performance in gene NER. The fine-tuned model achieves impressive F1 scores, precision, and recall, specifically 0.96, 0.95, and 0.98, surpassing previous models when it comes to recognizing the dvu (Desulfovibrio vulgaris Hildenborough) gene. In conclusion, this study underscores the pivotal role of domain-specific, custom datasets in gene NER tasks. It also highlights the significant potential in fine-tuning pre-trained language models such as BioBERT to enhance gene mention recognition.
Alain Bertrand Bomgni, Dialo Abdala, Bichar Dip Shrestha Gurung, Marcellin Nkenlifack, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2023 NLPADADE: Leveraging Natural Language Processing for Automated Detection of Adverse Drug Effects
abstract
Pharmacovigilance is a systematic and scientifically rigorous discipline that assumes responsibility for the safety of pharmaceuticals, with its primary objective being the mitigation of risks while optimizing the benefits associated with medication usage. This mission-critical undertaking plays an indispensable role in preserving public health. At its core, pharmacovigilance entails the methodical collection and proficient management of data pertaining to medication safety. Additionally, these activities encompass the vigilant scrutiny of data to detect emerging "signals" indicative of new or evolving safety concerns. The expert evaluation of this data facilitates well-informed decision-making regarding matters of drug safety. Furthermore, proactive risk management strategies are deployed to effectively mitigate potential associated risks. In the pursuit of proactive health protection, regulatory actions are swiftly executed. Concurrently, the World Health Organization (WHO) underscores the global importance of establishing a robust pharmacovigilance framework. It advocates for the establishment of a comprehensive pharmacovigilance system, defined as encompassing "the science and activities related to the detection, assessment, understanding, and prevention of adverse effects or any other problem related to drugs or any other healthcare product." In this dynamic landscape, a pivotal question arises: "How can the automated identification and extraction of references to diseases, medications, and adverse effects from clinical notes and biomedical literature be achieved?" Central to this discourse are adverse drug effects (ADEs), which present a formidable public health challenge, manifesting as a significant source of patient morbidity and mortality. To expedite the utilization of real-world data (RWD) for the enhancement of pharmacovigilance practices, our focus has gravitated towards the development of a high-performance natural language processing (NLP) model. This model aims to facilitate the rapid detection of potential ADEs linked to medications. Our innovative system, designed for the extraction of diseases and ADEs, leverages the synergy of an open-source NLP component system. The pinnacle of our achievement is the model obtained, which boasts a remarkable Score of 0.97 at step 1800. With a Precision of 1.00, a Recall of 0.9412, and an F-Score of 0.9697, our NLP model showcases its efficiency in extracting and identifying pertinent information from textual data. These results underscore the effectiveness of our approach in recognizing diseases and adverse effects related to medications, setting a new benchmark for pharmacovigilance practices.
Alain Bertrand Bomgni, Claude Epiphanie Mbotchack Ngale, Shiva Aryal, Marcellin Nkenlifack, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2023 Computational methods for biofouling and corrosion-resistant graphene nanocomposites. A transdisciplinary approach
abstract
The current study addresses the critical issue of bacterial adhesion and proliferation on surfaces, particularly in industries such as petroleum, marine, textiles, healthcare, food, and water management, with global biofilm-associated costs exceeding $5 trillion. Here we propose a transdisciplinary approach with computational methods to bring together the strengths, particularly advanced knowledge and cutting edge technologies of life sciences, material science, nanotechnology, engineering, and machine learning, with a goal of addressing vexing challenges facing microbiologically influenced corrosion. The current study focuses on developing biofouling resistant graphene nanocomposites with a thin silver film on nickel substrates against sulfate reducing bacteria (SRB). SEM images of biofilm formation after 30 days of exposure to SRB are analyzed using a pretrained deep learning architecture and otsu thresholding for semi-automated microbial segmentation. A pretrained multi-task deep learning model cellpose and segment anything model are employed to extract bacterial regions from the images, providing precise quantitative analysis of SRB cell surface coverage on graphene nanocomposites. This computational model establishes the interrelation between corrosion findings and the antibiofouling performance of graphene nanocomposites on nickel surfaces.
Ramesh Devadig, Bichar Dip Shrestha Gurung, Etienne Z. Gnimpieba, Bharat Jasthi, Venkataramana Gadhamshetty
BIBM3
2023 Utilizing XGBoost for the Prediction of Material Corrosion Rates from Embedded Tabular Data using Large Language Model
abstract
Microbial corrosion, scientifically referred to as microbial-induced corrosion (MIC), constitutes a noteworthy and frequently underestimated concern within diverse industrial domains. This phenomenon manifests when microorganisms, including bacteria, archaea, and fungi, engage with structural materials, resulting in the degradation of infrastructure and equipment. The accurate prognostication of material microbial corrosion rates is of upmost importance in the formulation of proactive strategies for maintenance and corrosion control. In this study, a novel methodology is introduced, which harnesses the capabilities of XGBoost, an advanced gradient boosting algorithm, for the precise prediction of material microbial corrosion rates. This predictive process is facilitated by employing tabular data that is intricately embedded within a comprehensive large language models (LLMs). The integration of tabular data into the language model yields a sophisticated contextual comprehension of the data, thereby augmenting the model's precision by its aptitude to discern intricate relationships and semantic nuances intrinsic to the tabular data.
Tuyen Do, Bichar Dip Shrestha Gurung, Shiva Aryal, Anup Khanal, Sandeep Chataut, Venkataramana Gadhamshetty, Carol Lushbough, Etienne Z. Gnimpieba
BIBM8
2023 Transformer in Microbial Image Analysis: A Comparative Exploration of TransUNet, UNet, and DoubleUNet for SEM Image Segmentation
abstract
The advent of transformer-based architectures such as TransUNet has revolutionized image segmentation as this approach combines the strengths of transformers for capturing contextual information with convolutional neural networks (CNNs) for localized feature identification. Microbes, known for their complex behaviors, present challenges in various fields, especially biomedicine. Image segmentation is crucial for ana-lyzing microbes, allowing quantitative analysis, growth tracking, and understanding host-pathogen interactions. This study is dedicated to a comparative analysis of TransUNet alongside two other popular segmentation methods, UNet and DoubleUNet, in the context of segmenting scanning electron microscope (SEM) images of microbes on layered graphene-nickel specimens. The TransUNet architecture employs a pre-defined ResNet-50 and Vision Transformer (ViT) as the encoder and a custom-built decoder trained on SEM data of Oleidesulfovibrio alaskensis (OA-G20) exposed to graphene-nickel specimens for 30 days. Using the Intersection Over Union (IoU) score as a performance metric, we observed that TransUNet achieved a maximum IoU of 79.58%, DoubleUNet exhibited a maximum IoU of 76.28%, and UNet attained a maximum IoU of 72.38%. We believe that this comparative study of the segmentation approach is invaluable for selecting the best model for the practitioner as per need. This study is the first step in our aim of developing an end-to-end framework with automated model selection based on dataset characteristics for microbial image segmentation.
Bichar Dip Shrestha Gurung, Anup Khanal, Timothy W. Hartman, Tuyen Do, Sandeep Chataut, Carol Lushbough, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM8
2023 SIMPL - An Application for Cell and Microbe Tracking Using Machine Learning
abstract
In the ever-evolving landscape of research and clinical practices, the role of microscope-based image capture and analysis cannot be overstated. While current methodologies for micrograph analysis offer substantial power, there exist critical gaps, particularly in user functionality and reproducibility. To bridge these gaps, we introduce the Smart Imaging of Micrographs Process and Labeling (SIMPL) system as an open-source, semi-automated framework designed for image and video capture, analysis, and particle tracking. SIMPL aims to meet the demands of high-throughput applications, especially for reproducible microbe tracking.
Sam Haas, Bichar Dip Shrestha Gurung, Timothy W. Hartman, Tuyen Do, Etienne Z. Gnimpieba
BIBM5
2023 Genome wide computational prediction and analysis of noncoding genome of corrosive biofilm forming Oleidesulfovibrio alaskensis G20
abstract
Biofilm, is a special-complex organization of bacterial cells with multiple layers, formed in certain environmental conditions. In the complex biofilm different bacterial cells perform different functions and contribute to overall biofilm activity, in different physiological states of growth, stress and signaling. Sulphate reducing bacteria (SRB) contributes to huge economic loss ($4B) causing microbial induced corrosion. To effectively combat the challenges posed by SRB, it is essential to understand their molecular mechanisms of biofilm formation and biocorrosion. Regulation of all key pathways and mechanisms involved in growth, stress management and biofilm formation performed by non-coding genomes. Here, in this study we have identified genome wide distribution of non-coding genome (ncRNAs) in a biofilm forming and corrosive SRB Oleidesulfovibrio alaskensis (earlier Desulfovibrio alaskensis) strain G20 (OA-G20). It is understood that regulatory small RNAs play key role in therefore objective of this study was to identify the noncoding RNAs in OA-G20 genome and unveil their role in biofilm formation and corrosive nature. This is the first study to identify ncRNAs in SRB. Here, we are presenting the results using computational methods to identify noncoding genome of OA-G20 and potential role in biofilm formation.
Ram Nageena Singh, Etienne Z. Gnimpieba, Rajesh Kumar Sani
BIBM2
2022 Attention model-based and multi-organism driven gene recognition from text: application to a microbial biofilm organism set
abstract
Nowadays, online databases such as PUBMED and PMC are experiencing an explosion of publications in the field of biomedical sciences. With so much information available online, one of the biggest challenges is managing all that raw, unstructured data and making it machine-readable. Name entity recognition is nowadays a prerequisite for data identification and extraction in biosciences. One of the areas that allows automatic extraction of information from biomedical literature today is Name Entity Recognition. Indeed, it makes it possible to simplify the workflow analysis and automatic extraction of name entities, thus improving the various existing models. There is in the literature a lot of tools for this purpose, but they are unable to extract microbial genes accurately. Moreover, current goal standard corpora such as BIOCREATIVE I to IV have limited representation of microbial knowledge. In this paper, we proposed a new method to recognize biofilm gene mentions from free text. This method relies on a context-specific dictionary to annotate a consistent corpus necessary to train an efficient recognition model. Indeed, this method provides a new workflow for dataset collection generation for microbial biofilm gene. Trained on a set of biofilm organisms our method achieves a score of up to 94%, outperforming state-of-the-art frameworks.
Alain Bertrand Bomgni, Ernest Basile Fotseu Fotseu, Daril Raoul Kengne Wambo, Rajesh Kumar Sani, Carol Lushbough, Etienne Z. Gnimpieba
BIBM6
2022 U-Net Based Image Segmentation Techniques for Development of Non-Biocidal Fouling-Resistant Ultra-Thin Two-dimensional (2D) Coatings
abstract
Bacterial adhesion to the metallic surfaces creates a complex biofilm network, resulting in many problems like corrosion and fouling. Precise quantitative analysis of the surface coverage of cells can be vital in decoding biofilm-related issues. We present a deep learning-based approach to automate the microbes’ segmentation from the Scanning Electron Microscope (SEM) images of biofilm developed on coated multilayer graphene nickel samples. We collected SEM images from multilayer graphene nickel exposed to Oleidesulfovibrioalaskensis (OA-G20) for 30 days. Then we manually annotated the microbes with the help of subject expertise and trained a deep-learning U-Net architecture. In order to deal with larger image sizes, we perform patched-based image training and techniques for predicting segmentation masks over the larger image. Intersection over Union (IOU) was calculated for the evaluation of the performance of the model. After training the image for multiple epochs and extracting the optimal model parameters from the learning curve, we were able to get 70.64% mean IOU score. The patched-based technique for image training and image inference showed promising output during the segmentation.
Bichar Dip Shrestha Gurung, Ramesh Devadig, Tuyen Do, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2022 Using BASIN-ML for Machine Learning-Based Statistical Analysis and Reporting for Biofilm Datasets
abstract
Biological and biomedical microscope image (bioimage) comparison remains useful to approach many research challenges—from biofilms to human diseases. This powerful technology allows researchers to provide the community with a quick visual snapshot of varying experimental conditions. But a two-condition comparison still relies on a researcher’s eyes to draw conclusions despite the availability of multiple— often complex—digital image analysis tools. Our Bioimage Analysis, Statistic, and Comparison (BASIN) software provides an easy, objective, reproducible comparison leveraging inferential statistics to bridge image data analysis with other biomedical data modalities such as gene expression. Users have access to a machine learning module to assist with image segmentation using modern, trainable algorithms. BASIN also provides several key data points including images’ object counts, net and mean pixel intensities, net and mean object surface areas, plus a variety of other potentially useful data. Hypothesis testing is performed on mean object intensities and surface areas using the statistical power of the R programming language. These features allow BASIN to extend the current scope of image comparison. It gives researchers a multi-model knowledge about matters such as drug protein marker response, the significance of cell population changes, and changes in cell morphology. To improve BASIN’s accessibility and transparency we implemented it in R using Shiny framework and provided both an online trial version and a customizable offline version. We also have a batch version to run on datasets with hundreds of biomedical images. BASIN workflows consist of five core modules including image upload, feature extraction, statistical analysis, visualization, and report generation.
Sam Haas, Timothy W. Hartman, Bichar Dip Shrestha Gurung, Tuyen Do, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM7
2022 Title: Challenges in single cells sequencing Microbial community and biofilm: A case of Oleidesulfovibrio alaskensis G20 NGS protocol
abstract
Biofilm, is a special-complex organization of bacterial cells with multiple layers, formed in certain environmental conditions. In the complex biofilm different bacterial (single) cells perform different functions and contribute to overall biofilm activity, where single bacterial cells are in different physiological states of growth, stress and signaling. Sulphate reducing bacteria (SRB) contributes to huge economic loss (${\$}$4B) causing microbial induced corrosion. To effectively combat the challenges posed by SRB, it is essential to understand their molecular mechanisms of biofilm formation and biocorrosion. Single Cell Genomics (single cell transcriptomics) is a technique which investigates the gene expression at the level of a single biological cell and could not be attainable in bulk analysis (RNASeq) could decipher the uniqueness of each cell in complex group. We are developing a NGS (oxford Nanopore)-workflow for single cell transcriptomics of biofilms using Oleidesulfovibrio alaskensis (earlier Desulfovibrio alaskensis) strain G20 (OA-G20) as a model. We encountered several challenges to design and develop the workflow such 1) absence of ploy(A) tail in mRNAs, 2) amount of total RNA, 3) half-lives of mRNAs, 4) perforation of cell membrane, 5) low copy number of mRNAs, 6) single cell suspension and cell sorting, 7) quantification of bacterial cells, 8) fixation of cells, 9) designing of probes, 10) hybridization of probes and 11) amplification of mRNAs. Here, we are presenting how we resolved some of these challenges to design and develop a single cell genomics (transcriptomics) workflow for SRB biofilms.
Ram Nageena Singh, Etienne Z. Gnimpieba, Rajesh Kumar Sani
BIBM2
2021 GenNER - A highly scalable and optimal NER method for text-based gene and protein recognition
abstract
Nowadays, there are a large number of models in the scientific literature capable of recognizing and extracting gene mentions from a given text. Several data sets have been developed to facilitate the learning process of these models. However, very few models are able to increase their knowledge and performance progressively from new annotated text but also to take into account the granularity of the input text of the model. Our proposed solution, GenNER, is a method for recognizing gene/protein mentions from free text. GenNER relies on continuous learning and a text granularization algorithm as input to the model, which allows it to achieve better performance. Its evaluation process was done around BioCreative II annotated datasets; we obtained an average F1-score of 0.9704, which outperforms current methods.
Ernest Basile Fotseu Fotseu, Thierry Kongne Nembot, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba, Alain Bertrand Bomgni
BIBM5
2021 Automatic Extension of Medical Subject Headings (MeSH) Thesaurus to Emerging Research
abstract
The proliferation of information technology infrastructure in recent decades has allowed for unprecedented ease of access to centrally-aggregated scholarly literature and scientific knowledge. This massive aggregation of knowledge requires an information retrieval infrastructure, to include formalized ontologies, that is engineered with careful consideration. A number of domains benefit from the use of hierarchical controlled vocabularies, which may be used to provide a rich set of descriptive terms for characterizing entities in a consistent manner. There are clear benefits to the creation and maintenance of these ontologies: search and retrieval is made easier and analyses of the contained entities are enabled that would not otherwise be possible. However, there may be the opportunity to decrease the manual burden of ontology creation and maintenance with automated methods that leverage natural language processing and other computational techniques. This work presents an automated ontology creation methodology, adapted and expanded from prior work [1], that can produce a topic hierarchy from natural language and may be used to assist in the creation of a novel ontology or the expansion of existing ontologies. The effectiveness of the proposed method is studied using two examples: immunology, an established biomedical domain and a prominent topic in MeSH, and graphene, from the 2D materials domain with wide-ranging biomedical applications, which also has a sparse presence in MeSH
William Gasper, Dario Ghersi, Etienne Z. Gnimpieba, Venkataramana Gadhamshetty, Parvathi Chundi
BIBM4
2021 Prediction of essential genes in G20 using machine learning model
abstract
Despite the exponential growth in bioscience data, one of the key challenges for machine learning engineers remains the incompleteness of bioscience dataset (biodata). For a specific bioscience problem such as (e.g. biofilm formation, drug response, organism survival), it is very difficult to find a good consistent dataset capturing the numerous variables involved in each of these processes. Each systems biology data point is measured with different protocols in different settings, making their integration hard and not reliable. This paper focuses on using machine learning (ML) models and data mining (DM) workflow to perform gene essential prediction in G20. Actually, developing next-generation and nano-scale coatings to control biofilm formation on technologically relevant materials is a great challenge today. This can help to control microbial corrosion on material or engineer better relevant material. To tackle this relevant problem, a detailed understanding of the bacterial survival mechanisms is crucial. Computational methods for predicting essential genes can make it easier and faster to obtain reliable results. Method: The main hypothesis of our work is that a minimal information-driven specific Machine Learning model can outperform an interesting prediction score. To reach our goal, we set up first a completed data mining workflow to extract gene features from G20. We then derive 10192 features from gene sequence and protein sequence divided into 25 relevant subgroups. From each subgroup, we build a couple of interesting machine learning models. Result: We identified 69 relevant subgroups of features using our features selection algorithm. We tested the model performance on each of these subgroups and our predictive result achieved up to 98% accuracy score. These subgroups of features can be used to assist researchers to select good variables for their respective experiments.
Thierry Kongne Nembot, Ernest Basile Fotseu Fotseu, Rajesh Kumar Sani, Etienne Z. Gnimpieba, Carol Lushbough, Alain Bertrand Bomgni
BIBM4
2021 Integration of text mining and biological network analysis to access essential genes in Desulfovibrio alaskensis G20
abstract
Essential genes are crucial for the survival and growth of any organism, and therefore alteration of such genes could result in unexpected behavioral change. Identification of essential genes and their role in functioning of organisms is a basic knowledge requirement for any research, which could be manipulated to understand the mechanisms of survival and growth [1]. Several decades have witnessed the virtues and iniquities of the gram-negative facultative anaerobes, sulfate reducing bacteria (SRB) in both ecological and commercial arena. Despite of relentless increase in the number of published articles that belong to diverse research areas-from industrial biotechnology (removal of heavy metals and waste valorization) to molecular biology (genetic architecture of the genes in biocorrosion and biofilm formation on metals), not much information about the essential genes of SRB community is known yet [2]. The Desulfovibrio alaskensis G20 (DA-G20) is a well-known SRB; its genes have been annotated but have large numbers that encode for hypothetical proteins. Till date no categorization is available for the genes of DA-G20 with reference to essentiality [3]. The in-vitro prediction of essential genes relies highly on the exhaustive multi-omics strategies. In the era of big-data and artificial intelligence research, demand of abstraction and interpretation of complex relationships of biological importance using text mining has increased. Therefore, to propose an alternative and economic method, text mining is a comparable method for the prediction of the essential genes. In this study, we reported the essential genes of DA-G20 using text mining and biological network analysis. Moreover, the present work provides a foundation for the expansion of genome wide investigation and identification of essential genes in prokaryotes using machine learning and data science approaches.
Priya Saxena, Abhilash Kumar Tripathi, Payal Thakur, Shailabh Rauniyar, Vinoj Gopalakrishnan, Ram Nageena Singh, Mathew A. Olakunle, Etienne Z. Gnimpieba, Rajesh Kumar Sani
BIBM8
2021 Identifying genes involved in biocorrosion from the literature using text-mining
abstract
Background: The central repository of scientific models from the literature; however, the manual knowledge is research publications which also plays a selection of pertinent information from databases like crucial role in communication within the scientific PubMed, PMC, Dimension, Google scholar, and community. The public repositories like PubMed, Semantic scholar can be tedious; therefore, a robust PMC, Dimension, Google scholar, and Semantic approach like text mining can be used for this process. scholar act as a storehouse of biological systems data. Text mining can be defined as a practical approach to A substantial amount of information can be recovered extracting biologically relevant information from the in a semi-structured form in the literature. The main growing amount of published literature. It comprises obstacle to large-scale analysis of this kind of data is three main tasks: information retrieval from relevant their highly unstructured and heterogenous format, documents, extraction of information of interest, and making it even harder to extract information contained data mining, which allows identifying new within the literature. Nonetheless, this information is associations among the extracted set of information. inherently helpful in a variety of genomics and Here, we show that a text mining approach can exploit systems biology contexts. For example, it is a standard large literature databases like PubMed and PMC to practice in the genomics community to manually extract genes/proteins related to biocorrosion by curate and extract literature-derived protein-protein Sulfate-reducing bacteria(SRB). The corrosion of metal due to microbial activity is known as biocorrosion or MIC(Microbial Induced corrosion). The primary class of bacteria associated with corrosion of metals in aquatic and terrestrial habitats is Sulfur Reducing Bacteria(SRB). Biocorrosion results from collaborative interactions between the metal surface, corrosion products, and bacterial cells and their metabolites. SRB are nonpathogenic and anaerobic bacteria, but SRB can act as a catalyst in the reduction reaction of sulfate to sulfide. It means they can make severe corrosion of metals in a water system by producing enzymes, which can accelerate the reduction of sulphate compounds to hydrogen sulfide. MIC is also known as metabolite corrosion or chemical microbially influenced corrosion (CMIC) owing to the generation of corrosive metabolite (hydrogen sulfide).In contrast, corrosion through direct withdrawal of electrons is called electrical microbial influenced corrosion. The presence of biofilm affects microbial corrosion; It is recognized that under the biofilm at the metal/biofilm interface, the concentrations of acidic metabolites are much greater, and their impact is amplified, leading to higher metal corrosion. It is also becoming apparent that one predominant mechanism of biocorrosion does not exist, and experimentally validating each of these theories can be laborious. Therefore, there is a need for an advanced technique for identifying genes and proteins of SRB involved in biocorrosion; this can help construct other biological processes, related pathways, and other processes associated with these genes.
Payal Thakur, Shailabh Rauniyar, Abhilash Kumar Tripathi, Priya Saxena, Vinoj Gopalakrishnan, Ram Nageena Singh, Mathew A. Olakunle, Etienne Z. Gnimpieba, Rajesh Kumar Sani
BIBM8
2021 Discovery of genes associated with sulfate-reducing bacteria biofilm using text mining and biological network analysis
abstract
Bacterial biofilms are complex surface attached communities of bacteria glued together by extracellular polymeric substance (EPS) matrix, secreted proteins, and extracellular DNAs [1]. Biofilm show reduced growth rates and metabolism. Biofilm formation is a survival mechanism that provides with better options compared to their planktonic counterparts. It impart bacterial communities stronger ability to grow in oligotrophic environments, greater access to nutritional resources, and enhanced syntropic interactions as well as greater tolerance towards environmental stress [2]. Biofilm play a detrimental role in many areas such as healthcare, food industry, water distribution systems, oil and gas industry etc. The composition of biofilm microbial community is varies depending on the environment in which it is formed. Biofilms are stratified formations where deeper layers maintain anoxic conditions. These anoxic niches promote the growth of certain specific groups, including sulfate reducing bacteria (SRB), that use the surface (usually metal) as resources for their survival.
Abhilash Kumar Tripathi, Priya Saxena, Payal Thakur, Shailabh Rauniyar, Vinoj Gopalakrishnan, Ram Nageena Singh, Mathew A. Olakunle, Etienne Z. Gnimpieba, Rajesh Kumar Sani
BIBM8
2017 NanoStringBioNet: Integrated R framework for bioscience knowledge discovery from NanoString nCounter data
abstract
The NanoString nCounter Analysis System is a medium-throughput gene expression quantification technique that is becoming increasingly popular in the fields of immunology and oncology due to its ease of use and sensitivity, particularly in the analysis of formalin-fixed paraffin embedded samples. Despite the growing interest in NanoString, systematic analysis frameworks for the reproducible analysis of nCounter data remain limited. NanoStringBioNet is a pair of R packages that form a semi-automatic, open source framework for integrative and systematic knowledge discovery from nCounter datasets. Using the NSData module, NanoStringBioNet preprocesses a raw NanoString dataset and stores it in Biobase format for sharability. Subsequently, the NSFunc module performs downstream analyses such as enrichment and network inference of stable differentially expressed gene clusters by leveraging existing data analysis tools and custom script.
Mariah M. Hoffman, Carrie J. Minette, Shanta M. Messerli, Ratan D. Bhardwaj, Etienne Z. Gnimpieba
BIBM5
2015 Life science data analysis workflow development using the bioextract server leveraging the iPlant collaborative cyberinfrastructure
abstract
Summary In order to handle the vast quantities of biological data gener6ated by high‐throughput experimental technologies, the BioExtract Server (bioextract.org) has leveraged iPlant Collaborative ( www.iplantcollaborative.org ) functionality to help address big data storage and analysis issues in the bioinformatics field. The BioExtract Server is a Web‐based, workflow‐enabling system that offers researchers a flexible environment for analyzing genomic data. It provides researchers with the ability to save a series of BioExtract Server tasks (e.g., query a data source, save a data extract, and execute an analytic tool) as a workflow and the opportunity for researchers to share their data extracts, analytic tools, and workflows with collaborators. The iPlant Collaborative is a community of researchers, educators, and students working to enrich science through the development of cyberinfrastructure—the physical computing resources, collaborative environment, virtual machine resources, and interoperable analysis software and data services—that are essential components of modern biology. The iPlant AGAVE Advanced Programming Interface, developed through the iPlant Collaborative, is a hosted, Software‐as‐a‐Service resource providing access to a collection of high performance computing and cloud resources. Leveraging AGAVE, the BioExtract Server gives researchers easy access to multiple high performance computers and delivers computation and storage as dynamically allocated resources via the Internet. © 2014 The Authors. Concurrency and Computation: Practice and Experience published by John Wiley & Sons Ltd.
Carol Lushbough, Etienne Z. Gnimpieba, Rion Dooley
Concurr. Comput. Pract. Exp.2
2013 BioExtract Server, a Web-based workflow enabling system, leveraging iPlant collaborative resources
abstract
In order to handle the vast quantities of biological data generated by high-throughput experimental technologies, the BioExtract Server (bioextract.org) has leveraged iPlant Collaborative (www.iplantcollaborative.org) functionality to help address big data storage and analysis issues in the bioinformatics field. The BioExtract Server is a Web-based, workflow-enabling system that offers researchers a flexible environment for analyzing genomic data. It provides researchers with the ability to save a series of BioExtract Server tasks (e.g. query a data source, save a data extract, and execute an analytic tool) as a workflow and the opportunity for researchers to share their data extracts, analytic tools and workflows with collaborators. The iPlant Collaborative is a community of researchers, educators, and students working to enrich science through the development of cyberinfrastructure - the physical computing resources, collaborative environment, virtual machine resources, and interoperable analysis software and data services - that are essential components of modern biology. The iPlant Agave API (Agave), developed through the iPlant Collaborative, is a hosted, Software-as-a-Service resource providing access to a collection of High Performance Computing (HPC) and Cloud resources [6]. Leveraging Agave, the BioExtract Server gives researchers easy access to multiple high performance computers and delivers computation and storage as dynamically allocated resources via the Internet.
Carol Lushbough, Etienne Z. Gnimpieba, Rion Dooley
CLUSTER2