Venkataramana Gadhamshetty

dblp:311/1009 · DBLP profile ↗
← Back
32ranked-venue papers
1as first author
31since 2021 · last 2025
0000-0002-8418-3515ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 32 · 1 first-author · 31 since 2021
YearPublicationVenuePosition
2025 Dimensional Reprojection and Sequential Modeling with Transformer Integration for Interesting Gene and Protein Recognition
abstract
Biomedical Named Entity Recognition (NER) is essential for structuring and extracting vital information from specialized medical texts, thereby improving research and diagnostics, particularly in emerging fields such as biofilm studies, where understanding gene-protein interactions is crucial for characterizing microbial communities and antimicrobial resistance mechanisms. This work presents an innovative hybrid architecture that integrates BioBERT's deep contextualization with HunFlair's sequential modeling capabilities through a novel dimensional reprojection mechanism. The architecture combines a specialized embedding layer (BioBERT dmis-lab/biobert-v1.1), optimized for understanding biomedical and biofilm-related contexts, with a sequential processing suite (BiLSTM-CRF) designed to accurately identify entities such as genes and proteins. A sophisticated dimensional reprojection layer (768$\boldsymbol{\rightarrow} \mathbf{4 2 9 6}$dimensions) employs a learned linear transformation to align and optimize information transfer between layers, enhancing overall performance without compromising structural coherence. We trained our model on 12 harmonized biomedical corpora containing gene and protein annotations related to biofilms and general biomedical domains, with fine-tuning using a learning rate of$5 \times 10^{-6}$over 10 epochs. Testing demonstrates that our model outperforms conventional architectures in biomedical named entity recognition, achieving F1 scores of 90.58 % on BC2GM ($\mathbf{+ 5. 4 3 \%}$compared to BioBERT), 90.70% on JNLPBA (+13.21% compared to HunFlair), 89.20 % on BioNLPCG$(+1.49 \%$compared to HunFlair), and 80.56 % on CRAFT ($+8.37 \%$compared to HunFlair). Precision scores reach 90.75% (BC2GM), 89.32% (JNLPBA), 89.03% (BioNLPCG), and 74.19% (CRAFT). Recall scores are particularly high: 90.41% (BC2GM), 92.12% (JNLPBA), 89.37% (BioNLPCG), and 88.13% (CRAFT), which is essential for comprehensive entity detection in biofilm research, where omitting a critical gene or protein could lead to gaps in understanding microbial mechanisms. Statistical validation confirms the significance of improvements$(\mathbf{p}<0.01)$. These results represent a notable advance over existing models, paving the way for future applications in extracting biofilm-related information from large text datasets and enabling the construction of biofilmspecific knowledge graphs. The code is publicly available to ensure reproducibility.
Alain Bertrand Bomgni, Feuzing Ntemma Donald, Shiva Aryal, Bichar Dip Shrestha Gurung, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2025 Minimal Features Subset Enabling Essential Gene Prediction Within and Between Organisms for Sulfate Reducing Bacteria Family
abstract
The identification of essential genes has garnered considerable attention from researchers in recent years. This process of identification uncovers minimal functional modules that enable the survival of an organism, making it of paramount importance in the fields of biomedicine and biotechnology. To address this challenging issue, computational methods have become increasingly utilized to complement experimental approaches, which tend to be intricate and costly. Various classifiers, based on the selection of feature sets, have been proposed and have shown promising results thus far. In this paper, leveraging 50 sulfate reducing bacteria (SRB) organisms - microbes frequently associated with biofilm formation, biofilm-driven corrosion, and complex microbial community dynamics; we aim to show that classifiers can achieve very good performance using only a minimal set of relevant features. Specifically, we demonstrate that classifier performance can be improved by considering minimal relevant features while taking into account the taxonomy of different organisms. A total of 37,500 features were generated from nucleotide and protein sequences of 41 SRB organisms to construct a machine learning model system aimed at predicting essential genes. Our feature engineering module identified 58 subsets of features. Through cross-validation, we achieved competitive intra-organism prediction performance. The best models obtained had an AUC of 0.99, precision of 0.99, recall of 0.99, and an F1-score of 0.99. Subsequently, this system was used to perform extra-organism (new organism not seen by the model) validation using nine left-out SRB organisms. The results obtained for these test organisms demonstrated the efficacy of our models with maximum precision, maximum recall, maximum F1-score, and maximum AUC equal to 0.99,$0.99,0.99$, and 0.97, respectively. Our approach has significantly outperformed previously proposed methods in terms of average metrics, indicating better generalization of the models. Finally, this approach allows researchers to evaluate the predicted result in the lab with fewer variables to consider in their experimental design.
Alain Bertrand Bomgni, Junior Basile Fofack, Shiva Aryal, Jerry Lonlac, Venkataramana Gadhamshetty, Etienne T. Gnimpieba
BIBM5
2025 Dual-Level Bayesian Predictive Modeling of Biological Nitrogen Fixation: The Effect of TPE Bayesian Optimization on Stacked Model Performance
abstract
Biological Nitrogen Fixation (BNF) is a key ecological process performed by diazotrophs, whose nitrogenase activity is often strictly regulated by environmental factors such as oxygen levels, metal availability, and carbon sources. Given this physiological complexity, accurately predicting nitrogenase activity from protein sequences remains a challenging problem in computational biology. Existing approaches such as Carmma and NFEmbed rely on features derived from large protein language models, yet their performance is often limited by suboptimal hyperparameter tuning. In this work, we introduce a dual-level predictive framework that integrates Bayesian Optimization using the Tree-structured Parzen Estimator (TPE) to systematically optimize both classification and regression models. For the classification task, our TPE-optimized XGBoost model achieves the highest overall performance, with an AUC of 0.9396, F1-score of 0.8675, accuracy of 0.8182, and recall of 0.9231, outperforming Carmma and matching or exceeding NFEmbed in key metrics. For the regression task, our dual-level TPEoptimized stacked SVR attains an$\mathrm{R}^{2}$of 0.6221 with reduced prediction error (MAE: 0.3044, RMSE: 0.5254), representing a substantial improvement over Carmma's SVR ($\mathrm{R}^{2} : 0.5572$) and offering competitive performance relative to NFEmbed. These findings demonstrate that TPE-based hyperparameter optimization—especially when applied at both the base- and meta-learner levels—significantly enhances the predictive reliability of complex machine learning architectures for BNF classification and activity estimation. This establishes systematic Bayesian optimization as a powerful strategy for advancing model accuracy, stability, and generalization in protein-sequence-based bioinformatics.
Alain Bertrand Bomgni, Dilane Sagueu Wakam, Ribot Fleury T. Ceskoutsé, Nick Klein, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2025 Heterogeneous Graph Network (HGN) for Binary Spatio-Chemical-Source Classification of Per- and Polyfluoroalkyl Substances (PFAS) Contamination: An Optimized Approach
abstract
Per- and Polyfluoroalkyl Substances (PFAS) pose a pervasive global environmental and public health challenge due to their persistence and widespread use. Accurate assessment of their contamination is crucial for risk assessment and remediation planning. Traditional machine learning (ML) and conventional geospatial models often fail to fully capture the complex, non-Euclidean interactions between chemical properties, environmental variables, geographical proximity (e.g., groundwater flow and atmospheric transport), and anthropogenic sources. This paper introduces an Optimized Heterogeneous Graph Network (HGN) designed as an alternative approach for the binary classification of PFAS contamination exceeding regulatory thresholds. The HGN models the assessment landscape as a graph where nodes represent sampling locations (spatial features), PFAS compounds (chemical features), and contaminant sources (e.g., industrial facilities), connected by weighted edges reflecting geographic distance, chemical similarity, and source associations. The model utilizes a multi-relational attention mechanism to differentially weigh node features and edge types, thus capturing intricate dependencies more effectively than existing ML approaches, such as ensemble models. Furthermore, we incorporate Bayesian Optimization (BO) for hyperparameter tuning, ensuring peak performance and model stability. Results show that the HGN framework offers a promising alternative in PFAS research by providing a superior representation of the underlying transport, exposure, and source mechanisms, potentially enabling efficient and actionable identification of contamination hotspots through binary outcomes.
Alain Bertrand Bomgni, Wilfried Loic Dnjomou Yonmba, Ribot Fleury T. Ceskoutsé, Nick Klein, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2025 A Convergence Roadmap for AI-Enabled, Omics-Guided Living-Interface Engineering. A National Science Foundation (NSF) National Research Traineeship (NRT) Initiative
abstract
Living-interface engineering seeks to understand and control what happens at the boundaries where materials and microbes meet regions that govern corrosion, filtration, implant safety, biomanufacturing, and water quality. These interfaces become far more complex when biofilms reshape surface chemistry, electron flow, and molecular transport. This paper presents the U.S. national roadmap for AI-enabled, omicsguided living-interface engineering, built on a decade of NSFand National Institute of Health (NIH)-supported research by the South Dakota team. The need for this roadmap is driven by industries whose interface-dependent failures and solutions represent hundreds of billions of dollars annually, with additional trillion-dollar opportunities emerging when microbiome science and AI converge. Our intention is to use this roadmap not only as the foundation for training students across South Dakota universities and collaborating institutions, but also to expand its reach through broader educational platforms such as IEEE workshops and ASCE EWRI short courses. The NSF NRT initiative formalizes this vision into an AI-enabled, digital-twin-supported training ecosystem for the next generation of convergent scientists.
Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM1
2025 Digital Twin COVID Tracker Using Wastewater Data: A Middle-School Led Study Within the U.S. NSF National Research Traineeship Program Framework
abstract
Wastewater infrastructure exists in every municipality across the United States and many other nations, offering a universal, non-invasive platform for community-level disease surveillance. Because viruses such as SARS-CoV-2 shed into wastewater days before symptoms appear, wastewater-based epidemiology (WBE) can provide crucial early-warning signals for public health. This study presents an AI-enabled Digital Twin prototype that predicts short-term COVID-19 trends using Center for Disease Control (CDC) wastewater viral activity data. Uniquely, this project was conceived and executed by middleschool first authors, highlighting the importance of early STEM engagement and intentional mentoring of young professionals on societally relevant environmental and health challenges. This work was conducted as part of our ongoing National Science Foundation (NSF) and National Institutes of Health (NIH) projects led by senior authors, which focus on convergence research and workforce development in AI-enabled, omics-guided living-interface engineering. Computational modeling, Jupyter Notebook workflow, GitHub integration, and cloud deployment were supported by graduate mentors, while system design, experimental logic, and interpretation were led by the student authors. The resulting platform, accessible through an interactive web app and QR-code interface, illustrates how guided, ageappropriate research experiences can empower middle-school students to explore wastewater informatics, digital twin concepts, machine learning, and epidemiological modeling. This work was also recognized with a 3rd-place award in the Sixth Grade Engineering Category at the 2025 High Plains Regional Science & Engineering Fair, highlighting both scientific merit and the broader impact of engaging middle-school students in societally relevant STEM research.
Isha Srikari Gadhamshetty, Jeanne Gnimpieba, Neha Sriveda Gadhamshetty, Shiva Aryal, Arun Kalaga, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2025 Functional Biofilm-Associated Microbial Markers in a Colorectal Cancer Cohort: Evidence for Quorum Sensing, Exopolysaccharide Production, and Biofilm-Driven Pathogenesis
abstract
Colorectal cancer (CRC) is increasingly recognized as a disease influenced by microbe-host interactions, particularly through mucosal biofilms that attach to the epithelial surface and alter the tumor microenvironment. While many microbiome studies focus on taxonomic composition, functional microbial pathways that drive biofilm formation may provide more reliable biomarkers and mechanistic insight. This study analyzes highvariance microbial enzyme signatures derived from a published CRC cohort to identify biofilm-associated genetic markers and to evaluate their roles in CRC biofilm development, quorum sensing signaling, and therapy response. The results highlight enrichment of LuxS/AI-2 quorum sensing enzymes, autoinducer synthases, and exopolysaccharide (EPS) biosynthesis pathways, suggesting a multi-species cooperative biofilm phenotype in CRC. These functional markers align with recent findings that Fusobacterium nucleatum-rich invasive biofilms are common in CRC [4]. The analysis supports four innovation claims: the identification of biofilm-associated functional markers in CRC; mechanistic pathways underlying CRC biofilm development; the impact of biofilm markers on therapy resistance; and additional diagnostic and therapeutic implications. Together, these results emphasize the value of functional biofilm biomarkers in understanding CRC progression and treatment outcomes.
Grace Goeden, Tuyen Do, Bichar Dip Shrestha Gurung, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM4
2025 Digital Twin Forecasting of Quorum-Sensing Associated Biofilm Microorganisms in Urban Wastewater Over a 30-Week Interval
abstract
Biofilms in wastewater systems contain dynamic microbial communities regulated by quorum sensing (QS), which governs adhesion, extracellular polymeric substance (EPS) production, stress tolerance, and developmental transitions. Forecasting QS-associated organisms is essential for anticipating biofilm formation and mitigating operational risks in wastewater infrastructure. In this study, we identified QS and biofilm-associated taxa present in a European wastewater metagenomic dataset: Vibrio harveyi, Vibrio parahaemolyticus, Pseudomonas aeruginosa, Escherichia coli, Salmonella typhi, Salmonella typhimurium, Bacillus subtilis, and Staphylococcus aureus. These organisms encode diverse QS and biofilm regulators across multiple bacterial lineages. From this broader QS-associated set, we developed an organism-specific, QS-aware digital twin focused solely on Pseudomonas, a dominant biofilm-forming genus in wastewater systems. Using the Q-net modeling framework, the model was trained on early-window observations and calibrated to perform short-horizon, end-point forecasting by predicting the final week of a 30-week interval. To ensure a stable conditioning window, the terminal signal was replicated prior to prediction. The digital twin demonstrated strong agreement between predicted and observed final-week abundance$(\mathbf{R}^{2}=0.98)$, highlighting its effectiveness for end-point microbial forecasting. This streamlined and interpretable framework supports rapid microbial surveillance and provides a scalable tool for biofilm-aware operational decision-making in wastewater systems.
Bichar Dip Shrestha Gurung, Tuyen Do, Shiva Aryal, Naina Maharjan, Dikshya Bhandari, Bipul Bhattarai, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM7
2025 Toward Autonomous Biofilm Engineering: A Vision for a Digital Twin-Based Lab Automation
abstract
Biofilm research is undergoing a significant transformation due to a growing need to address biological complexity, temporal dynamics, and environmental sensitivity. These growing needs demand experimental platforms that are far more autonomous, scalable, and computationally integrated than current microbiology lab workflows. This paper presents a vision for a digital twin-driven “cyber-physical laboratory” that combines modular robotics, automated imaging, an Internet of Things network, and AI-based phenotyping into a cohesive system for next-generation computational biofilm engineering. We outline an eight-component architecture comprising twins for physical, environmental, biological, protocol, data, AI/ML, control, and safety constraints. This architecture will enable virtual experimentation, predictive simulation, and self-correction, with the goal of interactive total lab automation and the “robot scientist”. Although we demonstrate early prototypes, the main objective of this paper is to articulate a forward-looking blueprint for how cyber-physical labs with digital twinning can fundamentally reshape biofilm research. Two use cases implementation in our lab shows the feasibility with up to 500 samples run per day (10,000 samples per week). By integrating real-time data streams with predictive biological models and closed-loop automated control, this architecture sets the stage for autonomous, reproducible, and computationally guided biofilm experimentation. We propose this platform as a foundational vision for the future of computational biofilm engineering.
Bichar Dip Shrestha Gurung, Timothy W. Hartman, Tuyen Do, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM4
2025 Development and OMICs Evaluation of Antibiofilm Material Using Hexanoic-Anhydride Modified Chitosan Electrospun Fibers Embedded with Silver Nanoparticles for Antibacterial Effects
abstract
Antimicrobial wound-healing materials are essential in clinical settings because wound infections present a significant challenge due to the attack of pathogenic bacteria, which can rapidly lead to biofilm formation and delayed healing. This work focused on developing the Ag-containing HexanoicAnhydride functionalized chitosan-based electrospun membrane nanocomposite (HA-CEMN) material to control bacterial growth. The Ag-containing HA-CEMN material was developed via electrospinning, and its surface was modified with hexanoic anhydride, and finally, the silver nanoparticles were generated within the nanofiber via an in-situ technique. The prepared Agcontaining HA-CEMN material offers a hydrated microenvironment around the wound area, promoting tissue regeneration. The Ag-containing HA-CEMN material was confirmed with Fourier Transform Infrared Spectroscopy, X-ray diffraction, Scanning electron microscopy, and Thermogravimetric analysis using. Finally, the antimicrobial properties of the prepared Ag-containing HA-CEMN material were studied via the zone of inhibition assay against Gram-negative and Gram-positive bacteria, followed with a mechanistic validation from OMICs data mining. The antimicrobial results demonstrated that the prepared materials show significant bacterial growth suppression towards the selected bacteria. The preliminary results highlight the potential of hexanoic-anhydride-modified Ag-containing HA-CEMN as a next-generation wound dressing with integrated antibacterial, antibiofilm, and regenerative properties.
Tippabattini Jayaramudu, Etienne Z. Gnimpieba, Venkataramana Gadhamshetty
BIBM3
2025 AI-Driven Sensor-Omics Framework for Detecting Antibiotic Resistance in Environmental Engineering Systems
abstract
The proliferation of antibiotic resistance genes (ARGs) in aquatic environments represents a critical One Health challenge, posing interconnected risks to human, animal, and ecosystem health. The spread of ARGs through surface waters can facilitate the emergence of multidrug-resistant pathogens, with far-reaching economic, social, and public health consequences. Traditional surveillance methods for ARG detection, including culture-based assays and molecular techniques such as PCR or metagenomics, are often limited by high costs, extended processing times, technical complexity, and regulatory constraints, hindering timely and comprehensive monitoring. To overcome these challenges, this study introduces the Hybrid Gen-erative-Discriminative Q-Network (Hybrid GDQ-net), a novel deep learning framework designed to integrate heterogeneous data from omics and sensor sources. The framework features a generative module that reconstructs latent representations from incomplete or noisy omics data, capturing hidden patterns that enhance predictive modeling. These latent features are then combined with a lightweight Q-network layer, which learns optimal classification policies through a reward-based training process, improving the accuracy and reliability of ARG detection across diverse environmental conditions. Evaluation of the model demonstrates strong predictive performance, achieving R2values of 0.81 when integrating omics and sensor data, 0.76 using only sensor data, and 0.73 with reduced sequencing coverage (60% of the dataset). These results highlight the ability of the Hybrid GDQ-net to maintain high accuracy even under limited or single-source data conditions, offering a scalable, time-efficient, and cost-effective solution for environmental monitoring. By enabling rapid and reliable ARG detection, this framework supports informed decision-making for water quality management, public health protection, and targeted interventions to mitigate the spread of antibiotic resistance in aquatic ecosystems.
Somayeh Falahati Khanaman, Etienne Z. Gnimpieba, Mengistu Geza Nisrani, Venkataramana Gadhamshetty
BIBM4
2025 Computational Validation of AlphaFold 3 as a Design Engine for Next Generation Aptamer Therapeutics
abstract
The advent of AlphaFold 3 (AF3) marks a pivotal shift in computational structural biology by extending deep learning capabilities beyond protein folding to encompass complex nucleic acid interactions. However, the reliability of this model in predicting the conformational dynamics of single-stranded oligonucleotides for therapeutic applications remains a critical area of investigation. This study rigorously benchmarks AF3 by evaluating its predictive fidelity across a curated dataset of high-affinity aptamers designed for two distinct pathological microenvironments: bacterial biofilms and solid tumors. We assessed the model's ability to resolve aptamer-target interfaces for biofilm disruption, specifically analyzing sequences that inhibit flagellar motility (Flagellin), block methicillin resistance mechanisms mediated by Penicillin-Binding Protein 2a (PBP2a), and disrupt glucan-mediated adhesion via Glucan-Binding Protein C (GbpC). Parallel evaluations were conducted on oncological aptamers designed to suppress angiogenesis via Platelet-Derived Growth Factor subunit B (PDGF-B) pathways and recognize diagnostic biomarkers such as Cancer Antigen 125 (CA-125) and Carcinoembryonic Antigen (CEA). By comparing predicted models against experimental baselines using Root-Mean-Square Deviation (RMSD), our results indicate that AF3 demonstrates high accuracy for rigid protein-aptamer complexes but exhibits variability when modeling interactions with small-molecule targets or functionalized conjugates. These findings establish AF3 as a potent hypothesis-generation engine for computational drug discovery that can significantly accelerate the design pipeline for next-generation nanotherapeutics.
Manish Rayamajhi, Shiva Aryal, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM3
2025 Emerging Biofilm Marker with Thin-Film (CuO) Surface Engineering for Applications in Electrochemical Sensors
abstract
Water quality monitoring strategies are challenged by sensor degradation caused by biofouling, corrosion, and fluctuating environmental conditions. Exposure to natural aqueous environments also leads to biofilm formation on the electrode surface, which alters its physicochemical properties and hinders charge transfer, thereby compromising sensor performance. These challenges hinder the accuracy, stability, and operational lifetime of electrochemical sensors deployed for continuous, in situ detection of contaminant. Developing non-invasive, thinfilm protective coatings that resist fouling while preserving electrochemical activity is therefore essential for reliable and sustainable water quality monitoring. Surface and structural characterizations through atomic force microscopy, Raman spectroscopy, and energy-dispersive X-ray spectroscopy confirmed coating uniformity and integrity. The detection of initial attachment during the biofilm formation, especially the quorum sensing biomarkers released is envisioned to provide insights on the inhibition of the fouling. The analysis of biofouling on the thin film coatings modified with Copper oxide (CuO) has been presented in this work. Ongoing efforts combine microscopy, spectroscopy, electrochemical, and Omics methods to assess fouling resistance and long-term signal stability. This study discusses the pros and cons of engineered thin-film coatings in enabling durable, high-performance sensors for next-generation environmental monitoring and intelligent water infrastructure systems. Future studies will integrate multi-omic information with materials informatics to elucidate foulant/microbe-surface interactions, optimize coating architectures, and enhance predictive modeling of fouling processes. We envision translating these thin-film sensor technologies beyond environmental systems toward biomedical and health monitoring platforms, where biointerface stability and selective sensing are equally critical.
Pawan Kumar Sapkota, Shiva Aryal, Bharat Jasthi, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM4
2025 A Physics-First, Physics-Gated AI Framework for Verifiable Biofilm Sensing
abstract
The Department of Defense (DoD) faces a significant challenge in the Verification and Validation (V&V) of non-deterministic, opaque AI, limiting the deployment of autonomous sensors. This challenge is acute in complex bio-electrochemical systems like biofilms. Current AI approaches for Electrochemical Impedance Spectroscopy (EIS) are bifurcated: (1) computationally intensive, non-interpretable neural networks, or (2) simplified models reliant on a priori ECM selection. We introduce the “Axiomatic-First Framework,” a three-stage methodology that converts first-principles physics into an auditable edge agent. Stage-1 establishes foundational integrity by using Molecular Dynamics (MD) as a physical constraint to model Nernst-Planck and Butler-Volmer physics, generating a “first-principles” dataset. Stage-2 trains a Physics-Informed Neural Network (PINN) “Teacher” whose loss function enforces the governing PDE residuals, yielding calibrated targets for an Equivalent-Circuit Model (ECM) vector$\theta=\left[R_{c t}, C_{d l}, Z_{W}\right]$. Stage-3 deploys a dual-mode “Student” AI: (i) a low-power “Physics-Gated Monitor” and (ii) an on-demand Spiking Neural Network (SNN) “Translator” that recovers$\theta$with uncertainty quantification. This architecture provides verifiable parameter recovery and auditable alarms, preserving low-power duty cycling.
Caine K. Shagla, Manoj Tripathi, Bichar Dip Shrestha Gurung, Etienne Z. Gnimpieba, Venkataramana Gadhamshetty
BIBM5
2025 Digital Biomonitoring of Microbial Communities: a Phenotype Feature-driven AutoML Framework for Bacteria Classification
abstract
Accurate identification of microbial species is essential for monitoring dynamic biological systems and guiding rapid industrial or clinical interventions. Conventional microbiological and genomic methods present an operational trade-off: traditional microscopy is subjective and slow, while high-fidelity sequencing lacks the speed and cost-effectiveness required for high-throughput, real-time screening. This study addresses the resulting “minimal data paradox,” where achieving high specieslevel accuracy requires data volumes that are not available in constrained environments. We introduce a rigorous benchmark of five state-of-the-art Convolutional Neural Networks (CNNs) to evaluate the feasibility of high-accuracy Microbe Species Prediction across scenarios of extreme data scarcity (using only 8 and 20 samples per class). Using transfer learning (TL) and synthetic data augmentation in 23 microbial species, we found that EfficientNetB0 consistently achieved the highest precision and resilience. In the 20 samples/class test, EfficientNetB0 reached 91.30% accuracy and an Area Under the Curve (AUC) of 0.998. Even under the extreme constraint of 8 samples/class, it maintained a superior validation accuracy of 86.96%, demonstrating profound resilience where other modern architectures failed. This benchmark establishes EfficientNetB0 as the optimal, lightweight model for subsequent high-throughput, in-lab operational validation.
Rupesh Kumar Yadav, Shiva Aryal, Bichar Dip Shrestha Gurung, Dikshya Bhandari, Graham Hartman, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2024 AIM-HKR: AI-Driven Molecular Retrosynthesis Using Heterogeneous Knowledge Representations
abstract
The synthesis of small molecules is a crucial task across multiple scientific domains, including drug discovery, materials science, and sustainable chemistry. As AI and Machine Learning (ML) continue to advance, these technologies offer transformative potential in molecular synthesis, enhancing efficiency and expanding synthetic possibilities. In particular, small molecule synthesis has profound implications for creating novel therapeutic agents, high-performance materials, and greener chemical processes, but a major challenge remains in designing efficient synthetic routes for target molecules. Retrosynthesis, the process of mapping out synthetic pathways by working backwards from a target molecule, represents a vital step; however, traditional retrosynthesis methods often struggle to predict complex or novel transformations. To address this limitation, we introduce AIM-HKR: AI-Driven Molecular Retrosynthesis Using Heterogeneous Knowledge Representations, a model that leverages graph neural network techniques to enhance retrosynthetic predictions. AIM-HKR integrates information from heterogeneous knowledge graphs, capturing the intricate relationships and analogical reasoning required for retrosynthesis. This model generates type-specific embeddings that reflect both network topology and semantic connections across different entity types, such as molecules, reactions, and functional groups. AIM-HKR’s unique capacity to leverage heterogeneous graphs allows it to propose synthetic pathways that extend beyond established precedents, enabling predictions of chemically feasible but previously uncharted transformations. We believe AIM-HKR has the potential to significantly advance molecular synthesis through AI-driven retrosynthesis, establishing a new paradigm in AI-assisted chemistry.
Alain Bertrand Bomgni, Ribot Fleury T. Ceskoutsé, Kevin Jordan Njike Njingang, Thomas Bouétou Bouétou, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2024 Enhancing Sulphate Reducing Bacteria Clustering by Implementing UMAP on Raman Spectroscopy Data
Md Hasanur Rahman 0004, Bichar Dip Shrestha Gurung, Manoj Tripathi, Alan Dalton, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2024 Biofilm marker discovery with cloud-based dockerized metagenomics analysis of microbial communities
abstract
In an environment, microbes often work in communities to achieve most of their essential functions, including the production of essential nutrients. Microbial biofilms are communities of microbes that attach to a nonliving or living surface by embedding themselves into a self-secreted matrix of extracellular polymeric substances. These communities work together to enhance their colonization of surfaces, produce essential nutrients, and achieve their essential functions for growth and survival. They often consist of diverse microbes including bacteria, viruses, and fungi. Biofilms play a critical role in influencing plant phenotypes and human microbial infections. Understanding how these biofilms impact plant health, human health, and the environment is important for analyzing genotype-phenotype-driven rule-of-life functions. Such fundamental knowledge can be used to precisely control the growth of biofilms on a given surface. Metagenomics is a powerful tool for analyzing biofilm genomes through function-based gene and protein sequence identification (functional metagenomics) and sequence-based function identification (sequence metagenomics). Metagenomic sequencing enables a comprehensive sampling of all genes in all organisms present within a biofilm sample. However, the complexity of biofilm metagenomic study warrants the increasing need to follow the Findability, Accessibility, Interoperability, and Reusable (FAIR) Guiding Principles for scientific data management. This will ensure that scientific findings can be more easily validated by the research community. This study proposes a dockerized, self-learning bioinformatics workflow to increase the community adoption of metagenomics toolkits in a metagenomics and meta-transcriptomics investigation. Our biofilm metagenomics workflow self-learning module includes integrated learning resources with an interactive dockerized workflow. This module will allow learners to analyze resources that are beneficial for aggregating knowledge about biofilm marker genes, proteins, and metabolic pathways as they define the composition of specific microbial communities. Cloud and dockerized technology can allow novice learners-even those with minimal knowledge in computer science-to use complicated bioinformatics tools. Our cloud-based, dockerized workflow splits biofilm microbiome metagenomics analyses into four easy-to-follow submodules. A variety of tools are built into each submodule. As students navigate these submodules, they learn about each tool used to accomplish the task. The downstream analysis is conducted using processed data obtained from online resources or raw data processed via Nextflow pipelines. This analysis takes place within Vertex AI's Jupyter notebook instance with R and Python kernels. Subsequently, results are stored and visualized in Google Cloud storage buckets, alleviating the computational burden on local resources. The result is a comprehensive tutorial that guides bioinformaticians of any skill level through the entire workflow. It enables them to comprehend and implement the necessary processes involved in this integrated workflow from start to finish. This manuscript describes the development of a resource module that is part of a learning platform named "NIGMS Sandbox for Cloud-based Learning" https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [1] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses.
Etienne Z. Gnimpieba, Timothy W. Hartman, Tuyen Do, Jessica Zylla, Shiva Aryal, Samuel J. Haas, Diing D. M. Agany, Bichar Dip Shrestha Gurung, Valena Doe, Zelaikha B. Yosufzai, Daniel Pan, Ross Campbell, Victor C. Huber, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Carol Lushbough
Briefings Bioinform.15
2023 Fine-tuning a pre-trained Transformers-based model for gene name entity recognition in biomedical text using a customized dataset: case of Desulfovibrio vulgaris Hildenborough
abstract
Gene Name Entity Recognition (NER) plays a crucial role in the realm of biomedical text mining by focusing on the identification and extraction of gene references from scientific literature. Recent advancements in the field-particularly the emergence of pre-trained transformer-based language models like BioBERT-have shown significant promise in the domain of biomedical NER. However, these models are often trained on existing, publicly available datasets, which may not fully capture the nuances of the domain or adequately cover less-studied genes. This study places its primary emphasis on fine-tuning BioBERT specifically for gene NER tasks. To address the limitations associated with current publicly available datasets, we have meticulously crafted a custom dataset. This dataset is thoughtfully constructed through the systematic collection and detailed annotation of a diverse range of biomedical literature from specialized sources. It intentionally includes genes that have been extensively researched, as well as those that have received limited attention in existing corpora. The fine-tuning process involves initializing the BioBERT model with pre-trained weights and then training it on our custom dataset using a sequence tagging approach. To enhance the model’s performance, we systematically explore various techniques, including data augmentation, entity-level features, and attention mechanisms. Additionally, we conduct rigorous hyperparameter optimization to maximize the model’s accuracy, precision, and recall in gene mention recognition. We thoroughly evaluate the performance of the fine-tuned BioBERT model through a comprehensive set of cross-validation experiments. The results highlight the effectiveness of our tailored dataset in enhancing BioBERT’s performance in gene NER. The fine-tuned model achieves impressive F1 scores, precision, and recall, specifically 0.96, 0.95, and 0.98, surpassing previous models when it comes to recognizing the dvu (Desulfovibrio vulgaris Hildenborough) gene. In conclusion, this study underscores the pivotal role of domain-specific, custom datasets in gene NER tasks. It also highlights the significant potential in fine-tuning pre-trained language models such as BioBERT to enhance gene mention recognition.
Alain Bertrand Bomgni, Dialo Abdala, Bichar Dip Shrestha Gurung, Marcellin Nkenlifack, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2023 NLPADADE: Leveraging Natural Language Processing for Automated Detection of Adverse Drug Effects
abstract
Pharmacovigilance is a systematic and scientifically rigorous discipline that assumes responsibility for the safety of pharmaceuticals, with its primary objective being the mitigation of risks while optimizing the benefits associated with medication usage. This mission-critical undertaking plays an indispensable role in preserving public health. At its core, pharmacovigilance entails the methodical collection and proficient management of data pertaining to medication safety. Additionally, these activities encompass the vigilant scrutiny of data to detect emerging "signals" indicative of new or evolving safety concerns. The expert evaluation of this data facilitates well-informed decision-making regarding matters of drug safety. Furthermore, proactive risk management strategies are deployed to effectively mitigate potential associated risks. In the pursuit of proactive health protection, regulatory actions are swiftly executed. Concurrently, the World Health Organization (WHO) underscores the global importance of establishing a robust pharmacovigilance framework. It advocates for the establishment of a comprehensive pharmacovigilance system, defined as encompassing "the science and activities related to the detection, assessment, understanding, and prevention of adverse effects or any other problem related to drugs or any other healthcare product." In this dynamic landscape, a pivotal question arises: "How can the automated identification and extraction of references to diseases, medications, and adverse effects from clinical notes and biomedical literature be achieved?" Central to this discourse are adverse drug effects (ADEs), which present a formidable public health challenge, manifesting as a significant source of patient morbidity and mortality. To expedite the utilization of real-world data (RWD) for the enhancement of pharmacovigilance practices, our focus has gravitated towards the development of a high-performance natural language processing (NLP) model. This model aims to facilitate the rapid detection of potential ADEs linked to medications. Our innovative system, designed for the extraction of diseases and ADEs, leverages the synergy of an open-source NLP component system. The pinnacle of our achievement is the model obtained, which boasts a remarkable Score of 0.97 at step 1800. With a Precision of 1.00, a Recall of 0.9412, and an F-Score of 0.9697, our NLP model showcases its efficiency in extracting and identifying pertinent information from textual data. These results underscore the effectiveness of our approach in recognizing diseases and adverse effects related to medications, setting a new benchmark for pharmacovigilance practices.
Alain Bertrand Bomgni, Claude Epiphanie Mbotchack Ngale, Shiva Aryal, Marcellin Nkenlifack, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM5
2023 Self-Supervised Scribble-based Segmentation of Single Cells in Biofilms
abstract
Supervised deep learning techniques have demonstrated remarkable performance in segmentation tasks in various domains including medical, biomaterial, and bioengineering fields. The advancement in imaging technologies has enabled the gathering of large amounts of raw data. Nevertheless, acquiring fully annotated ground-truth datasets for supervised deep learning-based segmentation tasks has remained a challenge attributed to the availability of expert resources and the laborious task of pixel-level annotation. Self-supervised learning(SSL) techniques have gained the momentum to learn representations from unlabeled datasets while also preserving the domain-specific information over transfer learning techniques thereby overcoming the challenge of annotating huge datasets. Weakly supervised learning(WSL) techniques have gained attention to address the annotation challenge by using weak annotations such as scribbles, dot annotations, and bounding boxes instead of pixel-wise annotations. In this work, we leveraged the advantages of both SSL and WSL to perform segmentation of single cells in optical images of biofilms attached on single-layer graphene-coated copper substrates guided by scribble annotations. Conditional random fields are applied as a post-processing to further improve segmentation performance. The method showed promising performance in segmenting single cells on biofilm images with scribble annotations that make up only 40-50% of pixel-level annotations.
Vidya Bommanapally, Mahadevan Subramaniam, Suvarna Talluri, Venkataramana Gadhamshetty, Juneau Jones
BIBM4
2023 Computational methods for biofouling and corrosion-resistant graphene nanocomposites. A transdisciplinary approach
abstract
The current study addresses the critical issue of bacterial adhesion and proliferation on surfaces, particularly in industries such as petroleum, marine, textiles, healthcare, food, and water management, with global biofilm-associated costs exceeding $5 trillion. Here we propose a transdisciplinary approach with computational methods to bring together the strengths, particularly advanced knowledge and cutting edge technologies of life sciences, material science, nanotechnology, engineering, and machine learning, with a goal of addressing vexing challenges facing microbiologically influenced corrosion. The current study focuses on developing biofouling resistant graphene nanocomposites with a thin silver film on nickel substrates against sulfate reducing bacteria (SRB). SEM images of biofilm formation after 30 days of exposure to SRB are analyzed using a pretrained deep learning architecture and otsu thresholding for semi-automated microbial segmentation. A pretrained multi-task deep learning model cellpose and segment anything model are employed to extract bacterial regions from the images, providing precise quantitative analysis of SRB cell surface coverage on graphene nanocomposites. This computational model establishes the interrelation between corrosion findings and the antibiofouling performance of graphene nanocomposites on nickel surfaces.
Ramesh Devadig, Bichar Dip Shrestha Gurung, Etienne Z. Gnimpieba, Bharat Jasthi, Venkataramana Gadhamshetty
BIBM5
2023 Utilizing XGBoost for the Prediction of Material Corrosion Rates from Embedded Tabular Data using Large Language Model
abstract
Microbial corrosion, scientifically referred to as microbial-induced corrosion (MIC), constitutes a noteworthy and frequently underestimated concern within diverse industrial domains. This phenomenon manifests when microorganisms, including bacteria, archaea, and fungi, engage with structural materials, resulting in the degradation of infrastructure and equipment. The accurate prognostication of material microbial corrosion rates is of upmost importance in the formulation of proactive strategies for maintenance and corrosion control. In this study, a novel methodology is introduced, which harnesses the capabilities of XGBoost, an advanced gradient boosting algorithm, for the precise prediction of material microbial corrosion rates. This predictive process is facilitated by employing tabular data that is intricately embedded within a comprehensive large language models (LLMs). The integration of tabular data into the language model yields a sophisticated contextual comprehension of the data, thereby augmenting the model's precision by its aptitude to discern intricate relationships and semantic nuances intrinsic to the tabular data.
Tuyen Do, Bichar Dip Shrestha Gurung, Shiva Aryal, Anup Khanal, Sandeep Chataut, Venkataramana Gadhamshetty, Carol Lushbough, Etienne Z. Gnimpieba
BIBM6
2023 Transformer in Microbial Image Analysis: A Comparative Exploration of TransUNet, UNet, and DoubleUNet for SEM Image Segmentation
abstract
The advent of transformer-based architectures such as TransUNet has revolutionized image segmentation as this approach combines the strengths of transformers for capturing contextual information with convolutional neural networks (CNNs) for localized feature identification. Microbes, known for their complex behaviors, present challenges in various fields, especially biomedicine. Image segmentation is crucial for ana-lyzing microbes, allowing quantitative analysis, growth tracking, and understanding host-pathogen interactions. This study is dedicated to a comparative analysis of TransUNet alongside two other popular segmentation methods, UNet and DoubleUNet, in the context of segmenting scanning electron microscope (SEM) images of microbes on layered graphene-nickel specimens. The TransUNet architecture employs a pre-defined ResNet-50 and Vision Transformer (ViT) as the encoder and a custom-built decoder trained on SEM data of Oleidesulfovibrio alaskensis (OA-G20) exposed to graphene-nickel specimens for 30 days. Using the Intersection Over Union (IoU) score as a performance metric, we observed that TransUNet achieved a maximum IoU of 79.58%, DoubleUNet exhibited a maximum IoU of 76.28%, and UNet attained a maximum IoU of 72.38%. We believe that this comparative study of the segmentation approach is invaluable for selecting the best model for the practitioner as per need. This study is the first step in our aim of developing an end-to-end framework with automated model selection based on dataset characteristics for microbial image segmentation.
Bichar Dip Shrestha Gurung, Anup Khanal, Timothy W. Hartman, Tuyen Do, Sandeep Chataut, Carol Lushbough, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM7
2023 Machine Learning-Assisted Optical Detection of Multilayer Hexagonal Boron Nitride for Enhanced Characterization and Analysis
abstract
Biofilms are ubiquitous in aqueous environments, exerting significant influence on diverse surfaces, including metals prone to microbiologically influenced corrosion (MIC). This multifaceted phenomenon demands interdisciplinary collaborations to combat its far-reaching implications. In this context, our research delves into the intricate characterization of twodimensional (2D) materials, particularly hexagonal boron nitride (hBN), which is crucial for advancing corrosion prevention coatings. The nanoscale dimensions of 2D materials pose challenges in microstructural analysis and defect identification, necessitating labor-intensive traditional techniques. To address these complexities, we utilized two unsupervised machine learning models, namely, (a) K-means clustering, and (b) Gaussian Mixture Model (GMM), which enabled clear differentiation between multilayer hBN (MLhBN) and cracks. Our approach will streamline the characterization process and facilitate the extraction of thin layers with enhanced accuracy.
Md Hasanur Rahman 0004, Vidya Bommanapally, Dilanga Abeyrathna, Md Ashaduzzman, Manoj Tripathi, Mahzuzah Zahan, Mahadevan Subramaniam, Venkataramana Gadhamshetty
BIBM8
2023 Artificial Intelligence-Driven Image Analysis of Bacterial Cells and Biofilms
abstract
The current study explores an artificial intelligence framework for measuring the structural features from microscopy images of the bacterial biofilms. Desulfovibrio alaskensis G20 (DA-G20) grown on mild steel surfaces is used as a model for sulfate reducing bacteria that are implicated in microbiologically influenced corrosion problems. Our goal is to automate the process of extracting the geometrical properties of the DA-G20 cells from the scanning electron microscopy (SEM) images, which is otherwise a laborious and costly process. These geometric properties are a biofilm phenotype that allow us to understand how the biofilm structurally adapts to the surface properties of the underlying metals, which can lead to better corrosion prevention solutions. We adapt two deep learning models: (a) a deep convolutional neural network (DCNN) model to achieve semantic segmentation of the cells, (d) a mask region-convolutional neural network (Mask R-CNN) model to achieve instance segmentation of the cells. These models are then integrated with moment invariants approach to measure the geometric characteristics of the segmented cells. Our numerical studies confirm that the Mask-RCNN and DCNN methods are 227x and 70x faster respectively, compared to the traditional method of manual identification and measurement of the cell geometric properties by the domain experts.
Shankarachary Ragi, Jamison Duckworth, Kalimuthu Jawaharraj, Parvathi Chundi, Venkataramana Gadhamshetty
IEEE ACM Trans. Comput. Biol. Bioinform.6
2022 Leveraging Weak annotations for Deep learning tasks on Biofilm Images
abstract
With the growth of technology capturing huge amounts of image data has become possible including electron microscopic images. Deep learning techniques have been thus drastically improved for analysing images due to their huge availability. Deep learning has been applied in various domains including medical, biomaterial, engineering fields t o analyze complex images. However, supervised deep learning techniques require huge amounts of annotated images. Annotating the images, specifically pixel wise annotations for segmentation tasks could be overwhelming and requires expert resources. Weakly-supervised learning has been popularly employed in such scenarios where weak labels are used for segmentation purposes. Also, self-supervised learning techniques have greatly reduced the amount of labeled data required to train a model for any downstream task. In this study, we would employ self-supervised learning technique followed by scribble supervision for performing biofilm segmentation on optical images. Our initial classification results and the proposed method for segmentation using scribble annotation are provided in this paper.
Vidya Bommanapally, Md Ashaduzzman, Mahadevan Subramaniam, Suvarna Talluri, Venkataramana Gadhamshetty
BIBM5
2022 U-Net Based Image Segmentation Techniques for Development of Non-Biocidal Fouling-Resistant Ultra-Thin Two-dimensional (2D) Coatings
abstract
Bacterial adhesion to the metallic surfaces creates a complex biofilm network, resulting in many problems like corrosion and fouling. Precise quantitative analysis of the surface coverage of cells can be vital in decoding biofilm-related issues. We present a deep learning-based approach to automate the microbes’ segmentation from the Scanning Electron Microscope (SEM) images of biofilm developed on coated multilayer graphene nickel samples. We collected SEM images from multilayer graphene nickel exposed to Oleidesulfovibrioalaskensis (OA-G20) for 30 days. Then we manually annotated the microbes with the help of subject expertise and trained a deep-learning U-Net architecture. In order to deal with larger image sizes, we perform patched-based image training and techniques for predicting segmentation masks over the larger image. Intersection over Union (IOU) was calculated for the evaluation of the performance of the model. After training the image for multiple epochs and extracting the optimal model parameters from the learning curve, we were able to get 70.64% mean IOU score. The patched-based technique for image training and image inference showed promising output during the segmentation.
Bichar Dip Shrestha Gurung, Ramesh Devadig, Tuyen Do, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM4
2022 Using BASIN-ML for Machine Learning-Based Statistical Analysis and Reporting for Biofilm Datasets
abstract
Biological and biomedical microscope image (bioimage) comparison remains useful to approach many research challenges—from biofilms to human diseases. This powerful technology allows researchers to provide the community with a quick visual snapshot of varying experimental conditions. But a two-condition comparison still relies on a researcher’s eyes to draw conclusions despite the availability of multiple— often complex—digital image analysis tools. Our Bioimage Analysis, Statistic, and Comparison (BASIN) software provides an easy, objective, reproducible comparison leveraging inferential statistics to bridge image data analysis with other biomedical data modalities such as gene expression. Users have access to a machine learning module to assist with image segmentation using modern, trainable algorithms. BASIN also provides several key data points including images’ object counts, net and mean pixel intensities, net and mean object surface areas, plus a variety of other potentially useful data. Hypothesis testing is performed on mean object intensities and surface areas using the statistical power of the R programming language. These features allow BASIN to extend the current scope of image comparison. It gives researchers a multi-model knowledge about matters such as drug protein marker response, the significance of cell population changes, and changes in cell morphology. To improve BASIN’s accessibility and transparency we implemented it in R using Shiny framework and provided both an online trial version and a customizable offline version. We also have a batch version to run on datasets with hundreds of biomedical images. BASIN workflows consist of five core modules including image upload, feature extraction, statistical analysis, visualization, and report generation.
Sam Haas, Timothy W. Hartman, Bichar Dip Shrestha Gurung, Tuyen Do, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2021 GenNER - A highly scalable and optimal NER method for text-based gene and protein recognition
abstract
Nowadays, there are a large number of models in the scientific literature capable of recognizing and extracting gene mentions from a given text. Several data sets have been developed to facilitate the learning process of these models. However, very few models are able to increase their knowledge and performance progressively from new annotated text but also to take into account the granularity of the input text of the model. Our proposed solution, GenNER, is a method for recognizing gene/protein mentions from free text. GenNER relies on continuous learning and a text granularization algorithm as input to the model, which allows it to achieve better performance. Its evaluation process was done around BioCreative II annotated datasets; we obtained an average F1-score of 0.9704, which outperforms current methods.
Ernest Basile Fotseu Fotseu, Thierry Kongne Nembot, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba, Alain Bertrand Bomgni
BIBM4
2021 Automatic Extension of Medical Subject Headings (MeSH) Thesaurus to Emerging Research
abstract
The proliferation of information technology infrastructure in recent decades has allowed for unprecedented ease of access to centrally-aggregated scholarly literature and scientific knowledge. This massive aggregation of knowledge requires an information retrieval infrastructure, to include formalized ontologies, that is engineered with careful consideration. A number of domains benefit from the use of hierarchical controlled vocabularies, which may be used to provide a rich set of descriptive terms for characterizing entities in a consistent manner. There are clear benefits to the creation and maintenance of these ontologies: search and retrieval is made easier and analyses of the contained entities are enabled that would not otherwise be possible. However, there may be the opportunity to decrease the manual burden of ontology creation and maintenance with automated methods that leverage natural language processing and other computational techniques. This work presents an automated ontology creation methodology, adapted and expanded from prior work [1], that can produce a topic hierarchy from natural language and may be used to assist in the creation of a novel ontology or the expansion of existing ontologies. The effectiveness of the proposed method is studied using two examples: immunology, an established biomedical domain and a prominent topic in MeSH, and graphene, from the 2D materials domain with wide-ranging biomedical applications, which also has a sparse presence in MeSH
William Gasper, Dario Ghersi, Etienne Z. Gnimpieba, Venkataramana Gadhamshetty, Parvathi Chundi
BIBM5
2020 A Thrifty Annotation Generation Approach for Semantic Segmentation of Biofilms
abstract
Recent advances in semantic segmentation using deep learning methods have achieved promising results on several benchmark datasets. However, the primary challenge involved in such segmentation approaches is the availability of applicable training data. Since only experts are equipped to effectively annotate (or label) any available data for training semantic segmentation networks, the effort and cost involved can be considerable, especially for larger datasets. In this paper, we aim to address this problem by proposing a Thrifty Annotation Generation approach that records high performance on segmentation networks with minimal expert effort and cost (intervention). We present a deep active learning framework that combines the use of marker-controlled watershed (MC-WS) algorithm to generate pseudo labels for segmentation networks (U-Net) and active learning to significantly minimize effort and cost by selecting only the most impactful training data for labeling. We built the initial U-Net model by generating pseudo labels for the training data using MC-WS. We then make use of the uncertainty information (entropy) of each image provided by the U-Net to determine the most uncertain or effective images for expert labeling. We evaluated the TAG approach using the 2012 ISBI Challenge dataset for 2D segmentation and a novel Biofilm dataset. Our approach achieved promising segmentation accuracy (IoU) and classification accuracy with minimal expert intervention. The results of our experiments also indicate that the TAG approach can be generalized to achieve high-performance segmentation results on any dataset using minimal expert effort and cost.
Adithi D. Chakravarthy, Parvathi Chundi, Mahadevan Subramaniam, Shankarachary Ragi, Venkataramana Gadhamshetty
BIBE5