EDBT 2026 Demo / reviewers in the wild / expert
Shiva Aryal
dblp:368/2205
· DBLP profile ↗
11ranked-venue papers
0as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dimensional Reprojection and Sequential Modeling with Transformer Integration for Interesting Gene and Protein RecognitionabstractBiomedical Named Entity Recognition (NER) is essential for structuring and extracting vital information from specialized medical texts, thereby improving research and diagnostics, particularly in emerging fields such as biofilm studies, where understanding gene-protein interactions is crucial for characterizing microbial communities and antimicrobial resistance mechanisms. This work presents an innovative hybrid architecture that integrates BioBERT's deep contextualization with HunFlair's sequential modeling capabilities through a novel dimensional reprojection mechanism. The architecture combines a specialized embedding layer (BioBERT dmis-lab/biobert-v1.1), optimized for understanding biomedical and biofilm-related contexts, with a sequential processing suite (BiLSTM-CRF) designed to accurately identify entities such as genes and proteins. A sophisticated dimensional reprojection layer (768$\boldsymbol{\rightarrow} \mathbf{4 2 9 6}$dimensions) employs a learned linear transformation to align and optimize information transfer between layers, enhancing overall performance without compromising structural coherence. We trained our model on 12 harmonized biomedical corpora containing gene and protein annotations related to biofilms and general biomedical domains, with fine-tuning using a learning rate of$5 \times 10^{-6}$over 10 epochs. Testing demonstrates that our model outperforms conventional architectures in biomedical named entity recognition, achieving F1 scores of 90.58 % on BC2GM ($\mathbf{+ 5. 4 3 \%}$compared to BioBERT), 90.70% on JNLPBA (+13.21% compared to HunFlair), 89.20 % on BioNLPCG$(+1.49 \%$compared to HunFlair), and 80.56 % on CRAFT ($+8.37 \%$compared to HunFlair). Precision scores reach 90.75% (BC2GM), 89.32% (JNLPBA), 89.03% (BioNLPCG), and 74.19% (CRAFT). Recall scores are particularly high: 90.41% (BC2GM), 92.12% (JNLPBA), 89.37% (BioNLPCG), and 88.13% (CRAFT), which is essential for comprehensive entity detection in biofilm research, where omitting a critical gene or protein could lead to gaps in understanding microbial mechanisms. Statistical validation confirms the significance of improvements$(\mathbf{p}<0.01)$. These results represent a notable advance over existing models, paving the way for future applications in extracting biofilm-related information from large text datasets and enabling the construction of biofilmspecific knowledge graphs. The code is publicly available to ensure reproducibility. Alain Bertrand Bomgni, Feuzing Ntemma Donald, Shiva Aryal, Bichar Dip Shrestha Gurung, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 3 |
| 2025 | Minimal Features Subset Enabling Essential Gene Prediction Within and Between Organisms for Sulfate Reducing Bacteria FamilyabstractThe identification of essential genes has garnered considerable attention from researchers in recent years. This process of identification uncovers minimal functional modules that enable the survival of an organism, making it of paramount importance in the fields of biomedicine and biotechnology. To address this challenging issue, computational methods have become increasingly utilized to complement experimental approaches, which tend to be intricate and costly. Various classifiers, based on the selection of feature sets, have been proposed and have shown promising results thus far. In this paper, leveraging 50 sulfate reducing bacteria (SRB) organisms - microbes frequently associated with biofilm formation, biofilm-driven corrosion, and complex microbial community dynamics; we aim to show that classifiers can achieve very good performance using only a minimal set of relevant features. Specifically, we demonstrate that classifier performance can be improved by considering minimal relevant features while taking into account the taxonomy of different organisms. A total of 37,500 features were generated from nucleotide and protein sequences of 41 SRB organisms to construct a machine learning model system aimed at predicting essential genes. Our feature engineering module identified 58 subsets of features. Through cross-validation, we achieved competitive intra-organism prediction performance. The best models obtained had an AUC of 0.99, precision of 0.99, recall of 0.99, and an F1-score of 0.99. Subsequently, this system was used to perform extra-organism (new organism not seen by the model) validation using nine left-out SRB organisms. The results obtained for these test organisms demonstrated the efficacy of our models with maximum precision, maximum recall, maximum F1-score, and maximum AUC equal to 0.99,$0.99,0.99$, and 0.97, respectively. Our approach has significantly outperformed previously proposed methods in terms of average metrics, indicating better generalization of the models. Finally, this approach allows researchers to evaluate the predicted result in the lab with fewer variables to consider in their experimental design. Alain Bertrand Bomgni, Junior Basile Fofack, Shiva Aryal, Jerry Lonlac, Venkataramana Gadhamshetty, Etienne T. Gnimpieba |
BIBM | 3 |
| 2025 | Digital Twin COVID Tracker Using Wastewater Data: A Middle-School Led Study Within the U.S. NSF National Research Traineeship Program FrameworkabstractWastewater infrastructure exists in every municipality across the United States and many other nations, offering a universal, non-invasive platform for community-level disease surveillance. Because viruses such as SARS-CoV-2 shed into wastewater days before symptoms appear, wastewater-based epidemiology (WBE) can provide crucial early-warning signals for public health. This study presents an AI-enabled Digital Twin prototype that predicts short-term COVID-19 trends using Center for Disease Control (CDC) wastewater viral activity data. Uniquely, this project was conceived and executed by middleschool first authors, highlighting the importance of early STEM engagement and intentional mentoring of young professionals on societally relevant environmental and health challenges. This work was conducted as part of our ongoing National Science Foundation (NSF) and National Institutes of Health (NIH) projects led by senior authors, which focus on convergence research and workforce development in AI-enabled, omics-guided living-interface engineering. Computational modeling, Jupyter Notebook workflow, GitHub integration, and cloud deployment were supported by graduate mentors, while system design, experimental logic, and interpretation were led by the student authors. The resulting platform, accessible through an interactive web app and QR-code interface, illustrates how guided, ageappropriate research experiences can empower middle-school students to explore wastewater informatics, digital twin concepts, machine learning, and epidemiological modeling. This work was also recognized with a 3rd-place award in the Sixth Grade Engineering Category at the 2025 High Plains Regional Science & Engineering Fair, highlighting both scientific merit and the broader impact of engaging middle-school students in societally relevant STEM research. Isha Srikari Gadhamshetty, Jeanne Gnimpieba, Neha Sriveda Gadhamshetty, Shiva Aryal, Arun Kalaga, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 4 |
| 2025 | Digital Twin Forecasting of Quorum-Sensing Associated Biofilm Microorganisms in Urban Wastewater Over a 30-Week IntervalabstractBiofilms in wastewater systems contain dynamic microbial communities regulated by quorum sensing (QS), which governs adhesion, extracellular polymeric substance (EPS) production, stress tolerance, and developmental transitions. Forecasting QS-associated organisms is essential for anticipating biofilm formation and mitigating operational risks in wastewater infrastructure. In this study, we identified QS and biofilm-associated taxa present in a European wastewater metagenomic dataset: Vibrio harveyi, Vibrio parahaemolyticus, Pseudomonas aeruginosa, Escherichia coli, Salmonella typhi, Salmonella typhimurium, Bacillus subtilis, and Staphylococcus aureus. These organisms encode diverse QS and biofilm regulators across multiple bacterial lineages. From this broader QS-associated set, we developed an organism-specific, QS-aware digital twin focused solely on Pseudomonas, a dominant biofilm-forming genus in wastewater systems. Using the Q-net modeling framework, the model was trained on early-window observations and calibrated to perform short-horizon, end-point forecasting by predicting the final week of a 30-week interval. To ensure a stable conditioning window, the terminal signal was replicated prior to prediction. The digital twin demonstrated strong agreement between predicted and observed final-week abundance$(\mathbf{R}^{2}=0.98)$, highlighting its effectiveness for end-point microbial forecasting. This streamlined and interpretable framework supports rapid microbial surveillance and provides a scalable tool for biofilm-aware operational decision-making in wastewater systems. Bichar Dip Shrestha Gurung, Tuyen Do, Shiva Aryal, Naina Maharjan, Dikshya Bhandari, Bipul Bhattarai, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 3 |
| 2025 | Computational Validation of AlphaFold 3 as a Design Engine for Next Generation Aptamer TherapeuticsabstractThe advent of AlphaFold 3 (AF3) marks a pivotal shift in computational structural biology by extending deep learning capabilities beyond protein folding to encompass complex nucleic acid interactions. However, the reliability of this model in predicting the conformational dynamics of single-stranded oligonucleotides for therapeutic applications remains a critical area of investigation. This study rigorously benchmarks AF3 by evaluating its predictive fidelity across a curated dataset of high-affinity aptamers designed for two distinct pathological microenvironments: bacterial biofilms and solid tumors. We assessed the model's ability to resolve aptamer-target interfaces for biofilm disruption, specifically analyzing sequences that inhibit flagellar motility (Flagellin), block methicillin resistance mechanisms mediated by Penicillin-Binding Protein 2a (PBP2a), and disrupt glucan-mediated adhesion via Glucan-Binding Protein C (GbpC). Parallel evaluations were conducted on oncological aptamers designed to suppress angiogenesis via Platelet-Derived Growth Factor subunit B (PDGF-B) pathways and recognize diagnostic biomarkers such as Cancer Antigen 125 (CA-125) and Carcinoembryonic Antigen (CEA). By comparing predicted models against experimental baselines using Root-Mean-Square Deviation (RMSD), our results indicate that AF3 demonstrates high accuracy for rigid protein-aptamer complexes but exhibits variability when modeling interactions with small-molecule targets or functionalized conjugates. These findings establish AF3 as a potent hypothesis-generation engine for computational drug discovery that can significantly accelerate the design pipeline for next-generation nanotherapeutics. Manish Rayamajhi, Shiva Aryal, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 2 |
| 2025 | Emerging Biofilm Marker with Thin-Film (CuO) Surface Engineering for Applications in Electrochemical SensorsabstractWater quality monitoring strategies are challenged by sensor degradation caused by biofouling, corrosion, and fluctuating environmental conditions. Exposure to natural aqueous environments also leads to biofilm formation on the electrode surface, which alters its physicochemical properties and hinders charge transfer, thereby compromising sensor performance. These challenges hinder the accuracy, stability, and operational lifetime of electrochemical sensors deployed for continuous, in situ detection of contaminant. Developing non-invasive, thinfilm protective coatings that resist fouling while preserving electrochemical activity is therefore essential for reliable and sustainable water quality monitoring. Surface and structural characterizations through atomic force microscopy, Raman spectroscopy, and energy-dispersive X-ray spectroscopy confirmed coating uniformity and integrity. The detection of initial attachment during the biofilm formation, especially the quorum sensing biomarkers released is envisioned to provide insights on the inhibition of the fouling. The analysis of biofouling on the thin film coatings modified with Copper oxide (CuO) has been presented in this work. Ongoing efforts combine microscopy, spectroscopy, electrochemical, and Omics methods to assess fouling resistance and long-term signal stability. This study discusses the pros and cons of engineered thin-film coatings in enabling durable, high-performance sensors for next-generation environmental monitoring and intelligent water infrastructure systems. Future studies will integrate multi-omic information with materials informatics to elucidate foulant/microbe-surface interactions, optimize coating architectures, and enhance predictive modeling of fouling processes. We envision translating these thin-film sensor technologies beyond environmental systems toward biomedical and health monitoring platforms, where biointerface stability and selective sensing are equally critical. Pawan Kumar Sapkota, Shiva Aryal, Bharat Jasthi, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 2 |
| 2025 | Digital Biomonitoring of Microbial Communities: a Phenotype Feature-driven AutoML Framework for Bacteria ClassificationabstractAccurate identification of microbial species is essential for monitoring dynamic biological systems and guiding rapid industrial or clinical interventions. Conventional microbiological and genomic methods present an operational trade-off: traditional microscopy is subjective and slow, while high-fidelity sequencing lacks the speed and cost-effectiveness required for high-throughput, real-time screening. This study addresses the resulting “minimal data paradox,” where achieving high specieslevel accuracy requires data volumes that are not available in constrained environments. We introduce a rigorous benchmark of five state-of-the-art Convolutional Neural Networks (CNNs) to evaluate the feasibility of high-accuracy Microbe Species Prediction across scenarios of extreme data scarcity (using only 8 and 20 samples per class). Using transfer learning (TL) and synthetic data augmentation in 23 microbial species, we found that EfficientNetB0 consistently achieved the highest precision and resilience. In the 20 samples/class test, EfficientNetB0 reached 91.30% accuracy and an Area Under the Curve (AUC) of 0.998. Even under the extreme constraint of 8 samples/class, it maintained a superior validation accuracy of 86.96%, demonstrating profound resilience where other modern architectures failed. This benchmark establishes EfficientNetB0 as the optimal, lightweight model for subsequent high-throughput, in-lab operational validation. Rupesh Kumar Yadav, Shiva Aryal, Bichar Dip Shrestha Gurung, Dikshya Bhandari, Graham Hartman, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 2 |
| 2024 | NFB-Checker: An AI/ML-Powered Microorganism Nitrogen Fixation Susceptibility Prediction from Gene CollectionabstractNitrogen fixation is a crucial process involved in many aspects of the ecological life cycle. However, few organismsparticularly kingdom bacteria-are known to fix N2in a form that could be used in agribusiness. The ability to identify an organism as an N2fixer is a novel case for more identification and application. This study employs functional genes and an XGBoost machine learning model to predict an organism's ability to fix nitrogen, focusing on key orthologs like nifH, nifD, and nifK, essential in the nitrogenase enzyme complex. The model, trained on a dataset of 923 N2fixers and 981 non-fixers, achieved an accuracy of 92.6% and an F1 score of 0.92. This research highlights the potential of machine learning in identifying genetic markers for N2fixation, providing an efficient alternative to traditional methods, and offers new insights for ecological and agricultural research. The inclusion of specific orthologs enhances the model's predictive accuracy, demonstrating the importance of targeted genetic markers in computational biology. Tuyen Do, Shiva Aryal, Bichar Dip Shrestha Gurung, Diing D. M. Agany, Nick Klein, Ruanbao Zhou, Rajesh Kumar Sani, Etienne Z. Gnimpieba |
BIBM | 2 |
| 2024 | Biofilm marker discovery with cloud-based dockerized metagenomics analysis of microbial communitiesabstractIn an environment, microbes often work in communities to achieve most of their essential functions, including the production of essential nutrients. Microbial biofilms are communities of microbes that attach to a nonliving or living surface by embedding themselves into a self-secreted matrix of extracellular polymeric substances. These communities work together to enhance their colonization of surfaces, produce essential nutrients, and achieve their essential functions for growth and survival. They often consist of diverse microbes including bacteria, viruses, and fungi. Biofilms play a critical role in influencing plant phenotypes and human microbial infections. Understanding how these biofilms impact plant health, human health, and the environment is important for analyzing genotype-phenotype-driven rule-of-life functions. Such fundamental knowledge can be used to precisely control the growth of biofilms on a given surface. Metagenomics is a powerful tool for analyzing biofilm genomes through function-based gene and protein sequence identification (functional metagenomics) and sequence-based function identification (sequence metagenomics). Metagenomic sequencing enables a comprehensive sampling of all genes in all organisms present within a biofilm sample. However, the complexity of biofilm metagenomic study warrants the increasing need to follow the Findability, Accessibility, Interoperability, and Reusable (FAIR) Guiding Principles for scientific data management. This will ensure that scientific findings can be more easily validated by the research community. This study proposes a dockerized, self-learning bioinformatics workflow to increase the community adoption of metagenomics toolkits in a metagenomics and meta-transcriptomics investigation. Our biofilm metagenomics workflow self-learning module includes integrated learning resources with an interactive dockerized workflow. This module will allow learners to analyze resources that are beneficial for aggregating knowledge about biofilm marker genes, proteins, and metabolic pathways as they define the composition of specific microbial communities. Cloud and dockerized technology can allow novice learners-even those with minimal knowledge in computer science-to use complicated bioinformatics tools. Our cloud-based, dockerized workflow splits biofilm microbiome metagenomics analyses into four easy-to-follow submodules. A variety of tools are built into each submodule. As students navigate these submodules, they learn about each tool used to accomplish the task. The downstream analysis is conducted using processed data obtained from online resources or raw data processed via Nextflow pipelines. This analysis takes place within Vertex AI's Jupyter notebook instance with R and Python kernels. Subsequently, results are stored and visualized in Google Cloud storage buckets, alleviating the computational burden on local resources. The result is a comprehensive tutorial that guides bioinformaticians of any skill level through the entire workflow. It enables them to comprehend and implement the necessary processes involved in this integrated workflow from start to finish. This manuscript describes the development of a resource module that is part of a learning platform named "NIGMS Sandbox for Cloud-based Learning" https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [1] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses. Etienne Z. Gnimpieba, Timothy W. Hartman, Tuyen Do, Jessica Zylla, Shiva Aryal, Samuel J. Haas, Diing D. M. Agany, Bichar Dip Shrestha Gurung, Valena Doe, Zelaikha B. Yosufzai, Daniel Pan, Ross Campbell, Victor C. Huber, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Carol Lushbough |
Briefings Bioinform. | 5 |
| 2023 | NLPADADE: Leveraging Natural Language Processing for Automated Detection of Adverse Drug EffectsabstractPharmacovigilance is a systematic and scientifically rigorous discipline that assumes responsibility for the safety of pharmaceuticals, with its primary objective being the mitigation of risks while optimizing the benefits associated with medication usage. This mission-critical undertaking plays an indispensable role in preserving public health. At its core, pharmacovigilance entails the methodical collection and proficient management of data pertaining to medication safety. Additionally, these activities encompass the vigilant scrutiny of data to detect emerging "signals" indicative of new or evolving safety concerns. The expert evaluation of this data facilitates well-informed decision-making regarding matters of drug safety. Furthermore, proactive risk management strategies are deployed to effectively mitigate potential associated risks. In the pursuit of proactive health protection, regulatory actions are swiftly executed. Concurrently, the World Health Organization (WHO) underscores the global importance of establishing a robust pharmacovigilance framework. It advocates for the establishment of a comprehensive pharmacovigilance system, defined as encompassing "the science and activities related to the detection, assessment, understanding, and prevention of adverse effects or any other problem related to drugs or any other healthcare product." In this dynamic landscape, a pivotal question arises: "How can the automated identification and extraction of references to diseases, medications, and adverse effects from clinical notes and biomedical literature be achieved?" Central to this discourse are adverse drug effects (ADEs), which present a formidable public health challenge, manifesting as a significant source of patient morbidity and mortality. To expedite the utilization of real-world data (RWD) for the enhancement of pharmacovigilance practices, our focus has gravitated towards the development of a high-performance natural language processing (NLP) model. This model aims to facilitate the rapid detection of potential ADEs linked to medications. Our innovative system, designed for the extraction of diseases and ADEs, leverages the synergy of an open-source NLP component system. The pinnacle of our achievement is the model obtained, which boasts a remarkable Score of 0.97 at step 1800. With a Precision of 1.00, a Recall of 0.9412, and an F-Score of 0.9697, our NLP model showcases its efficiency in extracting and identifying pertinent information from textual data. These results underscore the effectiveness of our approach in recognizing diseases and adverse effects related to medications, setting a new benchmark for pharmacovigilance practices. Alain Bertrand Bomgni, Claude Epiphanie Mbotchack Ngale, Shiva Aryal, Marcellin Nkenlifack, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 3 |
| 2023 | Utilizing XGBoost for the Prediction of Material Corrosion Rates from Embedded Tabular Data using Large Language ModelabstractMicrobial corrosion, scientifically referred to as microbial-induced corrosion (MIC), constitutes a noteworthy and frequently underestimated concern within diverse industrial domains. This phenomenon manifests when microorganisms, including bacteria, archaea, and fungi, engage with structural materials, resulting in the degradation of infrastructure and equipment. The accurate prognostication of material microbial corrosion rates is of upmost importance in the formulation of proactive strategies for maintenance and corrosion control. In this study, a novel methodology is introduced, which harnesses the capabilities of XGBoost, an advanced gradient boosting algorithm, for the precise prediction of material microbial corrosion rates. This predictive process is facilitated by employing tabular data that is intricately embedded within a comprehensive large language models (LLMs). The integration of tabular data into the language model yields a sophisticated contextual comprehension of the data, thereby augmenting the model's precision by its aptitude to discern intricate relationships and semantic nuances intrinsic to the tabular data. Tuyen Do, Bichar Dip Shrestha Gurung, Shiva Aryal, Anup Khanal, Sandeep Chataut, Venkataramana Gadhamshetty, Carol Lushbough, Etienne Z. Gnimpieba |
BIBM | 3 |