VLDB 2026 Research / reviewers in the wild / expert
Tuyen Do
dblp:337/4342
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0002-1817-4565ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Metagenomic Functional Signatures for Real-Time Predictive Monitoring of Cyanobacterial Bloom Assessment and Biofilm Marker Discovery in Managed Freshwater SystemsabstractHarmful algal blooms (HABs), driven by cyanobacteria such as Microcystis, pose a critical and escalating threat to global freshwater security. Standard environmental monitoring often provides reactive rather than predictive data, underestimating the true ecological state after cyanobacteria draw down nutrient levels, and fails to capture the risk of biofilm establishment within water infrastructure. To develop a proactive monitoring framework, we conducted a two-year longitudinal study (2024-2025) of Lake Mitchell, South Dakota, integrating environmental chemistry with advanced computational metagenomics. We used 16S/18S rRNA amplicon sequencing (QIIME2) to profile microbial community structure and Phylogenetic Investigation of Communities by Reconstruction of Unobserved States (PICRUSt2) to predict functional enzyme repertoires. Temporal analysis revealed a significant ecological trajectory toward eutrophic conditions in 2025, characterized by increased alkalinity (pH 8.58 to 8.79) and substantial dissolved nutrient depletion. Crucially, community shifts showed a sharp increase in the HAB-former Microcystis alongside the heterotrophic degrader Flavobacterium. Functional inference reinforced this finding, revealing the significant enrichment of metabolic pathways linked to stress and replication, which often precedes settlement and biofilm initiation: specifically, elevated ATP-dependent helicase (EC:3.6.4.12) and NADH dehydrogenase (EC:1.6.5.3). These enzyme signatures reflect an intensified state of oxidative stress handling and metabolic turnover, characteristic of established bloom communities with high surface-attachment potential. These findings show that predictive functional metagenomic signatures serve as strong and sensitive markers of bloom development and the potential for harmful bacteria to colonize water-treatmen infrastructure. Because these signatures operate independently of short-term nutrient fluctuations, they offer early insight into conditions that may lead to surface-associated microbial growth later in the system-information that is critical for protecting municipal water supplies and public health from HAB contamination. Tuyen Do, Naina Maharjan, Connor Johansen-Sallee, Bichar Dip Shrestha Gurung, Paula Mazzer, Etienne Z. Gnimpieba |
BIBM | 1 |
| 2025 | Functional Biofilm-Associated Microbial Markers in a Colorectal Cancer Cohort: Evidence for Quorum Sensing, Exopolysaccharide Production, and Biofilm-Driven PathogenesisabstractColorectal cancer (CRC) is increasingly recognized as a disease influenced by microbe-host interactions, particularly through mucosal biofilms that attach to the epithelial surface and alter the tumor microenvironment. While many microbiome studies focus on taxonomic composition, functional microbial pathways that drive biofilm formation may provide more reliable biomarkers and mechanistic insight. This study analyzes highvariance microbial enzyme signatures derived from a published CRC cohort to identify biofilm-associated genetic markers and to evaluate their roles in CRC biofilm development, quorum sensing signaling, and therapy response. The results highlight enrichment of LuxS/AI-2 quorum sensing enzymes, autoinducer synthases, and exopolysaccharide (EPS) biosynthesis pathways, suggesting a multi-species cooperative biofilm phenotype in CRC. These functional markers align with recent findings that Fusobacterium nucleatum-rich invasive biofilms are common in CRC [4]. The analysis supports four innovation claims: the identification of biofilm-associated functional markers in CRC; mechanistic pathways underlying CRC biofilm development; the impact of biofilm markers on therapy resistance; and additional diagnostic and therapeutic implications. Together, these results emphasize the value of functional biofilm biomarkers in understanding CRC progression and treatment outcomes. Grace Goeden, Tuyen Do, Bichar Dip Shrestha Gurung, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 2 |
| 2025 | Digital Twin Forecasting of Quorum-Sensing Associated Biofilm Microorganisms in Urban Wastewater Over a 30-Week IntervalabstractBiofilms in wastewater systems contain dynamic microbial communities regulated by quorum sensing (QS), which governs adhesion, extracellular polymeric substance (EPS) production, stress tolerance, and developmental transitions. Forecasting QS-associated organisms is essential for anticipating biofilm formation and mitigating operational risks in wastewater infrastructure. In this study, we identified QS and biofilm-associated taxa present in a European wastewater metagenomic dataset: Vibrio harveyi, Vibrio parahaemolyticus, Pseudomonas aeruginosa, Escherichia coli, Salmonella typhi, Salmonella typhimurium, Bacillus subtilis, and Staphylococcus aureus. These organisms encode diverse QS and biofilm regulators across multiple bacterial lineages. From this broader QS-associated set, we developed an organism-specific, QS-aware digital twin focused solely on Pseudomonas, a dominant biofilm-forming genus in wastewater systems. Using the Q-net modeling framework, the model was trained on early-window observations and calibrated to perform short-horizon, end-point forecasting by predicting the final week of a 30-week interval. To ensure a stable conditioning window, the terminal signal was replicated prior to prediction. The digital twin demonstrated strong agreement between predicted and observed final-week abundance$(\mathbf{R}^{2}=0.98)$, highlighting its effectiveness for end-point microbial forecasting. This streamlined and interpretable framework supports rapid microbial surveillance and provides a scalable tool for biofilm-aware operational decision-making in wastewater systems. Bichar Dip Shrestha Gurung, Tuyen Do, Shiva Aryal, Naina Maharjan, Dikshya Bhandari, Bipul Bhattarai, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 2 |
| 2025 | Toward Autonomous Biofilm Engineering: A Vision for a Digital Twin-Based Lab AutomationabstractBiofilm research is undergoing a significant transformation due to a growing need to address biological complexity, temporal dynamics, and environmental sensitivity. These growing needs demand experimental platforms that are far more autonomous, scalable, and computationally integrated than current microbiology lab workflows. This paper presents a vision for a digital twin-driven “cyber-physical laboratory” that combines modular robotics, automated imaging, an Internet of Things network, and AI-based phenotyping into a cohesive system for next-generation computational biofilm engineering. We outline an eight-component architecture comprising twins for physical, environmental, biological, protocol, data, AI/ML, control, and safety constraints. This architecture will enable virtual experimentation, predictive simulation, and self-correction, with the goal of interactive total lab automation and the “robot scientist”. Although we demonstrate early prototypes, the main objective of this paper is to articulate a forward-looking blueprint for how cyber-physical labs with digital twinning can fundamentally reshape biofilm research. Two use cases implementation in our lab shows the feasibility with up to 500 samples run per day (10,000 samples per week). By integrating real-time data streams with predictive biological models and closed-loop automated control, this architecture sets the stage for autonomous, reproducible, and computationally guided biofilm experimentation. We propose this platform as a foundational vision for the future of computational biofilm engineering. Bichar Dip Shrestha Gurung, Timothy W. Hartman, Tuyen Do, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 3 |
| 2024 | Toward Fine-Tuning Large Language Models in Ontology of Microbial Phenotypes ConstructionabstractOntologies are crucial for organizing domainspecific knowledge in biomedical fields, but their manual construction is time-consuming. This study explores the automation of ontology learning using large language models (LLMs) like BERT, RoBERTa, and DistilBERT, focusing on the Ontology of Microbial Phenotypes (OMP). We investigate three key tasks: (1) entity extraction, (2) relation extraction between entities, and (3) ontology verification. These tasks align with broader applications in biomedical annotation and named entity recognition (NER) by enabling the identification and structuring of key terms and relationships within microbial phenotypes. We evaluate LLMs in two scenarios: baseline performance using pre-trained models and fine-tuned performance after training on OMP-specific data. Our approach integrates spaCy for entity extraction, Llama 2 for relation identification, and LLMs for ontology verification. Experiments reveal that fine-tuned models significantly improve accuracy, precision, recall, and F1 scores, particularly for ontology verification. This research highlights the potential of LLMs to enhance ontology learning and support related biomedical applications like biofilm analysis, annotation, and NER, while emphasizing the value of expert curation. Anushuya Baidya, Tuyen Do, Etienne Z. Gnimpieba |
BIBM | 2 |
| 2024 | NFB-Checker: An AI/ML-Powered Microorganism Nitrogen Fixation Susceptibility Prediction from Gene CollectionabstractNitrogen fixation is a crucial process involved in many aspects of the ecological life cycle. However, few organismsparticularly kingdom bacteria-are known to fix N2in a form that could be used in agribusiness. The ability to identify an organism as an N2fixer is a novel case for more identification and application. This study employs functional genes and an XGBoost machine learning model to predict an organism's ability to fix nitrogen, focusing on key orthologs like nifH, nifD, and nifK, essential in the nitrogenase enzyme complex. The model, trained on a dataset of 923 N2fixers and 981 non-fixers, achieved an accuracy of 92.6% and an F1 score of 0.92. This research highlights the potential of machine learning in identifying genetic markers for N2fixation, providing an efficient alternative to traditional methods, and offers new insights for ecological and agricultural research. The inclusion of specific orthologs enhances the model's predictive accuracy, demonstrating the importance of targeted genetic markers in computational biology. Tuyen Do, Shiva Aryal, Bichar Dip Shrestha Gurung, Diing D. M. Agany, Nick Klein, Ruanbao Zhou, Rajesh Kumar Sani, Etienne Z. Gnimpieba |
BIBM | 1 |
| 2024 | Applying the DeepSeqDock Harmonization Framework on Pseudomonas Aeruginosa Bacterial Biofilm DataabstractIn the modern era, data availability has exponentially increased. From different universities, institutions, and laboratories across the globe, the wealth of information brings potential for integration and downstream analysis. This paper focuses specifically on the integration of RNA-sequencing data using data harmonization, a technique that reduces the inter-institutional batch effects that confound data quality with non-biological factors. One such framework is the DeepSeqDock framework, which creates data harmonization pipelines through training and tuning in specific domains and is evaluated on human RNA data from the SEQC. In this study, we present an alternate application of the adjusted DeepSeqDock framework for bacterial biofilm data. By adjusting for non-human data, we show the generalization of data harmonization and the DeepSeqDock for its applications in specific bacterial domains of interest for downstream analysis. Corey Zhao, Bichar Dip Shrestha Gurung, Timothy W. Hartman, Tuyen Do, Mariah M. Hoffman, Etienne Z. Gnimpieba |
BIBM | 4 |
| 2024 | Biofilm marker discovery with cloud-based dockerized metagenomics analysis of microbial communitiesabstractIn an environment, microbes often work in communities to achieve most of their essential functions, including the production of essential nutrients. Microbial biofilms are communities of microbes that attach to a nonliving or living surface by embedding themselves into a self-secreted matrix of extracellular polymeric substances. These communities work together to enhance their colonization of surfaces, produce essential nutrients, and achieve their essential functions for growth and survival. They often consist of diverse microbes including bacteria, viruses, and fungi. Biofilms play a critical role in influencing plant phenotypes and human microbial infections. Understanding how these biofilms impact plant health, human health, and the environment is important for analyzing genotype-phenotype-driven rule-of-life functions. Such fundamental knowledge can be used to precisely control the growth of biofilms on a given surface. Metagenomics is a powerful tool for analyzing biofilm genomes through function-based gene and protein sequence identification (functional metagenomics) and sequence-based function identification (sequence metagenomics). Metagenomic sequencing enables a comprehensive sampling of all genes in all organisms present within a biofilm sample. However, the complexity of biofilm metagenomic study warrants the increasing need to follow the Findability, Accessibility, Interoperability, and Reusable (FAIR) Guiding Principles for scientific data management. This will ensure that scientific findings can be more easily validated by the research community. This study proposes a dockerized, self-learning bioinformatics workflow to increase the community adoption of metagenomics toolkits in a metagenomics and meta-transcriptomics investigation. Our biofilm metagenomics workflow self-learning module includes integrated learning resources with an interactive dockerized workflow. This module will allow learners to analyze resources that are beneficial for aggregating knowledge about biofilm marker genes, proteins, and metabolic pathways as they define the composition of specific microbial communities. Cloud and dockerized technology can allow novice learners-even those with minimal knowledge in computer science-to use complicated bioinformatics tools. Our cloud-based, dockerized workflow splits biofilm microbiome metagenomics analyses into four easy-to-follow submodules. A variety of tools are built into each submodule. As students navigate these submodules, they learn about each tool used to accomplish the task. The downstream analysis is conducted using processed data obtained from online resources or raw data processed via Nextflow pipelines. This analysis takes place within Vertex AI's Jupyter notebook instance with R and Python kernels. Subsequently, results are stored and visualized in Google Cloud storage buckets, alleviating the computational burden on local resources. The result is a comprehensive tutorial that guides bioinformaticians of any skill level through the entire workflow. It enables them to comprehend and implement the necessary processes involved in this integrated workflow from start to finish. This manuscript describes the development of a resource module that is part of a learning platform named "NIGMS Sandbox for Cloud-based Learning" https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [1] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses. Etienne Z. Gnimpieba, Timothy W. Hartman, Tuyen Do, Jessica Zylla, Shiva Aryal, Samuel J. Haas, Diing D. M. Agany, Bichar Dip Shrestha Gurung, Valena Doe, Zelaikha B. Yosufzai, Daniel Pan, Ross Campbell, Victor C. Huber, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Carol Lushbough |
Briefings Bioinform. | 3 |
| 2023 | Utilizing XGBoost for the Prediction of Material Corrosion Rates from Embedded Tabular Data using Large Language ModelabstractMicrobial corrosion, scientifically referred to as microbial-induced corrosion (MIC), constitutes a noteworthy and frequently underestimated concern within diverse industrial domains. This phenomenon manifests when microorganisms, including bacteria, archaea, and fungi, engage with structural materials, resulting in the degradation of infrastructure and equipment. The accurate prognostication of material microbial corrosion rates is of upmost importance in the formulation of proactive strategies for maintenance and corrosion control. In this study, a novel methodology is introduced, which harnesses the capabilities of XGBoost, an advanced gradient boosting algorithm, for the precise prediction of material microbial corrosion rates. This predictive process is facilitated by employing tabular data that is intricately embedded within a comprehensive large language models (LLMs). The integration of tabular data into the language model yields a sophisticated contextual comprehension of the data, thereby augmenting the model's precision by its aptitude to discern intricate relationships and semantic nuances intrinsic to the tabular data. Tuyen Do, Bichar Dip Shrestha Gurung, Shiva Aryal, Anup Khanal, Sandeep Chataut, Venkataramana Gadhamshetty, Carol Lushbough, Etienne Z. Gnimpieba |
BIBM | 1 |
| 2023 | Transformer in Microbial Image Analysis: A Comparative Exploration of TransUNet, UNet, and DoubleUNet for SEM Image SegmentationabstractThe advent of transformer-based architectures such as TransUNet has revolutionized image segmentation as this approach combines the strengths of transformers for capturing contextual information with convolutional neural networks (CNNs) for localized feature identification. Microbes, known for their complex behaviors, present challenges in various fields, especially biomedicine. Image segmentation is crucial for ana-lyzing microbes, allowing quantitative analysis, growth tracking, and understanding host-pathogen interactions. This study is dedicated to a comparative analysis of TransUNet alongside two other popular segmentation methods, UNet and DoubleUNet, in the context of segmenting scanning electron microscope (SEM) images of microbes on layered graphene-nickel specimens. The TransUNet architecture employs a pre-defined ResNet-50 and Vision Transformer (ViT) as the encoder and a custom-built decoder trained on SEM data of Oleidesulfovibrio alaskensis (OA-G20) exposed to graphene-nickel specimens for 30 days. Using the Intersection Over Union (IoU) score as a performance metric, we observed that TransUNet achieved a maximum IoU of 79.58%, DoubleUNet exhibited a maximum IoU of 76.28%, and UNet attained a maximum IoU of 72.38%. We believe that this comparative study of the segmentation approach is invaluable for selecting the best model for the practitioner as per need. This study is the first step in our aim of developing an end-to-end framework with automated model selection based on dataset characteristics for microbial image segmentation. Bichar Dip Shrestha Gurung, Anup Khanal, Timothy W. Hartman, Tuyen Do, Sandeep Chataut, Carol Lushbough, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 4 |
| 2023 | SIMPL - An Application for Cell and Microbe Tracking Using Machine LearningabstractIn the ever-evolving landscape of research and clinical practices, the role of microscope-based image capture and analysis cannot be overstated. While current methodologies for micrograph analysis offer substantial power, there exist critical gaps, particularly in user functionality and reproducibility. To bridge these gaps, we introduce the Smart Imaging of Micrographs Process and Labeling (SIMPL) system as an open-source, semi-automated framework designed for image and video capture, analysis, and particle tracking. SIMPL aims to meet the demands of high-throughput applications, especially for reproducible microbe tracking. Sam Haas, Bichar Dip Shrestha Gurung, Timothy W. Hartman, Tuyen Do, Etienne Z. Gnimpieba |
BIBM | 4 |
| 2022 | U-Net Based Image Segmentation Techniques for Development of Non-Biocidal Fouling-Resistant Ultra-Thin Two-dimensional (2D) CoatingsabstractBacterial adhesion to the metallic surfaces creates a complex biofilm network, resulting in many problems like corrosion and fouling. Precise quantitative analysis of the surface coverage of cells can be vital in decoding biofilm-related issues. We present a deep learning-based approach to automate the microbes’ segmentation from the Scanning Electron Microscope (SEM) images of biofilm developed on coated multilayer graphene nickel samples. We collected SEM images from multilayer graphene nickel exposed to Oleidesulfovibrioalaskensis (OA-G20) for 30 days. Then we manually annotated the microbes with the help of subject expertise and trained a deep-learning U-Net architecture. In order to deal with larger image sizes, we perform patched-based image training and techniques for predicting segmentation masks over the larger image. Intersection over Union (IOU) was calculated for the evaluation of the performance of the model. After training the image for multiple epochs and extracting the optimal model parameters from the learning curve, we were able to get 70.64% mean IOU score. The patched-based technique for image training and image inference showed promising output during the segmentation. Bichar Dip Shrestha Gurung, Ramesh Devadig, Tuyen Do, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 3 |
| 2022 | Using BASIN-ML for Machine Learning-Based Statistical Analysis and Reporting for Biofilm DatasetsabstractBiological and biomedical microscope image (bioimage) comparison remains useful to approach many research challenges—from biofilms to human diseases. This powerful technology allows researchers to provide the community with a quick visual snapshot of varying experimental conditions. But a two-condition comparison still relies on a researcher’s eyes to draw conclusions despite the availability of multiple— often complex—digital image analysis tools. Our Bioimage Analysis, Statistic, and Comparison (BASIN) software provides an easy, objective, reproducible comparison leveraging inferential statistics to bridge image data analysis with other biomedical data modalities such as gene expression. Users have access to a machine learning module to assist with image segmentation using modern, trainable algorithms. BASIN also provides several key data points including images’ object counts, net and mean pixel intensities, net and mean object surface areas, plus a variety of other potentially useful data. Hypothesis testing is performed on mean object intensities and surface areas using the statistical power of the R programming language. These features allow BASIN to extend the current scope of image comparison. It gives researchers a multi-model knowledge about matters such as drug protein marker response, the significance of cell population changes, and changes in cell morphology. To improve BASIN’s accessibility and transparency we implemented it in R using Shiny framework and provided both an online trial version and a customizable offline version. We also have a batch version to run on datasets with hundreds of biomedical images. BASIN workflows consist of five core modules including image upload, feature extraction, statistical analysis, visualization, and report generation. Sam Haas, Timothy W. Hartman, Bichar Dip Shrestha Gurung, Tuyen Do, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba |
BIBM | 4 |