Carol Lushbough

dblp:84/6686 · also Carol M. Lushbough · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0003-3838-5690ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 2 first-author
YearPublicationVenuePosition
2024 Biofilm marker discovery with cloud-based dockerized metagenomics analysis of microbial communities
abstract
In an environment, microbes often work in communities to achieve most of their essential functions, including the production of essential nutrients. Microbial biofilms are communities of microbes that attach to a nonliving or living surface by embedding themselves into a self-secreted matrix of extracellular polymeric substances. These communities work together to enhance their colonization of surfaces, produce essential nutrients, and achieve their essential functions for growth and survival. They often consist of diverse microbes including bacteria, viruses, and fungi. Biofilms play a critical role in influencing plant phenotypes and human microbial infections. Understanding how these biofilms impact plant health, human health, and the environment is important for analyzing genotype-phenotype-driven rule-of-life functions. Such fundamental knowledge can be used to precisely control the growth of biofilms on a given surface. Metagenomics is a powerful tool for analyzing biofilm genomes through function-based gene and protein sequence identification (functional metagenomics) and sequence-based function identification (sequence metagenomics). Metagenomic sequencing enables a comprehensive sampling of all genes in all organisms present within a biofilm sample. However, the complexity of biofilm metagenomic study warrants the increasing need to follow the Findability, Accessibility, Interoperability, and Reusable (FAIR) Guiding Principles for scientific data management. This will ensure that scientific findings can be more easily validated by the research community. This study proposes a dockerized, self-learning bioinformatics workflow to increase the community adoption of metagenomics toolkits in a metagenomics and meta-transcriptomics investigation. Our biofilm metagenomics workflow self-learning module includes integrated learning resources with an interactive dockerized workflow. This module will allow learners to analyze resources that are beneficial for aggregating knowledge about biofilm marker genes, proteins, and metabolic pathways as they define the composition of specific microbial communities. Cloud and dockerized technology can allow novice learners-even those with minimal knowledge in computer science-to use complicated bioinformatics tools. Our cloud-based, dockerized workflow splits biofilm microbiome metagenomics analyses into four easy-to-follow submodules. A variety of tools are built into each submodule. As students navigate these submodules, they learn about each tool used to accomplish the task. The downstream analysis is conducted using processed data obtained from online resources or raw data processed via Nextflow pipelines. This analysis takes place within Vertex AI's Jupyter notebook instance with R and Python kernels. Subsequently, results are stored and visualized in Google Cloud storage buckets, alleviating the computational burden on local resources. The result is a comprehensive tutorial that guides bioinformaticians of any skill level through the entire workflow. It enables them to comprehend and implement the necessary processes involved in this integrated workflow from start to finish. This manuscript describes the development of a resource module that is part of a learning platform named "NIGMS Sandbox for Cloud-based Learning" https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [1] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses.
Etienne Z. Gnimpieba, Timothy W. Hartman, Tuyen Do, Jessica Zylla, Shiva Aryal, Samuel J. Haas, Diing D. M. Agany, Bichar Dip Shrestha Gurung, Valena Doe, Zelaikha B. Yosufzai, Daniel Pan, Ross Campbell, Victor C. Huber, Rajesh Kumar Sani, Venkataramana Gadhamshetty, Carol Lushbough
Briefings Bioinform.16
2023 Utilizing XGBoost for the Prediction of Material Corrosion Rates from Embedded Tabular Data using Large Language Model
abstract
Microbial corrosion, scientifically referred to as microbial-induced corrosion (MIC), constitutes a noteworthy and frequently underestimated concern within diverse industrial domains. This phenomenon manifests when microorganisms, including bacteria, archaea, and fungi, engage with structural materials, resulting in the degradation of infrastructure and equipment. The accurate prognostication of material microbial corrosion rates is of upmost importance in the formulation of proactive strategies for maintenance and corrosion control. In this study, a novel methodology is introduced, which harnesses the capabilities of XGBoost, an advanced gradient boosting algorithm, for the precise prediction of material microbial corrosion rates. This predictive process is facilitated by employing tabular data that is intricately embedded within a comprehensive large language models (LLMs). The integration of tabular data into the language model yields a sophisticated contextual comprehension of the data, thereby augmenting the model's precision by its aptitude to discern intricate relationships and semantic nuances intrinsic to the tabular data.
Tuyen Do, Bichar Dip Shrestha Gurung, Shiva Aryal, Anup Khanal, Sandeep Chataut, Venkataramana Gadhamshetty, Carol Lushbough, Etienne Z. Gnimpieba
BIBM7
2023 Transformer in Microbial Image Analysis: A Comparative Exploration of TransUNet, UNet, and DoubleUNet for SEM Image Segmentation
abstract
The advent of transformer-based architectures such as TransUNet has revolutionized image segmentation as this approach combines the strengths of transformers for capturing contextual information with convolutional neural networks (CNNs) for localized feature identification. Microbes, known for their complex behaviors, present challenges in various fields, especially biomedicine. Image segmentation is crucial for ana-lyzing microbes, allowing quantitative analysis, growth tracking, and understanding host-pathogen interactions. This study is dedicated to a comparative analysis of TransUNet alongside two other popular segmentation methods, UNet and DoubleUNet, in the context of segmenting scanning electron microscope (SEM) images of microbes on layered graphene-nickel specimens. The TransUNet architecture employs a pre-defined ResNet-50 and Vision Transformer (ViT) as the encoder and a custom-built decoder trained on SEM data of Oleidesulfovibrio alaskensis (OA-G20) exposed to graphene-nickel specimens for 30 days. Using the Intersection Over Union (IoU) score as a performance metric, we observed that TransUNet achieved a maximum IoU of 79.58%, DoubleUNet exhibited a maximum IoU of 76.28%, and UNet attained a maximum IoU of 72.38%. We believe that this comparative study of the segmentation approach is invaluable for selecting the best model for the practitioner as per need. This study is the first step in our aim of developing an end-to-end framework with automated model selection based on dataset characteristics for microbial image segmentation.
Bichar Dip Shrestha Gurung, Anup Khanal, Timothy W. Hartman, Tuyen Do, Sandeep Chataut, Carol Lushbough, Venkataramana Gadhamshetty, Etienne Z. Gnimpieba
BIBM6
2022 Attention model-based and multi-organism driven gene recognition from text: application to a microbial biofilm organism set
abstract
Nowadays, online databases such as PUBMED and PMC are experiencing an explosion of publications in the field of biomedical sciences. With so much information available online, one of the biggest challenges is managing all that raw, unstructured data and making it machine-readable. Name entity recognition is nowadays a prerequisite for data identification and extraction in biosciences. One of the areas that allows automatic extraction of information from biomedical literature today is Name Entity Recognition. Indeed, it makes it possible to simplify the workflow analysis and automatic extraction of name entities, thus improving the various existing models. There is in the literature a lot of tools for this purpose, but they are unable to extract microbial genes accurately. Moreover, current goal standard corpora such as BIOCREATIVE I to IV have limited representation of microbial knowledge. In this paper, we proposed a new method to recognize biofilm gene mentions from free text. This method relies on a context-specific dictionary to annotate a consistent corpus necessary to train an efficient recognition model. Indeed, this method provides a new workflow for dataset collection generation for microbial biofilm gene. Trained on a set of biofilm organisms our method achieves a score of up to 94%, outperforming state-of-the-art frameworks.
Alain Bertrand Bomgni, Ernest Basile Fotseu Fotseu, Daril Raoul Kengne Wambo, Rajesh Kumar Sani, Carol Lushbough, Etienne Z. Gnimpieba
BIBM5
2021 Prediction of essential genes in G20 using machine learning model
abstract
Despite the exponential growth in bioscience data, one of the key challenges for machine learning engineers remains the incompleteness of bioscience dataset (biodata). For a specific bioscience problem such as (e.g. biofilm formation, drug response, organism survival), it is very difficult to find a good consistent dataset capturing the numerous variables involved in each of these processes. Each systems biology data point is measured with different protocols in different settings, making their integration hard and not reliable. This paper focuses on using machine learning (ML) models and data mining (DM) workflow to perform gene essential prediction in G20. Actually, developing next-generation and nano-scale coatings to control biofilm formation on technologically relevant materials is a great challenge today. This can help to control microbial corrosion on material or engineer better relevant material. To tackle this relevant problem, a detailed understanding of the bacterial survival mechanisms is crucial. Computational methods for predicting essential genes can make it easier and faster to obtain reliable results. Method: The main hypothesis of our work is that a minimal information-driven specific Machine Learning model can outperform an interesting prediction score. To reach our goal, we set up first a completed data mining workflow to extract gene features from G20. We then derive 10192 features from gene sequence and protein sequence divided into 25 relevant subgroups. From each subgroup, we build a couple of interesting machine learning models. Result: We identified 69 relevant subgroups of features using our features selection algorithm. We tested the model performance on each of these subgroups and our predictive result achieved up to 98% accuracy score. These subgroups of features can be used to assist researchers to select good variables for their respective experiments.
Thierry Kongne Nembot, Ernest Basile Fotseu Fotseu, Rajesh Kumar Sani, Etienne Z. Gnimpieba, Carol Lushbough, Alain Bertrand Bomgni
BIBM5
2021 Ten simple rules to cultivate transdisciplinary collaboration in data science
abstract
Author(s): Sahneh, Faryad; Balk, Meghan A; Kisley, Marina; Chan, Chi-kwan; Fox, Mercury; Nord, Brian; Lyons, Eric; Swetnam, Tyson; Huppenkothen, Daniela; Sutherland, Will; Walls, Ramona L; Quinn, Daven P; Tarin, Tonantzin; LeBauer, David; Ribes, David; Birnie, Dunbar P; Lushbough, Carol; Carr, Eric; Nearing, Grey; Fischer, Jeremy; Tyle, Kevin; Carrasco, Luis; Lang, Meagan; Rose, Peter W; Rushforth, Richard R; Roy, Samapriya; Matheson, Thomas; Lee, Tina; Brown, C Titus; Teal, Tracy K; Papeș, Monica; Kobourov, Stephen; Merchant, Nirav | Editor(s): Schwartz, Russell
Faryad Sahneh, Meghan A. Balk, Marina Kisley, Chi-Kwan Chan, Mercury Fox, Brian Nord, Eric Lyons 0002, Tyson Lee Swetnam, Daniela Huppenkothen, Will Sutherland, Ramona L. Walls, Daven P. Quinn, Tonantzin Tarin, David S. LeBauer, David Ribes, Dunbar P. Birnie III, Carol Lushbough, Eric Carr, Grey Nearing, Jeremy Fischer, Kevin Tyle, Luis Carrasco, Meagan Lang, Peter W. Rose, Richard R. Rushforth, Samapriya Roy, Thomas Matheson, Tina Lee, C. Titus Brown, Tracy K. Teal, Monica Papes, Stephen G. Kobourov, Nirav C. Merchant
PLoS Comput. Biol.17
2015 Life science data analysis workflow development using the bioextract server leveraging the iPlant collaborative cyberinfrastructure
abstract
Summary In order to handle the vast quantities of biological data gener6ated by high‐throughput experimental technologies, the BioExtract Server (bioextract.org) has leveraged iPlant Collaborative ( www.iplantcollaborative.org ) functionality to help address big data storage and analysis issues in the bioinformatics field. The BioExtract Server is a Web‐based, workflow‐enabling system that offers researchers a flexible environment for analyzing genomic data. It provides researchers with the ability to save a series of BioExtract Server tasks (e.g., query a data source, save a data extract, and execute an analytic tool) as a workflow and the opportunity for researchers to share their data extracts, analytic tools, and workflows with collaborators. The iPlant Collaborative is a community of researchers, educators, and students working to enrich science through the development of cyberinfrastructure—the physical computing resources, collaborative environment, virtual machine resources, and interoperable analysis software and data services—that are essential components of modern biology. The iPlant AGAVE Advanced Programming Interface, developed through the iPlant Collaborative, is a hosted, Software‐as‐a‐Service resource providing access to a collection of high performance computing and cloud resources. Leveraging AGAVE, the BioExtract Server gives researchers easy access to multiple high performance computers and delivers computation and storage as dynamically allocated resources via the Internet. © 2014 The Authors. Concurrency and Computation: Practice and Experience published by John Wiley & Sons Ltd.
Carol Lushbough, Etienne Z. Gnimpieba, Rion Dooley
Concurr. Comput. Pract. Exp.1
2015 A genome-wide association study platform built on iPlant cyber-infrastructure
abstract
Summary We demonstrate a flexible genome‐wide association study platform built upon the iPlant Collaborative Cyber‐infrastructure. The platform supports big data management, sharing, and large‐scale study of both genotype and phenotype data on clusters. End users can add their own analysis tools and create customized analysis workflows through the graphical user interfaces in both iPlant Discovery Environment and BioExtract server. Copyright © 2014 John Wiley & Sons, Ltd.
Doreen Ware, Carol Lushbough, Nirav C. Merchant, Lincoln Stein
Concurr. Comput. Pract. Exp.3
2013 BioExtract Server, a Web-based workflow enabling system, leveraging iPlant collaborative resources
abstract
In order to handle the vast quantities of biological data generated by high-throughput experimental technologies, the BioExtract Server (bioextract.org) has leveraged iPlant Collaborative (www.iplantcollaborative.org) functionality to help address big data storage and analysis issues in the bioinformatics field. The BioExtract Server is a Web-based, workflow-enabling system that offers researchers a flexible environment for analyzing genomic data. It provides researchers with the ability to save a series of BioExtract Server tasks (e.g. query a data source, save a data extract, and execute an analytic tool) as a workflow and the opportunity for researchers to share their data extracts, analytic tools and workflows with collaborators. The iPlant Collaborative is a community of researchers, educators, and students working to enrich science through the development of cyberinfrastructure - the physical computing resources, collaborative environment, virtual machine resources, and interoperable analysis software and data services - that are essential components of modern biology. The iPlant Agave API (Agave), developed through the iPlant Collaborative, is a hosted, Software-as-a-Service resource providing access to a collection of High Performance Computing (HPC) and Cloud resources [6]. Leveraging Agave, the BioExtract Server gives researchers easy access to multiple high performance computers and delivers computation and storage as dynamically allocated resources via the Internet.
Carol Lushbough, Etienne Z. Gnimpieba, Rion Dooley
CLUSTER1
2013 Building an open Genome Wide Association Study (GWAS) platform
abstract
We demonstrated how a flexible Genome Wide Association Study (GWAS) platform was built using the iPlant Collaborative Cyber-infrastructure. The platform is open for end users to add additional analysis tools. With the platform, customized GWAS workflows can be built in both iPlant Discovery Environment and BioExtract server.
Doreen Ware, Nirav C. Merchant, Carol Lushbough
CLUSTER4
2010 BioExtract Server - An Integrated Workflow-Enabling System to Access and Analyze Heterogeneous, Distributed Biomolecular Data
abstract
Many in silico investigations in bioinformatics require access to multiple, distributed data sources and analytic tools. The requisite data sources may include large public data repositories, community databases, and project databases for use in domain-specific research. Different data sources frequently utilize distinct query languages and return results in unique formats, and therefore researchers must either rely upon a small number of primary data sources or become familiar with multiple query languages and formats. Similarly, the associated analytic tools often require specific input formats and produce unique outputs which make it difficult to utilize the output from one tool as input to another. The BioExtract Server (http://bioextract.org) is a Web-based data integration application designed to consolidate, analyze, and serve data from heterogeneous biomolecular databases in the form of a mash-up. The basic operations of the BioExtract Server allow researchers, via their Web browsers, to specify data sources, flexibly query data sources, apply analytic tools, download result sets, and store query results for later reuse. As a researcher works with the system, their "steps" are saved in the background. At any time, these steps can be preserved long-term as a workflow simply by providing a workflow name and description.
Carol Lushbough, Michael K. Bergman, Carolyn J. Lawrence-Dill, Douglas M. Jennewein, Volker Brendel
IEEE ACM Trans. Comput. Biol. Bioinform.1