Trupti Joshi

dblp:21/2104 · DBLP profile ↗
← Back
36ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0001-8944-4924ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 29 · 3 first-author · 1 since 2021Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2024 Impact of Extracellular Vesicles Derived from Human Placenta Cells on Neural Progenitor Cell Transcriptome Dynamics
abstract
The fetal brain relies on the placenta for essential developmental signals, forming what is known as the placenta-brain axis. We hypothesize that extracellular vesicles (EVs) released by trophoblast (TB) cells transport small RNAs and other molecules from the placenta to the brain, influencing its early development. In this study, we isolated EVs from human TB cells differentiated in vitro and from their parental induced pluripotent stem cells (iPSCs). Small RNA sequencing revealed that TB-derived EVs contain distinct miRNAs, including hsa-miR-149-3p and hsa-miR-302a-5p, as well as long non-coding RNAs (lncRNAs), which are enriched in neural tissue. These TB-derived EVs are efficiently internalized by human neural progenitor cells (NPCs), leading to the induction of transcripts associated with forebrain formation and neurogenesis. Our findings provide insights into the role of the placenta-brain axis and suggest that TB-derived EVs may be key players in fetal neurodevelopment, offering potential biomarkers and therapeutic targets for neurobehavioral disorders like autism spectrum disorders (ASD).
Pallav Singh, Jessica A. Kinkade, Teka Khan, Toshihiko Ezashi, Nathan J. Bivens, R. Michael Roberts, Cheryl S. Rosenfeld, Trupti Joshi
BIBM9
2023 Domain-Specific Topic Model for Knowledge Discovery in Computational and Data-Intensive Scientific Communities
abstract
Shortened time to knowledge discovery and adapting prior domain knowledge is a challenge for computational and data-intensive communities such as e.g., bioinformatics and neuroscience. The challenge for a domain scientist lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics when: investigating new methods, developing new tools, or integrating datasets. In this paper, we propose a novel "domain-specific topic model" (DSTM) to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplary scientific domains. Our DSTM is a generative model that extends the Latent Dirichlet Allocation (LDA) model and uses the Markov chain Monte Carlo (MCMC) algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include more than 25,000 of papers over the last ten years, featuring hundreds of tools and datasets that are commonly used in relevant studies. Evaluation experiments based on generalization and information retrieval metrics show that our model has better performance than the state-of-the-art baseline models for discovering highly-specific latent topics within a domain. Lastly, we demonstrate applications that benefit from our DSTM to discover intra-domain, cross-domain and trend knowledge patterns.
Yuanxun Zhang, Prasad Calyam, Trupti Joshi, Satish S. Nair, Dong Xu 0002
IEEE Trans. Knowl. Data Eng.3
2023 Knowledge-Engineered Multi-Cloud Resource Brokering for Application Workflow Optimization
abstract
Data-intensive application workflows benefit by leveraging cloud services to decrease execution times and increase data sharing. Cloud service providers (CSPs) have distinct capabilities and policies, and performance/cost of the cloud services are amongst the prime factors for CSP selection. However, workflow users who need brokering of cloud resources often lack expert guidance to handle the problem of overwhelming choice in CSP selection, and optimization to compensate for service dynamics. In this paper, we address the optimal resource selection problem using a multi-cloud resource broker viz., OnTimeURB that uses knowledge-engineering of user requirements and service capabilities across multiple CSPs. OnTimeURB is powered by integer linear programming and a Naive Bayes classifier to recommend optimal cloud template solutions by weighting performance, agility, cost, and security (PACS) factors. We evaluate the OnTimeURB recommendations with a catalog of bioinformatics application workflows using four CSP resources featuring more than 300 different instance configurations. Our evaluation results show the efficacy of OnTimeURB in creating consistently cost-effective and agile solutions compared to a state-of-the-art k-nearest neighbors (k-NN) approach. We also show that OnTimeURB has 91% success rate improvement in workflow execution times via cloud template recommendations over approaches that do not use knowledge-engineered multi-CSP resource brokering.
Prasad Calyam, Zhen Lyu, Songjie Wang, D. Yu. Chemodanov, Trupti Joshi
IEEE Trans. Netw. Serv. Manag.6
2021 Fuzzy-Engineered Multi-Cloud Resource Brokering for Data-intensive Applications
abstract
Multi-cloud resource brokering is becoming a critical requirement for applications that require high scale, diversity, and resilience. Applications demand timely selection of distributed data storage and computation platforms that span local private cloud resources as well as resources from multiple cloud service providers (CSPs). The distinct capabilities and policies, as well as performance/cost of the cloud services, are amongst the prime factors for CSP selection. However, application owners who need suitable cyber resources in community/public clouds, often have preliminary knowledge and preferences of certain CSPs. They also lack expert guidance to handle the problem of overwhelming resource choice from CSPs, and optimization to compensate for service dynamics. In this paper, we address this challenge of optimal resource selection while also leveraging limited user's expertise and preferences towards CSPs through multi-level fuzzy logic modeling based on convoluted factors of performance, agility, cost, and security. We evaluate the efficiency of our fuzzy-engineered resource brokering in improving allocation of resources as well as user satisfiability by using case studies and independent validations of CSPs evaluation.
Prasad Calyam, Zhen Lyu, Trupti Joshi
CCGRID4
2021 Recommender-as-a-service with chatbot guided domain-science knowledge discovery in a science gateway
abstract
Scientists in disciplines such as neuroscience and bioinformatics are increasingly relying on science gateways for experimentation on voluminous data, as well as analysis and visualization in multiple perspectives. Though current science gateways provide easy access to computing resources, datasets and tools specific to the disciplines, scientists often use slow and tedious manual efforts to perform knowledge discovery to accomplish their research/education tasks. Recommender systems can provide expert guidance and can help them to navigate and discover relevant publications, tools, data sets, or even automate cloud resource configurations suitable for a given scientific task. To realize the potential of integration of recommenders in science gateways in order to spur research productivity, we present a novel "OnTimeRecommend" recommender system. The OnTimeRecommend comprises of several integrated recommender modules implemented as microservices that can be augmented to a science gateway in the form of a recommender-as-a-service. The guidance for use of the recommender modules in a science gateway is aided by a chatbot plug-in viz., Vidura Advisor. To validate our OnTimeRecommend, we integrate and show benefits for both novice and expert users in domain-specific knowledge discovery within two exemplar science gateways, one in neuroscience (CyNeuro) and the other in bioinformatics (KBCommons).
Komal Bhupendra Vekaria, Prasad Calyam, Sai Swathi Sivarathri, Songjie Wang, Yuanxun Zhang, Dong Xu 0002, Trupti Joshi, Satish S. Nair
Concurr. Comput. Pract. Exp.9
2021 Multi-Cloud Performance and Security Driven Federated Workflow Management
abstract
Federated multi-cloud resource allocation for data-intensive application workflows is generally performed based on performance or quality of service (i.e., QSpecs) considerations. At the same time, end-to-end security requirements of these workflows across multiple domains are considered as an afterthought due to lack of standardized formalization methods. Consequently, diverse/heterogenous domain resource and security policies cause inter-conflicts between application's security and performance requirements that lead to sub-optimal resource allocations. In this paper, we present a joint performance and security-driven federated resource allocation scheme for data-intensive scientific applications. In order to aid joint resource brokering among multi-cloud domains with diverse/heterogenous security postures, we first define and characterize a data-intensive application's security specifications (i.e., SSpecs). Then we describe an alignment technique inspired by Portunes Algebra to homogenize the various domain resource policies (i.e., RSpecs) along an application's workflow lifecycle stages. Using such formalization and alignment, we propose a near optimal cost-aware joint QSpecs-SSpecs-driven, RSpecs-compliant resource allocation algorithm for multi-cloud computing resource domain/location selection as well as network path selection. We implement our security formalization, alignment, and allocation scheme as a framework, viz., “OnTimeURB” and validate it in a multi-cloud environment with exemplar data-intensive application workflows involving distributed computing and remote instrumentation use cases with different performance and security requirements.
Matthew Dickinson, Saptarshi Debroy, Prasad Calyam, Samaikya Valluripally, Yuanxun Zhang, Ronny Bazan Antequera, Trupti Joshi, Tommi A. White, Dong Xu 0002
IEEE Trans. Cloud Comput.7
2020 A multiomics discriminatory analysis approach to identify drought-related signatures in maize nodal roots
abstract
Maize is one of the major food crops grown in the continental US, and as such, major interest is directed towards understanding its adaptability to drought stress. Certain cultivars of maize have shown increased resistance to water shortages, by continuing to maintain root growth even when under severe water stress. To better understand the molecular mechanisms which lead to such adaptation, we analyzed multiomics datasets generated from the growth zone of nodal roots from FR697, an inbred line that shows a superior capacity for root growth maintenance under drought stress. We used a research pipeline consisting of a discriminatory multiomics data integration approach, which uses a combination of sparse Generalized Canonical Correlation Analysis (sGCCA) and generalized Partial Least Square (PLS) analysis instead of traditional “filter funnel” approaches, to incorporate all datasets into one holistic global network and form clusters spanning all omics levels. We then linked significant elements from these clusters to various observations associated with drought stress in the root tip samples, reinforced by their roles in biological pathways.
Sidharth Sen, Tyler McCubbin, Shannon K. King, Laura A. Greeley, Cheyenne Baker, Rachel Mertz, Nicole D. Niehues, Jonathon T. Stemmle, Felix B. Fristchi, David Braun, Scott C. Peck, Melvin J. Oliver, Robert E. Sharp, Trupti Joshi
BIBM15
2020 A Formative Usability Study to Improve Prescriptive Systems for Bioinformatics Big Data
abstract
Big data computation tools are vital for researchers and educators from various domains such as plant science, animal science, biomedical science and others. With the growing computational complexity of biology big data, advanced analytic systems, known as prescriptive systems, are being built using machine learning models to intelligently predict optimum computation solutions for users for better data analysis. However, lack of user-friendly prescriptive systems poses a critical roadblock to facilitating informed decision-making by users. In this paper, we detail a formative usability study to address the complexities faced by users while using prescriptive systems. Our usability research approach considers bioinformatics workflows and community cloud resources in the KBCommons framework's science gateway. The results show that recommendations from usability studies performed in iterations during the development of prescriptive systems can improve user experience, user satisfaction and help novice as well as expert users to make decisions in a well-informed manner.
Kanu Priya Singh, Shangman Li, Isa Jahnke, Zhen Lyu, Trupti Joshi, Prasad Calyam
BIBM6
2020 SNPViz v2.0: A web-based tool for enhanced haplotype analysis using large scale resequencing datasets and discovery of phenotypes causative gene using allelic variations
abstract
Single nucleotide polymorphisms (SNPs) and insertions/deletions (Indels) are widely spread across all chromosomes of the genome and act as biological markers, which aid in identification of genes associated with traits or phenotypes. With the advances in next-generation sequencing (NGS) technology, large amounts of SNPs and Indels data have become available, making it difficult to effectively perform analysis across multiple samples and intuitively integrate, compare and/or visualize them simultaneously. Genome-wide association studies (GWAS) is a widely used method to find genetic variations associated with a trait, but it lacks an efficient way to investigate genomic variant functions. To tackle these issues, we have developed SNPViz v2.0, a web-based tool to visualize large-scale haplotype blocks with detailed SNPs and Indels grouped by their chromosomal coordinates, along with their overlapping gene models, phenotype to genotype accuracies, Gene Ontology (GO) annotations, protein families (Pfam) annotations, genomic variant annotations, and their functional effects. Moreover, SNPViz v2.0 integrates several large scale soybean SNPs and Indels datasets including G. Soja, GWAS, NAM41, USB-15x, USB-40x, MSMC and Zhou302, available from multiple studies. SNPViz v2.0 is deployed for all organisms and available in both SoyKB and KBCommons frameworks. For soybean data only, the SNPViz 2.0 is publicly available at http://soykb.org/SNPViz2/. For other organisms such as Arabidopsis thaliana, Mus musculus and Zea mays, SNPViz 2.0 is publicly available in their respective knowledge bases at https://kbcommons.org.
Mária Skrabisová, Zhen Lyu, Yen On Chan, Kristin Bilyeu, Trupti Joshi
BIBM6
2020 A dynamic programing approach to integrate gene expression data and network information for pathway model generation
abstract
MOTIVATION: As large amounts of biological data continue to be rapidly generated, a major focus of bioinformatics research has been aimed toward integrating these data to identify active pathways or modules under certain experimental conditions or phenotypes. Although biologically significant modules can often be detected globally by many existing methods, it is often hard to interpret or make use of the results toward pathway model generation and testing. RESULTS: To address this gap, we have developed the IMPRes algorithm, a new step-wise active pathway detection method using a dynamic programing approach. IMPRes takes advantage of the existing pathway interaction knowledge in Kyoto Encyclopedia of Genes and Genomes. Omics data are then used to assign penalties to genes, interactions and pathways. Finally, starting from one or multiple seed genes, a shortest path algorithm is applied to detect downstream pathways that best explain the gene expression data. Since dynamic programing enables the detection one step at a time, it is easy for researchers to trace the pathways, which may lead to more accurate drug design and more effective treatment strategies. The evaluation experiments conducted on three yeast datasets have shown that IMPRes can achieve competitive or better performance than other state-of-the-art methods. Furthermore, a case study on human lung cancer dataset was performed and we provided several insights on genes and mechanisms involved in lung cancer, which had not been discovered before. AVAILABILITY AND IMPLEMENTATION: IMPRes visualization tool is available via web server at http://digbio.missouri.edu/impres. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yuexu Jiang, Yanchun Liang 0001, Duolin Wang, Dong Xu 0002, Trupti Joshi
Bioinform.5
2019 Mutational Forks: Inferring Deregulated Flow of Signal Transduction Based on Patient-Specific Mutations
abstract
The precise mechanism behind treatment resistance in cancer is still not fully understood. Despite advances in precision oncology, there is a lack of tools that help to understand a mechanistic picture of treatment resistance in cancer patients. Existing enrichment methods heavily rely on quantitative data and limited to analysis of differentially expressed genes, ignoring crucial players that might be involved in this process. In order to tackle treatment resistance, the precise identification of deregulated flow of signal transduction is critical. Here, we introduce a bioinformatics framework that is capable of inferring deregulated flow of signal transduction given evidence-based knowledge about pathway topology and patient-specific mutations. While testing the proposed pipeline on a case study, our algorithm was able to confirm findings from biological experiment, where KRAS mutant cells developed treatment resistance to MEK inhibitor. Our model provides a framework for mechanistic understanding of acquired treatment resistance, thus, equipped clinicians with tool for searching more accurate diagnostic clues in patients with non-trivial disease representations.
Olha Kholod, Zhen Lyu, Jonathan B. Mitchem, Peter J. Tonellato, Trupti Joshi, Dmitriy Shin
BIBM5
2019 OnTimeURB: Multi-Cloud Resource Brokering for Bioinformatics Workflows
abstract
Scientific workflows due to their data and memory intensive requirements are among the prime applications which benefit by leveraging cloud computing. However, Cloud service providers (CSPs) have distinct policies and service dynamics that present a problem of excess choice for users. Performance and cost of the cloud services are among the principal factors in CSP selection for scientific bioinformatics workflows. The workflows typically are based on private data, and require diverse cloud resources, thus often requiring synergistic services from multiple CSPs. In this paper, we address this challenge of multi-cloud resource selection using cloud template solutions based on user specifications. We propose an optimizer that incorporates a combinatorial optimization model built on performance, cost and CSPs interoperability factors. The optimizer is integrated within a novel resource broker (i.e., OnTimeURB) for prescriptive recommendations of template solutions with intuitive choices for users. We implement and evaluate the OnTimeURB recommendations framework with a catalog of bioinformatics workflow applications integrated within a KBCommons science gateway. The evaluation considered four CSP resources featuring more than 300 different machine configuration instances. Our evaluation results show that our OnTimeURB creates consistently more economical, performance optimized and practical cloud solutions compared to a k-nearest neighbors (k-NN) approach.
Zhen Lyu, Trupti Joshi, Prasad Calyam
BIBM3
2019 Multi-Cloud Performance and Security-driven Brokering for Bioinformatics Workflows
abstract
Data-intensive bioinformatics applications often use federated multi-cloud infrastructures to support compute-intensive processing needs. In this paper, we propose a Multi-Cloud Performance and Security (MCPS) Brokering framework within such federated multi-cloud infrastructures to allocate cloud resources to applications by satisfying their performance and security requirements.
Saptarshi Debroy, Prasad Calyam, Zhen Lyu, Trupti Joshi
ICNP5
2018 Integrative Analysis of DNA Methylation and RNA-Seq Data for Biomarker Detection of Endometriosis
Sadia Akter, John Bromfield, Katherine Pelch, Gil Wilshire, Sarah Crowder, Danny J. Schust, Bret Barrier, J. Wade Davis, Susan C. Nagel, Trupti Joshi
AMIA10
2018 A Data Mining Approach for Biomarker Discovery Using Transcriptomics in Endometriosis
Sadia Akter, Dong Xu 0002, Susan C. Nagel, Trupti Joshi
BIBM4
2018 Integrating Gene Expression Data and Pathway Knowledge for In Silico Hypothesis Generation with IMPRes
Yuexu Jiang, Duolin Wang, Dong Xu 0002, Trupti Joshi
BIBM4
2018 Knowledge Base Commons (KBCommons) v1.0: A multi OMICS' web-based data integration framework for biological discoveries
Zhen Lyu, Siva Ratna Kumari Narisetti, Dong Xu 0002, Trupti Joshi
BIBM5
2018 Domain-specific Topic Model for Knowledge Discovery through Conversational Agents in Data Intensive Scientific Communities
abstract
Machine learning techniques underlying Big Data analytics have the potential to benefit data intensive communities in e.g., bioinformatics and neuroscience domain sciences. Today's innovative advances in these domain communities are increasingly built upon multi-disciplinary knowledge discovery and cross-domain collaborations. Consequently, shortened time to knowledge discovery is a challenge when investigating new methods, developing new tools, or integrating datasets. The challenge for a domain scientist particularly lies in the actions to obtain guidance through query of massive information from diverse text corpus comprising of a wide-ranging set of topics. In this paper, we propose a novel "domain-specific topic model" (DSTM) that can drive conversational agents for users to discover latent knowledge patterns about relationships among research topics, tools and datasets from exemplar scientific domains. The goal of DSTM is to perform data mining to obtain meaningful guidance via a chatbot for domain scientists to choose the relevant tools or datasets pertinent to solving a computational and data intensive research problem at hand. Our DSTM is a Bayesian hierarchical model that extends the Latent Dirichlet Allocation (LDA) model and uses a Markov chain Monte Carlo algorithm to infer latent patterns within a specific domain in an unsupervised manner. We apply our DSTM to large collections of data from bioinformatics and neuroscience domains that include hundreds of papers from reputed journal archives, hundreds of tools and datasets. Through evaluation experiments with a perplexity metric, we show that our model has better generalization performance within a domain for discovering highly specific latent topics.
Yuanxun Zhang, Prasad Calyam, Trupti Joshi, Satish S. Nair, Dong Xu 0002
IEEE BigData3
2018 ADON: Application-Driven Overlay Network-as-a-Service for Data-Intensive Science
abstract
Campuses are increasingly adopting hybrid cloud architectures for supporting data-intensive science applications that require “on-demand” resources, which are not always available locally on-site. Policies at the campus edge for handling multiple such applications competing for remote resources can cause bottlenecks across applications. These bottlenecks can be proactively avoided with pertinent profiling, monitoring and control of application flows using software-defined networking and pertinent selection of local or remote compute resources. In this paper, we present an “application-driven overlay network-as-a-service” (ADON) that manages the hybrid cloud requirements of multiple applications in a scalable and extensible manner by allowing users to specify requirements of the application that are translated into the underlying network and compute provisioning requirements. Our solution involves scheduling transit selection, a cost optimized selection of site(s) for computation and traffic engineering at the campus-edge based upon real-time policy control that ensures prioritized application performance delivery for multi-tenant traffic profiles. We validate our ADON approach through an emulation study and through a wide-area overlay network testbed implementation across two campuses. Our workflow orchestration results show the ADON effectiveness in handling temporal behavior of multi-tenant traffic burst arrivals using profiles from a diverse set of actual data-intensive applications.
Ronny Bazan Antequera, Prasad Calyam, Saptarshi Debroy, Longhai Cui, Sripriya Seetharam, Matthew Dickinson, Trupti Joshi, Dong Xu 0002, Tsegereda Beyene
IEEE Trans. Cloud Comput.7
2017 A multi-omics informatics approach for identifying molecular mechanisms and biomarkers in clinical patients with endometriosis
abstract
Endometriosis is a complex gynecological disorder. The diagnostic process of endometriosis involves an invasive procedure thus delaying the diagnosis for about 10 years on average. Both DNA-methylation data and RNA-seq data has the potential to uncover molecular mechanisms of diseases. The objective of this project is to identify diagnostic molecular mechanisms of endometriosis using a multi-omics approach that will lead to noninvasive diagnostic procedure.
Sadia Akter, Gil Wilshire, J. Wade Davis, John Bromfield, Sarah Crowder, Trupti Joshi, Katherine Pelch, Danny J. Schust, Angela Meng, Bret Barrier, Susan C. Nagel
BIBM6
2017 IMPRes: Integrative MultiOmics pathway resolution algorithm and tool
abstract
A central goal of systems biology is to uncover the underlying functional architecture of the cell and study its mechanisms. To this end, large amounts of omics data are being rapidly generated, and a focus of bioinformatics research has been towards integrating these data to identify active pathways or modules under certain conditions. Many bioinformatics algorithms include optimization methods, statistical methods, and methods using interaction network topology attributes have been applied for this. Although biologically significant modules can often be detected globally by these methods, it is hard to interpret or make use of the results towards in silico hypothesis generation and testing. We propose a step-wise active pathway detection method (IMPRes) using a dynamic programming approach. First, we take advantage of the existing pathway interaction knowledge in KEGG to build a background network, and then starting from one or multiple receptors of a certain perturbation, we use transcriptomics data collected under these conditions to detect paths that best explain the variations of genes downstream. More other omics data will be integrated in the future. Since dynamic programming enables the detection one step a time, it is easy for biomedical researchers to trace the pathway and finally lead to more accurate drug design and more effective treatment strategies. Additionally, by adding protein-protein interactions in our method, the hypotheses that we generate do not merely utilize existing knowledge, but have potential to discover new knowledge. We have evaluated our method on a dataset of cell wall stress in yeast. The path we found highly agrees with the Cell Wall Integrity (CWI) pathway, which is the main signaling pathway involved in the regulation of cell wall stress responses. We have also compared with other methods on a yeast high osmolality stress dataset and achieved an overall better performance than some other methods. More experiments have been done on human cancer datasets and mouse datasets. Finally, the IMPRes web server is established to offer a simple interface for applying IMPRes. Users can upload their own data and obtain an interactive visualization of the resulting pathway map. Users can further filter or highlight interactions according to pathway information or relation types. All genes in the pathway map are listed with detailed annotations. The IMPRes web server is available at http://gene.rnet.missouri.edu/soykb_dev/IMPRes/.
Yuexu Jiang, Yanchun Liang 0001, Duolin Wang, Dong Xu 0002, Trupti Joshi
BIBM5
2017 Enabling precision medicine with CancerKB and KBCommons informatics framework
abstract
Precision Medicine is one of the most rapidly evolving areas within medicine, which allows utilization of patient's own genomic information in guiding personalized treatments and customized therapy. Applying the public data obtained from human genomic and cancer cell line studies in the diagnosis and treatment of diseases is a challenging task as these data are spread across different repositories with diverse formats and omics data types. There is an immediate need for platforms to make this integration and analysis seamless and readily accessible to researchers and clinicians. To achieve this, we have implemented Cancer Cell Line KB, a branch of the KBCommons for homoSapiens which focuses on automatically establishing a web-based resource for comprehensive multi-omics data access using our in-house developed KBCommons framework. It currently provides data for human genes / proteins, miRNA, and publicly available information for 1046 cancer cell lines and drug response datasets. Additionally, gene expression and de-identified clinical data from The Cancer Genome Atlas has also been added which includes 7706 patient samples for 20 cancer types, and 23368 genes expression as FPKM and FeatureCounts. Users can view these datasets using various analytical and graphical visualization tools in Cancer Cell Line KB and search by genes or cancer cell lines. We have also developed Principal Component Analysis (PCA) tool and patient clinical data browser to provide access to patient's genomic tests information and gene expression. In the future, heatmap will be developed to analysis the microarray gene expression data with hierarchical clustering method and more de-identified clinical data from cBioportal can be imported through WEB-API. We also plan to implement a tool to show how the patient's own genomic tests information compares with the population based datasets for that disease traits.
Zhen Lyu, Trupti Joshi
BIBM3
2017 Development of "KBCommons" - universal informatics framework for multi-omics translational research
abstract
Multi-level ‘OMICS’ data integration for multiple organisms has been one of the major challenges in the era of advanced next generation sequencing and high performance technologies. However, these data are often stored individually across different web resources based on data type and organism making it difficult to find and integrate them. There are many websites which stores different data types and display data in pie charts or plain text format but limit their data to only one fixed organism. Making it difficult for researchers working on other biological organisms including plants, animals, humans, and microbes have similar needs with multi-level omics data. These complex omics data requires extensive data management, exhaustive computational analysis, and effective integration to have a one-stop interactive, web-based portal to browse, access, analyze, integrate, visualize and share knowledge about genomics and molecular mechanisms, with ultimate links to phenotypes and traits. To achieve this, we have developed Knowledge Base Commons (KBCommons), a platform that automates the process of establishing the database and making the tools for organisms available via a dedicated web resource.
Siva Ratna Kumari Narisetti, Zhen Lyu, Trupti Joshi
BIBM4
2017 Development of an informatics analytics workflow for DAP-seq data exploration and validation for auxin response factors in maize
abstract
DNA affinity purification sequencing (DAP-seq) is a recently developed technique for transcription factor (TF) binding site discovery that produces datasets like ChIP-seq. A major advantage of the DAP-seq method is that it uses exogenously expressed TFs to directly interrogate genomic DNA, without the need for tagged transgenic lines or gene-specific antibodies while still capturing TF binding events in their genomic sequence context. To assess the accuracy of the DAP-seq, we utilized this method to generate genome wide binding profiles of maize AUXIN RESPONSE FACTORS (ARFs). ARFs are responsible for activating or repressing auxin response genes and play an important role in growth and developmental processes. This provides a typical scenario in which researchers would use DAP-seq to better understand how this important family of TFs regulates gene expression. The informatics analysis workflow consists of a selection of highly validated read aligners and transcription factor binding site prediction bioinformatics tools supported by in-house built custom python scripts. We investigate the accuracy of pattern mining underlying the ARF binding signatures with respect to the presence and position of conserved motifs. ARFs are known to bind as dimers to pairs of TGTC motifs as direct repeats, inverted repeats and everted repeats. Based on this knowledge, our workflow mines the DAP-seq datasets to find specific genomic regions with such signatures. After each round of mining, we validate the accuracy of results with known patterns and domain knowledge from our collaborators.
Sidharth Sen, Mary Galli, Andrea Gallavotti, Trupti Joshi
BIBM4
2017 SoyTSN: A web-based prediction tool for soybean tissue specific network within SoyKB
abstract
Soybean tissue-specific network helps identify and visualize the gene-gene relationships in various tissues [1]. We have built SoyTSN, a web-based tool for tissue-specific network prediction in soybean using 14 tissues RNA-Seq datasets including flower, root, nodule, leaf, stem, seed, etc. SoyTSN first combines multiple tissue specific RNA-Seq studies, and later Cross-Conditions Cluster Detection (C3D) algorithm [2] was applied to detect modules based on co-expression relationships across all the tissues. Following this soybean tissue-specific interactomes were inferred by combining tissue-specific expression and protein-protein interaction from the STRING database [3]. All these relationships between various soybean genes were collected and stored in Soybean Knowledge Base (SoyKB)[4] and can be queried using SoyTSN. For every query gene, SoyTSN computes and visualizes any of the 14 tissue specific networks both at the expression and interactome level. Users can compare the gene-gene relationship differences at different confidence levels across all the soybean tissues.
Juexin Wang, Zhen Lyu, Shakhawat Hossain, Gary Stacey, Dong Xu 0002, Trupti Joshi
BIBM6
2017 KBCommons: A multi 'OMICS' integrative framework for database and informatics tools
abstract
Advancement of next generation sequencing and high-throughput technologies has resulted in generation of multi-level of `OMICS' data for many organisms. However, these data are often individually scattered across different repositories based on data type, making it difficult to integrate them. We have addressed this issue through our in-house developed Soybean Knowledge Base[1,2](SoyKB) framework, a comprehensive web-based resource. It acts as a centralized repository for soybean multi-omics data, and is equipped with an array of bioinformatics analytical and graphical visualization tools. It is available at http://soykb.org and has proven to be a great success with more than 500 registered users. Users working on other biological organisms including plants, animals and biomedical diseases have similar needs and the developed framework can be expanded to make the visualization and analysis tools function for other organisms, without having to reinvent the wheel. To achieve this we have developed KBCommons, a platform that automates the process of establishing the database and making the tools for other organisms available via a dedicated web resource. It provides information for six entities including genes/proteins, microRNAs/sRNAs, metabolites, SNP, traits as well as plant introduction or strains/populations. It also incorporates several multi-omics datasets including transcriptomics, proteomics, metabolomics, epigenomics, molecular breeding and other types. We have currently expanded KBCommons framework and tools to Zea mays, Arabidopsis, Mus musculus and Homo sapiens. We have integrated various genomics dataset for maize including RNAseq B73 mutants and Tassel meristem from our collaborators. It provides a suite of tools such as the gene/metabolite pathway viewer, Protein Bio-Viewer, heatmaps, scatter plots and hierarchical clustering. It also provides access to PGen, Pegasus analytics workflows developed for genomics variations analysis. It also has suite of tools for differential expression analysis of transcriptomics and other multi-omics datasets including venn diagrams, volcano plots, function enrichment and gene modules.
Siva Ratna Kumari Narisetti, Zhen Lyu, Trupti Joshi
BIBM4
2017 MusiteDeep: a deep-learning framework for general and kinase-specific phosphorylation site prediction
abstract
MOTIVATION: Computational methods for phosphorylation site prediction play important roles in protein function studies and experimental design. Most existing methods are based on feature extraction, which may result in incomplete or biased features. Deep learning as the cutting-edge machine learning method has the ability to automatically discover complex representations of phosphorylation patterns from the raw sequences, and hence it provides a powerful tool for improvement of phosphorylation site prediction. RESULTS: We present MusiteDeep, the first deep-learning framework for predicting general and kinase-specific phosphorylation sites. MusiteDeep takes raw sequence data as input and uses convolutional neural networks with a novel two-dimensional attention mechanism. It achieves over a 50% relative improvement in the area under the precision-recall curve in general phosphorylation site prediction and obtains competitive results in kinase-specific prediction compared to other well-known tools on the benchmark data. AVAILABILITY AND IMPLEMENTATION: MusiteDeep is provided as an open-source tool available at https://github.com/duolinwang/MusiteDeep. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Duolin Wang, Wangren Qiu, Yanchun Liang 0001, Trupti Joshi, Dong Xu 0002
Bioinform.6
2016 End-to-End Security Formalization and Alignment for Federated Workflow Management
abstract
Traditionally, the allocation and dynamic adaptation of federated cyberinfrastructure resources residing across multiple domains for data-intensive application workflows have been performance or quality of service-centric (i.e., QSpecs), often compromising the end-to-end security requirements of scientific workflows. Lack of standardized formalization methods of the workflows' end-to-end security requirements, and diverse/heterogenous domain resource and security policies make inter-conflict characterization between application's security and performance requirements non-trivial, and leads to sub-optimal resource allocation. In this paper, we present a joint security and performance-driven federated resource allocation and adaptation scheme to define and characterize a data-intensive scientific application's security specifications (i.e., SSpecs). In order to aid security-driven resource brokering among domains with diverse security postures, we describe an alignment technique inspired by Portunes Algebra to combine domain-specific resource policies (i.e., RSpecs) along the application workflow life cycle. We use standardized guidelines that help in compute/storage resource domain/location selection as well as network path selection based on both application QSpecs and SSpecs. We implement our security formalization and alignment methods as a framework, viz., "OnTimeURB" and apply it on an exemplar Distributed Computing workflow to show the benefits of joint QSpecs-SSpecs-driven, RSpecs-compliant federated workflow management.
Matthew Dickinson, Saptarshi Debroy, Prasad Calyam, Samaikya Valluripally, Yuanxun Zhang, Trupti Joshi, Dong Xu 0002
CLOUD6
2016 Complexity Reduction and Visualization of RDF Knowledge Networks for Precision medicine
Zainab Al-Taie, Nattaphon Thanintorn, Ilker Ersoy, Richard D. Hammer, Dong Xu 0002, Trupti Joshi, Dmitriy Shin
AMIA6
2016 PGen: large-scale genomic variations analysis workflow and browser in SoyKB
abstract
BACKGROUND: With the advances in next-generation sequencing (NGS) technology and significant reductions in sequencing costs, it is now possible to sequence large collections of germplasm in crops for detecting genome-scale genetic variations and to apply the knowledge towards improvements in traits. To efficiently facilitate large-scale NGS resequencing data analysis of genomic variations, we have developed "PGen", an integrated and optimized workflow using the Extreme Science and Engineering Discovery Environment (XSEDE) high-performance computing (HPC) virtual system, iPlant cloud data storage resources and Pegasus workflow management system (Pegasus-WMS). The workflow allows users to identify single nucleotide polymorphisms (SNPs) and insertion-deletions (indels), perform SNP annotations and conduct copy number variation analyses on multiple resequencing datasets in a user-friendly and seamless way. RESULTS: We have developed both a Linux version in GitHub ( https://github.com/pegasus-isi/PGen-GenomicVariations-Workflow ) and a web-based implementation of the PGen workflow integrated within the Soybean Knowledge Base (SoyKB), ( http://soykb.org/Pegasus/index.php ). Using PGen, we identified 10,218,140 single-nucleotide polymorphisms (SNPs) and 1,398,982 indels from analysis of 106 soybean lines sequenced at 15X coverage. 297,245 non-synonymous SNPs and 3330 copy number variation (CNV) regions were identified from this analysis. SNPs identified using PGen from additional soybean resequencing projects adding to 500+ soybean germplasm lines in total have been integrated. These SNPs are being utilized for trait improvement using genotype to phenotype prediction approaches developed in-house. In order to browse and access NGS data easily, we have also developed an NGS resequencing data browser ( http://soykb.org/NGS_Resequence/NGS_index.php ) within SoyKB to provide easy access to SNP and downstream analysis results for soybean researchers. CONCLUSION: PGen workflow has been optimized for the most efficient analysis of soybean data using thorough testing and validation. This research serves as an example of best practices for development of genomics data analysis workflows by integrating remote HPC resources and efficient data management with ease of use for biological users. PGen workflow can also be easily customized for analysis of data in other species.
Saad M. Khan, Juexin Wang, Mats Rynge, Yuanxun Zhang, Shiyuan Chen, João V. Maldonado dos Santos, Babu Valliyodan, Prasad Calyam, Nirav C. Merchant, Henry T. Nguyen, Dong Xu 0002, Trupti Joshi
BMC Bioinform.14
2013 Soybean knowledge base (SoyKB): Bridging the gap between soybean translational genomics and breeding
abstract
Many genome-scale data are available in soybean including genomic sequence, transcriptomics (microarray, RNA-seq), proteomics and metabolomics datasets, together with growing knowledge of soybean in gene, microRNAs, pathways, and phenotypes. This represents rich and resourceful information which can provide valuable insights, if mined in an innovative and integrative manner and thus, the need for informatics resources to achieve that. Towards this we have developed Soybean Knowledge Base (SoyKB), a comprehensive all-inclusive web resource for soybean translational genomics and breeding. SoyKB handles the management and integration of soybean genomics and multi-omics data along with gene function annotations, biological pathway and trait information. It has many useful tools including Affymetrix probelD search, gene family search, multiple gene/metabolite analysis, motif analysis tool, protein 3D structure viewer and download/upload capacity for experimental data and annotations. It has a user-friendly web interface together with genome browser and pathway viewer, which display data in an intuitive manner to the soybean researchers, breeders and consumers. SoyKB has new innovative tools for soybean breeding including a graphical chromosome visualizer targeted towards ease of navigation for breeders. It integrates QTLs, traits, germplasm information along with genomic variation data such as single nucleotide polymorphisms (SNPs) and genome-wide association studies (GWAS) data from multiple genotypes, cultivars and G. soja. QTLs for multiple traits can be queried and visualized in the chromosome visualizer simultaneously and overlaid on top of the genes and other molecular markers as well as multi-omics experimental data for meaningful inferences. SoyKB can be publicly accessed at http://soykb.org.
Trupti Joshi, Michael R. Fitzpatrick, Shiyuan Chen, Ryan Z. Endacott, Eric C. Gaudiello, Gary Stacey, Henry T. Nguyen, Dong Xu 0002
BIBM1
2010 SoyMetDB: The soybean metabolome database
abstract
SoyMetDB is a metabolomic database for soybean, developed to target the growing needs of the soybean community. The goal is to provide a one-stop web resource for integrating, mining and visualizing soybean metabolomic data, including identification and expression of various metabolites across different experiments and time courses. It incorporates GC-MS and LC-MS based metabolite-profiling data dynamically linked to metabolite information from other public metabolomic databases, including HMDB and Knapsack. SoyMetDB includes Arabidopsis metabolomic data for cross-species comparisons and can retrieve information including the expression patterns of various experiments for complete or partial metabolite name queries. It also incorporates a pathway viewer tool integrating the data from various experimental conditions and presenting them on the pathways to highlight the expressed metabolite, and identifies the most highly represented pathways for multiple metabolite queries. SoyMetDB can be accessed at http://soymetdb.org.
Trupti Joshi, Qiuming Yao, D. Franklin Levi, Laurent Brechenmacher, Babu Valliyodan, Gary Stacey, Henry T. Nguyen, Dong Xu 0002
BIBM1
2010 Prediction of novel miRNAs and associated target genes in Glycine max
abstract
BACKGROUND: Small non-coding RNAs (21 to 24 nucleotides) regulate a number of developmental processes in plants and animals by silencing genes using multiple mechanisms. Among these, the most conserved classes are microRNAs (miRNAs) and small interfering RNAs (siRNAs), both of which are produced by RNase III-like enzymes called Dicers. Many plant miRNAs play critical roles in nutrient homeostasis, developmental processes, abiotic stress and pathogen responses. Currently, only 70 miRNA have been identified in soybean. METHODS: We utilized Illumina's SBS sequencing technology to generate high-quality small RNA (sRNA) data from four soybean (Glycine max) tissues, including root, seed, flower, and nodules, to expand the collection of currently known soybean miRNAs. We developed a bioinformatics pipeline using in-house scripts and publicly available structure prediction tools to differentiate the authentic mature miRNA sequences from other sRNAs and short RNA fragments represented in the public sequencing data. RESULTS: The combined sequencing and bioinformatics analyses identified 129 miRNAs based on hairpin secondary structure features in the predicted precursors. Out of these, 42 miRNAs matched known miRNAs in soybean or other species, while 87 novel miRNAs were identified. We also predicted the putative target genes of all identified miRNAs with computational methods and verified the predicted cleavage sites in vivo for a subset of these targets using the 5' RACE method. Finally, we also studied the relationship between the abundance of miRNA and that of the respective target genes by comparison to Solexa cDNA sequencing data. CONCLUSION: Our study significantly increased the number of miRNAs known to be expressed in soybean. The bioinformatics analysis provided insight on regulation patterns between the miRNAs and their predicted target genes expression. We also deposited the data in a soybean genome browser based on the UCSC Genome Browser architecture. Using the browser, we annotated the soybean data with miRNA sequences from four tissues and cDNA sequencing data. Overlaying these two datasets in the browser allows researchers to analyze the miRNA expression levels relative to that of the associated target genes. The browser can be accessed at http://digbio.missouri.edu/soybean_mirna/.
Trupti Joshi, Marc Libault, Dong-Hoon Jeong, Sunhee Park, Pamela J. Green, D. Janine Sherrier, Andrew Farmer, Greg May, Blake C. Meyers, Dong Xu 0002, Gary Stacey
BMC Bioinform.1
2006 Supervised Inference of Gene Regulatory Networks by Linear Programming
Yong Wang 0001, Trupti Joshi, Dong Xu 0002, Xiang-Sun Zhang, Luonan Chen
ICIC (3)2
2006 Inferring gene regulatory networks from multiple microarray datasets
abstract
MOTIVATION: Microarray gene expression data has increasingly become the common data source that can provide insights into biological processes at a system-wide level. One of the major problems with microarrays is that a dataset consists of relatively few time points with respect to a large number of genes, which makes the problem of inferring gene regulatory network an ill-posed one. On the other hand, gene expression data generated by different groups worldwide are increasingly accumulated on many species and can be accessed from public databases or individual websites, although each experiment has only a limited number of time-points. RESULTS: This paper proposes a novel method to combine multiple time-course microarray datasets from different conditions for inferring gene regulatory networks. The proposed method is called GNR (Gene Network Reconstruction tool) which is based on linear programming and a decomposition procedure. The method theoretically ensures the derivation of the most consistent network structure with respect to all of the datasets, thereby not only significantly alleviating the problem of data scarcity but also remarkably improving the prediction reliability. We tested GNR using both simulated data and experimental data in yeast and Arabidopsis. The result demonstrates the effectiveness of GNR in terms of predicting new gene regulatory relationship in yeast and Arabidopsis. AVAILABILITY: The software is available from http://zhangorup.aporc.org/bioinfo/grninfer/, http://digbio.missouri.edu/grninfer/ and http://intelligent.eic.osaka-sandai.ac.jp or upon request from the authors.
Yong Wang 0001, Trupti Joshi, Xiang-Sun Zhang, Dong Xu 0002, Luonan Chen
Bioinform.2
2003 Towards Automated Derivation of Biological Pathways Using High-Throughput Biological Data
abstract
Characterizing biological pathways at the genome scale is one of the most important and challenging tasks in the post genomic era. To address this challenge, we have developed a computational method to systematically and automatically derive partial biological pathways in yeast using high-throughput biological data, including yeast two hybrid data, protein complexes identified from mass spectroscopy, genetics interactions, and microarray gene expression data in yeast Saccharomyces cerevisiae. The inputs of the method are the upstream starting protein (e.g., a sensor of a signal) and the downstream terminal protein (e.g., a transcriptional factor that induces genes to respond the signal); the output of the method is the protein interaction chain between the two proteins. The high-throughput data are coded into a graph of interaction network, where each node represents a protein. The weight of an edge between two nodes models the "closeness" of the two represented proteins in the interaction network and it is defined by a rule-based formula according to the high-throughput data and modified by the protein function classification and subcellular localization information. The protein interaction cascade pathway in vivo is predicted as the shortest path identified from the graph of the interaction network using Dijkstra's algorithm. We have also developed a web server of this method (http://compbio.ornl.gov/structure/pathway) for public use. To our knowledge, our method is the first automated method to generally construct partial biological pathways using a suite of high-throughput biological data. This work demonstrates the proof of principle using computational approaches for discoveries of biological pathways with high-throughput data and biological annotation data.
Trupti Joshi, Ying Xu 0001, Dong Xu 0002
BIBE2