EDBT 2026 Demo / reviewers in the wild / expert
Ravi K. Madduri
dblp:02/1272
· DBLP profile ↗
36ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-2130-2887ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 7 since 2021Systems, architecture and hardware · 10 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 9 · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUsabstractMOTIVATION: Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. RESULTS: We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems. Alex Rodriguez, Youngdae Kim, Tarak Nath Nandi, Karl Keat, Rachit Kumar, Mitchell Conery, Rohan Bhukar, Molei Liu, John Hessington, Ketan Maheshwari, VA Million Veteran Program, Edmon Begoli, Georgia Tourassi, Pradeep Natarajan, Benjamin F. Voight, John Michael Gaziano, Scott M. Damrauer, Katherine P. Liao, Jennifer E. Huffman, Anurag Verma, Ravi K. Madduri |
Bioinform. | 22 |
| 2026 | Increasing value in the Veterans Affairs Healthcare System (VA) with precision health: a continuing landmark collaboration with the Department of EnergyabstractOBJECTIVE: Phase II of MVP-CHAMPION, a federal collaboration between the Veterans Affairs Healthcare System (VA) and the Department of Energy (DoE), leveraged large-scale clinical, geo-spatial, and genetic data with state-of-the-art artificial intelligence (AI), and high-performance computing (HPC) to improve value in healthcare. MATERIALS AND METHODS: Eight clinical priority projects for which AI was a critical missing capability were initiated to address: lung cancer screening (MVP 061), suicide risk screening (MVP 062), cardiovascular risk in obstructive sleep apnea (MVP 063), checkpoint inhibitor toxicity (MVP 064), heart failure (MVP 065), renal complications in diabetes (MVP 066), post COVID-19 sequelae (MVP 067), and antipsychotic medication toxicity (MVP 068). RESULTS: Building on a strong regulatory and administrative foundation, we developed multimorbidity-aware analytic frameworks, reusable computational tools, and analytic pipelines. These greatly facilitated identification of novel risk factors including genetic variants and specification of more discriminating prediction models. Novel genetic risk factors are informing development and repurposing of medications and discriminating prediction models promise to improve healthcare value. DISCUSSION: The research foundation developed in Phase I and extended in Phase II of MVP CHAMPION has supported an unprecedented federal collaboration and yielded significant scientific advances. Our clinical findings are poised for near-term application, while advances in machine learning and high-performance computing may accelerate the broader adoption of artificial intelligence in healthcare. CONCLUSION: This maturing VA-DoE federal collaboration is poised to transform the future of Veterans' healthcare and the broader national landscape of precision health. Amy Justice, Benjamin H. McMahon, Daniel A. Jacobson, Kelly Cho, Anuj J. Kapadia, Samuel M. Aguayo, Zeynep H. Gümüs, Ioana Danciu, Jean C. Beckham, Nathan A. Kimbrel, Silvia Crivelli, Eilis A. Boudreau, Patrick D. Finley, Alex K. Bryant, Shinjae Yoo, Jacob Joseph, Peter Reaven, Shiuh-Wen Luoh, Ravi K. Madduri, Ayman Fanous, Khushbu Agarwal, Harshini Mukundan, Sumitra Muralidhar |
J. Am. Medical Informatics Assoc. | 21 |
| 2025 | Advances in Appfl: a Comprehensive and Extensible Federated Learning FrameworkabstractFederated learning (FL) is a distributed machine learning paradigm enabling collaborative model training while preserving data privacy. In today's landscape, where most data is proprietary, confidential, and distributed, FL has become a promising approach to leverage such data effectively, particularly in sensitive domains such as medicine and the electric grid. Heterogeneity and security are the key challenges in FL, however, most existing FL frameworks either fail to address these challenges adequately or lack the flexibility to incorporate new solutions. To this end, we present the recent advances in developing Appfl, an extensible framework and benchmarking suite for federated learning, which offers comprehensive solutions for heterogeneity and security concerns, as well as user-friendly interfaces for integrating new algorithms or adapting to new applications. We demonstrate the capabilities of Appfl through extensive experiments evaluating various aspects of FL, including communication efficiency, privacy preservation, computational performance, and resource utilization. We further highlight the extensibility of Appfl through case studies in vertical, hierarchical, and decentralized FL. Appfl is fully open-sourced on Github at https://github.com/APPFL/APPFL. Zilinghan Li, Shilan He, Minseok Ryu, Kibaek Kim, Ravi K. Madduri |
CCGrid | 6 |
| 2025 | CADRE: Customizable Assurance of Data Readiness in Privacy-Preserving Federated LearningabstractPrivacy-Preserving Federated Learning (PPFL) is a decentralized machine learning approach where multiple clients train a model collaboratively. PPFL preserves the privacy and security of a client’s data without exchanging it. However, ensuring that data at each client is of high quality and ready for federated learning (FL) is a challenge due to restricted data access. In this paper, we introduce CADRE (Customizable Assurance of Data REadiness) for federated learning (FL), a novel framework that allows users to define custom data readiness (DR) metrics, rules, and remedies tailored to specific FL tasks. CADRE generates comprehensive DR reports based on the user-defined metrics, rules, and remedies to ensure datasets are prepared for FL while preserving privacy. We demonstrate a practical application of CADRE by integrating it into an existing PPFL framework. We conducted experiments across six datasets and addressed seven different DR issues. The results illustrate the versatility and effectiveness of CADRE in ensuring DR across various dimensions, including data quality, privacy, and fairness. This approach enhances the performance and reliability of FL models as well as utilizes valuable resources. Kaveen Hiniduma, Zilinghan Li, Aditya Sinha, Ravi K. Madduri, Surendra Byna |
eScience | 4 |
| 2025 | Cost-Aware Federated Learning on the CloudabstractWe introduce FedCostAware, a cost-aware scheduling algorithm designed to optimize synchronous federated learning (FL) on cloud spot instances, which addresses the challenges of training on spot instances and different client budgets by employing intelligent management of the lifecycle of spot instances. This approach minimizes idle resource time and overall expenses. Experiments on real-world medical datasets demonstrate that FedCostAware significantly reduces cloud computing costs compared to conventional spot and on-demand schemes, enhancing the accessibility and affordability of FL. Aditya Sinha, Zilinghan Li, Tingkai Liu, Volodymyr V. Kindratenko, Kibaek Kim, Ravi K. Madduri |
eScience | 6 |
| 2025 | Pathology Image Compression with Pre-trained Autoencoders
Srikar Yellapragada, Alexandros Graikos, Kostas Triaridis, Zilinghan Li, Tarak Nath Nandi, Ravi K. Madduri, Prateek Prasanna, Joel H. Saltz, Dimitris Samaras |
MICCAI (2) | 6 |
| 2024 | Privacy-Preserving Federated Learning for Science: Challenges and Research DirectionsabstractThis paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific artificial intelligence models, in particular, foundation models (FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy—an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains. Kibaek Kim, Raghavan Krishnan, Olivera Kotevska, Matthieu Dorier, Ravi K. Madduri, Minseok Ryu, Todd S. Munson, Robert B. Ross, Thomas Flynn 0001, Ai Kagawa, Byung-Jun Yoon, Christian Engelmann, Farzad Yousefian |
IEEE Big Data | 5 |
| 2024 | FedCompass: Efficient Cross-Silo Federated Learning on Heterogeneous Client Devices Using a Computing Power-Aware SchedulerabstractCross-silo federated learning offers a promising solution to collaboratively train robust and generalized AI models without compromising the privacy of local datasets, e.g., healthcare, financial, as well as scientific projects that lack a centralized data facility. Nonetheless, because of the disparity of computing resources among different clients (i.e., device heterogeneity), synchronous federated learning algorithms suffer from degraded efficiency when waiting for straggler clients. Similarly, asynchronous federated learning algorithms experience degradation in the convergence rate and final model accuracy on non-identically and independently distributed (non-IID) heterogeneous datasets due to stale local models and client drift. To address these limitations in cross-silo federated learning with heterogeneous clients and data, we propose FedCompass, an innovative semi-asynchronous federated learning algorithm with a computing power-aware scheduler on the server side, which adaptively assigns varying amounts of training tasks to different clients using the knowledge of the computing power of individual clients. FedCompass ensures that multiple locally trained models from clients are received almost simultaneously as a group for aggregation, effectively reducing the staleness of local models. At the same time, the overall training process remains asynchronous, eliminating prolonged waiting periods from straggler clients. Using diverse non-IID heterogeneous distributed datasets, we demonstrate that FedCompass achieves faster convergence and higher accuracy than other asynchronous algorithms while remaining more efficient than synchronous algorithms when performing federated learning on heterogeneous clients. The source code for FedCompass is available at https://github.com/APPFL/FedCompass. Zilinghan Li, Pranshu Chaturvedi, Shilan He, Volodymyr V. Kindratenko, Eliu A. Huerta, Kibaek Kim, Ravi K. Madduri |
ICLR | 9 |
| 2024 | AI Data Readiness Inspector (AIDRIN) for Quantitative Assessment of Data Readiness for AIabstractGarbage In Garbage Out is a universally agreed quote by computer scientists from various domains, including Artificial Intelligence (AI). As data is the fuel for AI, models trained on low-quality, biased data are often ineffective. Computer scientists who use AI invest a considerable amount of time and effort in preparing the data for AI. However, there are no standard methods or frameworks for assessing the “readiness” of data for AI. To provide a quantifiable assessment of the readiness of data for AI processes, we define parameters of AI data readiness and introduce AIDRIN (AI Data Readiness INspector). AIDRIN is a framework covering a broad range of readiness dimensions available in the literature that aid in evaluating the readiness of data quantitatively and qualitatively. AIDRIN uses metrics in traditional data quality assessment such as completeness, outliers, and duplicates for data evaluation. Furthermore, AIDRIN uses metrics specific to assess data for AI, such as feature importance, feature correlations, class imbalance, fairness, privacy, and FAIR (Findability, Accessibility, Interoperability, and Reusability) principle compliance. AIDRIN provides visualizations and reports to assist data scientists in further investigating the readiness of data. The AIDRIN framework enhances the efficiency of the machine learning pipeline to make informed decisions on data readiness for AI applications. Kaveen Hiniduma, Surendra Byna, Jean Luca Bez, Ravi K. Madduri |
SSDBM | 4 |
| 2023 | APPFLx: Providing Privacy-Preserving Cross-Silo Federated Learning as a ServiceabstractCross-silo privacy-preserving federated learning (PPFL) is a powerful tool to collaboratively train robust and generalized machine learning (ML) models without sharing sensitive (e.g., healthcare of financial) local data. To ease and accelerate the adoption of PPFL, we introduce APPFLx, a ready-to-use platform that provides privacy-preserving cross-silo federated learning as a service. APPFLx employs Globus authentication to allow users to easily and securely invite trustworthy collaborators for PPFL, implements several synchronous and asynchronous FL algorithms, streamlines the FL experiment launch process, and enables tracking and visualizing the life cycle of FL experiments, allowing domain experts and ML practitioners to easily orchestrate and evaluate cross-silo FL under one platform. APPFLx is available online at https://appflx.link Zilinghan Li, Shilan He, Pranshu Chaturvedi, Trung-Hieu Hoang, Minseok Ryu, Eliu A. Huerta, Volodymyr V. Kindratenko, Jordan D. Fuhrman, Maryellen L. Giger, Ryan Chard, Kibaek Kim, Ravi K. Madduri |
e-Science | 12 |
| 2020 | Using the FACE-IT portal and workflow engine for operational food quality prediction and assessment: An application to mussel farms monitoring in the Bay of Napoli, Italy
Raffaele Montella, Alison Brizius, Diana Di Luccio, Cheryl H. Porter, Joshua Elliott, Ravi K. Madduri, David Kelly, Angelo Riccio, Ian T. Foster |
Future Gener. Comput. Syst. | 6 |
| 2018 | Scalable pCT Image Reconstruction Delivered as a Cloud ServiceabstractWe describe a cloud-based medical image reconstruction service designed to meet a real-time and daily demand to reconstruct thousands of images from proton cancer treatment facilities worldwide. Rapid reconstruction of a three-dimensional Proton Computed Tomography (pCT) image can require the transfer of 100 GB of data and use of approximately 120 GPU-enabled compute nodes. The nature of proton therapy means that demand for such a service is sporadic and comes from potentially hundreds of clients worldwide. We thus explore the use of a commercial cloud as a scalable and cost-efficient platform for pCT reconstruction. To address the high performance requirements of this application we leverage Amazon Web Services' GPU-enabled cluster resources that are provisioned with high performance networks between nodes. To support episodic demand, we develop an on-demand multi-user provisioning service that can dynamically provision and resize clusters based on image reconstruction requirements, priorities, and wait times. We compare the performance of our pCT reconstruction service running on commercial cloud resources with that of the same application on dedicated local high performance computing resources. We show that we can achieve scalable and on-demand reconstruction of large scale pCT images for simultaneous multi-client requests, processing images in less than 10 minutes for less than $10 per image. Ryan Chard, Ravi K. Madduri, Nicholas T. Karonis, Kyle Chard, Kirk L. Duffin, Caesar E. Ordoñez, Thomas D. Uram, Justin Fleischauer, Ian T. Foster, Michael E. Papka, John Winans |
IEEE Trans. Cloud Comput. | 2 |
| 2017 | Developing a framework for digital objects in the Big Data to Knowledge (BD2K) commons: Report from the Commons Framework Pilots workshop
Kathleen M. Jagodnik, Simon Koplev, Sherry L. Jenkins, Lucila Ohno-Machado, Benedict Paten, Stephan C. Schürer, Michel Dumontier, Ruben Verborgh, Alex Bui, Peipei Ping, Neil J. McKenna, Ravi K. Madduri, Ajay Pillai, Avi Ma'ayan |
J. Biomed. Informatics | 12 |
| 2016 | I'll take that to go: Big data bags and minimal identifiers for exchange of large, complex datasetsabstractBig data workflows often require the assembly and exchange of complex, multi-element datasets. For example, in biomedical applications, the input to an analytic pipeline can be a dataset consisting thousands of images and genome sequences assembled from diverse repositories, requiring a description of the contents of the dataset in a concise and unambiguous form. Typical approaches to creating datasets for big data workflows assume that all data reside in a single location, requiring costly data marshaling and permitting errors of omission and commission because dataset members are not explicitly specified. We address these issues by proposing simple methods and tools for assembling, sharing, and analyzing large and complex datasets that scientists can easily integrate into their daily workflows. These tools combine a simple and robust method for describing data collections (BDBags), data descriptions (Research Objects), and simple persistent identifiers (Minids) to create a powerful ecosystem of tools and services for big data analysis and sharing. We present these tools and use biomedical case studies to illustrate their use for the rapid assembly, sharing, and analysis of large datasets. Kyle Chard, Mike D'Arcy, Benjamin D. Heavner, Ian T. Foster, Carl Kesselman, Ravi K. Madduri, Alexis A. Rodriguez, Stian Soiland-Reyes, Carole A. Goble, Kristi Clark, Eric W. Deutsch, Ivo D. Dinov, Nathan D. Price 0001, Arthur W. Toga |
IEEE BigData | 6 |
| 2016 | An Automated Tool Profiling Service for the CloudabstractCloud providers offer a diverse set of instance types with varying resource capacities, designed to meet the needs of a broad range of user requirements. While this flexibility is a major benefit of the cloud computing model, it also creates challenges when selecting the most suitable instance type for a given application. Sub-optimal instance selection can result in poor performance and/or increased cost, with significant impacts when applications are executed repeatedly. Yet selecting an optimal instance type is challenging, as each instance type can be configured differently, application performance is dependent on input data and configuration, and instance types and applications are frequently updated. We present a service that supports automatic profiling of application performance on different instance types to create rich application profiles that can be used for comparison, provisioning, and scheduling. This service can dynamically provision cloud instances, automatically deploy and contextualize applications, transfer input datasets, monitor execution performance, and create a composite profile with fine grained resource usage information. We use real usage data from four production genomics gateways and estimate the use of profiles in autonomic provisioning systems can decrease execution time by up to 15.7% and cost by up to 86.6%. Ryan Chard, Kyle Chard, Bryan C. K. Ng, Kris Bubendorfer, Alexis A. Rodriguez, Ravi K. Madduri, Ian T. Foster |
CCGrid | 6 |
| 2015 | Cost-Aware Elastic Cloud Provisioning for Scientific WorkloadsabstractCloud computing provides an efficient model to host and scale scientific applications. While cloud-based approaches can reduce costs as users pay only for the resources used, it is often challenging to scale execution both efficiently and cost-effectively. We describe here a cost-aware elastic cloud provisioner designed to elastically provision cloud infrastructure to execute analyses cost-effectively. The provisioner considers real-time spot instance prices across availability zones, leverages application profiles to optimize instance type selection, over-provisions resources to alleviate bottlenecks caused by oversubscribed instance types, and is capable of reverting to on-demand instances when spot prices exceed thresholds. We evaluate the usage of our cost-aware provisioner using four production scientific gateways and show that it can produce cost savings of up to 97.2% when compared to naive provisioning approaches. Ryan Chard, Kyle Chard, Kris Bubendorfer, Lukasz Lacinski, Ravi K. Madduri, Ian T. Foster |
CLOUD | 5 |
| 2015 | Cost-Aware Cloud ProvisioningabstractCloud computing is often suggested as a low-cost and scalable model for executing and scaling scientific analyses. However, while the benefits of cloud computing are frequently touted, there are inherent technical challenges associated with scaling execution efficiently and cost-effectively. We describe here a cost-aware elastic provisioner designed to dynamically and cost-effectively provision cloud infrastructure based on the requirements of user-submitted scientific workflows. Our provisioner is used in the Globus Galaxies platform -- a Software-as-a-Service provider of scientific analysis capabilities using commercial cloud infrastructure. Using workloads from production usage of this platform we investigate the performance of our provisioner in terms of cost, spot instance termination rate, and execution time. We demonstrate cost savings across six production gateways of up to 95% and 12% improvement in total execution time when compared to a worst case scenario using a single instance type in a single availability zone. Ryan Chard, Kyle Chard, Kris Bubendorfer, Lukasz Lacinski, Ravi K. Madduri, Ian T. Foster |
e-Science | 5 |
| 2015 | Consensus Genotyper for Exome Sequencing (CGES): improving the quality of exome variant genotypesabstractMOTIVATION: The development of cost-effective next-generation sequencing methods has spurred the development of high-throughput bioinformatics tools for detection of sequence variation. With many disparate variant-calling algorithms available, investigators must ask, 'Which method is best for my data?' Machine learning research has shown that so-called ensemble methods that combine the output of multiple models can dramatically improve classifier performance. Here we describe a novel variant-calling approach based on an ensemble of variant-calling algorithms, which we term the Consensus Genotyper for Exome Sequencing (CGES). CGES uses a two-stage voting scheme among four algorithm implementations. While our ensemble method can accept variants generated by any variant-calling algorithm, we used GATK2.8, SAMtools, FreeBayes and Atlas-SNP2 in building CGES because of their performance, widespread adoption and diverse but complementary algorithms. RESULTS: We apply CGES to 132 samples sequenced at the Hudson Alpha Institute for Biotechnology (HAIB, Huntsville, AL) using the Nimblegen Exome Capture and Illumina sequencing technology. Our sample set consisted of 40 complete trios, two families of four, one parent-child duo and two unrelated individuals. CGES yielded the fewest total variant calls (N(CGES) = 139° 897), the highest Ts/Tv ratio (3.02), the lowest Mendelian error rate across all genotypes (0.028%), the highest rediscovery rate from the Exome Variant Server (EVS; 89.3%) and 1000 Genomes (1KG; 84.1%) and the highest positive predictive value (PPV; 96.1%) for a random sample of previously validated de novo variants. We describe these and other quality control (QC) metrics from consensus data and explain how the CGES pipeline can be used to generate call sets of varying quality stringency, including consensus calls present across all four algorithms, calls that are consistent across any three out of four algorithms, calls that are consistent across any two out of four algorithms or a more liberal set of all calls made by any algorithm. AVAILABILITY AND IMPLEMENTATION: To enable accessible, efficient and reproducible analysis, we implement CGES both as a stand-alone command line tool available for download in GitHub and as a set of Galaxy tools and workflows configured to execute on parallel computers. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vassily Trubetskoy, Alexis A. Rodriguez, Uptal J. Dave, Nicholas Campbell, Emily L. Crawford, Edwin H. Cook Jr., James S. Sutcliffe, Ian T. Foster, Ravi K. Madduri, Nancy J. Cox, Lea K. Davis |
Bioinform. | 9 |
| 2015 | The Globus Galaxies platform: delivering science gateways as a serviceabstractSummary The use of public cloud computers to host sophisticated scientific data and software is transforming scientific practice by enabling broad access to capabilities previously available only to the few. The primary obstacle to more widespread use of public clouds to host scientific software (‘cloud‐based science gateways’) has thus far been the considerable gap between the specialized needs of science applications and the capabilities provided by cloud infrastructures. We describe here a domain‐independent, cloud‐based science gateway platform, the Globus Galaxies platform, which overcomes this gap by providing a set of hosted services that directly address the needs of science gateway developers. The design and implementation of this platform leverages our several years of experience with Globus Genomics, a cloud‐based science gateway that has served more than 200 genomics researchers across 30 institutions. Building on that foundation, we have implemented a platform that leverages the popular Galaxy system for application hosting and workflow execution; Globus services for data transfer, user and group management, and authentication; and a cost‐aware elastic provisioning model specialized for public cloud resources. We describe here the capabilities and architecture of this platform, present six scientific domains in which we have successfully applied it, report on user experiences, and analyze the economics of our deployments. Published 2015. This article is a U.S. Government work and is in the public domain in the USA. Ravi K. Madduri, Kyle Chard, Ryan Chard, Lukasz Lacinski, Alexis A. Rodriguez, Dinanath Sulakhe, David Kelly, Utpal J. Dave, Ian T. Foster |
Concurr. Comput. Pract. Exp. | 1 |
| 2015 | FACE-IT: A science gateway for food security researchabstractSummary Progress in sustainability science is hindered by challenges in creating and managing complex data acquisition, processing, simulation, post‐processing, and intercomparison pipelines. To address these challenges, we developed the Framework to Advance Climate, Economic, and Impact Investigations with Information Technology (FACE‐IT) for crop and climate impact assessments. This integrated data processing and simulation framework enables data ingest from geospatial archives; data regridding, aggregation, and other processing prior to simulation; large‐scale climate impact simulations with agricultural and other models, leveraging high‐performance and cloud computing; and post‐processing to produce aggregated yields and ensemble variables needed for statistics, for model intercomparison, and to connect biophysical models to global and regional economic models. FACE‐IT leverages the capabilities of the Globus Galaxies platform to enable the capture of workflows and outputs in well‐defined, reusable, and comparable forms. We describe FACE‐IT and applications within the Agricultural Model Intercomparison and Improvement Project and the Center for Robust Decision‐making on Climate and Energy Policy. Copyright © 2015 John Wiley & Sons, Ltd. Raffaele Montella, David Kelly, Wei Xiong 0003, Alison Brizius, Joshua Elliott, Ravi K. Madduri, Ketan Maheshwari, Cheryl H. Porter, Peter Vilter, Michael Wilde, Meng Zhang 0007, Ian T. Foster |
Concurr. Comput. Pract. Exp. | 6 |
| 2015 | Big biomedical data as the key resource for discovery scienceabstractModern biomedical data collection is generating exponentially more data in a multitude of formats. This flood of complex data poses significant opportunities to discover and understand the critical interplay among such diverse domains as genomics, proteomics, metabolomics, and phenomics, including imaging, biometrics, and clinical data. The Big Data for Discovery Science Center is taking an "-ome to home" approach to discover linkages between these disparate data sources by mining existing databases of proteomic and genomic data, brain images, and clinical assessments. In support of this work, the authors developed new technological capabilities that make it easy for researchers to manage, aggregate, manipulate, integrate, and model large amounts of distributed data. Guided by biological domain expertise, the Center's computational resources and software will reveal relationships and patterns, aiding researchers in identifying biomarkers for the most confounding conditions and diseases, such as Parkinson's and Alzheimer's. Arthur W. Toga, Ian T. Foster, Carl Kesselman, Ravi K. Madduri, Kyle Chard, Eric W. Deutsch, Nathan D. Price 0001, Gwênlyn Glusman, Benjamin D. Heavner, Ivo D. Dinov, Joseph Ames, John D. Van Horn, Roger Kramer, Leroy E. Hood |
J. Am. Medical Informatics Assoc. | 4 |
| 2014 | Experiences building Globus Genomics: a next-generation sequencing analysis service using Galaxy, Globus, and Amazon Web ServicesabstractWe describe Globus Genomics, a system that we have developed for rapid analysis of large quantities of next-generation sequencing (NGS) genomic data. This system achieves a high degree of end-to-end automation that encompasses every stage of data analysis including initial data retrieval from remote sequencing centers or storage (via the Globus file transfer system); specification, configuration, and reuse of multi-step processing pipelines (via the Galaxy workflow system); creation of custom Amazon Machine Images and on-demand resource acquisition via a specialized elastic provisioner (on Amazon EC2); and efficient scheduling of these pipelines over many processors (via the HTCondor scheduler). The system allows biomedical researchers to perform rapid analysis of large NGS datasets in a fully automated manner, without software installation or a need for any local computing infrastructure. We report performance and cost results for some representative workloads. Ravi K. Madduri, Dinanath Sulakhe, Lukasz Lacinski, Bo Liu 0010, Alexis A. Rodriguez, Kyle Chard, Utpal J. Dave, Ian T. Foster |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Cloud-based bioinformatics workflow platform for large-scale next-generation sequencing analyses
Bo Liu 0010, Ravi K. Madduri, Borja Sotomayor, Kyle Chard, Lukasz Lacinski, Utpal J. Dave, Jianqiang Li 0002, Ian T. Foster |
J. Biomed. Informatics | 2 |
| 2013 | Enabling multi-task computation on Galaxy-based gateways using swiftabstractThe Galaxy science portal is a popular gateway to data analysis and computational tools for a broad range of life sciences communities. While Galaxy enables users to overcome the complexities of integrating diverse tools into unified workflows, it has only limited capabilities to execute those tools on the parallel and often distributed high-performance resources that the life sciences fields increasingly requires. We outline here an approach to meet this pressing requirement with the Swift parallel scripting language and its distributed runtime system. Swift's model of computation - implicitly parallel functional dataflow - is an elemental abstraction to which the core computing model of Galaxy maps very closely. We describe an integration between Galaxy and Swift that is transforming Galaxy into a much more powerful science gateway, retaining its user-friendly nature while extending its power to execute highly scalable workflows on diverse parallel environments. Ketan Maheshwari, Alexis A. Rodriguez, David Kelly, Ravi K. Madduri, Justin M. Wozniak, Michael Wilde, Ian T. Foster |
CLUSTER | 4 |
| 2011 | Toward Semantics Empowered Biomedical Web ServicesabstractcaGrid has accumulated a repository of biomedical services, however, how a cancer researcher can find proper services in the caGrid when needed remains a big challenge. This research aims to enhance the cyber infrastructure of caGrid, by developing a mechanism that turns caGrid services into semantic-aware interoperable services. We proposed a service semantics model, and developed a technique that automatically extracts semantic metadata from static WSDL service descriptions. Such semantic information is stored as loosely coupled annotations that can be queried using semantic Web techniques, to enhance services discovery and composition. We also proposed a two-phase discovery technique that helps users quickly identify interested service operations. This paper also reports our examinations over available techniques and recommends a feasible infrastructure for biomedical service reuse. A prototyping system is developed as a proof of concept. Jia Zhang 0001, Ravi K. Madduri, Wei Tan 0001, Kevin Deichl, John Alexander 0002, Ian T. Foster |
ICWS | 2 |
| 2011 | Enabling collaborative research using the Biomedical Informatics Research Network (BIRN)abstractOBJECTIVE: As biomedical technology becomes increasingly sophisticated, researchers can probe ever more subtle effects with the added requirement that the investigation of small effects often requires the acquisition of large amounts of data. In biomedicine, these data are often acquired at, and later shared between, multiple sites. There are both technological and sociological hurdles to be overcome for data to be passed between researchers and later made accessible to the larger scientific community. The goal of the Biomedical Informatics Research Network (BIRN) is to address the challenges inherent in biomedical data sharing. MATERIALS AND METHODS: BIRN tools are grouped into 'capabilities' and are available in the areas of data management, data security, information integration, and knowledge engineering. BIRN has a user-driven focus and employs a layered architectural approach that promotes reuse of infrastructure. BIRN tools are designed to be modular and therefore can work with pre-existing tools. BIRN users can choose the capabilities most useful for their application, while not having to ensure that their project conforms to a monolithic architecture. RESULTS: BIRN has implemented a new software-based data-sharing infrastructure that has been put to use in many different domains within biomedicine. BIRN is actively involved in outreach to the broader biomedical community to form working partnerships. CONCLUSION: BIRN's mission is to provide capabilities and services related to data sharing to the biomedical research community. It does this by forming partnerships and solving specific, user-driven problems whose solutions are then available for use by other groups. Karl G. Helmer, José Luis Ambite, Joseph Ames, Rachana Ananthakrishnan, Gully A. P. C. Burns, Ann L. Chervenak, Ian T. Foster, Liming Lee, David B. Keator, Fabio Macciardi, Ravi K. Madduri, John-Paul Navarro, Steven G. Potkin, Bruce R. Rosen, Seth Ruffins, Robert Schuler, Jessica A. Turner, Arthur W. Toga, Christina Williams, Carl Kesselman |
J. Am. Medical Informatics Assoc. | 11 |
| 2010 | caGrid Workflow Toolkit: A Taverna based workflow tool for cancer GridabstractBACKGROUND: In biological and medical domain, the use of web services made the data and computation functionality accessible in a unified manner, which helped automate the data pipeline that was previously performed manually. Workflow technology is widely used in the orchestration of multiple services to facilitate in-silico research. Cancer Biomedical Informatics Grid (caBIG) is an information network enabling the sharing of cancer research related resources and caGrid is its underlying service-based computation infrastructure. CaBIG requires that services are composed and orchestrated in a given sequence to realize data pipelines, which are often called scientific workflows. RESULTS: CaGrid selected Taverna as its workflow execution system of choice due to its integration with web service technology and support for a wide range of web services, plug-in architecture to cater for easy integration of third party extensions, etc. The caGrid Workflow Toolkit (or the toolkit for short), an extension to the Taverna workflow system, is designed and implemented to ease building and running caGrid workflows. It provides users with support for various phases in using workflows: service discovery, composition and orchestration, data access, and secure service invocation, which have been identified by the caGrid community as challenging in a multi-institutional and cross-discipline domain. CONCLUSIONS: By extending the Taverna Workbench, caGrid Workflow Toolkit provided a comprehensive solution to compose and coordinate services in caGrid, which would otherwise remain isolated and disconnected from each other. Using it users can access more than 140 services and are offered with a rich set of features including discovery of data and analytical services, query and transfer of data, security protections for service invocations, state management in service interactions, and sharing of workflows, experiences and best practices. The proposed solution is general enough to be applicable and reusable within other service-computing infrastructures that leverage similar technology stack. Wei Tan 0001, Ravi K. Madduri, Aleksandra Nenadic, Stian Soiland-Reyes, Dinanath Sulakhe, Ian T. Foster, Carole A. Goble |
BMC Bioinform. | 2 |
| 2010 | A comparison of using Taverna and BPEL in building scientific workflows: the case of caGridabstractWith the emergence of "service oriented science," the need arises to orchestrate multiple services to facilitate scientific investigation-that is, to create "science workflows." We present here our findings in providing a workflow solution for the caGrid service-based grid infrastructure. We choose BPEL and Taverna as candidates, and compare their usability in the lifecycle of a scientific workflow, including workflow composition, execution, and result analysis. Our experience shows that BPEL as an imperative language offers a comprehensive set of modeling primitives for workflows of all flavors; while Taverna offers a dataflow model and a more compact set of primitives that facilitates dataflow modeling and pipelined execution. We hope that this comparison study not only helps researchers select a language or tool that meets their specific needs, but also offers some insight on how a workflow language and tool can fulfill the requirement of the scientific community. Wei Tan 0001, Paolo Missier, Ian T. Foster, Ravi K. Madduri, David De Roure, Carole A. Goble |
Concurr. Comput. Pract. Exp. | 4 |
| 2009 | Wrap Scientific Applications as WSRF Grid Services Using gRAVIabstractWeb service models are increasingly being used in the Grid community as way to create distributed applications exposing data and/or applications through self describing interfaces. Scientific research is one key field in which the benefits are apparent as individual services can be orchestrated into experimental workflows that model the research process and facilitate verification and extension. However, many applications are not web enabled and the task of creating services from scratch is cumbersome in part due to the range of complex technologies, tools, standards and languages involved. In this paper we present gRAVI, a WSRF Web service wrapping tool that allows scientists to rapidly expose applications, scripts and workflows as Web services. gRAVI generated services include GSI security, Grid scheduling, state notifications, persistence and data staging. All service code, scripts and definition files are created automatically without any developer input. gRAVI services are created in standard Grid Archive files and are able to be moved and deployed to any compliant container with no requirement for any gRAVI or Grid infrastructure on the target machine. gRAVI supports deployment to the open science cloud Nimbus, whilst also being able to parse Taverna workflow definition files to create strongly typed services. Kyle Chard, Wei Tan 0001, Joshua Boverhof, Ravi K. Madduri, Ian T. Foster |
ICWS | 4 |
| 2009 | Scientific Workflows as Services in caGrid: A Taverna and gRAVI ApproachabstractIn scientific collaboration platforms such as caGrid, workflow-as-a-service is a useful concept for various reasons, such as easy reuse of workflows, access to remote resources, security concerns, and improved execution performance. We propose a solution for facilitating workflow-as-a-service based on Taverna as the workflow engine and gRAVI as a service wrapping tool. We provide both a generic service to execute all Taverna workflows, and an easy-to-use tool (gRAVI-t) for users to wrap their workflows as workflow-specific services, without developing service code. The signature of the specific service is identical to the corresponding workflow's input/output definition and is therefore more self-explained to workflow users. These two categories of services are useful in different scenarios, respectively. We use a tumor analysis workflow as an example to demonstrate how the workflow-as-a-service approach benefits the execution performance. Finally a conclusion is drawn and future research opportunities are discussed. Wei Tan 0001, Kyle Chard, Dinanath Sulakhe, Ravi K. Madduri, Ian T. Foster, Stian Soiland-Reyes, Carole A. Goble |
ICWS | 4 |
| 2008 | Build Grid Enabled Scientific Workflows Using gRAVI and TavernaabstractScientific communities are increasingly exposing information and tools as online services in an effort to abstract complex scientific processes and large data sets. Clients are then able to access services without knowledge of their internal workings therefore simplifying the process of replicating scientific research. Taking a service-oriented approach to science (SOS) facilitates reuse, extension, and scalability of components, whist also making information and tools available to a wider audience. Scientific workflows play a key role in realizing SOS by orchestrating services into well formed logical pipelines created to model the requirements of complex scientific experiments. The task of developing such service-oriented infrastructures is not trivial as developers must create and deploy Web services and then coordinate multiple services into workflows. This paper presents an end-to-end approach for developing SOS-based workflows with the aim of simplifying development, deployment, and execution. In particular, we use gRAVI to wrap applications as WSRF Web services and Taverna to compose and execute workflows. The process is validated through the creation of a real world bioinformatics workflow involving multiple services and complex execution paths. Kyle Chard, Cem Onyuksel, Wei Tan 0001, Dinanath Sulakhe, Ravi K. Madduri, Ian T. Foster |
eScience | 5 |
| 2008 | Orchestrating caGrid Services in TavernaabstractcaBIGtrade (the cancer Biomedical Informatics Gridtrade) is an open-source, open-access information network enabling cancer researchers to share tools, data, applications, and technologies. caGrid is the underlying service-based grid software infrastructure for caBIG, integrating distributed data and analytic resources into a virtual collaborative platform for cancer research. Within caGrid, many cancer-related data analysis and aggregation tasks can make use of "canned" sets of service invocations, or workflows. As a result, there is a need to orchestrate the invocation of caGrid services through the use of both a workflow language and tooling. In this paper, we first explain why we select Taverna as a candidate for workflow authoring and invocation. We then review the development of Taverna plug-ins in general, and describe how we extend Taverna to use caGrid services. We then detail a real-world example and the lessons learned from our research. Finally we conclude with a summary and a description of potential next steps. Wei Tan 0001, Ravi K. Madduri, Kiran Keshav, Baris E. Suzek, Scott Oster, Ian T. Foster |
ICWS | 2 |
| 2008 | Model Formulation: caGrid 1.0: An Enterprise Grid Infrastructure for Biomedical ResearchabstractOBJECTIVE: To develop software infrastructure that will provide support for discovery, characterization, integrated access, and management of diverse and disparate collections of information sources, analysis methods, and applications in biomedical research. DESIGN: An enterprise Grid software infrastructure, called caGrid version 1.0 (caGrid 1.0), has been developed as the core Grid architecture of the NCI-sponsored cancer Biomedical Informatics Grid (caBIG) program. It is designed to support a wide range of use cases in basic, translational, and clinical research, including 1) discovery, 2) integrated and large-scale data analysis, and 3) coordinated study. MEASUREMENTS: The caGrid is built as a Grid software infrastructure and leverages Grid computing technologies and the Web Services Resource Framework standards. It provides a set of core services, toolkits for the development and deployment of new community provided services, and application programming interfaces for building client applications. RESULTS: The caGrid 1.0 was released to the caBIG community in December 2006. It is built on open source components and caGrid source code is publicly and freely available under a liberal open source license. The core software, associated tools, and documentation can be downloaded from the following URL: https://cabig.nci.nih.gov/workspaces/Architecture/caGrid. CONCLUSIONS: While caGrid 1.0 is designed to address use cases in cancer research, the requirements associated with discovery, analysis and integration of large scale data, and coordinated studies are common in other biomedical fields. In this respect, caGrid 1.0 is the realization of a framework that can benefit the entire biomedical community. Scott Oster, Stephen Langella, Shannon Hastings, David Ervin, Ravi K. Madduri, Joshua Phillips, Tahsin M. Kurç, Frank Siebenlist, Peter A. Covitz, Krishnakant Shanbhag, Ian T. Foster, Joel H. Saltz |
J. Am. Medical Informatics Assoc. | 5 |
| 2007 | caGrid 1.0: A Grid Enterprise Architecture for Cancer Research
Scott Oster, Stephen Langella, Shannon Hastings, David Ervin, Ravi K. Madduri, Tahsin M. Kurç, Frank Siebenlist, Peter A. Covitz, Krishnakant Shanbhag, Ian T. Foster, Joel H. Saltz |
AMIA | 5 |
| 2006 | Streamlining Grid Operations: Definition and Deployment of a Portal-based User Registration Service
Ian T. Foster, Veronika Nefedova, Mehran Ahsant, Rachana Ananthakrishnan, Liming Lee, Ravi K. Madduri, Olle Mulmo, Laura Pearlman, Frank Siebenlist |
J. Grid Comput. | 6 |
| 2002 | Reliable File Transfer in Grid EnvironmentsabstractGrid-based computing environments are becoming increasingly popular for scientific computing. One of the key issues for scientific computing is the efficient transfer of large amounts of data across the Grid. In this paper we present a reliable file transfer (RFT) service that significantly improves the efficiency of large-scale file transfer. RFT can detect a variety of failures and restart the file transfer from the point of failure. It also has capabilities for improving transfer performance through TCP tuning. Ravi K. Madduri, William E. Allcock |
LCN | 1 |