EDBT 2026 Demo / reviewers in the wild / expert
Enis Afgan
dblp:28/6086
· DBLP profile ↗
18ranked-venue papers
10as first author
2since 2021 · last 2024
0000-0003-0922-6711ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 2 since 2021Systems, architecture and hardware · 7 · 6 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Cloud and datacenter computing · 97% High-performance computing · 3% | |
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 66% Computational science and engineering · 34% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing
container orchestration |
0.8 | 1 | 2024 | Galaxy Helm chart: a standardized method for deploying production Galaxy servers · Bioinform. 2024 |
Cloud and datacenter computing › resource management
cloud resource management |
0.5 | 1 | 2021 | GalaxyCloudRunner: enhancing scalable computing for Galaxy · Bioinform. 2021 |
Cloud and datacenter computing
cloud security |
0.4 | 1 | 2020 | Cloud bursting galaxy: federated identity and access management · Bioinform. 2020 |
Bioinformatics and computational biology
computational infrastructure |
0.4 | 2 | 2024 | Galaxy Helm chart: a standardized method for deploying production Galaxy servers · Bioinform. 2024 GalaxyCloudRunner: enhancing scalable computing for Galaxy · Bioinform. 2021 |
Computational science and engineering
workflow management |
0.4 | 2 | 2014 | BioBlend.objects: metacomputing with Galaxy · Bioinform. 2014 BioBlend: automating pipeline analyses within Galaxy and CloudMan · Bioinform. 2013 |
Bioinformatics and computational biology
biomedical data analysis |
0.1 | 1 | 2020 | Cloud bursting galaxy: federated identity and access management · Bioinform. 2020 |
High-performance computing › distributed computing infrastructure
metacomputing |
0.1 | 1 | 2014 | BioBlend.objects: metacomputing with Galaxy · Bioinform. 2014 |
Methods — techniques the papers use, named apart from their topics
OpenID Connect · 0.9OAuth2 · 0.9object-oriented programming · 0.4API wrapping · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Galaxy Helm chart: a standardized method for deploying production Galaxy serversabstractMOTIVATION: The Galaxy application is a popular open-source framework for data intensive sciences, counting thousands of monthly users across more than 100 public servers. To support a growing number of users and a greater variety of use cases, the complexity of a production-grade Galaxy installation has also grown, requiring more administration effort. There is a need for a rapid and reproducible Galaxy deployment method that can be maintained at high-availability with minimal maintenance. RESULTS: We describe the Galaxy Helm chart that codifies all elements of a production-grade Galaxy installation into a single package. Deployable on Kubernetes clusters, the chart encapsulates supporting software services and implements the best-practices model for running Galaxy. It is also the most rapid method available for deploying a scalable, production-grade Galaxy instance on one's own infrastructure. The chart is highly configurable, allowing systems administrators to swap dependent services if desired. Notable uses of the chart include on-demand, fully-automated deployments on AnVIL, providing training infrastructure for the Bioconductor project, and as the AWS-recommended solution for running Galaxy on the Amazon cloud. AVAILABILITY AND IMPLEMENTATION: The source code for Galaxy Helm is available at https://github.com/galaxyproject/galaxy-helm, the corresponding Helm package at https://github.com/CloudVE/helm-charts, and the required Galaxy container image https://github.com/galaxyproject/galaxy-docker-k8s. Nuwan Goonasekera, Alexandru Mahmoud, Keith Suderman, Enis Afgan |
Bioinform. | 4 |
| 2021 | GalaxyCloudRunner: enhancing scalable computing for GalaxyabstractSUMMARY: The existence of more than 100 public Galaxy servers with service quotas is indicative of the need for an increased availability of compute resources for Galaxy to use. The GalaxyCloudRunner enables a Galaxy server to easily expand its available compute capacity by sending user jobs to cloud resources. User jobs are routed to the acquired resources based on a set of configurable rules and the resources can be dynamically acquired from any of four popular cloud providers (AWS, Azure, GCP or OpenStack) in an automated fashion. AVAILABILITY AND IMPLEMENTATION: GalaxyCloudRunner is implemented in Python and leverages Docker containers. The source code is MIT licensed and available at https://github.com/cloudve/galaxycloudrunner. The documentation is available at http://gcr.cloudve.org/. Nuwan Goonasekera, Alexandru Mahmoud, John Chilton, Enis Afgan |
Bioinform. | 4 |
| 2020 | Cloud bursting galaxy: federated identity and access managementabstractMOTIVATION: Large biomedical datasets, such as those from genomics and imaging, are increasingly being stored on commercial and institutional cloud computing platforms. This is because cloud-scale computing resources, from robust backup to high-speed data transfer to scalable compute and storage, are needed to make these large datasets usable. However, one challenge for large-scale biomedical data on the cloud is providing secure access, especially when datasets are distributed across platforms. While there are open Web protocols for secure authentication and authorization, these protocols are not in wide use in bioinformatics and are difficult to use for even technologically sophisticated users. RESULTS: We have developed a generic and extensible approach for securely accessing biomedical datasets distributed across cloud computing platforms. Our approach combines OpenID Connect and OAuth2, best-practice Web protocols for authentication and authorization, together with Galaxy (https://galaxyproject.org), a web-based computational workbench used by thousands of scientists across the world. With our enhanced version of Galaxy, users can access and analyze data distributed across multiple cloud computing providers without any special knowledge of access/authorization protocols. Our approach does not require users to share permanent credentials (e.g. username, password, API key), instead relying on automatically generated temporary tokens that refresh as needed. Our approach is generalizable to most identity providers and cloud computing platforms. To the best of our knowledge, Galaxy is the only computational workbench where users can access biomedical datasets across multiple cloud computing platforms using best-practice Web security approaches and thereby minimize risks of unauthorized data access and credential use. AVAILABILITY AND IMPLEMENTATION: Freely available for academic and commercial use under the open-source Academic Free License (https://opensource.org/licenses/AFL-3.0) from the following Github repositories: https://github.com/galaxyproject/galaxy and https://github.com/galaxyproject/cloudauthz. Vahid Jalili, Enis Afgan, James Taylor 0001, Jeremy Goecks |
Bioinform. | 2 |
| 2019 | Jetstream - Early operations performance, adoption, and impactsabstractSummary Jetstream is a first of its kind system for the NSF — a distributed production cloud resource. We review the purpose for creating Jetstream, discuss Jetstream's key characteristics, describe our experiences from the first year of maintaining an OpenStack‐based cloud environment, and share some of the early scientific impacts achieved by Jetstream users. Jetstream offers a unique capability within the XSEDE‐supported US national cyberinfrastructure, delivering interactive virtual machines (VMs) via the Atmosphere interface. As a multi‐region deployment that operates as an integrated system, Jetstream is proving effective in supporting modes and disciplines of research traditionally underrepresented on larger XSEDE‐supported clusters and supercomputers. Already, Jetstream has been used to perform research and education in biology, biochemistry, atmospheric science, earth science, and computer science. David Y. Hancock, Craig A. Stewart, Matthew W. Vaughn, Jeremy Fischer, John Michael Lowe, George W. Turner, Tyson Lee Swetnam, Tyler K. Chafin, Enis Afgan, Marlon E. Pierce, Winona Snapp-Childs |
Concurr. Comput. Pract. Exp. | 9 |
| 2019 | CloudLaunch: Discover and deploy cloud applications
Enis Afgan, Andrew Lonie, James Taylor 0001, Nuwan Goonasekera |
Future Gener. Comput. Syst. | 1 |
| 2018 | Federated Galaxy: Biomedical Computing at the FrontierabstractBiomedical data exploration requires integrative analyses of large datasets using a diverse ecosystem of tools. For more than a decade, the Galaxy project (https://galaxyproject.org) has provided researchers with a web-based, user-friendly, scalable data analysis framework complemented by a rich ecosystem of tools (https://usegalaxy.org/toolshed) used to perform genomic, proteomic, metabolomic, and imaging experiments. Galaxy can be deployed on the cloud (https://launch.usegalaxy.org), institutional computing clusters, and personal computers, or readily used on a number of public servers (e.g., https://usegalaxy.org). In this paper, we present our plan and progress towards creating Galaxy-as-a-Service-a federation of distributed data and computing resources into a panoptic analysis platform. Users can leverage a pool of public and institutional resources, in addition to plugging-in their private resources, helping answer the challenge of resource divergence across various Galaxy instances and enabling seamless analysis of biomedical data. Enis Afgan, Vahid Jalili, Nuwan Goonasekera, James Taylor 0001, Jeremy Goecks |
IEEE CLOUD | 1 |
| 2015 | Enabling cloud bursting for life sciences within GalaxyabstractSummary Fueled by the radically increased capacity to generate data over the past decade, the field of biomedical research has been constrained by the ability to analyze data. Galaxy, a Web‐based, open‐source data integration and analysis platform for life science research, has been democratizing access to data analysis tools. However, the scale of data and the scope of tools required have proven to be a significant challenge for any monolithic deployment of the Galaxy application. We have found that a distributed and federated approach to utilizing compute and storage resources is necessary. This paper describes the ongoing efforts in creating a ubiquitous platform capable of simultaneously utilizing dedicated as well as on‐demand cloud resources. Specifically, the requirements, process, and an implementation of a cloud‐bursting system are detailed. Copyright © 2010 John Wiley & Sons, Ltd. Enis Afgan, Nathan Coraor, John Chilton, Dannon Baker, James Taylor 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | BioBlend.objects: metacomputing with GalaxyabstractSUMMARY: BioBlend.objects is a new component of the BioBlend package, adding an object-oriented interface for the Galaxy REST-based application programming interface. It improves support for metacomputing on Galaxy entities by providing higher-level functionality and allowing users to more easily create programs to explore, query and create Galaxy datasets and workflows. AVAILABILITY AND IMPLEMENTATION: BioBlend.objects is available online at https://github.com/afgane/bioblend. The new object-oriented API is implemented by the galaxy/objects subpackage. Simone Leo, Luca Pireddu, Gianmauro Cuccuru, Luca Lianas, Nicola Soranzo, Enis Afgan, Gianluigi Zanetti |
Bioinform. | 6 |
| 2014 | Community-driven development for computational biology at Sprints, Hackathons and CodefestsabstractBACKGROUND: Computational biology comprises a wide range of technologies and approaches. Multiple technologies can be combined to create more powerful workflows if the individuals contributing the data or providing tools for its interpretation can find mutual understanding and consensus. Much conversation and joint investigation are required in order to identify and implement the best approaches. Traditionally, scientific conferences feature talks presenting novel technologies or insights, followed up by informal discussions during coffee breaks. In multi-institution collaborations, in order to reach agreement on implementation details or to transfer deeper insights in a technology and practical skills, a representative of one group typically visits the other. However, this does not scale well when the number of technologies or research groups is large. Conferences have responded to this issue by introducing Birds-of-a-Feather (BoF) sessions, which offer an opportunity for individuals with common interests to intensify their interaction. However, parallel BoF sessions often make it hard for participants to join multiple BoFs and find common ground between the different technologies, and BoFs are generally too short to allow time for participants to program together. RESULTS: This report summarises our experience with computational biology Codefests, Hackathons and Sprints, which are interactive developer meetings. They are structured to reduce the limitations of traditional scientific meetings described above by strengthening the interaction among peers and letting the participants determine the schedule and topics. These meetings are commonly run as loosely scheduled "unconferences" (self-organized identification of participants and topics for meetings) over at least two days, with early introductory talks to welcome and organize contributors, followed by intensive collaborative coding sessions. We summarise some prominent achievements of those meetings and describe differences in how these are organised, how their audience is addressed, and their outreach to their respective communities. CONCLUSIONS: Hackathons, Codefests and Sprints share a stimulating atmosphere that encourages participants to jointly brainstorm and tackle problems of shared interest in a self-driven proactive environment, as well as providing an opportunity for new participants to get involved in collaborative projects. Steffen Möller, Enis Afgan, Michael Banck, Raoul Jean Pierre Bonnal, Tim Booth, John Chilton, Peter J. A. Cock, Markus Gumbel, Nomi L. Harris, Richard C. G. Holland, Matús Kalas, László Kaján, Eri Kibukawa, David R. Powell, Pjotr Prins, Jacqueline Quinn, Olivier Sallou, Francesco Strozzi, Torsten Seemann, Clare Sloggett, Stian Soiland-Reyes, William Spooner, Sascha Steinbiss, Andreas Tille, Anthony J. Travis, Roman Guimera, Toshiaki Katayama, Brad A. Chapman |
BMC Bioinform. | 2 |
| 2013 | BioBlend: automating pipeline analyses within Galaxy and CloudManabstractUNLABELLED: We present BioBlend, a unified API in a high-level language (python) that wraps the functionality of Galaxy and CloudMan APIs. BioBlend makes it easy for bioinformaticians to automate end-to-end large data analysis, from scratch, in a way that is highly accessible to collaborators, by allowing them to both provide the required infrastructure and automate complex analyses over large datasets within the familiar Galaxy environment. AVAILABILITY AND IMPLEMENTATION: http://bioblend.readthedocs.org/. Automated installation of BioBlend is available via PyPI (e.g. pip install bioblend). Alternatively, the source code is available from the GitHub repository (https://github.com/afgane/bioblend) under the MIT open source license. The library has been tested and is working on Linux, Macintosh and Windows-based systems. Clare Sloggett, Nuwan Goonasekera, Enis Afgan |
Bioinform. | 3 |
| 2012 | Bio-Linux as a tool for bioinformatics trainingabstractBecause of the ever-increasing application of next-generation sequencing (NGS) in research, and the expectation of faster experiment turn-around, it is becoming unfeasible and unscalable for analysis to be done exclusively by existing trained bioinformaticians. Instead, researchers and bench biologists are performing at least parts of most analyses. In order for this to be realized, two conditions must be satisfied: (1) well designed and accessible tools need to be made available, and (2) researchers and biologists need to be trained to use such tools in order to confidently handle high volumes of NGS data. Bio-Linux is a fully featured, powerful, configurable and easy to maintain bioinformatics workstation and helps on both counts by offering well over one hundred bioinformatics tools packaged into a single distribution, easily accessible and readily usable. Bio-Linux is also accessible in the form of virtual images or on the cloud, thus providing researchers with immediate access to scalable compute infrastructure required to run the analysis. Furthermore this paper discusses how bioinformatics training on Bio-Linux is helping to bridge the data production and analysis gap. Tim Booth, Mesude Bicak, Hyun Soon Gweon, Dawn Field, Enis Afgan |
BIBE | 5 |
| 2012 | CloudMan as a platform for tool, data, and analysis distributionabstractBACKGROUND: Cloud computing provides an infrastructure that facilitates large scale computational analysis in a scalable, democratized fashion, However, in this context it is difficult to ensure sharing of an analysis environment and associated data in a scalable and precisely reproducible way. RESULTS: CloudMan (usecloudman.org) enables individual researchers to easily deploy, customize, and share their entire cloud analysis environment, including data, tools, and configurations. CONCLUSIONS: With the enabled customization and sharing of instances, CloudMan can be used as a platform for collaboration. The presented solution improves accessibility of cloud resources, tools, and data to the level of an individual researcher and contributes toward reproducibility and transparency of research solutions. Enis Afgan, Brad A. Chapman, James Taylor 0001 |
BMC Bioinform. | 1 |
| 2012 | A reference model for deploying applications in virtualized environmentsabstractModern scientific research has been revolutionized by the availability of powerful and flexible computational infrastructure. Virtualization has made it possible to acquire computational resources on demand. Establishing and enabling use of these environments is essential, but their widespread adoption will only succeed if they are transparently usable. Requiring changes to applications being deployed or requiring users to change how they utilize those applications represent barriers to the infrastructure acceptance. The problem lies in the process of deploying applications so that they can take advantage of the elasticity of the environment and deliver it transparently to users. Here, we describe a reference model for deploying applications into virtualized environments. The model is rooted in the low-level components common to a range of virtualized environments and it describes how to compose those otherwise dispersed components into a coherent unit. Use of the model enables applications to be deployed into the new environment without any modifications, it imposes minimal overhead on management of the infrastructure required to run the application, and yields a set of higher-level services as a byproduct of the component organization and the underlying infrastructure. We provide a fully functional sample application deployment and implement a framework for managing the overall application deployment. Enis Afgan, Dannon Baker, Anton Nekrutenko, James Taylor 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | Scheduling and planning job execution of loosely coupled applications
Enis Afgan, Purushotham V. Bangalore, Tibor Skala |
J. Supercomput. | 1 |
| 2011 | Application Information Services for distributed computing environments
Enis Afgan, Purushotham V. Bangalore, Karolj Skala |
Future Gener. Comput. Syst. | 1 |
| 2010 | Galaxy CloudMan: delivering cloud compute clustersabstractBACKGROUND: Widespread adoption of high-throughput sequencing has greatly increased the scale and sophistication of computational infrastructure needed to perform genomic research. An alternative to building and maintaining local infrastructure is "cloud computing", which, in principle, offers on demand access to flexible computational infrastructure. However, cloud computing resources are not yet suitable for immediate "as is" use by experimental biologists. RESULTS: We present a cloud resource management system that makes it possible for individual researchers to compose and control an arbitrarily sized compute cluster on Amazon's EC2 cloud infrastructure without any informatics requirements. Within this system, an entire suite of biological tools packaged by the NERC Bio-Linux team (http://nebc.nerc.ac.uk/tools/bio-linux) is available for immediate consumption. The provided solution makes it possible, using only a web browser, to create a completely configured compute cluster ready to perform analysis in less than five minutes. Moreover, we provide an automated method for building custom deployments of cloud resources. This approach promotes reproducibility of results and, if desired, allows individuals and labs to add or customize an otherwise available cloud system to better meet their needs. CONCLUSIONS: The expected knowledge and associated effort with deploying a compute cluster in the Amazon EC2 cloud is not trivial. The solution presented in this paper eliminates these barriers, making it possible for researchers to deploy exactly the amount of computing power they need, combined with a wealth of existing analysis software, to handle the ongoing data deluge. Enis Afgan, Dannon Baker, Nathan Coraor, Brad A. Chapman, Anton Nekrutenko, James Taylor 0001 |
BMC Bioinform. | 1 |
| 2009 | GridAtlas - A grid application and resource configuration repository and discovery serviceabstractAlthough access to grid resources is realized through a standardized interface, independent grid resources are not only managed autonomously but are also accessed as independent entities. Such environment results in configuration differences among individual resources forcing users that access those resources to deal with the variability in resource configurations. This behaviour breaks the concept of interpreting the grid as a unified entity and forces the users to think of the grid in terms of individual resources. Concretely, this variability is expressed through the requirement for the users to explicitly state application installation properties on individual resources during each job submission. This is a tedious, error-prone and unnecessary process that acts as a barrier in the use of the grid. In this paper, a tool named GridAtlas is presented that keeps up with the details of individual resource and application configurations and makes such data easily accessible from a well-known location through Web-service API calls or a Web interface. This paper describes the architecture of the GridAtlas service along with use cases where GridAtlas has been successfully applied and illustrates the benefit of such a service in real grid environments. Enis Afgan, Purushotham V. Bangalore, Dustin Duncan |
CLUSTER | 1 |
| 2007 | Performance Characterization of BLAST for the GridabstractBLAST is a commonly used bioinformatics application for performing query searches and analysis of biological data. As the amount of search data increases so does job search times. As means of reducing job turnaround times, scientists are resorting to new technologies such as grid computing to obtain needed computational and storage resources. Inherent with advent of new technologies, are additional complexities that arise, forcing scientists to deal with them. Grid computing exemplifies dynamic and transient state of heterogeneous resources that become a major obstacle in realizing user-desired levels of service. Many users do not realize that techniques applied in more traditional cluster environments do not simply transition into grid environment. This paper analyzes resource and application dependencies for BLAST in terms of job parameters that result in performance tradeoffs. We present a set of examples showing performance variability and point out a set of guidelines, which lead to establishing job performance tradeoffs. Enis Afgan, Purushotham V. Bangalore |
BIBE | 1 |