Enis Afgan

dblp:28/6086 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
2since 2021 · last 2024
0000-0003-0922-6711ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 2 since 2021Systems, architecture and hardware · 7 · 6 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 97% High-performance computing · 3%
Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 66% Computational science and engineering · 34%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
container orchestration
0.812024
Galaxy Helm chart: a standardized method for deploying production Galaxy servers · Bioinform. 2024
Cloud and datacenter computing › resource management
cloud resource management
0.512021
GalaxyCloudRunner: enhancing scalable computing for Galaxy · Bioinform. 2021
Cloud and datacenter computing
cloud security
0.412020
Cloud bursting galaxy: federated identity and access management · Bioinform. 2020
Bioinformatics and computational biology
computational infrastructure
0.422024
Galaxy Helm chart: a standardized method for deploying production Galaxy servers · Bioinform. 2024
GalaxyCloudRunner: enhancing scalable computing for Galaxy · Bioinform. 2021
Computational science and engineering
workflow management
0.422014
BioBlend.objects: metacomputing with Galaxy · Bioinform. 2014
BioBlend: automating pipeline analyses within Galaxy and CloudMan · Bioinform. 2013
Bioinformatics and computational biology
biomedical data analysis
0.112020
Cloud bursting galaxy: federated identity and access management · Bioinform. 2020
High-performance computing › distributed computing infrastructure
metacomputing
0.112014
BioBlend.objects: metacomputing with Galaxy · Bioinform. 2014

Methods — techniques the papers use, named apart from their topics

OpenID Connect · 0.9OAuth2 · 0.9object-oriented programming · 0.4API wrapping · 0.2
YearPublicationVenuePosition
2024 Galaxy Helm chart: a standardized method for deploying production Galaxy servers
abstract
MOTIVATION: The Galaxy application is a popular open-source framework for data intensive sciences, counting thousands of monthly users across more than 100 public servers. To support a growing number of users and a greater variety of use cases, the complexity of a production-grade Galaxy installation has also grown, requiring more administration effort. There is a need for a rapid and reproducible Galaxy deployment method that can be maintained at high-availability with minimal maintenance. RESULTS: We describe the Galaxy Helm chart that codifies all elements of a production-grade Galaxy installation into a single package. Deployable on Kubernetes clusters, the chart encapsulates supporting software services and implements the best-practices model for running Galaxy. It is also the most rapid method available for deploying a scalable, production-grade Galaxy instance on one's own infrastructure. The chart is highly configurable, allowing systems administrators to swap dependent services if desired. Notable uses of the chart include on-demand, fully-automated deployments on AnVIL, providing training infrastructure for the Bioconductor project, and as the AWS-recommended solution for running Galaxy on the Amazon cloud. AVAILABILITY AND IMPLEMENTATION: The source code for Galaxy Helm is available at https://github.com/galaxyproject/galaxy-helm, the corresponding Helm package at https://github.com/CloudVE/helm-charts, and the required Galaxy container image https://github.com/galaxyproject/galaxy-docker-k8s.
Nuwan Goonasekera, Alexandru Mahmoud, Keith Suderman, Enis Afgan
Bioinform.4
2021 GalaxyCloudRunner: enhancing scalable computing for Galaxy
abstract
SUMMARY: The existence of more than 100 public Galaxy servers with service quotas is indicative of the need for an increased availability of compute resources for Galaxy to use. The GalaxyCloudRunner enables a Galaxy server to easily expand its available compute capacity by sending user jobs to cloud resources. User jobs are routed to the acquired resources based on a set of configurable rules and the resources can be dynamically acquired from any of four popular cloud providers (AWS, Azure, GCP or OpenStack) in an automated fashion. AVAILABILITY AND IMPLEMENTATION: GalaxyCloudRunner is implemented in Python and leverages Docker containers. The source code is MIT licensed and available at https://github.com/cloudve/galaxycloudrunner. The documentation is available at http://gcr.cloudve.org/.
Nuwan Goonasekera, Alexandru Mahmoud, John Chilton, Enis Afgan
Bioinform.4
2020 Cloud bursting galaxy: federated identity and access management
abstract
MOTIVATION: Large biomedical datasets, such as those from genomics and imaging, are increasingly being stored on commercial and institutional cloud computing platforms. This is because cloud-scale computing resources, from robust backup to high-speed data transfer to scalable compute and storage, are needed to make these large datasets usable. However, one challenge for large-scale biomedical data on the cloud is providing secure access, especially when datasets are distributed across platforms. While there are open Web protocols for secure authentication and authorization, these protocols are not in wide use in bioinformatics and are difficult to use for even technologically sophisticated users. RESULTS: We have developed a generic and extensible approach for securely accessing biomedical datasets distributed across cloud computing platforms. Our approach combines OpenID Connect and OAuth2, best-practice Web protocols for authentication and authorization, together with Galaxy (https://galaxyproject.org), a web-based computational workbench used by thousands of scientists across the world. With our enhanced version of Galaxy, users can access and analyze data distributed across multiple cloud computing providers without any special knowledge of access/authorization protocols. Our approach does not require users to share permanent credentials (e.g. username, password, API key), instead relying on automatically generated temporary tokens that refresh as needed. Our approach is generalizable to most identity providers and cloud computing platforms. To the best of our knowledge, Galaxy is the only computational workbench where users can access biomedical datasets across multiple cloud computing platforms using best-practice Web security approaches and thereby minimize risks of unauthorized data access and credential use. AVAILABILITY AND IMPLEMENTATION: Freely available for academic and commercial use under the open-source Academic Free License (https://opensource.org/licenses/AFL-3.0) from the following Github repositories: https://github.com/galaxyproject/galaxy and https://github.com/galaxyproject/cloudauthz.
Vahid Jalili, Enis Afgan, James Taylor 0001, Jeremy Goecks
Bioinform.2
2019 Jetstream - Early operations performance, adoption, and impacts
abstract
Summary Jetstream is a first of its kind system for the NSF — a distributed production cloud resource. We review the purpose for creating Jetstream, discuss Jetstream's key characteristics, describe our experiences from the first year of maintaining an OpenStack‐based cloud environment, and share some of the early scientific impacts achieved by Jetstream users. Jetstream offers a unique capability within the XSEDE‐supported US national cyberinfrastructure, delivering interactive virtual machines (VMs) via the Atmosphere interface. As a multi‐region deployment that operates as an integrated system, Jetstream is proving effective in supporting modes and disciplines of research traditionally underrepresented on larger XSEDE‐supported clusters and supercomputers. Already, Jetstream has been used to perform research and education in biology, biochemistry, atmospheric science, earth science, and computer science.
David Y. Hancock, Craig A. Stewart, Matthew W. Vaughn, Jeremy Fischer, John Michael Lowe, George W. Turner, Tyson Lee Swetnam, Tyler K. Chafin, Enis Afgan, Marlon E. Pierce, Winona Snapp-Childs
Concurr. Comput. Pract. Exp.9
2019 CloudLaunch: Discover and deploy cloud applications
Enis Afgan, Andrew Lonie, James Taylor 0001, Nuwan Goonasekera
Future Gener. Comput. Syst.1
2018 Federated Galaxy: Biomedical Computing at the Frontier
abstract
Biomedical data exploration requires integrative analyses of large datasets using a diverse ecosystem of tools. For more than a decade, the Galaxy project (https://galaxyproject.org) has provided researchers with a web-based, user-friendly, scalable data analysis framework complemented by a rich ecosystem of tools (https://usegalaxy.org/toolshed) used to perform genomic, proteomic, metabolomic, and imaging experiments. Galaxy can be deployed on the cloud (https://launch.usegalaxy.org), institutional computing clusters, and personal computers, or readily used on a number of public servers (e.g., https://usegalaxy.org). In this paper, we present our plan and progress towards creating Galaxy-as-a-Service-a federation of distributed data and computing resources into a panoptic analysis platform. Users can leverage a pool of public and institutional resources, in addition to plugging-in their private resources, helping answer the challenge of resource divergence across various Galaxy instances and enabling seamless analysis of biomedical data.
Enis Afgan, Vahid Jalili, Nuwan Goonasekera, James Taylor 0001, Jeremy Goecks
IEEE CLOUD1
2015 Enabling cloud bursting for life sciences within Galaxy
abstract
Summary Fueled by the radically increased capacity to generate data over the past decade, the field of biomedical research has been constrained by the ability to analyze data. Galaxy, a Web‐based, open‐source data integration and analysis platform for life science research, has been democratizing access to data analysis tools. However, the scale of data and the scope of tools required have proven to be a significant challenge for any monolithic deployment of the Galaxy application. We have found that a distributed and federated approach to utilizing compute and storage resources is necessary. This paper describes the ongoing efforts in creating a ubiquitous platform capable of simultaneously utilizing dedicated as well as on‐demand cloud resources. Specifically, the requirements, process, and an implementation of a cloud‐bursting system are detailed. Copyright © 2010 John Wiley & Sons, Ltd.
Enis Afgan, Nathan Coraor, John Chilton, Dannon Baker, James Taylor 0001
Concurr. Comput. Pract. Exp.1
2014 BioBlend.objects: metacomputing with Galaxy
abstract
SUMMARY: BioBlend.objects is a new component of the BioBlend package, adding an object-oriented interface for the Galaxy REST-based application programming interface. It improves support for metacomputing on Galaxy entities by providing higher-level functionality and allowing users to more easily create programs to explore, query and create Galaxy datasets and workflows. AVAILABILITY AND IMPLEMENTATION: BioBlend.objects is available online at https://github.com/afgane/bioblend. The new object-oriented API is implemented by the galaxy/objects subpackage.
Simone Leo, Luca Pireddu, Gianmauro Cuccuru, Luca Lianas, Nicola Soranzo, Enis Afgan, Gianluigi Zanetti
Bioinform.6
2014 Community-driven development for computational biology at Sprints, Hackathons and Codefests
abstract
BACKGROUND: Computational biology comprises a wide range of technologies and approaches. Multiple technologies can be combined to create more powerful workflows if the individuals contributing the data or providing tools for its interpretation can find mutual understanding and consensus. Much conversation and joint investigation are required in order to identify and implement the best approaches. Traditionally, scientific conferences feature talks presenting novel technologies or insights, followed up by informal discussions during coffee breaks. In multi-institution collaborations, in order to reach agreement on implementation details or to transfer deeper insights in a technology and practical skills, a representative of one group typically visits the other. However, this does not scale well when the number of technologies or research groups is large. Conferences have responded to this issue by introducing Birds-of-a-Feather (BoF) sessions, which offer an opportunity for individuals with common interests to intensify their interaction. However, parallel BoF sessions often make it hard for participants to join multiple BoFs and find common ground between the different technologies, and BoFs are generally too short to allow time for participants to program together. RESULTS: This report summarises our experience with computational biology Codefests, Hackathons and Sprints, which are interactive developer meetings. They are structured to reduce the limitations of traditional scientific meetings described above by strengthening the interaction among peers and letting the participants determine the schedule and topics. These meetings are commonly run as loosely scheduled "unconferences" (self-organized identification of participants and topics for meetings) over at least two days, with early introductory talks to welcome and organize contributors, followed by intensive collaborative coding sessions. We summarise some prominent achievements of those meetings and describe differences in how these are organised, how their audience is addressed, and their outreach to their respective communities. CONCLUSIONS: Hackathons, Codefests and Sprints share a stimulating atmosphere that encourages participants to jointly brainstorm and tackle problems of shared interest in a self-driven proactive environment, as well as providing an opportunity for new participants to get involved in collaborative projects.
Steffen Möller, Enis Afgan, Michael Banck, Raoul Jean Pierre Bonnal, Tim Booth, John Chilton, Peter J. A. Cock, Markus Gumbel, Nomi L. Harris, Richard C. G. Holland, Matús Kalas, László Kaján, Eri Kibukawa, David R. Powell, Pjotr Prins, Jacqueline Quinn, Olivier Sallou, Francesco Strozzi, Torsten Seemann, Clare Sloggett, Stian Soiland-Reyes, William Spooner, Sascha Steinbiss, Andreas Tille, Anthony J. Travis, Roman Guimera, Toshiaki Katayama, Brad A. Chapman
BMC Bioinform.2
2013 BioBlend: automating pipeline analyses within Galaxy and CloudMan
abstract
UNLABELLED: We present BioBlend, a unified API in a high-level language (python) that wraps the functionality of Galaxy and CloudMan APIs. BioBlend makes it easy for bioinformaticians to automate end-to-end large data analysis, from scratch, in a way that is highly accessible to collaborators, by allowing them to both provide the required infrastructure and automate complex analyses over large datasets within the familiar Galaxy environment. AVAILABILITY AND IMPLEMENTATION: http://bioblend.readthedocs.org/. Automated installation of BioBlend is available via PyPI (e.g. pip install bioblend). Alternatively, the source code is available from the GitHub repository (https://github.com/afgane/bioblend) under the MIT open source license. The library has been tested and is working on Linux, Macintosh and Windows-based systems.
Clare Sloggett, Nuwan Goonasekera, Enis Afgan
Bioinform.3
2012 Bio-Linux as a tool for bioinformatics training
abstract
Because of the ever-increasing application of next-generation sequencing (NGS) in research, and the expectation of faster experiment turn-around, it is becoming unfeasible and unscalable for analysis to be done exclusively by existing trained bioinformaticians. Instead, researchers and bench biologists are performing at least parts of most analyses. In order for this to be realized, two conditions must be satisfied: (1) well designed and accessible tools need to be made available, and (2) researchers and biologists need to be trained to use such tools in order to confidently handle high volumes of NGS data. Bio-Linux is a fully featured, powerful, configurable and easy to maintain bioinformatics workstation and helps on both counts by offering well over one hundred bioinformatics tools packaged into a single distribution, easily accessible and readily usable. Bio-Linux is also accessible in the form of virtual images or on the cloud, thus providing researchers with immediate access to scalable compute infrastructure required to run the analysis. Furthermore this paper discusses how bioinformatics training on Bio-Linux is helping to bridge the data production and analysis gap.
Tim Booth, Mesude Bicak, Hyun Soon Gweon, Dawn Field, Enis Afgan
BIBE5
2012 CloudMan as a platform for tool, data, and analysis distribution
abstract
BACKGROUND: Cloud computing provides an infrastructure that facilitates large scale computational analysis in a scalable, democratized fashion, However, in this context it is difficult to ensure sharing of an analysis environment and associated data in a scalable and precisely reproducible way. RESULTS: CloudMan (usecloudman.org) enables individual researchers to easily deploy, customize, and share their entire cloud analysis environment, including data, tools, and configurations. CONCLUSIONS: With the enabled customization and sharing of instances, CloudMan can be used as a platform for collaboration. The presented solution improves accessibility of cloud resources, tools, and data to the level of an individual researcher and contributes toward reproducibility and transparency of research solutions.
Enis Afgan, Brad A. Chapman, James Taylor 0001
BMC Bioinform.1
2012 A reference model for deploying applications in virtualized environments
abstract
Modern scientific research has been revolutionized by the availability of powerful and flexible computational infrastructure. Virtualization has made it possible to acquire computational resources on demand. Establishing and enabling use of these environments is essential, but their widespread adoption will only succeed if they are transparently usable. Requiring changes to applications being deployed or requiring users to change how they utilize those applications represent barriers to the infrastructure acceptance. The problem lies in the process of deploying applications so that they can take advantage of the elasticity of the environment and deliver it transparently to users. Here, we describe a reference model for deploying applications into virtualized environments. The model is rooted in the low-level components common to a range of virtualized environments and it describes how to compose those otherwise dispersed components into a coherent unit. Use of the model enables applications to be deployed into the new environment without any modifications, it imposes minimal overhead on management of the infrastructure required to run the application, and yields a set of higher-level services as a byproduct of the component organization and the underlying infrastructure. We provide a fully functional sample application deployment and implement a framework for managing the overall application deployment.
Enis Afgan, Dannon Baker, Anton Nekrutenko, James Taylor 0001
Concurr. Comput. Pract. Exp.1
2012 Scheduling and planning job execution of loosely coupled applications
Enis Afgan, Purushotham V. Bangalore, Tibor Skala
J. Supercomput.1
2011 Application Information Services for distributed computing environments
Enis Afgan, Purushotham V. Bangalore, Karolj Skala
Future Gener. Comput. Syst.1
2010 Galaxy CloudMan: delivering cloud compute clusters
abstract
BACKGROUND: Widespread adoption of high-throughput sequencing has greatly increased the scale and sophistication of computational infrastructure needed to perform genomic research. An alternative to building and maintaining local infrastructure is "cloud computing", which, in principle, offers on demand access to flexible computational infrastructure. However, cloud computing resources are not yet suitable for immediate "as is" use by experimental biologists. RESULTS: We present a cloud resource management system that makes it possible for individual researchers to compose and control an arbitrarily sized compute cluster on Amazon's EC2 cloud infrastructure without any informatics requirements. Within this system, an entire suite of biological tools packaged by the NERC Bio-Linux team (http://nebc.nerc.ac.uk/tools/bio-linux) is available for immediate consumption. The provided solution makes it possible, using only a web browser, to create a completely configured compute cluster ready to perform analysis in less than five minutes. Moreover, we provide an automated method for building custom deployments of cloud resources. This approach promotes reproducibility of results and, if desired, allows individuals and labs to add or customize an otherwise available cloud system to better meet their needs. CONCLUSIONS: The expected knowledge and associated effort with deploying a compute cluster in the Amazon EC2 cloud is not trivial. The solution presented in this paper eliminates these barriers, making it possible for researchers to deploy exactly the amount of computing power they need, combined with a wealth of existing analysis software, to handle the ongoing data deluge.
Enis Afgan, Dannon Baker, Nathan Coraor, Brad A. Chapman, Anton Nekrutenko, James Taylor 0001
BMC Bioinform.1
2009 GridAtlas - A grid application and resource configuration repository and discovery service
abstract
Although access to grid resources is realized through a standardized interface, independent grid resources are not only managed autonomously but are also accessed as independent entities. Such environment results in configuration differences among individual resources forcing users that access those resources to deal with the variability in resource configurations. This behaviour breaks the concept of interpreting the grid as a unified entity and forces the users to think of the grid in terms of individual resources. Concretely, this variability is expressed through the requirement for the users to explicitly state application installation properties on individual resources during each job submission. This is a tedious, error-prone and unnecessary process that acts as a barrier in the use of the grid. In this paper, a tool named GridAtlas is presented that keeps up with the details of individual resource and application configurations and makes such data easily accessible from a well-known location through Web-service API calls or a Web interface. This paper describes the architecture of the GridAtlas service along with use cases where GridAtlas has been successfully applied and illustrates the benefit of such a service in real grid environments.
Enis Afgan, Purushotham V. Bangalore, Dustin Duncan
CLUSTER1
2007 Performance Characterization of BLAST for the Grid
abstract
BLAST is a commonly used bioinformatics application for performing query searches and analysis of biological data. As the amount of search data increases so does job search times. As means of reducing job turnaround times, scientists are resorting to new technologies such as grid computing to obtain needed computational and storage resources. Inherent with advent of new technologies, are additional complexities that arise, forcing scientists to deal with them. Grid computing exemplifies dynamic and transient state of heterogeneous resources that become a major obstacle in realizing user-desired levels of service. Many users do not realize that techniques applied in more traditional cluster environments do not simply transition into grid environment. This paper analyzes resource and application dependencies for BLAST in terms of job parameters that result in performance tradeoffs. We present a set of examples showing performance variability and point out a set of guidelines, which lead to establishing job performance tradeoffs.
Enis Afgan, Purushotham V. Bangalore
BIBE1