Thilina Gunarathne

dblp:35/7524 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 55% Cloud and datacenter computing · 33% High-performance computing · 12%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › data-parallel programming
mapreduce
0.222011
Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011
Twister: a runtime for iterative MapReduce · HPDC 2010
Cloud and datacenter computing
job scheduling
0.112011
Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011
High-performance computing › high-throughput computing
many-task computing
0.112011
Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011
Parallel and multicore computing › data parallelism
data-parallel applications
0.112010
Twister: a runtime for iterative MapReduce · HPDC 2010
Cloud and datacenter computing › cluster computing framework
mapreduce framework
0.112010
Cloud computing paradigms for pleasingly parallel biomedical applications · HPDC 2010
Parallel and multicore computing › parallel programming runtimes
mapreduce runtime
0.112010
Twister: a runtime for iterative MapReduce · HPDC 2010
Parallel and multicore computing
parallel programming runtimes
0.112010
Twister: a runtime for iterative MapReduce · HPDC 2010
Cloud and datacenter computing
utility computing
0.112010
Cloud computing paradigms for pleasingly parallel biomedical applications · HPDC 2010
Bioinformatics and computational biology › sequence analysis › sequence assembly
EST assembly
0.012011
Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011
Bioinformatics and computational biology
sequence alignment
0.012011
Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011
Bioinformatics and computational biology › sequence analysis
sequence assembly
0.012011
Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly
0.012010
Cloud computing paradigms for pleasingly parallel biomedical applications · HPDC 2010

Methods — techniques the papers use, named apart from their topics

virtualization · 0.2hadoop · 0.2MPI · 0.2DryadLINQ · 0.2iterative mapreduce · 0.1
YearPublicationVenuePosition
2014 Towards a Collective Layer in the Big Data Stack
abstract
We generalize MapReduce, Iterative MapReduce and data intensive MPI runtime as a layered Map-Collective architecture with Map-All Gather, Map-All Reduce, MapReduce Merge Broadcast and Map-Reduce Scatter patterns as the initial focus. Map-collectives improve the performance and efficiency of the computations while at the same time facilitating ease of use for the users. These collective primitives can be applied to multiple runtimes and we propose building high performance robust implementations that cross cluster and cloud systems. Here we present results for two collectives shared between Hadoop (where we term our extension H-Collectives) on clusters and the Twister4Azure Iterative MapReduce for the Azure Cloud. Our prototype implementations of Map-All Gather and Map-All Reduce primitives achieved up to 33% performance improvement for K-means Clustering and up to 50% improvement for Multi-Dimensional Scaling, while also improving the user friendliness. In some cases, use of Map-collectives virtually eliminated almost all the overheads of the computations.
Thilina Gunarathne, Judy Qiu, Dennis Gannon
CCGRID1
2013 Scalable parallel computing on clouds using Twister4Azure iterative MapReduce
Thilina Gunarathne, Bingjing Zhang, Tak-Lon Wu, Judy Qiu
Future Gener. Comput. Syst.1
2011 Cloud computing paradigms for pleasingly parallel biomedical applications
abstract
SUMMARY Cloud computing offers exciting new approaches for scientific computing that leverage major commercial players’ hardware and software investments in large‐scale data centers. Loosely coupled problems are very important in many scientific fields, and with the ongoing move towards data‐intensive computing, they are on the rise. There exist several different approaches to leveraging clouds and cloud‐oriented data processing frameworks to perform pleasingly parallel (also called embarrassingly parallel) computations. In this paper, we present three pleasingly parallel biomedical applications: (i) assembly of genome fragments; (ii) sequence alignment and similarity search; and (iii) dimension reduction in the analysis of chemical structures, which are implemented utilizing a cloud infrastructure service‐based utility computing models of Amazon Web Services ( http://Amazon.com Inc., Seattle, WA, USA) and Microsoft Windows Azure (Microsoft Corp., Redmond, WA, USA) as well as utilizing MapReduce‐based data processing frameworks Apache Hadoop (Apache Software Foundation, Los Angeles, CA, USA) and Microsoft DryadLINQ. We review and compare each of these frameworks, performing a comparative study among them based on performance, cost, and usability. High latency, eventually consistent cloud infrastructure service‐based frameworks that rely on off‐the‐node cloud storage were able to exhibit performance efficiencies and scalability comparable to the MapReduce‐based frameworks with local disk‐based storage for the applications considered. In this paper, we also analyze variations in cost among the different platform choices (e.g., Elastic Compute Cloud instance types), highlighting the importance of selecting an appropriate platform based on the nature of the computation. Copyright © 2011 John Wiley & Sons, Ltd.
Thilina Gunarathne, Tak-Lon Wu, Jong Choi 0001, Seung-Hee Bae, Judy Qiu
Concurr. Comput. Pract. Exp.1
2011 Cloud Technologies for Bioinformatics Applications
abstract
Executing large number of independent jobs or jobs comprising of large number of tasks that perform minimal intertask communication is a common requirement in many domains. Various technologies ranging from classic job schedulers to the latest cloud technologies such as MapReduce can be used to execute these "many-tasks” in parallel. In this paper, we present our experience in applying two cloud technologies Apache Hadoop and Microsoft DryadLINQ to two bioinformatics applications with the above characteristics. The applications are a pairwise Alu sequence alignment application and an Expressed Sequence Tag (EST) sequence assembly program. First, we compare the performance of these cloud technologies using the above applications and also compare them with traditional MPI implementation in one application. Next, we analyze the effect of inhomogeneous data on the scheduling mechanisms of the cloud technologies. Finally, we present a comparison of performance of the cloud technologies under virtual and nonvirtual hardware platforms.
Jaliya Ekanayake, Thilina Gunarathne, Judy Qiu
IEEE Trans. Parallel Distributed Syst.2
2010 MapReduce in the Clouds for Science
abstract
The utility computing model introduced by cloud computing combined with the rich set of cloud infrastructure services offers a very viable alternative to traditional servers and computing clusters. MapReduce distributed data processing architecture has become the weapon of choice for data-intensive analyses in the clouds and in commodity clusters due to its excellent fault tolerance features, scalability and the ease of use. Currently, there are several options for using MapReduce in cloud environments, such as using MapReduce as a service, setting up one's own MapReduce cluster on cloud instances, or using specialized cloud MapReduce runtimes that take advantage of cloud infrastructure services. In this paper, we introduce Azure MapReduce, a novel MapReduce runtime built using the Microsoft Azure cloud infrastructure services. Azure MapReduce architecture successfully leverages the high latency, eventually consistent, yet highly scalable Azure infrastructure services to provide an efficient, on demand alternative to traditional MapReduce clusters. Further we evaluate the use and performance of MapReduce frameworks, including Azure MapReduce, in cloud environments for scientific applications using sequence assembly and sequence alignment as use cases.
Thilina Gunarathne, Tak-Lon Wu, Judy Qiu, Geoffrey C. Fox
CloudCom1
2010 Twister: a runtime for iterative MapReduce
abstract
MapReduce programming model has simplified the implementation of many data parallel applications. The simplicity of the programming model and the quality of services provided by many implementations of MapReduce attract a lot of enthusiasm among distributed computing communities. From the years of experience in applying MapReduce to various scientific applications we identified a set of extensions to the programming model and improvements to its architecture that will expand the applicability of MapReduce to more classes of applications. In this paper, we present the programming model and the architecture of Twister an enhanced MapReduce runtime that supports iterative MapReduce computations efficiently. We also show performance comparisons of Twister with other similar runtimes such as Hadoop and DryadLINQ for large scale data parallel applications.
Jaliya Ekanayake, Bingjing Zhang, Thilina Gunarathne, Seung-Hee Bae, Judy Qiu, Geoffrey C. Fox
HPDC4
2010 Cloud computing paradigms for pleasingly parallel biomedical applications
abstract
Cloud computing offers exciting new approaches for scientific computing that leverages the hardware and software investments on large scale data centers by major commercial players. Loosely coupled problems are very important in many scientific fields and are on the rise with the ongoing move towards data intensive computing. There exist several approaches to leverage clouds & cloud oriented data processing frameworks to perform pleasingly parallel computations. In this paper we present two pleasingly parallel biomedical applications, 1) assembly of genome fragments 2) dimension reduction in the analysis of chemical structures, implemented utilizing cloud infrastructure service based utility computing models of Amazon AWS and Microsoft Windows Azure as well as utilizing MapReduce based data processing frameworks, Apache Hadoop and Microsoft DryadLINQ. We review and compare each of the frameworks and perform a comparative study among them based on performance, efficiency, cost and the usability. Cloud service based utility computing model and the managed parallelism (MapReduce) exhibited comparable performance and efficiencies for the applications we considered. We analyze the variations in cost between the different platform choices (eg: EC2 instance types), highlighting the need to select the appropriate platform based on the nature of the computation.
Thilina Gunarathne, Tak-Lon Wu, Judy Qiu, Geoffrey C. Fox
HPDC1
2010 Hybrid cloud and cluster computing paradigms for life science applications
abstract
BACKGROUND: Clouds and MapReduce have shown themselves to be a broadly useful approach to scientific computing especially for parallel data intensive applications. However they have limited applicability to some areas such as data mining because MapReduce has poor performance on problems with an iterative structure present in the linear algebra that underlies much data analysis. Such problems can be run efficiently on clusters using MPI leading to a hybrid cloud and cluster environment. This motivates the design and implementation of an open source Iterative MapReduce system Twister. RESULTS: Comparisons of Amazon, Azure, and traditional Linux and Windows environments on common applications have shown encouraging performance and usability comparisons in several important non iterative cases. These are linked to MPI applications for final stages of the data analysis. Further we have released the open source Twister Iterative MapReduce and benchmarked it against basic MapReduce (Hadoop) and MPI in information retrieval and life sciences applications. CONCLUSIONS: The hybrid cloud (MapReduce) and cluster (MPI) approach offers an attractive production environment while Twister promises a uniform programming environment for many Life Sciences applications. METHODS: We used commercial clouds Amazon and Azure and the NSF resource FutureGrid to perform detailed comparisons and evaluations of different approaches to data intensive computing. Several applications were developed in MPI, MapReduce and Twister in these different environments.
Judy Qiu, Jaliya Ekanayake, Thilina Gunarathne, Jong Choi 0001, Seung-Hee Bae, Bingjing Zhang, Tak-Lon Wu, Yang Ruan 0001, Saliya Ekanayake, Adam Hughes, Geoffrey C. Fox
BMC Bioinform.3
2009 Biomedical Case Studies in Data Intensive Computing
Geoffrey C. Fox, Xiaohong Qiu, Scott Beason, Jong Choi 0001, Jaliya Ekanayake, Thilina Gunarathne, Mina Rho, Haixu Tang, Neil Devadasan, Gilbert C. Liu
CloudCom6
2009 DryadLINQ for Scientific Analyses
abstract
Applying high level parallel runtimes to data/compute intensive applications is becoming increasingly common. The simplicity of the MapReduce programming model and the availability of open source MapReduce runtimes such as Hadoop, are attracting more users to the MapReduce programming model. Microsoft has released DryadLINQ for academic use, allowing users to experience a new programming model and a runtime that is capable of performing large scale data/compute intensive analyses. In this paper, we present our experience in applying DryadLINQ for a series of scientific data analysis applications, identify their mapping to the DryadLINQ programming model, and compare their performances with Hadoop implementations of the same applications.
Jaliya Ekanayake, Thilina Gunarathne, Geoffrey C. Fox, Atilla Soner Balkir, Christophe Poulain, Nelson Araujo, Roger S. Barga
eScience2
2009 Application of Management Frameworks to Manage Workflow-Based Systems: A Case Study on a Large Scale E-science Project
abstract
Management architectures are well discussed in the literature, but their application in real life settings has not been as well covered. Automatic management of a system involves many more complexities than closing the control-loop by reacting to sensor data and executing corrective actions. In this paper, we discuss those complexities and propose solutions to those problems on top of Hasthi management framework, where Hasthi is a robust, scalable, and distributed management framework that enables users to manage a system by enforcing management logic authored by users themselves. Furthermore, we present in detail a real life case study, which uses Hasthi to manage a large, SOA based, e-science cyberinfrastructure.
Srinath Perera, Suresh Marru, Thilina Gunarathne, Dennis Gannon, Beth Plale
ICWS3