EDBT 2026 Demo / reviewers in the wild / expert
Thilina Gunarathne
dblp:35/7524
· DBLP profile ↗
11ranked-venue papers
5as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 55% Cloud and datacenter computing · 33% High-performance computing · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › data-parallel programming
mapreduce |
0.2 | 2 | 2011 | Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011 Twister: a runtime for iterative MapReduce · HPDC 2010 |
Cloud and datacenter computing
job scheduling |
0.1 | 1 | 2011 | Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011 |
High-performance computing › high-throughput computing
many-task computing |
0.1 | 1 | 2011 | Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011 |
Parallel and multicore computing › data parallelism
data-parallel applications |
0.1 | 1 | 2010 | Twister: a runtime for iterative MapReduce · HPDC 2010 |
Cloud and datacenter computing › cluster computing framework
mapreduce framework |
0.1 | 1 | 2010 | Cloud computing paradigms for pleasingly parallel biomedical applications · HPDC 2010 |
Parallel and multicore computing › parallel programming runtimes
mapreduce runtime |
0.1 | 1 | 2010 | Twister: a runtime for iterative MapReduce · HPDC 2010 |
Parallel and multicore computing
parallel programming runtimes |
0.1 | 1 | 2010 | Twister: a runtime for iterative MapReduce · HPDC 2010 |
Cloud and datacenter computing
utility computing |
0.1 | 1 | 2010 | Cloud computing paradigms for pleasingly parallel biomedical applications · HPDC 2010 |
Bioinformatics and computational biology › sequence analysis › sequence assembly
EST assembly |
0.0 | 1 | 2011 | Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011 |
Bioinformatics and computational biology
sequence alignment |
0.0 | 1 | 2011 | Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011 |
Bioinformatics and computational biology › sequence analysis
sequence assembly |
0.0 | 1 | 2011 | Cloud Technologies for Bioinformatics Applications · IEEE Trans. Parallel Distributed Syst. 2011 |
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly |
0.0 | 1 | 2010 | Cloud computing paradigms for pleasingly parallel biomedical applications · HPDC 2010 |
Methods — techniques the papers use, named apart from their topics
virtualization · 0.2hadoop · 0.2MPI · 0.2DryadLINQ · 0.2iterative mapreduce · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Towards a Collective Layer in the Big Data StackabstractWe generalize MapReduce, Iterative MapReduce and data intensive MPI runtime as a layered Map-Collective architecture with Map-All Gather, Map-All Reduce, MapReduce Merge Broadcast and Map-Reduce Scatter patterns as the initial focus. Map-collectives improve the performance and efficiency of the computations while at the same time facilitating ease of use for the users. These collective primitives can be applied to multiple runtimes and we propose building high performance robust implementations that cross cluster and cloud systems. Here we present results for two collectives shared between Hadoop (where we term our extension H-Collectives) on clusters and the Twister4Azure Iterative MapReduce for the Azure Cloud. Our prototype implementations of Map-All Gather and Map-All Reduce primitives achieved up to 33% performance improvement for K-means Clustering and up to 50% improvement for Multi-Dimensional Scaling, while also improving the user friendliness. In some cases, use of Map-collectives virtually eliminated almost all the overheads of the computations. Thilina Gunarathne, Judy Qiu, Dennis Gannon |
CCGRID | 1 |
| 2013 | Scalable parallel computing on clouds using Twister4Azure iterative MapReduce
Thilina Gunarathne, Bingjing Zhang, Tak-Lon Wu, Judy Qiu |
Future Gener. Comput. Syst. | 1 |
| 2011 | Cloud computing paradigms for pleasingly parallel biomedical applicationsabstractSUMMARY Cloud computing offers exciting new approaches for scientific computing that leverage major commercial players’ hardware and software investments in large‐scale data centers. Loosely coupled problems are very important in many scientific fields, and with the ongoing move towards data‐intensive computing, they are on the rise. There exist several different approaches to leveraging clouds and cloud‐oriented data processing frameworks to perform pleasingly parallel (also called embarrassingly parallel) computations. In this paper, we present three pleasingly parallel biomedical applications: (i) assembly of genome fragments; (ii) sequence alignment and similarity search; and (iii) dimension reduction in the analysis of chemical structures, which are implemented utilizing a cloud infrastructure service‐based utility computing models of Amazon Web Services ( http://Amazon.com Inc., Seattle, WA, USA) and Microsoft Windows Azure (Microsoft Corp., Redmond, WA, USA) as well as utilizing MapReduce‐based data processing frameworks Apache Hadoop (Apache Software Foundation, Los Angeles, CA, USA) and Microsoft DryadLINQ. We review and compare each of these frameworks, performing a comparative study among them based on performance, cost, and usability. High latency, eventually consistent cloud infrastructure service‐based frameworks that rely on off‐the‐node cloud storage were able to exhibit performance efficiencies and scalability comparable to the MapReduce‐based frameworks with local disk‐based storage for the applications considered. In this paper, we also analyze variations in cost among the different platform choices (e.g., Elastic Compute Cloud instance types), highlighting the importance of selecting an appropriate platform based on the nature of the computation. Copyright © 2011 John Wiley & Sons, Ltd. Thilina Gunarathne, Tak-Lon Wu, Jong Choi 0001, Seung-Hee Bae, Judy Qiu |
Concurr. Comput. Pract. Exp. | 1 |
| 2011 | Cloud Technologies for Bioinformatics ApplicationsabstractExecuting large number of independent jobs or jobs comprising of large number of tasks that perform minimal intertask communication is a common requirement in many domains. Various technologies ranging from classic job schedulers to the latest cloud technologies such as MapReduce can be used to execute these "many-tasks” in parallel. In this paper, we present our experience in applying two cloud technologies Apache Hadoop and Microsoft DryadLINQ to two bioinformatics applications with the above characteristics. The applications are a pairwise Alu sequence alignment application and an Expressed Sequence Tag (EST) sequence assembly program. First, we compare the performance of these cloud technologies using the above applications and also compare them with traditional MPI implementation in one application. Next, we analyze the effect of inhomogeneous data on the scheduling mechanisms of the cloud technologies. Finally, we present a comparison of performance of the cloud technologies under virtual and nonvirtual hardware platforms. Jaliya Ekanayake, Thilina Gunarathne, Judy Qiu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2010 | MapReduce in the Clouds for ScienceabstractThe utility computing model introduced by cloud computing combined with the rich set of cloud infrastructure services offers a very viable alternative to traditional servers and computing clusters. MapReduce distributed data processing architecture has become the weapon of choice for data-intensive analyses in the clouds and in commodity clusters due to its excellent fault tolerance features, scalability and the ease of use. Currently, there are several options for using MapReduce in cloud environments, such as using MapReduce as a service, setting up one's own MapReduce cluster on cloud instances, or using specialized cloud MapReduce runtimes that take advantage of cloud infrastructure services. In this paper, we introduce Azure MapReduce, a novel MapReduce runtime built using the Microsoft Azure cloud infrastructure services. Azure MapReduce architecture successfully leverages the high latency, eventually consistent, yet highly scalable Azure infrastructure services to provide an efficient, on demand alternative to traditional MapReduce clusters. Further we evaluate the use and performance of MapReduce frameworks, including Azure MapReduce, in cloud environments for scientific applications using sequence assembly and sequence alignment as use cases. Thilina Gunarathne, Tak-Lon Wu, Judy Qiu, Geoffrey C. Fox |
CloudCom | 1 |
| 2010 | Twister: a runtime for iterative MapReduceabstractMapReduce programming model has simplified the implementation of many data parallel applications. The simplicity of the programming model and the quality of services provided by many implementations of MapReduce attract a lot of enthusiasm among distributed computing communities. From the years of experience in applying MapReduce to various scientific applications we identified a set of extensions to the programming model and improvements to its architecture that will expand the applicability of MapReduce to more classes of applications. In this paper, we present the programming model and the architecture of Twister an enhanced MapReduce runtime that supports iterative MapReduce computations efficiently. We also show performance comparisons of Twister with other similar runtimes such as Hadoop and DryadLINQ for large scale data parallel applications. Jaliya Ekanayake, Bingjing Zhang, Thilina Gunarathne, Seung-Hee Bae, Judy Qiu, Geoffrey C. Fox |
HPDC | 4 |
| 2010 | Cloud computing paradigms for pleasingly parallel biomedical applicationsabstractCloud computing offers exciting new approaches for scientific computing that leverages the hardware and software investments on large scale data centers by major commercial players. Loosely coupled problems are very important in many scientific fields and are on the rise with the ongoing move towards data intensive computing. There exist several approaches to leverage clouds & cloud oriented data processing frameworks to perform pleasingly parallel computations. In this paper we present two pleasingly parallel biomedical applications, 1) assembly of genome fragments 2) dimension reduction in the analysis of chemical structures, implemented utilizing cloud infrastructure service based utility computing models of Amazon AWS and Microsoft Windows Azure as well as utilizing MapReduce based data processing frameworks, Apache Hadoop and Microsoft DryadLINQ. We review and compare each of the frameworks and perform a comparative study among them based on performance, efficiency, cost and the usability. Cloud service based utility computing model and the managed parallelism (MapReduce) exhibited comparable performance and efficiencies for the applications we considered. We analyze the variations in cost between the different platform choices (eg: EC2 instance types), highlighting the need to select the appropriate platform based on the nature of the computation. Thilina Gunarathne, Tak-Lon Wu, Judy Qiu, Geoffrey C. Fox |
HPDC | 1 |
| 2010 | Hybrid cloud and cluster computing paradigms for life science applicationsabstractBACKGROUND: Clouds and MapReduce have shown themselves to be a broadly useful approach to scientific computing especially for parallel data intensive applications. However they have limited applicability to some areas such as data mining because MapReduce has poor performance on problems with an iterative structure present in the linear algebra that underlies much data analysis. Such problems can be run efficiently on clusters using MPI leading to a hybrid cloud and cluster environment. This motivates the design and implementation of an open source Iterative MapReduce system Twister. RESULTS: Comparisons of Amazon, Azure, and traditional Linux and Windows environments on common applications have shown encouraging performance and usability comparisons in several important non iterative cases. These are linked to MPI applications for final stages of the data analysis. Further we have released the open source Twister Iterative MapReduce and benchmarked it against basic MapReduce (Hadoop) and MPI in information retrieval and life sciences applications. CONCLUSIONS: The hybrid cloud (MapReduce) and cluster (MPI) approach offers an attractive production environment while Twister promises a uniform programming environment for many Life Sciences applications. METHODS: We used commercial clouds Amazon and Azure and the NSF resource FutureGrid to perform detailed comparisons and evaluations of different approaches to data intensive computing. Several applications were developed in MPI, MapReduce and Twister in these different environments. Judy Qiu, Jaliya Ekanayake, Thilina Gunarathne, Jong Choi 0001, Seung-Hee Bae, Bingjing Zhang, Tak-Lon Wu, Yang Ruan 0001, Saliya Ekanayake, Adam Hughes, Geoffrey C. Fox |
BMC Bioinform. | 3 |
| 2009 | Biomedical Case Studies in Data Intensive Computing
Geoffrey C. Fox, Xiaohong Qiu, Scott Beason, Jong Choi 0001, Jaliya Ekanayake, Thilina Gunarathne, Mina Rho, Haixu Tang, Neil Devadasan, Gilbert C. Liu |
CloudCom | 6 |
| 2009 | DryadLINQ for Scientific AnalysesabstractApplying high level parallel runtimes to data/compute intensive applications is becoming increasingly common. The simplicity of the MapReduce programming model and the availability of open source MapReduce runtimes such as Hadoop, are attracting more users to the MapReduce programming model. Microsoft has released DryadLINQ for academic use, allowing users to experience a new programming model and a runtime that is capable of performing large scale data/compute intensive analyses. In this paper, we present our experience in applying DryadLINQ for a series of scientific data analysis applications, identify their mapping to the DryadLINQ programming model, and compare their performances with Hadoop implementations of the same applications. Jaliya Ekanayake, Thilina Gunarathne, Geoffrey C. Fox, Atilla Soner Balkir, Christophe Poulain, Nelson Araujo, Roger S. Barga |
eScience | 2 |
| 2009 | Application of Management Frameworks to Manage Workflow-Based Systems: A Case Study on a Large Scale E-science ProjectabstractManagement architectures are well discussed in the literature, but their application in real life settings has not been as well covered. Automatic management of a system involves many more complexities than closing the control-loop by reacting to sensor data and executing corrective actions. In this paper, we discuss those complexities and propose solutions to those problems on top of Hasthi management framework, where Hasthi is a robust, scalable, and distributed management framework that enables users to manage a system by enforcing management logic authored by users themselves. Furthermore, we present in detail a real life case study, which uses Hasthi to manage a large, SOA based, e-science cyberinfrastructure. Srinath Perera, Suresh Marru, Thilina Gunarathne, Dennis Gannon, Beth Plale |
ICWS | 3 |