EDBT 2026 Demo / reviewers in the wild / expert
Madhusudhan Govindaraju
dblp:02/7038
· DBLP profile ↗
55ranked-venue papers
5as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 since 2021Software engineering, systems software and programming languages · 10 · 1 first-authorArtificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cost-Justified Multi-type Resource Fair Scheduling for Kubernetes
Madhusudhan Govindaraju, Sujoy Sikdar |
CCGrid | 2 |
| 2022 | A Study of Contributing Factors to Power Aware Vertical Scaling of Deadline Constrained ApplicationsabstractThe adoption of virtualization technologies in datacenters has increased dramatically in the past decade. Clouds have pivoted from being just an infrastructure rental to offering platforms and solutions, made possible by having several layers of abstraction, providing internal and external users the ability to focus on core business logic. Efficient resource management has in turn become salient in ensuring operational efficiency. In this work, we study key factors that can influence vertical scaling decisions, propose a policy to vertically scale deadline constrained applications and surface our findings from experimentation. We observe that (a) the duration for which an application is profiled has an almost cyclic influence on the accuracy of behavior predictions and is inversely proportional to the time spent consuming backlog, (b) the duration for which an application is scaled can help achieve up to a 9.6% and 4.2% reduction in the 75thand 95thpercentile of power usage respectively, (c) reducing the tolerance towards accrual of backlog influences the application execution time and can reduce the number of SLA violations by 50% or 100% at times and (d) increasing the time to deadline offers power saving opportunities and can help achieve a 9.3% improvement in the 75thpercentile of power usage. Pradyumna Kaushik, Srinidhi Raghavendra, Madhusudhan Govindaraju |
CLOUD | 3 |
| 2022 | QSketch: GPU-Aware Probabilistic Sketch Data StructuresabstractA fundamental problem in data analysis is determining the frequency distribution of items within a data set. A count-min sketch is a widely used data structure that can estimate such a distribution. General purpose computation with graphics processing units (GPUs) has become increasingly prevalent over the last decade or so. A GPU provides much higher instruction throughput and memory bandwidth than a typical CPU. However, achieving maximum throughput on GPUs for count-min sketch is challenging, especially for large, uniformly distributed data sets, because each operation needs several random memory accesses that are slow and inefficient on GPUs. In this paper, we propose a suite of novel count-min sketch structures called QSketch. QSketch structures need only one coalesced memory access on average. We also present methods and algorithms to improve the accuracy without decreasing the searching performance. We implemented the QSketch with CUDA 11 and evaluated its performance and accuracy on various GPU platforms. Huanyi Qin, Zongpai Zhang, Madhusudhan Govindaraju, Kenneth Chiu |
CCGRID | 4 |
| 2021 | Low Latency and High Throughput Write-Ahead Logging Using CAPI-FlashabstractHigh-velocity data imposes high durability overheads on Big Data technology components such as NoSQL data stores. In Apache Cassandra and MongoDB, widely used NoSQL solutions with high scalability and availability, write-ahead logging is used to provide durability. However, current write-ahead logging techniques are limited by the excessive overhead in the I/O subsystem. To address this performance gap, we have designed a novel CAPI-Flash based high performance durable logging mechanism for Apache Cassandra and MongoDB. We take advantage of the high throughput, low latency path to flash storage provided by the Coherent Accelerator Processor Interface (CAPI) on IBM POWER8 Systems. Our experimental results show that for insert-only workloads, CAPI-Flash logging provides up to 70 and 514 percent improvement in throughput compared to Cassandra and MongoDB’s durable alternatives, respectively. It also provides average of 45 percent increase in throughput with Cassandra and average of 115 percent increase in throughput with MongoDB for update-mostly and update-only workloads. Bedri Sendir, Madhusudhan Govindaraju, Rei Odaira, H. Peter Hofstee |
IEEE Trans. Cloud Comput. | 2 |
| 2019 | Exploring Potential for Non-Disruptive Vertical Auto Scaling and Resource Estimation in KubernetesabstractCloud platforms typically require users to provide resource requirements for applications so that resource managers can schedule containers with adequate allocations. However, the requirements for container resources often depend on numerous factors such as application input parameters, optimization flags, input files, and attributes that are specified for each run. So, it is complex for users to estimate the resource requirements for a given container accurately, leading to resource over-estimation that negatively affects overall utilization. We have designed a Resource Utilization Based Autoscaling System (RUBAS) that can dynamically adjust the allocation of containers running in a Kubernetes cluster. RUBAS improves upon the Kubernetes Vertical Pod Autoscaler (VPA) system non-disruptively by incorporating container migration. Our experiments use multiple scientific benchmarks. We analyze the allocation pattern of RUBAS with Kubernetes VPA. We compare the performance of container migration for in-place and remote node migration and we evaluate the overhead in RUBAS. Our results show that compared to Kubernetes VPA, RUBAS improves the CPU and memory utilization of the cluster by 10% and reduces the runtime by 15% with an overhead for each application ranging from 5% to 20%. Gourav Rattihalli, Madhusudhan Govindaraju, Devesh Tiwari |
CLOUD | 2 |
| 2019 | Towards Enabling Dynamic Resource Estimation and Correction for Improving Utilization in an Apache Mesos Cloud EnvironmentabstractAcademic cloud infrastructures require users to specify an estimate of their resource requirements. The resource usage for applications often depends on the input file sizes, parameters, optimization flags, and attributes, specified for each run. Incorrect estimation can result in low resource utilization of the entire infrastructure and long wait times for jobs in the queue. We have designed a Resource Utilization based Migration (RUMIG) system to address the resource estimation problem. We present the overall architecture of the two-stage elastic cluster design, the Apache Mesos-specific container migration system, and analyze the performance for several scientific workloads on three different cloud/cluster environments. In this paper we (b) present a design and implementation for container migration in a Mesos environment, (c) evaluate the effect of right-sizing and cluster elasticity on overall performance, (d) analyze different profiling intervals to determine the best fit, (e) determine the overhead of our profiling mechanism. Compared to the default use of Apache Mesos, in the best cases, RUMIG provides a gain of 65% in runtime (local cluster), 51% in CPU utilization in the Chameleon cloud, and 27% in memory utilization in the Jetstream cloud. Gourav Rattihalli, Madhusudhan Govindaraju, Devesh Tiwari |
CCGRID | 2 |
| 2018 | Analysis of Dynamically Switching Energy-Aware Scheduling Policies for Varying WorkloadsabstractOver the past decade, the compute resource requirements of applications has continued to grow. This leads to an increase in the required infrastructure that poses additional challenges to mitigating energy consumption and peak power draws. Apache Mesos has been a dominant player in the cluster management space, acting as a middleware between frameworks and the underlying infrastructure. Due to its high availability, fault-tolerance, and scalability, Mesos is used in the industry to run massive scale workloads. Datacenters typically handle continuous dynamically varying workloads, potentially leading to high peak power and energy consumption, in-turn leading to high cost of operation. Given a set of scheduling policies that have been proven to be beneficial in lowering peak power and/or energy consumption, we propose a mechanism of switching between these policies in order to adapt to the dynamic variation in the workloads and the changes in the state of the cluster. Our experiments show that adapting to the variation in the workload can lead to a 9.2% reduction in max peak power consumption, 6.9% in total energy consumption, and a 6.4% reduction in makespan when compared to using a single scheduling policy. Pradyumna Kaushik, Akash Kothawale, Renan Delvalle, Abhishek Jain 0006, Madhusudhan Govindaraju |
IEEE CLOUD | 5 |
| 2018 | Exploring the Fairness and Resource Distribution in an Apache Mesos EnvironmentabstractApache Mesos, a cluster-wide resource manager, is widely deployed in massive scale at several Clouds and Data Centers. Mesos aims to provide high cluster utilization via fine grained resource co-scheduling and resource fairness among multiple users through Dominant Resource Fairness (DRF) based allocation. DRF takes into account different resource types (CPU, Memory, Disk I/O) requested by each application and determines the share of each cluster resource that could be allocated to the applications. Mesos has adopted a two-level scheduling policy: (1) DRF to allocate resources to competing frameworks and (2) task level scheduling by each framework for the resources allocated during the previous step. We have conducted experiments in a local Mesos cluster when used with frameworks such as Apache Aurora, Marathon, and our own framework Scylla, to study resource fairness and cluster utilization. Experimental results show how informed decision regarding second level scheduling policy of frameworks and attributes like offer holding period, offer refusal cycle and task arrival rate can reduce unfair resource distribution. Bin-Packing scheduling policy on Scylla with Marathon can reduce unfair allocation from 38% to 3%. By reducing unused free resources in offers we bring down the unfairness from to 90% to 28%. We also show the effect of task arrival rate to reduce the unfairness from 23% to 7%. Pankaj Saha, Angel Beltre, Madhusudhan Govindaraju |
IEEE CLOUD | 3 |
| 2018 | CAPI-Flash Accelerated Persistent Read Cache for Apache CassandraabstractIn real-world NoSQL deployments, users have to trade off CPU, memory, I/O bandwidth and storage space to achieve the required performance and efficiency goals. Data compression is a vital component to improve storage space efficiency, but reading compressed data increases response time. Therefore, compressed data stores rely heavily on using the memory as a cache to speed up read operations. However, as large DRAM capacity is expensive, NoSQL databases have become costly to deploy and hard to scale. In our work, we present a persistent caching mechanism for Apache Cassandra on a high-throughput, low-latency FPGA-based NVMe Flash accelerator (CAPI-Flash), replacing Cassandra's in-memory cache. Because flash is dramatically less expensive per byte than DRAM, our caching mechanism provides Apache Cassandra with access to a large caching layer at lower cost. The experimental results show that for read-intensive workloads, our caching layer provides up to 85% improved throughput and also reduces CPU usage by 25% compared to default Cassandra. Bedri Sendir, Madhusudhan Govindaraju, Rei Odaira, H. Peter Hofstee |
IEEE CLOUD | 2 |
| 2017 | Electron: Towards Efficient Resource Management on Heterogeneous Clusters with Apache MesosabstractAs data centers continue to grow in scale, the resource management software needs to work closely with the hardware infrastructure to provide high utilization, performance, fault tolerance, and high availability. Apache Mesos has emerged as a leader in this space, providing an abstraction over the entire cluster, data center, or cloud to present a uniform view of all the resources. In addition, frameworks built on Mesos such as Apache Aurora, developed within Twitter and later contributed to the Apache Software Foundation, allow massive job submissions with heterogeneous resource requirements. The availability of such tools in the Open Source space, with proven record of large scale production use, make them suitable for research on how they can be adapted for use in campus-clusters and emerging cloud infrastructures for different workloads in both academia and industry. As data centers run these workloads and strive to maintain high utilization of their components, they suffer a significant cost in terms of energy and power consumption. To address this cost we have developed our own framework, Electron, for use with Mesos. Electron is designed to be configurable with heuristic-driven power capping policies along with different scheduling policies such as Bin Packing and First Fit. We characterize the performance of Electron, in comparison with the widely used Aurora framework. On average, our experiments show that Electron can reduce the 95th percentile of CPU and DRAM power usage by 27.89%, total energy consumption by 19.15%, average power consumption by 27.90%, and max peak power usage by 16.91%, while maintaining a similar makespan when compared to Aurora using the proper combination of power capping and scheduling policies. Renan Delvalle, Pradyumna Kaushik, Abhishek Jain 0006, Jessica Hartog, Madhusudhan Govindaraju |
CLOUD | 5 |
| 2016 | Exploring the Design Space for Optimizations with Apache Aurora and MesosabstractCloud infrastructures increasingly include a heterogeneous mix of components in terms of performance, power, and energy usage. As the size of cloud infrastructures grows, power consumption becomes a significant constraint. We use Apache Mesos and Apache Aurora, which provide massive scalability to Web-scale applications, to demonstrate how a policy driven approach involving bin-packing workloads according to their power profiles, instead of the default allocation by Mesos and Aurora, can effectively reduce the peak-power and energy usage as well as the node utilization, when workloads are co-scheduled. Our experimental results show reductions of 11% in peak power, 86% for total energy usage, and an increase in utilization of 148% for memory and 8% CPU for the different policies. Renan Delvalle, Gourav Rattihalli, Angel Beltre, Madhusudhan Govindaraju, Michael J. Lewis |
CLOUD | 4 |
| 2016 | Optimized Durable Commitlog for Apache Cassandra Using CAPI-FlashabstractHigh-velocity data imposes high durability overheads on Big Data technology components such as NoSQL data stores. In Apache Cassandra, a widely used NoSQL solution with high scalability and availability, write-ahead logging is used to support Commitlog operations, which in turn provides fault tolerance to applications. However, current write-ahead logging techniques are limited by the excessive overhead in the I/O subsystem. To address this performance gap, we have designed a novel CAPI-Flash based high performance durable Commitlog for Apache Cassandra. We take advantage of the high throughput, low latency path to flash storage provided by the Coherent Accelerator Processor Interface (CAPI) on IBM POWER8 Systems. Our experimental results show that for write-intensive workloads CAPI-Flash logging provides up to 107% improvement in throughput compared to Cassandra's durable alternative. We also provide 77% better throughput in update-mostly workloads. Bedri Sendir, Madhusudhan Govindaraju, Rei Odaira, H. Peter Hofstee |
CLOUD | 2 |
| 2016 | Integrating Apache Airavata with Docker, Marathon, and MesosabstractSummary Science Gateways provide scientists with tools for creating, executing, and monitoring scientific experiments on multiple resource infrastructures. Apache Airavata abstracts interactions between gateways and distributed computing infrastructures like Extreme Science and Engineering Discovery Environment, international grids, and campus clusters. Airavata consists of several component services such as the API server, Orchestrator, Workflow Interpreter, Credential Store, and Application Factory. In addition, Airavata uses third party software, including RabbbitMQ for messaging, MySQL for production database management, and Apache Zookeeper for internal communications. In this paper, we discuss our initial experiences with leveraging open source technologies to manage Airavata and its dependent components to deploy, detect, and restart failed components in an auto‐scaling platform. Such capabilities will allow Airavata services to be deployed in a wide area, large Virtual Machine (VM) based cluster, and a developer's laptop. The emerging technologies in cloud computing and Big Data that address these needs are the following: Docker, Marathon, and Apache Mesos. Docker is a Linux‐based lightweight container that allows different applications to run isolated from each other but safely share the machine's resources. Docker images of applications can be published in registries and retrieved for execution in the target infrastructures. Marathon provides a cluster‐wide init and control system for services, including Docker containers. Mesos provides a cluster‐wide framework to schedule tasks based on fine‐grained resource needs. Mesosphere provides the packages, scripts, and web interface to ease the use of these technologies. We present the design, experience, and lessons learned from integrating Mesos, Docker, and Marathon with Apache Airavata. Copyright © 2015 John Wiley & Sons, Ltd. Pankaj Saha, Madhusudhan Govindaraju, Suresh Marru, Marlon E. Pierce |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | Processing Cassandra Datasets with Hadoop-Streaming Based ApproachesabstractThe progressive transition in the nature of both scientific and industrial datasets has been the driving force behind the development and research interests in the NoSQL model. Loosely structured data poses a challenge to traditional data store systems, and when working with the NoSQL model, these systems are often considered impractical and costly. As the quantity and quality of unstructured data grows, so does the demand for a processing pipeline that is capable of seamlessly combining the NoSQL storage model and a “Big Data” processing platform such as MapReduce. Although MapReduce is the paradigm of choice for data-intensive computing, Java-based frameworks such as Hadoop require users to write MapReduce code in Java while Hadoop Streaming module allows users to define non-Java executables as map and reduce operations. When confronted with legacy C/C++ applications and other non-Java executables, there arises a further need to allow NoSQL data stores access to the features of Hadoop Streaming. We present approaches in solving the challenge of integrating NoSQL data stores with MapReduce under non-Java application scenarios, along with advantages and disadvantages of each approach. We compare Hadoop Streaming alongside our own streaming framework, MARISSA, to show performance implications of coupling NoSQL data stores like Cassandra with MapReduce frameworks that normally rely on file-system based data stores. Our experiments also include Hadoop-C*, which is a setup where a Hadoop cluster is co-located with a Cassandra cluster in order to process data using Hadoop with non-java executables. Elif Dede, Bedri Sendir, Pinar Kuzlu, J. Weachock, Madhusudhan Govindaraju, Lavanya Ramakrishnan |
IEEE Trans. Serv. Comput. | 5 |
| 2014 | Benchmarking MapReduce implementations under different application scenarios
Elif Dede, Zacharia Fadika, Madhusudhan Govindaraju, Lavanya Ramakrishnan |
Future Gener. Comput. Syst. | 3 |
| 2014 | MARIANE: Using MApReduce in HPC environments
Zacharia Fadika, Elif Dede, Madhusudhan Govindaraju, Lavanya Ramakrishnan |
Future Gener. Comput. Syst. | 3 |
| 2013 | An Evaluation of Cassandra for HadoopabstractIn the last decade, the increased use and growth of social media, unconventional web technologies, and mobile applications, have all encouraged development of a new breed of database models. NoSQL data stores target the unstructured data, which by nature is dynamic and a key focus area for "Big Data" research. New generation data can prove costly and unpractical to administer with SQL databases due to lack of structure, high scalability, and elasticity needs. NoSQL data stores such as MongoDB and Cassandra provide a desirable platform for fast and efficient data queries. This leads to increased importance in areas such as cloud applications, e-commerce, social media, bioinformatics, and materials science. In an effort to combine the querying capabilities of conventional database systems and the processing power of the MapReduce model, this paper presents a thorough evaluation of the Cassandra NoSQL database when used in conjunction with the Hadoop MapReduce engine. We characterize the performance for a wide range of representative use cases, and then compare, contrast, and evaluate so that application developers can make informed decisions based upon data size, cluster size, replication factor, and partitioning strategy to meet their performance needs. Elif Dede, Bedri Sendir, Pinar Kuzlu, Jessica Hartog, Madhusudhan Govindaraju |
IEEE CLOUD | 5 |
| 2013 | Processing HDF5 Datasets on Multi-core ArchitecturesabstractIn order to make scientific middleware and applications more scalable, there is a need to design them in such a way that they can utilize the evolving multi-core processor architectures available in grid and cloud computing environments. In this paper, we analyze various processing and scheduling techniques on multi-core architectures based on scientific data characteristics and access patterns. More specifically, we conduct fine-grained analysis of scientific datasets such as HDF5 to make effective processing and scheduling decisions in multi-threaded programming. We present performance analysis on how processing threads can be scheduled on multi-core nodes to enhance the performance of scientific applications that process HDF5 data. To accomplish this we introduce a dynamic marking scheme to keep track of the progress of threads on each core. This can be used to help determine work allocation, which results in a decrease in overall application execution time. Rajdeep Bhowmik, Jessica Hartog, Madhusudhan Govindaraju |
AINA | 3 |
| 2012 | Evaluating Hadoop for Data-Intensive Scientific OperationsabstractEmerging sensor networks, more capable instruments, and ever increasing simulation scales are generating data at a rate that exceeds our ability to effectively manage, curate, analyze, and share it. Data-intensive computing is expected to revolutionize the next-generation software stack. Hadoop, an open source implementation of the MapReduce model provides a way for large data volumes to be seamlessly processed through use of large commodity computers. The inherent parallelization, synchronization and fault-tolerance the model offers, makes it ideal for highly-parallel data-intensive applications. MapReduce and Hadoop have traditionally been used for web data processing and only recently been used for scientific applications. There is a limited understanding on the performance characteristics that scientific data intensive applications can obtain from MapReduce and Hadoop. Thus, it is important to evaluate Hadoop specifically for data-intensive scientific operations -- filter, merge and reorder-- to understand its various design considerations and performance trade-offs. In this paper, we evaluate Hadoop for these data operations in the context of High Performance Computing (HPC) environments to understand the impact of the file system, network and programming modes on performance. Zacharia Fadika, Madhusudhan Govindaraju, Shane Canon, Lavanya Ramakrishnan |
IEEE CLOUD | 2 |
| 2012 | Configuring a MapReduce Framework for Dynamic and Efficient Energy AdaptationabstractMapReduce has become a popular framework for Big Data applications. While MapReduce has received much praise for its scalability and efficiency, it has not been thoroughly evaluated for power consumption. Our goal with this paper is to explore the possibility of scheduling in a power-efficient manner without the need for expensive power monitors on every node. We begin by considering that no cluster is truly homogeneous with respect to energy consumption. From there we develop a MapReduce framework that can evaluate the current status of each node and dynamically react to estimated power usage. Inso doing, we shift power consumption work toward more energy efficient nodes which are currently consuming less power. Our work shows that given an ideal framework configuration, certain nodes may consume only 62.3% of the dynamic power they consumed when the same framework was configured as it would be in a traditional MapReduce implementation. Jessica Hartog, Zacharia Fadika, Elif Dede, Madhusudhan Govindaraju |
IEEE CLOUD | 4 |
| 2012 | MARLA: MapReduce for Heterogeneous ClustersabstractMapReduce has gradually become the framework of choice for "big data". The MapReduce model allows for efficient and swift processing of large scale data with a cluster of compute nodes. However, the efficiency here comes at a price. The performance of widely used MapReduce implementations such as Hadoop suffers in heterogeneous and load-imbalanced clusters. We show the disparity in performance between homogeneous and heterogeneous clusters in this paper to be high. Subsequently, we present MARLA, a MapReduce framework capable of performing well not only in homogeneous settings, but also when the cluster exhibits heterogeneous properties. We address the problems associated with existing MapReduce implementations affecting cluster heterogeneity, and subsequently present through MARLA the components and trade-offs necessary for better MapReduce performance in heterogeneous cluster and cloud environments. We quantify the performance gains exhibited by our approach against Apache Hadoop and MARIANE in data intensive and compute intensive applications. Zacharia Fadika, Elif Dede, Jessica Hartog, Madhusudhan Govindaraju |
CCGRID | 4 |
| 2012 | MARISSA: MApReduce Implementation for Streaming Science ApplicationsabstractMapReduce has since its inception been steadily gaining ground in various scientific disciplines ranging from space exploration to protein folding. The model poses a challenge for a wide range of current and legacy scientific applications for addressing their “Big Data” challenges. For example: MapRe-duce's best known implementation, Apache Hadoop, only offers native support for Java applications. While Hadoop streaming supports applications compiled in a variety of languages such as C, C++, Python and FORTRAN, streaming has shown to be a less efficient MapReduce alternative in terms of performance, and effectiveness. Additionally, Hadoop streaming offers lesser options than its native counterpart, and as such offers less flexibility along with a limited array of features for scientific software. The Hadoop File System (HDFS), a central pillar of Apache Hadoop is not a POSIX compliant file system. In this paper, we present an alternative framework to Hadoop streaming to address the needs of scientific applications: MARISSA (MApReduce Implementation for Streaming Science Applications). We describe MARISSA's design and explain how it expands the scientific applications that can benefit from the MapReduce model. We also compare and explain the performance gains of MARISSA over Hadoop streaming. Elif Dede, Zacharia Fadika, Jessica Hartog, Madhusudhan Govindaraju, Lavanya Ramakrishnan, Dan Gunter, Shane Canon |
eScience | 4 |
| 2012 | L2 Cache Performance Analysis and Optimizations for Processing HDF5 Data on Multi-core NodesabstractIt is important to design and develop scientific middleware libraries to harness the opportunities presented by emerging multi-core processors that are available on grid and cloud environments. Scientific middleware libraries not adhering or adapting to this programming paradigm can suffer from severe performance limitations while executing on emerging multi-core processors. In this paper, we focus on the utilization of a critical shared resource on chip multiprocessors (CMPs), the L2 cache. The way in which an application schedules and assigns processing work to each thread determines the access pattern of the shared L2 cache, which may result in either enhancing or diminishing the effects of memory latency on a multi-core processor. Therefore, while processing scientific datasets such as HDF5, it is essential to conduct fine-grained analysis of cache utilization, to make informed processing and scheduling decisions in multi-threaded programming. In this paper, using the TAU toolkit for performance feedback from dual- and quad-core machines, we analyze and recommend methods for effective scheduling of threads on multi-core nodes to augment the performance of scientific applications processing HDF5 data. We discuss the benefits that can be achieved by using L2 Cache-Affinity and L2 Balanced-Set based scheduling algorithms for improving L2 cache performance and effectively the overall execution time. Rajdeep Bhowmik, Madhusudhan Govindaraju |
ISPA | 2 |
| 2011 | DELMA: Dynamically ELastic MapReduce Framework for CPU-Intensive ApplicationsabstractSince its introduction, MapReduce implementations have been primarily focused towards static compute cluster sizes. In this paper, we introduce the concept of dynamic elasticity to MapReduce. We present the design decisions and implementation tradeoffs for DELMA, (Dynamically Elastic MapReduce), a framework that follows the MapReduce paradigm, just like Hadoop MapReduce, but that is capable of growing and shrinking its cluster size, as jobs are underway. In our study, we test DELMA in diverse performance scenarios, ranging from diverse node additions to node additions at various points in the application run-time with various dataset sizes. The applicability of the MapReduce paradigm extends far beyond its use with large-scale data intensive applications, and can also be brought to bear in processing long running distributed applications executing on small-sized clusters. In this work, we focus both on the performance of processing hierarchical data in distributed scientific applications, as well as the processing of smaller but demanding input sizes primarily used in small clusters. We run experiments for datasets that require CPU intensive processing, ranging in size from Millions of input data elements to process, up to over half a billion elements, and observe the positive scalability patterns exhibited by the system. We show that for such sizes, performance increases accordingly with data and cluster size increases. We conclude on the benefits of providing MapReduce with the capability of dynamically growing and shrinking its cluster configuration by adding and removing nodes during jobs, and explain the possibilities presented by this model. Zacharia Fadika, Madhusudhan Govindaraju |
CCGRID | 2 |
| 2011 | Automatic Creation of an Ontological Knowledge Base from Grid and Cloud-based WikipagesabstractAs the size of experimental data, documents, and web pages grows larger, it will be difficult for scientists to search for appropriate data and grid web-services through keyword-based grid search portal interfaces. A well designed and dynamically updated knowledge base can help understand user queries when they are structurally different, but semantically correlated, with actual data stored in a grid or cloud. In this paper we discuss the implementation of a knowledge base creation tool for grid and cloud based wikipages. In order to achieve accurate results in the query match-making process, we store the knowledge base in the rich OWL format and update it automatically with subsequent inference of new facts. Our framework extracts grid and cloud related wikipages and web pages using syntactic Link Grammar Parser and creates core ontology models specific to the grid and cloud domain. In this paper we describe two core components of the ontology generation framework - an information extraction framework and a core ontology model. We present the accuracy of search in terms of precision and recall and its relationship with the dynamically updated ontological domain knowledge base. Chaitali Gupta, Madhusudhan Govindaraju |
CloudCom | 2 |
| 2011 | Adapting MapReduce for HPC environmentsabstractMapReduce is increasingly gaining popularity as a programming model for use in large-scale distributed processing. The model is most widely used when implemented using the Hadoop Distributed File System (HDFS). The use of the HDFS, however, precludes the direct applicability of the model to HPC environments, which use high performance distributed file systems. In such distributed environments, the MapReduce model can rarely make use of full resources, as local disks may not be available for data placement on all the nodes. This work proposes a MapReduce implementation and design choices directly suitable for such HPC environments. Zacharia Fadika, Elif Dede, Madhusudhan Govindaraju, Lavanya Ramakrishnan |
HPDC | 3 |
| 2010 | Cache Performance Optimization for Processing XML-Based Application Data on Multi-core ProcessorsabstractThere is a critical need to develop new programming paradigms for grid middleware tools and applications to harness the opportunities presented by emerging multi-core processors. Implementations of grid middleware and applications that do not adapt to the programming paradigm when executing on emerging processors can severely impact the overall performance. In this paper we focus on the utilization of the L2 cache, which is a critical shared resource on Chip Multiprocessors. The access pattern of the shared L2 cache, which is dependent on how the application schedules and assigns processing work to each thread, can either enhance or undermine the ability to hide memory latency on a multi-core processor. None of the current grid simulators and emulators provides feedback and fine-grained performance data that is essential for a detailed analysis. In this paper, using the feedback from an emulation framework, we present performance analysis and provide recommendations on how processing threads can be scheduled on multi-core nodes to enhance the performance of a class of grid applications that requires processing of large-scale XML data. In particular, we discuss the gains associated with the use of the adaptations we have made to the Cache-Affinity and Balanced-Set scheduling algorithms to improve L2 cache performance, and hence the overall application execution time. Rajdeep Bhowmik, Madhusudhan Govindaraju |
CCGRID | 2 |
| 2010 | Framework for Efficient Indexing and Searching of Scientific MetadataabstractA seamless and intuitive data reduction capability for the vast amount of scientific metadata generated by experiments is critical to ensure effective use of the data by domain specific scientists. The portal environments and scientific gateways currently used by scientists provide search capability that is limited to the pre-defined pull-down menus and conditions set in the portal interface. Currently, data reduction can only be effectively achieved by scientists who have developed expertise in dealing with complex and disparate query languages. A common theme in our discussions with scientists is that data reduction capability, similar to web search in terms of ease-of-use, scalability, and freshness/accuracy of results, is a critical need that can greatly enhance the productivity and quality of scientific research. Most existing search tools are designed for exact string matching, but such matches are highly unlikely given the nature of metadata produced by instruments and a user’s inability to recall exact numbers to search in very large datasets. This paper presents research to locate metadata of interest within a range of values. To meet this goal, we leverage the use of XML in metadata description for scientific datasets, specifically the NeXus datasets generated by the SNS scientists. We have designed a scalable indexing structure for processing data reduction queries. Web semantics and ontology based methodologies are also employed to provide an elegant, intuitive, and powerful free-form query based data reduction interface to end users. Chaitali Gupta, Madhusudhan Govindaraju |
CCGRID | 2 |
| 2010 | LEMO-MR: Low Overhead and Elastic MapReduce Implementation Optimized for Memory and CPU-Intensive ApplicationsabstractSince its inception, MapReduce has frequently been associated with Hadoop and large-scale datasets. Its deployment at Amazon in the cloud, and its applications at Yahoo! and Face book for large-scale distributed document indexing and database building, among other tasks, have thrust MapReduce to the forefront of the data processing application domain. The applicability of the paradigm however extends far beyond its use with data intensive applications and disk based systems, and can also be brought to bear in processing small but CPU intensive distributed applications. In this work, we focus both on the performance of processing large-scale hierarchical data in distributed scientific applications, as well as the processing of smaller but demanding input sizes primarily used in diskless, and memory resident I/O systems. In this paper, we present LEMO-MR (Low overhead, Elastic, configurable for in-Memory applications, and on-Demand fault tolerance), an optimized implementation of MapReduce, for both on-disk and in-memory applications, describe its architecture and identify not only the necessary components of this model, but also trade offs and factors to be considered. We show the efficacy of our implementation in terms of potential speedup that can be achieved for representative data sets used by cloud applications. Finally, we quantify the performance gains exhibited by our MapReduce implementation over Apache Hadoop in a compute intensive environment. Zacharia Fadika, Madhusudhan Govindaraju |
CloudCom | 2 |
| 2009 | Performance enhancement with speculative execution based parallelism for processing large-scale xml-based application dataabstractWe present the design and implementation of a toolkit for processing large-scale XML datasets that utilizes the capabilities for parallelism that are available in the emerging multi-core architectures. Multi-core processors are expected to be widely available in research clusters and scientific desktops, and it is critical to harness the opportunities for parallelism in the middleware, instead of passing on the task to application programmers. An emerging trend is the use of XML as the data format for many distributed/grid applications, with the size of these documents ranging from tens of megabytes to hundreds of megabytes. Our earlier benchmarking results revealed that most of the widely available XML processing toolkits do not scale well for large sized XML data. A significant transformation is necessary in the design of XML processing for distributed applications so that the overall application turn-around time is not negatively affected by XML processing. We discuss XML processing using PiXiMaL, a parallel processing library for large-scale XML datasets. The parallelization approach is to build a DFA-based parser that recognizes a useful subset of the XML specification, and convert the DFA into an NFA that can be applied to an arbitrary subset of the input. Speculative NFAs are scheduled on available cores in a node to effectively utilize the processing capabilities and achieve overall performance gains. We evaluate the efficacy of this approach in terms of potential speedup that can be achieved for representative XML datasets. We also evaluate the effect of two different memory allocation libraries to quantify the memory-bottleneck as different cores access shared data structures. Michael R. Head, Madhusudhan Govindaraju |
HPDC | 2 |
| 2008 | Analysis of Cache Performance for Processing XML-Based Application Data on Multi-core ProcessorsabstractComputer architecture is now at an important juncture as single-core CPU power is expected to be nearly constant. The microprocessor industry is rapidly moving towards chip multi-processors (CMPs), commonly referred to as multi-core processors. The transition of CPUs from single to multi-core implementations requires a corresponding shift in the programming paradigm for grid and e-science libraries. Naive implementations of processing on multi-core systems can severely impact performance because of limitations of shared bus bandwidth, cache size and coherency, and communication between threads. To optimize the performance of e-science services, careful application of thread-level parallelism is needed. We study this problem in the context of processing XML data used in grid and e-science applications. The web services model, which strongly leverages XML, has been adopted as the basic architecture for grid and e-science services. As a result, the optimization of separate Web services applications is critical because Web services that are deployed in a longer chain of service processing events must guarantee minimal response times to ensure overall system performance. Our goal is to analyze and provide insightful feedback on cache behavior of each core and reveal performance limitations, bottlenecks, and multi-threaded optimization opportunities for processing XML data relevant to grid and e-science application data formats. We use a micro-architectural emulation framework, Multi-core Grid (McGrid), to generate performance data at various levels of granularity. We analyze cache behavior to quantify the exact gains and present recommendations for processing XML data in grid and e-science applications that will be deployed on emerging multi-core systems. Rajdeep Bhowmik, Madhusudhan Govindaraju |
eScience | 2 |
| 2008 | Semantic Framework for Free-Form Search of Grid ResourcesabstractThe model of free-form queries has been an enormous success for HTML-based search engines on the web. If the same free-form search is made available for grid services, it will serve as a powerful tool for scientists to retrieve information on resources, monitoring data, replica location sets, and meta-data on scientific data sets, in an intuitive manner. Current implementations of XML-based grid service descriptions require end users to have intimate knowledge of service descriptions, related toolkits, and query languages. We have developed a system that abstracts away these fundamental complexities and provides a simple free-form query based scientific discovery system for grid users. In this paper, we present the design and initial implementation results of our ontological framework that employs matching algorithms, automated extension of ontologies, and semantic Web and ontological concepts to match free-form queries with corresponding grid resource information stored in RDF/OWL format. Chaitali Gupta, Madhusudhan Govindaraju |
eScience | 2 |
| 2008 | Parallel Processing of Large-Scale XML-Based Application Documents on Multi-core Architectures with PiXiMaLabstractVery large scientific datasets are becoming increasingly available in XML formats. Our earlier benchmarking results show that parsing XML is a time consuming process when compared with binary formats optimized for largescale documents. This performance bottleneck will get exacerbated as size of XML data increases in e-science applications. Our focus in this paper is on addressing this performance bottleneck. In recent times, the microprocessor industry has made rapid strides towards chip multi processors (CMPs). The widely available XML parsers have been unable to take advantage of the opportunities presented by CMPs, instead, passing the task of parallelization to the application programmer. The paradigms used thus far to process large size XML documents on uniprocessors are not applicable for CMPs. We present the design, implementation, and performance analysis of PiXiMaL, a parallel processing library for large-scale XML-data files. In particular, we discuss an effective scheme to parallelize the tokenization process to achieve an overall performance increase when parsing large-scale XML documents that are increasingly in use today. Our approach is to build a DFA-based parser that recognizes a useful subset of the XML specification and converts the DFA into an NFA which can be applied on any subset of the input. Michael R. Head, Madhusudhan Govindaraju |
eScience | 2 |
| 2008 | Ontological framework for a free-form query based grid search engineabstractIf the model of free-form queries, which has proved successful for HTML based search on the Web, is made available for Grid services, it will serve as a powerful tool for scientists to retrieve information on resources, monitoring, replica location sets, and meta-data on scientific data sets, etc., in a seamless manner. To enable this vision, there is a critical need to design and develop tools that abstract away the fundamental complexity of XML based Grid specifications and toolkits, and provide an elegant, intuitive, simple, and powerful free-form query based invocation system to end users. Current implementations of XML-based Grid service descriptions require end users to have intimate knowledge of service descriptions, related toolkits, and query languages. We present our research project and initial results that employ self-learning mechanisms, matching algorithms and optimizations to match free-form user queries with corresponding operations in Grid services, and present the results to the end user. Our system uses Semantic Web concepts and Ontologies to automate discovery and matchmaking of Grid services. The research focus of this project is on the development of novel algorithms for matching user queries with correct operation names and quantifying the exact gains in accuracy due to knowledge acquisition. Chaitali Gupta, Rajdeep Bhowmik, Madhusudhan Govindaraju |
HPDC | 3 |
| 2008 | Optimizing XML processing for grid applications using an emulation frameworkabstractChip multi-processors (CMPs), commonly referred to as multi-core processors, are being widely adopted for deployment as part of the grid infrastructure. This change in computer architecture requires corresponding design modifications in programming paradigms, including grid middleware tools, to harness the opportunities presented by multi-core processors. Simple and naive implementations of grid middleware on multi-core systems can severely impact performance. This is because programming for CMPs requires special consideration for issues such as limitations of shared bus bandwidth, cache size and coherency, and communication between threads. The goal of developing an optimized multi-threaded grid middleware for emerging multi-core processors will be realized only if researchers and developers have access to an in-depth analysis of the impact of several low level microarchitectural parameters on performance. None of the current grid simulators and emulators provide feedback at the microarchitectural level, which is essential for such an analysis. In earlier work we presented our initial results on the design and implementation of such an emulation framework, Multi- core Grid (McGrid). In this paper we extend that work and present a performance study on the effect of cache coherency, scheduling of processing threads to take advantage of data available in the cache of each core, and read and write access patterns for shared data structures. We present the performance results, analysis, and recommendations based on experiments conducted using the McGrid framework for processing XML-based grid data and documents. Rajdeep Bhowmik, Chaitali Gupta, Madhusudhan Govindaraju, Aneesh Aggarwal |
IPDPS | 3 |
| 2008 | Grid-based research, development, and deployment in New York StateabstractIn this paper, we present cyberinfrastructure and grid computing efforts in New York State. In particular, we focus on fundamental efforts in Binghamton and Buffalo, including the design, development, and deployment of the New York State Grid, as well as a grass-roots New York State Initiative. Russ Miller, Jonathan J. Bednasz, Kenneth Chiu, Steven M. Gallo, Madhusudhan Govindaraju |
IPDPS | 5 |
| 2007 | Efficient XML-Based Grid Middleware Design for Multi-Core ProcessorsabstractChip multi-processors (CMPs), commonly referred to as multi-core processors, are being widely adopted for deployment as part of the grid infrastructure. In CMPs, multiple cores can independently execute different threads. This change in computer architecture requires corresponding design modifications in programming paradigms, including grid middleware tools, to harness the opportunities presented by multi- core processors. Simple and naive implementations of grid middleware on multi-core systems can severely impact performance. The goal of developing an optimized multi-threaded grid middleware for emerging multi-core processors will be realized only if researchers and developers have access to an in-depth analysis of the impact of several low level microarchitectural parameters on performance. None of the current grid simulators and emulators provides feedback at the micro-architectural level. We have designed an emulation framework, Multi-core Grid (McGrid), to analyze and provide insightful feedback on the performance limitations, bottlenecks, and optimization opportunities for grid middleware on multi-core systems. Rajdeep Bhowmik, Chaitali Gupta, Madhusudhan Govindaraju, Aneesh Aggarwal |
ICWS | 3 |
| 2007 | XML Schemas Based Flexible Distributed Code Generation FrameworkabstractTo leverage the strengths of different distributed systems, a user may want to develop an application that spans multiple frameworks. Currently, in such cases a user needs to use a different code generator for each one of the distributed frameworks that the application incorporates. Learning the details of the interface specification and code generation tool for each one of these distributed frameworks is tedious and error-prone Instead, it is desirable to present the user with a flexible code generation framework that leverages the wide variety of available XML-based tools and is capable of generating code for all distributed systems. Our work involves the design and implementation of an XML schema based toolkit that can serve as the universal code generation toolkit for distributed frameworks. Madhusudhan Govindaraju |
ICWS | 1 |
| 2007 | A Query-based System for Automatic Invocation of Web ServicesabstractThere is a critical need to design and develop tools that abstract away the fundamental complexity of XML based Web services specifications and toolkits, and provide an elegant, intuitive, simple, and powerful query based invocation system to end users. Web services based tools and standards have been designed to facilitate seamless integration and development for application developers. As a result, current implementations require the end user to have intimate knowledge of Web services and related toolkits, and users often play an informed role in the overall Web services execution process We employ a set of algorithms and optimizations to match user queries with corresponding operations in Web services, invoke the operations with the correct set of parameters, and present the results to the end user. Our system uses the Semantic Web and Ontologies in the process of automating Web services invocation and execution. Chaitali Gupta, Rajdeep Bhowmik, Michael R. Head, Madhusudhan Govindaraju, Weiyi Meng |
ICWS | 4 |
| 2007 | Search Algorithms for Discovery of Web ServicesabstractWeb services are designed to standardize interactions between heterogeneous applications using Internet technologies. Within the framework of Internet search technologies, web services provide structured channels to access search engines and web-accessible databases. Our work involves research in methods to discover web service description language (WSDL) documents, which provide interface formats, expected data-types, supported protocols and precise service endpoints. This project extends current discovery research through use of the Google web service, UDDI category searching, and private registry querying with preliminary experiments resulting in a very high percentage of success. The goal is to find WSDL documents for a given domain name, parse the desired service document to obtain invocation formats, and automatically invoke the web service. Contributions of this research will support enhancements of HTML-dependent search tools by providing access to data inaccessible through surface HTML interfaces. Janette Hicks, Madhusudhan Govindaraju, Weiyi Meng |
ICWS | 2 |
| 2007 | Securing Grid Data Transfer Services with Active Network PortalsabstractWidely available and utilized grid servers are vulnerable to a variety of threats from denial of service (DoS) attacks, overloading caused by flash crowds, and compromised client machines. The focus of our paper is the design, implementation and evaluation of a set of admission control policies that permit the server to maintain sustained throughput to legitimate clients even in the face of such overloads and attacks. We propose several schemes to effectively, and importantly in an automated fashion, deal with these attacks and overloads. We discuss how these schemes can be efficiently implemented on an active network adapter based gateway that controls access to a pool of backend data servers. Performance tests conducted on a system based on a dual-ported active NIC demonstrate that efficient optimization schemes can be implemented on such a gateway to minimize the grid service response time and to improve server throughputs under heavy loads and DoS attacks. Our results, using the gridFTP server available with Globus Toolkit 4.0.1, demonstrate that even in adverse load conditions, the response times can be maintained at a level similar to normal, low-load conditions. Onur Demir, Michael R. Head, Kanad Ghose, Madhusudhan Govindaraju |
IPDPS | 4 |
| 2007 | Improving Performance of Web Services Query Matchmaking with Automated Knowledge AcquisitionabstractThere is a critical need to design and develop tools that abstract away the fundamental complexity of XML-based Web services specifications and toolkits, and provide an elegant, intuitive, simple, and powerful query-based invocation system to end users. Web services based tools and standards have been designed to facilitate seamless integration and development for application developers. As a result, current implementations require the end user to have intimate knowledge of Web services and related toolkits, and users often play an informed role in the overall Web services execution process. We employ a self-learning mechanism and a set of algorithms and optimizations to match user queries with corresponding operations in Web services. Our system uses Semantic Web concepts and Ontologies in the process of automating Web services matchmaking. We present performance analysis of our system and quantify the exact gains in precision and recall due to the knowledge acquisition algorithms. Chaitali Gupta, Rajdeep Bhowmik, Michael R. Head, Madhusudhan Govindaraju, Weiyi Meng |
Web Intelligence | 4 |
| 2007 | Design and implementation issues for distributed CCA framework interoperabilityabstractAbstract Component frameworks, including those that support the Common Component Architecture (CCA), represent a promising approach to addressing the challenge of building and deploying high‐performance scientific applications in Grid environments, one that is being realized, for example, in our LegionCCA and XCAT‐C++ frameworks. The next step beyond building independent individual frameworks is making them interoperate. Component‐based applications should be able to transparently span multiple disjoint component frameworks with low overhead as compared with the same applications running within a single framework. Interoperable frameworks enable applications to take advantage of more resources, and to better match constituent parts to the underlying resources that best support them. The CCA specification does not prescribe a wire format for inter‐component calls in distributed frameworks, thereby promoting considerable flexibility and customization for the framework developer. This approach thus requires an additional specific strategy outside of the CCA to support interoperability between distributed frameworks. Mandating one common wire format, however, risks choosing the wrong format. We discuss in detail five underlying component framework interoperability requirements, and three general approaches to addressing them. We then discuss how the approaches can be applied to meet the requirements, and address the advantages, issues, and implications of doing so. This effectively defines a design space for framework interoperability approaches. We then address the communication interoperability in detail via a single multi‐protocol communication library called Proteus, and discuss how we have incorporated it into two distinct distributed framework implementations of the CCA specification: LegionCCA and XCAT‐C++. Copyright © 2006 John Wiley & Sons, Ltd. Madhusudhan Govindaraju, Michael J. Lewis, Kenneth Chiu |
Concurr. Comput. Pract. Exp. | 1 |
| 2006 | Exploring Remote Object Coherence in XMLWeb ServicesabstractObject-level coherence in distributed applications and systems has been studied extensively. Object coherence in platform-specific and tightly-coupled systems is achieved with binary serialization protocols to ensure data structures and object graphs are safely transmitted, manipulated, and stored. On the opposite side of the spectrum are platform-neutral Web services that embrace XML as a serialization protocol for building loosely coupled systems. The advantages of XML to connect heterogeneous systems are plenty, but rendering programming-language specific data structures and object graphs in text form incurs a performance hit and presents challenges for systems that require object coherence. Achieving the latter goal poses difficulties by a phenomenon that is sometimes referred to as the "impedance mismatch" between programming language data types and XML schema types. This paper examines the problem, debunks the O/X-mismatch controversy, and presents a mix of static/dynamic algorithms for accurate XML serialization. Experimental results show that the implementation in C/C++ is efficient and competitive to binary protocols. Application of the approach to other programming languages, such as Java, is also discussed Robert A. van Engelen, Madhusudhan Govindaraju, Wei Zhang 0020 |
ICWS | 2 |
| 2006 | Grid scheduling and protocols - Benchmarking XML processors for applications in grid web servicesabstractWeb services based specifications have emerged as the underlying architecture for core grid services and standards, such as WSRF. XML is inextricably inter-twined with Web services based specifications, and as a result the design and implementation of XML processing tools plays a significant role in grid applications. These applications use XML in a wide variety of ways, including workflow specifications, WS-Security based documents, service descriptions in WSDL, and on-the-wire format in SOAP-based communication. The application characteristics also vary widely in the use of XML messages in their performance, memory, size, and processing requirements. Numerous XML processing tools exist today, each of which is optimized for specific features. To make the right decisions, grid application and middleware developers must thus understand the complex dependencies between XML features and the application. We propose a standard benchmark suite for quantifying, comparing, and contrasting the performance of XML processors under a wide range of representative use cases. The benchmarks are defined by a set of XML schemas and conforming documents. To demonstrate the utility of the benchmarks and to provide a snapshot of the current XML implementation landscape, we report the performance of many different XML implementations, on the benchmarks, and draw conclusions about their current performance characteristics. We also present a brief analysis on the current shortcomings and required critical design changes for multi-threaded XML processing tools to run efficiently on emerging multi-core architectures.1 Michael R. Head, Madhusudhan Govindaraju, Robert A. van Engelen, Wei Zhang 0020 |
SC | 2 |
| 2005 | XCAT-C++: Design and Performance of a Distributed CCA Framework
Madhusudhan Govindaraju, Michael R. Head, Kenneth Chiu |
HiPC | 1 |
| 2005 | A Benchmark Suite for SOAP-based Communication in Grid Web ServicesabstractThe convergence of Web services and grid computing has promoted SOAP, a widely used Web services protocol, into a prominent protocol for a wide variety of grid applications. These applications differ widely in the characteristics of their respective SOAP messages, and also in their performance requirements. To make the right decisions, an application developer must thus understand the complex dependencies between the SOAP implementation and the application. We propose a standard benchmark suite for quantifying, comparing, and contrasting the performance of SOAP implementations under a wide range of representative use cases. The benchmarks are defined by a set of WSDL documents. To demonstrate the utility of the benchmarks and to provide a snapshot of the current SOAP implementation landscape, we report the performance of many different SOAP implementations (gSOAP, AxisJava, XSUL and bSOAP) on the benchmarks, and draw conclusions about their current performance characteristics. Michael R. Head, Madhusudhan Govindaraju, Aleksander Slominski, Pu Liu, Nayef Abu-Ghazaleh, Robert A. van Engelen, Kenneth Chiu, Michael J. Lewis |
SC | 2 |
| 2004 | Differential Serialization for Optimized SOAP Performance
Nayef Abu-Ghazaleh, Michael J. Lewis, Madhusudhan Govindaraju |
HPDC | 3 |
| 2003 | Merging the CCA Component Model with the OGSI FrameworkabstractThe most important recent development in Grid systems is the adoption of the Web Services model as its basic architecture. The result is called the Open Grid Services Architecture (OGSA). This paper describes a component framework for distributed Grid applications that is consistent with that model. The framework, called XCAT, is based on the U.S. Department of Energy Common Component Architecture (CCA) but with an implementation based on the standard Web Services stack. Using this framework, an application programmer can compose an application from a set of distributed components. The result is a set of Web Services that collectively represent the executing application instance. This paper describes the basic architecture of XCAT and the design issues to be considered for a component to serve as both a CCA and Open Grid Service Infrastructure (OGSI) service. Madhusudhan Govindaraju, Sriram Krishnan, Kenneth Chiu, Aleksander Slominski, Dennis Gannon, Randall Bramley |
CCGRID | 1 |
| 2002 | Investigating the Limits of SOAP Performance for Scientific ComputingabstractThe growing synergy between Web Services and Grid-based technologies will potentially enable profound, dynamic interactions between scientific applications dispersed in geographic, institutional, and conceptual space. Such deep interoperability requires the simplicity, robustness, and extensibility for which SOAP was conceived, thus making it a natural lingua franca. Concomitant with these advantages, however is a degree of inefficiency that may limit the applicability of SOAP to some situations. We investigate the limitations of SOAP for high-performance scientific computing. We analyze the processing of SOAP messages, and identify the issues of each stage. We present a high-performance SOAP implementation and a schema-specific parser based on the results of our investigation. After our SOAP optimizations are implemented, the most significant bottleneck is ASCII/double conversion. Instead of handling this using extensions to SOAP we recommend a multiprotocol approach that uses SOAP to negotiate faster binary protocols between messaging participants. Kenneth Chiu, Madhusudhan Govindaraju, Randall Bramley |
HPDC | 2 |
| 2002 | The Proteus multiprotocol message libraryabstractGrid systems span manifold organizations and application domains. Because this diverse environment inevitably engenders multiple protocols, interoperability mechanisms are crucial to seamless, pervasive access. This paper presents the design, rationale, and implementation of the Proteus multiprotocol library for integrating multiple message protocols, such as SOAP and JMS, within one system. Proteus decouples application code from protocol code at run-time, allowing clients to incorporate separately developed protocols without recompiling or halting. Through generic serialization, which separates the transfer syntax from the message type, protocols can also be added independently of serialization routines. We also show performance-enhancing mechanisms for Grid services that examine metadata, but pass actual data through opaquely (such as adapters). The interface provided to protocol implementors is general enough to support protocols as disparate as our current implementations: SOAP, JMS, and binary. Proteus is written in C++; a Java port is planned. Kenneth Chiu, Madhusudhan Govindaraju, Dennis Gannon |
SC | 2 |
| 2001 | The XCAT science portalabstractThe design and prototype implementation of the XCAT Grid Science Portal is described in this paper. The portal lets grid application programmers easily script complex distributed computations and package these applications with simple interfaces for others to use. Each application is packaged as a "notebook" which consists of web pages and editable parameterized scripts. The portal is a workstation-based specialized "personal" web server, capable of executing the application scripts and launching remote grid applications for the user. The portal server can receive event streams published by the application and grid resource information published by Network Weather Service (NWS) [32] or Autopilot [15] sensors. Notebooks can be "published" and stored in web based archives for others to retrieve and modify. The XCAT Grid Science Portal has been tested with various applications, including the distributed simulation of chemical processes in semiconductor manufacturing and collaboratory support for X-ray crystallographers. Sriram Krishnan, Randall Bramley, Dennis Gannon, Madhusudhan Govindaraju, Rahul Indurkar, Aleksander Slominski, Benjamin Temko, Jay Alameda, Richard C. Alkire, Timothy O. Drews, Eric Webb |
SC | 4 |
| 2000 | A Component based Services Architecture for Building Distributed ApplicationsabstractDescribes an approach to building a distributed software component system for scientific and engineering applications that is based on representing Computational Grid services as application-level software components. These Grid services provide tools such as registry and directory services, event services and remote component creation. While a service-based architecture for grids and other distributed systems is not new, this framework provides several unique features. First, the public interfaces to each software component are described as XML documents. This allows many adaptors and user interfaces to be generated from the specification dynamically. Second, this system is designed to exploit the resources of existing Grid infrastructures like Globus and Legion, and commercial Internet frameworks like e-speak. Third, and most important, the component-based design extends throughout the system. Hence, tools such as application builders, which allow users to select components, start them on remote resources, and connect and execute them, are also interchangeable software components. Consequently, it is possible to build distributed applications using a graphical "drag-and-drop" interface, a Web-based interface, a scripting language like Python, or an existing tool such as Matlab. Randall Bramley, Kenneth Chiu, Shridhar Diwan, Dennis Gannon, Madhusudhan Govindaraju, Nirmal Mukhi, Benjamin Temko, Madhuri Yechuri |
HPDC | 5 |
| 2000 | Requirements for and Evaluation of RMI Protocols for Scientific ComputingabstractDistributed software component architectures provide promising approach to the problem of building large scale, scientific Grid applications [18]. Communication in these component architectures is based on Remote Method Invocation (RMI) protocols that allow one software component to invoke the functionality of another. Examples include Java remote method invocation (Java RMI)[25] and the new Simple Object Access Protocol (SOAP) [15]. SOAP has the advantage that many programming languages and component frameworks can support it. This paper describes experiments showing that SOAP by itself is not efficient enough for large scale scientific applications. However, when it is embedded in multi-protocol RMI framework, SOAP can be effectively used as a universal control protocol, that can be swapped out by faster, more special purpose protocols when large data transfer speeds are needed. Madhusudhan Govindaraju, Aleksander Slominski, Venkatesh Choppella, Randall Bramley, Dennis Gannon |
SC | 1 |
| 1999 | CAT: A High Performance Distributed Component Architecture Toolkit for the GridabstractGrid systems, such as Globus, Legion and Globe, provide an infrastructure for implementing metacomputing over the Internet. The Component Architecture Toolkit (CAT) provides a software layer above the grid that facilitates programming and end-user interaction with the grid. Juan E. Villacis, Madhusudhan Govindaraju, David Stern, Andrew Whitaker, Fabian Breg, Prafulla Deuskar, Benjamin Temko, Dennis Gannon, Randall Bramley |
HPDC | 2 |