VLDB 2026 Research / reviewers in the wild / expert
Geoffrey C. Fox
dblp:f/GeoffreyFox · also Geoffrey Charles Fox
· DBLP profile ↗
250ranked-venue papers
52as first author
17since 2021 · last 2025
0000-0003-1017-1391ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 161 · 41 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 48 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 23 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 19 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSecurity and privacy · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep RC: A Scalable Data Engineering and Deep Learning Pipeline
Arup Kumar Sarker, Aymen Alsaadi, Alexander James Halpern, Prabhath Tangella, Mikhail Titov, Niranda Perera, Mills Staylor, Gregor von Laszewski, Shantenu Jha, Geoffrey C. Fox |
JSSPP | 10 |
| 2025 | IrrMap: A Large-Scale Comprehensive Dataset for Irrigation Method MappingabstractWe introduce IrrMap, the first large-scale dataset (1.1 million patches) for irrigation method mapping across regions. IrrMap consists of multi-resolution satellite imagery from LandSat and Sentinel, along with key auxiliary data such as crop type, land use, and vegetation indices. The dataset spans 1,668,899 farms and 11,443,492 acres across multiple western U.S. states from 2013 to 2023, providing a rich and diverse foundation for irrigation analysis and ensuring geospatial alignment and quality control. The dataset is ML-ready, with standardized 224×224 GeoTIFF patches, the multiple input modalities, carefully chosen train-test-split data, and accompanying dataloaders for seamless deep learning model training and benchmarking in irrigation mapping. The dataset is also accompanied by a complete pipeline for dataset generation, enabling researchers to extend IrrMap to new regions for irrigation data collection or adapt it with minimal effort for other similar applications in agricultural and geospatial analysis. We also analyze the irrigation method distribution across crop groups, spatial irrigation patterns (using Shannon diversity indices), and irrigated area variations for both LandSat and Sentinel, providing insights into regional and resolution-based differences. To promote further exploration, we openly release IrrMap, along with the derived datasets, benchmark models, and pipeline code, through a GitHub repository: https://github.com/Nibir088/IrrMap and Data repository: https://huggingface.co/Nibir/IrrMap, providing comprehensive documentation and implementation details. Nibir Chandra Mandal, Oishee Bintey Hoque, Abhijin Adiga, Samarth Swarup, Mandy L. Wilson, Lu Feng 0001, Yangfeng Ji, Miaomiao Zhang 0002, Geoffrey C. Fox, Madhav V. Marathe |
KDD (2) | 9 |
| 2025 | Surrogate modeling of Cellular-Potts agent-based models as a segmentation task using the U-Net neural network architectureabstractThe Cellular-Potts model is a powerful and ubiquitous framework for developing computational models for simulating complex multicellular biological systems. Cellular-Potts models (CPMs) are often computationally expensive due to the explicit modeling of interactions among large numbers of individual model agents and diffusive fields described by partial differential equations (PDEs). In this work, we develop a convolutional neural network (CNN) surrogate model using a U-Net architecture that accounts for periodic boundary conditions. We use this model to accelerate the evaluation of a mechanistic CPM previously used to investigate in vitro vasculogenesis. The surrogate model was trained to predict 100 computational steps ahead (Monte-Carlo steps, MCS), accelerating simulation evaluations by a factor of 562 times compared to single-core CPM code execution on CPU. Over short timescales of up to 3 recursive evaluations, or 300 MCS, our model captures the emergent behaviors demonstrated by the original Cellular-Potts model such as vessel sprouting, extension and anastomosis, and contraction of vascular lacunae. This approach demonstrates the potential for deep learning to serve as a step toward efficient surrogate models for CPM simulations, enabling faster evaluation of computationally expensive CPM simulations of biological processes. Tien Comlekoglu, Javier Quetzalcóatl Toledo-Marín, Tina Comlekoglu, Douglas W. DeSimone, Shayn M. Peirce, Geoffrey C. Fox, James A. Glazier |
PLoS Comput. Biol. | 6 |
| 2024 | AstroMAE: Redshift Prediction Using a Masked Autoencoder with a Novel Fine-Tuning ArchitectureabstractRedshift prediction is a fundamental task in astronomy, essential for understanding the expansion of the universe and determining the distances of astronomical objects. Accurate redshift prediction plays a crucial role in advancing our knowledge of the cosmos. Machine learning (ML) methods, renowned for their precision and speed, offer promising solutions for this complex task. However, traditional ML algorithms heavily depend on labeled data and task-specific feature extraction. To overcome these limitations, we introduce AstroMAE, an innovative approach that pretrains a vision transformer encoder using a masked autoencoder method on Sloan Digital Sky Survey (SDSS) images. This technique enables the encoder to capture the global patterns within the data without relying on labels. To the best of our knowledge, AstroMAE represents the first application of a masked autoencoder to astronomical data. By ignoring labels during the pretraining phase, the encoder gathers a general understanding of the data. The pretrained encoder is subsequently fine-tuned within a specialized architecture tailored for redshift prediction. We evaluate our model against various vision transformer architectures and CNN-based models, demonstrating the superior performance of AstroMAE’s pretrained model and fine-tuning architecture. Amirreza Dolatpour Fathkouhi, Geoffrey C. Fox |
e-Science | 2 |
| 2024 | Radical-Cylon: A Heterogeneous Data Pipeline for Scientific Computing
Arup Kumar Sarker, Aymen Alsaadi, Niranda Perera, Mills Staylor, Gregor von Laszewski, Matteo Turilli, Ozgur O. Kilic, Mikhail Titov, André Merzky, Shantenu Jha, Geoffrey C. Fox |
JSSPP | 11 |
| 2023 | Templated Hybrid Reusable Computational Analytics Workflow Management with Cloudmesh, Applied to the Deep Learning MLCommons Cloudmask ApplicationabstractIn this paper, we summarize our effort to create and utilize an integrated framework to coordinate computational AI analytics tasks with the help of a task and experiment management workflow system. Our design is based on a minimalistic approach while at the same time allowing access to hybrid computational resources offered through the owner's computer, HPC computing centers, cloud resources, and distributed systems in general. Access to this framework includes a GUI for monitoring and managing the workflow, a REST service, a command line interface, as well as a Python interface. It uses a template-based batch management system that, through configuration files, easily allows for the generation of reproducible experiments while creating permutations over selected experiment parameters as typical in deep learning applications. The resulting framework was developed for analytics workflows targeting MLCommons benchmarks of AI applications on hybrid computing resources, as well as an educational tool for teaching scientists and students sophisticated concepts to execute computations on resources ranging from a single computer to many thousands of computers as part of on-premise and cloud infrastructure. We demonstrate the usefulness of the tool while creating FAIR principle-based application accuracy benchmark generation for the MLCommons Science Working Group Cloudmask application. The code is available as an open-source project in GitHub and is based on an easy-to-enhance framework called Cloudmesh. It can be applied to other applications easily. Gregor von Laszewski, Jacques Phillipe Fleischer, Geoffrey C. Fox, Juri Papay, Samuel Jackson, Jeyan Thiyagalingam |
e-Science | 3 |
| 2023 | Accurate and Efficient Distributed COVID-19 Spread Prediction based on a Large-Scale Time-Varying People Mobility GraphabstractCompared to previous epidemics, COVID-19 spreads much faster in people gatherings. Thus, we need not only more accurate epidemic spread prediction considering the people gatherings but also more time-efficient prediction for taking actions (e.g., allocating medical equipments) in time. Motivated by this, we analyzed a time-varying people mobility graph of the United States (US) for one year and the effectiveness of previous methods in handling time-varying graphs. We identified several factors that influence COVID-19 spread and observed that some graph changes are transient, which degrades the effectiveness of the previous graph repartitioning and replication methods in distributed graph processing since they generate more time overhead than saved time. Based on the analysis, we propose an accurate and time-efficient Distributed Epidemic Spread Prediction system (DESP). First, DESP incorporates the factors into a previous prediction model to increase the prediction accuracy. Second, DESP conducts repartitioning and replication only when a graph change is stable for a certain time period (predicted using machine learning) to ensure the operation improves time-efficiency. We conducted extensive experiments on Amazon AWS based on real people movement datasets. Experimental results show DESP reduces communication time by up to 52%, while enhancing accuracy by up to 24% compared to existing methods. Sudipta Saha Shubha, Shohaib Mahmud, Haiying Shen, Geoffrey C. Fox, Madhav V. Marathe |
IPDPS | 4 |
| 2023 | In-depth analysis on parallel processing patterns for high-performance Dataframes
Niranda Perera, Arup Kumar Sarker, Mills Staylor, Gregor von Laszewski, Kaiying Shan, Supun Kamburugamuve, Chathura Widanage, Vibhatha Abeykoon, Thejaka Amila Kanewala, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 10 |
| 2022 | Hybrid Cloud and HPC Approach to High-Performance DataframesabstractData pre-processing is a fundamental component in any data-driven application. With the increasing complexity of data processing operations and volume of data, Cylon, a distributed dataframe system, is developed to facilitate data processing both as a standalone application and as a library, especially for Python applications. While Cylon shows promising performance results, we experienced difficulties trying to integrate with frameworks incompatible with the traditional Message Passing Interface (MPI). While MPI implementations encompass scalable and efficient c ommunication routines, their process launching mechanisms work well with mainstream HPC systems but are incompatible with some environments that adopt their own resource management systems. In this work, we alleviated this issue by directly integrating the Unified Communication X (UCX) framework, which supports a variety of classic HPC and non-HPC process-bootstrapping mechanisms as our communication framework. While we experimented with our methodology on Cylon, the same technique can be used to bring MPI communication to other applications that do not employ MPI’s built-in process management approach. Kaiying Shan, Niranda Perera, Damitha Lenadora, Tianle Zhong, Arup Kumar Sarker, Supun Kamburugamuve, Thejaka Amila Kanewala, Chathura Widanage, Geoffrey C. Fox |
IEEE Big Data | 9 |
| 2022 | Stochastic gradient descent-based support vector machines training optimization on Big Data and HPC frameworksabstractSummary Support vector machines (SVM) is a widely used machine learning algorithm. With the increasing amount of research data nowadays, understanding how to do efficient training is more important than ever. This article discusses the performance optimizations and benchmarks related to providing high‐performance support for SVM training. In this research, we have focused on a highly scalable gradient descent‐based approach to implementing the core SVM algorithm. In providing a scalable solution, we have designed optimized high‐performance computing and dataflow‐oriented SVM implementations. A high‐performance computing approach means the algorithm is implemented with the bulk synchronous parallel (BSP) model. In addition, we analyzed the language level optimizations and math kernel optimizations on a prominent HPC modeling programming language (C++) and dataflow modeling programming language (Java). In the experiments, we compared the performance of classic HPC models, classic dataflow models, and hybrid models designed on classic HPC and dataflow programming models. Our research illustrates a scientific approach in designing the SVM algorithm at scale in classic HPC, dataflow, and hybrid systems. Vibhatha Abeykoon, Geoffrey C. Fox, Saliya Ekanayake, Supun Kamburugamuve, Kannan Govindarajan, Pulasthi Wickramasinghe, Niranda Perera, Chathura Widanage, Ahmet Uyar, Gurhan Gunduz, Selahatin Akkas |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Twister2 Cross-platform resource scheduler for big dataabstractAbstract Twister2 is an open‐source big data hosting environment designed to process both batch and streaming data at scale. Twister2 runs jobs in both high‐performance computing (HPC) and big data clusters. It provides a cross‐platform resource scheduler to run jobs in diverse environments. Twister2 is designed with a layered architecture to support various clusters and big data problems. In this paper, we present the cross‐platform resource scheduler of Twister2. We identify required services and explain implementation details. We present job startup delays for single jobs and multiple concurrent jobs in Kubernetes and OpenMPI clusters. We compare job startup delays for Twister2 and Spark at a Kubernetes cluster. In addition, we compare the performance of terasort algorithm on Kubernetes and bare metal clusters at AWS cloud. Ahmet Uyar, Gurhan Gunduz, Supun Kamburugamuve, Pulasthi Wickramasinghe, Chathura Widanage, Kannan Govindarajan, Niranda Perera, Vibhatha Abeykoon, Selahattin Akkas, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 10 |
| 2022 | High-performance iterative dataflow abstractions in Twister2: TSetabstractSummary The dataflow model is gradually becoming the de facto standard for big data applications. While many popular frameworks are built around this model, very little research has been done on understanding its inner workings, which in turn has led to inefficiencies in existing frameworks. It is important to note that understanding the relationship between dataflow and high performance computing (HPC) building blocks allows us to address and alleviate many of these fundamental inefficiencies by learning from the extensive research literature in the HPC community. In this article, we present TSets, the dataflow abstraction of Twister2, which is a big data framework designed for high‐performance dataflow and iterative computations. We discuss the dataflow model adopted by TSets and the rationale behind implementing iteration handling at the worker level. Finally, we evaluate TSets to show the performance of the framework and the importance of the worker level iteration model. Pulasthi Wickramasinghe, Niranda Perera, Supun Kamburugamuve, Kannan Govindarajan, Vibhatha Abeykoon, Chathura Widanage, Ahmet Uyar, Gurhan Gunduz, Selahattin Akkas, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 10 |
| 2021 | HPTMT: Operator-Based Architecture for Scalable High-Performance Data-Intensive FrameworksabstractData-intensive applications impact many domains, and their steadily increasing size and complexity demands highperformance, highly usable environments. We integrate a set of ideas developed in various data science and data engineering frameworks. They employ a set of operators on specific data abstractions that include vectors, matrices, tensors, graphs, and tables. Our key concepts are inspired from systems like MPI, HPF (High-Performance Fortran), NumPy, Pandas, Spark, Modin, PyTorch, TensorFlow, RAPIDS(NVIDIA), and OneAPI (Intel). Further, it is crucial to support different languages in everyday use in the Big Data arena, including Python, R, C++, and Java. We note the importance of Apache Arrow and Parquet for enabling language agnostic high performance and interoperability. In this paper, we propose High-Performance Tensors, Matrices and Tables (HPTMT), an operator-based architecture for data-intensive applications, and identify the fundamental principles needed for performance and usability success. We illustrate these principles by a discussion of examples using our software environments, Cylon and Twister2 that embody HPTMT. Supun Kamburugamuve, Chathura Widanage, Niranda Perera, Vibhatha Abeykoon, Ahmet Uyar, Thejaka Amila Kanewala, Gregor von Laszewski, Geoffrey C. Fox |
CLOUD | 8 |
| 2021 | Using Cloudmesh GAS for Speedy Generation of Hybrid Multi-Cloud Auto Generated AI ServicesabstractToday’s problems require a plethora of analytics tasks to be conducted to tackle state-of-the-art computational challenges posed in society impacting many areas including health care, automotive, banking, natural language processing, image detection, and many more data analytics-related tasks. Sharing existing analytics functions allows reuse and reduces overall effort. However, integrating deployment frameworks in the age of cloud computing are often out of reach for domain experts. Simple frameworks are needed that allow even non-experts to deploy and host services in the cloud. To avoid vendor lock-in, we require a generalized composable analytics service framework that allows users to integrate their services and those offered in clouds, not only by one, but by many cloud compute and service providers.We report on work that we conducted to provide a service integration framework for composing generalized analytics frame-works on multi-cloud providers that we call our Generalized AI Service (GAS) Generator. We demonstrate the framework’s usability by showcasing useful analytics workflows on various cloud providers, including AWS, Azure, and Google, and edge computing IoT devices. The examples are based on Scikit learn so they can be used in educational settings, replicated, and expanded upon. Benchmarks are used to compare the different services and showcase general replicability. Gregor von Laszewski, Anthony Orlowski, Richard H. Otten, Reilly Markowitz, Sunny Gandhi, Adam Chai, Geoffrey C. Fox, Wo L. Chang |
COMPSAC | 7 |
| 2021 | Spatiotemporal Pattern Mining for Nowcasting Extreme Earthquakes in Southern CaliforniaabstractGeoscience and seismology have utilized the most advanced technologies and equipment to monitor seismic events globally from the past few decades. With the enormous amount of data, modern GPU-powered deep learning presents a promising approach to analyze data and discover patterns. In recent years, there are plenty of successful deep learning models for picking seismic waves. However, forecasting extreme earthquakes, which can cause disasters, is still an underdeveloped topic in history. Relevant research in spatiotemporal dynamics mining and forecasting has revealed some successful predictions, a crucial topic in many scientific research fields. Most studies of them have many successful applications of using deep neural networks. In Geology and Earth science studies, earthquake prediction is one of the world’s most challenging problems, about which cutting-edge deep learning technologies may help discover some valuable patterns. In this project, we propose a deep learning modeling approach, namely EQPred, to mine spatiotemporal patterns from data to nowcast extreme earthquakes by discovering visual dynamics in regional coarse-grained spatial grids over time. In this modeling approach, we use synthetic deep learning neural networks with domain knowledge in geoscience and seismology to exploit earthquake patterns for prediction using convolutional long short-term memory neural networks. Our experiments show a strong correlation between location prediction and magnitude prediction for earthquakes in Southern California. Ablation studies and visualization validate the effectiveness of the proposed modeling method. Geoffrey C. Fox |
e-Science | 2 |
| 2021 | CRYPTOGRU: Low Latency Privacy-Preserving Text Analysis With GRUabstractHomomorphic encryption (HE) and garbled circuit (GC) provide the protection for users' privacy.However, simply mixing the HE and GC in RNN models suffer from long inference latency due to slow activation functions.In this paper, we present a novel hybrid structure of HE and GC gated recurrent unit (GRU) network, CRYPTOGRU, for low-latency secure inferences.CRYPTOGRU replaces computationally expensive GC-based tanh with fast GC-based ReLU , and then quantizes sigmoid and ReLU to smaller bit-length to accelerate activations in a GRU.We evaluate CRYP-TOGRU with multiple GRU models trained on 4 public datasets.Experimental results show CRYPTOGRU achieves top-notch accuracy and improves the secure inference latency by up to 138× over one of the state-of-the-art secure networks on the Penn Treebank dataset. Qian Lou, Lei Jiang 0001, Geoffrey C. Fox |
EMNLP (1) | 4 |
| 2021 | Deep Tiered Image Segmentation for Detecting Internal ICE Layers in Radar ImageryabstractUnderstanding the structure of Earth’s polar ice sheets is important for modeling how global warming will impact polar ice and, in turn, the Earth’s climate. Ground-penetrating radar is able to collect observations of the internal structure of snow and ice, but the process of manually labeling these observations is slow and laborious. Recent work has developed automatic techniques for finding the boundaries between the ice and the bedrock, but finding internal layers – the subtle boundaries that indicate where one year’s ice accumulation ended and the next began – is much more challenging because the number of layers varies and the boundaries often merge and split. In this paper, we propose a novel deep neural network for solving a general class of tiered segmentation problems. We then apply it to detecting internal layers in polar ice, evaluating on a large-scale dataset of polar ice radar data with human-labeled annotations as ground truth. John Paden, Lora Koenig, Geoffrey C. Fox, David Crandall |
ICME | 5 |
| 2020 | A Fast, Scalable, Universal Approach For Distributed Data AggregationsabstractIn the current era of Big Data, data engineering has transformed into an essential field of study across many branches of science. Advancements in Artificial Intelligence (AI) have broadened the scope of data engineering and opened up new applications in both enterprise and research communities. Aggregations (also termed reduce in functional programming) are an integral functionality in these applications. They are traditionally aimed at generating meaningful information on large data-sets, and today, they are being used for engineering more effective features for complex AI models. Aggregations are usually carried out on top of data abstractions such as tables/ arrays and are combined with other operations such as grouping of values. There are frameworks that excel in the said domains individually. But, we believe that there is an essential requirement for a data analytics tool that can universally integrate with existing frameworks, and thereby increase the productivity and efficiency of the entire data analytics pipeline. Cylon endeavors to fulfill this void. In this paper, we present Cylon's fast and scalable aggregation operations implemented on top of a distributed in-memory table structure that universally integrates with existing frameworks. Niranda Perera, Vibhatha Abeykoon, Chathura Widanage, Supun Kamburugamuve, Thejaka Amila Kanewala, Pulasthi Wickramasinghe, Ahmet Uyar, Hasara Maithree, Damitha Lenadora, Geoffrey C. Fox |
IEEE BigData | 10 |
| 2020 | Taxonomic Classification of Objects with Convolutional Neural Networksabstractit is difficult to build a CNN model that can classify many classes at once. Therefore, this study does not want to make many classes recognizable at once using only one model but by taxonomic classification. This study suggests a method of dividing the large number of classes into different steps of each step using Taxonomic classification. We propose a method of classifying a large number of classes by dividing them into models for each step using taxonomic classification. Our method uses taxonomic classification to distribute the weights required for training and test step by step. This will save a lot of time than creating a one-level model. In addition, to detect objects in never trained categories, the result may come up to a certain step without retraining the model. This shows that part of the model can be recycled. In this study, we presented a way to distinguish large numbers of classes using taxonomic classification by using multiple datasets, such as PASCAL VOC2012, ILSVRC 2013 image data from ImageNet, and 102 Category Flower Dataset. SungRyeol Yang, Geoffrey C. Fox, Bokyoon Na |
IEEE BigData | 2 |
| 2020 | Glyph: Fast and Accurately Training Deep Neural Networks on Encrypted DataabstractBecause of the lack of expertise, to gain benefits from their data, average users have to upload their private data to cloud servers they may not trust. Due to legal or privacy constraints, most users are willing to contribute only their encrypted data, and lack interests or resources to join deep neural network (DNN) training in cloud. To train a DNN on encrypted data in a completely non-interactive way, a recent work proposes a fully homomorphic encryption (FHE)-based technique implementing all activations by \textit{Brakerski-Gentry-Vaikuntanathan} (BGV)-based lookup tables. However, such inefficient lookup-table-based activations significantly prolong private training latency of DNNs. In this paper, we propose, Glyph, an FHE-based technique to fast and accurately train DNNs on encrypted data by switching between TFHE (Fast Fully Homomorphic Encryption over the Torus) and BGV cryptosystems. Glyph uses logic-operation-friendly TFHE to implement nonlinear activations, while adopts vectorial-arithmetic-friendly BGV to perform multiply-accumulations (MACs). Glyph further applies transfer learning on DNN training to improve test accuracy and reduce the number of MACs between ciphertext and ciphertext in convolutional layers. Our experimental results show Glyph obtains state-of-the-art accuracy, and reduces training latency by 69%~99% over prior FHE-based privacy-preserving techniques on encrypted datasets. Qian Lou, Geoffrey C. Fox, Lei Jiang 0001 |
NeurIPS | 3 |
| 2020 | Twister2: Design of a big data toolkitabstractSummary Data‐driven applications are essential to handle the ever‐increasing volume, velocity, and veracity of data generated by sources such as the Web and Internet of Things (IoT) devices. Simultaneously, an event‐driven computational paradigm is emerging as the core of modern systems designed for database queries, data analytics, and on‐demand applications. Modern big data processing runtimes and asynchronous many task (AMT) systems from high performance computing (HPC) community have adopted dataflow event‐driven model. The services are increasingly moving to an event‐driven model in the form of Function as a Service (FaaS) to compose services. An event‐driven runtime designed for data processing consists of well‐understood components such as communication, scheduling, and fault tolerance. Different design choices adopted by these components determine the type of applications a system can support efficiently. We find that modern systems are limited to specific sets of applications because they have been designed with fixed choices that cannot be changed easily. In this paper, we present a loosely coupled component‐based design of a big data toolkit where each component can have different implementations to support various applications. Such a polymorphic design would allow services and data analytics to be integrated seamlessly and expand from edge to cloud to HPC environments. Supun Kamburugamuve, Kannan Govindarajan, Pulasthi Wickramasinghe, Vibhatha Abeykoon, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | Parallel performance of molecular dynamics trajectory analysisabstractSummary The performance of biomolecular molecular dynamics simulations has steadily increased on modern high‐performance computing resources but acceleration of the analysis of the output trajectories has lagged behind so that analyzing simulations is becoming a bottleneck. To close this gap, we studied the performance of trajectory analysis with message passing interface (MPI) parallelization and the PythonMDAnalysislibrary on three different Extreme Science and Engineering Discovery Environment (XSEDE) supercomputers where trajectories were read from a Lustre parallel file system. Strong scaling performance was impeded by stragglers, MPI processes that were slower than the typical process. Stragglers were less prevalent for compute‐bound workloads, thus pointing to file reading as a bottleneck for scaling. However, a more complicated picture emerged in which both the computation and the data ingestion exhibited close to ideal strong scaling behavior whereas stragglers were primarily caused by either large MPI communication costs or long times to open the single shared trajectory file. We improved overall strong scaling performance by either subfiling (splitting the trajectory into separate files) or MPI‐IO with parallel HDF5 trajectory files. The parallel HDF5 approach resulted in near ideal strong scaling on up to 384 cores (16 nodes), thus reducing trajectory analysis times by two orders of magnitude compared with the serial approach. Mahzad Khoshlessan, Ioannis Paraskevakos, Geoffrey C. Fox, Shantenu Jha, Oliver Beckstein |
Concurr. Comput. Pract. Exp. | 3 |
| 2020 | Research on the Architecture and its Implementation for Instrumentation and Measurement CloudabstractCloud computing has brought a new method of resource utilization and management. Nowadays some researchers are working on cloud-based instrumentation and measurement systems designated as Instrumentation and Measurement Clouds (IMCs). However, until now, no standard definition or detailed architecture with an implemented system for IMC has been presented. This paper adopts the philosophy of cloud computing and brings forward a relatively standard definition and a novel architecture for IMC. The architecture inherits many key features of cloud computing, such as service provision on demand, scalability and so on, for remote Instrumentation and Measurement (IM) resource utilization and management. In the architecture, instruments and sensors are virtualized into abstracted resources, and commonly used IM functions are wrapped into services. Users can use these resources and services on demand remotely. Platforms implemented under such architecture can reduce the investment for building IM systems greatly, enable remote sharing of IM resources, increase utilization efficiency of various resources, and facilitate processing and analyzing of Big Data from instruments and sensors. Practical systems with a typical application are implemented upon the architecture. Results demonstrate that the novel IMC architecture can provide a new effective and efficient framework for establishing IM systems. Hengjing He, Wei Zhao 0031, Songling Huang, Geoffrey C. Fox, Qing Wang 0014 |
IEEE Trans. Serv. Comput. | 4 |
| 2019 | Big Data Benchmarks of High-Performance Storage Systems on Commercial Bare Metal CloudsabstractBare metal servers are widely available on public clouds to provide direct access to hardware and the system configuration with high performance storage and network devices are well suited for big data applications. Highly-optimized server with additional CPU core count and dense storage may lead to better performance in certain workloads and to ensure responsiveness of deployed services. Recent work on Hadoop ecosystems has addressed the performance improvement of scale-up machines configured with SSD storage and increased network bandwidth. The paper evaluates big data processing on dedicated clusters and provides the performance analysis of NVMe devices and SSD block storage options available on Amazon, Google, Microsoft, and Oracle Clouds. We show the benchmark results along with the system performance tests as we want to demonstrate the compute resource requirements for large-scale applications. The system capacity and limits for the underlying servers are described along with the cost analysis of scaling workloads on these platforms. Hyungro Lee, Geoffrey C. Fox |
CLOUD | 2 |
| 2019 | Benchmarking Deep Learning for Time Series: Challenges and DirectionsabstractDeep learning for time series is an emerging area with close ties to industry, yet under represented in performance benchmarks for machine learning systems. In this paper, we present a landscape of deep learning applications applied to time series, and discuss the challenges and directions towards building a robust performance benchmark of deep learning workloads for time series data. Geoffrey C. Fox, Sergey Serebryakov, Ankur Mohan, Pawel M. Morkisz, Debojyoti Dutta |
IEEE BigData | 2 |
| 2019 | Performance Optimization on Model Synchronization in Parallel Stochastic Gradient Descent Based SVMabstractUnderstanding the bottlenecks in implementing stochastic gradient descent (SGD)-based distributed support vector machines (SVM) algorithm is important in training larger data sets. The communication time to do the model synchronization across the parallel processes is the main bottleneck that causes inefficiency in the training process. The model synchronization is directly affected by the mini-batch size of data processed before the global synchronization. In producing an efficient distributed model, the communication time in training model synchronization has to be as minimum as possible while retaining a high testing accuracy. The effect from model synchronization frequency over the convergence of the algorithm and accuracy of the generated model must be well understood to design an efficient distributed model. In this research, we identify the bottlenecks in model synchronization in parallel stochastic gradient descent (PSGD)-based SVM algorithm with respect to the training model synchronization frequency (MSF). Our research shows that by optimizing the MSF in the data sets that we used, a reduction of 98% in communication time can be gained (16x - 24x speed up) with respect to high-frequency model synchronization. The training model optimization discussed in this paper guarantees a higher accuracy than the sequential algorithm along with faster convergence. Vibhatha Abeykoon, Geoffrey C. Fox |
CCGRID | 2 |
| 2019 | Learning Everywhere: A Taxonomy for the Integration of Machine Learning and SimulationsabstractWe present a taxonomy of research on Machine Learning (ML) applied to enhance simulations together with a catalog of some activities. We cover eight patterns for the link of ML to the simulations or systems plus three algorithmic areas: particle dynamics, agent-based models and partial differential equations. The patterns are further divided into three action areas: Improving simulation with Configurations and Integration of Data, Learn Structure, Theory and Model for Simulation, and Learn to make Surrogates. Geoffrey C. Fox, Shantenu Jha |
eScience | 1 |
| 2019 | Understanding ML Driven HPC: Applications and InfrastructureabstractWe recently outlined the vision of "Learning Everywhere" which captures the possibility and impact of how learning methods and traditional HPC methods can be coupled together. A primary driver of such coupling is the promise that Machine Learning (ML) will give major performance improvements for traditional HPC simulations. Motivated by this potential, the ML around HPC class of integration is of particular significance. In a related follow-up paper, we provided an initial taxonomy for integrating learning around HPC methods. In this paper which is part of the Learning Everywhere series, we discuss ``how'' learning methods and HPC simulations are being integrated to enhance effective performance of computations. This paper describes several modes --- substitution, assimilation, and control, in which learning methods integrate with HPC simulations and provide representative applications in each mode. This paper discusses some open research questions and we hope will motivate and clear the ground for MLaroundHPC benchmarks. Shantenu Jha, Geoffrey C. Fox |
eScience | 2 |
| 2019 | Perspectives on High-Performance Computing in a Big Data WorldabstractHigh-Performance Computing (HPC) and Cyberinfrastructure have played a leadership role in computational science even since the start of the NSF computing centers program. Thirty years ago parallel computing was a centerpiece of computer science research. Naively Big Data surely requires HPC to be processed, and transformational Big Data technology such as Hadoop and Spark exploit parallelism to success. Nevertheless, the HPC community does not appear to be thriving as a leader in Data Science while parallel computing is no longer a centerpiece. Some reasons for this are the dominant presence of Industry in technology futures and the universal fascination with Artificial Intelligence and Machine Learning. Maybe the pendulum will swing back a bit, but I expect the "AI first" philosophy to dominate in the foreseeable future. Thus I describe a future where HPC thrives in collaboration with Industry and AI. In particular, I discuss the promise of MLforHPC (AI for systems) and HPCforML (systems for AI). Geoffrey C. Fox |
HPDC | 1 |
| 2019 | Advances in big data programming, system software and HPC convergence
Ching-Hsien Hsu, Geoffrey C. Fox, Geyong Min, Sugam Sharma |
J. Supercomput. | 2 |
| 2018 | Twister: Net - Communication Library for Big Data Processing in HPC and Cloud EnvironmentsabstractStreaming processing and batch data processing are the dominant forms of big data analytics today, with numerous systems such as Hadoop, Spark, and Heron designed to process the ever-increasing explosion of data. Generally, these systems are developed as single projects with aspects such as communication, task management, and data management integrated together. By contrast, we take a component-based approach to big data by developing the essential features of a big data system as independent components with polymorphic implementations to support different requirements. Consequently, we recognize the requirements of both dataflow used in popular Apache Systems and the Bulk Synchronous Processing communication style common in High-Performance Computing (HPC) for different applications. Message Passing Interface (MPI) implementations are dominant in HPC but there are no such standard libraries available for big data. Twister:Net is a stand-alone, highly optimized dataflow style parallel communication library which can be used by big data systems or advanced users. Twister:Net can work both in cloud environments using TCP or HPC environments using MPI implementations. This paper introduces Twister:Net and compares it with existing systems to highlight its design and performance. Supun Kamburugamuve, Pulasthi Wickramasinghe, Kannan Govindarajan, Ahmet Uyar, Gurhan Gunduz, Vibhatha Abeykoon, Geoffrey C. Fox |
IEEE CLOUD | 7 |
| 2018 | Evaluation of Production Serverless Computing EnvironmentsabstractServerless computing provides a small runtime container to execute lines of codes without infrastructure management which is similar to Platform as a Service (PaaS) but a functional level. Amazon started the event-driven compute named Lambda functions in 2014 with a 25 concurrent limitation, but it now supports at least a thousand of concurrent invocation to process event messages generated by resources like databases, storage and system logs. Other providers, i.e., Google, Microsoft, and IBM offer a dynamic scaling manager to handle parallel requests of stateless functions in which additional containers are provisioning on new compute nodes for distribution. However, while functions are often developed for microservices and lightweight workload, they are associated with distributed data processing using the concurrent invocations. We claim that the current serverless computing environments can support dynamic applications in parallel when a partitioned task is executable on a small function instance. We present results of throughput, network bandwidth, a file I/O and compute performance regarding the concurrent invocations. We deployed a series of functions for distributed data processing to address the elasticity and then demonstrated the differences between serverless computing and virtual machines for cost efficiency and resource utilization. Hyungro Lee, Kumar Satyam, Geoffrey C. Fox |
IEEE CLOUD | 3 |
| 2018 | Object Detection by a Super-Resolution Method and a Convolutional Neural NetworksabstractRecently with many blurless or slightly blurred images, convolutional neural networks classify objects with around 90 percent classification rates, even if there are variable sized images. However, small object regions or cropping of images make object detection or classification difficult and decreases the detection rates. In many methods related to convolutional neural network (CNN), Bilinear or Bicubic algorithms are popularly used to interpolate region of interests. To overcome the limitations of these algorithms, we introduce a super-resolution method applied to the cropped regions or candidates, and this leads to improve recognition rates for object detection and classification. Large object candidates comparable in size of the full image have good results for object detections using many popular conventional methods. However, for smaller region candidates, using our super-resolution preprocessing and region candidates, allows a CNN to outperform conventional methods in the number of detected objects when tested on the VOC2007 and MSO datasets. Bokyoon Na, Geoffrey C. Fox |
IEEE BigData | 2 |
| 2018 | Task-parallel Analysis of Molecular Dynamics TrajectoriesabstractDifferent parallel frameworks for implementing data analysis applications have been proposed by the HPC and Big Data communities. In this paper, we investigate three task-parallel frameworks: Spark, Dask and RADICAL-Pilot with respect to their ability to support data analytics on HPC resources and compare them to MPI. We investigate the data analysis requirements of Molecular Dynamics (MD) simulations which are significant consumers of supercomputing cycles, producing immense amounts of data. A typical large-scale MD simulation of a physical system of O(100k) atoms over μsecs can produce from O(10) GB to O(1000) GBs of data. We propose and evaluate different approaches for parallelization of a representative set of MD trajectory analysis algorithms, in particular the computation of path similarity and leaflet identification. We evaluate Spark, Dask and RADICAL-Pilot with respect to their abstractions and runtime engine capabilities to support these algorithms. We provide a conceptual basis for comparing and understanding different frameworks that enable users to select the optimal system for each application. We also provide a quantitative performance analysis of the different algorithms across the three frameworks. Ioannis Paraskevakos, André Luckow, Mahzad Khoshlessan, George Chantzialexiou, Thomas E. Cheatham, Oliver Beckstein, Geoffrey C. Fox, Shantenu Jha |
ICPP | 7 |
| 2018 | Automated Tracking of 2D and 3D Ice Radar Imagery Using Viterbi and TRW-SabstractWe present improvements to existing implementations of the Viterbi and TRW-S algorithms applied to ice-bottom layer tracking on 2D and 3D radar imagery, respectively. Along with an explanation of our modifications and the reasoning behind them, we present a comparison between our results, the results obtained with the original implementations, and those obtained with other proposed methods of performing ice-bottom layer tracking. Victor Berger, Shane Chu, David Crandall, John Paden, Geoffrey C. Fox |
IGARSS | 6 |
| 2018 | Deep Hybrid Wavelet Network for Ice Boundary Detection in Radra ImageryabstractThis paper proposes a deep convolutional neural network approach to detect Ice surface and bottom layers from radar imagery. Radar images are capable to penetrate the Ice surface and provide us with valuable information from the underlying layers of ice surface. In recent years, deep hierarchical learning techniques for object detection and segmentation greatly improved the performance of traditional techniques based on hand-crafted feature engineering. We designed a deep convolutional network to produce the images of surface and bottom ice boundary. Our network take advantage of undecimated wavelet transform to provide the higest level of information from radar images, as well as multilayer and multi-scale optimized architecture. In this work, radar images from 2009-2016 NASA Operation IceBridge Mission are used to train and test the network. Our network outperformed the state-of-the art accuracy. Hamid Kamangir, Maryam Rahnemoonfar, Dugan Dobbs, John Paden, Geoffrey C. Fox |
IGARSS | 5 |
| 2018 | Multi-task Spatiotemporal Neural Networks for Structured Surface ReconstructionabstractDeep learning methods have surpassed the performance of traditional techniques on a wide range of problems in computer vision, but nearly all of this work has studied consumer photos, where precisely correct output is often not critical. It is less clear how well these techniques may apply on structured prediction problems where fine-grained output with high precision is required, such as in scientific imaging domains. Here we consider the problem of segmenting echogram radar data collected from the polar ice sheets, which is challenging because segmentation boundaries are often very weak and there is a high degree of noise. We propose a multi-task spatiotemporal neural network that combines 3D ConvNets and Recurrent Neural Networks (RNNs) to estimate ice surface boundaries from sequences of tomographic radar images. We show that our model outperforms the state-of-the-art on this problem by (1) avoiding the need for hand-tuned parameters, (2) extracting multiple surfaces (ice-air and ice-bed) simultaneously, (3) requiring less non-visual metadata, and (4) being about 6 times faster. Chenyou Fan, John Paden, Geoffrey C. Fox, David Crandall |
WACV | 4 |
| 2017 | Conceptualizing a Computing Platform for Science Beyond 2020: To Cloudify HPC, or HPCify Clouds?abstractA primary challenge of the cyberinfrastructure research community is the need to define the Platforms for Science beyond 2020. We analyze major current trends and propose that in order to deliver the Platform for Science in 2020 the dominant research challenge is to manage the convergence of capabilities of traditional HPC systems with richness of Apache Big Data systems. In this vision paper, we purport to examine the relationship between infrastructure for data-intensive computing and that for High Performance Computing and examine possible "convergence" of capabilities. Geoffrey C. Fox, Shantenu Jha |
CLOUD | 1 |
| 2017 | Efficient Software Defined Systems Using Common Core ComponentsabstractWith advent of Docker containers, an application deployment using container images gains popularity over scientific communities and major cloud providers to ease building reproducible environments. While a single base image can be imported multiple times from different containers to reduce storage consumption by a sharing technique, copy-on-write, duplicates of package dependencies are often observed over containers. In this paper, we propose new approaches to the container image management for eliminating duplicated dependencies. We create Common Core Components (3C) to share package dependencies by version control system commands, submodules and merge. 3C with submodules provides a collection of required libraries and tools in a separate branch, while keeping their base image same. 3C with merge offers a new base image including domain specific components thereby reducing duplicates in similar base images. Container images built with 3C enable efficient and compact software defined systems and disclose security information for tracking Common Vulnerability and Exposure (CVE). As a result, building application environments with 3C-enabled container images consumes less storage compared to existing Docker images. Dependency information for vulnerability is provided in detail for further developments. Hyungro Lee, Geoffrey C. Fox |
CLOUD | 2 |
| 2017 | Automatic estimation of ice bottom surfaces from radar imageryabstractGround-penetrating radar on planes and satellites now makes it practical to collect 3D observations of the subsurface structure of the polar ice sheets, providing crucial data for understanding and tracking global climate change. But converting these noisy readings into useful observations is generally done by hand, which is impractical at a continental scale. In this paper, we propose a computer vision-based technique for extracting 3D ice-bottom surfaces by viewing the task as an inference problem on a probabilistic graphical model. We first generate a seed surface subject to a set of constraints, and then incorporate additional sources of evidence to refine it via discrete energy minimization. We evaluate the performance of the tracking algorithm on 7 topographic sequences (each with over 3000 radar images) collected from the Canadian Arctic Archipelago with respect to human-labeled ground truth. David Crandall, Geoffrey C. Fox, John Paden |
ICIP | 3 |
| 2017 | DEM extraction of the basal topography of the Canadian archipelago ICE caps via 2D automated layer-trackerabstractThe basal topography of most of the glaciers that drain the ice caps of the Canadian Arctic Archipelago is largely unknown. To measure the basal topography, NASA Operation IceBridge flew a radar depth sounder in a wide swath mode with three transmit beams to image the glacier beds during three flights over the archipelago in 2014. We describe the measurement setup of the radar system, the algorithms used to process the data to produce a 3D image of the glacier bed, show digital elevation model (DEM) results of the beds, and provide a basic assessment of the tracking algorithm used to extract the DEM. Mohanad Al-Ibadi, Jordan Sprick, Sravya Athinarapu, Theresa Stumpf, John Paden, Carlton J. Leuschen, Fernando Rodriguez-Morales, David Crandall, Geoffrey C. Fox, David Burgess, Martin Sharp, Luke Copland, Wesley Van Wychen |
IGARSS | 10 |
| 2017 | Automatic Ice thickness estimation in radar imagery based on charged particles conceptabstractAccelerated loss of ice from Greenland and Antarctica has been observed in recent decades. Ice thickness is a key factor in making predictions about the future of massive ice reservoirs and can be estimated by calculating the exact location of the ice surface and bottom in radar imagery. Identifying the locations of ice boundaries is typically performed manually which is a very time consuming procedure. Here we propose a novel approach which automatically detects the complex topology of ice surface and bottom boundaries based on charged particle concept. Here we first applied anisotropic diffusion to remove the noise and enhance the image. At the second step, we detected the contours in the image based on Coulomb's electrostatic law and the assumption that each pixel is an electrically charged particle. The final ice surface and bottom are detected based on the projection profile of the contours. The results are evaluated on a large dataset of airborne radar imagery collected during IceBridge mission over Antarctica and show promising results with respect to hand-labeled ground truth. Maryam Rahnemoonfar, Amin Abbasi Habashi, John Paden, Geoffrey C. Fox |
IGARSS | 4 |
| 2017 | Java Technologies for Real-Time and Embedded Systems (JTRES2013)abstractThis outline describes a special issue of papers from the 2013 workshop on Java Technologies for Real-Time and Embedded Systems 1. There are 2 papers in this special issue. The first paper 2 discusses software locking mechanisms that commonly protect shared resources for multithreaded applications. This mechanism can, especially in chip-multiprocessor systems, result in a large synchronization overhead. For real-time systems in particular, this overhead increases the worst-case execution time and may void a task set's schedulability. This paper presents 2 hardware locking mechanisms to reduce the worst-case time required to acquire and release synchronization locks. These solutions are implemented for the chip-multiprocessor version of the Java Optimized Processor. The 2 hardware locking mechanisms are compared with a software locking solution as well as the original locking system of the processor. The hardware cost and performance are evaluated for all presented locking mechanisms. The performance of the better performing hardware locks is comparable to the original single global lock when contending for the same lock. When several noncontending locks are used, the hardware locks enable true concurrency for critical sections. Benchmarks show that using the hardware locks yields performance ranging from no worse than the original locks to more than twice their best performance. This improvement can allow a larger number of real-time tasks to be reliably scheduled on a multiprocessor real-time platform. Safety Critical Java (SCJ) is a profile of the Real-Time Specification for Java that brings to the safety-critical industry the possibility of using Java. Safety Critical Java defines 3 compliance levels: Level 0, Level 1, and Level 2. The SCJ specification is clear on what constitutes a Level 2 application in its use of the defined API but not the occasions on which it should be used. The second paper 3 broadly classifies the features that are only available at Level 2 into 3 groups: nested mission sequencers, managed threads, and global scheduling across multiple processors. It then explores the first 2 groups to elicit programming requirements that they support. The paper identifies several areas where the SCJ specification needs modifications to support these requirements fully; these include support for terminating managed threads, the ability to set a deadline on the transition between missions, and augmentation of the mission sequencer concept to support composability of timing constraints. The paper also proposes simplifications to the termination protocol of missions and their mission sequencers. To illustrate the benefit of our changes, the paper presents excerpts from a formal model of SCJ Level 2 written in Circus, a state-rich process algebra for refinement. We thank Kelvin Nilsen and Fridtjof Siebert for their work on this special issue. Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Special issue on 12th international workshop on Java technologies for real-time and embedded systems (JTRES2014)abstractSpecial issue on 12th international workshop on Java technologies for real-time and embedded systems (JTRES2014) Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Automatic Ice Surface and Bottom Boundaries Estimation in Radar Imagery Based on Level-Set ApproachabstractAccelerated loss of ice from Greenland and Antarctica has been observed in recent decades. The melting of polar ice sheets and mountain glaciers has considerable influence on sea level rise in a changing climate. Ice thickness is a key factor in making predictions about the future of massive ice reservoirs. The ice thickness can be estimated by calculating the exact location of the ice surface and subglacial topography beneath the ice in radar imagery. Identifying the locations of ice surface and bottom is typically performed manually, which is a very time-consuming procedure. Here, we propose an approach, which automatically detects ice surface and bottom boundaries using distance-regularized level-set evolution. In this approach, the complex topology of ice surface and bottom boundary layers can be detected simultaneously by evolving an initial curve in the radar imagery. Using a distance-regularized term, the regularity of the level-set function is intrinsically maintained, which solves the reinitialization issues arising from conventional level-set approaches. The results are evaluated on a large data set of airborne radar imagery collected during a NASA IceBridge mission over Antarctica and show promising results with respect to manually picked data. Maryam Rahnemoonfar, Geoffrey C. Fox, Masoud Yari, John Paden |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Java thread and process performance for parallel machine learning on multicore HPC clustersabstractThe growing use of Big Data frameworks on large machines highlights the importance of performance issues and the value of High Performance Computing (HPC) technology. This paper looks carefully at three major frameworks Spark, Flink and Message Passing Interface (MPI) both in scaling across nodes and internally over the many cores inside modern nodes. We focus on the special challenges of the Java Virtual Machine (JVM) using an Intel Haswell HPC cluster with 24 cores per node. Two parallel machine learning algorithms, K-Means clustering and Multidimensional Scaling (MDS) are used in our performance studies. We identify three major issues - thread models, affinity patterns, and communication mechanisms - as factors affecting performance by large factors and show how to optimize them so that Java can match the performance of traditional HPC languages like C. Further we suggest approaches that preserve the user interface and elegant dataflow approach of Flink and Spark but modify the runtime so that these Big Data frameworks can achieve excellent performance and realize the goals of HPC-Big Data convergence. Saliya Ekanayake, Supun Kamburugamuve, Pulasthi Wickramasinghe, Geoffrey C. Fox |
IEEE BigData | 4 |
| 2016 | TSmap3D: Browser visualization of high dimensional time series dataabstractLarge volumes of high dimensional time series data are increasingly becoming commonplace, and the ability to project such data into three dimensional space to visually inspect them is an important capability for scientific exploration. Algorithms such as Multidimensional Scaling (MDS) and Principal Component Analysis (PCA) can be used to reduce high dimensional data into a lower dimensional space. The time sensitive nature of such data requires continuous processing in time windows and visualizations to be shown as moving plots. In this paper we present: 1. an MDS-based approach to project high dimensional time series data to 3D with automatic transformation to align successive data segments; 2. an open source commodity visualization of three-dimensional time series in web browser based on Three.js; and 3. An example based on stock market data. The paper discusses various options available when producing the visualizations and how one optimizes the heuristic methods based on experimental results. Supun Kamburugamuve, Pulasthi Wickramasinghe, Saliya Ekanayake, Chathuri Wimalasena, Milinda Pathirage, Geoffrey C. Fox |
IEEE BigData | 6 |
| 2016 | MGEScan: a Galaxy-based system for identifying retrotransposons in genomesabstractUNLABELLED: : MGEScan-long terminal repeat (LTR) and MGEScan-non-LTR are successfully used programs for identifying LTRs and non-LTR retrotransposons in eukaryotic genome sequences. However, these programs are not supported by easy-to-use interfaces nor well suited for data visualization in general data formats. Here, we present MGEScan, a user-friendly system that combines these two programs with a Galaxy workflow system accelerated with MPI and Python threading on compute clusters. MGEScan and Galaxy empower researchers to identify transposable elements in a graphical user interface with ready-to-use workflows. MGEScan also visualizes the custom annotation tracks for mobile genetic elements in public genome browsers. A maximum speed-up of 3.26× is attained for execution time using concurrent processing and MPI on four virtual cores. MGEScan provides four operational modes: as a command line tool, as a Galaxy Toolshed, on a Galaxy-based web server, and on a virtual cluster on the Amazon cloud. AVAILABILITY AND IMPLEMENTATION: MGEScan tutorials and source code are available at http://mgescan.readthedocs.org/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hyungro Lee, Minsu Lee 0001, Wazim Mohammed Ismail, Mina Rho, Geoffrey C. Fox, Sangyoon Oh 0001, Haixu Tang |
Bioinform. | 5 |
| 2016 | A novel digital information service for federating distributed digital entities
Ahmet Fatih Mustacoglu, Geoffrey C. Fox |
Inf. Syst. | 2 |
| 2016 | Energy-efficient multisite offloading policy using Markov decision process for mobile cloud computing
Mati B. Terefe, Heezin Lee, Nojung Heo, Geoffrey C. Fox, Sangyoon Oh 0001 |
Pervasive Mob. Comput. | 4 |
| 2015 | HPC-ABDS High Performance Computing Enhanced Apache Big Data StackabstractWe review the High Performance Computing Enhanced Apache Big Data Stack HPC-ABDS and summarize the capabilities in 21 identified architecture layers. These cover Message and Data Protocols, Distributed Coordination, Security & Privacy, Monitoring, Infrastructure Management, DevOps, Interoperability, File Systems, Cluster & Resource management, Data Transport, File management, NoSQL, SQL (NewSQL), Extraction Tools, Object-relational mapping, In-memory caching and databases, Inter-process Communication, Batch Programming model and Runtime, Stream Processing, High-level Programming, Application Hosting and PaaS, Libraries and Applications, Workflow and Orchestration. We summarize status of these layers focusing on issues of importance for data analytics. We highlight areas where HPC and ABDS have good opportunities for integration. Geoffrey C. Fox, Judy Qiu, Supun Kamburugamuve, Shantenu Jha, André Luckow |
CCGRID | 1 |
| 2015 | Data Science and Online EducationabstractWe discuss the Data Science program at Indiana University, which is offered in both traditional residential and online formats. We describe Data Science, our chosen curriculum and its motivation. We describe experience in online delivery for both traditional lectures and online programming laboratories, and discuss implications for the technology used. Geoffrey C. Fox, Siddharth Maini, Howard Rosenbaum, David J. Wild 0001 |
CloudCom | 1 |
| 2015 | Peer Comparison of XSEDE and NCAR Publication DataabstractWe present a framework that compares the publication impact based on a comprehensive peer analysis of papers produced by scientists using XSEDE and NCAR resources. The analysis is introducing a percentile ranking based approach of citations of the XSEDE and NCAR papers compared to peer publications in the same journal that do not use these resources. This analysis is unique in that it evaluates the impact of the two facilities by comparing the reported publications from them to their peers from within the same journal issue. From this analysis, we can see that papers that utilize XSEDE and NCAR resources are cited statistically significantly more often. Hence we find that reported publications indicate that XSEDE and NCAR resources exert a strong positive impact on scientific research. Gregor von Laszewski, Fugang Wang, Geoffrey C. Fox, David L. Hart, Thomas R. Furlani, Robert L. DeLeon, Steven M. Gallo |
CLUSTER | 3 |
| 2015 | Panel on Cloud and Internet-of-ThingsabstractThe Internet of Things broadly interpreted covers everything from monitoring sensors, smartphones that today have 10 "things" each, robots and surveillance systems. The smartphones capture both the Internet access for social media sites with 1.8 billion photos uploaded every day and the content of tweets and Facebook posts that are being analyzed to capture in real-time the sentiment and thoughts of people. There are many estimates for the potential size of the IoT with at least 20 Billion devices expected by 2020. As well as the consumer IoT there is also the Industrial Internet of Things IIoT delivering intelligent machines and revolutionary industrial systems (e.g. manufacturing and transportation) of every type. The Cloud is often viewed as the natural controller for IoT devices and new software models ("Map-Streaming") like Apache Storm are emerging. The panel will take a broad look at the future of IoT covering devices and their cloud support. Geoffrey C. Fox |
IC2E | 1 |
| 2015 | Supporting High Performance Molecular Dynamics in Virtualized Clusters using IOMMU, SR-IOV, and GPUDirectabstractCloud Infrastructure-as-a-Service paradigms have recently shown their utility for a vast array of computational problems, ranging from advanced web service architectures to high throughput computing. However, many scientific computing applications have been slow to adapt to virtualized cloud frameworks. This is due to performance impacts of virtualization technologies, coupled with the lack of advanced hardware support necessary for running many high performance scientific applications at scale. Andrew J. Younge, John Paul Walters, Stephen P. Crago, Geoffrey C. Fox |
VEE | 4 |
| 2015 | Evaluating ARM HPC clusters for scientific workloadsabstractSummary The power consumption of modern high‐performance computing (HPC) systems that are built using power hungry commodity servers is one of the major hurdles for achieving Exascale computation. Several efforts have been made by the HPC community to encourage the use of low‐powered system‐on‐chip (SoC) embedded processors in large‐scale HPC systems. These initiatives have successfully demonstrated the use of ARM SoCs in HPC systems, but there is still a need to analyze the viability of these systems for HPC platforms before a case can be made for Exascale computation. The major shortcomings of current ARM‐HPC evaluations include a lack of detailed insights about performance levels on distributed multicore systems and performance levels for benchmarking in large‐scale applications running on HPC. In this paper, we present a comprehensive evaluation of results that covers major aspects of server and HPC benchmarking for ARM‐based SoCs. For the experiments, we built an unconventional cluster of ARM Cortex‐A9s that is referred to as Weiser and ran single‐node benchmarks (STREAM, Sysbench, and PARSEC) and multi‐node scientific benchmarks (High‐performance Linpack (HPL), NASA Advanced Supercomputing (NAS) Parallel Benchmark, and Gadget‐2) in order to provide a baseline for performance limitations of the system. Based on the experimental results, we claim that the performance of ARM SoCs depends heavily on the memory bandwidth, network latency, application class, workload type, and support for compiler optimizations. During server‐based benchmarking, we observed that when performing memory intensive benchmarks for database transactions, x86 performed 12% better for multithreaded query processing. However, ARM performed four times better for performance to power ratios for a single core and 2.6 times better on four cores. We noticed that emulated double precision floating point in Java resulted in three to four times slower performance as compared with the performance in C for CPU‐bound benchmarks. Even though Intel x86 performed slightly better in computation‐oriented applications, ARM showed better scalability in I/O bound applications for shared memory benchmarks. We incorporated the support for ARM in the MPJ‐Express runtime and performed comparative analysis of two widely used message passing libraries. We obtained similar results for network bandwidth, large‐scale application scaling, floating‐point performance, and energy‐efficiency for clusters in message passing evaluations (NBP and Gadget 2 with MPJ‐Express and MPICH). Our findings can be used to evaluate the energy efficiency of ARM‐based clusters for server workloads and scientific workloads and to provide a guideline for building energy‐efficient HPC clusters. Copyright © 2015 John Wiley & Sons, Ltd. Maqbool Jahanzeb, Sangyoon Oh 0001, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 3 |
| 2014 | GPU Passthrough Performance: A Comparison of KVM, Xen, VMWare ESXi, and LXC for CUDA and OpenCL ApplicationsabstractAs more scientific workloads are moved into the cloud, the need for high performance accelerators increases. Accelerators such as GPUs offer improvements in both performance and power efficiency over traditional multi-core processors, however, their use in the cloud has been limited. Today, several common hypervisors support GPU passthrough, but their performance has not been systematically characterized. In this paper we show that low overhead GPU passthrough is achievable across 4 major hypervisors and two processor microarchitectures. We compare the performance of two generations of NVIDIA GPUs within the Xen, VMWare ESXi, and KVM hypervisors, and we also compare the performance to that of Linux Containers (LXC). We show that GPU passthrough to KVM achieves 98 -- 100\% of the base system's performance across two architectures, while Xen and VMWare achieve 96 -- 99\% of the base systems performance, respectively. In addition, we describe several valuable lessons learned through our analysis and share the advantages and disadvantages of each hypervisor/GPU passthrough solution. John Paul Walters, Andrew J. Younge, Dong-In Kang 0001, Ke Thia Yao, Mikyung Kang, Stephen P. Crago, Geoffrey C. Fox |
IEEE CLOUD | 7 |
| 2014 | Integration of Clustering and Multidimensional Scaling to Determine Phylogenetic Trees as Spherical Phylograms Visualized in 3 DimensionsabstractPhylogenetic analysis is commonly used to analyze genetic sequence data from fungal communities, while ordination and clustering techniques commonly are used to analyze sequence data from bacterial communities. However, few studies have attempted to link these two independent approaches. In this paper, we propose a method, which we call spherical phylogram (SP), to display the phylogenetic tree within the clustering and visualization result from a pipeline called DACIDR. In comparison with traditional tree display methods, the correlations between the tree and the clustering can be observed directly. In addition, we propose an algorithm called interpolative joining (IJ) to construct and visualize the SP in 3D space. In the experiments, we used the sum of branch lengths to quantify the general fit between the clustering and the phylogenetic tree in SP and Mantel tests to determine how well the same grouping of sequences was preserved between the clustering and the SP. Our results show that DACIDR has a classification accuracy that is similar to a phylogenetic tree generated using a multiple sequence alignment, while having much lower computational cost. Yang Ruan 0001, Geoffrey L. House, Saliya Ekanayake, Ursel Schutte, James D. Bever, Haixu Tang, Geoffrey C. Fox |
CCGRID | 7 |
| 2014 | Advanced Virtualization Techniques for High Performance Cloud CyberinfrastructureabstractWith the advent of virtualization and Infrastructure-as-a-Service (IaaS), the broader scientific computing community is considering the use of clouds frothier scientific computing needs. This is due to the relative scalability, ease of use, advanced user environment customization abilities, and the many novel computing paradigms available for data-intensive applications. However, there is still a notable gap that exists between the performance of IaaS when compared to typical high performance computing (HPC) resources, limiting the applicability of IaaS for many potential users. This work proposes to bridge the gap between supercomputing and clouds using a few key aspects. First, we evaluate current hypervisors and their viability to run HPC workloads within current infrastructure. Next, we illustrate a mechanism to enable advanced accelerators such as GPUs in a Virtual Machine that can significantly enhance scientific computing problems. Furthermore, we are also able to support high speed, low latency inter-node communication through the use of Infini Band within virtual machines. Upon evaluating these newfound features and leveraging the system within the Open Stack environment, we illustrate that cloud computing can perform at near-native speeds and support a broad range of scientific computing problems as never before. Andrew J. Younge, Geoffrey C. Fox |
CCGRID | 2 |
| 2014 | CINET 2.0: A CyberInfrastructure for Network ScienceabstractAnalysis of structural properties and dynamics of networks is currently a central topic in many disciplines including Social Sciences, Biology and Business. CINET, a cyber infrastructure for such studies, introduced the concept of supporting network analysis as a service. The basic idea is to allow experts in various disciplines to focus on obtaining domain-specific insights from the results of network analyses instead of worrying about programming details and allocation of computational resources needed to carry out the analyses. A basic version of CINET was released in May 2012. This paper discusses CINET 2.0, a significantly enhanced version that supports complex network analyses through a web portal. CINET 2.0 has already been used for teaching courses related to Network Science at several US universities. In this paper, we discuss how CINET 2.0 significantly extends CINET 1.0 through enhancements to some components and the addition of new components. Sherif Hanie El Meligy Abdelhamid, Md. Maksudul Alam, Richard A. Aló, S. M. Arifuzzaman, Pete Beckman, Tirtha Bhattacharjee, Md Hasanuzzaman Bhuiyan, Keith R. Bisset, Stephen G. Eubank, Albert C. Esterline, Edward A. Fox, Geoffrey C. Fox, S. M. Shamimul Hasan, Harshal Hayatnagarkar, Maleq Khan, Chris J. Kuhlman, Madhav V. Marathe, Natarajan Meghanathan, Henning S. Mortveit, Judy Qiu, S. S. Ravi, Zalia Shams, Ongard Sirisaengtaksin, Samarth Swarup, Anil Vullikanti, Tak-Lon Wu |
eScience | 12 |
| 2014 | Estimating bedrock and surface layer boundaries and confidence intervals in ice sheet radar imagery using MCMCabstractClimate models that predict polar ice sheet behavior require accurate measurements of the bedrock-ice and ice-air boundaries in ground-penetrating radar imagery. Identifying these features is typically performed by hand, which can be tedious and error prone. We propose an approach for automatically estimating layer boundaries by viewing this task as a probabilistic inference problem. Our solution uses Markov-Chain Monte Carlo to sample from the joint distribution over all possible layers conditioned on an image. Layer boundaries can then be estimated from the expectation over this distribution, and confidence intervals can be estimated from the variance of the samples. We evaluate the method on 560 echograms collected in Antarctica, and compare to a state-of-the-art technique with respect to hand-labeled images. These experiments show an approximately 50% reduction in error for tracing both bedrock and surface layers. Stefan Lee, Jerome E. Mitchell, David Crandall, Geoffrey C. Fox |
ICIP | 4 |
| 2014 | Visualizing the Protein Sequence UniverseabstractSUMMARY Modern biology is experiencing a rapid increase in data volumes that challenges our analytical skills and existing cyberinfrastructure. Exponential expansion of the protein sequence universe (PSU), the protein sequence space, together with the costs and complexities of manual curation creates a major bottleneck in life sciences research. Existing resources lack scalable visualization tools that are instrumental for functional annotation. Here, we describe a new visualization tool using multidimensional scaling to create a 3D embedding of the protein space. The advantages of the proposed PSU method include the ability to scale to large numbers of sequences, integrate different similarity measures with other functional and experimental data, and facilitate protein annotation. We applied the method to visualize the prokaryotic PSU using sequence alignment scores. As an annotation example, we used the interpolation approach to map the set of annotated archaeal proteins into the prokaryotic PSU. Transdisciplinary approaches akin to the one described in this paper are urgently needed to quickly and efficiently translate the influx of new data into tangible innovations and groundbreaking discoveries. Copyright © 2013 John Wiley & Sons, Ltd. Larissa Stanberry, Roger Higdon, Winston Haynes, Natali Kolker, William Broomall, Saliya Ekanayake, Adam Hughes, Yang Ruan 0001, Judy Qiu, Eugene Kolker, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 11 |
| 2014 | Effective real-time scheduling algorithm for cyber physical systems society
Sanghyuk Park, Jai-Hoon Kim, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 3 |
| 2014 | Grey Forecast model for accurate recommendation in presence of data sparsity and correlation
Zhen Chen 0001, Jiaxing Shang, Geoffrey C. Fox |
Knowl. Based Syst. | 4 |
| 2014 | A parallel clustering method combined information bottleneck theory and centroid-based clustering
Geoffrey C. Fox, Weidong Gu |
J. Supercomput. | 2 |
| 2013 | Parallel deterministic annealing clustering and its application to LC-MS data analysisabstractWe present a scalable parallel deterministic annealing formalism for clustering with cutoffs and position-dependent variances. We apply it to the “peak matching" problem of the precise identification of the common LC-MS peaks across a cohort of multiple biological samples in proteomic biomarker discovery. We reliably and automatically find tens of thousands of clusters starting with a single one that is split recursively as distance resolution is sharpened. We parallelize the algorithm and compare unconstrained and trimmed clusters using data from a human tuberculosis cohort. Geoffrey C. Fox, Deepak R. Mani, Saumyadipta Pyne |
IEEE BigData | 1 |
| 2013 | Data Science, Clouds and X-Informatics
Geoffrey C. Fox |
CLOSER | 1 |
| 2013 | Co-processing SPMD computation on CPUs and GPUs clusterabstractHeterogeneous parallel systems with multi processors and accelerators are becoming ubiquitous due to better cost-performance and energy-efficiency. These heterogeneous processor architectures have different instruction sets and are optimized for either task-latency or throughput purposes. Challenges occur in regard to programmability and performance when running SPMD tasks on heterogeneous devices. In order to meet these challenges, we implemented a parallel runtime system that used to co-process SPMD computation on CPUs and GPUs clusters. Furthermore, we are proposing an analytic model to automatically schedule SPMD tasks on heterogeneous clusters. Our analytic model is derived from the roofline model, and therefore it can be applied to a wider range of SPMD applications and hardware devices. The experimental results of the C-means, GMM, and GEMV show good speedup in practical heterogeneous cluster environments. Geoffrey C. Fox, Gregor von Laszewski, Arun Chauhan 0001 |
CLUSTER | 2 |
| 2013 | A Robust and Scalable Solution for Interpolative Multidimensional Scaling with WeightingabstractAdvances in modern bio-sequencing techniques have led to a proliferation of raw genomic data that enables an unprecedented opportunity for data mining. To analyze such large volume and high-dimensional scientific data, many high performance dimension reduction and clustering algorithms have been developed. Among the known algorithms, we use Multidimensional Scaling (MDS) to reduce the dimension of original data and Pair wise Clustering, and to classify the data. We have shown that interpolative MDS, which is an online technique for real-time streaming in Big Data, can be applied to get better performance on massive data. However, SMACOF MDS approach is only directly applicable to cases where all pair wise distances are used and where weight is one for each term. In this paper, we proposed a robust and scalable MDS and interpolation algorithm using Deterministic Annealing technique, to solve problems with either missing distances or a non-trivial weight function. We compared our method to three state-of-art techniques. By experimenting on three common types of bioinformatics dataset, the results illustrate that the precision of our algorithms are better than other algorithms, and the weighted solutions has a lower computational time cost as well. Yang Ruan 0001, Geoffrey C. Fox |
e-Science | 2 |
| 2013 | A semi-automatic approach for estimating near surface internal layers from snow radar imageryabstractThe near surface layer signatures in polar firn are preserved from the glaciological behaviors of past climate and are important to understanding the rapidly changing polar ice sheets. Identifying and tracing near surface internal layers in snow radar echograms can be used to produce high-resolution accumulation maps. This process is typically performed manually, which requires time-consuming, dense hand-selection and interpolation between sections, for each echogram. We have developed an approach for semi-automatically estimating near surface internal layers and have applied it to snow radar echograms acquired from Antarctica. Our solution utilizes an active contour (“snakes”) model to find high-intensity edges likely to correspond to layer boundaries, while simultaneously imposing constraints on smoothness of layer depth and parallelism among layers. Jerome E. Mitchell, David Crandall, Geoffrey C. Fox, John Paden |
IGARSS | 3 |
| 2013 | Toward a modular and efficient distribution for Web service handlersabstractSUMMARY Over the last few decades, distributed systems have architecturally evolved. One recent evolutionary step is SOA. The SOA model is perfectly engendered in Web services, which provide software platforms for building applications as services. Web services utilize supportive capabilities such as security, reliability, and monitoring. These capabilities are typically provisioned as handlers, which incrementally add new features. Even though handlers are very important, the method of utilization is crucial for obtaining potential benefits. Every attempt to support a service with an additional handler increases the chance of an overwhelmingly crowded handler chain. Moreover, a handler may become a bottleneck because of its comparably higher processing time. In this paper, we present the Distributed Handler Architecture to provide an efficient, scalable, and modular architecture. The performance and scalability benchmarks show that the distributed and parallel handler executions are very promising for suitable handler configurations. The paper is concluded with remarks on the fundamentals of a promising computing environment for Web service handlers. Copyright © 2012 John Wiley & Sons, Ltd. Beytullah Yildiz, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Recent work in utility and cloud computing
Geoffrey C. Fox, Shrideep Pallickara |
Future Gener. Comput. Syst. | 1 |
| 2012 | Abstract Image Management and Universal Image Registration for Cloud and HPC InfrastructuresabstractCloud computing has become an important driver for delivering infrastructure as a service (IaaS) to users with on-demand requests for customized environments and sophisticated software stacks. Within the FutureGrid (FG) project, we offer different IaaS frameworks as well as high performance computing infrastructures by allowing users to explore them as part of the FG testbed. To ease the use of these infrastructures, as part of performance experiments, we have designed an image management framework, which allows us to create user defined software stacks based on abstract image management and uniform image registration. Consequently, users can create their own customized environments very easily. The complex processes of the underlying infrastructures are managed by our sophisticated software tools and services. Besides being able to manage images for IaaS frameworks, we also allow the registration and deployment of images onto bare-metal by the user. This level of functionality is typically not offered in a HPC (high performance computing) infrastructure. However, our approach provides users with the ability to create their own environments changing the paradigm of administrator-controlled dynamic provisioning to user-controlled dynamic provisioning, which we also call raining. Thus, users obtain access to a testbed with the ability to manage state-of-the-art software stacks that would otherwise not be supported in typical compute centers. Security is also considered by vetting images before they are registered in a infrastructure. In this paper, we present the design of our image management framework and evaluate two of its major components. This includes the image creation and image registration. Our design and implementation can support the current FG user community interested in such capabilities. Gregor von Laszewski, Fugang Wang, Geoffrey C. Fox |
IEEE CLOUD | 4 |
| 2012 | Comparison of Multiple Cloud FrameworksabstractToday, many cloud Infrastructure as a Service(IaaS) frameworks exist. Users, developers, and administrators have to make a decision about which environment is best suited for them. Unfortunately, the comparison of such frameworks is difficult because either users do not have access to all of them or they are comparing the performance of such systems on different resources, which make it difficult to obtain objective comparisons. Hence, the community benefits from the availability of a testbed on which comparisons between the IaaS frameworks can be conducted. FutureGrid aims to offer a number of IaaS including Nimbus, Eucalyptus, OpenStack, and OpenNebula. One of the important features that FutureGrid provides is not only the ability to compare between IaaS frameworks, but also to compare them in regards to bare-metal and traditional high performance computing services. In this paper, we outline some of our initial findings by providing such a testbed. As one of our conclusions, we also present our work on making access to the various infrastructures on FutureGrid easier. Gregor von Laszewski, Fugang Wang, Geoffrey C. Fox |
IEEE CLOUD | 4 |
| 2012 | Improving MapReduce Performance in Heterogeneous Network Environments and Resource UtilizationabstractMapReduce is a widely-used model for data parallel applications. We found its resource utilization is inefficient when there are not enough tasks to fill all task slots as the resources "reserved" for idle slots are just wasted. We propose resource stealing which enables running tasks to steal the unutilized resources and return them when new tasks are assigned. It exploits the opportunistic use of the otherwise wasted resources to improve overall resource utilization and reduce job execution time. Besides, our practical use of Hadoop shows the current mechanism adopted to trigger speculative execution creates many unnecessary speculative tasks that are killed soon after creation as the original tasks complete earlier. To alleviate the issue, we propose Benefit Aware Speculative Execution which predicts the benefit of running new speculative tasks and greatly eliminates unnecessary runs. Finally, MapReduce is mainly optimized for homogeneous environments and its inefficiency in heterogeneous network environments has been observed in our experiments. We investigate network heterogeneity aware scheduling of both map and reduce tasks. Overall, our goal is to enhance Hadoop to cope with significant network heterogeneity and improve resource utilization. Zhenhua Guo 0004, Geoffrey C. Fox |
CCGRID | 2 |
| 2012 | Investigation of Data Locality in MapReduceabstractTraditional HPC architectures separate compute nodes and storage nodes, which are interconnected with high speed links to satisfy data access requirements in multi-user environments. However, the capacity of those high speed links is still much less than the aggregate bandwidth of all compute nodes. In Data Parallel Systems such as GFS/MapReduce, clusters are built with commodity hardware and each node takes the roles of both computation and storage, which makes it possible to bring compute to data. Data locality is a significant advantage of data parallel systems over traditional HPC systems. Good data locality reduces cross-switch network traffic - one of the bottlenecks in data-intensive computing. In this paper, we investigate data locality in depth. Firstly, we build a mathematical model of scheduling in MapReduce and theoretically analyze the impact on data locality of configuration factors, such as the numbers of nodes and tasks. Secondly, we find the default Hadoop scheduling is non-optimal and propose an algorithm that schedules multiple tasks simultaneously rather than one by one to give optimal data locality. Thirdly, we run extensive experiments to quantify performance improvement of our proposed algorithms, measure how different factors impact data locality, and investigate how data locality influences job execution time in both single-cluster and cross-cluster environments. Zhenhua Guo 0004, Geoffrey C. Fox |
CCGRID | 2 |
| 2012 | Improving Resource Utilization in MapReduceabstractMapReduce has been adopted widely in both academia and industry to run large-scale data parallel applications. In MapReduce, each slave node hosts a number of task slots to which tasks can be assigned. So they limit the maximum number of tasks that can execute concurrently on each node. When all task slots of a node are not used, the resources “reserved” for idle slots are unutilized. To improve resource utilization, we propose resource stealing to enable running tasks to steal resources reserved for idle slots and give them back proportionally whenever new tasks are assigned. Resource stealing makes the otherwise wasted resources get fully utilized without interfering with normal job scheduling. MapReduce uses speculative execution to improve fault tolerance. Current Hadoop implementation decides whether to run speculative tasks based on the progress rates of running tasks, which does not take into consideration the absolute progress of each task. We propose Benefit Aware Speculative Execution which evaluates the potential benefit of speculative tasks and eliminates unnecessary runs. We implement the proposed algorithms in Hadoop, and our experiments show that our algorithms can significantly shorten job execution time and reduce the number of non-beneficial speculative tasks. Zhenhua Guo 0004, Geoffrey C. Fox, Yang Ruan 0001 |
CLUSTER | 2 |
| 2012 | CINET: A cyberinfrastructure for network scienceabstractNetworks are an effective abstraction for representing real systems. Consequently, network science is increasingly used in academia and industry to solve problems in many fields. Computations that determine structure properties and dynamical behaviors of networks are useful because they give insights into the characteristics of real systems. We introduce a newly built and deployed cyberinfrastructure for network science (CINET) that performs such computations, with the following features: (i) it offers realistic networks from the literature and various random and deterministic network generators; (ii) it provides many algorithmic modules and measures to study and characterize networks; (iii) it is designed for efficient execution of complex algorithms on distributed high performance computers so that they scale to large networks; and (iv) it is hosted with web interfaces so that those without direct access to high performance computing resources and those who are not computing experts can still reap the system benefits. It is a combination of application design and cyberinfrastructure that makes these features possible. To our knowledge, these capabilities collectively make CINET novel. We describe the system and illustrative use cases, with a focus on the CINET user. Sherif Elmeligy Abdelhamid, Richard A. Aló, S. M. Arifuzzaman, Pete Beckman, Md Hasanuzzaman Bhuiyan, Keith R. Bisset, Edward A. Fox, Geoffrey C. Fox, Kevin Hall, S. M. Shamimul Hasan, Anurodh Joshi, Maleq Khan, Chris J. Kuhlman, Spencer J. Lee, Jonathan Leidig, Hemanth Makkapati, Madhav V. Marathe, Henning S. Mortveit, Judy Qiu, S. S. Ravi, Zalia Shams, Ongard Sirisaengtaksin, Rajesh Subbiah, Samarth Swarup, Nick Trebon, Anil Vullikanti |
eScience | 8 |
| 2012 | Mining hidden mixture context with ADIOS-P to improve predictive pre-fetcher accuracyabstractPredictive pre-fetcher, which predicts future data access events and loads the data before users requests, has been widely studied, especially in file systems or web contents servers, to reduce data load latency. Especially in scientific data visualization, pre-fetching can reduce the IO waiting time. In order to increase the accuracy, we apply a data mining technique to extract hidden information. More specifically, we apply a data mining technique for discovering the hidden contexts in data access patterns and make prediction based on the inferred context to boost the accuracy. In particular, we performed Probabilistic Latent Semantic Analysis (PLSA), a mixture model based algorithm popular in the text mining area, to mine hidden contexts from the collected user access patterns and, then, we run a predictor within the discovered context. We further improve PLSA by applying the Deterministic Annealing (DA) method to overcome the local optimum problem. In this paper we demonstrate how we can apply PLSA and DA optimization to mine hidden contexts from users data access patterns and improve predictive pre-fetcher performance. Jong Choi 0001, Hasan Abbasi, David Pugmire, Norbert Podhorszki, Scott Klasky, Cristian Capdevila, Manish Parashar, Matthew Wolf, Judy Qiu, Geoffrey C. Fox |
eScience | 10 |
| 2012 | Layer-finding in radar echograms using probabilistic graphical models
David Crandall, Geoffrey C. Fox, John Paden |
ICPR | 2 |
| 2012 | Cyberinfrastructure for eScience and eBusiness from Clouds to Exascale
Geoffrey C. Fox |
SECRYPT | 1 |
| 2012 | Interpolative multidimensional scaling techniques for the identification of clusters in very large sequence setsabstractBACKGROUND: Modern pyrosequencing techniques make it possible to study complex bacterial populations, such as 16S rRNA, directly from environmental or clinical samples without the need for laboratory purification. Alignment of sequences across the resultant large data sets (100,000+ sequences) is of particular interest for the purpose of identifying potential gene clusters and families, but such analysis represents a daunting computational task. The aim of this work is the development of an efficient pipeline for the clustering of large sequence read sets. METHODS: Pairwise alignment techniques are used here to calculate genetic distances between sequence pairs. These methods are pleasingly parallel and have been shown to more accurately reflect accurate genetic distances in highly variable regions of rRNA genes than do traditional multiple sequence alignment (MSA) approaches. By utilizing Needleman-Wunsch (NW) pairwise alignment in conjunction with novel implementations of interpolative multidimensional scaling (MDS), we have developed an effective method for visualizing massive biosequence data sets and quickly identifying potential gene clusters. RESULTS: This study demonstrates the use of interpolative MDS to obtain clustering results that are qualitatively similar to those obtained through full MDS, but with substantial cost savings. In particular, the wall clock time required to cluster a set of 100,000 sequences has been reduced from seven hours to less than one hour through the use of interpolative MDS. CONCLUSIONS: Although work remains to be done in selecting the optimal training set size for interpolative MDS, substantial computational cost savings will allow us to cluster much larger sequence sets in the future. Adam Hughes, Yang Ruan 0001, Saliya Ekanayake, Seung-Hee Bae, Qunfeng Dong, Mina Rho, Judy Qiu, Geoffrey C. Fox |
BMC Bioinform. | 8 |
| 2012 | Advanced theory and practice for high performance computing and communicationsabstractThis collection of papers was selected from those presented at the International Conference on High Performance Computing and Communications HPCC-09 1 and the International Conference on Information Security and AssuranceISA-09 2, Seoul, Korea, June 25–27, 2009. The papers 3-9 were all enhanced over the conference versions and separately reviewed. The collection is called ATPHPCC or advanced theory and practice for high performance computing and communications. I would like to thank Omer Rana and Tai-hoon Kim for their help in putting this special issue together. Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | Enabling hierarchical dissemination of streams in content distribution networksabstractSUMMARY In streaming systems the content distribution network routes streams based on interests registered by the consuming entities. In hierarchical streaming, the dissemination is also predicated on the resolution of hierarchical dependencies between various streams. Entities specify explicit wildcards, in addition to the implicit ones in place, to further control the types of streams within a given hierarchy that should be routed to them. This paper presents an analysis and performance evaluation of three different algorithms for hierarchical streaming. In our evaluation of these algorithms we are especially interested in three factors: performance, ability to cope with flux, and memory consumption. Comprehensive benchmarks for these algorithms, in this paper, will enable system designers to harness the best algorithm that satisfies their hierarchical streaming requirements. Copyright © 2011 John Wiley & Sons, Ltd. Shrideep Pallickara, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | Analysis of Virtualization Technologies for High Performance Computing EnvironmentsabstractAs Cloud computing emerges as a dominant paradigm in distributed systems, it is important to fully understand the underlying technologies that make Clouds possible. One technology, and perhaps the most important, is virtualization. Recently virtualization, through the use of hyper visors, has become widely used and well understood by many. However, there are a large spread of different hyper visors, each with their own advantages and disadvantages. This paper provides an in-depth analysis of some of today's commonly accepted virtualization technologies from feature comparison to performance analysis, focusing on the applicability to High Performance Computing environments using Future Grid resources. The results indicate virtualization sometimes introduces slight performance impacts depending on the hyper visor type, however the benefits of such technologies are profound and not all virtualization technologies are equal. From our experience, the KVM hyper visor is the optimal choice for supporting HPC applications within a Cloud infrastructure. Andrew J. Younge, Robert Henschel, James T. Brown, Gregor von Laszewski, Judy Qiu, Geoffrey C. Fox |
IEEE CLOUD | 6 |
| 2011 | FutureGrid Image Repository: A Generic Catalog and Storage System for Heterogeneous Virtual Machine ImagesabstractFuture Grid (FG) is an experimental, high-performance test bed that supports HPC, cloud and grid computing experiments for both application and computer scientist. Future Grid includes the use of virtualization technology to allow the support of a wide range of operating systems in order to include a test bed for various cloud computing infrastructure as a service frameworks. Therefore, efficient management of a variety of virtual machine images becomes a key issue. Current cloud frameworks do not provide a way to manage images for different IaaS frameworks. They typically provide their own image repositories, but in general they do not allow us to store the needed metadata to handle other IaaS images. We present a generic catalog and image repository to store images of any type. Our image repository has a convenient interface that distinguishes image types. Therefore, it is not only useful for Future Grid, but also for any application that needs to manage images. Gregor von Laszewski, Fugang Wang, Andrew J. Younge, Geoffrey C. Fox |
CloudCom | 5 |
| 2011 | Automatic Task Re-organization in MapReduceabstractMapReduce is increasingly considered as a useful parallel programming model for large-scale data processing. It exploits parallelism among execution of primitive map and reduce operations. Hadoop is an open source implementation of MapReduce that has been used in both academic research and industry production. However, its implementation strategy that one map task processes one data block limits the degree of concurrency and degrades performance because of inability to fully utilize available resources. In addition, its assumption that task execution time in each phase does not vary much does not always hold, which makes speculative execution useless. In this paper, we present mechanisms to dynamically split and consolidate tasks to cope with load balancing and break through the concurrency limit resulting from fixed task granularity. For single-job systems, two algorithms are proposed for circumstances where prior knowledge is known and unknown. For multi-job cases, we propose a modified shortest-job-first strategy, which minimizes job turnaround time theoretically when combined with task splitting. We compared the effectiveness of our approach to the default task scheduling strategy using both synthesized and trace-based workloads. Simulation results show that our approach improves performance significantly. Zhenhua Guo 0004, Marlon E. Pierce, Geoffrey C. Fox |
CLUSTER | 3 |
| 2011 | Modeling, simulation, and practice of floor control for synchronous and ubiquitous collaboration
Kangseok Kim, Geoffrey C. Fox |
Multim. Tools Appl. | 2 |
| 2010 | High Performance Dimension Reduction and Visualization for Large High-Dimensional Data AnalysisabstractLarge high dimension datasets are of growing importance in many fields and it is important to be able to visualize them for understanding the results of data mining approaches or just for browsing them in a way that distance between points in visualization (2D or 3D) space tracks that in original high dimensional space. Dimension reduction is a well understood approach but can be very time and memory intensive for large problems. Here we report on parallel algorithms for Scaling by MAjorizing a Complicated Function (SMACOF) to solve Multidimensional Scaling problem and Generative Topographic Mapping (GTM). The former is particularly time consuming with complexity that grows as square of data set size but has advantage that it does not require explicit vectors for dataset points but just measurement of inter-point dissimilarities. We compare SMACOF and GTM on a subset of the NIH PubChem database which has binary vectors of length 166 bits. We find good parallel performance for both GTM and SMACOF and strong correlation between the dimension-reduced PubChem data from these two methods. Jong Choi 0001, Seung-Hee Bae, Xiaohong Qiu, Geoffrey C. Fox |
CCGRID | 4 |
| 2010 | Performance of Windows Multicore Systems on Threading and MPIabstractWe present performance results on a Windows cluster with up to 768 cores using MPI and two variants of threading - CCR and TPL. CCR (Concurrency and Coordination Runtime) presents a message based interface while TPL (Task Parallel Library) allows for loops to be automatically parallelized. MPI is used between the cluster nodes (up to 32) and either threading or MPI for parallelism on the 24 cores of each node. We use a simple matrix multiplication kernel as well as a significant bioinformatics gene clustering application. We find that the two threading models offer similar performance with MPI outperforming both at low levels of parallelism but threading much better when the grain size (problem size per process) is small. We find better performance on Intel compared to AMD on comparable 24 core systems. We develop simple models for the performance of the clustering code. Judy Qiu, Scott Beason, Seung-Hee Bae, Saliya Ekanayake, Geoffrey C. Fox |
CCGRID | 5 |
| 2010 | Building a Distributed Block Storage System for Cloud InfrastructureabstractThe development of cloud infrastructures has stimulated interest in virtualized block storage systems, exemplified by Amazon Elastic Block Store (EBS), Eucalyptus’ EBS implementation, and the Virtual Block Store (VBS) system. Compared with other solutions, VBS is designed for flexibility, and can be extended to support various Virtual Machine Managers and Cloud platforms. However, due to its single-volume-server architecture, VBS has the problem of single point of failure and low scalability. This paper presents our latest improvements to VBS for solving these problems, including a new distributed architecture based on the Lustre file system, new workflows, better reliability and scalability, and read-only volume sharing. We call this improved implementation VBS-Lustre. Preliminary tests show that VBS-Lustre can provide both better throughput and higher scalability in multiple attachment scenarios than VBS. VBS-Lustre could potentially be applied to solve some challenges for current cluster file systems, such as metadata management and small file access. Xiaoming Gao, Marlon E. Pierce, Mike Lowe, Geoffrey C. Fox |
CloudCom | 5 |
| 2010 | MapReduce in the Clouds for ScienceabstractThe utility computing model introduced by cloud computing combined with the rich set of cloud infrastructure services offers a very viable alternative to traditional servers and computing clusters. MapReduce distributed data processing architecture has become the weapon of choice for data-intensive analyses in the clouds and in commodity clusters due to its excellent fault tolerance features, scalability and the ease of use. Currently, there are several options for using MapReduce in cloud environments, such as using MapReduce as a service, setting up one's own MapReduce cluster on cloud instances, or using specialized cloud MapReduce runtimes that take advantage of cloud infrastructure services. In this paper, we introduce Azure MapReduce, a novel MapReduce runtime built using the Microsoft Azure cloud infrastructure services. Azure MapReduce architecture successfully leverages the high latency, eventually consistent, yet highly scalable Azure infrastructure services to provide an efficient, on demand alternative to traditional MapReduce clusters. Further we evaluate the use and performance of MapReduce frameworks, including Azure MapReduce, in cloud environments for scientific applications using sequence assembly and sequence alignment as use cases. Thilina Gunarathne, Tak-Lon Wu, Judy Qiu, Geoffrey C. Fox |
CloudCom | 4 |
| 2010 | Applying Twister to Scientific ApplicationsabstractMany scientific applications suffer from the lack of a unified approach to support the management and efficient processing of large-scale data. The Twister MapReduce Framework, which not only supports the traditional MapReduce programming model but also extends it by allowing iterations, addresses these problems. This paper describes how Twister is applied to several kinds of scientific applications such as BLAST, MDS Interpolation and GTM Interpolation in a non-iterative style and to MDS without interpolation in an iterative style. The results show the applicability of Twister to data parallel and EM algorithms with small overhead and increased efficiency. Bingjing Zhang, Yang Ruan 0001, Tak-Lon Wu, Judy Qiu, Adam Hughes, Geoffrey C. Fox |
CloudCom | 6 |
| 2010 | Multidimensional Scaling by Deterministic Annealing with Iterative Majorization AlgorithmabstractMultidimensional Scaling (MDS) is a dimension reduction method for information visualization, which is set up as a non-linear optimization problem. It is applicable to many data intensive scientific problems including studies of DNA sequences but tends to get trapped in local minima. Deterministic Annealing (DA) has been applied to many optimization problems to avoid local minima. We apply DA approach to MDS problem in this paper and show that our proposed DA approach improves the mapping quality and shows high reliability in a variety of experimental results. Further its execution time is similar to that of the un-annealed approach. We use different data sets for comparing the proposed DA approach with both a well known algorithm called SMACOF and a MDS with distance smoothing method which aims to avoid local optima. Our proposed DA method outperforms SMACOF algorithm and the distance smoothing MDS algorithm in terms of the mapping quality and shows much less sensitivity with respect to initial configurations and stopping condition. We also investigate various temperature cooling parameters for our deterministic annealing method within an exponential cooling scheme. Seung-Hee Bae, Judy Qiu, Geoffrey C. Fox |
eScience | 3 |
| 2010 | Practice-centered e-science: a practice turn perspective on cyberinfrastructure designabstractCyberinfrastructure is a rapidly growing area of global research and funding with a history of emphasizing the role technology will play in changing scientific work practices. This paper proposes a practice-theoretic perspective that is informative to cyberinfrastructure research and design. To illustrate the relevancy of a practice-theoretic perspective to cyberinfrastructure, this paper presents a critical review of 160 cyberinfrastructure research papers and reports published in the last decade through a perspective of embodied practice. After relating common cyberinfrastructure research themes through a focus on embodied practice, we propose a series of early implications for design aimed at incorporating the lessons of embodied practice into the design and development of future cyberinfrastructure applications. Tyler Pace, Shaowen Bardzell, Geoffrey C. Fox |
GROUP | 3 |
| 2010 | Dimension reduction and visualization of large high-dimensional data via interpolationabstractThe recent explosion of publicly available biology gene sequences and chemical compounds offers an unprecedented opportunity for data mining. To make data analysis feasible for such vast volume and high-dimensional scientific data, we apply high performance dimension reduction algorithms. It facilitates the investigation of unknown structures in a three dimensional visualization. Among the known dimension reduction algorithms, we utilize the multidimensional scaling and generative topographic mapping algorithms to configure the given high-dimensional data into the target dimension. However, both algorithms require large physical memory as well as computational resources. Thus, the authors propose an interpolated approach to utilizing the mapping of only a subset of the given data. This approach effectively reduces computational complexity. With minor trade-off of approximation, interpolation method makes it possible to process millions of data points with modest amounts of computation and memory requirement. Since huge amount of data are dealt, we represent how to parallelize proposed interpolation algorithms, as well. For the evaluation of the interpolated MDS by STRESS criteria, it is necessary to compute symmetric all pairwise computation with only subset of required data per process, so we also propose a simple but efficient parallel mechanism for the symmetric all pairwise computation when only a subset of data is available to each process. Our experimental results illustrate that the quality of interpolated mapping results are comparable to the mapping results of original algorithm only. In parallel performance aspect, those interpolation methods are well parallelized with high efficiency. With the proposed interpolation method, we construct a configuration of two-million out-of-sample data into the target dimension, and the number of out-of-sample data can be increased further. Seung-Hee Bae, Jong Choi 0001, Judy Qiu, Geoffrey C. Fox |
HPDC | 4 |
| 2010 | Browsing large scale cheminformatics data with dimension reductionabstractVisualization of large-scale high dimensional data tool is highly valuable for scientific discovery in many fields. We presentPubChemBrowse, acustomizedvisualizationtoolfor cheminformatics research. It provides a novel 3D data point browser that displays complex properties of massive data on commodity clients. As in GIS browsers for Earth and Environment data, chemical compounds with similar properties are nearby in the browser. PubChemBrowse is built around in-househighperformanceparallel MDS(Multi-Dimensional Scaling) and GTM (Generative Topographic Mapping) services andsupports fast interaction with anexternalproperty database. These properties can be overlaid on 3D mapped compound space or queried for individual points. We prototype use with Chem2Bio2RDF system using SPARQLquery language to access over 20 publicly accessible bioinformatics databases. We describe our design and implementation of the integrated PubChemBrowse application and outline its use in drug discovery. The same core technologies can be used to develop similar high dimensional browsers in other scientific areas. Jong Choi 0001, Seung-Hee Bae, Judy Qiu, Geoffrey C. Fox, Bin Chen 0002, David J. Wild 0001 |
HPDC | 4 |
| 2010 | Twister: a runtime for iterative MapReduceabstractMapReduce programming model has simplified the implementation of many data parallel applications. The simplicity of the programming model and the quality of services provided by many implementations of MapReduce attract a lot of enthusiasm among distributed computing communities. From the years of experience in applying MapReduce to various scientific applications we identified a set of extensions to the programming model and improvements to its architecture that will expand the applicability of MapReduce to more classes of applications. In this paper, we present the programming model and the architecture of Twister an enhanced MapReduce runtime that supports iterative MapReduce computations efficiently. We also show performance comparisons of Twister with other similar runtimes such as Hadoop and DryadLINQ for large scale data parallel applications. Jaliya Ekanayake, Bingjing Zhang, Thilina Gunarathne, Seung-Hee Bae, Judy Qiu, Geoffrey C. Fox |
HPDC | 7 |
| 2010 | Cloud computing paradigms for pleasingly parallel biomedical applicationsabstractCloud computing offers exciting new approaches for scientific computing that leverages the hardware and software investments on large scale data centers by major commercial players. Loosely coupled problems are very important in many scientific fields and are on the rise with the ongoing move towards data intensive computing. There exist several approaches to leverage clouds & cloud oriented data processing frameworks to perform pleasingly parallel computations. In this paper we present two pleasingly parallel biomedical applications, 1) assembly of genome fragments 2) dimension reduction in the analysis of chemical structures, implemented utilizing cloud infrastructure service based utility computing models of Amazon AWS and Microsoft Windows Azure as well as utilizing MapReduce based data processing frameworks, Apache Hadoop and Microsoft DryadLINQ. We review and compare each of the frameworks and perform a comparative study among them based on performance, efficiency, cost and the usability. Cloud service based utility computing model and the managed parallelism (MapReduce) exhibited comparable performance and efficiencies for the applications we considered. We analyze the variations in cost between the different platform choices (eg: EC2 instance types), highlighting the need to select the appropriate platform based on the nature of the computation. Thilina Gunarathne, Tak-Lon Wu, Judy Qiu, Geoffrey C. Fox |
HPDC | 4 |
| 2010 | Algorithms and application for grids and cloudsabstractWe discuss the impact of clouds and grid technology on scientific computing using examples from a variety of fields -- especially the life sciences. We cover the impact of the growing importance of data analysis and note that it is more suitable for these modern architectures than the large simulations (particle dynamics and partial differential equation solution) that are mainstream use of large scale "massively parallel" supercomputers. The importance of grids is seen in the support of distributed data collection and archiving while clouds are and will replace grids for the large scale analysis of the data. Geoffrey C. Fox |
SPAA | 1 |
| 2010 | Hybrid cloud and cluster computing paradigms for life science applicationsabstractBACKGROUND: Clouds and MapReduce have shown themselves to be a broadly useful approach to scientific computing especially for parallel data intensive applications. However they have limited applicability to some areas such as data mining because MapReduce has poor performance on problems with an iterative structure present in the linear algebra that underlies much data analysis. Such problems can be run efficiently on clusters using MPI leading to a hybrid cloud and cluster environment. This motivates the design and implementation of an open source Iterative MapReduce system Twister. RESULTS: Comparisons of Amazon, Azure, and traditional Linux and Windows environments on common applications have shown encouraging performance and usability comparisons in several important non iterative cases. These are linked to MPI applications for final stages of the data analysis. Further we have released the open source Twister Iterative MapReduce and benchmarked it against basic MapReduce (Hadoop) and MPI in information retrieval and life sciences applications. CONCLUSIONS: The hybrid cloud (MapReduce) and cluster (MPI) approach offers an attractive production environment while Twister promises a uniform programming environment for many Life Sciences applications. METHODS: We used commercial clouds Amazon and Azure and the NSF resource FutureGrid to perform detailed comparisons and evaluations of different approaches to data intensive computing. Several applications were developed in MPI, MapReduce and Twister in these different environments. Judy Qiu, Jaliya Ekanayake, Thilina Gunarathne, Jong Choi 0001, Seung-Hee Bae, Bingjing Zhang, Tak-Lon Wu, Yang Ruan 0001, Saliya Ekanayake, Adam Hughes, Geoffrey C. Fox |
BMC Bioinform. | 12 |
| 2010 | Editorial for Economic Models and Algorithms for Grid SystemsabstractThis collection of papers was selected from those presented at the 1st Workshop on Economic Models and Algorithms for Grid Systems, held as part of the eighth IEEE International Conference Grid 2007, Austin, TX, September 19–21, 2007 1. The papers 2-5 were all enhanced over the conference versions and separately reviewed. Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 1 |
| 2010 | The Quakesim portal and services: new approaches to science gateway development techniquesabstractAbstract Traditional techniques in building science portals and gateways are being challenged by new techniques such as Web 2.0 and Cloud Computing. This paper discusses some of our efforts to evaluate these techniques as we evolve the QuakeSim architecture. We believe that architecturally both traditional and newer approaches for Gateways are very similar; thus, giving us a path for moving to hybrid approaches. In this paper, we specifically evaluate techniques for building interactive user interfaces that rely on remote services; architectural approaches for managing massive job submissions that can include both parallel and serial jobs; and an architectural prototype for building component‐based containers compatible with emerging standards. Copyright © 2009 John Wiley & Sons, Ltd. Marlon E. Pierce, Xiaoming Gao, Sangmi Lee Pallickara, Zhenhua Guo 0004, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 5 |
| 2010 | Special Issue: Advanced Scheduling Strategies and Grid Programming Environments
Bruno Schulze, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 2 |
| 2010 | Real-time performance analysis for publish/subscribe systems
Sangyoon Oh 0001, Jai-Hoon Kim, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 3 |
| 2009 | Biomedical Case Studies in Data Intensive Computing
Geoffrey C. Fox, Xiaohong Qiu, Scott Beason, Jong Choi 0001, Jaliya Ekanayake, Thilina Gunarathne, Mina Rho, Haixu Tang, Neil Devadasan, Gilbert C. Liu |
CloudCom | 1 |
| 2009 | Granules: A lightweight, streaming runtime for cloud computing with support, for Map-ReduceabstractCloud computing has gained significant traction in recent years. The Map-Reduce framework is currently the most dominant programming model in cloud computing settings. In this paper, we describe Granules, a lightweight, streaming-based runtime for cloud computing which incorporates support for the Map-Reduce framework. Granules provides rich lifecycle support for developing scientific applications with support for iterative, periodic and data driven semantics for individual computations and pipelines. We describe our support for variants of the Map-Reduce framework. The paper presents a survey of related work in this area. Finally, this paper describes our performance evaluation of various aspects of the system, including (where possible) comparisons with other comparable systems. Shrideep Pallickara, Jaliya Ekanayake, Geoffrey C. Fox |
CLUSTER | 3 |
| 2009 | DryadLINQ for Scientific AnalysesabstractApplying high level parallel runtimes to data/compute intensive applications is becoming increasingly common. The simplicity of the MapReduce programming model and the availability of open source MapReduce runtimes such as Hadoop, are attracting more users to the MapReduce programming model. Microsoft has released DryadLINQ for academic use, allowing users to experience a new programming model and a runtime that is capable of performing large scale data/compute intensive analyses. In this paper, we present our experience in applying DryadLINQ for a series of scientific data analysis applications, identify their mapping to the DryadLINQ programming model, and compare their performances with Hadoop implementations of the same applications. Jaliya Ekanayake, Thilina Gunarathne, Geoffrey C. Fox, Atilla Soner Balkir, Christophe Poulain, Nelson Araujo, Roger S. Barga |
eScience | 3 |
| 2009 | Dynamic Resource-Critical Workflow Scheduling in Heterogeneous Environments
Yili Gong, Marlon E. Pierce, Geoffrey C. Fox |
JSSPP | 3 |
| 2009 | Grids challenged by a Web 2.0 and multicore sandwichabstractAbstract We discuss the application of Web 2.0 to support scientific research (e‐Science) and related ‘e‐more or less anything’ applications. Web 2.0 offers interesting technical approaches (protocols, message formats, and programming tools) to build core e‐infrastructure (cyberinfrastructure) as well as many interesting services (Facebook, YouTube, Amazon S3/EC2, and Google maps) that can add value to e‐infrastructure projects. We discuss why some of the original Grid goals of linking the world's computer systems may not be so relevant today and that interoperability is needed at the data and not always at the infrastructure level. Web 2.0 may also support Parallel Programming 2.0—a better parallel computing software environment motivated by the need to run commodity applications on multicore chips. A ‘Grid on the chip’ will be a common use of future chips with tens or hundreds of cores. Copyright © 2008 John Wiley & Sons, Ltd. Geoffrey C. Fox, Marlon E. Pierce |
Concurr. Comput. Pract. Exp. | 1 |
| 2009 | Using clouds to provide grids with higher levels of abstraction and explicit support for usage modesabstractAbstract Grids in their current form of deployment and implementation have not been as successful as hoped in engendering distributed applications. Among other reasons, the level of detail that needs to be controlled for the successful development and deployment of applications remains too high. We argue that there is a need for higher levels of abstractions for current Grids. By introducing the relevant terminology, we try to understand Grids and Clouds as systems; we find this leads to a natural role for the concept of Affinity, and argue that this is a missing element in current Grids. Providing these affinities and higher‐level abstractions is consistent with the common concepts of Clouds. Thus this paper establishes how Clouds can be viewed as a logical and next higher‐level abstraction from Grids. Copyright © 2009 John Wiley & Sons, Ltd. Shantenu Jha, André Merzky, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 3 |
| 2009 | Using Web 2.0 for scientific applications and scientific communitiesabstractAbstract Web 2.0 approaches are revolutionizing the Internet, blurring lines between developers and users and enabling collaboration and social networks that scale into the millions of users. As discussed in our previous work, the core technologies of Web 2.0 effectively define a comprehensive distributed computing environment that parallels many of the more complicated service‐oriented systems such as Web service and Grid service architectures. In this paper we build upon this previous work to discuss the applications of Web 2.0 approaches to four different scenarios: client‐side JavaScript libraries for building and composing Grid services; integrating server‐side portlets with ‘rich client’ AJAX tools and Web services for analyzing Global Positioning System data; building and analyzing folksonomies of scientific user communities through social bookmarking; and applying microformats and GeoRSS to problems in scientific metadata description and delivery. Copyright © 2009 John Wiley & Sons, Ltd. Marlon E. Pierce, Geoffrey C. Fox, Jong Choi 0001, Zhenhua Guo 0004, Xiaoming Gao |
Concurr. Comput. Pract. Exp. | 2 |
| 2009 | Special Section: Third IEEE International Conference on e-Science and Grid Computing
Kenneth Chiu, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 2 |
| 2009 | Lease-based consistency schemes in the web environment
Byoung-Hoon Lee, Sung-Hwa Lim, Jai-Hoon Kim, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 4 |
| 2008 | MapReduce for Data Intensive Scientific AnalysesabstractMost scientific data analyses comprise analyzing voluminous data collected from various instruments. Efficient parallel/concurrent algorithms and frameworks are the key to meeting the scalability and performance requirements entailed in such scientific data analyses. The recently introduced MapReduce technique has gained a lot of attention from the scientific community for its applicability in large parallel data analyses. Although there are many evaluations of the MapReduce technique using large textual data collections, there have been only a few evaluations for scientific data analyses. The goals of this paper are twofold. First, we present our experience in applying the MapReduce technique for two scientific data analyses: (i) high energy physics data analyses; (ii) K-means clustering. Second, we present CGL-MapReduce, a streaming-based MapReduce implementation and compare its performance with Hadoop. Jaliya Ekanayake, Shrideep Pallickara, Geoffrey C. Fox |
eScience | 3 |
| 2008 | An Overview of the Granules Runtime for Cloud ComputingabstractIn this paper we present a short introduction to the granules system, which is a lightweight streaming-based runtime for cloud computing. This paper provides a summary of the capabilities supported by the runtime. Shrideep Pallickara, Jaliya Ekanayake, Geoffrey C. Fox |
eScience | 3 |
| 2008 | SALSA Project: Parallel Data Mining of GIS, Web, Medical, Physics, Chemical, and Biology DataabstractThe multicore revolution promises potentially hundreds of cores in desktop computers. The ever increasing number of cores per chip will be accompanied by a pervasive data deluge whose size will probably increase even faster than CPU core count over the next few years. This suggests the importance of parallel data analysis and data mining applications with good multicore, cluster and grid performance. The SALSA project at Community Grid Lab of Indiana University is looking to revolutionize the way software is written in parallel for real applications that advance scientific discovery and improve the quality of people's life. Xiaohong Qiu, Geoffrey C. Fox, Seung-Hee Bae, Jong Choi 0001, Jaliya Ekanayake, Yang Ruan 0001 |
eScience | 2 |
| 2008 | An Orchestration for Distributed Web Service HandlersabstractWeb service is a standardization effort to interoperate loosely-coupled applications. A Web service interaction benefits and sometimes requires additive functionalities, called as handlers. They contribute to build rich, modular and efficient Web services. However, the way of utilizing them is very crucial for the Web service architecture and its overall performance. Using distributed approach for the handler execution facilitates significantly to obtain full benefit from them. In this paper we describe an orchestration structure for the handlers to attain richer, more modular and efficient Web services. Beytullah Yildiz, Geoffrey C. Fox, Shrideep Pallickara |
ICIW | 2 |
| 2008 | XML Metadata ServicesabstractAbstract As service‐oriented architecture principles have gained importance, an emerging need has appeared for methodologies to locate desired services that provide access to their capability descriptions. These services must typically be assembled into short‐term service collections that, together with code execution services, are combined into a meta‐application to perform a particular task. To address the metadata requirements of these problems, we introduce a hybrid Information Service to manage both stateless and stateful (transient) metadata. We leverage the two widely used Web Service standards: Universal Description, Discovery and Integration (UDDI) and Web Services Context (WS‐Context), in our design. We describe our approach and experiences when designing ‘semantics’. We report the results from a prototype of the system that is applied to a mobile environment for optimizing Web Service communications. Copyright © 2007 John Wiley & Sons, Ltd. Mehmet S. Aktas, Geoffrey C. Fox, Marlon E. Pierce, Sangyoon Oh 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2008 | Building and applying geographical information system GridsabstractAbstract We discuss the development and application of Web‐service‐based geographical information system (GIS) Grids. Following the WS‐I+ approach of building Grids on Web service standards, we have developed data Grid components for archival and real‐time data, map generating services that can be used to build user interfaces, information services for storing both stateless and stateful metadata, and service orchestration and management tools. Our goal is to support dynamically assembled Grid service collections that combine both GIS services with more traditional Grid capabilities such as file transfer and remote code execution. We are applying these tools to problems in earthquake modeling and forecasting, but we are attempting to build general purpose tools by using and extending appropriate standards. Copyright © 2008 John Wiley & Sons, Ltd. Galip Aydin, Ahmet Sayar, Harshawardhan Gadgil, Mehmet S. Aktas, Geoffrey C. Fox, Sung Hoon Ko, Hasan Bulut, Marlon E. Pierce |
Concurr. Comput. Pract. Exp. | 5 |
| 2008 | Special Issue: The 2nd International Conference on Semantics, Knowledge and GridabstractThe International Conference on Semantics, Knowledge and Grid (SKG) is a cross-area international forum on semantic computing, knowledge networking, and grid computing.SKG promotes crossarea research and prods the development of relevant areas.Themes include the following aspects:• Semantics and Semantic Grid Geoffrey C. Fox, Hai Zhuge |
Concurr. Comput. Pract. Exp. | 1 |
| 2008 | Runtime support for scalable programming in Java
Sang Boem Lim, Han-Ku Lee, Bryan Carpenter, Geoffrey C. Fox |
J. Supercomput. | 4 |
| 2007 | Scalable, fault-tolerant management of Grid ServicesabstractThe service-oriented architecture has come a long way in solving the problem of reusability of existing software resources. Grid applications today are composed of a large number of loosely coupled services. While this has opened up new avenues for building large, complex applications, it has made the management of the application components a nontrivial task. Use of services existing on different platforms, implemented in different languages and presence of variety of network constraints further complicates management. This paper investigates problems that emerge when there is a need to uniformly manage a set of distributed services. We present a scalable, fault-tolerant management framework. Our empirical evaluation shows that the architecture adds an acceptable number of additional resources for providing scalable, fault-tolerant management framework. Harshawardhan Gadgil, Geoffrey C. Fox, Shrideep Pallickara, Marlon E. Pierce |
CLUSTER | 2 |
| 2007 | High Performance Multi-paradigm Messaging Runtime Integrating Grids and Multicore SystemsabstracteScience applications need to use distributed Grid environments where each component is an individual or cluster of multicore machines. These are expected to have 64-128 cores 5 years from now and need to support scalable parallelism. Users will want to compose heterogeneous components into single jobs and run seamlessly in both distributed fashion and on a future "Grid on a chip" with different subsets of cores supporting individual components. We support this with a simple programming model made up of two layers supporting traditional parallel and Grid programming paradigms (workflow) respectively. We examine for a parallel clustering application, the Concurrency and Coordination Runtime CCR from Microsoft as a multi-paradigm runtime that integrates the two layers. Our work uses managed code (C#) and for AMD and Intel processors shows around a factor of 5 better performance than Java. CCR has MPI pattern and dynamic threading latencies of a few microseconds that are competitive with the performance of standard MPI for C. Xiaohong Qiu, Geoffrey C. Fox, Huapeng Yuan, Seung-Hee Bae, George Chrysanthakopoulos, Henrik Frystyk Nielsen |
eScience | 2 |
| 2007 | Scalable, fault-tolerant management in a service oriented architectureabstractThe service-oriented architecture has come a long way in solving the problem of reusability of existing software resources. Grid applications today are composed of a large number of loosely coupled services. While this has opened up new avenues for building large, complex applications, it has made the management of the application components a non-trivial task. Management is further complicated when services exist on different platforms, are written in different languages, present in varying administrative domains restricted by firewalls and are susceptible to failure. This paper investigates problems that emerge when there is a need to uniformly manage a set of distributed services. We present a scalable, fault-tolerant management framework. Our empirical evaluation shows that the architecture adds an acceptable number of additional resources making the approach feasible. Harshawardhan Gadgil, Geoffrey C. Fox, Shrideep Pallickara, Marlon E. Pierce |
HPDC | 2 |
| 2007 | A Scalable Approach for the Secure and Authorized Tracking of the Availability of Entities in Distributed SystemsabstractAs the scale and proliferation of distributed applications continues to increase a need often arises to track the availability of entities that comprise the distributed system. An entity that is part of such a distributed system could be a resource, a service that provides a set of exposed capabilities, an application or a user. In this paper we present a transport-independent scheme for tracking the availability of entities in distributed systems. The scheme enforces the authorized generation and consumption of traces (encapsulating entity availability). The scheme also facilitates the secure distribution of traces while coping with some classes of denial of service attacks. Shrideep Pallickara, Jaliya Ekanayake, Geoffrey C. Fox |
IPDPS | 3 |
| 2007 | The Open Grid Computing Environments collaboration: portlets and services for science gatewaysabstractAbstract We review the efforts of the Open Grid Computing Environments collaboration. By adopting a general three‐tiered architecture based on common standards for portlets and Grid Web services, we can deliver numerous capabilities to science gateways from our diverse constituent efforts. In this paper, we discuss our support for standards‐based Grid portlets using the Velocity development environment. Our Grid portlets are based on abstraction layers provided by the Java CoG kit, which hide the differences of different Grid toolkits. Sophisticated services are decoupled from the portal container using Web service strategies. We describe advance information, semantic data, collaboration, and science application services developed by our consortium. Copyright © 2006 John Wiley & Sons, Ltd. Jay Alameda, Marcus Christie, Geoffrey C. Fox, Joe Futrelle, Dennis Gannon, Mihael Hategan, Gopi Kandaswamy, Gregor von Laszewski, Mehmet A. Nacar, Marlon E. Pierce, Eric Roberts 0002, Charles R. Severance, Mary P. Thomas |
Concurr. Comput. Pract. Exp. | 3 |
| 2007 | Management of real-time streaming data Grid servicesabstractAbstract We discuss our message‐based approach to managing real‐time data streams and building higher level services to produce and consume them. Our messaging system acts as a substrate that can be used to provide qualities of service to various streaming applications ranging from audio–video collaboration systems to sensor Grids. The messaging substrates are composed of distributed, hierarchically arranged message broker networks. Services such as filters are deployed along the edges of the network. We discuss the role of management systems for both broker networks and filter services: broker network topologies must be created and maintained, and distributed filters must be arranged in appropriate sequences. These managed broker networks may be applied to a wide range of problems. We discuss applications to audio–video collaboration in some detail and also describe applications to streaming Global Positioning System data streams. These provide specific application filters that can transform and republish message streams to the broker system. Copyright © 2006 John Wiley & Sons, Ltd. Geoffrey C. Fox, Galip Aydin, Hasan Bulut, Harshawardhan Gadgil, Shrideep Pallickara, Marlon E. Pierce, Wenjun Wu 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Special Issue: Progress of the Knowledge Grid
Geoffrey C. Fox, Xiaoping Sun |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Special Issue: Autonomous Grid ComputingabstractThis special issue selects high-quality papers from the 4th International Conference on Grid and Cooperative Computing ( Geoffrey C. Fox, Hai Zhuge |
Concurr. Comput. Pract. Exp. | 1 |
| 2007 | Fault tolerant high performance Information Services for dynamic collections of Grid and Web services
Mehmet S. Aktas, Geoffrey C. Fox, Marlon E. Pierce |
Future Gener. Comput. Syst. | 2 |
| 2007 | Optimizing Web Service messaging performance in mobile computing
Sangyoon Oh 0001, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 2 |
| 2007 | Special issue on "Grid Technologies"
George A. Gravvanis, John P. Morrison, Geoffrey C. Fox |
J. Supercomput. | 3 |
| 2006 | Special Issue: Workflow in Grid Systems
Geoffrey C. Fox, Dennis Gannon |
Concurr. Comput. Pract. Exp. | 1 |
| 2005 | A scripting based architecture for management of streams and services in real-time grid applicationsabstractRecent specifications such as WS-management and WS-distributed management have stressed the importance of management of resources and services and propose methods towards querying Web services to gather the meta-data associated with these services. Management often entails system setup, querying system metadata, manipulation of system parameters at runtime and taking actions based on the system parameters to tune system performance. Real-time applications require rapid deployment of application components and demand results in real time. In this paper we present the HPSearch system which enables dynamic management of the system including both streams and Web services, and rapid deployment of applications via a scripting interface. We illustrate the functioning of the system by modeling a data streaming application and rapidly deploying the system and application components. Harshawardhan Gadgil, Geoffrey C. Fox, Shrideep Pallickara, Marlon E. Pierce, Robert A. Granat |
CCGRID | 2 |
| 2005 | On the Discovery of Brokers in Distributed Messaging InfrastructuresabstractIncreasingly messaging infrastructures are being used to support the communication requirements of a wide variety of clients, services, and proxies thereto. Typically, for various reason this messaging infrastructure is a distributed one with multiple constituent brokers. In the paper we present our scheme for the discovery of brokers in distributed messaging infrastructures based on the publish/subscribe paradigm. We also include empirical results from our experiments related to the implementation of our scheme Shrideep Pallickara, Harshawardhan Gadgil, Geoffrey C. Fox |
CLUSTER | 3 |
| 2005 | Grids for the GiG and Real Time SimulationsabstractWe study the current architecture of the grid and Web services and that of the global information grid (GiG) with the Network Centric Operations and Warfare (NCOW) from the Department of Defense. We compare the GiG core enterprise services with those being developed for Grids (the open grid services architecture) and Web Services (so called WS-* specifications), identifying both similarities and differences. We discuss both modeling and simulation with HLA (high level architecture) and broad defense NCOW applications. We illustrate this analysis with an open geospatial community (OGC) compatible set of geographical information system grid services. We illustrate the use of grids to efficiently support realtime simulation by an application of grids to audio-video conferencing. Geoffrey C. Fox, Alex Ho, Shrideep Pallickara, Marlon E. Pierce, Wenjun Wu 0001 |
DS-RT | 1 |
| 2005 | Revisiting Distributed Simulation and the Grid: A PanelabstractSummary form only given. The grid, or grid computing, provides a new and unrivalled technology for large scale distributed simulation as it enables collaboration and the use of distributed computing resources. Last year at DS-RT 2004 a panel was convened to consider the impact of the grid on distributed simulation. Four members presented their views of this area and together they tried to identify the main research issues involved in applying grid technology to distributed simulation and the key future challenges that need to be solved to achieve this goal. These challenges included not only technical ones, but also social ones such as management methodology and the development of standards. This year we revisit this fast changing technology and ask the questions: 1) What major changes has grid computing had over the past year?; 2) How can distributed simulation and related applications benefit from grid computing?; 3) What are the barriers to this?; 4) Are there alternative new technologies that would be a better ROI for distributed simulation than grid computing?. Simon J. E. Taylor, Geoffrey C. Fox, Richard M. Fujimoto, J. Mark Pullen, David J. Roberts 0001, Georgios Theodoropoulos 0001 |
DS-RT | 2 |
| 2005 | On the Costs for Reliable Messaging in Web/Grid Service EnvironmentsabstractAs Web services have matured they have been substantially leveraged within the academic, research and business communities. An exemplar of this is the realignment, last year, of the dominant grid application framework - Open Grid Services Infrastucture (OGSI) - with the emerging consensus within the Web services community. Reliable messaging is an important component within the Web services stack. There are two competing, and very similar, specifications within this domain viz. WS-ReliableMessaging (WSRM) and WS-reliability (WSR); this work focuses on the WSRM specification. In this paper we provide an overview of the WSRM protocol, describe our implementation of WSRM, and present an analysis of the costs (in terms of latencies and memory utilizations) involved in the use of WSRM. Since WSRM is very similar to WS-reliability we expect the performance of WSRM to be very similar to that of WSR. We hope that the work presented here helps researchers and systems designers gauge the suitability of Web services based reliable messaging in their applications and also to make appropriate trade-offs, which includes inter alia interoperability, guarantees, quality of service and performance. Shrideep Pallickara, Geoffrey C. Fox, Beytullah Yildiz, Sangmi Lee Pallickara, Sima Patel, Damodar Yemme |
e-Science | 2 |
| 2005 | A Low-Level Communication Library for Java HPC
Sang Boem Lim, Bryan Carpenter, Geoffrey C. Fox, Han-Ku Lee |
ICA3PP | 3 |
| 2005 | GridFTP and Parallel TCP Support in NaradaBrokering
Sang Boem Lim, Geoffrey C. Fox, Ali Kaplan, Shrideep Pallickara, Marlon E. Pierce |
ICA3PP | 2 |
| 2005 | eSports: Collaborative and Synchronous Video Annotation System in Grid Computing EnvironmentabstractWe designed eSports - a collaborative and synchronous video annotation platform, which is to be used in Internet scale cross-platform grid computing environment to facilitate computer supported cooperative work (CSCW) in education settings such as distance sport coaching, distance classroom etc. Different from traditional multimedia annotation systems, eSports provides the capabilities to collaboratively and synchronously play and archive real time live video, to take snapshots, to annotate video snapshots using whiteboard and to play back the video annotations synchronized with original video streams. eSports is designed based on the grid based collaboration paradigm $the shared event model using NaradaBrokering, which is a publish/subscribe based distributed message passing and event notification system. In addition to elaborate the design and implementation of eSports, we analyze the potential use cases of eSports under different education settings. We believed that eSports is very useful to improve the online collaborative coaching and education. Gang Zhai, Geoffrey C. Fox, Marlon E. Pierce, Wenjun Wu 0001, Hasan Bulut |
ISM | 2 |
| 2005 | Collective Communications for Scalable Programming
Sang Boem Lim, Bryan Carpenter, Geoffrey C. Fox, Han-Ku Lee |
ISPA | 3 |
| 2005 | Data Integration Hub for a Hybrid Paper Search
Jungkee Kim, Geoffrey C. Fox, Seong Joon Yoo |
KES (3) | 2 |
| 2005 | Web Service Grids: an evolutionary approachabstractAbstract The U.K. e‐Science Programme is a £250 million, five‐year initiative which has funded over 100 projects. These application‐led projects are underpinned by an emerging set of core middleware services that allow the coordinated, collaborative use of distributed resources. This set of middleware services runs on top of the research network and beneath the applications we call the ‘Grid’. Grid middleware is currently in transition from pre‐Web Service versions to a new version based on Web Services. Unfortunately, only a very basic set of Web Services embodied in the Web Services Interoperability proposal, WS‐I, are agreed by most IT companies. IBM and others have submitted proposals for Web Services for Grids—the Web Services ResourceFramework and Web Services Notification specifications—to the OASIS organization for standardization. This process could take up to 12 months from March 2004 and the specifications are subject to debate and potentially significant changes. Since several significant U.K. e‐Science projects come to an end before the end of this process, the U.K. needs to develop a strategy that will protect the U.K.'s investment in Grid middleware by informing the Open Middleware Infrastructure Institute's (OMII) roadmap and U.K. middleware repository in Southampton. This paper sets out an evolutionary roadmap that will allow us to capture generic middleware components from projects in a form that will facilitate migration or interoperability with the emerging Grid Web Services standards and with ongoing OGSA developments. In this paper we therefore define a set of Web Services specifications, which we call ‘WS‐I+’ to reflect the fact that this is a larger set than currently accepted by WS‐I, that we believe will enable us to achieve the twin goals of capturing these components and facilitating migration to future standards. We believe that the extra Web Services specifications we have included in WS‐I+ are both helpful in building e‐Science Grids and likely to be widely accepted. Copyright © 2005 John Wiley & Sons, Ltd. Malcolm P. Atkinson 0001, David De Roure, Alistair N. Dunlop, Geoffrey C. Fox, Peter Henderson 0001, Anthony J. G. Hey, Norman W. Paton, Steven J. Newhouse, Savas Parastatidis, Anne E. Trefethen, Paul Watson 0001, Jim Webber |
Concurr. Pract. Exp. | 4 |
| 2005 | Special Issue: ACM 2002 Java Grande-ISCOPE Conference
Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 2005 | Towards enabling peer-to-peer GridsabstractAbstract In this paper we propose a peer‐to‐peer (P2P) Grid comprising resources such as relatively static clients, high‐end resources and a dynamic collection of multiple P2P subsystems. We investigate the architecture of the messaging and event service that will support such a hybrid environment. We designed a distributed publish–subscribe system NaradaBrokering for XML‐specified messages. NaradaBrokering provides support for centralized, distributed and P2P (via JXTA) interactions. Here we investigate and present our strategy for the integration of JXTA into NaradaBrokering. The resultant system naturally scales with multiple Peer Groups linked by NaradaBrokering. Copyright © 2005 John Wiley & Sons, Ltd. Geoffrey C. Fox, Shrideep Pallickara, Xi Rao |
Concurr. Pract. Exp. | 1 |
| 2005 | Special Issue: Grids and Web Services for e-ScienceabstractAbstract This editorial describes four papers that summarize key Grid technology capabilities to support distributed e‐Science applications. These papers discuss the Condor system supporting computing communities, the OGSA‐DAI service interfaces for databases, the WS‐I+ Grid Service profile and finally WS‐GAF (the Web Service Grid Application Framework). We discuss the confluence of mainstream IT industry development and the very latest science and computer science research and urge the communities to reach consensus rapidly. Agreement on a set of core Web Service standards is essential to allow developers to build Grids and distributed business and science applications with some assurance that their investment will not be obviated by the changing Web Service frameworks. Copyright © 2005 John Wiley & Sons, Ltd. Anthony J. G. Hey, Geoffrey C. Fox |
Concurr. Pract. Exp. | 2 |
| 2005 | Collective communication for the HPJava programming languageabstractAbstract This paper addresses functionality and implementation of a HPJava version of the Adlib collective communication library for data parallel programming. We begin by illustrating typical use of the library, through an example multigrid application. Then we describe implementation issues for the high‐level library. At a software engineering level, we illustrate how the primitives of the HPJava language assist in writing library methods whose implementation can be largely independent of the distribution format of the argument arrays. We also describe a low‐level API called mpjdev, which handles basic communication underlying the Adlib implementation. Finally we present some benchmark results, and some conclusions. Copyright © 2005 John Wiley & Sons, Ltd. Sang Boem Lim, Bryan Carpenter, Geoffrey C. Fox, Han-Ku Lee |
Concurr. Pract. Exp. | 3 |
| 2005 | Performance of a possible Grid message infrastructureabstractAbstract In this paper we present the results pertaining to the NaradaBrokering middleware infrastructure. NaradaBrokering is designed to run on a large network of cooperating broker nodes. NaradaBrokering capabilities include, among other things, support for a wide variety of transport protocols, Java Message Service compliance, support for routing JXTA interactions, support for audio/video conferencing applications and, finally, support for multiple constraint specification formats such as XPath, SQL and regular expression queries. This paper demonstrates the suitability of NaradaBrokering to a wide variety of applications and scenarios. Copyright © 2005 John Wiley & Sons, Ltd. Shrideep Pallickara, Geoffrey C. Fox, Ahmet Uyar, Xi Rao, David W. Walker, Beytullah Yildiz |
Concurr. Pract. Exp. | 2 |
| 2005 | Automating metadata Web service deployment for problem solving environments
Ozgur Balsoy, Ying Jin 0005, Galip Aydin, Marlon E. Pierce, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 5 |
| 2005 | Message-based cellular peer-to-peer grids: foundations for secure federation and autonomic services
Geoffrey C. Fox, Sang Lim, Shrideep Pallickara, Marlon E. Pierce |
Future Gener. Comput. Syst. | 1 |
| 2005 | Building Problem-Solving Environments with Application Web Service toolkits
Choon-Han Youn, Marlon E. Pierce, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 3 |
| 2005 | Deploying the NaradaBrokering Substrate in Aiding Efficient Web and Grid Service InteractionsabstractNaradaBrokering has been developed as the messaging infrastructure for collaboration, peer-to-peer, and Grid applications. The value of NaradaBrokering in the context of Grid and Web Services has been clear for some time. NaradaBrokering-combined with further extensions to, and testing of, its existing capabilities - can also take advantage of the maturing of Web Service standards and specifications to build very powerful general mechanisms to deploy and integrate it with general Web Services. This paper describes a framework to integrate the NaradaBrokering substrate with Web Services. Geoffrey C. Fox, Shrideep Pallickara |
Proc. IEEE | 1 |
| 2004 | Wireless Reliable Messaging Protocol for Web Services (WS-WRM)abstractBy employing Web services technology, the grid system is evolving to be more manageable service infrastructure, including lifetime management, discovery of characteristics, and notifications. Reliable messaging is one of the key issues addressed for quality of services in Web services. In this paper, we propose a reliable message scheme designed for mobile environments in the context of a Web services architecture; Web services-Wireless Reliable Messaging (WS-WRM). We also consider the federation issues with emerging specifications proposed by leading Web services standard groups. We eventually intend to extend the reliability to mobile end-nodes in a more efficient way. In this paper, we address the design issues and describe the detailed scheme of messaging architecture. Sangmi Lee, Geoffrey C. Fox |
ICWS | 2 |
| 2004 | Toward Flexible Messaging for SOAP-Based ServicesabstractNaradaBrokering provides a messaging abstraction that allows it to provide message-related capabilities in a transparent fashion. These capabilities include message-based security, time and causal ordering, compression, virtualization of transport protocol and addressing, and fault tolerance related functionalities. NaradaBrokering — combined with further extensions to its existing capabilities — can also take advantage of the maturing of Web Service specifications to build very powerful general mechanisms to deploy and integrate it with general Web services. In this paper we describe our strategy to interface NaradaBrokering with Web services. The strategy described in this paper will allow new, and existing, applications built around the Web Services Framework to leverage capabilities offered by the NaradaBrokering substrate without changes to the service implementations. Geoffrey C. Fox, Shrideep Pallickara, Savas Parastatidis |
SC | 1 |
| 2004 | Global multimedia collaboration systemabstractAbstract In order to build an integrated collaboration system over heterogeneous collaboration technologies, we propose a global multimedia collaboration system (Global‐MMCS) based on XGSP A/V Web‐Services framework. This system can integrate multiple A/V services, and support various collaboration clients and communities. Now the prototype is being developed and deployed across many universities in U.S.A. and China. Copyright © 2004 John Wiley & Sons, Ltd. Geoffrey C. Fox, Wenjun Wu 0001, Ahmet Uyar, Hasan Bulut, Shrideep Pallickara |
Concurr. Pract. Exp. | 1 |
| 2003 | A Web Service Approach to Universal Accessibility in Collaboration Services
Sangmi Lee, Sung Hoon Ko, Geoffrey C. Fox, Kangseok Kim, Sangyoon Oh 0001 |
ICWS | 3 |
| 2003 | Integration of SIP VoIP and Messaging Systems with AccessGrid and H.323
Wenjun Wu 0001, Ahmet Uyar, Hasan Bulut, Geoffrey C. Fox |
ICWS | 4 |
| 2003 | NaradaBrokering: A Distributed Middleware Framework and Architecture for Enabling Durable Peer-to-Peer Grids
Shrideep Pallickara, Geoffrey C. Fox |
Middleware | 2 |
| 2002 | A Batch Script Generator Web Service for Computational PortalsabstractComputational portal developers often reimplement functionality found in existing portals because there are no common discovery and access mechanisms in place to enable sharing of portal functions. The Web services architecture provides an implementation independent, protocol based mechanism for entities to find, share, and invoke remote services. We developed complementary web services at SDSC and IU that generate batch scripts for different batch queuing schedulers. Stephen A. Mock, Choon-Han Youn, Marlon E. Pierce, Geoffrey C. Fox, Mary P. Thomas |
HPDC | 4 |
| 2002 | Interoperable Web services for computational portalsabstractComputational web portals are designed to simplify access to diverse sets of high performance computing resources, typically through an interface to computational Grid tools. An important shortcoming of these portals is their lack of interoperable and reusable services. This paper presents an overview of research efforts undertaken by our group to build interoperating portal services around a Web Services model. We present a comprehensive view of an interoperable portal architecture, beginning with core portal services that can be used to build Application Web Services, which in turn may be aggregated and managed through portlet containers. Marlon E. Pierce, Geoffrey C. Fox, Choon-Han Youn, Stephen A. Mock, Kurt Mueller, Ozgur Balsoy |
SC | 2 |
| 2002 | Grid services for earthquake scienceabstractAbstract We describe an information system architecture for the ACES (Asia–Pacific Cooperation for Earthquake Simulation) community. It addresses several key features of the field—simulations at multiple scales that need to be coupled together; real‐time and archival observational data, which needs to be analyzed for patterns and linked to the simulations; a variety of important algorithms including partial differential equation solvers, particle dynamics, signal processing and data analysis; a natural three‐dimensional space (plus time) setting for both visualization and observations; the linkage of field to real‐time events both as an aid to crisis management and to scientific discovery. We also address the need to support education and research for a field whose computational sophistication is rapidly increasing and spans a broad range. The information system assumes that all significant data is defined by an XML layer which could be virtual, but whose existence ensures that all data is object‐based and can be accessed and searched in this form. The various capabilities needed by ACES are defined as grid services, which are conformant with emerging standards and implemented with different levels of fidelity and performance appropriate to the application. Grid Services can be composed in a hierarchical fashion to address complex problems. The real‐time needs of the field are addressed by high‐performance implementation of data transfer and simulation services. Further, the environment is linked to real‐time collaboration to support interactions between scientists in geographically distant locations. Copyright © 2002 John Wiley & Sons, Ltd. Geoffrey C. Fox, Sung Hoon Ko, Marlon E. Pierce, Ozgur Balsoy, Jake Kim, Sangmi Lee, Kangseok Kim, Sangyoon Oh 0001, Xi Rao, Mustafa Varank, Hasan Bulut, Gurhan Gunduz, Xiaohong Qiu, Shrideep Pallickara, Ahmet Uyar, Choon-Han Youn |
Concurr. Comput. Pract. Exp. | 1 |
| 2002 | An event service to support Grid computational environmentsabstractAbstract We believe that it is interesting to study the system and software architecture of environments which integrate the evolving ideas of computational Grids, distributed objects, Web services, peer‐to‐peer (P2P) networks and message‐oriented middleware. Such P2P Grids should seamlessly integrate users to themselves and to resources which are also linked to each other. We can abstract such environments as a distributed system of ‘clients’ which consist either of ‘users’ or ‘resources’ or proxies thereto. These clients must be linked together in a flexible, fault‐tolerant, efficient, high‐performance fashion. In this paper, we study the messaging or event system—termed Grid Event Service (GES)—that is appropriate to link the clients (both users and resources of course) together. For our purposes (registering, transporting and discovering information), events are just messages—typically with time stamps. The messaging system GES must scale over a wide variety of devices—from handheld computers at one extreme to high‐performance computers and sensors at the other. We have analyzed the requirements of several Grid services that could be built with this model, including computing and education and incorporated constraints of collaboration with a shared event model. We suggest that generalizing the well‐known publish–subscribe model is an attractive approach and here we study some of the issues to be addressed if this model is used in GES. Copyright © 2002 John Wiley & Sons, Ltd. Geoffrey C. Fox, Shrideep Pallickara |
Concurr. Comput. Pract. Exp. | 1 |
| 2002 | The Gateway computational Web portalabstractAbstract In this paper we describe the basic services and architecture of Gateway, a commodity‐based Web portal that provides secure remote access to unclassified Department of Defense computational resources. The portal consists of a dynamically generated, browser‐based user interface supplemented by client applications and a distributed middle tier, WebFlow. WebFlow provides a coarse‐grained approach to accessing both stand‐alone and Grid‐enabled back‐end computing resources. We describe in detail the implementation of basic portal features such as job submission, file transfer, and job monitoring and discuss how the portal addresses security requirements of the deployment centers. Finally, we outline future plans, including integration of Gateway with Department of Defense testbed Grids. Copyright © 2002 John Wiley & Sons, Ltd. Marlon E. Pierce, Choon-Han Youn, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 3 |
| 2001 | Special Issue: ACM 2000 Java Grande Conference
Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 1 |
| 2001 | Editorial: Concurrency and Computation: Practice and Experience
Geoffrey C. Fox, Anthony J. G. Hey |
Concurr. Comput. Pract. Exp. | 1 |
| 2000 | Metacomputing
Alexander Reinefeld, Geoffrey C. Fox, Domenico Laforenza, Edward Seidel |
Euro-Par | 2 |
| 2000 | Object serialization for marshaling data in a Java interface to MPIabstractSeveral Java bindings to Message Passing Interface (MPI) software have been developed recently. Message buffers have usually been restricted to arrays with elements of primitive type. We discuss adoption of the Java object serialization model for marshaling general communication data in MPI-like APIs. This approach is compared with a Java transcription of the standard MPI derived datatype mechanism. We describe an implementation of the mpiJava interface to MPI that incorporates automatic object serialization. Benchmark results confirm that current JDK implementations of serialization are not fast enough for high performance messaging applications. Means of solving this problem are discussed, and benchmarks for greatly improved schemes are presented. Copyright © 2000 John Wiley & Sons, Ltd. Bryan Carpenter, Geoffrey C. Fox, Sung Hoon Ko, Sang Lim |
Concurr. Pract. Exp. | 2 |
| 2000 | MPJ: MPI-like message passing for JavaabstractRecently, there has been a lot of interest in using Java for parallel programming. Efforts have been hindered by lack of standard Java parallel programming APIs. To alleviate this problem, various groups started projects to develop Java message passing systems modelled on the successful Message Passing Interface (MPI). Official MPI bindings are currently defined only for C, Fortran, and C++, so early MPI-like environments for Java have been divergent. This paper relates an effort undertaken by a working group of the Java Grande Forum, seeking a consensus on an MPI-like API, to enhance the viability of parallel programming using Java. Copyright © 2000 John Wiley & Sons, Ltd. Bryan Carpenter, Vladimir Getov, Glenn Judd, Anthony Skjellum, Geoffrey C. Fox |
Concurr. Pract. Exp. | 5 |
| 2000 | Special Issue: ACM 1999 Java Grande Conference (Editorial)abstractThis is the first part of a three-part special issue devoted to selected papers from the ACM 1999 Java Grande Conference, held in San Francisco on 12-14 June 1999.All the papers have been revised and re-refereed to ensure appropriate journal quality.This was the fourth of a series of meetings, exploring the use of the Java programming language for scientific and engineering computing and high-performance network computing-a range of applications that has been denoted with the epithet 'Grande'.The previous Java Grande workshops were held very successfully in Syracuse in 1996, Las Vegas in 1997, and in Palo Alto in 1998.The proceedings were also published in special issues of Concurrency: Practice and Experience (Volume 9, Issues 6, 11; Volume 10, Issues 11-13).The papers have been divided into three groups: (a) largely Java Virtual Machine, Language and compiler issues (this issue); (b) largely technology of general applicability (Volume 12, Issue 7); (c) largely distributed computing and applications (Volume 12, Issue 8).The Java Grande conference focuses on the use of Java in the broad area of high-performance computing; including engineering and scientific applications, simulations, data-intensive applications, and other emerging application areas that exploit parallel and distributed computing or combine communication and computing.We believe that Java will play an increasingly important role in these areas.The goal is to provide feedback to users and language developers on what is required to successfully deploy Java in a broad range of scientific and high-performance network computing systems.The 1998 workshop initiated the Java Grande Forum process which has organized activities aimed at making Java a superior programming environment for 'Grande applications'.This conference included a short forum meeting.Further details about related activities will be found at http://www.javagrande.org. Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 2000 | Special Issue: ACM 1999 Java Grande Conference (Editorial)
Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 2000 | Special Issue: ACM 1999 Java Grande Conference (Editorial)abstractThis is the third part of a three-part special issue devoted to selected papers from the ACM 1999 Java Grande Conference, held in San Francisco on 12-14 June 1999.All the papers have been revised and re-refereed to ensure appropriate journal quality.This was the fourth of a series of meetings, exploring the use of the Java programming language for scientific and engineering computing and high-performance network computing-a range of applications that has been denoted with the epithet 'Grande'.The previous Java Grande workshops were held very successfully in Syracuse in 1996, Las Vegas in 1997, and in Palo Alto in 1998.The proceedings were also published in special issues of Concurrency: Practice and Experience (Volume 9, Issues 6, 11; Volume 10, Issues 11-13).The papers have been divided into three groups: (a) largely Java Virtual Machine, Language and compiler issues (Volume 12, Issue 6); (b) largely technology of general applicability (Volume 12, Issue 7); (c) largely distributed computing and applications (this issue).The Java Grande conference focuses on the use of Java in the broad area of high-performance computing; including engineering and scientific applications, simulations, data-intensive applications, and other emerging application areas that exploit parallel and distributed computing or combine communication and computing.We believe that Java will play an increasingly important role in these areas.The goal is to provide feedback to users and language developers on what is required to successfully deploy Java in a broad range of scientific and high-performance network computing systems.The 1998 workshop initiated the Java Grande Forum process which has organized activities aimed at making Java a superior programming environment for 'Grande applications'.This conference included a short forum meeting.Further details about related activities will be found at http://www.javagrande.org. Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 2000 | The Gateway system: uniform web based access to remote resourcesabstractExploiting our experience developing the WebFlow system, we designed the Gateway system to provide seamless and secure access to computational resources at ASC MSRC. The Gateway follows our commodity components strategy and is implemented as a modern three-tier system. Tier 1 is a high-level front-end for visual programming, steering, run-time data analysis and visualization, built on top of the Web and OO commodity standards. Distributed object-based, scalable, and reusable Web server and Object broker middleware forms Tier 2. Back-end services comprise Tier 3. In particular, access to high-performance computational resources is provided by implementing the emerging standard for meta-computing API. Copyright © 2000 John Wiley & Sons, Ltd. Tomasz Haupt, Erol Akarsu, Geoffrey C. Fox, Choon-Han Youn |
Concurr. Pract. Exp. | 3 |
| 2000 | WebFlow: a framework for web based metacomputingabstractWe developed a platform independent, three-tier system, called WebFlow. The visual authoring tools implemented in the front end integrated with the middle-tier network of servers based on CORBA and following distributed object paradigm, facilitate seamless integration of commodity software components. We add high performance to commodity systems using Globus metacomputing toolkit as the backend. Tomasz Haupt, Erol Akarsu, Geoffrey C. Fox |
Future Gener. Comput. Syst. | 3 |
| 1999 | Using Gateway System to Provide a Desktop Access to High Performance Computational ResourcesabstractIn this paper, we discuss the use of Gateway for seamless desktop access to high performance resources. We illustrate our ideas with two Gateway applications that require access to remote resources: the Landscape Management System (LMS) and Quantum Simulations (QS). For LMS we use Gateway to retrieve data from many different sources as well as to allocate remote computational resources needed to solve the problem at hand. Gateway transparently controls the necessary data transfer between hosts for the user. Quantum Simulations requires access to HPCC resources and therefore we layered Gateway on top of the Globus metacomputing toolkit. This way Gateway plays the role of a job broker for Globus. Erol Akarsu, Geoffrey C. Fox, Tomasz Haupt, Alexey Kalinichenko, Kangseok Kim, Praveen Sheethaalnath, Choon-Han Youn |
HPDC | 2 |
| 1999 | A computing framework for integrating interactive visualization in HPCC applicationsabstractNetwork-based concurrent computing and interactive data visualization are two important components in industry applications of high-performance computing and communication. We propose an execution framework to build interactive remote visualization systems for real-world applications on heterogeneous parallel and distributed computers. Using a dataflow model of a commercial visualization software AVS in three case studies, we demonstrate a simple, effective, and modular approach to couple parallel simulation modules into an interactive remote visualization environment. The applications described in this paper are drawn from our industrial projects in financial modeling, computational electromagnetics and computational chemistry. Copyright © 1999 John Wiley & Sons, Ltd. Geoffrey C. Fox, Tseng-Hui Lin, Tomasz Haupt |
Concurr. Pract. Exp. | 2 |
| 1999 | Web based metacomputingabstractAbstract Programming tools that are simultaneously sustainable, highly functional, robust and easy to use have been hard to come by in the HPCC arena. This is partially due to the difficulty in developing sophisticated customized systems for what is a relatively small part of the worldwide computing enterprise. Thus, we have developed a new strategy – termed High Performance Commodity Computing (HPCC) [G. Fox, W. Furmanski, HPCC as high performance commodity computing, in: I. Foster, C. Kesselman (Eds.), Building National Grid, http://www.npac.syr.edu/users/gcf/HPcc/HPcc.html ] – which builds HPCC programming tools on top of the remarkable new software infrastructure being built for the commercial web and distributed object areas. We add high performance to commodity systems using multi-tier architecture with Globus metacomputing toolkit as the backend of a middle-tier of commodity web and object servers. We have demonstrated the fully functional prototype of WebFlow during Alliance’98 meeting. Tomasz Haupt, Erol Akarsu, Geoffrey C. Fox, Wojtek Furmanski |
Future Gener. Comput. Syst. | 3 |
| 1998 | Towards a Java Environment for SPMD Programming
Bryan Carpenter, Guansong Zhang, Geoffrey C. Fox, Yuhong Wen |
Euro-Par | 3 |
| 1998 | HPcc as High Performance Commodity Computing on Top of Integrated Java, CORBA, COM and Web Standards
Geoffrey C. Fox, Wojtek Furmanski, Tomasz Haupt, Erol Akarsu, Hasan Timucin Ozdemir |
Euro-Par | 1 |
| 1998 | Java data parallel extensions with runtime system supportabstractIn order to provide Java with the ability for supporting scientific parallel computing, we introduce a data parallel extension to Java language with runtime system support. We provide the distributed array extension to Java, and discuss the related operation and control over the new distributed array. Communication involving distributed arrays are handles through a standard of a collective communication library. We consider the programming in a Single Program Multiple Data (SPMD) model. Yuhong Wen, Bryan Carpenter, Geoffrey C. Fox, Guansong Zhang |
HiPC | 3 |
| 1998 | Language Bindings for a Data-Parallel RuntimeabstractThe NPAC kernel runtime, developed in the PCRC (Parallel Compiler Runtime Consortium) project, is a runtime library with special support for the High Performance Fortran data model. It provides array descriptors for a generalized class of HPF like distributed arrays, support for parallel access to their elements, and a rich library of collective communication and arithmetic operations for manipulating these arrays. The library has been successfully used as a component in experimental HPF translation systems. With prospects for early appearance of fully featured, efficient HPF compilers looking questionable, we discuss a class of more easily implementable data parallel language extensions that preserve many of the attractive features of HPF, while providing the programmer with direct access to runtime libraries such as the NPAC PCRC kernel. Bryan Carpenter, Geoffrey C. Fox, Donald Leskiw, Yuhong Wen, Guansong Zhang |
HIPS | 2 |
| 1998 | WebFlow - High-Level Programming Environment and Visual Authoring Toolkit for High Performance Distributed ComputingabstractWe developed a platform independent, three-tier system, called WebFlow. The visual authoring tools implemented in the front end integrated with the middle tier network of servers based on the industry standards and following distributed object paradigm, facilitate seamless integration of commodity software components. We add high performance to commodity systems using GLOBUS metacomputing toolkit as the backend. We have explained these ideas in general before, and here for the first time we describe a fully operational example which is expected to be deployed in an NCSA Alliance Grand Challenge. Erol Akarsu, Geoffrey C. Fox, Wojtek Furmanski, Tomasz Haupt |
SC | 2 |
| 1998 | DARP: Java-based data analysis and rapid prototyping environment for distributed high performance computationsabstractThe integration of a compiled and interpreted HPF gives us an opportunity to design a powerful application development environment targeted for high performance parallel and distributed systems. This Web based system follows a three-tier model. The Java front-end holds proxy objects which can be manipulated with an interpreted Web client (a Java applet) interacting dynamically with compiled code through a tier-2 server. Although targeted for the HPF back-end, the system's architecture is independent of the back-end language, and can be extended to support other high performance languages. © 1998 John Wiley & Sons, Ltd. Erol Akarsu, Geoffrey C. Fox, Tomasz Haupt |
Concurr. Pract. Exp. | 2 |
| 1998 | HPJava: Data Parallel Extensions to JavaabstractWe outline an extension of Java for programming with distributed arrays. The basic programming style is Single Program Multiple Data (SPMD), but parallel arrays are provided as new language primitives. Further extensions include three distributed control constructs, the most important being a data-parallel loop construct. Communications involving distributed arrays are handled through a standard library of collective operations. Because the underlying programming model is SPMD programming, direct calls to MPI or other communication packages are also allowed in an HPJava program. © 1998 John Wiley & Sons, Ltd. Bryan Carpenter, Guansong Zhang, Geoffrey C. Fox, Yuhong Wen |
Concurr. Pract. Exp. | 3 |
| 1998 | Java for High-performance Network Computing - Editorial
Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 1997 | Design Issues in Building Web-Based Programming EnvironmentsabstractWe exploited the recent advances in Internet connectivity and Web technologies for building Web-based parallel programming environments (WPPEs) that facilitate the development and execution of parallel programs on remote high-performance computers. A Web browser running on the user's machine provides a user-friendly interface to server-site user accounts and allows the use of parallel computing platforms and software in a convenient manner. The user may create, edit, and execute files through this Web browser interface. This new Web-based client-server architecture has the potential of being used as a future front-end to high-performance computer systems. We discuss the design and implementation of several prototype WPPEs that are currently in use at the Northeast Parallel Architectures Center and the Cornell Theory Center. These initial prototypes support high-level parallel programming with Fortran 90 and High Performance Fortran (HPF), as well as explicit low-level programming with Message Passing Interface (MPI). We detail the lessons learned during the development process and outline the tradeoffs of various design choices in the realization of the design. We especially concentrate on providing server-site user accounts, mechanisms to access those accounts through the Web, and the Web-related system security issues. Kivanç Dinçer, Geoffrey C. Fox |
HPDC | 2 |
| 1997 | A Comparison of Annealing Techniques for Academic Course Scheduling
M. A. Saleh Elmohamed, Paul D. Coddington, Geoffrey C. Fox |
PATAT | 3 |
| 1997 | Java enabling collaborative education, health care, and computingabstractWeb technologies – in particular linked Java servers and clients – allow new dynamic collaborative environments linking people and computers. We describe the architecture of a system, TANGOsim, that combines a Java collaborative environment with an executive providing general message filters, and an event-driven simulator. The initial application is to command and control, but we describe how this approach can also be used in other areas, such as health care, scientific visualization and (distance) education. © 1997 John Wiley & Sons, Ltd. Lukasz Beca, Geoffrey C. Fox, Tomasz Jurga, Konrad Olszewski, Marek Podgorny, Piotr Sokolowski, Krzysztof Walczak 0001 |
Concurr. Pract. Exp. | 3 |
| 1997 | WebFlow - a visual programming paradigm for Web/Java based coarse grain distributed computingabstractWe present here recent work at NPAC aimed at developing WebFlow – a general purpose Web-based visual interactive programming environment for coarse grain distributed computing. We follow the 3-tier architecture with the central control and integration WebVM layer in tier-2, interacting with the visual graph editor applets in tier-1 (front-end) and the legacy systems in tier-3. WebVM is given by a mesh of Java Web servers such as Jeeves from JavaSoft or Jigsaw from MIT/W3C. All system control structures are implemented as URL-addressable servlets which enable Web browser-based authoring, monitoring, publication, documentation and software distribution tools for distributed computing. We view WebFlow/WEbVM as a promising programming paradigm and co-ordination model for the exploding volume of Web/Java software, and we illustrate it in a set of ongoing application development activities. © 1997 John Wiley & Sons, Ltd. Dimple Bhatia, Vanco Burzevski, Maja Camuseva, Geoffrey C. Fox, Wojtek Furmanski, Girish Premchandran |
Concurr. Pract. Exp. | 4 |
| 1997 | Experiments with HP JavaabstractWe consider the possible role of Java as a language for high performance computing. After discussing reasons why Java may be a natural candidate for a portable parallel programming language, we describe several case studies. These cover Java socket programming, message-passing through a Java interface to MPI, and class libraries for data-parallel programming in Java. © 1997 John Wiley & Sons, Ltd. Bryan Carpenter, Yuh-Jye Chang, Geoffrey C. Fox, Donald Leskiw |
Concurr. Pract. Exp. | 3 |
| 1997 | A comparison of optimization heuristics for the data mapping problemabstractIn the paper we compare the performance of six heuristics with suboptimal solutions for the data distribution of two dimensional meshes that are used for the numerical solution of partial differential equations (PDEs) on multicomputers. The data mapping heuristics are evaluated with respect to seven criteria covering load balancing, interprocessor communication, flexibility and ease of use for a class of single-phase iterative PDE solvers. Our evaluation suggests that the simple and fast block distribution heuristic can be as effective as the other five complex and computational expensive algorithms. © 1997 by John Wiley & Sons, Ltd. Nikos Chrisochoides, Nashat Mansour, Geoffrey C. Fox |
Concurr. Pract. Exp. | 3 |
| 1997 | Using Java and JavaScript in the Virtual Programming Laboratory: a Web-based parallel programming environmentabstractThe Virtual Programming Laboratory (VPL) is a Web-based virtual programming environment built based on a client–server architecture. The system can be accessed on any platform (Unix, PC or Mac) using a standard Java-enabled browser. Software delivery over the Web imposes a novel set of constraints on design. We outline the tradeoffs in this design space, motivate the choices necessary to deliver an application, and detail the lessons learned in the process. We discuss the role of Java and other Web technologies in the realization of the design. VPL facilitates the development and execution of parallel programs. The initial prototype supports high-level parallel programming based on Fortran 90 and High Performance Fortran (HPF), as well as explicit low-level programming with the MPI message-passing interface. Supplementary Java-based platform-independent tools for data and performance visualization are an integral part of the VPL. Pablo SDDF trace files generated by the Pablo performance instrumentation system are used for post-mortem performance visualization. © 1997 John Wiley & Sons, Ltd. Kivanç Dinçer, Geoffrey C. Fox |
Concurr. Pract. Exp. | 2 |
| 1997 | Java for Computational Science and Engineering - Simulation and Modeling II (Editorial)abstractJava for Computational Science and Engineering -Simulation and Modeling IIWe are pleased to present a second set of papers discussing the role of Java in Science and Engineering Simulation.The first group of 14 publications was presented at a small workshop with 45 participants at Syracuse, 16-17 December 1996 and published in the June 97 issue of Concurrency: Practice and Experience.Here, we introduce 30 papers from a follow-up workshop with over 100 participants that was sponsored by ACM in Las Vegas on 21 June 1997.The growing interest in this field is also supported by an email discussion list and other materials collected at the web site http://www.npac.syr.edu/projects/javaforcse.Java and Web technology can be used in many areas of science and engineering computation.These include sophisticated user interfaces and coarse-grain integration of different modules in complex meta-applications.However, there also seems general agreement that it is possible to build powerful compilers that will allow Java coded programs to give performance that is very competitive with those written in C or Fortran.The first six papers ('Annotating the Java bytecodes in support of optimization' by Joseph Hummel, Ana Azevedo, David Kolson and Alexandru Nicolau; 'CACAO -A 64 bit JavaVM just-intime compiler' by Andreas Krall and Reinhard Grafl; 'A Java bytecode optimizer using side-effect analysis' by Lars R. Clausen; 'A prototype of Fortran-to-Java converter' by Geoffrey Fox, Xiaoming Li, Zheng Qiang and Wu Zhigang; 'Just-in-time optimizations for high-performance Java programs' by Michał Cierniak and Wei Li; 'Interactive simulations on the Web: Compiling NESL into Java' by Jonathan Hardwick, Girija Narlikar and Jay Sipelstein) address compiler issues while the next four ('A note on native level 1 BLAS in Java' by Aart J. C. Bik and Dennis B. Gannon; 'Towards automatic support of parallel sparse computation in Java with continuous compilation' by Rong- Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 1997 | Editorial: Java for computational science and engineering - simulation and modelingabstractWe are pleased to present a set of papers discussing the role of Java in Science and Engineering Simulation. These were presented at a small workshop with 45 participants at Syracuse on 16-17 December 1996. This was very successful, and a follow-up event will be sponsored by ACM in Las Vegas on 21 June 1997. The growing interest in this field is also supported by an email discussion list and other materials collected at the web site http://www.npac.syr.edu/projects/javaforcse. Java and Web technology can be used in many areas of science and engineering computation. These include sophisticated user interfaces and coarse-grain integration of different modules in complex meta-applications. However, also interesting (and controversial) is perhaps the use of Java as the language used for the computationally intense parts of a scientific code. All these areas were discussed at the workshop, with promising initial results and studies reported in each case. Again applications were described both for large-scale event-driven and time-stepped simulations and also for smaller client-side applets aimed at education. The appeal of Java as a simulation language includes its object-oriented characteristics, elegant applet software distribution model and natural support of graphical user interfaces. There are also non-technical reasons to think Java will be very important. In particular, one expects children to learn Java naturally as part of their Web experiences. On entering University, I find it hard to believe that many will be willing to switch from Java to Fortran77 or Fortran90. The papers in this issue fall into five areas. The first paper (‘Java for parallel computing and as a general language for scientific and engineering simulation and modeling’ by Geoffrey C. Fox and Wojtek Furmanski) is a general overview and the next three (‘Optimizing Java bytecodes’ by Michał Cierniak an Wei Li; ‘Optimizing Java: theory and practice’ by Zoran Budimlic and Ken Kennedy; ‘Technologies for ubiquitous supercomputing: a Java interface to the Nexus communication system’ by Ian Foster, George K. Thiruvathukal and Steven Tuecke) describe base Java technology from optimized compilation to linkage with communication infrastructure. The next two papers (‘Java simulations for physics education’ by Simeon Warner, Simon Catterall and Edward Lipson; ‘Using Java and JavaScript in the Virtual Programming Laboratory: a Web-based parallel programming environment’ by Kivanc Dincer and Geoffrey C. Fox) describe uses of Java in both science and computer science education. Then we have two papers (‘Java's role in distributed collaboration by Marina Chen and James Cowie’; ‘Java enabling collaborative education, health care, and computing’ by Lukasz Beca, Gang Cheng, Geoffrey C. Fox, Tomasz Jurga, Konrad Olszewski, Marek Podgorny, Piotr Sokolowski and Krzysztof Walczak) centred on the fascinating field of collaboration. The last six papers study the critical area of parallel and distributed computing in Java. These discuss world-wide computing (‘SuperWeb: research issues in Java-based global computing’ by Albert D. Alexandrov, Maximilian Ibel, Klaus E. Schauser and Chris J. Scheiman), large-scale software integration with Java servers (‘WebFlow – a visual programming paradigm for Web/Java based coarse grain distributed computing’ by Dimple Bhatia, Vanco Burzevski, Maja Camuseva, Geoffrey Fox, Wojtek Furmanski and Girish Premchandran) and mobility (‘Resource-aware metacomputing’ by Anurag Acharya, M. Ranganathan and Joel Saltz). These three distributed computing studies are contrasted with three on parallel computing: ‘Automatically exploiting implicit parallelism in Java’ by Aart J. C. Bik and Dennis B. Gannon on shared memory; ‘SPMD programming in Java’ by Susan Flynn Hummel, Ton Ngo and Harini Srinivasan on the SPMD style, and ‘Experiments with “HP Java”’ by Bryan Carpenter, Yuh-Jye Chang, Geoffrey Fox, Donald Leskiw and Xiaoming Li on classic distributed memory data parallelism. Currently, it appears that Java promises the computational scientist programming environments which have both attractive user interfaces and high-performance execution. An important purpose of the first workshop and the follow-up events is to get a broad input and study of the issues in this field so that we can guide the rapidly moving Java juggernaut to be maximally effective for scientific and engineering computation. © 1997 John Wiley & Sons, Ltd. Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 1997 | Java for parallel computing and as a general language for scientific and engineering simulation and modelingabstractWe discuss the role of Java and Web technologies for general simulation. We classify the classes of concurrency typical in problems and analyse separately the role of Java in user interfaces, coarse grain software integration and detailed computational kernels. We conclude that Java could become a major language for computational science, as it potentially offers good performance, excellent user interfaces and the advantages of object-oriented structure. © 1997 John Wiley & Sons, Ltd. Geoffrey C. Fox, Wojtek Furmanski |
Concurr. Pract. Exp. | 1 |
| 1997 | A Prototype of Fortran-to-Java ConverterabstractThis is a report on a prototype of a Fortran 77 to Java converter, f2j. Translation issues are identified, approaches are presented, a URL is provided for interested readers to download the package, and some unsolved problems are brought up. F2j allows value to be added to some of the investment on Fortran code – in particular, those well-established Fortran libraries for scientific and engineering computation. © 1997 John Wiley & Sons, Ltd. Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 1996 | High-Performance Fortran and Possible Extensions to Support Conjugate Gradient AlgorithmsabstractEvaluates the High Performance Fortran (HPF) language for the compact expression and efficient implementation of conjugate-gradient iterative matrix-solvers on high-performance computing and communications (HPCC) platforms. We discuss the use of intrinsic functions, data distribution directives and explicitly parallel constructs to optimize performance by minimizing communications requirements in a portable manner. We focus on implementations using the existing HPF definitions but also discuss issues arising that may influence a revised definition for HPF-2. Some of the codes discussed are available on the World Wide Web at http://www.npac.syr.edu/hpfa/, along with other educational and discussion material related to applications in HPF. Kivanç Dinçer, Geoffrey C. Fox, Kenneth A. Hawick |
HPDC | 2 |
| 1996 | Towards Web/Java-Based High Performance Distributed Computing-an Evolving Virtual MachineabstractDiscusses the emergent World Wide Web-based distributed environments for high-performance computing and communications (HPCC) on the National Information Infrastructure (NII) with the focus on Java as an enabling technology. We start with a review of the past, present and near-term future of the "Java phenomenon", exposed in the background of some related previous approaches towards a distributed interpretative virtual machine architecture. Next, we discuss the anticipated role of Java in building distributed Web-based computing environments. We outline an evolutionary path from the current Web technology "soup" towards "all-Java" systems and we illustrate this process in terms of the Northeast Parallel Architectures Center (NPAC) Web technology prototypes (WebVM, WebFlow, Bridge-based Collaboratory) and selected applications (CareWeb, 3D Visible Human). Geoffrey C. Fox, Wojtek Furmanski |
HPDC | 1 |
| 1996 | Particle-in-Cell Simulation Codes in High Performance FortranabstractParticle-in-Cell (PIC) plasma simulation codes model the interaction of charged particles with surrounding electrostatic and magnetic fields. Its computational requirements made it to be classified as one of the grand-challenge problems facing the high performance community. In this paper we present the implementation of 1-D and 2-D electrostatic PIC codes in High Performance Fortran(HPF) on a IBM SP-2. HPF expands Fortran 90 with data distribution and alignment directives and data parallel statements. It is a powerful language for writing portable and high performance programs across many platforms. We used one of the most successful commerical HPF compilers currently available in the market and augmented the compiler's missing HPF functions with extrinsic routines when necessary. We obtained near linear speed-up in all of our test cases. The performance of the HPF programs is comparable to the native message passing implementations of the same codes on the SP-2. Erol Akarsu, Kivanç Dinçer, Tomasz Haupt, Geoffrey C. Fox |
SC | 4 |
| 1996 | Building a World-Wide Virtual Machine Based on Web and HPCC TechnologiesabstractIn today's high performance computing arena, there is a strong trend toward building virtual computers from heterogeneous resources on a network. In this paper we describe our experiences in building a world-wide virtual machine (WWVM) based on emerging Web and existing HPCC technologies. We have constructed a Web-based parallel/distributed programming environment on top of this machine demonstrating MPI and PVM message-passing programs and High Performance Fortran programs. Alternatively, the WWVM can be configured as a metacomputer for the solution of metaproblems. Kivanç Dinçer, Geoffrey C. Fox |
SC | 2 |
| 1996 | Integrating multiple parallel programming paradigms in a dataflow-based software environmentabstractBy viewing different parallel programming paradigms as essentially heterogeneous approaches in mapping ‘real-world’ problems to parallel systems, the authors discuss methodologies in integrating multiple programming models on a massively parallel system such as Connection Machine CM5. Using a dataflow based integration model built in a visualization software AVS, the authors describe a simple, effective and modular way to couple sequential, data-parallel and explicit message-passing modules into an integrated parallel programming environment on a CM5. A case study in the area of numerical advection modeling is given to demonstrate the integration of data-parallel and message-passing modules in the proposed multi-paradigm programming environment. Geoffrey C. Fox |
Concurr. Pract. Exp. | 2 |
| 1996 | Benchmarking the computation and communication performance of the CM-5abstractThinking Machines' CM-5 machine is a distributed-memory, message-passing computer. In the paper we devise a performance benchmark for the base and vector units and the data communication networks of the CM-5 machine. We model the communication characteristics such as communication latency and bandwidths of point-to-point and global communication primitives. We show, on a simple Gaussian elimination code, that an accurate static performance estimation of parallel algorithms is possible by using those basic machine properties connected with computation, vectorization, communication and synchronization. Futhermore, we describe the embedding of meshes or hypercubes on the CM-5 fat-tree topology and illustrate the performance results of their basic communication primitives. Kivanç Dinçer, Zeki Bozkus, Sanjay Ranka, Geoffrey C. Fox |
Concurr. Pract. Exp. | 4 |
| 1996 | Constant Bit Rate Network Transmission of Variable Bit Rate Continuous Media in Video-On-Demand Servers
Juan Miguel del Rosario, Geoffrey C. Fox |
Multim. Tools Appl. | 2 |
| 1996 | Fast and parallel mapping algorithms for irregular problems
Chao-Wei Ou, Sanjay Ranka, Geoffrey C. Fox |
J. Supercomput. | 3 |
| 1995 | Software Tool Evaluation MethodologyabstractThe recent development of parallel and distributed computing software has introduced a variety of software tools that support several programming paradigms and languages. This variety of tools makes the selection of the best tool to run a given class of applications on a parallel or distributed system a non-trivial task that requires some investigation. We expect tool evaluation to receive more attention as the deployment and usage of distributed systems increases. In this paper, we present a multi-level evaluation methodology for parallel/distributed tools in which tools are evaluated from different perspectives. We apply our evaluation methodology to three message passing tools viz Express, p4, and PVM. The approach covers several important distributed systems platforms consisting of different computers (e.g., IBM-SP1, Alpha cluster, SUN workstations) interconnected by different types of networks (e.g., Ethernet, FDDI, ATM). Salim Hariri, Sungyong Park, Rajashekar Reddy, Mahesh Subramanyan, Rajesh Yadav, Geoffrey C. Fox, Manish Parashar |
ICDCS | 6 |
| 1995 | Distributed Information Management in the National HPCC Software Exchange
Shirley Browne, Jack J. Dongarra, Geoffrey C. Fox, Kenneth A. Hawick, Ken Kennedy, Rick L. Stevens, Robert Olson, Tom Rowan |
SC | 3 |
| 1995 | The Living Textbook and the K-12 Classroom of the FutureabstractThe Living Textbook creates a unique learning environment enabling teachers and students to use educational resources on multimedia information servers, supercomputers, parallel databases, and network testbeds. We have three innovative educational software applications running in our laboratory, and under test in the classroom. Our education-focused goal is to learn how new, learner-driven, explorative models of learning can be supported by these high bandwidth, interactive applications and ultimately how they will impact the classroom of the future. Kim Mills, Geoffrey C. Fox, Paul D. Coddington, Barbara Mihalas, Marek Podgorny, Barbara Shelly, Steven Bossert |
SC | 2 |
| 1995 | Complete exchange on the CM-5 and Touchstone Delta
Rajeev Thakur, Ravi Ponnusamy, Alok N. Choudhary, Geoffrey C. Fox |
J. Supercomput. | 4 |
| 1995 | Runtime Support and Compilation Methods for User-Specified Irregular Data DistributionsabstractThis paper describes two new ideas by which a High Performance Fortran compiler can deal with irregular computations effectively. The first mechanism invokes a user specified mapping procedure via a set of proposed compiler directives. The directives allow use of program arrays to describe graph connectivity, spatial location of array elements, and computational load. The second mechanism is a conservative method for compiling irregular loops in which dependence arises only due to reduction operations. This mechanism in many cases enables a compiler to recognize that it is possible to reuse previously computed information from inspectors (e.g., communication schedules, loop iteration partitions, and information that associates off-processor data copies with on-processor buffer locations). This paper also presents performance results for these mechanisms from a Fortran 90D compiler implementation.> Ravi Ponnusamy, Joel H. Saltz, Alok N. Choudhary, Yuan-Shin Hwang, Geoffrey C. Fox |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 1994 | The Virtual Computing EnvironmentabstractA network of supercomputers and high-performance workstations appears to be the only reasonable way to provide adequate computing resources for the Grand Challenge problems of the next century. Such a collection of computers and supporting software environments is called a virtual computing environment (VCE). The paper describes the motivation and goals of the VCE project, followed by a description of the system. The paper concentrates on the runtime aspects of the VCE, and concludes with a discussion of a small prototype system that has been built using the Isis distributed toolkit.> Philip Rousselle, Paul T. Tymann, Salim Hariri, Geoffrey C. Fox |
HPDC | 4 |
| 1994 | A Concurrent Multi Target Tracker: Benchmarking and PortabilityabstractWith the current advances in computing and network technology and software, the gap between parallel and distributed computing environment is gradually becoming narrower. Consequently, parallel programs run on parallel as well as distributed systems. However, programming and porting complex applications to such environment is challenging task and not well understood. In this paper, we use a concurrent multi target tracker as a running example to analyze and evaluate performance of two different parallel implementations on parallel and distributed systems. We have benchmarked both these implementations on different architectures that vary from a network of worksta-tions{SUN, IBM RS6000) to parallel computers (CM5, iPSC 860) using different parallel/distributed message passing tools{PVM, p4, EXPRESS). Salim Hariri, Rajesh Yadav, Balaji Thiagarajan, Sungyong Park, Mahesh Subramanyan, Rajashekar Reddy, Geoffrey C. Fox |
ICPP (3) | 7 |
| 1994 | A parallel Gauss-Seidel algorithm for sparse power system matricesabstractWe describe the implementation and performance of an efficient parallel Gauss-Seidel algorithm that has been developed for irregular, sparse matrices from electrical power systems applications. Although, Gauss-Seidel algorithms are inherently sequential, by performing specialized orderings on sparse matrices, it is possible to eliminate much of the data dependencies caused by precedence in the calculations. A two-part matrix ordering technique has been developed-first to partition the matrix into block-diagonal-bordered form using diakoptic techniques and then to multi-color the data in the last diagonal block using graph coloring techniques. The ordered matrices often have extensive parallelism, while maintaining the strict precedence relationships in the Gauss-Seidel algorithm. We present timing results for a parallel Gauss-Seidel solver implemented on the Thinking Machines CM-5 distributed memory multi-processor. The algorithm presented here requires active message remote procedure calls in order to minimize communications overhead and obtain good relative speedup.> D. P. Koester, Sanjay Ranka, Geoffrey C. Fox |
SC | 3 |
| 1994 | Interpreting the performance of HPF/Fortran 90DabstractWe present a novel interpretive approach for accurate and cost effective performance prediction in a high performance computing environment, and describe the design of a source driven HPF/Fortran 90D performance prediction framework based on this approach. The performance prediction framework has been implemented as part of a HPF/Fortran 90D application development environment. A set of benchmarking kernels and application codes are used to validate the accuracy, utility, usability, and cost effectiveness of the performance prediction framework. The use of the framework for selecting appropriate compiler directives and for application performance debugging is demonstrated.> Manish Parashar, Salim Hariri, Tomasz Haupt, Geoffrey C. Fox |
SC | 4 |
| 1994 | Communication system for high-performance distributed computingabstractAbstract With the current advances in computer and networking technology coupled with the availability of software tools for parallel and distributed computing, there has been increased interest in high‐performance distributed computing (HPDC). We envision that HPDC environments with supercomputing capabilities will be available in the near future. However, a number of issues have to be resolved before future network‐based applications can fully exploit the potential of the HPDC environment. In the paper we present an architecture for a high‐speed local area network and a communication system that provides HPDC applications with high bandwidth and low latency. We also characterize the message‐passing primitives required in HPDC applications and develop a communication protocol that implements these primitives efficiently. Salim Hariri, J.-B. Park, Manish Parashar, Geoffrey C. Fox |
Concurr. Pract. Exp. | 4 |
| 1994 | Allocating data to distributed-memory multiprocessors by genetic algorithmsabstractAbstract We present three genetic algorithms (GAs) for allocating irregular data sets to multiprocessors. These are a sequential hybrid GA, a coarse‐grain GA and a fine‐grain GA. The last two are based on models of natural evolution that are suitable for parallel implementation; they have been implemented on a hypercube and a Connection Machine. Experimental results show that the three GAs evolve good suboptimal solutions which are better than those produced by other methods. The GAs are also robust and do not show a bias towards particular problem configurations. The two parallel GAs have reasonable execution times, with the coarse‐grain GA producing better solutions for the allocation of loosely synchronous computations. Nashat Mansour, Geoffrey C. Fox |
Concurr. Pract. Exp. | 2 |
| 1994 | Hierarchical Scheduling of Dynamic Parallel Computaion on Hypercube MulticomputersabstractIn this paper a hierarchical task scheduling strategy for assigning parallel computations with dynamic structures to large hypercube multicomputers is proposed. Such computations represent a wide range of recursive and divide/conquer algorithms for which structure of the problem varies dynamically. To achieve load balancing and reduce processor contentions, the system is divided into multiple regions of processors for which the first level of scheduling is done by the host computer that spreads out the initial computations into these regions. The second level scheduling is done by a set of median processors of these regions which enable the processors of their regions to optimally balance the dynamically created load and to communicate with each other with reduced overhead. The results of an extensive simulation study are presented that exhibit the performance of the proposed strategy under different loading conditions, varying degrees of depth and parallelism, and communication costs. The proposed dual-level hierarchical scheduling is shown to outperform a well known distributed scheduling strategy. Ishfaq Ahmad 0001, Arif Ghafoor, Geoffrey C. Fox |
J. Parallel Distributed Comput. | 3 |
| 1994 | Compiling Fortran 90D/HPF for Distributed Memory MIMD ComputersabstractThis paper describes the design of the Fortran90D/HPF compiler, a source-to-source parallel compiler for distributed memory systems being developed at Syracuse University. Fortran 90D/HPF is a data parallel language with special directives to specify data alignment and distributions. A systematic methodology to process distribution directives of Fortran 90D/HPF is presented. Furthermore, techniques for data and computation partitioning, communication detection and generation, and the run-time support for the compiler are discussed. Finally, initial performance results for the compiler are presented. We believe that the methodology to process data distribution, computation partitioning, communication system design, and the overall compiler design can be used by the implementors of compilers for HPF. Zeki Bozkus, Alok N. Choudhary, Geoffrey C. Fox, Tomasz Haupt, Sanjay Ranka, Min-You Wu |
J. Parallel Distributed Comput. | 3 |
| 1994 | A Data Parallel Algorithm for Solving the Region Growing Problem on the Connection MachineabstractRegion growing is a general technique for image segmentation, where image characteristics are used to group adjacent pixels together to form regions. This paper presents a parallel algorithm for solving the region growing problem based on the split-and-merge approach, and uses it to test and compare various parallel architectures and programming models. The implementations were done on the Connection Machine, models CM-2 and CM-5, in the data parallel and message passing programming models. Randomization was introduced in breaking ties during merging to increase the degree of parallelism, and only one- and two-dimensional arrays of data were used in the implementations. Nawal Copty, Sanjay Ranka, Geoffrey C. Fox, Ravi V. Shankar |
J. Parallel Distributed Comput. | 3 |
| 1994 | Parallel physical optimization algorithms for allocating data to multicomputer nodes
Nashat Mansour, Geoffrey C. Fox |
J. Supercomput. | 2 |
| 1994 | Static and Run-Time Algorithms for All-to-Many Personalized Communication on Permutation NetworksabstractWith the advent of new routing methods, the distance that a message is sent is becoming relatively less and less important. Thus, assuming no link contention, permutation seems to be an efficient collective communication primitive. In this paper, we present several algorithms for decomposing all-to-many personalized communication into a set of disjoint partial permutations. We discuss several algorithms and study their effectiveness from the view of static scheduling as well as run-time scheduling. An approximate analysis shows that with n processors, and assuming that every processor sends and receives d messages to random destinations, our algorithm can perform the scheduling in O(dn In d) time, on average, and can use an expected number of d+log d partial permutations to carry out the communication. We present experimental results of our algorithms on the CM-5.> Sanjay Ranka, Jhy-Chun Wang, Geoffrey C. Fox |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 1993 | A Low-Latency Programming Interface and a Prototype Switch for Scalable High-Performance Distributed ComputingabstractThis paper discusses the architecture and performance of a prototype switch for interconnecting IBM RISC System/6000 workstations. The paper describes the interconnection architecture and performance on a cluster of four IBM RISC System 6000 model 340 workstations. It also describes the driver level software interface to the switch and the features incorporated to minimize communication overhead. The performance measurements cover communication latency and bandwidth. In addition, performance measurements of Express, a popular parallel-programming interface, are provided.> Taitin Chen, Jim Feeney, Geoffrey C. Fox, Gideon Frieder, Sanjay Ranka, Bill Wilhelm, Fang-Kuo Yu |
HPDC | 3 |
| 1993 | A Message Passing Interface for Parallel and Distributed ComputingabstractThe proliferation of high performance workstations and the emergence of high speed networks have attracted a lot of interest in parallel and distributed computing (PDC). The authors envision that PDC environments with supercomputing capabilities will be available in the near future. However, a number of hardware and software issues have to be resolved before the full potential of these PDC environments can be exploited. The presented research has the following objectives: (1) to characterize the message-passing primitives used in parallel and distributed computing; (2) to develop a communication protocol that supports PDC; and (3) to develop an architectural support for PDC over gigabit networks.> Salim Hariri, Jong Park, Fang-Kuo Yu, Manish Parashar, Geoffrey C. Fox |
HPDC | 5 |
| 1993 | Panel - Software Tools for High-Performance Distributed Computing
Vaidy S. Sunderam, Geoffrey C. Fox, Al Geist, William Gropp, Bob Harrison, Adam Kolawa, Michael J. Quinn, Anthony Skjellum |
HPDC | 2 |
| 1993 | Solving the Region Growing Problem on the Connection MachineabstractThis paper presents a parallel algorithm for solving the region growing problem based on the split and merge approach. The algorithm was implemented on the CM-2 and the CM-5 in the data parallel and message passing models. The performance of these implementations is examined and compared. Nawal Copty, Sanjay Ranka, Geoffrey C. Fox, Ravi V. Shankar |
ICPP (3) | 3 |
| 1993 | Graph Contraction for Physical Optimization Methods: A Quality-Cost Tradeoff for Mapping Data on Parallel ComputersabstractMapping data to parallel computers aims at minimizing the execution time of the associated application. However, it can take an unacceptable amount of time in comparison with the execution time of the application if the size of the problem is large. In this paper, first we motivate the case for graph contraction as a means for reducing the problem size. We restrict our discussion to applications where the problem domain can be described using a graph (e.g., computational fluid dynamics applications). Then we present a mapping-oriented Parallel Graph Contraction (PGC) heuristic algorithm that yields a smaller representation of the problem to which mapping is then applied. The mapping solution for the original problem is obtained by a straight-forward interpolation. We then present experimental results on using contracted graphs as inputs to two physical optimization methods; namely, Genetic Algorithm and Simulated Annealing. The experimental results show that the PGC algorithm still leads to a reasonably good quality mapping solutions to the original problem, while producing a substantial reduction in mapping time. Finally, we discuss the cost-quality tradeoffs in performing graph contraction. Nashat Mansour, Ravi Ponnusamy, Alok N. Choudhary, Geoffrey C. Fox |
International Conference on Supercomputing | 4 |
| 1993 | Fortran 90D/HPF compiler for distributed memory MIMD computers: design, implementation, and performance resultsabstract90D\HPF is a data parallel lanquage w~ih speczal directives to enable users to spectfy data a[ignment and distributions.This paper describes the design and implementation of a Fortran!)ODjHPF compiler.Techniques for data and computation partitioning, communication detect ton and generation, and the run-ttme support for the compiler are dtscussed.Finally, tn~txal performance results for the cornptler are presented.We belteve that the methodology to process data dtstributton, computation partittontngl conlmunz catton system design and the overall comptler destgn can be used by the implementors of HPF compzlers.1 Introduction Currently, distributed melmory machines are programmed using a node language and a message passing library.This process is tedious and error prone because the user must perform the task of data distribution and communication for non-local data access.There has been significant research in developing parallelizing compilers.In this approach, the compiler takes a sequential Fortran 77 program as input, applies a set of transformation rules, and produces a parallelized code for the target machine.However, a sequential language, such as Fortran 77, obscures the parallelism of a problem in sequential loops and other sequential constructs.This makes the potential parallelism of a program more difficult to detect by a parallelizing compiler.Therefore, compiling a sequential program into a parallel program is not a natural approach.An alternative approach is to use Zeki Bozkus, Alok N. Choudhary, Geoffrey C. Fox, Tomasz Haupt, Sanjay Ranka |
SC | 3 |
| 1993 | An interactive remote visualization environment for an electromagnetic scattering simulation on high performance computing systemabstractElectromagnetic scattering (EMS) simulation is an important computationally intensive application waihin the field of eleciromagnetics.Advances in high performance computing and communication (HPCC) and data visualization environ ment(DVE) provide new opportunities to visualize real-time simulation problems such as EMS which requzre stgntficant computational resources.In this work, an integrated interactive visuahzatzon environment was created for an EMS simulation, coupling a graphtcal user tnterface(GUI) for runtime simulation parameters input and 3D rendering output on a graphical workstation, with computational modu!es running on a parallel supercomputer and two workstations.Application Visualization System(AVS) was used as integrating software to facilitate both networking and scientific data visualization.Using the EMS simulation as a case study in this paper, we explore the AVS datafiow methodology to naturally integrate data visualization, parallel systems and heterogeneous computing.Major issues in integrating this remote visualization system are discussed, including task decomposition, system integration, concurrent control, and a high level DVE-based distributed programming model. Yinghua Lu, Geoffrey C. Fox, Kim Mills, Tomasz Haupt |
SC | 3 |
| 1993 | Common runtime support for high-performance parallel languagesabstractNo abstract available. Geoffrey C. Fox, Sanjay Ranka, Michael L. Scott, Allen D. Malony, James C. Browne, Marina C. Chen, Alok N. Choudhary, Thomas E. Cheatham, Janice E. Cuny, Rudolf Eigenmann, Amr F. Fahmy, Ian T. Foster, Dennis Gannon, Tomasz Haupt, Carl Kesselman, Charles Koelbel, Wei Li 0015, Monica S. Lam, Thomas J. LeBlanc, Jim Openshaw, David A. Padua, Constantine D. Polychronopoulos, Joel H. Saltz, Alan Sussman, Gil Weigand, Katherine A. Yelick |
SC | 1 |
| 1993 | Dependence analysis for outer loop parallelization of existing Fortran-77 programsabstractAbstract We have used six static parallelization tools on four Fortran‐77 programs used in physics simulations. We indicate areas where current tools have difficulties in recognizing parallelism, and illustrate these issues with simple examples. We suggest that a dynamic dependency analysis tool is needed to aid the user in the parallelization of dusty decks. Josef Stein, Geoffrey C. Fox |
Concurr. Pract. Exp. | 2 |
| 1993 | Experimental Performance Evaluation of the CM-5
Ravi Ponnusamy, Rajeev Thakur, Alok N. Choudhary, Kishore Velamakanni, Zeki Bozkus, Geoffrey C. Fox |
J. Parallel Distributed Comput. | 6 |
| 1993 | Constrained Clustering as an Optimization MethodabstractA deterministic annealing approach to clustering is derived on the basis of the principle of maximum entropy. This approach is independent of the initial state and produces natural hierarchical clustering solutions by going through a sequence of phase transitions. It is modified for a larger class of optimization problems by adding constraints to the free energy. The concept of constrained clustering is explained, and three examples are are given in which it is used to introduce deterministic annealing. The previous clustering method is improved by adding cluster mass variables and a total mass constraint. The traveling salesman problem is reformulated as constrained clustering, yielding the elastic net (EN) approach to the problem. More insight is gained by identifying a second Lagrange multiplier that is related to the tour length and can also be used to control the annealing process. The open path constraint formulation is shown to relate to dimensionality reduction by self-organization in unsupervised learning. A similar annealing procedure is applicable in this case as well.> Kenneth Rose, Eitan Gurewitz, Geoffrey C. Fox |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1992 | Allocation of Computations with Dynamic Structures on Hypercube Based Distributed SystemsabstractA dual-level dynamic load distribution strategy is proposed for allocating parallel computations with unpredictable structures to hypercube based distributed systems. Computations with dynamic structures represent a wide range of recursive and divide/conquer algorithms. The allocation strategy supports dynamic partitioning of these computations into communicating sub-tasks. Using the topological characteristics of hypercube networks, the system is divided into multiple regions of processors. The first level allocation is done by the central computer that spreads out the initial computations into these regions to reduce processor contention. The second level allocation is done by the median processors of these regions which enable the processors of their regions to optimally balance the dynamically created load and to communicate with each other with reduced overhead. The results of a simulation study are presented illustrating numerous examples that exhibit the performance of the proposed strategy under different loading conditions, varying degrees of depth and parallelism in the task graphs. The proposed allocation strategy is shown to outperform distributed load distribution.> Ishfaq Ahmad 0001, Arif Ghafoor, Geoffrey C. Fox |
HPDC | 3 |
| 1992 | A Requirement Analysis for High Performance Distributed Computing over LANsabstractWith the proliferation of high performance workstations and the current trend towards high speed communication networks. the cumulative computing power provided by a group of general purpose workstations is comparable to supercomputers. However a number of obstacles have to be overcome before the full potential of these network-based distributed systems can be exploited. This paper investigates the requirements of current workstation clusters interconnected by local area networks (LANs) which would allow them to be used as platforms for high performance distributed computing. The blocked LU decomposition of dense matrices is used as the running example in the presented study. Performance of this algorithm is measured on the iPSC/860 hypercube and on a set of homogeneous workstations (SUN SPARCstation 1+) interconnected by Ethernet. These measures are analyzed and a set of requirements are identified which would enable a network of workstations to deliver high performance distributed computing.> Manish Parashar, Salim Hariri, A. Gaber Mohamed, Geoffrey C. Fox |
HPDC | 4 |
| 1992 | On the Parallelization of Blocked LU Factorization Algorithms on Distributed Memory ArchitecturesabstractThe authors present the parallelization of blocked algorithms for LU factorization. They isolate problems inherent in sequential blocked algorithms and provide approaches to overcome them on distributed memory architectures. The performances of the parallelized versions of three blocked algorithms suited to column oriented Fortran are compared. Experiments are performed on the iPSC/860 hypercube. It is shown that it is not intuitively clear which algorithm might perform best on a given architecture; this is dependent on the problem size and the number of available parameters.> Gregor von Laszewski, Manish Parashar, A. Gaber Mohamed, Geoffrey C. Fox |
SC | 4 |
| 1992 | Scheduling Regular and Irregular Communication Patterns on the CM-5abstractThe authors study the communication characteristics of the CM-5 (Connection Machine 5) and the performance effects of scheduling regular and irregular communication patterns on the CM-5. They consider the scheduling of regular communication patterns such as complete exchange and broadcast. They have implemented four algorithms for complete exchange and studied their performances on a 2-D FFT (fast Fourier transform) algorithm. They have also implemented four algorithms for scheduling irregular communication patterns and studied their performance on the communication patterns of several synthetic as well as real problems such as the conjugate gradient solver and the Euler solver.> Ravi Ponnusamy, Rajeev Thakur, Alok N. Choudhary, Geoffrey C. Fox |
SC | 4 |
| 1992 | Allocating data to multicomputer nodes by physical optimization algorithms for loosely synchronous computationsabstractAbstract Three optimization methods derived from natural sciences are considered for allocating data to multicomputer nodes. These are simulated annealing, genetic algorithms and neural networks. A number of design choices and the addition of preprocessing and postprocessing steps lead to versions of the algorithms which differ in solution qualities and execution times. In this paper the performances of these versions are critically evaluated and compared for test cases with different features. The performance criteria are solution quality, execution time, robustness, bias and parallelizability. Experimental results show that the physical algorithms produce better solutions than those of recursive bisection methods and that they have diverse properties. Hence, different algorithms would be suitable for different applications. For example, the annealing and genetic algorithms produce better solutions and do not show a bias towards particular problem structures, but they are slower than the neural network algorithms. Preprocessing graph contraction is one of the additional steps suggested for the physical methods. It produces a significant reduction in execution time, which is necessary for their applicability to large problems. Nashat Mansour, Geoffrey C. Fox |
Concurr. Pract. Exp. | 2 |
| 1992 | Vector quantization by deterministic annealingabstractA deterministic annealing approach is suggested to search for the optimal vector quantizer given a set of training data. The problem is reformulated within a probabilistic framework. No prior knowledge is assumed on the source density, and the principle of maximum entropy is used to obtain the association probabilities at a given average distortion. The corresponding Lagrange multiplier is inversely related to the 'temperature' and is used to control the annealing process. In this process, as the temperature is lowered, the system undergoes a sequence of phase transitions when existing clusters split naturally, without use of heuristics. The resulting codebook is independent of the codebook used to initialize the iterations.> Kenneth Rose, Eitan Gurewitz, Geoffrey C. Fox |
IEEE Trans. Inf. Theory | 3 |
| 1991 | A Static Performance Estimator to Guide Data Partitioning DecisionsabstractThe choice of the data domain partitioning scheme is an important factor in determining the available parallelism and hence the performance of an application on a distributed memory multiprocessor.In this paper, we present a performance estimator for statically evaluating the relative efficiency of different data partitioning schemes for any given program on any given distributed memory multiprocessor.Our methlod is not based on a theoretical machine model, but ixnstead uses a set of kernel routinea to "train" the estimator for each target machine.We also describe a prototype implementation of this technique and discuss an experimental evaluation of its accuracy. Vasanth Balasundaram, Geoffrey C. Fox, Ken Kennedy, Ulrich Kremer |
PPoPP | 2 |
| 1991 | Two approaches to the concurrent implementation of the prime factor algorithm on a hypercubeabstractAbstract On sequential computers, the prime factor algorithm (PFA) allows the Computation of the discrete Fourier transform (DFT) with a higher efficiency than the traditional Cooley‐Tukey FFT algorithm (CTA). However, the PFA requires substantial data movement, which poses a challenging problem for distributed‐memory multi‐processor systems. In this paper, two approaches for a concurrent implementation of the PFA on these structures are presented. In the first approach, the concurrent PFA runs on all nodes of the multi‐processor system, which is inefficient on large configurations due to the large communication overhead. A second approach developed to reduce this bottleneck is also presented. These solutions have been benchmarked on Caltech hypercubes, and the performances achieved are reported. In both approaches, thecrystal_routeralgorithm was exploited as a concurrent technique for communicating data among nodes. Giovanni Aloisio, E. Lopinto, Nicola Veneziani, Geoffrey C. Fox, Jai Sam Kim |
Concurr. Pract. Exp. | 4 |
| 1991 | Physical computationabstractAbstract Physical computation embraces a variety of physical analogies used to tackle non‐traditional problems. We describe Monte Carlo and deterministic methods, including simulated annealing and neural networks. Applications include economic change in Eastern Europe, the travelling salesman problem, vehicle navigation, track finding, and parallel computer load balancing. We show how different problems are suitable for the different various approaches to optimizatton—there is no universally applicable method. Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 1991 | Achievements and prospects for parallel computingabstractAbstract Parallel computing works for the majority of large‐scale computations. The development of parallel hardware designs has been largely transferred to industry, while universities continue major research efforts Into better software environments. We describe a classification of problems and how different software models are needed for portable user‐friendly, high‐performance Implementations on parallel machines. The education of a new generation of computational scientists will be a major challenge to our universities. Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 1990 | A VLSI Neural Network for Color Constancy
John Allman, Geoffrey C. Fox, Rodney M. Goodman |
NIPS | 3 |
| 1990 | A deterministic annealing approach to clustering
Kenneth Rose, Eitan Gurewitz, Geoffrey C. Fox |
Pattern Recognit. Lett. | 3 |
| 1989 | Practical parallel supercomputing: examples from chemistry and physicsabstractWe use two large simulations, the chemical reaction dynamics of H + H2 and the collision of two galaxies to show that current parallel machines are capable of large supercomputer level calculations. We contrast the different architectural tradeoffs for these problems and draw some implications for future production parallel supercomputers. Geoffrey C. Fox, Paul G. Hipes, John K. Salmon |
SC | 1 |
| 1989 | Parallel Computing Comes of Age: Supercomputer Level Parallel Computations at CaltechabstractAbstract Parallel supercomputers are now in regular use at Caltech for several major scientific calculations. We use this experience to abstract a set of lessons for applications, decomposition, performance, hardware and software. We consider hypercubes, transputer arrays and the SIMD Connection Machine CM‐2 and AMT DAP. These are contrasted, where possible, with CRAY and other high performance conventional computers. Applications covered are lattice gauge theory, plasma physics, statistical and condensed matter physics, astronomical data analysis, quantum chemistry, graphics ray tracing, string dynamics, grain dynamics, astrophysical particle dynamics, computer chess and Kalman filters. Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 1989 | Code Generation by a Generalized Neural Network: General Principles and Elementary Examples
Geoffrey C. Fox, Jefferey G. Koller |
J. Parallel Distributed Comput. | 1 |
| 1988 | Issues in software development for concurrent computersabstractThree approaches to programming concurrent computers are described, with attention focused on MIMD machines. These approaches involve the user preparing: the whole program; large-grain objects; and fine-grain objects. Four architectures (two homogeneous and two hierarchical) are considered.> Geoffrey C. Fox |
COMPSAC | 1 |
| 1987 | Domain Decomposition in Distributed and Shared Memory Environments. I: A Uniform Decomposition and Performance Analysis for the NCUBE and JPL Mark IIIfp Hypercubes
Geoffrey C. Fox |
ICS | 1 |
| 1987 | Matrix algorithms on a hypercube I: Matrix multiplication
Geoffrey C. Fox, Steve W. Otto, Anthony J. G. Hey |
Parallel Comput. | 1 |