Alan Sussman

dblp:s/AlanSussman · DBLP profile ↗
← Back
77ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-1672-2879ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 47 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 1Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Modernizing the Introductory Computing Sequence: Integrating Parallel and Distributed Computing in CS1 and CS2
abstract
The rapid evolution of computing demands curricula that reflect modern practices, yet many CS1 and CS2 courses continue to emphasize only sequential programming. This NSF-funded project addresses that gap by designing and disseminating exemplar CS1 and CS2 courses that integrate parallel, distributed, and event-driven computing as core concepts. The materials include unplugged activities and programming labs for both C++ and Java. To ensure broad applicability and adoption, development occurred in collaboration with instructors from six diverse institutions who are now implementing the materials. Evaluation includes surveys, assignment-specific instruments, and cross-team analysis. This poster presents the project’s vision, methods, and resources, highlighting how others can adopt and adapt them to teach modern computing.
April Renee Crockett, David P. Bunde, Gerald C. Gannod, Sushil K. Prasad, Jaime Spacco, Alan Sussman, Neena Thota, Charles C. Weems, Ramachandran Vaidyanathan
SIGCSE (2)6
2026 Envisioning CS1 and CS2: The Future of Introductory Problem Solving and Programming
abstract
Computer Science education, and all education for that matter, is being disrupted by Generative AI. While there have been few truly transformational technologies similar to AI, other incremental but impactful advances have helped shape the computing ecosystem. Other recent examples include the transition to multicore systems (requiring the promotion of parallel computing from an elective topic), the shift to graphical interfaces (raising expectations for assignments and motivating the creation of Media Computation), and the emergence of object-oriented programming. In this Birds of a Feather Session, we ask the question ''How might we redesign our CS1 and CS2 courses to better prepare students for emerging and future computing paradigms while maintaining strong foundations in problem solving, programming, and computational thinking?'' Using collaborative brainstorming techniques, participants will create a list of potential future paradigms (either disruptive or incremental) that are relevant to CS1/CS2, and develop proposed roadmaps that identify how those paradigms can be leveraged as contexts for teaching the existing CS1 and CS2 courses within the CS2023 curriculum.
Gerald C. Gannod, David P. Bunde, April Renee Crockett, Alan Sussman, Sushil K. Prasad, Charles C. Weems, Ramachandran Vaidyanathan, Suzanne Matthews, Jaime Spacco
SIGCSE (2)4
2026 Modernizing the CS Introductory Sequence with Parallel and Distributed Computing (and some AI)
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, and thus it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning. Computer science education is still teaching a 20th century model of algorithmic problem solving, where sequence, branch, and loop are the only organizing principles needed for algorithms. We invest considerable time in showing how best to sequentially process large volumes of data. All computing devices that students use currently have multiple cores as well as a GPU in many cases. Most of their favorite applications use multiple cores and distributed resources. Often concurrency offers simpler solutions than sequential approaches. In this tutorial we overview key PDC concepts and provide examples of how they may naturally be incorporated in early computing classes. We lead participants through plugged and unplugged curriculum modules that have been successfully integrated and tested in existing computing classes at multiple institutions. We also discuss recent efforts at integrating AI methods, including LLMs, into early classes. In addition, we highlight other CDER activities for integration of PDC and AI into undergraduate computing curricula. Additional Information: No equipment or prior PDC experience is required, although a laptop that can run C++, Java and Python is recommended for following along with some code examples if desired.
Charles C. Weems, April Renee Crockett, David P. Bunde, Alan Sussman, Ramachandran Vaidyanathan, Sushil K. Prasad, Gerald C. Gannod, Jaime Spacco
SIGCSE (2)4
2025 Modernizing the CS Introductory Sequence with Parallel and Distributed Computing (and some AI)
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, so it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the beginning of their computing education. With all computing devices that students use having multiple cores as well as a GPU in many cases, many students' favorite applications use multiple cores and/or distributed processors. However, we are still teaching them to solve problems using only sequential thinking. Why?
Alan Sussman, Sushil K. Prasad, David P. Bunde, Jaime Spacco, Gerald C. Gannod, April Renee Crockett, Ramachandran Vaidyanathan
SIGCSE (2)1
2024 Adaptive Prefetching for Fine-grain Communication in PGAS Programs
abstract
Applications that require distributed-memory systems and exhibit irregular memory access patterns present both productivity and performance challenges. This is largely due to the fine-grain communication that arises from irregular memory accesses to distributed data. The Partitioned Global Address Space (PGAS) model provides a globally shared address space and one-sided communication, making it well suited for implementing irregular codes. However, while the PGAS model provides productivity advantages, irregular applications have difficulty achieving high performance due to the cost of fine-grain remote accesses. One way to improve the performance of such PGAS programs is to prefetch remote data before it is needed, thereby hiding communication latency. The challenges of applying prefetching are computing the prefetch distance (i.e., how far ahead to issue prefetches), determining when prefetching will be profitable, and modifying the program to perform prefetching. In this work, we present an adaptive prefetching optimization that can be applied to PGAS programs with irregular memory access patterns. We target the Chapel parallel programming language, which implements a PGAS model. Our optimization leverages runtime information to adjust the prefetch distance and to pause/resume prefetching when that is likely to provide performance gains. Furthermore, we develop compiler support to automatically apply the optimization without requiring user intervention. We evaluate the optimization across five different distributed-memory systems and four workloads. We observe runtime speed-ups of 0.5 – 377x for the Chapel workloads when compared to not prefetching, and 3 – 215x when compared to implementations written in UPC, OpenSHMEM, and one-sided MPI.
Thomas B. Rolinger, Alan Sussman
IPDPS2
2024 Integrating Parallel and Distributed Computing in Early Computing Classes
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, so it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning of their computing education. With all computing devices that students use currently having multiple cores as well as a GPU in many cases, many students' favorite applications use multiple cores and/or distributed processors. However, we are still teaching them to solve problems using only sequential thinking. Why?
Alan Sussman, Sushil K. Prasad, Charles C. Weems, Sheikh K. Ghafoor, Ramachandran Vaidyanathan
SIGCSE (2)1
2023 Integrating Parallel and Distributed Computing in Early Computing Classes
abstract
Parallel and distributed computing (PDC) has become pervasive in all aspects of computing, and thus it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning. Computer science education is still teaching to a 20th century model of algorithmic problem solving. Sequence, branch, and loop are taught in our early courses as the only organizing principles needed for algorithms, and we invest considerable time in showing how best to sequentially process large volumes of data. All computing devices that students use currently have multiple cores as well as a GPU in many cases. Most of their favorite applications use multiple cores and numbers of distributed processors. Often concurrency offers simpler solutions than sequential approaches. Industry is desperate for software engineers who think naturally in terms of exploiting these capabilities, rather than seeing them as an exotic upper-level topic that gets layered over a sequential solution. However, we are still teaching students to solve problems using sequential thinking. In this workshop we overview key PDC concepts and provide examples of how they may naturally be incorporated in early computing classes. We will introduce plugged and unplugged curriculum modules that have been successfully integrated in existing computing classes at multiple institutions. We will highlight the upcoming summer training workshop, for which we have funding to support attendance, as well as other CDER (Center for Parallel and Distributed Computing Curriculum Development and Educational Resources) activities.
Sheikh K. Ghafoor, Charles C. Weems, Alan Sussman, Ramachandran Vaidyanathan, Sushil K. Prasad
SIGCSE (2)3
2023 NSF/IEEE-TCPP Curriculum on Parallel and Distributed Computing for Undergraduates - Version II - Big Data, Energy, and Distributed Computing
abstract
This special session will report on the updated NSF/IEEE-TCPP Curriculum on Parallel and Distributed Computing released in Nov 2020 by the Center for Parallel and Distributed Computing Curriculum Development and Educational Resources (CDER). The purpose of the special session is to obtain SIGCSE community feedback on this curriculum in a highly interactive manner employing the hybrid modality and supported by a full-time CDER booth for the duration of SIGCSE. In this era of big data, cloud, and multi- and many-core systems, it is essential that the computer science (CS) and computer engineering (CE) graduates have basic skills in parallel and distributed computing (PDC). The topics are primarily organized into the areas of architecture, programming, and algorithms topics. A set of pervasive concepts that percolate across area boundaries are also identified. Version 1 of this curriculum was released in December 2012. That curriculum guideline has over 140 early adopter institutions worldwide and has been incorporated into the 2013 ACM/IEEE Computer Science curricula. This Version-II represents a major revision. The updates have focused on enhancing coverage related to the topical aspects of Big Data, Energy, and Distributed Computing.
Sushil K. Prasad, Charles C. Weems, Alan Sussman, Trilce Estrada, Ramachandran Vaidyanathan, Sheikh K. Ghafoor, Krishna Kant 0001, Craig B. Stunkel
SIGCSE (2)3
2022 VeloxDFS: Streaming Access to Distributed Datasets to Reduce Disk Seeks
Sunghwan Ahn, Hyeongjun Park, Vicente A. B. Sanchez, Deukyeon Hwang, Wonbae Kim, Alan Sussman, Beomseok Nam
CCGRID6
2022 ListDB: Union of Write-Ahead Logs and Persistent SkipLists for Incremental Checkpointing on Persistent Memory
Wonbae Kim, Chanyeol Park, Dongui Kim, Hyeongjun Park, Young-ri Choi, Alan Sussman, Beomseok Nam
OSDI6
2022 A codesign framework for online data analysis and reduction
abstract
Abstract Science applications preparing for the exascale era are increasingly exploring in situ computations comprising of simulation‐analysis‐reduction pipelines coupled in‐memory. Efficient composition and execution of such complex pipelines for a target platform is a codesign process that evaluates the impact and tradeoffs of various application‐ and system‐specific parameters. In this article, we describe a toolset for automating performance studies of composed HPC applications that perform online data reduction and analysis. We describe Cheetah, a new framework for composing parametric studies on coupled applications, and Savanna, a runtime engine for orchestrating and executing campaigns of codesign experiments. This toolset facilitates understanding the impact of various factors such as process placement, synchronicity of algorithms, and storage versus compute requirements for online analysis of large data. Ultimately, we aim to create a catalog of performance results that can help scientists understand tradeoffs when designing next‐generation simulations that make use of online processing techniques. We illustrate the design of Cheetah and Savanna, and present application examples that use this framework to conduct codesign studies on small clusters as well as leadership class supercomputers.
Kshitij Mehta, Bryce Allen, Matthew Wolf, Jeremy Logan, Eric Suchyta, Swati Singhal, Jong Choi 0001, Keichi Takahashi, Kevin A. Huck, Igor Yakushin, Alan Sussman, Todd S. Munson, Ian T. Foster, Scott Klasky
Concurr. Comput. Pract. Exp.11
2021 Optimizing Memory-Compute Colocation for Irregular Applications on a Migratory Thread Architecture
abstract
The movement of data between memory and processors has become a performance bottleneck for many applications. This is made worse for applications with sparse and irregular memory accesses, as they exhibit weak locality and make poor utilization of cache. As a result, colocating memory and compute is crucial for achieving high performance on irregular applications. There are two paradigms for memory-compute colocation. The first is the conventional approach of moving the data to the compute. The second paradigm is to move the compute to the data, which is less conventional and not as well understood. An example are migratory threads, which physically relocate upon remote accesses to the compute resource that hosts the data. In this paper, we explore the paradigm of moving compute to the data by optimizing memory-compute colocation for irregular applications on a migratory thread architecture. Our optimization method includes both initial data placement as well as data replication. We evaluate our optimization on sparse matrix-vector multiply (SpMV) and sparse matrix-matrix multiply (SpGEMM). Our results show that we can achieve speed-ups as high as 4.2x on SpMV and 6x on SpGEMM when compared to the default data layout. We also highlight that our optimization to improve memory-compute colocation can be applicable to both migratory threads and more conventional systems. To this end, we evaluate our optimization approach on a conventional compute cluster using the Chapel programming language. We demonstrate speed-ups as high as 18x for SpMV.
Thomas B. Rolinger, Christopher D. Krieger, Alan Sussman
IPDPS3
2019 Modernizing Early CS Courses with Parallel and Distributed Computing
abstract
Parallel and distributed computing (PDC) is now a pervasive aspect of deployed systems, and thus it is essential that students include parallelism and distribution in the computational thinking that they apply to problem solving, from the very beginning. Our students all have multicore laptops. Most of their favorite applications use vast numbers of distributed processors. Why are we still teaching them to solve problems using only sequential thinking? Come to this workshop to see how easy it is to open their eyes to exploiting concurrency in problem solving, starting in their earliest courses. You'll hear about and experience some unplugged activities, learn how to help students recognize examples of concurrency in the world around them, see how event driven user interfaces can easily exemplify issues related to multithreading, and how freely available libraries can be used to naturally exploit parallelism in working with large data structures. We will also highlight the two summer training programs that we are organizing, for which we have funding to support attendance by instructors. Having a laptop that can run Java and C++ will allow you to follow along with some code examples, but isn't necessary.
Sushil K. Prasad, Sheikh K. Ghafoor, Charles C. Weems, Alan Sussman
SIGCSE4
2016 Coordinated Collaborative Testing of Shared Software Components
abstract
Software developers commonly build their software systems by reusing other components developed and maintained by third-party developer groups. As the components evolve over time, new end-user machine configurations that contain new component versions will be added continuously for the potential user base. Therefore developers must test whether their components function correctly in the new configurations to ensure the quality of the overall systems. This would be achievable if developers could provision the configurations in house and conduct regression testing over the configurations. However, this is often very time-consuming and also there can be redundancy in test effort between developers when a common set of components is reused for providing the functionality of the systems. In this paper, we present a coordinated collaborative regression testing process for multiple developer groups. It involves a scheduling method for distributing test effort across the groups at component updates, with the objectives of reducing test redundancy between the groups and also shortening the time window in which compatibility faults are exposed to user community. The process is implemented on Conch, a collaborative test data repository and services we developed in our previous work. Conch has been modified to function as the test process coordinator, as well as the shared repository of test data. Our experiments over the 1.5-year evolution history of eleven components in the Ubuntu developer community show that developers can quickly discover compatibility faults by applying the coordinated process. Moreover, total testing time is comparable to the scenario where the developers conduct regression testing only at updates of their own components.
Teng Long 0003, Il-Chul Yoon, Adam A. Porter, Atif M. Memon, Alan Sussman
ICST5
2015 A Ping Too Far: Real World Network Latency Measurement
abstract
There is risk when doing experiments on real world systems and making real world measurements. In a theoretical system, a simulation, or a test bed, we can make assumptions about real world behaviors and gloss over potential problems. In a real world system, however, we have to solve those problems to get science done. Sometimes these are technical problems, like making sure a tool chain is consistent across many hosts. Sometimes these are policy problems, such as obtaining permission and access. The other users of a system have a reasonable expectation that their work will not be disrupted, so researchers studying those systems must be careful accommodate that expectation. Sometimes, systems researchers encounter these problems and a study must be changed or canceled, though the problem may have nothing to do with the study itself. Our goal was to make a high-quality all-to-all network latency map that captures features not present in existing latency data sets. We solved most of the technical problems needed to make the required measurements. We solved most policy problems too, often with complex technical solutions to get around a policy obstacle without creating disruptions. However, we eventually ran out of technical solutions: we ran afoul of a policy obstacle to a technical solution to a technical problem. There is no one to blame, and no one has done anything obviously wrong: the policy decision was a conservative one designed to protect the data of the regular users of the system. Although our resulting latency data set is not complete, we believe it still has value for various network applications. Our goal was to make a high-quality all-to-all network latency map that captures features not present in existing latency data sets. We solved most of the technical problems needed to make the required measurements. We solved most policy problems too, often with complex technical solutions to get around a policy obstacle without creating disruptions. However, we eventually ran out of technical solutions: we ran afoul of a policy obstacle to a technical solution to a technical problem. There is no one to blame, and no one has done anything obviously wrong: the policy decision was a conservative one designed to protect the data of the regular users of the system. Although our resulting latency data set is not complete, we believe it still has value for various network applications.
Gary Jackson, Peter J. Keleher, Alan Sussman
e-Science3
2014 Decentralized Scheduling and Load Balancing for Parallel Programs
abstract
We present a completely decentralized algorithm for parallel job scheduling and load balancing in distributed peer-to-peer environments. This algorithm is useful for meta-scheduling across known clusters and scheduling on desktop grids. To accomplish this, we build on previous work to route jobs to appropriate resources then use the new algorithm to start parallel jobs and balance load across the grid. We also discuss what constitutes useful clustering's for this algorithm as well as inherent scaling limitations. Ultimately, we show that our algorithm performs comparably to one using centralized load balancing with global up-to-date information. The principal contribution of this work is that the parallel job scheduling is completely decentralized, which is not featured in previous work, and enables reliable ad hoc sharing of distributed resources to run parallel computations.
Gary Jackson, Peter J. Keleher, Alan Sussman
CCGRID3
2014 NSF/IEEE-TCPP curriculum initiative on parallel and distributed computing: core topics for undergraduates (abstract only)
abstract
Parallelism pervades all aspects of modern computing, from in-home devices such as cell phones to large-scale supercomputers. Recognizing this - and motivated by the premise that every undergraduate student in a computer-related field should be prepared to cope with parallel computing - a working group sponsored by NSF and IEEE/TCPP, and interacting with the ACM CS2013 initiative, has developed guidelines for assimilating parallel and distributed computing (PDC) into the core undergraduate curriculum. Over 100 Early-Adopter institutions worldwide are currently modifying their computer-related curricula in response to the guidelines. Additionally, the CDER Center for Curriculum Development and Educational Resources, which grew out of the working group, is currently assembling a book of contributed essays on how to teach PDC topics in lower-level CS/CE courses, to fill the serious lack of textual material for students and instructors.
Sushil K. Prasad, Almadena Yu. Chtchelkanova, Arnold L. Rosenberg, Alan Sussman
SIGCSE5
2014 PStore: an efficient storage framework for managing scientific data
abstract
In this paper, we present the design, implementation, and evaluation of PStore, a no-overwrite storage framework for managing large volumes of array data generated by scientific simulations. PStore consists of two modules, a data ingestion module and a query processing module, that respectively address two of the key challenges in scientific simulation data management. The data ingestion module is geared toward handling the high volumes of simulation data generated at a very rapid rate, which often makes it impossible to offload the data onto storage devices; the module is responsible for selecting an appropriate compression scheme for the data at hand, chunking the data, and then compressing it before sending it to the storage nodes. On the other hand, the query processing module is in charge of efficiently executing different types of queries over the stored data; in this paper, we specifically focus on dicing (also called range) queries. PStore provides a suite of compression schemes that leverage, and in some cases extend, existing techniques to provide support for diverse scientific simulation data. To efficiently execute queries over such compressed data, PStore adopts and extends a two-level chunking scheme by incorporating the effect of compression, and hides expensive disk latencies for long running range queries by exploiting chunk prefetching. In addition, we also parallelize the query processing module to further speed up execution. We evaluate PStore on a 140 GB dataset obtained from real-world simulations using the regional climate model CWRF [5]. In this paper, we use both 3D and 4D datasets and demonstrate high performance through extensive experiments.
Souvik Bhattacherjee, Amol Deshpande, Alan Sussman
SSDBM3
2014 Exploiting multi-core nodes in peer-to-peer grids
Jaehwan Lee 0001, Peter J. Keleher, Alan Sussman
J. Parallel Distributed Comput.3
2014 Decentralized multi-attribute range search for resource discovery and load balancing
Jaehwan Lee 0001, Peter J. Keleher, Alan Sussman
J. Supercomput.3
2013 Decentralized Preemptive Scheduling Across Heterogeneous Multi-core Grid Resources
Arun Balasubramanian, Alan Sussman, Norman M. Sadeh
JSSPP2
2013 Testing component compatibility in evolving configurations
Il-Chul Yoon, Alan Sussman, Atif M. Memon, Adam A. Porter
Inf. Softw. Technol.2
2012 Overlap and Synergy in Testing Software Components across Loosely Coupled Communities
abstract
Component integration rather than from-scratch programming increasingly defines software development. As a result software developers often play diverse roles, including that of a component provider -- packaging a component for others to use, a component user -- integrating other providers' components into their software, and a component tester -- ensuring that other providers' components work as part of an integrated system. In this paper, we explore the conjecture that we can better utilize testing resources by focusing not just on individual component-based systems, but on groups of systems that form what we refer to as loosely-coupled software development communities, meaning a set of independently-managed systems that use many of the same components. We demonstrate that such communities do in fact exist, and that there are overlaps and synergies in their test efforts.
Teng Long 0003, Il-Chul Yoon, Adam A. Porter, Alan Sussman, Atif M. Memon
ISSRE4
2012 DEMB: Cache-Aware Scheduling for Distributed Query Processing
Youngmoon Eom, Alan Sussman, Beomseok Nam
JSSPP3
2012 Analyzing design choices for distributed multidimensional indexing
Beomseok Nam, Alan Sussman
J. Supercomput.2
2011 Supporting Computing Element Heterogeneity in P2P Grids
abstract
We propose resource discovery and load balancing techniques to accommodate computing nodes with many types of computing elements, such as multi-core CPUs and GPUs, in a peer-to-peer desktop grid architecture. Heterogeneous nodes can have multiple types of computing elements, and the performance and characteristics of each computing element can be very different. Our scheme takes into account these diverse aspects of heterogeneous nodes to maximize overall system throughput. However, straightforward methods of handling diverse computing elements that differ on many axes can result in high overheads, both in local state and in communication volume. We describe approaches that minimize messaging costs without sacrificing the failure resilience provided by an underlying peer-to-peer overlay network. Simulation results show that our scheme's load balancing performance is comparable to that of a centralized approach, that communication costs are reduced significantly compared to the existing system, and that failure resilience is not compromised.
Jaehwan Lee 0001, Peter J. Keleher, Alan Sussman
CLUSTER3
2011 Automatic Computer System Characterization for a Parallelizing Compiler
abstract
Effectively utilizing the compute power of modern multi-core machines is a challenging task for a programmer. Automated extraction of shared memory parallelism via powerful compiler transformations and optimizations is one means to such a goal. However, the effectiveness of such transformations is tied to detailed characteristics of the target computer system. In this paper, we describe an automated system for capturing such computer system characteristics that is based on prior work on various parts of the overall problem. The system characteristics measured include the number of available compute elements available to run threads, multiple memory hierarchy parameters, and functional unit latencies and bandwidths. We show experimental results on a wide range of compute platforms that validate the effectiveness of the overall approach.
Alan Sussman, Norman Lo
CLUSTER1
2011 Searching for Bandwidth-Constrained Clusters
abstract
Data-intensive distributed applications can increase their performance by running on a cluster of hosts connected via high-bandwidth interconnections. However, there is no effective method to find such a bandwidth-constrained cluster in a decentralized fashion. Our work is inspired by prior work that treats Internet bandwidth as an approximate tree metric space. This paper presents a decentralized, accurate, and efficient method to find a cluster of Internet hosts, given the desired cluster size and minimum interconnection bandwidth. We describe a centralized polynomial time algorithm for a tree metric space, along with a proof of correctness. We then provide a decentralized version of the algorithm. Simulation experiments with two real-world datasets confirm that our clustering approach achieves high accuracy and scalability. We also discuss the costs of decentralization and how the treeness of the dataset affects clustering accuracy.
Sukhyun Song, Peter J. Keleher, Alan Sussman
ICDCS3
2011 Decentralized, accurate, and low-cost network bandwidth prediction
abstract
The distributed nature of modern computing makes end-to-end prediction of network bandwidth increasingly important. Our work is inspired by prior work that treats the Internet and bandwidth as an approximate tree metric space. This paper presents a decentralized, accurate, and low cost system that predicts pairwise bandwidth between hosts. We describe an algorithm to construct a distributed tree that embeds bandwidth measurements. The correctness of the algorithm is provable when driven by precise measurements. We then describe three novel heuristics that achieve high accuracy for predicting bandwidth even with imprecise input data. Simulation experiments with a real-world dataset confirm that our approach shows high accuracy with low cost.
Sukhyun Song, Peter J. Keleher, Bobby Bhattacharjee, Alan Sussman
INFOCOM4
2011 NSF/IEEE-TCPP curriculum initiative on parallel and distributed computing: core topics for undergraduates
abstract
No abstract available.
Sushil K. Prasad, Almadena Yu. Chtchelkanova, Sajal K. Das 0001, Frank Dehne, Mohamed G. Gouda, Joseph F. JáJá, Krishna Kant 0001, Anita La Salle, Richard LeBlanc, Manish Lumsdaine, David A. Padua, Manish Parashar, Viktor Prasanna 0001, Yves Robert, Arnold L. Rosenberg, Sartaj Sahni, Behrooz A. Shirazi, Alan Sussman, Charles C. Weems, Jie Wu 0001
SIGCSE19
2010 Decentralized resource management for multi-core desktop grids
abstract
The majority of CPUs now sold contain multiple computing cores. However, current desktop grid computing systems either ignore the multiplicity of cores, or treat them as distinct, independent machines. The latter approach ignores the resource contention present between cores in a single CPU, while the former approach fails to take advantage of significant computing power. We propose a decentralized resource management framework for exploiting multi-core nodes in peer-to-peer grids. We present two new load-balancing schemes that explicitly account for the resource sharing and contention of multiple cores, and propose a simple simulation model that can represent a continuum of resource sharing among cores of a CPU. We use simulation to confirm that our two algorithms match jobs w ith computing nodes efficiently, and balance load during the lifetime of the computing jobs.
Jaehwan Lee 0001, Peter J. Keleher, Alan Sussman
IPDPS3
2010 Brief Announcement: Decentralized Network Bandwidth Prediction
Sukhyun Song, Peter J. Keleher, Bobby Bhattacharjee, Alan Sussman
DISC4
2010 Multiple query scheduling for distributed semantic caches
Beomseok Nam, Minho Shin, Henrique Andrade, Alan Sussman
J. Parallel Distributed Comput.4
2009 Prioritizing component compatibility tests via user preferences
abstract
Many software systems rely on third-party components during their build process. Because the components are constantly evolving, quality assurance demands that developers perform compatibility testing to ensure that their software systems build correctly over all deployable combinations of component versions, also called configurations. However, large software systems can have many configurations, and compatibility testing is often time and resource constrained. We present a prioritization mechanism that enhances compatibility testing by examining the ldquomost importantrdquo configurations first, while distributing the work over a cluster of computers. We evaluate our new approach on two large scientific middleware systems and examine tradeoffs between the new prioritization approach and a previously developed lowest-cost-configuration-first approach.
Il-Chul Yoon, Alan Sussman, Atif M. Memon, Adam A. Porter
ICSM2
2008 Effective and scalable software compatibility testing
abstract
Today's software systems are typically composed of multiple components, each with different versions. Software compatibility testing is a quality assurance task aimed at ensuring that multi-component based systems build and/or execute correctly across all their versions' combinations, or configurations. Because there are complex and changing interdependencies between components and their versions, and because there are such a large number of configurations, it is generally infeasible to test all potential configurations. Consequently, in practice, compatibility testing examines only a handful of default or popular configurations to detect problems; as a result costly errors can and do escape to the field.
Il-Chul Yoon, Alan Sussman, Atif M. Memon, Adam A. Porter
ISSTA2
2008 Trade-offs in matching jobs and balancing load for distributed desktop grids
Jik-Soo Kim, Beomseok Nam, Peter J. Keleher, Michael A. Marsh, Bobby Bhattacharjee, Alan Sussman
Future Gener. Comput. Syst.6
2007 Using content-addressable networks for load balancing in desktop grids
abstract
Desktop grids have evolved to combine Peer-to-Peer and Grid computing techniques to improve the robustness, reliability and scalability of job execution infrastructures. However, efficiently matching incoming jobs to available system resources and achieving good load balance in a fully decentralized and heterogeneous computing environment is a challenging problem. In this paper, we extend our prior work with a new decentralized algorithm for maintaining approximate global load information, and a job pushing mechanism that uses the global information to push jobs towards underutilized portions of the system. The resulting system more effectively balances load and improves overall system throughput. Through a comparative analysis of experimental results across different system configurations and job profiles, performed via simulation, we show that our system can reliably execute Grid applications on a distributed set of resources both with low cost and with good load balance.
Jik-Soo Kim, Peter J. Keleher, Michael A. Marsh, Bobby Bhattacharjee, Alan Sussman
HPDC5
2007 Creating a Robust Desktop Grid using Peer-to-Peer Services
abstract
The goal of the work described in this paper is to design and build a scalable infrastructure for executing grid applications on a widely distributed set of resources. Such grid infrastructure must be decentralized, robust, highly available, and scalable, while efficiently mapping application instances to available resources in the system. However, current desktop grid computing platforms are typically based on a client-server architecture, which has inherent shortcomings with respect to robustness, reliability and scalability. Fortunately, these problems can be addressed through the capabilities promised by new techniques and approaches in peer-to-peer (P2P) systems. By employing P2P services, our system allows users to submit jobs to be run in the system and to run jobs submitted by other users on any resources available in the system, essentially allowing a group of users to form an ad-hoc set of shared resources. The initial target application areas for the desktop grid system are in astronomy and space science simulation and data analysis.
Jik-Soo Kim, Beomseok Nam, Michael A. Marsh, Peter J. Keleher, Bobby Bhattacharjee, Derek Richardson, Dennis Wellnitz, Alan Sussman
IPDPS8
2007 Taking Advantage of Collective Operation Semantics for Loosely Coupled Simulations
abstract
Although a loosely coupled component-based framework offers flexibility and versatility for building and deploying large-scale multi-physics simulation systems, the performance of such a system can suffer from excessive buffering of data objects which may or may not be transferred between components. By taking advantages of the collective properties of parallel simulation components, which is common for data-parallel scientific applications, an optimization method, which we call buddy-help, can greatly enhance overall performance. Buddy-help can reduce the time taken for buffering operations in an exporting component, when there are timing differences across processes in the exporting component. The optimization enables skipping unnecessary buffering operations, once another process, which has already performed the collective export operation, has determined that the buffered data would never be needed. Because an analytical study would be very difficult due to the complexity of the overall coupled simulation system, the performance improvement enabled by buddy-help is investigated via a micro-benchmark specifically designed to illustrate the behavior of coupled simulation scenarios under which buddy-help can provide performance gains.
Joe Shang-Chieh Wu, Alan Sussman
IPDPS2
2007 Direct-dependency-based software compatibility testing
abstract
Software compatibility testing is an important quality assurance task aimed at ensuring that component-based software systems build and/or execute properly across a broad range of user system configurations. Because each configuration can involve multiple components with different versions, and because there are complex and changing interdependencies between components and their versions, it is generally infeasible to test all potential configurations. Therefore, compatibility testing usually means examining only a handful of default or popular configurations to detect problems, and as a result costly errors can and do escape to the field
Il-Chul Yoon, Alan Sussman, Atif M. Memon, Adam A. Porter
ASE2
2007 Principles for designing data-/compute-intensive distributed applications and middleware systems for heterogeneous environments
Jik-Soo Kim, Henrique Andrade, Alan Sussman
J. Parallel Distributed Comput.3
2007 Active semantic caching to optimize multidimensional data analysis in parallel and distributed environments
Henrique Andrade, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
Parallel Comput.3
2006 Improving Resiliency Using Capacity-Aware Multicast Tree in P2P-Based Streaming Environments
Eunseok Kim, Jiyong Jang, Sungyoung Park, Alan Sussman, Jae Soo Yoo
HPCC4
2006 DiST: fully decentralized indexing for querying distributed multidimensional datasets
abstract
Grid computing and peer-to-peer (P2P) systems are emerging as new paradigms for managing large scale distributed resources across wide area networks. While grid computing focuses on managing heterogeneous resources and relies on centralized managers for resource and data discovery, P2P systems target scalable, decentralized methods for publishing and searching for data. In large distributed systems, a centralized resource manager is a potential performance bottleneck and decentralization can help avoid this bottleneck, as is done in P2P systems. However, the query functionality provided by most existing P2P systems is very rudimentary, and is not directly applicable to grid resource management. In this paper, we propose a fully decentralized multidimensional indexing structure, called DiST, that operates in a fully distributed environment with no centralized control. In DiST, each data server only acquires information about data on other servers from executing and routing queries. We describe the DiST algorithms for maintaining the decentralized network of data servers, including adding and deleting servers, the query routing algorithm, and failure recovery algorithms. We also evaluate the performance of the decentralized scheme against a more structured hierarchical indexing scheme that we have previously shown to perform well in distributed grid environments.
Beomseok Nam, Alan Sussman
IPDPS2
2006 Data management and query - Multiple range query optimization with distributed cache indexing
abstract
MQO is a distributed multiple query processing middleware that can use resources available on the Grid to optimize query processing for data analysis and visualization applications. It does so by introducing one or more proxies that act as front-ends to a collection of backend servers. The basic idea behind this architecture is active semantic caching, whereby queries can leverage available cached results in the proxy either directly or through transformations. While this approach has been shown to speed up query evaluation under multi-client workloads, the caching infrastructure in the backend servers is not used well for query processing. Because this collective caching infrastructure scales with the number of servers, it is an important asset. In this paper, we describe a distributed multidimensional indexing scheme that enables the proxy to directly consider the cache contents available at the backend servers for query planning and scheduling. This approach is shown to produce better query plans and faster query response times as we experimentally demonstrate.
Beomseok Nam, Henrique Andrade, Alan Sussman
SC3
2006 Data redistribution and remote method invocation for coupled components
Felipe Bertrand, Randall Bramley, David E. Bernholdt, James Arthur Kohl, Alan Sussman, Jay Walter Larson, Kostadin Damevski
J. Parallel Distributed Comput.5
2005 Spatial indexing of distributed multidimensional datasets
abstract
While declustering methods for distributed multidimensional indexing of large datasets have been researched widely in the past, replication techniques for multidimensional indexes have not been investigated deeply. In general, a centralized index server may become the performance bottleneck in a wide area network rather than the data servers, since the index is likely to be accessed more often than any of the datasets in the servers. In this paper, we present two different multidimensional indexing algorithms for a distributed environment - a centralized global index and a two-level hierarchical index. Our experimental results show that the centralized scheme does not scale well for either insertion or searching the index. In order to improve the scalability of the index server, we have employed a replication protocol for both the centralized and two-level index schemes that allows some inconsistency between replicas without affecting correctness. Our experiments show that the two-level hierarchical index scheme shows better scalability for both building and searching the index than the non-replicated centralized index, but replication can make the centralized index faster than the two-level hierarchical index for searching in some cases.
Beomseok Nam, Alan Sussman
CCGRID2
2005 Query planning for the grid: adapting to dynamic resource availability
abstract
The availability of massive datasets, comprising sensor measurements or the results of scientific simulations, has had a significant impact on the methodology of scientific reasoning. Scientists require storage, bandwidth and computational capacity to query and analyze these datasets, to understand physical phenomena or to test hypotheses. This paper addresses the challenge of identifying and selecting resources to develop an evaluation plan for large scale data analysis queries when data processing capabilities and datasets are dispersed across nodes in one or more computing and storage clusters. We show that generating an optimal plan is hard and we propose heuristic techniques to find a good choice of resources. We also consider heuristics to cope with dynamic resource availability; in this situation we have stale information about reusable cached results (datasets) and the load on various nodes.
Henrique Andrade, Louiqa Raschid, Alan Sussman
CCGRID4
2005 A simulation and data analysis system for large-scale, data-driven oil reservoir simulation studies
abstract
Abstract The main goal of oil reservoir management is to provide more efficient, cost‐effective and environmentally safer production of oil from reservoirs. Numerical simulations can aid in the design and implementation of optimal production strategies. However, traditional simulation‐based approaches to optimizing reservoir management are rapidly overwhelmed by data volume when large numbers of realizations are sought using detailed geologic descriptions. In this paper, we describe a software architecture to facilitate large‐scale simulation studies, involving ensembles of long‐running simulations and analysis of vast volumes of output data. Copyright © 2005 John Wiley & Sons, Ltd.
Tahsin M. Kurç, Ümit V. Çatalyürek, Joel H. Saltz, Ryan Martino, Mary F. Wheeler, Malgorzata Peszynska, Alan Sussman, Christian Hansen 0002, Mrinal K. Sen, Roustam Seifoullaev, Paul L. Stoffa, Carlos Torres-Verdín, Manish Parashar
Concurr. Pract. Exp.8
2004 Time and space optimization for processing groups of multi-dimensional scientific queries
abstract
Data analysis applications in areas as diverse as remote sensing and telepathology require operating on and processing very large datasets. For such applications to execute efficiently, careful attention must be paid to the storage, retrieval, and manipulation of the datasets. This paper addresses the optimizations performed by a high performance database system that processes groups of data analysis requests for these applications, which we call queries. The system performs end-to-end processing of the requests, formulated as PostgreSQL declarative queries. The queries are converted into imperative descriptions, multiple imperative descriptions are merged into a single execution plan, the plan is optimized to decrease execution time via common compiler optimization techniques, and, finally, the plan is optimized to decrease memory consumption. The last two steps are experimentally shown to effectively reduc the amount of time required while conserving memory space as a group of queries is processed by the database.
Suresh Aryangat, Henrique Andrade, Alan Sussman
ICS3
2004 A Comparative Study of Spatial Indexing Techniques for Multidimensional Scientific Datasets
Beomseok Nam, Alan Sussman
SSDBM2
2004 Optimizing the Execution of Multiple Data Analysis Queries on Parallel and Distributed Environments
abstract
We investigate techniques for efficiently executing multiquery workloads from data and computation-intensive applications in parallel and/or distributed computing environments. In this context, we describe a database optimization framework that supports data and computation reuse, query scheduling, and active semantic caching to speed up the evaluation of multiquery workloads. Its most striking feature is the ability of optimizing the execution of queries in the presence of application-specific constructs by employing a customizable data and computation reuse model. Furthermore, we discuss how the proposed optimization model is flexible enough to work efficiently irrespective of the parallel/distributed environment underneath. In order to evaluate the proposed optimization techniques, we present experimental evidence using real data analysis applications. For this purpose, a common implementation for the queries under study was provided according to the database optimization framework and deployed on top of three distinct experimental configurations: a shared memory multiprocessor, a cluster of workstations, and a distributed computational Grid-like environment.
Henrique Andrade, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
IEEE Trans. Parallel Distributed Syst.3
2003 Improving Access to Multi-dimensional Self-describing Scientific Dataset
abstract
Applications that query into very large multidimensional datasets are becoming more common. Many self-describing scientific data file formats have also emerged, which have structural metadata to help navigate the multi-dimensional arrays that are stored in the files. The files may also contain application-specific semantic metadata. In this paper, we discuss efficient methods for performing searches for subsets of multi-dimensional data objects, using semantic information to build multidimensional indexes, and group data items into properly sized chunks to maximize disk I/O bandwidth. This work is the first step in the design and implementation of a generic indexing library that will work with various high-dimension scientific data file formats containing semantic information about the stored data. To validate the approach, we have implemented indexing structures for NASA remote sensing data stored in the HDF format with a specific schema (HDF-EOS), and show the performance improvements that are gained from indexing the datasets, compared to using the existing HDF library for accessing the data.
Beomseok Nam, Alan Sussman
CCGRID2
2003 A high performance multi-perspective vision studio
abstract
We describe a multi-perspective vision studio as a flexible high performance framework for solving complex image processing and machine vision problems on multi-view image sequences. The studio abstracts multi-view image data from image sequence acquisition facilities, stores and catalogs sequences in a high performance distributed database, allows customization of back-end processing services, and can serve custom client applications, thus helping make multi-view video sequence processing efficient and generic. To illustrate our approach, we describe two multi-perspective studio applications, and discuss performance and scalability results.
Eugene Borovikov, Alan Sussman
ICS2
2003 The virtual microscope
abstract
We present the design and implementation of the Virtual Microscope, a software system employing a client/server architecture to provide a realistic emulation of a high power light microscope. The system provides a form of completely digital telepathology, allowing simultaneous access to archived digital slide images by multiple clients. The main problem the system targets is storing and processing the extremely large quantities of data required to represent a collection of slides. The Virtual Microscope client software runs on the end user's PC or workstation, while database software for storing, retrieving and processing the microscope image data runs on a parallel computer or on a set of workstations at one or more potentially remote sites. We have designed and implemented two versions of the data server software. One implementation is a customization of a database system framework that is optimized for a tightly coupled parallel machine with attached local disks. The second implementation is component-based, and has been designed to accommodate access to and processing of data in a distributed, heterogeneous environment. We also have developed caching client software, implemented in Java, to achieve good response time and portability across different computer platforms. The performance results presented show that the Virtual Microscope systems scales well, so that many clients can be adequately serviced by an appropriately configured data server.
Ümit V. Çatalyürek, Michael D. Beynon, Chialin Chang, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
IEEE Trans. Inf. Technol. Biomed.5
2002 Multiple Query Optimization for Data Analysis Applications on Clusters of SMPs
abstract
This paper is concerned with the efficient execution of multiple query workloads on a cluster of SMPs. We target applications that access and manipulate large scientific datasets. Queries in these applications involve user-defined processing operations and distributed data structures to hold intermediate and final results. Our goal is to implement system components to leverage previously computed query results and to effectively utilize processing power and aggregated I/O bandwidth on SMP nodes so that both single queries and multi-query batches can be efficiently executed.
Henrique Andrade, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
CCGRID3
2002 Active Proxy-G: optimizing the query execution process in the grid
abstract
The Grid environment facilitates collaborative work and allows many users to query and process data over geographically dispersed data repositories. Over the past several years, there has been a growing interest in developing applications that interactively analyze datasets, potentially in a collaborative setting. We describe the Active Proxy-G service that is able to cache query results, use those results for answering new incoming queries, generate subqueries for the parts of a query that cannot be produced from the cache, and submit the subqueries for final processing at application servers that store the raw datasets. We present an experimental evaluation to illustrate the effects of various design tradeoffs. We also show the benefits that two real applications gain from using the middleware.
Henrique Andrade, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
SC3
2002 Executing multiple pipelined data analysis operations in the grid
abstract
Processing of data in many data analysis applications can be represented as an acyclic, coarse grain data flow, from data sources to the client. This paper is concerned with scheduling of multiple data analysis operations, each of which is represented as a pipelined chain of processing on data. We define the scheduling problem for effectively placing components onto Grid resources, and propose two scheduling algorithms. Experimental results are presented using a visualization application.
Matthew Spencer, Renato Ferreira 0001, Michael D. Beynon, Tahsin M. Kurç, Ümit V. Çatalyürek, Alan Sussman, Joel H. Saltz
SC6
2002 Optimizing execution of component-based applications using group instances
Michael D. Beynon, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
Future Gener. Comput. Syst.3
2002 Processing large-scale multi-dimensional data in parallel and distributed environments
Michael D. Beynon, Chialin Chang, Ümit V. Çatalyürek, Tahsin M. Kurç, Alan Sussman, Henrique Andrade, Renato Ferreira 0001, Joel H. Saltz
Parallel Comput.5
2001 Optimizing Execution of Component-based Applications using Group Instances
abstract
Research on programming models for developing applications in the Grid has proposed component-based models as a viable approach, in which an application is composed of multiple interacting computational objects. We have been developing a framework, called filter-stream programming, for building data-intensive applications that query, analyze and manipulate very large data sets in a distributed environment. In this model, the processing structure of an application is represented as a set of processing units, referred to as filters. We develop the problem of scheduling instances of a filter group. A filter group is a set of filters collectively performing a computation for an application. In particular we seek the answer to the following question: should a new instance be created, or an existing one reused? We experimentally investigate the effects of instantiating multiple filter groups on performance under varying application characteristics.
Michael D. Beynon, Alan Sussman, Tahsin M. Kurç, Joel H. Saltz
CCGRID2
2001 Efficient execution of multiple query workloads in data analysis applications
abstract
Applications that analyze, mine, and visualize large datasets are considered an important class of applications in many areas of science, engineering, and business. Queries commonly executed in data analysis applications often involve user-defined processing of data and application-specific data structures. If data analysis is employed in a collaborative environment, the data server should execute multiple such queries simultaneously to minimize the response time to clients. In this paper we present the design of a runtime system for executing multiple query workloads on a shared-memory machine. We describe experimental results using an application for browsing digitized microscopy images.
Henrique Andrade, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
SC3
2001 Distributed processing of very large datasets with DataCutter
Michael D. Beynon, Tahsin M. Kurç, Ümit V. Çatalyürek, Chialin Chang, Alan Sussman, Joel H. Saltz
Parallel Comput.5
2000 Optimizing Retrieval and Processing of Multi-Dimensional Scientific Datasets
abstract
We have developed the Active Data Repository (ADR), an infrastructure that integrates storage, retrieval, and processing of large multi-dimensional scientific datasets on distributed memory parallel machines with multiple disks attached to each node. In earlier work, we proposed three strategies for processing range queries within the ADR framework. Our experimental results show that the relative performance of the strategies changes under varying application characteristics and machine configurations. In this work we investigate approaches to guide and automate the selection of the best strategy for a given application and machine configuration. We describe analytical models to predict the relative performance of the strategies where input data elements are uniformly distributed in the attribute space of the output dataset, restricting the output dataset to be a regular d-dimensional array.
Chialin Chang, Tahsin M. Kurç, Alan Sussman, Joel H. Saltz
IPDPS3
1999 A High-Performance Database System for Managing Large Multi-resolution Medical Images
Tahsin M. Kurç, Michael D. Beynon, Chialin Chang, Renato Ferreira 0001, Benjamin B. Bederson, Joel H. Saltz, Alan Sussman
AMIA7
1999 Performance impact of proxies in data intensive client-server applications
abstract
Large client-server data intensive applications can place high demands on system and network resources.This is especially true when the connection between the client and server spans a widearea intemet link.In this paper, we describe our experience changing the typical client-server architecture for a class of data intensive applications.We show that given sufficient common interest among multiple clients, our enhancements reduce the response time per-client, the response-time variation and the amount of data sent across the wide-area link.In addition, we also see a reduction in server utilization which helps to improve server scalability.
Michael D. Beynon, Alan Sussman, Joel H. Saltz
International Conference on Supercomputing2
1999 Querying Very Large Multi-dimensional Datasets in ADR
abstract
Applications that make use of very large scientific datasets have become an increasingly important subset of scientific applications.In these applications, datasets are often multi-dimensional, i.e., data items are associated with points in a multi-dimensional attribute space, and access to data items is described by range queries.The basic processing involves mapping input data items to output data items, and some form of aggregation of all the input data items that project to the each output data item.We have developed an infrastructure, called the Active Data Repository (ADR), that integrates storage, retrieval and processing of multi-dimensional datasets on distributed-memory parallel architectures with multiple disks attached to each node.In this paper we address efficient execution of range queries on distributed memory parallel machines within ADR framework.We present three potential strategies, and evaluate them under different application scenarios and machine configurations.We present experimental results on the scalability and performance of the strategies on a 128-node IBM SP.
Tahsin M. Kurç, Chialin Chang, Renato Ferreira 0001, Alan Sussman, Joel H. Saltz
SC4
1998 Digital dynamic telepathology-the Virtual Microscope
Asmara Afework, Michael D. Beynon, Fabián E. Bustamante, Soon Cho, Angelo Demarzo, Renato Ferreira 0001, Mark Silberman, Joel H. Saltz, Alan Sussman, Hubert Tsang
AMIA10
1998 The Design and Evaluation of a High-Performance Earth Science Database
Carter Shock, Chialin Chang, Bongki Moon, Anurag Acharya 0001, Larry Davis 0001, Joel H. Saltz, Alan Sussman
Parallel Comput.7
1997 The Virtual Microscope
Renato Ferreira 0001, Bongki Moon, Jim Humphries, Alan Sussman, Joel H. Saltz, Angelo Demarzo
AMIA4
1997 Titan: A High-Performance Remote Sensing Database
abstract
There are two major challenges for a high performance remote sensing database. First, it must provide low latency retrieval of very large volumes of spatio temporal data. This requires effective declustering and placement of a multidimensional dataset onto a large disk farm. Second, the order of magnitude reduction in data size due to post processing makes it imperative, from a performance perspective, that the post processing be done on the machine that holds the data. This requires careful coordination of computation and data retrieval. The paper describes the design, implementation and evaluation of Titan, a parallel shared nothing database designed for handling remote sensing data. The computational platform for Titan is a 16 processor IBM SP-2 with four fast disks attached to each processor. Titan is currently operational and contains about 24 GB of AVHRR data from the NOAA-7 satellite. The experimental results show that Titan provides good performance for global queries and interactive response times for local queries.
Chialin Chang, Bongki Moon, Anurag Acharya 0001, Carter Shock, Alan Sussman, Joel H. Saltz
ICDE5
1996 Runtime Coupling of Data-Parallel Programs
abstract
\f e conslcier the problem of efficiently couphrrg multiple cfataparakl programs at Irrntime.We propose an approach that establishes mappings between data structures m different flat a-parallel programs and implements a user-specified com wmencv model.Mappings are established at rnntlrne and can be added and deleted while the programs being coupled are in execution.Mappings, or the icleutity of the processors involved, do not ha~,r TObe kno~vn at compile-time or even link-time.Pro,gramh ran be m adc to interact tvith different granularities of interaction without requiring any re-cociing.A-priori knowledge of consistency requirements allows buffering of data as well as concurrent execution of the coupled applications.Efficient data movement is achii=w=d b~-p] e-computing an optimized schedule.~~e describe our ])lorotype mrpiernent,atiou and evaluate lts performance us-LMga wt of svnthetrc "benchmarks.\Ve examme the varlat]on of performance with varlatlon in t he conslstenc} reqrrme-[nent W+ demonstrate That the cost of' the tlexibiht}-pro-~,lded hl our coupling scheme is not prolubltl~,e whel L ronpared with a monolithic program that performs the sam~ computatlou.1
M. Ranganathan, Anurag Acharya 0001, Guy Edjlali, Alan Sussman, Joel H. Saltz
International Conference on Supercomputing4
1995 An Integrated Runtime and Compile-Time Approach for Parallelizing Structured and Block Structured Applications
abstract
In compiling applications for distributed memory machines, runtime analysis is required when data to be communicated cannot be determined at compile-time. One such class of applications requiring runtime analysis is block structured codes. These codes employ multiple structured meshes, which may be nested (for multigrid codes) and/or irregularly coupled (called multiblock or irregularly coupled regular mesh problems). In this paper, we present runtime and compile-time analysis for compiling such applications on distributed memory parallel machines in an efficient and machine-independent fashion. We have designed and implemented a runtime library which supports the runtime analysis required. The library is currently implemented on several different systems. We have also developed compiler analysis for determining data access patterns at compile time and inserting calls to the appropriate runtime routines. Our methods can be used by compilers for HPF-like parallel programming languages in compiling codes in which data distribution, loop bounds and/or strides are unknown at compile-time. To demonstrate the efficacy of our approach, we have implemented our compiler analysis in the Fortran 90D/HPF compiler developed at Syracuse University. We have experimented with a multi-bloc Navier-Stokes solver template and a multigrid code. Our experimental results show that our primitives have low runtime communication overheads and the compiler parallelized codes perform within 20% of the codes parallelized by manually inserting calls to the runtime library.>
Gagan Agrawal, Alan Sussman, Joel H. Saltz
IEEE Trans. Parallel Distributed Syst.2
1994 High performance computing for land cover dynamics
abstract
Presents the overall goals of the authors' research program on the application of high performance computing to remote sensing applications, specifically applications in land cover dynamics. This involves developing scalable and portable programs for a variety of image and map data processing applications, eventually integrated with new models for parallel I/O of large scale images and maps. After an overview of the multiblock PARTI run time support system, the authors explain extensions made to that system to support image processing applications, and then present an example involving multiresolution image processing. Results of running the parallel code on both a TMC CM5 and an Intel Paragon are discussed.
Rahul Parulekar, Larry Davis 0001, Rama Chellappa, Joel H. Saltz, Alan Sussman, John Townshend
ICPR (3)5
1993 Compiler and runtime support for structured and block structured applications
abstract
Scientific and engineering applications often involve structured meshes. These meshes may be nested (for multigrid or adaptive codes) and/or irregularly coupled(called Irregularly CoupledRegular Meshes). We have designed and implemented a runtime library for parallelizing this general class of applications on distributed memory parallel machines in an efficient and machine independent manner. In this paper we present how this runtime library can be integrated with compilers for High PerformanceFortran (HPF) style parallel programming languages. We discuss how we have integrated this runtime library with the Fortran 90D compiler being developed at Syracuse University and provide experimental data on a block structured NavierStokes solver template and a small multigrid example parallelized using this compiler and run on an Intel iPSC/860. We show that the compiler parallelizedcode performs within 20% of the code parallelized by inserting calls to the runtime library manually.
Gagan Agrawal, Alan Sussman, Joel H. Saltz
SC2
1993 Common runtime support for high-performance parallel languages
abstract
No abstract available.
Geoffrey C. Fox, Sanjay Ranka, Michael L. Scott, Allen D. Malony, James C. Browne, Marina C. Chen, Alok N. Choudhary, Thomas E. Cheatham, Janice E. Cuny, Rudolf Eigenmann, Amr F. Fahmy, Ian T. Foster, Dennis Gannon, Tomasz Haupt, Carl Kesselman, Charles Koelbel, Wei Li 0015, Monica S. Lam, Thomas J. LeBlanc, Jim Openshaw, David A. Padua, Constantine D. Polychronopoulos, Joel H. Saltz, Alan Sussman, Gil Weigand, Katherine A. Yelick
SC24
1992 Model-Driven Mapping onto Distributed Memory Parallel Computers
abstract
The author addresses the problem of exploiting the parallelism available in a program to efficiently employ the resources of the target machine in the context of building a mapping compiler for a distributed memory parallel machine. He demonstrates the effectiveness of using execution models to select the best mapping technique from among those available for a given program segment on a particular machine. Through analysis of the execution models for several mapping techniques for one class of programs on a linear processor array, it is shown that selecting the best technique for a particular program instance can make a significant difference in performance. On the other hand, the results of benchmarks from a mapping compiler for the Warp systolic array machine show that the execution models considered are accurate enough to select the best mapping technique for a given program.>
Alan Sussman
SC1