Calin Cascaval

dblp:c/CalinCascaval · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0002-2780-6763ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
12 papers
Distributed systems · 41% Parallel and multicore computing · 34% High-performance computing · 7%
Computer networks
1 paper
Software-defined and programmable networks · 87% Network management and operations · 13%
Software engineering, system software, and programming languages
4 papers
Program verification · 58% Operating systems · 29% Compilers and program optimization · 11%
Theoretical computer science
1 paper
Distributed computing theory · 100%

Topics — the 30 heaviest of 42, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
clock synchronization
0.812024
Logical Synchrony and the Bittide Mechanism · IEEE Trans. Parallel Distributed Syst. 2024
Distributed systems
distributed coordination
0.812024
Logical Synchrony and the Bittide Mechanism · IEEE Trans. Parallel Distributed Syst. 2024
Parallel and multicore computing
parallel programming models
0.352009
How much parallelism is there in irregular applications? · PPoPP 2009
Performance without pain = productivity: data layout and collective communication in UPC · PPoPP 2008
Implicit parallelism with ordered transactions · PPoPP 2007
Software-defined and programmable networks › programmable data plane
p4 program verification
0.312018
p4v: practical verification for programmable data planes · SIGCOMM 2018
Software-defined and programmable networks
programmable data plane
0.312018
p4v: practical verification for programmable data planes · SIGCOMM 2018
Energy-efficient computing
clock management
0.212024
Logical Synchrony and the Bittide Mechanism · IEEE Trans. Parallel Distributed Syst. 2024
Operating systems
mobile systems
0.212013
ZOOMM: a parallel web browser engine for multicore mobile devices · PPoPP 2013
Parallel and multicore computing
speculative parallelization
0.222008
Modeling optimistic concurrency using quantitative dependence analysis · PPoPP 2008
Implicit parallelism with ordered transactions · PPoPP 2007
Parallel and multicore computing › parallel programming models › distributed memory programming models
partitioned global address space
0.122008
Performance without pain = productivity: data layout and collective communication in UPC · PPoPP 2008
Shared memory programming for large scale machines · PLDI 2006
Parallel and multicore computing › parallel computing › parallel applications
irregular applications
0.112009
How much parallelism is there in irregular applications? · PPoPP 2009
Parallel and multicore computing › loop transformation
loop parallelization
0.112009
Parallelization spectroscopy: analysis of thread-level parallelism in hpc programs · PPoPP 2009
Storage systems
data layout
0.112008
Performance without pain = productivity: data layout and collective communication in UPC · PPoPP 2008
Distributed systems › concurrency control
optimistic concurrency control
0.112008
Modeling optimistic concurrency using quantitative dependence analysis · PPoPP 2008
Parallel and multicore computing › parallelization strategies
implicit parallelism
0.112007
Implicit parallelism with ordered transactions · PPoPP 2007
Compilers and program optimization
optimizing compiler
0.112006
Shared memory programming for large scale machines · PLDI 2006
Memory systems
cache coherence
0.112006
Bulk Disambiguation of Speculative Threads in Multiprocessors · ISCA 2006
Parallel and multicore computing › transactional memory
hardware transactional memory
0.112006
Bulk Disambiguation of Speculative Threads in Multiprocessors · ISCA 2006
Parallel and multicore computing › speculative parallelization
thread-level speculation
0.112006
Bulk Disambiguation of Speculative Threads in Multiprocessors · ISCA 2006
Parallel and multicore computing
transactional memory
0.112006
Bulk Disambiguation of Speculative Threads in Multiprocessors · ISCA 2006
High-performance computing › supercomputing
bluegene/l
0.012002
An overview of the BlueGene/L Supercomputer · SC 2002
Memory systems › cache
cache organization
0.012002
Evaluation of a Multithreaded Architecture for Cellular Computing · HPCA 2002
Emerging computing paradigms › unconventional computing
cellular computing
0.012002
Evaluation of a Multithreaded Architecture for Cellular Computing · HPCA 2002
Memory systems
memory hierarchy
0.012002
Evaluation of a Multithreaded Architecture for Cellular Computing · HPCA 2002
Processor architecture and microarchitecture
multithreading
0.012002
Evaluation of a Multithreaded Architecture for Cellular Computing · HPCA 2002
High-performance computing
supercomputing
0.012002
An overview of the BlueGene/L Supercomputer · SC 2002
Processor architecture and microarchitecture › multiprocessor architecture
synchronization hardware
0.012002
Evaluation of a Multithreaded Architecture for Cellular Computing · HPCA 2002
Integrated circuit design
system-on-chip
0.012002
An overview of the BlueGene/L Supercomputer · SC 2002
High-performance computing
collective communication
0.012008
Performance without pain = productivity: data layout and collective communication in UPC · PPoPP 2008
Memory systems › memory access patterns
irregular memory access
0.012008
Modeling optimistic concurrency using quantitative dependence analysis · PPoPP 2008
Parallel and multicore computing › parallel programming models
shared-memory parallelization
0.011999
MATmarks: A Shared Memory Environment for MATLAB Programming · HPDC 1999

Methods — techniques the papers use, named apart from their topics

logical latency measure · 1.5buffer overflow/underflow prevention · 1.5formal verification · 0.7domain-specific optimization · 0.7resource preloading · 0.3concurrency management · 0.3parallelization technique classification · 0.1optimistic parallel execution · 0.1quantitative modeling · 0.1dependence analysis · 0.1transactional memory · 0.1annotation · 0.1asynchronous message · 0.1affinity test elimination · 0.1shared memory commands · 0.0
YearPublicationVenuePosition
2025 Timetide: A Programming Model for Logically Synchronous Distributed Systems
abstract
Massive strides in deterministic models have been made using synchronous languages. They are mainly focused on centralised applications, as the traditional approach is to compile away the concurrency. Time triggered languages such as Giotto and Lingua Franca are suitable for distribution albeit that they rely on physical clock synchronisation, which is both expensive and may suffer from scalability. Hence, deterministic programming of distributed systems remains challenging. We address the challenges of deterministic distribution by developing a novel multiclock semantics of synchronous programs. The developed semantics is amenable to seamless distribution. Moreover, our programming model, Timetide, alleviates the need for physical clock synchronisation by building on the recently proposed logical synchrony model for distributed systems. We discuss the important aspects of distributing computation, such as network communication delays, and explore the formal verification of Timetide programs. To the best of our knowledge, Timetide is the first multiclock synchronous language that is both amenable to distribution and formal verification without the need for physical clock synchronisation or clock gating.
Logan Kenwright, Partha S. Roop, Nathan Allen, Calin Cascaval, Avinash Malik
ACM Trans. Embed. Comput. Syst.4
2024 Logical Synchrony and the Bittide Mechanism
abstract
We introduce logical synchrony, a framework that allows distributed computing to be coordinated as tightly as in synchronous systems without the distribution of a global clock or any reference to universal time. We develop a model of events called a logical synchrony network, in which nodes correspond to processors and every node has an associated local clock which generates the events. We construct a measure of logical latency and develop its properties. A further model, called a multiclock network, is then analyzed and shown to be a refinement of the logical synchrony network. We present the bittide mechanism as an instantiation of multiclock networks, and discuss the clock control mechanism that ensures that buffers do not overflow or underflow. Finally we give conditions under which a logical synchrony network has an equivalent synchronous realization.
Sanjay Lall, Calin Cascaval, Martin Izzard, Tammo Spalink
IEEE Trans. Parallel Distributed Syst.2
2023 On Buffer Centering for Bittide Synchronization
abstract
We discuss distributed reframing control of bittide systems. In a bittide system, multiple processors synchronize by monitoring communication over the network. Processors remain in logical synchrony by controlling the timing of frame transmissions. The protocol for doing this relies upon an underlying dynamic control system where each node makes only local observations and performs no direct coordination with other nodes. In this paper we develop a control algorithm based on the idea of buffer centering, which allows all nodes to maintain small buffer offsets while also requiring very little state information. We demonstrate that with buffer centering we can achieve separate control of frequency and phase, allowing frequencies to be syntonized and also buffers to maintain desired offsets rather than combining their control via a proportional-integral controller. The minimalism of this approach offers the potential to simplify both boot processes and failure handling.
Sanjay Lall, Calin Cascaval, Martin Izzard, Tammo Spalink
CoDIT2
2018 p4v: practical verification for programmable data planes
abstract
We present the design and implementation of p4v, a practical tool for verifying data planes described using the P4 programming language. The design of p4v is based on classic verification techniques but adds several key innovations including a novel mechanism for incorporating assumptions about the control plane and domain-specific optimizations which are needed to scale to large programs. We present case studies showing that p4v verifies important properties and finds bugs in real-world programs. We conduct experiments to quantify the scalability of p4v on a wide range of additional examples. We show that with just a few hundred lines of control-plane annotations, p4v is able to verify critical safety properties for switch.p4, a program that implements the functionality of on a modern data center switch, in under three minutes.
Jed Liu, William T. Hallahan, Cole Schlesinger, Milad Sharif, Jeongkeun Lee, Robert Soulé, Han Wang 0009, Calin Cascaval, Nick McKeown, Nate Foster
SIGCOMM8
2014 Deoptimization for dynamic language JITs on typed, stack-based virtual machines
abstract
We are interested in implementing dynamic language runtimes on top of language-level virtual machines. Type specialization is a critical optimization for dynamic language runtimes: generic code that handles any type of data is replaced with specialized code for particular types observed during execution. However, types can change, and the runtime must recover whenever unexpected types are encountered. The state-of-the-art recovery mechanism is called deoptimization. Deoptimization is a well-known technique for dynamic language runtimes implemented in low-level languages like C. However, no dynamic language runtime implemented on top of a virtual machine such as the Common Language Runtime (CLR) or the Java Virtual Machine (JVM) uses deoptimization, because the implementation thereof used in low-level languages is not possible.
Madhukar N. Kedlaya, Behnam Robatmili, Calin Cascaval, Ben Hardekopf
VEE3
2014 MuscalietJS: rethinking layered dynamic web runtimes
abstract
Layered JavaScript engines, in which the JavaScript runtime is built on top another managed runtime, provide better extensibility and portability compared to traditional monolithic engines. In this paper, we revisit the design of layered JavaScript engines and propose a layered architecture, called MuscalietJS2, that splits the responsibilities of a JavaScript engine between a high-level, JavaScript-specific component and a low-level, language-agnostic .NET VM. To make up for the performance loss due to layering, we propose a two pronged approach: high-level JavaScript optimizations and exploitation of low-level VM features that produce very efficient code for hot functions. We demonstrate the validity of the MuscalietJS design through a comprehensive evaluation using both the Sunspider benchmarks and a set of web workloads. We demonstrate that our approach outperforms other layered engines such as IronJS and Rhino engines while providing extensibility, adaptability and portability.
Behnam Robatmili, Calin Cascaval, Mehrdad Reshadi, Madhukar N. Kedlaya, Seth Fowler, Vrajesh Bhavsar, Michael Weber 0002, Ben Hardekopf
VEE2
2013 Keynote talk: Parallel programming for mobile computing
abstract
Summary form only given. Personal computing is going mobile and applications are changing to adapt to take advantage of new opportunities offered by permanent availability and connectivity. Mobile devices are a significant departure from traditional computing. On one hand, they are very personal, always on, always connected. They promise to fulfill the promise of being the hub for our digital lives. On the other hand, they are much more constrained in terms of resources than desktops. Even though progress in their computing capabilities has been staggering, they continue to rely on battery power and are packaged in appealing packages that are a nightmare for thermal dissipation. In this talk I will present the challenges facing programmers for mobile devices driven by architectural and packaging constraints, as well as the changes in applications domains. I will give examples on how we used concurrency to improve performance and power efficiency, in a number of projects at Qualcomm Research, including the Zoomm parallel browser.
Calin Cascaval
PACT1
2013 ZOOMM: a parallel web browser engine for multicore mobile devices
abstract
We explore the challenges in expressing and managing concurrency in browsers on mobile devices. Browsers are complex applications that implement multiple standards, need to support legacy behavior, and are highly dynamic and interactive. We present ZOOMM, a highly concurrent web browser engine prototype and show how concurrency is effectively exploited at different levels: speed up computation performance, preload network resources, and preprocess resources outside the critical path of page loading. On a dual-core Android mobile device we demonstrate that ZOOMM is two times faster than the native WebKit based browser when loading the set of pages defined in the Vellamo benchmark.
Calin Cascaval, Seth Fowler, Pablo Montesinos, Wayne Piekarski, Mehrdad Reshadi, Behnam Robatmili, Michael Weber 0002, Vrajesh Bhavsar
PPoPP1
2009 Analytical Modeling of Pipeline Parallelism
abstract
Parallel programming is a requirement in the multi-core era. One of the most promising techniques to make parallel programming available for the general users is the use of parallel programming patterns. Functional pipeline parallelism is a pattern that is well suited for many emerging applications, such as streaming and "recognition, mining and synthesis" (RMS) workloads. In this paper we develop an analytical model for pipeline parallelism based on queueing theory. The model is useful to both characterize the performance and efficiency of existing implementations and to guide the design of new pipeline algorithms. We demonstrate the usefulness of the model by characterizing and optimizing two of the PARSEC benchmarks, ferret and dedup. We identified two issues with these codes: load imbalance and I/O bottlenecks. We addressed load imbalance using two techniques: i) parallel pipeline stage collapsing; and ii) dynamic scheduling. We implemented these optimizations using pthreads and the threading building blocks (TBB) libraries. We compare the performance of different alternatives and we note that the TBB implementation based on work stealing outperforms all other variants.
Angeles G. Navarro, Rafael Asenjo, Siham Tabik, Calin Cascaval
PACT4
2009 Load balancing using work-stealing for pipeline parallelism in emerging applications
abstract
Parallel programming is a requirement in the multi-core era. One of the most promising techniques to make parallel programming available for general users is the use of parallel programming patterns. Functional pipeline parallelism is a well suited pattern for many emerging applications, such as streaming and "Recognition, Mining and Synthesis" (RMS) workloads. In this paper we develop an analytical model for pipeline parallelism and use it to characterize and optimize two of the PARSEC benchmarks which use the parallel pipeline pattern, ferret and dedup. We identify two scalability limitations: load imbalance and I/O bottlenecks. We address load imbalance using two techniques: parallel pipeline stage collapsing and dynamic scheduling. We implemented these optimizations using Pthreads and the Threading Building Blocks (TBB) libraries. We compare predicted and measured performance of all these implementations on a large scale SMP machine and we note that the work-stealing TBB implementation outperforms all other variants.
Angeles G. Navarro, Rafael Asenjo, Siham Tabik, Calin Cascaval
ICS4
2009 Scalable RDMA performance in PGAS languages
abstract
Partitioned global address space (PGAS) languages provide a unique programming model that can span shared-memory multiprocessor (SMP) architectures, distributed memory machines, or cluster ofSMPs. Users can program large scale machines with easy-to-use, shared memory paradigms. In order to exploit large scale machines efficiently, PGAS language implementations and their runtime system must be designed for scalability and performance. The IBM XLUPC compiler and runtime system provide a scalable design through the use of the shared variable directory (SVD). The SVD stores meta-information needed to access shared data. It is dereferenced, in the worst case, for every shared memory access, thus exposing a potential performance problem. In this paper we present a cache of remote addresses as an optimization that will reduce the SVD access overhead and allow the exploitation of native (remote) direct memory accesses. It results in a significant performance improvement while maintaining the run-time portability and scalability.
Montse Farreras, Gheorghe Almási 0001, Calin Cascaval, Toni Cortes
IPDPS3
2009 Lonestar: A suite of parallel irregular programs
abstract
Until recently, parallel programming has largely focused on the exploitation of data-parallelism in dense matrix programs. However, many important application domains, including meshing, clustering, simulation, and machine learning, have very different algorithmic foundations: they require building, computing with, and modifying large sparse graphs. In the parallel programming literature, these types of applications are usually classified as irregular applications, and relatively little attention has been paid to them. To study and understand the patterns of parallelism and locality in sparse graph computations better, we are in the process of building the Lonestar benchmark suite. In this paper, we characterize the first five programs from this suite, which target domains like data mining, survey propagation, and design automation. We show that even such irregular applications often expose large amounts of parallelism in the form of amorphous data-parallelism. Our speedup numbers demonstrate that this new type of parallelism can successfully be exploited on modern multi-core machines.
Milind Kulkarni 0001, Martin Burtscher, Calin Cascaval, Keshav Pingali
ISPASS3
2009 Parallelization spectroscopy: analysis of thread-level parallelism in hpc programs
abstract
In this paper, we present a method - parallelization spectroscopy - for analyzing the thread-level parallelism available in production High Performance Computing (HPC) codes.We survey a number of techniques that are commonly used for parallelization and classify all the loops in the case study presented using a sensitivity metric: how likely is a particular technique is successful in parallelizing the loop.
Arun Kejariwal, Calin Cascaval
PPoPP2
2009 How much parallelism is there in irregular applications?
abstract
Irregular programs are programs organized around pointer-based data structures such as trees and graphs. Recent investigations by the Galois project have shown that many irregular programs have a generalized form of data-parallelism called amorphous data-parallelism. However, in many programs, amorphous data-parallelism cannot be uncovered using static techniques, and its exploitation requires runtime strategies such as optimistic parallel execution. This raises a natural question: how much amorphous data-parallelism actually exists in irregular programs?
Milind Kulkarni 0001, Martin Burtscher, R. Inkulu, Keshav Pingali, Calin Cascaval
PPoPP5
2009 Compiler and runtime techniques for software transactional memory optimization
abstract
Abstract Software transactional memory (STM) systems are an attractive environment to evaluate optimistic concurrency. We describe our experience of supporting and optimizing an STM system at both the managed runtime and compiler levels. We describe the design policies of our STM system and the statistics collected by the runtime to identify performance bottlenecks and guide tuning decisions. We present an initial work on supporting automatic instrumentation of the STM primitives for C/C++ and Java programs in the IBM XL compiler and J9 Java virtual machine. We evaluate and discuss the performance of several transactional programs running on our system. Copyright © 2008 John Wiley & Sons, Ltd.
Peng Wu 0001, Maged M. Michael, Christoph von Praun, Takuya Nakaike, Rajesh Bordawekar, Harold W. Cain, Calin Cascaval, Siddhartha Chatterjee, Stefanie Chiras, Mark F. Mergen, Michael F. Spear, Huayong Wang
Concurr. Comput. Pract. Exp.7
2008 Performance without pain = productivity: data layout and collective communication in UPC
abstract
The next generations of supercomputers are projected to have hundreds of thousands of processors. However, as the numbers of processors grow, the scalability of applications will be the dominant challenge. This forces us to reexamine some of our fundamental ways that we approach the design and use of parallel languages and runtime systems.
Rajesh Nishtala, Gheorghe Almási 0001, Calin Cascaval
PPoPP3
2008 Modeling optimistic concurrency using quantitative dependence analysis
abstract
This work presents a quantitative approach to analyze parallelization opportunities in programs with irregular memory access where potential data dependencies mask available parallelism. The model captures data and causal dependencies among critical sections as algorithmic properties and quantifies them as a density computed over the number of executed instructions. The model abstracts from runtime aspects such as scheduling, the number of threads, and concurrency control used in a particular parallelization.
Christoph von Praun, Rajesh Bordawekar, Calin Cascaval
PPoPP3
2007 Implicit parallelism with ordered transactions
abstract
Implicit Parallelism with Ordered Transactions (IPOT) is an extension of sequential or explicitly parallel programming models to support speculative parallelization. The key idea is to specify opportunities for parallelization in a sequential program using annotations similar to transactions. Unlike explicit parallelism, IPOT annotations do not require the absence of data dependence, since the parallelization relies on runtime support for speculative execution. IPOT as a parallel programming model is determinate, i.e., program semantics are independent of the thread scheduling. For optimization, non-determinism can be introduced selectively.
Christoph von Praun, Luis Ceze, Calin Cascaval
PPoPP3
2006 Bulk Disambiguation of Speculative Threads in Multiprocessors
abstract
Transactional Memory (TM), Thread-Level Speculation (TLS), and Checkpointed multiprocessors are three popular architectural techniques based on the execution of multiple, cooperating speculative threads. In these environments, correctly maintaining data dependences across threads requires mechanisms for disambiguating addresses across threads, invalidating stale cache state, and making committed state visible. These mechanisms are both conceptually involved and hard to implement. In this paper, we present Bulk, a novel approach to simplify these mechanisms. The idea is to hash-encode a thread’s access information in a concise signature, and then support in hardware signature operations that efficiently process sets of addresses. Such operations implement the mechanisms described. Bulk operations are inexact but correct, and provide substantial conceptual and implementation simplicity. We evaluate Bulk in the context of TLS using SPECint2000 codes and TM using multithreaded Java workloads. Despite its simplicity, Bulk has competitive performance with more complex schemes. We also find that signature configuration is a key design parameter.
Luis Ceze, James Tuck 0001, Josep Torrellas, Calin Cascaval
ISCA4
2006 Shared memory programming for large scale machines
abstract
This paper describes the design and implementation of a scalable run-time system and an optimizing compiler for Unified Parallel C (UPC). An experimental evaluation on BlueGene/L®, a distributed-memory machine, demonstrates that the combination of the compiler with the runtime system produces programs with performance comparable to that of efficient MPI programs and good performance scalability up to hundreds of thousands of processors.Our runtime system design solves the problem of maintaining shared object consistency efficiently in a distributed memory machine. Our compiler infrastructure simplifies the code generated for parallel loops in UPC through the elimination of affinity tests, eliminates several levels of indirection for accesses to segments of shared arrays that the compiler can prove to be local, and implements remote update operations through a lower-cost asynchronous message. The performance evaluation uses three well-known benchmarks --- HPC RandomAccess, HPC STREAM and NAS CG --- to obtain scaling and absolute performance numbers for these benchmarks on up to 131072 processors, the full BlueGene/L machine. These results were used to win the HPC Challenge Competition at SC05 in Seattle WA, demonstrating that PGAS languages support both productivity and performance.
Christopher Barton, Calin Cascaval, Gheorghe Almási 0001, Yili Zheng, Montse Farreras, Siddhartha Chatterjee, José Nelson Amaral
PLDI2
2003 An Overview of the Blue Gene/L System Software Organization
Gheorghe Almási 0001, Ralph Bellofatto, José R. Brunheroto, Calin Cascaval, José G. Castaños, Luis Ceze, Paul Crumley, C. Christopher Erway, Joseph Gagliano, Derek Lieber, Xavier Martorell, José E. Moreira, Alda Sanomiya, Karin Strauss
Euro-Par4
2003 Estimating cache misses and locality using stack distances
abstract
Cache behavior modeling is an important part of modern optimizing compilers. In this paper we present a method to estimate the number of cache misses, at compile time, using a machine independent model based on stack algorithms. Our algorithm computes the stack histograms symbolically, using data dependence distance vectors and is totally accurate when dependence distances are uniformly generated. The stack histogram models accurately fully associative caches with LRU replacement policy, and provides a very good approximation for set-associative caches and programs with non-constant dependence distances.The stack histogram is an accurate, machine-independent metric of locality. Compilers using this metric can evaluate optimizations with respect to memory behavior. We illustrate this use of the stack histogram by comparing three locality enhancing transformations: tiling, data shackling and the product-space transformation. Additionally, the stack histogram model can be used to compute optimal parameters for data locality transformations, such as the tile size for loop tiling.
Calin Cascaval, David A. Padua
ICS1
2002 Blue Gene/L, a System-On-A-Chip
abstract
Summary form only given. Large powerful networks coupled to state-of-the-art processors have traditionally dominated supercomputing. As technology advances, this approach is likely to be challenged by a more cost-effective System-On-A-Chip approach, with higher levels of system integration. The scalability of applications to architectures with tens to hundreds of thousands of processors is critical to the success of this approach. Significant progress has been made in mapping numerous compute-intensive applications, many of them grand challenges, to parallel architectures. Applications hoping to efficiently execute on future supercomputers of any architecture must be coded in a manner consistent with an enormous degree of parallelism. The BG/L program is developing a peak nominal 180 TFLOPS (360 TFLOPS for some applications) supercomputer to serve a broad range of science applications. BG/L generalizes QCDOC, the first System-On-A-Chip supercomputer that is expected in 2003. BG/L consists of 65,536 nodes, and contains five integrated networks: a 3D torus, a combining tree, a Gb Ethernet network, barrier/global interrupt network and JTAG.
George S. Almási, Daniel K. Beece, Ralph Bellofatto, Gyan Bhanot, Randy Bickford, Matthias A. Blumrich, Arthur A. Bright, José R. Brunheroto, Calin Cascaval, José G. Castaños, Luis Ceze, Paul Coteus, Siddhartha Chatterjee, Dong Chen 0005, George L.-T. Chiu, Thomas M. Cipolla, Paul Crumley, Alina Deutsch, Marc Boris Dombrowa, Wilm E. Donath, Maria Eleftheriou, Blake G. Fitch, Joseph Gagliano, Alan Gara, Robert S. Germain, Mark Giampapa, Manish Gupta 0002, Fred G. Gustavson, Shawn Hall, Ruud A. Haring, David F. Heidel, Philip Heidelberger, Lorraine M. Herger, Dirk Hoenicke, T. Jamal-Eddine, Gerard V. Kopcsay, Alphonso P. Lanzetta, Derek Lieber, M. Lu, Mark P. Mendell, Lawrence S. Mok, José E. Moreira, Ben J. Nathanson, Matthew Newton, Martin Ohmacht, Rick A. Rand, Richard D. Regan, Ramendra K. Sahoo, Alda Sanomiya, Eugen Schenfeld, Sarabjeet Singh, Peilin Song, Burkhard D. Steinmacher-Burow, Karin Strauss, Richard A. Swetz, Todd Takken, R. Brett Tremaine, Mickey Tsao, Pavlos Vranas, T. J. Christopher Ward, Michael E. Wazlowski, J. Brown, Thomas A. Liebsch, A. Schram, G. Ulsh
CLUSTER9
2002 Evaluation of a Multithreaded Architecture for Cellular Computing
abstract
Cyclops is a new architecture for high-performance parallel computers that is being developed at the IBM T. J. Watson Research Center. The basic cell of this architecture is a single-chip SMP (symmetric multiprocessor) system with multiple threads of execution, embedded memory and integrated communications hardware. Massive intra-chip parallelism is used to tolerate memory and functional unit latencies. Large systems with thousands of chips can be built by replicating this basic cell in a regular pattern. In this paper, we describe the Cyclops architecture and evaluate two of its new hardware features: a memory hierarchy with a flexible cache organization and fast barrier hardware. Our experiments with the STREAM benchmark show that a particular design can achieve a sustainable memory bandwidth of 40 GB/s, equal to the peak hardware bandwidth and similar to the performance of a 128-processor SGI Origin 3800. For small vectors, we have observed in-cache bandwidth above 80 GB/s. We also show that the fast barrier hardware can improve the performance of the Splash-2 FFT kernel by up to 10%. Our results demonstrate that the Cyclops approach of integrating a large number of simple processing elements and multiple memory banks in the same chip is an effective alternative for designing high-performance systems.
Calin Cascaval, José G. Castaños, Luis Ceze, Monty Denneau, Manish Gupta 0002, Derek Lieber, José E. Moreira, Karin Strauss, Henry S. Warren Jr.
HPCA1
2002 An overview of the BlueGene/L Supercomputer
abstract
This paper gives an overview of the BlueGene/L Supercomputer. This is a jointly funded research partnership between IBM and the Lawrence Livermore National Laboratory as part of the United States Department of Energy ASCI Advanced Architecture Research Program. Application performance and scaling studies have recently been initiated with partners at a number of academic and government institutions,including the San Diego Supercomputer Center and the California Institute of Technology. This massively parallel system of 65,536 nodes is based on a new architecture that exploits system-on-a-chip technology to deliver target peak processing power of 360 teraFLOPS (trillion floating-point operations per second). The machine is scheduled to be operational in the 2004-2005 time frame, at price/performance and power consumption/performance targets unobtainable with conventional architectures.
Narasimha R. Adiga, Gheorghe Almási 0001, George S. Almási, Yariv Aridor, Rajkishore Barik, Daniel K. Beece, Ralph Bellofatto, Gyan Bhanot, Randy Bickford, Matthias A. Blumrich, Arthur A. Bright, José R. Brunheroto, Calin Cascaval, José G. Castaños, Waiman Chan, Luis Ceze, Paul Coteus, Siddhartha Chatterjee, Dong Chen 0005, George L.-T. Chiu, Thomas M. Cipolla, Paul Crumley, K. M. Desai, Alina Deutsch, Tamar Domany, Marc Boris Dombrowa, Wilm E. Donath, Maria Eleftheriou, C. Christopher Erway, J. Esch, Blake G. Fitch, Joseph Gagliano, Alan Gara, Rahul Garg 0001, Robert S. Germain, Mark Giampapa, Balaji Gopalsamy, John A. Gunnels, Manish Gupta 0002, Fred G. Gustavson, Shawn Hall, Ruud A. Haring, David F. Heidel, Philip Heidelberger, Lorraine M. Herger, Dirk Hoenicke, R. D. Jackson, T. Jamal-Eddine, Gerard V. Kopcsay, Elie Krevat, Manish P. Kurhekar, Alphonso P. Lanzetta, Derek Lieber, L. K. Liu, M. Lu, Mark P. Mendell, A. Misra, Yosef Moatti, Lawrence S. Mok, José E. Moreira, Ben J. Nathanson, Matthew Newton, Martin Ohmacht, Adam J. Oliner, Vinayaka Pandit, R. B. Pudota, Rick A. Rand, Richard D. Regan, Bradley Rubin, Albert E. Ruehli, Silvius Vasile Rus, Ramendra K. Sahoo, Alda Sanomiya, Eugen Schenfeld, M. Sharma, Edi Shmueli, Sarabjeet Singh, Peilin Song, Vijay Srinivasan, Burkhard D. Steinmacher-Burow, Karin Strauss, Christopher W. Surovic, Richard A. Swetz, Todd Takken, R. Brett Tremaine, Mickey Tsao, Arun R. Umamaheshwaran, P. Verma, Pavlos Vranas, T. J. Christopher Ward, Michael E. Wazlowski, W. Barrett, C. Engel, B. Drehmel, B. Hilgart, D. Hill, F. Kasemkhani, David J. Krolak, Chun-Tao Li 0001, Thomas A. Liebsch, James A. Marcella, A. Muff, A. Okomo, M. Rouse, A. Schram, M. Tubbs, G. Ulsh, Charles D. Wait, J. Wittrup, Myung Bae, Kenneth A. Dockser, Lynn Kissel, Mark K. Seager, Jeffrey S. Vetter, K. Yates
SC13
2001 Demonstrating the scalability of a molecular dynamics application on a Petaflop computer
abstract
The IBM Blue Gene project has endeavored into the development of a cellular architecture computer with millions of concurrent threads of execution. One of the major challenges of this project is demonstrating that applications can successfully exploit this massive amount of parallelism. Starting from the sequential version of a well known molecular dynamics code, we developed a new application that exploits the multiple levels of parallelism in the Blue Gene cellular architecture. We perform both analytical and simulation studies of the behavior of this application when executed on a very large number of threads. As a result, we demonstrate that this class of applications can execute efficiently on a large cellular machine.
George S. Almási, Calin Cascaval, José G. Castaños, Monty Denneau, Wilm E. Donath, Maria Eleftheriou, Mark Giampapa, C. T. Howard Ho, Derek Lieber, José E. Moreira, Dennis M. Newns, Marc Snir, Henry S. Warren Jr.
ICS2
1999 MATmarks: A Shared Memory Environment for MATLAB Programming
abstract
MATmarks is an extension of the MATLAB tool that enables shared memory programming on a network of workstations by adding a small set of commands. The authors present a high level overview of the MATmarks system, the commands we added to MATLAB, and the performance gains we achieved as a result.
Gheorghe Almási 0001, Calin Cascaval, David A. Padua
HPDC2