EDBT 2026 Demo / reviewers in the wild / expert
Karin Strauss
dblp:60/4729
· DBLP profile ↗
44ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0002-8327-5477ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 2 first-authorSoftware engineering, systems software and programming languages · 15 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
27 papers |
Storage systems · 37% Memory systems · 28% Emerging computing paradigms · 7% | |
| Software engineering, system software, and programming languages
7 papers |
Concurrent programming · 80% Runtime systems and virtual machines · 13% Compilers and program optimization · 5% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 60% Computational science and engineering · 40% | |
| Computer networks
2 papers |
Content delivery and video streaming · 53% Wireless networking · 27% Edge and fog computing · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 50% Web and social media mining · 50% |
Topics — the 30 heaviest of 81, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
non-volatile memory |
1.1 | 7 | 2017 | Atomic In-place Updates for Non-volatile Main Memories with Kamino-Tx · EuroSys 2017 Approximate Storage in Solid-State Memories · ACM Trans. Comput. Syst. 2014 Using managed runtime systems to tolerate holes in wearable memories · PLDI 2013 |
Storage systems › energy-efficient storage
approximate storage |
0.9 | 4 | 2017 | Approximate Storage of Compressed and Encrypted Videos · ASPLOS 2017 High-Density Image Storage Using Approximate Memory Cells · ASPLOS 2016 Approximate Storage in Solid-State Memories · ACM Trans. Comput. Syst. 2014 |
Memory systems › non-volatile memory
phase change memory |
0.6 | 4 | 2014 | Approximate Storage in Solid-State Memories · ACM Trans. Comput. Syst. 2014 Using managed runtime systems to tolerate holes in wearable memories · PLDI 2013 Zombie memory: extending memory lifetime by reviving dead blocks · ISCA 2013 |
Storage systems › storage devices › molecular data storage
DNA storage |
0.6 | 2 | 2019 | DNA Data Storage and Hybrid Molecular-Electronic Computing · Proc. IEEE 2019 A DNA-Based Archival Storage System · ASPLOS 2016 |
Storage systems
storage reliability |
0.6 | 3 | 2017 | Approximate Storage of Compressed and Encrypted Videos · ASPLOS 2017 Approximate Storage in Solid-State Memories · ACM Trans. Comput. Syst. 2014 High-Density Image Storage Using Approximate Memory Cells · ASPLOS 2016 |
Emerging computing paradigms
molecular computing |
0.4 | 2 | 2019 | DNA Data Storage and Hybrid Molecular-Electronic Computing · Proc. IEEE 2019 DNA-based molecular architecture with spatially localized components · ISCA 2013 |
Computational science and engineering › fluid dynamics
microfluidics |
0.4 | 1 | 2019 | Puddle: A Dynamic, Error-Correcting, Full-Stack Microfluidics Platform · ASPLOS 2019 |
Embedded and real-time systems
cyber-physical system platforms |
0.4 | 1 | 2019 | Puddle: A Dynamic, Error-Correcting, Full-Stack Microfluidics Platform · ASPLOS 2019 |
Storage systems › storage devices
molecular data storage |
0.4 | 1 | 2019 | DNA Data Storage and Hybrid Molecular-Electronic Computing · Proc. IEEE 2019 |
Bioinformatics and computational biology › sequence analysis › sequence clustering
DNA sequence clustering |
0.3 | 1 | 2017 | Clustering Billions of Reads for DNA Data Storage · NIPS 2017 |
Bioinformatics and computational biology
DNA storage |
0.3 | 1 | 2017 | Clustering Billions of Reads for DNA Data Storage · NIPS 2017 |
Storage systems
atomic writes |
0.3 | 1 | 2017 | Atomic In-place Updates for Non-volatile Main Memories with Kamino-Tx · EuroSys 2017 |
Memory systems › non-volatile memory › persistent memory
byte-addressable persistent memory |
0.3 | 1 | 2017 | Atomic In-place Updates for Non-volatile Main Memories with Kamino-Tx · EuroSys 2017 |
Algorithms and data structures › sequence algorithms › string algorithms
edit distance |
0.3 | 1 | 2017 | Clustering Billions of Reads for DNA Data Storage · NIPS 2017 |
Image and video coding
image compression |
0.2 | 1 | 2016 | High-Density Image Storage Using Approximate Memory Cells · ASPLOS 2016 |
Storage systems
archival storage |
0.2 | 1 | 2016 | A DNA-Based Archival Storage System · ASPLOS 2016 |
Storage systems
key-value storage |
0.2 | 1 | 2016 | A DNA-Based Archival Storage System · ASPLOS 2016 |
Hardware reliability and fault tolerance
error correction |
0.2 | 2 | 2016 | Zombie memory: extending memory lifetime by reviving dead blocks · ISCA 2013 High-Density Image Storage Using Approximate Memory Cells · ASPLOS 2016 |
Information retrieval
search engines |
0.2 | 1 | 2015 | PocketTrend: Timely Identification and Delivery of Trending Search Content to Mobile Users · WWW 2015 |
Web and social media mining
trend detection |
0.2 | 1 | 2015 | PocketTrend: Timely Identification and Delivery of Trending Search Content to Mobile Users · WWW 2015 |
Concurrent programming › concurrency bugs
atomicity violation |
0.2 | 2 | 2010 | ColorSafe: architectural support for debugging and dynamically avoiding multi-variable atomicity violations · ISCA 2010 Atom-Aid: Detecting and Surviving Atomicity Violations · ISCA 2008 |
Concurrent programming › concurrency bug detection
data race detection |
0.2 | 2 | 2012 | RADISH: Always-on sound and complete race detection in software and hardware · ISCA 2012 ColorSafe: architectural support for debugging and dynamically avoiding multi-variable atomicity violations · ISCA 2010 |
Runtime systems and virtual machines
managed runtime |
0.2 | 1 | 2013 | Using managed runtime systems to tolerate holes in wearable memories · PLDI 2013 |
Emerging computing paradigms › molecular computing
DNA computing |
0.2 | 1 | 2013 | DNA-based molecular architecture with spatially localized components · ISCA 2013 |
Storage systems › flash and SSD
endurance management |
0.2 | 1 | 2013 | Zombie memory: extending memory lifetime by reviving dead blocks · ISCA 2013 |
Storage systems › flash and SSD
solid-state drive |
0.2 | 1 | 2013 | Approximate storage in solid-state memories · MICRO 2013 |
Wireless networking › mobile computing
mobile web browsing |
0.1 | 1 | 2012 | PocketWeb: instant web browsing for mobile devices · ASPLOS 2012 |
Content delivery and video streaming
prefetching |
0.1 | 1 | 2012 | PocketWeb: instant web browsing for mobile devices · ASPLOS 2012 |
Content delivery and video streaming › prefetching
web prefetching |
0.1 | 1 | 2012 | PocketWeb: instant web browsing for mobile devices · ASPLOS 2012 |
Concurrent programming › concurrency bug detection › data race detection
dynamic race detection |
0.1 | 1 | 2012 | RADISH: Always-on sound and complete race detection in software and hardware · ISCA 2012 |
Methods — techniques the papers use, named apart from their topics
edit distance approximation · 0.9distributed clustering · 0.9runtime resource management · 0.8computer vision-based error correction · 0.8simulation · 0.7error-tolerant storage · 0.6compression-aware storage · 0.6query log analysis · 0.4digital microfluidics · 0.4DNA synthesis · 0.4DNA sequencing · 0.4selective error correction · 0.2biasing · 0.2user study · 0.2page retirement · 0.2line-level failure handling · 0.2prefetching · 0.1energy-aware scheduling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Robust Digital Molecular Design of Binarized Neural NetworksabstractMolecular programming - a paradigm wherein molecules are engineered to perform computation - shows great potential for applications in nanotechnology, disease diagnostics and smart therapeutics. A key challenge is to identify systematic approaches for compiling abstract models of computation to molecules. Due to their wide applicability, one of the most useful abstractions to realize is neural networks. In prior work, real-valued weights were achieved by individually controlling the concentrations of the corresponding "weight" molecules. However, large-scale preparation of reactants with precise concentrations quickly becomes intractable. Here, we propose to bypass this fundamental problem using Binarized Neural Networks (BNNs), a model that is highly scalable in a molecular setting due to the small number of distinct weight values. We devise a noise-tolerant digital molecular circuit that compactly implements a majority voting operation on binary-valued inputs to compute the neuron output. The network is also rate-independent, meaning the speed at which individual reactions occur does not affect the computation, further increasing robustness to noise. We first demonstrate our design on the MNIST classification task by simulating the system as idealized chemical reactions. Next, we map the reactions to DNA strand displacement cascades, providing simulation results that demonstrate the practical feasibility of our approach. We perform extensive noise tolerance simulations, showing that digital molecular neurons are notably more robust to noise in the concentrations of chemical reactants compared to their analog counterparts. Finally, we provide initial experimental results of a single binarized neuron. Our work suggests a solid framework for building even more complex neural network computation. Johannes Linder, Yuan-Jyue Chen, Georg Seelig, Luis Ceze, Karin Strauss |
DNA | 6 |
| 2019 | Puddle: A Dynamic, Error-Correcting, Full-Stack Microfluidics PlatformabstractMicrofluidic devices promise to automate wetlab procedures by manipulating small chemical or biological samples. This technology comes in many varieties, all of which aim to save time, labor, and supplies by performing lab protocol steps typically done by a technician. However, existing microfluidic platforms remain some combination of inflexible, error-prone, prohibitively expensive, and difficult to program. We address these concerns with a full-stack digital microfluidic automation platform. Our main contribution is a runtime system that provides a high-level API for microfluidic manipulations. It manages fluidic resources dynamically, allowing programmers to freely mix regular computation with microfluidics, which results in more expressive programs than previous work. It also provides real-time error correction through a computer vision system, allowing robust execution on cheaper microfluidic hardware. We implement our stack on top of a low-cost droplet microfluidic device that we have developed. We evaluate our system with the fully-automated execution of polymerase chain reaction (PCR) and a DNA sequencing preparation protocol. These protocols demonstrate high-level programs that combine computational and fluidic operations such as input/output of reagents, heating of samples, and data analysis. We also evaluate the impact of automatic error correction on our system's reliability. Max Willsey, Ashley P. Stephenson, Chris Takahashi, Pranav Vaid, Bichlien Nguyen, Michal Piszczek, Christine Betts, Sharon Newman, Karin Strauss, Luis Ceze |
ASPLOS | 10 |
| 2019 | DNA Data Storage and Near-Molecule Processing for the Yottabyte Era
Karin Strauss, Luis Ceze |
CIDR | 1 |
| 2019 | Scaling Microfluidics to Complex, Dynamic Protocols: Invited PaperabstractMicrofluidic devices promise to automate wetlab procedures by manipulating small chemical or biological samples. We are developing a full-stack microfluidic automation platform that allows and allows users to scale up the complexity of microfluidic programming, encouraging them to mix fluidic manipulations with traditional programming. Puddle is a runtime system that provides a high-level API for microfluidic manipulations. It manages fluidic resources dynamically, allowing programmers to freely mix regular computation with microfluidics, resulting in more expressive programs. It also provides real-time error correction through a computer vision system, allowing robust execution on cheaper digital microfluidic hardware. We have been running Puddle on PurpleDrop, a new digital microfluidic device that is affordable and has novel features such as fully automated input/output of fluids. With this combination, we have demonstrated PCR with automated replenishment, a DNA sequencing preparation protocol, and the complete retrieval of digital data stored in dehydrated spots of DNA on the device's surface. Going forward, we see Puddle and PurpleDrop as part of a platform for further research. PurpleDrop is affordable and extensible, which makes a compelling case for adding new periferials or even scaling out by connecting multiple devices. And Puddle provides a flexible and abstract programming model that could enable microfluidic programs to run on different hardware targets (DMF or liquid handling robots), or even a combination thereof. Max Willsey, Ashley P. Stephenson, Chris Takahashi, Bichlien Nguyen, Karin Strauss, Luis Ceze |
ICCAD | 5 |
| 2019 | DNA Data Storage and Hybrid Molecular-Electronic ComputingabstractMoore's law may be slowing, but our ability to manipulate molecules is improving faster than ever. DNA could provide alternative substrates for computing and storage as existing ones approach physical limits. In this paper, we explore the implications of this trend in computer architecture. We present a computer systems perspective on molecular processing and storage, positing a hybrid molecular-electronic architecture that plays to the strengths of both domains. We cover the design and implementation of all stages of the pipeline: encoding, DNA synthesis, system integration with digital microfluidics, DNA sequencing (including emerging technologies such as nanopores), and decoding. We first draw on our experience designing a DNA-based archival storage system, which includes the largest demonstration to date of DNA digital data storage of over three billion nucleotides encoding over 400 MB of data. We then propose a more ambitious hybrid-electronic design that uses a molecular form of near-data processing for massive parallelism. We present a model that demonstrates the feasibility of these systems in the near future. We think the time is ripe to consider molecular storage seriously and explore system designs and architectural implications. Douglas M. Carmean, Luis Ceze, Georg Seelig, Kendall Stewart, Karin Strauss, Max Willsey |
Proc. IEEE | 5 |
| 2018 | A Content-Addressable DNA Database with Learned Sequence Encodings
Kendall Stewart, Yuan-Jyue Chen, David Ward, Georg Seelig, Karin Strauss, Luis Ceze |
DNA | 6 |
| 2017 | Approximate Storage of Compressed and Encrypted VideosabstractThe popularization of video capture devices has created strong storage demand for encoded videos. Approximate storage can ease this demand by enabling denser storage at the expense of occasional errors. Unfortunately, even minor storage errors, such as bit flips, can result in major visual damage in encoded videos. Similarly, video encryption, widely employed for privacy and digital rights management, may create long dependencies between bits that show little or no tolerance to storage errors. Djordje Jevdjic, Karin Strauss, Luis Ceze, Henrique S. Malvar |
ASPLOS | 2 |
| 2017 | Atomic In-place Updates for Non-volatile Main Memories with Kamino-TxabstractData structures for non-volatile memories have to be designed such that they can be atomically modified using transactions. Existing atomicity methods require data to be copied in the critical path which significantly increases the latency of transactions. These overheads are further amplified for transactions on byte-addressable persistent memories where often the byte ranges modified for data structure updates are significantly smaller compared to the granularity at which data can be efficiently copied and logged. We propose Kamino-Tx that provides a new way to perform transactional updates on non-volatile byte-addressable memories (NVM) without requiring any copying of data in the critical path. Kamino-Tx maintains an additional copy of data off the critical path to achieve atomicity. But in doing so Kamino-Tx has to overcome two important challenges of safety and minimizing NVM storage overhead. We propose a more dynamic approach to maintaining the additional copy of data to reduce storage overheads. To further mitigate the storage overhead of using Kamino-Tx in a replicated setting, we develop Kamino-Tx-Chain, a variant of Chain Replication where replicas perform in-place updates and do not maintain data copies locally; replicas in Kamino-Tx-Chain leverage other replicas as copies to roll back or forward for atomicity. Our results show that using Kamino-Tx increases throughput by up to 9.5x for unreplicated systems and up to 2.2x for replicated settings. Amir Saman Memaripour, Anirudh Badam, Amar Phanishayee, Yanqi Zhou, Ramnatthan Alagappan, Karin Strauss, Steven Swanson |
EuroSys | 6 |
| 2017 | Customizing Progressive JPEG for Efficient Image Storage
Eddie Q. Yan, Kaiyuan Zhang 0001, Xi Wang 0005, Karin Strauss, Luis Ceze |
HotStorage | 4 |
| 2017 | Clustering Billions of Reads for DNA Data StorageabstractStoring data in synthetic DNA offers the possibility of improving information density and durability by several orders of magnitude compared to current storage technologies. However, DNA data storage requires a computationally intensive process to retrieve the data. In particular, a crucial step in the data retrieval pipeline involves clustering billions of strings with respect to edit distance. Datasets in this domain have many notable properties, such as containing a very large number of small clusters that are well-separated in the edit distance metric space. In this regime, existing algorithms are unsuitable because of either their long running time or low accuracy. To address this issue, we present a novel distributed algorithm for approximately computing the underlying clusters. Our algorithm converges efficiently on any dataset that satisfies certain separability properties, such as those coming from DNA data storage systems. We also prove that, under these assumptions, our algorithm is robust to outliers and high levels of noise. We provide empirical justification of the accuracy, scalability, and convergence of our algorithm on real and synthetic data. Compared to the state-of-the-art algorithm for clustering DNA sequences, our algorithm simultaneously achieves higher accuracy and a 1000x speedup on three real datasets. Cyrus Rashtchian, Konstantin Makarychev, Miklós Z. Rácz, Siena Ang, Djordje Jevdjic, Sergey Yekhanin, Luis Ceze, Karin Strauss |
NIPS | 8 |
| 2016 | A DNA-Based Archival Storage SystemabstractDemand for data storage is growing exponentially, but the capacity of existing storage media is not keeping up. Using DNA to archive data is an attractive possibility because it is extremely dense, with a raw limit of 1 exabyte/mm3 (109 GB/mm3), and long-lasting, with observed half-life of over 500 years. This paper presents an architecture for a DNA-based archival storage system. It is structured as a key-value store, and leverages common biochemical techniques to provide random access. We also propose a new encoding scheme that offers controllable redundancy, trading off reliability for density. We demonstrate feasibility, random access, and robustness of the proposed encoding with wet lab experiments involving 151 kB of synthesized DNA and a 42 kB random-access subset, and simulation experiments of larger sets calibrated to the wet lab experiments. Finally, we highlight trends in biotechnology that indicate the impending practicality of DNA storage for much larger datasets. James Bornholt, Randolph Lopez, Douglas M. Carmean, Luis Ceze, Georg Seelig, Karin Strauss |
ASPLOS | 6 |
| 2016 | High-Density Image Storage Using Approximate Memory CellsabstractThis paper proposes tailoring image encoding for an approximate storage substrate. We demonstrate that indiscriminately storing encoded images in approximate memory generates unacceptable and uncontrollable quality degradation. The key finding is that errors in the encoded bit streams have non-uniform impact on the decoded image quality. We develop a methodology to determine the relative importance of encoded bits and store them in an approximate storage substrate. The storage cells are optimized to reduce error rate via biasing and are tuned to meet the desired reliability requirement via selective error correction. In a case study with the progressive transform codec (PTC), a precursor to JPEG XR, the proposed approximate image storage system exhibits a 2.7x increase in density of pixels per silicon volume under bounded error rates, and this achievement is additive to the storage savings of PTC compression. Karin Strauss, Luis Ceze, Henrique S. Malvar |
ASPLOS | 2 |
| 2015 | Toward accelerating deep learning at scale using specialized hardware in the datacenter
Kalin Ovtcharov, Olatunji Ruwase, Joo-Young Kim 0001, Jeremy Fowers, Karin Strauss, Eric S. Chung |
Hot Chips Symposium | 5 |
| 2015 | PocketTrend: Timely Identification and Delivery of Trending Search Content to Mobile UsersabstractTrending search topics cause unpredictable query load spikes that hurt the end-user search experience, particularly the mobile one, by introducing longer delays. To understand how trending search topics are formed and evolve over time, we analyze 21 million queries submitted during periods where popular events caused search query volume spikes. Based on our findings, we design and evaluate PocketTrend, a system that automatically detects trending topics in real time, identifies the search content associated to the topics, and then intelligently pushes this content to users in a timely manner. In that way, PocketTrend enables a client-side search engine that can instantly answer user queries related to trending events, while at the same time reducing the impact of these trends on the datacenter workload. Our results, using real mobile search logs, show that in the presence of a trending event, up to 13-17% of the overall search traffic can be eliminated from the datacenter, with as many as 19% of all users benefiting from PocketTrend. Gennady Pekhimenko, Dimitrios Lymberopoulos, Oriana Riva, Karin Strauss, Doug Burger |
WWW | 4 |
| 2014 | A High Memory Bandwidth FPGA Accelerator for Sparse Matrix-Vector MultiplicationabstractSparse matrix-vector multiplication (SMVM) is a crucial primitive used in a variety of scientific and commercial applications. Despite having significant parallelism, SMVM is a challenging kernel to optimize due to its irregular memory access characteristics. Numerous studies have proposed the use of FPGAs to accelerate SMVM implementations. However, most prior approaches focus on parallelizing multiply-accumulate operations within a single row of the matrix (which limits parallelism if rows are small) and/or make inefficient uses of the memory system when fetching matrix and vector elements. In this paper, we introduce an FPGA-optimized SMVM architecture and a novel sparse matrix encoding that explicitly exposes parallelism across rows, while keeping the hardware complexity and on-chip memory usage low. This system compares favorably with prior FPGA SMVM implementations. For the over 700 University of Florida sparse matrices we evaluated, it also performs within about two thirds of CPU SMVM performance on average, even though it has 2.4x lower DRAM memory bandwidth, and within almost one third of GPU SVMV performance on average, even at 9x lower memory bandwidth. Additionally, it consumes only 25W, for power efficiencies 2.6x and 2.3x higher than CPU and GPU, respectively, based on maximum device power. Jeremy Fowers, Kalin Ovtcharov, Karin Strauss, Eric S. Chung, Greg Stitt |
FCCM | 3 |
| 2014 | Data Race Detection with Minimal Hardware SupportabstractThis article presents AccessedBefore (AccB), an algorithm and its associated minimal hardware support to detect data races, and compares it with two widely known and used commercial tools: Helgrind, the data race detection tool included in the general purpose memory checking suite Valgrind, and Intel Thread Checker, now shipped as part of Intel Thread Inspector. It provides a performance overhead evaluation using current workloads, along with an analysis of AccB's scalability with the number of threads and workload input set size. It demonstrates that AccB is in the range of 2× to 11× faster than these two tools. Finally, it shows the complete proof that AccB is complete in that, for every static data race present in a program, there exists an instruction interleaving that would expose this data race such that AccB can detect it. Rekai González-Alberquilla, Fernando Emmanuel Frati, Luis Piñuel, Karin Strauss, Luis Ceze |
Comput. J. | 4 |
| 2014 | Approximate Storage in Solid-State MemoriesabstractMemories today expose an all-or-nothing correctness model that incurs significant costs in performance, energy, area, and design complexity. But not all applications need high-precision storage for all of their data structures all of the time. This article proposes mechanisms that enable applications to store data approximately and shows that doing so can improve the performance, lifetime, or density of solid-state memories. We propose two mechanisms. The first allows errors in multilevel cells by reducing the number of programming pulses used to write them. The second mechanism mitigates wear-out failures and extends memory endurance by mapping approximate data onto blocks that have exhausted their hardware error correction resources. Simulations show that reduced-precision writes in multilevel phase-change memory cells can be 1.7 × faster on average and using failed blocks can improve array lifetime by 23% on average with quality loss under 10%. Adrian Sampson, Jacob Nelson 0001, Karin Strauss, Luis Ceze |
ACM Trans. Comput. Syst. | 3 |
| 2013 | Reducing disruption from subtle information delivery during a conversation: mode and bandwidth investigationabstractWith proliferation of mobile devices that provide ubiquitous access to information, the question arises of how distracting processing information in social settings can be, especially during face-to-face conversations. However, relevant information presented at opportune moments may help enhance conversation quality. In this paper, we study how much information users can consume during a conversation and what information delivery mode, via audio or visual aids, helps them effectively conceal the fact that they are receiving information. We observe that users can internalize more information while still disguising this fact the best when information is delivered visually in batches (multiple pieces of information at a time) and perform better on both dimensions if information is delivered while they are not speaking. Interestingly, participants qualitatively did not prefer this mode as being the easiest to use, preferring modes that displayed one piece of information at a time. Eyal Ofek, Shamsi T. Iqbal, Karin Strauss |
CHI | 3 |
| 2013 | Zombie memory: extending memory lifetime by reviving dead blocksabstractZombie is an endurance management framework that enables a variety of error correction mechanisms to extend the lifetimes of memories that suffer from bit failures caused by wearout, such as phase-change memory (PCM). Zombie supports both single-level cell (SLC) and multi-level cell (MLC) variants. It extends the lifetime of blocks in working memory pages (primary blocks) by pairing them with spare blocks, i.e., working blocks in pages that have been disabled due to exhaustion of a single block's error correction resources, which would be 'dead' otherwise. Spare blocks adaptively provide error correction resources to primary blocks as failures accumulate over time. This reduces the waste caused by early block failures, making working blocks in discarded pages a useful resource. Even though we use PCM as the target technology, Zombie applies to any memory technology that suffers stuck-at cell failures. Rodolfo Azevedo, John D. Davis, Karin Strauss, Parikshit Gopalan, Mark S. Manasse, Sergey Yekhanin |
ISCA | 3 |
| 2013 | DNA-based molecular architecture with spatially localized componentsabstractPerforming computation inside living cells offers life-changing applications, from improved medical diagnostics to better cancer therapy to intelligent drugs. Due to its bio-compatibility and ease of engineering, one promising approach for performing in-vivo computation is DNA strand displacement. This paper introduces computer architects to DNA strand displacement "circuits", discusses associated architectural challenges, and proposes a new organization that provides practical composability. In particular, prior approaches rely mostly on stochastic interaction of freely diffusing components. This paper proposes practical spatial isolation of components, leading to more easily designed DNA-based circuits. DNA nanotechnology is currently at a turning point, with many proposed applications being realized [20, 9]. We believe that it is time for the computer architecture community to take notice and contribute. Richard A. Muscat, Karin Strauss, Luis Ceze, Georg Seelig |
ISCA | 2 |
| 2013 | Approximate storage in solid-state memoriesabstractMemories today expose an all-or-nothing correctness model that incurs significant costs in performance, energy, area, and design complexity. But not all applications need high-precision storage for all of their data structures all of the time. This paper proposes mechanisms that enable applications to store data approximately and shows that doing so can improve the performance, lifetime, or density of solid-state memories. We propose two mechanisms. The first allows errors in multi-level cells by reducing the number of programming pulses used to write them. The second mechanism mitigates wear-out failures and extends memory endurance by mapping approximate data onto blocks that have exhausted their hardware error correction resources. Simulations show that reduced-precision writes in multi-level phase-change memory cells can be 1.7x faster on average and using failed blocks can improve array lifetime by 23% on average with quality loss under 10%. Adrian Sampson, Jacob Nelson 0001, Karin Strauss, Luis Ceze |
MICRO | 3 |
| 2013 | Using managed runtime systems to tolerate holes in wearable memoriesabstractNew memory technologies, such as phase-change memory (PCM), promise denser and cheaper main memory, and are expected to displace DRAM. However, many of them experience permanent failures far more quickly than DRAM. DRAM mechanisms that handle permanent failures rely on very low failure rates and, if directly applied to PCM, are extremely inefficient: Discarding a page when the first line fails wastes 98% of the memory. Tiejun Gao, Karin Strauss, Steve Blackburn, Kathryn S. McKinley, Doug Burger, James R. Larus |
PLDI | 2 |
| 2012 | PocketWeb: instant web browsing for mobile devicesabstractThe high network latencies and limited battery life of mobile phones can make mobile web browsing a frustrating experience. In prior work, we proposed trading memory capacity for lower web access latency and a more convenient data transfer schedule from an energy perspective by prefetching slowly-changing data (search queries and results) nightly, when the phone is charging. However, most web content is intrinsically much more dynamic and may be updated multiple times a day, thus eliminating the effectiveness of periodic updates. Dimitrios Lymberopoulos, Oriana Riva, Karin Strauss, Akshay Mittal, Alexandros Ntoulas |
ASPLOS | 3 |
| 2012 | RADISH: Always-on sound and complete race detection in software and hardwareabstractData-race freedom is a valuable safety property for multithreaded programs that helps with catching bugs, simplifying memory consistency model semantics, and verifying and enforcing both atomicity and determinism. Unfortunately, existing software-only dynamic race detectors are precise but slow; proposals with hardware support offer higher performance but are imprecise. Both precision and performance are necessary to achieve the many advantages always-on dynamic race detection could provide. Joseph Devietti, Benjamin P. Wood, Karin Strauss, Luis Ceze, Dan Grossman, Shaz Qadeer |
ISCA | 3 |
| 2012 | Goldilocks and the two mobile devices: going beyond all-or-nothing access to a device's applicationsabstractMost mobile phones and tablets support only two access control device states: locked and unlocked. We investigated how well all or-nothing device access control meets the need of users by interviewing 20 participants who had both a smartphone and tablet. We find all-or-nothing device access control to be a remarkably poor fit with users' preferences. On both phones and tablets, participants wanted roughly half their applications to be available even when their device was locked and half protected by authentication. We also solicited participants' interest in new access control mechanisms designed specifically to facilitate device sharing. Fourteen participants out of 20 preferred these controls to existing security locks alone. Finally, we gauged participants' interest in using face and voice biometrics to authenticate to their mobile phone and tablets; participants were surprisingly receptive to biometrics, given that they were also aware of security and reliability limitations. Eiji Hayashi, Oriana Riva, Karin Strauss, A. J. Bernheim Brush, Stuart E. Schechter |
SOUPS | 3 |
| 2012 | Progressive Authentication: Deciding When to Authenticate on Mobile Phones
Oriana Riva, Karin Strauss, Dimitrios Lymberopoulos |
USENIX Security Symposium | 3 |
| 2011 | Pocket cloudletsabstractCloud services accessed through mobile devices suffer from high network access latencies and are constrained by energy budgets dictated by the devices' batteries. Radio and battery technologies will improve over time, but are still expected to be the bottlenecks in future systems. Non-volatile memories (NVM), however, may continue experiencing significant and steady improvements in density for at least ten more years. In this paper, we propose to leverage the abundance in memory capacity of mobile devices to mitigate latency and energy issues when accessing cloud services. Emmanouil Koukoumidis, Dimitrios Lymberopoulos, Karin Strauss, Jie Liu 0001, Doug Burger |
ASPLOS | 3 |
| 2011 | Accelerating Data Race Detection with Minimal Hardware Support
Rodrígo González-Alberquilla, Karin Strauss, Luis Ceze, Luis Piñuel |
Euro-Par (1) | 2 |
| 2011 | Preventing PCM banks from seizing too much powerabstractWidespread adoption of Phase Change Memory (PCM) requires solutions to several problems recently addressed in the literature, including limited endurance, increased write latencies, and system-level changes required to exploit non-volatility. One important difference between PCM and DRAM that has received less attention is the increased need for write power management. Writing to a PCM cell requires high current density over hundreds of nanoseconds, and hard limits on the number of simultaneous writes must be enforced to ensure correct operation, limiting write throughput and therefore overall performance. Because several wear reduction schemes only write those bits that need to be written, the amount of power required to write a cache line back to memory under such a system is now variable, which creates opportunity to reduce write power. This paper proposes policies that monitor the bits that have actually been changed over time, as opposed to simply those lines that are dirty. These polices can more effectively allocate power across the system to improve write concurrency. This method for allocating power across the memory subsystem is built on the idea of "power tokens," a transferable, but time-specific, allocation of power. The results show that with a storage overhead of 4.3% in the last-level cache, a power-aware memory system can improve the performance of multiprogrammed workloads by up to 84%. Andrew W. Hay, Karin Strauss, Timothy Sherwood, Gabriel H. Loh, Doug Burger |
MICRO | 2 |
| 2011 | The impact of memory models on software reliability in multiprocessorsabstractThe memory consistency model is a fundamental system property characterizing a multiprocessor. The relative merits of strict versus relaxed memory models have been widely debated in terms of their impact on performance, hardware complexity and programmability. This paper adds a new dimension to this discussion: the impact of memory models on software reliability. By allowing some instructions to reorder, weak memory models may expand the window between critical memory operations. This can increase the chance of an undesirable thread-interleaving, thus allowing an otherwise-unlikely concurrency bug to manifest. To explore this phenomenon, we define and study a probabilistic model of shared-memory parallel programs that takes into account such reordering. We use this model to formally derive bounds on the vulnerability to concurrency bugs of different memory models. Our results show that for 2 concurrent threads, weaker memory models do indeed have a higher likelihood of allowing bugs. On the other hand, we show that as the number of parallel, buggy threads increases, the gap between the different memory models becomes proportionally insignificant, and thus the importance of using a strict memory model diminishes. Alexander Jaffe, Thomas Moscibroda, Laura Effinger-Dean, Luis Ceze, Karin Strauss |
PODC | 5 |
| 2010 | ColorSafe: architectural support for debugging and dynamically avoiding multi-variable atomicity violationsabstractIn this paper, we propose ColorSafe, an architecture that detects and dynamically avoids single- and multi-variable atomicity violation bugs. The key idea is to group related data into colors and then monitor access interleavings in the space. This enables detection of atomicity violations involving any data of the same color. We leverage support for meta-data to maintain color information, and signatures to efficiently keep recent color access histories. ColorSafe dynamically avoids atomicity violations by inserting ephemeral transactions that prevent erroneous interleavings. ColorSafe has two modes of operation: (1)debugging mode makes detection more precise, producing fewer false positives and collecting more information; and, (2)deployment mode provides robust, efficient dynamic bug avoidance with less precise detection. This makes ColorSafe useful throughout the lifetime of programs, not just during development. Our results show that, in deployment mode, ColorSafe is able to successfully avoid the majority of multi-variable atomicity violations in bug kernels, as well as in large applications (Apache and MySQL). In debugging mode, ColorSafe detects bugs with few false positives. Brandon Lucia, Luis Ceze, Karin Strauss |
ISCA | 3 |
| 2010 | Conflict exceptions: simplifying concurrent language semantics with precise hardware exceptions for data-racesabstractWe argue in this paper that concurrency errors should be treated as exceptions, i.e., have fail-stop behavior and precise semantics. We propose an exception model based on conflict of synchronization free regions, which precisely detects a broad class of data-races. We show that our exceptions provide enough guarantees to simplify high-level programming language semantics and debugging, but are significantly cheaper to enforce than traditional data-race detection. To make the performance cost of enforcement negligible, we propose architecture support for accurately detecting and precisely delivering these exceptions. We evaluate the suitability of our model as well as the behavior of our architectural mechanisms using the PARSEC benchmark suite and commercial applications. Our results show that the exception model largely reflects how programmers are already writing code and that the main memory, traffic and performance overheads of the enforcement mechanisms we propose are very low. Brandon Lucia, Luis Ceze, Karin Strauss, Shaz Qadeer, Hans-Juergen Boehm |
ISCA | 3 |
| 2010 | Use ECP, not ECC, for hard failures in resistive memoriesabstractAs leakage and other charge storage limitations begin to impair the scalability of DRAM, non-volatile resistive memories are being developed as a potential replacement. Unfortunately, current error correction techniques are poorly suited to this emerging class of memory technologies. Unlike DRAM, PCM and other resistive memories have wear lifetimes, measured in writes, that are sufficiently short to make cell failures common during a system's lifetime. However, resistive memories are much less susceptible to transient faults than DRAM. The Hamming-based ECC codes used in DRAM are designed to handle transient faults with no effective lifetime limits, but ECC codes applied to resistive memories would wear out faster than the cells they are designed to repair. This paper evaluates Error-Correcting Pointers (ECP), a new approach to error correction optimized for memories in which errors are the result of permanent cell failures that occur, and are immediately detectable, at write time. ECP corrects errors by permanently encoding the locations of failed cells into a table and assigning cells to replace them. ECP provides longer lifetimes than previously proposed solutions with equivalent overhead. What's more, as the level of variance in cell lifetimes increases -- a likely consequence of further scalaing -- ECP's margin of improvement over existing schemes increases. Stuart E. Schechter, Gabriel H. Loh, Karin Strauss, Doug Burger |
ISCA | 3 |
| 2008 | Atom-Aid: Detecting and Surviving Atomicity ViolationsabstractWriting shared-memory parallel programs is error-prone. Among the concurrency errors that programmers often face are atomicity violations, which are especially challenging. They happen when programmers make incorrect assumptions about atomicity and fail to enclose memory accesses that should occur atomically inside the same critical section. If these accesses happen to be interleaved with conflicting accesses from different threads, the program might behave incorrectly. Recent architectural proposals arbitrarily group consecutive dynamic memory operations into atomic blocks to enforce memory ordering at a coarse grain. This provides what we call implicit atomicity, as the atomic blocks are not derived from explicit program annotations. In this paper, we make the fundamental observation that implicit atomicity probabilistically hides atomicity violations by reducing the number of interleaving opportunities between memory operations. We then propose Atom-Aid, which creates implicit atomic blocks intelligently instead of arbitrarily, dramatically reducing the probability that atomicity violations will manifest themselves. Atom-Aid is also able to report where atomicity violations might exist in the code, providing resilience and debuggability. We evaluate Atom-Aid using buggy code from applications including Apache, MySQL, and XMMS, showing that Atom-Aid virtually eliminates the manifestation of atomicity violations. Brandon Lucia, Joseph Devietti, Karin Strauss, Luis Ceze |
ISCA | 3 |
| 2007 | Uncorq: Unconstrained Snoop Request Delivery in Embedded-Ring MultiprocessorsabstractSnoopy cache coherence can be implemented in any physical network topology by embedding a logical unidirectional ring in the network. Control messages are forwarded using the ring, while other messages can use any path. While the resulting coherence protocols are inexpensive to implement, they enable many ways of overlapping multiple transactions that access the same line-making it hard to reason about correctness. Moreover, snoop requests are required to traverse the ring, therefore lengthening coherence transaction latencies. In this paper, we address these problems and make two main contributions. First, we introduce theorderinginvariant, which ensures the correct serialization of colliding transactions in embedded-ring protocols. Second, based on this invariant, we remove the requirement that snoop requests traverse the ring. Instead, they are delivered using any network path, as long as snoop responses - which are typically off the critical path - use the logical ring. This approach substantially reduces coherence transaction latency. We call the resulting protocolUncorq. We show that, on a 64-node chip multiprocessor (CMP), Uncorq improves the performance, on average, by 23% for SPLASH-2 applications and by 10% for commercial applications. With an additional simple prefetching optimization, the performance improvement is, on average, 26% for SPLASH-2 applications and 18% for commercial applications. Karin Strauss, Josep Torrellas |
MICRO | 1 |
| 2006 | Flexible Snooping: Adaptive Forwarding and Filtering of Snoops in Embedded-Ring MultiprocessorsabstractA simple and low-cost approach to supporting snoopy cache coherence is to logically embed a unidirectional ring in the network of a multiprocessor, and use it to transfer snoop messages. Other messages can use any link in the network. While this scheme works for any network topology, a naive implementation may result in long response times or in many snoop messages and snoop operations. To address this problem, this paper proposes flexible snooping algorithms, a family of adaptive forwarding and filtering snooping algorithms. In these algorithms, a node receiving a snoop request may either forward it to another node and then perform the snoop, or snoop and then forward it, or simply forward it without snooping. The resulting design space offers trade-offs in number of snoop operations and messages, response time, and energy consumption. Our analysis using SPLASH-2, SPECjbb, and SPECweb workloads finds several snooping algorithms that are more cost-effective than current ones. Specifically, our choice for a high-performance snooping algorithm is faster than the currently fastest algorithm while consuming 9-17% less energy; our choice for an energy-efficient algorithm is only 3-6% slower than the previous one while consuming 36-42% less energy Karin Strauss, Josep Torrellas |
ISCA | 1 |
| 2006 | POSH: a TLS compiler that exploits program structureabstractAs multi-core architectures with Thread-Level Speculation (TLS) are becoming better understood, it is important to focus on TLS compilation. TLS compilers are interesting in that, while they do not need to fully prove the independence of concurrent tasks, they make choices of where and when to generate speculative tasks that are crucial to overall TLS performance.This paper presents POSH, a new, fully automated TLS compiler built on top of gcc. POSH is based on two design decisions. First, to partition the code into tasks, it leverages the code structures created by the programmer, namely subroutines and loops. Second, it uses a simple profiling pass to discard ineffective tasks. With the code generated by POSH, a simulated TLS chip multiprocessor with 4 superscalar cores delivers an average speedup of 1.30 for the SPECint 2000 applications. Moreover, an estimated 26% of this speedup is a result of the implicit data prefetching provided by squashed tasks. Wei Liu 0014, James Tuck 0001, Luis Ceze, Wonsun Ahn, Karin Strauss, Jose Renau, Josep Torrellas |
PPoPP | 5 |
| 2006 | CAVA: Using checkpoint-assisted value prediction to hide L2 missesabstractModern superscalar processors often suffer long stalls because of load misses in on-chip L2 caches. To address this problem, we propose hiding L2 misses with Checkpoint-Assisted VAlue prediction (CAVA). On an L2 cache miss, a predicted value is returned to the processor. When the missing load finally reaches the head of the ROB, the processor checkpoints its state, retires the load, and speculatively uses the predicted value and continues execution. When the value in memory arrives at the L2 cache, it is compared to the predicted value. If the prediction was correct, speculation has succeeded and execution continues; otherwise, execution is rolled back and restarted from the checkpoint. CAVA uses fast checkpointing, speculative buffering, and a modest-sized value prediction structure that has about 50% accuracy. Compared to an aggressive superscalar processor, CAVA speeds up execution by up to 1.45 for SPECint applications and 1.58 for SPECfp applications, with a geometric mean of 1.14 for SPECint and 1.34 for SPECfp applications. We also evaluate an implementation of Runahead execution---a previously proposed scheme that does not perform value prediction and discards all work done between checkpoint and data reception from memory. Runahead execution speeds up execution by a geometric mean of 1.07 for SPECint and 1.18 for SPECfp applications, compared to the same baseline. Luis Ceze, Karin Strauss, James Tuck 0001, Josep Torrellas, Jose Renau |
ACM Trans. Archit. Code Optim. | 2 |
| 2005 | Thread-Level Speculation on a CMP can be energy efficientabstractChip Multiprocessors (CMP) with Thread-Level Speculation (TLS) have become the subject of intense research. However, TLS is suspected of being too energy inefficient to compete against conventional processors. In this paper, we refute this claim. To do so, we first identify the main sources of dynamic energy consumption in TLS. Then, we present simple energy-saving optimizations that cut the energy cost of TLS by over 60% on average with minimal performance impact. The resulting TLS CMP, populated with four 3-issue cores, speeds-up full SPECint 2000 codes by 1.27 on average, while keeping the fraction of the chip's energy consumption due to TLS to only 20%. Compared to a 6-issue superscalar at the same frequency, the TLS CMP is on average faster, while consuming only 85% of its total on-chip power. Jose Renau, Karin Strauss, Luis Ceze, Wei Liu 0014, Smruti R. Sarangi, James Tuck 0001, Josep Torrellas |
ICS | 2 |
| 2005 | Tasking with out-of-order spawn in TLS chip multiprocessors: microarchitecture and compilationabstractChip Multiprocessors (CMPs) are flexible, high-frequency platforms on which to support Thread-Level Speculation (TLS). However, for TLS to deliver on its promise, CMPs must exploit multiple sources of speculative task-level parallelism, including any nesting levels of both subroutines and loop iterations. Unfortunately, these environments are hard to support in decentralized CMP hardware: since tasks are spawned out-of-order and unpredictably, maintaining key TLS basics such as task ordering and efficient resource allocation is challenging.While the concept of out-of-order spawning is not new, this paper is the first to propose a set of microarchitectural mechanisms that, altogether, fundamentally enable fast TLS with out-of-order spawn in a CMP. Moreover, we develop a fully-automated TLS compiler for aggressive out-of-order spawn. With our mechanisms, a TLS CMP with four 4-issue cores achieves an average speedup of 1.30 for full SPECint 2000 applications; the corresponding speedup for in-order only spawn is 1.04. Overall, our mechanisms unlock the potential of TLS for the toughest applications. Jose Renau, James Tuck 0001, Wei Liu 0014, Luis Ceze, Karin Strauss, Josep Torrellas |
ICS | 5 |
| 2003 | An Overview of the Blue Gene/L System Software Organization
Gheorghe Almási 0001, Ralph Bellofatto, José R. Brunheroto, Calin Cascaval, José G. Castaños, Luis Ceze, Paul Crumley, C. Christopher Erway, Joseph Gagliano, Derek Lieber, Xavier Martorell, José E. Moreira, Alda Sanomiya, Karin Strauss |
Euro-Par | 14 |
| 2002 | Blue Gene/L, a System-On-A-ChipabstractSummary form only given. Large powerful networks coupled to state-of-the-art processors have traditionally dominated supercomputing. As technology advances, this approach is likely to be challenged by a more cost-effective System-On-A-Chip approach, with higher levels of system integration. The scalability of applications to architectures with tens to hundreds of thousands of processors is critical to the success of this approach. Significant progress has been made in mapping numerous compute-intensive applications, many of them grand challenges, to parallel architectures. Applications hoping to efficiently execute on future supercomputers of any architecture must be coded in a manner consistent with an enormous degree of parallelism. The BG/L program is developing a peak nominal 180 TFLOPS (360 TFLOPS for some applications) supercomputer to serve a broad range of science applications. BG/L generalizes QCDOC, the first System-On-A-Chip supercomputer that is expected in 2003. BG/L consists of 65,536 nodes, and contains five integrated networks: a 3D torus, a combining tree, a Gb Ethernet network, barrier/global interrupt network and JTAG. George S. Almási, Daniel K. Beece, Ralph Bellofatto, Gyan Bhanot, Randy Bickford, Matthias A. Blumrich, Arthur A. Bright, José R. Brunheroto, Calin Cascaval, José G. Castaños, Luis Ceze, Paul Coteus, Siddhartha Chatterjee, Dong Chen 0005, George L.-T. Chiu, Thomas M. Cipolla, Paul Crumley, Alina Deutsch, Marc Boris Dombrowa, Wilm E. Donath, Maria Eleftheriou, Blake G. Fitch, Joseph Gagliano, Alan Gara, Robert S. Germain, Mark Giampapa, Manish Gupta 0002, Fred G. Gustavson, Shawn Hall, Ruud A. Haring, David F. Heidel, Philip Heidelberger, Lorraine M. Herger, Dirk Hoenicke, T. Jamal-Eddine, Gerard V. Kopcsay, Alphonso P. Lanzetta, Derek Lieber, M. Lu, Mark P. Mendell, Lawrence S. Mok, José E. Moreira, Ben J. Nathanson, Matthew Newton, Martin Ohmacht, Rick A. Rand, Richard D. Regan, Ramendra K. Sahoo, Alda Sanomiya, Eugen Schenfeld, Sarabjeet Singh, Peilin Song, Burkhard D. Steinmacher-Burow, Karin Strauss, Richard A. Swetz, Todd Takken, R. Brett Tremaine, Mickey Tsao, Pavlos Vranas, T. J. Christopher Ward, Michael E. Wazlowski, J. Brown, Thomas A. Liebsch, A. Schram, G. Ulsh |
CLUSTER | 54 |
| 2002 | Evaluation of a Multithreaded Architecture for Cellular ComputingabstractCyclops is a new architecture for high-performance parallel computers that is being developed at the IBM T. J. Watson Research Center. The basic cell of this architecture is a single-chip SMP (symmetric multiprocessor) system with multiple threads of execution, embedded memory and integrated communications hardware. Massive intra-chip parallelism is used to tolerate memory and functional unit latencies. Large systems with thousands of chips can be built by replicating this basic cell in a regular pattern. In this paper, we describe the Cyclops architecture and evaluate two of its new hardware features: a memory hierarchy with a flexible cache organization and fast barrier hardware. Our experiments with the STREAM benchmark show that a particular design can achieve a sustainable memory bandwidth of 40 GB/s, equal to the peak hardware bandwidth and similar to the performance of a 128-processor SGI Origin 3800. For small vectors, we have observed in-cache bandwidth above 80 GB/s. We also show that the fast barrier hardware can improve the performance of the Splash-2 FFT kernel by up to 10%. Our results demonstrate that the Cyclops approach of integrating a large number of simple processing elements and multiple memory banks in the same chip is an effective alternative for designing high-performance systems. Calin Cascaval, José G. Castaños, Luis Ceze, Monty Denneau, Manish Gupta 0002, Derek Lieber, José E. Moreira, Karin Strauss, Henry S. Warren Jr. |
HPCA | 8 |
| 2002 | An overview of the BlueGene/L SupercomputerabstractThis paper gives an overview of the BlueGene/L Supercomputer. This is a jointly funded research partnership between IBM and the Lawrence Livermore National Laboratory as part of the United States Department of Energy ASCI Advanced Architecture Research Program. Application performance and scaling studies have recently been initiated with partners at a number of academic and government institutions,including the San Diego Supercomputer Center and the California Institute of Technology. This massively parallel system of 65,536 nodes is based on a new architecture that exploits system-on-a-chip technology to deliver target peak processing power of 360 teraFLOPS (trillion floating-point operations per second). The machine is scheduled to be operational in the 2004-2005 time frame, at price/performance and power consumption/performance targets unobtainable with conventional architectures. Narasimha R. Adiga, Gheorghe Almási 0001, George S. Almási, Yariv Aridor, Rajkishore Barik, Daniel K. Beece, Ralph Bellofatto, Gyan Bhanot, Randy Bickford, Matthias A. Blumrich, Arthur A. Bright, José R. Brunheroto, Calin Cascaval, José G. Castaños, Waiman Chan, Luis Ceze, Paul Coteus, Siddhartha Chatterjee, Dong Chen 0005, George L.-T. Chiu, Thomas M. Cipolla, Paul Crumley, K. M. Desai, Alina Deutsch, Tamar Domany, Marc Boris Dombrowa, Wilm E. Donath, Maria Eleftheriou, C. Christopher Erway, J. Esch, Blake G. Fitch, Joseph Gagliano, Alan Gara, Rahul Garg 0001, Robert S. Germain, Mark Giampapa, Balaji Gopalsamy, John A. Gunnels, Manish Gupta 0002, Fred G. Gustavson, Shawn Hall, Ruud A. Haring, David F. Heidel, Philip Heidelberger, Lorraine M. Herger, Dirk Hoenicke, R. D. Jackson, T. Jamal-Eddine, Gerard V. Kopcsay, Elie Krevat, Manish P. Kurhekar, Alphonso P. Lanzetta, Derek Lieber, L. K. Liu, M. Lu, Mark P. Mendell, A. Misra, Yosef Moatti, Lawrence S. Mok, José E. Moreira, Ben J. Nathanson, Matthew Newton, Martin Ohmacht, Adam J. Oliner, Vinayaka Pandit, R. B. Pudota, Rick A. Rand, Richard D. Regan, Bradley Rubin, Albert E. Ruehli, Silvius Vasile Rus, Ramendra K. Sahoo, Alda Sanomiya, Eugen Schenfeld, M. Sharma, Edi Shmueli, Sarabjeet Singh, Peilin Song, Vijay Srinivasan, Burkhard D. Steinmacher-Burow, Karin Strauss, Christopher W. Surovic, Richard A. Swetz, Todd Takken, R. Brett Tremaine, Mickey Tsao, Arun R. Umamaheshwaran, P. Verma, Pavlos Vranas, T. J. Christopher Ward, Michael E. Wazlowski, W. Barrett, C. Engel, B. Drehmel, B. Hilgart, D. Hill, F. Kasemkhani, David J. Krolak, Chun-Tao Li 0001, Thomas A. Liebsch, James A. Marcella, A. Muff, A. Okomo, M. Rouse, A. Schram, M. Tubbs, G. Ulsh, Charles D. Wait, J. Wittrup, Myung Bae, Kenneth A. Dockser, Lynn Kissel, Mark K. Seager, Jeffrey S. Vetter, K. Yates |
SC | 81 |