VLDB 2026 Research / reviewers in the wild / expert
Lance Hammond
dblp:08/4658
· DBLP profile ↗
8ranked-venue papers
3as first author
0since 2021 · last 2006
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-authorSoftware engineering, systems software and programming languages · 6 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Parallel and multicore computing · 50% Memory systems · 33% Processor architecture and microarchitecture · 16% | |
| Software engineering, system software, and programming languages
1 paper |
Concurrent programming · 77% Programming languages and type systems · 23% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
transactional memory |
0.1 | 2 | 2004 | Transactional Memory Coherence and Consistency · ISCA 2004 Programming with transactional coherence and consistency (TCC) · ASPLOS 2004 |
Concurrent programming
memory models |
0.0 | 1 | 2004 | Programming with transactional coherence and consistency (TCC) · ASPLOS 2004 |
Memory systems › memory consistency
memory consistency model |
0.0 | 1 | 2004 | Transactional Memory Coherence and Consistency · ISCA 2004 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2004 | Programming with transactional coherence and consistency (TCC) · ASPLOS 2004 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 2 | 1998 | Data Speculation Support for a Chip Multiprocessor · ASPLOS 1998 The Case for a Single-Chip Multiprocessor · ASPLOS 1996 |
Memory systems
cache coherence |
0.0 | 3 | 2004 | Transactional Memory Coherence and Consistency · ISCA 2004 Data Speculation Support for a Chip Multiprocessor · ASPLOS 1998 Evaluation of Design Alternatives for a Multiprocessor Microprocessor · ISCA 1996 |
Parallel and multicore computing › speculative parallelization
thread-level speculation |
0.0 | 1 | 1998 | Data Speculation Support for a Chip Multiprocessor · ASPLOS 1998 |
Memory systems
cache design |
0.0 | 1 | 1996 | Evaluation of Design Alternatives for a Multiprocessor Microprocessor · ISCA 1996 |
Memory systems › cache › multiprocessor cache
shared cache |
0.0 | 1 | 1996 | Evaluation of Design Alternatives for a Multiprocessor Microprocessor · ISCA 1996 |
Processor architecture and microarchitecture
superscalar processor |
0.0 | 1 | 1996 | The Case for a Single-Chip Multiprocessor · ASPLOS 1996 |
Memory systems › cache coherence
invalidation |
0.0 | 1 | 1996 | Evaluation of Design Alternatives for a Multiprocessor Microprocessor · ISCA 1996 |
Parallel and multicore computing › parallel architecture
on-chip parallelism |
0.0 | 1 | 1996 | The Case for a Single-Chip Multiprocessor · ASPLOS 1996 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 1996 | Evaluation of Design Alternatives for a Multiprocessor Microprocessor · ISCA 1996 |
Methods — techniques the papers use, named apart from their topics
transactional execution · 0.1speculative execution · 0.0hardware-controlled rollback · 0.0atomic transaction broadcast · 0.0speculation control handlers · 0.0full-system simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2006 | Executing Java programs with transactional memory
Brian D. Carlstrom, JaeWoong Chung, Hassan Chafi, Austen McDonald, Chi Cao Minh, Lance Hammond, Christoforos E. Kozyrakis, Kunle Olukotun |
Sci. Comput. Program. | 6 |
| 2005 | TAPE: a transactional application profiling environmentabstractTransactional Coherence and Consistency (TCC) provides a new parallel programming model that uses transactions as the basic unit of parallel work and communication. TCC simplifies the development of correct parallel code because hardware provides transaction atomicity and ordering. Nevertheless, the programmer or a dynamic compiler must still optimize the parallel code for performance.This paper presents TAPE, a hardware and software infrastructure for profiling in TCC systems. TAPE extends the hardware for transactional execution to identify performance impediments such as dependence violations, buffer overflows, and work imbalance. It filters infrequent events to reduce resource requirements and allows the programmer to focus on the most important bottlenecks. We demonstrate that TAPE introduces minimal die area and performance overhead and can be used continuously, even for production runs. Moreover, we demonstrate how to leverage the profiling information to guide optimization for a set of parallel applications. TAPE accurately identifies the source code location and type of the most important bottlenecks, allowing a programmer to achieve maximum parallel speedup with a few profiling steps. Hassan Chafi, Chi Cao Minh, Austen McDonald, Brian D. Carlstrom, JaeWoong Chung, Lance Hammond, Christoforos E. Kozyrakis, Kunle Olukotun |
ICS | 6 |
| 2004 | Programming with transactional coherence and consistency (TCC)abstractTransactional Coherence and Consistency (TCC) offers a way to simplify parallel programming by executing all code within transactions. In TCC systems, transactions serve as the fundamental unit of parallel work, communication and coherence. As each transaction completes, it writes all of its newly produced state to shared memory atomically, while restarting other processors that have speculatively read stale data. With this mechanism, a TCC-based system automatically handles data synchronization correctly, without programmer intervention. To gain the benefits of TCC, programs must be decomposed into transactions. We describe two basic programming language constructs for decomposing programs into transactions, a loop conversion syntax and a general transaction-forking mechanism. With these constructs, writing correct parallel programs requires only small, incremental changes to correct sequential programs. The performance of these programs may then easily be optimized, based on feedback from real program execution, using a few simple techniques. Lance Hammond, Brian D. Carlstrom, Vicky Wong, Ben Hertzberg, Michael K. Chen 0001, Christoforos E. Kozyrakis, Kunle Olukotun |
ASPLOS | 1 |
| 2004 | Transactional Memory Coherence and ConsistencyabstractIn this paper, we propose a new shared memory model: transactional memory coherence and consistency (TCC). TCC provides a model in which atomic transactions are always the basic unit of parallel work, communication, memory coherence, and memory reference consistency. TCC greatly simplifies parallel software by eliminating the need for synchronization using conventional locks and semaphores, along with their complexities. TCC hardware must combine all writes from each transaction region in a program into a single packet and broadcast this packet to the permanent shared memory state atomically as a large block. This simplifies the coherence hardware because it reduces the need for small, low-latency messages and completely eliminates the need for conventional snoopy cache coherence protocols, as multiple speculatively written versions of a cache line may safely coexist within the system. Meanwhile, automatic, hardware-controlled rollback of speculative transactions resolves any correctness violations that may occur when several processors attempt to read and write the same data simultaneously. The cost of this simplified scheme is higher interprocessor bandwidth. To explore the costs and benefits of TCC, we study the characteristics of an optimal transaction-based memory system, and examine how different design parameters could affect the performance of real systems. Across a spectrum of applications, the TCC model itself did not limit available parallelism. Most applications are easily divided into transactions requiring only small write buffers, on the order of 4-8 KB. The broadcast requirements of TCC are high, but are well within the capabilities of CMPs and small-scale SMPs with high-speed interconnects. Lance Hammond, Vicky Wong, Michael K. Chen 0001, Brian D. Carlstrom, John D. Davis, Ben Hertzberg, Manohar K. Prabhu, Honggo Wijaya, Christoforos E. Kozyrakis, Kunle Olukotun |
ISCA | 1 |
| 1999 | Improving the performance of speculatively parallel applications on the Hydra CMPabstractArticle Free Access Share on Improving the performance of speculatively parallel applications on the Hydra CMP Authors: Kunle Olukotun Computer Systems Laboratory, Stanford University, Stanford, CA Computer Systems Laboratory, Stanford University, Stanford, CAView Profile , Lance Hammond Computer Systems Laboratory, Stanford University, Stanford, CA Computer Systems Laboratory, Stanford University, Stanford, CAView Profile , Mark Willey Computer Systems Laboratory, Stanford University, Stanford, CA Computer Systems Laboratory, Stanford University, Stanford, CAView Profile Authors Info & Claims ICS '99: Proceedings of the 13th international conference on SupercomputingJune 1999 Pages 21–30https://doi.org/10.1145/305138.305155Published:01 May 1999Publication History 59citation482DownloadsMetricsTotal Citations59Total Downloads482Last 12 Months16Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Kunle Olukotun, Lance Hammond, Mark Willey |
International Conference on Supercomputing | 2 |
| 1998 | Data Speculation Support for a Chip MultiprocessorabstractThread-level speculation is a technique that enables parallel execution of sequential applications on a multiprocessor. This paper describes the complete implementation of the support for threadlevel speculation on the Hydra chip multiprocessor (CMP). The support consists of a number of software speculation control handlers and modifications to the shared secondary cache memory system of the CMP This support is evaluated using five representative integer applications. Our results show that the speculative support is only able to improve performance when there is a substantial amount of medium--grained loop-level parallelism in the application. When the granularity of parallelism is too small or there is little inherent parallelism in the application, the overhead of the software handlers overwhelms any potential performance benefits from speculative-thread parallelism. Overall, thread-level speculation still appears to be a promising approach for expanding the class of applications that can be automatically parallelized, but more hardware intensive implementations for managing speculation control are required to achieve performance improvements on a wide class of integer applications. Lance Hammond, Mark Willey, Kunle Olukotun |
ASPLOS | 1 |
| 1996 | The Case for a Single-Chip MultiprocessorabstractAdvances in IC processing allow for more microprocessor design options. The increasing gate density and cost of wires in advanced integrated circuit technologies require that we look for new ways to use their capabilities effectively. This paper shows that in advanced technologies it is possible to implement a single-chip multiprocessor in the same area as a wide issue superscalar processor. We find that for applications with little parallelism the performance of the two microarchitectures is comparable. For applications with large amounts of parallelism at both the fine and coarse grained levels, the multiprocessor microarchitecture outperforms the superscalar architecture by a significant margin. Single-chip multiprocessor architectures have the advantage in that they offer localized implementation of a high-clock rate processor for inherently sequential applications and low latency interprocessor communication for parallel applications. Kunle Olukotun, Basem A. Nayfeh, Lance Hammond, Kenneth G. Wilson, Kunyung Chang |
ASPLOS | 3 |
| 1996 | Evaluation of Design Alternatives for a Multiprocessor MicroprocessorabstractIn the future, advanced integrated circuit processing and packaging technology will allow for several design options for multiprocessor microprocessors. In this paper we consider three architectures: shared-primary cache, shared-secondary cache, and shared-memory. We evaluate these three architectures using a complete system simulation environment which models the CPU, memory hierarchy and I/O devices in sufficient detail to boot and run a commercial operating system. Within our simulation environment, we measure performance using representative hand and compiler generated parallel applications, and a multiprogramming workload. Our results show that when applications exhibit fine-grained sharing, both shared-primary and shared-secondary architectures perform similarly when the full costs of sharing the primary cache are included. Basem A. Nayfeh, Lance Hammond, Kunle Olukotun |
ISCA | 2 |