EDBT 2026 Demo / reviewers in the wild / expert
Kevin E. Moore
dblp:79/43
· DBLP profile ↗
7ranked-venue papers
1as first author
0since 2021 · last 2007
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-authorSoftware engineering, systems software and programming languages · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Parallel and multicore computing · 40% Memory systems · 32% Performance modeling and evaluation · 13% | |
| Software engineering, system software, and programming languages
1 paper |
Programming languages and type systems · 100% |
Topics — the 17 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
transactional memory |
0.2 | 4 | 2007 | Performance pathologies in hardware transactional memory · ISCA 2007 LogTM: log-based transactional memory · HPCA 2006 Supporting nested transactional memory in logTM · ASPLOS 2006 |
Memory systems
cache coherence |
0.2 | 4 | 2007 | LogTM-SE: Decoupling Hardware Transactional Memory from Caches · HPCA 2007 LogTM: log-based transactional memory · HPCA 2006 Supporting nested transactional memory in logTM · ASPLOS 2006 |
Parallel and multicore computing › transactional memory
conflict detection |
0.1 | 2 | 2006 | LogTM: log-based transactional memory · HPCA 2006 Supporting nested transactional memory in logTM · ASPLOS 2006 |
Performance modeling and evaluation
workload characterization |
0.1 | 2 | 2007 | Performance pathologies in hardware transactional memory · ISCA 2007 Memory System Behavior of Java-Based Middleware · HPCA 2003 |
Parallel and multicore computing › transactional memory
hardware transactional memory |
0.1 | 1 | 2007 | LogTM-SE: Decoupling Hardware Transactional Memory from Caches · HPCA 2007 |
Memory systems › cache coherence
directory-based coherence |
0.1 | 1 | 2006 | LogTM: log-based transactional memory · HPCA 2006 |
Memory systems › cache coherence › cache coherence protocol
MOESI |
0.1 | 1 | 2006 | LogTM: log-based transactional memory · HPCA 2006 |
Distributed systems › transaction processing
nested transactions |
0.1 | 1 | 2006 | Supporting nested transactional memory in logTM · ASPLOS 2006 |
Memory systems
cache |
0.0 | 1 | 2003 | Memory System Behavior of Java-Based Middleware · HPCA 2003 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2003 | Memory System Behavior of Java-Based Middleware · HPCA 2003 |
Performance modeling and evaluation › workload characterization
memory system behavior |
0.0 | 1 | 2003 | Memory System Behavior of Java-Based Middleware · HPCA 2003 |
Processor architecture and microarchitecture
multiprocessor architecture |
0.0 | 1 | 2000 | Timestamp snooping: an approach for extending SMPs · ASPLOS 2000 |
Processor architecture and microarchitecture › multiprocessor architecture
symmetric multiprocessor |
0.0 | 1 | 2000 | Timestamp snooping: an approach for extending SMPs · ASPLOS 2000 |
Parallel and multicore computing
speculative parallelization |
0.0 | 1 | 2007 | Performance pathologies in hardware transactional memory · ISCA 2007 |
Programming languages and type systems › module systems
software composition |
0.0 | 1 | 2006 | Supporting nested transactional memory in logTM · ASPLOS 2006 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2006 | LogTM: log-based transactional memory · HPCA 2006 |
Cloud and datacenter computing
server workloads |
0.0 | 1 | 2003 | Memory System Behavior of Java-Based Middleware · HPCA 2003 |
Methods — techniques the papers use, named apart from their topics
log-based transactional memory · 0.1escape actions · 0.1simulation · 0.1performance pathology analysis · 0.1version management · 0.1lazy cleanup · 0.1full-system simulation · 0.0benchmark characterization · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2007 | LogTM-SE: Decoupling Hardware Transactional Memory from CachesabstractThis paper proposes a hardware transactional memory (HTM) system called LogTM Signature Edition (LogTM-SE). LogTM-SE uses signatures to summarize a transactions read-and write-sets and detects conflicts on coherence requests (eager conflict detection). Transactions update memory "in place" after saving the old value in a per-thread memory log (eager version management). Finally, a transaction commits locally by clearing its signature, resetting the log pointer, etc., while aborts must undo the log. LogTM-SE achieves two key benefits. First, signatures and logs can be implemented without changes to highly-optimized cache arrays because LogTM-SE never moves cached data, changes a blocks cache state, or flash clears bits in the cache. Second, transactions are more easily virtualized because signatures and logs are software accessible, allowing the operating system and runtime to save and restore this state. In particular, LogTM-SE allows cache victimization, unbounded nesting (both open and closed), thread context switching and migration, and paging Luke Yen, Jayaram Bobba, Michael R. Marty, Kevin E. Moore, Haris Volos 0001, Mark D. Hill, Michael M. Swift, David A. Wood 0001 |
HPCA | 4 |
| 2007 | Performance pathologies in hardware transactional memoryabstractHardware Transactional Memory (HTM) systems reflect choices from three key design dimensions: conflict detection, version management, and conflict resolution. Previously proposed HTMs represent three points in this design space: lazy conflict detection, lazy version management, committer wins (LL); eager conflict detection, lazy version management, requester wins (EL); and eager conflict detection, eager version management, and requester stalls with conservative deadlock avoidance (EE). To isolate the effects of these high-level design decisions, we develop a common framework that abstracts away differences in cache write policies, interconnects, and ISA to compare these three design points. Not surprisingly, the relative performance of these systems depends on the workload. Under light transactional loads they perform similarly, but under heavy loads they differ by up to 80%. None of the systems performs best on all of our benchmarks. We identify seven performance pathologies-interactions between workload and system that degrade performance-as the root cause of many performance differences: FriendlyFire, StarvingWriter, SerializedCommit, FutileStall, StarvingElder, RestartConvoy, and DuelingUpgrades. We discuss when and on which systems these pathologies can occur and show that they actually manifest within TM workloads. The insight provided by these pathologies motivated four enhanced systems that often significantly reduce transactional memory overhead. Importantly, by avoiding transaction pathologies, each enhanced system performs well across our suite of benchmarks. Jayaram Bobba, Kevin E. Moore, Haris Volos 0001, Luke Yen, Mark D. Hill, Michael M. Swift, David A. Wood 0001 |
ISCA | 2 |
| 2006 | Supporting nested transactional memory in logTMabstractNested transactional memory (TM) facilitates software composition by letting one module invoke another without either knowing whether the other uses transactions. Closed nested transactions extend isolation of an inner transaction until the toplevel transaction commits. Implementations may flatten nested transactions into the top-level one, resulting in a complete abort on conflict, or allow partial abort of inner transactions. Open nested transactions allow a committing inner transaction to immediately release isolation, which increases parallelism and expressiveness at the cost of both software and hardware complexity.This paper extends the recently-proposed flat Log-based Transactional Memory (LogTM) with nested transactions. Flat LogTM saves pre-transaction values in a log, detects conflicts with read (R) and write (W) bits per cache block, and, on abort, invokes a software handler to unroll the log. Nested LogTM supports nesting by segmenting the log into a stack of activation records and modestly replicating R/W bits. To facilitate composition with nontransactional code, such as language runtime and operating system services, we propose escape actions that allow trusted code to run outside the confines of the transactional memory system. Michelle J. Moravan, Jayaram Bobba, Kevin E. Moore, Luke Yen, Mark D. Hill, Ben Liblit, Michael M. Swift, David A. Wood 0001 |
ASPLOS | 3 |
| 2006 | LogTM: log-based transactional memoryabstractTransactional memory (TM) simplifies parallel programming by guaranteeing that transactions appear to execute atomically and in isolation. Implementing these properties includes providing data version management for the simultaneous storage of both new (visible if the transaction commits) and old (retained if the transaction aborts) values. Most (hardware) TM systems leave old values "in place" (the target memory address) and buffer new values elsewhere until commit. This makes aborts fast, but penalizes (the much more frequent) commits. In this paper, we present a new implementation of transactional memory, log-based transactional memory (LogTM), that makes commits fast by storing old values to a per-thread log in cacheable virtual memory and storing new values in place. LogTM makes two additional contributions. First, LogTM extends a MOESI directory protocol to enable both fast conflict detection on evicted blocks and fast commit (using lazy cleanup). Second, LogTM handles aborts in (library) software with little performance penalty. Evaluations running micro- and SPLASH-2 benchmarks on a 32-way multiprocessor support our decision to optimize for commit by showing that only 1-2% of transactions abort. Kevin E. Moore, Jayaram Bobba, Michelle J. Moravan, Mark D. Hill, David A. Wood 0001 |
HPCA | 1 |
| 2005 | Exploring Processor Design Options for Java-Based MiddlewareabstractJava-based middleware is a rapidly growing workload for high-end server processors, particularly chip multiprocessors (CMP). To help architects design future microprocessors to run this important new workload, we provide a detailed characterization of two popular Java server benchmarks, ECperf and SPECjbb2000. We first estimate the amount of instruction-level parallelism in these workloads by simulating a very wide issue processor with perfect caches and perfect branch predictors. We then identify performance bottlenecks for these workloads on a more realistic processor by selectively idealizing individual processor structures. Finally, we combine our findings on available ILP in Java middleware with results from previous papers that characterize the availibility of TLP to investigate the optimal balance between ILP and TLP in CMPs. We find that, like other commercial workloads, Java middleware has only a small amount of instruction-level parallelism, even when run on very aggressive processors. When run on processors resembling currently available processors, the performance of Java middleware is limited by frequent traps, address translation and stalls in the memory system. We find that SPECjbb2000 differs from ECperf in two meaningful ways: (1) the performance of ECperf is affected much more by cache and TLB misses during instruction fetch and (2) SPECjbb2000 has more memory-level parallelism. Martin Karlsson, Erik Hagersten, Kevin E. Moore, David A. Wood 0001 |
ICPP | 3 |
| 2003 | Memory System Behavior of Java-Based MiddlewareabstractIn this paper, we present a detailed characterization of the memory system, behavior of ECperf and SPECjbb using both commercial server hardware and Simics full-system simulation. We find that the memory footprint and primary working sets of these workloads are small compared to other commercial workloads (e.g. on-line transaction processing), and that a large fraction of the working sets are shared between processors. We observed two key differences between ECperf and SPECjbb that highlight the importance of isolating the behavior of the middle tier. First, ECperf has a larger instruction footprint, resulting in much higher miss rates for intermediate-size instruction caches. Second, SPECjbb's data set size increases linearly as the benchmark scales up, while ECperf's remains roughly constant. This difference can lead to opposite conclusions on the design of multiprocessor memory systems, such as the utility of moderate sized (i.e. 1 MB) shared caches in a chip multiprocessor. Martin Karlsson, Kevin E. Moore, Erik Hagersten, David A. Wood 0001 |
HPCA | 2 |
| 2000 | Timestamp snooping: an approach for extending SMPs
Milo M. K. Martin, Daniel J. Sorin, Anastasia Ailamaki, Alaa R. Alameldeen, Ross M. Dickson, Carl J. Mauer, Kevin E. Moore, Manoj Plakal, Mark D. Hill, David A. Wood 0001 |
ASPLOS | 7 |