Ariel Eizenberg

dblp:180/8191 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3Software engineering, systems software and programming languages · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
4 papers
Concurrent programming · 87% Runtime systems and virtual machines · 10% Program analysis · 3%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 75% GPUs and heterogeneous computing · 20% Performance modeling and evaluation · 5%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Concurrent programming
concurrency bugs
0.932018
SOFRITAS: Serializable Ordering-Free Regions for Increasing Thread Atomicity Scalably · ASPLOS 2018
BARRACUDA: binary-level analysis of runtime RAces in CUDA programs · PLDI 2017
Remix: online detection and repair of cache contention for the JVM · PLDI 2016
Memory systems
cache coherence
0.522017
TMI: thread memory isolation for false sharing repair · MICRO 2017
Remix: online detection and repair of cache contention for the JVM · PLDI 2016
Concurrent programming › concurrency bug detection
atomicity violation detection
0.312018
SOFRITAS: Serializable Ordering-Free Regions for Increasing Thread Atomicity Scalably · ASPLOS 2018
Concurrent programming
memory models
0.312018
SOFRITAS: Serializable Ordering-Free Regions for Increasing Thread Atomicity Scalably · ASPLOS 2018
Concurrent programming › concurrency control
serializability
0.312018
SOFRITAS: Serializable Ordering-Free Regions for Increasing Thread Atomicity Scalably · ASPLOS 2018
Concurrent programming › concurrency bug detection
data race detection
0.312017
BARRACUDA: binary-level analysis of runtime RAces in CUDA programs · PLDI 2017
GPUs and heterogeneous computing › GPU programming
CUDA
0.312017
BARRACUDA: binary-level analysis of runtime RAces in CUDA programs · PLDI 2017
Memory systems › cache coherence
false sharing
0.312017
TMI: thread memory isolation for false sharing repair · MICRO 2017
Runtime systems and virtual machines › virtual machine implementation
java virtual machine
0.212016
Remix: online detection and repair of cache contention for the JVM · PLDI 2016
Memory systems › memory interference
cache contention
0.212016
Remix: online detection and repair of cache contention for the JVM · PLDI 2016
Program analysis
dynamic analysis
0.112017
BARRACUDA: binary-level analysis of runtime RAces in CUDA programs · PLDI 2017
Performance modeling and evaluation › performance monitoring
hardware performance counters
0.112016
Remix: online detection and repair of cache contention for the JVM · PLDI 2016

Methods — techniques the papers use, named apart from their topics

binary-level analysis · 0.6runtime detection · 0.5JIT compilation · 0.5runtime system · 0.3concurrency bug finding · 0.3
YearPublicationVenuePosition
2018 SOFRITAS: Serializable Ordering-Free Regions for Increasing Thread Atomicity Scalably
abstract
Correctly synchronizing multithreaded programs is challenging and errors can lead to program failures such as atomicity violations. Existing strong memory consistency models rule out some possible failures, but are limited by depending on programmer-defined locking code. We present the new Ordering-Free Region (OFR) serializability consistency model that ensures atomicity for OFRs, which are spans of dynamic instructions between consecutive ordering constructs (e.g., barriers), without breaking atomicity at lock operations. Our platform, Serializable Ordering-Free Regions for Increasing Thread Atomicity Scalably (SOFRITAS), ensures a C/C++ program's execution is equivalent to a serialization of OFRs by default. We build two systems that realize the SOFRITAS idea: a concurrency bug finding tool for testing called SOFRITEST, and a production runtime system called SOPRO. SOFRITEST uses OFRs to find concurrency bugs, including a multi-critical-section atomicity violation in memcached that weaker consistency models will miss. If OFR's are too coarse-grained, SOFRITEST suggests refinement annotations automatically. Our software-only SOPRO implementation has high performance, scales well with increased parallelism, and prevents failures despite bugs in locking code. SOFRITAS has an average overhead of just 1.59x on a single-threaded execution and 1.51x on sixteen threads, despite pthreads' much weaker memory model.
Christian DeLozier, Ariel Eizenberg, Brandon Lucia, Joseph Devietti
ASPLOS2
2018 SLIMFAST: Reducing Metadata Redundancy in Sound and Complete Dynamic Data Race Detection
abstract
Data races are one of the main culprits behind the complexity of multithreaded programming. Existing data race detectors require large amounts of metadata for each program variable to perform their analyses. The SLIMFAST system exploits the insight that there is a large amount of redundancy in this metadata: many program variables often have identical metadata state. By sharing metadata across variables, a large reduction in space usage can be realized. SLIMFAST uses immutable metadata to safely support metadata sharing across threads while also accelerating concurrency control. SLIMFAST's lossless metadata compression achieves these benefits while preserving soundness and completeness. Across a range of benchmarks from Java Grande, DaCapo, NAS Parallel Benchmarks and Oracle's BerkeleyDB Java Edition, SLIMFAST is able to reduce memory consumption by 1.83x on average, and up to 4.90x for some benchmarks, compared to the state-of-the-art FASTTRACK system. By improving cache locality and simplifying concurrency control, SLIMFAST also accelerates data race detection by 1.40x on average, and up to 8.8x for some benchmarks, compared to FASTTRACK.
Yuanfeng Peng, Christian DeLozier, Ariel Eizenberg, William Mansky, Joseph Devietti
IPDPS3
2017 TMI: thread memory isolation for false sharing repair
abstract
Cache contention in the form of false sharing and true sharing arises when threads overshare cache lines at high frequency. Such oversharing can reduce or negate the performance benefits of parallel execution. Prior systems for detecting and repairing cache contention lack efficiency in detection or repair, contain subtle memory consistency flaws, or require invasive changes to the program environment.
Christian DeLozier, Ariel Eizenberg, Shiliang Hu, Gilles Pokam, Joseph Devietti
MICRO2
2017 BARRACUDA: binary-level analysis of runtime RAces in CUDA programs
abstract
GPU programming models enable and encourage massively parallel programming with over a million threads, requiring extreme parallelism to achieve good performance. Massive parallelism brings significant correctness challenges by increasing the possibility for bugs as the number of thread interleavings balloons. Conventional dynamic safety analyses struggle to run at this scale.
Ariel Eizenberg, Yuanfeng Peng, Toma Pigli, William Mansky, Joseph Devietti
PLDI1
2016 Remix: online detection and repair of cache contention for the JVM
abstract
As ever more computation shifts onto multicore architectures, it is increasingly critical to find effective ways of dealing with multithreaded performance bugs like true and false sharing. Previous approaches to fixing false sharing in unmanaged languages have employed highly-invasive runtime program modifications. We observe that managed language runtimes, with garbage collection and JIT code compilation, present unique opportunities to repair such bugs directly, mirroring the techniques used in manual repairs. We present Remix, a modified version of the Oracle HotSpot JVM which can detect cache contention bugs and repair false sharing at runtime. Remix's detection mechanism leverages recent performance counter improvements on Intel platforms, which allow for precise, unobtrusive monitoring of cache contention at the hardware level. Remix can detect and repair known false sharing issues in the LMAX Disruptor high-performance inter-thread messaging library and the Spring Reactor event-processing framework, automatically providing 1.5-2x speedups over unoptimized code and matching the performance of hand-optimization. Remix also finds a new false sharing bug in SPECjvm2008, and uncovers a true sharing bug in the HotSpot JVM that, when fixed, improves the performance of three NAS Parallel Benchmarks by 7-25x. Remix incurs no statistically-significant performance overhead on other benchmarks that do not exhibit cache contention, making Remix practical for always-on use.
Ariel Eizenberg, Shiliang Hu, Gilles Pokam, Joseph Devietti
PLDI1