EDBT 2026 Demo / reviewers in the wild / expert
Zachary Benavides
dblp:180/8543
· DBLP profile ↗
5ranked-venue papers
4as first author
0since 2021 · last 2019
0000-0003-3479-4518ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorTheory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Performance modeling and evaluation · 49% Parallel and multicore computing · 26% Distributed systems · 25% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
bottleneck analysis |
0.2 | 1 | 2016 | Parallel Execution Profiles · HPDC 2016 |
Parallel and multicore computing › parallel computing › parallel program analysis
parallelism profiling |
0.2 | 1 | 2016 | Parallel Execution Profiles · HPDC 2016 |
Performance modeling and evaluation
profiling |
0.1 | 1 | 2019 | DProf: distributed profiler with strong guarantees · Proc. ACM Program. Lang. 2019 |
Parallel and multicore computing › thread-level parallelism
multithreaded applications |
0.1 | 1 | 2016 | Parallel Execution Profiles · HPDC 2016 |
Parallel and multicore computing › synchronization
thread synchronization |
0.1 | 1 | 2016 | Parallel Execution Profiles · HPDC 2016 |
Methods — techniques the papers use, named apart from their topics
timestamp synchronization · 0.4context-sensitive profiling · 0.4causal profiling · 0.4execution profiling · 0.2code region annotation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Annotation guided collection of context-sensitive parallel execution profiles
Zachary Benavides, Keval Vora, Rajiv Gupta 0001, Xiangyu Zhang 0001 |
Formal Methods Syst. Des. | 1 |
| 2019 | DProf: distributed profiler with strong guaranteesabstractPerformance analysis of a distributed system is typically achieved by collecting profiles whose underlying events are timestamped with unsynchronized clocks of multiple machines in the system. To allow comparison of timestamps taken at different machines, several timestamp synchronization algorithms have been developed. However, the inaccuracies associated with these algorithms can lead to inaccuracies in the final results of performance analysis. To address this problem, in this paper, we develop a system for constructing distributed performance profiles called DProf. At the core of DProf is a new timestamp synchronization algorithm, FreeZer, that tightly bounds the inaccuracy in a converted timestamp to a time interval. This not only allows timestamps from different machines to be compared, it also enables maintaining strong guarantees throughout the comparison which can be carefully transformed into guarantees for analysis results. To demonstrate the utility of DProf, we use it to implement dCSP and dCOZ that are accuracy bounded distributed versions of Context Sensitive Profiles and Causal Profiles developed for shared memory systems. While dCSP enables user to ascertain existence of a performance bottleneck, dCOZ estimates the expected performance benefit from eliminating that bottleneck. Experiments with three distributed applications on a cluster of heterogeneous machines validate that inferences via dCSP and dCOZ are highly accurate. Moreover, if FreeZer is replaced by two existing timestamp algorithms (linear regression & convex hull), the inferences provided by dCSP and dCOZ are severely degraded. Zachary Benavides, Keval Vora, Rajiv Gupta 0001 |
Proc. ACM Program. Lang. | 1 |
| 2018 | COMPI: Concolic Testing for MPI ApplicationsabstractMPI is widely used as the bedrock of HPC applications, but there are no effective systematic software testing techniques for MPI programs. In this paper we develop COMPI, the first practical concolic testing tool for MPI applications. COMPI tackles two major challenges. First, it provides an automated testing tool for MPI programs - it performs concolic execution on a single process and records branch coverage across all. Infusing MPI semantics such as MPI rank and MPI_COMM_WORLD into COMPI enables it to automatically direct testing with various processes' executions as well as automatically determine the total number of processes used in the testing. Second, COMPI employs three techniques to effectively control the cost of testing as too high a cost may prevent its adoption. By capping input values, COMPI is made practical as too large an input can make the testing extremely slow and sometimes even fail as memory needed could exceed the computing platform's memory limit. With two-way instrumentation, we reduce the unnecessary memory and I/O overhead of COMPI and the target program. With constraint set reduction, COMPI keeps significantly fewer constraints by removing redundant ones in the presence of loops so as to avoid redundant tests against these branches. Our evaluation of COMPI uncovered four new bugs in a complex application and achieved 69-86% branch coverage which far exceeds the 1.8-38% coverage achieved via random testing. Hongbo Li 0006, Sihuan Li, Zachary Benavides, Zizhong Chen, Rajiv Gupta 0001 |
IPDPS | 3 |
| 2017 | Annotation Guided Collection of Context-Sensitive Parallel Execution Profiles
Zachary Benavides, Rajiv Gupta 0001, Xiangyu Zhang 0001 |
RV | 1 |
| 2016 | Parallel Execution ProfilesabstractObserving the relative behavior of an application's threads is critical to identifying performance bottlenecks and understanding their root causes. We present parallel execution profiles (PEPs), which capture the relative behavior of parallel threads in terms of the user selected code regions they execute. The user annotates the program to identify code regions of interest. The PEP divides the execution time of a multithreaded application into time intervals or a sequence of frames during which the code regions being executed in parallel by application threads remain the same. PEPs can be easily analyzed to compute execution times spent by the application in interesting behavior states. This helps user understand the severity of common performance problems such as excessive waiting on events by threads, threads contending for locks, and the presence of straggler threads. Zachary Benavides, Rajiv Gupta 0001, Xiangyu Zhang 0001 |
HPDC | 1 |