Zachary Benavides

dblp:180/8543 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2019
0000-0003-3479-4518ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 49% Parallel and multicore computing · 26% Distributed systems · 25%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
bottleneck analysis
0.212016
Parallel Execution Profiles · HPDC 2016
Parallel and multicore computing › parallel computing › parallel program analysis
parallelism profiling
0.212016
Parallel Execution Profiles · HPDC 2016
Performance modeling and evaluation
profiling
0.112019
DProf: distributed profiler with strong guarantees · Proc. ACM Program. Lang. 2019
Parallel and multicore computing › thread-level parallelism
multithreaded applications
0.112016
Parallel Execution Profiles · HPDC 2016
Parallel and multicore computing › synchronization
thread synchronization
0.112016
Parallel Execution Profiles · HPDC 2016

Methods — techniques the papers use, named apart from their topics

timestamp synchronization · 0.4context-sensitive profiling · 0.4causal profiling · 0.4execution profiling · 0.2code region annotation · 0.2
YearPublicationVenuePosition
2019 Annotation guided collection of context-sensitive parallel execution profiles
Zachary Benavides, Keval Vora, Rajiv Gupta 0001, Xiangyu Zhang 0001
Formal Methods Syst. Des.1
2019 DProf: distributed profiler with strong guarantees
abstract
Performance analysis of a distributed system is typically achieved by collecting profiles whose underlying events are timestamped with unsynchronized clocks of multiple machines in the system. To allow comparison of timestamps taken at different machines, several timestamp synchronization algorithms have been developed. However, the inaccuracies associated with these algorithms can lead to inaccuracies in the final results of performance analysis. To address this problem, in this paper, we develop a system for constructing distributed performance profiles called DProf. At the core of DProf is a new timestamp synchronization algorithm, FreeZer, that tightly bounds the inaccuracy in a converted timestamp to a time interval. This not only allows timestamps from different machines to be compared, it also enables maintaining strong guarantees throughout the comparison which can be carefully transformed into guarantees for analysis results. To demonstrate the utility of DProf, we use it to implement dCSP and dCOZ that are accuracy bounded distributed versions of Context Sensitive Profiles and Causal Profiles developed for shared memory systems. While dCSP enables user to ascertain existence of a performance bottleneck, dCOZ estimates the expected performance benefit from eliminating that bottleneck. Experiments with three distributed applications on a cluster of heterogeneous machines validate that inferences via dCSP and dCOZ are highly accurate. Moreover, if FreeZer is replaced by two existing timestamp algorithms (linear regression & convex hull), the inferences provided by dCSP and dCOZ are severely degraded.
Zachary Benavides, Keval Vora, Rajiv Gupta 0001
Proc. ACM Program. Lang.1
2018 COMPI: Concolic Testing for MPI Applications
abstract
MPI is widely used as the bedrock of HPC applications, but there are no effective systematic software testing techniques for MPI programs. In this paper we develop COMPI, the first practical concolic testing tool for MPI applications. COMPI tackles two major challenges. First, it provides an automated testing tool for MPI programs - it performs concolic execution on a single process and records branch coverage across all. Infusing MPI semantics such as MPI rank and MPI_COMM_WORLD into COMPI enables it to automatically direct testing with various processes' executions as well as automatically determine the total number of processes used in the testing. Second, COMPI employs three techniques to effectively control the cost of testing as too high a cost may prevent its adoption. By capping input values, COMPI is made practical as too large an input can make the testing extremely slow and sometimes even fail as memory needed could exceed the computing platform's memory limit. With two-way instrumentation, we reduce the unnecessary memory and I/O overhead of COMPI and the target program. With constraint set reduction, COMPI keeps significantly fewer constraints by removing redundant ones in the presence of loops so as to avoid redundant tests against these branches. Our evaluation of COMPI uncovered four new bugs in a complex application and achieved 69-86% branch coverage which far exceeds the 1.8-38% coverage achieved via random testing.
Hongbo Li 0006, Sihuan Li, Zachary Benavides, Zizhong Chen, Rajiv Gupta 0001
IPDPS3
2017 Annotation Guided Collection of Context-Sensitive Parallel Execution Profiles
Zachary Benavides, Rajiv Gupta 0001, Xiangyu Zhang 0001
RV1
2016 Parallel Execution Profiles
abstract
Observing the relative behavior of an application's threads is critical to identifying performance bottlenecks and understanding their root causes. We present parallel execution profiles (PEPs), which capture the relative behavior of parallel threads in terms of the user selected code regions they execute. The user annotates the program to identify code regions of interest. The PEP divides the execution time of a multithreaded application into time intervals or a sequence of frames during which the code regions being executed in parallel by application threads remain the same. PEPs can be easily analyzed to compute execution times spent by the application in interesting behavior states. This helps user understand the severity of common performance problems such as excessive waiting on events by threads, threads contending for locks, and the presence of straggler threads.
Zachary Benavides, Rajiv Gupta 0001, Xiangyu Zhang 0001
HPDC1