Zarka Cvetanovic

dblp:42/2159 · DBLP profile ↗
← Back
12ranked-venue papers
9as first author
0since 2021 · last 2006
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 9 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Performance modeling and evaluation · 58% Parallel and multicore computing · 22% Memory systems · 8%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%

Topics — the 21 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › parallel system performance
multiprocessor performance evaluation
0.132003
Performance Analysis of the Alpha 21364-BAsed HP GS1280 Multiprocessor · ISCA 2003
Performance analysis of the Alpha 21264-based Compaq ES40 system · ISCA 2000
Perfect Benchmarks decomposition and performance on VAX multiprocessors · SC 1990
Performance modeling and evaluation › profiling
application profiling
0.112006
Performance tools - Performance tools for large-scale clusters · SC 2006
Performance modeling and evaluation
performance analysis tools
0.112006
Performance tools - Performance tools for large-scale clusters · SC 2006
Performance modeling and evaluation
workload characterization
0.142006
Performance tools - Performance tools for large-scale clusters · SC 2006
Performance Characterization of the Alpha 21164 Microprocessor Using TP and SPEC Workloads · HPCA 1996
Characterization of Alpha AXP Performance Using TP and SPEC Workloads · ISCA 1994
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor
0.012003
Performance Analysis of the Alpha 21364-BAsed HP GS1280 Multiprocessor · ISCA 2003
Parallel and multicore computing
parallel programming models
0.022000
Extending OpenMP for NUMA Machines · SC 2000
The Effects of Problem Partitioning, Allocation, and Granularity on the Performance of Multiple-Processor Systems · IEEE Trans. Computers 1987
Memory systems › non-uniform memory access
NUMA data placement
0.012000
Extending OpenMP for NUMA Machines · SC 2000
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP extensions
0.012000
Extending OpenMP for NUMA Machines · SC 2000
Processor architecture and microarchitecture
out-of-order execution
0.012000
Performance analysis of the Alpha 21264-based Compaq ES40 system · ISCA 2000
Interconnection networks and networks-on-chip › network topology
torus network
0.012003
Performance Analysis of the Alpha 21364-BAsed HP GS1280 Multiprocessor · ISCA 2003
Query processing and optimization
sorting
0.011994
AlphaSort: A RISC Machine Sort · SIGMOD Conference 1994
Performance modeling and evaluation
benchmarking
0.011994
AlphaSort: A RISC Machine Sort · SIGMOD Conference 1994
Performance modeling and evaluation › performance monitoring
hardware performance monitoring
0.011994
Characterization of Alpha AXP Performance Using TP and SPEC Workloads · ISCA 1994
Performance modeling and evaluation › benchmarking › algorithm benchmarking
sort benchmark
0.011994
AlphaSort: A RISC Machine Sort · SIGMOD Conference 1994
Compilers and program optimization
code generation
0.012000
Extending OpenMP for NUMA Machines · SC 2000
Memory systems
memory bandwidth
0.012000
Performance analysis of the Alpha 21264-based Compaq ES40 system · ISCA 2000
Parallel and multicore computing › parallelization strategies
parallel program decomposition
0.011990
Perfect Benchmarks decomposition and performance on VAX multiprocessors · SC 1990
Processor architecture and microarchitecture
superscalar processor
0.011996
Performance Characterization of the Alpha 21164 Microprocessor Using TP and SPEC Workloads · HPCA 1996
Performance modeling and evaluation
parallel system performance
0.011987
The Effects of Problem Partitioning, Allocation, and Granularity on the Performance of Multiple-Processor Systems · IEEE Trans. Computers 1987
Memory systems › cache
cache-aware algorithm design
0.011995
AlphaSort: A Cache-Sensitive Parallel External Sort · VLDB J. 1995
Memory systems
cache design
0.011990
Perfect Benchmarks decomposition and performance on VAX multiprocessors · SC 1990

Methods — techniques the papers use, named apart from their topics

profiling · 0.1hardware performance counters · 0.1high performance fortran directives · 0.1quicksort · 0.0file striping · 0.0measurement · 0.0parallel external sorting · 0.0cache-sensitive partitioning · 0.0replacement-selection · 0.0replacement selection · 0.0hardware monitor · 0.0
YearPublicationVenuePosition
2006 Performance tools - Performance tools for large-scale clusters
abstract
Identifying factors that limit large-scale cluster scalability remains a challenging area of research. Several efforts have focused on developing tools to address different aspects of this problem space. Presenters from both industry and research will describe their tools, including: monitoring cluster-wide utilizations of system components (CPUs, memory, I/O, interconnect), monitoring node-level load (memory bandwidth, cache/TLB misses, stall components), high-frequency, low-overhead, fine-grained application profiling.Our first goal is to discuss how such tools can be helpful for improving application performance, reducing hot-spots, and determining the best match between applications and platforms. The second goal is to solicit developers' and users' experiences with problems in this space and ideas and techniques that might help address them.
Zarka Cvetanovic
SC1
2004 Performance analysis tools for large-scale Linux clusters
abstract
As cluster computer environments increase in size and complexity, it is becoming more challenging to analyze and identify factors that limit performance and scalability. Easy-to-use tools that help identify such bottlenecks are crucial for tuning applications and configuring systems for best performance. We present a collection of visualization tools, which allow users to monitor load on all cluster components simultaneously, with negligible overhead, and no changes in the application. We include examples where the tools have been used to identify bottlenecks within a cluster and improve performance. We provide several examples of application profiles gathered using the tools and outline the methodology for projecting performance of future cluster platforms.
Zarka Cvetanovic
CLUSTER1
2003 Performance Analysis of the Alpha 21364-BAsed HP GS1280 Multiprocessor
abstract
This paper evaluates performance characteristics of the HP GS1280 shared memory multiprocessor system. The GS1280 system contains up to 64 Alpha 21364 CPUs connected together via a torus-based interconnect. We describe architectural features of the GS1280 system. We compare and contrast the GS1280 to the previousgeneration Alpha systems: AlphaServer GS320 and ES45/SC45. We further quantitatively show the performance effects of these features using application results and profiling data based on the built-in performance counters. We find that the HP GS1280 often provides 2 to 3 times the performance of the AlphaServer GS320 at similar clock frequencies. We find the key reasons for such performance gains are advances in memory, inter-processor, and I/O subsystem designs.
Zarka Cvetanovic
ISCA1
2000 Performance analysis of the Alpha 21264-based Compaq ES40 system
abstract
This paper evaluates performance characteristics of the Compaq ES40 shared memory multiprocessor. The ES40 system contains up to four Alpha 21264 CPU's together with a high-performance memory system. We qualitatively describe architectural features included in the 21264 microprocessor and the surrounding system chipset. We further quantitatively show the performance effects of these features using benchmark results and profiling data collected from industry-standard commercial and technical workloads. The profile data includes basic performance information - such as instructions per cycle, branch mispredicts, and cache misses - as well as other data that specifically characterizes the 21264. Wherever possible, we compare and contrast the ES40 to the AlphaServer 4100 - a previous-generation Alpha system containing four Alpha 21164 microprocessors - to highlight the architectural advances in the ES40. We find that the Compaq ES40 often provides 2 to 3 times the performance of the AlphaServer 4100 at similar clock frequencies. We also find that the ES40 memory system has about five times the memory bandwidth of the 4100. These performance improvements come from numerous microprocessor and platform enhancements, including out-of-order execution, branch prediction, functional units, and the memory system.
Zarka Cvetanovic, Richard E. Kessler
ISCA1
2000 Extending OpenMP for NUMA Machines
abstract
This paper describes extensions to OpenMP that implemen data placemen features needed for NUMA architectures. OpenMP is a collection of compiler directives and library routines used to write portable parallel programs for shared-memory architectures. Writing efficient parallel programs for NUMA architectures, which have characteristics of both shared-memory and distributed-memory architectures, requires that a programmer control the placement of data in memory and the placement of computations that operate on that data. Optimal performance is obtained when computations occur on processors that have fast access to the data needed by those computations. OpenMP-designed for shared-memory architectures-does not by itself address these issues. The extensions to OpenMP Fortran presented here have been mainly taken from High Performance Fortran. The paper describes some of the techniques that the Compaq Fortran compiler uses to generate efficient code based on these extensions. I also describes some additional compiler optimizations, and concludes with some preliminary results.
John Bircsak, Peter Craig, RaeLyn Crowell, Zarka Cvetanovic, Jonathan Harris, C. Alexander Nelson, Carl D. Offner
SC4
1996 Performance Characterization of the Alpha 21164 Microprocessor Using TP and SPEC Workloads
abstract
This paper compares the performance characteristics of the Alpha 21164 to the previous-generation 21064 microprocessor. Measurements on the 21164-based AlphaServer 8200 system are compared to the 21064-based DEC 7000 server using several commercial and technical workloads. The data analyzed includes cycles per instruction, multiple-issued instructions, branch predictions, stall components, cache misses, and instruction frequencies. The AlphaServer 8200 provides 2 to 3 times the performance of the DEC 7000 server based on the faster clock, larger on-chip cache, expanded multiple-issuing, and lower cache/memory latencies and higher bandwidth.
Zarka Cvetanovic, Dileep Bhandarkar
HPCA1
1995 AlphaSort: A Cache-Sensitive Parallel External Sort
Chris Nyberg, Tom Barclay, Zarka Cvetanovic, Jim Gray 0001, David B. Lomet
VLDB J.3
1994 Characterization of Alpha AXP Performance Using TP and SPEC Workloads
abstract
The characteristics of several commercial and technical workloads on the DEC 7000 AXP system are compared using built-in hardware monitors. The data analyzed include total instructions, cycles, multiple-issued instructions, stall components, cache misses, and instruction types. The data indicates that the two classes of workloads have vastly different characteristics and impose different requirements on the system design. Compared to VAX, Alpha AXP takes advantage of lower cycles per instruction and cycle time to achieve a significant performance advantage. The cache and memory interconnect subsystems are expected to play a crucial role in the performance of future systems. A simple model for evaluating the effects of various design tradeoffs based on the data collected by using hardware monitors is proposed.>
Zarka Cvetanovic, Dileep Bhandarkar
ISCA1
1994 AlphaSort: A RISC Machine Sort
abstract
A new sort algorithm, called AlphaSort, demonstrates that commodity processors and disks can handle commercial batch workloads. Using Alpha AXP processors, commodity memory, and arrays of SCSI disks, AlphaSort runs the industry-standard sort benchmark in seven seconds. This beats the best published record on a 32-cpu 32-disk Hypercube by 8:1. On another benchmark, AlphaSort sorted more than a gigabyte in a minute.AlphaSort is a cache-sensitive memory-intensive sort algorithm. It uses file striping to get high disk bandwidth. It uses QuickSort to generate runs and uses replacement-selection to merge the runs. It uses shared memory multiprocessors to break the sort into subsort chores.Because startup times are becoming a significant part of the total time, we propose two new benchmarks: (1) Minutesort: how much can you sort in a minute, and (2) DollarSort: how much can you sort for a dollar.
Chris Nyberg, Tom Barclay, Zarka Cvetanovic, Jim Gray 0001, David B. Lomet
SIGMOD Conference3
1991 Efficient decomposition and performance of parallel PDE, FFT, Monte Carlo simulations, simplex, and Sparse solvers
Zarka Cvetanovic, Edward G. Freedman, Charles Nofsinger
J. Supercomput.1
1990 Perfect Benchmarks decomposition and performance on VAX multiprocessors
abstract
The authors present the methods for decomposition and the performance of the Perfect Benchmark suite on two VAX multiprocessors. Results indicate that it was possible to obtain significant performance gains on both VAX multiprocessors, relative to a uniprocessor case. A methodology that can be applied to decompose other existing scientific and engineering applications is proposed. Guidelines for autodecomposing compiler improvements are provided. The authors discuss and quantify the effect of different cache designs on the multiprocessor performance.>
Zarka Cvetanovic, Edward G. Freedman, Charles Nofsinger
SC1
1987 The Effects of Problem Partitioning, Allocation, and Granularity on the Performance of Multiple-Processor Systems
abstract
In this paper we analyze the effects of the problem decomposition, the allocation of subproblems to processors, and the grain size of subproblems on the performance of a multiple- processor shared-memory architecture. Our results indicate that for algorithms where both the computation and the communication overhead can be fully decomposed among N processors, the speedup is a nondecreasing function of the level of granularity for arbitrary interconnection structure and allocation of subproblems to processors. For these algorithms, the speedup is an increasing function of the level of granularity provided that the interconnection bandwidth is greater than unity. If the bandwidth is equal to unity, then the speedup converges to the value equal to the ratio of processing time to communication time. For algorithms where the computation is decomposable but the communication overhead cannot be decomposed, the speedup is a nondecreasing function of the level of granularity for the best case bandwidth only. If the bandwidth is less than N, the speedup reaches its maximum and then decreases approaching zero as the level of granularity grows. For algorithms where the computation consists of parallel and serial sections of code and the communication overhead is fully decomposable, the speedup converges to a value inversely proportional to the fraction of time spent in the serial code even for the best case interconnection bandwidth.
Zarka Cvetanovic
IEEE Trans. Computers1