Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zoran Radovic

dblp:r/ZoranRadovic · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
0since 2021 · last 2006
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 53% Parallel and multicore computing · 32% Processor architecture and microarchitecture · 10%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
synchronization
0.122003
Hierarchical Backoff Locks for Nonuniform Communication Architectures · HPCA 2003
Efficient synchronization for nonuniform communication architectures · SC 2002
Memory systems
cache coherence
0.022003
Efficient synchronization for nonuniform communication architectures · SC 2002
Hierarchical Backoff Locks for Nonuniform Communication Architectures · HPCA 2003
Processor architecture and microarchitecture
multicore design
0.012002
Efficient synchronization for nonuniform communication architectures · SC 2002
Memory systems › cache coherence
cache coherence protocol
0.012001
Removing the overhead from software-based shared memory · SC 2001
Memory systems › shared memory
distributed shared memory
0.012001
Removing the overhead from software-based shared memory · SC 2001
Memory systems
shared memory
0.012001
Removing the overhead from software-based shared memory · SC 2001
Memory systems › memory architecture › parallel memory system
software shared memory
0.012001
Removing the overhead from software-based shared memory · SC 2001
Memory systems › non-uniform memory access
CC-NUMA
0.012003
Hierarchical Backoff Locks for Nonuniform Communication Architectures · HPCA 2003
Memory systems › memory hierarchy › cache hierarchy
non-uniform cache access
0.012002
Efficient synchronization for nonuniform communication architectures · SC 2002
Interconnection networks and networks-on-chip
cluster interconnect
0.012001
Removing the overhead from software-based shared memory · SC 2001
Interconnection networks and networks-on-chip › cluster interconnect
infiniband
0.012001
Removing the overhead from software-based shared memory · SC 2001

Methods — techniques the papers use, named apart from their topics

starvation avoidance · 0.0hierarchical backoff · 0.0synchronization primitives · 0.0RH lock · 0.0interrupt-free protocol processing · 0.0OS bypass · 0.0
YearPublicationVenuePosition
2006 TMA: a trap-based memory architecture
abstract
The advances in semiconductor technology have set the shared-memory server trend towards processors with multiple cores per die and multiple threads per core. We believe that this technology shift forces a reevaluation of how to interconnect multiple such chips to form larger systems.This paper argues that by adding support for coherence traps in future chip multiprocessors, large-scale server systems can be formed at a much lower cost. This is due to shorter design time, verification and time to market when compared to its traditional all-hardware counter part. In the proposed trap-based memory architecture (TMA), software trap handlers are responsible for obtaining read/write permission, whereas the coherence trap hardware is responsible for the actual permission check.In this paper we evaluate a TMA implementation (called TMA Lite) with a minimal amount of hardware extensions, all contained within the processor. The proposed mechanisms for coherence trap processing should not affect the critical path and have a negligible cost in terms of area and power for most processor designs.Our evaluation is based on detailed full system simulation using out-of-order processors with one or two dual-threaded cores per die as processing nodes. The results show that a TMA based distributed shared memory system can perform on par with a highly optimized hardware based design.
Håkan Zeffer, Zoran Radovic, Martin Karlsson, Erik Hagersten
ICS2
2006 Exploiting locality: a flexible DSM approach
abstract
No single coherence strategy suits all applications well. Many promising adaptive protocols and coherence predictors, capable of dynamically modifying the coherence strategy, have been suggested over the years. While most dynamic detection schemes rely on plentiful of dedicated hardware, the customization technique suggested in this paper requires no extra hardware support for its per-application coherence strategy. Instead, each application is profiled using a low-overhead profiling tool. The appropriate coherence flag setting, suggested by the profiling, is specified when the application is launched. We have compared the performance of a hardware DSM (Sun WildFire) to a software DSM (distributed shared memory) built with identical interconnect hardware and coherence strategy. With no support for flexibility, the software DSM runs on average 45 percent slower than the hardware DSM on the 12 studied applications, while the flexibility can get the software DSM within 11 percent. Our all-software system outperforms the hardware DSM on four applications
Håkan Zeffer, Zoran Radovic, Erik Hagersten
IPDPS2
2004 Exploiting Spatial Store Locality Through Permission Caching in Software DSMs
Håkan Zeffer, Zoran Radovic, Oskar Grenholm, Erik Hagersten
Euro-Par2
2003 THROOM - Supporting POSIX Multithreaded Binaries on a Cluster
Henrik Löf, Zoran Radovic, Erik Hagersten
Euro-Par2
2003 Hierarchical Backoff Locks for Nonuniform Communication Architectures
abstract
This paper identifies node affinity as an important property for scalable general-purpose locks. Nonuniform communication architectures (NUCA), for example CC-NUMA built from a few large nodes or from chip multiprocessors (CMP), have a lower penalty for reading data from a neighbor's cache than from a remote cache. Lock implementations that encourages handing over locks to neighbors will improve the lock handover time, as well as the access to the critical data guarded by the lock, but will also be vulnerable to starvation. We propose a set of simple software-based hierarchical backoff locks (HBO) that create node affinity in NUCA. A solution for lowering the risk of starvation is also suggested. The HBO locks are compared with other software-based lock implementations using simple benchmarks, and are shown to be very competitive for uncontested locks while being more than twice as fast for contended locks. An application study also demonstrates superior performance for applications with high lock contention and competitive performance for other programs.
Zoran Radovic, Erik Hagersten
HPCA1
2002 Efficient synchronization for nonuniform communication architectures
abstract
Scalable parallel computers are often nonuniform communication architectures (NUCAs), where the access time to other processor’s caches vary with their physical location. Still, few attempts of exploring cache-to-cache communication locality have been made. This paper introduces a new kind of synchronization primitives (lock-unlock) that favor neighboring processors when a lock is released. This improves the lock handover time as well as access time to the shared data of the critical region. A critical section guarded by our new RH lock takes less than half the time to execute compared with the same critical section guarded by any other lock on our NUCA hardware. The execution time for Raytrace with 28 processors was improved 2.23 - 4.68 times, while global traffic was dramatically decreased compared with all the other locks. The average execution time was improved 7 - 24% while the global traffic was decreased 8 - 28% for an average over the seven applications studied.
Zoran Radovic, Erik Hagersten
SC1
2001 Removing the overhead from software-based shared memory
abstract
The implementation presented in this paper---DSZOOM-WF---is a sequentially consistent, fine-grained distributed software-based shared memory. It demonstrates a protocol-handling overhead below a microsecond for all the actions involved in a remote load operation, to be compared to the fastest implementation to date of around ten microseconds.The all-software protocol is implemented assuming some basic low-level primitives in the cluster interconnect and an operating system bypass functionality, similar to the emerging InfiniBand standard. All interrupt- and/or poll-based asynchronous protocol processing is completely removed by running the entire coherence protocol in the requesting processor. This not only removes the asynchronous overhead, but also makes use of a processor that otherwise would stall. The technique is applicable to both page-based and fine-grain software-based shared memory.DSZOOM-WF consistently demonstrates performance comparable to hardware-based distributed shared memory implementations.
Zoran Radovic, Erik Hagersten
SC1