Vladimir Cakarevic

dblp:91/7638 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 46% Processor architecture and microarchitecture · 26% Performance modeling and evaluation · 22%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
multithreading
0.652016
Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors · IEEE Trans. Parallel Distributed Syst. 2013
Optimal task assignment in multithreaded processors: a statistical approach · ASPLOS 2012
Thread to strand binding of parallel network applications in massive multi-threaded systems · PPoPP 2010
Parallel and multicore computing › parallel scheduling
thread scheduling
0.532016
Thread Assignment in Multicore/Multithreaded Processors: A Statistical Approach · IEEE Trans. Computers 2016
Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors · IEEE Trans. Parallel Distributed Syst. 2013
Thread to strand binding of parallel network applications in massive multi-threaded systems · PPoPP 2010
Parallel and multicore computing › task allocation
thread placement
0.422016
Thread Assignment in Multicore/Multithreaded Processors: A Statistical Approach · IEEE Trans. Computers 2016
Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors · IEEE Trans. Parallel Distributed Syst. 2013
Performance modeling and evaluation › statistical analysis
extreme value theory
0.212016
Thread Assignment in Multicore/Multithreaded Processors: A Statistical Approach · IEEE Trans. Computers 2016
Performance modeling and evaluation
performance prediction
0.212016
Thread Assignment in Multicore/Multithreaded Processors: A Statistical Approach · IEEE Trans. Computers 2016
Electronic design automation › high-level synthesis
scheduling
0.212013
Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors · IEEE Trans. Parallel Distributed Syst. 2013
Processor architecture and microarchitecture
chip multiprocessor
0.122013
Characterizing the resource-sharing levels in the UltraSPARC T2 processor · MICRO 2009
Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors · IEEE Trans. Parallel Distributed Syst. 2013
Performance modeling and evaluation › statistical analysis
statistical modeling
0.112012
Optimal task assignment in multithreaded processors: a statistical approach · ASPLOS 2012
Parallel and multicore computing
task allocation
0.112012
Optimal task assignment in multithreaded processors: a statistical approach · ASPLOS 2012
Parallel and multicore computing
task scheduling
0.112012
Optimal task assignment in multithreaded processors: a statistical approach · ASPLOS 2012
Parallel and multicore computing › parallel scheduling
resource-aware scheduling
0.112009
Characterizing the resource-sharing levels in the UltraSPARC T2 processor · MICRO 2009
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.012010
Thread to strand binding of parallel network applications in massive multi-threaded systems · PPoPP 2010
Operating systems › resource management
load balancing
0.012009
Characterizing the resource-sharing levels in the UltraSPARC T2 processor · MICRO 2009
Operating systems › resource management › process management › CPU scheduling
thread scheduling
0.012009
Characterizing the resource-sharing levels in the UltraSPARC T2 processor · MICRO 2009

Methods — techniques the papers use, named apart from their topics

statistical sampling · 0.2sample pruning · 0.2extreme value theory · 0.2performance characterization · 0.2case study · 0.2load balancing · 0.2blackbox scheduler · 0.2statistical approach · 0.1scheduling · 0.1performance measurement · 0.1
YearPublicationVenuePosition
2016 Thread Assignment in Multicore/Multithreaded Processors: A Statistical Approach
abstract
The introduction of multicore/multithreaded processors, comprised of a large number of hardware contexts (virtual CPUs) that share resources at multiple levels, has made process scheduling, in particular assignment of running threads to available hardware contexts, an important aspect of system performance. Nevertheless, thread assignment of applications running on state-of-the art processors is an NP-complete problem. Over the years, numerous studies have proposed heuristic-based algorithms for thread assignment. Since the thread assignment problem is intractable, it is in general impossible to know the performance of the optimal assignment, so the room for improvement of a given algorithm is also unknown. It is therefore hard to decide whether to invest more effort and time to improve an algorithm that may already be close to optimal. In this paper, we present a statistical approach to the thread assignment problem. First, we present a method that predicts the performance of the optimal thread assignment, based on the observed performance of each thread assignment in a random sample. The method is based on Extreme Value Theory (EVT), a branch of statistics that analyses extreme deviations from the population mean. We also propose sample pruning, a method that significantly reduces the time required to apply the statistical method by reducing the number of candidate solutions that need to be measured. Finally, we show that, if no suitable heuristic-based algorithm is available, a sample of several thousand random thread assignments is enough to obtain, with high confidence, an assignment with performance close to optimal. The presented approach is architecture and application independent, and it can be used to address the thread assignment problem in various domains. It is especially well suited for systems in which the workload seldom changes. An example is network systems, which typically provide a constant set of services that are known in advance, with network applications performing a similar processing algorithm for each packet in the system. In this paper, we validate our methods with an industrial case study for a set of multithreaded network applications on an UltraSPARC T2 processor. This article is an extension of our previous work [44], which was published in Proceedings of 17th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS-2012).
Petar Radojkovic, Paul M. Carpenter, Miquel Moretó, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero
IEEE Trans. Computers4
2013 Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors
abstract
The introduction of multithreaded processors comprised of a large number of cores with many shared resources makes thread scheduling, and in particular optimal assignment of running threads to processor hardware contexts to become one of the most promising ways to improve the system performance. However, finding optimal thread assignments for workloads running in state-of-the-art multicore/multithreaded processors is an NP-complete problem. In this paper, we propose BlackBox scheduler, a systematic method for thread assignment of multithreaded network applications running on multicore/multithreaded processors. The method requires minimum information about the target processor architecture and no data about the hardware requirements of the applications under study. The proposed method is evaluated with an industrial case study for a set of multithreaded network applications running on the UltraSPARC T2 processor. In most of the experiments, the proposed thread assignment method detected the best actual thread assignment in the evaluation sample. The method improved the system performance from 5 to 48 percent with respect to load balancing algorithms used in state-of-the-art OSs, and up to 60 percent with respect to a naive thread assignment.
Petar Radojkovic, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero
IEEE Trans. Parallel Distributed Syst.2
2012 Optimal task assignment in multithreaded processors: a statistical approach
abstract
The introduction of massively multithreaded (MMT) processors, comprised of a large number of cores with many shared resources, has made task scheduling, in particular task to hardware thread assignment, one of the most promising ways to improve system performance. However, finding an optimal task assignment for a workload running on MMT processors is an NP-complete problem. Due to the fact that the performance of the best possible task assignment is unknown, the room for improvement of current task-assignment algorithms cannot be determined. This is a major problem for the industry because it could lead to: (1)~A waste of resources if excessive effort is devoted to improving a task assignment algorithm that already provides a performance that is close to the optimal one, or (2)~significant performance loss if insufficient effort is devoted to improving poorly-performing task assignment algorithms.
Petar Radojkovic, Vladimir Cakarevic, Miquel Moretó, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero
ASPLOS2
2010 Thread to strand binding of parallel network applications in massive multi-threaded systems
abstract
In processors with several levels of hardware resource sharing,like CMPs in which each core is an SMT, the scheduling process becomes more complex than in processors with a single level of resource sharing, such as pure-SMT or pure-CMP processors. Once the operating system selects the set of applications to simultaneously schedule on the processor (workload), each application/thread must be assigned to one of the hardware contexts(strands). We call this last scheduling step the Thread to Strand Binding or TSB. In this paper, we show that the TSB impact on the performance of processors with several levels of shared resources is high. We measure a variation of up to 59% between different TSBs of real multithreaded network applications running on the UltraSPARC T2 processor which has three levels of resource sharing. In our view, this problem is going to be more acute in future multithreaded architectures comprising more cores, more contexts per core, and more levels of resource sharing.
Petar Radojkovic, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero
PPoPP2
2009 Characterizing the resource-sharing levels in the UltraSPARC T2 processor
abstract
Thread level parallelism (TLP) has become a popular trend to improve processor performance, overcoming the limitations of extracting instruction level parallelism. Each TLP paradigm, such as Simultaneous Multithreading or Chip-Multiprocessors, provides di erent bene ts, which has motivated processor vendors to combine several TLP paradigms in each chip design. Even if most of these combined-TLP designs are homogeneous, they present di erent levels of hardware resource sharing, which introduces complexities on the operating system scheduling and load balancing.\nCommonly, processor designs provide two levels of resource sharing: Inter-core in which only the highest levels of the cache hierarchy are shared, and Intracore in which\nmost of the hardware resources of the core are shared . Recently, Sun Microsystems has released the UltraSPARC T2, a processor with three levels of hardware resource sharing:\nInterCore, IntraCore, and IntraPipe. In this work, we provide the rst characterization of a three-level resource sharing processor, the UltraSPARC T2, and we show how\nmulti-level resource sharing a ects the operating system design. We further identify the most critical hardware resources in the T2 and the characteristics of applications that are not sensitive to resource sharing. Finally, we present a case study in which we run a real multithreaded network application, showing that a resource sharing aware scheduler can improve the system throughput up to 55%.
Vladimir Cakarevic, Petar Radojkovic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero
MICRO1
2008 Measuring Operating System Overhead on CMT Processors
abstract
Numerous studies have shown that Operating System (OS)noise is one of the reasons for significant performance degradation in clustered architectures. Although many studies examine the OS noise for High Performance Computing (HPC),especially in multi-processor/core systems, most of them focus on 2- or 4-core systems.In this paper, we analyze the major sources of OS noise on a massive multithreading processor, the Sun UltraSPARC T1,running Linux and Solaris. Since a real system is too complex to analyze, we compare those results with a low-overhead runtime environment: the Netra Data Plane Software Suite (Netra DPS).Our results show that the overhead introduced by the OStimer interrupt in Linux and Solaris depends on the particular core and hardware context in which the application is running. This overhead is up to 30% when the application is executed on the same hardware context of the timer interrupt handler and up to 10% when the application and the timer interrupt handler run on different contexts but on the same core.We detect no overhead when the benchmark and the timer interrupt handler run on different cores of the processor.
Petar Radojkovic, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Roberto Gioiosa, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero
SBAC-PAD2