EDBT 2026 Demo / reviewers in the wild / expert
Vladimir Cakarevic
dblp:91/7638
· DBLP profile ↗
6ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Parallel and multicore computing · 46% Processor architecture and microarchitecture · 26% Performance modeling and evaluation · 22% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
multithreading |
0.6 | 5 | 2016 | Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors · IEEE Trans. Parallel Distributed Syst. 2013 Optimal task assignment in multithreaded processors: a statistical approach · ASPLOS 2012 Thread to strand binding of parallel network applications in massive multi-threaded systems · PPoPP 2010 |
Parallel and multicore computing › parallel scheduling
thread scheduling |
0.5 | 3 | 2016 | Thread Assignment in Multicore/Multithreaded Processors: A Statistical Approach · IEEE Trans. Computers 2016 Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors · IEEE Trans. Parallel Distributed Syst. 2013 Thread to strand binding of parallel network applications in massive multi-threaded systems · PPoPP 2010 |
Parallel and multicore computing › task allocation
thread placement |
0.4 | 2 | 2016 | Thread Assignment in Multicore/Multithreaded Processors: A Statistical Approach · IEEE Trans. Computers 2016 Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors · IEEE Trans. Parallel Distributed Syst. 2013 |
Performance modeling and evaluation › statistical analysis
extreme value theory |
0.2 | 1 | 2016 | Thread Assignment in Multicore/Multithreaded Processors: A Statistical Approach · IEEE Trans. Computers 2016 |
Performance modeling and evaluation
performance prediction |
0.2 | 1 | 2016 | Thread Assignment in Multicore/Multithreaded Processors: A Statistical Approach · IEEE Trans. Computers 2016 |
Electronic design automation › high-level synthesis
scheduling |
0.2 | 1 | 2013 | Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors · IEEE Trans. Parallel Distributed Syst. 2013 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2013 | Characterizing the resource-sharing levels in the UltraSPARC T2 processor · MICRO 2009 Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded Processors · IEEE Trans. Parallel Distributed Syst. 2013 |
Performance modeling and evaluation › statistical analysis
statistical modeling |
0.1 | 1 | 2012 | Optimal task assignment in multithreaded processors: a statistical approach · ASPLOS 2012 |
Parallel and multicore computing
task allocation |
0.1 | 1 | 2012 | Optimal task assignment in multithreaded processors: a statistical approach · ASPLOS 2012 |
Parallel and multicore computing
task scheduling |
0.1 | 1 | 2012 | Optimal task assignment in multithreaded processors: a statistical approach · ASPLOS 2012 |
Parallel and multicore computing › parallel scheduling
resource-aware scheduling |
0.1 | 1 | 2009 | Characterizing the resource-sharing levels in the UltraSPARC T2 processor · MICRO 2009 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.0 | 1 | 2010 | Thread to strand binding of parallel network applications in massive multi-threaded systems · PPoPP 2010 |
Operating systems › resource management
load balancing |
0.0 | 1 | 2009 | Characterizing the resource-sharing levels in the UltraSPARC T2 processor · MICRO 2009 |
Operating systems › resource management › process management › CPU scheduling
thread scheduling |
0.0 | 1 | 2009 | Characterizing the resource-sharing levels in the UltraSPARC T2 processor · MICRO 2009 |
Methods — techniques the papers use, named apart from their topics
statistical sampling · 0.2sample pruning · 0.2extreme value theory · 0.2performance characterization · 0.2case study · 0.2load balancing · 0.2blackbox scheduler · 0.2statistical approach · 0.1scheduling · 0.1performance measurement · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Thread Assignment in Multicore/Multithreaded Processors: A Statistical ApproachabstractThe introduction of multicore/multithreaded processors, comprised of a large number of hardware contexts (virtual CPUs) that share resources at multiple levels, has made process scheduling, in particular assignment of running threads to available hardware contexts, an important aspect of system performance. Nevertheless, thread assignment of applications running on state-of-the art processors is an NP-complete problem. Over the years, numerous studies have proposed heuristic-based algorithms for thread assignment. Since the thread assignment problem is intractable, it is in general impossible to know the performance of the optimal assignment, so the room for improvement of a given algorithm is also unknown. It is therefore hard to decide whether to invest more effort and time to improve an algorithm that may already be close to optimal. In this paper, we present a statistical approach to the thread assignment problem. First, we present a method that predicts the performance of the optimal thread assignment, based on the observed performance of each thread assignment in a random sample. The method is based on Extreme Value Theory (EVT), a branch of statistics that analyses extreme deviations from the population mean. We also propose sample pruning, a method that significantly reduces the time required to apply the statistical method by reducing the number of candidate solutions that need to be measured. Finally, we show that, if no suitable heuristic-based algorithm is available, a sample of several thousand random thread assignments is enough to obtain, with high confidence, an assignment with performance close to optimal. The presented approach is architecture and application independent, and it can be used to address the thread assignment problem in various domains. It is especially well suited for systems in which the workload seldom changes. An example is network systems, which typically provide a constant set of services that are known in advance, with network applications performing a similar processing algorithm for each packet in the system. In this paper, we validate our methods with an industrial case study for a set of multithreaded network applications on an UltraSPARC T2 processor. This article is an extension of our previous work [44], which was published in Proceedings of 17th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS-2012). Petar Radojkovic, Paul M. Carpenter, Miquel Moretó, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
IEEE Trans. Computers | 4 |
| 2013 | Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded ProcessorsabstractThe introduction of multithreaded processors comprised of a large number of cores with many shared resources makes thread scheduling, and in particular optimal assignment of running threads to processor hardware contexts to become one of the most promising ways to improve the system performance. However, finding optimal thread assignments for workloads running in state-of-the-art multicore/multithreaded processors is an NP-complete problem. In this paper, we propose BlackBox scheduler, a systematic method for thread assignment of multithreaded network applications running on multicore/multithreaded processors. The method requires minimum information about the target processor architecture and no data about the hardware requirements of the applications under study. The proposed method is evaluated with an industrial case study for a set of multithreaded network applications running on the UltraSPARC T2 processor. In most of the experiments, the proposed thread assignment method detected the best actual thread assignment in the evaluation sample. The method improved the system performance from 5 to 48 percent with respect to load balancing algorithms used in state-of-the-art OSs, and up to 60 percent with respect to a naive thread assignment. Petar Radojkovic, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | Optimal task assignment in multithreaded processors: a statistical approachabstractThe introduction of massively multithreaded (MMT) processors, comprised of a large number of cores with many shared resources, has made task scheduling, in particular task to hardware thread assignment, one of the most promising ways to improve system performance. However, finding an optimal task assignment for a workload running on MMT processors is an NP-complete problem. Due to the fact that the performance of the best possible task assignment is unknown, the room for improvement of current task-assignment algorithms cannot be determined. This is a major problem for the industry because it could lead to: (1)~A waste of resources if excessive effort is devoted to improving a task assignment algorithm that already provides a performance that is close to the optimal one, or (2)~significant performance loss if insufficient effort is devoted to improving poorly-performing task assignment algorithms. Petar Radojkovic, Vladimir Cakarevic, Miquel Moretó, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
ASPLOS | 2 |
| 2010 | Thread to strand binding of parallel network applications in massive multi-threaded systemsabstractIn processors with several levels of hardware resource sharing,like CMPs in which each core is an SMT, the scheduling process becomes more complex than in processors with a single level of resource sharing, such as pure-SMT or pure-CMP processors. Once the operating system selects the set of applications to simultaneously schedule on the processor (workload), each application/thread must be assigned to one of the hardware contexts(strands). We call this last scheduling step the Thread to Strand Binding or TSB. In this paper, we show that the TSB impact on the performance of processors with several levels of shared resources is high. We measure a variation of up to 59% between different TSBs of real multithreaded network applications running on the UltraSPARC T2 processor which has three levels of resource sharing. In our view, this problem is going to be more acute in future multithreaded architectures comprising more cores, more contexts per core, and more levels of resource sharing. Petar Radojkovic, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
PPoPP | 2 |
| 2009 | Characterizing the resource-sharing levels in the UltraSPARC T2 processorabstractThread level parallelism (TLP) has become a popular trend to improve processor performance, overcoming the limitations of extracting instruction level parallelism. Each TLP paradigm, such as Simultaneous Multithreading or Chip-Multiprocessors, provides di erent bene ts, which has motivated processor vendors to combine several TLP paradigms in each chip design. Even if most of these combined-TLP designs are homogeneous, they present di erent levels of hardware resource sharing, which introduces complexities on the operating system scheduling and load balancing.\nCommonly, processor designs provide two levels of resource sharing: Inter-core in which only the highest levels of the cache hierarchy are shared, and Intracore in which\nmost of the hardware resources of the core are shared . Recently, Sun Microsystems has released the UltraSPARC T2, a processor with three levels of hardware resource sharing:\nInterCore, IntraCore, and IntraPipe. In this work, we provide the rst characterization of a three-level resource sharing processor, the UltraSPARC T2, and we show how\nmulti-level resource sharing a ects the operating system design. We further identify the most critical hardware resources in the T2 and the characteristics of applications that are not sensitive to resource sharing. Finally, we present a case study in which we run a real multithreaded network application, showing that a resource sharing aware scheduler can improve the system throughput up to 55%. Vladimir Cakarevic, Petar Radojkovic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
MICRO | 1 |
| 2008 | Measuring Operating System Overhead on CMT ProcessorsabstractNumerous studies have shown that Operating System (OS)noise is one of the reasons for significant performance degradation in clustered architectures. Although many studies examine the OS noise for High Performance Computing (HPC),especially in multi-processor/core systems, most of them focus on 2- or 4-core systems.In this paper, we analyze the major sources of OS noise on a massive multithreading processor, the Sun UltraSPARC T1,running Linux and Solaris. Since a real system is too complex to analyze, we compare those results with a low-overhead runtime environment: the Netra Data Plane Software Suite (Netra DPS).Our results show that the overhead introduced by the OStimer interrupt in Linux and Solaris depends on the particular core and hardware context in which the application is running. This overhead is up to 30% when the application is executed on the same hardware context of the timer interrupt handler and up to 10% when the application and the timer interrupt handler run on different contexts but on the same core.We detect no overhead when the benchmark and the timer interrupt handler run on different cores of the processor. Petar Radojkovic, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Roberto Gioiosa, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
SBAC-PAD | 2 |