VLDB 2026 Research / reviewers in the wild / expert
Zoltan Majó
dblp:336/4275
· DBLP profile ↗
4ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 67% Memory systems · 33% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › locality optimization
data locality optimization |
0.2 | 1 | 2015 | A library for portable and composable data locality optimizations for NUMA systems · PPoPP 2015 |
Memory systems
non-uniform memory access |
0.2 | 1 | 2015 | A library for portable and composable data locality optimizations for NUMA systems · PPoPP 2015 |
Parallel and multicore computing
parallel programming models |
0.2 | 1 | 2015 | A library for portable and composable data locality optimizations for NUMA systems · PPoPP 2015 |
Operating systems › resource management › process management › CPU scheduling
task scheduling |
0.1 | 1 | 2015 | A library for portable and composable data locality optimizations for NUMA systems · PPoPP 2015 |
Methods — techniques the papers use, named apart from their topics
intel threading building blocks · 0.4composable optimization · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | A library for portable and composable data locality optimizations for NUMA systemsabstractMany recent multiprocessor systems are realized with a non-uniform memory architecture (NUMA) and accesses to remote memory locations take more time than local memory accesses. Optimizing NUMA memory system performance is difficult and costly for three principal reasons: (1) today's programming languages/libraries have no explicit support for NUMA systems, (2) NUMA optimizations are not~portable, and (3) optimizations are not~composable (i.e., they can become ineffective or worsen performance in environments that support composable parallel software). This paper presents TBB-NUMA, a parallel programming library based on Intel Threading Building Blocks (TBB) that supports portable and composable NUMA-aware programming. TBB-NUMA provides a model of task affinity that captures a programmer's insights on mapping tasks to resources. NUMA-awareness affects all layers of the library (i.e., resource management, task scheduling, and high-level parallel algorithm templates) and requires close coupling between all these layers. Optimizations implemented with TBB-NUMA (for a set of standard benchmark programs) result in up to 44% performance improvement over standard TBB, but more important, optimized programs are portable across different NUMA architectures and preserve data locality also when composed with other parallel computations. Zoltan Majó, Thomas R. Gross |
PPoPP | 1 |
| 2012 | Matching memory access patterns and data placement for NUMA systemsabstractMany recent multicore multiprocessors are based on a nonuniform memory architecture (NUMA). A mismatch between the data access patterns of programs and the mapping of data to memory incurs a high overhead, as remote accesses have higher latency and lower throughput than local accesses. This paper reports on a limit study that shows that many scientific loop-parallel programs include multiple, mutually incompatible data access patterns, therefore these programs encounter a high fraction of costly remote memory accesses. Matching the data distribution of a program to the individual data access patterns is possible, however it is difficult to find a data distribution that matches all access patterns. Zoltan Majó, Thomas R. Gross |
CGO | 1 |
| 2011 | Memory management in NUMA multicore systems: trapped between cache contention and interconnect overheadabstractMultiprocessors based on processors with multiple cores usually include a non-uniform memory architecture (NUMA); even current 2-processor systems with 8 cores exhibit non-uniform memory access times. As the cores of a processor share a common cache, the issues of memory management and process mapping must be revisited. We find that optimizing only for data locality can counteract the benefits of cache contention avoidance and vice versa. Therefore, system software must take both data locality and cache contention into account to achieve good performance, and memory management cannot be decoupled from process scheduling. We present a detailed analysis of a commercially available NUMA-multicore architecture, the Intel Nehalem. We describe two scheduling algorithms: maximum-local, which optimizes for maximum data locality, and its extension, N-MASS, which reduces data locality to avoid the performance degradation caused by cache contention. N-MASS is fine-tuned to support memory management on NUMA-multicores and improves performance up to 32%, and 7% on average, over the default setup in current Linux implementations. Zoltan Majó, Thomas R. Gross |
ISMM | 1 |
| 2011 | Memory system performance in a NUMA multicore multiprocessorabstractModern multicore processors with an on-chip memory controller form the base for NUMA (non-uniform memory architecture) multiprocessors. Each processor accesses part of the physical memory directly and has access to the other parts via the memory controller of other processors. These other processors are reached via the cross-processor interconnect. As a consequence a processor's memory controller must satisfy two kinds of requests: those that are generated by the local cores and those that arrive via the interconnect from other processors. On the other hand, a core (respectively the core's cache) can obtain data from multiple sources: data can be supplied by the local memory controller or by a remote memory controller on another processor. In this paper we experimentally analyze the behavior of the memory controllers of a commercial multicore processor, the Intel Xeon 5520 (Nehalem). We develop a simple model to characterize the sharing of local and remote memory bandwidth. The uneven treatment of local and remote accesses has implications for mapping applications onto such a NUMA multicore multiprocessor. Maximizing data locality does not always minimize execution time; it may be more advantageous to allocate data on a remote processor (and then to fetch these data via the cross-processor interconnect) than to store the data of all processes in local memory (and consequently over-loading the on-chip memory controller). Zoltan Majó, Thomas R. Gross |
SYSTOR | 1 |