EDBT 2026 Demo / reviewers in the wild / expert
Alejandro Duran
dblp:44/2291 · also Alex Duran
· DBLP profile ↗
13ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 100% | |
| Software engineering, system software, and programming languages
2 papers |
Program analysis · 78% Runtime systems and virtual machines · 22% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel programming models |
0.5 | 3 | 2018 | The Ongoing Evolution of OpenMP · Proc. IEEE 2018 The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009 An adaptive cut-off for task parallelism · SC 2008 |
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP |
0.3 | 1 | 2018 | The Ongoing Evolution of OpenMP · Proc. IEEE 2018 |
Parallel and multicore computing › parallel programming models
shared-memory parallelization |
0.3 | 1 | 2018 | The Ongoing Evolution of OpenMP · Proc. IEEE 2018 |
Parallel and multicore computing › parallel programming models
task parallelism |
0.2 | 2 | 2009 | The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009 An adaptive cut-off for task parallelism · SC 2008 |
Parallel and multicore computing › parallel programming models › task parallelism
OpenMP tasking |
0.1 | 2 | 2009 | The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009 An adaptive cut-off for task parallelism · SC 2008 |
Program analysis
performance analysis tools |
0.1 | 1 | 2018 | The Ongoing Evolution of OpenMP · Proc. IEEE 2018 |
Parallel and multicore computing › parallel programming runtimes
runtime systems and scheduling |
0.1 | 1 | 2008 | An adaptive cut-off for task parallelism · SC 2008 |
Runtime systems and virtual machines
parallel runtime systems |
0.0 | 1 | 2009 | The Design of OpenMP Tasks · IEEE Trans. Parallel Distributed Syst. 2009 |
Methods — techniques the papers use, named apart from their topics
task-based parallelism · 0.2runtime information collection · 0.1adaptive cut-off · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Multi-level spatial and temporal tiling for efficient HPC stencil computation on many-core processors with large shared caches
Charles Yount, Alejandro Duran, Josh Tobin |
Future Gener. Comput. Syst. | 2 |
| 2018 | The Ongoing Evolution of OpenMPabstractThis paper presents an overview of the past, present and future of the OpenMP application programming interface (API). While the API originally specified a small set of directives that guided shared memory fork-join parallelization of loops and program sections, OpenMP now provides a richer set of directives that capture a wide range of parallelization strategies that are not strictly limited to shared memory. As we look toward the future of OpenMP, we immediately see further evolution of the support for that range of parallelization strategies and the addition of direct support for debugging and performance analysis tools. Looking beyond the next major release of the specification of the OpenMP API, we expect the specification eventually to include support for more parallelization strategies and to embrace closer integration into its Fortran, C and, in particular, C++ base languages, which will likely require the API to adopt additional programming abstractions. Bronis R. de Supinski, Thomas Scogland, Alejandro Duran, Michael Klemm, Sergi Mateo, Stephen Olivier, Christian Terboven, Timothy G. Mattson |
Proc. IEEE | 3 |
| 2015 | Optimizing Overlapped Memory Accesses in User-directed VectorizationabstractCurrent processors incorporate wide and powerful vector units whose optimal exploitation is crucial to reach peak performance. However, present autovectorizing compilers fall short of that goal. Exploiting some vector instructions requires aggressive approaches that are not affordable in production compilers. Thus, advanced programmers pursuing the best performance from their applications are compelled to manually vectorize them using low-level SIMD intrinsics. Diego Caballero, Sara Royuela, Roger Ferrer, Alejandro Duran, Xavier Martorell |
ICS | 4 |
| 2012 | Productive Programming of GPU Clusters with OmpSsabstractClusters of GPUs are emerging as a new computational scenario. Programming them requires the use of hybrid models that increase the complexity of the applications, reducing the productivity of programmers. We present the implementation of OmpSs for clusters of GPUs, which supports asynchrony and heterogeneity for task parallelism. It is based on annotating a serial application with directives that are translated by the compiler. With it, the same program that runs sequentially in a node with a single GPU can run in parallel in multiple GPUs either local (single node) or remote (cluster of GPUs). Besides performing a task-based parallelization, the runtime system moves the data as needed between the different nodes and GPUs minimizing the impact of communication by using affinity scheduling, caching, and by overlapping communication with the computational task. We show several applications programmed with OmpSs and their performance with multiple GPUs in a local node and in remote nodes. The results show good tradeoff between performance and effort from the programmer. Javier Bueno, Judit Planas, Alejandro Duran, Rosa M. Badia, Xavier Martorell, Eduard Ayguadé, Jesús Labarta |
IPDPS | 3 |
| 2011 | Productive Cluster Programming with OmpSs
Javier Bueno, Luis Martinell, Alejandro Duran, Montse Farreras, Xavier Martorell, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta |
Euro-Par (1) | 3 |
| 2011 | Poster: programming clusters of GPUs with OMPSsabstractOmpSs is a programming model that provides an environment to develop parallel applications for cluster environments with heterogeneous architectures. Based on OpenMP and StarSs, it offers a set of compiler directives that can be used to annotate a sequential code. Additional features have been added to support the use of accelerators like GPUs. This schema offers a high productivity environment due to its simplicity compared to other models like MPI. Our current implementation has shown a good performance when running different benchmarks. Javier Bueno, Alejandro Duran, Xavier Martorell, Eduard Ayguadé, Rosa M. Badia, Jesús Labarta |
ICS | 2 |
| 2011 | Trace-driven simulation of multithreaded applicationsabstractOver the past few years, computer architecture research has moved towards execution-driven simulation, due to the inability of traces to capture timing-dependent thread execution interleaving. However, trace-driven simulation has many advantages over execution-driven that are being missed in multithreaded application simulations. We present a methodology to properly simulate multithreaded applications using trace-driven environments. We distinguish the intrinsic application behavior from the computation for managing parallelism. Application traces capture the intrinsic behavior in the sections of code that are independent from the dynamic multithreaded nature, and the points where parallelism-management computation occurs. The simulation framework is composed of a trace-driven simulation engine and a dynamic-behavior component that implements the parallelism-management operations for the application. Then, at simulation time, these operations are reproduced by invoking their implementation in the dynamic-behavior component. The decisions made by these operations are based on the simulated architecture, allowing to dynamically reschedule sections of code taken from the trace to the target simulated components. As the captured sections of code are independent from the parallel state of the application, they can be simulated on the trace-driven engine, while the parallelism-management operations, that require to be re-executed, are carried out by the execution-driven component, thus achieving the best of both trace- and execution-driven worlds. This simulation methodology creates several new research opportunities, including research on scheduling and other parallelism-management techniques for future architectures, and hardware support for programming models. Alejandro Rico, Alejandro Duran, Felipe Cabarcas, Yoav Etsion, Alex Ramírez, Mateo Valero |
ISPASS | 2 |
| 2009 | Barcelona OpenMP Tasks Suite: A Set of Benchmarks Targeting the Exploitation of Task Parallelism in OpenMPabstractTraditional parallel applications have exploited regular parallelism, based on parallel loops. Only a few applications exploit sections parallelism. With the release of the new OpenMP specification (3.0), this programming model supports tasking. Parallel tasks allow the exploitation of irregular parallelism, but there is a lack of benchmarks exploiting tasks in OpenMP. With the current (and projected) multicore architectures that offer many more alternatives to execute parallel applications than traditional SMP machines, this kind of parallelism is increasingly important. And so, the need to have some set of benchmarks to evaluate it. In this paper, we motivate the need of having such a benchmarks suite, for irregular and/or recursive task parallelism. We present our proposal, the Barcelona OpenMP Tasks Suite (BOTS), with a set of applications exploiting regular and irregular parallelism, based on tasks. We present an overall evaluation of the BOTS benchmarks in an Altix system and we discuss some of the different experiments that can be done with the different compilation and runtime alternatives of the benchmarks. Alejandro Duran, Xavier Teruel, Roger Ferrer, Xavier Martorell, Eduard Ayguadé |
ICPP | 1 |
| 2009 | The Design of OpenMP TasksabstractOpenMP has been very successful in exploiting structured parallelism in applications. With increasing application complexity, there is a growing need for addressing irregular parallelism in the presence of complicated control structures. This is evident in various efforts by the industry and research communities to provide a solution to this challenging problem. One of the primary goals of OpenMP 3.0 is to define a standard dialect to express and efficiently exploit unstructured parallelism. This paper presents the design of the OpenMP tasking model by members of the OpenMP 3.0 tasking sub-committee which was formed for this purpose. The paper summarizes the efforts of the sub-committee (spanning over two years) in designing, evaluating and seamlessly integrating the tasking model into the OpenMP specification. In this paper, we present the design goals and key features of the tasking model, including a rich set of examples and an in-depth discussion of the rationale behind various design choices. We compare a prototype implementation of the tasking model with existing models, and evaluate it on a wide range of applications. The comparison shows that the OpenMP tasking model provides expressiveness, flexibility, and huge potential for performance and scalability. Eduard Ayguadé, Nawal Copty, Alejandro Duran, Jay P. Hoeflinger, Federico Massaioli, Xavier Teruel, Priya Unnikrishnan, Guansong Zhang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2008 | An adaptive cut-off for task parallelismabstractIn task parallel languages, an important factor for achieving a good performance is the use of a cut-off technique to reduce the number of tasks created. Using a cut-off to avoid an excessive number of tasks helps the runtime system to reduce the total overhead associated with task creation, particularlt if the tasks are fine grain. Unfortunately, the best cut-off technique its usually dependent on the application structure or even the input data of the application. We propose a new cut-off technique that, using information from the application collected at runtime, decides which tasks should be pruned to improve the performance of the application. This technique does not rely on the programmer to determine the cut-off technique that is best suited for the application. We have implemented this cut-off in the context of the new OpenMP tasking model. Our evaluation, with a variety of applications, shows that our adaptive cut-off is able to make good decisions and most of the time matches the optimal cut-off that could be set by hand by a programmer. Alejandro Duran, Julita Corbalán, Eduard Ayguadé |
SC | 1 |
| 2006 | Techniques supporting threadprivate in OpenMPabstractThis paper presents the alternatives available to support threadprivate data in OpenMP and evaluates them. We show how current compilation systems rely on custom techniques for implementing thread-local data. But in fact the ELF binary specification currently supports data sections that become threadprivate by default. ELF naming for such areas is thread-local storage (TLS). Our experiments demonstrate that implementing threadprivate based on the TLS support is very easy, and more efficient. This proposal goes in the same line as the future implementation of OpenMP on the GNU compiler collection. In addition, our experience with the use of threadprivate in OpenMP applications shows that usually it is better to avoid it. This is because threadprivate variables reside in common blocks and they impede the compiler to fully optimize the code. So it is better to keep threadprivate as a temporary technique only to ease porting MPI codes to OpenMP. Xavier Martorell, Marc González 0001, Alejandro Duran, Jairo Balart, Roger Ferrer, Eduard Ayguadé, Jesús Labarta |
IPDPS | 3 |
| 2005 | Automatic thread distribution for nested parallelism in OpenMPabstractOpenMP is becoming the standard programming model for shared-memory parallel architectures. One of its most interesting features in the language is the support for nested parallelism. Previous research and parallelization experiences have shown the benefits of using nested parallelism as an alternative to combining several programming models such as MPI and OpenMP. However, all these works rely on the manual definition of an appropriate distribution of all the available thread across the different levels of parallelism. Some proposals have been made to extend the OpenMP language to allow the programmers to specify the thread distribution.This paper proposes a mechanism to dynamically compute the most appropriate thread distribution strategy. The mechanism is based on gathering information at runtime to derive the structure of the nested parallelism. This information is used to determine how the overall computation is distributed between the parallel branches in the outermost level of parallelism, which is constant in this work. According to this, threads in the innermost level of parallelism are distributed.The proposed mechanism is evaluated in two different environments: a research environment, the Nanos OpenMP research platform, and a commercial environment, the IBM XL runtime library. The performance numbers obtained validate the mechanism in both environments and they show the importance of selecting the proper amount of parallelism in the outer level. Alejandro Duran, Marc González 0001, Julita Corbalán |
ICS | 1 |
| 2004 | Dynamic Load Balancing of MPI+OpenMP ApplicationsabstractThe hybrid programming model MPI+OpenMP are useful to solve the problems of load balancing of parallel applications independently of the architecture. Typical approaches to balance parallel applications using two levels of parallelism or only MPI consist of including complex codes that dynamically detect which data domains are more computational intensive and either manually redistribute the allocated processors or manually redistribute data. This approach has two drawbacks: it is time consuming and it requires an expert in application analysis. In this paper we present an automatic and dynamic approach for load balancing MPI+OpenMP applications. The system calculates the percentage of load imbalance and decides a processor distribution for the MPI processes that eliminates the computational load imbalance. Results show that this method can balance effectively applications without analyzing nor modifying them and that in the cases that the application was well balanced does not incur in a great overhead for the dynamic instrumentation and analysis realized. Julita Corbalán, Alejandro Duran, Jesús Labarta |
ICPP | 2 |