Manuel Arenaz

dblp:14/3532 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
0since 2021 · last 2014
0000-0002-0195-969XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 7 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
dependence analysis
0.112008
XARK: An extensible framework for automatic recognition of computational kernels · ACM Trans. Program. Lang. Syst. 2008
Parallel and multicore computing
parallelizing compiler
0.012008
XARK: An extensible framework for automatic recognition of computational kernels · ACM Trans. Program. Lang. Syst. 2008

Methods — techniques the papers use, named apart from their topics

symbolic analysis · 0.2gated single assignment · 0.2demand-driven analysis · 0.2
YearPublicationVenuePosition
2014 A parallelizing compiler for multicore systems
abstract
This manuscript summarizes the main ideas introduced in [1]. We propose a compiler that automatically transforms a sequential application into a parallel counterpart for multicore processors. It is based on an intermediate representation, named KIR, which exposes multiple levels of parallelism and hides the complexity of the implementation details thanks to the domain-independent kernels (e.g., assignment, reduction). The effectiveness and performance of our approach, built on top of GCC, has been tested with a large variety of codes.
José M. Andión, Manuel Arenaz, Gabriel Rodríguez 0001, Juan Touriño
SCOPES2
2013 A multi-GPU shallow-water simulation with transport of contaminants
abstract
SUMMARY This work presents cost‐effective multi‐graphics processing unit (GPU) parallel implementations of a finite‐volume numerical scheme for solving pollutant transport problems in bidimensional domains. The fluid is modeled by 2D shallow‐water equations, whereas the transport of pollutant is modeled by a transport equation. The 2D domain is discretized using a first‐order Roe finite‐volume scheme. Specifically, this paper presents multi‐GPU implementations of both a solution that exploits recomputation on the GPU and an optimized solution that is based on a ghost cell decoupling approach. Our multi‐GPU implementations have been optimized using nonblocking communications, overlapping communications and computations and the application of ghost cell expansion to minimize communications. The fastest one reached a speedup of 78 × using four GPUs on an InfiniBand network with respect to a parallel execution on a multicore CPU with six cores and two‐way hyperthreading per core. Such performance, measured using a realistic problem, enabled the calculation of solutions not only in real time but also in orders of magnitude faster than the simulated time.Copyright © 2012 John Wiley & Sons, Ltd.
Moisés Viñas, Jacobo Lobeiras, Basilio B. Fraguela, Manuel Arenaz, Margarita Amor, José A. García, Manuel Jesús Castro Díaz, Ramón Doallo
Concurr. Comput. Pract. Exp.4
2013 A novel compiler support for automatic parallelization on multicore systems
José M. Andión, Manuel Arenaz, Gabriel Rodríguez 0001, Juan Touriño
Parallel Comput.2
2008 Efficiently Building the Gated Single Assignment Form in Codes with Pointers in Modern Optimizing Compilers
Manuel Arenaz, Pedro Amoedo, Juan Touriño
Euro-Par1
2008 XARK: An extensible framework for automatic recognition of computational kernels
abstract
The recognition of program constructs that are frequently used by software developers is a powerful mechanism for optimizing and parallelizing compilers to improve the performance of the object code. The development of techniques for automatic recognition of computational kernels such as inductions, reductions and array recurrences has been an intensive research area in the scope of compiler technology during the 90's. This article presents a new compiler framework that, unlike previous techniques that focus on specific and isolated kernels, recognizes a comprehensive collection of computational kernels that appear frequently in full-scale real applications. The XARK compiler operates on top of the Gated Single Assignment (GSA) form of a high-level intermediate representation (IR) of the source code. Recognition is carried out through a demand-driven analysis of this high-level IR at two different levels. First, the dependences between the statements that compose the strongly connected components (SCCs) of the data-dependence graph of the GSA form are analyzed. As a result of this intra-SCC analysis, the computational kernels corresponding to the execution of the statements of the SCCs are recognized. Second, the dependences between statements of different SCCs are examined in order to recognize more complex kernels that result from combining simpler kernels in the same code. Overall, the XARK compiler builds a hierarchical representation of the source code as kernels and dependence relationships between those kernels. This article describes in detail the collection of computational kernels recognized by the XARK compiler. Besides, the internals of the recognition algorithms are presented. The design of the algorithms enables to extend the recognition capabilities of XARK to cope with new kernels, and provides an advanced symbolic analysis framework to run other compiler techniques on demand. Finally, extensive experiments showing the effectiveness of XARK for a collection of benchmarks from different application domains are presented. In particular, the SparsKit-II library for the manipulation of sparse matrices, the Perfect benchmarks, the SPEC CPU2000 collection and the PLTMG package for solving elliptic partial differential equations are analyzed in detail.
Manuel Arenaz, Juan Touriño, Ramón Doallo
ACM Trans. Program. Lang. Syst.1
2007 Program Behavior Characterization Through Advanced Kernel Recognition
Manuel Arenaz, Juan Touriño, Ramón Doallo
Euro-Par1
2007 Automated and accurate cache behavior analysis for codes with irregular access patterns
abstract
Abstract The memory hierarchy plays an essential role in the performance of current computers, so good analysis tools that help in predicting and understanding its behavior are required. Analytical modeling is the ideal base for such tools if its traditional limitations in accuracy and scope of application can be overcome. While there has been extensive research on the modeling of codes with regular access patterns, less attention has been paid to codes with irregular patterns due to the increased difficulty in analyzing them. Nevertheless, many important applications exhibit this kind of pattern, and their lack of locality make them more cache‐demanding, which makes their study more relevant. The focus of this paper is the automation of the Probabilistic Miss Equations (PME) model, an analytical model of the cache behavior that provides fast and accurate predictions for codes with irregular access patterns. The information requirements of the PME model are defined and its integration in the XARK compiler, a research compiler oriented to automatic kernel recognition in scientific codes, is described. We show how to exploit the powerful information‐gathering capabilities provided by this compiler to allow the automated modeling of loop‐oriented scientific codes. Experimental results that validate the correctness of the automated PME model are also presented. Copyright © 2007 John Wiley & Sons, Ltd.
Diego Andrade, Manuel Arenaz, Basilio B. Fraguela, Juan Touriño, Ramón Doallo
Concurr. Comput. Pract. Exp.2
2007 Special Issue: Current Trends in Compilers for Parallel Computers
abstract
This special issue of Concurrency and Computation: Practice and Experience contains a selection of the papers presented at the 12th International Workshop on Compilers for Parallel Computers (CPC'2006), held in A Coruña, Spain, 9-11 January 2006.The CPC Workshop series is well established as an invitational workshop for leading research groups in the field (mainly from Europe, North America and Asia-Pacific) to provide a forum for exchanging and developing new ideas in compiler design for parallel systems and related topics.The Workshop series began in 1989 in Oxford, U.K., and continued every 18 months in a European city: Paris,
Juan Touriño, Basilio B. Fraguela, Ramón Doallo, Manuel Arenaz
Concurr. Comput. Pract. Exp.4
2004 Compiler Support for Parallel Code Generation through Kernel Recognition
abstract
Summary form only given. The automatic parallelization of loops that contain complex computations is still a challenge for current parallelizing compilers. The main limitations are related to the analysis of expressions that contain subscripted subscripts, and the analysis of conditional statements that introduce complex control flows at run-time. We use the term complex loop to designate loops with such characteristics. We describe the parallelization of sequential complex loop nests using a generic compiler framework (proposed in an earlier paper [Arenaz et al., ICS'2003] ) that accomplishes kernel recognition through the analysis of the gated single assignment program representation. Specifically, we focus on an extension of this framework that enables its use as a powerful tool for gathering source code information that is relevant for the parallelization of each computational kernel. A set of example codes are analyzed in detail to illustrate the potential of our approach. Experimental results using a benchmark suite of complex loop nests are also presented.
Manuel Arenaz, Juan Touriño, Ramón Doallo
IPDPS1
2004 An Inspector-Executor Algorithm for Irregular Assignment Parallelization
Manuel Arenaz, Juan Touriño, Ramón Doallo
ISPA1
2003 A GSA-based compiler infrastructure to extract parallelism from complex loops
abstract
This paper presents a new approach for the detection of coarse-grain parallelism in loop nests that contain complex computations, including subscripted subscripts as well as conditional statements that introduce complex control flows at run-time. The approach is based on the recognition of the computational kernels calculated in a loop without considering the semantics of the code. The detection is carried out on top of the Gated Single Assignment (GSA) program representation at two different levels. First, the use-def chains between the statements that compose the strongly connected components (SCCs) of the GSA use-def chain graph are analyzed (intra-SCC analysis). As a result, the kernel computed in each SCC is recognized. Second, the use-def chains between statements of different SCCs are examined (inter-SCC analysis). This second abstraction level enables the detection of more complex computational kernels by the compiler. A prototype was implemented using the infrastructure provided by the Polaris compiler. Experimental results that show the effectiveness of our approach for the detection of coarse-grain parallelism in a suite of real codes are presented.
Manuel Arenaz, Juan Touriño, Ramón Doallo
ICS1
2002 Towards Detection of Coarse-Grain Loop-Level Parallelism in Irregular Computations
Manuel Arenaz, Juan Touriño, Ramón Doallo
Euro-Par1
2001 Efficient parallel numerical solver for the elastohydrodynamic Reynolds-Hertz problem
Manuel Arenaz, Ramón Doallo, Juan Touriño, Carlos Vázquez 0002
Parallel Comput.1