Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Agustín Fernández

dblp:98/4686 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
5 papers
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
GPUs and heterogeneous computing · 37% Memory systems · 32% High-performance computing · 14%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
loop transformation
0.132002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Performance Evaluation of Tiling for the Register Level · HPCA 1998
Loop Transformation Using Nonunimodular Matrices · IEEE Trans. Parallel Distributed Syst. 1995
Compilers and program optimization
loop optimization
0.122003
A Cost-Effective Implementation of Multilevel Tiling · IEEE Trans. Parallel Distributed Syst. 2003
On the Performance of Hand vs. Automatically Optimized Numerical Codes · HPCA 2000
Compilers and program optimization › loop transformation
register tiling
0.122002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Performance Evaluation of Tiling for the Register Level · HPCA 1998
GPUs and heterogeneous computing
GPU architecture
0.112005
Shader Performance Analysis on a Modern GPU Architecture · MICRO 2005
Compilers and program optimization › loop transformation
multi-level tiling
0.012003
A Cost-Effective Implementation of Multilevel Tiling · IEEE Trans. Parallel Distributed Syst. 2003
Memory systems › cache
cache optimization
0.012003
A Cost-Effective Implementation of Multilevel Tiling · IEEE Trans. Parallel Distributed Syst. 2003
Compilers and program optimization
compiler optimization
0.012002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Compilers and program optimization › instruction scheduling
instruction-level parallelism
0.012002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Compilers and program optimization › loop optimization
loop tiling
0.012002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
High-performance computing › performance optimization
numerical program optimization
0.012000
On the Performance of Hand vs. Automatically Optimized Numerical Codes · HPCA 2000
GPUs and heterogeneous computing
GPU performance analysis
0.012005
Shader Performance Analysis on a Modern GPU Architecture · MICRO 2005
Performance modeling and evaluation › performance evaluation methodology
simulation and benchmarking
0.012005
Shader Performance Analysis on a Modern GPU Architecture · MICRO 2005
Compilers and program optimization › parallelization › automatic parallelization
loop parallelization
0.011995
Loop Transformation Using Nonunimodular Matrices · IEEE Trans. Parallel Distributed Syst. 1995
Parallel and multicore computing › loop transformation › loop parallelization
loop nest parallelization
0.012003
A Cost-Effective Implementation of Multilevel Tiling · IEEE Trans. Parallel Distributed Syst. 2003
Memory systems
cache
0.012002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Memory systems › data locality
data reuse
0.012002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Parallel and multicore computing
parallelizing compiler
0.011995
Loop Transformation Using Nonunimodular Matrices · IEEE Trans. Parallel Distributed Syst. 1995

Methods — techniques the papers use, named apart from their topics

software pipelining · 0.1tiling · 0.1affine loop bounds computation · 0.1unroll-and-jam · 0.1loop tiling · 0.1workload characterization · 0.1multilevel tiling · 0.1architectural simulation · 0.1hermite normal form · 0.0outer unrolling · 0.0inner unrolling · 0.0
YearPublicationVenuePosition
2011 A take-home exam to assess professional skills
abstract
Professional Skills, such as the ability to communicate effectively or the ability to gather and integrate information, are not easy to teach or to assess. A traditional exam is not the best way of assessing these skills because it is limited both by time and by the resources students are able to consult. Moreover, in a traditional exam it is difficult to assess if professional skills have been acquired in depth. In this paper we propose to substitute the traditional exam by a take-home exam in which students have more time to solve the questions and are not restricted by the sources they can consult, thereby providing a highly educational task in which students experience a deep learning process. We also analyze what kind of questions should be asked to evaluate professional skills, as well as analyzing the potential drawbacks of these kind of exams (such as inappropriate student behavior). Finally, we show the results of one subject at the Barcelona School of Informatics, in which the take-home exam replaced the traditional exam. This course has been taught over 11 terms with good results.
David López 0001, José-Lorenzo Cruz, Fermín Sánchez, Agustín Fernández
FIE4
2010 A SIMD-efficient 14 instruction shader program for high-throughput microtriangle rasterization
Jordi Roca, Victor Moya Del Barrio, Vicente Escandell, Albert Murciego, Agustín Fernández, Roger Espasa
Vis. Comput.6
2006 ATTILA: a cycle-level execution-driven simulator for modern GPU architectures
abstract
The present work presents a cycle-level execution-driven simulator for modern GPU architectures. We discuss the simulation model used for our GPU simulator, based in the concept of boxes and signals, and the relation between the timing simulator and the functional emulator. The simulation model we use helps to increase the accuracy and reduce the number of errors in the timing simulator while allowing for an easy extensibility of the simulated GPU architecture. We also introduce the OpenGL framework used to feed the simulator with traces from real applications (UT2004, Doom3) and a performance debugging tool (Signal Trace Visualizer). The presented ATTILA simulator supports the simulation of a whole range of GPU configurations and architectures, from the embedded segment to the high end PC segment, supporting both the unified and non unified shader architectural models.
Victor Moya Del Barrio, Jordi Roca, Agustín Fernández, Roger Espasa
ISPASS4
2005 A Single (Unified) Shader GPU Microarchitecture for Embedded Systems
Victor Moya Del Barrio, Jordi Roca, Agustín Fernández, Roger Espasa
HiPEAC4
2005 Shader Performance Analysis on a Modern GPU Architecture
abstract
This paper presents an analysis of the performance of the shader processing units in a modern graphics processor unit (GPU) architecture using real graphic applications. The architecture of a modern GPU is described and a simulator and associated framework used to evaluate the architecture is introduced. The paper analyses the effects in performance of different configurations of the shader processing units and compares a classic GPU with a unified shader GPU. The evaluated unified shader architecture proves to be 15% to 30% more efficient, in terms of area, with a 2% to 7% improvement in performance when compared with a similar nonunified architecture
Victor Moya Del Barrio, Jordi Roca, Agustín Fernández, Roger Espasa
MICRO4
2003 A Cost-Effective Implementation of Multilevel Tiling
abstract
This paper presents a new cost-effective algorithm to compute exact loop bounds when multilevel tiling is applied to a loop nest having affine functions as bounds (nonrectangular loop nest). Traditionally, exact loop bounds computation has not been performed because its complexity is doubly exponential on the number of loops in the multilevel tiled code and, therefore, for certain classes of loops (i.e., nonrectangular loop nests), can be extremely time consuming. Although computation of exact loop bounds is not very important when tiling only for cache levels, it is critical when tiling includes the register level. This paper presents an efficient implementation of multilevel tiling that computes exact loop bounds and has a much lower complexity than conventional techniques. To achieve this lower complexity, our technique deals simultaneously with all levels to be tiled, rather than applying tiling level by level as is usually done. For loop nests having very simple affine functions as bounds, results show that our method is between 15 and 28 times faster than conventional techniques. For loop nests caving not so simple bounds, we have measured speedups as high as 2,300. Additionally, our technique allows eliminating redundant bounds efficiently. Results show that eliminating redundant bounds in our method is between 22 and 11 times faster than in conventional techniques for typical linear algebra programs.
Marta Jiménez, José María Llabería, Agustín Fernández
IEEE Trans. Parallel Distributed Syst.3
2002 Register tiling in nonrectangular iteration spaces
abstract
Loop tiling is a well-known loop transformation generally used to expose coarse-grain parallelism and to exploit data reuse at the cache level. Tiling can also be used to exploit data reuse at the register level and to improve a program's ILP. However, previous proposals in the literature (as well as commercial compilers) are only able to perform multidimensional tiling for the register level when the iteration space is rectangular. In this article we present a new general algorithm to perform multidimensional tiling for the register level in both rectangular and nonrectangular iteration spaces. We also propose a simple heuristic to determine the tiling parameters at this level. Finally, we evaluate our method using as benchmarks typical linear algebra algorithms having nonrectangular iteration spaces and compare our proposal against hand-optimized vendor-supplied numerical libraries and against commercial compilers able to perform optimizing code transformations such as inner unrolling, unroll-and-jam, and software pipelining. Measurements were taken on three different superscalar microprocessors. Results will show that our method outperforms the native compilers (showing speedups of 2.5 in average) and matches the performance of vendor-supplied numerical libraries. The general conclusion is that compiler technology can make it possible for nonrectangular loop nests to achieve as high performance as hand-optimized codes.
Marta Jiménez, José María Llabería, Agustín Fernández
ACM Trans. Program. Lang. Syst.3
2000 On the Performance of Hand vs. Automatically Optimized Numerical Codes
abstract
In this paper, we compare automatic-optimized codes against hand-optimized codes. The automatic-optimized codes have been generated using our own developed tool that implements compiler techniques proposed in our previous work. Our compiler techniques focus on applying multilevel tiling to non-rectangular loop nests. This type of loop nests are commonly found in linear algebra algorithms, typically used in numerical codes. As hand-optimized codes, we use two different numerical libraries: the BLAS3 library provided by the manufacturers and the RISC-BLAS library proposed in Dayde and Duff (1998). Results will show how compiler technology can make it possible for non-rectangular loop nests to achieve as high performance as hand-optimized codes on modern microprocessors.
Marta Jiménez, José María Llabería, Agustín Fernández
HPCA3
1998 Performance Evaluation of Tiling for the Register Level
abstract
Tiling is a well-known loop transformation, which is basically used to expose coarse-grain parallelism and to exploit data reuse at the cache level. However, it can also be used to exploit data reuse at the register level and to improve programs's ILP. Previous work on tiling and also commercial compilers are able to perform tiling for the register level in more than one dimension when the iteration space is rectangular. Non-rectangular iteration spaces are commonly found in linear algebra algorithms or can arise as a result of applying previous transformations such as loop skewing. In this paper we evaluate the technique presented in Jimenez et al. (1996) which is able to perform tiling for the register level in more than one dimension in both rectangular and non-rectangular iteration spaces. We use typical linear algebra algorithms having non-rectangular iteration spaces as benchmarks and compare our proposal against commercial preprocessors able to perform optimizing code transformations such as inner unrolling, outer unrolling and software pipelining. We will also present quantitative data showing the benefits of tiling only for the register level, tiling only for the cache level and tiling for both levels simultaneously. Results measured on a ALPHA 21164 processor show that tiling for both cache and register levels improves upon commercial compilers and preprocessors by factors in the range of 1.3 to 6.3.
Marta Jiménez, José María Llabería, Agustín Fernández
HPCA3
1998 A General Algorithm for Tiling the Register Level
abstract
Tiling is a well-known loop transformation that can be used to exploit data reuse at the register level and to improve a program’s ILP. Previous work on tiling and also commercial compilers are able to perform tiling for the register level in more than one dimension when the iteration space is rectangular. However, they either cannot handle or can only handle limited cases of non-rectangular iteration spaces. Nonrectangular iteration spaces 1 are commonly found in linear algebra algorithms or can arise as a result of applying previous transformations such as loop skewing. In this paper we present a new general algorithm to perform tiling for the register level in more than one dimension in both rectangular and nonrectangular iteration spaces. Our method uses index set splitting to distinguish loop nests that traverse boundary tiles of the tiled iteration space from loop nests that traverse nonboundary tiles. We evaluate our method using as benchmarks typical linear algebra algorithms having non-rectangular iteration spaces. Results measured on both ALPHA 21064 and MIPS R10000 machines show that our method achieves speedups in the range of 1.11 to 5.96 over commercial compilers and preprocessors able to perform optimizing code transformations. 2.
Marta Jiménez, José María Llabería, Agustín Fernández, Enric Morancho
International Conference on Supercomputing3
1995 Loop Transformation Using Nonunimodular Matrices
abstract
Linear transformations are widely used to vectorize and parallelize loops. A subset of these transformations are unimodular transformations. When a unimodular transformation is used, the exact bounds of the transformed loop nest are easily computed and the steps of the loops are equal to 1. Unimodular loop transformations have been widely used since they permit the implementation of many useful loop transformations. Recently, nonunimodular transformations have been proposed to reduce communication requirements or to use the memory hierarchy efficiently. The methods used for unimodular transformations do not work in the case of nonunimodular transformations, since they do not produce the exact bounds of the transformed loop nest. In this paper, we present a method for nested loop transformation which gives the exact bounds for both unimodular and nonunimodular transformations. The basic idea is to use the Hermite Normal Form (HNF) of the transformation matrix.>
Agustín Fernández, José María Llabería, Miguel Valero-García
IEEE Trans. Parallel Distributed Syst.1
1992 Scheduling partitions in systolic algorithms
abstract
The authors present a technique for scheduling partitions in systolic algorithms (SA). This technique can be used in combination with any possible projection used for the problem dependent size SA design and with any possible spatial mapping used for the partitions. They also present the necessary code transformations to transform the sequential code into the code that is executed in a processing element (PE) of the systolic processor (SP). This technique is applied to the matrix by vector problem using a non-unimodular transformation matrix and taking into account input and output of data. Cut&pile spatial mapping is used for the partitions.>
Alvaro Suárez, José María Llabería, Agustín Fernández
ASAP3
1991 Transformation of systolic algorithms for interleaving partitions
abstract
A systematic method to map systolic problems onto multicomputers is presented. A systolic problem is a problem for which it is possible to design a systolic algorithm. This method selects and transforms the systolic algorithm into a parallel algorithm with high granularity. The communications requirements are reduced and the performance can be increased. The proposed scheme requires a classification of dependences and it is based on the interleaved execution of several partitions of the systolic algorithm. The code to be executed in a processing element of the multicomputer system is obtained through application of the proposed systematic transformations to the original sequential code. By applying this method to algebraic path problem (APP), the authors illustrate the main features, and several performance measures for a torus of transputers system are presented, considering the various algorithms which are unified by the APP.>
Agustín Fernández, José María Llabería, Juan J. Navarro, Miguel Valero-García
ASAP1
1991 Performance evaluation of transputer systems with linear algebra problems
Agustín Fernández, José María Llabería, Juan J. Navarro, Miguel Valero-García
Microprocessing and Microprogramming1