Marta Jiménez

dblp:29/5669 · also Marta Jiménez Castells · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
0since 2021 · last 2003
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
4 papers
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 61% High-performance computing · 27% Parallel and multicore computing · 12%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
loop optimization
0.122003
A Cost-Effective Implementation of Multilevel Tiling · IEEE Trans. Parallel Distributed Syst. 2003
On the Performance of Hand vs. Automatically Optimized Numerical Codes · HPCA 2000
Compilers and program optimization
loop transformation
0.122002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Performance Evaluation of Tiling for the Register Level · HPCA 1998
Compilers and program optimization › loop transformation
register tiling
0.122002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Performance Evaluation of Tiling for the Register Level · HPCA 1998
Compilers and program optimization › loop transformation
multi-level tiling
0.012003
A Cost-Effective Implementation of Multilevel Tiling · IEEE Trans. Parallel Distributed Syst. 2003
Memory systems › cache
cache optimization
0.012003
A Cost-Effective Implementation of Multilevel Tiling · IEEE Trans. Parallel Distributed Syst. 2003
Compilers and program optimization
compiler optimization
0.012002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Compilers and program optimization › instruction scheduling
instruction-level parallelism
0.012002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Compilers and program optimization › loop optimization
loop tiling
0.012002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
High-performance computing › performance optimization
numerical program optimization
0.012000
On the Performance of Hand vs. Automatically Optimized Numerical Codes · HPCA 2000
Parallel and multicore computing › loop transformation › loop parallelization
loop nest parallelization
0.012003
A Cost-Effective Implementation of Multilevel Tiling · IEEE Trans. Parallel Distributed Syst. 2003
Memory systems
cache
0.012002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002
Memory systems › data locality
data reuse
0.012002
Register tiling in nonrectangular iteration spaces · ACM Trans. Program. Lang. Syst. 2002

Methods — techniques the papers use, named apart from their topics

software pipelining · 0.1tiling · 0.1affine loop bounds computation · 0.1unroll-and-jam · 0.1loop tiling · 0.1multilevel tiling · 0.1outer unrolling · 0.0inner unrolling · 0.0
YearPublicationVenuePosition
2003 A Cost-Effective Implementation of Multilevel Tiling
abstract
This paper presents a new cost-effective algorithm to compute exact loop bounds when multilevel tiling is applied to a loop nest having affine functions as bounds (nonrectangular loop nest). Traditionally, exact loop bounds computation has not been performed because its complexity is doubly exponential on the number of loops in the multilevel tiled code and, therefore, for certain classes of loops (i.e., nonrectangular loop nests), can be extremely time consuming. Although computation of exact loop bounds is not very important when tiling only for cache levels, it is critical when tiling includes the register level. This paper presents an efficient implementation of multilevel tiling that computes exact loop bounds and has a much lower complexity than conventional techniques. To achieve this lower complexity, our technique deals simultaneously with all levels to be tiled, rather than applying tiling level by level as is usually done. For loop nests having very simple affine functions as bounds, results show that our method is between 15 and 28 times faster than conventional techniques. For loop nests caving not so simple bounds, we have measured speedups as high as 2,300. Additionally, our technique allows eliminating redundant bounds efficiently. Results show that eliminating redundant bounds in our method is between 22 and 11 times faster than in conventional techniques for typical linear algebra programs.
Marta Jiménez, José María Llabería, Agustín Fernández
IEEE Trans. Parallel Distributed Syst.1
2002 Register tiling in nonrectangular iteration spaces
abstract
Loop tiling is a well-known loop transformation generally used to expose coarse-grain parallelism and to exploit data reuse at the cache level. Tiling can also be used to exploit data reuse at the register level and to improve a program's ILP. However, previous proposals in the literature (as well as commercial compilers) are only able to perform multidimensional tiling for the register level when the iteration space is rectangular. In this article we present a new general algorithm to perform multidimensional tiling for the register level in both rectangular and nonrectangular iteration spaces. We also propose a simple heuristic to determine the tiling parameters at this level. Finally, we evaluate our method using as benchmarks typical linear algebra algorithms having nonrectangular iteration spaces and compare our proposal against hand-optimized vendor-supplied numerical libraries and against commercial compilers able to perform optimizing code transformations such as inner unrolling, unroll-and-jam, and software pipelining. Measurements were taken on three different superscalar microprocessors. Results will show that our method outperforms the native compilers (showing speedups of 2.5 in average) and matches the performance of vendor-supplied numerical libraries. The general conclusion is that compiler technology can make it possible for nonrectangular loop nests to achieve as high performance as hand-optimized codes.
Marta Jiménez, José María Llabería, Agustín Fernández
ACM Trans. Program. Lang. Syst.1
2000 On the Performance of Hand vs. Automatically Optimized Numerical Codes
abstract
In this paper, we compare automatic-optimized codes against hand-optimized codes. The automatic-optimized codes have been generated using our own developed tool that implements compiler techniques proposed in our previous work. Our compiler techniques focus on applying multilevel tiling to non-rectangular loop nests. This type of loop nests are commonly found in linear algebra algorithms, typically used in numerical codes. As hand-optimized codes, we use two different numerical libraries: the BLAS3 library provided by the manufacturers and the RISC-BLAS library proposed in Dayde and Duff (1998). Results will show how compiler technology can make it possible for non-rectangular loop nests to achieve as high performance as hand-optimized codes on modern microprocessors.
Marta Jiménez, José María Llabería, Agustín Fernández
HPCA1
1998 Performance Evaluation of Tiling for the Register Level
abstract
Tiling is a well-known loop transformation, which is basically used to expose coarse-grain parallelism and to exploit data reuse at the cache level. However, it can also be used to exploit data reuse at the register level and to improve programs's ILP. Previous work on tiling and also commercial compilers are able to perform tiling for the register level in more than one dimension when the iteration space is rectangular. Non-rectangular iteration spaces are commonly found in linear algebra algorithms or can arise as a result of applying previous transformations such as loop skewing. In this paper we evaluate the technique presented in Jimenez et al. (1996) which is able to perform tiling for the register level in more than one dimension in both rectangular and non-rectangular iteration spaces. We use typical linear algebra algorithms having non-rectangular iteration spaces as benchmarks and compare our proposal against commercial preprocessors able to perform optimizing code transformations such as inner unrolling, outer unrolling and software pipelining. We will also present quantitative data showing the benefits of tiling only for the register level, tiling only for the cache level and tiling for both levels simultaneously. Results measured on a ALPHA 21164 processor show that tiling for both cache and register levels improves upon commercial compilers and preprocessors by factors in the range of 1.3 to 6.3.
Marta Jiménez, José María Llabería, Agustín Fernández
HPCA1
1998 A General Algorithm for Tiling the Register Level
abstract
Tiling is a well-known loop transformation that can be used to exploit data reuse at the register level and to improve a program’s ILP. Previous work on tiling and also commercial compilers are able to perform tiling for the register level in more than one dimension when the iteration space is rectangular. However, they either cannot handle or can only handle limited cases of non-rectangular iteration spaces. Nonrectangular iteration spaces 1 are commonly found in linear algebra algorithms or can arise as a result of applying previous transformations such as loop skewing. In this paper we present a new general algorithm to perform tiling for the register level in more than one dimension in both rectangular and nonrectangular iteration spaces. Our method uses index set splitting to distinguish loop nests that traverse boundary tiles of the tiled iteration space from loop nests that traverse nonboundary tiles. We evaluate our method using as benchmarks typical linear algebra algorithms having non-rectangular iteration spaces. Results measured on both ALPHA 21064 and MIPS R10000 machines show that our method achieves speedups in the range of 1.11 to 5.96 over commercial compilers and preprocessors able to perform optimizing code transformations. 2.
Marta Jiménez, José María Llabería, Agustín Fernández, Enric Morancho
International Conference on Supercomputing1