Josep Llosa

dblp:92/1611 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
1since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
13 papers
Processor architecture and microarchitecture · 69% Memory systems · 12% Performance modeling and evaluation · 9%
Software engineering, system software, and programming languages
14 papers
Compilers and program optimization · 100%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
instruction scheduling
0.262004
Register Constrained Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2004
Reduced code size modulo scheduling in the absence of hardware support · MICRO 2002
Improved spill code generation for software pipelined loops · PLDI 2000
Compilers and program optimization › instruction scheduling › software pipelining
modulo scheduling
0.262004
Register Constrained Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2004
Reduced code size modulo scheduling in the absence of hardware support · MICRO 2002
Lifetime-Sensitive Modulo Scheduling in a Production Environment · IEEE Trans. Computers 2001
Compilers and program optimization › instruction scheduling
software pipelining
0.272004
Register Constrained Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2004
Lifetime-Sensitive Modulo Scheduling in a Production Environment · IEEE Trans. Computers 2001
Improved spill code generation for software pipelined loops · PLDI 2000
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
clustered VLIW
0.122010
CSMT: Simultaneous Multithreading for Clustered VLIW Processors · IEEE Trans. Computers 2010
Modulo scheduling with integrated register spilling for clustered VLIW architectures · MICRO 2001
Processor architecture and microarchitecture › instruction-level parallelism
VLIW
0.162001
Cost-Conscious Strategies to Increase Performance of Numerical Programs on Aggressive VLIW Architectures · IEEE Trans. Computers 2001
Modulo scheduling with integrated register spilling for clustered VLIW architectures · MICRO 2001
Distributed Modulo Scheduling · HPCA 1999
Processor architecture and microarchitecture › multithreading
simultaneous multithreading
0.112010
CSMT: Simultaneous Multithreading for Clustered VLIW Processors · IEEE Trans. Computers 2010
Compilers and program optimization
register allocation
0.152004
Register Constrained Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2004
Modulo Scheduling with Reduced Register Pressure · IEEE Trans. Computers 1998
Heuristics for Register-Constrained Software Pipelining · MICRO 1996
Processor architecture and microarchitecture
instruction-level parallelism
0.162001
Cost-Conscious Strategies to Increase Performance of Numerical Programs on Aggressive VLIW Architectures · IEEE Trans. Computers 2001
Widening Resources: A Cost-effective Technique for Aggressive ILP Architectures · MICRO 1998
Non-Consistent Dual Register Files to Reduce Register Pressure · HPCA 1995
Compilers and program optimization › register allocation
register pressure reduction
0.132004
Register Constrained Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2004
Modulo Scheduling with Reduced Register Pressure · IEEE Trans. Computers 1998
Hypernode reduction modulo scheduling · MICRO 1995
Compilers and program optimization › memory optimization
data locality optimization
0.122005
An accurate cost model for guiding data locality transformations · ACM Trans. Program. Lang. Syst. 2005
A fast and accurate framework to analyze and optimize cache memory behavior · ACM Trans. Program. Lang. Syst. 2004
Memory systems
cache
0.122005
A fast and accurate framework to analyze and optimize cache memory behavior · ACM Trans. Program. Lang. Syst. 2004
An accurate cost model for guiding data locality transformations · ACM Trans. Program. Lang. Syst. 2005
Compilers and program optimization › memory optimization
data layout transformation
0.112005
An accurate cost model for guiding data locality transformations · ACM Trans. Program. Lang. Syst. 2005
Compilers and program optimization › loop optimization
loop tiling
0.112005
An accurate cost model for guiding data locality transformations · ACM Trans. Program. Lang. Syst. 2005
Memory systems › cache
cache behavior
0.012004
A fast and accurate framework to analyze and optimize cache memory behavior · ACM Trans. Program. Lang. Syst. 2004
Performance modeling and evaluation › cache performance modeling
cache miss equation
0.012004
A fast and accurate framework to analyze and optimize cache memory behavior · ACM Trans. Program. Lang. Syst. 2004
Processor architecture and microarchitecture
out-of-order execution
0.012004
Out-of-Order Commit Processors · HPCA 2004
Performance modeling and evaluation
workload characterization
0.012004
A fast and accurate framework to analyze and optimize cache memory behavior · ACM Trans. Program. Lang. Syst. 2004
Compilers and program optimization
code size reduction
0.012002
Reduced code size modulo scheduling in the absence of hardware support · MICRO 2002
Processor architecture and microarchitecture
multithreading
0.012010
CSMT: Simultaneous Multithreading for Clustered VLIW Processors · IEEE Trans. Computers 2010
Parallel and multicore computing
thread-level parallelism
0.012010
CSMT: Simultaneous Multithreading for Clustered VLIW Processors · IEEE Trans. Computers 2010
Compilers and program optimization
numerical program optimization
0.012001
Cost-Conscious Strategies to Increase Performance of Numerical Programs on Aggressive VLIW Architectures · IEEE Trans. Computers 2001
Processor architecture and microarchitecture
instruction scheduling
0.012001
Modulo scheduling with integrated register spilling for clustered VLIW architectures · MICRO 2001
Processor architecture and microarchitecture › instruction scheduling
software pipelining
0.012001
Modulo scheduling with integrated register spilling for clustered VLIW architectures · MICRO 2001
Processor architecture and microarchitecture › register file
hierarchical register file
0.012000
Two-level hierarchical register file organization for VLIW processors · MICRO 2000
Processor architecture and microarchitecture › register file
register file organization
0.012000
Two-level hierarchical register file organization for VLIW processors · MICRO 2000
Processor architecture and microarchitecture
clustered architecture
0.011999
Distributed Modulo Scheduling · HPCA 1999
Parallel and multicore computing
software partitioning
0.011999
Distributed Modulo Scheduling · HPCA 1999
Memory systems › cache
cache miss
0.012005
An accurate cost model for guiding data locality transformations · ACM Trans. Program. Lang. Syst. 2005
Distributed systems › fault tolerance
checkpointing
0.012004
Out-of-Order Commit Processors · HPCA 2004
Processor architecture and microarchitecture
register file
0.011995
Non-Consistent Dual Register Files to Reduce Register Pressure · HPCA 1995

Methods — techniques the papers use, named apart from their topics

modulo scheduling · 0.2backtracking · 0.1cluster renaming · 0.1genetic algorithm · 0.1cost model · 0.1sampling · 0.1polyhedral analysis · 0.1speculative scheduling · 0.1scheduling heuristics · 0.1spilling · 0.0simulation · 0.0heuristics · 0.0
YearPublicationVenuePosition
2025 The European master for HPC curriculum
abstract
International audience
Pascal Bouvry, Mats Brorsson, Ramon Canal, Aryan Eftekhari, Siegfried Höfinger, Didier Smets, Harald Köstler, Tomás Kozubek, Ezhilmathi Krishnasamy, Josep Llosa, Alexandra Lukas-Rother, Xavier Martorell, Dirk Pleiter, Ana Proykova, Maria-Ribera Sancho, Olaf Schenk, Cristina Silvano
J. Parallel Distributed Comput.10
2010 A low cost split-issue technique to improve performance of SMT clustered VLIW processors
abstract
Very Long Instruction Word (VLIW) processors are a popular choice in embedded domain due to their hardware simplicity, low cost and low power consumption. Simultaneous MultiThreading (SMT) is a popular technique for improving processor performance. To maintain execution semantics, a VLIW instruction needs to be issued in entirety, which restricts the opportunities in SMT. Split-issue at operation-level is a technique that allows issuing a VLIW instruction in parts without breaking execution semantics. Issuing an instruction in parts allows non-conflicting part of an instruction to be issued along with other instructions and improves SMT performance. However, implementing split-issue at operation-level requires complex structures and is not practical for an embedded VLIW processor. This paper proposes cluster-level split-issue, which implements split-issue at a cluster-level boundary for clustered VLIW processors. Cluster-level split-issue has a very low hardware overhead in contrast to split-issue at operation-level. Experimental results show that cluster-level split-issue, despite being more restrictive than split-issue at operation-level, achieves similar performance and improves SMT performance significantly.
Manoj Gupta 0001, Fermín Sánchez, Josep Llosa
IPDPS3
2010 CSMT: Simultaneous Multithreading for Clustered VLIW Processors
abstract
Simultaneous MultiThreading (SMT) is a well-known technique that improves resource utilization by exploiting thread-level parallelism at the instruction grain level. However, implementing SMT for VLIWs requires complex structures, which is contrary to the VLIW philosophy of hardware simplicity. In this paper, we propose Cluster-level Simultaneous MultiThreading (CSMT) to allow some degree of SMT in clustered VLIW processors with low hardware cost and complexity. CSMT considers the set of operations that execute simultaneously in a given cluster as the assignment unit. To minimize cluster conflicts between threads, a very simple hardware-based cluster renaming mechanism is proposed. The hardware required to implement CSMT is cheap, realistic, and practical for a clustered VLIW processor. An analysis of the hardware required to implement CSMT shows that it is quite scalable, with up to eight threads easily supported at low hardware cost. The experimental results show that CSMT significantly improves performance when compared with other multithreading approaches suited for VLIW. For instance, with four threads, CSMT shows an average speedup of 110 percent over a single-thread VLIW architecture and 40 percent over Interleaved MultiThreading (IMT). In some cases, speedup can be as high as 225 percent over single-thread architecture and 84 percent over IMT.
Manoj Gupta 0001, Fermín Sánchez, Josep Llosa
IEEE Trans. Computers3
2009 Hybrid multithreading for VLIW processors
abstract
© ACM, 2009. This is the author's version of the work: http://doi.acm.org/10.1145/1629395.1629403
Manoj Gupta 0001, Fermín Sánchez, Josep Llosa
CASES3
2009 Thread Merging Schemes for Multithreaded Clustered VLIW Processors
abstract
Several multithreading techniques have been proposed to reduce the resource underutilization in very long instruction word (VLIW) processors. Simultaneous MultiThreading (SMT) is a popular technique which improves processor performance by issuing multiple instructions from different threads. SMT requires extra hardware to merge instructions from different threads. The complexity of this hardware increases substantially with the number of threads, limiting the number of threads that can be realistically supported to only 2. Cluster-level Simultaneous MultiThreading (CSMT) is a technique that merges instructions from threads at the cluster level. CSMT has a much lower merging hardware cost and can support a larger number of threads. However, CSMT performance is lower than SMT. In this paper, we evaluate several hardware designs that can support a high number of threads by using a merging scheme that combines both SMT and CSMT merging. For instance, one of the evaluated schemes, which merges the first 2 threads using SMT and the produced merging with other 2 threads by CSMT, achieves performance similar to supporting 4 threads by SMT but maintaining a reasonable merging hardware cost.
Manoj Gupta 0001, Fermín Sánchez, Josep Llosa
ICPP3
2007 Merge Logic for Clustered Multithreaded VLIW Processors
abstract
Clustered VLIW embedded processors have become widespread due to benefits of simple hardware and low power. Simultaneous MultiThreading (SMT) is a well known technique that uses thread level parallelism at the instruction grain level. However, implementing SMT for VLIW requires complex structures. CSMT (cluster-level simultaneous MultiThreading) allows some degree of SMT in clustered VLIW processors with minimal hardware cost and complexity. This paper deals with the hardware required to implement CSMT instruction merge logic on a clustered VLIW processor. The paper presents two implementations of CSMT merge logic and an analysis of both comparing design issues like delay and number of transistors required.
Manoj Gupta 0001, Fermín Sánchez, Josep Llosa
DSD3
2007 Silicon Compaction/Defragmentation for Partial Runtime Reconfiguration
abstract
The effective use of Run Time Reconfiguration (RTR) in modern FPGAs opens up new avenues to design area and power efficient high performance architectures. However the current design flow for exploiting RTR in designs, leads to the problem of silicon Defragmentation. We propose a silicon compaction/ defragmentation technique which works on already placed and routed modules to generate partial bitstreams (programming files) for the device. We have outlined a method which generates these partial bitstreams very fast taking into account the size and position of the "free" silicon when the device is in operation. The other advantage of this method is that the changes in the basic FPGA fabric needed to implement this defragmentation strategy are (almost) trivial.
Kolin Paul, Joël Porquet-Lupine, Josep Llosa
DSD3
2007 Cluster-level simultaneous multithreading for VLIW processors
abstract
Clustered VLIW embedded processors have become widespread due to benefits of simple hardware and low power. However, while some applications exhibit large amounts of instruction level parallelism (ILP) and benefit from very wide machines, others have little ILP, which wastes precious resources in wide processors. Simultaneous multithreading (SMT) is a well known technique that improves resource utilization by exploiting thread level parallelism at the instruction grain level. However, implementing SMT for VLIWs requires complex structures. In this paper, we propose CSMT (cluster-level simultaneous multithreading) to allow some degree of SMT in clustered VLIW processors with minimal hardware cost and complexity. CSMT considers the set of operations that execute simultaneously in a given cluster (named bundle) as the assignment unit. All bundles belonging to a VLIW instruction from a given thread are issued simultaneously. To minimize cluster conflicts between threads, a very simple hardware- based cluster renaming mechanism is proposed. The experimental results show that CSMT significantly improves ILP when compared with other multithreading approaches suited for VLIW. For instance, with 4 threads CSMT shows an average speedup of 113% over a single-thread VLIW architecture and 36% over interleaved multithreading (IMT). In some cases, speedup can be as high as 228% over single thread architecture and 97% over IMT.
Manoj Gupta 0001, Fermín Sánchez, Josep Llosa
ICCD3
2005 An accurate cost model for guiding data locality transformations
abstract
Caches have become increasingly important with the widening gap between main memory and processor speeds. Small and fast cache memories are designed to bridge this discrepancy. However, they are only effective when programs exhibit sufficient data locality.The performance of the memory hierarchy can be improved by means of data and loop transformations. Tiling is a loop transformation that aims at reducing capacity misses by shortening the reuse distance. Padding is a data layout transformation targeted to reduce conflict misses.This article presents an accurate cost model that describes misses across different hierarchy levels and considers the effects of other hardware components such as branch predictors. The cost model drives the application of tiling and padding transformations. We combine the cost model with a genetic algorithm to compute the tile and pad factors that enhance the program performance.To validate our strategy, we ran experiments for a set of benchmarks on a large set of modern architectures. Our results show that this scheme is useful to optimize programs' performance. When compared to previous approaches, we observe that with a reasonable compile-time overhead, our approach gives significant performance improvements for all studied kernels on all architectures.
Xavier Vera, Jaume Abella 0001, Josep Llosa, Antonio González 0001
ACM Trans. Program. Lang. Syst.3
2004 Out-of-Order Commit Processors
abstract
Modern out-of-order processors tolerate long latency memory operations by supporting a large number of in-flight instructions. This is particularly useful in numerical applications where branch speculation is normally not a problem and where the cache hierarchy is not capable of delivering the data soon enough. In order to support more in-flight instructions, several resources have to be up-sized, such as the reorder buffer (ROB), the general purpose instructions queues, the load/store queue and the number of physical registers in the processor. However, scaling-up the number of entries in these resources is impractical because of area, cycle time, and power consumption constraints. We propose to increase the capacity of future processors by augmenting the number of in-flight instructions. Instead of simply up-sizing resources, we push for new and novel microarchitectural structures that achieve the same performance benefits but with a much lower need for resources. Our main contribution is a new checkpointing mechanism that is capable of keeping thousands of in-flight instructions at a practically constant cost. We also propose a queuing mechanism that takes advantage of the differences in waiting time of the instructions in the flow. Using these two mechanisms our processor has a performance degradation of only 10% for SPEC2000fp over a conventional processor requiring more than an order of magnitude additional entries in the ROB and instruction queues, and about a 200% improvement over a current processor with a similar number of entries.
Adrián Cristal, Daniel Ortega, Josep Llosa, Mateo Valero
HPCA3
2004 A fast and accurate framework to analyze and optimize cache memory behavior
abstract
The gap between processor and main memory performance increases every year. In order to overcome this problem, cache memories are widely used. However, they are only effective when programs exhibit sufficient data locality. Compile-time program transformations can significantly improve the performance of the cache. To apply most of these transformations, the compiler requires a precise knowledge of the locality of the different sections of the code, both before and after being transformed.Cache miss equations (CMEs) allow us to obtain an analytical and precise description of the cache memory behavior for loop-oriented codes. Unfortunately, a direct solution of the CMEs is computationally intractable due to its NP-complete nature.This article proposes a fast and accurate approach to estimate the solution of the CMEs. We use sampling techniques to approximate the absolute miss ratio of each reference by analyzing a small subset of the iteration space. The size of the subset, and therefore the analysis time, is determined by the accuracy selected by the user. In order to reduce the complexity of the algorithm to solve CMEs, effective mathematical techniques have been developed to analyze the subset of the iteration space that is being considered. These techniques exploit some properties of the particular polyhedra represented by CMEs.
Xavier Vera, Nerina Bermudo, Josep Llosa, Antonio González 0001
ACM Trans. Program. Lang. Syst.3
2004 Register Constrained Modulo Scheduling
abstract
Software pipelining is an instruction scheduling technique that exploits the instruction level parallelism (ILP) available in loops by overlapping operations from various successive loop iterations. The main drawback of aggressive software pipelining techniques is their high register requirements. If the requirements exceed the number of registers available in the target architecture, some steps need to be applied to reduce the register pressure (incurring some performance degradation): reduce iteration overlapping or spilling some lifetimes to memory. In the first part, we propose a set of heuristics to improve the spilling process and to better decide between adding spill code or directly decreasing the execution rate of iterations. The experimental evaluation, over a large number of representative loops and for a processor configuration, reports an increase in performance by a factor of 1.29 and a reduction of memory traffic by a factor of 1.36. In the second part, we analyze the use of backtracking and propose a novel approach for simultaneous instruction scheduling and register spilling in modulo scheduling: MIPS (modulo scheduling with integrated register spilling). The experimental evaluation reports an increase in performance by a factor of 1.46 and a reduction of the memory traffic by a factor of 1.66 (or an additional 1.13 and 1.22 with regard to the proposal in the first part). These improvements are achieved at the expense of a reasonable increase in the compilation time.
Javier Zalamea, Josep Llosa, Eduard Ayguadé, Mateo Valero
IEEE Trans. Parallel Distributed Syst.2
2002 A comparative study of modulo scheduling techniques
abstract
Modulo Scheduling is an instruction scheduling technique that is used by many current compilers. Different approaches have been proposed in the past but there is not a quantitative comparison among them, using the same compiling platform, benchmarks and architectures.This paper presents a performance comparison of the most relevant Modulo Scheduling techniques, based on a detailed quantitative evaluation of them. The results point out which are the most effective techniques for different architectures, which is useful for compiler designers when choosing the most appropriate technique for a particular processor architecture.
Josep M. Codina, Josep Llosa, Antonio González 0001
ICS2
2002 Reduced code size modulo scheduling in the absence of hardware support
abstract
Modulo scheduling is a very effective instruction scheduling technique that exploits Instruction Level Parallelism (ILP) in loop bodies by overlapping the execution of successive iterations. Unfortunately, modulo scheduling has been shown to cause heavy code expansion. To avoid the penalties of code expansion, some processors have dedicated hardware support for modulo scheduled loops. However, this dedicated hardware support has a cost in chip area, cycle time, processor complexity, and compiler complexity. This paper shows that the right combination of scheduling heuristics combined with speculative modulo scheduling can significantly reduce code expansion. In addition, several code generation schema heuristics are proposed to further reduce code expansion. The evaluations show that loops can be effectively modulo scheduled with an average code expansion only 1.5 times the original loop size. Compared with a state of the art modulo scheduler, our code size sensitive heuristics reduce the size of embedded domain benchmarks binaries by 30% on average. While performance is mostly unchanged, some applications show speed-ups up to 20% due to a reduction in instruction cache capacity misses.
Josep Llosa, Stefan M. Freudenberger
MICRO1
2001 Modulo scheduling with integrated register spilling for clustered VLIW architectures
abstract
Clustering is a technique to decentralize the design of future wide issue VLIW cores and enable them to meet the technology constraints in terms of cycle time, area and power dissipation. In a clustered design, registers and functional units are grouped in clusters so that new instructions are needed to move data between them. New aggressive instruction scheduling techniques are required to minimize the negative effect of resource clustering and delays in moving data around. In this paper we present a novel software pipelining technique that performs instruction scheduling with reduced register requirements, register allocation, register spilling and inter-cluster communication in a single step. The algorithm uses limited backtracking to reconsider previously taken decisions. This backtracking provides the algorithm with additional possibilities for obtaining high throughput schedules with low spill code requirements for clustered architectures. We show that the proposed approach outperforms previously proposed techniques and that it is very scalable independently of the number of clusters, the number of communication buses and communication latency. The paper also includes an exploration of some parameters in the design of future clustered VLIW cores.
Javier Zalamea, Josep Llosa, Eduard Ayguadé, Mateo Valero
MICRO2
2001 Lifetime-Sensitive Modulo Scheduling in a Production Environment
abstract
This paper presents a novel software pipelining approach, which is called Swing Modulo Scheduling (SMS). It generates schedules that are near optimal in terms of initiation interval, register requirements, and stage count. Swing Modulo Scheduling is a heuristic approach that has a low computational cost. This paper first describes the technique and evaluates it for the Perfect Club benchmark suite on a generic VLIW architecture. SMS is compared with other heuristic methods, showing that it outperforms them in terms of the quality of the obtained schedules and compilation time. To further explore the effectiveness of SMS, the experience of incorporating it into a production quality compiler for the Equator MAP1000 processor is described; implementation issues are discussed, as well as modifications and improvements to the original algorithm. Finally, experimental results from using a set of industrial multimedia applications are presented.
Josep Llosa, Eduard Ayguadé, Antonio González 0001, Mateo Valero, Jason Eckhardt
IEEE Trans. Computers1
2001 Cost-Conscious Strategies to Increase Performance of Numerical Programs on Aggressive VLIW Architectures
abstract
Loops are the main time-consuming part of numerical applications. The performance of the loops is limited either by the resources offered by the architecture or by recurrences in the computation. To execute more operations per cycle, current processors are designed with growing degrees of resource replication (replication technique) for memory ports and functional units. However, the high cost in terms of area and cycle time of this technique precludes the use of high degrees of replication. High values for the cycle time may clearly offset any gain in terms of number of execution cycles. High values for the area may lead to an unimplementable configuration. An alternative to resource replication is resource widening (widening technique), which has also been used in some recent designs in which the width of the resources is increased (i.e., a single operation is performed over multiple data). Moreover, several general-purpose superscalar microprocessors have been implemented with multiply-add fused floating-point units (fusion technique), which reduces the latency of the combined operation and the number of resources used. The authors evaluate a broad set of VLIW processor design alternatives that combine the three techniques. We perform a technological projection for the next processor generations in order to foresee the possible implementable alternatives. From this study, we conclude that if the cost is taken into account, combining certain degrees of replication and widening in the hardware resources is more effective than applying only replication. Also, we confirm that multiply-add fused units will have a significant impact in raising the performance of future processor architectures with a reasonable increase in cost.
David López 0001, Josep Llosa, Mateo Valero, Eduard Ayguadé
IEEE Trans. Computers2
2000 A Fast and Accurate Approach to Analyze Cache Memory Behavior (Research Note)
Xavier Vera, Josep Llosa, Antonio González 0001, Nerina Bermudo
Euro-Par2
2000 An efficient solver for Cache Miss Equations
abstract
Cache Miss Equations (CME) (S. Ghosh et al., 1997) is a method that accurately describes the cache behavior by means of polyhedra. Even though the computation cost of generating CME is a linear function of the number of references, solving them is a very time consuming task and thus trying to study a whole program may be infeasible. The paper presents effective techniques that exploit some properties of the particular polyhedra generated by CME. Such techniques reduce the complexity of the algorithm to solve CME, which results in a significant speedup when compared with traditional methods. In particular, the proposed approach does not require the computation of the vertices of each polyhedron, which has an exponential complexity.
Nerina Bermudo, Xavier Vera, Antonio González 0001, Josep Llosa
ISPASS4
2000 Two-level hierarchical register file organization for VLIW processors
abstract
High-performance microprocessors are currently designed to exploit the inherent instruction level parallelism (ILP) available in most applications. The techniques used in their design and the aggressive scheduling techniques used to exploit this ILP tend to increase the register requirements of the loops. If more registers than those available in the architecture are required, some actions (such as spill code insertion) have to be applied to reduce this pressure, at the expense of some performance degradation. This degradation could be avoided if a high-capacity register file were included without causing a negative impact on the cycle time of the processor. The authors propose a two-level hierarchical register file organization for VLIW architectures that combines high capacity and low access time. For the configuration proposed in the paper, the new organization achieves a speed-up of 10-14% over a monolithic organization with 64 registers; it is obtained with a 43% (40%) reduction in area (peak power dissipation). Compared to a monolithic file with 32 registers, the speed-up is as much as 38% with just a 14% (4%) increase in area (peak power dissipation).
Javier Zalamea, Josep Llosa, Eduard Ayguadé, Mateo Valero
MICRO2
2000 Improved spill code generation for software pipelined loops
abstract
Software pipelining is a loop scheduling technique that extracts parallelism out of loops by overlapping the execution of several consecutive iterations. Due to the overlapping of iterations, schedules impose high register requirements during their execution. A schedule is valid if it requires at most the number of registers available in the target architecture. If not, its register requirements have to be reduced either by decreasing the iteration overlapping or by spilling registers to memory. In this paper we describe a set of heuristics to increase the quality of register-constrained modulo schedules. The heuristics decide between the two previous alternatives and define criteria for effectively selecting spilling candidates. The heuristics proposed for reducing the register pressure can be applied to any software pipelining technique. The proposals are evaluated using a register-conscious software pipeliner on a workbench composed of a large set of loops from the Perfect Club benchmark and a set of processor configurations. Proposals in this paper are compared against a previous proposal already described in the literature. For one of these processor configurations and the set of loops that do not fit in the available registers (32), a speed-up of 1.68 and a reduction of the memory traffic by a factor of 0.57 are achieved with an affordable increase in compilation time. For all the loops, this represents a speed-up of 1.38 and a reduction of the memory traffic by a factor of 0.7.
Javier Zalamea, Josep Llosa, Eduard Ayguadé, Mateo Valero
PLDI2
1999 Distributed Modulo Scheduling
abstract
Wide-issue ILP machines can be built using the VLIW approach as many of the hardware complexities found in superscalar processors can be transferred to the compiler. However, the scalability of VLIW architectures is still constrained by the size and number of ports of the register file required by a large number of functional units. Organizations composed of clusters of a few functional units and small private register files have been proposed to deal with this problem; an approach highly dependent on scheduling and partitioning strategies. The paper presents DMS, an algorithm that integrates modulo scheduling and code partitioning in a single procedure. Experimental results have shown that the algorithm is effective for configurations up to 8 clusters, or even more when targeting vectorizable loops.
Marcio Merino Fernandes, Josep Llosa, Nigel P. Topham
HPCA2
1999 Impact on Performance of Fused Multiply-Add Units in Aggressive VLIW Architectures
abstract
Loops are the main time consuming part of programs based on floating point computations. The performance of the loops is limited either by recurrences in the computation or by the resources offered by the architecture. Several general-purpose superscalar microprocessors have been implemented with multiply-add fused floating-point units, that reduces the latency of the combined operation and the number of resources used. This paper analyses the influence of these two factors in the instruction-level parallelism exploitable from loops executed on a broad set of future aggressive processor configurations. The estimation of implementation costs (area and cycle time) enables a fair comparison of these configurations in terms of final performance and implementation feasibility. The paper performs technological projection for the next years in order to foresee the possible implementable alternatives. From this study we conclude that multiply-add fused units may have a deep impact in raising the performance of future processor architectures with a reasonable increase in cost.
David López 0001, Josep Llosa, Eduard Ayguadé, Mateo Valero
ICPP2
1998 Resource Widening Versus Replication: Limits and Performance-cost Trade-off
David López 0001, Josep Llosa, Mateo Valero, Eduard Ayguadé
International Conference on Supercomputing2
1998 Widening Resources: A Cost-effective Technique for Aggressive ILP Architectures
abstract
The inherent instruction-level parallelism (ILP) of current applications (specially those based on floating point computations) has driven hardware designers and compilers writers to investigate aggressive techniques for exploiting program parallelism at the lowest level. To execute more operations per cycle, many processors are designed with growing degrees of resource replication (buses and functional units). However the high cost in terms of area and cycle time of this technique precludes the use of high degrees of replication. An alternative to resource replication is resource widening, that has also been used in some recent designs, in which the width of the resources is increased. In this paper we evaluate a broad set of design alternatives that combine both replication and widening. For each alternative we perform an estimation of the ILP limits (including the impact of spill code for several register file configurations) and the cost in terms of area and access time of the register file. We also perform a technological projection for the next 10 years in order to foresee the possible implementable alternatives. From this study we conclude that if the cost is taken into account, the best performance is obtained when combining certain degrees of replication and widening in the hardware resources. The results have been obtained from a large number of inner loops from numerical programs scheduled for VLIW architectures.
David López 0001, Josep Llosa, Mateo Valero, Eduard Ayguadé
MICRO2
1998 Modulo Scheduling with Reduced Register Pressure
abstract
Software pipelining is a scheduling technique that is used by some product compilers in order to expose more instruction level parallelism out of innermost loops. Module scheduling refers to a class of algorithms for software pipelining. Most previous research on module scheduling has focused on reducing the number of cycles between the initiation of consecutive iterations (which is termed II) but has not considered the effect of the register pressure of the produced schedules. The register pressure increases as the instruction level parallelism increases. When the register requirements of a schedule are higher than the available number of registers, the loop must be rescheduled perhaps with a higher II. Therefore, the register pressure has an important impact on the performance of a schedule. This paper presents a novel heuristic module scheduling strategy that tries to generate schedules with the lowest II, and, from all the possible schedules with such II, it tries to select that with the lowest register requirements. The proposed method has been implemented in an experimental compiler and has been tested for the Perfect Club benchmarks. The results show that the proposed method achieves an optimal II for at least 97.5 percent of the loops and its compilation time is comparable to a conventional top-down approach, whereas the register requirements are lower. In addition, the proposed method is compared with some other existing methods. The results indicate that the proposed method performs better than other heuristic methods and almost as well as linear programming methods, which obtain optimal solutions but are impractical for product compilers because their computing cost grows exponentially with the number of operations in the loop body.
Josep Llosa, Mateo Valero, Eduard Ayguadé, Antonio González 0001
IEEE Trans. Computers1
1997 Allocating Lifetimes to Queues in Software Pipelined Architectures
Marcio Merino Fernandes, Josep Llosa, Nigel P. Topham
Euro-Par2
1997 Increasing Memory Bandwidth with Wide Buses: Compiler, Hardware and Performance Trade-Offs
abstract
Article Increasing memory bandwidth with wide buses: compiler, hardware and performance trade-offs Share on Authors: David López Departament d'Arquitectura de Computadors, Universitat Politècnica de Catalunya, Campus Nord, Mòdul D6, Jordi Girona 1-3, 08034 Barcelona, Spain Departament d'Arquitectura de Computadors, Universitat Politècnica de Catalunya, Campus Nord, Mòdul D6, Jordi Girona 1-3, 08034 Barcelona, SpainView Profile , Mateo Valero Departament d'Arquitectura de Computadors, Universitat Politècnica de Catalunya, Campus Nord, Mòdul D6, Jordi Girona 1-3, 08034 Barcelona, Spain Departament d'Arquitectura de Computadors, Universitat Politècnica de Catalunya, Campus Nord, Mòdul D6, Jordi Girona 1-3, 08034 Barcelona, SpainView Profile , Josep Llosa Departament d'Arquitectura de Computadors, Universitat Politècnica de Catalunya, Campus Nord, Mòdul D6, Jordi Girona 1-3, 08034 Barcelona, Spain Departament d'Arquitectura de Computadors, Universitat Politècnica de Catalunya, Campus Nord, Mòdul D6, Jordi Girona 1-3, 08034 Barcelona, SpainView Profile , Eduard Ayguadé Departament d'Arquitectura de Computadors, Universitat Politècnica de Catalunya, Campus Nord, Mòdul D6, Jordi Girona 1-3, 08034 Barcelona, Spain Departament d'Arquitectura de Computadors, Universitat Politècnica de Catalunya, Campus Nord, Mòdul D6, Jordi Girona 1-3, 08034 Barcelona, SpainView Profile Authors Info & Claims ICS '97: Proceedings of the 11th international conference on SupercomputingJuly 1997 Pages 12–19https://doi.org/10.1145/263580.263585Online:11 July 1997Publication History 10citation232DownloadsMetricsTotal Citations10Total Downloads232Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
David López 0001, Mateo Valero, Josep Llosa, Eduard Ayguadé
International Conference on Supercomputing3
1996 Heuristics for Register-Constrained Software Pipelining
abstract
Software Pipelining is a loop scheduling technique that extracts parallelism from loops by overlapping the execution of several consecutive iterations. There has been a significant effort to produce throughput-optimal schedules under resource constraints, and more recently to produce throughput-optimal schedules with minimum register requirements. Unfortunately even a throughput-optimal schedule with minimum register requirements is useless if it requires more registers than those available in the target machine. This paper evaluates several techniques for producing register-constrained modulo schedules: increasing the initiation interval (II) and adding spill code. We show that, in general, increasing the II performs poorly and might not converge for some loops. The paper also presents an iterative spilling mechanism that can be applied to any software pipelining technique and proposes several heuristics in order to speed-up the scheduling process.
Josep Llosa, Mateo Valero, Eduard Ayguadé
MICRO1
1995 Non-Consistent Dual Register Files to Reduce Register Pressure
abstract
The continuous grow on instruction level parallelism offered by microprocessors requires a large register file and a large number of ports to access it. This paper presents the non-consistent dual register file, an alternative implementation and management of the register file. Non-consistent dual register files support the bandwidth demands and the high register requirements, penalizing neither access time nor implementation cost. The proposal is evaluated for software pipelined loops and compared against a unified register file. Empirical results show improvements on performance and a noticeable reduction of the density of memory traffic due to a reduction of the spill code. The spill code can in general increase the minimum initiation interval and decrease loop performance. Additional improvements can be obtained when the operations are scheduled having in mind the register file organization proposed.>
Josep Llosa, Mateo Valero, Eduard Ayguadé
HPCA1
1995 Hypernode reduction modulo scheduling
abstract
Software pipelining is a loop scheduling technique that extracts parallelism from loops by overlapping the execution of several consecutive iterations. Most prior scheduling research has focused on achieving minimum execution time, without regarding register requirements. Most strategies tend to stretch operand lifetimes because they schedule some operations too early or too late. The paper presents a novel strategy that simultaneously schedules some operations late and other operations early, minimizing all the stretchable dependencies and therefore reducing the registers required by the loop. The key of this strategy is a pre-ordering that selects the order in which the operations will be scheduled. The results show that the method described in this paper performs better than other heuristic methods and almost as well as a linear programming method but requiring much less time to produce the schedules.
Josep Llosa, Mateo Valero, Eduard Ayguadé, Antonio González 0001
MICRO1