Alex Aletà

dblp:55/2385 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
0since 2021 · last 2009
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Processor architecture and microarchitecture · 78% Reconfigurable computing and FPGAs · 19% Hardware reliability and fault tolerance · 3%
Software engineering, system software, and programming languages
5 papers
Compilers and program optimization · 100%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › clustered architecture
clustered microarchitecture
0.242009
AGAMOS: A Graph-Based Approach to Modulo Scheduling for Clustered Microarchitectures · IEEE Trans. Computers 2009
Removing communications in clustered microarchitectures through instruction replication · ACM Trans. Archit. Code Optim. 2004
Instruction Replication for Clustered Microarchitectures · MICRO 2003
Processor architecture and microarchitecture
instruction scheduling
0.122009
AGAMOS: A Graph-Based Approach to Modulo Scheduling for Clustered Microarchitectures · IEEE Trans. Computers 2009
Graph-partitioning based instruction scheduling for clustered processors · MICRO 2001
Compilers and program optimization
instruction scheduling
0.132004
Removing communications in clustered microarchitectures through instruction replication · ACM Trans. Archit. Code Optim. 2004
Instruction Replication for Clustered Microarchitectures · MICRO 2003
Graph-partitioning based instruction scheduling for clustered processors · MICRO 2001
Compilers and program optimization › instruction scheduling › software pipelining
modulo scheduling
0.122005
Demystifying on-the-fly spill code · PLDI 2005
Removing communications in clustered microarchitectures through instruction replication · ACM Trans. Archit. Code Optim. 2004
Reconfigurable computing and FPGAs
modulo scheduling
0.112009
AGAMOS: A Graph-Based Approach to Modulo Scheduling for Clustered Microarchitectures · IEEE Trans. Computers 2009
Compilers and program optimization
register allocation
0.112005
Demystifying on-the-fly spill code · PLDI 2005
Compilers and program optimization › register allocation
register pressure reduction
0.112005
Demystifying on-the-fly spill code · PLDI 2005
Processor architecture and microarchitecture
instruction-level parallelism
0.022005
Demystifying on-the-fly spill code · PLDI 2005
Removing communications in clustered microarchitectures through instruction replication · ACM Trans. Archit. Code Optim. 2004
Compilers and program optimization › loop transformation
loop scheduling
0.012009
AGAMOS: A Graph-Based Approach to Modulo Scheduling for Clustered Microarchitectures · IEEE Trans. Computers 2009
Processor architecture and microarchitecture › instruction scheduling
software pipelining
0.012005
Demystifying on-the-fly spill code · PLDI 2005
Hardware reliability and fault tolerance › soft errors
instruction duplication
0.012004
Removing communications in clustered microarchitectures through instruction replication · ACM Trans. Archit. Code Optim. 2004

Methods — techniques the papers use, named apart from their topics

instruction replication · 0.4pseudoschedules · 0.2multilevel graph partitioning · 0.2on-the-fly spilling · 0.1a posteriori spilling · 0.1modulo scheduling · 0.1spill code generation · 0.1register allocation · 0.1graph partitioning · 0.1
YearPublicationVenuePosition
2009 AGAMOS: A Graph-Based Approach to Modulo Scheduling for Clustered Microarchitectures
abstract
This paper presents AGAMOS, a technique to modulo schedule loops on clustered microarchitectures. The proposed scheme uses a multilevel graph partitioning strategy to distribute the workload among clusters and reduces the number of intercluster communications at the same time. Partitioning is guided by approximate schedules (i.e., pseudoschedules), which take into account all of the constraints that influence the final schedule. To further reduce the number of intercluster communications, heuristics for instruction replication are included. The proposed scheme is evaluated using the SPECfp95 programs. The described scheme outperforms a state-of-the-art scheduler for all programs and different cluster configurations. For some configurations, the speedup obtained when using this new scheme is greater than 40 percent, and for selected programs, performance can be more than doubled.
Alex Aletà, Josep M. Codina, F. Jesús Sánchez, Antonio González 0001, David R. Kaeli
IEEE Trans. Computers1
2007 Heterogeneous Clustered VLIW Microarchitectures
abstract
Increasing performance, while at the same time reducing power consumption, is a major design tradeoff in current microprocessors. In this paper, we investigate the potential of using a heterogeneous clustered VLIW microarchitecture. In the proposed microarchitecture, each cluster, the interconnection network and the supporting memory hierarchy can run at different frequencies and voltages. Some of the clusters can then be configured to be performance-oriented and run at high frequency, while the other clusters can be configured to be low-power-oriented and run at lower frequencies, thus reducing overall consumption. For this heterogeneous design to be effective, we need to select the most suitable frequencies and voltages for each component. We propose a scheme to choose these parameters based on a model that estimates the energy consumption and the execution time of floating-point codes at compile time. Finally, we present a modulo scheduling technique based on graph partitioning that exploits the opportunities presented on heterogeneous clustered microarchitectures. Results show that the Energy-Delay product (ED2) can be significantly reduced by 15% on average for a microarchitecture with 4-clusters and by as much as 35% for selected programs
Alex Aletà, Josep M. Codina, Antonio González 0001, David R. Kaeli
CGO1
2005 Demystifying on-the-fly spill code
abstract
Modulo scheduling is an effective code generation technique that exploits the parallelism in program loops by overlapping iterations. One drawback of this optimization is that register requirements increase significantly because values across different loop iterations can be live concurrently. One possible solution to reduce register pressure is to insert spill code to release registers. Spill code stores values to memory between the producer and consumer instructions.Spilling heuristics can be divided into two classes: 1) a posteriori approaches (spill code is inserted after scheduling the loop) or 2) on-the-fly approaches (spill code is inserted during loop scheduling). Recent studies have reported obtaining better results for spilling on-the-fly. In this work, we study both approaches and propose two new techniques, one for each approach. Our new algorithms try to address the drawbacks observed in previous proposals. We show that the new algorithms outperform previous techniques and, at the same time, reduce compilation time. We also show that, much to our surprise, a posteriori spilling can be in fact slitghtly more effective than on-the-fly spilling.
Alex Aletà, Josep M. Codina, Antonio González 0001, David R. Kaeli
PLDI1
2004 Removing communications in clustered microarchitectures through instruction replication
abstract
The need to communicate values between clusters can result in a significant performance loss for clustered microarchitectures. In this work, we describe an optimization technique that removes communications by selectively replicating an appropriate set of instructions. Instruction replication is done carefully because it might degrade performance due to the increased contention it can place on processor resources. The proposed scheme is built on top of a previously proposed state-of-the-art modulo-scheduling algorithm. Though this algorithm has been proved to be very effective at reducing communications, results show that the number of communications can be further decreased by around one-third through replication, which results in a significant speedup. IPC is increased by 25% on average for a four-cluster microarchitecture and by as much as 70% for selected programs. We also show that replicating appropriate sets of instructions is more effective than doubling the intercluster connection network bandwidth.
Alex Aletà, Josep M. Codina, Antonio González 0001, David R. Kaeli
ACM Trans. Archit. Code Optim.1
2003 Instruction Replication for Clustered Microarchitectures
abstract
This work presents a new compilation technique that uses instruction replication in order to reduce the number of communications executed on a clustered microarchitecture. For such architectures, the need to communicate values between clusters can result in a significant performance loss. Inter-cluster communications can be reduced by selectively replicating an appropriate set of instructions. However, instruction replication must be done carefully since it may also degrade performance due to the increased contention it can place on processor resources. The proposed scheme is built on top of a previously proposed state-of-the-art modulo scheduling algorithm that effectively reduces communications. Results show that the number of communications can decrease using replication, which results in significant speed-ups. IPC is increased by 25% on average for a 4-cluster microarchitecture and by as mush as 70% for selected programs.
Alex Aletà, Josep M. Codina, Antonio González 0001, David R. Kaeli
MICRO1
2001 Graph-partitioning based instruction scheduling for clustered processors
abstract
This paper presents a novel scheme to schedule loops for clustered microarchitectures. The scheme is based on a preliminary cluster assignment phase implemented through graph partitioning techniques followed by a scheduling phase that integrates register allocation and spill code generation. The graph partitioning scheme is shown to be very effective due to its global view of the whole code while the partition is generated. Results show a significant speedup when compared with previously proposed techniques. For some processor configuration the average speedup for the SPECfp95 is 23% with respect to the published scheme with the best performance. Besides, the proposed scheme is much faster (between 2-7 times, depending on the configuration).
Alex Aletà, Josep M. Codina, F. Jesús Sánchez, Antonio González 0001
MICRO1