VLDB 2026 Research / reviewers in the wild / expert
Javier Zalamea
dblp:34/1658
· DBLP profile ↗
4ranked-venue papers
4as first author
0since 2021 · last 2004
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Processor architecture and microarchitecture · 93% Energy-efficient computing · 7% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
instruction scheduling |
0.1 | 2 | 2004 | Register Constrained Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2004 Improved spill code generation for software pipelined loops · PLDI 2000 |
Compilers and program optimization › instruction scheduling
software pipelining |
0.1 | 2 | 2004 | Register Constrained Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2004 Improved spill code generation for software pipelined loops · PLDI 2000 |
Compilers and program optimization
register allocation |
0.1 | 2 | 2004 | Register Constrained Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2004 Modulo scheduling with integrated register spilling for clustered VLIW architectures · MICRO 2001 |
Compilers and program optimization › instruction scheduling › software pipelining
modulo scheduling |
0.0 | 1 | 2004 | Register Constrained Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2004 |
Compilers and program optimization › register allocation
register pressure reduction |
0.0 | 1 | 2004 | Register Constrained Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2004 |
Processor architecture and microarchitecture › instruction-level parallelism
VLIW |
0.0 | 2 | 2001 | Modulo scheduling with integrated register spilling for clustered VLIW architectures · MICRO 2001 Two-level hierarchical register file organization for VLIW processors · MICRO 2000 |
Processor architecture and microarchitecture › instruction-level parallelism › VLIW
clustered VLIW |
0.0 | 1 | 2001 | Modulo scheduling with integrated register spilling for clustered VLIW architectures · MICRO 2001 |
Processor architecture and microarchitecture
instruction scheduling |
0.0 | 1 | 2001 | Modulo scheduling with integrated register spilling for clustered VLIW architectures · MICRO 2001 |
Processor architecture and microarchitecture › instruction scheduling
software pipelining |
0.0 | 1 | 2001 | Modulo scheduling with integrated register spilling for clustered VLIW architectures · MICRO 2001 |
Processor architecture and microarchitecture › register file
hierarchical register file |
0.0 | 1 | 2000 | Two-level hierarchical register file organization for VLIW processors · MICRO 2000 |
Processor architecture and microarchitecture › register file
register file organization |
0.0 | 1 | 2000 | Two-level hierarchical register file organization for VLIW processors · MICRO 2000 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 2 | 2000 | Improved spill code generation for software pipelined loops · PLDI 2000 Two-level hierarchical register file organization for VLIW processors · MICRO 2000 |
Compilers and program optimization › register allocation
register spilling |
0.0 | 1 | 2001 | Modulo scheduling with integrated register spilling for clustered VLIW architectures · MICRO 2001 |
Energy-efficient computing
power management |
0.0 | 1 | 2000 | Two-level hierarchical register file organization for VLIW processors · MICRO 2000 |
Energy-efficient computing › memory energy efficiency
register file energy reduction |
0.0 | 1 | 2000 | Two-level hierarchical register file organization for VLIW processors · MICRO 2000 |
Methods — techniques the papers use, named apart from their topics
modulo scheduling · 0.1backtracking · 0.1heuristic spill candidate selection · 0.1spilling · 0.0heuristics · 0.0spill code insertion · 0.0instruction scheduling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2004 | Register Constrained Modulo SchedulingabstractSoftware pipelining is an instruction scheduling technique that exploits the instruction level parallelism (ILP) available in loops by overlapping operations from various successive loop iterations. The main drawback of aggressive software pipelining techniques is their high register requirements. If the requirements exceed the number of registers available in the target architecture, some steps need to be applied to reduce the register pressure (incurring some performance degradation): reduce iteration overlapping or spilling some lifetimes to memory. In the first part, we propose a set of heuristics to improve the spilling process and to better decide between adding spill code or directly decreasing the execution rate of iterations. The experimental evaluation, over a large number of representative loops and for a processor configuration, reports an increase in performance by a factor of 1.29 and a reduction of memory traffic by a factor of 1.36. In the second part, we analyze the use of backtracking and propose a novel approach for simultaneous instruction scheduling and register spilling in modulo scheduling: MIPS (modulo scheduling with integrated register spilling). The experimental evaluation reports an increase in performance by a factor of 1.46 and a reduction of the memory traffic by a factor of 1.66 (or an additional 1.13 and 1.22 with regard to the proposal in the first part). These improvements are achieved at the expense of a reasonable increase in the compilation time. Javier Zalamea, Josep Llosa, Eduard Ayguadé, Mateo Valero |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2001 | Modulo scheduling with integrated register spilling for clustered VLIW architecturesabstractClustering is a technique to decentralize the design of future wide issue VLIW cores and enable them to meet the technology constraints in terms of cycle time, area and power dissipation. In a clustered design, registers and functional units are grouped in clusters so that new instructions are needed to move data between them. New aggressive instruction scheduling techniques are required to minimize the negative effect of resource clustering and delays in moving data around. In this paper we present a novel software pipelining technique that performs instruction scheduling with reduced register requirements, register allocation, register spilling and inter-cluster communication in a single step. The algorithm uses limited backtracking to reconsider previously taken decisions. This backtracking provides the algorithm with additional possibilities for obtaining high throughput schedules with low spill code requirements for clustered architectures. We show that the proposed approach outperforms previously proposed techniques and that it is very scalable independently of the number of clusters, the number of communication buses and communication latency. The paper also includes an exploration of some parameters in the design of future clustered VLIW cores. Javier Zalamea, Josep Llosa, Eduard Ayguadé, Mateo Valero |
MICRO | 1 |
| 2000 | Two-level hierarchical register file organization for VLIW processorsabstractHigh-performance microprocessors are currently designed to exploit the inherent instruction level parallelism (ILP) available in most applications. The techniques used in their design and the aggressive scheduling techniques used to exploit this ILP tend to increase the register requirements of the loops. If more registers than those available in the architecture are required, some actions (such as spill code insertion) have to be applied to reduce this pressure, at the expense of some performance degradation. This degradation could be avoided if a high-capacity register file were included without causing a negative impact on the cycle time of the processor. The authors propose a two-level hierarchical register file organization for VLIW architectures that combines high capacity and low access time. For the configuration proposed in the paper, the new organization achieves a speed-up of 10-14% over a monolithic organization with 64 registers; it is obtained with a 43% (40%) reduction in area (peak power dissipation). Compared to a monolithic file with 32 registers, the speed-up is as much as 38% with just a 14% (4%) increase in area (peak power dissipation). Javier Zalamea, Josep Llosa, Eduard Ayguadé, Mateo Valero |
MICRO | 1 |
| 2000 | Improved spill code generation for software pipelined loopsabstractSoftware pipelining is a loop scheduling technique that extracts parallelism out of loops by overlapping the execution of several consecutive iterations. Due to the overlapping of iterations, schedules impose high register requirements during their execution. A schedule is valid if it requires at most the number of registers available in the target architecture. If not, its register requirements have to be reduced either by decreasing the iteration overlapping or by spilling registers to memory. In this paper we describe a set of heuristics to increase the quality of register-constrained modulo schedules. The heuristics decide between the two previous alternatives and define criteria for effectively selecting spilling candidates. The heuristics proposed for reducing the register pressure can be applied to any software pipelining technique. The proposals are evaluated using a register-conscious software pipeliner on a workbench composed of a large set of loops from the Perfect Club benchmark and a set of processor configurations. Proposals in this paper are compared against a previous proposal already described in the literature. For one of these processor configurations and the set of loops that do not fit in the available registers (32), a speed-up of 1.68 and a reduction of the memory traffic by a factor of 0.57 are achieved with an affordable increase in compilation time. For all the loops, this represents a speed-up of 1.38 and a reduction of the memory traffic by a factor of 0.7. Javier Zalamea, Josep Llosa, Eduard Ayguadé, Mateo Valero |
PLDI | 1 |