EDBT 2026 Demo / reviewers in the wild / expert
Raúl Martínez
dblp:78/980
· DBLP profile ↗
16ranked-venue papers
9as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 8 first-authorSoftware engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Processor architecture and microarchitecture · 67% Electronic design automation · 19% Interconnection networks and networks-on-chip · 11% | |
| Computer networks
1 paper |
Internet architecture and protocols · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Runtime systems and virtual machines · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
fused multiply-add |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Electronic design automation
hardware/software co-design |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Processor architecture and microarchitecture
instruction set architecture |
0.2 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 1 | 2009 | Boosting single-thread performance in multi-core systems through fine-grain multi-threading · ISCA 2009 |
Processor architecture and microarchitecture › multithreading
fine-grain multithreading |
0.1 | 1 | 2009 | Boosting single-thread performance in multi-core systems through fine-grain multi-threading · ISCA 2009 |
Processor architecture and microarchitecture
multithreading |
0.1 | 1 | 2009 | Boosting single-thread performance in multi-core systems through fine-grain multi-threading · ISCA 2009 |
Internet architecture and protocols
quality of service |
0.1 | 1 | 2008 | A Framework to Provide Quality of Service over Advanced Switching · IEEE Trans. Parallel Distributed Syst. 2008 |
Interconnection networks and networks-on-chip › switching
advanced switching |
0.1 | 1 | 2008 | A Framework to Provide Quality of Service over Advanced Switching · IEEE Trans. Parallel Distributed Syst. 2008 |
Runtime systems and virtual machines › binary translation
dynamic binary translation |
0.1 | 1 | 2014 | Speculative hardware/software co-designed floating-point multiply-add fusion · ASPLOS 2014 |
Parallel and multicore computing
thread-level parallelism |
0.0 | 1 | 2009 | Boosting single-thread performance in multi-core systems through fine-grain multi-threading · ISCA 2009 |
Interconnection networks and networks-on-chip › high-speed interconnect
PCIe interconnect |
0.0 | 1 | 2008 | A Framework to Provide Quality of Service over Advanced Switching · IEEE Trans. Parallel Distributed Syst. 2008 |
Methods — techniques the papers use, named apart from their topics
speculative instruction-fusion optimization · 0.4cycle-accurate simulation · 0.4simulation · 0.2amdahl's law · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Energy efficient torus networks with on/off links
Francisco J. Andujar, Salvador Coll, Marina Alonso, Juan-Miguel Martinez-Rubio, Pedro López 0001, José L. Sánchez 0002, Francisco J. Alfaro, Raúl Martínez |
J. Parallel Distributed Comput. | 8 |
| 2016 | Facing prefetching challenges in distributed shared memories for CMPs
Martí Torrents, Raúl Martínez, Carlos Molina 0004 |
J. Supercomput. | 2 |
| 2014 | Speculative hardware/software co-designed floating-point multiply-add fusionabstractA Fused Multiply-Add (FMA) instruction is currently available in many general-purpose processors. It increases performance by reducing latency of dependent operations and increases precision by computing the result as an indivisible operation with no intermediate rounding. However, since the arithmetic behavior of a single-rounding FMA operation is different than independent FP multiply followed by FP add instructions, some algorithms require significant revalidation and rewriting efforts to work as expected when they are compiled to operate with FMA--a cost that developers may not be willing to pay. Because of that, abundant legacy applications are not able to utilize FMA instructions. In this paper we propose a novel HW/SW collaborative technique that is able to efficiently execute workloads with increased utilization of FMA, by adding the option to get the same numerical result as separate FP multiply and FP add pairs. In particular, we extended the host ISA of a HW/SW co-designed processor with a new Combined Multiply-Add (CMA) instruction that performs an FMA operation with an intermediate rounding. This new instruction is used by a transparent dynamic translation software layer that uses a speculative instruction-fusion optimization to transform FP multiply and FP add sequences into CMA instructions. The FMA unit has been slightly modified to support both single-rounding and double-rounding fused instructions without increasing their latency and to provide a conservative fall-back path in case of mispeculation. Evaluation on a cycle-accurate timing simulator showed that CMA improved SPECfp performance by 6.3% and reduced executed instructions by 4.7%. Marc Lupon, Enric Gibert, Grigorios Magklis, Sridhar Samudrala, Raúl Martínez, Kyriakos Stavrou, David R. Ditzel |
ASPLOS | 5 |
| 2012 | Hardware implementation study of several new egress link scheduling algorithms
Raúl Martínez, José M. Claver, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Parallel Distributed Comput. | 1 |
| 2009 | Anaphase: A Fine-Grain Thread Decomposition Scheme for Speculative MultithreadingabstractIndustry is moving towards multi-core designs as we have hit the memory and power walls. Multi-core designs are very effective to exploit thread-level parallelism (TLP) but do not provide benefits when executing serial code (applications with low TLP, serial parts of a parallel application and legacy code). In this paper we propose Anaphase, a novel approach for speculative multithreading to improve single-thread performance in a multi-core design. The proposed technique is based on a graph partitioning technique which performs a decomposition of applications into speculative threads at instruction granularity. Moreover, the proposed technique leverages communications and pre-computation slices to deal with inter-thread dependences. Results presented in this paper show that this approach improves single-thread performance by 32% on average and up to 2.15x for some selected applications of the Spec2006 suite. In addition, the proposed technique outperforms by 21% on average schemes in which thread decomposition is performed at a coarser granularity. Carlos Madriles, Pedro López 0001, Josep M. Codina, Enric Gibert, Fernando Latorre, Alejandro Martínez, Raúl Martínez, Antonio González 0001 |
PACT | 7 |
| 2009 | Hardware Implementation Study of the SCFQ-CA and DRR-CA Scheduling Algorithms
Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002, José M. Claver |
Euro-Par | 1 |
| 2009 | Hardware Implementation Study of the Deficit Table Egress Link Scheduling AlgorithmabstractThe provision of quality of service (QoS) in computing and communication environments has increasingly focused the attention from academia and industry during the last decades. Some of the current interconnection technologies include hardware support that, adequately used, allows to offer QoS guarantees to the applications. The egress link scheduling algorithm is a key part of that support. Apart from providing a good performance in terms of, for example, good end-to-end delay (also called latency) and fair bandwidth allocation, an ideal scheduling algorithm implemented in a high-performance network with QoS support should satisfy other important property which is to have a low computational and implementation complexity. In this paper, we propose a specific implementation of the DTable scheduling algorithm and show estimates about its complexity in terms of silicon area and computation delay. In order to obtain these estimates, we have performed our own hardware implementation using the Handel-C language and employed the DK design suite tool from Celoxica. Raúl Martínez, José M. Claver, Francisco J. Alfaro, José L. Sánchez 0002 |
ICPP | 1 |
| 2009 | Boosting single-thread performance in multi-core systems through fine-grain multi-threadingabstractIndustry has shifted towards multi-core designs as we have hit the memory and power walls. However, single thread performance remains of paramount importance since some applications have limited thread-level parallelism (TLP), and even a small part with limited TLP impose important constraints to the global performance, as explained by Amdahl's law. Carlos Madriles, Pedro López 0001, Josep M. Codina, Enric Gibert, Fernando Latorre, Alejandro Martínez, Raúl Martínez, Antonio González 0001 |
ISCA | 7 |
| 2008 | A Framework to Provide Quality of Service over Advanced SwitchingabstractAdvanced Switching (AS) is a network technology that expands the capabilities of PCI-Express adding new features like peer-to-peer communication. Together, PCI Express and AS have the potential for building the next generation interconnects. Furthermore, the provision of Quality of Service (QoS) in computing and communication environments is currently the focus of much discussion and research in industry and academia. In this paper we propose a framework to provide QoS based on bandwidth, latency, and jitter over AS employing the mechanisms provided by AS. We also present several implementations for the output scheduling mechanism. Finally, we evaluate our proposals by simulation, comparing the performance of the schedulers that we propose and their implementation complexity. Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2007 | Comparing the latency performance of the DTable and DRR schedulersabstractA key component for networks with quality of service (QoS) support is the egress link scheduling algorithm. An ideal scheduling algorithm implemented in a high performance network with QoS support should satisfy two main properties: good end-to-end delay and implementation simplicity. The deficit round robin (DRR) algorithm is known to have a very little implementation complexity. However, depending on the situation, its latency performance can be very bad. On the other hand, table-based schedulers try to offer a simple implementation and good latency bounds. Some of the latest proposals of network technologies, like advanced switching and infiniband, include in their specifications one of these schedulers. However, these table-based schedulers do not work properly with variable packet sizes and face the problem of bounding the bandwidth and latency assignments. We have proposed a new table-based scheduler, which we have called deficit table (DTable) scheduler, that works properly with variable packet sizes. Moreover, we have proposed a methodology to configure this table-based scheduler to decouple the bounding of bandwidth and latency assignments. In this paper, we review these proposals and present simulation results that show that the DTable scheduler is able to provide a better latency performance than the DRR scheduler, with only a slightly higher implementation and computational complexity. Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
IPDPS | 1 |
| 2007 | A low-cost strategy to provide full QoS support in Advanced Switching networks
Alejandro Martínez, Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Syst. Archit. | 2 |
| 2006 | Improving the Flexibility of the Deficit Table Scheduler
Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
HiPC | 1 |
| 2006 | A Statistical Approach to Traffic Management in Source Routed Loss-Less Networks
Thomas Sødring, Raúl Martínez, Geir Horn |
HPCC | 2 |
| 2006 | Decoupling the Bandwidth and Latency Bounding for Table-based SchedulersabstractThe provision of quality of service (QoS) in computing and communication environments is currently the focus of much discussion and research in industry and academia. A key component for networks with QoS support is the output scheduling algorithm. Some of the latest network technology proposals define scheduling algorithms that use an arbitration table to select the next packet to be transmitted. These table-based schedulers are simple to implement and can offer good latency performance. However, the versions proposed until now do not work properly with variable packet sizes. Moreover, they face the problem of bounding the bandwidth and latency assignments. In this paper, we propose a new table-based scheduler, which we call deficit table (DTable), that works properly with variable packet sizes. We also propose a methodology to decouple the bandwidth and latency assignments Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
ICPP | 1 |
| 2006 | Studying Several Proposals for the Adaptation of the DTable Scheduler to Advanced Switching
Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
ISPA | 1 |
| 2006 | Implementing the Advanced Switching Minimum Bandwidth Egress Link SchedulerabstractAdvanced switching (AS) is a new fabric-interconnect technology that further enhances the capabilities of PCI Express, which is the next PCI generation. On the other hand, the provision of quality of service (QoS) in computing and communication environments is currently the focus of much discussion and research in industry and academia. One of the mechanisms that AS provides to support QoS is the minimum bandwidth egress link scheduler, or just MinBW scheduler. In this paper, we propose several implementations of the MinBW scheduler and compare their performance by simulation. These implementations fulfill all the properties that an AS MinBW scheduler must have, including the interaction with the AS link layer flow control Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
NCA | 1 |