EDBT 2026 Demo / reviewers in the wild / expert
Antonio Robles
dblp:230/9603 · also Antonio Robles Martinez
· DBLP profile ↗
41ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-0887-3258ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 1 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Memory systems · 60% Interconnection networks and networks-on-chip · 17% Processor architecture and microarchitecture · 14% |
Topics — the 24 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
cache coherence |
1.4 | 7 | 2018 | TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage Prediction · IEEE Trans. Parallel Distributed Syst. 2018 TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs · IEEE Trans. Parallel Distributed Syst. 2017 Efficient TLB-Based Detection of Private Pages in Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2016 |
Processor architecture and microarchitecture
chip multiprocessor |
0.5 | 2 | 2017 | TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs · IEEE Trans. Parallel Distributed Syst. 2017 Efficient TLB-Based Detection of Private Pages in Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2016 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.5 | 3 | 2018 | TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage Prediction · IEEE Trans. Parallel Distributed Syst. 2018 TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs · IEEE Trans. Parallel Distributed Syst. 2017 Efficient TLB-Based Detection of Private Pages in Chip Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2016 |
Memory systems › cache coherence
directory-based coherence |
0.3 | 2 | 2013 | Increasing the Effectiveness of Directory Caches by Avoiding the Tracking of Noncoherent Memory Blocks · IEEE Trans. Computers 2013 Extending Magny-Cours Cache Coherence · IEEE Trans. Computers 2012 |
Memory systems › cache coherence
data classification |
0.3 | 1 | 2017 | TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs · IEEE Trans. Parallel Distributed Syst. 2017 |
Memory systems › virtual memory management
TLB hierarchy |
0.3 | 1 | 2017 | TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBs · IEEE Trans. Parallel Distributed Syst. 2017 |
Interconnection networks and networks-on-chip › routing algorithms
fault-tolerant routing |
0.2 | 2 | 2012 | A Survey and Evaluation of Topology-Agnostic Deterministic Routing Algorithms · IEEE Trans. Parallel Distributed Syst. 2012 A Routing Methodology for Achieving Fault Tolerance in Direct Networks · IEEE Trans. Computers 2006 |
Processor architecture and microarchitecture
multicore design |
0.1 | 2 | 2018 | TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage Prediction · IEEE Trans. Parallel Distributed Syst. 2018 Increasing the Effectiveness of Directory Caches by Avoiding the Tracking of Noncoherent Memory Blocks · IEEE Trans. Computers 2013 |
Interconnection networks and networks-on-chip
congestion control |
0.1 | 1 | 2012 | Progressive Congestion Management Based on Packet Marking and Validation Techniques · IEEE Trans. Computers 2012 |
Interconnection networks and networks-on-chip › switching network
multistage interconnection network |
0.1 | 1 | 2012 | Progressive Congestion Management Based on Packet Marking and Validation Techniques · IEEE Trans. Computers 2012 |
Interconnection networks and networks-on-chip
routing algorithms |
0.1 | 1 | 2012 | A Survey and Evaluation of Topology-Agnostic Deterministic Routing Algorithms · IEEE Trans. Parallel Distributed Syst. 2012 |
Memory systems › cache coherence
directory cache |
0.1 | 1 | 2011 | Increasing the effectiveness of directory caches by deactivating coherence for private memory blocks · ISCA 2011 |
Distributed systems › mutual exclusion
starvation prevention |
0.1 | 1 | 2011 | Efficient and Scalable Starvation Prevention Mechanism for Token Coherence · IEEE Trans. Parallel Distributed Syst. 2011 |
Memory systems › cache coherence › cache coherence protocol
token coherence |
0.1 | 1 | 2011 | Efficient and Scalable Starvation Prevention Mechanism for Token Coherence · IEEE Trans. Parallel Distributed Syst. 2011 |
Performance modeling and evaluation › parallel performance evaluation
multicore scalability |
0.1 | 1 | 2018 | TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage Prediction · IEEE Trans. Parallel Distributed Syst. 2018 |
Interconnection networks and networks-on-chip › routing algorithms
adaptive routing |
0.1 | 1 | 2006 | A Routing Methodology for Achieving Fault Tolerance in Direct Networks · IEEE Trans. Computers 2006 |
Cloud and datacenter computing › resource management › shared resource management
deadlock avoidance |
0.0 | 1 | 2004 | An Effective Methodology to Improve the Performance of the Up*/Down* Routing Algorithm · IEEE Trans. Parallel Distributed Syst. 2004 |
Interconnection networks and networks-on-chip › network topology › static interconnection networks
irregular topology |
0.0 | 1 | 2004 | An Effective Methodology to Improve the Performance of the Up*/Down* Routing Algorithm · IEEE Trans. Parallel Distributed Syst. 2004 |
Interconnection networks and networks-on-chip
network topology |
0.0 | 1 | 2004 | An Effective Methodology to Improve the Performance of the Up*/Down* Routing Algorithm · IEEE Trans. Parallel Distributed Syst. 2004 |
Electronic design automation › physical design
routing |
0.0 | 1 | 2004 | An Effective Methodology to Improve the Performance of the Up*/Down* Routing Algorithm · IEEE Trans. Parallel Distributed Syst. 2004 |
Interconnection networks and networks-on-chip
cluster interconnect |
0.0 | 1 | 2012 | A Survey and Evaluation of Topology-Agnostic Deterministic Routing Algorithms · IEEE Trans. Parallel Distributed Syst. 2012 |
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor |
0.0 | 1 | 2011 | Increasing the effectiveness of directory caches by deactivating coherence for private memory blocks · ISCA 2011 |
Distributed systems
fault tolerance |
0.0 | 1 | 2006 | A Routing Methodology for Achieving Fault Tolerance in Direct Networks · IEEE Trans. Computers 2006 |
Parallel and multicore computing › parallel architecture
massively parallel processor |
0.0 | 1 | 2006 | A Routing Methodology for Achieving Fault Tolerance in Direct Networks · IEEE Trans. Computers 2006 |
Methods — techniques the papers use, named apart from their topics
simulation · 1.0token-based classification · 0.3cycle-accurate simulation · 0.3unicast messaging · 0.3operating system assisted coherence deactivation · 0.2survey · 0.1packet validation · 0.1explicit congestion notification · 0.1classification · 0.1checkpoint/restart · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Improving the Robustness of Redundant Execution with Register File RandomizationabstractStaggered Redundant execution (SRE) is a fault-tolerance mechanism that has been widely deployed in the context of safety-critical applications. SRE not only protects the system in the presence of faults but also helps relaxing safety requirements of individual elements. However, in this paper, we show that SRE does not effectively protect the system against a wide range of faults and thus, new mechanisms to increase the diversity of homogeneous cores are needed. In this paper, we propose Register File Randomization (RFR), a low-cost diversity mechanism that significantly increases the robustness of homogeneous multicores in front of common-cause faults (CCFs) and register file wearout. Our results show that RFR completely removes the failure rate for register file CCFs for certain workloads and reduces by a factor of 5X the impact of stress related register file aging for the workloads analysed. Our implementation requires less than 50 RTL lines of code and the area (FPGA logic) overhead of RFR is less than 0.2% of a 64-bit RISC-V core FPGA implementation. Ilya Tuzov, Pablo Andreu, Laura Medina, Tomás Picornell, Antonio Robles, Pedro López 0001, José Flich, Carles Hernández 0001 |
ICCAD | 5 |
| 2018 | TokenTLB+CUP: A Token-Based Page Classification with Cooperative Usage PredictionabstractDiscerning the private or shared condition of the data accessed by the applications is an increasingly decisive approach to achieving efficiency and scalability in multiand many-core systems. Since most memory accesses in both sequential and parallel applications are either private (accessed only by one core) or read-only (not written) data, devoting the full cost of coherence to every memory access results in sub-optimal performance and limits the scalability and efficiency of the multiprocessor. This paper introduces TokenTLB, a TLB-based page classification approach based on exchange and count of tokens. Token counting on TLBs is a natural and efficient way for classifying memory pages, and it does not require the use of complex and undesirable persistent requests or arbitration. In addition, classification is extended with Cooperative Usage Predictor (CUP), a token-based system-wide page usage predictor retrieved through TLB cooperation, in order to perform a classification unaffected by TLB size. Through cycle-accurate simulation we observed that TokenTLB spends 43.6 percent of cycles as private per page on average, and CUP further increases the time spent as private by 22.0 percent. CUP avoids 4 out of 5 TLB invalidations when compared to state-of-the-art predictors, thus proving far better prediction accuracy and making usage prediction an attractive mechanism for the first time. Albert Esteve, Alberto Ros 0001, Antonio Robles, María Engracia Gómez |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | TLB-Based Temporality-Aware Classification in CMPs with Multilevel TLBsabstractRecent proposals are based on classifying memory accesses into private or shared in order to process private accesses more efficiently and reduce coherence overhead. The classification mechanisms previously proposed are either not able to adapt to the dynamic sharing behavior of the applications or require frequent broadcast messages. Additionally, most of these classification approaches assume single-level translation lookaside buffers (TLBs). However, deeper and more efficient TLB hierarchies, such as the ones implemented in current commodity processors, have not been appropriately explored. This paper analyzes accurate classification mechanisms in multilevel TLB hierarchies. In particular, we propose an efficient data classification strategy for systems with distributed shared last-level TLBs. Our approach classifies data accounting for temporal private accesses and constrains TLB-related traffic by issuing unicast messages on first-level TLB misses. When our classification is employed to deactivate coherence for private data in directory-based protocols, it improves the directory efficiency and, consequently, reduces coherence traffic to merely 53.0 percent, on average. Additionally, it avoids some of the overheads of previous classification approaches for purely private TLBs, improving average execution time by nearly 9 percent for large-scale systems. Albert Esteve, Alberto Ros 0001, María Engracia Gómez, Antonio Robles, José Duato |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | TokenTLB: A Token-Based Page Classification ApproachabstractClassifying memory accesses into private or shared data has become a fundamental approach to achieving efficiency and scalability in multi- and many-core systems. Since most memory accesses in both sequential and parallel applications are either private (accessed only by one core) or read-only (not written) data, devoting the full cost of coherence to every memory access results in sub-optimal performance and limits the scalability and efficiency of the multiprocessor. Albert Esteve, Alberto Ros 0001, Antonio Robles, María Engracia Gómez, José Duato |
ICS | 3 |
| 2016 | Efficient TLB-Based Detection of Private Pages in Chip MultiprocessorsabstractMost of the data referenced by sequential and parallel applications running in current chip multiprocessors are referenced by a single thread, i.e., private. Recent proposals leverage this observation to improve many aspects of chip multiprocessors, such as reducing coherence overhead or the access latency to distributed caches. The effectiveness of those proposals depends to a large extent on the amount of detected private data. However, the mechanisms proposed so far do not consider neither thread migration nor the private use of data within different application phases. As a result, a considerable amount of private data is not detected. In order to increase the detection of private data, we propose a TLB-based mechanism that is able to account for both thread migration and application phases. Simulation results show that the average number of pages detected as private significantly increases from 43 percent in previous proposals up to 79 percent in ours while keeping a reasonable TLB miss rate. Furthermore, when our proposal is used to deactivate the coherence for private data in a directory protocol, it improves execution time by 13.5 percent, on average, with respect to previous techniques. Albert Esteve, Alberto Ros 0001, María Engracia Gómez, Antonio Robles, José Duato |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | Temporal-Aware Mechanism to Detect Private Data in Chip MultiprocessorsabstractMost of the data referenced by sequential and parallel applications running in current chip multiprocessors are referenced by only one thread and can be considered as private data. A lot of recent proposals leverage this observation to improve many aspects of chip multiprocessors, such as reducing coherence overhead or the access latency to distributed caches. The effectiveness of those proposals depend to a large extent on the amount of detected private data. However, the mechanisms proposed so far do not consider thread migration and the private use of data within different application phases. As a result, a considerable amount of data is not detected as private. In order to make this detection more accurate and reaching more significant improvements, we propose a mechanism that is able to account for both thread migration and private data within application phases. Simulation results for 16-core systems show that, thanks to our mechanism, the average number of pages detected as private significantly increases from 43% in previous proposals up to 74% in ours. Finally, when our detection mechanism is used to deactivate the coherence for private data in a directory protocol, our proposal improves execution time by 13% with respect to previous proposals. Alberto Ros 0001, Blas Cuesta, María Engracia Gómez, Antonio Robles, José Duato |
ICPP | 4 |
| 2013 | Increasing the Effectiveness of Directory Caches by Avoiding the Tracking of Noncoherent Memory BlocksabstractA key aspect in the design of efficient multiprocessor systems is the cache coherence protocol. Although directory-based protocols constitute the most scalable approach, the limited size of the directory caches together with the growing size of systems may cause frequent evictions and, consequently, the invalidation of cached blocks, which jeopardizes system performance. Directory caches keep track of every memory block stored in processor caches in order to provide coherent access to the shared memory. However, a significant fraction of the cached memory blocks do not require coherence maintenance (even in parallel applications) because they are either accessed by just one processor or they are never modified. In this paper, we propose to deactivate the coherence protocol for those blocks that do not require coherence. This deactivation means directory caches do not have to keep track of noncoherent blocks, which reduces directory cache occupancy and increases its effectiveness. Since the detection of noncoherent blocks is carried out by the operating system, our proposal only requires minor hardware modifications. Simulation results show that, thanks to our proposal, directory caches can avoid the tracking of about 66 percent (on average) of the blocks accessed by a wide range of applications, thereby improving the efficiency of directory caches. This contributes either to shortening the runtime of parallel applications by 15 percent (on average) while keeping directory cache size or to maintaining performance while using directory caches 16 times smaller. Blas Cuesta, Alberto Ros 0001, María Engracia Gómez, Antonio Robles, José Duato |
IEEE Trans. Computers | 4 |
| 2012 | Cache Miss Characterization in Hierarchical Large-Scale Cache-Coherent SystemsabstractThere is a growing trend towards developing large-scale cache-coherent systems by using commodity symmetric multiprocessors, which requires to extend their coherence protocol. In such systems, cache coherence transactions issued due to cache misses traverse interconnection networks with very different topologies and latencies. In this work, we perform a cache miss characterization aimed at analyzing the benefits that can be expected for a specialized coherence controller able to locally resolve cache misses, thus saving traffic across long-latency links. Results show that there is a high potential in reducing miss latency in these systems, and that this potential reduction grows as the number of nodes in the system increases. Particularly, in a system with just two boards 40% of the cache misses do not need the expensive inter-board communication. This percentage can increase up to 67.5% for an 8-board system. Alberto Ros 0001, Blas Cuesta, María Engracia Gómez, Antonio Robles, José Duato |
ISPA | 4 |
| 2012 | Switch-based packing technique to reduce traffic and latency in token coherence
Blas Cuesta, Antonio Robles, José Duato |
J. Parallel Distributed Comput. | 2 |
| 2012 | Progressive Congestion Management Based on Packet Marking and Validation TechniquesabstractCongestion management in multistage interconnection networks is a serious problem, which is not solved completely. In order to avoid the degradation of network performance when congestion appears, several congestion management mechanisms have been proposed. Most of these mechanisms are based on explicit congestion notification. For this purpose, switches detect congestion and depending on the applied strategy, packets are marked to warn the source hosts. In response, source hosts apply some corrective actions to adjust their packet injection rate. Although these proposals seem quite effective, they either exhibit some drawbacks or are partial solutions. Some of them introduce some penalties over the flows not responsible for congestion, whereas others can cope only with congestion situations that last for a short time. In this paper, we present an overview of the different strategies to detect and correct congestion in multistage interconnection networks, and propose a new mechanism referred to as Marking and Validation Congestion Management (MVCM), targeted to this kind of lossless networks, and based on a more refined packet marking strategy combined with a fair set of corrective actions, that makes the mechanism able to effectively manage congestion regardless of the congestion degree. Evaluation results show the effectiveness and robustness of the proposed mechanism. Joan-Lluís Ferrer, Elvira Baydal, Antonio Robles, Pedro López 0001, José Duato |
IEEE Trans. Computers | 3 |
| 2012 | Extending Magny-Cours Cache CoherenceabstractOne cost-effective way to meet the increasing demand for larger high-performance shared-memory servers is to build clusters with off-the-shelf processors connected with low-latency point-to-point interconnections like HyperTransport. Unfortunately, HyperTransport addressing limitations prevent building systems with more than eight nodes. While the recent High-Node Count HyperTransport specification overcomes this limitation, recently launched twelve-core Magny-Cours processors have already inherited it and provide only 3 bits to encode the pointers used by the directory cache which they include to increase the scalability of their coherence protocol. In this work, we propose and develop an external device to extend the coherence domain of Magny-Cours processors beyond the 8-node limit while maintaining the advantages provided by the directory cache. Evaluation results for systems with up to 32 nodes show that the performance offered by our solution scales with the number of nodes, enhancing the directory cache effectiveness by filtering additional messages. Particularly, we reduce execution time by 47 percent in a 32-die system with respect to the 8-die Magny-Cours configuration. Alberto Ros 0001, Blas Cuesta, Ricardo Fernández-Pascual, María Engracia Gómez, Manuel E. Acacio, Antonio Robles, José M. García 0001, José Duato |
IEEE Trans. Computers | 6 |
| 2012 | A Survey and Evaluation of Topology-Agnostic Deterministic Routing AlgorithmsabstractMost standard cluster interconnect technologies are flexible with respect to network topology. This has spawned a substantial amount of research on topology-agnostic routing algorithms, which make no assumption about the network structure, thus providing the flexibility needed to route on irregular networks. Actually, such an irregularity should be often interpreted as minor modifications of some regular interconnection pattern, such as those induced by faults. In fact, topology-agnostic routing algorithms are also becoming increasingly useful for networks on chip (NoCs), where faults may make the preferred 2D mesh topology irregular. Existing topology-agnostic routing algorithms were developed for varying purposes, giving them different and not always comparable properties. Details are scattered among many papers, each with distinct conditions, making comparison difficult. This paper presents a comprehensive overview of the known topology-agnostic routing algorithms. We classify these algorithms by their most important properties, and evaluate them consistently. This provides significant insight into the algorithms and their appropriateness for different on- and off-chip environments. José Flich, Tor Skeie, Andres Mejia, Olav Lysne, Pedro López 0001, Antonio Robles, José Duato, Michihiro Koibuchi, Tomas Rokicki, José Carlos Sancho |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2011 | Increasing the effectiveness of directory caches by deactivating coherence for private memory blocksabstractTo meet the demand for more powerful high-performance shared-memory servers, multiprocessor systems must incorporate efficient and scalable cache coherence protocols, such as those based on directory caches. However, the limited directory cache size of the increasingly larger systems may cause frequent evictions of directory entries and, consequently, invalidations of cached blocks, which severely degrades system performance. Blas Cuesta, Alberto Ros 0001, María Engracia Gómez, Antonio Robles, José Duato |
ISCA | 4 |
| 2011 | Efficient and Scalable Starvation Prevention Mechanism for Token CoherenceabstractToken Coherence is a cache coherence protocol that simultaneously captures the best attributes of the traditional approximations to coherence: direct communication between processors (like snooping-based protocols) and no reliance on bus-like interconnects (like directory-based protocols). This is possible thanks to a class of unordered requests that usually succeed in resolving the cache misses. The problem of the unordered requests is that they can cause protocol races, which prevent some misses from being resolved. To eliminate races and ensure the completion of the unresolved misses, Token Coherence uses a starvation prevention mechanism named persistent requests. This mechanism is extremely inefficient and, besides, it endangers the scalability of Token Coherence since it requires storage structures (at each node) whose size grows proportionally to the system size. While multiprocessors continue including an increasingly number of nodes, both the performance and scalability of cache coherence protocols will continue to be key aspects. In this work, we propose an alternative starvation prevention mechanism, named priority requests, that outperforms the persistent request one. This mechanism is able to reduce the application runtime more than 20 percent (on average) in a 64-processor system. Furthermore, thanks to the flexibility shown by priority requests, it is possible to drastically minimize its storage requirements, thereby improving the whole scalability of Token Coherence. Although this is achieved at the expense of a slight performance degradation, priority requests still outperform persistent requests significantly. Blas Cuesta, Antonio Robles, José Duato |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2010 | EMC2: Extending Magny-Cours coherence for large-scale serversabstractThe demand of larger and more powerful high-performance shared-memory servers is growing over the last few years. To meet this need, AMD has recently launched the twelve-core Magny-Cours processors. They include a directory cache (Probe Filter) that increases the scalability of the coherence protocol applied by Opterons, based on coherent Hyper Transport interconnect (cHT). cHT limits up to 8 the number of nodes that can be addressed. Recent High Node Count HT specification overcomes this limitation. However, the 3-bit pointer used by the Probe Filter prevents Magny-Cours-based servers from being built beyond 8 nodes. In this paper, we propose and develop an external logic to extend the coherence domain of Magny-Cours processors beyond the 8-node limit while maintaining the advantages provided by the Probe Filter. Evaluation results for up to a 32-node system show how the performance offered by our solution scales with the increment in the number of nodes, enhancing the Probe Filter effectiveness by filtering additional messages. Particularly, we reduce runtime by 47% in a 32-die system respect to the 8-die Magny-Cours system. Alberto Ros 0001, Blas Cuesta, Ricardo Fernández-Pascual, María Engracia Gómez, Manuel E. Acacio, Antonio Robles, José M. García 0001, José Duato |
HiPC | 6 |
| 2010 | A Scalable and Early Congestion Management Mechanism for MINsabstractSeveral packet marking-based mechanisms have been proposed to manage congestion in multistage interconnection networks. One of them, the MVCM mechanism obtains very good results for different network configurations and traffic loads. However, as MVCM applies full virtual output queuing at origin, its memory requirements may jeopardize its scalability. Additionally, the applied packet marking technique introduces certain delay to detect congestion. In this paper, we propose and evaluate the Scalable Early Congestion Management mechanism which eliminates the drawbacks exhibited by MVCM. The new mechanism replaces the full virtual output queuing at origin by either a partial virtual output queuing or a shared buffer, in order to reduce its memory requirements, thus making the mechanism scalable. Also, it applies an improved packet marking technique based on marking packets at output buffers regardless of their marking at input buffers, which simplifies the marking technique, allowing also a sooner detection of the root of a congestion tree. Joan-Lluís Ferrer, Elvira Baydal, Antonio Robles, Pedro López 0001, José Duato |
PDP | 3 |
| 2008 | On the Influence of the Packet Marking and Injection Control Schemes in Congestion Management for MINs
Joan-Lluís Ferrer, Elvira Baydal, Antonio Robles, Pedro López 0001, José Duato |
Euro-Par | 3 |
| 2008 | Switch-Based Packing Technique for Improving Token Coherence ScalabilityabstractTraditional cache coherence protocols either provide low latency cache misses (snooping protocols) or bandwidth efficiency (directory protocols). To simultaneously capture the best attributes of traditional protocols, Token Coherence has been recently proposed. This protocol can quickly resolve cache misses by transient requests. However, since transient requests are unordered messages, they may sometimes fail in solving cache misses mainly due to the occurrence of protocol races. Thus, when the completion of cache misses is not possible by transient requests, Token Coherence uses a starvation prevention mechanism to ensure their completion. Although several implementation options of starvation prevention mechanisms have been proposed, all of them are broadcast-based. This fact represents a large detriment to the Token Coherence scalability. To tackle this problem, in this work we apply a switch-based packing technique that alleviates the harm of broadcast messages and improves the protocol scalability. Blas Cuesta, Antonio Robles, José Duato |
PDCAT | 2 |
| 2008 | Improving Token Coherence by Multicast Coherence MessagesabstractToken coherence is a cache coherence protocol that joins the main advantages of traditional protocols. However, unlike them, token coherence does not handle messages in order, which may lead to races, causing some cache misses not to be solved. To assure their completion, an inefficient mechanism named persistent requests is used. Recently we have proposed the priority request mechanism to efficiently handle races. As acknowledgements are not required, a single node can solve several misses for the same memory block at the same time. When solving a lot of misses, the node may become a bottleneck. To avoid it, in this work we propose the multicast coherence message, which allows to simultaneously resolve several misses by using only one response message. It reduces the network traffic and the average response latency, improving significantly the overall performance. Blas Cuesta, Antonio Robles, José Duato |
PDP | 2 |
| 2008 | EST2uni: an open, parallel tool for automated EST analysis and database creation, with a data mining web interface and microarray expression data integrationabstractBACKGROUND: Expressed sequence tag (EST) collections are composed of a high number of single-pass, redundant, partial sequences, which need to be processed, clustered, and annotated to remove low-quality and vector regions, eliminate redundancy and sequencing errors, and provide biologically relevant information. In order to provide a suitable way of performing the different steps in the analysis of the ESTs, flexible computation pipelines adapted to the local needs of specific EST projects have to be developed. Furthermore, EST collections must be stored in highly structured relational databases available to researchers through user-friendly interfaces which allow efficient and complex data mining, thus offering maximum capabilities for their full exploitation. RESULTS: We have created EST2uni, an integrated, highly-configurable EST analysis pipeline and data mining software package that automates the pre-processing, clustering, annotation, database creation, and data mining of EST collections. The pipeline uses standard EST analysis tools and the software has a modular design to facilitate the addition of new analytical methods and their configuration. Currently implemented analyses include functional and structural annotation, SNP and microsatellite discovery, integration of previously known genetic marker data and gene expression results, and assistance in cDNA microarray design. It can be run in parallel in a PC cluster in order to reduce the time necessary for the analysis. It also creates a web site linked to the database, showing collection statistics, with complex query capabilities and tools for data mining and retrieval. CONCLUSION: The software package presented here provides an efficient and complete bioinformatics tool for the management of EST collections which is very easy to adapt to the local needs of different EST projects. The code is freely available under the GPL license and can be obtained at http://bioinf.comav.upv.es/est2uni. This site also provides detailed instructions for installation and configuration of the software package. The code is under active development to incorporate new analyses, methods, and algorithms as they are released by the bioinformatics community. Javier Forment, Francisco Gilabert Villamón, Antonio Robles, Vicente Conejero, Fernando Nuez, Jose M. Blanca |
BMC Bioinform. | 3 |
| 2007 | An Effective Starvation Avoidance Mechanism to Enhance the Token Coherence ProtocolabstractShared-memory multiprocessors are becoming to be formed by an increasingly larger number of nodes. In these systems, implementing cache coherence is a key issue. Token coherence is a low latency cache coherence protocol that avoids indirection for cache-to-cache misses and which does not require a totally-ordered interconnect. When races are rare, the protocol performs well thanks to the performance policy. Unfortunately, some medium/large systems and some applications that often access the same data simultaneously make races more common. As a result, the protocol does not perform as well as it could because it uses the persistent request mechanism to prevent starvation. This mechanism is too slow and inflexible because it overrides the performance policy. In consequence, the protocol slows down the system and does not take advantage of the flexibility and speed of the common case. We propose a new mechanism, namely priority requests, which replaces the persistent request one. Our mechanism solves races, while still respecting the performance policy, simply by ordering and giving a higher priority to requests suffering from starvation. Thus, our mechanism handles the tokens more efficiently and reduces the network traffic Blas Cuesta, Antonio Robles, José Duato |
PDP | 2 |
| 2007 | Congestion Management in MINs through Marked and Validated PacketsabstractCongestion management is a very critical problem tackled in interconnection networks for years but not solved yet. Although several mechanisms have been recently proposed for lossless multistage interconnection networks (MINs), they either have drawbacks or are partial solutions. Some of them introduce penalty over packets not really addressed to the hot-spots, whereas others can cope only with congestion situations that last a short time. In this paper, we propose an effective and efficient congestion management mechanism for lossless interconnection networks based on explicit congestion notification. The mechanism uses two different flags in ACK packets, a Marking Bit (MB) and a Validation Bit (VB), to detect congestion and warn the origin hosts. In this way, packets belonging to "coldflows" but stopped because of head-of-line (HOL) blocking can be distinguished from "hotflow" packets which are really causing congestion. In response, origin hosts can apply corrective actions only to the "hotflows", minimizing the negative impact on "coldflows"performance. Evaluation results show that the proposed congestion management strategy is able to avoid the degradation of network performance, regardless of traffic load and the location of the congestion in the network. Joan-Lluís Ferrer, Elvira Baydal, Antonio Robles, Pedro López 0001, José Duato |
PDP | 3 |
| 2006 | A Routing Methodology for Achieving Fault Tolerance in Direct NetworksabstractMassively parallel computing systems are being built with thousands of nodes. The interconnection network plays a key role for the performance of such systems. However, the high number of components significantly increases the probability of failure. Additionally, failures in the interconnection network may isolate a large fraction of the machine. It is therefore critical to provide an efficient fault-tolerant mechanism to keep the system running, even in the presence of faults. This paper presents a new fault-tolerant routing methodology that does not degrade performance in the absence of faults and tolerates a reasonably large number of faults without disabling any healthy node. In order to avoid faults, for some source-destination pairs, packets are first sent to an intermediate node and then from this node to the destination node. Fully adaptive routing is used along both subpaths. The methodology assumes a static fault model and the use of a checkpoint/restart mechanism. However, there are scenarios where the faults cannot be avoided solely by using an intermediate node. Thus, we also provide some extensions to the methodology. Specifically, we propose disabling adaptive routing and/or using misrouting on a per-packet basis. We also propose the use of more than one intermediate node for some paths. The proposed fault-tolerant routing methodology is extensively evaluated in terms of fault tolerance, complexity, and performance. María Engracia Gómez, Nils Agne Nordbotten, José Flich, Pedro López 0001, Antonio Robles, José Duato, Tor Skeie, Olav Lysne |
IEEE Trans. Computers | 5 |
| 2005 | Enforcing in-order packet delivery in system area networks with adaptive routing
Michihiro Koibuchi, José Flich, Antonio Robles, Pedro López 0001, José Duato |
J. Parallel Distributed Comput. | 4 |
| 2004 | A Methodology to Evaluate the Effectiveness of Traffic Balancing Algorithms
J. E. Villalobos, José L. Sánchez 0002, José A. Gámez 0001, José Carlos Sancho, Antonio Robles |
Euro-Par | 5 |
| 2004 | A New Adaptive Fault-Tolerant Routing Methodology for Direct Networks
María Engracia Gómez, José Duato, José Flich, Pedro López 0001, Antonio Robles, Nils Agne Nordbotten, Tor Skeie, Olav Lysne |
HiPC | 5 |
| 2004 | LASH-TOR: A Generic Transition-Oriented Routing Algorithm
Tor Skeie, Olav Lysne, José Flich, Pedro López 0001, Antonio Robles, José Duato |
ICPADS | 5 |
| 2004 | An Effective Fault-Tolerant Routing Methodology for Direct NetworksabstractCurrent massively parallel computing systems are being built with thousands of nodes, which significantly affect the probability of failure. M. E. Gomex proposed a methodology to design fault-tolerant routing algorithms for direct interconnection networks. The methodology uses a simple mechanism: for some source-destination pairs, packets are first forwarded to an intermediate node, and later, from this node to the destination node. Minimal adaptive routing is used along both subpaths. For those cases where the methodology cannot find a suitable intermediate node, it combines the use of intermediate nodes with two additional mechanisms: disabling adaptive routing and using misrouting on a per-packet basis. While the combination of these three mechanisms tolerates a large number of faults, each one requires adding some hardware support in the network and also introduces some overhead. In this paper, we perform an in-depth detailed analysis of the impact of these mechanisms on network behaviour. We analyze the impact of the three mechanisms separately and combined. The ultimate goal of this paper is to obtain a suitable combination of mechanisms that is able to meet the trade-off between fault-tolerance degree, routing complexity, and performance. María Engracia Gómez, José Flich, Pedro López 0001, Antonio Robles, José Duato, Nils Agne Nordbotten, Olav Lysne, Tor Skeie |
ICPP | 4 |
| 2004 | A Transition-Based Fault-Tolerant Routing Methodology for InfiniBand NetworksabstractSummary form only given. Currently, clusters of PCs are considered a cost-effective alternative to large parallel computers. As the number of elements increases in these systems, the probability of faults increases dramatically. Therefore, it is critical to keep the system running even in the presence of faults. The interconnection network plays a key role in its performance. InfiniBand (IBA) is a new standard interconnect suitable for clusters. Most of the fault-tolerant routing strategies proposed for massively parallel computers cannot be applied to IBA because routing and virtual channel transitions are deterministic, which prevents packets from avoiding the faults. A possible approach to provide fault-tolerance in IBA consists of using several disjoint paths between every source-destination pair of nodes and selecting the appropriate path at the source host. However, to this end, a routing algorithm able to provide enough disjoint paths, while still guaranteeing deadlock freedom, is required. We propose a simple and effective fault-tolerant methodology for IBA networks that can be applied to any network topology and meets the trade-off between fault-tolerance degree and the number of network resources devoted to it. Preliminary results show that the proposed methodology scales well and supports up to three faults in 2D and five in 3D tori using only two virtual channels. José Miguel Montañana, José Flich, Antonio Robles, Pedro López 0001, José Duato |
IPDPS | 3 |
| 2004 | A Fully Adaptive Fault-Tolerant Routing Methodology Based on Intermediate Nodes
Nils Agne Nordbotten, María Engracia Gómez, José Flich, Pedro López 0001, Antonio Robles, Tor Skeie, Olav Lysne, José Duato |
NPC | 5 |
| 2004 | An Effective Methodology to Improve the Performance of the Up*/Down* Routing AlgorithmabstractNetworks of workstations (NOWs) are being considered as a cost-effective alternative to parallel computers. Most NOWs are arranged as a switch-based network and provide mechanisms for discovering the network topology. Hence, they provide support for both regular and irregular topologies, which makes routing and deadlock avoidance quite complicated. Current proposals use the up*/down* routing algorithm to remove cyclic dependencies between channels and avoid deadlock. However, routing is considerably restricted and most messages must follow nonminimal paths, increasing latency and wasting resources. We propose and evaluate a simple and effective methodology to compute up*/down* routing tables. The new methodology is based on computing a depth-first search (DPS) spanning tree on the network graph that decreases the number of routing restrictions with respect to the breadth-first search (BFS) spanning tree used by the traditional methodology. Additionally, we propose different heuristic rules for computing the spanning trees to improve the efficiency of up*/down* routing. Evaluation results for several different topologies show that computing the up*/down* routing tables by using the new methodology increases throughput by a factor of up to 2.48 in large networks with respect to the traditional methodology, and also reduces latency significantly. José Carlos Sancho, Antonio Robles, José Duato |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2003 | Low-Fragmentation Mapping Strategies for Linear Forwarding Tables in InfiniBandTM
Pedro López 0001, José Flich, Antonio Robles |
Euro-Par | 3 |
| 2003 | Routing in InfiniBandTM Torus Network TopologieabstractInfiniBand is an interconnect standard for communication between processing nodes and I/O devices as well as for interprocessor communication (NOWs). The InfiniBand architecture (IBA) defines a switch-based network with point-to-point links whose topology can be established by the customer. When the performance is the primary concern regular topologies are preferred. Low-dimensional tori (2D and 3D) are some of the regular topologies most widely used in commercial parallel computers. Routing in torus requires the use of virtual channels. Although InfiniBand provides support for deterministic routing and virtual channels, they are selected at each switch by service level (SL) identifiers associated to packets and do not depend on packet destination. This makes routing algorithm implementation more complex. In particular, a large number of SLs may be required, which is a scarce resource. We analyze the way several routing strategies can be applied in tori InfiniBand networks, also evaluating their resource requirements. In particular, we analyze and compare the well-known e-cube and up*/down* routing algorithms and the flexible routing algorithm recently proposed José Carlos Sancho, Antonio Robles, Pedro López 0001, José Flich, José Duato |
ICPP | 2 |
| 2003 | Supporting adaptive routing in IBA switches
José Flich, Antonio Robles, Pedro López 0001, José Duato |
J. Syst. Archit. | 3 |
| 2002 | Evaluation of Routing Algorithms for InfiniBand Networks (Research Note)
María Engracia Gómez, José Flich, Antonio Robles, Pedro López 0001, José Duato |
Euro-Par | 3 |
| 2002 | Effective Methodology for Deadlock-Free Minimal Routing in InfiniBand NetworksabstractThe InfiniBand Architecture (IBA) defines a switch-based network with point-to-point links whose topology is arbitrarily established by the customer. We propose a simple and effective methodology for designing deadlock-free routing strategies that are able to route packets through minimal paths in InfiniBand networks. This methodology can meet the trade-off between network performance and the number of resources dedicated to deadlock avoidance. Evaluation results show that the resulting routing strategies significantly outperform up*/down* routing. In particular, throughput improvement ranges, on average, from 1.33 for small networks to 4.05 for large networks. Also, it is shown that just two virtual lanes and three service levels are enough to achieve more than 80% of the throughput improvement achieved by the best proposed routing strategy (the one that always provides minimal paths without limiting the number of resources). José Carlos Sancho, Antonio Robles, José Flich, Pedro López 0001, José Duato |
ICPP | 2 |
| 2001 | Topic 12: Routing and Communication in Interconnection Networks
Ramón Beivide, Chris R. Jesshope, Antonio Robles, Cruz Izu |
Euro-Par | 3 |
| 2001 | Effective Strategy to Compute Forwarding Tables for InfiniBand NetworksabstractInfiniBand is very likely to become the facto standard for communication between processing nodes and I/O devices as well as for interprocessor communication. The InifiniBand Architecture (IBA) defines a switch-based network with point-to-point links that support any topology defined by the user. Routing in IBA is distributed based on forwarding tables, and only considers the packet destination ID for routing within subnets. Up*/down* routing is the simplest and most popular routing algorithm for irregular topologies. Unfortunately, up*/down* routing cannot be used in IBA switches because it may leads to deadlock. In this paper we address this issue, proposing an easy-to-implement strategy to complete up*/down* forwarding tables for IBA switches that guarantees deadlock freedom, and is effective whatever the methodology applied to compute up*/down* routing tables. Preliminary evaluation results modeling an InfiniBand network at register transfer level show that the proposed strategy allows up*/down* routing algorithms to be implemented on InfiniBand networks with minimal performance degradation. José Carlos Sancho, Antonio Robles, José Duato |
ICPP | 2 |
| 2001 | A Comparison of Router Architectures for Virtual Cut-Through and Wormhole Switching in a NOW Environment
José Duato, Antonio Robles, Federico Silla, Ramón Beivide |
J. Parallel Distributed Comput. | 2 |
| 2000 | Improving the Up*/Down* Routing Scheme for Networks of Workstations
José Carlos Sancho, Antonio Robles |
Euro-Par | 2 |
| 1998 | Improving Performance of Networks of Workstations by using Disha ConcurrentabstractNetworks of workstations are currently emerging as a cost-effective alternative to parallel computers. Recently, deadlock recovery techniques have been shown to be an alternative to deadlock avoidance. Disha Concurrent is a progressive deadlock recovery scheme able to simultaneously redirect several deadlocked messages through a deadlock-free lane. Unlike deadlock avoidance techniques, Disha provides true fully adaptive routing without using virtual channels to guarantee deadlock freedom. In this paper, we analyze the application of Disha to networks of workstations. We propose an implementation of Disha on irregular networks that allows concurrent deadlock recovery proving that this implementation is always able to recover from deadlock. A new switch organization and a new flow control protocol are proposed to support Disha. Performance evaluation results show that applying Disha to irregular networks increases network throughput by a factor of up to 3.5, and also reduces latency with regard to other routing algorithms based on deadlock avoidance techniques. Federico Silla, Antonio Robles, José Duato |
ICPP | 2 |