Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Fernando Vallejo

dblp:32/3504 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
0since 2021 · last 2016
0000-0003-3343-6479ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 1 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Interconnection networks and networks-on-chip · 58% Electronic design automation · 18% Parallel and multicore computing · 16%
Software engineering, system software, and programming languages
1 paper
Operating systems · 77% Concurrent programming · 23%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation › physical design
routing
0.322016
Assessing the Suitability of King Topologies for Interconnection Networks · IEEE Trans. Parallel Distributed Syst. 2016
A new routing mechanism for networks with irregular topology · SC 2001
Interconnection networks and networks-on-chip
network topology
0.212016
Assessing the Suitability of King Topologies for Interconnection Networks · IEEE Trans. Parallel Distributed Syst. 2016
Interconnection networks and networks-on-chip
routing algorithms
0.212016
Assessing the Suitability of King Topologies for Interconnection Networks · IEEE Trans. Parallel Distributed Syst. 2016
Interconnection networks and networks-on-chip › routing algorithms
fault-tolerant routing
0.122008
Immunet: Dependable Routing for Interconnection Networks with Arbitrary Topology · IEEE Trans. Computers 2008
Immunet: A Cheap and Robust Fault-Tolerant Packet Routing Mechanism · ISCA 2004
Parallel and multicore computing › synchronization
reader-writer locks
0.112010
Architectural Support for Fair Reader-Writer Locking · MICRO 2010
Parallel and multicore computing
synchronization
0.112010
Architectural Support for Fair Reader-Writer Locking · MICRO 2010
Distributed systems
fault tolerance
0.122008
Immunet: Dependable Routing for Interconnection Networks with Arbitrary Topology · IEEE Trans. Computers 2008
Immunet: A Cheap and Robust Fault-Tolerant Packet Routing Mechanism · ISCA 2004
Interconnection networks and networks-on-chip › network reconfiguration
dynamic network reconfiguration
0.112008
Immunet: Dependable Routing for Interconnection Networks with Arbitrary Topology · IEEE Trans. Computers 2008
Interconnection networks and networks-on-chip › routing algorithms
interconnection routing
0.112008
Immunet: Dependable Routing for Interconnection Networks with Arbitrary Topology · IEEE Trans. Computers 2008
Interconnection networks and networks-on-chip
network reconfiguration
0.012004
Immunet: A Cheap and Robust Fault-Tolerant Packet Routing Mechanism · ISCA 2004
Processor architecture and microarchitecture
multicore design
0.012010
Architectural Support for Fair Reader-Writer Locking · MICRO 2010
Interconnection networks and networks-on-chip › routing algorithms › adaptive routing
fully adaptive routing
0.012001
A new routing mechanism for networks with irregular topology · SC 2001
Interconnection networks and networks-on-chip › network topology › static interconnection networks
irregular topology
0.012001
A new routing mechanism for networks with irregular topology · SC 2001
Operating systems › multiprocessing
multiprocessor operating system
0.011994
Shared Memory Multimicroprocessor Operating System with an Extended Petri Net Model · IEEE Trans. Parallel Distributed Syst. 1994
Parallel and multicore computing › programming models
event-driven programming
0.011994
Shared Memory Multimicroprocessor Operating System with an Extended Petri Net Model · IEEE Trans. Parallel Distributed Syst. 1994
Parallel and multicore computing
parallel programming models
0.011994
Shared Memory Multimicroprocessor Operating System with an Extended Petri Net Model · IEEE Trans. Parallel Distributed Syst. 1994

Methods — techniques the papers use, named apart from their topics

topological analysis · 0.2simulation · 0.2local information routing · 0.1hardware implementation · 0.1virtual cut-through · 0.0restricted packet injection · 0.0pseudo-hamiltonian cycle · 0.0extended petri net · 0.0
YearPublicationVenuePosition
2016 Assessing the Suitability of King Topologies for Interconnection Networks
abstract
In the late years many different interconnection networks have been used with two main tendencies. One is characterized by the use of high-degree routers with long wires while the other uses routers of much smaller degree. The latter rely on two-dimensional mesh and torus topologies with shorter local links. This paper focuses on doubling the degree of common 2D meshes and tori while still preserving an attractive layout for VLSI design. By adding a set of diagonal links in one direction, diagonal networks are obtained. By adding a second set of links, networks of degree eight are built, named king networks. This research presents a comprehensive study of these networks which includes a topological analysis, the proposal of appropriate routing procedures and an empirical evaluation. King networks exhibit a number of attractive characteristics which translate to reduced execution times of parallel applications. For example, the execution times NPB suite are reduced up to a 30 percent. In addition, this work reveals other properties of king networks such as perfect partitioning that deserves further attention for its convenient exploitation in forthcoming high-performance parallel systems.
Esteban Stafford, José Luis Bosque, Carmen Martínez 0001, Fernando Vallejo, Ramón Beivide, Cristobal Camarero, Emilio Castillo
IEEE Trans. Parallel Distributed Syst.4
2013 Advanced Switching Mechanisms for Forthcoming On-Chip Networks
abstract
Many current VLSI on-chip multiprocessors and systems-on-chip employ point-to-point switched interconnection networks. Rings and 2D-meshes are among the most popular interconnection topologies for these increasingly important onchip networks. Nevertheless, rings cannot scale beyond dozens of nodes and meshes are asymmetric. Two of the key features of square 2D-tori are their scalability and symmetry. As higher scalability is demanded by the increasing number of cores (or specialized units) integrated on a chip and symmetry is critical for high-performance and load balancing, we concentrate on 2D-tori. However, most popular deadlock-free routing mechanisms are based on Dimension Order Routing (DOR) which breaks the torus symmetry when managing adversarial traffic patterns. This paper analyzes this problem and its consequences. After that, it proposes a new deadlock-free fully adaptive minimal routing, denoted as σDOR, that preserves tori symmetry under any load. It uses just two virtual channels to avoid DOR-induced asymmetry, the same as in previous competitive proposals. σDOR exhibits better behavior than any of previous solutions as it allows packets to dynamically adapt to local congestion. Experimental results evidence the superior performance of our mechanism, confirming the negative impact of DOR asymmetry.
Emilio Castillo, Cristobal Camarero, Esteban Stafford, Fernando Vallejo, José Luis Bosque, Ramón Beivide
DSD4
2010 A First Approach to King Topologies for On-Chip Networks
Esteban Stafford, José Luis Bosque, Carmen Martínez 0001, Fernando Vallejo, Ramón Beivide, Cristobal Camarero
Euro-Par (2)4
2010 Architectural Support for Fair Reader-Writer Locking
abstract
Many shared-memory parallel systems use lock-based synchronization mechanisms to provide mutual exclusion or reader-writer access to memory locations. Software locks are inefficient either in memory usage, lock transfer time, or both. Proposed hardware locking mechanisms are either too specific (for example, requiring static assignment of threads to cores and vice-versa), support a limited number of concurrent locks, require tag values to be associated with every memory location, rely on the low latencies of single-chip multicore designs or are slow in adversarial cases such as suspended threads in a lock queue. Additionally, few proposals cover reader-writer locks and their associated fairness issues. In this paper we introduce the Lock Control Unit (LCU) which is an acceleration mechanism collocated with each core to explicitly handle fast reader-writer locking. By associating a unique thread-id to each lock request we decouple the hardware lock from the requestor core. This provides correct and efficient execution in the presence of thread migration. By making the LCU logic autonomous from the core, it seamlessly handles thread preemption. Our design offers richer semantics than previous proposals, such as try lock support while providing direct core-to-core transfers. We evaluate our proposal with micro benchmarks, a fine-grain Software Transactional Memory system and programs from the Parsec and Splash parallel benchmark suites. The lock transfer time decreases in up to 30% when compared to previous hardware proposals. Transactional Memory systems limited by reader-locking congestion boost up to 3x while still preserving graceful fairness and starvation freedom properties. Finally, commonly used applications achieve speedups up to a 7% when compared to software models.
Enrique Vallejo 0001, Ramón Beivide, Adrián Cristal, Tim Harris 0001, Fernando Vallejo, Osman S. Unsal, Mateo Valero
MICRO5
2009 Light NUCA: A proposal for bridging the inter-cache latency gap
abstract
To deal with the “memory wall” problem, microprocessors include large secondary on-chip caches. But as these caches enlarge, they originate a new latency gap between them and fast L1 caches (inter-cache latency gap). Recently, Non-Uniform Cache Architectures (NUCAs) have been proposed to sustain the size growth trend of secondary caches that is threatened by wire-delay problems. NUCAs are size-oriented, and they were not conceived to close the inter-cache latency gap. To tackle this problem, we propose Light NUCAs (L-NUCAs) leveraging on-chip wire density to interconnect small tiles through specialized networks, which convey packets with distributed and dynamic routing. Our design reduces the tile delay (cache access plus one-hop routing) to a single processor cycle and places cache lines at a finer granularity than conventional caches, reducing cache latency. Our evaluations show that in general, an L-NUCA improves simultaneously performance, energy, and area when integrated into both conventional or D-NUCA hierarchies.
Darío Suárez Gracia, Teresa Monreal Arnal, Fernando Vallejo, Ramón Beivide, Víctor Viñals
DATE3
2008 Graph-based metrics over QAM constellations
abstract
In order to propose a new metric over QAM constellations, diagonal Gaussian graphs defined over quotients of the Gaussian integers are introduced in this paper. Distance properties of the constellations are detailed by means of the vertex-to-vertex distribution of this family of graphs. Moreover, perfect codes for this metric are considered. Finally, notable subgraphs of diagonal Gaussian graphs are studied which leads to relate the proposed metric to other well-known graph-based metrics such as the Lee distance.
Carmen Martínez 0001, Esteban Stafford, Ramón Beivide, Cristobal Camarero, Fernando Vallejo, Ernst M. Gabidulin
ISIT5
2008 Immunet: Dependable Routing for Interconnection Networks with Arbitrary Topology
abstract
A complete mechanism for tolerating multiple failures in parallel computer systems, denoted as Immunet, is described in this paper. Immunet can be applied to arbitrary topologies, either regular or irregular, exhibiting in both cases graceful performance degradation. Provided that the network remains connected, Immunet is able to deal with any number of failures regardless of their spatial and temporal distribution. Our mechanism operates on the basis of a dynamic network reconfiguration in response to failures. The network reconfiguration only employs local information recorded at the router nodes which leads to a highly scalable system. In addition, its low cost and overhead permit a practicable hardware implementation. Finaly, Immunet could allow circumvent failures transparently to applications running on a parallel system because it does not require dropping in-flight traffic. Only packets stored in or traveling through a broken component should be recovered by higher system levels.
Valentin Puente, José-Ángel Gregorio, Fernando Vallejo, Ramón Beivide
IEEE Trans. Computers3
2006 High-performance adaptive routing for networks with arbitrary topology
Valentin Puente, José-Ángel Gregorio, Fernando Vallejo, Ramón Beivide, Cruz Izu
J. Syst. Archit.3
2004 Load Unbalance in k-ary n-Cube Networks
José Miguel-Alonso, José-Ángel Gregorio, Valentin Puente, Fernando Vallejo, Ramón Beivide
Euro-Par4
2004 Immunet: A Cheap and Robust Fault-Tolerant Packet Routing Mechanism
abstract
A new and efficient mechanism to tolerate failures in interconnection networks for parallel and distributed computers, denoted as Immunet, is presented in this work. In the presence of failures, Immunet automatically reacts with a hardware reconfiguration of the surviving network resources. Immunet has four important advantages over previous fault-tolerant switching mechanisms. Its low hardware costs minimize the overhead that the network must support in absence of faults. As long as the network remains connected, Immunet can tolerate any number of failures regardless of their spatial and temporal combinations. The resulting communication infrastructure provides optimized adaptive minimal routing over the surviving topology. The system behavior under successive failures exhibits graceful performance degradation. Immunet reconfiguration can be totally transparent to the applications running on the parallel system as they will only be affected by the loss of those data packets circulating through the broken components. The rest of the packets will suffer only a tolerable delay induced by the time employed to perform the automatic network reconfiguration. Descriptions of the hardware network architecture and detailed synthetic and execution-driven simulations will demonstrate the benefits of Immunet.
Valentin Puente, José-Ángel Gregorio, Fernando Vallejo, Ramón Beivide
ISCA3
2002 Modeling of interconnection subsystems for massively parallel computers
José-Ángel Gregorio, Ramón Beivide, Fernando Vallejo
Perform. Evaluation3
2001 A new routing mechanism for networks with irregular topology
abstract
Selecting a Pseudo-Hamiltonian cycle in any irregular network and applying a restricted packet injection mechanism to avoid the exhaustion of the storage resources, a new fully adaptive routing algorithm has been developed and tested. Our new routing mechanism outperforms the most relevant routing proposals for networks with irregular topology. In all the tested cases a significant improvement has been obtained. The most spectacular gains were obtained for big networks. For a 512-node network, uniform traffic, and virtual cut-through flow control, our mechanism can outperform, in some cases, the classic up*/down* algorithm by almost a factor of 2.
Valentin Puente, José-Ángel Gregorio, Ramón Beivide, Fernando Vallejo, Andres Ibañez
SC4
2001 The Adaptive Bubble Router
Valentin Puente, Cruz Izu, Ramón Beivide, José-Ángel Gregorio, Fernando Vallejo, J. M. Prellezo
J. Parallel Distributed Comput.5
2000 Improving parallel system performance by changing the arrangement of the network links
abstract
The Midimew network is an excellent contender for implementing the communication subsystem of a high performance computer. This network is an optimal 2D topology in the sense there are no other symmetric direct networks of degree 4 with a lower average distance or diameter. In fact, it reduces the diameter of the well known torus network by approximately □2. Although the topology was proposed and analyzed a decade ago, the lack of simple deadlock avoidance mechanisms prevented its utilization up to date. This study solved this drawback by applying the Bubble switching mechanism, a low cost deadlock-avoidance strategy developed by the authors. Moreover, by using routing tables we can configure our Virtual Cut-Through adaptive router to implement either a torus or a Midimew network. Thus, we can exploit the topological advantages of Midimew networks by simply changing the disposition of the wrap-around connections of its torus counterpart, without increasing the network implementation cost. To prove this assertion, we have carried out a thorough evaluation, from the hardware cost of the router to the parallel system performance under real loads.
Valentin Puente, Cruz Izu, José-Ángel Gregorio, Ramón Beivide, J. M. Prellezo, Fernando Vallejo
ICS6
1999 Low-level router design and its impact on supercomputer system performance
abstract
Supercomputer performance is highly dependent on its interconnection subsystem design.In this paper we study how different architectural approaches for router design impact into system performance when running real parallel applications.A thorough methodology has been employed to quantify this impact.Architectural router decisions have been chosen taking into account the constraints of the underlying VLSI technology.After that, an exhaustive evaluation of the interconnection network under standard synthetic traffic has been carried out.Finally, an execution-driven simulation environment has been used to assess the consequences of several router designs on the performance of the entire machine.We will show that low-level decisions, as the adequate selection of router's arbiter, significantly reduce the execution time of parallel applications.To illustrate the effects of the router architecture on system performance two benchmarks were selected: Radix and MPSD. IntroductionIn the field of high-performance computing, distributed shared-memory multiprocessors (DSMS) are becoming widespread.These parallel computers implement a single address space, either with coherent caches (SGI Origin 2000 [13]) or without them (Cray T3E 1181).The communication time involved on fetching remote data is one of the main overheads which limits the performance of many parallel applications.Moreover, cc-NUMA machines impose additional overheads due to synchronization amongst processes and coherence maintenance.As processor computing power increases, communication performance should increase accordingly in order to adequately balance the system.
Valentin Puente, José-Ángel Gregorio, Cruz Izu, Ramón Beivide, Fernando Vallejo
International Conference on Supercomputing5
1997 A flow control mechanism to avoid message deadlock in k-ary n-cube networks
abstract
We propose a flow control algorithm for k-ary n-cube networks which avoids the deadlock problems without using virtual channels. Some basic definitions and theorems are proposed in order to establish the necessary and sufficient conditions to verify that an algorithm is deadlock-free. Our proposal is based on a restriction of the virtual cut-through flow control rather than of the routing algorithm and it can be applied both over central buffers or edge buffers. A minimum free buffer space of two packets is required. The implementation complexity of the router according to Chien's (1993) model, is much easier and faster than using virtual channels. Network simulations considering the router complexity show the performance achieved by this new algorithm. The results display a latency improvement of 20% to 35% compared with the use of virtual channels depending on the load of the network.
Carmen Carrión 0001, Ramón Beivide, José-Ángel Gregorio, Fernando Vallejo
HiPC4
1995 Petri Net Modeling of Interconnection Networks for Massively Parallel Architectures
abstract
The analysls, design and evaluation of the interconnection subsystem for massively parallel arch i~ectures is norm ally carried out using computer simulation tools, requiring
José-Ángel Gregorio, Fernando Vallejo, Ramón Beivide, Carmen Carrión 0001
International Conference on Supercomputing2
1994 Shared Memory Multimicroprocessor Operating System with an Extended Petri Net Model
abstract
We propose a methodology for programming multiprocessor event-driven systems. This methodology is based on two programming levels: the task level, which involves programming the basic actions that may be executed in the system as units with a single control thread; and the job level, on which parallel programs to be executed by the complete multiprocessor system are developed. We also present the structure and implementation of an operating system designed as the programming support for software development under the proposed methodology. The model that has been chosen for the representation of the system software is based on an extended Petri net, which provides a well-established conceptual model for the development of the tasks, thus allowing a totally independent and generic development. This model also facilitates job-level programming, since the Petri net is a very powerful description tool for the parallel program.>
Fernando Vallejo, José-Ángel Gregorio, Michael González Harbour, José M. Drake
IEEE Trans. Parallel Distributed Syst.1