Miguel Valero-García

dblp:79/6916 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
0since 2021 · last 2002
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Interconnection networks and networks-on-chip · 51% Embedded and real-time systems · 30% Parallel and multicore computing · 7%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Interconnection networks and networks-on-chip › graph embedding
network topology embedding
0.022002
Hypercube Algorithms on Mesh Connected Multicomputers · IEEE Trans. Parallel Distributed Syst. 2002
Executing Algorithms with Hypercube Topology on Torus Multicomputers · IEEE Trans. Parallel Distributed Syst. 1995
Embedded and real-time systems › real-time communication
message scheduling
0.012002
Hypercube Algorithms on Mesh Connected Multicomputers · IEEE Trans. Parallel Distributed Syst. 2002
Compilers and program optimization › parallelization › automatic parallelization
loop parallelization
0.011995
Loop Transformation Using Nonunimodular Matrices · IEEE Trans. Parallel Distributed Syst. 1995
Compilers and program optimization
loop transformation
0.011995
Loop Transformation Using Nonunimodular Matrices · IEEE Trans. Parallel Distributed Syst. 1995
Interconnection networks and networks-on-chip › network topology › static interconnection networks
mesh and torus networks
0.012002
Hypercube Algorithms on Mesh Connected Multicomputers · IEEE Trans. Parallel Distributed Syst. 2002
Processor architecture and microarchitecture › parallel computer organization
systolic array design
0.011989
Systematic Hardware Adaptation of Systolic Algorithms · ISCA 1989
Parallel and multicore computing › parallel algorithms › parallel algorithm design
parallel algorithm mapping
0.011995
Executing Algorithms with Hypercube Topology on Torus Multicomputers · IEEE Trans. Parallel Distributed Syst. 1995
Parallel and multicore computing
parallelizing compiler
0.011995
Loop Transformation Using Nonunimodular Matrices · IEEE Trans. Parallel Distributed Syst. 1995
Processor architecture and microarchitecture › pipelining
pipelined functional units
0.011989
Systematic Hardware Adaptation of Systolic Algorithms · ISCA 1989

Methods — techniques the papers use, named apart from their topics

message scheduling · 0.0embedding · 0.0communication pipelining · 0.0hermite normal form · 0.0graph embedding · 0.0transformation rules · 0.0
YearPublicationVenuePosition
2002 Hypercube Algorithms on Mesh Connected Multicomputers
abstract
A new methodology named CALMANT (CC-cube Algorithms on Meshes and Tori) for mapping a type of algorithm that we call CC-cube algorithm onto multicomputers with hypercube, mesh, or torus interconnection topology is proposed. This methodology is suitable when the initial problem can be expressed as a set of processes that communicate through a hypercube topology (a CC-cube algorithm). There are many important algorithms that fit into the CC-cube type. CALMANT is based on three different techniques: (a) the standard embedding to assign the processes of the algorithm to the nodes of the mesh multicomputer; (b) the communication pipelining technique to increase the level of communication parallelism inherent in the CC-cube algorithms; and (c) optimal message-scheduling algorithms proposed in this work in order to avoid conflicts and minimizing in this way the communication time. Although CALMANT is proposed for multicomputers with different interconnection network topologies, the paper only focuses on the particular case of meshes.
Luis Díaz de Cerio, Miguel Valero-García, Antonio González 0001
IEEE Trans. Parallel Distributed Syst.2
2001 Implementing the one-sided Jacobi method on a 2D/3D mesh multicomputer
Dolors Royo, Miguel Valero-García, Antonio González 0001
Parallel Comput.2
2000 Complete Exchange Algorithms for Meshes and Tori Using a Systematic Approach (Research Note)
Luis Díaz de Cerio, Miguel Valero-García, Antonio González 0001
Euro-Par2
1999 Low Communication Overhead Jacobi Algorithms for Eigenvalues Computation on Hypercubes
Dolors Royo, Antonio González 0001, Miguel Valero-García
J. Supercomput.3
1998 Divide-and-Conquer Algorithms on Two-Dimensional Meshes
Miguel Valero-García, Antonio González 0001, Luis Díaz de Cerio, Dolors Royo
Euro-Par1
1998 A Method for Exploiting Communication/Computation Overlap in Hypercubes
Luis Díaz de Cerio, Miguel Valero-García, Antonio González 0001
Parallel Comput.2
1997 A Methodology for User-Oriented Scalability Analysis
abstract
Scalability analysis provides information about the effectiveness of increasing the number of resources of a parallel system. Several methods have been proposed which use different approaches to provide this information. This paper presents a family of analysis methods oriented to the user. The methods in this family should assist the user in estimating the benefits when increasing the system size. The key issue in the proposal is the appropriate combination of a scaling model, which reflects the way the users utilize an increasing number of resources, and a figure of merit that the user wants to improve with the larger system. Another important element in the proposal is the approach to characterize the scalability, which enables quick visual analyses and comparisons. Finally, three concrete examples of methods belonging to the proposed family are introduced in this paper.
Dolors Royo, Miguel Valero-García, Antonio González 0001, Carme Mari
ASAP2
1995 Loop Transformation Using Nonunimodular Matrices
abstract
Linear transformations are widely used to vectorize and parallelize loops. A subset of these transformations are unimodular transformations. When a unimodular transformation is used, the exact bounds of the transformed loop nest are easily computed and the steps of the loops are equal to 1. Unimodular loop transformations have been widely used since they permit the implementation of many useful loop transformations. Recently, nonunimodular transformations have been proposed to reduce communication requirements or to use the memory hierarchy efficiently. The methods used for unimodular transformations do not work in the case of nonunimodular transformations, since they do not produce the exact bounds of the transformed loop nest. In this paper, we present a method for nested loop transformation which gives the exact bounds for both unimodular and nonunimodular transformations. The basic idea is to use the Hermite Normal Form (HNF) of the transformation matrix.>
Agustín Fernández, José María Llabería, Miguel Valero-García
IEEE Trans. Parallel Distributed Syst.3
1995 Executing Algorithms with Hypercube Topology on Torus Multicomputers
abstract
Many parallel algorithms use hypercubes as the communication topology among their processes. When such algorithms are executed on hypercube multicomputers the communication cost is kept minimum since processes can be allocated to processors in such a way that only communication between neighbor processors is required. However, the scalability of hypercube multicomputers is constrained by the fact that the interconnection cost-per-node increases with the total number of nodes. From scalability point of view, meshes and toruses are more interesting classes of interconnection topologies. This paper focuses on the execution of algorithms with hypercube communication topology on multicomputers with mesh or torus interconnection topologies. The proposed approach is based on looking at different embeddings of hypercube graphs onto mesh or torus graphs. The paper concentrates on toruses since an already known embedding, which is called standard embedding, is optimal for meshes. In this paper, an embedding of hypercubes onto toruses of any given dimension is proposed. This novel embedding is called xor embedding. The paper presents a set of performance figures for both the standard and the xor embeddings and shows that the latter outperforms the former for any torus. In addition, it is proven that for a one-dimensional torus (a ring) the xor embedding is optimal in the sense that it minimizes the execution time of a class of parallel algorithms with hypercube topology. This class of algorithms is frequently found in real applications, such as FFT and some class of sorting algorithms.>
Antonio González 0001, Miguel Valero-García, Luis Díaz de Cerio
IEEE Trans. Parallel Distributed Syst.2
1993 The Xor embedding: An embedding of hypercubes onto rings and toruses
abstract
Many parallel algorithms use hypercubes as the communication topology among processes, which make them suitable to be executed on a hypercube multicomputer. In this way the communication cost is kept to a minimum since processes can be allocated to processors in such a way that only communication between neighbor processors is required. However, the scalability of hypercube multicomputer is constrained by the fact that the interconnection cost per node increases with the total number of nodes. From the point of view of scalability, meshes and toruses are a more interesting class of interconnection topologies. In this paper the authors propose an embedding of hypercubes onto toruses of any given dimension, incuding one-dimensional toruses which are also called rings. They also prove that the embedding is optimal in the sense that it minimizes the execution time on a ring of a class of parallel algorithms frequently found in real applications, such as FFT and some class of sorting algorithms.>
Antonio González 0001, Miguel Valero-García
ASAP2
1991 Transformation of systolic algorithms for interleaving partitions
abstract
A systematic method to map systolic problems onto multicomputers is presented. A systolic problem is a problem for which it is possible to design a systolic algorithm. This method selects and transforms the systolic algorithm into a parallel algorithm with high granularity. The communications requirements are reduced and the performance can be increased. The proposed scheme requires a classification of dependences and it is based on the interleaved execution of several partitions of the systolic algorithm. The code to be executed in a processing element of the multicomputer system is obtained through application of the proposed systematic transformations to the original sequential code. By applying this method to algebraic path problem (APP), the authors illustrate the main features, and several performance measures for a torus of transputers system are presented, considering the various algorithms which are unified by the APP.>
Agustín Fernández, José María Llabería, Juan J. Navarro, Miguel Valero-García
ASAP4
1991 Performance evaluation of transputer systems with linear algebra problems
Agustín Fernández, José María Llabería, Juan J. Navarro, Miguel Valero-García
Microprocessing and Microprogramming4
1990 Implementation of systolic algorithms using pipelined functional units
abstract
The authors present a method to implement systolic algorithms (SAs) using pipelined functional units (PFUs). This kind of unit makes it possible to improve the throughput of a processor because of the possibility of initiating a new operation before the previous one has been completed. The method permits transformation of a SA so that it can be efficiently executed using PFUs. The method is based on two temporal transformations (slowdown and retiming) and one spatial transformation (coalescing). The temporal transformations permit the modification of the SA in such a way that dependences established by the PFU are preserved. The spatial transformation improves the hardware utilization. The method was applied to 1-D SAs with data contraflow. To demonstrate the effectiveness of the method, the authors describe an efficient implementation of a non-time-homogeneous SA with data contraflow for QR decomposition.>
Miguel Valero-García, Juan J. Navarro, José María Llabería, Mateo Valero
ASAP1
1989 Systematic Hardware Adaptation of Systolic Algorithms
abstract
In this paper we propose a methodology to adapt Systolic Algorithms to the hardware selected for their implementation. Systolic Algorithms obtained can be efficiently implemented using Pipelined Functional Units. The methodology is based on two transformation rules. These rules are applied to an initial Systolic Algorithm, possibly obtained through one of the design methodologies proposed by other authors. Parameters for these transformations are obtained from the specification of the hardware to be used. The methodology has been particularized in the case of one-dimensional Systolic Algorithms with data contraflow.
Miguel Valero-García, Juan J. Navarro, José María Llabería, Mateo Valero
ISCA1