EDBT 2026 Demo / reviewers in the wild / expert
Miguel Valero-García
dblp:79/6916
· DBLP profile ↗
14ranked-venue papers
3as first author
0since 2021 · last 2002
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Interconnection networks and networks-on-chip · 51% Embedded and real-time systems · 30% Parallel and multicore computing · 7% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Interconnection networks and networks-on-chip › graph embedding
network topology embedding |
0.0 | 2 | 2002 | Hypercube Algorithms on Mesh Connected Multicomputers · IEEE Trans. Parallel Distributed Syst. 2002 Executing Algorithms with Hypercube Topology on Torus Multicomputers · IEEE Trans. Parallel Distributed Syst. 1995 |
Embedded and real-time systems › real-time communication
message scheduling |
0.0 | 1 | 2002 | Hypercube Algorithms on Mesh Connected Multicomputers · IEEE Trans. Parallel Distributed Syst. 2002 |
Compilers and program optimization › parallelization › automatic parallelization
loop parallelization |
0.0 | 1 | 1995 | Loop Transformation Using Nonunimodular Matrices · IEEE Trans. Parallel Distributed Syst. 1995 |
Compilers and program optimization
loop transformation |
0.0 | 1 | 1995 | Loop Transformation Using Nonunimodular Matrices · IEEE Trans. Parallel Distributed Syst. 1995 |
Interconnection networks and networks-on-chip › network topology › static interconnection networks
mesh and torus networks |
0.0 | 1 | 2002 | Hypercube Algorithms on Mesh Connected Multicomputers · IEEE Trans. Parallel Distributed Syst. 2002 |
Processor architecture and microarchitecture › parallel computer organization
systolic array design |
0.0 | 1 | 1989 | Systematic Hardware Adaptation of Systolic Algorithms · ISCA 1989 |
Parallel and multicore computing › parallel algorithms › parallel algorithm design
parallel algorithm mapping |
0.0 | 1 | 1995 | Executing Algorithms with Hypercube Topology on Torus Multicomputers · IEEE Trans. Parallel Distributed Syst. 1995 |
Parallel and multicore computing
parallelizing compiler |
0.0 | 1 | 1995 | Loop Transformation Using Nonunimodular Matrices · IEEE Trans. Parallel Distributed Syst. 1995 |
Processor architecture and microarchitecture › pipelining
pipelined functional units |
0.0 | 1 | 1989 | Systematic Hardware Adaptation of Systolic Algorithms · ISCA 1989 |
Methods — techniques the papers use, named apart from their topics
message scheduling · 0.0embedding · 0.0communication pipelining · 0.0hermite normal form · 0.0graph embedding · 0.0transformation rules · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2002 | Hypercube Algorithms on Mesh Connected MulticomputersabstractA new methodology named CALMANT (CC-cube Algorithms on Meshes and Tori) for mapping a type of algorithm that we call CC-cube algorithm onto multicomputers with hypercube, mesh, or torus interconnection topology is proposed. This methodology is suitable when the initial problem can be expressed as a set of processes that communicate through a hypercube topology (a CC-cube algorithm). There are many important algorithms that fit into the CC-cube type. CALMANT is based on three different techniques: (a) the standard embedding to assign the processes of the algorithm to the nodes of the mesh multicomputer; (b) the communication pipelining technique to increase the level of communication parallelism inherent in the CC-cube algorithms; and (c) optimal message-scheduling algorithms proposed in this work in order to avoid conflicts and minimizing in this way the communication time. Although CALMANT is proposed for multicomputers with different interconnection network topologies, the paper only focuses on the particular case of meshes. Luis Díaz de Cerio, Miguel Valero-García, Antonio González 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2001 | Implementing the one-sided Jacobi method on a 2D/3D mesh multicomputer
Dolors Royo, Miguel Valero-García, Antonio González 0001 |
Parallel Comput. | 2 |
| 2000 | Complete Exchange Algorithms for Meshes and Tori Using a Systematic Approach (Research Note)
Luis Díaz de Cerio, Miguel Valero-García, Antonio González 0001 |
Euro-Par | 2 |
| 1999 | Low Communication Overhead Jacobi Algorithms for Eigenvalues Computation on Hypercubes
Dolors Royo, Antonio González 0001, Miguel Valero-García |
J. Supercomput. | 3 |
| 1998 | Divide-and-Conquer Algorithms on Two-Dimensional Meshes
Miguel Valero-García, Antonio González 0001, Luis Díaz de Cerio, Dolors Royo |
Euro-Par | 1 |
| 1998 | A Method for Exploiting Communication/Computation Overlap in Hypercubes
Luis Díaz de Cerio, Miguel Valero-García, Antonio González 0001 |
Parallel Comput. | 2 |
| 1997 | A Methodology for User-Oriented Scalability AnalysisabstractScalability analysis provides information about the effectiveness of increasing the number of resources of a parallel system. Several methods have been proposed which use different approaches to provide this information. This paper presents a family of analysis methods oriented to the user. The methods in this family should assist the user in estimating the benefits when increasing the system size. The key issue in the proposal is the appropriate combination of a scaling model, which reflects the way the users utilize an increasing number of resources, and a figure of merit that the user wants to improve with the larger system. Another important element in the proposal is the approach to characterize the scalability, which enables quick visual analyses and comparisons. Finally, three concrete examples of methods belonging to the proposed family are introduced in this paper. Dolors Royo, Miguel Valero-García, Antonio González 0001, Carme Mari |
ASAP | 2 |
| 1995 | Loop Transformation Using Nonunimodular MatricesabstractLinear transformations are widely used to vectorize and parallelize loops. A subset of these transformations are unimodular transformations. When a unimodular transformation is used, the exact bounds of the transformed loop nest are easily computed and the steps of the loops are equal to 1. Unimodular loop transformations have been widely used since they permit the implementation of many useful loop transformations. Recently, nonunimodular transformations have been proposed to reduce communication requirements or to use the memory hierarchy efficiently. The methods used for unimodular transformations do not work in the case of nonunimodular transformations, since they do not produce the exact bounds of the transformed loop nest. In this paper, we present a method for nested loop transformation which gives the exact bounds for both unimodular and nonunimodular transformations. The basic idea is to use the Hermite Normal Form (HNF) of the transformation matrix.> Agustín Fernández, José María Llabería, Miguel Valero-García |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 1995 | Executing Algorithms with Hypercube Topology on Torus MulticomputersabstractMany parallel algorithms use hypercubes as the communication topology among their processes. When such algorithms are executed on hypercube multicomputers the communication cost is kept minimum since processes can be allocated to processors in such a way that only communication between neighbor processors is required. However, the scalability of hypercube multicomputers is constrained by the fact that the interconnection cost-per-node increases with the total number of nodes. From scalability point of view, meshes and toruses are more interesting classes of interconnection topologies. This paper focuses on the execution of algorithms with hypercube communication topology on multicomputers with mesh or torus interconnection topologies. The proposed approach is based on looking at different embeddings of hypercube graphs onto mesh or torus graphs. The paper concentrates on toruses since an already known embedding, which is called standard embedding, is optimal for meshes. In this paper, an embedding of hypercubes onto toruses of any given dimension is proposed. This novel embedding is called xor embedding. The paper presents a set of performance figures for both the standard and the xor embeddings and shows that the latter outperforms the former for any torus. In addition, it is proven that for a one-dimensional torus (a ring) the xor embedding is optimal in the sense that it minimizes the execution time of a class of parallel algorithms with hypercube topology. This class of algorithms is frequently found in real applications, such as FFT and some class of sorting algorithms.> Antonio González 0001, Miguel Valero-García, Luis Díaz de Cerio |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 1993 | The Xor embedding: An embedding of hypercubes onto rings and torusesabstractMany parallel algorithms use hypercubes as the communication topology among processes, which make them suitable to be executed on a hypercube multicomputer. In this way the communication cost is kept to a minimum since processes can be allocated to processors in such a way that only communication between neighbor processors is required. However, the scalability of hypercube multicomputer is constrained by the fact that the interconnection cost per node increases with the total number of nodes. From the point of view of scalability, meshes and toruses are a more interesting class of interconnection topologies. In this paper the authors propose an embedding of hypercubes onto toruses of any given dimension, incuding one-dimensional toruses which are also called rings. They also prove that the embedding is optimal in the sense that it minimizes the execution time on a ring of a class of parallel algorithms frequently found in real applications, such as FFT and some class of sorting algorithms.> Antonio González 0001, Miguel Valero-García |
ASAP | 2 |
| 1991 | Transformation of systolic algorithms for interleaving partitionsabstractA systematic method to map systolic problems onto multicomputers is presented. A systolic problem is a problem for which it is possible to design a systolic algorithm. This method selects and transforms the systolic algorithm into a parallel algorithm with high granularity. The communications requirements are reduced and the performance can be increased. The proposed scheme requires a classification of dependences and it is based on the interleaved execution of several partitions of the systolic algorithm. The code to be executed in a processing element of the multicomputer system is obtained through application of the proposed systematic transformations to the original sequential code. By applying this method to algebraic path problem (APP), the authors illustrate the main features, and several performance measures for a torus of transputers system are presented, considering the various algorithms which are unified by the APP.> Agustín Fernández, José María Llabería, Juan J. Navarro, Miguel Valero-García |
ASAP | 4 |
| 1991 | Performance evaluation of transputer systems with linear algebra problems
Agustín Fernández, José María Llabería, Juan J. Navarro, Miguel Valero-García |
Microprocessing and Microprogramming | 4 |
| 1990 | Implementation of systolic algorithms using pipelined functional unitsabstractThe authors present a method to implement systolic algorithms (SAs) using pipelined functional units (PFUs). This kind of unit makes it possible to improve the throughput of a processor because of the possibility of initiating a new operation before the previous one has been completed. The method permits transformation of a SA so that it can be efficiently executed using PFUs. The method is based on two temporal transformations (slowdown and retiming) and one spatial transformation (coalescing). The temporal transformations permit the modification of the SA in such a way that dependences established by the PFU are preserved. The spatial transformation improves the hardware utilization. The method was applied to 1-D SAs with data contraflow. To demonstrate the effectiveness of the method, the authors describe an efficient implementation of a non-time-homogeneous SA with data contraflow for QR decomposition.> Miguel Valero-García, Juan J. Navarro, José María Llabería, Mateo Valero |
ASAP | 1 |
| 1989 | Systematic Hardware Adaptation of Systolic AlgorithmsabstractIn this paper we propose a methodology to adapt Systolic Algorithms to the hardware selected for their implementation. Systolic Algorithms obtained can be efficiently implemented using Pipelined Functional Units. The methodology is based on two transformation rules. These rules are applied to an initial Systolic Algorithm, possibly obtained through one of the design methodologies proposed by other authors. Parameters for these transformations are obtained from the specification of the hardware to be used. The methodology has been particularized in the case of one-dimensional Systolic Algorithms with data contraflow. Miguel Valero-García, Juan J. Navarro, José María Llabería, Mateo Valero |
ISCA | 1 |