VLDB 2026 Research / reviewers in the wild / expert
Juan Carlos Díaz Martín
dblp:55/5695
· DBLP profile ↗
14ranked-venue papers
2as first author
0since 2021 · last 2020
0000-0002-8435-3844ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 62% Parallel and multicore computing · 38% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation › cost modeling
communication cost modeling |
0.3 | 1 | 2017 | Model-Based Estimation of the Communication Cost of Hybrid Data-Parallel Applications on Heterogeneous Clusters · IEEE Trans. Parallel Distributed Syst. 2017 |
Parallel and multicore computing › data parallelism
data-parallel applications |
0.1 | 1 | 2017 | Model-Based Estimation of the Communication Cost of Hybrid Data-Parallel Applications on Heterogeneous Clusters · IEEE Trans. Parallel Distributed Syst. 2017 |
Parallel and multicore computing
task partitioning |
0.1 | 1 | 2017 | Model-Based Estimation of the Communication Cost of Hybrid Data-Parallel Applications on Heterogeneous Clusters · IEEE Trans. Parallel Distributed Syst. 2017 |
Methods — techniques the papers use, named apart from their topics
t-lop model · 0.3functional performance models · 0.3HLogGP model · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Training deep neural networks: a static load balancing approach
Sergio Moreno-Álvarez, Juan Mario Haut, Mercedes Eugenia Paoletti, Juan A. Rico-Gallego, Juan Carlos Díaz Martín, Javier Plaza |
J. Supercomput. | 5 |
| 2020 | A tool to assess the communication cost of parallel kernels on heterogeneous platforms
Juan A. Rico-Gallego, Sergio Moreno-Álvarez, Juan Carlos Díaz Martín, Alexey L. Lastovetsky |
J. Supercomput. | 3 |
| 2019 | Analytical Communication Performance Models as a metric in the partitioning of data-parallel kernels on heterogeneous platforms
Juan A. Rico-Gallego, Juan Carlos Díaz Martín, Carmen Calvo-Jurado, Sergio Moreno-Álvarez, Juan-Luis García Zapata |
J. Supercomput. | 2 |
| 2017 | Formal modeling and performance evaluation of a run-time rank remapping technique in Broadcast, Allgather and Allreduce MPI collective operationsabstractMPI collective operations are implemented using a variety of algorithms which define different communication patterns between the ranks involved in the operation. The performance of these algorithms in multi-core clusters highly depends on the mapping of the ranks to the system processors due to the uneven capabilities of shared memory and network channels. The hierarchical design of these algorithms contributes to use optimally the communication channels. Nevertheless, common hierarchical algorithms have shown themselves, for some collectives as allgather, inefficient and even impracticable. This paper analyzed the reasons for that and works out an alternate approach through performance modeling. Such approach, departing from the a priori knowledge of a regular mapping as round-robin or sequential, and keeping the original algorithm unmodified, switches at run time the rank to process mapping into another regular mapping that reduces network traffic. The methodology is evaluated with three collectives and their underlying algorithms, showing speedups of up to 5x in the Binomial Tree or 3x in Ring algorithms compared to unfavorable mappings. Jesús M. Álvarez-Llorente, Juan Carlos Díaz Martín, Juan A. Rico-Gallego |
CCGrid | 2 |
| 2017 | Model-Based Estimation of the Communication Cost of Hybrid Data-Parallel Applications on Heterogeneous ClustersabstractHeterogeneous systems composed of CPUs and accelerators sharing communication channels of different performance are getting mainstream in HPC but, at the same time, they show a complexity that makes it difficult to optimize the deployment of a data parallel application. Recent analytical tools such as Functional Performance Models, combined with advanced partitioning algorithms, manage to achieve a balanced configuration by distributing the workload unevenly, according to the performance of the different processing units. Unfortunately, such uneven distribution of the computation load leads to communication unbalances that, very often, render worthless the previous workload balancing efforts. Finding the optimal communication scheme without expensive testing on the executing platform requires an analytical approach to the estimation of the communication cost of different configurations of the application. With this goal in mind, we propose and discuss an extension of the t-Lop communication performance model to cover heterogeneous architectures. In order to provide a quantitative assessment of this extended model, we conduct experiments with two representative computational kernels, the SUMMA algorithm and the 2D wave equation solver. The t-Lop predictions are compared against the HLogGP model and the observed costs for a variety of configurations, hardware resources and problem sizes. Juan A. Rico-Gallego, Alexey L. Lastovetsky, Juan Carlos Díaz Martín |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Extending τ-Lop to model concurrent MPI communications in multicore clusters
Juan A. Rico-Gallego, Juan Carlos Díaz Martín, Alexey L. Lastovetsky |
Future Gener. Comput. Syst. | 2 |
| 2015 | τ-Lop: Modeling performance of shared memory MPI
Juan A. Rico-Gallego, Juan Carlos Díaz Martín |
Parallel Comput. | 2 |
| 2013 | On the performance of concurrent transfers in collective algorithmsabstractInter- and intra-machine MPI collective operations in current multicore clusters are essentially different, and therefore their performance modelling ask for different approaches. Inside a multicore each individual message transmission in a collective operation flow in parallel with others, but sharing the bandwidth of the main memory channel. Current models ignore this issue, making errors like giving the same cost estimation to quite different collective algorithms. We outline a new cost model focused on shared channels. Juan A. Rico-Gallego, Juan Carlos Díaz Martín |
EuroMPI | 2 |
| 2012 | Improving Collectives by User Buffer Relocation
Juan A. Rico-Gallego, Juan Carlos Díaz Martín, Carolina Gómes-Tostón Gutierrez, Álvaro Cortés Fácila |
EuroMPI | 2 |
| 2012 | A geometric algorithm for winding number computation with complexity analysis
Juan-Luis García Zapata, Juan Carlos Díaz Martín |
J. Complex. | 2 |
| 2011 | Performance Evaluation of Thread-Based MPI in Shared Memory
Juan A. Rico-Gallego, Juan Carlos Díaz Martín |
EuroMPI | 2 |
| 2008 | A Network Service for DSP Multicomputers
Juan A. Rico-Gallego, Jesús M. Álvarez-Llorente, Juan Carlos Díaz Martín, Francisco J. Perogil-Duque |
ICA3PP | 3 |
| 2001 | Robust voice recognition as a distributed serviceabstractCurrent voice recognition systems tend to be implemented as embedded proprietary solutions. This model is not suitable for the growing complexities of present and future developments: It is single-user, it is non portable, and it assumes the workstation model, where all the CPU resources are supposed to he locally available. This work researches how a high performance speech recognition system can be redesigned and implemented as a time-critical network service with three main design goals: Scalability, predictability and POSIX portability. The whole idea has been tested by rebuilding IVORY, a well known robust desktop voice recognition methodology, as a distributed service. Also, potential fields of application are identified. Juan Carlos Díaz Martín, José Manuel Rodríguez García, Juan-Luis García Zapata, Pedro Gómez-Vilda |
ETFA (2) | 1 |
| 2001 | DIARCA: a component approach to voice recognitionabstractAbstract Current voice recognition systems tend to be implemented asa PC desktop facility. This model is not suitable for thegrowing complexities of present and future developments: Itis single-user, it is non portable, and it assumes theworkstation model, where all the CPU resources are supposedto be locally available. This work researches how a highperformance speech recognition system can be redesigned andimplemented as a time-critical network service shared throughordinary data transmission media with three main designgoals: Scalability, predictability and POSIX portability. Thewhole idea has been tested by rebuilding IVORY, a wellknown robust desktop voice recognition methodology, as adistributed component. 1. Introduction While Speech Processing and Recognition is a fieldexperiencing a rapid and promising expansion, the operating-system environments for the desktop PC still typically lack oftrue real-time support. To overcome this limitation, currentspeech recognition systems are confident on the workstationprinciple: all the CPU resources are always available to theapplication where they are embedded. This approach shows amain limitation: Its growing computational complexity.IVORY ([1], [7]), a stand-alone speech recognition system ofisolated words, gives figures of computational complexityaround 21 Mflop/s. Though this load is easily assumed bycurrent CPU's, continuous speech can raise the computingpower demand one order of magnitude. Noise cancellationdemands up to five or six times the power of the recognitionitself. Furthermore, new applications of speech processingdemand much more computing power. For instance, tracking asingle speaker by the Microphone Arrays technique shows acomputational complexity near 166 Mflop/s ([6]). Thoughtoday's PC microprocessors claim peak execution ratesexceeding 1 Gflop/s, regular DSP algorithms rarely result insuch a high performance. In our view, desktop speechprocessing is -and will always be- strongly limited by itscomputational complexity, nowadays constrained to thecomputing power of the average personal computer. This work was founded by CICYT and Junta de Extremaduraunder the TIC99-0609 (DIARCA) and CICYTEX IPRR98A039 projects respectively.Distributed computing should change this scenery.Ongoing developments on component based softwareengineering makes possible to envision a remote service ofDSP computing power for speech processing. It would allowto bring both to the current desktop PC and to the futureinternet appliances the more advanced developments on thefield. This work investigates the distribution of speechrecognition in the context of DIARCA, a research projectwhose aim is two-fold. Firstly, to distribute IVORY with threedesign goals: Scalability, predictability and POSIX portability.Secondly, to extend the results in order to support microphonearray developments. This work is about the first goal.Figure 1. IVORY: a Robust Voice Recognition system Juan Carlos Díaz Martín, Juan-Luis García Zapata, José Manuel Rodríguez García, José F. Álvarez Salgado, Pablo Espada Bueno, Pedro Gómez-Vilda |
INTERSPEECH | 1 |