EDBT 2026 Demo / reviewers in the wild / expert
José Carlos Cabaleiro
dblp:24/6490 · also José Carlos Cabaleiro Domínguez
· DBLP profile ↗
35ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0002-5674-5162ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 5 since 2021Security and privacy · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A new thread-level speculative automatic parallelization model and library based on duplicate code executionabstractAbstract Loop-efficient automatic parallelization has become increasingly relevant due to the growing number of cores in current processors and the programming effort needed to parallelize codes in these systems efficiently. However, automatic tools fail to extract all the available parallelism in irregular loops with indirections, race conditions or potential data dependency violations, among many other possible causes. One of the successful ways to automatically parallelize these loops is the use of speculative parallelization techniques. This paper presents a new model and the corresponding C++ library that supports the speculative automatic parallelization of loops in shared memory systems, seeking competitive performance and scalability while keeping user effort to a minimum. The primary speculative strategy consists of redundantly executing chunks of loop iterations in a duplicate fashion. Namely, each chunk is executed speculatively in parallel to obtain results as soon as possible and sequentially in a different thread to validate the speculative results. The implementation uses C++11 threads and it makes intensive use of templates and advanced multithreading techniques. An evaluation based on various benchmarks confirms that our proposal provides a competitive level of performance and scalability. Millán Álvarez Martínez, Basilio B. Fraguela, José Carlos Cabaleiro, Francisco F. Rivera |
J. Supercomput. | 3 |
| 2023 | Digital forensic analysis of the private mode of browsers on AndroidabstractThe smartphone has become an essential electronic device in our daily lives. We carry our most precious and important data on it, from family videos of the last few years to credit card information so that we can pay with our phones. In addition, in recent years, mobile devices have become the preferred device for surfing the web, already representing more than 50% of Internet traffic. As one of the devices we spend the most time with throughout the day, it is not surprising that we are increasingly demanding a higher level of privacy. One of the measures introduced to help us protect our data by isolating certain activities on the Internet is the private mode integrated in most modern browsers. Of course, this feature is not new, and has been available on desktop platforms for more than a decade. Reviewing the literature, one can find several studies that test the correct functioning of the private mode on the desktop. However, the number of studies conducted on mobile devices is incredibly small. And not only is it small, but also most of them perform the tests using various emulators or virtual machines running obsolete versions of Android. Therefore, in this paper we apply the methodology we presented in a previous work to Google Chrome, Brave, Mozilla Firefox, and Tor Browser running on a tablet with Android 13 and on two virtual devices created with Android Emulator. The results confirm that these browsers do not store information about the browsing performed in private mode in the file system. However, the analysis of the volatile memory made it possible to recover the username and password used to log in to a website or the keywords typed in a search engine, even after the devices had been rebooted. Xosé Fernández-Fuentes, Tomás F. Pena, José Carlos Cabaleiro |
Comput. Secur. | 3 |
| 2022 | Digital forensic analysis methodology for private browsing: Firefox and Chrome on Linux as a case studyabstractThe web browser has become one of the basic tools of everyday life. A tool that is increasingly used to manage personal information. This has led to the introduction of new privacy options by the browsers, including private mode. In this paper, a methodology to explore the effectiveness of the private mode included in most browsers is proposed. A browsing session was designed and conducted in Mozilla Firefox and Google Chrome running on four different Linux environments. After analyzing the information written to disk and the information available in memory, it can be observed that Firefox and Chrome did not store any browsing-related information on the hard disk. However, memory analysis reveals that a large amount of information could be retrieved in some of the environments tested. For example, for the case where the browsers were executed in a VMware virtual machine, it was possible to retrieve most of the actions performed, from the keywords entered in a search field to the username and password entered to log in to a website, even after restarting the computer. In contrast, when Firefox was run on a slightly hardened non-virtualized Linux, it was not possible to retrieve any browsing-related artifacts after the browser was closed. Xosé Fernández-Fuentes, Tomás F. Pena, José Carlos Cabaleiro |
Comput. Secur. | 3 |
| 2022 | CIMAR, NIMAR, and LMMA: Novel algorithms for thread and memory migrations in user space on NUMA systems using hardware countersabstractThis paper introduces two novel algorithms for thread migrations, named CIMAR (Core-aware Interchange and Migration Algorithm with performance Record –IMAR–) and NIMAR (Node-aware IMAR), and a new algorithm for the migration of memory pages, LMMA (Latency-based Memory pages Migration Algorithm), in the context of Non-Uniform Memory Access (NUMA) systems. This kind of system has complex memory hierarchies that present a challenging problem in extracting the best possible performance, where thread and memory mapping play a critical role. The presented algorithms gather and process the information provided by hardware counters to make decisions about the migrations to be performed, trying to find the optimal mapping. They have been implemented as a user space tool that looks for improving the system performance, particularly in, but not restricted to, scenarios where multiple programs with different characteristics are running. This approach has the advantage of not requiring any modification on the target programs or the Linux kernel while keeping a low overhead. Two different benchmark suites have been used to validate our algorithms: The NAS parallel benchmark, mainly devoted to computational routines, and the LevelDB database benchmark focused on read–write operations. These benchmarks allow us to illustrate the influence of our proposal in these two important types of codes. Note that those codes are state-of-the-art implementations of the routines, so few improvements could be initially expected. Experiments have been designed and conducted to emulate three different scenarios: a single program running in the system with full resources, an interactive server where multiple programs run concurrently varying the availability of resources, and a queue of tasks where granted resources are limited. The proposed algorithms have been able to produce significant benefits, especially in systems with higher latency penalties for remote accesses. When more than one benchmark is executed simultaneously, performance improvements have been obtained, reducing execution times up to 60%. In this kind of situation, the behaviour of the system is more critical, and the NUMA topology plays a more relevant role. Even in the worst case, when isolated benchmarks are executed using the whole system, that is, just one task at a time, the performance is not degraded. Ruben Laso, Oscar G. Lorenzo, José Carlos Cabaleiro, Tomás F. Pena, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera |
Future Gener. Comput. Syst. | 3 |
| 2022 | A highly optimized skeleton for unbalanced and deep divide-and-conquer algorithms on multi-core clustersabstractAbstract Efficiently implementing the divide-and-conquer pattern of parallelism in distributed memory systems is very relevant, given its ubiquity, and difficult, given its recursive nature and the need to exchange tasks and data among the processors. This task is noticeably further complicated in the presence of multi-core systems, where hybrid parallelism must be exploited to attain the best performance, and when unbalanced and deep workloads are considered, as additional measures must be taken to load balance and avoid deep recursion problems. In this manuscript a parallel skeleton that fulfills all these requirements while providing high levels of usability is presented. In fact, the evaluation shows that our proposal is on average 415.32% faster than MPI codes and 229.18% faster than MPI + OpenMP benchmarks, while offering an average improvement in the programmability metrics of 131.04% over MPI alternatives and 155.18% over MPI + OpenMP solutions. Millán Álvarez Martínez, Basilio B. Fraguela, José Carlos Cabaleiro |
J. Supercomput. | 3 |
| 2021 | LBMA and IMAR2: Weighted lottery based migration strategies for NUMA multiprocessing serversabstractSummary Multicore NUMA systems present on‐board memory hierarchies and communication networks that influence performance when executing shared memory parallel codes. Characterizing this influence is complex, and understanding the effect of particular hardware configurations on different codes is of paramount importance. In this article, monitoring information extracted from hardware counters at runtime is used to characterize the behavior of each thread for an arbitrary number of multithreaded processes running in a multiprocessing environment. This characterization is given in terms of number of operations per second, operational intensity, and latency of memory accesses. We propose a runtime tool, executed in user space, that uses this information to guide two different thread migration strategies for improving execution efficiency by increasing locality and affinity without requiring any modification in the running codes. Different configurations of NAS Parallel OpenMP benchmarks running concurrently on multicore NUMA systems were used to validate the benefits of our proposal, in which up to four processes are running simultaneously. In more than the 95% of the executions of our tool, results outperform those of the operating system (OS) and produces up to 38% improvement in execution time over the OS for heterogeneous workloads, under different and realistic locality and affinity scenarios. Ruben Laso, Oscar G. Lorenzo, Francisco F. Rivera, José Carlos Cabaleiro, Tomás F. Pena, Juan Ángel Lorenzo del Castillo |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | IHP: a dynamic heterogeneous parallel scheme for iterative or time-step methods - image denoising as case study
Ruben Laso, José Carlos Cabaleiro, Francisco F. Rivera, M. Carmen Muñiz, José A. Álvarez-Dios |
J. Supercomput. | 2 |
| 2020 | Next-generation big data federation access control: A reference model
Feras M. Awaysheh, Mamoun Alazab, Maanak Gupta, Tomás F. Pena, José Carlos Cabaleiro |
Future Gener. Comput. Syst. | 5 |
| 2020 | TrustE-VC: Trustworthy Evaluation Framework for Industrial Connected Vehicles in the CloudabstractThe integration between cloud computing and vehicular ad hoc networks, namely, vehicular clouds (VCs), has become a significant research area. This integration was proposed to accelerate the adoption of intelligent transportation systems. The trustworthiness in VCs is expected to carry more computing capabilities that manage large-scale collected data. This trend requires a security evaluation framework that ensures data privacy protection, integrity of information, and availability of resources. To the best of our knowledge, this is the first study that proposes a robust trustworthiness evaluation of vehicular cloud for security criteria evaluation and selection. This article proposes three-level security features in order to develop effectiveness and trustworthiness in VCs. To assess and evaluate these security features, our evaluation framework consists of three main interconnected components: 1) an aggregation of the security evaluation values of the security criteria for each level; 2) a fuzzy multicriteria decision-making algorithm; and 3) a simple additive weight associated with the importance-performance analysis and performance rate to visualize the framework findings. The evaluation results of the security criteria based on the average performance rate and global weight suggest that data residency, data privacy, and data ownership are the most pressing challenges in assessing data protection in a VC environment. Overall, this article paves the way for a secure VC using an evaluation of effective security features and underscores directions and challenges facing the VC community. This article sheds light on the importance of security by design, emphasizing multiple layers of security when implementing industrial VCs. Mohammad Aladwan, Feras M. Awaysheh, Sadi Alawadi, Mamoun Alazab, Tomás F. Pena, José Carlos Cabaleiro |
IEEE Trans. Ind. Informatics | 6 |
| 2019 | Poster: A Pluggable Authentication Module for Big Data Federation ArchitectureabstractThis paper intends to propose a trustworthy model for authenticating users and services over a Big Data Federation deployment architecture. The main goal of this model is to provide a Single-Sign-on (SSO) approach for the latest Hadoop 3.x platform. To achieve this, a conceptual model is proposed combining Hadoop access control primitives and the Apache Knox framework. The paper provides various insights regarding the latest ongoing developments and open challenges in this domain. Feras M. Awaysheh, José Carlos Cabaleiro, Tomás F. Pena, Mamoun Alazab |
SACMAT | 2 |
| 2017 | EME: An Automated, Elastic and Efficient Prototype for Provisioning Hadoop Clusters On-demand
Feras M. Awaysheh, Tomás F. Pena, José Carlos Cabaleiro |
CLOSER | 3 |
| 2017 | Landing sites detection using LiDAR data on manycore systems
Oscar G. Lorenzo, Jorge Martínez Sánchez, David López Vilariño, Tomás F. Pena, José Carlos Cabaleiro, Francisco F. Rivera |
J. Supercomput. | 5 |
| 2014 | Multiobjective optimization technique based on monitoring information to increase the performance of thread migration on multicoresabstractMulticore systems present on-board memory hierarchies and communication networks that influence their performance when they execute shared memory parallel codes. Characterizing this influence is complex, and understanding the effect of particular hardware configurations on different codes is of paramount importance. In this paper, monitoring information extracted from hardware counters in runtime is used to characterize the behaviour of each thread in the parallel code in terms of three values: the number of floating point operations per second, the operational intensity, and the memory access latency. Note that these values characterize the Roofline Model with the inclusion of additional information about memory access latencies. We propose to use this information to guide thread migration strategies that improve the efficiency of the execution of the code by increasing locality and affinity. The idea behind this proposal is to use these three values as objective functions to be optimized as a multiobjective optimization problem. The proposed technique is an iterative method inspired in evolutive optimization algorithms. To this end, an individual utility function is defined to represent the relative importance of these values. This function is a weighted product that can be considered as representative of the performance of each parallel thread. Different configurations of the SAXPY and SDOT kernels on multicores were used to validate the benefits of the proposed thread migration strategies. The results show that our strategy produces improvements up to 25% in scenarios where locality and affinity are low, and negligible degradation is observed when they are high. The use of hardware counters produces low overheads when extracting monitoring information. Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Francisco F. Rivera |
CLUSTER | 3 |
| 2014 | A hardware counter-based toolkit for the analysis of memory accesses in SMPsabstractSUMMARY In this paper, a set of three hardware counter (HC)‐based tools to characterise memory access of parallel codes in Symmetric Multiprocessors (SMPs) is presented. This toolkit simplifies accessing and programming HCs, which are included in modern microprocessors. Hardware counters are used to obtain information about memory accesses in a parallel code at very low cost. This information is presented to the user in a friendly way. The first tool can be used to automatically monitor the memory accesses of a system and to analyse a code even if the source is not available. The second tool allows the user to insert in a source code, in a simple and transparent way, the instructions needed to monitor and manage HCs. This way, specific parts of the code can be analysed. The user can either add appropriate directives to a C code or use a graphical interface to select those parts of the code to be analysed. The tool takes this source file and automatically adds the monitoring code. The third tool takes the information gathered by the aforementioned tools, processes it and displays it graphically. This tool shows the information in a comprehensive and simple way, allowing the user to adjust the level of detail. The aim of these tools was to characterise the memory accesses of parallel codes in multicore systems, in which the cache hierarchy can greatly influence the performance. For illustrative purposes, these tools were used to carry out two case studies, a sparse matrix vector product and a dot product. These studies have been made in two different environments. Anyway, they can be used in almost any system as long as the necessary HCs are available.Copyright © 2013 John Wiley & Sons, Ltd. Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera |
Concurr. Comput. Pract. Exp. | 3 |
| 2014 | Modeling the performance of parallel applications using model selection techniquesabstractSUMMARY Nowadays, parallel architectures are changing so fast that there is a need for scalable and efficient tools to analyze and predict the performance of parallel applications. Analytical models are proved to be a useful approximation for characterizing parallel algorithms, but developing accurate analytical models is a hard issue, and, in general, they provide coarse performance predictions due to their intrinsic lack of accuracy. In this paper, we describe in detail the Tools for Instrumentation and Analysis (TIA) framework, an easy‐to‐use tool that automatically obtains accurate performance models by means of analytical expressions. This framework automatizes most of its internal tasks, reducing opportunities for human error, and it only requires the user to focus on the metrics and execution parameters that might influence the performance, those that should be considered in the modeling process. Its main advantage over other tools is that TIA uses model selection techniques that allow the automation of the modeling process. As a case of study, the use of TIA to obtain analytical models of different implementations of the broadcast collective communication in a cluster of multicores is shown. The results obtained by TIA are evaluated and compared with theoretical approaches based on the LogGP model. Copyright © 2013 John Wiley & Sons, Ltd. Diego Rodríguez Martínez, Vicente Blanco 0001, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera |
Concurr. Comput. Pract. Exp. | 3 |
| 2014 | 3DyRM: a dynamic roofline model including memory latency information
Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Francisco F. Rivera |
J. Supercomput. | 3 |
| 2013 | A flexible and dynamic page migration infrastructure based on hardware counters
Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, Francisco F. Rivera, Tomás F. Pena, José Carlos Cabaleiro |
J. Supercomput. | 5 |
| 2012 | Hardware Counters Based Analysis of Memory Accesses in SMPsabstractModern microprocessors incorporate Hardware Counters (HC) that provide useful information with low overhead. HC are not commonly used because of the lack of tools to get their information in an easy way. In this paper, a set of tools to simplify the accessing and programming of Intel Itanium 2 ™EARs (Event Address Registers) is presented. The aim of these tools is to characterise the memory accesses of parallel codes, in multicore systems, in which the cache hierarchy can greatly influence the performance. The first tool allows the user to insert in the code, in a simple and transparent way, the instructions needed to monitor and manage hardware counters. Two versions of this tool have been implemented. The first one is a command line tool that takes as input a C source file with appropriate directives and outputs it with the monitoring code added. The other one is a graphical interface that allows the user to select the parts of the code to analise. The second tool takes the information gathered by the monitored parallel code provided by the hardware counters and displays it graphically. This tool shows the information in a comprehensive but simple way, allowing the user to adjust the level of detail. These tools were used to carry out a study of parallel irregular codes. Although this study has been made in a specific environment, the tools here presented can be used in any system as long as it is based on hardware counters present in current processors. Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera |
ISPA | 3 |
| 2012 | Model Selection to Characterize Performance Using Genetic AlgorithmsabstractThe TIA modeling framework provides analytical models of the performance of parallel applications. The resulting models are obtained using model selection techniques and are accurate enough for various purposes. Its main drawback is that the completion time depends on the number of candidate models and, in some situations, it becomes critical. In this work, a genetic algorithm is proposed for reducing the time for searching of the best candidate model. The use of this genetic algorithm to obtain the performance model of the linear implementation of the broadcast collective communication in a cluster of multicores is shown. Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001 |
ISPA | 2 |
| 2012 | A Graphical Tool for Performance Analysis of Multicore Systems Based on the Roofline ModelabstractA tool to characterize the performance of parallel codes on multicore systems is presented in this paper. This tool allows the user to define the Roofline Model of the target system, to execute the code under study and to represent the performance results in the roofline plot. The final product is an easy to use tool to provide an insightful model which allows to determine, at a glance, performance issues like load balance, locality and those related to thread and memory allocation. Results show that this model provides practical information of the effects that degrade the performance of a code and gives hints to improve it. Francisco F. Rivera, Ramón Iglesias, Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, Tomás F. Pena, José Carlos Cabaleiro |
ISPA | 6 |
| 2011 | Estimating the effect of cache misses on the performance of parallel applications using analytical modelsabstractIn this paper a methodology to characterize the influence of cache misses on the performance of parallel applications is presented. This methodology is based on analytical models provided by the TIA framework. This framework obtains analytical models of given observable quantities by instrumenting the source code and applying model selection techniques. In particular, two metrics related with the performance are considered in this work: the number of cache misses and the elapsed time. Based on both models, the influence in terms of execution time due to the cache misses can be inferred. Two different versions of the parallel product of dense matrices are used as case of study. Diego Rodríguez Martínez, Vicente Blanco 0001, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera |
AICCSA | 3 |
| 2011 | Using accurate AIC-based performance models to improve the scheduling of parallel applications
Diego Rodríguez Martínez, Julio L. Albín, Tomás F. Pena, José Carlos Cabaleiro, Francisco F. Rivera, Vicente Blanco 0001 |
J. Supercomput. | 4 |
| 2011 | Analyzing the execution of sparse matrix-vector product on the Finisterrae SMP-NUMA system
Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Dora Blanco Heras, José Carlos Cabaleiro, Tomás F. Pena |
J. Supercomput. | 4 |
| 2010 | Performance Modeling of MPI Applications Using Model Selection TechniquesabstractA new method for obtaining models of the performance of parallel applications based on statistical analysis is presented in this paper. This method is based on the Akaike's information criterion (AIC) that provides an objective mechanism to rank different models by means of an experimental data fit. The input of the modeling process is a set of variables and parameters that can a priori influence the performance of the application. This set can be provided by the user. Using this information, the method automatically generates a set of candidate models. These models are fit to the experimental data and the AIC score of each model is calculated. The model with the best AIC score is selected as the best model. Also, using the AIC scores of all candidate models, useful statistical information is provided to help the user to evaluate the quality of the selected model, as well as indications of how to interactively improve this modeling process. As a first case of study, statistical models obtained for different implementations of the broadcast collective communication in Open MPI are shown. These models are very accurate, exceeding its adjustment to theoretical approaches based on the LogGP model. Finally, the NAS Parallel Benchmark is also characterized using this new method with good results in terms of accuracy. Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001 |
PDP | 2 |
| 2009 | Accurate analytical performance model of communications in MPI applicationsabstractThis paper presents a new LogP-based model, called LoOgGP, which allows an accurate characterization of MPI applications based on microbenchmark measurements. This new model is an extension of LogP for long messages in which both overhead and gap parameters perform a linear dependency with message size. The LoOgGP model has been fully integrated into a modelling framework to obtain statistical models of parallel applications, providing the analyst with an easy and automatic tool for LoOgGP parameter set assessment to characterize communications. The use of LoOgGP model to obtain a statistical performance model of an image deconvolution application is illustrated as a case of study. Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001 |
IPDPS | 2 |
| 2009 | Increasing data reuse of sparse algebra codes on simultaneous multithreading architecturesabstractAbstract In this paper the problem of the locality of sparse algebra codes on simultaneous multithreading (SMT) architectures is studied. In these kind of architectures many hardware structures are dynamically shared among the running threads. This puts a lot of stress on the memory hierarchy, and a poor locality, both inter‐thread and intra‐thread, may become a major bottleneck in the performance of a code. This behavior is even more pronounced when the code is irregular, which is the case of sparse matrix ones. Therefore, techniques that increase the locality of irregular codes on SMT architectures are important to achieve high performance. This paper proposes a data reordering technique specially tuned for these kind of architectures and codes. It is based on a locality model developed by the authors in previous works. The technique has been tested, first, using a simulator of a SMT architecture, and subsequently, on a real architecture as Intel's Hyper‐Threading. Important reductions in the number of cache misses have been achieved, even when the number of running threads grows. When applying the locality improvement technique, we also decrease the total execution time and improve the scalability of the code. Copyright © 2009 John Wiley & Sons, Ltd. Juan Carlos Pichel, Dora Blanco Heras, José Carlos Cabaleiro, Francisco F. Rivera |
Concurr. Comput. Pract. Exp. | 3 |
| 2007 | Software Tools for Performance Modeling of Parallel ProgramsabstractThis paper presents a framework based on a user driven methodology to obtain analytical models of MPI applications on parallel systems in a systematic and easy to use way. This methodology consists of two stages. In the first one, instrumentation of the source code is performed using CALL, which is a profiling tool for interacting with the code in an easy, simple and direct way. New features are added to CALL to obtain different performance metrics and store the performance information in XML files. Using this information, an analytical model of the performance behavior is obtained in the second stage by means of R, a language and environment for statistical analysis. The structure of the whole framework is detailed in this paper, and some selected examples are used to show its practical use. Diego Rodríguez Martínez, Vicente Blanco 0001, Marcos Boullón-Magán, José Carlos Cabaleiro, Casiano Rodríguez, Francisco F. Rivera |
IPDPS | 4 |
| 2005 | Performance optimization of irregular codes based on the combination of reordering and blocking techniques
Juan Carlos Pichel, Dora Blanco Heras, José Carlos Cabaleiro, Francisco F. Rivera |
Parallel Comput. | 3 |
| 2004 | Performance Prediction for Parallel Iterative Solvers
Vicente Blanco 0001, Patricia González, José Carlos Cabaleiro, Dora Blanco Heras, Tomás F. Pena, Juan J. Pombo, Francisco F. Rivera |
J. Supercomput. | 3 |
| 2003 | AVISPA: visualizing the performance prediction of parallel iterative solvers
Vicente Blanco 0001, Patricia González, José Carlos Cabaleiro, Dora Blanco Heras, Tomás F. Pena, Juan J. Pombo, Francisco F. Rivera |
Future Gener. Comput. Syst. | 3 |
| 2001 | Modeling and improving locality for the sparse-matrix-vector product on cache memories
Dora Blanco Heras, Vicente Blanco 0001, José Carlos Cabaleiro, Francisco F. Rivera |
Future Gener. Comput. Syst. | 3 |
| 2001 | Modeling data locality for the sparse matrix-vector product using distance measures
Dora Blanco Heras, José Carlos Cabaleiro, Francisco F. Rivera |
Parallel Comput. | 2 |
| 2001 | Parallel Computation of Wavelet Transforms Using the Lifting Scheme
Patricia González, José Carlos Cabaleiro, Tomás F. Pena |
J. Supercomput. | 2 |
| 2000 | On parallel solvers for sparse triangular systems
Patricia González, José Carlos Cabaleiro, Tomás F. Pena |
J. Syst. Archit. | 2 |
| 1990 | Systolic architecture for the calculation of the correlation coefficients
Emilio L. Zapata, José Carlos Cabaleiro, Ramón Doallo, Francisco Argüello |
Microprocessing and Microprogramming | 2 |