EDBT 2026 Demo / reviewers in the wild / expert
Francisco F. Rivera
dblp:68/5325 · also Francisco Fernandez Rivera
· DBLP profile ↗
51ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0002-6728-9350ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Assessing Intel OneAPI capabilities and cloud-performance for heterogeneous computingabstractAbstract This work presents a performance-oriented study of a heterogeneous application developed with Intel OneAPI to solve two well-known diffusion problems: heat diffusion and image denoising. We have explored CPU+iGPU and CPU+FPGA schemes, applying dynamic load balancing and conducting experiments on Intel DevCloud. The results demonstrate that the CPU+iGPU scheme outperforms the execution times achieved by the fastest device when the problem is sufficiently computationally demanding. We also found that the performance of the CPU+FPGA scheme is heavily affected by bandwidth limitations and specific strategies to manage memory efficiently are required. Moreover, it was demonstrated that dynamic workload balancing is crucial due to possible performance fluctuations in any of the implicated devices. In conclusion, Intel OneAPI provides a helpful tool for multi-platform development using a unique high-level language, DPC++. However, developing specific code for each platform is necessary to achieve optimal performance. Silvia R. Alcaraz, Ruben Laso, Oscar G. Lorenzo, David López Vilariño, Tomás F. Pena, Francisco F. Rivera |
J. Supercomput. | 6 |
| 2024 | A new thread-level speculative automatic parallelization model and library based on duplicate code executionabstractAbstract Loop-efficient automatic parallelization has become increasingly relevant due to the growing number of cores in current processors and the programming effort needed to parallelize codes in these systems efficiently. However, automatic tools fail to extract all the available parallelism in irregular loops with indirections, race conditions or potential data dependency violations, among many other possible causes. One of the successful ways to automatically parallelize these loops is the use of speculative parallelization techniques. This paper presents a new model and the corresponding C++ library that supports the speculative automatic parallelization of loops in shared memory systems, seeking competitive performance and scalability while keeping user effort to a minimum. The primary speculative strategy consists of redundantly executing chunks of loop iterations in a duplicate fashion. Namely, each chunk is executed speculatively in parallel to obtain results as soon as possible and sequentially in a different thread to validate the speculative results. The implementation uses C++11 threads and it makes intensive use of templates and advanced multithreading techniques. An evaluation based on various benchmarks confirms that our proposal provides a competitive level of performance and scalability. Millán Álvarez Martínez, Basilio B. Fraguela, José Carlos Cabaleiro, Francisco F. Rivera |
J. Supercomput. | 4 |
| 2022 | CIMAR, NIMAR, and LMMA: Novel algorithms for thread and memory migrations in user space on NUMA systems using hardware countersabstractThis paper introduces two novel algorithms for thread migrations, named CIMAR (Core-aware Interchange and Migration Algorithm with performance Record –IMAR–) and NIMAR (Node-aware IMAR), and a new algorithm for the migration of memory pages, LMMA (Latency-based Memory pages Migration Algorithm), in the context of Non-Uniform Memory Access (NUMA) systems. This kind of system has complex memory hierarchies that present a challenging problem in extracting the best possible performance, where thread and memory mapping play a critical role. The presented algorithms gather and process the information provided by hardware counters to make decisions about the migrations to be performed, trying to find the optimal mapping. They have been implemented as a user space tool that looks for improving the system performance, particularly in, but not restricted to, scenarios where multiple programs with different characteristics are running. This approach has the advantage of not requiring any modification on the target programs or the Linux kernel while keeping a low overhead. Two different benchmark suites have been used to validate our algorithms: The NAS parallel benchmark, mainly devoted to computational routines, and the LevelDB database benchmark focused on read–write operations. These benchmarks allow us to illustrate the influence of our proposal in these two important types of codes. Note that those codes are state-of-the-art implementations of the routines, so few improvements could be initially expected. Experiments have been designed and conducted to emulate three different scenarios: a single program running in the system with full resources, an interactive server where multiple programs run concurrently varying the availability of resources, and a queue of tasks where granted resources are limited. The proposed algorithms have been able to produce significant benefits, especially in systems with higher latency penalties for remote accesses. When more than one benchmark is executed simultaneously, performance improvements have been obtained, reducing execution times up to 60%. In this kind of situation, the behaviour of the system is more critical, and the NUMA topology plays a more relevant role. Even in the worst case, when isolated benchmarks are executed using the whole system, that is, just one task at a time, the performance is not degraded. Ruben Laso, Oscar G. Lorenzo, José Carlos Cabaleiro, Tomás F. Pena, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera |
Future Gener. Comput. Syst. | 6 |
| 2021 | LBMA and IMAR2: Weighted lottery based migration strategies for NUMA multiprocessing serversabstractSummary Multicore NUMA systems present on‐board memory hierarchies and communication networks that influence performance when executing shared memory parallel codes. Characterizing this influence is complex, and understanding the effect of particular hardware configurations on different codes is of paramount importance. In this article, monitoring information extracted from hardware counters at runtime is used to characterize the behavior of each thread for an arbitrary number of multithreaded processes running in a multiprocessing environment. This characterization is given in terms of number of operations per second, operational intensity, and latency of memory accesses. We propose a runtime tool, executed in user space, that uses this information to guide two different thread migration strategies for improving execution efficiency by increasing locality and affinity without requiring any modification in the running codes. Different configurations of NAS Parallel OpenMP benchmarks running concurrently on multicore NUMA systems were used to validate the benefits of our proposal, in which up to four processes are running simultaneously. In more than the 95% of the executions of our tool, results outperform those of the operating system (OS) and produces up to 38% improvement in execution time over the OS for heterogeneous workloads, under different and realistic locality and affinity scenarios. Ruben Laso, Oscar G. Lorenzo, Francisco F. Rivera, José Carlos Cabaleiro, Tomás F. Pena, Juan Ángel Lorenzo del Castillo |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | IHP: a dynamic heterogeneous parallel scheme for iterative or time-step methods - image denoising as case study
Ruben Laso, José Carlos Cabaleiro, Francisco F. Rivera, M. Carmen Muñiz, José A. Álvarez-Dios |
J. Supercomput. | 3 |
| 2017 | Landing sites detection using LiDAR data on manycore systems
Oscar G. Lorenzo, Jorge Martínez Sánchez, David López Vilariño, Tomás F. Pena, José Carlos Cabaleiro, Francisco F. Rivera |
J. Supercomput. | 6 |
| 2014 | Multiobjective optimization technique based on monitoring information to increase the performance of thread migration on multicoresabstractMulticore systems present on-board memory hierarchies and communication networks that influence their performance when they execute shared memory parallel codes. Characterizing this influence is complex, and understanding the effect of particular hardware configurations on different codes is of paramount importance. In this paper, monitoring information extracted from hardware counters in runtime is used to characterize the behaviour of each thread in the parallel code in terms of three values: the number of floating point operations per second, the operational intensity, and the memory access latency. Note that these values characterize the Roofline Model with the inclusion of additional information about memory access latencies. We propose to use this information to guide thread migration strategies that improve the efficiency of the execution of the code by increasing locality and affinity. The idea behind this proposal is to use these three values as objective functions to be optimized as a multiobjective optimization problem. The proposed technique is an iterative method inspired in evolutive optimization algorithms. To this end, an individual utility function is defined to represent the relative importance of these values. This function is a weighted product that can be considered as representative of the performance of each parallel thread. Different configurations of the SAXPY and SDOT kernels on multicores were used to validate the benefits of the proposed thread migration strategies. The results show that our strategy produces improvements up to 25% in scenarios where locality and affinity are low, and negligible degradation is observed when they are high. The use of hardware counters produces low overheads when extracting monitoring information. Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Francisco F. Rivera |
CLUSTER | 5 |
| 2014 | A hardware counter-based toolkit for the analysis of memory accesses in SMPsabstractSUMMARY In this paper, a set of three hardware counter (HC)‐based tools to characterise memory access of parallel codes in Symmetric Multiprocessors (SMPs) is presented. This toolkit simplifies accessing and programming HCs, which are included in modern microprocessors. Hardware counters are used to obtain information about memory accesses in a parallel code at very low cost. This information is presented to the user in a friendly way. The first tool can be used to automatically monitor the memory accesses of a system and to analyse a code even if the source is not available. The second tool allows the user to insert in a source code, in a simple and transparent way, the instructions needed to monitor and manage HCs. This way, specific parts of the code can be analysed. The user can either add appropriate directives to a C code or use a graphical interface to select those parts of the code to be analysed. The tool takes this source file and automatically adds the monitoring code. The third tool takes the information gathered by the aforementioned tools, processes it and displays it graphically. This tool shows the information in a comprehensive and simple way, allowing the user to adjust the level of detail. The aim of these tools was to characterise the memory accesses of parallel codes in multicore systems, in which the cache hierarchy can greatly influence the performance. For illustrative purposes, these tools were used to carry out two case studies, a sparse matrix vector product and a dot product. These studies have been made in two different environments. Anyway, they can be used in almost any system as long as the necessary HCs are available.Copyright © 2013 John Wiley & Sons, Ltd. Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera |
Concurr. Comput. Pract. Exp. | 6 |
| 2014 | Modeling the performance of parallel applications using model selection techniquesabstractSUMMARY Nowadays, parallel architectures are changing so fast that there is a need for scalable and efficient tools to analyze and predict the performance of parallel applications. Analytical models are proved to be a useful approximation for characterizing parallel algorithms, but developing accurate analytical models is a hard issue, and, in general, they provide coarse performance predictions due to their intrinsic lack of accuracy. In this paper, we describe in detail the Tools for Instrumentation and Analysis (TIA) framework, an easy‐to‐use tool that automatically obtains accurate performance models by means of analytical expressions. This framework automatizes most of its internal tasks, reducing opportunities for human error, and it only requires the user to focus on the metrics and execution parameters that might influence the performance, those that should be considered in the modeling process. Its main advantage over other tools is that TIA uses model selection techniques that allow the automation of the modeling process. As a case of study, the use of TIA to obtain analytical models of different implementations of the broadcast collective communication in a cluster of multicores is shown. The results obtained by TIA are evaluated and compared with theoretical approaches based on the LogGP model. Copyright © 2013 John Wiley & Sons, Ltd. Diego Rodríguez Martínez, Vicente Blanco 0001, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera |
Concurr. Comput. Pract. Exp. | 5 |
| 2014 | Using sampled information: is it enough for the sparse matrix-vector product locality optimization?abstractSUMMARY One of the main factors that affect the performance of the sparse matrix–vector product (SpMV) is the low data reuse caused by the irregular and indirect memory access patterns. Different strategies to deal with this problem such as data reordering techniques have been proposed. The computational cost of these techniques is typically high because they consider all the nonzeros of the sparse matrix in order to find an appropriate permutation of rows and columns that improves the SpMV performance. In this paper, we analyze the possibility of increasing the locality of the SpMV using incomplete information in the reordering process. This partial information comes as a consequence of considering only a subset of the nonzero elements of the matrix. These nonzeros are obtained from the original matrix through a sampling process. In particular, two different sampling methods have been considered: a random sampling and an event‐based sampling using hardware counters. We have detected that a small number of samples is enough to obtain quality reorderings. As a consequence, using sampling‐based reorderings leads to noticeable performance improvements with respect to the non‐reordered matrices, reaching speedup values up to 2.1 × . In addition, an important reduction in the computational time required by the reordering technique has been observed. Copyright © 2012 John Wiley & Sons, Ltd. Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera, Dora Blanco Heras, Tomás F. Pena |
Concurr. Comput. Pract. Exp. | 3 |
| 2014 | 3DyRM: a dynamic roofline model including memory latency information
Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Francisco F. Rivera |
J. Supercomput. | 5 |
| 2013 | Sparse matrix-vector multiplication on the Single-Chip Cloud Computer many-core processor
Juan Carlos Pichel, Francisco F. Rivera |
J. Parallel Distributed Comput. | 2 |
| 2013 | A flexible and dynamic page migration infrastructure based on hardware counters
Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, Francisco F. Rivera, Tomás F. Pena, José Carlos Cabaleiro |
J. Supercomput. | 3 |
| 2012 | Hardware Counters Based Analysis of Memory Accesses in SMPsabstractModern microprocessors incorporate Hardware Counters (HC) that provide useful information with low overhead. HC are not commonly used because of the lack of tools to get their information in an easy way. In this paper, a set of tools to simplify the accessing and programming of Intel Itanium 2 ™EARs (Event Address Registers) is presented. The aim of these tools is to characterise the memory accesses of parallel codes, in multicore systems, in which the cache hierarchy can greatly influence the performance. The first tool allows the user to insert in the code, in a simple and transparent way, the instructions needed to monitor and manage hardware counters. Two versions of this tool have been implemented. The first one is a command line tool that takes as input a C source file with appropriate directives and outputs it with the monitoring code added. The other one is a graphical interface that allows the user to select the parts of the code to analise. The second tool takes the information gathered by the monitored parallel code provided by the hardware counters and displays it graphically. This tool shows the information in a comprehensive but simple way, allowing the user to adjust the level of detail. These tools were used to carry out a study of parallel irregular codes. Although this study has been made in a specific environment, the tools here presented can be used in any system as long as it is based on hardware counters present in current processors. Oscar G. Lorenzo, Tomás F. Pena, José Carlos Cabaleiro, Juan Carlos Pichel, Juan Ángel Lorenzo del Castillo, Francisco F. Rivera |
ISPA | 6 |
| 2012 | Model Selection to Characterize Performance Using Genetic AlgorithmsabstractThe TIA modeling framework provides analytical models of the performance of parallel applications. The resulting models are obtained using model selection techniques and are accurate enough for various purposes. Its main drawback is that the completion time depends on the number of candidate models and, in some situations, it becomes critical. In this work, a genetic algorithm is proposed for reducing the time for searching of the best candidate model. The use of this genetic algorithm to obtain the performance model of the linear implementation of the broadcast collective communication in a cluster of multicores is shown. Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001 |
ISPA | 4 |
| 2012 | A Graphical Tool for Performance Analysis of Multicore Systems Based on the Roofline ModelabstractA tool to characterize the performance of parallel codes on multicore systems is presented in this paper. This tool allows the user to define the Roofline Model of the target system, to execute the code under study and to represent the performance results in the roofline plot. The final product is an easy to use tool to provide an insightful model which allows to determine, at a glance, performance issues like load balance, locality and those related to thread and memory allocation. Results show that this model provides practical information of the effects that degrade the performance of a code and gives hints to improve it. Francisco F. Rivera, Ramón Iglesias, Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, Tomás F. Pena, José Carlos Cabaleiro |
ISPA | 1 |
| 2011 | Estimating the effect of cache misses on the performance of parallel applications using analytical modelsabstractIn this paper a methodology to characterize the influence of cache misses on the performance of parallel applications is presented. This methodology is based on analytical models provided by the TIA framework. This framework obtains analytical models of given observable quantities by instrumenting the source code and applying model selection techniques. In particular, two metrics related with the performance are considered in this work: the number of cache misses and the elapsed time. Based on both models, the influence in terms of execution time due to the cache misses can be inferred. Two different versions of the parallel product of dense matrices are used as case of study. Diego Rodríguez Martínez, Vicente Blanco 0001, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera |
AICCSA | 5 |
| 2011 | Using accurate AIC-based performance models to improve the scheduling of parallel applications
Diego Rodríguez Martínez, Julio L. Albín, Tomás F. Pena, José Carlos Cabaleiro, Francisco F. Rivera, Vicente Blanco 0001 |
J. Supercomput. | 5 |
| 2010 | Lessons Learnt Porting Parallelisation Techniques for Irregular Codes to NUMA SystemsabstractThis work presents a study undertaken to characterise the behaviour of some parallelisation techniques for irregular codes, previously developed for SMP architectures, on a several-node SMP NUMA system. The main objective is to determine the performance effect of bus contention and cache coherency in such a complex architecture. Results show that: (1) cores which share a socket can be considered as independent processors in this context; (2) for big data sizes, the effect of sharing a bus degrades the performance but masks the cache coherency effects and (3) the NUMA-ratio is a critical factor on irregular codes. These results allow us to study the effect in performance of the thread-to-core mappings and memory allocation policies. Juan Ángel Lorenzo del Castillo, Juan Carlos Pichel, David LaFrance-Linden, Francisco F. Rivera, David E. Singh |
PDP | 4 |
| 2010 | Performance Modeling of MPI Applications Using Model Selection TechniquesabstractA new method for obtaining models of the performance of parallel applications based on statistical analysis is presented in this paper. This method is based on the Akaike's information criterion (AIC) that provides an objective mechanism to rank different models by means of an experimental data fit. The input of the modeling process is a set of variables and parameters that can a priori influence the performance of the application. This set can be provided by the user. Using this information, the method automatically generates a set of candidate models. These models are fit to the experimental data and the AIC score of each model is calculated. The model with the best AIC score is selected as the best model. Also, using the AIC scores of all candidate models, useful statistical information is provided to help the user to evaluate the quality of the selected model, as well as indications of how to interactively improve this modeling process. As a first case of study, statistical models obtained for different implementations of the broadcast collective communication in Open MPI are shown. These models are very accurate, exceeding its adjustment to theoretical approaches based on the LogGP model. Finally, the NAS Parallel Benchmark is also characterized using this new method with good results in terms of accuracy. Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001 |
PDP | 4 |
| 2009 | Accurate analytical performance model of communications in MPI applicationsabstractThis paper presents a new LogP-based model, called LoOgGP, which allows an accurate characterization of MPI applications based on microbenchmark measurements. This new model is an extension of LogP for long messages in which both overhead and gap parameters perform a linear dependency with message size. The LoOgGP model has been fully integrated into a modelling framework to obtain statistical models of parallel applications, providing the analyst with an easy and automatic tool for LoOgGP parameter set assessment to characterize communications. The use of LoOgGP model to obtain a statistical performance model of an image deconvolution application is illustrated as a case of study. Diego Rodríguez Martínez, José Carlos Cabaleiro, Tomás F. Pena, Francisco F. Rivera, Vicente Blanco 0001 |
IPDPS | 4 |
| 2009 | On the Influence of Thread Allocation for Irregular Codes in NUMA SystemsabstractThis work presents a study undertaken to characterise the FINISTERRAE supercomputer, one of the biggest NUMA systems in Europe. The main objective was to determine the performance effect of bus contention and cache coherency as well as the suitability of porting strategies regarding irregular codes in such a complex architecture. Results show that: (1) cores which share a socket can be considered as independent processors in this context; (2) for big data sizes, the effect of sharing a bus degrades the final performance but masks the cache coherency effects; (3) the NUMA factor (remote to local memory latency ratio) is an important factor on irregular codes and (4) the default kernel allocation policy is not optimal in this system. These results allow us to understand the behaviour of thread-to-core mappings and memory allocation policies. Juan Ángel Lorenzo del Castillo, Francisco F. Rivera, Petr Tuma 0001, Juan Carlos Pichel |
PDCAT | 2 |
| 2009 | Increasing data reuse of sparse algebra codes on simultaneous multithreading architecturesabstractAbstract In this paper the problem of the locality of sparse algebra codes on simultaneous multithreading (SMT) architectures is studied. In these kind of architectures many hardware structures are dynamically shared among the running threads. This puts a lot of stress on the memory hierarchy, and a poor locality, both inter‐thread and intra‐thread, may become a major bottleneck in the performance of a code. This behavior is even more pronounced when the code is irregular, which is the case of sparse matrix ones. Therefore, techniques that increase the locality of irregular codes on SMT architectures are important to achieve high performance. This paper proposes a data reordering technique specially tuned for these kind of architectures and codes. It is based on a locality model developed by the authors in previous works. The technique has been tested, first, using a simulator of a SMT architecture, and subsequently, on a real architecture as Intel's Hyper‐Threading. Important reductions in the number of cache misses have been achieved, even when the number of running threads grows. When applying the locality improvement technique, we also decrease the total execution time and improve the scalability of the code. Copyright © 2009 John Wiley & Sons, Ltd. Juan Carlos Pichel, Dora Blanco Heras, José Carlos Cabaleiro, Francisco F. Rivera |
Concurr. Comput. Pract. Exp. | 4 |
| 2008 | Topic 3: Scheduling and Load Balancing
Dieter Kranzlmüller, Uwe Schwiegelshohn, Yves Robert, Francisco F. Rivera |
Euro-Par | 4 |
| 2007 | Software Tools for Performance Modeling of Parallel ProgramsabstractThis paper presents a framework based on a user driven methodology to obtain analytical models of MPI applications on parallel systems in a systematic and easy to use way. This methodology consists of two stages. In the first one, instrumentation of the source code is performed using CALL, which is a profiling tool for interacting with the code in an easy, simple and direct way. New features are added to CALL to obtain different performance metrics and store the performance information in XML files. Using this information, an analytical model of the performance behavior is obtained in the second stage by means of R, a language and environment for statistical analysis. The structure of the whole framework is detailed in this paper, and some selected examples are used to show its practical use. Diego Rodríguez Martínez, Vicente Blanco 0001, Marcos Boullón-Magán, José Carlos Cabaleiro, Casiano Rodríguez, Francisco F. Rivera |
IPDPS | 6 |
| 2007 | An Inspector/Executor Based Strategy to Efficiently Parallelize N-Body Simulation Programs on Shared Memory SystemsabstractReordering of data is becoming more and more significant in order to achieve a higher performance in memory data access and, particularly, in program runtime. This fact becomes specially important in parallel applications that are executed in shared memory systems. This work presents a new parallelizing, run time strategy for irregular structures associated to N-Body problem simulation algorithms. Such strategy, so-called STPCLS (Step Classification), is based on the inspector-executor paradigm. It has been tested in a shared memory system using a significant set of irregular loops. The outcomes show that the efficiency of our solution is high, and the benefits overcome the overheads imposed by our algorithm. Juan Ángel Lorenzo del Castillo, Julio L. Albín, Tomás F. Pena, Francisco F. Rivera, David E. Singh |
ISPDC | 4 |
| 2006 | Image segmentation based on merging of sub-optimal segmentations
Juan Carlos Pichel, David E. Singh, Francisco F. Rivera |
Pattern Recognit. Lett. | 3 |
| 2005 | Performance optimization of irregular codes based on the combination of reordering and blocking techniques
Juan Carlos Pichel, Dora Blanco Heras, José Carlos Cabaleiro, Francisco F. Rivera |
Parallel Comput. | 4 |
| 2004 | Performance Prediction for Parallel Iterative Solvers
Vicente Blanco 0001, Patricia González, José Carlos Cabaleiro, Dora Blanco Heras, Tomás F. Pena, Juan J. Pombo, Francisco F. Rivera |
J. Supercomput. | 7 |
| 2003 | Efficient Dynamic Load Balancing Strategies for Parallel Active Set Optimization Methods
Inmaculada Pardines, Francisco F. Rivera |
Euro-Par | 2 |
| 2003 | Increasing the Parallelism of Irregular Loops with Dependences
David E. Singh, María J. Martín, Francisco F. Rivera |
Euro-Par | 3 |
| 2003 | AVISPA: visualizing the performance prediction of parallel iterative solvers
Vicente Blanco 0001, Patricia González, José Carlos Cabaleiro, Dora Blanco Heras, Tomás F. Pena, Juan J. Pombo, Francisco F. Rivera |
Future Gener. Comput. Syst. | 7 |
| 2003 | Research Article: A GIS-embedded system to support land consolidation plans in GaliciaabstractLand consolidation is a strategic instrument for rural planning and thus economic development in the Spanish region of Galicia. This paper describes an experimental system embedded in a GIS environment to aid rural engineers to develop land consolidation plans. The system supports all the stages of the plan and many functionalities are implemented as heuristic processes based on expert knowledge and advice. The overall aim is to overcome administrative and technical problems of traditional consolidation procedures. The system provides an integrated framework for the management of spatial and administrative consolidation information. It also includes optimization-based algorithms for the automated generation of multiple alternative parcel reallocations, as well as an environment to refine and objectively evaluate the proposed solutions. These key capabilities result in a powerful tool for decision making that dramatically reduces the time and cost of land consolidation plans. Pilot experiences in two consolidation zones of Galicia assess the feasibility and effectiveness of the system. Juan Touriño, Jorge Parapar, Ramón Doallo, Marcos Boullón-Magán, Francisco F. Rivera, Javier D. Bruguera, Xesús P. González, Rafael Crecente-Maseda |
Int. J. Geogr. Inf. Sci. | 5 |
| 2003 | High performance air pollution modeling for a power plant environment
María J. Martín, David E. Singh, José Carlos Mouriño, Francisco F. Rivera, Ramón Doallo, Javier D. Bruguera |
Parallel Comput. | 4 |
| 2002 | Improving Locality in the Parallelization of Doacross Loops (Research Note)
María J. Martín, David E. Singh, Juan Touriño, Francisco F. Rivera |
Euro-Par | 4 |
| 2002 | Exploiting Locality in the Run-Time Parallelization of Irregular LoopsabstractThe goal of this work is the efficient parallel execution of loops with indirect array accesses, in order to be embedded in a parallelizing compiler framework. In this kind of loop pattern, dependences can not always be determined at compile-time as, in many cases, they involve input data that are only known at run-time and/or the access pattern is too complex to be analyzed In this paper we propose runtime strategies for the parallelization of these loops. Our approaches focus not only on extracting parallelism among iterations of the loop, but also on exploiting data access locality to improve memory hierarchy behavior and, thus, the overall program speedup. Two strategies are proposed one based on graph partitioning techniques and other based on a block-cyclic distribution. Experimental results show that both strategies are complementary and the choice of the best alternative depends on some features of the loop pattern. María J. Martín, David E. Singh, Juan Touriño, Francisco F. Rivera |
ICPP | 4 |
| 2001 | Modeling and improving locality for the sparse-matrix-vector product on cache memories
Dora Blanco Heras, Vicente Blanco 0001, José Carlos Cabaleiro, Francisco F. Rivera |
Future Gener. Comput. Syst. | 4 |
| 2001 | Modeling data locality for the sparse matrix-vector product using distance measures
Dora Blanco Heras, José Carlos Cabaleiro, Francisco F. Rivera |
Parallel Comput. | 3 |
| 1999 | Scheduling of Algorithms Based on Elimination Trees on NUMA Systems
María J. Martín, Inmaculada Pardines, Francisco F. Rivera |
Euro-Par | 3 |
| 1994 | Combining static and dynamic scheduling on distributed-memory multiprocessorsabstractLoops are a large source of parallelism for many numerical applications. An important issue in the parallel execution of loops is how to schedule them so that the workload is well balanced among the processors. Most existing loop scheduling algorithms were designed for shared-memory multiprocessors, with uniform memory access costs. These approaches are not suitable for distributed-memory multiprocessors where data locality is a major concern and communication costs are high. This paper presents a new scheduling algorithm in which data locality is taken into account. Our approach combines both worlds, static and dynamic scheduling, in a two-level (overlapped) fashion. This way data locality is considered and communication costs are limited. The performance of the new algorithm is evaluated on a CM-5 message-passing distributed-memory multiprocessor. Oscar G. Plata, Francisco F. Rivera |
International Conference on Supercomputing | 2 |
| 1992 | Image reconstruction on hypercube computers: Application to electron microscopy
Emilio L. Zapata, José Ignacio Benavides Benítez, Francisco F. Rivera, Javier D. Bruguera, Tomás F. Pena, José María Carazo |
Signal Process. | 3 |
| 1991 | Modified Gram-Schmidt QR Factorization on Hypercube SIMD Computers
Emilio L. Zapata, J. A. Lamas, Francisco F. Rivera, Oscar G. Plata |
J. Parallel Distributed Comput. | 3 |
| 1990 | ACLE: A Software Package for SIMD Computer SimulationabstractThis paper describes ACLE (Array C Language Emulator), a software package comprising an ACLAN-to-C translator and a library of simulation routines enabling the execution of programs written in ACLAN to be simulated on a conventional sequential computer. Array C LANguage (ACLAN) is a machine-independent programming language that extends C by endowing it with structures for programming array processors. ACLAN was successfully proven by developing many parallel algorithms for hypercube computers. An algorithmic solution for mapping algorithms onto these computers is explained. Oscar G. Plata, Javier D. Bruguera, Francisco F. Rivera, Ramón Doallo, Emilio L. Zapata |
Comput. J. | 3 |
| 1990 | Parallel Squared Error Clustering on Hypercube Arrays
Francisco F. Rivera, Emilio L. Zapata |
J. Parallel Distributed Comput. | 1 |
| 1990 | Multidimensional fast Hartley transform onto SIMD hypercubes
Emilio L. Zapata, Francisco Argüello, Francisco F. Rivera, Javier D. Bruguera |
Microprocessing and Microprogramming | 3 |
| 1990 | Gaussian elimination with pivoting on hypercubes
Francisco F. Rivera, Ramón Doallo, Javier D. Bruguera, Emilio L. Zapata, Richard L. Peskin |
Parallel Comput. | 1 |
| 1990 | Cluster validity based on the hard tendency of the fuzzy classification
Francisco F. Rivera, Emilio L. Zapata, José María Carazo |
Pattern Recognit. Lett. | 1 |
| 1990 | Image template matching on hypercube SIMD computers
Emilio L. Zapata, José Ignacio Benavides Benítez, Oscar G. Plata, Francisco F. Rivera, José María Carazo |
Signal Process. | 4 |
| 1989 | A parallel markovian model reliability algorithm for hypercube networks
Emilio L. Zapata, Javier D. Bruguera, Oscar G. Plata, Francisco F. Rivera |
Microprocessing and Microprogramming | 4 |
| 1989 | Parallel fuzzy clustering on fixed size hypercube SIMD computers
Emilio L. Zapata, Francisco F. Rivera, Oscar G. Plata |
Parallel Comput. | 2 |
| 1988 | A VLSI systolic architecture for fuzzy clustering
Emilio L. Zapata, Ramón Doallo, Francisco F. Rivera |
Microprocess. Microprogramming | 3 |