José M. García 0001

dblp:99/1774-1 · also José Manuel García 0001, José Manuel García Carrasco · DBLP profile ↗
← Back
97ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0002-6388-2835ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 73 · 5 since 2021Artificial intelligence and machine learning · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Computer networks · 1Security and privacy · 1
YearPublicationVenuePosition
2026 Characterization of machine learning compilers for LLM inference on NVIDIA GPUs
abstract
Abstract AI inference is conflicted between Performance, developer Productivity, and device Portability–the P3 problem. Machine learning compilers (MLCs) aim to address this, but their ecosystem is fragmented, with tools that each prioritize a different issue. This paper evaluates the deployment trade-offs of PyTorch-based LLMs on NVIDIA GPUs using four intertwined prominent MLC tools: , TensorRT, XLA, and ONNX Runtime. A dual methodology is used, leveraging synthetic PyTorch models to isolate optimizations and end-to-end benchmarks with State-of-the-Art (SOTA) models (TinyLlama-1.1B, Llama-2-7B) to measure real-world performance. Findings reveal that the peak performance of Ahead-Of-Time (AOT) compilation requires architecture-specific tools such as TensorRT-LLM, which are necessary for SOTA LLMs but are unusable for PyTorch models. As for Just-In-Time (JIT) solutions such as and its backends, they are flexible and portable, compatible with all tested models, but they do not consistently accelerate LLMs; therefore, the choice of MLC depends on P3 considerations and model architecture.
Alejandro Carmona-Martínez, Gregorio Bernabé, José M. García 0001
J. Supercomput.3
2025 A Real Time Cardiomyopathy Detection Tool Using Ml Ensemble Models
abstract
Left Ventricular noncompaction (LVNC) is a recently classified form of cardiomyopathy. Although various methods have been proposed for accurately quantifying trabeculae in the left ventricle (LV), consensus on the optimal approach remains elusive. Previous research introduced DL‐LVTQ, a deep learning solution for trabecular quantification based on a UNet 2D convolutional neural network (CNN) architecture and a graphical user interface (GUI) to streamline its use in clinical workflows. Building on this foundation, this work presents LVNC detector, an enhanced application designed to support cardiologists in the automated diagnosis of LVNC. The application integrates two segmentation models: DL‐LVTQ and ViTUNet, the latter inspired by modern hybrid architectures combining convolutional neural networks (CNNs) and transformer‐based designs. These models, implemented within an ensemble framework, leverage advancements in deep learning to improve the accuracy and robustness of magnetic resonance imaging (MRI) segmentation. Key innovations include multithreading to optimize model loading times and ensemble methods to enhance segmentation consistency across MRI slices. Additionally, the platform‐independent design ensures compatibility with Windows and Linux, eliminating complex setup requirements. The LVNC detector delivers an efficient and user‐friendly solution for LVNC diagnosis. It enables real‐time performance and allows cardiologists to select and compare segmentation models for improved diagnostic outcomes. This work demonstrates how state‐of‐the‐art machine learning techniques can seamlessly integrate into clinical practice to reduce human error and expedite diagnostic processes.
Salvador de Haro, Esteban Becerra, Pilar González-Férez, José M. García 0001, Gregorio Bernabé
IET Softw.4
2024 POAS: a framework for exploiting accelerator level parallelism in heterogeneous environments
abstract
Abstract In the era of heterogeneous computing, a new paradigm called accelerator level parallelism (ALP) has emerged. In ALP, accelerators are used concurrently to provide unprecedented levels of performance and energy efficiency. To reach that there are many problems to be solved, one of the most challenging being co-execution. In this paper, we present a new scheduling framework called POAS, a general method for providing co-execution to applications. Our proposal consists of four steps: predict, optimize, adapt and schedule. With POAS, an unseen application can be executed concurrently in ALP with little effort. We evaluate POAS on a heterogeneous environment consisting of CPUs, GPUs (CUDA cores), and XPUs (Tensor cores) on two different fields, namely linear algebra (matrix multiplication benchmark) and deep learning (convolution benchmark). Our experiments prove that POAS provides excellent performance and completes the tasks within a time very close to the optimal time for the hardware and applications used, with a negligible execution time overhead. Moreover, the POAS predictor performed exceptionally well, achieving very low RMSE values for both use cases. Therefore, POAS can be a valuable tool for fully exploiting ALP and improving overall performance over offloading in heterogeneous settings.
Pablo Antonio Martínez, Gregorio Bernabé, José M. García 0001
J. Supercomput.3
2023 Matching Linear Algebra and Tensor Code to Specialized Hardware Accelerators
abstract
Dedicated tensor accelerators demonstrate the importance of linear algebra in modern applications. Such accelerators have the potential for impressive performance gains, but require programmers to rewrite code using vendor APIs - a barrier to wider scale adoption. Recent work overcomes this by matching and replacing patterns within code, but such approaches are fragile and fail to cope with the diversity of real-world codes.
Pablo Antonio Martínez, Jackson Woodruff, Jordi Armengol-Estapé, Gregorio Bernabé, José M. García 0001, Michael F. P. O'Boyle
CC5
2023 On the representativeness and stability of a set of EFMs
abstract
MOTIVATION: Elementary flux modes are a well-known tool for analyzing metabolic networks. The whole set of elementary flux modes (EFMs) cannot be computed in most genome-scale networks due to their large cardinality. Therefore, different methods have been proposed to compute a smaller subset of EFMs that can be used for studying the structure of the network. These latter methods pose the problem of studying the representativeness of the calculated subset. In this article, we present a methodology to tackle this problem. RESULTS: We have introduced the concept of stability for a particular network parameter and its relation to the representativeness of the EFM extraction method studied. We have also defined several metrics to study and compare the EFM biases. We have applied these techniques to compare the relative behavior of previously proposed methods in two case studies. Furthermore, we have presented a new method for the EFM computation (PiEFM), which is more stable (less biased) than previous ones, has suitable representativeness measures, and exhibits better variability in the extracted EFMs. AVAILABILITY AND IMPLEMENTATION: Software and additional material are freely available at https://github.com/biogacop/PiEFM.
Francisco Guil 0002, José F. Hidalgo, José M. García 0001
Bioinform.3
2022 Applying Intel's oneAPI to a machine learning case study
abstract
Abstract Different technologies and approaches exist to work around the performance portability problem. Companies and academia work together to find a way to preserve performance across heterogeneous hardware using a unified language, one language to rule them all. Intel's oneAPI appears with this idea in mind. In this article, we try the new Intel solution to approach heterogeneous programming, choosing machine learning as our case study. More precisely, we choose Caffe, a machine learning framework that was created six years ago. Nevertheless, how would it be to make Caffe again from the beginning, using a fresh new technology like oneAPI? In terms of not only the ease of programming‐because only one source code would be needed to deploy Caffe to CPUs, GPUs, FPGAs, and accelerators (platforms that oneAPI currently supports)‐but also performance, where oneAPI may be capable of taking advantage of specific hardware automatically. Is Intel's oneAPI ready to take the leap?
Pablo Antonio Martínez, Biagio Peccerillo, Sandro Bartolini, José M. García 0001, Gregorio Bernabé
Concurr. Comput. Pract. Exp.4
2022 HDNN: a cross-platform MLIR dialect for deep neural networks
abstract
Abstract This paper presents HDNN, a proof-of-concept MLIR dialect for cross-platform computing specialized in deep neural networks. As target devices, HDNN supports CPUs, GPUs and TPUs. In this paper, we provide a comprehensive description of the HDNN dialect, outlining how this novel approach aims to solve the $$P^3$$ P 3 problem of parallel programming (portability, productivity, and performance). An HDNN program is device-agnostic, i.e., only the device specifier has to be changed to run a given workload in one device or another. Moreover, HDNN has been designed to be a domain-specific language, which ultimately helps programming productivity. Finally, HDNN relies on optimized libraries for heavy, performance-critical workloads. HDNN has been evaluated against other state-of-the-art machine learning frameworks on all the hardware platforms achieving excellent performance. We conclude that the ideas and concepts used in HDNN can be crucial for designing future generation compilers and programming languages to overcome the challenges of the forthcoming heterogeneous computing era.
Pablo Antonio Martínez, Gregorio Bernabé, José M. García 0001
J. Supercomput.3
2021 ACOTSP-MF: A memory-friendly and highly scalable ACOTSP approach
abstract
Ant Colony Optimization (ACO) is a population-based meta-heuristic inspired by the social behavior of ants. It is successfully applied in solving many NP-hard problems, such as the Traveling Salesman Problem (TSP). Large-sized instances pose two memory problems to the ACOTSP algorithm: the memory size and the memory bandwidth. This work has focused on developing ACOTSP-MF, a new ACOTSP algorithm proposed to adequately manage the memory issues that arise while solving large TSP instances. ACOTSP-MF uses the nearest neighbor list, introducing a novel class of cities, the backup cities, while grouping cities into three classes: the nearest neighbor cities, the backup cities, and the rest of the cities (the majority). ACOTSP-MF also modifies how the base ACOTSP carries out the tour construction and pheromone update phases depending on the group to which a city belongs. This way, ACOTSP-MF reduces both the memory requirements of its data structures (from O(n∗n) to O(n)), and the memory bandwidth needs (thanks to better exploitation of the memory data locality). In this paper, we have carried out an in-depth analysis of ACOTSP-MF performance for medium and large TSP instances, covering vectorization and scalability issues and showing its main bottlenecks. For medium-size instances, the paper reports speedup factors of 20-500X for the rl11849 instance compared to the base ACOTSP version. ACOTSP-MF is intended and especially adequate for large-size instances. In this context, the paper reports excellent execution time for the Tour Construction phase, with less than 500 ms per iteration for the earring-200k instance. Finally, a study about the solution quality of ACOTSP-MF has been included, showing that ACOTSP-MF paired with local search offers high solution quality (within 2% of the best-known solution).
Pablo Antonio Martínez, José M. García 0001
Eng. Appl. Artif. Intell.2
2021 Deploying deep learning approaches to left ventricular non-compaction measurement
Jesús M. Rodríguez-de-Vera, Josefa González-Carrillo, José M. García 0001, Gregorio Bernabé
J. Supercomput.3
2020 Boosting the extraction of elementary flux modes in genome-scale metabolic networks using the linear programming approach
abstract
MOTIVATION: Elementary flux modes (EFMs) are a key tool for analyzing genome-scale metabolic networks, and several methods have been proposed to compute them. Among them, those based on solving linear programming (LP) problems are known to be very efficient if the main interest lies in computing large enough sets of EFMs. RESULTS: Here, we propose a new method called EFM-Ta that boosts the efficiency rate by analyzing the information provided by the LP solver. We base our method on a further study of the final tableau of the simplex method. By performing additional elementary steps and avoiding trivial solutions consisting of two cycles, we obtain many more EFMs for each LP problem posed, improving the efficiency rate of previously proposed methods by more than one order of magnitude. AVAILABILITY AND IMPLEMENTATION: Software is freely available at https://github.com/biogacop/Boost_LP_EFM. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Francisco Guil 0002, José F. Hidalgo, José M. García 0001
Bioinform.3
2020 High-throughput fuzzy clustering on heterogeneous architectures
Juan M. Cebrian, Baldomero Imbernon, Jesús A. Soto, José M. García 0001, José M. Cecilia
Future Gener. Comput. Syst.4
2020 Re-engineering the ant colony optimization for CMP architectures
José M. Cecilia, José M. García 0001
J. Supercomput.2
2019 Efficient, semantics-rich transformation and integration of large datasets
José Antonio Bernabé-Díaz, María Del Carmen Legaz-García, José M. García 0001, Jesualdo Tomás Fernández-Breis
Expert Syst. Appl.3
2018 Application of High Performance Computing Techniques to the Semantic Data Transformation
José Antonio Bernabé-Díaz, María Del Carmen Legaz-García, José M. García 0001, Jesualdo Tomás Fernández-Breis
WorldCIST (1)3
2017 Multi-objective evolutionary feature selection for online sales forecasting
Fernando Jiménez, Gracia Sánchez, José M. García 0001, Guido Sciavicco, Luis Miralles Pechuán
Neurocomputing3
2017 A methodology based on Deep Learning for advert value calculation in CPM, CPC and CPA networks
Luis Miralles Pechuán, Dafne Rosso, Fernando Jiménez, José M. García 0001
Soft Comput.4
2015 TreeEFM: calculating elementary flux modes using linear optimization in a tree-based algorithm
abstract
MOTIVATION: Elementary flux modes (EFMs) analysis constitutes a fundamental tool in systems biology. However, the efficient calculation of EFMs in genome-scale metabolic networks (GSMNs) is still a challenge. We present a novel algorithm that uses a linear programming-based tree search and efficiently enumerates a subset of EFMs in GSMNs. RESULTS: Our approach is compared with the EFMEvolver approach, demonstrating a significant improvement in computation time. We also validate the usefulness of our new approach by studying the acetate overflow metabolism in the Escherichia coli bacteria. To do so, we computed 1 million EFMs for each energetic amino acid and then analysed the relevance of each energetic amino acid based on gene/protein expression data and the obtained EFMs. We found good agreement between previous experiments and the conclusions reached using EFMs. Finally, we also analysed the performance of our approach when applied to large GSMNs. AVAILABILITY AND IMPLEMENTATION: The stand-alone software TreeEFM is implemented in C++ and interacts with the open-source linear solver COIN-OR Linear program Solver (CLP).
Jon Pey, Juan A. Villar, Luis Tobalina, Alberto Rezola, José M. García 0001, John E. Beasley, Francisco J. Planes
Bioinform.5
2015 Soft-error mitigation by means of decoupled transactional memory threads
Daniel Sánchez 0004, Juan M. Cebrian, José M. García 0001, Juan L. Aragón
Distributed Comput.3
2015 ICCI: In-Cache Coherence Information
abstract
In this paper we introduce ICCI, a new cache organization that leverages shared cache resources and flat coherence protocols to provide inexpensive hardware cache coherence for large core counts (e.g., 512), achieving execution times close to a non-scalable sparse directory while noticeably reducing the energy consumption of the memory system. Very simple changes in the system with respect to traditional bit-vector directories are enough to implement ICCI. Moreover, ICCI does not introduce any storage overhead with respect to a broadcast-based protocol, yet it provides large storage space for coherence information. ICCI makes smarter use of cache resources by dynamically allowing last-level cache entries to store blocks or sharing codes. This way, just the minimum number of directory entries required at runtime are allocated. Besides, ICCI suffers a negligible amount of directory-induced invalidations. Results for a 512-core CMP show that ICCI reduces the energy consumption of the memory system by up to 48 percent compared to a tag-embedded directory, up to 15 percent compared to a sparse directory, and up to 8 percent compared to the state-of-the-art Scalable Coherence Directory which ICCI also outperforms in execution time. In addition, ICCI can be used in combination with elaborated sharing codes to apply it to extremely large core counts. We also show analytically that ICCI’s dynamic allocation of entries makes it a suitable candidate to store coherence information efficiently for very large core counts (e.g., over 200K cores), based on the observation that data sharing makes fewer directory entries necessary per core as core count increases.
Antonio García-Guirado, Ricardo Fernández-Pascual, José M. García 0001
IEEE Trans. Computers3
2015 Adaptive Selection of Cache Indexing Bits for Removing Conflict Misses
abstract
The design of cache memories is a crucial part of the design cycle of a modern processor, since they are able to bridge the performance gap between the processor and the memory. Unfortunately, caches with low degrees of associativity suffer a large amount of conflict misses. Although by increasing their associativity a significant fraction of these misses can be removed, this comes at a high cost in both power, area, and access time. In this work, we address the problem of high number of conflict misses in low-associative caches, by proposing an indexing policy that adaptively selects the bits from the block address used to index the cache. The basic premise of this work is that the non-uniformity in the set usage is caused by a poor selection of the indexing bits. Instead, by selecting at run time those bits that disperse the working set more evenly across the available sets, a large fraction of the conflict misses (85 percent, on average) can be removed. This leads to IPC improvements of 10.9 percent for the SPEC CPU2006 benchmark suite. By having less accesses in the L2 cache, our proposal also reduces the energy consumption of the cache hierarchy by 13.2 percent. These benefits come with a negligible area overhead.
Alberto Ros 0001, Polychronis Xekalakis, Marcelo Cintra, Manuel E. Acacio, José M. García 0001
IEEE Trans. Computers5
2014 Exploiting silicon photonics for energy-efficient heterogeneous parallel architectures
abstract
Welcome to this special issue of the journal Concurrency and Computation: Practice and Experience on Exploiting Silicon Photonics for Energy-Efficient Heterogeneous Parallel Architectures, which contains five original manuscripts that cover a complete range of perspectives. Silicon photonics is undoubtedly expected to play a big role in the evolution in board, cross-chip, interposer-level and on-chip interconnection for low-power and/or high-performance computer systems spanning from high-end embedded devices (e.g., tablets and smartphones) and other System-on-Chips (SoCs), up to chips for the High Performance Computing (HPC) domain. The unique features of photonics (e.g., extreme low-latency, end-to-end transmission, high bandwidth density and passive long-range propagation) have the potential constitute a discontinuity element able to modify the expected shape of future computer systems from the design point of view and also from the programmability and/or runtime management perspectives. Summarizing, silicon photonics can bring innovations and benefits into current and foreseeable computing systems directly, due to their intrinsic features, but also indirectly enabling the evolution toward architectures, runtime and resource management approaches that maximize the photonic raw technological opportunities and lead to more efficient overall designs, otherwise impossible. For instance, the extreme low transmission latency (i.e., group velocity of light into silicon, about 15 ps/mm) can potentially allow a higher number of architectural modules to be close each other and thus to enable their effective tight cooperation and communication. However, computer architecture, as well as network on- and off-chip, designs needs to be adapted to extract maximum benefits from the photonic technology, which exposes other substantial differences compared to what designers are well accustomed to. For example, at the moment optical interconnection is end-to-end by nature therefore much of the knowledge and solutions based on store-and-forward paradigm cannot be directly transferred and exploited. However, propagation into a silicon waveguide can occur with limited losses (e.g., even less than 1 dB/cm) over on-chip or interposer distances without signal regeneration needs. In brief, in this arena, new ad-hoc solutions need to be pursued. Then, despite optical communication is very well established and all the involved elements (such as modulators, detectors, waveguides and resonators) have been extensively researched on, silicon photonics applied to computing systems is still in its infancy. Consequently, researchers have already highlighted a deep interaction between design choices at very different layers of abstraction. For this reason, nowadays the whole spectrum of layers, from physical concerns about optical structures on silicon (e.g., module layout and modeling to expose interactions and, for instance, to evaluate and limit insertion losses) up to network issues (e.g., connectivity and topologies, bandwidth and latency) and even computer architecture choices (e.g., memory coherency and consistency models, memory hierarchy and parallelism management), need to be studied with a strong multidisciplinary approach. This special issue contributes to this promising field with extended and carefully reviewed versions of selected papers from the First International Workshop on Exploiting Silicon Photonics for Energy-Efficient Heterogeneous Parallel Architectures (SiPhotonics'14), which was held in Vienna (Austria) as part of the 9th HiPEAC conference on High Performance and Embedded Architecture and Compilers. Therefore, due to the peculiarity of the depicted scenario, the papers of this number address a complete range of perspectives to silicon photonics applied to computing systems, from raw technology issues and solutions up to studies at the overall system level of modern multi-/many-core systems, both from academic and industrial researchers working in this area. We start this special issue with the paper entitled Optical Crossbars on Chip, A Comparative Study based on Worst-Case Losses. In this paper, Le Beux et al., 1 study the worst-case losses for possible crossbar implementations depending on three key design factors: network topology, considered layout and insertion losses induced by the fabrication process. They compare different implementations relying on matrix, multistage and ring-based network topologies, finding that ring-based networks yield the most power-efficient solution. The paper Capturing the Sensitivity of Optical Network Quality Metrics to its Network Interface Parameters by Ortin et al., 2 addresses the network interface architecture (NI) required to support optical communications on the silicon chip. The paper proposes a complete network interface architecture for wavelength-routed optical NoCs, by coping with the intricacy of some specific issues such as flow control, buffering strategy and dual-clock domains, deadlock avoidance, serialization, and above all, the co-design around the requirements of a cache coherency protocol. The most important conclusion is that NI design and optimization perhaps has now higher priority over the relentless search for improvements in individual optical devices. The emerging of circuit-level simulators for photonic integrated circuits (PICs) is driven by recent developments in technologies for integration of large-scale monolithic PICs in both, silicon and InP technologies. Arellano et al., 3 present their solution for modeling PICs in the framework of the circuit-level simulation tool VPIcomponentMakerTMPhotonic Circuits. In their paper The Power of Circuit Simulations for Designing Photonic Integrated Circuits, they demonstrate the combination of different simulation approaches in time domain, frequency domain and time-and-frequency domain (TFDM) for fast and accurate simulations. This is particularly crucial for being able to model and design more and more complex optical circuits like the ones that could be needed to be employed in, and/or between, chip multiprocessors. In the paper entitled Managing Resources Dynamically in Hybrid Photonic-Electronic NoCs, García-Guirado et al., 4 present novel fine-grain policies to manage the photonic resources in a tiled-CMP scenario. The objective is to maintain the optical channel in the load condition that allows it to deliver best performance. Their policies are dynamic and base their decisions on parameters such as message size, ring availability and distance between endpoints, at the message level. The resulting network behavior is also fairer to all cores, reducing processor idle time thanks to faster thread synchronization, improving performance and reducing both the overall network latency and energy consumption when compared to the same CMP without the photonic ring. Finally, the paper Towards Zero Latency Photonic Switching in Shared Memory Networks explores techniques which intelligently use information from the memory hierarchy to predict communication in order to setup photonic circuits with reduced or eliminated arbitration latency in case of reconcilable optical networks. Madarbux et al., 5 present a switch scheduling algorithm which arbitrates on a per memory transaction basis and holds open photonic circuits to exploit temporal locality, showing that this can reduce the average arbitration latency overhead and eliminate arbitration latency altogether for many of memory transactions. We would like to thank all the authors, reviewers and editors involved in the elaboration of this special issue, including also the reviewers that were involved in the SiPhotonics'14 workshop, where short versions of the papers were previously selected. We are especially grateful to Profs. Geoffrey C. Fox and David W. Walker, editors of the journal, for approving this special issue and for his help along the process of its preparation.
Sandro Bartolini, José M. García 0001
Concurr. Comput. Pract. Exp.2
2014 Managing resources dynamically in hybrid photonic-electronic networks-on-chip
abstract
SUMMARY Nanophotonics promises to solve the scalability problems of current electrical interconnects thanks to its low sensitivity to distance in terms of latency and energy consumption. Before this technology reaches maturity, hybrid photonic‐electronic networks will be a viable alternative. Ideally, ordinary electrical meshes and ring‐based photonic networks should cooperate to minimize overall latency and energy consumption, but currently, we lack mechanisms to do this efficiently. In this paper, we present novel fine‐grain policies to manage the photonic resources in a tiled chip multiprocessor (CMP) scenario. Our policies are dynamic and base their decisions on parameters such as message size, ring availability, and distance between endpoints, at the message level. The resulting network behavior is also fairer to all cores, reducing processor idle time thanks to faster thread synchronization. All these policies improve performance when compared to the same CMP without the photonic ring, and the most elaborate ones reduce the overall network latency by 50%, execution time by 36%, and network energy consumption by 52% on average, in a 16‐core CMP for the PARSEC benchmark suite. Larger hybrid networks with 64 endpoints for 256‐core CMPs, based on Corona and Firefly designs, also show far superior throughput and lower latency if managed by one of the proposed policies. Copyright © 2014 John Wiley & Sons, Ltd.
Antonio García-Guirado, Ricardo Fernández-Pascual, José M. García 0001, Sandro Bartolini
Concurr. Comput. Pract. Exp.3
2014 Toward energy efficiency in heterogeneous processors: findings on virtual screening methods
abstract
ABSTRACT The integration of the latest breakthroughs in computational modeling and high performance computing (HPC) has leveraged advances in the fields of healthcare and drug discovery, among others. By integrating all these developments together, scientists are creating new exciting personal therapeutic strategies for living longer that were unimaginable not that long ago. However, we are witnessing the biggest revolution in HPC in the last decade. Several graphics processing unit architectures have established their niche in the HPC arena but at the expense of an excessive power and heat. A solution for this important problem is based on heterogeneity. In this paper, we analyze power consumption on heterogeneous systems, benchmarking a bioinformatics kernel within the framework of virtual screening methods. Cores and frequencies are tuned to further improve the performance or energy efficiency on those architectures. Our experimental results show that targeted low‐cost systems are the lowest power consumption platforms, although the most energy efficient platform and the best suited for performance improvement is the Kepler GK110 graphics processing unit from Nvidia by using compute unified device architecture. Finally, the open computing language version of virtual screening shows a remarkable performance penalty compared with its compute unified device architecture counterpart. Copyright © 2013 John Wiley & Sons, Ltd.
Ginés D. Guerrero, Juan M. Cebrian, Horacio Emilio Pérez Sánchez, José M. García 0001, Manuel Ujaldon, José M. Cecilia
Concurr. Comput. Pract. Exp.4
2014 A performance/cost model for a CUDA drug discovery application on physical and public cloud infrastructures
abstract
SUMMARY Virtual Screening (VS) methods can considerably aid drug discovery research, predicting how ligands interact with drug targets. BINDSURF is an efficient and fast blind VS methodology for the determination of protein binding sites, depending on the ligand, using the massively parallel architecture of graphics processing units(GPUs) for fast unbiased prescreening of large ligand databases. In this contribution, we provide a performance/cost model for the execution of this application on both local system and public cloud infrastructures. With our model, it is possible to determine which is the best infrastructure to use in terms of execution time and costs for any given problem to be solved by BINDSURF. Conclusions obtained from our study can be extrapolated to other GPU‐based VS methodologies.Copyright © 2013 John Wiley & Sons, Ltd.
Ginés D. Guerrero, Richard M. Wallace, José Luis Vázquez-Poletti, José M. Cecilia, José M. García 0001, Daniel Mozos, Horacio Emilio Pérez Sánchez
Concurr. Comput. Pract. Exp.5
2014 Evaluating the SAT problem on P systems for different high-performance architectures
José M. Cecilia, José M. García 0001, Ginés D. Guerrero, Manuel Ujaldon
J. Supercomput.2
2014 Comparative evaluation of platforms for parallel Ant Colony Optimization
Ginés D. Guerrero, José M. Cecilia, Antonio Llanes, José M. García 0001, Martyn Amos, Manuel Ujaldon
J. Supercomput.4
2014 ZEBRA: Data-Centric Contention Management in Hardware Transactional Memory
abstract
Transactional contention management policies show considerable variation in relative performance with changing workload characteristics. Consequently, incorporation of fixed-policy Transactional Memory (TM) in general purpose computing systems is suboptimal by design and renders such systems susceptible to pathologies. Of particular concern are Hardware TM (HTM) systems where traditional designs have hardwired policies in silicon. Adaptive HTMs hold promise, but pose major challenges in terms of design and verification costs. In this paper, we present the ZEBRA HTM design, which lays down a simple yet high-performance approach to implement adaptive contention management in hardware. Prior work in this area has associated contention with transactional code blocks. However, we discover that by associating contention with data (cache blocks) accessed by transactional code rather than the code block itself, we achieve a neat match in granularity with that of the cache coherence protocol. This leads to a design that is very simple and yet able to track closely or exceed the performance of the best performing policy for a given workload. ZEBRA, therefore, brings together the inherent benefits of traditional eager HTMs-parallel commits-and lazy HTMs-good optimistic concurrency without deadlock avoidance mechanisms-, combining them into a low-complexity design.
J. Rubén Titos Gil, Anurag Negi, Manuel E. Acacio, José M. García 0001, Per Stenström
IEEE Trans. Parallel Distributed Syst.4
2013 Improving drug discovery using a neural networks based parallel scoring function
abstract
Virtual Screening (VS) methods can considerably aid clinical research, predicting how ligands interact with drug targets. Most VS methods suppose a unique binding site for the target, but it has been demonstrated that diverse ligands interact with unrelated parts of the target and many VS methods do not take into account this relevant fact. This problem is circumvented by a novel VS methodology named BINDSURF that scans the whole protein surface to find new hotspots, where ligands might potentially interact with, and which is implemented in massively parallel Graphics Processing Units, allowing fast processing of large ligand databases. BINDSURF can thus be used in drug discovery, drug design, drug repurposing and therefore helps considerably in clinical research. However, the accuracy of most VS methods is constrained by limitations in the scoring function that describes biomolecular interactions, and even nowadays these uncertainties are not completely understood. In order to solve this problem, we propose a novel approach where neural networks are trained with databases of known active (drugs) and inactive compounds, and later used to improve VS predictions.
Horacio Emilio Pérez Sánchez, Ginés D. Guerrero, José M. García 0001, Jorge Peña-García, José M. Cecilia, Gaspar Cano, Sergio Orts, José García Rodríguez 0001
IJCNN3
2013 Enhancing data parallelism for Ant Colony Optimization on GPUs
José M. Cecilia, José M. García 0001, Andy Nisbet, Martyn Amos, Manuel Ujaldon
J. Parallel Distributed Comput.2
2013 Adaptive Neuromorphic Architecture (ANA)
Frank Wang, Leon O. Chua, Xiao Yang 0006, Na Helian, Ronald Tetzlaff, Torsten Schmidt, Caroline Li, José M. García 0001, Wanlong Chen, Dominique F. Chu
Neural Networks8
2013 Modeling the impact of permanent faults in caches
abstract
The traditional performance cost benefits we have enjoyed for decades from technology scaling are challenged by several critical constraints including reliability. Increases in static and dynamic variations are leading to higher probability of parametric and wear-out failures and are elevating reliability into a prime design constraint. In particular, SRAM cells used to build caches that dominate the processor area are usually minimum sized and more prone to failure. It is therefore of paramount importance to develop effective methodologies that facilitate the exploration of reliability techniques for caches. To this end, we present an analytical model that can determine for a given cache configuration, address trace, and random probability of permanent cell failure the exact expected miss rate and its standard deviation when blocks with faulty bits are disabled. What distinguishes our model is that it is fully analytical, it avoids the use of fault maps, and yet, it is both exact and simpler than previous approaches. The analytical model is used to produce the miss-rate trends ( expected miss-rate ) for future technology nodes for both uncorrelated and clustered faults. Some of the key findings based on the proposed model are (i) block disabling has a negligible impact on the expected miss-rate unless probability of failure is equal or greater than 2.6e-4, (ii) the fault map methodology can accurately calculate the expected miss-rate as long as 1,000 to 10,000 fault maps are used, and (iii) the expected miss-rate for execution of parallel applications increases with the number of threads and is more pronounced for a given probability of failure as compared to sequential execution.
Daniel Sánchez 0004, Yiannakis Sazeides, Juan M. Cebrian, José M. García 0001, Juan L. Aragón
ACM Trans. Archit. Code Optim.4
2013 Enhancing GPU parallelism in nature-inspired algorithms
José M. Cecilia, Andy Nisbet, Martyn Amos, José M. García 0001, Manuel Ujaldon
J. Supercomput.4
2013 Efficient Eager Management of Conflicts for Scalable Hardware Transactional Memory
abstract
The efficient management of conflicts among concurrent transactions constitutes a key aspect that hardware transactional memory (HTM) systems must achieve. Scalable HTM proposals so far inherit the cache-based style of conflict detection typically found in bus-based systems, largely unaware of the interactions between transactions and directory coherence. In this paper, we demonstrate that the traditional approach of detecting conflicts at the private cache levels is inefficient when used in the context of a directory protocol. We find that the use of the directory as a mere router of coherence requests restricts the throughput of conflict detection, and show how it becomes a bottleneck under high contention. This paper proposes a scheme for conflict detection that decouples conflict detection from cache coherence in order to overcome pathological situations that degrade the performance of an eager HTM system. Our scheme places bookkeeping metadata at the directory, introducing it as a separate hardware module that leaves the coherence protocol unmodified. In comparison to a state-of-the-art eager HTM system, our design handles contention more efficiently, minimizes the performance degradation of false positives for signatures of similar hardware cost, and reduces the network traffic generated.
J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001
IEEE Trans. Parallel Distributed Syst.3
2013 Eager Beats Lazy: Improving Store Management in Eager Hardware Transactional Memory
abstract
Hardware transactional memory (HTM) designs are very sensitive to the manner in which speculative updates from transactions are handled in the system. This study highlights how the lack of effective techniques for store management results in a quick degradation in the performance of eager HTM systems with increasing contention and, thus, lends credence to the belief that eager designs do not perform as well as their lazy counterparts when conflicts abound. In this work, we present two simple ways to improve handling of speculative stores--a way to effectively manage lines that exhibit migratory sharing and a way to hide store latency, particularly for those stores that target contended cache lines owned by other concurrent transactions. These two mechanisms yield substantial improvements in execution time when running applications with high contention, allowing eager designs to exceed the performance of lazy ones. Interestingly, the benefits that accrue from these enhancements can be at par with those achieved using more complex system-wide HTM techniques. Coupled with the fact that eager designs are easier to integrate into cache coherent architectures than lazy ones, we claim that with judicious management of stores they represent a more compelling design alternative.
J. Rubén Titos Gil, Anurag Negi, Manuel E. Acacio, José M. García 0001, Per Stenström
IEEE Trans. Parallel Distributed Syst.4
2012 π-TM: Pessimistic invalidation for scalable lazy hardware transactional memory
abstract
Lazy hardware transactional memory has been shown to be more efficient at extracting available concurrency than its eager counterpart. However, it poses scalability challenges at commit time as existence of conflicts among concurrent transactions is not known prior to commit. Non-conflicting transactions may have to wait before committing, severely affecting performance in certain workloads. Early conflict detection can be employed to allow such transactions to commit simultaneously. In this paper we show that the potential of this technique has not yet been fully utilized, with design choices in prior work severely burdening common-case transactional execution to avoid some relatively uncommon correctness concerns. The paper quantifies the severity of the problem and develops μ-TM, an early conflict detection - lazy conflict resolution design. This design highlights how, with modest extensions to existing directory-based coherence protocols, information regarding possible conflicts can be effectively used to achieve true parallelism at commit without burdening the common-case. We leverage the observation that contention is typically seen on only a small fraction of shared data accessed by coarse-grained transactions. Pessimistic invalidation of such lines when committing or aborting, therefore, enables fast common-case execution. Our results show that μ-TM performs consistently well and, in particular, far better than previous work on early conflict detection in lazy HTM. We also identify a pathological scenario that lazy designs with early conflict detection suffer from and propose a simple hardware workaround to sidestep it.
Anurag Negi, J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001, Per Stenström
HPCA4
2012 ASCIB: adaptive selection of cache indexing bits for removing conflict misses
abstract
The design of cache memories is a crucial part of the design cycle of a modern processor. Unfortunately, caches with low degrees of associativity suffer a large amount of conflict misses, while high-associative caches consume more power per access. We propose ASCIB, a simple technique able to dynamically adjust the bits used for cache indexing so as to minimize conflict misses. By selecting at run time the bits that disperse the working set more evenly across the available sets, ASCIB removes 73% of the conflict misses on average. This results in an improvement in energy efficiency by 17% on average.
Alberto Ros 0001, Polychronis Xekalakis, Marcelo Cintra, Manuel E. Acacio, José M. García 0001
ISLPED5
2012 Parallelization of Virtual Screening in Drug Discovery on Massively Parallel Architectures
abstract
The current trend in medical research for the discovery of new drugs is the use of Virtual Screening (VS) methods. In these methods, the calculation of the non-bonded interactions, such as electrostatics or van der Waals forces, plays an important role, representing up to 80% of the total execution time. These kernels are computational intensive and massively parallel in nature, and thus they are well suited to be accelerated on parallel architectures. In this work, we discuss the effective parallelization of the non-bonded electrostatic interactions kernel for VS on three different parallel architectures: a shared memory system, a distributed memory system, and a Graphics Processing Units (GPUs). For an efficient handling of the computational intensive and massively parallelism of this kernel, we enable different data policies on those architectures to take advantage of all computational resources offered by them. Four implementations are provided based on MPI, OpenMP, Hybrid MPI Open MP and CUDA programming models. The sequential implementation is defeated by a wide margin by all parallel implementations, obtaining up to 72x speed-up factor on the shared memory system through OpenMP, up to 60x and229x speed-ups factors on the distributed memory system for the MPI implementation and the Hybrid MPI-Open MP implementation respectively, and finally, up to 213x speedup factor for the CUDA implementation on the GPU architecture to offer the best alternative in terms of performance/cost ratio.
Ginés D. Guerrero, Horacio Emilio Pérez Sánchez, José M. Cecilia, José M. García 0001
PDP4
2012 Accelerating Fibre Orientation Estimation from Diffusion Weighted Magnetic Resonance Imaging Using GPUs
abstract
Diffusion Weighted Magnetic Resonance Imaging (DW-MRI) and tractography approaches are the only tools that can be utilized to estimate structural connections between different brain areas, non-invasively and in-vivo. A first step that is commonly utilized in these techniques includes the estimation of the underlying fibre orientations and their uncertainty in each voxel of the image. A popular method to achieve that is implemented in the FSL software, provided by the FMRIB Centre at University of Oxford, and is based on a Bayesian inference framework. Despite its popularity, the approach has high computational demands, taking normally more than 24 hours for analyzing a single subject. In this paper, we present a GPU-optimized version of the FSL tool that estimates fibre orientations. We report up to 85x of speed-up factor between the GPU and its sequential counterpart CPU-based version.
Moisés Hernández, Ginés D. Guerrero, José M. Cecilia, José M. García 0001, Alberto Inuggi, Stamatios N. Sotiropoulos
PDP4
2012 High-Throughput parallel blind Virtual Screening using BINDSURF
abstract
BACKGROUND: Virtual Screening (VS) methods can considerably aid clinical research, predicting how ligands interact with drug targets. Most VS methods suppose a unique binding site for the target, usually derived from the interpretation of the protein crystal structure. However, it has been demonstrated that in many cases, diverse ligands interact with unrelated parts of the target and many VS methods do not take into account this relevant fact. RESULTS: We present BINDSURF, a novel VS methodology that scans the whole protein surface in order to find new hotspots, where ligands might potentially interact with, and which is implemented in last generation massively parallel GPU hardware, allowing fast processing of large ligand databases. CONCLUSIONS: BINDSURF is an efficient and fast blind methodology for the determination of protein binding sites depending on the ligand, that uses the massively parallel architecture of GPUs for fast pre-screening of large ligand databases. Its results can also guide posterior application of more detailed VS methods in concrete binding sites of proteins, and its utilization can aid in drug discovery, design, repurposing and therefore help considerably in clinical research.
Irene Sánchez-Linares, Horacio Emilio Pérez Sánchez, José M. Cecilia, José M. García 0001
BMC Bioinform.4
2012 The GPU on the simulation of cellular computing models
José M. Cecilia, José M. García 0001, Ginés D. Guerrero, Miguel A. Martínez-del-Amor, Mario J. Pérez-Jiménez, Manuel Ujaldon
Soft Comput.2
2012 DAPSCO: Distance-aware partially shared cache organization
abstract
Many-core tiled CMP proposals often assume a partially shared last level cache (LLC) since this provides a good compromise between access latency and cache utilization. In this paper, we propose a novel way to map memory addresses to LLC banks that takes into account the average distance between the banks and the tiles that access them. Contrary to traditional approaches, our mapping does not group the tiles in clusters within which all the cores access the same bank for the same addresses. Instead, two neighboring cores access different sets of banks minimizing the average distance travelled by the cache requests. Results for a 64-core CMP show that our proposal improves both execution time and the energy consumed by the network by 13% when compared to a traditional mapping. Moreover, our proposal comes at a negligible cost in terms of hardware and its benefits in both energy and execution time increase with the number of cores.
Antonio García-Guirado, Ricardo Fernández-Pascual, Alberto Ros 0001, José M. García 0001
ACM Trans. Archit. Code Optim.4
2012 Hardware transactional memory with software-defined conflicts
abstract
In this paper we investigate the benefits of turning the concept of transactional conflict from its traditionally fixed definition into a variable one that can be dynamically controlled in software. We propose the extension of the atomic language construct with an attribute that specifies the definition of conflict, so that programmers can write code which adjusts what kinds of conflicts are to be detected, relaxing or tightening the conditions according to the forms of interference that can be tolerated by a particular algorithm. Using this performance-motivated construct, specific conflict information can be associated with portions of code, as each transaction is provided with a local definition that applies while it executes. We find that defining conflicts in software makes possible the removal of dependencies which arise as a result of the coarse synchronization style encouraged by the TM programming model. We illustrate the use of the proposed construct in a variety of use cases with real applications, showing how programmers can take advantage of their knowledge about the problem and other global information not available at run-time. We describe how to implement a hardware TM design that utilizes this software construct. Our experiments reveal that leveraging software-defined conflicts, the programmer is able to achieve significant reductions in the number of aborts--over 50% for most applications. At 16 threads, our system with software-defined conflicts outperforms LogTM-SE in nearly all benchmarks, reaching an average reduction in execution time of 18%.
J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001, Tim Harris 0001, Adrián Cristal, Osman S. Unsal, Ibrahim Hur, Mateo Valero
ACM Trans. Archit. Code Optim.3
2012 Extending Magny-Cours Cache Coherence
abstract
One cost-effective way to meet the increasing demand for larger high-performance shared-memory servers is to build clusters with off-the-shelf processors connected with low-latency point-to-point interconnections like HyperTransport. Unfortunately, HyperTransport addressing limitations prevent building systems with more than eight nodes. While the recent High-Node Count HyperTransport specification overcomes this limitation, recently launched twelve-core Magny-Cours processors have already inherited it and provide only 3 bits to encode the pointers used by the directory cache which they include to increase the scalability of their coherence protocol. In this work, we propose and develop an external device to extend the coherence domain of Magny-Cours processors beyond the 8-node limit while maintaining the advantages provided by the directory cache. Evaluation results for systems with up to 32 nodes show that the performance offered by our solution scales with the number of nodes, enhancing the directory cache effectiveness by filtering additional messages. Particularly, we reduce execution time by 47 percent in a 32-die system with respect to the 8-die Magny-Cours configuration.
Alberto Ros 0001, Blas Cuesta, Ricardo Fernández-Pascual, María Engracia Gómez, Manuel E. Acacio, Antonio Robles, José M. García 0001, José Duato
IEEE Trans. Computers7
2012 Stencil computations on heterogeneous platforms for the Jacobi method: GPUs versus Cell BE
José M. Cecilia, José L. Abellán, Juan Fernández Peinador, Manuel E. Acacio, José M. García 0001, Manuel Ujaldon
J. Supercomput.5
2012 A fault-tolerant architecture for parallel applications in tiled-CMPs
Daniel Sánchez 0004, Juan L. Aragón, José M. García 0001
J. Supercomput.3
2011 Pi-TM: Pessimistic Invalidation for Scalable Lazy Hardware Transactional Memory
abstract
Lazy hardware transactional memory (HTM) allows better utilization of available concurrency in transactional workloads than eager HTM, but poses challenges at commit time due to the requirement of en-masse publication of speculative updates to global system state. Early conflict detection can be employed in lazy HTM designs to allow non-conflicting transactions to commit in parallel. Though this has the potential to improve performance, it has not been utilized effectively so far. Prior work in the area burdens common-case transactional execution severely to avoid some relatively uncommon correctness concerns. In this work we investigate this problem and introduce a novel design, π-TM, which eliminates this problem. π-TM uses modest extensions to existing directory-based cache coherence protocols to keep a record of conflicting cache lines as a transaction executes. This information allows a consistent cache state to be maintained when transactions commit or abort. We observe that contention is typically seen only on a small fraction of shared data accessed by coarse-grained transactions. In π-TM early conflict detection mechanisms imply additional work only when such contention actually exists. Thus, the design is able to avoid expensive core-to-core and core-to-directory communication for a large part of transactionally accessed data. Our evaluation shows major performance gains when compared to other HTM designs in this class and competitive performance when compared to more complex lazy commit schemes.
Anurag Negi, Per Stenström, J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001
PACT5
2011 Energy-Efficient Cache Coherence Protocols in Chip-Multiprocessors for Server Consolidation
abstract
As the number of cores in a chip increases, power consumption is becoming a major constraint in the design of chip multiprocessors. At the same time, server consolidation is gaining importance to take advantage of such a number of cores. Our goal is to alleviate this constraint by reducing the power consumption of chip multiprocessors used for consolidated workloads by means of the cache coherence protocol. For this, we statically divide the chip in areas, which allows us to reduce the directory overhead needed to support coherence and to reduce the network traffic. This translates into less power consumption without performance degradation. Cache coherence is maintained per area and pointers are used to link the areas, thereby achieving isolation among virtual machines and savings in memory requirements. Additionally, the coherence protocol dynamically selects one node per area as responsible for providing the data on a cache miss, thus lessening the average cache miss latency and the traffic among areas. Compared to a highly-optimized directory implementation, the leakage power consumption is reduced by 54% and the dynamic power consumption of the caches and the network-on-chip by up to 38% for a 64-tile chip multiprocessor with 4 virtual machines, showing no performance degradation.
Antonio García-Guirado, Ricardo Fernández-Pascual, Alberto Ros 0001, José M. García 0001
ICPP4
2011 Eager Meets Lazy: The Impact of Write-Buffering on Hardware Transactional Memory
abstract
Hardware transactional memory (HTM) systems have been studied extensively along the dimensions of speculative versioning and contention management policies. The relative performance of several designs policies has been discussed at length in prior work within the framework of scalable chip-multiprocessing systems. Yet, the impact of simple structural optimizations like write-buffering has not been investigated and performance deviations due to the presence or absence of these optimizations remains unclear. This lack of insight into the effective use and impact of these interfacial structures between the processor core and the coherent memory hierarchy forms the crux of the problem we study in this paper. Through detailed modeling of various write-buffering configurations we show that they play a major role in determining the overall performance of a practical HTM system. Our study of both eager and lazy conflict resolution mechanisms in a scalable parallel architecture notes a remarkable convergence of the performance of these two diametrically opposite design points when write buffers are introduced and used well to support the common case. Mitigation of redundant actions, fewer invalidations on abort, latency-hiding and prefetch effects contribute towards reducing execution times for transactions. Shorter transaction durations also imply a lower contention probability, thereby amplifying gains even further. The insights, related to the interplay between buffering mechanisms, system policies and workload characteristics, contained in this paper clearly distinguish gains in performance to be had from write-buffering from those that can be ascribed to HTM policy. We believe that this information would facilitate sound design decisions when incorporating HTMs into parallel architectures.
Anurag Negi, J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001, Per Stenström
ICPP4
2011 ZEBRA: a data-centric, hybrid-policy hardware transactional memory design
abstract
Hardware Transactional Memory (HTM) systems, in prior research, have either fixed policies of conflict resolution and data versioning for the entire system or allowed a degree of flexibility at the level of transactions. Unfortunately, this results in susceptibility to pathologies, lower average performance over diverse workload characteristics or high design complexity. In this work we explore a new dimension along which flexibility in policy can be introduced. Recognizing the fact that contention is more a property of data rather than that of an atomic code block, we develop an HTM system that allows selection of versioning and conflict resolution policies at the granularity of cache lines. We discover that this neat match in granularity with that of the cache coherence protocol results in a design that is very simple and yet able to track closely or exceed the performance of the best performing policy for a given workload. It also brings together the benefits of parallel commits (inherent in traditional eager HTMs) and good optimistic concurrency without deadlock avoidance mechanisms (inherent in lazy HTMs), with little increase in complexity.
J. Rubén Titos Gil, Anurag Negi, Manuel E. Acacio, José M. García 0001, Per Stenström
ICS4
2011 An analytical model for the calculation of the Expected Miss Ratio in faulty caches
abstract
Technology scaling improvement is affecting the reliability of ICs due to increases in static and dynamic variations as well as wear-out failures. This is particularly true for caches that dominate the area of modern processors and are built with minimum-sized, but prone to failure, SRAM cells. Our attempt to address this cache reliability challenge is an analytical model for determining the implications on cache miss-rate of block-disabling due to random cell failure. The proposed model is distinct from previous work in that is an exact model rather than an approximation and yet it is simpler than previous work. Its simplicity stems from the lack of fault-maps in the analysis. The model capabilities are illustrated through a study of cache miss-rate trends in future technology nodes. The model is also used to determine the accuracy of a random fault map methodology. The analysis reveals, for the assumptions, programs and cache configuration used in this study, a surprising result: a relative small number of random fault maps, 100–1000, is sufficient to obtain accurate mean and standard-deviation values for the miss-rate. Additional investigation revealed that the cause of this behavior is a high correlation between the number of accesses and access distribution between cache sets.
Daniel Sánchez 0004, Yiannakis Sazeides, Juan L. Aragón, José M. García 0001
IOLTS4
2011 Leakage-efficient design of value predictors through state and non-state preserving techniques
Juan M. Cebrian, Juan L. Aragón, José M. García 0001, Stefanos Kaxiras
J. Supercomput.3
2010 EMC2: Extending Magny-Cours coherence for large-scale servers
abstract
The demand of larger and more powerful high-performance shared-memory servers is growing over the last few years. To meet this need, AMD has recently launched the twelve-core Magny-Cours processors. They include a directory cache (Probe Filter) that increases the scalability of the coherence protocol applied by Opterons, based on coherent Hyper Transport interconnect (cHT). cHT limits up to 8 the number of nodes that can be addressed. Recent High Node Count HT specification overcomes this limitation. However, the 3-bit pointer used by the Probe Filter prevents Magny-Cours-based servers from being built beyond 8 nodes. In this paper, we propose and develop an external logic to extend the coherence domain of Magny-Cours processors beyond the 8-node limit while maintaining the advantages provided by the Probe Filter. Evaluation results for up to a 32-node system show how the performance offered by our solution scales with the increment in the number of nodes, enhancing the Probe Filter effectiveness by filtering additional messages. Particularly, we reduce runtime by 47% in a 32-die system respect to the 8-die Magny-Cours system.
Alberto Ros 0001, Blas Cuesta, Ricardo Fernández-Pascual, María Engracia Gómez, Manuel E. Acacio, Antonio Robles, José M. García 0001, José Duato
HiPC7
2010 A log-based redundant architecture for reliable parallel computation
abstract
CMOS scaling exacerbates hardware errors making reliability a big concern for recent and future microarchitecture designs. Mechanisms to provide fault tolerance in architectures must accomplish several objectives such as low performance degradation, power consumption and area overhead. Several studies have been already proposed to provide fault tolerance for parallel codes. However, these proposals are usually implemented over non-realistic environments including the use of shared-buses among processors or modifying highly optimized hardware designs such as caches. Our main design goal is to provide transient fault detection and recovery while modifying hardware as less as possible. To this end, we propose LBRA based on a Hardware Transactional Memory (HTM) architecture in which two redundant threads successfully detects and recovers from transient faults, assuring a consistent view of the memory by means of a pair-shared cacheable virtual memory log which keeps the computation results. Results show that our log-based mechanism introduces a small performance degradation of 5% in a non-faulty scenario. Additionally, we show that LBRA supports huge fault rates such as 100 faults per million of cycles with low additional performance degradation.
Daniel Sánchez 0004, Juan L. Aragón, José M. García 0001
HiPC3
2010 Analyzing Cache Coherence Protocols for Server Consolidation
abstract
Server consolidation is commonly used today to make the most out of all the cores of a chip multiprocessor by running several virtual machines (VMs) on it. Cache coherence protocols can be adapted to take advantage of such an scenario. In this line, Virtual Hierarchies (VHs) use two levels of cache coherence in a consolidated server. They isolate the coherence actions of each VM and improve performance by maximizing the number of memory accesses serviced by caches within the VM. In this paper we show how hierarchical protocols with no single ordering point for the requests, such as VHs in the form currently proposed, are prone to deadlocks. Besides, when memory deduplication is used, VHs cannot take advantage of memory deduplication at the cache level, both because deduplicated data is reduplicated in cache, and because accesses to deduplicated data often require the access to the cache tiles used by a different VM by means of broadcast. We analyze all these problems and we propose solutions for them, showing the actual performance of these protocols, and giving some insights for the future development of coherence protocols optimized for server consolidation.
Antonio García-Guirado, Ricardo Fernández-Pascual, José M. García 0001
SBAC-PAD3
2010 Simulation of P systems with active membranes on CUDA
abstract
P systems or Membrane Systems provide a high-level computational modelling framework that combines the structure and dynamic aspects of biological systems in a relevant and understandable way. They are inherently parallel and non-deterministic computing devices. In this article, we discuss the motivation, design principles and key of the implementation of a simulator for the class of recognizer P systems with active membranes running on a (GPU). We compare our parallel simulator for GPUs to the simulator developed for a single central processing unit (CPU), showing that GPUs are better suited than CPUs to simulate P systems due to their highly parallel nature.
José M. Cecilia, José M. García 0001, Ginés D. Guerrero, Miguel A. Martínez-del-Amor, Ignacio Pérez-Hurtado, Mario J. Pérez-Jiménez
Briefings Bioinform.2
2010 A scalable organization for distributed directories
Alberto Ros 0001, Manuel E. Acacio, José M. García 0001
J. Syst. Archit.3
2010 Dealing with Transient Faults in the Interconnection Network of CMPs at the Cache Coherence Level
abstract
The importance of transient faults is predicted to grow due to current technology trends of increased scale of integration. One of the components that will be significantly affected by transient faults is the interconnection network of chip multiprocessors (CMPs). To deal efficiently with these faults and differently from other authors, we propose to use fault-tolerant cache coherence protocols that ensure the correct execution of programs when not all messages are correctly delivered. We describe the extensions made to a directory-based cache coherence protocol to provide fault tolerance and provide a modified set of token counting rules which are useful to design fault-tolerant token-based cache coherence protocols. We compare the directory-based fault-tolerant protocol with a token-based fault-tolerant one. We also show how to adjust the fault tolerance parameters to achieve the desired level of fault tolerance and measure the overhead achieved to be able to support very high fault rates. Simulation results using a set of scientific, multimedia, and commercial applications show that the fault tolerance measures have virtually no impact on execution time with respect to a non-fault-tolerant protocol. Additionally, our protocols can support very high rates of transient faults at the cost of slightly increased network traffic.
Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José Duato
IEEE Trans. Parallel Distributed Syst.2
2010 A Direct Coherence Protocol for Many-Core Chip Multiprocessors
abstract
Future many-core CMP designs that will integrate tens of processor cores on-chip will be constrained by area and power. Area constraints make impractical the use of a bus or a crossbar as the on-chip interconnection network, and tiled CMPs organized around a direct interconnection network will probably be the architecture of choice. Power constraints make impractical to rely on broadcasts (as, for example, Token-CMP does) or any other brute-force method for keeping cache coherence, and directory-based cache coherence protocols are currently being employed. Unfortunately, directory protocols introduce indirection to access directory information, which negatively impacts performance. In this work, we present DiCo-CMP, a novel cache coherence protocol especially suited to future many-core tiled CMP architectures. In DiCo-CMP, the task of storing up-to-date sharing information and ensuring ordered accesses for every memory block is assigned to the cache that must provide the block on a miss. Therefore, DiCo-CMP reduces the miss latency compared to a directory protocol by sending requests directly to the cache that provides the block in a cache miss. These latency reductions result in improvements in execution time of up to 6 percent, on average, over a directory protocol. In comparison with Token-CMP, our protocol only sends one request message for each cache miss, as such is able to reduce network traffic by 43 percent.
Alberto Ros 0001, Manuel E. Acacio, José M. García 0001
IEEE Trans. Parallel Distributed Syst.3
2009 Dealing with Traffic-Area Trade-Off in Direct Coherence Protocols for Many-Core CMPs
Alberto Ros 0001, Manuel E. Acacio, José M. García 0001
APPT3
2009 REPAS: Reliable Execution for Parallel ApplicationS in Tiled-CMPs
Daniel Sánchez 0004, Juan L. Aragón, José M. García 0001
Euro-Par3
2009 Distance-aware round-robin mapping for large NUCA caches
abstract
In many-core architectures, memory blocks are commonly assigned to the banks of a NUCA cache by following a physical mapping. This mapping assigns blocks to cache banks in a round-robin fashion, thus neglecting the distance between the cores that most frequently access every block and the corresponding NUCA bank for the block. This issue impacts both cache access latency and the amount of on-chip network traffic generated. On the other hand, first-touch mapping policies, which take into account distance, can lead to an unbalanced utilization of cache banks, and consequently, to an increased number of expensive off-chip accesses. In this work, we propose the distance-aware round-robin mapping policy, an OS-managed policy which addresses the trade-off between cache access latency and number of off-chip accesses. Our policy tries to map the pages accessed by a core to its closest (local) bank, like in a first-touch policy. However, our policy also introduces an upper bound on the deviation of the distribution of memory pages among cache banks, which lessens the number of off-chip accesses. This tradeoff is addressed without requiring any extra hardware structure. We also show that the private cache indexing commonly used in many-core architectures is not the most appropriate for OS-managed distance-aware mapping policies, and propose to employ different bits for such indexing. Using GEMS simulator we show that our proposal obtains average improvements of 11% for parallel applications and 14% for multi-programmed workloads in terms of execution time, and significant reductions in network traffic, over a traditional physical mapping. Moreover, when compared to a first-touch mapping policy, our proposal improves average execution time by 5% for parallel applications and 6% for multi-programmed workloads, slightly increasing on-chip network traffic.
Alberto Ros 0001, Marcelo Cintra, Manuel E. Acacio, José M. García 0001
HiPC4
2009 Efficient microarchitecture policies for accurately adapting to power constraints
abstract
In the past years dynamic voltage and frequency scaling (DVFS) has been an effective technique that allowed microprocessors to match a predefined power budget. However, as process technology shrinks, DVFS becomes less effective (because of the increasing leakage power) and it is getting closer to a point where DVFS won't be useful at all (when static power exceeds dynamic power). In this paper we propose the use of microarchitectural techniques to accurately match a power constraint while maximizing the energy efficiency of the processor. We predict the processor power consumption at a basic block level, using the consumed power translated into tokens to select between different power-saving micro-architectural techniques. These techniques are orthogonal to DVFS so they can be simultaneously applied. We propose a two-level approach where DVFS acts as a coarse-grained technique to lower the average power while microarchitectural techniques remove all the power spikes efficiently. Experimental results show that the use of power-saving microarchitectural techniques in conjunction with DVFS is up to six times more precise, in terms of total energy consumed (area) over the power budget, than using DVFS alone for matching a predefined power budget. Furthermore, in a near future DVFS will become DFS because lowering the supply voltage will be too expensive in terms of leakage power. At that point, the use of power-saving microarchitectural techniques will become even more energy efficient.
Juan M. Cebrian, Juan L. Aragón, José M. García 0001, Pavlos Petoumenos, Stefanos Kaxiras
IPDPS3
2009 Speculation-based conflict resolution in hardware transactional memory
abstract
Conflict management is a key design dimension of hardware transactional memory (HTM) systems, and the implementation of efficient mechanisms for detection and resolution becomes critical when conflicts are not a rare event. Current designs address this problem from two opposite perspectives, namely, lazy and eager schemes. While the former approach is based on an purely optimistic view that is not well-suited when conflicts become frequent, the latter results too pessimistic because resolves conflicts too conservatively, often limiting concurrency unnecessarily. In this paper, we present a hybrid, pseudo-optimistic scheme of conflict resolution for HTM systems that recaptures the concept of speculation to allow transactions to continue their execution past conflicting accesses. Simulation results show that our proposal is capable of combining the advantages of both classical approaches. For the STAMP transactional benchmarks, our hybrid scheme outperforms both eager and lazy systems with average reductions in execution time of 8 and 17%, respectively, and it decreases network traffic by another 17% compared to the eager policy.
J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001
IPDPS3
2009 Extending SRT for parallel applications in tiled-CMP architectures
abstract
Reliability has become a first-class consideration issue for architects along with performance and energy-efficiency. The increasing scaling technology and subsequent supply voltage reductions are increasing the susceptibility of architectures to soft errors. However, mechanisms to achieve full coverage to errors usually degrade performance in an unacceptable way for the majority of common users. Simultaneous and Redundantly Threaded (SRT) is a fault tolerant architecture in which pairs of threads in a SMT core redundantly execute the same program instructions. In this paper, we study the under-explored architectural support of SRT to reliably execute shared-memory applications. We show how atomic operations induce a serialization point between master and slave threads. This bottleneck has an impact of 34% in execution speed for several parallel scientific benchmarks. We propose an alternative mechanism in which the L1 cache is updated by master's stores before verification reducing the overhead up to 21%. Our approach also outperforms other recent proposals such as DCC with a decrease of 8% in execution speed.
Daniel Sánchez 0004, Juan L. Aragón, José M. García 0001
IPDPS3
2009 A lossy 3D wavelet transform for high-quality compression of medical video
Gregorio Bernabé, José M. García 0001, José González 0002
J. Syst. Softw.2
2008 A fault-tolerant directory-based cache coherence protocol for CMP architectures
abstract
Current technology trends of increased scale of integration are pushing CMOS technology into the deep-submicron domain, enabling the creation of chips with a significantly greater number of transistors but also more prone to transient failures. Hence, computer architects will have to consider reliability as a prime concern for future chip-multiprocessor designs (CMPs). Since the interconnection network of future CMPs will use a significant portion of the chip real state, it will be especially affected by transient failures. We propose to deal with this kind of failures at the level of the cache coherence protocol instead of ensuring the reliability of the network itself. Particularly, we have extended a directory-based cache coherence protocol to ensure correct program semantics even in presence of transient failures in the interconnection network. Additionally, we show that our proposal has virtually no impact on execution time with respect to a non fault-tolerant protocol, and just entails modest hardware and network traffic overhead.
Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José Duato
DSN2
2008 Fault-Tolerant Cache Coherence Protocols for CMPs: Evaluation and Trade-Offs
Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José Duato
HiPC2
2008 Directory-Based Conflict Detection in Hardware Transactional Memory
J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001
HiPC3
2008 DiCo-CMP: Efficient cache coherency in tiled CMP architectures
abstract
Future CMP designs that will integrate tens of processor cores on-chip will be constrained by area and power. Area constraints make impractical the use of a bus or a crossbar as the on-chip interconnection network, and tiled CMPs organized around a direct interconnection network will probably be the architecture of choice. Power constraints make impractical to rely on broadcasts (as Token-CMP does) or any other brute-force method for keeping cache coherence, and directory-based cache coherence protocols are currently being employed. Unfortunately, directory protocols introduce indirection to access directory information, which negatively impacts performance. In this work, we present DiCo-CMP, a novel cache coherence protocol especially suited to future tiled CMP architectures. In DiCo- CMP the role of storing up-to-date sharing information and ensuring totally ordered accesses for every memory block is assigned to the cache that must provide the block on a miss. Therefore, DiCo-CMP reduces the miss latency compared to a directory protocol by sending coherence messages directly from the requesting caches to those that must observe them (as it would be done in brute-force protocols), and reduces the network traffic compared to Token-CMP (and consequently, power consumption in the interconnection network) by sending just one request message for each miss. Using an extended version of GEMS simulator we show that DiCo-CMP achieves improvements in execution time of up to 8% on average over a directory protocol, and reductions in terms of network traffic of up to 42% on average compared to Token-CMP.
Alberto Ros 0001, Manuel E. Acacio, José M. García 0001
IPDPS3
2008 Characterization of Conflicts in Log-Based Transactional Memory (LogTM)
abstract
The difficulty of multithreaded programming remains a major obstacle for programmers to fully exploit multicore chips. Transactional memory has been proposed as an abstraction capable of ameliorating the challenges of traditional lock-based parallel programming. Hardware transactional memory (HTM) systems implement the necessary mechanisms to provide transactional semantics efficiently. In order to keep hardware simple, current HTM designs apply fixed policies that aim at optimizing the most expected application behaviour, and many of these proposals explicitly assume that commits will be clearly more frequent than aborts in future transactional workloads. This paper shows that some applications developed under the TM programming model are by nature prone to experience many conflicts. As a result, aborted transactions can get to be common and may seriously hurt performance. Our characterization, performed with truly transactional benchmarks on the LogTM system, shows that certain programs composed by large transactions suffer indeed very high abort rates. Thus, if TM is to unburden developers from the programmability-performance trade-off, HTM systems must obtain good performance levels in the presence of frequent aborts, requiring more flexible policies of data versioning as well as more sophisticated recovery schemes.
J. Rubén Titos Gil, Manuel E. Acacio, José M. García 0001
PDP3
2008 Two proposals for the inclusion of directory information in the last-level private caches of glueless shared-memory multiprocessors
Alberto Ros 0001, Ricardo Fernández-Pascual, Manuel E. Acacio, José M. García 0001
J. Parallel Distributed Comput.4
2008 Extending the TokenCMP Cache Coherence Protocol for Low Overhead Fault Tolerance in CMP Architectures
abstract
It is widely accepted that transient failures will appear more frequently in chips designed in the near future due to several factors such as the increased integration scale. On the other hand, Chip-multiprocessors (CMP) that integrate several processor cores in a single chip are nowadays the best alternative to more efficient use of the increasing number of transistors that can be placed in a single die. Hence, it is necessary to design new techniques to deal with these faults to be able to build sufficiently reliable Chip Multiprocessors (CMPs). In this work, we present a coherence protocol aimed at dealing with transient failures that affect the interconnection network of a CMP, thus assuming that the network is no longer reliable. In particular, our proposal extends a token-based cache coherence protocol so that no data can be lost and no deadlock can occur due to any dropped message. Using GEMS full system simulator, we compare our proposal against TokenCMP. We show that in absence of failures our proposal does not introduce overhead in terms of increased execution time over TokenCMP. Additionally, our protocol can tolerate message loss rates much higher than those likely to be found in the real world without increasing execution time more than 15%.
Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José Duato
IEEE Trans. Parallel Distributed Syst.2
2007 Direct Coherence: Bringing Together Performance and Scalability in Shared-Memory Multiprocessors
Alberto Ros 0001, Manuel E. Acacio, José M. García 0001
HiPC3
2007 A Low Overhead Fault Tolerant Coherence Protocol for CMP Architectures
abstract
It is widely accepted that transient failures will appear more frequently in chips designed in the near future due to several factors such as the increased integration scale. On the other hand, chip-multiprocessors (CMP) that integrate several processor cores in a single chip are nowadays the best alternative to more efficient use of the increasing number of transistors that can be placed in a single die. Hence, it is necessary to design new techniques to deal with these faults to be able to build sufficiently reliable chip multiprocessors (CMPs). In this work, we present a coherence protocol aimed at dealing with transient failures that affect the interconnection network of a CMP, thus assuming that the network is no longer reliable. In particular, our proposal extends a token-based cache coherence protocol so that no data can be lost and no deadlock can occur due to any dropped message. Using GEMS full system simulator, we compare our proposal against a similar protocol without fault tolerance (TOKENCMP). We show that in absence of failures our proposal does not introduce overhead in terms of increased execution time over TOKENCMP. Additionally, our protocol can tolerate message loss rates much higher than those likely to be found in the real world without increasing execution time more than 15%
Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José Duato
HPCA2
2007 Leakage Energy Reduction in Value Predictors through Static Decay
abstract
As process technology advances toward deep submicron (below 90 nm), static power becomes a new challenge to address for energy-efficient high performance processors, especially for large on-chip array structures such as caches and prediction tables. Value prediction appeared as an effective way of increasing processor performance by overcoming data dependences, but at the risk of becoming a thermal hot spot due to the additional power dissipation. This paper proposes the design of low-leakage value predictors by applying static decay techniques in order to disable unused entries from the prediction tables. We explore decay strategies suited for the three most common Value predictors (STP, FCM and DFCM) studying the particular tradeoffs for these prediction structures. Our mechanism reduces VP leakage energy efficiently without compromising VP accuracy nor processor performance. Results show average leakage energy reductions of 52%, 65% and 75% for the STP, DFCM and FCM value predictors, respectively.
Juan M. Cebrian, Juan L. Aragón, José M. García 0001
IPDPS3
2007 Aspect-Oriented Programing Techniques to support Distribution, Fault Tolerance, and Load Balancing in the CORBA-LC Component Model
abstract
The design and implementation of distributed High Performance Computing (HPC) applications is becoming harder as the scale and number of distributed resources and application is growing. Programming abstractions, libraries and frameworks are needed to better overcome that complexity. Moreover, when Quality of Service (QoS) requirements such as load balancing, efficient resource usage and fault tolerance have to be met, the resulting code is harder to develop, maintain, and reuse, as the code for providing the QoS requirements gets normally mixed with the functionality code. Component Technology, on the other hand, allows a better modularity and reusability of applications and even a better support for the development of distributed applications, as those applications can be partitioned in terms of components installed and running (deployed) in the different hosts participating in the system. Components also have requirements in forms of the aforementioned non-functional aspects. In our approach, the code for ensuring these aspects can be automatically generated based on the requirements stated by components and applications, thus leveraging the component implementer of having to deal with these non-functional aspects. In this paper we present the characteristics and the convenience of the generated code for dealing with load balancing, distribution, and fault-tolerance aspects in the context of CORBA-LC. CORBA-LC is a lightweight distributed reflective component model based on CORBA that imposes a peer network model in which the whole network acts as a repository for managing and assigning the whole set of resources: components, CPU cycles, memory, etc.
Diego Sevilla Ruiz, José M. García 0001, Antonio F. Skarmeta
NCA2
2007 An efficient implementation of a 3D wavelet transform based encoder on hyper-threading technology
Gregorio Bernabé, Ricardo Fernández-Pascual, José M. García 0001, Manuel E. Acacio, José González 0002
Parallel Comput.3
2007 The Design of New Journaling File Systems: The DualFS Case
abstract
This paper describes the foundation, design, implementation, and evaluation of DualFS, a new high-performance journaling file system which has the same consistency guarantees as traditional journaling file systems but a greater performance. DualFS places data and metadata in different devices (usually, two partitions of the same storage device) and manages them in very different ways. The metadata device is organized as a log-structured file system, whereas the data device is organized as groups. The new design allows DualFS not only to recover the consistency quickly after a system crash, but also to improve the overall file system performance. We have evaluated DualFS and we have found that it greatly reduces the total I/O time taken by the file system in most workloads as compared to other file systems (Ext2, Ext3, ReiserFS, XFS, and JFS). The work carried out has also allowed us to draw some lessons which ought to be taken into account when implementing new file systems
Juan Piernas, Toni Cortes, José M. García 0001
IEEE Trans. Computers3
2005 A Novel Lightweight Directory Architecture for Scalable Shared-Memory Multiprocessors
Alberto Ros 0001, Manuel E. Acacio, José M. García 0001
Euro-Par3
2005 Memory Subsystem Characterization in a 16-Core Snoop-Based Chip-Multiprocessor Architecture
Francisco J. Villa, Manuel E. Acacio, José M. García 0001
HPCC3
2005 Evaluating IA-32 web servers through simics: a practical experience
Francisco J. Villa, Manuel E. Acacio, José M. García 0001
J. Syst. Archit.3
2005 A Two-Level Directory Architecture for Highly Scalable cc-NUMA Multiprocessors
abstract
One important issue the designer of a scalable shared-memory multiprocessor must deal with is the amount of extra memory required to store the directory information. It is desirable that the directory memory overhead be kept as low as possible, and that it scales very slowly with the size of the machine. Unfortunately, current directory architectures provide scalability at the expense of performance. This work presents a scalable directory architecture that significantly reduces the size of the directory for large-scale configurations of a multiprocessor without degrading performance. First, we propose multilayer clustering as an effective approach to reduce the width of directory entries. Based on this concept, we derive three new compressed sharing codes, some of them with a space complexity of O(log/sub 2/(log/sub 2/(N))) for an N-node system. Then, we present a novel two-level directory architecture to eliminate the penalty caused by compressed directories in general. The proposed organization consists of a small full-map first-level directory (which provides precise information for the most recently referenced lines) and a compressed second-level directory (which provides in-excess information for all the lines). The proposals are evaluated based on extensive execution-driven simulations (using RSIM) of a 64-node cc-NUMA multiprocessor. Results demonstrate that a system with a two-level directory architecture achieves the same performance as a multiprocessor with a big and nonscalable full-map directory, with a very significant reduction of the memory overhead.
Manuel E. Acacio, José González 0002, José M. García 0001, José Duato
IEEE Trans. Parallel Distributed Syst.3
2004 An Architecture for High-Performance Scalable Shared-Memory Multiprocessors Exploiting On-Chip Integration
abstract
Recent technology improvements allow multiprocessor designers to put some key components inside the processor chip, such as the memory controller, the coherence hardware, and the network interface/router. In this paper, we exploit such integration scale, presenting a novel node architecture aimed at reducing the long L2 miss latencies and the memory overhead of using directories that characterize cc-NUMA machines and limit their scalability. Our proposal replaces the traditional directory with a novel three-level directory architecture, as well as it adds a small shared data cache to each of the nodes of a multiprocessor system. Due to their small size, the first-level directory and the shared data cache are integrated into the processor chip in every node, which enhances performance by saving accesses to the slower main memory. Scalability is guaranteed by having the second and third-level directories out of the processor chip and using compressed data structures. A taxonomy of the L2 misses, according to the actions performed by the directory to satisfy them, is also presented. Using execution-driven simulations, we show that significant latency reductions can be obtained by using the proposed node architecture, which translates into reductions of more than 30 percent in several cases in the application execution time.
Manuel E. Acacio, José González 0002, José M. García 0001, José Duato
IEEE Trans. Parallel Distributed Syst.3
2003 Real-Time Extraction of Colored Segments for Robot Visual Navigation
Pedro E. López-de-Teruel, Alberto Ruiz, Ginés García-Mateos, José M. García 0001
ICVS4
2002 DualFS: a new journaling file system without meta-data duplication
abstract
In this paper we introduce DualFS, a new high performance journaling file system that puts data and meta-data on different devices (usually, two partitions on the same disk or on different disks), and manages them in very different ways. Unlike other journaling file systems, DualFS has only one copy of every meta-data block. This copy is in the meta-data device, a log which is used by DualFS both to read and to write meta-data blocks. By avoiding a time-expensive extra copy of meta-data blocks, DualFS can achieve a good performance as compared to other journaling file systems. Indeed, we have implemented a DualFS prototype, which has been evaluated with microbenchmarks and macrobenchmarks, and we have found that DualFS greatly reduces the total I/O time taken by the file system in most cases (up to 97%), whereas it slightly increases the total I/O time only in a few and limited cases.
Juan Piernas, Toni Cortes, José M. García 0001
ICS3
2002 Improving the Performance of Real-Time Communication Services on High-Speed LANs under Topology Changes
abstract
In this paper, we propose and evaluate a new protocol that provides topology change- and fault-tolerant real-time communication services on NOW and clusters. This protocol overcomes the main drawback of our previously proposed protocol, called Dynamically Re-established Real-Time Channels (DRRTC), which is physically limited by the number of virtual channels per port. The new protocol allows different real-time channels to share the same virtual channel. In this way, the new protocol allows us to establish a greater number of real-time channels than the previous one. Moreover, its only limitation is the bandwidth devoted to real-time traffic. However, this introduces two new problems that are successfully managed by the new protocol: the existence of cyclic dependencies among different real-time channels and the increased complexity of deadline requirements. We present and analyze the performance evaluation results when a single switch or a single link is deactivated/activated for different topologies and workloads. The new protocol overwhelms the DRRTC protocol while guaranteeing deadline requirements and channel recovery.
Juan Fernández Peinador, José M. García 0001, José Duato
LCN2
2002 Owner prediction for accelerating cache-to-cache transfer misses in a cc-NUMA architecture
abstract
Cache misses for which data must be obtained from a remote cache (cache-to-cache transfer misses) account for an important fraction of the total miss rate. Unfortunately, cc-NUMA designs put the access to the directory information into the critical path of 3-hop misses, which significantly penalizes them compared to SMP designs. This work studies the use of owner prediction as a means of providing cc-NUMA multiprocessors with a more efficient support for cache-to-cache transfer misses. Our proposal comprises an effective prediction scheme as well as a coherence protocol designed to support the use of prediction. Results indicate that owner prediction can significantly reduce the latency of cache-to-cache transfer misses, which translates into speed-ups on application performance up to 12%. In order to also accelerate most of those 3-hop misses that are either not predicted or mispredicted, the inclusion of a small and fast directory cache in every node is evaluated, leading to improvements up to 16% on the final performance.
Manuel E. Acacio, José González 0002, José M. García 0001, José Duato
SC3
2002 MPI-Delphi: an MPI implementation for visual programming environments and heterogeneous computing
Manuel E. Acacio, Óscar Cánovas Reverte, José M. García 0001, Pedro E. López-de-Teruel
Future Gener. Comput. Syst.3
2001 On Deadlock Frequency during Dynamic Reconfiguration in NOWs
Lorenzo Fernández Maimó, José M. García 0001, Rafael Casado
Euro-Par2
2001 CORBA Lightweight Components : A Model for Distributed Component-BasedHeterogeneous Computation
Diego Sevilla Ruiz, José M. García 0001, Antonio F. Skarmeta
Euro-Par2
2001 Confidence Estimation for Branch Prediction Reversal
Juan L. Aragón, José González 0002, José M. García 0001, Antonio González 0001
HiPC3
2001 Performance Evaluation of Real-Time Communication Services on High-Speed LANs under Topology Changes
Juan Fernández Peinador, José M. García 0001, José Duato
HiPC2
2001 A New Scalable Directory Architecture for Large-Scale Multiprocessors
abstract
The memory overhead introduced by directories constitutes a major hurdle in the scalability of cc-NUMA architectures, which makes the shared-memory paradigm unfeasible for very large-scale systems. This work is focused on improving the scalability of shared-memory multiprocessors by significantly reducing the size of the directory. We propose multilayer clustering as an effective approach to reduce the directory-entry width. Detailed evaluation for 64 processors shows that using this approach we can drastically reduce the memory overhead, while suffering a performance degradation we similar to previous compressed schemes (such as Coarse Vector). In addition, a novel two-level directory architecture is proposed in order to eliminate the penalty caused by these compressed directories. This organization consists of a small Full-Map first-level directory (which provides precise information for the most recently referenced lines) and a compressed second-level directory (which provides in-excess information). Results show that a system with this directory architecture can achieve the same performance as a multiprocessor with a big and non-scalable Full-Map directory with a very significant reduction of the memory overhead.
Manuel E. Acacio, José González 0002, José M. García 0001, José Duato
HPCA3
2001 Selective Branch Prediction Reversal By Correlating with Data Values and Control Flow
abstract
Branch prediction is one of the main hurdles in the roadmap towards deeper pipelines and higher clock frequencies. This work presents a new approach to enhancing current branch predictors: Selective Branch Prediction Reversal. The rationale behind this proposal is the fact that many branch mispredictions can be avoided if branch prediction is selectively reversed. We present a Branch Prediction Reversal Unit (BPRU) that selectively reverses branch predictions by correlating with the predicted values of the branch inputs, in addition to recent control flow. As a case study, we have included the BPRU in an already proposed branch predictor, the Branch Predictor through Value Prediction (BPVP). The effect is a reduction by half in its original misprediction rate. We have also measured the improvement when the BPRU engine is used in a hybrid scheme composed of a BPVP and a gshare predictor. Results using immediate updates show average reductions in misprediction rate ranging from 7% to 14%. Performance evaluation of the proposed BPRU in a 20-stage superscalar processor shows an IPC improvement of up to 9%.
Juan L. Aragón, José González 0002, José M. García 0001, Antonio González 0001
ICCD3
2001 A New Approach to Provide Real-Time Services on High-Speed Local Area Networks
abstract
In the past few years, networks of workstations (NOWs) and clusters, based on high-speed local area networks (LANs), have emerged as a serious alternative to supercomputers and high-performance servers. Meanwhile, applications demanding real-time network services have also suffered a substantial growth. In order to use NOWs for distributed real-time processing, a topology change and faulttolerant mechanism that guarantees the maximum latency or the minimum bandwidth in the worst case must be provided. Up to now, the backup channel protocol (BCP), based on real-time channels, provides fault-tolerant realtime services. But in this approach, fault tolerance is limited by the alternative paths provided by the routing function to establish the backup channels and topology change tolerance is not supported. On the other hand, dynamic reconfiguration updates the routing tables without stopping user traffic when a topology change or fault occurs. However, dynamic reconfiguration by itself does not provide neither quality of service nor real-time services, but it provides support for an additional mechanism designed to meet realtime requirements.
Joaquin Fernández, José M. García 0001, José Duato
IPDPS2
2000 A Parallel Algorithm for Tracking of Segments in Noisy Edge Images
abstract
We present a parallel implementation of a probabilistic algorithm for real time tracking of segments in noisy edge images. Given an initial solution-a set of segments that reasonably describe the input binary edge image-,the algorithm efficiently updates the parameters of these segments to track the movements of objects in the image in successive image frames. The proposed method is based on the EM algorithm-a technique for parameter estimation of statistical distributions in presence of incomplete data-,used here to estimate the parameters of a mixture density. The algorithm is highly susceptible of parallelization, because of the uncoupled nature of the computations needed on its main data structures. This property is exploited in order to make an efficient version for parallel distributed memory environments, under the message passing paradigm. We carefully describe the details of the implementation, and finally, we show an evaluation of the algorithm in a NOW (network of workstations), using the standard message passing interface (MPI) library. Our evaluation shows that the reached speedup is very close to the ideal optimum.
Pedro E. López-de-Teruel, Alberto Ruiz, José M. García 0001
ICPR3
2000 Dynamic reconfiguration of node location in wormhole networks
José L. Sánchez 0002, José M. García 0001
J. Syst. Archit.2