EDBT 2026 Demo / reviewers in the wild / expert
Alex Ramírez
dblp:39/2928 · also Alex Ramírez Bellido
· DBLP profile ↗
68ranked-venue papers
10as first author
1since 2021 · last 2021
0000-0003-3317-4616ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 54 · 8 first-author · 1 since 2021Software engineering, systems software and programming languages · 14 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
23 papers |
Processor architecture and microarchitecture · 26% GPUs and heterogeneous computing · 14% Performance modeling and evaluation · 12% | |
| Software engineering, system software, and programming languages
10 papers |
Compilers and program optimization · 67% Operating systems · 33% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 30 heaviest of 63, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › datacenter architecture
datacenter acceleration |
0.5 | 1 | 2021 | Warehouse-scale video acceleration: co-design and deployment in the wild · ASPLOS 2021 |
GPUs and heterogeneous computing › GPU cache
GPU cache architecture |
0.3 | 1 | 2017 | Beyond the socket: NUMA-aware GPUs · MICRO 2017 |
Interconnection networks and networks-on-chip › interconnect architecture
GPU interconnect |
0.3 | 1 | 2017 | Beyond the socket: NUMA-aware GPUs · MICRO 2017 |
Processor architecture and microarchitecture
instruction fetch |
0.3 | 6 | 2009 | DIA: A Complexity-Effective Decoding Architecture · IEEE Trans. Computers 2009 Enlarging Instruction Streams · IEEE Trans. Computers 2007 Fetching instruction streams · MICRO 2002 |
Energy-efficient computing
energy-efficient system design |
0.2 | 1 | 2016 | The mont-blanc prototype: an alternative approach for HPC systems · SC 2016 |
GPUs and heterogeneous computing
GPU scheduling |
0.2 | 1 | 2014 | Enabling preemptive multiprogramming on GPUs · ISCA 2014 |
GPUs and heterogeneous computing
GPU sharing |
0.2 | 1 | 2014 | Enabling preemptive multiprogramming on GPUs · ISCA 2014 |
Processor architecture and microarchitecture › microprocessor
commodity processors |
0.2 | 1 | 2013 | Supercomputing with commodity CPUs: are mobile SoCs ready for HPC? · SC 2013 |
Processor architecture and microarchitecture › multithreading
simultaneous multithreading |
0.2 | 3 | 2006 | Predictable Performance in SMT Processors: Synergy between the OS and SMTs · IEEE Trans. Computers 2006 Dynamically Controlled Resource Allocation in SMT Processors · MICRO 2004 A Low-Complexity, High-Performance Fetch Unit for Simultaneous Multithreading Processors · HPCA 2004 |
Memory systems
on-chip memory |
0.2 | 2 | 2012 | DMA++: on the fly data realignment for on-chip memories · HPCA 2010 DMA++: On the Fly Data Realignment for On-Chip Memories · IEEE Trans. Computers 2012 |
Bioinformatics and computational biology › gene regulation
regulatory region identification |
0.1 | 1 | 2012 | ReLA, a local alignment search tool for the identification of distal and proximal gene regulatory regions and their conserved transcription factor binding sites · Bioinform. 2012 |
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction |
0.1 | 1 | 2012 | ReLA, a local alignment search tool for the identification of distal and proximal gene regulatory regions and their conserved transcription factor binding sites · Bioinform. 2012 |
Performance modeling and evaluation › simulation
architectural simulation |
0.1 | 1 | 2012 | On the simulation of large-scale architectures using multiple application abstraction levels · ACM Trans. Archit. Code Optim. 2012 |
Performance modeling and evaluation › statistical analysis
extreme value theory |
0.1 | 1 | 2012 | Kernel Partitioning of Streaming Applications: A Statistical Approach to an NP-complete Problem · MICRO 2012 |
Reconfigurable computing and FPGAs
kernel partitioning |
0.1 | 1 | 2012 | Kernel Partitioning of Streaming Applications: A Statistical Approach to an NP-complete Problem · MICRO 2012 |
Performance modeling and evaluation › simulation › processor simulation
multicore simulation |
0.1 | 1 | 2012 | On the simulation of large-scale architectures using multiple application abstraction levels · ACM Trans. Archit. Code Optim. 2012 |
Performance modeling and evaluation
performance prediction |
0.1 | 1 | 2012 | Kernel Partitioning of Streaming Applications: A Statistical Approach to an NP-complete Problem · MICRO 2012 |
Parallel and multicore computing › parallel programming models
stream programming |
0.1 | 1 | 2012 | Kernel Partitioning of Streaming Applications: A Statistical Approach to an NP-complete Problem · MICRO 2012 |
Compilers and program optimization
code layout optimization |
0.1 | 5 | 2005 | Software Trace Cache · IEEE Trans. Computers 2005 Instruction fetch architectures and code layout optimizations · Proc. IEEE 2001 Code layout optimizations for transaction processing workloads · ISCA 2001 |
Processor architecture and microarchitecture › instruction fetch
fetch architecture |
0.1 | 3 | 2004 | A low-complexity fetch architecture for high-performance superscalar processors · ACM Trans. Archit. Code Optim. 2004 A Low-Complexity, High-Performance Fetch Unit for Simultaneous Multithreading Processors · HPCA 2004 Fetching instruction streams · MICRO 2002 |
Hardware accelerators and domain-specific architectures › accelerator architecture
accelerator local memory |
0.1 | 1 | 2010 | DMA++: on the fly data realignment for on-chip memories · HPCA 2010 |
Embedded and real-time systems
DMA transfer |
0.1 | 1 | 2010 | DMA++: on the fly data realignment for on-chip memories · HPCA 2010 |
Processor architecture and microarchitecture
out-of-order execution |
0.1 | 1 | 2010 | Task Superscalar: An Out-of-Order Task Pipeline · MICRO 2010 |
Parallel and multicore computing › parallel programming models › task parallelism
out-of-order task execution |
0.1 | 1 | 2010 | Task Superscalar: An Out-of-Order Task Pipeline · MICRO 2010 |
Parallel and multicore computing › parallelization strategies
task-level parallelism |
0.1 | 1 | 2010 | Task Superscalar: An Out-of-Order Task Pipeline · MICRO 2010 |
Performance modeling and evaluation
workload characterization |
0.1 | 2 | 2018 | vbench: Benchmarking Video Transcoding in the Cloud · ASPLOS 2018 Code layout optimizations for transaction processing workloads · ISCA 2001 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2018 | vbench: Benchmarking Video Transcoding in the Cloud · ASPLOS 2018 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2018 | vbench: Benchmarking Video Transcoding in the Cloud · ASPLOS 2018 |
Processor architecture and microarchitecture
branch prediction |
0.1 | 3 | 2009 | Prophet/Critic Hybrid Branch Prediction · ISCA 2004 DIA: A Complexity-Effective Decoding Architecture · IEEE Trans. Computers 2009 Enlarging Instruction Streams · IEEE Trans. Computers 2007 |
Processor architecture and microarchitecture › instruction fetch
decoded instruction cache |
0.1 | 1 | 2009 | DIA: A Complexity-Effective Decoding Architecture · IEEE Trans. Computers 2009 |
Methods — techniques the papers use, named apart from their topics
hardware-software co-design · 0.5performance evaluation · 0.4microarchitectural profiling · 0.3statistical sampling · 0.3smith-waterman algorithm · 0.3local alignment search · 0.3extreme value theory · 0.3dynamic interconnect policy · 0.3cache policy · 0.3scalability analysis · 0.2benchmarking · 0.2application abstraction level · 0.1hardware realignment unit · 0.1trace cache · 0.1instruction fetch policy · 0.1basic block chaining · 0.1simulation · 0.0branch prediction · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Warehouse-scale video acceleration: co-design and deployment in the wildabstractVideo sharing (e.g., YouTube, Vimeo, Facebook, TikTok) accounts for the majority of internet traffic, and video processing is also foundational to several other key workloads (video conferencing, virtual/augmented reality, cloud gaming, video in Internet-of-Things devices, etc.). The importance of these workloads motivates larger video processing infrastructures and – with the slowing of Moore’s law – specialized hardware accelerators to deliver more computing at higher efficiencies. This paper describes the design and deployment, at scale, of a new accelerator targeted at warehouse-scale video transcoding. We present our hardware design including a new accelerator building block – the video coding unit (VCU) – and discuss key design trade-offs for balanced systems at data center scale and co-designing accelerators with large-scale distributed software systems. We evaluate these accelerators “in the wild" serving live data center jobs, demonstrating 20-33x improved efficiency over our prior well-tuned non-accelerated baseline. Our design also enables effective adaptation to changing bottlenecks and improved failure management, and new workload capabilities not otherwise possible with prior systems. To the best of our knowledge, this is the first work to discuss video acceleration at scale in large warehouse-scale environments. Parthasarathy Ranganathan, Daniel Stodolsky, Jeff Calow, Jeremy Dorfman, Marisabel Guevara, Clinton Wills Smullen IV, Aki Kuusela, Raghu Balasubramanian, Sandeep Bhatia, Prakash Chauhan, Anna Cheung, In Suk Chong, Niranjani Dasharathi, Brian Fosco, Samuel Foss, Ben Gelb, Sara J. Gwin, Yoshiaki Hase, Dake He, Richard Ho 0001, Roy W. Huffman Jr., Elisha Indupalli, Indira Jayaram, Poonacha Kongetira, Cho Mon Kyaw, Aaron Laursen, Fong Lou, Kyle Lucke, J. P. Maaninen, Ramon Macias, Maire Mahony, David Alexander Munday, Srikanth Muroor, Narayana Penukonda, Eric Perkins-Argueta, Devin Persaud, Alex Ramírez, Ville-Mikko Rautio, Yolanda Ripley, Amir Salek, Sathish Sekar, Sergey N. Sokolov, Rob Springer, Don Stark 0002, Mercedes Tan, Mark Wachsler, Andrew C. Walton, David A. Wickeraad, Alvin Wijaya, Hon Kwan Wu |
ASPLOS | 39 |
| 2018 | vbench: Benchmarking Video Transcoding in the CloudabstractThis paper presents vbench, a publicly available benchmark for cloud video services. We are the first study, to the best of our knowledge, to characterize the emerging video-as-a-service workload. Unlike prior video processing benchmarks, vbench's videos are algorithmically selected to represent a large commercial corpus of millions of videos. Reflecting the complex infrastructure that processes and hosts these videos, vbench includes carefully constructed metrics and baselines. The combination of validated corpus, baselines, and metrics reveal nuanced tradeoffs between speed, quality, and compression. We demonstrate the importance of video selection with a microarchitectural study of cache, branch, and SIMD behavior. vbench reveals trends from the commercial corpus that are not visible in other video corpuses. Our experiments with GPUs under vbench's scoring scenarios reveal that context is critical: GPUs are well suited for live-streaming, while for video-on-demand shift costs from compute to storage and network. Counterintuitively, they are not viable for popular videos, for which highly compressed, high quality copies are required. We instead find that popular videos are currently well-served by the current trajectory of software encoders. Andrea Lottarini, Alex Ramírez, Joel Coburn, Martha A. Kim, Parthasarathy Ranganathan, Daniel Stodolsky, Mark Wachsler |
ASPLOS | 2 |
| 2017 | Sharing the instruction cache among lean cores on an asymmetric CMP for HPC applicationsabstractHigh performance computing (HPC) applications have parallel code sections that must scale to large numbers of cores, which makes them sensitive to serial regions. Current supercomputing systems with heterogeneous or asymmetric CMPs (ACMP) combine few high-performance big cores for serial regions, together with many low-power lean cores for throughput computing. The low requirements of HPC applications in the core front-end lead some designs, such as SMT and GPU cores, to share front-end structures including the instruction cache (I-cache). However, little work exists to analyze the benefit of sharing the I-cache among full cores, which seems compelling as a solution to reduce silicon area and power. This paper analyzes the performance, power and area impact of such a design on an ACMP with one high-performance core and multiple low-power cores. Having identified that multiple cores run the same code during parallel regions, the lean cores share the I-cache with the intent of benefiting from mutual prefetching, without increasing the average access latency. Our exploration of the multiple parameters finds the sweet spot on a wide interconnect to access the shared I-cache and the inclusion of a few line buffers to provide the required bandwidth and latency to sustain performance. The projections with McPAT and a rich set of HPC benchmarks show 11% area savings with a 5% energy reduction at no performance cost. Ugljesa Milic, Alejandro Rico, Paul M. Carpenter, Alex Ramírez |
ISPASS | 4 |
| 2017 | Beyond the socket: NUMA-aware GPUsabstractGPUs achieve high throughput and power efficiency by employing many small single instruction multiple thread (SIMT) cores. To minimize scheduling logic and performance variance they utilize a uniform memory system and leverage strong data parallelism exposed via the programming model. With Moore's law slowing, for GPUs to continue scaling performance (which largely depends on SIMT core count) they are likely to embrace multi-socket designs where transistors are more readily available. However when moving to such designs, maintaining the illusion of a uniform memory system is increasingly difficult. In this work we investigate multi-socket non-uniform memory access (NUMA) GPU designs and show that significant changes are needed to both the GPU interconnect and cache architectures to achieve performance scalability. We show that application phase effects can be exploited allowing GPU sockets to dynamically optimize their individual interconnect and cache policies, minimizing the impact of NUMA effects. Our NUMA-aware GPU outperforms a single GPU by 1.5×, 2.3×, and 3.2× while achieving 89%, 84%, and 76% of theoretical application scalability in 2, 4, and 8 sockets designs respectively. Implementable today, NUMA-aware multi-socket GPUs may be a promising candidate for scaling GPU performance beyond a single socket. Ugljesa Milic, Oreste Villa, Evgeny Bolotin, Akhil Arunkumar, Eiman Ebrahimi, Aamer Jaleel, Alex Ramírez, David W. Nellans |
MICRO | 7 |
| 2016 | The mont-blanc prototype: an alternative approach for HPC systemsabstractHigh-performance computing (HPC) is recognized as one of the pillars for further progress in science, industry, medicine, and education. Current HPC systems are being developed to overcome emerging architectural challenges in order to reach Exascale level of performance, projected for the year 2020. The much larger embedded and mobile market allows for rapid development of intellectual property (IP) blocks and provides more flexibility in designing an application-specific system-on-chip (SoC), in turn providing the possibility in balancing performance, energy-efficiency, and cost. In the Mont-Blanc project, we advocate for HPC systems being built from such commodity IP blocks, currently used in embedded and mobile SoCs. As a first demonstrator of such an approach, we present the Mont-Blanc prototype; the first HPC system built with commodity SoCs, memories, and network interface cards (NICs) from the embedded and mobile domain, and off-the-shelf HPC networking, storage, cooling, and integration solutions. We present the system's architecture and evaluate both performance and energy efficiency. Further, we compare the system's abilities against a production level supercomputer. At the end, we discuss parallel scalability and estimate the maximum scalability point of this approach across a set of applications. Nikola Rajovic, Alejandro Rico, Filippo Mantovani, Daniel Ruiz 0003, Josep Oriol Vilarrubi, Constantino Gómez, Luna Backes, Diego Nieto, Harald Servat, Xavier Martorell, Jesús Labarta, Eduard Ayguadé, Chris Adeniyi-Jones, Said Derradji, Hervé Gloaguen, Piero Lanucara, Nico Sanna, Jean-François Méhaut, Kevin Pouget, Brice Videau, Eric Boyer, Momme Allalen, Axel Auweter, David Brayford, Daniele Tafani, Volker Weinberg, Dirk Brömmel, René Halver, Jan H. Meinke, Ramón Beivide, Mariano Benito, Enrique Vallejo 0001, Mateo Valero, Alex Ramírez |
SC | 34 |
| 2015 | Exploring multiple sleep modes in on/off based energy efficient HPC networksabstractEnergy efficiency is one of the key challenges in high-performance computing (HPC). The current target of 1 ExaFlop in 20 MW requires a ten-fold improvement in energy efficiency, which is only possible through significant improvements in the energy efficiency throughout the system. Interconnects are particularly inefficient, since their links are always on, consuming full power in order to provide low latency, even though the average interconnect utilization is low. To address the above, the Ethernet standards committee in-charge of 40/100/400Gb Ethernet has opted to include protocols that define low power modes, specifically Fast-Wake, alongside the older Deep-Sleep, to make interconnect links energy proportional. With these standards ratified as recently as March 2014, it is unclear how these low power modes can be used in HPC. While energy efficiency is critical, techniques with excessive performance overheads are unlikely to be adopted in HPC. To this end, this paper performs the first detailed analysis of Fast-Wake mode for link energy savings in the context of HPC. Our results show that a combination of Fast-Wake and Deep-Sleep can reduce link energy savings by up to 70% with less than 1% performance overheads. However, we show how the parameters of these low power modes must be carefully configured to obtain the right trade-offs in energy and performance. We believe that our analysis could benefit interconnect vendors looking to use these low power modes for deployment in HPC. Karthikeyan P. Saravanan, Paul M. Carpenter, Alex Ramírez |
ICCD | 3 |
| 2015 | Limpio: LIghtweight MPI instrumentatiOnabstractCharacterization of high-performance computing applications often has to be done without access to the source code. Computer architects, therefore, have a narrowed choice of instrumentation tools. Moreover, potentially large amount of collected data can prohibit creating a full time stamped event trace and analyzing it post-mortem. This paper describes Limpio -- a Light weight MPI instrumentation framework, that allows dynamic instrumentation of user-selected MPI calls, and customization of data gathering, analysis and visualization. Milan Pavlovic, Milan Radulovic, Alex Ramírez, Petar Radojkovic |
ICPC | 3 |
| 2014 | A performance perspective on energy efficient HPC linksabstractEnergy costs are an increasing part of the total cost of ownership of HPC systems. As HPC systems become increasingly energy proportional in an effort to reduce energy costs, interconnect links stand out for their inefficiency. Commodity interconnect links remain 'always-on', consuming full power even when no data is being transmitted. Although various techniques have been proposed towards energy-proportional interconnects, they are often too conservative or are not focused toward HPC. Aggressive techniques for interconnect energy savings are often not applied to HPC, in particular, because they may incur excessive performance overheads. Any energy-saving technique will only be adopted in HPC if there is no significant impact on performance, which is still the primary design objective. Karthikeyan P. Saravanan, Paul M. Carpenter, Alex Ramírez |
ICS | 3 |
| 2014 | Energy Efficient HPC on Embedded SoCs: Optimization Techniques for Mali GPUabstractA lot of effort from academia and industry has been invested in exploring the suitability of low-power embedded technologies for HPC. Although state-of-the-art embedded systems-on-chip (SoCs) inherently contain GPUs that could be used for HPC, their performance and energy capabilities have never been evaluated. Two reasons contribute to the above. Primarily, embedded GPUs until now, have not supported 64-bit floating point arithmetic - a requirement for HPC. Secondly, embedded GPUs did not provide support for parallel programming languages such as OpenCL and CUDA. However, the situation is changing, and the latest GPUs integrated in embedded SoCs do support 64-bit floating point precision and parallel programming models. In this paper, we analyze performance and energy advantages of embedded GPUs for HPC. In particular, we analyze ARM Mali-T604 GPU - the first embedded GPUs with OpenCL Full Profile support. We identify, implement and evaluate software optimization techniques for efficient utilization of the ARM Mali GPU Compute Architecture. Our results show that, HPC benchmarks running on the ARM Mali-T604 GPU integrated into Exynos 5250 SoC, on average, achieve speed-up of 8.7X over a single Cortex-A15 core, while consuming only 32% of the energy. Overall results show that embedded GPUs have performance and energy qualities that make them candidates for future HPC systems. Ivan Grasso, Petar Radojkovic, Nikola Rajovic, Isaac Gelado, Alex Ramírez |
IPDPS | 5 |
| 2014 | Enabling preemptive multiprogramming on GPUsabstractGPUs are being increasingly adopted as compute accelerators in many domains, spanning environments from mobile systems to cloud computing. These systems are usually running multiple applications, from one or several users. However GPUs do not provide the support for resource sharing traditionally expected in these scenarios. Thus, such systems are unable to provide key multiprogrammed workload requirements, such as responsiveness, fairness or quality of service. In this paper, we propose a set of hardware extensions that allow GPUs to efficiently support multiprogrammed GPU workloads. We argue for preemptive multitasking and design two preemption mechanisms that can be used to implement GPU scheduling policies. We extend the architecture to allow concurrent execution of GPU kernels from different user processes and implement a scheduling policy that dynamically distributes the GPU cores among concurrently running kernels, according to their priorities. We extend the NVIDIA GK110 (Kepler) like GPU architecture with our proposals and evaluate them on a set of multiprogrammed workloads with up to eight concurrent processes. Our proposals improve execution time of high-priority processes by 15.6x, the average application turnaround time between 1.5x to 2x, and system fairness up to 3.4x. Ivan Tanasic, Isaac Gelado, Javier Cabezas, Alex Ramírez, Nacho Navarro, Mateo Valero |
ISCA | 4 |
| 2014 | Tibidabo: Making the case for an ARM-based HPC system
Nikola Rajovic, Alejandro Rico, Nikola Puzovic, Chris Adeniyi-Jones, Alex Ramírez |
Future Gener. Comput. Syst. | 5 |
| 2013 | Experiences with mobile processors for energy efficient HPCabstractThe performance of High Performance Computing (HPC) systems is already limited by their power consumption. The majority of top HPC systems today are built from commodity server components that were designed for maximizing the compute performance. The Mont-Blanc project aims at using low-power parts from the mobile domain for HPC. In this paper we present our first experiences with the use of mobile processors and accelerators for the HPC domain based on the research that was performed in the project. We show initial evaluation of NVIDIA Tegra 2 and Tegra 3 mobile SoCs and the NVIDIA Quadro 1000M GPU with a set of HPC micro-benchmarks to evaluate their potential for energy-efficient HPC. Nikola Rajovic, Alejandro Rico, James Vipond, Isaac Gelado, Nikola Puzovic, Alex Ramírez |
DATE | 6 |
| 2013 | Data placement in HPC architectures with heterogeneous off-chip memoryabstractThe performance of HPC applications is often bounded by the underlying memory system's performance. The trend of increasing the number of cores on a chip imposes even higher memory bandwidth and capacity requirements. The limitations of traditional memory technologies are pushing research in the direction of hybrid memory systems that, besides DRAM, include one or more modules based on some of the higher-density non-volatile memory technologies, where one of them will provide the required bandwidth, while the other will provide the required capacity for the application. This creates many challenges with data placement and migration policies between the modules of such hybrid memory system. In this paper, we propose an architecture with a hybrid memory design that places two technologically different memory modules in a flat address space. On such system, we evaluate several HPC workloads against different data placement and migration policies, compare their performance by means of execution time and the number of non-volatile memory writes, and consider how it can be applied to the future HPC architectures. Our results show that the hybrid memory system with dynamic page migration and limited DRAM capacity, can achieve performance that is comparable to a hypothetical, hard to implement, DRAM-only system. Milan Pavlovic, Nikola Puzovic, Alex Ramírez |
ICCD | 3 |
| 2013 | Programmable and Scalable Reductions on ClustersabstractReductions matter and they are here to stay. Wide adoption of parallel processing hardware in a broad range of computer applications has encouraged recent research efforts on their efficient parallelization. Furthermore, trends towards high productivity languages in mainstream computing increases the demand for efficient programming support. In this paper we present a new approach on parallel reductions for distributed memory systems that provides both scalability and programmability. Using OmpSs, a task-based parallel programming model, the developer has the ability to express scalable reductions through a single pragma annotation. This pragma annotation is applicable for tasks as well as for work-sharing constructs (with implicit tasking) and instructs the compiler to generate the required runtime calls. The supporting runtime handles data and task distribution, parallel execution and data reduction. Scalability is achieved through a software cache that maximizes local and temporal data reuse and allows overlapped computation and communication. Results confirm scalability for up to 32 12-core cluster nodes. Jan Ciesko, Javier Bueno, Nikola Puzovic, Alex Ramírez, Rosa M. Badia, Jesús Labarta |
IPDPS | 4 |
| 2013 | Trace filtering of multithreaded applications for CMP memory simulationabstractRecent works have shown that modelling the performance of out-of-order superscalar cores is doable using filtered memory traces for single thread simulations. However, those techniques do not account for cache coherence actions so they cannot be used reliably in multithreaded scenarios. In this paper, we leverage the structure of parallel applications to propose a simulation methodology that enables the use of filtered memory traces for the simulation of multithreaded applications on multicore architectures. In our experiments our proposal reduced the simulation error of state-of-the-art techniques in 39% on average, while only losing 9.5% of simulation speedup. Alejandro Rico, Alex Ramírez, Mateo Valero |
ISPASS | 2 |
| 2013 | Power/performance evaluation of energy efficient Ethernet (EEE) for High Performance ComputingabstractAs of June 2012, 41 % of all systems in the TOP5OO use Gigabit Ethernet. Ethernet has been a strong contender in the HPC interconnect market for its competitive performance and low cost. However, until recently, little emphasis has been thrown upon bringing about energy efficient HPC interconnects. To illustrate, in a majority if not all Ethernet based systems, the transmitter and receiver operate at full power regardless of any data transmission between them, leading to power inefficiency. The recent standard IEEE 802.3az, Energy Efficient Ethernet (EEE), approved in 2010, solves the above conundrum by introducing "Low-Power-Idle", dynamically turning off unused links to save interconnect power. In this paper, we present the first analysis of Energy Efficient Ethernet in the domain of HPC, examining its potential for power savings. Unlike previous proposals, we present a detailed analysis of the impact of additional latency overhead introduced by EEE, using multiple simulated systems running actual HPC application traces. We propose the use of "Power-Down Threshold", as a possible add-on to EEE to mitigate its on/off transition overhead. We find that EEE brings about link power savings of about 70% by switching off links, but at the cost of performance, leading to increased power consumption of the overall system by l5% (average). In contrast, using our proposed "Power-Down Threshold", we demonstrate reduced on/off transition overhead, from 25% to 2%, translating to overall system power savings of about 7.5%. Furthermore, in this work we point out relevant design decisions for future vendors intending to deploy EEE solutions for their HPC systems. Karthikeyan P. Saravanan, Paul M. Carpenter, Alex Ramírez |
ISPASS | 3 |
| 2013 | Supercomputing with commodity CPUs: are mobile SoCs ready for HPC?abstractIn the late 1990s, powerful economic forces led to the adoption of commodity desktop processors in high-performance computing. This transformation has been so effective that the June 2013 TOP500 list is still dominated by x86. Nikola Rajovic, Paul M. Carpenter, Isaac Gelado, Nikola Puzovic, Alex Ramírez, Mateo Valero |
SC | 5 |
| 2012 | Topic 16: GPU and Accelerators Computing
Alex Ramírez, Dimitrios S. Nikolopoulos, David R. Kaeli, Satoshi Matsuoka |
Euro-Par | 1 |
| 2012 | Kernel Partitioning of Streaming Applications: A Statistical Approach to an NP-complete ProblemabstractOne of the greatest challenges in computer architecture is how to write efficient, portable, and correct software for multi-core processors. A promising approach is to expose more parallelism to the compiler, through the use of domain-specific languages. The compiler can then perform complex transformations that the programmer would otherwise have had to do. Many important applications related to audio and video encoding, software radio and signal processing have regular behavior that can be represented using a stream programming language. When written in such a language, a portable stream program can be automatically mapped by the stream compiler onto multicore hardware. One of the most difficult tasks of the stream compiler is partitioning the stream program into software threads. The choice of partition significantly affects performance, but finding the optimal partition is an NP-complete problem. This paper presents a method, based on Extreme Value theory (EVT), that statistically estimates the performance of the optimal partition. Knowing the optimal performance improves the evaluation of any partitioning algorithm, and it is the most important piece of information when deciding whether an existing algorithm should be enhanced. We use the method to evaluate a recently-published partitioning algorithm based on a heuristic. We further analyze how the statistical method is affected by the choice of sampling method, and we recommend how sampling should be done. Finally, since a heuristic-based algorithm may not always be available, the user may try to find a good partition by picking the best from a random sample. We analyze whether this approach is likely to find a good partition. To the best of our knowledge, this study is the first application of EVT to a graph partitioning problem. Petar Radojkovic, Paul M. Carpenter, Miquel Moretó, Alex Ramírez, Francisco J. Cazorla |
MICRO | 4 |
| 2012 | ReLA, a local alignment search tool for the identification of distal and proximal gene regulatory regions and their conserved transcription factor binding sitesabstractAbstract Motivation: The prediction and annotation of the genomic regions involved in gene expression has been largely explored. Most of the energy has been devoted to the development of approaches that detect transcription start sites, leaving the identification of regulatory regions and their functional transcription factor binding sites (TFBSs) largely unexplored and with important quantitative and qualitative methodological gaps. Results: We have developed ReLA (for REgulatory region Local Alignment tool), a unique tool optimized with the Smith–Waterman algorithm that allows local searches of conserved TFBS clusters and the detection of regulatory regions proximal to genes and enhancer regions. ReLA's performance shows specificities of 81 and 50% when tested on experimentally validated proximal regulatory regions and enhancers, respectively. Availability: The source code of ReLA's is freely available and can be remotely used through our web server under http://www.bsc.es/cg/rela. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Santi González, Bàrbara Montserrat-Sentís, Friman Sánchez, Montserrat Puiggròs, Enrique Blanco, Alex Ramírez, David Torrents |
Bioinform. | 6 |
| 2012 | On the simulation of large-scale architectures using multiple application abstraction levelsabstractSimulation is a key tool for computer architecture research. In particular, cycle-accurate simulators are extremely important for microarchitecture exploration and detailed design decisions, but they are slow and, so, not suitable for simulating large-scale architectures, nor are they meant for this. Moreover, microarchitecture design decisions are irrelevant, or even misleading, for early processor design stages and high-level explorations. This allows one to raise the abstraction level of the simulated architecture, and also the application abstraction level, as it does not necessarily have to be represented as an instruction stream. In this paper we introduce a definition of different application abstraction levels, and how these are employed in TaskSim, a multi-core architecture simulator, to provide several architecture modeling abstractions, and simulate large-scale architectures with hundreds of cores. We compare the simulation speed of these abstraction levels to the ones in existing simulation tools, and also evaluate their utility and accuracy. Our simulations show that a very high-level abstraction, which may be even faster than native execution, is useful for scalability studies on parallel applications; and that just simulating explicit memory transfers, we achieve accurate simulations for architectures using non-coherent scratchpad memories, with just a 25x slowdown compared to native execution. Furthermore, we revisit trace memory simulation techniques, that are more abstract than instruction-by-instruction simulations and provide an 18x simulation speedup. Alejandro Rico, Felipe Cabarcas, Carlos Villavieja, Milan Pavlovic, Augusto Vega, Yoav Etsion, Alex Ramírez, Mateo Valero |
ACM Trans. Archit. Code Optim. | 7 |
| 2012 | DMA++: On the Fly Data Realignment for On-Chip MemoriesabstractMultimedia extensions based on Single-Instruction Multiple-Data (SIMD) units are widespread. They have been used, for some time, in processors and accelerators (e.g., the Cell SPEs). SIMD units usually have significant memory alignment constraints in order to meet power requirements and design simplicity. This increases the complexity of the code generated by the compiler as, in the general case, the compiler cannot be sure of the proper alignment of data. For that, the ISA provides either unaligned memory load and store instructions, or a special set of instructions to perform realignments in software. In this paper, we propose a hardware realignment unit that takes advantage of the DMA transfers needed in accelerators with local memories. While the data are being transferred, it is realigned on the fly by our realignment unit, and stored at the desired alignment in the accelerator memory. This mechanism can help programmers to better organize data in the accelerator memory so that the accelerator can possibly access the data with no special instructions. Finally, the data are realigned properly also when put back to main memory. Our experiments with nine applications show that with our approach, the bandwidth of the DMA transfers is not penalized. Nikola Vujic, Felipe Cabarcas, Marc González 0001, Alex Ramírez, Xavier Martorell, Eduard Ayguadé |
IEEE Trans. Computers | 4 |
| 2011 | DiDi: Mitigating the Performance Impact of TLB Shootdowns Using a Shared TLB DirectoryabstractTranslation Look aside Buffers (TLBs) are ubiquitously used in modern architectures to cache virtual-to-physical mappings and, as they are looked up on every memory access, are paramount to performance scalability. The emergence of chip-multiprocessors (CMPs) with per-core TLBs, has brought the problem of TLB coherence to front stage. TLBs are kept coherent at the software-level by the operating system (OS). Whenever the OS modifies page permissions in a page table, it must initiate a coherency transaction among TLBs, a process known as a TLB shoot down. Current CMPs rely on the OS to approximate the set of TLBs caching a mapping and synchronize TLBs using costly Inter-Proceessor Interrupts (IPIs) and software handlers. In this paper, we characterize the impact of TLB shoot downs on multiprocessor performance and scalability, and present the design of a scalable TLB coherency mechanism. First, we show that both TLB shoot down cost and frequency increase with the number of processors and project that software-based TLB shoot downs would thwart the performance of large multiprocessors. We then present a scalable architectural mechanism that couples a shared TLB directory with load/store queue support for lightweight TLB invalidation, and thereby eliminates the need for costly IPIs. Finally, we show that the proposed mechanism reduces the fraction of machine cycles wasted on TLB shoot downs by an order of magnitude. Carlos Villavieja, Vasileios Karakostas, Lluís Vilanova, Yoav Etsion, Alex Ramírez, Avi Mendelson, Nacho Navarro, Adrián Cristal, Osman S. Unsal |
PACT | 5 |
| 2011 | Scaling HMMER Performance on Multicore ArchitecturesabstractIn bioinformatics, protein sequence alignment is one of the fundamental tasks that scientists perform. Since the growth of biological data is exponential, there is an ever-increasing demand for computational power. While current processor technology is shifting towards the use of multicores, the mapping and parallelization of applications has become a critical issue. In order to keep up with the processing demands, applications' bottlenecks to performance need to be found and properly addressed. In this paper we study the parallelism and performance scalability of HMMER, a bioinformatics application to perform sequence alignment. After our study of the bottlenecks in a HMMER version ported to the Cell processor, we present two optimized versions to improve scalability in a larger multicore architecture. We use a simulator that allows us to model a system with up to 512 processors and study the performance of the three parallel versions of HMMER. Results show that removing the I/O bottleneck improves performance by 3X and 2.4X for a short and a long HMM query respectively. Additionally, by offloading the sequence pre-formatting to the worker cores, larger speedups of up to 27X and 7X are achieved. Compared to using a single worker processor, up to 156X speedup is obtained when using 256 cores. Sebastián Isaza, Ernst Houtgast, Friman Sánchez, Alex Ramírez, Georgi Gaydadjiev |
CISIS | 4 |
| 2011 | FELI: HW/SW Support for On-Chip Distributed Shared Memory in Multicores
Carlos Villavieja, Yoav Etsion, Alex Ramírez, Nacho Navarro |
Euro-Par (1) | 3 |
| 2011 | Trace-driven simulation of multithreaded applicationsabstractOver the past few years, computer architecture research has moved towards execution-driven simulation, due to the inability of traces to capture timing-dependent thread execution interleaving. However, trace-driven simulation has many advantages over execution-driven that are being missed in multithreaded application simulations. We present a methodology to properly simulate multithreaded applications using trace-driven environments. We distinguish the intrinsic application behavior from the computation for managing parallelism. Application traces capture the intrinsic behavior in the sections of code that are independent from the dynamic multithreaded nature, and the points where parallelism-management computation occurs. The simulation framework is composed of a trace-driven simulation engine and a dynamic-behavior component that implements the parallelism-management operations for the application. Then, at simulation time, these operations are reproduced by invoking their implementation in the dynamic-behavior component. The decisions made by these operations are based on the simulated architecture, allowing to dynamically reschedule sections of code taken from the trace to the target simulated components. As the captured sections of code are independent from the parallel state of the application, they can be simulated on the trace-driven engine, while the parallelism-management operations, that require to be re-executed, are carried out by the execution-driven component, thus achieving the best of both trace- and execution-driven worlds. This simulation methodology creates several new research opportunities, including research on scheduling and other parallelism-management techniques for future architectures, and hardware support for programming models. Alejandro Rico, Alejandro Duran, Felipe Cabarcas, Yoav Etsion, Alex Ramírez, Mateo Valero |
ISPASS | 5 |
| 2011 | Scalable multicore architectures for long DNA sequence comparisonabstractSUMMARY Biological sequence comparison is one of the most important tasks in Bioinformatics. Owing to the fast growth of databases that contain biological information, sequence comparison represents an important challenge for high‐performance computing, especially when very long sequences are compared, i.e. the complete genome of several organisms. The Smith–Waterman (SW) algorithm is an exact method based on dynamic programming to quantify local similarity between sequences. The inherent large parallelism of the algorithm makes it ideal for architectures supporting multiple dimensions of parallelism (TLP, DLP and ILP). Concurrently, there is a paradigm shift towards chip multiprocessors in computer architecture, which offer a huge amount of potential performance that can only be exploited efficiently if applications are effectively mapped and parallelized. In this work, we analyze how large‐scale biology sequence comparison takes advantage of the current and future multicore architectures. Our starting point is the performance analysis of the current multicore IBM Cell B.E. processor; we analyze two different SW implementations on the Cell B.E. Then, using simulation tools, we study the performance scalability when a many‐core architecture is used for performing long DNA sequence comparison. We investigate the efficient memory organization that delivers the maximum bandwidth with the minimum cost. Our results show that a heterogeneous architecture can be an efficient alternative to execute challenging bioinformatic workloads. Copyright © 2011 John Wiley & Sons, Ltd. Friman Sánchez, Felipe Cabarcas, Alex Ramírez, Mateo Valero |
Concurr. Comput. Pract. Exp. | 3 |
| 2010 | Scalability Analysis of Progressive Alignment on a MulticoreabstractSequence alignment is a fundamental instrument in bioinformatics. In recent years, numerous proposals have been addressing the problem of accelerating this class of applications. This, due to the rapid growth of sequence databases in combination with the high computational demands imposed by the algorithms. In this paper we focus on the analysis of the progressive alignment in ClustalW, a widely used program for performing multiple sequence alignment. We have parallelized ClustalW for the Cell processor architecture and have carefully analyzed the scalability of its different phases with both the number of cores used and the input size. Experimental results show that computing profile scores scales well up to 16 SPE cores. With the increase of the input size, profiles initialization in the PPE core becomes the predominant bottleneck. Sebastián Isaza, Friman Sánchez, Georgi Gaydadjiev, Alex Ramírez, Mateo Valero |
CISIS | 4 |
| 2010 | Empowering Business Students - Using Web 2.0 Tools in the Classroom
Alex Ramírez, Shaobo Ji, Robert J. Riordan, Frank Ulbrich, Michael J. Hine |
CSEDU (1) | 1 |
| 2010 | Starsscheck: A Tool to Find Errors in Task-Based Parallel Programs
Paul M. Carpenter, Alex Ramírez, Eduard Ayguadé |
Euro-Par (1) | 2 |
| 2010 | Long DNA Sequence Comparison on Multicore Architectures
Friman Sánchez, Felipe Cabarcas, Alex Ramírez, Mateo Valero |
Euro-Par (2) | 3 |
| 2010 | Buffer Sizing for Self-timed Stream Programs on Heterogeneous Distributed Memory Multiprocessors
Paul M. Carpenter, Alex Ramírez, Eduard Ayguadé |
HiPEAC | 2 |
| 2010 | DMA++: on the fly data realignment for on-chip memoriesabstractMultimedia extensions based on Single-Instruction Multiple-Data (SIMD) units are widespread. They are used both in processors and accelerators (e.g., the Cell SPEs), since some time ago. SIMD units have usually big memory alignment constraints in order to meet power requirements and design simplicity. This increases the complexity of the code generated by the compiler, as in the general case, the compiler cannot be sure of the proper alignment of data. For that, the ISA provides either unaligned memory load and store instructions, or a special set of instructions to perform the realignments in software. In this paper, we propose a hardware realignment unit that takes advantage of the DMA transfers needed in accelerators with local memories. While the data is being transferred, it is realigned on the fly by our realignment unit, and stored with the proper alignment in the accelerator memory. The accelerator can then access the data with no special instructions. Finally, the data is realigned properly also when put back to main memory. Our experiments with four applications show that with our approach, the bandwidth of the DMA transfers is not penalized. And the performance of the synthetic benchmarks shows that aligned code is 1.5 to 2 times better with respect using unaligned code. Nikola Vujic, Marc González 0001, Felipe Cabarcas, Alex Ramírez, Xavier Martorell, Eduard Ayguadé |
HPCA | 4 |
| 2010 | Task Superscalar: An Out-of-Order Task PipelineabstractWe present \emph{Task Super scalar}, an abstraction of instruction-level out-of-order pipeline that operates at the task-level. Like ILP pipelines, which uncover parallelism in a sequential instruction stream, task super scalar uncovers task-level parallelism among tasks generated by a sequential thread. Utilizing intuitive programmer annotations of task inputs and outputs, the task super scalar pipeline dynamically detects inter-task data dependencies, identifies task-level parallelism, and executes tasks out-of-order. Furthermore, we propose a design for a distributed task super scalar pipeline front end, that can be embedded into any many core fabric, and manages cores as functional units. We show that our proposed mechanism is capable of driving hundreds of cores simultaneously with non-speculative tasks, which allows our pipeline to sustain work windows consisting of tens of thousands of tasks. We further show that our pipeline can maintain a decode rate faster than 60ns per task and dynamically uncover data dependencies among as many as ~50,000 in-flight tasks, using 7MB of on-chip eDRAM storage. This configuration achieves speedups of 95-255x (average 183x) over sequential execution for nine scientific benchmarks, running on a simulated CMP with 256 cores. Task super scalar thus enables programmers to exploit many core systems effectively, while simultaneously simplifying their programming model. Yoav Etsion, Felipe Cabarcas, Alejandro Rico, Alex Ramírez, Rosa M. Badia, Eduard Ayguadé, Jesús Labarta, Mateo Valero |
MICRO | 4 |
| 2009 | Mapping stream programs onto heterogeneous multiprocessor systemsabstractThis paper presents a partitioning and allocation algorithm\nfor an iterative stream compiler, targeting heterogeneous\nmultiprocessors with constrained distributed memory and\nany communications topology. We introduce a novel definition\nof connectedness that enables the algorithm to model\nthe capabilities of the compiler. The algorithm uses convexity\nand connectedness constraints to produce partitions that\nare easier to compile and require short pipelines. Software\npipelining is an effective transformation, but it increases\nmemory footprint and latency, and has a startup overhead.\nOur algorithm takes account of these downstream costs.\nWe show results for the StreamIt 2.1.1 benchmarks for\nan SMP, 2 × 2 mesh, SMP plus accelerator, and IBM QS20\nblade, which has two Cell processors. Our results show that\nthe average performance is within 5% of the unrestricted\noptimum found using a brute force search, while seldom requiring\nsoftware pipelining. The heuristic is robust, and fast\nenough to be inside the feedback loop of an iterative compiler. Paul M. Carpenter, Alex Ramírez, Eduard Ayguadé |
CASES | 2 |
| 2009 | Parallel H.264 Decoding on an Embedded Multicore Processor
Arnaldo Azevedo, Cor Meenderinck, Ben H. H. Juurlink, Andrei Sergeevich Terechko, Jan Hoogerbrugge, Mauricio Alvarez-Mesa, Alex Ramírez |
HiPEAC | 7 |
| 2009 | Scalability of Macroblock-level Parallelism for H.264 DecodingabstractThis paper investigates the scalability of MacroBlock (MB) level parallelization of the H.264 decoder for High Definition (HD) applications. The study includes three parts. First, a formal model for predicting the maximum performance that can be obtained taking into account variable processing time of tasks and thread synchronization overhead. Second, an implementation on a real multiprocessor architecture including a comparison of different scheduling strategies and a profiling analysis for identifying the performance bottlenecks. Finally, a trace-driven simulation methodology has been used for identifying the opportunities of acceleration for removing the main bottlenecks. It includes the acceleration potential for the entropy decoding stage and thread synchronization and scheduling. Our study presents a quantitative analysis of the main bottlenecks of the application and estimates the acceleration levels that are required to make the MB-level parallel decoder scalable. Mauricio Alvarez-Mesa, Alex Ramírez, Arnaldo Azevedo, Cor Meenderinck, Ben H. H. Juurlink, Mateo Valero |
ICPADS | 2 |
| 2009 | Thread to Core Assignment in SMT On-Chip MultiprocessorsabstractState-of-the-art high-performance processors like the IBM POWER5 and Intel i7 show a trend in industry towards on-chip Multiprocessors (CMP) involving Simultaneous Multithreading (SMT) in each core. In these processors, the way in which applications are assigned to cores plays a key role in the performance of each application and the overall system performance. In this paper we show that the system throughput highly depends on the Thread to Core Assignment (TCA), regardless the SMT Instruction Fetch (IFetch) Policy implemented in the cores. Our results indicate that a good TCA can improve the results of any underlying IFetch Policy, yielding speedups of up to 28%. Given the relevance of TCA, we propose an algorithm to manage it in CMP+SMT processors. The proposed throughput-oriented TCA Algorithm takes into account the workload characteristics and the underlying SMT IFetch Policy. Our results show that the TCA Algorithm obtains thread-to-core assignments 3% close to the optimal assignation for each case, yielding system throughput improvements up to 21%. Carmelo Acosta, Francisco J. Cazorla, Alex Ramírez, Mateo Valero |
SBAC-PAD | 3 |
| 2009 | DIA: A Complexity-Effective Decoding ArchitectureabstractFast instruction decoding is a true challenge for the design of CISC microprocessors implementing variable-length instructions. A well-known solution to overcome this problem is caching decoded instructions in a hardware buffer. Fetching already decoded instructions avoids the need for decoding them again, improving processor performance. However, introducing such special--purpose storage in the processor design involves an important increase in the fetch architecture complexity. In this paper, we propose a novel decoding architecture that reduces the fetch engine implementation cost. Instead of using a special-purpose hardware buffer, our proposal stores frequently decoded instructions in the memory hierarchy. The address where the decoded instructions are stored is kept in the branch prediction mechanism, enabling it to guide our decoding architecture. This makes it possible for the processor front end to fetch already decoded instructions from the memory instead of the original nondecoded instructions. Our results show that using our decoding architecture, a state-of-the-art superscalar processor achieves competitive performance improvements, while requiring less chip area and energy consumption in the fetch architecture than a hardware code caching mechanism. Oliverio J. Santana, Ayose Falcón, Alex Ramírez, Mateo Valero |
IEEE Trans. Computers | 3 |
| 2008 | MLP-Aware Dynamic Cache Partitioning
Miquel Moretó, Francisco J. Cazorla, Alex Ramírez, Mateo Valero |
HiPEAC | 3 |
| 2008 | MFLUSH: Handling Long-Latency Loads in SMT On-Chip MultiprocessorsabstractNowadays, there is a clear trend in industry towards employing the growing amount of transistors on chip in replicating execution cores (CMP), where each core is simultaneous multithreading (SMT). State-of-the-art high-performance processors like the IBM POWER5 and POWER6 corroborate this CMP+SMT trend. Within each SMT core any of the well-known SMT mechanisms may be applied to face SMT related challenges. Among them, probably the most important issue in an SMT execution pipeline concerns the instruction fetch (IFetch) Policy. The FLUSH IFetch Policy represents a choice for throughput-oriented scenarios. It handles L2 cache misses in order to avoid hardware resource monopolization by any given execution thread; involving an additional energy cost via instruction refetching. However, the new constraints imposed by the CMP+SMT scenario may affect well-known SMT mechanisms, like the FLUSH mechanism. In this paper we revisit the FLUSH mechanism and analyze its application in the emerging CMP+SMT scenario. The included analysis points out the new difficulties to be faced by the FLUSH mechanism in the emerging CMP+SMT scenario. Then we propose a novel IFetch Policy designed to cope with the CMP+SMT scenario: the MFLUSH. We also include a complete evaluation of the MFLUSH policy, both in terms of throughput and energy consumption. Our results indicate that the MFLUSH, specifically designed for the emerging CMP+SMT scenario, succeeds not only in overcoming the specific CMP+SMT constraints but also allowing a 20% energy consumption reduction without a significant system throughput loss. Carmelo Acosta, Francisco J. Cazorla, Alex Ramírez, Mateo Valero |
ICPP | 3 |
| 2008 | Analysis of video filtering on the cell processorabstractIn this paper an analysis of bi-dimensional video altering on the cell broadband engine processor is presented. To evaluate the processor, a highly adaptive altering algorithm was chosen: the deblocking filter of the H.264 video compression standard. The baseline version is a scalar implementation extracted from the FFMPEG H.264 decoder. The scalar version was vectorized using the SIMD instructions of the cell synergistic processing element (SPE) and with AltiVec instructions for the Power Processor Element. Results show that approximately one third of the processing time of the SPE SIMD version is used for transposition and data packing and unpacking. Despite the required SIMD overhead and the high adaptivity of the kernel, the SIMD version of the kernel is 2.6 times faster than the scalar versions. Arnaldo Azevedo, Cor Meenderinck, Ben H. H. Juurlink, Mauricio Alvarez-Mesa, Alex Ramírez |
ISCAS | 5 |
| 2007 | MLP-Aware Dynamic Cache Partitioning
Miquel Moretó, Francisco J. Cazorla, Alex Ramírez, Mateo Valero |
PACT | 3 |
| 2007 | Performance Impact of Unaligned Memory Operations in SIMD Extensions for Video Codec ApplicationsabstractAlthough SIMD extensions are a cost effective way to exploit the data level parallelism present in most media applications, we will show that they had have a very limited memory architecture with a weak support for unaligned memory accesses. In video codec, and other applications, the overhead for accessing unaligned positions without an efficient architecture support has a big performance penalty and in some cases makes vectorization counter-productive. In this paper we analyze the performance impact of extending the Altivec SIMD ISA with unaligned memory operations. Results show that for several kernels in the H.264/AVC media codec, unaligned access support provides a speedup up to 3.8times compared to the plain SIMD version, translating into an average of 1.2times in the entire application. In addition to providing a significant performance advantage, the use of unaligned memory instructions makes programming SIMD code much easier both for the manual developer and the auto vectorizing compiler Mauricio Alvarez-Mesa, Esther Salamí, Alex Ramírez, Mateo Valero |
ISPASS | 3 |
| 2007 | Performance Analysis of Cell Broadband Engine for High Memory Bandwidth ApplicationsabstractThe cell broadband engine (CBE) is designed to be a general purpose platform exposing an enormous arithmetic performance due to its eight SIMD-only synergistic processor elements (SPEs), capable of achieving 134.4 GFLOPS (16.8 GFLOPS * 8) at 2.1 GHz, and a 64-bit power processor element (PPE). Each SPE has a 256Kb non-coherent local memory, and communicates to other SPEs and main memory through its DMA controller. CBE main memory is connected to all the CBE processor elements (PPE and SPEs) through the element interconnect bus (EIB), which has a 134.4 GB/s bandwidth performance peak at half the processor speed. Therefore, CBE platform is suitable to be used by applications using MPI and streaming programming models with a potential high performance peak. In this paper we focus on the communication part of those applications, and measure the actual memory bandwidth that each of the CBE processor components can sustain. We have measured the sustained bandwidth between PPE and memory, SPE and memory, two individual SPEs to determine if this bandwidth depends on their physical location, pairs of SPEs to achieve maximum bandwidth in nearly-ideal conditions, and in a cycle of SPEs representing a streaming kind of computation. Our results on a real machine show that following some strict programming rules, individual SPE to SPE communication almost achieves the peak bandwidth when using the DMA controllers to transfer memory chunks of at least 1024 Bytes. In addition, SPE to memory bandwidth should be considered in streaming programming. For instance, implementing two data streams using 4 SPEs each can be more efficient than having a single data stream using the 8 SPEs Daniel Jiménez-González, Xavier Martorell, Alex Ramírez |
ISPASS | 3 |
| 2007 | Enlarging Instruction StreamsabstractThe stream fetch engine is a high-performance fetch architecture based on the concept of instruction stream. We call stream to a sequence of instructions from the target of a taken branch to the next taken branch, potentially containing multiple basic blocks. The long size of instruction streams makes it possible for the stream fetch engine to provide high fetch bandwidth and to hide the branch predictor access latency, leading to performance results close to a trace cache at lower implementation cost and complexity. Therefore, enlarging instruction streams is an excellent way for improving the stream fetch engine. In this paper, we present several hardware and software mechanisms focused on enlarging those streams that finalize at particular branch types. However, our results point out that focusing on particular branch types is not a good strategy due to Amdahl's law. Consequently, we propose the multiple stream predictor, a novel mechanism that deals with all branch types by combining single streams into long virtual streams. This proposal tolerates the prediction table access latency without requiring the complexity caused by additional hardware mechanisms like prediction overriding. Moreover, it provides high performance results, which are comparable to state-of-the-art fetch architectures, but with a simpler design that consumes less energy. Oliverio J. Santana, Alex Ramírez, Mateo Valero |
IEEE Trans. Computers | 2 |
| 2006 | Branch predictor guided instruction decodingabstractFast instruction decoding is a challenge for the design of CISC microprocessors. A well-known solution to overcome this problem is using a trace cache. It stores and fetches already decoded instructions, avoiding the need for decoding them again. However, implementing a trace cache involves an important increase in the fetch architecture complexity. In this paper, we propose a novel decoding architecture that reduces the fetch engine implementation cost. Instead of using a special-purpose buffer like the trace cache, our proposal stores frequently decoded instructions in the memory hierarchy. The address where the decoded instructions are stored is kept in the branch prediction mechanism, enabling it to guide our decoding architecture. This makes it possible for the processor front-end to fetch already decoded instructions from memory instead of the original non-decoded instructions. Our results show that an 8-wide superscalar processor achieves an average 14% performance improvement by using our decoding architecture. This improvement is comparable to the one achieved by using the more complex trace cache, while requiring 16% less chip area and 21% less energy consumption in the fetch architecture. Oliverio J. Santana, Ayose Falcón, Alex Ramírez, Mateo Valero |
PACT | 3 |
| 2006 | Predictable Performance in SMT Processors: Synergy between the OS and SMTsabstractCurrent operating systems (OS) perceive the different contexts of simultaneous multithreaded (SMT) processors as multiple independent processing units, although, in reality, threads executed in these units compete for the same hardware resources. Furthermore, hardware resources are assigned to threads implicitly as determined by the SMT instruction fetch (Ifetch) policy, without the control of the OS. Both factors cause a lack of control over how individual threads are executed, which can frustrate the work of the job scheduler. This presents a problem for general purpose systems, where the OS job scheduler cannot enforce priorities, and also for embedded systems, where it would be difficult to guarantee worst-case execution times. In this paper, we propose a novel strategy that enables a two-way interaction between the OS and the SMT processor and allows the OS to run jobs at a certain percentage of their maximum speed, regardless of the workload in which these jobs are executed. In contrast to previous approaches, our approach enables the OS to run time-critical jobs without dedicating all internal resources to them so that non-time-critical jobs can make significant progress as well and without significantly compromising overall throughput. In fact, our mechanism, in addition to fulfilling OS requirements, achieves 90 percent of the throughput of one of the best currently known fetch policies for SMTs. Francisco J. Cazorla, Peter M. W. Knijnenburg, Rizos Sakellariou, Enrique Fernández, Alex Ramírez, Mateo Valero |
IEEE Trans. Computers | 5 |
| 2005 | Architectural support for real-time task scheduling in SMT processorsabstractIn Simultaneous Multithreaded (SMT) architectures most hardware resources are shared between threads. This provides a good cost/performance trade-off which renders these architectures suitable for use in embedded systems. However, since threads share many resources, they also interfere with each other. As a result, execution times of applications become highly unpredictable and dependent on the context in which an application is executed. Obviously, this poses problems if an SMT is to be used in a real-time system.In this paper, we propose two novel hardware mechanisms that can be used to reduce this performance variability. In contrast to previous approaches, our proposed mechanisms do not need any information beyond the information already known by traditional job schedulers. Nor do they require extensive profiling of workloads to determine optimal schedules. Our mechanisms are based on dynamic resource partitioning. The OS level job scheduler needs to be slightly adapted in order to provide the hardware resource allocator some information on how this resource partitioning needs to be done. We show that our mechanisms provide high stability for SMT architectures to be used in real-time systems: the real time benchmarks we used meet their deadlines in more than 98% of the cases considered while the other thread in the workload still achieves high throughput. Francisco J. Cazorla, Peter M. W. Knijnenburg, Rizos Sakellariou, Enrique Fernández, Alex Ramírez, Mateo Valero |
CASES | 5 |
| 2005 | A Complexity-Effective Simultaneous Multithreading ArchitectureabstractDifferent applications may exhibit radically different behaviors and thus have very different requirements in terms of hardware support. In simultaneous multithreading (SMT) architectures, the hardware is shared among multiple running applications in order to better profit from it. However, current architectures are designed for the common case, and try to satisfy a number of different application classes with a single design. That is, current designs are usually overdesigned for most cases, obtaining high performance, but wasting a lot of resources to do so. In this paper we present an alternative SMT architecture, the heterogeneously distributed SMT (hdSMT). Our architecture is based in a novel combination of SMT and clustering techniques in a heterogeneity-aware fashion. The hardware is designed to match the heterogeneous application behavior with the statically and heterogeneously partitioned resources. Such a design is aimed for minimizing the amount of resources wasted to achieve a given performance rate. On top of our statically partitioned architecture, we propose a heuristic policy to map threads to clusters so that each cluster matches the characteristics of the running threads and overall hardware usage is optimized. We compare our hdSMT architecture with a monolithic SMT processor, where all threads compete for the same resources, and with a homogeneous clustered SMT, where resources are statically and equally partitioned across clusters. Our results show that hdSMT architectures obtain an average improvement of 13% and 14% in optimizing performance per area over monolithic SMT and homogeneously clustered SMT respectively. Carmelo Acosta, Ayose Falcón, Alex Ramírez, Mateo Valero |
ICPP | 3 |
| 2005 | On the Scalability of 1- and 2-Dimensional SIMD Extensions for Multimedia ApplicationsabstractSIMD extensions are the most common technique used in current processors for multimedia computing. In order to obtain more performance for emerging applications SIMD extensions need to be scaled. In this paper we perform a scalability analysis of SIMD extensions for multimedia applications. Scaling a 1-dimensional extension, like Intel MMX, was compared to scaling a 2-dimensional (matrix) extension. Evaluations have demonstrated that the 2-d architecture is able to use more parallel hardware than the 1-d extension. Speed-ups over a 2-way superscalar processor with MMX-like extension go up to 4X for kernels and up to 3.3X for complete applications and the matrix architecture can deliver, in some cases, more performance with simpler processor configurations. The experiments also show that the scaled matrix architecture is reaching the limits of the DLP available in the internal loops of common multimedia kernels Friman Sánchez, Mauricio Alvarez-Mesa, Esther Salamí, Alex Ramírez, Mateo Valero |
ISPASS | 4 |
| 2005 | Software Trace CacheabstractWe explore the use of compiler optimizations, which optimize the layout of instructions in memory. The target is to enable the code to make better use of the underlying hardware resources regardless of the specific details of the processor/architecture in order to increase fetch performance. The Software Trace Cache (STC) is a code layout algorithm with a broader target than previous layout optimizations. We target not only an improvement in the instruction cache hit rate, but also an increase in the effective fetch width of the fetch engine. The STC algorithm organizes basic blocks into chains trying to make sequentially executed basic blocks reside in consecutive memory positions, then maps the basic block chains in memory to minimize conflict misses in the important sections of the program. We evaluate and analyze in detail the impact of the STC, and code layout optimizations in general, on the three main aspects of fetch performance; the instruction cache hit rate, the effective fetch width, and the branch prediction accuracy. Our results show that layout optimized, codes have some special characteristics that make them more amenable for high-performance instruction fetch. They have a very high rate of not-taken branches and execute long chains of sequential instructions; also, they make very effective use of instruction cache lines, mapping only useful instructions which will execute close in time, increasing both spatial and temporal locality. Alex Ramírez, Josep Lluís Larriba-Pey, Mateo Valero |
IEEE Trans. Computers | 1 |
| 2004 | Implicit vs. Explicit Resource Allocation in SMT ProcessorsabstractIn a simultaneous multithreaded (SMT) architecture, the front end of a superscalar is adapted in order to be able to fetch from several threads while the back end is shared among the threads. In this paper, we describe different resource sharing models in SMT processors. We show that explicit resource allocation can improve SMT performance. In addition, it enables SMTs to solve other QoS requirements, not realizable before. Francisco J. Cazorla, Peter M. W. Knijnenburg, Rizos Sakellariou, Enrique Fernández, Alex Ramírez, Mateo Valero |
DSD | 5 |
| 2004 | Feasibility of QoS for SMT
Francisco J. Cazorla, Peter M. W. Knijnenburg, Rizos Sakellariou, Enrique Fernández, Alex Ramírez, Mateo Valero |
Euro-Par | 5 |
| 2004 | A Low-Complexity, High-Performance Fetch Unit for Simultaneous Multithreading ProcessorsabstractSimultaneous multithreading (SMT) is an architectural technique that allows for the parallel execution of several threads simultaneously. Fetch performance has been identified as the most important bottleneck for SMT processors. The commonly adopted solution has been fetching from more than one thread each cycle. Recent studies have proposed a plethora of fetch policies to deal with fetch priority among threads, trying to increase fetch performance. We demonstrate that the simultaneous sharing of the fetch unit, apart from increasing the complexity of the fetch unit, can be counterproductive in terms of performance. We evaluate the use of high-performance fetch units in the context of SMT. Our new fetch architecture proposal allows us to feed an 8-way processor fetching from a single thread each cycle, reducing complexity, and increasing the usefulness of proposed fetch policies. Our results show that using new high-performance fetch units, like the FTB or the stream fetch, provides higher performance than fetching from two threads using common SMT fetch architectures. Furthermore, our results show that our design obtains better average performance for any kind of workloads (both ILP and memory bounded benchmarks), in contrast to previously proposed solutions. Ayose Falcón, Alex Ramírez, Mateo Valero |
HPCA | 2 |
| 2004 | DCache Warn: An I-Fetch Policy to Increase SMT EfficiencyabstractSummary form only given. Simultaneous multithreading (SMT) processors increase performance by executing instructions from multiple threads simultaneously. These threads share the processor's resources, but also compete for them. In this environment, a thread missing in the L2 cache may allocate a large number of resources for a long time, causing other threads to run much slower than they could. To prevent this problem we should know in advance if a thread is going to miss in the L2 cache. L1 misses are a clear indicator of a possible L2 miss. However, to stall a thread on every L1 miss is too severe, because not all L1 misses lead to an L2 miss, and this would cause an unnecessary stall and resource under-use. Also, to wait until an L2 miss is declared and squash the thread to free up the allocated resources is too expensive in terms of complexity and reexecuted instructions. We propose a novel fetch policy, which we call DWarn. DWarn uses L1 misses as indicators of L2 misses, giving higher priority to threads with no outstanding L1 misses. DWarn acts on L1 misses, before L2 misses happen in a controlled manner to reduce resource under-use and to avoid harming a thread when L1 misses do not lead to L2 misses. Our results show that DWarn outperforms previously proposed policies, in both throughput and fairness, while requiring fewer resources and avoiding instruction reexecution. Francisco J. Cazorla, Alex Ramírez, Mateo Valero, Enrique Fernández |
IPDPS | 2 |
| 2004 | Prophet/Critic Hybrid Branch PredictionabstractThis paper introduces the prophet/critic hybrid conditional branch predictor, which has two component predictors that play the role of either prophet or critic. The prophet is a conventional predictor that uses branch history to predict the direction of the current branch. Further accesses of the prophet yield predictions for the branches following the current one. Predictions for the current branch and the ones that follow are collectively known as the branch's future. They are actually a prophecy, or predicted branch future. The critic uses both the branch's history and future to give a critique of the prophet's prediction fo the current branch. The critique, either agree or disagree, is used to generate the final prediction for the branch. Our results show an 8K + 8K byte prophet/critic hybrid has 39% fewer mispredicts than a 16K byte 2Bc - gskew predictor-a predictor similar to that of the proposed Compaq* Alpha* EV8 processor - across a wide range of applications. The distance between pipeline flushes due to mispredicts increases from one flush per 418 micro-operations (uops) to one per 680 uops. For gcc, the percentage of mispredicted branches drops from 3.11% to 1.23%. On a machine based on the Intel/spl reg/ Pentium/spl reg/ 4 processor, this improves uPC (Uops Per Cycle) by 7.8% (18% for gcc) and reduces the number of uops fetched (along both correct and incorrect paths) by 8.6%. Ayose Falcón, Jared Stark, Alex Ramírez, Konrad Lai, Mateo Valero |
ISCA | 3 |
| 2004 | Dynamically Controlled Resource Allocation in SMT ProcessorsabstractSMT processors increase performance by executing instructions from several threads simultaneously. These threads use the resources of the processor better by sharing them but, at the same time, threads are competing for these resources. The way critical resources are distributed among threads determines the final performance. Currently, processor resources are distributed among threads as determined by the fetch policy that decides which threads enter the processor to compete for resources. However, current fetch policies only use indirect indicators of resource usage in their decision, which can lead to resource monopolization by a single thread or to resource waste when no thread can use them. Both situations can harm performance and happen, for example, after an L2 cache miss. In this paper, we introduce the concept of dynamic resource control in SMT processors. Using this concept, we propose a novel resource allocation policy for SMT processors. This policy directly monitors the usage of resources by each thread and guarantees that all threads get their fair share of the critical shared resources, avoiding monopolization. We also define a mechanism to allow a thread to borrow resources from another thread if that thread does not require them, thereby reducing resource under-use. Simulation results show that our dynamic resource allocation policy outperforms a static resource allocation policy by 8%, on average. It also improves the best dynamic resource-conscious fetch policies like FLUSH++ by 4%, on average, using the harmonic mean as a metric. This indicates that our policy does not obtain the ILP boost by unfairly running high ILP threads over slow memory-bounded threads. Instead, it achieves a better throughput-fairness balance. Francisco J. Cazorla, Alex Ramírez, Mateo Valero, Enrique Fernández |
MICRO | 2 |
| 2004 | A low-complexity fetch architecture for high-performance superscalar processorsabstractFetch engine performance is a key topic in superscalar processors, since it limits the instruction-level parallelism that can be exploited by the execution core. In the search of high performance, the fetch engine has evolved toward more efficient designs, but its complexity has also increased.In this paper, we present the stream fetch engine, a novel architecture based on the execution of long streams of sequential instructions, taking maximum advantage of code layout optimizations. We describe our design in detail, showing that it achieves high fetch performance, while requiring less complexity than other state-of-the-art fetch architectures. Oliverio J. Santana, Alex Ramírez, Josep Lluís Larriba-Pey, Mateo Valero |
ACM Trans. Archit. Code Optim. | 2 |
| 2002 | A Comparative Study of Redundancy in Trace Caches (Research Note)
Hans Vandierendonck, Alex Ramírez, Koen De Bosschere, Mateo Valero |
Euro-Par | 2 |
| 2002 | Fetching instruction streamsabstractFetch performance is a very important factor because it effectively limits the overall processor performance. However there is little performance advantage in increasing front-end performance beyond what the back-end can consume. For each processor design, the target is to build the best possible fetch engine for the required performance level. A fetch engine will be better if it provides better performance, but also if it takes fewer resources, requires less chip area, or consumes less power. In this paper we propose a novel fetch architecture based on the execution of long streams of sequential instructions, taking maximum advantage of code layout optimizations. We describe our architecture in detail, and show that it requires less complexity and resources than other high performance fetch architectures like the trace cache, while providing a high fetch performance suitable for wide-issue superscalar processors. Our results show that using our fetch architecture and code layout optimizations obtains 10% higher performance than the EV8 fetch architecture, and 4% higher than the FTB architecture using state-of-the-art branch predictors, while being only 1.5% slower than the trace cache. Even in the absence of code layout optimizations, fetching instruction streams is still 10% faster than the EV8, and only 4% slower than the trace cache. Fetching instruction streams effectively exploits the special characteristics of layout optimized codes to provide a high fetch performance, close to that of a trace cache, but has a much lower cost and complexity, similar to that of a basic block architecture. Alex Ramírez, Oliverio J. Santana, Josep Lluís Larriba-Pey, Mateo Valero |
MICRO | 1 |
| 2001 | Branch Prediction Using Profile Data
Alex Ramírez, Josep Lluís Larriba-Pey, Mateo Valero |
Euro-Par | 1 |
| 2001 | Code layout optimizations for transaction processing workloadsabstractCommercial applications such as databases and Web servers constitute the most important market segment for high-performance servers. Among these applications, on-line transaction processing (OLTP) workloads provide a challenging set of requirements for system designs since they often exhibit inefficient executions dominated by a large memory stall component. This behavior arises from large instruction and data footprints and high communication miss rates. A number of recent studies have characterized the behavior of commercial workloads and proposed architectural features to improve their performance. However, there has been little research on the impact of software and compiler-level optimizations for improving the behavior of such workloads. Alex Ramírez, Luiz André Barroso, Kourosh Gharachorloo, Robert S. Cohn, Josep Lluís Larriba-Pey, P. Geoffrey Lowney, Mateo Valero |
ISCA | 1 |
| 2001 | Instruction fetch architectures and code layout optimizationsabstractThe design of higher performance processors has been following two major trends: increasing the pipeline depth to allow faster clock rates, and widening the pipeline to allow parallel execution of more instructions. Designing a higher performance processor implies balancing all the pipeline stages to ensure that overall performance is not dominated by any of them. This means that a faster execution engine also requires a faster fetch engine, to ensure that it is possible to read and decode enough instructions to keep the pipeline full and the functional units busy. This paper explores the challenges faced by the instruction fetch stage for a variety of processor designs, from early pipelined processors, to the more aggressive wide issue superscalars. We describe the different fetch engines proposed in the literature, the performance issues involved, and some of the proposed improvements. We also show how compiler techniques that optimize the layout of the code in memory can be used to improve the fetch performance of the different engines described Overall, we show how instruction fetch has evolved from fetching one instruction every few cycles, to fetching one instruction per cycle, to fetching a full basic block per cycle, to several basic blocks per cycle: the evolution of the mechanism surrounding the instruction cache, and the different compiler optimizations used to better employ these mechanisms. Alex Ramírez, Josep Lluís Larriba-Pey, Mateo Valero |
Proc. IEEE | 1 |
| 2000 | On the Performance of Fetch Engines Running DSS Workloads
Carlos Navarro, Alex Ramírez, Josep Lluís Larriba-Pey, Mateo Valero |
Euro-Par | 2 |
| 2000 | Trace Cache Redundancy: Red & Blue TracesabstractThe objective of this paper is to improve the use of the hardware resources of the trace cache mechanism, reducing the implementation cost with no performance degradation. We achieve that by eliminating the replication of traces between the instruction cache and the trace cache. As we show, the trace cache mechanism is generating a high degree of redundancy between the traces stored in the trace cache and those built by the compiler, already present in the instruction cache. Furthermore, code reordering techniques like the software trace cache arrange the basic blocks in a program so that the fall-through path is the most common, effectively increasing this trace redundancy. We propose selective trace storage to avoid trace redundancy between the trace cache and the instruction cache. A simple modification of the fill unit allows the trace cache to store only those traces containing taken branches, which can nor be obtained in a single cycle from the instruction cache. Our results show that selective trace storage and the software trace cache used on a 32 entry trace cache (2 KB) perform as well as a 2048 entry trace cache (128 KB) without the enhancements. This shows that the cooperation between hardware and software is crucial to improve the performance and reduce the requirements of hardware mechanisms in the fetch engine. Alex Ramírez, Josep Lluís Larriba-Pey, Mateo Valero |
HPCA | 1 |
| 1999 | Optimization of Instruction Fetch for Decision Support WorkloadsabstractInstruction fetch bandwidth is feared to be a major limiting factor to the performance of future wide-issue aggressive superscalars. In this paper, we focus on database applications running decision support workloads. We characterize the locality patterns of ia database kernel and find frequently executed paths. Using this information, we propose an algorithm to lay out the basic blocks for improved I-fetch. Our results show a miss reduction of 60-98% for realistic I-cache sizes and a doubling of the number of instructions executed between taken branches. As a consequence, we increase the fetch bandwith provided by an aggressive sequential fetch unit from 5.8 for the original code to 10.6 using our proposed layout. Our software scheme combines well with hardware schemes like a trace cache providing up to 12.1 instruction per cycle, suggesting that commercial workloads may be amenable to the aggressive I-fetch of future superscalars. Alex Ramírez, Josep Lluís Larriba-Pey, Carlos Navarro, Xavi Serrano, Mateo Valero, Josep Torrellas |
ICPP | 1 |
| 1999 | Software trace cacheabstractIn this paper we address the important problem of instruction fetch for future wide issue superscalar processors.Our approach focuses on understanding the interaction between software and hardware techniques targeting an increase in the instruction fetch bandwidth.That is the objective, for instance, of the Hardware Trace Cache (HTC).We design a profile based code reordering technique which targets a maximization of the sequentiality of instructions, while still trying to minimize instruction cache misses.We call our software approach, Software Xrace Cache {STC).We evaluate our software approach, and then compare it with the HTC and the combination of both techniques.Our results show that for large codes with few loops and deterministic execution sequences like databases and some SPEC-INT codes, the STC offers similar, or better, results than a HTC.Moreover, when combining the software and hardware approaches, we obtain encouraging results: the STC and a small HTC offer similar performance to a much larger HTC alone.1 Introduction Instruction fetch bandwidth may become a major limiting factor for future aggressive wide-issue superscalars.Consequently, it is crucial to develop software and hardware techniques that interact to deliver multiple basic blocks to the processor every cycle.Unfortunately, for many important codes, this is hard to do.For instance, database codes and several integer SPEC *This research has been SuDDOrted bv CICYT erant TIC-0511-98 (UPC authors), the Generaiitat de Catalunyagrants ACI 97-26 (Josep L. Larriba-Pey and Josep Torrellas) and 1998FI-00306-APTIND (Alex Ramiree), the Commission for Cultural, Educational and Scientific Exchange'between the United States of .America and Spain Alex Ramírez, Josep Lluís Larriba-Pey, Carlos Navarro, Josep Torrellas, Mateo Valero |
International Conference on Supercomputing | 1 |