VLDB 2026 Research / reviewers in the wild / expert
Fernando Castro
dblp:46/4601
· DBLP profile ↗
16ranked-venue papers
4as first author
1since 2021 · last 2022
0000-0002-2773-3023ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 71% Processor architecture and microarchitecture · 20% Energy-efficient computing · 7% | |
| Software engineering, system software, and programming languages
2 papers |
Operating systems · 100% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Operating systems
resource management |
0.6 | 1 | 2022 | LFOC+: A Fair OS-Level Cache-Clustering Policy for Commodity Multicore Systems · IEEE Trans. Computers 2022 |
Memory systems
cache |
0.6 | 1 | 2022 | LFOC+: A Fair OS-Level Cache-Clustering Policy for Commodity Multicore Systems · IEEE Trans. Computers 2022 |
Memory systems › cache management › cache partitioning
last-level cache partitioning |
0.6 | 1 | 2022 | LFOC+: A Fair OS-Level Cache-Clustering Policy for Commodity Multicore Systems · IEEE Trans. Computers 2022 |
Processor architecture and microarchitecture
multicore design |
0.2 | 1 | 2022 | LFOC+: A Fair OS-Level Cache-Clustering Policy for Commodity Multicore Systems · IEEE Trans. Computers 2022 |
Operating systems › resource management › process management › CPU scheduling
thread scheduling |
0.2 | 1 | 2013 | Delivering fairness and priority enforcement on asymmetric multicore systems via OS scheduling · SIGMETRICS 2013 |
Processor architecture and microarchitecture
out-of-order execution |
0.2 | 2 | 2009 | Replacing Associative Load Queues: A Timing-Centric Approach · IEEE Trans. Computers 2009 DMDC: Delayed Memory Dependence Checking through Age-Based Filtering · MICRO 2006 |
Energy-efficient computing › low-power design
low-power processor design |
0.1 | 2 | 2009 | Replacing Associative Load Queues: A Timing-Centric Approach · IEEE Trans. Computers 2009 DMDC: Delayed Memory Dependence Checking through Age-Based Filtering · MICRO 2006 |
Processor architecture and microarchitecture › multicore design › heterogeneous multicore
asymmetric multicore |
0.0 | 1 | 2013 | Delivering fairness and priority enforcement on asymmetric multicore systems via OS scheduling · SIGMETRICS 2013 |
Energy-efficient computing
power management |
0.0 | 1 | 2006 | DMDC: Delayed Memory Dependence Checking through Age-Based Filtering · MICRO 2006 |
Methods — techniques the papers use, named apart from their topics
simulation · 1.1performance counters · 1.1scheduler implementation · 0.3experimental evaluation · 0.3timing-based dependence checking · 0.1hash table · 0.1age-based filtering · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | LFOC+: A Fair OS-Level Cache-Clustering Policy for Commodity Multicore SystemsabstractCommodity multicore systems are increasingly adopting hardware support that enables the system software to partition the last-level cache (LLC). This support makes it possible for the operating system (OS) to mitigate shared-resource contention effects on multicores by assigning different co-running applications to various cache partitions. Cache-clustering strategies have emerged as a way to improve throughput and fairness on platforms with cache-partitioning support. Unlike strict cache-partitioning, which allocates separate cache partitions to each application, cache-clustering allows partitions to be shared by several applications. In this article we propose LFOC+, a fair OS-level cache-clustering policy for commodity multicores. LFOC+ tries to mimic the behavior of the optimal cache-clustering solution for fairness, which we could obtain for different workloads by using a simulation tool. Our strategy continuously gathers data from performance counters to classify applications based on the degree of cache sensitivity and contentiouness, and separates cache-sensitive applications from aggressor programs to improve fairness, while providing acceptable throughput. We implemented LFOC+ in the Linux kernel and evaluated it on a system featuring an Intel Skylake processor, where we compare its effectiveness to that of four state-of-the-art policies. Our analysis reveals that LFOC+ brings a higher reduction in unfairness, and constitutes a lightweight cache-clustering policy. Juan Carlos Saez, Fernando Castro, Graziano Fanizzi, Manuel Prieto 0001 |
IEEE Trans. Computers | 2 |
| 2020 | Enabling performance portability of data-parallel OpenMP applications on asymmetric multicore processorsabstractAsymmetric multicore processors (AMPs) couple high-performance big cores and low-power small cores with the same instruction-set architecture but different features, such as clock frequency or microarchitecture. Previous work has shown that asymmetric designs may deliver higher energy efficiency than symmetric multicores for diverse workloads. Despite their benefits, AMPs pose significant challenges to runtime systems of parallel programming models. While previous work has mainly explored how to efficiently execute task-based parallel applications on AMPs, via enhancements in the runtime system, improving the performance of unmodified data-parallel applications on these architectures is still a big challenge. In this work we analyze the particular case of loop-based OpenMP applications, which are widely used today in scientific and engineering domains, and constitute the dominant application type in many parallel benchmark suites used for performance evaluation on multicore systems. We observed that conventional loop-scheduling OpenMP approaches are unable to efficiently cope with the load imbalance that naturally stems from the different performance delivered by big and small cores. Juan Carlos Saez, Fernando Castro, Manuel Prieto 0001 |
ICPP | 2 |
| 2019 | LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity MulticoresabstractMulticore processors constitute the main architecture choice for modern computing systems in different market segments. Despite their benefits, the contention that naturally appears when multiple applications compete for the use of shared resources among cores, such as the last-level cache (LLC), may lead to substantial performance degradation. This may have a negative impact on key system aspects such as throughput and fairness. Assigning the various applications in the workload to separate LLC partitions with possibly different sizes, has been proven effective to mitigate shared-resource contention effects. Adrian Garcia-Garcia, Juan Carlos Saez, Fernando Castro, Manuel Prieto 0001 |
ICPP | 3 |
| 2018 | Reuse Detector: Improving the Management of STT-RAM SLLCsabstractVarious constraints of Static Random Access Memory (SRAM) are leading to consider new memory technologies as candidates for building on-chip shared last-level caches (SLLCs). Spin-Transfer Torque RAM (STT-RAM) is currently postulated as the prime contender due to its better energy efficiency, smaller die footprint and higher scalability. However, STT-RAM also exhibits some drawbacks, like slow and energy-hungry write operations that need to be mitigated before it can be used in SLLCs for the next generation of computers. In this work, we address these shortcomings by leveraging a new management mechanism for STT-RAM SLLCs. This approach is based on the previous observation that although the stream of references arriving at the SLLC of a Chip MultiProcessor (CMP) exhibits limited temporal locality, it does exhibit reuse locality, i.e. those blocks referenced several times manifest high probability of forthcoming reuse. As such, conventional STT-RAM SLLC management mechanisms, mainly focused on exploiting temporal locality, result in low efficient behavior. In this paper, we employ a cache management mechanism that selects the contents of the SLLC aimed to exploit reuse locality instead of temporal locality. Specifically, our proposal consists in the inclusion of a Reuse Detector (RD) between private cache levels and the STT-RAM SLLC. Its mission is to detect blocks that do not exhibit reuse, in order to avoid their insertion in the SLLC, hence reducing the number of write operations and the energy consumption in the STT-RAM. Our evaluation, using multiprogrammed workloads in quad-core, eight-core and 16-core systems, reveals that our scheme reports on average, energy reductions in the SLLC in the range of 37–30%, additional energy savings in the main memory in the range of 6–8% and performance improvements of 3% (quad-core), 7% (eight-core) and 14% (16-core) compared with an STT-RAM SLLC baseline where no RD is employed. More importantly, our approach outperforms DASCA, the state-of-the-art STT-RAM SLLC management, reporting—depending on the specific scenario and the kind of applications used—SLLC energy savings in the range of 4–11% higher than those of DASCA, delivering higher performance in the range of 1.5–14% and additional improvements in DRAM energy consumption in the range of 2–9% higher than DASCA. Roberto Rodríguez-Rodríguez, Fernando Castro, Pablo Ibáñez 0001, Daniel Chaver, Víctor Viñals, Juan Carlos Saez, Manuel Prieto 0001, Luis Piñuel, Teresa Monreal Arnal, José María Llabería |
Comput. J. | 3 |
| 2017 | PMCTrack: Delivering Performance Monitoring Counter Support to the OS SchedulerabstractHardware performance monitoring counters (PMCs) have proven effective in characterizing application performance. Because PMCs can only be accessed directly at the OS privilege level, kernel-level tools must be developed to enable the end-user and userspace programs to access PMCs. A large body of work has demonstrated that the OS can perform effective runtime optimizations in multicore systems by leveraging performance-counter data. Special attention has been paid to optimizations in the OS scheduler. While existing performance monitoring tools greatly simplify the collection of PMC application data from userspace, they do not provide an architecture-agnostic kernel-level mechanism that is capable of exposing high-level PMC metrics to OS components, such as the scheduler. As a result, the implementation of PMC-based OS scheduling schemes is typically tied to specific processor models. To address this shortcoming we present PMCTrack, a novel tool for the Linux kernel that provides a simple architecture-independent mechanism that makes it possible for the OS scheduler to access per-thread PMC data. Despite being an OS-oriented tool, PMCTrack still allows the gathering of monitoring data from userspace, enabling kernel developers to carry out the necessary offline analysis and debugging to assist them during the scheduler design process. In addition, the tool provides both the OS and the user-space PMCTrack components with other insightful metrics available in modern processors and which are not directly exposed as PMCs, such as cache occupancy or energy consumption. This information is also of great value when it comes to analyzing the potential benefits of novel scheduling policies on real systems. In this paper, we analyze different case studies that demonstrate the flexibility, simplicity and powerful features of PMCTrack. Juan Carlos Saez, Adrian Pousa, Roberto Rodríguez-Rodríguez, Fernando Castro, Manuel Prieto 0001 |
Comput. J. | 4 |
| 2017 | Towards completely fair scheduling on asymmetric single-ISA multicore processors
Juan Carlos Saez, Adrian Pousa, Fernando Castro, Daniel Chaver, Manuel Prieto 0001 |
J. Parallel Distributed Comput. | 3 |
| 2015 | Write-Aware Replacement Policies for PCM-Based SystemsabstractThe gap between processor and memory speeds is one of the greatest challenges that current designers face in order to develop more powerful computer systems. In addition, the scalability of the Dynamic Random Access Memory (DRAM) technology is very limited nowadays, leading one to consider new memory technologies as candidates for the replacement of conventional DRAM. Phase-Change Memory (PCM) is currently postulated as the prime contender due to its higher scalability and lower leakage. However, compared with DRAM, PCM also exhibits some drawbacks, like lower endurance or higher dynamic energy consumption and write latency, that need to be mitigated before it can be used as the main memory technology for the next generation of computers. This work addresses the PCM endurance constraint. For this purpose, we present an analysis of conventional cache replacement policies in terms of the amount of writebacks to main memory that they imply and we also propose some new replacement algorithms for the last-level cache (LLC) with the goal of cutting down the write traffic to memory and consequently, to increase PCM lifetime without degrading system performance. In this paper, we target general purpose processors provided with this kind of non-volatile main memory and we exhaustively evaluate our proposed policies in both single- and multi-core environments. Experimental results show that, on average, compared with a conventional Least Recently Used (LRU) algorithm, some of our proposals manage to reduce the amount of writes to main memory up to 20–30% depending on the scenario evaluated, which leads to memory endurance extensions of up to 20–45%, also reducing the energy consumption in the memory hierarchy by up to 9% and hardly degrading performance. Roberto Rodríguez-Rodríguez, Fernando Castro, Daniel Chaver, Rekai González-Alberquilla, Luis Piñuel, Francisco Tirado |
Comput. J. | 2 |
| 2013 | Reducing writes in phase-change memory environments by using efficient cache replacement policiesabstractPhase Change Memory (PCM) is currently postulated as the best alternative for replacing Dynamic Random Access Memory (DRAM) as the technology used for implementing main memories, thanks to its significant advantages such as good scalability and low leakage. However, PCM also presents some drawbacks compared to DRAM, like its lower endurance. This work presents a behavior analysis of conventional cache replacement policies in terms of the amount of writes to main memory. Besides, new last level cache (LLC) replacement algorithms are exposed, aimed at reducing the number of writes to PCM and hence increasing its lifetime, without significantly degrading system performance. Roberto Rodríguez-Rodríguez, Fernando Castro, Daniel Chaver, Luis Piñuel, Francisco Tirado |
DATE | 2 |
| 2013 | Delivering fairness and priority enforcement on asymmetric multicore systems via OS schedulingabstractSymmetric-ISA (instruction set architecture) asymmetric-performance multicore processors (AMPs) were shown to deliver higher performance per watt and area than symmetric CMPs for applications with diverse architectural requirements. So, it is likely that future multicore processors will combine big power-hungry fast cores and small low-power slow ones. In this paper, we propose a novel thread scheduling algorithm that aims to improve the throughput-fairness trade-off on AMP systems. Our experimental evaluation on real hardware and using scheduler implementations on a general-purpose operating system, reveals that our proposal delivers a better throughput-fairness trade-off than previous schedulers for a wide variety of multi-application workloads including single-threaded and multithreaded applications. Juan Carlos Saez, Fernando Castro, Daniel Chaver, Manuel Prieto 0001 |
SIGMETRICS | 2 |
| 2012 | OpenIRS-UCM: an open-source multi-platform for interactive response systemsabstractInteractive Response Systems (IRS) have been gaining acceptance within the educational community in recent years and a clear proof is the growing number of commercial systems available today in the market. However, most solutions are based on systems which are closed, rigid and dependent on proprietary keypad or platform. We have developed OpenIRS-UCM, a free teaching tool for interactive polling that solves these drawbacks. It is an open source software so it allows the development of new functions by anybody. It has a friendly interface that anyone without high computer skills can use. It enables the coexistence of several commercial clickers simultaneously with smart-phones, tablets or other modern electronic devices. It is developed in Java, thus its use is not restricted to systems based on Microsoft Windows and it is independent of any proprietary software. Carlos García 0001, Fernando Castro, José Ignacio Gómez, Christian Tenllado, Daniel Chaver, José Antonio López Orozco |
ITiCSE | 2 |
| 2010 | Stack filter: Reducing L1 data cache power consumption
Rodrígo González-Alberquilla, Fernando Castro, Luis Piñuel, Francisco Tirado |
J. Syst. Archit. | 2 |
| 2009 | Using age registers for a simple load-store queue filtering
Fernando Castro, Daniel Chaver, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado |
J. Syst. Archit. | 1 |
| 2009 | Replacing Associative Load Queues: A Timing-Centric ApproachabstractOne of the main challenges of modern processor design is the implementation of a scalable and efficient mechanism to detect memory access order violations as a result of out-of-order execution. Traditional age-ordered associative load queues are complex, inefficient, and power hungry. In this paper, we introduce two new dependence checking schemes with different design tradeoffs, but both explicitly rely on timing information as a primary instrument to rule out dependence violation. Our timing-centric designs operate at a fraction of the energy cost of an associative LQ and achieve the same functionality with an insignificant performance impact on average. Studies with parallel benchmarks also show that they are equally effective and efficient in a chip-multiprocessor environment. Fernando Castro, Regana Noor, Alok Garg, Daniel Chaver, Michael C. Huang 0001, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado |
IEEE Trans. Computers | 1 |
| 2006 | Substituting associative load queue with simple hash tables in out-of-order microprocessorsabstractBuffering more in-flight instructions in an out-of-order microprocessor is a straightforward and effective method to help tolerate the long latencies generally associated with off-chip memory accesses. One of the main challenges of buffering a large number of instructions, however, is the implementation of a scalable and efficient mechanism to detect memory access order violations as a result of out-of-order scheduling of load and store instructions. Traditional CAM-based associative queues can be very slow and energy consuming. In this paper, instead of using the traditional age-based load queue to record load addresses, we explicitly record age information in address-indexed hash tables to achieve the same functionality of detecting premature loads. This alternative design eliminates associative searches and significantly reduces the energy consumption of the load queue. With simple techniques to reduce the number of false positives, performance degradation is kept at a minimum. Alok Garg, Fernando Castro, Michael C. Huang 0001, Daniel Chaver, Luis Piñuel, Manuel Prieto 0001 |
ISLPED | 2 |
| 2006 | DMDC: Delayed Memory Dependence Checking through Age-Based FilteringabstractOne of the main challenges of modern processor design is the implementation of a scalable and efficient mechanism to detect memory access order violations as a result of out-of-order execution of memory instructions. Traditional CAM-based associative queues can be very slow and energy hungry. In this paper we introduce two new management schemes. The first one is a filtering scheme based on simple age-tracking. This scheme can easily avoid 95-98% of associative load queue (LQ) searches using only a few registers. This translates into significant power savings. More importantly, however, this filtering makes our second scheme, delayed memory dependence checking (DMDC), practical. With a small hash table, DMDC completely avoids the need for an associative LQ and relies on indexing-based checking at the commit phase and hence cuts the energy spent on LQ by an average of 95%. At an average of about 0.3%, the performance impact is negligible. When the energy cost of the increased execution time is factored in, the processor still makes net energy savings of about 3-8%, depending on the configuration and the applications Fernando Castro, Luis Piñuel, Daniel Chaver, Manuel Prieto 0001, Michael C. Huang 0001, Francisco Tirado |
MICRO | 1 |
| 2005 | Load-Store Queue Management: an Energy-Efficient Design Based on a State-Filtering MechanismabstractModern microprocessors incorporate sophisticated techniques to allow early execution of loads without compromising program correctness. To do so, the structures that hold the memory instructions (load and store queues) implement several complex mechanisms to dynamically resolve the memory-based dependences. Our main objective in this paper is to design an efficient LQ-SQ structure, which saves energy without sacrificing much performance. We propose a new design that divides the load queue into two structures, a conventional associative queue and a simpler FIFO queue that does not allow associative searching. A dependence predictor predicts whether a load instruction has a memory dependence on any inflight store instruction. If so, the load is sent to the conventional associative queue. Otherwise, it is sent to the non-associative queue which can only detect dependence in an inexact and conservative way. In addition, the load will not check the store queue at execution time. These measures combined save energy consumption. We explore different predictor designs and runtime policies. Our experiments indicate that such a design can reduce the energy consumption in the load-store queue by 35-50% with an insignificant performance penalty of about 1%. When the energy cost of the increased execution time is factored in, the processor still makes net energy savings of about 3-4%. Fernando Castro, Daniel Chaver, Luis Piñuel, Manuel Prieto 0001, Francisco Tirado, Michael C. Huang 0001 |
ICCD | 1 |