VLDB 2026 Research / reviewers in the wild / expert
José A. Joao
dblp:49/5144
· DBLP profile ↗
16ranked-venue papers
4as first author
2since 2021 · last 2025
0000-0002-3571-5562ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 8 · 4 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
13 papers |
Memory systems · 55% Processor architecture and microarchitecture · 27% Parallel and multicore computing · 12% | |
| Software engineering, system software, and programming languages
5 papers |
Runtime systems and virtual machines · 70% Operating systems · 20% Program analysis · 9% |
Topics — the 30 heaviest of 36, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory access
atomic memory operation |
1.5 | 2 | 2025 | A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets · MICRO 2025 DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory Operations · ISCA 2023 |
Memory systems
cache coherence |
1.5 | 2 | 2025 | A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets · MICRO 2025 DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory Operations · ISCA 2023 |
Processor architecture and microarchitecture › multi-chip architecture
chiplet architecture |
0.9 | 1 | 2025 | A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets · MICRO 2025 |
Parallel and multicore computing
synchronization |
0.7 | 1 | 2023 | DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory Operations · ISCA 2023 |
Energy-efficient computing
thermal management |
0.4 | 1 | 2019 | DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019 |
Processor architecture and microarchitecture
multicore design |
0.3 | 2 | 2023 | DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory Operations · ISCA 2023 Bottleneck identification and scheduling in multithreaded applications · ASPLOS 2012 |
Memory systems
integrity tree |
0.3 | 1 | 2018 | Morphable Counters: Enabling Compact Integrity Trees For Low-Overhead Secure Memories · MICRO 2018 |
Memory systems
secure memory |
0.3 | 1 | 2018 | Morphable Counters: Enabling Compact Integrity Trees For Low-Overhead Secure Memories · MICRO 2018 |
Processor architecture and microarchitecture › chip multiprocessor
heterogeneous chip multiprocessor |
0.3 | 2 | 2013 | Utility-based acceleration of multithreaded applications on asymmetric CMPs · ISCA 2013 Bottleneck identification and scheduling in multithreaded applications · ASPLOS 2012 |
Memory systems › memory hierarchy
cache hierarchy |
0.3 | 1 | 2025 | A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets · MICRO 2025 |
Memory systems › cache coherence
directory |
0.3 | 1 | 2025 | A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets · MICRO 2025 |
Processor architecture and microarchitecture
branch prediction |
0.2 | 3 | 2009 | Virtual Program Counter (VPC) Prediction: Very Low Cost Indirect Branch Prediction Using Conditional Branch Prediction Hardware · IEEE Trans. Computers 2009 VPC prediction: reducing the cost of indirect branches via hardware-based dynamic devirtualization · ISCA 2007 Diverge-Merge Processor (DMP): Dynamic Predicated Execution of Complex Control-Flow Graphs Based on Frequently Executed Paths · MICRO 2006 |
Processor architecture and microarchitecture › branch prediction
indirect branch prediction |
0.2 | 2 | 2009 | Virtual Program Counter (VPC) Prediction: Very Low Cost Indirect Branch Prediction Using Conditional Branch Prediction Hardware · IEEE Trans. Computers 2009 VPC prediction: reducing the cost of indirect branches via hardware-based dynamic devirtualization · ISCA 2007 |
Parallel and multicore computing › parallel scheduling
thread scheduling |
0.2 | 1 | 2013 | Utility-based acceleration of multithreaded applications on asymmetric CMPs · ISCA 2013 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2011 | Data marshaling for multi-core architectures · ISCA 2010 Parallel application memory scheduling · MICRO 2011 |
Parallel and multicore computing › parallel computing › parallel application performance
multithreaded application performance |
0.1 | 1 | 2012 | Bottleneck identification and scheduling in multithreaded applications · ASPLOS 2012 |
Memory systems › memory controller
memory scheduling |
0.1 | 1 | 2011 | Parallel application memory scheduling · MICRO 2011 |
Memory systems
cache |
0.1 | 1 | 2019 | DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019 |
Memory systems › cache management
cache capacity management |
0.1 | 1 | 2019 | DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019 |
Memory systems › memory hierarchy › cache hierarchy
last-level cache |
0.1 | 1 | 2019 | DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019 |
Memory systems
cache management |
0.1 | 1 | 2010 | Data marshaling for multi-core architectures · ISCA 2010 |
Processor architecture and microarchitecture › chip multiprocessor
inter-core communication |
0.1 | 1 | 2010 | Data marshaling for multi-core architectures · ISCA 2010 |
Hardware security and side channels › memory security
memory encryption and integrity |
0.1 | 1 | 2018 | Morphable Counters: Enabling Compact Integrity Trees For Low-Overhead Secure Memories · MICRO 2018 |
Runtime systems and virtual machines
garbage collection |
0.1 | 1 | 2009 | Flexible reference-counting-based hardware acceleration for garbage collection · ISCA 2009 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.1 | 1 | 2008 | Improving the performance of object-oriented languages with dynamic predication of indirect jumps · ASPLOS 2008 |
Processor architecture and microarchitecture › branch prediction
hard-to-predict branch |
0.1 | 1 | 2006 | Diverge-Merge Processor (DMP): Dynamic Predicated Execution of Complex Control-Flow Graphs Based on Frequently Executed Paths · MICRO 2006 |
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution |
0.1 | 1 | 2006 | Diverge-Merge Processor (DMP): Dynamic Predicated Execution of Complex Control-Flow Graphs Based on Frequently Executed Paths · MICRO 2006 |
Electronic design automation › timing analysis
critical path analysis |
0.0 | 1 | 2013 | Utility-based acceleration of multithreaded applications on asymmetric CMPs · ISCA 2013 |
Operating systems › resource management › process management › CPU scheduling
thread scheduling |
0.0 | 1 | 2012 | Bottleneck identification and scheduling in multithreaded applications · ASPLOS 2012 |
Parallel and multicore computing
thread-level parallelism |
0.0 | 1 | 2011 | Parallel application memory scheduling · MICRO 2011 |
Methods — techniques the papers use, named apart from their topics
tracing · 0.9prediction · 0.9simulation · 0.7thermal headroom modeling · 0.4dynamic utility prediction · 0.4dynamic predication · 0.3critical-path measurement · 0.3cooperative software-hardware mechanism · 0.3utility-based acceleration · 0.2software/hardware cooperative scheduling · 0.2reference counting · 0.1hardware acceleration · 0.1branch prediction · 0.1hardware-based dynamic devirtualization · 0.1control-flow hint · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A. Delegato: Locality-Aware Atomic Memory Operations on ChipletsabstractThe irruption of chiplet-based architectures has been a game changer, enabling higher transistor integration and core counts in a single socket.However, chiplets impose higher and non-uniform memory access (NUMA) latencies than monolithic integration.This harms the efficiency of atomic memory operations (AMOs), which are fundamental to implementing fine-grained synchronization and concurrent data structures on large systems.AMOs are executed either near the core (near) or at a remote location within the cache hierarchy (far).On near AMOs, the core's private cache fetches the target cache line in exclusiveness to modify it locally.Near AMOs cause significant data movement between private caches, especially harming parallel applications' performance on chiplet-based architectures.Alternatively, far AMOs can alleviate the communication overhead by reducing data movement between processing elements.However, current multicore architectures only support one type of far AMO, which sends all updates to a single serialization point (centralized AMOs).This work introduces two new types of far AMOs, delegated and migrating, that execute AMOs remotely without centralizing updates in a single point of the cache hierarchy.Combining centralized, delegated, and migrating AMOs allows the directory to select the best location to execute AMOs.Moreover, we propose Delegato, a tracing optimization to effectively transport usage information from private caches to the directory to predict the best atomic type to issue accurately.Additionally, we design a simple predictor on Víctor Soria 0001, Adrià Armejach, Tiago Rogério Mück, Darío Suárez Gracia, José A. Joao, Miquel Moretó |
MICRO | 5 |
| 2023 | DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory OperationsabstractWith increasing core counts in modern multi-core designs, the overhead of synchronization jeopardizes the scalability and efficiency of parallel applications. To mitigate these overheads, modern cache-coherent protocols offer support for Atomic Memory Operations (AMOs) that can be executed near-core (near) or remotely in the on-chip memory hierarchy (far). Víctor Soria 0001, Adrià Armejach, Tiago Rogério Mück, Darío Suárez Gracia, José A. Joao, Alejandro Rico, Miquel Moretó |
ISCA | 5 |
| 2019 | DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal ManagementabstractSprinting is a class of mechanisms that provides a short but significant performance boost while temporarily exceeding the thermal design point. We propose DynaSprint, a software runtime that manages sprints by dynamically predicting utility and modeling thermal headroom. Moreover, we propose a new sprint mechanism for caches, increasing capacity briefly for enhanced performance. For a system that extends last-level cache capacity from 2MB to 4MB per core and can absorb 10J of heat, DynaSprint-guided cache sprints improve performance by 17% on average and by up to 40% over a non-sprinting system. These performance outcomes, within 95% of an oracular policy, are possible because DynaSprint accurately predicts phase behavior and sprint utility. Ziqiang Huang, José A. Joao, Alejandro Rico, Andrew D. Hilton, Benjamin C. Lee |
MICRO | 2 |
| 2018 | BUQS: Battery- and user-aware QoS scaling for interactive mobile devicesabstractBattery life has become one of major concerns for mobile user experience. Existing approaches for balancing device quality-of-service (QoS) and energy often over- or under-provision available battery capacity, or do not properly account for the non-obvious impact of QoS and battery state on actual user experience. In this paper, we propose BUQS, Battery- and User-aware QoS Scaling to maximize user experience under desired battery lifetime goals by leveraging insights about mobile device users. BUQS continually evaluates optimal QoS based on battery status and varying user expectations, and it dynamically adjusts the device service level to maximize user experience and simultaneously meet the battery lifetime requirements. BUQS recognizes dependence of user experience on both device use and battery life using an extended battery-aware quality-of-experience (QoE) model. Furthermore, BUQS learns user's behavior to predict energy demands in time and proactively rebalance energy to improve user experience. Experimental results show that our approach delivers about 30% higher QoE than the state-of-the-art QoS and energy balancing approach. Wooseok Lee, Reena Panda, Dam Sunwoo, José A. Joao, Andreas Gerstlauer, Lizy Kurian John |
ASP-DAC | 4 |
| 2018 | Accelerating Synchronization in Graph Analytics Using Moving Compute to Data Model on Tilera TILE-Gx72abstractThe shared memory cache coherence paradigm is prevalent in modern multicores. However, as the number of cores increases, synchronization between threads limits performance scaling. Hardware-based core-to-core explicit messaging has been incorporated as an auxiliary communication capability to the shared memory cache coherence paradigm in the Tilera TILE-Gx72 multicore. We propose to utilize the auxiliary explicit messaging capability to build a moving computation to data model that accelerates synchronization using fine-grain serialization of critical code regions at dedicated cores. The proposed communication model exploits data locality and improves performance over both spin-lock and atomic instruction based synchronization methods for a set of parallelized graph analytic benchmarks executing on real world graphs. Experimental results show an average 34% better performance over spin-locks, and 15% over atomic instructions at 64 cores setup on TILE-Gx72. Halit Dogan, Masab Ahmad, José A. Joao, Omer Khan |
ICCD | 3 |
| 2018 | Morphable Counters: Enabling Compact Integrity Trees For Low-Overhead Secure MemoriesabstractSecuring off-chip main memory is essential for protection from adversaries with physical access to systems. However, current secure-memory designs incur considerable performance overheads - a major cause being the multiple memory accesses required for traversing an integrity-tree, that provides protection against man-in-the-middle attacks or replay attacks. In this paper, we provide a scalable solution to this problem by proposing a compact integrity tree design that requires fewer memory accesses for its traversal. We enable this by proposing new storage-efficient representations for the counters used for encryption and integrity-tree in secure memories. Our Morphable Counters are more cacheable on-chip, as they provide more counters per cacheline than existing split counters. Additionally, they incur lower overheads due to counter-overflows, by dynamically switching between counter representations based on usage pattern. We show that using Morphable Counters enables a 128-ary integrity-tree, that can improve performance by 6.3% on average (up to 28.3%) and reduce system energy-delay product by 8.8% on average, compared to an aggressive baseline using split counters with a 64-ary integrity-tree. These benefits come without any additional storage or reduction in security and are derived from our compact counter representation, that reduces the integrity-tree size for a 16GB memory from 4MB in the baseline to 1MB. Compared to recently proposed VAULT, our design provides a speedup of 13.5% on average (up to 47.4%). Gururaj Saileshwar, Prashant J. Nair, Prakash Ramrakhyani, Wendy Elsasser, José A. Joao, Moinuddin K. Qureshi |
MICRO | 5 |
| 2013 | Utility-based acceleration of multithreaded applications on asymmetric CMPsabstractAsymmetric Chip Multiprocessors (ACMPs) are becoming a reality. ACMPs can speed up parallel applications if they can identify and accelerate code segments that are critical for performance. Proposals already exist for using coarse-grained thread scheduling and fine-grained bottleneck acceleration. Unfortunately, there have been no proposals offered thus far to decide which code segments to accelerate in cases where both coarse-grained thread scheduling and fine-grained bottleneck acceleration could have value. This paper proposes Utility-Based Acceleration of Multithreaded Applications on Asymmetric CMPs (UBA), a cooperative software/hardware mechanism for identifying and accelerating the most likely critical code segments from a set of multithreaded applications running on an ACMP. The key idea is a new Utility of Acceleration metric that quantifies the performance benefit of accelerating a bottleneck or a thread by taking into account both the criticality and the expected speedup. UBA outperforms the best of two state-of-the-art mechanisms by 11% for single application workloads and by 7% for two-application workloads on an ACMP with 52 small cores and 3 large cores. José A. Joao, M. Aater Suleman, Onur Mutlu, Yale N. Patt |
ISCA | 1 |
| 2012 | Bottleneck identification and scheduling in multithreaded applicationsabstractPerformance of multithreaded applications is limited by a variety of bottlenecks, e.g. critical sections, barriers and slow pipeline stages. These bottlenecks serialize execution, waste valuable execution cycles, and limit scalability of applications. This paper proposes Bottleneck Identification and Scheduling in Multithreaded Applications (BIS), a cooperative software-hardware mechanism to identify and accelerate the most critical bottlenecks. BIS identifies which bottlenecks are likely to reduce performance by measuring the number of cycles threads have to wait for each bottleneck, and accelerates those bottlenecks using one or more fast cores on an Asymmetric Chip Multi-Processor (ACMP). Unlike previous work that targets specific bottlenecks, BIS can identify and accelerate bottlenecks regardless of their type. We compare BIS to four previous approaches and show that it outperforms the best of them by 15% on average. BIS' performance improvement increases as the number of cores and the number of fast cores in the system increase. José A. Joao, M. Aater Suleman, Onur Mutlu, Yale N. Patt |
ASPLOS | 1 |
| 2011 | Parallel application memory schedulingabstractA primary use of chip-multiprocessor (CMP) systems is to speed up a single application by exploiting thread-level parallelism. In such systems, threads may slow each other down by issuing memory requests that interfere in the shared memory subsystem. This inter-thread memory system interference can significantly degrade parallel application performance. Better memory request scheduling may mitigate such performance degradation. However, previously proposed memory scheduling algorithms for CMPs are designed for multi-programmed workloads where each core runs an independent application, and thus do not take into account the inter-dependent nature of threads in a parallel application. Eiman Ebrahimi, Rustam Miftakhutdinov, Chris Fallin, Chang Joo Lee, José A. Joao, Onur Mutlu, Yale N. Patt |
MICRO | 5 |
| 2010 | Data marshaling for multi-core architecturesabstractPrevious research has shown that Staged Execution (SE), i.e., dividing a program into segments and executing each segment at the core that has the data and/or functionality to best run that segment, can improve performance and save power. However, SE's benefit is limited because most segments access inter-segment data, i.e., data generated by the previous segment. When consecutive segments run on different cores, accesses to inter-segment data incur cache misses, thereby reducing performance. This paper proposes Data Marshaling (DM), a new technique to eliminate cache misses to inter-segment data. DM uses profiling to identify instructions that generate inter-segment data, and adds only 96 bytes/core of storage overhead. We show that DM significantly improves the performance of two promising Staged Execution models, Accelerated Critical Sections and producer-consumer pipeline parallelism, on both homogeneous and heterogeneous multi-core systems. In both models, DM can achieve almost all of the potential of ideally eliminating cache misses to inter-segment data. DM's performance benefit increases with the number of cores. M. Aater Suleman, Onur Mutlu, José A. Joao, Khubaib, Yale N. Patt |
ISCA | 3 |
| 2009 | Flexible reference-counting-based hardware acceleration for garbage collectionabstractLanguages featuring automatic memory management (garbage collection) are increasingly used to write all kinds of applications because they provide clear software engineering and security advantages. Unfortunately, garbage collection imposes a toll on performance and introduces pause times, making such languages less attractive for high-performance or real-time applications. Much progress has been made over the last five decades to reduce the overhead of garbage collection, but it remains significant. José A. Joao, Onur Mutlu, Yale N. Patt |
ISCA | 1 |
| 2009 | Virtual Program Counter (VPC) Prediction: Very Low Cost Indirect Branch Prediction Using Conditional Branch Prediction HardwareabstractIndirect branches have become increasingly common in modular programs written in modern object-oriented languages and virtual-machine-based runtime systems. Unfortunately, the prediction accuracy of indirect branches has not improved as much as that of conditional branches. Furthermore, previously proposed indirect branch predictors usually require a significant amount of extra hardware storage and complexity, which makes them less attractive to implement. This paper proposes a new technique for handling indirect branches, called Virtual Program Counter (VPC) prediction. The key idea of VPC prediction is to use the existing conditional branch prediction hardware to predict indirect branch targets, avoiding the need for a separate storage structure. Our comprehensive evaluation shows that VPC prediction improves average performance by 26.7 percent and reduces average energy consumption by 19 percent compared to a commonly used branch target buffer based predictor on 12 indirect branch intensive C/C++ applications. Moreover, VPC prediction improves the average performance of the full set of object-oriented Java DaCapo applications by 21.9 percent, while reducing their average energy consumption by 22 percent. We show that VPC prediction can be used with any existing conditional branch prediction mechanism and that the accuracy of VPC prediction improves when a more accurate conditional branch predictor is used. Hyesoon Kim, José A. Joao, Onur Mutlu, Chang Joo Lee, Yale N. Patt, Robert S. Cohn |
IEEE Trans. Computers | 2 |
| 2008 | Improving the performance of object-oriented languages with dynamic predication of indirect jumpsabstractIndirect jump instructions are used to implement increasingly-common programming constructs such as virtual function calls, switch-case statements, jump tables, and interface calls. The performance impact of indirect jumps is likely to increase because indirect jumps with multiple targets are difficult to predict even with specialized hardware. José A. Joao, Onur Mutlu, Hyesoon Kim, Rishi Agarwal, Yale N. Patt |
ASPLOS | 1 |
| 2007 | Profile-assisted Compiler Support for Dynamic Predication in Diverge-Merge ProcessorsabstractDynamic predication has been proposed to reduce the branch misprediction penalty due to hard-to-predict branch instructions. A proposed dynamic predication architecture, the diverge-merge processor (DMP), provides large performance improvements by dynamically predicating a large set of complex control-flow graphs that result in branch mispredictions. DMP requires significant support from a profiling compiler to determine which branch instructions and control-flow structures can be dynamically predicated. However, previous work on dynamic predication did not extensively examine the tradeoffs involved in profiling and code generation for dynamic predication architectures. This paper describes compiler support for obtaining high performance in the diverge-merge processor. We describe new profile-driven algorithms and heuristics to select branch instructions that are suitable and profitable for dynamic predication. We also develop a new profile-based analytical cost-benefit model to estimate, at compile-time, the performance benefits of the dynamic predication of different types of control-flow structures including complex hammocks and loops. Our evaluations show that DMP can provide 20.4% average performance improvement over a conventional processor on SPEC integer benchmarks with our optimized compiler algorithms, whereas the average performance improvement of the best-performing alternative simple compiler algorithm is 4.5%. We also find that, with the proposed algorithms, DMP performance is not significantly affected by the differences in profile- and run-time input data sets Hyesoon Kim, José A. Joao, Onur Mutlu, Yale N. Patt |
CGO | 2 |
| 2007 | VPC prediction: reducing the cost of indirect branches via hardware-based dynamic devirtualizationabstractIndirect branches have become increasingly common in modular programs written in modern object-oriented languages and virtual machine based runtime systems. Unfortunately, the prediction accuracy of indirect branches has not improved as much as that of conditional branches. Furthermore, previously proposed indirect branch predictors usually require a significant amount of extra hardware storage and complexity, which makes them less attractive to implement. Hyesoon Kim, José A. Joao, Onur Mutlu, Chang Joo Lee, Yale N. Patt, Robert S. Cohn |
ISCA | 2 |
| 2006 | Diverge-Merge Processor (DMP): Dynamic Predicated Execution of Complex Control-Flow Graphs Based on Frequently Executed PathsabstractThis paper proposes a new processor architecture for handling hard-to-predict branches, the diverge-merge processor (DMP). The goal of this paradigm is to eliminate branch mispredictions due to hard-to-predict dynamic branches by dynamically predicating them without requiring ISA support for predicate registers and predicated instructions. To achieve this without incurring large hardware cost and complexity, the compiler provides control-flow information by hints and the processor dynamically predicates instructions only on frequently executed program paths. The key insight behind DMP is that most control-flow graphs look and behave like simple hammock (if-else) structures when only frequently executed paths in the graphs are considered. Therefore, DMP can dynamically predicate a much larger set of branches than simple hammock branches. Our evaluations show that DMP out performs a baseline processor with an aggressive branch predictor by 19.3% on average over SPEC integer 95 and 2000 benchmarks, through a reduction of 38% in pipeline flushes due to branch mispredictions, while consuming 9.0% less energy. We also compare DMP with previously proposed predication and dual-path/multipath execution paradigms in terms of performance, complexity, and energy consumption, and find that DMP is the highest performance and also the most energy-efficient design Hyesoon Kim, José A. Joao, Onur Mutlu, Yale N. Patt |
MICRO | 2 |