José A. Joao

dblp:49/5144 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
2since 2021 · last 2025
0000-0002-3571-5562ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 8 · 4 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
13 papers
Memory systems · 55% Processor architecture and microarchitecture · 27% Parallel and multicore computing · 12%
Software engineering, system software, and programming languages
5 papers
Runtime systems and virtual machines · 70% Operating systems · 20% Program analysis · 9%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory access
atomic memory operation
1.522025
A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets · MICRO 2025
DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory Operations · ISCA 2023
Memory systems
cache coherence
1.522025
A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets · MICRO 2025
DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory Operations · ISCA 2023
Processor architecture and microarchitecture › multi-chip architecture
chiplet architecture
0.912025
A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets · MICRO 2025
Parallel and multicore computing
synchronization
0.712023
DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory Operations · ISCA 2023
Energy-efficient computing
thermal management
0.412019
DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019
Processor architecture and microarchitecture
multicore design
0.322023
DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory Operations · ISCA 2023
Bottleneck identification and scheduling in multithreaded applications · ASPLOS 2012
Memory systems
integrity tree
0.312018
Morphable Counters: Enabling Compact Integrity Trees For Low-Overhead Secure Memories · MICRO 2018
Memory systems
secure memory
0.312018
Morphable Counters: Enabling Compact Integrity Trees For Low-Overhead Secure Memories · MICRO 2018
Processor architecture and microarchitecture › chip multiprocessor
heterogeneous chip multiprocessor
0.322013
Utility-based acceleration of multithreaded applications on asymmetric CMPs · ISCA 2013
Bottleneck identification and scheduling in multithreaded applications · ASPLOS 2012
Memory systems › memory hierarchy
cache hierarchy
0.312025
A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets · MICRO 2025
Memory systems › cache coherence
directory
0.312025
A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets · MICRO 2025
Processor architecture and microarchitecture
branch prediction
0.232009
Virtual Program Counter (VPC) Prediction: Very Low Cost Indirect Branch Prediction Using Conditional Branch Prediction Hardware · IEEE Trans. Computers 2009
VPC prediction: reducing the cost of indirect branches via hardware-based dynamic devirtualization · ISCA 2007
Diverge-Merge Processor (DMP): Dynamic Predicated Execution of Complex Control-Flow Graphs Based on Frequently Executed Paths · MICRO 2006
Processor architecture and microarchitecture › branch prediction
indirect branch prediction
0.222009
Virtual Program Counter (VPC) Prediction: Very Low Cost Indirect Branch Prediction Using Conditional Branch Prediction Hardware · IEEE Trans. Computers 2009
VPC prediction: reducing the cost of indirect branches via hardware-based dynamic devirtualization · ISCA 2007
Parallel and multicore computing › parallel scheduling
thread scheduling
0.212013
Utility-based acceleration of multithreaded applications on asymmetric CMPs · ISCA 2013
Processor architecture and microarchitecture
chip multiprocessor
0.122011
Data marshaling for multi-core architectures · ISCA 2010
Parallel application memory scheduling · MICRO 2011
Parallel and multicore computing › parallel computing › parallel application performance
multithreaded application performance
0.112012
Bottleneck identification and scheduling in multithreaded applications · ASPLOS 2012
Memory systems › memory controller
memory scheduling
0.112011
Parallel application memory scheduling · MICRO 2011
Memory systems
cache
0.112019
DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019
Memory systems › cache management
cache capacity management
0.112019
DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.112019
DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management · MICRO 2019
Memory systems
cache management
0.112010
Data marshaling for multi-core architectures · ISCA 2010
Processor architecture and microarchitecture › chip multiprocessor
inter-core communication
0.112010
Data marshaling for multi-core architectures · ISCA 2010
Hardware security and side channels › memory security
memory encryption and integrity
0.112018
Morphable Counters: Enabling Compact Integrity Trees For Low-Overhead Secure Memories · MICRO 2018
Runtime systems and virtual machines
garbage collection
0.112009
Flexible reference-counting-based hardware acceleration for garbage collection · ISCA 2009
Processor architecture and microarchitecture
instruction-level parallelism
0.112008
Improving the performance of object-oriented languages with dynamic predication of indirect jumps · ASPLOS 2008
Processor architecture and microarchitecture › branch prediction
hard-to-predict branch
0.112006
Diverge-Merge Processor (DMP): Dynamic Predicated Execution of Complex Control-Flow Graphs Based on Frequently Executed Paths · MICRO 2006
Processor architecture and microarchitecture › instruction-level parallelism
predicated execution
0.112006
Diverge-Merge Processor (DMP): Dynamic Predicated Execution of Complex Control-Flow Graphs Based on Frequently Executed Paths · MICRO 2006
Electronic design automation › timing analysis
critical path analysis
0.012013
Utility-based acceleration of multithreaded applications on asymmetric CMPs · ISCA 2013
Operating systems › resource management › process management › CPU scheduling
thread scheduling
0.012012
Bottleneck identification and scheduling in multithreaded applications · ASPLOS 2012
Parallel and multicore computing
thread-level parallelism
0.012011
Parallel application memory scheduling · MICRO 2011

Methods — techniques the papers use, named apart from their topics

tracing · 0.9prediction · 0.9simulation · 0.7thermal headroom modeling · 0.4dynamic utility prediction · 0.4dynamic predication · 0.3critical-path measurement · 0.3cooperative software-hardware mechanism · 0.3utility-based acceleration · 0.2software/hardware cooperative scheduling · 0.2reference counting · 0.1hardware acceleration · 0.1branch prediction · 0.1hardware-based dynamic devirtualization · 0.1control-flow hint · 0.1
YearPublicationVenuePosition
2025 A. Delegato: Locality-Aware Atomic Memory Operations on Chiplets
abstract
The irruption of chiplet-based architectures has been a game changer, enabling higher transistor integration and core counts in a single socket.However, chiplets impose higher and non-uniform memory access (NUMA) latencies than monolithic integration.This harms the efficiency of atomic memory operations (AMOs), which are fundamental to implementing fine-grained synchronization and concurrent data structures on large systems.AMOs are executed either near the core (near) or at a remote location within the cache hierarchy (far).On near AMOs, the core's private cache fetches the target cache line in exclusiveness to modify it locally.Near AMOs cause significant data movement between private caches, especially harming parallel applications' performance on chiplet-based architectures.Alternatively, far AMOs can alleviate the communication overhead by reducing data movement between processing elements.However, current multicore architectures only support one type of far AMO, which sends all updates to a single serialization point (centralized AMOs).This work introduces two new types of far AMOs, delegated and migrating, that execute AMOs remotely without centralizing updates in a single point of the cache hierarchy.Combining centralized, delegated, and migrating AMOs allows the directory to select the best location to execute AMOs.Moreover, we propose Delegato, a tracing optimization to effectively transport usage information from private caches to the directory to predict the best atomic type to issue accurately.Additionally, we design a simple predictor on
Víctor Soria 0001, Adrià Armejach, Tiago Rogério Mück, Darío Suárez Gracia, José A. Joao, Miquel Moretó
MICRO5
2023 DynAMO: Improving Parallelism Through Dynamic Placement of Atomic Memory Operations
abstract
With increasing core counts in modern multi-core designs, the overhead of synchronization jeopardizes the scalability and efficiency of parallel applications. To mitigate these overheads, modern cache-coherent protocols offer support for Atomic Memory Operations (AMOs) that can be executed near-core (near) or remotely in the on-chip memory hierarchy (far).
Víctor Soria 0001, Adrià Armejach, Tiago Rogério Mück, Darío Suárez Gracia, José A. Joao, Alejandro Rico, Miquel Moretó
ISCA5
2019 DynaSprint: Microarchitectural Sprints with Dynamic Utility and Thermal Management
abstract
Sprinting is a class of mechanisms that provides a short but significant performance boost while temporarily exceeding the thermal design point. We propose DynaSprint, a software runtime that manages sprints by dynamically predicting utility and modeling thermal headroom. Moreover, we propose a new sprint mechanism for caches, increasing capacity briefly for enhanced performance. For a system that extends last-level cache capacity from 2MB to 4MB per core and can absorb 10J of heat, DynaSprint-guided cache sprints improve performance by 17% on average and by up to 40% over a non-sprinting system. These performance outcomes, within 95% of an oracular policy, are possible because DynaSprint accurately predicts phase behavior and sprint utility.
Ziqiang Huang, José A. Joao, Alejandro Rico, Andrew D. Hilton, Benjamin C. Lee
MICRO2
2018 BUQS: Battery- and user-aware QoS scaling for interactive mobile devices
abstract
Battery life has become one of major concerns for mobile user experience. Existing approaches for balancing device quality-of-service (QoS) and energy often over- or under-provision available battery capacity, or do not properly account for the non-obvious impact of QoS and battery state on actual user experience. In this paper, we propose BUQS, Battery- and User-aware QoS Scaling to maximize user experience under desired battery lifetime goals by leveraging insights about mobile device users. BUQS continually evaluates optimal QoS based on battery status and varying user expectations, and it dynamically adjusts the device service level to maximize user experience and simultaneously meet the battery lifetime requirements. BUQS recognizes dependence of user experience on both device use and battery life using an extended battery-aware quality-of-experience (QoE) model. Furthermore, BUQS learns user's behavior to predict energy demands in time and proactively rebalance energy to improve user experience. Experimental results show that our approach delivers about 30% higher QoE than the state-of-the-art QoS and energy balancing approach.
Wooseok Lee, Reena Panda, Dam Sunwoo, José A. Joao, Andreas Gerstlauer, Lizy Kurian John
ASP-DAC4
2018 Accelerating Synchronization in Graph Analytics Using Moving Compute to Data Model on Tilera TILE-Gx72
abstract
The shared memory cache coherence paradigm is prevalent in modern multicores. However, as the number of cores increases, synchronization between threads limits performance scaling. Hardware-based core-to-core explicit messaging has been incorporated as an auxiliary communication capability to the shared memory cache coherence paradigm in the Tilera TILE-Gx72 multicore. We propose to utilize the auxiliary explicit messaging capability to build a moving computation to data model that accelerates synchronization using fine-grain serialization of critical code regions at dedicated cores. The proposed communication model exploits data locality and improves performance over both spin-lock and atomic instruction based synchronization methods for a set of parallelized graph analytic benchmarks executing on real world graphs. Experimental results show an average 34% better performance over spin-locks, and 15% over atomic instructions at 64 cores setup on TILE-Gx72.
Halit Dogan, Masab Ahmad, José A. Joao, Omer Khan
ICCD3
2018 Morphable Counters: Enabling Compact Integrity Trees For Low-Overhead Secure Memories
abstract
Securing off-chip main memory is essential for protection from adversaries with physical access to systems. However, current secure-memory designs incur considerable performance overheads - a major cause being the multiple memory accesses required for traversing an integrity-tree, that provides protection against man-in-the-middle attacks or replay attacks. In this paper, we provide a scalable solution to this problem by proposing a compact integrity tree design that requires fewer memory accesses for its traversal. We enable this by proposing new storage-efficient representations for the counters used for encryption and integrity-tree in secure memories. Our Morphable Counters are more cacheable on-chip, as they provide more counters per cacheline than existing split counters. Additionally, they incur lower overheads due to counter-overflows, by dynamically switching between counter representations based on usage pattern. We show that using Morphable Counters enables a 128-ary integrity-tree, that can improve performance by 6.3% on average (up to 28.3%) and reduce system energy-delay product by 8.8% on average, compared to an aggressive baseline using split counters with a 64-ary integrity-tree. These benefits come without any additional storage or reduction in security and are derived from our compact counter representation, that reduces the integrity-tree size for a 16GB memory from 4MB in the baseline to 1MB. Compared to recently proposed VAULT, our design provides a speedup of 13.5% on average (up to 47.4%).
Gururaj Saileshwar, Prashant J. Nair, Prakash Ramrakhyani, Wendy Elsasser, José A. Joao, Moinuddin K. Qureshi
MICRO5
2013 Utility-based acceleration of multithreaded applications on asymmetric CMPs
abstract
Asymmetric Chip Multiprocessors (ACMPs) are becoming a reality. ACMPs can speed up parallel applications if they can identify and accelerate code segments that are critical for performance. Proposals already exist for using coarse-grained thread scheduling and fine-grained bottleneck acceleration. Unfortunately, there have been no proposals offered thus far to decide which code segments to accelerate in cases where both coarse-grained thread scheduling and fine-grained bottleneck acceleration could have value. This paper proposes Utility-Based Acceleration of Multithreaded Applications on Asymmetric CMPs (UBA), a cooperative software/hardware mechanism for identifying and accelerating the most likely critical code segments from a set of multithreaded applications running on an ACMP. The key idea is a new Utility of Acceleration metric that quantifies the performance benefit of accelerating a bottleneck or a thread by taking into account both the criticality and the expected speedup. UBA outperforms the best of two state-of-the-art mechanisms by 11% for single application workloads and by 7% for two-application workloads on an ACMP with 52 small cores and 3 large cores.
José A. Joao, M. Aater Suleman, Onur Mutlu, Yale N. Patt
ISCA1
2012 Bottleneck identification and scheduling in multithreaded applications
abstract
Performance of multithreaded applications is limited by a variety of bottlenecks, e.g. critical sections, barriers and slow pipeline stages. These bottlenecks serialize execution, waste valuable execution cycles, and limit scalability of applications. This paper proposes Bottleneck Identification and Scheduling in Multithreaded Applications (BIS), a cooperative software-hardware mechanism to identify and accelerate the most critical bottlenecks. BIS identifies which bottlenecks are likely to reduce performance by measuring the number of cycles threads have to wait for each bottleneck, and accelerates those bottlenecks using one or more fast cores on an Asymmetric Chip Multi-Processor (ACMP). Unlike previous work that targets specific bottlenecks, BIS can identify and accelerate bottlenecks regardless of their type. We compare BIS to four previous approaches and show that it outperforms the best of them by 15% on average. BIS' performance improvement increases as the number of cores and the number of fast cores in the system increase.
José A. Joao, M. Aater Suleman, Onur Mutlu, Yale N. Patt
ASPLOS1
2011 Parallel application memory scheduling
abstract
A primary use of chip-multiprocessor (CMP) systems is to speed up a single application by exploiting thread-level parallelism. In such systems, threads may slow each other down by issuing memory requests that interfere in the shared memory subsystem. This inter-thread memory system interference can significantly degrade parallel application performance. Better memory request scheduling may mitigate such performance degradation. However, previously proposed memory scheduling algorithms for CMPs are designed for multi-programmed workloads where each core runs an independent application, and thus do not take into account the inter-dependent nature of threads in a parallel application.
Eiman Ebrahimi, Rustam Miftakhutdinov, Chris Fallin, Chang Joo Lee, José A. Joao, Onur Mutlu, Yale N. Patt
MICRO5
2010 Data marshaling for multi-core architectures
abstract
Previous research has shown that Staged Execution (SE), i.e., dividing a program into segments and executing each segment at the core that has the data and/or functionality to best run that segment, can improve performance and save power. However, SE's benefit is limited because most segments access inter-segment data, i.e., data generated by the previous segment. When consecutive segments run on different cores, accesses to inter-segment data incur cache misses, thereby reducing performance. This paper proposes Data Marshaling (DM), a new technique to eliminate cache misses to inter-segment data. DM uses profiling to identify instructions that generate inter-segment data, and adds only 96 bytes/core of storage overhead. We show that DM significantly improves the performance of two promising Staged Execution models, Accelerated Critical Sections and producer-consumer pipeline parallelism, on both homogeneous and heterogeneous multi-core systems. In both models, DM can achieve almost all of the potential of ideally eliminating cache misses to inter-segment data. DM's performance benefit increases with the number of cores.
M. Aater Suleman, Onur Mutlu, José A. Joao, Khubaib, Yale N. Patt
ISCA3
2009 Flexible reference-counting-based hardware acceleration for garbage collection
abstract
Languages featuring automatic memory management (garbage collection) are increasingly used to write all kinds of applications because they provide clear software engineering and security advantages. Unfortunately, garbage collection imposes a toll on performance and introduces pause times, making such languages less attractive for high-performance or real-time applications. Much progress has been made over the last five decades to reduce the overhead of garbage collection, but it remains significant.
José A. Joao, Onur Mutlu, Yale N. Patt
ISCA1
2009 Virtual Program Counter (VPC) Prediction: Very Low Cost Indirect Branch Prediction Using Conditional Branch Prediction Hardware
abstract
Indirect branches have become increasingly common in modular programs written in modern object-oriented languages and virtual-machine-based runtime systems. Unfortunately, the prediction accuracy of indirect branches has not improved as much as that of conditional branches. Furthermore, previously proposed indirect branch predictors usually require a significant amount of extra hardware storage and complexity, which makes them less attractive to implement. This paper proposes a new technique for handling indirect branches, called Virtual Program Counter (VPC) prediction. The key idea of VPC prediction is to use the existing conditional branch prediction hardware to predict indirect branch targets, avoiding the need for a separate storage structure. Our comprehensive evaluation shows that VPC prediction improves average performance by 26.7 percent and reduces average energy consumption by 19 percent compared to a commonly used branch target buffer based predictor on 12 indirect branch intensive C/C++ applications. Moreover, VPC prediction improves the average performance of the full set of object-oriented Java DaCapo applications by 21.9 percent, while reducing their average energy consumption by 22 percent. We show that VPC prediction can be used with any existing conditional branch prediction mechanism and that the accuracy of VPC prediction improves when a more accurate conditional branch predictor is used.
Hyesoon Kim, José A. Joao, Onur Mutlu, Chang Joo Lee, Yale N. Patt, Robert S. Cohn
IEEE Trans. Computers2
2008 Improving the performance of object-oriented languages with dynamic predication of indirect jumps
abstract
Indirect jump instructions are used to implement increasingly-common programming constructs such as virtual function calls, switch-case statements, jump tables, and interface calls. The performance impact of indirect jumps is likely to increase because indirect jumps with multiple targets are difficult to predict even with specialized hardware.
José A. Joao, Onur Mutlu, Hyesoon Kim, Rishi Agarwal, Yale N. Patt
ASPLOS1
2007 Profile-assisted Compiler Support for Dynamic Predication in Diverge-Merge Processors
abstract
Dynamic predication has been proposed to reduce the branch misprediction penalty due to hard-to-predict branch instructions. A proposed dynamic predication architecture, the diverge-merge processor (DMP), provides large performance improvements by dynamically predicating a large set of complex control-flow graphs that result in branch mispredictions. DMP requires significant support from a profiling compiler to determine which branch instructions and control-flow structures can be dynamically predicated. However, previous work on dynamic predication did not extensively examine the tradeoffs involved in profiling and code generation for dynamic predication architectures. This paper describes compiler support for obtaining high performance in the diverge-merge processor. We describe new profile-driven algorithms and heuristics to select branch instructions that are suitable and profitable for dynamic predication. We also develop a new profile-based analytical cost-benefit model to estimate, at compile-time, the performance benefits of the dynamic predication of different types of control-flow structures including complex hammocks and loops. Our evaluations show that DMP can provide 20.4% average performance improvement over a conventional processor on SPEC integer benchmarks with our optimized compiler algorithms, whereas the average performance improvement of the best-performing alternative simple compiler algorithm is 4.5%. We also find that, with the proposed algorithms, DMP performance is not significantly affected by the differences in profile- and run-time input data sets
Hyesoon Kim, José A. Joao, Onur Mutlu, Yale N. Patt
CGO2
2007 VPC prediction: reducing the cost of indirect branches via hardware-based dynamic devirtualization
abstract
Indirect branches have become increasingly common in modular programs written in modern object-oriented languages and virtual machine based runtime systems. Unfortunately, the prediction accuracy of indirect branches has not improved as much as that of conditional branches. Furthermore, previously proposed indirect branch predictors usually require a significant amount of extra hardware storage and complexity, which makes them less attractive to implement.
Hyesoon Kim, José A. Joao, Onur Mutlu, Chang Joo Lee, Yale N. Patt, Robert S. Cohn
ISCA2
2006 Diverge-Merge Processor (DMP): Dynamic Predicated Execution of Complex Control-Flow Graphs Based on Frequently Executed Paths
abstract
This paper proposes a new processor architecture for handling hard-to-predict branches, the diverge-merge processor (DMP). The goal of this paradigm is to eliminate branch mispredictions due to hard-to-predict dynamic branches by dynamically predicating them without requiring ISA support for predicate registers and predicated instructions. To achieve this without incurring large hardware cost and complexity, the compiler provides control-flow information by hints and the processor dynamically predicates instructions only on frequently executed program paths. The key insight behind DMP is that most control-flow graphs look and behave like simple hammock (if-else) structures when only frequently executed paths in the graphs are considered. Therefore, DMP can dynamically predicate a much larger set of branches than simple hammock branches. Our evaluations show that DMP out performs a baseline processor with an aggressive branch predictor by 19.3% on average over SPEC integer 95 and 2000 benchmarks, through a reduction of 38% in pipeline flushes due to branch mispredictions, while consuming 9.0% less energy. We also compare DMP with previously proposed predication and dual-path/multipath execution paradigms in terms of performance, complexity, and energy consumption, and find that DMP is the highest performance and also the most energy-efficient design
Hyesoon Kim, José A. Joao, Onur Mutlu, Yale N. Patt
MICRO2