EDBT 2026 Demo / reviewers in the wild / expert
Clay Hughes
dblp:64/1853 · also Clayton Hughes
· DBLP profile ↗
10ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-9071-2487ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ASLink: Modeling Multi-GPU Execution in Accel-SimabstractGraphical processing units (GPUs) are widely used in numerous modern application domains, including modeling and simulation, machine learning, and data analytics. Many applications such as recommendation models and graph neural networks benefit from the use of multiple GPUs to scale up the size of the workload and increase throughput. While current open-sourced GPU architectural simulators can model multiGPU workloads, doing so remains inefficient and challenging, limiting their broader applicability across various application domains. On the other hand, simulation tools used by the industry are often closed-source, thus hindering efforts to democratize architectural research. This paper proposes an open-source simulator design, ASLink, that extends Accel-sim to support multiGPU configurations. We highlight the limitations of popular state-of-the-art GPU architecture simulators and propose mechanisms to improve user experience and modeling fidelity in multi-GPU systems. Finally, we validate our proposed infrastructure against kernels representative of real world workloads. Christin Bose, Cesar Avalos, Junrui Pan 0001, Yechen Liu, Mahmoud Khairy, Clay Hughes, Timothy G. Rogers |
ISPASS | 6 |
| 2024 | SEFsim: A Statistically-Guided Fast DRAM SimulatorabstractIn academia and industry, computer architects rely heavily on performance models for design space exploration. However, performance models are now experiencing longer simulation times due to the increasing design complexity of modern computing systems. DDR memory, a critical component in a computing system, requires an accurate performance model to properly evaluate the instructions per cycle (IPC). However, a detailed DRAM simulator models each DDR event and, therefore, contributes a considerable simulation time. This paper proposes Satistically-guided Epoch-evolving Fixed-latency Simulator (SEFsim), an approximate and fast DRAM simulation model, to significantly improve the simulation speed. The key design principle of SEFsim is to statistically capture the performance model of DRAM using a large number of patterns, enabling the model to accurately predict the latency and behavior of new workloads. Based on our evaluation using a detailed memory model and 10 workloads, SEFsim captures the original model with 96.16% accuracy while speeding up the simulation by 10.3X and 8.25 % in the standalone and full system evaluations, respectively. Debpratim Adak, Hyokeun Lee, Ben Feinberg, Gwendolyn Voskuilen, Clay Hughes, Huiyang Zhou, Amro Awad |
ISPASS | 5 |
| 2022 | Minerva: Rethinking Secure Architectures for the Era of Fabric-Attached Memory ArchitecturesabstractFabric-attached memory (FAM) is proposed to enable the seamless integration of directly accessible memory modules attached to the shared system fabric, which will provide future systems with flexible memory integration options, mitigate underutilization, and facilitate data sharing. Recently proposed interconnects, such as Gen-Z and Compute Express Link (CXL), define security, correctness, and performance requirements of fabric-attached devices, including memory. These initiatives are supported by most major system and processor vendors, bringing widespread adoption of FAM-enabled systems one step closer to reality and security concerns to the forefront. This paper discusses the challenges for adapting secure memory implementations to FAM-enabled systems for the first time in literature. Specifically, we observe that handling the security metadata used to protect fabric-attached memories needs to be done deliberately to eliminate unintentional integrity check failures and/or security vulnerabilities, caused by an inconsistent view of the shared security metadata across nodes. Our scheme, Minerva, elegantly adapts secure memory implementations to support FAM-enabled systems with negligible performance over-heads (3.8% of an ideal scheme), compared to the performance overhead (99.5% of an ideal scheme) for a scheme that uses conventional invalidation-based cache coherence to ensure the consistency of security metadata across nodes. Mazen Al-Wadi, Rujia Wang, David Mohaisen, Clay Hughes, Simon D. Hammond, Amro Awad |
IPDPS | 4 |
| 2022 | "Smarter" NICs for faster molecular dynamics: a case studyabstractThis work evaluates the benefits of using a “smart” network interface card (SmartNIC) as a compute accelerator for the example of the MiniMD molecular dynamics proxy application. The accelerator is NVIDIA's BlueField-2 card, which includes an 8-core Arm processor along with a small amount of DRAM and storage. We test the networking and data movement performance of these cards compared to a standard Intel server host using microbenchmarks and MiniMD. In MiniMD, we identify two distinct classes of computation, namely core computation and maintenance computation, which are executed in sequence. We restructure the algorithm and code to weaken this dependence and increase task parallelism, thereby making it possible to increase utilization of the BlueField-2 concurrently with the host. We evaluate our implementation on a cluster consisting of 16 dual-socket Intel Broadwell host nodes with one BlueField-2 per host-node. Our results show that while the overall compute performance of BlueField-2 is limited, using them with a modified MiniMD algorithm allows for up to 20% speedup over the host CPU baseline with no loss in simulation accuracy. Sara Karamati, Clay Hughes, Karl S. Hemmert, Ryan E. Grant, Whit Schonbein, Scott Levy, Thomas M. Conte, Jeffrey Young 0001, Richard W. Vuduc |
IPDPS | 2 |
| 2021 | Stealth-Persist: Architectural Support for Persistent Applications in Hybrid Memory SystemsabstractNon-volatile memories (NVMs) have the characteristics of both traditional storage systems (persistent) and traditional memory systems (byte-addressable). However, they suffer from high write latency and have a limited write endurance. Researchers have proposed hybrid memory systems that combine DRAM and NVM, utilizing the lower latency of the DRAM to hide some of the shortcomings of the NVM - improving system's performance by caching resident NVM data in the DRAM. However, this can nullify the persistency of the cached pages, leading to a question of trade-offs in terms of performance and reliability. In this paper, we propose Stealth-Persist, a novel architecture support feature that allows applications that need persistence to run in the DRAM while maintaining the persistency features provided by the NVM. Stealth-Persist creates the illusion of a persistent memory for the application to use, while utilizing the DRAM for performance optimizations. Our experimental results show that Stealth-Persist improves the performance by 42.02% for persistent applications. Mazen Al-Wadi, Vamsee Reddy Kommareddy, Clay Hughes, Simon D. Hammond, Amro Awad |
HPCA | 3 |
| 2021 | DeACT: Architecture-Aware Virtual Memory Support for Fabric Attached Memory SystemsabstractThe exponential growth of data has driven technology providers to develop new protocols, such as cache coherent interconnects and memory semantic fabrics, to help users and facilities leverage advances in memory technologies to satisfy these growing memory and storage demands. Using these new protocols, fabric-attached memories (FAM) can be directly attached to a system interconnect and be easily integrated with a variety of processing elements (PEs). Moreover, systems that support FAM can be smoothly upgraded and allow multiple PEs to share the FAM memory pools using well-defined protocols. The sharing of FAM between PEs allows efficient data sharing, improves memory utilization, reduces cost by allowing flexible integration of different PEs and memory modules from several vendors, and makes it easier to upgrade the system. One promising use-case for FAMs is in High-Performance Compute (HPC) systems, where the underutilization of memory is a major challenge. However, adopting FAMs in HPC systems brings new challenges. In addition to cost, flexibility, and efficiency, one particular problem that requires rethinking is virtual memory support for security and performance. To address these challenges, this paper presents decoupled access control and address translation (DeACT), a novel virtual memory implementation that supports HPC systems equipped with FAM. Compared to the state-of-the-art two-level translation approach, DeACT achieves speedup of up to 4.59× (1.8× on average) without compromising security. Vamsee Reddy Kommareddy, Clay Hughes, Simon D. Hammond, Amro Awad |
HPCA | 2 |
| 2018 | Osiris: A Low-Cost Mechanism to Enable Restoration of Secure Non-Volatile MemoriesabstractWith Non-Volatile Memories (NVMs) beginning to enter the mainstream computing market, it is time to consider how to secure NVM-equipped computing systems. Recent Meltdown and Spectre attacks are evidence that security must be intrinsic to computing systems and not added as an afterthought. Processor vendors are taking the first steps and are beginning to build security primitives into commodity processors. One security primitive that is associated with the use of emerging NVMs is memory encryption. Memory encryption, while necessary, is very challenging when used with NVMs because it exacerbates the write endurance problem. Secure architectures use cryptographic metadata that must be persisted and restored to allow secure recovery of data in the event of power-loss. Specifically, encryption counters must be persistent to enable secure and functional recovery of an interrupted system. However, the cost of ensuring and maintaining persistence for these counters can be significant. In this paper, we propose a novel scheme to maintain encryption counters without the need for frequent updates. Our new memory controller design, Osiris, repurposes memory Error-Correction Codes (ECCs) to enable fast restoration and recovery of encryption counters. To evaluate our design, we use Gem5 to run eight memory-intensive workloads selected from SPEC2006 and U.S. Department of Energy (DoE) proxy applications. Compared to a write-through counter-cache scheme, on average, Osiris can reduce 48.7% of the memory writes (increase lifetime by 1.95x), and reduce the performance overhead from 51.5% (for write-through) to only 5.8%. Furthermore, without the need for backup battery or extra power-supply hold-up time, Osiris performs better than a battery-backed write-back (5.8% vs. 6.6% overhead) and has less write-traffic (2.6% vs. 5.9% overhead). Mao Ye 0008, Clay Hughes, Amro Awad |
MICRO | 2 |
| 2011 | Optimizing throughput/power trade-offs in hardware transactional memory using DVFS and intelligent schedulingabstractPower has emerged as a first-order design constraint in modern processors and has energized microarchitecture researchers to produce a growing number of power optimization proposals. Almost in tandem with the move toward more energy-efficient designs, architects have been increasing the number of processing elements (PEs) on a single chip and promoting the concept of running multithreaded workloads. Nevertheless, software is still lagging behind and is often unable to exploit these additional resources -- giving rise to transactional memory. Transactional memory is a promising programming abstraction that makes it easier for programmers to exploit the resources available in many- core processor systems by removing some of the complexity associated with traditional lock-based programming. This paper proposes new techniques to merge the power and transactional memory domains. Clay Hughes, Tao Li 0006 |
ICS | 1 |
| 2009 | TransMetric: architecture independent workload characterization for transactional memory benchmarksabstractTransactional memory (TM) has emerged as a parallel programming paradigm for multi-core processors yet there is no standardized set of metrics with which to describe their behavior. In this work, we propose a set of transaction-oriented workload characteristics that can accurately capture the behavior of transactional memory programs. We apply principle component analysis and clustering algorithms to analyze the proposed transactional workload characteristics and show that these characteristics are architecturally independent James Poe, Clay Hughes, Tao Li 0006 |
ICS | 2 |
| 2009 | TransPlant: A parameterized methodology for generating transactional memory workloadsabstractTransactional memory provides a means to bridge the discrepancy between programmer productivity and the difficulty in exploiting thread-level parallelism gains offered by emerging chip multiprocessors. Because the hardware has outpaced the software, there are very few modern multithreaded benchmarks available and even fewer for transactional memory researchers. This hurdle must be overcome for transactional memory research to mature and to gain widespread acceptance. Currently, for performance evaluations, most researchers rely on manually converted lock-based multithreaded workloads or the small group of programs written explicitly for transactional memory. Using converted benchmarks is problematic because they have been tuned so well that they may not be representative of how a programmer will actually use transactional memory. Hand coding stressor benchmarks is unattractive because it is tedious and time consuming. A new parameterized methodology that can automatically generate a program based on the desired high-level program characteristics benefits the transactional memory community. In this work, we propose techniques to generate parameterized transactional memory benchmarks based on a feature set, decoupled from the underlying transactional model. Using principle component analysis, clustering, and raw transactional performance metrics, we show that TransPlant can generate benchmarks with features that lie outside the boundary occupied by these traditional benchmarks. We also show how TransPlant can mimic the behavior of SPLASH-2 and STAMP transactional memory workloads. The program generation methods proposed here will help transactional memory architects select a robust set of programs for quick design evaluations. James Poe, Clay Hughes, Tao Li 0006 |
MASCOTS | 2 |