EDBT 2026 Demo / reviewers in the wild / expert
Greg Byrd
dblp:61/10504 · also Gregory T. Byrd
· DBLP profile ↗
25ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-3647-8738ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 first-author · 3 since 2021Computer networks · 4Software engineering, systems software and programming languages · 4 · 1 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VirtShield: A Security Evaluation Framework for Virtualized and Containerized SystemsabstractVirtualization is a foundational technology in modern cloud computing. However, it is subject to security threats such as malicious co-located tenants, hypervisor vulnerabilities, and side-channel attacks. Such a threat is countered by deploying advanced and complex security solutions that have significant performance overhead.Prior work on VMs and containers has mainly evaluated basic security solutions, such as firewalls, using narrow performance metrics and synthetic models within limited evaluation frameworks. These studies often overlook advanced security modules in both user and kernel space, lack flexibility to incorporate emerging features, and fail to capture detailed system-level impacts. To address these gaps, we present VirtShield, an open-source framework for unified security testing in VMs and containers that mimics realistic cloud infrastructures. VirtShield supports advanced security modules across user and kernel space, providing rich, system-level performance metrics for comprehensive evaluation.Our evaluation shows that containers generally outperform VMs due to their lower virtualization overhead, achieving a throughput of 9.38 Gb/s compared to 1.98 Gb/s for VMs. However, VMs are comparable for kernel-space deployments, as Docker utilizes the shared kernel space Docker bridge, which can result in packet congestion. In latency-sensitive workloads, VM access latency (14.91 ms) is comparable to Docker (12.86 ms). In storage benchmarks (FIO), however, VMs outperform Docker due to the overhead of Docker’s layered, copy-on-write file system, whereas VMs leverage optimized virtual block devices with near-native I/O performance. These results highlight essential trade-offs in partitioning security workloads between user and kernel space, as well as across containerized and virtualized environments. Faiz Alam, Mohammed Mubeen Mifthak, Sahil Purohit, Md Shadab, Greg Byrd, Khaled Harfoush |
CCNC | 5 |
| 2023 | PreFlush: Lightweight Hardware Prediction Mechanism for Cache Line Flush and WritebackabstractNon-Volatile Main Memory (NVMM) technologies make it possible for applications to permanently store data in memory. To do so, they need to make sure that updates to persistent data comply with the crash consistency model, which often involves explicitly flushing a dirty cache line after a store and then waiting for the flush operation to complete using a store fence. While cache line flush and write back instructions can complete in the background, fence instructions expose the latency of flushing to the critical path of the program's execution, incurring significant overheads. If flush operations are started earlier, the penalty of fences can be significantly reduced. We propose PreFlush, a lightweight and transparent hardware mechanism that predicts when a cache line flush or write back is needed and speculatively performs the operation early. Since we speculatively perform the flush, we add hardware to handle flush misspeculation to ensure correct execution of the code without the need for any complex recovery mechanisms. Our PreFlush design is transparent to the programmer (i.e. it requires no modification on existing NVMM-enabled code). Our results show that PreFlush can improve performance by up to 25% (15.7% average) for the WHISPER NVM benchmark suite and loop-based matrix microbenchmarks. Hussein Elnawawy, James Tuck 0001, Greg Byrd |
PACT | 3 |
| 2023 | lfbench: a lock-free microbenchmark suiteabstractIn this work, we present lfbench: a microbenchmark suite intended as a one-stop shop representing all the popular lock-free data structures. Lock-free programming is very complex and so hard that there hasn’t been a generalized lockfree algorithm designed; instead, lock-free data structures are individually developed and optimized for the specific use-cases. In spite of this difficulty, lock-free programs are indispensable; OS kernel codes, popular databases, networking buffers, and so forth, all rely on lock-free data structures for the performance and scalability they provide. We attempt for the first time to bring all the popular lock-free data structures under one roof, primarily to enable development of new WW semantics needed for easy lock-free programming and help evaluate the same. Additionally, the benchmark suite can be used for:1)Performance analysis of any new S/W algorithms/ libraries developed.2)Building blocks for complex multi-threaded applications. Mahita Nagabhiru, Greg Byrd |
ISPASS | 2 |
| 2022 | CAPI-Precis: Towards a Compute-Centric Interface for Coherent Shared Memory AcceleratorsabstractEmerging shared memory accelerator interfaces promote a tighter coupling between traditional general-purpose processing cores and accelerator units through cache-coherence and shared virtual address space capabilities. However, different interface standards solving similar problems often require custom designs and optimizations depending on the adopted interface. This work introduces CAPI-Precis, an abstract layer between CAPI, a cache-coherent interface standard proposed by IBM, and the Accelerator Functional Unit (AFU). CAPI-Precis provides a Compute-Centric FIFO-based paradigm with the shared memory accelerator interface, hiding CAPI complexities and latency requirements in an abstract layer focusing on optimized, efficient, and scalable AFUs. Such a layer adapts to other shared memory interfaces, such as CCIX or CXL, with minimal overhead in area and performance while preserving the algorithm logic design. Abdullah T. Mughrabi, Greg Byrd |
FPT | 2 |
| 2021 | QPR: Quantizing PageRank with Coherent Shared Memory AcceleratorsabstractGraph algorithms often require fine-grained, random access across substantially large data structures. Previous work on FPGA-based acceleration has required significant preprocessing and restructuring to transform the memory access patterns into a streaming format that is more friendly to of fchip hardware. However, the emergence of cache-coherent shared memory interfaces, such as CAPI, allows designers to more easily work with the natural in-memory organization of the data. This paper introduces a vertex-centric shared-memory accelerator for the PageRank algorithm, optimized for high performance while effectively using coherent caching on the FPGA hardware. The proposed design achieves up to 14.9x speedups by selectively caching graph data for the accelerator while taking into account locality and reuse, compared to naively using the shared address space access and DRAM only. We also introduce PageRank Quantization, an innovative technique to represent page-ranks with 32-bit quantized fixed-point values. This approach is up to 1.5x faster than 64-bit fixed-point while keeping precision within a tolerable error margin. As a result, we maintain both the hardware scalability of fixed-point representation and the cache performance of 32-bit floating-point. Abdullah T. Mughrabi, Mohannad Ibrahim, Greg Byrd |
IPDPS | 3 |
| 2020 | Quantum Circuits for Dynamic Runtime Assertions in Quantum ComputationabstractIn this paper, we propose quantum circuits for runtime assertions, which can be used for both software debugging and error detection. Runtime assertion is challenging in quantum computing for two key reasons. First, a quantum bit (qubit) cannot be copied, which is known as the non-cloning theorem. Second, when a qubit is measured, its superposition state collapses into a classical state, losing the inherent parallel information. In this paper, we overcome these challenges with runtime computation through ancilla qubits, which are used to indirectly collect the information of the qubits of interest. We design quantum circuits to assert classical states, entanglement, and superposition states. Our experimental results show that they are effective in debugging as well as improving the success rate for various quantum algorithms on IBM Q quantum computers. Ji Liu 0007, Greg Byrd, Huiyang Zhou |
ASPLOS | 2 |
| 2019 | Diligent TLBs: a mechanism for exploiting heterogeneity in TLB miss behaviorabstractModern workloads such as graph analytics, sparse matrix multiplication, and in-memory key-value stores use very large datasets and typically have non-uniform memory access patterns which defy traditional concepts of locality. Moreover, many of these algorithms simultaneously use multiple data structures that have very distinct access patterns to the corresponding pages, leading to heterogeneity in TLB behavior. Our intuition suggests that these two factors make it important to architect a heterogeneity-aware TLB hierarchy. Hussein Elnawawy, Rangeen Basu Roy Chowdhury, Amro Awad, Greg Byrd |
ICS | 4 |
| 2019 | Programming quantum computers: a primer with IBM Q and D-Wave exercisesabstractThis tutorial provides a hands-on introduction to quantum computing. It will feature the three pillars, architectures, programming, and algorithms/applications of quantum computing. Its focus is on the applicability of problems to quantum computing from a practical point, with only the necessary foundational coverage of the physics and theoretical aspects to understand quantum computing. Simulation software will be utilized complemented by access to actual quantum computers to prototype problem solutions. This should develop a better understanding of how problems are transformed into quantum algorithms and what programming language support is best suited for a given application area. As a first of its kind, to the best of our knowledge, the tutorial includes hands-on programming experience with IBM Q and D-Wave hardware. Frank Mueller 0001, Greg Byrd, Patrick Dreher |
PPoPP | 2 |
| 2011 | A Canonical Multicore Architecture for Network RoutersabstractThere has been a significant increase in the Internet dynamics in the past decade. This has put tremendous pressure on the performance of routing protocols as they need to keep updating their routing information with every network change across the globe. With the growth of Internet, Border Gateway Protocol (BGP) has become a critical routing application. Good performance of BGP on network processors directly translates to better convergence time for route changes on the Internet, leading to reduced data loss on the network. BGP is the ubiquitous routing protocol on the Internet core, and hence analyzing its performance and exploring avenues for speeding it up can greatly help in improving the responsiveness and reliability of the Internet. In this paper, we investigate the use of multicore as the compute platform for routing protocols using BGP as a representative application. We discuss two different schemes for parallelizing BGP and analyze the performance of both serial and parallel BGP implementations on a fully configurable multicore simulation environment. Subsequently, we analyze the architectural bottlenecks in the conventional multicore systems which limit the speedup that can be achieved by software parallelism alone, and propose a canonical multicore architecture for routing protocols, which can be used for future routing processor designs. The analysis and proposed schemes in this paper would greatly help in understanding the behavior of BGP, thereby assisting in design and development of next generation network processors. Sabina Grover, Abhishek Dhanotia, Greg Byrd |
ANCS | 3 |
| 2011 | Welcome to ICCD 2011!abstractOn behalf of the organizing and program committee, we would like to welcome you to the 29thIEEE International Conference on Computer Design 2011. The International Conference on Computer Design (ICCD) encompasses a wide range of technical topics and provides an ideal environment to discuss practical and theoretical work that enables cross-pollination. The ICCD venue and program reflect this goal. This year the conference is being held at the beautiful campus of the University of Massachusetts at Amherst, United States. Georgi Gaydadjiev, Sofiène Tahar, Greg Byrd, Klaus Schneider 0001 |
ICCD | 3 |
| 2009 | Limited early value communication to improve performance of transactional memoryabstractParallel programming is receiving renewed attention with the advent of multi-core CPU architectures. The Transactional Memory (TM) paradigm has the potential to provide good speedup and make parallel programming easier to adopt. Under low contention, it has been shown that TM programs can outperform standard lock-based programs. However, under high contention, performance of TM programs can degrade. Previous work has shown that we can use either data forwarding or value prediction to improve performance under high contention. Both these techniques demand significant changes to the architecture and coherence protocol above and beyond those required by TM. Salil Mohan Pant, Greg Byrd |
ICS | 2 |
| 2008 | Exploiting producer patterns and L2 cache for timely dependence-based prefetchingabstractThis paper proposes an architecture that efficiently prefetches for loads whose effective addresses are directly dependent on previously-loaded values. This dependence-based prefetching scheme covers most frequently missed loads in programs that contain linked data structures (LDS). For timely prefetches, memory access patterns of producing loads are dynamically learned. These patterns (such as strides) are used to prefetch well ahead of the consumer load. The proposed prefetcher is placed near the processor core and targets L1 cache misses, because removing L1 cache misses has greater performance potential than removing L2 cache misses. We also examine how to capture pointers in LDS with pure hardware implementation. We find that the space requirement can be reduced, compared to previous work, if we selectively record patterns. Still, to make the prefetching scheme generally applicable, a large table is required for storing pointers. We show that storing the prefetch table in a partition of the L2 cache outperforms using the L2 cache conventionally. Chungsoo Lim, Greg Byrd |
ICCD | 2 |
| 2006 | High-throughput sketch update on a low-power stream processorabstractSketch algorithms are widely used for many networking applications, such as identifying frequent items, top-k flows, and traffic anomalies. This paper explores the implementation of the Count-Min sketch update using Indexed SRF accesses on a SIMD stream processor (Imagine). Both the sketch data structure and the packet stream are modeled as streams, and in-lane accesses to the stream register file (SRF) support concurrent updates without explicit synchronization. The 500-MHz stream processor is capable of supporting sketch update at 10 Gbps throughput for minimum-sized IP packets. This is nearly the same performance as the 1.4-GHz Intel IXP2800 (13 Gbps), using significantly less power (2.89W vs. 21W). Yu-Kuen Lai, Greg Byrd |
ANCS | 2 |
| 2006 | Stream-Based Implementation of Hash Functions for Multi-Gigabit Message Authentication CodesabstractStream processing architectures have been proposed as efficient and flexible platforms for network packet processing. As part of an investigation into stream-based network processors, we have implemented MMH, a family of almost-universal hash functions for message authentication, on a SIMD stream processor (Imagine). The hash computation over an entire packet is a good fit for the stream programming model, with an abundance of producer-consumer locality: hash values are computed and stored in the stream register file (SRF), then used for calculating new hash values repeatedly. By using eight VLIW clusters, the construction is performed in a multi-SIMD fashion, achieving multi-gigabit-per-second throughput with a collision probability on the order of 2~120 Yu-Kuen Lai, Greg Byrd |
PDCAT | 2 |
| 2005 | Trust-Based Secure Workflow Path Construction
Mine Altunay, Douglas E. Brown, Greg Byrd, Ralph A. Dean |
ICSOC | 3 |
| 2005 | Evaluation of Mutual Trust during MatchmakingabstractThe authors introduced a new service discovery and matchmaking architecture, layered on top of Globus MDS3, that integrates mutual trust evaluations into the matchmaking process. The architecture adopts a symmetric approach, and checks trust policies of both grid users and resources without requiring policy disclosures. This approach eliminates run-time security failures arising from incompatible user/resource pairs, seamlessly integrates user-side authorization tools with the matchmaking process, and protects naive grid users by allowing a security principal to define policies that control the list of discoverable resources. Mine Altunay, Douglas E. Brown, Greg Byrd, Ralph A. Dean |
Peer-to-Peer Computing | 3 |
| 2003 | Slipstream Execution Mode for CMP-Based MultiprocessorsabstractScalability of applications on distributed shared-memory (DSM) multiprocessors is limited by communication overheads. At some point, using more processors to increase parallelism yields diminishing returns or even degrades performance. When increasing concurrency is futile, we propose an additional mode of execution, called slipstream mode, that instead enlists extra processors to assist parallel tasks by reducing perceived overheads. We consider DSM multiprocessors built from dual-processor chip multiprocessor (CMP) nodes with shared L2 cache. A task is allocated on one processor of each CMP node. The other processor of each node executes a reduced version of the same task. The reduced version skips shared-memory stores and synchronization, running ahead of the true task. Even with the skipped operations, the reduced task makes accurate forward progress and generates an accurate reference stream, because branches and addresses depend primarily on private data. Slipstream execution mode yields two benefits. First, the reduced task prefetches data on behalf of the true task. Second, reduced tasks provide a detailed picture of future reference behavior, enabling a number of optimizations aimed at accelerating coherence events, e.g., self-invalidation. For multiprocessor systems with up to 16 CMP nodes, slipstream mode outperforms running one or two conventional tasks per CMP in 7 out of 9 parallel scientific benchmarks. Slipstream mode is 12-19% faster with prefetching only and up to 29% faster with self-invalidation enabled. Khaled Z. Ibrahim, Greg Byrd, Eric Rotenberg |
HPCA | 2 |
| 2003 | Design and implementation of Acceptance Monitor for building intrusion tolerant systemsabstractAbstract Intrusion detection research has so far concentrated on techniques that effectively identify the malicious behaviors. No assurance can be assumed once the system is compromised. Intrusion tolerance, however, focuses on providing minimal level of services, even when some components have been partially compromised. The challenges here are how to take advantage of fault tolerant techniques in the intrusion tolerant system context and how to deal with possible unknown attacks and compromised components so as to continue providing the service. This paper presents our work on applying one important fault tolerance technique, acceptance testing, for building scalable intrusion tolerant systems. First, we propose a general methodology for designing acceptance testing. An Acceptance Monitor architecture is proposed to apply various tests for detecting the compromises based on the impact of the attacks. Second, we make a comprehensive vulnerability analysis on typical commercial‐off‐the‐shelf (COTS) Web servers. Various acceptance testing modules are implemented to show the effectiveness of the proposed approach. By utilizing the fault tolerance techniques on intrusion tolerance system, we provide a mechanism for building reliable distributed services that are more resistant to both known and unknown attacks. Copyright © 2003 John Wiley & Sons, Ltd. Feiyi Wang, Greg Byrd |
Softw. Pract. Exp. | 3 |
| 2001 | Design and implementation of acceptance monitor for building scalable intrusion tolerant systemabstractIntrusion detection research has so far mostly concentrated on techniques that effectively identify malicious behavior. No assurance can be assumed once the system is compromised. Intrusion tolerance, on the other hand, focuses on providing minimal level of services even when some components have been partially compromised. The challenges here are how to take advantage of fault tolerant techniques in the intrusion tolerant system context and how to deal with possible unknown attacks and compromised components so as to continue providing the service. This paper presents our work on applying one important fault tolerance technique, acceptance testing, for building scalable intrusion tolerant systems. First, we propose a general methodology for designing acceptance tests. An acceptance monitor architecture is proposed to apply various tests for detecting compromises based on the impact of the attacks. Second, we make a comprehensive vulnerability analysis on typical commercial-off-the-shelf (COTS) Web servers. Various acceptance testing modules are implemented to show the effectiveness of the proposed approach. By utilizing the fault tolerance techniques on intrusion tolerance system, we provide a mechanism for building reliable distributed services that are more resistant to both known and unknown attacks. Feiyi Wang, Greg Byrd |
ICCCN | 3 |
| 2001 | On the Exploitation of Value Predication and Producer Identification to Reduce Barrier Synchronization TimeabstractBarrier synchronization is a source of inefficiency in many parallel programs, due to the association of many producer-consumer relations in with one synchronization variable. This inefficiency may consume a significant percentage of total execution time, especially as we increase the degree of parallelism while maintaining the problem size. Barrier synchronization wait time can be hidden by speculatively executing instructions after the barrier. The speculative execution must not violate the dependencies imposed by the program. Dependency violation causes rollback, incurring a penalty that may exceed the benefit of speculation. In this work, we investigate how to reduce the probability of rollback through the use of two different techniques: value prediction and producer identification. The first technique tries to break the dependency between the running processes. The second technique tries to respect only true dependencies by transforming the barrier synchronization into per-variable flags. Simulation results using scientific benchmarks mostly SPLASH-2, indicate that producer identification promises a greater potential reduction in synchronization time, close to actual dependency, and maintains rollback percentage below 10% for most benchmarks. Khaled Z. Ibrahim, Greg Byrd |
IPDPS | 2 |
| 2001 | Practical Experiences with ATM Encryption
Greg Byrd, Nathan Hillery, Jim Symon |
NDSS | 1 |
| 1999 | Producer-consumer communication in distributed shared memory multiprocessorsabstractThe shared memory abstraction supported by hardware based distributed shared memory (DSM) multiprocessors is an inherently consumer driven means of communication. When a process requires data, it retrieves them from the global shared memory. In distributed cache coherent systems, the data may reside in a remote memory module or in the producer's cache. Producer initiated mechanisms reduce communication latency by sending data to the consumer as soon as they are produced. We classify producer initiated mechanisms as implicit or explicit, according to whether the producer must know the identity of the consumer when data are transmitted. Explicit schemes include data forwarding and message passing. Implicit schemes include update based coherence, selective updates, and cache based locks. Several of these mechanisms are evaluated for performance and sensitivity to network parameters, using a common simulated architecture and a set of application kernel benchmarks. StreamLine, a cache based message passing mechanism, provides the best performance on the benchmarks with regular communication patterns. Forwarding write and cache based locks are also among the best performing producer initiated mechanisms. Consumer initiated prefetch, however, has good average performance and is the least expensive to implement. Greg Byrd, Michael J. Flynn |
Proc. IEEE | 1 |
| 1995 | Design of a key agile cryptographic system for OC-12c rate ATMabstractThe paper describes an experimental key agile cryptographic system under design at MCNC. The system is compatible with ATM local- and wide-area networks. The system establishes and manages secure connections between hosts in a manner which is transparent to the end users and compatible with existing public network standards. A Cryptographic Unit supports hardware encryption and decryption at the ATM protocol layer. The system is SONET compatible and operates full duplex at the OC-12c rate (622 Mbps). Separate encryption keys are negotiated for each secure connection. Each Cryptographic Unit can manage more than 65,000 active secure connections. The Cryptographic Unit can be connected either in a security gateway mode referred to as a 'bump-in-the-fiber' or as a direct ATM host interface. Authentication and access control are implemented through a certificate-based system. The current status of the system is that hardware and software detail designs have been completed. An early version of the key management software has been completed and demonstrated. Hardware fabrication and systems integration are expected to take place over the next several months. Once completed the proof-of concept system will be used to explore issues of privacy, access control and authentication in relation to communications over emerging public networks.> Daniel S. Stevenson, Nathan Hillery, Greg Byrd, Fengmin Gong, Dan Winkelstein |
NDSS | 3 |
| 1991 | Streamline: Cache-Based Message Passing in Scalable Multiprocessors
Greg Byrd, Bruce Delagi |
ICPP (1) | 1 |
| 1989 | Multicast Communication in Multiprocessor Systems
Greg Byrd, Nakul P. Saraiya, Bruce Delagi |
ICPP (1) | 1 |