EDBT 2026 Demo / reviewers in the wild / expert
Yuval Tamir
dblp:02/4293
· DBLP profile ↗
30ranked-venue papers
10as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 9 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 1 since 2021Security and privacy · 4Computer networks · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Operating systems · 60% Concurrent programming · 26% Program analysis · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
11 papers |
Distributed systems · 44% Memory systems · 33% Hardware reliability and fault tolerance · 11% | |
| Computer networks
2 papers |
Network management and operations · 100% |
Topics — the 30 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Operating systems › resource management
memory management |
0.9 | 1 | 2025 | CortenMM: Efficient Memory Management with Strong Correctness Guarantees · SOSP 2025 |
Operating systems › resource management › memory management
page table management |
0.9 | 1 | 2025 | CortenMM: Efficient Memory Management with Strong Correctness Guarantees · SOSP 2025 |
Memory systems
virtual memory management |
0.9 | 1 | 2025 | CortenMM: Efficient Memory Management with Strong Correctness Guarantees · SOSP 2025 |
Network management and operations
configuration verification |
0.6 | 2 | 2021 | Campion: debugging router configuration differences · SIGCOMM 2021 Finding Network Misconfigurations by Automatic Template Inference · NSDI 2020 |
Distributed systems
fault tolerance |
0.6 | 1 | 2022 | RRC: Responsive Replicated Containers · USENIX ATC 2022 |
Distributed systems
replication |
0.6 | 1 | 2022 | RRC: Responsive Replicated Containers · USENIX ATC 2022 |
Network management and operations › configuration verification
misconfiguration detection |
0.4 | 1 | 2020 | Finding Network Misconfigurations by Automatic Template Inference · NSDI 2020 |
Network management and operations
network configuration |
0.4 | 1 | 2020 | Finding Network Misconfigurations by Automatic Template Inference · NSDI 2020 |
Concurrent programming
concurrency bugs |
0.4 | 1 | 2019 | PUSh: Data Race Detection Based on Hardware-Supported Prevention of Unintended Sharing · MICRO 2019 |
Concurrent programming › concurrency bug detection
data race detection |
0.4 | 1 | 2019 | PUSh: Data Race Detection Based on Hardware-Supported Prevention of Unintended Sharing · MICRO 2019 |
Program analysis
dynamic analysis |
0.4 | 1 | 2019 | PUSh: Data Race Detection Based on Hardware-Supported Prevention of Unintended Sharing · MICRO 2019 |
Hardware reliability and fault tolerance
fault injection |
0.2 | 1 | 2015 | Fault Injection in Virtualized Systems - Challenges and Applications · IEEE Trans. Dependable Secur. Comput. 2015 |
Cloud and datacenter computing
virtualization |
0.2 | 1 | 2015 | Fault Injection in Virtualized Systems - Challenges and Applications · IEEE Trans. Dependable Secur. Comput. 2015 |
Network management and operations › fault management
fault diagnosis |
0.1 | 1 | 2021 | Campion: debugging router configuration differences · SIGCOMM 2021 |
Hardware reliability and fault tolerance
dependability analysis |
0.1 | 1 | 2015 | Fault Injection in Virtualized Systems - Challenges and Applications · IEEE Trans. Dependable Secur. Comput. 2015 |
Interconnection networks and networks-on-chip › switching network
multistage interconnection network |
0.0 | 4 | 1993 | Symmetric Crossbar Arbiters for VLSI Communication Switches · IEEE Trans. Parallel Distributed Syst. 1993 Dynamically-Allocated Multi-Queue Buffers for VLSI Communication Switches · IEEE Trans. Computers 1992 Support for High-Priority Traffic in VLSI Communication Switches · RTSS 1988 |
Interconnection networks and networks-on-chip › switch architecture
buffer design |
0.0 | 2 | 1992 | Dynamically-Allocated Multi-Queue Buffers for VLSI Communication Switches · IEEE Trans. Computers 1992 Support for High-Priority Traffic in VLSI Communication Switches · RTSS 1988 |
Interconnection networks and networks-on-chip
switch architecture |
0.0 | 2 | 1992 | Dynamically-Allocated Multi-Queue Buffers for VLSI Communication Switches · IEEE Trans. Computers 1992 Support for High-Priority Traffic in VLSI Communication Switches · RTSS 1988 |
Hardware reliability and fault tolerance › error detection
concurrent error detection |
0.0 | 2 | 1990 | High-Performance Fault-Tolerant VLSI Systems Using Micro Rollback · IEEE Trans. Computers 1990 Design and Application of Self-Testing Comparators Implemented with MOS PLA's · IEEE Trans. Computers 1984 |
Integrated circuit design › VLSI design
VLSI processor design |
0.0 | 1 | 1990 | High-Performance Fault-Tolerant VLSI Systems Using Micro Rollback · IEEE Trans. Computers 1990 |
Interconnection networks and networks-on-chip
interconnection networks |
0.0 | 1 | 1988 | Support for High-Priority Traffic in VLSI Communication Switches · RTSS 1988 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 1987 | A block-and-actions generator as an alternative to a simulator for collecting architecture measurements · PLDI 1987 |
Performance modeling and evaluation › simulation
simulation-based evaluation |
0.0 | 1 | 1993 | Symmetric Crossbar Arbiters for VLSI Communication Switches · IEEE Trans. Parallel Distributed Syst. 1993 |
Performance modeling and evaluation › simulation › communication system simulation
network simulation |
0.0 | 1 | 1992 | Dynamically-Allocated Multi-Queue Buffers for VLSI Communication Switches · IEEE Trans. Computers 1992 |
Processor architecture and microarchitecture › register file
register file management |
0.0 | 1 | 1983 | Strategies for Managing the Register File in RISC · IEEE Trans. Computers 1983 |
Processor architecture and microarchitecture › register file
register window |
0.0 | 1 | 1983 | Strategies for Managing the Register File in RISC · IEEE Trans. Computers 1983 |
Processor architecture and microarchitecture › instruction set architecture › RISC
RISC processor |
0.0 | 1 | 1983 | Strategies for Managing the Register File in RISC · IEEE Trans. Computers 1983 |
Parallel and multicore computing
multicomputer |
0.0 | 1 | 1988 | High-Performance Multi-Queue Buffers for VLSI Communication Switches · ISCA 1988 |
Interconnection networks and networks-on-chip › switching
virtual cut-through switching |
0.0 | 1 | 1988 | High-Performance Multi-Queue Buffers for VLSI Communication Switches · ISCA 1988 |
Integrated circuit design
VLSI design |
0.0 | 1 | 1984 | Design and Application of Self-Testing Comparators Implemented with MOS PLA's · IEEE Trans. Computers 1984 |
Methods — techniques the papers use, named apart from their topics
concurrency bug analysis · 1.7structural equivalence checking · 0.5modular analysis · 0.5template inference · 0.4intended sharing specification · 0.4hardware-supported prevention · 0.4hypervisor-based injection · 0.2fault injection · 0.2simulation · 0.0network simulation · 0.0circuit simulation · 0.0queueing simulation · 0.0hardware rollback mechanism · 0.0microarchitecture design · 0.0assembly-to-assembly translation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CortenMM: Efficient Memory Management with Strong Correctness GuaranteesabstractModern memory management systems suffer from poor performance and subtle concurrency bugs, slowing down applications while introducing security vulnerabilities. We observe that both issues stem from the conventional design of memory management systems with two levels of abstraction: a software-level abstraction (e.g., VMA trees in Linux) and a hardware-level abstraction (typically, page tables). This design increases portability but requires correctly and efficiently synchronizing two drastically different and complex data structures, which is generally challenging. Junyang Zhang 0003, Xiangcan Xu, Yonghao Zou, Xinyi Wan 0001, Siyuan Wang 0026, Di Wang 0017, Hao Chen 0023, Lin Huang 0005, Shoumeng Yan, Yuval Tamir, Yingwei Luo, Xiaolin Wang 0001, Huashan Yu, Zhenlin Wang 0003, Hongliang Tian, Diyu Zhou |
SOSP | 13 |
| 2022 | RRC: Responsive Replicated Containers
Diyu Zhou, Yuval Tamir |
USENIX ATC | 2 |
| 2021 | Campion: debugging router configuration differencesabstractWe present a new approach for debugging two router configurations that are intended to be behaviorally equivalent. Existing router verification techniques cannot identify all differences or localize those differences to relevant configuration lines. Our approach addresses these limitations through a _modular_ analysis, which separately analyzes pairs of corresponding configuration components. It handles all router components that affect routing and forwarding, including configuration for BGP, OSPF, static routes, route maps and ACLs. Further, for many configuration components our modular approach enables simple _structural equivalence_ checks to be used without additional loss of precision versus modular semantic checks, aiding both efficiency and error localization. We implemented this approach in the tool Campion and applied it to debugging pairs of backup routers from different manufacturers and validating replacement of critical routers. Campion analyzed 30 proposed router replacements in a production cloud network and proactively detected four configuration bugs, including a route reflector bug that could have caused a severe outage. Campion also found multiple differences between backup routers from different vendors in a university network. These were undetected for three years, and depended on subtle semantic differences that the operators said they were "highly unlikely" to detect by "just eyeballing the configs." Alan Tang, Siva Kesava Reddy K., Ryan Beckett, Ennan Zhai, Matt Brown, Todd D. Millstein, Yuval Tamir, George Varghese |
SIGCOMM | 7 |
| 2020 | Fault-Tolerant Containers Using NiLiConabstractMany services deployed in the cloud require high reliability and must thus survive machine failures. Providing such fault tolerance transparently, without requiring application modifications, has motivated extensive research on replicating virtual machines (VMs). Cloud computing typically relies on VMs or containers to provide an isolation and multitenancy layer. Containers have advantages over VMs in smaller size, faster startup, and avoiding the need to manage updates of multiple VMs. This paper reports on the design, implementation, and evaluation of NiLiCon — a transparent container replication mechanism for fault tolerance. To the best of our knowledge, NiLiCon is the first implementation of container replication, demonstrating that it can be used for transparent deployment of critical services in the cloud.NiLiCon is based on high-frequency asynchronous incremental checkpointing to a warm spare, as previously used for VMs. The challenge to accomplishing this is that, compared to VMs, there is much tighter coupling between the container state and the state of the underlying platform. NiLiCon meets this challenge, eliminating the need to deploy services in VMs, with performance overheads that are competitive with those of similar VM replication mechanisms. Specifically, with the seven benchmarks used in the evaluation, the performance overhead of NiLiCon is in the range of 19%-67%. For fail-stop faults, the recovery rate is 100%. Diyu Zhou, Yuval Tamir |
IPDPS | 2 |
| 2020 | Finding Network Misconfigurations by Automatic Template Inference
Siva Kesava Reddy K., Alan Tang, Ryan Beckett, Karthick Jayaraman, Todd D. Millstein, Yuval Tamir, George Varghese |
NSDI | 6 |
| 2019 | PUSh: Data Race Detection Based on Hardware-Supported Prevention of Unintended SharingabstractSome of the most difficult to find bugs in multi-threaded programs are caused by unintended sharing, leading to data races. The detection of data races can be facilitated by requiring programmers to explicitly specify any intended sharing and then verifying compliance with these intentions. We present a novel dynamic checker based on this approach, called PUSh. Diyu Zhou, Yuval Tamir |
MICRO | 2 |
| 2018 | Fast Hypervisor Recovery Without RebootabstractSystem recovery latency is decreased by using microreboot to reboot only the failed component instead of the entire system. For large, complex components, such as hypervisors, even the latency of microreboot is unacceptably high in important deployment scenarios. We investigate an alternative component-level recovery mechanism, which we call microreset, that can achieve dramatically lower recovery latency for some such components. Instead of component reboot, microreset quickly resets the component to a quiescent state that is highly likely to be valid and where the component is ready to handle new or retried interactions with the rest of the system. We present a recovery mechanism for the Xen hypervisor, called NiLiHype, based on microreset. We show that, compared to microreboot-based hypervisor recovery, NiLiHype achieves nearly the same recovery success rate but with a recovery latency that is shorter by a factor of over 30. Diyu Zhou, Yuval Tamir |
DSN | 2 |
| 2015 | Fault Injection in Virtualized Systems - Challenges and ApplicationsabstractWe analyze the interaction between system virtualization and fault injection: (i) use of virtualization to facilitate fault injection into non-virtualized systems, and (ii) use of fault injection to evaluate the dependability of virtualized systems. We explore the benefits of using virtualization for fault injection and discuss the challenges of implementing fault injection in virtualized systems along with resolutions to those challenges. For experimental evaluation, we use a test platform that consists of the Gigan fault injector, that we have developed, with the Xen virtual machine monitor. We evaluate the degree to which fault injection results obtained from running the target system in a virtual machine are comparable to running the target system on bare hardware. We compare results when injection is done from within the target system versus from the hosting hypervisor. We evaluate the performance benefits of leveraging system virtualization for fault injection. Finally, we demonstrate the capabilities of our injector and highlight the benefits of leveraging system virtualization for fault injection by describing deployments of Gigan to evaluate both non-virtualized and virtualized systems. Michael Le, Yuval Tamir |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2011 | Resilient Virtual ClustersabstractClusters of computers can provide, in aggregate, reliable services despite the failure of individual computers. System-level virtualization is widely used to consolidate the workload of multiple physical systems as multiple virtual machines (VMs) on a single physical computer. A single physical computer thus forms a \fIvirtual cluster\fP of VMs. A key difficulty with virtualization is that the failure of the virtualization infrastructure (VI) often leads to the failure of multiple VMs. This is likely to overload "cluster computing" resiliency mechanisms, typically designed to tolerate the failure of only a single node at a time. By supporting recovery from failure of key VI components, we have enhanced the resiliency of a VI (Xen), thus enabling the use of existing "cluster computing" techniques to provide resilient virtual clusters. In the overwhelming majority of cases, these enhancements allow recovery from errors in the VI to be accomplished without the failure of more than a single VM. The resulting resiliency of the virtual cluster is demonstrated by running two existing "cluster computing" systems while subjecting the VI to injected faults. Michael Le, Israel Hsu, Yuval Tamir |
PRDC | 3 |
| 2011 | ReHype: enabling VM survival across hypervisor failuresabstractWith existing virtualized systems, hypervisor failures lead to overall system failure and the loss of all the work in progress of virtual machines (VMs) running on the system. We introduce ReHype, a mechanism for recovery from hypervisor failures by booting a new instance of the hypervisor while preserving the state of running VMs. VMs are stalled during the hypervisor reboot and resume normal execution once the new hypervisor instance is running. Hypervisor failures can lead to arbitrary state corruption and inconsistencies throughout the system. ReHype deals with the challenge of protecting the recovered hypervisor instance from such corrupted state and resolving inconsistencies between different parts of hypervisor state as well as between the hypervisor and VMs and between the hypervisor and the hardware. We have implemented ReHype for the Xen hypervisor. The implementation was done incrementally, using results from fault injection experiments to identify the sources of dangerous state corruption and inconsistencies. The implementation of ReHype involved only 880 LOC added or modified in Xen. The memory space overhead of ReHype is only 2.1MB for a pristine copy of the hypervisor code and static data plus a small reserved memory area. The fault injection campaigns used to evaluate the effectiveness of ReHype involved a system with multiple VMs running I/O and hypercall-intensive benchmarks. Our experimental results show that the ReHype prototype can successfully recover from over 90% of detected hypervisor failures. Michael Le, Yuval Tamir |
VEE | 2 |
| 2009 | Maintaining Network QoS Across NIC Device Driver Failures Using VirtualizationabstractDevice driver failures have been shown to be a major cause of system failures. Network services stress NIC device drivers, increasing the probability of NIC driver bugs being manifested as server failures. System virtualization is increasingly used for server consolidation and management. The isolated driver domain (IDD) architecture used by several virtual machine monitors, such as Xen, forms a natural foundation for making systems resilient to NIC driver failures. In order to realize this potential, recovery must be fast enough to maintain QoS for network services across NIC driver failures. We show that the standard Xen configuration, enhanced with simple detection and recovery mechanisms, cannot provide such QoS. However, with NIC drivers isolated in two virtual machines, in a primary/warm-spare configuration, the system can recover from an overwhelming majority of NIC driver failures in under 10 ms. Michael Le, Andrew Gallagher, Yuval Tamir, Yoshio Turner |
NCA | 3 |
| 2009 | CoRAL: A transparent fault-tolerant web service
Navid Aghdaie, Yuval Tamir |
J. Syst. Softw. | 2 |
| 2007 | Deadlock-free connection-based adaptive routing with dynamic virtual circuits
Yoshio Turner, Yuval Tamir |
J. Parallel Distributed Comput. | 2 |
| 2005 | Understanding the energy efficiency of SMT and CMP with multiclusteringabstractIn this paper we study the energy efficiency of SMT and CMP with multiclustering. Through a detailed design space exploration, we show that clustering closes the energy efficiency gap between SMT and CMP at equal performance points. Specifically, we show that the energy efficiency of CMP compared to SMT at a given performance decreases from a maximum of 25% in a monolithic processor case to 6% when the processor resources are clustered. By carefully considering floorplans, we show that this is, in part, enabled by the small energy consumption (less than 3%) of the interconnection buses required for clustering, even with SMT. As the gap narrows, we show that the efficiency of SMT versus CMP depends on the contribution of leakage energy: at lower leakage, the CMP tends to be better than the SMT, while the SMT outperforms the CMP at higher leakage levels. We demonstrate these results over a wide range of performance and machine configurations Jason Cong, Ashok Jagannathan, Glenn Reinman, Yuval Tamir |
ISLPED | 4 |
| 2002 | Design and Validation of Portable Communication Infrastructure for Fault-Tolerant Cluster MiddlewareabstractWe describe the communication infrastructure (CI) for our fault-tolerant cluster middleware, which is optimized for two classes of communication: for the applications and for the cluster management middleware. This CI was designed for portability and for efficient operation on top of modern user-level message passing mechanisms. We present a functional fault model for the CI and show how platform-specific faults map to this fault model. Based on this fault model, we have developed a fault injection scheme that is integrated with the CI and is thus portable across different communication technologies. We have used fault injection to validate and evaluate the implementation of the CI itself as well as the cluster management middleware in the presence of communication faults. Ming Li 0020, Wenchao Tao, Daniel Goldberg 0002, Israel Hsu, Yuval Tamir |
CLUSTER | 5 |
| 2002 | Implementation and evaluation of transparent fault-tolerant Web service with kernel-level supportabstractMost of the techniques used for increasing the availability of Web services do not provide fault tolerance for requests being processed at the time of server failure. Other schemes require deterministic servers or changes to the Web client. These limitations are unacceptable for many current and future applications of the Web. We have developed an efficient implementation of a client-transparent mechanism for providing fault-tolerant Web service that does not have the limitations mentioned above. The scheme is based on a hot standby backup server that maintains logs of requests and replies. The implementation includes modifications to the Linux kernel and to the Apache Web server, using their respective module mechanisms. We describe the implementation and present an evaluation of the impact of the backup scheme in terms of throughput, latency, and CPU processing cycles overhead. Navid Aghdaie, Yuval Tamir |
ICCCN | 2 |
| 1994 | Coordinated Checkpointing-Rollback Error Recovery for Distributed Shared Memory MulticomputersabstractMost recovery schemes that have been proposed for Distributed Shared Memory (DSM) systems require unnecessarily high checkpointing frequency and checkpoint traffic, which are sensitive to the frequency of interprocess communication in the applications. For message-passing systems, low overhead error recovery based on coordinated checkpointing allows the frequency of checkpointing to be determined only by the reliability requirements of the application. Efficient adaptation of this approach to DSM multicomputers is complicated by the absence of explicit messages in DSM systems, the presence of a shared and partially replicated address space, and the presence of a distributed coherency directory. We present solutions to these issues, and propose an error recovery scheme based on coordinated checkpointing and rollback for DSM multicomputers. Our performance evaluation based on trace-driven simulations indicates that this scheme incurs less checkpoint traffic than recovery schemes previously proposed for DSM systems.> G. Janakiraman, Yuval Tamir |
SRDS | 2 |
| 1993 | Symmetric Crossbar Arbiters for VLSI Communication SwitchesabstractThe design and implementation of symmetric crossbar arbiters are addressed. Several arbiter designs are compared based on simulations of a multistage interconnection network. These simulations demonstrate the influence of the switch arbitration policy on network throughput, average latency, and worst-case latency. It is shown that some natural designs result in poor system performance and/or slow implementations. Two efficient arbiter implementations are proposed. Based on network simulations, VLSI implementation, and circuit simulation, it is shown that these arbiters achieve nearly optimal system performance without becoming the critical path that limits the system clock.> Yuval Tamir, Hsin-Chou Chi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1992 | Hardware Support for High-Priority Traffic in VLSI Communication Swithes
Yuval Tamir, Gregory L. Frazier |
J. Parallel Distributed Comput. | 1 |
| 1992 | Hierarchical Coherency Management for Shared Virtual Memory Multicomputers
Yuval Tamir, G. Janakiraman |
J. Parallel Distributed Comput. | 1 |
| 1992 | Dynamically-Allocated Multi-Queue Buffers for VLSI Communication SwitchesabstractSmall n*n switches are key components of interconnection networks used in multiprocessors and multicomputers. The architecture of these n*n switches, particularly their internal buffers, is critical for achieving high-throughput low-latency communication with cost-effective implementations. Several buffer structures are discussed and compared in terms of implementation complexity, inter-switch handshaking requirements, and their ability to deal with variations in traffic patterns and message lengths. A design for buffers that provide non-FIFO message handling and efficient storage allocation for variable size packets using linked lists managed by a simple on-chip controller is presented. The new buffer design is evaluated by comparing it to several alternative designs in the context of a multistage interconnection network. The modeling and simulation show that the new buffer outperforms alternative buffers and can thus be used to improve the performance of a wide variety of systems currently using less efficient buffers.> Yuval Tamir, Gregory L. Frazier |
IEEE Trans. Computers | 1 |
| 1991 | Decomposed Arbiters for Large Crossbars with Multi-Queue Input BuffersabstractCrossbars are key components of communication switches used to construct multiprocessor interconnection networks. For a fixed number of nodes, larger crossbars result in reduced probability of conflicts and allow packets to traverse the network in fewer hops. However, increasing the size of the crossbar also increases the delay of the arbiter used to resolve conflicting requests. The increased arbitration delay can lead to overall poor network performance. The impact of the increased arbitration delay can be mitigated by decomposing the arbitration process into multiple steps, such that some requests can be granted before the arbitration of the entire crossbar is complete. The design of such decomposed arbiters for larger crossbars is discussed. The focus is on crossbars with multi-queue buffers at their inputs. Such buffers have been shown to provide significantly higher performance than conventional FIFO buffers.> Hsin-Chou Chi, Yuval Tamir |
ICCD | 2 |
| 1990 | High-Performance Fault-Tolerant VLSI Systems Using Micro RollbackabstractA technique called micro rollback, which allows most of the performance penalty for concurrent error detection to be eliminated, is presented. Detection is performed in parallel with the transmission of information between modules, thus removing the delay for detection from the critical path. Erroneous information may thus reach its destination module several clock cycles before an error indication. Operations performed on this erroneous information are undone using a hardware mechanism for fast rollback of a few cycles. The implementation of a VLSI processor capable of micro rollback is discussed, as well as several critical issues related to its use in a complete system.> Yuval Tamir, Marc Tremblay |
IEEE Trans. Computers | 1 |
| 1989 | The design and implementation of a multi-queue buffer for VLSI communication switchesabstractThe micro-architecture and VLSI implementation of a dynamically allocated multiqueue (DAMQ) buffer are presented. Design tradeoffs for the DAMQ buffer's datapath are discussed and a floorplan and the timing of the major functional units are presented. It is shown that in VLSI switches, with buffers than can store multiple packets, additional chip area is better used for the control of DAMQ buffers than for increased buffer space in simpler FIFO buffers.> Gregory L. Frazier, Yuval Tamir |
ICCD | 2 |
| 1988 | High-Performance Multi-Queue Buffers for VLSI Communication SwitchesabstractA type of buffer called a dynamically allocated multiqueue (DAMQ) buffer, designed for use in n*n switches, is presented. This buffer provides efficient handling of variable-length packets and the forwarding of packets in non-FIFO (first-in-first-out) order. The microarchitecture of the DAMQ buffer and its controller is described in the context of the ComCoBB communication coprocessor for multicomputers. The DAMQ buffer can be efficiently implemented in LVSI to support packet transmission and reception at the rate of one byte per clock cycle. With a hardwired linked-list manager and a fast-routing mechanism, the ComCoBB chip will support virtual cut-through of messages with a latency of four cycles. The performance of the DAMQ buffer is compared with that of three alternative buffers in the context of a multistage interconnection network. Simulations show that for uniform traffic the DAMQ buffer results in significantly lower latencies and higher maximal throughput than other designs with the same total buffer storage capacity.> Yuval Tamir, Gregory L. Frazier |
ISCA | 1 |
| 1988 | Support for High-Priority Traffic in VLSI Communication SwitchesabstractThe design of small n*n switches that can be used to construct communication networks that provide low-latency communication for high-priority traffic, which is required for both multistage interconnection networks used in multiprocessors and direct networks used in multicomputers. The focus is on the design of the internal buffers, specifically on the design of buffers that provide non-FIFO handling of messages. Alternative designs and configurations are evaluated in the context of a multistage interconnection network. Simulation shows that a slightly modified version of the recently introduced dynamically allocated multiqueue buffer can provide superior support for high-priority traffic.> Yuval Tamir, Gregory L. Frazier |
RTSS | 1 |
| 1987 | A Software-Based Hardware Fault Tolerance Scheme for Multicomputers
Yuval Tamir, Eli Gafni |
ICPP | 1 |
| 1987 | A block-and-actions generator as an alternative to a simulator for collecting architecture measurementsabstractTo design a new processor or to modify an existing one, designers need to gather data to estimate the influence of specific architecture features on the performance of the proposed machine (PM). To obtain this data, it is necessary to measure on an existing machine (EM) the dynamic behavior of typical programs. Traditionally, simulators have been used to obtain measurements for PMs. Since several hundred EM instructions are required to decode, interpret, and measure each simulated (PM) instruction, the simulation time of typical programs is prohibitively large. Thus, designers tend to simulate only small programs and the results obtained might not be representative of a real system behavior. In this paper we present an alternative tool for collecting architecture measurements: the Block-and-Actions Generator (BKGEN). BKGEN produces a version of the program being measured which is directly executable by the EM. This executable version is obtained directly with the EM compiler or with the PM compiler and a assembly-to-assembly translator. The choice between these alternatives depends on the EM and PM compiler technology and the type of measurements to be obtained. BKGEN also collects the PM events to be measured (called actions). Each EM block of instructions is associated with a PM block of actions so that when the program is executed, it collects the measurements associated with the PM. The main advantage of BKGEN is that the execution time is substantially reduced compared to the execution time of a simulator while collecting similar data. Thus, large typical programs (compilers, assemblers, word processors, ...) can be used by the designer to obtain meaningful measurements. Miquel Huguet, Tomás Lang, Yuval Tamir |
PLDI | 3 |
| 1984 | Design and Application of Self-Testing Comparators Implemented with MOS PLA'sabstractA high probability of detecting errors caused by hardware faults is an essential property of any fault-tolerant system. VLSI technology makes the use of duplication and matching for error detection practical and attractive. A critical circuit in this context is a self-testing comparator. Faults in the comparator must be detected so that they do not mask discrepancies between the duplicated modules. Yuval Tamir, Carlo H. Séquin |
IEEE Trans. Computers | 1 |
| 1983 | Strategies for Managing the Register File in RISCabstractThe RISC (reduced instruction set computer) architecture attempts to achieve high performance without resorting to complex instructions and irregular pipelining schemes. One of the novel features of this architecture is a large register file which is used to minimize the overhead involved in procedure calls and returns. This paper investigates several strategies for managing this register file. The costs of practical strategies are compared with a lower bound on this management overhead, obtained from a theoretical optimal strategy, for several register file sizes. Yuval Tamir, Carlo H. Séquin |
IEEE Trans. Computers | 1 |