EDBT 2026 Demo / reviewers in the wild / expert
Matthew K. Farrens
dblp:41/4899
· DBLP profile ↗
40ranked-venue papers
9as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 9 first-author · 1 since 2021Software engineering, systems software and programming languages · 6 · 5 first-authorComputer networks · 4Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
21 papers |
Performance modeling and evaluation · 44% Interconnection networks and networks-on-chip · 13% Processor architecture and microarchitecture · 13% | |
| Computer networks
2 papers |
Software-defined and programmable networks · 40% Transport protocols and congestion control · 34% Network performance modeling · 26% |
Topics — the 30 heaviest of 59, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Software-defined and programmable networks
openflow |
0.2 | 1 | 2014 | Simultaneously Reducing Latency and Power Consumption in OpenFlow Switches · IEEE/ACM Trans. Netw. 2014 |
Performance modeling and evaluation › simulation › architectural simulation
full-system simulation |
0.2 | 1 | 2014 | PDG_GEN: A Methodology for Fast and Accurate Simulation of On-Chip Networks · IEEE Trans. Computers 2014 |
Performance modeling and evaluation › simulation › communication system simulation
network simulation |
0.2 | 1 | 2014 | PDG_GEN: A Methodology for Fast and Accurate Simulation of On-Chip Networks · IEEE Trans. Computers 2014 |
Performance modeling and evaluation
simulation |
0.2 | 1 | 2014 | PDG_GEN: A Methodology for Fast and Accurate Simulation of On-Chip Networks · IEEE Trans. Computers 2014 |
Performance modeling and evaluation › simulation › discrete-event simulation
trace-driven simulation |
0.2 | 1 | 2014 | PDG_GEN: A Methodology for Fast and Accurate Simulation of On-Chip Networks · IEEE Trans. Computers 2014 |
Transport protocols and congestion control › transport protocols
rate-based transport protocol |
0.1 | 1 | 2011 | Introspective end-system modeling to optimize the transfer time of rate based protocols · HPDC 2011 |
Electronic design automation › hardware verification and test
fault modeling |
0.1 | 1 | 2011 | Resilient microring resonator based photonic networks · MICRO 2011 |
Interconnection networks and networks-on-chip
optical interconnection networks |
0.1 | 1 | 2011 | Resilient microring resonator based photonic networks · MICRO 2011 |
Interconnection networks and networks-on-chip
optical network-on-chip |
0.1 | 1 | 2011 | Addressing system-level trimming issues in on-chip nanophotonic networks · HPCA 2011 |
Energy-efficient computing
power management |
0.1 | 1 | 2014 | Simultaneously Reducing Latency and Power Consumption in OpenFlow Switches · IEEE/ACM Trans. Netw. 2014 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 3 | 1999 | Exploiting ILP in Page-based Intelligent Memory · MICRO 1999 Techniques for extracting instruction level parallelism on MIMD architectures · MICRO 1993 MISC: a Multiple Instruction Stream Computer · MICRO 1992 |
Memory systems
cache management |
0.0 | 2 | 2000 | Eager writeback - a technique for improving bandwidth utilization · MICRO 2000 A modified approach to data cache management · MICRO 1995 |
Transport protocols and congestion control › transport protocols
UDP |
0.0 | 1 | 2011 | Introspective end-system modeling to optimize the transfer time of rate based protocols · HPDC 2011 |
Processor architecture and microarchitecture
branch prediction |
0.0 | 2 | 2000 | Branch Transition Rate: A New Metric for Improved Branch Classification Analysis · HPCA 2000 A comparision of superscalar and decoupled access/execute architectures · MICRO 1993 |
Performance modeling and evaluation › simulation
simulation-based evaluation |
0.0 | 1 | 2011 | Addressing system-level trimming issues in on-chip nanophotonic networks · HPCA 2011 |
Processor architecture and microarchitecture › branch prediction
branch classification |
0.0 | 1 | 2000 | Branch Transition Rate: A New Metric for Improved Branch Classification Analysis · HPCA 2000 |
Memory systems
cache |
0.0 | 1 | 2000 | Eager writeback - a technique for improving bandwidth utilization · MICRO 2000 |
Performance modeling and evaluation › simulation › analog and hybrid computer simulation
hybrid simulation |
0.0 | 1 | 2000 | HLS: combining statistical and symbolic simulation to guide microprocessor designs · ISCA 2000 |
Performance modeling and evaluation › simulation
processor simulation |
0.0 | 1 | 2000 | HLS: combining statistical and symbolic simulation to guide microprocessor designs · ISCA 2000 |
Processor architecture and microarchitecture › branch prediction
two-level adaptive branch prediction |
0.0 | 1 | 2000 | Branch Transition Rate: A New Metric for Improved Branch Classification Analysis · HPCA 2000 |
Processor architecture and microarchitecture
value prediction |
0.0 | 1 | 2000 | HLS: combining statistical and symbolic simulation to guide microprocessor designs · ISCA 2000 |
Memory systems › processing-in-memory
intelligent memory |
0.0 | 1 | 1999 | Exploiting ILP in Page-based Intelligent Memory · MICRO 1999 |
Memory systems
processing-in-memory |
0.0 | 1 | 1999 | Exploiting ILP in Page-based Intelligent Memory · MICRO 1999 |
Memory systems › memory management
virtual memory |
0.0 | 3 | 1992 | Modifying VM hardware to reduce address pin requirements · MICRO 1992 A partitioned translation lookaside buffer approach to reducing address bandwith · ISCA 1992 Dynamic Base Register Caching: A Technique for Reducing Address Bus Width · ISCA 1991 |
Memory systems
cache design |
0.0 | 2 | 1994 | A Study of Single-Chip Processor/Cache Organizations for Large Numbers of Transistors · ISCA 1994 Improving Performance of Small On-Chip Instruction Caches · ISCA 1989 |
Processor architecture and microarchitecture
instruction fetch |
0.0 | 2 | 1993 | A comparision of superscalar and decoupled access/execute architectures · MICRO 1993 Improving Performance of Small On-Chip Instruction Caches · ISCA 1989 |
Performance modeling and evaluation
workload characterization |
0.0 | 2 | 2000 | Branch Transition Rate: A New Metric for Improved Branch Classification Analysis · HPCA 2000 An Analysis of the Information Content of Address Reference Streams · MICRO 1991 |
Memory systems › cache › CPU cache
data cache |
0.0 | 1 | 1995 | A modified approach to data cache management · MICRO 1995 |
Processor architecture and microarchitecture
multicore design |
0.0 | 2 | 1993 | MISC: a Multiple Instruction Stream Computer · MICRO 1992 Techniques for extracting instruction level parallelism on MIMD architectures · MICRO 1993 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 1994 | A Study of Single-Chip Processor/Cache Organizations for Large Numbers of Transistors · ISCA 1994 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.4packet prediction · 0.4trace-based simulation · 0.2dependency inference · 0.2thermal simulation · 0.1sliding window scheme · 0.1retransmission · 0.1introspective end-system modeling · 0.1fault simulation · 0.1error correction · 0.1trace-driven simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Leveraging Network Delay Variability to Improve QoE of Latency Critical ServicesabstractEven as cloud providers offer strict guarantees on the intra-cloud delay of requests for Latency-Critical (LC) Services, a high external network delay can result in a large end-to-end delay, causing a low user Quality of Experience (QoE). Furthermore, due to the variability in the external network delay, there is a disconnect between the user’s QoE and the cloud guaranteed service level objective (SLO). Specifically, a request that meets the SLO, can have a high or low QoE depending on the external network delay. In this work we propose a usercentric End-to-end Service Level Objective (ESLO), an extension of the traditional cloud-centric SLO, that guarantees stricter bounds on end-to-end delay and thereby achieving a higher QoE. We show how the variability in the external network delay can be both addressed and leveraged to meet the ESLO and improve server utilization. We propose ESLO-aware extensions to the Kubernetes infrastructure, that uses information about the external network delay and its distribution - (a) to reduce the number of QoE-violating responses by using deadline-based scheduling at the service instances, and (b) to appropriately scale service instances with load. We implement the ESLO-aware framework on the NSF Chameleon cloud testbed and present experimental results demonstrating the benefit of the proposed paradigm. Sambit Kumar Shukla, Matthew K. Farrens |
NAS | 2 |
| 2020 | HCAPP: Scalable Power Control for Heterogeneous 2.5D Integrated SystemsabstractPackage pin allocation is becoming a key bottleneck in the capabilities of designs due to the increased bandwidth requirements. 2.5D integration compounds these package-level requirements while introducing an increased number of compute units within the package. We propose a decentralized power control implementation called Heterogeneous Constant Average Power Processing (HCAPP) to maintain the power limit while maximizing the efficiency of the package pins allocated for power. HCAPP uses a hardware-based decentralized design to handle fast power limits, maintain scalability and enable simplified control for heterogeneous systems while maximizing performance. As extensions, we evaluate a software interface and the impact of different accelerator designs. Overall, HCAPP achieves 7% speedup over a RAPL-like implementation. The power utilization improves from 79.7% (RAPL-like) to 93.9% (HCAPP) with this design. A priority-based static software control methodology alongside HCAPP provides average speedups of 8.3% (CPU), 5.4% (GPU), and 12% (Accelerator) for the prioritized component compared to the unprioritized version. Kramer Straube, Jason Lowe-Power, Christopher Nitta, Matthew K. Farrens, Venkatesh Akella |
ICPP | 4 |
| 2018 | Improving Provisioned Power Efficiency in HPC Systems with GPU-CAPPabstractIn this paper we propose a microarchitectural technique called GPU Constant Average Power Processing (GPU-CAPP) that improves the power utilization of power provisioning-limited systems by using provisioned power as much as possible to accelerate computation on parallel work-loads. GPU-CAPP uses a flexible, decentralized control to ensure fast response times and the scalability required for increasingly parallel GPU designs. We use GPGPU-Sim and GPUWattch to simulate GPU-CAPP and evaluate its capabilities on a subset of the Rodinia benchmark suite. Overall, GPU-CAPP enables speedup by an average of 26% and 12% over equivalent fixed frequency systems at two power targets. Kramer Straube, Jason Lowe-Power, Christopher Nitta, Matthew K. Farrens, Venkatesh Akella |
HiPC | 4 |
| 2017 | Improving Execution Time of Parallel Programs on Large Scale Chip Multiprocessors with Constant Average Power ProcessingabstractIn this paper we propose a microarchitectural technique called Constant Average Power Processing (CAPP) that reduces the execution time of parallel programs by dynamically detecting the power slack at runtime and directing it to specific core(s) that are the bottleneck at any given time. The key insight of this work is that by sensing the current, communicating it to the global controller and adjusting the cores' frequencies, it is possible to maintain a constant power level in a distributed and scalable manner. We evaluate the potential benefits and scalability of the proposed technique on a set of synthetic benchmarks and compare the results with related work such as Running Average Power Limit (RAPL). Kramer Straube, Christopher Nitta, Rajeevan Amirtharajah, Matthew K. Farrens, Venkatesh Akella |
ICCD | 4 |
| 2016 | Improving network performance on multicore systems: Impact of core affinities on high throughput flows
Nathan Hanford, Vishal Ahuja, Matthew K. Farrens, Dipak Ghosal, Mehmet Balman, Eric Pouyoul, Brian Tierney |
Future Gener. Comput. Syst. | 3 |
| 2015 | Performance Analysis of Real-Time Covert Timing Channel Detection Using a Parallel System
Ross K. Gegan, Rennie Archibald, Matthew K. Farrens, Dipak Ghosal |
NSS | 3 |
| 2014 | Impact of the end-system and affinities on the throughput of high-speed flowsabstractNetwork throughput is scaling "up" to higher data transfer rates while processors are scaling "out" to multiple cores. As a result, network adapter "offloads" and performance "tuning" have received a good deal of attention lately. However, much of this attention is focused on the "how" and not the "why" of performance efficiency. There are two types of efficiencies that we have found particularly intriguing: First, processor core "affinity," or "binding" is fundamentally the choice of which processor core or cores handle certain tasks in a network- or I/O-heavy application running on a MIMD machine. Second, Ethernet "pause frames" slightly violate the "end-to-end" nature of TCP/IP in order to perform link-to-link flow control. The goal of our research is to delve deeper into why these tuning suggestions and this offload exist, and how they affect the end-to-end performance and efficiency of a single, large TCP flow. Nathan Hanford, Vishal Ahuja, Matthew K. Farrens, Dipak Ghosal, Mehmet Balman, Eric Pouyoul, Brian Tierney |
ANCS | 3 |
| 2014 | PDG_GEN: A Methodology for Fast and Accurate Simulation of On-Chip NetworksabstractWith the advent of large scale chip multiprocessors, there is growing interest in the design and analysis of on-chip networks. Full-system simulation is the most accurate way to perform such an analysis, but unfortunately it is very slow and thus limits design space exploration. To overcome this problem researchers frequently use trace-based simulation to study different network topologies and properties, which can be done much faster. Unfortunately, unless the traces that are used include information about dependencies between packets, trace-based simulations can lead one to draw incorrect conclusions about network performance metrics such as average packet latency and overall execution time. The primary contributions of this work are to demonstrate the importance of including dependency information in traces, and to present PDG_GEN, an inference-based technique for identifying and including dependencies in traces. This technique uses traces obtained from multiple full-system simulations of an application of interest to infer dependency information between packets and augment traces with this information. On the SPLASH-2 benchmark suite, PDG_GEN is 2.3 times more accurate at predicting overall execution time and almost 4,000 times more accurate at predicting average packet latency than traditional trace-based methods. Kevin Macdonald, Christopher Nitta, Matthew K. Farrens, Venkatesh Akella |
IEEE Trans. Computers | 3 |
| 2014 | Simultaneously Reducing Latency and Power Consumption in OpenFlow SwitchesabstractThe Ethernet switch is a primary building block for today's enterprise networks and data centers. As network technologies converge upon a single Ethernet fabric, there is ongoing pressure to improve the performance and efficiency of the switch while maintaining flexibility and a rich set of packet processing features. The OpenFlow architecture aims to provide flexibility and programmable packet processing to meet these converging needs. Of the many ways to create an OpenFlow switch, a popular choice is to make heavy use of ternary content addressable memories (TCAMs). Unfortunately, TCAMs can consume a considerable amount of power and, when used to match flows in an OpenFlow switch, put a bound on switch latency. In this paper, we propose enhancing an OpenFlow Ethernet switch with per-port packet prediction circuitry in order to simultaneously reduce latency and power consumption without sacrificing rich policy-based forwarding enabled by the OpenFlow architecture. Packet prediction exploits the temporal locality in network communications to predict the flow classification of incoming packets. When predictions are correct, latency can be reduced, and significant power savings can be achieved from bypassing the full lookup process. Simulation studies using actual network traces indicate that correct prediction rates of 97% are achievable using only a small amount of prediction circuitry per port. These studies also show that prediction circuitry can help reduce the power consumed by a lookup process that includes a TCAM by 92% and simultaneously reduce the latency of a cut-through switch by 66%. Paul Congdon, Prasant Mohapatra, Matthew K. Farrens, Venkatesh Akella |
IEEE/ACM Trans. Netw. | 3 |
| 2012 | Cache-aware affinitization on commodity multicores for high-speed network flowsabstractFor a given TCP or UDP flow, protocol processing of incoming packets is performed on the core that receives the interrupt, while the user-space application which consumes the data may run on the same or a different core. If the cores are not the same, additional costs due to context switches, cache misses, and the movement of data between the caches of the cores may occur. The magnitude of this cost depends upon the processor affinity of the user-space process relative to the network stack. In this paper we present a prototype implementation of a tool which enables the application processing and protocol processing to occur on cores which share the lowest cache level. The Cache-Aware Affinity Deamon (CAAD) analyzes the topology of the die and the NIC characteristics and conveys information to the sender which allows the entire end-to-end path for each new flow to be be managed and controlled. This is done in a light-weight manner for both uni and bi-directional flows. Measurements show that for bulk data transfers using commodity multicore machines, the use of CAAD improves the overall TCP throughput by as much as 31%, and reduces the cache miss rate as much as 37.5%. GridFTP combined with CAAD improves the download time for big file transfers by up to 18%. Vishal Ahuja, Matthew K. Farrens, Dipak Ghosal |
ANCS | 2 |
| 2012 | Minimizing the Data Transfer Time Using Multicore End-System Aware Flow BifurcationabstractData centers are being deployed in a wide variety of environments (cloud computing, scientific, financial, defense, etc.). When geographically distributed, these data centers must transmit and receive growing volumes of data. In order to avoid congestion in the public internet, most use high speed dedicated optical networks, which can be thought of as private highways for carrying data. In this work, we examined the impact of such high speed network traffic on a commodity multicore machine, and identified a number of scenarios that cause packet loss and degraded throughput due to an end-system inability to consume incoming data fast enough. We show that high speed single flow traffic nullifies the benefits of multicore systems and multiqueue NICs, and we propose an end-system aware flow bifurcation technique to optimize the data transfer time using rate based protocols. Using introspective end-system modeling, we determine the optimal number of parallel flows required to utilize the available bandwidth, and the optimal rate for each of the flows. We compare our approach with GridFTP, which is a widely used data transfer protocol in computational grids, and show that our approach performs better (particularly when the end-system losses are in the receive ring buffer.). Vishal Ahuja, Dipak Ghosal, Matthew K. Farrens |
CCGRID | 3 |
| 2012 | DCAF - A Directly Connected Arbitration-Free Photonic Crossbar for Energy-Efficient High Performance ComputingabstractDCAF is a directly connected arbitration free photonic crossbar that is realized by taking advantage of multiple photonic layers connected with photonic vias. In order to evaluate DCAF we developed a detailed implementation model for the network and analyzed the power and performance on a variety of benchmarks, including SPLASH-2 and synthetic traces. Our results demonstrate that the overhead required by arbitration is non-trivial, especially at high loads. Eliminating the need for arbitration, sizing the buffers carefully and retransmitting lost packets when there is contention results in a 44% reduction in average packet latency without additional power overhead. We also use an analytical model for ScaLAPACK QR decomposition and find that a 64 processor DCAF could outperform a 1024 node cluster connected with 40Gbps links on matrices up to 500MB in size. Christopher Nitta, Matthew K. Farrens, Venkatesh Akella |
IPDPS | 2 |
| 2011 | Addressing system-level trimming issues in on-chip nanophotonic networksabstractThe basic building block of on-chip nanophotonic interconnects is the microring resonator, and these resonators change their resonant wavelengths due to variations in temperature - a problem that can be addressed using a technique called ”trimming”, which involves correcting the drift via heating and/or current injection. Thus far system researchers have modeled trimming as a per ring fixed cost. In this work we show that at the system level using a fixed cost model is inappropriate - our simulations demonstrate that the cost of heating has a non-linear relationship with the number of rings, and also that current injection can lead to thermal runaway. We show that a very narrow Temperature Control Window (TCW) must be maintained in order for the network to work as desired. However, by exploiting the group drift property of co-located rings, it is possible to create a sliding window scheme which can increase the TCW. We also show that partially athermal rings can alleviate but not eliminate the problem. Christopher Nitta, Matthew K. Farrens, Venkatesh Akella |
HPCA | 2 |
| 2011 | Introspective end-system modeling to optimize the transfer time of rate based protocolsabstractThe transmission capacity of today's high-speed networks is often greater than the capacity of an end-system (such as a server or a remote client) to consume the incoming data. The mismatch between the network and the end-system, which can be exacerbated by high end-system workloads, will result in incoming packets being dropped at different points in the packet receiving process. In particular, a packet may be dropped in the NIC, in the kernel ring buffer, and (for rate based protocols) in the socket buffer. To provide reliable data transfers, these losses require retransmissions, and if the loss rate is high enough result in longer download times. In this paper, we focus on UDP-like rate based transport protocols, and address the question of how best to estimate the rate at which the end-system can consume data which minimizes the overall transfer time of a file. Vishal Ahuja, Amitabha Banerjee, Matthew K. Farrens, Dipak Ghosal, Giuseppe Serazzi |
HPDC | 3 |
| 2011 | Resilient microring resonator based photonic networksabstractMicroring resonator-based photonic interconnects are being considered for both on-chip and off-chip communication in order to satisfy the power and bandwidth requirements of future large scale chip multiprocessors. However, microring resonators are prone to malfunction due to fabrication errors, and they are also extremely sensitive to fluctuations in temperature. In this paper we derive a fault model for microring based optical links that can be used by computer architects to make informed design choices. We evaluate different schemes for improving resilience, such as retransmission versus error-correction, using an optical fault simulator based on our fault model. We show how meeting a target mean time between failures (MTBF) affects the choice of resilience scheme - our investigation indicates that until fault rates are in the range of 10−21 to 10−24 per cycle, error detection/correction schemes will be needed in order to meet a 1M hour MTBF. We also evaluate how the resilience scheme impacts the performance of the link, which will help an architect choose the appropriate scheme based on the throughput requirements of a particular design. Christopher Nitta, Matthew K. Farrens, Venkatesh Akella |
MICRO | 2 |
| 2011 | Inferring packet dependencies to improve trace based simulation of on-chip networksabstractWith the advent of large scale chip-level multiprocessors, there is a growing interest in the design and analysis of on-chip networks. The use of full system simulation is the most accurate way to perform such an analysis, but unfortunately it is very slow and thus limits design space exploration. In order to overcome this problem researchers frequently use trace based simulation to study different network topologies and properties, which can be done much faster. Unfortunately, unless the traces that are used include information about dependencies between messages (packets), trace based simulation can lead one to draw incorrect conclusions about network performance metrics such as latency and overall execution time. In this paper we will demonstrate the importance of including dependency information in traces, as well as present an inference-based technique for identifying and including dependencies, and show that using these augmented traces results in much better simulation accuracy without excessively extending simulation time. Christopher Nitta, Kevin Macdonald, Matthew K. Farrens, Venkatesh Akella |
NOCS | 3 |
| 2010 | Performance Evaluation of a Multicore System with Optically Connected Memory ModulesabstractOver the past years, there have been impressive advances in bandwidth, latency, and scalability of on-chip networks. However, unless the off-chip network bandwidth and latency are also improved, we might have unbalanced systems which will limit the improvements to overall system performance. In this paper, we show how dense wavelength-division multiplexing (DWDM) -based optical interconnects could be used to emulate multiple buses in a fully-buffered DIMM (FB-DIMM) -like memory system to improve both bandwidth and latency. We evaluate an optically connected memory using full-system simulations of an 8-core system running memory-intensive multithreaded workloads. We show that for the FFT benchmark, optically connected memory can reduce the average memory request latency by 29% compared to a single-channel DDR3 SDRAM system and provide an overall performance speedup of 1.20. We also show that at least two DDR3 memory channels are needed to match the performance of a single optical bus, which demonstrates the advantage of optical interconnects in terms of savings in the number of pins required. Paul Vincent Mejia, Rajeevan Amirtharajah, Matthew K. Farrens, Venkatesh Akella |
NOCS | 3 |
| 2008 | Packet prediction for speculative cut-through switchingabstractThe amount of intelligent packet processing in an Ethernet switch continues to grow, in order to support of embedded applications such as network security, load balancing and quality of service assurance. This increased packet processing is contributing to greater per-packet latency through the switch. Paul Congdon, Matthew K. Farrens, Prasant Mohapatra |
ANCS | 2 |
| 2008 | Design and evaluation of an optical CPU-DRAM interconnectabstractWe present OCDIMM (Optically Connected DIMM), a CPU-DRAM interface that uses multiwavelength optical interconnects. We show that OCDIMM is more scalable and offers higher bandwidth and lower latency than FBDIMM (Fully-Buffered DIMM), a state-of-the-art electrical alternative. Though OCDIMM is more power efficient than FBDIMM, we show that ultimately the total power consumption in the memory subsystem is a key impediment to scalability and thus to achieving truly balanced computing systems in the terascale era. Amit Hadke, Tony Benavides, Rajeevan Amirtharajah, Matthew K. Farrens, Venkatesh Akella |
ICCD | 4 |
| 2008 | Techniques for increasing effective data bandwidthabstractIn this paper we examine techniques for increasing the effective bandwidth of the microprocessor off-chip interconnect. We focus on mechanisms that are orthogonal to other techniques currently being studied (3-D fabrication, optical interconnect, etc.) Using a range of full-system simulations we study the distribution of values being transferred to and from memory, and find that (as expected) high entropy data such as floating point numbers have limited compressibility, but that other data types offer more potential for compression. By using a simple heuristic to classify the contents of a cache line and providing different compression schemes for each classification, we show it is possible to provide overall compression at a cache line granularity comparable to that obtained by using a much more complex Lempel-Ziv-Welch algorithm. Christopher Nitta, Matthew K. Farrens |
ICCD | 2 |
| 2000 | The Decoupled-Style Prefetch Architecture (Research Note)
Kevin D. Rich, Matthew K. Farrens |
Euro-Par | 2 |
| 2000 | Code Partitioning in Decoupled Compilers
Kevin D. Rich, Matthew K. Farrens |
Euro-Par | 2 |
| 2000 | Branch Transition Rate: A New Metric for Improved Branch Classification AnalysisabstractRecent studies have shown significantly improved branch prediction through the use of branch classification. By separating static branches into groups, or classes, with similar dynamic behavior predictors may be selected that are best suited for each class. Previous methods have classified branches according to taken rate (or bias). We propose a new metric for branch classification: branch transition rate, which is defined as the number of times a branch changes direction between taken and not taken during execution. We show that transition rate is a more appropriate indicator of branch behavior than taken rate for determining predictor performance. When both metrics are combined, an even clearer picture of dynamic branch behavior emerges, in which expected predictor performance for a branch is closely correlated with its combined taken and transition rate class. Using this classification, a small group of branches is identified for which two-level predictors are ineffective. Michael Haungs, Phil Sallee, Matthew K. Farrens |
HPCA | 3 |
| 2000 | HLS: combining statistical and symbolic simulation to guide microprocessor designsabstractAs microprocessors continue to evolve, many optimizations reach a point of diminishing returns. We introduce HLS, a hybrid processor simulator which uses statistical models and symbolic execution to evaluate design alternatives. This simulation methodology allows for quick and accurate contour maps to be generated of the performance space spanned by design parameters. We validate the accuracy of HLS through correlation with existing cycle-by-cycle simulation techniques and current generation hardware. We demonstrate. The power of HLS by exploring design spaces defined by two parameters: code properties and value prediction. These examples motivate how HLS can be used to set design goals and individual component performance targets. Mark Oskin, Fred Chong, Matthew K. Farrens |
ISCA | 3 |
| 2000 | Eager writeback - a technique for improving bandwidth utilizationabstractModern high-performance processors utilize multi-level cache structures to help tolerate the increasing latency of main memory. Most of these caches employ either a writeback or a write-through strategy to deal with store operations. Write-through caches propagate data to more distant memory levels at the time each store occurs, which requires a very large bandwidth between the memory hierarchy levels. Writeback caches can significantly reduce the bandwidth requirements between caches and memory by marking cache lines as dirty when stores are processed and writing those lines to the memory system only when that dirty line is evicted. Unfortunately, for applications that experience significant numbers of cache misses due to streaming data, writeback cache designs can degrade overall system performance by clustering bus activity when dirty lines contend with data being fetched into the cache. In this paper we present a new technique called Eager Writeback, which re-distributes and balances memory traffic by writing and "cleaning" dirty cache lines prior to their eviction. Eager Writeback can be viewed as a compromise between write-through and writeback policies, in which dirty lines are written later than write-through, but prior to writeback. We will show that this approach can reduce the large number of writes seen in a write-through design, while avoiding the performance degradation caused by clustering bus traffic in a writeback approach. Hsien-Hsin S. Lee, Gary S. Tyson, Matthew K. Farrens |
MICRO | 3 |
| 1999 | Exploiting ILP in Page-based Intelligent MemoryabstractThis study compares the speed, area, and power of different implementations of Active Pages, an intelligent memory system which helps bridge the growing gap between processor and memory performance by associating simple functions with each page of data. Previous investigations have shown up to 1000X speedups using a block of reconfigurable logic to implement these functions next to each subarray on a DRAM chip. In this study, we show that instruction-level parallelism, not hardware specialization, is the key to the previous success with reconfigurable logic. In order to demonstrate this fact, an Active Page implementation based upon a simplified VLIW processor was developed. Unlike conventional VLIW processors, power and area constraints lead to a design which has a small number of pipeline stages. Our results demonstrate that a four-wide VLIW processor attains comparable performance to that of pure FPGA logic but requires significantly less area and power. Mark Oskin, Justin Hensley, Diana Franklin, Fred Chong, Matthew K. Farrens, Aneet Chopra |
MICRO | 5 |
| 1998 | Utilizing Reuse Information in Data Cache ManagementabstractAs microprocessor speeds continue to outgrow memory subsystem speeds, minimizing the average data access time grows in importance. As current data caches are often poorly and inefficiently managed, a good management technique can improve the average data access time. This paper presents a comparative evaluation of two approaches that utilize reuse information for more efficiently managing the firstlevel cache. While one approach is based on the effective address of the data being referenced, the other uses the program counter of the memory instruction generating the reference. Our evaluations show that using effective address reuse information performs better than using program counter reuse information. In addition, we show that the Victim cache performs best for multi-lateral caches with a direct-mapped main cache and high L2 cache latency, while the NTS (effective-addressbased) approach performs better as the L2 latency decreases or the associativity of the main cache increases. Jude A. Rivers, Edward S. Tam, Gary S. Tyson, Edward S. Davidson, Matthew K. Farrens |
International Conference on Supercomputing | 5 |
| 1995 | A modified approach to data cache managementabstractAs processor performance continues to improve, more emphasis must be placed on the performance of the memory system. In this paper, a detailed characterization of data cache behavior for individual load instructions is given. We show that by selectively applying cache line allocation according the characteristics of individual load instructions, overall performance can be improved for both the data cache and the memory system. This approach can improve some aspects of memory performance by as much as 60 percent on existing executables. Gary S. Tyson, Matthew K. Farrens, John Matthews, Andrew R. Pleszkun |
MICRO | 2 |
| 1994 | A Study of Single-Chip Processor/Cache Organizations for Large Numbers of TransistorsabstractPresents a trace-driven simulation-based study of a wide range of cache configurations and processor counts. This study was undertaken in an attempt to help answer the question of how best to allocate large numbers of transistors, a question that is rapidly increasing in importance as transistor densities continue to climb. At what point does continuing to increase the size of the on-chip first level cache cease to provide sufficient increases in hit rate and become prohibitively difficult to access in a single cycle? In order to compare different configurations, the concept of an Equivalent Cache Transistor is presented. Results indicate that the access time of the first-level data cache is more important than the size. In addition, it appears that once approximately 15 million transistors become available, a two processor configuration is preferable to a single processor with correspondingly larger caches.> Matthew K. Farrens, Gary S. Tyson, Andrew R. Pleszkun |
ISCA | 1 |
| 1993 | A comparision of superscalar and decoupled access/execute architecturesabstractEven with a very accurate dynamic branch predictor, a superscalar processor must predict instruction fetch addresses no later than the first pipeline stage to avoid suffering pipeline bubbles every time a branch is taken. Unfortunately, branch addresses generally are not known prior to instruction decode. Therefore, some indirect technique is required to identify a branch instruction and enable branch prediction while the branch instruction is being fetched. This is the branch identification problem. Intel Pentium adopts a scheme that solves this problem; however, its scheme assumes an issue rate of two instructions per cycle. An aggressive superscalar processor, issuing more than two instructions per cycle, cannot effectively use that scheme. The authors propose and compare two viable schemes for solving the branch identification problem for wide-issue superscalar processors.> Matthew K. Farrens, Pius Ng, Phil Nico |
MICRO | 1 |
| 1993 | Techniques for extracting instruction level parallelism on MIMD architecturesabstractExtensive research has been done on extracting parallelism from single instruction stream processors. The authors present some results of an investigation into ways to modify MIMD architectures to allow them to extract the instruction level parallelism achieved by current superscalar and VLIW machines. A new architecture is proposed which utilizes the advantages of a multiple instruction stream design while addressing some of the limitations that have prevented MIMD architectures from performing ILP operation. A new code scheduling mechanism is described to support this new architecture by partitioning instructions across multiple processing elements in order to exploit this level of parallelism.> Gary S. Tyson, Matthew K. Farrens |
MICRO | 2 |
| 1992 | A partitioned translation lookaside buffer approach to reducing address bandwithabstractSimulations indicate a simple modification of existing virtual memory hardware can significantly reduce the number of pins required to transmit address information from processor to off-chip memory. This modification consists of partitioning a TLB so that virtual page numbers are stored in a cache on the processor and corresponding real page numbers are sotred in registers at the memory, making it possible to transmit a small register index instead of the entire real page number. Matthew K. Farrens, Arvin Park, Rob Fanfelle, Pius Ng, Gary S. Tyson |
ISCA | 1 |
| 1992 | Modifying VM hardware to reduce address pin requirements
Matthew K. Farrens, Arvin Park, Gary S. Tyson |
MICRO | 1 |
| 1992 | MISC: a Multiple Instruction Stream ComputerabstractThis paper describes a single chip Multiple Instruction Stream Computer (MISC) capable of extracting instruction level parallelism from a broad spectrum of programs. The MISC architecture uses multiple asynchronous processing elements to separate a program into streams that can be executed in parallel, and integrates a conflict-free message passing system into the lowest level of the processor design to facilitate low latency intra-MISC communication. This approach allows for increased machine parallelism with minimal code expansion, and provides an alternative approach to single instruction stream multi-issue machines such as SuperScalar and VLIW. # # 1. Introduction The goal of most high-performance computers is to maximize the amount of work that can be done per unit time. This quantity (work) can be expressed by the following equation [HePa90]: work = clock rate× instruction count 1 ############### × Clocks per Instruction 1 ################### A number of different approa... Gary S. Tyson, Matthew K. Farrens, Andrew R. Pleszkun |
MICRO | 2 |
| 1991 | Dynamic Base Register Caching: A Technique for Reducing Address Bus WidthabstractWhenaddress reference degrees of spatial and temporal higher order address lines carry streams exhibit high locality, many of the redundant information.By caching the higher order portions of address references in a set of dynamically allocated base registers, it becomes possible to transmit small register indices between the processor and memory instead of the high order address bits themselves.Trace driven simulations indicate that this technique can significantly reduce processor-to-memory address bus width without an appreciable loss in performance, fhereby increasing available processor bandwidth.Our resulfs imply that as much as 25% of the available 1/0 bandwidth of a processor is used less than 1% of the time. Matthew K. Farrens, Arvin Park |
ISCA | 1 |
| 1991 | Strategies for Achieving Improved Processor ThroughputabstractDeeply pipelined processors have relatively low issue rates due to dependencies between instructions. In this paper we examine the possibility of interleaving a second stream of instructions into the pipeline, which would issue instructions during the cycles the first stream was unable to. Such an interleaving has the potential to significantly increase the throughput of a processor without seriously imparing the execution of either process. We propose a dynamic interleaving of at most 2 instructions streams, which share the the pipelined functional units of a machine. To support the interleaving of 2 instruction streams a number of interleaving policies are described and discused. Finally, the amount of improvement in processor throughput is evaluated by simulating the interleaving policies for several machine variants. 1. Introduction An important metric for evaluating processor performance, especially in a multiprocessing context, is the throughput rate of a processor (defined as h... Matthew K. Farrens, Andrew R. Pleszkun |
ISCA | 1 |
| 1991 | An Analysis of the Information Content of Address Reference StreamsabstractWe analyze the information content of several address reference streams.Our results indicate that a new scheme, based on Dynamic Huffman Coding [Vitt87], can encode a typical 32 bit address in four to seven bits.Unlike previous schemes used to estimate the information content of address words [HaDa771 ~arnm77], our scheme is completely on-line and does not rely on preeomputation of address transition probabilities.Our results imply that at least 83% of address bits in the traces we studied contain redundant information.Although our coding scheme is too complex and computationally expensive to implement in practice, it provides a lower bound on the bandwidth that can be achieved by practical compression schemes.Through use of these address compression techniques, the number of bus lines and 1/0 pins required to transmit address information between processor and memory can be ptly reduced. Jeffrey C. Becker, Arvin Park, Matthew K. Farrens |
MICRO | 3 |
| 1991 | Workload and Implementation Considerations for Dynamic Base Register CachingabstractDynamic Base Register Caching (DBRC) FaFa90] lJhPa91] has been shown to be a useful technique for significantly reducing processor to memory address bandwidth.By caching the higher order portions of memory addresses in a set of dynamically allocated base registers, only small register indices need to be transmitted between the processor and memory instead of the high order address bits themselves.In this paper we present the results of trace driven simulations which indicate that DR13C can facilitate the provision of separate paths for instructions Matthew K. Farrens, Arvin Park |
MICRO | 1 |
| 1991 | Alleviation of tree saturation in multistage interconnection networksabstractThis paper presents an examination of two distinct but complementary extensions of previous work on hot spot contention in multistage interconnection networks. The first extension focuses on the use of larger queues at the memory modules than are traditionally studied. The second extension explores a simple feedback damping scheme, which we refer to as bleeding, which allows selected processors to ignore feedback information. The impact of memory queue size, feedback threshold value, and bleeding on system performance (specifically maximum bandwidth per processor) is evaluated by analyzing the results of extensive network simulations. These results indicate that combining these approaches can significantly improve the effective bandwidth of a multistage interconnection network. 1. Introduction In order to achieve the goal of teraflop computing speeds by the end of the century, the number of processing nodes in a multiprocessor will have to increase substantially. However, as the numb... Matthew K. Farrens, Brad Wetmore, Allison Woodruff |
SC | 1 |
| 1989 | Improving Performance of Small On-Chip Instruction CachesabstractMost current single-chip processors employ an on-chip instruction cache to improve performance. A miss in this instruction cache will cause an external memory reference which must compete with data references for access to the external memory, thus affecting the overall performance of the processor. One common way to reduce the number of off-chip instruction requests is to increase the size of the on-chip cache. An alternative approach is presented in this paper, in which a combination of an instruction cache, instruction queue and instruction queue buffer is used to achieve the same effect with a much smaller instruction cache size. Such an approach is significant for emerging technologies where high circuit densities are initially difficult to achieve yet a high level of performance is desired, or for more mature technologies where chip area can be used to provide more functionality. The viability of this approach is demonstrated by its implementation in an existing single-chip processor. Matthew K. Farrens, Andrew R. Pleszkun |
ISCA | 1 |