VLDB 2026 Research / reviewers in the wild / expert
Al Davis
dblp:d/AlDavis · also Alan L. Davis
· DBLP profile ↗
42ranked-venue papers
6as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 2 first-authorSoftware engineering, systems software and programming languages · 12 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
19 papers |
Memory systems · 47% Interconnection networks and networks-on-chip · 13% Hardware reliability and fault tolerance · 8% | |
| Computer networks
1 paper |
Optical networks · 100% |
Topics — the 30 heaviest of 54, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Interconnection networks and networks-on-chip › routing algorithms
adaptive routing |
0.5 | 2 | 2019 | Practical and efficient incremental adaptive routing for HyperX networks · SC 2019 HyperX: topology, routing, and packaging of efficient large-scale networks · SC 2009 |
Parallel and multicore computing
load balancing |
0.4 | 1 | 2019 | Practical and efficient incremental adaptive routing for HyperX networks · SC 2019 |
Memory systems
DRAM |
0.4 | 4 | 2016 | Staged Reads: Mitigating the impact of DRAM writes on DRAM reads · HPCA 2012 Micro-pages: increasing DRAM efficiency with locality-aware data placement · ASPLOS 2010 A unified memory network architecture for in-memory computing in commodity servers · MICRO 2016 |
Memory systems › DRAM
DRAM architecture |
0.3 | 2 | 2012 | LOT-ECC: Localized and tiered reliability mechanisms for commodity memory systems · ISCA 2012 Rethinking DRAM design and organization for energy-constrained multi-cores · ISCA 2010 |
Memory systems
in-memory computing |
0.2 | 1 | 2016 | A unified memory network architecture for in-memory computing in commodity servers · MICRO 2016 |
Memory systems › memory disaggregation
memory network |
0.2 | 1 | 2016 | A unified memory network architecture for in-memory computing in commodity servers · MICRO 2016 |
Hardware reliability and fault tolerance › error correction
error-correcting codes |
0.2 | 2 | 2014 | LOT-ECC: Localized and tiered reliability mechanisms for commodity memory systems · ISCA 2012 MemZip: Exploring unconventional benefits from memory compression · HPCA 2014 |
Memory systems › memory compression
hardware compressed memory |
0.2 | 1 | 2014 | MemZip: Exploring unconventional benefits from memory compression · HPCA 2014 |
Memory systems
memory compression |
0.2 | 1 | 2014 | MemZip: Exploring unconventional benefits from memory compression · HPCA 2014 |
Energy-efficient computing
memory energy efficiency |
0.2 | 1 | 2014 | MemZip: Exploring unconventional benefits from memory compression · HPCA 2014 |
Memory systems › DRAM › DRAM microarchitecture
rank subsetting |
0.2 | 1 | 2014 | MemZip: Exploring unconventional benefits from memory compression · HPCA 2014 |
Memory systems
3d-stacked memory |
0.2 | 1 | 2013 | Quantifying the relationship between the power delivery network and architectural policies in a 3D-stacked memory device · MICRO 2013 |
Electronic design automation › power integrity
IR-drop |
0.2 | 1 | 2013 | Quantifying the relationship between the power delivery network and architectural policies in a 3D-stacked memory device · MICRO 2013 |
Integrated circuit design
power delivery network |
0.2 | 1 | 2013 | Quantifying the relationship between the power delivery network and architectural policies in a 3D-stacked memory device · MICRO 2013 |
Hardware reliability and fault tolerance › error-correcting codes for memory
chipkill correct |
0.1 | 1 | 2012 | LOT-ECC: Localized and tiered reliability mechanisms for commodity memory systems · ISCA 2012 |
Memory systems › memory controller
DRAM scheduling |
0.1 | 1 | 2012 | Staged Reads: Mitigating the impact of DRAM writes on DRAM reads · HPCA 2012 |
Hardware reliability and fault tolerance
memory reliability |
0.1 | 1 | 2012 | LOT-ECC: Localized and tiered reliability mechanisms for commodity memory systems · ISCA 2012 |
Interconnection networks and networks-on-chip › switch architecture
high-radix switch |
0.1 | 1 | 2011 | The role of optics in future high radix switch design · ISCA 2011 |
Memory systems
memory interface |
0.1 | 1 | 2011 | Combining memory and a controller with photonics through 3D-stacking to enable scalable and energy-efficient systems · ISCA 2011 |
Performance modeling and evaluation › simulation › architectural simulation
cycle-accurate simulation |
0.1 | 1 | 2019 | Practical and efficient incremental adaptive routing for HyperX networks · SC 2019 |
Memory systems › DRAM › DRAM microarchitecture
row buffer management |
0.1 | 1 | 2010 | Micro-pages: increasing DRAM efficiency with locality-aware data placement · ASPLOS 2010 |
Rendering
ray tracing |
0.1 | 1 | 2009 | StreamRay: a stream filtering architecture for coherent ray tracing · ASPLOS 2009 |
GPUs and heterogeneous computing
graphics accelerator |
0.1 | 1 | 2009 | StreamRay: a stream filtering architecture for coherent ray tracing · ASPLOS 2009 |
Interconnection networks and networks-on-chip
network topology |
0.1 | 1 | 2009 | HyperX: topology, routing, and packaging of efficient large-scale networks · SC 2009 |
Electronic design automation › physical design
routing |
0.1 | 1 | 2009 | HyperX: topology, routing, and packaging of efficient large-scale networks · SC 2009 |
Memory systems
memory controller |
0.1 | 3 | 2012 | LOT-ECC: Localized and tiered reliability mechanisms for commodity memory systems · ISCA 2012 Design of a Parallel Vector Access Unit for SDRAM Memory Systems · HPCA 2000 Impulse: Building a Smarter Memory Controller · HPCA 1999 |
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2008 | Corona: System Implications of Emerging Nanophotonic Technology · ISCA 2008 |
Interconnection networks and networks-on-chip › optical interconnection networks
nanophotonic interconnect |
0.1 | 1 | 2008 | Corona: System Implications of Emerging Nanophotonic Technology · ISCA 2008 |
Energy-efficient computing
power management |
0.1 | 2 | 2013 | Quantifying the relationship between the power delivery network and architectural policies in a 3D-stacked memory device · MICRO 2013 Rethinking DRAM design and organization for energy-constrained multi-cores · ISCA 2010 |
Storage systems
i/o architecture |
0.1 | 1 | 2006 | Design Trade-Offs for User-Level I/O Architectures · IEEE Trans. Computers 2006 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.5power gating · 0.2event-driven simulation · 0.2distance-aware selective compression · 0.2metadata placement · 0.2compressed data placement · 0.2pin count modeling · 0.2IR-drop analysis · 0.2checksum codes · 0.1area overhead analysis · 0.1technology scaling analysis · 0.1power and performance comparison · 0.1stream filtering · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Gen-Z Chipsetfor Exascale FabricsabstractThis article consists of a collection of slides from the author's conference presentation. Patrick Knebel, Daniel A. Berkram, Al Davis, Darel Emmot, Paolo Faraboschi, Gary Gostin |
Hot Chips Symposium | 3 |
| 2019 | Practical and efficient incremental adaptive routing for HyperX networksabstractIn efforts to increase performance and reduce cost, modern low-diameter networks are designed for average case traffic and rely on non-minimal adaptive routing for network load-balancing when adversarial traffic patterns are encountered. Source adaptive routing is the predominant method for adaptive routing even though it presents many deficiencies related to making global decisions based solely on local information. In contrast, incremental adaptive routing, which performs an adaptive decision at every hop, is able to increase throughput and reduce latency by overcoming the deficiencies of source adaptive routing. We present two incremental adaptive routing algorithms for HyperX which are the first to be fully implementable in modern high-radix router architectures and interconnection network protocols. Using cycle accurate simulations of a 4,096 node network, our evaluation shows these algorithms are able to exceed the performance of prior work by as much as 4x with synthetic traffic and 25% with 27-point stencil traffic. Nic McDonald, Mikhail Isaev, Adriana Flores, Al Davis, John Kim 0001 |
SC | 4 |
| 2018 | SuperSim: Extensible Flit-Level Simulation of Large-Scale Interconnection NetworksabstractThe interconnection networks of modern largescale computing systems are quickly increasing in size and complexity to keep up with the demand for computing capability. These systems rely heavily on complex router microarchitectures and intelligent adaptive routing algorithms structured for cost-optimized low-diameter networks. These technologies need to be properly modeled and evaluated during design space exploration and for performance characterization of the system. We present SuperSim, an open-source flit-level interconnection network simulator that enables focused evaluation of issues related to designing and deploying large-scale highperformance networks. SuperSim is a programmer-centric simulation framework explicitly designed to be flexibly extended and is supported by a number of tools making it easy to use and allowing users to model systems quickly. In this work we show the results for simulation case studies demonstrating the power of SuperSim to uncover otherwise overlooked details in large-scale interconnection networks. Nic McDonald, Adriana Flores, Al Davis, Mikhail Isaev, John Kim 0001, Doug Gibson |
ISPASS | 3 |
| 2016 | A unified memory network architecture for in-memory computing in commodity serversabstractIn-memory computing is emerging as a promising paradigm in commodity servers to accelerate data-intensive processing by striving to keep the entire dataset in DRAM. To address the tremendous pressure on the main memory system, discrete memory modules can be networked together to form a memory pool, enabled by recent trends towards richer memory interfaces (e.g. Hybrid Memory Cubes, or HMCs). Such an inter-memory network provides a scalable fabric to expand memory capacity, but still suffers from long multi-hop latency, limited bandwidth, and high power consumption — problems that will continue to exacerbate as the gap between interconnect and transistor performance grows. Moreover, inside each memory module, an intra-memory network (NoC) is typically employed to connect different memory partitions. Without careful design, the back-pressure inside the memory modules can further propagate to the inter-memory network to cause a performance bottleneck. To address these problems, we propose co-optimization of intra- and inter-memory network. First, we re-organize the intra-memory network structure, and provide a smart I/O interface to reuse the intra-memory NoC as the network switches for inter-memory communication, thus forming a unified memory network. Based on this architecture, we further optimize the inter-memory network for both high performance and lower energy, including a distance-aware selective compression scheme to drastically reduce communication burden, and a light-weight power-gating algorithm to turn off under-utilized links while guaranteeing a connected graph and deadlock-free routing. We develop an event-driven simulator to model our proposed architectures. Experiment results based on both synthetic traffic and real big-data workloads show that our unified memory network architecture can achieve 75.1% average memory access latency reduction and 22.1% total memory energy saving. Jia Zhan, Itir Akgun, Jishen Zhao, Al Davis, Paolo Faraboschi, Yuangang Wang, Yuan Xie 0001 |
MICRO | 4 |
| 2015 | Memory Considerations for Low Energy Ray TracingabstractAbstract We propose two hardware mechanisms to decrease energy consumption on massively parallel graphics processors for ray tracing. First, we use a streaming data model and configure part of the L2 cache into a ray stream memory to enable efficient data processing through ray reordering. This increases L1 hit rates and reduces off‐chip memory energy substantially through better management of off‐chip memory access patterns. To evaluate this model, we augment our architectural simulator with a detailed memory system simulation that includes accurate control, timing and power models for memory controllers and off‐chip dynamic random‐access memory . These details change the results significantly over previous simulations that used a simpler model of off‐chip memory, indicating that this type of memory system simulation is important for realistic simulations that involve external memory. Secondly, we employ reconfigurable special‐purpose pipelines that are constructed dynamically under program control. These pipelines use shared execution units that can be configured to support the common compute kernels that are the foundation of the ray tracing algorithm. This reduces the overhead incurred by on‐chip memory and register accesses. These two synergistic features yield a ray tracing architecture that reduces energy by optimizing both on‐chip and off‐chip memory activity when compared to a more traditional approach. Daniel M. Kopta, Konstantin Shkurko, Josef B. Spjut, Erik Brunvand, Al Davis |
Comput. Graph. Forum | 5 |
| 2014 | MemZip: Exploring unconventional benefits from memory compressionabstractMemory compression has been proposed and deployed in the past to grow the capacity of a memory system and reduce page fault rates. Compression also has secondary benefits: it can reduce energy and bandwidth demands. However, most prior mechanisms have been designed to focus on the capacity metric and few prior works have attempted to explicitly reduce energy or bandwidth. Further, mechanisms that focus on the capacity metric also require complex logic to locate the requested data in memory. In this paper, we design a highly simple compressed memory architecture that does not target the capacity metric. Instead, it focuses on complexity, energy, bandwidth, and reliability. It relies on rank subsetting and a careful placement of compressed data and metadata to achieve these benefits. Further, the space made available via compression is used to boost other metrics - the space can be used to implement stronger error correction codes or energy-efficient data encodings. The best performing MemZip configuration yields a 45% performance improvement and 57% memory energy reduction, compared to an uncompressed non-sub-ranked baseline. Another energy-optimized configuration yields a 29.8% performance improvement and a 79% memory energy reduction, relative to the same baseline. Ali Shafiee, Meysam Taassori, Rajeev Balasubramonian, Al Davis |
HPCA | 4 |
| 2014 | NDC: Analyzing the impact of 3D-stacked memory+logic devices on MapReduce workloadsabstractWhile Processing-in-Memory has been investigated for decades, it has not been embraced commercially. A number of emerging technologies have renewed interest in this topic. In particular, the emergence of 3D stacking and the imminent release of Micron's Hybrid Memory Cube device have made it more practical to move computation near memory. However, the literature is missing a detailed analysis of a killer application that can leverage a Near Data Computing (NDC) architecture. This paper focuses on in-memory MapReduce workloads that are commercially important and are especially suitable for NDC because of their embarrassing parallelism and largely localized memory accesses. The NDC architecture incorporates several simple processing cores on a separate, non-memory die in a 3D-stacked memory package; these cores can perform Map operations with efficient memory access and without hitting the bandwidth wall. This paper describes and evaluates a number of key elements necessary in realizing efficient NDC operation: (i) low-EPI cores, (ii) long daisy chains of memory devices, (iii) the dynamic activation of cores and SerDes links. Compared to a baseline that is heavily optimized for MapReduce execution, the NDC design yields up to 15X reduction in execution time and 18X reduction in system energy. Seth H. Pugsley, Jeffrey Jestes, Rajeev Balasubramonian, Vijayalakshmi Srinivasan, Alper Buyuktosunoglu, Al Davis, Feifei Li 0001 |
ISPASS | 7 |
| 2013 | Quantifying the relationship between the power delivery network and architectural policies in a 3D-stacked memory deviceabstractMany of the pins on a modern chip are used for power delivery. If fewer pins were used to supply the same current, the wires and pins used for power delivery would have to carry larger currents over longer distances. This results in an "IR-drop" problem, where some of the voltage is dropped across the long resistive wires making up the power delivery network, and the eventual circuits experience fluctuations in their supplied voltage. The same problem also manifests if the pin count is the same, but the current draw is higher. IR-drop can be especially problematic in 3D DRAM devices because (i) low cost (few pins and TSVs) is a high priority, (ii) 3D-stacking increases current draw within the package without providing proportionate room for more pins, and (iii) TSVs add to the resistance of the power delivery network. Manjunath Shevgoor, Jung-Sik Kim, Niladrish Chatterjee, Rajeev Balasubramonian, Al Davis, Aniruddha N. Udipi |
MICRO | 5 |
| 2012 | The role of photonics in future data centersabstractThe most prolific information appliance today is a mobile device, typically a phone or a tablet. The near ubiquity of the internet makes it very easy to access vast amounts of information from these mobile devices. The storage and processing capability to support this infrastructure is increasingly in the datacenter and both the number of these warehouse scale computational (WSC) facilities and their size is increasing. Compound annual growth rates (CAGR) of both storage requirements and internet traffic is approximately 50%. Even more alarming is that in 2011, 2% of the total energy used in the United States was consumed by this information technology. Power and the attendant cooling requirements fundamentally affect both the cost of our IT infrastructure as well as what can be competitively designed. As VLSI technology scales, the power and speed of the transistors scale nicely but wires scale less favorably and the gap grows with every new process step. The telecom industry has long recognized the benefits of optical as opposed to electrical communication over long distances. However the definition of "long" changes with the signalling rate. Electrical signalling at high speeds consumes too much power, has significant signal integrity problems which limit bandwidth and at the WSC scale these issues border on catastrophic for future datacenters. Recent advances in silicon nanophotonics may well provide a solution to this problem. This talk will provide a brief introduction to silicon nanophotonic devices and then delve into how this technology can be used in the construction of more energy efficient datacenters of the future and discuss where this technology will be most useful and quantify the benefits for intra-datacenter networks. Al Davis |
ACM Great Lakes Symposium on VLSI | 1 |
| 2012 | Staged Reads: Mitigating the impact of DRAM writes on DRAM readsabstractMain memory latencies have always been a concern for system performance. Given that reads are on the critical path for CPU progress, reads must be prioritized over writes. However, writes must be eventually processed and they often delay pending reads. In fact, a single channel in the main memory system offers almost no parallelism between reads and writes. This is because a single off-chip memory bus is shared by reads and writes and the direction of the bus has to be explicitly turned around when switching from writes to reads. This is an expensive operation and its cost is amortized by carrying out a burst of writes or reads every time the bus direction is switched. As a result, no reads can be processed while a memory channel is busy servicing writes. This paper proposes a novel mechanism to boost read-write parallelism and perform useful components of read operations even when the memory system is busy performing writes. If some of the banks are busy servicing writes, we start issuing reads to the other idle banks. The results of these reads are stored in a few registers near the memory chip's I/O pads. These results are quickly returned immediately following the bus turnaround. The process is referred to as a Staged Read because it decouples a single read operation into two stages, with the first step being performed in parallel with writes. This innovation can also be viewed as a form of prefetch that is internal to a memory chip. The proposed technique works best when there is bank imbalance in the write stream. We also introduce a write scheduling algorithm that artificially creates bank imbalance and allows useful read operations to be performed during the write drain. Across a suite of memory-intensive workloads, we show that Staged Reads can boost throughput by up to 33% (average 7%) with an average DRAM access latency improvement of 17%, while incurring a very small cost (0.25%) in terms of memory chip area. The throughput improvements are even greater when considering write-intensive workloads (average 11%) or future systems (average 12%). Niladrish Chatterjee, Naveen Muralimanohar, Rajeev Balasubramonian, Al Davis, Norman P. Jouppi |
HPCA | 4 |
| 2012 | LOT-ECC: Localized and tiered reliability mechanisms for commodity memory systemsabstractMemory system reliability is a serious and growing concern in modern servers. Existing chipkill-level memory protection mechanisms suffer from several draw-backs. They activate a large number of chips on every memory access - this increases energy consumption, and reduces performance due to the reduction in rank-level parallelism. Additionally, they increase access granularity, resulting in wasted bandwidth in the absence of sufficient access locality. They also restrict systems to use narrow-I/O ×4 devices, which are known to be less energy-efficient than the wider ×8 DRAM devices. In this paper, we present LOT-ECC, a localized and multi-tiered protection scheme that attempts to solve these problems. We separate error detection and error correction functionality, and employ simple checksum and parity codes effectively to provide strong fault-tolerance, while simultaneously simplifying implementation. Data and codes are localized to the same DRAM row to improve access efficiency. We use system firmware to store correction codes in DRAM data memory and modify the memory controller to handle data mapping. We thus build an effective fault-tolerance mechanism that provides strong reliability guarantees, activates as few chips as possible (reducing power consumption by up to 44.8% and reducing latency by up to 46.9%), and reduces circuit complexity, all while working with commodity DRAMs and operating systems. Finally, we propose the novel concept of a heterogeneous DIMM that enables the extension of LOT-ECC to ×16 and wider DRAM parts. Aniruddha N. Udipi, Naveen Muralimanohar, Rajeev Balasubramonian, Al Davis, Norman P. Jouppi |
ISCA | 4 |
| 2012 | Leveraging Heterogeneity in DRAM Main Memories to Accelerate Critical Word AccessabstractThe DRAM main memory system in modern servers is largely homogeneous. In recent years, DRAM manufacturers have produced chips with vastly differing latency and energy characteristics. This provides the opportunity to build a heterogeneous main memory system where different parts of the address space can yield different latencies and energy per access. The limited prior work in this area has explored smart placement of pages with high activities. In this paper, we propose a novel alternative to exploit DRAM heterogeneity. We observe that the critical word in a cache line can be easily recognized beforehand and placed in a low-latency region of the main memory. Other non-critical words of the cache line can be placed in a low-energy region. We design an architecture that has low complexity and that can accelerate the transfer of the critical word by tens of cycles. For our benchmark suite, we show an average performance improvement of 12.9% and an accompanying memory energy reduction of 15%. Niladrish Chatterjee, Manjunath Shevgoor, Rajeev Balasubramonian, Al Davis, Zhen Fang 0002, Ramesh Illikkal, Ravi R. Iyer 0001 |
MICRO | 4 |
| 2012 | Fast, effective BVH updates for animated scenesabstractBounding volume hierarchies (BVHs) are a popular acceleration structure choice for animated scenes rendered with ray tracing. This is due to the relative simplicity of refitting bounding volumes around moving geometry. However, the quality of such a refitted tree can degrade rapidly if objects in the scene deform or rearrange significantly as the animation progresses, resulting in dramatic increases in rendering times and a commensurate reduction in the frame rate. The BVH could be rebuilt on every frame, but this could take significant time. We present a method to efficiently extend refitting for animated scenes with tree rotations, a technique previously proposed for off-line improvement of BVH quality for static scenes. Tree rotations are local restructuring operations which can mitigate the effects that moving primitives have on BVH quality by rearranging nodes in the tree during each refit rather than triggering a full rebuild. The result is a fast, lightweight, incremental update algorithm that requires negligible memory, has minor update times, parallelizes easily, avoids significant degradation in tree quality or the need for rebuilding, and maintains fast rendering times. We show that our method approaches or exceeds the frame rates of other techniques and is consistently among the best options regardless of the animated scene. Daniel M. Kopta, Thiago Ize, Josef B. Spjut, Erik Brunvand, Al Davis, Andrew Kensler |
I3D | 5 |
| 2011 | Prediction Based DRAM Row-Buffer Management in the Many-Core EraabstractModern processors are experiencing interleaved memory access streams from different threads/cores, reducing the spatial locality that is seen at the memory controller, making the combined stream appear increasingly random. Traditional methods for exploiting locality at the DRAM level, such as open-page and timer-based policies, become less effective as the number of threads accessing memory increases. Employing closed-page policies in such systems can improve performance but it eliminates any possibility of exploiting locality. In this paper, we build upon the key insight that a history-based predictor that tracks the number of accesses to a given DRAM page is a much better indicator of DRAM locality than timer based policies. We extend prior work to propose a simple Access Based Predictor (ABP) that tracks limited access history at the page level to determine page closure decisions, and does so with much smaller storage overhead than previously proposed policies. We show that ABP, with additional optimizations, can improve system throughput by 12.3% and 21.6% over open and closed-page policies, respectively. The proposed ABP requires 20 KB of storage overhead and is outside the critical path of memory access. Manu Awasthi, David W. Nellans, Rajeev Balasubramonian, Al Davis |
PACT | 4 |
| 2011 | The role of optics in future high radix switch designabstractFor large-scale networks, high-radix switches reduce hop and switch count, which decreases latency and power. The ITRS projections for signal-pin count and per-pin bandwidth are nearly flat over the next decade, so increased radix in electronic switches will come at the cost of less per-port bandwidth. Silicon nanophotonic technology provides a long-term solution to this problem. We first compare the use of photonic I/O against an all-electrical, Cray YARC inspired baseline. We compare the power and performance of switches of radix 64, 100, and 144 in the 45, 32, and 22 nm technology steps. In addition with the greater off-chip bandwidth enabled by photonics, the high power of electrical components inside the switch becomes a problem beyond radix 64. Nathan L. Binkert, Al Davis, Norman P. Jouppi, Moray McLaren, Naveen Muralimanohar, Robert Schreiber, Jung Ho Ahn |
ISCA | 2 |
| 2011 | Combining memory and a controller with photonics through 3D-stacking to enable scalable and energy-efficient systemsabstractIt is well-known that memory latency, energy, capacity, bandwidth, and scalability will be critical bottlenecks in future large-scale systems. This paper addresses these problems, focusing on the interface between the compute cores and memory, comprising the physical interconnect and the memory access protocol. For the physical interconnect, we study the prudent use of emerging silicon-photonic technology to reduce energy consumption and improve capacity scaling. We conclude that photonics are effective primarily to improve socket-edge bandwidth by breaking the pin barrier, and for use on heavily utilized links. For the access protocol, we propose a novel packet based interface that relinquishes most of the tight control that the memory controller holds in current systems and allows the memory modules to be more autonomous, improving flexibility and interoperability. The key enabler here is the introduction of a 3D-stacked interface die that allows both these optimizations without modifying commodity memory dies. The interface die handles all conversion between optics and electronics, as well as all low-level memory device control functionality. Communication beyond the interface die is fully electrical, with TSVs between dies and low-swing wires on-die. We show that such an approach results in substantially lowered energy consumption, reduced latency, better scalability to large capacities, and better support for heterogeneity and interoperability. Aniruddha N. Udipi, Naveen Muralimanohar, Rajeev Balasubramonian, Al Davis, Norman P. Jouppi |
ISCA | 4 |
| 2010 | Handling the problems and opportunities posed by multiple on-chip memory controllersabstractModern processors such as Tilera's Tile64, Intel's Nehalem, and AMD's Opteron are migrating memory controllers (MCs) on-chip, while maintaining a large, flat memory address space. This trend to utilize multiple MC's will likely continue and a core or socket will consequently need to route memory requests to the appropriate MC via an inter- or intra-socket interconnect fabric similar to AMD's HyperTransport(TM), or Intel's Quick-Path Interconnect(TM). Such systems are therefore subject to non-uniform memory access (NUMA) latencies because of the time spent traveling to remote MCs. Each MC will act as the gateway to a particular piece of the physical memory. Data placement will therefore become increasingly critical in minimizing memory access latencies. Manu Awasthi, David W. Nellans, Kshitij Sudan, Rajeev Balasubramonian, Al Davis |
PACT | 5 |
| 2010 | Micro-pages: increasing DRAM efficiency with locality-aware data placementabstractPower consumption and DRAM latencies are serious concerns in modern chip-multiprocessor (CMP or multi-core) based compute systems. The management of the DRAM row buffer can significantly impact both power consumption and latency. Modern DRAM systems read data from cell arrays and populate a row buffer as large as 8 KB on a memory request. But only a small fraction of these bits are ever returned back to the CPU. This ends up wasting energy and time to read (and subsequently write back) bits which are used rarely. Traditionally, an open-page policy has been used for uni-processor systems and it has worked well because of spatial and temporal locality in the access stream. In future multi-core processors, the possibly independent access streams of each core are interleaved, thus destroying the available locality and significantly under-utilizing the contents of the row buffer. In this work, we attempt to improve row-buffer utilization for future multi-core systems. Kshitij Sudan, Niladrish Chatterjee, David W. Nellans, Manu Awasthi, Rajeev Balasubramonian, Al Davis |
ASPLOS | 6 |
| 2010 | Photonics and future datacenter networks
Al Davis |
Hot Chips Symposium | 1 |
| 2010 | Efficient MIMD architectures for high-performance ray tracingabstractRay tracing efficiently models complex illumination effects to improve visual realism in computer graphics. Typical modern GPUs use wide SIMD processing, and have achieved impressive performance for a variety of graphics processing including ray tracing. However, SIMD efficiency can be reduced due to the divergent branching and memory access patterns that are common in ray tracing codes. This paper explores an alternative approach using MIMD processing cores custom-designed for ray tracing. By relaxing the requirement that instruction paths be synchronized as in SIMD, caches and less frequently used area expensive functional units may be more effectively shared. Heavy resource sharing provides significant area savings while still maintaining a high MIMD issue rate from our numerous light-weight cores. This paper explores the design space of this architecture and compares performance to the best reported results for a GPU ray tracer and a parallel ray tracer using general purpose cores. We show an overall performance that is six to ten times higher in a similar die area. Daniel M. Kopta, Josef B. Spjut, Erik Brunvand, Al Davis |
ICCD | 4 |
| 2010 | Rethinking DRAM design and organization for energy-constrained multi-coresabstractDRAM vendors have traditionally optimized the cost-per-bit metric, often making design decisions that incur energy penalties. A prime example is the overfetch feature in DRAM, where a single request activates thousands of bit-lines in many DRAM chips, only to return a single cache line to the CPU. The focus on cost-per-bit is questionable in modern-day servers where operating costs can easily exceed the purchase cost. Modern technology trends are also placing very different demands on the memory system: (i)queuing delays are a significant component of memory access time, (ii) there is a high energy premium for the level of reliability expected for business-critical computing, and (iii) the memory access stream emerging from multi-core systems exhibits limited locality. All of these trends necessitate an overhaul of DRAM architecture, even if it means a slight compromise in the cost-per-bit metric. Aniruddha N. Udipi, Naveen Muralimanohar, Niladrish Chatterjee, Rajeev Balasubramonian, Al Davis, Norman P. Jouppi |
ISCA | 5 |
| 2009 | StreamRay: a stream filtering architecture for coherent ray tracingabstractThe wide availability of commodity graphics processors has made real-time graphics an intrinsic component of the human/computer interface. These graphics cores accelerate the z-buffer algorithm and provide a highly interactive experience at a relatively low cost. However, many applications in entertainment, science, and industry require high quality lighting effects such as accurate shadows, reflection, and refraction. These effects can be difficult to achieve with z-buffer algorithms but are straightforward to implement using ray tracing. Although ray tracing is computationally more complex, the algorithm exhibits excellent scaling and parallelism properties. Nevertheless, ray tracing memory access patterns are difficult to predict and the parallelism speedup promise is therefore hard to achieve. Karthik Ramani, Christiaan P. Gribble, Al Davis |
ASPLOS | 3 |
| 2009 | HyperX: topology, routing, and packaging of efficient large-scale networksabstractIn the push to achieve exascale performance, systems will grow to over 100,000 sockets, as growing cores-per-socket and improved single-core performance provide only part of the speedup needed. These systems will need affordable interconnect structures that scale to this level. To meet the need, we consider an extension of the hypercube and flattened butterfly topologies, the HyperX, and give an adaptive routing algorithm, DAL. HyperX takes advantage of high-radix switch components that integrated photonics will make available. Our main contributions include a formal descriptive framework, enabling a search method that finds optimal HyperX configurations; DAL; and a low cost packaging strategy for an exascale HyperX. Simulations show that HyperX can provide performance as good as a folded Clos, with fewer switches. We also describe a HyperX packaging scheme that reduces system cost. Our analysis of efficiency, performance, and packaging demonstrates that the HyperX is a strong competitor for exascale networks. Jung Ho Ahn, Nathan L. Binkert, Al Davis, Moray McLaren, Robert S. Schreiber |
SC | 3 |
| 2008 | Corona: System Implications of Emerging Nanophotonic TechnologyabstractWe expect that many-core microprocessors will push performance per chip from the 10 gigaflop to the 10 teraflop range in the coming decade. To support this increased performance, memory and inter-core bandwidths will also have to scale by orders of magnitude. Pin limitations, the energy cost of electrical signaling, and the non-scalability of chip-length global wires are significant bandwidth impediments. Recent developments in silicon nanophotonic technology have the potential to meet these off- and on-stack bandwidth requirements at acceptable power levels. Corona is a 3 D many-core architecture that uses nanophotonic communication for both inter-core communication and off-stack communication to memory or I/O devices. Its peak floating-point performance is 10 teraflops. Dense wavelength division multiplexed optically connected memory modules provide 10 terabyte per second memory bandwidth. A photonic crossbar fully interconnects its 256 low-power multithreaded cores at 20 terabyte per second bandwidth. We have simulated a 1024 thread Corona system running synthetic benchmarks and scaled versions of the SPLASH-2 benchmark suite. We believe that in comparison with an electrically-connected many-core alternative that uses the same on-stack interconnect power, Corona can provide 2 to 6 times more performance on many memory intensive workloads, while simultaneously reducing power. Dana Vantrease, Robert Schreiber, Matteo Monchiero, Moray McLaren, Norman P. Jouppi, Marco Fiorentino, Al Davis, Nathan L. Binkert, Raymond G. Beausoleil, Jung Ho Ahn |
ISCA | 7 |
| 2008 | Fast ray tracing and the potential effects on graphics and gaming courses
Peter Shirley, Kelvin Sung, Erik Brunvand, Al Davis, Steven G. Parker, Solomon Boulos |
Comput. Graph. | 4 |
| 2007 | Application driven embedded system design: a face recognition case studyabstractThe key to increasing performance without a commensurate increase in power consumption in modern processors lies in increasing both parallelism and core specialization. Core specialization has been employed in the embedded space and is likely to play an important role in future heterogeneous multi-core architectures as well. In this paper, the face recognition application domain is employed as a case study to showcase an architectural design methodology which generates a specialized core with high performance and very low powercharacteristics. Specifically, we create "ASIC-like" execution flows to sustain the high memory parallelism generated within the core. The price of this benefit is a significant increase in compilation complexity. The crux of the problem is the need to co-schedule the often conflicting constraints of data access, data movement, and computation. A modular compiler approach that employs integer linear programming (ILP) based "interconnect-aware" instruction and data scheduling techniques to solve this problem is then described. The resulting core running the compiled code delivers a 1.65x throughput improvement over a high performance processor (Pentium 4) while simultaneously achieving an 80x energy-delay improvement over an energy-efficient processor (XScale) and performs real-time face recognition at embedded power budgets. Karthik Ramani, Al Davis |
CASES | 2 |
| 2007 | REFS Keynote: "Requirements for Services: Does it Make Sense?"abstractSummary form only given. Since the dawn of computers, we have struggled with the best ways to record requirements for our software-based systems. Many thousands of papers have been written that describe "new" ways to discover, prune, write, interrelate, test, and manage changes to, software requirements. However, there are two trends (one historic and one future) worth examining more closely: (1) Historic. "Systems" and "products" have been around for many centuries before the advent of computers; requirements have only become important since computers because software has given us so much more flexibility in the features we provide to our customers. Before software, products were constructed of plastic, metal, and wood and little flexibility existed. (2) Future. With the advent of widespread broadband access to the internet, more and more companies are discovering the economies of "delivering" software not as a product but as a service. This relatively new "Software as a Service" (SOS) business model raises the issue of whether the lessons we have learned about discovering, pruning, writing, interrelating, testing, and managing changes to, software requirements still apply. But more importantly, just as software has given us incredible flexibility in the features of our products (no longer built exclusively of physical materials), now software has given us the same kind of flexibility in the features of our services (no longer based solely on human delivery). What lessons still apply? Is the business of requirements engineering unchanged? Or do new principles apply to requirements for services? Al Davis |
COMPSAC (2) | 1 |
| 2006 | Design Trade-Offs for User-Level I/O ArchitecturesabstractTo address the growing I/O bottleneck, next-generation distributed I/O architectures employ scalable point-to-point interconnects and minimize operating system overhead by providing user-level access to the I/O subsystem. Reduced I/O overhead allows I/O intensive applications to efficiently employ latency hiding techniques for improved throughput. This paper presents the design of a novel scalable user-level I/O architecture and evaluates the impact of various architectural mechanisms in terms of overall performance improvement. Results demonstrate that eliminating data movement across protection domains is the dominant contributor to improved scalability. Eliminating system call and interrupt overhead only has a small additional benefit that may not justify the additional hardware support required. While this evaluation is based on one specific design, the conclusions can be generalized to other user-level I/O architectures Lambert Schaelicke, Al Davis |
IEEE Trans. Computers | 2 |
| 2005 | Keynote: Just Enough Requirements Management for Web Engineering
Al Davis |
ICWE | 1 |
| 2004 | A low power architecture for embedded perceptionabstractRecognizing speech, gestures, and visual features are important interface capabilities for future embedded mobile systems. Unfortunately, the real-time performance requirements of complex perception applications cannot be met by current embedded processors and often even exceed the performance of high performance microprocessors whose energy consumption far exceeds embedded energy budgets. Though custom ASICs provide a solution to this problem, they incur expensive and lengthy design cycles and are inflexible. This paper introduces a VLIW perception processor which uses a combination of clustered function units, compiler controlled dataflow and compiler controlled clock-gating in conjunction with a scratch-pad memory system to achieve high performance for perceptual algorithms at low energy consumption. The architecture is evaluated using ten benchmark applications taken from complex speech and visual feature recognition, security, and signal processing domains. The energy-delay product of a 0.13μ implementation of this architecture is compared against ASICs and general purpose processors. Using a combination of Spice simulations and real processor power measurements, we show that the cluster running at 1 GHz clock frequency outperforms a 2.4 GHz Pentium 4 by a factor of 1.75 while simultaneously achieving 159 times better energy delay product than a low power Intel XScale embedded processor. Binu K. Mathew, Al Davis, Michael A. Parker |
CASES | 2 |
| 2004 | Energy efficient cluster co-processors [3G wireless applications]abstractNew 3G wireless algorithms require more performance than can be currently provided by embedded processors. ASICs provide the necessary performance but are costly to design and sacrifice generality. This paper introduces a clustered VLIW coprocessor approach that organizes the execution and storage resources differently than a traditional general-purpose processor or DSP. The execution units of the coprocessor are clustered and embedded in a rich set of communication resources. Fine grain control of these resources is imposed by a wide-word horizontal micro-code program. The advantages of this approach are quantified on a suite of six algorithms that are taken from both traditional DSP applications and from the new 3G cellular telephony domain. The result is surprising. The execution clusters retain much of the generality of a conventional processor while simultaneously improving performance by one to two orders of magnitude and by reducing energy-delay by three to four orders of magnitude when compared to a conventional embedded processor such as the Intel XScale. Ali Ibrahim, Michael A. Parker, Al Davis |
ICASSP (5) | 3 |
| 2003 | A low-power accelerator for the SPHINX 3 speech recognition systemabstractAccurate real-time speech recognition is not currently possible in the mobile embedded space where the need for natural voice interfaces is clearly important. The continuous nature of speech recognition coupled with an inherently large working set creates significant cache interference with other processes. Hence real-time recognition is problematic even on high-performance general-purpose platforms. This paper provides a detailed analysis of CMU's latest speech recognizer (Sphinx 3.2), identifies three distinct processing phases, and quantifies the architectural requirements for each phase. Several optimizations are then described which expose parallelism and drastically reduce the bandwidth and power requirements for real-time recognition. A special-purpose accelerator for the dominant Gaussiann probability phase is developed for a 0.25μ CMOS process which is then analyzed and compared with Sphinx's measured energy and performance on a 0.13μ 2.4 GHz Pentium 4 system. The results show an improvement in power consumption by a factor of 29 at equivalent processing throughput. However after normalizing for process, the special-purpose approach has twice the throughput, and consumes 104 times less energy than the general-purpose processor. The energy-delay product is a better comparison metric due to the inherent design trade-offs between energy consumption and performance. The energy-delay product of the special-purpose approach is 196 times better than the Pentium 4. These results provide strong evidence that real-time large vocabulary speech recognition can be done within a power budget commensurate with embedded processing using today's technology. Binu K. Mathew, Al Davis, Zhen Fang 0002 |
CASES | 2 |
| 2001 | Requirements Triage: The Most Important Part of Software Engineeringand the Most Ignored
Al Davis |
SEKE | 1 |
| 2000 | Design of a Parallel Vector Access Unit for SDRAM Memory SystemsabstractWe are attacking the memory bottleneck by building a "smart" memory controller that improves effective memory bandwidth, bus utilization, and cache efficiency by letting applications dictate how their data is accessed and cached. This paper describes a parallel vector access unit (PVA), the vector memory subsystem that efficiently "gathers" sparse, strided data structures in parallel on a multi-bank SDRAM memory. We have validated our PVA design via gate-level simulation, and have evaluated its performance via functional simulation and formal analysis. On unit-stride vectors, PVA performance equals or exceeds that of an SDRAM system optimized for cache line fills. On vectors with larger strides, the PVA is up to 32.8 times faster. Our design is up to 3.3 times faster than a pipelined, serial SDRAM memory system that gathers sparse vector data, and the gathering mechanism is two to five times faster than in other PVAs with similar goals. Our PVA only slightly increases hardware complexity with respect to these other systems, and the scalable design is appropriate for a range of computing platforms, from vector supercomputers to commodity PCs. Binu K. Mathew, Sally A. McKee, John B. Carter, Al Davis |
HPCA | 4 |
| 2000 | Profiling I/O Interrupts in Modern ArchitecturesabstractAs applications grow increasingly communication-oriented, interrupt performance quickly becomes a crucial component of high performance I/O system design. At the same time, accurately measuring interrupt handler performance is difficult with the traditional simulation, instrumentation, or statistical sampling approaches. One of the most important components of interrupt performance is cache behavior. This paper presents a portable method for measuring the cache effects of I/O interrupt handling using hardware performance counters. The method is demonstrated on two commercial platforms with different architectures, the SGI Origin 200 and the Sun Ultra-1. This case study uses the methodology to measure the overhead of the two most common forms of interrupts: disk and network interrupts. It demonstrates that the method works well and is reasonably robust. In addition, the results show that network interrupts have larger cache footprints than disk interrupts, and behave very differently on both platforms, due to significant differences in OS organization. Lambert Schaelicke, Al Davis, Sally A. McKee |
MASCOTS | 2 |
| 2000 | Algorithmic foundations for a parallel vector access memory systemabstractThis paper presents mathematical foundations for the design of a memory controller subcomponent that helps to bridge the processor/memory performance gap for applications with strided access patterns. The Parallel Vector Access (PVA) unit exploits the regularity of vectors or streams to access them efficiently in parallel on a multi-bank SDRAM memory system. The PVA unit performs scatter/gather operations so that only the elements accessed by the application are transmitted across the system bus. Vector operations are broadcast in parallel to all memory banks, each of which implements an efficient algorithm to determine which vector elements it holds. Earlier performance evaluations have demonstrated that our PVA implementation loads elements up to 32.8 times faster than a conventional memory system and 3.3 times faster than a pipelined vector unit, without hurting the performance of normal cache-line fills. Here we present the underlying PVA algorithms for both word interleaved and cache-line inter-leaved memory systems. Binu K. Mathew, Sally A. McKee, John B. Carter, Al Davis |
SPAA | 4 |
| 1999 | Impulse: Building a Smarter Memory ControllerabstractImpulse is a new memory system architecture that adds two important features to a traditional memory controller. First, Impulse supports application-specific optimizations through configurable physical address remapping. By remapping physical addresses, applications control how their data is accessed and cached, improving their cache and bus utilization. Second, Impulse supports prefetching at the memory controller, which can hide much of the latency of DRAM accesses. In this paper we describe the design of the Impulse architecture, and show how an Impulse memory system can be used to improve the performance of memory-bound programs. For the NAS conjugate gradient benchmark, Impulse improves performance by 67%. Because it requires no modification to processor, cache, or bus designs, Impulse can be adopted in conventional systems. In addition to scientific applications, we expect that Impulse will benefit regularly strided memory-bound applications of commercial importance, such as database and multimedia programs. John B. Carter, Wilson C. Hsieh, Leigh Stoller, Mark R. Swanson, Lixin Zhang 0002, Erik Brunvand, Al Davis, Chen-Chi Kuo, Ravindra Kuramkote, Michael A. Parker, Lambert Schaelicke, Terry Tateyama |
HPCA | 7 |
| 1998 | Improving I/O Performance with a Conditional Store BufferabstractMicroprocessor I/O performance is becoming increasingly critical in order to support efficient communication interfaces as modern microprocessors continue to be used in a variety of multiprocessor configurations. Numerous performance enhancements have been made to improve processor performance by improving the latency and bandwidth to main memory or creating efficient mechanisms to hide main memory latency. These include speculative out of order instruction execution, lock-up free caches, and improved memory bus designs. Sadly these improvements are not directly applicable to improved I/O system performance and may even complicate high performance I/O system design. This paper introduces and analyzes the design of a simple mechanism called the conditional store buffer. The conditional score buffer improves I/O write performance by making better use of the system bus to increase effective I/O bandwidth, while greatly reducing synchronization overhead. The cost is a minor increase in hardware complexity. Lambert Schaelicke, Al Davis |
MICRO | 2 |
| 1996 | Components of Congestion ControlabstractThis paper presents three original and complementary ideas on flow control mechanisms for packet switched interconnects, backpressure flow control, alpha message scheduling and balanced injection.These three components were adopted for the I'ed-.Y fabric.Each of the three addresses performance problems caused by a particular characteristic of realistic network workloads. Ludmila Cherkasova, Al Davis, Robin Hodgson, Vadim E. Kotov, Ian N. Robinson, Tomas Rokicki |
SPAA | 2 |
| 1993 | The Post Office experience: designing a large asynchronous chip
Bill Coates 0001, Al Davis, Ken Stevens |
Integr. | 2 |
| 1986 | The Post Office-Communication Support for Distributed Ensemble Architectures
Kenneth S. Stevens, Shane V. Robison, Al Davis |
ICDCS | 3 |
| 1985 | The Architecture of the FAIM-1 Symbolic Multiprocessing System
Al Davis, Shane V. Robison |
IJCAI | 1 |