Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

David Grant

dblp:56/1597 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
2since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 6 first-author · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 25% Embedded and real-time systems · 25% Energy-efficient computing · 25%
Computer networks
1 paper
Physical-layer communications · 100%
Software engineering, system software, and programming languages
1 paper
Debugging and program repair · 77% Software maintenance and evolution · 23%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Embedded and real-time systems
digital twin
0.812024
A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale · SC 2024
Energy-efficient computing › thermal management
liquid cooling
0.812024
A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale · SC 2024
High-performance computing
supercomputing
0.812024
A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale · SC 2024
Physical-layer communications
interference alignment
0.312017
Feasibility of Single-Beam Interference Alignment in Multi-Carrier Interference Channels · IEEE Trans. Inf. Theory 2017
Physical-layer communications › MIMO
interference channel
0.312017
Feasibility of Single-Beam Interference Alignment in Multi-Carrier Interference Channels · IEEE Trans. Inf. Theory 2017
Performance modeling and evaluation
simulation
0.212024
A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale · SC 2024
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.112011
A CAD framework for Malibu: an FPGA with time-multiplexed coarse-grained elements · FPGA 2011
Electronic design automation
logic synthesis
0.112011
A CAD framework for Malibu: an FPGA with time-multiplexed coarse-grained elements · FPGA 2011
Reconfigurable computing and FPGAs › dynamic reconfiguration
time-multiplexed FPGA
0.112011
A CAD framework for Malibu: an FPGA with time-multiplexed coarse-grained elements · FPGA 2011
Debugging and program repair
fault localization
0.112009
Debugging in the (very) large: ten years of implementation and experience · SOSP 2009
Electronic design automation
physical design
0.012011
A CAD framework for Malibu: an FPGA with time-multiplexed coarse-grained elements · FPGA 2011
Electronic design automation › physical design
placement and routing
0.012011
A CAD framework for Malibu: an FPGA with time-multiplexed coarse-grained elements · FPGA 2011
Software maintenance and evolution
bug triage
0.012009
Debugging in the (very) large: ten years of implementation and experience · SOSP 2009

Methods — techniques the papers use, named apart from their topics

thermo-fluidic modeling · 0.8power simulation · 0.8augmented reality · 0.8linear algebra · 0.6algebraic geometry · 0.6progressive data collection · 0.2error statistics · 0.2bucketing · 0.2verilog compilation · 0.1configuration bitstream generation · 0.1
YearPublicationVenuePosition
2024 A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale
abstract
We present ExaDigiT, an open-source framework for developing comprehensive digital twins of liquid-cooled supercomputers. It integrates three main modules: (1) a resource allocator and power simulator, (2) a transient thermo-fluidic cooling model, and (3) an augmented reality model of the supercomputer and central energy plant. The framework enables the study of “what-if” scenarios, system optimizations, and virtual prototyping of future systems. Using Frontier as a case study, we demonstrate the framework’s capabilities by replaying six months of system telemetry for systematic verification and validation. Such a comprehensive analysis of a liquid-cooled exascale supercomputer is the first of its kind. ExaDigiT elucidates complex transient cooling system dynamics, runs synthetic or real workloads, and predicts energy losses due to rectification and voltage conversion. Throughout our paper, we present lessons learned to benefit HPC practitioners developing similar digital twins. We envision the digital twin will be a key enabler for sustainable, energy-efficient supercomputing.
Wesley Brewer, Matthias Maiterth, Rafal P. Wojda, Sedrick Bouknight, Jesse Hines, Woong Shin, Scott Greenwood, David Grant, Wesley Williams, Feiyi Wang
SC9
2021 Cooling the Data Center: Design of a Mechanical Controls Owner Project Requirements (OPR) Template
abstract
As the power demands of supercomputers continue to grow, so do the demands of the mechanical cooling systems that support the infrastructure and building in which the supercomputers reside. Planning for the cooling systems and their mechanical controls are an intrinsic part of any new supercomputer installation. To support the design and commissioning of the mechanical control systems, the Energy Efficient High Performance Computing Working Group (EE HPC WG) Cooling Controls is developing a template for an OPR (Owner Project Requirements) document. The design of the template, while pursued by a small team, leveraged the expertise of the broad membership of the EE HPC WG through surveys and feedback sessions. As a result, the OPR template includes not only a suggested structure, and a checklist of topics that a site might consider including in the document, but also many real-world examples of how the topics were addressed in previous projects. Finally, the Mechanical Controls OPR template is being developed in parallel with other templates focused on other aspects of a supercomputer installation. These templates are intended to improve the efficiency and comprehensiveness of the programming (pre-design) phase of project execution and provide the engineering and design team with better clarity of the facility infrastructure capabilities, expandability, and performance requirements. The intended result is improved construction documents (Basis-of-Design, drawings and specifications) that support the project goals and objectives including reliability, resiliency, and energy efficiency expectations of the HPC facility.
Stefan A. Robila, David Grant, Chris DePrater, Vali Sorell, Terry L. Rodgers, David Martinez, Shlomo Novotny
CLUSTER2
2017 Feasibility of Single-Beam Interference Alignment in Multi-Carrier Interference Channels
abstract
Sun and Luo recently showed that if the vector-space single-beam interference alignment problem for a K-user, L-carrier interference channel is feasible, then K ≤ 2L - 2. We prove the converse, that if K ≤ 2L - 2, then the problem is feasible, i.e., that the requisite beamformers do exist.
David Grant, Mahesh K. Varanasi
IEEE Trans. Inf. Theory1
2013 Realizing the strategic potential of e-HRM
David Grant, Sue Newell
J. Strateg. Inf. Syst.1
2011 A CAD framework for Malibu: an FPGA with time-multiplexed coarse-grained elements
abstract
Modern FPGAs are used to implement a wide range of circuits, many of which have coarse-grained and fine-grained components. The ever-increasing size of these circuits places great demand on CAD tools to synthesize circuits faster and without loss in quality. Synthesizing coarse-grained components onto fine-grained FPGA resources is inefficient, and past attempts to optimize FPGAs for word-oriented datapaths have met with limited success. This paper presents a CAD flow to fully compile Verilog into a configuration bitstream for a new type of FPGA with time-multiplexed coarse-grained resources. We demonstrate two approaches with gains of 61x and 42x in synthesis time on average compared to QuartusII, but due to time-multiplexing and current synthesis limitations we achieve circuit speeds of 14x and 8.5x slower on average. We show the tools can also trade density for maximum clock frequency.
David Grant, Chris C. Wang, Guy Lemieux
FPGA1
2009 Rapid synthesis and simulation of computational circuits in an MPPA
abstract
A computational circuit is custom-designed hardware which promises to offer maximum speedup of computationally intensive software algorithms. However, the practical needs to manage development cost and many low-level physical design details erodes much of the potential speedup by distracting attention away from high-level architectural design. Instead, designers need an inexpensive, processor-like platform where computational circuits can be rapidly synthesized and simulated. This enables rapid architectural evolution and mitigates the risk of producing custom hardware. In this paper we present a tool flow (RVETool) for compiling computational circuits into a massively parallel processor array (MPPA). We demonstrate the CAD runtime is on average 70x faster than FPGA tools, with a circuit speed 6.4x slower than FPGA devices. Unlike the fixed logic capacity of FPGAs, RVETool can trade area for simulation performance by targeting a wide range of processor cores.
David Grant, Graeme Smecher, Guy Lemieux, Rosemary Francis
FPT1
2009 Debugging in the (very) large: ten years of implementation and experience
abstract
Windows Error Reporting (WER) is a distributed system that automates the processing of error reports coming from an installed base of a billion machines. WER has collected billions of error reports in ten years of operation. It collects error data automatically and classifies errors into buckets, which are used to prioritize developer effort and report fixes to users. WER uses a progressive approach to data collection, which minimizes overhead for most reports yet allows developers to collect detailed information when needed. WER takes advantage of its scale to use error statistics as a tool in debugging; this allows developers to isolate bugs that could not be found at smaller scale. WER has been designed for large scale: one pair of database servers can record all the errors that occur on all Windows computers worldwide.
Kirk Glerum, Kinshuman Kinshumann, Steve Greenberg, Gabriel Aul, Vince R. Orgovan, Greg Nichols, David Grant, Gretchen Loihle, Galen C. Hunt
SOSP7
2008 Perturb+mutate: Semisynthetic circuit generation for incremental placement and routing
abstract
CAD tool designers are always searching for more benchmark circuits to stress their software. In this article we present a heuristic method to generate benchmark circuits specially suited for incremental place-and-route tools. The method removes part of a real circuit and replaces it with an altered version of the same circuit to mimic an incremental design change. The alteration consists of two steps: mutate followed by perturb . The perturb step exactly preserves as many circuit characteristics as possible. While perturbing, reproduction of interconnect locality, a characteristic that is difficult to measure reliably or reproduce exactly, is controlled using a new technique, ancestor depth control (ADC). Perturbing with ADC produces circuits with postrouting properties that match the best techniques known to-date. The mutate step produces targetted mutations resulting in controlled changes to specific circuit properties (while keeping other properties constant). We demonstrate one targetted mutation heuristic, scale, to significantly change circuit size with little change to other circuit characteristics. The method is simple enough for inclusion in a CAD tool directly, and fast enough for use in on-the-fly benchmark generation.
David Grant, Guy Lemieux
ACM Trans. Reconfigurable Technol. Syst.1
2006 Semi-Synthetic Circuit Generation Using Graph Monomorphism for Testing Incremental Placement and Incremental Routing Tools
abstract
FPGA architects are always searching for more benchmark circuits to stress CAD tools and device architectures. In this paper we present a new method to generate benchmark circuits by removing part of a real circuit and replacing it with a synthetic clone. This replacement or stitching process can easily introduce combinational loops if the synthetic circuit contains an input-to-output dependence that was not in the original subcircuit it is replacing. We show that this can be expressed as the graph monomorphism problem, and that a solution to that problem gives a precise stitching assignment that is cycle-free. This technique can be used to create new benchmark circuits that are identical to the original circuit except for small, local changes. The resulting semi-synthetic benchmarks are ideal for testing incremental place and route tools.
David Grant, Scott Chin, Guy Lemieux
FPL1
2006 Perturber: semi-synthetic circuit generation using ancestor control for testing incremental place and route
abstract
FPGA architects are always searching for more benchmark circuits to stress CAD tools and device architectures. In this paper we present a new heuristic to generate benchmark circuits specifically for incremental place and route tools. The method removes part of a real circuit and replaces it with a modified version of the same circuit to mimic an incremental design change. The generation procedure exactly preserves key circuit characteristics and achieves a post-routing channel width, critical path, and wire length that closely approximates those of the original circuit. Additionally, the method is fast and thus is suitable for use in on-the-fly benchmark generation
David Grant, Guy Lemieux
FPT1
1994 A High-Speed Integrated Hamming Neural Classifier
abstract
This paper describes a fixed-weight Hamming binary neural classifier chip suitable for applications where high-speed operation is required. The circuit uses switched current-mode techniques throughout and achieves classification in one forward pass through the network. The chip was realised in 2.4 /spl mu/m n-well CMOS technology and was designed to recognise the integers 0-9. It achieved a classification rate of 10 MHz without any observable tendency to misclassification or instability.>
David Grant, Paul Houselander
ISCAS1