Todd Massengill

dblp:175/6098 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0008-6924-4763ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Reconfigurable computing and FPGAs · 51% Cloud and datacenter computing · 25% Hardware accelerators and domain-specific architectures · 23%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 6 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
FPGA-based neural network accelerator
0.312018
A Configurable Cloud-Scale DNN Processor for Real-Time AI · ISCA 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit
0.312018
A Configurable Cloud-Scale DNN Processor for Real-Time AI · ISCA 2018
Reconfigurable computing and FPGAs
FPGA design flow
0.312026
Hyperscale FPGA Engineering Systems at Microsoft · FPGA 2026
Cloud and datacenter computing › datacenter architecture
datacenter acceleration
0.212016
Agile Co-Design for a Reconfigurable Datacenter · FPGA 2016
Reconfigurable computing and FPGAs
FPGA accelerator
0.212016
A cloud-scale acceleration architecture · MICRO 2016
Reconfigurable computing and FPGAs
datacenter FPGA deployment
0.112016
Agile Co-Design for a Reconfigurable Datacenter · FPGA 2016

Methods — techniques the papers use, named apart from their topics

continuous integration · 1.0automated regression testing · 1.0synthesis specialization · 0.7distributed microarchitecture · 0.7SIMD ISA · 0.7measurement · 0.5deployment · 0.5hardware-software co-design · 0.2agile development · 0.2
YearPublicationVenuePosition
2026 Hyperscale FPGA Engineering Systems at Microsoft
abstract
Microsoft has deployed FPGAs at hyperscale for over a decade, powering diverse application domains and products. While the underlying EDA tool flow remains familiar (synthesis, place & route, and verification), the engineering system that supports FPGA development at Microsoft looks nothing like a traditional hardware flow. Instead, it borrows heavily from modern cloud-scale software practices: Git for version control, Azure DevOps for automated pipelines, extensive regression suites, and daily compiles, effectively adapting the software mantra of ''ship every day'' to the hardware world as ''tape-out every day.''
Rob Rydberg, Madison N. Emas, John Demme, Ana Ibarra, Kara Kagi, Brandon Klouchek, Abhijeet Lawande, Todd Massengill, David J. Powers, Andrew Putnam
FPGA8
2018 A Configurable Cloud-Scale DNN Processor for Real-Time AI
abstract
Interactive AI-powered services require low-latency evaluation of deep neural network (DNN) models-aka ""real-time AI"". The growing demand for computationally expensive, state-of-the-art DNNs, coupled with diminishing performance gains of general-purpose architectures, has fueled an explosion of specialized Neural Processing Units (NPUs). NPUs for interactive services should satisfy two requirements: (1) execution of DNN models with low latency, high throughput, and high efficiency, and (2) flexibility to accommodate evolving state-of-the-art models (e.g., RNNs, CNNs, MLPs) without costly silicon updates. This paper describes the NPU architecture for Project Brainwave, a production-scale system for real-time AI. The Brainwave NPU achieves more than an order of magnitude improvement in latency and throughput over state-of-the-art GPUs on large RNNs at a batch size of 1. The NPU attains this performance using a single-threaded SIMD ISA paired with a distributed microarchitecture capable of dispatching over 7M operations from a single instruction. The spatially distributed microarchitecture, scaled up to 96,000 multiply-accumulate units, is supported by hierarchical instruction decoders and schedulers coupled with thousands of independently addressable high-bandwidth on-chip memories, and can transparently exploit many levels of fine-grain SIMD parallelism. When targeting an FPGA, microarchitectural parameters such as native datapaths and numerical precision can be "synthesis specialized" to models at compile time, enabling atypically high FPGA performance competitive with hardened NPUs. When running on an Intel Stratix 10 280 FPGA, the Brainwave NPU achieves performance ranging from ten to over thirty-five teraflops, with no batching, on large, memory-intensive RNNs.
Jeremy Fowers, Kalin Ovtcharov, Michael Papamichael, Todd Massengill, Daniel Lo, Shlomi Alkalay, Michael Haselman, Logan Adams, Mahdi Ghandi, Stephen Heil, Prerak Patel, Adam Sapek, Gabriel Weisz, Lisa Woods, Sitaram Lanka, Steven K. Reinhardt, Adrian M. Caulfield, Eric S. Chung, Doug Burger
ISCA4
2016 Agile Co-Design for a Reconfigurable Datacenter
abstract
In 2015, a team of software and hardware developers at Microsoft shipped the world?s first commercial search engine accelerated using FPGAs in the datacenter. During the sprint to production, new algorithms in the Bing ranking service were ported into FPGAs and deployed to a production bed within several weeks of conception, leading to significant gains in latency and throughput. The fast turnaround time of new features demanded by an agile software culture would not have been possible without a disciplined and effective approach to co-design in the datacenter. This talk will describe some of the learnings and best practices developed from this unique experience.
Shlomi Alkalay, Hari Angepat, Adrian M. Caulfield, Eric S. Chung, Oren Firestein, Michael Haselman, Stephen Heil, Kyle Holohan, Matt Humphrey, Tamás Juhász, Puneet Kaur, Sitaram Lanka, Daniel Lo, Todd Massengill, Kalin Ovtcharov, Michael Papamichael, Andrew Putnam, Raja Seera, Rimon Tadros, Jason Thong, Lisa Woods, Derek Chiou, Doug Burger
FPGA14
2016 A cloud-scale acceleration architecture
abstract
Hyperscale datacenter providers have struggled to balance the growing need for specialized hardware (efficiency) with the economic benefits of homogeneity (manageability). In this paper we propose a new cloud architecture that uses reconfigurable logic to accelerate both network plane functions and applications. This Configurable Cloud architecture places a layer of reconfigurable logic (FPGAs) between the network switches and the servers, enabling network flows to be programmably transformed at line rate, enabling acceleration of local applications running on the server, and enabling the FPGAs to communicate directly, at datacenter scale, to harvest remote FPGAs unused by their local servers. We deployed this design over a production server bed, and show how it can be used for both service acceleration (Web search ranking) and network acceleration (encryption of data in transit at high-speeds). This architecture is much more scalable than prior work which used secondary rack-scale networks for inter-FPGA communication. By coupling to the network plane, direct FPGA-to-FPGA messages can be achieved at comparable latency to previous work, without the secondary network. Additionally, the scale of direct inter-FPGA messaging is much larger. The average round-trip latencies observed in our measurements among 24, 1000, and 250,000 machines are under 3, 9, and 20 microseconds, respectively. The Configurable Cloud architecture has been deployed at hyperscale in Microsoft's production datacenters worldwide.
Adrian M. Caulfield, Eric S. Chung, Andrew Putnam, Hari Angepat, Jeremy Fowers, Michael Haselman, Stephen Heil, Matt Humphrey, Puneet Kaur, Joo-Young Kim 0001, Daniel Lo, Todd Massengill, Kalin Ovtcharov, Michael Papamichael, Lisa Woods, Sitaram Lanka, Derek Chiou, Doug Burger
MICRO12