John Demme

dblp:31/9796 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
1since 2021 · last 2026
0009-0006-4329-6755ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Reconfigurable computing and FPGAs · 53% Performance modeling and evaluation · 24% Cloud and datacenter computing · 21%
Network and information security
3 papers
Hardware security and side channels · 62% Malware analysis · 23% Privacy and data protection · 15%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Software engineering, system software, and programming languages
1 paper
Program analysis · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA design flow
0.312026
Hyperscale FPGA Engineering Systems at Microsoft · FPGA 2026
Hardware security and side channels
microarchitectural side channel
0.322012
TimeWarp: Rethinking timekeeping and performance monitoring mechanisms to mitigate side-channel attacks · ISCA 2012
Side-channel vulnerability factor: A metric for measuring information leakage · ISCA 2012
Performance modeling and evaluation
workload characterization
0.322012
Approximate graph clustering for program characterization · ACM Trans. Archit. Code Optim. 2012
Rapid identification of architectural bottlenecks via precise event counting · ISCA 2011
Cloud and datacenter computing › datacenter architecture
datacenter acceleration
0.212014
A reconfigurable fabric for accelerating large-scale datacenter services · ISCA 2014
Reconfigurable computing and FPGAs › FPGA architecture
FPGA fabric
0.212014
A reconfigurable fabric for accelerating large-scale datacenter services · ISCA 2014
Malware analysis › malware detection
online malware detection
0.212013
On the feasibility of online malware detection with performance counters · ISCA 2013
Data mining
clustering
0.112012
Approximate graph clustering for program characterization · ACM Trans. Archit. Code Optim. 2012
Data mining › clustering
graph clustering
0.112012
Approximate graph clustering for program characterization · ACM Trans. Archit. Code Optim. 2012
Privacy and data protection › information leakage
information leakage metric
0.112012
Side-channel vulnerability factor: A metric for measuring information leakage · ISCA 2012
Hardware security and side channels
side-channel attack
0.112012
Side-channel vulnerability factor: A metric for measuring information leakage · ISCA 2012
Hardware security and side channels
side-channel countermeasures
0.112012
TimeWarp: Rethinking timekeeping and performance monitoring mechanisms to mitigate side-channel attacks · ISCA 2012
Program analysis
graph-based analysis
0.112012
Approximate graph clustering for program characterization · ACM Trans. Archit. Code Optim. 2012
Performance modeling and evaluation
performance monitoring
0.112012
TimeWarp: Rethinking timekeeping and performance monitoring mechanisms to mitigate side-channel attacks · ISCA 2012
Performance modeling and evaluation › performance monitoring
hardware performance counters
0.112011
Rapid identification of architectural bottlenecks via precise event counting · ISCA 2011
Performance modeling and evaluation › workload characterization
parallel program behavior
0.112011
Rapid identification of architectural bottlenecks via precise event counting · ISCA 2011
Cloud and datacenter computing
datacenter services
0.112014
A reconfigurable fabric for accelerating large-scale datacenter services · ISCA 2014
Malware analysis › malware detection evasion
antivirus evasion
0.012013
On the feasibility of online malware detection with performance counters · ISCA 2013
Memory systems
cache side channel
0.012012
Side-channel vulnerability factor: A metric for measuring information leakage · ISCA 2012
Memory systems
on-chip memory
0.012012
Side-channel vulnerability factor: A metric for measuring information leakage · ISCA 2012

Methods — techniques the papers use, named apart from their topics

continuous integration · 1.0automated regression testing · 1.0approximate graph clustering · 0.4timekeeping limitation · 0.3side-channel vulnerability factor · 0.3performance counter modification · 0.3performance counters · 0.2machine learning · 0.2virtualization · 0.1hardware performance counters · 0.1
YearPublicationVenuePosition
2026 Hyperscale FPGA Engineering Systems at Microsoft
abstract
Microsoft has deployed FPGAs at hyperscale for over a decade, powering diverse application domains and products. While the underlying EDA tool flow remains familiar (synthesis, place & route, and verification), the engineering system that supports FPGA development at Microsoft looks nothing like a traditional hardware flow. Instead, it borrows heavily from modern cloud-scale software practices: Git for version control, Azure DevOps for automated pipelines, extensive regression suites, and daily compiles, effectively adapting the software mantra of ''ship every day'' to the hardware world as ''tape-out every day.''
Rob Rydberg, Madison N. Emas, John Demme, Ana Ibarra, Kara Kagi, Brandon Klouchek, Abhijeet Lawande, Todd Massengill, David J. Powers, Andrew Putnam
FPGA3
2015 A silicon anti-virus engine
John Demme, Simha Sethumadhavan, Salvatore J. Stolfo
Hot Chips Symposium2
2015 Increasing reconfigurability with memristive interconnects
abstract
The design of on-chip interconnects is largely governed by the size and power of the devices being connected. While large components like memory controllers, video decode accelerators, and cores can afford the overhead of a large packet switching NoC router, smaller components like adders or other ALUs cannot. Instead, they are typically connected via simple wires, limiting their runtime reconfigurability. The notable exception - FPGAs - use an interconnect which allows extreme reconfigurability, but the FPGA pays for it in area, power, and latency costs. Less costly reconfigurable interconnects, therefore, could allow hardware designers to expose more reconfigurability while limiting area and power costs. This paper presents the design of a high-radix circuit switching crossbar design using memristors. This design utilizes Phase Change Memory (PCM), overcoming some of its limitations such as leakage power and low voltage operation. The very small size of memristors shrinks the area, power, and latency of crossbars by up to 16x, 4.4x, and 2.4x, respectively, leaving little interconnect overhead but wiring overhead. As a tool for designers, memristive interconnects offer significant potential to increase runtime design flexibility.
John Demme, Bipin Rajendran, Steven M. Nowick, Simha Sethumadhavan
ICCD1
2014 A reconfigurable fabric for accelerating large-scale datacenter services
abstract
Datacenter workloads demand high computational capabilities, flexibility, power efficiency, and low cost. It is challenging to improve all of these factors simultaneously. To advance datacenter capabilities beyond what commodity server designs can provide, we have designed and built a composable, reconfigurable fabric to accelerate portions of large-scale software services. Each instantiation of the fabric consists of a 6×8 2-D torus of high-end Stratix V FPGAs embedded into a half-rack of 48 machines. One FPGA is placed into each server, accessible through PCIe, and wired directly to other FPGAs with pairs of 10 Gb SAS cables. In this paper, we describe a medium-scale deployment of this fabric on a bed of 1,632 servers, and measure its efficacy in accelerating the Bing web search engine. We describe the requirements and architecture of the system, detail the critical engineering challenges and solutions needed to make the system robust in the presence of failures, and measure the performance, power, and resilience of the system when ranking candidate documents. Under high load, the largescale reconfigurable fabric improves the ranking throughput of each server by a factor of 95% for a fixed latency distribution—or, while maintaining equivalent throughput, reduces the tail latency by 29%.
Andrew Putnam, Adrian M. Caulfield, Eric S. Chung, Derek Chiou, Kypros Constantinides, John Demme, Hadi Esmaeilzadeh, Jeremy Fowers, Gopi Prashanth Gopal, Jan Gray, Michael Haselman, Scott Hauck, Stephen Heil, Amir Hormati, Joo-Young Kim 0001, Sitaram Lanka, James R. Larus, Eric Peterson, Simon Pope, Aaron Smith, Jason Thong, Phillip Yi Xiao, Doug Burger
ISCA6
2013 On the feasibility of online malware detection with performance counters
abstract
The proliferation of computers in any domain is followed by the proliferation of malware in that domain. Systems, including the latest mobile platforms, are laden with viruses, rootkits, spyware, adware and other classes of malware. Despite the existence of anti-virus software, malware threats persist and are growing as there exist a myriad of ways to subvert anti-virus (AV) software. In fact, attackers today exploit bugs in the AV software to break into systems.
John Demme, Matthew Maycock, Jared Schmitz, Adam Waksman, Simha Sethumadhavan, Salvatore J. Stolfo
ISCA1
2012 Side-channel vulnerability factor: A metric for measuring information leakage
abstract
There have been many attacks that exploit side-effects of program execution to expose secret information and many proposed countermeasures to protect against these attacks. However there is currently no systematic, holistic methodology for understanding information leakage. As a result, it is not well known how design decisions affect information leakage or the vulnerability of systems to side-channel attacks. In this paper, we propose a metric for measuring information leakage called the Side-channel Vulnerability Factor (SVF). SVF is based on our observation that all side-channel attacks ranging from physical to microarchitectural to software rely on recognizing leaked execution patterns. SVF quantifies patterns in attackers' observations and measures their correlation to the victim's actual execution patterns and in doing so captures systems' vulnerability to side-channel attacks. In a detailed case study of on-chip memory systems, SVF measurements help expose unexpected vulnerabilities in whole-system designs and shows how designers can make performance-security trade-offs. Thus, SVF provides a quantitative approach to secure computer architecture.
John Demme, Robert Martin, Adam Waksman, Simha Sethumadhavan
ISCA1
2012 TimeWarp: Rethinking timekeeping and performance monitoring mechanisms to mitigate side-channel attacks
abstract
Over the past two decades, several microarchitectural side channels have been exploited to create sophisticated security attacks. Solutions to this problem have mainly focused on fixing the source of leaks either by limiting the flow of information through the side channel by modifying hardware, or by refactoring vulnerable software to protect sensitive data from leaking. These solutions are reactive and not preventative: while the modifications may protect against a single attack, they do nothing to prevent future side channel attacks that exploit other microarchitectural side channels or exploit the same side channel in a novel way. In this paper we present a general mitigation strategy that focuses on the infrastructure used to measure side channel leaks rather than the source of leaks, and thus applies to all known and unknown microarchitectural side channel leaks. Our approach is to limit the fidelity of fine grain timekeeping and performance counters, making it difficult for an attacker to distinguish between different microarchitectural events, thus thwarting attacks. We demonstrate the strength of our proposed security modifications, and validate that our changes do not break existing software. Our proposed changes require minor - or in some cases, no - hardware modifications and do not result in any substantial performance degradation, yet offer the most comprehensive protection against microarchitectural side channels to date.
Robert Martin, John Demme, Simha Sethumadhavan
ISCA2
2012 Approximate graph clustering for program characterization
abstract
An important aspect of system optimization research is the discovery of program traits or behaviors. In this paper, we present an automated method of program characterization which is able to examine and cluster program graphs, i.e., dynamic data graphs or control flow graphs. Our novel approximate graph clustering technology allows users to find groups of program fragments which contain similar code idioms or patterns in data reuse, control flow, and context. Patterns of this nature have several potential applications including development of new static or dynamic optimizations to be implemented in software or in hardware. For the SPEC CPU 2006 suite of benchmarks, our results show that approximate graph clustering is effective at grouping behaviorally similar functions. Graph based clustering also produces clusters that are more homogeneous than previously proposed non-graph based clustering methods. Further qualitative analysis of the clustered functions shows that our approach is also able to identify some frequent unexploited program behaviors. These results suggest that our approximate graph clustering methods could be very useful for program characterization.
John Demme, Simha Sethumadhavan
ACM Trans. Archit. Code Optim.1
2011 Rapid identification of architectural bottlenecks via precise event counting
abstract
On-chip performance counters play a vital role in computer architecture research due to their ability to quickly provide insights into application behaviors that are time consuming to characterize with traditional methods. The usefulness of modern performance counters, however, is limited by inefficient techniques used today to access them. Current access techniques rely on imprecise sampling or heavyweight kernel interaction forcing users to choose between precision or speed and thus restricting the use of performance counter hardware. In this paper, we describe new methods that enable precise, lightweight interfacing to on-chip performance counters. These low-overhead techniques allow precise reading of virtualized counters in low tens of nanoseconds, which is one to two orders of magnitude faster than current access techniques. Further, these tools provide several fresh insights on the behavior of modern parallel programs such as MySQL and Firefox, which were previously obscured (or impossible to obtain) by existing methods for characterization. Based on case studies with our new access methods, we discuss seven implications for computer architects in the cloud era and three methods for enhancing hardware counters further. Taken together, these observations have the potential to open up new avenues for architecture research.
John Demme, Simha Sethumadhavan
ISCA1