EDBT 2026 Demo / reviewers in the wild / expert
Trivikram Krishnamurthy
dblp:242/8938
· DBLP profile ↗
1ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 50% Processor architecture and microarchitecture · 25% Cloud and datacenter computing · 25% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › front-end
front-end stalls |
0.4 | 1 | 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computers · ISCA 2019 |
Memory systems › cache › CPU cache
instruction cache |
0.4 | 1 | 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computers · ISCA 2019 |
Memory systems › cache
instruction cache miss |
0.4 | 1 | 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computers · ISCA 2019 |
Cloud and datacenter computing › datacenter architecture
warehouse-scale computer |
0.4 | 1 | 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computers · ISCA 2019 |
Compilers and program optimization
code layout optimization |
0.1 | 1 | 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computers · ISCA 2019 |
Methods — techniques the papers use, named apart from their topics
hardware performance monitoring · 0.8control flow graph analysis · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | AsmDB: understanding and mitigating front-end stalls in warehouse-scale computersabstractThe large instruction working sets of private and public cloud workloads lead to frequent instruction cache misses and costs in the millions of dollars. While prior work has identified the growing importance of this problem, to date, there has been little analysis of where the misses come from, and what the opportunities are to improve them. To address this challenge, this paper makes three contributions. First, we present the design and deployment of a new, always-on, fleet-wide monitoring system, AsmDB, that tracks front-end bottlenecks. AsmDB uses hardware support to collect bursty execution traces, fleet-wide temporal and spatial sampling, and sophisticated offline post-processing to construct full-program dynamic control-flow graphs. Second, based on a longitudinal analysis of AsmDB data from real-world online services, we present two detailed insights on the sources of front-end stalls: (1) cold code that is brought in along with hot code leads to significant cache fragmentation and a corresponding large number of instruction cache misses; (2) distant branches and calls that are not amenable to traditional cache locality or next-line prefetching strategies account for a large fraction of cache misses. Third, we prototype two optimizations that target these insights. For misses caused by fragmentation, we focus on memcmp, one of the hottest functions contributing to cache misses, and show how fine-grained layout optimizations lead to significant benefits. For misses at the targets of distant jumps, we propose new hardware support for software code prefetching and prototype a new feedback-directed compiler optimization that combines static program flow analysis with dynamic miss profiles to demonstrate significant benefits for several large warehouse-scale workloads. Improving upon prior work, our proposal avoids invasive hardware modifications by prefetching via software in an efficient and scalable way. Simulation results show that such an approach can eliminate up to 96% of instruction cache misses with negligible overheads. Grant Ayers, Nayana P. Nagendra, David I. August, Hyoun Kyu Cho, Svilen Kanev, Christoforos E. Kozyrakis, Trivikram Krishnamurthy, Heiner Litz, Tipp Moseley, Parthasarathy Ranganathan |
ISCA | 7 |