EDBT 2026 Demo / reviewers in the wild / expert
Tanvir Ahmed Khan 0001
dblp:172/7471-1
· DBLP profile ↗
20ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0003-3741-1455ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 4 first-author · 14 since 2021Software engineering, systems software and programming languages · 8 · 3 first-author · 7 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Wax: Optimizing Data Center Applications With Stale Profile
Tawhid Bhuiyan, Sumya Hoque, Angelica Aparecida Moreira, Tanvir Ahmed Khan 0001 |
ASPLOS (2) | 4 |
| 2026 | CXLMemSim: Practical Performance Simulation and Characterization of CXL 3.0 Memory SystemsabstractCompute Express Link (CXL) 3.0 can turn stranded DRAM into a composable datacenter memory fabric, but hardware is scarce and cycle-accurate simulation is prohibitively slow for full workloads. CXLMemSim is an open-source software CXL memory simulator that keeps computation at native speed and models only CXL-affected memory behavior. Across SPEC CPU2017, Memcached, wrf, vector search, and llama.cpp, CXLMemSim stays within 11–15% of hardware and gem5 while completing a Redis-scale run in 45 minutes instead of gem5’s 73 hours. The public release at SlugLab/CXLMemSim provides PEBS/LBR/eBPF tracing, JSON topology descriptions, and reusable migration/cache policy hooks for CXL design-space exploration. Yiwei Yang 0002, Shri Vishakh Devanand, Brian Zhao, Yusheng Zheng, Pooneh Safayenikoo, Tanvir Ahmed Khan 0001, Andi Quinn |
HPDC | 6 |
| 2026 | NanoTag: Systems Support for Efficient Byte-Granular Overflow Detection on ARM MTE
Hang Ye 0010, Joseph Devietti, Suman Jana, Tanvir Ahmed Khan 0001 |
SP | 5 |
| 2025 | From Optimal to Practical: Efficient Micro-op Cache Replacement Policies for Data Center ApplicationsabstractOptimizing the CPU frontend has become crucial for modern processors with intricate instruction decoding logic, especially for efficiently running planet-scale data center applications. Micro-operation (micro-op) cache is a key unit to help improve the energy efficiency of the CPU frontend. Unfortunately, we find that data center applications suffer from frequent micro-op cache misses due to the lack of an effective micro-op cache replacement policy. Developing micro-op cache-specific replacement policies is challenging, as there currently does not exist an optimal theoretical solution akin to Belady’s algorithm for conventional caches. As a result, it is unknown by how much replacement policies can be improved and how to get there. To address these challenges, we introduce FLACK, a new near-optimal offline policy that considers the key features of the micro-op cache, such as variable and disproportional costs of micro-op cache misses and partial hits. We show that FLACK substantially outperforms Belady’s algorithm, thus establishing a new baseline for micro-op cache replacement policies. We then design FURBYS, a practical policy that mimics FLACK via profile-guided methods. FURBYS has three key components to perform cache replacement decisions: (1) it uses profiles of the whole-execution hit/miss behavior, (2) it detects locally (transiently) hot data, and (3) it selectively ignores data with profiled low hit rates. We evaluate FLACK and FURBYS using 11 data center applications and find that FLACK demonstrates an average bound of 30.21% miss reduction, achieving 4.46% greater miss reduction than Belady’s algorithm. Our practical policy, FURBYS, provides 14.34% average miss reduction compared to LRU, which is $1.84 \times$ greater than the current state-of-the-art replacement policy, contributing to 3.10% of performance-perwatt improvement for the CPU core. On average, in terms of miss reduction and IPC gain, FURBYS is equivalent to LRU policy on $1.5 \times$ micro-op cache sizes (up to $2 \times$), demonstrating the effectiveness of the proposed replacement policy. Kan Zhu, Yilong Zhao 0002, Peter Braun 0005, Tanvir Ahmed Khan 0001, Heiner Litz, Baris Kasikci, Shuwen Deng |
HPCA | 5 |
| 2024 | RPG2: Robust Profile-Guided Runtime Prefetch GenerationabstractData cache prefetching is a well-established optimization to overcome the limits of the cache hierarchy and keep the processor pipeline fed with data. In principle, accurate, well-timed prefetches can sidestep the majority of cache misses and dramatically improve performance. In practice, however, it is challenging to identify which data to prefetch and when to do so. In particular, data can be easily requested too early, causing eviction of useful data from the cache, or requested too late, failing to avoid cache misses. Competition for limited off-chip memory bandwidth must also be balanced between prefetches and a program's regular "demand" accesses. Due to these challenges, prefetching can both help and hurt performance, and the outcome can depend on program structure, decisions about what to prefetch and when to do it, and, as we demonstrate in a series of experiments, program input, processor microarchitecture, and their interaction as well. Nathan Sobotka, Soyoon Park, Saba Jamilan, Tanvir Ahmed Khan 0001, Baris Kasikci, Gilles Pokam, Heiner Litz, Joseph Devietti |
ASPLOS (2) | 5 |
| 2024 | UDP: Utility-Driven Fetch Directed Instruction PrefetchingabstractDatacenter applications exhibit large instruction footprints causing significant instruction cache misses and, as a result, frontend stalls. To address this issue, instruction prefetching mechanisms have been proposed, including state-of-the-art techniques such as fetch-directed instruction prefetching. However, our study shows that existing implementations still fall far short of an ideal system with a perfect instruction cache. In particular, up to $588.47 \%$ of potential IPC speedup of existing processors hides due to frontend stalls, and these frontend stalls are due to inaccurate and untimely instruction prefetches. We quantify the impact of these individual effects, observing that applications exhibit different characteristics that call for adaptive application-specific optimizations. Based on these insights, we propose two novel mechanisms, UDP and UFTQ, to improve the accuracy of FDIP without negatively affecting timeliness while leveraging prefetches on the wrong path. We evaluate our technique on 10 data center workloads showing a maximal IPC improvement of $16.1 \%$ and an average IPC improvement of $3.6 \%$. Our techniques only introduce moderate hardware modifications and a storage cost of 8 KB. Surim Oh, Mingsheng Xu, Tanvir Ahmed Khan 0001, Baris Kasikci, Heiner Litz |
ISCA | 3 |
| 2023 | PEDAL: A Power Efficient GCN Accelerator with Multiple DAtafLowsabstractGraphs are ubiquitous in many application domains due to their ability to describe structural relations. Graph Convolutional Networks (GCNs) have emerged in recent years and are rapidly being adopted due to their capability to perform Machine Learning (ML) tasks on graph-structured data. GCN exhibits irregular memory accesses due to the lack of locality when accessing graph-structured data. This makes it hard for general-purpose architectures like CPUs and GPUs to fully utilize their computing resources. In this paper, we propose PEDAL, a power-efficient accelerator for GCN inference supporting multiple dataflows. PEDAL chooses the best-fit dataflow and phase ordering based on input graph characteristics and GCN algorithm, achieving both efficiency and flexibility. To achieve both high power efficiency and performance, PEDAL features a light-weight processing element design. PEDAL achieves 144.5x, 9.4x, and 2.6x speedup compared to CPU, GPU, and HyGCN, respectively, and 8856x, 1606x, 8.4x, and 1.8x better power efficiency compared to CPU, GPU, HyGCN, and EnGN, respectively. Alireza Khadem, Xin He 0011, Nishil Talati, Tanvir Ahmed Khan 0001, Trevor N. Mudge |
DATE | 5 |
| 2022 | APT-GET: profile-guided timely software prefetchingabstractPrefetching which predicts future memory accesses and preloads them from main memory, is a widely-adopted technique to overcome the processor-memory performance gap. Unfortunately, hardware prefetchers implemented in today's processors cannot identify complex and irregular memory access patterns exhibited by modern data-driven applications and hence developers need to rely on software prefetching techniques. We investigate the challenges of enabling effective, automated software data prefetching. Our investigation reveals that the state-of-the-art compiler-based prefetching mechanism falls short in achieving high performance due to its static nature. Based on this insight, we design APT-GET, a novel profile-guided technique that ensures prefetch timeliness by leveraging dynamic execution time information. APT-GET leverages efficient hardware support such as Intel's Last Branch Record (LBR), for collecting application execution profiles with negligible overhead to characterize the execution time of loads. APT-GET then introduces a novel analytical model to find the optimal prefetch-distance and prefetch injection site based on the collected profile to enable timely prefetches. We study APT-GET in the context of 10 real-world applications and demonstrate that it achieves a speedup of up to 1.98× and of 1.30× on average. By ensuring prefetch timeliness, APT-GET improves the performance by 1.25× over the state-of-the-art software data prefetching mechanism. Saba Jamilan, Tanvir Ahmed Khan 0001, Grant Ayers, Baris Kasikci, Heiner Litz |
EuroSys | 2 |
| 2022 | Thermometer: profile-guided btb replacement for data center applicationsabstractModern processors employ a decoupled frontend with Fetch Directed Instruction Prefetching (FDIP) to avoid frontend stalls in data center applications. However, the large branch footprint of data center applications precipitates frequent Branch Target Buffer (BTB) misses that prohibit FDIP from eliminating more than 40% of all frontend stalls. We find that the state-of-the-art BTB optimization techniques (e.g., BTB prefetching and replacement mechanisms) cannot eliminate these misses due to their inadequate understanding of branch reuse behavior in data center applications. Shixin Song, Tanvir Ahmed Khan 0001, Sara Mahdizadeh-Shahri, Akshitha Sriraman, Niranjan Soundararajan, Sreenivas Subramoney, Daniel A. Jiménez, Heiner Litz, Baris Kasikci |
ISCA | 2 |
| 2022 | Whisper: Profile-Guided Branch Misprediction Elimination for Data Center ApplicationsabstractModern data center applications experience frequent branch mispredictions– degrading performance, increasing cost, and reducing energy efficiency in data centers. Even the state-of the-art branch predictor, TAGE-SC-L, suffers from an average branch Mispredictions Per Kilo Instructions (branch-MPKI) of 3.0 (0.5-7.2) for these applications since their large code footprints exhaust TAGE-SC-L’s intended capacity. In this work, we propose Whisper, a novel profile-guided mechanism to avoid branch mispredictions. Whisper investigates the in-production profile of data center applications to identify precise program contexts that lead to branch mispredictions. Corresponding prediction hints are then inserted into code to strategically avoid those mispredictions during program execution. Whisper presents three novel profile-guided techniques: (1) hashed history correlation which efficiently encodes hard-to-predict correlations in branch history using lightweight Boolean formulas, (2) randomized formula testing which selects a locally-optimal Boolean formula from a randomly selected subset of possible formulas to predict a branch, and (3) the extension of Read-Once Monotone Boolean Formulas with Implication and Converse Non-Implication to improve the branch history coverage of these formulas with minimal overhead. We evaluate Whisper on 12 widely-used data center applications and demonstrate that Whisper enables traditional branch predictors to achieve a speedup close to that of an ideal branch predictor. Specifically, Whisper achieves an average speedup of 2.8% (0.4%-4.6%) by reducing 16.8% (1.7%-32.4%) of branch mispredictions over TAGE-SC-L and outperforms the state-of the-art profile-guided branch prediction mechanisms by 7.9% on average. Tanvir Ahmed Khan 0001, Muhammed Ugur, Krishnendra Nathella, Dam Sunwoo, Heiner Litz, Daniel A. Jiménez, Baris Kasikci |
MICRO | 1 |
| 2022 | OCOLOS: Online COde Layout OptimizationSabstractThe processor front-end has become an increasingly important bottleneck in recent years due to growing application code footprints, particularly in data centers. First-level instruction caches and branch prediction engines have not been able to keep up with this code growth, leading to more front-end stalls and lower Instructions Per Cycle (IPC). Profile-guided optimizations performed by compilers represent a promising approach, as they rearrange code to maximize instruction cache locality and branch prediction efficiency along a relatively small number of hot code paths. However, these optimizations require continuous profiling and rebuilding of applications to ensure that the code layout matches the collected profiles. If an application’s code is frequently updated, it becomes challenging to map profiling data from a previous version onto the latest version, leading to ignored profiling data and missed optimization opportunities.In this paper, we propose OCOLOS, the first online code layout optimization system for unmodified applications written in unmanaged languages. OCOLOS allows profile-guided optimization to be performed on a running process, instead of being performed offline and requiring the application to be re-launched. By running online, profile data is always relevant to the current execution and always maps perfectly to the running code. OCOLOS demonstrates how to achieve robust online code replacement in complex multithreaded applications like MySQL and MongoDB, without requiring any application changes. Our experiments show that OCOLOS can accelerate MySQL by up to $1.41 \times $, the Verilator hardware simulator by up to $2.20 \times $, and a build of the Clang compiler by up to $1.14 \times $. Tanvir Ahmed Khan 0001, Gilles Pokam, Baris Kasikci, Heiner Litz, Joseph Devietti |
MICRO | 2 |
| 2021 | Rethinking File Mapping for Persistent Memory
Ian Neal, Gefei Zuo, Eric Shiple, Tanvir Ahmed Khan 0001, Youngjin Kwon, Simon Peter 0001, Baris Kasikci |
FAST | 4 |
| 2021 | Ripple: Profile-Guided Instruction Cache Replacement for Data Center ApplicationsabstractModern data center applications exhibit deep software stacks, resulting in large instruction footprints that frequently cause instruction cache misses degrading performance, cost, and energy efficiency. Although numerous mechanisms have been proposed to mitigate instruction cache misses, they still fall short of ideal cache behavior, and furthermore, introduce significant hardware overheads. We first investigate why existing I-cache miss mitigation mechanisms achieve sub-optimal performance for data center applications. We find that widely-studied instruction prefetchers fall short due to wasteful prefetch-induced cache line evictions that are not handled by existing replacement policies. Existing replacement policies are unable to mitigate wasteful evictions since they lack complete knowledge of a data center application’s complex program behavior.To make existing replacement policies aware of these eviction-inducing program behaviors, we propose Ripple, a novel software-only technique that profiles programs and uses program context to inform the underlying replacement policy about efficient replacement decisions. Ripple carefully identifies program con-texts that lead to I-cache misses and sparingly injects "cache line eviction" instructions in suitable program locations at link time. We evaluate Ripple using nine popular data center applications and demonstrate that Ripple enables any replacement policy to achieve speedup that is closer to that of an ideal I-cache. Specifically, Ripple achieves an average performance improvement of 1.6% (up to 2.13%) over prior work due to a mean 19% (up to 28.6%) I-cache miss reduction. Tanvir Ahmed Khan 0001, Akshitha Sriraman, Joseph Devietti, Gilles Pokam, Heiner Litz, Baris Kasikci |
ISCA | 1 |
| 2021 | Twig: Profile-Guided BTB Prefetching for Data Center ApplicationsabstractModern data center applications have deep software stacks, with instruction footprints that are orders of magnitude larger than typical instruction cache (I-cache) sizes. To efficiently prefetch instructions into the I-cache despite large application footprints, modern server-class processors implement a decoupled frontend with Fetch Directed Instruction Prefetching (FDIP). In this work, we first characterize the limitations of a decoupled frontend processor with FDIP and find that FDIP suffers from significant Branch Target Buffer (BTB) misses. We also find that existing techniques (e.g., stream prefetchers and predecoders) are unable to mitigate these misses, as they rely on an incomplete understanding of a program’s branching behavior. Tanvir Ahmed Khan 0001, Akshitha Sriraman, Niranjan Soundararajan, Rakesh Kumar 0003, Joseph Devietti, Sreenivas Subramoney, Gilles Pokam, Heiner Litz, Baris Kasikci |
MICRO | 1 |
| 2021 | PDede: Partitioned, Deduplicated, Delta Branch Target BufferabstractDue to large instruction footprints, contemporary data center applications suffer from frequent frontend stalls. Despite being a significant contributor to these stalls, the Branch Target Buffer (BTB) has received less attention compared to other frontend structures such as the instruction cache. While prior works have looked at enhancing the BTB through more efficient replacement policies and prefetching policies, a thorough analysis into optimizing the BTB’s storage efficiency is missing. In this work, we analyze BTB accesses for a large number (100+) of frontend bound applications to understand their branch target characteristics. This analysis, provides three significant observations about the nature of branch targets: (1) a significant number of branch instructions have the same branch target, (2) a significant number of branch targets share the same page address, and (3) a significant percentage of branch instructions and their targets are located on the same page. Furthermore, we observe that while applications’ address spaces are sparsely populated, they exhibit spatial locality within and across pages. We refer to these multi-page addresses as regions and we show that applications traverse a significantly smaller number of regions than pages. Based on these insights, we propose PDede, an efficient re-design of the BTB micro-architecture that improves storage efficiency by removing redundancy among branches and their targets. PDede introduces three techniques, (a) BTB Partitioning, (b) Branch Target Deduplication, and (c) Delta Branch Target Encoding to reduce BTB miss induced frontend stalls. We evaluate PDede across 100+ applications, spanning several usage scenarios, and show that it provides an average 14.4% (up to 76%) IPC speedup by reducing BTB misses by 54.7% on average (and up to 99.8%). Niranjan Soundararajan, Peter Braun 0005, Tanvir Ahmed Khan 0001, Baris Kasikci, Heiner Litz, Sreenivas Subramoney |
MICRO | 3 |
| 2021 | DMon: Efficient Detection and Correction of Data Locality Problems Using Selective Profiling
Tanvir Ahmed Khan 0001, Ian Neal, Gilles Pokam, Barzan Mozafari, Baris Kasikci |
OSDI | 1 |
| 2020 | I-SPY: Context-Driven Conditional Instruction Prefetching with CoalescingabstractModern data center applications have rapidly expanding instruction footprints that lead to frequent instruction cache misses, increasing cost and degrading data center performance and energy efficiency. Mitigating instruction cache misses is challenging since existing techniques (1) require significant hardware modifications, (2) expect impractical on-chip storage, or (3) prefetch instructions based on inaccurate understanding of program miss behavior. To overcome these limitations, we first investigate the challenges of effective instruction prefetching. We then use insights derived from our investigation to develop I-SPY, a novel profile-driven prefetching technique. I-SPY uses dynamic miss profiles to drive an offline analysis of I-cache miss behavior, which it uses to inform prefetching decisions. Two key techniques underlie I-SPY's design: (1) conditional prefetching, which only prefetches instructions if the program context is known to lead to misses, and (2) prefetch coalescing, which merges multiple prefetches of non-contiguous cache lines into a single prefetch instruction. I-SPY exposes these techniques via a family of light-weight hardware code prefetch instructions. We study I-SPY in the context of nine data center applications and show that it provides an average of 15.5% (up to 45.9%) speedup and 95.9% (and up to 98.4%) reduction in instruction cache misses, outperforming the state-of-the-art prefetching technique by 22.5%. We show that I-SPY achieves performance improvements that are on average 90.5% of the performance of an ideal cache with no misses. Tanvir Ahmed Khan 0001, Akshitha Sriraman, Joseph Devietti, Gilles Pokam, Heiner Litz, Baris Kasikci |
MICRO | 1 |
| 2019 | Huron: hybrid false sharing detection and repairabstractWriting efficient multithreaded code that can leverage the full parallelism of underlying hardware is difficult. A key impediment is insidious cache contention issues, such as false sharing. False sharing occurs when multiple threads from different cores access disjoint portions of the same cache line, causing it to go back and forth between the caches of different cores and leading to substantial slowdown. Tanvir Ahmed Khan 0001, Gilles Pokam, Barzan Mozafari, Baris Kasikci |
PLDI | 1 |
| 2019 | Enhancing throughput in multi-radio cognitive radio networks
Tanvir Ahmed Khan 0001, A. B. M. Alim Al Islam |
Wirel. Networks | 1 |
| 2015 | Towards exploiting a synergy between cognitive and multi-radio networkingabstractDynamic spectrum access through Cognitive Radio Networks (CRNs) and exploiting multiple radios on a single node are two different well accepted techniques for enhancing network performance. However, simultaneous usage of both the techniques, i.e., augmenting dynamic spectrum access with multiple radios, is yet to be investigated in the literature. Therefore, in this paper, we investigate simultaneous usage of both the techniques. In our investigation, we perform rigorous ns-2 simulation. Simulation results reveal a key finding - augmenting spectrum harvesting with multiple radios makes throughput worse, however, can improve delay. Besides, as the simulation results do not reveal micro-level aspects of the improved delay performance, we perform mathematical modeling of the delay to do so. Further, we present numerical results based on the models to demonstrate the micro-level aspects. Tanvir Ahmed Khan 0001, Chowdhury Sayeed Hyder, A. B. M. Alim Al Islam |
WiMob | 1 |