VLDB 2026 Research / reviewers in the wild / expert
Karthik Ganesan 0002
dblp:83/6972-2
· DBLP profile ↗
6ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0002-2541-1549ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Embedded and real-time systems · 24% Storage systems · 17% Cloud and datacenter computing · 17% | |
| Computer networks
1 paper |
Internet of things and sensor networks · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 9 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › flash and SSD
flash memory |
0.8 | 1 | 2024 | FlipBit: Approximate Flash Memory for IoT Devices · HPCA 2024 |
Embedded and real-time systems
intermittent computing |
0.7 | 2 | 2019 | The What's Next Intermittent Computing Architecture · HPCA 2019 The EH Model: Early Design Space Exploration of Intermittent Processor Architectures · MICRO 2018 |
Cloud and datacenter computing
datacenter RPC |
0.6 | 1 | 2022 | ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure Calls · MICRO 2022 |
Distributed systems
remote procedure call |
0.6 | 1 | 2022 | ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure Calls · MICRO 2022 |
Embedded and real-time systems
energy harvesting devices |
0.4 | 1 | 2019 | The What's Next Intermittent Computing Architecture · HPCA 2019 |
Emerging computing paradigms
approximate computing |
0.3 | 2 | 2024 | FlipBit: Approximate Flash Memory for IoT Devices · HPCA 2024 The What's Next Intermittent Computing Architecture · HPCA 2019 |
Energy-efficient computing
energy harvesting |
0.3 | 1 | 2018 | The EH Model: Early Design Space Exploration of Intermittent Processor Architectures · MICRO 2018 |
Internet of things and sensor networks › resource-constrained devices
energy-constrained iot device |
0.2 | 1 | 2024 | FlipBit: Approximate Flash Memory for IoT Devices · HPCA 2024 |
Memory systems
non-volatile memory |
0.1 | 1 | 2018 | The EH Model: Early Design Space Exploration of Intermittent Processor Architectures · MICRO 2018 |
Methods — techniques the papers use, named apart from their topics
approximate write · 1.50→1 transition avoidance · 1.5software-hardware co-design · 1.1proactive scheduling · 1.1low-overhead messaging · 1.1hardware migration primitives · 1.1subword vectorization · 0.4subword pipelining · 0.4skim points · 0.4EH model · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HeadSkip: Characterizing and Accelerating 'Small' Language Models on Edge CPUsabstractRecent years have seen a surge in the popularity of large language models. Such models are typically run in datacenters using high-end GPUs. However, due to latency and privacy concerns, companies have proposed ‘small’ language models (SLMs), designed to run locally on edge devices. We characterize two notable SLMs (i.e., Gemma 3 and Qwen3) on an edge CPU and find that the final layer accounts for up to $31.5 \%$ of the total SLM inference runtime. This final layer maps the output of the last decoder block to a specific token in the model’s vocabulary. Therefore, this layer’s runtime scales with the vocabulary size, which can be very large even for SLMs (e.g., 262k tokens for Gemma 3). Building on this observation, we present HeadSkip, two complimentary techniques to speed-up the final layer. First, we execute the final layer in chunks in order from most frequently to least frequently seen tokens. As we run, we check the logit output for tokens in a chunk and if any exceed a pre-set threshold, we output the token with the largest logit and skip executing the rest of the layer, obtaining a significant speedup. We perform this re-ordering statically based on the frequency of words in English. However, this does not capture words which are generally rare but may appear frequently in some contexts. For such words, a static order may still execute most of the layer until we check this ‘infrequent’ token, diminishing our speedup gains. To avoid this, HeadSkip also caches tokens which appear in the current context. This helps us achieve even greater speed-ups (up to $45 \%$) on modern SLMs, when evaluated on two edge CPUs. Ian Sartor, Karthik Ganesan 0002, Andreas Moshovos |
ISPASS | 2 |
| 2024 | FlipBit: Approximate Flash Memory for IoT DevicesabstractIoT devices commonly use flash memory for both data and code storage. Flash memory consumes a significant portion of the overall energy of such devices. This is problematic because IoT devices are energy constrained due to their reliance on batteries or energy harvesting. To save energy, we leverage a unique property of flash memory; write operations take unequal amounts of energy depending on if we are flipping a 1 → 0 versus a 0 → 1. We exploit this asymmetry to reduce energy consumption with FLIPBIT, a hardware-software approximation approach that limits costly 0→1 transitions in flash. Instead of performing an exact write, we write an approximated value that avoids any costly 0→1 bit flips. Using FLIPBIT, we reduce the mean energy used by flash by 68% on video streaming applications while maintaining 42 dB PSNR. On machine learning models, we reduce energy by an average of 39% and up to 71% with only a 1% accuracy loss. Additionally, by reducing the number of program-erase cycles, we increase the flash lifetime by 68%. Alexander Buck, Karthik Ganesan 0002, Natalie D. Enright Jerger |
HPCA | 2 |
| 2022 | ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure CallsabstractOnline services in modern datacenters use Remote Procedure Calls (RPCs) to communicate between different software layers. Despite RPCs using just a few small functions, inefficient RPC handling can cause delays to propagate across the system and degrade end-to-end performance. Prior work has reduced RPC processing time to less than 1 $\mu$ s, which now shifts the bottleneck to the scheduling of RPCs. Existing RPC schedulers suffer from either high overheads, inability to effectively utilize high core-count CPUs or do not adaptively fit different traffic patterns. To address these shortcomings, we present ALTOCUMULUS,1a scalable, software-hardware codesign to schedule RPCs at nanosecond scales. ALTOCUMULUS provides a proactive scheduling scheme and low-overhead messaging mechanism on top of a decentralized user runtime. ALTOCUMULUS also offers direct access from the user space to a set of simple hardware primitives to quickly migrate long-latency RPCs. We evaluate ALTOCUMULUS with synthetic workloads and an end-to-end in-memory key-value store application under real-world traffic patterns. ALTOCUMULUS improves throughput by 1.3-24.6$\times$ under a 99thpercentile latencythpercentile latency $\lt 8.5\mu \mathrm{s}$.1Automatic Concurrent Migration Load-balancing Strategy (AutoCuMuLuS), homophonic with “altocumulus” as a type of clouds in meteorology, fragmented to separate patches or nodes. Jiechen Zhao 0002, Iris Uwizeyimana, Karthik Ganesan 0002, Mark C. Jeffrey, Natalie D. Enright Jerger |
MICRO | 3 |
| 2019 | The What's Next Intermittent Computing ArchitectureabstractEnergy-harvesting devices operate under extremely tight energy constraints. Ensuring forward progress under frequent power outages is paramount. Applications running on these devices are typically amenable to approximation, offering new opportunities to provide better forward progress between power outages. We propose What's Next (WN), a set of anytime approximation techniques for energy harvesting: subword pipelining, subword vectorization and skim points. Skim points fundamentally decouple the checkpoint location from the recovery location upon a power outage. Ultimately, WN transforms processing on energy-harvesting devices from all-or-nothing to as-is computing. We enable an approximate (yet acceptable) result sooner and proceed to the next task when power is restored rather than resume processing from a checkpoint to yield the perfect output. WN yields speedups of 2.26x and 3.02x on non-volatile and checkpoint-based volatile processors, while still producing high-quality outputs. Karthik Ganesan 0002, Joshua San Miguel, Natalie D. Enright Jerger |
HPCA | 1 |
| 2018 | The EH Model: Early Design Space Exploration of Intermittent Processor ArchitecturesabstractEnergy-harvesting devices—which operate solely on energy collected from their environment—have brought forth a new paradigm of intermittent computing. These devices succumb to frequent power outages that would cause conventional systems to be stuck in a perpetual loop of restarting computation and never making progress. Ensuring forward progress in an intermittent execution model requires saving state in nonvolatile memory (backup) and potentially re-executing from the last saved state upon a power loss (restore). The interplay between spending energy on useful processing and spending energy on these necessary overheads yield unexpected trade-offs. To facilitate early design space exploration, the field of intermittent computing requires better models for 1) generalizing and reasoning about these trade-offs and 2) helping architects and programmers in making early-stage design decisions. We propose the EH model, which characterizes an intermittent system's ability to maximize how much of its available energy is spent on useful processor execution. The model parametrizes the energy costs associated with intermittent execution to allow an intuitive understanding of how forward progress can change. We use the EH model to explore how forward progress is impacted with the frequency of backups and the energy cost of backups and restores. We validate the EH model with hardware measurements on an MSP430 and characterize its parameters via simulation. We also demonstrate how architects and programmers can use the model to explore the design space of intermittent processors, derive insights, and model new optimizations that are unique to intermittent processor architectures. Joshua San Miguel, Karthik Ganesan 0002, Mario Badr, Chunqiu Xia, Rose Li, Hsuan Hsiao, Natalie D. Enright Jerger |
MICRO | 2 |
| 2017 | Measuring the Power-Constrained Performance and Energy Gap between FPGAs and Processors (Abstract Only)
Andy Gean Ye, Karthik Ganesan 0002 |
FPGA | 2 |