Karthik Ganesan 0002

dblp:83/6972-2 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0002-2541-1549ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Embedded and real-time systems · 24% Storage systems · 17% Cloud and datacenter computing · 17%
Computer networks
1 paper
Internet of things and sensor networks · 100%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 9 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › flash and SSD
flash memory
0.812024
FlipBit: Approximate Flash Memory for IoT Devices · HPCA 2024
Embedded and real-time systems
intermittent computing
0.722019
The What's Next Intermittent Computing Architecture · HPCA 2019
The EH Model: Early Design Space Exploration of Intermittent Processor Architectures · MICRO 2018
Cloud and datacenter computing
datacenter RPC
0.612022
ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure Calls · MICRO 2022
Distributed systems
remote procedure call
0.612022
ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure Calls · MICRO 2022
Embedded and real-time systems
energy harvesting devices
0.412019
The What's Next Intermittent Computing Architecture · HPCA 2019
Emerging computing paradigms
approximate computing
0.322024
FlipBit: Approximate Flash Memory for IoT Devices · HPCA 2024
The What's Next Intermittent Computing Architecture · HPCA 2019
Energy-efficient computing
energy harvesting
0.312018
The EH Model: Early Design Space Exploration of Intermittent Processor Architectures · MICRO 2018
Internet of things and sensor networks › resource-constrained devices
energy-constrained iot device
0.212024
FlipBit: Approximate Flash Memory for IoT Devices · HPCA 2024
Memory systems
non-volatile memory
0.112018
The EH Model: Early Design Space Exploration of Intermittent Processor Architectures · MICRO 2018

Methods — techniques the papers use, named apart from their topics

approximate write · 1.50→1 transition avoidance · 1.5software-hardware co-design · 1.1proactive scheduling · 1.1low-overhead messaging · 1.1hardware migration primitives · 1.1subword vectorization · 0.4subword pipelining · 0.4skim points · 0.4EH model · 0.3
YearPublicationVenuePosition
2026 HeadSkip: Characterizing and Accelerating 'Small' Language Models on Edge CPUs
abstract
Recent years have seen a surge in the popularity of large language models. Such models are typically run in datacenters using high-end GPUs. However, due to latency and privacy concerns, companies have proposed ‘small’ language models (SLMs), designed to run locally on edge devices. We characterize two notable SLMs (i.e., Gemma 3 and Qwen3) on an edge CPU and find that the final layer accounts for up to $31.5 \%$ of the total SLM inference runtime. This final layer maps the output of the last decoder block to a specific token in the model’s vocabulary. Therefore, this layer’s runtime scales with the vocabulary size, which can be very large even for SLMs (e.g., 262k tokens for Gemma 3). Building on this observation, we present HeadSkip, two complimentary techniques to speed-up the final layer. First, we execute the final layer in chunks in order from most frequently to least frequently seen tokens. As we run, we check the logit output for tokens in a chunk and if any exceed a pre-set threshold, we output the token with the largest logit and skip executing the rest of the layer, obtaining a significant speedup. We perform this re-ordering statically based on the frequency of words in English. However, this does not capture words which are generally rare but may appear frequently in some contexts. For such words, a static order may still execute most of the layer until we check this ‘infrequent’ token, diminishing our speedup gains. To avoid this, HeadSkip also caches tokens which appear in the current context. This helps us achieve even greater speed-ups (up to $45 \%$) on modern SLMs, when evaluated on two edge CPUs.
Ian Sartor, Karthik Ganesan 0002, Andreas Moshovos
ISPASS2
2024 FlipBit: Approximate Flash Memory for IoT Devices
abstract
IoT devices commonly use flash memory for both data and code storage. Flash memory consumes a significant portion of the overall energy of such devices. This is problematic because IoT devices are energy constrained due to their reliance on batteries or energy harvesting. To save energy, we leverage a unique property of flash memory; write operations take unequal amounts of energy depending on if we are flipping a 1 → 0 versus a 0 → 1. We exploit this asymmetry to reduce energy consumption with FLIPBIT, a hardware-software approximation approach that limits costly 0→1 transitions in flash. Instead of performing an exact write, we write an approximated value that avoids any costly 0→1 bit flips. Using FLIPBIT, we reduce the mean energy used by flash by 68% on video streaming applications while maintaining 42 dB PSNR. On machine learning models, we reduce energy by an average of 39% and up to 71% with only a 1% accuracy loss. Additionally, by reducing the number of program-erase cycles, we increase the flash lifetime by 68%.
Alexander Buck, Karthik Ganesan 0002, Natalie D. Enright Jerger
HPCA2
2022 ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure Calls
abstract
Online services in modern datacenters use Remote Procedure Calls (RPCs) to communicate between different software layers. Despite RPCs using just a few small functions, inefficient RPC handling can cause delays to propagate across the system and degrade end-to-end performance. Prior work has reduced RPC processing time to less than 1 $\mu$ s, which now shifts the bottleneck to the scheduling of RPCs. Existing RPC schedulers suffer from either high overheads, inability to effectively utilize high core-count CPUs or do not adaptively fit different traffic patterns. To address these shortcomings, we present ALTOCUMULUS,1a scalable, software-hardware codesign to schedule RPCs at nanosecond scales. ALTOCUMULUS provides a proactive scheduling scheme and low-overhead messaging mechanism on top of a decentralized user runtime. ALTOCUMULUS also offers direct access from the user space to a set of simple hardware primitives to quickly migrate long-latency RPCs. We evaluate ALTOCUMULUS with synthetic workloads and an end-to-end in-memory key-value store application under real-world traffic patterns. ALTOCUMULUS improves throughput by 1.3-24.6$\times$ under a 99thpercentile latencythpercentile latency $\lt 8.5\mu \mathrm{s}$.1Automatic Concurrent Migration Load-balancing Strategy (AutoCuMuLuS), homophonic with “altocumulus” as a type of clouds in meteorology, fragmented to separate patches or nodes.
Jiechen Zhao 0002, Iris Uwizeyimana, Karthik Ganesan 0002, Mark C. Jeffrey, Natalie D. Enright Jerger
MICRO3
2019 The What's Next Intermittent Computing Architecture
abstract
Energy-harvesting devices operate under extremely tight energy constraints. Ensuring forward progress under frequent power outages is paramount. Applications running on these devices are typically amenable to approximation, offering new opportunities to provide better forward progress between power outages. We propose What's Next (WN), a set of anytime approximation techniques for energy harvesting: subword pipelining, subword vectorization and skim points. Skim points fundamentally decouple the checkpoint location from the recovery location upon a power outage. Ultimately, WN transforms processing on energy-harvesting devices from all-or-nothing to as-is computing. We enable an approximate (yet acceptable) result sooner and proceed to the next task when power is restored rather than resume processing from a checkpoint to yield the perfect output. WN yields speedups of 2.26x and 3.02x on non-volatile and checkpoint-based volatile processors, while still producing high-quality outputs.
Karthik Ganesan 0002, Joshua San Miguel, Natalie D. Enright Jerger
HPCA1
2018 The EH Model: Early Design Space Exploration of Intermittent Processor Architectures
abstract
Energy-harvesting devices—which operate solely on energy collected from their environment—have brought forth a new paradigm of intermittent computing. These devices succumb to frequent power outages that would cause conventional systems to be stuck in a perpetual loop of restarting computation and never making progress. Ensuring forward progress in an intermittent execution model requires saving state in nonvolatile memory (backup) and potentially re-executing from the last saved state upon a power loss (restore). The interplay between spending energy on useful processing and spending energy on these necessary overheads yield unexpected trade-offs. To facilitate early design space exploration, the field of intermittent computing requires better models for 1) generalizing and reasoning about these trade-offs and 2) helping architects and programmers in making early-stage design decisions. We propose the EH model, which characterizes an intermittent system's ability to maximize how much of its available energy is spent on useful processor execution. The model parametrizes the energy costs associated with intermittent execution to allow an intuitive understanding of how forward progress can change. We use the EH model to explore how forward progress is impacted with the frequency of backups and the energy cost of backups and restores. We validate the EH model with hardware measurements on an MSP430 and characterize its parameters via simulation. We also demonstrate how architects and programmers can use the model to explore the design space of intermittent processors, derive insights, and model new optimizations that are unique to intermittent processor architectures.
Joshua San Miguel, Karthik Ganesan 0002, Mario Badr, Chunqiu Xia, Rose Li, Hsuan Hsiao, Natalie D. Enright Jerger
MICRO2
2017 Measuring the Power-Constrained Performance and Energy Gap between FPGAs and Processors (Abstract Only)
Andy Gean Ye, Karthik Ganesan 0002
FPGA2