VLDB 2026 Research / reviewers in the wild / expert
Shvetank Prakash
dblp:310/1595
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-0748-5157ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lifetime-Aware Design for Item-Level Intelligence at the Extreme EdgeabstractWe present FlexiFlow, a lifetime-aware design framework for item-level intelligence (ILI) where computation is integrated directly into disposable products like food packaging and medical patches. Our framework leverages natively flexible electronics which offer significantly lower costs than silicon but are limited to kHz speeds and several thousands of gates. Our insight is that unlike traditional computing with more uniform deployment patterns, ILI applications exhibit 1000× variation in operational lifetime, fundamentally changing optimal architectural design decisions when considering trillion-item deployment scales. To enable holistic design and optimization, we model the trade-offs between embodied carbon footprint and operational carbon footprint based on application-specific lifetimes. The framework includes: (1) FlexiBench, a workload suite targeting sustainability applications from spoilage detection to health monitoring; (2) FlexiBits, area-optimized RISC-V cores with 1/4/8-bit datapaths achieving 2.65× to 3.50× better energy efficiency per workload execution; and (3) a carbon-aware model that selects optimal architectures based on deployment characteristics. We show that lifetime-aware microarchitectural design can reduce carbon footprint by 1.62×, while algorithmic decisions can reduce carbon footprint by 14.5×. We validate our approach through the first tape-out using a PDK for flexible electronics with fully open-source tools, achieving 30.9\,kHz operation. FlexiFlow enables exploration of computing at the Extreme Edge where conventional design methodologies must be reevaluated to account for new constraints and considerations. FlexiFlow is available at https://github.com/harvard-edge/FlexiFlow. Shvetank Prakash, Andrew Cheng, Olof Kindgren, Ashiq Ahamed, Graham Knight, Jedrzej Kufel, Francisco Rodriguez, Arya Tschand, David Kong 0001, Mariam Elgamal, Jerry Huang, Emma Chen, Gage Hills, Richard Price, Emre Ozer 0001, Vijay Janapa Reddi |
ASPLOS (2) | 1 |
| 2025 | 333-eDRAM - 3T Embedded DRAM Leveraging Monolithic 3D Integration of 3 Transistor Types: IGZO, Carbon Nanotube and Silicon FETsabstractThe memory wall is a major bottleneck for continuing to improve the energy efficiency of computing systems. To overcome this challenge, various nanomaterials, devices, circuits, architectures, and three-dimensional (3D) integration techniques are under development for future memory solutions. However, major trade-offs exist when designing memories to achieve high on-chip memory capacity, high retention time, high endurance, low access times, low access energy, and low static leakage power. We present an energy- and area-efficient embedded DRAM memory architecture (quantified by EADP: the product of total energy consumption, circuit area footprint and application execution time) that leverages monolithic threedimensional (3D) integration of three types of field-effect transistors (FETs): (i) Indium Gallium Zinc Oxide (IGZO) FETs for ultra-low off-state leakage currents enabling high retention time DRAM; (ii) Carbon Nanotube FETs (CNFETs) for high on-state drive currents leading to fast access times; and (iii) Silicon CMOS for its combined energy efficiency and low off-state leakage current (for memory peripheral circuits implemented on the bottom physical circuit layer). Our resulting 333-eDRAM achieves each of the following simultaneously, which we quantify and describe how to co-optimize in this paper: high density, high retention time, high endurance, low access times, low access energy, and low static leakage power. We show full physical layout designs detailing how to implement 333-eDRAM and quantify EADP for an ARM Cortex-M0 processor + on-chip 333-eDRAM implemented at a 7 nm technology node, running applications from the Embench benchmark suite. Using cycleaccurate simulations of applications, SPICE circuit simulations, compact models calibrated to experimental data, and detailed full physical layout designs of 333-eDRAM memories, we show that on average (across 16 Embench benchmarks), ARM CortexM0 + IGZO/CNT/Si 333-eDRAM offers $1.96 \times$ better EDP and $5.15 \times$ better EADP than ARM Cortex-M0 + Silicon eDRAM. David Kong 0001, Shvetank Prakash, Jedrzej Kufel, Georgios Kyriazidis, Yasmine Omri, David Verity, Vijay Janapa Reddi, Gage Hills |
DAC | 2 |
| 2023 | CFU Playground: Want a faster ML processor? Do it yourself!abstractThe rise of machine learning (ML) has necessitated the development of innovative processing engines. However, devel-opment of specialized hardware accelerators can incur enormous one-time engineering expenses that should be avoided in low-cost embedded ML systems. In addition, embedded systems have tight resource constraints that prevent them from affording the “full-blown” machine learning (ML) accelerators seen in many cloud environments. In embedded situations, a custom function unit (CFU) that is more lightweight is preferable. We offer CFU Playground, an open-source toolchain for accelerating embedded machine learning (ML) on FPGAs through the use of CFUs. Shvetank Prakash, Tim Callahan, Joseph Bushagour, Colby R. Banbury, Alan V. Green, Pete Warden, Tim Ansell, Vijay Janapa Reddi |
DATE | 1 |
| 2023 | ArchGym: An Open-Source Gymnasium for Machine Learning Assisted Architecture DesignabstractMachine learning (ML) has become a prevalent approach to tame the complexity of design space exploration for domain-specific architectures. While appealing, using ML for design space exploration poses several challenges. First, it is not straightforward to identify the most suitable algorithm from an ever-increasing pool of ML methods. Second, assessing the trade-offs between performance and sample efficiency across these methods is inconclusive. Finally, the lack of a holistic framework for fair, reproducible, and objective comparison across these methods hinders the progress of adopting ML-aided architecture design space exploration and impedes creating repeatable artifacts. To mitigate these challenges, we introduce ArchGym, an open-source gymnasium and easy-to-extend framework that connects a diverse range of search algorithms to architecture simulators. To demonstrate its utility, we evaluate ArchGym across multiple vanilla and domain-specific search algorithms in the design of a custom memory controller, deep neural network accelerators, and a custom SoC for AR/VR workloads, collectively encompassing over 21K experiments. The results suggest that with an unlimited number of samples, ML algorithms are equally favorable to meet the user-defined target specification if its hyperparameters are tuned thoroughly; no one solution is necessarily better than another (e.g., reinforcement learning vs. Bayesian methods). We coin the term "hyperparameter lottery" to describe the relatively probable chance for a search algorithm to find an optimal design provided meticulously selected hyperparameters. Additionally, the ease of data collection and aggregation in ArchGym facilitates research in ML-aided architecture design space exploration. As a case study, we show this advantage by developing a proxy cost model with an RMSE of 0.61% that offers a 2,000-fold reduction in simulation time. Code and data for ArchGym is available at https://bit.ly/ArchGym. Srivatsan Krishnan, Amir Yazdanbakhsh, Shvetank Prakash, Jason Jabbour, Ikechukwu Uchendu, Susobhan Ghosh, Behzad Boroujerdian, Daniel Richins, Devashree Tripathy, Aleksandra Faust, Vijay Janapa Reddi |
ISCA | 3 |
| 2023 | CFU Playground: Full-Stack Open-Source Framework for Tiny Machine Learning (TinyML) Acceleration on FPGAsabstractNeed for the efficient processing of neural networks has given rise to the development of hardware accelerators. The increased adoption of specialized hardware has highlighted the need for more agile design flows for hardware-software co-design and domain-specific optimizations. In this paper, we present CFU Playground— a full-stack open-source framework that enables rapid and iterative design and evaluation of machine learning (ML) accelerators for embedded ML systems. Our tool provides a completely open-source end-to-end flow for hardwaresoftware co-design on FPGAs and future systems research. This full-stack framework gives the users access to explore experimental and bespoke architectures that are customized and co-optimized for embedded ML. Our rapid, deploy-profileoptimization feedback loop lets ML hardware and software developers achieve significant returns out of a relatively small investment in customization. Using CFU Playground’s design and evaluation loop, we show substantial speedups between $55 \times$ and $75 \times$. The soft CPU coupled with the accelerator opens up a new, rich design space between the two components that we explore in an automated fashion using Vizier, an open-source black-box optimization service. Shvetank Prakash, Tim Callahan, Joseph Bushagour, Colby R. Banbury, Alan V. Green, Pete Warden, Tim Ansell, Vijay Janapa Reddi |
ISPASS | 1 |