David Kong 0001

dblp:242/2819-1 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0002-6117-2925ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Lifetime-Aware Design for Item-Level Intelligence at the Extreme Edge
abstract
We present FlexiFlow, a lifetime-aware design framework for item-level intelligence (ILI) where computation is integrated directly into disposable products like food packaging and medical patches. Our framework leverages natively flexible electronics which offer significantly lower costs than silicon but are limited to kHz speeds and several thousands of gates. Our insight is that unlike traditional computing with more uniform deployment patterns, ILI applications exhibit 1000× variation in operational lifetime, fundamentally changing optimal architectural design decisions when considering trillion-item deployment scales. To enable holistic design and optimization, we model the trade-offs between embodied carbon footprint and operational carbon footprint based on application-specific lifetimes. The framework includes: (1) FlexiBench, a workload suite targeting sustainability applications from spoilage detection to health monitoring; (2) FlexiBits, area-optimized RISC-V cores with 1/4/8-bit datapaths achieving 2.65× to 3.50× better energy efficiency per workload execution; and (3) a carbon-aware model that selects optimal architectures based on deployment characteristics. We show that lifetime-aware microarchitectural design can reduce carbon footprint by 1.62×, while algorithmic decisions can reduce carbon footprint by 14.5×. We validate our approach through the first tape-out using a PDK for flexible electronics with fully open-source tools, achieving 30.9\,kHz operation. FlexiFlow enables exploration of computing at the Extreme Edge where conventional design methodologies must be reevaluated to account for new constraints and considerations. FlexiFlow is available at https://github.com/harvard-edge/FlexiFlow.
Shvetank Prakash, Andrew Cheng, Olof Kindgren, Ashiq Ahamed, Graham Knight, Jedrzej Kufel, Francisco Rodriguez, Arya Tschand, David Kong 0001, Mariam Elgamal, Jerry Huang, Emma Chen, Gage Hills, Richard Price, Emre Ozer 0001, Vijay Janapa Reddi
ASPLOS (2)9
2025 333-eDRAM - 3T Embedded DRAM Leveraging Monolithic 3D Integration of 3 Transistor Types: IGZO, Carbon Nanotube and Silicon FETs
abstract
The memory wall is a major bottleneck for continuing to improve the energy efficiency of computing systems. To overcome this challenge, various nanomaterials, devices, circuits, architectures, and three-dimensional (3D) integration techniques are under development for future memory solutions. However, major trade-offs exist when designing memories to achieve high on-chip memory capacity, high retention time, high endurance, low access times, low access energy, and low static leakage power. We present an energy- and area-efficient embedded DRAM memory architecture (quantified by EADP: the product of total energy consumption, circuit area footprint and application execution time) that leverages monolithic threedimensional (3D) integration of three types of field-effect transistors (FETs): (i) Indium Gallium Zinc Oxide (IGZO) FETs for ultra-low off-state leakage currents enabling high retention time DRAM; (ii) Carbon Nanotube FETs (CNFETs) for high on-state drive currents leading to fast access times; and (iii) Silicon CMOS for its combined energy efficiency and low off-state leakage current (for memory peripheral circuits implemented on the bottom physical circuit layer). Our resulting 333-eDRAM achieves each of the following simultaneously, which we quantify and describe how to co-optimize in this paper: high density, high retention time, high endurance, low access times, low access energy, and low static leakage power. We show full physical layout designs detailing how to implement 333-eDRAM and quantify EADP for an ARM Cortex-M0 processor + on-chip 333-eDRAM implemented at a 7 nm technology node, running applications from the Embench benchmark suite. Using cycleaccurate simulations of applications, SPICE circuit simulations, compact models calibrated to experimental data, and detailed full physical layout designs of 333-eDRAM memories, we show that on average (across 16 Embench benchmarks), ARM CortexM0 + IGZO/CNT/Si 333-eDRAM offers $1.96 \times$ better EDP and $5.15 \times$ better EADP than ARM Cortex-M0 + Silicon eDRAM.
David Kong 0001, Shvetank Prakash, Jedrzej Kufel, Georgios Kyriazidis, Yasmine Omri, David Verity, Vijay Janapa Reddi, Gage Hills
DAC1
2025 Quantifying Trade-Offs in Power, Performance, Area, and Total Carbon Footprint of Future Three-Dimensional Integrated Computing Systems
abstract
To address computing's carbon footprint challenge, designers of computing systems are beginning to consider carbon footprint as a first-class figure of merit, alongside conventional metrics such as power, performance, and area. To account for total carbon$(\text{tC})$footprint of a computing system, carbon footprint models must consider both embodied carbon$(\mathrm{C}_{\text{embodied}})$due to emissions during manufacturing, and operational carbon$(\mathbf{C}_{\text{operational}})$from day-to-day use. Models for$(\mathbf{C}_{\text{operational}})$are relatively mature due to the direct relationship between$(\mathbf{C}_{\text{operational}})$and energy consumed while computing. In contrast, models for$\mathrm{C}_{\text{embodied}}$primarily focus on today's silicon-based technologies, not capturing the wide range of beyond-Si technologies that are actively being developed for future computing systems, including emerging nanomaterials, emerging memory devices, and various three-dimensional (3D) integration techniques.$\mathbf{C}_{\text {embodied }}$models for emerging technologies are essential for accurately predicting which technology directions to pursue without exacerbating computing's carbon footprint. In this paper, we (1) develop$\mathbf{C}_{\text {embodied }}$models for$\mathbf{3D}$-integrated computing systems that leverage emerging nanotechnologies. We analyze an example fabrication process that is highly promising for energy-efficient computing:$3\mathbf{D}$integration of carbon nanotube field-effect transistors (CNFETs) and indium gallium zinc oxide (IGZO) FETs fabricated directly on top of Si CMOS at a 7 nm technology node. We show that$\mathbf{C}_{\text{embodied}}$of this process is, on average (considering various energy grids),$1.31\times$higher per wafer vs. a baseline 7 nm node Si CMOS process. (2) As a case study, we quantify tradeoffs in power, performance, area, and tC footprint for an embedded system comprising an ARM Cortex-M0 processor and embedded DRAM, implemented in each of the above processes. For a representative lifetime of the system (running applications from the Embench suite for 2 hours per day over 24 months, with a clock frequency of 500 MHz), we show that the 3D IGZO/CNFET/Si implementation is 1.02 × more carbon-efficient per good die (considering yield) vs. the baseline Si implementation, quantified by the product of tC and application execution time$(tCDP$, an effective metric of carbon efficiency). (3) Finally, we show techniques to quantify carbon efficiency benefits of future computing systems, even when there is uncertainty in carbon footprint models. Specifically, we show how to robustly compare$\text{tCDP}$for multiple computing systems, given underlying uncertainty in$\mathbf{C}_{\text{embodied}}$, computing system lifetime, carbon intensity (in equivalent grams of CO2emissions per unit energy consumption), and yield.
Danielle Grey-Stewart, David Kong 0001, Mariam Elgamal, Georgios Kyriazidis, Jalil Morris, Gage Hills
DATE2