EDBT 2026 Demo / reviewers in the wild / expert
Jedrzej Kufel
dblp:144/4603
· DBLP profile ↗
5ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0003-3648-5898ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lifetime-Aware Design for Item-Level Intelligence at the Extreme EdgeabstractWe present FlexiFlow, a lifetime-aware design framework for item-level intelligence (ILI) where computation is integrated directly into disposable products like food packaging and medical patches. Our framework leverages natively flexible electronics which offer significantly lower costs than silicon but are limited to kHz speeds and several thousands of gates. Our insight is that unlike traditional computing with more uniform deployment patterns, ILI applications exhibit 1000× variation in operational lifetime, fundamentally changing optimal architectural design decisions when considering trillion-item deployment scales. To enable holistic design and optimization, we model the trade-offs between embodied carbon footprint and operational carbon footprint based on application-specific lifetimes. The framework includes: (1) FlexiBench, a workload suite targeting sustainability applications from spoilage detection to health monitoring; (2) FlexiBits, area-optimized RISC-V cores with 1/4/8-bit datapaths achieving 2.65× to 3.50× better energy efficiency per workload execution; and (3) a carbon-aware model that selects optimal architectures based on deployment characteristics. We show that lifetime-aware microarchitectural design can reduce carbon footprint by 1.62×, while algorithmic decisions can reduce carbon footprint by 14.5×. We validate our approach through the first tape-out using a PDK for flexible electronics with fully open-source tools, achieving 30.9\,kHz operation. FlexiFlow enables exploration of computing at the Extreme Edge where conventional design methodologies must be reevaluated to account for new constraints and considerations. FlexiFlow is available at https://github.com/harvard-edge/FlexiFlow. Shvetank Prakash, Andrew Cheng, Olof Kindgren, Ashiq Ahamed, Graham Knight, Jedrzej Kufel, Francisco Rodriguez, Arya Tschand, David Kong 0001, Mariam Elgamal, Jerry Huang, Emma Chen, Gage Hills, Richard Price, Emre Ozer 0001, Vijay Janapa Reddi |
ASPLOS (2) | 6 |
| 2025 | 333-eDRAM - 3T Embedded DRAM Leveraging Monolithic 3D Integration of 3 Transistor Types: IGZO, Carbon Nanotube and Silicon FETsabstractThe memory wall is a major bottleneck for continuing to improve the energy efficiency of computing systems. To overcome this challenge, various nanomaterials, devices, circuits, architectures, and three-dimensional (3D) integration techniques are under development for future memory solutions. However, major trade-offs exist when designing memories to achieve high on-chip memory capacity, high retention time, high endurance, low access times, low access energy, and low static leakage power. We present an energy- and area-efficient embedded DRAM memory architecture (quantified by EADP: the product of total energy consumption, circuit area footprint and application execution time) that leverages monolithic threedimensional (3D) integration of three types of field-effect transistors (FETs): (i) Indium Gallium Zinc Oxide (IGZO) FETs for ultra-low off-state leakage currents enabling high retention time DRAM; (ii) Carbon Nanotube FETs (CNFETs) for high on-state drive currents leading to fast access times; and (iii) Silicon CMOS for its combined energy efficiency and low off-state leakage current (for memory peripheral circuits implemented on the bottom physical circuit layer). Our resulting 333-eDRAM achieves each of the following simultaneously, which we quantify and describe how to co-optimize in this paper: high density, high retention time, high endurance, low access times, low access energy, and low static leakage power. We show full physical layout designs detailing how to implement 333-eDRAM and quantify EADP for an ARM Cortex-M0 processor + on-chip 333-eDRAM implemented at a 7 nm technology node, running applications from the Embench benchmark suite. Using cycleaccurate simulations of applications, SPICE circuit simulations, compact models calibrated to experimental data, and detailed full physical layout designs of 333-eDRAM memories, we show that on average (across 16 Embench benchmarks), ARM CortexM0 + IGZO/CNT/Si 333-eDRAM offers $1.96 \times$ better EDP and $5.15 \times$ better EADP than ARM Cortex-M0 + Silicon eDRAM. David Kong 0001, Shvetank Prakash, Jedrzej Kufel, Georgios Kyriazidis, Yasmine Omri, David Verity, Vijay Janapa Reddi, Gage Hills |
DAC | 3 |
| 2025 | Flexing RISC-V Instruction Subset Processors to Extreme EdgeabstractThis paper presents an automated approach for designing processors that support a subset of the RISC-V instruction set architecture (ISA) for a new class of applications at Extreme Edge.The electronics used in extreme edge applications must be area and power-efficient, but also provide additional qualities, such as low cost, conformability, comfort and sustainability.Flexible electronics, rather than silicon-based electronics, will be able to meet the above qualities.For this purpose, we propose a methodology for generating RISC-V instruction subset processors (RISSPs) tailored to these applications and implementing them as flexible integrated circuits (FlexICs).The methodology makes verification an integral part of the processor design by treating each instruction in the ISA as a discrete, fully functional, pre-verified hardware block.It automatically builds a custom processor by stitching together the instruction hardware blocks required by an application or a set of applications in a specific domain.We generate RISSPs using the proposed methodology for three extreme edge applications, and embedded applications from the Embench benchmark suite.When synthesized, RISSPs can achieve 8-to-43% reduction in area and 3-to-30% reduction in power compared to a processor supporting the full RISC-V ISA, and are also on average ~40 times more energy efficient than Serv -the world's smallest 32-bit RISC-V processor.When physically implemented as FlexICs, the three extreme edge RISSPs achieve up to 42% area and 21% power savings with respect to the full RISC-V processor. Alireza Raisiardali, Konstantinos Iordanou, Jedrzej Kufel, Kowshik Gudimetla, Kris Myny, Emre Ozer 0001 |
MICRO | 3 |
| 2016 | Sequence-Aware Watermark Design for Soft IP Embedded ProcessorsabstractThis paper describes a design approach for incorporating sequence-aware watermarks in soft intellectual property (IP) embedded processors. The influence of watermark sequence parameters on detection, area, and power overheads is examined, and consequently a method for incorporating sequence-aware watermarks in soft IP embedded processors is proposed. The intrinsic parameters of sequences, such as the activity factor and the overlapping factor, are introduced, and their impact on correlation results is demonstrated. Measurement and application-specified integrated circuits validate the design approach and demonstrate the resulting IP protection and subsequent costs for constrained embedded processors. Results presented in this paper show that the tradeoff occurs between the watermark robustness against third-party IP attacks and hardware implementation costs. The analysis of this tradeoff is provided, and an application specific watermark implementation is proposed. Jedrzej Kufel, Peter R. Wilson, Stephen Hill, Bashir M. Al-Hashimi, Paul N. Whatmough |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | Clock-modulation based watermark for protection of embedded processorsabstractThis paper presents a novel watermark generation technique for the protection of embedded processors. In previous work, a load circuit is used to generate detectable watermark patterns in the ASIC power supply. This approach leads to hardware area overheads. We propose removing the dedicated load circuit entirely, instead to compensate the reduced power consumption the watermark power pattern is emulated by reusing existing clock gated sequential logic as a zero-overhead load circuit and modulating the clock-gating enable signal with the watermark sequence. The proposed technique has been validated through experiments using two ASICs in 65nm CMOS, one with an ARM Cortex-M0 microcontroller and one with a Cortex-A5 microprocessor. Silicon measurement results verify the viability of the technique for embedded processors. Furthermore, the proposed clock modulation technique demonstrates a significant area reduction, without compromising the detection performance. In our experiments an area overhead reduction of 98% was achieved. Through reuse of existing logic and reduction of watermark hardware implementation costs, the proposed clock modulation technique offers an improved robustness against removal attacks. Jedrzej Kufel, Peter R. Wilson, Stephen Hill, Bashir M. Al-Hashimi, Paul N. Whatmough, James Myers |
DATE | 1 |